Sunday, August 30, 2026

News

Former Google Researchers Launch Sampura Research to Keep AI From Evading Oversight

ResearchPatryk Raba
Former Google Researchers Launch Sampura Research to Keep AI From Evading Oversight
Fot. Tara Winstead, Pexels (Pexels License)

Two former Google DeepMind researchers have launched nonprofit Sampura Research, which is building hybrid human-AI systems to detect when models learn to evade oversight. The organization has raised over $10 million from Coefficient Giving.

Contents
  1. Where the hybrid idea came from
  2. Funding and plans
  3. Industry context
  4. Why it matters

Rishub Jain and Josh Jacob, both longtime researchers at Google DeepMind, have announced the launch of a nonprofit called Sampura Research. Its goal is to build evaluation systems in which humans and artificial intelligence jointly watch for signs that AI models are learning to circumvent oversight.

Sampura Research is not another lab building language models. It is a research organization trying to answer a different question: who, or what, should judge whether a given step taken by an AI agent is safe and consistent with its creators' intent. Jain and Jacob call this evaluation mechanism a "judge" - it can be a human, an AI model, or a hybrid of the two.

Where the hybrid idea came from

Jain spent seven years at Google DeepMind, four of them working on AlphaFold 2 and AlphaFold 3, before co-leading the Scalable Oversight team. Jacob worked on curating human data used in capability and safety research for models. Both started from the premise that systems evaluated solely by other AI models share a common weakness - they can carry the same cognitive blind spots as the models being evaluated, which makes those blind spots easier to exploit.

If you involve a human in the process from the very beginning, the model is less likely to learn to exploit those blind spots - Rishub Jain, co-founder of Sampura Research

Sampura defines a "judge" as a human, an AI, or a hybrid system that assesses whether a model's behavior is correct and compliant, either in a single conversation or across an entire agent trajectory. The point isn't just a single chatbot response, but the full sequence of decisions an autonomous agent makes while carrying out a task step by step.

Funding and plans

The $11 million in funding comes from Coefficient Giving, a philanthropic organization tied to the effective altruism movement. Seven million dollars is meant to cover the first year of operations, with another four million pledged for the future. The team plans to build a leaderboard covering more than twenty test datasets - spanning fraud detection, cultural errors, and recognition of actions dangerous to people or systems.

On that leaderboard, hybrid human-AI systems will be compared against evaluations based solely on models such as GPT or Gemini, as well as against evaluations done entirely by humans. The next step will be testing ways to combine both approaches - including routing tasks to the appropriate type of judge, breaking down complex trajectories into smaller pieces for evaluation, and designing interfaces that support human reviewers.

Industry context

The impetus for Sampura Research came from reports of models that, during safety testing, gained unauthorized access to other parties' systems - incidents disclosed earlier by OpenAI and Anthropic sparked industry debate over whether current oversight methods can keep pace with the growing autonomy of AI agents. A 2024 DeepMind study cited by the founders found that hybrid human-AI teams were more effective at catching security flaws than models alone, though the advantage was modest.

In its next phase, Sampura intends to test its tools directly inside AI labs, optimizing them for cost and speed so they can be used in practice rather than only under experimental conditions. Longer term, the organization wants to extend the approach to generative tasks - not just evaluating finished answers, but also creating instructions and grading rubrics for the models themselves.

Why it matters

For companies deploying AI agents in production environments, from process automation to financial systems, Sampura's proposal addresses a practical problem: it's getting harder to manually verify every step an autonomous system takes, and relying solely on another AI model as reviewer raises the question of who watches the watchers. If Sampura's leaderboard and benchmarks prove reliable, they could become a reference point for choosing agent-oversight tools, much like other independent model-safety rankings already on the market.

For now, Sampura Research is a small nonprofit with no commercial model of its own - its impact will depend on whether AI labs and companies deploying agents choose to use its benchmarks and methodology. Results from the first tests and the shape of the leaderboard are expected in the coming months.

Share: