AI at Prova

Put the power to find out in more hands.

Our purpose is to democratize the ability to find out what works in social policy.

The next advance in how we help people could begin in a neighborhood program, a school, or a public agency. More of those ideas deserve the chance to be tested and developed.

AI opens a remarkable possibility: putting serious research within reach of many more people working for social progress. We’re building Prova to make that possibility real.

01 · The opening

More complex investigations are coming within reach.

Serious inquiry brings many forms of work together: reasoning through a question, pursuing evidence, designing a study, building its instruments, and checking what the findings establish.

AI’s growing ability to contribute across that work changes what people and institutions can undertake. Prova’s founding thesis is that expanding this capacity can change the organization of social research.

Capabilities that extend the work

Evidence available by August 2026
01 · Reasoning

Work through a difficult question.

Weigh arguments, examine assumptions, and connect the implications.

Conceptual Reasoning Index · points

GPT-4o27.1

2024 model

Claude Opus 573.6

August 2026 result

Opus 5: 95% confidence interval ±2.1 points.

Redwood Research / Anthropic
02 · Evidence gathering

Pursue evidence through tools.

Search across sources, follow leads, and locate hard-to-find information.

BrowseComp · questions answered correctly

Deep Research51.5%

April 2025 result

GPT-5.6 Sol90.4%

July 2026 result

1,266 questions. Reported agent systems use different configurations.

OpenAI evaluations
03 · Analysis

Carry out a demanding analysis.

Work with spreadsheets, documents, and code to resolve a question.

AA-AnalystAgent · correct on all five attempts

60%48 of 80 questions

Gemini 3.7 Flash · high reasoning

80 questions across 14 domains. August 13, 2026 result.

Artificial Analysis

Bring the capabilities into one investigation.

  1. Frame the question
  2. Pursue evidence
  3. Design & build
  4. Analyze & check
Three separate benchmarks, each measuring a particular capability on its own scale. The investigation below them illustrates Prova’s intended application; research outcomes require their own evaluation.
Read the measures and sources
  1. Reasoning. The Conceptual Reasoning Index combines argument assessment, consistency of beliefs, and decision-theoretic reasoning. Both models are scored on the same index; GPT-4o’s score is retrospective. The August 12 release reports results through August 10. Exact scores are available in the results table. Index points are not a ratio scale of intelligence.
  2. Evidence gathering. BrowseComp tests difficult fact-finding with short, verifiable answers. Deep Research’s April 2025 result and Sol’s July 2026 result use different trained systems and configurations. The latter is the single-agent result, available by August. This is not a controlled estimate of model-only improvement.
  3. Analysis. AA-AnalystAgent uses tools for code execution, web fetching, and image viewing. A question counts as solved only when all five independent attempts are correct (pass⁵). The squares show the aggregate share, not individual question identities. See the benchmark methodology and August model evaluation.
Supporting evidence · METRHow complex can the work become?

From minutes to multi-hour problems.

METR measures the difficulty of a task by how long a human expert would need to complete it. Its horizon estimates show the task duration at which an AI agent is predicted to succeed at a given rate.

Predicted task success
Point estimate 95% confidence interval
  1. GPT-4March 2023
    4 min95% CI: 1.9 min–8 min
  2. GPT-4oMay 2024
    7 min95% CI: 4 min–13 min
  3. Claude 3.7 SonnetFebruary 2025
    1.0 hr95% CI: 33 min–1.7 hr
  4. o3April 2025
    2.0 hr95% CI: 1.2 hr–3.2 hr
  5. GPT-5August 2025
    3.4 hr95% CI: 1.9 hr–6.8 hr
  6. Claude Opus 4.6February 2026
    12.0 hr95% CI: 5.3 hr–60.6 hr
Human-expert task duration · logarithmic scale · selected model generations from Time Horizon 1.1. Marks show estimates and 95% confidence intervals.

At 50% predicted success, Claude Opus 4.6’s task horizon is about 12.0 hr. The confidence interval is 5.3 hr–60.6 hr.

The tasks primarily cover software engineering, machine learning, and cybersecurity. They are self-contained problems with clear success criteria; the results do not measure complete social-research projects.

Source series updated May 8, 2026. This is evidence available by August, with no August measurement added or extrapolated. METR’s methods · Published data

Build the capacity to use them together.

Prova brings these developing capabilities together with organized research knowledge and the digital material institutions already produce. The aim is to investigate more demanding questions, build what a study needs, and support learning as a program develops. Lower costs can widen access to that capacity.

02 · What we’re building

AI created the intellectual foundation.

Prova is an AI-native company. AI created the organized body of research knowledge and relationships at the heart of our proprietary systems, connecting methods with the people, resources, and conditions a study depends on.

We use AI with that structure throughout research, design, analysis, and implementation. The question, evidence, local conditions, and professional expertise give the work its direction.

The Prova research foundationTwo sources of developing capability
Frontier models

Reasoning, research & software.

Models contribute language, reasoning, coding, and tool use.

AI created this foundation

Prova’s organized research knowledge.

Methods & measures
People & conditions
Relationships & standards
Applied through our systems

Research put to work.

  • Studies
  • Findings
  • Instruments
  • Software
Professional judgment throughout.The question · evidence · local conditions
Explore how the capability can develop
The model and the research foundation can develop separately. Their combination creates room for more demanding work.
What is a compiled prior?

Our starting research knowledge, organized for use. AI created this body of methods, assumptions, and relationships. It remains open to examination, correction, and revision. The bands above illustrate its scope.

A company organized around that capability.

AI helped make Prova itself possible, expanding what its founder could undertake in developing the systems and building an organization around them. It continues to contribute to strategy, internal operations, software, and product development.

Research and product development can both strengthen the capability. A study can reveal a missing method; a tool can make that method practical to use. Prova is being built to connect these forms of work.

Change the question. See the research change.

Are participants in work six months later?

  1. Design

    Plan a follow-up

    Measure employment six months after the program.

  2. Build

    Create the instruments

    Develop a survey and a collection tool.

  3. Carry out

    Reach participants

    Run the follow-up and track missing responses.

  4. Learn

    Understand the outcome

    Describe employment among respondents and examine missing follow-up.

Illustrative study choices. The question changes the design, the work involved, and what the findings can establish.

03 · An invitation to researchers

Making research knowledge a testable source of AI capability.

We are exploring how much of the capability required for serious social research can be built into an explicit, revisable body of knowledge and relationships, then carried across generations of AI models.

Maintaining this knowledge separately from the model makes the structure itself available for examination and improvement. For AI and ML researchers, that creates a concrete experimental agenda.

Research agenda · Proposed evaluation

Test what the structure contributes.

Same held-out cases. Same model.Match tool access and context budgets; record actual resource use.
A · Task baseline

A strong task-prompted model

Case materials and a carefully developed task prompt.

B · Knowledge as text

The same knowledge, in prose

Matched facts and relational statements, presented as text.

C · Structured knowledge

Explicit relationships

The matched knowledge, organized into an explicit structure.

Blinded expert assessmentPredeclared criteria · Cases held out from development
  • Design quality
  • Source fidelity
  • Expert corrections
  • Total effort

Does the organization of knowledge add value beyond access to it?

Match factual and relational content and coverage in B and C. A tests what adding research knowledge contributes.

Proposed experiments. No performance advantage or completed evaluation is asserted.

Another question concerns how the work preserves agreed analytical commitments when results or incentives make them inconvenient. We invite collaborators to help establish which improvements are real, which transfer, and what deserves to be built next.

Related research on compound AI systems and DSPy provides a broader setting for this work. The comparisons above are Prova’s proposed research agenda.

04 · The responsibility that comes with it

The same capabilities raise the stakes.

Greater reach, lower costs, and continuing learning bring real possibilities. Each also asks something of how the work is designed and judged. A mistake in a reusable research structure can travel into later studies and tools.

Select a capability to explore the discipline it requires.

Investigate more

Extend research, design, and implementation.

Pollution

Weak work at volume. Errors carried into later work.

Reach more programs

Bring serious inquiry within reach.

Dismissal

A low price mistaken for a low standard.

Keep learning

Make inquiry part of ordinary operations.

Surveillance & target chasing

People lose a say. The score replaces the goal.

Examine what gets reused.

Check sources, methods, and tools. Challenge assumptions before they travel into another study.

The standard follows the question and its intended use.

Institutions also need arrangements in which an unwelcome finding can lead to useful action.

More evidence has limited value if a program cannot change, a funder cannot reconsider, or people cannot challenge how they are represented. The ability to learn depends on the authority, incentives, and relationships around the work.

05 · The larger ambition

Knowledge and capability should accumulate.

A study can leave a better instrument, a useful method, or a finding another team can examine. Every investigation should give the next one a stronger start.

Advances in models and improvements in maintained research knowledge offer two ways for the capability to develop. Establishing what can travel between questions and settings is part of the work ahead.

The starting pointOne investigation

↓ Can leave useful assets

Better instruments
Reviewed findings
Reusable tools

↓ To inform further work

Refine the instrument and reuse the tools to investigate what changed.
The ambition. Reuse requires review and permission; client-private information stays within its agreed boundaries.

The longer horizon reaches across institutions.

We want useful knowledge to survive individual projects, leadership changes, and funding cycles. The ambition is a growing capacity for institutions to cooperate, test ideas, learn from unsuccessful attempts, and act on what they discover.

For capital committed to social progress, the question is what becomes possible when more institutions can sustain inquiry. For Prova, the task is to build a company whose research, systems, and products help bring that future within reach.

Build with Prova

Help develop the capacity to find out.

We welcome foundations and investors interested in developing this capability, researchers who want to test it, and institutions willing to shape its real-world use. The next stage is to establish quality, transfer, total cost, and the conditions for responsible adoption.