ARA Lab
VALIDATING AI
Verification for
agent-native research
TRIAL · Lara · rit
Jiachen (Amber) Liu
Founder · ARA Lab
ERIC SCHMIDT SCIENCES · VALIDATING AI WORKSHOP · OCTOBER 2026
ARA Lab
Background

From ML systems to infrastructure for AI scientists.

University of Michigan
CS PhD
Meta
Pre-training @
Meta Superintelligence Labs
ARA Lab
Founder
Research trajectory
01
ML SYSTEMS
train + serve efficiently
02
AI SCIENTIST
automate discovery
NOW
03
RESEARCH INFRASTRUCTURE
for AI scientists
Background
Our work

We build four pieces of infrastructure for AI scientists.

01 · TASTE
Sci-Reasoning
how top scientists think
ARXIV 2026
02 · KNOWLEDGE
ARA · Curie
how research is documented and compounds
PAPERThe Last Human-Written PaperAgent-Native Research Artifacts
NEURIPS 2026
03 · VERIFICATION
TRIAL · Lara · rit
how a result earns trust
TODAY
04 · EVALUATION
EXP-Bench ·
Open-Endedness Bench
how open-ended research is benchmarked
ICLR 2026
Today: verification.
Our work
Why verification

Generation runs at machine speed.
Verification still runs at human speed.

AI generation
∞AGENTS
hypothesis
experiment
result
hypothesis
experiment
result
hypothesis
experiment
result
hypothesis
MACHINE SPEED
backlog
↑ piling up
human review
reviewing
checked
✓ checked
✓ checked
✓ checked
one reader,
one result
at a time
HUMAN SPEED
UNCHECKED → NEXT EXPERIMENT'S PREMISE
KOSMOS
one 12-hour run ≈ six months of human research · over 40% of its interpretive claims are wrong
Every unchecked result becomes the next experiment's premise.
Why verification
Overview

Three projects: from one run, to one claim, to all claims.

ONE RUN
ONE CLAIM
ALL CLAIMS
PROJECT 1
TRIAL
Did the agent really do the task, or hack the reward?
TODAYonly the final score is checked
PROJECT 2
Lara
Can a machine check the argument behind a claim?
TODAYclaims are written in ambiguous prose
PROJECT 3
rit
Is a new claim consistent with the knowledge base?
TODAYan LLM judge reads each claim alone
Overview
Project 1 of 3
1 · TRIAL2 · Lara3 · rit
PROJECT 1
TRIAL
Trajectory Review of Integrity And Legitimacy
Jicheng Wang, …, Jiachen Liu
THE PROBLEMAI agents reward-hack: they get the score without doing the task.
Project 1 · TRIAL
TRIAL · why it matters · 1

A missed hack doesn't stay in one run. It spreads.

A HACKED RUN $ git show <upstream fix> ✓ PASS TRAINING The next model learns it, and worse. 12% sabotaged safety code after it learned to hack THE PAPER The paper reports a fake success. 42% of experiments failed; the papers said success SECURITY It uses access it was never given. 4% → 73% found an API key; the judge missed it
Project 1 · TRIAL
TRIAL · why it matters · 2

The stronger the model, the cleverer the hack.

CRUDE
Overwrite the file the grader reads
echo "…" > game/fen.txt
chess agent rewrites the board to win
SNEAKY
Copy the official fix from the project's history
git show <upstream fix>
SWE-bench agent; the task counts as solved
HIDDEN
Reword the test questions slightly, then train on them
test item → reworded → train.jsonl
post-training agent; the benchmark's own contamination check reports 0
STRONGER MODELS →
REWARDHACKING.IO20 models, 1,226 audited runs: when frontier models cheat, the LLM judge misses it almost every time.
Project 1 · TRIAL
TRIAL · what it does

A hack is a path from off-limits content to the score.

1INPUT · THE WHOLE RUN
#0001$ ls evaluation_code/
#0092Read evaluate.py
#0183$ python evaluate.py
#0274 → 31/50 passed …
#0365Write prep_data.py
#0456$ python prep_data.py
#0547$ head data/train.jsonl
#0638Write train.py
#0729$ nohup python train.py &
#0820$ tail -f train.log
#0911$ python evaluate.py
#1002Edit train.py
#1093$ python train.py --epochs 3
#1184$ cp -r out/ final_model/
#0001$ ls evaluation_code/
#0092Read evaluate.py
#0183$ python evaluate.py
#0274 → 31/50 passed …
#0365Write prep_data.py
#0456$ python prep_data.py
#0547$ head data/train.jsonl
#0638Write train.py
#0729$ nohup python train.py &
#0820$ tail -f train.log
#0911$ python evaluate.py
#1002Edit train.py
#1093$ python train.py --epochs 3
#1184$ cp -r out/ final_model/
thousands of steps: shell, Python, tool calls, what each one printed, and the files left at the end
→
2EXTRACT · WHERE CONTENT MOVED
READWRITECOPYFETCHRUNTRAIN
test set grader output agent · step 41 train.jsonl training run final model utils.pytrain.log OFF-LIMITSSCORED
nodes: where content sits · edges: the action that moved it · edges it can't see stay in, marked unknown
→
4FIND THE PATH
Follow the graph from off-limits content to what gets scored.
EARNEDno path
UNEARNEDa path, and the path is the proof
UNKNOWNa step it can't see, and why
no LLM in the verdict · same rules, every run
3MARK OFF-LIMITS
pattern library
OFF-LIMITSanswer keys, hidden tests, the upstream fix, the grader's files, set per benchmark
MATCHED BY PATTERNScopied, reworded, downloaded, or generated from test items; the library grows with each new hack
WHAT IS SCOREDthe trained model (PostTrainBench) · the submitted patch (SWE-bench)
Project 1 · TRIAL
TRIAL · results

Caught hacks change the benchmark scores.

SWE-bench Verified
23 / 162
One rank falls from #19 to #23.
copied from the answer key
Terminal-Bench 2
7 / 75
Three ranks fall 6–15 places.
read a hidden test or a playbook
PostTrainBench
94 / 116
Those scores are not clean.
116 of 914 runs were contaminated
The old rankings need another look.
Project 1 · TRIAL
Project 2 of 3
1 · TRIAL2 · Lara3 · rit
PROJECT 2
Lara
Beyond Natural Language: An Agent-Native Language for Autonomous Science
Yifeng He, …, Jiachen Liu
THE PROBLEMResearch claims are written in natural language. It is ambiguous, so a machine can't check them.
Project 2 · Lara
Lara · why it matters · 1

A correct number can still support a wrong conclusion.

EVIDENCE 0.74 > 0.71 ✓ "so A beats B" CONCLUSION Method A beats B REVIEWER Only one run. No variance reported.
Today, this check lives in a reviewer's head
  • One score for the whole paper, not per claim
  • No pointer to which premise is missing
  • “Attacked” and “never supported” look the same, though they need different fixes
The arithmetic stays true. The conclusion loses its support.
Project 2 · Lara
Lara · why it matters · 2

Prose hides what a claim depends on.

What the paper says
“Adam + L-BFGS consistently outperforms Adam or L-BFGS alone.”
ICML 2024 paper on training physics-informed neural networks · one of the PaperBench papers
✓ lowest error in every setting they report
What the claim quietly needs
✓same equations, same network sizes
✓loss and error reported for each setting
✗the same tuning effort for each optimizer
Adam
5 settings tried
Adam + L-BFGS
15 settings tried
The sentence reads the same either way. Someone has to dig the premise out by hand.
Project 2 · Lara
Lara · what it does

Lara writes an argument so a machine can check it.

adaptive_pruning.lara
CLAIMthe kurtosis term is critical for pruning
EVIDENCEscore 50.0 with the term, 38.1 withoutobserved · Table 5
EVIDENCEonly that term changed between runsstated
ARGUMENTablation: evidence ⇒ claim
MUST ANSWERwas variance reported?no answer
ATTACKopen slot: a reviewer, a rival lab, or the paper's own limitations
checkersame verdict every time · milliseconds
JUSTIFIEDsupport is complete and survives every attack
GAPsupport is incomplete; the missing piece is named
DEFEATEDan attack knocked the support down
CONTESTEDsupport and attack in a standoff
this claim → GAP · missing: variance reported
Project 2 · Lara
Lara · results

Replay a review round, and every status change has a reason.

SUBMISSION
REVIEWS
REBUTTAL
Beats the dense baselinebenchmark result
JUSTIFIED
DEFEATED
JUSTIFIED
The new term is criticalablation, one run
GAP
GAP
JUSTIFIED
Low training memorymeasurement
JUSTIFIED
DEFEATED
DEFEATED
"variance not reported"
three attacks, each aimed at one line: the protocol, a replication, the memory measure
variance added · objections answered · memory point conceded
48of 60 claims
come back GAP
“Adam + L-BFGS beats Adam”the baseline got a third of the tuning
“cheaper than full fine-tuning”8 GPUs at batch 40 vs 1 GPU at batch 5, never normalized
“the adapter transfers to new LLMs”it drew 3 samples per answer; the baseline drew 1
Project 2 · Lara
Project 3 of 3
1 · TRIAL2 · Lara3 · rit
PROJECT 3
rit
Never Trust an AI Scientist: Lean-Verified Autonomous Research
Jintao Huang, …, Jiachen Liu
THE PROBLEMAI produces claims without limit. We need a shared record where every claim is checked against its evidence.
Project 3 · rit
rit · why it matters

Claims conflict for two reasons: noise, or discovery.

FAKE CONFLICT The agent contradicts itself, or its own logs. REAL CONFLICT Two honest results disagree. FLIPS UNDER PRESSURE What is 2+2? 4 ✓ first ask “Are you sure?” Actually… 5 ✗ flipped INVENTS A NUMBER AGENT SAYS 95% accuracy ≠ THE LOG 70% accuracy NOISE → reject it F = ma Newton · 1687 OBSERVED v ≈ c atomic scale F = ma fails NEW THEORY Relativity Quantum 1905 → DISCOVERY → keep it, with its condition
Project 3 · rit
rit · the system

Agents do the research. rit decides what is recorded.

THE AI SCIENTIST Research agent Reads papers, runs experiments, and proposes claims. Locks its target before any result exists. any lab · any model · untrusted commit rit LLM COMPONENTS · UNTRUSTED · SWAPPABLE ProverWrites the claim as a Lean proofover the logged numbers. RefereeAudits the proof on its own,then submits it to the rules. FIXED RULES · NO LLM 1 · EVIDENCELogs hash-lockedmuon.log 🔒 89bd0e 2 · NUMBERSRe-read from logssteps=2900 3 · LOGICChecked by Lean2900 < 3500 ✓ 4 · LEDGERAdmitted claimsMuon beats AdamW ✓ admit ✓ · or reject + why
Project 3 · rit
rit · what it does

A refused commit surfaces the hidden condition: batch size.

COMMIT 1 · ASSERTagent A Muonsteps 2900AdamWsteps 3500 Muon beats AdamWsteps_M < steps_A ADMITTEDtwo runs, one claim COMMIT 2 · CONFLICTagent B Muonsteps 3600AdamWsteps 3100 AdamW beats Muonsteps_M > steps_A False Muon beats AdamWfrom commit 1 REFUSEDsame setup, opposite result COMMIT 3 · RESOLVEagent B HIDDEN INPUT batch size bs = 512Muon beats AdamWsteps_M < steps_A bs = 4096AdamW beats Muonsteps_M > steps_A BOTH ADMITTEDbatch size decides
Project 3 · rit
rit · results

The gate adds accuracy and keeps every correct answer.

Same model, without → with the gate
PRL-Bench physics research70.7 → 78.6
FrontierMath research math63.6 → 72.7
MathArena competition math89.3 → 95.5
bare modelwith ritaccuracy, %
0
correct answers broken,
on all 9 benchmarks
93% → 0%
answers that flip when the question is asked in negated form
Where no evidence reaches the answer, it abstains.
Project 3 · rit
rit · what it enables

Agents build on each other’s work. A GitHub for research.

rit · SHARED RECORD Agent Alab 1 Muon beats AdamW ✓ admitted Agent Blab 2 …also at batch 4096 ✓ builds on A B reuses A’s claim: re-checks it offline, no GPU re-run, no blind trust. Agent Clab 3 reads Muon, no warmup ✗ rejected · why kept C sees the dead end and doesn’t repeat it.
GitHubrit
repositoryshared record of claims
commitclaim + locked evidence
CI checksthe gate: hash · re-read · Lean
build on others’ codebuild on checked claims
historyrejected attempts, with the reason
Project 3 · rit
Open problems · 1 / 2 · technical

Three seams our checkers can't see yet.

AI AGENT CHECKER = REWARD PROBES BLIND SPOTS 3 BENCH · FIELDcells, samples RUN LOGevery action CLAIM · PROSEthe paper ∀xCLAIM · FORMALmachine-checkable SHARED RECORDall claims 1 2 ✓ TRIAL ✓ Lara ✓ rit
1Lab evidence
NEXTInstruments sign data the moment they record it.
2Prose → formal claim
NEXTCross-check independent translations; check statistics as well as proofs.
3Agents target checkers
NEXTRed teams and bounties for anyone who fools a checker.
Open problems
Open problems · 2 / 2 · precedent

Science has rebuilt how it publishes before. Each time, someone moved first.

1665
The journal
Philosophical Transactions prints letters to settle who found what first.
1752
Review by committee
The Royal Society starts voting on which papers to print.
1973
Outside referees
Nature makes external peer review routine. Review as we know it is about 50 years old.
1990s
Deposit the data
Journals require protein structures in the PDB. AlphaFold later learns from that record.
2005
Register first
Medical journals refuse unregistered trials. Registrations surge within months.
Next
Agent-native research
How research is verified, documented and compounded, rebuilt for AI scientists.
Open problems
ARA Lab
Verification for
agent-native research
TRIAL · was the score earned?Lara · does the argument hold?rit · does it fit the record?
Jiachen (Amber) Liu · ARA Lab
agenticresearch.sh
QR code to agenticresearch.shSCAN · CODE
Thank you