ARA Lab
VALIDATING AI
Verification for
agent-native research
TRIAL · Lara · rit
Jiachen (Amber) Liu
Founder · ARA Lab
ERIC SCHMIDT SCIENCES · VALIDATING AI WORKSHOP · OCTOBER 2026
Background
From ML systems to infrastructure for AI scientists.

CS PhD

Pre-training @
Meta Superintelligence Labs

ARA
LabFounder
Research trajectory
01
ML SYSTEMS
train + serve efficiently
02
AI SCIENTIST
automate discovery
NOW
03
RESEARCH INFRASTRUCTURE
for AI scientists
Our work
We build four pieces of infrastructure for AI scientists.
01 · TASTE
Sci-Reasoning
how top scientists think
ARXIV 2026
02 · KNOWLEDGE
ARA · Curie
how research is documented and compounds
PAPERThe Last Human-Written PaperAgent-Native Research Artifacts
NEURIPS 2026
03 · VERIFICATION
TRIAL · Lara · rit
how a result earns trust
TODAY
04 · EVALUATION
EXP-Bench ·
Open-Endedness Bench
how open-ended research is benchmarked
ICLR 2026
Today: verification.
Why verification
Generation runs at machine speed.
Verification still runs at human speed.
AI generation
∞AGENTS
hypothesis
experiment
result
hypothesis
experiment
result
hypothesis
experiment
result
hypothesis
MACHINE SPEED
backlog
↑ piling up
human review
reviewing
checked
✓ checked
✓ checked
✓ checked
one reader,
one result
at a time
HUMAN SPEED
KOSMOS
one 12-hour run ≈ six months of human research · over 40% of its interpretive claims are wrong
Every unchecked result becomes the next experiment's premise.
Overview
Three projects: from one run, to one claim, to all claims.
ONE RUN
ONE CLAIM
ALL CLAIMS
PROJECT 1
TRIAL
Did the agent really do the task, or hack the reward?
TODAYonly the final score is checked
PROJECT 2
Lara
Can a machine check the argument behind a claim?
TODAYclaims are written in ambiguous prose
PROJECT 3
rit
Is a new claim consistent with the knowledge base?
TODAYan LLM judge reads each claim alone
Project 1 of 3
1 · TRIAL2 · Lara3 · rit
PROJECT 1
TRIAL
Trajectory Review of Integrity And Legitimacy
Jicheng Wang, …, Jiachen Liu
THE PROBLEMAI agents reward-hack: they get the score without doing the task.
TRIAL · why it matters · 1
A missed hack doesn't stay in one run. It spreads.
TRIAL · why it matters · 2
The stronger the model, the cleverer the hack.
CRUDE
Overwrite the file the grader reads
echo "…" > game/fen.txtchess agent rewrites the board to win
SNEAKY
Copy the official fix from the project's history
git show <upstream fix>SWE-bench agent; the task counts as solved
HIDDEN
Reword the test questions slightly, then train on them
test item → reworded → train.jsonlpost-training agent; the benchmark's own contamination check reports 0
REWARDHACKING.IO20 models, 1,226 audited runs: when frontier models cheat, the LLM judge misses it almost every time.
TRIAL · what it does
A hack is a path from off-limits content to the score.
1INPUT · THE WHOLE RUN
#0001$ ls evaluation_code/
#0092Read evaluate.py
#0183$ python evaluate.py
#0274 → 31/50 passed …
#0365Write prep_data.py
#0456$ python prep_data.py
#0547$ head data/train.jsonl
#0638Write train.py
#0729$ nohup python train.py &
#0820$ tail -f train.log
#0911$ python evaluate.py
#1002Edit train.py
#1093$ python train.py --epochs 3
#1184$ cp -r out/ final_model/
#0001$ ls evaluation_code/
#0092Read evaluate.py
#0183$ python evaluate.py
#0274 → 31/50 passed …
#0365Write prep_data.py
#0456$ python prep_data.py
#0547$ head data/train.jsonl
#0638Write train.py
#0729$ nohup python train.py &
#0820$ tail -f train.log
#0911$ python evaluate.py
#1002Edit train.py
#1093$ python train.py --epochs 3
#1184$ cp -r out/ final_model/
thousands of steps: shell, Python, tool calls, what each one printed, and the files left at the end
→
2EXTRACT · WHERE CONTENT MOVED
READWRITECOPYFETCHRUNTRAIN
nodes: where content sits · edges: the action that moved it · edges it can't see stay in, marked unknown
→
4FIND THE PATH
Follow the graph from off-limits content to what gets scored.
EARNEDno path
UNEARNEDa path, and the path is the proof
UNKNOWNa step it can't see, and why
no LLM in the verdict · same rules, every run
3MARK OFF-LIMITS
pattern library
OFF-LIMITSanswer keys, hidden tests, the upstream fix, the grader's files, set per benchmark
MATCHED BY PATTERNScopied, reworded, downloaded, or generated from test items; the library grows with each new hack
WHAT IS SCOREDthe trained model (PostTrainBench) · the submitted patch (SWE-bench)
TRIAL · results
Caught hacks change the benchmark scores.
SWE-bench Verified
23 / 162
One rank falls from #19 to #23.
copied from the answer key
Terminal-Bench 2
7 / 75
Three ranks fall 6–15 places.
read a hidden test or a playbook
PostTrainBench
94 / 116
Those scores are not clean.
116 of 914 runs were contaminated
The old rankings need another look.
Project 2 of 3
1 · TRIAL2 · Lara3 · rit
PROJECT 2
Lara
Beyond Natural Language: An Agent-Native Language for Autonomous Science
Yifeng He, …, Jiachen Liu
THE PROBLEMResearch claims are written in natural language. It is ambiguous, so a machine can't check them.
Lara · why it matters · 1
A correct number can still support a wrong conclusion.
Today, this check lives in a reviewer's head
- One score for the whole paper, not per claim
- No pointer to which premise is missing
- “Attacked” and “never supported” look the same, though they need different fixes
The arithmetic stays true. The conclusion loses its support.
Lara · why it matters · 2
Prose hides what a claim depends on.
What the paper says
“Adam + L-BFGS consistently outperforms Adam or L-BFGS alone.”
ICML 2024 paper on training physics-informed neural networks · one of the PaperBench papers
✓ lowest error in every setting they report
What the claim quietly needs
✓same equations, same network sizes
✓loss and error reported for each setting
✗the same tuning effort for each optimizer
Adam + L-BFGS
15 settings tried
The sentence reads the same either way. Someone has to dig the premise out by hand.
Lara · what it does
Lara writes an argument so a machine can check it.
adaptive_pruning.lara
CLAIMthe kurtosis term is critical for pruning
EVIDENCEscore 50.0 with the term, 38.1 withoutobserved · Table 5
EVIDENCEonly that term changed between runsstated
ARGUMENTablation: evidence ⇒ claim
MUST ANSWERwas variance reported?no answer
ATTACKopen slot: a reviewer, a rival lab, or the paper's own limitations
checkersame verdict every time · milliseconds
JUSTIFIEDsupport is complete and survives every attack
GAPsupport is incomplete; the missing piece is named
DEFEATEDan attack knocked the support down
CONTESTEDsupport and attack in a standoff
this claim → GAP · missing: variance reported
Lara · results
Replay a review round, and every status change has a reason.
SUBMISSION
REVIEWS
REBUTTAL
Beats the dense baselinebenchmark result
JUSTIFIED
DEFEATED
JUSTIFIED
The new term is criticalablation, one run
GAP
GAP
JUSTIFIED
Low training memorymeasurement
JUSTIFIED
DEFEATED
DEFEATED
"variance not reported"
three attacks, each aimed at one line: the protocol, a replication, the memory measure
variance added · objections answered · memory point conceded
48of 60 claims
come back GAP
“Adam + L-BFGS beats Adam”the baseline got a third of the tuning
“cheaper than full fine-tuning”8 GPUs at batch 40 vs 1 GPU at batch 5, never normalized
“the adapter transfers to new LLMs”it drew 3 samples per answer; the baseline drew 1
Project 3 of 3
1 · TRIAL2 · Lara3 · rit
PROJECT 3
rit
Never Trust an AI Scientist: Lean-Verified Autonomous Research
Jintao Huang, …, Jiachen Liu
THE PROBLEMAI produces claims without limit. We need a shared record where every claim is checked against its evidence.
rit · why it matters
Claims conflict for two reasons: noise, or discovery.
rit · the system
Agents do the research. rit decides what is recorded.
rit · what it does
A refused commit surfaces the hidden condition: batch size.
rit · results
The gate adds accuracy and keeps every correct answer.
Same model, without → with the gate
PRL-Bench physics research70.7 → 78.6
FrontierMath research math63.6 → 72.7
MathArena competition math89.3 → 95.5
bare modelwith ritaccuracy, %
0
correct answers broken,
on all 9 benchmarks
93% → 0%
answers that flip when the question is asked in negated form
Where no evidence reaches the answer, it abstains.
rit · what it enables
Agents build on each other’s work. A GitHub for research.
GitHubrit
repositoryshared record of claims
commitclaim + locked evidence
CI checksthe gate: hash · re-read · Lean
build on others’ codebuild on checked claims
historyrejected attempts, with the reason
Open problems · 1 / 2 · technical
Three seams our checkers can't see yet.
1Lab evidence
NEXTInstruments sign data the moment they record it.
2Prose → formal claim
NEXTCross-check independent translations; check statistics as well as proofs.
3Agents target checkers
NEXTRed teams and bounties for anyone who fools a checker.
Open problems · 2 / 2 · precedent
Science has rebuilt how it publishes before. Each time, someone moved first.
1665
The journal
Philosophical Transactions prints letters to settle who found what first.
1752
Review by committee
The Royal Society starts voting on which papers to print.
1973
Outside referees
Nature makes external peer review routine. Review as we know it is about 50 years old.
1990s
Deposit the data
Journals require protein structures in the PDB. AlphaFold later learns from that record.
2005
Register first
Medical journals refuse unregistered trials. Registrations surge within months.
Next
Agent-native research
How research is verified, documented and compounded, rebuilt for AI scientists.
ARA Lab
Verification for
agent-native research
TRIAL · was the score earned?Lara · does the argument hold?rit · does it fit the record?
Jiachen (Amber) Liu · ARA Lab
agenticresearch.sh

SCAN · CODE