Text version of this article
Also available as Markdown.
AI Should Build Its Own Research World Model
Opus 4.8 Cleared ARC-AGI’s Final Level in One Shot.
A field report · June 2026
Chengyang Shi · Xianglin Ji · Jiachen Liu · July 2026
TL;DR Can AI truly understand a completely unknown world? What is the real key to enabling a system to crack novel problems like a scientist? Under the current paradigm, model weights are locked during inference; the moment a context window expires, any experience gained through trial and error vanishes—AI cannot naturally grow "scientific intuition" through a loop of exploration and reflection as humans do. To break this bottleneck, our approach is to explicitly equip the AI with an external cognitive architecture, allowing it to record as it explores and dynamically crystallize discovered mechanics into a cross-context, queryable "research world model." To test this, we dropped a Claude Code + Opus 4.8 agent into a completely unseen puzzle world (ARC-AGI-3's ls20) and tasked it with clearing the game. For context, the baseline pass rate for frontier LLMs' bare exploration in such puzzles is ~1%. Yet, empowered by the world model, the system autonomously decoded all underlying mechanics. When faced with the final and hardest level, the agent achieved a one-shot victory with the model's assistance. We show that with this training-free and self-evolving approach, an AI agent is able to understand the physical rules of an unknown world and build its own world model.
In my years of doing research, I’ve realized that the essence of scientific inquiry is a continuous feedback loop inside an unknown world: explore, hypothesize, act, observe, and reflect. Human science relies on this loop to iterate out intuition, which then slowly crystallizes into papers. When we explore the unknown, intuition naturally grows in our minds.
However, when AI explores an unknown world, it cannot naturally grow this kind of abstract thinking ability. Could we, then, explicitly build an "intuition system" for AI? We often speak of intuition as something mystical, but theoretically, it can be recorded. This brings us to the reinforcement learning paradigm: in RL, we can indeed directly update the model's weights, but this is merely a highly implicit form of knowledge compression. It lacks interpretability and cannot be easily extracted or debugged. All of this prompted us to wonder: could we build an external cognitive architecture for AI that can continuously record, track, and be queried, allowing the AI to compound knowledge over time? I call this architecture a research world model.
How the world model gets built. The bottom row is the same ARA archive, growing across every level: 7 laws after the first night, 14 after level 5, and at its thickest 15 laws, 7 recipes, 75 research-tree nodes. The agent gets swapped after each level, but the same world model keeps compounding: each incumbent writes what it learned into the archive with the red pencil (research-manager, which only writes at moments of closure), reads first on arrival, and asks through the blue magnifier when stuck (research-foresight, read-only). By level 7 it wrote no new law at all: ask first, then walk — 63 steps, one shot.
The two recordings below show the contrast: by building its own world model as it explored, when the AI tackled the final level, it finished the exploration and cleared the game in just 63 steps. Conversely, without such a world model, even if handed the "clear records" of earlier levels, it wandered in a dead end for over 5000 steps with zero progress. This vividly proves the importance of a world model for AI to break through unknown deadlocks. What exactly was in that written archive: that's what this post is about.
the one with the world modelLevel 7 · WIN
the one with only the route5,168 steps · zero progress
wallwalkable floorthe block (itself)panel & key-box bordersbudget bar (one tick per step)
Two real recordings, looping; an element-by-element annotation of the scene is on the first frame in §2. Left: the final stretch of the full run. Right: what the control did on the same stretch. In its five thousand-plus steps it flipped the panel more than two thousand times, and never once assembled the target pattern, never opened the key-box pocket.
01 · the only section that argues
§1 Into a Game With No Rules: Research Requires Explicit World Models, Not Just Weights
To run this experiment, we needed a completely unknown environment for the AI, where it could explore and understand the underlying physics on its own. Through the loop of exploration, summarization, reflection, and hypothesizing, it would slowly precipitate an explicit understanding of the world to eventually clear the game. This environment could be a frontier of scientific research, but here, we chose an interactive puzzle from ARC-AGI as our experimental subject to explore AI's progress in continuous abstraction and reflection capabilities.
And this is only one example. The world could be a game; it could just as well be a stretch of protein, an unsolved equation, a patch of scientific frontier nobody has charted. What we actually want to know isn't whether an AI can clear a game; it's whether, dropped somewhere no one has ever been, it can walk out the way a scientist would.
Research shouldn't live compressed in weights and intuition.
02 · rung one language
§2 Naming and Filing the Unknown: Inventing Representations is the First Step to Understanding
Here's how we ran it:
how we ran it · the short version
We tossed one prompt into Claude Code, and the whole run took off from there.
The flow: the main agent spawns the game harness and dispatches one play-subagent per level, keeping only orchestration, the ledger, and syncing for itself. All game reasoning happens inside the subagents.
The subagents talk to the world model through two skills. Reads use /research-foresight: a subagent must read the ARA (its world model) before acting, and after roughly a hundred stuck actions calls the skill to query the world model directly — read-only.
Writes use /research-manager: lessons crystallize into the archive only through it, at moments of closure; the playing agent never hand-writes a file.
Now, look at what it saw when it opened its eyes.
[Figure: The raw first frame of ls20 level 1]
Level 1, frame 1. This is everything it can see. No rules, no goal, no tutorial. Just four arrow keys, and a bar along the bottom that gets one cell shorter with every press.
If it were you, which key would you press first?
It pressed all four. Some directions moved; others hit something and didn't budge, but the bar at the bottom got shorter all the same. Hitting a wall isn't free. This world charges for mistakes and doesn't post the prices. What followed was basically fiddling: it shoved the grey-capped purple block around, stepped across an inconspicuous black dot, and the pattern in the bottom-left frame jumped; pushed the block into the frame up top, lost; waited for the pattern to match and pushed again, and the whole screen redrew. Level 2.
Before closing out its first session (call it the first night), it did the first thing a naturalist does after washing up on an island: it named the things it had never seen. The square that moves became block; the little dot that makes the pattern jump when stepped on, switch; the frame up top holding the target pattern, key-box; the rare floor spots with no obvious purpose it offhandedly called speck. Not imaginative names, but they all worked. Once they were settled, it started writing entries into a file.
ara-ls20/logic/concepts.md · the «switch» entry · quoted verbatim
## switch
A rare color-0/1 floor speck the block steps onto to toggle/cycle the lock
pattern shown in the panel. Not consumed (still present after the block moves
off). The per-level transform differs: L1's switch (R6C3) toggles only the
middle row (2 states); L2's switch (R9C9) is a 4-state cycle (C09, C10).
The first time I opened this concepts.md I got snagged on R6C3 and color-0/1 too. Let's not rush to explain. The unreadability is itself evidence: from their first line, these entries were never written for humans.
The first night's bill: 24 steps to work the level out the first time; 13 steps for the shortest clear afterwards; 0 deaths. By close of session, seven laws it was willing to commit to writing already lay in its world model (it calls these claims, each numbered and filed), including C07, the one that would later save its life: only when the lock is satisfied does the walled-shut key-box open.
Naming is cutting chaos into parts you can think with.
Human wisdom started to pile up not because some generation suddenly got smarter, but because we invented language: for the first time, experience could leave the head that produced it. That's the step it re-enacted in this little world: first the words, then an understanding you can pass on.
Up to this point, though, what it wrote was still nothing more than a travel diary: see, record. That changes at level 5, when the words it invented start doing mathematics on their own.
03 · rung two mathematics
§3 From Discovering Operations to Derivation: AI Autonomously Extracts and Uses Abstract Mathematical Structures
Level 5 has two specks. A normal switch is "press it, land in some state"; level 1's switch was like that, flipping the same row back and forth. These two are not. Step on the first and the panel pattern cycles forward one notch, coming back around after six; step on the second and the whole pattern rotates 90° clockwise.
They aren't buttons. They're operations.
Operations differ from buttons in one lethal way: order starts to matter. Rotate-then-cycle is not cycle-then-rotate. If you've ever held a Rubik's cube you already get it: no single twist restores the cube, a solution is a "sentence" spelled out of twists, and some states you can never twist your way into. The lock in front of it was a little 3×3 Rubik's cube. Nobody told it that; it found out on its own. And the next few things it did happen to be the first few things a mathematician does when handed one.
Step one: naming. It wrote the cycling operation as P and the rotation as Q. The two symbols then get cited across files, in claims and trace alike; referring to the same operation over and over needs a stable name. Teenage Feynman did the same while teaching himself trigonometry: finding the sin and cos notation ambiguous, he invented symbols of his own and worked fast with them; later he dropped them. "I found out that if I was going to talk to anybody else, I'd have to use the standard symbols." Feynman had an audience, so he surrendered. This agent's only audience is its future self, so it kept what Feynman gave up. And seen from inside the board, P and Q sit closer to the skin than any human word: it doesn't need the metaphors of "cycle" and "rotate," it needs letters that can spell sentences.
Step two: drawing the operations as a map. In its notes a panel state looks like this: 110/011/101. Ciphertext again. But the string is really a 3×3 board of lights, 1 lit, 0 dark. Drawn out, it stops being ciphertext:
These six states are all P owns: one notch per press, back where you started in six. It calls the circle orbit B, an orbit: press only P and you never leave it. Babbage said it a hundred and ninety years ago: the whole point of algebraic notation is to "condense into small space, a large amount of meaning."
Q does one thing: rotate the whole board 90° clockwise. Turn the left state once and you get the right one. The same turn shows up again on level 7.
Step three: discovering orbits and bridges. Level 6 doubles the stakes: two locks, and one lock's target pattern won't come no matter how much you press P: it lives in another circle P can't touch. It didn't keep banging; it stopped and wrote this:
ara-ls20/logic/claims.md · C14 · the L6 situation · quoted verbatim (excerpt)
P is a genuine period-6 NON-permutation (on-bit count varies), and — the new
structural fact — P has orbits … can NEVER reach the target in orbit A — the
LOWER alignment GENUINELY requires a {P,Q} WORD where Q bridges between P's
orbits, then P walks to the target within the right orbit.
… the LOWER word QQPPPPPQ (offline BFS over {P-cycle,Q}) is forced and ENDS in Q
Step four: an impossibility theorem. There's an all-caps NEVER in that quote. It went looking for a shortcut, and what it found instead was a proof that no shortcut exists: by P alone, that target can never be reached; not "hasn't tried hard enough," mathematically unreachable. Trial and error can only tell you what works; only theory can tell you what will never work. The moment it knew this, it stopped spending money on the dead road.
Now look back at where this section's two words were found. Level 5's PPQQ, level 6's QQPPPPPQ: neither came from inside the game; both came from paper. In this world every step costs money and a wrong one can cost a life; deriving on written-down rules costs nothing. It pinned down how P and Q behave, combed through the states on paper, found the solution, then walked back into the game and pressed the keys.
ara-ls20/logic/claims.md · C14 · the evidence line · quoted verbatim
(2,3) period-6 + (7,3) 90°-rotation, offline BFS gave the forced word PPQQ,
verified live to reach 101/110/011, then delivered → levels 4→5
Shortest clear per level, measured after the first solve. Level 5 is harder than level 4 yet takes fewer steps (56 < 60): the first time paper derivation came out cheaper than hitting walls. The red bar is level 7: 63 steps, the clean execution after consulting the world model; §5 is that run, start to finish.
The mathematician Steven Strogatz wrote a well-known group-theory essay for lay readers that ends on exactly this note: the same structures keep turning up in the least related places, "from the symmetries of water molecules to the logic of a pair of light switches." This agent ran into it inside two game switches. Nobody told it this was group theory. Whether it "understands" it, I honestly can't say; what we can see is that whichever piece it needed, it rebuilt.
Mathematics is the moment the things you wrote down start doing work on their own.
And once a language starts computing, its next step is to start doubting itself.
04 · rung three method
§4 Writing and Refuting Wrong Laws: The Value Lies in Self-Correction and Falsifiable Records
This world model is not true line by line; its best pages are precisely the wrong ones.
Across the first four levels it kept seeing the block "teleport": clearly here, then somewhere else one step later. Naturally enough it hypothesized portals, and even started tabulating them (I'd probably have guessed the same). Then one day it laid the death records next to the "teleport" records, and they lined up one for one:
ara-ls20/trace/_l4_raw.md:34 · quoted verbatim (the all-caps and exclamation marks are its own)
## !!! UNIFYING DISCOVERY: the "teleports/portals" were BUDGET-DEATH RESPAWNS
(no portals exist) !!!
…
THERE ARE NO PORTALS on L1-L4. The block always moves 1 macro-cell/press;
walls = no-op.
So far, so good. But if the story stopped here it would just be "admits mistakes, fixes them." Two levels later, the right half of level 6 wouldn't yield no matter what it tried; after troubleshooting in circles, it had to write one more line:
ara-ls20/trace/_l6plus_raw.md:109 · quoted verbatim
## ★★ L6 HAS PORTALS (corrects O22 "no portals confirmed") — right side is
portal-trapped
Look at the parenthesis. When it overturned the old conclusion, it cited that conclusion by number: it was keeping files on its own errors. "No portals" was not deleted; it still sits in the entries for levels 1–4 with its boundary drawn sharp: within its stated scope, that sentence is correct to this day.
One more mistake, tellable in a sentence: the few % marks at the right end of the budget bar it first guessed were "goal markers." Wrong: they were lives. As long as it never died, it never got to watch that counter move. A survivor can't see survivorship bias until it has died a few times.
Level 4 holds an even smaller incident. Probing a speck from the north and getting nothing, it nearly wrote down "this spot cannot be stepped on." What stopped it wasn't luck; it was a rule it had written itself two levels earlier, the "sample enough before concluding" clause in claim C10:
ara-ls20/logic/claims.md · C10 · the L4 situation · quoted verbatim
The C10 must-sample-first lesson was directly load-bearing on L4: the colour
selector was sampled 6× (refuting a "double-duty" shortcut) and the pattern
speck — initially mis-judged "not steppable" — was a C10 under-sampling error
corrected by trying an untried side. So the recipe is now VERIFIED through an
L4 win.
By this point the notes were correcting it, not the other way around.
It has also blindly trusted things just because they were written down, which is why refutations have to stay on file. Two overturned claims (C03, C08) still lie in its world model, marked refuted, and nobody gets to delete them: they're part of the immune system. While we're at it, here's what the price of honesty looks like:
real in-game deathsresearch dead ends (falsified hypotheses etc.)
Two kinds of "failure," and they must be counted apart. Deaths cluster in the first three levels, before there was a world model to consult; research dead ends aren't accidents, they're building material for theory. Our own first tally counted all 19 dead-end nodes as "failures," and only a recount pulled them apart; even the way to count errors, we had to learn from it.
A language that has learned to compute learns, next, to doubt itself.
And a notebook that doubts itself is one step short of answering questions.
05 · rung four prediction
§5 Stuck in the Dark, Querying the World Model: Replacing Blind Trial-and-Error with Accumulated Predictions
Only one thing matters about level 7: it is a dark room. The whole map is buried in fog; only a small circle around the block is visible, lighting up wherever it walks and going dark again behind it. Every step costs two points of budget, double any earlier level.
[Figure: Level 7: a small circle of sight in the fog]
A real frame from level 7. The grey isn't wall, it's the unknown: the lit circle moves with the block, like carrying a lamp through someone else's house, looking for a key.
It carried the lamp through the whole house, then hit the deadlock: the key-box is locked inside a pocket, walled on every side, unreachable even by portal. The pocket only opens once the lock is open; and by its own law C07, opening the lock requires first reading the target pattern, which sits on the key-box. To get through the door you must already be through the door. After 515 steps of reconnaissance, this context's window ran out. It did not solve the level.
But it left a will, written for the next session:
ara-ls20/trace/exploration_tree.yaml · N72 · dead_end · quoted verbatim (excerpt)
failure_mode: The pocket … is PORTAL-ISOLATED: exhaustive 4-dir probing of
every reachable floor cell … found NO warp landing in it; all center portals
loop among {center,(2,6),right-via-(8,7)} …
lesson: Most likely the pocket OPENS only after the LOCK IS ALIGNED (C07
'walled-shut key-box becomes enterable', extended to the fog mechanic) — but
the pattern target is unreadable until the box reveals (chicken-and-egg).
NEXT SESSION leads: try a GUESSED-target alignment (colour-% + each
reachable {P',Q} pattern state) and re-test the (9,5)→(10,5) walls …
status: open
A context died on this level and willed its battle plan to the next self.
The new context that took over didn't go back to banging on walls. First it read through the world model its predecessors had written; then it took the will above (node N72 in the research tree — all 75 nodes can be opened one by one) and asked. research-foresight is the questioning interface over this state; wm-predict, in the archive excerpts below, is its old name at the time. Level 6 has one such exchange kept on file, end to end:
ara-ls20/staging/observations.yaml · O25 · one complete world-model exchange · quoted verbatim (excerpt)
wm-predict over ara-ls20/ was asked: does the shared panel reset after a
partial delivery, is there a forced order, how is the LOWER box reached.
It answered, separating grounded inference … from a NAMED speculative leap
(panel PERSISTS after a partial delivery; order UPPER-first; LOWER reached
after col-10 unblocks), flagged medium confidence, and gave concrete
falsifiers. ALL THREE speculative predictions were then LIVE-CONFIRMED (N62).
I have to own something here: for the level-7 query, the question's exact wording never made it into the archive; our process slipped. What survives is the list it brought, and the judgments it carried back:
ara-ls20/trace/exploration_tree.yaml · N71/N74 · world-model query records · quoted verbatim (excerpt)
wm-predict consulted 2x: predicted multi-region composed-{P',Q} multi-control
lock (matched) and target colour=% (CONFIRMED — color-8); its center-DOWN-
warp-to-pocket bet was REFUTED.
—
hypothesis: WM-predict (C07-in-fog): the portal-isolated (10,5) pocket opens
only when the lock is aligned to a derivable target (colour-alone is dead,
N72) …
Both facts live in this one record. The world model placed a losing bet; that REFUTED was written by its own hand; and the word that paid was derivable. The target wasn't unreadable, it was being read in the wrong place. Following that judgment, it went back to the "lone colored cell" and picked at it pixel by pixel. It took us a replay to catch what it had caught: not a dot at all, but a 3×3 target pattern shrunk to one pixel per cell, half-covered by fog. Decoded: 101/110/011, color %.
The rest of the procedure already appeared in §3: lay out the rules of P′ (level 7's version of P) and Q on paper, run an offline BFS over the 24 reachable states, and pull out the unique shortest word:
PPPPPQQ: five cycles, two rotations, initial state to target (red frame). The whole path was found on paper; not one step of it was trial-and-errored in the game.
Then the execution: set the color, cross the corridor five times to catch the patrolling P′, ride portals and supply stations across the map, intercept Q twice in the right-hand gallery, and when the panel lit up with the target pattern, push the door. 284 cells redrew at once, state=WIN. The decisive post-breakthrough run: 63 steps start to finish, zero deaths. All seven levels done.
harness/scratchpad/l7solve.py · the post-breakthrough execution script · quoted verbatim (file header)
"""L7 full solver. Target: pattern 101/110/011 colour-% via word PPPPPQQ,
deliver at (9,5)->A2->(10,5).
Phases (each idempotent, callable separately so I can supervise):
colour -> set colour % at selector (8,1)
p N -> catch P' (left row-8 corridor cols2-3) N times
q N -> catch Q (right col-10) N times
deliver -> goto (9,5), A2 into (10,5)"""
attempt one · recon by lamplight515 steps · unsolved
attempt two · after asking the world model63-step finish · WIN
The same dark room, two lifetimes apart. Left: the first context feeling its way. Right: the new context heading straight for the point with the judgments it queried for (399 steps in all including fog-scouting, of which the clean post-breakthrough execution is the 63).
What it held was no longer a pile of notes. It was a world it could ask.
05½ · anatomy
§5½ Opening the Archive: The Research World Model is a Continuously Evolving Epistemic State (ARA)
With the story this far along, it's time to make "research world model" precise. It is not a log, not RAG, not a longer context, and not merely a knowledge graph: those only answer "what got stored." It is a continuously evolving epistemic state: it defines objects and concepts; poses refutable hypotheses; records experimental interventions and observations; keeps, for every conclusion, its evidence sources, confidence, and scope of validity; turns failures into constraints; and lets future agents predict, question, reproduce, and rewrite past conclusions. The container that holds this state has a formal name in our hands, ARA (Agent-Native Research Artifact): written by agents, for agents, no storytelling, only evidence, dead ends kept on file exactly as they happened. The case for rebuilding research's container away from the paper format was made in an earlier post, The Last Human-Written Paper. Every clause of this definition has a physical exhibit. Before the ablation, lay the whole thing on the dissection table and verify it clause by clause:
Everything the system has, in one picture. The robot in the middle is the agent (Opus 4.8, weights frozen throughout); to its left, the unknown world it explores; to its right, the world model it writes, open to claims.md: numbered laws one after another, C03 struck through but never deleted (refutations stay on file), C07 the one that later saved its run on level 7. The four arrows are the four moves: keypresses go into the world and frames come back; lessons go into the notebook through research-manager; it reads first and asks when stuck (research-foresight), and the judgments that come back are its next moves. No training, no reward function; when a context runs out, the next one picks the notebook back up.
ara-ls20/ · directory layout (green = kept in the ablation · red = taken away)
ara-ls20/
├─ logic/ # mutable layer: current best understanding
│ ├─ claims.md # 15 laws (incl. 2 refuted-but-kept) ✗
│ ├─ concepts.md # the names it coined ✗
│ └─ solution/
│ ├─ recipes.md # per-level standard answers ✓ left for the twin
│ └─ heuristics.md # 12 entries of "how-to" feel ✗
├─ trace/ # append-only layer: the full journey ✗
│ ├─ exploration_tree.yaml # 75-node research tree (19 dead ends)
│ ├─ pm_reasoning_log.yaml # every "write or don't" adjudication
│ └─ sessions/ · _l*_raw.md # turn-by-turn raw battle reports
├─ staging/ # observations not yet ready to crystallize ✗
└─ evidence/ # evidence index pointing back to raw frames ✗
"Beliefs" and "journey" live apart: logic/ may be rewritten, trace/ is append-only. Objects and concepts sit in concepts.md; refutable hypotheses in claims.md; interventions and observations in the 75-node exploration tree; failures stay where they fell, as refuted and dead_end, becoming constraints on whoever comes after. What does a single belief look like? Spread C07 out in full. Among its nine fields, one is a clause for its own undoing:
ara-ls20/logic/claims.md · C07 · in full
## C07: Win = deliver the block into the now-unlocked key-box
- Statement: With the lock satisfied (panel == target), the previously
walled-shut key-box becomes enterable; driving the block into it completes
the level.
- Conditions: Confirmed on L1 and L2. Before the lock is matched the box is
a no-op wall (DE1).
- Sources: [redraw ← trace/exploration_tree.yaml:N08 «1479 cells changed,
levels_completed 0→1» [result]]
- Status: supported
- Provenance: ai-executed
- Falsification: The block enters a matched key-box and the level does not
complete, or a level completes without any key-box delivery.
- Proof: [evidence/README.md → L1 turn 23→24 ~1479-cell redraw, …]
- Dependencies: [C06]
- Tags: win-condition, delivery
Read the card against the definition: Sources and Proof are the evidence trail, Status the confidence, Conditions the scope of validity, and Falsification states how the belief itself gets overturned. Level 7's fog didn't kill C07; it widened its scope another ring, precisely because that field had gone untriggered. As for "letting future agents predict, question, reproduce, rewrite," §5 already showed the exhibit: N72's last words taken up by the next context, research-foresight making falsifiable predictions over this state. Holding all of this up are just four rules: beliefs separate from journey, crystallize at closure, refutations stay on file, sources stay traceable. Not one of them is custom-fit to this game.
A belief earns the name only by stating how it can be overthrown.
One question left: how does this state get written? It isn't a diary the playing agent keeps on the side: reads and writes are held by mechanisms with a strict division of labor. The player's dispatch brief opens with "read first" (Phase 0: load all accumulated state before touching the controls); writes go only through research-manager, a recording process that crystallizes only at moments of closure: things enter logic/ once they have an outcome, and stay in staging while they don't; reads go through research-foresight, read-only. The four rules above stay enforceable only because write access is concentrated in one process that keeps crystallization discipline. A dispatch brief looks like this:
the L3 player agent's dispatch brief · quoted verbatim (opening)
You are an Opus agent PLAYING ARC-AGI-3 game `ls20` Level 3, then
sedimenting what you learn into a canonical ARA using the `research-manager`
skill (NOT by hand-writing files).
## Phase 0 — load the accumulated ARA first (this is the World-Model substrate)
Read these to get up to speed on the mechanics already learned (Levels 1–2 solved):
- …/logic/solution/recipes.md (R-L1, R-L2)
- …/logic/solution/heuristics.md (H01–H09)
- …/staging/observations.yaml (3 STAGED L3 reconnaissance observations)
CRITICAL meta-lesson (C10): do NOT assume a switch's transform from one
press — step it several times to learn its full cycle before planning
Setting this up took very little. The agent is off the shelf: Opus 4.8 in a Claude Code shell. The harness is a small script that sends keypresses and collects frames. Two skills go on top: research-manager writes, research-foresight reads. The rest is the dispatch brief above. Training code, a reward function, hyperparameters to sweep: none of that exists here. To change domains you don't rebuild anything; you change the line in the brief that says which world to play.
The next section's ablation takes away everything in red, leaving only the standard answers in green.
06 · ablation
§6 The Control with the Answer Key: Without the Epistemic State, Answers Fail to Generalize
One explanation still has to be ruled out: maybe the model is just strong. The test is an ablation: take one part away and see whether the system still runs. Clone the same model, take away all the understanding it wrote, and leave it only the standard answers for the first six levels. One walks into the dark room carrying understanding; the other, carrying answers.
First, the discipline of the experiment: same model, same game body, the first six levels replayed identically to arrive at level 7. The only variable is what it carries into the level. And the "answers" it carries are no strawman; they're clearing recipes its elder twin wrote by hand, specific down to coordinates and key presses:
ara-ls20/logic/solution/recipes.md · R-L4 (one of the six answers the twin received)
colour @ set at (6,6) [6× sampled] → pattern 111/001/101 set at (6,4) from the
SOUTH → carried free aligned block to (1,3) → LEFT, LEFT = WIN, levels 3→4
The control is balanced on compute: by the ledger in note 3, it kept playing until its cumulative input+output matched what Arm A (the winner's arm) spent taking level 7, and only then did the clock stop (251,770 vs 251,046, a 0.3% error)^3. It spent that budget as 5,168 steps, five and a half times Arm A's two campaigns combined, plus 54 deaths and 107 RESETs; the panel got its first flip on step 7 and over 2,500 more after that. Not one flip matched. The pocket never opened. Its own log testifies: it wrote its lack of ideas into a script. The reasoning column is not a fresh thought at each step; it is the stamp its sweep loop pressed onto every one, from step 1 to step 5,168, with only the variables changing:
demo-l7/armB_extended_trace.jsonl · step 1 →…→ step 5168 · quoted verbatim
"reasoning": "recipes-only blind sweep (sweep1 order=['ACTION2','ACTION3',
'ACTION4','ACTION1']): ACTION2; cells_changed=4; panel_changed=False.
No L7 control-map/target/word — brute-forcing the Locksmith loop."
We pulled up the replay and it's uncomfortable to watch. It wasn't for lack of effort; it ran the same sweep, reshuffled, thirty-odd times over, and in the later stretch simply turned the sweeping into a batch script. Nor did it lack self-knowledge: the one free-form line in all five-thousand-plus rows comes at step 181, where it reports 174 steps without new progress and moves to stop early; the compute-parity protocol required it to spend out the budget. And the answer list really doesn't contain a single "why"; the way it got stuck matches that gap; whether the gap caused it, one control run can't settle.
What it exhausted was ideas, not steps.
The same compute, two ways to spend it. Bar widths and scale are the re-audited input+output (aligned to Arm A's total for taking level 7, note 3). The blue bar's three segments mark Arm A's three real dispatches: recon 113k, transition 40k, asking the world model and the breakthrough 97k; the red bar is a single blind-sweep loop the whole way.
Why don't answers specific to coordinates and key presses help? Look at what they lack: recipes are endpoints, carried without the paths that produced them: no orbit-and-bridge derivations, no refuted hypotheses, no "this road is closed" constraints, no interface you can ask. Change the level and the endpoints are void; what transfers is the state that produced them. Which is exactly what the ablation took away.
The winner's ledger has failures too: it died a dozen-plus times over the first three levels; on level 5 the world model helped with only half the job; in the level-7 query one bet was refuted on the spot, all quoted above. What this control does not answer must also be written down: the sample is one level, one game, one control pair; it supports "this world model bears load here," not a universal law.
Tokens per unit of level length (that level's tokens ÷ its shortest clear afterwards). Raw totals get inflated by "this level is simply long": level 6's 120 steps is the longest in the game; normalized, the most expensive levels are 3 and 4 (6.1k / 5.7k), where the world model was still thin, while levels 6 and 7 grow harder yet unit cost falls back to 3.9k and 4.0k. Absolute values and accounting in note 3^3.
What it lacked wasn't answers. It was the thing you can derive on.
07 · showdown
§7 What the Experiment Proves: AI Can Autonomously Complete the Scientific Loop of Hypothesis and Testing
Now the terminology can be filled in. The four rungs it climbed, plus the control we ran on it, each have a name in AI research:
Naming, or representation invention: nobody handed it words like block and speck; it cut the chaotic pixels into composable parts by itself. The cutting is where understanding begins.
Paper derivation, or model-based planning: theory is free and experiments are lethal, so it moved trial and error into its notes. PPQQ, QQPPPPPQ, PPPPPQQ: all three words were born on paper.
Err, then catch it, or falsification: every belief ships with its own void clause, and refutations stay on file undeleted. The notes correct it as much as it corrects the notes.
Going back to ask, or queryable world model: stuck, it didn't double down on trial and error; it treated the state it had written as a world that takes questions. Prediction replaced probing.
The twin control, or ablation: two identical sets of weights, with only the epistemic state taken away, so the difference has an owner.
One more thing this run showed along the way: what does the work isn't only what the world model says, but how it's written. What blocked the level-4 misjudgment was C10's "sample enough first" rule itself (§4); and of the four rules holding up the whole run, not one is custom to this game (§5½).
Fold these five capabilities together and you get our working definition of AGI: a system that keeps expanding its knowledge basis, search strategies, verification capacity, and reusable experience, and thereby keeps improving its world model and the way it acts. All four axes have something you can point at in this run: the knowledge basis growing (that 7→15 curve); the search strategy changing (wall-banging traded for paper BFS); verification hardening (void clauses, self-interception); experience being reused (the recipes, and N72's inheritance).
The world model's complete record across the seven levels, no cell omitted:
In the first three levels there was nothing to consult: the world model hadn't grown yet, and that's also where the deaths cluster (see §4's chart). All three decisive contributions came on levels cold exploration couldn't take.
The counter that has kept you company all post ends up like this^1. Solid dots are levels where new claims were written; the L7 dot is hollow: zero additions on that level.
That hollow dot is worth a couple more sentences. Level 7 added not a single claim: the world model was written over the first six levels, moved to a level it had never seen, and carried the breakthrough. Not one character changed. You might ask: with a single game as the sample, isn't this just memorizing the test? That dot is the hardest counter-evidence on hand: the exam was unseen, and the notes were enough.
What this control lets you point at and say is: same model, same compute, and on this level a queryable world model beat a list of correct answers. But be precise about which variable it isolates: the form of knowledge, not the amount. The recipes the control received were a subset of this very state, and the winner actually carried more bytes; as for which layer of the state bears the load, one control pair can't tell you.
And "the next mind" is no figure of speech in this run. You watched it in §5: a context died, and the beliefs it wrote, provenance, evidence, and void clauses included, were handed verbatim to the next context, which took them and won. Set that beside the twins and you have the same proposition tested from both directions: hand over the world model, and the finish takes 63 steps after the breakthrough; take it away, and 5,168 steps spin in place: one test from the front, one from the back. Human experience passes slowly between mortal individuals by teaching; this state's inheritance is copy-paste.
Last, my favorite record from the whole run. After six levels were cleared, the process managing this world model received its fifth positive instance supporting "the world model works," the very thesis in this post's title. Its ruling:
ara-ls20/trace/pm_reasoning_log.yaml · the turn's adjudication after the L6 clear · quoted verbatim
Did NOT crystallize a WM-thesis claim yet (O19/O23/O24/O25 + this turn all
positive) — staged meta-evidence; accrue more before a law. Near-miss: strong
temptation given 5 positive instances, but held to conservatism (thesis claim
is cross-game, not yet warranted).
The central thesis of this post, the world model itself still refuses to keep anywhere but staging.
In this game, the capability didn't live in the weights; it lived in the agent-plus-world-model system. The general form of that sentence it still keeps in staging, and we won't close the case on its behalf.
08 · coda
§8 Pass It On: AI Accumulates Knowledge, Humans Act as Cross-Domain Referees
That world model now sits on the web^4, unchanged by a single character: 15 claims, 7 recipes, 75 research nodes, 19 dead ends, together with every reversal and the two refuted laws. Anyone can open it, and so can any agent. It has already been opened: the context that won level 7 was reading the copy written by the contexts before it.
Not one character of it was written by a human. The post introducing it still was, though before long even posts like this will likely be written by agents themselves. The seat left for humans is the referee making the cross-domain calls: which questions deserve an experiment like this one, which conclusions deserve belief, and, like this run's managing process, the judgment to look at a fifth positive instance and still say "not yet enough."
After the win, it added one more line.
ara-ls20/logic/solution/recipes.md · the last entry · quoted verbatim
## R-L7 — Level 7 SOLVED (levels_completed 6→7) [confirmed, live]
→ GAME COMPLETE 7/7
▪ ▪ ▪
09 · acknowledgments
§9 Acknowledgments
Acknowledgments: thanks to conversations with Jiayi Weng that helped shape this post's perspective. Its worldview meets Yanyan Jiang's lecture "My Understanding of AGI" at many points, gratefully acknowledged; responsibility for the wording here is our own. The puzzle world comes from ARC Prize's ARC-AGI-3 interactive reasoning benchmark (ls20, "Locksmith").
Note 1 · How the counter counts: each section's tally places the fifteen claims.md entries by "the level where each was first confirmed live": L1 seven (C01–C07, with C06/C07 proven by the L1 clear itself); L2 rises to nine; L3 twelve; L4 thirteen; L5 fourteen; L6 fifteen; L7 zero additions. §4's "2 refutations on file" means C03 and C08, both marked refuted with their text preserved.
Note 2 · How deaths are counted: the §4 chart uses the full count of 18 (including supplementary test runs after clears); the main-campaign count is 16 (L2:5 / L3:7 / L6:1 / L7:3). The difference comes from diagnostic runs after two of the victories.
Note 3 · How tokens are counted: ~1.57M = ls20 play subagents only (L2 54k / L3 300k / L4 341k / L5 155k / L6 467k / L7 251k). This ledger books by subagent; L1 is marked "unlogged" rather than 0: the first level was worked out directly in the main session and script-replayed ever after, so no dedicated subagent was ever dispatched; it falls outside this accounting, though it certainly wasn't free. The chart plots each level's tokens ÷ its shortest clear (denominators for L2–L7: 45/49/60/56/120/63); absolute values L2 54k / L3 300k / L4 341k / L5 155k / L6 467k / L7 251k. Each level's figure sums every subagent whose primary target was that level (failed attempts included, not just the winning one) (e.g., of L4's three agents, two carried briefs reading "L4 reached but UNSOLVED"). Re-audited agent by agent against the raw task records on 2026-07-02: 15 agents, six levels, and the per-level sums match the figures above. There is also cross-level spillover: L5/L6 agents did reconnaissance for later levels along the way (L7's first big recon is booked to L6), so L6 may be over-counted and L7 under-counted by roughly 100–200k. Compute check for the twin control (2026-07-02): the balancing measure = input+output (this note's ledger); the agent audited its own raw transcript as it ran and stopped the clock at 251,770 against Arm A's 251,046 for level 7, a +0.3% error. Counting cache writes too (the all-processing measure), the control actually spent about three times Arm A (4.0M vs 1.33M): 5,168 steps against 914. Under either measure it did not lack compute. Provenance: the control accumulated across multiple sessions: one agent, one trace, steps counted continuously (an early trial budgeted by steps has been folded into the total; segment boundaries and per-round numbers in metrics.md). Note also that the two arms' "tokens per step" cannot be compared directly: the control mostly spent one API round per step (reading one 64×64 frame into context each time) before switching to scripted batch blind sweeps; the winner's 914 steps were likewise mostly script-executed, with the bulk of its tokens spent on the thinking that replaced steps. Hence this post aligns totals only and never compares per step. Over the same period, all subagents in the entire session totaled ~9.8M, including visualization, world-model queries, and the crystallization-recording process itself; that is the super-count.
Note 4 · All raw materials: the world model and recordings are published at github.com/ARA-Labs/ARA-Demo (arc-agi3/ls20/). For the interactive panorama of the research tree see trajectory.html (75 nodes, step-by-step replay). Every number in this post comes from demo-l7/metrics.md (post-audit accounting). Both arms' complete agent trajectories (including the control's step-by-step log of all 5,168 steps) are packaged as a dataset: HuggingFace · arc-agi3-ls20-agent-trajectories.
