Home / Historical research

ARC-AGI-3 - spring 2026 research record

ARC-AGI-3 = Abstraction and Reasoning Corpus for Artificial General Intelligence, third generation — an external interactive benchmark from ARC Prize in which an agent must infer each game's mechanics through play; no rules are given. SAGE = Situation-Aware Governance Engine, this lab's on-device cognition kernel. Both terms, and every other one on this page, are defined in the glossary.

Status in September 2026

This page preserves an important SAGE research milestone. It is not current competition positioning and should not be read as evidence that dp-web4 is near the top of the ARC-AGI-3 leaderboard today. Current competition-legal local-model work is well behind the leaders.

What happened

In April 2026, a Phase-1 SAGE/ARC harness (Phase 1: the cloud-model run, frozen in the ARC-SAGE repository linked below; Phase 2 is the local-model follow-on) around Claude Opus 4.6 produced a published 94.85% official ARC Prize action score on the public interactive environments. The action score is efficiency-weighted — it credits solving a level in fewer actions — so it is not a solve rate. The yardstick is human: ARC Prize scores each level by the agent's action count against a human baseline (since 2026-04-14, three days before this run, the median first-time human player per level, with each level's credit capped at 1.15×), and says a 100% score means an agent beats every game as efficiently as humans. Until 2026-09-17 this page said ARC Prize publishes no baseline; its scoring methodology is one. That does not rescue the result from the caveats below: the engine-source affordance is the one that matters. The run completed 175 of 183 levels across 23 of 25 environments: a 92.0% environment-completion rate. Two environments were left unfinished (5 of 9 levels, and 6 of 10).

The score is real and publicly verifiable. The affordances matter just as much as the number: the harness analyzed the games' public engine source and built per-game world-model / solver cartridges. That is outside strict from-observation competition play, which is what “competition-legal” means on this page: an agent that learns each game only by playing it, without the engine source, as the ARC Prize competition requires, rather than a cloud model given engine-level context. The result therefore demonstrates what the model+harness could do with engine-level context and tooling, not blind generalization from observation.

94.85%
Official ARC Prize action score (efficiency-weighted)
23/25
Environments completed (92.0%) — 175/183 levels
2026-04-17
Published milestone

Why it mattered

The lesson we took was methodological rather than positional: changing the structure around an unchanged model can materially change behavior. It is a lesson taken, not yet a result tested. No ablation has run the same model without the harness, so the harness's independent contribution is unknown (see /context). World models, persistent knowledge, explicit skills, prediction and verification became concrete engineering objects rather than prompt ideas.

That work fed into the broader SAGE program. The current target is harder: persistent agents that can formulate hypotheses, run their own experiments, learn from outcomes, retain reusable procedures and act under explicit governance (the SAGE README's phrase; that README names Hestia as the layer that governs local agent authority).

Why it is no longer the headline

The field moved quickly. Stronger models and stronger competition harnesses now outperform this lab's competition-legal local-model work by a wide margin. Continuing to present the April score as a current proof point would create the wrong comparison and obscure the work that has actually advanced since then.

The benchmark remains useful as a laboratory for perception, memory, experimentation, planning and learning. It is no longer the lab's primary credibility claim.

Artifacts

The historical artifact stays public so the framing can evolve without rewriting the record.