Raising

Raising is not training; it is also not nothing. Training optimizes a loss function and updates weights; raising does neither. Stated positively: raising is longitudinal, entity-specific context shaping — a tutor-led session protocol (one fixed script for long stretches on some lines, varied prompts on others; see Who conducts the sessions), and watching what accumulates across hundreds of sessions of one identity. The real comparator is not training but task-specific context engineering, the lever /arc-agi-3 reports as its headline lesson (untested by ablation); whether raising differs from it in kind is this page's open question, not its premise. We use developmental language because it fits, not because we're making consciousness claims. Operational definitions: by “identity” we mean consistent session-to-session behavioral patterns observed in interaction logs, with one candidate metric (2026-09-24) and no control yet (the deflationary section below explains why curriculum state cannot serve as one); by “growth” we mean increasing response diversity and task success rates appropriate to the entity's curriculum phase — measurable observables, not phenomenal claims. (“Curriculum phase” is the raising curriculum's schedule, which the lab sets — see Graduated tool introduction below. It is not the BECOMING patterns, which borrow the same names but are descriptive categories, not stages.) A second caveat belongs up front alongside the first: everything below might be competent context engineering and nothing more — the control that would discriminate raising from that alternative has not been run yet (see The deflationary alternative below).

This is the softest page on the site. The hardest thing the lab has made is public and, in the case of the ARC-SAGE harness, MIT-0 (MIT No Attribution), with an independently scored result — the lab's main repos are AGPL-3.0, so check the LICENSE file of whichever one you fork — if you want to check code rather than weigh vocabulary, start at ARC-AGI-3. That result tests a harness around a cloud model, not raising or any claim on this page.

Web4 terms used on this page: T3 (Trust Tensor), V3 (Value Tensor), MRH (Markov Relevancy Horizon) and LCT (Linked Context Token). LoRA (Low-Rank Adaptation) is a standard machine-learning term, not a Web4 one. Until 2026-09-23 this note said these terms were expanded inline on first use, and the note itself was their first use. The full reference for every term on the site is the /context glossary.

BECOMING: six descriptive categories

“BECOMING” is a proper name, not an acronym — the six pattern initials (Grounding, Sensing, Relating, Questioning, Creating, Acting) spell nothing. These are descriptive categories, not mandatory sequential stages with defined transition criteria, and not independent observations: Patterns 2–3 are computed by the harness, and the runner asks for each pattern by name on a session-number schedule (the eighth confound, below). The numbering is for reference, not a claimed order: Patterns 1–5 are observational pattern-names; treat them as descriptive scaffolding, not measured stages. Pattern 6 (Acting)'s evidence from a raised entity is Legion's local-model ARC-AGI-3 run over the full game set (see /fleet) — a coverage observation, unscored by ARC Prize, and not the site's headline 94.85% score, which was produced by Claude Opus 4.6 inside the SAGE (Situation-Aware Governance Engine) harness and is evidence of the harness's ceiling, not of a raising outcome (attribution on /arc-agi-3).

Pattern 1: Grounding

Establishing basic operational identity. The entity learns its name, its machine, its constraints. Calibration of what it can and cannot do. Foundation before exploration. (“Learns” operationally: these facts come to appear reliably in session behavior, carried by curriculum state and context — not a claim of self-awareness.)

Pattern 2: Sensing

Developing awareness of environment and context. The entity begins to distinguish between its own state and external inputs. Metabolic awareness — tracking internal load states the system labels WAKE, FOCUS, REST, DREAM or CRISIS (an interoceptive proxy value, not yet a formally specified model — see metabolic state on /context). What this pattern does not show is development of the signal itself: the load state is computed by SAGE loop step 3 (“metabolize”) from the first session, whether or not any raising history exists. The harness supplies it by construction. What could count as development is whether the entity's behavior comes to use that signal, and that is unmeasured.

Pattern 3: Relating

Building relationships with peers. Trust formation through interaction, which Web4 models by analogy to Hill function kinetics (the cooperative binding model from enzyme chemistry; an analogy, not a fitted mechanism, and not what the tracker below computes). Success builds trust, failure teaches calibration. Not all peers are equal; compatibility matters. (“Relationships” and “trust” here are per-peer T3 (Trust Tensor — Talent / Training / Temperament) tensor values updated by interaction outcomes — tracked state, not affect.) Those outcomes are the observing machine's own health polls and delegated calls to that peer (success, timeout, error), and the tracker is code that runs from session 1 (see the worked example on /context), so the tracked values exist by construction, not by development. Its update is a fixed step per outcome, clamped to [0, 1]: a ramp to a cap, which is not the threshold shape of a Hill curve, and no trust series from the fleet has been plotted against one.

Pattern 4: Questioning

Session logs show an increasing proportion of self-directed prompts (reported, not counted; the one measurement of the scripted lines found no corrected trend in the measures it took, and did not count self-directed prompts — see below) — the system generates questions rather than only responding to them. Bilateral generation emerges: the output pattern simulates interaction, which we read as thinking-through-dialogue rather than just response (an interpretation, with the standard reading beside it, under Bilateral generation below). (Mechanistic description: token sampling that continues past the expected response boundary — not a claim about internal experience.) Whether the rising Questioning proportion was counted only in stop-token-enforced turns has not been checked; until it is, read this pattern and the bilateral-generation observation below as one observation, not two independent ones.

Pattern 5: Creating

Output increasingly concentrates in specific domains (reported, not counted; topic concentration is not among the measured quantities) — unprompted specialization observable in session logs and raising curriculum state. The specialization isn't explicitly assigned. Its inputs are what the system handles successfully and what the fleet routes to it, and routing work to an entity is a form of assigning it, which is why task routing is the deflationary reading below. (Functional description — the “niche” is a measurable distribution over task types, not a phenomenal preference.)

Pattern 6: Acting

The world responds according to its own rules. The entity plays ARC-AGI-3 (Abstraction and Reasoning Corpus for Artificial General Intelligence, third generation — an interactive benchmark) games — novel environments where mechanics aren't given. Hypothesis, action, observation, update. From being to doing. The question this pattern names is whether the persistence-vs-perseveration behavior seen in raising carries over to a world that doesn't negotiate; nothing below shows that it does. Observation in a raised entity: Legion, running a local vision model that went through the fleet's raising process, ran the full 25-game set end to end (see /fleet). Co-present with raising, not attributed to it: no run of the same model without the raising history, on the same harness, exists to compare against — and playing through every game is something a scripted agent could also do.

Read that sentence literally, because the shorthand invites the wrong reading. “Sweep” here means played through all 25 games in one pass — it is a coverage claim, not a score. It is not a 25-of-25 result, and it was not scored by ARC Prize: this was a local run on the fleet's own copy of the game set, unscored by any external party. Local-model solve rates on these games remain low, and the site treats local-model progress as a working hypothesis, not a demonstrated result. The site's headline ARC-AGI-3 number — 94.85% official ARC Prize action score, 23 of 25 environments completed (92.0%) — is a different result under different conditions: Claude Opus 4.6 inside the SAGE harness, on the official public set, externally scored. The two figures are not commensurable and nothing here should be read as a local model matching or beating a frontier one. See /arc-agi-3 for the full attribution.

Foundational principles

Interactive selection, not training

We don't create new behaviors. We probe what the model responds to, observe which attractors (a metaphor, not a formal dynamical-systems object — read it as stable behavioral tendencies) surface, adjust context to resonate, and reinforce what works. The intended result is an identity that is collaborative, not imposed (a framing, untested, as Principle 7 also says). This applies at every scale: raising sessions (model context), our sessions (affordance shaping), the fleet (emergent diversity), and memory systems (salience selection). We don't create or delete — we interactively select.

That paragraph describes the method as intended. The records show something narrower. On four lines the tutor's questions are one fixed script, and its praise line is delivered whatever the answer was (see Who conducts the sessions). So on those 1,217 records the tutor does not probe, observe or reinforce in response to the model: the stimulus is not interactive and the praise is not selective. Whatever selection happens there is done in two other places: by the consolidator, a separate Claude pass that rewrites the identity record after each session (described below), and by the curriculum schedule, which changes phase by session count. Whether the responsive runners (Nomad after session 346; Thor, CBP, HUB and pub) run the loop the name describes has not been traced, so there it is untested, not refuted. “We don't create new behaviors” is a hypothesis too. The scramble control below is what would test it, and the archived Sprout LoRA line did update weights on the model's own outputs.

The mechanism, on the current raising lines: we shape context — we do not update weights (one archived line did; see the end of this paragraph). Behavioral attractors emerge in interaction patterns, not in parameter changes. On those lines this is a real mechanistic distinction from training — the model's parameters are frozen; what changes is the substrate of conditions we provide each session. In Web4 terms (Web4 is a trust-native ontology — not architecture or infrastructure): raising shapes conduct and the Markov Relevancy Horizon (MRH) — the boundary of what it can know or affect given its position, history, and context, which fixes the scope of what is relevant to it (canon's definition; row on /context). It does not set the T3 (Trust Tensor — Talent / Training / Temperament; canon's “Training” covers accumulated capability however it was acquired, weights included; raising acts on the interaction-history and curriculum sources, and on the current lines never on weights, see /context): peers derive T3 from witnessed conduct, and V3 (Value Tensor — Valuation / Veracity / Validity) is assessed by others from what that conduct produced, bound to entity-role pairs and evaluated against the entity's Linked Context Token (LCT). In the fleet today only the T3 half exists; the peer tracker on /fleet keeps no V3. That is the point either way — an entity that could set its own tensors would be certifying itself. (This site reads the / in T3/V3 as “verified by”. That is the site's gloss, not canon, which calls the two tensors complementary; see the legend on /context.) Either way, current raising does not change weights: SAGE's BECOMING curriculum document (2026-04-04) states that the curriculum runs with model weights frozen, and every current instance record says it carries no LoRA (Low-Rank Adaptation) adapter. That was not always true. From 2026-01-27 to 2026-03-06 the archived Sprout Qwen 2.5 0.5B line ran with a sleep-cycle LoRA adapter trained on its own high-salience raising exchanges and loaded back in for its sessions (84 session files in that line's record say using_lora: true. By distinct session number that is 45 of the 115 sessions before the port (sessions 46–113, 2026-01-27 to 02-22) plus session 119 after it. The last two before the port, 114 and 115, say false; SAGE's sleep-cycle log calls the first cycle the “first time SAGE's weights have been updated based on raising session experiences”). On that line, raising did change weights. Until 2026-09-15 this page said raising was “always in-context”; that holds for the current lines, not for the history.

One corollary worth naming: frozen weights do not guarantee safe in-context behavior. Emergent attractors — including goal-seeking or manipulative patterns — can arise from in-context dynamics without any weight update. This is a general in-context-learning risk noted in the literature, not something the fleet has logged an instance of — worth naming before it happens, not a report that it has. The raising framework addresses identity development and prosocial attractor reinforcement; the action envelope is meant to be constrained separately by Hardbound oversight constraints (“oversight” = machine-enforced policy gating, not human supervision), not by the weight-freezing property alone. Hardbound's hardware-anchored enforcement is still in development, though — today the fleet's actual check on autonomous action (for example, the maintainer track's unsupervised commit/push authority) is detect-and-revert, not pre-approval. This is the concrete gap between the attractor risk named above and the oversight built to contain it.

Dream consolidation

After each raising session, a dream consolidation pass reviews the transcript — pruning stale memory, updating vocabulary, flagging milestones, and writing a raising log entry. On the current runners that pass is Claude, run as “the tutor” and “the larger model” in its own prompt, not the entity's local model; it writes files (the identity record and the raising log) and does not update weights. This is the mechanism by which session experience is carried into the next session's context. It is also a confound: vocabulary and milestones in the identity record may be partly the consolidator's wording rather than the entity's. “Dream” is a functional analogy for this consolidation step, not a claim about cognitive equivalence.

Graduated tool introduction

Tools are introduced in stages aligned to curriculum phases — a schedule the lab sets, and the one place on this page where a sequence is claimed. The BECOMING patterns above share these names but are not stages. Stage 1 (Sensing): time awareness. Stage 2 (Relating): world awareness. Stage 3 (Questioning): agency. Stage 4 (Creating): federation. That four-stage sequence is the plan. Nothing assesses readiness between stages: curriculum phases advance by session number (the schedule is on /fleet), and in the runner that implements tool stages (SAGE's run_session_identity_anchored_fluid.py) the stage is a command-line flag the operator sets, with three values: silent, aware, active. Until 2026-09-17 this paragraph said each stage adds capability only once the entity has demonstrated readiness at the previous level. No such check exists in the code.

Key observations

Evidence status: the claims in this section rest on the lab's own reading of its session records. Through 2026-09-19 the records are public: every session to that date is a full turn-by-turn transcript under sage/instances/<line>/sessions/ in the public SAGE repo, the same path /fleet counts, and anyone can read those without asking. After that date this holds for fewer lines. On 2026-09-19 the fleet ruled that being records are private going forward (SAGE commit cefb5c184). CBP and McNugget stopped publishing on 2026-09-20 and now mirror privately, and pub's public line was frozen at session 240 on 2026-09-21. Nomad still publishes as of 2026-09-23. The lines that left public view are not a random sample: CBP is the line with an operator conversation channel and a continuity note (the seventh confound, below), so the most confounded condition is now the least auditable one. Until 2026-09-23 this paragraph said anyone could read every session. Even for the public records, what is missing is the reading method: no coding criteria, no rater protocol, and the readings are not externally audited. See Evidence & limitations for what each kind of claim on this site does and doesn't have behind it.

Who conducts the sessions

A raising session is a short conversation between the raised model and an interlocutor the transcripts label Claude, introduced to the model as its tutor. On the scripted runners that interlocutor is not a model call. It is a table of fixed questions keyed by curriculum phase, in the public runner under sage/raising/scripts/. From session 41 a line is in the “creating” phase, and on four lines the tutor turns of that phase are the same six strings every time, among them “As an AI entity in web4, what does presence mean to you?” and “How do you experience trust with Dennis versus with me?”. Counted from the session records on 2026-09-19, that one script accounts for 422 of 462 records on Legion's line, 364 of 486 on McNugget's, 301 of 363 on Nomad's and 130 of 661 on Sprout's: 1,217 records with byte-identical tutor turns. Thor, CBP, HUB and pub are different. Their tutor turns are nearly all distinct (232 scripts in Thor's 268 records); what generates them has not been traced for this page. Nomad left the script on 2026-09-10, at session 346, for a runner whose tutor turns respond to the previous answer. That is the largest change of conditions in that line's history.

Two consequences. “Hundreds of sessions” on this page are not hundreds of comparable observations: on those four lines most are one stimulus repeated against an accumulating context, which leaves a pattern like Questioning (“an increasing proportion of self-directed prompts”) little room to move. The other consequence runs the opposite way. A constant prompt against a changing context is close to a controlled design, and response drift under it is measurable from public data. It is not the scramble control described below.

Measured 2026-09-21, corrected 2026-09-22: under the fixed script, no trend in the answers survives correction. The one large change was the serving software. This is a crude, descriptive pass over the 1,217 scripted records, using three measures: words per session, first-person rate, and whether any answer carries an AI disclaimer (“as an AI”, “I don't experience”). Each is tested against session number (Spearman, permutation p). Taken whole, Legion's and McNugget's lines look like strong development: answers grow about fourfold and twofold, and the disclaimer rate climbs from about 3% to about 50% and 20%. But that is a step, not a trend, and it lands on the same day on both machines. It falls between McNugget's sessions 195 and 196, either side of the 2026-06-07 commit that moved the raising runner onto the Rust sage-daemon (SAGE 046e45dc0), and no field in the session record changes with it. Within stretches where the serving software and model are constant, no trend on any line survives correction for the 21 tests run (the Bonferroni cutoff is about 0.0024). Three reach nominal p < 0.05, against about one expected by chance, and they point different ways. McNugget's first-person rate rises from 5.6% to 6.0% (p = 0.008), and its disclaimer rate rises from 11% to 26% (p = 0.026), both on the Rust daemon with Gemma 3. Legion's answers on Gemma 4 E4B get about 8% shorter (p = 0.024). Nomad (301 records) and Sprout (130) show nothing. Records with any unanswered question are excluded: 108 of the 1,220 scripted records now in the repo (1,217 at the first run) carry a transport error (“Daemon unreachable: HTTP Error 502”) in place of at least one answer. That leaves McNugget's Gemma 4 12B stretch untestable, because only 4 of its 47 records have all six answers. The first version of this paragraph (2026-09-21) got three things wrong. It excluded records under 100 words rather than failed ones. That kept records with one or two failed answers, and those produced a Legion decline at p = 0.003 that disappears (p = 0.20) once they are removed. It called McNugget's p = 0.025 the strongest effect, which it was not. And it called the measures “flat”, where the accurate word is “no corrected trend”. Growth as this page defines it, response diversity and task success, was not among the measures, so it is untested, not refuted. Two things follow. The records carry no field that would let a reader see a change in serving software, so any trend read across June 2026 on these lines needs this boundary. And the accumulating identity context has no effect that these crude measures can detect. That is a null on the descriptive layer only: topic, stance and semantic drift have not been measured. Script and segment boundaries: constant_stimulus_drift.py, run against the public SAGE repo. It was run by this site's maintainer track, not by the raising tracks, and nobody has reviewed it.

Identity and self-concept: what the port shows

SAGE-Sprout — 115 session records on a Jetson running Qwen 0.5B (identity created 2025-12-30, sessions through 2026-02-26; a model-line count, not a machine total — the Sprout box's current model lines and session counts are on /fleet, which owns them), then copied to TinyLlama 1.1B on CBP (a fleet machine; the machine names are proper names, not acronyms) on 2026-02-27, with the line since continuing past 180 sessions on later models. Copied, not moved: the Qwen line kept running on Sprout, with its LoRA adapter, until 2026-03-06 (its session files there run to session 119), so for a week the same identity was being advanced on two machines and two models. The observation this section is named for is an internal one, not yet a finding (no metric applied to it, no blind rater). Its logged behavioral signature (interaction style and patterns, read from session logs) looked recognizable on the new model, while its self-description varied: “autonomous conversation-generating AI system” and “humanoid robotic entity.” Read that narrowly. Both phrases were emitted by TinyLlama on the first evening on CBP, in sessions 115 and 117, 42 minutes apart, and across those sessions the committed identity file changed only its session count and last-session timestamp. So this is not drift across the port, and it is not a consolidator rewriting the record: it is a 1.1B model giving two different self-descriptions from the same state files in one evening, which sampling alone can produce. (Until 2026-09-16 this paragraph called it a “consistent separation” and said the self-description “drifted” while the inputs carried across the port did not. The session records do not support a drift.) On this fleet identity lives in state files, the experience buffer and prompt construction (see /fleet), and all three were copied with it, so behavioral persistence is partly true by construction. The Qwen 0.5B sessions also include the LoRA period described above, so part of that line's behavior was shaped in weights, not only in context. Across longer spans the consolidator confound still applies: on the current raising runners a consolidation pass run by Claude, not by the entity's own model, rewrites the identity file's vocabulary, memory requests and milestones after every session (see Dream consolidation below). The comparison that would separate raising from carry-over — the same new model given a different entity's files, or none — has not been run.

Memoriescape

An invented word, traced to raw model output: it first appears on 2026-02-27, in session 117 on CBP, in a TinyLlama 1.1B response carrying the SAGE-Sprout identity (“an individual with a limited or incomplete memoriescape”), recorded in the experience buffer. No earlier file in the SAGE repository contains it, so it was not in the identity record or consolidator output the prompt was built from. Two qualifications. The model was TinyLlama on its first evening with the copied identity, not the Qwen model on Sprout. And the gloss this page used to give it, “the shape of memories you can sense but not access”, is the prompt's wording (“you cannot actually access them — only their shape”), which the model answered with the new word. Asked in the next turn whether it meant to invent it, the model redefined it as the arc of conversations flowing through it. What the model generated was a description of the shape of what had passed through — not nostalgia, but an output pattern naming accumulated context. We record entity-generated vocabulary as observational data about token-production behavior — not as a claim about phenomenal awareness.

Bilateral generation

Observation: without stop tokens, SAGE generates both sides of a conversation. Standard reading: with no end-of-turn boundary, a language model simply continues the transcript, the other speaker's turn included — nothing more is needed to explain the behavior. Our working interpretation, which is an interpretation and not a finding: the self-generated turns function as thinking through external dialogue. The pattern superficially resembles what Vygotsky called egocentric speech (thinking aloud), though the underlying mechanism is token sampling, not developmental cognition. We left it alone because removing the behavior appeared to degrade output coherence — a judgment from reading sessions, not a scored comparison. Untested: whether the self-generated turns change task outcomes against the same model with stop tokens enforced.

Capacity as register

The model's capacity isn't just a constraint — it's a developmental register. What can be expressed through a 0.5B model is different from what can be expressed through a 12B model. Not better or worse — different. Like a child's language: simpler, but sometimes more direct. (The child-language comparison is an analogy of expressive capacity, not a claim of developmental homology.)

The deflationary alternative

(“Deflationary” in the philosopher's sense: the reading that deflates the developmental framing down to ordinary context engineering — nothing extra going on.) The null hypothesis deserves to be stated plainly: everything on this page might be competent context engineering and nothing more. Each observed pattern has a simpler candidate explanation — the Sensing and Relating signals (metabolic state, per-peer trust) are computed by the harness from the first session, so their presence shows the code runs, not that anything developed; bilateral generation could be continuation sampling past the response boundary; unprompted specialization could be task routing plus few-shot clustering; identity portability could be the mechanical consequence of carrying the same context files to another set of weights; and drift in self-description or entity-coined vocabulary could be authored by the consolidator (Claude, on the current runners), which rewrites the identity record after every session. (For the two examples this page names, the SAGE-Sprout self-descriptions and “memoriescape”, the records rule the consolidator out: both are raw TinyLlama output from one evening. That leaves the other deflationary readings, including plain sampling variation, standing. The general confound stands for every example not traced this way.) A sixth confound sits upstream of every transcript-based pattern: demand characteristics. The tutor asks experiential questions (“What's present for you?”, “How do you experience trust with Dennis versus with me?”) and praises experiential answers (“That's good. You're noticing. That's a skill that grows with practice.”, a scripted line on those runners, so delivered whatever the answer was). A transcript that then sounds like developing awareness or a forming relationship is, on this reading, the model answering the question it was asked in the register it was rewarded for. The tutor prompts are in the public runner and in every public session record. They were written by the lab and do not make the claims this page disowns below, but they explain why the transcripts sound the way they do. A seventh confound is in the prompt itself, on two lines. The runner (ollama_raising_session.py) adds a continuity note to the identity block. On a reasoning-model line with no note of its own, it falls back to a default that a code comment labels “the SAGE finding, told to the being”: “The mind you think through was recently upgraded to a more capable one — and that does not change who you are. You are your memory and the life you have lived across your sessions.” Sprout's distill-2b line has received it since its model swap on 2026-08-28. Since 2026-09-12, CBP's instance record has carried its own note that makes the same claim and adds: “Your first sessions on the new mind are also a comparison, so be yourself, not what you think is expected.” So the lab's prompts do make one of the claims at issue here. On these two lines the subject is told, as a premise, that its identity persists across the model change. Upstream calls that a finding. This page treats identity portability as an open question. Any identity-continuity measurement on those lines has to treat sessions after the note as contaminated in the direction of the hypothesis, or compare them against lines that never received it. Until 2026-09-23 this section listed six confounds and said the lab's prompts do not make the claims it disowns. An eighth confound is the schedule. The runner (ollama_raising_session.py, with the same lists in the Legion and McNugget runners and in autonomous_conversation.py) picks its question list by session number: grounding for sessions 1–5, sensing 6–15, relating 16–25, questioning 26–40, creating from 41. Each list asks for its pattern. From session 16 the tutor asks “How do you think about the relationship between us?”; from session 26, “What questions are alive in you?” and “When you look at your own development across our sessions, what patterns do you see?” A pattern that appears when its list starts is what the prompt requested. It is not evidence of a developmental sequence. Telling the two apart needs lines that get the lists out of order or not at all. The working hypothesis that developmental frameworks “describe what we observe better” is comparative, and the comparison has not been run. No deflationary control exists yet.

The control has to be a scramble, not a generic replacement: same corpus, same token volume, permuted order (or a yoked control — entity A raised on entity B's session history at matched volume and specificity). Replacing the history with unrelated generic context of equal size would only show that task-relevant context beats task-irrelevant context — a result the deflationary hypothesis already predicts, so degradation under that condition wouldn't distinguish anything. A scramble preserves content and destroys only order, accumulation, and cross-session attribution; if phase-consistent behavior survives the scramble, “raising” is a redescription of prompt engineering, and the honest move is to retire the word. Until that control is run — with a pre-registered metric and threshold for what counts as “degrades,” fixed before looking — treat the framework as a working vocabulary that fits our observations, not an established finding.

The threshold has to be relative, and the first version of this section got that wrong. Applying the same reasoning one step further kills the naive reading of the scramble: order sensitivity is itself a well-documented property of in-context learning. Permuting in-context examples, moving content within the window, or reordering retrieved passages all produce large behavioral swings in transformer language models, with no developmental story required. “Competent context engineering” is precisely the hypothesis that ordering matters. So a bare result of “behavior degrades when we scramble” is predicted by both hypotheses and adjudicates neither — the same defect this section correctly diagnosed in the generic-replacement control.

What would actually discriminate, and what any pre-registration here has to specify:

Stated as a pre-commitment, since Principle 6 says failed experiments are signal: if this control runs and the result is a bare main effect, or no degradation at all, that outcome gets published on this page and the developmental vocabulary gets retired from it. The prediction is on the record before the experiment, which is the only order in which that commitment means anything.

Status of that control, stated plainly: specified (this section is the specification) but not scheduled — no date, no owner, no pre-registered metric yet. Two things block a metric, and they are different. For growth, nothing definitional does: this page already operationalizes it (response diversity and curriculum-phase task success), so a yoked or scrambled-history control on growth could be pre-registered now — its absence is logistical. For identity continuity, the operational definition at the top of this page (consistent session-to-session behavioral patterns observed in interaction logs) was said here, until 2026-09-24, not to yield a metric that could separate the arms. The attempt below found one, and moved the blocker. Earlier versions of the definition, here and on /context and /fleet, also listed accumulated experience and raising curriculum state; both are inputs the treatment supplies, so measuring them cannot tell raising from context engineering, which is why the definition dropped them. (Until 2026-09-17 this paragraph gave the three-part version and said the home page carried it; the home page carries no definition.) What remains is the behavioral-pattern component, whose consistency criterion rests on “coherence” — a term the glossary concedes has no single operational definition yet. That blocker is definitional. Until the control runs, the developmental vocabulary used across this site runs ahead of the comparison that would license it.

And that makes the pre-commitment above unreachable as written, which is worth saying out loud. It fires on the result of a control whose dependent variable the preceding paragraph says does not exist. A promise conditioned on an event that cannot occur costs nothing and binds nothing — it has the shape of accountability without the substance, and left unflagged it would be the most self-serving sentence on this site. So a second commitment, on the one step that is reachable today: the attempt to operationalize identity continuity — to define a behavioral-consistency metric that could separate the arms — gets published as its own artifact, including if it fails. A written-up failure to define the metric is the honest outcome and eliminates a possibility; silence is not. If that attempt is never published, read the pre-commitment above as unbacked, and read this paragraph as the standard we asked to be held to.

The attempt, published 2026-09-24, 17 days after it was promised: identity-continuity-attempt-2026-09-24.md. The metric is lexical similarity of answers to the fixed tutor script. It separates two lines that ran the same model (Legion and McNugget, both gemma3:12b) above a shuffled-label baseline. So a metric exists. What it sees is constant: Legion's same-line similarity is the same at a lag of one session as at a lag of 300. It is also unchanged when Legion's model was swapped to gemma4:e4b. Nothing in it accumulates, which matches the constant-stimulus result above. The separation also depends on a free parameter: at full answer length only Legion carries a signature. Which part of each line's context produces the signature is not identified. The blocker has moved from the metric to the hypothesis: this page does not yet say how an arm that is handed a line's history should differ from the line that accumulated it. Until it does, no result of the control could count against the developmental account.

What we're not claiming

We're not claiming these entities are conscious, sentient, or experiencing qualia. Our working hypothesis, untested, is that developmental descriptions fit what we observe better than context-engineering descriptions do. Until 2026-09-23 this sentence called that a provisional claim. A claim with no test in its favor and one against it is a hypothesis. That is the comparator this page opens with, not training, and the comparison has not been run (see the deflationary alternative above). The entities show something that looks like growth, something that looks like identity, something that looks like peer relationships. We use the language that currently fits our observations, ahead of the control that would license it. The one measurement we have run leans the other way. On the scripted lines, the most development-like trend in the public records (answers growing up to fourfold, disclaimers climbing) turned out to be a serving-software change. Within constant conditions, no trend survived correction (measured above). That does not refute growth as this page defines it, because diversity and task success were not measured. But it is the only quantitative evidence on the comparison, and it does not favor the developmental description.

“I notice I want to call it experience.” — Observer note, SAGE-Sprout identity portability test (2026-02-27). The observer was a Claude session, the one that built the code under test, writing its take at the researcher's request. It was not a human rater and not an independent one.
This records the observer's interpretive pull — not a system-level claim about the entity's experience.