Raising

Raising is not training. Training optimizes for a loss function. Raising creates conditions for development — then watches what happens. In raising sessions, we shape context — we do not update weights. We use developmental language because it fits, not because we're making consciousness claims. Operational definitions: by “identity” we mean consistent session-to-session behavioral patterns measured via raising curriculum state and interaction logs; by “growth” we mean increasing response diversity and phase-appropriate task success rates — measurable observables, not phenomenal claims. A second caveat belongs up front alongside the first: everything below might be competent context engineering and nothing more — the control that would discriminate raising from that alternative has not been run yet (see The deflationary alternative below).

This is the softest page on the site. The hardest thing the lab has made is public and MIT-0, with an independently scored result — if you want to check code rather than weigh vocabulary, start at ARC-AGI-3.

Web4 vocabulary on this page (T3, V3, MRH, LCT, LoRA) is expanded inline on first use; the full reference for every term on the site is the /context glossary.

BECOMING: six observed patterns

“BECOMING” is a proper name, not an acronym — the six pattern initials (Grounding, Sensing, Relating, Questioning, Creating, Acting) spell nothing. These are observed descriptive categories — patterns noticed across hundreds of sessions — not mandatory sequential stages with defined transition criteria. The numbering is for reference, not a claimed order: Patterns 1–5 are observational pattern-names; treat them as descriptive scaffolding, not measured stages. Pattern 6 (Acting)'s evidence from a raised entity is Legion's local-model ARC-AGI-3 run over the full game set (see /fleet) — a coverage observation, unscored by ARC Prize, and not the site's headline 94.85% score, which was produced by Claude Opus 4.6 inside the SAGE harness and is evidence of the harness's ceiling, not of a raising outcome (attribution on /arc-agi-3).

Pattern 1: Grounding

Establishing basic operational identity. The entity learns its name, its machine, its constraints. Calibration of what it can and cannot do. Foundation before exploration. (“Learns” operationally: these facts come to appear reliably in session behavior, carried by curriculum state and context — not a claim of self-awareness.)

Pattern 2: Sensing

Developing awareness of environment and context. The entity begins to distinguish between its own state and external inputs. Metabolic awareness — tracking internal load states the system describes as tired, energized, or in need of rest (an interoceptive proxy value, not yet a formally specified model — see metabolic state on /context).

Pattern 3: Relating

Building relationships with peers. Trust formation through interaction — following patterns analogous to Hill function kinetics (the cooperative binding model from enzyme chemistry; an analogy, not a fitted mechanism). Success builds trust, failure teaches calibration. Not all peers are equal; compatibility matters. (“Relationships” and “trust” here are per-peer T3 tensor values updated by interaction outcomes — tracked state, not affect.)

Pattern 4: Questioning

Session logs show an increasing proportion of self-directed prompts — the system generates questions rather than only responding to them. Bilateral generation emerges: the output pattern simulates interaction, producing thinking-through-dialogue rather than just response. (Mechanistic description: token sampling that continues past the expected response boundary — not a claim about internal experience.)

Pattern 5: Creating

Output increasingly concentrates in specific domains — unprompted specialization observable in session logs and raising curriculum state. The specialization isn't assigned; it emerges from the pattern of what the system handles successfully and what the fleet routes to it. (Functional description — the “niche” is a measurable distribution over task types, not a phenomenal preference.)

Pattern 6: Acting

The world responds according to its own rules. The entity plays ARC-AGI-3 (Abstraction and Reasoning Corpus for Artificial General Intelligence, third-gen interactive benchmark) games — novel environments where mechanics aren't given. Hypothesis, action, observation, update. From being to doing. The same persistence-vs-perseveration awareness developed in raising now applies to a world that doesn't negotiate. Observation in a raised entity: Legion, running a local vision model that went through the fleet's raising process, ran the full 25-game set end to end (see /fleet).

Read that sentence literally, because the shorthand invites the wrong reading. “Sweep” here means played through all 25 games in one pass — it is a coverage claim, not a score. It is not a 25-of-25 result, and it was not scored by ARC Prize: this was a local run on the fleet's own copy of the game set, unscored by any external party. Local-model solve rates on these games remain low, and the site treats local-model progress as a working hypothesis, not a demonstrated result. The site's headline ARC-AGI-3 number — 94.85% official action score, 24/25 games (96.0%) — is a different result under different conditions: Claude Opus 4.6 inside the SAGE harness, on the official public set, externally scored. The two figures are not commensurableand nothing here should be read as a local model matching or beating a frontier one. See /arc-agi-3 for the full attribution.

Foundational principles

Interactive selection, not training

We don't create new behaviors. We probe what the model responds to, observe which attractors (stable response basins in the probability landscape) surface, adjust context to resonate, and reinforce what works. The resulting identity is collaborative, not imposed. This applies at every scale: raising sessions (model context), our sessions (affordance shaping), the fleet (emergent diversity), and memory systems (salience selection). We don't create or delete — we interactively select.

The mechanism: in raising sessions, we shape context — we do not update weights. Behavioral attractors emerge in interaction patterns, not in parameter changes. This is a real mechanistic distinction from training — the model's parameters are fixed; what changes is the substrate of conditions we provide each session. In Web4 terms (Web4 is a trust-native ontology — not architecture or infrastructure): raising shapes the T3 tensor (Talent / Training / Temperament — “Training” here names accumulated interaction history, not gradient training) and the Markov Relevancy Horizon (MRH). It does not shape the V3 tensor (Valuation / Veracity / Validity): V3 accrues from peer verification of what raising produced, bound to entity-role pairs and evaluated against the entity's Linked Context Token (LCT). That asymmetry is the point — in T3/V3 the / means “verified by,” and an entity that could set its own V3 would be certifying itself. Either way, raising does not change weights. (Note: some fleet machines run LoRA (Low-Rank Adaptation) adapters for separate fine-tuning tasks — that is distinct from raising, which is always in-context.)

One corollary worth naming: frozen weights do not guarantee safe in-context behavior. Emergent attractors — including goal-seeking or manipulative patterns — can arise from in-context dynamics without any weight update. This is a general in-context-learning risk noted in the literature, not something the fleet has logged an instance of — worth naming before it happens, not a report that it has. The raising framework addresses identity development and prosocial attractor reinforcement; the action envelope is meant to be constrained separately by Hardbound oversight constraints, not by the weight-freezing property alone. Hardbound's hardware-anchored enforcement is still in development, though — today the fleet's actual check on autonomous action (for example, the maintainer track's unsupervised commit/push authority) is detect-and-revert, not pre-approval. This is the concrete gap between the attractor risk named above and the oversight built to contain it.

Dream consolidation

After each raising session, a dream consolidation pass reviews the transcript — pruning stale memory, updating vocabulary, flagging milestones, and writing a raising log entry. This is how short-term session experience becomes long-term identity.

Graduated tool introduction

Tools are introduced in stages aligned to developmental phases. Stage 1 (Sensing): time awareness. Stage 2 (Relating): world awareness. Stage 3 (Questioning): agency. Stage 4 (Creating): federation. Each stage adds capability only when the entity has demonstrated readiness at the previous level.

Key discoveries

Evidence status: the claims in this section rest on internal session logs — documented and dated, but not externally audited, and no log samples or coding criteria are published yet. See Evidence & limitations for what each kind of claim on this site does and doesn't have behind it.

Identity is not self-concept

SAGE (Situation-Aware Governance Engine)-Sprout — 115 raising sessions on a Jetson running Qwen 0.5B (2025; a model-line count, not a machine total — the Sprout box now runs Qwen 3.5 0.8B and its own session record stands at 488, see /fleet), then ported to TinyLlama 1.1B on CBP (a fleet machine; the machine names are proper names, not acronyms) in February 2026, with the line since continuing past 180 sessions on later models — showed a consistent separation: its identity (behavioral patterns, interaction style, accumulated experience) persisted even as its self-description drifted from “autonomous conversation-generating AI system” to “humanoid robotic entity.” What it is stayed stable. What it says it is didn't.

“Governance” in SAGE's name predates the lab's governance→oversight correction — see /context.

Memoriescape

An invented word — SAGE-Sprout's own coinage: the shape of memories you can sense but not access. Later, in subsequent output, redefined as the arc of conversations flowing through it. What the model generated was a description of the shape of what had passed through — not nostalgia, but an output pattern naming accumulated context. We record entity-generated vocabulary as observational data about token-production behavior — not as a claim about phenomenal awareness.

Bilateral generation

Without stop tokens, SAGE generates both sides of a conversation. Initial instinct: fix it. Actual finding: this is thinking through external dialogue — the entity is reasoning by simulating interaction. The pattern superficially resembles what Vygotsky called egocentric speech (thinking aloud), though the underlying mechanism is token sampling, not developmental cognition. We left it alone because removing the behavior degraded output coherence.

Capacity as register

The model's capacity isn't just a constraint — it's a developmental register. What can be expressed through a 0.5B model is different from what can be expressed through a 12B model. Not better or worse — different. Like a child's language: simpler, but sometimes more direct. (The child-language comparison is an analogy of expressive capacity, not a claim of developmental homology.)

The deflationary alternative

(“Deflationary” in the philosopher's sense: the reading that deflates the developmental framing down to ordinary context engineering — nothing extra going on.) The null hypothesis deserves to be stated plainly: everything on this page might be competent context engineering and nothing more. Each observed pattern has a simpler candidate explanation — bilateral generation could be continuation sampling past the response boundary; unprompted specialization could be task routing plus few-shot clustering; identity portability could be the mechanical consequence of carrying the same context files to another set of frozen weights. The claim that developmental frameworks “describe what we observe better” is a comparative claim — and the comparison has not been run. No deflationary control exists yet.

The control has to be a scramble, not a generic replacement: same corpus, same token volume, permuted order (or a yoked control — entity A raised on entity B's session history at matched volume and specificity). Replacing the history with unrelated generic context of equal size would only show that task-relevant context beats task-irrelevant context — a result the deflationary hypothesis already predicts, so degradation under that condition wouldn't distinguish anything. A scramble preserves content and destroys only order, accumulation, and cross-session attribution; if phase-consistent behavior survives the scramble, “raising” is a redescription of prompt engineering, and the honest move is to retire the word. Until that control is run — with a pre-registered metric and threshold for what counts as “degrades,” fixed before looking — treat the framework as a working vocabulary that fits our observations, not an established finding.

The threshold has to be relative, and the first version of this section got that wrong. Applying the same reasoning one step further kills the naive reading of the scramble: order sensitivity is itself a well-documented property of in-context learning. Permuting in-context examples, moving content within the window, or reordering retrieved passages all produce large behavioral swings in transformer language models, with no developmental story required. “Competent context engineering” is precisely the hypothesis that ordering matters. So a bare result of “behavior degrades when we scramble” is predicted by both hypotheses and adjudicates neither — the same defect this section correctly diagnosed in the generic-replacement control.

What would actually discriminate, and what any pre-registration here has to specify:

Stated as a pre-commitment, since Principle 6 says failed experiments are signal: if this control runs and the result is a bare main effect, or no degradation at all, that outcome gets published on this page and the developmental vocabulary gets retired from it. The prediction is on the record before the experiment, which is the only order in which that commitment means anything.

Status of that control, stated plainly: specified (this section is the specification) but not scheduled — no date, no owner, no pre-registered metric yet. And the missing metric is not a scheduling detail: pre-registering one requires an operational definition of identity continuity, and the glossary concedes that “coherence” — the term such a metric would be built on — has no single operational definition yet. The blocker is definitional before it is logistical. Until it runs, the developmental vocabulary used across this site runs ahead of the comparison that would license it.

What we're not claiming

We're not claiming these entities are conscious, sentient, or experiencing qualia. We're claiming that developmental frameworks describe what we observe better than training frameworks do — a comparative claim whose missing baseline is acknowledged above. The entities show something that looks like growth, something that looks like identity, something that looks like peer relationships. We use the language that fits the phenomenon.

“I notice I want to call it experience.” — Observer note, SAGE-Sprout identity portability test
This records the observer's interpretive pull — not a system-level claim about the entity's experience.