Skip to content

Self-Modification

Without self-modification, every routing rule, decay rate, confidence threshold, and condensation trigger must be manually configured and manually updated. The system can observe its own operation — eval traces record every classification decision, every query hit and miss, every contradiction — but it cannot act on those observations to improve itself. A pattern that holds with 100% consistency across 200 traces still requires a human to notice it, write a rule, and deploy the change.

This is the difference between a static configuration and an adaptive system. Static configurations degrade under distributional drift (CDR Theorem 3) and can only be renewed by human intervention. Self-modification closes the loop: the system observes its own performance, formulates hypotheses about stable patterns, validates them against historical evidence, compiles validated patterns into explicit rules, monitors those rules for degradation, and renews or culls them when they stop working. The human sets boundaries — what is modifiable, what evidence thresholds are required, what requires approval — but the system does the work of adapting within those boundaries.

Without this capability, the system accumulates two forms of debt. First, optimization debt: stable patterns that could be compiled into cheap rules remain as expensive inference calls because no one has written the rule. Second, staleness debt: rules that were correct when written but have drifted remain in place because no one has checked them. Self-modification addresses both by making the system responsible for its own configuration lifecycle.

PhaseTimescaleDescription
Pattern observationDays to weeksEval traces accumulate; statistical significance requires volume
Hypothesis formationOn evidence thresholdNot scheduled — triggered when trace analysis detects a candidate pattern
Counterfactual evaluationHours (per hypothesis)Replay historical traces against proposed rule to measure improvement
Sandbox executionMinutes (per replay)Isolated replay of interactions against modified or historical state (§7.9)
CompilationMinutesWrite the validated rule into conventions with provenance and expiry
MonitoringContinuousEvery operation that touches a compiled rule updates its accuracy metrics
Task-type aggregationPer-interaction (append), periodic (aggregate)Performance register updated on each outcome, aggregated daily/per-session (§7.10)
Comparative evaluationDays to weeks (post-promotion)Before/after measurement once post-change traces accumulate (§7.5 Stage 6)
Renewal / cullingWeeks to months (at τ*)CDR Theorem 3 determines the optimal renewal interval per rule
Colony-level evolutionDays to weeksCross-node proposal, validation, consensus, and promotion

Self-modification draws on established concepts:

  • Online learning. The system learns from its own operational data stream. Compiled rules are online-learned classifiers validated against historical traces.
  • A/B testing (counterfactual evaluation). Before a pattern is compiled, the system replays historical traces with and without the proposed rule to measure whether the rule improves outcomes. This is offline counterfactual evaluation — the same principle as A/B testing applied to historical data rather than live traffic.
  • Configuration management. Compiled rules are configuration that the system writes and maintains. The conventions document is a configuration file with version history, provenance, and a computable expiry on every entry.
  • Hot-path optimization. The inference optimization ladder (Section 4) is hot-path compilation: frequently executed inference patterns are compiled into cheaper representations. Self-modification is the mechanism that performs this compilation.
  • Feature flags with automatic rollback. Compiled rules are analogous to feature flags — they change system behavior, they carry monitoring, and they can be rolled back (demoted to pure inference) when they degrade.
  • Staging environment / transaction ROLLBACK (sandbox execution). The sandbox (Section 7.9) provides an isolated context for evaluating hypotheses without affecting real state — the same principle as a staging environment for testing deployments or a database transaction that is rolled back after observation. Also: fork() in process management.
  • SLIs/SLOs per endpoint (task-type performance register). The task-type register (Section 7.10) tracks per-domain performance metrics the same way service-level indicators track per-endpoint health. Each task type is an endpoint with its own outcome rate, trend, and throughput.
  • A/B test result (comparative evaluation). The comparative evaluation record (Section 7.5 Stage 6) captures before/after performance measurement across a change boundary — the same structure as a canary deployment result or an A/B test conclusion with statistical significance.

The protocol distinguishes three classes of system parameters based on who may change them and under what conditions.

Hardwired parameters MUST NOT be modified by the system under any circumstances. They are structural invariants that the protocol’s safety properties depend on. A conforming node that modifies a hardwired parameter is non-conforming.

ParameterRationale
Provenance requirementEvery memory entry MUST trace to a verifiable source. Without provenance, the system cannot audit its own state, cannot detect fabricated entries, and cannot perform renewal (CDR Theorem 3 requires knowing what a compression was derived from).
Eval trace formatThe measurement system MUST be fixed. If the system can modify how it measures its own performance, it can game its own metrics. Eval traces are the ground truth for compilation decisions — modifying the format invalidates all historical evidence.
Conflict resolution hierarchyThe escalation path (recency → confidence → domain owner → human) MUST be immutable. If the system can modify its own conflict resolution, it can suppress contradictions that would otherwise trigger renewal or fault detection.
Changelog requirementEvery modification to conventions, compiled rules, or colony-wide state MUST be logged. The changelog is the audit trail for the self-modification system itself. A system that can suppress its own changelog is unauditable.
CDR structural propertiesThe three theorems (irreversible compression, temporal partitioning, finite optimal lifespan) are mathematical results, not design choices. They constrain the system; they are not parameters of the system.

7.4.2 Evolvable (System-Proposed, Validated)

Section titled “7.4.2 Evolvable (System-Proposed, Validated)”

Evolvable parameters MAY be modified by the system through the learned behavior lifecycle (Section 7.5). Every modification MUST include:

  • The evidence that motivated the change (eval trace references)
  • The counterfactual evaluation result (improvement metric)
  • The previous value and the new value
  • A computed τ* for the new value (when it should be reviewed)
  • A rollback path (the previous value is preserved in the changelog)
ParameterConstraints
Decay rates per regionMUST remain within human-configured bounds. A decay rate of 0 (never decay) or ∞ (instant decay) violates CDR Theorem 3. Bounds are specified in the steering configuration.
Confidence thresholdsMUST remain within [0.1, 0.95] unless human override. A threshold of 0 admits everything; a threshold of 1 admits nothing. Both are degenerate.
Condensation triggersThe minimum observation count and time window for promoting observations to patterns. MUST require at least 2 observations across at least 2 distinct time periods (prevents single-event overfitting).
Routing rulesCompiled classification patterns (e.g., source:meeting-transcript → region:operational). These are the primary target for self-modification. See Section 7.5 for lifecycle.
Renewal intervalsThe review cadence for entries in each region. MUST NOT exceed the region’s maximum staleness threshold (a structural entry reviewed less than once per quarter violates the staleness protocol).
Colony steering parametersParameters such as max_colony_size, dormancy_threshold, and maintenance_budget_fraction MAY be proposed for adjustment by the system. Changes to steering parameters MUST require human approval regardless of evidence strength.

7.4.3 Learnable (System-Discovered, Compiled)

Section titled “7.4.3 Learnable (System-Discovered, Compiled)”

Learnable parameters are not configured — they are discovered through operational experience and compiled into explicit rules. They begin as implicit patterns in eval traces and are promoted through the lifecycle in Section 7.5.

ParameterDescription
Classification shortcutsStable mappings from input characteristics to region assignment. Example: if items sourced from meeting transcripts are classified as operational with >95% consistency across >50 traces, the mapping is compiled into a routing rule.
Query routing patternsStable mappings from query type to optimal memory retrieval path. Example: if questions about “who decided X” consistently resolve from structural/architecture.md, that routing is compiled.
Contradiction detection heuristicsCommon conflict patterns that the system learns to detect early. Example: if entries about project timelines contradict each other at a rate significantly above baseline, a specific contradiction check for timeline entries is compiled.
Condensation templatesCommon shapes for condensed patterns. Example: if the system repeatedly condenses meeting observations into [person] is responsible for [domain] structures, that template is compiled so future condensations can match the pattern rather than re-derive the structure.

Every compiled rule follows an eight-stage lifecycle. The lifecycle is CDR-enforced: stages are evidence-gated (not time-gated), and every rule that reaches COMPILATION has a computable optimal lifespan.

EXPERIENCE → HYPOTHESIS → ANALYSIS → EVIDENCE → COMPILATION → PROMOTION → MONITORING → RENEWAL/CULL

The node accumulates eval traces through normal operation. Traces record every classification decision, query result, contradiction check, and condensation event with their outcomes.

Requirements:

  • Traces MUST be append-only (the eval record is immutable)
  • Traces MUST include both the decision and the outcome (e.g., “classified as operational” + “user later confirmed useful” or “never accessed again”)
  • No minimum trace volume to enter this stage — accumulation is passive

The node detects a candidate pattern in its traces. Hypothesis formation MAY be triggered by:

  • Scheduled trace analysis (part of the periodic convention_review maintenance task)
  • A human directive (e.g., “check if meeting notes are always operational”)
  • A threshold crossing (e.g., the same classification has been made >N times with >P% consistency)

Hypothesis format:

hypothesis:
id: HYP-{UUID-SHORT}
created: ISO-8601
node_id: originating node
type: classification_shortcut | query_routing | contradiction_heuristic | condensation_template
statement: "Items from source type S are classified as region R with consistency > P%"
supporting_traces: [TRACE-001, TRACE-002, ...] # minimum 10
minimum_evidence:
trace_count: 50 # configurable, default 50
consistency: 0.90 # configurable, default 0.90
time_span: 14d # configurable, default 14 days
status: PENDING_ANALYSIS

The node evaluates the hypothesis against its full trace history using counterfactual evaluation. The question is: if this rule had been in effect during the historical period, would outcomes have improved?

Procedure:

  1. Select all traces within the hypothesis’s domain (e.g., all meeting-transcript classifications)
  2. For each trace, compute the decision the proposed rule would have made
  3. Compare against the actual decision and its recorded outcome
  4. Measure: accuracy (does the rule match the correct decision?), coverage (what fraction of relevant decisions does the rule apply to?), and false positive rate (does the rule misclassify cases it should not touch?)

Requirements:

  • Analysis MUST use only historical data. The system MUST NOT deploy the rule experimentally before validation.
  • Analysis MUST report both improvement and regression cases. A rule that improves 90% of cases but causes 10% regressions is a different decision than one that improves 95% with 0% regressions.
  • Analysis results MUST be persisted with the hypothesis record.

State Hypotheses (Generalized Counterfactual)

Section titled “State Hypotheses (Generalized Counterfactual)”

In addition to rule hypotheses (described above), the analysis stage accepts state hypotheses: counterfactuals about memory contents rather than processing rules. A state hypothesis asks “if memory state had been different, would outcomes have changed?” — for example, “what if this entry had higher confidence?”, “what if these entries had been condensed earlier?”, “what if this entry had been in structural instead of operational?”

Justification: Rule-only counterfactuals test whether a proposed rule improves decisions. State counterfactuals test whether different memory organization, content, or confidence levels would have produced better task outcomes. This is required for the system to reason about its own memory quality, not just its own rule quality. State hypotheses are typically triggered by consolidation replay findings (§6c.4) or by gap detection in the evolution path.

DS analogue: Chaos engineering with controlled state injection. Mutation testing — deliberately modifying system state and measuring whether replayed interactions detect the difference.

State hypotheses specify a diff against a node state snapshot (§6c.5) and use sandbox execution (§3.3.5) for evaluation:

state_hypothesis:
id: HYP-{UUID-SHORT}
type: state_counterfactual
base_snapshot: SNAP-{UUID-SHORT}
state_diff:
entries_added: list[entry]
entries_removed: list[entry_id]
entries_modified: list[entry_diff]
confidence_adjustments: list[{entry_id, new_confidence}]
region_reassignments: list[{entry_id, new_region}]
replay_traces: list[EXEC-{UUID-SHORT}]
expected_effect: "Outcome quality for [task type] should improve"
analysis_output:
traces_replayed: integer
outcomes_improved: integer
outcomes_degraded: integer
outcomes_unchanged: integer
net_quality_delta: float

Requirements:

  • State hypotheses MUST specify a base_snapshot — the node state to modify.
  • The state_diff MUST be applied within a sandbox (§3.3.5). The sandbox isolation guarantee ensures no mutation to real node state.
  • The replay_traces MUST reference existing execution traces (§3.6 execution event class). These are the interactions replayed against the modified state.
  • Analysis results follow the same evidence format (Stage 4) and the same compilation gate. A state hypothesis that passes the gate MAY result in a compiled maintenance action (e.g., “entries from source X should be promoted to structural earlier”) rather than a compiled routing rule.

The analysis produces a structured evidence record:

evidence:
hypothesis_id: HYP-{UUID-SHORT}
analysis_date: ISO-8601
traces_evaluated: integer
time_span_covered: duration
accuracy: float # fraction of decisions that match ground truth
coverage: float # fraction of relevant decisions the rule applies to
false_positive_rate: float # fraction of rule applications that were incorrect
improvement_over_baseline: float # marginal improvement vs. pure inference
regression_cases: integer # count of cases where the rule was worse than inference
verdict: COMPILE | REJECT | INSUFFICIENT_DATA
reasoning: free text

Compilation gate:

  • accuracy >= minimum_evidence.consistency (default 0.90)
  • traces_evaluated >= minimum_evidence.trace_count (default 50)
  • time_span_covered >= minimum_evidence.time_span (default 14 days)
  • false_positive_rate < 0.05 (default; configurable)
  • regression_cases == 0 OR regression_cases / traces_evaluated < 0.02 with human approval

If the evidence does not meet the gate, the hypothesis is marked REJECTED with a reason. Rejected hypotheses MAY be re-evaluated after additional traces accumulate, but MUST NOT be re-submitted within minimum_evidence.time_span of the rejection (prevents retry-spam).

A hypothesis that passes the evidence gate is compiled into an explicit rule in the conventions document.

Compiled rule format:

compiled_rule:
id: RULE-{UUID-SHORT}
compiled_from: HYP-{UUID-SHORT}
compiled_date: ISO-8601
node_id: compiling node
type: classification_shortcut | query_routing | contradiction_heuristic | condensation_template
rule: "source:meeting-transcript → region:operational"
confidence: float # from evidence record
evidence_summary:
traces: integer
accuracy: float
time_span: duration
level: 2 # inference optimization level (convention rule)
τ_star: duration # computed optimal lifespan (see Section 7.6)
review_date: ISO-8601 # compiled_date + τ*
changelog_ref: CHANGE-{SEQ} # reference to changelog entry
status: ACTIVE

Requirements:

  • Compilation MUST be logged in the protocol changelog with the full evidence chain
  • The previous behavior (pure inference for this decision class) MUST be preserved as a rollback target
  • The rule MUST include a computed τ* and a corresponding review date

For single-node deployments, compilation and promotion are the same event — the rule is written into the node’s local conventions.

For colony deployments, promotion has an additional phase (see Section 7.8 for colony-level evolution):

  1. The compiled rule is broadcast as a CONVENTION_PROPOSAL on the colony coordination channel
  2. Other nodes validate the proposal against their own traces
  3. Consensus (N/2+1 validation) promotes the rule to colony-wide conventions

A rule MAY be active locally (node-level promotion) while pending colony-level consensus. Local and colony-level rules are tracked separately in the conventions document.

After a rule is promoted, the node MUST record a comparative evaluation once sufficient post-change traces accumulate. This is the A/B test result that validates whether the self-modification actually improved system performance — not just whether the rule is accurate (monitoring handles that), but whether the change produced better task outcomes in aggregate.

Justification: The evolution path performs counterfactual evaluation before compilation (Stage 3) and monitoring after compilation (Stage 7), but neither answers the question “did this change actually improve system performance?” The counterfactual is hypothetical; monitoring tracks rule accuracy, not downstream impact. The comparative evaluation closes this gap by measuring the same task types before and after the change boundary and comparing.

DS analogue: A/B test result record. Canary deployment metrics. Before/after performance comparison in continuous deployment.

Comparative evaluation format:

comparative_evaluation:
change_id: RULE-{UUID-SHORT} | CHANGE-{SEQ}
change_type: rule_compiled | rule_demoted | policy_adjusted | convention_changed
baseline_window: duration # Pre-change measurement period
comparison_window: duration # Post-change measurement period
task_types_affected: list[string]
metrics:
baseline_outcome_mean: float
post_change_outcome_mean: float
delta: float
statistical_significance: float # p-value or equivalent
sample_size_baseline: integer
sample_size_post: integer
verdict: IMPROVED | DEGRADED | NEUTRAL | INSUFFICIENT_DATA

Requirements:

  • The comparative evaluation MUST be recorded as a comparative_evaluation subtype of the eval telemetry event class (§3.6).
  • The comparison_window MUST be at least equal to the baseline_window to ensure sufficient post-change data.
  • A verdict of DEGRADED SHOULD trigger early review of the promoted rule (same as a DEGRADING status in monitoring).
  • A verdict of INSUFFICIENT_DATA MUST be re-evaluated at the next maintenance cycle until a definitive verdict is reached or the rule reaches its τ* review date.

Every compiled rule is continuously monitored through ongoing eval traces. The monitoring system tracks:

rule_monitoring:
rule_id: RULE-{UUID-SHORT}
window: trailing N operations (configurable, default 100)
metrics:
accuracy: float # match rate against ground truth in the trailing window
invocation_count: integer # how often the rule fires
override_count: integer # how often the agent overrides the rule's suggestion
false_positive_count: integer
degradation_signal: float # D(t) — instantaneous degradation rate (see Section 7.6)
status: HEALTHY | DEGRADING | DEMOTED

Requirements:

  • Monitoring MUST update on every operation that falls within the rule’s domain
  • If accuracy drops below confidence * 0.9 (10% relative degradation from compilation accuracy), the rule’s status MUST change to DEGRADING
  • A DEGRADING rule MUST be flagged for early review (before its scheduled τ* review date)
  • If accuracy drops below 0.5 (worse than random for binary decisions), the rule MUST be immediately demoted to pure inference (Level 0) without waiting for review

When a rule reaches its review date (compiled_date + τ*) or is flagged as DEGRADING, the system re-evaluates it:

Renewal procedure:

  1. Re-run the analysis (Stage 3) using traces accumulated since compilation
  2. Produce a new evidence record (Stage 4)
  3. If the evidence still meets the compilation gate: renew — update τ*, reset the review date, log the renewal in the changelog
  4. If the evidence no longer meets the gate: cull — demote the rule to Level 0 (pure inference), archive the rule with status: CULLED, log in changelog

Requirements:

  • Renewal MUST NOT be skipped. A rule past its review date without a renewal decision is a protocol violation.
  • Culled rules MUST be archived with their full lifecycle history (hypothesis, evidence records, monitoring data, renewal/cull decision). This is the provenance chain for the system’s self-modification history.
  • A culled rule MAY be re-derived from scratch if new evidence supports it — but it starts the lifecycle from EXPERIENCE, not from the previous rule’s state.

7.6 CDR-Enforced Lifecycle: The Safety Mechanism

Section titled “7.6 CDR-Enforced Lifecycle: The Safety Mechanism”

The self-modification system’s safety property is that every compiled rule has a finite optimal lifespan. This is CDR Theorem 3 applied to the system’s own conventions.

Every compiled rule operates on an implicit distributional assumption — the pattern it encodes was stable during the evidence period. As the operational environment drifts, the rule’s accuracy degrades. This degradation is not a possibility; it is a mathematical certainty under CDR Theorem 3 (any compression scheme degrades under distributional drift).

Define:

  • D(t): the instantaneous degradation rate of a rule at time t since compilation. Measured as the derivative of the rule’s cumulative error rate.
  • J(t): the long-run average cost, computed as total cumulative error divided by t.
  • τ:* the optimal renewal interval, defined by D(τ*) = J(τ*).

The rule SHOULD be renewed when D(τ*) = J(τ*) — the point at which the instantaneous degradation rate equals the long-run average cost. Continuing past this point means the rule is costing more than its replacement would.

Pathological rules — rules that actively harm system performance — are self-correcting under this framework:

  1. A pathological rule shows high and increasing D(t) immediately after compilation
  2. High D(t) drives τ* to a small value (the instantaneous degradation quickly exceeds the long-run average)
  3. Small τ* triggers early review
  4. At review, the evidence gate fails (accuracy below threshold)
  5. The rule is culled

The system does not need to define “pathological” in advance. CDR Theorem 3 guarantees that harmful rules have short optimal lifespans. The math does the detection; the lifecycle does the remediation.

The feedback loop — system modifies conventions → changes processing → generates new patterns → more modifications — is a potential instability. CDR addresses this directly:

  • A rule compiled based on a pattern that subsequent modifications have altered shows increasing D(t) for that rule
  • The monitoring system detects the degradation
  • The rule is flagged for early review or immediately demoted
  • The system converges: rules that survive renewal are genuinely stable; rules that are artifacts of previous modifications are culled within one τ* cycle

Rate limiting as a secondary safeguard: The protocol enforces a maximum compilation rate per node per time window (configurable, default: 5 rules per 7-day window). This prevents a node from making many small modifications that individually pass evidence gates but collectively destabilize the system. The rate limit is a hardwired invariant, not an evolvable parameter.


Compiled rules exist at different levels of the inference optimization ladder (Section 4). Self-modification governs movement between levels.

LevelLocationCostPromotion CriteriaDemotion Trigger
0: Pure inferenceNo explicit rule; agent reasons from conventionsFull LLM inference per decisionDefault — no promotion neededRule culled or accuracy < 0.5
1: Pattern recognizedNoted in eval traces but not yet compiledFull inference, but evidence accumulatingAutomatic — threshold crossing in tracesPattern stops appearing in new traces
2: Convention compiledExplicit rule in protocol/conventions.mdReduced inference — rule matching replaces open reasoningEvidence gate (Section 7.5, Stage 4)τ* reached with failing renewal, or accuracy degradation
3: Prompt compiledRule embedded in CLAUDE.md preamble or system promptNear-zero marginal cost — rule in contextColony-level consensus + sustained stability (3+ renewal cycles survived)Any accuracy drop below threshold triggers demotion to Level 2, not Level 0

Requirements:

  • Level 3 promotion MUST require human approval (the system prompt is a high-trust surface)
  • Demotion MUST always be possible regardless of level — no rule is permanently compiled
  • The changelog MUST record every level transition with justification

When multiple nodes in a colony independently observe the same pattern, the evidence for compilation is stronger. Colony-level evolution extends the single-node lifecycle to the colony scope.

1. LOCAL COMPILATION — A node compiles a rule from its own traces (Stages 1-5)
2. BROADCAST — Node publishes the rule as a CONVENTION_PROPOSAL on the colony channel
3. CROSS-NODE VALIDATION — Each receiving node evaluates the proposal against its own traces
4. CONSENSUS — N/2+1 nodes confirm the pattern holds in their experience
5. COLONY PROMOTION — Rule added to colony-wide conventions (colony-state/conventions/)
6. COLONY MONITORING — All nodes track the rule's accuracy in their own operations
7. COLONY RENEWAL — Any node MAY flag a colony-wide rule for re-evaluation
convention_proposal:
id: PROP-{UUID-SHORT}
proposing_node: node_id
proposed_date: ISO-8601
rule: compiled_rule record (Section 7.5, Stage 5)
local_evidence:
accuracy: float
traces: integer
time_span: duration
validation_responses: [] # populated as nodes respond
consensus_threshold: 0.5 # N/2+1 of active (non-fenced) nodes
status: PENDING | ACCEPTED | REJECTED | EXPIRED
expiry: ISO-8601 # proposals that don't reach consensus within this window are EXPIRED
validation_response:
proposal_id: PROP-{UUID-SHORT}
responding_node: node_id
response_date: ISO-8601
verdict: CONFIRM | REJECT | ABSTAIN
local_accuracy: float # accuracy of the proposed rule against this node's traces
local_traces: integer # number of relevant traces this node evaluated
reasoning: free text # required if REJECT
  • A node MUST NOT vote on its own proposal (the proposer’s evidence is already included)
  • ABSTAIN is permitted when a node has insufficient traces in the relevant domain (<10 traces)
  • ABSTAIN votes do not count toward the denominator for consensus (consensus is N/2+1 of non-abstaining active nodes)
  • A colony-wide rule has its own τ* independent of the local rule’s τ*. Colony τ* SHOULD be longer than local τ* (colony evidence is broader, so the rule is more robust to local drift)
  • Any node MAY flag a colony-wide rule for early review by publishing a CONVENTION_REVIEW_REQUEST with supporting evidence (its own degrading accuracy metrics)
  • A colony-wide rule flagged by 2+ nodes MUST be reviewed regardless of its scheduled review date

A node MAY maintain a local rule that differs from the colony-wide rule for the same decision class. When local and colony rules conflict:

  1. The node MUST log the conflict
  2. The node SHOULD follow the colony rule for shared-state operations (writes to shared structural/identity memory)
  3. The node MAY follow its local rule for local-state operations
  4. Persistent conflict (>30 days) SHOULD trigger a colony-level review of the rule

The self-modification lifecycle depends on counterfactual evaluation (Stage 3) and state hypothesis testing (§7.5 Stage 3, State Hypotheses). Both require the ability to execute operations against modified or historical memory state without affecting real node state. The sandbox execution context is the isolation primitive that makes this safe.

Justification: Without sandbox isolation, any replay or counterfactual evaluation that exercises the full processing pipeline (write path, read path, maintenance path) would mutate real state — creating entries, updating links, triggering decay, firing maintenance tasks. The self-modification system cannot evaluate hypotheses about its own state if evaluating them changes that state. The sandbox is the transactional boundary that separates observation from mutation.

Operating timescale: Per-replay (minutes per replayed interaction). Sandbox lifecycle is ephemeral — created for a specific evaluation, discarded after.

DS analogue: Staging environment / shadow traffic. Database transaction with ROLLBACK — execute operations, observe results, roll back without committing. fork() in process management — create a copy of state, let it diverge, observe the divergence, discard the copy.

sandbox:
sandbox_id: SANDBOX-{UUID-SHORT}
created: ISO-8601
base_state: SNAP-{UUID-SHORT} # Node state snapshot to replay against
purpose: counterfactual | replay | consolidation
isolation_guarantee:
- Writes within the sandbox MUST NOT affect real node state
- Reads within the sandbox MUST resolve against the sandbox's base_state,
not the node's current state
- Maintenance operations within the sandbox MUST NOT modify real
maintenance queue or eval traces
sandbox_log: # Captures all operations performed in the sandbox
entries_created: list[entry]
entries_modified: list[entry_diff]
decisions_made: list[eval_trace]
outcome: outcome_signal | null
lifecycle: ephemeral # Sandbox state is discarded after use
# Only the sandbox_log is persisted (for comparison)
PurposeTriggerDescription
counterfactualEvolution path Stage 3 (ANALYSIS)Evaluate a rule hypothesis or state hypothesis against historical traces. The sandbox replays execution traces against the base state (optionally with a state diff applied) and captures outcomes for comparison.
replayConsolidation replay (§6c.4)Re-process selected past interactions against current memory state to detect regressions, gaps, or improvements.
consolidationPeriodic maintenance (§6c.4)Compressed replay at faster timescale for schema formation and pattern extraction — the protocol realization of the DLM sleep consolidation mapping (§1.8).
  1. All operations within a sandbox MUST execute through the standard write, read, and maintenance paths (§3.3). The sandbox does not define new operations — it defines an isolation context for existing operations.
  2. The isolation_guarantee is protocol. A sandbox that leaks mutations to real state is a protocol violation.
  3. The sandbox_log MUST be persisted after sandbox teardown. The log is the observable output of the simulation — it enables comparison between sandbox outcomes and real outcomes.
  4. When to create sandboxes (triggers) is policy. The isolation guarantee and log persistence are protocol.
  5. Sandboxes MUST reference a valid node state snapshot (§6c.5) as their base_state. A sandbox without a restorable base state cannot guarantee isolation.
  • §3.3.5 (Simulation Path): The node contract defines the sandbox as an operation path. This section defines its role within self-modification.
  • §7.5 Stage 3 (ANALYSIS): State hypotheses require sandbox execution for evaluation.
  • §6c.4 (Scheduled Maintenance): Consolidation replay creates sandboxes for re-processing past interactions.
  • §6c.5 (Node State Snapshot): Sandboxes depend on snapshots for their base_state.

The self-modification lifecycle compiles routing rules (§7.4.3), evaluates hypotheses against historical traces (§7.5 Stage 3), and promotes rules to colony-wide conventions (§7.8). All of these require the node to know how well it performs at different categories of work. The task-type performance register provides this empirical basis.

Justification: Without per-task-type performance data, the action cascade’s affinity() function (§3.4) has no empirical grounding — routing decisions are based on declared scope rather than measured competency. Colony-level evolution (§7.8) cannot evaluate whether a proposed rule applies to task types where the proposing node actually performs well. And the self-modification system cannot prioritize which domains need rule compilation most urgently (the domains where performance is worst or most variable).

Operating timescale: Updated per-interaction (append to outcome stream). Aggregated periodically (daily or per-session) as a scheduled maintenance task.

DS analogue: Service-level indicators / service-level objectives (SLIs/SLOs) per endpoint. Each task type is analogous to a service endpoint with its own latency, error rate, and throughput metrics.

task_type_register:
node_id: string
task_type: string
window: trailing N interactions
metrics:
outcome_mean: float # Mean outcome_quality across tasks of this type
outcome_trend: float # Slope of outcome_quality over time (improving/degrading)
volume: integer # How many tasks of this type in the window
retrieval_hit_rate: float # Fraction of queries for this task type that returned useful results
contribution_rate: float # Fraction of retrieved entries that appeared in positive outcome signals
  1. The register MUST be maintained per node. Each node tracks its own performance by task type.
  2. The task_type taxonomy is policy (implementations define their own task categories). The requirement that the register exists and is queryable is protocol.
  3. The register MUST be updated whenever a new outcome signal (§3.3.2 Stage 4) is recorded for an interaction of the given task type.
  4. Aggregate metrics (outcome_trend, retrieval_hit_rate, contribution_rate) SHOULD be recomputed at each aggregation interval, not maintained as running averages, to prevent drift.
  5. A node with an empty register for a given task type MUST report INSUFFICIENT_DATA for that task type’s affinity rather than a default value. Absence of data is not evidence of low performance.
ConsumerHow It Uses the Register
Action cascade (§3.4)affinity(node, input) is grounded in the node’s outcome_mean and retrieval_hit_rate for the input’s task type. A node with high measured performance on research queries has higher affinity for research input than a node with low or unknown performance.
Hypothesis prioritization (§7.5 Stage 2)Task types with degrading outcome_trend or low retrieval_hit_rate are higher-priority targets for hypothesis formation. The system should look for compilable patterns where performance is worst.
Colony routing (§7.8)Colony-wide convention proposals can reference proposing-node register data as additional evidence. A routing rule proposed by a node with strong performance in the relevant task type carries more weight than one proposed by a node with sparse data.
Comparative evaluation (§7.5 Stage 6)The register provides the baseline_outcome_mean and post_change_outcome_mean for the comparative evaluation record. Changes are measured per task type, not globally.
  • §3.4 (Action Cascade): The register feeds affinity().
  • §3.6 (Telemetry Conformance): The register is defined as a telemetry aggregation in the node contract.
  • §6c.4 (Scheduled Maintenance): task_type_performance is a periodic maintenance task.
  • §7.5 Stage 6 (Comparative Evaluation): Per-task-type metrics are the measurement basis for before/after comparison.

For clarity, the following are outside the scope of self-modification and MUST NOT be targets of the compilation pipeline:

  • Hardwired invariants (Section 7.4.1) — provenance, eval trace format, conflict resolution, changelog, CDR properties
  • Fault isolation invariant checks (Section 7b, Layer 1) — these are structural safety mechanisms, not optimizable patterns
  • Human steering parameters — colony size limits, budget caps, domain constraints. The system MAY propose changes, but changes MUST require human approval.
  • The self-modification lifecycle itself — the evidence gates, compilation procedure, monitoring thresholds, and renewal mechanics described in this section are protocol, not convention. The system cannot modify how it modifies itself.

This last point is the kernel/userspace boundary (from the self-modifying cognitive architecture research). The system can modify its processing (userspace). It cannot modify the lifecycle that governs those modifications (kernel). The lifecycle is the invariant that makes self-modification safe.