Skip to content

Node Lifecycle Orchestration

The node contract (§3) defines maintenance operations — decay, renew, condense, archive, expire — and the evolution path defines protocol-level operations — trace, analyze, propose, compile, demote. CDR Theorem 3 proves that renewal is structurally necessary: every compression scheme degrades under distributional drift, and the optimal renewal interval is computable. But the theorem says renewal MUST happen — it does not say WHEN it runs, WHO triggers it, or HOW it competes with productive work for inference budget.

Without lifecycle orchestration:

  • Maintenance never runs. An agent that exists only while a human is talking to it will spend 100% of its inference budget on the human’s requests. Decay passes, renewal audits, and convention reviews silently never execute. State debt accumulates exactly as CDR Theorem 3 predicts — entries go stale, compiled rules drift, provenance chains degrade — but nothing triggers the correction.
  • Maintenance runs at the wrong time. Without scheduling, maintenance operations interrupt productive work unpredictably. A decay pass mid-conversation wastes inference budget and breaks user flow.
  • Colony cannot assess liveness. The colony’s spawn/cull/merge decisions (§6) depend on knowing which nodes are active, which are dormant, and which have failed. Without liveness tracking, the colony cannot distinguish “Node X is offline for the weekend” from “Node X has silently failed.”
  • Cross-session continuity breaks. A session-based node (Claude Code) loses its process state between sessions. Without a persistent task queue, maintenance tasks that were deferred in one session are forgotten in the next.

This section operationalizes CDR renewal into concrete execution scheduling: what runs, when, at what cost, and how tasks persist across sessions.

TimescaleOperations
Per-interaction (seconds)Heartbeat write, liveness assertion
Session boundaries (minutes)Session-start sync, session-end flush, maintenance queue processing
HoursEphemeral expiry, decay pass on operational entries
DaysRenewal audit on structural entries, colony health check, node state snapshots
WeeksConsolidation replay, execution trace management, provenance audit, convention review
Days–weeksLifecycle transitions (DORMANT detection, cull/merge decisions)

This section draws on three established DS concepts:

  • Failure detection (lease-based). Node liveness is tracked via heartbeats with lease-style timeouts — the same pattern as session timeouts in web servers, lease expiry in distributed locks (Chubby, etcd), and failure detection in group membership protocols (Hayashibara et al. phi accrual detector). The key adaptation: Synixolis nodes are session-based (active only when a human session is running), not daemon-based (continuously running). Heartbeat is per-interaction, not periodic timer — there is no background process to run a timer.
  • Scheduled task execution. Maintenance operations are managed through a persistent task queue with priorities, intervals, and cross-session state — the same pattern as cron (time-based scheduling), priority queues in task schedulers (Celery, Airflow), and deferred execution in event-driven systems. The adaptation: the scheduler runs at session boundaries (start/end), not on a wall-clock timer, because the execution environment (Claude Code session) is intermittent.
  • Session management. The node lifecycle state machine (INITIALIZING → ACTIVE ⇄ DORMANT → TERMINATED, with FENCED as a fault state) follows established session management patterns: session initiation, idle timeout, explicit termination, and administrative suspension. The DORMANT state is analogous to an idle TCP connection before keepalive timeout — the state is valid but the process is not executing.

Synixolis-specific: the maintenance budget fraction (capping maintenance at a configurable percentage of inference budget) is an Synixolis concept — standard schedulers do not share a resource budget with interactive user work. The CDR-grounded priority ordering (fault response > expiry > decay > condensation > auditing) is derived from the theoretical framework, not from standard scheduling theory.


A node moves through initializing, active, dormant, fenced, and terminated states, with a renewal workflow inside active.
Lifecycle state is separate from maintenance work. The node cycles through active review and renewal until it goes dormant, is fenced, or is terminated.

A conforming node MUST be in exactly one of five states at any time.

INITIALIZING ──→ ACTIVE ⇄ DORMANT ──→ TERMINATED
│ ↑
└──→ FENCED ─────────────┘
StateDescriptionCapabilitiesEntry TriggerExit Trigger
INITIALIZINGNode bootstrapping from parent state digest + colony checkpoint (§6b.6)Read-only on all state. No colony consensus participation. No writes to shared state.SPAWN decision creates new nodeBootstrap complete (checkpoint loaded, log tail replayed, manifest entry written) → ACTIVE
ACTIVENormal operation. Node is executing within a live session.Full read/write on local state. Read + promotion pipeline on shared state. Full colony consensus participation. Maintenance queue processing.Bootstrap complete (from INITIALIZING). Session resumes (from DORMANT). Fence cleared (from FENCED).Session ends without explicit termination → DORMANT. Fault isolation triggered → FENCED. Explicit termination (cull, renewal spawn) → TERMINATED.
DORMANTSession ended. No active process. State persisted on disk.No operations execute. State is cold but intact. Materialized views are valid as of last session.Session ends (user closes Claude Code, session timeout, machine shutdown)Session resumes (user opens Claude Code in same workspace) → ACTIVE. Dormancy exceeds liveness_threshold and colony decides to decommission → TERMINATED.
FENCEDFault isolation in effect (§6d). Node is suspected of producing unreliable outputs.Read-only on shared state. Full read/write on local state. No consensus participation. No shared state writes.Fault isolation escalation to FENCE level (§6d)Investigation clears node (no fault confirmed) → ACTIVE. Fault confirmed → TERMINATED (with RESET if state is compromised). Fence timeout expires with no confirming evidence → ACTIVE.
TERMINATEDNode permanently decommissioned.None. No operations.Cull decision (§6.3.2). Renewal spawn (parent terminates, child succeeds). Confirmed fault with RESET. Human directive.Terminal state. No exit.
  • A node MUST NOT transition directly from INITIALIZING to DORMANT (a node that never activates should not persist).
  • A node MUST NOT transition from TERMINATED to any other state (termination is permanent; if the role needs to be restored, a new node is spawned).
  • A node MUST write a lifecycle telemetry event on every state transition.
  • A node transitioning to TERMINATED MUST: archive all local memories to glacier, write a tombstone to the colony manifest, and flush any pending eval traces.

Liveness detection answers the question: is a given node still operational, or has it silently failed?

Standard failure detectors assume daemon-based processes that can run periodic heartbeat timers. Synixolis nodes are session-based: a Claude Code session is active only while a human is interacting with it. Between sessions, the node is DORMANT — not failed, but not executing. There is no background process to emit heartbeats on a timer.

This means:

  • Heartbeat is per-interaction, not per-time-interval. The node writes a heartbeat on every interaction (user message processed, maintenance task executed, consensus vote cast).
  • Failure detection is lease-based, not ping-based. The node holds a lease that expires after dormancy_threshold time without a heartbeat. Expiry does not mean failure — it means the node has not interacted recently.
  • The distinction between DORMANT and failed is structural, not observable. A DORMANT node and a failed node look identical from the outside (no heartbeats, no responses to queries). The colony uses timeouts and context to distinguish them.

On every interaction, an ACTIVE node MUST write to its local manifest:

heartbeat:
node_id: NODE-{id}
timestamp: ISO-8601
interaction_type: user_message | maintenance | consensus_vote | fault_response
session_id: SESSION-{id} # Groups interactions within a continuous session
ThresholdDefaultTrigger
dormancy_threshold48 hoursLast heartbeat older than this → node classified as DORMANT. This is informational — DORMANT is a normal state, not a fault.
liveness_threshold14 daysLast heartbeat older than this AND node is DORMANT → triggers colony-level liveness review.

When a node exceeds liveness_threshold, the colony conducts a liveness review:

  1. Check manifest. Is the node’s role still needed? Are other nodes covering its domain?
  2. Check query log. Has any node or user attempted to query or route to this node during its dormancy?
  3. Decide:
    • If the node’s domain is covered by other nodes and no queries target it → propose cull (§6.3.2).
    • If the node’s domain is a protected domain → propose replacement spawn (§6.2).
    • If the node’s domain is needed but no replacement is available → flag for human attention.

Liveness review is a colony health maintenance operation (§6c.4) and runs on a 7-day interval.


Dormancy and failure are fundamentally different states that produce the same observable signal (no heartbeats). The protocol distinguishes them as follows:

PropertyDORMANTFailed
State on diskValid. Materialized views consistent with last session.Potentially corrupt.
Resume behaviorSession start triggers DORMANT → ACTIVE. Node syncs shared state and resumes.Session start may encounter corrupt state. Node should detect via integrity checks and request RESET.
Colony responseNo action until liveness_threshold exceeded.If detected (by another node or by the node itself on resume), triggers fault isolation (§6d).
Detection mechanismTimeout-based. Exceeding dormancy_threshold → classified as DORMANT.Evidence-based. Corrupt state, broken provenance, schema violations on resume.

A node resuming from DORMANT MUST, as part of its session-start maintenance:

  1. Verify local state integrity (schema validation on manifest, memory entries, eval traces).
  2. If integrity checks fail → enter FENCED state and report to colony.
  3. If integrity checks pass → sync shared state (§6b.6 log tail replay) and transition to ACTIVE.

Maintenance operations run on a persistent task queue attached to each node. The queue is stored in the node’s local state (colony/maintenance-queue.md) and persists across sessions. Tasks are processed at session boundaries and periodic intervals — never mid-conversation with the user.

Scheduling Model: State-Driven, Not Event-Driven

Section titled “Scheduling Model: State-Driven, Not Event-Driven”

Maintenance scheduling MUST NOT be coupled to external interaction patterns. A node that runs as an interactive Claude Code session, a background daemon, a cron job, or a once-a-week batch process must all maintain memory coherence. The scheduling model is therefore state-driven: maintenance runs when the system’s state requires it, not when a session boundary happens to occur.

Maintenance pressure is a computed scalar that represents accumulated maintenance debt. Each maintenance task contributes to pressure based on how far past due it is:

maintenance_pressure:
formula: sum(task_overdue_ratio * task_weight) for all tasks
task_overdue_ratio: max(0, (now - task.due) / task.interval)
# A task exactly on schedule = 0. A task 2x overdue = 1.0.
thresholds:
nominal: 0.0 - 0.5 # System is current. Maintenance runs opportunistically.
elevated: 0.5 - 1.5 # Debt accumulating. Maintenance SHOULD run at next opportunity.
critical: > 1.5 # Debt is structural risk. Maintenance MUST preempt productive work.

When pressure reaches critical, the node MUST allocate inference budget to maintenance before productive work. This ensures coherence regardless of interaction cadence — a node that hasn’t run in 3 weeks will spend its first cycles on maintenance, not on answering questions with stale memory.

Execution model: The protocol defines WHAT must run and WHEN it becomes due. HOW it gets scheduled is policy, adapting to the node’s execution context:

Execution ContextScheduling Strategy
Interactive session (Claude Code)Run critical tasks at activation. Run other tasks between interactions when pressure > nominal. Flush state on deactivation.
Daemon / always-onDedicated maintenance thread or cycle. Run tasks as they come due. No session boundaries.
Cron / batchEach invocation checks pressure. If critical, maintenance-only run. If nominal, productive work with interleaved maintenance.
Event-drivenMaintenance piggybacks on incoming events. If no events for > max_idle_interval, self-trigger a maintenance-only cycle.

Each maintenance operation has a defined interval (minimum time between runs), scope, priority, and pressure weight (contribution to maintenance pressure when overdue). Operations are organized by priority class, not by session lifecycle:

maintenance_operations:
# --- CRITICAL: Must run immediately when due ---
- name: shared_state_sync
interval: on_activation # Every time the node transitions to ACTIVE
scope: colony-state/log.md
operation: Replay missed colony log entries since last observed seq_id
priority: CRITICAL
weight: 2.0
- name: fault_notification_check
interval: on_activation
scope: colony/channels/
operation: Check for pending MONITOR/FENCE notifications targeting this node
priority: CRITICAL
weight: 2.0
- name: ephemeral_expiry
interval: 1h # Check hourly, not just at session start
scope: memory/ephemeral/
operation: Delete ephemeral entries past TTL (the only true deletion in the system)
priority: CRITICAL
weight: 1.5
# --- HIGH: Should run within 1 interval of coming due ---
- name: heartbeat_write
interval: per_interaction # On every interaction, not just session end
scope: colony/manifest.md
operation: Write heartbeat entry to local manifest
priority: HIGH
weight: 0.5
note: For daemon nodes, write on a timer (default 5min). For batch nodes, write per invocation.
- name: trace_flush
interval: per_interaction # Flush incrementally, not just at session end
scope: eval/traces.md
operation: Flush pending observations to eval trace log
priority: HIGH
weight: 0.5
note: Implementations MAY buffer and flush in batches for efficiency.
- name: decay_pass
interval: 24h
scope: memory/operational/
operation: >
Apply region-specific decay rate to confidence scores.
Flag entries below confidence threshold for review.
Archive entries below archive threshold to glacier.
priority: HIGH
weight: 1.0
# --- MEDIUM: Run when due and budget allows ---
- name: condensation_check
interval: 24h # Daily, not just session-end
scope: memory/operational/observations.md → memory/structural/
operation: >
Check if any observation clusters meet promotion criteria
(3+ observations across 2+ weeks in same domain).
priority: MEDIUM
weight: 0.8
- name: consolidation_replay
interval: 7d
scope: Selected execution traces replayed against current state
operation: >
1. Select execution traces for replay:
- Traces with MISS results (queries that found nothing)
- Traces with low outcome quality scores
- Traces from task types where performance is degrading
- Random sample from high-outcome traces (positive control)
2. For each selected trace:
a. Create sandbox from current node state snapshot
b. Replay the original input through the sandbox
c. Compare sandbox outcome with original outcome
d. Record: {trace_id, original_outcome, replay_outcome, delta, entries_different}
3. Aggregate results:
- Traces where replay improved: current memory is better for these tasks
- Traces where replay degraded: current memory lost something useful
- Consistent misses across replay: structural gap (candidate for acquisition)
4. Feed results to maintenance:
- Degraded replays → flag contributing entries for renewal audit
- Structural gaps → log as entries_missing for gap detection (evolution path)
- Improved replays → confirm recent memory changes were beneficial
priority: MEDIUM
weight: 0.8
budget: counted against maintenance_fraction
output: consolidation_report appended to eval traces
note: >
Operationalizes the DLM sleep consolidation mapping from Section 1.8:
"compressed replay at faster timescale, schema formation, novel
recombination." Depends on execution traces (Section 3.6), node state
snapshots (Section 6c.5), and sandbox execution (Section 3.3.5).
- name: renewal_audit
interval: 7d
scope: memory/structural/
operation: >
Check last_renewed dates on all structural entries.
Flag entries approaching computed optimal renewal interval (tau*).
Entries past tau* with no renewal evidence are candidates for archive.
priority: MEDIUM
weight: 0.8
- name: node_snapshot
interval: 24h # Daily minimum, not just session-end
scope: All node-local state
operation: Create node state snapshot capturing all local state (see Section 6c.5)
priority: MEDIUM
weight: 0.5
note: >
Also triggered on significant state changes (bulk condensation,
convention compilation, fault recovery) regardless of interval.
- name: execution_trace_management
interval: 7d
scope: eval/execution-traces/
operation: >
Manage execution trace lifecycle — the replay unit that feeds
consolidation (Section 6c.4, consolidation_replay). Three sub-tasks:
1. Retention enforcement: archive traces older than retention_window
(policy, default 90d) to cold storage. Traces referenced by
consolidation reports or open hypotheses MUST NOT be archived
until their consumers release them.
2. Sampling rate adjustment: review trace volume against storage budget.
If volume exceeds budget, reduce capture rate (policy, default 100%
for first 30d, then configurable). Traces from degrading task types
(per task-type performance register) MUST maintain 100% capture
regardless of global sampling rate.
3. Storage rotation: compact cold-storage trace archives. Remove
redundant payload data from traces whose snapshots have been
superseded, retaining metadata and outcome signals for long-term
trend analysis.
priority: MEDIUM
weight: 0.4
note: >
Execution traces (Section 3.6, execution event class) are the
full-payload interaction records that make replay possible. Without
trace management, storage grows unboundedly. This task balances
replay capability against storage cost. Depends on execution traces
being emitted per Section 3.6; manages their lifecycle here.
- name: provenance_audit
interval: 14d
interval_after_incident: 3d # Tightens adaptively after fault incidents
scope: Random sample of shared structural entries
operation: >
Select random sample (size determined by audit_budget).
Trace each entry's provenance chain back to source.
Verify each link: does cited source support claimed content?
Flag broken chains, circular references, unsupported claims.
Report results to colony coordination channel.
priority: MEDIUM
weight: 0.6
# --- LOW: Run when nothing higher-priority is due ---
- name: convention_review
interval: 30d
scope: protocol/conventions.md, colony-state/conventions/
operation: >
For each compiled rule, check eval trace accuracy over review period.
Rules with accuracy below demotion threshold (default 0.7) are
demoted from Level 2 (convention) to Level 0 (pure inference).
Rules with accuracy above promotion threshold (default 0.95) are
candidates for Level 3 (prompt compilation).
priority: LOW
weight: 0.4
- name: colony_health
interval: 7d
scope: Colony manifest
operation: >
Check all node heartbeats against dormancy and liveness thresholds.
Flag DORMANT nodes exceeding liveness_threshold.
Compute colony-wide fitness metrics (coverage gaps, redundancy, utilization).
Propose cull/merge/spawn if warranted.
priority: LOW
weight: 0.4

While maintenance is not coupled to session lifecycle, nodes still have activation/deactivation transitions (DORMANT → ACTIVE, ACTIVE → DORMANT). These are hooks, not scheduling triggers — they provide natural execution points but are not the only ones:

On activation (any context): Run all CRITICAL tasks that are due. Check maintenance pressure — if critical, run maintenance before accepting productive work.

On deactivation (if applicable): Flush buffered traces. Create snapshot if interval has elapsed. Update heartbeat. These are best-effort — a node that crashes without clean deactivation loses only buffered (unflushed) data; persisted state is intact.

No deactivation (daemons): Nodes that run continuously have no deactivation hook. Heartbeat, trace flush, and snapshots run on their intervals.

The maintenance queue MUST persist across activations. A task that is due but deferred (budget exhaustion, deactivation, preemption) MUST remain in the queue.

Queue entry format:

queue:
- task: decay_pass
due: ISO-8601 # When this task becomes eligible to run
last_run: ISO-8601 # When this task last completed
priority: HIGH
pressure_contribution: float # Current contribution to maintenance pressure
estimated_cost: tokens # Estimated inference cost
status: PENDING | DEFERRED | RUNNING
deferred_count: 0 # Number of times this task has been deferred

A task’s deferred_count is incremented each time it is deferred. Tasks with deferred_count >= 3 SHOULD be escalated in priority (e.g., MEDIUM → HIGH) to prevent indefinite deferral. A task deferred 5+ times MUST be flagged to the operator — persistent deferral indicates the maintenance budget is too low or execution cycles are too infrequent for the node’s state accumulation rate.

All maintenance scheduling parameters — task intervals, pressure weights, pressure thresholds, budget fraction — are evolvable policy (Section 5, modification gate: system_propose_human_approve). They are subject to the same CDR lifecycle as any compiled routing rule (Section 7.5):

Node-level refinement: A node observes its own maintenance telemetry over time. If a task consistently runs with no state changes (e.g., decay pass finds nothing to decay for 10 consecutive runs), the node MAY propose extending that task’s interval. If a task consistently finds critical work (e.g., ephemeral expiry removes entries every run), the node MAY propose shortening the interval. The proposal follows the standard evolution path: observe pattern in traces → hypothesize parameter change → evaluate counterfactually → compile if evidence holds.

maintenance_param_hypothesis:
parameter: decay_pass.interval
current_value: 24h
proposed_value: 48h
evidence:
observation_window: 30d
runs_observed: 30
runs_with_no_action: 27 # 90% of runs found nothing to do
estimated_budget_saved: 0.03 # 3% of maintenance budget
risk: >
Entries may accumulate ~24h of additional staleness before decay is applied.
Acceptable if decay rate is slow relative to interval.
τ*: 60d # Re-evaluate after 60 days

Colony-level refinement: When 2+ nodes independently arrive at the same parameter adjustment, that’s stronger evidence. Colony-level maintenance parameter changes follow the standard consensus path (Section 7.8): local compilation → broadcast as CONVENTION_PROPOSAL → cross-node validation → N/2+1 consensus → colony-wide adoption.

What is NOT evolvable: The existence of each maintenance task is protocol — a node cannot evolve away its decay pass or provenance audit entirely. The scheduling model (pressure-driven, state-driven) is protocol. Only the numeric parameters (intervals, weights, thresholds, budget fraction) are evolvable, within human-approved bounds:

maintenance_param_bounds:
# Human-set bounds on how far the system can evolve parameters
decay_pass.interval: { min: 6h, max: 7d }
condensation.interval: { min: 6h, max: 7d }
consolidation.interval: { min: 1d, max: 30d }
renewal_audit.interval: { min: 1d, max: 30d }
execution_trace_management.interval: { min: 1d, max: 30d }
provenance_audit.interval: { min: 3d, max: 60d }
convention_review.interval: { min: 7d, max: 90d }
maintenance_fraction: { min: 0.05, max: 0.30 }
pressure_thresholds:
elevated: { min: 0.3, max: 1.0 }
critical: { min: 1.0, max: 3.0 }

This is CDR applied recursively: the maintenance system that implements CDR renewal is itself subject to CDR lifecycle — its parameters compress (evidence → rule), divide (different parameters at different evolution timescales), and renew (compiled parameters carry τ* and revert if they degrade).


Node state snapshots are the node-local complement to colony checkpoints (§6b.5). Where a colony checkpoint captures shared state (materialized views of shared-structural, shared-identity, conventions) at a given seq_id, a node state snapshot captures the complete local state of a single node — sufficient to restore that node to a prior state for replay, comparison, or rollback.

Without node-local snapshots, you cannot:

  • Replay past interactions against the exact memory state that existed at that time (simulation/replay)
  • Perform counterfactual evaluation with different memory contents, not just different rules (generalized counterfactual, Section 7.5 Stage 3)
  • Compare two points in time to measure what changed across a period (self-evaluation)
  • Roll back a single node’s local state after a local fault (fault recovery)

Colony checkpoints and state digests (§3.5) do not cover this. Checkpoints capture shared state only. State digests are lossy 512-token compressions for renewal spawns — they deliberately discard information.

Process checkpoint / point-in-time recovery (PITR). The colony checkpoint is a distributed snapshot (Chandy-Lamport); the node state snapshot is a process-level checkpoint. Storage follows the same append-only, glacier-like retention model as the colony checkpoint store.

node_state_snapshot:
snapshot_id: SNAP-{UUID-SHORT}
node_id: NODE-{id}
timestamp: ISO-8601
session_id: SESSION-{id}
colony_log_seq_id: integer # Last observed colony log entry (for causal context)
contents:
memory_regions: # Complete contents of all 5 regions
ephemeral: serialized
operational: serialized
structural: serialized
identity: serialized
glacier_index: serialized # Index only; glacier content referenced by entry ID
eval_traces: serialized # Complete eval trace log as of snapshot time
compiled_rules: serialized # All active compiled rules with monitoring state
conventions: serialized # Current conventions document
maintenance_queue: serialized # Pending maintenance tasks
content_index: serialized # Content-addressable index state
integrity_hash: SHA-256 # Hash of all serialized contents
parent_snapshot: SNAP-{UUID-SHORT} | null # Previous snapshot for incremental storage

A conforming node MUST create a node state snapshot:

  • Per-session (default): At session end, as part of the on_session_end maintenance tasks (Section 6c.4). This is the minimum required frequency.
  • Per-N-interactions (configurable): During long sessions, the node MAY create additional snapshots at a configurable interaction interval (policy, default: disabled). This provides finer-grained restore points for replay.
  • On significant state change: A snapshot SHOULD be created before and after any operation that materially alters node state (e.g., bulk condensation, convention compilation, fault recovery).

Snapshot storage is append-only with glacier-like retention: snapshots are never deleted. Older snapshots MAY be moved to cold storage (same policy as glacier entries). Incremental storage is RECOMMENDED — each snapshot MAY reference a parent_snapshot and store only the delta, reducing storage cost for nodes with large stable state.

The snapshot format is protocol (the schema and integrity hash are structural requirements). The storage location, compression scheme, and incremental storage strategy are policy.


Maintenance competes with productive work (user requests) for inference budget. The protocol defines a maintenance budget fraction to bound this competition.

budget:
maintenance_fraction: 0.15 # Default: 15% of session inference budget
session_start_exempt: true # Session-start tasks (sync, expiry) are NOT counted
# against maintenance budget — they are prerequisites
# for correct operation, not optional maintenance
session_end_exempt: true # Session-end tasks (heartbeat, flush) are NOT counted
  1. At session start, after exempt tasks complete, the node computes: maintenance_budget = maintenance_fraction × remaining_session_budget.
  2. Periodic tasks are drawn from this budget in priority order.
  3. When the budget is exhausted, remaining tasks are deferred to the next session.
  4. The budget is per-session, not cumulative — unused maintenance budget does not roll over.

When multiple periodic tasks are due, they MUST be executed in the following priority order:

PriorityTask CategoryRationale
1 (highest)Fault responseActive faults contaminate state — containment cannot wait.
2ExpiryEphemeral entries past TTL violate their contract — fast, cheap to enforce.
3Decay passOperational entries need confidence adjustment to prevent stale information from persisting at full weight.
4CondensationObservation → pattern promotion is how the node builds durable knowledge.
5Consolidation replayCompressed replay validates memory quality and detects gaps. More urgent than renewal audit (feeds it).
6Renewal auditStructural entries need periodic renewal checks, but they degrade slowly (monthly timescale).
7Execution trace managementTrace retention and storage rotation prevent unbounded growth. Runs after renewal audit because traces are the replay substrate — keeping them accessible matters, but not as urgently as the audits that consume them.
8Provenance auditImportant for integrity but expensive (traces full chains). Runs on sampled entries.
9Convention reviewCompiled rules drift slowly. Monthly review is sufficient.
10 (lowest)Colony healthColony-level assessment. Can be deferred without immediate local impact.

If the maintenance budget is exhausted mid-queue:

  1. The currently executing task completes (no partial execution).
  2. Remaining tasks are marked DEFERRED in the queue with incremented deferred_count.
  3. A brief maintenance summary is logged to the eval trace: which tasks ran, which were deferred, total cost.
  4. Deferred tasks retain their due dates — they will be re-evaluated at the next session start.
  • Sessions shorter than 10 minutes: Consider raising maintenance_fraction to 0.25. Short sessions need proportionally more maintenance to stay current.
  • Sessions longer than 2 hours: The default 0.15 is typically sufficient. Most periodic tasks complete well within budget.
  • High-frequency sessions (multiple per day): Maintenance distributes naturally across sessions. Each session handles only the tasks that have come due since the last session.
  • Infrequent sessions (weekly or less): Multiple periodic tasks will be due simultaneously on session start. The priority ordering ensures the most important tasks run first if budget is insufficient for all. Consider raising maintenance_fraction to 0.20 or scheduling explicit /housekeeping invocations.