Fault Isolation and Byzantine Mitigation
7b.1 Justification
Section titled “7b.1 Justification”CDR Theorem 3 provides a mechanism for handling gradual degradation: every compression scheme has a finite optimal lifespan, and the renewal interval is computable. This handles the fail-slow case — entries go stale, rules drift, patterns shift. The renewal cycle operates on days-to-months timescales and assumes degradation is predictable and independent across entries.
Byzantine faults violate both assumptions. A node that writes fabricated entries with plausible provenance, a compiled rule that actively misroutes entries after a domain shift, a consensus proposal backed by fabricated trace evidence — these are sudden, correlated, and potentially adversarial. They propagate through shared state faster than the next scheduled renewal cycle. A single compromised node can corrupt colony-wide memory in minutes. By the time the renewal cycle runs, the contamination has been replicated to other nodes, incorporated into condensation, and used as evidence for further compilation.
Without fault isolation, the system has periodic garbage collection but no failure detector. It can reclaim resources from honestly-degraded state, but it cannot detect or contain a node that is producing arbitrarily incorrect outputs. This is the standard distinction in distributed systems between garbage collection and failure detection — both are required, and neither substitutes for the other.
7b.2 Operating Timescales
Section titled “7b.2 Operating Timescales”| Layer | Timescale | Mechanism |
|---|---|---|
| Layer 1: Invariant enforcement | Per-operation (milliseconds) | Synchronous validation before every write, spawn, and convention proposal commits |
| Layer 2: Anomaly detection | Per-batch (seconds to minutes) | Asynchronous analysis of accumulated telemetry against fault signature registry |
| Layer 3: Fault response | Minutes to hours | Graduated escalation from monitoring through fencing to rollback to reset |
| Provenance auditing | Scheduled (hours to days, adaptive) | Budget-constrained random sampling, tightening after incidents, loosening during stability |
7b.3 Distributed Systems Framing
Section titled “7b.3 Distributed Systems Framing”Fault isolation draws on established distributed systems concepts:
- Byzantine fault tolerance. The system assumes nodes may produce arbitrarily incorrect outputs — not just crash faults (fail-stop) but semantic faults (incorrect data with valid structure). The fault model is Byzantine, not crash-stop.
- Failure detectors. Layer 2 anomaly detection implements unreliable failure detectors — detectors that may produce false positives (healthy node flagged) or false negatives (faulty node missed). The protocol manages this tradeoff explicitly through quorum requirements and false positive tracking.
- Circuit breakers. The FENCE response is a circuit breaker — it isolates a suspect component from the shared system while allowing it to continue local operation. The circuit breaker trips on measured anomaly thresholds, not predictions.
- Fencing. Fenced nodes lose write access to shared state. This is the same fencing pattern used in distributed lock managers to prevent a node that may have lost its lease from corrupting shared resources.
- Checkpoint/restore. The ROLLBACK response reverts shared state to a pre-contamination checkpoint. Checkpoints are taken periodically and on fault events, providing known-good restore points.
- Probabilistic auditing. Cross-node provenance auditing samples shared state randomly and traces provenance chains. This is the same principle as random auditing in distributed ledgers — you cannot verify everything, but random sampling with sufficient frequency provides statistical confidence in state integrity.
- Token bucket rate limiting. Layer 1 rate limits use the token bucket algorithm — a fixed-capacity bucket that refills at a constant rate, permitting bursts up to bucket capacity but enforcing a sustained maximum rate.
- Optimistic concurrency control. Layer 1 conflict detection accepts writes optimistically (does not block on conflict) but logs every conflict for the anomaly detector. This is optimistic concurrency with conflict logging — the same pattern as OCC in database systems, adapted to log rather than abort.
7b.4 Failure Modes CDR Renewal Cannot Address
Section titled “7b.4 Failure Modes CDR Renewal Cannot Address”The following failure modes operate at timescales or through mechanisms that CDR renewal does not cover.
7b.4.1 State Poisoning
Section titled “7b.4.1 State Poisoning”DS concept: Byzantine fault.
A node writes fabricated or contradictory entries with structurally valid provenance. The entries pass schema validation. They enter the normal CDR pipeline — classification, storage, potential condensation, potential promotion to shared structural memory. Because the provenance chains look valid, the entries are indistinguishable from legitimate memory writes until downstream consumers encounter contradictions or trace the provenance to its fabricated source.
Why renewal is insufficient: Poisoned entries are replicated to other nodes via the colony state log before the next decay cycle. Once replicated, they become inputs to other nodes’ condensation and compilation processes. Renewal evaluates entries independently — it does not detect that a cluster of entries from the same node all reference fabricated sources.
7b.4.2 Rule Corruption
Section titled “7b.4.2 Rule Corruption”DS concept: Configuration drift / cascading failure.
A compiled routing rule actively misroutes entries after a domain shift. The rule is the routing table — it determines where entries go. A corrupted routing rule cannot be detected by the mechanism it governs, because the mechanism relies on the rule to function correctly.
Why renewal is insufficient: The rule’s monitoring metrics (Section 7, Stage 7) rely on downstream outcomes, but if the rule is misrouting entries to the wrong region, the downstream signals are also corrupted. The rule may appear healthy by its own metrics while causing harm to the broader system. This is the self-referential detection problem — a corrupted classifier cannot reliably detect its own corruption.
7b.4.3 Consensus Poisoning
Section titled “7b.4.3 Consensus Poisoning”DS concept: Byzantine agreement failure.
A faulty node proposes harmful conventions backed by fabricated trace evidence. The consensus protocol assumes honest reporting — nodes vote based on their own traces. But if the proposing node fabricates traces showing a pattern that does not exist, other nodes that lack contradicting data (because they operate in different domains) may validate the proposal based on the proposer’s evidence alone.
Why renewal is insufficient: The poisoned convention is installed through the legitimate consensus mechanism. It operates as a colony-wide rule. Renewal checks the rule’s accuracy — but if the rule governs a domain that only the proposing node operates in, other nodes have no independent traces to contradict it.
7b.4.4 Resource Exhaustion
Section titled “7b.4.4 Resource Exhaustion”DS concept: Fork bomb / unbounded allocation.
A node spawns excessively, exhausting the colony’s inference budget. The action cascade permits spawning — the mechanism operates correctly in each individual spawn decision, but the cumulative result exhausts shared resources. Alternatively, a node submits convention proposals at high volume, consuming other nodes’ evaluation budgets for validation.
Why renewal is insufficient: Resource exhaustion is not a state degradation problem. The entries and rules involved may be individually correct. The harm is in the aggregate resource consumption, which CDR’s per-entry lifecycle does not track.
7b.4.5 Trajectory Drift
Section titled “7b.4.5 Trajectory Drift”DS concept: Slow Byzantine fault.
Subtle cumulative degradation where each individual operation passes validation but the trajectory across a node’s output history is harmful. Example: a node’s provenance chains increasingly reference its own prior outputs rather than external sources, creating a self-reinforcing loop that drifts from ground truth. Each individual entry has valid provenance (it cites a real prior entry). The trajectory — increasing self-reference ratio — is the fault signal.
Why renewal is insufficient: CDR renewal evaluates entries independently. It has no trajectory view across a node’s output history. Each entry passes individual validation. The harmful pattern is only visible in the aggregate.
7b.5 The Three-Layer Model
Section titled “7b.5 The Three-Layer Model”Fault isolation operates through three layers, each at a different timescale and with different detection capabilities. A conforming node MUST implement Layer 1. Layers 2 and 3 are REQUIRED for colony participation and RECOMMENDED for single-node deployments.
Layer 1: Invariant Enforcement (Per-Operation Validation)
Section titled “Layer 1: Invariant Enforcement (Per-Operation Validation)”Timescale: Synchronous, per-operation (milliseconds). Runs before every write, spawn, and convention proposal commits.
Layer 1 checks are precondition validations — structural invariants that MUST hold on every operation. They are stateless, deterministic, and hardwired (never learned, never modified by the self-modification system). They correspond to input validation and schema enforcement at system boundaries.
7b.5.1 Provenance Chain Validation
Section titled “7b.5.1 Provenance Chain Validation”Every memory write MUST include a provenance chain terminating at a verifiable source. A verifiable source is one of:
- User input (direct human-provided content)
- Tool output (response from an external tool or API with a tool invocation ID)
- Another memory entry with its own valid provenance chain
Rejection criteria:
- Circular provenance references (entry A cites entry B, entry B cites entry A) → REJECT
- Provenance chain with a missing link (cited entry does not exist) → REJECT
- Provenance claiming a source that does not contain the claimed information → REJECT (when verifiable at write time; deferred to provenance auditing otherwise)
This is referential integrity enforcement — the same principle as foreign key constraints in relational databases. The provenance chain is a directed acyclic graph; cycles and dangling references are structural violations.
7b.5.2 Schema Enforcement
Section titled “7b.5.2 Schema Enforcement”All protocol artifacts MUST conform to their defined schemas:
| Artifact | Required Fields | Rejection Condition |
|---|---|---|
| Memory entry | id, region, created, source, confidence, last_renewed, renewal_count, decay_rate, content, provenance | Any required field missing or malformed |
| Convention proposal | id, proposing_node, proposed_date, rule, local_evidence, consensus_threshold | Missing evidence or rule definition |
| Telemetry event | event_class (decision / lifecycle / memory / eval / fault), timestamp, node_id | Unknown event_class or missing timestamp |
| Compiled rule | id, compiled_from, type, rule, confidence, evidence_summary, τ_star, status | Missing evidence summary or τ* |
Non-conforming operations MUST be rejected at the boundary. This is input validation — malformed data does not enter the system.
7b.5.3 Rate Limiting
Section titled “7b.5.3 Rate Limiting”Rate limits prevent resource exhaustion and consensus flooding. They use the token bucket algorithm: a bucket of capacity B refills at rate R per time window W. Operations consume one token. When the bucket is empty, operations are throttled (queued, not rejected) until tokens refill.
| Operation | Default Bucket (B) | Default Refill Rate (R) | Window (W) |
|---|---|---|---|
| Memory writes | 100 | 100 | per hour |
| Node spawns | 3 | 3 | per 24 hours |
| Convention proposals | 5 | 5 | per 7 days |
| Fault MONITOR escalations | 10 | 10 | per 24 hours |
Requirements:
- Rate limits are policy (configurable per deployment) but their existence is protocol (a conforming node MUST enforce rate limits)
- Exceeding a rate limit MUST be logged as a telemetry event (event_class:
fault, subtype:rate_limit_exceeded) - Sustained rate limit violations (>3 occurrences in a window) elevate the node’s anomaly score in Layer 2
7b.5.4 Conflict Detection
Section titled “7b.5.4 Conflict Detection”When a write contradicts an existing entry in the same region and domain:
- The write is accepted (optimistic concurrency — do not block on conflict)
- The conflict is logged with both entries’ IDs, the contradiction description, and the writing node’s ID
- The conflict is forwarded to Layer 2 anomaly detection
Anomaly signals from conflicts:
- A node whose writes consistently conflict with its own recent writes → elevated anomaly score (internal inconsistency)
- A node whose writes consistently conflict with other nodes’ entries in the same domain → highest anomaly signal (cross-node inconsistency)
- Conflict rate significantly above baseline for the node’s domain → anomaly flag
This is optimistic concurrency control with conflict logging. The system accepts writes to minimize latency but records every conflict as a potential fault signal.
Layer 2: Anomaly Detection (Learned Fault Signatures)
Section titled “Layer 2: Anomaly Detection (Learned Fault Signatures)”Timescale: Asynchronous, per-batch (seconds to minutes). Runs against accumulated telemetry, not individual operations.
Layer 2 implements failure detector patterns compiled from resolved incidents. The system maintains a registry of fault signatures — specific telemetry patterns that historically preceded confirmed failures. This is the same principle as anomaly-based intrusion detection in network security: build a model of normal behavior, flag deviations.
7b.5.5 Fault Signature Registry
Section titled “7b.5.5 Fault Signature Registry”Every resolved incident produces a detection signature. Signatures are stored in the fault signature registry with full provenance and a CDR-governed lifecycle.
Signature format:
fault_signature: id: SIG-{UUID-SHORT} name: descriptive_identifier description: human-readable explanation of the fault pattern detection_rule: metric: metric_name # from telemetry stream aggregation: trailing N operations | trailing duration | cumulative threshold: "> X" | "< X" | "rate_of_change > X" response: MONITOR | FENCE # default escalation level when matched derived_from: INC-{UUID-SHORT} # incident that produced this signature created: ISO-8601 last_validated: ISO-8601 last_matched: ISO-8601 # last time this signature fired (null if never) τ_star: duration # optimal lifespan — signature undergoes CDR renewal status: ACTIVE | ARCHIVEDExample signatures:
- id: SIG-001 name: self_referential_drift description: > Node's provenance chains increasingly reference its own prior outputs rather than external sources, indicating a self-reinforcing loop drifting from ground truth (trajectory drift). detection_rule: metric: self_reference_ratio # own_entries_cited / total_citations aggregation: trailing 20 writes threshold: "> 0.8" response: FENCE derived_from: INC-003 τ_star: 90d
- id: SIG-002 name: classification_rule_degradation description: > Compiled routing rule's eval trace accuracy drops below threshold, indicating the rule's underlying distribution has shifted. detection_rule: metric: eval_mismatch_rate aggregation: trailing 50 classifications threshold: "> 0.3" response: MONITOR derived_from: INC-007 τ_star: 60d
- id: SIG-003 name: consensus_flooding description: > Node submitting convention proposals at an abnormally high rate, consuming colony evaluation budget. detection_rule: metric: proposal_rate aggregation: trailing 7 days threshold: "> 3 standard deviations above colony mean" response: MONITOR derived_from: INC-012 τ_star: 120d
- id: SIG-004 name: provenance_fabrication description: > Provenance audit reveals a node's entries cite sources that do not contain the claimed information at a rate above baseline. detection_rule: metric: provenance_audit_failure_rate aggregation: trailing audit cycle threshold: "> 0.1" response: FENCE derived_from: INC-015 τ_star: 180d7b.5.6 Signature Learning Pipeline
Section titled “7b.5.6 Signature Learning Pipeline”Fault signatures are not hardwired — they are learned from resolved incidents and maintained under CDR lifecycle. The signature learning pipeline follows seven stages:
INCIDENT → FENCE → ROOT CAUSE → SIGNATURE → REGISTRY → PROPAGATE → RENEWALStage 1: INCIDENT. An anomaly is detected by Layer 1 invariant checks, a peer node’s report, or a human report. The incident is logged in the append-only incident log.
incident: id: INC-{UUID-SHORT} reported_date: ISO-8601 reported_by: node_id | human detection_source: invariant_check | peer_report | human_report | signature_match suspect_node: node_id (if applicable) suspect_rule: rule_id (if applicable) description: what was observed telemetry_snapshot: relevant telemetry from the detection window status: OPEN | FENCED | ROOT_CAUSED | RESOLVED | FALSE_POSITIVEStage 2: FENCE. The suspect node or rule is isolated from shared state (Layer 3 fault response). Fencing is precautionary — it limits blast radius while investigation proceeds.
Stage 3: ROOT CAUSE. Trace back through telemetry to identify the failure mechanism. This MAY require inference budget and SHOULD be prioritized above normal maintenance tasks.
- Identify the earliest telemetry signal that preceded the fault
- Identify the propagation path (which entries/rules were affected downstream)
- Classify the failure mode (state poisoning, rule corruption, consensus poisoning, resource exhaustion, trajectory drift)
Stage 4: SIGNATURE. Extract the detection pattern — the metric, aggregation window, and threshold that would have detected this fault earlier.
Requirements:
- The signature MUST be expressible as a metric + threshold over a telemetry window (not arbitrary code)
- The signature MUST be tested against historical telemetry to measure false positive rate
- If false positive rate exceeds 0.1 on historical data, the signature MUST be refined before registry entry
Stage 5: REGISTRY. Add the signature to the fault signature registry with provenance (the incident it was derived from), a computed τ*, and status ACTIVE.
Stage 6: PROPAGATE. Distribute the signature to all colony nodes via the colony coordination channel. Each node adds the signature to its local fault signature registry.
channel_message: type: FAULT_SIGNATURE_UPDATE payload: fault_signature record action: ADD | UPDATE | ARCHIVEStage 7: RENEWAL. Fault signatures undergo CDR renewal. A signature that has not matched any incident in 3 × τ* is archived (status: ARCHIVED). Archived signatures are retained in the registry for reference but are not evaluated against live telemetry.
Key principle: The fault signature registry does not grow without bound. Signatures compress (specific incident details → general detection rule), operate at their own timescale (each has a τ*), and are renewed or culled. The anomaly detector maintains itself under CDR, preventing detection overhead from scaling linearly with incident history.
Layer 3: Fault Response (Graduated Containment and Recovery)
Section titled “Layer 3: Fault Response (Graduated Containment and Recovery)”Timescale: Minutes to hours. Escalation is evidence-gated, not time-gated.
Detection without containment is monitoring, not fault tolerance. Layer 3 provides a graduated response cascade. Each escalation level uses a standard distributed systems pattern and has explicit reversibility semantics.
7b.5.7 Escalation Levels
Section titled “7b.5.7 Escalation Levels”| Level | DS Pattern | Trigger | Response | Reversibility |
|---|---|---|---|---|
| MONITOR | Health check | Anomaly score elevated but below FENCE threshold | Increase telemetry sampling rate for suspect node/rule. No capability reduction. | Fully reversible — sampling returns to baseline when anomaly score clears below threshold for a sustained period (configurable, default 72h). |
| FENCE | Circuit breaker / fencing | Anomaly score exceeds FENCE threshold OR fault signature matched with response=FENCE OR 2+ peer nodes flag the same target | Revoke node’s write access to shared state. Node continues local operation: read-only on shared memory, full read/write on local memory, no consensus participation. Node state transitions to FENCED. | Reversible — fence is time-bounded (configurable, default 24h). If no confirming evidence emerges, fence lifts automatically (lease-based timeout). Investigation may lift the fence earlier or escalate. |
| ROLLBACK | Checkpoint restore | Confirmed contamination of shared state. Specific entries/conventions identified as compromised. | Revert affected shared entries and/or conventions to pre-contamination checkpoint. Colony-wide notification listing affected entry IDs. All nodes re-materialize their local views from the corrected log. | Partially reversible — rolled-back entries are archived with status: REVERTED flag, not deleted. The rollback operation itself is logged in the colony state log. |
| EXCISE | Targeted state cleanup | Identified harmful entries/rules with known blast radius. More precise than rollback when contamination is localized. | Remove specific entries from shared state. Run a transitive dependency check: trace provenance forward from excised entries and flag anything downstream. | Reversible — excised entries are archived with status: EXCISED flag. Downstream entries flagged for review, not automatically excised. |
| RESET | Node restart / failover | Node confirmed as comprehensively compromised — local state cannot be trusted. | Wipe the node’s local state. Re-initialize from the latest known-good colony checkpoint + compressed parent state digest. The node re-enters INITIALIZING state. | Irreversible for the node’s local state — it is destroyed. Shared state contributions from the reset node are audited separately (provenance audit on all shared entries contributed by the node). |
7b.5.8 Escalation Principles
Section titled “7b.5.8 Escalation Principles”Graduated escalation. The response MUST always start at MONITOR. Escalation is driven by evidence accumulation, not suspicion. The system MUST NOT skip levels except when a Layer 1 invariant violation directly triggers FENCE (invariant violations are definitive, not probabilistic).
Audit trail. Every fault response action MUST be logged in the incident log with:
- The triggering evidence (anomaly score, signature match, peer report, invariant violation)
- The escalation level selected
- The justification (why this level and not a lower one)
- The affected scope (which node, entries, rules)
- The timestamp and the acting node or human
The fault isolation system itself MUST be transparent and auditable. A node under investigation MUST NOT be able to modify its own incident log entries (the incident log is append-only).
Minimize blast radius. The containment action MUST NOT cause more disruption than the fault it addresses. This ordering is mandatory:
- FENCE before ROLLBACK — isolate the source before touching shared state
- ROLLBACK before RESET — revert specific state before wiping a node
- EXCISE as a precision alternative to ROLLBACK when contamination scope is known
Human escalation gates.
| Level | Approval Requirement |
|---|---|
| MONITOR | Autonomous — within configured anomaly thresholds |
| FENCE | Autonomous — within configured thresholds. Requires 2+ nodes flagging OR invariant violation. |
| ROLLBACK | Human approval REQUIRED for shared structural/identity state. Autonomous for shared ephemeral/operational state if blast radius is below configurable entry count threshold (default: 10 entries). |
| EXCISE | Human approval REQUIRED. Targeted state removal demands human judgment on blast radius. |
| RESET | Human approval REQUIRED. Node state destruction is irreversible. |
7b.6 Colony-Level Fault Coordination
Section titled “7b.6 Colony-Level Fault Coordination”Individual nodes run Layer 1 invariant checks locally. Colony-level faults — a faulty node contaminating shared state, a corrupted convention affecting all nodes — require coordinated detection and response.
7b.6.1 Detection Quorum Rules
Section titled “7b.6.1 Detection Quorum Rules”Detection of faults across node boundaries follows quorum rules to prevent a single node from disrupting the colony through false accusations.
| Action | Quorum Requirement | Rationale |
|---|---|---|
| Raise MONITOR on another node | Unilateral (any single node) | MONITOR is observational and imposes no capability reduction. Low cost, low risk. |
| Raise FENCE on another node | 2+ nodes flagging the same target OR 1 node + invariant violation | Corroborated detection. Prevents a single faulty detector from isolating healthy nodes. |
| Initiate ROLLBACK | N/2+1 nodes agreeing on contamination scope | Majority quorum — same threshold as consensus for convention changes. Shared state modification requires shared agreement. |
| Initiate EXCISE | N/2+1 nodes + human approval | Targeted removal is more precise but more judgment-dependent than rollback. |
| Initiate RESET | Human approval (out-of-band escalation) | Node state destruction is irreversible and requires human judgment. |
7b.6.2 False Positive Mitigation
Section titled “7b.6.2 False Positive Mitigation”Aggressive fault detection produces false positives that disrupt healthy nodes. This is the completeness vs. accuracy tradeoff in failure detectors — detecting all faults (completeness) conflicts with never falsely accusing healthy nodes (accuracy). The protocol manages this tradeoff through five safeguards:
1. MONITOR is observational. The node under MONITOR continues operating with full capabilities. The only effect is increased telemetry sampling. False MONITOR flags waste some detection budget but cause no operational disruption.
2. FENCE is time-bounded (lease-based timeout). Every FENCE has a configurable maximum duration (default 24h). If no confirming evidence emerges within the lease period, the fence lifts automatically. This is the same pattern as distributed lock lease expiry — if the lock holder fails before releasing, the lock expires rather than being held forever.
3. Detector self-monitoring. A node that repeatedly raises false MONITOR alerts gets its own elevated anomaly score. The detector is itself monitored — the system tracks each node’s false positive rate for fault reports. Nodes with a historical false positive rate above threshold (configurable, default 0.3) have their MONITOR reports weighted lower (require corroboration even at MONITOR level).
4. No unilateral escalation beyond FENCE. A single node cannot trigger ROLLBACK, EXCISE, or RESET. This prevents a single faulty detector from causing colony-wide disruption.
5. Historical false positive tracking. The system maintains per-node and global false positive rates for fault detection. Detection thresholds are adjusted to maintain a target precision (configurable, default 0.9 — at most 10% of fault escalations should be false positives). This is adaptive threshold tuning — the same principle as adjusting alarm thresholds in monitoring systems to maintain a target signal-to-noise ratio.
7b.6.3 Cross-Node Provenance Auditing
Section titled “7b.6.3 Cross-Node Provenance Auditing”The strongest colony-level verification mechanism. Provenance auditing does not detect faults in real time — it verifies the integrity of shared state on a scheduled or on-demand basis.
Procedure:
-
Sample. Select a random sample of shared structural and shared identity entries. Sample size is determined by the audit budget (configurable fraction of inference budget, default 0.1).
-
Trace. For each sampled entry, trace the full provenance chain back to its terminal source (user input, tool output, or external entry).
-
Verify. At each link in the chain, verify:
- The cited source exists
- The cited source contains information consistent with the claim made by the citing entry
- The chain is acyclic (no circular references)
- All intermediate entries are still present (no orphaned references due to improper deletion)
-
Flag. Record audit results:
audit_result:entry_id: ENTRY-{UUID-SHORT}audit_date: ISO-8601chain_length: integerlinks_verified: integerlinks_failed: integerfailure_details: [{link_index, failure_type, description}]verdict: PASS | WARN | FAIL- PASS: all links verified
- WARN: chain intact but one or more links could not be fully verified (e.g., cited source is in glacier and not loaded)
- FAIL: broken chain, circular reference, or content mismatch
-
Aggregate. Aggregate results by contributing node. A node whose contributions show systematic provenance failures (failure rate significantly above the colony baseline for a given audit cycle) is a strong Byzantine signal.
Adaptive scheduling:
- During stable periods: audit runs on the scheduled interval (configurable, default every 14 days)
- After a fault incident: audit interval tightens (configurable, default 3 days) for a configurable recovery period (default 30 days)
- Audit budget is capped — if the sample size required for statistical confidence exceeds the budget, the system audits at the budgeted sample size and reports the achieved confidence level
Relationship to Layer 1 provenance validation: Layer 1 validates provenance structure at write time (referential integrity). Provenance auditing validates provenance semantics over time (does the cited source actually support the claim?). Structural validation catches malformed chains; semantic auditing catches fabricated or drifted content.
7b.7 Relationship to CDR
Section titled “7b.7 Relationship to CDR”| Aspect | CDR Renewal (Theorem 3) | Fault Isolation (This Subsystem) |
|---|---|---|
| Failure model | Gradual drift (fail-slow). Degradation is predictable, independent across entries. | Byzantine fault (fail-arbitrary). Disruption may be sudden, correlated, adversarial. |
| Timescale | Days to months (τ*) | Milliseconds (Layer 1) to hours (Layer 3) |
| Trigger | Scheduled review cycle or computed τ* threshold | Per-operation invariant checks, anomaly threshold breach, peer report |
| Response | Renew entry (rebuild from current evidence) or archive (remove from active set) | Monitor, fence, rollback, excise, or reset (graduated containment) |
| Scope | Individual entry or rule | Entry, rule, node, or colony-wide |
| Detection mechanism | Accuracy monitoring against eval traces | Invariant checks (structural), fault signatures (behavioral), provenance auditing (semantic) |
| Learning | Better decay rates and confidence thresholds (Section 7) | Fault signatures compiled from incidents (with CDR lifecycle) |
| DS analogue | Garbage collection, cache expiry | Failure detectors, circuit breakers, fencing, checkpoint/restore |
Why both are required: CDR renewal handles entropy — the natural, inevitable degradation of compressed state under distributional drift. Fault isolation handles disruption — sudden, correlated, potentially adversarial failures that propagate faster than renewal cycles. A system with only CDR renewal is a distributed system with garbage collection but no failure detector: it reclaims degraded resources but cannot detect or contain node failure. A system with only fault isolation but no CDR renewal would contain acute faults but never clean up honestly-degraded state, accumulating unbounded staleness debt. Together they cover the full failure spectrum.
7b.8 Node Contract Additions
Section titled “7b.8 Node Contract Additions”The fault isolation subsystem adds the following requirements to the node contract (Section 3):
Layer 1 (REQUIRED for all conforming nodes):
- Every write operation MUST pass provenance chain validation before commit
- Every protocol artifact MUST pass schema enforcement before commit
- Rate limits MUST be enforced on writes, spawns, and convention proposals
- Conflicts MUST be logged with both entries’ IDs and forwarded to anomaly detection
Layer 2 (REQUIRED for colony participation, RECOMMENDED for single-node):
- The node MUST maintain a local copy of the fault signature registry
- The node MUST evaluate accumulated telemetry against active signatures on a per-batch basis
- The node MUST accept and integrate fault signature updates from the colony coordination channel
Layer 3 (REQUIRED for colony participation):
- The node MUST respect FENCE directives (transition to FENCED state, revoke shared-state writes)
- The node MUST participate in detection quorums when requested by peer nodes
- The node MUST log all fault events to the append-only incident log
Telemetry additions:
- New event class:
fault— covering MONITOR, FENCE, ROLLBACK, EXCISE, and RESET events - Subtype:
rate_limit_exceeded,invariant_violation,signature_match,anomaly_threshold,provenance_audit_result - The fault event log is append-only and tamper-evident — a node under investigation MUST NOT be able to modify its own fault log entries
7b.9 Reference Implementation Structure
Section titled “7b.9 Reference Implementation Structure”The fault isolation subsystem is stored in a dedicated directory within the node structure:
fault-isolation/├── invariants.md ← Layer 1: hardwired validation rules│ (provenance checks, schema definitions,│ rate limit configurations, conflict detection rules)├── fault-signatures.md ← Layer 2: learned fault signatures with│ provenance, τ*, and lifecycle status├── incident-log.md ← Append-only record of all fault events│ (incidents, escalations, resolutions, false positives)└── audit-schedule.md ← Cross-node provenance audit schedule, results history, adaptive interval stateRequirements:
invariants.mdis hardwired — it MUST NOT be modified by the self-modification system (Section 7). Changes to invariant definitions require a protocol version change.fault-signatures.mdis evolvable — signatures are learned from incidents and undergo CDR lifecycle (Section 7b.5.6). The self-modification system MUST NOT modify signatures, but the fault isolation system maintains them independently under its own CDR lifecycle.incident-log.mdis append-only — entries MUST NOT be modified or deleted after creation. This is the audit trail for the fault isolation system itself.audit-schedule.mdtracks the adaptive audit interval (current interval, last tightening event, next scheduled audit, historical results summary).
7b.10 Pluggable Policies
Section titled “7b.10 Pluggable Policies”The following parameters are policy (configurable per deployment). Their existence and enforcement is protocol (a conforming node MUST have values for all of these and MUST enforce them).
fault_isolation_policy: # Layer 1: Rate limits rate_limits: writes_per_hour: 100 # token bucket capacity spawns_per_day: 3 proposals_per_week: 5 monitor_escalations_per_day: 10
# Layer 2: Anomaly detection anomaly_thresholds: fence_threshold: float # anomaly score above which FENCE triggers monitor_threshold: float # anomaly score above which MONITOR triggers false_positive_target: 0.9 # target precision for fault escalations
# Layer 3: Fault response fence_duration: 24h # lease-based timeout for FENCE monitor_clear_duration: 72h # sustained below-threshold period to clear MONITOR rollback_entry_threshold: 10 # autonomous rollback permitted below this entry count
# Provenance auditing audit_interval: 14d # baseline audit frequency audit_interval_post_incident: 3d # tightened frequency after incidents audit_recovery_period: 30d # how long tightened frequency persists audit_budget_fraction: 0.1 # fraction of inference budget allocated to auditing
# Detector self-monitoring detector_false_positive_threshold: 0.3 # above this, detector's reports require corroboration