Skip to content

Data Model

Status: Designed, not yet implemented. Object store and search exist from the build system; event log and policy engine are new.

The data model has four components: objects, an event log, reactive policies, and procedures.

Immutable, content-addressed, with provenance:

FieldTypeDescription
idhashContent hash (SHA256)
typeenumtranscript, episode, observation, pattern, policy, procedure, …
authorstringWhich agent or principal created this
sourcelist[hash]Input object IDs consumed
createdtimestampCreation time
regionenumephemeral, operational, structural, identity, glacier
confidencefloatDecays over time
contenttextPayload

Five regions, matched to information timescales: ephemeral (hours), operational (days–weeks), structural (months–quarters), identity (permanent until changed), glacier (archived, indexed, retrievable).

Append-only. Every mutation is a log entry:

seq_id | operation | object_id | author | causal_deps | timestamp

Operations: WRITE, ARCHIVE, PROMOTE, INVALIDATE, CONVENTION_CHANGE

Materialized views — search index, context documents, conventions directory — are derived from the log and can be rebuilt by replaying.

Tradeoff: Full replay capability requires retaining the complete log. Storage cost scales linearly with mutation count. Compaction strategies are not yet specified.

Governance is expressed as policy documents stored as objects:

# Policy: Episode Summarization
trigger: new object of type `transcript`
procedure: summarize_episode
author: any agent with role `synthesizer`
output_type: episode
output_region: operational
evidence_requirement: minimum 1 source object

An agent reads the policy, checks if the trigger condition is met, executes the referenced procedure, writes the result. Dependency structure is implicit in the trigger conditions rather than declared as a DAG.

Instead of Python transform classes, procedures are markdown documents:

# Procedure: Summarize Episode
Given one or more transcript objects, produce a single episode summary.
## Inputs
- 1+ objects of type `transcript`
## Prompt
[prompt template with {{placeholders}}]
## Output schema
- type: episode
- region: operational
- confidence: 0.8
- source: [input object ids]

The procedure is a versioned object in the store. Changing the procedure changes the behavior of any agent that reads it.

Open question: How reliably can LLM agents parse and execute arbitrary procedure documents? This is an empirical question the implementation needs to answer.

Objects carry confidence scores based on provenance:

SourceInitial confidence
Human-provided1.0
Raw ingestion0.9
Cross-agent corroborated0.7
Single-agent synthesis0.5

Confidence decays over time. The rate is region-specific. These initial values are defaults — whether they are well-calibrated is an empirical question.