Skip to content

GRIP: State-Tracking Context Degradation

Measuring when lossy memory compression outperforms raw context stuffing

Status: WIP — sweep data collected, compression crossover experiment next

There exists a context length N* where a lossy compressed representation of information (with error rate R) outperforms stuffing the full raw context into the model’s window. N* is predictable from two independently measurable quantities:

  1. The stuffing degradation curve — accuracy as a function of context length for a fixed task
  2. The compression error rate R — the fraction of facts lost or corrupted by compression

When context stuffing accuracy drops below (1 − R), compression wins. The crossover point N* and its dependence on task type, model size, and compression quality is the core finding. For agent memory systems, this answers: when should I stop stuffing context and start compressing?

Existing work establishes that context degradation is universal (Chroma Context Rot, 18 frontier models tested) and that multi-fact reasoning degrades faster than retrieval (BABILong, FLenQA). But nobody plots the crossover curve — the point where lossy compression outperforms raw stuffing — or frames compression error rate as a variable in that comparison.

We use a minimal state-tracking language with three statement types operating over 26 single-letter variables (a–z) and 16 color values:

SET x = value — assign
UNSET x — remove
RULE IF x = value THEN y = value — persistent rule
QUERY x — what is x's current value?

Statements are generated randomly (85% SET, 10% RULE IF, 5% UNSET) with uniform random variable and value selection. After generation, a random variable is chosen to query. Ground truth is computed by a deterministic interpreter that executes all statements and evaluates rules to a fixed point. No LLM in the scoring loop.

The mixed grammar makes this substantially harder than pure SET tracking. Rules create derived state — the value of a queried variable might not come from any SET statement, but from a rule chain that fires based on other variables’ values. UNSETs can invalidate rule preconditions. The model must track explicit assignments, rule activations, and state invalidations simultaneously.

For the SET-only variant, we identified three factors that control task difficulty:

  • C — context length (number of statements). Proxy for total tokens.
  • T — target rate (fraction of statements that assign the queried variable). Controls how many competing writes the model must filter through.
  • P — last position (the queried variable’s final assignment must occur at or before this fraction of the context). Controls where the answer is buried.

Results: 750-trial sweep on Claude Haiku 4.5

Section titled “Results: 750-trial sweep on Claude Haiku 4.5”

We swept C × T × P with 5 trials per cell (750 total API calls, ~$3).

Two charts showing accuracy degradation. Left: by target rate T, showing higher T degrades faster. Right: by position P, showing P=5% (beginning) is most robust.
Left: higher target rates degrade faster. Right: beginning positions (P=5%) are most robust — classic “lost in the middle” effect.

Key observations:

  • C=100 (500 tokens): 99% accuracy. The task is trivially solvable at short context.
  • C=500–1000 (2.5K–5K tokens): 45–64%. Rapid degradation, driven primarily by T.
  • C=5000+ (25K+ tokens): 6–15%. Near-floor for most parameter combinations.
  • T is the dominant factor. At any C, low T (few target writes) passes while high T (many competing writes) fails.
  • P=5% (beginning) is most robust. The model attends to the start of context better than the middle, confirming the “lost in the middle” effect from Liu et al. 2023.
Three heatmaps showing accuracy by T (target rate) × P (position) at C=500, C=1000, and C=5000. Clear gradient from top-right (hard) to bottom-left (easy).
Accuracy by target rate × position at three context sizes. Green = high accuracy, red = low. The gradient runs from easy (low T, early P) to hard (high T, middle P).

The heatmaps show the interaction between T and P:

  • Low T + any P: mostly green (easy). Few competing writes means the model can find the answer regardless of position.
  • High T + P=50%: deep red. Many competing writes with the answer in the middle is the hardest combination.
  • High T + P=5%: still partially green. Even with heavy interference, the model retrieves values from the start of context.

1. Variable-name disjointness eliminates retrieval interference

Section titled “1. Variable-name disjointness eliminates retrieval interference”

Our earlier design used separate variable pools for signal and noise (e.g., signal uses permit, noise uses treasury). This made the task trivially solvable at all context sizes — the model just scans for the variable name mentioned in the query. Same-variable writes are essential for creating genuine interference.

2. Turn-boundary fragmentation is catastrophic

Section titled “2. Turn-boundary fragmentation is catastrophic”

When signal facts are split across separate user/assistant conversation turns (with noise turns between them), chain reasoning breaks at ~274 tokens — not a context length effect, but a turn structure effect. Single-turn presentation is required to isolate context degradation from conversation structure artifacts.

Context length alone doesn’t determine difficulty. The number of competing writes to the target variable (T × C) is what drives degradation. This means the benchmark can be calibrated for any model by adjusting T while holding C fixed.

The SET-only variant showed the role of T and P clearly, but the task was too simple — the oracle state table (26 key-value pairs) is trivially small, so compression is never under real pressure. Adding RULE IF and UNSET creates a richer state-tracking task where the compressed representation itself requires reasoning to compute.

Degradation curve: 100% at 10 statements, drops to 3% at 5000 statements. Clean monotonic decline.
Mixed grammar degradation: 100 trials per point, 95% confidence intervals shown. The model tracks state perfectly up to ~50 statements, then degrades monotonically to 5% at 5000 statements (~27K tokens).
Statements~TokensAccuracy (n=100)95% CI
10–5056–276100%—
7541297%±3.3%
10054894%±4.7%
15081989%±6.1%
2001.1K91%±5.6%
3001.6K82%±7.5%
5002.7K79%±8.0%
7504.1K67%±9.2%
10005.4K48%±9.8%
15008.2K39%±9.6%
200010.9K26%±8.6%
300016.3K14%±6.8%
400021.8K10%±5.9%
500027.2K5%±4.3%

1,600 total API calls across 16 context sizes. This is the task and the curve we’ll use for the compression crossover analysis. The degradation is driven by three interacting factors: more statements mean more competing variable writes (retrieval interference), more rules to evaluate (reasoning load), and more UNSETs that invalidate derived state (temporal tracking). A compressor for this task must reason about the rule graph — it cannot simply record the last SET per variable.

For the simpler SET-only variant, the optimal compression is a state table — the most recent value per variable (26 entries). We evaluate three compressors that model distinct failure modes over this representation:

  1. Oracle — exact final state. The task-sufficient upper bound.
  2. Drop — oracle with entries removed at rate p_d. Models omission (facts forgotten).
  3. Substitute — oracle with values replaced at rate p_s. Models distortion (facts corrupted).

These are not intended to model realistic summarization systems. They are controlled loss channels that isolate the two most basic ways compression can fail: forgetting and misremembering.

Three panels: drop compressor crossover, substitute compressor crossover, and compressor degradation curves.
Left: drop compressor vs stuffing. Center: substitute compressor vs stuffing. Right: both compressors’ accuracy vs corruption rate, with oracle ceiling.

Drop vs substitute behavior: The substitute compressor is more robust at low corruption rates — 100% accuracy through p_s=20%, because the model only needs q’s entry to be correct (probability 1−p_s). The drop compressor degrades earlier because a missing entry for q yields no answer at all. At high rates (50%+) both converge toward zero.

Stuffing curves by T overlaid with drop and substitute compressor baselines, showing crossover points.
Per-T stuffing curves vs drop (red) and substitute (blue) compressor baselines. The harder the retrieval (higher T), the sooner compression wins.
CompressorRateAccuracyT=0.004T=0.01T=0.04T=0.1T=0.2
Dropp_d=10%100%500300200150100
Dropp_d=20%60%50005000500400300
Dropp_d=50%40%5000500010001000500
Subp_s=10%100%500300200150100
Subp_s=20%100%500300200150100
Subp_s=30%80%10001000350200200
Subp_s=50%40%5000500010001000500

N* in SET statements. Multiply by ~5 for approximate token count.

Key finding: a substitute compressor that corrupts 20% of its state entries still achieves 100% accuracy — and outperforms stuffing beyond 150–500 statements (750–2500 tokens) depending on interference level. Even a drop compressor that loses 10% of entries beats stuffing at the same threshold. The crossover happens remarkably early.

Run the same stuffing + compression sweeps on frontier models (GPT-5-mini, Kimi K2, Gemini 2.5 Flash) to see how N* varies with model capability. Hypothesis: larger models have higher N*, meaning they benefit from stuffing longer before compression wins.

Real memory systems don’t compress everything uniformly. The realistic scenario: old memories are compressed (lossy), recent memories are raw. Test hybrid prompts: compressed state for the first Z% of context + uncompressed raw statements after.

Fourth independent variable: rate at which non-target statements reuse the target’s current value. High F creates value-level confounders beyond variable-name interference.

github.com/marklubin/grip-bench