GRIP: State-Tracking Context Degradation
GRIP: State-Tracking Context Degradation
Section titled “GRIP: State-Tracking Context Degradation”Measuring when lossy memory compression outperforms raw context stuffing
Status: WIP — sweep data collected, compression crossover experiment next
Thesis
Section titled “Thesis”There exists a context length N* where a lossy compressed representation of information (with error rate R) outperforms stuffing the full raw context into the model’s window. N* is predictable from two independently measurable quantities:
- The stuffing degradation curve — accuracy as a function of context length for a fixed task
- The compression error rate R — the fraction of facts lost or corrupted by compression
When context stuffing accuracy drops below (1 − R), compression wins. The crossover point N* and its dependence on task type, model size, and compression quality is the core finding. For agent memory systems, this answers: when should I stop stuffing context and start compressing?
Why this matters
Section titled “Why this matters”Existing work establishes that context degradation is universal (Chroma Context Rot, 18 frontier models tested) and that multi-fact reasoning degrades faster than retrieval (BABILong, FLenQA). But nobody plots the crossover curve — the point where lossy compression outperforms raw stuffing — or frames compression error rate as a variable in that comparison.
The task
Section titled “The task”We use a minimal state-tracking language with three statement types operating over 26 single-letter variables (a–z) and 16 color values:
SET x = value — assignUNSET x — removeRULE IF x = value THEN y = value — persistent ruleQUERY x — what is x's current value?Statements are generated randomly (85% SET, 10% RULE IF, 5% UNSET) with uniform random variable and value selection. After generation, a random variable is chosen to query. Ground truth is computed by a deterministic interpreter that executes all statements and evaluates rules to a fixed point. No LLM in the scoring loop.
The mixed grammar makes this substantially harder than pure SET tracking. Rules create derived state — the value of a queried variable might not come from any SET statement, but from a rule chain that fires based on other variables’ values. UNSETs can invalidate rule preconditions. The model must track explicit assignments, rule activations, and state invalidations simultaneously.
Independent variables
Section titled “Independent variables”For the SET-only variant, we identified three factors that control task difficulty:
- C — context length (number of statements). Proxy for total tokens.
- T — target rate (fraction of statements that assign the queried variable). Controls how many competing writes the model must filter through.
- P — last position (the queried variable’s final assignment must occur at or before this fraction of the context). Controls where the answer is buried.
Results: 750-trial sweep on Claude Haiku 4.5
Section titled “Results: 750-trial sweep on Claude Haiku 4.5”We swept C × T × P with 5 trials per cell (750 total API calls, ~$3).
Degradation curves
Section titled “Degradation curves”
Key observations:
- C=100 (500 tokens): 99% accuracy. The task is trivially solvable at short context.
- C=500–1000 (2.5K–5K tokens): 45–64%. Rapid degradation, driven primarily by T.
- C=5000+ (25K+ tokens): 6–15%. Near-floor for most parameter combinations.
- T is the dominant factor. At any C, low T (few target writes) passes while high T (many competing writes) fails.
- P=5% (beginning) is most robust. The model attends to the start of context better than the middle, confirming the “lost in the middle” effect from Liu et al. 2023.
T × P heatmaps
Section titled “T × P heatmaps”
The heatmaps show the interaction between T and P:
- Low T + any P: mostly green (easy). Few competing writes means the model can find the answer regardless of position.
- High T + P=50%: deep red. Many competing writes with the answer in the middle is the hardest combination.
- High T + P=5%: still partially green. Even with heavy interference, the model retrieves values from the start of context.
What we learned so far
Section titled “What we learned so far”1. Variable-name disjointness eliminates retrieval interference
Section titled “1. Variable-name disjointness eliminates retrieval interference”Our earlier design used separate variable pools for signal and noise (e.g., signal uses permit, noise uses treasury). This made the task trivially solvable at all context sizes — the model just scans for the variable name mentioned in the query. Same-variable writes are essential for creating genuine interference.
2. Turn-boundary fragmentation is catastrophic
Section titled “2. Turn-boundary fragmentation is catastrophic”When signal facts are split across separate user/assistant conversation turns (with noise turns between them), chain reasoning breaks at ~274 tokens — not a context length effect, but a turn structure effect. Single-turn presentation is required to isolate context degradation from conversation structure artifacts.
3. The task scales via T × C interaction
Section titled “3. The task scales via T × C interaction”Context length alone doesn’t determine difficulty. The number of competing writes to the target variable (T × C) is what drives degradation. This means the benchmark can be calibrated for any model by adjusting T while holding C fixed.
Mixed grammar degradation curve
Section titled “Mixed grammar degradation curve”The SET-only variant showed the role of T and P clearly, but the task was too simple — the oracle state table (26 key-value pairs) is trivially small, so compression is never under real pressure. Adding RULE IF and UNSET creates a richer state-tracking task where the compressed representation itself requires reasoning to compute.

| Statements | ~Tokens | Accuracy (n=100) | 95% CI |
|---|---|---|---|
| 10–50 | 56–276 | 100% | — |
| 75 | 412 | 97% | ±3.3% |
| 100 | 548 | 94% | ±4.7% |
| 150 | 819 | 89% | ±6.1% |
| 200 | 1.1K | 91% | ±5.6% |
| 300 | 1.6K | 82% | ±7.5% |
| 500 | 2.7K | 79% | ±8.0% |
| 750 | 4.1K | 67% | ±9.2% |
| 1000 | 5.4K | 48% | ±9.8% |
| 1500 | 8.2K | 39% | ±9.6% |
| 2000 | 10.9K | 26% | ±8.6% |
| 3000 | 16.3K | 14% | ±6.8% |
| 4000 | 21.8K | 10% | ±5.9% |
| 5000 | 27.2K | 5% | ±4.3% |
1,600 total API calls across 16 context sizes. This is the task and the curve we’ll use for the compression crossover analysis. The degradation is driven by three interacting factors: more statements mean more competing variable writes (retrieval interference), more rules to evaluate (reasoning load), and more UNSETs that invalidate derived state (temporal tracking). A compressor for this task must reason about the rule graph — it cannot simply record the last SET per variable.
Compression crossover: SET-only variant
Section titled “Compression crossover: SET-only variant”For the simpler SET-only variant, the optimal compression is a state table — the most recent value per variable (26 entries). We evaluate three compressors that model distinct failure modes over this representation:
- Oracle — exact final state. The task-sufficient upper bound.
- Drop — oracle with entries removed at rate p_d. Models omission (facts forgotten).
- Substitute — oracle with values replaced at rate p_s. Models distortion (facts corrupted).
These are not intended to model realistic summarization systems. They are controlled loss channels that isolate the two most basic ways compression can fail: forgetting and misremembering.
Compressor degradation
Section titled “Compressor degradation”
Drop vs substitute behavior: The substitute compressor is more robust at low corruption rates — 100% accuracy through p_s=20%, because the model only needs q’s entry to be correct (probability 1−p_s). The drop compressor degrades earlier because a missing entry for q yields no answer at all. At high rates (50%+) both converge toward zero.
Crossover by target rate
Section titled “Crossover by target rate”
N* crossover table
Section titled “N* crossover table”| Compressor | Rate | Accuracy | T=0.004 | T=0.01 | T=0.04 | T=0.1 | T=0.2 |
|---|---|---|---|---|---|---|---|
| Drop | p_d=10% | 100% | 500 | 300 | 200 | 150 | 100 |
| Drop | p_d=20% | 60% | 5000 | 5000 | 500 | 400 | 300 |
| Drop | p_d=50% | 40% | 5000 | 5000 | 1000 | 1000 | 500 |
| Sub | p_s=10% | 100% | 500 | 300 | 200 | 150 | 100 |
| Sub | p_s=20% | 100% | 500 | 300 | 200 | 150 | 100 |
| Sub | p_s=30% | 80% | 1000 | 1000 | 350 | 200 | 200 |
| Sub | p_s=50% | 40% | 5000 | 5000 | 1000 | 1000 | 500 |
N* in SET statements. Multiply by ~5 for approximate token count.
Key finding: a substitute compressor that corrupts 20% of its state entries still achieves 100% accuracy — and outperforms stuffing beyond 150–500 statements (750–2500 tokens) depending on interference level. Even a drop compressor that loses 10% of entries beats stuffing at the same threshold. The crossover happens remarkably early.
Next steps
Section titled “Next steps”Model comparison
Section titled “Model comparison”Run the same stuffing + compression sweeps on frontier models (GPT-5-mini, Kimi K2, Gemini 2.5 Flash) to see how N* varies with model capability. Hypothesis: larger models have higher N*, meaning they benefit from stuffing longer before compression wins.
Partial compression (hybrid)
Section titled “Partial compression (hybrid)”Real memory systems don’t compress everything uniformly. The realistic scenario: old memories are compressed (lossy), recent memories are raw. Test hybrid prompts: compressed state for the first Z% of context + uncompressed raw statements after.
Value reuse rate F
Section titled “Value reuse rate F”Fourth independent variable: rate at which non-target statements reuse the target’s current value. High F creates value-level confounders beyond variable-name interference.
github.com/marklubin/grip-bench
References
Section titled “References”- FLenQA / Same Task, More Tokens (Levy et al., 2024) — gold standard methodology for isolating context effects
- RULER (Hsieh et al., 2024) — variable tracking task that inspired our design
- BABILong (Kuratov et al., 2024) — multi-hop reasoning in long context
- Lost in the Middle (Liu et al., 2023) — positional bias in retrieval
- Context Length Alone Hurts (2025) — length as causal degradation factor
- Chroma Context Rot (2025) — universal degradation across 18 models