Retain Leaderboard
Which model should I use for retain() and observation consolidation?
Results from the previous LoComo-based methodology are on the legacy leaderboard.|Want another model here, or spotted a deployment issue?Open a GitHub issue
Rank | Model | Provider | Total Score Quality + Efficiency + Conformance + Speed + Cost | Quality BEAM extraction accuracy | Reasoning Effort tested | Serving Context · hardware | Efficiency Accuracy per stored token | Speed Latency + Throughput | Cost $ per 1M tokens | JSON Conformance Valid JSON / total tests |
|---|---|---|---|---|---|---|---|---|---|---|
1🏆 | GPT-5.6 Luna (low reasoning) | OpenAI | 60.4 | 62.5 50% accuracy | low | Hindsight v0.9.2 | 51.8 152.1k tokens stored | 47.2 11.2s · 340 tok/s | 55.6 $0.20/$1.20 | 100.0 50/50 tests |
2🥈 | gemini-3.7-flash (high reasoning) | Google | 59.4 | 59.5 49% accuracy | high | API provider Hindsight v0.9.2 | 64.2 89.1k tokens stored | 44.6 12.4s · 407 tok/s | 27.6 $0.75/$3.75 | 100.0 50/50 tests |
3🥉 | Qwen3.6 35B-A3B | GCP (self-hosted) | 59.4 | 59.5 49% accuracy | off | 128K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 56.1 124.7k tokens stored | 46.7 11.4s · 671 tok/s | 60.6 $0.15/$1.00 | 96.0 48/50 tests |
4 | gpt-oss 20B (low reasoning) | GCP (self-hosted) | 57.8 | 56.3 48% accuracy | low | 128K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 42.6 209.2k tokens stored | 66.4 5.1s · 520 tok/s | 79.7 $0.08/$0.35 | 98.0 49/50 tests |
5 | Granite 4.2 3B | GCP (self-hosted) | 56.4 | 56.3 48% accuracy | off | 128K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 40.5 228.5k tokens stored | 51.3 9.5s · 375 tok/s | 93.2 $0.02/$0.11 | 96.0 48/50 tests |
6 | DeepSeek V4 Flash (low reasoning) | GCP (self-hosted) | 53.8 | 56.3 48% accuracy | low | 8× B200 Hindsight v0.9.2 | 56.6 119.0k tokens stored | 18.0 45.7s · 307 tok/s | 64.5 $0.22/$0.66 | 74.0 37/50 tests |
7 | Qwen3.8 27B (low reasoning) | GCP (self-hosted) | 53.4 | 56.3 48% accuracy | low | 128K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 51.7 145.3k tokens stored | 24.6 30.7s · 233 tok/s | 37.0 $0.42/$2.55 | 100.0 50/50 tests |
8 | Gemma 4 12B | GCP (self-hosted) | 52.9 | 47.0 44% accuracy | off | 256K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 50.0 143.3k tokens stored | 59.8 6.7s · 330 tok/s | 80.0 $0.10/$0.30 | 94.0 47/50 tests |
9 | gemini-3.7-flash (low reasoning) | Google | 51.4 | 47.0 44% accuracy | low | API provider Hindsight v0.9.2 | 60.8 92.3k tokens stored | 46.3 11.6s · 421 tok/s | 27.6 $0.75/$3.75 | 100.0 50/50 tests |
10 | Qwen3.8 27B | GCP (self-hosted) | 49.9 | 47.0 44% accuracy | off | 128K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 59.0 99.6k tokens stored | 30.1 23.2s · 254 tok/s | 37.0 $0.42/$2.55 | 100.0 50/50 tests |
11 | Qwen3.8 27B (medium reasoning) | GCP (self-hosted) | 48.1 | 50.0 45% accuracy | medium | 128K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 45.6 175.7k tokens stored | 21.5 36.5s · 239 tok/s | 37.0 $0.42/$2.55 | 100.0 50/50 tests |
12 | GPT-5.6 Luna (reasoning off) | OpenAI | 45.8 | 40.5 41% accuracy | off | Hindsight v0.9.2 | 42.4 182.8k tokens stored | 52.4 9.1s · 380 tok/s | 55.6 $0.20/$1.20 | 100.0 50/50 tests |
13 | Gemma 4 E4B | GCP (self-hosted) | 43.4 | 37.5 40% accuracy | off | 128K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 34.8 245.3k tokens stored | 42.7 13.4s · 346 tok/s | 93.5 $0.02/$0.10 | 100.0 50/50 tests |
14 | gpt-oss 20B (medium reasoning) | GCP (self-hosted) | 42.3 | 40.5 41% accuracy | medium | 128K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 37.5 224.7k tokens stored | 15.3 55.2s · 391 tok/s | 79.7 $0.08/$0.35 | 100.0 50/50 tests |
15 | Granite 4.2 8B | GCP (self-hosted) | 35.4 | 25.0 35% accuracy | off | 128K ctx 1× RTX PRO 6000 Hindsight v0.9.2 | 46.1 133.9k tokens stored | 29.8 23.6s · 155 tok/s | 90.9 $0.05/$0.10 | 74.0 37/50 tests |
About This Benchmark
| Metric | Value shown | How it is measured |
|---|---|---|
| Quality | % accuracy | Extraction accuracy on a frozen subset of the BEAM long-term memory benchmark (ICLR 2026): 4 conversations from the 128K-token tier, 80 questions tagged across ten ability categories (temporal reasoning, knowledge update, abstention, and others). Hover the Quality cell for a model's per-ability breakdown. The model under test powers the retain step (fact extraction and schema structuring). The answer context is built from the extracted facts only. Source chunks are excluded on purpose: they contain the raw conversation, and including them lets a weak extractor score like a strong one. Production recall does return chunks, so this setting is harder than production; the number measures extraction quality, not end-user accuracy. Answer generation and judging use a fixed gemini-3.7-flash, which is also a ranked model on this board. |
| Efficiency | accuracy / stored tokens | Quality accuracy divided by the tokens the model wrote into memory during ingestion (accuracy points per 1,000 stored fact tokens). This counters verbose extraction: a model that copies the conversation near-verbatim into its facts would score well on facts-only accuracy but pays for the footprint here. Stored tokens also drive real recall and storage cost. |
| Speed | latency (s) · tok/s | Mean end-to-end latency per request (arithmetic average across all successful requests) and output throughput (tokens/second), measured during the fact-extraction benchmark. Results will vary depending on your network conditions, geographic proximity to the provider's servers, and server-side load at the time of testing. Note: this benchmark does not enforce or simulate rate limits. Actual throughput may be lower in production depending on your subscription tier and the provider's rate-limiting policies. |
| Cost | $ input / $ output per 1M tokens | Published list prices (USD per million tokens) for input and output tokens, as advertised by each provider at the time of testing. Prices may have changed since then. Subscription-based models are scored separately on value relative to their monthly fee. Local models have no per-token cost and always score 100 on this dimension. |
| JSON Conformance | success / total tests | The fraction of fact-extraction requests that returned valid JSON conforming to the required schema. A failed request is one that timed out, returned an HTTP error, or produced malformed / schema-invalid JSON. The canonical suite is 50 extraction tests at concurrency 4; rows measured under older conditions show their own test count and are being re-run. |
| Total Score | 0 – 100 | Weighted composite: Quality 60% + Efficiency 20% + Speed 10% + JSON Conformance 5% + Cost 5%. Each dimension is normalised to a 0–100 scale before weighting. Quality maps accuracy onto that scale with fixed anchors, 25% → 0and 65% → 100, so accuracy differences carry the weight the raw range would compress. Speed uses the formula100 × 10 / (10 + latency_s)so a 10-second response scores 50; Cost uses100 × 0.001 / (0.001 + cost_per_req)so only genuinely free models approach 100. |


