Retain Leaderboard

Which model should I use for retain() and observation consolidation?

Results from the previous LoComo-based methodology are on the legacy leaderboard.|Want another model here, or spotted a deployment issue?Open a GitHub issue
Rank
Model
Provider
Total Score
Quality + Efficiency + Conformance + Speed + Cost
Quality
BEAM extraction accuracy
Reasoning
Effort tested
Serving
Context · hardware
Efficiency
Accuracy per stored token
Speed
Latency + Throughput
Cost
$ per 1M tokens
JSON Conformance
Valid JSON / total tests
1🏆
GPT-5.6 Luna (low reasoning)
OpenAIOpenAI
60.4
62.5
50% accuracy
low
Hindsight v0.9.2
51.8
152.1k tokens stored
47.2
11.2s · 340 tok/s
55.6
$0.20/$1.20
100.0
50/50 tests
2🥈
gemini-3.7-flash (high reasoning)
GoogleGoogle
59.4
59.5
49% accuracy
high
API provider
Hindsight v0.9.2
64.2
89.1k tokens stored
44.6
12.4s · 407 tok/s
27.6
$0.75/$3.75
100.0
50/50 tests
3🥉
Qwen3.6 35B-A3B
GCP (self-hosted)GCP (self-hosted)
59.4
59.5
49% accuracy
off
128K ctx
1× RTX PRO 6000
Hindsight v0.9.2
56.1
124.7k tokens stored
46.7
11.4s · 671 tok/s
60.6
$0.15/$1.00
96.0
48/50 tests
4
gpt-oss 20B (low reasoning)
GCP (self-hosted)GCP (self-hosted)
57.8
56.3
48% accuracy
low
128K ctx
1× RTX PRO 6000
Hindsight v0.9.2
42.6
209.2k tokens stored
66.4
5.1s · 520 tok/s
79.7
$0.08/$0.35
98.0
49/50 tests
5
Granite 4.2 3B
GCP (self-hosted)GCP (self-hosted)
56.4
56.3
48% accuracy
off
128K ctx
1× RTX PRO 6000
Hindsight v0.9.2
40.5
228.5k tokens stored
51.3
9.5s · 375 tok/s
93.2
$0.02/$0.11
96.0
48/50 tests
6
DeepSeek V4 Flash (low reasoning)
GCP (self-hosted)GCP (self-hosted)
53.8
56.3
48% accuracy
low
8× B200
Hindsight v0.9.2
56.6
119.0k tokens stored
18.0
45.7s · 307 tok/s
64.5
$0.22/$0.66
74.0
37/50 tests
7
Qwen3.8 27B (low reasoning)
GCP (self-hosted)GCP (self-hosted)
53.4
56.3
48% accuracy
low
128K ctx
1× RTX PRO 6000
Hindsight v0.9.2
51.7
145.3k tokens stored
24.6
30.7s · 233 tok/s
37.0
$0.42/$2.55
100.0
50/50 tests
8
Gemma 4 12B
GCP (self-hosted)GCP (self-hosted)
52.9
47.0
44% accuracy
off
256K ctx
1× RTX PRO 6000
Hindsight v0.9.2
50.0
143.3k tokens stored
59.8
6.7s · 330 tok/s
80.0
$0.10/$0.30
94.0
47/50 tests
9
gemini-3.7-flash (low reasoning)
GoogleGoogle
51.4
47.0
44% accuracy
low
API provider
Hindsight v0.9.2
60.8
92.3k tokens stored
46.3
11.6s · 421 tok/s
27.6
$0.75/$3.75
100.0
50/50 tests
10
Qwen3.8 27B
GCP (self-hosted)GCP (self-hosted)
49.9
47.0
44% accuracy
off
128K ctx
1× RTX PRO 6000
Hindsight v0.9.2
59.0
99.6k tokens stored
30.1
23.2s · 254 tok/s
37.0
$0.42/$2.55
100.0
50/50 tests
11
Qwen3.8 27B (medium reasoning)
GCP (self-hosted)GCP (self-hosted)
48.1
50.0
45% accuracy
medium
128K ctx
1× RTX PRO 6000
Hindsight v0.9.2
45.6
175.7k tokens stored
21.5
36.5s · 239 tok/s
37.0
$0.42/$2.55
100.0
50/50 tests
12
GPT-5.6 Luna (reasoning off)
OpenAIOpenAI
45.8
40.5
41% accuracy
off
Hindsight v0.9.2
42.4
182.8k tokens stored
52.4
9.1s · 380 tok/s
55.6
$0.20/$1.20
100.0
50/50 tests
13
Gemma 4 E4B
GCP (self-hosted)GCP (self-hosted)
43.4
37.5
40% accuracy
off
128K ctx
1× RTX PRO 6000
Hindsight v0.9.2
34.8
245.3k tokens stored
42.7
13.4s · 346 tok/s
93.5
$0.02/$0.10
100.0
50/50 tests
14
gpt-oss 20B (medium reasoning)
GCP (self-hosted)GCP (self-hosted)
42.3
40.5
41% accuracy
medium
128K ctx
1× RTX PRO 6000
Hindsight v0.9.2
37.5
224.7k tokens stored
15.3
55.2s · 391 tok/s
79.7
$0.08/$0.35
100.0
50/50 tests
15
Granite 4.2 8B
GCP (self-hosted)GCP (self-hosted)
35.4
25.0
35% accuracy
off
128K ctx
1× RTX PRO 6000
Hindsight v0.9.2
46.1
133.9k tokens stored
29.8
23.6s · 155 tok/s
90.9
$0.05/$0.10
74.0
37/50 tests

About This Benchmark

MetricValue shownHow it is measured
Quality% accuracyExtraction accuracy on a frozen subset of the BEAM long-term memory benchmark (ICLR 2026): 4 conversations from the 128K-token tier, 80 questions tagged across ten ability categories (temporal reasoning, knowledge update, abstention, and others). Hover the Quality cell for a model's per-ability breakdown. The model under test powers the retain step (fact extraction and schema structuring). The answer context is built from the extracted facts only. Source chunks are excluded on purpose: they contain the raw conversation, and including them lets a weak extractor score like a strong one. Production recall does return chunks, so this setting is harder than production; the number measures extraction quality, not end-user accuracy. Answer generation and judging use a fixed gemini-3.7-flash, which is also a ranked model on this board.
Efficiencyaccuracy / stored tokensQuality accuracy divided by the tokens the model wrote into memory during ingestion (accuracy points per 1,000 stored fact tokens). This counters verbose extraction: a model that copies the conversation near-verbatim into its facts would score well on facts-only accuracy but pays for the footprint here. Stored tokens also drive real recall and storage cost.
Speedlatency (s) · tok/sMean end-to-end latency per request (arithmetic average across all successful requests) and output throughput (tokens/second), measured during the fact-extraction benchmark. Results will vary depending on your network conditions, geographic proximity to the provider's servers, and server-side load at the time of testing.

Note: this benchmark does not enforce or simulate rate limits. Actual throughput may be lower in production depending on your subscription tier and the provider's rate-limiting policies.
Cost$ input / $ output per 1M tokensPublished list prices (USD per million tokens) for input and output tokens, as advertised by each provider at the time of testing. Prices may have changed since then. Subscription-based models are scored separately on value relative to their monthly fee. Local models have no per-token cost and always score 100 on this dimension.
JSON Conformancesuccess / total testsThe fraction of fact-extraction requests that returned valid JSON conforming to the required schema. A failed request is one that timed out, returned an HTTP error, or produced malformed / schema-invalid JSON. The canonical suite is 50 extraction tests at concurrency 4; rows measured under older conditions show their own test count and are being re-run.
Total Score0 – 100Weighted composite: Quality 60% + Efficiency 20% + Speed 10% + JSON Conformance 5% + Cost 5%. Each dimension is normalised to a 0–100 scale before weighting. Quality maps accuracy onto that scale with fixed anchors, 25% → 0and 65% → 100, so accuracy differences carry the weight the raw range would compress. Speed uses the formula100 × 10 / (10 + latency_s)so a 10-second response scores 50; Cost uses100 × 0.001 / (0.001 + cost_per_req)so only genuinely free models approach 100.