docs: update README benchmarks to current temporal-reasoning results (#6182)

This commit is contained in:
Kartik
2026-07-09 15:57:05 +05:30
committed by GitHub
parent b26469e006
commit 5dbf071356
+5 -5
View File
@@ -46,12 +46,12 @@
| Benchmark | Old | New | Tokens | Latency p50 | | Benchmark | Old | New | Tokens | Latency p50 |
| --- | --- | --- | --- | --- | | --- | --- | --- | --- | --- |
| **LoCoMo** | 71.4 | **91.6** | 7.0K | 0.88s | | **LoCoMo** | 71.4 | **92.5** | 7.0K | 0.88s |
| **LongMemEval** | 67.8 | **94.8** | 6.8K | 1.09s | | **LongMemEval** | 67.8 | **94.4** | 6.8K | 1.09s |
| **BEAM (1M)** | — | **64.1** | 6.7K | 1.00s | | **BEAM (1M)** | — | **64.1** | 6.7K | 1.00s |
| **BEAM (10M)** | — | **48.6** | 6.9K | 1.05s | | **BEAM (10M)** | — | **48.6** | 6.9K | 1.05s |
All benchmarks run on the same production-representative model stack. Single-pass retrieval (one call, no agentic loops). All benchmarks run on the same production-representative model stack. Single-pass retrieval (one call, no agentic loops) at a top_200 retrieval budget. Scores reflect Mem0's managed platform, which includes proprietary optimizations not available in the open-source SDK; open-source users should expect directionally similar gains but not identical numbers.
**What changed:** **What changed:**
- **Single-pass ADD-only extraction** -- one LLM call, no UPDATE/DELETE. Memories accumulate; nothing is overwritten. - **Single-pass ADD-only extraction** -- one LLM call, no UPDATE/DELETE. Memories accumulate; nothing is overwritten.
@@ -63,8 +63,8 @@ All benchmarks run on the same production-representative model stack. Single-pas
See the [migration guide](https://docs.mem0.ai/migration/oss-v2-to-v3) for upgrade instructions. The [evaluation framework](https://github.com/mem0ai/memory-benchmarks) is open-sourced so anyone can reproduce the numbers. See the [migration guide](https://docs.mem0.ai/migration/oss-v2-to-v3) for upgrade instructions. The [evaluation framework](https://github.com/mem0ai/memory-benchmarks) is open-sourced so anyone can reproduce the numbers.
## Research Highlights ## Research Highlights
- **91.6 on LoCoMo** -- +20 points over the previous algorithm - **92.5 on LoCoMo** -- +21 points over the previous algorithm
- **94.8 on LongMemEval** -- +27 points, with +53.6 on assistant memory recall - **94.4 on LongMemEval** -- +27 points, with 98.2 on assistant memory recall
- **64.1 on BEAM (1M)** -- production-scale memory evaluation at 1M tokens - **64.1 on BEAM (1M)** -- production-scale memory evaluation at 1M tokens
- [Read the full paper](https://mem0.ai/research) - [Read the full paper](https://mem0.ai/research)