Compare commits

..

56 Commits

Author SHA1 Message Date
David Shrader 6ddf1669f4 fix(vector_stores): point fastembed-missing warning at mem0ai[extras] (#5622) 2026-06-24 14:47:02 +05:30
David Shrader d3d2e89fd5 fix(llms): skip JSON response_format for Groq compound models (#5513) 2026-06-24 14:42:22 +05:30
Xiaoju 43175d85f2 fix: preserve empty Azure AI Search update values (#5524)
Signed-off-by: Xiaoju <xiaojuchh@gmail.com>
2026-06-24 14:41:55 +05:30
Kartik 661ecb9f0f docs(hermes): update for v3 API, OSS self-hosted mode, and new tools (#5807) 2026-06-24 11:59:13 +05:30
冯基魁 09f181c577 fix(server): fetch filtered dashboard memories beyond default page (#5753) 2026-06-24 11:49:06 +05:30
Zaid 3497f26a00 fix: add auto_refresh option for OpenSearch Serverless compatibility (#3893)
Co-authored-by: Zaid Malhis <Malhis@users.noreply.github.com>
2026-06-24 11:48:52 +05:30
Bartok e9c0547423 fix(ts-sdk): preserve customCategories names through key conversion (#5741) 2026-06-24 11:15:19 +05:30
Yash Singh 25bc1b7426 fix(memory): guard against malformed image_url in parse_vision_messages (#5631) 2026-06-24 11:10:36 +05:30
Hrushikesh Yadav fee344db85 fix(chroma): wrap scalar vector_id in list for delete() (#5703) 2026-06-24 10:54:36 +05:30
Bartok b611f69381 fix(chroma): wrap update() ids/embeddings/metadatas in lists (#5757) 2026-06-24 10:54:15 +05:30
Muhammad Furqan c2862831db fix(reranker): log reranking failures instead of swallowing them silently (#5717) 2026-06-24 10:44:37 +05:30
홍찬희 1678e682ee fix(ts-sdk): check message.role instead of content for system messages (#3921) 2026-06-23 17:18:00 +05:30
Bartok ced4af681f fix(claude-plugin): rerank auto-injected memory context by default (#5690) 2026-06-23 16:52:35 +05:30
Hrushikesh Yadav 565db27121 fix(milvus,baidu): sanitize filter values to prevent expression injection (#5746) 2026-06-23 16:51:05 +05:30
Hrushikesh Yadav c0ac9f81fa fix(milvus): wrap scalar vector_id in list for delete() (#5704) 2026-06-23 16:48:47 +05:30
Yash Singh 879c68555c fix(memory): return attributed_to from get/get_all/search (#5629) 2026-06-23 16:43:30 +05:30
Abhishek Chauhan c2e723352e fix(ts-oss): return attributedTo from get/search/getAll (#5675) 2026-06-23 16:42:57 +05:30
Hrushikesh Yadav 716f021df8 fix(pinecone): map all comparison operators in _create_filter() (#5707) 2026-06-23 16:38:28 +05:30
Jiangtian Feng 7fa996261d perf: batch BM25 sparse encoding in Qdrant insert (#5592) 2026-06-23 16:29:16 +05:30
Yash Singh 87bd2d91e0 fix(llms): preserve reasoning fields in base-to-provider config conversion (#5638) 2026-06-23 16:27:43 +05:30
Yash Singh 15a930dac2 fix(llms): pass configured anthropic_base_url to the Anthropic client (#5626) 2026-06-23 16:26:22 +05:30
Bartok fa9abc77a6 fix(ts-oss): honor configured baseURL in AnthropicLLM (#5740) 2026-06-23 16:25:51 +05:30
Hrushikesh Yadav 4e448269bc fix(mongodb): reject dict filter values to prevent NoSQL operator injection (#5748) 2026-06-22 18:09:01 +05:30
Hrushikesh Yadav 29d131f7aa fix(chroma): return None instead of {} from _generate_where_clause for empty filters (#5713) 2026-06-22 12:05:28 +05:30
Bartok 42fe129330 fix(opensearch): return [[]] from list() error path to honor list() contract (#5727) 2026-06-22 11:58:37 +05:30
Hrushikesh Yadav bd5996f41e fix(pinecone): return [[]] from list() error path instead of dict (#5706) 2026-06-22 11:57:52 +05:30
Bartok 299c423213 fix(faiss): return [[]] for uninitialized index to honor list() contract (#5725) 2026-06-22 11:57:06 +05:30
Hrushikesh Yadav ce0531a13e fix(mongodb): wrap list() return in outer list to match interface contract (#5729) 2026-06-22 11:54:16 +05:30
Lucas Kim 513b56159f fix(embeddings): forward embedding_dims to Titan V2 in AWS Bedrock embedder (#5671)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 11:49:23 +05:30
Yash Singh 1676b3168d fix(server): return 404/400 instead of 502 for not-found and invalid input (#5634) 2026-06-22 11:47:48 +05:30
Davide Leopardi 8a786bf72d fix(azure): stop mutating and corrupting caller messages in content rewrite (#5731) 2026-06-22 11:45:46 +05:30
Yash Singh a48f34cf77 fix(vector_stores): deep-copy Redis DEFAULT_FIELDS so instances keep distinct dims (#5633) 2026-06-22 11:29:25 +05:30
Hrushikesh Yadav 650b734b1b fix: reset() only drops history table, leaving stale messages (#5541) 2026-06-22 11:28:17 +05:30
youneshima 871a1de7d2 docs: fix add memory v3 behavior (#5694) 2026-06-21 12:18:58 -07:00
fran3cc e615cc66de docs: rebrand Keywords AI integration to Respan (#5098)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Kartik <kartik.labhshetwar@mem0.ai>
2026-06-20 23:10:15 +05:30
Harsh Vardhan Gupta ca86a164bd fix(pi-agent-plugin): resolve undici CVE-2026-9697 / CVE-2026-9678 (#5669) 2026-06-19 16:21:18 +05:30
Kartik 2ac3f3956a fix(cli): pass telemetry context via stdin instead of argv (#5668)
Co-authored-by: JunghwanNA <70629228+shaun0927@users.noreply.github.com>
2026-06-19 13:57:53 +05:30
Yash Raj Pandey 7a9f03af3f fix(llms,embeddings): repair HTTP proxy support (httpx>=0.28) and preserve proxies in LlmFactory (#5447)
Co-authored-by: kartik-mem0 <kartik.labhshetwar@mem0.ai>
2026-06-19 13:45:21 +05:30
Yash Singh 48f1d6f010 fix(reranker): clamp out-of-range LLM scores instead of mis-parsing them (#5635) 2026-06-19 12:49:50 +05:30
Yash Singh f0ccd99924 fix(vertex): pass required vectors arg in list and similarity search (#5627) 2026-06-19 12:44:48 +05:30
ly-wang19 c5971193a2 fix(vector_stores): return None from Redis.get() for missing IDs (#5625)
Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com>
Co-authored-by: Kartik <kartik.labhshetwar@mem0.ai>
2026-06-19 12:25:24 +05:30
Yash Singh ff53fd60b7 fix(graph): keep distinct entities that share a substring prefix (#5630) 2026-06-19 12:23:15 +05:30
Yash Singh ffa334537a fix(client): check HTTP status before parsing ping response in _validate_api_key (#5639) 2026-06-19 12:19:13 +05:30
Yash Singh bd7ce2c13c fix(vector_stores): drop stray print in Weaviate list_cols (#5637) 2026-06-19 12:07:10 +05:30
Yash Singh 5d767219ff fix(reranker): export all five rerankers from package root (#5636) 2026-06-19 12:06:12 +05:30
Haochen 6b744845c3 fix(oss-ts): preserve message roles in extraction input so assistant facts aren't attributed to the user (#5643) 2026-06-19 09:17:52 +05:30
Yash Singh 5b4478458b fix(server): return 404 not 500 for malformed api key id on revoke (#5640) 2026-06-18 17:45:05 +05:30
Bartok 1751e7bff9 fix(ts-oss): reject empty/blank messages in Memory.add() to prevent hallucinated memories (#5545) 2026-06-18 17:36:19 +05:30
Bartok 7ae6a8c36a fix(memory): guard entity embed_batch count mismatch in v3 add pipeline (#5604)
Co-authored-by: Kartik <kartik.labhshetwar@mem0.ai>
2026-06-18 17:22:44 +05:30
Harsh Vardhan Gupta 1dcee153b9 fix(deps): patch js-yaml, ai, python-dotenv vulnerabilities (CVE-2026-53550, CVE-2025-48985, CVE-2026-28684) (#5641)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-18 17:13:15 +05:30
Md. Mamun Hossain 4065e846f6 fix(ts-sdk): prevent hallucinated memories on empty messages payload … (#5613) 2026-06-18 17:01:02 +05:30
Hrushikesh Yadav 466249113c fix: async delete_all race condition corrupts entity store linked_memory_ids (#5553) 2026-06-18 16:57:03 +05:30
Alok Tripathi 3e2ae734e7 feat(embeddings): add native embed_batch to 5 embedders (LMStudio, Together, HuggingFace, VertexAI, GoogleGenAI) (#5609) 2026-06-18 16:46:31 +05:30
Harsh Vardhan Gupta 96b31c4bc0 fix(form-data): upgrade to >=4.0.6 across pnpm workspaces (CVE-2026-12143) (#5618)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 13:44:05 +05:30
Rocke Dong 0117d5838b fix(server): use 127.0.0.1 in dashboard healthcheck to avoid IPv6 localhost resolution (#5612) 2026-06-18 11:53:06 +05:30
Abhishek Chauhan 42a3b4043c fix(ts-sdk): preserve user metadata keys across the case-conversion round-trip (#5515) 2026-06-18 11:35:20 +05:30
143 changed files with 3434 additions and 300 deletions
+9
View File
@@ -5,6 +5,15 @@ All notable changes to `@mem0/cli` are documented here.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [0.2.9] — 2026-06-19
### Security
- Telemetry no longer passes the Mem0 API key to its child process via
command-line arguments. The context is now sent over stdin, so the key is no
longer visible in the process list (`ps`, `/proc/<pid>/cmdline`, Activity
Monitor). Fixes #4862.
## [0.2.8] — 2026-06-01
### Security
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@mem0/cli",
"version": "0.2.8",
"version": "0.2.9",
"description": "The official CLI for mem0 — the memory layer for AI agents",
"type": "module",
"bin": {
+5 -5
View File
@@ -145,11 +145,11 @@ export function captureEvent(
anonDistinctIdToAlias: anonIdToAlias,
};
const child = spawn(
process.execPath,
[SENDER_SCRIPT, JSON.stringify(context)],
{ detached: true, stdio: "ignore" },
);
const child = spawn(process.execPath, [SENDER_SCRIPT], {
detached: true,
stdio: ["pipe", "ignore", "ignore"],
});
child.stdin?.end(JSON.stringify(context));
child.unref();
} catch {
/* silently swallow */
+28 -2
View File
@@ -1,7 +1,8 @@
/**
* Standalone telemetry sender — runs as a detached child process.
*
* Usage: node telemetry-sender.cjs '<json context>'
* Usage: node telemetry-sender.cjs (JSON context is read from stdin; a single
* argv argument is still accepted as a legacy fallback)
*
* This script is spawned by telemetry.captureEvent() and runs independently
* of the parent CLI process. It:
@@ -19,6 +20,31 @@
const https = require("https");
const fs = require("fs");
function loadContext() {
return new Promise((resolve, reject) => {
if (process.argv[2]) {
try {
resolve(JSON.parse(process.argv[2]));
} catch (err) {
reject(err);
}
return;
}
let data = "";
process.stdin.setEncoding("utf8");
process.stdin.on("data", (chunk) => (data += chunk));
process.stdin.on("end", () => {
try {
resolve(JSON.parse(data));
} catch (err) {
reject(err);
}
});
process.stdin.on("error", reject);
});
}
function httpsRequest(url, method, headers, body) {
return new Promise((resolve, reject) => {
const u = new URL(url);
@@ -108,7 +134,7 @@ async function sendIdentifyEvent(ctx, payload, anonId) {
}
async function main() {
const ctx = JSON.parse(process.argv[2]);
const ctx = await loadContext();
const payload = ctx.payload;
if (ctx.needsEmail && ctx.mem0ApiKey) {
+59
View File
@@ -0,0 +1,59 @@
import { beforeEach, describe, expect, it, vi } from "vitest";
const mockLoadConfig = vi.fn();
const mockSaveConfig = vi.fn();
const mockSpawn = vi.fn();
vi.mock("../src/config.js", () => ({
CONFIG_FILE: "/tmp/mem0-config.json",
loadConfig: mockLoadConfig,
saveConfig: mockSaveConfig,
}));
vi.mock("node:child_process", () => ({
spawn: mockSpawn,
}));
describe("captureEvent", () => {
beforeEach(() => {
vi.resetModules();
mockLoadConfig.mockReset();
mockSaveConfig.mockReset();
mockSpawn.mockReset();
delete process.env.MEM0_TELEMETRY;
});
it("pipes the telemetry context through stdin instead of argv", async () => {
mockLoadConfig.mockReturnValue({
platform: {
apiKey: "m0-node-secret",
baseUrl: "https://api.mem0.ai",
userEmail: "",
},
telemetry: {
anonymousId: "cli-anon-node",
},
});
const stdin = { end: vi.fn() };
const child = { stdin, unref: vi.fn() };
mockSpawn.mockReturnValue(child);
const { captureEvent } = await import("../src/telemetry.js");
captureEvent("node_test_event", { case: "stdin-secret" });
expect(mockSpawn).toHaveBeenCalledTimes(1);
const [execPath, args, options] = mockSpawn.mock.calls[0];
expect(execPath).toBe(process.execPath);
expect(args).toHaveLength(1);
expect(String(args[0])).toContain("telemetry-sender.cjs");
expect(JSON.stringify(args)).not.toContain("m0-node-secret");
expect(options).toMatchObject({ detached: true, stdio: ["pipe", "ignore", "ignore"] });
expect(stdin.end).toHaveBeenCalledTimes(1);
const payload = JSON.parse(stdin.end.mock.calls[0][0]);
expect(payload.mem0ApiKey).toBe("m0-node-secret");
expect(payload.payload.event).toBe("node_test_event");
expect(child.unref).toHaveBeenCalledTimes(1);
});
});
+13
View File
@@ -5,6 +5,19 @@ All notable changes to `mem0-cli` (Python) are documented here.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [0.2.8] — 2026-06-19
### Security
- Telemetry no longer passes the Mem0 API key to its child process via
command-line arguments. The context is now sent over stdin, so the key is no
longer visible in the process list (`ps`, `/proc/<pid>/cmdline`, Activity
Monitor). Fixes #4862.
### Fixed
- `__version__` now matches the packaged version (was stale at 0.2.4).
## [0.2.7] — 2026-05-20
### Added
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "mem0-cli"
version = "0.2.7"
version = "0.2.8"
description = "The official CLI for mem0 — the memory layer for AI agents"
readme = "README.md"
license = "Apache-2.0"
+1 -1
View File
@@ -1,3 +1,3 @@
"""mem0 CLI — the command-line interface for the mem0 memory layer."""
__version__ = "0.2.4"
__version__ = "0.2.8"
+9 -2
View File
@@ -137,12 +137,19 @@ def capture_event(
"anon_distinct_id_to_alias": anon_id_to_alias,
}
subprocess.Popen(
[sys.executable, "-m", "mem0_cli.telemetry_sender", json.dumps(context)],
child = subprocess.Popen(
[sys.executable, "-m", "mem0_cli.telemetry_sender"],
stdin=subprocess.PIPE,
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
start_new_session=True,
close_fds=True,
text=True,
)
if child.stdin:
with contextlib.suppress(Exception):
child.stdin.write(json.dumps(context))
with contextlib.suppress(Exception):
child.stdin.close()
except Exception:
pass
+13 -2
View File
@@ -1,6 +1,7 @@
"""Standalone telemetry sender — runs as a detached subprocess.
Usage: python -m mem0_cli.telemetry_sender '<json context>'
Usage: python -m mem0_cli.telemetry_sender (JSON context is read from stdin;
a single argv argument is still accepted as a legacy fallback)
This module is spawned by telemetry.capture_event() and runs independently
of the parent CLI process. It:
@@ -20,8 +21,18 @@ import sys
import urllib.request
def _load_context() -> dict:
"""Load telemetry context from stdin, falling back to argv for compatibility."""
raw = ""
if not sys.stdin.isatty():
raw = sys.stdin.read().strip()
if not raw and len(sys.argv) > 1:
raw = sys.argv[1]
return json.loads(raw)
def main() -> None:
ctx = json.loads(sys.argv[1])
ctx = _load_context()
payload = ctx["payload"]
if ctx.get("needs_email") and ctx.get("mem0_api_key"):
+80
View File
@@ -0,0 +1,80 @@
"""Tests for telemetry subprocess secret handling."""
from __future__ import annotations
import io
import json
import subprocess
import sys
from mem0_cli.config import Mem0Config, save_config
from mem0_cli.telemetry import capture_event
from mem0_cli.telemetry_sender import _load_context
class _CaptureStdin:
def __init__(self):
self.buffer = ""
self.closed = False
def write(self, value: str) -> None:
self.buffer += value
def close(self) -> None:
self.closed = True
class _DummyProcess:
def __init__(self):
self.stdin = _CaptureStdin()
def test_capture_event_writes_context_to_stdin_not_argv(isolate_config, monkeypatch):
config = Mem0Config()
config.platform.api_key = "m0-test-secret"
config.telemetry.anonymous_id = "cli-anon-test"
save_config(config)
captured: dict[str, object] = {}
proc = _DummyProcess()
def fake_popen(args, **kwargs):
captured["args"] = args
captured["kwargs"] = kwargs
return proc
monkeypatch.setattr("mem0_cli.telemetry.subprocess.Popen", fake_popen)
capture_event("unit_test_event", {"case": "stdin-secret"})
argv = captured["args"]
assert argv == [sys.executable, "-m", "mem0_cli.telemetry_sender"]
assert all("m0-test-secret" not in arg for arg in argv)
kwargs = captured["kwargs"]
assert kwargs["stdin"] == subprocess.PIPE
assert kwargs["text"] is True
ctx = json.loads(proc.stdin.buffer)
assert ctx["mem0_api_key"] == "m0-test-secret"
assert ctx["payload"]["event"] == "unit_test_event"
assert proc.stdin.closed
def test_load_context_reads_from_stdin(monkeypatch):
monkeypatch.setattr("sys.argv", ["telemetry_sender"])
monkeypatch.setattr("sys.stdin", io.StringIO('{"payload": {"event": "stdin"}}'))
ctx = _load_context()
assert ctx["payload"]["event"] == "stdin"
def test_load_context_falls_back_to_argv(monkeypatch):
monkeypatch.setattr("sys.argv", ["telemetry_sender", '{"payload": {"event": "argv"}}'])
monkeypatch.setattr("sys.stdin", io.StringIO(""))
ctx = _load_context()
assert ctx["payload"]["event"] == "argv"
@@ -56,6 +56,30 @@ config = {
}
```
### Configuration Options
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `collection_name` | string | required | Name of the OpenSearch index |
| `host` | string | required | OpenSearch endpoint URL |
| `port` | int | 9200 | Port number |
| `http_auth` | object | None | Authentication credentials (e.g., AWSV4SignerAuth) |
| `embedding_model_dims` | int | 1536 | Dimension of embedding vectors |
| `use_ssl` | bool | False | Enable SSL/TLS connection |
| `verify_certs` | bool | False | Verify SSL certificates |
| `auto_refresh` | bool | False | Automatically refresh index after insert. OpenSearch refreshes every ~1 second by default, so this is rarely needed. |
<Note>
The defaults above match a local OpenSearch instance. The AWS OpenSearch Serverless
example earlier on this page intentionally overrides them with `port=443`, `use_ssl=True`,
and `verify_certs=True`, which are required when connecting to a Serverless collection.
</Note>
<Note>
For **AWS OpenSearch Serverless**, keep `auto_refresh=False` (the default).
The `indices.refresh()` API is not supported on Serverless collections.
</Note>
### Add Memories
```python
+11 -11
View File
@@ -21,7 +21,7 @@ Adding memory is how Mem0 captures useful details from a conversation so your ag
- **Messages** – The ordered list of user/assistant turns you send to `add`.
- **Infer** – Controls whether Mem0 extracts structured memories (`infer=True`, default) or stores raw messages.
- **Metadata** – Optional filters (e.g., `{"category": "movie_recommendations"}`) that improve retrieval later.
- **User / Session identifiers** – `user_id`, `agent_id`, or `run_id` that scope the memory for future searches.
- **User / Session identifiers** – `user_id`, `agent_id`, `app_id`, or `run_id` that scope the memory for future searches.
## How does it work?
@@ -30,22 +30,22 @@ Mem0 offers two flows:
- **Mem0 Platform** – Fully managed API with dashboard and scaling.
- **Mem0 Open Source** – Local SDK that you run in your own environment.
Both flows take the same payload and pass it through the same pipeline.
Both flows take the same payload and add memories through an additive pipeline.
<Steps>
<Step title="Information extraction">
Mem0 sends the messages through an LLM that pulls out key facts, decisions, or preferences to remember.
</Step>
<Step title="Conflict resolution">
Existing memories are checked for duplicates or contradictions so the latest truth wins.
<Step title="Additive storage">
New memories are added without overwriting or deleting existing memories.
</Step>
<Step title="Storage">
The resulting memories land in managed vector storage so future searches return them quickly.
<Step title="Retrieval">
Future searches rank the most relevant memories for the query.
</Step>
</Steps>
<Warning>
Duplicate protection only runs during that conflict-resolution step when you let Mem0 infer memories (`infer=True`, the default). If you switch to `infer=False`, Mem0 stores your payload exactly as provided, so duplicates will land. Mixing both modes for the same fact will save it twice.
When you switch to `infer=False`, Mem0 stores your payload exactly as provided, so duplicates can land. Mixing both modes for the same fact can save it twice.
</Warning>
You trigger this pipeline with a single `add` call—no manual orchestration needed.
@@ -80,13 +80,13 @@ const messages = [
];
await client.add(messages, {
user_id: "alice",
userId: "alice",
});
```
</CodeGroup>
<Info icon="check">
Expect a `memory_id` (or list of IDs) in the response. Check the Mem0 dashboard to confirm the new entry under the correct user.
Expect a `status: "PENDING"` response with an `event_id`. Poll `GET /v1/event/{event_id}/` to confirm completion.
</Info>
## Add with Mem0 Open Source
@@ -138,7 +138,7 @@ const result = memory.add(messages, {
</Tip>
<Warning>
If you do choose `infer=False`, keep it consistent. Raw inserts skip conflict resolution, so a later `infer=True` call with the same content will create a second memory instead of updating the first.
If you do choose `infer=False`, keep it consistent. Raw inserts skip inference, so a later `infer=True` call with the same content can create a second memory.
</Warning>
## When Should You Add Memory?
@@ -167,7 +167,7 @@ For full list of supported fields, required formats, and advanced options, see t
| Capability | Mem0 Platform | Mem0 OSS |
| --- | --- | --- |
| Conflict resolution | Automatic with dashboard visibility | SDK handles merges locally; you control storage |
| Add behavior | ADD-only; memories accumulate | ADD-only; you control storage |
| Rate limits | Managed quotas per workspace | Limited by your hardware and provider APIs |
| Dashboard visibility | Yes — inspect memories visually | Inspect via CLI, logs, or custom UI |
+1 -1
View File
@@ -441,7 +441,7 @@
"integrations/flowise",
"integrations/langchain-tools",
"integrations/agentops",
"integrations/keywords",
"integrations/respan",
"integrations/raycast"
]
}
+6 -4
View File
@@ -309,19 +309,21 @@ Here are the available integrations for Mem0:
</Card>
<Card
title="Keywords AI"
title="Respan"
icon={
<svg
xmlns="http://www.w3.org/2000/svg"
width="24"
height="24"
viewBox="0 0 24 24"
viewBox="0 0 200 200"
fill="none"
>
<path fill-rule="evenodd" clip-rule="evenodd" d="M9.07513 1.1863C9.21663 1.07722 9.39144 1.01009 9.56624 1.01009C9.83261 1.01009 10.0823 1.12756 10.2405 1.33734L15.0101 7.4964V12.4136L16.4335 13.8401C16.7582 14.1673 16.7582 14.7043 16.4335 15.0316C16.1089 15.3588 15.5762 15.3588 15.2515 15.0316L13.3453 13.1016V8.07538L8.92529 2.36944V2.36105C8.64228 2.00024 8.70887 1.4716 9.07513 1.1863ZM18.976 14.4133C18.8344 14.3778 18.7003 14.3042 18.5894 14.1925L16.9163 12.5059C16.7249 12.3129 16.6416 12.0528 16.6749 11.8094V6.88385H16.6499L11.8553 0.691225C11.7282 0.529117 11.6716 0.333133 11.6803 0.140562C11.134 0.0481292 10.5726 0 10 0C4.47715 0 0 4.47715 0 10C0 15.5228 4.47715 20 10 20C13.9387 20 17.3456 17.7229 18.976 14.4133Z" fill="currentColor"></path>
<path d="M2.00635 190.234V9.76584H53.3558V29.5101H26.7223V170.562H53.3558V190.234H2.00635Z" fill="currentColor"></path>
<path d="M120.692 160.902C116.383 160.902 112.691 159.387 109.612 156.357C106.535 153.327 105.02 149.633 105.067 145.277C105.02 141.016 106.535 137.37 109.612 134.34C112.691 131.309 116.383 129.794 120.692 129.794C124.859 129.794 128.481 131.309 131.559 134.34C134.684 137.37 136.27 141.016 136.317 145.277C136.27 148.166 135.512 150.793 134.045 153.161C132.624 155.528 130.73 157.422 128.362 158.842C126.042 160.216 123.486 160.902 120.692 160.902Z" fill="currentColor"></path>
<path d="M197.993 9.76584V190.234H146.643V170.562H173.278V29.5101H146.643V9.76584H197.993Z" fill="currentColor"></path>
</svg>
}
href="/integrations/keywords"
href="/integrations/respan"
>
Build AI applications with persistent memory and comprehensive LLM observability.
</Card>
+164 -37
View File
@@ -1,35 +1,42 @@
---
title: Hermes Agent
description: "Add long-term memory to Hermes agents using Mem0 as a pluggable memory provider with automatic background sync and zero-latency prefetch."
description: "Add long-term memory to Hermes agents with Mem0, on managed Mem0 Cloud or fully self-hosted (OSS), with automatic background sync and zero-latency prefetch."
---
Add long-term memory to [Hermes Agent](https://github.com/NousResearch/hermes-agent) — a self-improving AI agent CLI by Nous Research. Hermes has a pluggable memory system, and Mem0 is one of the supported providers. Once enabled, Mem0 automatically learns facts from your conversations and surfaces relevant ones before each turn — all without slowing down the chat.
Add long-term memory to [Hermes Agent](https://github.com/NousResearch/hermes-agent), a self-improving AI agent CLI by Nous Research. Hermes has a pluggable memory system, and Mem0 is one of the supported providers. Once enabled, Mem0 learns facts from your conversations and surfaces relevant ones before each turn, without slowing down the chat.
## Overview
You can run Mem0 in two ways:
Hermes runs a built-in memory system (file-based `MEMORY.md` and `USER.md`) alongside one external provider. When Mem0 is active, it works additively with the built-in system at three key moments in every conversation turn:
- **Platform mode** (default): managed Mem0 Cloud. Add your API key and you are ready.
- **OSS mode**: fully self-hosted with your own LLM, embedder, and vector store. No data leaves your machine.
### 1. Before the Agent Responds (Prefetch)
## How It Works
When you send a message, Hermes checks if it already has cached Mem0 search results from the previous turn. If so, those memories are injected into the system prompt so the LLM can see them. This is **zero-latency** — no waiting for an API call.
Hermes runs a built-in memory system (file-based `MEMORY.md` and `USER.md`) alongside one external provider. When Mem0 is active, it works additively with the built-in system at three points in every conversation turn.
### 2. After the Agent Responds (Sync)
### 1. Before the agent responds (prefetch)
Once the LLM finishes responding, Hermes sends the `(user message, assistant response)` pair to Mem0's API in a **background thread**. Mem0's server-side LLM automatically extracts facts (e.g., "user prefers Python", "user works at Acme Corp") — you don't have to tell it what to remember.
When you send a message, Hermes checks for cached Mem0 search results from the previous turn. If they exist, those memories are injected into the system prompt so the model can see them. This is zero-latency, with no waiting on an API call.
### 3. Background Prefetch for Next Turn
### 2. After the agent responds (sync)
At the same time as sync, Hermes kicks off a background search on Mem0 to pre-load relevant memories for the next turn. By the time you type your next message, the memories are already cached.
Once the model finishes, Hermes sends the `(user message, assistant response)` pair to Mem0 in a background thread. Mem0 extracts facts automatically (for example, "user prefers Python" or "user works at Acme Corp"), so you never have to tell it what to remember. Each write is tagged with the gateway channel it came from.
### 3. Background prefetch for the next turn
At the same time, Hermes runs a background search to pre-load relevant memories for your next message. By the time you type, the results are already cached.
## Agent Tools
When Mem0 is active, the LLM gets three extra tools it can call during conversations:
When Mem0 is active, the model gets five tools it can call during a conversation:
| Tool | Description |
|------|-------------|
| `mem0_profile` | Fetch all stored memories about the user |
| `mem0_search` | Semantic search through memories (supports optional reranking via `rerank` and `top_k` parameters) |
| `mem0_conclude` | Store a specific fact verbatim — uses `infer=False` so no server-side LLM extraction happens |
| Tool | Description | Parameters |
|------|-------------|------------|
| `mem0_list` | List all stored memories, for a full overview | `page`, `page_size` (default 100, max 200) |
| `mem0_search` | Semantic search by meaning, ranked by relevance | `query` (required), `top_k` (default 10, max 50), `rerank` (default `true`, Platform mode only) |
| `mem0_add` | Store a fact verbatim, with no LLM extraction | `content` (required) |
| `mem0_update` | Update a memory's text by ID | `memory_id`, `text` (both required) |
| `mem0_delete` | Delete a memory by ID | `memory_id` (required) |
## Installation
@@ -40,17 +47,19 @@ curl -fsSL https://raw.githubusercontent.com/NousResearch/hermes-agent/main/scri
source ~/.bashrc
```
The `mem0ai` Python package is automatically installed when you enable the Mem0 provider — no manual pip install needed.
The `mem0ai` package is installed automatically when you enable the Mem0 provider, so there is no manual pip step. OSS providers may need extra packages (for example `qdrant-client`, `psycopg2-binary`, or `ollama`), which the setup flow installs for you when you pick them.
## Setup
## Platform Setup
### Option 1: Interactive Setup Wizard (Recommended)
Platform mode uses managed Mem0 Cloud and is the fastest way to start.
### Option 1: Interactive wizard (recommended)
```bash
hermes memory setup
```
Select **mem0** as the provider and enter your Mem0 API key when prompted. The wizard writes your config to `~/.hermes/mem0.json`.
Select **mem0**, choose **Platform**, and paste your API key when prompted. The wizard writes the non-secret settings to `~/.hermes/mem0.json` and keeps the key in `~/.hermes/.env`.
<Note>Get your API key from <a href="https://app.mem0.ai?utm_source=oss&utm_medium=integration-hermes" rel="nofollow">app.mem0.ai</a>.</Note>
@@ -68,33 +77,151 @@ memory:
provider: mem0
```
That's it — Mem0 runs automatically from this point.
That's it. Mem0 runs automatically from here.
## Configuration Options
## OSS (Self-Hosted) Setup
Configuration is stored in `~/.hermes/mem0.json`. Values can also be set via environment variables.
OSS mode runs Mem0 entirely on your own infrastructure: your LLM, your embedder, and your vector store. No data is sent to Mem0 Cloud, and no Mem0 API key is required.
| Key | Env Variable | Default | Description |
|-----|-------------|---------|-------------|
| `api_key` | `MEM0_API_KEY` | — | **Required.** Mem0 Platform API key |
| `user_id` | `MEM0_USER_ID` | `hermes-user` | User identifier for scoping memories |
| `agent_id` | `MEM0_AGENT_ID` | `hermes` | Agent identifier |
| `rerank` | — | `true` | Enable reranking for memory recall |
### Interactive
```bash
hermes memory setup
# Select "mem0", then "Open Source (self-hosted)"
# Follow the prompts for LLM, embedder, and vector store
```
### With flags
```bash
hermes memory setup mem0 --mode oss \
--oss-llm openai --oss-llm-key sk-... \
--oss-vector qdrant
```
### Supported providers
| Component | Providers |
|-----------|-----------|
| LLM | `openai` (default model `gpt-5-mini`), `ollama` (local, default `llama3.1:8b`) |
| Embedder | `openai` (default `text-embedding-3-small`), `ollama` (local, default `nomic-embed-text`) |
| Vector store | `qdrant` (local path or server), `pgvector` |
### Flag reference
| Flag | Description |
|------|-------------|
| `--mode` | `platform` or `oss` |
| `--oss-llm` | LLM provider (`openai` or `ollama`, default `openai`) |
| `--oss-llm-key` | LLM API key (for `openai`) |
| `--oss-llm-model` | Override the LLM model |
| `--oss-llm-url` | LLM base URL (for `ollama` or a custom endpoint) |
| `--oss-embedder` | Embedder provider (default `openai`) |
| `--oss-embedder-key` | Embedder API key |
| `--oss-vector` | Vector store (`qdrant` or `pgvector`, default `qdrant`) |
| `--oss-vector-path` | Local Qdrant storage path |
| `--oss-vector-host`, `--oss-vector-port` | PGVector or remote Qdrant host and port |
| `--oss-vector-user`, `--oss-vector-password`, `--oss-vector-dbname` | PGVector connection details |
| `--user-id` | Canonical user identifier |
| `--dry-run` | Preview the resolved config without writing it |
## Switching Modes
You can move between Platform and OSS at any time. Run the setup command again, or edit `~/.hermes/mem0.json` directly.
```bash
# Platform to OSS
hermes memory setup mem0 --mode oss --oss-llm-key sk-...
# OSS to Platform
hermes memory setup mem0 --mode platform --api-key sk-...
# Preview without writing anything
hermes memory setup mem0 --mode oss --oss-llm-key sk-... --dry-run
```
A self-hosted `~/.hermes/mem0.json` looks like this:
```json
{
"mode": "oss",
"oss": {
"llm": {"provider": "openai", "config": {"model": "gpt-5-mini"}},
"embedder": {"provider": "openai", "config": {"model": "text-embedding-3-small"}},
"vector_store": {"provider": "qdrant", "config": {"path": "~/.hermes/mem0_qdrant"}}
}
}
```
## Configuration
Behavioral settings live in `~/.hermes/mem0.json` and are written for you by `hermes memory setup`. Only the secret `MEM0_API_KEY` belongs in `~/.hermes/.env`.
| Key | Default | Description |
|-----|---------|-------------|
| `mode` | `platform` | `platform` (Mem0 Cloud) or `oss` (self-hosted) |
| `api_key` | none | Mem0 Platform API key, required in Platform mode. Stored in `.env` as `MEM0_API_KEY` |
| `user_id` | `hermes-user` | Identifier that scopes memories. See cross-channel behavior below |
| `agent_id` | `hermes` | Agent identifier attached to writes |
| `rerank` | `true` | Rerank search results for relevance (Platform mode only) |
### Cross-channel memories
Hermes can run from the CLI and from gateways like Telegram, Slack, and Discord. The `user_id` setting controls how memories are scoped across them:
- **Set a `user_id`** and it applies to every gateway, so one person gets a single merged memory store no matter where they talk to the agent.
- **Leave it unset** (or at the default `hermes-user`) and each gateway uses its own native id, keeping per-platform memories separate.
Either way, every write is tagged with `metadata.channel` (for example `telegram` or `cli`), so per-channel views are still possible at query time.
## Reliability
- **Circuit Breaker** — If Mem0's API fails 5 times in a row, Hermes stops calling it for 2 minutes, then retries. The agent keeps working fine without memory during that time.
- **Non-blocking** — All Mem0 API calls happen in background daemon threads. A slow or failed API call never blocks your conversation.
- **Thread-safe** — The Mem0 client uses lazy initialization with locking, safe for concurrent access.
- **Circuit breaker**: if Mem0 fails five times in a row, Hermes pauses calls for two minutes, then retries. The agent keeps working without memory during that window. Expected client errors, like a 404 on a missing memory id, do not count toward tripping the breaker.
- **Non-blocking**: every Mem0 call runs in a background daemon thread, so a slow or failed call never blocks your conversation.
- **Thread-safe**: the client uses lazy initialization with locking, and the background sync and prefetch threads are guarded so concurrent gateway messages cannot produce duplicate memories.
## Troubleshooting
### "Mem0 temporarily unavailable"
The circuit breaker tripped after five consecutive failures and resets after two minutes.
- **Platform mode**: check your API key and internet connection.
- **OSS mode**: make sure your vector store (Qdrant or PGVector) is running and reachable.
### OSS: vector store connection refused
```bash
# Local Qdrant: confirm the storage path is writable
ls -la ~/.hermes/mem0_qdrant
# Qdrant server: confirm it is reachable
curl http://localhost:6333/healthz
# PGVector: confirm PostgreSQL is accepting connections
pg_isready -h localhost -p 5432
```
### OSS: Ollama not reachable
```bash
curl http://localhost:11434/api/tags
```
### Memories not appearing
- `mem0_add` stores text verbatim with no extraction. Ordinary conversation turns are extracted automatically by the background sync.
- Search is semantic, so try a broader query.
- Confirm `user_id` is the same across sessions (check `~/.hermes/mem0.json`).
## Key Features
1. **Zero-Latency Recall** — Memories are prefetched in the background and cached, ready before you type
2. **Server-side Extraction** — Mem0's API automatically extracts and deduplicates facts from each exchange
3. **Non-blocking** — All API calls run in background daemon threads
4. **Fault Tolerant** — Circuit breaker ensures the agent works even if Mem0 is temporarily unreachable
5. **Additive Memory** — Works alongside Hermes' built-in file-based memory system (MEMORY.md, USER.md)
1. **Two ways to run**: managed Platform or fully self-hosted OSS, switchable at any time.
2. **Zero-latency recall**: memories are prefetched in the background and cached before you type.
3. **Automatic extraction**: Mem0 extracts and deduplicates facts from each exchange for you.
4. **Non-blocking and fault tolerant**: background threads plus a circuit breaker keep the agent responsive even when Mem0 is unreachable.
5. **Additive memory**: works alongside Hermes' built-in file memory (`MEMORY.md`, `USER.md`).
<CardGroup cols={2}>
<Card title="OpenClaw Integration" icon={<svg width="24" height="24" viewBox="0 0 500 500" fill="none" xmlns="http://www.w3.org/2000/svg"><path fill-rule="evenodd" d="m153.5 173.5q24.62 1.46 46 13.5 12.11 8.1 17.5 21.5 0.74 2.45 0.5 5 0.09 0.81 1 1 1.48-4.9 1-10 5.04 10.48 1.5 22-9.81 27.86-35.5 42.5-26.17 14.97-56 19.5-2.77-0.4-2 1 2.86 1.27 6 1 25.64 1.53 48.5-10 0.34 10.08 2 20 1.08 5.76 5 10 1 1.5 0 3-31.11 20.84-68.5 17.5-23.7-5.7-32.5-28.5-4.39-9.18-3.5-19 15.41 6.23 32 4.5-20.68-6.39-39-18-34.81-27.22-12.5-65.5 11.84-14.83 29-23 4.21 7.66 11.5 12.5 3 1 6 0-26.04-34.62-29-78-0.13-8.46 2-16.5 1 6.5 2 13 3.43 39.53 24.5 73 2.03 2.28 4.5 4 0.5-1.25 1-2.5-1.27-6.54-5-12 0.5-0.75 1-1.5 9.72-3.43 20-4 0.55 10.34 8 17.5 1.94 0.74 4 0.5-17.8-64.6 16.5-122 0.98-1.79 1.5 0-28.21 56.64-13.5 118 1.08 1.43 2.5 0.5 2.21-4.98 2-10.5z" fill="currentColor"/><path fill-rule="evenodd" d="m454.5 97.5q-1.33 11.18-8.5 20-21.81 26.28-55.5 32-1.11-0.2-2 0.5 2.31 2.82 5.5 4.5 1 2 0 4-9.56 11.3-19.5 20 19.71-8.72 31-27 2.68-0.43 5 1-14.24 30.97-48 36.5-9.93 1.71-20 1.5-6.8-0.48-13 1 5.81 6.92 14 11-10.78 16.03-27 26.5 27.16-7.4 38-33.5 4.34 1.35 9 1-9.08 23.84-33 33.5-18.45 6.41-38 7 22.59 8.92 45-1 12.05-5.52 24-11 9.01-1.79 17 2.5 5.28-4.38 11-8 12.8-6.07 27-5 0 0.5 0 1-19.34 2.69-34 15.5 0.5 0.25 1 0.5 17.79-8.09 36-15 2.71-0.79 5-2 2.5-1 5-2 5.53-4.04 11-8 11.7-4.18 24-6.5 7.78-1.36 15 1.5-2.97 18.45-13.5 34-34.92 49.37-94.5 62.5-59.27 12.45-108-23-15.53-12.52-21.5-31.5-2.47-14.26 4-27-3.15 24.41 14 42-4.92-10.28-7-22-1.97-17.63 7-33 47.28-69.5 125.5-100 15.86-3.42 32-5.5 18.63-1.47 37 1.5z" fill="currentColor"/><path fill-rule="evenodd" d="m231.5 238.5q1.31-0.2 2 1-3.13 28.62 15 51-16.25 6.75-27-7.5-1-1-2 0 14.73 29.34 46 18.5 1.79 0.52 0 1.5-37.63 16.82-50.5-22.5-5.1-26.48 16.5-42z" fill="currentColor"/><path fill-rule="evenodd" d="m203.5 266.5q1.31-0.2 2 1-2.48 22.08 12 39-6.99 1.35-14 0.5 4.59 4.08 10 7-8.71 0.28-14.5-6.5-16.98-22.76 4.5-41z" fill="currentColor"/><path fill-rule="evenodd" d="m58.5 284.5q9.6-2.17 14.5 6 5.15 14.18-1 28-11.05-13.14-27.5-17.5 5.15-9.9 14-16.5z" fill="currentColor"/><path fill-rule="evenodd" d="m56.5 313.5q3.43 5.43 8 10-4.88 0.44-8 4-1.11-0.2-2 0.5 28.91 1.65 38 28.5 0.45 3.16-1 6-11.02-7.01-23-12.5-4.75-3.75-9.5-7.5 1.47 7.42 7 13 8.34 27.18 32 43 0.99 2.41-1.5 3.5-40.25 5.58-66.5-25.5-15.67-22.01-8-48 10.46-23.87 34.5-15z" fill="currentColor"/><path fill-rule="evenodd" d="m198.5 319.5q1.44 0.68 2.5 2 2.41 8.23 6 16 1.2 2.64-0.5 5-30.65 21.41-68 18.5-25.16-6.17-32.5-30.5 6.96 4.99 15.5 6.5 8.99 0.75 18 0.5 16.25 2.38 32-2.5 15.9-3.94 27-15.5z" fill="currentColor"/><path fill-rule="evenodd" d="m239.5 342.5q7.02-0.25 14 0.5 4.46 1.06 8 3.5-5.2 2.35-10 5.5-3.88 4.65-9 7.5-9.89-3.09-9.5-13 2.36-3.63 6.5-4z" fill="currentColor"/><path fill-rule="evenodd" d="m214.5 349.5q5.96 7.2 13.5 13 1 1 0 2-28.58 23.34-65.5 20.5-18.15-4.24-27.5-19.5 1.13 0.94 2.5 1.5 14.7 1.42 29-1.5 26.57-0.52 48-16z" fill="currentColor"/><path fill-rule="evenodd" d="m302.5 373.5q0.21 2.44-2 3.5-28.69 7.6-50.5-12.5-0.06-6.71 6.5-9 4.45-0.75 9-1 22.26 2.27 37 19z" fill="currentColor"/><path fill-rule="evenodd" d="m232.5 365.5q17.6 6.19 10.5 23-10.6 10.42-25.5 11.5-25.94 3.21-49-9 36.75-1.65 64-25.5z" fill="currentColor"/><path fill-rule="evenodd" d="m113.5 367.5q7.7-0.01 9.5 7-9.69 7.19-18.5 15.5-7.23 5.76-5.5-3.5 3.12-12.84 14.5-19z" fill="currentColor"/><path fill-rule="evenodd" d="m126.5 380.5q7.88-0.4 12 6.5-8.5 7.25-17 14.5-5.62-12.55 5-21z" fill="currentColor"/><path fill-rule="evenodd" d="m283.5 385.5q3.22 2.95 7 5.5 2.8 4.03 6 7.5 0.42 2.77-2 4-15.5-9.75-31-19.5-1.79-0.98 0-1.5 9.96 2.49 20 4z" fill="currentColor"/></svg>} href="/integrations/openclaw">
@@ -1,22 +1,22 @@
---
title: Keywords AI
description: "Combine Mem0 persistent memory with Keywords AI observability for tracked, cost-optimized AI applications."
title: Respan
description: "Combine Mem0 persistent memory with Respan observability for tracked, cost-optimized AI applications."
---
Build AI applications with persistent memory and comprehensive LLM observability by integrating Mem0 with Keywords AI.
Build AI applications with persistent memory and comprehensive LLM observability by integrating Mem0 with Respan.
## Overview
Mem0 is a self-improving memory layer for LLM applications, enabling personalized AI experiences that save costs and delight users. Keywords AI provides complete LLM observability.
Mem0 is a self-improving memory layer for LLM applications, enabling personalized AI experiences that save costs and delight users. Respan (formerly Keywords AI) provides complete LLM observability.
Combining Mem0 with Keywords AI allows you to:
Combining Mem0 with Respan allows you to:
1. Add persistent memory to your AI applications
2. Track interactions across sessions
3. Monitor memory usage and retrieval with Keywords AI observability
3. Monitor memory usage and retrieval with Respan observability
4. Optimize token usage and reduce costs
<Note>
You can get your Mem0 API key from the <a href="https://app.mem0.ai/?utm_source=oss&utm_medium=integration-keywords" rel="nofollow">Mem0 dashboard</a>.
You can get your Mem0 API key from the <a href="https://app.mem0.ai/?utm_source=oss&utm_medium=integration-respan" rel="nofollow">Mem0 dashboard</a>.
</Note>
## Setup and Configuration
@@ -24,7 +24,7 @@ You can get your Mem0 API key from the <a href="https://app.mem0.ai/?utm_source=
Install the necessary libraries:
```bash
pip install mem0ai keywordsai-sdk
pip install mem0ai openai
```
Set up your environment variables:
@@ -34,13 +34,13 @@ import os
# Set your API keys
os.environ["MEM0_API_KEY"] = "your-mem0-api-key"
os.environ["KEYWORDSAI_API_KEY"] = "your-keywords-api-key"
os.environ["KEYWORDSAI_BASE_URL"] = "https://api.keywordsai.co/api/"
os.environ["RESPAN_API_KEY"] = "your-respan-api-key"
os.environ["RESPAN_BASE_URL"] = "https://api.respan.ai/api/"
```
## Basic Integration Example
Here's a simple example of using Mem0 with Keywords AI:
Here's a simple example of using Mem0 with Respan:
```python
from mem0 import Memory
@@ -48,17 +48,17 @@ import os
# Configuration
api_key = os.getenv("MEM0_API_KEY")
keywordsai_api_key = os.getenv("KEYWORDSAI_API_KEY")
base_url = os.getenv("KEYWORDSAI_BASE_URL") # "https://api.keywordsai.co/api/"
respan_api_key = os.getenv("RESPAN_API_KEY")
base_url = os.getenv("RESPAN_BASE_URL") # "https://api.respan.ai/api/"
# Set up Mem0 with Keywords AI as the LLM provider
# Set up Mem0 with Respan as the LLM provider
config = {
"llm": {
"provider": "openai",
"config": {
"model": "gpt-5-mini",
"temperature": 0.0,
"api_key": keywordsai_api_key,
"api_key": respan_api_key,
"openai_base_url": base_url,
},
}
@@ -79,7 +79,7 @@ print(result)
## Advanced Integration with OpenAI SDK
For more advanced use cases, you can integrate Keywords AI with Mem0 through the OpenAI SDK:
For more advanced use cases, you can integrate Respan with Mem0 through the OpenAI SDK:
```python
from openai import OpenAI
@@ -88,8 +88,8 @@ import json
# Initialize client
client = OpenAI(
api_key=os.environ.get("KEYWORDSAI_API_KEY"),
base_url=os.environ.get("KEYWORDSAI_BASE_URL"),
api_key=os.environ.get("RESPAN_API_KEY"),
base_url=os.environ.get("RESPAN_BASE_URL"),
)
# Sample conversation messages
@@ -118,18 +118,18 @@ response = client.chat.completions.create(
print(json.dumps(response.model_dump(), indent=4))
```
For detailed information on this integration, refer to the official [Keywords AI Mem0 integration documentation](https://docs.keywordsai.co/integration/development-frameworks/mem0).
For detailed information on this integration, refer to the official [Respan Mem0 integration documentation](https://www.respan.ai/docs/integrations/mem0).
## Key Features
1. **Memory Integration**: Store and retrieve relevant information from past interactions
2. **LLM Observability**: Track memory usage and retrieval patterns with Keywords AI
2. **LLM Observability**: Track memory usage and retrieval patterns with Respan
3. **Session Persistence**: Maintain context across multiple user sessions
4. **Cost Optimization**: Reduce token usage through efficient memory retrieval
## Conclusion
Integrating Mem0 with Keywords AI provides a powerful combination for building AI applications with persistent memory and comprehensive observability. This integration enables more personalized user experiences while providing insights into your application's memory usage.
Integrating Mem0 with Respan provides a powerful combination for building AI applications with persistent memory and comprehensive observability. This integration enables more personalized user experiences while providing insights into your application's memory usage.
<CardGroup cols={2}>
<Card title="OpenAI Agents SDK" icon="cube" href="/integrations/openai-agents-sdk">
@@ -139,4 +139,3 @@ Integrating Mem0 with Keywords AI provides a powerful combination for building A
Monitor agent performance with AgentOps
</Card>
</CardGroup>
+1 -1
View File
@@ -286,7 +286,7 @@ If the user is on a pre-current major (Python < 2, TS < 3, or Platform `output_f
- [Dify](https://docs.mem0.ai/integrations/dify) [Both]: Use when the user is on Dify LLMOps.
- [Flowise](https://docs.mem0.ai/integrations/flowise) [Both]: Use when the user is on Flowise no-code.
- [AgentOps](https://docs.mem0.ai/integrations/agentops) [Both]: Use when tracking agent observability with memory metadata.
- [Keywords AI](https://docs.mem0.ai/integrations/keywords) [Both]: Use when monitoring with Keywords AI.
- [Respan](https://docs.mem0.ai/integrations/respan) [Both]: Use when monitoring Mem0 with Respan (formerly Keywords AI) LLM observability.
- [Raycast](https://docs.mem0.ai/integrations/raycast) [Both]: Use when the user wants quick memory access via Raycast.
## Cookbooks
+1 -1
View File
@@ -23,7 +23,7 @@
"@types/js-cookie": "^3.0.6",
"@types/react-syntax-highlighter": "^15.5.13",
"@types/uuid": "^10.0.0",
"ai": "^4.1.46",
"ai": "^5.0.52",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"js-cookie": "^3.0.6",
+1 -1
View File
@@ -18,7 +18,7 @@
"@radix-ui/react-scroll-area": "^1.2.0",
"@radix-ui/react-select": "^2.1.2",
"@radix-ui/react-slot": "^1.1.0",
"ai": "4.1.42",
"ai": "^5.0.52",
"buffer": "^6.0.3",
"class-variance-authority": "^0.7.0",
"clsx": "^2.1.1",
+1 -1
View File
@@ -18,7 +18,7 @@
"@radix-ui/react-scroll-area": "^1.2.0",
"@radix-ui/react-select": "^2.1.2",
"@radix-ui/react-slot": "^1.1.0",
"ai": "4.1.42",
"ai": "^5.0.52",
"buffer": "^6.0.3",
"class-variance-authority": "^0.7.0",
"clsx": "^2.1.1",
@@ -7,12 +7,31 @@ All pre-fetch hooks use this instead of duplicating urllib boilerplate.
from __future__ import annotations
import json
import os
import urllib.request
SEARCH_URL = "https://api.mem0.ai/v3/memories/search/"
SEARCH_TIMEOUT = 5
def should_rerank() -> bool:
"""Whether auto-injection searches should request Platform reranking.
The REST search endpoint does not rerank when ``rerank`` is omitted, so
auto-injected context is ordered by raw vector similarity and the single
most relevant memory can fall outside the injected top_k window. We default
reranking ON for the hook-driven injection path (the extra ~150-200ms is
well within the hook's curl budget) and let users opt out via MEM0_RERANK.
MEM0_RERANK is read case-insensitively; ``0``, ``false``, ``no``, and
``off`` disable reranking. Anything else (including unset) enables it.
"""
raw = os.environ.get("MEM0_RERANK")
if raw is None:
return True
return raw.strip().lower() not in ("0", "false", "no", "off", "")
def _do_search(api_key: str, payload: dict) -> list[dict]:
body = json.dumps(payload).encode()
req = urllib.request.Request(
@@ -23,7 +23,7 @@ sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from _formatting import TYPE_ICONS, format_age
from _identity import resolve_api_key, resolve_user_id
from _project import resolve_project_id
from _search import search_memories
from _search import search_memories, should_rerank
FILE_READ_GATE_MIN_BYTES = 1500
MAX_RESULTS = 5
@@ -93,6 +93,7 @@ def search_file_context(
api_key, user_id, project_id, query,
top_k=MAX_RESULTS, threshold=0.3,
global_search=global_search,
rerank=should_rerank(),
)
results = results[:MAX_RESULTS]
@@ -77,15 +77,16 @@ RESULTS=$(PYTHONPATH="$SCRIPT_DIR" MEM0_SEARCH_QUERY="$ERROR_QUERY" MEM0_SEARCH_
python3 -c "
import os, sys
sys.path.insert(0, os.environ.get('PYTHONPATH', '.'))
from _search import search_memories, format_results_for_context
from _search import search_memories, format_results_for_context, should_rerank
api_key = os.environ.get('MEM0_API_KEY', '')
user_id = os.environ.get('MEM0_SEARCH_USER', 'default')
project_id = os.environ.get('MEM0_PROJECT_ID', 'unknown')
query = os.environ.get('MEM0_SEARCH_QUERY', '')
rerank = should_rerank()
r1 = search_memories(api_key, user_id, project_id, query, metadata_type='anti_pattern', top_k=3)
r2 = search_memories(api_key, user_id, project_id, query, metadata_type='bug_fix', top_k=3)
r1 = search_memories(api_key, user_id, project_id, query, metadata_type='anti_pattern', top_k=3, rerank=rerank)
r2 = search_memories(api_key, user_id, project_id, query, metadata_type='bug_fix', top_k=3, rerank=rerank)
seen = set()
combined = []
@@ -120,14 +120,15 @@ if [ -n "$HAS_RESUME" ]; then
RESUME_RESULTS=$(PYTHONPATH="$SCRIPT_DIR" MEM0_SEARCH_USER="$USER_ID" python3 -c "
import os, sys
sys.path.insert(0, os.environ.get('PYTHONPATH', '.'))
from _search import search_memories, format_results_for_context
from _search import search_memories, format_results_for_context, should_rerank
api_key = os.environ.get('MEM0_API_KEY', '')
user_id = os.environ.get('MEM0_SEARCH_USER', 'default')
project_id = os.environ.get('MEM0_PROJECT_ID', 'unknown')
rerank = should_rerank()
state = search_memories(api_key, user_id, project_id, 'session state current task', metadata_type='session_state', top_k=3)
decisions = search_memories(api_key, user_id, project_id, 'recent decisions and learnings', metadata_type='decision', top_k=3)
state = search_memories(api_key, user_id, project_id, 'session state current task', metadata_type='session_state', top_k=3, rerank=rerank)
decisions = search_memories(api_key, user_id, project_id, 'recent decisions and learnings', metadata_type='decision', top_k=3, rerank=rerank)
all_r = state + decisions
seen = set()
@@ -101,6 +101,67 @@ def test_search_memories_no_api_key_returns_empty():
assert results == []
def test_search_memories_omits_rerank_by_default():
"""Regression for #5684: rerank must not be sent unless requested."""
from _search import search_memories
captured_body = {}
def mock_urlopen(req, timeout=None):
captured_body.update(json.loads(req.data.decode()))
resp = MagicMock()
resp.read.return_value = json.dumps({"results": []}).encode()
resp.__enter__ = lambda s: s
resp.__exit__ = MagicMock(return_value=False)
return resp
with patch("urllib.request.urlopen", side_effect=mock_urlopen):
search_memories("key", "user", "proj", "query")
assert "rerank" not in captured_body
def test_search_memories_forwards_rerank_true():
"""Regression for #5684: rerank=True must reach the request body so the
REST endpoint actually reranks (it does not rerank when omitted)."""
from _search import search_memories
captured_body = {}
def mock_urlopen(req, timeout=None):
captured_body.update(json.loads(req.data.decode()))
resp = MagicMock()
resp.read.return_value = json.dumps({"results": []}).encode()
resp.__enter__ = lambda s: s
resp.__exit__ = MagicMock(return_value=False)
return resp
with patch("urllib.request.urlopen", side_effect=mock_urlopen):
search_memories("key", "user", "proj", "query", rerank=True)
assert captured_body.get("rerank") is True
def test_should_rerank_defaults_true(monkeypatch):
"""Regression for #5684: auto-injection reranks by default."""
from _search import should_rerank
monkeypatch.delenv("MEM0_RERANK", raising=False)
assert should_rerank() is True
def test_should_rerank_opt_out_values(monkeypatch):
from _search import should_rerank
for falsey in ("0", "false", "False", "NO", "off", ""):
monkeypatch.setenv("MEM0_RERANK", falsey)
assert should_rerank() is False, falsey
for truthy in ("1", "true", "yes", "on"):
monkeypatch.setenv("MEM0_RERANK", truthy)
assert should_rerank() is True, truthy
def test_format_results_for_context():
from _search import format_results_for_context
+1
View File
@@ -64,6 +64,7 @@
},
"pnpm": {
"overrides": {
"form-data@<4.0.6": ">=4.0.6",
"protobufjs@<7.5.5": "^7.5.5",
"vite": "^8.0.5",
"langsmith@<0.6.0": "^0.6.0",
+6 -5
View File
@@ -5,6 +5,7 @@ settings:
excludeLinksFromLockfile: false
overrides:
form-data@<4.0.6: '>=4.0.6'
protobufjs@<7.5.5: ^7.5.5
vite: ^8.0.5
langsmith@<0.6.0: ^0.6.0
@@ -1109,8 +1110,8 @@ packages:
form-data-encoder@1.7.2:
resolution: {integrity: sha512-qfqtYan3rxrnCk1VYaA4H+Ms9xdpPqvLZa6xmMgFvhO32x7/3J/ExcTd6qpxM0vH2GdMI+poehyBZvqfMTto8A==}
form-data@4.0.5:
resolution: {integrity: sha512-8RipRLol37bNs2bhoV67fiTEvdTrbMUYcFTiy3+wuuOnUog2QBHCZWXDRijWQfAkhBj2Uf5UnVaiWwA5vdd82w==}
form-data@4.0.6:
resolution: {integrity: sha512-vKatAh4SlVfgbv+YtmhiRjhEMJsYpsG1Y2rMQtR+SVSbytsSD1YGzDIcrAJmdFec88u/+VoGmxnl+80gL1tRCQ==}
engines: {node: '>= 6'}
formdata-node@4.4.1:
@@ -2728,7 +2729,7 @@ snapshots:
'@types/node-fetch@2.6.13':
dependencies:
'@types/node': 22.19.20
form-data: 4.0.5
form-data: 4.0.6
'@types/node@18.19.130':
dependencies:
@@ -2870,7 +2871,7 @@ snapshots:
axios@1.17.0:
dependencies:
follow-redirects: 1.16.0
form-data: 4.0.5
form-data: 4.0.6
https-proxy-agent: 5.0.1
proxy-from-env: 2.1.0
transitivePeerDependencies:
@@ -3131,7 +3132,7 @@ snapshots:
form-data-encoder@1.7.2: {}
form-data@4.0.5:
form-data@4.0.6:
dependencies:
asynckit: 0.4.0
combined-stream: 1.0.8
@@ -13,6 +13,7 @@ onlyBuiltDependencies:
- protobufjs
overrides:
"form-data@<4.0.6": ">=4.0.6"
"protobufjs@<7.5.5": "^7.5.5"
"vite": "^8.0.5"
"langsmith@<0.6.0": "^0.6.0"
+3 -1
View File
@@ -71,8 +71,10 @@
},
"pnpm": {
"overrides": {
"form-data@<4.0.6": ">=4.0.6",
"uuid@<11.1.1": ">=11.1.1",
"esbuild": ">=0.28.1"
"esbuild": ">=0.28.1",
"undici@>=8.0.0 <8.5.0": ">=8.5.0"
}
}
}
+11 -9
View File
@@ -5,8 +5,10 @@ settings:
excludeLinksFromLockfile: false
overrides:
form-data@<4.0.6: '>=4.0.6'
uuid@<11.1.1: '>=11.1.1'
esbuild: '>=0.28.1'
undici@>=8.0.0 <8.5.0: '>=8.5.0'
importers:
@@ -1368,8 +1370,8 @@ packages:
form-data-encoder@1.7.2:
resolution: {integrity: sha512-qfqtYan3rxrnCk1VYaA4H+Ms9xdpPqvLZa6xmMgFvhO32x7/3J/ExcTd6qpxM0vH2GdMI+poehyBZvqfMTto8A==}
form-data@4.0.5:
resolution: {integrity: sha512-8RipRLol37bNs2bhoV67fiTEvdTrbMUYcFTiy3+wuuOnUog2QBHCZWXDRijWQfAkhBj2Uf5UnVaiWwA5vdd82w==}
form-data@4.0.6:
resolution: {integrity: sha512-vKatAh4SlVfgbv+YtmhiRjhEMJsYpsG1Y2rMQtR+SVSbytsSD1YGzDIcrAJmdFec88u/+VoGmxnl+80gL1tRCQ==}
engines: {node: '>= 6'}
formdata-node@4.4.1:
@@ -2353,8 +2355,8 @@ packages:
resolution: {integrity: sha512-4yqz8a3n5HmGTlsbADNtr/dJlhkh/55Rq798G6ibiULcXbDtaLpTl1pvdqcbFfeoj3iSi52lePFM7h9H21cw/A==}
engines: {node: '>=18.17'}
undici@8.3.0:
resolution: {integrity: sha512-TkUDgb6tl7KOGZ+7e8E3d2FYgUQgF6z5YypqjWmixVQSQERFcVrVg0ySADm2LVLRh5ljAaHTCR5Fmz3Q34rB7Q==}
undici@8.5.0:
resolution: {integrity: sha512-xamtWoB1EshgjpmlXd7GGm2VfdDtw1+rD8uhry8pSNW3If6S8E0m2T2+orSKeZXEn/aPJMviCpDBA65WJt8zhg==}
engines: {node: '>=22.19.0'}
util-deprecate@1.0.2:
@@ -2937,7 +2939,7 @@ snapshots:
minimatch: 10.2.5
proper-lockfile: 4.1.2
typebox: 1.1.38
undici: 8.3.0
undici: 8.5.0
yaml: 2.9.0
optionalDependencies:
'@mariozechner/clipboard': 0.3.9
@@ -3474,7 +3476,7 @@ snapshots:
'@types/node-fetch@2.6.13':
dependencies:
'@types/node': 25.9.2
form-data: 4.0.5
form-data: 4.0.6
'@types/node@18.19.130':
dependencies:
@@ -3598,7 +3600,7 @@ snapshots:
axios@1.17.0:
dependencies:
follow-redirects: 1.16.0
form-data: 4.0.5
form-data: 4.0.6
https-proxy-agent: 5.0.1
proxy-from-env: 2.1.0
transitivePeerDependencies:
@@ -3891,7 +3893,7 @@ snapshots:
form-data-encoder@1.7.2: {}
form-data@4.0.5:
form-data@4.0.6:
dependencies:
asynckit: 0.4.0
combined-stream: 1.0.8
@@ -4894,7 +4896,7 @@ snapshots:
undici@6.26.0: {}
undici@8.3.0: {}
undici@8.5.0: {}
util-deprecate@1.0.2: {}
@@ -2,5 +2,7 @@ packages:
- '.'
overrides:
"form-data@<4.0.6": ">=4.0.6"
"uuid@<11.1.1": ">=11.1.1"
"esbuild": ">=0.28.1"
"undici@>=8.0.0 <8.5.0": ">=8.5.0"
+1
View File
@@ -80,6 +80,7 @@
],
"overrides": {
"glob@>=10.2.0 <10.5.0": "^10.5.0",
"js-yaml@<=4.1.1": ">=4.2.0",
"minimatch@<3.1.3": "^3.1.3",
"minimatch@>=5.0.0 <5.1.8": "^5.1.8",
"minimatch@>=9.0.0 <9.0.7": "^9.0.7",
+9 -23
View File
@@ -6,6 +6,7 @@ settings:
overrides:
glob@>=10.2.0 <10.5.0: ^10.5.0
js-yaml@<=4.1.1: '>=4.2.0'
minimatch@<3.1.3: ^3.1.3
minimatch@>=5.0.0 <5.1.8: ^5.1.8
minimatch@>=9.0.0 <9.0.7: ^9.0.7
@@ -793,8 +794,8 @@ packages:
arg@4.1.3:
resolution: {integrity: sha512-58S9QDqG0Xx27YwPSt9fJxivjYl432YCwfDMfZ+71RAqUrZef7LrKQZ3LHLOwCS4FLNBplP533Zx895SeOCHvA==}
argparse@1.0.10:
resolution: {integrity: sha512-o5Roy6tNG4SL/FOkCAN6RzjiakZS25RLYFrcMttJqbdd8BWrnA+fGz57iN5Pb06pvBGvl5gQ0B48dJlslXvoTg==}
argparse@2.0.1:
resolution: {integrity: sha512-8+9WqebbFzpX9OR+Wa6O29asIogeRMzcGtAINdpMHHyAg10f05aSFVBbcEqGf/PXw1EjAZ+q2/bEBg3DvurK3Q==}
babel-jest@29.7.0:
resolution: {integrity: sha512-BrvGY3xZSwEcCzKvKsCi2GgHqDqsYkOP4/by5xCgIwGXQxIEh+8ew3gmrE1y7XRR6LHZIj6yLYnUi/mm2KXKBg==}
@@ -1025,11 +1026,6 @@ packages:
resolution: {integrity: sha512-UpzcLCXolUWcNu5HtVMHYdXJjArjsF9C0aNnquZYY4uW/Vu0miy5YoWvbV345HauVvcAUnpRuhMMcqTcGOY2+w==}
engines: {node: '>=8'}
esprima@4.0.1:
resolution: {integrity: sha512-eGuFFw7Upda+g4p+QHvnW0RyTX/SVeJBDM/gCtMARO0cLuT2HcEKnTPvhjV6aGeqrCB/sbNop0Kszm0jsaWU4A==}
engines: {node: '>=4'}
hasBin: true
eventsource-parser@3.1.0:
resolution: {integrity: sha512-kJezFj9YFAMLeORyi7aCLxLbD5/qWMQnoMVlVPyHIll7lgRJCc3JVln9Vgl9nwQi0YkMnhdGTMNn7CkRRAptMg==}
engines: {node: '>=18.0.0'}
@@ -1351,8 +1347,8 @@ packages:
js-tokens@4.0.0:
resolution: {integrity: sha512-RdJUflcE3cUzKiMqQgsCu06FPu9UdIJO0beYbPhHN4k6apgJtifcoCtT9bcxOpYBtpD2kCM6Sbzg4CausW/PKQ==}
js-yaml@3.14.2:
resolution: {integrity: sha512-PMSmkqxr106Xa156c2M265Z+FTrPl+oxd/rgOQy2tijQeK5TxQ43psO1ZCwhVOSdnn+RzkzlRz/eY4BgJBYVpg==}
js-yaml@4.2.0:
resolution: {integrity: sha512-ePWsvanv0DWuDRsW8dnt+R4jQ31SCRCQ7hhNcPXZPsoBZiemuZNYGf7adZdqX2D86j6rvKp3RpCxVTSb8WQlOw==}
hasBin: true
jsesc@3.1.0:
@@ -1654,9 +1650,6 @@ packages:
resolution: {integrity: sha512-i5uvt8C3ikiWeNZSVZNWcfZPItFQOsYTUAOkcUPGd8DqDy1uOUikjt5dG+uRlwyvR108Fb9DOd4GvXfT0N2/uQ==}
engines: {node: '>= 12'}
sprintf-js@1.0.3:
resolution: {integrity: sha512-D9cPgkvLlV3t3IzL0D0YLvGA9Ahk4PcvVwUbN0dSGr1aP0Nrt4AEnTUbuGvquEC0mA64Gqt1fzirlRs5ibXx8g==}
stack-utils@2.0.6:
resolution: {integrity: sha512-XlkWvfIm6RmsWtNJx+uqtKLS8eqFbxUg0ZzLXqY0caEy9l7hruX8IpiDnjsLavoBgqCCR71TqWO8MaXYheJ3RQ==}
engines: {node: '>=10'}
@@ -2226,7 +2219,7 @@ snapshots:
camelcase: 5.3.1
find-up: 4.1.0
get-package-type: 0.1.0
js-yaml: 3.14.2
js-yaml: 4.2.0
resolve-from: 5.0.0
'@istanbuljs/schema@0.1.6': {}
@@ -2605,9 +2598,7 @@ snapshots:
arg@4.1.3: {}
argparse@1.0.10:
dependencies:
sprintf-js: 1.0.3
argparse@2.0.1: {}
babel-jest@29.7.0(@babel/core@7.29.7):
dependencies:
@@ -2857,8 +2848,6 @@ snapshots:
escape-string-regexp@2.0.0: {}
esprima@4.0.1: {}
eventsource-parser@3.1.0: {}
execa@5.1.1:
@@ -3355,10 +3344,9 @@ snapshots:
js-tokens@4.0.0: {}
js-yaml@3.14.2:
js-yaml@4.2.0:
dependencies:
argparse: 1.0.10
esprima: 4.0.1
argparse: 2.0.1
jsesc@3.1.0: {}
@@ -3628,8 +3616,6 @@ snapshots:
source-map@0.7.6: {}
sprintf-js@1.0.3: {}
stack-utils@2.0.6:
dependencies:
escape-string-regexp: 2.0.0
@@ -7,6 +7,7 @@ onlyBuiltDependencies:
overrides:
"glob@>=10.2.0 <10.5.0": "^10.5.0"
"js-yaml@<=4.1.1": ">=4.2.0"
"minimatch@<3.1.3": "^3.1.3"
"minimatch@>=5.0.0 <5.1.8": "^5.1.8"
"minimatch@>=9.0.0 <9.0.7": "^9.0.7"
+1
View File
@@ -139,6 +139,7 @@
"better-sqlite3"
],
"overrides": {
"form-data@<4.0.6": ">=4.0.6",
"picomatch@<2.3.2": "^2.3.2",
"picomatch@>=4.0.0 <4.0.4": "^4.0.4",
"jws@4.0.0": "4.0.1",
+6 -5
View File
@@ -5,6 +5,7 @@ settings:
excludeLinksFromLockfile: false
overrides:
form-data@<4.0.6: '>=4.0.6'
picomatch@<2.3.2: ^2.3.2
picomatch@>=4.0.0 <4.0.4: ^4.0.4
jws@4.0.0: 4.0.1
@@ -1554,8 +1555,8 @@ packages:
form-data-encoder@1.7.2:
resolution: {integrity: sha512-qfqtYan3rxrnCk1VYaA4H+Ms9xdpPqvLZa6xmMgFvhO32x7/3J/ExcTd6qpxM0vH2GdMI+poehyBZvqfMTto8A==}
form-data@4.0.5:
resolution: {integrity: sha512-8RipRLol37bNs2bhoV67fiTEvdTrbMUYcFTiy3+wuuOnUog2QBHCZWXDRijWQfAkhBj2Uf5UnVaiWwA5vdd82w==}
form-data@4.0.6:
resolution: {integrity: sha512-vKatAh4SlVfgbv+YtmhiRjhEMJsYpsG1Y2rMQtR+SVSbytsSD1YGzDIcrAJmdFec88u/+VoGmxnl+80gL1tRCQ==}
engines: {node: '>= 6'}
formdata-node@4.4.1:
@@ -3972,7 +3973,7 @@ snapshots:
'@types/node-fetch@2.6.13':
dependencies:
'@types/node': 22.19.21
form-data: 4.0.5
form-data: 4.0.6
'@types/node@18.19.130':
dependencies:
@@ -4078,7 +4079,7 @@ snapshots:
axios@1.17.0:
dependencies:
follow-redirects: 1.16.0
form-data: 4.0.5
form-data: 4.0.6
https-proxy-agent: 5.0.1
proxy-from-env: 2.1.0
transitivePeerDependencies:
@@ -4561,7 +4562,7 @@ snapshots:
form-data-encoder@1.7.2: {}
form-data@4.0.5:
form-data@4.0.6:
dependencies:
asynckit: 0.4.0
combined-stream: 1.0.8
+1
View File
@@ -6,6 +6,7 @@ onlyBuiltDependencies:
- better-sqlite3
overrides:
"form-data@<4.0.6": ">=4.0.6"
"picomatch@<2.3.2": "^2.3.2"
"picomatch@>=4.0.0 <4.0.4": "^4.0.4"
"jws@4.0.0": "4.0.1"
+5
View File
@@ -253,6 +253,11 @@ export default class MemoryClient {
messages: Array<Message>,
options: AddMemoryOptions & Record<string, any> = {},
): Promise<Array<Memory>> {
// Tightly scoped validation guard to resolve #5465
if (!messages || (Array.isArray(messages) && messages.length === 0)) {
throw new Error("Cannot process an empty messages payload.");
}
if (this.telemetryId === "") await this.ping();
const payload = this._preparePayload(messages, options);
+3 -1
View File
@@ -169,7 +169,9 @@ export interface PaginatedMemories {
export interface ProjectResponse {
customInstructions?: string;
customCategories?: string[];
// The API returns category objects (`[{ "<name>": "<description>" }]`),
// not bare strings (see issue #5738).
customCategories?: custom_categories[];
[key: string]: any;
}
@@ -59,16 +59,15 @@ describe("MemoryClient - add()", () => {
expect(getFetchBody(call!).user_id).toBe("user_1");
});
test("sends empty messages array without crashing", async () => {
const extra = new Map<string, { status: number; body: unknown }>();
extra.set("/v3/memories/add/", { status: 200, body: [] });
const mock = setupMockFetch(extra);
test("throws an error when given an empty messages array", async () => {
setupMockFetch();
const client = new MemoryClient({ apiKey: TEST_API_KEY });
await client.add([], { userId: "u1" });
const call = findFetchCall(mock, "/v3/memories/add/", "POST");
expect(getFetchBody(call!).messages).toEqual([]);
//Asserts that the validation guard catches the empty input early
await expect(client.add([], { userId: "u1" })).rejects.toThrow(
"Cannot process an empty messages payload.",
);
});
});
+148
View File
@@ -0,0 +1,148 @@
import { camelToSnakeKeys, snakeToCamelKeys } from "../utils";
describe("camelToSnakeKeys / snakeToCamelKeys", () => {
it("converts SDK-defined keys between camelCase and snake_case", () => {
expect(camelToSnakeKeys({ userId: "u1", agentId: "a1" })).toEqual({
user_id: "u1",
agent_id: "a1",
});
expect(snakeToCamelKeys({ user_id: "u1", agent_id: "a1" })).toEqual({
userId: "u1",
agentId: "a1",
});
});
it("preserves logical operator keys (OR/AND/NOT)", () => {
expect(camelToSnakeKeys({ OR: [{ userId: "u1" }] })).toEqual({
OR: [{ user_id: "u1" }],
});
});
describe("user-controlled metadata blob (issue #5055)", () => {
it("does not camelize snake_case keys inside metadata on read", () => {
const apiResponse = {
id: "mem-1",
user_id: "u1",
metadata: { message_id: "x", some_custom_key: "y" },
};
expect(snakeToCamelKeys(apiResponse)).toEqual({
id: "mem-1",
userId: "u1",
metadata: { message_id: "x", some_custom_key: "y" },
});
});
it("does not snake_case camelCase keys inside metadata on write", () => {
const payload = {
userId: "u1",
metadata: { messageId: "x", someCustomKey: "y" },
};
expect(camelToSnakeKeys(payload)).toEqual({
user_id: "u1",
metadata: { messageId: "x", someCustomKey: "y" },
});
});
it("round-trips arbitrary metadata keys losslessly", () => {
const metadata = {
message_id: "abc",
camelKey: 1,
nested: { deep_snake: true, deepCamel: false },
arr: [{ inner_key: 1 }],
};
const roundTripped = snakeToCamelKeys(
camelToSnakeKeys({ userId: "u1", metadata }),
);
expect(roundTripped.metadata).toEqual(metadata);
});
it("preserves metadata nested inside an array of results", () => {
const apiResponse = {
results: [
{ id: "1", metadata: { message_id: "x" } },
{ id: "2", metadata: { another_key: "z" } },
],
};
expect(snakeToCamelKeys(apiResponse)).toEqual({
results: [
{ id: "1", metadata: { message_id: "x" } },
{ id: "2", metadata: { another_key: "z" } },
],
});
});
});
describe("user-controlled structuredDataSchema blob (issue #5055)", () => {
it("converts the outer key but leaves user field names on write", () => {
expect(
camelToSnakeKeys({
structuredDataSchema: { firstName: "string", lastName: "string" },
}),
).toEqual({
// outer SDK key is snake_cased, user-defined field names are not
structured_data_schema: { firstName: "string", lastName: "string" },
});
});
it("converts the outer key but leaves user field names on read", () => {
expect(
snakeToCamelKeys({
structured_data_schema: { first_name: "string", last_name: "string" },
}),
).toEqual({
structuredDataSchema: { first_name: "string", last_name: "string" },
});
});
});
describe("user-controlled customCategories names (issue #5738)", () => {
it("converts the outer key but leaves multi-word category names on write", () => {
expect(
camelToSnakeKeys({
customCategories: [
{ work_life_balance: "desc" },
{ AIResearch: "desc" },
],
}),
).toEqual({
// outer SDK key is snake_cased, user-defined category names are not
custom_categories: [
{ work_life_balance: "desc" },
{ AIResearch: "desc" },
],
});
});
it("converts the outer key but leaves category names verbatim on read", () => {
expect(
snakeToCamelKeys({
custom_categories: [
{ work_life_balance: "desc" },
{ AIResearch: "desc" },
],
}),
).toEqual({
customCategories: [
{ work_life_balance: "desc" },
{ AIResearch: "desc" },
],
});
});
it("round-trips category names losslessly (write then read)", () => {
const customCategories = [
{ work_life_balance: "balance between work and life" },
{ AIResearch: "artificial intelligence research" },
];
const roundTripped = snakeToCamelKeys(
camelToSnakeKeys({ customCategories }),
);
expect(roundTripped.customCategories).toEqual(customCategories);
});
});
});
+26 -2
View File
@@ -14,9 +14,32 @@ function snakeToCamel(str: string): string {
return str.replace(/_([a-z])/g, (_, letter) => letter.toUpperCase());
}
/**
* Keys whose values are user-controlled, opaque blobs. Their nested keys must
* be passed through verbatim — converting them would silently rewrite the
* user's own keys and break round-trips (see issue #5055).
*
* The check runs against the source key, so a multi-word key must be listed in
* both casings to be covered in both directions: the camelCase form for the
* outbound `camelToSnakeKeys` path and the snake_case form for the inbound
* `snakeToCamelKeys` path. `metadata` is spelled identically in both, so one
* entry suffices; `structuredDataSchema` needs both.
*/
const OPAQUE_VALUE_KEYS = new Set([
"metadata",
"structuredDataSchema",
"structured_data_schema",
// Custom-category names are user-controlled keys (`[{ "<name>": "<desc>" }]`).
// Listed in both casings so they round-trip verbatim in both directions
// (see issue #5738; same class as `metadata`/`structuredDataSchema`).
"customCategories",
"custom_categories",
]);
/**
* Recursively converts all keys of an object from camelCase to snake_case.
* Used for converting user-facing camelCase params to API snake_case payloads.
* Values under {@link OPAQUE_VALUE_KEYS} (e.g. `metadata`) are left untouched.
*/
export function camelToSnakeKeys(obj: any): any {
if (obj === null || obj === undefined || typeof obj !== "object") return obj;
@@ -26,7 +49,7 @@ export function camelToSnakeKeys(obj: any): any {
return Object.fromEntries(
Object.entries(obj).map(([key, value]) => [
camelToSnake(key),
camelToSnakeKeys(value),
OPAQUE_VALUE_KEYS.has(key) ? value : camelToSnakeKeys(value),
]),
);
}
@@ -34,6 +57,7 @@ export function camelToSnakeKeys(obj: any): any {
/**
* Recursively converts all keys of an object from snake_case to camelCase.
* Used for converting API snake_case responses to user-facing camelCase.
* Values under {@link OPAQUE_VALUE_KEYS} (e.g. `metadata`) are left untouched.
*/
export function snakeToCamelKeys(obj: any): any {
if (obj === null || obj === undefined || typeof obj !== "object") return obj;
@@ -43,7 +67,7 @@ export function snakeToCamelKeys(obj: any): any {
return Object.fromEntries(
Object.entries(obj).map(([key, value]) => [
snakeToCamel(key),
snakeToCamelKeys(value),
OPAQUE_VALUE_KEYS.has(key) ? value : snakeToCamelKeys(value),
]),
);
}
+7 -1
View File
@@ -14,7 +14,13 @@ export class AnthropicLLM implements LLM {
if (!apiKey) {
throw new Error("Anthropic API key is required");
}
this.client = new Anthropic({ apiKey });
// Forward baseURL to the client when set so proxy/gateway users are
// honored (parity with the OpenAI provider and the Python fix in #5626).
const clientArgs: { apiKey: string; baseURL?: string } = { apiKey };
if (config.baseURL) {
clientArgs.baseURL = config.baseURL;
}
this.client = new Anthropic(clientArgs);
this.model = config.model || "claude-sonnet-4-6";
// Defaults mirror the Python provider's AnthropicConfig
// (max_tokens=2000, temperature=0.1, top_p omitted).
+38 -3
View File
@@ -619,6 +619,25 @@ export class Memory {
"messages is required and cannot be undefined or null. Provide a string or array of messages.",
);
}
if (Array.isArray(messages)) {
if (messages.length === 0) {
throw new Error(
"messages array cannot be empty. Provide at least one message with non-empty content.",
);
}
const allBlank = messages.every(
(m) => typeof m.content === "string" && m.content.trim() === "",
);
if (allBlank) {
throw new Error(
"messages array cannot contain only blank content. Provide at least one message with non-empty content.",
);
}
} else if (messages.trim() === "") {
throw new Error(
"messages string cannot be empty. Provide non-empty content.",
);
}
const temporalUsageNotice = detectTemporalUsageFromMetadata(
config?.metadata,
@@ -698,7 +717,7 @@ export class Memory {
if (!infer) {
const returnedMemories: MemoryItem[] = [];
for (const message of messages) {
if (message.content === "system") {
if (message.role === "system") {
continue;
}
const memoryId = await this.createMemory(
@@ -731,7 +750,13 @@ export class Memory {
// getLastMessages not supported — proceed without context
}
}
const parsedMessages = messages.map((m) => m.content).join("\n");
// Preserve role on the messages being extracted so the prompt's role-aware
// logic and the required `attributed_to` output have the speaker to work
// with. Matches the Python oss `parse_messages` helper (`role: content`);
// without this, assistant statements get attributed to the user.
const parsedMessages = messages
.map((m) => `${m.role}: ${m.content}`)
.join("\n");
// Phase 1: Existing memory retrieval
const queryEmbedding = await this.embedder.embed(parsedMessages);
@@ -1164,7 +1189,13 @@ export class Memory {
}
}
const result = { ...memoryItem, ...filters };
const result = {
...memoryItem,
...filters,
...(memory.payload.attributedTo && {
attributedTo: memory.payload.attributedTo,
}),
};
await this._displayFirstRunNotice("get");
return result;
}
@@ -1428,6 +1459,7 @@ export class Memory {
...(payload.user_id && { user_id: payload.user_id }),
...(payload.agent_id && { agent_id: payload.agent_id }),
...(payload.run_id && { run_id: payload.run_id }),
...(payload.attributedTo && { attributedTo: payload.attributedTo }),
...(scored.scoreDetails && { score_details: scored.scoreDetails }),
};
});
@@ -1662,6 +1694,9 @@ export class Memory {
...(mem.payload.user_id && { user_id: mem.payload.user_id }),
...(mem.payload.agent_id && { agent_id: mem.payload.agent_id }),
...(mem.payload.run_id && { run_id: mem.payload.run_id }),
...(mem.payload.attributedTo && {
attributedTo: mem.payload.attributedTo,
}),
}));
const result = { results };
+1
View File
@@ -82,6 +82,7 @@ export interface MemoryItem {
updatedAt?: string;
score?: number;
metadata?: Record<string, any>;
attributedTo?: string;
}
export interface SearchFilters {
+5 -3
View File
@@ -31,9 +31,11 @@ const parse_vision_messages = async (messages: Message[]) => {
typeof message.content === "object" &&
message.content.type === "image_url"
) {
const description = await get_image_description(
message.content.image_url.url,
);
const imageUrl = message.content.image_url?.url;
if (!imageUrl) {
throw new Error("image_url content part is missing image_url.url");
}
const description = await get_image_description(imageUrl);
new_message.content =
typeof description === "string"
? description
+33 -4
View File
@@ -4,17 +4,46 @@
*/
const mockCreate = jest.fn();
const mockConstructor = jest.fn();
jest.mock("@anthropic-ai/sdk", () => {
return jest.fn().mockImplementation(() => ({
messages: { create: mockCreate },
}));
return jest.fn().mockImplementation((args) => {
mockConstructor(args);
return { messages: { create: mockCreate } };
});
});
import { AnthropicLLM } from "../src/llms/anthropic";
describe("AnthropicLLM (unit)", () => {
beforeEach(() => mockCreate.mockClear());
beforeEach(() => {
mockCreate.mockClear();
mockConstructor.mockClear();
});
// Regression #5665: a configured baseURL must reach the Anthropic client so
// proxy/gateway users are not silently bypassed (TS parity with #5626).
it("forwards baseURL to the Anthropic client when set", () => {
new AnthropicLLM({
apiKey: "test-key",
baseURL: "https://proxy.example/v1",
});
expect(mockConstructor).toHaveBeenCalledTimes(1);
const ctorArgs = mockConstructor.mock.calls[0][0];
expect(ctorArgs.apiKey).toBe("test-key");
expect(ctorArgs.baseURL).toBe("https://proxy.example/v1");
});
// When no baseURL is configured the client must not receive a baseURL key
// (so the SDK default endpoint is used).
it("does NOT set baseURL when none is configured", () => {
new AnthropicLLM({ apiKey: "test-key" });
expect(mockConstructor).toHaveBeenCalledTimes(1);
const ctorArgs = mockConstructor.mock.calls[0][0];
expect(ctorArgs.baseURL).toBeUndefined();
});
it("returns text when no tools are provided and model returns a text block", async () => {
mockCreate.mockResolvedValueOnce({
+16
View File
@@ -126,6 +126,22 @@ describe("Memory - add()", () => {
expect(result.results.length).toBeGreaterThan(0);
});
test("preserves message roles in the extraction prompt (## New Messages)", async () => {
// The OpenAI LLM mock echoes the `## New Messages` section of the prompt back
// as the extracted text, so the stored memory reveals what the LLM received.
// Roles must survive into that section, otherwise the prompt's role-aware
// logic and required `attributed_to` output have no speaker to attribute to
// and assistant statements get stored as user facts.
const messages = [
{ role: "user", content: "I want to sleep earlier." },
{ role: "assistant", content: "Aim for 00:30 sleep / 08:30 wake." },
];
const result: SearchResult = await memory.add(messages, { userId });
const seen = result.results.map((r) => r.memory).join("\n");
expect(seen).toContain("user: I want to sleep earlier.");
expect(seen).toContain("assistant: Aim for 00:30 sleep / 08:30 wake.");
});
test("works with agentId instead of userId", async () => {
const result: SearchResult = await memory.add("test", {
agentId: "agent_1",
+39
View File
@@ -361,6 +361,45 @@ describe("Memory - search()", () => {
});
});
// ─── attributedTo (#5666) ────────────────────────────────
describe("Memory - attributedTo round-trip (#5666)", () => {
let memory: Memory;
const userId = `attributed_test_${Date.now()}`;
let id: string;
beforeAll(async () => {
memory = createMemory();
// The mocked LLM tags every extracted fact with attributed_to: "user".
const addResult: SearchResult = await memory.add("I love AI", { userId });
id = addResult.results[0].id;
});
afterAll(async () => {
await memory.reset();
});
test("get() surfaces attributedTo", async () => {
const item: MemoryItem | null = await memory.get(id);
expect(item!.attributedTo).toBe("user");
});
test("getAll() surfaces attributedTo", async () => {
const result: SearchResult = await memory.getAll({
filters: { user_id: userId },
});
expect(result.results[0].attributedTo).toBe("user");
});
test("search() surfaces attributedTo", async () => {
const result: SearchResult = await memory.search("AI", {
filters: { user_id: userId },
});
expect(result.results.length).toBeGreaterThan(0);
expect(result.results[0].attributedTo).toBe("user");
});
});
// ─── history() ───────────────────────────────────────────
describe("Memory - history()", () => {
@@ -84,6 +84,24 @@ describe("Memory Input Validation", () => {
memory.add(null, { userId: testUserId }),
).rejects.toThrow("messages is required");
});
it("should throw error when messages is an empty array", async () => {
await expect(memory.add([], { userId: testUserId })).rejects.toThrow(
"messages array cannot be empty",
);
});
it("should throw error when messages array contains only blank content", async () => {
await expect(
memory.add([{ role: "user", content: " " }], { userId: testUserId }),
).rejects.toThrow("messages array cannot contain only blank content");
});
it("should throw error when messages is an empty string", async () => {
await expect(memory.add(" ", { userId: testUserId })).rejects.toThrow(
"messages string cannot be empty",
);
});
});
describe("search() threshold validation", () => {
+4 -4
View File
@@ -151,10 +151,10 @@ class MemoryClient:
try:
params = self._prepare_params()
response = self.client.get("/v1/ping/", params=params)
data = response.json()
response.raise_for_status()
data = response.json()
if data.get("org_id") and data.get("project_id"):
self.org_id = data.get("org_id")
self.project_id = data.get("project_id")
@@ -1044,10 +1044,10 @@ class AsyncMemoryClient:
},
params=params,
)
data = response.json()
response.raise_for_status()
data = response.json()
if data.get("org_id") and data.get("project_id"):
self.org_id = data.get("org_id")
self.project_id = data.get("project_id")
+3 -4
View File
@@ -2,9 +2,8 @@ import os
from abc import ABC
from typing import Dict, Optional, Union
import httpx
from mem0.configs.base import AzureConfig
from mem0.utils.http import build_http_client
class BaseEmbedderConfig(ABC):
@@ -81,7 +80,8 @@ class BaseEmbedderConfig(ABC):
self.embedding_dims = embedding_dims
# AzureOpenAI specific
self.http_client = httpx.Client(proxies=http_client_proxies) if http_client_proxies else None
self.http_client_proxies = http_client_proxies
self.http_client = build_http_client(http_client_proxies)
# Ollama specific
self.ollama_base_url = ollama_base_url
@@ -109,4 +109,3 @@ class BaseEmbedderConfig(ABC):
self.aws_secret_access_key = aws_secret_access_key
self.aws_session_token = aws_session_token
self.aws_region = aws_region or os.environ.get("AWS_REGION") or "us-west-2"
+3 -2
View File
@@ -1,7 +1,7 @@
from abc import ABC
from typing import Dict, Optional, Union
import httpx
from mem0.utils.http import build_http_client
class BaseLlmConfig(ABC):
@@ -74,4 +74,5 @@ class BaseLlmConfig(ABC):
self.vision_details = vision_details
self.reasoning_effort = reasoning_effort
self.is_reasoning_model = is_reasoning_model
self.http_client = httpx.Client(proxies=http_client_proxies) if http_client_proxies else None
self.http_client_proxies = http_client_proxies
self.http_client = build_http_client(http_client_proxies)
+6
View File
@@ -18,6 +18,12 @@ class OpenSearchConfig(BaseModel):
"RequestsHttpConnection", description="Connection class for OpenSearch"
)
pool_maxsize: int = Field(20, description="Maximum number of connections in the pool")
auto_refresh: bool = Field(
False,
description="Automatically refresh index after insert operations to make documents "
"immediately searchable. Disabled by default for OpenSearch Serverless compatibility. "
"OpenSearch automatically refreshes indices every ~1 second, so most users don't need this.",
)
@model_validator(mode="before")
@classmethod
+5
View File
@@ -70,6 +70,11 @@ class AWSBedrockEmbedding(EmbeddingBase):
else:
# Amazon and other providers
input_body["inputText"] = text
# Titan Text Embeddings V2 accepts an optional output dimension
# (256/512/1024). Only forward embedding_dims when the user set it,
# mirroring the OpenAI embedder's guarded `dimensions` pass-through.
if self.config.embedding_dims is not None and "v2" in self.config.model:
input_body["dimensions"] = self.config.embedding_dims
body = json.dumps(input_body)
+17
View File
@@ -37,3 +37,20 @@ class GoogleGenAIEmbedding(EmbeddingBase):
response = self.client.models.embed_content(model=self.config.model, contents=text, config=config)
return response.embeddings[0].values
def embed_batch(self, texts, memory_action="add"):
if not texts:
return []
config = types.EmbedContentConfig(output_dimensionality=self.config.embedding_dims)
MAX_BATCH = 100
all_embeddings = []
for i in range(0, len(texts), MAX_BATCH):
chunk = [t.replace("\n", " ") for t in texts[i : i + MAX_BATCH]]
response = self.client.models.embed_content(model=self.config.model, contents=chunk, config=config)
all_embeddings.extend(e.values for e in response.embeddings)
if len(all_embeddings) != len(texts):
raise ValueError(
f"Gemini embed_batch() returned {len(all_embeddings)} embeddings for {len(texts)} texts "
f"using model '{self.config.model}'"
)
return all_embeddings
+22
View File
@@ -42,3 +42,25 @@ class HuggingFaceEmbedding(EmbeddingBase):
).data[0].embedding
else:
return self.model.encode(text, convert_to_numpy=True).tolist()
def embed_batch(self, texts, memory_action="add"):
if not texts:
return []
if self.config.huggingface_base_url:
response = self.client.embeddings.create(input=texts, model=self.config.model, **self.config.model_kwargs)
sorted_data = sorted(response.data, key=lambda x: x.index)
embeddings = [item.embedding for item in sorted_data]
if len(embeddings) != len(texts):
raise ValueError(
f"HuggingFace embed_batch() returned {len(embeddings)} embeddings for {len(texts)} texts"
f" using model '{self.config.model}'"
)
return embeddings
else:
result = self.model.encode(texts, convert_to_numpy=True).tolist()
if len(result) != len(texts):
raise ValueError(
f"HuggingFace embed_batch() returned {len(result)} embeddings for {len(texts)} texts"
f" using model '{self.config.model}'"
)
return result
+14
View File
@@ -27,3 +27,17 @@ class LMStudioEmbedding(EmbeddingBase):
"""
text = text.replace("\n", " ")
return self.client.embeddings.create(input=[text], model=self.config.model).data[0].embedding
def embed_batch(self, texts, memory_action="add"):
if not texts:
return []
cleaned = [t.replace("\n", " ") for t in texts]
response = self.client.embeddings.create(input=cleaned, model=self.config.model)
sorted_data = sorted(response.data, key=lambda x: x.index)
embeddings = [item.embedding for item in sorted_data]
if len(embeddings) != len(texts):
raise ValueError(
f"LM Studio embed_batch() returned {len(embeddings)} embeddings for {len(texts)} texts"
f" using model '{self.config.model}'"
)
return embeddings
+13
View File
@@ -29,3 +29,16 @@ class TogetherEmbedding(EmbeddingBase):
"""
return self.client.embeddings.create(model=self.config.model, input=text).data[0].embedding
def embed_batch(self, texts, memory_action="add"):
if not texts:
return []
response = self.client.embeddings.create(model=self.config.model, input=texts)
sorted_data = sorted(response.data, key=lambda x: x.index)
embeddings = [item.embedding for item in sorted_data]
if len(embeddings) != len(texts):
raise ValueError(
f"Together embed_batch() returned {len(embeddings)} embeddings for {len(texts)} texts"
f" using model '{self.config.model}'"
)
return embeddings
+21
View File
@@ -62,3 +62,24 @@ class VertexAIEmbedding(EmbeddingBase):
embeddings = self.model.get_embeddings(texts=[text_input], output_dimensionality=self.config.embedding_dims)
return embeddings[0].values
def embed_batch(self, texts, memory_action="add"):
if not texts:
return []
embedding_type = "SEMANTIC_SIMILARITY"
if memory_action is not None:
if memory_action not in self.embedding_types:
raise ValueError(f"Invalid memory action: {memory_action}")
embedding_type = self.embedding_types[memory_action]
all_embeddings = []
for i in range(0, len(texts), 250):
chunk = texts[i : i + 250]
inputs = [TextEmbeddingInput(text=t, task_type=embedding_type) for t in chunk]
results = self.model.get_embeddings(texts=inputs, output_dimensionality=self.config.embedding_dims)
all_embeddings.extend(r.values for r in results)
if len(all_embeddings) != len(texts):
raise ValueError(
f"Vertex AI embed_batch() returned {len(all_embeddings)} embeddings for {len(texts)} texts"
f" using model '{self.config.model}'"
)
return all_embeddings
+6 -2
View File
@@ -29,7 +29,7 @@ class AnthropicLLM(LLMBase):
top_k=config.top_k,
enable_vision=config.enable_vision,
vision_details=config.vision_details,
http_client_proxies=config.http_client,
http_client_proxies=config.http_client_proxies,
)
super().__init__(config)
@@ -38,7 +38,11 @@ class AnthropicLLM(LLMBase):
self.config.model = "claude-sonnet-4-6"
api_key = self.config.api_key or os.getenv("ANTHROPIC_API_KEY")
self.client = anthropic.Anthropic(api_key=api_key)
base_url = self.config.anthropic_base_url or os.getenv("ANTHROPIC_BASE_URL")
client_kwargs = {"api_key": api_key}
if base_url:
client_kwargs["base_url"] = base_url
self.client = anthropic.Anthropic(**client_kwargs)
def _get_common_params(self, **kwargs) -> Dict:
"""Get common parameters, avoiding sending both temperature and top_p together.
+28 -7
View File
@@ -1,3 +1,4 @@
import copy
import json
import os
from typing import Dict, List, Optional, Union
@@ -32,7 +33,7 @@ class AzureOpenAILLM(LLMBase):
enable_vision=config.enable_vision,
vision_details=config.vision_details,
reasoning_effort=getattr(config, 'reasoning_effort', None),
http_client_proxies=config.http_client,
http_client_proxies=config.http_client_proxies,
is_reasoning_model=getattr(config, 'is_reasoning_model', None),
)
@@ -69,6 +70,26 @@ class AzureOpenAILLM(LLMBase):
default_headers=default_headers,
)
@staticmethod
def _rewrite_assistant_keyword(messages):
"""
Return a copy of ``messages`` with the word "assistant" replaced by "ai"
in the last message's textual content.
Azure's content management policy can flag the literal word "assistant",
which makes ``add`` fail (see issue #2636). The rewrite targets that
trigger without mutating the caller's messages and without assuming the
content is a string, so multimodal (list) content passes through untouched.
"""
if not messages:
return messages
messages = copy.deepcopy(messages)
last_content = messages[-1].get("content")
if isinstance(last_content, str):
messages[-1]["content"] = last_content.replace("assistant", "ai")
return messages
def _parse_response(self, response, tools):
"""
Process the response based on whether tools are used or not.
@@ -121,14 +142,14 @@ class AzureOpenAILLM(LLMBase):
str: The generated response.
"""
user_prompt = messages[-1]["content"]
user_prompt = user_prompt.replace("assistant", "ai")
messages[-1]["content"] = user_prompt
# Azure's "Indirect Attacks" content filter can flag the literal word
# "assistant" in the prompt, so it is rewritten to "ai" before the request.
# Work on a copy so the caller's messages are left untouched and string-only
# content is handled without breaking multimodal (list) content.
messages = self._rewrite_assistant_keyword(messages)
params = self._get_supported_params(messages=messages, **kwargs)
# Add model and messages
params.update({
"model": self.config.model,
+26 -5
View File
@@ -1,3 +1,4 @@
import copy
import json
import os
from typing import Dict, List, Optional
@@ -66,11 +67,11 @@ class AzureOpenAIStructuredLLM(LLMBase):
str: The generated response.
"""
user_prompt = messages[-1]["content"]
user_prompt = user_prompt.replace("assistant", "ai")
messages[-1]["content"] = user_prompt
# Azure's "Indirect Attacks" content filter can flag the literal word
# "assistant" in the prompt, so it is rewritten to "ai" before the request.
# Work on a copy so the caller's messages are left untouched and string-only
# content is handled without breaking multimodal (list) content.
messages = self._rewrite_assistant_keyword(messages)
is_reasoning = self._is_reasoning_model(self.config.model)
params = {
@@ -101,6 +102,26 @@ class AzureOpenAIStructuredLLM(LLMBase):
response = self.client.chat.completions.create(**params)
return self._parse_response(response, tools)
@staticmethod
def _rewrite_assistant_keyword(messages):
"""
Return a copy of ``messages`` with the word "assistant" replaced by "ai"
in the last message's textual content.
Azure's content management policy can flag the literal word "assistant",
which makes ``add`` fail (see issue #2636). The rewrite targets that
trigger without mutating the caller's messages and without assuming the
content is a string, so multimodal (list) content passes through untouched.
"""
if not messages:
return messages
messages = copy.deepcopy(messages)
last_content = messages[-1].get("content")
if isinstance(last_content, str):
messages[-1]["content"] = last_content.replace("assistant", "ai")
return messages
def _parse_response(self, response, tools):
"""
Process the response based on whether tools are used or not.
+1 -1
View File
@@ -28,7 +28,7 @@ class DeepSeekLLM(LLMBase):
top_k=config.top_k,
enable_vision=config.enable_vision,
vision_details=config.vision_details,
http_client_proxies=config.http_client,
http_client_proxies=config.http_client_proxies,
)
super().__init__(config)
+33 -2
View File
@@ -1,6 +1,7 @@
import json
import logging
import os
from typing import Dict, List, Optional
from typing import Dict, List, Optional, Union
try:
from groq import Groq
@@ -11,6 +12,8 @@ from mem0.configs.llms.base import BaseLlmConfig
from mem0.llms.base import LLMBase
from mem0.memory.utils import extract_json
logger = logging.getLogger(__name__)
class GroqLLM(LLMBase):
def __init__(self, config: Optional[BaseLlmConfig] = None):
@@ -22,6 +25,24 @@ class GroqLLM(LLMBase):
api_key = self.config.api_key or os.getenv("GROQ_API_KEY")
self.client = Groq(api_key=api_key)
@staticmethod
def _supports_json_mode(model: Optional[Union[str, Dict]]) -> bool:
"""
Groq's compound agentic systems (e.g. ``groq/compound``, ``groq/compound-mini``)
do not support the JSON ``response_format`` and return empty or non-JSON content
when it is requested. See https://console.groq.com/docs/structured-outputs.
Non-string models (the config allows a dict) are assumed to support JSON mode,
preserving prior behavior.
"""
if not isinstance(model, str):
return True
# Strip provider prefixes (e.g. "groq/compound-mini" -> "compound-mini"),
# mirroring the _is_reasoning_model heuristic in LLMBase, so the match
# targets the compound family rather than any name containing the substring.
base_model = model.lower().rsplit("/", 1)[-1]
return not base_model.startswith("compound")
def _parse_response(self, response, tools):
"""
Process the response based on whether tools are used or not.
@@ -79,7 +100,17 @@ class GroqLLM(LLMBase):
"top_p": self.config.top_p,
}
if response_format:
params["response_format"] = response_format
requests_json = isinstance(response_format, dict) and response_format.get("type") in (
"json_object",
"json_schema",
)
if requests_json and not self._supports_json_mode(self.config.model):
logger.debug(
f"Model '{self.config.model}' does not support JSON response_format; "
"sending the request without it."
)
else:
params["response_format"] = response_format
if tools:
params["tools"] = tools
params["tool_choice"] = tool_choice
+1 -1
View File
@@ -27,7 +27,7 @@ class LMStudioLLM(LLMBase):
top_k=config.top_k,
enable_vision=config.enable_vision,
vision_details=config.vision_details,
http_client_proxies=config.http_client,
http_client_proxies=config.http_client_proxies,
)
super().__init__(config)
+1 -1
View File
@@ -28,7 +28,7 @@ class MiniMaxLLM(LLMBase):
top_k=config.top_k,
enable_vision=config.enable_vision,
vision_details=config.vision_details,
http_client_proxies=config.http_client,
http_client_proxies=config.http_client_proxies,
)
super().__init__(config)
+1 -1
View File
@@ -30,7 +30,7 @@ class OllamaLLM(LLMBase):
top_k=config.top_k,
enable_vision=config.enable_vision,
vision_details=config.vision_details,
http_client_proxies=config.http_client,
http_client_proxies=config.http_client_proxies,
)
super().__init__(config)
+1 -1
View File
@@ -30,7 +30,7 @@ class OpenAILLM(LLMBase):
enable_vision=config.enable_vision,
vision_details=config.vision_details,
reasoning_effort=getattr(config, 'reasoning_effort', None),
http_client_proxies=config.http_client,
http_client_proxies=config.http_client_proxies,
is_reasoning_model=getattr(config, 'is_reasoning_model', None),
)
+1 -1
View File
@@ -28,7 +28,7 @@ class VllmLLM(LLMBase):
top_k=config.top_k,
enable_vision=config.enable_vision,
vision_details=config.vision_details,
http_client_proxies=config.http_client,
http_client_proxies=config.http_client_proxies,
)
super().__init__(config)
+63 -15
View File
@@ -993,6 +993,17 @@ class Memory(MemoryBase):
except Exception:
entity_embeddings.append(None)
if len(entity_embeddings) != len(ordered_keys):
logger.warning(
"embed_batch returned %d vectors for %d entity texts — "
"padding/truncating to avoid dropping entity links",
len(entity_embeddings),
len(ordered_keys),
)
entity_embeddings = list(entity_embeddings[: len(ordered_keys)])
entity_embeddings += [None] * (len(ordered_keys) - len(entity_embeddings))
# Filter out entities with failed embeddings
valid = [(i, k) for i, k in enumerate(ordered_keys) if entity_embeddings[i] is not None]
if valid:
@@ -1091,6 +1102,7 @@ class Memory(MemoryBase):
"run_id",
"actor_id",
"role",
"attributed_to",
]
core_and_promoted_keys = {"data", "hash", "created_at", "updated_at", "id", "text_lemmatized", "attributed_to", *promoted_payload_keys}
@@ -1204,6 +1216,7 @@ class Memory(MemoryBase):
"run_id",
"actor_id",
"role",
"attributed_to",
]
core_and_promoted_keys = {"data", "hash", "created_at", "updated_at", "id", "text_lemmatized", "attributed_to", *promoted_payload_keys}
@@ -1539,6 +1552,7 @@ class Memory(MemoryBase):
"run_id",
"actor_id",
"role",
"attributed_to",
]
core_and_promoted_keys = {"data", "hash", "created_at", "updated_at", "id", "text_lemmatized", "attributed_to", *promoted_payload_keys}
@@ -1831,8 +1845,9 @@ class Memory(MemoryBase):
try:
existing_memory = self.vector_store.get(vector_id=memory_id)
except Exception:
# Backing-store failure, not a bad memory_id: re-raise the original so the REST layer maps it to 5xx, not 4xx.
logger.error(f"Error getting memory with ID {memory_id} during update.")
raise ValueError(f"Error getting memory with ID {memory_id}. Please provide a valid 'memory_id'")
raise
if existing_memory is None:
raise ValueError(f"Memory with id {memory_id} not found. Please provide a valid 'memory_id'")
@@ -1923,10 +1938,8 @@ class Memory(MemoryBase):
"""
logger.warning("Resetting all memories")
if hasattr(self.db, "connection") and self.db.connection:
self.db.connection.execute("DROP TABLE IF EXISTS history")
self.db.connection.close()
self.db.reset()
self.db.close()
self.db = SQLiteManager(self.config.history_db_path)
if hasattr(self.vector_store, "reset"):
@@ -2076,6 +2089,27 @@ class AsyncMemory(MemoryBase):
except Exception as e:
logger.warning(f"Entity upsert failed for '{entity_text}' (async): {e}")
async def _bulk_clear_entity_store(self, filters):
"""Delete all entity records matching the given scope filters.
Used by delete_all to avoid the race condition that occurs when
concurrent _delete_memory coroutines each try to read-modify-write
the same entity rows' linked_memory_ids lists.
"""
if self._entity_store is None:
return
search_filters = {k: v for k, v in filters.items() if k in ("user_id", "agent_id", "run_id") and v}
try:
listed = await asyncio.to_thread(self.entity_store.list, filters=search_filters, top_k=10000)
rows = listed[0] if isinstance(listed, (list, tuple)) and listed and isinstance(listed[0], list) else listed
for row in rows or []:
try:
await asyncio.to_thread(self.entity_store.delete, vector_id=row.id)
except Exception as e:
logger.debug(f"Bulk entity delete failed for id={row.id}: {e}")
except Exception as e:
logger.warning(f"Bulk entity store cleanup failed: {e}")
async def _remove_memory_from_entity_store(self, memory_id, filters):
"""Async variant of `Memory._remove_memory_from_entity_store`."""
if self._entity_store is None:
@@ -2499,6 +2533,16 @@ class AsyncMemory(MemoryBase):
except Exception:
entity_embeddings.append(None)
if len(entity_embeddings) != len(ordered_keys):
logger.warning(
"embed_batch returned %d vectors for %d entity texts — "
"padding/truncating to avoid dropping entity links",
len(entity_embeddings),
len(ordered_keys),
)
entity_embeddings = list(entity_embeddings[: len(ordered_keys)])
entity_embeddings += [None] * (len(ordered_keys) - len(entity_embeddings))
valid = [(i, k) for i, k in enumerate(ordered_keys) if entity_embeddings[i] is not None]
if valid:
valid_indices, valid_keys = zip(*valid)
@@ -2597,6 +2641,7 @@ class AsyncMemory(MemoryBase):
"run_id",
"actor_id",
"role",
"attributed_to",
]
core_and_promoted_keys = {"data", "hash", "created_at", "updated_at", "id", "text_lemmatized", "attributed_to", *promoted_payload_keys}
@@ -2710,6 +2755,7 @@ class AsyncMemory(MemoryBase):
"run_id",
"actor_id",
"role",
"attributed_to",
]
core_and_promoted_keys = {"data", "hash", "created_at", "updated_at", "id", "text_lemmatized", "attributed_to", *promoted_payload_keys}
@@ -3051,6 +3097,7 @@ class AsyncMemory(MemoryBase):
"run_id",
"actor_id",
"role",
"attributed_to",
]
core_and_promoted_keys = {"data", "hash", "created_at", "updated_at", "id", "text_lemmatized", "attributed_to", *promoted_payload_keys}
@@ -3233,10 +3280,13 @@ class AsyncMemory(MemoryBase):
delete_tasks = []
for memory in memories[0]:
delete_tasks.append(self._delete_memory(memory.id))
delete_tasks.append(self._delete_memory(memory.id, skip_entity_cleanup=True))
results = await asyncio.gather(*delete_tasks, return_exceptions=True)
if self._entity_store is not None:
await self._bulk_clear_entity_store(filters)
errors = [r for r in results if isinstance(r, BaseException)]
if errors:
logger.warning("Failed to delete %d out of %d memories", len(errors), len(results))
@@ -3363,8 +3413,9 @@ class AsyncMemory(MemoryBase):
try:
existing_memory = await asyncio.to_thread(self.vector_store.get, vector_id=memory_id)
except Exception:
# Backing-store failure, not a bad memory_id: re-raise the original so the REST layer maps it to 5xx, not 4xx.
logger.error(f"Error getting memory with ID {memory_id} during update.")
raise ValueError(f"Error getting memory with ID {memory_id}. Please provide a valid 'memory_id'")
raise
if existing_memory is None:
raise ValueError(f"Memory with id {memory_id} not found. Please provide a valid 'memory_id'")
@@ -3418,7 +3469,7 @@ class AsyncMemory(MemoryBase):
return memory_id
async def _delete_memory(self, memory_id, existing_memory=None):
async def _delete_memory(self, memory_id, existing_memory=None, skip_entity_cleanup=False):
logger.info(f"Deleting memory with {memory_id=}")
if existing_memory is None:
existing_memory = await asyncio.to_thread(self.vector_store.get, vector_id=memory_id)
@@ -3444,9 +3495,8 @@ class AsyncMemory(MemoryBase):
is_deleted=1,
)
# Entity-store cleanup: strip this memory's id from any entity records
# that linked to it. Non-fatal — the helper swallows errors.
await self._remove_memory_from_entity_store(memory_id, session_filters)
if not skip_entity_cleanup:
await self._remove_memory_from_entity_store(memory_id, session_filters)
return memory_id
@@ -3465,10 +3515,8 @@ class AsyncMemory(MemoryBase):
if hasattr(self.vector_store, "client") and hasattr(self.vector_store.client, "close"):
await asyncio.to_thread(self.vector_store.client.close)
if hasattr(self.db, "connection") and self.db.connection:
await asyncio.to_thread(lambda: self.db.connection.execute("DROP TABLE IF EXISTS history"))
await asyncio.to_thread(self.db.connection.close)
await asyncio.to_thread(self.db.reset)
await asyncio.to_thread(self.db.close)
self.db = SQLiteManager(self.config.history_db_path)
self.vector_store = VectorStoreFactory.create(
+3 -3
View File
@@ -324,7 +324,9 @@ class SQLiteManager:
]
def reset(self) -> None:
"""Drop and recreate the history and messages tables."""
"""Drop both tables. Caller is expected to replace this instance."""
if not self.connection:
raise RuntimeError("Cannot reset a closed SQLiteManager")
with self._lock:
try:
self.connection.execute("BEGIN")
@@ -335,8 +337,6 @@ class SQLiteManager:
self.connection.execute("ROLLBACK")
logger.error(f"Failed to reset tables: {e}")
raise
self._create_history_table()
self._create_messages_table()
def close(self) -> None:
if self.connection:
+4 -1
View File
@@ -206,7 +206,10 @@ def parse_vision_messages(messages, llm=None, vision_details="auto"):
elif isinstance(content, dict) and content.get("type") == "image_url":
if llm is None:
continue
image_url = content["image_url"]["url"]
image_url_obj = content.get("image_url")
image_url = image_url_obj.get("url") if isinstance(image_url_obj, dict) else None
if not image_url:
raise ValueError("image_url content part is missing image_url.url")
try:
description = get_image_description(image_url, llm, vision_details)
returned_messages.append({"role": role, "content": description})
+11 -1
View File
@@ -4,6 +4,16 @@ Reranker implementations for mem0 search functionality.
from .base import BaseReranker
from .cohere_reranker import CohereReranker
from .huggingface_reranker import HuggingFaceReranker
from .llm_reranker import LLMReranker
from .sentence_transformer_reranker import SentenceTransformerReranker
from .zero_entropy_reranker import ZeroEntropyReranker
__all__ = ["BaseReranker", "CohereReranker", "SentenceTransformerReranker"]
__all__ = [
"BaseReranker",
"CohereReranker",
"HuggingFaceReranker",
"LLMReranker",
"SentenceTransformerReranker",
"ZeroEntropyReranker",
]
+5 -1
View File
@@ -1,3 +1,4 @@
import logging
import os
from typing import List, Dict, Any
@@ -9,6 +10,8 @@ try:
except ImportError:
COHERE_AVAILABLE = False
logger = logging.getLogger(__name__)
class CohereReranker(BaseReranker):
"""Cohere-based reranker implementation."""
@@ -78,8 +81,9 @@ class CohereReranker(BaseReranker):
return reranked_docs
except Exception:
except Exception as e:
# Fallback to original order if reranking fails
logger.warning("Cohere reranking failed, falling back to original order: %s", e)
for doc in documents:
doc['rerank_score'] = 0.0
final_top_k = top_k or self.config.top_k
+5 -1
View File
@@ -1,3 +1,4 @@
import logging
from typing import List, Dict, Any, Union
import numpy as np
@@ -12,6 +13,8 @@ try:
except ImportError:
TRANSFORMERS_AVAILABLE = False
logger = logging.getLogger(__name__)
class HuggingFaceReranker(BaseReranker):
"""HuggingFace Transformers based reranker implementation."""
@@ -139,8 +142,9 @@ class HuggingFaceReranker(BaseReranker):
return reranked_docs
except Exception:
except Exception as e:
# Fallback to original order if reranking fails
logger.warning("HuggingFace reranking failed, falling back to original order: %s", e)
for doc in documents:
doc['rerank_score'] = 0.0
final_top_k = top_k or self.config.top_k
+10 -6
View File
@@ -1,3 +1,4 @@
import logging
import re
from typing import Any, Dict, List, Union
@@ -6,6 +7,8 @@ from mem0.configs.rerankers.llm import LLMRerankerConfig
from mem0.reranker.base import BaseReranker
from mem0.utils.factory import LlmFactory
logger = logging.getLogger(__name__)
class LLMReranker(BaseReranker):
"""LLM-based reranker implementation."""
@@ -90,14 +93,14 @@ class LLMReranker(BaseReranker):
def _extract_score(self, response_text: str) -> float:
"""Extract numerical score from LLM response."""
# Look for decimal numbers between 0.0 and 1.0
pattern = r'\b([01](?:\.\d+)?)\b'
matches = re.findall(pattern, response_text)
# Prefer a decimal, fall back to an integer, then clamp: out-of-range outputs
# like "2.0"/"5" become 1.0 instead of being mis-parsed into a stray 0/1 digit.
matches = re.findall(r'-?\d+\.\d+', response_text) or re.findall(r'-?\d+', response_text)
if matches:
score = float(matches[0])
return min(max(score, 0.0), 1.0) # Clamp between 0.0 and 1.0
# Fallback: return 0.5 if no valid score found
return 0.5
@@ -151,8 +154,9 @@ class LLMReranker(BaseReranker):
scored_doc['rerank_score'] = score
scored_docs.append(scored_doc)
except Exception:
except Exception as e:
# Fallback: assign neutral score if scoring fails
logger.warning("LLM reranking failed for a document, assigning neutral score: %s", e)
scored_doc = doc.copy()
scored_doc['rerank_score'] = 0.5
scored_docs.append(scored_doc)
@@ -1,3 +1,4 @@
import logging
from typing import List, Dict, Any, Union
import numpy as np
@@ -11,6 +12,8 @@ try:
except ImportError:
SENTENCE_TRANSFORMERS_AVAILABLE = False
logger = logging.getLogger(__name__)
class SentenceTransformerReranker(BaseReranker):
"""Sentence Transformer based reranker implementation."""
@@ -102,8 +105,9 @@ class SentenceTransformerReranker(BaseReranker):
return reranked_docs
except Exception:
except Exception as e:
# Fallback to original order if reranking fails
logger.warning("SentenceTransformer reranking failed, falling back to original order: %s", e)
for doc in documents:
doc['rerank_score'] = 0.0
final_top_k = top_k or self.config.top_k
+5 -1
View File
@@ -1,3 +1,4 @@
import logging
import os
from typing import List, Dict, Any
@@ -9,6 +10,8 @@ try:
except ImportError:
ZERO_ENTROPY_AVAILABLE = False
logger = logging.getLogger(__name__)
class ZeroEntropyReranker(BaseReranker):
"""Zero Entropy-based reranker implementation."""
@@ -89,8 +92,9 @@ class ZeroEntropyReranker(BaseReranker):
return reranked_docs
except Exception:
except Exception as e:
# Fallback to original order if reranking fails
logger.warning("Zero Entropy reranking failed, falling back to original order: %s", e)
for doc in documents:
doc['rerank_score'] = 0.0
final_top_k = top_k or self.config.top_k
+8 -2
View File
@@ -352,6 +352,12 @@ def _extract_entities_from_doc(doc) -> List[Tuple[str, str]]:
best[k] = (t, e)
deduped = list(best.values())
# Remove entities that are substrings of longer entities
# Remove entities that are whole-word substrings of longer entities.
# Word-boundary anchoring avoids dropping distinct entities that only share a
# leading substring (e.g. "Sam" must survive alongside "Samsung").
all_lower = [e[1].lower() for e in deduped]
return [(t, e) for t, e in deduped if not any(e.lower() != o and e.lower() in o for o in all_lower)]
return [
(t, e)
for t, e in deduped
if not any(e.lower() != o and re.search(rf"\b{re.escape(e.lower())}\b", o) for o in all_lower)
]
+10 -1
View File
@@ -1,4 +1,5 @@
import importlib
import inspect
from typing import Dict, Optional, Union
from mem0.configs.embeddings.base import BaseEmbedderConfig
@@ -100,8 +101,16 @@ class LlmFactory:
"top_k": config.top_k,
"enable_vision": config.enable_vision,
"vision_details": config.vision_details,
"http_client_proxies": config.http_client,
"http_client_proxies": config.http_client_proxies,
}
# Only forward reasoning fields to provider configs that accept them
# (explicitly or via **kwargs); others would raise on unexpected kwargs.
params = inspect.signature(config_class).parameters
accepts_kwargs = any(p.kind == p.VAR_KEYWORD for p in params.values())
if accepts_kwargs or "reasoning_effort" in params:
config_dict["reasoning_effort"] = config.reasoning_effort
if accepts_kwargs or "is_reasoning_model" in params:
config_dict["is_reasoning_model"] = config.is_reasoning_model
config_dict.update(kwargs)
config = config_class(**config_dict)
else:
+13
View File
@@ -0,0 +1,13 @@
from typing import Dict, Optional, Union
import httpx
def build_http_client(http_client_proxies: Optional[Union[Dict, str]]) -> Optional[httpx.Client]:
if not http_client_proxies:
return None
if isinstance(http_client_proxies, dict):
return httpx.Client(
mounts={scheme: httpx.HTTPTransport(proxy=url) for scheme, url in http_client_proxies.items()}
)
return httpx.Client(proxy=http_client_proxies)
+2 -2
View File
@@ -298,9 +298,9 @@ class AzureAISearch(VectorStoreBase):
payload (Dict, optional): Updated payload.
"""
document = {"id": vector_id}
if vector:
if vector is not None:
document["vector"] = vector
if payload:
if payload is not None:
json_payload = json.dumps(payload)
document["payload"] = json_payload
for field in ["user_id", "run_id", "agent_id"]:
+13 -2
View File
@@ -1,4 +1,5 @@
import logging
import re
import time
from typing import Dict, Optional
@@ -47,6 +48,8 @@ class OutputData(BaseModel):
class BaiduDB(VectorStoreBase):
_SAFE_FILTER_KEY = re.compile(r"^[a-zA-Z_][a-zA-Z0-9_]*$")
def __init__(
self,
endpoint: str,
@@ -404,8 +407,16 @@ class BaiduDB(VectorStoreBase):
"""
conditions = []
for key, value in filters.items():
if not self._SAFE_FILTER_KEY.match(key):
raise ValueError(f"Invalid filter key: {key!r}")
if isinstance(value, str):
conditions.append(f'metadata["{key}"] = "{value}"')
else:
escaped = value.replace("\\", "\\\\").replace('"', '\\"')
conditions.append(f'metadata["{key}"] = "{escaped}"')
elif isinstance(value, (int, float, bool)):
conditions.append(f'metadata["{key}"] = {value}')
else:
raise ValueError(
f"Filter value for {key!r} must be str, int, float, or bool, "
f"got {type(value).__name__}"
)
return " AND ".join(conditions)
+8 -4
View File
@@ -169,7 +169,7 @@ class ChromaDB(VectorStoreBase):
Args:
vector_id (str): ID of the vector to delete.
"""
self.collection.delete(ids=vector_id)
self.collection.delete(ids=[vector_id])
def update(
self,
@@ -185,7 +185,11 @@ class ChromaDB(VectorStoreBase):
vector (Optional[List[float]], optional): Updated vector. Defaults to None.
payload (Optional[Dict], optional): Updated payload. Defaults to None.
"""
self.collection.update(ids=vector_id, embeddings=vector, metadatas=payload)
self.collection.update(
ids=[vector_id],
embeddings=[vector] if vector is not None else None,
metadatas=[payload] if payload is not None else None,
)
def get(self, vector_id: str) -> Optional[OutputData]:
"""
@@ -258,7 +262,7 @@ class ChromaDB(VectorStoreBase):
dict[str, any]: Properly formatted where clause for ChromaDB.
"""
if where is None:
return {}
return None
def convert_condition(key: str, value: any) -> dict:
"""Convert universal filter format to ChromaDB format."""
@@ -352,7 +356,7 @@ class ChromaDB(VectorStoreBase):
# Return appropriate format based on number of conditions
if len(processed_filters) == 0:
return {}
return None
elif len(processed_filters) == 1:
return processed_filters[0]
else:
+1 -1
View File
@@ -603,7 +603,7 @@ class FAISS(VectorStoreBase):
List[OutputData]: List of vectors.
"""
if self.index is None:
return []
return [[]]
results = []
count = 0
+14 -3
View File
@@ -1,4 +1,5 @@
import logging
import re
from typing import Dict, Optional
from pydantic import BaseModel
@@ -143,6 +144,8 @@ class MilvusDB(VectorStoreBase):
data = [_build_record(idx, embedding, metadata) for idx, embedding, metadata in zip(ids, vectors, payloads)]
self.client.insert(collection_name=self.collection_name, data=data, **kwargs)
_SAFE_FILTER_KEY = re.compile(r"^[a-zA-Z_][a-zA-Z0-9_]*$")
def _create_filter(self, filters: dict):
"""Prepare filters for efficient query.
@@ -154,10 +157,18 @@ class MilvusDB(VectorStoreBase):
"""
operands = []
for key, value in filters.items():
if not self._SAFE_FILTER_KEY.match(key):
raise ValueError(f"Invalid filter key: {key!r}")
if isinstance(value, str):
operands.append(f'(metadata["{key}"] == "{value}")')
else:
escaped = value.replace("\\", "\\\\").replace('"', '\\"')
operands.append(f'(metadata["{key}"] == "{escaped}")')
elif isinstance(value, (int, float, bool)):
operands.append(f'(metadata["{key}"] == {value})')
else:
raise ValueError(
f"Filter value for {key!r} must be str, int, float, or bool, "
f"got {type(value).__name__}"
)
return " and ".join(operands)
@@ -262,7 +273,7 @@ class MilvusDB(VectorStoreBase):
Args:
vector_id (str): ID of the vector to delete.
"""
self.client.delete(collection_name=self.collection_name, ids=vector_id)
self.client.delete(collection_name=self.collection_name, ids=[vector_id])
def update(self, vector_id=None, vector=None, payload=None):
"""
+28 -2
View File
@@ -148,6 +148,22 @@ class MongoDB(VectorStoreBase):
except PyMongoError as e:
logger.error(f"Error inserting data: {e}")
@staticmethod
def _validate_filter_value(key: str, value: Any) -> None:
"""Reject values that could inject MongoDB query operators (e.g. $ne, $gt)."""
if isinstance(value, dict):
raise ValueError(
f"Filter value for {key!r} must be a scalar (str, int, float, bool), "
f"not a dict. Dicts may contain MongoDB query operators."
)
if isinstance(value, list):
for item in value:
if isinstance(item, dict):
raise ValueError(
f"Filter list for {key!r} contains a dict, "
f"which may contain MongoDB query operators."
)
def search(self, query: str, vectors: List[float], top_k=5, filters: Optional[Dict] = None) -> List[OutputData]:
"""
Search for similar vectors using the vector search index.
@@ -167,6 +183,10 @@ class MongoDB(VectorStoreBase):
logger.error(f"Index '{self.index_name}' does not exist.")
return []
if filters:
for key, value in filters.items():
self._validate_filter_value(key, value)
results = []
try:
collection = self.client[self.db_name][self.collection_name]
@@ -215,6 +235,9 @@ class MongoDB(VectorStoreBase):
Returns:
List[OutputData]: Search results, or None if Atlas Search index is not available.
"""
if filters:
for key, value in filters.items():
self._validate_filter_value(key, value)
try:
collection = self.client[self.db_name][self.collection_name]
search_index_name = f"{self.collection_name}_text_search_index"
@@ -369,6 +392,9 @@ class MongoDB(VectorStoreBase):
Returns:
List[OutputData]: List of vectors.
"""
if filters:
for key, value in filters.items():
self._validate_filter_value(key, value)
try:
query = {}
if filters:
@@ -382,10 +408,10 @@ class MongoDB(VectorStoreBase):
cursor = self.collection.find(query).limit(top_k)
results = [OutputData(id=str(doc["_id"]), score=None, payload=doc.get("payload")) for doc in cursor]
logger.info(f"Retrieved {len(results)} documents from collection '{self.collection_name}'.")
return results
return [results]
except PyMongoError as e:
logger.error(f"Error listing documents: {e}")
return []
return [[]]
def reset(self):
"""Reset the index by deleting and recreating it."""
+11 -3
View File
@@ -39,6 +39,8 @@ class OpenSearchDB(VectorStoreBase):
self.collection_name = config.collection_name
self.embedding_model_dims = config.embedding_model_dims
self.auto_refresh = config.auto_refresh
self.create_col(self.collection_name, self.embedding_model_dims)
def create_index(self) -> None:
@@ -148,8 +150,6 @@ class OpenSearchDB(VectorStoreBase):
}
try:
self.client.index(index=self.collection_name, body=body)
# Force refresh to make documents immediately searchable for tests
self.client.indices.refresh(index=self.collection_name)
results.append(
OutputData(
@@ -162,6 +162,14 @@ class OpenSearchDB(VectorStoreBase):
logger.error(f"Error inserting vector {id_}: {e}", exc_info=True)
raise
# Refresh once after the full batch (not per document) if explicitly enabled.
# Disabled by default for Serverless compatibility: OpenSearch Serverless does not
# support the indices.refresh() API, and refreshing per document would cause a
# cluster-level I/O stall on every insert.
# See: https://docs.aws.amazon.com/opensearch-service/latest/developerguide/serverless-genref.html
if self.auto_refresh:
self.client.indices.refresh(index=self.collection_name)
return results
def search(
@@ -371,7 +379,7 @@ class OpenSearchDB(VectorStoreBase):
return [results] # VectorStore expects tuple/list format
except Exception as e:
logger.error(f"Error listing vectors: {e}", exc_info=True)
return []
return [[]]
def reset(self):
"""Reset the index by deleting and recreating it."""
+21 -3
View File
@@ -186,6 +186,17 @@ class PineconeDB(VectorStoreBase):
return result
OPERATOR_MAP = {
"eq": "$eq",
"ne": "$ne",
"gt": "$gt",
"gte": "$gte",
"lt": "$lt",
"lte": "$lte",
"in": "$in",
"nin": "$nin",
}
def _create_filter(self, filters: Optional[Dict]) -> Dict:
"""
Create a filter dictionary from the provided filters.
@@ -196,8 +207,15 @@ class PineconeDB(VectorStoreBase):
pinecone_filter = {}
for key, value in filters.items():
if isinstance(value, dict) and "gte" in value and "lte" in value:
pinecone_filter[key] = {"$gte": value["gte"], "$lte": value["lte"]}
if isinstance(value, dict):
condition = {}
for op, operand in value.items():
pc_op = self.OPERATOR_MAP.get(op)
if pc_op:
condition[pc_op] = operand
else:
condition[f"${op}"] = operand
pinecone_filter[key] = condition
else:
pinecone_filter[key] = {"$eq": value}
@@ -391,7 +409,7 @@ class PineconeDB(VectorStoreBase):
return [results]
except Exception as e:
logger.error(f"Error listing vectors: {e}")
return {"points": [], "next_page_token": None}
return [[]]
def count(self) -> int:
"""
+44 -7
View File
@@ -98,7 +98,10 @@ class Qdrant(VectorStoreBase):
self._bm25_encoder = SparseTextEmbedding(model_name="Qdrant/bm25")
logger.info("BM25 encoder loaded (fastembed Qdrant/bm25)")
except ImportError:
logger.warning("fastembed not installed — BM25 keyword search disabled. Install with: pip install fastembed")
logger.warning(
"fastembed not installed - BM25 keyword search disabled. "
'Install it with: pip install "mem0ai[extras]"'
)
self._bm25_encoder = False # sentinel: tried and failed
except Exception as e:
logger.warning(f"Failed to load BM25 encoder: {e}")
@@ -191,6 +194,43 @@ class Qdrant(VectorStoreBase):
ids (list, optional): List of IDs corresponding to vectors. Defaults to None.
"""
logger.info(f"Inserting {len(vectors)} vectors into collection {self.collection_name}")
# Pre-compute BM25 sparse vectors in a single batch call. fastembed's
# embed() accepts a list of texts, so batching avoids per-row encoder
# overhead (model dispatch, tokenizer setup, etc.).
bm25_sparse_vectors: list[Optional[SparseVector]] = [None] * len(vectors)
if self._has_bm25_slot and payloads:
texts_for_bm25: list[str] = []
indices_for_bm25: list[int] = []
for idx, payload in enumerate(payloads):
text = payload.get("text_lemmatized") or payload.get("data", "")
if text:
texts_for_bm25.append(text)
indices_for_bm25.append(idx)
if texts_for_bm25:
encoder = self._get_bm25_encoder()
if encoder is not None:
try:
sparse_results = list(encoder.embed(texts_for_bm25))
if len(sparse_results) != len(texts_for_bm25):
logger.warning(
f"BM25 batch returned {len(sparse_results)} results for "
f"{len(texts_for_bm25)} texts; falling back to per-row encoding"
)
raise ValueError("count mismatch")
for i, sparse in enumerate(sparse_results):
bm25_sparse_vectors[indices_for_bm25[i]] = SparseVector(
indices=sparse.indices.tolist(),
values=sparse.values.tolist(),
)
except Exception as e:
# Fall back to per-row encoding so a single bad input
# doesn't drop BM25 for the whole batch.
logger.debug(f"Batch BM25 encoding failed, falling back to per-row: {e}")
for i, text in enumerate(texts_for_bm25):
bm25_sparse_vectors[indices_for_bm25[i]] = self._encode_bm25(text)
points = []
for idx, vector in enumerate(vectors):
payload = payloads[idx] if payloads else {}
@@ -198,12 +238,8 @@ class Qdrant(VectorStoreBase):
# Build named vectors: dense + optional BM25 sparse (only if collection has the slot).
named_vectors = {"": vector}
if self._has_bm25_slot:
text_for_bm25 = payload.get("text_lemmatized") or payload.get("data", "")
if text_for_bm25:
sparse = self._encode_bm25(text_for_bm25)
if sparse is not None:
named_vectors["bm25"] = sparse
if self._has_bm25_slot and bm25_sparse_vectors[idx] is not None:
named_vectors["bm25"] = bm25_sparse_vectors[idx]
points.append(PointStruct(id=point_id, vector=named_vectors, payload=payload))
@@ -476,6 +512,7 @@ class Qdrant(VectorStoreBase):
if self._has_bm25_slot:
text_for_bm25 = payload.get("text_lemmatized") or payload.get("data", "")
if text_for_bm25:
# Single-item update: per-row encoding is correct here; see insert() for the batch path.
sparse = self._encode_bm25(text_for_bm25)
if sparse is not None:
named_vectors["bm25"] = sparse
+7 -3
View File
@@ -1,3 +1,4 @@
import copy
import json
import logging
from datetime import datetime, timezone
@@ -65,7 +66,7 @@ class RedisDB(VectorStoreBase):
"prefix": f"mem0:{collection_name}",
}
fields = DEFAULT_FIELDS.copy()
fields = copy.deepcopy(DEFAULT_FIELDS)
fields[-1]["attrs"]["dims"] = embedding_model_dims
self.schema = {"index": index_schema, "fields": fields}
@@ -98,8 +99,9 @@ class RedisDB(VectorStoreBase):
"prefix": f"mem0:{collection_name}",
}
# Copy the default fields and update the vector field with the specified dimensions
fields = DEFAULT_FIELDS.copy()
# Deep-copy the default fields so mutating the nested vector attrs never
# leaks into the module global or other instances.
fields = copy.deepcopy(DEFAULT_FIELDS)
fields[-1]["attrs"]["dims"] = embedding_dims
fields[-1]["attrs"]["distance_metric"] = distance_metric
@@ -263,6 +265,8 @@ class RedisDB(VectorStoreBase):
def get(self, vector_id):
result = self.index.fetch(vector_id)
if result is None:
return None
payload = {
"hash": result["hash"],
"data": result["memory"],
@@ -487,7 +487,7 @@ class GoogleMatchingEngine(VectorStoreBase):
# Use a large top_k if none specified
search_limit = top_k if top_k is not None else 10000
results = self.search(query=zero_vector, top_k=search_limit, filters=filters)
results = self.search(query="", vectors=zero_vector, top_k=search_limit, filters=filters)
logger.debug("Found %d results", len(results))
return [results] # Wrap in extra array to match interface
@@ -620,7 +620,7 @@ class GoogleMatchingEngine(VectorStoreBase):
logger.debug("Filter: %s", filter)
embedding = self.embedder.embed_query(query)
results = self.search(query=embedding, top_k=k, filters=filter)
results = self.search(query=query, vectors=embedding, top_k=k, filters=filter)
docs_and_scores = [
(Document(page_content=result.payload.get("text", ""), metadata=result.payload), result.score)
-1
View File
@@ -342,7 +342,6 @@ class Weaviate(VectorStoreBase):
"""
collections = self.client.collections.list_all()
logger.debug(f"collections: {collections}")
print(f"collections: {collections}")
return {"collections": [{"name": col.name} for col in collections]}
def delete_col(self):
+1 -1
View File
@@ -1,7 +1,7 @@
fastapi>=0.68.0
uvicorn>=0.15.0
sqlalchemy>=1.4.0
python-dotenv>=0.19.0
python-dotenv>=1.2.2
alembic>=1.7.0
psycopg2-binary>=2.9.0
python-multipart>=0.0.27
+1
View File
@@ -81,6 +81,7 @@
"sharp"
],
"overrides": {
"form-data@<4.0.6": ">=4.0.6",
"immutable@>=5.0.0 <5.1.5": "^5.1.5",
"picomatch@<2.3.2": "^2.3.2",
"minimatch@>=9.0.0 <9.0.7": "^9.0.7",
+5 -4
View File
@@ -5,6 +5,7 @@ settings:
excludeLinksFromLockfile: false
overrides:
form-data@<4.0.6: '>=4.0.6'
immutable@>=5.0.0 <5.1.5: ^5.1.5
picomatch@<2.3.2: ^2.3.2
minimatch@>=9.0.0 <9.0.7: ^9.0.7
@@ -1476,8 +1477,8 @@ packages:
debug:
optional: true
form-data@4.0.5:
resolution: {integrity: sha512-8RipRLol37bNs2bhoV67fiTEvdTrbMUYcFTiy3+wuuOnUog2QBHCZWXDRijWQfAkhBj2Uf5UnVaiWwA5vdd82w==}
form-data@4.0.6:
resolution: {integrity: sha512-vKatAh4SlVfgbv+YtmhiRjhEMJsYpsG1Y2rMQtR+SVSbytsSD1YGzDIcrAJmdFec88u/+VoGmxnl+80gL1tRCQ==}
engines: {node: '>= 6'}
fraction.js@5.3.4:
@@ -3032,7 +3033,7 @@ snapshots:
axios@1.17.0:
dependencies:
follow-redirects: 1.16.0
form-data: 4.0.5
form-data: 4.0.6
https-proxy-agent: 5.0.1
proxy-from-env: 2.1.0
transitivePeerDependencies:
@@ -3235,7 +3236,7 @@ snapshots:
follow-redirects@1.16.0: {}
form-data@4.0.5:
form-data@4.0.6:
dependencies:
asynckit: 0.4.0
combined-stream: 1.0.8
+1
View File
@@ -6,6 +6,7 @@ onlyBuiltDependencies:
- sharp
overrides:
"form-data@<4.0.6": ">=4.0.6"
"immutable@>=5.0.0 <5.1.5": "^5.1.5"
"picomatch@<2.3.2": "^2.3.2"
"minimatch@>=9.0.0 <9.0.7": "^9.0.7"
+1
View File
@@ -17,6 +17,7 @@ dependencies = [
"qdrant-client>=1.12.0",
"pydantic>=2.7.3",
"openai>=1.90.0",
"httpx>=0.28.0",
"posthog>=7.14.0",
"pytz>=2024.1",
"sqlalchemy>=2.0.31",

Some files were not shown because too many files have changed in this diff Show More