Compare commits

..

1 Commits

Author SHA1 Message Date
Saket Aryan df70d7833f fix(plugins): stamp surface identity at record time, not send time
Six defects in 0.3.x plugin telemetry. Defects 1, 2 and 6 were not three bugs:
they were one spool protocol getting three properties wrong.

Identity was decided by the wrong process. `harness` was stamped in record(),
correctly, but `source` was read from a module global in flush() — so whichever
process drained the spool named every event in it. Two processes never call
init(): mcp_server.py, and the detached `python3 telemetry.py` sender that
spawn_flush() starts. record() now stamps source beside harness, and the build
generates core/_harness_id.py per host so identity resolves with no init() call
at all. That also unifies two defaults that disagreed (`<host>_plugin` vs
`MEM0_<HOST>_PLUGIN`), which could yield three source values for one plugin.

Ownership was inferred, not held. Path.replace is os.rename, which preserves
mtime, so a claim made after a quiet minute inherited the spool's age and was
stealable the instant it existed. Claims are touched on creation and the
per-batch rewrite doubles as a lease heartbeat.

Progress was not durable. flush() returned on the first failed batch without
truncating, so the retry re-posted from index 0 — 150 events delivered 250
times. It now rewrites the claim with the unsent remainder after every batch,
bounding a crash to one repeated batch, and each event carries a uuid.

Parked batches starved. They were only reachable when no spool existed, and
because sessions keep recording there usually was one, so a batch parked by a
failed send waited until the 7-day expiry deleted it unsent — despite its own
presence being what starts the sender. flush() drains them in the same run, and
expiry now applies only after a genuine retry has failed.

code.install counted upgrades and repeat sessions. is_first_run() read the
identity file, which only a successful flush writes, so an offline user recorded
an install every session forever. A dedicated install-state.json is claimed
atomically at record time; a non-empty data directory reads as an upgrade.

The docs called this anonymous. Every event carries the account email, and the
hashes were unsalted SHA-256 over a git remote URL or an absolute path
containing the username. READMEs, the module docstring and a new docs section
now say what the code does, and repo/session digests are salted per install.

A cached email outlived an API key change. It is now re-resolved when the key's
fingerprint differs, and $identify aliases anonymous->email only — aliasing one
account to another merges person profiles irreversibly.

All six shipped green because the shared core's only tests lived under one host,
behind a conftest that calls init() at import. Core behaviour was never
exercised uninitialised. Adds agent-plugin-core/tests with no init, including
subprocess tests and coverage for the portable plugin, which has no flush worker
and would pass a native-only test vacuously.

Also puts the three surface headers on the SDKs, CLIs and integrations, and
corrects a README claiming ZAPIER/STRANDS were already in the platform allowlist.

Verified: 59 core tests, 203 claude-code, 11 cursor, 5 codex, 2 kimi, 6
antigravity. ruff and compileall clean. --check clean for all six hosts.
TypeScript changes are not typechecked locally (deps not installed).

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:14:11 +05:30
155 changed files with 680 additions and 4762 deletions
+1 -1
View File
@@ -12,7 +12,7 @@
"name": "mem0",
"source": "./integrations/claude-code-plugin",
"description": "Cross-session memory and token savings for coding agents.",
"version": "0.3.2"
"version": "0.3.1"
}
]
}
+1 -1
View File
@@ -12,7 +12,7 @@
"name": "mem0",
"source": "./integrations/cursor-plugin",
"description": "Cross-session memory and token savings for coding agents.",
"version": "0.3.2"
"version": "0.3.1"
}
]
}
@@ -32,9 +32,6 @@ jobs:
- name: Type check
run: bun run type-check
- name: Test
run: bun test
- name: Build
run: bun run build
+1 -1
View File
@@ -5,7 +5,7 @@
{
"id": "mem0",
"displayName": "Mem0",
"version": "0.3.2",
"version": "0.3.1",
"description": "Cross-session memory and token savings for coding agents.",
"homepage": "https://mem0.ai",
"keywords": ["memory", "personalization", "mcp", "semantic-search"],
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@mem0/cli",
"version": "0.2.14",
"version": "0.2.13",
"description": "The official CLI for mem0 — the memory layer for AI agents",
"type": "module",
"bin": {
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "mem0-cli"
version = "0.2.13"
version = "0.2.12"
description = "The official CLI for mem0 — the memory layer for AI agents"
readme = "README.md"
license = "Apache-2.0"
+1 -1
View File
@@ -1,3 +1,3 @@
"""mem0 CLI — the command-line interface for the mem0 memory layer."""
__version__ = "0.2.13"
__version__ = "0.2.12"
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "API Reference Overview"
sidebarTitle: "Overview"
title: "Overview"
seo:
title: "API Reference Overview - Mem0"
icon: "terminal"
iconType: "solid"
description: "REST APIs for memory management, search, and entity operations"
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Delete Memory API Endpoint"
sidebarTitle: "Delete Memory"
title: 'Delete Memory'
seo:
title: "Delete Memory API Endpoint - Mem0"
description: "Delete a single memory by its unique memory ID from the Mem0 platform using the DELETE endpoint."
openapi: delete /v1/memories/{memory_id}/
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Update Memory API Endpoint"
sidebarTitle: "Update Memory"
title: 'Update Memory'
seo:
title: "Update Memory API Endpoint - Mem0"
description: "Update the content, metadata, timestamp, or expiration date of a single memory by its unique ID using the PUT endpoint."
openapi: put /v1/memories/{memory_id}/
---
@@ -1,6 +1,7 @@
---
title: "Add Organization Member API Endpoint"
sidebarTitle: "Add Member"
title: 'Add Member'
seo:
title: "Add Organization Member API Endpoint - Mem0"
description: "Add a new member to an organization with a specified role such as READER or OWNER access level."
openapi: post /api/v1/orgs/organizations/{org_id}/members/
---
@@ -1,6 +1,7 @@
---
title: "Get Organization Members API Endpoint"
sidebarTitle: "Get Members"
title: 'Get Members'
seo:
title: "Get Organization Members API Endpoint - Mem0"
description: "Retrieve a list of all members belonging to a specific organization on the Mem0 platform."
openapi: get /api/v1/orgs/organizations/{org_id}/members/
---
@@ -1,6 +1,7 @@
---
title: "Add Project Member API Endpoint"
sidebarTitle: "Add Member"
title: 'Add Member'
seo:
title: "Add Project Member API Endpoint - Mem0"
description: "Add a new member to a project with a specified role such as READER or OWNER access level."
openapi: post /api/v1/orgs/organizations/{org_id}/projects/{project_id}/members/
---
@@ -1,6 +1,7 @@
---
title: "Get Project Members API Endpoint"
sidebarTitle: "Get Members"
title: 'Get Members'
seo:
title: "Get Project Members API Endpoint - Mem0"
description: "Retrieve a list of all members belonging to a specific project on the Mem0 platform."
openapi: get /api/v1/orgs/organizations/{org_id}/projects/{project_id}/members/
---
+32 -93
View File
@@ -7,14 +7,6 @@ mode: "wide"
<Tabs>
<Tab title="Python">
<Update label="2026-09-18" description="v2.1.0">
**Improvements:**
- **Client:** Requests now carry three surface-identity headers so the platform can tell which product made a call. `X-Mem0-Source` names the surface and `X-Application` the host app it runs inside, both set-once so a wrapper that already declared its identity keeps it. `X-Mem0-Client` is append-only and carries `name/version` per layer, outermost first, so a plugin calling this SDK reports the whole chain rather than only the last speaker. `MEM0_SOURCE`, `MEM0_APPLICATION` and `MEM0_CLIENT_STACK` set them from the environment for wrappers that cannot pass options ([#7326](https://github.com/mem0ai/mem0/pull/7326))
- **Client:** The client stack is bounded by dropping whole entries rather than slicing characters, and this SDK's own entry is the reserved one. Truncating the joined string could sever an identifier mid-name and the platform parsed the fragment as a real client ([#7326](https://github.com/mem0ai/mem0/pull/7326))
</Update>
<Update label="2026-09-02" description="v2.0.20">
**Improvements:**
@@ -1235,14 +1227,6 @@ See the [OSS v2 to v3 migration guide](https://docs.mem0.ai/migration/oss-v2-to-
<Tab title="TypeScript">
<Update label="2026-09-18" description="v3.2.0">
**Improvements:**
- **Client:** Requests now carry `X-Mem0-Source`, `X-Application` and `X-Mem0-Client`, matching the Python SDK. The first two are set-once so an outer wrapper keeps its identity; the third is append-only and reports the whole layer chain. Read from `MEM0_SOURCE`, `MEM0_APPLICATION` and `MEM0_CLIENT_STACK` when set ([#7326](https://github.com/mem0ai/mem0/pull/7326))
- **Client:** The SDK version in `X-Mem0-Client` is injected at build time rather than hardcoded, so it cannot go stale at the next release ([#7326](https://github.com/mem0ai/mem0/pull/7326))
</Update>
<Update label="2026-09-02" description="v3.1.8">
**Improvements:**
@@ -1880,13 +1864,6 @@ See the [OSS v2 to v3 migration guide](https://docs.mem0.ai/migration/oss-v2-to-
<Tab title="CLI">
<Update label="2026-09-18" description="Python v0.2.13 / Node v0.2.14">
**Improvements:**
- **Client:** Requests now carry the three surface-identity headers (`X-Mem0-Source`, `X-Application`, `X-Mem0-Client`) introduced in the Python and TypeScript SDKs, so the platform can attribute calls made through the CLI to the correct surface and version ([#7326](https://github.com/mem0ai/mem0/pull/7326))
</Update>
<Update label="2026-08-24" description="Python v0.2.12 / Node v0.2.13">
**New Features:**
@@ -2099,7 +2076,7 @@ A full-featured command-line interface for Mem0, available in both Python and No
- New Git repository writes use a hash of the remote identity in `agent_id`. Search and explicit shared-memory deletion include both current and legacy repository IDs within the repository's `app_id`. Existing memories are not rewritten. Legacy IDs retain their original ambiguity for matching owner/repository names on different Git hosts.
**Packaging:**
- Claude Code, Cursor, Codex, Kimi, Antigravity, and the portable Python bundle are versioned at `0.3.1`. OpenCode and DeepSeek Harness are `0.3.0`; Pi Agent is `0.3.0`; OpenClaw is `1.1.0`. Each host's changes and upgrade considerations are listed in its tab.
- Claude Code, Cursor, Codex, Kimi, Antigravity, and the portable Python bundle are versioned at `0.3.1`. OpenCode, Pi Agent, and DeepSeek Harness are `0.3.0`; OpenClaw is `1.1.0`. Each host's changes and upgrade considerations are listed in its tab.
- Python and TypeScript CI run their respective runtime suites. Package checks build the installable artifacts, check generated-file consistency, and reject TypeScript output that still imports monorepo source.
[#7203](https://github.com/mem0ai/mem0/pull/7203)
@@ -2394,14 +2371,9 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
<Tab title="Claude Code">
<Update label="2026-09-18" description="Claude Code plugin v0.3.2">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Sidekick is now available only in Claude Code, with Sonnet, worktree isolation, and parent memories.
Sidekick is now available only in Claude Code, with Sonnet, worktree isolation, and parent memories.
</Update>
@@ -2424,14 +2396,9 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
<Tab title="Cursor">
<Update label="2026-09-18" description="Cursor plugin v0.3.2">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Removes Sidekick and its start/stop hooks. Memory capture, search, and six skills remain available.
Removes Sidekick and its start/stop hooks. Memory capture, search, and six skills remain available.
</Update>
@@ -2454,14 +2421,9 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
<Tab title="Codex">
<Update label="2026-09-18" description="Codex plugin v0.3.2">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Renames shared tracking to use subagent terminology. Native subagent memory support remains available.
Renames shared tracking to use subagent terminology. Native subagent memory support remains available.
</Update>
@@ -2481,16 +2443,32 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
</Tab>
<Tab title="OpenCode">
<Tab title="Agent Plugins v1">
<Update label="2026-09-18" description="OpenCode plugin v0.4.0">
<Update label="Unreleased" description="Sidekick availability">
**Changes:**
- **Telemetry:** The PostHog `source` tag changed from the literal `"plugin"` to `OPENCODE_PLUGIN`, and `project_hash` is now salted. Saved PostHog insights filtering on `source = "plugin"` will stop matching new events; historical data is unaffected ([#7322](https://github.com/mem0ai/mem0/pull/7322))
- **Config:** A new `keyFingerprint` key appears in the install-count deduplication logic; installs are now counted once per key rather than on every activation ([#7325](https://github.com/mem0ai/mem0/pull/7325))
Sidekick is available only in Claude Code, not in the portable package.
</Update>
<Update label="2026-09-08" description="Portable Mem0 plugin v0.3.1">
**Added:**
- One portable package at `integrations/mem0-agent-plugin/`, using the Agent Plugins 1.0.0 root `plugin.json`, `mcp.json`, and fixed `skills/` locations.
- Ships a local, read-only `search_memories` server and the six shared memory skills. Uses `PLUGIN_ROOT` for bundled files and `PLUGIN_DATA` for persistent plugin state; all package files remain inside the installable directory.
**Packaging:**
- Generated from the shared Python runtime and skill templates. Builds validate the manifest, MCP configuration, skills, and generated-file consistency.
- Host lifecycle hooks and native Sidekick declarations remain in the native plugin packages; the portable package does not provide automatic lifecycle capture or host-specific subagent isolation. Its bundled remember skill cannot persist a new memory on its own because the portable package has no capture hooks or write tool.
[#7203](https://github.com/mem0ai/mem0/pull/7203)
</Update>
</Tab>
<Tab title="OpenCode">
<Update label="2026-09-08" description="OpenCode plugin v0.3.0">
**Changed:**
@@ -2589,14 +2567,9 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
<Tab title="Antigravity">
<Update label="2026-09-18" description="Antigravity plugin v0.3.2">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Removes Sidekick. Memory capture, search, and six skills remain available.
Removes Sidekick. Memory capture, search, and six skills remain available.
</Update>
@@ -2692,14 +2665,9 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="Kimi">
<Update label="2026-09-18" description="Kimi Code plugin v0.3.2">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Removes Sidekick and its start/stop hooks. Memory capture, recall, and six skills remain available.
Removes Sidekick and its start/stop hooks. Memory capture, recall, and six skills remain available.
</Update>
@@ -2737,14 +2705,6 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="OpenClaw">
<Update label="2026-09-18" description="openclaw-mem0 v1.2.0">
**Changes:**
- **Config:** Added `keyFingerprint` to the config schema for install-count deduplication; installs are now counted once per key rather than on every activation ([#7325](https://github.com/mem0ai/mem0/pull/7325))
- **Telemetry:** Events are no longer delivered twice, and the `plugin_version` field now reflects the plugin that produced the event ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
</Update>
<Update label="2026-09-08" description="openclaw-mem0 v1.1.0">
**Changed:**
@@ -3033,13 +2993,6 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="Pi Agent">
<Update label="2026-09-18" description="Pi Agent plugin v0.3.1">
**Improvements:**
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events, and now attribute each event to the plugin that produced it. The `plugin_version` wire field reflects the fixed release ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
</Update>
<Update label="2026-09-08" description="Pi Agent plugin v0.3.0">
**Changed:**
@@ -3153,13 +3106,6 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="DeepSeek Harness">
<Update label="2026-09-18" description="deepseek-plugin v0.3.1">
**Improvements:**
- **Telemetry:** Rebuild with the fixed shared telemetry core from `agent-plugin-core`. Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
</Update>
<Update label="2026-09-08" description="deepseek-plugin v0.3.0">
**Added:**
@@ -3204,13 +3150,6 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="Vercel AI SDK">
<Update label="2026-09-18" description="Vercel AI SDK v3.0.3">
**Improvements:**
- **Client:** Inherits the three surface-identity headers (`X-Mem0-Source`, `X-Application`, `X-Mem0-Client`) from the TypeScript SDK bump, so platform calls made through the Vercel AI SDK provider are now correctly attributed ([#7326](https://github.com/mem0ai/mem0/pull/7326))
</Update>
<Update label="2026-08-24" description="Vercel AI SDK v3.0.2">
**Security:**
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Embedder Configuration Reference"
sidebarTitle: "Configurations"
title: Configurations
seo:
title: "Embedder Configuration Reference - Mem0"
description: "Reference for embedder configuration options in Mem0, including provider selection and model settings."
---
@@ -1,6 +1,7 @@
---
title: "AWS Bedrock as Embedding Provider"
sidebarTitle: "AWS Bedrock"
title: AWS Bedrock
seo:
title: "AWS Bedrock as Embedding Provider - Mem0"
description: "Configure AWS Bedrock as an embedding provider in Mem0 with IAM credentials and boto3 authentication."
---
@@ -1,6 +1,7 @@
---
title: "Azure OpenAI as Embedding Provider"
sidebarTitle: "Azure OpenAI"
title: Azure OpenAI
seo:
title: "Azure OpenAI as Embedding Provider - Mem0"
description: "Configure Azure OpenAI as an embedding provider in Mem0 with API key, deployment, and endpoint settings."
---
@@ -1,6 +1,7 @@
---
title: "Google AI as Embedding Provider"
sidebarTitle: "Google AI"
title: Google AI
seo:
title: "Google AI as Embedding Provider - Mem0"
description: "Configure Google AI as an embedding provider in Mem0 using Gemini models and the GOOGLE_API_KEY variable."
---
@@ -1,6 +1,7 @@
---
title: "LangChain as Embedding Provider"
sidebarTitle: "LangChain"
title: LangChain
seo:
title: "LangChain as Embedding Provider - Mem0"
description: "Use LangChain as an embedding provider in Mem0 to access a wide range of models through a unified interface."
---
@@ -1,6 +1,7 @@
---
title: "LM Studio as Embedding Provider"
sidebarTitle: "LM Studio"
title: "LM Studio"
seo:
title: "LM Studio as Embedding Provider - Mem0"
description: "Configure LM Studio as an embedding provider in Mem0 for local embedding generation with models like nomic-embed-text."
---
You can use embedding models from LM Studio to run Mem0 locally.
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Ollama as Embedding Provider"
sidebarTitle: "Ollama"
title: "Ollama"
seo:
title: "Ollama as Embedding Provider - Mem0"
description: "Configure Ollama as an embedding provider in Mem0 to generate embeddings locally using open-source models."
---
You can use embedding models from Ollama to run Mem0 locally.
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "OpenAI as Embedding Provider"
sidebarTitle: "OpenAI"
title: OpenAI
seo:
title: "OpenAI as Embedding Provider - Mem0"
description: "Configure OpenAI as an embedding provider in Mem0 using models like text-embedding-3-large for vector generation."
---
@@ -1,6 +1,7 @@
---
title: "Together AI as Embedding Provider"
sidebarTitle: "Together"
title: Together
seo:
title: "Together AI as Embedding Provider - Mem0"
description: "Configure Together AI as an embedding provider in Mem0 with support for 1024-dimensional embedding models."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Embedding Providers Overview"
sidebarTitle: Overview
title: Overview
seo:
title: "Embedding Providers Overview - Mem0"
description: "Overview of all supported embedding model providers in Mem0, including OpenAI, Azure, Ollama, and more."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "LLM Configuration Reference"
sidebarTitle: "Configurations"
title: Configurations
seo:
title: "LLM Configuration Reference - Mem0"
description: "Reference for LLM configuration options in Mem0 for Python and TypeScript, including value precedence rules."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "AWS Bedrock as LLM Provider"
sidebarTitle: "AWS Bedrock"
title: AWS Bedrock
seo:
title: "AWS Bedrock as LLM Provider - Mem0"
description: "Configure AWS Bedrock as an LLM provider in Mem0 with IAM authentication and Claude model support."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Azure OpenAI as LLM Provider"
sidebarTitle: "Azure OpenAI"
title: Azure OpenAI
seo:
title: "Azure OpenAI as LLM Provider - Mem0"
description: "Configure Azure OpenAI as an LLM provider in Mem0 with Azure Identity authentication and deployment settings."
---
+1 -1
View File
@@ -3,7 +3,7 @@ title: DeepSeek
description: "Configure DeepSeek as an LLM provider in Mem0 with API key setup and optional custom endpoint configuration."
---
To use DeepSeek LLM models, you have to set the `DEEPSEEK_API_KEY` environment variable. You can also optionally set `DEEPSEEK_API_BASE` if you need to use a different API endpoint (defaults to `https://api.deepseek.com`).
To use DeepSeek LLM models, you have to set the `DEEPSEEK_API_KEY` environment variable. You can also optionally set `DEEPSEEK_API_BASE` if you need to use a different API endpoint (defaults to "https://api.deepseek.com").
## Usage
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Google AI as LLM Provider"
sidebarTitle: "Google AI"
title: Google AI
seo:
title: "Google AI as LLM Provider - Mem0"
description: "Configure Google Gemini as an LLM provider in Mem0 using the google.genai SDK and GOOGLE_API_KEY variable."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "LangChain as LLM Provider"
sidebarTitle: "LangChain"
title: LangChain
seo:
title: "LangChain as LLM Provider - Mem0"
description: "Use LangChain as an LLM provider in Mem0 to integrate with various chat models through a unified interface."
---
+4 -3
View File
@@ -1,6 +1,7 @@
---
title: "LM Studio as LLM Provider"
sidebarTitle: "LM Studio"
title: LM Studio
seo:
title: "LM Studio as LLM Provider - Mem0"
description: "Configure LM Studio as an LLM provider in Mem0 for running local language models via an OpenAI-compatible API."
---
@@ -77,7 +78,7 @@ m.add(messages, user_id="alice123", metadata={"category": "movies"})
To use LM Studio, you need to:
1. Download and install [LM Studio](https://lmstudio.ai/)
2. Start a local server from the "Server" tab
3. Set the appropriate `lmstudio_base_url` in your configuration (default is usually `http://localhost:1234/v1`)
3. Set the appropriate `lmstudio_base_url` in your configuration (default is usually http://localhost:1234/v1)
</Note>
## Config
+1 -1
View File
@@ -3,7 +3,7 @@ title: MiniMax
description: "Configure MiniMax as an LLM provider in Mem0 with API key setup and optional custom endpoint configuration."
---
To use MiniMax LLM models, you have to set the `MINIMAX_API_KEY` environment variable. You can also optionally set `MINIMAX_API_BASE` if you need to use a different API endpoint (defaults to `https://api.minimax.io/v1`).
To use MiniMax LLM models, you have to set the `MINIMAX_API_KEY` environment variable. You can also optionally set `MINIMAX_API_BASE` if you need to use a different API endpoint (defaults to "https://api.minimax.io/v1").
## Usage
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Ollama as LLM Provider"
sidebarTitle: "Ollama"
title: Ollama
seo:
title: "Ollama as LLM Provider - Mem0"
description: "Configure Ollama as an LLM provider in Mem0 for running local language models with tool-calling support."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "OpenAI as LLM Provider"
sidebarTitle: "OpenAI"
title: OpenAI
seo:
title: "OpenAI as LLM Provider - Mem0"
description: "Configure OpenAI as an LLM provider in Mem0 with support for GPT models and Openrouter compatibility."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Together AI as LLM Provider"
sidebarTitle: "Together"
title: Together
seo:
title: "Together AI as LLM Provider - Mem0"
description: "Configure Together AI as an LLM provider in Mem0 with API key setup and optional custom endpoint configuration."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "xAI Grok as LLM Provider"
sidebarTitle: "xAI"
title: xAI
seo:
title: "xAI Grok as LLM Provider - Mem0"
description: "Configure xAI Grok models as an LLM provider in Mem0 with API key setup and usage examples."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "LLM Providers Overview"
sidebarTitle: Overview
title: Overview
seo:
title: "LLM Providers Overview - Mem0"
description: "Overview of all supported LLM providers in Mem0, including OpenAI, Anthropic, Groq, Ollama, and more."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Reranker Providers Overview"
sidebarTitle: "Overview"
title: Overview
seo:
title: "Reranker Providers Overview - Mem0"
description: 'Pick the right reranker path to boost Mem0 search relevance.'
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Vector Store Configuration Reference"
sidebarTitle: "Configurations"
title: Configurations
seo:
title: "Vector Store Configuration Reference - Mem0"
description: "Reference for vector database configuration options in Mem0, including provider selection and connection settings."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "LangChain as Vector Store Provider"
sidebarTitle: "LangChain"
title: LangChain
seo:
title: "LangChain as Vector Store Provider - Mem0"
description: "Use LangChain as a unified vector store provider in Mem0 to access multiple vector databases through one interface."
---
@@ -102,7 +102,7 @@ Here are the parameters available for configuring Upstash Vector:
| `url` | URL for the Upstash Vector index | `None` |
| `token` | Token for the Upstash Vector index | `None` |
| `client` | An `upstash_vector.Index` instance | `None` |
| `collection_name` | The default namespace used | `"mem0"` |
| `collection_name` | The default namespace used | `""` |
| `enable_embeddings` | Whether to use Upstash embeddings | `False` |
<Note>
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Vector Store Providers Overview"
sidebarTitle: "Overview"
title: Overview
seo:
title: "Vector Store Providers Overview - Mem0"
description: "Overview of all supported vector databases in Mem0, including Qdrant, Chroma, PGVector, Pinecone, Oracle, and more."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Cookbooks and Tutorials"
sidebarTitle: "Overview"
title: Overview
seo:
title: "Cookbooks and Tutorials - Mem0"
description: "Browse cookbook examples and tutorials for building AI applications with Mem0, from companion chatbots to AI agents."
---
@@ -1,6 +1,7 @@
---
title: "Delete Memory Operation"
sidebarTitle: "Delete Memory"
title: Delete Memory
seo:
title: "Delete Memory Operation - Mem0"
description: Remove memories from Mem0 either individually, in bulk, or via filters.
icon: "trash"
iconType: "solid"
@@ -1,6 +1,7 @@
---
title: "Update Memory Operation"
sidebarTitle: "Update Memory"
title: Update Memory
seo:
title: "Update Memory Operation - Mem0"
description: Modify an existing memory by updating its content or metadata.
icon: "pen-to-square"
iconType: "solid"
-1
View File
@@ -41,7 +41,6 @@
"pages": [
"platform/quickstart",
"platform/overview",
"platform/copilot",
"platform/agent-signup",
"vibecoding",
"platform/cli",
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Integrations Overview"
sidebarTitle: "Overview"
title: Overview
seo:
title: "Integrations Overview - Mem0"
description: "Overview of Mem0 integrations with popular AI frameworks and tools for persistent memory and context management."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "AWS Bedrock Integration"
sidebarTitle: "AWS Bedrock"
title: AWS Bedrock
seo:
title: "AWS Bedrock Integration with Mem0"
description: "Use Mem0 with AWS Bedrock and OpenSearch Service for cloud-native persistent semantic memory storage."
---
+3 -8
View File
@@ -183,16 +183,11 @@ the same way the Python SDK and the CLI attribute theirs. Without a key they
are sent under a random per-machine id.
Each event carries the event name, the plugin version, the harness it ran in,
your OS and Python version, and per-event properties describing what happened:
timings, counts, coarse outcome and failure labels, and which model was
configured. Repository and session identifiers are hashed with a random salt
generated on your machine, so they cannot be linked back to a repository name
your OS and Python version, timings, counts, and a coarse failure label.
Repository and session identifiers are hashed with a random salt generated on
your machine and never sent, so they cannot be linked back to a repository name
or path.
The exact set is enforced in code rather than by this list: every property is
filtered through a denylist of sensitive keys and credential-shaped values are
redacted before anything is sent.
Prompts, memory text, queries, file paths, repository names, and API keys are
never sent.
+1 -1
View File
@@ -3,7 +3,7 @@ title: Codex
description: "Add persistent memory to OpenAI Codex with automatic capture, automatic recall, a search tool, and six memory skills."
---
Add persistent memory to [**OpenAI Codex**](https://openai.com/codex/) with the Mem0 plugin. Codex forgets everything between tasks. This plugin fixes that by connecting to Mem0's cloud memory layer via MCP, automatically capturing learnings at key lifecycle points, and retrieving relevant context on the first prompt of a session. Codex can use the search tool for recall later in the session.
Add persistent memory to [**OpenAI Codex**](https://openai.com/index/codex/) with the Mem0 plugin. Codex forgets everything between tasks. This plugin fixes that by connecting to Mem0's cloud memory layer via MCP, automatically capturing learnings at key lifecycle points, and retrieving relevant context on the first prompt of a session. Codex can use the search tool for recall later in the session.
<Info>Current plugin version: `0.3.1`.</Info>
+1 -1
View File
@@ -114,7 +114,7 @@ Both tools also accept optional per-call `userId`, `agentId`, and `runId` params
## Telemetry
Writes are tagged `source="DEEPSEEK_HARNESS"` so Mem0 can attribute usage to this integration. Usage events include operation names, durations, result counts, and coarse failure kinds. They are **not anonymous**: when an API key is configured they are sent under your Mem0 account email, the same way the SDK attributes its own. Queries, memory text, entity IDs, and API keys are never included. Set `MEM0_TELEMETRY=false` to opt out.
Writes are tagged `source="DEEPSEEK_HARNESS"` so Mem0 can attribute usage to this integration. Anonymous usage events include operation names, durations, result counts, and coarse failure kinds. Queries, memory text, entity IDs, and API keys are never included. Set `MEM0_TELEMETRY=false` to opt out.
<Note>
This plugin is a developer preview and tracks the evolving DeepSeek Harness plugin API.
+1 -1
View File
@@ -23,7 +23,7 @@ npm install -g flowise
npx flowise start
```
2. Access to the Flowise UI at `http://localhost:3000`
2. Access to the Flowise UI at http://localhost:3000
3. Basic familiarity with [Flowise's LLM orchestration](https://flowiseai.com/#features) concepts
## Setup and Configuration
+26 -61
View File
@@ -1,35 +1,39 @@
---
title: Hermes Agent
description: "Add long-term memory to Hermes agents using Mem0 Platform, a self-hosted server, or local OSS mode with background fact extraction."
description: "Add long-term memory to Hermes agents with Mem0, on managed Mem0 Cloud or fully self-hosted (OSS), with automatic background sync and zero-latency prefetch."
---
Add long-term memory to [Hermes Agent](https://github.com/NousResearch/hermes-agent), a self-improving AI agent CLI by Nous Research. Hermes has a pluggable memory system, and Mem0 is one of the supported providers. Once enabled, Mem0 learns facts from your conversations and surfaces relevant ones for the current question, without slowing down the chat.
Add long-term memory to [Hermes Agent](https://github.com/NousResearch/hermes-agent), a self-improving AI agent CLI by Nous Research. Hermes has a pluggable memory system, and Mem0 is one of the supported providers. Once enabled, Mem0 learns facts from your conversations and surfaces relevant ones before each turn, without slowing down the chat.
You can run Mem0 in three ways:
You can run Mem0 in two ways:
- **Platform mode** (default): managed Mem0 Cloud. Add your API key and you are ready.
- **Self-hosted server mode**: point the plugin at a Mem0 server you run yourself (the Docker-shipped server). The plugin only talks HTTP to your server.
- **OSS mode**: run Mem0 in-process with your own LLM, embedder, and vector store. No Mem0 server required.
- **OSS mode**: fully self-hosted with your own LLM, embedder, and vector store. No data leaves your machine.
## How It Works
Hermes runs a built-in memory system (file-based `MEMORY.md` and `USER.md`) alongside one external provider. When Mem0 is active, it works additively with the built-in system at two points in every conversation turn.
Hermes runs a built-in memory system (file-based `MEMORY.md` and `USER.md`) alongside one external provider. When Mem0 is active, it works additively with the built-in system at three points in every conversation turn.
### 1. Current-turn recall (bounded wait)
### 1. Before the agent responds (prefetch)
When you send a message, Hermes searches your stored memories for the current question and waits up to 3 seconds for results. If they arrive in time, they are injected into the system prompt so the model can see them. If the backend is slower, Hermes skips the injection and the model can still call `mem0_search` itself — so a slow backend never blocks a turn.
When you send a message, Hermes checks for cached Mem0 search results from the previous turn. If they exist, those memories are injected into the system prompt so the model can see them. This is zero-latency, with no waiting on an API call.
### 2. Background fact extraction (sync)
### 2. After the agent responds (sync)
Once the model finishes, Hermes sends the `(user message, assistant response)` pair to Mem0 in a background thread. Mem0 extracts facts automatically (for example, "user prefers Python" or "user works at Acme Corp"), so you never have to tell it what to remember. Each write is tagged with the gateway channel it came from.
### 3. Background prefetch for the next turn
At the same time, Hermes runs a background search to pre-load relevant memories for your next message. By the time you type, the results are already cached.
## Agent Tools
When Mem0 is active, the model gets four tools it can call during a conversation:
When Mem0 is active, the model gets five tools it can call during a conversation:
| Tool | Description | Parameters |
|------|-------------|------------|
| `mem0_search` | Semantic search by meaning, ranked by relevance | `query` (required), `top_k` (default 10, max 50), `rerank` (default `false`, Platform mode only) |
| `mem0_list` | List all stored memories, for a full overview | `page`, `page_size` (default 100, max 200) |
| `mem0_search` | Semantic search by meaning, ranked by relevance | `query` (required), `top_k` (default 10, max 50), `rerank` (default `true`, Platform mode only) |
| `mem0_add` | Store a fact verbatim, with no LLM extraction | `content` (required) |
| `mem0_update` | Update a memory's text by ID | `memory_id`, `text` (both required) |
| `mem0_delete` | Delete a memory by ID | `memory_id` (required) |
@@ -75,36 +79,6 @@ memory:
That's it. Mem0 runs automatically from here.
## Self-Hosted Server Setup
Run the [Mem0 server](https://github.com/mem0ai/mem0/tree/main/server) (FastAPI + pgvector) from its Docker image and point the plugin at it. Unlike OSS mode, the plugin just talks HTTP to your server.
### Interactive
```bash
hermes memory setup
# Select "mem0", then "Self-hosted server", and enter the server URL
```
### With flags
```bash
hermes memory setup mem0 --mode selfhosted \
--host http://localhost:8888 \
--api-key your-admin-api-key
```
### With environment variables
```bash
echo "MEM0_HOST=http://localhost:8888" >> ~/.hermes/.env
echo "MEM0_API_KEY=your-admin-api-key" >> ~/.hermes/.env
```
Then start a fresh Hermes session and call `mem0_search` — it connects to your server. The plugin authenticates with `X-API-Key` and uses the server's `/search` and `/memories` routes. The API key is optional only for servers running with `AUTH_DISABLED`.
<Note>Setting `host` routes to the self-hosted server automatically. Don't combine it with `mode: oss` — OSS takes precedence and ignores `host`.</Note>
## OSS (Self-Hosted) Setup
OSS mode runs Mem0 entirely on your own infrastructure: your LLM, your embedder, and your vector store. No data is sent to Mem0 Cloud, and no Mem0 API key is required.
@@ -137,20 +111,15 @@ hermes memory setup mem0 --mode oss \
| Flag | Description |
|------|-------------|
| `--mode` | `platform`, `selfhosted`, or `oss` |
| `--api-key` | Platform API key, or the admin key of a self-hosted server |
| `--host` | Self-hosted server URL (with `--mode selfhosted`) |
| `--mode` | `platform` or `oss` |
| `--oss-llm` | LLM provider (`openai` or `ollama`, default `openai`) |
| `--oss-llm-key` | LLM API key (for `openai`) |
| `--oss-llm-model` | Override the LLM model |
| `--oss-llm-url` | LLM base URL (for `ollama` or a custom endpoint) |
| `--oss-embedder` | Embedder provider (default `openai`) |
| `--oss-embedder-key` | Embedder API key |
| `--oss-embedder-model` | Override the embedder model |
| `--oss-embedder-url` | Embedder base URL (for `ollama` or a custom endpoint) |
| `--oss-vector` | Vector store (`qdrant` or `pgvector`, default `qdrant`) |
| `--oss-vector-path` | Local Qdrant storage path |
| `--oss-vector-url` | Qdrant server URL |
| `--oss-vector-host`, `--oss-vector-port` | PGVector or remote Qdrant host and port |
| `--oss-vector-user`, `--oss-vector-password`, `--oss-vector-dbname` | PGVector connection details |
| `--user-id` | Canonical user identifier |
@@ -158,7 +127,7 @@ hermes memory setup mem0 --mode oss \
## Switching Modes
You can move between the three modes at any time. Run the setup command again, or edit `~/.hermes/mem0.json` directly.
You can move between Platform and OSS at any time. Run the setup command again, or edit `~/.hermes/mem0.json` directly.
```bash
# Platform to OSS
@@ -167,9 +136,6 @@ hermes memory setup mem0 --mode oss --oss-llm-key sk-...
# OSS to Platform
hermes memory setup mem0 --mode platform --api-key sk-...
# Platform to a self-hosted server
hermes memory setup mem0 --mode selfhosted --host http://localhost:8888
# Preview without writing anything
hermes memory setup mem0 --mode oss --oss-llm-key sk-... --dry-run
```
@@ -180,7 +146,7 @@ A self-hosted `~/.hermes/mem0.json` looks like this:
{
"mode": "oss",
"oss": {
"llm": {"provider": "openai", "config": {"model": "gpt-5-mini", "is_reasoning_model": true}},
"llm": {"provider": "openai", "config": {"model": "gpt-5-mini"}},
"embedder": {"provider": "openai", "config": {"model": "text-embedding-3-small"}},
"vector_store": {"provider": "qdrant", "config": {"path": "~/.hermes/mem0_qdrant"}}
}
@@ -193,12 +159,11 @@ Behavioral settings live in `~/.hermes/mem0.json` and are written for you by `he
| Key | Default | Description |
|-----|---------|-------------|
| `mode` | `platform` | `platform` (Mem0 Cloud) or `oss` (self-managed, in-process). Self-hosted server routing is set via `host` |
| `host` | none | Self-hosted Mem0 server URL. When set, the plugin talks HTTP to your server instead of the cloud |
| `api_key` | none | Mem0 Platform API key, or the admin key of a self-hosted server. Stored in `.env` as `MEM0_API_KEY` |
| `mode` | `platform` | `platform` (Mem0 Cloud) or `oss` (self-hosted) |
| `api_key` | none | Mem0 Platform API key, required in Platform mode. Stored in `.env` as `MEM0_API_KEY` |
| `user_id` | `hermes-user` | Identifier that scopes memories. See cross-channel behavior below |
| `agent_id` | `hermes` | Agent identifier attached to writes |
| `rerank` | `false` | Rerank search results for relevance (Platform mode only) |
| `rerank` | `true` | Rerank search results for relevance (Platform mode only) |
### Cross-channel memories
@@ -209,11 +174,12 @@ Hermes can run from the CLI and from gateways like Telegram, Slack, and Discord.
Either way, every write is tagged with `metadata.channel` (for example `telegram` or `cli`), so per-channel views are still possible at query time.
## Reliability
- **Circuit breaker**: if Mem0 fails five times in a row, Hermes pauses calls for two minutes, then retries. The agent keeps working without memory during that window. Expected client errors, like a 404 on a missing memory id, do not count toward tripping the breaker.
- **Non-blocking**: fact extraction runs in a background daemon thread, and current-turn recall waits at most 3 seconds, so a slow or failed call never blocks your conversation.
- **Thread-safe**: the client uses lazy initialization with locking, and the background sync and recall threads are guarded so concurrent gateway messages cannot produce duplicate memories.
- **Non-blocking**: every Mem0 call runs in a background daemon thread, so a slow or failed call never blocks your conversation.
- **Thread-safe**: the client uses lazy initialization with locking, and the background sync and prefetch threads are guarded so concurrent gateway messages cannot produce duplicate memories.
## Troubleshooting
@@ -222,7 +188,6 @@ Either way, every write is tagged with `metadata.channel` (for example `telegram
The circuit breaker tripped after five consecutive failures and resets after two minutes.
- **Platform mode**: check your API key and internet connection.
- **Self-hosted server mode**: check that the server is running and reachable at the configured `host` URL.
- **OSS mode**: make sure your vector store (Qdrant or PGVector) is running and reachable.
### OSS: vector store connection refused
@@ -252,8 +217,8 @@ curl http://localhost:11434/api/tags
## Key Features
1. **Three ways to run**: managed Platform, a self-hosted server, or fully local OSS, switchable at any time.
2. **Current-turn recall**: memories for the current question are injected within a 3-second window, with `mem0_search` as the model's own backstop.
1. **Two ways to run**: managed Platform or fully self-hosted OSS, switchable at any time.
2. **Zero-latency recall**: memories are prefetched in the background and cached before you type.
3. **Automatic extraction**: Mem0 extracts and deduplicates facts from each exchange for you.
4. **Non-blocking and fault tolerant**: background threads plus a circuit breaker keep the agent responsive even when Mem0 is unreachable.
5. **Additive memory**: works alongside Hermes' built-in file memory (`MEMORY.md`, `USER.md`).
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "LangChain Integration"
sidebarTitle: "Langchain"
title: Langchain
seo:
title: "LangChain Integration with Mem0"
description: "Build personalized AI agents using LangChain for conversation flow and Mem0 for long-term memory retention."
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "n8n Integration"
sidebarTitle: "n8n"
title: n8n
seo:
title: "n8n Integration with Mem0"
description: "Add long-term memory to n8n workflows and AI Agents with the Mem0 community node, no code required."
---
+1 -3
View File
@@ -481,9 +481,7 @@ Plugin config is stored in `~/.openclaw/openclaw.json` with file permissions `0o
### Telemetry
Usage telemetry (PostHog) is enabled by default to help improve the plugin. No conversation content or memory values are included, only event counts (recall, capture, tool usage, CLI commands).
These events are **not anonymous**. OpenClaw does not send your account email the way the SDK does, but it does send an unsalted SHA-256 hash of it, falling back to a hash of the API key and then to a random per-machine id. Mem0 holds the email the hash is derived from, so the hash identifies your account rather than concealing it. The first run that resolves an account also emits a PostHog `$identify`, which permanently merges any earlier random id into that identity.
Anonymous usage telemetry (PostHog) is enabled by default to help improve the plugin. No conversation content or memory values are included, only event counts (recall, capture, tool usage, CLI commands).
To opt out, set the environment variable:
-1
View File
@@ -171,7 +171,6 @@ If the user is on a pre-current major (Python < 2, TS < 3, or a Platform call st
- [Introduction](https://docs.mem0.ai/introduction) [Both]: Use when the user wants a one-page overview of how memory fits between the LLM and the app.
- [Vibe Code with Mem0](https://docs.mem0.ai/vibecoding) [Both]: Use when the user is in Claude Code, Cursor, or Windsurf and wants memory wired into their editor.
- [Platform Overview](https://docs.mem0.ai/platform/overview) [Platform]: Use when the user picks the managed product - 4-line integration, hosted API, dashboard.
- [Mem0 Copilot](https://docs.mem0.ai/platform/copilot) [Platform]: Use when inspecting project memories, reviewing configuration changes, or testing extraction in the dashboard.
- [Sign up as an agent](https://docs.mem0.ai/platform/agent-signup) [Platform]: Use when an AI agent needs to mint a Mem0 API key autonomously - four commands, no email or dashboard, human claims ownership later.
- [Platform vs Open Source](https://docs.mem0.ai/platform/platform-vs-oss) [Both]: Use when the user is deciding between managed and self-hosted.
- [Platform Quickstart](https://docs.mem0.ai/platform/quickstart) [Platform]: Use for the first Platform integration - API key plus `MemoryClient.add/search`.
@@ -1,6 +1,7 @@
---
title: "Open Source Custom Instructions"
sidebarTitle: "Custom Instructions"
title: Custom Instructions
seo:
title: "Open Source Custom Instructions - Mem0"
description: Tailor fact extraction so Mem0 stores only the details you care about.
icon: "wand-magic-sparkles"
---
@@ -1,6 +1,7 @@
---
title: "Open Source Multimodal Support"
sidebarTitle: "Multimodal Support"
title: Multimodal Support
seo:
title: "Open Source Multimodal Support - Mem0"
description: Capture and recall memories from both text and images.
icon: "image"
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Open Source Features Overview"
sidebarTitle: "Overview"
title: "Overview"
seo:
title: "Open Source Features Overview - Mem0"
description: "Self-hosting features that extend Mem0 beyond basic memory storage"
icon: "list"
---
+1 -1
View File
@@ -56,7 +56,7 @@ The Mem0 REST API server exposes every OSS memory operation over HTTP. Run it al
make bootstrap # starts Compose, creates an admin, issues the first API key
```
Or to start the stack only and finish setup via the browser wizard at `http://localhost:3000`:
Or to start the stack only and finish setup via the browser wizard at http://localhost:3000:
```bash
cd server
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Open Source Overview"
sidebarTitle: "Overview"
title: "Overview"
seo:
title: "Mem0 Open Source Overview"
description: "Self-host Mem0 with full control over your infrastructure and data"
icon: "house"
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "CLI for Terminal Memory Management"
sidebarTitle: "CLI"
title: CLI
seo:
title: "Mem0 CLI for Terminal Memory Management"
description: "Manage memories from your terminal, for both humans and AI agents."
icon: "terminal"
iconType: "solid"
-194
View File
@@ -1,194 +0,0 @@
---
title: "Mem0 Copilot"
description: "Inspect project memories, review configuration changes, and test extraction from the Mem0 dashboard."
icon: "message"
---
Copilot is an AI assistant in the [Mem0 dashboard](https://app.mem0.ai/dashboard/copilot). Ask it to inspect stored memories, suggest extraction rules and categories, change project settings, or explain the platform SDKs and APIs.
You describe the task in chat. Copilot uses tools to read project data or propose changes. In **Review changes** mode, you approve each change before it runs.
## Start with your project
1. Sign in to the dashboard and select the organization and project you want to work on.
2. Open **Copilot** in the sidebar.
3. Select **Review changes** below the message box.
4. Ask: “Show my current project settings and explain what each one does.”
5. Expand an activity card, such as **Read project config**, to see its **Input** and **Result**. Check these results when reviewing an answer or confirming a change.
You can ask about settings or SDK usage with an empty project. Suggestions based on stored data need enough project history to analyze.
## Inspect stored memories
Start with a project overview, then narrow the question:
| Prompt | What to inspect |
| --- | --- |
| “Summarize what this project has stored and how it is categorized.” | The memories and category patterns Copilot reads. Check which records support its summary. |
| “Analyze the memories for user alice.” | Facts stored for `alice`, their categories, and unwanted information. To investigate a missing fact, also provide the original input. |
| “Search alice's memories for dietary preferences.” | Memories relevant to the query within that user's scope. |
Copilot can list memories across the project. Searching by meaning needs a specific User, Agent, App, or Run ID. Select one in **Scope** or name it in your prompt.
Memory lists are paginated, and suggestions use samples. Check the activity results to see which records were read. Ask for more pages when you need to inspect the rest.
### Choose the memory scope
Open **Scope** below the message box. Select an existing ID or type one and choose it.
| Control | Identifies |
| --- | --- |
| **User** | The end user whose memories you want to inspect, such as `alice`. |
| **Agent** | An AI agent associated with the memories. |
| **App** | An application associated with the memories. |
| **Run** | A particular execution or session associated with the memories. |
A field set to **any** adds no filter for that entity type. **Clear scope** removes the selections. Scope gives Copilot default IDs to use; you can request a different entity in a message. Check the activity's **Input** to confirm which IDs it used. Adding a test memory requires a **User** ID, even when another entity is selected.
<Note>
Scope selects memories, not separate settings. Categories, extraction instructions, memory depth, and multilingual settings apply to the selected project. Category and extraction suggestions also analyze project data, regardless of the selected entity.
</Note>
See [Entity-scoped memory](/platform/features/entity-scoped-memory) for how these IDs organize memories in your application.
## Review and apply changes
The mode control below the message box determines when changes run:
- **Review changes**: Copilot pauses before changing project configuration or adding a test memory. Read the proposal and choose whether to apply it.
- **Auto-apply**: Copilot can change settings and add test memories without asking for approval. Check the selected project and scope before using it.
Ask Copilot to show the current settings before requesting an update. In the approval card, review every listed field. Approving the card applies the whole proposal, including fields you did not edit.
| Action | Effect |
| --- | --- |
| **Apply change** | Approve the proposed write. Inspect the subsequent result to confirm it succeeded. |
| **Edit** | When offered, edit extraction instructions or category names and descriptions. Choose **Apply with edits** to submit your revision. |
| **Discard edits** | Return to the original proposal without applying it. |
| **Decline** | Reject this proposal without applying it. |
| **Ask for something else** | Enter feedback and choose **Send**. This declines the current proposal and sends your feedback as a new message. Review the next proposal before applying it. |
Not every field has an inline editor. If a proposal includes both extraction instructions and categories, only the instructions get an editor. The editor cannot apply blank instructions or an empty category list. Use **Ask for something else** to change other fields or request separate proposals.
<Warning>
A generated prompt profile contains separate prompts for extraction and summaries. If your project uses one, saving extraction instructions through Copilot deactivates it. Review the replacement rules and test the affected behavior before using them in production.
</Warning>
## What Copilot cannot do
Copilot cannot perform these actions, even in **Auto-apply** mode:
- Create or delete API keys.
- Directly edit or delete existing memories.
- Export project data.
- Invite or remove organization or project members, or change their roles and permissions.
- Create or delete organizations or projects.
- Change billing details or subscription plans.
Use the dashboard or the relevant platform API for these tasks. Copilot can explain the steps or point you to documentation, but it cannot carry out the actions.
Copilot supports the managed Mem0 platform. It does not help configure or use the [self-hosted open-source library](/open-source/overview).
## Example 1: Stop storing small talk, then test extraction
This walkthrough uses a project where you have noticed greetings or small talk in stored memories. Use a development project when trying configuration changes, since a test user does not isolate project settings.
<Steps>
<Step title="Inspect the problem">
Select the project, set **User** to `alice`, and ask:
> Analyze the memories for user alice.
Inspect the returned memories. Identify examples of small talk you want to exclude and useful facts you still want to keep.
</Step>
<Step title="Request a targeted change">
With **Review changes** selected, ask:
> Update my extraction instructions to ignore greetings and small talk. Keep the existing rules for durable user preferences.
Check for a **Read project config** activity before reviewing the update. If it is missing, ask Copilot to read the current instructions first. If analysis reports insufficient data, ask it to use the rule you provided. You can also set the rule directly through [Custom instructions](/platform/features/custom-instructions).
</Step>
<Step title="Review the proposal">
Review the full replacement text. Keep existing rules your application needs. For this test, the relevant rules could look like:
```text
Remember the following:
* The user's dietary preferences and food restrictions.
Don't remember the following:
* Greetings and small talk.
```
Use **Edit** to adjust the wording, or **Ask for something else** to request a revision.
Choose **Apply change** or **Apply with edits**, then inspect the result to confirm the instructions were saved.
</Step>
<Step title="Add a test conversation">
Choose a new **User** ID for this test, such as `copilot-extraction-test-01`, and clear any other entity selections. Ask:
> Add this as a user message for copilot-extraction-test-01: “Hi! How's your day? I prefer vegetarian meals and avoid peanuts.”
Review the test-memory proposal and choose **Apply change**.
<Warning>
Adding a test memory writes to the selected project. It is not a dry run. Use a dedicated test user so you can find and remove the test data afterward.
</Warning>
</Step>
<Step title="Wait for extraction, then check the result">
Expand **Add test memory** and inspect **Result**. A `PENDING` status means extraction is still running. Use the returned `event_id` with the [Get Event API](/api-reference/events/get-event) to check progress. Wait for `SUCCEEDED` before evaluating the output. If the event is `FAILED`, inspect its details before retrying.
Then ask:
> List all memories for user copilot-extraction-test-01. Then search that user's memories for dietary preferences.
Check that the memories retain the vegetarian preference and peanut restriction, and exclude the greeting. Inspect the full list as well as search results: a relevant search can hide unwanted small-talk records.
If the output is wrong, describe the mismatch and review another instruction change. Use a new test user for the next attempt so earlier memories do not affect the result. Remove test data afterward through the dashboard or [memory deletion API](/api-reference/memory/delete-memories).
</Step>
</Steps>
Test with new inputs after changing settings. Updating configuration is not a cleanup operation for existing memories. See [Custom instructions](/platform/features/custom-instructions) for guidance on writing extraction rules.
## Example 2: Suggest categories and extraction instructions
These workflows sample the selected project's data. Generating a suggestion does not update settings by itself. Copilot can then propose applying it, which follows the selected review mode. Use **Review changes** to inspect suggestions before they are saved.
| Workflow | Data used | Example prompt |
| --- | --- | --- |
| Custom categories | Stored memories, excluding deleted memories. | “Suggest custom categories from this project's memories.” |
| Extraction instructions | Successful requests to add conversation data, including requests that produced no memories. | “Compare recent add requests with the extracted memories and suggest better extraction instructions.” |
Category suggestions identify recurring themes. Review the names, descriptions, and examples. Applying a category list replaces the project's previous list; it does not add to it or re-tag existing memories. See [Custom categories](/platform/features/custom-categories) for project and per-call behavior.
Extraction suggestions compare conversation inputs with what was extracted. One add request can produce several memories or none, so the number of add requests is different from the number of stored memories. Keep your application's existing requirements when reviewing the proposal.
Both workflows require enough project data. If a suggestion fails because there is too little data, check the activity's **Result** for the current and required counts. You can still inspect memories, ask SDK/API questions, or provide your own rules instead of requesting a data-based suggestion.
## Example 3: Adjust detail and language
Ask Copilot to read the current settings, then request the change you need:
- **Memory depth** controls the level of detail: **Less**, **Medium**, or **More detailed**. Try “Show my current memory depth, then propose More detailed memories.” For deciding *which facts* to retain, use extraction instructions.
- **Multilingual** behavior preserves the user's original language. Try “Enable multilingual behavior so memories preserve the language of the input.” Test it with a new conversation in the language your application uses.
Both are project settings. Review the proposed values and test a fresh input after applying them. See [Organization and project settings](/api-reference/organizations-projects) for configuration through the API.
## Example 4: Ask SDK and API questions
Include your language and the task, for example:
> Show me how to search memories for user alice with the managed Mem0 Python SDK. Link the documentation you used.
Copilot can look up the official platform documentation. Check the **Browse mem0 docs** and **Read docs page** activities and open the cited pages. If an answer has no sources, ask for them before using the example. The [Platform quickstart](/platform/quickstart) covers adding and searching memories in code.
## Resume a conversation and check message limits
Use **History** to reopen your 50 most recently active chats for the selected project. Chats belong to the person who created them; other project members cannot open them. Use **New chat** to start a separate conversation, or **Delete chat** in history to remove one. Deleting a chat does not undo settings changes or remove test memories.
Message allowances depend on the organization's plan and are shared across its projects and members. They reset at the start of each calendar month in UTC.
For plans with a limit, a usage notice appears once 80% of the allowance is used. Below that point, no counter is shown. At the limit, new messages are disabled until the allowance resets or the plan is upgraded. Follow the notice's upgrade link to review plan options.
Sending feedback through **Ask for something else** counts as a new message. Approving or declining a saved proposal without feedback does not use another message. Test additions and searches also use the platform APIs and remain subject to their normal quotas.
@@ -1,6 +1,7 @@
---
title: "Platform Custom Instructions"
sidebarTitle: "Custom Instructions"
title: Custom Instructions
seo:
title: "Platform Custom Instructions - Mem0"
description: 'Control how Mem0 extracts and stores memories using natural language guidelines'
---
@@ -1,6 +1,7 @@
---
title: "Platform Multimodal Support"
sidebarTitle: "Multimodal Support"
title: Multimodal Support
seo:
title: "Platform Multimodal Support - Mem0"
description: Integrate images and documents into your interactions with Mem0
---
+3 -2
View File
@@ -1,6 +1,7 @@
---
title: "Platform Overview"
sidebarTitle: "Overview"
title: "Overview"
seo:
title: "Mem0 Platform Overview"
description: "Managed memory layer for AI agents, production-ready in minutes"
icon: "cloud"
---
+1 -1
View File
@@ -28,7 +28,7 @@
"clsx": "^2.1.1",
"js-cookie": "^3.0.6",
"lucide-react": "^0.477.0",
"next": "15.5.24",
"next": "15.5.21",
"react": "^19.0.0",
"react-dom": "^19.0.0",
"react-markdown": "^10.0.1",
+6 -25
View File
@@ -60,37 +60,18 @@ and the rules on them are what keep one layer from erasing another:
| `X-Application` | the host app it runs inside | **set-once** — write only if absent |
| `X-Mem0-Client` | `name/version`, outermost first | **append-only** — add yourself, never replace |
Set-once means check-then-set, never assignment. An integration that wraps the
SDK is the outermost layer and sets the source; the SDK underneath defers to it.
Assignment is exactly how every agent plugin came to be indistinguishable from
every other one at the platform.
Set-once means `setdefault`, never assignment. An integration that wraps the
SDK is the outermost layer and sets the source; the SDK underneath must defer to
it. Assignment is exactly how every agent plugin came to be indistinguishable
from every other one at the platform.
How to declare it from an integration, in order of preference:
1. Send the headers yourself, if you make the HTTP call directly.
2. Pass `source` in the call options, if you go through an SDK.
3. Set `MEM0_SOURCE` / `MEM0_APPLICATION` / `MEM0_CLIENT_STACK` in the
environment before constructing the client. The SDKs read these and defer to
anything already present.
Append-only applies where a stack can actually form: an SDK handed a client that
already carries `X-Mem0-Client` appends itself rather than replacing. An SDK
constructed with no outer context simply reports itself, which is correct — it
is the outermost layer in that process.
Append-only means a plugin calling the Python SDK produces
`mem0-plugin/0.3.1, mem0-python/2.0.19`, so neither layer can erase the other.
The backend recognizes a fixed list of source values and buckets everything else
into `OTHERS`. A new value has to land in the platform's `EventSource` enum, so
do not invent one without that change going in too.
`X-Application` is allowlisted the same way, and this one has a rule of its own:
**omit the header when you do not know the host.** A value outside the allowlist
is discarded server-side, so guessing produces an event that claims an
attribution we do not actually have. The portable bundle is the case that
matters. It runs in whatever editor a user drops it into, so its build leaves
`PLATFORM_APPLICATION` empty and `memory_core` sends no header at all, while the
native bundles each name the host they were generated for. If you add a build
target, decide which of those two it is.
## Adding an integration
1. For a native coding-agent host, add `integrations/<name>-plugin/` with `plugin-build.json`, its manifest, and a thin adapter, then generate its shared runtime. Portable clients use the single `mem0-agent-plugin/` package. Independent TypeScript integrations stay self-contained and import shared lifecycle behavior from `agent-plugin-core/typescript/`.
+3 -13
View File
@@ -81,23 +81,15 @@ def replace_output(staged: Path, output: Path) -> Path:
return output
def _render_harness_id(host: str, *, portable: bool = False) -> str:
def _render_harness_id(host: str) -> str:
"""Emit core/_harness_id.py for one host.
Carries both vocabularies from a single definition: the PostHog `source` tag
and the platform's X-Mem0-Source / X-Application pair. Keeping them together
is what stops the two from drifting into separate vocabularies for the same
thing.
The portable bundle runs in whatever editor a user drops it into, so it does
not know its host and must not guess one. HARNESS_ID stays "coding-agent",
which is true and useful for grouping in PostHog, but PLATFORM_APPLICATION is
left empty: X-Application names a real host app, is checked against an
allowlist server-side, and a value that is always discarded is worse than no
value -- it reads like an attribution we have and do not.
"""
tag = host.upper().replace("-", "_") + "_PLUGIN"
application = "" if portable else host
return (
'"""Generated by integrations/agent-plugin-core/build/build.py. Do not edit."""\n'
"\n"
@@ -106,10 +98,8 @@ def _render_harness_id(host: str, *, portable: bool = False) -> str:
"\n"
"# Platform-side vocabulary (mem0_event.source + X-Application). The whole\n"
"# plugin family is one source; which editor it runs in is the application.\n"
"# An empty application means the host is unknown, and memory_core omits\n"
"# the header entirely rather than sending a placeholder.\n"
'PLATFORM_SOURCE = "MEM0_PLUGIN"\n'
f'PLATFORM_APPLICATION = "{application}"\n'
f'PLATFORM_APPLICATION = "{host}"\n'
)
@@ -132,7 +122,7 @@ def _bundle_python(
# to call telemetry.init(). mcp_server.py and the detached telemetry.py sender
# never did, which is how MCP searches reported harness=generic and every
# batch they drained was labelled MEM0_PLUGIN regardless of the real host.
(core / "_harness_id.py").write_text(_render_harness_id(host, portable=portable), encoding="utf-8")
(core / "_harness_id.py").write_text(_render_harness_id(host), encoding="utf-8")
values = {
"PLUGIN_ROOT": plugin_root,
@@ -290,11 +290,6 @@ def run(
if args.plugin_data_dir:
os.environ[data_dir_env] = args.plugin_data_dir
# Snapshot BEFORE anything writes to the data dir: cache_plugin_api_key
# writes `api-key` and EvidenceStore creates `evidence.sqlite3`, so asking
# after them always saw content and every fresh install reported an upgrade.
data_dir_was_empty = telemetry.data_dir_was_empty()
cache_plugin_api_key()
if args.action == "session-start":
clear_stale_api_key_cache()
@@ -312,7 +307,7 @@ def run(
if args.action == "session-start":
# Claims the marker atomically and says which event to record, so a
# second session starting alongside this one cannot record it too.
first_event = telemetry.claim_install(was_empty=data_dir_was_empty)
first_event = telemetry.claim_install()
if first_event == "install":
telemetry.record("install")
elif first_event == "upgrade":
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.2"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -2008,12 +2008,9 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
# Top level, not metadata: the backend reads `source` from the body,
# query string or X-Mem0-Source header, never from metadata. The
# harness tag stays in metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
@@ -49,7 +49,6 @@ except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
_salt_cache: str = ""
_harness: str = _DEFAULT_HARNESS
_source_tag: str = _DEFAULT_SOURCE_TAG
_PRIVATE_KEYS = {
@@ -105,11 +104,6 @@ MAX_CLAIM_ATTEMPTS = 3
# Parked claims drained per run, after the live spool. Bounded so a long backlog
# cannot turn one flush into an unbounded send loop.
MAX_PARKED_PER_RUN = 3
# Added to the wait before a released claim becomes reclaimable, per attempt
# already spent. Releasing straight to "reclaimable now" let two senders burn the
# whole budget within seconds of one another on a single momentary failure, and
# discard a batch a retry a minute later would have delivered.
RETRY_COOLDOWN_SECONDS = 60
def is_enabled() -> bool:
@@ -127,98 +121,15 @@ def _digest(value: str, length: int = 16) -> str:
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
def _salt_path() -> Path:
return memory_core.data_dir() / "telemetry-salt"
def _install_salt() -> str:
"""Random per-install salt, created once and memoized for the process.
Deliberately its own file, claimed with O_CREAT|O_EXCL, rather than a key in
the identity file. Three reasons, all of which produced wrong data when this
lived in the identity dict:
- Hooks are short-lived separate processes firing on every tool call, and
people run more than one agent window. A read-modify-write would let each
process mint its own salt, so one repository would hash several ways in the
window before a writer won.
- resolve_distinct_id holds a copy of the identity dict across a network call
to /v1/ping/, so whichever write landed second erased the other's key —
losing either the salt (repo_hash changes mid-stream) or the email (a
second $identify, splitting the person).
- Touching the identity file from record() would create it, and is_first_run
keys off that file, so recording an event would silently suppress the
install event.
Published atomically, and there is deliberately no derived fallback. Creating
the file with O_CREAT|O_EXCL and then writing into it leaves a window where
the file exists and is empty, and a concurrent hook that reads it in that
window gets nothing. Falling back to a digest of the path would hand that
process a salt an attacker can compute, memoized for its whole run, which is
the privacy control this function exists to provide silently turning itself
off under load. The salt is written to a private temp file first and linked
into place, so the name either does not exist or already has the full value.
Returns "" when it genuinely cannot persist. Callers omit the hash entirely
rather than emit an unsalted one.
"""
global _salt_cache
if _salt_cache:
return _salt_cache
path = _salt_path()
# Read before writing. Hooks are separate processes firing on every tool
# call, so all but the first find the salt already published; going straight
# to create-fsync-link-unlink meant every one of them paid an fsync to
# discover that, on a path whose whole promise is appending a line and
# returning.
try:
_salt_cache = path.read_text(encoding="utf-8").strip()
if _salt_cache:
return _salt_cache
except OSError:
pass
temporary = path.with_name(f"{path.name}.{os.getpid()}.tmp")
try:
path.parent.mkdir(parents=True, exist_ok=True)
handle = os.open(temporary, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
with os.fdopen(handle, "w", encoding="utf-8") as stream:
stream.write(uuid.uuid4().hex)
stream.flush()
os.fsync(stream.fileno())
try:
# Atomic claim: fails if another process already published one.
# os.link rather than replace, which would clobber theirs.
os.link(temporary, path)
except FileExistsError:
pass
except OSError:
# No hardlinks here (some network mounts, some container volumes).
# Claim the name directly instead. That reopens the empty-file
# window, but the window is now benign: a reader that lands in it
# gets "" and omits the hash for that process rather than caching a
# guessable one. Losing the hashes on every run of an entire
# filesystem is the worse failure.
try:
fallback = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
with os.fdopen(fallback, "w", encoding="utf-8") as stream:
stream.write(temporary.read_text(encoding="utf-8"))
except OSError:
pass
except OSError:
pass
finally:
try:
temporary.unlink()
except OSError:
pass
try:
_salt_cache = path.read_text(encoding="utf-8").strip()
except OSError:
_salt_cache = ""
return _salt_cache
"""Random per-install salt, created on first use and kept in the identity file."""
identity = _read_identity()
salt = identity.get("salt")
if not salt:
salt = uuid.uuid4().hex
identity["salt"] = salt
_write_identity(identity)
return salt
def _scoped_digest(value: str, length: int = 16) -> str:
@@ -230,17 +141,10 @@ def _scoped_digest(value: str, length: int = 16) -> str:
privacy control without the salt. Salting per install keeps every
within-account join the analytics actually use and gives up only
cross-machine joins on the same repository, which nothing computes.
Returns "" when there is no salt, so record() omits the property. An
unsalted digest over this input space is close to plaintext, and emitting one
under a name that implies it is hashed is worse than sending nothing.
"""
if not value:
return ""
salt = _install_salt()
if not salt:
return ""
return hashlib.sha256(f"{salt}:{value}".encode("utf-8")).hexdigest()[:length]
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
def _safe_value(value: Any) -> Any:
@@ -302,25 +206,6 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
return created
def _rotate_anonymous_id(identity: dict[str, str]) -> str:
"""Mint a fresh anonymous id because the account context is gone.
The previous id may already have been merged into a person profile by an
$identify, and that merge is permanent. Reusing it after a logout or a key
change attributes everything that follows to the account that just went
away, which is the same misattribution the key fingerprint exists to stop,
only arriving through the anonymous path instead.
`aliased` is cleared with it: the new id has never been merged, so it is
eligible to be aliased into whatever account comes next.
"""
created = f"code-anon-{uuid.uuid4().hex}"
identity["anonymous_id"] = created
identity.pop("aliased", None)
_write_identity(identity)
return created
def _install_state_path() -> Path:
return memory_core.data_dir() / "install-state.json"
@@ -336,35 +221,18 @@ def is_first_run() -> bool:
return not _install_state_path().exists()
def data_dir_was_empty() -> bool:
"""Whether the data directory is untouched. Call BEFORE anything writes to it.
hook_runner reaches claim_install() only after cache_plugin_api_key() has
written `api-key` and EvidenceStore() has created `evidence.sqlite3`, so
asking at claim time always saw content and every fresh install reported an
upgrade. The caller snapshots this at the top of the run instead.
"""
return not _data_dir_has_content()
def claim_install(was_empty: bool | None = None) -> str | None:
def claim_install() -> str | None:
"""Claim the one install/upgrade record for this machine, atomically.
Returns the event to record ("install" or "upgrade"), or None if another
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
cannot both win.
`was_empty` must come from data_dir_was_empty() called before this process
wrote anything. Omitting it falls back to checking now, which is only
correct for a caller that has touched nothing.
"""
if not is_enabled():
# Never consume the one-shot claim while the user is opted out, or they
# would silently lose their install event if they later opt in.
return None
path = _install_state_path()
upgrading = not (data_dir_was_empty() if was_empty is None else was_empty)
# A fresh install has an empty data directory. Anything already there —
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
# before the marker is created, since creating it would itself be content.
upgrading = _data_dir_has_content()
try:
path.parent.mkdir(parents=True, exist_ok=True)
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
@@ -382,14 +250,6 @@ def claim_install(was_empty: bool | None = None) -> str | None:
},
stream,
)
# Durable before this returns. The O_EXCL open is what makes the
# claim exclusive, so it cannot be replaced by a temp-and-rename
# without losing that, which leaves the content as the thing to make
# safe. A kill between the open and this fsync used to leave a marker
# that exists but parses to nothing: is_first_run reads it as claimed
# and claim_version_change cannot read a version out of it.
stream.flush()
os.fsync(stream.fileno())
except OSError:
pass
return "upgrade" if upgrading else "install"
@@ -406,19 +266,6 @@ def _data_dir_has_content() -> bool:
return False
def _repair_install_state(path: Path) -> None:
"""Rewrite an unparseable marker so version tracking can resume."""
try:
temporary = path.with_suffix(f".{os.getpid()}.tmp")
temporary.write_text(
json.dumps({"plugin_version": memory_core.PLUGIN_VERSION, "repaired_at": memory_core.utc_now()}),
encoding="utf-8",
)
temporary.replace(path)
except OSError:
pass
def claim_version_change() -> str | None:
"""Return the previously recorded version if it differs, updating the marker.
@@ -429,47 +276,20 @@ def claim_version_change() -> str | None:
path = _install_state_path()
try:
state = json.loads(path.read_text(encoding="utf-8"))
except OSError:
except (OSError, json.JSONDecodeError):
return None
except json.JSONDecodeError:
# A crash between O_EXCL and the write leaves an empty marker. Left
# alone it disables every future upgrade event on this machine, because
# claim_install sees the file and this function cannot parse it.
state = None
if not isinstance(state, dict):
_repair_install_state(path)
return None
previous = str(state.get("plugin_version") or "")
if not previous or previous == memory_core.PLUGIN_VERSION:
return None
# Claim the transition with an exclusive sentinel before rewriting the
# marker. A plain read-modify-write let every concurrently starting session
# observe the old version and each record its own upgrade — and the first
# session after a version bump is exactly when several agent windows restart
# together.
sentinel = path.with_name(f"upgraded-{memory_core.PLUGIN_VERSION}")
try:
os.close(os.open(sentinel, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600))
except FileExistsError:
return None
except OSError:
return None
state["plugin_version"] = memory_core.PLUGIN_VERSION
state["upgraded_at"] = memory_core.utc_now()
temporary = path.with_suffix(f".{os.getpid()}.tmp")
try:
temporary = path.with_suffix(f".{os.getpid()}.tmp")
temporary.write_text(json.dumps(state), encoding="utf-8")
temporary.replace(path)
except OSError:
# Release the claim. The marker still records the old version, so
# without this the sentinel makes claim_version_change return early on
# every later run and this version's upgrade is never recorded again.
for leftover in (sentinel, temporary):
try:
leftover.unlink()
except OSError:
pass
return None
return previous
@@ -503,17 +323,10 @@ def record(
os=sys.platform,
python_version=platform.python_version(),
)
# Assigned only when the digest is real. _scoped_digest returns "" when
# the salt could not be persisted, and an empty property is worse than an
# absent one: it survives the None filter below and reads as a value.
if repo is not None:
repo_hash = _scoped_digest(getattr(repo, "identity", ""))
if repo_hash:
properties["repo_hash"] = repo_hash
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
if session_id:
session_hash = _scoped_digest(session_id)
if session_hash:
properties["session_hash"] = session_hash
properties["session_hash"] = _scoped_digest(session_id)
line = json.dumps(
{
"event": f"{EVENT_PREFIX}.{event}",
@@ -583,18 +396,9 @@ def _claim_name(attempt: int = 0) -> str:
def _claim_attempt(claim: Path) -> int:
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape.
Anchored on field position, not on a leading "a": the legacy shape is
``telemetry-<pid>-<hex>.sending`` and a hex id such as ``a1234567`` would
otherwise parse as attempt 1234567 and be discarded unsent on the first
flush after an upgrade.
"""
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
parts = stem.split("-")
if len(parts) != 4:
return 0
tail = parts[3]
tail = stem.rsplit("-", 1)[-1]
if tail.startswith("a") and tail[1:].isdigit():
return int(tail[1:])
return 0
@@ -628,42 +432,6 @@ def _claim_spool() -> Path | None:
return _claim_parked(directory)
def _sweep_debris(directory: Path) -> None:
"""Remove files nothing else will ever pick up again.
*.partial is a temp file orphaned by a crash between write and rename.
*.corrupt is a batch quarantined for undecodable content. No glob in this
module matches either, so without this they accumulate on disk for the life
of the install.
Quarantined batches are kept far longer than debris: they are the only
evidence left of events that could not be delivered, and someone diagnosing
a report of missing telemetry has to be able to find one.
"""
now = time.time()
for debris in directory.glob("telemetry-*.partial"):
try:
if now - debris.stat().st_mtime > CLAIM_STALE_SECONDS:
debris.unlink()
except OSError:
continue
for quarantined in directory.glob("telemetry-*.corrupt"):
try:
if now - quarantined.stat().st_mtime > CLAIM_EXPIRY_SECONDS:
quarantined.unlink()
except OSError:
continue
# The same reasoning covers *.tmp. _write_identity and _install_salt both
# create one and unlink it in a finally, which a SIGKILL skips, and no glob
# in this module matches the leftovers either.
for temporary in directory.glob("telemetry-*.tmp"):
try:
if now - temporary.stat().st_mtime > CLAIM_STALE_SECONDS:
temporary.unlink()
except OSError:
continue
def _claim_parked(directory: Path) -> Path | None:
"""Take the oldest abandoned claim, if any lease has actually expired.
@@ -679,24 +447,15 @@ def _claim_parked(directory: Path) -> Path | None:
age = now - orphan.stat().st_mtime
except OSError:
continue
if age < CLAIM_STALE_SECONDS:
# Someone else holds a live lease on it. This check has to come
# first. Claiming a file bumps its attempt count and refreshes its
# mtime, so a sender that has just taken the final attempt looks
# exhausted to everyone else while it is actively draining. Judging
# exhaustion before liveness let a second sender unlink a batch out
# from under its owner, losing every event in it.
continue
# Attempts, not age. Every re-claim touches the mtime and every release
# backdates it by a fixed amount, so age is pinned near the stale
# threshold and never reaches the expiry. Age stays only as a backstop
# for files that never carried an attempt marker.
if _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS or age > CLAIM_EXPIRY_SECONDS:
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
try:
orphan.unlink()
except OSError:
pass
continue
if age < CLAIM_STALE_SECONDS:
# Someone else holds a live lease on it.
continue
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
try:
orphan.replace(claim)
@@ -730,14 +489,10 @@ def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
return True
temporary = claim.with_suffix(f".{os.getpid()}.partial")
try:
payload = "".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining)
# fsync before the rename: without it the rename can land while the
# bytes have not, and the claim comes back empty or truncated after a
# crash. _drain then reads zero events and unlinks it.
with open(temporary, "w", encoding="utf-8") as handle:
handle.write(payload)
handle.flush()
os.fsync(handle.fileno())
temporary.write_text(
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
encoding="utf-8",
)
temporary.replace(claim)
_touch(claim)
return True
@@ -761,11 +516,7 @@ def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
if not _rewrite_claim(claim, remaining):
return
try:
# Backdate past the stale threshold so the next flush can pick it up,
# minus a cooldown that grows with the attempts already spent. Clamped so
# the mtime never lands in the future, which would read as a live lease.
cooldown = min(_claim_attempt(claim) * RETRY_COOLDOWN_SECONDS, CLAIM_STALE_SECONDS)
released = time.time() - CLAIM_STALE_SECONDS - 1 + cooldown
released = time.time() - CLAIM_STALE_SECONDS - 1
os.utime(claim, (released, released))
except OSError:
pass
@@ -812,56 +563,23 @@ def resolve_distinct_id() -> tuple[str, str]:
fingerprint = _digest(key) if key else ""
email = identity.get("email", "")
if email and fingerprint:
recorded = identity.get("key_fingerprint", "")
if recorded == fingerprint:
return email, ""
if not recorded:
# Rows written before fingerprints existed. Verify rather than
# adopt: a key changed before the upgrade would otherwise bind the
# new key to the previous account's email, permanently, and the
# fingerprint would then agree with itself forever after.
verified = _resolve_email(key)
if not verified:
# Offline, firewalled, or the API is down. Keep the previous
# behaviour and retry on the next flush rather than dropping a
# real account attribution. Safe because the same network that
# failed /v1/ping/ is about to fail the PostHog POST, so nothing
# is delivered under the unverified identity in the meantime.
return email, ""
identity["email"] = verified
identity["key_fingerprint"] = fingerprint
_write_identity(identity)
return verified, ""
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
return email, ""
if not key:
# No key to verify the account with; do not keep attributing to it.
if email:
identity.pop("email", None)
identity.pop("key_fingerprint", None)
return _rotate_anonymous_id(identity), ""
_write_identity(identity)
return anonymous_id(identity), ""
resolved = _resolve_email(key)
if not resolved:
# The key changed and will not resolve (revoked, offline, API down).
# Reaching here with an email means the recorded fingerprint disagreed,
# so the key really did change. Drop the account and rotate: the stored
# anonymous id may already be merged into that account's person, and
# reusing it would keep the events on the profile we are trying to
# leave.
if email:
identity.pop("email", None)
identity.pop("key_fingerprint", None)
return _rotate_anonymous_id(identity), ""
return anonymous_id(identity), ""
return (email, "") if email else (anonymous_id(identity), "")
# Alias only when going anonymous -> email for the first time. Once an anon
# id has been merged into an account it must never be offered again: an
# alias naming an already-identified id is what could link two real people.
previous = "" if (email or identity.get("aliased")) else identity.get("anonymous_id", "")
if previous:
identity["aliased"] = True
# Alias only when going anonymous -> email for the first time.
previous = "" if email else identity.get("anonymous_id", "")
identity["email"] = resolved
identity["key_fingerprint"] = fingerprint
_write_identity(identity)
@@ -881,7 +599,6 @@ def flush() -> int:
# Parked batches used to starve behind the live spool indefinitely. Bounded
# per run so a long backlog cannot turn one flush into an unbounded loop.
directory = memory_core.data_dir()
_sweep_debris(directory)
for _ in range(MAX_PARKED_PER_RUN):
parked = _claim_parked(directory)
if parked is None:
@@ -902,28 +619,8 @@ def _drain(claim: Path | None) -> tuple[int, bool]:
return 0, True
try:
lines = claim.read_text(encoding="utf-8").splitlines()
except ValueError:
# UnicodeDecodeError from a torn write: the content is unrecoverable, so
# quarantine rather than retry. flush() runs from a bare `finally:` in
# flush_worker, so raising here also skips the handoff cleanup, and an
# undecodable file would otherwise be re-read on every flush forever.
# Reported as delivered because there is nothing left to deliver and the
# rest of the run should continue.
try:
claim.replace(claim.with_suffix(".corrupt"))
except OSError:
try:
claim.unlink()
except OSError:
pass
return 0, True
except OSError:
# Could not read it, which is not the same as having nothing to send.
# The file is left exactly where it is: a vanished or briefly unreadable
# claim is retryable, and quarantining it here would discard events over
# a transient filesystem error. Reported as undelivered so the run stops
# instead of counting a batch nothing was posted from as delivered.
return 0, False
return 0, True
events = []
for line in lines:
try:
@@ -933,15 +630,8 @@ def _drain(claim: Path | None) -> tuple[int, bool]:
if isinstance(value, dict) and value.get("event"):
events.append(value)
if not events:
# Only delete when the file really is empty. A non-empty file that
# parses to nothing is a torn write, and its contents are the unsent
# remainder — deleting it is the data loss this PR exists to prevent.
try:
empty = claim.stat().st_size == 0
except OSError:
empty = True
try:
claim.replace(claim.with_suffix(".corrupt")) if not empty else claim.unlink()
claim.unlink()
except OSError:
pass
return 0, True
@@ -991,13 +681,8 @@ def _drain(claim: Path | None) -> tuple[int, bool]:
return sent, False
sent += len(chunk)
# Record progress and refresh the lease after each successful batch, so
# a crash repeats at most one batch instead of the entire file. If the
# rewrite fails the claim still holds delivered events, so stop rather
# than carry on as though progress were recorded — continuing is how the
# duplicate delivery this PR fixes would come back.
if not _rewrite_claim(claim, events[start + len(chunk) :]):
_release_claim(claim, events[start + len(chunk) :])
return sent, False
# a crash repeats at most one batch instead of the entire file.
_rewrite_claim(claim, events[start + len(chunk) :])
return sent, True
@@ -7,8 +7,8 @@ disable-model-invocation: true
# Pause memory capture
To pause (hooks stop capturing and sending session content; a minimal
telemetry ping still fires at session start, under your Mem0 account email,
unless `MEM0_TELEMETRY=false`):
anonymous telemetry ping still fires at session start unless
`MEM0_TELEMETRY=false`):
```bash
python3 "{{PLUGIN_ROOT}}/core/memory_cli.py" --harness "{{HARNESS_ID}}" {{PLUGIN_DATA_ARG}} pause
@@ -70,38 +70,6 @@ def test_portable_bundle_is_conformant_and_self_contained(tmp_path: Path) -> Non
assert not any(path.is_symlink() for path in root.rglob("*"))
def _harness_identity(root: Path) -> dict[str, str]:
"""Read the generated core/_harness_id.py without importing it."""
values: dict[str, str] = {}
for line in (root / "core" / "_harness_id.py").read_text(encoding="utf-8").splitlines():
if "=" in line and not line.lstrip().startswith("#"):
name, _, raw = line.partition("=")
values[name.strip()] = raw.strip().strip('"')
return values
def test_the_portable_bundle_declares_no_host_application(tmp_path: Path) -> None:
"""It runs in whatever editor a user drops it into, so it cannot know the host.
X-Application is allowlisted server-side. A guessed value is silently dropped
there, which is the worst outcome: the wire says we know the host and the
stored event says we do not.
"""
identity = _harness_identity(build("mem0-agent-plugin", "portable", tmp_path / "portable"))
assert identity["PLATFORM_APPLICATION"] == ""
# The PostHog-side label is still useful for grouping and stays populated.
assert identity["HARNESS_ID"] == "coding-agent"
assert identity["PLATFORM_SOURCE"] == "MEM0_PLUGIN"
@pytest.mark.parametrize("host", ["claude-code", "cursor", "codex", "kimi", "antigravity"])
def test_a_native_bundle_names_the_host_it_was_built_for(host: str, tmp_path: Path) -> None:
identity = _harness_identity(build(host, "native", tmp_path / host))
assert identity["PLATFORM_APPLICATION"] == host
@pytest.mark.parametrize("host", ["claude-code", "cursor", "codex", "kimi", "antigravity"])
def test_native_bundle_is_self_contained(host: str, tmp_path: Path) -> None:
root = build(host, "native", tmp_path / host)
@@ -25,32 +25,15 @@ pytestmark = pytest.mark.skipif(not HOST_CORE.exists(), reason="claude-code-plug
@pytest.fixture()
def telemetry(tmp_path, monkeypatch):
# CI runs this directory and claude-code-plugin/tests in ONE pytest process,
# and that suite's conftest sets MEM0_TELEMETRY=false at import, process-wide.
# Without this the whole file silently no-ops: record() returns early and
# every assertion sees an empty spool. Do not rely on ambient env.
monkeypatch.setenv("MEM0_TELEMETRY", "true")
monkeypatch.setenv("MEM0_CODE_DATA_DIR", str(tmp_path / "data"))
monkeypatch.syspath_prepend(str(HOST_CORE))
# Save and RESTORE rather than delete. claude-code-plugin/tests/conftest.py
# imports memory_core once at collection and calls configure_harness() on it;
# dropping the module left a later re-import with default harness config, so
# tests in that suite failed depending on collection order.
names = ("telemetry", "memory_core", "_harness_id")
saved = {name: sys.modules.get(name) for name in names}
for name in names:
for name in ("telemetry", "memory_core", "_harness_id"):
sys.modules.pop(name, None)
module = importlib.import_module("telemetry")
monkeypatch.setattr(module, "resolve_distinct_id", lambda: ("tester@example.com", ""))
try:
yield module
finally:
for name in names:
sys.modules.pop(name, None)
if saved[name] is not None:
sys.modules[name] = saved[name]
yield module
for name in ("telemetry", "memory_core", "_harness_id"):
sys.modules.pop(name, None)
def _delivered(payloads):
@@ -104,67 +87,6 @@ def test_a_fresh_claim_is_not_immediately_stealable(telemetry):
assert telemetry._claim_parked(first.parent) is None
def test_a_live_final_attempt_is_not_deleted_by_another_sender(telemetry):
"""Review finding: exhaustion was judged before liveness, so owners lost batches.
Claiming a parked file bumps its attempt count and refreshes its mtime. Once
the count reaches the budget, the owner draining it looked exhausted to every
other sender, which unlinked the file out from under it. Everything in that
batch was gone, which is precisely the loss this PR exists to stop.
"""
telemetry.record("search", reason="owned-by-the-first-sender")
spool = telemetry._spool_path()
stale = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
os.utime(spool, (stale, stale))
claim = telemetry._claim_spool()
assert claim is not None
# Walk it to the final attempt, ageing it each round so it can be re-claimed.
# _claim_spool hands back a0 and _release_claim keeps the name, so it takes
# one full round per attempt to reach the budget.
for _ in range(telemetry.MAX_CLAIM_ATTEMPTS):
# Carry the marker through each rewrite so the final assertion proves the
# events survived, not merely that some file with the right name did.
telemetry._release_claim(claim, [{"event": "code.search", "uuid": "owned-by-the-first-sender"}])
parked = sorted(claim.parent.glob("telemetry-*.sending"))
assert parked, "the batch was dropped while still inside its budget"
os.utime(parked[0], (stale, stale))
claim = telemetry._claim_parked(claim.parent)
assert claim is not None
assert telemetry._claim_attempt(claim) >= telemetry.MAX_CLAIM_ATTEMPTS
assert claim.exists()
# The owner is draining it right now: fresh mtime, live lease.
second_sender = telemetry._claim_parked(claim.parent)
assert second_sender is None, "a second sender took a batch under a live lease"
assert claim.exists(), "a second sender deleted a batch its owner was draining"
assert "owned-by-the-first-sender" in claim.read_text(encoding="utf-8")
def test_an_exhausted_batch_is_still_discarded_once_its_lease_lapses(telemetry):
"""The liveness check must defer the cleanup, not cancel it.
Guards the obvious over-correction: skipping live claims is only safe if an
abandoned one at the same attempt count is still reaped on a later run.
"""
telemetry.record("search")
spool = telemetry._spool_path()
stale = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
os.utime(spool, (stale, stale))
claim = telemetry._claim_spool()
assert claim is not None
exhausted = claim.parent / telemetry._claim_name(telemetry.MAX_CLAIM_ATTEMPTS)
claim.replace(exhausted)
os.utime(exhausted, (stale, stale))
assert telemetry._claim_parked(exhausted.parent) is None
assert not exhausted.exists(), "an abandoned exhausted batch was left behind forever"
def test_a_parked_batch_is_drained_behind_the_live_spool(telemetry):
"""Defect 6: parked claims were only reachable when no spool existed.
@@ -190,19 +112,16 @@ def test_a_parked_batch_is_drained_behind_the_live_spool(telemetry):
assert names == {"code.parked", "code.fresh"}
def test_a_batch_is_retried_until_the_budget_is_spent_not_discarded(telemetry):
"""Expiry discards what failed repeatedly, not what merely sat for a while.
The budget is the attempt count, because age cannot be one: every re-claim
touches the mtime and every release backdates it, so age never accumulates.
"""
def test_an_untried_batch_is_not_expired_by_age_alone(telemetry):
"""Expiry should discard what failed, not what never got a turn."""
telemetry.record("parked")
telemetry._post = lambda payload, url: False
telemetry.flush()
parked = list(telemetry.memory_core.data_dir().glob("telemetry-*.sending"))
assert len(parked) == 1
assert telemetry._claim_attempt(parked[0]) < telemetry.MAX_CLAIM_ATTEMPTS
ancient = time.time() - (telemetry.CLAIM_EXPIRY_SECONDS + 60)
os.utime(parked[0], (ancient, ancient))
sent: list[dict] = []
telemetry._post = lambda payload, url: sent.append(payload) or True
@@ -235,212 +154,11 @@ def test_progress_is_recorded_after_every_batch(telemetry):
assert json.loads(remaining[0])["properties"]["index"] == 200
def test_the_heartbeat_actually_refreshes_the_lease(telemetry):
def test_the_heartbeat_stays_well_inside_the_lease(telemetry):
"""The claim rewrite doubles as the lease heartbeat.
Previously asserted `SEND_TIMEOUT * 4 < CLAIM_STALE_SECONDS`, which compares
two constants and executes none of the code under test. Drive the real
rewrite and watch the mtime move instead.
_post makes a single attempt with SEND_TIMEOUT and no retry, so a heartbeat
lands at least that often. If a retry loop is ever added to _post, this is
the assertion that catches a sender losing its claim mid-flight.
"""
for index in range(150):
telemetry.record("search", index=index)
claim = telemetry._claim_spool()
assert claim is not None
stale = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
os.utime(claim, (stale, stale))
assert time.time() - claim.stat().st_mtime > telemetry.CLAIM_STALE_SECONDS
telemetry._rewrite_claim(claim, [{"event": "code.x", "properties": {}}])
assert time.time() - claim.stat().st_mtime < telemetry.CLAIM_STALE_SECONDS
def test_an_undeliverable_batch_is_eventually_given_up_on(telemetry):
"""Expiry has to be reachable from a state the state machine can produce.
It was not: every re-claim touched the mtime and every release backdated it
by a fixed amount, so age hovered near the stale threshold and the 7-day
expiry never fired. An undeliverable batch lived on disk forever, and
spawn_flush saw it and started a sender on every hook.
"""
telemetry.record("doomed")
telemetry._post = lambda payload, url: False
directory = telemetry.memory_core.data_dir()
for _ in range(telemetry.MAX_CLAIM_ATTEMPTS + 3):
telemetry.flush()
# Attempts now carry a cooldown, so a released claim is not instantly
# reclaimable. Age it to stand in for the wall time a real retry waits;
# without this the loop spins inside one cooldown and proves nothing.
for parked in directory.glob("telemetry-*.sending"):
stale = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
os.utime(parked, (stale, stale))
leftover = list(directory.glob("telemetry-*.sending"))
assert leftover == [], f"batch never given up on: {[p.name for p in leftover]}"
def test_a_batch_that_cannot_be_read_is_not_counted_as_delivered(telemetry):
"""Review finding: a read failure reported 'everything delivered'.
Nothing was posted, so calling it delivered lets flush() carry on to other
claims as though this batch had arrived, and hides the failure from the one
signal that says the run went badly. It also must not quarantine: a briefly
unreadable file is retryable, and moving it to .corrupt discards the events
over a transient filesystem error, because nothing ever re-globs .corrupt.
"""
telemetry.record("search")
spool = telemetry._spool_path()
stale = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
os.utime(spool, (stale, stale))
claim = telemetry._claim_spool()
assert claim is not None
original = Path.read_text
def unreadable(self, *args, **kwargs):
if self == claim:
raise OSError(5, "I/O error")
return original(self, *args, **kwargs)
Path.read_text = unreadable
try:
sent, delivered = telemetry._drain(claim)
finally:
Path.read_text = original
assert sent == 0
assert delivered is False, "an unread batch was reported as delivered"
assert claim.exists(), "a transient read error discarded the batch"
assert not list(claim.parent.glob("*.corrupt")), "quarantined over a transient error"
def test_undecodable_content_is_still_quarantined_and_the_run_continues(telemetry):
"""The other half: genuinely unrecoverable content must not block the run.
Guards the over-correction. If every read problem returned undelivered, one
torn file would stop every later claim on every flush, forever.
"""
telemetry.record("search")
spool = telemetry._spool_path()
stale = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
os.utime(spool, (stale, stale))
claim = telemetry._claim_spool()
assert claim is not None
claim.write_bytes(b"\xff\xfe torn \x00 write")
sent, delivered = telemetry._drain(claim)
assert (sent, delivered) == (0, True)
assert not claim.exists()
assert list(claim.parent.glob("*.corrupt")), "unrecoverable content was not quarantined"
def test_retries_are_spread_over_real_time_not_burned_at_once(telemetry):
"""Review finding: releasing straight to reclaimable spent the budget instantly.
Two senders hitting one momentary failure could walk a batch from attempt 0
to the limit within seconds and discard it, when a retry a minute later would
have delivered. Each release now has to age past a cooldown that grows with
the attempts already spent.
"""
telemetry.record("doomed")
telemetry._post = lambda payload, url: False
directory = telemetry.memory_core.data_dir()
telemetry.flush()
parked = list(directory.glob("telemetry-*.sending"))
assert parked, "the batch was discarded on its first failure"
assert telemetry._claim_attempt(parked[0]) == 0
# Second sender, immediately: the cooldown has not elapsed, so it must not
# be able to spend another attempt.
telemetry.flush()
still = list(directory.glob("telemetry-*.sending"))
assert len(still) == 1
assert telemetry._claim_attempt(still[0]) <= 1, "burned attempts without waiting"
def test_a_legacy_claim_filename_is_not_mistaken_for_a_huge_attempt_count(telemetry):
"""The old shape is telemetry-<pid>-<hex>.sending, and hex can start with 'a'."""
assert telemetry._claim_attempt(Path("telemetry-999-deadbeef.sending")) == 0
assert telemetry._claim_attempt(Path("telemetry-999-a1234567.sending")) == 0
assert telemetry._claim_attempt(Path("telemetry-999-deadbeef-a2.sending")) == 2
def test_a_torn_claim_is_quarantined_not_deleted(telemetry):
"""A non-empty file that parses to nothing is the remainder, not garbage."""
telemetry.record("search")
claim = telemetry._claim_spool()
claim.write_bytes(b"\xff\xfe not utf-8 at all")
stale = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
os.utime(claim, (stale, stale))
sent = telemetry.flush()
assert sent == 0
assert not claim.exists()
quarantined = list(telemetry.memory_core.data_dir().glob("*.corrupt"))
assert len(quarantined) == 1, "torn claim was destroyed instead of kept"
def test_a_failed_rewrite_stops_instead_of_redelivering(telemetry):
"""Ignoring the rewrite result reintroduced the duplicates this PR fixes."""
for index in range(250):
telemetry.record("search", index=index)
telemetry._rewrite_claim = lambda claim, remaining: False
delivered = []
telemetry._post = lambda payload, url: delivered.extend(payload.get("batch", [])) or True
telemetry.flush()
assert len(delivered) == 100, f"kept going after a failed rewrite: {len(delivered)}"
def test_partial_files_are_swept(telemetry):
"""Nothing else globs *.partial, so a crash mid-rename orphans one forever."""
data_dir = telemetry.memory_core.data_dir()
data_dir.mkdir(parents=True, exist_ok=True)
debris = data_dir / "telemetry-1-abc-a0.1.partial"
debris.write_text("x", encoding="utf-8")
old = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
os.utime(debris, (old, old))
telemetry.flush()
assert not debris.exists()
def test_quarantined_batches_are_eventually_collected(telemetry):
"""Nothing re-globs .corrupt, so without a sweep they live on disk forever.
Kept much longer than .partial debris on purpose: a quarantined batch is the
only remaining evidence of events that could not be delivered.
"""
directory = telemetry.memory_core.data_dir()
directory.mkdir(parents=True, exist_ok=True)
fresh = directory / "telemetry-1-aaaaaaaa-a0.corrupt"
old = directory / "telemetry-2-bbbbbbbb-a0.corrupt"
for path in (fresh, old):
path.write_text("torn", encoding="utf-8")
expired = time.time() - (telemetry.CLAIM_EXPIRY_SECONDS + 60)
os.utime(old, (expired, expired))
telemetry._sweep_debris(directory)
assert fresh.exists(), "a recent quarantine was discarded before anyone could look at it"
assert not old.exists(), "an expired quarantine was left on disk forever"
def test_temp_files_orphaned_by_a_kill_are_collected(telemetry):
"""_write_identity and _install_salt unlink in a finally, which SIGKILL skips."""
directory = telemetry.memory_core.data_dir()
directory.mkdir(parents=True, exist_ok=True)
orphan = directory / "telemetry-salt.999.tmp"
orphan.write_text("abandoned", encoding="utf-8")
stale = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
os.utime(orphan, (stale, stale))
telemetry._sweep_debris(directory)
assert not orphan.exists(), "a killed process left a temp file on disk forever"
assert telemetry.SEND_TIMEOUT * 4 < telemetry.CLAIM_STALE_SECONDS
@@ -177,98 +177,3 @@ def test_source_tag_defaults_agree_between_the_two_modules():
)
left, right = out.split()
assert left == right == "KIMI_PLUGIN"
def test_the_plugin_declares_its_surface_in_the_body_and_the_headers():
"""Body and headers both, because only the body works on every backend."""
core = _core_dir("claude-code-plugin")
if not core.exists():
pytest.skip("claude-code-plugin is not built in this tree")
with tempfile.TemporaryDirectory() as tmp:
out = _run(
core,
Path(tmp),
"import json, memory_core\n"
"h = memory_core.platform_headers('k')\n"
"print(json.dumps({'source': h.get('X-Mem0-Source'),"
" 'app': h.get('X-Application'),"
" 'client': h.get('X-Mem0-Client'),"
" 'auth': h.get('Authorization'),"
" 'ctype': h.get('Content-Type')}))",
)
headers = json.loads(out)
assert headers["source"] == "MEM0_PLUGIN"
assert headers["app"] == "claude-code"
assert headers["client"].startswith("mem0-plugin/")
# The transport headers the three call sites relied on must survive.
assert headers["auth"] == "Token k"
assert headers["ctype"] == "application/json"
def _session_start(core: Path, data_dir: Path) -> list[str]:
"""Drive the real hook_runner session-start path and return lifecycle events."""
recorded = "\n".join(
[
"import io, json, sys",
f"sys.path.insert(0, {str(core)!r})",
"import telemetry, hook_runner",
"seen = []",
"telemetry.record = lambda event, **kw: seen.append(event) or None",
"telemetry.spawn_flush = lambda: False",
# run() reads sys.argv through argparse; it takes no positional args.
"sys.argv = ['hook_runner', 'session-start']",
"sys.stdin = io.StringIO('{}')",
"hook_runner.run()",
"print(json.dumps([e for e in seen if e in ('install', 'upgrade')]))",
]
)
import json as _json
return _json.loads(_run(core, data_dir, recorded) or "[]")
def test_a_fresh_install_reports_install_not_upgrade():
"""The decision must survive the writes hook_runner does before asking.
claim_install() is reached only after cache_plugin_api_key() has written
`api-key` and EvidenceStore() has created `evidence.sqlite3`. Asking "is the
data dir empty" at that point always saw content, so code.install could
never fire and every new user was counted as an upgrade.
"""
core = _core_dir("claude-code-plugin")
if not core.exists():
pytest.skip("claude-code-plugin is not built in this tree")
with tempfile.TemporaryDirectory() as tmp:
data_dir = Path(tmp) / "data"
assert _session_start(core, data_dir) == ["install"]
def test_the_lifecycle_event_fires_exactly_once():
core = _core_dir("claude-code-plugin")
if not core.exists():
pytest.skip("claude-code-plugin is not built in this tree")
with tempfile.TemporaryDirectory() as tmp:
data_dir = Path(tmp) / "data"
first = _session_start(core, data_dir)
second = _session_start(core, data_dir)
third = _session_start(core, data_dir)
assert first == ["install"]
assert second == []
assert third == []
def test_an_existing_data_dir_reports_upgrade():
core = _core_dir("claude-code-plugin")
if not core.exists():
pytest.skip("claude-code-plugin is not built in this tree")
with tempfile.TemporaryDirectory() as tmp:
data_dir = Path(tmp) / "data"
data_dir.mkdir(parents=True)
# A 0.2.x leftover: the data dir survives the upgrade.
(data_dir / "requirements.txt").write_text("mem0ai\n", encoding="utf-8")
assert _session_start(core, data_dir) == ["upgrade"]
@@ -1,5 +1,3 @@
import { randomUUID } from "node:crypto";
import { redactSecrets } from "./lifecycle.ts";
const POSTHOG_API_KEY = "phc_hgJkUVJFYtmaJqrvf6CYN67TIQ8yhXAkWzUn9AMU4yX";
@@ -76,108 +74,34 @@ export function errorKind(error: unknown): string {
return error instanceof Error ? error.constructor.name : "other";
}
// Delivery is retried in memory, not spooled to disk, and that is a decision
// rather than an omission. The Python core spools because its hooks are separate
// processes that fire per tool call and exit immediately, so nothing survives
// without a file. These plugins are loaded into a host that lives for a whole
// session, so re-queueing covers the same transient failures without the claim
// and lease machinery a correct cross-process spool needs. What that leaves
// uncovered is narrow: a session that both starts and ends with no connectivity.
const RETRY_BACKOFF_CEILING_MS = 60_000;
// Consecutive failed flushes before the queue is dropped. Deliberately NOT the
// same thing as Python's budget, which rides in the claim filename and so
// follows one batch: this counter lives in the closure and counts the outage,
// not the payload. Events captured between attempts join the same queue and go
// with it. Per-batch accounting would need an attempt count on every event, and
// the queue is already bounded, so the simpler rule is the one in force here.
// Without any bound a payload the server will never accept is retried for the
// whole session and, now that the backlog is preferred over new events, holds
// the queue against everything behind it.
const MAX_DELIVERY_ATTEMPTS = 5;
export function createTelemetry(config: TelemetryConfig) {
let queue: Record<string, unknown>[] = [];
let timer: ReturnType<typeof setInterval> | undefined;
let consecutiveFailures = 0;
let retryNotBefore = 0;
let exitFlushAttempted = false;
let flushing = false;
const flushThreshold = config.flushThreshold ?? 10;
const maxQueueSize = config.maxQueueSize ?? 100;
const deliver = config.delivery ?? (async (batch: Record<string, unknown>[]) => {
const response = await fetch(POSTHOG_BATCH_URL, {
await fetch(POSTHOG_BATCH_URL, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ api_key: POSTHOG_API_KEY, batch }),
signal: AbortSignal.timeout(3_000),
});
// fetch only rejects on a network-level failure. Without this check a 500,
// a 503 or a 429 resolved normally and the batch was counted as delivered
// and dropped, which is the likelier outage than a refused connection.
// Any non-2xx is retried, matching the Python core: the backoff and the
// queue bound contain a payload that will never be accepted, because the
// re-queued batch sits at the front and is the first thing evicted.
if (!response.ok) throw new Error(`posthog responded ${response.status}`);
});
async function flush(force = false): Promise<void> {
// One at a time. Two overlapping flushes each detach the queue and each
// prepend their own batch back on failure, so the later batch lands in front
// of the earlier one and the truncation then drops the OLDER events first,
// inverting the priority the failure path exists to establish. A second
// caller returns immediately; the queue waits for the next flush.
if (flushing) return;
async function flush(): Promise<void> {
if (!queue.length) return;
// `force` skips the cooldown. beforeExit is the last chance this process
// gets, and gating it on the same backoff meant that after any failure the
// exit flush did nothing and the queue died with the process, which is the
// loss this whole mechanism exists to prevent.
if (!force && Date.now() < retryNotBefore) return;
const batch = queue;
queue = [];
flushing = true;
try {
await deliver(batch);
consecutiveFailures = 0;
retryNotBefore = 0;
} catch {
// Put it back. Detaching the batch and swallowing the error deleted the
// events outright, so any blip silently dropped telemetry with nothing
// recording that it had happened. Every event carries a uuid, so a retry
// that duplicates one PostHog already accepted is collapsed there.
//
consecutiveFailures += 1;
if (consecutiveFailures >= MAX_DELIVERY_ATTEMPTS) {
// Give up on the queue so a failing outage cannot hold it for the
// session. This drops whatever is queued now, which includes events
// captured during the outage, not only the batch that kept failing.
consecutiveFailures = 0;
retryNotBefore = 0;
return;
}
// Keep the FRONT on overflow, so the batch being retried survives and a
// new event is what gets dropped. Matches the Python core, where record()
// refuses new events once the spool is full rather than evicting the
// backlog. Keeping the newest would throw away exactly the events this
// retry exists to save.
queue = [...batch, ...queue].slice(0, maxQueueSize);
retryNotBefore = Date.now() + Math.min(2 ** consecutiveFailures * 1_000, RETRY_BACKOFF_CEILING_MS);
} finally {
flushing = false;
// Telemetry must never affect plugin behavior.
}
}
function beforeExit(): void {
// Once, and only once. Node re-emits beforeExit whenever the handler
// schedules more async work, so an unconditional forced flush looped until
// the attempt budget was spent: five attempts against a 3s delivery timeout
// is fifteen seconds added to the shutdown of whatever editor or CLI is
// hosting this. The backoff used to end that loop after one attempt, and
// removing it for the forced path removed the only thing bounding it.
if (exitFlushAttempted) return;
exitFlushAttempted = true;
void flush(true);
void flush();
}
function build(event: string, properties: Record<string, unknown> = {}): Record<string, unknown> | null {
@@ -188,15 +112,6 @@ export function createTelemetry(config: TelemetryConfig) {
return {
event: config.eventName?.(event) ?? event,
distinct_id: distinctId,
// Stamped once, at capture. This is what makes retrying safe: a batch
// re-sent after a failure carries the same ids, so PostHog collapses
// anything it already accepted instead of counting it twice.
uuid: randomUUID(),
// Capture time, not ingestion time. Events now sit through backoff and
// across a whole outage, so without this PostHog records them whenever
// delivery happened to succeed. It also matters for the uuid dedupe
// above, whose key includes the event date.
timestamp: new Date().toISOString(),
properties: {
...safeProperties(properties),
...safeProperties(config.commonProperties ?? {}),
@@ -219,10 +134,8 @@ export function createTelemetry(config: TelemetryConfig) {
try {
const payload = build(event, properties);
if (!payload) return;
// Full means drop this event, not evict the backlog. Same rule as the
// failure path above and as Python's record().
if (queue.length >= maxQueueSize) return;
queue.push(payload);
if (queue.length > maxQueueSize) queue = queue.slice(-maxQueueSize);
if (!timer) {
timer = setInterval(() => void flush(), config.flushIntervalMs ?? 5_000);
timer.unref?.();
@@ -236,8 +149,6 @@ export function createTelemetry(config: TelemetryConfig) {
function resetForTesting(): void {
queue = [];
consecutiveFailures = 0;
retryNotBefore = 0;
if (timer) clearInterval(timer);
timer = undefined;
process.off("beforeExit", beforeExit);
@@ -122,250 +122,3 @@ test("error classification does not expose messages", () => {
assert.equal(errorKind(new Error("request timeout")), "timeout");
assert.equal(errorKind(new Error("fetch failed")), "network");
});
test("a failed delivery keeps the batch instead of deleting it", async () => {
// The defect: the queue was detached before the await and the error swallowed,
// so one blip destroyed the events with nothing recording that it happened.
const attempts: Record<string, unknown>[][] = [];
let failNext = true;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d",
flushThreshold: 1000,
delivery: async (batch) => {
attempts.push(batch);
if (failNext) throw new Error("network down");
},
});
telemetry.capture("one");
telemetry.capture("two");
await telemetry.flush();
assert.equal(attempts.length, 1);
assert.equal(telemetry.queueForTesting().length, 2, "events were dropped on failure");
failNext = false;
// Backoff is in force, so wait it out the way wall time would.
await new Promise((resolve) => setTimeout(resolve, 2_100));
await telemetry.flush();
assert.equal(attempts.length, 2, "never retried");
assert.equal(telemetry.queueForTesting().length, 0);
telemetry.resetForTesting();
});
test("a retried event carries the same uuid so PostHog can collapse it", async () => {
const attempts: Record<string, unknown>[][] = [];
let failNext = true;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d",
flushThreshold: 1000,
delivery: async (batch) => {
attempts.push(batch);
if (failNext) throw new Error("network down");
},
});
telemetry.capture("once");
await telemetry.flush();
failNext = false;
await new Promise((resolve) => setTimeout(resolve, 2_100));
await telemetry.flush();
assert.equal(attempts.length, 2);
const first = attempts[0][0].uuid;
assert.ok(first, "events carry no uuid, so a retry would double count");
assert.equal(attempts[1][0].uuid, first, "retry minted a new uuid");
telemetry.resetForTesting();
});
test("repeated failures back off instead of retrying every flush", async () => {
let calls = 0;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d",
flushThreshold: 1000,
delivery: async () => { calls += 1; throw new Error("blocked"); },
});
telemetry.capture("one");
await telemetry.flush();
await telemetry.flush();
await telemetry.flush();
assert.equal(calls, 1, "a blocked host was hammered on every flush");
assert.equal(telemetry.queueForTesting().length, 1, "the event was lost while backing off");
telemetry.resetForTesting();
});
test("a full queue drops the new event and keeps the batch being retried", async () => {
// Python's record() refuses new events once the spool is full rather than
// evicting the backlog. Keeping the newest here would throw away exactly the
// events the retry exists to save.
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d",
flushThreshold: 1000, maxQueueSize: 3,
delivery: async () => { throw new Error("down"); },
});
// Fill past the cap BEFORE the flush, so the re-queue actually has to truncate.
// Capturing only two left the queue empty at re-queue time and the slice on the
// failure path never ran, which is the half that decides the direction.
telemetry.capture("a");
telemetry.capture("b");
telemetry.capture("c");
await telemetry.flush();
telemetry.capture("d");
telemetry.capture("e");
const events = telemetry.queueForTesting().map((e) => (e as any).event);
assert.equal(events.length, 3, "queue grew past maxQueueSize");
assert.deepEqual(events, ["a", "b", "c"], "the retried batch was evicted instead of the new events");
telemetry.resetForTesting();
});
test("the exit-time flush ignores the backoff", async () => {
// beforeExit is the last chance the process gets. Gating it on the same
// cooldown meant that after any failure it did nothing and the queue died.
let attempts = 0;
let failing = true;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: async () => { attempts += 1; if (failing) throw new Error("down"); },
});
telemetry.capture("a");
await telemetry.flush();
assert.equal(attempts, 1);
failing = false;
await telemetry.flush();
assert.equal(attempts, 1, "the backoff should still hold for an ordinary flush");
await telemetry.flush(true);
assert.equal(attempts, 2, "the exit flush was suppressed by the backoff");
assert.equal(telemetry.queueForTesting().length, 0);
telemetry.resetForTesting();
});
test("every event carries a capture-time timestamp", async () => {
const sent: Record<string, unknown>[][] = [];
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: async (batch) => { sent.push(batch); },
});
telemetry.capture("a");
const capturedAt = Date.now();
await new Promise((resolve) => setTimeout(resolve, 50));
await telemetry.flush();
const stamped = sent[0][0].timestamp as string;
assert.ok(stamped, "no timestamp, so PostHog would record delivery time");
assert.ok(Math.abs(Date.parse(stamped) - capturedAt) < 1_000, "not capture time");
telemetry.resetForTesting();
});
test("a batch the server will never accept is eventually given up on", async () => {
let attempts = 0;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: async () => { attempts += 1; throw new Error("permanently bad"); },
});
telemetry.capture("doomed");
for (let i = 0; i < 8; i += 1) await telemetry.flush(true);
assert.ok(attempts <= 6, `retried ${attempts} times with no cap`);
assert.equal(telemetry.queueForTesting().length, 0, "a doomed batch held the queue forever");
telemetry.resetForTesting();
});
test("an HTTP error response is a failure, not a delivery", async () => {
// fetch only rejects on a network-level failure, so a 500 used to resolve
// normally and the batch was dropped as delivered. Exercises the real default
// delivery path rather than an injected one, which is where this hid.
const realFetch = globalThis.fetch;
let calls = 0;
globalThis.fetch = (async () => {
calls += 1;
return new Response("upstream is unwell", { status: 503 });
}) as typeof fetch;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
});
try {
telemetry.capture("during.outage");
await telemetry.flush();
assert.equal(calls, 1, "never reached the network");
assert.equal(telemetry.queueForTesting().length, 1, "a 503 was counted as delivered");
} finally {
globalThis.fetch = realFetch;
telemetry.resetForTesting();
}
});
test("a 2xx is a delivery", async () => {
const realFetch = globalThis.fetch;
globalThis.fetch = (async () => new Response("ok", { status: 200 })) as typeof fetch;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
});
try {
telemetry.capture("fine");
await telemetry.flush();
assert.equal(telemetry.queueForTesting().length, 0, "a good response did not clear the queue");
} finally {
globalThis.fetch = realFetch;
telemetry.resetForTesting();
}
});
test("the exit flush is attempted once, not until the budget is spent", async () => {
// Node re-emits beforeExit whenever the handler schedules async work, so an
// unconditional forced flush looped until MAX_DELIVERY_ATTEMPTS. Against the
// real 3s delivery timeout that is fifteen seconds added to a host's shutdown.
let attempts = 0;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: async () => { attempts += 1; throw new Error("down"); },
});
telemetry.capture("a");
const handlers = process.listeners("beforeExit");
const ours = handlers[handlers.length - 1] as () => void;
ours();
ours();
ours();
await new Promise((resolve) => setTimeout(resolve, 20));
assert.equal(attempts, 1, `exit flush ran ${attempts} times`);
telemetry.resetForTesting();
});
test("overlapping flushes do not reorder the backlog behind newer events", async () => {
// Each flush detaches the queue and prepends its own batch back on failure, so
// two in flight at once put the LATER batch in front of the earlier one. The
// truncation then drops the older events first, inverting the priority the
// failure path exists to establish.
let release: (() => void)[] = [];
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: () => new Promise((_resolve, reject) => { release.push(() => reject(new Error("down"))); }),
});
telemetry.capture("first");
const a = telemetry.flush();
telemetry.capture("second");
const b = telemetry.flush();
release.forEach((fn) => fn());
await Promise.all([a, b]);
const events = telemetry.queueForTesting().map((e) => (e as any).event);
assert.equal(release.length, 1, "a second delivery started while one was in flight");
assert.deepEqual(events, ["first", "second"], `backlog reordered: ${events.join(",")}`);
telemetry.resetForTesting();
});
@@ -5,7 +5,5 @@ SOURCE_TAG = "ANTIGRAVITY_PLUGIN"
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
# plugin family is one source; which editor it runs in is the application.
# An empty application means the host is unknown, and memory_core omits
# the header entirely rather than sending a placeholder.
PLATFORM_SOURCE = "MEM0_PLUGIN"
PLATFORM_APPLICATION = "antigravity"
@@ -290,11 +290,6 @@ def run(
if args.plugin_data_dir:
os.environ[data_dir_env] = args.plugin_data_dir
# Snapshot BEFORE anything writes to the data dir: cache_plugin_api_key
# writes `api-key` and EvidenceStore creates `evidence.sqlite3`, so asking
# after them always saw content and every fresh install reported an upgrade.
data_dir_was_empty = telemetry.data_dir_was_empty()
cache_plugin_api_key()
if args.action == "session-start":
clear_stale_api_key_cache()
@@ -312,7 +307,7 @@ def run(
if args.action == "session-start":
# Claims the marker atomically and says which event to record, so a
# second session starting alongside this one cannot record it too.
first_event = telemetry.claim_install(was_empty=data_dir_was_empty)
first_event = telemetry.claim_install()
if first_event == "install":
telemetry.record("install")
elif first_event == "upgrade":
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.2"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -2008,12 +2008,9 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
# Top level, not metadata: the backend reads `source` from the body,
# query string or X-Mem0-Source header, never from metadata. The
# harness tag stays in metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
+39 -354
View File
@@ -49,7 +49,6 @@ except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
_salt_cache: str = ""
_harness: str = _DEFAULT_HARNESS
_source_tag: str = _DEFAULT_SOURCE_TAG
_PRIVATE_KEYS = {
@@ -105,11 +104,6 @@ MAX_CLAIM_ATTEMPTS = 3
# Parked claims drained per run, after the live spool. Bounded so a long backlog
# cannot turn one flush into an unbounded send loop.
MAX_PARKED_PER_RUN = 3
# Added to the wait before a released claim becomes reclaimable, per attempt
# already spent. Releasing straight to "reclaimable now" let two senders burn the
# whole budget within seconds of one another on a single momentary failure, and
# discard a batch a retry a minute later would have delivered.
RETRY_COOLDOWN_SECONDS = 60
def is_enabled() -> bool:
@@ -127,98 +121,15 @@ def _digest(value: str, length: int = 16) -> str:
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
def _salt_path() -> Path:
return memory_core.data_dir() / "telemetry-salt"
def _install_salt() -> str:
"""Random per-install salt, created once and memoized for the process.
Deliberately its own file, claimed with O_CREAT|O_EXCL, rather than a key in
the identity file. Three reasons, all of which produced wrong data when this
lived in the identity dict:
- Hooks are short-lived separate processes firing on every tool call, and
people run more than one agent window. A read-modify-write would let each
process mint its own salt, so one repository would hash several ways in the
window before a writer won.
- resolve_distinct_id holds a copy of the identity dict across a network call
to /v1/ping/, so whichever write landed second erased the other's key —
losing either the salt (repo_hash changes mid-stream) or the email (a
second $identify, splitting the person).
- Touching the identity file from record() would create it, and is_first_run
keys off that file, so recording an event would silently suppress the
install event.
Published atomically, and there is deliberately no derived fallback. Creating
the file with O_CREAT|O_EXCL and then writing into it leaves a window where
the file exists and is empty, and a concurrent hook that reads it in that
window gets nothing. Falling back to a digest of the path would hand that
process a salt an attacker can compute, memoized for its whole run, which is
the privacy control this function exists to provide silently turning itself
off under load. The salt is written to a private temp file first and linked
into place, so the name either does not exist or already has the full value.
Returns "" when it genuinely cannot persist. Callers omit the hash entirely
rather than emit an unsalted one.
"""
global _salt_cache
if _salt_cache:
return _salt_cache
path = _salt_path()
# Read before writing. Hooks are separate processes firing on every tool
# call, so all but the first find the salt already published; going straight
# to create-fsync-link-unlink meant every one of them paid an fsync to
# discover that, on a path whose whole promise is appending a line and
# returning.
try:
_salt_cache = path.read_text(encoding="utf-8").strip()
if _salt_cache:
return _salt_cache
except OSError:
pass
temporary = path.with_name(f"{path.name}.{os.getpid()}.tmp")
try:
path.parent.mkdir(parents=True, exist_ok=True)
handle = os.open(temporary, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
with os.fdopen(handle, "w", encoding="utf-8") as stream:
stream.write(uuid.uuid4().hex)
stream.flush()
os.fsync(stream.fileno())
try:
# Atomic claim: fails if another process already published one.
# os.link rather than replace, which would clobber theirs.
os.link(temporary, path)
except FileExistsError:
pass
except OSError:
# No hardlinks here (some network mounts, some container volumes).
# Claim the name directly instead. That reopens the empty-file
# window, but the window is now benign: a reader that lands in it
# gets "" and omits the hash for that process rather than caching a
# guessable one. Losing the hashes on every run of an entire
# filesystem is the worse failure.
try:
fallback = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
with os.fdopen(fallback, "w", encoding="utf-8") as stream:
stream.write(temporary.read_text(encoding="utf-8"))
except OSError:
pass
except OSError:
pass
finally:
try:
temporary.unlink()
except OSError:
pass
try:
_salt_cache = path.read_text(encoding="utf-8").strip()
except OSError:
_salt_cache = ""
return _salt_cache
"""Random per-install salt, created on first use and kept in the identity file."""
identity = _read_identity()
salt = identity.get("salt")
if not salt:
salt = uuid.uuid4().hex
identity["salt"] = salt
_write_identity(identity)
return salt
def _scoped_digest(value: str, length: int = 16) -> str:
@@ -230,17 +141,10 @@ def _scoped_digest(value: str, length: int = 16) -> str:
privacy control without the salt. Salting per install keeps every
within-account join the analytics actually use and gives up only
cross-machine joins on the same repository, which nothing computes.
Returns "" when there is no salt, so record() omits the property. An
unsalted digest over this input space is close to plaintext, and emitting one
under a name that implies it is hashed is worse than sending nothing.
"""
if not value:
return ""
salt = _install_salt()
if not salt:
return ""
return hashlib.sha256(f"{salt}:{value}".encode("utf-8")).hexdigest()[:length]
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
def _safe_value(value: Any) -> Any:
@@ -302,25 +206,6 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
return created
def _rotate_anonymous_id(identity: dict[str, str]) -> str:
"""Mint a fresh anonymous id because the account context is gone.
The previous id may already have been merged into a person profile by an
$identify, and that merge is permanent. Reusing it after a logout or a key
change attributes everything that follows to the account that just went
away, which is the same misattribution the key fingerprint exists to stop,
only arriving through the anonymous path instead.
`aliased` is cleared with it: the new id has never been merged, so it is
eligible to be aliased into whatever account comes next.
"""
created = f"code-anon-{uuid.uuid4().hex}"
identity["anonymous_id"] = created
identity.pop("aliased", None)
_write_identity(identity)
return created
def _install_state_path() -> Path:
return memory_core.data_dir() / "install-state.json"
@@ -336,35 +221,18 @@ def is_first_run() -> bool:
return not _install_state_path().exists()
def data_dir_was_empty() -> bool:
"""Whether the data directory is untouched. Call BEFORE anything writes to it.
hook_runner reaches claim_install() only after cache_plugin_api_key() has
written `api-key` and EvidenceStore() has created `evidence.sqlite3`, so
asking at claim time always saw content and every fresh install reported an
upgrade. The caller snapshots this at the top of the run instead.
"""
return not _data_dir_has_content()
def claim_install(was_empty: bool | None = None) -> str | None:
def claim_install() -> str | None:
"""Claim the one install/upgrade record for this machine, atomically.
Returns the event to record ("install" or "upgrade"), or None if another
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
cannot both win.
`was_empty` must come from data_dir_was_empty() called before this process
wrote anything. Omitting it falls back to checking now, which is only
correct for a caller that has touched nothing.
"""
if not is_enabled():
# Never consume the one-shot claim while the user is opted out, or they
# would silently lose their install event if they later opt in.
return None
path = _install_state_path()
upgrading = not (data_dir_was_empty() if was_empty is None else was_empty)
# A fresh install has an empty data directory. Anything already there —
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
# before the marker is created, since creating it would itself be content.
upgrading = _data_dir_has_content()
try:
path.parent.mkdir(parents=True, exist_ok=True)
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
@@ -382,14 +250,6 @@ def claim_install(was_empty: bool | None = None) -> str | None:
},
stream,
)
# Durable before this returns. The O_EXCL open is what makes the
# claim exclusive, so it cannot be replaced by a temp-and-rename
# without losing that, which leaves the content as the thing to make
# safe. A kill between the open and this fsync used to leave a marker
# that exists but parses to nothing: is_first_run reads it as claimed
# and claim_version_change cannot read a version out of it.
stream.flush()
os.fsync(stream.fileno())
except OSError:
pass
return "upgrade" if upgrading else "install"
@@ -406,19 +266,6 @@ def _data_dir_has_content() -> bool:
return False
def _repair_install_state(path: Path) -> None:
"""Rewrite an unparseable marker so version tracking can resume."""
try:
temporary = path.with_suffix(f".{os.getpid()}.tmp")
temporary.write_text(
json.dumps({"plugin_version": memory_core.PLUGIN_VERSION, "repaired_at": memory_core.utc_now()}),
encoding="utf-8",
)
temporary.replace(path)
except OSError:
pass
def claim_version_change() -> str | None:
"""Return the previously recorded version if it differs, updating the marker.
@@ -429,47 +276,20 @@ def claim_version_change() -> str | None:
path = _install_state_path()
try:
state = json.loads(path.read_text(encoding="utf-8"))
except OSError:
except (OSError, json.JSONDecodeError):
return None
except json.JSONDecodeError:
# A crash between O_EXCL and the write leaves an empty marker. Left
# alone it disables every future upgrade event on this machine, because
# claim_install sees the file and this function cannot parse it.
state = None
if not isinstance(state, dict):
_repair_install_state(path)
return None
previous = str(state.get("plugin_version") or "")
if not previous or previous == memory_core.PLUGIN_VERSION:
return None
# Claim the transition with an exclusive sentinel before rewriting the
# marker. A plain read-modify-write let every concurrently starting session
# observe the old version and each record its own upgrade — and the first
# session after a version bump is exactly when several agent windows restart
# together.
sentinel = path.with_name(f"upgraded-{memory_core.PLUGIN_VERSION}")
try:
os.close(os.open(sentinel, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600))
except FileExistsError:
return None
except OSError:
return None
state["plugin_version"] = memory_core.PLUGIN_VERSION
state["upgraded_at"] = memory_core.utc_now()
temporary = path.with_suffix(f".{os.getpid()}.tmp")
try:
temporary = path.with_suffix(f".{os.getpid()}.tmp")
temporary.write_text(json.dumps(state), encoding="utf-8")
temporary.replace(path)
except OSError:
# Release the claim. The marker still records the old version, so
# without this the sentinel makes claim_version_change return early on
# every later run and this version's upgrade is never recorded again.
for leftover in (sentinel, temporary):
try:
leftover.unlink()
except OSError:
pass
return None
return previous
@@ -503,17 +323,10 @@ def record(
os=sys.platform,
python_version=platform.python_version(),
)
# Assigned only when the digest is real. _scoped_digest returns "" when
# the salt could not be persisted, and an empty property is worse than an
# absent one: it survives the None filter below and reads as a value.
if repo is not None:
repo_hash = _scoped_digest(getattr(repo, "identity", ""))
if repo_hash:
properties["repo_hash"] = repo_hash
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
if session_id:
session_hash = _scoped_digest(session_id)
if session_hash:
properties["session_hash"] = session_hash
properties["session_hash"] = _scoped_digest(session_id)
line = json.dumps(
{
"event": f"{EVENT_PREFIX}.{event}",
@@ -583,18 +396,9 @@ def _claim_name(attempt: int = 0) -> str:
def _claim_attempt(claim: Path) -> int:
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape.
Anchored on field position, not on a leading "a": the legacy shape is
``telemetry-<pid>-<hex>.sending`` and a hex id such as ``a1234567`` would
otherwise parse as attempt 1234567 and be discarded unsent on the first
flush after an upgrade.
"""
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
parts = stem.split("-")
if len(parts) != 4:
return 0
tail = parts[3]
tail = stem.rsplit("-", 1)[-1]
if tail.startswith("a") and tail[1:].isdigit():
return int(tail[1:])
return 0
@@ -628,42 +432,6 @@ def _claim_spool() -> Path | None:
return _claim_parked(directory)
def _sweep_debris(directory: Path) -> None:
"""Remove files nothing else will ever pick up again.
*.partial is a temp file orphaned by a crash between write and rename.
*.corrupt is a batch quarantined for undecodable content. No glob in this
module matches either, so without this they accumulate on disk for the life
of the install.
Quarantined batches are kept far longer than debris: they are the only
evidence left of events that could not be delivered, and someone diagnosing
a report of missing telemetry has to be able to find one.
"""
now = time.time()
for debris in directory.glob("telemetry-*.partial"):
try:
if now - debris.stat().st_mtime > CLAIM_STALE_SECONDS:
debris.unlink()
except OSError:
continue
for quarantined in directory.glob("telemetry-*.corrupt"):
try:
if now - quarantined.stat().st_mtime > CLAIM_EXPIRY_SECONDS:
quarantined.unlink()
except OSError:
continue
# The same reasoning covers *.tmp. _write_identity and _install_salt both
# create one and unlink it in a finally, which a SIGKILL skips, and no glob
# in this module matches the leftovers either.
for temporary in directory.glob("telemetry-*.tmp"):
try:
if now - temporary.stat().st_mtime > CLAIM_STALE_SECONDS:
temporary.unlink()
except OSError:
continue
def _claim_parked(directory: Path) -> Path | None:
"""Take the oldest abandoned claim, if any lease has actually expired.
@@ -679,24 +447,15 @@ def _claim_parked(directory: Path) -> Path | None:
age = now - orphan.stat().st_mtime
except OSError:
continue
if age < CLAIM_STALE_SECONDS:
# Someone else holds a live lease on it. This check has to come
# first. Claiming a file bumps its attempt count and refreshes its
# mtime, so a sender that has just taken the final attempt looks
# exhausted to everyone else while it is actively draining. Judging
# exhaustion before liveness let a second sender unlink a batch out
# from under its owner, losing every event in it.
continue
# Attempts, not age. Every re-claim touches the mtime and every release
# backdates it by a fixed amount, so age is pinned near the stale
# threshold and never reaches the expiry. Age stays only as a backstop
# for files that never carried an attempt marker.
if _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS or age > CLAIM_EXPIRY_SECONDS:
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
try:
orphan.unlink()
except OSError:
pass
continue
if age < CLAIM_STALE_SECONDS:
# Someone else holds a live lease on it.
continue
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
try:
orphan.replace(claim)
@@ -730,14 +489,10 @@ def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
return True
temporary = claim.with_suffix(f".{os.getpid()}.partial")
try:
payload = "".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining)
# fsync before the rename: without it the rename can land while the
# bytes have not, and the claim comes back empty or truncated after a
# crash. _drain then reads zero events and unlinks it.
with open(temporary, "w", encoding="utf-8") as handle:
handle.write(payload)
handle.flush()
os.fsync(handle.fileno())
temporary.write_text(
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
encoding="utf-8",
)
temporary.replace(claim)
_touch(claim)
return True
@@ -761,11 +516,7 @@ def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
if not _rewrite_claim(claim, remaining):
return
try:
# Backdate past the stale threshold so the next flush can pick it up,
# minus a cooldown that grows with the attempts already spent. Clamped so
# the mtime never lands in the future, which would read as a live lease.
cooldown = min(_claim_attempt(claim) * RETRY_COOLDOWN_SECONDS, CLAIM_STALE_SECONDS)
released = time.time() - CLAIM_STALE_SECONDS - 1 + cooldown
released = time.time() - CLAIM_STALE_SECONDS - 1
os.utime(claim, (released, released))
except OSError:
pass
@@ -812,56 +563,23 @@ def resolve_distinct_id() -> tuple[str, str]:
fingerprint = _digest(key) if key else ""
email = identity.get("email", "")
if email and fingerprint:
recorded = identity.get("key_fingerprint", "")
if recorded == fingerprint:
return email, ""
if not recorded:
# Rows written before fingerprints existed. Verify rather than
# adopt: a key changed before the upgrade would otherwise bind the
# new key to the previous account's email, permanently, and the
# fingerprint would then agree with itself forever after.
verified = _resolve_email(key)
if not verified:
# Offline, firewalled, or the API is down. Keep the previous
# behaviour and retry on the next flush rather than dropping a
# real account attribution. Safe because the same network that
# failed /v1/ping/ is about to fail the PostHog POST, so nothing
# is delivered under the unverified identity in the meantime.
return email, ""
identity["email"] = verified
identity["key_fingerprint"] = fingerprint
_write_identity(identity)
return verified, ""
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
return email, ""
if not key:
# No key to verify the account with; do not keep attributing to it.
if email:
identity.pop("email", None)
identity.pop("key_fingerprint", None)
return _rotate_anonymous_id(identity), ""
_write_identity(identity)
return anonymous_id(identity), ""
resolved = _resolve_email(key)
if not resolved:
# The key changed and will not resolve (revoked, offline, API down).
# Reaching here with an email means the recorded fingerprint disagreed,
# so the key really did change. Drop the account and rotate: the stored
# anonymous id may already be merged into that account's person, and
# reusing it would keep the events on the profile we are trying to
# leave.
if email:
identity.pop("email", None)
identity.pop("key_fingerprint", None)
return _rotate_anonymous_id(identity), ""
return anonymous_id(identity), ""
return (email, "") if email else (anonymous_id(identity), "")
# Alias only when going anonymous -> email for the first time. Once an anon
# id has been merged into an account it must never be offered again: an
# alias naming an already-identified id is what could link two real people.
previous = "" if (email or identity.get("aliased")) else identity.get("anonymous_id", "")
if previous:
identity["aliased"] = True
# Alias only when going anonymous -> email for the first time.
previous = "" if email else identity.get("anonymous_id", "")
identity["email"] = resolved
identity["key_fingerprint"] = fingerprint
_write_identity(identity)
@@ -881,7 +599,6 @@ def flush() -> int:
# Parked batches used to starve behind the live spool indefinitely. Bounded
# per run so a long backlog cannot turn one flush into an unbounded loop.
directory = memory_core.data_dir()
_sweep_debris(directory)
for _ in range(MAX_PARKED_PER_RUN):
parked = _claim_parked(directory)
if parked is None:
@@ -902,28 +619,8 @@ def _drain(claim: Path | None) -> tuple[int, bool]:
return 0, True
try:
lines = claim.read_text(encoding="utf-8").splitlines()
except ValueError:
# UnicodeDecodeError from a torn write: the content is unrecoverable, so
# quarantine rather than retry. flush() runs from a bare `finally:` in
# flush_worker, so raising here also skips the handoff cleanup, and an
# undecodable file would otherwise be re-read on every flush forever.
# Reported as delivered because there is nothing left to deliver and the
# rest of the run should continue.
try:
claim.replace(claim.with_suffix(".corrupt"))
except OSError:
try:
claim.unlink()
except OSError:
pass
return 0, True
except OSError:
# Could not read it, which is not the same as having nothing to send.
# The file is left exactly where it is: a vanished or briefly unreadable
# claim is retryable, and quarantining it here would discard events over
# a transient filesystem error. Reported as undelivered so the run stops
# instead of counting a batch nothing was posted from as delivered.
return 0, False
return 0, True
events = []
for line in lines:
try:
@@ -933,15 +630,8 @@ def _drain(claim: Path | None) -> tuple[int, bool]:
if isinstance(value, dict) and value.get("event"):
events.append(value)
if not events:
# Only delete when the file really is empty. A non-empty file that
# parses to nothing is a torn write, and its contents are the unsent
# remainder — deleting it is the data loss this PR exists to prevent.
try:
empty = claim.stat().st_size == 0
except OSError:
empty = True
try:
claim.replace(claim.with_suffix(".corrupt")) if not empty else claim.unlink()
claim.unlink()
except OSError:
pass
return 0, True
@@ -991,13 +681,8 @@ def _drain(claim: Path | None) -> tuple[int, bool]:
return sent, False
sent += len(chunk)
# Record progress and refresh the lease after each successful batch, so
# a crash repeats at most one batch instead of the entire file. If the
# rewrite fails the claim still holds delivered events, so stop rather
# than carry on as though progress were recorded — continuing is how the
# duplicate delivery this PR fixes would come back.
if not _rewrite_claim(claim, events[start + len(chunk) :]):
_release_claim(claim, events[start + len(chunk) :])
return sent, False
# a crash repeats at most one batch instead of the entire file.
_rewrite_claim(claim, events[start + len(chunk) :])
return sent, True
@@ -1,6 +1,6 @@
{
"id": "mem0",
"version": "0.3.2",
"version": "0.3.1",
"homepage": "https://docs.mem0.ai/integrations/antigravity",
"native": {
"pluginRoot": "${ANTIGRAVITY_PLUGIN_ROOT}",
@@ -7,8 +7,8 @@ disable-model-invocation: true
# Pause memory capture
To pause (hooks stop capturing and sending session content; a minimal
telemetry ping still fires at session start, under your Mem0 account email,
unless `MEM0_TELEMETRY=false`):
anonymous telemetry ping still fires at session start unless
`MEM0_TELEMETRY=false`):
```bash
python3 "${ANTIGRAVITY_PLUGIN_ROOT}/core/memory_cli.py" --harness "antigravity" pause
@@ -1,6 +1,6 @@
{
"name": "mem0",
"version": "0.3.2",
"version": "0.3.1",
"description": "Cross-session memory and token savings for coding agents.",
"author": {
"name": "Mem0"
+1 -5
View File
@@ -137,8 +137,6 @@ Local data lives in `${CLAUDE_PLUGIN_DATA}`:
- `flush-worker.log`: whether memory creation succeeded
- `plugin-errors.log`: hook errors (no credentials)
- `telemetry.jsonl` / `telemetry-identity.json`: usage events and the id they are sent under
- `telemetry-salt`: random per-install salt for the repo and session hashes
- `install-state.json`: records that install has been counted once on this machine
Mem0 receives captured user messages, Claude's answers, sidekick assignments and completed responses, and changed file paths. When a failed command is recorded, extraction can also include bounded command details and results. Complete files and general tool output stay on your machine. Values that look like credentials are redacted before anything is sent.
@@ -148,9 +146,7 @@ Usage events (which hook ran, timing, result counts, failure types) so Mem0 can
**These events are not anonymous.** When an API key is configured — which installing the plugin requires — events are sent under your Mem0 account email, the same way the Python SDK and the CLI attribute theirs. Without a key they are sent under a random per-machine id.
What each event carries: the event name, the plugin version, the harness it ran in, your OS and Python version, and per-event properties describing what happened — timings, counts, coarse outcome and failure labels, and which model was configured. Repository and session identifiers are hashed with a random salt generated on your machine, so they cannot be linked back to a repository name or path.
Rather than restate a list that drifts, the exact set is enforced in code: `telemetry.record` filters every property through a denylist of sensitive keys and redacts credential-shaped values. See `_PRIVATE_KEYS` in `core/telemetry.py`.
What each event carries: the event name, the plugin version, the harness it ran in, your OS and Python version, timings, counts, and a coarse failure label. Repository and session identifiers are hashed with a random salt generated on your machine and never sent, so they cannot be linked back to a repository name or path.
Prompts, memory text, queries, file paths, repository names, and API keys are never sent.
@@ -5,7 +5,5 @@ SOURCE_TAG = "CLAUDE_CODE_PLUGIN"
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
# plugin family is one source; which editor it runs in is the application.
# An empty application means the host is unknown, and memory_core omits
# the header entirely rather than sending a placeholder.
PLATFORM_SOURCE = "MEM0_PLUGIN"
PLATFORM_APPLICATION = "claude-code"
@@ -290,11 +290,6 @@ def run(
if args.plugin_data_dir:
os.environ[data_dir_env] = args.plugin_data_dir
# Snapshot BEFORE anything writes to the data dir: cache_plugin_api_key
# writes `api-key` and EvidenceStore creates `evidence.sqlite3`, so asking
# after them always saw content and every fresh install reported an upgrade.
data_dir_was_empty = telemetry.data_dir_was_empty()
cache_plugin_api_key()
if args.action == "session-start":
clear_stale_api_key_cache()
@@ -312,7 +307,7 @@ def run(
if args.action == "session-start":
# Claims the marker atomically and says which event to record, so a
# second session starting alongside this one cannot record it too.
first_event = telemetry.claim_install(was_empty=data_dir_was_empty)
first_event = telemetry.claim_install()
if first_event == "install":
telemetry.record("install")
elif first_event == "upgrade":
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.2"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -2008,12 +2008,9 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
# Top level, not metadata: the backend reads `source` from the body,
# query string or X-Mem0-Source header, never from metadata. The
# harness tag stays in metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
+39 -354
View File
@@ -49,7 +49,6 @@ except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
_salt_cache: str = ""
_harness: str = _DEFAULT_HARNESS
_source_tag: str = _DEFAULT_SOURCE_TAG
_PRIVATE_KEYS = {
@@ -105,11 +104,6 @@ MAX_CLAIM_ATTEMPTS = 3
# Parked claims drained per run, after the live spool. Bounded so a long backlog
# cannot turn one flush into an unbounded send loop.
MAX_PARKED_PER_RUN = 3
# Added to the wait before a released claim becomes reclaimable, per attempt
# already spent. Releasing straight to "reclaimable now" let two senders burn the
# whole budget within seconds of one another on a single momentary failure, and
# discard a batch a retry a minute later would have delivered.
RETRY_COOLDOWN_SECONDS = 60
def is_enabled() -> bool:
@@ -127,98 +121,15 @@ def _digest(value: str, length: int = 16) -> str:
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
def _salt_path() -> Path:
return memory_core.data_dir() / "telemetry-salt"
def _install_salt() -> str:
"""Random per-install salt, created once and memoized for the process.
Deliberately its own file, claimed with O_CREAT|O_EXCL, rather than a key in
the identity file. Three reasons, all of which produced wrong data when this
lived in the identity dict:
- Hooks are short-lived separate processes firing on every tool call, and
people run more than one agent window. A read-modify-write would let each
process mint its own salt, so one repository would hash several ways in the
window before a writer won.
- resolve_distinct_id holds a copy of the identity dict across a network call
to /v1/ping/, so whichever write landed second erased the other's key —
losing either the salt (repo_hash changes mid-stream) or the email (a
second $identify, splitting the person).
- Touching the identity file from record() would create it, and is_first_run
keys off that file, so recording an event would silently suppress the
install event.
Published atomically, and there is deliberately no derived fallback. Creating
the file with O_CREAT|O_EXCL and then writing into it leaves a window where
the file exists and is empty, and a concurrent hook that reads it in that
window gets nothing. Falling back to a digest of the path would hand that
process a salt an attacker can compute, memoized for its whole run, which is
the privacy control this function exists to provide silently turning itself
off under load. The salt is written to a private temp file first and linked
into place, so the name either does not exist or already has the full value.
Returns "" when it genuinely cannot persist. Callers omit the hash entirely
rather than emit an unsalted one.
"""
global _salt_cache
if _salt_cache:
return _salt_cache
path = _salt_path()
# Read before writing. Hooks are separate processes firing on every tool
# call, so all but the first find the salt already published; going straight
# to create-fsync-link-unlink meant every one of them paid an fsync to
# discover that, on a path whose whole promise is appending a line and
# returning.
try:
_salt_cache = path.read_text(encoding="utf-8").strip()
if _salt_cache:
return _salt_cache
except OSError:
pass
temporary = path.with_name(f"{path.name}.{os.getpid()}.tmp")
try:
path.parent.mkdir(parents=True, exist_ok=True)
handle = os.open(temporary, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
with os.fdopen(handle, "w", encoding="utf-8") as stream:
stream.write(uuid.uuid4().hex)
stream.flush()
os.fsync(stream.fileno())
try:
# Atomic claim: fails if another process already published one.
# os.link rather than replace, which would clobber theirs.
os.link(temporary, path)
except FileExistsError:
pass
except OSError:
# No hardlinks here (some network mounts, some container volumes).
# Claim the name directly instead. That reopens the empty-file
# window, but the window is now benign: a reader that lands in it
# gets "" and omits the hash for that process rather than caching a
# guessable one. Losing the hashes on every run of an entire
# filesystem is the worse failure.
try:
fallback = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
with os.fdopen(fallback, "w", encoding="utf-8") as stream:
stream.write(temporary.read_text(encoding="utf-8"))
except OSError:
pass
except OSError:
pass
finally:
try:
temporary.unlink()
except OSError:
pass
try:
_salt_cache = path.read_text(encoding="utf-8").strip()
except OSError:
_salt_cache = ""
return _salt_cache
"""Random per-install salt, created on first use and kept in the identity file."""
identity = _read_identity()
salt = identity.get("salt")
if not salt:
salt = uuid.uuid4().hex
identity["salt"] = salt
_write_identity(identity)
return salt
def _scoped_digest(value: str, length: int = 16) -> str:
@@ -230,17 +141,10 @@ def _scoped_digest(value: str, length: int = 16) -> str:
privacy control without the salt. Salting per install keeps every
within-account join the analytics actually use and gives up only
cross-machine joins on the same repository, which nothing computes.
Returns "" when there is no salt, so record() omits the property. An
unsalted digest over this input space is close to plaintext, and emitting one
under a name that implies it is hashed is worse than sending nothing.
"""
if not value:
return ""
salt = _install_salt()
if not salt:
return ""
return hashlib.sha256(f"{salt}:{value}".encode("utf-8")).hexdigest()[:length]
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
def _safe_value(value: Any) -> Any:
@@ -302,25 +206,6 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
return created
def _rotate_anonymous_id(identity: dict[str, str]) -> str:
"""Mint a fresh anonymous id because the account context is gone.
The previous id may already have been merged into a person profile by an
$identify, and that merge is permanent. Reusing it after a logout or a key
change attributes everything that follows to the account that just went
away, which is the same misattribution the key fingerprint exists to stop,
only arriving through the anonymous path instead.
`aliased` is cleared with it: the new id has never been merged, so it is
eligible to be aliased into whatever account comes next.
"""
created = f"code-anon-{uuid.uuid4().hex}"
identity["anonymous_id"] = created
identity.pop("aliased", None)
_write_identity(identity)
return created
def _install_state_path() -> Path:
return memory_core.data_dir() / "install-state.json"
@@ -336,35 +221,18 @@ def is_first_run() -> bool:
return not _install_state_path().exists()
def data_dir_was_empty() -> bool:
"""Whether the data directory is untouched. Call BEFORE anything writes to it.
hook_runner reaches claim_install() only after cache_plugin_api_key() has
written `api-key` and EvidenceStore() has created `evidence.sqlite3`, so
asking at claim time always saw content and every fresh install reported an
upgrade. The caller snapshots this at the top of the run instead.
"""
return not _data_dir_has_content()
def claim_install(was_empty: bool | None = None) -> str | None:
def claim_install() -> str | None:
"""Claim the one install/upgrade record for this machine, atomically.
Returns the event to record ("install" or "upgrade"), or None if another
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
cannot both win.
`was_empty` must come from data_dir_was_empty() called before this process
wrote anything. Omitting it falls back to checking now, which is only
correct for a caller that has touched nothing.
"""
if not is_enabled():
# Never consume the one-shot claim while the user is opted out, or they
# would silently lose their install event if they later opt in.
return None
path = _install_state_path()
upgrading = not (data_dir_was_empty() if was_empty is None else was_empty)
# A fresh install has an empty data directory. Anything already there —
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
# before the marker is created, since creating it would itself be content.
upgrading = _data_dir_has_content()
try:
path.parent.mkdir(parents=True, exist_ok=True)
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
@@ -382,14 +250,6 @@ def claim_install(was_empty: bool | None = None) -> str | None:
},
stream,
)
# Durable before this returns. The O_EXCL open is what makes the
# claim exclusive, so it cannot be replaced by a temp-and-rename
# without losing that, which leaves the content as the thing to make
# safe. A kill between the open and this fsync used to leave a marker
# that exists but parses to nothing: is_first_run reads it as claimed
# and claim_version_change cannot read a version out of it.
stream.flush()
os.fsync(stream.fileno())
except OSError:
pass
return "upgrade" if upgrading else "install"
@@ -406,19 +266,6 @@ def _data_dir_has_content() -> bool:
return False
def _repair_install_state(path: Path) -> None:
"""Rewrite an unparseable marker so version tracking can resume."""
try:
temporary = path.with_suffix(f".{os.getpid()}.tmp")
temporary.write_text(
json.dumps({"plugin_version": memory_core.PLUGIN_VERSION, "repaired_at": memory_core.utc_now()}),
encoding="utf-8",
)
temporary.replace(path)
except OSError:
pass
def claim_version_change() -> str | None:
"""Return the previously recorded version if it differs, updating the marker.
@@ -429,47 +276,20 @@ def claim_version_change() -> str | None:
path = _install_state_path()
try:
state = json.loads(path.read_text(encoding="utf-8"))
except OSError:
except (OSError, json.JSONDecodeError):
return None
except json.JSONDecodeError:
# A crash between O_EXCL and the write leaves an empty marker. Left
# alone it disables every future upgrade event on this machine, because
# claim_install sees the file and this function cannot parse it.
state = None
if not isinstance(state, dict):
_repair_install_state(path)
return None
previous = str(state.get("plugin_version") or "")
if not previous or previous == memory_core.PLUGIN_VERSION:
return None
# Claim the transition with an exclusive sentinel before rewriting the
# marker. A plain read-modify-write let every concurrently starting session
# observe the old version and each record its own upgrade — and the first
# session after a version bump is exactly when several agent windows restart
# together.
sentinel = path.with_name(f"upgraded-{memory_core.PLUGIN_VERSION}")
try:
os.close(os.open(sentinel, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600))
except FileExistsError:
return None
except OSError:
return None
state["plugin_version"] = memory_core.PLUGIN_VERSION
state["upgraded_at"] = memory_core.utc_now()
temporary = path.with_suffix(f".{os.getpid()}.tmp")
try:
temporary = path.with_suffix(f".{os.getpid()}.tmp")
temporary.write_text(json.dumps(state), encoding="utf-8")
temporary.replace(path)
except OSError:
# Release the claim. The marker still records the old version, so
# without this the sentinel makes claim_version_change return early on
# every later run and this version's upgrade is never recorded again.
for leftover in (sentinel, temporary):
try:
leftover.unlink()
except OSError:
pass
return None
return previous
@@ -503,17 +323,10 @@ def record(
os=sys.platform,
python_version=platform.python_version(),
)
# Assigned only when the digest is real. _scoped_digest returns "" when
# the salt could not be persisted, and an empty property is worse than an
# absent one: it survives the None filter below and reads as a value.
if repo is not None:
repo_hash = _scoped_digest(getattr(repo, "identity", ""))
if repo_hash:
properties["repo_hash"] = repo_hash
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
if session_id:
session_hash = _scoped_digest(session_id)
if session_hash:
properties["session_hash"] = session_hash
properties["session_hash"] = _scoped_digest(session_id)
line = json.dumps(
{
"event": f"{EVENT_PREFIX}.{event}",
@@ -583,18 +396,9 @@ def _claim_name(attempt: int = 0) -> str:
def _claim_attempt(claim: Path) -> int:
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape.
Anchored on field position, not on a leading "a": the legacy shape is
``telemetry-<pid>-<hex>.sending`` and a hex id such as ``a1234567`` would
otherwise parse as attempt 1234567 and be discarded unsent on the first
flush after an upgrade.
"""
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
parts = stem.split("-")
if len(parts) != 4:
return 0
tail = parts[3]
tail = stem.rsplit("-", 1)[-1]
if tail.startswith("a") and tail[1:].isdigit():
return int(tail[1:])
return 0
@@ -628,42 +432,6 @@ def _claim_spool() -> Path | None:
return _claim_parked(directory)
def _sweep_debris(directory: Path) -> None:
"""Remove files nothing else will ever pick up again.
*.partial is a temp file orphaned by a crash between write and rename.
*.corrupt is a batch quarantined for undecodable content. No glob in this
module matches either, so without this they accumulate on disk for the life
of the install.
Quarantined batches are kept far longer than debris: they are the only
evidence left of events that could not be delivered, and someone diagnosing
a report of missing telemetry has to be able to find one.
"""
now = time.time()
for debris in directory.glob("telemetry-*.partial"):
try:
if now - debris.stat().st_mtime > CLAIM_STALE_SECONDS:
debris.unlink()
except OSError:
continue
for quarantined in directory.glob("telemetry-*.corrupt"):
try:
if now - quarantined.stat().st_mtime > CLAIM_EXPIRY_SECONDS:
quarantined.unlink()
except OSError:
continue
# The same reasoning covers *.tmp. _write_identity and _install_salt both
# create one and unlink it in a finally, which a SIGKILL skips, and no glob
# in this module matches the leftovers either.
for temporary in directory.glob("telemetry-*.tmp"):
try:
if now - temporary.stat().st_mtime > CLAIM_STALE_SECONDS:
temporary.unlink()
except OSError:
continue
def _claim_parked(directory: Path) -> Path | None:
"""Take the oldest abandoned claim, if any lease has actually expired.
@@ -679,24 +447,15 @@ def _claim_parked(directory: Path) -> Path | None:
age = now - orphan.stat().st_mtime
except OSError:
continue
if age < CLAIM_STALE_SECONDS:
# Someone else holds a live lease on it. This check has to come
# first. Claiming a file bumps its attempt count and refreshes its
# mtime, so a sender that has just taken the final attempt looks
# exhausted to everyone else while it is actively draining. Judging
# exhaustion before liveness let a second sender unlink a batch out
# from under its owner, losing every event in it.
continue
# Attempts, not age. Every re-claim touches the mtime and every release
# backdates it by a fixed amount, so age is pinned near the stale
# threshold and never reaches the expiry. Age stays only as a backstop
# for files that never carried an attempt marker.
if _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS or age > CLAIM_EXPIRY_SECONDS:
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
try:
orphan.unlink()
except OSError:
pass
continue
if age < CLAIM_STALE_SECONDS:
# Someone else holds a live lease on it.
continue
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
try:
orphan.replace(claim)
@@ -730,14 +489,10 @@ def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
return True
temporary = claim.with_suffix(f".{os.getpid()}.partial")
try:
payload = "".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining)
# fsync before the rename: without it the rename can land while the
# bytes have not, and the claim comes back empty or truncated after a
# crash. _drain then reads zero events and unlinks it.
with open(temporary, "w", encoding="utf-8") as handle:
handle.write(payload)
handle.flush()
os.fsync(handle.fileno())
temporary.write_text(
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
encoding="utf-8",
)
temporary.replace(claim)
_touch(claim)
return True
@@ -761,11 +516,7 @@ def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
if not _rewrite_claim(claim, remaining):
return
try:
# Backdate past the stale threshold so the next flush can pick it up,
# minus a cooldown that grows with the attempts already spent. Clamped so
# the mtime never lands in the future, which would read as a live lease.
cooldown = min(_claim_attempt(claim) * RETRY_COOLDOWN_SECONDS, CLAIM_STALE_SECONDS)
released = time.time() - CLAIM_STALE_SECONDS - 1 + cooldown
released = time.time() - CLAIM_STALE_SECONDS - 1
os.utime(claim, (released, released))
except OSError:
pass
@@ -812,56 +563,23 @@ def resolve_distinct_id() -> tuple[str, str]:
fingerprint = _digest(key) if key else ""
email = identity.get("email", "")
if email and fingerprint:
recorded = identity.get("key_fingerprint", "")
if recorded == fingerprint:
return email, ""
if not recorded:
# Rows written before fingerprints existed. Verify rather than
# adopt: a key changed before the upgrade would otherwise bind the
# new key to the previous account's email, permanently, and the
# fingerprint would then agree with itself forever after.
verified = _resolve_email(key)
if not verified:
# Offline, firewalled, or the API is down. Keep the previous
# behaviour and retry on the next flush rather than dropping a
# real account attribution. Safe because the same network that
# failed /v1/ping/ is about to fail the PostHog POST, so nothing
# is delivered under the unverified identity in the meantime.
return email, ""
identity["email"] = verified
identity["key_fingerprint"] = fingerprint
_write_identity(identity)
return verified, ""
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
return email, ""
if not key:
# No key to verify the account with; do not keep attributing to it.
if email:
identity.pop("email", None)
identity.pop("key_fingerprint", None)
return _rotate_anonymous_id(identity), ""
_write_identity(identity)
return anonymous_id(identity), ""
resolved = _resolve_email(key)
if not resolved:
# The key changed and will not resolve (revoked, offline, API down).
# Reaching here with an email means the recorded fingerprint disagreed,
# so the key really did change. Drop the account and rotate: the stored
# anonymous id may already be merged into that account's person, and
# reusing it would keep the events on the profile we are trying to
# leave.
if email:
identity.pop("email", None)
identity.pop("key_fingerprint", None)
return _rotate_anonymous_id(identity), ""
return anonymous_id(identity), ""
return (email, "") if email else (anonymous_id(identity), "")
# Alias only when going anonymous -> email for the first time. Once an anon
# id has been merged into an account it must never be offered again: an
# alias naming an already-identified id is what could link two real people.
previous = "" if (email or identity.get("aliased")) else identity.get("anonymous_id", "")
if previous:
identity["aliased"] = True
# Alias only when going anonymous -> email for the first time.
previous = "" if email else identity.get("anonymous_id", "")
identity["email"] = resolved
identity["key_fingerprint"] = fingerprint
_write_identity(identity)
@@ -881,7 +599,6 @@ def flush() -> int:
# Parked batches used to starve behind the live spool indefinitely. Bounded
# per run so a long backlog cannot turn one flush into an unbounded loop.
directory = memory_core.data_dir()
_sweep_debris(directory)
for _ in range(MAX_PARKED_PER_RUN):
parked = _claim_parked(directory)
if parked is None:
@@ -902,28 +619,8 @@ def _drain(claim: Path | None) -> tuple[int, bool]:
return 0, True
try:
lines = claim.read_text(encoding="utf-8").splitlines()
except ValueError:
# UnicodeDecodeError from a torn write: the content is unrecoverable, so
# quarantine rather than retry. flush() runs from a bare `finally:` in
# flush_worker, so raising here also skips the handoff cleanup, and an
# undecodable file would otherwise be re-read on every flush forever.
# Reported as delivered because there is nothing left to deliver and the
# rest of the run should continue.
try:
claim.replace(claim.with_suffix(".corrupt"))
except OSError:
try:
claim.unlink()
except OSError:
pass
return 0, True
except OSError:
# Could not read it, which is not the same as having nothing to send.
# The file is left exactly where it is: a vanished or briefly unreadable
# claim is retryable, and quarantining it here would discard events over
# a transient filesystem error. Reported as undelivered so the run stops
# instead of counting a batch nothing was posted from as delivered.
return 0, False
return 0, True
events = []
for line in lines:
try:
@@ -933,15 +630,8 @@ def _drain(claim: Path | None) -> tuple[int, bool]:
if isinstance(value, dict) and value.get("event"):
events.append(value)
if not events:
# Only delete when the file really is empty. A non-empty file that
# parses to nothing is a torn write, and its contents are the unsent
# remainder — deleting it is the data loss this PR exists to prevent.
try:
empty = claim.stat().st_size == 0
except OSError:
empty = True
try:
claim.replace(claim.with_suffix(".corrupt")) if not empty else claim.unlink()
claim.unlink()
except OSError:
pass
return 0, True
@@ -991,13 +681,8 @@ def _drain(claim: Path | None) -> tuple[int, bool]:
return sent, False
sent += len(chunk)
# Record progress and refresh the lease after each successful batch, so
# a crash repeats at most one batch instead of the entire file. If the
# rewrite fails the claim still holds delivered events, so stop rather
# than carry on as though progress were recorded — continuing is how the
# duplicate delivery this PR fixes would come back.
if not _rewrite_claim(claim, events[start + len(chunk) :]):
_release_claim(claim, events[start + len(chunk) :])
return sent, False
# a crash repeats at most one batch instead of the entire file.
_rewrite_claim(claim, events[start + len(chunk) :])
return sent, True
@@ -1,6 +1,6 @@
{
"id": "mem0",
"version": "0.3.2",
"version": "0.3.1",
"homepage": "https://docs.mem0.ai/integrations/claude-code",
"native": {
"pluginRoot": "${CLAUDE_PLUGIN_ROOT}",
@@ -7,8 +7,8 @@ disable-model-invocation: true
# Pause memory capture
To pause (hooks stop capturing and sending session content; a minimal
telemetry ping still fires at session start, under your Mem0 account email,
unless `MEM0_TELEMETRY=false`):
anonymous telemetry ping still fires at session start unless
`MEM0_TELEMETRY=false`):
```bash
python3 "${CLAUDE_PLUGIN_ROOT}/core/memory_cli.py" --harness "claude-code" --plugin-data-dir "${CLAUDE_PLUGIN_DATA}" pause
@@ -1,7 +1,6 @@
from __future__ import annotations
import json
import re
import os
import sqlite3
import subprocess
@@ -3461,12 +3460,7 @@ def test_automatic_flush_can_be_disabled_for_external_harnesses(isolated_env):
def test_version_is_single_sourced():
manifest = json.loads((PLUGIN_ROOT / ".claude-plugin" / "plugin.json").read_text())
assert manifest["name"] == "mem0"
# Compared against PLUGIN_VERSION, never a literal. A hardcoded version here
# was one more place to edit on every release, inside the test asserting the
# version is single-sourced, and it caught nothing that the agreement checks
# below do not: fifteen places set to the same wrong value would still pass.
assert re.fullmatch(r"\d+\.\d+\.\d+", memory_core.PLUGIN_VERSION), memory_core.PLUGIN_VERSION
assert manifest["version"] == memory_core.PLUGIN_VERSION
assert manifest["version"] == memory_core.PLUGIN_VERSION == "0.3.1"
root = REPOSITORY_ROOT
for mp in (root / "marketplace.json", root / ".claude-plugin" / "marketplace.json"):
entry = next(p for p in json.loads(mp.read_text())["plugins"] if p["name"] == "mem0")
@@ -265,109 +265,6 @@ def test_an_unresolvable_key_falls_back_to_the_anonymous_id(isolated_env, monkey
assert telemetry.resolve_distinct_id()[0].startswith("code-anon-")
def test_logging_out_does_not_leave_events_on_the_previous_account(isolated_env, monkeypatch):
"""Review finding: clearing the email kept an id already merged into a person.
The anonymous id is offered to PostHog as $anon_distinct_id on first sign-in,
and that merge is permanent. Keeping it after the key goes away means every
later anonymous event lands on the account that just left.
"""
# Run anonymously first, which is the only way an id exists to be merged.
merged = telemetry.anonymous_id()
monkeypatch.setenv("MEM0_API_KEY", "key-for-account-a")
with patch.object(telemetry, "_resolve_email", lambda key: "a@example.com"):
identified, alias = telemetry.resolve_distinct_id()
assert identified == "a@example.com"
assert alias == merged, "the anonymous id was merged into this account"
monkeypatch.delenv("MEM0_API_KEY", raising=False)
after_logout, logout_alias = telemetry.resolve_distinct_id()
assert after_logout.startswith("code-anon-")
assert after_logout != merged, "reused an id already merged into the previous account"
assert logout_alias == ""
assert "aliased" not in telemetry._read_identity(), "rotated id must be aliasable again"
def test_a_changed_key_that_will_not_resolve_rotates_the_anonymous_id(isolated_env, monkeypatch):
"""Same leak by the other route: fingerprint disagrees and the lookup fails."""
merged = telemetry.anonymous_id()
monkeypatch.setenv("MEM0_API_KEY", "key-for-account-a")
with patch.object(telemetry, "_resolve_email", lambda key: "a@example.com"):
telemetry.resolve_distinct_id()
monkeypatch.setenv("MEM0_API_KEY", "key-for-account-b")
with patch.object(telemetry, "_resolve_email", lambda key: ""):
after, alias = telemetry.resolve_distinct_id()
assert after.startswith("code-anon-")
assert after != merged
assert alias == ""
assert "email" not in telemetry._read_identity()
def test_a_legacy_cached_email_is_verified_before_the_key_is_bound(isolated_env, monkeypatch):
"""Review finding: a key changed before upgrading bound the wrong account.
Rows written before fingerprints existed carry an email and no fingerprint.
Adopting the current key without checking pinned that key to the previous
account's email, and every run after that agreed with itself.
"""
telemetry._write_identity({"email": "old@example.com", "anonymous_id": "code-anon-seed"})
monkeypatch.setenv("MEM0_API_KEY", "key-for-account-b")
with patch.object(telemetry, "_resolve_email", lambda key: "new@example.com"):
resolved, alias = telemetry.resolve_distinct_id()
assert resolved == "new@example.com"
assert alias == "", "email to email must never alias; it merges two real people"
stored = telemetry._read_identity()
assert stored["email"] == "new@example.com"
assert stored["key_fingerprint"] == telemetry._digest("key-for-account-b")
def test_a_legacy_row_keeps_working_when_the_account_cannot_be_checked(isolated_env, monkeypatch):
"""Firewalled users must not lose attribution, and must not bind unverified.
The same network that fails /v1/ping/ fails the PostHog POST, so nothing is
delivered under the unverified identity while this holds.
"""
telemetry._write_identity({"email": "old@example.com"})
monkeypatch.setenv("MEM0_API_KEY", "key-for-account-b")
with patch.object(telemetry, "_resolve_email", lambda key: ""):
resolved, _ = telemetry.resolve_distinct_id()
assert resolved == "old@example.com"
assert "key_fingerprint" not in telemetry._read_identity(), "bound an unverified key"
def test_a_failed_upgrade_claim_can_be_retried(isolated_env, monkeypatch):
"""Review finding: a failed rewrite left the sentinel and suppressed forever.
claim_version_change returns early on FileExistsError, and the marker still
holds the old version, so the upgrade for that version was never recorded
again on that machine.
"""
telemetry.claim_install()
state_path = memory_core.data_dir() / "install-state.json"
state = json.loads(state_path.read_text())
state["plugin_version"] = "0.0.1-old"
state_path.write_text(json.dumps(state), encoding="utf-8")
real_replace = Path.replace
def failing_replace(self, target):
raise OSError("disk full")
monkeypatch.setattr(Path, "replace", failing_replace)
assert telemetry.claim_version_change() is None
monkeypatch.setattr(Path, "replace", real_replace)
assert telemetry.claim_version_change() == "0.0.1-old", "sentinel suppressed the retry"
def test_first_run_is_not_flipped_by_writing_the_identity_file(isolated_env):
"""The identity file is written by a successful flush, not by recording.
@@ -418,114 +315,3 @@ def test_spawn_flush_does_nothing_without_a_spool(isolated_env):
with patch.object(telemetry.subprocess, "Popen") as popen:
assert telemetry.spawn_flush() is True
popen.assert_called_once()
def test_salt_is_stable_across_processes(isolated_env):
"""Hooks are separate short-lived processes; one repo must hash one way.
An unlocked read-modify-write let each process mint its own salt, so a
repository hashed several ways in the window before one writer won.
"""
import subprocess as sp
core = str(Path(__file__).resolve().parents[1] / "core")
script = (
f"import sys; sys.path.insert(0, {core!r})\n"
"import telemetry\n"
"print(telemetry._install_salt())"
)
env = {**os.environ, "MEM0_CODE_DATA_DIR": str(memory_core.data_dir())}
salts = {
sp.run([sys.executable, "-c", script], capture_output=True, text=True, env=env).stdout.strip()
for _ in range(4)
}
assert len(salts) == 1, f"one repo hashed {len(salts)} ways: {salts}"
def test_salt_does_not_touch_the_identity_file(isolated_env):
"""The identity file is is_first_run's marker and the sender's email store.
Writing the salt into it would create it from record(), suppressing the
install event, and would race resolve_distinct_id, which holds a stale copy
of that dict across a network call.
"""
telemetry._install_salt()
assert not telemetry._identity_path().exists()
def test_no_salt_means_no_hash_rather_than_an_unsalted_one(isolated_env, monkeypatch):
"""A read-only data dir drops the property; it must not emit a weak digest.
The previous fallback was a digest of the salt file's own path, which an
attacker can compute, memoized for the whole process. A property named
repo_hash carrying an effectively unsalted digest is worse than no property:
it reads as protected and is not.
"""
telemetry._salt_cache = ""
monkeypatch.setattr(telemetry.os, "open", lambda *a, **k: (_ for _ in ()).throw(OSError("read-only")))
assert telemetry._install_salt() == ""
assert telemetry._scoped_digest("git@github.com:acme/secret.git") == ""
def test_a_half_written_salt_is_never_visible_to_another_process(isolated_env, monkeypatch):
"""The window this closes: file created, value not yet written.
O_CREAT|O_EXCL then write leaves the name present and empty in between. A
hook reading it there used to get "", fall back to the path digest and cache
that for its whole run, so the same repo hashed two ways depending on timing.
Publishing by link means the name either does not exist or is complete.
"""
telemetry._salt_cache = ""
salt_path = telemetry._salt_path()
observed = []
real_link = telemetry.os.link
def observing_link(source, target):
# Stand where the racing reader stands: after the temp file is written,
# before the real name exists.
observed.append(salt_path.exists())
return real_link(source, target)
monkeypatch.setattr(telemetry.os, "link", observing_link)
salt = telemetry._install_salt()
assert observed == [False], "the salt name existed before it held a value"
assert len(salt) == 32
assert salt_path.read_text(encoding="utf-8").strip() == salt
def test_a_filesystem_without_hardlinks_still_gets_a_salt(isolated_env, monkeypatch):
"""Publishing by link must not become a silent loss of the hashes.
Some network mounts and container volumes reject os.link. Returning ""
there would drop repo_hash and session_hash on every run for that whole
cohort, which is a bigger loss than the narrow race the link closes.
"""
telemetry._salt_cache = ""
monkeypatch.setattr(
telemetry.os, "link", lambda src, dst: (_ for _ in ()).throw(OSError(38, "not implemented"))
)
salt = telemetry._install_salt()
assert len(salt) == 32, "no salt on a filesystem without hardlinks"
assert telemetry._salt_path().read_text(encoding="utf-8").strip() == salt
assert telemetry._scoped_digest("git@github.com:acme/x.git") != ""
assert not list(telemetry._salt_path().parent.glob("telemetry-salt.*.tmp"))
def test_a_concurrent_writer_does_not_clobber_the_published_salt(isolated_env):
"""Second process to finish must adopt the first one's salt, not replace it.
os.link rather than os.replace is what makes losing the race harmless.
"""
telemetry._salt_cache = ""
first = telemetry._install_salt()
telemetry._salt_cache = ""
second = telemetry._install_salt()
assert second == first
assert not list(telemetry._salt_path().parent.glob("telemetry-salt.*.tmp")), "temp file left behind"
@@ -1,6 +1,6 @@
{
"name": "mem0",
"version": "0.3.2",
"version": "0.3.1",
"description": "Cross-session memory and token savings for coding agents.",
"author": { "name": "Mem0", "email": "support@mem0.ai" },
"homepage": "https://docs.mem0.ai/integrations/codex",
@@ -5,7 +5,5 @@ SOURCE_TAG = "CODEX_PLUGIN"
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
# plugin family is one source; which editor it runs in is the application.
# An empty application means the host is unknown, and memory_core omits
# the header entirely rather than sending a placeholder.
PLATFORM_SOURCE = "MEM0_PLUGIN"
PLATFORM_APPLICATION = "codex"
@@ -290,11 +290,6 @@ def run(
if args.plugin_data_dir:
os.environ[data_dir_env] = args.plugin_data_dir
# Snapshot BEFORE anything writes to the data dir: cache_plugin_api_key
# writes `api-key` and EvidenceStore creates `evidence.sqlite3`, so asking
# after them always saw content and every fresh install reported an upgrade.
data_dir_was_empty = telemetry.data_dir_was_empty()
cache_plugin_api_key()
if args.action == "session-start":
clear_stale_api_key_cache()
@@ -312,7 +307,7 @@ def run(
if args.action == "session-start":
# Claims the marker atomically and says which event to record, so a
# second session starting alongside this one cannot record it too.
first_event = telemetry.claim_install(was_empty=data_dir_was_empty)
first_event = telemetry.claim_install()
if first_event == "install":
telemetry.record("install")
elif first_event == "upgrade":
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.2"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -2008,12 +2008,9 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
# Top level, not metadata: the backend reads `source` from the body,
# query string or X-Mem0-Source header, never from metadata. The
# harness tag stays in metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,

Some files were not shown because too many files have changed in this diff Show More