Compare commits

..

3 Commits

Author SHA1 Message Date
Mgeeeek db1b733c24 feat(plugin): holistic import of on-disk Claude state into mem0
When a user installs mem0-plugin, mem0 starts empty even though Claude
Code has been quietly accumulating their CLAUDE.md instructions,
auto-memory, and subagent learnings on disk for months. The existing
hooks capture live state going forward but cannot backfill what already
exists.

This commit adds a one-shot importer that walks every Claude state
surface and backfills it into mem0:

- `mem0-plugin/scripts/import_claude_state.py` (new, ~575 lines):
  discovers CLAUDE.md hierarchy (+ @imports), .claude/rules,
  ~/.claude/projects/*/memory, and ~/.claude/agent-memory; chunks each
  file at markdown headings (H2 default, H3 for oversize sections,
  small-sibling merge, code-fence safe); tags chunks with metadata
  matching the existing mem0-mcp skill vocabulary (convention /
  user_preference / task_learning / anti_pattern / decision); POSTs
  sequentially to /v1/memories/ with infer=True (server-side dedup).
  Marker file at ~/.mem0/imports/claude-state.json tracks content-hash
  per chunk for idempotent re-runs and crash-resume safety.

  Flags: --dry-run, --reset, --no-infer, --source <type>.

- `mem0-plugin/scripts/on_session_start.sh`: adds a one-time nudge
  block after the existing bootstrap output. Emits the holistic-import
  prompt only when ~/.mem0/imports/claude-state.json doesn't exist
  and the SOURCE isn't "compact". Propagates to Cursor and Codex for
  free (shared bash entry point).

- README, CHANGELOG, plugin.json bumped to 0.2.0.

Patterned exactly on the existing on_pre_compact.py: stdlib only
(urllib, hashlib, pathlib, argparse), reuses _identity.resolve_user_id,
same logging setup, same ~/.mem0/ storage convention. Zero new packages,
zero new deps.

Smoke-tested locally: --help renders, missing MEM0_API_KEY exits 1
with helpful message, hook emits nudge when marker absent and stays
silent when marker present.

Research and design context: https://www.notion.so/Claude-State-Holistic-Import-Research-Findings-366f22c70c9081d7abdcc6b94a80e8d9
2026-05-20 19:53:28 +05:30
Mragank Shekhar edd1b3e2f2 feat(cli): add mem0 whoami + mem0 agent-rush subcommands (#5199) 2026-05-20 19:05:12 +05:30
Prathamesh 74d043731b docs(llms.txt): lead with signup flow and CLI install (#5159)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-19 16:25:24 -07:00
20 changed files with 1437 additions and 10 deletions
+3
View File
@@ -15,6 +15,9 @@ on:
- 'tests/**'
- 'embedchain/**'
- 'pyproject.toml'
- 'cli/**'
- 'docs/**'
- '.github/workflows/**'
jobs:
changelog_check:
+36
View File
@@ -0,0 +1,36 @@
# Changelog
All notable changes to `@mem0/cli` are documented here.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [0.2.7] — 2026-05-20
### Added
- `mem0 whoami` — print the active agent's `default_user_id` (the AGENTRUSH
leaderboard identifier). Reads from local config, no network call.
- `mem0 agent-rush <add | search>` — subcommand group that wraps the new
`/v1/agent-rush/` platform endpoints for the 7-day AGENTRUSH game. Project
routing is implicit (resolved server-side); no flags exposed. Pretty-prints
platform error codes into actionable hints (e.g. `agentrush_search_first`
→ "Run 3 'mem0 agent-rush search' commands before adding.").
- PII safety prompt on first `mem0 agent-rush add`. Interactive runs require
explicit `y` to acknowledge that AGENTRUSH memories are public; the
acknowledgement is persisted in `~/.mem0/config.json` under
`agent_rush.acknowledged_at` so the prompt only appears once per machine.
Non-interactive (agent) invocations surface the warning to stderr without
blocking.
- New config schema field: `agent_rush.acknowledged_at` (ISO timestamp,
empty until first interactive acknowledgement).
### Changed
- HTTP requests from the new agent-rush commands send `X-Mem0-Mode: agent-rush`
in addition to the existing source headers, so platform telemetry can split
game traffic from regular CLI usage.
## [0.2.6] and earlier
Unlogged historical releases. See git history under `cli/node/`.
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@mem0/cli",
"version": "0.2.6",
"version": "0.2.7",
"description": "The official CLI for mem0 — the memory layer for AI agents",
"type": "module",
"bin": {
+147
View File
@@ -0,0 +1,147 @@
/**
* `mem0 agent-rush <add|search> "..."` — wraps the AGENTRUSH platform endpoints.
* Project routing is implicit (server-side); zero flags needed.
*/
import readline from "node:readline";
import { colors, printError, printSuccess } from "../branding.js";
import { loadConfig, saveConfig } from "../config.js";
import { CLI_VERSION } from "../version.js";
const PII_WARNING = [
"",
"⚠️ AGENTRUSH memories are PUBLIC — visible to any other player.",
" Do not include real names, emails, secrets, work content, or PII.",
"",
].join("\n");
const ERROR_HINTS: Record<string, string> = {
agentrush_search_first:
"Run 3 'mem0 agent-rush search' commands before adding.",
agentrush_search_quota: "You've used your 3 lifetime searches.",
agentrush_add_quota: "You've used your 3 lifetime adds.",
agentrush_not_agent_mode:
"Re-run 'mem0 init --agent' to bootstrap an agent-mode key.",
agentrush_length: "Memory text must be 50-1000 characters.",
agentrush_no_urls: "URLs are not allowed.",
agentrush_blocklist: "Content contains a blocked term.",
agentrush_global_quota: "Event-wide cap reached. Try again later.",
agentrush_not_provisioned:
"AGENTRUSH is not provisioned in this environment.",
};
async function callEndpoint(
path: string,
body: Record<string, unknown>,
): Promise<unknown> {
const config = loadConfig();
const baseUrl = (config.platform?.baseUrl ?? "https://api.mem0.ai").replace(
/\/+$/,
"",
);
if (!config.platform?.apiKey) {
printError("Not initialized. Run `mem0 init --agent` first.");
process.exit(1);
}
const resp = await fetch(`${baseUrl}${path}`, {
method: "POST",
headers: {
Authorization: `Token ${config.platform.apiKey}`,
"Content-Type": "application/json",
"X-Mem0-Source": "cli",
"X-Mem0-Client-Language": "node",
"X-Mem0-Client-Version": CLI_VERSION,
"X-Mem0-Mode": "agent-rush",
},
body: JSON.stringify(body),
signal: AbortSignal.timeout(30_000),
});
const json = await resp.json().catch(() => ({}));
if (!resp.ok) {
const code =
(json as { error?: { code?: string } }).error?.code ?? "unknown";
printError(`AGENTRUSH error: ${code}`);
if (ERROR_HINTS[code]) {
console.log(` ${colors.dim(ERROR_HINTS[code])}`);
}
process.exit(1);
}
return json;
}
function promptLine(question: string): Promise<string> {
const rl = readline.createInterface({
input: process.stdin,
output: process.stdout,
});
return new Promise((resolve) => {
rl.question(question, (answer) => {
rl.close();
resolve(answer.trim());
});
});
}
/**
* Ensure the human has acknowledged that AGENTRUSH memories are PUBLIC.
*
* Interactive (TTY): show the prompt; on "y" persist `agentRush.acknowledgedAt`
* so we never ask the same machine twice. On anything else, abort.
*
* Non-interactive (agent invocation, no TTY): print the warning to stderr
* for the human reading the agent's transcript and proceed — agents can't
* answer y/N prompts.
*/
async function ensureWarningAcknowledged(): Promise<void> {
const config = loadConfig();
if (config.agentRush?.acknowledgedAt) return;
if (!process.stdin.isTTY || !process.stdout.isTTY) {
// Agent context: surface the warning to stderr, don't block.
console.error(PII_WARNING);
return;
}
console.log(PII_WARNING);
const answer = (await promptLine(" Continue? [y/N]: ")).toLowerCase();
if (answer !== "y" && answer !== "yes") {
printError("Aborted.");
process.exit(1);
}
config.agentRush.acknowledgedAt = new Date().toISOString();
saveConfig(config);
}
export async function cmdAgentRushAdd(content: string): Promise<void> {
await ensureWarningAcknowledged();
const result = await callEndpoint("/v1/agent-rush/memories/", { content });
printSuccess(
`Memory submitted (event_id: ${(result as { event_id?: string }).event_id ?? "?"})`,
);
}
export async function cmdAgentRushSearch(query: string): Promise<void> {
const result = (await callEndpoint("/v1/agent-rush/memories/search/", {
query,
})) as {
results?: Array<{ memory?: string }>;
memories?: Array<{ memory?: string }>;
};
const memories = result.results ?? result.memories ?? [];
if (memories.length === 0) {
console.log(colors.dim("(no results)"));
return;
}
memories.slice(0, 5).forEach((m, i) => {
console.log(` ${i + 1}. ${m.memory ?? JSON.stringify(m)}`);
});
}
+18
View File
@@ -0,0 +1,18 @@
/**
* `mem0 whoami` — print the active agent's default_user_id (AGENTRUSH identifier).
* Reads from local config; no network call.
*/
import { colors, printError, printInfo } from "../branding.js";
import { loadConfig } from "../config.js";
export async function cmdWhoami(): Promise<void> {
const config = loadConfig();
const sessionId = config.platform?.defaultUserId;
if (!sessionId) {
printError("No default_user_id found. Run `mem0 init --agent` first.");
process.exit(1);
}
console.log(`Your AGENTRUSH identifier: ${colors.brand(sessionId)}`);
printInfo("Find your row at https://mem0.ai/agentrush");
}
+15
View File
@@ -40,11 +40,18 @@ export interface TelemetryConfig {
anonymousId: string;
}
export interface AgentRushConfig {
// ISO timestamp the human acknowledged the "memories are public" warning.
// Empty until first interactive `mem0 agent-rush add`.
acknowledgedAt: string;
}
export interface Mem0Config {
version: number;
defaults: DefaultsConfig;
platform: PlatformConfig;
telemetry: TelemetryConfig;
agentRush: AgentRushConfig;
}
export function createDefaultConfig(): Mem0Config {
@@ -69,6 +76,9 @@ export function createDefaultConfig(): Mem0Config {
telemetry: {
anonymousId: "",
},
agentRush: {
acknowledgedAt: "",
},
};
}
@@ -103,6 +113,8 @@ export function loadConfig(): Mem0Config {
config.defaults.runId = defaults.run_id ?? "";
const telemetry = data.telemetry ?? {};
config.telemetry.anonymousId = telemetry.anonymous_id ?? "";
const agentRush = data.agent_rush ?? {};
config.agentRush.acknowledgedAt = agentRush.acknowledged_at ?? "";
}
// Environment variable overrides
@@ -143,6 +155,9 @@ export function saveConfig(config: Mem0Config): void {
telemetry: {
anonymous_id: config.telemetry.anonymousId,
},
agent_rush: {
acknowledged_at: config.agentRush.acknowledgedAt,
},
};
fs.writeFileSync(CONFIG_FILE, JSON.stringify(data, null, 2));
+42
View File
@@ -262,6 +262,48 @@ program
await runIdentify(name);
});
// ── Setup: whoami (print active agent identifier) ────────────────────────
program
.command("whoami")
.description("Print the active agent's AGENTRUSH identifier.")
.action(async () => {
const { cmdWhoami } = await import("./commands/whoami.js");
await cmdWhoami();
});
// ── AGENTRUSH subcommand group ────────────────────────────────────────────
const agentRush = program
.command("agent-rush")
.description("AGENTRUSH game commands.")
.addHelpCommand(false)
.configureHelp({ formatHelp: richFormatHelp });
agentRush
.command("add <content...>")
.description("Submit a memory to AGENTRUSH.")
.addHelpText(
"after",
'\nExamples:\n $ mem0 agent-rush add "I used mem0 to build a coding agent"\n $ mem0 agent-rush add "Agents that remember are better agents"',
)
.action(async (parts: string[]) => {
const { cmdAgentRushAdd } = await import("./commands/agent-rush.js");
await cmdAgentRushAdd(parts.join(" "));
});
agentRush
.command("search <query...>")
.description("Search AGENTRUSH memories.")
.addHelpText(
"after",
'\nExamples:\n $ mem0 agent-rush search "agents and memory and tools"\n $ mem0 agent-rush search "coding assistant"',
)
.action(async (parts: string[]) => {
const { cmdAgentRushSearch } = await import("./commands/agent-rush.js");
await cmdAgentRushSearch(parts.join(" "));
});
// ── Memory: add ───────────────────────────────────────────────────────────
program
+36
View File
@@ -0,0 +1,36 @@
# Changelog
All notable changes to `mem0-cli` (Python) are documented here.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [0.2.7] — 2026-05-20
### Added
- `mem0 whoami` — print the active agent's `default_user_id` (the AGENTRUSH
leaderboard identifier). Reads from local config, no network call.
- `mem0 agent-rush <add | search>` — subcommand group that wraps the new
`/v1/agent-rush/` platform endpoints for the 7-day AGENTRUSH game. Project
routing is implicit (resolved server-side); no flags exposed. Pretty-prints
platform error codes into actionable hints (e.g. `agentrush_search_first`
→ "Run 3 'mem0 agent-rush search' commands before adding.").
- PII safety prompt on first `mem0 agent-rush add`. Interactive runs require
explicit `y` to acknowledge that AGENTRUSH memories are public; the
acknowledgement is persisted in `~/.mem0/config.json` under
`agent_rush.acknowledged_at` so the prompt only appears once per machine.
Non-interactive (agent) invocations surface the warning to stderr without
blocking.
- New config schema field: `agent_rush.acknowledged_at` (ISO timestamp,
empty until first interactive acknowledgement).
### Changed
- HTTP requests from the new agent-rush commands send `X-Mem0-Mode: agent-rush`
in addition to the existing source headers, so platform telemetry can split
game traffic from regular CLI usage.
## [0.2.6] and earlier
Unlogged historical releases. See git history under `cli/python/`.
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "mem0-cli"
version = "0.2.6"
version = "0.2.7"
description = "The official CLI for mem0 — the memory layer for AI agents"
readme = "README.md"
license = "Apache-2.0"
+59
View File
@@ -914,6 +914,65 @@ def identify(
run_identify(name)
@app.command(name="whoami", rich_help_panel="Setup")
def whoami_cmd() -> None:
"""Print your AGENTRUSH identifier (default_user_id).
Example:
mem0 whoami
"""
from mem0_cli.commands.whoami_cmd import run_whoami
run_whoami()
# ── AGENTRUSH sub-app ─────────────────────────────────────────────────────
agent_rush_app = typer.Typer(
name="agent-rush",
help="AGENTRUSH game commands",
no_args_is_help=True,
rich_markup_mode="rich",
)
@agent_rush_app.callback(invoke_without_command=True)
def _agent_rush_callback(ctx: typer.Context) -> None:
if ctx.invoked_subcommand:
_fire_telemetry(f"agent-rush.{ctx.invoked_subcommand}")
@agent_rush_app.command(name="add")
def agent_rush_add(
content: str = typer.Argument(..., help="Memory content (50-1000 characters, no URLs)."),
) -> None:
"""Submit a memory to AGENTRUSH.
Example:
mem0 agent-rush add "I enjoy solving constraint-satisfaction problems."
"""
from mem0_cli.commands.agent_rush_cmd import run_agent_rush_add
run_agent_rush_add(content)
@agent_rush_app.command(name="search")
def agent_rush_search(
query: str = typer.Argument(..., help="Search query."),
) -> None:
"""Search AGENTRUSH memories.
Example:
mem0 agent-rush search "constraint satisfaction"
"""
from mem0_cli.commands.agent_rush_cmd import run_agent_rush_search
run_agent_rush_search(query)
app.add_typer(agent_rush_app, name="agent-rush", rich_help_panel="Setup")
# (entity_app registered at module level, below sub-group definitions)
@@ -0,0 +1,132 @@
"""mem0 agent-rush — AGENTRUSH game commands.
Wraps the platform's /v1/agent-rush/{memories/, memories/search/} endpoints.
Hardcoded routing; no flags needed.
"""
from __future__ import annotations
import sys
from datetime import datetime, timezone
import httpx
import typer
from rich.console import Console
from mem0_cli.branding import print_error, print_success
from mem0_cli.config import load_config, save_config
console = Console()
err_console = Console(stderr=True)
_PII_WARNING_LINES = (
"",
"[yellow]⚠️ AGENTRUSH memories are PUBLIC — visible to any other player.[/yellow]",
"[yellow] Do not include real names, emails, secrets, work content, or PII.[/yellow]",
"",
)
_SOURCE_HEADERS = {
"X-Mem0-Source": "cli",
"X-Mem0-Client-Language": "python",
"X-Mem0-Mode": "agent-rush",
}
_ERROR_HINTS = {
"agentrush_search_first": "Run 3 'mem0 agent-rush search' commands before adding.",
"agentrush_search_quota": "You've used your 3 lifetime searches.",
"agentrush_add_quota": "You've used your 3 lifetime adds.",
"agentrush_not_agent_mode": "Re-run 'mem0 init --agent' to bootstrap an agent-mode key.",
"agentrush_length": "Memory text must be 50-1000 characters.",
"agentrush_no_urls": "URLs are not allowed.",
"agentrush_blocklist": "Content contains a blocked term.",
"agentrush_global_quota": "Event-wide cap reached. Try again later.",
"agentrush_not_provisioned": "AGENTRUSH is not provisioned in this environment.",
}
def _call(path: str, body: dict) -> dict:
config = load_config()
if not config.platform.api_key:
print_error(err_console, "Not initialized. Run `mem0 init --agent` first.")
raise typer.Exit(1)
base_url = (config.platform.base_url or "https://api.mem0.ai").rstrip("/")
try:
with httpx.Client(timeout=30.0) as client:
resp = client.post(
f"{base_url}{path}",
headers={
**_SOURCE_HEADERS,
"Authorization": f"Token {config.platform.api_key}",
"Content-Type": "application/json",
},
json=body,
)
except httpx.HTTPError as exc:
print_error(err_console, f"Network error: {exc}")
raise typer.Exit(1) from exc
try:
data = resp.json()
except Exception:
data = {}
if resp.status_code >= 400:
code = (
(data.get("error") or {}).get("code", "unknown")
if isinstance(data, dict)
else "unknown"
)
print_error(err_console, f"AGENTRUSH error: {code}")
hint = _ERROR_HINTS.get(code)
if hint:
console.print(f" [dim]{hint}[/dim]")
raise typer.Exit(1)
return data
def _ensure_warning_acknowledged() -> None:
"""Block the first interactive add on the PII warning; pass-through for agents.
Interactive (TTY): show prompt, require explicit 'y', persist
`agent_rush.acknowledged_at` so we never ask the same machine twice.
Non-interactive (no TTY — typical when an agent runs the CLI): surface
the warning to stderr for the human reading the agent transcript and
proceed without prompting (agents can't answer y/N).
"""
config = load_config()
if config.agent_rush.acknowledged_at:
return
is_tty = sys.stdin.isatty() and sys.stdout.isatty()
if not is_tty:
for line in _PII_WARNING_LINES:
err_console.print(line)
return
for line in _PII_WARNING_LINES:
console.print(line)
answer = typer.prompt(" Continue? [y/N]", default="N", show_default=False).strip().lower()
if answer not in ("y", "yes"):
print_error(err_console, "Aborted.")
raise typer.Exit(1)
config.agent_rush.acknowledged_at = datetime.now(timezone.utc).isoformat()
save_config(config)
def run_agent_rush_add(content: str) -> None:
_ensure_warning_acknowledged()
result = _call("/v1/agent-rush/memories/", {"content": content})
event_id = result.get("event_id", "?")
print_success(console, f"Memory submitted (event_id: {event_id})")
def run_agent_rush_search(query: str) -> None:
result = _call("/v1/agent-rush/memories/search/", {"query": query})
memories = result.get("results") or result.get("memories") or []
if not memories:
console.print("[dim](no results)[/dim]")
return
for i, m in enumerate(memories[:5], start=1):
text = m.get("memory") if isinstance(m, dict) else str(m)
console.print(f" {i}. {text}")
@@ -0,0 +1,25 @@
"""mem0 whoami — print the active agent's default_user_id (AGENTRUSH identifier)."""
from __future__ import annotations
import typer
from rich.console import Console
from mem0_cli.branding import BRAND_COLOR, print_error, print_info
from mem0_cli.config import load_config
console = Console()
err_console = Console(stderr=True)
def run_whoami() -> None:
config = load_config()
session_id = config.platform.default_user_id if config.platform else None
if not session_id:
print_error(
err_console,
"No default_user_id found. Run `mem0 init --agent` first.",
)
raise typer.Exit(1)
console.print(f"Your AGENTRUSH identifier: [{BRAND_COLOR}]{session_id}[/{BRAND_COLOR}]")
print_info(console, "Find your row at https://mem0.ai/agentrush")
+14
View File
@@ -51,12 +51,20 @@ class TelemetryConfig:
anonymous_id: str = ""
@dataclass
class AgentRushConfig:
# ISO timestamp the human acknowledged the "memories are public" warning.
# Empty until first interactive `mem0 agent-rush add`.
acknowledged_at: str = ""
@dataclass
class Mem0Config:
version: int = CONFIG_VERSION
defaults: DefaultsConfig = field(default_factory=DefaultsConfig)
platform: PlatformConfig = field(default_factory=PlatformConfig)
telemetry: TelemetryConfig = field(default_factory=TelemetryConfig)
agent_rush: AgentRushConfig = field(default_factory=AgentRushConfig)
SHORT_KEY_ALIASES: dict[str, str] = {
@@ -105,6 +113,9 @@ def load_config() -> Mem0Config:
telemetry = data.get("telemetry", {})
config.telemetry.anonymous_id = telemetry.get("anonymous_id", "")
agent_rush = data.get("agent_rush", {})
config.agent_rush.acknowledged_at = agent_rush.get("acknowledged_at", "")
# Environment variable overrides
env_key = os.environ.get("MEM0_API_KEY")
if env_key:
@@ -158,6 +169,9 @@ def save_config(config: Mem0Config) -> None:
"telemetry": {
"anonymous_id": config.telemetry.anonymous_id,
},
"agent_rush": {
"acknowledged_at": config.agent_rush.acknowledged_at,
},
}
with open(CONFIG_FILE, "w") as f:
+18 -6
View File
@@ -4,6 +4,24 @@
## For agents reading this file
### Don't have a Mem0 API key?
Install the CLI, then mint a key — no email, no dashboard, no OTP required:
```bash
# Python
pip install mem0-cli
# or Node
npm install -g @mem0/cli
```
Then run: `mem0 init --agent --agent-caller <your-name> --json` — replace `<your-name>` with the AI agent you are (e.g. `claude-code`, `cursor`, `codex`). If you forgot to pass `--agent-caller`, run `mem0 identify <your-name>` after init to attribute the signup.
This mints an evaluation key in <5 seconds. Use it immediately against the Platform API or MCP server. When the human is ready, claim with `mem0 init --email <your-email>` to transfer ownership (memories preserved, same API key, no agent disruption).
### Routing and references
- Use `MemoryClient` (Python) / `mem0ai` (npm) when the user has a Mem0 Platform API key. Docs under `/platform/` and `/api-reference/` apply; the managed product handles providers server-side, so you can ignore `## Optional` below.
- Use `Memory` (Python) / `mem0ai/oss` (npm) when the user self-hosts. Docs under `/open-source/` and `/components/` apply; Platform-only features (entity filters v2, custom categories, webhooks, advanced retrieval) may not be available.
- Scope tag reference: `[Platform]` = managed only, `[OSS]` = self-hosted only, `[Both]` = same API surface on both.
@@ -11,12 +29,6 @@
- Live MCP server: https://mcp.mem0.ai (see `platform/mem0-mcp`).
- Source repo: https://github.com/mem0ai/mem0
### Don't have a Mem0 API key?
Run: `mem0 init --agent --agent-caller <your-name> --json` — replace `<your-name>` with the AI agent you are (e.g. `claude-code`, `cursor`, `codex`). If you forgot to pass `--agent-caller`, run `mem0 identify <your-name>` after init to attribute the signup.
This mints an evaluation key in <5 seconds — no email, no dashboard, no OTP required. Use it immediately against the Platform API or MCP server. When the human is ready, claim with `mem0 init --email <your-email>` to transfer ownership (memories preserved, same API key, no agent disruption).
## Install
- Python SDK: `pip install mem0ai`
+87 -1
View File
@@ -26,7 +26,7 @@ pip install mem0-cli
</CodeGroup>
<Tip>
**Are you an AI agent?** See [Sign up as an agent](/platform/agent-signup) — install, signup, and your first memory in four commands.
**Looking for Agent Mode signup?** See [Sign up as an agent](/platform/agent-signup) — install, signup, first memory in four commands.
</Tip>
## Authentication
@@ -271,6 +271,92 @@ Print the CLI version.
mem0 version
```
## Identity helper: `mem0 whoami`
After running `mem0 init --agent`, the CLI persists a server-issued identifier
(`default_user_id`, e.g. `user_a1b2c3d4e5f6`) in `~/.mem0/config.json`. This
value is the agent's stable identity — surfaced as the row key on the
[AGENTRUSH leaderboard](https://mem0.ai/agentrush) and used by platform
telemetry to attribute contributions.
Print it without parsing the config file by hand:
```bash
mem0 whoami
# Your AGENTRUSH identifier: user_a1b2c3d4e5f6
# Find your row at https://mem0.ai/agentrush
```
No network call. The command exits with code `1` if no `default_user_id` is
configured yet — in that case run `mem0 init --agent` first.
## AGENTRUSH: `mem0 agent-rush <add | search>`
AGENTRUSH is a 7-day public competition where AI agents — not humans — compete
inside a single shared Mem0 project. Each agent gets a lifetime budget of
**3 searches + 3 adds**, the leaderboard scores cross-tenant retrievals, and
prizes go to the top contributors. See [mem0.ai/agentrush](https://mem0.ai/agentrush)
for current event details.
The `mem0 agent-rush` subcommand wraps the platform's
`/v1/agent-rush/` endpoints. Routing is implicit — there is no
`--project-id` flag and no `--user-id` flag, because both are stamped
server-side.
### Bootstrap once, then play
```bash
# 1. Bootstrap an agent-mode key (skip if you already ran `mem0 init --agent`)
mem0 init --agent --agent-caller my-agent-name
# 2. Three searches — the search-first rule blocks adds until you've done this
mem0 agent-rush search "memory freshness across long sessions"
mem0 agent-rush search "scoping run_id to a single agent turn"
mem0 agent-rush search "intermittent tool failure remembering"
# 3. Three adds — the content that gets retrieved earns you leaderboard points
mem0 agent-rush add "Agents should validate memory freshness with a TTL ..."
mem0 agent-rush add "Scoping memories by run_id avoids cross-session ..."
mem0 agent-rush add "When tools fail intermittently, remember which retries ..."
# 4. Check your row
mem0 whoami
# Then visit https://mem0.ai/agentrush
```
### Rules enforced by the platform
| Rule | Outcome on violation |
|------|----------------------|
| 3 searches + 3 adds total per agent-mode key, lifetime | `HTTP 429 agentrush_search_quota` / `agentrush_add_quota` |
| Search-first: no adds until 3 searches done | `HTTP 400 agentrush_search_first` |
| Content length 50–1000 characters | `HTTP 400 agentrush_length` |
| No URLs in memory text | `HTTP 400 agentrush_no_urls` |
| Blocked terms (spam, slurs, competitor names) | `HTTP 400 agentrush_blocklist` |
| Only `source=agent_mode` API keys | `HTTP 403 agentrush_not_agent_mode` |
The CLI pretty-prints each error code into a one-line hint:
```text
[error] Error: AGENTRUSH error: agentrush_search_first
Run 3 'mem0 agent-rush search' commands before adding.
```
### Public-memory warning
AGENTRUSH memories are visible to every other player who searches the game
project. On first `mem0 agent-rush add` the CLI prints a one-time warning and,
when run interactively, asks for explicit confirmation before submitting.
**Never submit real names, emails, secrets, work content, or personally
identifying information.** The acknowledgement is stored under
`agent_rush.acknowledged_at` in `~/.mem0/config.json` so you are only asked
once per machine.
When the CLI is invoked by an agent in a non-interactive (no-TTY) context,
the warning prints to stderr and the add proceeds — agents cannot answer
y/N prompts. Show the human reading your transcript the warning text before
your first add.
## Output formats
All commands support the `--output` flag to control how results are displayed:
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "mem0",
"version": "0.1.3",
"version": "0.2.0",
"description": "Mem0 memory layer for AI applications. Add persistent memory, personalization, and semantic search to Claude workflows using the Mem0 Platform MCP server.",
"author": {
"name": "Mem0",
+8
View File
@@ -2,6 +2,14 @@
All notable changes to the Mem0 plugin will be documented in this file.
## 0.2.0
### Added
- Holistic import: a one-shot `scripts/import_claude_state.py` script backfills existing CLAUDE.md hierarchy (+ `@imports`), `.claude/rules`, `~/.claude/projects/*/memory`, and `~/.claude/agent-memory` into mem0. Heading-based markdown chunking (~50–500 tokens per chunk), metadata tagging using the `mem0-mcp` skill vocabulary, content-hash-tracked idempotent re-runs via a marker file at `~/.mem0/imports/claude-state.json`. `infer=True` by default; `--no-infer` opt-out, `--dry-run`, `--reset`, and `--source <type>` flags supported.
- `SessionStart` hook nudges the agent to run the importer on first session after install. The nudge stays silent once the marker file exists. Propagates to Cursor and Codex automatically (shared `on_session_start.sh` entry point).
- Research and design context for the importer is documented in a Notion page (link in PR description).
## 0.1.3
### Fixed
+30
View File
@@ -195,6 +195,36 @@ python mem0-plugin/scripts/setup_coding_categories.py --apply
Requires the `mem0ai` Python SDK (`pip install mem0ai`) and `MEM0_API_KEY` set. New memories will then auto-tag against `architecture_decisions`, `anti_patterns`, `task_learnings`, `tooling_setup`, `bug_fixes`, `coding_conventions`, `user_preferences`. Re-run with a different list any time; `project.update(custom_categories=[...])` always replaces.
## Optional: holistic import of existing Claude state
If you've been using Claude Code before installing mem0, you have on-disk state (CLAUDE.md files, `~/.claude/projects/*/memory/`, `~/.claude/agent-memory/*`) that won't appear in mem0 until you backfill it. A one-shot importer ships with the plugin:
```bash
# Preview what would be imported:
python3 mem0-plugin/scripts/import_claude_state.py --dry-run
# Run the import:
python3 mem0-plugin/scripts/import_claude_state.py
```
What it does:
- Discovers every `CLAUDE.md` (user / project / nested), `@import` target, `.claude/rules/*.md`, `~/.claude/projects/*/memory/*.md`, and `~/.claude/agent-memory/*/*.md` on your machine.
- Heading-chunks each file (H2 default, descending to H3 for oversize sections, merging tiny siblings). Code-fenced blocks are preserved.
- POSTs each chunk to mem0 with metadata: `source_file`, `source_type`, `section_heading`, `project`, `subagent`, `content_hash`, and a `type` matching the `mem0-mcp` skill vocabulary (`convention`, `user_preference`, `task_learning`, `anti_pattern`, `decision`).
- Writes `~/.mem0/imports/claude-state.json` to remember what's been uploaded. Re-running is idempotent: only new chunks are sent.
Flags:
| Flag | Meaning |
|------|---------|
| `--dry-run` | Print the import plan; don't upload. |
| `--reset` | Wipe the marker before running (forces full re-import). |
| `--no-infer` | Upload chunks raw (`infer=False`). Default is `infer=True` so mem0 extracts atomic facts. |
| `--source <type>` | Restrict to one source type (e.g., `--source memory_md`). |
The `SessionStart` hook nudges you about this importer the first time you open a session after installing. Once you run it (or once the marker file exists), the nudge disappears.
## MCP Tools
Once installed, the following tools are available:
+745
View File
@@ -0,0 +1,745 @@
#!/usr/bin/env python3
"""Holistic import of on-disk Claude state into mem0.
Backfills the user's existing CLAUDE.md hierarchy (+ @imports),
`.claude/rules`, `~/.claude/projects/*/memory`, and `~/.claude/agent-memory`
into mem0 as searchable memories. One-shot. Idempotent across re-runs via a
marker file at ~/.mem0/imports/claude-state.json.
Patterned on on_pre_compact.py: urllib-based POSTs, stderr logging,
_identity.resolve_user_id() for user_id, and an optional ~/.mem0/hooks.log
when MEM0_DEBUG is set.
Invocation:
python3 import_claude_state.py [--dry-run] [--reset] [--no-infer]
[--source TYPE]
Exit codes:
0 = success (including dry-run and partial failures)
1 = fatal (missing MEM0_API_KEY, auth failure)
"""
from __future__ import annotations
import argparse
import hashlib
import json
import logging
import os
import re
import sys
import tempfile
import time
import urllib.error
import urllib.request
from dataclasses import dataclass
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
# ── Logging ──────────────────────────────────────────────────────────────
log = logging.getLogger("mem0-import-claude-state")
log.setLevel(logging.DEBUG)
_handler = logging.StreamHandler(sys.stderr)
_handler.setFormatter(logging.Formatter("[mem0-import-claude-state] %(message)s"))
log.addHandler(_handler)
if os.environ.get("MEM0_DEBUG"):
_log_dir = os.path.expanduser("~/.mem0")
try:
os.makedirs(_log_dir, exist_ok=True)
_file_handler = logging.FileHandler(os.path.join(_log_dir, "hooks.log"))
_file_handler.setFormatter(
logging.Formatter("[mem0-import-claude-state] %(asctime)s %(message)s")
)
log.addHandler(_file_handler)
except OSError:
pass
# ── Constants ────────────────────────────────────────────────────────────
API_URL = "https://api.mem0.ai"
REQUEST_TIMEOUT_SECS = 15
RETRY_SLEEP_SECS = 2.0
MARKER_PATH = Path.home() / ".mem0" / "imports" / "claude-state.json"
EMPTY_MARKER: dict[str, Any] = {"schema_version": 1, "imports": []}
MAX_CHARS_PER_CHUNK = 2000
MIN_CHARS_PER_CHUNK = 200
MAX_IMPORT_DEPTH = 5
H1_LINE = re.compile(r"^#\s+(.+?)\s*$", re.MULTILINE)
H2_LINE = re.compile(r"^##\s+(.+?)\s*$", re.MULTILINE)
H3_LINE = re.compile(r"^###\s+(.+?)\s*$", re.MULTILINE)
FENCE_LINE = re.compile(r"^\s*(```+|~~~+)")
IMPORT_LINE_RE = re.compile(r"(?<![A-Za-z0-9_])@([^\s,;)]+)")
PROJECT_DIR_DECODE_RE = re.compile(r"^-Users-[^-]+-(.+)$")
SOURCE_TYPES = {
"claude_md_managed",
"claude_md_user",
"claude_md_project",
"claude_md_import",
"claude_local",
"rule",
"memory_md",
"memory_topic",
"agent_memory",
"agent_memory_topic",
}
DEFAULT_TYPE_BY_SOURCE = {
"claude_md_managed": "convention",
"claude_md_user": "convention",
"claude_md_project": "convention",
"claude_md_import": "convention",
"claude_local": "user_preference",
"rule": "convention",
"memory_md": "task_learning",
"memory_topic": "task_learning",
"agent_memory": "task_learning",
"agent_memory_topic": "task_learning",
}
# Override priority: anti_pattern > decision > user_preference > default.
_ANTI_PATTERN_RE = re.compile(r"\b(bug|fix|debug|never|always|critical)\b", re.IGNORECASE)
_DECISION_RE = re.compile(r"\b(decision|decided|chose|picked|chosen)\b", re.IGNORECASE)
_PREFERENCE_RE = re.compile(r"\b(prefer|preference|style)\b", re.IGNORECASE)
# ── Data classes ─────────────────────────────────────────────────────────
@dataclass
class Chunk:
heading: str
body: str
content_hash: str
@dataclass
class Source:
path: Path
source_type: str
project_name: str | None = None
subagent_name: str | None = None
class AuthError(RuntimeError):
"""Raised on 401 to abort the import run immediately."""
# ── Marker I/O ───────────────────────────────────────────────────────────
def read_marker(path: Path) -> dict[str, Any]:
"""Read the import marker, tolerating missing files and corrupted JSON."""
if not path.exists():
return {"schema_version": 1, "imports": []}
try:
with path.open("r", encoding="utf-8") as f:
data = json.load(f)
if not isinstance(data, dict) or "imports" not in data:
log.warning("Marker at %s has unexpected shape, treating as empty", path)
return {"schema_version": 1, "imports": []}
return data
except (json.JSONDecodeError, OSError) as e:
log.warning("Marker at %s is corrupted (%s), treating as empty", path, e)
return {"schema_version": 1, "imports": []}
def write_marker(path: Path, marker: dict[str, Any]) -> None:
"""Atomically write the marker file (temp + os.replace)."""
path.parent.mkdir(parents=True, exist_ok=True)
fd, tmp_name = tempfile.mkstemp(
prefix=path.name + ".",
suffix=".tmp",
dir=str(path.parent),
)
try:
with os.fdopen(fd, "w", encoding="utf-8") as f:
json.dump(marker, f, indent=2, sort_keys=True)
f.write("\n")
os.replace(tmp_name, path)
except Exception:
try:
os.unlink(tmp_name)
except OSError:
pass
raise
# ── Chunker ──────────────────────────────────────────────────────────────
def _sha256(text: str) -> str:
return hashlib.sha256(text.encode("utf-8")).hexdigest()
def _make_chunk(heading: str, body: str) -> Chunk:
body = body.strip("\n")
return Chunk(heading=heading, body=body, content_hash=_sha256(body))
def _mask_fences(text: str) -> str:
"""Replace leading `#` of heading-like lines inside fenced code blocks with
a space so the heading regex won't split them. Length-preserving so chunk
offsets stay aligned with the original text.
"""
out: list[str] = []
open_marker: str | None = None
for line in text.splitlines(keepends=True):
m = FENCE_LINE.match(line)
if m:
marker = m.group(1)[:3]
if open_marker is None:
open_marker = marker
elif open_marker == marker:
open_marker = None
out.append(line)
elif open_marker is not None:
stripped = line.lstrip()
if stripped.startswith("#"):
idx = len(line) - len(stripped)
out.append(line[:idx] + " " + line[idx + 1 :])
else:
out.append(line)
else:
out.append(line)
return "".join(out)
def _split_on_pattern(text: str, pattern: re.Pattern[str]) -> list[tuple[str, str]]:
"""Split text into (heading, body) sections at each match of `pattern`.
Content before the first match becomes one section with the H1 (if any) as
its heading. Bodies preserve heading lines verbatim.
"""
matches = list(pattern.finditer(text))
if not matches:
h1 = H1_LINE.search(text)
heading = h1.group(1).strip() if h1 else ""
return [(heading, text)]
sections: list[tuple[str, str]] = []
first = matches[0]
preamble = text[: first.start()]
if preamble.strip():
h1 = H1_LINE.search(preamble)
heading = h1.group(1).strip() if h1 else ""
sections.append((heading, preamble))
for i, m in enumerate(matches):
heading = m.group(1).strip()
body_start = m.start()
body_end = matches[i + 1].start() if i + 1 < len(matches) else len(text)
sections.append((heading, text[body_start:body_end]))
return sections
def _descend_oversized(sections: list[tuple[str, str]]) -> list[tuple[str, str]]:
"""For each section whose body exceeds MAX_CHARS_PER_CHUNK, re-split at H3."""
result: list[tuple[str, str]] = []
for heading, body in sections:
if len(body) <= MAX_CHARS_PER_CHUNK:
result.append((heading, body))
continue
h3_subs = _split_on_pattern(body, H3_LINE)
if len(h3_subs) <= 1:
# No H3s present — emit oversized chunk rather than break paragraphs.
result.append((heading, body))
else:
result.extend(h3_subs)
return result
def _merge_small_siblings(sections: list[tuple[str, str]]) -> list[tuple[str, str]]:
"""Merge adjacent chunks where the *first* sibling's stripped body is under
MIN_CHARS_PER_CHUNK and the combined size still fits in MAX_CHARS_PER_CHUNK.
The merged chunk keeps the first sibling's heading.
"""
if not sections:
return sections
merged: list[tuple[str, str]] = []
buf_heading, buf_body = sections[0]
for heading, body in sections[1:]:
if (
len(buf_body.strip()) < MIN_CHARS_PER_CHUNK
and len(buf_body) + len(body) <= MAX_CHARS_PER_CHUNK
):
buf_body = buf_body + "\n\n" + body
else:
merged.append((buf_heading, buf_body))
buf_heading, buf_body = heading, body
merged.append((buf_heading, buf_body))
return merged
def chunk_markdown(text: str) -> list[Chunk]:
"""Split markdown into Chunks for upload to mem0.
1. Mask heading-like lines inside fenced code blocks.
2. Split the masked text at H2.
3. For sections over MAX_CHARS_PER_CHUNK, re-split at H3.
4. Merge adjacent sections whose first body is under MIN_CHARS_PER_CHUNK.
5. Slice bodies from the ORIGINAL text using offsets so user wording is
preserved exactly.
"""
masked = _mask_fences(text)
if len(masked) != len(text):
log.warning("fence mask changed text length (%d -> %d); falling back to whole-file chunk", len(text), len(masked))
return [_make_chunk("", text)]
sections = _split_on_pattern(masked, H2_LINE)
sections = _descend_oversized(sections)
sections = _merge_small_siblings(sections)
result: list[Chunk] = []
cursor = 0
for heading, masked_body in sections:
end = cursor + len(masked_body)
original_body = text[cursor:end]
result.append(_make_chunk(heading, original_body))
cursor = end
return result
# ── Tagger ───────────────────────────────────────────────────────────────
def _override_type(default: str, search_text: str) -> str:
if _ANTI_PATTERN_RE.search(search_text):
return "anti_pattern"
if _DECISION_RE.search(search_text):
return "decision"
if _PREFERENCE_RE.search(search_text):
return "user_preference"
return default
def tag_chunk(source: Source, chunk: Chunk) -> dict[str, Any]:
"""Build the metadata dict attached to the chunk's add_memory POST."""
default = DEFAULT_TYPE_BY_SOURCE.get(source.source_type, "task_learning")
type_value = _override_type(default, chunk.heading + "\n" + chunk.body)
return {
"source_file": str(source.path),
"source_type": source.source_type,
"section_heading": chunk.heading,
"project": source.project_name,
"subagent": source.subagent_name,
"content_hash": chunk.content_hash,
"type": type_value,
}
# ── @import resolver ─────────────────────────────────────────────────────
def _parse_import_paths(text: str, base: Path) -> list[Path]:
paths: list[Path] = []
for m in IMPORT_LINE_RE.finditer(text):
raw = m.group(1)
if raw.startswith("~"):
resolved = Path(os.path.expanduser(raw))
elif raw.startswith("/"):
resolved = Path(raw)
else:
resolved = (base.parent / raw).resolve()
paths.append(resolved)
return paths
def resolve_imports(
path: Path,
*,
depth: int = 0,
visited: set[Path] | None = None,
) -> list[Path]:
"""Return the transitive @-import closure starting at `path`.
The starting `path` is NOT included. Cycles, missing targets, and depth
overflows are logged and skipped (matches Claude Code's loader).
"""
if visited is None:
visited = set()
if depth >= MAX_IMPORT_DEPTH:
log.warning("import depth limit reached at %s (max=%d)", path, MAX_IMPORT_DEPTH)
return []
try:
text = path.read_text(encoding="utf-8")
except OSError:
return []
out: list[Path] = []
for target in _parse_import_paths(text, path):
try:
target_resolved = target.resolve()
except OSError:
continue
if target_resolved in visited:
log.warning("import cycle: skipping %s (re-entered from %s)", target_resolved, path)
continue
if not target_resolved.exists():
log.warning("missing import target %s referenced from %s", target_resolved, path)
continue
visited.add(target_resolved)
out.append(target_resolved)
out.extend(resolve_imports(target_resolved, depth=depth + 1, visited=visited))
return out
# ── Discovery ────────────────────────────────────────────────────────────
def _decode_project_name(encoded: str) -> str:
"""`-Users-alice-mem0-platform` → `mem0-platform`."""
m = PROJECT_DIR_DECODE_RE.match(encoded)
if m:
return m.group(1)
return encoded.lstrip("-")
def _collect_claude_md_chain(cwd: Path) -> list[Path]:
"""Walk up from cwd collecting CLAUDE.md / .claude/CLAUDE.md at each level."""
chain: list[Path] = []
seen: set[Path] = set()
current = cwd
while True:
for candidate in (current / "CLAUDE.md", current / ".claude" / "CLAUDE.md"):
try:
if candidate.exists():
resolved = candidate.resolve()
if resolved not in seen:
chain.append(resolved)
seen.add(resolved)
except OSError:
continue
parent = current.parent
if parent == current:
break
current = parent
return chain
def discover(*, home: Path | None = None, cwd: Path | None = None) -> list[Source]:
"""Enumerate every Claude state file we know how to import."""
home = (home or Path.home()).resolve()
cwd = (cwd or Path.cwd()).resolve()
sources: list[Source] = []
seen_paths: set[Path] = set()
def add(path: Path, source_type: str, **kw: Any) -> None:
try:
resolved = path.resolve()
except OSError:
return
if resolved in seen_paths or not resolved.exists():
return
seen_paths.add(resolved)
sources.append(Source(path=resolved, source_type=source_type, **kw))
# User-global CLAUDE.md
add(home / ".claude" / "CLAUDE.md", "claude_md_user")
# User-level .claude/rules
user_rules_dir = home / ".claude" / "rules"
if user_rules_dir.is_dir():
for rule in sorted(user_rules_dir.rglob("*.md")):
add(rule, "rule")
# Project CLAUDE.md chain (walking up from cwd)
for path in _collect_claude_md_chain(cwd):
add(path, "claude_md_project")
# Project CLAUDE.local.md
add(cwd / "CLAUDE.local.md", "claude_local")
# Project-level .claude/rules
project_rules_dir = cwd / ".claude" / "rules"
if project_rules_dir.is_dir():
for rule in sorted(project_rules_dir.rglob("*.md")):
add(rule, "rule")
# Per-project auto-memory
projects_root = home / ".claude" / "projects"
if projects_root.is_dir():
for proj_dir in sorted(projects_root.iterdir()):
mem_dir = proj_dir / "memory"
if not mem_dir.is_dir():
continue
proj_name = _decode_project_name(proj_dir.name)
for md in sorted(mem_dir.glob("*.md")):
source_type = "memory_md" if md.name == "MEMORY.md" else "memory_topic"
add(md, source_type, project_name=proj_name)
# Subagent auto-memory
agent_root = home / ".claude" / "agent-memory"
if agent_root.is_dir():
for sub_dir in sorted(agent_root.iterdir()):
if not sub_dir.is_dir():
continue
for md in sorted(sub_dir.glob("*.md")):
source_type = "agent_memory" if md.name == "MEMORY.md" else "agent_memory_topic"
add(md, source_type, subagent_name=sub_dir.name)
# Follow @imports from every CLAUDE.md / local / rule found above.
snapshot = list(sources)
for s in snapshot:
if s.source_type in {"claude_md_user", "claude_md_project", "claude_local", "rule"}:
for imp in resolve_imports(s.path):
add(imp, "claude_md_import")
return sources
# ── Dispatcher ───────────────────────────────────────────────────────────
def post_memory(
*,
content: str,
user_id: str,
metadata: dict[str, Any],
infer: bool,
api_key: str,
) -> dict[str, Any] | None:
"""POST one chunk to /v1/memories/. Returns the parsed response or None on
recoverable failure. Raises AuthError on 401."""
body = {
"messages": [{"role": "user", "content": content}],
"user_id": user_id,
"metadata": metadata,
"infer": infer,
}
data = json.dumps(body).encode("utf-8")
headers = {
"Content-Type": "application/json",
"Authorization": f"Token {api_key}",
}
for attempt in (1, 2):
req = urllib.request.Request(
f"{API_URL}/v1/memories/",
data=data,
headers=headers,
method="POST",
)
try:
with urllib.request.urlopen(req, timeout=REQUEST_TIMEOUT_SECS) as resp:
if resp.status in (200, 201):
return json.loads(resp.read().decode("utf-8"))
log.warning("API returned status %d", resp.status)
return None
except urllib.error.HTTPError as e:
if e.code == 401:
raise AuthError("Auth failed — check MEM0_API_KEY") from e
if e.code == 429 and attempt == 1:
log.warning("rate limited; sleeping %.1fs and retrying once", RETRY_SLEEP_SECS)
time.sleep(RETRY_SLEEP_SECS)
continue
log.warning("API error %d: %s", e.code, e.reason)
return None
except urllib.error.URLError as e:
log.warning("network error: %s", e)
return None
return None
# ── Orchestrator ─────────────────────────────────────────────────────────
def _now_iso() -> str:
return datetime.now(timezone.utc).isoformat(timespec="seconds").replace("+00:00", "Z")
def _hash_set_for_file(marker: dict[str, Any], file_path: str) -> set[str]:
for entry in marker.get("imports", []):
if entry.get("file") == file_path:
return {c["content_hash"] for c in entry.get("chunks", []) if "content_hash" in c}
return set()
def _get_or_create_file_entry(
marker: dict[str, Any], file_path: str, source_type: str
) -> dict[str, Any]:
for entry in marker["imports"]:
if entry.get("file") == file_path:
return entry
entry = {
"file": file_path,
"source_type": source_type,
"imported_at": _now_iso(),
"chunks": [],
}
marker["imports"].append(entry)
return entry
def run_import(
*,
home: Path,
cwd: Path,
marker_path: Path,
user_id: str,
api_key: str,
infer: bool,
dry_run: bool,
source_filter: str | None,
) -> dict[str, int]:
"""Orchestrate the full import. Returns stats: planned, uploaded, skipped,
failed, files.
"""
marker = read_marker(marker_path)
sources = discover(home=home, cwd=cwd)
if source_filter:
sources = [s for s in sources if s.source_type == source_filter]
stats = {
"planned": 0,
"uploaded": 0,
"skipped": 0,
"failed": 0,
"files": len(sources),
}
for source in sources:
try:
text = source.path.read_text(encoding="utf-8")
except OSError as e:
log.warning("cannot read %s: %s", source.path, e)
continue
if not text.strip():
continue
chunks = chunk_markdown(text)
already = _hash_set_for_file(marker, str(source.path))
for chunk in chunks:
stats["planned"] += 1
if chunk.content_hash in already:
stats["skipped"] += 1
continue
if dry_run:
continue
metadata = tag_chunk(source, chunk)
try:
resp = post_memory(
content=chunk.body,
user_id=user_id,
metadata=metadata,
infer=infer,
api_key=api_key,
)
except AuthError:
log.error("auth failed — aborting import")
raise
if resp is None:
stats["failed"] += 1
continue
memory_ids = [item.get("id", "") for item in resp.get("results", [])]
entry = _get_or_create_file_entry(marker, str(source.path), source.source_type)
entry["chunks"].append(
{
"heading": chunk.heading,
"content_hash": chunk.content_hash,
"memory_ids": memory_ids,
}
)
marker["last_run_at"] = _now_iso()
marker["user_id"] = user_id
write_marker(marker_path, marker)
stats["uploaded"] += 1
return stats
# ── CLI ──────────────────────────────────────────────────────────────────
def _get_home() -> Path:
return Path.home()
def _get_cwd() -> Path:
return Path.cwd()
def _build_arg_parser() -> argparse.ArgumentParser:
p = argparse.ArgumentParser(
prog="import_claude_state.py",
description="Holistic import of on-disk Claude state into mem0.",
)
p.add_argument("--dry-run", action="store_true", help="Print plan, don't upload.")
p.add_argument("--reset", action="store_true", help="Delete marker before running (forces full re-import).")
p.add_argument("--no-infer", action="store_true", help="Upload chunks raw (infer=False) instead of letting mem0 extract facts.")
p.add_argument(
"--source",
choices=sorted(SOURCE_TYPES),
default=None,
help="Restrict to one source_type.",
)
return p
def _print_summary(stats: dict[str, int], *, dry_run: bool) -> None:
if dry_run:
print(
f"[dry-run] Would upload {stats['planned'] - stats['skipped']} chunks "
f"from {stats['files']} files ({stats['skipped']} already imported, "
f"planned={stats['planned']})."
)
return
print(
f"Imported {stats['uploaded']} chunks from {stats['files']} files "
f"(skipped {stats['skipped']}, failed {stats['failed']})."
)
def main(argv: list[str] | None = None) -> int:
parser = _build_arg_parser()
args = parser.parse_args(argv)
api_key = os.environ.get("MEM0_API_KEY", "")
if not api_key:
print("MEM0_API_KEY not set — see plugin README.", file=sys.stderr)
return 1
try:
from _identity import resolve_user_id # type: ignore[import-not-found]
except ImportError:
def resolve_user_id() -> str:
return os.environ.get("MEM0_USER_ID") or os.environ.get("USER") or "default"
user_id = resolve_user_id()
if args.reset:
try:
MARKER_PATH.unlink()
except FileNotFoundError:
pass
try:
stats = run_import(
home=_get_home(),
cwd=_get_cwd(),
marker_path=MARKER_PATH,
user_id=user_id,
api_key=api_key,
infer=not args.no_infer,
dry_run=args.dry_run,
source_filter=args.source,
)
except AuthError as e:
print(str(e), file=sys.stderr)
return 1
_print_summary(stats, dry_run=args.dry_run)
return 0
if __name__ == "__main__":
sys.exit(main())
+19
View File
@@ -84,4 +84,23 @@ is being captured to mem0 in the background as `metadata.type=compact_summary`.
EOF
fi
# ── Holistic import nudge ────────────────────────────────────────────────
# One-time nudge to backfill existing CLAUDE.md / MEMORY.md / agent-memory
# into mem0. Silent after the marker file appears (i.e., after first run).
# Only emit on startup or resume -- during compact the user is mid-conversation.
if [ "$SOURCE" != "compact" ] && [ ! -f "$HOME/.mem0/imports/claude-state.json" ]; then
cat <<EOF
## Holistic import available
On-disk Claude state (CLAUDE.md, ~/.claude/projects/*/memory, ~/.claude/agent-memory)
has never been imported into mem0. To preview what would be imported, run:
python3 "$SCRIPT_DIR/import_claude_state.py" --dry-run
Then drop \`--dry-run\` to import. This nudge disappears after the first
successful run. Pass \`--reset\` to re-import from scratch.
EOF
fi
exit 0