fix(plugins): stamp surface identity at record time, not send time
Six defects in 0.3.x plugin telemetry. Defects 1, 2 and 6 were not three bugs: they were one spool protocol getting three properties wrong. Identity was decided by the wrong process. `harness` was stamped in record(), correctly, but `source` was read from a module global in flush() — so whichever process drained the spool named every event in it. Two processes never call init(): mcp_server.py, and the detached `python3 telemetry.py` sender that spawn_flush() starts. record() now stamps source beside harness, and the build generates core/_harness_id.py per host so identity resolves with no init() call at all. That also unifies two defaults that disagreed (`<host>_plugin` vs `MEM0_<HOST>_PLUGIN`), which could yield three source values for one plugin. Ownership was inferred, not held. Path.replace is os.rename, which preserves mtime, so a claim made after a quiet minute inherited the spool's age and was stealable the instant it existed. Claims are touched on creation and the per-batch rewrite doubles as a lease heartbeat. Progress was not durable. flush() returned on the first failed batch without truncating, so the retry re-posted from index 0 — 150 events delivered 250 times. It now rewrites the claim with the unsent remainder after every batch, bounding a crash to one repeated batch, and each event carries a uuid. Parked batches starved. They were only reachable when no spool existed, and because sessions keep recording there usually was one, so a batch parked by a failed send waited until the 7-day expiry deleted it unsent — despite its own presence being what starts the sender. flush() drains them in the same run, and expiry now applies only after a genuine retry has failed. code.install counted upgrades and repeat sessions. is_first_run() read the identity file, which only a successful flush writes, so an offline user recorded an install every session forever. A dedicated install-state.json is claimed atomically at record time; a non-empty data directory reads as an upgrade. The docs called this anonymous. Every event carries the account email, and the hashes were unsalted SHA-256 over a git remote URL or an absolute path containing the username. READMEs, the module docstring and a new docs section now say what the code does, and repo/session digests are salted per install. A cached email outlived an API key change. It is now re-resolved when the key's fingerprint differs, and $identify aliases anonymous->email only — aliasing one account to another merges person profiles irreversibly. All six shipped green because the shared core's only tests lived under one host, behind a conftest that calls init() at import. Core behaviour was never exercised uninitialised. Adds agent-plugin-core/tests with no init, including subprocess tests and coverage for the portable plugin, which has no flush worker and would pass a native-only test vacuously. Also puts the three surface headers on the SDKs, CLIs and integrations, and corrects a README claiming ZAPIER/STRANDS were already in the platform allowlist. Verified: 59 core tests, 203 claude-code, 11 cursor, 5 codex, 2 kimi, 6 antigravity. ruff and compileall clean. --check clean for all six hosts. TypeScript changes are not typechecked locally (deps not installed). Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
This commit is contained in:
@@ -31,7 +31,8 @@ export class PlatformBackend implements Backend {
|
||||
this.headers = {
|
||||
Authorization: `Token ${config.apiKey}`,
|
||||
"Content-Type": "application/json",
|
||||
"X-Mem0-Source": "cli",
|
||||
"X-Mem0-Source": "CLI",
|
||||
"X-Mem0-Client": `mem0-cli-node/${CLI_VERSION}`,
|
||||
"X-Mem0-Client-Language": "node",
|
||||
"X-Mem0-Client-Version": CLI_VERSION,
|
||||
};
|
||||
|
||||
@@ -27,7 +27,8 @@ class PlatformBackend(Backend):
|
||||
headers={
|
||||
"Authorization": f"Token {config.api_key}",
|
||||
"Content-Type": "application/json",
|
||||
"X-Mem0-Source": "cli",
|
||||
"X-Mem0-Source": "CLI",
|
||||
"X-Mem0-Client": f"mem0-cli-python/{__version__}",
|
||||
"X-Mem0-Client-Language": "python",
|
||||
"X-Mem0-Client-Version": __version__,
|
||||
},
|
||||
|
||||
@@ -172,6 +172,31 @@ claude plugin update mem0@mem0-plugins --scope user
|
||||
| Sidekick won't start | Must be in a Git repo. Check that your Claude Code version supports plugin agents and worktrees. |
|
||||
| Remove the plugin | `claude plugin uninstall mem0@mem0-plugins` |
|
||||
|
||||
## Telemetry
|
||||
|
||||
The plugin sends usage events (which hook ran, timing, result counts, failure
|
||||
types) so Mem0 can see what's used and what's breaking.
|
||||
|
||||
These events are **not anonymous**. When an API key is configured, which
|
||||
installing the plugin requires, they are sent under your Mem0 account email,
|
||||
the same way the Python SDK and the CLI attribute theirs. Without a key they
|
||||
are sent under a random per-machine id.
|
||||
|
||||
Each event carries the event name, the plugin version, the harness it ran in,
|
||||
your OS and Python version, timings, counts, and a coarse failure label.
|
||||
Repository and session identifiers are hashed with a random salt generated on
|
||||
your machine and never sent, so they cannot be linked back to a repository name
|
||||
or path.
|
||||
|
||||
Prompts, memory text, queries, file paths, repository names, and API keys are
|
||||
never sent.
|
||||
|
||||
Turn it off:
|
||||
|
||||
```bash
|
||||
export MEM0_TELEMETRY=false
|
||||
```
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Mem0 MCP Setup" icon="puzzle-piece" href="/platform/mem0-mcp">
|
||||
Detailed MCP configuration for all clients
|
||||
|
||||
@@ -49,6 +49,29 @@ Run the type check after every TypeScript change: `pnpm run typecheck` or `tsc -
|
||||
- **`zapier-mem0/`** is a Zapier Platform CLI app: add, search, get, delete. It deploys to Zapier, not npm, so it is **not** in the release router. Deploy it with `gh workflow run zapier-mem0-cd.yml --ref main` (needs the `ZAPIER_DEPLOY_KEY` secret).
|
||||
- **`mem0-strands/`** is a native Strands `MemoryStore` (Python, published to PyPI as `mem0-strands`). It plugs into the Strands `MemoryManager` for automatic recall and server-side extraction, over the hosted Mem0 platform or self-hosted Mem0 OSS. The package lives under `mem0-strands/python/`.
|
||||
|
||||
## Surface attribution
|
||||
|
||||
Every integration tells the Mem0 platform which surface it is. Three headers,
|
||||
and the rules on them are what keep one layer from erasing another:
|
||||
|
||||
| Header | Carries | Rule |
|
||||
|--------|---------|------|
|
||||
| `X-Mem0-Source` | one canonical source value | **set-once** — write only if absent |
|
||||
| `X-Application` | the host app it runs inside | **set-once** — write only if absent |
|
||||
| `X-Mem0-Client` | `name/version`, outermost first | **append-only** — add yourself, never replace |
|
||||
|
||||
Set-once means `setdefault`, never assignment. An integration that wraps the
|
||||
SDK is the outermost layer and sets the source; the SDK underneath must defer to
|
||||
it. Assignment is exactly how every agent plugin came to be indistinguishable
|
||||
from every other one at the platform.
|
||||
|
||||
Append-only means a plugin calling the Python SDK produces
|
||||
`mem0-plugin/0.3.1, mem0-python/2.0.19`, so neither layer can erase the other.
|
||||
|
||||
The backend recognizes a fixed list of source values and buckets everything else
|
||||
into `OTHERS`. A new value has to land in the platform's `EventSource` enum, so
|
||||
do not invent one without that change going in too.
|
||||
|
||||
## Adding an integration
|
||||
|
||||
1. For a native coding-agent host, add `integrations/<name>-plugin/` with `plugin-build.json`, its manifest, and a thin adapter, then generate its shared runtime. Portable clients use the single `mem0-agent-plugin/` package. Independent TypeScript integrations stay self-contained and import shared lifecycle behavior from `agent-plugin-core/typescript/`.
|
||||
@@ -59,3 +82,4 @@ Run the type check after every TypeScript change: `pnpm run typecheck` or `tsc -
|
||||
5. If it is a Claude Code or editor marketplace plugin, register the generated native bundle path in the applicable marketplace files. Preserve the existing public plugin name.
|
||||
6. Document it under `docs/integrations/` and add the page to `docs/docs.json` and `docs/llms.txt`.
|
||||
7. Add rows to the table above and to the CI/CD tables in [`../.github/AGENTS.md`](../.github/AGENTS.md).
|
||||
8. Send the three headers in [Surface attribution](#surface-attribution), and land the matching `EventSource` value on the platform in the same week. Until it exists, your traffic reports as `OTHERS`.
|
||||
|
||||
@@ -81,6 +81,28 @@ def replace_output(staged: Path, output: Path) -> Path:
|
||||
return output
|
||||
|
||||
|
||||
def _render_harness_id(host: str) -> str:
|
||||
"""Emit core/_harness_id.py for one host.
|
||||
|
||||
Carries both vocabularies from a single definition: the PostHog `source` tag
|
||||
and the platform's X-Mem0-Source / X-Application pair. Keeping them together
|
||||
is what stops the two from drifting into separate vocabularies for the same
|
||||
thing.
|
||||
"""
|
||||
tag = host.upper().replace("-", "_") + "_PLUGIN"
|
||||
return (
|
||||
'"""Generated by integrations/agent-plugin-core/build/build.py. Do not edit."""\n'
|
||||
"\n"
|
||||
f'HARNESS_ID = "{host}"\n'
|
||||
f'SOURCE_TAG = "{tag}"\n'
|
||||
"\n"
|
||||
"# Platform-side vocabulary (mem0_event.source + X-Application). The whole\n"
|
||||
"# plugin family is one source; which editor it runs in is the application.\n"
|
||||
'PLATFORM_SOURCE = "MEM0_PLUGIN"\n'
|
||||
f'PLATFORM_APPLICATION = "{host}"\n'
|
||||
)
|
||||
|
||||
|
||||
def _bundle_python(
|
||||
staged: Path,
|
||||
host: str,
|
||||
@@ -96,6 +118,12 @@ def _bundle_python(
|
||||
continue
|
||||
shutil.copy2(source, core / source.name)
|
||||
|
||||
# Generated per host so identity does not depend on an entrypoint remembering
|
||||
# to call telemetry.init(). mcp_server.py and the detached telemetry.py sender
|
||||
# never did, which is how MCP searches reported harness=generic and every
|
||||
# batch they drained was labelled MEM0_PLUGIN regardless of the real host.
|
||||
(core / "_harness_id.py").write_text(_render_harness_id(host), encoding="utf-8")
|
||||
|
||||
values = {
|
||||
"PLUGIN_ROOT": plugin_root,
|
||||
"PLUGIN_DATA": "${PLUGIN_DATA}",
|
||||
|
||||
@@ -305,8 +305,19 @@ def run(
|
||||
return 0
|
||||
|
||||
if args.action == "session-start":
|
||||
if telemetry.is_first_run():
|
||||
# Claims the marker atomically and says which event to record, so a
|
||||
# second session starting alongside this one cannot record it too.
|
||||
first_event = telemetry.claim_install()
|
||||
if first_event == "install":
|
||||
telemetry.record("install")
|
||||
elif first_event == "upgrade":
|
||||
# First run after a build that never wrote the marker; the
|
||||
# predecessor version was never recorded anywhere.
|
||||
telemetry.record("upgrade", from_version="pre-0.3")
|
||||
else:
|
||||
previous = telemetry.claim_version_change()
|
||||
if previous:
|
||||
telemetry.record("upgrade", from_version=previous)
|
||||
recovered = recover_pending_handoffs()
|
||||
record_session_start(store, hook_input)
|
||||
if recovered:
|
||||
|
||||
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
|
||||
return batches
|
||||
|
||||
|
||||
# Platform surface attribution. Read from the generated per-host module so a new
|
||||
# entrypoint is correct without remembering to configure anything.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
except ImportError:
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
|
||||
def platform_headers(key: str) -> dict[str, str]:
|
||||
"""Auth plus the three surface-identity headers.
|
||||
|
||||
X-Mem0-Source and X-Application are set-once by contract: this is the
|
||||
outermost layer, so it sets them, and nothing below may overwrite them.
|
||||
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
|
||||
"""
|
||||
headers = {
|
||||
"Authorization": f"Token {key}",
|
||||
"Content-Type": "application/json",
|
||||
"X-Mem0-Source": _PLATFORM_SOURCE,
|
||||
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
|
||||
}
|
||||
if _PLATFORM_APPLICATION:
|
||||
headers["X-Application"] = _PLATFORM_APPLICATION
|
||||
return headers
|
||||
|
||||
|
||||
def _request_json(
|
||||
url: str, key: str, payload: dict[str, Any], timeout: float
|
||||
) -> tuple[dict[str, Any] | list[Any], int, int]:
|
||||
@@ -1807,7 +1835,7 @@ def _request_json(
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
data=raw,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="POST",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1834,7 +1862,7 @@ def _get_json(
|
||||
) -> tuple[dict[str, Any] | list[Any], int]:
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="GET",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1980,6 +2008,10 @@ def flush_session(
|
||||
"user_id": write_user,
|
||||
"app_id": repo.app_id,
|
||||
"run_id": session_id,
|
||||
# Top level, not metadata: the backend reads `source` from the body,
|
||||
# query string or X-Mem0-Source header, never from metadata. The
|
||||
# harness tag stays in metadata as hook provenance.
|
||||
"source": _PLATFORM_SOURCE,
|
||||
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
|
||||
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
|
||||
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
|
||||
@@ -2523,7 +2555,7 @@ def _collect_memory_ids(
|
||||
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
|
||||
request = urllib.request.Request(
|
||||
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="DELETE",
|
||||
)
|
||||
try:
|
||||
|
||||
@@ -1,5 +1,9 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Anonymous usage telemetry for Mem0 agent plugins.
|
||||
"""Usage telemetry for Mem0 agent plugins.
|
||||
|
||||
Events are linked to your Mem0 account email when an API key is configured, and
|
||||
to a random per-machine id otherwise. Not anonymous — the Python SDK and CLI
|
||||
attribute the same way.
|
||||
|
||||
Hooks run on a 3-6 second budget and fire on every tool call, so recording never
|
||||
touches the network: `record` appends one JSON line to a local spool and returns.
|
||||
@@ -9,7 +13,8 @@ started once per session and again from the flush worker that is already detache
|
||||
Pure stdlib, matching the rest of the plugin. Opt out with MEM0_TELEMETRY=false.
|
||||
|
||||
Never sends prompts, memory text, queries, file paths, repository names, or API
|
||||
keys: only event names, durations, counts, coarse outcomes, and salted hashes.
|
||||
keys: only event names, durations, counts, coarse outcomes, and repo/session
|
||||
identifiers hashed with a random per-install salt.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -29,8 +34,23 @@ from typing import Any
|
||||
|
||||
import memory_core
|
||||
|
||||
_harness: str = "generic"
|
||||
_source_tag: str = "MEM0_PLUGIN"
|
||||
# Seeded from the per-host module the build generates into core/. Two processes
|
||||
# in this pipeline never call init() — mcp_server.py, and the detached
|
||||
# `python3 telemetry.py` sender that spawn_flush() starts — so a module default
|
||||
# was what every one of their events got labelled with.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import HARNESS_ID as _DEFAULT_HARNESS
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
from _harness_id import SOURCE_TAG as _DEFAULT_SOURCE_TAG
|
||||
except ImportError:
|
||||
_DEFAULT_HARNESS = "generic"
|
||||
_DEFAULT_SOURCE_TAG = "MEM0_PLUGIN"
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
_harness: str = _DEFAULT_HARNESS
|
||||
_source_tag: str = _DEFAULT_SOURCE_TAG
|
||||
_PRIVATE_KEYS = {
|
||||
"apikey",
|
||||
"authorization",
|
||||
@@ -56,10 +76,19 @@ _PRIVATE_KEYS = {
|
||||
}
|
||||
|
||||
|
||||
def init(harness: str = "generic", source_tag: str = "") -> None:
|
||||
def init(harness: str = "", source_tag: str = "") -> None:
|
||||
"""Override the generated identity. Optional — core/_harness_id.py is the default.
|
||||
|
||||
The fallback shape matches memory_core.configure_harness's (``<HOST>_PLUGIN``).
|
||||
It used to be ``MEM0_<HOST>_PLUGIN`` here and ``<host>_plugin`` there, which
|
||||
meant one plugin could emit three different source values depending on which
|
||||
process happened to send the batch.
|
||||
"""
|
||||
global _harness, _source_tag
|
||||
_harness = harness
|
||||
_source_tag = source_tag or f"MEM0_{harness.upper().replace('-', '_')}_PLUGIN"
|
||||
_harness = harness or _DEFAULT_HARNESS
|
||||
_source_tag = source_tag or (
|
||||
f"{_harness.upper().replace('-', '_')}_PLUGIN" if harness else _DEFAULT_SOURCE_TAG
|
||||
)
|
||||
|
||||
POSTHOG_API_KEY = "phc_hgJkUVJFYtmaJqrvf6CYN67TIQ8yhXAkWzUn9AMU4yX"
|
||||
POSTHOG_CAPTURE_URL = "https://us.i.posthog.com/i/v0/e/"
|
||||
@@ -70,6 +99,11 @@ BATCH_SIZE = 100
|
||||
SEND_TIMEOUT = 5
|
||||
CLAIM_STALE_SECONDS = 120
|
||||
CLAIM_EXPIRY_SECONDS = 7 * 24 * 60 * 60
|
||||
# A batch is only discarded once it has genuinely been retried this many times.
|
||||
MAX_CLAIM_ATTEMPTS = 3
|
||||
# Parked claims drained per run, after the live spool. Bounded so a long backlog
|
||||
# cannot turn one flush into an unbounded send loop.
|
||||
MAX_PARKED_PER_RUN = 3
|
||||
|
||||
|
||||
def is_enabled() -> bool:
|
||||
@@ -83,9 +117,36 @@ def is_enabled() -> bool:
|
||||
|
||||
|
||||
def _digest(value: str, length: int = 16) -> str:
|
||||
"""Unsalted digest. Only for values that are already secrets (API keys)."""
|
||||
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _install_salt() -> str:
|
||||
"""Random per-install salt, created on first use and kept in the identity file."""
|
||||
identity = _read_identity()
|
||||
salt = identity.get("salt")
|
||||
if not salt:
|
||||
salt = uuid.uuid4().hex
|
||||
identity["salt"] = salt
|
||||
_write_identity(identity)
|
||||
return salt
|
||||
|
||||
|
||||
def _scoped_digest(value: str, length: int = 16) -> str:
|
||||
"""Salted digest for values drawn from a guessable space.
|
||||
|
||||
repo.identity is a git remote URL, or ``local:<absolute path>`` when there is
|
||||
no remote — which normally contains the account username. Sixteen unsalted
|
||||
hex characters over that input space is enumerable, so this is not a
|
||||
privacy control without the salt. Salting per install keeps every
|
||||
within-account join the analytics actually use and gives up only
|
||||
cross-machine joins on the same repository, which nothing computes.
|
||||
"""
|
||||
if not value:
|
||||
return ""
|
||||
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _safe_value(value: Any) -> Any:
|
||||
if isinstance(value, str):
|
||||
return memory_core.redact(value)
|
||||
@@ -145,9 +206,92 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
|
||||
return created
|
||||
|
||||
|
||||
def _install_state_path() -> Path:
|
||||
return memory_core.data_dir() / "install-state.json"
|
||||
|
||||
|
||||
def is_first_run() -> bool:
|
||||
"""Whether this machine has never recorded a plugin event before."""
|
||||
return not _identity_path().exists()
|
||||
"""Whether install has never been recorded on this machine.
|
||||
|
||||
Deliberately NOT the identity file. That file is only written by a
|
||||
successful flush, so an offline or firewalled user recorded code.install on
|
||||
every single session, forever — and every 0.2.x user recorded one on their
|
||||
first 0.3.x session because 0.2.x never wrote it at all.
|
||||
"""
|
||||
return not _install_state_path().exists()
|
||||
|
||||
|
||||
def claim_install() -> str | None:
|
||||
"""Claim the one install/upgrade record for this machine, atomically.
|
||||
|
||||
Returns the event to record ("install" or "upgrade"), or None if another
|
||||
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
|
||||
cannot both win.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
# A fresh install has an empty data directory. Anything already there —
|
||||
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
|
||||
# before the marker is created, since creating it would itself be content.
|
||||
upgrading = _data_dir_has_content()
|
||||
try:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
|
||||
except FileExistsError:
|
||||
return None
|
||||
except OSError:
|
||||
return None
|
||||
try:
|
||||
with os.fdopen(handle, "w", encoding="utf-8") as stream:
|
||||
json.dump(
|
||||
{
|
||||
"plugin_version": memory_core.PLUGIN_VERSION,
|
||||
"installed_at": memory_core.utc_now(),
|
||||
"upgraded": upgrading,
|
||||
},
|
||||
stream,
|
||||
)
|
||||
except OSError:
|
||||
pass
|
||||
return "upgrade" if upgrading else "install"
|
||||
|
||||
|
||||
def _data_dir_has_content() -> bool:
|
||||
"""Whether anything predates this session in the plugin data directory."""
|
||||
try:
|
||||
for entry in memory_core.data_dir().iterdir():
|
||||
if entry.name != "install-state.json":
|
||||
return True
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def claim_version_change() -> str | None:
|
||||
"""Return the previously recorded version if it differs, updating the marker.
|
||||
|
||||
Only meaningful once the marker exists — the first transition into 0.3.x has
|
||||
no recorded predecessor and reports "pre-0.3" instead. Claiming by rewriting
|
||||
the marker means the next session sees no change and records nothing.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
try:
|
||||
state = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return None
|
||||
if not isinstance(state, dict):
|
||||
return None
|
||||
previous = str(state.get("plugin_version") or "")
|
||||
if not previous or previous == memory_core.PLUGIN_VERSION:
|
||||
return None
|
||||
state["plugin_version"] = memory_core.PLUGIN_VERSION
|
||||
state["upgraded_at"] = memory_core.utc_now()
|
||||
try:
|
||||
temporary = path.with_suffix(f".{os.getpid()}.tmp")
|
||||
temporary.write_text(json.dumps(state), encoding="utf-8")
|
||||
temporary.replace(path)
|
||||
except OSError:
|
||||
return None
|
||||
return previous
|
||||
|
||||
|
||||
def record(
|
||||
@@ -168,19 +312,25 @@ def record(
|
||||
except OSError:
|
||||
pass
|
||||
properties = _safe_value(properties)
|
||||
# Stamped in the RECORDING process, beside harness. `source` used to be
|
||||
# read in the sending process from a module global, so whichever process
|
||||
# drained the spool named every event in it. flush() spreads per-event
|
||||
# properties last, so this now wins over any sender's default.
|
||||
properties.update(
|
||||
harness=_harness,
|
||||
source=_source_tag,
|
||||
plugin_version=memory_core.PLUGIN_VERSION,
|
||||
os=sys.platform,
|
||||
python_version=platform.python_version(),
|
||||
)
|
||||
if repo is not None:
|
||||
properties["repo_hash"] = _digest(getattr(repo, "identity", ""))
|
||||
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
|
||||
if session_id:
|
||||
properties["session_hash"] = _digest(session_id)
|
||||
properties["session_hash"] = _scoped_digest(session_id)
|
||||
line = json.dumps(
|
||||
{
|
||||
"event": f"{EVENT_PREFIX}.{event}",
|
||||
"uuid": str(uuid.uuid4()),
|
||||
"timestamp": memory_core.utc_now(),
|
||||
"properties": {
|
||||
key: value for key, value in properties.items() if value is not None
|
||||
@@ -239,38 +389,139 @@ def spawn_flush() -> bool:
|
||||
return False
|
||||
|
||||
|
||||
def _claim_name(attempt: int = 0) -> str:
|
||||
"""Claim filename. The attempt count rides in the name so the 7-day expiry
|
||||
only ever discards a batch that was actually retried and failed."""
|
||||
return f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}-a{attempt}.sending"
|
||||
|
||||
|
||||
def _claim_attempt(claim: Path) -> int:
|
||||
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
|
||||
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
|
||||
tail = stem.rsplit("-", 1)[-1]
|
||||
if tail.startswith("a") and tail[1:].isdigit():
|
||||
return int(tail[1:])
|
||||
return 0
|
||||
|
||||
|
||||
def _touch(path: Path) -> None:
|
||||
"""Refresh mtime so a claim's age measures time since it was claimed.
|
||||
|
||||
``Path.replace`` is ``os.rename``, which preserves mtime — so a claim created
|
||||
after a quiet minute inherited the spool's last-write time and looked
|
||||
abandoned the instant it was made. A second sender would then take it over
|
||||
while the first was still posting, and both would deliver the batch.
|
||||
"""
|
||||
try:
|
||||
os.utime(path, None)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _claim_spool() -> Path | None:
|
||||
"""Rename the spool aside so exactly one sender owns each batch."""
|
||||
directory = memory_core.data_dir()
|
||||
claim = directory / f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}.sending"
|
||||
claim = directory / _claim_name()
|
||||
spool = _spool_path()
|
||||
try:
|
||||
spool.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
pass
|
||||
return _claim_parked(directory)
|
||||
|
||||
|
||||
def _claim_parked(directory: Path) -> Path | None:
|
||||
"""Take the oldest abandoned claim, if any lease has actually expired.
|
||||
|
||||
Kept separate from the live spool so flush() can drain both in one run.
|
||||
Previously parked batches were only reachable when no spool existed at all,
|
||||
and because sessions keep recording there usually was one — so a batch
|
||||
parked by a failed send waited until the 7-day expiry deleted it unsent,
|
||||
even though its own presence is what started the sender.
|
||||
"""
|
||||
now = time.time()
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending")):
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending"), key=_safe_mtime):
|
||||
try:
|
||||
age = now - orphan.stat().st_mtime
|
||||
except OSError:
|
||||
continue
|
||||
if age > CLAIM_EXPIRY_SECONDS:
|
||||
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
|
||||
try:
|
||||
orphan.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
continue
|
||||
if age < CLAIM_STALE_SECONDS:
|
||||
# Someone else holds a live lease on it.
|
||||
continue
|
||||
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
|
||||
try:
|
||||
orphan.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
def _safe_mtime(path: Path) -> float:
|
||||
try:
|
||||
return path.stat().st_mtime
|
||||
except OSError:
|
||||
return 0.0
|
||||
|
||||
|
||||
def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
|
||||
"""Persist the unsent remainder, atomically, and refresh the lease.
|
||||
|
||||
Called after every successful batch. Two jobs: a retry resumes where the
|
||||
send stopped instead of re-posting from the top, and the rewrite doubles as
|
||||
the lease heartbeat, so a slow sender does not have its claim stolen
|
||||
mid-flight. Interval is one batch, well inside CLAIM_STALE_SECONDS.
|
||||
"""
|
||||
if not remaining:
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return True
|
||||
temporary = claim.with_suffix(f".{os.getpid()}.partial")
|
||||
try:
|
||||
temporary.write_text(
|
||||
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
|
||||
encoding="utf-8",
|
||||
)
|
||||
temporary.replace(claim)
|
||||
_touch(claim)
|
||||
return True
|
||||
except OSError:
|
||||
try:
|
||||
temporary.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
|
||||
"""Persist the remainder and drop the lease, because this sender has given up.
|
||||
|
||||
Distinct from the per-batch heartbeat: heartbeating on the way out would
|
||||
make an abandoned batch look actively owned for a further
|
||||
CLAIM_STALE_SECONDS, delaying the retry for no reason. Ageing it past the
|
||||
threshold lets the next flush pick it up immediately, while the attempt
|
||||
count in the filename still bounds how many times that can happen.
|
||||
"""
|
||||
if not _rewrite_claim(claim, remaining):
|
||||
return
|
||||
try:
|
||||
released = time.time() - CLAIM_STALE_SECONDS - 1
|
||||
os.utime(claim, (released, released))
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _resolve_email(key: str) -> str:
|
||||
"""Trade the API key for the account email so events join other Mem0 surfaces."""
|
||||
url = os.environ.get("MEM0_API_URL", memory_core.DEFAULT_API_URL).rstrip("/") + "/v1/ping/"
|
||||
@@ -300,34 +551,76 @@ def _post(payload: dict[str, Any], url: str) -> bool:
|
||||
|
||||
|
||||
def resolve_distinct_id() -> tuple[str, str]:
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any."""
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any.
|
||||
|
||||
The second value becomes a PostHog $identify alias. It is ONLY ever an
|
||||
anonymous id: aliasing one account email to another merges two real person
|
||||
profiles and cannot be undone, so a key that now belongs to a different
|
||||
account re-resolves with no alias.
|
||||
"""
|
||||
identity = _read_identity()
|
||||
email = identity.get("email", "")
|
||||
if email:
|
||||
return email, ""
|
||||
key = memory_core.api_key()
|
||||
fingerprint = _digest(key) if key else ""
|
||||
email = identity.get("email", "")
|
||||
|
||||
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
|
||||
return email, ""
|
||||
|
||||
if not key:
|
||||
# No key to verify the account with; do not keep attributing to it.
|
||||
if email:
|
||||
identity.pop("email", None)
|
||||
identity.pop("key_fingerprint", None)
|
||||
_write_identity(identity)
|
||||
return anonymous_id(identity), ""
|
||||
email = _resolve_email(key)
|
||||
if not email:
|
||||
return anonymous_id(identity), ""
|
||||
previous = identity.get("anonymous_id", "")
|
||||
identity["email"] = email
|
||||
|
||||
resolved = _resolve_email(key)
|
||||
if not resolved:
|
||||
return (email, "") if email else (anonymous_id(identity), "")
|
||||
|
||||
# Alias only when going anonymous -> email for the first time.
|
||||
previous = "" if email else identity.get("anonymous_id", "")
|
||||
identity["email"] = resolved
|
||||
identity["key_fingerprint"] = fingerprint
|
||||
_write_identity(identity)
|
||||
return email, previous
|
||||
return resolved, previous
|
||||
|
||||
|
||||
def flush() -> int:
|
||||
"""Drain claimed spools to PostHog and return the number of events sent."""
|
||||
"""Drain the live spool, then any parked claims, and return events sent."""
|
||||
if not is_enabled():
|
||||
return 0
|
||||
claim = _claim_spool()
|
||||
sent, delivered = _drain(_claim_spool())
|
||||
if not delivered:
|
||||
# The network is failing. Retrying other batches now would only burn
|
||||
# their attempt budget against the same broken connection.
|
||||
return sent
|
||||
|
||||
# Parked batches used to starve behind the live spool indefinitely. Bounded
|
||||
# per run so a long backlog cannot turn one flush into an unbounded loop.
|
||||
directory = memory_core.data_dir()
|
||||
for _ in range(MAX_PARKED_PER_RUN):
|
||||
parked = _claim_parked(directory)
|
||||
if parked is None:
|
||||
break
|
||||
count, delivered = _drain(parked)
|
||||
sent += count
|
||||
if not delivered:
|
||||
break
|
||||
return sent
|
||||
|
||||
|
||||
def _drain(claim: Path | None) -> tuple[int, bool]:
|
||||
"""Post one claimed batch file, recording progress after every batch.
|
||||
|
||||
Returns (events sent, whether everything was delivered).
|
||||
"""
|
||||
if claim is None:
|
||||
return 0
|
||||
return 0, True
|
||||
try:
|
||||
lines = claim.read_text(encoding="utf-8").splitlines()
|
||||
except OSError:
|
||||
return 0
|
||||
return 0, True
|
||||
events = []
|
||||
for line in lines:
|
||||
try:
|
||||
@@ -341,7 +634,7 @@ def flush() -> int:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return 0
|
||||
return 0, True
|
||||
|
||||
distinct_id, aliased_anonymous_id = resolve_distinct_id()
|
||||
if aliased_anonymous_id:
|
||||
@@ -360,12 +653,17 @@ def flush() -> int:
|
||||
|
||||
sent = 0
|
||||
for start in range(0, len(events), BATCH_SIZE):
|
||||
chunk = events[start : start + BATCH_SIZE]
|
||||
batch = [
|
||||
{
|
||||
"event": event["event"],
|
||||
"distinct_id": distinct_id,
|
||||
# Carried through from record() so a resend can be collapsed.
|
||||
"uuid": event.get("uuid"),
|
||||
"timestamp": event.get("timestamp"),
|
||||
"properties": {
|
||||
# Fallback only: events recorded by a build before source
|
||||
# moved into record() have none of their own.
|
||||
"source": _source_tag,
|
||||
"language": "python",
|
||||
"$process_person_profile": False,
|
||||
@@ -373,16 +671,19 @@ def flush() -> int:
|
||||
**(event.get("properties") or {}),
|
||||
},
|
||||
}
|
||||
for event in events[start : start + BATCH_SIZE]
|
||||
for event in chunk
|
||||
]
|
||||
if not _post({"api_key": POSTHOG_API_KEY, "batch": batch}, POSTHOG_BATCH_URL):
|
||||
return sent
|
||||
sent += len(batch)
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return sent
|
||||
# Keep only what has not been delivered, and release the lease.
|
||||
# Previously the whole file was kept and the retry re-posted every
|
||||
# batch, including the ones that had already arrived.
|
||||
_release_claim(claim, events[start:])
|
||||
return sent, False
|
||||
sent += len(chunk)
|
||||
# Record progress and refresh the lease after each successful batch, so
|
||||
# a crash repeats at most one batch instead of the entire file.
|
||||
_rewrite_claim(claim, events[start + len(chunk) :])
|
||||
return sent, True
|
||||
|
||||
|
||||
def main() -> int:
|
||||
|
||||
@@ -0,0 +1,164 @@
|
||||
"""Delivery semantics of the telemetry spool: no duplicates, no starvation.
|
||||
|
||||
These run against a built host's core in-process (not a subprocess) because they
|
||||
need to inject failures into ``_post``. The identity tests next door cover the
|
||||
uninitialised-process case that needs a real interpreter.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
CORE_ROOT = Path(__file__).resolve().parents[1]
|
||||
REPOSITORY_ROOT = CORE_ROOT.parents[1]
|
||||
HOST_CORE = REPOSITORY_ROOT / "integrations" / "claude-code-plugin" / "core"
|
||||
|
||||
pytestmark = pytest.mark.skipif(not HOST_CORE.exists(), reason="claude-code-plugin is not built")
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def telemetry(tmp_path, monkeypatch):
|
||||
monkeypatch.setenv("MEM0_CODE_DATA_DIR", str(tmp_path / "data"))
|
||||
monkeypatch.syspath_prepend(str(HOST_CORE))
|
||||
for name in ("telemetry", "memory_core", "_harness_id"):
|
||||
sys.modules.pop(name, None)
|
||||
module = importlib.import_module("telemetry")
|
||||
monkeypatch.setattr(module, "resolve_distinct_id", lambda: ("tester@example.com", ""))
|
||||
yield module
|
||||
for name in ("telemetry", "memory_core", "_harness_id"):
|
||||
sys.modules.pop(name, None)
|
||||
|
||||
|
||||
def _delivered(payloads):
|
||||
return [event for payload in payloads if "batch" in payload for event in payload["batch"]]
|
||||
|
||||
|
||||
def test_a_partial_failure_does_not_redeliver_what_already_arrived(telemetry):
|
||||
"""Defect 2a: flush kept the whole claim on failure and retried from the top.
|
||||
|
||||
150 events across two batches, the second failing, previously delivered 250.
|
||||
"""
|
||||
for index in range(150):
|
||||
telemetry.record("search", index=index)
|
||||
|
||||
sent: list[dict] = []
|
||||
calls = {"n": 0}
|
||||
|
||||
def flaky(payload, url):
|
||||
calls["n"] += 1
|
||||
if calls["n"] == 2: # second batch fails
|
||||
return False
|
||||
sent.append(payload)
|
||||
return True
|
||||
|
||||
telemetry._post = flaky
|
||||
telemetry.flush()
|
||||
|
||||
telemetry._post = lambda payload, url: sent.append(payload) or True
|
||||
telemetry.flush()
|
||||
|
||||
events = _delivered(sent)
|
||||
assert len(events) == 150
|
||||
assert len({event["uuid"] for event in events}) == 150
|
||||
|
||||
|
||||
def test_a_fresh_claim_is_not_immediately_stealable(telemetry):
|
||||
"""Defect 2b: rename preserves mtime, so a claim inherited the spool's age.
|
||||
|
||||
With the last write older than the stale threshold, a claim made now looked
|
||||
abandoned the instant it existed and a second sender took it over.
|
||||
"""
|
||||
telemetry.record("search")
|
||||
spool = telemetry._spool_path()
|
||||
old = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
|
||||
os.utime(spool, (old, old))
|
||||
|
||||
first = telemetry._claim_spool()
|
||||
assert first is not None
|
||||
|
||||
# A second sender starting right now must find nothing to take.
|
||||
assert telemetry._claim_parked(first.parent) is None
|
||||
|
||||
|
||||
def test_a_parked_batch_is_drained_behind_the_live_spool(telemetry):
|
||||
"""Defect 6: parked claims were only reachable when no spool existed.
|
||||
|
||||
Because sessions keep recording there usually was one, so a batch parked by
|
||||
a failed send waited until the 7-day expiry deleted it unsent — even though
|
||||
its own presence is what starts the sender.
|
||||
"""
|
||||
telemetry.record("parked")
|
||||
telemetry._post = lambda payload, url: False
|
||||
telemetry.flush()
|
||||
|
||||
parked = list(telemetry.memory_core.data_dir().glob("telemetry-*.sending"))
|
||||
assert len(parked) == 1
|
||||
old = time.time() - (telemetry.CLAIM_STALE_SECONDS + 60)
|
||||
os.utime(parked[0], (old, old))
|
||||
|
||||
telemetry.record("fresh")
|
||||
sent: list[dict] = []
|
||||
telemetry._post = lambda payload, url: sent.append(payload) or True
|
||||
telemetry.flush()
|
||||
|
||||
names = {event["event"] for event in _delivered(sent)}
|
||||
assert names == {"code.parked", "code.fresh"}
|
||||
|
||||
|
||||
def test_an_untried_batch_is_not_expired_by_age_alone(telemetry):
|
||||
"""Expiry should discard what failed, not what never got a turn."""
|
||||
telemetry.record("parked")
|
||||
telemetry._post = lambda payload, url: False
|
||||
telemetry.flush()
|
||||
|
||||
parked = list(telemetry.memory_core.data_dir().glob("telemetry-*.sending"))
|
||||
assert len(parked) == 1
|
||||
ancient = time.time() - (telemetry.CLAIM_EXPIRY_SECONDS + 60)
|
||||
os.utime(parked[0], (ancient, ancient))
|
||||
|
||||
sent: list[dict] = []
|
||||
telemetry._post = lambda payload, url: sent.append(payload) or True
|
||||
telemetry.flush()
|
||||
|
||||
assert [event["event"] for event in _delivered(sent)] == ["code.parked"]
|
||||
|
||||
|
||||
def test_progress_is_recorded_after_every_batch(telemetry):
|
||||
"""A crash repeats at most one batch, not the whole file."""
|
||||
for index in range(250):
|
||||
telemetry.record("search", index=index)
|
||||
|
||||
calls = {"n": 0}
|
||||
|
||||
def die_after_two(payload, url):
|
||||
calls["n"] += 1
|
||||
if calls["n"] > 2:
|
||||
return False
|
||||
return True
|
||||
|
||||
telemetry._post = die_after_two
|
||||
telemetry.flush()
|
||||
|
||||
parked = list(telemetry.memory_core.data_dir().glob("telemetry-*.sending"))
|
||||
assert len(parked) == 1
|
||||
remaining = parked[0].read_text(encoding="utf-8").strip().splitlines()
|
||||
# Two batches of 100 landed; only the last 50 should still be pending.
|
||||
assert len(remaining) == 50
|
||||
assert json.loads(remaining[0])["properties"]["index"] == 200
|
||||
|
||||
|
||||
def test_the_heartbeat_stays_well_inside_the_lease(telemetry):
|
||||
"""The claim rewrite doubles as the lease heartbeat.
|
||||
|
||||
_post makes a single attempt with SEND_TIMEOUT and no retry, so a heartbeat
|
||||
lands at least that often. If a retry loop is ever added to _post, this is
|
||||
the assertion that catches a sender losing its claim mid-flight.
|
||||
"""
|
||||
assert telemetry.SEND_TIMEOUT * 4 < telemetry.CLAIM_STALE_SECONDS
|
||||
@@ -0,0 +1,179 @@
|
||||
"""Core telemetry behaviour with NO telemetry.init(), in a real subprocess.
|
||||
|
||||
Why this file exists
|
||||
--------------------
|
||||
``telemetry.py`` lives in ``agent-plugin-core/python/`` but its only tests lived
|
||||
under ``claude-code-plugin/tests/``, behind a ``conftest.py`` that calls
|
||||
``configure_harness()`` and ``telemetry.init()`` at import. Core behaviour was
|
||||
therefore only ever exercised inside an already-configured module.
|
||||
|
||||
Two processes in the real pipeline never call ``init()``:
|
||||
|
||||
- ``mcp_server.py``, which records every manual search;
|
||||
- the detached ``python3 telemetry.py`` sender that ``spawn_flush()`` starts at
|
||||
session start, after every skill command, and when the MCP server exits.
|
||||
|
||||
Both fell back to module defaults, so MCP searches reported ``harness=generic``
|
||||
and everything that sender delivered was labelled ``MEM0_PLUGIN`` regardless of
|
||||
which of the six plugins produced it. The suite stayed green throughout.
|
||||
|
||||
These tests run in a fresh interpreter with no conftest, against a built host
|
||||
bundle, which is the only arrangement that can catch that class of bug.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
CORE_ROOT = Path(__file__).resolve().parents[1]
|
||||
REPOSITORY_ROOT = CORE_ROOT.parents[1]
|
||||
HOSTS = {
|
||||
"claude-code": ("claude-code-plugin", "CLAUDE_CODE_PLUGIN"),
|
||||
"cursor": ("cursor-plugin", "CURSOR_PLUGIN"),
|
||||
"codex": ("codex-plugin", "CODEX_PLUGIN"),
|
||||
"kimi": ("kimi-plugin", "KIMI_PLUGIN"),
|
||||
"antigravity": ("antigravity-plugin", "ANTIGRAVITY_PLUGIN"),
|
||||
# Portable: no flush_worker and no hook_runner, so its ONLY sender is the
|
||||
# uninitialised telemetry.py. A native-only test passes here vacuously.
|
||||
"coding-agent": ("mem0-agent-plugin", "CODING_AGENT_PLUGIN"),
|
||||
}
|
||||
|
||||
|
||||
def _core_dir(directory: str) -> Path:
|
||||
return REPOSITORY_ROOT / "integrations" / directory / "core"
|
||||
|
||||
|
||||
def _run(core: Path, data_dir: Path, body: str) -> str:
|
||||
"""Execute `body` in a fresh interpreter with only the host's core on sys.path."""
|
||||
script = f"import sys; sys.path.insert(0, {str(core)!r})\n{body}"
|
||||
result = subprocess.run(
|
||||
[sys.executable, "-c", script],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
env={
|
||||
"MEM0_CODE_DATA_DIR": str(data_dir),
|
||||
"PATH": "/usr/bin:/bin",
|
||||
"HOME": str(data_dir),
|
||||
},
|
||||
)
|
||||
assert result.returncode == 0, result.stderr
|
||||
return result.stdout.strip()
|
||||
|
||||
|
||||
@pytest.mark.parametrize("harness,spec", sorted(HOSTS.items()))
|
||||
def test_identity_resolves_without_init(harness, spec):
|
||||
"""Every built host knows what it is with no configuration call at all."""
|
||||
directory, source_tag = spec
|
||||
core = _core_dir(directory)
|
||||
if not core.exists():
|
||||
pytest.skip(f"{directory} is not built in this tree")
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
out = _run(
|
||||
core,
|
||||
Path(tmp),
|
||||
"import telemetry; print(telemetry._harness, telemetry._source_tag)",
|
||||
)
|
||||
assert out == f"{harness} {source_tag}"
|
||||
|
||||
|
||||
def test_mcp_server_records_the_real_harness():
|
||||
"""mcp_server imports telemetry and never initialises it (server.py has no init).
|
||||
|
||||
Its recorded events used to carry harness=generic for every plugin.
|
||||
"""
|
||||
core = _core_dir("claude-code-plugin")
|
||||
if not core.exists():
|
||||
pytest.skip("claude-code-plugin is not built in this tree")
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
data_dir = Path(tmp)
|
||||
_run(
|
||||
core,
|
||||
data_dir,
|
||||
"import mcp_server, telemetry; telemetry.record('search', trigger='mcp-search')",
|
||||
)
|
||||
spooled = (data_dir / "telemetry.jsonl").read_text(encoding="utf-8").strip()
|
||||
|
||||
event = json.loads(spooled)
|
||||
assert event["properties"]["harness"] == "claude-code"
|
||||
assert event["properties"]["source"] == "CLAUDE_CODE_PLUGIN"
|
||||
|
||||
|
||||
def test_the_detached_sender_does_not_relabel_events():
|
||||
"""`python3 telemetry.py` is the sender spawn_flush() starts, and never inits.
|
||||
|
||||
source is stamped at record time now, so which process sends is irrelevant.
|
||||
"""
|
||||
core = _core_dir("claude-code-plugin")
|
||||
if not core.exists():
|
||||
pytest.skip("claude-code-plugin is not built in this tree")
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
data_dir = Path(tmp)
|
||||
_run(core, data_dir, "import telemetry; telemetry.record('search')")
|
||||
|
||||
captured = data_dir / "captured.json"
|
||||
# Drain with a fresh, unconfigured interpreter, capturing the payload
|
||||
# instead of posting it.
|
||||
_run(
|
||||
core,
|
||||
data_dir,
|
||||
"import json, telemetry\n"
|
||||
"sent = []\n"
|
||||
"telemetry._post = lambda payload, url: sent.append(payload) or True\n"
|
||||
"telemetry.flush()\n"
|
||||
f"open({str(captured)!r}, 'w').write(json.dumps(sent))",
|
||||
)
|
||||
payloads = json.loads(captured.read_text(encoding="utf-8"))
|
||||
|
||||
batches = [p for p in payloads if "batch" in p]
|
||||
assert batches, "nothing was sent"
|
||||
properties = batches[0]["batch"][0]["properties"]
|
||||
assert properties["source"] == "CLAUDE_CODE_PLUGIN"
|
||||
assert properties["harness"] == "claude-code"
|
||||
|
||||
|
||||
def test_every_event_carries_a_uuid_for_dedupe():
|
||||
core = _core_dir("claude-code-plugin")
|
||||
if not core.exists():
|
||||
pytest.skip("claude-code-plugin is not built in this tree")
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
data_dir = Path(tmp)
|
||||
_run(core, data_dir, "import telemetry; telemetry.record('search'); telemetry.record('flush')")
|
||||
lines = (data_dir / "telemetry.jsonl").read_text(encoding="utf-8").strip().splitlines()
|
||||
|
||||
ids = [json.loads(line)["uuid"] for line in lines]
|
||||
assert len(ids) == 2
|
||||
assert len(set(ids)) == 2
|
||||
|
||||
|
||||
def test_source_tag_defaults_agree_between_the_two_modules():
|
||||
"""configure_harness and telemetry.init must derive the same tag.
|
||||
|
||||
They disagreed: `<host>_plugin` in one and `MEM0_<HOST>_PLUGIN` in the other,
|
||||
so one plugin could emit three different source values depending on which
|
||||
process sent the batch.
|
||||
"""
|
||||
core = _core_dir("claude-code-plugin")
|
||||
if not core.exists():
|
||||
pytest.skip("claude-code-plugin is not built in this tree")
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
out = _run(
|
||||
core,
|
||||
Path(tmp),
|
||||
"import memory_core, telemetry\n"
|
||||
"memory_core.configure_harness('kimi')\n"
|
||||
"telemetry.init(harness='kimi')\n"
|
||||
"print(memory_core.harness_config()['source_tag'].upper(), telemetry._source_tag)",
|
||||
)
|
||||
left, right = out.split()
|
||||
assert left == right == "KIMI_PLUGIN"
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Generated by integrations/agent-plugin-core/build/build.py. Do not edit."""
|
||||
|
||||
HARNESS_ID = "antigravity"
|
||||
SOURCE_TAG = "ANTIGRAVITY_PLUGIN"
|
||||
|
||||
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
|
||||
# plugin family is one source; which editor it runs in is the application.
|
||||
PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
PLATFORM_APPLICATION = "antigravity"
|
||||
@@ -305,8 +305,19 @@ def run(
|
||||
return 0
|
||||
|
||||
if args.action == "session-start":
|
||||
if telemetry.is_first_run():
|
||||
# Claims the marker atomically and says which event to record, so a
|
||||
# second session starting alongside this one cannot record it too.
|
||||
first_event = telemetry.claim_install()
|
||||
if first_event == "install":
|
||||
telemetry.record("install")
|
||||
elif first_event == "upgrade":
|
||||
# First run after a build that never wrote the marker; the
|
||||
# predecessor version was never recorded anywhere.
|
||||
telemetry.record("upgrade", from_version="pre-0.3")
|
||||
else:
|
||||
previous = telemetry.claim_version_change()
|
||||
if previous:
|
||||
telemetry.record("upgrade", from_version=previous)
|
||||
recovered = recover_pending_handoffs()
|
||||
record_session_start(store, hook_input)
|
||||
if recovered:
|
||||
|
||||
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
|
||||
return batches
|
||||
|
||||
|
||||
# Platform surface attribution. Read from the generated per-host module so a new
|
||||
# entrypoint is correct without remembering to configure anything.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
except ImportError:
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
|
||||
def platform_headers(key: str) -> dict[str, str]:
|
||||
"""Auth plus the three surface-identity headers.
|
||||
|
||||
X-Mem0-Source and X-Application are set-once by contract: this is the
|
||||
outermost layer, so it sets them, and nothing below may overwrite them.
|
||||
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
|
||||
"""
|
||||
headers = {
|
||||
"Authorization": f"Token {key}",
|
||||
"Content-Type": "application/json",
|
||||
"X-Mem0-Source": _PLATFORM_SOURCE,
|
||||
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
|
||||
}
|
||||
if _PLATFORM_APPLICATION:
|
||||
headers["X-Application"] = _PLATFORM_APPLICATION
|
||||
return headers
|
||||
|
||||
|
||||
def _request_json(
|
||||
url: str, key: str, payload: dict[str, Any], timeout: float
|
||||
) -> tuple[dict[str, Any] | list[Any], int, int]:
|
||||
@@ -1807,7 +1835,7 @@ def _request_json(
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
data=raw,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="POST",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1834,7 +1862,7 @@ def _get_json(
|
||||
) -> tuple[dict[str, Any] | list[Any], int]:
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="GET",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1980,6 +2008,10 @@ def flush_session(
|
||||
"user_id": write_user,
|
||||
"app_id": repo.app_id,
|
||||
"run_id": session_id,
|
||||
# Top level, not metadata: the backend reads `source` from the body,
|
||||
# query string or X-Mem0-Source header, never from metadata. The
|
||||
# harness tag stays in metadata as hook provenance.
|
||||
"source": _PLATFORM_SOURCE,
|
||||
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
|
||||
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
|
||||
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
|
||||
@@ -2523,7 +2555,7 @@ def _collect_memory_ids(
|
||||
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
|
||||
request = urllib.request.Request(
|
||||
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="DELETE",
|
||||
)
|
||||
try:
|
||||
|
||||
@@ -1,5 +1,9 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Anonymous usage telemetry for Mem0 agent plugins.
|
||||
"""Usage telemetry for Mem0 agent plugins.
|
||||
|
||||
Events are linked to your Mem0 account email when an API key is configured, and
|
||||
to a random per-machine id otherwise. Not anonymous — the Python SDK and CLI
|
||||
attribute the same way.
|
||||
|
||||
Hooks run on a 3-6 second budget and fire on every tool call, so recording never
|
||||
touches the network: `record` appends one JSON line to a local spool and returns.
|
||||
@@ -9,7 +13,8 @@ started once per session and again from the flush worker that is already detache
|
||||
Pure stdlib, matching the rest of the plugin. Opt out with MEM0_TELEMETRY=false.
|
||||
|
||||
Never sends prompts, memory text, queries, file paths, repository names, or API
|
||||
keys: only event names, durations, counts, coarse outcomes, and salted hashes.
|
||||
keys: only event names, durations, counts, coarse outcomes, and repo/session
|
||||
identifiers hashed with a random per-install salt.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -29,8 +34,23 @@ from typing import Any
|
||||
|
||||
import memory_core
|
||||
|
||||
_harness: str = "generic"
|
||||
_source_tag: str = "MEM0_PLUGIN"
|
||||
# Seeded from the per-host module the build generates into core/. Two processes
|
||||
# in this pipeline never call init() — mcp_server.py, and the detached
|
||||
# `python3 telemetry.py` sender that spawn_flush() starts — so a module default
|
||||
# was what every one of their events got labelled with.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import HARNESS_ID as _DEFAULT_HARNESS
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
from _harness_id import SOURCE_TAG as _DEFAULT_SOURCE_TAG
|
||||
except ImportError:
|
||||
_DEFAULT_HARNESS = "generic"
|
||||
_DEFAULT_SOURCE_TAG = "MEM0_PLUGIN"
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
_harness: str = _DEFAULT_HARNESS
|
||||
_source_tag: str = _DEFAULT_SOURCE_TAG
|
||||
_PRIVATE_KEYS = {
|
||||
"apikey",
|
||||
"authorization",
|
||||
@@ -56,10 +76,19 @@ _PRIVATE_KEYS = {
|
||||
}
|
||||
|
||||
|
||||
def init(harness: str = "generic", source_tag: str = "") -> None:
|
||||
def init(harness: str = "", source_tag: str = "") -> None:
|
||||
"""Override the generated identity. Optional — core/_harness_id.py is the default.
|
||||
|
||||
The fallback shape matches memory_core.configure_harness's (``<HOST>_PLUGIN``).
|
||||
It used to be ``MEM0_<HOST>_PLUGIN`` here and ``<host>_plugin`` there, which
|
||||
meant one plugin could emit three different source values depending on which
|
||||
process happened to send the batch.
|
||||
"""
|
||||
global _harness, _source_tag
|
||||
_harness = harness
|
||||
_source_tag = source_tag or f"MEM0_{harness.upper().replace('-', '_')}_PLUGIN"
|
||||
_harness = harness or _DEFAULT_HARNESS
|
||||
_source_tag = source_tag or (
|
||||
f"{_harness.upper().replace('-', '_')}_PLUGIN" if harness else _DEFAULT_SOURCE_TAG
|
||||
)
|
||||
|
||||
POSTHOG_API_KEY = "phc_hgJkUVJFYtmaJqrvf6CYN67TIQ8yhXAkWzUn9AMU4yX"
|
||||
POSTHOG_CAPTURE_URL = "https://us.i.posthog.com/i/v0/e/"
|
||||
@@ -70,6 +99,11 @@ BATCH_SIZE = 100
|
||||
SEND_TIMEOUT = 5
|
||||
CLAIM_STALE_SECONDS = 120
|
||||
CLAIM_EXPIRY_SECONDS = 7 * 24 * 60 * 60
|
||||
# A batch is only discarded once it has genuinely been retried this many times.
|
||||
MAX_CLAIM_ATTEMPTS = 3
|
||||
# Parked claims drained per run, after the live spool. Bounded so a long backlog
|
||||
# cannot turn one flush into an unbounded send loop.
|
||||
MAX_PARKED_PER_RUN = 3
|
||||
|
||||
|
||||
def is_enabled() -> bool:
|
||||
@@ -83,9 +117,36 @@ def is_enabled() -> bool:
|
||||
|
||||
|
||||
def _digest(value: str, length: int = 16) -> str:
|
||||
"""Unsalted digest. Only for values that are already secrets (API keys)."""
|
||||
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _install_salt() -> str:
|
||||
"""Random per-install salt, created on first use and kept in the identity file."""
|
||||
identity = _read_identity()
|
||||
salt = identity.get("salt")
|
||||
if not salt:
|
||||
salt = uuid.uuid4().hex
|
||||
identity["salt"] = salt
|
||||
_write_identity(identity)
|
||||
return salt
|
||||
|
||||
|
||||
def _scoped_digest(value: str, length: int = 16) -> str:
|
||||
"""Salted digest for values drawn from a guessable space.
|
||||
|
||||
repo.identity is a git remote URL, or ``local:<absolute path>`` when there is
|
||||
no remote — which normally contains the account username. Sixteen unsalted
|
||||
hex characters over that input space is enumerable, so this is not a
|
||||
privacy control without the salt. Salting per install keeps every
|
||||
within-account join the analytics actually use and gives up only
|
||||
cross-machine joins on the same repository, which nothing computes.
|
||||
"""
|
||||
if not value:
|
||||
return ""
|
||||
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _safe_value(value: Any) -> Any:
|
||||
if isinstance(value, str):
|
||||
return memory_core.redact(value)
|
||||
@@ -145,9 +206,92 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
|
||||
return created
|
||||
|
||||
|
||||
def _install_state_path() -> Path:
|
||||
return memory_core.data_dir() / "install-state.json"
|
||||
|
||||
|
||||
def is_first_run() -> bool:
|
||||
"""Whether this machine has never recorded a plugin event before."""
|
||||
return not _identity_path().exists()
|
||||
"""Whether install has never been recorded on this machine.
|
||||
|
||||
Deliberately NOT the identity file. That file is only written by a
|
||||
successful flush, so an offline or firewalled user recorded code.install on
|
||||
every single session, forever — and every 0.2.x user recorded one on their
|
||||
first 0.3.x session because 0.2.x never wrote it at all.
|
||||
"""
|
||||
return not _install_state_path().exists()
|
||||
|
||||
|
||||
def claim_install() -> str | None:
|
||||
"""Claim the one install/upgrade record for this machine, atomically.
|
||||
|
||||
Returns the event to record ("install" or "upgrade"), or None if another
|
||||
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
|
||||
cannot both win.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
# A fresh install has an empty data directory. Anything already there —
|
||||
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
|
||||
# before the marker is created, since creating it would itself be content.
|
||||
upgrading = _data_dir_has_content()
|
||||
try:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
|
||||
except FileExistsError:
|
||||
return None
|
||||
except OSError:
|
||||
return None
|
||||
try:
|
||||
with os.fdopen(handle, "w", encoding="utf-8") as stream:
|
||||
json.dump(
|
||||
{
|
||||
"plugin_version": memory_core.PLUGIN_VERSION,
|
||||
"installed_at": memory_core.utc_now(),
|
||||
"upgraded": upgrading,
|
||||
},
|
||||
stream,
|
||||
)
|
||||
except OSError:
|
||||
pass
|
||||
return "upgrade" if upgrading else "install"
|
||||
|
||||
|
||||
def _data_dir_has_content() -> bool:
|
||||
"""Whether anything predates this session in the plugin data directory."""
|
||||
try:
|
||||
for entry in memory_core.data_dir().iterdir():
|
||||
if entry.name != "install-state.json":
|
||||
return True
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def claim_version_change() -> str | None:
|
||||
"""Return the previously recorded version if it differs, updating the marker.
|
||||
|
||||
Only meaningful once the marker exists — the first transition into 0.3.x has
|
||||
no recorded predecessor and reports "pre-0.3" instead. Claiming by rewriting
|
||||
the marker means the next session sees no change and records nothing.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
try:
|
||||
state = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return None
|
||||
if not isinstance(state, dict):
|
||||
return None
|
||||
previous = str(state.get("plugin_version") or "")
|
||||
if not previous or previous == memory_core.PLUGIN_VERSION:
|
||||
return None
|
||||
state["plugin_version"] = memory_core.PLUGIN_VERSION
|
||||
state["upgraded_at"] = memory_core.utc_now()
|
||||
try:
|
||||
temporary = path.with_suffix(f".{os.getpid()}.tmp")
|
||||
temporary.write_text(json.dumps(state), encoding="utf-8")
|
||||
temporary.replace(path)
|
||||
except OSError:
|
||||
return None
|
||||
return previous
|
||||
|
||||
|
||||
def record(
|
||||
@@ -168,19 +312,25 @@ def record(
|
||||
except OSError:
|
||||
pass
|
||||
properties = _safe_value(properties)
|
||||
# Stamped in the RECORDING process, beside harness. `source` used to be
|
||||
# read in the sending process from a module global, so whichever process
|
||||
# drained the spool named every event in it. flush() spreads per-event
|
||||
# properties last, so this now wins over any sender's default.
|
||||
properties.update(
|
||||
harness=_harness,
|
||||
source=_source_tag,
|
||||
plugin_version=memory_core.PLUGIN_VERSION,
|
||||
os=sys.platform,
|
||||
python_version=platform.python_version(),
|
||||
)
|
||||
if repo is not None:
|
||||
properties["repo_hash"] = _digest(getattr(repo, "identity", ""))
|
||||
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
|
||||
if session_id:
|
||||
properties["session_hash"] = _digest(session_id)
|
||||
properties["session_hash"] = _scoped_digest(session_id)
|
||||
line = json.dumps(
|
||||
{
|
||||
"event": f"{EVENT_PREFIX}.{event}",
|
||||
"uuid": str(uuid.uuid4()),
|
||||
"timestamp": memory_core.utc_now(),
|
||||
"properties": {
|
||||
key: value for key, value in properties.items() if value is not None
|
||||
@@ -239,38 +389,139 @@ def spawn_flush() -> bool:
|
||||
return False
|
||||
|
||||
|
||||
def _claim_name(attempt: int = 0) -> str:
|
||||
"""Claim filename. The attempt count rides in the name so the 7-day expiry
|
||||
only ever discards a batch that was actually retried and failed."""
|
||||
return f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}-a{attempt}.sending"
|
||||
|
||||
|
||||
def _claim_attempt(claim: Path) -> int:
|
||||
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
|
||||
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
|
||||
tail = stem.rsplit("-", 1)[-1]
|
||||
if tail.startswith("a") and tail[1:].isdigit():
|
||||
return int(tail[1:])
|
||||
return 0
|
||||
|
||||
|
||||
def _touch(path: Path) -> None:
|
||||
"""Refresh mtime so a claim's age measures time since it was claimed.
|
||||
|
||||
``Path.replace`` is ``os.rename``, which preserves mtime — so a claim created
|
||||
after a quiet minute inherited the spool's last-write time and looked
|
||||
abandoned the instant it was made. A second sender would then take it over
|
||||
while the first was still posting, and both would deliver the batch.
|
||||
"""
|
||||
try:
|
||||
os.utime(path, None)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _claim_spool() -> Path | None:
|
||||
"""Rename the spool aside so exactly one sender owns each batch."""
|
||||
directory = memory_core.data_dir()
|
||||
claim = directory / f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}.sending"
|
||||
claim = directory / _claim_name()
|
||||
spool = _spool_path()
|
||||
try:
|
||||
spool.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
pass
|
||||
return _claim_parked(directory)
|
||||
|
||||
|
||||
def _claim_parked(directory: Path) -> Path | None:
|
||||
"""Take the oldest abandoned claim, if any lease has actually expired.
|
||||
|
||||
Kept separate from the live spool so flush() can drain both in one run.
|
||||
Previously parked batches were only reachable when no spool existed at all,
|
||||
and because sessions keep recording there usually was one — so a batch
|
||||
parked by a failed send waited until the 7-day expiry deleted it unsent,
|
||||
even though its own presence is what started the sender.
|
||||
"""
|
||||
now = time.time()
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending")):
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending"), key=_safe_mtime):
|
||||
try:
|
||||
age = now - orphan.stat().st_mtime
|
||||
except OSError:
|
||||
continue
|
||||
if age > CLAIM_EXPIRY_SECONDS:
|
||||
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
|
||||
try:
|
||||
orphan.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
continue
|
||||
if age < CLAIM_STALE_SECONDS:
|
||||
# Someone else holds a live lease on it.
|
||||
continue
|
||||
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
|
||||
try:
|
||||
orphan.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
def _safe_mtime(path: Path) -> float:
|
||||
try:
|
||||
return path.stat().st_mtime
|
||||
except OSError:
|
||||
return 0.0
|
||||
|
||||
|
||||
def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
|
||||
"""Persist the unsent remainder, atomically, and refresh the lease.
|
||||
|
||||
Called after every successful batch. Two jobs: a retry resumes where the
|
||||
send stopped instead of re-posting from the top, and the rewrite doubles as
|
||||
the lease heartbeat, so a slow sender does not have its claim stolen
|
||||
mid-flight. Interval is one batch, well inside CLAIM_STALE_SECONDS.
|
||||
"""
|
||||
if not remaining:
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return True
|
||||
temporary = claim.with_suffix(f".{os.getpid()}.partial")
|
||||
try:
|
||||
temporary.write_text(
|
||||
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
|
||||
encoding="utf-8",
|
||||
)
|
||||
temporary.replace(claim)
|
||||
_touch(claim)
|
||||
return True
|
||||
except OSError:
|
||||
try:
|
||||
temporary.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
|
||||
"""Persist the remainder and drop the lease, because this sender has given up.
|
||||
|
||||
Distinct from the per-batch heartbeat: heartbeating on the way out would
|
||||
make an abandoned batch look actively owned for a further
|
||||
CLAIM_STALE_SECONDS, delaying the retry for no reason. Ageing it past the
|
||||
threshold lets the next flush pick it up immediately, while the attempt
|
||||
count in the filename still bounds how many times that can happen.
|
||||
"""
|
||||
if not _rewrite_claim(claim, remaining):
|
||||
return
|
||||
try:
|
||||
released = time.time() - CLAIM_STALE_SECONDS - 1
|
||||
os.utime(claim, (released, released))
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _resolve_email(key: str) -> str:
|
||||
"""Trade the API key for the account email so events join other Mem0 surfaces."""
|
||||
url = os.environ.get("MEM0_API_URL", memory_core.DEFAULT_API_URL).rstrip("/") + "/v1/ping/"
|
||||
@@ -300,34 +551,76 @@ def _post(payload: dict[str, Any], url: str) -> bool:
|
||||
|
||||
|
||||
def resolve_distinct_id() -> tuple[str, str]:
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any."""
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any.
|
||||
|
||||
The second value becomes a PostHog $identify alias. It is ONLY ever an
|
||||
anonymous id: aliasing one account email to another merges two real person
|
||||
profiles and cannot be undone, so a key that now belongs to a different
|
||||
account re-resolves with no alias.
|
||||
"""
|
||||
identity = _read_identity()
|
||||
email = identity.get("email", "")
|
||||
if email:
|
||||
return email, ""
|
||||
key = memory_core.api_key()
|
||||
fingerprint = _digest(key) if key else ""
|
||||
email = identity.get("email", "")
|
||||
|
||||
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
|
||||
return email, ""
|
||||
|
||||
if not key:
|
||||
# No key to verify the account with; do not keep attributing to it.
|
||||
if email:
|
||||
identity.pop("email", None)
|
||||
identity.pop("key_fingerprint", None)
|
||||
_write_identity(identity)
|
||||
return anonymous_id(identity), ""
|
||||
email = _resolve_email(key)
|
||||
if not email:
|
||||
return anonymous_id(identity), ""
|
||||
previous = identity.get("anonymous_id", "")
|
||||
identity["email"] = email
|
||||
|
||||
resolved = _resolve_email(key)
|
||||
if not resolved:
|
||||
return (email, "") if email else (anonymous_id(identity), "")
|
||||
|
||||
# Alias only when going anonymous -> email for the first time.
|
||||
previous = "" if email else identity.get("anonymous_id", "")
|
||||
identity["email"] = resolved
|
||||
identity["key_fingerprint"] = fingerprint
|
||||
_write_identity(identity)
|
||||
return email, previous
|
||||
return resolved, previous
|
||||
|
||||
|
||||
def flush() -> int:
|
||||
"""Drain claimed spools to PostHog and return the number of events sent."""
|
||||
"""Drain the live spool, then any parked claims, and return events sent."""
|
||||
if not is_enabled():
|
||||
return 0
|
||||
claim = _claim_spool()
|
||||
sent, delivered = _drain(_claim_spool())
|
||||
if not delivered:
|
||||
# The network is failing. Retrying other batches now would only burn
|
||||
# their attempt budget against the same broken connection.
|
||||
return sent
|
||||
|
||||
# Parked batches used to starve behind the live spool indefinitely. Bounded
|
||||
# per run so a long backlog cannot turn one flush into an unbounded loop.
|
||||
directory = memory_core.data_dir()
|
||||
for _ in range(MAX_PARKED_PER_RUN):
|
||||
parked = _claim_parked(directory)
|
||||
if parked is None:
|
||||
break
|
||||
count, delivered = _drain(parked)
|
||||
sent += count
|
||||
if not delivered:
|
||||
break
|
||||
return sent
|
||||
|
||||
|
||||
def _drain(claim: Path | None) -> tuple[int, bool]:
|
||||
"""Post one claimed batch file, recording progress after every batch.
|
||||
|
||||
Returns (events sent, whether everything was delivered).
|
||||
"""
|
||||
if claim is None:
|
||||
return 0
|
||||
return 0, True
|
||||
try:
|
||||
lines = claim.read_text(encoding="utf-8").splitlines()
|
||||
except OSError:
|
||||
return 0
|
||||
return 0, True
|
||||
events = []
|
||||
for line in lines:
|
||||
try:
|
||||
@@ -341,7 +634,7 @@ def flush() -> int:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return 0
|
||||
return 0, True
|
||||
|
||||
distinct_id, aliased_anonymous_id = resolve_distinct_id()
|
||||
if aliased_anonymous_id:
|
||||
@@ -360,12 +653,17 @@ def flush() -> int:
|
||||
|
||||
sent = 0
|
||||
for start in range(0, len(events), BATCH_SIZE):
|
||||
chunk = events[start : start + BATCH_SIZE]
|
||||
batch = [
|
||||
{
|
||||
"event": event["event"],
|
||||
"distinct_id": distinct_id,
|
||||
# Carried through from record() so a resend can be collapsed.
|
||||
"uuid": event.get("uuid"),
|
||||
"timestamp": event.get("timestamp"),
|
||||
"properties": {
|
||||
# Fallback only: events recorded by a build before source
|
||||
# moved into record() have none of their own.
|
||||
"source": _source_tag,
|
||||
"language": "python",
|
||||
"$process_person_profile": False,
|
||||
@@ -373,16 +671,19 @@ def flush() -> int:
|
||||
**(event.get("properties") or {}),
|
||||
},
|
||||
}
|
||||
for event in events[start : start + BATCH_SIZE]
|
||||
for event in chunk
|
||||
]
|
||||
if not _post({"api_key": POSTHOG_API_KEY, "batch": batch}, POSTHOG_BATCH_URL):
|
||||
return sent
|
||||
sent += len(batch)
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return sent
|
||||
# Keep only what has not been delivered, and release the lease.
|
||||
# Previously the whole file was kept and the retry re-posted every
|
||||
# batch, including the ones that had already arrived.
|
||||
_release_claim(claim, events[start:])
|
||||
return sent, False
|
||||
sent += len(chunk)
|
||||
# Record progress and refresh the lease after each successful batch, so
|
||||
# a crash repeats at most one batch instead of the entire file.
|
||||
_rewrite_claim(claim, events[start + len(chunk) :])
|
||||
return sent, True
|
||||
|
||||
|
||||
def main() -> int:
|
||||
|
||||
@@ -136,13 +136,19 @@ Local data lives in `${CLAUDE_PLUGIN_DATA}`:
|
||||
- `pending/`: sessions waiting to be sent to Mem0 (retried after interruption)
|
||||
- `flush-worker.log`: whether memory creation succeeded
|
||||
- `plugin-errors.log`: hook errors (no credentials)
|
||||
- `telemetry.jsonl` / `telemetry-identity.json`: anonymous usage events
|
||||
- `telemetry.jsonl` / `telemetry-identity.json`: usage events and the id they are sent under
|
||||
|
||||
Mem0 receives captured user messages, Claude's answers, sidekick assignments and completed responses, and changed file paths. When a failed command is recorded, extraction can also include bounded command details and results. Complete files and general tool output stay on your machine. Values that look like credentials are redacted before anything is sent.
|
||||
|
||||
## Telemetry
|
||||
|
||||
Anonymous usage events (which hook ran, timing, result counts, failure types) so Mem0 can identify what's used and what's breaking. Repo and session IDs are hashed before leaving your machine. Prompts, memory text, file paths, tool output, and API keys are never sent.
|
||||
Usage events (which hook ran, timing, result counts, failure types) so Mem0 can identify what's used and what's breaking.
|
||||
|
||||
**These events are not anonymous.** When an API key is configured — which installing the plugin requires — events are sent under your Mem0 account email, the same way the Python SDK and the CLI attribute theirs. Without a key they are sent under a random per-machine id.
|
||||
|
||||
What each event carries: the event name, the plugin version, the harness it ran in, your OS and Python version, timings, counts, and a coarse failure label. Repository and session identifiers are hashed with a random salt generated on your machine and never sent, so they cannot be linked back to a repository name or path.
|
||||
|
||||
Prompts, memory text, queries, file paths, repository names, and API keys are never sent.
|
||||
|
||||
Turn it off:
|
||||
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Generated by integrations/agent-plugin-core/build/build.py. Do not edit."""
|
||||
|
||||
HARNESS_ID = "claude-code"
|
||||
SOURCE_TAG = "CLAUDE_CODE_PLUGIN"
|
||||
|
||||
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
|
||||
# plugin family is one source; which editor it runs in is the application.
|
||||
PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
PLATFORM_APPLICATION = "claude-code"
|
||||
@@ -305,8 +305,19 @@ def run(
|
||||
return 0
|
||||
|
||||
if args.action == "session-start":
|
||||
if telemetry.is_first_run():
|
||||
# Claims the marker atomically and says which event to record, so a
|
||||
# second session starting alongside this one cannot record it too.
|
||||
first_event = telemetry.claim_install()
|
||||
if first_event == "install":
|
||||
telemetry.record("install")
|
||||
elif first_event == "upgrade":
|
||||
# First run after a build that never wrote the marker; the
|
||||
# predecessor version was never recorded anywhere.
|
||||
telemetry.record("upgrade", from_version="pre-0.3")
|
||||
else:
|
||||
previous = telemetry.claim_version_change()
|
||||
if previous:
|
||||
telemetry.record("upgrade", from_version=previous)
|
||||
recovered = recover_pending_handoffs()
|
||||
record_session_start(store, hook_input)
|
||||
if recovered:
|
||||
|
||||
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
|
||||
return batches
|
||||
|
||||
|
||||
# Platform surface attribution. Read from the generated per-host module so a new
|
||||
# entrypoint is correct without remembering to configure anything.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
except ImportError:
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
|
||||
def platform_headers(key: str) -> dict[str, str]:
|
||||
"""Auth plus the three surface-identity headers.
|
||||
|
||||
X-Mem0-Source and X-Application are set-once by contract: this is the
|
||||
outermost layer, so it sets them, and nothing below may overwrite them.
|
||||
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
|
||||
"""
|
||||
headers = {
|
||||
"Authorization": f"Token {key}",
|
||||
"Content-Type": "application/json",
|
||||
"X-Mem0-Source": _PLATFORM_SOURCE,
|
||||
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
|
||||
}
|
||||
if _PLATFORM_APPLICATION:
|
||||
headers["X-Application"] = _PLATFORM_APPLICATION
|
||||
return headers
|
||||
|
||||
|
||||
def _request_json(
|
||||
url: str, key: str, payload: dict[str, Any], timeout: float
|
||||
) -> tuple[dict[str, Any] | list[Any], int, int]:
|
||||
@@ -1807,7 +1835,7 @@ def _request_json(
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
data=raw,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="POST",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1834,7 +1862,7 @@ def _get_json(
|
||||
) -> tuple[dict[str, Any] | list[Any], int]:
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="GET",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1980,6 +2008,10 @@ def flush_session(
|
||||
"user_id": write_user,
|
||||
"app_id": repo.app_id,
|
||||
"run_id": session_id,
|
||||
# Top level, not metadata: the backend reads `source` from the body,
|
||||
# query string or X-Mem0-Source header, never from metadata. The
|
||||
# harness tag stays in metadata as hook provenance.
|
||||
"source": _PLATFORM_SOURCE,
|
||||
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
|
||||
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
|
||||
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
|
||||
@@ -2523,7 +2555,7 @@ def _collect_memory_ids(
|
||||
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
|
||||
request = urllib.request.Request(
|
||||
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="DELETE",
|
||||
)
|
||||
try:
|
||||
|
||||
@@ -1,5 +1,9 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Anonymous usage telemetry for Mem0 agent plugins.
|
||||
"""Usage telemetry for Mem0 agent plugins.
|
||||
|
||||
Events are linked to your Mem0 account email when an API key is configured, and
|
||||
to a random per-machine id otherwise. Not anonymous — the Python SDK and CLI
|
||||
attribute the same way.
|
||||
|
||||
Hooks run on a 3-6 second budget and fire on every tool call, so recording never
|
||||
touches the network: `record` appends one JSON line to a local spool and returns.
|
||||
@@ -9,7 +13,8 @@ started once per session and again from the flush worker that is already detache
|
||||
Pure stdlib, matching the rest of the plugin. Opt out with MEM0_TELEMETRY=false.
|
||||
|
||||
Never sends prompts, memory text, queries, file paths, repository names, or API
|
||||
keys: only event names, durations, counts, coarse outcomes, and salted hashes.
|
||||
keys: only event names, durations, counts, coarse outcomes, and repo/session
|
||||
identifiers hashed with a random per-install salt.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -29,8 +34,23 @@ from typing import Any
|
||||
|
||||
import memory_core
|
||||
|
||||
_harness: str = "generic"
|
||||
_source_tag: str = "MEM0_PLUGIN"
|
||||
# Seeded from the per-host module the build generates into core/. Two processes
|
||||
# in this pipeline never call init() — mcp_server.py, and the detached
|
||||
# `python3 telemetry.py` sender that spawn_flush() starts — so a module default
|
||||
# was what every one of their events got labelled with.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import HARNESS_ID as _DEFAULT_HARNESS
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
from _harness_id import SOURCE_TAG as _DEFAULT_SOURCE_TAG
|
||||
except ImportError:
|
||||
_DEFAULT_HARNESS = "generic"
|
||||
_DEFAULT_SOURCE_TAG = "MEM0_PLUGIN"
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
_harness: str = _DEFAULT_HARNESS
|
||||
_source_tag: str = _DEFAULT_SOURCE_TAG
|
||||
_PRIVATE_KEYS = {
|
||||
"apikey",
|
||||
"authorization",
|
||||
@@ -56,10 +76,19 @@ _PRIVATE_KEYS = {
|
||||
}
|
||||
|
||||
|
||||
def init(harness: str = "generic", source_tag: str = "") -> None:
|
||||
def init(harness: str = "", source_tag: str = "") -> None:
|
||||
"""Override the generated identity. Optional — core/_harness_id.py is the default.
|
||||
|
||||
The fallback shape matches memory_core.configure_harness's (``<HOST>_PLUGIN``).
|
||||
It used to be ``MEM0_<HOST>_PLUGIN`` here and ``<host>_plugin`` there, which
|
||||
meant one plugin could emit three different source values depending on which
|
||||
process happened to send the batch.
|
||||
"""
|
||||
global _harness, _source_tag
|
||||
_harness = harness
|
||||
_source_tag = source_tag or f"MEM0_{harness.upper().replace('-', '_')}_PLUGIN"
|
||||
_harness = harness or _DEFAULT_HARNESS
|
||||
_source_tag = source_tag or (
|
||||
f"{_harness.upper().replace('-', '_')}_PLUGIN" if harness else _DEFAULT_SOURCE_TAG
|
||||
)
|
||||
|
||||
POSTHOG_API_KEY = "phc_hgJkUVJFYtmaJqrvf6CYN67TIQ8yhXAkWzUn9AMU4yX"
|
||||
POSTHOG_CAPTURE_URL = "https://us.i.posthog.com/i/v0/e/"
|
||||
@@ -70,6 +99,11 @@ BATCH_SIZE = 100
|
||||
SEND_TIMEOUT = 5
|
||||
CLAIM_STALE_SECONDS = 120
|
||||
CLAIM_EXPIRY_SECONDS = 7 * 24 * 60 * 60
|
||||
# A batch is only discarded once it has genuinely been retried this many times.
|
||||
MAX_CLAIM_ATTEMPTS = 3
|
||||
# Parked claims drained per run, after the live spool. Bounded so a long backlog
|
||||
# cannot turn one flush into an unbounded send loop.
|
||||
MAX_PARKED_PER_RUN = 3
|
||||
|
||||
|
||||
def is_enabled() -> bool:
|
||||
@@ -83,9 +117,36 @@ def is_enabled() -> bool:
|
||||
|
||||
|
||||
def _digest(value: str, length: int = 16) -> str:
|
||||
"""Unsalted digest. Only for values that are already secrets (API keys)."""
|
||||
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _install_salt() -> str:
|
||||
"""Random per-install salt, created on first use and kept in the identity file."""
|
||||
identity = _read_identity()
|
||||
salt = identity.get("salt")
|
||||
if not salt:
|
||||
salt = uuid.uuid4().hex
|
||||
identity["salt"] = salt
|
||||
_write_identity(identity)
|
||||
return salt
|
||||
|
||||
|
||||
def _scoped_digest(value: str, length: int = 16) -> str:
|
||||
"""Salted digest for values drawn from a guessable space.
|
||||
|
||||
repo.identity is a git remote URL, or ``local:<absolute path>`` when there is
|
||||
no remote — which normally contains the account username. Sixteen unsalted
|
||||
hex characters over that input space is enumerable, so this is not a
|
||||
privacy control without the salt. Salting per install keeps every
|
||||
within-account join the analytics actually use and gives up only
|
||||
cross-machine joins on the same repository, which nothing computes.
|
||||
"""
|
||||
if not value:
|
||||
return ""
|
||||
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _safe_value(value: Any) -> Any:
|
||||
if isinstance(value, str):
|
||||
return memory_core.redact(value)
|
||||
@@ -145,9 +206,92 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
|
||||
return created
|
||||
|
||||
|
||||
def _install_state_path() -> Path:
|
||||
return memory_core.data_dir() / "install-state.json"
|
||||
|
||||
|
||||
def is_first_run() -> bool:
|
||||
"""Whether this machine has never recorded a plugin event before."""
|
||||
return not _identity_path().exists()
|
||||
"""Whether install has never been recorded on this machine.
|
||||
|
||||
Deliberately NOT the identity file. That file is only written by a
|
||||
successful flush, so an offline or firewalled user recorded code.install on
|
||||
every single session, forever — and every 0.2.x user recorded one on their
|
||||
first 0.3.x session because 0.2.x never wrote it at all.
|
||||
"""
|
||||
return not _install_state_path().exists()
|
||||
|
||||
|
||||
def claim_install() -> str | None:
|
||||
"""Claim the one install/upgrade record for this machine, atomically.
|
||||
|
||||
Returns the event to record ("install" or "upgrade"), or None if another
|
||||
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
|
||||
cannot both win.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
# A fresh install has an empty data directory. Anything already there —
|
||||
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
|
||||
# before the marker is created, since creating it would itself be content.
|
||||
upgrading = _data_dir_has_content()
|
||||
try:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
|
||||
except FileExistsError:
|
||||
return None
|
||||
except OSError:
|
||||
return None
|
||||
try:
|
||||
with os.fdopen(handle, "w", encoding="utf-8") as stream:
|
||||
json.dump(
|
||||
{
|
||||
"plugin_version": memory_core.PLUGIN_VERSION,
|
||||
"installed_at": memory_core.utc_now(),
|
||||
"upgraded": upgrading,
|
||||
},
|
||||
stream,
|
||||
)
|
||||
except OSError:
|
||||
pass
|
||||
return "upgrade" if upgrading else "install"
|
||||
|
||||
|
||||
def _data_dir_has_content() -> bool:
|
||||
"""Whether anything predates this session in the plugin data directory."""
|
||||
try:
|
||||
for entry in memory_core.data_dir().iterdir():
|
||||
if entry.name != "install-state.json":
|
||||
return True
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def claim_version_change() -> str | None:
|
||||
"""Return the previously recorded version if it differs, updating the marker.
|
||||
|
||||
Only meaningful once the marker exists — the first transition into 0.3.x has
|
||||
no recorded predecessor and reports "pre-0.3" instead. Claiming by rewriting
|
||||
the marker means the next session sees no change and records nothing.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
try:
|
||||
state = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return None
|
||||
if not isinstance(state, dict):
|
||||
return None
|
||||
previous = str(state.get("plugin_version") or "")
|
||||
if not previous or previous == memory_core.PLUGIN_VERSION:
|
||||
return None
|
||||
state["plugin_version"] = memory_core.PLUGIN_VERSION
|
||||
state["upgraded_at"] = memory_core.utc_now()
|
||||
try:
|
||||
temporary = path.with_suffix(f".{os.getpid()}.tmp")
|
||||
temporary.write_text(json.dumps(state), encoding="utf-8")
|
||||
temporary.replace(path)
|
||||
except OSError:
|
||||
return None
|
||||
return previous
|
||||
|
||||
|
||||
def record(
|
||||
@@ -168,19 +312,25 @@ def record(
|
||||
except OSError:
|
||||
pass
|
||||
properties = _safe_value(properties)
|
||||
# Stamped in the RECORDING process, beside harness. `source` used to be
|
||||
# read in the sending process from a module global, so whichever process
|
||||
# drained the spool named every event in it. flush() spreads per-event
|
||||
# properties last, so this now wins over any sender's default.
|
||||
properties.update(
|
||||
harness=_harness,
|
||||
source=_source_tag,
|
||||
plugin_version=memory_core.PLUGIN_VERSION,
|
||||
os=sys.platform,
|
||||
python_version=platform.python_version(),
|
||||
)
|
||||
if repo is not None:
|
||||
properties["repo_hash"] = _digest(getattr(repo, "identity", ""))
|
||||
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
|
||||
if session_id:
|
||||
properties["session_hash"] = _digest(session_id)
|
||||
properties["session_hash"] = _scoped_digest(session_id)
|
||||
line = json.dumps(
|
||||
{
|
||||
"event": f"{EVENT_PREFIX}.{event}",
|
||||
"uuid": str(uuid.uuid4()),
|
||||
"timestamp": memory_core.utc_now(),
|
||||
"properties": {
|
||||
key: value for key, value in properties.items() if value is not None
|
||||
@@ -239,38 +389,139 @@ def spawn_flush() -> bool:
|
||||
return False
|
||||
|
||||
|
||||
def _claim_name(attempt: int = 0) -> str:
|
||||
"""Claim filename. The attempt count rides in the name so the 7-day expiry
|
||||
only ever discards a batch that was actually retried and failed."""
|
||||
return f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}-a{attempt}.sending"
|
||||
|
||||
|
||||
def _claim_attempt(claim: Path) -> int:
|
||||
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
|
||||
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
|
||||
tail = stem.rsplit("-", 1)[-1]
|
||||
if tail.startswith("a") and tail[1:].isdigit():
|
||||
return int(tail[1:])
|
||||
return 0
|
||||
|
||||
|
||||
def _touch(path: Path) -> None:
|
||||
"""Refresh mtime so a claim's age measures time since it was claimed.
|
||||
|
||||
``Path.replace`` is ``os.rename``, which preserves mtime — so a claim created
|
||||
after a quiet minute inherited the spool's last-write time and looked
|
||||
abandoned the instant it was made. A second sender would then take it over
|
||||
while the first was still posting, and both would deliver the batch.
|
||||
"""
|
||||
try:
|
||||
os.utime(path, None)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _claim_spool() -> Path | None:
|
||||
"""Rename the spool aside so exactly one sender owns each batch."""
|
||||
directory = memory_core.data_dir()
|
||||
claim = directory / f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}.sending"
|
||||
claim = directory / _claim_name()
|
||||
spool = _spool_path()
|
||||
try:
|
||||
spool.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
pass
|
||||
return _claim_parked(directory)
|
||||
|
||||
|
||||
def _claim_parked(directory: Path) -> Path | None:
|
||||
"""Take the oldest abandoned claim, if any lease has actually expired.
|
||||
|
||||
Kept separate from the live spool so flush() can drain both in one run.
|
||||
Previously parked batches were only reachable when no spool existed at all,
|
||||
and because sessions keep recording there usually was one — so a batch
|
||||
parked by a failed send waited until the 7-day expiry deleted it unsent,
|
||||
even though its own presence is what started the sender.
|
||||
"""
|
||||
now = time.time()
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending")):
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending"), key=_safe_mtime):
|
||||
try:
|
||||
age = now - orphan.stat().st_mtime
|
||||
except OSError:
|
||||
continue
|
||||
if age > CLAIM_EXPIRY_SECONDS:
|
||||
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
|
||||
try:
|
||||
orphan.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
continue
|
||||
if age < CLAIM_STALE_SECONDS:
|
||||
# Someone else holds a live lease on it.
|
||||
continue
|
||||
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
|
||||
try:
|
||||
orphan.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
def _safe_mtime(path: Path) -> float:
|
||||
try:
|
||||
return path.stat().st_mtime
|
||||
except OSError:
|
||||
return 0.0
|
||||
|
||||
|
||||
def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
|
||||
"""Persist the unsent remainder, atomically, and refresh the lease.
|
||||
|
||||
Called after every successful batch. Two jobs: a retry resumes where the
|
||||
send stopped instead of re-posting from the top, and the rewrite doubles as
|
||||
the lease heartbeat, so a slow sender does not have its claim stolen
|
||||
mid-flight. Interval is one batch, well inside CLAIM_STALE_SECONDS.
|
||||
"""
|
||||
if not remaining:
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return True
|
||||
temporary = claim.with_suffix(f".{os.getpid()}.partial")
|
||||
try:
|
||||
temporary.write_text(
|
||||
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
|
||||
encoding="utf-8",
|
||||
)
|
||||
temporary.replace(claim)
|
||||
_touch(claim)
|
||||
return True
|
||||
except OSError:
|
||||
try:
|
||||
temporary.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
|
||||
"""Persist the remainder and drop the lease, because this sender has given up.
|
||||
|
||||
Distinct from the per-batch heartbeat: heartbeating on the way out would
|
||||
make an abandoned batch look actively owned for a further
|
||||
CLAIM_STALE_SECONDS, delaying the retry for no reason. Ageing it past the
|
||||
threshold lets the next flush pick it up immediately, while the attempt
|
||||
count in the filename still bounds how many times that can happen.
|
||||
"""
|
||||
if not _rewrite_claim(claim, remaining):
|
||||
return
|
||||
try:
|
||||
released = time.time() - CLAIM_STALE_SECONDS - 1
|
||||
os.utime(claim, (released, released))
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _resolve_email(key: str) -> str:
|
||||
"""Trade the API key for the account email so events join other Mem0 surfaces."""
|
||||
url = os.environ.get("MEM0_API_URL", memory_core.DEFAULT_API_URL).rstrip("/") + "/v1/ping/"
|
||||
@@ -300,34 +551,76 @@ def _post(payload: dict[str, Any], url: str) -> bool:
|
||||
|
||||
|
||||
def resolve_distinct_id() -> tuple[str, str]:
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any."""
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any.
|
||||
|
||||
The second value becomes a PostHog $identify alias. It is ONLY ever an
|
||||
anonymous id: aliasing one account email to another merges two real person
|
||||
profiles and cannot be undone, so a key that now belongs to a different
|
||||
account re-resolves with no alias.
|
||||
"""
|
||||
identity = _read_identity()
|
||||
email = identity.get("email", "")
|
||||
if email:
|
||||
return email, ""
|
||||
key = memory_core.api_key()
|
||||
fingerprint = _digest(key) if key else ""
|
||||
email = identity.get("email", "")
|
||||
|
||||
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
|
||||
return email, ""
|
||||
|
||||
if not key:
|
||||
# No key to verify the account with; do not keep attributing to it.
|
||||
if email:
|
||||
identity.pop("email", None)
|
||||
identity.pop("key_fingerprint", None)
|
||||
_write_identity(identity)
|
||||
return anonymous_id(identity), ""
|
||||
email = _resolve_email(key)
|
||||
if not email:
|
||||
return anonymous_id(identity), ""
|
||||
previous = identity.get("anonymous_id", "")
|
||||
identity["email"] = email
|
||||
|
||||
resolved = _resolve_email(key)
|
||||
if not resolved:
|
||||
return (email, "") if email else (anonymous_id(identity), "")
|
||||
|
||||
# Alias only when going anonymous -> email for the first time.
|
||||
previous = "" if email else identity.get("anonymous_id", "")
|
||||
identity["email"] = resolved
|
||||
identity["key_fingerprint"] = fingerprint
|
||||
_write_identity(identity)
|
||||
return email, previous
|
||||
return resolved, previous
|
||||
|
||||
|
||||
def flush() -> int:
|
||||
"""Drain claimed spools to PostHog and return the number of events sent."""
|
||||
"""Drain the live spool, then any parked claims, and return events sent."""
|
||||
if not is_enabled():
|
||||
return 0
|
||||
claim = _claim_spool()
|
||||
sent, delivered = _drain(_claim_spool())
|
||||
if not delivered:
|
||||
# The network is failing. Retrying other batches now would only burn
|
||||
# their attempt budget against the same broken connection.
|
||||
return sent
|
||||
|
||||
# Parked batches used to starve behind the live spool indefinitely. Bounded
|
||||
# per run so a long backlog cannot turn one flush into an unbounded loop.
|
||||
directory = memory_core.data_dir()
|
||||
for _ in range(MAX_PARKED_PER_RUN):
|
||||
parked = _claim_parked(directory)
|
||||
if parked is None:
|
||||
break
|
||||
count, delivered = _drain(parked)
|
||||
sent += count
|
||||
if not delivered:
|
||||
break
|
||||
return sent
|
||||
|
||||
|
||||
def _drain(claim: Path | None) -> tuple[int, bool]:
|
||||
"""Post one claimed batch file, recording progress after every batch.
|
||||
|
||||
Returns (events sent, whether everything was delivered).
|
||||
"""
|
||||
if claim is None:
|
||||
return 0
|
||||
return 0, True
|
||||
try:
|
||||
lines = claim.read_text(encoding="utf-8").splitlines()
|
||||
except OSError:
|
||||
return 0
|
||||
return 0, True
|
||||
events = []
|
||||
for line in lines:
|
||||
try:
|
||||
@@ -341,7 +634,7 @@ def flush() -> int:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return 0
|
||||
return 0, True
|
||||
|
||||
distinct_id, aliased_anonymous_id = resolve_distinct_id()
|
||||
if aliased_anonymous_id:
|
||||
@@ -360,12 +653,17 @@ def flush() -> int:
|
||||
|
||||
sent = 0
|
||||
for start in range(0, len(events), BATCH_SIZE):
|
||||
chunk = events[start : start + BATCH_SIZE]
|
||||
batch = [
|
||||
{
|
||||
"event": event["event"],
|
||||
"distinct_id": distinct_id,
|
||||
# Carried through from record() so a resend can be collapsed.
|
||||
"uuid": event.get("uuid"),
|
||||
"timestamp": event.get("timestamp"),
|
||||
"properties": {
|
||||
# Fallback only: events recorded by a build before source
|
||||
# moved into record() have none of their own.
|
||||
"source": _source_tag,
|
||||
"language": "python",
|
||||
"$process_person_profile": False,
|
||||
@@ -373,16 +671,19 @@ def flush() -> int:
|
||||
**(event.get("properties") or {}),
|
||||
},
|
||||
}
|
||||
for event in events[start : start + BATCH_SIZE]
|
||||
for event in chunk
|
||||
]
|
||||
if not _post({"api_key": POSTHOG_API_KEY, "batch": batch}, POSTHOG_BATCH_URL):
|
||||
return sent
|
||||
sent += len(batch)
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return sent
|
||||
# Keep only what has not been delivered, and release the lease.
|
||||
# Previously the whole file was kept and the retry re-posted every
|
||||
# batch, including the ones that had already arrived.
|
||||
_release_claim(claim, events[start:])
|
||||
return sent, False
|
||||
sent += len(chunk)
|
||||
# Record progress and refresh the lease after each successful batch, so
|
||||
# a crash repeats at most one batch instead of the entire file.
|
||||
_rewrite_claim(claim, events[start + len(chunk) :])
|
||||
return sent, True
|
||||
|
||||
|
||||
def main() -> int:
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
@@ -198,9 +199,11 @@ def test_a_stale_claim_is_reclaimed(isolated_env, monkeypatch):
|
||||
telemetry.record("search")
|
||||
orphan = telemetry._claim_spool()
|
||||
assert orphan is not None
|
||||
monkeypatch.setattr(
|
||||
telemetry.time, "time", lambda: orphan.stat().st_mtime + telemetry.CLAIM_STALE_SECONDS + 1
|
||||
)
|
||||
# Frozen rather than re-stat'd per call: flush() drains the live spool and
|
||||
# then looks for parked claims in the same run, so by the second look this
|
||||
# file no longer exists.
|
||||
stale_now = orphan.stat().st_mtime + telemetry.CLAIM_STALE_SECONDS + 1
|
||||
monkeypatch.setattr(telemetry.time, "time", lambda: stale_now)
|
||||
|
||||
with patch.object(telemetry, "_post", lambda payload, url: True):
|
||||
assert telemetry.flush() == 1
|
||||
@@ -210,9 +213,13 @@ def test_an_expired_claim_is_dropped(isolated_env, monkeypatch):
|
||||
telemetry.record("search")
|
||||
orphan = telemetry._claim_spool()
|
||||
assert orphan is not None
|
||||
monkeypatch.setattr(
|
||||
telemetry.time, "time", lambda: orphan.stat().st_mtime + telemetry.CLAIM_EXPIRY_SECONDS + 1
|
||||
)
|
||||
expired_now = orphan.stat().st_mtime + telemetry.CLAIM_EXPIRY_SECONDS + 1
|
||||
monkeypatch.setattr(telemetry.time, "time", lambda: expired_now)
|
||||
# Expiry now only discards a batch that was genuinely retried and failed,
|
||||
# so age alone is not enough — age it past the attempt budget too.
|
||||
retried = orphan.parent / orphan.name.replace("-a0.", f"-a{telemetry.MAX_CLAIM_ATTEMPTS}.")
|
||||
orphan.replace(retried)
|
||||
os.utime(retried, (expired_now, expired_now - telemetry.CLAIM_EXPIRY_SECONDS - 1))
|
||||
assert telemetry._claim_spool() is None
|
||||
assert not list(memory_core.data_dir().glob("telemetry-*.sending"))
|
||||
|
||||
@@ -258,12 +265,47 @@ def test_an_unresolvable_key_falls_back_to_the_anonymous_id(isolated_env, monkey
|
||||
assert telemetry.resolve_distinct_id()[0].startswith("code-anon-")
|
||||
|
||||
|
||||
def test_is_first_run_flips_after_the_first_identity_write(isolated_env):
|
||||
def test_first_run_is_not_flipped_by_writing_the_identity_file(isolated_env):
|
||||
"""The identity file is written by a successful flush, not by recording.
|
||||
|
||||
Keying first-run off it meant an offline user recorded code.install on every
|
||||
session forever, and every 0.2.x user recorded one on their first 0.3.x run.
|
||||
"""
|
||||
assert telemetry.is_first_run()
|
||||
telemetry.anonymous_id()
|
||||
assert telemetry.is_first_run()
|
||||
|
||||
|
||||
def test_claiming_install_ends_first_run(isolated_env):
|
||||
assert telemetry.claim_install() == "install"
|
||||
assert not telemetry.is_first_run()
|
||||
|
||||
|
||||
def test_install_can_only_be_claimed_once(isolated_env):
|
||||
"""Two sessions starting together must not both record an install."""
|
||||
assert telemetry.claim_install() == "install"
|
||||
assert telemetry.claim_install() is None
|
||||
|
||||
|
||||
def test_a_populated_data_dir_reads_as_an_upgrade(isolated_env):
|
||||
"""A fresh install has an empty data directory; anything else predates it."""
|
||||
data_dir = memory_core.data_dir()
|
||||
data_dir.mkdir(parents=True, exist_ok=True)
|
||||
(data_dir / "requirements.txt").write_text("mem0ai\n", encoding="utf-8")
|
||||
assert telemetry.claim_install() == "upgrade"
|
||||
|
||||
|
||||
def test_a_version_change_is_claimed_once(isolated_env):
|
||||
telemetry.claim_install()
|
||||
state_path = memory_core.data_dir() / "install-state.json"
|
||||
state = json.loads(state_path.read_text())
|
||||
state["plugin_version"] = "0.0.1-old"
|
||||
state_path.write_text(json.dumps(state), encoding="utf-8")
|
||||
|
||||
assert telemetry.claim_version_change() == "0.0.1-old"
|
||||
assert telemetry.claim_version_change() is None
|
||||
|
||||
|
||||
def test_spawn_flush_does_nothing_without_a_spool(isolated_env):
|
||||
with patch.object(telemetry.subprocess, "Popen") as popen:
|
||||
assert telemetry.spawn_flush() is False
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Generated by integrations/agent-plugin-core/build/build.py. Do not edit."""
|
||||
|
||||
HARNESS_ID = "codex"
|
||||
SOURCE_TAG = "CODEX_PLUGIN"
|
||||
|
||||
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
|
||||
# plugin family is one source; which editor it runs in is the application.
|
||||
PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
PLATFORM_APPLICATION = "codex"
|
||||
@@ -305,8 +305,19 @@ def run(
|
||||
return 0
|
||||
|
||||
if args.action == "session-start":
|
||||
if telemetry.is_first_run():
|
||||
# Claims the marker atomically and says which event to record, so a
|
||||
# second session starting alongside this one cannot record it too.
|
||||
first_event = telemetry.claim_install()
|
||||
if first_event == "install":
|
||||
telemetry.record("install")
|
||||
elif first_event == "upgrade":
|
||||
# First run after a build that never wrote the marker; the
|
||||
# predecessor version was never recorded anywhere.
|
||||
telemetry.record("upgrade", from_version="pre-0.3")
|
||||
else:
|
||||
previous = telemetry.claim_version_change()
|
||||
if previous:
|
||||
telemetry.record("upgrade", from_version=previous)
|
||||
recovered = recover_pending_handoffs()
|
||||
record_session_start(store, hook_input)
|
||||
if recovered:
|
||||
|
||||
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
|
||||
return batches
|
||||
|
||||
|
||||
# Platform surface attribution. Read from the generated per-host module so a new
|
||||
# entrypoint is correct without remembering to configure anything.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
except ImportError:
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
|
||||
def platform_headers(key: str) -> dict[str, str]:
|
||||
"""Auth plus the three surface-identity headers.
|
||||
|
||||
X-Mem0-Source and X-Application are set-once by contract: this is the
|
||||
outermost layer, so it sets them, and nothing below may overwrite them.
|
||||
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
|
||||
"""
|
||||
headers = {
|
||||
"Authorization": f"Token {key}",
|
||||
"Content-Type": "application/json",
|
||||
"X-Mem0-Source": _PLATFORM_SOURCE,
|
||||
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
|
||||
}
|
||||
if _PLATFORM_APPLICATION:
|
||||
headers["X-Application"] = _PLATFORM_APPLICATION
|
||||
return headers
|
||||
|
||||
|
||||
def _request_json(
|
||||
url: str, key: str, payload: dict[str, Any], timeout: float
|
||||
) -> tuple[dict[str, Any] | list[Any], int, int]:
|
||||
@@ -1807,7 +1835,7 @@ def _request_json(
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
data=raw,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="POST",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1834,7 +1862,7 @@ def _get_json(
|
||||
) -> tuple[dict[str, Any] | list[Any], int]:
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="GET",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1980,6 +2008,10 @@ def flush_session(
|
||||
"user_id": write_user,
|
||||
"app_id": repo.app_id,
|
||||
"run_id": session_id,
|
||||
# Top level, not metadata: the backend reads `source` from the body,
|
||||
# query string or X-Mem0-Source header, never from metadata. The
|
||||
# harness tag stays in metadata as hook provenance.
|
||||
"source": _PLATFORM_SOURCE,
|
||||
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
|
||||
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
|
||||
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
|
||||
@@ -2523,7 +2555,7 @@ def _collect_memory_ids(
|
||||
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
|
||||
request = urllib.request.Request(
|
||||
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="DELETE",
|
||||
)
|
||||
try:
|
||||
|
||||
@@ -1,5 +1,9 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Anonymous usage telemetry for Mem0 agent plugins.
|
||||
"""Usage telemetry for Mem0 agent plugins.
|
||||
|
||||
Events are linked to your Mem0 account email when an API key is configured, and
|
||||
to a random per-machine id otherwise. Not anonymous — the Python SDK and CLI
|
||||
attribute the same way.
|
||||
|
||||
Hooks run on a 3-6 second budget and fire on every tool call, so recording never
|
||||
touches the network: `record` appends one JSON line to a local spool and returns.
|
||||
@@ -9,7 +13,8 @@ started once per session and again from the flush worker that is already detache
|
||||
Pure stdlib, matching the rest of the plugin. Opt out with MEM0_TELEMETRY=false.
|
||||
|
||||
Never sends prompts, memory text, queries, file paths, repository names, or API
|
||||
keys: only event names, durations, counts, coarse outcomes, and salted hashes.
|
||||
keys: only event names, durations, counts, coarse outcomes, and repo/session
|
||||
identifiers hashed with a random per-install salt.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -29,8 +34,23 @@ from typing import Any
|
||||
|
||||
import memory_core
|
||||
|
||||
_harness: str = "generic"
|
||||
_source_tag: str = "MEM0_PLUGIN"
|
||||
# Seeded from the per-host module the build generates into core/. Two processes
|
||||
# in this pipeline never call init() — mcp_server.py, and the detached
|
||||
# `python3 telemetry.py` sender that spawn_flush() starts — so a module default
|
||||
# was what every one of their events got labelled with.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import HARNESS_ID as _DEFAULT_HARNESS
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
from _harness_id import SOURCE_TAG as _DEFAULT_SOURCE_TAG
|
||||
except ImportError:
|
||||
_DEFAULT_HARNESS = "generic"
|
||||
_DEFAULT_SOURCE_TAG = "MEM0_PLUGIN"
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
_harness: str = _DEFAULT_HARNESS
|
||||
_source_tag: str = _DEFAULT_SOURCE_TAG
|
||||
_PRIVATE_KEYS = {
|
||||
"apikey",
|
||||
"authorization",
|
||||
@@ -56,10 +76,19 @@ _PRIVATE_KEYS = {
|
||||
}
|
||||
|
||||
|
||||
def init(harness: str = "generic", source_tag: str = "") -> None:
|
||||
def init(harness: str = "", source_tag: str = "") -> None:
|
||||
"""Override the generated identity. Optional — core/_harness_id.py is the default.
|
||||
|
||||
The fallback shape matches memory_core.configure_harness's (``<HOST>_PLUGIN``).
|
||||
It used to be ``MEM0_<HOST>_PLUGIN`` here and ``<host>_plugin`` there, which
|
||||
meant one plugin could emit three different source values depending on which
|
||||
process happened to send the batch.
|
||||
"""
|
||||
global _harness, _source_tag
|
||||
_harness = harness
|
||||
_source_tag = source_tag or f"MEM0_{harness.upper().replace('-', '_')}_PLUGIN"
|
||||
_harness = harness or _DEFAULT_HARNESS
|
||||
_source_tag = source_tag or (
|
||||
f"{_harness.upper().replace('-', '_')}_PLUGIN" if harness else _DEFAULT_SOURCE_TAG
|
||||
)
|
||||
|
||||
POSTHOG_API_KEY = "phc_hgJkUVJFYtmaJqrvf6CYN67TIQ8yhXAkWzUn9AMU4yX"
|
||||
POSTHOG_CAPTURE_URL = "https://us.i.posthog.com/i/v0/e/"
|
||||
@@ -70,6 +99,11 @@ BATCH_SIZE = 100
|
||||
SEND_TIMEOUT = 5
|
||||
CLAIM_STALE_SECONDS = 120
|
||||
CLAIM_EXPIRY_SECONDS = 7 * 24 * 60 * 60
|
||||
# A batch is only discarded once it has genuinely been retried this many times.
|
||||
MAX_CLAIM_ATTEMPTS = 3
|
||||
# Parked claims drained per run, after the live spool. Bounded so a long backlog
|
||||
# cannot turn one flush into an unbounded send loop.
|
||||
MAX_PARKED_PER_RUN = 3
|
||||
|
||||
|
||||
def is_enabled() -> bool:
|
||||
@@ -83,9 +117,36 @@ def is_enabled() -> bool:
|
||||
|
||||
|
||||
def _digest(value: str, length: int = 16) -> str:
|
||||
"""Unsalted digest. Only for values that are already secrets (API keys)."""
|
||||
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _install_salt() -> str:
|
||||
"""Random per-install salt, created on first use and kept in the identity file."""
|
||||
identity = _read_identity()
|
||||
salt = identity.get("salt")
|
||||
if not salt:
|
||||
salt = uuid.uuid4().hex
|
||||
identity["salt"] = salt
|
||||
_write_identity(identity)
|
||||
return salt
|
||||
|
||||
|
||||
def _scoped_digest(value: str, length: int = 16) -> str:
|
||||
"""Salted digest for values drawn from a guessable space.
|
||||
|
||||
repo.identity is a git remote URL, or ``local:<absolute path>`` when there is
|
||||
no remote — which normally contains the account username. Sixteen unsalted
|
||||
hex characters over that input space is enumerable, so this is not a
|
||||
privacy control without the salt. Salting per install keeps every
|
||||
within-account join the analytics actually use and gives up only
|
||||
cross-machine joins on the same repository, which nothing computes.
|
||||
"""
|
||||
if not value:
|
||||
return ""
|
||||
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _safe_value(value: Any) -> Any:
|
||||
if isinstance(value, str):
|
||||
return memory_core.redact(value)
|
||||
@@ -145,9 +206,92 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
|
||||
return created
|
||||
|
||||
|
||||
def _install_state_path() -> Path:
|
||||
return memory_core.data_dir() / "install-state.json"
|
||||
|
||||
|
||||
def is_first_run() -> bool:
|
||||
"""Whether this machine has never recorded a plugin event before."""
|
||||
return not _identity_path().exists()
|
||||
"""Whether install has never been recorded on this machine.
|
||||
|
||||
Deliberately NOT the identity file. That file is only written by a
|
||||
successful flush, so an offline or firewalled user recorded code.install on
|
||||
every single session, forever — and every 0.2.x user recorded one on their
|
||||
first 0.3.x session because 0.2.x never wrote it at all.
|
||||
"""
|
||||
return not _install_state_path().exists()
|
||||
|
||||
|
||||
def claim_install() -> str | None:
|
||||
"""Claim the one install/upgrade record for this machine, atomically.
|
||||
|
||||
Returns the event to record ("install" or "upgrade"), or None if another
|
||||
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
|
||||
cannot both win.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
# A fresh install has an empty data directory. Anything already there —
|
||||
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
|
||||
# before the marker is created, since creating it would itself be content.
|
||||
upgrading = _data_dir_has_content()
|
||||
try:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
|
||||
except FileExistsError:
|
||||
return None
|
||||
except OSError:
|
||||
return None
|
||||
try:
|
||||
with os.fdopen(handle, "w", encoding="utf-8") as stream:
|
||||
json.dump(
|
||||
{
|
||||
"plugin_version": memory_core.PLUGIN_VERSION,
|
||||
"installed_at": memory_core.utc_now(),
|
||||
"upgraded": upgrading,
|
||||
},
|
||||
stream,
|
||||
)
|
||||
except OSError:
|
||||
pass
|
||||
return "upgrade" if upgrading else "install"
|
||||
|
||||
|
||||
def _data_dir_has_content() -> bool:
|
||||
"""Whether anything predates this session in the plugin data directory."""
|
||||
try:
|
||||
for entry in memory_core.data_dir().iterdir():
|
||||
if entry.name != "install-state.json":
|
||||
return True
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def claim_version_change() -> str | None:
|
||||
"""Return the previously recorded version if it differs, updating the marker.
|
||||
|
||||
Only meaningful once the marker exists — the first transition into 0.3.x has
|
||||
no recorded predecessor and reports "pre-0.3" instead. Claiming by rewriting
|
||||
the marker means the next session sees no change and records nothing.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
try:
|
||||
state = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return None
|
||||
if not isinstance(state, dict):
|
||||
return None
|
||||
previous = str(state.get("plugin_version") or "")
|
||||
if not previous or previous == memory_core.PLUGIN_VERSION:
|
||||
return None
|
||||
state["plugin_version"] = memory_core.PLUGIN_VERSION
|
||||
state["upgraded_at"] = memory_core.utc_now()
|
||||
try:
|
||||
temporary = path.with_suffix(f".{os.getpid()}.tmp")
|
||||
temporary.write_text(json.dumps(state), encoding="utf-8")
|
||||
temporary.replace(path)
|
||||
except OSError:
|
||||
return None
|
||||
return previous
|
||||
|
||||
|
||||
def record(
|
||||
@@ -168,19 +312,25 @@ def record(
|
||||
except OSError:
|
||||
pass
|
||||
properties = _safe_value(properties)
|
||||
# Stamped in the RECORDING process, beside harness. `source` used to be
|
||||
# read in the sending process from a module global, so whichever process
|
||||
# drained the spool named every event in it. flush() spreads per-event
|
||||
# properties last, so this now wins over any sender's default.
|
||||
properties.update(
|
||||
harness=_harness,
|
||||
source=_source_tag,
|
||||
plugin_version=memory_core.PLUGIN_VERSION,
|
||||
os=sys.platform,
|
||||
python_version=platform.python_version(),
|
||||
)
|
||||
if repo is not None:
|
||||
properties["repo_hash"] = _digest(getattr(repo, "identity", ""))
|
||||
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
|
||||
if session_id:
|
||||
properties["session_hash"] = _digest(session_id)
|
||||
properties["session_hash"] = _scoped_digest(session_id)
|
||||
line = json.dumps(
|
||||
{
|
||||
"event": f"{EVENT_PREFIX}.{event}",
|
||||
"uuid": str(uuid.uuid4()),
|
||||
"timestamp": memory_core.utc_now(),
|
||||
"properties": {
|
||||
key: value for key, value in properties.items() if value is not None
|
||||
@@ -239,38 +389,139 @@ def spawn_flush() -> bool:
|
||||
return False
|
||||
|
||||
|
||||
def _claim_name(attempt: int = 0) -> str:
|
||||
"""Claim filename. The attempt count rides in the name so the 7-day expiry
|
||||
only ever discards a batch that was actually retried and failed."""
|
||||
return f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}-a{attempt}.sending"
|
||||
|
||||
|
||||
def _claim_attempt(claim: Path) -> int:
|
||||
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
|
||||
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
|
||||
tail = stem.rsplit("-", 1)[-1]
|
||||
if tail.startswith("a") and tail[1:].isdigit():
|
||||
return int(tail[1:])
|
||||
return 0
|
||||
|
||||
|
||||
def _touch(path: Path) -> None:
|
||||
"""Refresh mtime so a claim's age measures time since it was claimed.
|
||||
|
||||
``Path.replace`` is ``os.rename``, which preserves mtime — so a claim created
|
||||
after a quiet minute inherited the spool's last-write time and looked
|
||||
abandoned the instant it was made. A second sender would then take it over
|
||||
while the first was still posting, and both would deliver the batch.
|
||||
"""
|
||||
try:
|
||||
os.utime(path, None)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _claim_spool() -> Path | None:
|
||||
"""Rename the spool aside so exactly one sender owns each batch."""
|
||||
directory = memory_core.data_dir()
|
||||
claim = directory / f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}.sending"
|
||||
claim = directory / _claim_name()
|
||||
spool = _spool_path()
|
||||
try:
|
||||
spool.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
pass
|
||||
return _claim_parked(directory)
|
||||
|
||||
|
||||
def _claim_parked(directory: Path) -> Path | None:
|
||||
"""Take the oldest abandoned claim, if any lease has actually expired.
|
||||
|
||||
Kept separate from the live spool so flush() can drain both in one run.
|
||||
Previously parked batches were only reachable when no spool existed at all,
|
||||
and because sessions keep recording there usually was one — so a batch
|
||||
parked by a failed send waited until the 7-day expiry deleted it unsent,
|
||||
even though its own presence is what started the sender.
|
||||
"""
|
||||
now = time.time()
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending")):
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending"), key=_safe_mtime):
|
||||
try:
|
||||
age = now - orphan.stat().st_mtime
|
||||
except OSError:
|
||||
continue
|
||||
if age > CLAIM_EXPIRY_SECONDS:
|
||||
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
|
||||
try:
|
||||
orphan.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
continue
|
||||
if age < CLAIM_STALE_SECONDS:
|
||||
# Someone else holds a live lease on it.
|
||||
continue
|
||||
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
|
||||
try:
|
||||
orphan.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
def _safe_mtime(path: Path) -> float:
|
||||
try:
|
||||
return path.stat().st_mtime
|
||||
except OSError:
|
||||
return 0.0
|
||||
|
||||
|
||||
def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
|
||||
"""Persist the unsent remainder, atomically, and refresh the lease.
|
||||
|
||||
Called after every successful batch. Two jobs: a retry resumes where the
|
||||
send stopped instead of re-posting from the top, and the rewrite doubles as
|
||||
the lease heartbeat, so a slow sender does not have its claim stolen
|
||||
mid-flight. Interval is one batch, well inside CLAIM_STALE_SECONDS.
|
||||
"""
|
||||
if not remaining:
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return True
|
||||
temporary = claim.with_suffix(f".{os.getpid()}.partial")
|
||||
try:
|
||||
temporary.write_text(
|
||||
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
|
||||
encoding="utf-8",
|
||||
)
|
||||
temporary.replace(claim)
|
||||
_touch(claim)
|
||||
return True
|
||||
except OSError:
|
||||
try:
|
||||
temporary.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
|
||||
"""Persist the remainder and drop the lease, because this sender has given up.
|
||||
|
||||
Distinct from the per-batch heartbeat: heartbeating on the way out would
|
||||
make an abandoned batch look actively owned for a further
|
||||
CLAIM_STALE_SECONDS, delaying the retry for no reason. Ageing it past the
|
||||
threshold lets the next flush pick it up immediately, while the attempt
|
||||
count in the filename still bounds how many times that can happen.
|
||||
"""
|
||||
if not _rewrite_claim(claim, remaining):
|
||||
return
|
||||
try:
|
||||
released = time.time() - CLAIM_STALE_SECONDS - 1
|
||||
os.utime(claim, (released, released))
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _resolve_email(key: str) -> str:
|
||||
"""Trade the API key for the account email so events join other Mem0 surfaces."""
|
||||
url = os.environ.get("MEM0_API_URL", memory_core.DEFAULT_API_URL).rstrip("/") + "/v1/ping/"
|
||||
@@ -300,34 +551,76 @@ def _post(payload: dict[str, Any], url: str) -> bool:
|
||||
|
||||
|
||||
def resolve_distinct_id() -> tuple[str, str]:
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any."""
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any.
|
||||
|
||||
The second value becomes a PostHog $identify alias. It is ONLY ever an
|
||||
anonymous id: aliasing one account email to another merges two real person
|
||||
profiles and cannot be undone, so a key that now belongs to a different
|
||||
account re-resolves with no alias.
|
||||
"""
|
||||
identity = _read_identity()
|
||||
email = identity.get("email", "")
|
||||
if email:
|
||||
return email, ""
|
||||
key = memory_core.api_key()
|
||||
fingerprint = _digest(key) if key else ""
|
||||
email = identity.get("email", "")
|
||||
|
||||
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
|
||||
return email, ""
|
||||
|
||||
if not key:
|
||||
# No key to verify the account with; do not keep attributing to it.
|
||||
if email:
|
||||
identity.pop("email", None)
|
||||
identity.pop("key_fingerprint", None)
|
||||
_write_identity(identity)
|
||||
return anonymous_id(identity), ""
|
||||
email = _resolve_email(key)
|
||||
if not email:
|
||||
return anonymous_id(identity), ""
|
||||
previous = identity.get("anonymous_id", "")
|
||||
identity["email"] = email
|
||||
|
||||
resolved = _resolve_email(key)
|
||||
if not resolved:
|
||||
return (email, "") if email else (anonymous_id(identity), "")
|
||||
|
||||
# Alias only when going anonymous -> email for the first time.
|
||||
previous = "" if email else identity.get("anonymous_id", "")
|
||||
identity["email"] = resolved
|
||||
identity["key_fingerprint"] = fingerprint
|
||||
_write_identity(identity)
|
||||
return email, previous
|
||||
return resolved, previous
|
||||
|
||||
|
||||
def flush() -> int:
|
||||
"""Drain claimed spools to PostHog and return the number of events sent."""
|
||||
"""Drain the live spool, then any parked claims, and return events sent."""
|
||||
if not is_enabled():
|
||||
return 0
|
||||
claim = _claim_spool()
|
||||
sent, delivered = _drain(_claim_spool())
|
||||
if not delivered:
|
||||
# The network is failing. Retrying other batches now would only burn
|
||||
# their attempt budget against the same broken connection.
|
||||
return sent
|
||||
|
||||
# Parked batches used to starve behind the live spool indefinitely. Bounded
|
||||
# per run so a long backlog cannot turn one flush into an unbounded loop.
|
||||
directory = memory_core.data_dir()
|
||||
for _ in range(MAX_PARKED_PER_RUN):
|
||||
parked = _claim_parked(directory)
|
||||
if parked is None:
|
||||
break
|
||||
count, delivered = _drain(parked)
|
||||
sent += count
|
||||
if not delivered:
|
||||
break
|
||||
return sent
|
||||
|
||||
|
||||
def _drain(claim: Path | None) -> tuple[int, bool]:
|
||||
"""Post one claimed batch file, recording progress after every batch.
|
||||
|
||||
Returns (events sent, whether everything was delivered).
|
||||
"""
|
||||
if claim is None:
|
||||
return 0
|
||||
return 0, True
|
||||
try:
|
||||
lines = claim.read_text(encoding="utf-8").splitlines()
|
||||
except OSError:
|
||||
return 0
|
||||
return 0, True
|
||||
events = []
|
||||
for line in lines:
|
||||
try:
|
||||
@@ -341,7 +634,7 @@ def flush() -> int:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return 0
|
||||
return 0, True
|
||||
|
||||
distinct_id, aliased_anonymous_id = resolve_distinct_id()
|
||||
if aliased_anonymous_id:
|
||||
@@ -360,12 +653,17 @@ def flush() -> int:
|
||||
|
||||
sent = 0
|
||||
for start in range(0, len(events), BATCH_SIZE):
|
||||
chunk = events[start : start + BATCH_SIZE]
|
||||
batch = [
|
||||
{
|
||||
"event": event["event"],
|
||||
"distinct_id": distinct_id,
|
||||
# Carried through from record() so a resend can be collapsed.
|
||||
"uuid": event.get("uuid"),
|
||||
"timestamp": event.get("timestamp"),
|
||||
"properties": {
|
||||
# Fallback only: events recorded by a build before source
|
||||
# moved into record() have none of their own.
|
||||
"source": _source_tag,
|
||||
"language": "python",
|
||||
"$process_person_profile": False,
|
||||
@@ -373,16 +671,19 @@ def flush() -> int:
|
||||
**(event.get("properties") or {}),
|
||||
},
|
||||
}
|
||||
for event in events[start : start + BATCH_SIZE]
|
||||
for event in chunk
|
||||
]
|
||||
if not _post({"api_key": POSTHOG_API_KEY, "batch": batch}, POSTHOG_BATCH_URL):
|
||||
return sent
|
||||
sent += len(batch)
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return sent
|
||||
# Keep only what has not been delivered, and release the lease.
|
||||
# Previously the whole file was kept and the retry re-posted every
|
||||
# batch, including the ones that had already arrived.
|
||||
_release_claim(claim, events[start:])
|
||||
return sent, False
|
||||
sent += len(chunk)
|
||||
# Record progress and refresh the lease after each successful batch, so
|
||||
# a crash repeats at most one batch instead of the entire file.
|
||||
_rewrite_claim(claim, events[start + len(chunk) :])
|
||||
return sent, True
|
||||
|
||||
|
||||
def main() -> int:
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Generated by integrations/agent-plugin-core/build/build.py. Do not edit."""
|
||||
|
||||
HARNESS_ID = "cursor"
|
||||
SOURCE_TAG = "CURSOR_PLUGIN"
|
||||
|
||||
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
|
||||
# plugin family is one source; which editor it runs in is the application.
|
||||
PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
PLATFORM_APPLICATION = "cursor"
|
||||
@@ -305,8 +305,19 @@ def run(
|
||||
return 0
|
||||
|
||||
if args.action == "session-start":
|
||||
if telemetry.is_first_run():
|
||||
# Claims the marker atomically and says which event to record, so a
|
||||
# second session starting alongside this one cannot record it too.
|
||||
first_event = telemetry.claim_install()
|
||||
if first_event == "install":
|
||||
telemetry.record("install")
|
||||
elif first_event == "upgrade":
|
||||
# First run after a build that never wrote the marker; the
|
||||
# predecessor version was never recorded anywhere.
|
||||
telemetry.record("upgrade", from_version="pre-0.3")
|
||||
else:
|
||||
previous = telemetry.claim_version_change()
|
||||
if previous:
|
||||
telemetry.record("upgrade", from_version=previous)
|
||||
recovered = recover_pending_handoffs()
|
||||
record_session_start(store, hook_input)
|
||||
if recovered:
|
||||
|
||||
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
|
||||
return batches
|
||||
|
||||
|
||||
# Platform surface attribution. Read from the generated per-host module so a new
|
||||
# entrypoint is correct without remembering to configure anything.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
except ImportError:
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
|
||||
def platform_headers(key: str) -> dict[str, str]:
|
||||
"""Auth plus the three surface-identity headers.
|
||||
|
||||
X-Mem0-Source and X-Application are set-once by contract: this is the
|
||||
outermost layer, so it sets them, and nothing below may overwrite them.
|
||||
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
|
||||
"""
|
||||
headers = {
|
||||
"Authorization": f"Token {key}",
|
||||
"Content-Type": "application/json",
|
||||
"X-Mem0-Source": _PLATFORM_SOURCE,
|
||||
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
|
||||
}
|
||||
if _PLATFORM_APPLICATION:
|
||||
headers["X-Application"] = _PLATFORM_APPLICATION
|
||||
return headers
|
||||
|
||||
|
||||
def _request_json(
|
||||
url: str, key: str, payload: dict[str, Any], timeout: float
|
||||
) -> tuple[dict[str, Any] | list[Any], int, int]:
|
||||
@@ -1807,7 +1835,7 @@ def _request_json(
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
data=raw,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="POST",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1834,7 +1862,7 @@ def _get_json(
|
||||
) -> tuple[dict[str, Any] | list[Any], int]:
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="GET",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1980,6 +2008,10 @@ def flush_session(
|
||||
"user_id": write_user,
|
||||
"app_id": repo.app_id,
|
||||
"run_id": session_id,
|
||||
# Top level, not metadata: the backend reads `source` from the body,
|
||||
# query string or X-Mem0-Source header, never from metadata. The
|
||||
# harness tag stays in metadata as hook provenance.
|
||||
"source": _PLATFORM_SOURCE,
|
||||
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
|
||||
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
|
||||
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
|
||||
@@ -2523,7 +2555,7 @@ def _collect_memory_ids(
|
||||
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
|
||||
request = urllib.request.Request(
|
||||
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="DELETE",
|
||||
)
|
||||
try:
|
||||
|
||||
@@ -1,5 +1,9 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Anonymous usage telemetry for Mem0 agent plugins.
|
||||
"""Usage telemetry for Mem0 agent plugins.
|
||||
|
||||
Events are linked to your Mem0 account email when an API key is configured, and
|
||||
to a random per-machine id otherwise. Not anonymous — the Python SDK and CLI
|
||||
attribute the same way.
|
||||
|
||||
Hooks run on a 3-6 second budget and fire on every tool call, so recording never
|
||||
touches the network: `record` appends one JSON line to a local spool and returns.
|
||||
@@ -9,7 +13,8 @@ started once per session and again from the flush worker that is already detache
|
||||
Pure stdlib, matching the rest of the plugin. Opt out with MEM0_TELEMETRY=false.
|
||||
|
||||
Never sends prompts, memory text, queries, file paths, repository names, or API
|
||||
keys: only event names, durations, counts, coarse outcomes, and salted hashes.
|
||||
keys: only event names, durations, counts, coarse outcomes, and repo/session
|
||||
identifiers hashed with a random per-install salt.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -29,8 +34,23 @@ from typing import Any
|
||||
|
||||
import memory_core
|
||||
|
||||
_harness: str = "generic"
|
||||
_source_tag: str = "MEM0_PLUGIN"
|
||||
# Seeded from the per-host module the build generates into core/. Two processes
|
||||
# in this pipeline never call init() — mcp_server.py, and the detached
|
||||
# `python3 telemetry.py` sender that spawn_flush() starts — so a module default
|
||||
# was what every one of their events got labelled with.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import HARNESS_ID as _DEFAULT_HARNESS
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
from _harness_id import SOURCE_TAG as _DEFAULT_SOURCE_TAG
|
||||
except ImportError:
|
||||
_DEFAULT_HARNESS = "generic"
|
||||
_DEFAULT_SOURCE_TAG = "MEM0_PLUGIN"
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
_harness: str = _DEFAULT_HARNESS
|
||||
_source_tag: str = _DEFAULT_SOURCE_TAG
|
||||
_PRIVATE_KEYS = {
|
||||
"apikey",
|
||||
"authorization",
|
||||
@@ -56,10 +76,19 @@ _PRIVATE_KEYS = {
|
||||
}
|
||||
|
||||
|
||||
def init(harness: str = "generic", source_tag: str = "") -> None:
|
||||
def init(harness: str = "", source_tag: str = "") -> None:
|
||||
"""Override the generated identity. Optional — core/_harness_id.py is the default.
|
||||
|
||||
The fallback shape matches memory_core.configure_harness's (``<HOST>_PLUGIN``).
|
||||
It used to be ``MEM0_<HOST>_PLUGIN`` here and ``<host>_plugin`` there, which
|
||||
meant one plugin could emit three different source values depending on which
|
||||
process happened to send the batch.
|
||||
"""
|
||||
global _harness, _source_tag
|
||||
_harness = harness
|
||||
_source_tag = source_tag or f"MEM0_{harness.upper().replace('-', '_')}_PLUGIN"
|
||||
_harness = harness or _DEFAULT_HARNESS
|
||||
_source_tag = source_tag or (
|
||||
f"{_harness.upper().replace('-', '_')}_PLUGIN" if harness else _DEFAULT_SOURCE_TAG
|
||||
)
|
||||
|
||||
POSTHOG_API_KEY = "phc_hgJkUVJFYtmaJqrvf6CYN67TIQ8yhXAkWzUn9AMU4yX"
|
||||
POSTHOG_CAPTURE_URL = "https://us.i.posthog.com/i/v0/e/"
|
||||
@@ -70,6 +99,11 @@ BATCH_SIZE = 100
|
||||
SEND_TIMEOUT = 5
|
||||
CLAIM_STALE_SECONDS = 120
|
||||
CLAIM_EXPIRY_SECONDS = 7 * 24 * 60 * 60
|
||||
# A batch is only discarded once it has genuinely been retried this many times.
|
||||
MAX_CLAIM_ATTEMPTS = 3
|
||||
# Parked claims drained per run, after the live spool. Bounded so a long backlog
|
||||
# cannot turn one flush into an unbounded send loop.
|
||||
MAX_PARKED_PER_RUN = 3
|
||||
|
||||
|
||||
def is_enabled() -> bool:
|
||||
@@ -83,9 +117,36 @@ def is_enabled() -> bool:
|
||||
|
||||
|
||||
def _digest(value: str, length: int = 16) -> str:
|
||||
"""Unsalted digest. Only for values that are already secrets (API keys)."""
|
||||
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _install_salt() -> str:
|
||||
"""Random per-install salt, created on first use and kept in the identity file."""
|
||||
identity = _read_identity()
|
||||
salt = identity.get("salt")
|
||||
if not salt:
|
||||
salt = uuid.uuid4().hex
|
||||
identity["salt"] = salt
|
||||
_write_identity(identity)
|
||||
return salt
|
||||
|
||||
|
||||
def _scoped_digest(value: str, length: int = 16) -> str:
|
||||
"""Salted digest for values drawn from a guessable space.
|
||||
|
||||
repo.identity is a git remote URL, or ``local:<absolute path>`` when there is
|
||||
no remote — which normally contains the account username. Sixteen unsalted
|
||||
hex characters over that input space is enumerable, so this is not a
|
||||
privacy control without the salt. Salting per install keeps every
|
||||
within-account join the analytics actually use and gives up only
|
||||
cross-machine joins on the same repository, which nothing computes.
|
||||
"""
|
||||
if not value:
|
||||
return ""
|
||||
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _safe_value(value: Any) -> Any:
|
||||
if isinstance(value, str):
|
||||
return memory_core.redact(value)
|
||||
@@ -145,9 +206,92 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
|
||||
return created
|
||||
|
||||
|
||||
def _install_state_path() -> Path:
|
||||
return memory_core.data_dir() / "install-state.json"
|
||||
|
||||
|
||||
def is_first_run() -> bool:
|
||||
"""Whether this machine has never recorded a plugin event before."""
|
||||
return not _identity_path().exists()
|
||||
"""Whether install has never been recorded on this machine.
|
||||
|
||||
Deliberately NOT the identity file. That file is only written by a
|
||||
successful flush, so an offline or firewalled user recorded code.install on
|
||||
every single session, forever — and every 0.2.x user recorded one on their
|
||||
first 0.3.x session because 0.2.x never wrote it at all.
|
||||
"""
|
||||
return not _install_state_path().exists()
|
||||
|
||||
|
||||
def claim_install() -> str | None:
|
||||
"""Claim the one install/upgrade record for this machine, atomically.
|
||||
|
||||
Returns the event to record ("install" or "upgrade"), or None if another
|
||||
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
|
||||
cannot both win.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
# A fresh install has an empty data directory. Anything already there —
|
||||
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
|
||||
# before the marker is created, since creating it would itself be content.
|
||||
upgrading = _data_dir_has_content()
|
||||
try:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
|
||||
except FileExistsError:
|
||||
return None
|
||||
except OSError:
|
||||
return None
|
||||
try:
|
||||
with os.fdopen(handle, "w", encoding="utf-8") as stream:
|
||||
json.dump(
|
||||
{
|
||||
"plugin_version": memory_core.PLUGIN_VERSION,
|
||||
"installed_at": memory_core.utc_now(),
|
||||
"upgraded": upgrading,
|
||||
},
|
||||
stream,
|
||||
)
|
||||
except OSError:
|
||||
pass
|
||||
return "upgrade" if upgrading else "install"
|
||||
|
||||
|
||||
def _data_dir_has_content() -> bool:
|
||||
"""Whether anything predates this session in the plugin data directory."""
|
||||
try:
|
||||
for entry in memory_core.data_dir().iterdir():
|
||||
if entry.name != "install-state.json":
|
||||
return True
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def claim_version_change() -> str | None:
|
||||
"""Return the previously recorded version if it differs, updating the marker.
|
||||
|
||||
Only meaningful once the marker exists — the first transition into 0.3.x has
|
||||
no recorded predecessor and reports "pre-0.3" instead. Claiming by rewriting
|
||||
the marker means the next session sees no change and records nothing.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
try:
|
||||
state = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return None
|
||||
if not isinstance(state, dict):
|
||||
return None
|
||||
previous = str(state.get("plugin_version") or "")
|
||||
if not previous or previous == memory_core.PLUGIN_VERSION:
|
||||
return None
|
||||
state["plugin_version"] = memory_core.PLUGIN_VERSION
|
||||
state["upgraded_at"] = memory_core.utc_now()
|
||||
try:
|
||||
temporary = path.with_suffix(f".{os.getpid()}.tmp")
|
||||
temporary.write_text(json.dumps(state), encoding="utf-8")
|
||||
temporary.replace(path)
|
||||
except OSError:
|
||||
return None
|
||||
return previous
|
||||
|
||||
|
||||
def record(
|
||||
@@ -168,19 +312,25 @@ def record(
|
||||
except OSError:
|
||||
pass
|
||||
properties = _safe_value(properties)
|
||||
# Stamped in the RECORDING process, beside harness. `source` used to be
|
||||
# read in the sending process from a module global, so whichever process
|
||||
# drained the spool named every event in it. flush() spreads per-event
|
||||
# properties last, so this now wins over any sender's default.
|
||||
properties.update(
|
||||
harness=_harness,
|
||||
source=_source_tag,
|
||||
plugin_version=memory_core.PLUGIN_VERSION,
|
||||
os=sys.platform,
|
||||
python_version=platform.python_version(),
|
||||
)
|
||||
if repo is not None:
|
||||
properties["repo_hash"] = _digest(getattr(repo, "identity", ""))
|
||||
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
|
||||
if session_id:
|
||||
properties["session_hash"] = _digest(session_id)
|
||||
properties["session_hash"] = _scoped_digest(session_id)
|
||||
line = json.dumps(
|
||||
{
|
||||
"event": f"{EVENT_PREFIX}.{event}",
|
||||
"uuid": str(uuid.uuid4()),
|
||||
"timestamp": memory_core.utc_now(),
|
||||
"properties": {
|
||||
key: value for key, value in properties.items() if value is not None
|
||||
@@ -239,38 +389,139 @@ def spawn_flush() -> bool:
|
||||
return False
|
||||
|
||||
|
||||
def _claim_name(attempt: int = 0) -> str:
|
||||
"""Claim filename. The attempt count rides in the name so the 7-day expiry
|
||||
only ever discards a batch that was actually retried and failed."""
|
||||
return f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}-a{attempt}.sending"
|
||||
|
||||
|
||||
def _claim_attempt(claim: Path) -> int:
|
||||
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
|
||||
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
|
||||
tail = stem.rsplit("-", 1)[-1]
|
||||
if tail.startswith("a") and tail[1:].isdigit():
|
||||
return int(tail[1:])
|
||||
return 0
|
||||
|
||||
|
||||
def _touch(path: Path) -> None:
|
||||
"""Refresh mtime so a claim's age measures time since it was claimed.
|
||||
|
||||
``Path.replace`` is ``os.rename``, which preserves mtime — so a claim created
|
||||
after a quiet minute inherited the spool's last-write time and looked
|
||||
abandoned the instant it was made. A second sender would then take it over
|
||||
while the first was still posting, and both would deliver the batch.
|
||||
"""
|
||||
try:
|
||||
os.utime(path, None)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _claim_spool() -> Path | None:
|
||||
"""Rename the spool aside so exactly one sender owns each batch."""
|
||||
directory = memory_core.data_dir()
|
||||
claim = directory / f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}.sending"
|
||||
claim = directory / _claim_name()
|
||||
spool = _spool_path()
|
||||
try:
|
||||
spool.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
pass
|
||||
return _claim_parked(directory)
|
||||
|
||||
|
||||
def _claim_parked(directory: Path) -> Path | None:
|
||||
"""Take the oldest abandoned claim, if any lease has actually expired.
|
||||
|
||||
Kept separate from the live spool so flush() can drain both in one run.
|
||||
Previously parked batches were only reachable when no spool existed at all,
|
||||
and because sessions keep recording there usually was one — so a batch
|
||||
parked by a failed send waited until the 7-day expiry deleted it unsent,
|
||||
even though its own presence is what started the sender.
|
||||
"""
|
||||
now = time.time()
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending")):
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending"), key=_safe_mtime):
|
||||
try:
|
||||
age = now - orphan.stat().st_mtime
|
||||
except OSError:
|
||||
continue
|
||||
if age > CLAIM_EXPIRY_SECONDS:
|
||||
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
|
||||
try:
|
||||
orphan.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
continue
|
||||
if age < CLAIM_STALE_SECONDS:
|
||||
# Someone else holds a live lease on it.
|
||||
continue
|
||||
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
|
||||
try:
|
||||
orphan.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
def _safe_mtime(path: Path) -> float:
|
||||
try:
|
||||
return path.stat().st_mtime
|
||||
except OSError:
|
||||
return 0.0
|
||||
|
||||
|
||||
def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
|
||||
"""Persist the unsent remainder, atomically, and refresh the lease.
|
||||
|
||||
Called after every successful batch. Two jobs: a retry resumes where the
|
||||
send stopped instead of re-posting from the top, and the rewrite doubles as
|
||||
the lease heartbeat, so a slow sender does not have its claim stolen
|
||||
mid-flight. Interval is one batch, well inside CLAIM_STALE_SECONDS.
|
||||
"""
|
||||
if not remaining:
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return True
|
||||
temporary = claim.with_suffix(f".{os.getpid()}.partial")
|
||||
try:
|
||||
temporary.write_text(
|
||||
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
|
||||
encoding="utf-8",
|
||||
)
|
||||
temporary.replace(claim)
|
||||
_touch(claim)
|
||||
return True
|
||||
except OSError:
|
||||
try:
|
||||
temporary.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
|
||||
"""Persist the remainder and drop the lease, because this sender has given up.
|
||||
|
||||
Distinct from the per-batch heartbeat: heartbeating on the way out would
|
||||
make an abandoned batch look actively owned for a further
|
||||
CLAIM_STALE_SECONDS, delaying the retry for no reason. Ageing it past the
|
||||
threshold lets the next flush pick it up immediately, while the attempt
|
||||
count in the filename still bounds how many times that can happen.
|
||||
"""
|
||||
if not _rewrite_claim(claim, remaining):
|
||||
return
|
||||
try:
|
||||
released = time.time() - CLAIM_STALE_SECONDS - 1
|
||||
os.utime(claim, (released, released))
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _resolve_email(key: str) -> str:
|
||||
"""Trade the API key for the account email so events join other Mem0 surfaces."""
|
||||
url = os.environ.get("MEM0_API_URL", memory_core.DEFAULT_API_URL).rstrip("/") + "/v1/ping/"
|
||||
@@ -300,34 +551,76 @@ def _post(payload: dict[str, Any], url: str) -> bool:
|
||||
|
||||
|
||||
def resolve_distinct_id() -> tuple[str, str]:
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any."""
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any.
|
||||
|
||||
The second value becomes a PostHog $identify alias. It is ONLY ever an
|
||||
anonymous id: aliasing one account email to another merges two real person
|
||||
profiles and cannot be undone, so a key that now belongs to a different
|
||||
account re-resolves with no alias.
|
||||
"""
|
||||
identity = _read_identity()
|
||||
email = identity.get("email", "")
|
||||
if email:
|
||||
return email, ""
|
||||
key = memory_core.api_key()
|
||||
fingerprint = _digest(key) if key else ""
|
||||
email = identity.get("email", "")
|
||||
|
||||
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
|
||||
return email, ""
|
||||
|
||||
if not key:
|
||||
# No key to verify the account with; do not keep attributing to it.
|
||||
if email:
|
||||
identity.pop("email", None)
|
||||
identity.pop("key_fingerprint", None)
|
||||
_write_identity(identity)
|
||||
return anonymous_id(identity), ""
|
||||
email = _resolve_email(key)
|
||||
if not email:
|
||||
return anonymous_id(identity), ""
|
||||
previous = identity.get("anonymous_id", "")
|
||||
identity["email"] = email
|
||||
|
||||
resolved = _resolve_email(key)
|
||||
if not resolved:
|
||||
return (email, "") if email else (anonymous_id(identity), "")
|
||||
|
||||
# Alias only when going anonymous -> email for the first time.
|
||||
previous = "" if email else identity.get("anonymous_id", "")
|
||||
identity["email"] = resolved
|
||||
identity["key_fingerprint"] = fingerprint
|
||||
_write_identity(identity)
|
||||
return email, previous
|
||||
return resolved, previous
|
||||
|
||||
|
||||
def flush() -> int:
|
||||
"""Drain claimed spools to PostHog and return the number of events sent."""
|
||||
"""Drain the live spool, then any parked claims, and return events sent."""
|
||||
if not is_enabled():
|
||||
return 0
|
||||
claim = _claim_spool()
|
||||
sent, delivered = _drain(_claim_spool())
|
||||
if not delivered:
|
||||
# The network is failing. Retrying other batches now would only burn
|
||||
# their attempt budget against the same broken connection.
|
||||
return sent
|
||||
|
||||
# Parked batches used to starve behind the live spool indefinitely. Bounded
|
||||
# per run so a long backlog cannot turn one flush into an unbounded loop.
|
||||
directory = memory_core.data_dir()
|
||||
for _ in range(MAX_PARKED_PER_RUN):
|
||||
parked = _claim_parked(directory)
|
||||
if parked is None:
|
||||
break
|
||||
count, delivered = _drain(parked)
|
||||
sent += count
|
||||
if not delivered:
|
||||
break
|
||||
return sent
|
||||
|
||||
|
||||
def _drain(claim: Path | None) -> tuple[int, bool]:
|
||||
"""Post one claimed batch file, recording progress after every batch.
|
||||
|
||||
Returns (events sent, whether everything was delivered).
|
||||
"""
|
||||
if claim is None:
|
||||
return 0
|
||||
return 0, True
|
||||
try:
|
||||
lines = claim.read_text(encoding="utf-8").splitlines()
|
||||
except OSError:
|
||||
return 0
|
||||
return 0, True
|
||||
events = []
|
||||
for line in lines:
|
||||
try:
|
||||
@@ -341,7 +634,7 @@ def flush() -> int:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return 0
|
||||
return 0, True
|
||||
|
||||
distinct_id, aliased_anonymous_id = resolve_distinct_id()
|
||||
if aliased_anonymous_id:
|
||||
@@ -360,12 +653,17 @@ def flush() -> int:
|
||||
|
||||
sent = 0
|
||||
for start in range(0, len(events), BATCH_SIZE):
|
||||
chunk = events[start : start + BATCH_SIZE]
|
||||
batch = [
|
||||
{
|
||||
"event": event["event"],
|
||||
"distinct_id": distinct_id,
|
||||
# Carried through from record() so a resend can be collapsed.
|
||||
"uuid": event.get("uuid"),
|
||||
"timestamp": event.get("timestamp"),
|
||||
"properties": {
|
||||
# Fallback only: events recorded by a build before source
|
||||
# moved into record() have none of their own.
|
||||
"source": _source_tag,
|
||||
"language": "python",
|
||||
"$process_person_profile": False,
|
||||
@@ -373,16 +671,19 @@ def flush() -> int:
|
||||
**(event.get("properties") or {}),
|
||||
},
|
||||
}
|
||||
for event in events[start : start + BATCH_SIZE]
|
||||
for event in chunk
|
||||
]
|
||||
if not _post({"api_key": POSTHOG_API_KEY, "batch": batch}, POSTHOG_BATCH_URL):
|
||||
return sent
|
||||
sent += len(batch)
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return sent
|
||||
# Keep only what has not been delivered, and release the lease.
|
||||
# Previously the whole file was kept and the retry re-posted every
|
||||
# batch, including the ones that had already arrived.
|
||||
_release_claim(claim, events[start:])
|
||||
return sent, False
|
||||
sent += len(chunk)
|
||||
# Record progress and refresh the lease after each successful batch, so
|
||||
# a crash repeats at most one batch instead of the entire file.
|
||||
_rewrite_claim(claim, events[start + len(chunk) :])
|
||||
return sent, True
|
||||
|
||||
|
||||
def main() -> int:
|
||||
|
||||
@@ -86,9 +86,9 @@ Per-call `userId` overrides are rejected unless the operator enables `allowUserO
|
||||
|
||||
## Telemetry
|
||||
|
||||
Writes are tagged `source="DEEPSEEK_HARNESS"` so Mem0's backend can attribute usage to this integration. For it to surface by name (rather than bucketing into `OTHERS`), `DEEPSEEK_HARNESS` must be present in the backend's `KNOWN_EVENT_SOURCES` allowlist, a one-line platform change matching the existing `ZAPIER` / `STRANDS` sources.
|
||||
Writes are tagged `source="DEEPSEEK_HARNESS"`, which the Mem0 backend recognizes so usage surfaces by name rather than bucketing into `OTHERS`.
|
||||
|
||||
The plugin also sends anonymous usage events (which tool ran, duration, result counts, coarse failure kind) so Mem0 can tell how the plugin is used and where it breaks. Queries, memory text, and entity ids are never sent. Turn it off with `MEM0_TELEMETRY=false`.
|
||||
The plugin also sends usage events (which tool ran, duration, result counts, coarse failure kind) so Mem0 can tell how the plugin is used and where it breaks. These are **not anonymous**: when an API key is configured they are sent under your Mem0 account email, the same way the SDK attributes its own. Queries, memory text, and entity ids are never sent. Turn it off with `MEM0_TELEMETRY=false`.
|
||||
|
||||
## Status
|
||||
|
||||
|
||||
@@ -26,10 +26,8 @@ export const name = "mem0";
|
||||
export const inject = ["tools", "systemPrompt"];
|
||||
|
||||
// Tags writes so Mem0's backend attributes them to this integration in
|
||||
// telemetry. The backend keeps recognized values via its KNOWN_EVENT_SOURCES
|
||||
// allowlist; unknown values bucket into "OTHERS", so "DEEPSEEK_HARNESS" must be
|
||||
// added to that allowlist for usage to surface by name (a one-line backend PR,
|
||||
// same pattern as the ZAPIER / STRANDS sources).
|
||||
// telemetry. The backend's KNOWN_EVENT_SOURCES allowlist recognizes this value;
|
||||
// anything outside it buckets into "OTHERS".
|
||||
const SOURCE = "DEEPSEEK_HARNESS";
|
||||
|
||||
const DEFAULT_SEARCH_LIMIT = 10;
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Generated by integrations/agent-plugin-core/build/build.py. Do not edit."""
|
||||
|
||||
HARNESS_ID = "kimi"
|
||||
SOURCE_TAG = "KIMI_PLUGIN"
|
||||
|
||||
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
|
||||
# plugin family is one source; which editor it runs in is the application.
|
||||
PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
PLATFORM_APPLICATION = "kimi"
|
||||
@@ -305,8 +305,19 @@ def run(
|
||||
return 0
|
||||
|
||||
if args.action == "session-start":
|
||||
if telemetry.is_first_run():
|
||||
# Claims the marker atomically and says which event to record, so a
|
||||
# second session starting alongside this one cannot record it too.
|
||||
first_event = telemetry.claim_install()
|
||||
if first_event == "install":
|
||||
telemetry.record("install")
|
||||
elif first_event == "upgrade":
|
||||
# First run after a build that never wrote the marker; the
|
||||
# predecessor version was never recorded anywhere.
|
||||
telemetry.record("upgrade", from_version="pre-0.3")
|
||||
else:
|
||||
previous = telemetry.claim_version_change()
|
||||
if previous:
|
||||
telemetry.record("upgrade", from_version=previous)
|
||||
recovered = recover_pending_handoffs()
|
||||
record_session_start(store, hook_input)
|
||||
if recovered:
|
||||
|
||||
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
|
||||
return batches
|
||||
|
||||
|
||||
# Platform surface attribution. Read from the generated per-host module so a new
|
||||
# entrypoint is correct without remembering to configure anything.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
except ImportError:
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
|
||||
def platform_headers(key: str) -> dict[str, str]:
|
||||
"""Auth plus the three surface-identity headers.
|
||||
|
||||
X-Mem0-Source and X-Application are set-once by contract: this is the
|
||||
outermost layer, so it sets them, and nothing below may overwrite them.
|
||||
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
|
||||
"""
|
||||
headers = {
|
||||
"Authorization": f"Token {key}",
|
||||
"Content-Type": "application/json",
|
||||
"X-Mem0-Source": _PLATFORM_SOURCE,
|
||||
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
|
||||
}
|
||||
if _PLATFORM_APPLICATION:
|
||||
headers["X-Application"] = _PLATFORM_APPLICATION
|
||||
return headers
|
||||
|
||||
|
||||
def _request_json(
|
||||
url: str, key: str, payload: dict[str, Any], timeout: float
|
||||
) -> tuple[dict[str, Any] | list[Any], int, int]:
|
||||
@@ -1807,7 +1835,7 @@ def _request_json(
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
data=raw,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="POST",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1834,7 +1862,7 @@ def _get_json(
|
||||
) -> tuple[dict[str, Any] | list[Any], int]:
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="GET",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1980,6 +2008,10 @@ def flush_session(
|
||||
"user_id": write_user,
|
||||
"app_id": repo.app_id,
|
||||
"run_id": session_id,
|
||||
# Top level, not metadata: the backend reads `source` from the body,
|
||||
# query string or X-Mem0-Source header, never from metadata. The
|
||||
# harness tag stays in metadata as hook provenance.
|
||||
"source": _PLATFORM_SOURCE,
|
||||
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
|
||||
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
|
||||
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
|
||||
@@ -2523,7 +2555,7 @@ def _collect_memory_ids(
|
||||
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
|
||||
request = urllib.request.Request(
|
||||
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="DELETE",
|
||||
)
|
||||
try:
|
||||
|
||||
@@ -1,5 +1,9 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Anonymous usage telemetry for Mem0 agent plugins.
|
||||
"""Usage telemetry for Mem0 agent plugins.
|
||||
|
||||
Events are linked to your Mem0 account email when an API key is configured, and
|
||||
to a random per-machine id otherwise. Not anonymous — the Python SDK and CLI
|
||||
attribute the same way.
|
||||
|
||||
Hooks run on a 3-6 second budget and fire on every tool call, so recording never
|
||||
touches the network: `record` appends one JSON line to a local spool and returns.
|
||||
@@ -9,7 +13,8 @@ started once per session and again from the flush worker that is already detache
|
||||
Pure stdlib, matching the rest of the plugin. Opt out with MEM0_TELEMETRY=false.
|
||||
|
||||
Never sends prompts, memory text, queries, file paths, repository names, or API
|
||||
keys: only event names, durations, counts, coarse outcomes, and salted hashes.
|
||||
keys: only event names, durations, counts, coarse outcomes, and repo/session
|
||||
identifiers hashed with a random per-install salt.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -29,8 +34,23 @@ from typing import Any
|
||||
|
||||
import memory_core
|
||||
|
||||
_harness: str = "generic"
|
||||
_source_tag: str = "MEM0_PLUGIN"
|
||||
# Seeded from the per-host module the build generates into core/. Two processes
|
||||
# in this pipeline never call init() — mcp_server.py, and the detached
|
||||
# `python3 telemetry.py` sender that spawn_flush() starts — so a module default
|
||||
# was what every one of their events got labelled with.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import HARNESS_ID as _DEFAULT_HARNESS
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
from _harness_id import SOURCE_TAG as _DEFAULT_SOURCE_TAG
|
||||
except ImportError:
|
||||
_DEFAULT_HARNESS = "generic"
|
||||
_DEFAULT_SOURCE_TAG = "MEM0_PLUGIN"
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
_harness: str = _DEFAULT_HARNESS
|
||||
_source_tag: str = _DEFAULT_SOURCE_TAG
|
||||
_PRIVATE_KEYS = {
|
||||
"apikey",
|
||||
"authorization",
|
||||
@@ -56,10 +76,19 @@ _PRIVATE_KEYS = {
|
||||
}
|
||||
|
||||
|
||||
def init(harness: str = "generic", source_tag: str = "") -> None:
|
||||
def init(harness: str = "", source_tag: str = "") -> None:
|
||||
"""Override the generated identity. Optional — core/_harness_id.py is the default.
|
||||
|
||||
The fallback shape matches memory_core.configure_harness's (``<HOST>_PLUGIN``).
|
||||
It used to be ``MEM0_<HOST>_PLUGIN`` here and ``<host>_plugin`` there, which
|
||||
meant one plugin could emit three different source values depending on which
|
||||
process happened to send the batch.
|
||||
"""
|
||||
global _harness, _source_tag
|
||||
_harness = harness
|
||||
_source_tag = source_tag or f"MEM0_{harness.upper().replace('-', '_')}_PLUGIN"
|
||||
_harness = harness or _DEFAULT_HARNESS
|
||||
_source_tag = source_tag or (
|
||||
f"{_harness.upper().replace('-', '_')}_PLUGIN" if harness else _DEFAULT_SOURCE_TAG
|
||||
)
|
||||
|
||||
POSTHOG_API_KEY = "phc_hgJkUVJFYtmaJqrvf6CYN67TIQ8yhXAkWzUn9AMU4yX"
|
||||
POSTHOG_CAPTURE_URL = "https://us.i.posthog.com/i/v0/e/"
|
||||
@@ -70,6 +99,11 @@ BATCH_SIZE = 100
|
||||
SEND_TIMEOUT = 5
|
||||
CLAIM_STALE_SECONDS = 120
|
||||
CLAIM_EXPIRY_SECONDS = 7 * 24 * 60 * 60
|
||||
# A batch is only discarded once it has genuinely been retried this many times.
|
||||
MAX_CLAIM_ATTEMPTS = 3
|
||||
# Parked claims drained per run, after the live spool. Bounded so a long backlog
|
||||
# cannot turn one flush into an unbounded send loop.
|
||||
MAX_PARKED_PER_RUN = 3
|
||||
|
||||
|
||||
def is_enabled() -> bool:
|
||||
@@ -83,9 +117,36 @@ def is_enabled() -> bool:
|
||||
|
||||
|
||||
def _digest(value: str, length: int = 16) -> str:
|
||||
"""Unsalted digest. Only for values that are already secrets (API keys)."""
|
||||
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _install_salt() -> str:
|
||||
"""Random per-install salt, created on first use and kept in the identity file."""
|
||||
identity = _read_identity()
|
||||
salt = identity.get("salt")
|
||||
if not salt:
|
||||
salt = uuid.uuid4().hex
|
||||
identity["salt"] = salt
|
||||
_write_identity(identity)
|
||||
return salt
|
||||
|
||||
|
||||
def _scoped_digest(value: str, length: int = 16) -> str:
|
||||
"""Salted digest for values drawn from a guessable space.
|
||||
|
||||
repo.identity is a git remote URL, or ``local:<absolute path>`` when there is
|
||||
no remote — which normally contains the account username. Sixteen unsalted
|
||||
hex characters over that input space is enumerable, so this is not a
|
||||
privacy control without the salt. Salting per install keeps every
|
||||
within-account join the analytics actually use and gives up only
|
||||
cross-machine joins on the same repository, which nothing computes.
|
||||
"""
|
||||
if not value:
|
||||
return ""
|
||||
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _safe_value(value: Any) -> Any:
|
||||
if isinstance(value, str):
|
||||
return memory_core.redact(value)
|
||||
@@ -145,9 +206,92 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
|
||||
return created
|
||||
|
||||
|
||||
def _install_state_path() -> Path:
|
||||
return memory_core.data_dir() / "install-state.json"
|
||||
|
||||
|
||||
def is_first_run() -> bool:
|
||||
"""Whether this machine has never recorded a plugin event before."""
|
||||
return not _identity_path().exists()
|
||||
"""Whether install has never been recorded on this machine.
|
||||
|
||||
Deliberately NOT the identity file. That file is only written by a
|
||||
successful flush, so an offline or firewalled user recorded code.install on
|
||||
every single session, forever — and every 0.2.x user recorded one on their
|
||||
first 0.3.x session because 0.2.x never wrote it at all.
|
||||
"""
|
||||
return not _install_state_path().exists()
|
||||
|
||||
|
||||
def claim_install() -> str | None:
|
||||
"""Claim the one install/upgrade record for this machine, atomically.
|
||||
|
||||
Returns the event to record ("install" or "upgrade"), or None if another
|
||||
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
|
||||
cannot both win.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
# A fresh install has an empty data directory. Anything already there —
|
||||
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
|
||||
# before the marker is created, since creating it would itself be content.
|
||||
upgrading = _data_dir_has_content()
|
||||
try:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
|
||||
except FileExistsError:
|
||||
return None
|
||||
except OSError:
|
||||
return None
|
||||
try:
|
||||
with os.fdopen(handle, "w", encoding="utf-8") as stream:
|
||||
json.dump(
|
||||
{
|
||||
"plugin_version": memory_core.PLUGIN_VERSION,
|
||||
"installed_at": memory_core.utc_now(),
|
||||
"upgraded": upgrading,
|
||||
},
|
||||
stream,
|
||||
)
|
||||
except OSError:
|
||||
pass
|
||||
return "upgrade" if upgrading else "install"
|
||||
|
||||
|
||||
def _data_dir_has_content() -> bool:
|
||||
"""Whether anything predates this session in the plugin data directory."""
|
||||
try:
|
||||
for entry in memory_core.data_dir().iterdir():
|
||||
if entry.name != "install-state.json":
|
||||
return True
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def claim_version_change() -> str | None:
|
||||
"""Return the previously recorded version if it differs, updating the marker.
|
||||
|
||||
Only meaningful once the marker exists — the first transition into 0.3.x has
|
||||
no recorded predecessor and reports "pre-0.3" instead. Claiming by rewriting
|
||||
the marker means the next session sees no change and records nothing.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
try:
|
||||
state = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return None
|
||||
if not isinstance(state, dict):
|
||||
return None
|
||||
previous = str(state.get("plugin_version") or "")
|
||||
if not previous or previous == memory_core.PLUGIN_VERSION:
|
||||
return None
|
||||
state["plugin_version"] = memory_core.PLUGIN_VERSION
|
||||
state["upgraded_at"] = memory_core.utc_now()
|
||||
try:
|
||||
temporary = path.with_suffix(f".{os.getpid()}.tmp")
|
||||
temporary.write_text(json.dumps(state), encoding="utf-8")
|
||||
temporary.replace(path)
|
||||
except OSError:
|
||||
return None
|
||||
return previous
|
||||
|
||||
|
||||
def record(
|
||||
@@ -168,19 +312,25 @@ def record(
|
||||
except OSError:
|
||||
pass
|
||||
properties = _safe_value(properties)
|
||||
# Stamped in the RECORDING process, beside harness. `source` used to be
|
||||
# read in the sending process from a module global, so whichever process
|
||||
# drained the spool named every event in it. flush() spreads per-event
|
||||
# properties last, so this now wins over any sender's default.
|
||||
properties.update(
|
||||
harness=_harness,
|
||||
source=_source_tag,
|
||||
plugin_version=memory_core.PLUGIN_VERSION,
|
||||
os=sys.platform,
|
||||
python_version=platform.python_version(),
|
||||
)
|
||||
if repo is not None:
|
||||
properties["repo_hash"] = _digest(getattr(repo, "identity", ""))
|
||||
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
|
||||
if session_id:
|
||||
properties["session_hash"] = _digest(session_id)
|
||||
properties["session_hash"] = _scoped_digest(session_id)
|
||||
line = json.dumps(
|
||||
{
|
||||
"event": f"{EVENT_PREFIX}.{event}",
|
||||
"uuid": str(uuid.uuid4()),
|
||||
"timestamp": memory_core.utc_now(),
|
||||
"properties": {
|
||||
key: value for key, value in properties.items() if value is not None
|
||||
@@ -239,38 +389,139 @@ def spawn_flush() -> bool:
|
||||
return False
|
||||
|
||||
|
||||
def _claim_name(attempt: int = 0) -> str:
|
||||
"""Claim filename. The attempt count rides in the name so the 7-day expiry
|
||||
only ever discards a batch that was actually retried and failed."""
|
||||
return f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}-a{attempt}.sending"
|
||||
|
||||
|
||||
def _claim_attempt(claim: Path) -> int:
|
||||
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
|
||||
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
|
||||
tail = stem.rsplit("-", 1)[-1]
|
||||
if tail.startswith("a") and tail[1:].isdigit():
|
||||
return int(tail[1:])
|
||||
return 0
|
||||
|
||||
|
||||
def _touch(path: Path) -> None:
|
||||
"""Refresh mtime so a claim's age measures time since it was claimed.
|
||||
|
||||
``Path.replace`` is ``os.rename``, which preserves mtime — so a claim created
|
||||
after a quiet minute inherited the spool's last-write time and looked
|
||||
abandoned the instant it was made. A second sender would then take it over
|
||||
while the first was still posting, and both would deliver the batch.
|
||||
"""
|
||||
try:
|
||||
os.utime(path, None)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _claim_spool() -> Path | None:
|
||||
"""Rename the spool aside so exactly one sender owns each batch."""
|
||||
directory = memory_core.data_dir()
|
||||
claim = directory / f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}.sending"
|
||||
claim = directory / _claim_name()
|
||||
spool = _spool_path()
|
||||
try:
|
||||
spool.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
pass
|
||||
return _claim_parked(directory)
|
||||
|
||||
|
||||
def _claim_parked(directory: Path) -> Path | None:
|
||||
"""Take the oldest abandoned claim, if any lease has actually expired.
|
||||
|
||||
Kept separate from the live spool so flush() can drain both in one run.
|
||||
Previously parked batches were only reachable when no spool existed at all,
|
||||
and because sessions keep recording there usually was one — so a batch
|
||||
parked by a failed send waited until the 7-day expiry deleted it unsent,
|
||||
even though its own presence is what started the sender.
|
||||
"""
|
||||
now = time.time()
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending")):
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending"), key=_safe_mtime):
|
||||
try:
|
||||
age = now - orphan.stat().st_mtime
|
||||
except OSError:
|
||||
continue
|
||||
if age > CLAIM_EXPIRY_SECONDS:
|
||||
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
|
||||
try:
|
||||
orphan.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
continue
|
||||
if age < CLAIM_STALE_SECONDS:
|
||||
# Someone else holds a live lease on it.
|
||||
continue
|
||||
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
|
||||
try:
|
||||
orphan.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
def _safe_mtime(path: Path) -> float:
|
||||
try:
|
||||
return path.stat().st_mtime
|
||||
except OSError:
|
||||
return 0.0
|
||||
|
||||
|
||||
def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
|
||||
"""Persist the unsent remainder, atomically, and refresh the lease.
|
||||
|
||||
Called after every successful batch. Two jobs: a retry resumes where the
|
||||
send stopped instead of re-posting from the top, and the rewrite doubles as
|
||||
the lease heartbeat, so a slow sender does not have its claim stolen
|
||||
mid-flight. Interval is one batch, well inside CLAIM_STALE_SECONDS.
|
||||
"""
|
||||
if not remaining:
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return True
|
||||
temporary = claim.with_suffix(f".{os.getpid()}.partial")
|
||||
try:
|
||||
temporary.write_text(
|
||||
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
|
||||
encoding="utf-8",
|
||||
)
|
||||
temporary.replace(claim)
|
||||
_touch(claim)
|
||||
return True
|
||||
except OSError:
|
||||
try:
|
||||
temporary.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
|
||||
"""Persist the remainder and drop the lease, because this sender has given up.
|
||||
|
||||
Distinct from the per-batch heartbeat: heartbeating on the way out would
|
||||
make an abandoned batch look actively owned for a further
|
||||
CLAIM_STALE_SECONDS, delaying the retry for no reason. Ageing it past the
|
||||
threshold lets the next flush pick it up immediately, while the attempt
|
||||
count in the filename still bounds how many times that can happen.
|
||||
"""
|
||||
if not _rewrite_claim(claim, remaining):
|
||||
return
|
||||
try:
|
||||
released = time.time() - CLAIM_STALE_SECONDS - 1
|
||||
os.utime(claim, (released, released))
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _resolve_email(key: str) -> str:
|
||||
"""Trade the API key for the account email so events join other Mem0 surfaces."""
|
||||
url = os.environ.get("MEM0_API_URL", memory_core.DEFAULT_API_URL).rstrip("/") + "/v1/ping/"
|
||||
@@ -300,34 +551,76 @@ def _post(payload: dict[str, Any], url: str) -> bool:
|
||||
|
||||
|
||||
def resolve_distinct_id() -> tuple[str, str]:
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any."""
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any.
|
||||
|
||||
The second value becomes a PostHog $identify alias. It is ONLY ever an
|
||||
anonymous id: aliasing one account email to another merges two real person
|
||||
profiles and cannot be undone, so a key that now belongs to a different
|
||||
account re-resolves with no alias.
|
||||
"""
|
||||
identity = _read_identity()
|
||||
email = identity.get("email", "")
|
||||
if email:
|
||||
return email, ""
|
||||
key = memory_core.api_key()
|
||||
fingerprint = _digest(key) if key else ""
|
||||
email = identity.get("email", "")
|
||||
|
||||
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
|
||||
return email, ""
|
||||
|
||||
if not key:
|
||||
# No key to verify the account with; do not keep attributing to it.
|
||||
if email:
|
||||
identity.pop("email", None)
|
||||
identity.pop("key_fingerprint", None)
|
||||
_write_identity(identity)
|
||||
return anonymous_id(identity), ""
|
||||
email = _resolve_email(key)
|
||||
if not email:
|
||||
return anonymous_id(identity), ""
|
||||
previous = identity.get("anonymous_id", "")
|
||||
identity["email"] = email
|
||||
|
||||
resolved = _resolve_email(key)
|
||||
if not resolved:
|
||||
return (email, "") if email else (anonymous_id(identity), "")
|
||||
|
||||
# Alias only when going anonymous -> email for the first time.
|
||||
previous = "" if email else identity.get("anonymous_id", "")
|
||||
identity["email"] = resolved
|
||||
identity["key_fingerprint"] = fingerprint
|
||||
_write_identity(identity)
|
||||
return email, previous
|
||||
return resolved, previous
|
||||
|
||||
|
||||
def flush() -> int:
|
||||
"""Drain claimed spools to PostHog and return the number of events sent."""
|
||||
"""Drain the live spool, then any parked claims, and return events sent."""
|
||||
if not is_enabled():
|
||||
return 0
|
||||
claim = _claim_spool()
|
||||
sent, delivered = _drain(_claim_spool())
|
||||
if not delivered:
|
||||
# The network is failing. Retrying other batches now would only burn
|
||||
# their attempt budget against the same broken connection.
|
||||
return sent
|
||||
|
||||
# Parked batches used to starve behind the live spool indefinitely. Bounded
|
||||
# per run so a long backlog cannot turn one flush into an unbounded loop.
|
||||
directory = memory_core.data_dir()
|
||||
for _ in range(MAX_PARKED_PER_RUN):
|
||||
parked = _claim_parked(directory)
|
||||
if parked is None:
|
||||
break
|
||||
count, delivered = _drain(parked)
|
||||
sent += count
|
||||
if not delivered:
|
||||
break
|
||||
return sent
|
||||
|
||||
|
||||
def _drain(claim: Path | None) -> tuple[int, bool]:
|
||||
"""Post one claimed batch file, recording progress after every batch.
|
||||
|
||||
Returns (events sent, whether everything was delivered).
|
||||
"""
|
||||
if claim is None:
|
||||
return 0
|
||||
return 0, True
|
||||
try:
|
||||
lines = claim.read_text(encoding="utf-8").splitlines()
|
||||
except OSError:
|
||||
return 0
|
||||
return 0, True
|
||||
events = []
|
||||
for line in lines:
|
||||
try:
|
||||
@@ -341,7 +634,7 @@ def flush() -> int:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return 0
|
||||
return 0, True
|
||||
|
||||
distinct_id, aliased_anonymous_id = resolve_distinct_id()
|
||||
if aliased_anonymous_id:
|
||||
@@ -360,12 +653,17 @@ def flush() -> int:
|
||||
|
||||
sent = 0
|
||||
for start in range(0, len(events), BATCH_SIZE):
|
||||
chunk = events[start : start + BATCH_SIZE]
|
||||
batch = [
|
||||
{
|
||||
"event": event["event"],
|
||||
"distinct_id": distinct_id,
|
||||
# Carried through from record() so a resend can be collapsed.
|
||||
"uuid": event.get("uuid"),
|
||||
"timestamp": event.get("timestamp"),
|
||||
"properties": {
|
||||
# Fallback only: events recorded by a build before source
|
||||
# moved into record() have none of their own.
|
||||
"source": _source_tag,
|
||||
"language": "python",
|
||||
"$process_person_profile": False,
|
||||
@@ -373,16 +671,19 @@ def flush() -> int:
|
||||
**(event.get("properties") or {}),
|
||||
},
|
||||
}
|
||||
for event in events[start : start + BATCH_SIZE]
|
||||
for event in chunk
|
||||
]
|
||||
if not _post({"api_key": POSTHOG_API_KEY, "batch": batch}, POSTHOG_BATCH_URL):
|
||||
return sent
|
||||
sent += len(batch)
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return sent
|
||||
# Keep only what has not been delivered, and release the lease.
|
||||
# Previously the whole file was kept and the retry re-posted every
|
||||
# batch, including the ones that had already arrived.
|
||||
_release_claim(claim, events[start:])
|
||||
return sent, False
|
||||
sent += len(chunk)
|
||||
# Record progress and refresh the lease after each successful batch, so
|
||||
# a crash repeats at most one batch instead of the entire file.
|
||||
_rewrite_claim(claim, events[start + len(chunk) :])
|
||||
return sent, True
|
||||
|
||||
|
||||
def main() -> int:
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
"""Generated by integrations/agent-plugin-core/build/build.py. Do not edit."""
|
||||
|
||||
HARNESS_ID = "coding-agent"
|
||||
SOURCE_TAG = "CODING_AGENT_PLUGIN"
|
||||
|
||||
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
|
||||
# plugin family is one source; which editor it runs in is the application.
|
||||
PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
PLATFORM_APPLICATION = "coding-agent"
|
||||
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
|
||||
return batches
|
||||
|
||||
|
||||
# Platform surface attribution. Read from the generated per-host module so a new
|
||||
# entrypoint is correct without remembering to configure anything.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
except ImportError:
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
|
||||
def platform_headers(key: str) -> dict[str, str]:
|
||||
"""Auth plus the three surface-identity headers.
|
||||
|
||||
X-Mem0-Source and X-Application are set-once by contract: this is the
|
||||
outermost layer, so it sets them, and nothing below may overwrite them.
|
||||
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
|
||||
"""
|
||||
headers = {
|
||||
"Authorization": f"Token {key}",
|
||||
"Content-Type": "application/json",
|
||||
"X-Mem0-Source": _PLATFORM_SOURCE,
|
||||
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
|
||||
}
|
||||
if _PLATFORM_APPLICATION:
|
||||
headers["X-Application"] = _PLATFORM_APPLICATION
|
||||
return headers
|
||||
|
||||
|
||||
def _request_json(
|
||||
url: str, key: str, payload: dict[str, Any], timeout: float
|
||||
) -> tuple[dict[str, Any] | list[Any], int, int]:
|
||||
@@ -1807,7 +1835,7 @@ def _request_json(
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
data=raw,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="POST",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1834,7 +1862,7 @@ def _get_json(
|
||||
) -> tuple[dict[str, Any] | list[Any], int]:
|
||||
request = urllib.request.Request(
|
||||
url,
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="GET",
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
@@ -1980,6 +2008,10 @@ def flush_session(
|
||||
"user_id": write_user,
|
||||
"app_id": repo.app_id,
|
||||
"run_id": session_id,
|
||||
# Top level, not metadata: the backend reads `source` from the body,
|
||||
# query string or X-Mem0-Source header, never from metadata. The
|
||||
# harness tag stays in metadata as hook provenance.
|
||||
"source": _PLATFORM_SOURCE,
|
||||
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
|
||||
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
|
||||
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
|
||||
@@ -2523,7 +2555,7 @@ def _collect_memory_ids(
|
||||
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
|
||||
request = urllib.request.Request(
|
||||
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
|
||||
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
|
||||
headers=platform_headers(key),
|
||||
method="DELETE",
|
||||
)
|
||||
try:
|
||||
|
||||
@@ -1,5 +1,9 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Anonymous usage telemetry for Mem0 agent plugins.
|
||||
"""Usage telemetry for Mem0 agent plugins.
|
||||
|
||||
Events are linked to your Mem0 account email when an API key is configured, and
|
||||
to a random per-machine id otherwise. Not anonymous — the Python SDK and CLI
|
||||
attribute the same way.
|
||||
|
||||
Hooks run on a 3-6 second budget and fire on every tool call, so recording never
|
||||
touches the network: `record` appends one JSON line to a local spool and returns.
|
||||
@@ -9,7 +13,8 @@ started once per session and again from the flush worker that is already detache
|
||||
Pure stdlib, matching the rest of the plugin. Opt out with MEM0_TELEMETRY=false.
|
||||
|
||||
Never sends prompts, memory text, queries, file paths, repository names, or API
|
||||
keys: only event names, durations, counts, coarse outcomes, and salted hashes.
|
||||
keys: only event names, durations, counts, coarse outcomes, and repo/session
|
||||
identifiers hashed with a random per-install salt.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -29,8 +34,23 @@ from typing import Any
|
||||
|
||||
import memory_core
|
||||
|
||||
_harness: str = "generic"
|
||||
_source_tag: str = "MEM0_PLUGIN"
|
||||
# Seeded from the per-host module the build generates into core/. Two processes
|
||||
# in this pipeline never call init() — mcp_server.py, and the detached
|
||||
# `python3 telemetry.py` sender that spawn_flush() starts — so a module default
|
||||
# was what every one of their events got labelled with.
|
||||
try: # pragma: no cover - absent only in the un-built shared source tree
|
||||
from _harness_id import HARNESS_ID as _DEFAULT_HARNESS
|
||||
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
|
||||
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
|
||||
from _harness_id import SOURCE_TAG as _DEFAULT_SOURCE_TAG
|
||||
except ImportError:
|
||||
_DEFAULT_HARNESS = "generic"
|
||||
_DEFAULT_SOURCE_TAG = "MEM0_PLUGIN"
|
||||
_PLATFORM_SOURCE = "MEM0_PLUGIN"
|
||||
_PLATFORM_APPLICATION = ""
|
||||
|
||||
_harness: str = _DEFAULT_HARNESS
|
||||
_source_tag: str = _DEFAULT_SOURCE_TAG
|
||||
_PRIVATE_KEYS = {
|
||||
"apikey",
|
||||
"authorization",
|
||||
@@ -56,10 +76,19 @@ _PRIVATE_KEYS = {
|
||||
}
|
||||
|
||||
|
||||
def init(harness: str = "generic", source_tag: str = "") -> None:
|
||||
def init(harness: str = "", source_tag: str = "") -> None:
|
||||
"""Override the generated identity. Optional — core/_harness_id.py is the default.
|
||||
|
||||
The fallback shape matches memory_core.configure_harness's (``<HOST>_PLUGIN``).
|
||||
It used to be ``MEM0_<HOST>_PLUGIN`` here and ``<host>_plugin`` there, which
|
||||
meant one plugin could emit three different source values depending on which
|
||||
process happened to send the batch.
|
||||
"""
|
||||
global _harness, _source_tag
|
||||
_harness = harness
|
||||
_source_tag = source_tag or f"MEM0_{harness.upper().replace('-', '_')}_PLUGIN"
|
||||
_harness = harness or _DEFAULT_HARNESS
|
||||
_source_tag = source_tag or (
|
||||
f"{_harness.upper().replace('-', '_')}_PLUGIN" if harness else _DEFAULT_SOURCE_TAG
|
||||
)
|
||||
|
||||
POSTHOG_API_KEY = "phc_hgJkUVJFYtmaJqrvf6CYN67TIQ8yhXAkWzUn9AMU4yX"
|
||||
POSTHOG_CAPTURE_URL = "https://us.i.posthog.com/i/v0/e/"
|
||||
@@ -70,6 +99,11 @@ BATCH_SIZE = 100
|
||||
SEND_TIMEOUT = 5
|
||||
CLAIM_STALE_SECONDS = 120
|
||||
CLAIM_EXPIRY_SECONDS = 7 * 24 * 60 * 60
|
||||
# A batch is only discarded once it has genuinely been retried this many times.
|
||||
MAX_CLAIM_ATTEMPTS = 3
|
||||
# Parked claims drained per run, after the live spool. Bounded so a long backlog
|
||||
# cannot turn one flush into an unbounded send loop.
|
||||
MAX_PARKED_PER_RUN = 3
|
||||
|
||||
|
||||
def is_enabled() -> bool:
|
||||
@@ -83,9 +117,36 @@ def is_enabled() -> bool:
|
||||
|
||||
|
||||
def _digest(value: str, length: int = 16) -> str:
|
||||
"""Unsalted digest. Only for values that are already secrets (API keys)."""
|
||||
return hashlib.sha256(value.encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _install_salt() -> str:
|
||||
"""Random per-install salt, created on first use and kept in the identity file."""
|
||||
identity = _read_identity()
|
||||
salt = identity.get("salt")
|
||||
if not salt:
|
||||
salt = uuid.uuid4().hex
|
||||
identity["salt"] = salt
|
||||
_write_identity(identity)
|
||||
return salt
|
||||
|
||||
|
||||
def _scoped_digest(value: str, length: int = 16) -> str:
|
||||
"""Salted digest for values drawn from a guessable space.
|
||||
|
||||
repo.identity is a git remote URL, or ``local:<absolute path>`` when there is
|
||||
no remote — which normally contains the account username. Sixteen unsalted
|
||||
hex characters over that input space is enumerable, so this is not a
|
||||
privacy control without the salt. Salting per install keeps every
|
||||
within-account join the analytics actually use and gives up only
|
||||
cross-machine joins on the same repository, which nothing computes.
|
||||
"""
|
||||
if not value:
|
||||
return ""
|
||||
return hashlib.sha256(f"{_install_salt()}:{value}".encode("utf-8")).hexdigest()[:length]
|
||||
|
||||
|
||||
def _safe_value(value: Any) -> Any:
|
||||
if isinstance(value, str):
|
||||
return memory_core.redact(value)
|
||||
@@ -145,9 +206,92 @@ def anonymous_id(identity: dict[str, str] | None = None) -> str:
|
||||
return created
|
||||
|
||||
|
||||
def _install_state_path() -> Path:
|
||||
return memory_core.data_dir() / "install-state.json"
|
||||
|
||||
|
||||
def is_first_run() -> bool:
|
||||
"""Whether this machine has never recorded a plugin event before."""
|
||||
return not _identity_path().exists()
|
||||
"""Whether install has never been recorded on this machine.
|
||||
|
||||
Deliberately NOT the identity file. That file is only written by a
|
||||
successful flush, so an offline or firewalled user recorded code.install on
|
||||
every single session, forever — and every 0.2.x user recorded one on their
|
||||
first 0.3.x session because 0.2.x never wrote it at all.
|
||||
"""
|
||||
return not _install_state_path().exists()
|
||||
|
||||
|
||||
def claim_install() -> str | None:
|
||||
"""Claim the one install/upgrade record for this machine, atomically.
|
||||
|
||||
Returns the event to record ("install" or "upgrade"), or None if another
|
||||
session already claimed it. O_CREAT|O_EXCL so two sessions starting together
|
||||
cannot both win.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
# A fresh install has an empty data directory. Anything already there —
|
||||
# a 0.2.x venv, an evidence db, a spool — means this is an upgrade. Read
|
||||
# before the marker is created, since creating it would itself be content.
|
||||
upgrading = _data_dir_has_content()
|
||||
try:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
handle = os.open(path, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
|
||||
except FileExistsError:
|
||||
return None
|
||||
except OSError:
|
||||
return None
|
||||
try:
|
||||
with os.fdopen(handle, "w", encoding="utf-8") as stream:
|
||||
json.dump(
|
||||
{
|
||||
"plugin_version": memory_core.PLUGIN_VERSION,
|
||||
"installed_at": memory_core.utc_now(),
|
||||
"upgraded": upgrading,
|
||||
},
|
||||
stream,
|
||||
)
|
||||
except OSError:
|
||||
pass
|
||||
return "upgrade" if upgrading else "install"
|
||||
|
||||
|
||||
def _data_dir_has_content() -> bool:
|
||||
"""Whether anything predates this session in the plugin data directory."""
|
||||
try:
|
||||
for entry in memory_core.data_dir().iterdir():
|
||||
if entry.name != "install-state.json":
|
||||
return True
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def claim_version_change() -> str | None:
|
||||
"""Return the previously recorded version if it differs, updating the marker.
|
||||
|
||||
Only meaningful once the marker exists — the first transition into 0.3.x has
|
||||
no recorded predecessor and reports "pre-0.3" instead. Claiming by rewriting
|
||||
the marker means the next session sees no change and records nothing.
|
||||
"""
|
||||
path = _install_state_path()
|
||||
try:
|
||||
state = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return None
|
||||
if not isinstance(state, dict):
|
||||
return None
|
||||
previous = str(state.get("plugin_version") or "")
|
||||
if not previous or previous == memory_core.PLUGIN_VERSION:
|
||||
return None
|
||||
state["plugin_version"] = memory_core.PLUGIN_VERSION
|
||||
state["upgraded_at"] = memory_core.utc_now()
|
||||
try:
|
||||
temporary = path.with_suffix(f".{os.getpid()}.tmp")
|
||||
temporary.write_text(json.dumps(state), encoding="utf-8")
|
||||
temporary.replace(path)
|
||||
except OSError:
|
||||
return None
|
||||
return previous
|
||||
|
||||
|
||||
def record(
|
||||
@@ -168,19 +312,25 @@ def record(
|
||||
except OSError:
|
||||
pass
|
||||
properties = _safe_value(properties)
|
||||
# Stamped in the RECORDING process, beside harness. `source` used to be
|
||||
# read in the sending process from a module global, so whichever process
|
||||
# drained the spool named every event in it. flush() spreads per-event
|
||||
# properties last, so this now wins over any sender's default.
|
||||
properties.update(
|
||||
harness=_harness,
|
||||
source=_source_tag,
|
||||
plugin_version=memory_core.PLUGIN_VERSION,
|
||||
os=sys.platform,
|
||||
python_version=platform.python_version(),
|
||||
)
|
||||
if repo is not None:
|
||||
properties["repo_hash"] = _digest(getattr(repo, "identity", ""))
|
||||
properties["repo_hash"] = _scoped_digest(getattr(repo, "identity", ""))
|
||||
if session_id:
|
||||
properties["session_hash"] = _digest(session_id)
|
||||
properties["session_hash"] = _scoped_digest(session_id)
|
||||
line = json.dumps(
|
||||
{
|
||||
"event": f"{EVENT_PREFIX}.{event}",
|
||||
"uuid": str(uuid.uuid4()),
|
||||
"timestamp": memory_core.utc_now(),
|
||||
"properties": {
|
||||
key: value for key, value in properties.items() if value is not None
|
||||
@@ -239,38 +389,139 @@ def spawn_flush() -> bool:
|
||||
return False
|
||||
|
||||
|
||||
def _claim_name(attempt: int = 0) -> str:
|
||||
"""Claim filename. The attempt count rides in the name so the 7-day expiry
|
||||
only ever discards a batch that was actually retried and failed."""
|
||||
return f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}-a{attempt}.sending"
|
||||
|
||||
|
||||
def _claim_attempt(claim: Path) -> int:
|
||||
"""Attempts recorded in a claim filename; 0 for the pre-attempt-count shape."""
|
||||
stem = claim.name[: -len(".sending")] if claim.name.endswith(".sending") else claim.name
|
||||
tail = stem.rsplit("-", 1)[-1]
|
||||
if tail.startswith("a") and tail[1:].isdigit():
|
||||
return int(tail[1:])
|
||||
return 0
|
||||
|
||||
|
||||
def _touch(path: Path) -> None:
|
||||
"""Refresh mtime so a claim's age measures time since it was claimed.
|
||||
|
||||
``Path.replace`` is ``os.rename``, which preserves mtime — so a claim created
|
||||
after a quiet minute inherited the spool's last-write time and looked
|
||||
abandoned the instant it was made. A second sender would then take it over
|
||||
while the first was still posting, and both would deliver the batch.
|
||||
"""
|
||||
try:
|
||||
os.utime(path, None)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _claim_spool() -> Path | None:
|
||||
"""Rename the spool aside so exactly one sender owns each batch."""
|
||||
directory = memory_core.data_dir()
|
||||
claim = directory / f"telemetry-{os.getpid()}-{uuid.uuid4().hex[:8]}.sending"
|
||||
claim = directory / _claim_name()
|
||||
spool = _spool_path()
|
||||
try:
|
||||
spool.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
pass
|
||||
return _claim_parked(directory)
|
||||
|
||||
|
||||
def _claim_parked(directory: Path) -> Path | None:
|
||||
"""Take the oldest abandoned claim, if any lease has actually expired.
|
||||
|
||||
Kept separate from the live spool so flush() can drain both in one run.
|
||||
Previously parked batches were only reachable when no spool existed at all,
|
||||
and because sessions keep recording there usually was one — so a batch
|
||||
parked by a failed send waited until the 7-day expiry deleted it unsent,
|
||||
even though its own presence is what started the sender.
|
||||
"""
|
||||
now = time.time()
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending")):
|
||||
for orphan in sorted(directory.glob("telemetry-*.sending"), key=_safe_mtime):
|
||||
try:
|
||||
age = now - orphan.stat().st_mtime
|
||||
except OSError:
|
||||
continue
|
||||
if age > CLAIM_EXPIRY_SECONDS:
|
||||
if age > CLAIM_EXPIRY_SECONDS and _claim_attempt(orphan) >= MAX_CLAIM_ATTEMPTS:
|
||||
try:
|
||||
orphan.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
continue
|
||||
if age < CLAIM_STALE_SECONDS:
|
||||
# Someone else holds a live lease on it.
|
||||
continue
|
||||
claim = orphan.parent / _claim_name(_claim_attempt(orphan) + 1)
|
||||
try:
|
||||
orphan.replace(claim)
|
||||
_touch(claim)
|
||||
return claim
|
||||
except OSError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
def _safe_mtime(path: Path) -> float:
|
||||
try:
|
||||
return path.stat().st_mtime
|
||||
except OSError:
|
||||
return 0.0
|
||||
|
||||
|
||||
def _rewrite_claim(claim: Path, remaining: list[dict[str, Any]]) -> bool:
|
||||
"""Persist the unsent remainder, atomically, and refresh the lease.
|
||||
|
||||
Called after every successful batch. Two jobs: a retry resumes where the
|
||||
send stopped instead of re-posting from the top, and the rewrite doubles as
|
||||
the lease heartbeat, so a slow sender does not have its claim stolen
|
||||
mid-flight. Interval is one batch, well inside CLAIM_STALE_SECONDS.
|
||||
"""
|
||||
if not remaining:
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return True
|
||||
temporary = claim.with_suffix(f".{os.getpid()}.partial")
|
||||
try:
|
||||
temporary.write_text(
|
||||
"".join(json.dumps(event, separators=(",", ":"), default=str) + "\n" for event in remaining),
|
||||
encoding="utf-8",
|
||||
)
|
||||
temporary.replace(claim)
|
||||
_touch(claim)
|
||||
return True
|
||||
except OSError:
|
||||
try:
|
||||
temporary.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return False
|
||||
|
||||
|
||||
def _release_claim(claim: Path, remaining: list[dict[str, Any]]) -> None:
|
||||
"""Persist the remainder and drop the lease, because this sender has given up.
|
||||
|
||||
Distinct from the per-batch heartbeat: heartbeating on the way out would
|
||||
make an abandoned batch look actively owned for a further
|
||||
CLAIM_STALE_SECONDS, delaying the retry for no reason. Ageing it past the
|
||||
threshold lets the next flush pick it up immediately, while the attempt
|
||||
count in the filename still bounds how many times that can happen.
|
||||
"""
|
||||
if not _rewrite_claim(claim, remaining):
|
||||
return
|
||||
try:
|
||||
released = time.time() - CLAIM_STALE_SECONDS - 1
|
||||
os.utime(claim, (released, released))
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def _resolve_email(key: str) -> str:
|
||||
"""Trade the API key for the account email so events join other Mem0 surfaces."""
|
||||
url = os.environ.get("MEM0_API_URL", memory_core.DEFAULT_API_URL).rstrip("/") + "/v1/ping/"
|
||||
@@ -300,34 +551,76 @@ def _post(payload: dict[str, Any], url: str) -> bool:
|
||||
|
||||
|
||||
def resolve_distinct_id() -> tuple[str, str]:
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any."""
|
||||
"""Return the PostHog distinct id and the anonymous id it replaced, if any.
|
||||
|
||||
The second value becomes a PostHog $identify alias. It is ONLY ever an
|
||||
anonymous id: aliasing one account email to another merges two real person
|
||||
profiles and cannot be undone, so a key that now belongs to a different
|
||||
account re-resolves with no alias.
|
||||
"""
|
||||
identity = _read_identity()
|
||||
email = identity.get("email", "")
|
||||
if email:
|
||||
return email, ""
|
||||
key = memory_core.api_key()
|
||||
fingerprint = _digest(key) if key else ""
|
||||
email = identity.get("email", "")
|
||||
|
||||
if email and identity.get("key_fingerprint", "") == fingerprint and fingerprint:
|
||||
return email, ""
|
||||
|
||||
if not key:
|
||||
# No key to verify the account with; do not keep attributing to it.
|
||||
if email:
|
||||
identity.pop("email", None)
|
||||
identity.pop("key_fingerprint", None)
|
||||
_write_identity(identity)
|
||||
return anonymous_id(identity), ""
|
||||
email = _resolve_email(key)
|
||||
if not email:
|
||||
return anonymous_id(identity), ""
|
||||
previous = identity.get("anonymous_id", "")
|
||||
identity["email"] = email
|
||||
|
||||
resolved = _resolve_email(key)
|
||||
if not resolved:
|
||||
return (email, "") if email else (anonymous_id(identity), "")
|
||||
|
||||
# Alias only when going anonymous -> email for the first time.
|
||||
previous = "" if email else identity.get("anonymous_id", "")
|
||||
identity["email"] = resolved
|
||||
identity["key_fingerprint"] = fingerprint
|
||||
_write_identity(identity)
|
||||
return email, previous
|
||||
return resolved, previous
|
||||
|
||||
|
||||
def flush() -> int:
|
||||
"""Drain claimed spools to PostHog and return the number of events sent."""
|
||||
"""Drain the live spool, then any parked claims, and return events sent."""
|
||||
if not is_enabled():
|
||||
return 0
|
||||
claim = _claim_spool()
|
||||
sent, delivered = _drain(_claim_spool())
|
||||
if not delivered:
|
||||
# The network is failing. Retrying other batches now would only burn
|
||||
# their attempt budget against the same broken connection.
|
||||
return sent
|
||||
|
||||
# Parked batches used to starve behind the live spool indefinitely. Bounded
|
||||
# per run so a long backlog cannot turn one flush into an unbounded loop.
|
||||
directory = memory_core.data_dir()
|
||||
for _ in range(MAX_PARKED_PER_RUN):
|
||||
parked = _claim_parked(directory)
|
||||
if parked is None:
|
||||
break
|
||||
count, delivered = _drain(parked)
|
||||
sent += count
|
||||
if not delivered:
|
||||
break
|
||||
return sent
|
||||
|
||||
|
||||
def _drain(claim: Path | None) -> tuple[int, bool]:
|
||||
"""Post one claimed batch file, recording progress after every batch.
|
||||
|
||||
Returns (events sent, whether everything was delivered).
|
||||
"""
|
||||
if claim is None:
|
||||
return 0
|
||||
return 0, True
|
||||
try:
|
||||
lines = claim.read_text(encoding="utf-8").splitlines()
|
||||
except OSError:
|
||||
return 0
|
||||
return 0, True
|
||||
events = []
|
||||
for line in lines:
|
||||
try:
|
||||
@@ -341,7 +634,7 @@ def flush() -> int:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return 0
|
||||
return 0, True
|
||||
|
||||
distinct_id, aliased_anonymous_id = resolve_distinct_id()
|
||||
if aliased_anonymous_id:
|
||||
@@ -360,12 +653,17 @@ def flush() -> int:
|
||||
|
||||
sent = 0
|
||||
for start in range(0, len(events), BATCH_SIZE):
|
||||
chunk = events[start : start + BATCH_SIZE]
|
||||
batch = [
|
||||
{
|
||||
"event": event["event"],
|
||||
"distinct_id": distinct_id,
|
||||
# Carried through from record() so a resend can be collapsed.
|
||||
"uuid": event.get("uuid"),
|
||||
"timestamp": event.get("timestamp"),
|
||||
"properties": {
|
||||
# Fallback only: events recorded by a build before source
|
||||
# moved into record() have none of their own.
|
||||
"source": _source_tag,
|
||||
"language": "python",
|
||||
"$process_person_profile": False,
|
||||
@@ -373,16 +671,19 @@ def flush() -> int:
|
||||
**(event.get("properties") or {}),
|
||||
},
|
||||
}
|
||||
for event in events[start : start + BATCH_SIZE]
|
||||
for event in chunk
|
||||
]
|
||||
if not _post({"api_key": POSTHOG_API_KEY, "batch": batch}, POSTHOG_BATCH_URL):
|
||||
return sent
|
||||
sent += len(batch)
|
||||
try:
|
||||
claim.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
return sent
|
||||
# Keep only what has not been delivered, and release the lease.
|
||||
# Previously the whole file was kept and the retry re-posted every
|
||||
# batch, including the ones that had already arrived.
|
||||
_release_claim(claim, events[start:])
|
||||
return sent, False
|
||||
sent += len(chunk)
|
||||
# Record progress and refresh the lease after each successful batch, so
|
||||
# a crash repeats at most one batch instead of the entire file.
|
||||
_rewrite_claim(claim, events[start + len(chunk) :])
|
||||
return sent, True
|
||||
|
||||
|
||||
def main() -> int:
|
||||
|
||||
@@ -79,9 +79,11 @@ the tool share one Mem0 backend and namespace.
|
||||
|
||||
## Telemetry
|
||||
|
||||
The store sends anonymous usage events (store configuration, operation, duration,
|
||||
result counts, coarse failure kind) over the Mem0 SDK's existing telemetry client,
|
||||
tagged `source="STRANDS"`. Queries, memory text, message content, entity ids, and
|
||||
The store sends usage events (store configuration, operation, duration, result
|
||||
counts, coarse failure kind) over the Mem0 SDK's existing telemetry client,
|
||||
tagged `source="STRANDS"`. These are **not anonymous**: when an API key is
|
||||
configured they are sent under your Mem0 account email, the same way the SDK
|
||||
attributes its own. Queries, memory text, message content, entity ids, and
|
||||
metadata are never sent. Turn it off with `MEM0_TELEMETRY=false`.
|
||||
|
||||
## Development
|
||||
|
||||
@@ -22,6 +22,10 @@ export function registerCommands(
|
||||
const pluralize = (n: number, one: string, many: string): string =>
|
||||
`${n} ${n === 1 ? one : many}`;
|
||||
|
||||
// Surface attribution on the wire. This was previously only a PostHog property,
|
||||
// so the platform saw these calls as generic SDK traffic.
|
||||
const PLATFORM_SOURCE = "PI_AGENT";
|
||||
|
||||
const searchMemories = async (query: string, scope: Scope) => {
|
||||
const filters = resolveSearchFilters(scope, getScopeCtx());
|
||||
const result = await mem0.search(query, {
|
||||
@@ -29,7 +33,8 @@ export function registerCommands(
|
||||
threshold: config.searchThreshold,
|
||||
topK: SEARCH_TOP_K,
|
||||
rerank: true,
|
||||
});
|
||||
source: PLATFORM_SOURCE,
|
||||
} as never);
|
||||
return result.results ?? [];
|
||||
};
|
||||
|
||||
@@ -45,7 +50,7 @@ export function registerCommands(
|
||||
const addParams = resolveAddParams(config.defaultScope, getScopeCtx());
|
||||
const result = await mem0.add(
|
||||
[{ role: "user", content: text }],
|
||||
{ ...addParams, customCategories: DEFAULT_CUSTOM_CATEGORIES, infer: false },
|
||||
{ ...addParams, customCategories: DEFAULT_CUSTOM_CATEGORIES, infer: false, source: PLATFORM_SOURCE } as never,
|
||||
);
|
||||
captureCommandEvent("mem0-remember", {}, telemetryCtx);
|
||||
|
||||
|
||||
@@ -1,3 +1,5 @@
|
||||
const PROVIDER_VERSION = "3.0.2";
|
||||
|
||||
import { LanguageModelV3Prompt } from '@ai-sdk/provider';
|
||||
import { Mem0ConfigSettings } from './mem0-types';
|
||||
import { loadApiKey } from '@ai-sdk/provider-utils';
|
||||
@@ -277,7 +279,11 @@ const searchInternalMemories = async (query: string, config?: Mem0ConfigSettings
|
||||
method: 'POST',
|
||||
headers: {
|
||||
Authorization: `Token ${apiKey}`,
|
||||
'Content-Type': 'application/json'
|
||||
'Content-Type': 'application/json',
|
||||
// Surface attribution. Set-once by contract: this wrapper is the
|
||||
// outermost layer on these raw fetch calls.
|
||||
'X-Mem0-Source': 'VERCEL_AI_SDK',
|
||||
'X-Mem0-Client': `mem0-vercel-ai-provider/${PROVIDER_VERSION}`
|
||||
},
|
||||
body: JSON.stringify(body),
|
||||
};
|
||||
@@ -331,7 +337,11 @@ const updateMemories = async (messages: Array<Message>, config?: Mem0ConfigSetti
|
||||
method: 'POST',
|
||||
headers: {
|
||||
Authorization: `Token ${apiKey}`,
|
||||
'Content-Type': 'application/json'
|
||||
'Content-Type': 'application/json',
|
||||
// Surface attribution. Set-once by contract: this wrapper is the
|
||||
// outermost layer on these raw fetch calls.
|
||||
'X-Mem0-Source': 'VERCEL_AI_SDK',
|
||||
'X-Mem0-Client': `mem0-vercel-ai-provider/${PROVIDER_VERSION}`
|
||||
},
|
||||
body: JSON.stringify(body),
|
||||
};
|
||||
|
||||
@@ -95,6 +95,38 @@ interface ClientIdentity {
|
||||
const IDENTITY_CACHE_MAX_DEFAULT = 50;
|
||||
const identityByCredentials = new Map<string, Promise<ClientIdentity>>();
|
||||
|
||||
const SDK_VERSION = "3.1.8";
|
||||
|
||||
/**
|
||||
* Surface-identity headers.
|
||||
*
|
||||
* X-Mem0-Source and X-Application are SET-ONCE by contract: whichever layer is
|
||||
* outermost sets them and nothing below overwrites, so a plugin wrapping this
|
||||
* SDK keeps its own identity. X-Mem0-Client is APPEND-ONLY - every layer adds
|
||||
* itself, so the platform sees the whole stack and not just the last speaker.
|
||||
*/
|
||||
function surfaceHeaders(): Record<string, string> {
|
||||
const env: Record<string, string | undefined> =
|
||||
typeof process !== "undefined" && process.env ? process.env : {};
|
||||
const existing = (env.MEM0_CLIENT_STACK ?? "").trim();
|
||||
const entries = existing
|
||||
? existing
|
||||
.split(",")
|
||||
.map((part) => part.trim())
|
||||
.filter(Boolean)
|
||||
: [];
|
||||
entries.push(`mem0-js/${SDK_VERSION}`);
|
||||
|
||||
const headers: Record<string, string> = {
|
||||
"X-Mem0-Client": entries.slice(0, 4).join(", ").slice(0, 200),
|
||||
};
|
||||
const source = (env.MEM0_SOURCE ?? "").trim();
|
||||
if (source) headers["X-Mem0-Source"] = source;
|
||||
const application = (env.MEM0_APPLICATION ?? "").trim();
|
||||
if (application) headers["X-Application"] = application;
|
||||
return headers;
|
||||
}
|
||||
|
||||
export default class MemoryClient {
|
||||
apiKey: string;
|
||||
host: string;
|
||||
@@ -129,6 +161,7 @@ export default class MemoryClient {
|
||||
this.headers = {
|
||||
Authorization: `Token ${this.apiKey}`,
|
||||
"Content-Type": "application/json",
|
||||
...surfaceHeaders(),
|
||||
};
|
||||
|
||||
this.client = axios.create({
|
||||
|
||||
+48
-18
@@ -79,6 +79,50 @@ def _maybe_alias_anon_to_email(user_email):
|
||||
logger.debug("Failed to alias anon telemetry to %r: %s", user_email, e)
|
||||
|
||||
|
||||
def _sdk_version() -> str:
|
||||
"""Resolved here rather than imported from the package root, which would cycle."""
|
||||
try:
|
||||
import importlib.metadata
|
||||
|
||||
return importlib.metadata.version("mem0ai")
|
||||
except Exception:
|
||||
return "unknown"
|
||||
|
||||
|
||||
def _client_headers(api_key: str, user_id: str) -> Dict[str, str]:
|
||||
"""Auth plus surface-identity headers.
|
||||
|
||||
X-Mem0-Source and X-Application are SET-ONCE by contract: whichever layer is
|
||||
outermost sets them, and nothing below overwrites. A plugin or harness that
|
||||
wraps this SDK therefore keeps its own identity — it declares via MEM0_SOURCE
|
||||
/ MEM0_APPLICATION and the SDK defers.
|
||||
|
||||
X-Mem0-Client is APPEND-ONLY: every layer adds itself, so the platform sees
|
||||
the whole stack rather than only whoever spoke last.
|
||||
"""
|
||||
headers = {
|
||||
"Authorization": f"Token {api_key}",
|
||||
"Mem0-User-ID": user_id,
|
||||
"X-Mem0-Client": _client_stack(),
|
||||
}
|
||||
source = os.getenv("MEM0_SOURCE", "").strip()
|
||||
if source:
|
||||
headers["X-Mem0-Source"] = source
|
||||
application = os.getenv("MEM0_APPLICATION", "").strip()
|
||||
if application:
|
||||
headers["X-Application"] = application
|
||||
return headers
|
||||
|
||||
|
||||
def _client_stack() -> str:
|
||||
"""This SDK appended to any stack an outer layer already declared."""
|
||||
existing = os.getenv("MEM0_CLIENT_STACK", "").strip()
|
||||
mine = f"mem0-python/{_sdk_version()}"
|
||||
entries = [part.strip() for part in existing.split(",") if part.strip()] if existing else []
|
||||
entries.append(mine)
|
||||
return ", ".join(entries[:4])[:200]
|
||||
|
||||
|
||||
class MemoryClient:
|
||||
"""Client for interacting with the Mem0 API.
|
||||
|
||||
@@ -129,19 +173,11 @@ class MemoryClient:
|
||||
self.client = client
|
||||
# Ensure the client has the correct base_url and headers
|
||||
self.client.base_url = httpx.URL(self.host)
|
||||
self.client.headers.update(
|
||||
{
|
||||
"Authorization": f"Token {self.api_key}",
|
||||
"Mem0-User-ID": self.user_id,
|
||||
}
|
||||
)
|
||||
self.client.headers.update(_client_headers(self.api_key, self.user_id))
|
||||
else:
|
||||
self.client = httpx.Client(
|
||||
base_url=self.host,
|
||||
headers={
|
||||
"Authorization": f"Token {self.api_key}",
|
||||
"Mem0-User-ID": self.user_id,
|
||||
},
|
||||
headers=_client_headers(self.api_key, self.user_id),
|
||||
timeout=300,
|
||||
)
|
||||
self.user_email = self._validate_api_key()
|
||||
@@ -1027,10 +1063,7 @@ class AsyncMemoryClient:
|
||||
else:
|
||||
self.async_client = httpx.AsyncClient(
|
||||
base_url=self.host,
|
||||
headers={
|
||||
"Authorization": f"Token {self.api_key}",
|
||||
"Mem0-User-ID": self.user_id,
|
||||
},
|
||||
headers=_client_headers(self.api_key, self.user_id),
|
||||
timeout=300,
|
||||
)
|
||||
|
||||
@@ -1053,10 +1086,7 @@ class AsyncMemoryClient:
|
||||
params = self._prepare_params()
|
||||
response = requests.get(
|
||||
f"{self.host}/v1/ping/",
|
||||
headers={
|
||||
"Authorization": f"Token {self.api_key}",
|
||||
"Mem0-User-ID": self.user_id,
|
||||
},
|
||||
headers=_client_headers(self.api_key, self.user_id),
|
||||
params=params,
|
||||
)
|
||||
response.raise_for_status()
|
||||
|
||||
Reference in New Issue
Block a user