Commit Graph

66 Commits

Author SHA1 Message Date
Saket Aryan 7bcb9bd124 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-15 00:32:58 +05:30
Saket Aryan 59627f2a6c Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time
# Conflicts:
#	integrations/agent-plugin-core/python/telemetry.py
#	integrations/antigravity-plugin/core/telemetry.py
#	integrations/claude-code-plugin/core/telemetry.py
#	integrations/codex-plugin/core/telemetry.py
#	integrations/cursor-plugin/core/telemetry.py
#	integrations/kimi-plugin/core/telemetry.py
#	integrations/mem0-agent-plugin/core/telemetry.py
2026-09-15 00:32:51 +05:30
Saket Aryan bd17f2b8c9 fix(plugins): make the retry budget reachable and the rewrite durable
Review found that the first cut traded the duplicate-delivery bug for a worse
one, and disproved its own load-bearing safety claim by experiment.

Expiry was unreachable. _claim_parked touched the mtime on every re-claim and
_release_claim backdated to exactly now minus the stale threshold, so a file's
age hovered around 121 seconds and never approached the 7-day expiry. The
attempt count in the filename therefore bounded nothing: an undeliverable batch
(revoked key, proxy 403, oversized event) lived on disk forever, and because
spawn_flush starts a sender whenever a .sending file exists, it spawned a
detached Python process on every hook, MCP call and CLI invocation, forever.
The old code self-healed here, so this was a regression. Expiry now gates on the
attempt budget, which is the thing that actually accumulates; age stays only as
a backstop for files that never carried an attempt marker.

The attempt parser sniffed for a leading "a", which also matches a hex id like
a1234567, so a legacy telemetry-<pid>-<hex>.sending file parsed as attempt
1234567 and was deleted unsent on the first flush after upgrade — precisely the
population this PR is meant to protect. Anchored on field position instead.

The rewrite was not durable: no fsync before the rename, and _drain unlinked any
claim that parsed to zero events. A crash between write and rename left the
claim empty, and the next flush deleted it. Now fsynced, and a non-empty file
that parses to nothing is quarantined as .corrupt rather than destroyed.

read_text raises UnicodeDecodeError on a torn file, which `except OSError` does
not catch. flush() runs from a bare `finally:` in flush_worker, so the exception
also skipped the handoff cleanup and left it stuck in .running.

The per-batch rewrite's return value was discarded, so a failed rewrite let the
loop continue as though progress had been recorded — reintroducing the exact
duplicate delivery this PR exists to fix.

.partial files orphaned by a crash between write and rename matched no glob in
the module and were never cleaned up.

Also replaces the heartbeat test, which asserted `SEND_TIMEOUT * 4 <
CLAIM_STALE_SECONDS` — two constants, executing none of the code under test. It
now drives the real rewrite and watches the mtime move. New tests cover expiry
being reachable, legacy filename parsing, torn-claim quarantine, failed-rewrite
behaviour and debris sweeping.

64 core tests, 199 host tests.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:30:58 +05:30
Saket Aryan 95d4fc27e2 fix(plugins): make the telemetry salt stable, its own file, and memoized
Review found three ways the first cut produced worse data than no salt at all.
All three came from keeping the salt as a key in the identity dict and doing an
unlocked read-modify-write.

Hooks are short-lived separate processes firing on every tool call, and people
run more than one agent window, so several processes would read {}, each mint
its own uuid4, and each hash with it. One repository hashed several ways in the
window before a writer won.

resolve_distinct_id holds a copy of that same dict across a network call to
/v1/ping/ with a 5s timeout, so whichever write landed second erased the other's
key: losing the salt changes repo_hash mid-stream, losing the email fires a
second $identify and splits the person.

_write_identity swallows OSError, and nothing memoized, so on a read-only or
full data directory every single event got a brand-new random salt — unbounded
cardinality in PostHog, which is strictly worse than the unsalted value it
replaced.

The salt now lives in its own file claimed with O_CREAT|O_EXCL, so exactly one
process wins and the losers read the winner's value, and it is memoized per
process. When it cannot be persisted the fallback is derived from the data
directory path: stable for the machine rather than random per call.

Its own file also means record() no longer creates telemetry-identity.json as a
side effect. is_first_run keys off that file, so the first cut would have
silently suppressed the install event — a production metric change hidden in a
docs PR.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:25:12 +05:30
Saket Aryan cfdfe40e09 fix(plugins): stop delivering telemetry events twice, and stop losing parked ones
Two defects, one cause: the spool protocol infers ownership instead of holding
it, and never records progress.

Duplicate delivery after a partial failure. flush() posts the claim in batches of
100 and returns on the first failure, keeping the whole file. The retry then
posts every batch again, including the ones that already arrived — 150 recorded
events were delivered 250 times. Progress is now written back to the claim after
each successful batch, so a retry resumes where the send stopped and a crash
repeats at most one batch.

Duplicate delivery when two senders overlap. spool.replace(claim) is os.rename,
which preserves mtime, so a claim created after a quiet minute inherited the
spool's last-write time and looked abandoned the instant it existed. A second
sender starting while the first was still posting took it over and sent it too —
most likely at session end, when the MCP server's exit sender and the SessionEnd
flush worker both drain. Claims are now touched at claim time, and the per-batch
rewrite doubles as a lease heartbeat. _post makes one attempt with SEND_TIMEOUT
and no retry, so a heartbeat lands well inside the 120s lease; a test asserts
that margin so adding a retry loop to _post cannot silently break it.

Parked batches starved. _claim_spool only looked at parked .sending files when
no spool existed, and because sessions keep recording there usually was one — so
a batch parked by a failed send waited until the 7-day expiry deleted it unsent,
despite its own presence being what starts the sender in the first place.
flush() now drains the live spool and then parked claims in the same run, oldest
first, bounded. Expiry applies only after a genuine retry has failed, with the
attempt count carried in the filename.

A sender that gives up releases its lease rather than heartbeating on the way
out, so the next run picks the batch up promptly instead of waiting a full stale
window for a batch nobody is working on. A failing send stops the run, so one
broken connection cannot burn every parked batch's attempt budget at once.

Two existing tests asserted the old lifecycle and are updated in place, each
with a comment saying what changed.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:18:21 +05:30
Saket Aryan f9c566aa16 fix(plugins): report the plugin that produced the event, not the one that sent it
harness is set when an event is recorded; source was set when its batch was
sent. Both came from module globals that stay at "generic" and "MEM0_PLUGIN"
until telemetry.init() runs, and two processes in the pipeline never run it:

- `python3 telemetry.py`, the detached sender spawn_flush() starts at session
  start, after every skill command, and when the MCP server exits. Everything it
  delivered was labelled source=MEM0_PLUGIN. Only batches flush_worker.py
  happened to drain got the real host.
- mcp_server.py, which records every manual search as harness=generic.

All six Python plugins ship the same files, so source could not tell any of them
apart and MCP searches from every plugin landed in one generic bucket. The
portable plugin is worse: it has no flush_worker at all, so its only sender is
the uninitialised one and 100% of its events were mislabelled.

Two changes. record() stamps source beside harness, so the sending process stops
mattering — flush() already spreads per-event properties last, so a per-event
source wins over any sender default. And the build generates core/_harness_id.py
per host, seeding both modules at import, so identity no longer depends on an
entrypoint remembering to call init(). The build already computed HARNESS_ID and
spent it only on skill templating, and bundle_drift already diffs core/
byte-for-byte, so --check catches drift for free.

Deliberately not adding MEM0_PLUGIN_HARNESS to the six manifests: they sit
outside the --sync and --check boundary, which is the property that caused this.

Also unifies two defaults that disagreed. configure_harness derived
`<host>_plugin` while telemetry.init derived `MEM0_<HOST>_PLUGIN`, so a third
value existed. It was unreachable only because hook_runner never calls flush();
moving source into record() would have made it live.

Events now carry a uuid so a resend can be collapsed.

The suite stayed green through all of this because the only tests live under one
host, behind a conftest that calls init() at import. New tests run in real
subprocesses with no init, and cover the portable plugin, which would pass a
native-only test vacuously.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:17:53 +05:30
Saket Aryan 0d37619f24 fix(plugins): say what telemetry actually sends, and salt the hashes
The plugin README promises "anonymous usage events" and the telemetry module's
docstring says it sends only "salted hashes". Neither is true.

resolve_distinct_id() exchanges the API key for the account email and sends that
as the distinct_id on every event. Installing the plugin requires an API key, so
this is nearly every user. That is probably the behaviour we want — the Python
SDK and the CLI attribute the same way — but the description has to match it.

repo_hash and session_hash were unsalted SHA-256 cut to 16 hex characters.
repo.identity is a git remote URL, or `local:<absolute path>` when there is no
remote, which normally contains the account username. Sixteen unsalted hex
characters over that input space is enumerable, so the hash was not a privacy
control at all.

Salted per install, with the salt kept in the identity file. That preserves
every within-account join the analytics actually use and gives up only
cross-machine joins on the same repository, which nothing computes. Since the
distinct_id is already the email, the hash was never buying privacy from us —
only from whoever obtains the data later, which is exactly what the salt fixes.

Also corrects deepseek-plugin's README and source comment, which told readers
ZAPIER and STRANDS were already in the backend's KNOWN_EVENT_SOURCES allowlist.
Neither was.

Adds a Telemetry section to docs/integrations/claude-code.mdx, which had none.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:17:26 +05:30
Harsh Vardhan Gupta c7ee362aff fix(security): resolve 12 Vanta/Dependabot vulnerabilities across 6 pnpm workspaces + poetry.lock (#7280)
Co-authored-by: kartik-mem0 <kartik.labhshetwar@mem0.ai>
2026-09-11 16:07:57 +05:30
Kartik d873892dad feat(plugins)!: make Sidekick exclusive to Claude Code (#7278) 2026-09-10 20:51:50 +05:30
Kartik 02f7a9b2c4 docs: align agent plugin guides with shared runtime behavior (#7269) 2026-09-09 01:03:26 +05:30
Kartik 73e7b8763a refactor(integrations): shared agent plugin runtimes and native adapters (#7203) 2026-09-08 23:32:25 +05:30
Kartik 71fba8d464 feat(claude-code-plugin): move the Claude Code plugin to its own integration and ship it as 0.3.0 (#7106) 2026-09-01 02:34:45 +05:30
Kartik 0070e08e01 feat(deepseek-plugin,mem0-strands): add usage telemetry (#7110) 2026-08-27 13:48:30 +05:30
Kartik b717e38785 refactor(integrations): rename dsh-mem0 to deepseek-plugin, strands-mem0 to mem0-strands (#7098) 2026-08-24 19:21:23 +05:30
Kartik 4ddee9c51d chore(release): bump SDK, CLI, and plugin versions; add Strands, DeepSeek Harness, and Kimi changelogs (#7097) 2026-08-24 18:10:44 +05:30
Himanshu 7e09615571 feat(integrations): dsh-mem0 — Mem0 as a native DeepSeek Harness plugin (#7027)
Co-authored-by: kartik-mem0 <kartik.labhshetwar@mem0.ai>
2026-08-24 14:03:50 +05:30
Himanshu 8d5b7865bd feat(integrations): strands-mem0 | Mem0 as a native Strands MemoryStore (#7021) 2026-08-22 19:04:24 +05:30
Harsh Vardhan Gupta 5af797834c fix(security): resolve 17 Vanta/Dependabot HIGH+CRITICAL vulnerabilities across 5 pnpm workspaces (#7032) 2026-08-21 17:41:41 +05:30
Himanshu 1de6499b8a fix(integrations/zapier): address Zapier publishing review (#6985) 2026-08-20 18:53:06 +05:30
Kartik 530d802b55 fix(plugins): bug-bash fixes for Cursor, Codex, Antigravity, and a Claude.ai docs page (#6948) 2026-08-20 15:56:17 +05:30
Kartik 290de24bb8 feat(ci): gate pull requests on an accepted issue (#6894) 2026-08-14 17:05:27 +05:30
Kartik a10c0cd030 fix(mem0-plugin): stop search errors from looking like empty results (#6898) 2026-08-14 16:57:01 +05:30
Kartik 96d45b78c7 fix(mem0-plugin): drop unused pytest import breaking make lint on main (#6937) 2026-08-13 16:51:01 +05:30
Himanshu ba2fb9f4c3 feat(kimi): Mem0 plugin for Kimi Code (MCP + skills + auto-capture) (#6919) 2026-08-13 16:15:04 +05:30
Harsh Vardhan Gupta 4debc58a83 fix(security): patch 8 HIGH + 18 MEDIUM Vanta vulnerabilities across 4 pnpm workspaces (#6847) 2026-08-07 19:04:33 +05:30
Himanshu 3f39fba28f fix(n8n): MIT license + themed icons for verified-node vetting (#6804)
Co-authored-by: kartik-mem0 <kartik.labhshetwar@mem0.ai>
2026-08-06 00:00:13 +05:30
Kartik 1112be3e5e chore(release): bump SDK, CLI, and plugin versions (#6800) 2026-08-05 00:16:26 +05:30
Himanshu b54710a3c3 fix(zapier): require user_id on Add Memory (#6790) 2026-08-04 18:49:35 +05:30
Himanshu 4cfa98f626 chore(n8n): route package contact to integrations@mem0.ai (#6791) 2026-08-04 17:39:20 +05:30
Himanshu 965140eb19 fix(zapier): add root index.js entry shim so deployed app resolves (#6789) 2026-08-04 12:48:46 +05:30
Kartik c90bdbdce0 feat(cli): Platform option parity across Python and Node CLIs (MEM-5893) (#6696) 2026-08-03 17:10:44 +05:30
Kartik 50bdaaea0c chore: bump versions and update changelog for Python 2.0.15, TypeScript 3.1.3, and plugin releases (#6715) 2026-08-01 20:26:31 +05:30
Kartik 07e58c54ae fix: align default model names with SDK defaults (#6704) 2026-08-01 01:30:57 +05:30
Harsh Vardhan Gupta 9c2d6222ce fix(security): patch 32 HIGH + 57 MEDIUM Vanta vulnerabilities across 6 pnpm workspaces (#6639)
Co-authored-by: kartik-mem0 <kartik.labhshetwar@mem0.ai>
2026-07-30 15:20:13 +05:30
Kartik 790e190486 chore(n8n): release 0.1.1 via CD for npm provenance (#6685) 2026-07-30 13:54:37 +05:30
Himanshu d4869d24ec feat(integrations): n8n community node for Mem0 (#6517)
Co-authored-by: kartik-mem0 <kartik.labhshetwar@mem0.ai>
2026-07-29 23:03:03 +05:30
Kartik 1ac3aa7256 fix(zapier): raise the add_memory poll budget past the real API latency tail (#6680) 2026-07-29 21:32:22 +05:30
Himanshu e168d48e04 feat(integrations): Zapier app for Mem0 (#6518)
Co-authored-by: kartik-mem0 <kartik.labhshetwar@mem0.ai>
2026-07-29 20:56:44 +05:30
Kartik ea2ee07586 chore: remove OpenMemory from the monorepo (#6530) 2026-07-29 15:10:32 +05:30
Kartik 5e7adc4d12 chore: update changelog, bump versions to Python 2.0.13, TypeScript 3.1.1, OpenCode plugin 0.2.2 (#6504) 2026-07-22 23:27:47 +05:30
Rod Boev b05cce581b fix(opencode-plugin): recover MEM0_API_KEY from shell profiles (#6404) 2026-07-20 20:41:41 +05:30
Kartik ccbe5861a1 docs: remove criteria retrieval docs for non-existent feature (#6282) 2026-07-14 20:06:05 +05:30
Kartik d6d2588ef5 fix(mem0-plugin): store assistant-authored summaries with role="assistant" (#6316) 2026-07-14 20:03:36 +05:30
Kartik 8488abe603 docs: correct custom categories, per-call custom_categories is supported (#6218) 2026-07-10 20:08:32 +05:30
Kartik 41c8f00851 chore(integrations): plugin updates, pi-agent auto-recall, and version bumps (#6011) 2026-07-01 20:57:32 +05:30
Kartik c325bd3b8e docs(changelog): consolidate per-package changelogs into the SDK changelog page (#6007) 2026-06-30 14:09:41 +05:30
Terrasse cc59d122db docs: remove instructions for unavailable Cursor marketplace plugin (#5971) 2026-06-29 20:35:07 +05:30
Harsh Vardhan Gupta bbbfcfea07 fix(deps): bump undici to >=6.27.0 (CVE-2026-12151) (#5861) 2026-06-26 15:32:09 +05:30
Rod Boev 1f66aadfa3 fix(openclaw): normalize Windows skill-loader URLs before fileURLToPath (#5679) 2026-06-25 16:36:11 +05:30
Kartik ac8f862ff7 fix(mem0-plugin): store files_touched as a list to stop double JSON-encoding (#5806) 2026-06-25 09:06:11 +05:30