Review point. The portable bundle is built with host "coding-agent", and the
build wrote that straight into PLATFORM_APPLICATION, so every portable install
sent X-Application: coding-agent.
That value is not in the platform's allowlist, so it was already being dropped
server-side. The effect was the worst of both: the wire claimed we knew the
editor, the stored event recorded that we did not, and nothing said which was
right. An absent header says the same thing honestly and costs a lookup.
HARNESS_ID stays "coding-agent". It is the PostHog-side label, it is true, and
grouping portable installs together there is useful.
Native bundles are unchanged apart from the regenerated comment. Covered by two
new build tests: portable declares no application, and each native names the
host it was generated for.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
The sweep in aa770aa6 deliberately left this page alone, reasoning that
hashing the email is materially different from sending it. Reading
integrations/openclaw/telemetry.ts does not support that: distinctId() is an
unsalted sha256 of the account email, and Mem0 holds the emails it is derived
from, so recovering the account is a table join. resolveEmail() also rewrites
already-queued events onto that id, and identifyAnonymous() fires a PostHog
$identify that merges the prior random id into it for good.
That is pseudonymous, not anonymous, and it is the same mismatch between the
stated privacy posture and the wire format that this stack exists to close.
The opt-out is unchanged and still correct.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
The pr4 -> pr5 merge resolved eleven documentation files to the pre-sweep
side, undoing aa770aa6 in its entirety. Because the stack lands in order,
main would have taken the corrected wording in pr1 and then had it reverted
by pr5, leaving the shipped claim wrong again:
- all six hosts' pause skill back to "a minimal anonymous telemetry ping"
- docs/integrations/deepseek-plugin.mdx back to "Anonymous usage events"
- integrations/zapier-mem0/README.md back to advertising telemetry the app
does not have, with an MEM0_TELEMETRY opt-out that controls nothing
- the README data-dir listing back to omitting telemetry-salt and
install-state.json, both of which this stack creates
- the README and docs property lists back to the enumeration that drifts
Restored verbatim from pr4. No file here is in pr5's scope, and the full
pr4..pr5 diff is now surface headers only.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
CI caught what I could not check locally: `pnpm exec tsc --noEmit` fails with
error TS2353: Object literal may only specify known properties,
and 'source' does not exist in type 'SearchMemoryOptions'.
Adding `source` to SearchMemoryOptions in mem0-ts does not help here. pi-agent
resolves `mem0ai` from npm, so it typechecks against the published 3.1.8 types,
not this repo's source. The declaration still belongs in mem0-ts for the next
release; this call site needs to compile today.
Widened by exactly that one property rather than restoring `as never`, which
was the original objection: a blanket cast also disabled checking of filters,
threshold, topK and rerank on the same literal. `source` reaches the wire
through the SDK's camelToSnakeKeys spread either way.
Verified by installing the package deps and running the real gates: tsc clean,
build clean. vercel-ai-sdk typechecks clean too. Also carries the spool-test
environment fix that had not been committed in this worktree.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
The first pass fixed the plugin README and the module docstring but left the
same claim standing everywhere else.
- docs/integrations/deepseek-plugin.mdx still said "Anonymous usage events".
The TS SDK's telemetryId is the raw account email, so it is not anonymous.
- The pause skill told users a "minimal anonymous telemetry ping" fires while
paused. Same ping, same email. Corrected in the template, which regenerates
into all six hosts.
- integrations/zapier-mem0/README.md advertised telemetry the app does not have:
there is no telemetry code in it at all. It now says what is actually true,
that its requests carry source="ZAPIER".
- The data directory listing is presented as exhaustive and had gone stale
against this stack's two new files, telemetry-salt and install-state.json.
Also replaced the property enumeration in both the README and the docs page.
Review pointed out it omitted the configured model name among others — writing
a fresh exhaustive list in a PR whose whole purpose is making docs match code
reproduces the defect being fixed. It now describes the shape and points at
where the rule is actually enforced, so it cannot drift again.
Deliberately unchanged: docs/integrations/openclaw.mdx. OpenClaw hashes the
email rather than sending it, which is materially different from the plugin and
the SDK, so its claim is not wrong in the same way.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
CI runs agent-plugin-core/tests and claude-code-plugin/tests in one pytest
process. Two things only show up in that combined run, so the suites passed
locally and failed on every push.
claude-code-plugin/tests/conftest.py sets MEM0_TELEMETRY=false at import, which
is process-wide. record() then returns early and every assertion in
test_spool_delivery.py saw an empty spool — nine failures, all reported as
"recorded nothing" rather than as a disabled feature. The fixture now pins
MEM0_TELEMETRY rather than trusting whatever collected first.
The fixture also dropped telemetry/memory_core/_harness_id from sys.modules on
teardown. That conftest imports memory_core once at collection and calls
configure_harness() on it, so a later re-import got a fresh module with default
harness config and test_memory_core failed depending on collection order. The
fixture now saves and restores those entries instead of deleting them.
Verified with CI's exact command rather than the narrower path I had been
running: 266 passed, 8 skipped.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
The keep-both conflict resolution split a function body. Rebuilt from both
merge parents so the header-contract test and the session-start tests are each
intact.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Review found the contract documented but not implemented, and one client path
missed entirely.
AsyncMemoryClient's custom-client branch still carried the old literal header
dict, so `AsyncMemoryClient(client=...)` sent no surface identity at all — the
exact asymmetry this work set out to remove.
Both custom-client branches also used a blanket headers.update(), which
overwrites. That is the one code path where an outer layer's identity can
physically be present, and it was the one path that erased it. They now
check-then-set the identity headers and append to an existing client stack,
which is what set-once and append-only were supposed to mean.
AGENTS.md claimed a plugin calling the Python SDK produces
`mem0-plugin/0.3.1, mem0-python/2.0.19`. Nothing in the repo sets the env vars
that would make that happen, so the concatenation was unreachable. Replaced with
the three ways an integration can actually declare itself, in preference order.
memory_core's comment said the backend reads X-Mem0-Source. That is only true
from the platform release shipping alongside this, and a reader would otherwise
trust it and build header-only attribution that silently does nothing — which is
how vercel-ai-sdk was written in the first cut. Corrected in all seven copies,
and the body value is what makes attribution work against either backend.
mem0-ts hardcoded SDK_VERSION = "3.1.8" while the repo already injects
__MEM0_SDK_VERSION__ via tsup, the same mechanism telemetry.ts uses. The
hardcode was correct only until the next release bump.
Dropped both `as never` casts in pi-agent. They suppressed an excess-property
error but also disabled checking of every other option at those call sites, so a
typo in filters or threshold would have compiled. SearchMemoryOptions now
declares `source` instead.
Stack truncation cut mid-identifier, leaving a fragment that parses as a real
client name. It now drops whole entries.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Review found the headline fix inverted: code.install could never fire, so every
fresh install reported an upgrade and the two cohorts became indistinguishable —
strictly worse than the bug being fixed.
hook_runner reaches claim_install() only after cache_plugin_api_key() has
written `api-key` and EvidenceStore() has created `evidence.sqlite3` and its WAL
files. Asking "is the data directory empty" at that point always saw content.
The caller now snapshots emptiness at the top of the run, before anything
writes, and passes it in.
Also caught by review, all in the same file:
- claim_version_change was an unsynchronized read-modify-write, so several
concurrently starting sessions each observed the old version and each recorded
an upgrade. The first session after a version bump is exactly when a user's
open agent windows all restart together. The transition is now claimed with an
exclusive per-version sentinel.
- A crash between O_EXCL and the write left an empty marker, which disabled
every future upgrade event on that machine: claim_install saw the file and
claim_version_change could not parse it. An unparseable marker is now
repaired.
- claim_install consumed the one-shot claim even under MEM0_TELEMETRY=false, so
a user who opted out for their first sessions would never report install after
opting in.
- Existing users have an email but no key fingerprint, so the fast path always
missed and every flush paid an uncached /v1/ping/ — a 5s timeout each time for
the offline users this stack keeps citing. Legacy rows now adopt the current
key's fingerprint instead of re-resolving.
- A key that will not resolve (revoked, offline) kept attributing to the
previous account's email, which is the bug this was meant to fix. It now falls
back to the anonymous id.
- The anonymous id was never rotated, so once it had been merged into one
account it was still offered as the alias for the next one. An alias naming an
already-identified id is what could link two real people; it is now offered
once.
The gap that let this ship was that no test drove hook_runner's session-start
path — the decision was only ever tested by calling claim_install() directly on
a directory nothing had touched. Adds subprocess tests that run the real
entrypoint: fresh install, exactly-once, and an existing data dir.
62 core tests, 203 host tests.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Review found that the first cut traded the duplicate-delivery bug for a worse
one, and disproved its own load-bearing safety claim by experiment.
Expiry was unreachable. _claim_parked touched the mtime on every re-claim and
_release_claim backdated to exactly now minus the stale threshold, so a file's
age hovered around 121 seconds and never approached the 7-day expiry. The
attempt count in the filename therefore bounded nothing: an undeliverable batch
(revoked key, proxy 403, oversized event) lived on disk forever, and because
spawn_flush starts a sender whenever a .sending file exists, it spawned a
detached Python process on every hook, MCP call and CLI invocation, forever.
The old code self-healed here, so this was a regression. Expiry now gates on the
attempt budget, which is the thing that actually accumulates; age stays only as
a backstop for files that never carried an attempt marker.
The attempt parser sniffed for a leading "a", which also matches a hex id like
a1234567, so a legacy telemetry-<pid>-<hex>.sending file parsed as attempt
1234567 and was deleted unsent on the first flush after upgrade — precisely the
population this PR is meant to protect. Anchored on field position instead.
The rewrite was not durable: no fsync before the rename, and _drain unlinked any
claim that parsed to zero events. A crash between write and rename left the
claim empty, and the next flush deleted it. Now fsynced, and a non-empty file
that parses to nothing is quarantined as .corrupt rather than destroyed.
read_text raises UnicodeDecodeError on a torn file, which `except OSError` does
not catch. flush() runs from a bare `finally:` in flush_worker, so the exception
also skipped the handoff cleanup and left it stuck in .running.
The per-batch rewrite's return value was discarded, so a failed rewrite let the
loop continue as though progress had been recorded — reintroducing the exact
duplicate delivery this PR exists to fix.
.partial files orphaned by a crash between write and rename matched no glob in
the module and were never cleaned up.
Also replaces the heartbeat test, which asserted `SEND_TIMEOUT * 4 <
CLAIM_STALE_SECONDS` — two constants, executing none of the code under test. It
now drives the real rewrite and watches the mtime move. New tests cover expiry
being reachable, legacy filename parsing, torn-claim quarantine, failed-rewrite
behaviour and debris sweeping.
64 core tests, 199 host tests.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Review found three ways the first cut produced worse data than no salt at all.
All three came from keeping the salt as a key in the identity dict and doing an
unlocked read-modify-write.
Hooks are short-lived separate processes firing on every tool call, and people
run more than one agent window, so several processes would read {}, each mint
its own uuid4, and each hash with it. One repository hashed several ways in the
window before a writer won.
resolve_distinct_id holds a copy of that same dict across a network call to
/v1/ping/ with a 5s timeout, so whichever write landed second erased the other's
key: losing the salt changes repo_hash mid-stream, losing the email fires a
second $identify and splits the person.
_write_identity swallows OSError, and nothing memoized, so on a read-only or
full data directory every single event got a brand-new random salt — unbounded
cardinality in PostHog, which is strictly worse than the unsalted value it
replaced.
The salt now lives in its own file claimed with O_CREAT|O_EXCL, so exactly one
process wins and the losers read the winner's value, and it is memoized per
process. When it cannot be persisted the fallback is derived from the data
directory path: stable for the machine rather than random per call.
Its own file also means record() no longer creates telemetry-identity.json as a
side effect. is_first_run keys off that file, so the first cut would have
silently suppressed the install event — a production metric change hidden in a
docs PR.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Nothing on the wire said which Mem0 surface made a call. Both SDKs sent only an
auth header, so the platform saw python-httpx and axios and attributed every
plugin, wrapper and direct API user to one undifferentiated bucket. Version was
unknowable, which is what gates every deprecation decision.
Three headers, and the rules on them are the point:
- X-Mem0-Source and X-Application are SET-ONCE. Whichever layer is outermost
sets them; nothing below overwrites. A plugin wrapping the SDK keeps its own
identity instead of being renamed by the transport underneath it.
- X-Mem0-Client is APPEND-ONLY. A plugin calling the Python SDK produces
`mem0-plugin/0.3.1, mem0-python/2.0.19`, so neither layer can erase the other.
Deliberately not User-Agent: proxies rewrite it, and we have already met a WAF
that 403s on it.
The plugin core also hoists `source` out of metadata to the top level, which is
where the backend actually reads it. It sat in metadata, which get_event_source
never consults, so all six plugins arrived indistinguishable from a raw SDK call
no matter what they set. The harness tag stays in metadata as hook provenance.
pi-agent had PI_AGENT as a PostHog property only and never sent it on the wire.
vercel-ai-sdk sent nothing at all from its raw fetch calls.
Values must exist in the platform's EventSource enum or they bucket to OTHERS,
so integrations/AGENTS.md now states the contract and the "adding an
integration" checklist requires landing the platform value in the same week.
Pairs with mem0ai/platform#3602, which recognizes these values.
TypeScript changes are not typechecked locally — deps are not installed for
those packages. CI covers them.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
code.install counted upgrades and repeat sessions. Session start records install
whenever is_first_run() is true, and that only checked whether
telemetry-identity.json exists. Recording install does not create that file —
only the first successful flush does. So install fired for every 0.2.x user on
their first 0.3.x session (0.2.x never wrote the file, and the data directory
survives the upgrade), again for any session starting before that first flush
finished, and — this is the part that makes it unbounded rather than a race —
on every single session, forever, for anyone whose flush never succeeds. An
offline or firewalled user reported a new install every time they opened an
editor, which is exactly the population hardest to see in the data.
A dedicated install-state.json is now claimed with O_CREAT|O_EXCL at the moment
install is recorded, so two sessions starting together cannot both win, and the
marker is not coupled to identity. Deliberately not the identity file: writing
that from a recording process would race the sender, which writes it during
resolve_distinct_id, and overloading it is what caused this.
Upgrade detection keys on the data directory already having content. A fresh
install has an empty one; anything else predates this session. That is a firmer
predicate than looking for 0.2.x's venv/ and requirements.txt, which is a guess
about files another part of the plugin may or may not have written and only ever
works for this one upgrade. The version is stored in the marker so later changes
record code.upgrade with a real from_version.
A cached email outlived an API key change. resolve_distinct_id kept the first
email it resolved and never looked again, so switching to a key from another
account kept attributing events to the previous one. It now stores a fingerprint
of the key the email came from and re-resolves when the current key differs, and
falls back to the anonymous id when no key is configured rather than continuing
to attribute to an account it cannot verify.
The dangerous part is the alias. resolve_distinct_id's second return value
becomes a PostHog $identify with $anon_distinct_id, and aliasing one account
email to another merges two real person profiles irreversibly. The re-resolve
path returns no alias; aliasing runs anonymous to email only, and never
email to email.
One existing test asserted that is_first_run flips when the identity file is
written, which is the defect itself. Rewritten, along with coverage for atomic
claiming, upgrade detection and version changes.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
Two defects, one cause: the spool protocol infers ownership instead of holding
it, and never records progress.
Duplicate delivery after a partial failure. flush() posts the claim in batches of
100 and returns on the first failure, keeping the whole file. The retry then
posts every batch again, including the ones that already arrived — 150 recorded
events were delivered 250 times. Progress is now written back to the claim after
each successful batch, so a retry resumes where the send stopped and a crash
repeats at most one batch.
Duplicate delivery when two senders overlap. spool.replace(claim) is os.rename,
which preserves mtime, so a claim created after a quiet minute inherited the
spool's last-write time and looked abandoned the instant it existed. A second
sender starting while the first was still posting took it over and sent it too —
most likely at session end, when the MCP server's exit sender and the SessionEnd
flush worker both drain. Claims are now touched at claim time, and the per-batch
rewrite doubles as a lease heartbeat. _post makes one attempt with SEND_TIMEOUT
and no retry, so a heartbeat lands well inside the 120s lease; a test asserts
that margin so adding a retry loop to _post cannot silently break it.
Parked batches starved. _claim_spool only looked at parked .sending files when
no spool existed, and because sessions keep recording there usually was one — so
a batch parked by a failed send waited until the 7-day expiry deleted it unsent,
despite its own presence being what starts the sender in the first place.
flush() now drains the live spool and then parked claims in the same run, oldest
first, bounded. Expiry applies only after a genuine retry has failed, with the
attempt count carried in the filename.
A sender that gives up releases its lease rather than heartbeating on the way
out, so the next run picks the batch up promptly instead of waiting a full stale
window for a batch nobody is working on. A failing send stops the run, so one
broken connection cannot burn every parked batch's attempt budget at once.
Two existing tests asserted the old lifecycle and are updated in place, each
with a comment saying what changed.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
harness is set when an event is recorded; source was set when its batch was
sent. Both came from module globals that stay at "generic" and "MEM0_PLUGIN"
until telemetry.init() runs, and two processes in the pipeline never run it:
- `python3 telemetry.py`, the detached sender spawn_flush() starts at session
start, after every skill command, and when the MCP server exits. Everything it
delivered was labelled source=MEM0_PLUGIN. Only batches flush_worker.py
happened to drain got the real host.
- mcp_server.py, which records every manual search as harness=generic.
All six Python plugins ship the same files, so source could not tell any of them
apart and MCP searches from every plugin landed in one generic bucket. The
portable plugin is worse: it has no flush_worker at all, so its only sender is
the uninitialised one and 100% of its events were mislabelled.
Two changes. record() stamps source beside harness, so the sending process stops
mattering — flush() already spreads per-event properties last, so a per-event
source wins over any sender default. And the build generates core/_harness_id.py
per host, seeding both modules at import, so identity no longer depends on an
entrypoint remembering to call init(). The build already computed HARNESS_ID and
spent it only on skill templating, and bundle_drift already diffs core/
byte-for-byte, so --check catches drift for free.
Deliberately not adding MEM0_PLUGIN_HARNESS to the six manifests: they sit
outside the --sync and --check boundary, which is the property that caused this.
Also unifies two defaults that disagreed. configure_harness derived
`<host>_plugin` while telemetry.init derived `MEM0_<HOST>_PLUGIN`, so a third
value existed. It was unreachable only because hook_runner never calls flush();
moving source into record() would have made it live.
Events now carry a uuid so a resend can be collapsed.
The suite stayed green through all of this because the only tests live under one
host, behind a conftest that calls init() at import. New tests run in real
subprocesses with no init, and cover the portable plugin, which would pass a
native-only test vacuously.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
The plugin README promises "anonymous usage events" and the telemetry module's
docstring says it sends only "salted hashes". Neither is true.
resolve_distinct_id() exchanges the API key for the account email and sends that
as the distinct_id on every event. Installing the plugin requires an API key, so
this is nearly every user. That is probably the behaviour we want — the Python
SDK and the CLI attribute the same way — but the description has to match it.
repo_hash and session_hash were unsalted SHA-256 cut to 16 hex characters.
repo.identity is a git remote URL, or `local:<absolute path>` when there is no
remote, which normally contains the account username. Sixteen unsalted hex
characters over that input space is enumerable, so the hash was not a privacy
control at all.
Salted per install, with the salt kept in the identity file. That preserves
every within-account join the analytics actually use and gives up only
cross-machine joins on the same repository, which nothing computes. Since the
distinct_id is already the email, the hash was never buying privacy from us —
only from whoever obtains the data later, which is exactly what the salt fixes.
Also corrects deepseek-plugin's README and source comment, which told readers
ZAPIER and STRANDS were already in the backend's KNOWN_EVENT_SOURCES allowlist.
Neither was.
Adds a Telemetry section to docs/integrations/claude-code.mdx, which had none.
Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb