Commit Graph

5 Commits

Author SHA1 Message Date
Saket Aryan 5104cc3276 fix(plugins): repair the merged test file and regenerate bundles
The keep-both conflict resolution split a function body. Rebuilt from both
merge parents so the header-contract test and the session-start tests are each
intact.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:33:49 +05:30
Saket Aryan c17d336fc0 Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers
# Conflicts:
#	integrations/agent-plugin-core/tests/test_uninitialised_identity.py
2026-09-15 00:33:21 +05:30
Saket Aryan 47ce17c21b fix(integrations): apply the header contract the docs described
Review found the contract documented but not implemented, and one client path
missed entirely.

AsyncMemoryClient's custom-client branch still carried the old literal header
dict, so `AsyncMemoryClient(client=...)` sent no surface identity at all — the
exact asymmetry this work set out to remove.

Both custom-client branches also used a blanket headers.update(), which
overwrites. That is the one code path where an outer layer's identity can
physically be present, and it was the one path that erased it. They now
check-then-set the identity headers and append to an existing client stack,
which is what set-once and append-only were supposed to mean.

AGENTS.md claimed a plugin calling the Python SDK produces
`mem0-plugin/0.3.1, mem0-python/2.0.19`. Nothing in the repo sets the env vars
that would make that happen, so the concatenation was unreachable. Replaced with
the three ways an integration can actually declare itself, in preference order.

memory_core's comment said the backend reads X-Mem0-Source. That is only true
from the platform release shipping alongside this, and a reader would otherwise
trust it and build header-only attribution that silently does nothing — which is
how vercel-ai-sdk was written in the first cut. Corrected in all seven copies,
and the body value is what makes attribution work against either backend.

mem0-ts hardcoded SDK_VERSION = "3.1.8" while the repo already injects
__MEM0_SDK_VERSION__ via tsup, the same mechanism telemetry.ts uses. The
hardcode was correct only until the next release bump.

Dropped both `as never` casts in pi-agent. They suppressed an excess-property
error but also disabled checking of every other option at those call sites, so a
typo in filters or threshold would have compiled. SearchMemoryOptions now
declares `source` instead.

Stack truncation cut mid-identifier, leaving a fragment that parses as a real
client name. It now drops whole entries.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:32:39 +05:30
Saket Aryan 2c885fdcd7 fix(plugins): make code.install reachable, and stop pinging on every flush
Review found the headline fix inverted: code.install could never fire, so every
fresh install reported an upgrade and the two cohorts became indistinguishable —
strictly worse than the bug being fixed.

hook_runner reaches claim_install() only after cache_plugin_api_key() has
written `api-key` and EvidenceStore() has created `evidence.sqlite3` and its WAL
files. Asking "is the data directory empty" at that point always saw content.
The caller now snapshots emptiness at the top of the run, before anything
writes, and passes it in.

Also caught by review, all in the same file:

- claim_version_change was an unsynchronized read-modify-write, so several
  concurrently starting sessions each observed the old version and each recorded
  an upgrade. The first session after a version bump is exactly when a user's
  open agent windows all restart together. The transition is now claimed with an
  exclusive per-version sentinel.
- A crash between O_EXCL and the write left an empty marker, which disabled
  every future upgrade event on that machine: claim_install saw the file and
  claim_version_change could not parse it. An unparseable marker is now
  repaired.
- claim_install consumed the one-shot claim even under MEM0_TELEMETRY=false, so
  a user who opted out for their first sessions would never report install after
  opting in.
- Existing users have an email but no key fingerprint, so the fast path always
  missed and every flush paid an uncached /v1/ping/ — a 5s timeout each time for
  the offline users this stack keeps citing. Legacy rows now adopt the current
  key's fingerprint instead of re-resolving.
- A key that will not resolve (revoked, offline) kept attributing to the
  previous account's email, which is the bug this was meant to fix. It now falls
  back to the anonymous id.
- The anonymous id was never rotated, so once it had been merged into one
  account it was still offered as the alias for the next one. An alias naming an
  already-identified id is what could link two real people; it is now offered
  once.

The gap that let this ship was that no test drove hook_runner's session-start
path — the decision was only ever tested by calling claim_install() directly on
a directory nothing had touched. Adds subprocess tests that run the real
entrypoint: fresh install, exactly-once, and an existing data dir.

62 core tests, 203 host tests.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:31:52 +05:30
Saket Aryan f9c566aa16 fix(plugins): report the plugin that produced the event, not the one that sent it
harness is set when an event is recorded; source was set when its batch was
sent. Both came from module globals that stay at "generic" and "MEM0_PLUGIN"
until telemetry.init() runs, and two processes in the pipeline never run it:

- `python3 telemetry.py`, the detached sender spawn_flush() starts at session
  start, after every skill command, and when the MCP server exits. Everything it
  delivered was labelled source=MEM0_PLUGIN. Only batches flush_worker.py
  happened to drain got the real host.
- mcp_server.py, which records every manual search as harness=generic.

All six Python plugins ship the same files, so source could not tell any of them
apart and MCP searches from every plugin landed in one generic bucket. The
portable plugin is worse: it has no flush_worker at all, so its only sender is
the uninitialised one and 100% of its events were mislabelled.

Two changes. record() stamps source beside harness, so the sending process stops
mattering — flush() already spreads per-event properties last, so a per-event
source wins over any sender default. And the build generates core/_harness_id.py
per host, seeding both modules at import, so identity no longer depends on an
entrypoint remembering to call init(). The build already computed HARNESS_ID and
spent it only on skill templating, and bundle_drift already diffs core/
byte-for-byte, so --check catches drift for free.

Deliberately not adding MEM0_PLUGIN_HARNESS to the six manifests: they sit
outside the --sync and --check boundary, which is the property that caused this.

Also unifies two defaults that disagreed. configure_harness derived
`<host>_plugin` while telemetry.init derived `MEM0_<HOST>_PLUGIN`, so a third
value existed. It was unreachable only because hook_runner never calls flush();
moving source into record() would have made it live.

Events now carry a uuid so a resend can be collapsed.

The suite stayed green through all of this because the only tests live under one
host, behind a conftest that calls init() at import. New tests run in real
subprocesses with no init, and cover the portable plugin, which would pass a
native-only test vacuously.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:17:53 +05:30