Compare commits

...

59 Commits

Author SHA1 Message Date
Saket Aryan af7dfc8f64 merge main into pr5 after #7322 was squash-merged
Same squash divergence as pr2 and pr4. Nine conflicts, three of them real and
six generated.

build.py: kept this branch's side, which carries the portable-bundle fix main
does not have. Regenerated all six _harness_id.py from it rather than resolving
them by hand, and verified the outcome: the portable bundle declares no
application and each native one still names its host.

deepseek-plugin/src/index.ts: kept this branch's side. Main has the comment
claiming the backend allowlist already recognizes DEEPSEEK_HARNESS, which is not
true until mem0ai/platform#3602 ships; this branch carries the correction.

test_uninitialised_identity.py: append-only, as on pr4. Our side kept whole.

Bundles clean for all six hosts, 318 passed 8 skipped, deepseek and pi-agent
suites green.
2026-09-18 13:47:56 +05:30
Saket Aryan 6aa1f60503 fix(client): drop the unused import CI's lint caught
Left behind when the async construction test was removed. ruff check now passes
on mem0/ and tests/.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:55:17 +05:30
Saket Aryan 3d0521cad4 Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-17 19:44:14 +05:30
Saket Aryan d0799f1cb5 Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-17 19:44:14 +05:30
Saket Aryan afc3d02a80 fix(plugins): sweep temp files a killed process left behind
Review finding, and the adjacent case to the one this PR already fixed.
_sweep_debris collects *.partial and *.corrupt; _write_identity and
_install_salt both create telemetry-*.<pid>.tmp and unlink it in a finally,
which a SIGKILL skips. Its own docstring reasoning, that no glob in the module
matches them so nothing else ever will, applies equally.

Collected on the stale window rather than the expiry window: unlike a quarantined
batch a temp file carries nothing worth keeping for diagnosis.

276 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:44:14 +05:30
Saket Aryan df4b88685b fix(client): repair the missed _bounded_stack call site, and test that a client constructs
Reported as a blocker by an independent re-review, and correctly: the previous
commit changed _bounded_stack to take the caller's entries and our own entry
separately, updated _apply_client_headers, and missed _client_stack. That runs on
every construction path, so every MemoryClient(...) raised TypeError. A total SDK
outage, introduced by the fix for a cosmetic truncation bug.

The whole suite stayed green because nothing constructed a client. That is the
actual defect here, so the test file exists as much for the gap as for the bug:
it builds a client, asserts the header reaches it, and covers the two bounding
rules directly. Confirmed it fails against the broken call site and passes
against the repaired one.

AsyncMemoryClient is deliberately not constructed: its validation path makes a
real request to /v1/ping/, and a unit test needing the network is worse than
none. It shares _client_headers with the sync client, which is what the
construction test guards.

tests/test_client_surface_headers.py 5 passed, plugin suites 312 passed 8
skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:40:17 +05:30
Saket Aryan 3a72dfdc52 fix(integrations): reserve our own slot in the client stack, and drop whole entries
Two review findings on this PR.

All three client-stack implementations appended our entry and then trimmed to
four, so whenever a caller already sent four entries the one dropped was exactly
the one the function exists to add. We vanished from our own stack while every
caller claim survived. The character cap was worse: slicing the joined string
severs an identifier, and the platform parses the fragment as a real client, so a
truncated tail arrives as a client literally named "me". Both caps now drop whole
entries and the reserved slot is ours, in the Python SDK, the TypeScript SDK and
pi-agent. mcp-server has the same fix on the platform branch.

The deepseek comment claimed the backend's allowlist recognizes DEEPSEEK_HARNESS.
This PR introduced that wording, replacing a neutral one. It is not true until
mem0ai/platform#3602 ships, so it now states the dependency.

Two pi-agent tests: our entry survives a full caller stack, and every surviving
entry is whole rather than a severed tail. Python side verified directly, a
4-entry caller stack keeps mem0-python and long entries are dropped whole.

312 passed 8 skipped, pi-agent 96, bundles clean, TS SDK builds.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:29:13 +05:30
Saket Aryan 8e59181edf Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-17 19:28:11 +05:30
Saket Aryan a0bbf748c2 Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-17 19:28:11 +05:30
Saket Aryan 4edfaabc06 fix(plugins): collect quarantined batches instead of leaving them on disk forever
Review finding. _sweep_debris globbed only *.partial. The *.corrupt files this
PR writes when a batch cannot be decoded are matched by no glob in the module, so
they accumulated for the life of the install.

Collected on the expiry window rather than the stale window, deliberately: a
quarantined batch is the only remaining evidence of events that could not be
delivered, so someone chasing a report of missing telemetry has to be able to
find a recent one. Debris keeps the short window; it carries nothing.

One test, asserting both halves: a recent quarantine survives and an expired one
does not.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:28:08 +05:30
Saket Aryan 74ca467b28 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-17 19:27:43 +05:30
Saket Aryan 8abedca24a Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-17 19:27:43 +05:30
Saket Aryan c043e97673 fix(plugins): read the salt before minting one, and stop overclaiming backend support
Two findings from an independent review of this branch.

_install_salt went straight to create, fsync, link, unlink on every call. All but
the first process finds the salt already published, so each hook paid an fsync to
discover that, on a path documented as appending a line and returning. Hooks are
separate processes firing on every tool call inside a few-second budget.
Measured: cold process one fsync, warm process zero, same salt.

The deepseek README said the backend recognizes DEEPSEEK_HARNESS so usage
surfaces by name. It does not yet. That value, along with STRANDS, ZAPIER,
MEM0_PLUGIN, PI_AGENT and VERCEL_AI_SDK, buckets into OTHERS until
mem0ai/platform#3602 ships, so the README now states the dependency and links it.
The neighbouring comment in mem0-strands was already accurate and is unchanged:
it says recognized values live in the allowlist without claiming this one is in
it.

265 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:27:39 +05:30
Saket Aryan 92aa8a1de1 fix(vercel-ai-sdk): inject the provider version at build instead of hardcoding it
Review finding from @karthik-indla on this PR. src/mem0-utils.ts carried
PROVIDER_VERSION = "3.0.2" as a literal. It matches package.json today and
misreports the client version from the next release bump onwards, which is the
one thing X-Mem0-Client exists to carry. Every other client in the repo injects
at build: mem0-ts via __MEM0_SDK_VERSION__, the Python side via
importlib.metadata, the CLIs via __CLI_VERSION__.

tsup now defines __MEM0_PROVIDER_VERSION__ from package.json, and the source
falls back to "dev" only when run unbundled, such as in tests. Verified in the
built bundle: PROVIDER_VERSION resolves to "3.0.2" and no placeholder survives.

resolveJsonModule is enabled alongside it, matching mem0-ts, because tsc
--noEmit covers tsup.config.ts and the package.json import fails without it.

Type check clean, build clean. The jest suite's failure is pre-existing and
unrelated: it requires a live MEM0_API_KEY, confirmed by running it on a stashed
tree. Python side 311 passed, 8 skipped; pi-agent 94 passed.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 22:16:20 +05:30
Saket Aryan 2266eccf07 fix(plugins): make the install marker durable before claim_install returns
Review nit from @karthik-indla on this PR. A hard kill between the O_EXCL open
and the buffered write reaching disk left a marker that exists but parses to
nothing: is_first_run reads it as claimed, so that install is never counted, and
claim_version_change cannot read a version out of it.

Not temp-and-rename, which is what the equivalent fixes in this stack use: the
O_EXCL open is what makes this claim exclusive across concurrently starting
sessions, and a rename would clobber rather than lose the race. The content is
the part that needed making safe, so it is flushed and fsynced before the call
returns.

The recovery path stays as the backstop: _repair_install_state already rewrites
an unparseable marker so version tracking resumes.

304 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 22:15:07 +05:30
Saket Aryan 75a2952004 Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-16 22:15:07 +05:30
Saket Aryan 8a014f299a Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-16 22:14:45 +05:30
Saket Aryan 40287f6f04 fix(plugins): a batch that could not be read is not a delivered batch
Two review findings from @karthik-indla on this PR.

_drain returned (0, True) on any read failure, so a claim nothing was posted
from counted as fully delivered. flush() then carried on to the next claim as
though this one had arrived, and the single signal that says the run went badly
never fired. The two cases are now separated: undecodable content is still
quarantined and reported delivered, because there is nothing left to send and
the rest of the run should continue, while an OSError leaves the file exactly
where it is and reports undelivered. Quarantining there would discard events
over a transient filesystem error, and nothing ever re-globs .corrupt.

Retries had no time backoff. _release_claim backdated straight to
immediately-reclaimable, so two senders meeting one momentary failure could walk
a batch from attempt 0 to the limit within seconds and discard it, when a retry
a minute later would have delivered. Releases now carry a cooldown that grows
with the attempts already spent, clamped so the mtime never lands in the future
and reads as a live lease.

Four tests: an unreadable batch is neither delivered nor quarantined,
undecodable content still is quarantined so one torn file cannot block every
later claim, and attempts cannot be burned without waiting. The expiry test now
ages the file between flushes, which is the wall time a real retry waits.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 22:14:42 +05:30
Saket Aryan dddeb4572b Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-16 21:08:46 +05:30
Saket Aryan 9f8b106f8e Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-16 21:08:45 +05:30
Saket Aryan ed7b09884c Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-16 21:08:45 +05:30
Saket Aryan 0377b9a85e Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-16 21:08:45 +05:30
Saket Aryan 3fd4949040 fix(plugins): keep the salt working where hardlinks are not supported
Self-review of the atomic-publish fix. Some network mounts and container volumes
reject os.link, and the outer handler swallowed that into "no salt", which meant
repo_hash and session_hash were dropped on every run for that whole cohort. The
race being closed is narrow; losing the hashes for an entire filesystem is not a
fair trade.

Falls back to claiming the name with O_CREAT|O_EXCL and writing, which is what
this did before. The empty-file window reopens there, but it is benign now: a
reader landing in it gets "" and omits the hash for that process rather than
caching a guessable path digest, which was the actual defect.

266 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 21:08:43 +05:30
Saket Aryan 64a01ab31e fix(pi-agent): attribute the shared client, not two command call sites
Review finding from @kartik-mem0 on this PR.

Only the explicit slash commands set a source. Automatic recall at entry.ts:80,
the capture path, the memory tools and deletion all go through the same client
constructed at entry.ts:31 with no attribution, so everything except the
commands still reached the platform as generic SDK traffic. That is most of the
plugin's traffic.

Identity is now stamped once on the shared client. Deliberately by mutating
client.headers rather than through the SDK's MEM0_SOURCE environment support:
this package pins mem0ai ^3.0.7, the installed build has no such support, and
setting an environment variable it does not read would have looked like a fix
and changed nothing. Every request method in the published client sends
this.headers, so this covers all of them and keeps working when the SDK gains
the env path.

Set-once and append-only are preserved, so a wrapper that already named a
surface keeps it and the client stack accumulates rather than being replaced.
PLATFORM_SOURCE moves into the new module and commands.ts imports it; the body
source stays on those two calls because that is what the backend reads when the
header is absent.

Five tests for the header contract, plus the existing suite: 94 pi-agent tests
pass and the tsup DTS build is clean. Python side 307 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 20:37:57 +05:30
Saket Aryan 9524c238db Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-16 20:36:27 +05:30
Saket Aryan 6d89b3b33e fix(plugins): rotate the anonymous id when the account goes, and verify legacy rows
Three review findings from @kartik-mem0 on this PR.

Anonymous id reuse, reported twice and one defect. The id is offered to PostHog
as $anon_distinct_id on first sign-in and that merge is permanent, so keeping it
after a logout or a key change puts every later anonymous event on the account
that just left. It is now rotated on both routes, and 'aliased' is cleared with
it so the fresh id can be merged into whatever account comes next. Rotation is
deliberately not triggered by a plain lookup failure with no cached email: there
is no previous account to leak to, and churning ids there would fragment the
person for anyone offline on first run.

Legacy rows are verified instead of adopted. A row written before fingerprints
existed carries an email and no fingerprint; adopting the current key bound that
key to the previous account's email permanently, and every run after agreed with
itself. It now resolves once and takes the answer. If the lookup fails it keeps
the cached email and retries next flush rather than dropping a real attribution,
which is safe because the network that failed /v1/ping/ is about to fail the
PostHog POST too. My original comment justifying the shortcut claimed the check
would cost a request on every flush forever; that was wrong, the fingerprint is
stored after one success.

A failed upgrade claim is released. The sentinel was created before the marker
rewrite and left behind if the rewrite failed, so claim_version_change returned
early on every later run and that version's upgrade was never recorded again.

Five tests, covering both rotation routes, legacy verification, the firewalled
legacy case, and retrying a failed upgrade claim.

300 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 20:36:23 +05:30
Saket Aryan 88bd5f164f Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-16 20:35:23 +05:30
Saket Aryan d8c99fb405 fix(plugins): check the lease before judging a claim exhausted
Review finding from @kartik-mem0 on this PR, and the most serious one: it loses
events, which is what this PR exists to prevent.

_claim_parked judged exhaustion before liveness. Claiming a parked file bumps
its attempt count and refreshes its mtime, so the moment a sender takes the
final attempt the file looks exhausted to every other sender while its owner is
actively draining it. The second sender unlinked it, and everything in that
batch was gone.

The liveness check now runs first, so a batch under a live lease is skipped
whatever its attempt count. The cleanup is deferred, not cancelled: once the
lease lapses, the same exhausted file is reaped on a later run.

Two tests. The first walks a batch to the final attempt and asserts a second
sender neither takes it nor deletes it, and that the events are still in it. The
second asserts an abandoned exhausted batch is still discarded once its lease
lapses, which is the over-correction to guard against. Confirmed the first fails
against the previous ordering.

288 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 20:35:13 +05:30
Saket Aryan 2ed501a43c Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-16 20:33:57 +05:30
Saket Aryan 26760b00b1 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-16 20:33:57 +05:30
Saket Aryan 7710a4e180 fix(plugins): publish the salt atomically, and omit the hash when there is none
Review finding from @kartik-mem0 on this PR.

O_CREAT|O_EXCL then write leaves a window where the salt file exists and is
empty. Hooks are short-lived processes firing on every tool call and people run
several agent windows, so a concurrent reader lands in that window, reads
nothing, and falls back to a digest of the salt file's own path, memoized for
its whole run. That path is guessable, so the race silently replaced the privacy
control with something an attacker can compute, and hashed the same repository
two ways depending on timing.

The value is now written to a private temp file, fsynced, and published with
os.link, which is atomic and fails if another process already published one.
Link rather than replace, so losing the race adopts their salt instead of
clobbering it. The temp file is removed either way.

The derived fallback is gone rather than fixed. _scoped_digest returns "" when
there is no salt and record() omits the property, because an unsalted digest
over a git remote or a home-directory path is close to plaintext, and shipping
one under a name that says hash is worse than sending nothing.

Three tests: the racing reader never sees the name half-written, a second writer
adopts the first's salt and leaves no temp file, and an unwritable data
directory drops the property instead of emitting a weak one. The old test
asserted the fallback behaviour and is replaced.

265 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 20:33:54 +05:30
Saket Aryan f80d21fa06 fix(plugins): do not claim a host application the portable bundle cannot know
Review point. The portable bundle is built with host "coding-agent", and the
build wrote that straight into PLATFORM_APPLICATION, so every portable install
sent X-Application: coding-agent.

That value is not in the platform's allowlist, so it was already being dropped
server-side. The effect was the worst of both: the wire claimed we knew the
editor, the stored event recorded that we did not, and nothing said which was
right. An absent header says the same thing honestly and costs a lookup.

HARNESS_ID stays "coding-agent". It is the PostHog-side label, it is true, and
grouping portable installs together there is useful.

Native bundles are unchanged apart from the regenerated comment. Covered by two
new build tests: portable declares no application, and each native names the
host it was generated for.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 02:49:11 +05:30
Saket Aryan 1304ffd4a1 Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-15 22:44:20 +05:30
Saket Aryan b66b70c212 Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-15 22:44:20 +05:30
Saket Aryan f2f58b5a64 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-15 22:44:20 +05:30
Saket Aryan 422a9caf8d Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-15 22:44:20 +05:30
Saket Aryan f8a9d524be docs(openclaw): correct the anonymity claim to match how it identifies events
The sweep in aa770aa6 deliberately left this page alone, reasoning that
hashing the email is materially different from sending it. Reading
integrations/openclaw/telemetry.ts does not support that: distinctId() is an
unsalted sha256 of the account email, and Mem0 holds the emails it is derived
from, so recovering the account is a table join. resolveEmail() also rewrites
already-queued events onto that id, and identifyAnonymous() fires a PostHog
$identify that merges the prior random id into it for good.

That is pseudonymous, not anonymous, and it is the same mismatch between the
stated privacy posture and the wire format that this stack exists to close.
The opt-out is unchanged and still correct.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 22:44:17 +05:30
Saket Aryan 2097e29edb docs(plugins): restore the telemetry sweep this branch's merge reverted
The pr4 -> pr5 merge resolved eleven documentation files to the pre-sweep
side, undoing aa770aa6 in its entirety. Because the stack lands in order,
main would have taken the corrected wording in pr1 and then had it reverted
by pr5, leaving the shipped claim wrong again:

- all six hosts' pause skill back to "a minimal anonymous telemetry ping"
- docs/integrations/deepseek-plugin.mdx back to "Anonymous usage events"
- integrations/zapier-mem0/README.md back to advertising telemetry the app
  does not have, with an MEM0_TELEMETRY opt-out that controls nothing
- the README data-dir listing back to omitting telemetry-salt and
  install-state.json, both of which this stack creates
- the README and docs property lists back to the enumeration that drifts

Restored verbatim from pr4. No file here is in pr5's scope, and the full
pr4..pr5 diff is now surface headers only.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 22:37:16 +05:30
Saket Aryan e9cbc626c0 fix(pi-agent): widen the search options by one property instead of to never
CI caught what I could not check locally: `pnpm exec tsc --noEmit` fails with

    error TS2353: Object literal may only specify known properties,
    and 'source' does not exist in type 'SearchMemoryOptions'.

Adding `source` to SearchMemoryOptions in mem0-ts does not help here. pi-agent
resolves `mem0ai` from npm, so it typechecks against the published 3.1.8 types,
not this repo's source. The declaration still belongs in mem0-ts for the next
release; this call site needs to compile today.

Widened by exactly that one property rather than restoring `as never`, which
was the original objection: a blanket cast also disabled checking of filters,
threshold, topK and rerank on the same literal. `source` reaches the wire
through the SDK's camelToSnakeKeys spread either way.

Verified by installing the package deps and running the real gates: tsc clean,
build clean. vercel-ai-sdk typechecks clean too. Also carries the spool-test
environment fix that had not been committed in this worktree.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:42:44 +05:30
Saket Aryan 3b43c78f7b Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-15 00:37:46 +05:30
Saket Aryan e6179ac447 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-15 00:37:40 +05:30
Saket Aryan 8783c590a0 Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-15 00:37:38 +05:30
Saket Aryan aa770aa652 docs(plugins): finish the telemetry sweep across the remaining surfaces
The first pass fixed the plugin README and the module docstring but left the
same claim standing everywhere else.

- docs/integrations/deepseek-plugin.mdx still said "Anonymous usage events".
  The TS SDK's telemetryId is the raw account email, so it is not anonymous.
- The pause skill told users a "minimal anonymous telemetry ping" fires while
  paused. Same ping, same email. Corrected in the template, which regenerates
  into all six hosts.
- integrations/zapier-mem0/README.md advertised telemetry the app does not have:
  there is no telemetry code in it at all. It now says what is actually true,
  that its requests carry source="ZAPIER".
- The data directory listing is presented as exhaustive and had gone stale
  against this stack's two new files, telemetry-salt and install-state.json.

Also replaced the property enumeration in both the README and the docs page.
Review pointed out it omitted the configured model name among others — writing
a fresh exhaustive list in a PR whose whole purpose is making docs match code
reproduces the defect being fixed. It now describes the shape and points at
where the rule is actually enforced, so it cannot drift again.

Deliberately unchanged: docs/integrations/openclaw.mdx. OpenClaw hashes the
email rather than sending it, which is materially different from the plugin and
the SDK, so its claim is not wrong in the same way.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:37:33 +05:30
Saket Aryan 1282b46f9f Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-15 00:36:29 +05:30
Saket Aryan 349f77e556 fix(plugins): stop the new spool tests depending on ambient state
CI runs agent-plugin-core/tests and claude-code-plugin/tests in one pytest
process. Two things only show up in that combined run, so the suites passed
locally and failed on every push.

claude-code-plugin/tests/conftest.py sets MEM0_TELEMETRY=false at import, which
is process-wide. record() then returns early and every assertion in
test_spool_delivery.py saw an empty spool — nine failures, all reported as
"recorded nothing" rather than as a disabled feature. The fixture now pins
MEM0_TELEMETRY rather than trusting whatever collected first.

The fixture also dropped telemetry/memory_core/_harness_id from sys.modules on
teardown. That conftest imports memory_core once at collection and calls
configure_harness() on it, so a later re-import got a fresh module with default
harness config and test_memory_core failed depending on collection order. The
fixture now saves and restores those entries instead of deleting them.

Verified with CI's exact command rather than the narrower path I had been
running: 266 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:36:23 +05:30
Saket Aryan 5104cc3276 fix(plugins): repair the merged test file and regenerate bundles
The keep-both conflict resolution split a function body. Rebuilt from both
merge parents so the header-contract test and the session-start tests are each
intact.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:33:49 +05:30
Saket Aryan c17d336fc0 Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers
# Conflicts:
#	integrations/agent-plugin-core/tests/test_uninitialised_identity.py
2026-09-15 00:33:21 +05:30
Saket Aryan faa3f029a1 Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-15 00:33:09 +05:30
Saket Aryan 7bcb9bd124 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-15 00:32:58 +05:30
Saket Aryan 59627f2a6c Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time
# Conflicts:
#	integrations/agent-plugin-core/python/telemetry.py
#	integrations/antigravity-plugin/core/telemetry.py
#	integrations/claude-code-plugin/core/telemetry.py
#	integrations/codex-plugin/core/telemetry.py
#	integrations/cursor-plugin/core/telemetry.py
#	integrations/kimi-plugin/core/telemetry.py
#	integrations/mem0-agent-plugin/core/telemetry.py
2026-09-15 00:32:51 +05:30
Saket Aryan 47ce17c21b fix(integrations): apply the header contract the docs described
Review found the contract documented but not implemented, and one client path
missed entirely.

AsyncMemoryClient's custom-client branch still carried the old literal header
dict, so `AsyncMemoryClient(client=...)` sent no surface identity at all — the
exact asymmetry this work set out to remove.

Both custom-client branches also used a blanket headers.update(), which
overwrites. That is the one code path where an outer layer's identity can
physically be present, and it was the one path that erased it. They now
check-then-set the identity headers and append to an existing client stack,
which is what set-once and append-only were supposed to mean.

AGENTS.md claimed a plugin calling the Python SDK produces
`mem0-plugin/0.3.1, mem0-python/2.0.19`. Nothing in the repo sets the env vars
that would make that happen, so the concatenation was unreachable. Replaced with
the three ways an integration can actually declare itself, in preference order.

memory_core's comment said the backend reads X-Mem0-Source. That is only true
from the platform release shipping alongside this, and a reader would otherwise
trust it and build header-only attribution that silently does nothing — which is
how vercel-ai-sdk was written in the first cut. Corrected in all seven copies,
and the body value is what makes attribution work against either backend.

mem0-ts hardcoded SDK_VERSION = "3.1.8" while the repo already injects
__MEM0_SDK_VERSION__ via tsup, the same mechanism telemetry.ts uses. The
hardcode was correct only until the next release bump.

Dropped both `as never` casts in pi-agent. They suppressed an excess-property
error but also disabled checking of every other option at those call sites, so a
typo in filters or threshold would have compiled. SearchMemoryOptions now
declares `source` instead.

Stack truncation cut mid-identifier, leaving a fragment that parses as a real
client name. It now drops whole entries.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:32:39 +05:30
Saket Aryan 2c885fdcd7 fix(plugins): make code.install reachable, and stop pinging on every flush
Review found the headline fix inverted: code.install could never fire, so every
fresh install reported an upgrade and the two cohorts became indistinguishable —
strictly worse than the bug being fixed.

hook_runner reaches claim_install() only after cache_plugin_api_key() has
written `api-key` and EvidenceStore() has created `evidence.sqlite3` and its WAL
files. Asking "is the data directory empty" at that point always saw content.
The caller now snapshots emptiness at the top of the run, before anything
writes, and passes it in.

Also caught by review, all in the same file:

- claim_version_change was an unsynchronized read-modify-write, so several
  concurrently starting sessions each observed the old version and each recorded
  an upgrade. The first session after a version bump is exactly when a user's
  open agent windows all restart together. The transition is now claimed with an
  exclusive per-version sentinel.
- A crash between O_EXCL and the write left an empty marker, which disabled
  every future upgrade event on that machine: claim_install saw the file and
  claim_version_change could not parse it. An unparseable marker is now
  repaired.
- claim_install consumed the one-shot claim even under MEM0_TELEMETRY=false, so
  a user who opted out for their first sessions would never report install after
  opting in.
- Existing users have an email but no key fingerprint, so the fast path always
  missed and every flush paid an uncached /v1/ping/ — a 5s timeout each time for
  the offline users this stack keeps citing. Legacy rows now adopt the current
  key's fingerprint instead of re-resolving.
- A key that will not resolve (revoked, offline) kept attributing to the
  previous account's email, which is the bug this was meant to fix. It now falls
  back to the anonymous id.
- The anonymous id was never rotated, so once it had been merged into one
  account it was still offered as the alias for the next one. An alias naming an
  already-identified id is what could link two real people; it is now offered
  once.

The gap that let this ship was that no test drove hook_runner's session-start
path — the decision was only ever tested by calling claim_install() directly on
a directory nothing had touched. Adds subprocess tests that run the real
entrypoint: fresh install, exactly-once, and an existing data dir.

62 core tests, 203 host tests.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:31:52 +05:30
Saket Aryan bd17f2b8c9 fix(plugins): make the retry budget reachable and the rewrite durable
Review found that the first cut traded the duplicate-delivery bug for a worse
one, and disproved its own load-bearing safety claim by experiment.

Expiry was unreachable. _claim_parked touched the mtime on every re-claim and
_release_claim backdated to exactly now minus the stale threshold, so a file's
age hovered around 121 seconds and never approached the 7-day expiry. The
attempt count in the filename therefore bounded nothing: an undeliverable batch
(revoked key, proxy 403, oversized event) lived on disk forever, and because
spawn_flush starts a sender whenever a .sending file exists, it spawned a
detached Python process on every hook, MCP call and CLI invocation, forever.
The old code self-healed here, so this was a regression. Expiry now gates on the
attempt budget, which is the thing that actually accumulates; age stays only as
a backstop for files that never carried an attempt marker.

The attempt parser sniffed for a leading "a", which also matches a hex id like
a1234567, so a legacy telemetry-<pid>-<hex>.sending file parsed as attempt
1234567 and was deleted unsent on the first flush after upgrade — precisely the
population this PR is meant to protect. Anchored on field position instead.

The rewrite was not durable: no fsync before the rename, and _drain unlinked any
claim that parsed to zero events. A crash between write and rename left the
claim empty, and the next flush deleted it. Now fsynced, and a non-empty file
that parses to nothing is quarantined as .corrupt rather than destroyed.

read_text raises UnicodeDecodeError on a torn file, which `except OSError` does
not catch. flush() runs from a bare `finally:` in flush_worker, so the exception
also skipped the handoff cleanup and left it stuck in .running.

The per-batch rewrite's return value was discarded, so a failed rewrite let the
loop continue as though progress had been recorded — reintroducing the exact
duplicate delivery this PR exists to fix.

.partial files orphaned by a crash between write and rename matched no glob in
the module and were never cleaned up.

Also replaces the heartbeat test, which asserted `SEND_TIMEOUT * 4 <
CLAIM_STALE_SECONDS` — two constants, executing none of the code under test. It
now drives the real rewrite and watches the mtime move. New tests cover expiry
being reachable, legacy filename parsing, torn-claim quarantine, failed-rewrite
behaviour and debris sweeping.

64 core tests, 199 host tests.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:30:58 +05:30
Saket Aryan 95d4fc27e2 fix(plugins): make the telemetry salt stable, its own file, and memoized
Review found three ways the first cut produced worse data than no salt at all.
All three came from keeping the salt as a key in the identity dict and doing an
unlocked read-modify-write.

Hooks are short-lived separate processes firing on every tool call, and people
run more than one agent window, so several processes would read {}, each mint
its own uuid4, and each hash with it. One repository hashed several ways in the
window before a writer won.

resolve_distinct_id holds a copy of that same dict across a network call to
/v1/ping/ with a 5s timeout, so whichever write landed second erased the other's
key: losing the salt changes repo_hash mid-stream, losing the email fires a
second $identify and splits the person.

_write_identity swallows OSError, and nothing memoized, so on a read-only or
full data directory every single event got a brand-new random salt — unbounded
cardinality in PostHog, which is strictly worse than the unsalted value it
replaced.

The salt now lives in its own file claimed with O_CREAT|O_EXCL, so exactly one
process wins and the losers read the winner's value, and it is memoized per
process. When it cannot be persisted the fallback is derived from the data
directory path: stable for the machine rather than random per call.

Its own file also means record() no longer creates telemetry-identity.json as a
side effect. is_first_run keys off that file, so the first cut would have
silently suppressed the install event — a production metric change hidden in a
docs PR.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:25:12 +05:30
Saket Aryan bc5526f13d feat(integrations): declare which surface each client is, and its version
Nothing on the wire said which Mem0 surface made a call. Both SDKs sent only an
auth header, so the platform saw python-httpx and axios and attributed every
plugin, wrapper and direct API user to one undifferentiated bucket. Version was
unknowable, which is what gates every deprecation decision.

Three headers, and the rules on them are the point:

- X-Mem0-Source and X-Application are SET-ONCE. Whichever layer is outermost
  sets them; nothing below overwrites. A plugin wrapping the SDK keeps its own
  identity instead of being renamed by the transport underneath it.
- X-Mem0-Client is APPEND-ONLY. A plugin calling the Python SDK produces
  `mem0-plugin/0.3.1, mem0-python/2.0.19`, so neither layer can erase the other.

Deliberately not User-Agent: proxies rewrite it, and we have already met a WAF
that 403s on it.

The plugin core also hoists `source` out of metadata to the top level, which is
where the backend actually reads it. It sat in metadata, which get_event_source
never consults, so all six plugins arrived indistinguishable from a raw SDK call
no matter what they set. The harness tag stays in metadata as hook provenance.

pi-agent had PI_AGENT as a PostHog property only and never sent it on the wire.
vercel-ai-sdk sent nothing at all from its raw fetch calls.

Values must exist in the platform's EventSource enum or they bucket to OTHERS,
so integrations/AGENTS.md now states the contract and the "adding an
integration" checklist requires landing the platform value in the same week.

Pairs with mem0ai/platform#3602, which recognizes these values.

TypeScript changes are not typechecked locally — deps are not installed for
those packages. CI covers them.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:19:00 +05:30
Saket Aryan 0d2b20c03d fix(plugins): count installs once, and re-resolve the email when the key changes
code.install counted upgrades and repeat sessions. Session start records install
whenever is_first_run() is true, and that only checked whether
telemetry-identity.json exists. Recording install does not create that file —
only the first successful flush does. So install fired for every 0.2.x user on
their first 0.3.x session (0.2.x never wrote the file, and the data directory
survives the upgrade), again for any session starting before that first flush
finished, and — this is the part that makes it unbounded rather than a race —
on every single session, forever, for anyone whose flush never succeeds. An
offline or firewalled user reported a new install every time they opened an
editor, which is exactly the population hardest to see in the data.

A dedicated install-state.json is now claimed with O_CREAT|O_EXCL at the moment
install is recorded, so two sessions starting together cannot both win, and the
marker is not coupled to identity. Deliberately not the identity file: writing
that from a recording process would race the sender, which writes it during
resolve_distinct_id, and overloading it is what caused this.

Upgrade detection keys on the data directory already having content. A fresh
install has an empty one; anything else predates this session. That is a firmer
predicate than looking for 0.2.x's venv/ and requirements.txt, which is a guess
about files another part of the plugin may or may not have written and only ever
works for this one upgrade. The version is stored in the marker so later changes
record code.upgrade with a real from_version.

A cached email outlived an API key change. resolve_distinct_id kept the first
email it resolved and never looked again, so switching to a key from another
account kept attributing events to the previous one. It now stores a fingerprint
of the key the email came from and re-resolves when the current key differs, and
falls back to the anonymous id when no key is configured rather than continuing
to attribute to an account it cannot verify.

The dangerous part is the alias. resolve_distinct_id's second return value
becomes a PostHog $identify with $anon_distinct_id, and aliasing one account
email to another merges two real person profiles irreversibly. The re-resolve
path returns no alias; aliasing runs anonymous to email only, and never
email to email.

One existing test asserted that is_first_run flips when the identity file is
written, which is the defect itself. Rewritten, along with coverage for atomic
claiming, upgrade detection and version changes.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:18:38 +05:30
Saket Aryan cfdfe40e09 fix(plugins): stop delivering telemetry events twice, and stop losing parked ones
Two defects, one cause: the spool protocol infers ownership instead of holding
it, and never records progress.

Duplicate delivery after a partial failure. flush() posts the claim in batches of
100 and returns on the first failure, keeping the whole file. The retry then
posts every batch again, including the ones that already arrived — 150 recorded
events were delivered 250 times. Progress is now written back to the claim after
each successful batch, so a retry resumes where the send stopped and a crash
repeats at most one batch.

Duplicate delivery when two senders overlap. spool.replace(claim) is os.rename,
which preserves mtime, so a claim created after a quiet minute inherited the
spool's last-write time and looked abandoned the instant it existed. A second
sender starting while the first was still posting took it over and sent it too —
most likely at session end, when the MCP server's exit sender and the SessionEnd
flush worker both drain. Claims are now touched at claim time, and the per-batch
rewrite doubles as a lease heartbeat. _post makes one attempt with SEND_TIMEOUT
and no retry, so a heartbeat lands well inside the 120s lease; a test asserts
that margin so adding a retry loop to _post cannot silently break it.

Parked batches starved. _claim_spool only looked at parked .sending files when
no spool existed, and because sessions keep recording there usually was one — so
a batch parked by a failed send waited until the 7-day expiry deleted it unsent,
despite its own presence being what starts the sender in the first place.
flush() now drains the live spool and then parked claims in the same run, oldest
first, bounded. Expiry applies only after a genuine retry has failed, with the
attempt count carried in the filename.

A sender that gives up releases its lease rather than heartbeating on the way
out, so the next run picks the batch up promptly instead of waiting a full stale
window for a batch nobody is working on. A failing send stops the run, so one
broken connection cannot burn every parked batch's attempt budget at once.

Two existing tests asserted the old lifecycle and are updated in place, each
with a comment saying what changed.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:18:21 +05:30
Saket Aryan f9c566aa16 fix(plugins): report the plugin that produced the event, not the one that sent it
harness is set when an event is recorded; source was set when its batch was
sent. Both came from module globals that stay at "generic" and "MEM0_PLUGIN"
until telemetry.init() runs, and two processes in the pipeline never run it:

- `python3 telemetry.py`, the detached sender spawn_flush() starts at session
  start, after every skill command, and when the MCP server exits. Everything it
  delivered was labelled source=MEM0_PLUGIN. Only batches flush_worker.py
  happened to drain got the real host.
- mcp_server.py, which records every manual search as harness=generic.

All six Python plugins ship the same files, so source could not tell any of them
apart and MCP searches from every plugin landed in one generic bucket. The
portable plugin is worse: it has no flush_worker at all, so its only sender is
the uninitialised one and 100% of its events were mislabelled.

Two changes. record() stamps source beside harness, so the sending process stops
mattering — flush() already spreads per-event properties last, so a per-event
source wins over any sender default. And the build generates core/_harness_id.py
per host, seeding both modules at import, so identity no longer depends on an
entrypoint remembering to call init(). The build already computed HARNESS_ID and
spent it only on skill templating, and bundle_drift already diffs core/
byte-for-byte, so --check catches drift for free.

Deliberately not adding MEM0_PLUGIN_HARNESS to the six manifests: they sit
outside the --sync and --check boundary, which is the property that caused this.

Also unifies two defaults that disagreed. configure_harness derived
`<host>_plugin` while telemetry.init derived `MEM0_<HOST>_PLUGIN`, so a third
value existed. It was unreachable only because hook_runner never calls flush();
moving source into record() would have made it live.

Events now carry a uuid so a resend can be collapsed.

The suite stayed green through all of this because the only tests live under one
host, behind a conftest that calls init() at import. New tests run in real
subprocesses with no init, and cover the portable plugin, which would pass a
native-only test vacuously.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:17:53 +05:30
Saket Aryan 0d37619f24 fix(plugins): say what telemetry actually sends, and salt the hashes
The plugin README promises "anonymous usage events" and the telemetry module's
docstring says it sends only "salted hashes". Neither is true.

resolve_distinct_id() exchanges the API key for the account email and sends that
as the distinct_id on every event. Installing the plugin requires an API key, so
this is nearly every user. That is probably the behaviour we want — the Python
SDK and the CLI attribute the same way — but the description has to match it.

repo_hash and session_hash were unsalted SHA-256 cut to 16 hex characters.
repo.identity is a git remote URL, or `local:<absolute path>` when there is no
remote, which normally contains the account username. Sixteen unsalted hex
characters over that input space is enumerable, so the hash was not a privacy
control at all.

Salted per install, with the salt kept in the identity file. That preserves
every within-account join the analytics actually use and gives up only
cross-machine joins on the same repository, which nothing computes. Since the
distinct_id is already the email, the hash was never buying privacy from us —
only from whoever obtains the data later, which is exactly what the salt fixes.

Also corrects deepseek-plugin's README and source comment, which told readers
ZAPIER and STRANDS were already in the backend's KNOWN_EVENT_SOURCES allowlist.
Neither was.

Adds a Telemetry section to docs/integrations/claude-code.mdx, which had none.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:17:26 +05:30
31 changed files with 812 additions and 57 deletions
+2 -1
View File
@@ -31,7 +31,8 @@ export class PlatformBackend implements Backend {
this.headers = {
Authorization: `Token ${config.apiKey}`,
"Content-Type": "application/json",
"X-Mem0-Source": "cli",
"X-Mem0-Source": "CLI",
"X-Mem0-Client": `mem0-cli-node/${CLI_VERSION}`,
"X-Mem0-Client-Language": "node",
"X-Mem0-Client-Version": CLI_VERSION,
};
+2 -1
View File
@@ -27,7 +27,8 @@ class PlatformBackend(Backend):
headers={
"Authorization": f"Token {config.api_key}",
"Content-Type": "application/json",
"X-Mem0-Source": "cli",
"X-Mem0-Source": "CLI",
"X-Mem0-Client": f"mem0-cli-python/{__version__}",
"X-Mem0-Client-Language": "python",
"X-Mem0-Client-Version": __version__,
},
+43
View File
@@ -49,6 +49,48 @@ Run the type check after every TypeScript change: `pnpm run typecheck` or `tsc -
- **`zapier-mem0/`** is a Zapier Platform CLI app: add, search, get, delete. It deploys to Zapier, not npm, so it is **not** in the release router. Deploy it with `gh workflow run zapier-mem0-cd.yml --ref main` (needs the `ZAPIER_DEPLOY_KEY` secret).
- **`mem0-strands/`** is a native Strands `MemoryStore` (Python, published to PyPI as `mem0-strands`). It plugs into the Strands `MemoryManager` for automatic recall and server-side extraction, over the hosted Mem0 platform or self-hosted Mem0 OSS. The package lives under `mem0-strands/python/`.
## Surface attribution
Every integration tells the Mem0 platform which surface it is. Three headers,
and the rules on them are what keep one layer from erasing another:
| Header | Carries | Rule |
|--------|---------|------|
| `X-Mem0-Source` | one canonical source value | **set-once** — write only if absent |
| `X-Application` | the host app it runs inside | **set-once** — write only if absent |
| `X-Mem0-Client` | `name/version`, outermost first | **append-only** — add yourself, never replace |
Set-once means check-then-set, never assignment. An integration that wraps the
SDK is the outermost layer and sets the source; the SDK underneath defers to it.
Assignment is exactly how every agent plugin came to be indistinguishable from
every other one at the platform.
How to declare it from an integration, in order of preference:
1. Send the headers yourself, if you make the HTTP call directly.
2. Pass `source` in the call options, if you go through an SDK.
3. Set `MEM0_SOURCE` / `MEM0_APPLICATION` / `MEM0_CLIENT_STACK` in the
environment before constructing the client. The SDKs read these and defer to
anything already present.
Append-only applies where a stack can actually form: an SDK handed a client that
already carries `X-Mem0-Client` appends itself rather than replacing. An SDK
constructed with no outer context simply reports itself, which is correct — it
is the outermost layer in that process.
The backend recognizes a fixed list of source values and buckets everything else
into `OTHERS`. A new value has to land in the platform's `EventSource` enum, so
do not invent one without that change going in too.
`X-Application` is allowlisted the same way, and this one has a rule of its own:
**omit the header when you do not know the host.** A value outside the allowlist
is discarded server-side, so guessing produces an event that claims an
attribution we do not actually have. The portable bundle is the case that
matters. It runs in whatever editor a user drops it into, so its build leaves
`PLATFORM_APPLICATION` empty and `memory_core` sends no header at all, while the
native bundles each name the host they were generated for. If you add a build
target, decide which of those two it is.
## Adding an integration
1. For a native coding-agent host, add `integrations/<name>-plugin/` with `plugin-build.json`, its manifest, and a thin adapter, then generate its shared runtime. Portable clients use the single `mem0-agent-plugin/` package. Independent TypeScript integrations stay self-contained and import shared lifecycle behavior from `agent-plugin-core/typescript/`.
@@ -59,3 +101,4 @@ Run the type check after every TypeScript change: `pnpm run typecheck` or `tsc -
5. If it is a Claude Code or editor marketplace plugin, register the generated native bundle path in the applicable marketplace files. Preserve the existing public plugin name.
6. Document it under `docs/integrations/` and add the page to `docs/docs.json` and `docs/llms.txt`.
7. Add rows to the table above and to the CI/CD tables in [`../.github/AGENTS.md`](../.github/AGENTS.md).
8. Send the three headers in [Surface attribution](#surface-attribution), and land the matching `EventSource` value on the platform in the same week. Until it exists, your traffic reports as `OTHERS`.
+13 -3
View File
@@ -81,15 +81,23 @@ def replace_output(staged: Path, output: Path) -> Path:
return output
def _render_harness_id(host: str) -> str:
def _render_harness_id(host: str, *, portable: bool = False) -> str:
"""Emit core/_harness_id.py for one host.
Carries both vocabularies from a single definition: the PostHog `source` tag
and the platform's X-Mem0-Source / X-Application pair. Keeping them together
is what stops the two from drifting into separate vocabularies for the same
thing.
The portable bundle runs in whatever editor a user drops it into, so it does
not know its host and must not guess one. HARNESS_ID stays "coding-agent",
which is true and useful for grouping in PostHog, but PLATFORM_APPLICATION is
left empty: X-Application names a real host app, is checked against an
allowlist server-side, and a value that is always discarded is worse than no
value -- it reads like an attribution we have and do not.
"""
tag = host.upper().replace("-", "_") + "_PLUGIN"
application = "" if portable else host
return (
'"""Generated by integrations/agent-plugin-core/build/build.py. Do not edit."""\n'
"\n"
@@ -98,8 +106,10 @@ def _render_harness_id(host: str) -> str:
"\n"
"# Platform-side vocabulary (mem0_event.source + X-Application). The whole\n"
"# plugin family is one source; which editor it runs in is the application.\n"
"# An empty application means the host is unknown, and memory_core omits\n"
"# the header entirely rather than sending a placeholder.\n"
'PLATFORM_SOURCE = "MEM0_PLUGIN"\n'
f'PLATFORM_APPLICATION = "{host}"\n'
f'PLATFORM_APPLICATION = "{application}"\n'
)
@@ -122,7 +132,7 @@ def _bundle_python(
# to call telemetry.init(). mcp_server.py and the detached telemetry.py sender
# never did, which is how MCP searches reported harness=generic and every
# batch they drained was labelled MEM0_PLUGIN regardless of the real host.
(core / "_harness_id.py").write_text(_render_harness_id(host), encoding="utf-8")
(core / "_harness_id.py").write_text(_render_harness_id(host, portable=portable), encoding="utf-8")
values = {
"PLUGIN_ROOT": plugin_root,
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
return batches
# Platform surface attribution. Read from the generated per-host module so a new
# entrypoint is correct without remembering to configure anything.
try: # pragma: no cover - absent only in the un-built shared source tree
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
def platform_headers(key: str) -> dict[str, str]:
"""Auth plus the three surface-identity headers.
X-Mem0-Source and X-Application are set-once by contract: this is the
outermost layer, so it sets them, and nothing below may overwrite them.
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
"""
headers = {
"Authorization": f"Token {key}",
"Content-Type": "application/json",
"X-Mem0-Source": _PLATFORM_SOURCE,
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
}
if _PLATFORM_APPLICATION:
headers["X-Application"] = _PLATFORM_APPLICATION
return headers
def _request_json(
url: str, key: str, payload: dict[str, Any], timeout: float
) -> tuple[dict[str, Any] | list[Any], int, int]:
@@ -1807,7 +1835,7 @@ def _request_json(
request = urllib.request.Request(
url,
data=raw,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="POST",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1834,7 +1862,7 @@ def _get_json(
) -> tuple[dict[str, Any] | list[Any], int]:
request = urllib.request.Request(
url,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="GET",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1980,6 +2008,13 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
@@ -2523,7 +2558,7 @@ def _collect_memory_ids(
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
request = urllib.request.Request(
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="DELETE",
)
try:
@@ -70,6 +70,38 @@ def test_portable_bundle_is_conformant_and_self_contained(tmp_path: Path) -> Non
assert not any(path.is_symlink() for path in root.rglob("*"))
def _harness_identity(root: Path) -> dict[str, str]:
"""Read the generated core/_harness_id.py without importing it."""
values: dict[str, str] = {}
for line in (root / "core" / "_harness_id.py").read_text(encoding="utf-8").splitlines():
if "=" in line and not line.lstrip().startswith("#"):
name, _, raw = line.partition("=")
values[name.strip()] = raw.strip().strip('"')
return values
def test_the_portable_bundle_declares_no_host_application(tmp_path: Path) -> None:
"""It runs in whatever editor a user drops it into, so it cannot know the host.
X-Application is allowlisted server-side. A guessed value is silently dropped
there, which is the worst outcome: the wire says we know the host and the
stored event says we do not.
"""
identity = _harness_identity(build("mem0-agent-plugin", "portable", tmp_path / "portable"))
assert identity["PLATFORM_APPLICATION"] == ""
# The PostHog-side label is still useful for grouping and stays populated.
assert identity["HARNESS_ID"] == "coding-agent"
assert identity["PLATFORM_SOURCE"] == "MEM0_PLUGIN"
@pytest.mark.parametrize("host", ["claude-code", "cursor", "codex", "kimi", "antigravity"])
def test_a_native_bundle_names_the_host_it_was_built_for(host: str, tmp_path: Path) -> None:
identity = _harness_identity(build(host, "native", tmp_path / host))
assert identity["PLATFORM_APPLICATION"] == host
@pytest.mark.parametrize("host", ["claude-code", "cursor", "codex", "kimi", "antigravity"])
def test_native_bundle_is_self_contained(host: str, tmp_path: Path) -> None:
root = build(host, "native", tmp_path / host)
@@ -179,6 +179,33 @@ def test_source_tag_defaults_agree_between_the_two_modules():
assert left == right == "KIMI_PLUGIN"
def test_the_plugin_declares_its_surface_in_the_body_and_the_headers():
"""Body and headers both, because only the body works on every backend."""
core = _core_dir("claude-code-plugin")
if not core.exists():
pytest.skip("claude-code-plugin is not built in this tree")
with tempfile.TemporaryDirectory() as tmp:
out = _run(
core,
Path(tmp),
"import json, memory_core\n"
"h = memory_core.platform_headers('k')\n"
"print(json.dumps({'source': h.get('X-Mem0-Source'),"
" 'app': h.get('X-Application'),"
" 'client': h.get('X-Mem0-Client'),"
" 'auth': h.get('Authorization'),"
" 'ctype': h.get('Content-Type')}))",
)
headers = json.loads(out)
assert headers["source"] == "MEM0_PLUGIN"
assert headers["app"] == "claude-code"
assert headers["client"].startswith("mem0-plugin/")
# The transport headers the three call sites relied on must survive.
assert headers["auth"] == "Token k"
assert headers["ctype"] == "application/json"
def _session_start(core: Path, data_dir: Path) -> list[str]:
"""Drive the real hook_runner session-start path and return lifecycle events."""
recorded = "\n".join(
@@ -5,5 +5,7 @@ SOURCE_TAG = "ANTIGRAVITY_PLUGIN"
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
# plugin family is one source; which editor it runs in is the application.
# An empty application means the host is unknown, and memory_core omits
# the header entirely rather than sending a placeholder.
PLATFORM_SOURCE = "MEM0_PLUGIN"
PLATFORM_APPLICATION = "antigravity"
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
return batches
# Platform surface attribution. Read from the generated per-host module so a new
# entrypoint is correct without remembering to configure anything.
try: # pragma: no cover - absent only in the un-built shared source tree
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
def platform_headers(key: str) -> dict[str, str]:
"""Auth plus the three surface-identity headers.
X-Mem0-Source and X-Application are set-once by contract: this is the
outermost layer, so it sets them, and nothing below may overwrite them.
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
"""
headers = {
"Authorization": f"Token {key}",
"Content-Type": "application/json",
"X-Mem0-Source": _PLATFORM_SOURCE,
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
}
if _PLATFORM_APPLICATION:
headers["X-Application"] = _PLATFORM_APPLICATION
return headers
def _request_json(
url: str, key: str, payload: dict[str, Any], timeout: float
) -> tuple[dict[str, Any] | list[Any], int, int]:
@@ -1807,7 +1835,7 @@ def _request_json(
request = urllib.request.Request(
url,
data=raw,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="POST",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1834,7 +1862,7 @@ def _get_json(
) -> tuple[dict[str, Any] | list[Any], int]:
request = urllib.request.Request(
url,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="GET",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1980,6 +2008,13 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
@@ -2523,7 +2558,7 @@ def _collect_memory_ids(
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
request = urllib.request.Request(
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="DELETE",
)
try:
@@ -5,5 +5,7 @@ SOURCE_TAG = "CLAUDE_CODE_PLUGIN"
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
# plugin family is one source; which editor it runs in is the application.
# An empty application means the host is unknown, and memory_core omits
# the header entirely rather than sending a placeholder.
PLATFORM_SOURCE = "MEM0_PLUGIN"
PLATFORM_APPLICATION = "claude-code"
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
return batches
# Platform surface attribution. Read from the generated per-host module so a new
# entrypoint is correct without remembering to configure anything.
try: # pragma: no cover - absent only in the un-built shared source tree
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
def platform_headers(key: str) -> dict[str, str]:
"""Auth plus the three surface-identity headers.
X-Mem0-Source and X-Application are set-once by contract: this is the
outermost layer, so it sets them, and nothing below may overwrite them.
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
"""
headers = {
"Authorization": f"Token {key}",
"Content-Type": "application/json",
"X-Mem0-Source": _PLATFORM_SOURCE,
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
}
if _PLATFORM_APPLICATION:
headers["X-Application"] = _PLATFORM_APPLICATION
return headers
def _request_json(
url: str, key: str, payload: dict[str, Any], timeout: float
) -> tuple[dict[str, Any] | list[Any], int, int]:
@@ -1807,7 +1835,7 @@ def _request_json(
request = urllib.request.Request(
url,
data=raw,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="POST",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1834,7 +1862,7 @@ def _get_json(
) -> tuple[dict[str, Any] | list[Any], int]:
request = urllib.request.Request(
url,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="GET",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1980,6 +2008,13 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
@@ -2523,7 +2558,7 @@ def _collect_memory_ids(
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
request = urllib.request.Request(
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="DELETE",
)
try:
@@ -5,5 +5,7 @@ SOURCE_TAG = "CODEX_PLUGIN"
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
# plugin family is one source; which editor it runs in is the application.
# An empty application means the host is unknown, and memory_core omits
# the header entirely rather than sending a placeholder.
PLATFORM_SOURCE = "MEM0_PLUGIN"
PLATFORM_APPLICATION = "codex"
+38 -3
View File
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
return batches
# Platform surface attribution. Read from the generated per-host module so a new
# entrypoint is correct without remembering to configure anything.
try: # pragma: no cover - absent only in the un-built shared source tree
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
def platform_headers(key: str) -> dict[str, str]:
"""Auth plus the three surface-identity headers.
X-Mem0-Source and X-Application are set-once by contract: this is the
outermost layer, so it sets them, and nothing below may overwrite them.
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
"""
headers = {
"Authorization": f"Token {key}",
"Content-Type": "application/json",
"X-Mem0-Source": _PLATFORM_SOURCE,
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
}
if _PLATFORM_APPLICATION:
headers["X-Application"] = _PLATFORM_APPLICATION
return headers
def _request_json(
url: str, key: str, payload: dict[str, Any], timeout: float
) -> tuple[dict[str, Any] | list[Any], int, int]:
@@ -1807,7 +1835,7 @@ def _request_json(
request = urllib.request.Request(
url,
data=raw,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="POST",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1834,7 +1862,7 @@ def _get_json(
) -> tuple[dict[str, Any] | list[Any], int]:
request = urllib.request.Request(
url,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="GET",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1980,6 +2008,13 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
@@ -2523,7 +2558,7 @@ def _collect_memory_ids(
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
request = urllib.request.Request(
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="DELETE",
)
try:
@@ -5,5 +5,7 @@ SOURCE_TAG = "CURSOR_PLUGIN"
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
# plugin family is one source; which editor it runs in is the application.
# An empty application means the host is unknown, and memory_core omits
# the header entirely rather than sending a placeholder.
PLATFORM_SOURCE = "MEM0_PLUGIN"
PLATFORM_APPLICATION = "cursor"
+38 -3
View File
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
return batches
# Platform surface attribution. Read from the generated per-host module so a new
# entrypoint is correct without remembering to configure anything.
try: # pragma: no cover - absent only in the un-built shared source tree
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
def platform_headers(key: str) -> dict[str, str]:
"""Auth plus the three surface-identity headers.
X-Mem0-Source and X-Application are set-once by contract: this is the
outermost layer, so it sets them, and nothing below may overwrite them.
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
"""
headers = {
"Authorization": f"Token {key}",
"Content-Type": "application/json",
"X-Mem0-Source": _PLATFORM_SOURCE,
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
}
if _PLATFORM_APPLICATION:
headers["X-Application"] = _PLATFORM_APPLICATION
return headers
def _request_json(
url: str, key: str, payload: dict[str, Any], timeout: float
) -> tuple[dict[str, Any] | list[Any], int, int]:
@@ -1807,7 +1835,7 @@ def _request_json(
request = urllib.request.Request(
url,
data=raw,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="POST",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1834,7 +1862,7 @@ def _get_json(
) -> tuple[dict[str, Any] | list[Any], int]:
request = urllib.request.Request(
url,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="GET",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1980,6 +2008,13 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
@@ -2523,7 +2558,7 @@ def _collect_memory_ids(
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
request = urllib.request.Request(
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="DELETE",
)
try:
+3 -2
View File
@@ -26,8 +26,9 @@ export const name = "mem0";
export const inject = ["tools", "systemPrompt"];
// Tags writes so Mem0's backend attributes them to this integration in
// telemetry. The backend's KNOWN_EVENT_SOURCES allowlist recognizes this value;
// anything outside it buckets into "OTHERS".
// telemetry. Values outside the backend's KNOWN_EVENT_SOURCES allowlist bucket
// into "OTHERS"; this one is added by mem0ai/platform#3602 and reads as OTHERS
// until that ships.
const SOURCE = "DEEPSEEK_HARNESS";
const DEFAULT_SEARCH_LIMIT = 10;
@@ -5,5 +5,7 @@ SOURCE_TAG = "KIMI_PLUGIN"
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
# plugin family is one source; which editor it runs in is the application.
# An empty application means the host is unknown, and memory_core omits
# the header entirely rather than sending a placeholder.
PLATFORM_SOURCE = "MEM0_PLUGIN"
PLATFORM_APPLICATION = "kimi"
+38 -3
View File
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
return batches
# Platform surface attribution. Read from the generated per-host module so a new
# entrypoint is correct without remembering to configure anything.
try: # pragma: no cover - absent only in the un-built shared source tree
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
def platform_headers(key: str) -> dict[str, str]:
"""Auth plus the three surface-identity headers.
X-Mem0-Source and X-Application are set-once by contract: this is the
outermost layer, so it sets them, and nothing below may overwrite them.
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
"""
headers = {
"Authorization": f"Token {key}",
"Content-Type": "application/json",
"X-Mem0-Source": _PLATFORM_SOURCE,
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
}
if _PLATFORM_APPLICATION:
headers["X-Application"] = _PLATFORM_APPLICATION
return headers
def _request_json(
url: str, key: str, payload: dict[str, Any], timeout: float
) -> tuple[dict[str, Any] | list[Any], int, int]:
@@ -1807,7 +1835,7 @@ def _request_json(
request = urllib.request.Request(
url,
data=raw,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="POST",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1834,7 +1862,7 @@ def _get_json(
) -> tuple[dict[str, Any] | list[Any], int]:
request = urllib.request.Request(
url,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="GET",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1980,6 +2008,13 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
@@ -2523,7 +2558,7 @@ def _collect_memory_ids(
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
request = urllib.request.Request(
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="DELETE",
)
try:
@@ -5,5 +5,7 @@ SOURCE_TAG = "CODING_AGENT_PLUGIN"
# Platform-side vocabulary (mem0_event.source + X-Application). The whole
# plugin family is one source; which editor it runs in is the application.
# An empty application means the host is unknown, and memory_core omits
# the header entirely rather than sending a placeholder.
PLATFORM_SOURCE = "MEM0_PLUGIN"
PLATFORM_APPLICATION = "coding-agent"
PLATFORM_APPLICATION = ""
@@ -1800,6 +1800,34 @@ def extraction_message_batches(
return batches
# Platform surface attribution. Read from the generated per-host module so a new
# entrypoint is correct without remembering to configure anything.
try: # pragma: no cover - absent only in the un-built shared source tree
from _harness_id import PLATFORM_APPLICATION as _PLATFORM_APPLICATION
from _harness_id import PLATFORM_SOURCE as _PLATFORM_SOURCE
except ImportError:
_PLATFORM_SOURCE = "MEM0_PLUGIN"
_PLATFORM_APPLICATION = ""
def platform_headers(key: str) -> dict[str, str]:
"""Auth plus the three surface-identity headers.
X-Mem0-Source and X-Application are set-once by contract: this is the
outermost layer, so it sets them, and nothing below may overwrite them.
X-Mem0-Client is append-only — anything downstream adds itself to the tail.
"""
headers = {
"Authorization": f"Token {key}",
"Content-Type": "application/json",
"X-Mem0-Source": _PLATFORM_SOURCE,
"X-Mem0-Client": f"mem0-plugin/{PLUGIN_VERSION}",
}
if _PLATFORM_APPLICATION:
headers["X-Application"] = _PLATFORM_APPLICATION
return headers
def _request_json(
url: str, key: str, payload: dict[str, Any], timeout: float
) -> tuple[dict[str, Any] | list[Any], int, int]:
@@ -1807,7 +1835,7 @@ def _request_json(
request = urllib.request.Request(
url,
data=raw,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="POST",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1834,7 +1862,7 @@ def _get_json(
) -> tuple[dict[str, Any] | list[Any], int]:
request = urllib.request.Request(
url,
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="GET",
)
with urllib.request.urlopen(request, timeout=timeout) as response:
@@ -1980,6 +2008,13 @@ def flush_session(
"user_id": write_user,
"app_id": repo.app_id,
"run_id": session_id,
# Top level, not metadata: the backend reads `source` from the body or
# the query string, never from metadata, which is where this used to
# sit. The X-Mem0-Source header is also read, but only from the
# platform release that ships alongside this change, so the body value
# is what makes attribution work on both. The harness tag stays in
# metadata as hook provenance.
"source": _PLATFORM_SOURCE,
"metadata": {**metadata, "author": write_user, "dirs": directory_chain(repo)},
"agent_custom_instructions": PROJECT_MEMORY_INSTRUCTIONS,
"custom_instructions": PERSONAL_MEMORY_INSTRUCTIONS,
@@ -2523,7 +2558,7 @@ def _collect_memory_ids(
def _delete_memory(api_url: str, key: str, memory_id: str) -> bool:
request = urllib.request.Request(
f"{api_url}/v1/memories/{urllib.parse.quote(memory_id)}/",
headers={"Authorization": f"Token {key}", "Content-Type": "application/json"},
headers=platform_headers(key),
method="DELETE",
)
try:
@@ -0,0 +1,80 @@
import { describe, expect, it } from "vitest";
import { applySurfaceHeaders, PLATFORM_APPLICATION, PLATFORM_SOURCE } from "./attribution.ts";
function client(headers: Record<string, string> = {}) {
return { headers: { Authorization: "Token k", ...headers } } as never;
}
describe("applySurfaceHeaders", () => {
it("stamps the shared client so every path is attributed, not just commands", () => {
const mem0 = client();
applySurfaceHeaders(mem0);
const headers = (mem0 as unknown as { headers: Record<string, string> }).headers;
expect(headers["X-Mem0-Source"]).toBe(PLATFORM_SOURCE);
expect(headers["X-Application"]).toBe(PLATFORM_APPLICATION);
expect(headers["X-Mem0-Client"]).toMatch(/^mem0-pi-agent\//);
expect(headers.Authorization).toBe("Token k");
});
it("defers to a surface an outer wrapper already declared", () => {
const mem0 = client({ "X-Mem0-Source": "OPENCLAW", "X-Application": "vscode" });
applySurfaceHeaders(mem0);
const headers = (mem0 as unknown as { headers: Record<string, string> }).headers;
expect(headers["X-Mem0-Source"]).toBe("OPENCLAW");
expect(headers["X-Application"]).toBe("vscode");
});
it("appends to the client stack rather than replacing it", () => {
const mem0 = client({ "X-Mem0-Client": "openclaw/2.1.0" });
applySurfaceHeaders(mem0);
const headers = (mem0 as unknown as { headers: Record<string, string> }).headers;
expect(headers["X-Mem0-Client"]).toMatch(/^openclaw\/2\.1\.0, mem0-pi-agent\//);
});
it("treats a blank header as absent", () => {
const mem0 = client({ "X-Mem0-Source": " " });
applySurfaceHeaders(mem0);
const headers = (mem0 as unknown as { headers: Record<string, string> }).headers;
expect(headers["X-Mem0-Source"]).toBe(PLATFORM_SOURCE);
});
it("bounds the stack so a long chain cannot grow the header without limit", () => {
const mem0 = client({ "X-Mem0-Client": "a/1, b/1, c/1, d/1, e/1" });
applySurfaceHeaders(mem0);
const headers = (mem0 as unknown as { headers: Record<string, string> }).headers;
expect(headers["X-Mem0-Client"].split(",").length).toBeLessThanOrEqual(4);
expect(headers["X-Mem0-Client"].length).toBeLessThanOrEqual(200);
});
});
describe("client stack bounding", () => {
it("keeps our own entry when the caller already filled the stack", () => {
// The defect: pushing then trimming to four dropped exactly the entry this
// function exists to add, so we vanished from our own stack.
const mem0 = client({ "X-Mem0-Client": "a/1, b/2, c/3, d/4" });
applySurfaceHeaders(mem0);
const stack = (mem0 as unknown as { headers: Record<string, string> }).headers["X-Mem0-Client"];
expect(stack).toMatch(/mem0-pi-agent\//);
expect(stack.split(",").length).toBeLessThanOrEqual(4);
});
it("drops whole entries at the character cap, never a fragment", () => {
const long = `${"n".repeat(90)}/1.0, ${"m".repeat(90)}/1.0, ${"o".repeat(90)}/1.0`;
const mem0 = client({ "X-Mem0-Client": long });
applySurfaceHeaders(mem0);
const stack = (mem0 as unknown as { headers: Record<string, string> }).headers["X-Mem0-Client"];
expect(stack.length).toBeLessThanOrEqual(200);
expect(stack.endsWith("/0.0.0") || /mem0-pi-agent\/[\w.\-]+$/.test(stack)).toBe(true);
// Every surviving entry is whole: name/version, no severed tail.
for (const entry of stack.split(",")) {
expect(entry.trim()).toMatch(/^[^/]+\/[^/]+$/);
}
});
});
@@ -0,0 +1,65 @@
import type MemoryClient from "mem0ai";
import * as fs from "node:fs";
/** Surface identity for this plugin, as the platform's EventSource knows it. */
export const PLATFORM_SOURCE = "PI_AGENT";
/** Host app the plugin runs inside. Allowlisted server-side. */
export const PLATFORM_APPLICATION = "pi";
const PLUGIN_VERSION = (() => {
try {
return JSON.parse(
fs.readFileSync(new URL("../package.json", import.meta.url), "utf-8"),
).version as string;
} catch {
return "unknown";
}
})();
const MAX_STACK_ENTRIES = 4;
const MAX_STACK_CHARS = 200;
/**
* Append our own entry and bound the result, dropping WHOLE entries.
*
* Neither cap cuts characters: slicing the joined string severs an identifier
* and leaves a fragment the platform parses as a real client name. And the
* reserved slot is ours, since it is the only entry this layer can vouch for.
*/
function boundedStack(callerEntries: string[], own: string): string {
const kept: string[] = [];
let budget = MAX_STACK_CHARS - own.length;
for (const entry of callerEntries.slice(0, MAX_STACK_ENTRIES - 1)) {
const cost = entry.length + ", ".length;
if (cost > budget) break;
budget -= cost;
kept.push(entry);
}
return [...kept, own].join(", ");
}
/**
* Stamp surface identity onto the shared client, once, at construction.
*
* Tagging individual call sites was not enough: automatic recall, capture, the
* memory tools and deletion all go through this same client, so everything
* except the explicit slash commands reached the platform as generic SDK
* traffic. Every request method in the SDK sends `this.headers`, so setting
* them here covers all of them.
*
* X-Mem0-Source and X-Application are set-once, so a wrapper that already named
* a surface keeps it. X-Mem0-Client is append-only, so the platform sees the
* whole chain rather than only the last speaker.
*/
export function applySurfaceHeaders(client: MemoryClient): void {
const headers = client.headers as Record<string, string>;
if (!headers["X-Mem0-Source"]?.trim()) headers["X-Mem0-Source"] = PLATFORM_SOURCE;
if (!headers["X-Application"]?.trim()) headers["X-Application"] = PLATFORM_APPLICATION;
const existing = (headers["X-Mem0-Client"] ?? "")
.split(",")
.map((part) => part.trim())
.filter(Boolean);
headers["X-Mem0-Client"] = boundedStack(existing, `mem0-pi-agent/${PLUGIN_VERSION}`);
}
+14 -2
View File
@@ -1,3 +1,4 @@
import type { SearchMemoryOptions } from "mem0ai";
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import type MemoryClient from "mem0ai";
import type { Mem0Config, ScopeContext, Scope } from "./types.ts";
@@ -5,6 +6,12 @@ import { DEFAULT_CUSTOM_CATEGORIES } from "./types.ts";
import { resolveSearchFilters, resolveAddParams } from "./memory/scoping.ts";
import { formatMemoryList, formatMemoryCompact, groupByCategory } from "./memory/formatting.ts";
import { captureCommandEvent } from "./telemetry.ts";
import { PLATFORM_SOURCE } from "./attribution.ts";
// Wire identity is set once on the shared client in entry.ts, which covers
// every path including recall, capture, tools and deletion. It stays in the
// body of the two calls below as well: body `source` is what the backend reads
// when the header is absent.
const SEARCH_TOP_K = 10;
@@ -29,7 +36,12 @@ export function registerCommands(
threshold: config.searchThreshold,
topK: SEARCH_TOP_K,
rerank: true,
});
source: PLATFORM_SOURCE,
// Widened by exactly this one property. `source` reaches the wire via the
// SDK's camelToSnakeKeys spread, but it is absent from SearchMemoryOptions
// in the published mem0ai types. A blanket `as never` would also disable
// checking of filters, threshold, topK and rerank above.
} as SearchMemoryOptions & { source: string });
return result.results ?? [];
};
@@ -45,7 +57,7 @@ export function registerCommands(
const addParams = resolveAddParams(config.defaultScope, getScopeCtx());
const result = await mem0.add(
[{ role: "user", content: text }],
{ ...addParams, customCategories: DEFAULT_CUSTOM_CATEGORIES, infer: false },
{ ...addParams, customCategories: DEFAULT_CUSTOM_CATEGORIES, infer: false, source: PLATFORM_SOURCE },
);
captureCommandEvent("mem0-remember", {}, telemetryCtx);
@@ -10,6 +10,7 @@ import { captureEvent } from "./telemetry.ts";
import * as os from "node:os";
import type { ScopeContext } from "./types.ts";
import { createMemoryLifecycle } from "../../agent-plugin-core/typescript/src/lifecycle.ts";
import { applySurfaceHeaders } from "./attribution.ts";
export { buildRecallContext } from "../../agent-plugin-core/typescript/src/lifecycle.ts";
@@ -29,6 +30,11 @@ export default function mem0Extension(pi: ExtensionAPI): void {
}
const mem0 = new MemoryClient({ apiKey: config.apiKey });
// Every path below shares this client: automatic recall, capture, the memory
// tools and deletion as well as the slash commands. Attribution belongs here
// rather than on individual calls, or everything except the commands reports
// as generic SDK traffic.
applySurfaceHeaders(mem0);
const scopeCtx: ScopeContext = {
userId: resolveUserId(config.userId),
+17 -2
View File
@@ -1,3 +1,10 @@
declare const __MEM0_PROVIDER_VERSION__: string | undefined;
// Replaced at build time by tsup `define`. The fallback only applies when the
// source is run unbundled, such as in tests.
const PROVIDER_VERSION =
typeof __MEM0_PROVIDER_VERSION__ !== "undefined" ? __MEM0_PROVIDER_VERSION__ : "dev";
import { LanguageModelV3Prompt } from '@ai-sdk/provider';
import { Mem0ConfigSettings } from './mem0-types';
import { loadApiKey } from '@ai-sdk/provider-utils';
@@ -277,7 +284,11 @@ const searchInternalMemories = async (query: string, config?: Mem0ConfigSettings
method: 'POST',
headers: {
Authorization: `Token ${apiKey}`,
'Content-Type': 'application/json'
'Content-Type': 'application/json',
// Surface attribution. Set-once by contract: this wrapper is the
// outermost layer on these raw fetch calls.
'X-Mem0-Source': 'VERCEL_AI_SDK',
'X-Mem0-Client': `mem0-vercel-ai-provider/${PROVIDER_VERSION}`
},
body: JSON.stringify(body),
};
@@ -331,7 +342,11 @@ const updateMemories = async (messages: Array<Message>, config?: Mem0ConfigSetti
method: 'POST',
headers: {
Authorization: `Token ${apiKey}`,
'Content-Type': 'application/json'
'Content-Type': 'application/json',
// Surface attribution. Set-once by contract: this wrapper is the
// outermost layer on these raw fetch calls.
'X-Mem0-Source': 'VERCEL_AI_SDK',
'X-Mem0-Client': `mem0-vercel-ai-provider/${PROVIDER_VERSION}`
},
body: JSON.stringify(body),
};
+1
View File
@@ -12,6 +12,7 @@
"noUnusedLocals": false,
"noUnusedParameters": false,
"preserveWatchOutput": true,
"resolveJsonModule": true,
"skipLibCheck": true,
"strict": true,
"types": ["@types/node", "jest"],
@@ -1,4 +1,5 @@
import { defineConfig } from 'tsup'
import pkg from './package.json'
export default defineConfig([
{
@@ -6,5 +7,11 @@ export default defineConfig([
entry: ['src/index.ts'],
format: ['cjs', 'esm'],
sourcemap: true,
// Injected rather than written in the source. A hardcoded literal matches
// package.json on the day it is written and misreports the client version
// from the next release bump onwards. Same mechanism as mem0-ts.
define: {
__MEM0_PROVIDER_VERSION__: JSON.stringify(pkg.version),
},
},
])
+61
View File
@@ -95,6 +95,66 @@ interface ClientIdentity {
const IDENTITY_CACHE_MAX_DEFAULT = 50;
const identityByCredentials = new Map<string, Promise<ClientIdentity>>();
declare const __MEM0_SDK_VERSION__: string | undefined;
// Injected by tsup (see mem0-ts/tsup.config.ts `define`), the same mechanism
// telemetry.ts already uses. A hardcoded literal goes stale at the next release
// bump and then misreports the client version forever.
const SDK_VERSION =
typeof __MEM0_SDK_VERSION__ !== "undefined" ? __MEM0_SDK_VERSION__ : "dev";
const MAX_STACK_ENTRIES = 4;
const MAX_STACK_CHARS = 200;
/**
* Append our own entry and bound the result, dropping WHOLE entries.
*
* Neither cap cuts characters: slicing the joined string severs an identifier
* and leaves a fragment the platform parses as a real client name. And the
* reserved slot is ours. Pushing first and then trimming to four dropped exactly
* the entry this exists to add whenever a caller already sent four, so we
* vanished from our own stack while every caller claim survived.
*/
function boundedStack(callerEntries: string[], own: string): string {
const kept: string[] = [];
let budget = MAX_STACK_CHARS - own.length;
for (const entry of callerEntries.slice(0, MAX_STACK_ENTRIES - 1)) {
const cost = entry.length + ", ".length;
if (cost > budget) break;
budget -= cost;
kept.push(entry);
}
return [...kept, own].join(", ");
}
/**
* Surface-identity headers.
*
* X-Mem0-Source and X-Application are SET-ONCE by contract: whichever layer is
* outermost sets them and nothing below overwrites, so a plugin wrapping this
* SDK keeps its own identity. X-Mem0-Client is APPEND-ONLY - every layer adds
* itself, so the platform sees the whole stack and not just the last speaker.
*/
function surfaceHeaders(): Record<string, string> {
const env: Record<string, string | undefined> =
typeof process !== "undefined" && process.env ? process.env : {};
const existing = (env.MEM0_CLIENT_STACK ?? "").trim();
const entries = existing
? existing
.split(",")
.map((part) => part.trim())
.filter(Boolean)
: [];
const headers: Record<string, string> = {
"X-Mem0-Client": boundedStack(entries, `mem0-js/${SDK_VERSION}`),
};
const source = (env.MEM0_SOURCE ?? "").trim();
if (source) headers["X-Mem0-Source"] = source;
const application = (env.MEM0_APPLICATION ?? "").trim();
if (application) headers["X-Application"] = application;
return headers;
}
export default class MemoryClient {
apiKey: string;
host: string;
@@ -129,6 +189,7 @@ export default class MemoryClient {
this.headers = {
Authorization: `Token ${this.apiKey}`,
"Content-Type": "application/json",
...surfaceHeaders(),
};
this.client = axios.create({
+3
View File
@@ -30,6 +30,9 @@ export interface SearchMemoryOptions {
showExpired?: boolean;
referenceDate?: string | number;
keywordSearch?: boolean;
/** Surface that produced the call, e.g. "OPENCLAW". Must be a value the
* backend's EventSource enum knows, or it buckets into OTHERS. */
source?: string;
}
export interface GetAllMemoryOptions {
+94 -24
View File
@@ -79,6 +79,95 @@ def _maybe_alias_anon_to_email(user_email):
logger.debug("Failed to alias anon telemetry to %r: %s", user_email, e)
def _sdk_version() -> str:
"""Resolved here rather than imported from the package root, which would cycle."""
try:
import importlib.metadata
return importlib.metadata.version("mem0ai")
except Exception:
return "unknown"
def _apply_client_headers(client: Any, api_key: str, user_id: str) -> None:
"""Merge our headers into a caller-supplied client without erasing theirs.
A wrapper may hand us a client already carrying its own X-Mem0-Source or a
partial X-Mem0-Client stack. Blanket update() replaced both, which is the
opposite of the set-once / append-only contract: the outermost layer is the
one whose identity should survive.
"""
existing = client.headers
mine = _client_headers(api_key, user_id)
outer_stack = existing.get("X-Mem0-Client")
if outer_stack:
entries = [part.strip() for part in str(outer_stack).split(",") if part.strip()]
mine["X-Mem0-Client"] = _bounded_stack(entries, f"mem0-python/{_sdk_version()}")
for name, value in mine.items():
if name in ("X-Mem0-Source", "X-Application") and existing.get(name):
continue
existing[name] = value
MAX_STACK_ENTRIES = 4
MAX_STACK_CHARS = 200
def _bounded_stack(caller_entries, own: str) -> str:
"""Append our own entry and bound the result, dropping WHOLE entries.
Two rules, and the second is the one that was wrong. Neither cap cuts
characters: a blunt slice severs an identifier and leaves a fragment that
parses as a real client name. And the reserved slot is OURS. Appending first
and then trimming to four dropped exactly the entry this function exists to
add, every time a caller already sent four, so the SDK vanished from its own
stack while the caller's claims all survived.
"""
kept = []
budget = MAX_STACK_CHARS - len(own)
for entry in list(caller_entries)[: MAX_STACK_ENTRIES - 1]:
cost = len(entry) + len(", ")
if cost > budget:
break
budget -= cost
kept.append(entry)
return ", ".join(kept + [own])
def _client_headers(api_key: str, user_id: str) -> Dict[str, str]:
"""Auth plus surface-identity headers.
X-Mem0-Source and X-Application are SET-ONCE by contract: whichever layer is
outermost sets them, and nothing below overwrites. A plugin or harness that
wraps this SDK therefore keeps its own identity — it declares via MEM0_SOURCE
/ MEM0_APPLICATION and the SDK defers.
X-Mem0-Client is APPEND-ONLY: every layer adds itself, so the platform sees
the whole stack rather than only whoever spoke last.
"""
headers = {
"Authorization": f"Token {api_key}",
"Mem0-User-ID": user_id,
"X-Mem0-Client": _client_stack(),
}
source = os.getenv("MEM0_SOURCE", "").strip()
if source:
headers["X-Mem0-Source"] = source
application = os.getenv("MEM0_APPLICATION", "").strip()
if application:
headers["X-Application"] = application
return headers
def _client_stack() -> str:
"""This SDK appended to any stack an outer layer already declared."""
existing = os.getenv("MEM0_CLIENT_STACK", "").strip()
entries = [part.strip() for part in existing.split(",") if part.strip()] if existing else []
return _bounded_stack(entries, f"mem0-python/{_sdk_version()}")
class MemoryClient:
"""Client for interacting with the Mem0 API.
@@ -129,19 +218,11 @@ class MemoryClient:
self.client = client
# Ensure the client has the correct base_url and headers
self.client.base_url = httpx.URL(self.host)
self.client.headers.update(
{
"Authorization": f"Token {self.api_key}",
"Mem0-User-ID": self.user_id,
}
)
_apply_client_headers(self.client, self.api_key, self.user_id)
else:
self.client = httpx.Client(
base_url=self.host,
headers={
"Authorization": f"Token {self.api_key}",
"Mem0-User-ID": self.user_id,
},
headers=_client_headers(self.api_key, self.user_id),
timeout=300,
)
self.user_email = self._validate_api_key()
@@ -1018,19 +1099,11 @@ class AsyncMemoryClient:
self.async_client = client
# Ensure the client has the correct base_url and headers
self.async_client.base_url = httpx.URL(self.host)
self.async_client.headers.update(
{
"Authorization": f"Token {self.api_key}",
"Mem0-User-ID": self.user_id,
}
)
_apply_client_headers(self.async_client, self.api_key, self.user_id)
else:
self.async_client = httpx.AsyncClient(
base_url=self.host,
headers={
"Authorization": f"Token {self.api_key}",
"Mem0-User-ID": self.user_id,
},
headers=_client_headers(self.api_key, self.user_id),
timeout=300,
)
@@ -1053,10 +1126,7 @@ class AsyncMemoryClient:
params = self._prepare_params()
response = requests.get(
f"{self.host}/v1/ping/",
headers={
"Authorization": f"Token {self.api_key}",
"Mem0-User-ID": self.user_id,
},
headers=_client_headers(self.api_key, self.user_id),
params=params,
)
response.raise_for_status()
+63
View File
@@ -0,0 +1,63 @@
"""Surface-identity headers, and that a client can be constructed at all.
The construction test exists because it was not there: a signature change to
_bounded_stack missed the _client_stack call site, every MemoryClient(...) raised
TypeError, and the whole suite stayed green because nothing built one.
"""
import os
from unittest.mock import patch
from mem0.client.main import _bounded_stack, _client_headers, _client_stack
def test_a_client_can_be_constructed():
from mem0 import MemoryClient
# _validate_api_key normally populates org/project from the API response;
# stubbing it leaves them None, which a later accessor rejects. Set them the
# way a real validation would. This test is about construction reaching the
# header stage at all.
def _stub(self):
self.org_id, self.project_id = "org", "proj"
with patch.object(MemoryClient, "_validate_api_key", _stub):
client = MemoryClient(api_key="m0-test")
assert client.client.headers["X-Mem0-Client"].startswith("mem0-python/")
# AsyncMemoryClient is deliberately not constructed here: its validation path
# makes a real request to /v1/ping/, and a unit test that needs the network is
# worse than none. It shares _client_headers with the sync client, which is the
# code the construction test above actually guards.
def test_headers_carry_this_sdk():
headers = _client_headers("m0-test", "u1")
assert headers["X-Mem0-Client"].startswith("mem0-python/")
def test_our_entry_survives_a_caller_that_already_filled_the_stack():
# Appending first and trimming to four dropped exactly the entry the
# function exists to add.
stack = _bounded_stack(["a/1", "b/2", "c/3", "d/4"], "mem0-python/9.9.9")
assert "mem0-python/9.9.9" in stack
assert len(stack.split(",")) <= 4
def test_the_character_cap_drops_whole_entries_not_characters():
long_entries = [f"{'n' * 90}/1.0", f"{'m' * 90}/1.0", "c/3"]
stack = _bounded_stack(long_entries, "mem0-python/9.9.9")
assert len(stack) <= 200
assert stack.endswith("mem0-python/9.9.9")
for entry in stack.split(","):
assert entry.strip().count("/") == 1, f"severed entry: {entry!r}"
def test_an_outer_stack_is_appended_to_not_replaced():
with patch.dict(os.environ, {"MEM0_CLIENT_STACK": "openclaw/2.1.0"}):
stack = _client_stack()
assert stack.startswith("openclaw/2.1.0")
assert "mem0-python/" in stack