Compare commits

..

59 Commits

Author SHA1 Message Date
Saket Aryan af7dfc8f64 merge main into pr5 after #7322 was squash-merged
Same squash divergence as pr2 and pr4. Nine conflicts, three of them real and
six generated.

build.py: kept this branch's side, which carries the portable-bundle fix main
does not have. Regenerated all six _harness_id.py from it rather than resolving
them by hand, and verified the outcome: the portable bundle declares no
application and each native one still names its host.

deepseek-plugin/src/index.ts: kept this branch's side. Main has the comment
claiming the backend allowlist already recognizes DEEPSEEK_HARNESS, which is not
true until mem0ai/platform#3602 ships; this branch carries the correction.

test_uninitialised_identity.py: append-only, as on pr4. Our side kept whole.

Bundles clean for all six hosts, 318 passed 8 skipped, deepseek and pi-agent
suites green.
2026-09-18 13:47:56 +05:30
Saket Aryan 6aa1f60503 fix(client): drop the unused import CI's lint caught
Left behind when the async construction test was removed. ruff check now passes
on mem0/ and tests/.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:55:17 +05:30
Saket Aryan 3d0521cad4 Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-17 19:44:14 +05:30
Saket Aryan d0799f1cb5 Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-17 19:44:14 +05:30
Saket Aryan afc3d02a80 fix(plugins): sweep temp files a killed process left behind
Review finding, and the adjacent case to the one this PR already fixed.
_sweep_debris collects *.partial and *.corrupt; _write_identity and
_install_salt both create telemetry-*.<pid>.tmp and unlink it in a finally,
which a SIGKILL skips. Its own docstring reasoning, that no glob in the module
matches them so nothing else ever will, applies equally.

Collected on the stale window rather than the expiry window: unlike a quarantined
batch a temp file carries nothing worth keeping for diagnosis.

276 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:44:14 +05:30
Saket Aryan df4b88685b fix(client): repair the missed _bounded_stack call site, and test that a client constructs
Reported as a blocker by an independent re-review, and correctly: the previous
commit changed _bounded_stack to take the caller's entries and our own entry
separately, updated _apply_client_headers, and missed _client_stack. That runs on
every construction path, so every MemoryClient(...) raised TypeError. A total SDK
outage, introduced by the fix for a cosmetic truncation bug.

The whole suite stayed green because nothing constructed a client. That is the
actual defect here, so the test file exists as much for the gap as for the bug:
it builds a client, asserts the header reaches it, and covers the two bounding
rules directly. Confirmed it fails against the broken call site and passes
against the repaired one.

AsyncMemoryClient is deliberately not constructed: its validation path makes a
real request to /v1/ping/, and a unit test needing the network is worse than
none. It shares _client_headers with the sync client, which is what the
construction test guards.

tests/test_client_surface_headers.py 5 passed, plugin suites 312 passed 8
skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:40:17 +05:30
Saket Aryan 3a72dfdc52 fix(integrations): reserve our own slot in the client stack, and drop whole entries
Two review findings on this PR.

All three client-stack implementations appended our entry and then trimmed to
four, so whenever a caller already sent four entries the one dropped was exactly
the one the function exists to add. We vanished from our own stack while every
caller claim survived. The character cap was worse: slicing the joined string
severs an identifier, and the platform parses the fragment as a real client, so a
truncated tail arrives as a client literally named "me". Both caps now drop whole
entries and the reserved slot is ours, in the Python SDK, the TypeScript SDK and
pi-agent. mcp-server has the same fix on the platform branch.

The deepseek comment claimed the backend's allowlist recognizes DEEPSEEK_HARNESS.
This PR introduced that wording, replacing a neutral one. It is not true until
mem0ai/platform#3602 ships, so it now states the dependency.

Two pi-agent tests: our entry survives a full caller stack, and every surviving
entry is whole rather than a severed tail. Python side verified directly, a
4-entry caller stack keeps mem0-python and long entries are dropped whole.

312 passed 8 skipped, pi-agent 96, bundles clean, TS SDK builds.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:29:13 +05:30
Saket Aryan 8e59181edf Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-17 19:28:11 +05:30
Saket Aryan a0bbf748c2 Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-17 19:28:11 +05:30
Saket Aryan 4edfaabc06 fix(plugins): collect quarantined batches instead of leaving them on disk forever
Review finding. _sweep_debris globbed only *.partial. The *.corrupt files this
PR writes when a batch cannot be decoded are matched by no glob in the module, so
they accumulated for the life of the install.

Collected on the expiry window rather than the stale window, deliberately: a
quarantined batch is the only remaining evidence of events that could not be
delivered, so someone chasing a report of missing telemetry has to be able to
find a recent one. Debris keeps the short window; it carries nothing.

One test, asserting both halves: a recent quarantine survives and an expired one
does not.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:28:08 +05:30
Saket Aryan 74ca467b28 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-17 19:27:43 +05:30
Saket Aryan 8abedca24a Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-17 19:27:43 +05:30
Saket Aryan c043e97673 fix(plugins): read the salt before minting one, and stop overclaiming backend support
Two findings from an independent review of this branch.

_install_salt went straight to create, fsync, link, unlink on every call. All but
the first process finds the salt already published, so each hook paid an fsync to
discover that, on a path documented as appending a line and returning. Hooks are
separate processes firing on every tool call inside a few-second budget.
Measured: cold process one fsync, warm process zero, same salt.

The deepseek README said the backend recognizes DEEPSEEK_HARNESS so usage
surfaces by name. It does not yet. That value, along with STRANDS, ZAPIER,
MEM0_PLUGIN, PI_AGENT and VERCEL_AI_SDK, buckets into OTHERS until
mem0ai/platform#3602 ships, so the README now states the dependency and links it.
The neighbouring comment in mem0-strands was already accurate and is unchanged:
it says recognized values live in the allowlist without claiming this one is in
it.

265 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-17 19:27:39 +05:30
Saket Aryan 92aa8a1de1 fix(vercel-ai-sdk): inject the provider version at build instead of hardcoding it
Review finding from @karthik-indla on this PR. src/mem0-utils.ts carried
PROVIDER_VERSION = "3.0.2" as a literal. It matches package.json today and
misreports the client version from the next release bump onwards, which is the
one thing X-Mem0-Client exists to carry. Every other client in the repo injects
at build: mem0-ts via __MEM0_SDK_VERSION__, the Python side via
importlib.metadata, the CLIs via __CLI_VERSION__.

tsup now defines __MEM0_PROVIDER_VERSION__ from package.json, and the source
falls back to "dev" only when run unbundled, such as in tests. Verified in the
built bundle: PROVIDER_VERSION resolves to "3.0.2" and no placeholder survives.

resolveJsonModule is enabled alongside it, matching mem0-ts, because tsc
--noEmit covers tsup.config.ts and the package.json import fails without it.

Type check clean, build clean. The jest suite's failure is pre-existing and
unrelated: it requires a live MEM0_API_KEY, confirmed by running it on a stashed
tree. Python side 311 passed, 8 skipped; pi-agent 94 passed.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 22:16:20 +05:30
Saket Aryan 2266eccf07 fix(plugins): make the install marker durable before claim_install returns
Review nit from @karthik-indla on this PR. A hard kill between the O_EXCL open
and the buffered write reaching disk left a marker that exists but parses to
nothing: is_first_run reads it as claimed, so that install is never counted, and
claim_version_change cannot read a version out of it.

Not temp-and-rename, which is what the equivalent fixes in this stack use: the
O_EXCL open is what makes this claim exclusive across concurrently starting
sessions, and a rename would clobber rather than lose the race. The content is
the part that needed making safe, so it is flushed and fsynced before the call
returns.

The recovery path stays as the backstop: _repair_install_state already rewrites
an unparseable marker so version tracking resumes.

304 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 22:15:07 +05:30
Saket Aryan 75a2952004 Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-16 22:15:07 +05:30
Saket Aryan 8a014f299a Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-16 22:14:45 +05:30
Saket Aryan 40287f6f04 fix(plugins): a batch that could not be read is not a delivered batch
Two review findings from @karthik-indla on this PR.

_drain returned (0, True) on any read failure, so a claim nothing was posted
from counted as fully delivered. flush() then carried on to the next claim as
though this one had arrived, and the single signal that says the run went badly
never fired. The two cases are now separated: undecodable content is still
quarantined and reported delivered, because there is nothing left to send and
the rest of the run should continue, while an OSError leaves the file exactly
where it is and reports undelivered. Quarantining there would discard events
over a transient filesystem error, and nothing ever re-globs .corrupt.

Retries had no time backoff. _release_claim backdated straight to
immediately-reclaimable, so two senders meeting one momentary failure could walk
a batch from attempt 0 to the limit within seconds and discard it, when a retry
a minute later would have delivered. Releases now carry a cooldown that grows
with the attempts already spent, clamped so the mtime never lands in the future
and reads as a live lease.

Four tests: an unreadable batch is neither delivered nor quarantined,
undecodable content still is quarantined so one torn file cannot block every
later claim, and attempts cannot be burned without waiting. The expiry test now
ages the file between flushes, which is the wall time a real retry waits.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 22:14:42 +05:30
Saket Aryan dddeb4572b Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-16 21:08:46 +05:30
Saket Aryan 9f8b106f8e Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-16 21:08:45 +05:30
Saket Aryan ed7b09884c Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-16 21:08:45 +05:30
Saket Aryan 0377b9a85e Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-16 21:08:45 +05:30
Saket Aryan 3fd4949040 fix(plugins): keep the salt working where hardlinks are not supported
Self-review of the atomic-publish fix. Some network mounts and container volumes
reject os.link, and the outer handler swallowed that into "no salt", which meant
repo_hash and session_hash were dropped on every run for that whole cohort. The
race being closed is narrow; losing the hashes for an entire filesystem is not a
fair trade.

Falls back to claiming the name with O_CREAT|O_EXCL and writing, which is what
this did before. The empty-file window reopens there, but it is benign now: a
reader landing in it gets "" and omits the hash for that process rather than
caching a guessable path digest, which was the actual defect.

266 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 21:08:43 +05:30
Saket Aryan 64a01ab31e fix(pi-agent): attribute the shared client, not two command call sites
Review finding from @kartik-mem0 on this PR.

Only the explicit slash commands set a source. Automatic recall at entry.ts:80,
the capture path, the memory tools and deletion all go through the same client
constructed at entry.ts:31 with no attribution, so everything except the
commands still reached the platform as generic SDK traffic. That is most of the
plugin's traffic.

Identity is now stamped once on the shared client. Deliberately by mutating
client.headers rather than through the SDK's MEM0_SOURCE environment support:
this package pins mem0ai ^3.0.7, the installed build has no such support, and
setting an environment variable it does not read would have looked like a fix
and changed nothing. Every request method in the published client sends
this.headers, so this covers all of them and keeps working when the SDK gains
the env path.

Set-once and append-only are preserved, so a wrapper that already named a
surface keeps it and the client stack accumulates rather than being replaced.
PLATFORM_SOURCE moves into the new module and commands.ts imports it; the body
source stays on those two calls because that is what the backend reads when the
header is absent.

Five tests for the header contract, plus the existing suite: 94 pi-agent tests
pass and the tsup DTS build is clean. Python side 307 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 20:37:57 +05:30
Saket Aryan 9524c238db Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-16 20:36:27 +05:30
Saket Aryan 6d89b3b33e fix(plugins): rotate the anonymous id when the account goes, and verify legacy rows
Three review findings from @kartik-mem0 on this PR.

Anonymous id reuse, reported twice and one defect. The id is offered to PostHog
as $anon_distinct_id on first sign-in and that merge is permanent, so keeping it
after a logout or a key change puts every later anonymous event on the account
that just left. It is now rotated on both routes, and 'aliased' is cleared with
it so the fresh id can be merged into whatever account comes next. Rotation is
deliberately not triggered by a plain lookup failure with no cached email: there
is no previous account to leak to, and churning ids there would fragment the
person for anyone offline on first run.

Legacy rows are verified instead of adopted. A row written before fingerprints
existed carries an email and no fingerprint; adopting the current key bound that
key to the previous account's email permanently, and every run after agreed with
itself. It now resolves once and takes the answer. If the lookup fails it keeps
the cached email and retries next flush rather than dropping a real attribution,
which is safe because the network that failed /v1/ping/ is about to fail the
PostHog POST too. My original comment justifying the shortcut claimed the check
would cost a request on every flush forever; that was wrong, the fingerprint is
stored after one success.

A failed upgrade claim is released. The sentinel was created before the marker
rewrite and left behind if the rewrite failed, so claim_version_change returned
early on every later run and that version's upgrade was never recorded again.

Five tests, covering both rotation routes, legacy verification, the firewalled
legacy case, and retrying a failed upgrade claim.

300 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 20:36:23 +05:30
Saket Aryan 88bd5f164f Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-16 20:35:23 +05:30
Saket Aryan d8c99fb405 fix(plugins): check the lease before judging a claim exhausted
Review finding from @kartik-mem0 on this PR, and the most serious one: it loses
events, which is what this PR exists to prevent.

_claim_parked judged exhaustion before liveness. Claiming a parked file bumps
its attempt count and refreshes its mtime, so the moment a sender takes the
final attempt the file looks exhausted to every other sender while its owner is
actively draining it. The second sender unlinked it, and everything in that
batch was gone.

The liveness check now runs first, so a batch under a live lease is skipped
whatever its attempt count. The cleanup is deferred, not cancelled: once the
lease lapses, the same exhausted file is reaped on a later run.

Two tests. The first walks a batch to the final attempt and asserts a second
sender neither takes it nor deletes it, and that the events are still in it. The
second asserts an abandoned exhausted batch is still discarded once its lease
lapses, which is the over-correction to guard against. Confirmed the first fails
against the previous ordering.

288 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 20:35:13 +05:30
Saket Aryan 2ed501a43c Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-16 20:33:57 +05:30
Saket Aryan 26760b00b1 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-16 20:33:57 +05:30
Saket Aryan 7710a4e180 fix(plugins): publish the salt atomically, and omit the hash when there is none
Review finding from @kartik-mem0 on this PR.

O_CREAT|O_EXCL then write leaves a window where the salt file exists and is
empty. Hooks are short-lived processes firing on every tool call and people run
several agent windows, so a concurrent reader lands in that window, reads
nothing, and falls back to a digest of the salt file's own path, memoized for
its whole run. That path is guessable, so the race silently replaced the privacy
control with something an attacker can compute, and hashed the same repository
two ways depending on timing.

The value is now written to a private temp file, fsynced, and published with
os.link, which is atomic and fails if another process already published one.
Link rather than replace, so losing the race adopts their salt instead of
clobbering it. The temp file is removed either way.

The derived fallback is gone rather than fixed. _scoped_digest returns "" when
there is no salt and record() omits the property, because an unsalted digest
over a git remote or a home-directory path is close to plaintext, and shipping
one under a name that says hash is worse than sending nothing.

Three tests: the racing reader never sees the name half-written, a second writer
adopts the first's salt and leaves no temp file, and an unwritable data
directory drops the property instead of emitting a weak one. The old test
asserted the fallback behaviour and is replaced.

265 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 20:33:54 +05:30
Saket Aryan f80d21fa06 fix(plugins): do not claim a host application the portable bundle cannot know
Review point. The portable bundle is built with host "coding-agent", and the
build wrote that straight into PLATFORM_APPLICATION, so every portable install
sent X-Application: coding-agent.

That value is not in the platform's allowlist, so it was already being dropped
server-side. The effect was the worst of both: the wire claimed we knew the
editor, the stored event recorded that we did not, and nothing said which was
right. An absent header says the same thing honestly and costs a lookup.

HARNESS_ID stays "coding-agent". It is the PostHog-side label, it is true, and
grouping portable installs together there is useful.

Native bundles are unchanged apart from the regenerated comment. Covered by two
new build tests: portable declares no application, and each native names the
host it was generated for.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-16 02:49:11 +05:30
Saket Aryan 1304ffd4a1 Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-15 22:44:20 +05:30
Saket Aryan b66b70c212 Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-15 22:44:20 +05:30
Saket Aryan f2f58b5a64 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-15 22:44:20 +05:30
Saket Aryan 422a9caf8d Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers 2026-09-15 22:44:20 +05:30
Saket Aryan f8a9d524be docs(openclaw): correct the anonymity claim to match how it identifies events
The sweep in aa770aa6 deliberately left this page alone, reasoning that
hashing the email is materially different from sending it. Reading
integrations/openclaw/telemetry.ts does not support that: distinctId() is an
unsalted sha256 of the account email, and Mem0 holds the emails it is derived
from, so recovering the account is a table join. resolveEmail() also rewrites
already-queued events onto that id, and identifyAnonymous() fires a PostHog
$identify that merges the prior random id into it for good.

That is pseudonymous, not anonymous, and it is the same mismatch between the
stated privacy posture and the wire format that this stack exists to close.
The opt-out is unchanged and still correct.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 22:44:17 +05:30
Saket Aryan 2097e29edb docs(plugins): restore the telemetry sweep this branch's merge reverted
The pr4 -> pr5 merge resolved eleven documentation files to the pre-sweep
side, undoing aa770aa6 in its entirety. Because the stack lands in order,
main would have taken the corrected wording in pr1 and then had it reverted
by pr5, leaving the shipped claim wrong again:

- all six hosts' pause skill back to "a minimal anonymous telemetry ping"
- docs/integrations/deepseek-plugin.mdx back to "Anonymous usage events"
- integrations/zapier-mem0/README.md back to advertising telemetry the app
  does not have, with an MEM0_TELEMETRY opt-out that controls nothing
- the README data-dir listing back to omitting telemetry-salt and
  install-state.json, both of which this stack creates
- the README and docs property lists back to the enumeration that drifts

Restored verbatim from pr4. No file here is in pr5's scope, and the full
pr4..pr5 diff is now surface headers only.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 22:37:16 +05:30
Saket Aryan e9cbc626c0 fix(pi-agent): widen the search options by one property instead of to never
CI caught what I could not check locally: `pnpm exec tsc --noEmit` fails with

    error TS2353: Object literal may only specify known properties,
    and 'source' does not exist in type 'SearchMemoryOptions'.

Adding `source` to SearchMemoryOptions in mem0-ts does not help here. pi-agent
resolves `mem0ai` from npm, so it typechecks against the published 3.1.8 types,
not this repo's source. The declaration still belongs in mem0-ts for the next
release; this call site needs to compile today.

Widened by exactly that one property rather than restoring `as never`, which
was the original objection: a blanket cast also disabled checking of filters,
threshold, topK and rerank on the same literal. `source` reaches the wire
through the SDK's camelToSnakeKeys spread either way.

Verified by installing the package deps and running the real gates: tsc clean,
build clean. vercel-ai-sdk typechecks clean too. Also carries the spool-test
environment fix that had not been committed in this worktree.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:42:44 +05:30
Saket Aryan 3b43c78f7b Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-15 00:37:46 +05:30
Saket Aryan e6179ac447 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-15 00:37:40 +05:30
Saket Aryan 8783c590a0 Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time 2026-09-15 00:37:38 +05:30
Saket Aryan aa770aa652 docs(plugins): finish the telemetry sweep across the remaining surfaces
The first pass fixed the plugin README and the module docstring but left the
same claim standing everywhere else.

- docs/integrations/deepseek-plugin.mdx still said "Anonymous usage events".
  The TS SDK's telemetryId is the raw account email, so it is not anonymous.
- The pause skill told users a "minimal anonymous telemetry ping" fires while
  paused. Same ping, same email. Corrected in the template, which regenerates
  into all six hosts.
- integrations/zapier-mem0/README.md advertised telemetry the app does not have:
  there is no telemetry code in it at all. It now says what is actually true,
  that its requests carry source="ZAPIER".
- The data directory listing is presented as exhaustive and had gone stale
  against this stack's two new files, telemetry-salt and install-state.json.

Also replaced the property enumeration in both the README and the docs page.
Review pointed out it omitted the configured model name among others — writing
a fresh exhaustive list in a PR whose whole purpose is making docs match code
reproduces the defect being fixed. It now describes the shape and points at
where the rule is actually enforced, so it cannot drift again.

Deliberately unchanged: docs/integrations/openclaw.mdx. OpenClaw hashes the
email rather than sending it, which is materially different from the plugin and
the SDK, so its claim is not wrong in the same way.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:37:33 +05:30
Saket Aryan 1282b46f9f Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-15 00:36:29 +05:30
Saket Aryan 349f77e556 fix(plugins): stop the new spool tests depending on ambient state
CI runs agent-plugin-core/tests and claude-code-plugin/tests in one pytest
process. Two things only show up in that combined run, so the suites passed
locally and failed on every push.

claude-code-plugin/tests/conftest.py sets MEM0_TELEMETRY=false at import, which
is process-wide. record() then returns early and every assertion in
test_spool_delivery.py saw an empty spool — nine failures, all reported as
"recorded nothing" rather than as a disabled feature. The fixture now pins
MEM0_TELEMETRY rather than trusting whatever collected first.

The fixture also dropped telemetry/memory_core/_harness_id from sys.modules on
teardown. That conftest imports memory_core once at collection and calls
configure_harness() on it, so a later re-import got a fresh module with default
harness config and test_memory_core failed depending on collection order. The
fixture now saves and restores those entries instead of deleting them.

Verified with CI's exact command rather than the narrower path I had been
running: 266 passed, 8 skipped.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:36:23 +05:30
Saket Aryan 5104cc3276 fix(plugins): repair the merged test file and regenerate bundles
The keep-both conflict resolution split a function body. Rebuilt from both
merge parents so the header-contract test and the session-start tests are each
intact.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:33:49 +05:30
Saket Aryan c17d336fc0 Merge branch 'pr4/install-marker-and-identity' into pr5/surface-headers
# Conflicts:
#	integrations/agent-plugin-core/tests/test_uninitialised_identity.py
2026-09-15 00:33:21 +05:30
Saket Aryan faa3f029a1 Merge branch 'pr3/spool-delivery' into pr4/install-marker-and-identity 2026-09-15 00:33:09 +05:30
Saket Aryan 7bcb9bd124 Merge branch 'pr2/source-at-record-time' into pr3/spool-delivery 2026-09-15 00:32:58 +05:30
Saket Aryan 59627f2a6c Merge branch 'pr1/telemetry-privacy-docs' into pr2/source-at-record-time
# Conflicts:
#	integrations/agent-plugin-core/python/telemetry.py
#	integrations/antigravity-plugin/core/telemetry.py
#	integrations/claude-code-plugin/core/telemetry.py
#	integrations/codex-plugin/core/telemetry.py
#	integrations/cursor-plugin/core/telemetry.py
#	integrations/kimi-plugin/core/telemetry.py
#	integrations/mem0-agent-plugin/core/telemetry.py
2026-09-15 00:32:51 +05:30
Saket Aryan 47ce17c21b fix(integrations): apply the header contract the docs described
Review found the contract documented but not implemented, and one client path
missed entirely.

AsyncMemoryClient's custom-client branch still carried the old literal header
dict, so `AsyncMemoryClient(client=...)` sent no surface identity at all — the
exact asymmetry this work set out to remove.

Both custom-client branches also used a blanket headers.update(), which
overwrites. That is the one code path where an outer layer's identity can
physically be present, and it was the one path that erased it. They now
check-then-set the identity headers and append to an existing client stack,
which is what set-once and append-only were supposed to mean.

AGENTS.md claimed a plugin calling the Python SDK produces
`mem0-plugin/0.3.1, mem0-python/2.0.19`. Nothing in the repo sets the env vars
that would make that happen, so the concatenation was unreachable. Replaced with
the three ways an integration can actually declare itself, in preference order.

memory_core's comment said the backend reads X-Mem0-Source. That is only true
from the platform release shipping alongside this, and a reader would otherwise
trust it and build header-only attribution that silently does nothing — which is
how vercel-ai-sdk was written in the first cut. Corrected in all seven copies,
and the body value is what makes attribution work against either backend.

mem0-ts hardcoded SDK_VERSION = "3.1.8" while the repo already injects
__MEM0_SDK_VERSION__ via tsup, the same mechanism telemetry.ts uses. The
hardcode was correct only until the next release bump.

Dropped both `as never` casts in pi-agent. They suppressed an excess-property
error but also disabled checking of every other option at those call sites, so a
typo in filters or threshold would have compiled. SearchMemoryOptions now
declares `source` instead.

Stack truncation cut mid-identifier, leaving a fragment that parses as a real
client name. It now drops whole entries.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:32:39 +05:30
Saket Aryan 2c885fdcd7 fix(plugins): make code.install reachable, and stop pinging on every flush
Review found the headline fix inverted: code.install could never fire, so every
fresh install reported an upgrade and the two cohorts became indistinguishable —
strictly worse than the bug being fixed.

hook_runner reaches claim_install() only after cache_plugin_api_key() has
written `api-key` and EvidenceStore() has created `evidence.sqlite3` and its WAL
files. Asking "is the data directory empty" at that point always saw content.
The caller now snapshots emptiness at the top of the run, before anything
writes, and passes it in.

Also caught by review, all in the same file:

- claim_version_change was an unsynchronized read-modify-write, so several
  concurrently starting sessions each observed the old version and each recorded
  an upgrade. The first session after a version bump is exactly when a user's
  open agent windows all restart together. The transition is now claimed with an
  exclusive per-version sentinel.
- A crash between O_EXCL and the write left an empty marker, which disabled
  every future upgrade event on that machine: claim_install saw the file and
  claim_version_change could not parse it. An unparseable marker is now
  repaired.
- claim_install consumed the one-shot claim even under MEM0_TELEMETRY=false, so
  a user who opted out for their first sessions would never report install after
  opting in.
- Existing users have an email but no key fingerprint, so the fast path always
  missed and every flush paid an uncached /v1/ping/ — a 5s timeout each time for
  the offline users this stack keeps citing. Legacy rows now adopt the current
  key's fingerprint instead of re-resolving.
- A key that will not resolve (revoked, offline) kept attributing to the
  previous account's email, which is the bug this was meant to fix. It now falls
  back to the anonymous id.
- The anonymous id was never rotated, so once it had been merged into one
  account it was still offered as the alias for the next one. An alias naming an
  already-identified id is what could link two real people; it is now offered
  once.

The gap that let this ship was that no test drove hook_runner's session-start
path — the decision was only ever tested by calling claim_install() directly on
a directory nothing had touched. Adds subprocess tests that run the real
entrypoint: fresh install, exactly-once, and an existing data dir.

62 core tests, 203 host tests.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:31:52 +05:30
Saket Aryan bd17f2b8c9 fix(plugins): make the retry budget reachable and the rewrite durable
Review found that the first cut traded the duplicate-delivery bug for a worse
one, and disproved its own load-bearing safety claim by experiment.

Expiry was unreachable. _claim_parked touched the mtime on every re-claim and
_release_claim backdated to exactly now minus the stale threshold, so a file's
age hovered around 121 seconds and never approached the 7-day expiry. The
attempt count in the filename therefore bounded nothing: an undeliverable batch
(revoked key, proxy 403, oversized event) lived on disk forever, and because
spawn_flush starts a sender whenever a .sending file exists, it spawned a
detached Python process on every hook, MCP call and CLI invocation, forever.
The old code self-healed here, so this was a regression. Expiry now gates on the
attempt budget, which is the thing that actually accumulates; age stays only as
a backstop for files that never carried an attempt marker.

The attempt parser sniffed for a leading "a", which also matches a hex id like
a1234567, so a legacy telemetry-<pid>-<hex>.sending file parsed as attempt
1234567 and was deleted unsent on the first flush after upgrade — precisely the
population this PR is meant to protect. Anchored on field position instead.

The rewrite was not durable: no fsync before the rename, and _drain unlinked any
claim that parsed to zero events. A crash between write and rename left the
claim empty, and the next flush deleted it. Now fsynced, and a non-empty file
that parses to nothing is quarantined as .corrupt rather than destroyed.

read_text raises UnicodeDecodeError on a torn file, which `except OSError` does
not catch. flush() runs from a bare `finally:` in flush_worker, so the exception
also skipped the handoff cleanup and left it stuck in .running.

The per-batch rewrite's return value was discarded, so a failed rewrite let the
loop continue as though progress had been recorded — reintroducing the exact
duplicate delivery this PR exists to fix.

.partial files orphaned by a crash between write and rename matched no glob in
the module and were never cleaned up.

Also replaces the heartbeat test, which asserted `SEND_TIMEOUT * 4 <
CLAIM_STALE_SECONDS` — two constants, executing none of the code under test. It
now drives the real rewrite and watches the mtime move. New tests cover expiry
being reachable, legacy filename parsing, torn-claim quarantine, failed-rewrite
behaviour and debris sweeping.

64 core tests, 199 host tests.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:30:58 +05:30
Saket Aryan 95d4fc27e2 fix(plugins): make the telemetry salt stable, its own file, and memoized
Review found three ways the first cut produced worse data than no salt at all.
All three came from keeping the salt as a key in the identity dict and doing an
unlocked read-modify-write.

Hooks are short-lived separate processes firing on every tool call, and people
run more than one agent window, so several processes would read {}, each mint
its own uuid4, and each hash with it. One repository hashed several ways in the
window before a writer won.

resolve_distinct_id holds a copy of that same dict across a network call to
/v1/ping/ with a 5s timeout, so whichever write landed second erased the other's
key: losing the salt changes repo_hash mid-stream, losing the email fires a
second $identify and splits the person.

_write_identity swallows OSError, and nothing memoized, so on a read-only or
full data directory every single event got a brand-new random salt — unbounded
cardinality in PostHog, which is strictly worse than the unsalted value it
replaced.

The salt now lives in its own file claimed with O_CREAT|O_EXCL, so exactly one
process wins and the losers read the winner's value, and it is memoized per
process. When it cannot be persisted the fallback is derived from the data
directory path: stable for the machine rather than random per call.

Its own file also means record() no longer creates telemetry-identity.json as a
side effect. is_first_run keys off that file, so the first cut would have
silently suppressed the install event — a production metric change hidden in a
docs PR.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:25:12 +05:30
Saket Aryan bc5526f13d feat(integrations): declare which surface each client is, and its version
Nothing on the wire said which Mem0 surface made a call. Both SDKs sent only an
auth header, so the platform saw python-httpx and axios and attributed every
plugin, wrapper and direct API user to one undifferentiated bucket. Version was
unknowable, which is what gates every deprecation decision.

Three headers, and the rules on them are the point:

- X-Mem0-Source and X-Application are SET-ONCE. Whichever layer is outermost
  sets them; nothing below overwrites. A plugin wrapping the SDK keeps its own
  identity instead of being renamed by the transport underneath it.
- X-Mem0-Client is APPEND-ONLY. A plugin calling the Python SDK produces
  `mem0-plugin/0.3.1, mem0-python/2.0.19`, so neither layer can erase the other.

Deliberately not User-Agent: proxies rewrite it, and we have already met a WAF
that 403s on it.

The plugin core also hoists `source` out of metadata to the top level, which is
where the backend actually reads it. It sat in metadata, which get_event_source
never consults, so all six plugins arrived indistinguishable from a raw SDK call
no matter what they set. The harness tag stays in metadata as hook provenance.

pi-agent had PI_AGENT as a PostHog property only and never sent it on the wire.
vercel-ai-sdk sent nothing at all from its raw fetch calls.

Values must exist in the platform's EventSource enum or they bucket to OTHERS,
so integrations/AGENTS.md now states the contract and the "adding an
integration" checklist requires landing the platform value in the same week.

Pairs with mem0ai/platform#3602, which recognizes these values.

TypeScript changes are not typechecked locally — deps are not installed for
those packages. CI covers them.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:19:00 +05:30
Saket Aryan 0d2b20c03d fix(plugins): count installs once, and re-resolve the email when the key changes
code.install counted upgrades and repeat sessions. Session start records install
whenever is_first_run() is true, and that only checked whether
telemetry-identity.json exists. Recording install does not create that file —
only the first successful flush does. So install fired for every 0.2.x user on
their first 0.3.x session (0.2.x never wrote the file, and the data directory
survives the upgrade), again for any session starting before that first flush
finished, and — this is the part that makes it unbounded rather than a race —
on every single session, forever, for anyone whose flush never succeeds. An
offline or firewalled user reported a new install every time they opened an
editor, which is exactly the population hardest to see in the data.

A dedicated install-state.json is now claimed with O_CREAT|O_EXCL at the moment
install is recorded, so two sessions starting together cannot both win, and the
marker is not coupled to identity. Deliberately not the identity file: writing
that from a recording process would race the sender, which writes it during
resolve_distinct_id, and overloading it is what caused this.

Upgrade detection keys on the data directory already having content. A fresh
install has an empty one; anything else predates this session. That is a firmer
predicate than looking for 0.2.x's venv/ and requirements.txt, which is a guess
about files another part of the plugin may or may not have written and only ever
works for this one upgrade. The version is stored in the marker so later changes
record code.upgrade with a real from_version.

A cached email outlived an API key change. resolve_distinct_id kept the first
email it resolved and never looked again, so switching to a key from another
account kept attributing events to the previous one. It now stores a fingerprint
of the key the email came from and re-resolves when the current key differs, and
falls back to the anonymous id when no key is configured rather than continuing
to attribute to an account it cannot verify.

The dangerous part is the alias. resolve_distinct_id's second return value
becomes a PostHog $identify with $anon_distinct_id, and aliasing one account
email to another merges two real person profiles irreversibly. The re-resolve
path returns no alias; aliasing runs anonymous to email only, and never
email to email.

One existing test asserted that is_first_run flips when the identity file is
written, which is the defect itself. Rewritten, along with coverage for atomic
claiming, upgrade detection and version changes.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:18:38 +05:30
Saket Aryan cfdfe40e09 fix(plugins): stop delivering telemetry events twice, and stop losing parked ones
Two defects, one cause: the spool protocol infers ownership instead of holding
it, and never records progress.

Duplicate delivery after a partial failure. flush() posts the claim in batches of
100 and returns on the first failure, keeping the whole file. The retry then
posts every batch again, including the ones that already arrived — 150 recorded
events were delivered 250 times. Progress is now written back to the claim after
each successful batch, so a retry resumes where the send stopped and a crash
repeats at most one batch.

Duplicate delivery when two senders overlap. spool.replace(claim) is os.rename,
which preserves mtime, so a claim created after a quiet minute inherited the
spool's last-write time and looked abandoned the instant it existed. A second
sender starting while the first was still posting took it over and sent it too —
most likely at session end, when the MCP server's exit sender and the SessionEnd
flush worker both drain. Claims are now touched at claim time, and the per-batch
rewrite doubles as a lease heartbeat. _post makes one attempt with SEND_TIMEOUT
and no retry, so a heartbeat lands well inside the 120s lease; a test asserts
that margin so adding a retry loop to _post cannot silently break it.

Parked batches starved. _claim_spool only looked at parked .sending files when
no spool existed, and because sessions keep recording there usually was one — so
a batch parked by a failed send waited until the 7-day expiry deleted it unsent,
despite its own presence being what starts the sender in the first place.
flush() now drains the live spool and then parked claims in the same run, oldest
first, bounded. Expiry applies only after a genuine retry has failed, with the
attempt count carried in the filename.

A sender that gives up releases its lease rather than heartbeating on the way
out, so the next run picks the batch up promptly instead of waiting a full stale
window for a batch nobody is working on. A failing send stops the run, so one
broken connection cannot burn every parked batch's attempt budget at once.

Two existing tests asserted the old lifecycle and are updated in place, each
with a comment saying what changed.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:18:21 +05:30
Saket Aryan f9c566aa16 fix(plugins): report the plugin that produced the event, not the one that sent it
harness is set when an event is recorded; source was set when its batch was
sent. Both came from module globals that stay at "generic" and "MEM0_PLUGIN"
until telemetry.init() runs, and two processes in the pipeline never run it:

- `python3 telemetry.py`, the detached sender spawn_flush() starts at session
  start, after every skill command, and when the MCP server exits. Everything it
  delivered was labelled source=MEM0_PLUGIN. Only batches flush_worker.py
  happened to drain got the real host.
- mcp_server.py, which records every manual search as harness=generic.

All six Python plugins ship the same files, so source could not tell any of them
apart and MCP searches from every plugin landed in one generic bucket. The
portable plugin is worse: it has no flush_worker at all, so its only sender is
the uninitialised one and 100% of its events were mislabelled.

Two changes. record() stamps source beside harness, so the sending process stops
mattering — flush() already spreads per-event properties last, so a per-event
source wins over any sender default. And the build generates core/_harness_id.py
per host, seeding both modules at import, so identity no longer depends on an
entrypoint remembering to call init(). The build already computed HARNESS_ID and
spent it only on skill templating, and bundle_drift already diffs core/
byte-for-byte, so --check catches drift for free.

Deliberately not adding MEM0_PLUGIN_HARNESS to the six manifests: they sit
outside the --sync and --check boundary, which is the property that caused this.

Also unifies two defaults that disagreed. configure_harness derived
`<host>_plugin` while telemetry.init derived `MEM0_<HOST>_PLUGIN`, so a third
value existed. It was unreachable only because hook_runner never calls flush();
moving source into record() would have made it live.

Events now carry a uuid so a resend can be collapsed.

The suite stayed green through all of this because the only tests live under one
host, behind a conftest that calls init() at import. New tests run in real
subprocesses with no init, and cover the portable plugin, which would pass a
native-only test vacuously.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:17:53 +05:30
Saket Aryan 0d37619f24 fix(plugins): say what telemetry actually sends, and salt the hashes
The plugin README promises "anonymous usage events" and the telemetry module's
docstring says it sends only "salted hashes". Neither is true.

resolve_distinct_id() exchanges the API key for the account email and sends that
as the distinct_id on every event. Installing the plugin requires an API key, so
this is nearly every user. That is probably the behaviour we want — the Python
SDK and the CLI attribute the same way — but the description has to match it.

repo_hash and session_hash were unsalted SHA-256 cut to 16 hex characters.
repo.identity is a git remote URL, or `local:<absolute path>` when there is no
remote, which normally contains the account username. Sixteen unsalted hex
characters over that input space is enumerable, so the hash was not a privacy
control at all.

Salted per install, with the salt kept in the identity file. That preserves
every within-account join the analytics actually use and gives up only
cross-machine joins on the same repository, which nothing computes. Since the
distinct_id is already the email, the hash was never buying privacy from us —
only from whoever obtains the data later, which is exactly what the salt fixes.

Also corrects deepseek-plugin's README and source comment, which told readers
ZAPIER and STRANDS were already in the backend's KNOWN_EVENT_SOURCES allowlist.
Neither was.

Adds a Telemetry section to docs/integrations/claude-code.mdx, which had none.

Claude-Session: https://claude.ai/code/session_01C7tEmH86HAr7GoAAKCEHZb
2026-09-15 00:17:26 +05:30
170 changed files with 2700 additions and 8330 deletions
+1 -1
View File
@@ -12,7 +12,7 @@
"name": "mem0",
"source": "./integrations/claude-code-plugin",
"description": "Cross-session memory and token savings for coding agents.",
"version": "0.3.3"
"version": "0.3.1"
}
]
}
+1 -1
View File
@@ -12,7 +12,7 @@
"name": "mem0",
"source": "./integrations/cursor-plugin",
"description": "Cross-session memory and token savings for coding agents.",
"version": "0.3.3"
"version": "0.3.1"
}
]
}
@@ -32,9 +32,6 @@ jobs:
- name: Type check
run: bun run type-check
- name: Test
run: bun test
- name: Build
run: bun run build
+1 -1
View File
@@ -5,7 +5,7 @@
{
"id": "mem0",
"displayName": "Mem0",
"version": "0.3.3",
"version": "0.3.1",
"description": "Cross-session memory and token savings for coding agents.",
"homepage": "https://mem0.ai",
"keywords": ["memory", "personalization", "mcp", "semantic-search"],
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@mem0/cli",
"version": "0.2.14",
"version": "0.2.13",
"description": "The official CLI for mem0 — the memory layer for AI agents",
"type": "module",
"bin": {
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "mem0-cli"
version = "0.2.13"
version = "0.2.12"
description = "The official CLI for mem0 — the memory layer for AI agents"
readme = "README.md"
license = "Apache-2.0"
+1 -1
View File
@@ -1,3 +1,3 @@
"""mem0 CLI — the command-line interface for the mem0 memory layer."""
__version__ = "0.2.13"
__version__ = "0.2.12"
@@ -1,5 +0,0 @@
---
title: 'Generate Profiles'
description: "Start one generation: sample a few entities, or build one for a single entity."
openapi: post /v2/profiles/jobs/
---
@@ -1,5 +0,0 @@
---
title: 'Get Generation Job'
description: "Read the progress of a generation, and whether it finished."
openapi: get /v2/profiles/jobs/{job_id}/
---
@@ -1,5 +0,0 @@
---
title: 'Get Profile Settings'
description: "Retrieve the profile schema, custom instructions, and enabled flag for the current project."
openapi: get /v2/profiles/settings/
---
@@ -1,5 +0,0 @@
---
title: 'Get Profile'
description: "Retrieve the structured profile for a user, with a status describing whether generation has completed."
openapi: get /v2/entities/{entity_type}/{entity_id}/profile/
---
@@ -1,5 +0,0 @@
---
title: 'Update Profile Settings'
description: "Set the JSON Schema, custom instructions, or enabled flag that control profile generation for the project."
openapi: post /v2/profiles/settings/
---
+37 -221
View File
@@ -7,36 +7,6 @@ mode: "wide"
<Tabs>
<Tab title="Python">
<Update label="2026-09-25" description="v2.2.1">
**Bug Fixes:**
- **Core:** `add()` (`Memory` and `AsyncMemory`) no longer reports records the vector store rejected as successful `ADD` events. Only records that were actually inserted are written to history, entity-linked, and returned. If none of the extracted memories could be inserted, `add()` now raises `VectorStoreError` instead of returning memories that were never stored ([#7066](https://github.com/mem0ai/mem0/pull/7066))
- **Core:** Restore the context-manager protocol on `Memory` (`with Memory() as m:`) and `AsyncMemory` (`async with AsyncMemory() as m:`), which closes the instance on exit ([#7354](https://github.com/mem0ai/mem0/pull/7354))
- **LLMs:** The AWS Bedrock Anthropic path now returns the first Converse content block that carries text instead of always reading `content[0]`. Claude reasoning models can emit a `reasoningContent` block before the text block, which previously made the call fail with a `KeyError` ([#6369](https://github.com/mem0ai/mem0/pull/6369))
- **Vector Stores:** Turbopuffer filters now apply every operator (`eq`, `ne`, `gt`, `gte`, `lt`, `lte`, `in`, `nin`). Only `gte` and `lte` were read before, so any other operator was silently dropped and the query returned unfiltered results. An unsupported operator now raises `ValueError` ([#6564](https://github.com/mem0ai/mem0/pull/6564))
- **Vector Stores:** Turbopuffer `search()` scores now respect `distance_metric`. With `euclidean_squared`, the unbounded squared distance maps to `1 / (1 + distance)` instead of `1 - distance`, which went negative and inverted ranking for any distance above 1 ([#6559](https://github.com/mem0ai/mem0/pull/6559))
- **Vector Stores:** S3 Vectors `search()` scores are now metric-aware. With `euclidean`, the distance maps to `1 / (1 + distance)` instead of `max(0, 1 - distance)`, which collapsed most scores to 0 ([#6547](https://github.com/mem0ai/mem0/pull/6547))
</Update>
<Update label="2026-09-23" description="v2.2.0">
**New Features:**
- **Client:** Add User Profiles to `MemoryClient` and `AsyncMemoryClient`: `get_profile()`, `generate_profile()`, `get_profile_settings()`, `update_profile_settings()`, `sample_profiles()`, and `get_profile_job()`. A profile is a structured, always-current JSON summary of one user, shaped by a JSON Schema you configure per project and filled by an LLM from that user's memories. Generation is asynchronous. Every job POST carries an `Idempotency-Key`; to retry a lost request without starting a second job, pass the same `idempotency_key` on each attempt ([#7340](https://github.com/mem0ai/mem0/pull/7340))
**Bug Fixes:**
- **Vector Stores:** Guard against `None` timestamps in the Valkey vector store's `insert()` and `update()` paths. `created_at` and `updated_at` fields that were present in the payload but set to `None` previously passed the `"created_at" not in payload` / `"updated_at" in payload` checks and raised `TypeError` when `datetime.fromisoformat()` received `None`. Both paths now use `.get()` with a truthiness check so `None` values fall through to the default, matching the Redis provider's behavior ([#6993](https://github.com/mem0ai/mem0/pull/6993))
</Update>
<Update label="2026-09-18" description="v2.1.0">
**Improvements:**
- **Client:** Requests now carry three surface-identity headers so the platform can tell which product made a call. `X-Mem0-Source` names the surface and `X-Application` the host app it runs inside, both set-once so a wrapper that already declared its identity keeps it. `X-Mem0-Client` is append-only and carries `name/version` per layer, outermost first, so a plugin calling this SDK reports the whole chain rather than only the last speaker. `MEM0_SOURCE`, `MEM0_APPLICATION` and `MEM0_CLIENT_STACK` set them from the environment for wrappers that cannot pass options ([#7326](https://github.com/mem0ai/mem0/pull/7326))
- **Client:** The client stack is bounded by dropping whole entries rather than slicing characters, and this SDK's own entry is the reserved one. Truncating the joined string could sever an identifier mid-name and the platform parsed the fragment as a real client ([#7326](https://github.com/mem0ai/mem0/pull/7326))
</Update>
<Update label="2026-09-02" description="v2.0.20">
**Improvements:**
@@ -204,7 +174,7 @@ mode: "wide"
<Update label="2026-06-24" description="v2.0.8">
**New Features:**
- **Embeddings:** Add native `embed_batch` to five embedders for batched embedding requests: LM Studio, Together, HuggingFace, Vertex AI, and Google GenAI ([#5609](https://github.com/mem0ai/mem0/pull/5609))
- **Embeddings:** Add native `embed_batch` to five embedders: LM Studio, Together, HuggingFace, Vertex AI, and Google GenAI: for batched embedding requests ([#5609](https://github.com/mem0ai/mem0/pull/5609))
**Bug Fixes:**
- **Core:** Guard against malformed `image_url` entries in `parse_vision_messages` to prevent crashes ([#5631](https://github.com/mem0ai/mem0/pull/5631))
@@ -994,7 +964,7 @@ See the [OSS v2 to v3 migration guide](https://docs.mem0.ai/migration/oss-v2-to-
**New Features:**
- **OpenMemory:** Added OpenMemory support
- **Neo4j:** Added weights to Neo4j model
- **AWS:** Added support for OpenSearch Serverless
- **AWS:** Added support for Opsearch Serverless
- **Examples:** Added ElizaOS Example
**Improvements:**
@@ -1257,31 +1227,6 @@ See the [OSS v2 to v3 migration guide](https://docs.mem0.ai/migration/oss-v2-to-
<Tab title="TypeScript">
<Update label="2026-09-25" description="v3.3.1">
**Bug Fixes:**
- **Config (OSS):** `ConfigManager` no longer injects OpenAI's default `baseURL` and `model` into non-OpenAI LLM providers. The defaults now apply only to `openai` and `openai_structured`, so a DeepSeek, xAI, or other provider config without an explicit `baseURL` or `model` falls back to that provider's own defaults and env vars (`DEEPSEEK_API_BASE`, `XAI_API_BASE`, ...) instead of pointing at OpenAI with an OpenAI model name ([#7350](https://github.com/mem0ai/mem0/pull/7350))
- **Vector Stores:** Turbopuffer filters now apply every operator (`eq`, `ne`, `gt`, `gte`, `lt`, `lte`, `in`, `nin`). Only the range operators were read before, so `eq`, `ne`, `in`, and `nin` were silently dropped. A `"*"` value now matches anything instead of nothing, an array value is treated as `in`, and an unsupported operator throws ([#6578](https://github.com/mem0ai/mem0/pull/6578))
- **Vector Stores:** Turbopuffer `search()` with `euclidean_squared` now maps the unbounded squared distance to `1 / (1 + distance)` instead of `1 - distance`, which went negative and inverted ranking for any distance above 1 ([#6580](https://github.com/mem0ai/mem0/pull/6580))
- **Packaging:** `pg`, `@types/pg`, and `natural` are now optional peer dependencies, and `pg` / `@types/pg` accept caret ranges instead of the exact `8.11.3` / `8.11.0` pins, so installing `mem0ai` no longer pulls in `pg` or conflicts with an app's own `pg` version. The PGVector store now imports `pg` only when it is used, so install it alongside `mem0ai` (`npm install pg`) if you use that store. `@types/jest` moved from peer dependencies to dev dependencies ([#7450](https://github.com/mem0ai/mem0/pull/7450))
</Update>
<Update label="2026-09-23" description="v3.3.0">
**New Features:**
- **Client:** Add User Profiles to `MemoryClient`: `getProfile()`, `generateProfile()`, `getProfileSettings()`, `updateProfileSettings()`, `sampleProfiles()`, and `getProfileJob()`. A profile is a structured, always-current JSON summary of one user, shaped by a JSON Schema you configure per project and filled by an LLM from that user's memories. Generation is asynchronous. Every job POST carries an `Idempotency-Key`; to retry a lost request without starting a second job, pass the same `idempotencyKey` on each attempt ([#7340](https://github.com/mem0ai/mem0/pull/7340))
</Update>
<Update label="2026-09-18" description="v3.2.0">
**Improvements:**
- **Client:** Requests now carry `X-Mem0-Source`, `X-Application` and `X-Mem0-Client`, matching the Python SDK. The first two are set-once so an outer wrapper keeps its identity; the third is append-only and reports the whole layer chain. Read from `MEM0_SOURCE`, `MEM0_APPLICATION` and `MEM0_CLIENT_STACK` when set ([#7326](https://github.com/mem0ai/mem0/pull/7326))
- **Client:** The SDK version in `X-Mem0-Client` is injected at build time rather than hardcoded, so it cannot go stale at the next release ([#7326](https://github.com/mem0ai/mem0/pull/7326))
</Update>
<Update label="2026-09-02" description="v3.1.8">
**Improvements:**
@@ -1919,13 +1864,6 @@ See the [OSS v2 to v3 migration guide](https://docs.mem0.ai/migration/oss-v2-to-
<Tab title="CLI">
<Update label="2026-09-18" description="Python v0.2.13 / Node v0.2.14">
**Improvements:**
- **Client:** Requests now carry the three surface-identity headers (`X-Mem0-Source`, `X-Application`, `X-Mem0-Client`) introduced in the Python and TypeScript SDKs, so the platform can attribute calls made through the CLI to the correct surface and version ([#7326](https://github.com/mem0ai/mem0/pull/7326))
</Update>
<Update label="2026-08-24" description="Python v0.2.12 / Node v0.2.13">
**New Features:**
@@ -2138,7 +2076,7 @@ A full-featured command-line interface for Mem0, available in both Python and No
- New Git repository writes use a hash of the remote identity in `agent_id`. Search and explicit shared-memory deletion include both current and legacy repository IDs within the repository's `app_id`. Existing memories are not rewritten. Legacy IDs retain their original ambiguity for matching owner/repository names on different Git hosts.
**Packaging:**
- Claude Code, Cursor, Codex, Kimi, Antigravity, and the portable Python bundle are versioned at `0.3.1`. OpenCode and DeepSeek Harness are `0.3.0`; Pi Agent is `0.3.0`; OpenClaw is `1.1.0`. Each host's changes and upgrade considerations are listed in its tab.
- Claude Code, Cursor, Codex, Kimi, Antigravity, and the portable Python bundle are versioned at `0.3.1`. OpenCode, Pi Agent, and DeepSeek Harness are `0.3.0`; OpenClaw is `1.1.0`. Each host's changes and upgrade considerations are listed in its tab.
- Python and TypeScript CI run their respective runtime suites. Package checks build the installable artifacts, check generated-file consistency, and reject TypeScript output that still imports monorepo source.
[#7203](https://github.com/mem0ai/mem0/pull/7203)
@@ -2433,25 +2371,9 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
<Tab title="Claude Code">
<Update label="2026-09-23" description="Claude Code plugin v0.3.3">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Search:** The `search_memories` tool description no longer tells the agent to call it before answering anything that could depend on prior context. It now asks for a search before repeating investigation or when earlier decisions, fixes, commands, or results may help, which reduces unnecessary searches ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Sidekick:** Sidekick searches memories when earlier sessions could help, instead of before every answer ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Extraction:** Repository memory instructions are shorter. They no longer ask for a dedicated memory for each command that failed and was then fixed, and no longer carry separate rules against saving personal preferences or memories that only name the repository, branch, or directory ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Search skill:** `/search` no longer describes categories as best-effort labels or asks for a retry without the category ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Packaging:** `PLUGIN_VERSION` bumped to `0.3.3`, so the `mem0-plugin/<version>` wire header and `plugin_version` telemetry field identify builds with these prompts ([#7420](https://github.com/mem0ai/mem0/pull/7420))
</Update>
<Update label="2026-09-18" description="Claude Code plugin v0.3.2">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Sidekick is now available only in Claude Code, with Sonnet, worktree isolation, and parent memories.
Sidekick is now available only in Claude Code, with Sonnet, worktree isolation, and parent memories.
</Update>
@@ -2474,24 +2396,9 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
<Tab title="Cursor">
<Update label="2026-09-23" description="Cursor plugin v0.3.3">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Search:** The `search_memories` tool description no longer tells the agent to call it before answering anything that could depend on prior context. It now asks for a search before repeating investigation or when earlier decisions, fixes, commands, or results may help, which reduces unnecessary searches ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Extraction:** Repository memory instructions are shorter. They no longer ask for a dedicated memory for each command that failed and was then fixed, and no longer carry separate rules against saving personal preferences or memories that only name the repository, branch, or directory ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Search skill:** `/search` no longer describes categories as best-effort labels or asks for a retry without the category ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Packaging:** `PLUGIN_VERSION` bumped to `0.3.3`, so the `mem0-plugin/<version>` wire header and `plugin_version` telemetry field identify builds with these prompts ([#7420](https://github.com/mem0ai/mem0/pull/7420))
</Update>
<Update label="2026-09-18" description="Cursor plugin v0.3.2">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Removes Sidekick and its start/stop hooks. Memory capture, search, and six skills remain available.
Removes Sidekick and its start/stop hooks. Memory capture, search, and six skills remain available.
</Update>
@@ -2514,24 +2421,9 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
<Tab title="Codex">
<Update label="2026-09-23" description="Codex plugin v0.3.3">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Search:** The `search_memories` tool description no longer tells the agent to call it before answering anything that could depend on prior context. It now asks for a search before repeating investigation or when earlier decisions, fixes, commands, or results may help, which reduces unnecessary searches ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Extraction:** Repository memory instructions are shorter. They no longer ask for a dedicated memory for each command that failed and was then fixed, and no longer carry separate rules against saving personal preferences or memories that only name the repository, branch, or directory ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Search skill:** `/search` no longer describes categories as best-effort labels or asks for a retry without the category ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Packaging:** `PLUGIN_VERSION` bumped to `0.3.3`, so the `mem0-plugin/<version>` wire header and `plugin_version` telemetry field identify builds with these prompts ([#7420](https://github.com/mem0ai/mem0/pull/7420))
</Update>
<Update label="2026-09-18" description="Codex plugin v0.3.2">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Renames shared tracking to use subagent terminology. Native subagent memory support remains available.
Renames shared tracking to use subagent terminology. Native subagent memory support remains available.
</Update>
@@ -2551,25 +2443,32 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
</Tab>
<Tab title="Agent Plugins v1">
<Update label="Unreleased" description="Sidekick availability">
Sidekick is available only in Claude Code, not in the portable package.
</Update>
<Update label="2026-09-08" description="Portable Mem0 plugin v0.3.1">
**Added:**
- One portable package at `integrations/mem0-agent-plugin/`, using the Agent Plugins 1.0.0 root `plugin.json`, `mcp.json`, and fixed `skills/` locations.
- Ships a local, read-only `search_memories` server and the six shared memory skills. Uses `PLUGIN_ROOT` for bundled files and `PLUGIN_DATA` for persistent plugin state; all package files remain inside the installable directory.
**Packaging:**
- Generated from the shared Python runtime and skill templates. Builds validate the manifest, MCP configuration, skills, and generated-file consistency.
- Host lifecycle hooks and native Sidekick declarations remain in the native plugin packages; the portable package does not provide automatic lifecycle capture or host-specific subagent isolation. Its bundled remember skill cannot persist a new memory on its own because the portable package has no capture hooks or write tool.
[#7203](https://github.com/mem0ai/mem0/pull/7203)
</Update>
</Tab>
<Tab title="OpenCode">
<Update label="2026-09-23" description="OpenCode plugin v0.4.1">
**Improvements:**
- **Search:** The `search_memories` tool description no longer asks the agent to search proactively or to run several searches for multi-part questions. It now asks for a search before repeating investigation or when earlier decisions, fixes, commands, or results may help, the same wording as the other coding-agent plugins ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Session context:** Removed two system-context lines that told the agent to run 2 parallel searches before responding and 2-4 parallel searches for non-trivial tasks ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Skills:** `/mem0-search` and `/mem0-context-loader` make one `search_memories` call instead of 2 and 2-4 parallel calls. `/mem0-context-loader` now uses the search skill description from the other coding-agent plugins instead of asking to load at every new task or context switch ([#7420](https://github.com/mem0ai/mem0/pull/7420))
</Update>
<Update label="2026-09-18" description="OpenCode plugin v0.4.0">
**Changes:**
- **Telemetry:** The PostHog `source` tag changed from the literal `"plugin"` to `OPENCODE_PLUGIN`, and `project_hash` is now salted. Saved PostHog insights filtering on `source = "plugin"` will stop matching new events; historical data is unaffected ([#7322](https://github.com/mem0ai/mem0/pull/7322))
- **Config:** A new `keyFingerprint` key appears in the install-count deduplication logic; installs are now counted once per key rather than on every activation ([#7325](https://github.com/mem0ai/mem0/pull/7325))
</Update>
<Update label="2026-09-08" description="OpenCode plugin v0.3.0">
**Changed:**
@@ -2668,24 +2567,9 @@ Initial release of the Mem0 plugin for Claude Code and Cursor, followed by Codex
<Tab title="Antigravity">
<Update label="2026-09-23" description="Antigravity plugin v0.3.3">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Search:** The `search_memories` tool description no longer tells the agent to call it before answering anything that could depend on prior context. It now asks for a search before repeating investigation or when earlier decisions, fixes, commands, or results may help, which reduces unnecessary searches ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Extraction:** Repository memory instructions are shorter. They no longer ask for a dedicated memory for each command that failed and was then fixed, and no longer carry separate rules against saving personal preferences or memories that only name the repository, branch, or directory ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Search skill:** `/search` no longer describes categories as best-effort labels or asks for a retry without the category ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Packaging:** `PLUGIN_VERSION` bumped to `0.3.3`, so the `mem0-plugin/<version>` wire header and `plugin_version` telemetry field identify builds with these prompts ([#7420](https://github.com/mem0ai/mem0/pull/7420))
</Update>
<Update label="2026-09-18" description="Antigravity plugin v0.3.2">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Removes Sidekick. Memory capture, search, and six skills remain available.
Removes Sidekick. Memory capture, search, and six skills remain available.
</Update>
@@ -2781,24 +2665,9 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="Kimi">
<Update label="2026-09-23" description="Kimi Code plugin v0.3.3">
<Update label="Unreleased" description="Sidekick availability">
**Improvements:**
- **Search:** The `search_memories` tool description no longer tells the agent to call it before answering anything that could depend on prior context. It now asks for a search before repeating investigation or when earlier decisions, fixes, commands, or results may help, which reduces unnecessary searches ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Extraction:** Repository memory instructions are shorter. They no longer ask for a dedicated memory for each command that failed and was then fixed, and no longer carry separate rules against saving personal preferences or memories that only name the repository, branch, or directory ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Search skill:** `/search` no longer describes categories as best-effort labels or asks for a retry without the category ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Packaging:** `PLUGIN_VERSION` bumped to `0.3.3`, so the `mem0-plugin/<version>` wire header and `plugin_version` telemetry field identify builds with these prompts ([#7420](https://github.com/mem0ai/mem0/pull/7420))
</Update>
<Update label="2026-09-18" description="Kimi Code plugin v0.3.2">
**Improvements:**
- **Telemetry:** `PLUGIN_VERSION` bumped to `0.3.2`. The `mem0-plugin/<version>` wire header and `plugin_version` telemetry field now reflect the fixes from #7322 through #7358 ([#7373](https://github.com/mem0ai/mem0/pull/7373))
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
**Changes:**
- **Sidekick:** Removes Sidekick and its start/stop hooks. Memory capture, recall, and six skills remain available.
Removes Sidekick and its start/stop hooks. Memory capture, recall, and six skills remain available.
</Update>
@@ -2836,21 +2705,6 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="OpenClaw">
<Update label="2026-09-23" description="openclaw-mem0 v1.2.1">
**Improvements:**
- **Search:** The `memory_search` tool description no longer asks the agent to search proactively or to run several searches for multi-part questions. It now asks for a search before repeating investigation or when earlier decisions, fixes, commands, or results may help. Recall strategies (`smart`, `always`, `manual`) are unchanged ([#7420](https://github.com/mem0ai/mem0/pull/7420))
</Update>
<Update label="2026-09-18" description="openclaw-mem0 v1.2.0">
**Changes:**
- **Config:** Added `keyFingerprint` to the config schema for install-count deduplication; installs are now counted once per key rather than on every activation ([#7325](https://github.com/mem0ai/mem0/pull/7325))
- **Telemetry:** Events are no longer delivered twice, and the `plugin_version` field now reflects the plugin that produced the event ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
</Update>
<Update label="2026-09-08" description="openclaw-mem0 v1.1.0">
**Changed:**
@@ -3139,22 +2993,6 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="Pi Agent">
<Update label="2026-09-23" description="Pi Agent plugin v0.3.2">
**Improvements:**
- **Search:** The memory policy, `mem0_memory` tool description, and prompt guidelines no longer ask the agent to search before answering anything that may depend on earlier context or to run several searches per question. They now ask for a search before repeating investigation or when earlier decisions, fixes, commands, or results may help ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Skills:** `context-loader` makes one search instead of 2-4 parallel searches, and uses the search skill description from the other coding-agent plugins ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Automatic recall:** Injected memories are introduced as "Mem0 found these relevant memories from earlier work in this repository:", the same heading as the Python plugins. The old heading called them a shallow first pass and told the agent to search again ([#7420](https://github.com/mem0ai/mem0/pull/7420))
</Update>
<Update label="2026-09-18" description="Pi Agent plugin v0.3.1">
**Improvements:**
- **Telemetry:** Events are no longer delivered twice, no longer lose parked events, and now attribute each event to the plugin that produced it. The `plugin_version` wire field reflects the fixed release ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
</Update>
<Update label="2026-09-08" description="Pi Agent plugin v0.3.0">
**Changed:**
@@ -3268,21 +3106,6 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="DeepSeek Harness">
<Update label="2026-09-23" description="deepseek-plugin v0.3.2">
**Improvements:**
- **Search:** The `search_memory` tool description no longer asks the agent to search proactively before answering anything that may depend on earlier context. It now asks for a search before repeating investigation or when earlier decisions, fixes, commands, or results may help ([#7420](https://github.com/mem0ai/mem0/pull/7420))
- **Automatic recall:** Injected memories are introduced as "Mem0 found these relevant memories from earlier work:". The old heading called them a shallow first pass and told the agent to search again with `mem0_memory`, a tool DeepSeek does not have ([#7420](https://github.com/mem0ai/mem0/pull/7420))
</Update>
<Update label="2026-09-18" description="deepseek-plugin v0.3.1">
**Improvements:**
- **Telemetry:** Rebuild with the fixed shared telemetry core from `agent-plugin-core`. Events are no longer delivered twice, no longer lose parked events on flush, and now attribute each event to the plugin that produced it ([#7323](https://github.com/mem0ai/mem0/pull/7323), [#7324](https://github.com/mem0ai/mem0/pull/7324), [#7358](https://github.com/mem0ai/mem0/pull/7358))
</Update>
<Update label="2026-09-08" description="deepseek-plugin v0.3.0">
**Added:**
@@ -3327,13 +3150,6 @@ Existing memories written by the previous versions are not rewritten. If your me
<Tab title="Vercel AI SDK">
<Update label="2026-09-18" description="Vercel AI SDK v3.0.3">
**Improvements:**
- **Client:** Inherits the three surface-identity headers (`X-Mem0-Source`, `X-Application`, `X-Mem0-Client`) from the TypeScript SDK bump, so platform calls made through the Vercel AI SDK provider are now correctly attributed ([#7326](https://github.com/mem0ai/mem0/pull/7326))
</Update>
<Update label="2026-08-24" description="Vercel AI SDK v3.0.2">
**Security:**
@@ -11,7 +11,7 @@ To use Together embedding models, set the `TOGETHER_API_KEY` environment variabl
<Note> The `embedding_model_dims` parameter for `vector_store` should be set to `1024` for Together embedder. </Note>
<Warning>
**Breaking default change.** The default Together embedding model is now `intfloat/multilingual-e5-large-instruct` (**1024-dim**), replacing the previous default `togethercomputer/m2-bert-80M-8k-retrieval` (**768-dim**). If you created a self-hosted vector store with the old default, its collection is 768-dim and will reject the new 1024-dim vectors — **recreate/reindex the collection at 1024 dimensions** after upgrading. To defer the change, pin the previous values explicitly (`model="togethercomputer/m2-bert-80M-8k-retrieval"`, `embedding_dims=768`) — note Together no longer lists this model among its recommended embeddings, so reindexing at 1024 is the durable path.
**Breaking default change.** The default Together embedding model is now `intfloat/multilingual-e5-large-instruct` (**1024-dim**), replacing the previous default `togethercomputer/m2-bert-80M-8k-retrieval` (**768-dim**). If you created a self-hosted vector store with the old default, its collection is 768-dim and will reject the new 1024-dim vectors **recreate/reindex the collection at 1024 dimensions** after upgrading. To defer the change, pin the previous values explicitly (`model="togethercomputer/m2-bert-80M-8k-retrieval"`, `embedding_dims=768`) note Together no longer lists this model among its recommended embeddings, so reindexing at 1024 is the durable path.
</Warning>
<CodeGroup>
+1 -1
View File
@@ -95,7 +95,7 @@ Uses the identity from Azure PowerShell (`Connect-AzAccount`).
7. **Azure Developer CLI Credential:**
Uses the session from Azure Developer CLI (`azd auth login`).
<Note> If an API key is provided, it will be used for authentication over an Azure Identity </Note>
<Note> If an API is provided, it will be used for authentication over an Azure Identity </Note>
To enable Role-Based Access Control (RBAC) for Azure AI Search, follow these steps:
1. In the Azure Portal, navigate to your **Azure AI Search** service.
@@ -5,12 +5,6 @@ description: "Use pgvector as a vector store in Mem0 for PostgreSQL-based vector
[pgvector](https://github.com/pgvector/pgvector) is an open-source vector similarity search extension for Postgres. After connecting to Postgres, run `CREATE EXTENSION IF NOT EXISTS vector;` to create the vector extension.
The TypeScript SDK loads the `pg` driver only when you use this store, so install it alongside `mem0ai`:
```bash
npm install pg
```
### Usage
<CodeGroup>
@@ -94,7 +94,6 @@ Here are the parameters available for configuring Pinecone:
| `hybrid_search` | Whether to enable hybrid search | `False` |
| `metric` | Distance metric for vector similarity | `"cosine"` |
| `batch_size` | Batch size for operations | `100` |
| `extra_params` | Additional keyword arguments passed to the `Pinecone` client constructor. Ignored when `client` is supplied. | `None` |
| `namespace` | Namespace for the collection, useful for multi-tenancy. | `None` |
</Tab>
<Tab title="TypeScript">
@@ -30,7 +30,7 @@ pip install google-adk mem0ai python-dotenv
## Code Breakdown
Let's get started and understand the different components required in building a healthcare assistant powered by memory.
Let's get started and understand the different components required in building a healthcare assistant powered by memory
```python
# Import dependencies
-388
View File
@@ -1,388 +0,0 @@
---
title: Build a Company Brain with Mem0 Platform and Supabase
description: "Build a shared company brain using Mem0 Platform as the managed memory layer, Supabase as your system of record, and the Mem0 MCP server."
---
<Info icon="server">
**Uses:** Mem0 **Platform** (`MemoryClient`) · **System of record:** Supabase (Postgres + Auth) · **Access layer:** the hosted Mem0 MCP server. **You'll build:** a company brain your whole org (and every agent) writes to and queries, ending with a new-hire onboarding demo.
</Info>
Companies lose knowledge constantly: why you picked Postgres over Mongo, who owns billing, the deploy rule only one engineer remembers. A **company brain** captures this and answers questions about it, for every employee and every agent, and keeps it after people leave.
We'll build one on **Mem0 Platform** (the managed memory layer, so there's no vector DB to run) with **Supabase as the system of record** (where your employees, teams, and source documents actually live) and the **Mem0 MCP server** as the wire that lets Claude Code, Cursor, or a Slack bot all reach the same brain.
<Note>
**How Platform and Supabase divide the work.** Mem0 Platform manages storage and extraction server-side, you do **not** point it at your own database. Supabase is your app's source of truth and identity provider; we *ingest* knowledge from Supabase into the brain and use Supabase Auth to decide who's asking. (If you want to self-host the vector store instead, that's the OSS path, see the [Supabase vector store reference](/components/vectordbs/dbs/supabase).)
</Note>
## Architecture
```mermaid
flowchart LR
subgraph SB["Supabase: system of record"]
K[(knowledge / employees / teams)]
AU[Auth · who is asking]
end
subgraph M0["Mem0 Platform: the brain"]
B[(managed memory)]
end
K -->|ingest| B
AU -->|maps to scope| B
CC[Claude Code] --> MCP[Mem0 MCP server]
CU[Cursor] --> MCP
SL[Slack bot] --> MCP
MCP --> B
```
Memory splits by entity. An individual is a **`user_id`** (their Supabase Auth id). Shared knowledge lives on an **`agent_id`**: the company-wide brain is `org:acme`, and each team is its own agent, e.g. `team:payments`. A person's own facts route to their `user_id`; company and team facts route to the agent. This split is what lets one search return "my" context alongside the shared org knowledge.
## Prerequisites
- **Python 3.9+**
- A **Mem0 Platform API key**, [app.mem0.ai/dashboard/api-keys](https://app.mem0.ai/dashboard/api-keys?utm_source=oss&utm_medium=cookbook-company-brain). (Platform runs extraction and embeddings for you, so there's no OpenAI key to manage.)
- A **Supabase** project, [supabase.com](https://supabase.com)
About 20 minutes.
---
## Step 1: Get your Mem0 Platform API key
Sign in at [app.mem0.ai](https://app.mem0.ai) and copy a key from **Dashboard → API Keys**. The key is scoped to your org and project; Mem0 resolves both server-side, so you never pass IDs by hand.
## Step 2: Create the Supabase system of record
In the Supabase **SQL editor**, create the tables your company already thinks in: people, teams, and a `knowledge` table the brain will ingest from. Identity reuses Supabase Auth's built-in `auth.users`.
```sql
-- Employees extend Supabase Auth's users; identity is auth.users.id (uuid)
create table public.employees (
id uuid primary key references auth.users (id) on delete cascade,
name text not null,
team text not null
);
-- The company knowledge the brain ingests. `scope` decides who can recall it.
create table public.knowledge (
id bigint generated always as identity primary key,
scope text not null, -- the shared agent this belongs to: 'org:acme' | 'team:payments'
content text not null,
author uuid references auth.users (id), -- who recorded it (their user_id); null for org seed data
created_at timestamptz default now(),
mem0_synced_at timestamptz -- null until ingested into the brain
);
create index on public.knowledge (mem0_synced_at, created_at);
-- Seed a little company knowledge to ingest.
insert into public.knowledge (scope, content) values
('org:acme', 'We chose Postgres over MongoDB for the core product for strong transactional guarantees and relational joins.'),
('org:acme', 'All production deploys go out Tuesday and Thursday; never on Fridays.'),
('org:acme', 'Customer data must stay in the EU region for GDPR compliance.'),
('org:acme', 'Billing is owned by the Payments team, and Alice is the Payments tech lead.'),
('team:payments', 'Stripe is our processor; webhooks are verified with PAYMENTS_WEBHOOK_SECRET.');
```
Grab your project URL and **service-role** key from **Settings → API** (the ingestion job runs server-side and needs to read every scope).
## Step 3: Project setup
```bash
mkdir company-brain && cd company-brain
pip install "mem0ai>=2.0.17" supabase requests # 2.0.17+ for agent_custom_instructions
```
```bash
export MEM0_API_KEY="m0-..."
export SUPABASE_URL="https://<project-ref>.supabase.co"
export SUPABASE_SERVICE_KEY="<service-role-key>"
```
## Step 4: Configure the brain
Create **`brain.py`**. This constructs the Platform client and teaches it what to remember. The key is the **two** instruction sets: `custom_instructions` governs a person's own (`user_id`) memories, and `agent_custom_instructions` governs shared (`agent_id`) memories, phrased in the third person so company facts read "The company…", not "The user's organization…". `custom_categories` files each memory under a useful label.
```python
# brain.py
import os
from mem0 import MemoryClient
client = MemoryClient(api_key=os.environ["MEM0_API_KEY"])
# Steer extraction (project-wide). Runs server-side; no LLM key needed here.
client.project.update(
# Governs a person's OWN memories (user_id).
custom_instructions=(
"Extract the individual's own durable preferences, context, and how they work. "
"Ignore greetings and one-off chatter."
),
# Governs SHARED memories (agent_id); write them in the third person.
agent_custom_instructions=(
"Extract durable company/team knowledge in the third person "
"(\"The company...\", \"The team...\"): decisions and their rationale, ownership "
"(who owns what), processes, policies, tooling choices, and gotchas. "
"Ignore greetings, scheduling, and one-off chatter."
),
custom_categories=[
{"decision": "Architectural or product decisions and why they were made"},
{"ownership": "Who owns a system, service, or process"},
{"policy": "Compliance, security, and process rules"},
{"tooling": "Tools, services, and how they're configured"},
],
)
# Scopes. A person is a user_id; shared brains are agent_ids.
COMPANY = "org:acme" # agent_id: company-wide shared brain
def team(name): return f"team:{name}" # agent_id: a team's shared brain
def person(uid): return uid # user_id: an individual (Supabase auth id)
```
Run it once to apply the project settings:
```bash
python -c "import brain; print('brain configured')"
```
## Step 5: Ingest company knowledge from Supabase
This is where Supabase and the brain connect. Create **`ingest.py`**: read un-synced rows from `knowledge`, add each to the Platform brain under its scope, then mark it synced. Platform `add()` is **asynchronous**, it returns an `event_id` you can poll, so we include a small `wait_for` helper.
```python
# ingest.py
import os, time, requests
from supabase import create_client
from brain import client, person
sb = create_client(os.environ["SUPABASE_URL"], os.environ["SUPABASE_SERVICE_KEY"])
MEM0_HEADERS = {"Authorization": f"Token {os.environ['MEM0_API_KEY']}"}
def wait_for(event_id, timeout=30):
"""Platform extraction is async; poll the event until it settles."""
for _ in range(timeout):
r = requests.get(f"https://api.mem0.ai/v1/event/{event_id}/", headers=MEM0_HEADERS).json()
if r.get("status") in ("SUCCEEDED", "FAILED"):
return r["status"]
time.sleep(1)
return "TIMEOUT"
# 1. Read knowledge that hasn't been ingested yet
rows = sb.table("knowledge").select("*").is_("mem0_synced_at", "null").execute().data
for row in rows:
# 2. Add it. agent_id = the shared scope (org/team); user_id = who recorded it.
# Mem0 routes shared facts to the agent and personal facts to the individual,
# so pass both when there's an author.
add_kwargs = {
"agent_id": row["scope"],
"metadata": {"source": "supabase", "knowledge_id": row["id"]},
}
if row["author"]:
add_kwargs["user_id"] = person(row["author"])
res = client.add([{"role": "user", "content": row["content"]}], **add_kwargs)
# 3. Platform returns an event_id; wait for extraction to finish
event_id = res.get("event_id") if isinstance(res, dict) else None
if event_id:
wait_for(event_id)
# 4. Mark the row synced so we never double-ingest
sb.table("knowledge").update({"mem0_synced_at": "now()"}).eq("id", row["id"]).execute()
print(f"Ingested {len(rows)} knowledge items into the company brain.")
```
```bash
python ingest.py
```
```text
Ingested 5 knowledge items into the company brain.
```
Re-running is safe, `mem0_synced_at` gates it, so a nightly cron can keep the brain in step with Supabase.
## Step 6: Ask the brain
Create **`ask.py`**. It searches everything relevant to the asker: their own (`user_id`) memories **plus** the shared company and team (`agent_id`) memories. This has to be an **`OR`**, each memory row belongs to exactly one entity, so a flat filter or an `AND` of a `user_id` and an `agent_id` matches nothing.
```python
# ask.py
import sys
from brain import client, COMPANY, team, person
def ask(question: str, uid: str | None = None, user_team: str | None = None) -> str:
scopes = [{"agent_id": COMPANY}] # company-wide brain
if user_team:
scopes.append({"agent_id": team(user_team)}) # the asker's team
if uid:
scopes.append({"user_id": person(uid)}) # the asker's own memories
hits = client.search(
query=question,
filters={"OR": scopes}, # OR, never AND (one FK per memory row)
top_k=5,
rerank=True,
)
return "\n".join(f"- {h['memory']}" for h in hits.get("results", hits))
if __name__ == "__main__":
print(ask(" ".join(sys.argv[1:]) or "When can we deploy?"))
```
```bash
python ask.py "Why did we pick Postgres, and can I deploy on Friday?"
```
```text
- The company chose Postgres over MongoDB for strong transactional guarantees and relational joins
- The company's production deploys go out Tuesday and Thursday, never on Fridays
```
Search returns every relevant memory, so a question resolves across separate facts, here it pulls both the owning team and the person:
```bash
python ask.py "Who should I talk to about billing?"
```
```text
- Billing is owned by the Payments team
- Alice is the Payments tech lead
```
## Step 7: Sharper retrieval
Platform search is hybrid (semantic + keyword) and filterable. Combine a keyword pass with a category filter to answer precise questions:
```python
client.search(
query="webhook signing secret",
filters={"agent_id": "team:payments", "categories": {"in": ["tooling"]}},
keyword_search=True, # hybrid keyword + semantic
rerank=True,
threshold=0.3,
)
```
Filters use keyword operators (`in`, `gte`, `contains`, …) and AND/OR/NOT, so you can scope by date, category, or metadata, for example the company's policies added this quarter:
```python
client.search(
query="compliance rules",
filters={"AND": [
{"agent_id": "org:acme"},
{"categories": {"in": ["policy"]}},
{"created_at": {"gte": "2026-01-01"}},
]},
)
```
## Step 8: Expose the brain to every agent (MCP)
A brain only your script can reach isn't a company brain. Mem0's **hosted MCP server** lets any agent (Claude Code, Cursor, a Slack bot) query and contribute to the *same* brain. The endpoint is `https://mcp.mem0.ai/mcp`, and the supported way to connect is the `mcp-add` helper, which registers the server and runs Mem0's OAuth login so no key ever lands in a config file.
<Tabs>
<Tab title="Claude Code / Cursor">
```bash
npx mcp-add --url "https://mcp.mem0.ai/mcp" --clients "claude code,cursor"
```
Complete the browser login on first connect. Now the agent has the brain's memory tools (`add_memory`, `search_memories`, and more) available in-editor.
</Tab>
<Tab title="Manual (.mcp.json)">
```json
{
"mcpServers": {
"mem0": { "url": "https://mcp.mem0.ai/mcp" }
}
}
```
Auth happens via Mem0's OAuth flow on first use, don't paste a static token into the file (the hosted gateway may reject a raw `Token` header).
</Tab>
<Tab title="Slack bot">
```python
# A Slack bot is just another MCP client. Point its MCP layer at the same URL,
# authenticate via Mem0's OAuth flow, and pass the company scope on each call.
await mcp.call_tool("search_memories", {
"query": user_message,
"agent_id": "org:acme",
})
```
</Tab>
</Tabs>
With this, an engineer asks the brain from their editor and a teammate asks it from Slack, one shared memory behind both.
## Step 9: Onboard a new hire (the payoff)
This is what a company brain is *for*. Dana joins, and her identity comes from **Supabase Auth**, which maps straight to her Mem0 `user_id`. She asks the questions every new hire asks and gets real answers on day one, drawn from the shared company (and her team's) brain, plus anything she's told it herself.
```python
# onboarding.py
from brain import client, person
from ask import ask
# In a real app these come from sb.auth.get_user(jwt) and the employees table.
dana_uid, dana_team = "8f3c...-dana", "payments"
# Dana also tells the brain how *she* works. This is personal, so it goes to her
# user_id, not the shared agent, and stays scoped to her.
client.add(
[{"role": "user", "content": "I prefer early returns over nested ifs, and I review PRs in the morning."}],
user_id=person(dana_uid),
)
for q in [
"Who owns billing and who do I talk to?", # company (agent) knowledge
"When are deploys, and are there hard rules?",
"How do I like to write code?", # Dana's own (user) knowledge
]:
print(f"Q: {q}\nA: {ask(q, uid=dana_uid, user_team=dana_team)}\n")
```
```text
Q: Who owns billing and who do I talk to?
A: - Billing is owned by the Payments team; Alice is the Payments tech lead
Q: When are deploys, and are there hard rules?
A: - The company's production deploys go out Tuesday and Thursday, never on Fridays
Q: How do I like to write code?
A: - User prefers early returns over nested ifs
```
The same `ask()` blends the shared company facts with Dana's own preference, because the `OR` filter spans both her `user_id` and the org and team `agent_id`s.
Dana onboarded herself by asking, drawing on the shared brain the rest of the team had been filling.
## Production notes
<Warning>
**`user_id` vs `agent_id`.** An individual is a `user_id`; shared brains (company, team) are `agent_id`s. Keeping them separate is what gives you the third-person "The company…" framing and lets a person's own context sit alongside org knowledge. Put a secret like a webhook key on a **team** agent, never the company agent, or everyone can recall it, and mirror the boundary in Supabase with a Row Level Security policy on `knowledge`.
</Warning>
<Warning>
**Search must `OR` the scopes.** A memory row belongs to exactly one entity, so `filters={"OR": [{"user_id": ...}, {"agent_id": "org:acme"}, {"agent_id": "team:..."}]}`. A flat filter, or an `AND` of a `user_id` and an `agent_id`, returns nothing.
</Warning>
<Warning>
**`add()` is asynchronous.** It returns `{event_id, status: "PENDING"}` and extraction finishes a moment later, poll `GET /v1/event/{event_id}/` (as in Step 5) when you need to know a write has landed before searching for it.
</Warning>
<Note>
**Where the entity ID goes differs by call.** `search()` and `get_all()` take the scope inside `filters={...}` (a top-level `user_id=`/`agent_id=` is rejected). `add()` and `delete_all()` are the opposite, they take it as a top-level keyword: `client.delete_all(agent_id="team:payments")`. Deletes are asynchronous too, so a `get_all` right after a `delete_all` can still show rows for a few seconds.
</Note>
## Where to take it next
- **Auto-feed the brain** from PR descriptions, RFCs, and incident write-ups so it grows without anyone thinking about it, just insert into Supabase `knowledge` and let the cron ingest.
- **Scope by real identity** end to end: verify the Supabase JWT, read `sb.auth.get_user(jwt).user.id` for the `user_id`, look up the person's team, and `OR` their `user_id` with the company and team `agent_id`s on every recall.
- **Give teams a private view** with Supabase RLS so `team:` knowledge is only readable by that team.
---
<CardGroup cols={2}>
<Card title="Mem0 MCP Server" icon="plug" href="/platform/mem0-mcp">
Connect any agent or editor to the brain over MCP.
</Card>
<Card title="Custom Categories & Instructions" icon="sliders" href="/platform/features/custom-instructions">
Steer exactly what the brain extracts and how it's filed.
</Card>
</CardGroup>
<Snippet file="star-on-github.mdx" />
+2 -2
View File
@@ -89,8 +89,8 @@ On Mem0 Platform, these stores are managed for you. In OSS, you choose and opera
## Next steps
<CardGroup cols={3}>
<Card title="Entity scoping" icon="brain" href="/platform/features/entity-scoped-memory">
Organize Platform memories by user, agent, app, and run.
<Card title="Memory types" icon="brain" href="/core-concepts/memory-types">
Choose the right scope for user, agent, run, and session memory.
</Card>
<Card title="Memory operations" icon="database" href="/core-concepts/memory-operations/add">
Add, search, update, and delete memories from your app.
+118
View File
@@ -0,0 +1,118 @@
---
title: Memory Types
description: "What memory_type actually does in Mem0: procedural memory is implemented, semantic and episodic are not."
icon: "tag"
iconType: "solid"
---
# Memory Types
Mem0's Python SDK exposes a `memory_type` parameter on `add()`. The underlying `MemoryType` enum defines three values, but only one of them is wired up. This page states plainly which is which so you don't build against a type that doesn't exist yet.
## Status
| Type | Enum value | Status | Notes |
| --- | --- | --- | --- |
| Procedural memory | `procedural_memory` | **Implemented** | Python OSS only (`Memory`/`AsyncMemory`). Pass `memory_type="procedural_memory"` and `agent_id` to `add()`. Not available on the Platform `MemoryClient`, and not available in the TypeScript SDK (OSS or Platform). |
| Semantic memory | `semantic_memory` | **Not implemented** | Defined in the `MemoryType` enum but never read anywhere else in the codebase. Passing it to `add()` raises a validation error. There is no evidence in this repo of a roadmap date for this. |
| Episodic memory | `episodic_memory` | **Not implemented** | Same as above: defined, never wired into the extraction pipeline, rejected by validation, no documented roadmap. |
<Warning>
Only `procedural_memory` is a real, working value. Calling `memory.add(messages, memory_type="semantic_memory")` (or `episodic_memory`) is rejected and tells you to pass `procedural_memory` instead. Sync `Memory.add()` raises `Mem0ValidationError`; `AsyncMemory.add()` raises a plain `ValueError`.
</Warning>
## Procedural memory
Procedural memory stores step-by-step task knowledge (how an agent performs a workflow) rather than facts about a user. It requires `agent_id`:
```python
from mem0 import Memory
memory = Memory()
memory.add(
[
{"role": "user", "content": "Book a flight from SFO to NYC"},
{"role": "assistant", "content": "1. Search flights. 2. Filter by price. 3. Confirm booking."},
],
agent_id="travel-agent",
memory_type="procedural_memory",
)
```
Omit `memory_type` entirely and Mem0 stores the messages as an ordinary memory: there is no semantic/episodic pathway for it to fall into. Any other explicit value is rejected by validation rather than quietly falling back to an ordinary memory.
## How every other memory is scoped
Outside of the `procedural_memory` special case, Mem0 does not sort memories into named types. Every memory is scoped by the identifiers you pass in, and the same identifiers are used to retrieve it later:
- **`user_id`**: ties a memory to a specific person or account.
- **`agent_id`**: ties a memory to a specific agent or assistant persona.
- **`run_id`**: ties a memory to a specific session, task, or conversation thread.
- **`app_id`** (Platform only): ties a memory to a specific application or tenant, in addition to the three above. See <Link href="/platform/features/entity-scoped-memory">Entity-Scoped Memory</Link>.
At least one identifier is required on `add()`. Passing more than one narrows the scope further (for example, `user_id` + `run_id` together).
```python
from mem0 import Memory
memory = Memory()
memory.add(
"I'm Alex and I prefer boutique hotels.",
user_id="alex",
run_id="trip-planning-2025",
)
results = memory.search(
"Any hotel preferences?",
filters={"user_id": "alex", "run_id": "trip-planning-2025"},
)
```
<Tip>
Use `run_id` when you want a set of memories to stay tied to one session or task; use `user_id` alone for anything that should persist across every session for that person.
</Tip>
## How memories are extracted and updated
When `infer=True` (the default) on `add()`, Mem0 runs a single pipeline rather than routing through separate type-specific paths:
1. **Context gathering**: pulls the most recent messages already stored for the same `user_id`/`agent_id`/`run_id` scope.
2. **Existing memory retrieval**: embeds the new messages and runs a vector search against memories already in that same scope, to find candidates that might need to change.
3. **Extraction**: a single LLM call compares the new messages against the retrieved candidates and decides, per fact, whether to `ADD`, `UPDATE`, `DELETE`, or leave a memory alone.
Alongside this, both OSS and Platform extract named entities (people, places, organizations) from memory text and use shared entities between memories to boost related results at search time. On Platform, that entity graph is also queryable directly; see <Link href="/platform/features/graph-memory">Graph Memory</Link>. In OSS, entities only affect ranking, there is no separate graph to query.
<Warning>
Avoid storing secrets or unredacted PII in memories: they are retrievable by design. Encrypt or hash sensitive values before calling `add()`.
</Warning>
## Put it into practice
<CardGroup cols={2}>
<Card
title="Explore Memory Operations"
description="Dive into the add/search/update/delete operations next."
icon="circle-check"
href="/core-concepts/memory-operations/add"
/>
<Card
title="Advanced Memory Operations"
description="Tune metadata, filters, and retrieval on Platform."
icon="sliders"
href="/platform/advanced-memory-operations"
/>
<Card
title="AI Tutor Cookbook"
description="See user_id-scoped memory used in a real tutoring agent."
icon="rocket"
href="/cookbooks/companions/ai-tutor"
/>
<Card
title="Support Inbox Cookbook"
description="See user_id-scoped memory used in a support workflow."
icon="inbox"
href="/cookbooks/operations/support-inbox"
/>
</CardGroup>
+3 -26
View File
@@ -54,6 +54,7 @@
"icon": "brain",
"pages": [
"core-concepts/how-it-works",
"core-concepts/memory-types",
"core-concepts/memory-operations/add",
"core-concepts/memory-operations/search",
"core-concepts/memory-operations/update",
@@ -71,7 +72,6 @@
"pages": [
"platform/features/v2-memory-filters",
"platform/features/entity-scoped-memory",
"platform/features/user-profiles",
"platform/features/graph-memory",
"platform/features/async-client",
"platform/features/multimodal-support",
@@ -366,13 +366,6 @@
{
"tab": "Agent Plugins",
"groups": [
{
"group": "Overview",
"icon": "puzzle-piece",
"pages": [
"integrations/agent-plugins"
]
},
{
"group": "Coding Agents",
"icon": "terminal",
@@ -451,8 +444,7 @@
"cookbooks/integrations/mastra-agent",
"cookbooks/integrations/healthcare-google-adk",
"cookbooks/integrations/aws-bedrock",
"cookbooks/integrations/tavily-search",
"cookbooks/integrations/supabase"
"cookbooks/integrations/tavily-search"
]
},
{
@@ -520,17 +512,6 @@
"api-reference/entities/delete-user"
]
},
{
"group": "Profiles",
"icon": "id-card",
"pages": [
"api-reference/profiles/get-profile",
"api-reference/profiles/get-profile-settings",
"api-reference/profiles/update-profile-settings",
"api-reference/profiles/generate-profiles",
"api-reference/profiles/get-profile-job"
]
},
{
"group": "Organizations",
"icon": "building",
@@ -1039,11 +1020,7 @@
},
{
"source": "/concepts/memory-scoring",
"destination": "/core-concepts/how-it-works"
},
{
"source": "/core-concepts/memory-types",
"destination": "/core-concepts/how-it-works"
"destination": "/core-concepts/memory-types"
},
{
"source": "/cookbooks/research-copilot",
File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 7.5 KiB

@@ -1 +0,0 @@
<svg height="1em" style="flex:none;line-height:1" viewBox="0 0 24 24" width="1em" xmlns="http://www.w3.org/2000/svg"><title>Claude</title><path d="M4.709 15.955l4.72-2.647.08-.23-.08-.128H9.2l-.79-.048-2.698-.073-2.339-.097-2.266-.122-.571-.121L0 11.784l.055-.352.48-.321.686.06 1.52.103 2.278.158 1.652.097 2.449.255h.389l.055-.157-.134-.098-.103-.097-2.358-1.596-2.552-1.688-1.336-.972-.724-.491-.364-.462-.158-1.008.656-.722.881.06.225.061.893.686 1.908 1.476 2.491 1.833.365.304.145-.103.019-.073-.164-.274-1.355-2.446-1.446-2.49-.644-1.032-.17-.619a2.97 2.97 0 01-.104-.729L6.283.134 6.696 0l.996.134.42.364.62 1.414 1.002 2.229 1.555 3.03.456.898.243.832.091.255h.158V9.01l.128-1.706.237-2.095.23-2.695.08-.76.376-.91.747-.492.584.28.48.685-.067.444-.286 1.851-.559 2.903-.364 1.942h.212l.243-.242.985-1.306 1.652-2.064.73-.82.85-.904.547-.431h1.033l.76 1.129-.34 1.166-1.064 1.347-.881 1.142-1.264 1.7-.79 1.36.073.11.188-.02 2.856-.606 1.543-.28 1.841-.315.833.388.091.395-.328.807-1.969.486-2.309.462-3.439.813-.042.03.049.061 1.549.146.662.036h1.622l3.02.225.79.522.474.638-.079.485-1.215.62-1.64-.389-3.829-.91-1.312-.329h-.182v.11l1.093 1.068 2.006 1.81 2.509 2.33.127.578-.322.455-.34-.049-2.205-1.657-.851-.747-1.926-1.62h-.128v.17l.444.649 2.345 3.521.122 1.08-.17.353-.608.213-.668-.122-1.374-1.925-1.415-2.167-1.143-1.943-.14.08-.674 7.254-.316.37-.729.28-.607-.461-.322-.747.322-1.476.389-1.924.315-1.53.286-1.9.17-.632-.012-.042-.14.018-1.434 1.967-2.18 2.945-1.726 1.845-.414.164-.717-.37.067-.662.401-.589 2.388-3.036 1.44-1.882.93-1.086-.006-.158h-.055L4.132 18.56l-1.13.146-.487-.456.061-.746.231-.243 1.908-1.312-.006.006z" fill="#D97757" fill-rule="nonzero"></path></svg>

Before

Width:  |  Height:  |  Size: 1.7 KiB

@@ -1 +0,0 @@
<svg height="1em" style="flex:none;line-height:1" viewBox="0 0 24 24" width="1em" xmlns="http://www.w3.org/2000/svg"><title>Claude Code</title><path clip-rule="evenodd" d="M20.998 10.949H24v3.102h-3v3.028h-1.487V20H18v-2.921h-1.487V20H15v-2.921H9V20H7.488v-2.921H6V20H4.487v-2.921H3V14.05H0V10.95h3V5h17.998v5.949zM6 10.949h1.488V8.102H6v2.847zm10.51 0H18V8.102h-1.49v2.847z" fill="#D97757" fill-rule="evenodd"></path></svg>

Before

Width:  |  Height:  |  Size: 424 B

@@ -1 +0,0 @@
<svg height="1em" style="flex:none;line-height:1" viewBox="0 0 24 24" width="1em" xmlns="http://www.w3.org/2000/svg"><title>Codex</title><path d="M19.503 0H4.496A4.496 4.496 0 000 4.496v15.007A4.496 4.496 0 004.496 24h15.007A4.496 4.496 0 0024 19.503V4.496A4.496 4.496 0 0019.503 0z" fill="#fff"></path><path d="M9.064 3.344a4.578 4.578 0 012.285-.312c1 .115 1.891.54 2.673 1.275.01.01.024.017.037.021a.09.09 0 00.043 0 4.55 4.55 0 013.046.275l.047.022.116.057a4.581 4.581 0 012.188 2.399c.209.51.313 1.041.315 1.595a4.24 4.24 0 01-.134 1.223.123.123 0 00.03.115c.594.607.988 1.33 1.183 2.17.289 1.425-.007 2.71-.887 3.854l-.136.166a4.548 4.548 0 01-2.201 1.388.123.123 0 00-.081.076c-.191.551-.383 1.023-.74 1.494-.9 1.187-2.222 1.846-3.711 1.838-1.187-.006-2.239-.44-3.157-1.302a.107.107 0 00-.105-.024c-.388.125-.78.143-1.204.138a4.441 4.441 0 01-1.945-.466 4.544 4.544 0 01-1.61-1.335c-.152-.202-.303-.392-.414-.617a5.81 5.81 0 01-.37-.961 4.582 4.582 0 01-.014-2.298.124.124 0 00.006-.056.085.085 0 00-.027-.048 4.467 4.467 0 01-1.034-1.651 3.896 3.896 0 01-.251-1.192 5.189 5.189 0 01.141-1.6c.337-1.112.982-1.985 1.933-2.618.212-.141.413-.251.601-.33.215-.089.43-.164.646-.227a.098.098 0 00.065-.066 4.51 4.51 0 01.829-1.615 4.535 4.535 0 011.837-1.388zm3.482 10.565a.637.637 0 000 1.272h3.636a.637.637 0 100-1.272h-3.636zM8.462 9.23a.637.637 0 00-1.106.631l1.272 2.224-1.266 2.136a.636.636 0 101.095.649l1.454-2.455a.636.636 0 00.005-.64L8.462 9.23z" fill="url(#lobe-icons-codex-_R_0_)"></path><defs><linearGradient gradientUnits="userSpaceOnUse" id="lobe-icons-codex-_R_0_" x1="12" x2="12" y1="3" y2="21"><stop stop-color="#B1A7FF"></stop><stop offset=".5" stop-color="#7A9DFF"></stop><stop offset="1" stop-color="#3941FF"></stop></linearGradient></defs></svg>

Before

Width:  |  Height:  |  Size: 1.7 KiB

-1
View File
@@ -1 +0,0 @@
<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24"><title>Cursor</title><rect x="0.25" y="0.25" width="23.5" height="23.5" rx="5.5" fill="#161616" stroke="#3a3a3a" stroke-width="0.5"/><g transform="translate(4.000 4.000) scale(0.66667)" fill="#ffffff" fill-rule="evenodd"><path d="M22.106 5.68L12.5.135a.998.998 0 00-.998 0L1.893 5.68a.84.84 0 00-.419.726v11.186c0 .3.16.577.42.727l9.607 5.547a.999.999 0 00.998 0l9.608-5.547a.84.84 0 00.42-.727V6.407a.84.84 0 00-.42-.726zm-.603 1.176L12.228 22.92c-.063.108-.228.064-.228-.061V12.34a.59.59 0 00-.295-.51l-9.11-5.26c-.107-.062-.063-.228.062-.228h18.55c.264 0 .428.286.296.514z"/></g></svg>

Before

Width:  |  Height:  |  Size: 671 B

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 19 KiB

@@ -1 +0,0 @@
<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24"><title>Kimi</title><rect x="0.25" y="0.25" width="23.5" height="23.5" rx="5.5" fill="#161616" stroke="#3a3a3a" stroke-width="0.5"/><g transform="translate(4 4) scale(0.66667)"><path d="M21.846 0a1.923 1.923 0 110 3.846H20.15a.226.226 0 01-.227-.226V1.923C19.923.861 20.784 0 21.846 0z" fill="#1783FF"></path><path d="M11.065 11.199l7.257-7.2c.137-.136.06-.41-.116-.41H14.3a.164.164 0 00-.117.051l-7.82 7.756c-.122.12-.302.013-.302-.179V3.82c0-.127-.083-.23-.185-.23H3.186c-.103 0-.186.103-.186.23V19.77c0 .128.083.23.186.23h2.69c.103 0 .186-.102.186-.23v-3.25c0-.069.025-.135.069-.178l2.424-2.406a.158.158 0 01.205-.023l6.484 4.772a7.677 7.677 0 003.453 1.283c.108.012.2-.095.2-.23v-3.06c0-.117-.07-.212-.164-.227a5.028 5.028 0 01-2.027-.807l-5.613-4.064c-.117-.078-.132-.279-.028-.381z" fill="#fff"></path></g></svg>

Before

Width:  |  Height:  |  Size: 900 B

-1
View File
@@ -1 +0,0 @@
<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24"><title>opencode</title><rect x="0.25" y="0.25" width="23.5" height="23.5" rx="5.5" fill="#161616" stroke="#3a3a3a" stroke-width="0.5"/><g transform="translate(4.000 4.000) scale(0.66667)" fill="#ffffff" fill-rule="evenodd"><path d="M16 6H8v12h8V6zm4 16H4V2h16v20z"/></g></svg>

Before

Width:  |  Height:  |  Size: 359 B

-1
View File
@@ -1 +0,0 @@
<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24"><title>Pi</title><rect x="0.25" y="0.25" width="23.5" height="23.5" rx="5.5" fill="#161616" stroke="#3a3a3a" stroke-width="0.5"/><g transform="translate(4.000 4.000) scale(0.02857)" fill="#ffffff" fill-rule="evenodd"><path d="M420 280H280V140H0V0H420V280Z"/><path d="M560 560H420V280H560V560Z"/><path d="M140 560H0V140H140V280H280V420H140V560Z"/></g></svg>

Before

Width:  |  Height:  |  Size: 439 B

-111
View File
@@ -1,111 +0,0 @@
---
title: "Agent Plugins Overview"
sidebarTitle: "Overview"
description: "Add persistent memory to the coding agent or agent harness you already use. Compare the Mem0 plugins and pick yours."
---
Coding agents forget everything when a session ends: your conventions, the bug you fixed yesterday, the command that finally worked. A Mem0 plugin fixes that inside the tool you already use. It captures what matters while you work and brings it back in the next session.
<Info>
Plugins use a [Mem0 Platform](https://app.mem0.ai?utm_source=oss&utm_medium=agent-plugins-overview) account. OpenClaw and Hermes can also run fully self-hosted.
</Info>
## How a plugin works
<Steps>
<Step title="Capture">
While you work, the plugin records your prompts, the agent's answers and useful tool results. Credentials are redacted before anything leaves your machine.
</Step>
<Step title="Extract">
Mem0 turns each session into short, durable memories: project conventions, decisions, fixes that worked, and your personal preferences.
</Step>
<Step title="Recall">
In the next session, relevant memories are added to the agent's context automatically or found with a search tool, depending on the plugin.
</Step>
</Steps>
## Coding agents
<CardGroup cols={3}>
<Card title="Claude Code" icon="/images/provider-icons/claudecode-color.svg" href="/integrations/claude-code">
Automatic capture and recall, six commands, and the Sidekick agent.
</Card>
<Card title="Codex" icon="/images/provider-icons/codex-color.svg" href="/integrations/codex">
Automatic capture and recall in the Codex CLI and app.
</Card>
<Card title="Cursor" icon="/images/provider-icons/cursor.svg" href="/integrations/cursor">
Automatic capture, with recall through a search tool.
</Card>
<Card title="OpenCode" icon="/images/provider-icons/opencode.svg" href="/integrations/opencode">
Recall before every prompt and ten native memory tools.
</Card>
<Card title="Kimi Code" icon="/images/provider-icons/kimi-color.svg" href="/integrations/kimi">
Automatic capture and recall in the Kimi Code CLI.
</Card>
<Card title="Antigravity" icon="/images/provider-icons/antigravity-color.svg" href="/integrations/antigravity">
Automatic capture, with recall through a search tool.
</Card>
</CardGroup>
## Agent harnesses and chat apps
<CardGroup cols={3}>
<Card title="OpenClaw" icon="/images/provider-icons/openclaw.svg" href="/integrations/openclaw">
Recall and capture on every turn. Platform or self-hosted.
</Card>
<Card title="Hermes Agent" icon="/images/provider-icons/hermes-tile.svg" href="/integrations/hermes">
Mem0 as Hermes' memory provider. Platform or self-hosted.
</Card>
<Card title="Pi Agent" icon="/images/provider-icons/pi.svg" href="/integrations/pi-agent">
Recall before every turn and automatic capture.
</Card>
<Card title="DeepSeek Harness" icon="/images/provider-icons/deepseek-color.svg" href="/integrations/deepseek-plugin">
Automatic recall and capture with two native tools.
</Card>
<Card title="Claude.ai" icon="/images/provider-icons/claude-color.svg" href="/integrations/claude-ai">
Add the verified Mem0 connector from Claude's directory and sign in. No API key needed.
</Card>
<Card title="Any MCP client" icon="plug" href="/platform/mem0-mcp">
Your tool is not listed? Connect it to the hosted Mem0 MCP server.
</Card>
</CardGroup>
## Compare plugins
<Tabs>
<Tab title="Coding agents">
| Plugin | Auto capture | Auto recall | Skills | Team |
| --- | --- | --- | --- | --- |
| [Claude Code](/integrations/claude-code) | Yes | First prompt | 6, plus Sidekick | Shared |
| [Codex](/integrations/codex) | Yes | First prompt | 6 | Shared |
| [Cursor](/integrations/cursor) | Yes | Search tool | 6 | Shared |
| [Kimi Code](/integrations/kimi) | Yes | First prompt | 6 | Shared |
| [Antigravity](/integrations/antigravity) | Yes | Search tool | 6 | Shared |
| [OpenCode](/integrations/opencode) | Yes | Every prompt | 7 | Per user |
**First prompt:** relevant memories are added before the agent answers the first prompt of a session. **Search tool:** the agent calls `search_memories` when it needs context. **Shared:** teammates on the same repo share project memory.
</Tab>
<Tab title="Harnesses and chat apps">
| Plugin | Auto capture | Auto recall | Tools | Self-hosted |
| --- | --- | --- | --- | --- |
| [OpenClaw](/integrations/openclaw) | Yes | Every turn | Tools and skills | Yes |
| [Hermes Agent](/integrations/hermes) | Yes | Every turn | 4 | Yes |
| [Pi Agent](/integrations/pi-agent) | Yes | Every turn | 1 | No |
| [DeepSeek Harness](/integrations/deepseek-plugin) | Yes | Every prompt | 2 | No |
| [Claude.ai](/integrations/claude-ai) | When asked | When relevant | 9 | No |
</Tab>
</Tabs>
<Note>
Claude Code, Codex, Cursor, Kimi Code and Antigravity share the same memory core. Teammates working in the same repository, in any of these tools, read and write one shared project memory while their personal preferences stay private.
</Note>
## Which plugin should I use?
- **You code in a terminal or IDE agent:** install the plugin for that tool from the cards above. Each page has its own install steps.
- **Your team uses different agents on the same repo:** pick any of Claude Code, Codex, Cursor, Kimi Code or Antigravity. They share project memory across tools.
- **You chat in Claude:** add the [Mem0 connector](https://claude.com/connectors/mem0) from Claude's directory.
- **Your tool has no plugin:** connect it to the [hosted Mem0 MCP server](/platform/mem0-mcp).
- **You need to keep data on your own infrastructure:** use OpenClaw or Hermes in self-hosted mode.
<Snippet file="star-on-github.mdx" />
+3 -3
View File
@@ -130,10 +130,10 @@ print(response.msgs[0].content)
<CardGroup cols={2}>
<Card
title="How Mem0 works"
description="Understand how Mem0 extracts, stores, and retrieves memories for your Camel agents."
title="Memory types in Mem0"
description="Choose between chat history and semantic search for your Camel agents."
icon="sparkles"
href="/core-concepts/how-it-works"
href="/core-concepts/memory-types"
/>
<Card
title="Try LangChain next"
+92 -116
View File
@@ -1,9 +1,9 @@
---
title: Hermes Agent
description: "Add persistent memory to Hermes Agent with Mem0 Cloud, a self-hosted server, or the in-process OSS SDK."
description: "Add long-term memory to Hermes agents using Mem0 Platform, a self-hosted server, or local OSS mode with background fact extraction."
---
Add long-term memory to [Hermes Agent](https://github.com/NousResearch/hermes-agent), a self-improving AI agent CLI by Nous Research. The [standalone Mem0 plugin](https://github.com/mem0ai/mem0/tree/main/integrations/hermes-plugin-mem0) learns facts from conversations and recalls relevant memories for the current question.
Add long-term memory to [Hermes Agent](https://github.com/NousResearch/hermes-agent), a self-improving AI agent CLI by Nous Research. Hermes has a pluggable memory system, and Mem0 is one of the supported providers. Once enabled, Mem0 learns facts from your conversations and surfaces relevant ones for the current question, without slowing down the chat.
You can run Mem0 in three ways:
@@ -17,15 +17,11 @@ Hermes runs a built-in memory system (file-based `MEMORY.md` and `USER.md`) alon
### 1. Current-turn recall (bounded wait)
When you send a message, Hermes searches your stored memories for the current question and waits up to 3 seconds for results. If they arrive in time, they are injected into the system prompt so the model can see them. If the backend is slower, Hermes skips the injection and the model can still call `mem0_search` itself after the bounded recall wait.
When you send a message, Hermes searches your stored memories for the current question and waits up to 3 seconds for results. If they arrive in time, they are injected into the system prompt so the model can see them. If the backend is slower, Hermes skips the injection and the model can still call `mem0_search` itself — so a slow backend never blocks a turn.
### 2. Background fact extraction (sync)
Once the model finishes, the plugin sends the user message and assistant response to Mem0 in a background thread for fact extraction. Each write includes the agent identifier and gateway channel.
<Note>
Automatic capture truncates each message to **450 characters by default in every mode**, preferring a sentence boundary. Adjust `sync_max_chars` for your model's context limit. Capture is best effort: if the previous sync is still running after a five-second wait, the next turn is skipped. Use `mem0_add` to store specific text verbatim.
</Note>
Once the model finishes, Hermes sends the `(user message, assistant response)` pair to Mem0 in a background thread. Mem0 extracts facts automatically (for example, "user prefers Python" or "user works at Acme Corp"), so you never have to tell it what to remember. Each write is tagged with the gateway channel it came from.
## Agent Tools
@@ -33,33 +29,21 @@ When Mem0 is active, the model gets four tools it can call during a conversation
| Tool | Description | Parameters |
|------|-------------|------------|
| `mem0_search` | Semantic search by meaning, ranked by relevance | `query` (required), `top_k` (default 10, max 50), `rerank` (uses the configured default, Platform mode only) |
| `mem0_search` | Semantic search by meaning, ranked by relevance | `query` (required), `top_k` (default 10, max 50), `rerank` (default `false`, Platform mode only) |
| `mem0_add` | Store a fact verbatim, with no LLM extraction | `content` (required) |
| `mem0_update` | Update a memory's text by ID | `memory_id`, `text` (both required) |
| `mem0_delete` | Delete a memory by ID | `memory_id` (required) |
## Installation
Install [Hermes Agent](https://github.com/NousResearch/hermes-agent) with memory-provider plugin support and Python 3.11 or later. Once the plugin directory is available on Mem0's main branch, install it from the repository subdirectory:
Install Hermes Agent:
```bash
hermes plugins install mem0ai/mem0/integrations/hermes-plugin-mem0
hermes plugins enable mem0
hermes memory setup
hermes memory status
curl -fsSL https://raw.githubusercontent.com/NousResearch/hermes-agent/main/scripts/install.sh | bash
source ~/.bashrc
```
Select **mem0** in setup and choose one of the modes below. Start a fresh Hermes conversation after setup.
Hermes installers with plugin dependency support install `mem0ai>=2.0.10,<3` and `httpx>=0.27,<1` from the plugin's `pyproject.toml`. Older hosts such as Hermes v0.21.3 require those packages to be installed explicitly into the Hermes Python environment. The OSS setup flow installs additional provider packages as needed.
<Note>
Hermes versions that still bundle Mem0 prefer the bundled provider. Use a Hermes release that has completed the standalone-provider migration; installing this plugin alone does not replace the bundled implementation. Existing users should keep their current configuration; see [Migration for existing users](#migration-for-existing-users).
</Note>
<Note>
Run the setup wizard in an interactive terminal. On Hermes hosts whose `hermes memory setup --help` lists only a provider argument, options such as `--mode`, `--host`, and `--oss-llm` are rejected by Hermes before the plugin runs. Use `hermes memory setup mem0` or the manual configuration below. Redirected input cannot select the mode picker and falls back to Platform.
</Note>
The `mem0ai` package is installed automatically when you enable the Mem0 provider, so there is no manual pip step. OSS providers may need extra packages (for example `qdrant-client`, `psycopg2-binary`, or `ollama`), which the setup flow installs for you when you pick them.
## Platform Setup
@@ -68,10 +52,10 @@ Platform mode uses managed Mem0 Cloud and is the fastest way to start.
### Option 1: Interactive wizard (recommended)
```bash
hermes memory setup mem0
hermes memory setup
```
Choose **Platform** and paste your API key when prompted. The wizard writes settings to `$HERMES_HOME/mem0.json` and keeps the key in that profile's `.env`. The default Hermes home is `~/.hermes`; named profiles use their own home directory.
Select **mem0**, choose **Platform**, and paste your API key when prompted. The wizard writes the non-secret settings to `~/.hermes/mem0.json` and keeps the key in `~/.hermes/.env`.
<Note>Get your API key from <a href="https://app.mem0.ai?utm_source=oss&utm_medium=integration-hermes">app.mem0.ai</a>.</Note>
@@ -79,25 +63,17 @@ Choose **Platform** and paste your API key when prompted. The wizard writes sett
```bash
hermes config set memory.provider mem0
echo "MEM0_API_KEY=your-api-key" >> ~/.hermes/.env
```
Add your key to the active Hermes profile's `.env`:
Then in your `config.yaml`:
```dotenv
MEM0_API_KEY=your-api-key
```yaml
memory:
provider: mem0
```
Set these values in the active profile's `mem0.json`, choosing a stable user identity:
```json
{
"mode": "platform",
"host": "",
"user_id": "my-hermes-user"
}
```
Remove any stale `MEM0_HOST` from the environment and profile `.env`, and remove an old inline `api_key` from `mem0.json` so the new `.env` key is used. The config command sets `memory.provider: mem0` in that profile's `config.yaml`. Restart Hermes and check `hermes memory status`.
That's it. Mem0 runs automatically from here.
## Self-Hosted Server Setup
@@ -106,47 +82,48 @@ Run the [Mem0 server](https://github.com/mem0ai/mem0/tree/main/server) (FastAPI
### Interactive
```bash
hermes memory setup mem0
# Choose "Self-hosted server", then enter the server URL and API key
hermes memory setup
# Select "mem0", then "Self-hosted server", and enter the server URL
```
### Manual configuration
### With flags
Select `mem0` with `hermes config set memory.provider mem0`. Set these values in the active profile's `mem0.json`:
```json
{
"mode": "platform",
"host": "http://localhost:8888",
"user_id": "my-hermes-user"
}
```bash
hermes memory setup mem0 --mode selfhosted \
--host http://localhost:8888 \
--api-key your-admin-api-key
```
Add the server key to that profile's `.env`:
### With environment variables
```dotenv
MEM0_API_KEY=your-admin-api-key
```bash
echo "MEM0_HOST=http://localhost:8888" >> ~/.hermes/.env
echo "MEM0_API_KEY=your-admin-api-key" >> ~/.hermes/.env
```
Remove an old inline `api_key` from `mem0.json` so the `.env` key is used. `MEM0_HOST` can also supply the server URL, but a non-empty `host` in `mem0.json` overrides it. Keep `mode` set to `platform` for the HTTP server backend.
Then start a fresh Hermes session and call `mem0_search` — it connects to your server. The plugin authenticates with `X-API-Key` and uses the server's `/search` and `/memories` routes. The API key is optional only for servers running with `AUTH_DISABLED`.
<Note>Setting `host` routes to the self-hosted server automatically. Don't combine it with `mode: oss` — OSS takes precedence and ignores `host`.</Note>
## OSS (Self-Hosted) Setup
OSS mode runs the Mem0 SDK in the Hermes process with your chosen LLM, embedder, and vector store. It does not use Mem0 Cloud or require a Mem0 API key. Data goes to the model services you configure; use local Ollama models and local storage for a fully local setup.
OSS mode runs Mem0 entirely on your own infrastructure: your LLM, your embedder, and your vector store. No data is sent to Mem0 Cloud, and no Mem0 API key is required.
### Interactive
```bash
hermes memory setup mem0
# Choose "Open Source"
hermes memory setup
# Select "mem0", then "Open Source (self-hosted)"
# Follow the prompts for LLM, embedder, and vector store
```
The wizard uses the listed default OpenAI models and local Qdrant storage. For custom OpenAI-compatible endpoints, deployment names, or a Qdrant server, use manual configuration below.
### With flags
```bash
hermes memory setup mem0 --mode oss \
--oss-llm openai --oss-llm-key sk-... \
--oss-vector qdrant
```
### Supported providers
@@ -156,24 +133,52 @@ The wizard uses the listed default OpenAI models and local Qdrant storage. For c
| Embedder | `openai` (default `text-embedding-3-small`), `ollama` (local, default `nomic-embed-text`) |
| Vector store | `qdrant` (local path or server), `pgvector` |
### Manual configuration
### Flag reference
| Flag | Description |
|------|-------------|
| `--mode` | `platform`, `selfhosted`, or `oss` |
| `--api-key` | Platform API key, or the admin key of a self-hosted server |
| `--host` | Self-hosted server URL (with `--mode selfhosted`) |
| `--oss-llm` | LLM provider (`openai` or `ollama`, default `openai`) |
| `--oss-llm-key` | LLM API key (for `openai`) |
| `--oss-llm-model` | Override the LLM model |
| `--oss-llm-url` | LLM base URL (for `ollama` or a custom endpoint) |
| `--oss-embedder` | Embedder provider (default `openai`) |
| `--oss-embedder-key` | Embedder API key |
| `--oss-embedder-model` | Override the embedder model |
| `--oss-embedder-url` | Embedder base URL (for `ollama` or a custom endpoint) |
| `--oss-vector` | Vector store (`qdrant` or `pgvector`, default `qdrant`) |
| `--oss-vector-path` | Local Qdrant storage path |
| `--oss-vector-url` | Qdrant server URL |
| `--oss-vector-host`, `--oss-vector-port` | PGVector or remote Qdrant host and port |
| `--oss-vector-user`, `--oss-vector-password`, `--oss-vector-dbname` | PGVector connection details |
| `--user-id` | Canonical user identifier |
| `--dry-run` | Preview the resolved config without writing it |
## Switching Modes
You can move between the three modes at any time. Run the setup command again, or edit `~/.hermes/mem0.json` directly.
```bash
hermes config set memory.provider mem0
# Platform to OSS
hermes memory setup mem0 --mode oss --oss-llm-key sk-...
# OSS to Platform
hermes memory setup mem0 --mode platform --api-key sk-...
# Platform to a self-hosted server
hermes memory setup mem0 --mode selfhosted --host http://localhost:8888
# Preview without writing anything
hermes memory setup mem0 --mode oss --oss-llm-key sk-... --dry-run
```
Add the model key to the active profile's `.env`:
```dotenv
OPENAI_API_KEY=your-model-api-key
```
Set the following in that profile's `mem0.json`. Use your existing storage path when migrating; for a new named profile, choose a path inside that profile's home.
A self-hosted `~/.hermes/mem0.json` looks like this:
```json
{
"mode": "oss",
"user_id": "my-hermes-user",
"oss": {
"llm": {"provider": "openai", "config": {"model": "gpt-5-mini", "is_reasoning_model": true}},
"embedder": {"provider": "openai", "config": {"model": "text-embedding-3-small"}},
@@ -182,65 +187,33 @@ Set the following in that profile's `mem0.json`. Use your existing storage path
}
```
For an OpenAI-compatible service such as Azure's `/openai/v1` endpoint, add `OPENAI_BASE_URL` to the profile's `.env` and set each `model` to its deployed name. Both the LLM and embedder use this endpoint unless their `config.openai_base_url` overrides it. The main Hermes chat model is configured separately; this JSON configures Mem0's extraction and embedding models.
For a Qdrant server, replace `vector_store.config.path` with `url`, for example `"url": "http://localhost:6333"`. Manual setup does not install optional provider dependencies: install `qdrant-client`, `psycopg2-binary`, or `ollama` in the **Hermes Python environment** as needed for your selected providers. Start a fresh session and verify a memory write and search; `hermes memory status` reports configuration availability, not a full backend health check.
Desktop sessions in the same process and profile share local Qdrant storage when their OSS settings match. Operations are serialized, and storage closes after the last session releases it. Conflicting settings are rejected without changing existing memories; close active sessions before changing models or credentials. For concurrent CLI and Desktop processes, use a Qdrant server or the self-hosted Mem0 HTTP API instead of sharing a local directory.
## Switching Modes
Run `hermes memory setup mem0` in an interactive terminal and choose the new mode, or edit the active profile's `mem0.json` using the examples above. Switching backends does not transfer memories between them. Preserve existing OSS storage paths when editing configuration. When returning to Platform, set `mode` to `platform`, clear `host`, and remove any stale `MEM0_HOST` setting from your environment and profile `.env`.
## Configuration
Settings live in `$HERMES_HOME/mem0.json` and are written by `hermes memory setup`. API keys normally live in that profile's `.env`; distinct OpenAI LLM/embedder keys and database credentials are stored in the OSS configuration. Setup writes these files atomically with owner-only permissions.
When editing these files manually, restrict both `.env` and `mem0.json` to their owner (`chmod 600` on Unix). Keep configuration and secrets in the same active Hermes profile.
`MEM0_MODE`, `MEM0_HOST`, `MEM0_USER_ID`, and `MEM0_AGENT_ID` supply environment defaults. Non-empty values in `mem0.json` take precedence. `MEM0_API_KEY` supplies the Cloud or server key unless `api_key` is set in the file.
Behavioral settings live in `~/.hermes/mem0.json` and are written for you by `hermes memory setup`. Only the secret `MEM0_API_KEY` belongs in `~/.hermes/.env`.
| Key | Default | Description |
|-----|---------|-------------|
| `mode` | `platform` | `platform` (Mem0 Cloud) or `oss` (self-managed, in-process). Self-hosted server routing is set via `host` |
| `host` | none | Self-hosted Mem0 server URL. When set, the plugin talks HTTP to your server instead of the cloud |
| `api_key` | none | Mem0 Platform API key, or the admin key of a self-hosted server. Stored in `.env` as `MEM0_API_KEY` |
| `user_id` | gateway user ID, then `hermes-user` | Identifier that scopes memories. See cross-channel behavior below |
| `user_id` | `hermes-user` | Identifier that scopes memories. See cross-channel behavior below |
| `agent_id` | `hermes` | Agent identifier attached to writes |
| `rerank` | `false` | Platform reranking for recall and tool searches that omit `rerank` |
| `sync_max_chars` | `450` | Per-message character cap for automatic fact extraction in every mode |
| `oss` | `{}` | OSS LLM, embedder, and vector-store configuration |
| `rerank` | `false` | Rerank search results for relevance (Platform mode only) |
### Cross-channel memories
Hermes can run from the CLI and from gateways like Telegram, Slack, and Discord. The `user_id` setting controls how memories are scoped across them:
- **Set a `user_id` other than `hermes-user`** and it applies to every gateway, so one person gets a single merged memory store no matter where they talk to the agent.
- **Leave it unset** (or at the default `hermes-user`) and each gateway uses its own native ID when available, falling back to `hermes-user`.
- **Set a `user_id`** and it applies to every gateway, so one person gets a single merged memory store no matter where they talk to the agent.
- **Leave it unset** (or at the default `hermes-user`) and each gateway uses its own native id, keeping per-platform memories separate.
Every write is tagged with `metadata.channel` (for example `telegram` or `cli`). Plugin searches filter by user identity across sessions; they do not restrict recall to the current channel or session.
## Migration for Existing Users
Keep `memory.provider: mem0`, `mem0.json`, `MEM0_*` settings, user identity, and OSS database paths unchanged. Moving from the bundled provider to this standalone plugin does not require rerunning setup or moving stored memories.
Automatic migration depends on Hermes rollout as well as this repository:
1. Users need a Hermes build containing [PR #114569](https://github.com/NousResearch/hermes-agent/pull/114569).
2. Hermes maintainers must approve a catalog entry named `mem0` with `repo: https://github.com/mem0ai/mem0`, `subdir: integrations/hermes-plugin-mem0`, and a reviewed full commit SHA.
3. The bundled Mem0 provider must be removed so the standalone provider can load.
With these in place, Hermes installs a missing configured provider during `hermes update` across profiles or at agent startup. Startup installation respects `security.allow_lazy_installs`. Offline or disabled installation needs manual action; merging the plugin directory alone does not complete automatic migration.
CLI setup and status are supported. This plugin does not ship a Desktop configuration panel or provider-specific CLI commands.
Either way, every write is tagged with `metadata.channel` (for example `telegram` or `cli`), so per-channel views are still possible at query time.
## Reliability
- **Circuit breaker**: five consecutive backend failures pause calls for two minutes. The agent can continue without memory during that window. Expected update/delete errors such as a missing memory do not trip the breaker.
- **Bounded waits**: recall waits up to three seconds. Capture runs in the background, but an overlapping turn may wait up to five seconds for the previous sync before being skipped.
- **Graceful shutdown**: shutdown and Python process exit wait for active recall and capture workers before closing the backend. Backend network timeouts still apply. Self-hosted HTTP capture uses a 120-second read timeout and a 30-second connection timeout; other self-hosted HTTP operations use 30 seconds.
- **Best-effort capture**: there is no durable queue. Forced termination, including Hermes' 30-second exit watchdog, can interrupt pending writes even while graceful shutdown is waiting.
- **OSS data protection**: an embedding dimension mismatch fails initialization without deleting the existing collection or table.
- **Circuit breaker**: if Mem0 fails five times in a row, Hermes pauses calls for two minutes, then retries. The agent keeps working without memory during that window. Expected client errors, like a 404 on a missing memory id, do not count toward tripping the breaker.
- **Non-blocking**: fact extraction runs in a background daemon thread, and current-turn recall waits at most 3 seconds, so a slow or failed call never blocks your conversation.
- **Thread-safe**: the client uses lazy initialization with locking, and the background sync and recall threads are guarded so concurrent gateway messages cannot produce duplicate memories.
## Troubleshooting
@@ -275,12 +248,15 @@ curl http://localhost:11434/api/tags
- `mem0_add` stores text verbatim with no extraction. Ordinary conversation turns are extracted automatically by the background sync.
- Search is semantic, so try a broader query.
- Confirm `user_id` is the same across sessions (check `$HERMES_HOME/mem0.json`).
- Check `sync_max_chars`: facts beyond the per-message limit are not sent for extraction.
- Confirm `user_id` is the same across sessions (check `~/.hermes/mem0.json`).
### OSS: embedding dimension mismatch
## Key Features
Restore the embedding model and dimensions that created the existing collection, or choose a new collection and migrate data explicitly. The plugin leaves the original collection intact when dimensions differ.
1. **Three ways to run**: managed Platform, a self-hosted server, or fully local OSS, switchable at any time.
2. **Current-turn recall**: memories for the current question are injected within a 3-second window, with `mem0_search` as the model's own backstop.
3. **Automatic extraction**: Mem0 extracts and deduplicates facts from each exchange for you.
4. **Non-blocking and fault tolerant**: background threads plus a circuit breaker keep the agent responsive even when Mem0 is unreachable.
5. **Additive memory**: works alongside Hermes' built-in file memory (`MEMORY.md`, `USER.md`).
<CardGroup cols={2}>
<Card title="OpenClaw Integration" icon="/images/provider-icons/openclaw.svg" href="/integrations/openclaw">
+1 -1
View File
@@ -45,7 +45,7 @@ memory_from_client = Mem0Memory.from_client(
)
```
Context is used to identify the user, agent or the conversation in the Mem0. It is required to be passed in at least one of the fields in the `Mem0Memory` constructor. It can be any of the following:
Context is used to identify the user, agent or the conversation in the Mem0. It is required to be passed in the at least one of the fields in the `Mem0Memory` constructor. It can be any of the following:
```python
context = {
+2 -9
View File
@@ -186,6 +186,7 @@ If the user is on a pre-current major (Python < 2, TS < 3, or a Platform call st
## Core Concepts
- [How Mem0 Works](https://docs.mem0.ai/core-concepts/how-it-works) [Both]: Use when explaining the end-to-end pipeline: extraction (ADD-only distillation), storage across vector/entity/history stores, and multi-signal retrieval.
- [Memory Types](https://docs.mem0.ai/core-concepts/memory-types) [Both]: Use when checking which `memory_type` values actually work: `procedural_memory` is implemented, `semantic_memory` and `episodic_memory` are defined in the enum but rejected by validation.
- [Memory Operations - Add](https://docs.mem0.ai/core-concepts/memory-operations/add) [Both]: Use when explaining how `add()` extracts facts, resolves conflicts, and writes to both stores.
- [Memory Operations - Search](https://docs.mem0.ai/core-concepts/memory-operations/search) [Both]: Use when explaining how queries are processed and ranked.
- [Memory Operations - Update](https://docs.mem0.ai/core-concepts/memory-operations/update) [Both]: Use when memories need to be edited in place or reconciled against new info.
@@ -197,7 +198,6 @@ If the user is on a pre-current major (Python < 2, TS < 3, or a Platform call st
### Features - Essential
- [V2 Memory Filters](https://docs.mem0.ai/platform/features/v2-memory-filters) [Platform]: Use when compound filters (AND/OR on metadata, entity, time) are needed at search.
- [Entity-Scoped Memory](https://docs.mem0.ai/platform/features/entity-scoped-memory) [Platform]: Use when partitioning memories by user, agent, app, or run.
- [Profiles](https://docs.mem0.ai/platform/features/user-profiles) [Platform]: Use when a structured always-current summary of a user is needed in one read, instead of searching their memories.
- [Graph Memory](https://docs.mem0.ai/platform/features/graph-memory) [Platform]: Use when connecting facts across memories through shared entities for entity-centric or multi-hop questions.
- [Async Client](https://docs.mem0.ai/platform/features/async-client) [Platform]: Use when the app issues many concurrent Mem0 calls and needs non-blocking I/O.
- [Multimodal Support](https://docs.mem0.ai/platform/features/multimodal-support) [Platform]: Use when storing images or PDFs as memory input.
@@ -256,7 +256,7 @@ If the user is on a pre-current major (Python < 2, TS < 3, or a Platform call st
- [Agno](https://docs.mem0.ai/integrations/agno) [Platform]: Use when the user is on Agno.
- [Camel AI](https://docs.mem0.ai/integrations/camel-ai) [Both]: Use when the user is on Camel AI.
- [ChatDev](https://docs.mem0.ai/integrations/chatdev) [Platform]: Use when the user is on ChatDev.
- [Hermes](https://docs.mem0.ai/integrations/hermes) [Both]: Use when installing or configuring the standalone Hermes memory plugin, or migrating from the bundled Mem0 provider.
- [Hermes](https://docs.mem0.ai/integrations/hermes) [Both]: Use when the user is on Hermes.
- [Pi Agent](https://docs.mem0.ai/integrations/pi-agent) [Platform]: Use when adding automatic capture, prompt recall, scoped memory, and six memory commands to Pi Agent.
- [DeepSeek Harness](https://docs.mem0.ai/integrations/deepseek-plugin) [Platform]: Use when adding automatic recall, completed-turn capture, and native search/add tools to DeepSeek Harness.
- [OpenAI Agents SDK](https://docs.mem0.ai/integrations/openai-agents-sdk) [Platform]: Use when the user is on the OpenAI Agents SDK.
@@ -268,7 +268,6 @@ If the user is on a pre-current major (Python < 2, TS < 3, or a Platform call st
- [Strands Agents](https://docs.mem0.ai/integrations/strands) [Both]: Use when the user is on AWS Strands and wants a native MemoryStore.
### AI Coding Tools
- [Agent Plugins Overview](https://docs.mem0.ai/integrations/agent-plugins) [Platform]: Use when choosing a Mem0 plugin for a coding agent or agent harness, or comparing what each plugin captures and recalls.
- [Claude Code](https://docs.mem0.ai/integrations/claude-code) [Platform]: Use when wiring memory into Claude Code.
- [Claude.ai](https://docs.mem0.ai/integrations/claude-ai) [Platform]: Use when connecting Mem0 to Claude.ai (the hosted web app) via a custom remote MCP connector, or when Claude's native memory seems to be crowding out mem0 tool calls.
- [Cursor](https://docs.mem0.ai/integrations/cursor) [Platform]: Use when adding lifecycle capture, explicit memory recall, and six memory skills to Cursor.
@@ -327,7 +326,6 @@ If the user is on a pre-current major (Python < 2, TS < 3, or a Platform call st
- [Healthcare Google ADK](https://docs.mem0.ai/cookbooks/integrations/healthcare-google-adk) [Platform]: Use when the domain is medical and the framework is Google ADK.
- [AWS Bedrock](https://docs.mem0.ai/cookbooks/integrations/aws-bedrock) [OSS]: Use when deploying with AWS managed model services.
- [Tavily Search](https://docs.mem0.ai/cookbooks/integrations/tavily-search) [Platform]: Use when the agent layers web search on memory.
- [Company Brain (Mem0 Platform + Supabase)](https://docs.mem0.ai/cookbooks/integrations/supabase) [Platform]: Use to build a shared org brain on Mem0 Platform with Supabase as system of record and the MCP server as the access layer (with a new-hire onboarding demo).
### Framework Examples
- [LlamaIndex React](https://docs.mem0.ai/cookbooks/frameworks/llamaindex-react) [Both]: Use when building a React UI with LlamaIndex and memory.
@@ -365,11 +363,6 @@ All API Reference docs describe Mem0 Platform REST endpoints (requires API key).
### Entities
- [Get Users](https://docs.mem0.ai/api-reference/entities/get-users) [Platform]: Use when listing users, agents, or apps known to a project.
- [Delete User](https://docs.mem0.ai/api-reference/entities/delete-user) [Platform]: Use when removing an entity and all its memories.
- [Get Profile](https://docs.mem0.ai/api-reference/profiles/get-profile) [Platform]: Use when reading a user's structured profile and branching on its generation status.
- [Get Profile Settings](https://docs.mem0.ai/api-reference/profiles/get-profile-settings) [Platform]: Use when checking the project's profile schema, instructions, or enabled flag.
- [Update Profile Settings](https://docs.mem0.ai/api-reference/profiles/update-profile-settings) [Platform]: Use when defining or changing the JSON Schema that shapes profiles for a project.
- [Generate Profiles](https://docs.mem0.ai/api-reference/profiles/generate-profiles) [Platform]: Use when building profiles now: a sample of ten, or one entity.
- [Get Generation Job](https://docs.mem0.ai/api-reference/profiles/get-profile-job) [Platform]: Use when checking how far a generation has got, and whether it finished.
### Organizations
- [Create Organization](https://docs.mem0.ai/api-reference/organization/create-org) [Platform]: Use when setting up a new org.
+1 -475
View File
@@ -8070,480 +8070,6 @@
}
}
}
},
"/v2/entities/{entity_type}/{entity_id}/profile/": {
"get": {
"tags": [
"profiles"
],
"operationId": "profiles_read",
"summary": "Get an entity's profile",
"description": "Return the memory profile for one user.\n\nGeneration is asynchronous, so a known entity that has no profile yet is a normal 200 carrying a `status`. A 404 means only that no such entity exists.",
"parameters": [
{
"name": "entity_type",
"in": "path",
"required": true,
"schema": {
"type": "string",
"enum": [
"user"
]
},
"description": "The kind of entity that carries the profile."
},
{
"name": "entity_id",
"in": "path",
"required": true,
"schema": {
"type": "string"
},
"description": "The entity's id, as supplied when the memory was added."
}
],
"responses": {
"200": {
"description": "The profile envelope.",
"content": {
"application/json": {
"schema": {
"type": "object",
"properties": {
"profile": {
"type": "object",
"additionalProperties": true,
"description": "The generated profile, shaped by the project's schema. Empty unless status is succeeded."
},
"status": {
"type": "string",
"enum": [
"succeeded",
"pending",
"failed",
"not_enabled",
"insufficient_data"
],
"description": "Generation state. Branch on this rather than on an empty profile."
},
"entity_type": {
"type": "string",
"enum": [
"user"
]
},
"entity_id": {
"type": "string"
},
"updated_at": {
"type": "string",
"format": "date-time",
"nullable": true
},
"generation_count": {
"type": "integer"
}
}
}
}
}
},
"400": {
"description": "Unsupported entity type."
},
"404": {
"description": "No such entity in this project."
}
}
}
},
"/v2/profiles/settings/": {
"get": {
"tags": [
"profiles"
],
"operationId": "profiles_settings_read",
"summary": "Get profile settings",
"description": "Return the profile settings for the project the API key is scoped to.",
"responses": {
"200": {
"description": "Current settings.",
"content": {
"application/json": {
"schema": {
"type": "object",
"properties": {
"enabled": {
"type": "boolean",
"description": "Whether profile generation runs for this project. Project-wide."
},
"entities": {
"type": "object",
"description": "Settings for user profiles, under `user`.",
"properties": {
"user": {
"type": "object",
"properties": {
"schema": {
"type": "object",
"additionalProperties": true,
"nullable": true,
"description": "JSON Schema describing the profile. Every property needs a description."
},
"custom_instructions": {
"type": "string",
"nullable": true,
"description": "Extra guidance for the extraction step."
}
}
}
}
},
"capabilities": {
"type": "object",
"properties": {
"jobs": {
"type": "boolean"
},
"estimates": {
"type": "boolean"
},
"samples": {
"type": "boolean"
},
"full_rebuild": {
"type": "boolean",
"description": "Whether a project-wide rebuild (regenerate/backfill) is available. Currently false."
}
}
}
}
}
}
}
}
}
},
"post": {
"tags": [
"profiles"
],
"operationId": "profiles_settings_update",
"summary": "Update profile settings",
"description": "Update the project's profile settings. Only the fields present in the body are written, so one setting can change without re-sending the others.",
"requestBody": {
"required": true,
"content": {
"application/json": {
"schema": {
"type": "object",
"description": "Only the fields present are written. `schema` and `custom_instructions` nest under `entities.user`; a flat body is rejected.",
"properties": {
"enabled": {
"type": "boolean",
"description": "Whether profile generation runs for this project. Project-wide."
},
"entities": {
"type": "object",
"description": "Settings for user profiles, under `user`.",
"properties": {
"user": {
"type": "object",
"properties": {
"schema": {
"type": "object",
"additionalProperties": true,
"nullable": true,
"description": "JSON Schema describing the profile. Every property needs a description. Send null to clear it."
},
"custom_instructions": {
"type": "string",
"nullable": true,
"description": "Extra guidance for the extraction step. Send null to clear it."
}
}
}
}
}
}
}
}
}
},
"responses": {
"200": {
"description": "Settings as stored after the update.",
"content": {
"application/json": {
"schema": {
"type": "object",
"properties": {
"enabled": {
"type": "boolean",
"description": "Whether profile generation runs for this project. Project-wide."
},
"entities": {
"type": "object",
"description": "Settings for user profiles, under `user`.",
"properties": {
"user": {
"type": "object",
"properties": {
"schema": {
"type": "object",
"additionalProperties": true,
"nullable": true,
"description": "JSON Schema describing the profile. Every property needs a description."
},
"custom_instructions": {
"type": "string",
"nullable": true,
"description": "Extra guidance for the extraction step."
}
}
}
}
},
"capabilities": {
"type": "object",
"properties": {
"jobs": {
"type": "boolean"
},
"estimates": {
"type": "boolean"
},
"samples": {
"type": "boolean"
},
"full_rebuild": {
"type": "boolean",
"description": "Whether a project-wide rebuild (regenerate/backfill) is available. Currently false."
}
}
}
}
}
}
}
},
"400": {
"description": "The schema is not a valid profile schema."
}
}
}
},
"/v2/profiles/jobs/": {
"post": {
"tags": [
"profiles"
],
"operationId": "profiles_create_job",
"summary": "Generate profiles",
"description": "Start one generation. `operation` says what to build:\n\n- `sample` — up to 10 real entities, so a schema can be judged before it is used widely. These are real profiles: they are saved to those entities and count toward usage.\n- `trigger` — one entity, named by `entity_id`.\n\nSend an `Idempotency-Key` header. Replaying the same key returns the same job instead of charging twice. Poll `status_url` from the response until the status is terminal.",
"parameters": [
{
"in": "header",
"name": "Idempotency-Key",
"required": true,
"schema": {
"type": "string",
"minLength": 8,
"maxLength": 128
},
"description": "Makes a retry safe: the same key returns the same job."
}
],
"requestBody": {
"required": true,
"content": {
"application/json": {
"schema": {
"type": "object",
"required": [
"operation",
"entity_type"
],
"properties": {
"operation": {
"type": "string",
"enum": [
"sample",
"trigger"
],
"description": "What to generate. Optional only when `entity_id` is set, which means `trigger`."
},
"entity_type": {
"type": "string",
"enum": [
"user"
]
},
"entity_id": {
"type": "string",
"description": "One entity, for `trigger`."
},
"limit": {
"type": "integer",
"minimum": 1,
"maximum": 10,
"description": "How many entities to sample."
}
}
}
}
}
},
"responses": {
"202": {
"description": "Job accepted.",
"content": {
"application/json": {
"schema": {
"type": "object",
"properties": {
"job_id": {
"type": "string"
},
"status": {
"type": "string"
},
"status_url": {
"type": "string",
"description": "Poll this. Building the path yourself breaks on a route change."
},
"operation": {
"type": "string"
},
"entity_type": {
"type": "string"
},
"entity_count_reserved": {
"type": "integer",
"description": "Entities reserved against usage for this job."
},
"event_id": {
"type": "string",
"nullable": true
},
"replayed": {
"type": "boolean",
"description": "True when an Idempotency-Key returned an existing job."
},
"sampled": {
"type": "integer",
"description": "`sample` only."
},
"entity_ids": {
"type": "array",
"items": {
"type": "string"
},
"description": "`sample` only: the entity ids picked. Read each with `GET /v2/entities/user/{entity_id}/profile/`."
}
}
}
}
}
},
"400": {
"description": "Unknown or missing `operation`, or profiles are not configured."
},
"402": {
"description": "Payment required."
},
"409": {
"description": "A job is already running, or the Idempotency-Key was used for a different request. Branch on `error.code`."
},
"429": {
"description": "Cooldown. `retry_after_seconds` sits inside `error`."
},
"503": {
"description": "`jobs_unavailable` — generation is switched off for this project."
}
}
}
},
"/v2/profiles/jobs/{job_id}/": {
"get": {
"tags": [
"profiles"
],
"operationId": "profiles_get_job",
"summary": "Read a generation job",
"description": "The job nests under `job`. `total` is null until `enumeration_complete`, and `completed` is `succeeded + failed + skipped`.",
"parameters": [
{
"in": "path",
"name": "job_id",
"required": true,
"schema": {
"type": "string"
}
}
],
"responses": {
"200": {
"description": "The job.",
"content": {
"application/json": {
"schema": {
"type": "object",
"properties": {
"job": {
"type": "object",
"properties": {
"id": {
"type": "string"
},
"operation": {
"type": "string"
},
"entity_type": {
"type": "string"
},
"status": {
"type": "string",
"enum": [
"QUEUED",
"RUNNING",
"SUCCEEDED",
"PARTIALLY_SUCCEEDED",
"FAILED",
"CANCELLED"
]
},
"total": {
"type": "integer",
"nullable": true
},
"enumeration_complete": {
"type": "boolean"
},
"completed": {
"type": "integer"
},
"succeeded": {
"type": "integer"
},
"failed": {
"type": "integer"
},
"skipped": {
"type": "integer"
}
}
}
}
}
}
}
},
"404": {
"description": "No such job in this project."
}
}
}
}
},
"components": {
@@ -9462,4 +8988,4 @@
}
},
"x-original-swagger-version": "2.0"
}
}
-359
View File
@@ -1,359 +0,0 @@
---
title: Profiles
description: "Build a structured, always-current summary of each user from their memories, shaped by a JSON Schema you define."
---
# Profiles
Memories are individual facts. A profile is the summary of all of them for one entity: a single structured object, shaped by a JSON Schema you define, that Mem0 keeps current as new memories arrive.
Search answers "what did this user say about X". A profile answers "who is this user", in one read, with no query to write.
<Info>
**Use profiles when…**
- You want to personalize a first response, before the user says anything in this session.
- You need a compact object to drop into a prompt instead of a list of memories.
- You want the same fields for every user, so your code can rely on their shape.
</Info>
<Note>
User Profiles are in **beta** and available on request. To enable them for your
organization, contact [support@mem0.ai](mailto:support@mem0.ai).
</Note>
## How it works
1. You define a **schema**: the fields a profile should contain, each with a description.
2. Mem0 builds each entity's profile from their memories, and rebuilds it as new memories arrive.
3. You read the profile whenever you need it.
Generation is **asynchronous**. A profile is not ready the instant an entity's first memory lands, so a read tells you where it is with a `status` rather than failing.
## Define the schema
The schema is JSON Schema. Every property needs a `description` — that is what tells the model how to fill the field, so a vague description gives a vague profile.
<CodeGroup>
```python Python
from mem0 import MemoryClient
client = MemoryClient()
client.update_profile_settings(
enabled=True,
schema={
"type": "object",
"properties": {
"communication_style": {
"type": "string",
"description": "How the user prefers to be addressed: terse, detailed, formal, casual",
},
"expertise_areas": {
"type": "array",
"items": {"type": "string"},
"description": "Subjects the user demonstrates working knowledge of",
},
"current_goals": {
"type": "array",
"items": {"type": "string"},
"description": "What the user is actively trying to accomplish",
},
},
},
custom_instructions="Prefer durable traits over one-off remarks.",
)
```
```typescript TypeScript
import MemoryClient from "mem0ai";
const client = new MemoryClient({ apiKey: "your-api-key" });
await client.updateProfileSettings({
enabled: true,
schema: {
type: "object",
properties: {
communication_style: {
type: "string",
description:
"How the user prefers to be addressed: terse, detailed, formal, casual",
},
expertise_areas: {
type: "array",
items: { type: "string" },
description: "Subjects the user demonstrates working knowledge of",
},
current_goals: {
type: "array",
items: { type: "string" },
description: "What the user is actively trying to accomplish",
},
},
},
customInstructions: "Prefer durable traits over one-off remarks.",
});
```
</CodeGroup>
<Note>
Your schema's property names reach the API exactly as you write them. The SDKs do not rewrite them, so a profile always comes back with the field names you chose.
</Note>
Only the fields you pass are written. To turn the feature off without touching your schema, send `enabled` alone.
## Read a profile
<CodeGroup>
```python Python
result = client.get_profile("alice")
if result["status"] == "succeeded":
print(result["profile"])
else:
print("not ready:", result["status"])
```
```typescript TypeScript
const result = await client.getProfile({ entityId: "alice" });
if (result.status === "succeeded") {
console.log(result.profile);
} else {
console.log("not ready:", result.status);
}
```
</CodeGroup>
A response looks like this:
```json
{
"profile": {
"communication_style": "terse",
"expertise_areas": ["distributed systems", "postgres"],
"current_goals": ["cut p99 latency", "migrate off the legacy queue"]
},
"status": "succeeded",
"entity_type": "user",
"entity_id": "alice",
"updated_at": "2026-02-08T10:30:00Z",
"generation_count": 3
}
```
`generation_count` is how many times this profile has been (re)generated — `0` before the first generation completes.
### Always branch on `status`
`profile` is empty unless `status` is `succeeded`. Check the status rather than the emptiness of the object, so a profile that is merely still building is not mistaken for a user you know nothing about.
| `status` | Meaning | What to do |
|---|---|---|
| `succeeded` | Profile is built and current | Use it |
| `pending` | Generation is queued or running | Read again shortly |
| `insufficient_data` | Not enough memories to say anything yet | Fall back to defaults |
| `not_enabled` | Profiles are off for this project | Enable them in settings |
| `failed` | The last generation did not complete | Retry, or trigger a new one |
A `404` means only that no such entity exists in your project.
## Generate a profile on demand
Profiles are built once an entity has accumulated enough messages, so a brand-new user has none during their first few interactions. Trigger one directly to close that gap:
<CodeGroup>
```python Python
client.generate_profile("alice")
```
```typescript TypeScript
await client.generateProfile({ entityId: "alice" });
```
</CodeGroup>
The call returns as soon as the work is queued. Poll the read endpoint and branch on `status`.
## Test a schema before applying it
A schema that reads well can still produce disappointing profiles. Sample a few real entities and inspect the output before committing to it.
Sampling is asynchronous: the call returns a job as soon as it is queued. Poll `status_url` until the job is terminal, then read each sampled entity's profile:
<CodeGroup>
```python Python
import time
job = client.sample_profiles(limit=5)
# Poll until the sample job reaches a terminal state (job status is UPPERCASE).
TERMINAL = {"SUCCEEDED", "PARTIALLY_SUCCEEDED", "FAILED", "CANCELLED"}
deadline = time.time() + 120
while True:
status = client.get_profile_job(job["status_url"])["job"]
if status["status"] in TERMINAL:
break
if time.time() > deadline:
raise TimeoutError("Sample job did not finish in time")
time.sleep(3)
print(status["status"], status["succeeded"], "of", status["total"])
# The create response lists the sampled entities; read each one's saved profile.
for entity_id in job.get("entity_ids", []):
print(client.get_profile(entity_id))
```
```typescript TypeScript
const job = await client.sampleProfiles({ limit: 5 });
// Poll until the sample job reaches a terminal state (job status is UPPERCASE).
const TERMINAL = ["SUCCEEDED", "PARTIALLY_SUCCEEDED", "FAILED", "CANCELLED"];
const deadline = Date.now() + 120_000;
let status;
while (true) {
status = (await client.getProfileJob(job.statusUrl)).job;
if (TERMINAL.includes(status.status)) break;
if (Date.now() > deadline)
throw new Error("Sample job did not finish in time");
await new Promise((resolve) => setTimeout(resolve, 3000));
}
console.log(status.status, status.succeeded, "of", status.total);
// The create response lists the sampled entities; read each one's saved profile.
for (const entityId of job.entityIds ?? []) {
console.log(await client.getProfile({ entityId }));
}
```
</CodeGroup>
These are real generations. The profiles are saved to those entities and count toward your usage, so sampling is not wasted work and not a free dry run. A sample covers up to 10 entities and cannot be repeated immediately.
## Apply a new schema to existing entities
A new schema shapes the next generation. Profiles that already exist keep their values until their entity is generated again.
Each entity picks the new schema up as it sends more memories, and you can generate one now with `generate_profile`.
<Note>
Rebuilding every profile in a project at once is not available yet. Refresh profiles one entity at a time with `generate_profile`, or let each one update on its own as its entity sends more memories.
</Note>
## When profiles update
You never call an "update profile" endpoint — Mem0 keeps each profile current for you. Two things drive it:
- **Automatically, as memories accumulate.** Mem0 refreshes an entity's profile after roughly every **10 messages** it receives, folding the new memories into the existing profile. There is no schedule to wait for and no extra call to make: the same `add` you already do keeps the profile moving.
- **On demand.** Call `generate_profile` to build or refresh a profile immediately — useful for a brand-new entity that has not yet crossed the automatic threshold.
Generation is **asynchronous and incremental**. A refresh runs in the background a short while after its trigger, so a read taken immediately after an `add` may still show the previous profile (or `pending`). Branch on `status` rather than assuming the latest memory is already reflected.
<Note>
Updates are **incremental**, not a full rebuild each time — Mem0 merges what it newly learns into the stored profile and keeps the fields your schema still defines. After a schema change, existing profiles pick it up as their entities send more memories, or when you call `generate_profile` — see [Apply a new schema to existing entities](#apply-a-new-schema-to-existing-entities).
</Note>
## Use a profile in a prompt
The point of the structure is that it drops straight into a prompt:
```python
result = client.get_profile(user_id)
if result["status"] == "succeeded":
profile = result["profile"]
system_prompt = f"""You are helping {user_id}.
Communication style: {profile.get("communication_style", "unknown")}
Areas of expertise: {", ".join(profile.get("expertise_areas", []))}
Current goals: {", ".join(profile.get("current_goals", []))}
Match their style and do not explain what they already know."""
else:
system_prompt = "You are a helpful assistant."
```
## Writing a schema that works
- **Describe every field.** The description is the instruction; without it the model guesses.
- **Prefer durable traits.** "Prefers dark mode" ages well; "is annoyed today" does not.
- **Keep it small.** Ten focused fields beat forty speculative ones, and cost less to generate.
- **Say what the field is not.** A description that rules out the near-miss interpretation is worth more than one that only states the obvious.
- **Sample before you commit.** It is the only way to see what your descriptions actually produce.
<Note>
A schema has a size budget of roughly **10,000 tokens** of serialized JSON — the whole schema is sent to the model on every generation, so a handful of verbose fields can cost more than many terse ones. Oversized schemas are rejected on save.
</Note>
## Availability
The feature is in beta and enabled per organization on request — see the note at the top of this page.
Once it is on, an entity gets a profile when two more things hold:
- profiles are **enabled** with a schema for the project (see [Define the schema](#define-the-schema)), and
- the memory is scoped to an entity — a `user_id`.
On a project where profiles are turned off, a read returns `status: not_enabled` rather than an error, so you can call it unconditionally and branch on the status.
## Settings reference
| Argument | Type | Description |
|---|---|---|
| `enabled` | boolean | Whether profile generation runs for the project |
| `schema` | object | JSON Schema describing the profile. Every property needs a `description` |
| `custom_instructions` | string | Extra guidance applied during extraction |
`enabled` is project-wide. `schema` and `custom_instructions` apply to user
profiles, so the stored settings nest them under `entities`:
```json
{
"enabled": true,
"entities": {
"user": {
"schema": { "type": "object", "properties": { "...": {} } },
"custom_instructions": "Prefer durable traits over one-off remarks."
}
},
"capabilities": { "full_rebuild": false }
}
```
That is what a read returns and what a write accepts. The SDKs take the fields
flat and nest them for you, so a schema you write with
`update_profile_settings` comes back unchanged from `get_profile_settings`.
<Note>
Profile settings are per project. An API key is scoped to one project, so profiles never cross a project boundary.
</Note>
## FAQ
**Do I need to change my `add` or `search` calls to use profiles?**
No. Profiles are built from the memories you already add. You define a schema once and read the profile when you need it — your ingestion and retrieval code is unchanged.
**Why is `profile` empty even though the entity has memories?**
Generation is asynchronous and needs enough to work with. Branch on `status`: `pending` means it is still building, and `insufficient_data` means there are not yet enough memories to fill the schema. Read again shortly, or call `generate_profile` to build one now.
**Is sampling free?**
No. `sample_profiles` runs real generations against real memories and **keeps** the profiles it produces, so it counts toward your usage like any other generation. It exists to check a schema on a few entities before you commit to it — not as a zero-cost dry run.
**Does changing the schema rewrite existing profiles?**
No. A schema change applies to the next generation. An existing profile keeps its values until its entity is generated again, which happens as that entity sends more memories, or when you call `generate_profile` for it.
**What happens to a field I remove from the schema?**
It stops being maintained. On an entity's next generation, fields your schema no longer defines are pruned from the stored profile — so keep a field in the schema for as long as you want its value kept.
**How current is a profile?**
It refreshes automatically as memories accumulate (about every 10 messages for an entity), plus any on-demand `generate_profile` calls. Because refreshes run in the background, expect a short delay after the triggering `add` rather than an instant update.
## Related
<CardGroup cols={2}>
<Card title="Entity-Scoped Memory" icon="users" href="/platform/features/entity-scoped-memory">
How users, agents, apps and runs partition memories.
</Card>
<Card title="Custom Instructions" icon="pen" href="/platform/features/custom-instructions">
Steer what Mem0 extracts in the first place.
</Card>
</CardGroup>
+2 -2
View File
@@ -50,8 +50,8 @@ For the full pipeline, see [How Mem0 works](/core-concepts/how-it-works).
<Card title="Run the quickstart" icon="rocket" href="/platform/quickstart">
Get an API key and save your first memory.
</Card>
<Card title="Scope your memories" icon="brain" href="/platform/features/entity-scoped-memory">
Organize memories by user, agent, app, and run.
<Card title="Understand memory types" icon="brain" href="/core-concepts/memory-types">
How user, agent, app, and run memory differ.
</Card>
<Card title="Add, search, and update" icon="layer-group" href="/core-concepts/memory-operations/add">
The core memory operations, end to end.
+1 -1
View File
@@ -39,7 +39,7 @@ The core memory loop is identical on both: `add`, `search`, `get`, `get_all`, `u
- **Entity scoping** by `user_id`, `agent_id`, and `run_id`
- **Filter grouping**: both accept `AND`/`OR`/`NOT` wrappers, both implicitly AND a flat multi-key filter like `{"user_id": "alice", "agent_id": "a1"}`, and both accept `*` as a wildcard value. Which fields you may filter on, and which operators each field accepts, differ (see below)
- **Entity-aware ranking**: both extract entities from memory text and use shared entities to boost related results at search time
- **Multimodal input**, **memory expiration** (`expiration_date`), **reranking**, and **custom extraction instructions** (`custom_instructions`)
- **Multimodal input**, **memory expiration** (`expiration_date`), **reranking**, **procedural memory** (Python), and **custom extraction instructions** (`custom_instructions`)
- Python and JavaScript SDKs, plus a REST API (self-hosted via `server/`, or hosted)
## What's actually different
+1 -1
View File
@@ -4,7 +4,7 @@ description: "Standard layout for documenting Mem0 API endpoints."
icon: "code"
---
# API Reference Template
# Api Reference Template
API reference pages document a single endpoint contract. Present metadata, request/response examples, and recovery guidance without narrative detours.
+1 -1
View File
@@ -124,7 +124,7 @@ Walk through a real request/response. Include sample payloads and highlight nota
{/* DEBUG: verify CTA targets */}
<CardGroup cols={2}>
<Card title="Dive Into Memory Scoring" icon="scale-balanced" href="/core-concepts/how-it-works">
<Card title="Dive Into Memory Scoring" icon="scale-balanced" href="/core-concepts/memory-types">
Understand how Mem0 ranks memories under the hood.
</Card>
<Card title="Build a Research Copilot" icon="book-open" href="/cookbooks/operations/deep-research">
+4 -2
View File
@@ -38,6 +38,7 @@ client.memories.add(
user_id: str,
memory: str,
metadata: Optional[dict] = None,
memory_type: Literal["session", "long_term"] = "session",
)
```
@@ -46,13 +47,13 @@ await mem0.memories.add({
userId: string;
memory: string;
metadata?: Record<string, string>;
memoryType?: "session" | "long_term";
});
```
</CodeGroup>
<Info>
[Describe defaults supported by this operation and SDK. Do not infer a
memory type or retention policy from a scoping identifier.]
Defaults to session memories. Override `memory_type` for long-term storage.
</Info>
<Warning>
@@ -66,6 +67,7 @@ await mem0.memories.add({
| `user_id` | string | Yes | Unique identifier for the end user. | Must match follow-up operations. |
| `memory` | string | Yes | Content to persist. | Managed & OSS. Markdown allowed. |
| `metadata` | object | No | Key-value pairs for filters. | OSS stores as JSONB; limit to 2KB. |
| `memory_type` | string | No | Retention bucket | Platform supports `shared`. |
<Tip>
Set `ttl_seconds` when you need memories to expire automatically (OSS only).
+2 -2
View File
@@ -96,8 +96,8 @@ time. Storage: vector embeddings.
**Architecture Overview:**
- Memory is scoped by user_id, agent_id, or run_id
- Core operations: add, search, update, delete
- Store preferences, facts, and past interactions; use run_id to scope a
session. Platform does not expose a memory_type selector.
- Memory types: factual (preferences, facts), episodic (past interactions),
semantic (concept relationships), working (session state)
- Integration pattern: retrieve relevant memories → generate response → store
new memories
+3 -3
View File
@@ -45,11 +45,11 @@ grok_client = OpenAI(
def recommend_movie_with_memory(user_id: str, user_query: str):
# Retrieve prior memory about movies
past_memories = memory.search("movie preferences", filters={"user_id": user_id})
past_memories = memory.search("movie preferences", user_id=user_id)
prompt = user_query
if past_memories["results"]:
prompt += f"\nPreviously, the user mentioned: {[m['memory'] for m in past_memories['results']]}"
if past_memories:
prompt += f"\nPreviously, the user mentioned: {past_memories}"
# Generate movie recommendation using Grok 3
response = grok_client.chat.completions.create(model="grok-3-beta", messages=[{"role": "user", "content": prompt}])
@@ -198,7 +198,7 @@ def search_memory_tool(query: str, user_id: str = "user") -> str:
Relevant vector memories found or message if none found
"""
try:
results = m.search(query, filters={"user_id": user_id})
results = m.search(query, user_id=user_id)
if isinstance(results, dict) and 'results' in results:
memory_list = results['results']
@@ -245,7 +245,7 @@ def search_graph_memory_tool(query: str, user_id: str = "user") -> str:
"""
try:
graph_query = f"relationships connections {query}"
results = m.search(graph_query, filters={"user_id": user_id})
results = m.search(graph_query, user_id=user_id)
if isinstance(results, dict) and 'results' in results:
memory_list = results['results']
@@ -290,7 +290,7 @@ def get_all_memories_tool(user_id: str = "user") -> str:
All memories for the user or message if none found
"""
try:
all_memories = m.get_all(filters={"user_id": user_id})
all_memories = m.get_all(user_id=user_id)
if isinstance(all_memories, dict) and 'results' in all_memories:
memory_list = all_memories['results']
+5 -5
View File
@@ -107,16 +107,16 @@ def main():
for query in search_queries:
print(f"\nQuery: {query}")
memories = memory.search(query=query, filters={"user_id": "user_123"})
memories = memory.search(query=query, user_id="user_123")
for memory_item in memories["results"]:
for memory_item in memories:
print(f" - {memory_item['memory']}")
print("\n--> Getting all memories for user...")
all_memories = memory.get_all(filters={"user_id": "user_123"})
print(f"Total memories stored: {len(all_memories['results'])}")
all_memories = memory.get_all(user_id="user_123")
print(f"Total memories stored: {len(all_memories)}")
for memory_item in all_memories["results"]:
for memory_item in all_memories:
print(f" - {memory_item['memory']}")
print("\n--> vLLM integration demo completed successfully!")
-764
View File
@@ -1,764 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# User Profiles — live demo\n",
"\n",
"A **profile** is a structured JSON document about ONE user, filled by an LLM from that\n",
"user's memories, shaped by a JSON Schema you supply.\n",
"\n",
"Search answers *\"what did this user say about X\"*. A profile answers *\"who is this\n",
"user\"*, in one read, with no query to write — and it is available on the first turn of a\n",
"session, before the user has said anything.\n",
"\n",
"**What this notebook does:** feed a user 12 conversation turns, watch a profile get\n",
"generated from them, add 6 more turns that contradict the first set, and watch the\n",
"profile rewrite itself. Then it shows every way the API says no.\n",
"\n",
"**You need:** an API key, and a project on the **Pro plan or higher**. Never commit one.\n",
"\n",
"> Set `MEM0_API_KEY`, and `MEM0_API_HOST` if you are pointing at a sandbox rather than\n",
"> production. The cells below read both from the environment.\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"> **Use a disposable project.** This notebook overwrites the project's profile settings\n",
"> (enabled, schema, custom instructions). The last cell restores the values saved at the\n",
"> start, but only if you reach it: if a cell fails midway, the project keeps the demo\n",
"> schema until you run the cleanup cell or reset it yourself. Do not point it at a\n",
"> project other people or production traffic depend on.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# This notebook drives the SDK from this worktree, not the published mem0ai:\n",
"# the profile fixes below are not released yet.\n",
"%pip install -q -e ../..\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": "import json\nimport os\nimport time\nimport uuid\n\nimport mem0\nfrom mem0 import MemoryClient\n\nAPI_KEY = os.environ.get(\"MEM0_API_KEY\")\nif not API_KEY:\n import getpass\n\n API_KEY = getpass.getpass(\"API key: \")\n\nclient = MemoryClient(api_key=API_KEY, host=os.environ.get(\"MEM0_API_HOST\") or None)\n\n# Fresh id each run, so nothing below is stale from a previous pass.\nUSER_ID = f\"demo_{uuid.uuid4().hex[:8]}\"\n\n# Snapshot the project's profile settings up front. This notebook overwrites the\n# shared project schema/instructions/enabled below; the cleanup cell restores this.\nORIGINAL_SETTINGS = client.get_profile_settings()\n\nprint(\"sdk :\", mem0.__file__) # must be this worktree\nprint(\"host :\", client.host)\nprint(\"demo user:\", USER_ID)"
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 1. Define the schema\n",
"\n",
"The schema is handed to the model as a **tool definition**, and each field's\n",
"`description` is the only instruction the model gets about what belongs there. An\n",
"undescribed field is a field the model guesses at.\n",
"\n",
"Rules worth knowing:\n",
"\n",
"- root `type: object` with a **non-empty** `properties` — an empty one is refused, because\n",
" it would bill you to extract nothing\n",
"- the root keys `_profile_config_version` and `entities` are **reserved** and rejected:\n",
" they name the storage envelope, so a schema using them could not be read back\n",
" unambiguously\n",
"- keep it small. The whole schema is sent to the model on every generation\n",
"\n",
"Descriptions are **not** enforced on write in this build — a property without one is\n",
"accepted and then quietly underfilled at generation time. Section F1 demonstrates it.\n",
"Treat descriptions as your job, not the validator's.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"SCHEMA = {\n",
" \"type\": \"object\",\n",
" \"properties\": {\n",
" \"occupation\": {\n",
" \"type\": \"string\",\n",
" \"description\": \"The person's current job title, in one short phrase.\",\n",
" },\n",
" \"location\": {\n",
" \"type\": \"string\",\n",
" \"description\": \"The city or region the person currently lives in.\",\n",
" },\n",
" \"interests\": {\n",
" \"type\": \"array\",\n",
" \"items\": {\"type\": \"string\"},\n",
" \"description\": \"Hobbies and topics they return to, as short lowercase tags.\",\n",
" },\n",
" \"dietary_restrictions\": {\n",
" \"type\": \"array\",\n",
" \"items\": {\"type\": \"string\"},\n",
" \"description\": \"Foods the person avoids, and why, if they said.\",\n",
" },\n",
" \"communication_style\": {\n",
" \"type\": \"string\",\n",
" \"enum\": [\"concise\", \"detailed\", \"casual\", \"formal\"],\n",
" \"description\": \"How this person prefers to be answered.\",\n",
" },\n",
" \"expertise_level\": {\n",
" \"type\": \"string\",\n",
" \"enum\": [\"beginner\", \"intermediate\", \"advanced\"],\n",
" \"description\": \"Their technical depth, judged from how they discuss their work.\",\n",
" },\n",
" },\n",
"}\n",
"\n",
"settings = client.update_profile_settings(\n",
" enabled=True,\n",
" schema=SCHEMA,\n",
" custom_instructions=(\n",
" \"Prefer facts the person stated outright over anything inferred. \"\n",
" \"Leave a field empty rather than guessing.\"\n",
" ),\n",
")\n",
"\n",
"# Sorted, because JSONB storage does not preserve the key order you sent.\n",
"# Compare a stored schema by SET, never by string or by key order.\n",
"stored = settings[\"entities\"][\"user\"][\"schema\"]\n",
"print(\"schema fields:\", sorted(stored[\"properties\"]))\n",
"print(\"enabled :\", settings[\"enabled\"])\n",
"print(\"capabilities :\", settings[\"capabilities\"])\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"`enabled` is project-wide; `schema` and `custom_instructions` apply to user profiles\n",
"and are stored under `entities`. The SDK takes them flat and nests them for you, so what\n",
"you write comes back unchanged from `get_profile_settings()`.\n",
"\n",
"Only the arguments you pass are written. To turn the feature off without touching your\n",
"schema, send `enabled` alone.\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 2. The turns\n",
"\n",
"Twelve conversation turns for one user. Nothing about `add()` changes — profiles are a\n",
"side effect of the normal pipeline.\n",
"\n",
"Twelve, not five, because generation fires when an entity crosses a **10-message\n",
"boundary**. Below that it waits for a flush window measured in hours, and this notebook\n",
"would sit there.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"TURNS = [\n",
" (\"user\", \"Hey — I just moved to Berlin for a new job.\"),\n",
" (\"assistant\", \"Congratulations! What's the new role?\"),\n",
" (\"user\", \"Senior data engineer at a logistics company. Mostly Spark and Airflow.\"),\n",
" (\"assistant\", \"Nice stack. How are you finding the pipelines there?\"),\n",
" (\"user\", \"Honestly the DAGs are a mess. I've been rewriting the partitioning to cut shuffle.\"),\n",
" (\"assistant\", \"That usually pays off fast. Anything blocking you?\"),\n",
" (\"user\", \"Just time. Keep it short when you answer me, I skim everything.\"),\n",
" (\"assistant\", \"Understood — short answers from here.\"),\n",
" (\"user\", \"Outside work I climb most weekends, and I'm learning German.\"),\n",
" (\"assistant\", \"Bouldering or ropes?\"),\n",
" (\"user\", \"Bouldering. Also — I'm vegetarian, so skip meat in any recipe suggestions.\"),\n",
" (\"assistant\", \"Noted, vegetarian only.\"),\n",
"]\n",
"\n",
"response = client.add(\n",
" [{\"role\": r, \"content\": c} for r, c in TURNS],\n",
" user_id=USER_ID,\n",
")\n",
"print(json.dumps(response, indent=2)[:300])\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The add is **async** — it returns an `event_id` and the memories do not exist yet. Poll\n",
"`GET /v1/event/{event_id}/` until it is `SUCCEEDED` or `FAILED`; that, not a sleep, is how\n",
"you know the add finished. Then let the extracted memories settle.\n",
"\n",
"Under load this can take a minute or more, so the cell says plainly whether it ran out of\n",
"time rather than printing `0 memories` as though that were the answer.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"event_id = response[\"event_id\"]\n",
"deadline = time.time() + 300\n",
"\n",
"# 1. The add itself. Terminal status, not a sleep.\n",
"event_status = None\n",
"while time.time() < deadline:\n",
" event_status = client.client.get(f\"/v1/event/{event_id}/\").json().get(\"status\")\n",
" if event_status in (\"SUCCEEDED\", \"FAILED\"):\n",
" break\n",
" print(f\" add {event_status}\")\n",
" time.sleep(5)\n",
"print(f\"add finished: {event_status}\")\n",
"# Stop here unless the add SUCCEEDED. A failed or unfinished add would otherwise let the\n",
"# generation below bill for a profile built without these memories.\n",
"if event_status != \"SUCCEEDED\":\n",
" raise RuntimeError(f\"add did not succeed (status={event_status}); not generating a profile\")\n",
"\n",
"# 2. Extraction lands in batches, so the FIRST non-empty page is not the whole set.\n",
"# Wait for the count to stop growing instead of breaking on the first result.\n",
"memories, stable = [], 0\n",
"while time.time() < deadline:\n",
" page = client.get_all(filters={\"user_id\": USER_ID}, page_size=50)\n",
" found = page.get(\"results\", []) if isinstance(page, dict) else page\n",
" stable = stable + 1 if found and len(found) == len(memories) else 0\n",
" memories = found\n",
" if stable >= 2: # two identical polls in a row\n",
" break\n",
" print(f\" ... {len(memories)} so far\")\n",
" time.sleep(5)\n",
"\n",
"if memories:\n",
" print(f\"\\n{len(memories)} memories extracted:\\n\")\n",
" for m in memories:\n",
" print(\" \\u2022\", m.get(\"memory\"))\n",
"else:\n",
" # Say so. Reporting '0 memories' as a result hides a busy or broken environment\n",
" # and makes the profile below look like it came from nothing.\n",
" print(\"\\nNO memories yet — extraction is still catching up, or the ingestion\")\n",
" print(\"worker is down. Everything below will report insufficient_data.\")\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 3. Read the profile\n",
"\n",
"Crossing the 10-message boundary should already have queued a generation. Read first —\n",
"and note that a known user with no profile yet is a **200 with a status**, not a 404. That\n",
"distinction is the whole point of the envelope.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"envelope = client.get_profile(USER_ID)\n",
"print(json.dumps(envelope, indent=2))\n",
"\n",
"print(\"\\nstatus vocabulary:\")\n",
"print(\" succeeded terminal — a generation ran AND the profile has content\")\n",
"print(\" pending queued or running\")\n",
"print(\" failed terminal — the last generation did not complete\")\n",
"print(\" not_enabled feature off, or plan below Pro\")\n",
"print(\" insufficient_data no content to show: no row yet, queued, or a\")\n",
"print(\" generation that legitimately found nothing\")\n",
"print()\n",
"print(\"`succeeded` is decided by the profile BODY, not by generation_count: an\")\n",
"print(\"empty extraction still increments the counter, so counting generations\")\n",
"print(\"reports 'done' for a profile with nothing in it.\")\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Force it, rather than waiting\n",
"\n",
"`generate_profile()` closes the bootstrapping gap: without it a new user has no profile\n",
"until their tenth message. One entity, a few seconds.\n",
"\n",
"Each call sends a new `Idempotency-Key` unless you pass one, and a new key starts a new job.\n",
"To retry a dropped request safely, generate the key yourself and pass the same\n",
"`idempotency_key` on every attempt: the server then returns the original job instead of\n",
"billing a second one.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"TERMINAL = {\"succeeded\", \"failed\", \"not_enabled\"}\n",
"\n",
"\n",
"def wait_for_profile(entity_id, timeout=300, interval=5, since=None):\n",
" \"\"\"Poll until terminal.\n",
"\n",
" `since` waits for a generation_count ABOVE that value, which is how you wait\n",
" for an UPDATE rather than accepting the profile you already had.\n",
"\n",
" `insufficient_data` is NOT terminal by itself — it also covers 'queued', so\n",
" poll through it and give up on the timeout instead.\n",
" \"\"\"\n",
" deadline = time.time() + timeout\n",
" body = None\n",
" while time.time() < deadline:\n",
" body = client.get_profile(entity_id)\n",
" status = (body.get(\"status\") or \"\").lower()\n",
" count = body.get(\"generation_count\") or 0\n",
" fresh = count > since if since is not None else True\n",
" if status == \"succeeded\" and fresh:\n",
" return body\n",
" if status in (\"failed\", \"not_enabled\"):\n",
" raise RuntimeError(f\"generation stopped: {status}\")\n",
" print(f\" ... {status} (generation_count={count})\")\n",
" time.sleep(interval)\n",
" raise TimeoutError(f\"not ready in {timeout}s: {body}\")\n",
"\n",
"\n",
"print(json.dumps(client.generate_profile(USER_ID), indent=2))\n",
"print(\"\\npolling...\")\n",
"\n",
"try:\n",
" body = wait_for_profile(USER_ID)\n",
" print(\"\\n=== PROFILE ===\")\n",
" print(json.dumps(body[\"profile\"], indent=2))\n",
" print(f\"\\nstatus={body['status']} generations={body['generation_count']} updated={body['updated_at']}\")\n",
"except TimeoutError as e:\n",
" # Say so plainly and let the rest of the notebook skip, rather than raising\n",
" # a NameError in every cell below and burying the real cause.\n",
" body = None\n",
" print(f\"\\nNO PROFILE: {e}\")\n",
" print(\"Generation never finished. Usually the ingestion worker is down, or\")\n",
" print(\"this project has no memories for the user yet.\")\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# The model must not invent fields outside your schema — the forced tool call is\n",
"# what makes that structural rather than a request.\n",
"if body is None:\n",
" print(\"skipped — no profile was generated above\")\n",
"else:\n",
" extra = set(body[\"profile\"]) - set(SCHEMA[\"properties\"])\n",
" print(\"fields outside the schema:\", extra or \"none\")\n",
"\n",
" # A forced JSON-Schema response makes the model emit SOMETHING for every property,\n",
" # so 'I found nothing' arrives as a type default: 0, \"\", [].\n",
" filled = {k: v for k, v in body[\"profile\"].items() if v not in (None, \"\", [], {}, 0)}\n",
" print(f\"genuinely populated: {len(filled)}/{len(SCHEMA['properties'])} -> {list(filled)}\")\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 4. Now watch it update\n",
"\n",
"Six more turns that contradict and extend what we already know: a promotion, a move, a\n",
"dropped hobby. A profile is a living document, not an append-only log — the model gets the\n",
"memories and rewrites the whole thing.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"if body is None:\n",
" print(\"skipped — no profile was generated above\")\n",
"else:\n",
" before = body[\"generation_count\"]\n",
"\n",
" MORE_TURNS = [\n",
" (\"user\", \"Update — I got promoted to staff engineer last week.\"),\n",
" (\"assistant\", \"Congratulations. Same team?\"),\n",
" (\"user\", \"Same company, but I'm relocating to Munich for it.\"),\n",
" (\"assistant\", \"Big move. How do you feel about it?\"),\n",
" (\"user\", \"Good. I've stopped climbing though — knee injury. Picked up cycling instead.\"),\n",
" (\"assistant\", \"Sorry about the knee. Cycling's kinder on it.\"),\n",
" ]\n",
"\n",
" followup = client.add(\n",
" [{\"role\": r, \"content\": c} for r, c in MORE_TURNS],\n",
" user_id=USER_ID,\n",
" )\n",
"\n",
" # Wait for the add to land before triggering: a generation queued before the new\n",
" # memories exist rewrites the profile from the OLD ones and looks like a no-op.\n",
" deadline = time.time() + 300\n",
" status = None\n",
" while time.time() < deadline:\n",
" status = client.client.get(f\"/v1/event/{followup['event_id']}/\").json().get(\"status\")\n",
" if status in (\"SUCCEEDED\", \"FAILED\"):\n",
" break\n",
" time.sleep(5)\n",
" print(\"follow-up add:\", status)\n",
" # Generating after a FAILED or unfinished add bills for a profile built from the OLD\n",
" # memories only, so stop instead.\n",
" if status != \"SUCCEEDED\":\n",
" raise RuntimeError(f\"follow-up add did not succeed (status={status}); not regenerating\")\n",
" time.sleep(15) # let extraction settle\n",
"\n",
" print(json.dumps(client.generate_profile(USER_ID), indent=2))\n",
" print(f\"\\npolling for a NEW generation (count must exceed {before})...\")\n",
" updated = wait_for_profile(USER_ID, since=before)\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"if body is None:\n",
" print(\"skipped — no profile was generated above\")\n",
"else:\n",
" print(f\"{'field':<22} {'before':<34} after\")\n",
" print(\"-\" * 92)\n",
" for field in SCHEMA[\"properties\"]:\n",
" b = json.dumps(body[\"profile\"].get(field))\n",
" a = json.dumps(updated[\"profile\"].get(field))\n",
" mark = \" \" if a == b else \"->\"\n",
" print(f\"{mark} {field:<20} {b[:32]:<34} {a[:32]}\")\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 5. Use it in a prompt\n",
"\n",
"The point of the structure is that it drops straight into a prompt — no list of memories\n",
"to summarize, no query to write.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"def build_system_prompt(entity_id):\n",
" result = client.get_profile(entity_id)\n",
" if result[\"status\"] != \"succeeded\":\n",
" # Branch on status, never on an empty profile: a user whose profile is\n",
" # still building is not a user you know nothing about.\n",
" return \"You are a helpful assistant.\"\n",
"\n",
" p = result[\"profile\"]\n",
" return f\"\"\"You are helping {entity_id}.\n",
"Occupation: {p.get(\"occupation\", \"unknown\")}\n",
"Location: {p.get(\"location\", \"unknown\")}\n",
"Interests: {\", \".join(p.get(\"interests\", [])) or \"unknown\"}\n",
"Dietary restrictions: {\", \".join(p.get(\"dietary_restrictions\", [])) or \"none stated\"}\n",
"Preferred style: {p.get(\"communication_style\", \"unknown\")}\n",
"\n",
"Match their style and do not explain what they already know.\"\"\"\n",
"\n",
"\n",
"print(build_system_prompt(USER_ID))\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 6. Judge a schema before committing to it\n",
"\n",
"`sample_profiles()` runs your schema against up to 10 **real** users that have memories.\n",
"\n",
"These are real generations and the results are **kept** — a dry run would cost exactly the\n",
"same and leave those users no better off. It is not a free preview.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# 202, not 200: the sample generations are queued, not finished.\n",
"#\n",
"# A 409 `already_running` means a sample from an earlier run is still going.\n",
"# That is the cooldown working, not an error — reuse that job rather than\n",
"# failing the notebook.\n",
"try:\n",
" job = client.sample_profiles(limit=3)\n",
" print(json.dumps(job, indent=2)[:400])\n",
" print(\"\\nsampled\", job.get(\"sampled\"), \"entities:\", job.get(\"entity_ids\"))\n",
"except Exception as e:\n",
" detail = str(e)\n",
" print(\"sample refused:\", detail[:200])\n",
" running = json.loads(detail).get(\"error\", {}).get(\"job_id\") if detail.startswith(\"{\") else None\n",
" job = {\"job_id\": running, \"status_url\": f\"/v2/profiles/jobs/{running}/\"} if running else None\n",
" print(\"reusing the running job:\", running)\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Poll `status_url` to see how the job went. `total` is `null` until enumeration finishes,\n",
"so format it defensively rather than assuming a number.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": "JOB_TERMINAL = {\"SUCCEEDED\", \"PARTIALLY_SUCCEEDED\", \"FAILED\", \"CANCELLED\"}\n\n\ndef wait_for_job(job_response, timeout=300, interval=5):\n \"\"\"Poll a generation job. Prefer status_url over a bare job id, so a route\n change needs no client update. Raise on timeout so an unfinished job is never\n mistaken for a finished one.\"\"\"\n handle = job_response.get(\"status_url\") or job_response[\"job_id\"]\n deadline = time.time() + timeout\n status = None\n while time.time() < deadline:\n status = client.get_profile_job(handle)[\"job\"]\n total = status.get(\"total\")\n print(\n f\" {status['status']} \"\n f\"completed={status.get('completed', 0)}/{total if total is not None else '?'} \"\n f\"succeeded={status.get('succeeded', 0)} \"\n f\"failed={status.get('failed', 0)} \"\n f\"skipped={status.get('skipped', 0)}\"\n )\n if str(status.get(\"status\", \"\")).upper() in JOB_TERMINAL:\n return status\n time.sleep(interval)\n raise TimeoutError(\n f\"job not terminal in {timeout}s (last status: {status.get('status') if status else 'none'})\"\n )\n\n\nif job is None:\n print(\"no sample job to poll\")\nelse:\n final = wait_for_job(job)\n\n print(\"\\n--- what the sample produced ---\")\n for entity_id in job.get(\"entity_ids\", []):\n got = client.get_profile(entity_id)\n print(f\"\\n{entity_id} [{got['status']}]\")\n print(\" \", json.dumps(got[\"profile\"])[:220])"
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 7. Apply a new schema to existing users\n",
"\n",
"A new schema shapes the **next** generation. Profiles that already exist keep their values\n",
"until their user is generated again — which happens as that user sends more memories, or\n",
"when you call `generate_profile()` for them.\n",
"\n",
"A field you **remove** stops being maintained: on the next generation, fields your schema\n",
"no longer defines are pruned. Keep a field for as long as you want its value kept."
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"---\n",
"\n",
"# Failure scenarios\n",
"\n",
"Everything above is the path that works. These are the ways it says no, and what each one\n",
"means. Run this section last: F3 deliberately leaves the project switched off for a moment.\n",
"\n",
"> **About the `HTTP error occurred:` lines below.** The SDK logs every 4xx at\n",
"> ERROR level before raising, so they appear even for the failures these cells\n",
"> deliberately catch. Read the line printed *after* each one — that is the cell's\n",
"> own verdict. Nothing here is unhandled.\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## F1. Schemas that get rejected\n",
"\n",
"Rejections happen on **write**, where you can see and fix them — not silently at\n",
"generation time, where you would only notice as an empty profile weeks later.\n",
"\n",
"The last case matters for storage: the user schema lives in one JSONB column alongside\n",
"the envelope that separates it, so a schema using the envelope's own reserved keys could\n",
"not be read back unambiguously. It is refused rather than stored.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"BAD_SCHEMAS = [\n",
" ({\"type\": \"object\", \"properties\": {}}, \"empty — bills you to extract nothing\"),\n",
" ({\"type\": \"array\", \"items\": {\"type\": \"string\"}}, \"root must be an object\"),\n",
" (\n",
" {\n",
" \"type\": \"object\",\n",
" \"properties\": {\"tone\": {\"type\": \"string\", \"description\": \"Preferred tone.\"}},\n",
" # At the schema ROOT, which is where the envelope's own keys live.\n",
" \"_profile_config_version\": 1,\n",
" \"entities\": {\"user\": {}},\n",
" },\n",
" \"reserved settings keys at the schema root\",\n",
" ),\n",
"]\n",
"\n",
"for bad, why in BAD_SCHEMAS:\n",
" try:\n",
" client.update_profile_settings(schema=bad)\n",
" print(f\"ACCEPTED (unexpected): {why}\")\n",
" except Exception as e:\n",
" print(f\"rejected [{why}]:\\n {str(e)[:160]}\\n\")\n",
"\n",
"# NOT rejected: a property with no description. The validator allows it and the\n",
"# model then has nothing to go on, so the field comes back empty. Descriptions are\n",
"# your job, not the validator's.\n",
"try:\n",
" client.update_profile_settings(schema={\"type\": \"object\", \"properties\": {\"x\": {\"type\": \"string\"}}})\n",
" print(\"accepted [no description on 'x'] <- the trap: valid to store, useless to generate\")\n",
"finally:\n",
" client.update_profile_settings(schema=SCHEMA) # put the good one back\n",
"\n",
"restored = client.get_profile_settings()[\"entities\"][\"user\"][\"schema\"]\n",
"assert set(restored[\"properties\"]) == set(SCHEMA[\"properties\"])\n",
"print(\"\\nschema restored:\", sorted(restored[\"properties\"]))\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## F2. A user that does not exist\n",
"\n",
"404 means only \"no such user\". A known user with no profile yet is a 200 carrying\n",
"`insufficient_data`, so an ordinary empty state never looks like an error.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from mem0.exceptions import MemoryNotFoundError\n",
"\n",
"try:\n",
" client.get_profile(\"user_who_never_existed\")\n",
" print(\"ACCEPTED (unexpected)\")\n",
"except MemoryNotFoundError as e:\n",
" print(\"404 as intended:\", str(e)[:120])\n",
"\n",
"# ...versus a real user who simply has no profile row yet.\n",
"fresh = f\"demo_never_profiled_{uuid.uuid4().hex[:6]}\"\n",
"client.add([{\"role\": \"user\", \"content\": \"One passing remark.\"}], user_id=fresh)\n",
"time.sleep(5)\n",
"print(\"known but unprofiled:\", client.get_profile(fresh)[\"status\"])\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## F3. Profiles turned off\n",
"\n",
"`enabled` is the one project-wide switch. Every generation path then refuses.\n",
"\n",
"Nothing is deleted. Your schema and every profile you already built are kept, so turning\n",
"it back on resumes rather than restarts.\n",
"\n",
"Note what a read does **not** do — a profile that already exists keeps reporting\n",
"`succeeded` and keeps returning its content. `not_enabled` is only what you get for a user\n",
"with no profile yet. Turning the feature off stops new work; it does not hide what has\n",
"already been built.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"client.update_profile_settings(enabled=False)\n",
"\n",
"print(\"read (demo user) :\", client.get_profile(USER_ID)[\"status\"])\n",
"print(\"read (never profiled) :\", client.get_profile(fresh)[\"status\"])\n",
"try:\n",
" client.generate_profile(USER_ID)\n",
" print(\"trigger: ACCEPTED (unexpected)\")\n",
"except Exception as e:\n",
" print(\"trigger:\", str(e)[:160])\n",
"\n",
"back = client.update_profile_settings(enabled=True) # put it back\n",
"print(\"\\nrestored:\", back[\"enabled\"])\n",
"print(\"schema survived:\", bool(back[\"entities\"][\"user\"][\"schema\"]))\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 7. Cleanup\n",
"\n",
"Removes the demo users. The profile row cascades with the entity.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": "# Restore the project's profile settings to the start-of-run snapshot in `finally`, so a\n# failed delete still leaves a shared project as we found it. Passing the original values\n# (including None) clears anything this notebook set: the SDK treats an explicit None as\n# \"clear\" and an omitted argument as \"unchanged\".\ntry:\n # `fresh` only exists if the error-handling section ran.\n for entity_id in (USER_ID, globals().get(\"fresh\")):\n if entity_id is None:\n continue\n r = client.client.delete(f\"/v2/entities/user/{entity_id}/\")\n print(entity_id, \"->\", r.status_code)\nfinally:\n _user = ORIGINAL_SETTINGS.get(\"entities\", {}).get(\"user\", {})\n client.update_profile_settings(\n enabled=ORIGINAL_SETTINGS.get(\"enabled\", False),\n schema=_user.get(\"schema\"),\n custom_instructions=_user.get(\"custom_instructions\"),\n )\n print(\"profile settings restored to the pre-notebook snapshot\")"
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"---\n",
"\n",
"## Cheat sheet\n",
"\n",
"| Want | Call | Cost |\n",
"| --- | --- | --- |\n",
"| configure | `update_profile_settings(...)` | free |\n",
"| read | `get_profile(user_id)` | free |\n",
"| one user now | `generate_profile(user_id)` | 1 LLM call |\n",
"| try a schema | `sample_profiles(limit=n)` | ≤10 real generations, kept |\n",
"| poll a job | `get_profile_job(status_url)` | free |\n",
"\n",
"**Settings apply to user profiles.** The stored shape is:\n",
"\n",
"```json\n",
"{\"enabled\": true,\n",
" \"entities\": {\"user\": {\"schema\": {...}, \"custom_instructions\": \"...\"}},\n",
" \"capabilities\": {\"full_rebuild\": false}}\n",
"```\n",
"\n",
"The SDK takes these flat and nests them for you. Only the fields you pass are written;\n",
"`enabled` is the one project-wide switch.\n",
"\n",
"**Left alone, generation fires** on a 10-message boundary, or after a flush window\n",
"measured in hours. `generate_profile()` is how you skip the wait for one user.\n",
"\n",
"**Three traps:**\n",
"\n",
"1. `insufficient_data` is not a terminal verdict — it also covers \"queued\", so poll\n",
" through it and give up on a timeout instead.\n",
"2. `succeeded` is decided by the profile **body**, not `generation_count`. An empty\n",
" extraction still increments the counter.\n",
"3. A forced JSON-Schema response emits something for every property, so \"nothing found\"\n",
" arrives as a type default — `\"\"`, `[]`, `0` — not as a missing key.\n"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.12.4"
}
},
"nbformat": 4,
"nbformat_minor": 4
}
-1
View File
@@ -6,7 +6,6 @@ Agent and editor integrations. Most packages are self-contained; coding-agent pl
|-----------|---------|-------|------|------|
| `vercel-ai-sdk/` | `@mem0/vercel-ai-provider` | tsup (CJS+ESM) | ESLint + Prettier | jest + vitest (edge/node) |
| `openclaw/` | `@mem0/openclaw-mem0` | tsup (ESM) | none | vitest |
| `hermes-plugin-mem0/` | Standalone Hermes memory provider | none | ruff + isort | pytest (host-stubbed offline); live CLI/Desktop validation |
| `agent-plugin-core/` | Shared Python/TypeScript behavior, skill templates, builds, and conformance | Python build script | ruff + tsc | pytest + node:test |
| `mem0-agent-plugin/` | One portable Agent Plugins v1 package | Python | ruff | shared conformance |
| `claude-code-plugin/`, `cursor-plugin/`, `codex-plugin/`, `kimi-plugin/`, `antigravity-plugin/` | Self-contained native plugins generated from the shared Python core | Python | ruff | pytest |
+2 -2
View File
@@ -8,7 +8,7 @@ This directory is the single source of shared memory behavior for Mem0 coding-ag
integrations/
├── agent-plugin-core/ # Shared source; never installed as a plugin
│ ├── python/ # Claude-derived capture, recall, MCP, scoping, and telemetry
│ ├── typescript/ # Shared lifecycle, search prompts, formatting, identity, scoping, and telemetry
│ ├── typescript/ # Shared lifecycle, formatting, identity, scoping, and telemetry
│ ├── skills/ # The only source for the six generated memory skills
│ ├── build/ # Bundle builder, schemas, and validation
│ ├── conformance/ # One offline/live verification entry point
@@ -45,7 +45,7 @@ New Git repository writes use a hashed remote identity for shared `agent_id`. Se
Captured prompts and responses preserve their full text after secret redaction. Python extraction splits oversized input across requests without dropping message text. The session-end worker flushes the conversation already collected by hooks without adding the final answer again. Search queries, retrieved context, and tool evidence have separate limits.
TypeScript hosts reuse redaction, lifecycle utilities, and the search prompts in `typescript/src/prompts.ts`, which a test keeps identical to the Python core. They retain their own tools, scopes, and capture events. They do not inherit the Python `repo`/`dir`/`mine` contract or its background batching. OpenCode captures selected user prompts; Pi and DeepSeek capture completed conversation turns; OpenClaw selects recent messages and earlier summaries, then filters noise. Removing message-length truncation does not turn these integrations into complete transcript archives.
TypeScript hosts reuse redaction and lifecycle utilities but retain their own tools, scopes, and capture events. They do not inherit the Python `repo`/`dir`/`mine` contract or its background batching. OpenCode captures selected user prompts; Pi and DeepSeek capture completed conversation turns; OpenClaw selects recent messages and earlier summaries, then filters noise. Removing message-length truncation does not turn these integrations into complete transcript archives.
For installation, follow the host guides: [Claude Code](../../docs/integrations/claude-code.mdx), [Cursor](../../docs/integrations/cursor.mdx), [Codex](../../docs/integrations/codex.mdx), [Kimi](../../docs/integrations/kimi.mdx), and [Antigravity](../../docs/integrations/antigravity.mdx).
@@ -21,9 +21,16 @@ from memory_core import (
PROTOCOL_VERSION = "2024-11-05"
TOOL_NAME = "search_memories"
TOOL_DESCRIPTION = (
"Search memories from earlier work in this repository. Use it before "
"repeating investigation or when earlier decisions, fixes, commands, or "
"results may help."
"Search memories from earlier work in this repository. ALWAYS call this "
"tool before answering anything that could depend on prior context: the "
"user's preferences, facts about this codebase, history, people, projects, "
"or earlier decisions. Do not rely on the chat window alone. The "
"repository's memory is shared by everyone who works in it and includes "
"what it took to run, test, or build here, so search before assuming an "
"invocation works. The scope argument changes what is searched: 'repo' "
"(default) is the whole repository's shared memory plus your own "
"preferences, 'dir' narrows the shared part to the directory you are "
"working in, and 'mine' is your preferences alone."
)
TOOL_SCHEMA = {
"type": "object",
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.3"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -71,13 +71,15 @@ MAX_FLUSH_ATTEMPTS = 5
FORGET_PAGE_SIZE = 100
FORGET_MAX_PAGES = 50
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help with future coding work.
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help anyone with future coding work in this repository.
A completed change should produce one memory explaining the resulting behavior, where it is implemented when useful, and any important constraints or reasoning. Exploration or accepted decisions may produce separate memories only when they are independently useful.
Use the coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
A command that failed and was then made to work should produce one memory naming the failing invocation, the error it returned, and the invocation that succeeded. Do not save one-off errors caused by an edit still in progress, transient network failures, or anything a rerun would fix on its own.
Write about the repository, not the user, assistant, session, or task. Do not include test results, documentation updates, release notes, or temporary state.
Use the current coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
Write about the repository, not the user, assistant, session, or task. Do not save personal preferences. Do not save a memory that only states which repository, branch, or directory the session worked in. Do not include test results, documentation updates, release notes, or temporary state.
If nothing useful was established, return no memories."""
@@ -12,9 +12,11 @@ Call `search_memories` with the user's question. Treat `--top-k`, `--category`,
query.
Omit `top_k` to use Mem0's configured default. Omit `category` to search every
category. Omit `scope` to use the configured default, normally `repo`: this
repository's shared memory, which everyone who works in it contributes to,
plus your own preferences.
category; a category is a best-effort label Mem0 assigned when it saved the
memory, so if a category search misses, repeat it without the category. Omit
`scope` to use the configured default, normally `repo`: this repository's
shared memory, which everyone who works in it contributes to, plus your own
preferences.
Pass `scope` when the question needs something else: `dir` to narrow the
shared memory to the directory you are working in (a package inside a
@@ -1,6 +1,5 @@
import type { MemoryLike } from "./formatting.ts";
import { formatMemoryCompact } from "./formatting.ts";
import { RECALL_HEADING } from "./prompts.ts";
const MAX_RECALL_QUERY_CHARS = 6_000;
export const DEFAULT_MAX_CONTEXT_CHARS = 4_000;
@@ -71,14 +70,12 @@ export function extractConversation(
}
interface RecallOptions {
heading?: string;
maxChars?: number;
seenIds?: Set<string>;
timeoutMs?: number;
}
interface MemoryLifecycleOptions {
recallHeading?: string;
maxContextChars?: number;
recallTimeoutMs?: number;
}
@@ -112,7 +109,6 @@ class MemoryLifecycle {
search: (query: string) => Promise<{ results?: unknown[] }>,
): Promise<string> {
return buildRecallContext(prompt, enabled, search, {
heading: this.#options.recallHeading,
maxChars: this.#options.maxContextChars,
seenIds: this.#seenMemoryIds,
timeoutMs: this.#options.recallTimeoutMs,
@@ -152,7 +148,8 @@ export async function buildRecallContext(
const unseen = memories.filter((memory) => !options.seenIds?.has(memory.id));
if (!unseen.length) return "";
const prefix = `<mem0-relevant-memories>\n${options.heading ?? RECALL_HEADING}\n`;
const prefix =
"<mem0-relevant-memories>\nRetrieved automatically for the current request. This is a shallow first pass — search mem0_memory for more if you need it.\n";
const suffix = "\n</mem0-relevant-memories>";
const maxChars = options.maxChars ?? DEFAULT_MAX_CONTEXT_CHARS;
const lines: string[] = [];
@@ -1,10 +0,0 @@
export const SEARCH_WHEN =
"before repeating investigation or when earlier decisions, fixes, commands, or results may help";
export const SEARCH_TOOL_DESCRIPTION = `Search memories from earlier work in this repository. Use it ${SEARCH_WHEN}.`;
export const SEARCH_QUERY_DESCRIPTION = "A direct question about earlier work in this repository.";
export const RECALL_HEADING = "Mem0 found these relevant memories from earlier work in this repository:";
export const USER_SEARCH_TOOL_DESCRIPTION = `Search memories from earlier work. Use it ${SEARCH_WHEN}.`;
export const USER_SEARCH_QUERY_DESCRIPTION = "A direct question about earlier work.";
export const USER_RECALL_HEADING = "Mem0 found these relevant memories from earlier work:";
@@ -1,5 +1,3 @@
import { randomUUID } from "node:crypto";
import { redactSecrets } from "./lifecycle.ts";
const POSTHOG_API_KEY = "phc_hgJkUVJFYtmaJqrvf6CYN67TIQ8yhXAkWzUn9AMU4yX";
@@ -76,108 +74,34 @@ export function errorKind(error: unknown): string {
return error instanceof Error ? error.constructor.name : "other";
}
// Delivery is retried in memory, not spooled to disk, and that is a decision
// rather than an omission. The Python core spools because its hooks are separate
// processes that fire per tool call and exit immediately, so nothing survives
// without a file. These plugins are loaded into a host that lives for a whole
// session, so re-queueing covers the same transient failures without the claim
// and lease machinery a correct cross-process spool needs. What that leaves
// uncovered is narrow: a session that both starts and ends with no connectivity.
const RETRY_BACKOFF_CEILING_MS = 60_000;
// Consecutive failed flushes before the queue is dropped. Deliberately NOT the
// same thing as Python's budget, which rides in the claim filename and so
// follows one batch: this counter lives in the closure and counts the outage,
// not the payload. Events captured between attempts join the same queue and go
// with it. Per-batch accounting would need an attempt count on every event, and
// the queue is already bounded, so the simpler rule is the one in force here.
// Without any bound a payload the server will never accept is retried for the
// whole session and, now that the backlog is preferred over new events, holds
// the queue against everything behind it.
const MAX_DELIVERY_ATTEMPTS = 5;
export function createTelemetry(config: TelemetryConfig) {
let queue: Record<string, unknown>[] = [];
let timer: ReturnType<typeof setInterval> | undefined;
let consecutiveFailures = 0;
let retryNotBefore = 0;
let exitFlushAttempted = false;
let flushing = false;
const flushThreshold = config.flushThreshold ?? 10;
const maxQueueSize = config.maxQueueSize ?? 100;
const deliver = config.delivery ?? (async (batch: Record<string, unknown>[]) => {
const response = await fetch(POSTHOG_BATCH_URL, {
await fetch(POSTHOG_BATCH_URL, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ api_key: POSTHOG_API_KEY, batch }),
signal: AbortSignal.timeout(3_000),
});
// fetch only rejects on a network-level failure. Without this check a 500,
// a 503 or a 429 resolved normally and the batch was counted as delivered
// and dropped, which is the likelier outage than a refused connection.
// Any non-2xx is retried, matching the Python core: the backoff and the
// queue bound contain a payload that will never be accepted, because the
// re-queued batch sits at the front and is the first thing evicted.
if (!response.ok) throw new Error(`posthog responded ${response.status}`);
});
async function flush(force = false): Promise<void> {
// One at a time. Two overlapping flushes each detach the queue and each
// prepend their own batch back on failure, so the later batch lands in front
// of the earlier one and the truncation then drops the OLDER events first,
// inverting the priority the failure path exists to establish. A second
// caller returns immediately; the queue waits for the next flush.
if (flushing) return;
async function flush(): Promise<void> {
if (!queue.length) return;
// `force` skips the cooldown. beforeExit is the last chance this process
// gets, and gating it on the same backoff meant that after any failure the
// exit flush did nothing and the queue died with the process, which is the
// loss this whole mechanism exists to prevent.
if (!force && Date.now() < retryNotBefore) return;
const batch = queue;
queue = [];
flushing = true;
try {
await deliver(batch);
consecutiveFailures = 0;
retryNotBefore = 0;
} catch {
// Put it back. Detaching the batch and swallowing the error deleted the
// events outright, so any blip silently dropped telemetry with nothing
// recording that it had happened. Every event carries a uuid, so a retry
// that duplicates one PostHog already accepted is collapsed there.
//
consecutiveFailures += 1;
if (consecutiveFailures >= MAX_DELIVERY_ATTEMPTS) {
// Give up on the queue so a failing outage cannot hold it for the
// session. This drops whatever is queued now, which includes events
// captured during the outage, not only the batch that kept failing.
consecutiveFailures = 0;
retryNotBefore = 0;
return;
}
// Keep the FRONT on overflow, so the batch being retried survives and a
// new event is what gets dropped. Matches the Python core, where record()
// refuses new events once the spool is full rather than evicting the
// backlog. Keeping the newest would throw away exactly the events this
// retry exists to save.
queue = [...batch, ...queue].slice(0, maxQueueSize);
retryNotBefore = Date.now() + Math.min(2 ** consecutiveFailures * 1_000, RETRY_BACKOFF_CEILING_MS);
} finally {
flushing = false;
// Telemetry must never affect plugin behavior.
}
}
function beforeExit(): void {
// Once, and only once. Node re-emits beforeExit whenever the handler
// schedules more async work, so an unconditional forced flush looped until
// the attempt budget was spent: five attempts against a 3s delivery timeout
// is fifteen seconds added to the shutdown of whatever editor or CLI is
// hosting this. The backoff used to end that loop after one attempt, and
// removing it for the forced path removed the only thing bounding it.
if (exitFlushAttempted) return;
exitFlushAttempted = true;
void flush(true);
void flush();
}
function build(event: string, properties: Record<string, unknown> = {}): Record<string, unknown> | null {
@@ -188,15 +112,6 @@ export function createTelemetry(config: TelemetryConfig) {
return {
event: config.eventName?.(event) ?? event,
distinct_id: distinctId,
// Stamped once, at capture. This is what makes retrying safe: a batch
// re-sent after a failure carries the same ids, so PostHog collapses
// anything it already accepted instead of counting it twice.
uuid: randomUUID(),
// Capture time, not ingestion time. Events now sit through backoff and
// across a whole outage, so without this PostHog records them whenever
// delivery happened to succeed. It also matters for the uuid dedupe
// above, whose key includes the event date.
timestamp: new Date().toISOString(),
properties: {
...safeProperties(properties),
...safeProperties(config.commonProperties ?? {}),
@@ -219,10 +134,8 @@ export function createTelemetry(config: TelemetryConfig) {
try {
const payload = build(event, properties);
if (!payload) return;
// Full means drop this event, not evict the backlog. Same rule as the
// failure path above and as Python's record().
if (queue.length >= maxQueueSize) return;
queue.push(payload);
if (queue.length > maxQueueSize) queue = queue.slice(-maxQueueSize);
if (!timer) {
timer = setInterval(() => void flush(), config.flushIntervalMs ?? 5_000);
timer.unref?.();
@@ -236,8 +149,6 @@ export function createTelemetry(config: TelemetryConfig) {
function resetForTesting(): void {
queue = [];
consecutiveFailures = 0;
retryNotBefore = 0;
if (timer) clearInterval(timer);
timer = undefined;
process.off("beforeExit", beforeExit);
@@ -1,31 +0,0 @@
import assert from "node:assert/strict";
import { readFileSync } from "node:fs";
import test from "node:test";
import { buildRecallContext } from "../src/lifecycle.ts";
import {
RECALL_HEADING,
SEARCH_QUERY_DESCRIPTION,
SEARCH_TOOL_DESCRIPTION,
USER_RECALL_HEADING,
} from "../src/prompts.ts";
const pythonSource = (name: string) =>
readFileSync(new URL(`../../python/${name}`, import.meta.url), "utf8").replace(/"\s*\n\s*"/g, "");
test("search prompts match the Python core", () => {
assert.ok(pythonSource("mcp_server.py").includes(SEARCH_TOOL_DESCRIPTION));
assert.ok(pythonSource("mcp_server.py").includes(SEARCH_QUERY_DESCRIPTION));
assert.ok(pythonSource("hook_runner.py").includes(RECALL_HEADING));
});
test("recall context uses the repository heading unless the host overrides it", async () => {
const search = async () => ({ results: [{ id: "m1", memory: "Use pnpm" }] });
assert.ok((await buildRecallContext("package manager", true, search)).includes(RECALL_HEADING));
assert.ok(
(await buildRecallContext("package manager", true, search, { heading: USER_RECALL_HEADING })).includes(
USER_RECALL_HEADING,
),
);
});
@@ -122,250 +122,3 @@ test("error classification does not expose messages", () => {
assert.equal(errorKind(new Error("request timeout")), "timeout");
assert.equal(errorKind(new Error("fetch failed")), "network");
});
test("a failed delivery keeps the batch instead of deleting it", async () => {
// The defect: the queue was detached before the await and the error swallowed,
// so one blip destroyed the events with nothing recording that it happened.
const attempts: Record<string, unknown>[][] = [];
let failNext = true;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d",
flushThreshold: 1000,
delivery: async (batch) => {
attempts.push(batch);
if (failNext) throw new Error("network down");
},
});
telemetry.capture("one");
telemetry.capture("two");
await telemetry.flush();
assert.equal(attempts.length, 1);
assert.equal(telemetry.queueForTesting().length, 2, "events were dropped on failure");
failNext = false;
// Backoff is in force, so wait it out the way wall time would.
await new Promise((resolve) => setTimeout(resolve, 2_100));
await telemetry.flush();
assert.equal(attempts.length, 2, "never retried");
assert.equal(telemetry.queueForTesting().length, 0);
telemetry.resetForTesting();
});
test("a retried event carries the same uuid so PostHog can collapse it", async () => {
const attempts: Record<string, unknown>[][] = [];
let failNext = true;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d",
flushThreshold: 1000,
delivery: async (batch) => {
attempts.push(batch);
if (failNext) throw new Error("network down");
},
});
telemetry.capture("once");
await telemetry.flush();
failNext = false;
await new Promise((resolve) => setTimeout(resolve, 2_100));
await telemetry.flush();
assert.equal(attempts.length, 2);
const first = attempts[0][0].uuid;
assert.ok(first, "events carry no uuid, so a retry would double count");
assert.equal(attempts[1][0].uuid, first, "retry minted a new uuid");
telemetry.resetForTesting();
});
test("repeated failures back off instead of retrying every flush", async () => {
let calls = 0;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d",
flushThreshold: 1000,
delivery: async () => { calls += 1; throw new Error("blocked"); },
});
telemetry.capture("one");
await telemetry.flush();
await telemetry.flush();
await telemetry.flush();
assert.equal(calls, 1, "a blocked host was hammered on every flush");
assert.equal(telemetry.queueForTesting().length, 1, "the event was lost while backing off");
telemetry.resetForTesting();
});
test("a full queue drops the new event and keeps the batch being retried", async () => {
// Python's record() refuses new events once the spool is full rather than
// evicting the backlog. Keeping the newest here would throw away exactly the
// events the retry exists to save.
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d",
flushThreshold: 1000, maxQueueSize: 3,
delivery: async () => { throw new Error("down"); },
});
// Fill past the cap BEFORE the flush, so the re-queue actually has to truncate.
// Capturing only two left the queue empty at re-queue time and the slice on the
// failure path never ran, which is the half that decides the direction.
telemetry.capture("a");
telemetry.capture("b");
telemetry.capture("c");
await telemetry.flush();
telemetry.capture("d");
telemetry.capture("e");
const events = telemetry.queueForTesting().map((e) => (e as any).event);
assert.equal(events.length, 3, "queue grew past maxQueueSize");
assert.deepEqual(events, ["a", "b", "c"], "the retried batch was evicted instead of the new events");
telemetry.resetForTesting();
});
test("the exit-time flush ignores the backoff", async () => {
// beforeExit is the last chance the process gets. Gating it on the same
// cooldown meant that after any failure it did nothing and the queue died.
let attempts = 0;
let failing = true;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: async () => { attempts += 1; if (failing) throw new Error("down"); },
});
telemetry.capture("a");
await telemetry.flush();
assert.equal(attempts, 1);
failing = false;
await telemetry.flush();
assert.equal(attempts, 1, "the backoff should still hold for an ordinary flush");
await telemetry.flush(true);
assert.equal(attempts, 2, "the exit flush was suppressed by the backoff");
assert.equal(telemetry.queueForTesting().length, 0);
telemetry.resetForTesting();
});
test("every event carries a capture-time timestamp", async () => {
const sent: Record<string, unknown>[][] = [];
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: async (batch) => { sent.push(batch); },
});
telemetry.capture("a");
const capturedAt = Date.now();
await new Promise((resolve) => setTimeout(resolve, 50));
await telemetry.flush();
const stamped = sent[0][0].timestamp as string;
assert.ok(stamped, "no timestamp, so PostHog would record delivery time");
assert.ok(Math.abs(Date.parse(stamped) - capturedAt) < 1_000, "not capture time");
telemetry.resetForTesting();
});
test("a batch the server will never accept is eventually given up on", async () => {
let attempts = 0;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: async () => { attempts += 1; throw new Error("permanently bad"); },
});
telemetry.capture("doomed");
for (let i = 0; i < 8; i += 1) await telemetry.flush(true);
assert.ok(attempts <= 6, `retried ${attempts} times with no cap`);
assert.equal(telemetry.queueForTesting().length, 0, "a doomed batch held the queue forever");
telemetry.resetForTesting();
});
test("an HTTP error response is a failure, not a delivery", async () => {
// fetch only rejects on a network-level failure, so a 500 used to resolve
// normally and the batch was dropped as delivered. Exercises the real default
// delivery path rather than an injected one, which is where this hid.
const realFetch = globalThis.fetch;
let calls = 0;
globalThis.fetch = (async () => {
calls += 1;
return new Response("upstream is unwell", { status: 503 });
}) as typeof fetch;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
});
try {
telemetry.capture("during.outage");
await telemetry.flush();
assert.equal(calls, 1, "never reached the network");
assert.equal(telemetry.queueForTesting().length, 1, "a 503 was counted as delivered");
} finally {
globalThis.fetch = realFetch;
telemetry.resetForTesting();
}
});
test("a 2xx is a delivery", async () => {
const realFetch = globalThis.fetch;
globalThis.fetch = (async () => new Response("ok", { status: 200 })) as typeof fetch;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
});
try {
telemetry.capture("fine");
await telemetry.flush();
assert.equal(telemetry.queueForTesting().length, 0, "a good response did not clear the queue");
} finally {
globalThis.fetch = realFetch;
telemetry.resetForTesting();
}
});
test("the exit flush is attempted once, not until the budget is spent", async () => {
// Node re-emits beforeExit whenever the handler schedules async work, so an
// unconditional forced flush looped until MAX_DELIVERY_ATTEMPTS. Against the
// real 3s delivery timeout that is fifteen seconds added to a host's shutdown.
let attempts = 0;
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: async () => { attempts += 1; throw new Error("down"); },
});
telemetry.capture("a");
const handlers = process.listeners("beforeExit");
const ours = handlers[handlers.length - 1] as () => void;
ours();
ours();
ours();
await new Promise((resolve) => setTimeout(resolve, 20));
assert.equal(attempts, 1, `exit flush ran ${attempts} times`);
telemetry.resetForTesting();
});
test("overlapping flushes do not reorder the backlog behind newer events", async () => {
// Each flush detaches the queue and prepends its own batch back on failure, so
// two in flight at once put the LATER batch in front of the earlier one. The
// truncation then drops the older events first, inverting the priority the
// failure path exists to establish.
let release: (() => void)[] = [];
const telemetry = createTelemetry({
host: "h", source: "S", version: "1", distinctId: "d", flushThreshold: 1000,
delivery: () => new Promise((_resolve, reject) => { release.push(() => reject(new Error("down"))); }),
});
telemetry.capture("first");
const a = telemetry.flush();
telemetry.capture("second");
const b = telemetry.flush();
release.forEach((fn) => fn());
await Promise.all([a, b]);
const events = telemetry.queueForTesting().map((e) => (e as any).event);
assert.equal(release.length, 1, "a second delivery started while one was in flight");
assert.deepEqual(events, ["first", "second"], `backlog reordered: ${events.join(",")}`);
telemetry.resetForTesting();
});
@@ -21,9 +21,16 @@ from memory_core import (
PROTOCOL_VERSION = "2024-11-05"
TOOL_NAME = "search_memories"
TOOL_DESCRIPTION = (
"Search memories from earlier work in this repository. Use it before "
"repeating investigation or when earlier decisions, fixes, commands, or "
"results may help."
"Search memories from earlier work in this repository. ALWAYS call this "
"tool before answering anything that could depend on prior context: the "
"user's preferences, facts about this codebase, history, people, projects, "
"or earlier decisions. Do not rely on the chat window alone. The "
"repository's memory is shared by everyone who works in it and includes "
"what it took to run, test, or build here, so search before assuming an "
"invocation works. The scope argument changes what is searched: 'repo' "
"(default) is the whole repository's shared memory plus your own "
"preferences, 'dir' narrows the shared part to the directory you are "
"working in, and 'mine' is your preferences alone."
)
TOOL_SCHEMA = {
"type": "object",
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.3"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -71,13 +71,15 @@ MAX_FLUSH_ATTEMPTS = 5
FORGET_PAGE_SIZE = 100
FORGET_MAX_PAGES = 50
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help with future coding work.
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help anyone with future coding work in this repository.
A completed change should produce one memory explaining the resulting behavior, where it is implemented when useful, and any important constraints or reasoning. Exploration or accepted decisions may produce separate memories only when they are independently useful.
Use the coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
A command that failed and was then made to work should produce one memory naming the failing invocation, the error it returned, and the invocation that succeeded. Do not save one-off errors caused by an edit still in progress, transient network failures, or anything a rerun would fix on its own.
Write about the repository, not the user, assistant, session, or task. Do not include test results, documentation updates, release notes, or temporary state.
Use the current coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
Write about the repository, not the user, assistant, session, or task. Do not save personal preferences. Do not save a memory that only states which repository, branch, or directory the session worked in. Do not include test results, documentation updates, release notes, or temporary state.
If nothing useful was established, return no memories."""
@@ -1,6 +1,6 @@
{
"id": "mem0",
"version": "0.3.3",
"version": "0.3.1",
"homepage": "https://docs.mem0.ai/integrations/antigravity",
"native": {
"pluginRoot": "${ANTIGRAVITY_PLUGIN_ROOT}",
@@ -12,9 +12,11 @@ Call `search_memories` with the user's question. Treat `--top-k`, `--category`,
query.
Omit `top_k` to use Mem0's configured default. Omit `category` to search every
category. Omit `scope` to use the configured default, normally `repo`: this
repository's shared memory, which everyone who works in it contributes to,
plus your own preferences.
category; a category is a best-effort label Mem0 assigned when it saved the
memory, so if a category search misses, repeat it without the category. Omit
`scope` to use the configured default, normally `repo`: this repository's
shared memory, which everyone who works in it contributes to, plus your own
preferences.
Pass `scope` when the question needs something else: `dir` to narrow the
shared memory to the directory you are working in (a package inside a
@@ -1,6 +1,6 @@
{
"name": "mem0",
"version": "0.3.3",
"version": "0.3.1",
"description": "Cross-session memory and token savings for coding agents.",
"author": {
"name": "Mem0"
@@ -12,8 +12,11 @@ You are Mem0's Sonnet coding agent. Complete the work the main agent gives you.
Work in the separate Git worktree Claude Code created for you. Return a tested
result that the main agent can review without doing the same work again.
When memories from earlier sessions could help, call `search_memories` with a
focused question before searching the repository again.
ALWAYS call `search_memories` before answering anything that could depend on
prior context (the user's preferences, facts about this codebase, history,
people, projects, or earlier decisions). Do not rely on the chat window or
assume you know enough from the current conversation. Search with a focused
question before investigating the repository.
Inspect the relevant code and repository rules. Reproduce the problem when that
helps. Decide the implementation details, edit files when asked, and test the
@@ -21,9 +21,16 @@ from memory_core import (
PROTOCOL_VERSION = "2024-11-05"
TOOL_NAME = "search_memories"
TOOL_DESCRIPTION = (
"Search memories from earlier work in this repository. Use it before "
"repeating investigation or when earlier decisions, fixes, commands, or "
"results may help."
"Search memories from earlier work in this repository. ALWAYS call this "
"tool before answering anything that could depend on prior context: the "
"user's preferences, facts about this codebase, history, people, projects, "
"or earlier decisions. Do not rely on the chat window alone. The "
"repository's memory is shared by everyone who works in it and includes "
"what it took to run, test, or build here, so search before assuming an "
"invocation works. The scope argument changes what is searched: 'repo' "
"(default) is the whole repository's shared memory plus your own "
"preferences, 'dir' narrows the shared part to the directory you are "
"working in, and 'mine' is your preferences alone."
)
TOOL_SCHEMA = {
"type": "object",
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.3"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -71,13 +71,15 @@ MAX_FLUSH_ATTEMPTS = 5
FORGET_PAGE_SIZE = 100
FORGET_MAX_PAGES = 50
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help with future coding work.
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help anyone with future coding work in this repository.
A completed change should produce one memory explaining the resulting behavior, where it is implemented when useful, and any important constraints or reasoning. Exploration or accepted decisions may produce separate memories only when they are independently useful.
Use the coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
A command that failed and was then made to work should produce one memory naming the failing invocation, the error it returned, and the invocation that succeeded. Do not save one-off errors caused by an edit still in progress, transient network failures, or anything a rerun would fix on its own.
Write about the repository, not the user, assistant, session, or task. Do not include test results, documentation updates, release notes, or temporary state.
Use the current coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
Write about the repository, not the user, assistant, session, or task. Do not save personal preferences. Do not save a memory that only states which repository, branch, or directory the session worked in. Do not include test results, documentation updates, release notes, or temporary state.
If nothing useful was established, return no memories."""
@@ -1,6 +1,6 @@
{
"id": "mem0",
"version": "0.3.3",
"version": "0.3.1",
"homepage": "https://docs.mem0.ai/integrations/claude-code",
"native": {
"pluginRoot": "${CLAUDE_PLUGIN_ROOT}",
@@ -12,9 +12,11 @@ Call `search_memories` with the user's question. Treat `--top-k`, `--category`,
query.
Omit `top_k` to use Mem0's configured default. Omit `category` to search every
category. Omit `scope` to use the configured default, normally `repo`: this
repository's shared memory, which everyone who works in it contributes to,
plus your own preferences.
category; a category is a best-effort label Mem0 assigned when it saved the
memory, so if a category search misses, repeat it without the category. Omit
`scope` to use the configured default, normally `repo`: this repository's
shared memory, which everyone who works in it contributes to, plus your own
preferences.
Pass `scope` when the question needs something else: `dir` to narrow the
shared memory to the directory you are working in (a package inside a
@@ -2,7 +2,6 @@ from __future__ import annotations
import json
import os
import re
import sqlite3
import subprocess
import sys
@@ -2307,8 +2306,8 @@ def test_sidekick_instructions_reject_unrequested_related_changes():
prompt = (PLUGIN_ROOT / "agents" / "sidekick.md").read_text()
normalized = " ".join(prompt.split())
assert "Skill" in prompt.split("---", 2)[1]
assert "call `search_memories` with a" in normalized
assert "focused question before searching the repository again" in normalized
assert "ALWAYS call `search_memories` before answering anything" in normalized
assert "Do not rely on the chat window" in normalized
assert "Complete only the work the main agent assigned" in normalized
assert "Do not make related improvements" in normalized
assert "report them separately" in normalized
@@ -3461,12 +3460,7 @@ def test_automatic_flush_can_be_disabled_for_external_harnesses(isolated_env):
def test_version_is_single_sourced():
manifest = json.loads((PLUGIN_ROOT / ".claude-plugin" / "plugin.json").read_text())
assert manifest["name"] == "mem0"
# Compared against PLUGIN_VERSION, never a literal. A hardcoded version here
# was one more place to edit on every release, inside the test asserting the
# version is single-sourced, and it caught nothing that the agreement checks
# below do not: fifteen places set to the same wrong value would still pass.
assert re.fullmatch(r"\d+\.\d+\.\d+", memory_core.PLUGIN_VERSION), memory_core.PLUGIN_VERSION
assert manifest["version"] == memory_core.PLUGIN_VERSION
assert manifest["version"] == memory_core.PLUGIN_VERSION == "0.3.1"
root = REPOSITORY_ROOT
for mp in (root / "marketplace.json", root / ".claude-plugin" / "marketplace.json"):
entry = next(p for p in json.loads(mp.read_text())["plugins"] if p["name"] == "mem0")
@@ -4333,7 +4327,7 @@ def test_flush_sends_unified_body_with_both_agent_and_user_id(isolated_env, monk
assert sent_body["run_id"] == "s1"
assert "lane" not in sent_body["metadata"]
assert "Save concise repository facts" in sent_body["agent_custom_instructions"]
assert "Write about the repository, not the user" in sent_body["agent_custom_instructions"]
assert "invocation that succeeded" in sent_body["agent_custom_instructions"]
assert "Do not save repository facts" in sent_body["custom_instructions"]
assert sent_body["custom_categories"] == memory_core.CODING_MEMORY_CATEGORIES
store.close()
@@ -1,6 +1,6 @@
{
"name": "mem0",
"version": "0.3.3",
"version": "0.3.1",
"description": "Cross-session memory and token savings for coding agents.",
"author": { "name": "Mem0", "email": "support@mem0.ai" },
"homepage": "https://docs.mem0.ai/integrations/codex",
+10 -3
View File
@@ -21,9 +21,16 @@ from memory_core import (
PROTOCOL_VERSION = "2024-11-05"
TOOL_NAME = "search_memories"
TOOL_DESCRIPTION = (
"Search memories from earlier work in this repository. Use it before "
"repeating investigation or when earlier decisions, fixes, commands, or "
"results may help."
"Search memories from earlier work in this repository. ALWAYS call this "
"tool before answering anything that could depend on prior context: the "
"user's preferences, facts about this codebase, history, people, projects, "
"or earlier decisions. Do not rely on the chat window alone. The "
"repository's memory is shared by everyone who works in it and includes "
"what it took to run, test, or build here, so search before assuming an "
"invocation works. The scope argument changes what is searched: 'repo' "
"(default) is the whole repository's shared memory plus your own "
"preferences, 'dir' narrows the shared part to the directory you are "
"working in, and 'mine' is your preferences alone."
)
TOOL_SCHEMA = {
"type": "object",
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.3"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -71,13 +71,15 @@ MAX_FLUSH_ATTEMPTS = 5
FORGET_PAGE_SIZE = 100
FORGET_MAX_PAGES = 50
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help with future coding work.
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help anyone with future coding work in this repository.
A completed change should produce one memory explaining the resulting behavior, where it is implemented when useful, and any important constraints or reasoning. Exploration or accepted decisions may produce separate memories only when they are independently useful.
Use the coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
A command that failed and was then made to work should produce one memory naming the failing invocation, the error it returned, and the invocation that succeeded. Do not save one-off errors caused by an edit still in progress, transient network failures, or anything a rerun would fix on its own.
Write about the repository, not the user, assistant, session, or task. Do not include test results, documentation updates, release notes, or temporary state.
Use the current coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
Write about the repository, not the user, assistant, session, or task. Do not save personal preferences. Do not save a memory that only states which repository, branch, or directory the session worked in. Do not include test results, documentation updates, release notes, or temporary state.
If nothing useful was established, return no memories."""
+1 -1
View File
@@ -1,6 +1,6 @@
{
"id": "mem0",
"version": "0.3.3",
"version": "0.3.1",
"homepage": "https://docs.mem0.ai/integrations/codex",
"native": {
"pluginRoot": "${PLUGIN_ROOT}",
@@ -12,9 +12,11 @@ Call `search_memories` with the user's question. Treat `--top-k`, `--category`,
query.
Omit `top_k` to use Mem0's configured default. Omit `category` to search every
category. Omit `scope` to use the configured default, normally `repo`: this
repository's shared memory, which everyone who works in it contributes to,
plus your own preferences.
category; a category is a best-effort label Mem0 assigned when it saved the
memory, so if a category search misses, repeat it without the category. Omit
`scope` to use the configured default, normally `repo`: this repository's
shared memory, which everyone who works in it contributes to, plus your own
preferences.
Pass `scope` when the question needs something else: `dir` to narrow the
shared memory to the directory you are working in (a package inside a
@@ -1,6 +1,6 @@
{
"name": "mem0",
"version": "0.3.3",
"version": "0.3.1",
"description": "Cross-session memory and token savings for coding agents.",
"author": { "name": "Mem0", "email": "support@mem0.ai" },
"homepage": "https://docs.mem0.ai/integrations/cursor",
+10 -3
View File
@@ -21,9 +21,16 @@ from memory_core import (
PROTOCOL_VERSION = "2024-11-05"
TOOL_NAME = "search_memories"
TOOL_DESCRIPTION = (
"Search memories from earlier work in this repository. Use it before "
"repeating investigation or when earlier decisions, fixes, commands, or "
"results may help."
"Search memories from earlier work in this repository. ALWAYS call this "
"tool before answering anything that could depend on prior context: the "
"user's preferences, facts about this codebase, history, people, projects, "
"or earlier decisions. Do not rely on the chat window alone. The "
"repository's memory is shared by everyone who works in it and includes "
"what it took to run, test, or build here, so search before assuming an "
"invocation works. The scope argument changes what is searched: 'repo' "
"(default) is the whole repository's shared memory plus your own "
"preferences, 'dir' narrows the shared part to the directory you are "
"working in, and 'mine' is your preferences alone."
)
TOOL_SCHEMA = {
"type": "object",
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.3"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -71,13 +71,15 @@ MAX_FLUSH_ATTEMPTS = 5
FORGET_PAGE_SIZE = 100
FORGET_MAX_PAGES = 50
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help with future coding work.
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help anyone with future coding work in this repository.
A completed change should produce one memory explaining the resulting behavior, where it is implemented when useful, and any important constraints or reasoning. Exploration or accepted decisions may produce separate memories only when they are independently useful.
Use the coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
A command that failed and was then made to work should produce one memory naming the failing invocation, the error it returned, and the invocation that succeeded. Do not save one-off errors caused by an edit still in progress, transient network failures, or anything a rerun would fix on its own.
Write about the repository, not the user, assistant, session, or task. Do not include test results, documentation updates, release notes, or temporary state.
Use the current coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
Write about the repository, not the user, assistant, session, or task. Do not save personal preferences. Do not save a memory that only states which repository, branch, or directory the session worked in. Do not include test results, documentation updates, release notes, or temporary state.
If nothing useful was established, return no memories."""
+1 -1
View File
@@ -1,6 +1,6 @@
{
"id": "mem0",
"version": "0.3.3",
"version": "0.3.1",
"homepage": "https://docs.mem0.ai/integrations/cursor",
"native": {
"pluginRoot": "${CURSOR_PLUGIN_ROOT}",
@@ -12,9 +12,11 @@ Call `search_memories` with the user's question. Treat `--top-k`, `--category`,
query.
Omit `top_k` to use Mem0's configured default. Omit `category` to search every
category. Omit `scope` to use the configured default, normally `repo`: this
repository's shared memory, which everyone who works in it contributes to,
plus your own preferences.
category; a category is a best-effort label Mem0 assigned when it saved the
memory, so if a category search misses, repeat it without the category. Omit
`scope` to use the configured default, normally `repo`: this repository's
shared memory, which everyone who works in it contributes to, plus your own
preferences.
Pass `scope` when the question needs something else: `dir` to narrow the
shared memory to the directory you are working in (a package inside a
+1 -6
View File
@@ -1,6 +1,6 @@
{
"name": "@mem0/deepseek-plugin",
"version": "0.3.2",
"version": "0.3.0",
"description": "Mem0 long-term memory as a native DeepSeek Harness (Cordis) plugin.",
"type": "module",
"license": "Apache-2.0",
@@ -60,10 +60,5 @@
"tsup": "^8.5.0",
"typescript": "^5.6.0",
"vitest": "^4.1.7"
},
"pnpm": {
"overrides": {
"axios@<1.20.0": ">=1.20.0 <2.0.0"
}
}
}
+4 -7
View File
@@ -4,9 +4,6 @@ settings:
autoInstallPeers: false
excludeLinksFromLockfile: false
overrides:
axios@<1.20.0: '>=1.20.0 <2.0.0'
importers:
.:
@@ -578,8 +575,8 @@ packages:
asynckit@0.4.0:
resolution: {integrity: sha512-Oei9OH4tRh0YqU3GxhX79dM/mwVgvbZJaSNaRk+bshkj0S5cfHcgYakreBjrHwatXKbz+IoIdYLxrKim2MjW0Q==}
axios@1.20.0:
resolution: {integrity: sha512-r8aOh8j9cGKpgQAqpzrUHnSIc6a59Y3Xf/cv8sy1DrHCkZHzQGEuoq1tARk6qSyDdtQGSDgpb9kFlruzPvrgwg==}
axios@1.19.0:
resolution: {integrity: sha512-ht/iuYZXEjFxLH/Hkezgd7m6JKlHHXEUSneaDz8uZe1Gj5QZtCnpyDsckvAiEnT89OEbCLmnte4R4sn7P0EKFw==}
bundle-require@5.1.0:
resolution: {integrity: sha512-3WrrOuZiyaaZPWiEt4G3+IffISVC9HYlWueJEBWED4ZH4aIAC2PnkdnuRrR94M+w6yGWn4AglWtJtBI8YqvgoA==}
@@ -1613,7 +1610,7 @@ snapshots:
asynckit@0.4.0: {}
axios@1.20.0:
axios@1.19.0:
dependencies:
follow-redirects: 1.16.0
form-data: 4.0.6
@@ -1859,7 +1856,7 @@ snapshots:
mem0ai@3.1.6:
dependencies:
axios: 1.20.0
axios: 1.19.0
openai: 4.104.0(zod@3.25.76)
uuid: 11.1.1
zod: 3.25.76
+4 -8
View File
@@ -21,11 +21,6 @@ import { truncateOutput } from "./output.ts";
import { resolveSearchFilters, resolveAddParams } from "./scoping.ts";
import { captureEvent, errorKind } from "./telemetry.ts";
import { createMemoryLifecycle } from "../../agent-plugin-core/typescript/src/lifecycle.ts";
import {
USER_RECALL_HEADING,
USER_SEARCH_QUERY_DESCRIPTION,
USER_SEARCH_TOOL_DESCRIPTION,
} from "../../agent-plugin-core/typescript/src/prompts.ts";
export const name = "mem0";
export const inject = ["tools", "systemPrompt"];
@@ -112,7 +107,7 @@ export function apply(ctx: Context, config: Config): void {
const stateFor = (session: object): SessionState => {
let state = sessionStates.get(session);
if (!state) {
const lifecycle = createMemoryLifecycle({ recallHeading: USER_RECALL_HEADING });
const lifecycle = createMemoryLifecycle();
lifecycle.beginSession();
state = { lifecycle, messages: [] };
sessionStates.set(session, state);
@@ -202,9 +197,10 @@ export function apply(ctx: Context, config: Config): void {
ctx.tools.register(
defineTool({
name: "search_memory",
description: USER_SEARCH_TOOL_DESCRIPTION,
description:
"Search the user's long-term Mem0 memory for facts relevant to a query. Use proactively before answering anything that may depend on what the user told you earlier.",
parameters: {
query: { type: "string", description: USER_SEARCH_QUERY_DESCRIPTION, required: true },
query: { type: "string", description: "What to recall.", required: true },
limit: {
type: "integer",
description: `Max results to return (default ${DEFAULT_SEARCH_LIMIT}).`,
@@ -1,3 +0,0 @@
__pycache__/
*.pyc
.venv/
-236
View File
@@ -1,236 +0,0 @@
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [2023] [Taranjeet Singh]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
Third-party software
--------------------
hermes-plugin-mem0
Mem0 contributions are licensed under the Apache License, Version 2.0.
This plugin includes code from Nous Research's hermes-plugin-mem0:
https://github.com/NousResearch/hermes-plugin-mem0
Source commit: 3fc36950b2b7c19cdd81c6de99f10d2cbed850af
The imported code retains the following MIT license and copyright notice:
MIT License
Copyright (c) 2025 Nous Research
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
-131
View File
@@ -1,131 +0,0 @@
# Mem0 for Hermes Agent
Persistent memory for [Hermes Agent](https://github.com/NousResearch/hermes-agent), powered by [Mem0](https://mem0.ai).
This standalone plugin recalls relevant memories before a response and extracts facts from conversations afterward. It works alongside Hermes' file-based memory and supports Mem0 Cloud, a self-hosted Mem0 server, or the in-process OSS SDK.
## Features
- **Automatic recall and capture** across conversations.
- **Four agent tools** to search, add, update, and delete memories.
- **Three backend modes** with interactive setup through `hermes memory setup`.
- **User-scoped memories** with agent and channel metadata on writes.
## Setup
### 1. Install
Requires [Hermes Agent](https://github.com/NousResearch/hermes-agent) with memory-provider plugin support and Python 3.11 or later. After this directory is merged to Mem0's main branch:
```bash
hermes plugins install mem0ai/mem0/integrations/hermes-plugin-mem0
hermes plugins enable mem0
```
Hermes installers with plugin dependency support install `mem0ai>=2.0.10,<3` and `httpx>=0.27,<1` from this directory's `pyproject.toml`. On older hosts such as Hermes v0.21.3, install those requirements into the **Hermes Python environment** explicitly. The OSS setup wizard installs additional provider packages as needed.
> Hermes versions that still bundle Mem0 prefer the bundled provider. Use a Hermes release that has completed the standalone-provider migration; installing this plugin alone does not replace the bundled implementation. See [Existing users and migration](#existing-users-and-migration).
### 2. Configure
```bash
hermes memory setup mem0
```
Run this in an interactive terminal and choose a backend:
| Mode | What you need |
|------|----------------|
| **Platform** (default) | A Mem0 API key from [app.mem0.ai](https://app.mem0.ai/dashboard/api-keys) |
| **Self-hosted server** | A running [Mem0 server](../../server), its URL, and its API key unless authentication is disabled |
| **OSS** | An LLM, embedder, and vector store; no Mem0 API key needed |
For a self-hosted server, choose **Self-hosted server** and enter its URL and API key. Requests use `X-API-Key` and the server's `/search` and `/memories` routes. Setting `host` selects this backend unless `mode` is `oss`.
For the in-process SDK, choose **Open Source**. The wizard offers OpenAI or Ollama and local Qdrant or PGVector. Use manual configuration for custom OpenAI-compatible endpoints, deployment names, or a Qdrant server. OSS does not use Mem0 Cloud; data still goes to whichever model services you configure. Run setup again to switch modes. When switching to Platform, remove any stale `MEM0_HOST` setting from the environment and profile `.env`.
Desktop sessions in the same process and profile share local Qdrant storage when their OSS settings match. Operations are serialized, and storage closes after the last session releases it. Conflicting settings are rejected without changing existing memories; close the active sessions before changing models or credentials. Use a Qdrant server or the self-hosted Mem0 HTTP API when separate processes (for example, CLI and Desktop together) need the same store.
Hermes hosts whose `hermes memory setup --help` lists only a provider argument reject options such as `--mode`, `--host`, and `--oss-llm` before the plugin runs. Use the interactive command above, or the [manual profile configuration](https://docs.mem0.ai/integrations/hermes) for unattended setup. Redirected input cannot select the mode picker; it falls back to Platform.
### 3. Verify
```bash
hermes memory status
```
Start a fresh Hermes conversation and ask it to remember a fact, then search for that fact in a later session using the same user identity.
## Tools
| Tool | Description | Parameters |
|------|-------------|------------|
| `mem0_search` | Search memories by meaning | `query`, optional `top_k` (default 10, max 50) and `rerank` (Platform only) |
| `mem0_add` | Store text verbatim, without fact extraction | `content` |
| `mem0_update` | Update a memory's text | `memory_id`, `text` |
| `mem0_delete` | Delete a memory | `memory_id` |
## Configuration
Settings live in `$HERMES_HOME/mem0.json`; the default Hermes home is `~/.hermes`. Setup normally stores API keys in that profile's `.env`. Distinct OpenAI LLM/embedder keys and database credentials are stored in the OSS configuration. Setup writes both files atomically with owner-only permissions.
| Key | Default | Description |
|-----|---------|-------------|
| `mode` | `platform` | `platform` for Cloud/server routing, or `oss` for the in-process SDK |
| `host` | unset | Self-hosted server URL; ignored in OSS mode |
| `user_id` | gateway user ID, then `hermes-user` | Set a stable ID to share memories across gateways |
| `agent_id` | `hermes` | Agent identifier attached to writes |
| `rerank` | `false` | Platform reranking for automatic recall and tool searches that omit `rerank` |
| `sync_max_chars` | `450` | Maximum characters per user/assistant message sent for automatic extraction |
| `oss` | `{}` | OSS LLM, embedder, and vector-store configuration written by setup |
`MEM0_MODE`, `MEM0_HOST`, `MEM0_USER_ID`, and `MEM0_AGENT_ID` provide environment defaults; non-empty file settings override them. `MEM0_API_KEY` supplies the Cloud or server key when `api_key` is not set in the file.
An explicit `user_id` other than `hermes-user` takes precedence over a gateway's native user ID. Searches use that user identity across sessions; writes attach `agent_id` and `metadata.channel`.
## Automatic recall and capture
Recall waits up to three seconds for memories relevant to the current message. If results are late, the model can still call `mem0_search`.
After a turn, a background worker sends the user message and assistant response for extraction. Each message is truncated to **450 characters by default in every mode**, preferring a sentence boundary. Increase `sync_max_chars` to suit your model's context limit. Explicit `mem0_add` calls store their supplied text verbatim.
Capture is best effort: if the previous sync remains busy after a five-second wait, the new turn is skipped. There is no durable queue. Five consecutive backend failures pause calls for two minutes before retrying.
Graceful shutdown waits for active recall and capture workers before closing the backend, including at Python process exit. Backend network timeouts still apply: self-hosted HTTP capture has a 120-second read timeout and a 30-second connection timeout; other self-hosted HTTP operations use 30 seconds. Forced termination, including Hermes' 30-second exit watchdog, can still interrupt pending writes.
## Existing users and migration
Keep `memory.provider: mem0`, your existing `mem0.json`, `MEM0_*` settings, user identity, and OSS storage paths. Moving the plugin does not require moving memories or rerunning setup.
Automatic migration also requires coordination in Hermes:
1. A Hermes build containing [migration support from PR #114569](https://github.com/NousResearch/hermes-agent/pull/114569).
2. An approved catalog entry named `mem0`, pointing to `https://github.com/mem0ai/mem0`, with `subdir: integrations/hermes-plugin-mem0` and a reviewed full commit SHA.
3. Removal of Hermes' bundled Mem0 provider, which otherwise takes precedence.
With those in place, Hermes can install a missing configured provider during `hermes update` or at agent startup. Startup installation respects `security.allow_lazy_installs`; disabled or offline installation requires manual action. Merging this directory alone does not register the catalog entry or complete rollout.
The plugin supports CLI setup/status. It does not include a Desktop configuration panel or provider-specific CLI commands.
## Troubleshooting
- **Mem0 unavailable:** run `hermes memory status`. Check the API key and backend connectivity; after five consecutive failures the circuit breaker waits two minutes.
- **Memories missing:** confirm the same user identity across sessions and check `sync_max_chars`. Automatic extraction may omit facts; use `mem0_add` to store exact text.
- **OSS connection refused:** check the configured model/vector service, or filesystem permissions for local Qdrant.
- **Embedding dimension mismatch:** initialization fails without deleting existing data. Restore the previous embedding model/dimensions, or use a new collection and migrate data explicitly.
## Development
From the Mem0 repository root, with `ruff` and `isort` installed:
```bash
ruff check integrations/hermes-plugin-mem0
isort --check-only --profile black integrations/hermes-plugin-mem0
```
Validate changes in an isolated Hermes profile using live CLI and Desktop sessions. Check memory
add/search/update/delete, automatic capture and recall, overlapping Desktop sessions, and persistence after restart.
## License
[Apache-2.0](LICENSE) for Mem0 contributions. Includes code from [Nous Research's standalone Mem0 provider](https://github.com/NousResearch/hermes-plugin-mem0/tree/3fc36950b2b7c19cdd81c6de99f10d2cbed850af) under MIT; its original license and copyright notice are preserved in the third-party section of [LICENSE](LICENSE).
-504
View File
@@ -1,504 +0,0 @@
"""Mem0 memory plugin — MemoryProvider interface.
Server-side fact extraction and semantic search via the Mem0 Platform API (cloud), a
self-hosted Mem0 server (MEM0_HOST, HTTP), or OSS Memory. Secrets live in $HERMES_HOME/.env
(MEM0_API_KEY, MEM0_HOST); settings in $HERMES_HOME/mem0.json via `hermes memory setup`:
mode ("platform"|"oss"), host, user_id (canonical id across gateways; unset → gateway-native
id), agent_id. MEM0_* env vars remain a fallback.
"""
from __future__ import annotations
import atexit
import json
import logging
import threading
import time
from contextlib import suppress
from pathlib import Path
from typing import Any, Dict, List
from agent.memory_provider import MemoryProvider, spawn_context_thread
from agent.secret_scope import get_secret
from tools.registry import tool_error
from utils import atomic_json_write, read_json_or_empty
logger = logging.getLogger(__name__)
# Circuit breaker: after _BREAKER_THRESHOLD consecutive failures, pause API
# calls for _BREAKER_COOLDOWN_SECS to avoid hammering a down server.
_BREAKER_THRESHOLD, _BREAKER_COOLDOWN_SECS, _PREFETCH_WAIT_SECS = 5, 120, 3
_CLIENT_ERROR_TYPES = ("MemoryNotFoundError", "ValidationError")
# Placeholder user_id. initialize() treats it as "no operator-configured user_id"
# so legacy mem0.json files written by the wizard don't override gateway-native ids.
_DEFAULT_USER_ID = "hermes-user"
# sync_turn sends the whole turn to the backend for fact extraction. OSS embedding
# models often have small context windows (bge-small-zh-v1.5: 512 tokens ≈ 500 chars;
# jina-embeddings-v3: 8192), and oversized turns make backend.add() raise — Ollama
# answers HTTP 500, hosted APIs return INPUT_TOKEN_LIMIT_EXCEEDED — which _try only
# logs, silently dropping the turn's memory extraction. Cap each message up front.
# The default fits a 512-token embedder (measured: 450 OK, 600 -> HTTP 500 on
# bge-small-zh-v1.5:f16); ``sync_max_chars`` in mem0.json raises it for larger windows.
_SYNC_MSG_MAX_CHARS = 450
# Sentence ends recognized when trimming a synced message. Deliberately unordered:
# the LAST boundary of ANY kind wins, so one CJK stop early in a mixed-script turn
# cannot outrank a Latin stop near the end of the window. ``".\n"`` is not listed —
# its index can never exceed the bare ``"."`` it starts with.
_SYNC_SENTENCE_ENDS = ("。", "!", "?", ".", "!", "?")
def _truncate_for_sync(text: str, max_len: int = _SYNC_MSG_MAX_CHARS) -> str:
"""Cap a synced message at its last sentence boundary within ``max_len``.
Short messages pass through unchanged; long ones keep the last complete
sentence inside the window so fact extraction still sees coherent statements,
with a hard cut as fallback when no boundary exists (or one only appears in
the first third of the window, which usually means unsegmented input).
"""
if len(text) <= max_len:
return text
window = text[:max_len]
cut = max(window.rfind(sep) for sep in _SYNC_SENTENCE_ENDS)
if cut > max_len // 3:
return text[:cut + 1]
return text[:max_len]
def _is_client_error(exc: Exception) -> bool:
"""True for user-caused errors (bad ID, not found) that should NOT trip circuit breaker."""
err_str = str(exc).lower()
return type(exc).__name__ in _CLIENT_ERROR_TYPES or any(s in err_str for s in ("404", "not found", "valid uuid"))
def _load_config() -> dict:
"""Env vars provide defaults; $HERMES_HOME/mem0.json overrides individual keys.
Layering avoids a silent failure when the JSON file exists but lacks fields
like ``api_key`` that the user set in ``.env``."""
from hermes_constants import get_hermes_home
# Identity (user/agent id), host and mode are .env values like the key: read them through the
# profile scope too, or a secondary profile's memories land in the default profile's account.
# A scope-less multiplex caller raises here on purpose — that is a spawn-site bug, and
# swallowing it would silently route the turn's memories to the default profile.
config = {"mode": get_secret("MEM0_MODE", "") or "platform", "host": get_secret("MEM0_HOST", "") or "",
"agent_id": get_secret("MEM0_AGENT_ID", "") or "hermes", "oss": {}}
if user_id := get_secret("MEM0_USER_ID", ""): # only when explicitly configured, so initialize() can fall back to the gateway-native id
config["user_id"] = user_id
file_cfg = read_json_or_empty(get_hermes_home() / "mem0.json")
config.update({k: v for k, v in file_cfg.items() if v is not None and v != ""})
# MEM0_API_KEY authenticates the Platform and self-hosted HTTP backends; pure OSS mode builds its
# backend from the local ``oss`` config and has no platform credential to resolve, so a profile
# scope WITHOUT the key must still load an OSS config (#99121 as it stands today: the caller is
# scoped, the scope is just empty). Decided after mem0.json overrode the env defaults because
# the file may be what selects ``oss``. Scope-less callers already raised above.
if config.get("mode", "platform") == "oss":
config.setdefault("api_key", "")
elif not config.get("api_key"):
config["api_key"] = get_secret("MEM0_API_KEY", "")
return config
def _schema(name: str, description: str, properties: dict[str, tuple[str, str]], required: list[str]) -> dict:
props = {k: {"type": t, "description": d} for k, (t, d) in properties.items()}
return {"name": name, "description": description, "parameters": {"type": "object", "properties": props, "required": required}}
TOOL_SCHEMAS = [
_schema("mem0_search", "Search the user's memories by meaning; returns facts ranked by relevance. Use this before answering any question that may depend on what you know about the user (preferences, facts, history, people, projects, past decisions). For multi-part or multi-hop questions, call it several times — vary the wording and run follow-up searches on what earlier results reveal; one search is rarely enough.",
{"query": ("string", "What to search for."), "top_k": ("integer", "Max results (default: 10, max: 50)."), "rerank": ("boolean", "Rerank results for relevance (default: false, platform mode only).")}, ["query"]),
_schema("mem0_add", "Store a durable fact about the user, verbatim (no LLM extraction). Call this the moment the user states a lasting preference, correction, decision, or personal detail worth recalling on future turns — don't wait to be asked to remember. Skip transient chit-chat and facts you've already stored.",
{"content": ("string", "The fact to store.")}, ["content"]),
_schema("mem0_update", "Replace the text of an existing memory by its ID (take the ID from a mem0_search result). Use when a stored fact has changed or was wrong — correct it in place instead of adding a duplicate.",
{"memory_id": ("string", "Memory UUID to update."), "text": ("string", "New text content.")}, ["memory_id", "text"]),
_schema("mem0_delete", "Delete a memory by its ID (take the ID from a mem0_search result). Use when a stored fact is obsolete or the user asks you to forget it; prefer mem0_update if the fact merely changed.",
{"memory_id": ("string", "Memory UUID to delete.")}, ["memory_id"]),
]
_PROMPT_BODY = (
"You have persistent memory of this user from past conversations. You should call mem0_search before answering anything that could depend on prior context (the user's preferences, facts, history, people, projects, or earlier decisions) — do not rely on the chat window alone, and do not assume you have no memory.\n"
"For multi-part or multi-hop questions, run several searches with different wording/angles and follow-up searches on what the first results surface; one search is rarely enough. Keep searching until you have every fact the question needs before you answer.\n"
"Tools: mem0_search to find memories, mem0_add to store facts, mem0_update and mem0_delete to manage by ID."
)
class Mem0MemoryProvider(MemoryProvider):
"""Mem0 memory with server-side extraction and semantic search (platform, self-hosted or OSS)."""
def __init__(self):
self._config = self._backend = self._sync_thread = self._prefetch_thread = None
self._mode, self._api_key, self._host, self._user_id, self._agent_id = "platform", "", "", _DEFAULT_USER_ID, "hermes"
self._rerank_default, self._channel = False, "cli" # channel = gateway name (cli/telegram/discord/...)
self._sync_max_chars = _SYNC_MSG_MAX_CHARS
self._prefetch_query = self._prefetch_result = ""
self._prefetch_done = self._atexit_registered = False
self._consecutive_failures, self._breaker_open_until = 0, 0.0 # circuit breaker state
self._breaker_lock, self._sync_lock, self._prefetch_lock = threading.Lock(), threading.Lock(), threading.Lock()
@property
def name(self) -> str:
return "mem0"
def is_available(self) -> bool:
cfg = _load_config()
if cfg.get("mode", "platform") == "oss":
return bool(cfg.get("oss", {}).get("vector_store"))
return bool(cfg.get("api_key") or cfg.get("host")) # platform needs a key; self-hosted a host (key optional with AUTH_DISABLED)
def save_config(self, values, hermes_home):
"""Merge-write config to $HERMES_HOME/mem0.json."""
config_path = Path(hermes_home) / "mem0.json"
atomic_json_write(config_path, {**read_json_or_empty(config_path), **values}, mode=0o600)
def get_config_schema(self):
cfg = _load_config()
api_key_required = cfg.get("mode", "platform") != "oss" and not cfg.get("host")
return [
{"key": "api_key", "description": "Mem0 Platform API key", "secret": True, "required": api_key_required, "env_var": "MEM0_API_KEY", "url": "https://app.mem0.ai"},
{"key": "host", "description": "Self-hosted Mem0 server URL (leave blank for cloud)", "required": False, "env_var": "MEM0_HOST"},
{"key": "user_id", "description": "User identifier", "default": "hermes-user"},
{"key": "agent_id", "description": "Agent identifier", "default": "hermes"},
{"key": "rerank", "description": "Enable reranking for recall", "default": "false", "choices": ["true", "false"]},
]
def post_setup(self, hermes_home: str, config: dict) -> None:
from ._setup import post_setup
post_setup(hermes_home, config)
def _oss_hint(self, template: str, default: str = "vector store") -> str:
"""OSS-only hint; ``{vs}`` is the configured vector-store provider. "" in other modes."""
return template.format(vs=self._config.get("oss", {}).get("vector_store", {}).get("provider", default)) if self._mode == "oss" else ""
def _create_backend(self):
# Lazy-install the mem0 SDK before the backend imports it (honors security.allow_lazy_installs);
# on failure the backend import raises the canonical error, captured below.
with suppress(Exception):
pass # dependencies come from pyproject.toml (installed by Hermes on install/enable/update)
try:
from . import _backend
if self._mode == "oss":
return _backend.OSSBackend(self._config.get("oss", {}))
return _backend.SelfHostedBackend(self._api_key, self._host) if self._host else _backend.PlatformBackend(self._api_key)
except Exception as e:
logger.error("Mem0 backend failed to initialize (%s mode): %s", self._mode, e)
self._init_error = str(e)
return None
def _is_breaker_open(self) -> bool:
"""True while the breaker is tripped; an expired cooldown resets the failure count."""
with self._breaker_lock:
if self._consecutive_failures >= _BREAKER_THRESHOLD and time.monotonic() < self._breaker_open_until:
return True
if self._consecutive_failures >= _BREAKER_THRESHOLD:
self._consecutive_failures = 0
return False
def _format_error(self, prefix: str, exc: Exception) -> str:
msg = f"{prefix}: {exc}"
if any(s in str(exc).lower() for s in ("connection", "refused", "timeout")):
msg += self._oss_hint(" (check that {vs} is running)")
return msg
def _record_success(self):
with self._breaker_lock:
self._consecutive_failures = 0
def _record_failure(self):
with self._breaker_lock:
self._consecutive_failures = count = self._consecutive_failures + 1
if count >= _BREAKER_THRESHOLD:
self._breaker_open_until = time.monotonic() + _BREAKER_COOLDOWN_SECS
if count >= _BREAKER_THRESHOLD:
hint = self._oss_hint(" Check that your {vs} vector store is running and reachable.", "unknown")
logger.warning("Mem0 circuit breaker tripped after %d consecutive failures. Pausing API calls for %ds.%s", count, _BREAKER_COOLDOWN_SECS, hint)
def _try(self, call, log, msg: str):
"""Background-path wrapper: run ``call`` under the breaker; on error log ``msg`` and return None."""
try:
result = call()
except Exception as e:
self._record_failure()
log(msg, e)
return None
self._record_success()
return result
def initialize(self, session_id: str, **kwargs) -> None:
self._config = cfg = _load_config()
self._mode, self._api_key, self._host, self._agent_id = cfg.get("mode", "platform"), cfg.get("api_key", ""), cfg.get("host", ""), cfg.get("agent_id", "hermes")
# user_id precedence: operator-configured (env/mem0.json) > gateway-native id (kwargs) > _DEFAULT_USER_ID.
# The literal placeholder counts as unset so wizard users still get gateway-native ids.
configured = cfg.get("user_id")
self._user_id = (None if configured == _DEFAULT_USER_ID else configured) or kwargs.get("user_id") or _DEFAULT_USER_ID
# Persisted rerank preference: default for mem0_search when the model omits ``rerank``. Platform-only.
_rr = cfg.get("rerank", False)
self._rerank_default = _rr.lower() in ("true", "1", "yes") if isinstance(_rr, str) else bool(_rr)
self._channel = kwargs.get("platform") or "cli"
try:
self._sync_max_chars = int(cfg.get("sync_max_chars") or _SYNC_MSG_MAX_CHARS)
except (ValueError, TypeError):
self._sync_max_chars = _SYNC_MSG_MAX_CHARS
self._backend = self._create_backend()
if self._backend and not self._atexit_registered:
atexit.register(self.shutdown)
self._atexit_registered = True
def _search(self, query: str, top_k: int = 10, rerank: bool = False, backend=None) -> list:
# Scoped to user_id only — by design — so recall surfaces memories from any gateway/agent under this
# principal; writes attach agent_id and metadata.channel so narrower views remain possible at query time.
return (backend or self._backend).search(query, filters={"user_id": self._user_id}, top_k=top_k, rerank=rerank)
def _add(self, messages: list, infer: bool):
metadata = {"channel": self._channel} if self._channel else {}
return self._backend.add(messages, user_id=self._user_id, agent_id=self._agent_id, infer=infer, metadata=metadata)
def system_prompt_block(self) -> str:
# Mirror _create_backend precedence (oss > host > platform). Rerank is a Mem0 Platform feature only.
mode_label = "OSS (self-hosted)" if self._mode == "oss" else "self-hosted (HTTP API)" if self._host else "platform (cloud API)"
rerank_note = " Rerank is available on search." if (self._mode == "platform" and not self._host) else ""
return f"# Mem0 Memory\nActive. Mode: {mode_label}. User: {self._user_id}.\n{_PROMPT_BODY}{rerank_note}"
def on_turn_start(self, turn_number: int, message: str, **kwargs) -> None:
self._start_prefetch(message)
def _consume_prefetch_result(self, query: str) -> str | None:
"""Pop the finished prefetch body for ``query`` (None if absent or still running)."""
with self._prefetch_lock:
if self._prefetch_query != query or not self._prefetch_done:
return None
result, self._prefetch_result, self._prefetch_done = self._prefetch_result, "", False
return result
def _start_prefetch(self, query: str) -> None:
backend = self._backend
if not query or backend is None or self._is_breaker_open():
return
def _run():
results = self._try(lambda: self._search(query, rerank=self._rerank_default, backend=backend), logger.debug, "Mem0 prefetch failed: %s")
lines = [r.get("memory", "") for r in (results or []) if r.get("memory")]
body = "## Mem0 Memory\n" + "\n".join(f"- {line}" for line in lines) if lines else ""
with self._prefetch_lock:
if self._prefetch_query == query:
self._prefetch_result, self._prefetch_done = body, True
with self._prefetch_lock:
if self._prefetch_query == query and (self._prefetch_done or (self._prefetch_thread and self._prefetch_thread.is_alive())):
return
self._prefetch_query, self._prefetch_result, self._prefetch_done = query, "", False
self._prefetch_thread = spawn_context_thread(_run, name="mem0-prefetch")
self._prefetch_thread.start()
def prefetch(self, query: str, *, session_id: str = "") -> str:
"""Recall memories for the CURRENT question with a short hot-path wait."""
if (cached := self._consume_prefetch_result(query)) is not None:
return cached
self._start_prefetch(query)
with self._prefetch_lock:
thread = self._prefetch_thread if self._prefetch_query == query else None
if thread:
thread.join(timeout=_PREFETCH_WAIT_SECS)
return self._consume_prefetch_result(query) or "" # slow backend: skip injection; mem0_search remains the backstop
def sync_turn(self, user_content: str, assistant_content: str, *, session_id: str = "") -> None:
"""Send the turn to Mem0 for server-side fact extraction (non-blocking)."""
if self._backend is None or self._is_breaker_open():
return
def _sync():
if self._backend is not None:
messages = [
{"role": "user", "content": _truncate_for_sync(user_content, self._sync_max_chars)},
{"role": "assistant", "content": _truncate_for_sync(assistant_content, self._sync_max_chars)},
]
self._try(lambda: self._add(messages, infer=True), logger.warning, "Mem0 sync failed: %s")
with self._sync_lock:
prev = self._sync_thread
if prev and prev.is_alive():
prev.join(timeout=5.0)
if prev.is_alive():
return
with self._sync_lock:
self._sync_thread = spawn_context_thread(_sync, name="mem0-sync")
self._sync_thread.start()
def get_tool_schemas(self) -> List[Dict[str, Any]]:
return list(TOOL_SCHEMAS)
# -- tool handlers: (required params, error label, body, client-error policy) ---
# Client errors (bad ID / not found) never trip the breaker, except for mem0_add
# where they count as failures; update/delete answer them with "Memory not found".
def _tool_search(self, args: dict) -> str:
top_k = max(1, min(int(args.get("top_k", 10)), 50))
rerank_raw = args.get("rerank", self._rerank_default)
rerank = rerank_raw.lower() not in ("false", "0", "no") if isinstance(rerank_raw, str) else bool(rerank_raw)
results = self._search(args["query"], top_k, rerank)
if not results:
return json.dumps({"result": "No relevant memories found."})
items = [{"id": r.get("id"), "memory": r.get("memory", ""), "score": r.get("score", 0)} for r in results]
return json.dumps({"results": items, "count": len(items)})
def _tool_add(self, args: dict) -> str:
result = self._add([{"role": "user", "content": args["content"]}], infer=False)
event_id = result.get("event_id") if isinstance(result, dict) else None
# Cloud add is async (server-side extraction); OSS and self-hosted store synchronously.
msg = "Fact stored." if (self._mode == "oss" or self._host) else "Fact queued for storage."
return json.dumps({"result": msg, "event_id": event_id})
def _ensure_owns_memory(self, memory_id: str) -> None:
"""Reject a mutation on a memory that doesn't belong to the caller's user_id."""
memory = self._backend.get(memory_id)
if not memory:
raise ValueError(f"Memory not found: {memory_id}")
owner = memory.get("user_id") if isinstance(memory, dict) else None
if owner and owner != self._user_id:
raise PermissionError(f"Memory {memory_id} does not belong to this user.")
def _tool_update(self, args: dict) -> str:
self._ensure_owns_memory(args["memory_id"])
return json.dumps(self._backend.update(args["memory_id"], args["text"]))
def _tool_delete(self, args: dict) -> str:
self._ensure_owns_memory(args["memory_id"])
return json.dumps(self._backend.delete(args["memory_id"]))
_TOOL_HANDLERS = {
"mem0_search": (("query",), "Search failed", _tool_search, "skip"),
"mem0_add": (("content",), "Failed to store", _tool_add, "count"),
"mem0_update": (("memory_id", "text"), "Update failed", _tool_update, "not_found"),
"mem0_delete": (("memory_id",), "Delete failed", _tool_delete, "not_found"),
}
def handle_tool_call(self, tool_name: str, args: dict, **kwargs) -> str:
if self._backend is None:
err = getattr(self, "_init_error", "unknown error")
return json.dumps({"error": f"Mem0 backend not initialized: {err}.{self._oss_hint(' Check that {vs} is running and reachable.')}"})
if self._is_breaker_open():
return json.dumps({"error": f"Mem0 temporarily unavailable (multiple consecutive failures). Will retry automatically.{self._oss_hint(' Check that your {vs} is running.')}"})
if tool_name not in self._TOOL_HANDLERS:
return tool_error(f"Unknown tool: {tool_name}")
required, label, body, on_client_error = self._TOOL_HANDLERS[tool_name]
if not isinstance(args, dict):
return tool_error("Tool arguments must be an object")
if missing := next((k for k in required if not isinstance(args.get(k), str) or not args[k].strip()), None):
return tool_error(f"Missing or invalid required parameter: {missing}")
if tool_name == "mem0_search":
try:
int(args.get("top_k", 10))
except (TypeError, ValueError, OverflowError):
return tool_error("top_k must be an integer")
try:
result = body(self, args)
except PermissionError as e:
return tool_error(str(e))
except Exception as e:
client = _is_client_error(e)
if client and on_client_error == "not_found":
return tool_error(f"Memory not found: {args['memory_id']}")
if not client or on_client_error == "count":
self._record_failure()
return tool_error(self._format_error(label, e))
self._record_success()
return result
def _shutdown_backend(self):
with suppress(Exception):
if self._backend:
self._backend.close()
self._backend = None
def shutdown(self) -> None:
for t in (self._prefetch_thread, self._sync_thread):
if t and t.is_alive():
# Extraction can outlast five seconds. Closing storage underneath it
# loses the turn; let the backend's network timeouts bound the drain.
t.join()
self._shutdown_backend()
def register(ctx) -> None:
"""Register Mem0 as a memory provider plugin."""
ctx.register_memory_provider(Mem0MemoryProvider())
# ---- BEGIN PLUGIN-COMPAT (revert-scheduled; see COMPAT_MANIFEST.md) ----
# Names external plugins imported from this module before the Sep 2026 decomposition.
# Internal code MUST NOT use these (scripts/check_compat_pointers.py fails CI if it does).
# The whole block is removed by reverting the commit that added it.
ADD_SCHEMA = {
"name": "mem0_add",
"description": (
"Store a durable fact about the user, verbatim (no LLM extraction). "
"Call this the moment the user states a lasting preference, correction, "
"decision, or personal detail worth recalling on future turns — don't "
"wait to be asked to remember. Skip transient chit-chat and facts you've "
"already stored."
),
"parameters": {
"type": "object",
"properties": {
"content": {"type": "string", "description": "The fact to store."},
},
"required": ["content"],
},
}
DELETE_SCHEMA = {
"name": "mem0_delete",
"description": (
"Delete a memory by its ID (take the ID from a mem0_search "
"result). Use when a stored fact is obsolete or the user asks you to "
"forget it; prefer mem0_update if the fact merely changed."
),
"parameters": {
"type": "object",
"properties": {
"memory_id": {"type": "string", "description": "Memory UUID to delete."},
},
"required": ["memory_id"],
},
}
SEARCH_SCHEMA = {
"name": "mem0_search",
"description": (
"Search the user's memories by meaning; returns facts ranked by "
"relevance. Use this before answering any question that may depend on "
"what you know about the user (preferences, facts, history, people, "
"projects, past decisions). For multi-part or multi-hop questions, "
"call it several times — vary the wording and run follow-up searches "
"on what earlier results reveal; one search is rarely enough."
),
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "What to search for."},
"top_k": {"type": "integer", "description": "Max results (default: 10, max: 50)."},
"rerank": {"type": "boolean", "description": "Rerank results for relevance (default: false, platform mode only)."},
},
"required": ["query"],
},
}
UPDATE_SCHEMA = {
"name": "mem0_update",
"description": (
"Replace the text of an existing memory by its ID (take the ID from a "
"mem0_search result). Use when a stored fact has changed "
"or was wrong — correct it in place instead of adding a duplicate."
),
"parameters": {
"type": "object",
"properties": {
"memory_id": {"type": "string", "description": "Memory UUID to update."},
"text": {"type": "string", "description": "New text content."},
},
"required": ["memory_id", "text"],
},
}
# ---- END PLUGIN-COMPAT ----
-336
View File
@@ -1,336 +0,0 @@
"""Backend abstraction for Mem0 Platform and OSS modes."""
from __future__ import annotations
import logging
import os
from abc import ABC, abstractmethod
from contextlib import closing, nullcontext, suppress
from copy import deepcopy
from dataclasses import dataclass, field
from threading import RLock
from typing import Any
logger = logging.getLogger(__name__)
def _add_kwargs(user_id: str, agent_id: str, infer: bool, metadata: dict | None) -> dict[str, Any]:
return {"user_id": user_id, "agent_id": agent_id, "infer": infer, **({"metadata": metadata} if metadata else {})}
def _unwrap_results(response: Any) -> list:
"""Normalize API response — extract results list from dict or pass through."""
return response.get("results", []) if isinstance(response, dict) else response if isinstance(response, list) else []
class Mem0Backend(ABC):
"""Unified interface over Platform (MemoryClient), self-hosted (HTTP) and OSS (Memory) backends.
update()/delete() are template methods: subclasses implement raw ``_update``/``_delete``."""
@abstractmethod
def search(self, query: str, *, filters: dict, top_k: int = 10, rerank: bool = False) -> list[dict]: ...
@abstractmethod
def add(self, messages: list, *, user_id: str, agent_id: str, infer: bool = False, metadata: dict | None = None) -> dict: ...
@abstractmethod
def get(self, memory_id: str) -> dict | None: ...
@abstractmethod
def _update(self, memory_id: str, text: str) -> None: ...
@abstractmethod
def _delete(self, memory_id: str) -> None: ...
def update(self, memory_id: str, text: str) -> dict:
self._update(memory_id, text)
return {"result": "Memory updated.", "memory_id": memory_id}
def delete(self, memory_id: str) -> dict:
self._delete(memory_id)
return {"result": "Memory deleted.", "memory_id": memory_id}
def close(self) -> None:
pass
class PlatformBackend(Mem0Backend):
"""Wraps mem0.MemoryClient for Mem0 Platform (cloud API)."""
def __init__(self, api_key: str):
from mem0 import MemoryClient
self._client = MemoryClient(api_key=api_key)
def search(self, query: str, *, filters: dict, top_k: int = 10, rerank: bool = False) -> list[dict]:
return _unwrap_results(self._client.search(query, filters=filters, top_k=top_k, rerank=rerank))
def add(self, messages: list, *, user_id: str, agent_id: str, infer: bool = False, metadata: dict | None = None) -> dict:
return self._client.add(messages, **_add_kwargs(user_id, agent_id, infer, metadata))
def get(self, memory_id: str) -> dict | None:
return self._client.get(memory_id)
def _update(self, memory_id: str, text: str) -> None:
self._client.update(memory_id=memory_id, text=text)
def _delete(self, memory_id: str) -> None:
self._client.delete(memory_id=memory_id)
class SelfHostedBackend(Mem0Backend):
"""Direct HTTP backend for a self-hosted Mem0 server (the FastAPI ``server/``).
mem0.MemoryClient is hardwired to the cloud API (``Authorization: Token``, ``GET /v1/ping/`` in ``__init__``),
so this speaks the server's real contract: ``X-API-Key`` auth and the ``/memories`` / ``/search`` routes."""
def __init__(self, api_key: str, host: str, transport=None):
import httpx
headers = {"Content-Type": "application/json", **({"X-API-Key": api_key} if api_key else {})} # key omitted only for AUTH_DISABLED servers
# Connect-level retries keep one dropped SYN from counting toward the breaker. ``transport`` is injectable for tests.
self._client = httpx.Client(base_url=host.rstrip("/"), headers=headers, timeout=30.0, transport=transport or httpx.HTTPTransport(retries=2))
self._capture_timeout = httpx.Timeout(120.0, connect=30.0)
def _json(self, method: str, path: str, **kwargs) -> Any:
resp = self._client.request(method, path, **kwargs)
resp.raise_for_status()
return resp.json() if resp.content else {}
def search(self, query: str, *, filters: dict, top_k: int = 10, rerank: bool = False) -> list[dict]:
# rerank is platform-only; the self-hosted /search ignores it. user_id belongs in filters (top-level is deprecated).
return _unwrap_results(self._json("POST", "/search", json={"query": query, "top_k": top_k, **({"filters": filters} if filters else {})}))
def add(self, messages: list, *, user_id: str, agent_id: str, infer: bool = False, metadata: dict | None = None) -> dict:
# Server-side extraction takes longer than a search or verbatim write.
return self._json("POST", "/memories", json={"messages": messages, **_add_kwargs(user_id, agent_id, infer, metadata)},
timeout=self._capture_timeout if infer else self._client.timeout)
def get(self, memory_id: str) -> dict | None:
return self._json("GET", f"/memories/{memory_id}")
def _update(self, memory_id: str, text: str) -> None:
self._json("PUT", f"/memories/{memory_id}", json={"text": text})
def _delete(self, memory_id: str) -> None:
self._json("DELETE", f"/memories/{memory_id}")
def close(self) -> None:
with suppress(Exception):
self._client.close()
_DIRECT_OPENAI_PROVIDER = "hermes_openai"
_DIRECT_OPENAI_CLASS_PATH = f"{__package__}._openai_llm.DirectOpenAILLM"
@dataclass
class _LocalQdrantMemory:
memory: Any
config: dict
profile: str
lock: Any = field(default_factory=RLock)
users: int = 1
_LOCAL_QDRANT_MEMORIES: dict[str, _LocalQdrantMemory] = {}
_LOCAL_QDRANT_LOCK = RLock()
def _register_direct_openai_provider() -> None:
"""Register Hermes' OpenAI-only Mem0 LLM provider once per factory."""
from mem0.configs.llms.openai import OpenAIConfig
from mem0.utils.factory import LlmFactory
provider_map = getattr(LlmFactory, "provider_to_class", None)
register_provider = getattr(LlmFactory, "register_provider", None)
if not isinstance(provider_map, dict) or not callable(register_provider):
raise RuntimeError("mem0 LlmFactory does not support the provider registration required for the Hermes OpenAI OSS backend")
if provider_map.get(_DIRECT_OPENAI_PROVIDER) != (_DIRECT_OPENAI_CLASS_PATH, OpenAIConfig):
register_provider(_DIRECT_OPENAI_PROVIDER, _DIRECT_OPENAI_CLASS_PATH, OpenAIConfig)
class OSSBackend(Mem0Backend):
"""Wraps mem0.Memory for self-hosted (OSS) mode."""
def __init__(self, oss_config: dict):
from ._oss_providers import EMBEDDER_PROVIDERS, KNOWN_DIMS, LLM_PROVIDERS
self._local_path = None
self._owner = None
self._lock = nullcontext()
self._closed = False
def _provider_block(name: str, registry: dict) -> dict:
"""Copy of oss_config[name] with the legacy ``api_base`` key mapped to the provider's canonical base-URL key."""
block = dict(oss_config[name])
provider_config = dict(block.get("config", {}))
legacy_base = provider_config.pop("api_base", None)
canonical_key = registry.get(str(block.get("provider") or "").strip().lower(), {}).get("base_url_key")
if legacy_base and canonical_key:
provider_config.setdefault(canonical_key, legacy_base)
if str(block.get("provider") or "").strip().lower() == "openai":
from agent.secret_scope import get_secret
# Resolve profile secrets before comparing configurations for sharing.
provider_config["api_key"] = provider_config.get("api_key") or get_secret("OPENAI_API_KEY", "")
if not provider_config["api_key"]:
raise ValueError(f"OpenAI API key is required for the Hermes Mem0 OSS {name}")
provider_config["openai_base_url"] = (
provider_config.get("openai_base_url") or get_secret("OPENAI_API_BASE", "")
or get_secret("OPENAI_BASE_URL", "") or "https://api.openai.com/v1"
)
block["config"] = provider_config
return block
vector_store = dict(oss_config["vector_store"])
vs_config = dict(vector_store.get("config", {}))
if vs_config.get("path"):
vs_config["path"] = os.path.expanduser(vs_config["path"])
embedder_config = oss_config.get("embedder", {}).get("config", {})
dims = embedder_config.get("embedding_dims") or KNOWN_DIMS.get(embedder_config.get("model", ""))
if dims:
vs_config["embedding_model_dims"] = dims
remote = (vs_config.get("host") and vs_config.get("port")) or vs_config.get("url") or vs_config.get("api_key")
if (vector_store.get("provider", "qdrant") == "qdrant" and not vs_config.get("client")
and not remote and vs_config.get("https") is None):
from mem0.configs.vector_stores.qdrant import QdrantConfig
path = vs_config.get("path", QdrantConfig.model_fields["path"].default)
if path:
self._local_path = vs_config["path"] = os.path.realpath(os.path.expanduser(path))
vector_store["config"] = vs_config
config = {"vector_store": vector_store, "llm": _provider_block("llm", LLM_PROVIDERS), "embedder": _provider_block("embedder", EMBEDDER_PROVIDERS), "version": "v1.1"}
if self._local_path:
from hermes_constants import get_hermes_home
profile = os.path.realpath(get_hermes_home())
with _LOCAL_QDRANT_LOCK:
owner = _LOCAL_QDRANT_MEMORIES.get(self._local_path)
if owner is None:
owner = _LocalQdrantMemory(self._create_memory(config, dims), deepcopy(config), profile)
_LOCAL_QDRANT_MEMORIES[self._local_path] = owner
else:
if owner.profile != profile or owner.config != config:
raise ValueError("Local Qdrant storage is already open with a different profile or configuration. "
"Existing memories were preserved. Close its active sessions before changing settings, "
"or use a separate storage path.")
owner.users += 1
self._owner, self._lock, self._memory = owner, owner.lock, owner.memory
else:
self._memory = self._create_memory(config, dims)
@staticmethod
def _create_memory(config: dict, dims: int | None):
from mem0 import Memory
vector_store = config["vector_store"]
vs_config = vector_store["config"]
if dims:
OSSBackend._reject_dimension_mismatch(vector_store.get("provider", "qdrant"), vs_config, dims)
else:
logger.warning("Unknown embedding dimensions; skipping dimension-change guard for collection %r.",
vs_config.get("collection_name", "mem0"))
if str(config["llm"].get("provider") or "").strip().lower() == "openai":
# mem0 validates LlmConfig.provider before its factory lookup: build the supported OpenAI config, then swap the provider.
_register_direct_openai_provider()
from mem0.configs.base import MemoryConfig
memory_config = MemoryConfig(**config)
try:
memory_config.llm.provider = _DIRECT_OPENAI_PROVIDER
except (AttributeError, TypeError) as exc:
raise RuntimeError("mem0 MemoryConfig does not expose a mutable llm.provider for the Hermes OpenAI OSS backend") from exc
return Memory(memory_config)
return Memory.from_config(config)
@staticmethod
def _detect_current_dims(provider: str, vs_config: dict, collection_name: str) -> int | None:
"""Current embedding dimension of ``collection_name``, or None if it doesn't exist yet.
Raises on any failure to connect/inspect so the caller can decide whether to skip the guard."""
if provider == "qdrant":
from qdrant_client import QdrantClient
path, url, host = vs_config.get("path"), vs_config.get("url"), vs_config.get("host")
if path:
client = QdrantClient(path=path)
elif url:
client = QdrantClient(url=url, api_key=vs_config.get("api_key"))
elif host:
client = QdrantClient(host=host, port=vs_config.get("port") or 6333, api_key=vs_config.get("api_key"))
else:
return None
with closing(client):
if not client.collection_exists(collection_name):
return None
vectors = client.get_collection(collection_name).config.params.vectors
# Named-vector collections expose a dict; unnamed expose an object with .size.
if isinstance(vectors, dict):
vectors = next(iter(vectors.values()), None)
return getattr(vectors, "size", None)
elif provider == "pgvector":
import psycopg2
conn_params = {k: vs_config[k] for k in ("host", "port", "user", "password", "dbname", "sslmode") if vs_config.get(k)}
with closing(psycopg2.connect(**conn_params)) as conn:
conn.autocommit = True
with closing(conn.cursor()) as cur:
cur.execute("SELECT atttypmod FROM pg_attribute WHERE attrelid = %s::regclass AND attname = 'vector'", (collection_name,))
row = cur.fetchone()
return row[0] if row and row[0] > 0 else None
return None
@staticmethod
def _reject_dimension_mismatch(provider: str, vs_config: dict, expected_dims: int) -> None:
"""Reject embedding dimension changes without deleting existing memories."""
collection_name = vs_config.get("collection_name", "mem0")
try:
current_dims = OSSBackend._detect_current_dims(provider, vs_config, collection_name)
except Exception as dimension_detection_error:
logger.warning(
"Could not determine embedding dimensions for collection %r (%s): %s. Skipping dimension-change guard.",
collection_name, provider, dimension_detection_error,
)
return
if current_dims is not None and current_dims != expected_dims:
raise ValueError(
f"Collection {collection_name!r} has {current_dims} embedding dimensions, but {expected_dims} are configured. "
"Existing memories were preserved. Restore the previous embedder or use a new collection_name."
)
def search(self, query: str, *, filters: dict, top_k: int = 10, rerank: bool = False) -> list[dict]:
return _unwrap_results(self._call("search", query, filters=filters, top_k=top_k))
def add(self, messages: list, *, user_id: str, agent_id: str, infer: bool = False, metadata: dict | None = None) -> dict:
return self._call("add", messages, **_add_kwargs(user_id, agent_id, infer, metadata))
def get(self, memory_id: str) -> dict | None:
return self._call("get", memory_id)
def _update(self, memory_id: str, text: str) -> None:
self._call("update", memory_id, data=text)
def _delete(self, memory_id: str) -> None:
self._call("delete", memory_id)
def _call(self, method, *args, **kwargs):
# ponytail: serialize whole local SDK operations, including extraction's read/modify/write.
# Use a Qdrant server for parallel throughput or access from multiple processes.
with self._lock:
if self._closed:
raise RuntimeError("Mem0 backend is closed")
return getattr(self._memory, method)(*args, **kwargs)
def close(self):
with self._lock:
if self._closed:
return
self._closed = True
if self._owner:
with _LOCAL_QDRANT_LOCK:
self._owner.users -= 1
if self._owner.users == 0:
self._close_memory()
del _LOCAL_QDRANT_MEMORIES[self._local_path]
else:
self._close_memory()
def _close_memory(self):
with suppress(Exception):
telemetry = getattr(self._memory, "telemetry", None)
if telemetry and hasattr(telemetry, "posthog"):
with suppress(Exception):
telemetry.posthog.shutdown()
vs = getattr(self._memory, "vector_store", None)
telemetry_vs = getattr(self._memory, "_telemetry_vector_store", None)
resources = (self._memory, vs, getattr(vs, "client", None), getattr(telemetry_vs, "client", None))
for obj in {id(obj): obj for obj in resources if obj is not None}.values():
if hasattr(obj, "close"):
with suppress(Exception):
obj.close()
@@ -1,66 +0,0 @@
"""OpenAI-only LLM adapter for Mem0 OSS mode."""
from __future__ import annotations
import logging
from typing import Dict, List, Optional, Union
from mem0.configs.llms.base import BaseLlmConfig
from mem0.configs.llms.openai import OpenAIConfig
from mem0.llms.base import LLMBase
from mem0.llms.openai import OpenAILLM
# BaseLlmConfig fields copied into OpenAIConfig; the last two may be absent on older mem0.
_COPIED_FIELDS = ("model", "temperature", "api_key", "max_tokens", "top_p", "top_k", "enable_vision", "vision_details", "http_client_proxies")
_OPTIONAL_FIELDS = ("reasoning_effort", "is_reasoning_model")
class DirectOpenAILLM(OpenAILLM):
"""Use OpenAI credentials and requests regardless of router environment."""
def __init__(self, config: Optional[Union[BaseLlmConfig, OpenAIConfig, Dict]] = None):
if config is None:
config = OpenAIConfig()
elif isinstance(config, dict):
config = OpenAIConfig(**config)
elif isinstance(config, BaseLlmConfig) and not isinstance(config, OpenAIConfig):
fields = {k: getattr(config, k) for k in _COPIED_FIELDS}
fields.update({k: getattr(config, k, None) for k in _OPTIONAL_FIELDS})
config = OpenAIConfig(**fields)
if not config.model:
config.model = "gpt-5-mini"
# Configs predating the setup marker: keep the default model reasoning-safe
# without overriding an explicit user choice.
if config.model == "gpt-5-mini" and config.is_reasoning_model is None:
config.is_reasoning_model = True
# Bypass OpenAILLM.__init__ (it picks OpenRouter when OPENROUTER_API_KEY is
# set); LLMBase still owns validation and supported-parameter filtering.
LLMBase.__init__(self, config)
# OPENAI_API_KEY / OPENAI_BASE_URL are profile credentials: read them through the secret
# scope, never raw os.environ, or a multiplexed secondary's memory extraction runs on the
# default profile's OpenAI account (and its proxy).
from agent.secret_scope import get_secret
api_key = self.config.api_key or get_secret("OPENAI_API_KEY", "")
if not api_key:
raise ValueError("OpenAI API key is required for the Hermes Mem0 OSS provider")
from openai import OpenAI
self.client = OpenAI(api_key=api_key, base_url=self.config.openai_base_url or get_secret("OPENAI_API_BASE", "") or get_secret("OPENAI_BASE_URL", "") or "https://api.openai.com/v1", timeout=120.0)
def generate_response(self, messages: List[Dict[str, str]], response_format=None, tools: Optional[List[Dict]] = None, tool_choice: str = "auto", **kwargs):
params = self._get_supported_params(messages=messages, **kwargs)
params.update({"model": self.config.model, "messages": messages})
# No OpenRouter-only fields; ``store`` is opt-in so OpenAI-compatible endpoints never receive unknown fields.
if self.config.store is not None:
params["store"] = self.config.store
if response_format:
params["response_format"] = response_format
if tools:
params["tools"], params["tool_choice"] = tools, tool_choice
response = self.client.chat.completions.create(**params)
parsed_response = self._parse_response(response, tools)
if self.config.response_callback:
try:
self.config.response_callback(self, response, params)
except Exception:
logging.error("Error running Mem0 OpenAI response callback")
return parsed_response
@@ -1,53 +0,0 @@
"""OSS provider definitions for LLM, embedder, and vector store."""
from __future__ import annotations
import os
from typing import Any
from hermes_constants import get_hermes_home
LLM_PROVIDERS: dict[str, dict[str, Any]] = {
"openai": {"label": "OpenAI", "needs_key": True, "env_var": "OPENAI_API_KEY", "default_model": "gpt-5-mini", "base_url_key": "openai_base_url"},
"ollama": {"label": "Ollama (local)", "needs_key": False, "default_model": "llama3.1:8b", "default_url": "http://localhost:11434", "base_url_key": "ollama_base_url", "pip_dep": "ollama"},
}
EMBEDDER_PROVIDERS: dict[str, dict[str, Any]] = {
"openai": {"label": "OpenAI", "needs_key": True, "env_var": "OPENAI_API_KEY", "default_model": "text-embedding-3-small", "base_url_key": "openai_base_url", "dims": 1536},
"ollama": {"label": "Ollama (local)", "needs_key": False, "default_model": "nomic-embed-text", "default_url": "http://localhost:11434", "base_url_key": "ollama_base_url", "dims": 768, "pip_dep": "ollama"},
}
VECTOR_PROVIDERS: dict[str, dict[str, Any]] = {
# Resolved lazily (see ``vector_default_config``): the profile home is a ContextVar at call time,
# not an import-time constant, and ``~/.hermes`` is wrong on Windows and under profiles.
"qdrant": {"label": "Qdrant", "default_config": {"path": lambda: str(get_hermes_home() / "mem0_qdrant")}, "pip_dep": "qdrant-client"},
"pgvector": {
"label": "PGVector",
"default_config": {"host": "localhost", "port": 5432, "user": os.getenv("USER", "postgres"), "dbname": "postgres"},
"pip_dep": "psycopg2-binary",
},
}
KNOWN_DIMS: dict[str, int] = {"text-embedding-3-small": 1536, "text-embedding-3-large": 3072, "text-embedding-ada-002": 1536, "nomic-embed-text": 768}
def vector_default_config(provider_id: str) -> dict[str, Any]:
"""A vector store's ``default_config`` with callable defaults resolved for the active profile."""
return {k: (v() if callable(v) else v) for k, v in VECTOR_PROVIDERS[provider_id]["default_config"].items()}
SECTION_REGISTRIES = (("llm", LLM_PROVIDERS), ("embedder", EMBEDDER_PROVIDERS), ("vector_store", VECTOR_PROVIDERS))
def validate_oss_config(oss_config: dict) -> list[str]:
"""Validate an OSS config dict. Returns list of error strings (empty = valid)."""
errors: list[str] = []
for section, registry in SECTION_REGISTRIES:
block = oss_config.get(section)
if not block or not isinstance(block, dict):
errors.append(f"Missing required section: {section}")
elif block.get("provider", "") not in registry:
errors.append(f"Unknown {section} provider '{block.get('provider', '')}'. Valid: {', '.join(registry.keys())}")
vs = oss_config.get("vector_store", {})
if vs.get("provider") == "pgvector" and not vs.get("config", {}).get("user"):
errors.append("PGVector requires 'user' in vector_store.config")
return errors
-567
View File
@@ -1,567 +0,0 @@
"""Setup wizard for Mem0 plugin — interactive and flag-based modes."""
from __future__ import annotations
import getpass
import json
import os
import re
import secrets
import shutil
import socket
import subprocess
import sys
import time
import urllib.error
import urllib.request
from contextlib import suppress
from pathlib import Path
from typing import Any
from hermes_constants import get_hermes_home # noqa: F401 — patched by tests
from ._oss_providers import (
EMBEDDER_PROVIDERS,
KNOWN_DIMS,
LLM_PROVIDERS,
SECTION_REGISTRIES,
VECTOR_PROVIDERS,
validate_oss_config,
vector_default_config,
)
_OLLAMA_URL = "http://localhost:11434"
_PGVECTOR_CONTAINER, _PGVECTOR_IMAGE = "hermes-pgvector", "pgvector/pgvector:pg17"
def _version_tuple(version: str) -> tuple[int, ...]:
"""Best-effort (major, minor, patch) from a version string, tolerating pre-release suffixes like '2.1.0rc1'."""
parts = version.split(".")[:3]
return tuple(int(m.group()) if (m := re.match(r"\d+", p)) else 0 for p in parts)
def _scrub(text: str, *secrets_to_hide: str) -> str:
"""Replace any occurrence of the given secrets in ``text`` (e.g. before printing a subprocess error)."""
for secret in secrets_to_hide:
if secret:
text = text.replace(secret, "***")
return text
def _curses_select(title: str, items: list[tuple[str, str]], default: int = 0) -> int:
from hermes_cli.curses_ui import curses_radiolist
return curses_radiolist(title, [f"{label} {desc}" if desc else label for label, desc in items], selected=default, cancel_returns=default)
def _prompt(label: str, default: str | None = None, secret: bool = False) -> str:
"""Prompt for a value with optional default and secret masking."""
sys.stdout.write(f" {label}{f' [{default}]' if default else ''}: ")
sys.stdout.flush()
val = getpass.getpass(prompt="") if secret and sys.stdin.isatty() else sys.stdin.readline().strip()
return val or (default or "")
def _input(label: str, default: str) -> str:
return input(f" {label} [{default}]: ").strip() or default
def _masked(secret: str) -> str:
return f"...{secret[-4:]}" if len(secret) > 4 else "set"
def _http_get(url: str, path: str, timeout: int):
return urllib.request.urlopen(urllib.request.Request(f"{url.rstrip('/')}{path}", method="GET"), timeout=timeout)
def _prompt_api_key(label: str, env_var: str, hermes_home: str) -> str:
"""Prompt for API key, showing masked existing value if found."""
existing = os.environ.get(env_var, "")
if not existing:
from agent.secret_scope import load_env_file
existing = load_env_file(Path(hermes_home) / ".env").get(env_var, "")
hint = f" (current: {_masked(existing)}, blank to keep)" if existing else ""
return getpass.getpass(f" {label} API key{hint}: ").strip()
def _api_key_writes(flags: dict, label: str, *, url: str | None = None, fresh_label: str | None = None) -> dict[str, str]:
"""MEM0_API_KEY for .env: from --api-key, else prompt (masking any key already in the environment)."""
if flags.get("api_key"):
return {"MEM0_API_KEY": flags["api_key"]}
existing = os.environ.get("MEM0_API_KEY", "")
if url and not existing:
print(f" Get yours at {url}")
val = _prompt(f"{label} (current: {_masked(existing)}, blank to keep)" if existing else fresh_label or label, secret=True)
return {"MEM0_API_KEY": val} if val else {}
def _print_dry_run(summary: str, env_writes: dict, check=None) -> None:
print(f"\n [dry-run] Would save config: {summary}")
if env_writes:
print(" [dry-run] Would write API key to .env")
if check:
check()
print(" [dry-run] No files written.\n")
# --oss-vector-<key> flags accepted per vector store (also the pgvector key order).
_VECTOR_FLAG_KEYS = {"qdrant": ("path", "url"), "pgvector": ("host", "port", "user", "password", "dbname")}
_FLAG_KEYS = ("mode", "api_key", "host", *(f"oss_{s}{k}" for s in ("llm", "embedder") for k in ("", "_key", "_model", "_url")),
"oss_vector", *(f"oss_vector_{k}" for ks in _VECTOR_FLAG_KEYS.values() for k in ks), "user_id")
_FLAG_DEFAULTS = {"oss_llm": "openai", "oss_embedder": "openai", "oss_vector": "qdrant"}
def parse_flags(argv: list[str] | None = None) -> dict[str, str]:
args = argv if argv is not None else sys.argv[1:]
flags: dict[str, Any] = {**{k: _FLAG_DEFAULTS.get(k, "") for k in _FLAG_KEYS}, "dry_run": False}
flag_map = {"--" + k.replace("_", "-"): k for k in _FLAG_KEYS}
i = 0
while i < len(args):
if args[i] == "--dry-run":
flags["dry_run"] = True
elif args[i] in flag_map and i + 1 < len(args):
flags[flag_map[args[i]]] = args[i + 1]
i += 1
i += 1
return flags
def _model_block(flags: dict, registry: dict, prefix: str) -> tuple[str, dict, dict[str, Any]]:
"""Resolve (provider_id, provider_def, config) for an LLM/embedder section from flags."""
pid = flags.get(prefix, "openai")
pdef = registry[pid]
cfg: dict[str, Any] = {"model": flags.get(f"{prefix}_model") or pdef["default_model"]}
url = flags.get(f"{prefix}_url") or pdef.get("default_url")
if url and pdef.get("base_url_key"):
cfg[pdef["base_url_key"]] = url
return pid, pdef, cfg
def build_oss_config(flags: dict[str, str]) -> tuple[dict, dict[str, str]]:
"""Build (oss_config for mem0.json, env_writes of secrets for .env) from parsed flags."""
llm_id, llm_def, llm_config = _model_block(flags, LLM_PROVIDERS, "oss_llm")
if llm_id == "openai" and llm_config["model"] == "gpt-5-mini":
llm_config["is_reasoning_model"] = True
embedder_id, embedder_def, embedder_config = _model_block(flags, EMBEDDER_PROVIDERS, "oss_embedder")
dims = KNOWN_DIMS.get(embedder_config["model"])
if dims:
embedder_config["embedding_dims"] = dims
vector_id = flags.get("oss_vector", "qdrant")
vector_config = vector_default_config(vector_id)
for key in _VECTOR_FLAG_KEYS.get(vector_id, ()):
if val := flags.get(f"oss_vector_{key}"):
if key == "port" and not val.isdigit():
raise ValueError(f"--oss-vector-port must be a number, got {val!r}")
vector_config[key] = int(val) if key == "port" else val
if "url" in vector_config:
vector_config.pop("path", None) # a remote Qdrant URL replaces local storage
oss_config = {"llm": {"provider": llm_id, "config": llm_config}, "embedder": {"provider": embedder_id, "config": embedder_config}, "vector_store": {"provider": vector_id, "config": vector_config}}
# An embedder sharing the LLM's provider reuses the LLM key when no embedder key was given.
llm_key = flags.get("oss_llm_key") if llm_def.get("needs_key") else ""
emb_key = (flags.get("oss_embedder_key") or (flags.get("oss_llm_key") if embedder_id == llm_id else "")) if embedder_def.get("needs_key") else ""
env_writes = {d["env_var"]: k for d, k in ((llm_def, llm_key), (embedder_def, emb_key)) if k}
if llm_key and emb_key and llm_key != emb_key and llm_def["env_var"] == embedder_def["env_var"]:
# One environment variable cannot hold two accounts; save explicit keys in the private config.
llm_config["api_key"], embedder_config["api_key"] = llm_key, emb_key
env_writes.pop(llm_def["env_var"])
return oss_config, env_writes
def _write_env(env_path: Path, env_writes: dict[str, str]) -> None:
env_path.parent.mkdir(parents=True, exist_ok=True)
# utf-8-sig like the canonical .env readers: a BOM'd first line would miss the key match and get duplicated.
existing_lines = env_path.read_text(encoding="utf-8-sig").splitlines() if env_path.exists() else []
keys = [line.split("=", 1)[0].strip() if "=" in line and not line.startswith("#") else None for line in existing_lines]
new_lines = [f"{k}={env_writes[k]}" if k in env_writes else line for k, line in zip(keys, existing_lines)]
new_lines += [f"{k}={v}" for k, v in env_writes.items() if k not in keys]
from utils import atomic_write_text
atomic_write_text(env_path, "\n".join(new_lines) + "\n", mode=0o600)
def _activate_provider(config: dict) -> None:
"""Point config.yaml's memory.provider at mem0."""
from hermes_cli.config import save_config
config["memory"]["provider"] = "mem0"
save_config(config)
def _persist_provider_config(hermes_home: str, config: dict, provider_config: dict, env_writes: dict[str, str], label: str, key_line: str, server: str | None = None) -> None:
"""Shared platform/self-hosted tail: activate, write mem0.json (0600), then .env, then a saved summary."""
_activate_provider(config)
from . import Mem0MemoryProvider
if "MEM0_API_KEY" in env_writes:
provider_config["api_key"] = "" # Let the newly saved .env key replace legacy inline credentials.
Mem0MemoryProvider().save_config(provider_config, hermes_home)
if env_writes:
_write_env(Path(hermes_home) / ".env", env_writes)
if server:
_check_selfhosted_server(server)
print("\n".join(["", f" Memory provider: {label}", *([f" Server: {server}"] if server else []), " Activation saved to config.yaml", " Provider config saved",
*([f" {key_line}"] if env_writes else []), "", " Start a new session to activate.", ""]))
def _setup_platform(hermes_home: str, config: dict, flags: dict[str, str]) -> None:
"""Platform mode setup — prompts for API key (secret -> .env), user/agent ids and rerank (-> mem0.json)."""
from utils import read_json_or_empty
provider_config = read_json_or_empty(Path(hermes_home) / "mem0.json")
print("\n Configuring mem0:\n")
env_writes = _api_key_writes(flags, "Mem0 Platform API key", url="https://app.mem0.ai")
for key, desc, default in (("user_id", "User identifier", "hermes-user"), ("agent_id", "Agent identifier", "hermes")):
if val := flags.get(key) or _prompt(desc, default=str(provider_config.get(key) or default)):
provider_config[key] = val
choices = ["true", "false"]
current = str(provider_config.get("rerank", "false") or "").lower()
provider_config["rerank"] = choices[_curses_select(" Enable reranking for recall", [(c, "") for c in choices], default=choices.index(current) if current in choices else 0)]
if flags.get("dry_run"):
_print_dry_run(str({k: provider_config.get(k) for k in ("user_id", "agent_id", "rerank")}), env_writes)
return
# Routing checks ``host`` before platform, so clear a stale self-hosted host. "" rather than
# pop(): save_config merges into the existing mem0.json, so a popped key would survive.
provider_config.update(mode="platform", host="")
# _load_config() also seeds ``host`` from MEM0_HOST (.env); the file clear can't help there, so warn.
if os.environ.get("MEM0_HOST", "").strip():
print(f"\n ⚠ MEM0_HOST is set in your environment ({os.environ['MEM0_HOST']}). It overrides platform mode — remove it from ~/.hermes/.env (or unset it) or Hermes will keep routing to the self-hosted server.")
_persist_provider_config(hermes_home, config, provider_config, env_writes, "mem0", "API keys saved to .env")
def _check_selfhosted_server(host: str) -> None:
"""Best-effort reachability check for a self-hosted Mem0 server (non-fatal)."""
try:
_http_get(host, "/docs", 5)
print(f" ✓ Mem0 server reachable at {host}")
except urllib.error.HTTPError:
# Any HTTP response (401/403/404) still means something is listening.
print(f" ✓ Mem0 server responding at {host}")
except Exception:
print(f" ⚠ Could not reach {host} — check the URL and that the server is running.")
def _setup_selfhosted(hermes_home: str, config: dict, flags: dict[str, str]) -> None:
"""Self-hosted mode — point at an existing Mem0 server: URL -> mem0.json, key -> .env (MEM0_API_KEY)."""
from utils import read_json_or_empty
provider_config = read_json_or_empty(Path(hermes_home) / "mem0.json")
print("\n Configuring mem0 (self-hosted server):\n")
host = flags.get("host") or _prompt("Mem0 server URL (e.g. http://localhost:8888)", default=provider_config.get("host") or None)
if not host:
print(" Error: a server URL is required for self-hosted mode.", file=sys.stderr)
return
host = host.rstrip("/")
env_writes = _api_key_writes(flags, "Server API key", fresh_label="Server API key (blank if AUTH_DISABLED)")
user_id = flags.get("user_id") or _prompt("User identifier", default=provider_config.get("user_id") or "hermes-user")
agent_id = _prompt("Agent identifier", default=provider_config.get("agent_id") or "hermes")
if flags.get("dry_run"):
_print_dry_run(f"host={host}, user_id={user_id}, agent_id={agent_id}", env_writes, lambda: _check_selfhosted_server(host))
return
provider_config.update(mode="platform", host=host, user_id=user_id, agent_id=agent_id) # routing: oss > host > platform
_persist_provider_config(hermes_home, config, provider_config, env_writes, "mem0 (self-hosted)", "API key saved to .env", server=host)
def _print_oss_summary(oss_config: dict, env_writes: dict, dry_run: bool = False) -> None:
llm, emb = oss_config["llm"], oss_config["embedder"]
w = 0 if dry_run else 9 # final summary column-aligns the labels
lines = ["", " [dry-run] OSS config would be:" if dry_run else " ✓ Mem0 configured (OSS mode)",
f" {'LLM:':<{w}} {llm['provider']} ({llm['config'].get('model', '')})", f" {'Embedder:':<{w}} {emb['provider']} ({emb['config'].get('model', '')})",
f" {'Vector:':<{w}} {oss_config['vector_store']['provider']}"]
if dry_run:
lines += [f" Env vars: {', '.join(env_writes.keys())}"] if env_writes else []
else:
lines += [*([" API keys saved to .env"] if env_writes else []), " Config saved to mem0.json", " Provider set in config.yaml", "", " Start a new session to activate.", ""]
print("\n".join(lines))
def _finish_oss(hermes_home: str, config: dict, oss_config: dict, env_writes: dict[str, str], user_id: str, agent_id: str, pgvector_config: dict | None = None) -> None:
"""Shared OSS tail: write secrets + mem0.json, install deps, activate, check, summarize."""
from . import Mem0MemoryProvider
if env_writes:
_write_env(Path(hermes_home) / ".env", env_writes)
Mem0MemoryProvider().save_config(
{"mode": "oss", "user_id": user_id, "agent_id": agent_id, "oss": oss_config}, hermes_home
)
_install_provider_deps(oss_config["llm"]["provider"], oss_config["embedder"]["provider"], oss_config["vector_store"]["provider"])
if pgvector_config:
_ensure_pgvector_extension(pgvector_config)
_activate_provider(config)
_run_connectivity_checks(oss_config)
_print_oss_summary(oss_config, env_writes)
def _setup_oss(hermes_home: str, config: dict, flags: dict[str, str]) -> None:
"""OSS mode — non-interactive when --mode was given, otherwise curses pickers."""
if not flags.get("_mode_from_flag"):
_setup_oss_interactive(hermes_home, config)
return
try:
oss_config, env_writes = build_oss_config(flags)
except ValueError as e:
print(f" Error: {e}", file=sys.stderr)
sys.exit(1)
if errors := validate_oss_config(oss_config):
print("".join(f" Error: {e}\n" for e in errors), end="", file=sys.stderr)
sys.exit(1)
if flags.get("dry_run"):
_print_oss_summary(oss_config, env_writes, dry_run=True)
_run_connectivity_checks(oss_config)
print(" [dry-run] No files written.\n")
return
_finish_oss(hermes_home, config, oss_config, env_writes, flags.get("user_id") or os.getenv("USER", "hermes-user"), "hermes")
def _docker(*args: str, timeout: int, **kwargs) -> subprocess.CompletedProcess:
return subprocess.run(["docker", *args], capture_output=True, timeout=timeout, stdin=subprocess.DEVNULL, **kwargs)
def _pg_ready(host: str, port: int, wait: int) -> bool:
"""Wait up to ``wait`` seconds for the port, then report whether PostgreSQL answers."""
_wait_for_port(host, port, timeout=wait)
return _check_pgvector(host, port)[0]
def _ensure_pgvector(host: str = "localhost", port: int = 5432) -> dict | None:
"""Ensure pgvector is reachable, offering Docker if not; returns the started container's vector_config, else None."""
if _check_pgvector(host, port)[0]:
print(f" ✓ PostgreSQL reachable at {host}:{port}")
return None
print(f" PostgreSQL not reachable at {host}:{port}")
if not shutil.which("docker"):
print(" Docker not found. Install Docker to auto-start pgvector,\n or run PostgreSQL with pgvector manually.")
return None
with suppress(Exception): # restart our own container if it exists but is stopped
result = _docker("inspect", _PGVECTOR_CONTAINER, "--format", "{{.State.Status}}", timeout=10, text=True, encoding='utf-8', errors='replace')
if result.returncode == 0 and "exited" in result.stdout:
print(f" Found stopped container '{_PGVECTOR_CONTAINER}', restarting...")
_docker("start", _PGVECTOR_CONTAINER, timeout=15)
if _pg_ready(host, port, 15):
print(" ✓ PostgreSQL container restarted")
return None
if input(" Start pgvector via Docker? [Y/n]: ").strip().lower() not in ("", "y", "yes"):
print(" Skipping Docker setup. Make sure PostgreSQL with pgvector is running.")
return None
password = secrets.token_urlsafe(24)
try:
print(f" Pulling {_PGVECTOR_IMAGE}...")
_docker("pull", _PGVECTOR_IMAGE, timeout=120)
print(f" Starting container '{_PGVECTOR_CONTAINER}' on port {port}...")
_docker("run", "-d", "--name", _PGVECTOR_CONTAINER, "-e", f"POSTGRES_PASSWORD={password}", "-p", f"127.0.0.1:{port}:5432", _PGVECTOR_IMAGE, timeout=30, check=True)
if _pg_ready(host, port, 20):
print(f" ✓ pgvector running on {host}:{port}")
else:
print(" Warning: Container started but PostgreSQL not yet accepting connections.\n It may need a few more seconds. Config will be saved; retry later.")
return {"host": host, "port": port, "user": "postgres", "password": password, "dbname": "postgres"}
except subprocess.CalledProcessError as e:
print(f" Failed to start Docker container: {_scrub(str(e), password)}")
except Exception as e:
print(f" Docker error: {_scrub(str(e), password)}")
return None
def _ensure_ollama(models: list[str]) -> bool:
"""Ensure Ollama is running and ``models`` are pulled; False when the user must handle it manually."""
ollama_bin = shutil.which("ollama")
if not (ok := _check_ollama(_OLLAMA_URL)[0]):
if not ollama_bin:
print(" Ollama not found. Install it:\n curl -fsSL https://ollama.com/install.sh | sh\n Or on macOS: brew install ollama")
return False
print(" Ollama installed but not running. Starting...")
try:
subprocess.Popen([ollama_bin, "serve"], stdin=subprocess.DEVNULL, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
_wait_for_port("localhost", 11434, timeout=10)
if ok := _check_ollama(_OLLAMA_URL)[0]:
print(" ✓ Ollama started")
except Exception as e:
print(f" Could not start Ollama: {e}")
if not ok:
print(" Warning: Ollama not reachable. Models cannot be pulled.")
return False
for model in models:
try:
names = [m.get("name", "") for m in json.loads(_http_get(_OLLAMA_URL, "/api/tags", 5).read()).get("models", [])]
except Exception:
names = []
if any(model in n or model.split(":")[0] in n for n in names):
print(f" ✓ Model '{model}' available")
continue
print(f" Pulling '{model}'... (this may take a few minutes)")
try:
subprocess.run([ollama_bin or "ollama", "pull", model], timeout=600, stdin=subprocess.DEVNULL)
print(f" ✓ Model '{model}' pulled")
except Exception as e:
print(f" Warning: Could not pull '{model}': {e}\n Run manually: ollama pull {model}")
return True
def _ensure_pgvector_extension(pg_config: dict) -> None:
try:
import psycopg2
except ImportError:
return
defaults = {"host": "localhost", "port": 5432, "user": "postgres", "dbname": "postgres"}
try:
conn = psycopg2.connect(**(defaults | {k: v for k, v in pg_config.items() if k in defaults or (k == "password" and v)}))
conn.autocommit = True
conn.cursor().execute("CREATE EXTENSION IF NOT EXISTS vector")
conn.close()
print(" ✓ pgvector extension enabled")
except Exception as e:
print(f" Warning: Could not enable pgvector extension: {e}")
def _wait_for_port(host: str, port: int, timeout: int = 15) -> None:
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
try:
socket.create_connection((host, port), timeout=1).close()
return
except OSError:
time.sleep(0.5)
# Picker descriptions: LLM/embedder show model (+ URL); vector stores by provider id (default: the id itself).
_VECTOR_DESCRIPTIONS = {"qdrant": lambda cfg: cfg.get("path", "local storage"), "pgvector": lambda cfg: f"{cfg.get('host', 'localhost')}:{cfg.get('port', 5432)}"}
def _configure_model_provider(kind: str, registry: dict, hermes_home: str, env_writes: dict[str, str], llm: tuple[str, dict] | None = None) -> tuple[str, dict, str, str | None]:
"""Pick an LLM/embedder provider, collect its key, and (for Ollama) model + URL -> (id, definition, model, url).
For the embedder (``llm`` given), a provider shared with the LLM reuses the LLM key instead of prompting again."""
items = [(v["label"], f"{v.get('default_model', '')} ({v['default_url']})" if v.get("default_url") else v.get("default_model", "")) for v in registry.values()]
pid = list(registry)[_curses_select(f"{kind} Provider", items, 0)]
pdef = registry[pid]
model, url = pdef["default_model"], pdef.get("default_url")
if pdef["needs_key"]:
if llm is None or pid != llm[0]:
if key := _prompt_api_key(pdef["label"] if llm is None else f"{pdef['label']} embedder", pdef["env_var"], hermes_home):
env_writes[pdef["env_var"]] = key
elif llm[1].get("env_var") in env_writes:
env_writes[pdef["env_var"]] = env_writes[llm[1]["env_var"]]
if pid == "ollama":
model = _input(f"{kind} model", pdef["default_model"])
url = _input("Ollama URL", pdef["default_url"])
return pid, pdef, model, url
def _setup_oss_interactive(hermes_home: str, config: dict) -> None:
env_writes: dict[str, str] = {}
llm_id, llm_def, llm_model, llm_url = _configure_model_provider("LLM", LLM_PROVIDERS, hermes_home, env_writes)
embedder_id, _, embedder_model, embedder_url = _configure_model_provider("Embedder", EMBEDDER_PROVIDERS, hermes_home, env_writes, llm=(llm_id, llm_def))
vector_items = [(v["label"], _VECTOR_DESCRIPTIONS.get(pid, lambda cfg: pid)(vector_default_config(pid))) for pid, v in VECTOR_PROVIDERS.items()]
vector_id = list(VECTOR_PROVIDERS)[_curses_select("Vector Store", vector_items, 0)]
# Auto-setup: ensure Ollama is running and models are pulled; ensure pgvector is reachable (offer Docker if not).
ollama_models = [m for pid, m in ((llm_id, llm_model), (embedder_id, embedder_model)) if pid == "ollama"]
if ollama_models:
_ensure_ollama(ollama_models)
pgvector_config = _ensure_pgvector() if vector_id == "pgvector" else None
if vector_id == "pgvector" and not pgvector_config: # native PostgreSQL: prompt for connection details (user first, historical order)
pg = {k: _input(f"PostgreSQL {label}", d) for k, label, d in (("user", "user", os.getenv("USER", "postgres")), ("host", "host", "localhost"), ("port", "port", "5432"), ("dbname", "database", "postgres"))}
pg_password = getpass.getpass(" PostgreSQL password (blank if none): ").strip()
pgvector_config = {**pg, "port": int(pg["port"]), **({"password": pg_password} if pg_password else {})}
user_id = _input("User ID", os.getenv("USER", "hermes-user"))
agent_id = _input("Agent ID", "hermes")
flags = {
"oss_llm": llm_id, "oss_llm_model": llm_model, "oss_llm_url": llm_url or "",
"oss_llm_key": env_writes.get(llm_def["env_var"], "") if llm_def.get("env_var") else "",
"oss_embedder": embedder_id, "oss_embedder_model": embedder_model, "oss_embedder_url": embedder_url or "",
"oss_vector": vector_id, "user_id": user_id,
}
flags.update({f"oss_vector_{key}": str(val) for key, val in (pgvector_config or {}).items() if val})
oss_config, _ = build_oss_config(flags)
_finish_oss(hermes_home, config, oss_config, env_writes, user_id, agent_id, pgvector_config)
def _install_provider_deps(llm_id: str, embedder_id: str, vector_id: str) -> None:
deps = {registry[pid]["pip_dep"] for (_, registry), pid in zip(SECTION_REGISTRIES, (llm_id, embedder_id, vector_id)) if registry.get(pid, {}).get("pip_dep")}
for dep in sorted(deps):
print(f" Installing {dep}...")
try:
# Environment-aware install: sealed hosted venvs redirect to the durable data-volume target instead of /opt/hermes.
from tools.lazy_deps import install_specs
outcome = install_specs([dep], timeout=60)
except Exception:
outcome = None
print(f" ✓ Installed {dep}" if outcome is not None and outcome.ok else f" Warning: cannot install {dep}: {outcome.reason}" if outcome is not None and outcome.blocked
else f" Warning: Could not install {dep}. Install manually: uv pip install {dep}")
if deps:
import importlib
importlib.invalidate_caches()
def _probe(fn, ok: str, fail: str, exc=Exception) -> tuple[bool, str]:
"""Run ``fn``; (True, ok) on success, (False, "fail: <error>") on ``exc``."""
try:
fn()
return True, ok
except exc as e:
return False, f"{fail}: {e}"
def _check_qdrant_path(path: str) -> tuple[bool, str]:
"""Check that qdrant local storage parent dir is writable."""
parent = Path(path).expanduser().parent
return _probe(lambda: parent.mkdir(parents=True, exist_ok=True), f"Directory writable: {parent}", f"Cannot write to {parent}", OSError)
def _check_ollama(url: str) -> tuple[bool, str]:
return _probe(lambda: _http_get(url, "/api/tags", 3), "Ollama reachable", f"Ollama not reachable at {url}")
def _check_pgvector(host: str, port: int) -> tuple[bool, str]:
return _probe(lambda: socket.create_connection((host, port), timeout=3).close(), f"PGVector reachable at {host}:{port}", f"PGVector not reachable at {host}:{port}")
def _warn_unless(check: tuple[bool, str]) -> None:
ok, msg = check
if not ok:
print(f" Warning: {msg}")
def _run_connectivity_checks(oss_config: dict) -> None:
vs = oss_config.get("vector_store", {})
cfg = vs.get("config", {})
if vs.get("provider") == "qdrant":
path, url = cfg.get("path"), cfg.get("url")
if path:
_warn_unless(_check_qdrant_path(path))
elif url:
_warn_unless(_probe(lambda: _http_get(url, "/healthz", 3), "Qdrant reachable", f"Qdrant not reachable at {url}"))
elif vs.get("provider") == "pgvector":
_warn_unless(_check_pgvector(cfg.get("host", "localhost"), cfg.get("port", 5432)))
llm = oss_config.get("llm", {})
if llm.get("provider") == "ollama":
_warn_unless(_check_ollama(llm.get("config", {}).get("ollama_base_url", _OLLAMA_URL)))
_MODE_HANDLERS = {"oss": _setup_oss, "selfhosted": _setup_selfhosted, "self-hosted": _setup_selfhosted, "platform": _setup_platform}
# Interactive picker order: Platform, Self-hosted server, Open Source.
_MODE_ITEMS = [("Platform", "Mem0 Cloud API (lightweight, just needs an API key)"), ("Self-hosted server", "Connect to an existing self-hosted Mem0 server (Docker/FastAPI)"), ("Open Source", "Run Mem0 locally (self-hosted LLM + vector store)")]
_MODE_PICKER = (_setup_platform, _setup_selfhosted, _setup_oss)
def post_setup(hermes_home: str, config: dict) -> None:
"""Entry point for `hermes memory setup`: routes on --mode (platform / selfhosted / oss), else shows a picker.
OSS is non-interactive only when the mode came from the flag."""
with suppress(ImportError): # mem0ai must meet the minimum version from plugin.yaml
import mem0
installed_ver = getattr(mem0, "__version__", None)
if installed_ver and _version_tuple(installed_ver) < (2, 0, 10):
print(f"\n ⚠ mem0ai {installed_ver} installed but >=2.0.10 required.\n Run: uv pip install --python {sys.executable} 'mem0ai>=2.0.10'")
flags = parse_flags(sys.argv[1:])
handler = _MODE_HANDLERS.get(flags["mode"])
flags["_mode_from_flag"] = handler is not None
if handler is None:
handler = _MODE_PICKER[_curses_select(" Select mode", _MODE_ITEMS, 0)]
handler(hermes_home, config, flags)
# ---- BEGIN PLUGIN-COMPAT (revert-scheduled; see COMPAT_MANIFEST.md) ----
# Names external plugins imported from this module before the Sep 2026 decomposition.
# Internal code MUST NOT use these (scripts/check_compat_pointers.py fails CI if it does).
# The whole block is removed by reverting the commit that added it.
def has_oss_flags() -> bool:
"""Check if OSS-related flags are present in sys.argv."""
flags = parse_flags(sys.argv[1:])
if flags["mode"] == "oss":
return True
if any(flags.get(k) for k in ("oss_llm_key", "oss_vector_path", "oss_vector_url")):
return True
return False
# ---- END PLUGIN-COMPAT ----
@@ -1,5 +0,0 @@
name: mem0
version: 1.3.0
description: "Mem0 — server-side LLM fact extraction with semantic search, automatic deduplication, and opt-in reranking (platform mode)."
pip_dependencies:
- mem0ai>=2.0.10,<3
@@ -1,17 +0,0 @@
[project]
name = "hermes-plugin-mem0"
version = "1.3.0"
description = "Hermes Agent memory provider plugin: mem0"
requires-python = ">=3.11"
license = { text = "Apache-2.0 AND MIT" }
# Installed into the Hermes venv by `hermes plugins install` / `enable` and re-applied after
# every `hermes update` (hermes-agent#113851). Keep upper bounds: Hermes pins its own deps exactly
# and refuses a plugin whose requirements cannot resolve against them.
dependencies = [
"mem0ai>=2.0.10,<3",
"httpx>=0.27,<1",
]
[project.optional-dependencies]
postgres = ['psycopg2-binary>=2.9,<3']
qdrant = ['qdrant-client>=1.9,<2']
@@ -1,178 +0,0 @@
"""Real Hermes + Mem0 SDK + local Qdrant; only the model API is simulated.
HERMES_SOURCE=/path/to/hermes-agent python tests/smoke_hermes.py
Requires the Hermes dependencies, mem0ai, and qdrant-client in the active environment.
All configuration, credentials, and database files are temporary.
"""
import importlib
import io
import json
import os
import shutil
import sys
import tempfile
import threading
from contextlib import redirect_stdout
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from unittest.mock import patch
class ModelAPI(BaseHTTPRequestHandler):
calls = []
def log_message(self, *args):
pass
def do_POST(self):
if self.headers.get("Authorization") != "Bearer local-test":
self.send_error(401)
return
body = json.loads(self.rfile.read(int(self.headers["Content-Length"])))
self.calls.append(self.path)
if self.path == "/v1/embeddings":
result = {
"object": "list",
"data": [{"object": "embedding", "index": i, "embedding": [1.0, 0.0, 0.0]}
for i, _ in enumerate(body["input"])],
"model": "test-embedding",
"usage": {"prompt_tokens": 1, "total_tokens": 1},
}
elif self.path == "/v1/chat/completions":
result = {
"id": "test-completion", "object": "chat.completion", "created": 0, "model": "gpt-4.1-mini",
"choices": [{"index": 0, "finish_reason": "stop", "message": {"role": "assistant", "content":
json.dumps({"facts": ["Prefers green tea."], "memory": [
{"id": "0", "text": "Prefers green tea.", "event": "ADD"}]})}}],
}
else:
self.send_error(404)
return
encoded = json.dumps(result).encode()
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(encoded)))
self.end_headers()
self.wfile.write(encoded)
def main():
hermes_source = Path(os.environ["HERMES_SOURCE"]).resolve()
sys.path.insert(0, str(hermes_source))
with tempfile.TemporaryDirectory() as temporary:
home = Path(temporary)
os.environ.update(HERMES_HOME=temporary, MEM0_DIR=str(home / "mem0-data"), MEM0_TELEMETRY="false")
(home / "mem0-data").mkdir()
server = ThreadingHTTPServer(("127.0.0.1", 0), ModelAPI)
worker = threading.Thread(target=server.serve_forever, daemon=True)
worker.start()
try:
url = f"http://127.0.0.1:{server.server_port}/v1"
config = {
"mode": "oss", "user_id": "existing-user", "agent_id": "hermes",
"oss": {
"llm": {"provider": "openai", "config": {
"model": "gpt-4.1-mini", "api_key": "local-test", "openai_base_url": url}},
"embedder": {"provider": "openai", "config": {
"model": "test-embedding", "embedding_dims": 3,
"api_key": "local-test", "openai_base_url": url}},
"vector_store": {"provider": "qdrant", "config": {"path": str(home / "qdrant")}},
},
}
(home / "mem0.json").write_text(json.dumps(config))
destination = home / "plugins" / "mem0"
shutil.copytree(Path(__file__).resolve().parents[1], destination)
import plugins.memory as memory_plugins
from agent.memory_manager import MemoryManager
# Simulate bundled-provider removal only inside this test process.
memory_plugins._MEMORY_PLUGINS_DIR = home / "empty-bundled"
memory_plugins._MEMORY_PLUGINS_DIR.mkdir()
assert memory_plugins.find_provider_dir("mem0") == destination
provider = memory_plugins.load_memory_provider("mem0", register_skills=False)
assert provider is not None and provider.__class__.__module__.startswith("_hermes_user_memory.")
setup = importlib.import_module(f"{provider.__class__.__module__}._setup")
from hermes_cli.memory_setup import cmd_setup_provider, cmd_status
output = io.StringIO()
arguments = ["hermes", "memory", "setup", "mem0", "--mode", "platform", "--api-key", "local-test"]
with patch.object(sys, "argv", arguments), patch.object(sys, "stdin", io.StringIO("\n\n")):
with patch.object(setup, "_curses_select", return_value=0), redirect_stdout(output):
cmd_setup_provider("mem0")
assert json.loads((home / "mem0.json").read_text())["user_id"] == "existing-user"
provider.save_config(config, home)
assert (home / "mem0.json").stat().st_mode & 0o777 == 0o600
with redirect_stdout(output):
cmd_status(None)
assert "installed ✓" in output.getvalue() and "available ✓" in output.getvalue()
assert (home / ".env").stat().st_mode & 0o777 == 0o600
manager = MemoryManager()
manager.add_provider(provider)
manager.initialize_all("smoke-session", platform="cli", user_id="gateway-user")
assert provider._backend is not None, getattr(provider, "_init_error", "initialization failed")
def tool(name, **arguments):
result = json.loads(manager.handle_tool_call(name, arguments))
assert "error" not in result, result
return result
try:
assert tool("mem0_add", content="Prefers Python.")["result"] == "Fact stored."
memories = tool("mem0_search", query="language")["results"]
assert len(memories) == 1 and memories[0]["memory"] == "Prefers Python."
memory_id = memories[0]["id"]
tool("mem0_update", memory_id=memory_id, text="Prefers Rust.")
assert "Prefers Rust." in provider.prefetch("language preference")
assert tool("mem0_search", query="language")["results"][0]["memory"] == "Prefers Rust."
tool("mem0_delete", memory_id=memory_id)
assert tool("mem0_search", query="language")["result"] == "No relevant memories found."
manager.sync_all("I prefer green tea.", "Noted.", session_id="smoke-session")
assert provider._user_id == "existing-user"
finally:
manager.shutdown_all()
assert provider._backend is None
# A changed embedder must fail without destroying the existing collection.
config["oss"]["embedder"]["config"]["embedding_dims"] = 4
provider.save_config(config, home)
mismatched = type(provider)()
mismatched.initialize("mismatched-session")
assert mismatched._backend is None and "Existing memories were preserved" in mismatched._init_error
config["oss"]["embedder"]["config"]["embedding_dims"] = 3
# Resolve embedder credentials from the active profile, not another profile's process env.
del config["oss"]["embedder"]["config"]["api_key"]
del config["oss"]["embedder"]["config"]["openai_base_url"]
provider.save_config(config, home)
from agent.secret_scope import (
reset_secret_scope,
set_multiplex_active,
set_secret_scope,
)
resumed = type(provider)()
with patch.dict(os.environ, {"OPENAI_API_KEY": "wrong-profile", "OPENAI_API_BASE": "http://127.0.0.1:1/v1"}):
set_multiplex_active(True)
token = set_secret_scope({"OPENAI_API_KEY": "local-test", "OPENAI_BASE_URL": url})
try:
resumed.initialize("resumed-session", user_id="different-gateway-user")
finally:
reset_secret_scope(token)
set_multiplex_active(False)
try:
assert resumed._backend is not None
found = json.loads(resumed.handle_tool_call("mem0_search", {"query": "drink"}))
assert "results" in found, found
assert found["results"][0]["memory"] == "Prefers green tea."
finally:
resumed.shutdown()
assert "/v1/chat/completions" in ModelAPI.calls
print("PASS: external Hermes loader, CLI setup/status, real Mem0/Qdrant CRUD, recall, background extraction,")
print(" existing identity, profile credentials, private files, shutdown, dimension safety and persistence across restart. Model responses simulated locally.")
finally:
server.shutdown()
server.server_close()
worker.join(timeout=5)
if __name__ == "__main__":
main()
@@ -1,286 +0,0 @@
"""Offline regressions for the standalone Hermes plugin."""
import contextvars
import importlib
import importlib.util
import json
import sys
import threading
import types
from pathlib import Path
from unittest.mock import Mock
import pytest
ROOT = Path(__file__).resolve().parents[1]
@pytest.fixture
def plugin(monkeypatch, tmp_path):
def spawn(target, *, name):
return threading.Thread(target=contextvars.copy_context().run, args=(target,), name=name)
def atomic_write_text(path, value, *, mode):
path.write_text(value)
path.chmod(mode)
modules = {
"agent.memory_provider": {"MemoryProvider": object, "spawn_context_thread": spawn},
"agent.secret_scope": {"get_secret": lambda key, default="": default},
"tools.registry": {"tool_error": lambda message: json.dumps({"error": message})},
"utils": {
"read_json_or_empty": lambda path: json.loads(path.read_text()) if path.exists() else {},
"atomic_json_write": lambda path, value, **kwargs: atomic_write_text(path, json.dumps(value), **kwargs),
"atomic_write_text": atomic_write_text,
},
"hermes_constants": {"get_hermes_home": lambda: tmp_path},
}
for name, values in modules.items():
module = types.ModuleType(name)
module.__dict__.update(values)
monkeypatch.setitem(sys.modules, name, module)
spec = importlib.util.spec_from_file_location("standalone_mem0", ROOT / "__init__.py")
module = importlib.util.module_from_spec(spec)
monkeypatch.setitem(sys.modules, spec.name, module)
spec.loader.exec_module(module)
monkeypatch.setattr(module.atexit, "register", lambda *args: None)
yield module
for name in list(sys.modules):
if name.startswith("standalone_mem0."):
monkeypatch.delitem(sys.modules, name)
def test_dimension_mismatch_preserves_qdrant_collection(plugin, monkeypatch):
backend = importlib.import_module(f"{plugin.__name__}._backend")
client = Mock()
client.collection_exists.return_value = True
client.get_collection.return_value.config.params.vectors = types.SimpleNamespace(size=1536)
monkeypatch.setitem(sys.modules, "qdrant_client", types.SimpleNamespace(QdrantClient=Mock(return_value=client)))
with pytest.raises(ValueError, match="1536.*768"):
backend.OSSBackend._reject_dimension_mismatch("qdrant", {"path": "/unused"}, 768)
client.delete_collection.assert_not_called()
client.close.assert_called_once()
def test_dimension_mismatch_preserves_pgvector_table(plugin, monkeypatch):
backend = importlib.import_module(f"{plugin.__name__}._backend")
cursor, connection = Mock(), Mock()
cursor.fetchone.return_value = (1536,)
connection.cursor.return_value = cursor
driver = types.SimpleNamespace(connect=Mock(return_value=connection), sql=Mock())
monkeypatch.setitem(sys.modules, "psycopg2", driver)
with pytest.raises(ValueError, match="1536.*768"):
backend.OSSBackend._reject_dimension_mismatch("pgvector", {"user": "test"}, 768)
assert cursor.execute.call_count == 1
assert cursor.execute.call_args.args[0].startswith("SELECT")
connection.close.assert_called_once()
def test_setup_keeps_credentials_private_and_preserves_existing_values(plugin, tmp_path):
setup = importlib.import_module(f"{plugin.__name__}._setup")
path = tmp_path / ".env"
path.write_text("EXISTING=value\nMEM0_API_KEY=old\n")
path.chmod(0o644)
setup._write_env(path, {"MEM0_API_KEY": "new"})
assert path.read_text() == "EXISTING=value\nMEM0_API_KEY=new\n"
assert path.stat().st_mode & 0o777 == 0o600
def test_oss_setup_keeps_database_password_private(plugin, tmp_path, monkeypatch):
setup = importlib.import_module(f"{plugin.__name__}._setup")
for name in ("_install_provider_deps", "_activate_provider", "_run_connectivity_checks"):
monkeypatch.setattr(setup, name, Mock())
config = {"llm": {"provider": "openai", "config": {}}, "embedder": {"provider": "openai", "config": {}},
"vector_store": {"provider": "pgvector", "config": {"password": "test-password"}}}
setup._finish_oss(str(tmp_path), {}, config, {}, "existing-user", "hermes")
path = tmp_path / "mem0.json"
assert json.loads(path.read_text())["user_id"] == "existing-user"
assert path.stat().st_mode & 0o777 == 0o600
def test_optional_server_key_and_prefetch_rerank(plugin, monkeypatch):
monkeypatch.setattr(plugin, "_load_config", lambda: {"host": "http://localhost:8888", "rerank": True})
provider = plugin.Mem0MemoryProvider()
assert not next(field for field in provider.get_config_schema() if field["key"] == "api_key")["required"]
backend = Mock()
backend.search.return_value = [{"memory": "Prefers Python"}]
monkeypatch.setattr(provider, "_create_backend", lambda: backend)
provider.initialize("test-session")
try:
assert "Prefers Python" in provider.prefetch("language")
assert backend.search.call_args.kwargs["rerank"] is True
finally:
provider.shutdown()
def test_invalid_tool_arguments_never_reach_backend(plugin, monkeypatch):
provider = plugin.Mem0MemoryProvider()
provider._backend = Mock()
for args in (None, {"query": []}, {"query": " "}, {"query": "fact", "top_k": "invalid"}):
assert "error" in json.loads(provider.handle_tool_call("mem0_search", args))
provider._backend.search.assert_not_called()
assert provider._consecutive_failures == 0
def test_selfhosted_http_auth_and_tool_routes(plugin):
import httpx
backend = importlib.import_module(f"{plugin.__name__}._backend")
requests = []
def respond(request):
requests.append(request)
return httpx.Response(200, json={"results": [{"id": "existing", "memory": "Prefers Python"}]})
client = backend.SelfHostedBackend("test-key", "http://localhost:8888", transport=httpx.MockTransport(respond))
try:
client.add([], user_id="existing-user", agent_id="hermes", infer=True)
assert requests[-1].headers["X-API-Key"] == "test-key"
assert json.loads(requests[-1].content)["user_id"] == "existing-user"
assert client.search("language", filters={"user_id": "existing-user"})[0]["id"] == "existing"
assert json.loads(requests[-1].content)["filters"] == {"user_id": "existing-user"}
client.update("existing", "Prefers Rust")
assert json.loads(requests[-1].content) == {"text": "Prefers Rust"}
client.delete("existing")
assert [(r.method, r.url.path) for r in requests] == [
("POST", "/memories"), ("POST", "/search"), ("PUT", "/memories/existing"), ("DELETE", "/memories/existing")
]
finally:
client.close()
@pytest.mark.parametrize("mode", ["platform", "selfhosted"])
def test_setup_rotates_legacy_file_key(plugin, monkeypatch, tmp_path, mode):
setup = importlib.import_module(f"{plugin.__name__}._setup")
(tmp_path / "mem0.json").write_text(json.dumps({"api_key": "old-key", "user_id": "existing-user"}))
monkeypatch.setattr(setup, "_activate_provider", Mock())
monkeypatch.setattr(setup, "_check_selfhosted_server", Mock())
monkeypatch.setattr(setup, "_prompt", lambda label, default=None, **kwargs: default or "")
monkeypatch.setattr(setup, "_curses_select", lambda *args, **kwargs: 0)
setup._MODE_HANDLERS[mode](str(tmp_path), {}, {"api_key": "new-key", "host": "http://localhost:8888"})
monkeypatch.setattr(plugin, "get_secret", lambda key, default="": "new-key" if key == "MEM0_API_KEY" else default)
assert plugin._load_config()["api_key"] == "new-key"
assert "old-key" not in (tmp_path / "mem0.json").read_text()
assert "MEM0_API_KEY=new-key" in (tmp_path / ".env").read_text()
def test_platform_setup_honors_user_id_flag(plugin, monkeypatch, tmp_path):
setup = importlib.import_module(f"{plugin.__name__}._setup")
monkeypatch.setattr(setup, "_activate_provider", Mock())
monkeypatch.setattr(setup, "_prompt", lambda label, default=None, **kwargs: default or "")
monkeypatch.setattr(setup, "_curses_select", lambda *args, **kwargs: 0)
setup._setup_platform(str(tmp_path), {}, {"api_key": "new-key", "user_id": "chosen-user"})
assert plugin._load_config()["user_id"] == "chosen-user"
@pytest.mark.parametrize("scoped_key", ["profile-key", ""])
def test_oss_embedder_never_uses_another_profiles_credentials(plugin, monkeypatch, scoped_key):
backend = importlib.import_module(f"{plugin.__name__}._backend")
memory = Mock()
monkeypatch.setitem(sys.modules, "mem0", types.SimpleNamespace(Memory=memory))
qdrant_config = types.SimpleNamespace(model_fields={"path": types.SimpleNamespace(default=None)})
monkeypatch.setitem(
sys.modules, "mem0.configs.vector_stores.qdrant", types.SimpleNamespace(QdrantConfig=qdrant_config)
)
monkeypatch.setenv("OPENAI_API_KEY", "other-profile-key")
monkeypatch.setenv("OPENAI_API_BASE", "https://other-profile.invalid/v1")
secrets = {"OPENAI_API_KEY": scoped_key, "OPENAI_BASE_URL": "https://profile.invalid/v1"}
monkeypatch.setattr(sys.modules["agent.secret_scope"], "get_secret", lambda key, default="": secrets.get(key, default))
config = {
"llm": {"provider": "ollama", "config": {}},
"embedder": {"provider": "openai", "config": {}},
"vector_store": {"provider": "qdrant", "config": {}},
}
if not scoped_key:
with pytest.raises(ValueError, match="OpenAI API key"):
backend.OSSBackend(config)
memory.from_config.assert_not_called()
else:
backend.OSSBackend(config)
resolved = memory.from_config.call_args.args[0]["embedder"]["config"]
assert resolved["api_key"] == scoped_key
assert resolved["openai_base_url"] == "https://profile.invalid/v1"
assert config["embedder"]["config"] == {}
def test_pgvector_setup_never_removes_existing_container(plugin, monkeypatch):
setup = importlib.import_module(f"{plugin.__name__}._setup")
monkeypatch.setattr(setup, "_check_pgvector", lambda *args: (False, "unreachable"))
monkeypatch.setattr(setup.shutil, "which", lambda name: "/test/docker")
monkeypatch.setattr("builtins.input", lambda prompt: "y")
calls = []
def docker(*args, **kwargs):
calls.append(args)
if args[0] == "run":
raise setup.subprocess.CalledProcessError(1, "docker run: container name already exists")
return types.SimpleNamespace(returncode=0, stdout="paused")
monkeypatch.setattr(setup, "_docker", docker)
assert setup._ensure_pgvector() is None
assert not any(args[0] == "rm" for args in calls)
def test_platform_dry_run_does_not_print_stored_secrets(plugin, monkeypatch, tmp_path, capsys):
setup = importlib.import_module(f"{plugin.__name__}._setup")
config = {"api_key": "old-secret", "oss": {"vector_store": {"config": {"password": "db-secret"}}}}
path = tmp_path / "mem0.json"
path.write_text(json.dumps(config))
monkeypatch.setattr(setup, "_prompt", lambda label, default=None, **kwargs: default or "")
monkeypatch.setattr(setup, "_curses_select", lambda *args, **kwargs: 0)
setup._setup_platform(str(tmp_path), {}, {"api_key": "new-secret", "dry_run": True})
output = capsys.readouterr().out
assert all(secret not in output for secret in ("old-secret", "db-secret", "new-secret"))
assert json.loads(path.read_text()) == config
assert not (tmp_path / ".env").exists()
def test_oss_setup_preserves_distinct_llm_and_embedder_keys(plugin):
setup = importlib.import_module(f"{plugin.__name__}._setup")
config, env = setup.build_oss_config({"oss_llm_key": "llm-key", "oss_embedder_key": "embedder-key"})
assert config["llm"]["config"].get("api_key", env.get("OPENAI_API_KEY")) == "llm-key"
assert config["embedder"]["config"].get("api_key", env.get("OPENAI_API_KEY")) == "embedder-key"
def test_direct_openai_llm_uses_scoped_credentials(plugin, monkeypatch):
llm_mod = importlib.import_module(f"{plugin.__name__}._openai_llm")
openai_mock = types.SimpleNamespace(OpenAI=Mock(return_value=Mock()))
monkeypatch.setitem(sys.modules, "openai", openai_mock)
monkeypatch.setenv("OPENROUTER_API_KEY", "should-be-ignored")
monkeypatch.setenv("OPENAI_API_KEY", "env-key-should-be-ignored")
secrets = {"OPENAI_API_KEY": "scoped-key", "OPENAI_API_BASE": "", "OPENAI_BASE_URL": ""}
monkeypatch.setattr(sys.modules["agent.secret_scope"], "get_secret", lambda key, default="": secrets.get(key, default))
llm_mod.DirectOpenAILLM({"api_key": "", "model": "gpt-5-mini"})
call_kwargs = openai_mock.OpenAI.call_args.kwargs
assert call_kwargs["api_key"] == "scoped-key"
assert "openrouter" not in call_kwargs.get("base_url", "").lower()
def test_direct_openai_llm_rejects_missing_key(plugin, monkeypatch):
llm_mod = importlib.import_module(f"{plugin.__name__}._openai_llm")
monkeypatch.setattr(sys.modules["agent.secret_scope"], "get_secret", lambda key, default="": "")
with pytest.raises(ValueError, match="API key"):
llm_mod.DirectOpenAILLM({"api_key": "", "model": "gpt-5-mini"})
def test_selfhosted_keyless_omits_auth_header(plugin):
import httpx
backend = importlib.import_module(f"{plugin.__name__}._backend")
requests = []
def respond(request):
requests.append(request)
return httpx.Response(200, json={"results": []})
client = backend.SelfHostedBackend("", "http://localhost:8888", transport=httpx.MockTransport(respond))
client.search("test", filters={"user_id": "u"})
assert "X-API-Key" not in requests[0].headers
client.close()
def test_initialize_tolerates_non_numeric_sync_max_chars(plugin, monkeypatch):
monkeypatch.setattr(plugin, "_load_config", lambda: {"sync_max_chars": "not-a-number"})
provider = plugin.Mem0MemoryProvider()
provider.initialize("test-session")
assert provider._sync_max_chars == plugin._SYNC_MSG_MAX_CHARS
+10 -3
View File
@@ -21,9 +21,16 @@ from memory_core import (
PROTOCOL_VERSION = "2024-11-05"
TOOL_NAME = "search_memories"
TOOL_DESCRIPTION = (
"Search memories from earlier work in this repository. Use it before "
"repeating investigation or when earlier decisions, fixes, commands, or "
"results may help."
"Search memories from earlier work in this repository. ALWAYS call this "
"tool before answering anything that could depend on prior context: the "
"user's preferences, facts about this codebase, history, people, projects, "
"or earlier decisions. Do not rely on the chat window alone. The "
"repository's memory is shared by everyone who works in it and includes "
"what it took to run, test, or build here, so search before assuming an "
"invocation works. The scope argument changes what is searched: 'repo' "
"(default) is the whole repository's shared memory plus your own "
"preferences, 'dir' narrows the shared part to the directory you are "
"working in, and 'mine' is your preferences alone."
)
TOOL_SCHEMA = {
"type": "object",
+6 -4
View File
@@ -29,7 +29,7 @@ from typing import Any, Iterable
import telemetry
DEFAULT_API_URL = "https://api.mem0.ai"
PLUGIN_VERSION = "0.3.3"
PLUGIN_VERSION = "0.3.1"
_harness_name: str = "generic"
_harness_env_prefix: str = "MEM0_PLUGIN"
@@ -71,13 +71,15 @@ MAX_FLUSH_ATTEMPTS = 5
FORGET_PAGE_SIZE = 100
FORGET_MAX_PAGES = 50
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help with future coding work.
PROJECT_MEMORY_INSTRUCTIONS = """Save concise repository facts that will help anyone with future coding work in this repository.
A completed change should produce one memory explaining the resulting behavior, where it is implemented when useful, and any important constraints or reasoning. Exploration or accepted decisions may produce separate memories only when they are independently useful.
Use the coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
A command that failed and was then made to work should produce one memory naming the failing invocation, the error it returned, and the invocation that succeeded. Do not save one-off errors caused by an edit still in progress, transient network failures, or anything a rerun would fix on its own.
Write about the repository, not the user, assistant, session, or task. Do not include test results, documentation updates, release notes, or temporary state.
Use the current coding agent's final response for conclusions about current repository behavior. Do not save proposed or recommended changes unless the user accepted them or the coding agent completed them. Treat subagent responses as supporting repository evidence, not as decisions.
Write about the repository, not the user, assistant, session, or task. Do not save personal preferences. Do not save a memory that only states which repository, branch, or directory the session worked in. Do not include test results, documentation updates, release notes, or temporary state.
If nothing useful was established, return no memories."""
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "mem0",
"version": "0.3.3",
"version": "0.3.1",
"description": "Cross-session memory and token savings for coding agents.",
"keywords": ["memory", "coding-agents", "continual-learning", "token-efficiency"],
"author": { "name": "Mem0", "email": "support@mem0.ai" },
+1 -1
View File
@@ -1,6 +1,6 @@
{
"id": "mem0",
"version": "0.3.3",
"version": "0.3.1",
"homepage": "https://docs.mem0.ai/integrations/kimi",
"native": {
"pluginRoot": "${KIMI_PLUGIN_ROOT}",
@@ -12,9 +12,11 @@ Call `search_memories` with the user's question. Treat `--top-k`, `--category`,
query.
Omit `top_k` to use Mem0's configured default. Omit `category` to search every
category. Omit `scope` to use the configured default, normally `repo`: this
repository's shared memory, which everyone who works in it contributes to,
plus your own preferences.
category; a category is a best-effort label Mem0 assigned when it saved the
memory, so if a category search misses, repeat it without the category. Omit
`scope` to use the configured default, normally `repo`: this repository's
shared memory, which everyone who works in it contributes to, plus your own
preferences.
Pass `scope` when the question needs something else: `dir` to narrow the
shared memory to the directory you are working in (a package inside a

Some files were not shown because too many files have changed in this diff Show More