---
title: Hermes Agent
description: "Add long-term memory to Hermes agents using Mem0 Platform, a self-hosted server, or local OSS mode with background fact extraction."
---
Add long-term memory to [Hermes Agent](https://github.com/NousResearch/hermes-agent), a self-improving AI agent CLI by Nous Research. Hermes has a pluggable memory system, and Mem0 is one of the supported providers. Once enabled, Mem0 learns facts from your conversations and surfaces relevant ones for the current question, without slowing down the chat.
You can run Mem0 in three ways:
- **Platform mode** (default): managed Mem0 Cloud. Add your API key and you are ready.
- **Self-hosted server mode**: point the plugin at a Mem0 server you run yourself (the Docker-shipped server). The plugin only talks HTTP to your server.
- **OSS mode**: run Mem0 in-process with your own LLM, embedder, and vector store. No Mem0 server required.
## Requirements and compatibility
Use Python 3.11+ and a recent Hermes release. The host contract was checked against
Hermes `v0.21.3` (`v2026.9.14`) and main at
`c62bd9f2078a946108f1c9d9b24bf118963277ef`. Earlier host versions have not been validated.
Both checked Hermes revisions still bundle Mem0. Hermes gives bundled memory providers
precedence over a same-named user plugin. Installing or enabling this plugin does not
replace that bundled copy. Use the isolated preview instructions in the
[plugin README](https://github.com/mem0ai/mem0/tree/main/integrations/hermes-plugin#local-worktree-preview)
to try this version in a separate Hermes checkout. Once Hermes removes its bundled Mem0
provider, the user-installed copy loads directly.
Keep `memory.provider: mem0`, `$HERMES_HOME/mem0.json`, and your existing `MEM0_*` variables.
There is no memory migration. An explicitly configured user ID continues to share memories
across hosts; otherwise the provider uses the gateway user ID, then `hermes-user`.
## How It Works
Hermes runs a built-in memory system (file-based `MEMORY.md` and `USER.md`) alongside one external provider. When Mem0 is active, it works additively with the built-in system at two points in every conversation turn.
### 1. Current-turn recall (bounded wait)
When you send a message, Hermes searches your stored memories for the current question and waits up to 3 seconds for results. If they arrive in time, they are injected into the system prompt so the model can see them. If the backend is slower, Hermes skips the injection and the model can still call `mem0_search` itself — so a slow backend never blocks a turn.
### 2. Background fact extraction (sync)
Once the model finishes, Hermes sends the `(user message, assistant response)` pair to Mem0 in a background thread. Mem0 extracts facts automatically (for example, "user prefers Python" or "user works at Acme Corp"), so you never have to tell it what to remember. Each write is tagged with the gateway channel it came from and the Hermes session as top-level `run_id`. Recall remains scoped to the user across sessions.
The plugin uses the shared `agent-plugin-core` redaction and token-aware batching. Full non-empty user and assistant text is preserved after known-secret redaction; oversized messages are split across requests instead of truncated. Raw tool results and the full historical transcript are not captured. Pattern-based redaction cannot recognize every possible secret.
## Agent Tools
When Mem0 is active, the model gets four tools it can call during a conversation:
| Tool | Description | Parameters |
|------|-------------|------------|
| `mem0_search` | Semantic search by meaning, ranked by relevance | `query` (required), `top_k` (default 10, max 50), `rerank` (default `false`, Platform mode only) |
| `mem0_add` | Store a fact after known-secret redaction, with no LLM extraction | `content` (required) |
| `mem0_update` | Update a memory's text by ID | `memory_id`, `text` (both required) |
| `mem0_delete` | Delete a memory by ID | `memory_id` (required) |
## Installation
Install Hermes Agent:
```bash
curl -fsSL https://raw.githubusercontent.com/NousResearch/hermes-agent/main/scripts/install.sh | bash
source ~/.bashrc
```
After this integration is published to the Mem0 repository:
```bash
hermes plugins install mem0ai/mem0/integrations/hermes-plugin
hermes plugins enable mem0
hermes memory setup
hermes memory status
```
For unpublished worktree changes, use the README's local installation instructions instead.
The bundled-provider precedence described above applies to both installation methods.
Recent Hermes development versions install the plugin's declared dependencies automatically.
On Hermes `v0.21.3`, install them into the environment that runs Hermes:
```bash
HERMES_REPO=/path/to/hermes-agent
uv pip install --python "$HERMES_REPO/.venv/bin/python" 'mem0ai>=2.0.10,<3' 'httpx>=0.27,<1'
```
OSS providers may need extra packages such as `qdrant-client`, `psycopg2-binary`, or `ollama`, which the setup flow installs when you select them.
## Platform Setup
Platform mode uses managed Mem0 Cloud and is the fastest way to start.
### Option 1: Interactive wizard (recommended)
```bash
hermes memory setup
```
Select **mem0**, choose **Platform**, and paste your API key when prompted. The wizard writes the non-secret settings to `~/.hermes/mem0.json` and keeps the key in `~/.hermes/.env`.
Get your API key from app.mem0.ai.
### Option 2: Manual Configuration
```bash
hermes config set memory.provider mem0
echo "MEM0_API_KEY=your-api-key" >> ~/.hermes/.env
```
Then in your `config.yaml`:
```yaml
memory:
provider: mem0
```
That's it. Mem0 runs automatically from here.
## Self-Hosted Server Setup
Run the [Mem0 server](https://github.com/mem0ai/mem0/tree/main/server) (FastAPI + pgvector) from its Docker image and point the plugin at it. Unlike OSS mode, the plugin just talks HTTP to your server.
### Interactive
```bash
hermes memory setup
# Select "mem0", then "Self-hosted server", and enter the server URL
```
### With flags
```bash
hermes memory setup mem0 --mode selfhosted \
--host http://localhost:8888 \
--api-key your-admin-api-key
```
### With environment variables
```bash
echo "MEM0_HOST=http://localhost:8888" >> ~/.hermes/.env
echo "MEM0_API_KEY=your-admin-api-key" >> ~/.hermes/.env
```
Then start a fresh Hermes session and call `mem0_search` — it connects to your server. The plugin authenticates with `X-API-Key` and uses the server's `/search` and `/memories` routes. The API key is optional only for servers running with `AUTH_DISABLED`.
Setting `host` routes to the self-hosted server automatically. Don't combine it with `mode: oss` — OSS takes precedence and ignores `host`.
## OSS (Self-Hosted) Setup
OSS mode runs Mem0 entirely on your own infrastructure: your LLM, your embedder, and your vector store. No data is sent to Mem0 Cloud, and no Mem0 API key is required.
### Interactive
```bash
hermes memory setup
# Select "mem0", then "Open Source (self-hosted)"
# Follow the prompts for LLM, embedder, and vector store
```
### With flags
```bash
hermes memory setup mem0 --mode oss \
--oss-llm openai --oss-llm-key sk-... \
--oss-vector qdrant
```
### Supported providers
| Component | Providers |
|-----------|-----------|
| LLM | `openai` (default model `gpt-5-mini`), `ollama` (local, default `llama3.1:8b`) |
| Embedder | `openai` (default `text-embedding-3-small`), `ollama` (local, default `nomic-embed-text`) |
| Vector store | `qdrant` (local path or server), `pgvector` |
### Flag reference
| Flag | Description |
|------|-------------|
| `--mode` | `platform`, `selfhosted`, or `oss` |
| `--api-key` | Platform API key, or the admin key of a self-hosted server |
| `--host` | Self-hosted server URL (with `--mode selfhosted`) |
| `--oss-llm` | LLM provider (`openai` or `ollama`, default `openai`) |
| `--oss-llm-key` | LLM API key (for `openai`) |
| `--oss-llm-model` | Override the LLM model |
| `--oss-llm-url` | LLM base URL (for `ollama` or a custom endpoint) |
| `--oss-embedder` | Embedder provider (default `openai`) |
| `--oss-embedder-key` | Embedder API key |
| `--oss-embedder-model` | Override the embedder model |
| `--oss-embedder-url` | Embedder base URL (for `ollama` or a custom endpoint) |
| `--oss-vector` | Vector store (`qdrant` or `pgvector`, default `qdrant`) |
| `--oss-vector-path` | Local Qdrant storage path |
| `--oss-vector-url` | Qdrant server URL |
| `--oss-vector-host`, `--oss-vector-port` | PGVector or remote Qdrant host and port |
| `--oss-vector-user`, `--oss-vector-password`, `--oss-vector-dbname` | PGVector connection details |
| `--user-id` | Canonical user identifier |
| `--dry-run` | Preview the resolved config without writing it |
## Switching Modes
You can move between the three modes at any time. Run the setup command again, or edit `~/.hermes/mem0.json` directly.
```bash
# Platform to OSS
hermes memory setup mem0 --mode oss --oss-llm-key sk-...
# OSS to Platform
hermes memory setup mem0 --mode platform --api-key sk-...
# Platform to a self-hosted server
hermes memory setup mem0 --mode selfhosted --host http://localhost:8888
# Preview without writing anything
hermes memory setup mem0 --mode oss --oss-llm-key sk-... --dry-run
```
A self-hosted `~/.hermes/mem0.json` looks like this:
```json
{
"mode": "oss",
"oss": {
"llm": {"provider": "openai", "config": {"model": "gpt-5-mini", "is_reasoning_model": true}},
"embedder": {"provider": "openai", "config": {"model": "text-embedding-3-small"}},
"vector_store": {"provider": "qdrant", "config": {"path": "~/.hermes/mem0_qdrant"}}
}
}
```
## Configuration
Behavioral settings live in `~/.hermes/mem0.json` and are written for you by `hermes memory setup`. Only the secret `MEM0_API_KEY` belongs in `~/.hermes/.env`.
| Key | Default | Description |
|-----|---------|-------------|
| `mode` | `platform` | `platform` (Mem0 Cloud) or `oss` (self-managed, in-process). Self-hosted server routing is set via `host` |
| `host` | none | Self-hosted Mem0 server URL. When set, the plugin talks HTTP to your server instead of the cloud |
| `api_key` | none | Mem0 Platform API key, or the admin key of a self-hosted server. Stored in `.env` as `MEM0_API_KEY` |
| `user_id` | Gateway user ID, then `hermes-user` | Identifier that scopes memories. See cross-channel behavior below |
| `agent_id` | `hermes` | Agent identifier attached to writes |
| `rerank` | `false` | Rerank search results for relevance (Platform mode only) |
| `sync_max_chars` | Uncapped for platform/HTTP; `450` for OSS | Positive maximum per chunk; longer text is split without dropping its tail. `0` disables the character cap |
### Cross-channel memories
Hermes can run from the CLI and from gateways like Telegram, Slack, and Discord. The `user_id` setting controls how memories are scoped across them:
- **Set a `user_id`** and it applies to every gateway, so one person gets a single merged memory store no matter where they talk to the agent.
- **Leave it unset** (or at the default `hermes-user`) and each gateway uses its own native id, keeping per-platform memories separate.
Either way, every write is tagged with `metadata.channel` (for example `telegram` or `cli`), so per-channel views are still possible at query time.
## Reliability
- **Circuit breaker**: if Mem0 fails five times in a row, Hermes pauses calls for two minutes, then retries. The agent keeps working without memory during that window. Expected client errors, like a 404 on a missing memory id, do not count toward tripping the breaker.
- **Non-blocking**: fact extraction runs in a background daemon thread, and current-turn recall waits at most 3 seconds, so a slow or failed call never blocks your conversation.
- **Capture queue**: overlapping turns are queued in memory instead of skipped while an earlier turn is syncing. Failed requests are logged without durable retries. Shutdown waits at most five seconds; a process exit can lose pending turns.
- **Existing collections**: an OSS embedding-dimension mismatch fails initialization without deleting vectors. Restore the original embedding configuration or select a new collection name.
## Troubleshooting
### "Mem0 temporarily unavailable"
The circuit breaker tripped after five consecutive failures and resets after two minutes.
- **Platform mode**: check your API key and internet connection.
- **Self-hosted server mode**: check that the server is running and reachable at the configured `host` URL.
- **OSS mode**: make sure your vector store (Qdrant or PGVector) is running and reachable.
### OSS: vector store connection refused
```bash
# Local Qdrant: confirm the storage path is writable
ls -la ~/.hermes/mem0_qdrant
# Qdrant server: confirm it is reachable
curl http://localhost:6333/healthz
# PGVector: confirm PostgreSQL is accepting connections
pg_isready -h localhost -p 5432
```
### OSS: Ollama not reachable
```bash
curl http://localhost:11434/api/tags
```
### Memories not appearing
- `mem0_add` stores text after known-secret redaction with no extraction. Ordinary conversation turns are extracted automatically by the background sync.
- Search is semantic, so try a broader query.
- Confirm `user_id` is the same across sessions (check `~/.hermes/mem0.json`).
## Key Features
1. **Three ways to run**: managed Platform, a self-hosted server, or fully local OSS, switchable at any time.
2. **Current-turn recall**: memories for the current question are injected within a 3-second window, with `mem0_search` as the model's own backstop.
3. **Automatic extraction**: Mem0 extracts and deduplicates facts from each exchange for you.
4. **Non-blocking and fault tolerant**: background threads plus a circuit breaker keep the agent responsive even when Mem0 is unreachable.
5. **Additive memory**: works alongside Hermes' built-in file memory (`MEMORY.md`, `USER.md`).
Add memory to OpenClaw agents with auto-recall and auto-capture
Get your API key and explore the Mem0 dashboard