feat(oss): add reranker + per-search rerank to the TypeScript OSS SDK (#6055)

This commit is contained in:
Kartik
2026-07-09 22:18:38 +05:30
committed by GitHub
parent 2a4aa232b2
commit 6a801bfe2f
32 changed files with 2173 additions and 19 deletions
+28 -1
View File
@@ -26,7 +26,7 @@ All rerankers share these common configuration parameters:
| Parameter | Description | Type | Default |
| -------------------- | -------------------------------------------- | ------ | ----------------------- |
| `model` | Cohere rerank model | `str` | `"rerank-english-v3.0"` |
| `model` | Cohere rerank model | `str` | `"rerank-v3.5"` |
| `api_key` | Cohere API key | `str` | `None` |
| `return_documents` | Whether to return document texts in response | `bool` | `False` |
| `max_chunks_per_doc` | Maximum chunks per document | `int` | `None` |
@@ -103,3 +103,30 @@ config = {
}
}
```
## TypeScript SDK
The self-hosted [TypeScript SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) supports the same five providers. Config keys are camelCase (`apiKey`, `topK`, `maxLength`) and each provider's SDK is a peer dependency you install per reranker.
| Provider | Install | Default model | Key config fields |
| --- | --- | --- | --- |
| `cohere` | `pnpm add cohere-ai` | `rerank-v3.5` | `apiKey`, `model`, `topK` |
| `zero_entropy` | `pnpm add zeroentropy` | `zerank-1` | `apiKey`, `model`, `topK` |
| `sentence_transformer` | `pnpm add @huggingface/transformers` | `Xenova/ms-marco-MiniLM-L-6-v2` | `model`, `device`, `maxLength`, `normalize`, `topK` |
| `huggingface` | `pnpm add @huggingface/transformers` | `Xenova/bge-reranker-base` | `model`, `device`, `maxLength`, `normalize`, `topK` |
| `llm_reranker` | — (uses your LLM provider's own SDK) | `openai` / `gpt-4o-mini` | `provider`, `model`, `apiKey`, `llm` (nested override), `topK` |
```typescript
import { Memory } from "mem0ai/oss";
const memory = new Memory({
reranker: {
provider: "zero_entropy",
config: { apiKey: process.env.ZERO_ENTROPY_API_KEY, topK: 5 },
},
});
```
<Note>
The local cross-encoder providers (`sentence_transformer`, `huggingface`) run on [Transformers.js](https://huggingface.co/docs/transformers.js) and default to ONNX (`Xenova/*`) model mirrors, so Python default model strings must be swapped for their ONNX equivalents. `batchSize` and `showProgressBar` are accepted for parity with Python but are no-ops in the TypeScript runtime. See the [reranker feature guide](/open-source/features/reranker-search#typescript-sdk) for full examples.
</Note>
+35 -7
View File
@@ -9,9 +9,9 @@ Cohere provides enterprise-grade reranking models with excellent multilingual su
Cohere offers several reranking models:
- **`rerank-english-v3.0`**: Latest English reranker with best performance
- **`rerank-multilingual-v3.0`**: Multilingual support for global applications
- **`rerank-english-v2.0`**: Previous generation English reranker
- **`rerank-v3.5`** (default): Latest reranker, multilingual, best performance
- **`rerank-english-v3.0`**: Previous generation, English only
- **`rerank-multilingual-v3.0`**: Previous generation, multilingual
## Installation
@@ -41,7 +41,7 @@ config = {
"reranker": {
"provider": "cohere",
"config": {
"model": "rerank-english-v3.0",
"model": "rerank-v3.5",
"api_key": "your-cohere-api-key", # or set COHERE_API_KEY
"top_k": 5,
"return_documents": False,
@@ -53,6 +53,34 @@ config = {
memory = Memory.from_config(config)
```
## TypeScript (self-hosted)
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) ships the Cohere reranker. Config keys are camelCase, it defaults to the `rerank-v3.5` model, and you opt in per search with `rerank: true`.
```bash
pnpm add cohere-ai
```
```typescript
import { Memory } from "mem0ai/oss";
const memory = new Memory({
reranker: {
provider: "cohere",
config: {
apiKey: process.env.COHERE_API_KEY, // or set COHERE_API_KEY
// model: "rerank-v3.5", // default
topK: 5,
},
},
});
const results = await memory.search("What is the user's profession?", {
filters: { userId: "bob" },
rerank: true,
});
```
## Environment Variables
Set your API key as an environment variable:
@@ -77,7 +105,7 @@ config = {
"rerank": {
"provider": "cohere",
"config": {
"model": "rerank-english-v3.0",
"model": "rerank-v3.5",
"top_k": 3
}
}
@@ -124,7 +152,7 @@ config = {
| Parameter | Description | Type | Default |
| -------------------- | -------------------------------- | ------ | ----------------------- |
| `model` | Cohere rerank model to use | `str` | `"rerank-english-v3.0"` |
| `model` | Cohere rerank model to use | `str` | `"rerank-v3.5"` |
| `api_key` | Cohere API key | `str` | `None` |
| `top_k` | Maximum documents to return | `int` | `None` |
| `return_documents` | Whether to return document texts | `bool` | `False` |
@@ -139,7 +167,7 @@ config = {
## Best Practices
1. **Model Selection**: Use `rerank-english-v3.0` for English, `rerank-multilingual-v3.0` for other languages
1. **Model Selection**: `rerank-v3.5` handles English and multilingual workloads; pin an older `v3.0` model only if you need to reproduce prior results
2. **Batch Processing**: Process multiple queries efficiently
3. **Error Handling**: Implement retry logic for production systems
4. **Monitoring**: Track reranking performance and costs
@@ -57,6 +57,40 @@ config = {
}
```
## TypeScript (self-hosted)
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) runs this reranker locally with [Transformers.js](https://huggingface.co/docs/transformers.js) — the same cross-encoder path as `sentence_transformer`, just a different default model. It executes ONNX weights, so the default is the ONNX mirror `Xenova/bge-reranker-base`. Point `model` at any ONNX-exported reranker on the Hub (a raw `BAAI/bge-reranker-*` PyTorch checkpoint will not load in this runtime).
```bash
pnpm add @huggingface/transformers
```
```typescript
import { Memory } from "mem0ai/oss";
const memory = new Memory({
reranker: {
provider: "huggingface",
config: {
// model: "Xenova/bge-reranker-base", // default (ONNX)
device: "cpu", // "cpu" | "wasm" | "webgpu"
maxLength: 512, // max tokens per query-document pair
normalize: true, // sigmoid-normalize logits to [0, 1] (default)
topK: 5,
},
},
});
const results = await memory.search("What are the user's interests?", {
filters: { userId: "alice" },
rerank: true,
});
```
<Note>
`batchSize` and `showProgressBar` are accepted for parity with the Python SDK but are no-ops in the TypeScript runtime. `trust_remote_code` and `model_kwargs` are Python-only.
</Note>
## Popular Models
### BGE Rerankers (Recommended)
@@ -67,6 +67,43 @@ config = {
}
```
## TypeScript (self-hosted)
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) ships the LLM reranker under the provider name `llm_reranker`. It does **not** reuse the Memory's main `llm` instance — it builds its own LLM from the reranker's own config, defaulting to `openai` / `gpt-4o-mini`. Set `provider`/`model`/`apiKey` directly on `config`, or nest a fully separate `config.llm: { provider, config }` (its `provider`/`config` take priority over the top-level fields, which only backfill values missing from the nested config).
```typescript
import { Memory } from "mem0ai/oss";
const memory = new Memory({
reranker: {
provider: "llm_reranker",
config: { apiKey: process.env.OPENAI_API_KEY },
},
});
const results = await memory.search("What movies do I like?", {
filters: { userId: "alice" },
rerank: true,
});
```
To rerank with a different LLM provider than the Memory's main `llm`, nest it under `config.llm`:
```typescript
const memory = new Memory({
llm: { provider: "openai", config: { apiKey: process.env.OPENAI_API_KEY } },
reranker: {
provider: "llm_reranker",
config: {
llm: {
provider: "anthropic",
config: { apiKey: process.env.ANTHROPIC_API_KEY },
},
},
},
});
```
## Supported LLM Providers
### OpenAI
@@ -54,6 +54,40 @@ config = {
memory = Memory.from_config(config)
```
## TypeScript (self-hosted)
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) runs this reranker locally with [Transformers.js](https://huggingface.co/docs/transformers.js). Because it executes ONNX weights, the default model is the ONNX mirror of the Python default — `Xenova/ms-marco-MiniLM-L-6-v2`. Point `model` at any ONNX-exported cross-encoder on the Hub (a raw `cross-encoder/...` PyTorch checkpoint will not load in this runtime).
```bash
pnpm add @huggingface/transformers
```
```typescript
import { Memory } from "mem0ai/oss";
const memory = new Memory({
reranker: {
provider: "sentence_transformer",
config: {
// model: "Xenova/ms-marco-MiniLM-L-6-v2", // default (ONNX)
device: "cpu", // "cpu" | "wasm" | "webgpu"
maxLength: 512, // max tokens per query-document pair
normalize: true, // sigmoid-normalize logits to [0, 1] (default)
topK: 5,
},
},
});
const results = await memory.search("What books does the user like?", {
filters: { userId: "charlie" },
rerank: true,
});
```
<Note>
`batchSize` and `showProgressBar` are accepted for parity with the Python SDK but are no-ops in the TypeScript runtime — a search reranks a small candidate set in a single in-process forward pass. The model downloads once and is cached in-process.
</Note>
## GPU Acceleration
For better performance, use GPU acceleration:
@@ -50,6 +50,34 @@ config = {
memory = Memory.from_config(config)
```
## TypeScript (self-hosted)
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) ships the Zero Entropy reranker under the same provider name as Python, `zero_entropy`. It reads the key from config or `ZERO_ENTROPY_API_KEY` and defaults to the `zerank-1` model.
```bash
pnpm add zeroentropy
```
```typescript
import { Memory } from "mem0ai/oss";
const memory = new Memory({
reranker: {
provider: "zero_entropy",
config: {
apiKey: process.env.ZERO_ENTROPY_API_KEY,
// model: "zerank-1", // default (or "zerank-1-small")
topK: 5,
},
},
});
const results = await memory.search("What Italian food does the user like?", {
filters: { userId: "alice" },
rerank: true,
});
```
## Environment Variables
Set your API key as an environment variable:
+2 -2
View File
@@ -47,7 +47,7 @@ config = {
"reranker": {
"provider": "cohere",
"config": {
"model": "rerank-english-v3.0",
"model": "rerank-v3.5",
"top_n": 10,
"max_chunks_per_doc": 10, # Limit chunk processing
"return_documents": False # Reduce response size
@@ -280,7 +280,7 @@ config = {
```python
def benchmark_rerankers():
configs = [
{"provider": "cohere", "model": "rerank-english-v3.0"},
{"provider": "cohere", "model": "rerank-v3.5"},
{"provider": "sentence_transformer", "model": "cross-encoder/ms-marco-MiniLM-L-6-v2"},
{"provider": "huggingface", "model": "BAAI/bge-reranker-base"}
]
+4
View File
@@ -19,6 +19,10 @@ Reranking trades extra latency for better precision. Start once you have baselin
<Card title="Zero Entropy" icon="/images/provider-icons/zeroentropy.svg" href="/components/rerankers/models/zero_entropy" />
</CardGroup>
<Note>
All five rerankers are available in both the Python and the [TypeScript](/open-source/features/reranker-search#typescript-sdk) self-hosted SDKs. Each provider page has a **TypeScript (self-hosted)** section with the camelCase config.
</Note>
## Reranking Workflow
<CardGroup cols={3}>