feat(oss): add reranker + per-search rerank to the TypeScript OSS SDK (#6055)
This commit is contained in:
@@ -26,7 +26,7 @@ All rerankers share these common configuration parameters:
|
||||
|
||||
| Parameter | Description | Type | Default |
|
||||
| -------------------- | -------------------------------------------- | ------ | ----------------------- |
|
||||
| `model` | Cohere rerank model | `str` | `"rerank-english-v3.0"` |
|
||||
| `model` | Cohere rerank model | `str` | `"rerank-v3.5"` |
|
||||
| `api_key` | Cohere API key | `str` | `None` |
|
||||
| `return_documents` | Whether to return document texts in response | `bool` | `False` |
|
||||
| `max_chunks_per_doc` | Maximum chunks per document | `int` | `None` |
|
||||
@@ -103,3 +103,30 @@ config = {
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## TypeScript SDK
|
||||
|
||||
The self-hosted [TypeScript SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) supports the same five providers. Config keys are camelCase (`apiKey`, `topK`, `maxLength`) and each provider's SDK is a peer dependency you install per reranker.
|
||||
|
||||
| Provider | Install | Default model | Key config fields |
|
||||
| --- | --- | --- | --- |
|
||||
| `cohere` | `pnpm add cohere-ai` | `rerank-v3.5` | `apiKey`, `model`, `topK` |
|
||||
| `zero_entropy` | `pnpm add zeroentropy` | `zerank-1` | `apiKey`, `model`, `topK` |
|
||||
| `sentence_transformer` | `pnpm add @huggingface/transformers` | `Xenova/ms-marco-MiniLM-L-6-v2` | `model`, `device`, `maxLength`, `normalize`, `topK` |
|
||||
| `huggingface` | `pnpm add @huggingface/transformers` | `Xenova/bge-reranker-base` | `model`, `device`, `maxLength`, `normalize`, `topK` |
|
||||
| `llm_reranker` | — (uses your LLM provider's own SDK) | `openai` / `gpt-4o-mini` | `provider`, `model`, `apiKey`, `llm` (nested override), `topK` |
|
||||
|
||||
```typescript
|
||||
import { Memory } from "mem0ai/oss";
|
||||
|
||||
const memory = new Memory({
|
||||
reranker: {
|
||||
provider: "zero_entropy",
|
||||
config: { apiKey: process.env.ZERO_ENTROPY_API_KEY, topK: 5 },
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
<Note>
|
||||
The local cross-encoder providers (`sentence_transformer`, `huggingface`) run on [Transformers.js](https://huggingface.co/docs/transformers.js) and default to ONNX (`Xenova/*`) model mirrors, so Python default model strings must be swapped for their ONNX equivalents. `batchSize` and `showProgressBar` are accepted for parity with Python but are no-ops in the TypeScript runtime. See the [reranker feature guide](/open-source/features/reranker-search#typescript-sdk) for full examples.
|
||||
</Note>
|
||||
|
||||
@@ -9,9 +9,9 @@ Cohere provides enterprise-grade reranking models with excellent multilingual su
|
||||
|
||||
Cohere offers several reranking models:
|
||||
|
||||
- **`rerank-english-v3.0`**: Latest English reranker with best performance
|
||||
- **`rerank-multilingual-v3.0`**: Multilingual support for global applications
|
||||
- **`rerank-english-v2.0`**: Previous generation English reranker
|
||||
- **`rerank-v3.5`** (default): Latest reranker, multilingual, best performance
|
||||
- **`rerank-english-v3.0`**: Previous generation, English only
|
||||
- **`rerank-multilingual-v3.0`**: Previous generation, multilingual
|
||||
|
||||
## Installation
|
||||
|
||||
@@ -41,7 +41,7 @@ config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"model": "rerank-english-v3.0",
|
||||
"model": "rerank-v3.5",
|
||||
"api_key": "your-cohere-api-key", # or set COHERE_API_KEY
|
||||
"top_k": 5,
|
||||
"return_documents": False,
|
||||
@@ -53,6 +53,34 @@ config = {
|
||||
memory = Memory.from_config(config)
|
||||
```
|
||||
|
||||
## TypeScript (self-hosted)
|
||||
|
||||
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) ships the Cohere reranker. Config keys are camelCase, it defaults to the `rerank-v3.5` model, and you opt in per search with `rerank: true`.
|
||||
|
||||
```bash
|
||||
pnpm add cohere-ai
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { Memory } from "mem0ai/oss";
|
||||
|
||||
const memory = new Memory({
|
||||
reranker: {
|
||||
provider: "cohere",
|
||||
config: {
|
||||
apiKey: process.env.COHERE_API_KEY, // or set COHERE_API_KEY
|
||||
// model: "rerank-v3.5", // default
|
||||
topK: 5,
|
||||
},
|
||||
},
|
||||
});
|
||||
|
||||
const results = await memory.search("What is the user's profession?", {
|
||||
filters: { userId: "bob" },
|
||||
rerank: true,
|
||||
});
|
||||
```
|
||||
|
||||
## Environment Variables
|
||||
|
||||
Set your API key as an environment variable:
|
||||
@@ -77,7 +105,7 @@ config = {
|
||||
"rerank": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"model": "rerank-english-v3.0",
|
||||
"model": "rerank-v3.5",
|
||||
"top_k": 3
|
||||
}
|
||||
}
|
||||
@@ -124,7 +152,7 @@ config = {
|
||||
|
||||
| Parameter | Description | Type | Default |
|
||||
| -------------------- | -------------------------------- | ------ | ----------------------- |
|
||||
| `model` | Cohere rerank model to use | `str` | `"rerank-english-v3.0"` |
|
||||
| `model` | Cohere rerank model to use | `str` | `"rerank-v3.5"` |
|
||||
| `api_key` | Cohere API key | `str` | `None` |
|
||||
| `top_k` | Maximum documents to return | `int` | `None` |
|
||||
| `return_documents` | Whether to return document texts | `bool` | `False` |
|
||||
@@ -139,7 +167,7 @@ config = {
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Model Selection**: Use `rerank-english-v3.0` for English, `rerank-multilingual-v3.0` for other languages
|
||||
1. **Model Selection**: `rerank-v3.5` handles English and multilingual workloads; pin an older `v3.0` model only if you need to reproduce prior results
|
||||
2. **Batch Processing**: Process multiple queries efficiently
|
||||
3. **Error Handling**: Implement retry logic for production systems
|
||||
4. **Monitoring**: Track reranking performance and costs
|
||||
|
||||
@@ -57,6 +57,40 @@ config = {
|
||||
}
|
||||
```
|
||||
|
||||
## TypeScript (self-hosted)
|
||||
|
||||
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) runs this reranker locally with [Transformers.js](https://huggingface.co/docs/transformers.js) — the same cross-encoder path as `sentence_transformer`, just a different default model. It executes ONNX weights, so the default is the ONNX mirror `Xenova/bge-reranker-base`. Point `model` at any ONNX-exported reranker on the Hub (a raw `BAAI/bge-reranker-*` PyTorch checkpoint will not load in this runtime).
|
||||
|
||||
```bash
|
||||
pnpm add @huggingface/transformers
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { Memory } from "mem0ai/oss";
|
||||
|
||||
const memory = new Memory({
|
||||
reranker: {
|
||||
provider: "huggingface",
|
||||
config: {
|
||||
// model: "Xenova/bge-reranker-base", // default (ONNX)
|
||||
device: "cpu", // "cpu" | "wasm" | "webgpu"
|
||||
maxLength: 512, // max tokens per query-document pair
|
||||
normalize: true, // sigmoid-normalize logits to [0, 1] (default)
|
||||
topK: 5,
|
||||
},
|
||||
},
|
||||
});
|
||||
|
||||
const results = await memory.search("What are the user's interests?", {
|
||||
filters: { userId: "alice" },
|
||||
rerank: true,
|
||||
});
|
||||
```
|
||||
|
||||
<Note>
|
||||
`batchSize` and `showProgressBar` are accepted for parity with the Python SDK but are no-ops in the TypeScript runtime. `trust_remote_code` and `model_kwargs` are Python-only.
|
||||
</Note>
|
||||
|
||||
## Popular Models
|
||||
|
||||
### BGE Rerankers (Recommended)
|
||||
|
||||
@@ -67,6 +67,43 @@ config = {
|
||||
}
|
||||
```
|
||||
|
||||
## TypeScript (self-hosted)
|
||||
|
||||
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) ships the LLM reranker under the provider name `llm_reranker`. It does **not** reuse the Memory's main `llm` instance — it builds its own LLM from the reranker's own config, defaulting to `openai` / `gpt-4o-mini`. Set `provider`/`model`/`apiKey` directly on `config`, or nest a fully separate `config.llm: { provider, config }` (its `provider`/`config` take priority over the top-level fields, which only backfill values missing from the nested config).
|
||||
|
||||
```typescript
|
||||
import { Memory } from "mem0ai/oss";
|
||||
|
||||
const memory = new Memory({
|
||||
reranker: {
|
||||
provider: "llm_reranker",
|
||||
config: { apiKey: process.env.OPENAI_API_KEY },
|
||||
},
|
||||
});
|
||||
|
||||
const results = await memory.search("What movies do I like?", {
|
||||
filters: { userId: "alice" },
|
||||
rerank: true,
|
||||
});
|
||||
```
|
||||
|
||||
To rerank with a different LLM provider than the Memory's main `llm`, nest it under `config.llm`:
|
||||
|
||||
```typescript
|
||||
const memory = new Memory({
|
||||
llm: { provider: "openai", config: { apiKey: process.env.OPENAI_API_KEY } },
|
||||
reranker: {
|
||||
provider: "llm_reranker",
|
||||
config: {
|
||||
llm: {
|
||||
provider: "anthropic",
|
||||
config: { apiKey: process.env.ANTHROPIC_API_KEY },
|
||||
},
|
||||
},
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
## Supported LLM Providers
|
||||
|
||||
### OpenAI
|
||||
|
||||
@@ -54,6 +54,40 @@ config = {
|
||||
memory = Memory.from_config(config)
|
||||
```
|
||||
|
||||
## TypeScript (self-hosted)
|
||||
|
||||
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) runs this reranker locally with [Transformers.js](https://huggingface.co/docs/transformers.js). Because it executes ONNX weights, the default model is the ONNX mirror of the Python default — `Xenova/ms-marco-MiniLM-L-6-v2`. Point `model` at any ONNX-exported cross-encoder on the Hub (a raw `cross-encoder/...` PyTorch checkpoint will not load in this runtime).
|
||||
|
||||
```bash
|
||||
pnpm add @huggingface/transformers
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { Memory } from "mem0ai/oss";
|
||||
|
||||
const memory = new Memory({
|
||||
reranker: {
|
||||
provider: "sentence_transformer",
|
||||
config: {
|
||||
// model: "Xenova/ms-marco-MiniLM-L-6-v2", // default (ONNX)
|
||||
device: "cpu", // "cpu" | "wasm" | "webgpu"
|
||||
maxLength: 512, // max tokens per query-document pair
|
||||
normalize: true, // sigmoid-normalize logits to [0, 1] (default)
|
||||
topK: 5,
|
||||
},
|
||||
},
|
||||
});
|
||||
|
||||
const results = await memory.search("What books does the user like?", {
|
||||
filters: { userId: "charlie" },
|
||||
rerank: true,
|
||||
});
|
||||
```
|
||||
|
||||
<Note>
|
||||
`batchSize` and `showProgressBar` are accepted for parity with the Python SDK but are no-ops in the TypeScript runtime — a search reranks a small candidate set in a single in-process forward pass. The model downloads once and is cached in-process.
|
||||
</Note>
|
||||
|
||||
## GPU Acceleration
|
||||
|
||||
For better performance, use GPU acceleration:
|
||||
|
||||
@@ -50,6 +50,34 @@ config = {
|
||||
memory = Memory.from_config(config)
|
||||
```
|
||||
|
||||
## TypeScript (self-hosted)
|
||||
|
||||
The [TypeScript OSS SDK](/open-source/features/reranker-search#typescript-sdk) (`mem0ai/oss`) ships the Zero Entropy reranker under the same provider name as Python, `zero_entropy`. It reads the key from config or `ZERO_ENTROPY_API_KEY` and defaults to the `zerank-1` model.
|
||||
|
||||
```bash
|
||||
pnpm add zeroentropy
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { Memory } from "mem0ai/oss";
|
||||
|
||||
const memory = new Memory({
|
||||
reranker: {
|
||||
provider: "zero_entropy",
|
||||
config: {
|
||||
apiKey: process.env.ZERO_ENTROPY_API_KEY,
|
||||
// model: "zerank-1", // default (or "zerank-1-small")
|
||||
topK: 5,
|
||||
},
|
||||
},
|
||||
});
|
||||
|
||||
const results = await memory.search("What Italian food does the user like?", {
|
||||
filters: { userId: "alice" },
|
||||
rerank: true,
|
||||
});
|
||||
```
|
||||
|
||||
## Environment Variables
|
||||
|
||||
Set your API key as an environment variable:
|
||||
|
||||
@@ -47,7 +47,7 @@ config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"model": "rerank-english-v3.0",
|
||||
"model": "rerank-v3.5",
|
||||
"top_n": 10,
|
||||
"max_chunks_per_doc": 10, # Limit chunk processing
|
||||
"return_documents": False # Reduce response size
|
||||
@@ -280,7 +280,7 @@ config = {
|
||||
```python
|
||||
def benchmark_rerankers():
|
||||
configs = [
|
||||
{"provider": "cohere", "model": "rerank-english-v3.0"},
|
||||
{"provider": "cohere", "model": "rerank-v3.5"},
|
||||
{"provider": "sentence_transformer", "model": "cross-encoder/ms-marco-MiniLM-L-6-v2"},
|
||||
{"provider": "huggingface", "model": "BAAI/bge-reranker-base"}
|
||||
]
|
||||
|
||||
@@ -19,6 +19,10 @@ Reranking trades extra latency for better precision. Start once you have baselin
|
||||
<Card title="Zero Entropy" icon="/images/provider-icons/zeroentropy.svg" href="/components/rerankers/models/zero_entropy" />
|
||||
</CardGroup>
|
||||
|
||||
<Note>
|
||||
All five rerankers are available in both the Python and the [TypeScript](/open-source/features/reranker-search#typescript-sdk) self-hosted SDKs. Each provider page has a **TypeScript (self-hosted)** section with the camelCase config.
|
||||
</Note>
|
||||
|
||||
## Reranking Workflow
|
||||
|
||||
<CardGroup cols={3}>
|
||||
|
||||
Reference in New Issue
Block a user