Aspen theme for 1.x (#3473)
This commit is contained in:
@@ -28,8 +28,8 @@ See the list of supported LLMs below.
|
||||
<Card title="Together" href="/components/llms/models/together" />
|
||||
<Card title="Groq" href="/components/llms/models/groq" />
|
||||
<Card title="Litellm" href="/components/llms/models/litellm" />
|
||||
<Card title="Mistral AI" href="/components/llms/models/mistral_ai" />
|
||||
<Card title="Google AI" href="/components/llms/models/google_ai" />
|
||||
<Card title="Mistral AI" href="/components/llms/models/mistral_AI" />
|
||||
<Card title="Google AI" href="/components/llms/models/google_AI" />
|
||||
<Card title="AWS bedrock" href="/components/llms/models/aws_bedrock" />
|
||||
<Card title="DeepSeek" href="/components/llms/models/deepseek" />
|
||||
<Card title="xAI" href="/components/llms/models/xAI" />
|
||||
|
||||
@@ -0,0 +1,145 @@
|
||||
---
|
||||
title: Configuration
|
||||
icon: "gear"
|
||||
iconType: "solid"
|
||||
---
|
||||
|
||||
## How to define configurations?
|
||||
|
||||
The `reranker` configuration is defined as an object with two main keys:
|
||||
- `provider`: The name of the reranker provider (e.g., "cohere", "sentence_transformer", "huggingface", "llm_reranker")
|
||||
- `config`: A nested dictionary containing provider-specific settings
|
||||
|
||||
## Basic Configuration
|
||||
|
||||
Here's how to configure a reranker with Mem0:
|
||||
|
||||
```python
|
||||
from mem0 import Memory
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"api_key": "your-api-key",
|
||||
"top_n": 10,
|
||||
"model": "rerank-english-v3.0"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
memory = Memory.from_config(config)
|
||||
```
|
||||
|
||||
## Configuration Parameters
|
||||
|
||||
| Parameter | Description | Required | Default |
|
||||
|-----------|-------------|----------|---------|
|
||||
| `provider` | Reranker provider name | Yes | - |
|
||||
| `config` | Provider-specific configuration | Yes | - |
|
||||
|
||||
### Common Config Parameters
|
||||
|
||||
| Parameter | Description | Providers |
|
||||
|-----------|-------------|-----------|
|
||||
| `api_key` | API key for the service | Cohere, Hugging Face |
|
||||
| `model` | Model name to use | All |
|
||||
| `top_n` | Number of results to return | All |
|
||||
| `device` | Device to run on (cpu/cuda/mps) | Sentence Transformer, Hugging Face |
|
||||
|
||||
## Provider-Specific Examples
|
||||
|
||||
### Cohere
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"api_key": "your-cohere-api-key",
|
||||
"model": "rerank-english-v3.0",
|
||||
"top_n": 5
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Sentence Transformer
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"device": "cpu",
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Hugging Face
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"api_key": "your-hf-token",
|
||||
"model": "BAAI/bge-reranker-large",
|
||||
"top_n": 8
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### LLM Reranker
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "openai",
|
||||
"config": {
|
||||
"model": "gpt-4",
|
||||
"api_key": "your-openai-key"
|
||||
}
|
||||
},
|
||||
"top_n": 5
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Advanced Configuration
|
||||
|
||||
You can combine rerankers with other components:
|
||||
|
||||
```python
|
||||
config = {
|
||||
"llm": {
|
||||
"provider": "openai",
|
||||
"config": {
|
||||
"model": "gpt-4",
|
||||
"api_key": "your-openai-key"
|
||||
}
|
||||
},
|
||||
"vector_store": {
|
||||
"provider": "qdrant",
|
||||
"config": {
|
||||
"collection_name": "memories",
|
||||
"host": "localhost",
|
||||
"port": 6333
|
||||
}
|
||||
},
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"api_key": "your-cohere-key",
|
||||
"model": "rerank-english-v3.0",
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
For provider-specific configuration details, visit the individual reranker pages.
|
||||
@@ -0,0 +1,217 @@
|
||||
---
|
||||
title: Custom Prompts
|
||||
icon: "pencil"
|
||||
iconType: "solid"
|
||||
---
|
||||
|
||||
When using LLM rerankers, you can customize the prompts used for ranking to better suit your specific use case and domain.
|
||||
|
||||
## Default Prompt
|
||||
|
||||
The default LLM reranker prompt is designed to be general-purpose:
|
||||
|
||||
```
|
||||
Given a query and a list of memory entries, rank the memory entries based on their relevance to the query.
|
||||
Rate each memory on a scale of 1-10 where 10 is most relevant.
|
||||
|
||||
Query: {query}
|
||||
|
||||
Memory entries:
|
||||
{memories}
|
||||
|
||||
Provide your ranking as a JSON array with scores for each memory.
|
||||
```
|
||||
|
||||
## Custom Prompt Configuration
|
||||
|
||||
You can provide a custom prompt template when configuring the LLM reranker:
|
||||
|
||||
```python
|
||||
from mem0 import Memory
|
||||
|
||||
custom_prompt = """
|
||||
You are an expert at ranking memories for a personal AI assistant.
|
||||
Given a user query and a list of memory entries, rank each memory based on:
|
||||
1. Direct relevance to the query
|
||||
2. Temporal relevance (recent memories may be more important)
|
||||
3. Emotional significance
|
||||
4. Actionability
|
||||
|
||||
Query: {query}
|
||||
User Context: {user_context}
|
||||
|
||||
Memory entries:
|
||||
{memories}
|
||||
|
||||
Rate each memory from 1-10 and provide reasoning.
|
||||
Return as JSON: {{"rankings": [{{"index": 0, "score": 8, "reason": "..."}}]}}
|
||||
"""
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "openai",
|
||||
"config": {
|
||||
"model": "gpt-4",
|
||||
"api_key": "your-openai-key"
|
||||
}
|
||||
},
|
||||
"custom_prompt": custom_prompt,
|
||||
"top_n": 5
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
memory = Memory.from_config(config)
|
||||
```
|
||||
|
||||
## Prompt Variables
|
||||
|
||||
Your custom prompt can use the following variables:
|
||||
|
||||
| Variable | Description |
|
||||
|----------|-------------|
|
||||
| `{query}` | The search query |
|
||||
| `{memories}` | The list of memory entries to rank |
|
||||
| `{user_id}` | The user ID (if available) |
|
||||
| `{user_context}` | Additional user context (if provided) |
|
||||
|
||||
## Domain-Specific Examples
|
||||
|
||||
### Customer Support
|
||||
```python
|
||||
customer_support_prompt = """
|
||||
You are ranking customer support conversation memories.
|
||||
Prioritize memories that:
|
||||
- Relate to the current customer issue
|
||||
- Show previous resolution patterns
|
||||
- Indicate customer preferences or constraints
|
||||
|
||||
Query: {query}
|
||||
Customer Context: Previous interactions with this customer
|
||||
|
||||
Memories:
|
||||
{memories}
|
||||
|
||||
Rank each memory 1-10 based on support relevance.
|
||||
"""
|
||||
```
|
||||
|
||||
### Educational Content
|
||||
```python
|
||||
educational_prompt = """
|
||||
Rank these learning memories for a student query.
|
||||
Consider:
|
||||
- Prerequisite knowledge requirements
|
||||
- Learning progression and difficulty
|
||||
- Relevance to current learning objectives
|
||||
|
||||
Student Query: {query}
|
||||
Learning Context: {user_context}
|
||||
|
||||
Available memories:
|
||||
{memories}
|
||||
|
||||
Score each memory for educational value (1-10).
|
||||
"""
|
||||
```
|
||||
|
||||
### Personal Assistant
|
||||
```python
|
||||
personal_assistant_prompt = """
|
||||
Rank personal memories for relevance to the user's query.
|
||||
Consider:
|
||||
- Recent vs. historical importance
|
||||
- Personal preferences and habits
|
||||
- Contextual relationships between memories
|
||||
|
||||
Query: {query}
|
||||
Personal context: {user_context}
|
||||
|
||||
Memories to rank:
|
||||
{memories}
|
||||
|
||||
Provide relevance scores (1-10) with brief explanations.
|
||||
"""
|
||||
```
|
||||
|
||||
## Advanced Prompt Techniques
|
||||
|
||||
### Multi-Criteria Ranking
|
||||
```python
|
||||
multi_criteria_prompt = """
|
||||
Evaluate memories using multiple criteria:
|
||||
|
||||
1. RELEVANCE (40%): How directly related to the query
|
||||
2. RECENCY (20%): How recent the memory is
|
||||
3. IMPORTANCE (25%): Personal or business significance
|
||||
4. ACTIONABILITY (15%): How useful for next steps
|
||||
|
||||
Query: {query}
|
||||
Context: {user_context}
|
||||
|
||||
Memories:
|
||||
{memories}
|
||||
|
||||
For each memory, provide:
|
||||
- Overall score (1-10)
|
||||
- Breakdown by criteria
|
||||
- Final ranking recommendation
|
||||
|
||||
Format: JSON with detailed scoring
|
||||
"""
|
||||
```
|
||||
|
||||
### Contextual Ranking
|
||||
```python
|
||||
contextual_prompt = """
|
||||
Consider the following context when ranking memories:
|
||||
- Current user situation: {user_context}
|
||||
- Time of day: {current_time}
|
||||
- Recent activities: {recent_activities}
|
||||
|
||||
Query: {query}
|
||||
|
||||
Rank these memories considering both direct relevance and contextual appropriateness:
|
||||
{memories}
|
||||
|
||||
Provide contextually-aware relevance scores (1-10).
|
||||
"""
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Be Specific**: Clearly define what makes a memory relevant for your use case
|
||||
2. **Use Examples**: Include examples in your prompt for better model understanding
|
||||
3. **Structure Output**: Specify the exact JSON format you want returned
|
||||
4. **Test Iteratively**: Refine your prompt based on actual ranking performance
|
||||
5. **Consider Token Limits**: Keep prompts concise while being comprehensive
|
||||
|
||||
## Prompt Testing
|
||||
|
||||
You can test different prompts by comparing ranking results:
|
||||
|
||||
```python
|
||||
# Test multiple prompt variations
|
||||
prompts = [
|
||||
default_prompt,
|
||||
custom_prompt_v1,
|
||||
custom_prompt_v2
|
||||
]
|
||||
|
||||
for i, prompt in enumerate(prompts):
|
||||
config["reranker"]["config"]["custom_prompt"] = prompt
|
||||
memory = Memory.from_config(config)
|
||||
|
||||
results = memory.search("test query", user_id="test_user")
|
||||
print(f"Prompt {i+1} results: {results}")
|
||||
```
|
||||
|
||||
## Common Issues
|
||||
|
||||
- **Too Long**: Keep prompts under token limits for your chosen LLM
|
||||
- **Too Vague**: Be specific about ranking criteria
|
||||
- **Inconsistent Format**: Ensure JSON output format is clearly specified
|
||||
- **Missing Context**: Include relevant variables for your use case
|
||||
@@ -0,0 +1,116 @@
|
||||
---
|
||||
title: Cohere
|
||||
---
|
||||
|
||||
Cohere provides state-of-the-art reranking models that can significantly improve the relevance of search results. Cohere's rerankers are optimized for various languages and use cases.
|
||||
|
||||
## Usage
|
||||
|
||||
To use Cohere's reranker with Mem0:
|
||||
|
||||
```python
|
||||
import os
|
||||
from mem0 import Memory
|
||||
|
||||
os.environ["COHERE_API_KEY"] = "your-cohere-api-key"
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"api_key": "your-cohere-api-key", # Can also use environment variable
|
||||
"model": "rerank-english-v3.0",
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
memory = Memory.from_config(config)
|
||||
|
||||
# Use memory as usual
|
||||
memory.add("I love playing basketball", user_id="alice")
|
||||
memory.add("I enjoy watching movies", user_id="alice")
|
||||
|
||||
# Search will now use Cohere reranking
|
||||
results = memory.search("What sports does Alice like?", user_id="alice")
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
| Parameter | Description | Default |
|
||||
|-----------|-------------|---------|
|
||||
| `api_key` | Cohere API key | Required |
|
||||
| `model` | Cohere rerank model | `rerank-english-v3.0` |
|
||||
| `top_n` | Number of results to return | `10` |
|
||||
|
||||
## Available Models
|
||||
|
||||
- `rerank-english-v3.0`: Latest English reranking model
|
||||
- `rerank-multilingual-v3.0`: Multilingual reranking model
|
||||
- `rerank-english-v2.0`: Previous English model
|
||||
- `rerank-multilingual-v2.0`: Previous multilingual model
|
||||
|
||||
## Example with Different Models
|
||||
|
||||
### English Reranker
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"api_key": "your-cohere-api-key",
|
||||
"model": "rerank-english-v3.0",
|
||||
"top_n": 5
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Multilingual Reranker
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"api_key": "your-cohere-api-key",
|
||||
"model": "rerank-multilingual-v3.0",
|
||||
"top_n": 8
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Environment Variables
|
||||
|
||||
You can set your Cohere API key as an environment variable:
|
||||
|
||||
```bash
|
||||
export COHERE_API_KEY="your-cohere-api-key"
|
||||
```
|
||||
|
||||
Then use the config without specifying the API key:
|
||||
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"model": "rerank-english-v3.0",
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Getting Your API Key
|
||||
|
||||
1. Sign up at [Cohere](https://cohere.ai/)
|
||||
2. Navigate to the API keys section in your dashboard
|
||||
3. Generate a new API key
|
||||
4. Use this key in your configuration
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
- Cohere rerankers work best with 10-100 candidate documents
|
||||
- Higher `top_n` values provide more comprehensive reranking but may increase latency
|
||||
- The v3.0 models generally provide better performance than v2.0 models
|
||||
@@ -0,0 +1,352 @@
|
||||
---
|
||||
title: Hugging Face Reranker
|
||||
description: 'Access thousands of reranking models from Hugging Face Hub'
|
||||
icon: "face-smile"
|
||||
iconType: "solid"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
The Hugging Face reranker provider gives you access to thousands of reranking models available on the Hugging Face Hub. This includes popular models like BAAI's BGE rerankers and other state-of-the-art cross-encoder models.
|
||||
|
||||
## Configuration
|
||||
|
||||
### Basic Setup
|
||||
|
||||
```python
|
||||
from mem0 import Memory
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"device": "cpu"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
m = Memory.from_config(config)
|
||||
```
|
||||
|
||||
### Configuration Parameters
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `model` | str | Required | Hugging Face model identifier |
|
||||
| `device` | str | "cpu" | Device to run model on ("cpu", "cuda", "mps") |
|
||||
| `batch_size` | int | 32 | Batch size for processing |
|
||||
| `max_length` | int | 512 | Maximum input sequence length |
|
||||
| `trust_remote_code` | bool | False | Allow remote code execution |
|
||||
|
||||
### Advanced Configuration
|
||||
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-large",
|
||||
"device": "cuda",
|
||||
"batch_size": 16,
|
||||
"max_length": 512,
|
||||
"trust_remote_code": False,
|
||||
"model_kwargs": {
|
||||
"torch_dtype": "float16"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Popular Models
|
||||
|
||||
### BGE Rerankers (Recommended)
|
||||
|
||||
```python
|
||||
# Base model - good balance of speed and quality
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"device": "cuda"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
# Large model - better quality, slower
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-large",
|
||||
"device": "cuda"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
# v2 models - latest improvements
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-v2-m3",
|
||||
"device": "cuda"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Multilingual Models
|
||||
|
||||
```python
|
||||
# Multilingual BGE reranker
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-v2-multilingual",
|
||||
"device": "cuda"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Domain-Specific Models
|
||||
|
||||
```python
|
||||
# For code search
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "microsoft/codebert-base",
|
||||
"device": "cuda"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
# For biomedical content
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "dmis-lab/biobert-base-cased-v1.1",
|
||||
"device": "cuda"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Usage Examples
|
||||
|
||||
### Basic Usage
|
||||
|
||||
```python
|
||||
from mem0 import Memory
|
||||
|
||||
m = Memory.from_config(config)
|
||||
|
||||
# Add some memories
|
||||
m.add("I love hiking in the mountains", user_id="alice")
|
||||
m.add("Pizza is my favorite food", user_id="alice")
|
||||
m.add("I enjoy reading science fiction books", user_id="alice")
|
||||
|
||||
# Search with reranking
|
||||
results = m.search(
|
||||
"What outdoor activities do I enjoy?",
|
||||
user_id="alice",
|
||||
rerank=True
|
||||
)
|
||||
|
||||
for result in results["results"]:
|
||||
print(f"Memory: {result['memory']}")
|
||||
print(f"Score: {result['score']:.3f}")
|
||||
```
|
||||
|
||||
### Batch Processing
|
||||
|
||||
```python
|
||||
# Process multiple queries efficiently
|
||||
queries = [
|
||||
"What are my hobbies?",
|
||||
"What food do I like?",
|
||||
"What books interest me?"
|
||||
]
|
||||
|
||||
results = []
|
||||
for query in queries:
|
||||
result = m.search(query, user_id="alice", rerank=True)
|
||||
results.append(result)
|
||||
```
|
||||
|
||||
## Performance Optimization
|
||||
|
||||
### GPU Acceleration
|
||||
|
||||
```python
|
||||
# Use GPU for better performance
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"device": "cuda",
|
||||
"batch_size": 64, # Increase batch size for GPU
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Memory Optimization
|
||||
|
||||
```python
|
||||
# For limited memory environments
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"device": "cpu",
|
||||
"batch_size": 8, # Smaller batch size
|
||||
"max_length": 256, # Shorter sequences
|
||||
"model_kwargs": {
|
||||
"torch_dtype": "float16" # Half precision
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Model Comparison
|
||||
|
||||
| Model | Size | Quality | Speed | Memory | Best For |
|
||||
|-------|------|---------|-------|---------|----------|
|
||||
| bge-reranker-base | 278M | Good | Fast | Low | General use |
|
||||
| bge-reranker-large | 560M | Better | Medium | Medium | High quality needs |
|
||||
| bge-reranker-v2-m3 | 568M | Best | Medium | Medium | Latest improvements |
|
||||
| bge-reranker-v2-multilingual | 568M | Good | Medium | Medium | Multiple languages |
|
||||
|
||||
## Error Handling
|
||||
|
||||
```python
|
||||
try:
|
||||
results = m.search(
|
||||
"test query",
|
||||
user_id="alice",
|
||||
rerank=True
|
||||
)
|
||||
except Exception as e:
|
||||
print(f"Reranking failed: {e}")
|
||||
# Fall back to vector search only
|
||||
results = m.search(
|
||||
"test query",
|
||||
user_id="alice",
|
||||
rerank=False
|
||||
)
|
||||
```
|
||||
|
||||
## Custom Models
|
||||
|
||||
### Using Private Models
|
||||
|
||||
```python
|
||||
# Use a private model from Hugging Face
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "your-org/custom-reranker",
|
||||
"device": "cuda",
|
||||
"use_auth_token": "your-hf-token"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Local Model Path
|
||||
|
||||
```python
|
||||
# Use a locally downloaded model
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "/path/to/local/model",
|
||||
"device": "cuda"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Choose the Right Model**: Balance quality vs speed based on your needs
|
||||
2. **Use GPU**: Significantly faster than CPU for larger models
|
||||
3. **Optimize Batch Size**: Tune based on your hardware capabilities
|
||||
4. **Monitor Memory**: Watch GPU/CPU memory usage with large models
|
||||
5. **Cache Models**: Download once and reuse to avoid repeated downloads
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Common Issues
|
||||
|
||||
**Out of Memory Error**
|
||||
```python
|
||||
# Reduce batch size and sequence length
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"batch_size": 4,
|
||||
"max_length": 256
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Model Download Issues**
|
||||
```python
|
||||
# Set cache directory
|
||||
import os
|
||||
os.environ["TRANSFORMERS_CACHE"] = "/path/to/cache"
|
||||
|
||||
# Or use offline mode
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"local_files_only": True
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**CUDA Not Available**
|
||||
```python
|
||||
import torch
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"device": "cuda" if torch.cuda.is_available() else "cpu"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Reranker Overview" icon="sort" href="/components/rerankers/overview">
|
||||
Learn about reranking concepts
|
||||
</Card>
|
||||
<Card title="Configuration Guide" icon="gear" href="/components/rerankers/config">
|
||||
Detailed configuration options
|
||||
</Card>
|
||||
</CardGroup>
|
||||
@@ -0,0 +1,491 @@
|
||||
---
|
||||
title: LLM Reranker
|
||||
description: 'Use any language model as a reranker with custom prompts'
|
||||
icon: "robot"
|
||||
iconType: "solid"
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
The LLM reranker allows you to use any supported language model as a reranker. This approach uses prompts to instruct the LLM to score and rank memories based on their relevance to the query. While slower than specialized rerankers, it offers maximum flexibility and can be fine-tuned with custom prompts.
|
||||
|
||||
## Configuration
|
||||
|
||||
### Basic Setup
|
||||
|
||||
```python
|
||||
from mem0 import Memory
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "openai",
|
||||
"config": {
|
||||
"model": "gpt-4",
|
||||
"api_key": "your-openai-api-key"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
m = Memory.from_config(config)
|
||||
```
|
||||
|
||||
### Configuration Parameters
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `llm` | dict | Required | LLM configuration object |
|
||||
| `top_k` | int | 10 | Number of results to rerank |
|
||||
| `temperature` | float | 0.0 | LLM temperature for consistency |
|
||||
| `custom_prompt` | str | None | Custom reranking prompt |
|
||||
| `score_range` | tuple | (0, 10) | Score range for relevance |
|
||||
|
||||
### Advanced Configuration
|
||||
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "anthropic",
|
||||
"config": {
|
||||
"model": "claude-3-sonnet-20240229",
|
||||
"api_key": "your-anthropic-api-key"
|
||||
}
|
||||
},
|
||||
"top_k": 15,
|
||||
"temperature": 0.0,
|
||||
"score_range": (1, 5),
|
||||
"custom_prompt": """
|
||||
Rate the relevance of each memory to the query on a scale of 1-5.
|
||||
Consider semantic similarity, context, and practical utility.
|
||||
Only provide the numeric score.
|
||||
"""
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Supported LLM Providers
|
||||
|
||||
### OpenAI
|
||||
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "openai",
|
||||
"config": {
|
||||
"model": "gpt-4",
|
||||
"api_key": "your-openai-api-key",
|
||||
"temperature": 0.0
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Anthropic
|
||||
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "anthropic",
|
||||
"config": {
|
||||
"model": "claude-3-sonnet-20240229",
|
||||
"api_key": "your-anthropic-api-key"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Ollama (Local)
|
||||
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "ollama",
|
||||
"config": {
|
||||
"model": "llama2",
|
||||
"ollama_base_url": "http://localhost:11434"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Azure OpenAI
|
||||
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "azure_openai",
|
||||
"config": {
|
||||
"model": "gpt-4",
|
||||
"api_key": "your-azure-api-key",
|
||||
"azure_endpoint": "https://your-resource.openai.azure.com/",
|
||||
"azure_deployment": "gpt-4-deployment"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Custom Prompts
|
||||
|
||||
### Default Prompt Behavior
|
||||
|
||||
The default prompt asks the LLM to score relevance on a 0-10 scale:
|
||||
|
||||
```
|
||||
Given a query and a memory, rate how relevant the memory is to answering the query.
|
||||
Score from 0 (completely irrelevant) to 10 (perfectly relevant).
|
||||
Only provide the numeric score.
|
||||
|
||||
Query: {query}
|
||||
Memory: {memory}
|
||||
Score:
|
||||
```
|
||||
|
||||
### Custom Prompt Examples
|
||||
|
||||
#### Domain-Specific Scoring
|
||||
|
||||
```python
|
||||
custom_prompt = """
|
||||
You are a medical information specialist. Rate how relevant each memory is for answering the medical query.
|
||||
Consider clinical accuracy, specificity, and practical applicability.
|
||||
Rate from 1-10 where:
|
||||
- 1-3: Irrelevant or potentially harmful
|
||||
- 4-6: Somewhat relevant but incomplete
|
||||
- 7-8: Relevant and helpful
|
||||
- 9-10: Highly relevant and clinically useful
|
||||
|
||||
Query: {query}
|
||||
Memory: {memory}
|
||||
Score:
|
||||
"""
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "openai",
|
||||
"config": {
|
||||
"model": "gpt-4",
|
||||
"api_key": "your-api-key"
|
||||
}
|
||||
},
|
||||
"custom_prompt": custom_prompt
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Contextual Relevance
|
||||
|
||||
```python
|
||||
contextual_prompt = """
|
||||
Rate how well this memory answers the specific question asked.
|
||||
Consider:
|
||||
- Direct relevance to the question
|
||||
- Completeness of information
|
||||
- Recency and accuracy
|
||||
- Practical usefulness
|
||||
|
||||
Rate 1-5:
|
||||
1 = Not relevant
|
||||
2 = Slightly relevant
|
||||
3 = Moderately relevant
|
||||
4 = Very relevant
|
||||
5 = Perfectly answers the question
|
||||
|
||||
Query: {query}
|
||||
Memory: {memory}
|
||||
Score:
|
||||
"""
|
||||
```
|
||||
|
||||
#### Conversational Context
|
||||
|
||||
```python
|
||||
conversation_prompt = """
|
||||
You are helping evaluate which memories are most useful for a conversational AI assistant.
|
||||
Rate how helpful this memory would be for generating a relevant response.
|
||||
|
||||
Consider:
|
||||
- Direct relevance to user's intent
|
||||
- Emotional appropriateness
|
||||
- Factual accuracy
|
||||
- Conversation flow
|
||||
|
||||
Rate 0-10:
|
||||
Query: {query}
|
||||
Memory: {memory}
|
||||
Score:
|
||||
"""
|
||||
```
|
||||
|
||||
## Usage Examples
|
||||
|
||||
### Basic Usage
|
||||
|
||||
```python
|
||||
from mem0 import Memory
|
||||
|
||||
m = Memory.from_config(config)
|
||||
|
||||
# Add memories
|
||||
m.add("I'm allergic to peanuts", user_id="alice")
|
||||
m.add("I love Italian food", user_id="alice")
|
||||
m.add("I'm vegetarian", user_id="alice")
|
||||
|
||||
# Search with LLM reranking
|
||||
results = m.search(
|
||||
"What foods should I avoid?",
|
||||
user_id="alice",
|
||||
rerank=True
|
||||
)
|
||||
|
||||
for result in results["results"]:
|
||||
print(f"Memory: {result['memory']}")
|
||||
print(f"LLM Score: {result['score']:.2f}")
|
||||
```
|
||||
|
||||
### Batch Processing with Error Handling
|
||||
|
||||
```python
|
||||
def safe_llm_rerank_search(query, user_id, max_retries=3):
|
||||
for attempt in range(max_retries):
|
||||
try:
|
||||
return m.search(query, user_id=user_id, rerank=True)
|
||||
except Exception as e:
|
||||
print(f"Attempt {attempt + 1} failed: {e}")
|
||||
if attempt == max_retries - 1:
|
||||
# Fall back to vector search
|
||||
return m.search(query, user_id=user_id, rerank=False)
|
||||
|
||||
# Use the safe function
|
||||
results = safe_llm_rerank_search("What are my preferences?", "alice")
|
||||
```
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
### Speed vs Quality Trade-offs
|
||||
|
||||
| Model Type | Speed | Quality | Cost | Best For |
|
||||
|------------|-------|---------|------|----------|
|
||||
| GPT-3.5 Turbo | Fast | Good | Low | High-volume applications |
|
||||
| GPT-4 | Medium | Excellent | Medium | Quality-critical applications |
|
||||
| Claude 3 Sonnet | Medium | Excellent | Medium | Balanced performance |
|
||||
| Ollama Local | Variable | Good | Free | Privacy-sensitive applications |
|
||||
|
||||
### Optimization Strategies
|
||||
|
||||
```python
|
||||
# Fast configuration for high-volume use
|
||||
fast_config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "openai",
|
||||
"config": {
|
||||
"model": "gpt-3.5-turbo",
|
||||
"api_key": "your-api-key"
|
||||
}
|
||||
},
|
||||
"top_k": 5, # Limit candidates
|
||||
"temperature": 0.0
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
# High-quality configuration
|
||||
quality_config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "openai",
|
||||
"config": {
|
||||
"model": "gpt-4",
|
||||
"api_key": "your-api-key"
|
||||
}
|
||||
},
|
||||
"top_k": 15,
|
||||
"temperature": 0.0
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Advanced Use Cases
|
||||
|
||||
### Multi-Step Reasoning
|
||||
|
||||
```python
|
||||
reasoning_prompt = """
|
||||
Evaluate this memory's relevance using multi-step reasoning:
|
||||
|
||||
1. What is the main intent of the query?
|
||||
2. What key information does the memory contain?
|
||||
3. How directly does the memory address the query?
|
||||
4. What additional context might be needed?
|
||||
|
||||
Based on this analysis, rate relevance 1-10:
|
||||
|
||||
Query: {query}
|
||||
Memory: {memory}
|
||||
|
||||
Analysis:
|
||||
Step 1 (Intent):
|
||||
Step 2 (Information):
|
||||
Step 3 (Directness):
|
||||
Step 4 (Context):
|
||||
Final Score:
|
||||
"""
|
||||
```
|
||||
|
||||
### Comparative Ranking
|
||||
|
||||
```python
|
||||
comparative_prompt = """
|
||||
You will see a query and multiple memories. Rank them in order of relevance.
|
||||
Consider which memories best answer the question and would be most helpful.
|
||||
|
||||
Query: {query}
|
||||
|
||||
Memories to rank:
|
||||
{memories}
|
||||
|
||||
Provide scores 1-10 for each memory, considering their relative usefulness.
|
||||
"""
|
||||
```
|
||||
|
||||
### Emotional Intelligence
|
||||
|
||||
```python
|
||||
emotional_prompt = """
|
||||
Consider both factual relevance and emotional appropriateness.
|
||||
Rate how suitable this memory is for responding to the user's query.
|
||||
|
||||
Factors to consider:
|
||||
- Factual accuracy and relevance
|
||||
- Emotional tone and sensitivity
|
||||
- User's likely emotional state
|
||||
- Appropriateness of response
|
||||
|
||||
Query: {query}
|
||||
Memory: {memory}
|
||||
Emotional Context: {context}
|
||||
Score (1-10):
|
||||
"""
|
||||
```
|
||||
|
||||
## Error Handling and Fallbacks
|
||||
|
||||
```python
|
||||
class RobustLLMReranker:
|
||||
def __init__(self, primary_config, fallback_config=None):
|
||||
self.primary = Memory.from_config(primary_config)
|
||||
self.fallback = Memory.from_config(fallback_config) if fallback_config else None
|
||||
|
||||
def search(self, query, user_id, max_retries=2):
|
||||
# Try primary LLM reranker
|
||||
for attempt in range(max_retries):
|
||||
try:
|
||||
return self.primary.search(query, user_id=user_id, rerank=True)
|
||||
except Exception as e:
|
||||
print(f"Primary reranker attempt {attempt + 1} failed: {e}")
|
||||
|
||||
# Try fallback reranker
|
||||
if self.fallback:
|
||||
try:
|
||||
return self.fallback.search(query, user_id=user_id, rerank=True)
|
||||
except Exception as e:
|
||||
print(f"Fallback reranker failed: {e}")
|
||||
|
||||
# Final fallback: vector search only
|
||||
return self.primary.search(query, user_id=user_id, rerank=False)
|
||||
|
||||
# Usage
|
||||
primary_config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {"llm": {"provider": "openai", "config": {"model": "gpt-4"}}}
|
||||
}
|
||||
}
|
||||
|
||||
fallback_config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {"llm": {"provider": "openai", "config": {"model": "gpt-3.5-turbo"}}}
|
||||
}
|
||||
}
|
||||
|
||||
reranker = RobustLLMReranker(primary_config, fallback_config)
|
||||
results = reranker.search("What are my preferences?", "alice")
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Use Specific Prompts**: Tailor prompts to your domain and use case
|
||||
2. **Set Temperature to 0**: Ensure consistent scoring across runs
|
||||
3. **Limit Top-K**: Don't rerank too many candidates to control costs
|
||||
4. **Implement Fallbacks**: Always have a backup plan for API failures
|
||||
5. **Monitor Costs**: Track API usage, especially with expensive models
|
||||
6. **Cache Results**: Consider caching reranking results for repeated queries
|
||||
7. **Test Prompts**: Experiment with different prompts to find what works best
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Common Issues
|
||||
|
||||
**Inconsistent Scores**
|
||||
- Set temperature to 0.0
|
||||
- Use more specific prompts
|
||||
- Consider using multiple calls and averaging
|
||||
|
||||
**API Rate Limits**
|
||||
- Implement exponential backoff
|
||||
- Use cheaper models for high-volume scenarios
|
||||
- Add retry logic with delays
|
||||
|
||||
**Poor Ranking Quality**
|
||||
- Refine your custom prompt
|
||||
- Try different LLM models
|
||||
- Add examples to your prompt
|
||||
|
||||
## Next Steps
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Custom Prompts Guide" icon="pencil" href="/components/rerankers/custom-prompts">
|
||||
Learn to craft effective reranking prompts
|
||||
</Card>
|
||||
<Card title="Performance Optimization" icon="bolt" href="/components/rerankers/optimization">
|
||||
Optimize LLM reranker performance
|
||||
</Card>
|
||||
</CardGroup>
|
||||
@@ -0,0 +1,162 @@
|
||||
---
|
||||
title: Sentence Transformer
|
||||
---
|
||||
|
||||
Sentence Transformer rerankers use cross-encoder models that are specifically designed for ranking tasks. These models can run locally and provide good reranking performance without external API calls.
|
||||
|
||||
## Usage
|
||||
|
||||
To use Sentence Transformer reranker with Mem0:
|
||||
|
||||
```python
|
||||
from mem0 import Memory
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"device": "cpu",
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
memory = Memory.from_config(config)
|
||||
|
||||
# Use memory as usual
|
||||
memory.add("I love playing basketball", user_id="alice")
|
||||
memory.add("I enjoy watching movies", user_id="alice")
|
||||
|
||||
# Search will now use Sentence Transformer reranking
|
||||
results = memory.search("What sports does Alice like?", user_id="alice")
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
| Parameter | Description | Default |
|
||||
|-----------|-------------|---------|
|
||||
| `model` | Sentence Transformer cross-encoder model | `cross-encoder/ms-marco-MiniLM-L-6-v2` |
|
||||
| `device` | Device to run on (`cpu`, `cuda`, `mps`) | `cpu` |
|
||||
| `top_n` | Number of results to return | `10` |
|
||||
|
||||
## Popular Models
|
||||
|
||||
### Lightweight Models
|
||||
- `cross-encoder/ms-marco-MiniLM-L-6-v2`: Fast and efficient
|
||||
- `cross-encoder/ms-marco-MiniLM-L-4-v2`: Even faster, slightly lower accuracy
|
||||
- `cross-encoder/ms-marco-MiniLM-L-2-v2`: Fastest, good for real-time applications
|
||||
|
||||
### High-Performance Models
|
||||
- `cross-encoder/ms-marco-electra-base`: Better accuracy, larger model
|
||||
- `ms-marco-MiniLM-L-12-v2`: Balanced performance and speed
|
||||
- `cross-encoder/qnli-electra-base`: Good for question-answering tasks
|
||||
|
||||
## Device Configuration
|
||||
|
||||
### CPU Usage
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"device": "cpu",
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### GPU Usage (CUDA)
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "cross-encoder/ms-marco-electra-base",
|
||||
"device": "cuda",
|
||||
"top_n": 15
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Apple Silicon (MPS)
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"device": "mps",
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Installation
|
||||
|
||||
The sentence-transformers library is required:
|
||||
|
||||
```bash
|
||||
pip install sentence-transformers
|
||||
```
|
||||
|
||||
For GPU support with CUDA:
|
||||
```bash
|
||||
pip install sentence-transformers torch
|
||||
```
|
||||
|
||||
## Performance Optimization
|
||||
|
||||
### Model Selection
|
||||
- Use MiniLM models for faster inference
|
||||
- Use larger models (electra-base) for better accuracy
|
||||
- Consider the trade-off between speed and quality
|
||||
|
||||
### Device Optimization
|
||||
- Use GPU (`cuda` or `mps`) for larger models
|
||||
- CPU is sufficient for MiniLM models
|
||||
- Batch processing improves GPU utilization
|
||||
|
||||
### Memory Considerations
|
||||
```python
|
||||
# For memory-constrained environments
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-2-v2", # Smallest model
|
||||
"device": "cpu",
|
||||
"top_n": 5 # Fewer results to process
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Custom Models
|
||||
|
||||
You can use any Sentence Transformer cross-encoder model:
|
||||
|
||||
```python
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "your-custom-model-name",
|
||||
"device": "cpu",
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Advantages
|
||||
|
||||
- **Local Processing**: No external API calls required
|
||||
- **Privacy**: Data stays on your infrastructure
|
||||
- **Cost Effective**: No per-request charges
|
||||
- **Fast**: Especially with GPU acceleration
|
||||
- **Customizable**: Can fine-tune on your specific data
|
||||
@@ -0,0 +1,312 @@
|
||||
---
|
||||
title: Performance Optimization
|
||||
icon: "bolt"
|
||||
iconType: "solid"
|
||||
---
|
||||
|
||||
Optimizing reranker performance is crucial for maintaining fast search response times while improving result quality. This guide covers best practices for different reranker types.
|
||||
|
||||
## General Optimization Principles
|
||||
|
||||
### Candidate Set Size
|
||||
The number of candidates sent to the reranker significantly impacts performance:
|
||||
|
||||
```python
|
||||
# Optimal candidate sizes for different rerankers
|
||||
config_map = {
|
||||
"cohere": {"initial_candidates": 100, "top_n": 10},
|
||||
"sentence_transformer": {"initial_candidates": 50, "top_n": 10},
|
||||
"huggingface": {"initial_candidates": 30, "top_n": 5},
|
||||
"llm_reranker": {"initial_candidates": 20, "top_n": 5}
|
||||
}
|
||||
```
|
||||
|
||||
### Batching Strategy
|
||||
Process multiple queries efficiently:
|
||||
|
||||
```python
|
||||
# Configure for batch processing
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"batch_size": 16, # Process multiple candidates at once
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Provider-Specific Optimizations
|
||||
|
||||
### Cohere Optimization
|
||||
|
||||
```python
|
||||
# Optimized Cohere configuration
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"model": "rerank-english-v3.0",
|
||||
"top_n": 10,
|
||||
"max_chunks_per_doc": 10, # Limit chunk processing
|
||||
"return_documents": False # Reduce response size
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Best Practices:**
|
||||
- Use v3.0 models for better speed/accuracy balance
|
||||
- Limit candidates to 100 or fewer
|
||||
- Cache API responses when possible
|
||||
- Monitor API rate limits
|
||||
|
||||
### Sentence Transformer Optimization
|
||||
|
||||
```python
|
||||
# Performance-optimized configuration
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"device": "cuda", # Use GPU when available
|
||||
"batch_size": 32,
|
||||
"top_n": 10,
|
||||
"max_length": 512 # Limit input length
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Device Optimization:**
|
||||
```python
|
||||
import torch
|
||||
|
||||
# Auto-detect best device
|
||||
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"device": device,
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Hugging Face Optimization
|
||||
|
||||
```python
|
||||
# Optimized for Hugging Face models
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "huggingface",
|
||||
"config": {
|
||||
"model": "BAAI/bge-reranker-base",
|
||||
"use_fp16": True, # Half precision for speed
|
||||
"max_length": 512,
|
||||
"batch_size": 8,
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### LLM Reranker Optimization
|
||||
|
||||
```python
|
||||
# Optimized LLM reranker configuration
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "llm_reranker",
|
||||
"config": {
|
||||
"llm": {
|
||||
"provider": "openai",
|
||||
"config": {
|
||||
"model": "gpt-3.5-turbo", # Faster than gpt-4
|
||||
"temperature": 0, # Deterministic results
|
||||
"max_tokens": 500 # Limit response length
|
||||
}
|
||||
},
|
||||
"batch_ranking": True, # Rank multiple at once
|
||||
"top_n": 5, # Fewer results for faster processing
|
||||
"timeout": 10 # Request timeout
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Performance Monitoring
|
||||
|
||||
### Latency Tracking
|
||||
```python
|
||||
import time
|
||||
from mem0 import Memory
|
||||
|
||||
def measure_reranker_performance(config, queries, user_id):
|
||||
memory = Memory.from_config(config)
|
||||
|
||||
latencies = []
|
||||
for query in queries:
|
||||
start_time = time.time()
|
||||
results = memory.search(query, user_id=user_id)
|
||||
latency = time.time() - start_time
|
||||
latencies.append(latency)
|
||||
|
||||
return {
|
||||
"avg_latency": sum(latencies) / len(latencies),
|
||||
"max_latency": max(latencies),
|
||||
"min_latency": min(latencies)
|
||||
}
|
||||
```
|
||||
|
||||
### Memory Usage Monitoring
|
||||
```python
|
||||
import psutil
|
||||
import os
|
||||
|
||||
def monitor_memory_usage():
|
||||
process = psutil.Process(os.getpid())
|
||||
return {
|
||||
"memory_mb": process.memory_info().rss / 1024 / 1024,
|
||||
"memory_percent": process.memory_percent()
|
||||
}
|
||||
```
|
||||
|
||||
## Caching Strategies
|
||||
|
||||
### Result Caching
|
||||
```python
|
||||
from functools import lru_cache
|
||||
import hashlib
|
||||
|
||||
class CachedReranker:
|
||||
def __init__(self, config):
|
||||
self.memory = Memory.from_config(config)
|
||||
self.cache_size = 1000
|
||||
|
||||
@lru_cache(maxsize=1000)
|
||||
def search_cached(self, query_hash, user_id):
|
||||
return self.memory.search(query, user_id=user_id)
|
||||
|
||||
def search(self, query, user_id):
|
||||
query_hash = hashlib.md5(f"{query}_{user_id}".encode()).hexdigest()
|
||||
return self.search_cached(query_hash, user_id)
|
||||
```
|
||||
|
||||
### Model Caching
|
||||
```python
|
||||
# Pre-load models to avoid initialization overhead
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"cache_folder": "/path/to/model/cache",
|
||||
"device": "cuda"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Parallel Processing
|
||||
|
||||
### Async Configuration
|
||||
```python
|
||||
import asyncio
|
||||
from mem0 import Memory
|
||||
|
||||
async def parallel_search(config, queries, user_id):
|
||||
memory = Memory.from_config(config)
|
||||
|
||||
# Process multiple queries concurrently
|
||||
tasks = [
|
||||
memory.search_async(query, user_id=user_id)
|
||||
for query in queries
|
||||
]
|
||||
|
||||
results = await asyncio.gather(*tasks)
|
||||
return results
|
||||
```
|
||||
|
||||
## Hardware Optimization
|
||||
|
||||
### GPU Configuration
|
||||
```python
|
||||
# Optimize for GPU usage
|
||||
import torch
|
||||
|
||||
if torch.cuda.is_available():
|
||||
torch.cuda.set_per_process_memory_fraction(0.8) # Reserve GPU memory
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"device": "cuda",
|
||||
"model": "cross-encoder/ms-marco-electra-base",
|
||||
"batch_size": 64, # Larger batch for GPU
|
||||
"fp16": True # Half precision
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### CPU Optimization
|
||||
```python
|
||||
import torch
|
||||
|
||||
# Optimize CPU threading
|
||||
torch.set_num_threads(4) # Adjust based on your CPU
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "sentence_transformer",
|
||||
"config": {
|
||||
"device": "cpu",
|
||||
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
|
||||
"num_workers": 4 # Parallel processing
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Benchmarking Different Configurations
|
||||
|
||||
```python
|
||||
def benchmark_rerankers():
|
||||
configs = [
|
||||
{"provider": "cohere", "model": "rerank-english-v3.0"},
|
||||
{"provider": "sentence_transformer", "model": "cross-encoder/ms-marco-MiniLM-L-6-v2"},
|
||||
{"provider": "huggingface", "model": "BAAI/bge-reranker-base"}
|
||||
]
|
||||
|
||||
test_queries = ["sample query 1", "sample query 2", "sample query 3"]
|
||||
|
||||
results = {}
|
||||
for config in configs:
|
||||
provider = config["provider"]
|
||||
performance = measure_reranker_performance(
|
||||
{"reranker": {"provider": provider, "config": config}},
|
||||
test_queries,
|
||||
"test_user"
|
||||
)
|
||||
results[provider] = performance
|
||||
|
||||
return results
|
||||
```
|
||||
|
||||
## Production Best Practices
|
||||
|
||||
1. **Model Selection**: Choose the right balance of speed vs. accuracy
|
||||
2. **Resource Allocation**: Monitor CPU/GPU usage and memory consumption
|
||||
3. **Error Handling**: Implement fallbacks for reranker failures
|
||||
4. **Load Balancing**: Distribute reranking load across multiple instances
|
||||
5. **Monitoring**: Track latency, throughput, and error rates
|
||||
6. **Caching**: Cache frequent queries and model predictions
|
||||
7. **Batch Processing**: Group similar queries for efficient processing
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
title: Overview
|
||||
icon: "info"
|
||||
iconType: "solid"
|
||||
---
|
||||
|
||||
Rerankers enhance the quality of search results by re-ordering the initial retrieval results using more sophisticated scoring mechanisms. They act as a secondary ranking layer that can significantly improve the relevance of retrieved memories.
|
||||
|
||||
## How Rerankers Work
|
||||
|
||||
1. **Initial Retrieval**: Vector search returns candidate memories based on semantic similarity
|
||||
2. **Reranking**: The reranker evaluates and re-scores these candidates using more complex criteria
|
||||
3. **Final Results**: Returns the top-k memories with improved relevance ordering
|
||||
|
||||
## Benefits
|
||||
|
||||
- **Improved Precision**: Better ranking of relevant memories
|
||||
- **Context Awareness**: More sophisticated understanding of query-memory relationships
|
||||
- **Performance**: Can improve results without changing the underlying vector store
|
||||
|
||||
## Supported Rerankers
|
||||
|
||||
Mem0 supports several reranker models:
|
||||
|
||||
<CardGroup cols={2}>
|
||||
<Card title="Cohere" href="/components/rerankers/models/cohere" />
|
||||
<Card title="Sentence Transformer" href="/components/rerankers/models/sentence_transformer" />
|
||||
<Card title="Hugging Face" href="/components/rerankers/models/huggingface" />
|
||||
<Card title="LLM Reranker" href="/components/rerankers/models/llm_reranker" />
|
||||
</CardGroup>
|
||||
|
||||
## Usage
|
||||
|
||||
Rerankers are configured as part of the memory configuration:
|
||||
|
||||
```python
|
||||
from mem0 import Memory
|
||||
|
||||
config = {
|
||||
"reranker": {
|
||||
"provider": "cohere",
|
||||
"config": {
|
||||
"api_key": "your-api-key",
|
||||
"top_n": 10
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
memory = Memory.from_config(config)
|
||||
```
|
||||
|
||||
For detailed configuration options, see the [Config](./config) page.
|
||||
Reference in New Issue
Block a user