Mem0 1.0.0 (#3545)

This commit is contained in:
Parshva Daftari
2025-10-16 15:50:20 +05:30
committed by GitHub
parent 8f5151c344
commit 394203d1b5
77 changed files with 3445 additions and 990 deletions
+108 -77
View File
@@ -1,116 +1,147 @@
---
title: Cohere
description: 'Reranking with Cohere'
icon: "building"
iconType: "solid"
---
Cohere provides state-of-the-art reranking models that can significantly improve the relevance of search results. Cohere's rerankers are optimized for various languages and use cases.
Cohere provides enterprise-grade reranking models with excellent multilingual support and production-ready performance.
## Usage
## Models
To use Cohere's reranker with Mem0:
Cohere offers several reranking models:
```python
import os
- **`rerank-english-v3.0`**: Latest English reranker with best performance
- **`rerank-multilingual-v3.0`**: Multilingual support for global applications
- **`rerank-english-v2.0`**: Previous generation English reranker
## Installation
```bash
pip install cohere
```
## Configuration
```python Python
from mem0 import Memory
os.environ["COHERE_API_KEY"] = "your-cohere-api-key"
config = {
"vector_store": {
"provider": "chroma",
"config": {
"collection_name": "my_memories",
"path": "./chroma_db"
}
},
"llm": {
"provider": "openai",
"config": {
"model": "gpt-4o-mini"
}
},
"reranker": {
"provider": "cohere",
"config": {
"api_key": "your-cohere-api-key", # Can also use environment variable
"model": "rerank-english-v3.0",
"top_n": 10
"api_key": "your-cohere-api-key", # or set COHERE_API_KEY
"top_k": 5,
"return_documents": False,
"max_chunks_per_doc": None
}
}
}
memory = Memory.from_config(config)
```
## Environment Variables
Set your API key as an environment variable:
```bash
export COHERE_API_KEY="your-api-key"
```
## Usage Example
```python Python
import os
from mem0 import Memory
# Set API key
os.environ["COHERE_API_KEY"] = "your-api-key"
# Initialize memory with Cohere reranker
config = {
"vector_store": {"provider": "chroma"},
"llm": {"provider": "openai", "config": {"model": "gpt-4o-mini"}},
"rerank": {
"provider": "cohere",
"config": {
"model": "rerank-english-v3.0",
"top_k": 3
}
}
}
memory = Memory.from_config(config)
# Use memory as usual
memory.add("I love playing basketball", user_id="alice")
memory.add("I enjoy watching movies", user_id="alice")
# Add memories
messages = [
{"role": "user", "content": "I work as a data scientist at Microsoft"},
{"role": "user", "content": "I specialize in machine learning and NLP"},
{"role": "user", "content": "I enjoy playing tennis on weekends"}
]
# Search will now use Cohere reranking
results = memory.search("What sports does Alice like?", user_id="alice")
memory.add(messages, user_id="bob")
# Search with reranking
results = memory.search("What is the user's profession?", user_id="bob")
for result in results['results']:
print(f"Memory: {result['memory']}")
print(f"Vector Score: {result['score']:.3f}")
print(f"Rerank Score: {result['rerank_score']:.3f}")
print()
```
## Configuration
## Multilingual Support
| Parameter | Description | Default |
|-----------|-------------|---------|
| `api_key` | Cohere API key | Required |
| `model` | Cohere rerank model | `rerank-english-v3.0` |
| `top_n` | Number of results to return | `10` |
For multilingual applications, use the multilingual model:
## Available Models
- `rerank-english-v3.0`: Latest English reranking model
- `rerank-multilingual-v3.0`: Multilingual reranking model
- `rerank-english-v2.0`: Previous English model
- `rerank-multilingual-v2.0`: Previous multilingual model
## Example with Different Models
### English Reranker
```python
```python Python
config = {
"reranker": {
"rerank": {
"provider": "cohere",
"config": {
"api_key": "your-cohere-api-key",
"model": "rerank-english-v3.0",
"top_n": 5
}
}
}
```
### Multilingual Reranker
```python
config = {
"reranker": {
"provider": "cohere",
"config": {
"api_key": "your-cohere-api-key",
"model": "rerank-multilingual-v3.0",
"top_n": 8
"top_k": 5
}
}
}
```
## Environment Variables
## Configuration Parameters
You can set your Cohere API key as an environment variable:
| Parameter | Description | Type | Default |
|-----------|-------------|------|---------|
| `model` | Cohere rerank model to use | `str` | `"rerank-english-v3.0"` |
| `api_key` | Cohere API key | `str` | `None` |
| `top_k` | Maximum documents to return | `int` | `None` |
| `return_documents` | Whether to return document texts | `bool` | `False` |
| `max_chunks_per_doc` | Maximum chunks per document | `int` | `None` |
```bash
export COHERE_API_KEY="your-cohere-api-key"
```
## Features
Then use the config without specifying the API key:
- **High Quality**: Enterprise-grade relevance scoring
- **Multilingual**: Support for 100+ languages
- **Scalable**: Production-ready with high throughput
- **Reliable**: SLA-backed service with 99.9% uptime
```python
config = {
"reranker": {
"provider": "cohere",
"config": {
"model": "rerank-english-v3.0",
"top_n": 10
}
}
}
```
## Best Practices
## Getting Your API Key
1. Sign up at [Cohere](https://cohere.ai/)
2. Navigate to the API keys section in your dashboard
3. Generate a new API key
4. Use this key in your configuration
## Performance Considerations
- Cohere rerankers work best with 10-100 candidate documents
- Higher `top_n` values provide more comprehensive reranking but may increase latency
- The v3.0 models generally provide better performance than v2.0 models
1. **Model Selection**: Use `rerank-english-v3.0` for English, `rerank-multilingual-v3.0` for other languages
2. **Batch Processing**: Process multiple queries efficiently
3. **Error Handling**: Implement retry logic for production systems
4. **Monitoring**: Track reranking performance and costs
+224
View File
@@ -0,0 +1,224 @@
---
title: LLM as Reranker
description: 'Flexible reranking using LLMs'
icon: "robot"
iconType: "solid"
---
LLM-based reranker provides maximum flexibility by using any Large Language Model to score document relevance. This approach allows for custom prompts and domain-specific scoring logic.
## Supported LLM Providers
Any LLM provider supported by Mem0 can be used for reranking:
- **OpenAI**: GPT-4, GPT-3.5-turbo, etc.
- **Anthropic**: Claude models
- **Together**: Open-source models
- **Groq**: Fast inference
- **Ollama**: Local models
- And more...
## Configuration
```python Python
from mem0 import Memory
config = {
"vector_store": {
"provider": "chroma",
"config": {
"collection_name": "my_memories",
"path": "./chroma_db"
}
},
"llm": {
"provider": "openai",
"config": {
"model": "gpt-4o-mini"
}
},
"rerank": {
"provider": "llm",
"config": {
"model": "gpt-4o-mini",
"provider": "openai",
"api_key": "your-openai-api-key", # or set OPENAI_API_KEY
"top_k": 5,
"temperature": 0.0
}
}
}
memory = Memory.from_config(config)
```
## Custom Scoring Prompt
You can provide a custom prompt for relevance scoring:
```python Python
custom_prompt = """You are a relevance scoring assistant. Rate how well this document answers the query.
Query: "{query}"
Document: "{document}"
Score from 0.0 to 1.0 where:
- 1.0: Perfect match, directly answers the query
- 0.8-0.9: Highly relevant, good match
- 0.6-0.7: Moderately relevant, partial match
- 0.4-0.5: Slightly relevant, limited useful information
- 0.0-0.3: Not relevant or no useful information
Provide only a single numerical score between 0.0 and 1.0."""
config["rerank"]["config"]["scoring_prompt"] = custom_prompt
```
## Usage Example
```python Python
import os
from mem0 import Memory
# Set API key
os.environ["OPENAI_API_KEY"] = "your-api-key"
# Initialize memory with LLM reranker
config = {
"vector_store": {"provider": "chroma"},
"llm": {"provider": "openai", "config": {"model": "gpt-4o-mini"}},
"rerank": {
"provider": "llm",
"config": {
"model": "gpt-4o-mini",
"provider": "openai",
"temperature": 0.0
}
}
}
memory = Memory.from_config(config)
# Add memories
messages = [
{"role": "user", "content": "I'm learning Python programming"},
{"role": "user", "content": "I find object-oriented programming challenging"},
{"role": "user", "content": "I love hiking in national parks"}
]
memory.add(messages, user_id="david")
# Search with LLM reranking
results = memory.search("What programming topics is the user studying?", user_id="david")
for result in results['results']:
print(f"Memory: {result['memory']}")
print(f"Vector Score: {result['score']:.3f}")
print(f"Rerank Score: {result['rerank_score']:.3f}")
print()
```
```text Output
Memory: I'm learning Python programming
Vector Score: 0.856
Rerank Score: 0.920
Memory: I find object-oriented programming challenging
Vector Score: 0.782
Rerank Score: 0.850
```
## Domain-Specific Scoring
Create specialized scoring for your domain:
```python Python
medical_prompt = """You are a medical relevance expert. Score how relevant this medical record is to the clinical query.
Clinical Query: "{query}"
Medical Record: "{document}"
Consider:
- Clinical relevance and accuracy
- Patient safety implications
- Diagnostic value
- Treatment relevance
Score from 0.0 to 1.0. Provide only the numerical score."""
config = {
"rerank": {
"provider": "llm",
"config": {
"model": "gpt-4o-mini",
"provider": "openai",
"scoring_prompt": medical_prompt,
"temperature": 0.0
}
}
}
```
## Multiple LLM Providers
Use different LLM providers for reranking:
```python Python
# Using Anthropic Claude
anthropic_config = {
"rerank": {
"provider": "llm",
"config": {
"model": "claude-3-haiku-20240307",
"provider": "anthropic",
"temperature": 0.0
}
}
}
# Using local Ollama model
ollama_config = {
"rerank": {
"provider": "llm",
"config": {
"model": "llama2:7b",
"provider": "ollama",
"temperature": 0.0
}
}
}
```
## Configuration Parameters
| Parameter | Description | Type | Default |
|-----------|-------------|------|---------|
| `model` | LLM model to use for scoring | `str` | `"gpt-4o-mini"` |
| `provider` | LLM provider name | `str` | `"openai"` |
| `api_key` | API key for the LLM provider | `str` | `None` |
| `top_k` | Maximum documents to return | `int` | `None` |
| `temperature` | Temperature for LLM generation | `float` | `0.0` |
| `max_tokens` | Maximum tokens for LLM response | `int` | `100` |
| `scoring_prompt` | Custom prompt template | `str` | Default prompt |
## Advantages
- **Maximum Flexibility**: Custom prompts for any use case
- **Domain Expertise**: Leverage LLM knowledge for specialized domains
- **Interpretability**: Understand scoring through prompt engineering
- **Multi-criteria**: Score based on multiple relevance factors
## Considerations
- **Latency**: Higher latency than specialized rerankers
- **Cost**: LLM API costs per reranking operation
- **Consistency**: May have slight variations in scoring
- **Prompt Engineering**: Requires careful prompt design
## Best Practices
1. **Temperature**: Use 0.0 for consistent scoring
2. **Prompt Design**: Be specific about scoring criteria
3. **Token Efficiency**: Keep prompts concise to reduce costs
4. **Caching**: Cache results for repeated queries when possible
5. **Fallback**: Handle API errors gracefully
@@ -1,162 +1,161 @@
---
title: Sentence Transformer
description: 'Local reranking with HuggingFace cross-encoder models'
icon: "server"
iconType: "solid"
---
Sentence Transformer rerankers use cross-encoder models that are specifically designed for ranking tasks. These models can run locally and provide good reranking performance without external API calls.
Sentence Transformer reranker provides local reranking using HuggingFace cross-encoder models, perfect for privacy-focused deployments where you want to keep data on-premises.
## Usage
## Models
To use Sentence Transformer reranker with Mem0:
Any HuggingFace cross-encoder model can be used. Popular choices include:
```python
- **`cross-encoder/ms-marco-MiniLM-L-6-v2`**: Default, good balance of speed and accuracy
- **`cross-encoder/ms-marco-TinyBERT-L-2-v2`**: Fastest, smaller model size
- **`cross-encoder/ms-marco-electra-base`**: Higher accuracy, larger model
- **`cross-encoder/stsb-distilroberta-base`**: Good for semantic similarity tasks
## Installation
```bash
pip install sentence-transformers
```
## Configuration
```python Python
from mem0 import Memory
config = {
"reranker": {
"vector_store": {
"provider": "chroma",
"config": {
"collection_name": "my_memories",
"path": "./chroma_db"
}
},
"llm": {
"provider": "openai",
"config": {
"model": "gpt-4o-mini"
}
},
"rerank": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
"device": "cpu",
"top_n": 10
"device": "cpu", # or "cuda" for GPU
"batch_size": 32,
"show_progress_bar": False,
"top_k": 5
}
}
}
memory = Memory.from_config(config)
```
## GPU Acceleration
For better performance, use GPU acceleration:
```python Python
config = {
"rerank": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
"device": "cuda", # Use GPU
"batch_size": 64 # high batch size for high memory GPUs
}
}
}
```
## Usage Example
```python Python
from mem0 import Memory
# Initialize memory with local reranker
config = {
"vector_store": {"provider": "chroma"},
"llm": {"provider": "openai", "config": {"model": "gpt-4o-mini"}},
"rerank": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
"device": "cpu"
}
}
}
memory = Memory.from_config(config)
# Use memory as usual
memory.add("I love playing basketball", user_id="alice")
memory.add("I enjoy watching movies", user_id="alice")
# Add memories
messages = [
{"role": "user", "content": "I love reading science fiction novels"},
{"role": "user", "content": "My favorite author is Isaac Asimov"},
{"role": "user", "content": "I also enjoy watching sci-fi movies"}
]
# Search will now use Sentence Transformer reranking
results = memory.search("What sports does Alice like?", user_id="alice")
```
memory.add(messages, user_id="charlie")
## Configuration
# Search with local reranking
results = memory.search("What books does the user like?", user_id="charlie")
| Parameter | Description | Default |
|-----------|-------------|---------|
| `model` | Sentence Transformer cross-encoder model | `cross-encoder/ms-marco-MiniLM-L-6-v2` |
| `device` | Device to run on (`cpu`, `cuda`, `mps`) | `cpu` |
| `top_n` | Number of results to return | `10` |
## Popular Models
### Lightweight Models
- `cross-encoder/ms-marco-MiniLM-L-6-v2`: Fast and efficient
- `cross-encoder/ms-marco-MiniLM-L-4-v2`: Even faster, slightly lower accuracy
- `cross-encoder/ms-marco-MiniLM-L-2-v2`: Fastest, good for real-time applications
### High-Performance Models
- `cross-encoder/ms-marco-electra-base`: Better accuracy, larger model
- `ms-marco-MiniLM-L-12-v2`: Balanced performance and speed
- `cross-encoder/qnli-electra-base`: Good for question-answering tasks
## Device Configuration
### CPU Usage
```python
config = {
"reranker": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
"device": "cpu",
"top_n": 10
}
}
}
```
### GPU Usage (CUDA)
```python
config = {
"reranker": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-electra-base",
"device": "cuda",
"top_n": 15
}
}
}
```
### Apple Silicon (MPS)
```python
config = {
"reranker": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
"device": "mps",
"top_n": 10
}
}
}
```
## Installation
The sentence-transformers library is required:
```bash
pip install sentence-transformers
```
For GPU support with CUDA:
```bash
pip install sentence-transformers torch
```
## Performance Optimization
### Model Selection
- Use MiniLM models for faster inference
- Use larger models (electra-base) for better accuracy
- Consider the trade-off between speed and quality
### Device Optimization
- Use GPU (`cuda` or `mps`) for larger models
- CPU is sufficient for MiniLM models
- Batch processing improves GPU utilization
### Memory Considerations
```python
# For memory-constrained environments
config = {
"reranker": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-MiniLM-L-2-v2", # Smallest model
"device": "cpu",
"top_n": 5 # Fewer results to process
}
}
}
for result in results['results']:
print(f"Memory: {result['memory']}")
print(f"Vector Score: {result['score']:.3f}")
print(f"Rerank Score: {result['rerank_score']:.3f}")
print()
```
## Custom Models
You can use any Sentence Transformer cross-encoder model:
You can use any HuggingFace cross-encoder model:
```python
```python Python
# Using a different model
config = {
"reranker": {
"provider": "sentence_transformer",
"rerank": {
"provider": "sentence_transformer",
"config": {
"model": "your-custom-model-name",
"device": "cpu",
"top_n": 10
"model": "cross-encoder/stsb-distilroberta-base",
"device": "cpu"
}
}
}
```
## Configuration Parameters
| Parameter | Description | Type | Default |
|-----------|-------------|------|---------|
| `model` | HuggingFace cross-encoder model name | `str` | `"cross-encoder/ms-marco-MiniLM-L-6-v2"` |
| `device` | Device to run model on (`cpu`, `cuda`, etc.) | `str` | `None` |
| `batch_size` | Batch size for processing documents | `int` | `32` |
| `show_progress_bar` | Show progress bar during processing | `bool` | `False` |
| `top_k` | Maximum documents to return | `int` | `None` |
## Advantages
- **Local Processing**: No external API calls required
- **Privacy**: Data stays on your infrastructure
- **Cost Effective**: No per-request charges
- **Fast**: Especially with GPU acceleration
- **Customizable**: Can fine-tune on your specific data
- **Privacy**: Complete local processing, no external API calls
- **Cost**: No per-token charges after initial model download
- **Customization**: Use any HuggingFace cross-encoder model
- **Offline**: Works without internet connection after model download
## Performance Considerations
- **First Run**: Model download may take time initially
- **Memory Usage**: Models require GPU/CPU memory
- **Batch Size**: Optimize batch size based on available memory
- **Device**: GPU acceleration significantly improves speed
## Best Practices
1. **Model Selection**: Choose model based on accuracy vs speed requirements
2. **Device Management**: Use GPU when available for better performance
3. **Batch Processing**: Process multiple documents together for efficiency
4. **Memory Monitoring**: Monitor system memory usage with larger models
@@ -0,0 +1,119 @@
---
title: Zero Entropy
description: 'Neural reranking with Zero Entropy'
icon: "sparkles"
iconType: "solid"
---
[Zero Entropy](https://www.zeroentropy.dev) provides neural reranking models that significantly improve search relevance with fast performance.
## Models
Zero Entropy offers two reranking models:
- **`zerank-1`**: Flagship state-of-the-art reranker (non-commercial license)
- **`zerank-1-small`**: Open-source model (Apache 2.0 license)
## Installation
```bash
pip install zeroentropy
```
## Configuration
```python Python
from mem0 import Memory
config = {
"vector_store": {
"provider": "chroma",
"config": {
"collection_name": "my_memories",
"path": "./chroma_db"
}
},
"llm": {
"provider": "openai",
"config": {
"model": "gpt-4o-mini"
}
},
"rerank": {
"provider": "zero_entropy",
"config": {
"model": "zerank-1", # or "zerank-1-small"
"api_key": "your-zero-entropy-api-key", # or set ZERO_ENTROPY_API_KEY
"top_k": 5
}
}
}
memory = Memory.from_config(config)
```
## Environment Variables
Set your API key as an environment variable:
```bash
export ZERO_ENTROPY_API_KEY="your-api-key"
```
## Usage Example
```python Python
import os
from mem0 import Memory
# Set API key
os.environ["ZERO_ENTROPY_API_KEY"] = "your-api-key"
# Initialize memory with Zero Entropy reranker
config = {
"vector_store": {"provider": "chroma"},
"llm": {"provider": "openai", "config": {"model": "gpt-4o-mini"}},
"rerank": {"provider": "zero_entropy", "config": {"model": "zerank-1"}}
}
memory = Memory.from_config(config)
# Add memories
messages = [
{"role": "user", "content": "I love Italian pasta, especially carbonara"},
{"role": "user", "content": "Japanese sushi is also amazing"},
{"role": "user", "content": "I enjoy cooking Mediterranean dishes"}
]
memory.add(messages, user_id="alice")
# Search with reranking
results = memory.search("What Italian food does the user like?", user_id="alice")
for result in results['results']:
print(f"Memory: {result['memory']}")
print(f"Vector Score: {result['score']:.3f}")
print(f"Rerank Score: {result['rerank_score']:.3f}")
print()
```
## Configuration Parameters
| Parameter | Description | Type | Default |
|-----------|-------------|------|---------|
| `model` | Model to use: `"zerank-1"` or `"zerank-1-small"` | `str` | `"zerank-1"` |
| `api_key` | Zero Entropy API key | `str` | `None` |
| `top_k` | Maximum documents to return after reranking | `int` | `None` |
## Performance
- **Fast**: Optimized neural architecture for low latency
- **Accurate**: State-of-the-art relevance scoring
- **Cost-effective**: ~$0.025/1M tokens processed
## Best Practices
1. **Model Selection**: Use `zerank-1` for best quality, `zerank-1-small` for faster processing
2. **Batch Size**: Process multiple queries together when possible
3. **Top-k Limiting**: Set reasonable `top_k` values (5-20) for best performance
4. **API Key Management**: Use environment variables for secure key storage