Compare commits

..

8 Commits

Author SHA1 Message Date
Deven Patel c0b5e93967 [Feature] add google ai embedder (#1019)
Co-authored-by: Deven Patel <deven298@yahoo.com>
2023-12-18 13:58:01 +05:30
Sidharth Mohanty 6983ebba49 Chainlit + Embedchain Integration (Example) (#1020) 2023-12-18 13:20:56 +05:30
Sidharth Mohanty 9943d1e015 [chore] Remove deployment_name for openai embedder and update docs (#1022) 2023-12-18 08:50:49 +05:30
Sidharth Mohanty b348251484 [Docs] Don't need open ai key for gpt4all model (#1018) 2023-12-16 08:44:47 +05:30
xuxiang e719b5bac3 [Bug-fix] Fix error caused by executing Repo.clone_from twice (#1015)
Co-authored-by: xuxiang <xuxiang@aliyun.com>
2023-12-16 08:33:07 +05:30
Sidharth Mohanty 54f43215cd Docs for directory data loader usage (#1016) 2023-12-16 07:59:05 +05:30
Deven Patel b246d9823e [Docs] Documentation updates (#1014) 2023-12-15 17:02:50 +05:30
Deven Patel 65c8dd445b Update README to show package downloads (#1013) 2023-12-15 11:05:13 +05:30
31 changed files with 372 additions and 62 deletions
+3
View File
@@ -28,6 +28,9 @@
<a href="https://codecov.io/gh/embedchain/embedchain">
<img src="https://codecov.io/gh/embedchain/embedchain/graph/badge.svg?token=EMRRHZXW1Q" alt="codecov">
</a>
<a href="https://pepy.tech/project/embedchain">
<img src="https://static.pepy.tech/badge/embedchain" alt="Downloads">
</a>
</p>
<hr />
+5
View File
@@ -6,3 +6,8 @@ llm:
temperature: 0.9
top_p: 1.0
stream: false
embedder:
provider: google
config:
model: models/embedding-001
@@ -1,6 +1,9 @@
<p>If you can't find the specific data source, please feel free to request through one of the following channels and help us prioritize.</p>
<CardGroup cols={2}>
<Card title="Google Form" icon="file" href="https://forms.gle/NDRCKsRpUHsz2Wcm8" color="#7387d0">
Fill out this form
</Card>
<Card title="Slack" icon="slack" href="https://join.slack.com/t/embedchain/shared_invite/zt-22uwz3c46-Zg7cIh5rOBteT_xe1jwLDw" color="#4A154B">
Let us know on our slack community
</Card>
+19
View File
@@ -0,0 +1,19 @@
---
title: 🗑 delete
---
`delete_chat_history()` method allows you to delete all previous messages in a chat history.
## Usage
```python
from embedchain import Pipeline as App
app = App()
app.add("https://www.forbes.com/profile/elon-musk")
app.chat("What is the net worth of Elon Musk?")
app.delete_chat_history()
```
+1 -1
View File
@@ -14,4 +14,4 @@ app.add("https://www.forbes.com/profile/elon-musk")
# Reset the app
app.reset()
```
```
@@ -0,0 +1,41 @@
---
title: '📁 Directory'
---
To use an entire directory as data source, just add `data_type` as `directory` and pass in the path of the local directory.
### Without customization
```python
import os
from embedchain import Pipeline as App
os.environ["OPENAI_API_KEY"] = "sk-xxx"
app = App()
app.add("./elon-musk", data_type="directory")
response = app.query("list all files")
print(response)
# Answer: Files are elon-musk-1.txt, elon-musk-2.pdf.
```
### Customization
```python
import os
from embedchain import Pipeline as App
from embedchain.loaders.directory_loader import DirectoryLoader
os.environ["OPENAI_API_KEY"] = "sk-xxx"
lconfig = {
"recursive": True,
"extensions": [".txt"]
}
loader = DirectoryLoader(config=lconfig)
app = App()
app.add("./elon-musk", loader=loader)
response = app.query("what are all the files related to?")
print(response)
# Answer: The files are related to Elon Musk.
```
@@ -29,6 +29,7 @@ Embedchain comes with built-in support for various data sources. We handle the c
<Card title="⚙️ Custom" href="/components/data-sources/custom"></Card>
<Card title="📝 Substack" href="/components/data-sources/substack"></Card>
<Card title="🐝 Beehiiv" href="/components/data-sources/beehiiv"></Card>
<Card title="📁 Directory" href="/components/data-sources/directory"></Card>
</CardGroup>
<br/ >
+29
View File
@@ -8,6 +8,7 @@ Embedchain supports several embedding models from the following providers:
<CardGroup cols={4}>
<Card title="OpenAI" href="#openai"></Card>
<Card title="GoogleAI" href="#google-ai"></Card>
<Card title="Azure OpenAI" href="#azure-openai"></Card>
<Card title="GPT4All" href="#gpt4all"></Card>
<Card title="Hugging Face" href="#hugging-face"></Card>
@@ -44,6 +45,34 @@ embedder:
</CodeGroup>
## Google AI
To use Google AI embedding function, you have to set the `GOOGLE_API_KEY` environment variable. You can obtain the Google API key from the [Google Maker Suite](https://makersuite.google.com/app/apikey)
<CodeGroup>
```python main.py
import os
from embedchain import Pipeline as App
os.environ["GOOGLE_API_KEY"] = "xxx"
app = App.from_config(config_path="config.yaml")
```
```yaml config.yaml
embedder:
provider: google
config:
model: 'models/embedding-001'
task_type: "retrieval_document"
title: "Embeddings for Embedchain"
```
</CodeGroup>
<br/>
<Note>
For more details regarding the Google AI embedding model, please refer to the [Google AI documentation](https://ai.google.dev/tutorials/python_quickstart#use_embeddings).
</Note>
## Azure OpenAI
To use Azure OpenAI embedding model, you have to set some of the azure openai related environment variables as given in the code block below:
+7 -1
View File
@@ -72,7 +72,6 @@ To use Google AI model, you have to set the `GOOGLE_API_KEY` environment variabl
import os
from embedchain import Pipeline as App
os.environ["OPENAI_API_KEY"] = "sk-xxxx"
os.environ["GOOGLE_API_KEY"] = "xxx"
app = App.from_config(config_path="config.yaml")
@@ -96,6 +95,13 @@ llm:
temperature: 0.5
top_p: 1
stream: false
embedder:
provider: google
config:
model: 'models/embedding-001'
task_type: "retrieval_document"
title: "Embeddings for Embedchain"
```
</CodeGroup>
@@ -1,5 +1,5 @@
---
title: 🔎 Examples
title: Notebooks & Replits
---
# Explore awesome apps
+31 -4
View File
@@ -4,7 +4,7 @@ description: 'Collections of all the frequently asked questions'
---
<AccordionGroup>
<Accordion title="Does Embedchain support OpenAI's Assistant APIs?">
Yes, it does. Please refer to the [OpenAI Assistant docs page](/get-started/openai-assistant).
Yes, it does. Please refer to the [OpenAI Assistant docs page](/examples/openai-assistant).
</Accordion>
<Accordion title="How to use MistralAI language model?">
Use the model provided on huggingface: `mistralai/Mistral-7B-v0.1`
@@ -90,11 +90,8 @@ llm:
<CodeGroup>
```python main.py
import os
from embedchain import Pipeline as App
os.environ['OPENAI_API_KEY'] = 'xxx'
# load llm configuration from opensource.yaml file
app = App.from_config(config_path="opensource.yaml")
```
@@ -116,6 +113,36 @@ embedder:
```
</CodeGroup>
</Accordion>
<Accordion title="How to stream response while using OpenAI model in Embedchain?">
You can achieve this by setting `stream` to `true` in the config file.
<CodeGroup>
```yaml openai.yaml
llm:
provider: openai
config:
model: 'gpt-3.5-turbo'
temperature: 0.5
max_tokens: 1000
top_p: 1
stream: true
```
```python main.py
import os
from embedchain import Pipeline as App
os.environ['OPENAI_API_KEY'] = 'sk-xxx'
app = App.from_config(config_path="openai.yaml")
app.add("https://www.forbes.com/profile/elon-musk")
response = app.query("What is the net worth of Elon Musk?")
# response will be streamed in stdout as it is generated.
```
</CodeGroup>
</Accordion>
</AccordionGroup>
+68
View File
@@ -0,0 +1,68 @@
---
title: '⛓️ Chainlit'
description: 'Integrate with Chainlit to create LLM chat apps'
---
In this example, we will learn how to use Chainlit and Embedchain together
## Setup
First, install the required packages:
```bash
pip install embedchain chainlit
```
## Create a Chainlit app
Create a new file called `app.py` and add the following code:
```python
import chainlit as cl
from embedchain import Pipeline as App
import os
os.environ["OPENAI_API_KEY"] = "sk-xxx"
@cl.on_chat_start
async def on_chat_start():
app = App.from_config(config={
'app': {
'config': {
'name': 'chainlit-app'
}
},
'llm': {
'config': {
'stream': True,
}
}
})
# import your data here
app.add("https://www.forbes.com/profile/elon-musk/")
app.collect_metrics = False
cl.user_session.set("app", app)
@cl.on_message
async def on_message(message: cl.Message):
app = cl.user_session.get("app")
msg = cl.Message(content="")
for chunk in await cl.make_async(app.chat)(message.content):
await msg.stream_token(chunk)
await msg.send()
```
## Run the app
```
chainlit run app.py
```
## Try it out
Open the app in your browser and start chatting with it!
![chainlit-demo](https://github.com/embedchain/embedchain/assets/73601258/d6635624-5cdb-485b-bfbd-3b7c8f18bfff)
+14 -4
View File
@@ -76,7 +76,8 @@
{
"group": "🔗 Integrations",
"pages": [
"integration/langsmith"
"integration/langsmith",
"integration/chainlit"
]
},
"get-started/faq"
@@ -116,7 +117,8 @@
"components/data-sources/discourse",
"components/data-sources/substack",
"components/data-sources/discord",
"components/data-sources/beehiiv"
"components/data-sources/beehiiv",
"components/data-sources/directory"
]
},
"components/data-sources/data-type-handling"
@@ -136,6 +138,7 @@
{
"group": "Examples",
"pages": [
"examples/notebooks-and-replits",
{
"group": "REST API Service",
"pages": [
@@ -183,7 +186,8 @@
"api-reference/pipeline/chat",
"api-reference/pipeline/search",
"api-reference/pipeline/deploy",
"api-reference/pipeline/reset"
"api-reference/pipeline/reset",
"api-reference/pipeline/delete"
]
},
"api-reference/store/openai-assistant",
@@ -233,5 +237,11 @@
},
"api": {
"baseUrl": "http://localhost:8080"
}
},
"redirects": [
{
"source": "/changelog/command-line",
"destination": "/get-started/introduction"
}
]
}
+1 -1
View File
@@ -25,7 +25,7 @@ class ChunkerConfig(BaseConfig):
self.min_chunk_size = min_chunk_size
if self.min_chunk_size >= self.chunk_size:
raise ValueError(f"min_chunk_size {min_chunk_size} should be less than chunk_size {chunk_size}")
if self.min_chunk_size <= self.chunk_overlap:
if self.min_chunk_size < self.chunk_overlap:
logging.warn(
f"min_chunk_size {min_chunk_size} should be greater than chunk_overlap {chunk_overlap}, otherwise it is redundant." # noqa:E501
)
+18
View File
@@ -0,0 +1,18 @@
from typing import Optional
from embedchain.config.embedder.base import BaseEmbedderConfig
from embedchain.helpers.json_serializable import register_deserializable
@register_deserializable
class GoogleAIEmbedderConfig(BaseEmbedderConfig):
def __init__(
self,
model: Optional[str] = None,
deployment_name: Optional[str] = None,
task_type: Optional[str] = None,
title: Optional[str] = None,
):
super().__init__(model, deployment_name)
self.task_type = task_type or "retrieval_document"
self.title = title or "Embeddings for Embedchain"
+3 -2
View File
@@ -650,7 +650,7 @@ class EmbedChain(JSONSerializable):
self.db.reset()
self.cursor.execute("DELETE FROM data_sources WHERE pipeline_id = ?", (self.config.id,))
self.connection.commit()
self.delete_history()
self.delete_chat_history()
# Send anonymous telemetry
self.telemetry.capture(event_name="reset", properties=self._telemetry_props)
@@ -661,5 +661,6 @@ class EmbedChain(JSONSerializable):
display_format=display_format,
)
def delete_history(self):
def delete_chat_history(self):
self.llm.memory.delete_chat_history(app_id=self.config.id)
self.llm.update_history(app_id=self.config.id)
+31
View File
@@ -0,0 +1,31 @@
from typing import Optional
import google.generativeai as genai
from chromadb import EmbeddingFunction, Embeddings
from embedchain.config.embedder.google import GoogleAIEmbedderConfig
from embedchain.embedder.base import BaseEmbedder
from embedchain.models import VectorDimensions
class GoogleAIEmbeddingFunction(EmbeddingFunction):
def __init__(self, config: Optional[GoogleAIEmbedderConfig] = None) -> None:
super().__init__()
self.config = config or GoogleAIEmbedderConfig()
def __call__(self, input: str) -> Embeddings:
model = self.config.model
title = self.config.title
task_type = self.config.task_type
embeddings = genai.embed_content(model=model, content=input, task_type=task_type, title=title)
return embeddings["embedding"]
class GoogleAIEmbedder(BaseEmbedder):
def __init__(self, config: Optional[GoogleAIEmbedderConfig] = None):
super().__init__(config)
embedding_fn = GoogleAIEmbeddingFunction(config=config)
self.set_embedding_fn(embedding_fn=embedding_fn)
vector_dimension = VectorDimensions.GOOGLE_AI.value
self.set_vector_dimension(vector_dimension=vector_dimension)
+2
View File
@@ -47,11 +47,13 @@ class EmbedderFactory:
"huggingface": "embedchain.embedder.huggingface.HuggingFaceEmbedder",
"openai": "embedchain.embedder.openai.OpenAIEmbedder",
"vertexai": "embedchain.embedder.vertexai.VertexAIEmbedder",
"google": "embedchain.embedder.google.GoogleAIEmbedder",
}
provider_to_config_class = {
"azure_openai": "embedchain.config.embedder.base.BaseEmbedderConfig",
"openai": "embedchain.config.embedder.base.BaseEmbedderConfig",
"gpt4all": "embedchain.config.embedder.base.BaseEmbedderConfig",
"google": "embedchain.config.embedder.google.GoogleAIEmbedderConfig",
}
@classmethod
+4 -21
View File
@@ -48,8 +48,7 @@ class BaseLlm(JSONSerializable):
def update_history(self, app_id: str):
"""Update class history attribute with history in memory (for chat method)"""
chat_history = self.memory.get_recent_memories(app_id=app_id, num_rounds=10)
if chat_history:
self.set_history([str(history) for history in chat_history])
self.set_history([str(history) for history in chat_history])
def add_history(self, app_id: str, question: str, answer: str, metadata: Optional[Dict[str, Any]] = None):
chat_message = ChatMessage()
@@ -147,21 +146,7 @@ class BaseLlm(JSONSerializable):
logging.info(f"Access search to get answers for {input_query}")
return search.run(input_query)
def _stream_query_response(self, answer: Any) -> Generator[Any, Any, None]:
"""Generator to be used as streaming response
:param answer: Answer chunk from llm
:type answer: Any
:yield: Answer chunk from llm
:rtype: Generator[Any, Any, None]
"""
streamed_answer = ""
for chunk in answer:
streamed_answer = streamed_answer + chunk
yield chunk
logging.info(f"Answer: {streamed_answer}")
def _stream_chat_response(self, answer: Any) -> Generator[Any, Any, None]:
def _stream_response(self, answer: Any) -> Generator[Any, Any, None]:
"""Generator to be used as streaming response
:param answer: Answer chunk from llm
@@ -221,7 +206,7 @@ class BaseLlm(JSONSerializable):
logging.info(f"Answer: {answer}")
return answer
else:
return self._stream_query_response(answer)
return self._stream_response(answer)
finally:
if config:
# Restore previous config
@@ -270,14 +255,12 @@ class BaseLlm(JSONSerializable):
return prompt
answer = self.get_answer_from_llm(prompt)
if isinstance(answer, str):
logging.info(f"Answer: {answer}")
return answer
else:
# this is a streamed response and needs to be handled differently.
return self._stream_chat_response(answer)
return self._stream_response(answer)
finally:
if config:
# Restore previous config
+14 -14
View File
@@ -1,7 +1,7 @@
import importlib
import logging
import os
from typing import Optional
from typing import Any, Generator, Optional, Union
import google.generativeai as genai
@@ -30,22 +30,22 @@ class GoogleLlm(BaseLlm):
def get_llm_model_answer(self, prompt):
if self.config.system_prompt:
raise ValueError("GoogleLlm does not support `system_prompt`")
return GoogleLlm._get_answer(prompt, self.config)
response = self._get_answer(prompt)
return response
@staticmethod
def _get_answer(prompt: str, config: BaseLlmConfig):
model_name = config.model or "gemini-pro"
def _get_answer(self, prompt: str) -> Union[str, Generator[Any, Any, None]]:
model_name = self.config.model or "gemini-pro"
logging.info(f"Using Google LLM model: {model_name}")
model = genai.GenerativeModel(model_name=model_name)
generation_config_params = {
"candidate_count": 1,
"max_output_tokens": config.max_tokens,
"temperature": config.temperature or 0.5,
"max_output_tokens": self.config.max_tokens,
"temperature": self.config.temperature or 0.5,
}
if config.top_p >= 0.0 and config.top_p <= 1.0:
generation_config_params["top_p"] = config.top_p
if self.config.top_p >= 0.0 and self.config.top_p <= 1.0:
generation_config_params["top_p"] = self.config.top_p
else:
raise ValueError("`top_p` must be > 0.0 and < 1.0")
@@ -54,11 +54,11 @@ class GoogleLlm(BaseLlm):
response = model.generate_content(
prompt,
generation_config=generation_config,
stream=config.stream,
stream=self.config.stream,
)
if config.stream:
for chunk in response:
yield chunk.text
if self.config.stream:
# TODO: Implement streaming
response.resolve()
return response.text
else:
return response.text
-1
View File
@@ -85,7 +85,6 @@ class GithubLoader(BaseLoader):
logging.info("Fetch completed.")
else:
logging.info("Cloning repository...")
Repo.clone_from(repo_url, local_path)
repo = Repo.clone_from(repo_url, local_path)
logging.info("Clone completed.")
return repo.head.commit.tree
+1
View File
@@ -7,3 +7,4 @@ class VectorDimensions(Enum):
OPENAI = 1536
VERTEX_AI = 768
HUGGING_FACE = 384
GOOGLE_AI = 768
+2 -2
View File
@@ -411,14 +411,14 @@ def validate_config(config_data):
Optional("config"): object, # TODO: add particular config schema for each provider
},
Optional("embedder"): {
Optional("provider"): Or("openai", "gpt4all", "huggingface", "vertexai", "azure_openai"),
Optional("provider"): Or("openai", "gpt4all", "huggingface", "vertexai", "azure_openai", "google"),
Optional("config"): {
Optional("model"): Optional(str),
Optional("deployment_name"): Optional(str),
},
},
Optional("embedding_model"): {
Optional("provider"): Or("openai", "gpt4all", "huggingface", "vertexai", "azure_openai"),
Optional("provider"): Or("openai", "gpt4all", "huggingface", "vertexai", "azure_openai", "google"),
Optional("config"): {
Optional("model"): str,
Optional("deployment_name"): str,
+1
View File
@@ -0,0 +1 @@
.chainlit
+17
View File
@@ -0,0 +1,17 @@
## Chainlit + Embedchain Demo
In this example, we will learn how to use Chainlit and Embedchain together
## Setup
First, install the required packages:
```bash
pip install -r requirements.txt
```
## Run the app locally,
```
chainlit run app.py
```
+35
View File
@@ -0,0 +1,35 @@
import chainlit as cl
from embedchain import Pipeline as App
import os
os.environ["OPENAI_API_KEY"] = "sk-xxx"
@cl.on_chat_start
async def on_chat_start():
app = App.from_config(config={
'app': {
'config': {
'name': 'chainlit-app'
}
},
'llm': {
'config': {
'stream': True,
}
}
})
# import your data here
app.add("https://www.forbes.com/profile/elon-musk/")
app.collect_metrics = False
cl.user_session.set("app", app)
@cl.on_message
async def on_message(message: cl.Message):
app = cl.user_session.get("app")
msg = cl.Message(content="")
for chunk in await cl.make_async(app.chat)(message.content):
await msg.stream_token(chunk)
await msg.send()
+15
View File
@@ -0,0 +1,15 @@
# Welcome to Embedchain! 🚀
Hello! 👋 Excited to see you join us. With Embedchain and Chainlit, create ChatGPT like apps effortlessly.
## Quick Start 🌟
- **Embedchain Docs:** Get started with our comprehensive [Embedchain Documentation](https://docs.embedchain.ai/) 📚
- **Discord Community:** Join our discord [Embedchain Discord](https://discord.gg/CUU9FPhRNt) to ask questions, share your projects, and connect with other developers! 💬
- **UI Guide**: Master Chainlit with [Chainlit Documentation](https://docs.chainlit.io/) ⛓️
Happy building with Embedchain! 🎉
## Customize welcome screen
Edit chainlit.md in your project root to change this welcome message.
+2
View File
@@ -0,0 +1,2 @@
chainlit==0.7.700
embedchain==0.1.31
-1
View File
@@ -90,7 +90,6 @@
" provider: openai\n",
" config:\n",
" model: text-embedding-ada-002\n",
" deployment_name: ec_embeddings_ada_002\n",
"\"\"\"\n",
"\n",
"# Write the multi-line string to a YAML file\n",
+1 -1
View File
@@ -1,6 +1,6 @@
[tool.poetry]
name = "embedchain"
version = "0.1.33"
version = "0.1.34"
description = "Data platform for LLMs - Load, index, retrieve and sync any unstructured data"
authors = [
"Taranjeet Singh <taranjeet@embedchain.ai>",
+2 -8
View File
@@ -38,15 +38,9 @@ def test_is_get_llm_model_answer_implemented():
assert llm.get_llm_model_answer() == "Implemented"
def test_stream_query_response(base_llm):
def test_stream_response(base_llm):
answer = ["Chunk1", "Chunk2", "Chunk3"]
result = list(base_llm._stream_query_response(answer))
assert result == answer
def test_stream_chat_response(base_llm):
answer = ["Chunk1", "Chunk2", "Chunk3"]
result = list(base_llm._stream_chat_response(answer))
result = list(base_llm._stream_response(answer))
assert result == answer