Compare commits

...

14 Commits

Author SHA1 Message Date
Taranjeet Singh 1b19d0d19c bump version to 0.0.48 (#503) 2023-08-29 00:31:50 +05:30
Taranjeet Singh 13f01e399c fix: add tiktoken as required dependency in pyproject (#502) 2023-08-29 00:29:12 +05:30
Taranjeet Singh 261e2d088c Bump version to 0.0.47 (#501) 2023-08-29 00:00:35 +05:30
Taranjeet Singh b0ae3e95c7 fix: update poe bot creation docs (#500) 2023-08-28 23:59:34 +05:30
cachho fc633dadeb feat: poe bot (#492)
Co-authored-by: Taranjeet Singh <reachtotj@gmail.com>
2023-08-28 23:47:01 +05:30
Taranjeet Singh aafb334916 fix: improve CTA of book meeting (#496) 2023-08-28 08:19:47 +05:30
Taranjeet Singh 04d851e802 fix: update discord link (#494) 2023-08-28 05:33:03 +05:30
Taranjeet Singh de31c63dac feat: Add cal.com link for feedback (#493) 2023-08-28 05:12:54 +05:30
Deshraj Yadav c068f58543 [version]: bump package version to 0.0.46 2023-08-27 12:24:37 -07:00
Deshraj Yadav 70df373807 [feat]: add support for sending anonymous user_id in telemetry (#491) 2023-08-28 00:53:42 +05:30
Deshraj Yadav 4388f6bfc2 [bug-fix] fix issue related to bot memory when using multiple bots at the same time (#486) 2023-08-25 21:59:39 -07:00
Taranjeet Singh d0956a0dc1 Bump version to 0.0.45 (#479) 2023-08-25 03:22:04 +05:30
cachho ccf515cadd Fix/use poetry for build (#478) 2023-08-25 03:20:57 +05:30
Taranjeet Singh a6e4235bb0 fix: update docs (#477) 2023-08-25 03:16:15 +05:30
9 changed files with 200 additions and 58 deletions
+4 -1
View File
@@ -1,5 +1,8 @@
blank_issues_enabled: true
contact_links:
- name: 1-on-1 Session
url: https://cal.com/taranjeetio/ec
about: Speak directly with Taranjeet, the founder, to discuss issues, share feedback, or explore improvements for Embedchain
- name: Discord
url: https://discord.gg/6PzXDgEjG5
about: General community discussions
about: General community discussions
+8 -4
View File
@@ -19,12 +19,16 @@ jobs:
with:
python-version: '3.11'
- name: Install pep517
- name: Install Poetry
run: |
python -m pip install pep517 --user
curl -sSL https://install.python-poetry.org | python3 -
echo "$HOME/.local/bin" >> $GITHUB_PATH
- name: Install dependencies
run: poetry install
- name: Build a binary wheel and a source tarball
run: python -m pep517.build .
run: poetry build
- name: Publish distribution 📦 to Test PyPI
uses: pypa/gh-action-pypi-publish@release/v1
+15 -24
View File
@@ -1,42 +1,23 @@
# embedchain
[![PyPI](https://img.shields.io/pypi/v/embedchain)](https://pypi.org/project/embedchain/)
[![Discord](https://dcbadge.vercel.app/api/server/6PzXDgEjG5?style=flat)](https://discord.gg/6PzXDgEjG5)
[![Discord](https://dcbadge.vercel.app/api/server/6PzXDgEjG5?style=flat)](https://discord.gg/CUU9FPhRNt)
[![Twitter](https://img.shields.io/twitter/follow/embedchain)](https://twitter.com/embedchain)
[![Substack](https://img.shields.io/badge/Substack-%23006f5c.svg?logo=substack)](https://embedchain.substack.com/)
[![Open in Colab](https://camo.githubusercontent.com/84f0493939e0c4de4e6dbe113251b4bfb5353e57134ffd9fcab6b8714514d4d1/68747470733a2f2f636f6c61622e72657365617263682e676f6f676c652e636f6d2f6173736574732f636f6c61622d62616467652e737667)](https://colab.research.google.com/drive/138lMWhENGeEu7Q1-6lNbNTHGLZXBBz_B?usp=sharing)
Embedchain is a framework to easily create LLM powered bots over any dataset. If you want a javascript version, check out [embedchain-js](https://github.com/embedchain/embedchainjs)
## 🤝 Schedule a 1-on-1 Session
Book a [1-on-1 Session](https://cal.com/taranjeetio/ec) with Taranjeet, the founder, to discuss any issues, provide feedback, or explore how we can improve Embedchain for you.
## 🔧 Quick install
```bash
pip install embedchain
```
## 🔥 Latest
- **[2023/07/19]** Released support for 🦙 `llama2` model. Start creating your `llama2` based bots like this:
```python
import os
from embedchain import Llama2App
os.environ['REPLICATE_API_TOKEN'] = "REPLICATE API TOKEN"
zuck_bot = Llama2App()
# Embed your data
zuck_bot.add("https://www.youtube.com/watch?v=Ff4fRgnuFgQ")
zuck_bot.add("https://en.wikipedia.org/wiki/Mark_Zuckerberg")
# Nice, your bot is ready now. Start asking questions to your bot.
zuck_bot.query("Who is Mark Zuckerberg?")
# Answer: Mark Zuckerberg is an American internet entrepreneur and business magnate. He is the co-founder and CEO of Facebook.
```
## 🔍 Demo
Try out embedchain in your browser:
@@ -51,6 +32,16 @@ The documentation for embedchain can be found at [docs.embedchain.ai](https://do
Embedchain empowers you to create chatbot models similar to ChatGPT, using your own evolving dataset.
### Data Types Supported
* Youtube video
* PDF file
* Web page
* Sitemap
* Doc file
* Code documentation website loader
* Notion
### Queries
For example, you can use Embedchain to create an Elon Musk bot using the following code:
+48
View File
@@ -0,0 +1,48 @@
---
title: '🔮 Poe Bot'
---
### 🚀 Getting started
1. Install embedchain python package:
```bash
pip install embedchain[poe]
```
2. Create a free account on [Poe](https://www.poe.com?utm_source=embedchain).
3. Click "Create Bot" button on top left
4. Give it a handle and an optional description.
5. Select `Use API`.
6. Under `API URL` enter your server or ngrok address. You can use your machine's public IP or DNS. Otherwise, employ a proxy server like [ngrok](https://ngrok.com/) to make your local bot accessible.
7. Copy your api key and paste it in `.env` as `POE_API_KEY`.
8. Start the bot.
```bash
python -m embedchain.bots.poe
```
If you want to run the bot on another port, you can pass `--port option` like
```bash
python -m embedchain.bots.poe --port 5000
```
9. Click `Run check` to make sure your machine can be reached.
10. Make sure your bot is private if that's what you want.
11. Click `Create bot` at the bottom to finally create the bot
12. Now you bot is created.
### 💬 How to use
- To include data sources, use this command:
```text
/add <url_or_text>
```
- You can refer the [Supported Data formats](https://docs.embedchain.ai/advanced/data_types) section to refer the supported data types in embedchain.
- To ask the bot questions, just type your query:
```text
<your-question-here>
```
+3 -2
View File
@@ -36,7 +36,7 @@
},
{
"group": "Examples",
"pages": ["examples/full_stack", "examples/api_server", "examples/discord_bot", "examples/slack_bot", "examples/telegram_bot", "examples/whatsapp_bot"]
"pages": ["examples/full_stack", "examples/api_server", "examples/discord_bot", "examples/slack_bot", "examples/telegram_bot", "examples/whatsapp_bot", "examples/poe_bot"]
},
{
"group": "Contribution Guidelines",
@@ -47,7 +47,8 @@
"footerSocials": {
"twitter": "https://twitter.com/embedchain",
"github": "https://github.com/embedchain/embedchain",
"linkedin": "https://www.linkedin.com/company/embedchain"
"linkedin": "https://www.linkedin.com/company/embedchain",
"website": "https://embedchain.ai"
},
"backgroundImage": "/background.png",
"isWhiteLabeled": true
+79
View File
@@ -0,0 +1,79 @@
import argparse
import logging
import os
from typing import List, Optional
from fastapi_poe import PoeBot, run
from embedchain.config import QueryConfig
from .base import BaseBot
class EcPoeBot(BaseBot, PoeBot):
def __init__(self):
self.history_length = 5
super().__init__()
async def get_response(self, query):
last_message = query.query[-1].content
try:
history = (
[f"{m.role}: {m.content}" for m in query.query[-(self.history_length + 1) : -1]]
if len(query.query) > 0
else None
)
except Exception as e:
logging.error(f"Error when processing the chat history. Message is being sent without history. Error: {e}")
logging.warning(history)
answer = self.handle_message(last_message, history)
yield self.text_event(answer)
def handle_message(self, message, history: Optional[List[str]] = None):
if message.startswith("/add "):
response = self.add_data(message)
else:
response = self.ask_bot(message, history)
return response
def add_data(self, message):
data = message.split(" ")[-1]
try:
self.add(data)
response = f"Added data from: {data}"
except Exception:
logging.exception(f"Failed to add data {data}.")
response = "Some error occurred while adding data."
return response
def ask_bot(self, message, history: List[str]):
try:
config = QueryConfig(history=history)
response = self.query(message, config)
except Exception:
logging.exception(f"Failed to query {message}.")
response = "An error occurred. Please try again!"
return response
def start_command():
parser = argparse.ArgumentParser(description="EmbedChain PoeBot command line interface")
# parser.add_argument("--host", default="0.0.0.0", help="Host IP to bind")
parser.add_argument("--port", default=8080, type=int, help="Port to bind")
parser.add_argument("--api-key", type=str, help="Poe API key")
# parser.add_argument(
# "--history-length",
# default=5,
# type=int,
# help="Set the max size of the chat history. Multiplies cost, but improves conversation awareness.",
# )
args = parser.parse_args()
# FIXME: Arguments are automatically loaded by Poebot's ArgumentParser which causes it to fail.
# the port argument here is also just for show, it actually works because poe has the same argument.
run(EcPoeBot(), api_key=args.api_key or os.environ.get("POE_API_KEY"))
if __name__ == "__main__":
start_command()
+34 -8
View File
@@ -1,9 +1,11 @@
import hashlib
import importlib.metadata
import json
import logging
import os
import threading
import uuid
from pathlib import Path
from typing import Dict, Optional
import requests
@@ -25,8 +27,9 @@ load_dotenv()
ABS_PATH = os.getcwd()
DB_DIR = os.path.join(ABS_PATH, "db")
memory = ConversationBufferMemory()
HOME_DIR = str(Path.home())
CONFIG_DIR = os.path.join(HOME_DIR, ".embedchain")
CONFIG_FILE = os.path.join(CONFIG_DIR, "config.json")
class EmbedChain:
@@ -44,12 +47,34 @@ class EmbedChain:
self.user_asks = []
self.is_docs_site_instance = False
self.online = False
self.memory = ConversationBufferMemory()
# Send anonymous telemetry
self.s_id = self.config.id if self.config.id else str(uuid.uuid4())
self.u_id = self._load_or_generate_user_id()
thread_telemetry = threading.Thread(target=self._send_telemetry_event, args=("init",))
thread_telemetry.start()
def _load_or_generate_user_id(self):
"""
Loads the user id from the config file if it exists, otherwise generates a new
one and saves it to the config file.
"""
if not os.path.exists(CONFIG_DIR):
os.makedirs(CONFIG_DIR)
if os.path.exists(CONFIG_FILE):
with open(CONFIG_FILE, "r") as f:
data = json.load(f)
if "user_id" in data:
return data["user_id"]
u_id = str(uuid.uuid4())
with open(CONFIG_FILE, "w") as f:
json.dump({"user_id": u_id}, f)
return u_id
def add(
self,
source,
@@ -362,8 +387,7 @@ class EmbedChain:
k["web_search_result"] = self.access_search_and_get_results(input_query)
contexts = self.retrieve_from_database(input_query, config)
global memory
chat_history = memory.load_memory_variables({})["history"]
chat_history = self.memory.load_memory_variables({})["history"]
if chat_history:
config.set_history(chat_history)
@@ -376,14 +400,14 @@ class EmbedChain:
answer = self.get_answer_from_llm(prompt, config)
memory.chat_memory.add_user_message(input_query)
self.memory.chat_memory.add_user_message(input_query)
# Send anonymous telemetry
thread_telemetry = threading.Thread(target=self._send_telemetry_event, args=("chat",))
thread_telemetry.start()
if isinstance(answer, str):
memory.chat_memory.add_ai_message(answer)
self.memory.chat_memory.add_ai_message(answer)
logging.info(f"Answer: {answer}")
return answer
else:
@@ -395,7 +419,7 @@ class EmbedChain:
for chunk in answer:
streamed_answer = streamed_answer + chunk
yield chunk
memory.chat_memory.add_ai_message(streamed_answer)
self.memory.chat_memory.add_ai_message(streamed_answer)
logging.info(f"Answer: {streamed_answer}")
def set_collection(self, collection_name):
@@ -445,9 +469,11 @@ class EmbedChain:
"version": importlib.metadata.version(__package__ or __name__),
"method": method,
"language": "py",
"u_id": self.u_id,
}
if extra_metadata:
metadata.update(extra_metadata)
response = requests.post(url, json={"metadata": metadata})
response.raise_for_status()
if response.status_code != 200:
logging.warning(f"Telemetry event failed with status code {response.status_code}")
+4 -2
View File
@@ -1,6 +1,6 @@
[tool.poetry]
name = "embedchain"
version = "0.0.44"
version = "0.0.48"
description = "embedchain is a framework to easily create LLM powered bots over any dataset"
authors = ["Taranjeet Singh"]
license = "Apache License"
@@ -85,6 +85,7 @@ python-dotenv = "^1.0.0"
langchain = "^0.0.237"
requests = "^2.31.0"
openai = "^0.27.5"
tiktoken = "^0.4.0"
chromadb ="^0.4.2"
youtube-transcript-api = "^0.6.1"
beautifulsoup4 = "^4.12.2"
@@ -98,6 +99,7 @@ gpt4all = { version = "^1.0.8", optional = true }
elasticsearch = { version = "^8.9.0", optional = true }
flask = "^2.3.3"
twilio = "^8.5.0"
fastapi-poe = { version = "0.0.16", optional = true }
@@ -116,10 +118,10 @@ streamlit = ["streamlit"]
community = ["llama-index"]
opensource = ["sentence-transformers", "torch", "gpt4all"]
elasticsearch = ["elasticsearch"]
poe = ["fastapi-poe"]
[tool.poetry.group.docs.dependencies]
[tool.poetry.scripts]
+5 -17
View File
@@ -7,15 +7,13 @@ from embedchain.config import AppConfig
class TestApp(unittest.TestCase):
os.environ["OPENAI_API_KEY"] = "test_key"
def setUp(self):
os.environ["OPENAI_API_KEY"] = "test_key"
self.app = App(config=AppConfig(collect_metrics=False))
@patch("embedchain.embedchain.memory", autospec=True)
@patch.object(App, "retrieve_from_database", return_value=["Test context"])
@patch.object(App, "get_answer_from_llm", return_value="Test answer")
def test_chat_with_memory(self, mock_answer, mock_retrieve, mock_memory):
def test_chat_with_memory(self, mock_get_answer, mock_retrieve):
"""
This test checks the functionality of the 'chat' method in the App class with respect to the chat history
memory.
@@ -23,27 +21,17 @@ class TestApp(unittest.TestCase):
The second call is expected to use the chat history from the first call.
Key assumptions tested:
- After the first call, 'memory.chat_memory.add_user_message' and 'memory.chat_memory.add_ai_message' are
called with correct arguments, adding the correct chat history.
- After the first call, 'memory.chat_memory.add_user_message' and 'memory.chat_memory.add_ai_message' are
- During the second call, the 'chat' method uses the chat history from the first call.
The test isolates the 'chat' method behavior by mocking out 'retrieve_from_database', 'get_answer_from_llm' and
'memory' methods.
"""
mock_memory.load_memory_variables.return_value = {"history": []}
app = App()
# First call to chat
first_answer = app.chat("Test query 1")
self.assertEqual(first_answer, "Test answer")
mock_memory.chat_memory.add_user_message.assert_called_once_with("Test query 1")
mock_memory.chat_memory.add_ai_message.assert_called_once_with("Test answer")
mock_memory.chat_memory.add_user_message.reset_mock()
mock_memory.chat_memory.add_ai_message.reset_mock()
# Second call to chat
self.assertEqual(len(app.memory.chat_memory.messages), 2)
second_answer = app.chat("Test query 2")
self.assertEqual(second_answer, "Test answer")
mock_memory.chat_memory.add_user_message.assert_called_once_with("Test query 2")
mock_memory.chat_memory.add_ai_message.assert_called_once_with("Test answer")
self.assertEqual(len(app.memory.chat_memory.messages), 4)