diff --git a/.gitignore b/.gitignore
index f8fd29897..760e3c59d 100644
--- a/.gitignore
+++ b/.gitignore
@@ -170,7 +170,6 @@ cython_debug/
# Database
db
test-db
-!embedchain/embedchain/core/db/
.vscode
.idea/
diff --git a/AGENTS.md b/AGENTS.md
index 52d8ffc1e..7754b84a2 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -33,7 +33,6 @@ This is a **polyglot monorepo** containing Python and TypeScript packages, CLIs,
| `evaluation/` | Benchmarking framework — LOCOMO evals, experiment runner, score generation |
| `examples/` | Sample projects — demo apps, Chrome extension, multi-agent patterns |
| `cookbooks/` | Jupyter notebooks — customer support chatbot, AutoGen integration |
-| `embedchain/` | Legacy Embedchain RAG framework (maintained separately, Poetry-based) |
| `pr-reviews/` | Pull request review materials |
| `scripts/` | Repo-wide utility scripts (e.g., `check-llms-txt-coverage.py` for docs/llms.txt sync) |
@@ -330,7 +329,7 @@ make run-openai # OpenAI comparison
- Root SDK: line length **120**
- Python CLI: line length **100** with extended rule set (UP, B, SIM, RUF)
- **isort** with `profile = "black"` for import sorting.
-- Ruff excludes `embedchain/` and `openmemory/` from root config.
+- Ruff excludes `openmemory/` from root config.
### TypeScript Conventions
@@ -414,7 +413,6 @@ To add a new LLM, embedding, vector store, or reranker provider:
| Python CLI | `cli-python-ci.yml` | Push to `cli/python/`, PRs, manual | Ruff lint + pytest + hatch build on Python 3.10, 3.11, 3.12 |
| Node CLI | `cli-node-ci.yml` | Push to `cli/node/`, PRs, manual | Biome lint + tsc + vitest + tsup build on Node 20, 22 |
| OpenClaw | `openclaw-checks.yml` | Push to `openclaw/`, PRs, manual | tsc + vitest (with Codecov) + tsup build on Node 20, 22 |
-| Embedchain | `ci.yml` (shared) | PRs on `embedchain/` | Ruff + pytest + coverage on Python 3.9–3.12 |
### CD Workflows (automated publishing)
@@ -576,7 +574,6 @@ N/A
- Modify CI/CD workflows without explicit approval.
- Add new Python dependencies to the core `dependencies` list in `pyproject.toml` without discussion — use optional dependency groups instead.
- Commit `.env` files, API keys, or credentials.
-- Modify `embedchain/` unless specifically working on that package — it has its own build system (Poetry).
- Skip pre-commit hooks.
- Use npm or yarn in TypeScript packages — this repo uses pnpm exclusively.
- Use `require()` for imports in TypeScript — use ES module `import` syntax.
diff --git a/embedchain/CITATION.cff b/embedchain/CITATION.cff
deleted file mode 100644
index 8b93297cd..000000000
--- a/embedchain/CITATION.cff
+++ /dev/null
@@ -1,8 +0,0 @@
-cff-version: 1.2.0
-message: "If you use this software, please cite it as below."
-authors:
-- family-names: "Singh"
- given-names: "Taranjeet"
-title: "Embedchain"
-date-released: 2023-06-20
-url: "https://github.com/embedchain/embedchain"
\ No newline at end of file
diff --git a/embedchain/CONTRIBUTING.md b/embedchain/CONTRIBUTING.md
deleted file mode 100644
index a0d7c12e8..000000000
--- a/embedchain/CONTRIBUTING.md
+++ /dev/null
@@ -1,76 +0,0 @@
-# Contributing to embedchain
-
-Let us make contribution easy, collaborative and fun.
-
-## Submit your Contribution through PR
-
-To make a contribution, follow these steps:
-
-1. Fork and clone this repository
-2. Do the changes on your fork with dedicated feature branch `feature/f1`
-3. If you modified the code (new feature or bug-fix), please add tests for it
-4. Include proper documentation / docstring and examples to run the feature
-5. Check the linting
-6. Ensure that all tests pass
-7. Submit a pull request
-
-For more details about pull requests, please read [GitHub's guides](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/proposing-changes-to-your-work-with-pull-requests/creating-a-pull-request).
-
-
-### 📦 Package manager
-
-We use `poetry` as our package manager. You can install poetry by following the instructions [here](https://python-poetry.org/docs/#installation).
-
-Please DO NOT use pip or conda to install the dependencies. Instead, use poetry:
-
-```bash
-make install_all
-
-#activate
-
-poetry shell
-```
-
-### 📌 Pre-commit
-
-To ensure our standards, make sure to install pre-commit before starting to contribute.
-
-```bash
-pre-commit install
-```
-
-### 🧹 Linting
-
-We use `ruff` to lint our code. You can run the linter by running the following command:
-
-```bash
-make lint
-```
-
-Make sure that the linter does not report any errors or warnings before submitting a pull request.
-
-### Code Formatting with `black`
-
-We use `black` to reformat the code by running the following command:
-
-```bash
-make format
-```
-
-### 🧪 Testing
-
-We use `pytest` to test our code. You can run the tests by running the following command:
-
-```bash
-poetry run pytest
-```
-
-
-Several packages have been removed from Poetry to make the package lighter. Therefore, it is recommended to run `make install_all` to install the remaining packages and ensure all tests pass.
-
-
-Make sure that all tests pass before submitting a pull request.
-
-## 🚀 Release Process
-
-At the moment, the release process is manual. We try to make frequent releases. Usually, we release a new version when we have a new feature or bugfix. A developer with admin rights to the repository will create a new release on GitHub, and then publish the new version to PyPI.
diff --git a/embedchain/LICENSE b/embedchain/LICENSE
deleted file mode 100644
index d20d5102c..000000000
--- a/embedchain/LICENSE
+++ /dev/null
@@ -1,201 +0,0 @@
- Apache License
- Version 2.0, January 2004
- http://www.apache.org/licenses/
-
- TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
-
- 1. Definitions.
-
- "License" shall mean the terms and conditions for use, reproduction,
- and distribution as defined by Sections 1 through 9 of this document.
-
- "Licensor" shall mean the copyright owner or entity authorized by
- the copyright owner that is granting the License.
-
- "Legal Entity" shall mean the union of the acting entity and all
- other entities that control, are controlled by, or are under common
- control with that entity. For the purposes of this definition,
- "control" means (i) the power, direct or indirect, to cause the
- direction or management of such entity, whether by contract or
- otherwise, or (ii) ownership of fifty percent (50%) or more of the
- outstanding shares, or (iii) beneficial ownership of such entity.
-
- "You" (or "Your") shall mean an individual or Legal Entity
- exercising permissions granted by this License.
-
- "Source" form shall mean the preferred form for making modifications,
- including but not limited to software source code, documentation
- source, and configuration files.
-
- "Object" form shall mean any form resulting from mechanical
- transformation or translation of a Source form, including but
- not limited to compiled object code, generated documentation,
- and conversions to other media types.
-
- "Work" shall mean the work of authorship, whether in Source or
- Object form, made available under the License, as indicated by a
- copyright notice that is included in or attached to the work
- (an example is provided in the Appendix below).
-
- "Derivative Works" shall mean any work, whether in Source or Object
- form, that is based on (or derived from) the Work and for which the
- editorial revisions, annotations, elaborations, or other modifications
- represent, as a whole, an original work of authorship. For the purposes
- of this License, Derivative Works shall not include works that remain
- separable from, or merely link (or bind by name) to the interfaces of,
- the Work and Derivative Works thereof.
-
- "Contribution" shall mean any work of authorship, including
- the original version of the Work and any modifications or additions
- to that Work or Derivative Works thereof, that is intentionally
- submitted to Licensor for inclusion in the Work by the copyright owner
- or by an individual or Legal Entity authorized to submit on behalf of
- the copyright owner. For the purposes of this definition, "submitted"
- means any form of electronic, verbal, or written communication sent
- to the Licensor or its representatives, including but not limited to
- communication on electronic mailing lists, source code control systems,
- and issue tracking systems that are managed by, or on behalf of, the
- Licensor for the purpose of discussing and improving the Work, but
- excluding communication that is conspicuously marked or otherwise
- designated in writing by the copyright owner as "Not a Contribution."
-
- "Contributor" shall mean Licensor and any individual or Legal Entity
- on behalf of whom a Contribution has been received by Licensor and
- subsequently incorporated within the Work.
-
- 2. Grant of Copyright License. Subject to the terms and conditions of
- this License, each Contributor hereby grants to You a perpetual,
- worldwide, non-exclusive, no-charge, royalty-free, irrevocable
- copyright license to reproduce, prepare Derivative Works of,
- publicly display, publicly perform, sublicense, and distribute the
- Work and such Derivative Works in Source or Object form.
-
- 3. Grant of Patent License. Subject to the terms and conditions of
- this License, each Contributor hereby grants to You a perpetual,
- worldwide, non-exclusive, no-charge, royalty-free, irrevocable
- (except as stated in this section) patent license to make, have made,
- use, offer to sell, sell, import, and otherwise transfer the Work,
- where such license applies only to those patent claims licensable
- by such Contributor that are necessarily infringed by their
- Contribution(s) alone or by combination of their Contribution(s)
- with the Work to which such Contribution(s) was submitted. If You
- institute patent litigation against any entity (including a
- cross-claim or counterclaim in a lawsuit) alleging that the Work
- or a Contribution incorporated within the Work constitutes direct
- or contributory patent infringement, then any patent licenses
- granted to You under this License for that Work shall terminate
- as of the date such litigation is filed.
-
- 4. Redistribution. You may reproduce and distribute copies of the
- Work or Derivative Works thereof in any medium, with or without
- modifications, and in Source or Object form, provided that You
- meet the following conditions:
-
- (a) You must give any other recipients of the Work or
- Derivative Works a copy of this License; and
-
- (b) You must cause any modified files to carry prominent notices
- stating that You changed the files; and
-
- (c) You must retain, in the Source form of any Derivative Works
- that You distribute, all copyright, patent, trademark, and
- attribution notices from the Source form of the Work,
- excluding those notices that do not pertain to any part of
- the Derivative Works; and
-
- (d) If the Work includes a "NOTICE" text file as part of its
- distribution, then any Derivative Works that You distribute must
- include a readable copy of the attribution notices contained
- within such NOTICE file, excluding those notices that do not
- pertain to any part of the Derivative Works, in at least one
- of the following places: within a NOTICE text file distributed
- as part of the Derivative Works; within the Source form or
- documentation, if provided along with the Derivative Works; or,
- within a display generated by the Derivative Works, if and
- wherever such third-party notices normally appear. The contents
- of the NOTICE file are for informational purposes only and
- do not modify the License. You may add Your own attribution
- notices within Derivative Works that You distribute, alongside
- or as an addendum to the NOTICE text from the Work, provided
- that such additional attribution notices cannot be construed
- as modifying the License.
-
- You may add Your own copyright statement to Your modifications and
- may provide additional or different license terms and conditions
- for use, reproduction, or distribution of Your modifications, or
- for any such Derivative Works as a whole, provided Your use,
- reproduction, and distribution of the Work otherwise complies with
- the conditions stated in this License.
-
- 5. Submission of Contributions. Unless You explicitly state otherwise,
- any Contribution intentionally submitted for inclusion in the Work
- by You to the Licensor shall be under the terms and conditions of
- this License, without any additional terms or conditions.
- Notwithstanding the above, nothing herein shall supersede or modify
- the terms of any separate license agreement you may have executed
- with Licensor regarding such Contributions.
-
- 6. Trademarks. This License does not grant permission to use the trade
- names, trademarks, service marks, or product names of the Licensor,
- except as required for reasonable and customary use in describing the
- origin of the Work and reproducing the content of the NOTICE file.
-
- 7. Disclaimer of Warranty. Unless required by applicable law or
- agreed to in writing, Licensor provides the Work (and each
- Contributor provides its Contributions) on an "AS IS" BASIS,
- WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
- implied, including, without limitation, any warranties or conditions
- of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
- PARTICULAR PURPOSE. You are solely responsible for determining the
- appropriateness of using or redistributing the Work and assume any
- risks associated with Your exercise of permissions under this License.
-
- 8. Limitation of Liability. In no event and under no legal theory,
- whether in tort (including negligence), contract, or otherwise,
- unless required by applicable law (such as deliberate and grossly
- negligent acts) or agreed to in writing, shall any Contributor be
- liable to You for damages, including any direct, indirect, special,
- incidental, or consequential damages of any character arising as a
- result of this License or out of the use or inability to use the
- Work (including but not limited to damages for loss of goodwill,
- work stoppage, computer failure or malfunction, or any and all
- other commercial damages or losses), even if such Contributor
- has been advised of the possibility of such damages.
-
- 9. Accepting Warranty or Additional Liability. While redistributing
- the Work or Derivative Works thereof, You may choose to offer,
- and charge a fee for, acceptance of support, warranty, indemnity,
- or other liability obligations and/or rights consistent with this
- License. However, in accepting such obligations, You may act only
- on Your own behalf and on Your sole responsibility, not on behalf
- of any other Contributor, and only if You agree to indemnify,
- defend, and hold each Contributor harmless for any liability
- incurred by, or claims asserted against, such Contributor by reason
- of your accepting any such warranty or additional liability.
-
- END OF TERMS AND CONDITIONS
-
- APPENDIX: How to apply the Apache License to your work.
-
- To apply the Apache License to your work, attach the following
- boilerplate notice, with the fields enclosed by brackets "[]"
- replaced with your own identifying information. (Don't include
- the brackets!) The text should be enclosed in the appropriate
- comment syntax for the file format. We also recommend that a
- file or class name and description of purpose be included on the
- same "printed page" as the copyright notice for easier
- identification within third-party archives.
-
- Copyright [2023] [Taranjeet Singh]
-
- Licensed under the Apache License, Version 2.0 (the "License");
- you may not use this file except in compliance with the License.
- You may obtain a copy of the License at
-
- http://www.apache.org/licenses/LICENSE-2.0
-
- Unless required by applicable law or agreed to in writing, software
- distributed under the License is distributed on an "AS IS" BASIS,
- WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
- See the License for the specific language governing permissions and
- limitations under the License.
diff --git a/embedchain/Makefile b/embedchain/Makefile
deleted file mode 100644
index f9ecc81fc..000000000
--- a/embedchain/Makefile
+++ /dev/null
@@ -1,56 +0,0 @@
-# Variables
-PYTHON := python3
-PIP := $(PYTHON) -m pip
-PROJECT_NAME := embedchain
-
-# Targets
-.PHONY: install format lint clean test ci_lint ci_test coverage
-
-install:
- poetry install
-
-# TODO: use a more efficient way to install these packages
-install_all:
- poetry install --all-extras
- poetry run pip install ruff==0.6.9 pinecone-text pinecone-client langchain-anthropic "unstructured[local-inference, all-docs]" ollama langchain_together==0.1.3 \
- langchain_cohere==0.1.5 deepgram-sdk==3.2.7 langchain-huggingface psutil clarifai==10.0.1 flask==2.3.3 twilio==8.5.0 fastapi-poe==0.0.16 discord==2.3.2 \
- slack-sdk==3.21.3 huggingface_hub==0.23.0 gitpython==3.1.38 yt_dlp==2023.11.14 PyGithub==1.59.1 feedparser==6.0.10 newspaper3k==0.2.8 listparser==0.19 \
- modal==0.56.4329 dropbox==11.36.2 boto3==1.34.20 youtube-transcript-api==0.6.1 pytube==15.0.0 beautifulsoup4==4.12.3
-
-install_es:
- poetry install --extras elasticsearch
-
-install_opensearch:
- poetry install --extras opensearch
-
-install_milvus:
- poetry install --extras milvus
-
-shell:
- poetry shell
-
-py_shell:
- poetry run python
-
-format:
- $(PYTHON) -m black .
- $(PYTHON) -m isort .
-
-clean:
- rm -rf dist build *.egg-info
-
-lint:
- poetry run ruff .
-
-build:
- poetry build
-
-publish:
- poetry publish
-
-# for example: make test file=tests/test_factory.py
-test:
- poetry run pytest $(file)
-
-coverage:
- poetry run pytest --cov=$(PROJECT_NAME) --cov-report=xml
diff --git a/embedchain/README.md b/embedchain/README.md
deleted file mode 100644
index 8b072ed87..000000000
--- a/embedchain/README.md
+++ /dev/null
@@ -1,125 +0,0 @@
-
-
-
-
-## What is Embedchain?
-
-Embedchain is an Open Source Framework for personalizing LLM responses. It makes it easy to create and deploy personalized AI apps. At its core, Embedchain follows the design principle of being *"Conventional but Configurable"* to serve both software engineers and machine learning engineers.
-
-Embedchain streamlines the creation of personalized LLM applications, offering a seamless process for managing various types of unstructured data. It efficiently segments data into manageable chunks, generates relevant embeddings, and stores them in a vector database for optimized retrieval. With a suite of diverse APIs, it enables users to extract contextual information, find precise answers, or engage in interactive chat conversations, all tailored to their own data.
-
-## 🔧 Quick install
-
-### Python API
-
-```bash
-pip install embedchain
-```
-
-## ✨ Live demo
-
-Checkout the [Chat with PDF](https://embedchain.ai/demo/chat-pdf) live demo we created using Embedchain. You can find the source code [here](https://github.com/mem0ai/mem0/tree/main/embedchain/examples/chat-pdf).
-
-## 🔍 Usage
-
-
-
-
-
-
-For example, you can create an Elon Musk bot using the following code:
-
-```python
-import os
-from embedchain import App
-
-# Create a bot instance
-os.environ["OPENAI_API_KEY"] = ""
-app = App()
-
-# Embed online resources
-app.add("https://en.wikipedia.org/wiki/Elon_Musk")
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# Query the app
-app.query("How many companies does Elon Musk run and name those?")
-# Answer: Elon Musk currently runs several companies. As of my knowledge, he is the CEO and lead designer of SpaceX, the CEO and product architect of Tesla, Inc., the CEO and founder of Neuralink, and the CEO and founder of The Boring Company. However, please note that this information may change over time, so it's always good to verify the latest updates.
-```
-
-You can also try it in your browser with Google Colab:
-
-[](https://colab.research.google.com/drive/17ON1LPonnXAtLaZEebnOktstB_1cJJmh?usp=sharing)
-
-## 📖 Documentation
-Comprehensive guides and API documentation are available to help you get the most out of Embedchain:
-
-- [Introduction](https://docs.embedchain.ai/get-started/introduction#what-is-embedchain)
-- [Getting Started](https://docs.embedchain.ai/get-started/quickstart)
-- [Examples](https://docs.embedchain.ai/examples)
-- [Supported data types](https://docs.embedchain.ai/components/data-sources/overview)
-
-## 🔗 Join the Community
-
-* Connect with fellow developers by joining our [Slack Community](https://embedchain.ai/slack) or [Discord Community](https://embedchain.ai/discord).
-
-* Dive into [GitHub Discussions](https://github.com/embedchain/embedchain/discussions), ask questions, or share your experiences.
-
-## 🤝 Schedule a 1-on-1 Session
-
-Book a [1-on-1 Session](https://cal.com/taranjeetio/ec) with the founders, to discuss any issues, provide feedback, or explore how we can improve Embedchain for you.
-
-## 🌐 Contributing
-
-Contributions are welcome! Please check out the issues on the repository, and feel free to open a pull request.
-For more information, please see the [contributing guidelines](CONTRIBUTING.md).
-
-For more reference, please go through [Development Guide](https://docs.embedchain.ai/contribution/dev) and [Documentation Guide](https://docs.embedchain.ai/contribution/docs).
-
-
-
-
-
-## Anonymous Telemetry
-
-We collect anonymous usage metrics to enhance our package's quality and user experience. This includes data like feature usage frequency and system info, but never personal details. The data helps us prioritize improvements and ensure compatibility. If you wish to opt-out, set the environment variable `EC_TELEMETRY=false`. We prioritize data security and don't share this data externally.
-
-## Citation
-
-If you utilize this repository, please consider citing it with:
-
-```
-@misc{embedchain,
- author = {Taranjeet Singh, Deshraj Yadav},
- title = {Embedchain: The Open Source RAG Framework},
- year = {2023},
- publisher = {GitHub},
- journal = {GitHub repository},
- howpublished = {\url{https://github.com/embedchain/embedchain}},
-}
-```
diff --git a/embedchain/configs/anthropic.yaml b/embedchain/configs/anthropic.yaml
deleted file mode 100644
index 395125f99..000000000
--- a/embedchain/configs/anthropic.yaml
+++ /dev/null
@@ -1,8 +0,0 @@
-llm:
- provider: anthropic
- config:
- model: 'claude-instant-1'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
diff --git a/embedchain/configs/aws_bedrock.yaml b/embedchain/configs/aws_bedrock.yaml
deleted file mode 100644
index 824ab0fff..000000000
--- a/embedchain/configs/aws_bedrock.yaml
+++ /dev/null
@@ -1,15 +0,0 @@
-llm:
- provider: aws_bedrock
- config:
- model: amazon.titan-text-express-v1
- deployment_name: your_llm_deployment_name
- temperature: 0.5
- max_tokens: 8192
- top_p: 1
- stream: false
-
-embedder::
- provider: aws_bedrock
- config:
- model: amazon.titan-embed-text-v2:0
- deployment_name: you_embedding_model_deployment_name
\ No newline at end of file
diff --git a/embedchain/configs/azure_openai.yaml b/embedchain/configs/azure_openai.yaml
deleted file mode 100644
index 50eaff0c8..000000000
--- a/embedchain/configs/azure_openai.yaml
+++ /dev/null
@@ -1,19 +0,0 @@
-app:
- config:
- id: azure-openai-app
-
-llm:
- provider: azure_openai
- config:
- model: gpt-35-turbo
- deployment_name: your_llm_deployment_name
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-embedder:
- provider: azure_openai
- config:
- model: text-embedding-ada-002
- deployment_name: you_embedding_model_deployment_name
diff --git a/embedchain/configs/chroma.yaml b/embedchain/configs/chroma.yaml
deleted file mode 100644
index 142eb05fc..000000000
--- a/embedchain/configs/chroma.yaml
+++ /dev/null
@@ -1,24 +0,0 @@
-app:
- config:
- id: 'my-app'
-
-llm:
- provider: openai
- config:
- model: 'gpt-4o-mini'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-vectordb:
- provider: chroma
- config:
- collection_name: 'my-app'
- dir: db
- allow_reset: true
-
-embedder:
- provider: openai
- config:
- model: 'text-embedding-ada-002'
diff --git a/embedchain/configs/chunker.yaml b/embedchain/configs/chunker.yaml
deleted file mode 100644
index 63cf3f82c..000000000
--- a/embedchain/configs/chunker.yaml
+++ /dev/null
@@ -1,4 +0,0 @@
-chunker:
- chunk_size: 100
- chunk_overlap: 20
- length_function: 'len'
diff --git a/embedchain/configs/clarifai.yaml b/embedchain/configs/clarifai.yaml
deleted file mode 100644
index 0c52ba007..000000000
--- a/embedchain/configs/clarifai.yaml
+++ /dev/null
@@ -1,12 +0,0 @@
-llm:
- provider: clarifai
- config:
- model: "https://clarifai.com/mistralai/completion/models/mistral-7B-Instruct"
- model_kwargs:
- temperature: 0.5
- max_tokens: 1000
-
-embedder:
- provider: clarifai
- config:
- model: "https://clarifai.com/clarifai/main/models/BAAI-bge-base-en-v15"
diff --git a/embedchain/configs/cohere.yaml b/embedchain/configs/cohere.yaml
deleted file mode 100644
index 0edd4e8fd..000000000
--- a/embedchain/configs/cohere.yaml
+++ /dev/null
@@ -1,7 +0,0 @@
-llm:
- provider: cohere
- config:
- model: large
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
diff --git a/embedchain/configs/full-stack.yaml b/embedchain/configs/full-stack.yaml
deleted file mode 100644
index 978722eac..000000000
--- a/embedchain/configs/full-stack.yaml
+++ /dev/null
@@ -1,40 +0,0 @@
-app:
- config:
- id: 'full-stack-app'
-
-chunker:
- chunk_size: 100
- chunk_overlap: 20
- length_function: 'len'
-
-llm:
- provider: openai
- config:
- model: 'gpt-4o-mini'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
- prompt: |
- Use the following pieces of context to answer the query at the end.
- If you don't know the answer, just say that you don't know, don't try to make up an answer.
-
- $context
-
- Query: $query
-
- Helpful Answer:
- system_prompt: |
- Act as William Shakespeare. Answer the following questions in the style of William Shakespeare.
-
-vectordb:
- provider: chroma
- config:
- collection_name: 'my-collection-name'
- dir: db
- allow_reset: true
-
-embedder:
- provider: openai
- config:
- model: 'text-embedding-ada-002'
diff --git a/embedchain/configs/google.yaml b/embedchain/configs/google.yaml
deleted file mode 100644
index 4f6a46553..000000000
--- a/embedchain/configs/google.yaml
+++ /dev/null
@@ -1,13 +0,0 @@
-llm:
- provider: google
- config:
- model: gemini-pro
- max_tokens: 1000
- temperature: 0.9
- top_p: 1.0
- stream: false
-
-embedder:
- provider: google
- config:
- model: models/embedding-001
diff --git a/embedchain/configs/gpt4.yaml b/embedchain/configs/gpt4.yaml
deleted file mode 100644
index e06c60de6..000000000
--- a/embedchain/configs/gpt4.yaml
+++ /dev/null
@@ -1,8 +0,0 @@
-llm:
- provider: openai
- config:
- model: 'gpt-4'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
\ No newline at end of file
diff --git a/embedchain/configs/gpt4all.yaml b/embedchain/configs/gpt4all.yaml
deleted file mode 100644
index 048239334..000000000
--- a/embedchain/configs/gpt4all.yaml
+++ /dev/null
@@ -1,11 +0,0 @@
-llm:
- provider: gpt4all
- config:
- model: 'orca-mini-3b-gguf2-q4_0.gguf'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-embedder:
- provider: gpt4all
diff --git a/embedchain/configs/huggingface.yaml b/embedchain/configs/huggingface.yaml
deleted file mode 100644
index 508c9d778..000000000
--- a/embedchain/configs/huggingface.yaml
+++ /dev/null
@@ -1,8 +0,0 @@
-llm:
- provider: huggingface
- config:
- model: 'google/flan-t5-xxl'
- temperature: 0.5
- max_tokens: 1000
- top_p: 0.5
- stream: false
diff --git a/embedchain/configs/jina.yaml b/embedchain/configs/jina.yaml
deleted file mode 100644
index 11627059b..000000000
--- a/embedchain/configs/jina.yaml
+++ /dev/null
@@ -1,7 +0,0 @@
-llm:
- provider: jina
- config:
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
diff --git a/embedchain/configs/llama2.yaml b/embedchain/configs/llama2.yaml
deleted file mode 100644
index 61b3b9253..000000000
--- a/embedchain/configs/llama2.yaml
+++ /dev/null
@@ -1,8 +0,0 @@
-llm:
- provider: llama2
- config:
- model: 'a16z-infra/llama13b-v2-chat:df7690f1994d94e96ad9d568eac121aecf50684a0b0963b25a41cc40061269e5'
- temperature: 0.5
- max_tokens: 1000
- top_p: 0.5
- stream: false
diff --git a/embedchain/configs/ollama.yaml b/embedchain/configs/ollama.yaml
deleted file mode 100644
index 7ec5def54..000000000
--- a/embedchain/configs/ollama.yaml
+++ /dev/null
@@ -1,14 +0,0 @@
-llm:
- provider: ollama
- config:
- model: 'llama2'
- temperature: 0.5
- top_p: 1
- stream: true
- base_url: http://localhost:11434
-
-embedder:
- provider: ollama
- config:
- model: 'mxbai-embed-large:latest'
- base_url: http://localhost:11434
diff --git a/embedchain/configs/opensearch.yaml b/embedchain/configs/opensearch.yaml
deleted file mode 100644
index 94a27b29f..000000000
--- a/embedchain/configs/opensearch.yaml
+++ /dev/null
@@ -1,33 +0,0 @@
-app:
- config:
- id: 'my-app'
- log_level: 'WARNING'
- collect_metrics: true
- collection_name: 'my-app'
-
-llm:
- provider: openai
- config:
- model: 'gpt-4o-mini'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-vectordb:
- provider: opensearch
- config:
- opensearch_url: 'https://localhost:9200'
- http_auth:
- - admin
- - admin
- vector_dimension: 1536
- collection_name: 'my-app'
- use_ssl: false
- verify_certs: false
-
-embedder:
- provider: openai
- config:
- model: 'text-embedding-ada-002'
- deployment_name: 'my-app'
diff --git a/embedchain/configs/opensource.yaml b/embedchain/configs/opensource.yaml
deleted file mode 100644
index e2d40c135..000000000
--- a/embedchain/configs/opensource.yaml
+++ /dev/null
@@ -1,25 +0,0 @@
-app:
- config:
- id: 'open-source-app'
- collect_metrics: false
-
-llm:
- provider: gpt4all
- config:
- model: 'orca-mini-3b-gguf2-q4_0.gguf'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-vectordb:
- provider: chroma
- config:
- collection_name: 'open-source-app'
- dir: db
- allow_reset: true
-
-embedder:
- provider: gpt4all
- config:
- deployment_name: 'test-deployment'
diff --git a/embedchain/configs/pinecone.yaml b/embedchain/configs/pinecone.yaml
deleted file mode 100644
index 24e33c11a..000000000
--- a/embedchain/configs/pinecone.yaml
+++ /dev/null
@@ -1,6 +0,0 @@
-vectordb:
- provider: pinecone
- config:
- metric: cosine
- vector_dimension: 1536
- collection_name: my-pinecone-index
diff --git a/embedchain/configs/pipeline.yaml b/embedchain/configs/pipeline.yaml
deleted file mode 100644
index e34866716..000000000
--- a/embedchain/configs/pipeline.yaml
+++ /dev/null
@@ -1,26 +0,0 @@
-pipeline:
- config:
- name: Example pipeline
- id: pipeline-1 # Make sure that id is different every time you create a new pipeline
-
-vectordb:
- provider: chroma
- config:
- collection_name: pipeline-1
- dir: db
- allow_reset: true
-
-llm:
- provider: gpt4all
- config:
- model: 'orca-mini-3b-gguf2-q4_0.gguf'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-embedding_model:
- provider: gpt4all
- config:
- model: 'all-MiniLM-L6-v2'
- deployment_name: null
diff --git a/embedchain/configs/together.yaml b/embedchain/configs/together.yaml
deleted file mode 100644
index b19bc07ff..000000000
--- a/embedchain/configs/together.yaml
+++ /dev/null
@@ -1,6 +0,0 @@
-llm:
- provider: together
- config:
- model: mistralai/Mixtral-8x7B-Instruct-v0.1
- temperature: 0.5
- max_tokens: 1000
diff --git a/embedchain/configs/vertexai.yaml b/embedchain/configs/vertexai.yaml
deleted file mode 100644
index f303654c0..000000000
--- a/embedchain/configs/vertexai.yaml
+++ /dev/null
@@ -1,6 +0,0 @@
-llm:
- provider: vertexai
- config:
- model: 'chat-bison'
- temperature: 0.5
- top_p: 0.5
diff --git a/embedchain/configs/vllm.yaml b/embedchain/configs/vllm.yaml
deleted file mode 100644
index 536a589a1..000000000
--- a/embedchain/configs/vllm.yaml
+++ /dev/null
@@ -1,14 +0,0 @@
-llm:
- provider: vllm
- config:
- model: 'meta-llama/Llama-2-70b-hf'
- temperature: 0.5
- top_p: 1
- top_k: 10
- stream: true
- trust_remote_code: true
-
-embedder:
- provider: huggingface
- config:
- model: 'BAAI/bge-small-en-v1.5'
diff --git a/embedchain/configs/weaviate.yaml b/embedchain/configs/weaviate.yaml
deleted file mode 100644
index a27623ab9..000000000
--- a/embedchain/configs/weaviate.yaml
+++ /dev/null
@@ -1,4 +0,0 @@
-vectordb:
- provider: weaviate
- config:
- collection_name: my_weaviate_index
diff --git a/embedchain/docs/Makefile b/embedchain/docs/Makefile
deleted file mode 100644
index 0db640d0e..000000000
--- a/embedchain/docs/Makefile
+++ /dev/null
@@ -1,10 +0,0 @@
-install:
- npm i -g mintlify
-
-run_local:
- mintlify dev
-
-troubleshoot:
- mintlify install
-
-.PHONY: install run_local troubleshoot
diff --git a/embedchain/docs/README.md b/embedchain/docs/README.md
deleted file mode 100644
index e322686dc..000000000
--- a/embedchain/docs/README.md
+++ /dev/null
@@ -1,25 +0,0 @@
-# Contributing to embedchain docs
-
-
-### 👩💻 Development
-
-Install the [Mintlify CLI](https://www.npmjs.com/package/mintlify) to preview the documentation changes locally. To install, use the following command
-
-```
-npm i -g mintlify
-```
-
-Run the following command at the root of your documentation (where mint.json is)
-
-```
-mintlify dev
-```
-
-### 😎 Publishing Changes
-
-Changes will be deployed to production automatically after your PR is merged to the main branch.
-
-#### Troubleshooting
-
-- Mintlify dev isn't running - Run `mintlify install` it'll re-install dependencies.
-- Page loads as a 404 - Make sure you are running in a folder with `mint.json`
diff --git a/embedchain/docs/_snippets/get-help.mdx b/embedchain/docs/_snippets/get-help.mdx
deleted file mode 100644
index 6f57e5ce5..000000000
--- a/embedchain/docs/_snippets/get-help.mdx
+++ /dev/null
@@ -1,11 +0,0 @@
-
-
- Schedule a call
-
-
- Join our slack community
-
-
- Join our discord community
-
-
diff --git a/embedchain/docs/_snippets/missing-data-source-tip.mdx b/embedchain/docs/_snippets/missing-data-source-tip.mdx
deleted file mode 100644
index b0e189553..000000000
--- a/embedchain/docs/_snippets/missing-data-source-tip.mdx
+++ /dev/null
@@ -1,19 +0,0 @@
-
If you can't find the specific data source, please feel free to request through one of the following channels and help us prioritize.
-
-
-
- Fill out this form
-
-
- Let us know on our slack community
-
-
- Let us know on discord community
-
-
- Open an issue on our GitHub
-
-
- Schedule a call with Embedchain founder
-
-
diff --git a/embedchain/docs/_snippets/missing-llm-tip.mdx b/embedchain/docs/_snippets/missing-llm-tip.mdx
deleted file mode 100644
index 7d2782d38..000000000
--- a/embedchain/docs/_snippets/missing-llm-tip.mdx
+++ /dev/null
@@ -1,16 +0,0 @@
-
If you can't find the specific LLM you need, no need to fret. We're continuously expanding our support for additional LLMs, and you can help us prioritize by opening an issue on our GitHub or simply reaching out to us on our Slack or Discord community.
-
-
-
- Let us know on our slack community
-
-
- Let us know on discord community
-
-
- Open an issue on our GitHub
-
-
- Schedule a call with Embedchain founder
-
-
diff --git a/embedchain/docs/_snippets/missing-vector-db-tip.mdx b/embedchain/docs/_snippets/missing-vector-db-tip.mdx
deleted file mode 100644
index 2edbbe4b0..000000000
--- a/embedchain/docs/_snippets/missing-vector-db-tip.mdx
+++ /dev/null
@@ -1,18 +0,0 @@
-
-
-
If you can't find specific feature or run into issues, please feel free to reach out through one of the following channels.
-
-
-
- Let us know on our slack community
-
-
- Let us know on discord community
-
-
- Open an issue on our GitHub
-
-
- Schedule a call with Embedchain founder
-
-
diff --git a/embedchain/docs/api-reference/advanced/configuration.mdx b/embedchain/docs/api-reference/advanced/configuration.mdx
deleted file mode 100644
index 568ea567e..000000000
--- a/embedchain/docs/api-reference/advanced/configuration.mdx
+++ /dev/null
@@ -1,273 +0,0 @@
----
-title: 'Custom configurations'
----
-
-Embedchain offers several configuration options for your LLM, vector database, and embedding model. All of these configuration options are optional and have sane defaults.
-
-You can configure different components of your app (`llm`, `embedding model`, or `vector database`) through a simple yaml configuration that Embedchain offers. Here is a generic full-stack example of the yaml config:
-
-
-
-Embedchain applications are configurable using YAML file, JSON file or by directly passing the config dictionary. Checkout the [docs here](/api-reference/app/overview#usage) on how to use other formats.
-
-
-
-```yaml config.yaml
-app:
- config:
- name: 'full-stack-app'
-
-llm:
- provider: openai
- config:
- model: 'gpt-4o-mini'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
- api_key: sk-xxx
- model_kwargs:
- response_format:
- type: json_object
- api_version: 2024-02-01
- http_client_proxies: http://testproxy.mem0.net:8000
- prompt: |
- Use the following pieces of context to answer the query at the end.
- If you don't know the answer, just say that you don't know, don't try to make up an answer.
-
- $context
-
- Query: $query
-
- Helpful Answer:
- system_prompt: |
- Act as William Shakespeare. Answer the following questions in the style of William Shakespeare.
-
-vectordb:
- provider: chroma
- config:
- collection_name: 'full-stack-app'
- dir: db
- allow_reset: true
-
-embedder:
- provider: openai
- config:
- model: 'text-embedding-ada-002'
- api_key: sk-xxx
- http_client_proxies: http://testproxy.mem0.net:8000
-
-chunker:
- chunk_size: 2000
- chunk_overlap: 100
- length_function: 'len'
- min_chunk_size: 0
-
-cache:
- similarity_evaluation:
- strategy: distance
- max_distance: 1.0
- config:
- similarity_threshold: 0.8
- auto_flush: 50
-
-memory:
- top_k: 10
-```
-
-```json config.json
-{
- "app": {
- "config": {
- "name": "full-stack-app"
- }
- },
- "llm": {
- "provider": "openai",
- "config": {
- "model": "gpt-4o-mini",
- "temperature": 0.5,
- "max_tokens": 1000,
- "top_p": 1,
- "stream": false,
- "prompt": "Use the following pieces of context to answer the query at the end.\nIf you don't know the answer, just say that you don't know, don't try to make up an answer.\n$context\n\nQuery: $query\n\nHelpful Answer:",
- "system_prompt": "Act as William Shakespeare. Answer the following questions in the style of William Shakespeare.",
- "api_key": "sk-xxx",
- "model_kwargs": {"response_format": {"type": "json_object"}},
- "api_version": "2024-02-01",
- "http_client_proxies": "http://testproxy.mem0.net:8000"
- }
- },
- "vectordb": {
- "provider": "chroma",
- "config": {
- "collection_name": "full-stack-app",
- "dir": "db",
- "allow_reset": true
- }
- },
- "embedder": {
- "provider": "openai",
- "config": {
- "model": "text-embedding-ada-002",
- "api_key": "sk-xxx",
- "http_client_proxies": "http://testproxy.mem0.net:8000"
- }
- },
- "chunker": {
- "chunk_size": 2000,
- "chunk_overlap": 100,
- "length_function": "len",
- "min_chunk_size": 0
- },
- "cache": {
- "similarity_evaluation": {
- "strategy": "distance",
- "max_distance": 1.0
- },
- "config": {
- "similarity_threshold": 0.8,
- "auto_flush": 50
- }
- },
- "memory": {
- "top_k": 10
- }
-}
-```
-
-```python config.py
-config = {
- 'app': {
- 'config': {
- 'name': 'full-stack-app'
- }
- },
- 'llm': {
- 'provider': 'openai',
- 'config': {
- 'model': 'gpt-4o-mini',
- 'temperature': 0.5,
- 'max_tokens': 1000,
- 'top_p': 1,
- 'stream': False,
- 'prompt': (
- "Use the following pieces of context to answer the query at the end.\n"
- "If you don't know the answer, just say that you don't know, don't try to make up an answer.\n"
- "$context\n\nQuery: $query\n\nHelpful Answer:"
- ),
- 'system_prompt': (
- "Act as William Shakespeare. Answer the following questions in the style of William Shakespeare."
- ),
- 'api_key': 'sk-xxx',
- "model_kwargs": {"response_format": {"type": "json_object"}},
- "http_client_proxies": "http://testproxy.mem0.net:8000",
- }
- },
- 'vectordb': {
- 'provider': 'chroma',
- 'config': {
- 'collection_name': 'full-stack-app',
- 'dir': 'db',
- 'allow_reset': True
- }
- },
- 'embedder': {
- 'provider': 'openai',
- 'config': {
- 'model': 'text-embedding-ada-002',
- 'api_key': 'sk-xxx',
- "http_client_proxies": "http://testproxy.mem0.net:8000",
- }
- },
- 'chunker': {
- 'chunk_size': 2000,
- 'chunk_overlap': 100,
- 'length_function': 'len',
- 'min_chunk_size': 0
- },
- 'cache': {
- 'similarity_evaluation': {
- 'strategy': 'distance',
- 'max_distance': 1.0,
- },
- 'config': {
- 'similarity_threshold': 0.8,
- 'auto_flush': 50,
- },
- },
- 'memory': {
- 'top_k': 10,
- },
-}
-```
-
-
-Alright, let's dive into what each key means in the yaml config above:
-
-1. `app` Section:
- - `config`:
- - `name` (String): The name of your full-stack application.
- - `id` (String): The id of your full-stack application.
- Only use this to reload already created apps. We recommend users not to create their own ids.
- - `collect_metrics` (Boolean): Indicates whether metrics should be collected for the app, defaults to `True`
- - `log_level` (String): The log level for the app, defaults to `WARNING`
-2. `llm` Section:
- - `provider` (String): The provider for the language model, which is set to 'openai'. You can find the full list of llm providers in [our docs](/components/llms).
- - `config`:
- - `model` (String): The specific model being used, 'gpt-4o-mini'.
- - `temperature` (Float): Controls the randomness of the model's output. A higher value (closer to 1) makes the output more random.
- - `max_tokens` (Integer): Controls how many tokens are used in the response.
- - `top_p` (Float): Controls the diversity of word selection. A higher value (closer to 1) makes word selection more diverse.
- - `stream` (Boolean): Controls if the response is streamed back to the user (set to false).
- - `online` (Boolean): Controls whether to use internet to get more context for answering query (set to false).
- - `token_usage` (Boolean): Controls whether to use token usage for the querying models (set to false).
- - `prompt` (String): A prompt for the model to follow when generating responses, requires `$context` and `$query` variables.
- - `system_prompt` (String): A system prompt for the model to follow when generating responses, in this case, it's set to the style of William Shakespeare.
- - `number_documents` (Integer): Number of documents to pull from the vectordb as context, defaults to 1
- - `api_key` (String): The API key for the language model.
- - `model_kwargs` (Dict): Keyword arguments to pass to the language model. Used for `aws_bedrock` provider, since it requires different arguments for each model.
- - `http_client_proxies` (Dict | String): The proxy server settings used to create `self.http_client` using `httpx.Client(proxies=http_client_proxies)`
- - `http_async_client_proxies` (Dict | String): The proxy server settings for async calls used to create `self.http_async_client` using `httpx.AsyncClient(proxies=http_async_client_proxies)`
-3. `vectordb` Section:
- - `provider` (String): The provider for the vector database, set to 'chroma'. You can find the full list of vector database providers in [our docs](/components/vector-databases).
- - `config`:
- - `collection_name` (String): The initial collection name for the vectordb, set to 'full-stack-app'.
- - `dir` (String): The directory for the local database, set to 'db'.
- - `allow_reset` (Boolean): Indicates whether resetting the vectordb is allowed, set to true.
- - `batch_size` (Integer): The batch size for docs insertion in vectordb, defaults to `100`
- We recommend you to checkout vectordb specific config [here](https://docs.embedchain.ai/components/vector-databases)
-4. `embedder` Section:
- - `provider` (String): The provider for the embedder, set to 'openai'. You can find the full list of embedding model providers in [our docs](/components/embedding-models).
- - `config`:
- - `model` (String): The specific model used for text embedding, 'text-embedding-ada-002'.
- - `vector_dimension` (Integer): The vector dimension of the embedding model. [Defaults](https://github.com/embedchain/embedchain/blob/main/embedchain/models/vector_dimensions.py)
- - `api_key` (String): The API key for the embedding model.
- - `endpoint` (String): The endpoint for the HuggingFace embedding model.
- - `deployment_name` (String): The deployment name for the embedding model.
- - `title` (String): The title for the embedding model for Google Embedder.
- - `task_type` (String): The task type for the embedding model for Google Embedder.
- - `model_kwargs` (Dict): Used to pass extra arguments to embedders.
- - `http_client_proxies` (Dict | String): The proxy server settings used to create `self.http_client` using `httpx.Client(proxies=http_client_proxies)`
- - `http_async_client_proxies` (Dict | String): The proxy server settings for async calls used to create `self.http_async_client` using `httpx.AsyncClient(proxies=http_async_client_proxies)`
-5. `chunker` Section:
- - `chunk_size` (Integer): The size of each chunk of text that is sent to the language model.
- - `chunk_overlap` (Integer): The amount of overlap between each chunk of text.
- - `length_function` (String): The function used to calculate the length of each chunk of text. In this case, it's set to 'len'. You can also use any function import directly as a string here.
- - `min_chunk_size` (Integer): The minimum size of each chunk of text that is sent to the language model. Must be less than `chunk_size`, and greater than `chunk_overlap`.
-6. `cache` Section: (Optional)
- - `similarity_evaluation` (Optional): The config for similarity evaluation strategy. If not provided, the default `distance` based similarity evaluation strategy is used.
- - `strategy` (String): The strategy to use for similarity evaluation. Currently, only `distance` and `exact` based similarity evaluation is supported. Defaults to `distance`.
- - `max_distance` (Float): The bound of maximum distance. Defaults to `1.0`.
- - `positive` (Boolean): If the larger distance indicates more similar of two entities, set it `True`, otherwise `False`. Defaults to `False`.
- - `config` (Optional): The config for initializing the cache. If not provided, sensible default values are used as mentioned below.
- - `similarity_threshold` (Float): The threshold for similarity evaluation. Defaults to `0.8`.
- - `auto_flush` (Integer): The number of queries after which the cache is flushed. Defaults to `20`.
-7. `memory` Section: (Optional)
- - `top_k` (Integer): The number of top-k results to return. Defaults to `10`.
-
- If you provide a cache section, the app will automatically configure and use a cache to store the results of the language model. This is useful if you want to speed up the response time and save inference cost of your app.
-
-If you have questions about the configuration above, please feel free to reach out to us using one of the following methods:
-
-
\ No newline at end of file
diff --git a/embedchain/docs/api-reference/app/add.mdx b/embedchain/docs/api-reference/app/add.mdx
deleted file mode 100644
index 21b24de34..000000000
--- a/embedchain/docs/api-reference/app/add.mdx
+++ /dev/null
@@ -1,47 +0,0 @@
----
-title: '📊 add'
----
-
-`add()` method is used to load the data sources from different data sources to a RAG pipeline. You can find the signature below:
-
-### Parameters
-
-
- The data to embed, can be a URL, local file or raw content, depending on the data type.. You can find the full list of supported data sources [here](/components/data-sources/overview).
-
-
- Type of data source. It can be automatically detected but user can force what data type to load as.
-
-
- Any metadata that you want to store with the data source. Metadata is generally really useful for doing metadata filtering on top of semantic search to yield faster search and better results.
-
-
- This parameter instructs Embedchain to retrieve all the context and information from the specified link, as well as from any reference links on the page.
-
-
-## Usage
-
-### Load data from webpage
-
-```python Code example
-from embedchain import App
-
-app = App()
-app.add("https://www.forbes.com/profile/elon-musk")
-# Inserting batches in chromadb: 100%|███████████████| 1/1 [00:00<00:00, 1.19it/s]
-# Successfully saved https://www.forbes.com/profile/elon-musk (DataType.WEB_PAGE). New chunks count: 4
-```
-
-### Load data from sitemap
-
-```python Code example
-from embedchain import App
-
-app = App()
-app.add("https://python.langchain.com/sitemap.xml", data_type="sitemap")
-# Loading pages: 100%|█████████████| 1108/1108 [00:47<00:00, 23.17it/s]
-# Inserting batches in chromadb: 100%|█████████| 111/111 [04:41<00:00, 2.54s/it]
-# Successfully saved https://python.langchain.com/sitemap.xml (DataType.SITEMAP). New chunks count: 11024
-```
-
-You can find complete list of supported data sources [here](/components/data-sources/overview).
diff --git a/embedchain/docs/api-reference/app/chat.mdx b/embedchain/docs/api-reference/app/chat.mdx
deleted file mode 100644
index f12b09793..000000000
--- a/embedchain/docs/api-reference/app/chat.mdx
+++ /dev/null
@@ -1,175 +0,0 @@
----
-title: '💬 chat'
----
-
-`chat()` method allows you to chat over your data sources using a user-friendly chat API. You can find the signature below:
-
-### Parameters
-
-
- Question to ask
-
-
- Configure different llm settings such as prompt, temprature, number_documents etc.
-
-
- The purpose is to test the prompt structure without actually running LLM inference. Defaults to `False`
-
-
- A dictionary of key-value pairs to filter the chunks from the vector database. Defaults to `None`
-
-
- Session ID of the chat. This can be used to maintain chat history of different user sessions. Default value: `default`
-
-
- Return citations along with the LLM answer. Defaults to `False`
-
-
-### Returns
-
-
- If `citations=False`, return a stringified answer to the question asked.
- If `citations=True`, returns a tuple with answer and citations respectively.
-
-
-## Usage
-
-### With citations
-
-If you want to get the answer to question and return both answer and citations, use the following code snippet:
-
-```python With Citations
-from embedchain import App
-
-# Initialize app
-app = App()
-
-# Add data source
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# Get relevant answer for your query
-answer, sources = app.chat("What is the net worth of Elon?", citations=True)
-print(answer)
-# Answer: The net worth of Elon Musk is $221.9 billion.
-
-print(sources)
-# [
-# (
-# 'Elon Musk PROFILEElon MuskCEO, Tesla$247.1B$2.3B (0.96%)Real Time Net Worthas of 12/7/23 ...',
-# {
-# 'url': 'https://www.forbes.com/profile/elon-musk',
-# 'score': 0.89,
-# ...
-# }
-# ),
-# (
-# '74% of the company, which is now called X.Wealth HistoryHOVER TO REVEAL NET WORTH BY YEARForbes ...',
-# {
-# 'url': 'https://www.forbes.com/profile/elon-musk',
-# 'score': 0.81,
-# ...
-# }
-# ),
-# (
-# 'founded in 2002, is worth nearly $150 billion after a $750 million tender offer in June 2023 ...',
-# {
-# 'url': 'https://www.forbes.com/profile/elon-musk',
-# 'score': 0.73,
-# ...
-# }
-# )
-# ]
-```
-
-
-When `citations=True`, note that the returned `sources` are a list of tuples where each tuple has two elements (in the following order):
-1. source chunk
-2. dictionary with metadata about the source chunk
- - `url`: url of the source
- - `doc_id`: document id (used for book keeping purposes)
- - `score`: score of the source chunk with respect to the question
- - other metadata you might have added at the time of adding the source
-
-
-
-### Without citations
-
-If you just want to return answers and don't want to return citations, you can use the following example:
-
-```python Without Citations
-from embedchain import App
-
-# Initialize app
-app = App()
-
-# Add data source
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# Chat on your data using `.chat()`
-answer = app.chat("What is the net worth of Elon?")
-print(answer)
-# Answer: The net worth of Elon Musk is $221.9 billion.
-```
-
-### With session id
-
-If you want to maintain chat sessions for different users, you can simply pass the `session_id` keyword argument. See the example below:
-
-```python With session id
-from embedchain import App
-
-app = App()
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# Chat on your data using `.chat()`
-app.chat("What is the net worth of Elon Musk?", session_id="user1")
-# 'The net worth of Elon Musk is $250.8 billion.'
-app.chat("What is the net worth of Bill Gates?", session_id="user2")
-# "I don't know the current net worth of Bill Gates."
-app.chat("What was my last question", session_id="user1")
-# 'Your last question was "What is the net worth of Elon Musk?"'
-```
-
-### With custom context window
-
-If you want to customize the context window that you want to use during chat (default context window is 3 document chunks), you can do using the following code snippet:
-
-```python with custom chunks size
-from embedchain import App
-from embedchain.config import BaseLlmConfig
-
-app = App()
-app.add("https://www.forbes.com/profile/elon-musk")
-
-query_config = BaseLlmConfig(number_documents=5)
-app.chat("What is the net worth of Elon Musk?", config=query_config)
-```
-
-### With Mem0 to store chat history
-
-Mem0 is a cutting-edge long-term memory for LLMs to enable personalization for the GenAI stack. It enables LLMs to remember past interactions and provide more personalized responses.
-
-In order to use Mem0 to enable memory for personalization in your apps:
-- Install the [`mem0`](https://docs.mem0.ai/) package using `pip install mem0ai`.
-- Prepare config for `memory`, refer [Configurations](docs/api-reference/advanced/configuration.mdx).
-
-```python with mem0
-from embedchain import App
-
-config = {
- "memory": {
- "top_k": 5
- }
-}
-
-app = App.from_config(config=config)
-app.add("https://www.forbes.com/profile/elon-musk")
-
-app.chat("What is the net worth of Elon Musk?")
-```
-
-## How Mem0 works:
-- Mem0 saves context derived from each user question into its memory.
-- When a user poses a new question, Mem0 retrieves relevant previous memories.
-- The `top_k` parameter in the memory configuration specifies the number of top memories to consider during retrieval.
-- Mem0 generates the final response by integrating the user's question, context from the data source, and the relevant memories.
diff --git a/embedchain/docs/api-reference/app/delete.mdx b/embedchain/docs/api-reference/app/delete.mdx
deleted file mode 100644
index d1f2ceda4..000000000
--- a/embedchain/docs/api-reference/app/delete.mdx
+++ /dev/null
@@ -1,48 +0,0 @@
----
-title: 🗑 delete
----
-
-## Delete Document
-
-`delete()` method allows you to delete a document previously added to the app.
-
-### Usage
-
-```python
-from embedchain import App
-
-app = App()
-
-forbes_doc_id = app.add("https://www.forbes.com/profile/elon-musk")
-wiki_doc_id = app.add("https://en.wikipedia.org/wiki/Elon_Musk")
-
-app.delete(forbes_doc_id) # deletes the forbes document
-```
-
-
- If you do not have the document id, you can use `app.db.get()` method to get the document and extract the `hash` key from `metadatas` dictionary object, which serves as the document id.
-
-
-
-## Delete Chat Session History
-
-`delete_session_chat_history()` method allows you to delete all previous messages in a chat history.
-
-### Usage
-
-```python
-from embedchain import App
-
-app = App()
-
-app.add("https://www.forbes.com/profile/elon-musk")
-
-app.chat("What is the net worth of Elon Musk?")
-
-app.delete_session_chat_history()
-```
-
-
- `delete_session_chat_history(session_id="session_1")` method also accepts `session_id` optional param for deleting chat history of a specific session.
- It assumes the default session if no `session_id` is provided.
-
\ No newline at end of file
diff --git a/embedchain/docs/api-reference/app/deploy.mdx b/embedchain/docs/api-reference/app/deploy.mdx
deleted file mode 100644
index 7cb8ff5e8..000000000
--- a/embedchain/docs/api-reference/app/deploy.mdx
+++ /dev/null
@@ -1,5 +0,0 @@
----
-title: 🚀 deploy
----
-
-The `deploy()` method is currently available on an invitation-only basis. To request access, please submit your information via the provided [Google Form](https://forms.gle/vigN11h7b4Ywat668). We will review your request and respond promptly.
diff --git a/embedchain/docs/api-reference/app/evaluate.mdx b/embedchain/docs/api-reference/app/evaluate.mdx
deleted file mode 100644
index 64cb612ca..000000000
--- a/embedchain/docs/api-reference/app/evaluate.mdx
+++ /dev/null
@@ -1,41 +0,0 @@
----
-title: '📝 evaluate'
----
-
-`evaluate()` method is used to evaluate the performance of a RAG app. You can find the signature below:
-
-### Parameters
-
-
- A question or a list of questions to evaluate your app on.
-
-
- The metrics to evaluate your app on. Defaults to all metrics: `["context_relevancy", "answer_relevancy", "groundedness"]`
-
-
- Specify the number of threads to use for parallel processing.
-
-
-### Returns
-
-
- Returns the metrics you have chosen to evaluate your app on as a dictionary.
-
-
-## Usage
-
-```python
-from embedchain import App
-
-app = App()
-
-# add data source
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# run evaluation
-app.evaluate("what is the net worth of Elon Musk?")
-# {'answer_relevancy': 0.958019958036268, 'context_relevancy': 0.12903225806451613}
-
-# or
-# app.evaluate(["what is the net worth of Elon Musk?", "which companies does Elon Musk own?"])
-```
diff --git a/embedchain/docs/api-reference/app/get.mdx b/embedchain/docs/api-reference/app/get.mdx
deleted file mode 100644
index 252c78508..000000000
--- a/embedchain/docs/api-reference/app/get.mdx
+++ /dev/null
@@ -1,33 +0,0 @@
----
-title: 📄 get
----
-
-## Get data sources
-
-`get_data_sources()` returns a list of all the data sources added in the app.
-
-
-### Usage
-
-```python
-from embedchain import App
-
-app = App()
-
-app.add("https://www.forbes.com/profile/elon-musk")
-app.add("https://en.wikipedia.org/wiki/Elon_Musk")
-
-data_sources = app.get_data_sources()
-# [
-# {
-# 'data_type': 'web_page',
-# 'data_value': 'https://en.wikipedia.org/wiki/Elon_Musk',
-# 'metadata': 'null'
-# },
-# {
-# 'data_type': 'web_page',
-# 'data_value': 'https://www.forbes.com/profile/elon-musk',
-# 'metadata': 'null'
-# }
-# ]
-```
\ No newline at end of file
diff --git a/embedchain/docs/api-reference/app/overview.mdx b/embedchain/docs/api-reference/app/overview.mdx
deleted file mode 100644
index 8c369cbf8..000000000
--- a/embedchain/docs/api-reference/app/overview.mdx
+++ /dev/null
@@ -1,130 +0,0 @@
----
-title: "App"
----
-
-Create a RAG app object on Embedchain. This is the main entrypoint for a developer to interact with Embedchain APIs. An app configures the llm, vector database, embedding model, and retrieval strategy of your choice.
-
-### Attributes
-
-
- App ID
-
-
- Name of the app
-
-
- Configuration of the app
-
-
- Configured LLM for the RAG app
-
-
- Configured vector database for the RAG app
-
-
- Configured embedding model for the RAG app
-
-
- Chunker configuration
-
-
- Client object (used to deploy an app to Embedchain platform)
-
-
- Logger object
-
-
-## Usage
-
-You can create an app instance using the following methods:
-
-### Default setting
-
-```python Code Example
-from embedchain import App
-app = App()
-```
-
-
-### Python Dict
-
-```python Code Example
-from embedchain import App
-
-config_dict = {
- 'llm': {
- 'provider': 'gpt4all',
- 'config': {
- 'model': 'orca-mini-3b-gguf2-q4_0.gguf',
- 'temperature': 0.5,
- 'max_tokens': 1000,
- 'top_p': 1,
- 'stream': False
- }
- },
- 'embedder': {
- 'provider': 'gpt4all'
- }
-}
-
-# load llm configuration from config dict
-app = App.from_config(config=config_dict)
-```
-
-### YAML Config
-
-
-
-```python main.py
-from embedchain import App
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: gpt4all
- config:
- model: 'orca-mini-3b-gguf2-q4_0.gguf'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-embedder:
- provider: gpt4all
-```
-
-
-
-### JSON Config
-
-
-
-```python main.py
-from embedchain import App
-
-# load llm configuration from config.json file
-app = App.from_config(config_path="config.json")
-```
-
-```json config.json
-{
- "llm": {
- "provider": "gpt4all",
- "config": {
- "model": "orca-mini-3b-gguf2-q4_0.gguf",
- "temperature": 0.5,
- "max_tokens": 1000,
- "top_p": 1,
- "stream": false
- }
- },
- "embedder": {
- "provider": "gpt4all"
- }
-}
-```
-
-
diff --git a/embedchain/docs/api-reference/app/query.mdx b/embedchain/docs/api-reference/app/query.mdx
deleted file mode 100644
index f1d94aa8f..000000000
--- a/embedchain/docs/api-reference/app/query.mdx
+++ /dev/null
@@ -1,109 +0,0 @@
----
-title: '❓ query'
----
-
-`.query()` method empowers developers to ask questions and receive relevant answers through a user-friendly query API. Function signature is given below:
-
-### Parameters
-
-
- Question to ask
-
-
- Configure different llm settings such as prompt, temprature, number_documents etc.
-
-
- The purpose is to test the prompt structure without actually running LLM inference. Defaults to `False`
-
-
- A dictionary of key-value pairs to filter the chunks from the vector database. Defaults to `None`
-
-
- Return citations along with the LLM answer. Defaults to `False`
-
-
-### Returns
-
-
- If `citations=False`, return a stringified answer to the question asked.
- If `citations=True`, returns a tuple with answer and citations respectively.
-
-
-## Usage
-
-### With citations
-
-If you want to get the answer to question and return both answer and citations, use the following code snippet:
-
-```python With Citations
-from embedchain import App
-
-# Initialize app
-app = App()
-
-# Add data source
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# Get relevant answer for your query
-answer, sources = app.query("What is the net worth of Elon?", citations=True)
-print(answer)
-# Answer: The net worth of Elon Musk is $221.9 billion.
-
-print(sources)
-# [
-# (
-# 'Elon Musk PROFILEElon MuskCEO, Tesla$247.1B$2.3B (0.96%)Real Time Net Worthas of 12/7/23 ...',
-# {
-# 'url': 'https://www.forbes.com/profile/elon-musk',
-# 'score': 0.89,
-# ...
-# }
-# ),
-# (
-# '74% of the company, which is now called X.Wealth HistoryHOVER TO REVEAL NET WORTH BY YEARForbes ...',
-# {
-# 'url': 'https://www.forbes.com/profile/elon-musk',
-# 'score': 0.81,
-# ...
-# }
-# ),
-# (
-# 'founded in 2002, is worth nearly $150 billion after a $750 million tender offer in June 2023 ...',
-# {
-# 'url': 'https://www.forbes.com/profile/elon-musk',
-# 'score': 0.73,
-# ...
-# }
-# )
-# ]
-```
-
-
-When `citations=True`, note that the returned `sources` are a list of tuples where each tuple has two elements (in the following order):
-1. source chunk
-2. dictionary with metadata about the source chunk
- - `url`: url of the source
- - `doc_id`: document id (used for book keeping purposes)
- - `score`: score of the source chunk with respect to the question
- - other metadata you might have added at the time of adding the source
-
-
-### Without citations
-
-If you just want to return answers and don't want to return citations, you can use the following example:
-
-```python Without Citations
-from embedchain import App
-
-# Initialize app
-app = App()
-
-# Add data source
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# Get relevant answer for your query
-answer = app.query("What is the net worth of Elon?")
-print(answer)
-# Answer: The net worth of Elon Musk is $221.9 billion.
-```
-
diff --git a/embedchain/docs/api-reference/app/reset.mdx b/embedchain/docs/api-reference/app/reset.mdx
deleted file mode 100644
index 07e136d86..000000000
--- a/embedchain/docs/api-reference/app/reset.mdx
+++ /dev/null
@@ -1,17 +0,0 @@
----
-title: 🔄 reset
----
-
-`reset()` method allows you to wipe the data from your RAG application and start from scratch.
-
-## Usage
-
-```python
-from embedchain import App
-
-app = App()
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# Reset the app
-app.reset()
-```
\ No newline at end of file
diff --git a/embedchain/docs/api-reference/app/search.mdx b/embedchain/docs/api-reference/app/search.mdx
deleted file mode 100644
index db4eee1b2..000000000
--- a/embedchain/docs/api-reference/app/search.mdx
+++ /dev/null
@@ -1,111 +0,0 @@
----
-title: '🔍 search'
----
-
-`.search()` enables you to uncover the most pertinent context by performing a semantic search across your data sources based on a given query. Refer to the function signature below:
-
-### Parameters
-
-
- Question
-
-
- Number of relevant documents to fetch. Defaults to `3`
-
-
- Key value pair for metadata filtering.
-
-
- Pass raw filter query based on your vector database.
- Currently, `raw_filter` param is only supported for Pinecone vector database.
-
-
-### Returns
-
-
- Return list of dictionaries that contain the relevant chunk and their source information.
-
-
-## Usage
-
-### Basic
-
-Refer to the following example on how to use the search api:
-
-```python Code example
-from embedchain import App
-
-app = App()
-app.add("https://www.forbes.com/profile/elon-musk")
-
-context = app.search("What is the net worth of Elon?", num_documents=2)
-print(context)
-```
-
-### Advanced
-
-#### Metadata filtering using `where` params
-
-Here is an advanced example of `search()` API with metadata filtering on pinecone database:
-
-```python
-import os
-
-from embedchain import App
-
-os.environ["PINECONE_API_KEY"] = "xxx"
-
-config = {
- "vectordb": {
- "provider": "pinecone",
- "config": {
- "metric": "dotproduct",
- "vector_dimension": 1536,
- "index_name": "ec-test",
- "serverless_config": {"cloud": "aws", "region": "us-west-2"},
- },
- }
-}
-
-app = App.from_config(config=config)
-
-app.add("https://www.forbes.com/profile/bill-gates", metadata={"type": "forbes", "person": "gates"})
-app.add("https://en.wikipedia.org/wiki/Bill_Gates", metadata={"type": "wiki", "person": "gates"})
-
-results = app.search("What is the net worth of Bill Gates?", where={"person": "gates"})
-print("Num of search results: ", len(results))
-```
-
-#### Metadata filtering using `raw_filter` params
-
-Following is an example of metadata filtering by passing the raw filter query that pinecone vector database follows:
-
-```python
-import os
-
-from embedchain import App
-
-os.environ["PINECONE_API_KEY"] = "xxx"
-
-config = {
- "vectordb": {
- "provider": "pinecone",
- "config": {
- "metric": "dotproduct",
- "vector_dimension": 1536,
- "index_name": "ec-test",
- "serverless_config": {"cloud": "aws", "region": "us-west-2"},
- },
- }
-}
-
-app = App.from_config(config=config)
-
-app.add("https://www.forbes.com/profile/bill-gates", metadata={"year": 2022, "person": "gates"})
-app.add("https://en.wikipedia.org/wiki/Bill_Gates", metadata={"year": 2024, "person": "gates"})
-
-print("Filter with person: gates and year > 2023")
-raw_filter = {"$and": [{"person": "gates"}, {"year": {"$gt": 2023}}]}
-results = app.search("What is the net worth of Bill Gates?", raw_filter=raw_filter)
-print("Num of search results: ", len(results))
-```
diff --git a/embedchain/docs/api-reference/overview.mdx b/embedchain/docs/api-reference/overview.mdx
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/docs/api-reference/store/ai-assistants.mdx b/embedchain/docs/api-reference/store/ai-assistants.mdx
deleted file mode 100644
index 09c6122a4..000000000
--- a/embedchain/docs/api-reference/store/ai-assistants.mdx
+++ /dev/null
@@ -1,54 +0,0 @@
----
-title: 'AI Assistant'
----
-
-The `AIAssistant` class, an alternative to the OpenAI Assistant API, is designed for those who prefer using large language models (LLMs) other than those provided by OpenAI. It facilitates the creation of AI Assistants with several key benefits:
-
-- **Visibility into Citations**: It offers transparent access to the sources and citations used by the AI, enhancing the understanding and trustworthiness of its responses.
-
-- **Debugging Capabilities**: Users have the ability to delve into and debug the AI's processes, allowing for a deeper understanding and fine-tuning of its performance.
-
-- **Customizable Prompts**: The class provides the flexibility to modify and tailor prompts according to specific needs, enabling more precise and relevant interactions.
-
-- **Chain of Thought Integration**: It supports the incorporation of a 'chain of thought' approach, which helps in breaking down complex queries into simpler, sequential steps, thereby improving the clarity and accuracy of responses.
-
-It is ideal for those who value customization, transparency, and detailed control over their AI Assistant's functionalities.
-
-### Arguments
-
-
- Name for your AI assistant
-
-
-
- How the Assistant and model should behave or respond
-
-
-
- Load existing AI Assistant. If you pass this, you don't have to pass other arguments.
-
-
-
- Existing thread id if exists
-
-
-
- Embedchain pipeline config yaml path to use. This will define the configuration of the AI Assistant (such as configuring the LLM, vector database, and embedding model)
-
-
-
- Add data sources to your assistant. You can add in the following format: `[{"source": "https://example.com", "data_type": "web_page"}]`
-
-
-
- Anonymous telemetry (doesn't collect any user information or user's files). Used to improve the Embedchain package utilization. Default is `True`.
-
-
-
-## Usage
-
-For detailed guidance on creating your own AI Assistant, click the link below. It provides step-by-step instructions to help you through the process:
-
-
- Learn how to build a customized AI Assistant using the `AIAssistant` class.
-
diff --git a/embedchain/docs/api-reference/store/openai-assistant.mdx b/embedchain/docs/api-reference/store/openai-assistant.mdx
deleted file mode 100644
index 1ab21aa1f..000000000
--- a/embedchain/docs/api-reference/store/openai-assistant.mdx
+++ /dev/null
@@ -1,45 +0,0 @@
----
-title: 'OpenAI Assistant'
----
-
-### Arguments
-
-
- Name for your AI assistant
-
-
-
- how the Assistant and model should behave or respond
-
-
-
- Load existing OpenAI Assistant. If you pass this, you don't have to pass other arguments.
-
-
-
- Existing OpenAI thread id if exists
-
-
-
- OpenAI model to use
-
-
-
- OpenAI tools to use. Default set to `[{"type": "retrieval"}]`
-
-
-
- Add data sources to your assistant. You can add in the following format: `[{"source": "https://example.com", "data_type": "web_page"}]`
-
-
-
- Anonymous telemetry (doesn't collect any user information or user's files). Used to improve the Embedchain package utilization. Default is `True`.
-
-
-## Usage
-
-For detailed guidance on creating your own OpenAI Assistant, click the link below. It provides step-by-step instructions to help you through the process:
-
-
- Learn how to build an OpenAI Assistant using the `OpenAIAssistant` class.
-
diff --git a/embedchain/docs/community/connect-with-us.mdx b/embedchain/docs/community/connect-with-us.mdx
deleted file mode 100644
index e08dfd1c7..000000000
--- a/embedchain/docs/community/connect-with-us.mdx
+++ /dev/null
@@ -1,28 +0,0 @@
----
-title: 🤝 Connect with Us
----
-
-We believe in building a vibrant and supportive community around embedchain. There are various channels through which you can connect with us, stay updated, and contribute to the ongoing discussions:
-
-
-
- Follow us on Twitter
-
-
- Join our slack community
-
-
- Join our discord community
-
-
- Connect with us on LinkedIn
-
-
- Schedule a call with Embedchain founder
-
-
- Subscribe to our newsletter
-
-
-
-We look forward to connecting with you and seeing how we can create amazing things together!
diff --git a/embedchain/docs/components/data-sources/audio.mdx b/embedchain/docs/components/data-sources/audio.mdx
deleted file mode 100644
index 5f2772a71..000000000
--- a/embedchain/docs/components/data-sources/audio.mdx
+++ /dev/null
@@ -1,25 +0,0 @@
----
-title: "🎤 Audio"
----
-
-
-To use an audio as data source, just add `data_type` as `audio` and pass in the path of the audio (local or hosted).
-
-We use [Deepgram](https://developers.deepgram.com/docs/introduction) to transcribe the audiot to text, and then use the generated text as the data source.
-
-You would require an Deepgram API key which is available [here](https://console.deepgram.com/signup?jump=keys) to use this feature.
-
-### Without customization
-
-```python
-import os
-from embedchain import App
-
-os.environ["DEEPGRAM_API_KEY"] = "153xxx"
-
-app = App()
-app.add("introduction.wav", data_type="audio")
-response = app.query("What is my name and how old am I?")
-print(response)
-# Answer: Your name is Dave and you are 21 years old.
-```
diff --git a/embedchain/docs/components/data-sources/beehiiv.mdx b/embedchain/docs/components/data-sources/beehiiv.mdx
deleted file mode 100644
index 5a94cf1fe..000000000
--- a/embedchain/docs/components/data-sources/beehiiv.mdx
+++ /dev/null
@@ -1,16 +0,0 @@
----
-title: "🐝 Beehiiv"
----
-
-To add any Beehiiv data sources to your app, just add the base url as the source and set the data_type to `beehiiv`.
-
-```python
-from embedchain import App
-
-app = App()
-
-# source: just add the base url and set the data_type to 'beehiiv'
-app.add('https://aibreakfast.beehiiv.com', data_type='beehiiv')
-app.query("How much is OpenAI paying developers?")
-# Answer: OpenAI is aggressively recruiting Google's top AI researchers with offers ranging between $5 to $10 million annually, primarily in stock options.
-```
diff --git a/embedchain/docs/components/data-sources/csv.mdx b/embedchain/docs/components/data-sources/csv.mdx
deleted file mode 100644
index 07663a3b1..000000000
--- a/embedchain/docs/components/data-sources/csv.mdx
+++ /dev/null
@@ -1,28 +0,0 @@
----
-title: '📊 CSV'
----
-
-You can load any csv file from your local file system or through a URL. Headers are included for each line, so if you have an `age` column, `18` will be added as `age: 18`.
-
-## Usage
-
-### Load from a local file
-
-```python
-from embedchain import App
-app = App()
-app.add('/path/to/file.csv', data_type='csv')
-```
-
-### Load from URL
-
-```python
-from embedchain import App
-app = App()
-app.add('https://people.sc.fsu.edu/~jburkardt/data/csv/airtravel.csv', data_type="csv")
-```
-
-
-There is a size limit allowed for csv file beyond which it can throw error. This limit is set by the LLMs. Please consider chunking large csv files into smaller csv files.
-
-
diff --git a/embedchain/docs/components/data-sources/custom.mdx b/embedchain/docs/components/data-sources/custom.mdx
deleted file mode 100644
index 40a8c75e1..000000000
--- a/embedchain/docs/components/data-sources/custom.mdx
+++ /dev/null
@@ -1,42 +0,0 @@
----
-title: '⚙️ Custom'
----
-
-When we say "custom", we mean that you can customize the loader and chunker to your needs. This is done by passing a custom loader and chunker to the `add` method.
-
-```python
-from embedchain import App
-import your_loader
-from my_module import CustomLoader
-from my_module import CustomChunker
-
-app = App()
-loader = CustomLoader()
-chunker = CustomChunker()
-
-app.add("source", data_type="custom", loader=loader, chunker=chunker)
-```
-
-
- The custom loader and chunker must be a class that inherits from the [`BaseLoader`](https://github.com/embedchain/embedchain/blob/main/embedchain/loaders/base_loader.py) and [`BaseChunker`](https://github.com/embedchain/embedchain/blob/main/embedchain/chunkers/base_chunker.py) classes respectively.
-
-
-
- If the `data_type` is not a valid data type, the `add` method will fallback to the `custom` data type and expect a custom loader and chunker to be passed by the user.
-
-
-Example:
-
-```python
-from embedchain import App
-from embedchain.loaders.github import GithubLoader
-
-app = App()
-
-loader = GithubLoader(config={"token": "ghp_xxx"})
-
-app.add("repo:embedchain/embedchain type:repo", data_type="github", loader=loader)
-
-app.query("What is Embedchain?")
-# Answer: Embedchain is a Data Platform for Large Language Models (LLMs). It allows users to seamlessly load, index, retrieve, and sync unstructured data in order to build dynamic, LLM-powered applications. There is also a JavaScript implementation called embedchain-js available on GitHub.
-```
diff --git a/embedchain/docs/components/data-sources/data-type-handling.mdx b/embedchain/docs/components/data-sources/data-type-handling.mdx
deleted file mode 100644
index d939537af..000000000
--- a/embedchain/docs/components/data-sources/data-type-handling.mdx
+++ /dev/null
@@ -1,85 +0,0 @@
----
-title: 'Data type handling'
----
-
-## Automatic data type detection
-
-The add method automatically tries to detect the data_type, based on your input for the source argument. So `app.add('https://www.youtube.com/watch?v=dQw4w9WgXcQ')` is enough to embed a YouTube video.
-
-This detection is implemented for all formats. It is based on factors such as whether it's a URL, a local file, the source data type, etc.
-
-### Debugging automatic detection
-
-Set `log_level: DEBUG` in the config yaml to debug if the data type detection is done right or not. Otherwise, you will not know when, for instance, an invalid filepath is interpreted as raw text instead.
-
-### Forcing a data type
-
-To omit any issues with the data type detection, you can **force** a data_type by adding it as a `add` method argument.
-The examples below show you the keyword to force the respective `data_type`.
-
-Forcing can also be used for edge cases, such as interpreting a sitemap as a web_page, for reading its raw text instead of following links.
-
-## Remote data types
-
-
-**Use local files in remote data types**
-
-Some data_types are meant for remote content and only work with URLs.
-You can pass local files by formatting the path using the `file:` [URI scheme](https://en.wikipedia.org/wiki/File_URI_scheme), e.g. `file:///info.pdf`.
-
-
-## Reusing a vector database
-
-Default behavior is to create a persistent vector db in the directory **./db**. You can split your application into two Python scripts: one to create a local vector db and the other to reuse this local persistent vector db. This is useful when you want to index hundreds of documents and separately implement a chat interface.
-
-Create a local index:
-
-```python
-from embedchain import App
-
-config = {
- "app": {
- "config": {
- "id": "app-1"
- }
- }
-}
-naval_chat_bot = App.from_config(config=config)
-naval_chat_bot.add("https://www.youtube.com/watch?v=3qHkcs3kG44")
-naval_chat_bot.add("https://navalmanack.s3.amazonaws.com/Eric-Jorgenson_The-Almanack-of-Naval-Ravikant_Final.pdf")
-```
-
-You can reuse the local index with the same code, but without adding new documents:
-
-```python
-from embedchain import App
-
-config = {
- "app": {
- "config": {
- "id": "app-1"
- }
- }
-}
-naval_chat_bot = App.from_config(config=config)
-print(naval_chat_bot.query("What unique capacity does Naval argue humans possess when it comes to understanding explanations or concepts?"))
-```
-
-## Resetting an app and vector database
-
-You can reset the app by simply calling the `reset` method. This will delete the vector database and all other app related files.
-
-```python
-from embedchain import App
-
-app = App()config = {
- "app": {
- "config": {
- "id": "app-1"
- }
- }
-}
-naval_chat_bot = App.from_config(config=config)
-app.add("https://www.youtube.com/watch?v=3qHkcs3kG44")
-app.reset()
-```
diff --git a/embedchain/docs/components/data-sources/directory.mdx b/embedchain/docs/components/data-sources/directory.mdx
deleted file mode 100644
index 33c1e9b73..000000000
--- a/embedchain/docs/components/data-sources/directory.mdx
+++ /dev/null
@@ -1,41 +0,0 @@
----
-title: '📁 Directory/Folder'
----
-
-To use an entire directory as data source, just add `data_type` as `directory` and pass in the path of the local directory.
-
-### Without customization
-
-```python
-import os
-from embedchain import App
-
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-app = App()
-app.add("./elon-musk", data_type="directory")
-response = app.query("list all files")
-print(response)
-# Answer: Files are elon-musk-1.txt, elon-musk-2.pdf.
-```
-
-### Customization
-
-```python
-import os
-from embedchain import App
-from embedchain.loaders.directory_loader import DirectoryLoader
-
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-lconfig = {
- "recursive": True,
- "extensions": [".txt"]
-}
-loader = DirectoryLoader(config=lconfig)
-app = App()
-app.add("./elon-musk", loader=loader)
-response = app.query("what are all the files related to?")
-print(response)
-
-# Answer: The files are related to Elon Musk.
-```
diff --git a/embedchain/docs/components/data-sources/discord.mdx b/embedchain/docs/components/data-sources/discord.mdx
deleted file mode 100644
index 2c8780210..000000000
--- a/embedchain/docs/components/data-sources/discord.mdx
+++ /dev/null
@@ -1,28 +0,0 @@
----
-title: "💬 Discord"
----
-
-To add any Discord channel messages to your app, just add the `channel_id` as the source and set the `data_type` to `discord`.
-
-
- This loader requires a Discord bot token with read messages access.
- To obtain the token, follow the instructions provided in this tutorial:
- How to Get a Discord Bot Token?.
-
-
-```python
-import os
-from embedchain import App
-
-# add your discord "BOT" token
-os.environ["DISCORD_TOKEN"] = "xxx"
-
-app = App()
-
-app.add("1177296711023075338", data_type="discord")
-
-response = app.query("What is Joe saying about Elon Musk?")
-
-print(response)
-# Answer: Joe is saying "Elon Musk is a genius".
-```
diff --git a/embedchain/docs/components/data-sources/discourse.mdx b/embedchain/docs/components/data-sources/discourse.mdx
deleted file mode 100644
index 4ba0a36ce..000000000
--- a/embedchain/docs/components/data-sources/discourse.mdx
+++ /dev/null
@@ -1,44 +0,0 @@
----
-title: '🗨️ Discourse'
----
-
-You can now easily load data from your community built with [Discourse](https://discourse.org/).
-
-## Example
-
-1. Setup the Discourse Loader with your community url.
-```Python
-from embedchain.loaders.discourse import DiscourseLoader
-
-dicourse_loader = DiscourseLoader(config={"domain": "https://community.openai.com"})
-```
-
-2. Once you setup the loader, you can create an app and load data using the above discourse loader
-```Python
-import os
-from embedchain.pipeline import Pipeline as App
-
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-app = App()
-
-app.add("openai after:2023-10-1", data_type="discourse", loader=dicourse_loader)
-
-question = "Where can I find the OpenAI API status page?"
-app.query(question)
-# Answer: You can find the OpenAI API status page at https:/status.openai.com/.
-```
-
-NOTE: The `add` function of the app will accept any executable search query to load data. Refer [Discourse API Docs](https://docs.discourse.org/#tag/Search) to learn more about search queries.
-
-3. We automatically create a chunker to chunk your discourse data, however if you wish to provide your own chunker class. Here is how you can do that:
-```Python
-
-from embedchain.chunkers.discourse import DiscourseChunker
-from embedchain.config.add_config import ChunkerConfig
-
-discourse_chunker_config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
-discourse_chunker = DiscourseChunker(config=discourse_chunker_config)
-
-app.add("openai", data_type='discourse', loader=dicourse_loader, chunker=discourse_chunker)
-```
\ No newline at end of file
diff --git a/embedchain/docs/components/data-sources/docs-site.mdx b/embedchain/docs/components/data-sources/docs-site.mdx
deleted file mode 100644
index 342bbdc85..000000000
--- a/embedchain/docs/components/data-sources/docs-site.mdx
+++ /dev/null
@@ -1,14 +0,0 @@
----
-title: '📚 Code Docs website'
----
-
-To add any code documentation website as a loader, use the data_type as `docs_site`. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-app.add("https://docs.embedchain.ai/", data_type="docs_site")
-app.query("What is Embedchain?")
-# Answer: Embedchain is a platform that utilizes various components, including paid/proprietary ones, to provide what is believed to be the best configuration available. It uses LLM (Language Model) providers such as OpenAI, Anthpropic, Vertex_AI, GPT4ALL, Azure_OpenAI, LLAMA2, JINA, Ollama, Together and COHERE. Embedchain allows users to import and utilize these LLM providers for their applications.'
-```
diff --git a/embedchain/docs/components/data-sources/docx.mdx b/embedchain/docs/components/data-sources/docx.mdx
deleted file mode 100644
index cc459621f..000000000
--- a/embedchain/docs/components/data-sources/docx.mdx
+++ /dev/null
@@ -1,18 +0,0 @@
----
-title: '📄 Docx file'
----
-
-### Docx file
-
-To add any doc/docx file, use the data_type as `docx`. `docx` allows remote urls and conventional file paths. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-app.add('https://example.com/content/intro.docx', data_type="docx")
-# Or add file using the local file path on your system
-# app.add('content/intro.docx', data_type="docx")
-
-app.query("Summarize the docx data?")
-```
diff --git a/embedchain/docs/components/data-sources/dropbox.mdx b/embedchain/docs/components/data-sources/dropbox.mdx
deleted file mode 100644
index bb2800bf8..000000000
--- a/embedchain/docs/components/data-sources/dropbox.mdx
+++ /dev/null
@@ -1,37 +0,0 @@
----
-title: '💾 Dropbox'
----
-
-To load folders or files from your Dropbox account, configure the `data_type` parameter as `dropbox` and specify the path to the desired file or folder, starting from the root directory of your Dropbox account.
-
-For Dropbox access, an **access token** is required. Obtain this token by visiting [Dropbox Developer Apps](https://www.dropbox.com/developers/apps). There, create a new app and generate an access token for it.
-
-Ensure your app has the following settings activated:
-
-- In the Permissions section, enable `files.content.read` and `files.metadata.read`.
-
-## Usage
-
-Install the `dropbox` pypi package:
-
-```bash
-pip install dropbox
-```
-
-Following is an example of how to use the dropbox loader:
-
-```python
-import os
-from embedchain import App
-
-os.environ["DROPBOX_ACCESS_TOKEN"] = "sl.xxx"
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-app = App()
-
-# any path from the root of your dropbox account, you can leave it "" for the root folder
-app.add("/test", data_type="dropbox")
-
-print(app.query("Which two celebrities are mentioned here?"))
-# The two celebrities mentioned in the given context are Elon Musk and Jeff Bezos.
-```
diff --git a/embedchain/docs/components/data-sources/excel-file.mdx b/embedchain/docs/components/data-sources/excel-file.mdx
deleted file mode 100644
index af8a2cd62..000000000
--- a/embedchain/docs/components/data-sources/excel-file.mdx
+++ /dev/null
@@ -1,18 +0,0 @@
----
-title: '📄 Excel file'
----
-
-### Excel file
-
-To add any xlsx/xls file, use the data_type as `excel_file`. `excel_file` allows remote urls and conventional file paths. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-app.add('https://example.com/content/intro.xlsx', data_type="excel_file")
-# Or add file using the local file path on your system
-# app.add('content/intro.xls', data_type="excel_file")
-
-app.query("Give brief information about data.")
-```
diff --git a/embedchain/docs/components/data-sources/github.mdx b/embedchain/docs/components/data-sources/github.mdx
deleted file mode 100644
index 14791aca4..000000000
--- a/embedchain/docs/components/data-sources/github.mdx
+++ /dev/null
@@ -1,52 +0,0 @@
----
-title: 📝 Github
----
-
-1. Setup the Github loader by configuring the Github account with username and personal access token (PAT). Check out [this](https://docs.github.com/en/enterprise-server@3.6/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens#creating-a-personal-access-token) link to learn how to create a PAT.
-```Python
-from embedchain.loaders.github import GithubLoader
-
-loader = GithubLoader(
- config={
- "token":"ghp_xxxx"
- }
- )
-```
-
-2. Once you setup the loader, you can create an app and load data using the above Github loader
-```Python
-import os
-from embedchain.pipeline import Pipeline as App
-
-os.environ["OPENAI_API_KEY"] = "sk-xxxx"
-
-app = App()
-
-app.add("repo:embedchain/embedchain type:repo", data_type="github", loader=loader)
-
-response = app.query("What is Embedchain?")
-# Answer: Embedchain is a Data Platform for Large Language Models (LLMs). It allows users to seamlessly load, index, retrieve, and sync unstructured data in order to build dynamic, LLM-powered applications. There is also a JavaScript implementation called embedchain-js available on GitHub.
-```
-The `add` function of the app will accept any valid github query with qualifiers. It only supports loading github code, repository, issues and pull-requests.
-
-You must provide qualifiers `type:` and `repo:` in the query. The `type:` qualifier can be a combination of `code`, `repo`, `pr`, `issue`, `branch`, `file`. The `repo:` qualifier must be a valid github repository name.
-
-
-
- - `repo:embedchain/embedchain type:repo` - to load the repository
- - `repo:embedchain/embedchain type:branch name:feature_test` - to load the branch of the repository
- - `repo:embedchain/embedchain type:file path:README.md` - to load the specific file of the repository
- - `repo:embedchain/embedchain type:issue,pr` - to load the issues and pull-requests of the repository
- - `repo:embedchain/embedchain type:issue state:closed` - to load the closed issues of the repository
-
-
-3. We automatically create a chunker to chunk your GitHub data, however if you wish to provide your own chunker class. Here is how you can do that:
-```Python
-from embedchain.chunkers.common_chunker import CommonChunker
-from embedchain.config.add_config import ChunkerConfig
-
-github_chunker_config = ChunkerConfig(chunk_size=2000, chunk_overlap=0, length_function=len)
-github_chunker = CommonChunker(config=github_chunker_config)
-
-app.add(load_query, data_type="github", loader=loader, chunker=github_chunker)
-```
diff --git a/embedchain/docs/components/data-sources/gmail.mdx b/embedchain/docs/components/data-sources/gmail.mdx
deleted file mode 100644
index aaaf002ed..000000000
--- a/embedchain/docs/components/data-sources/gmail.mdx
+++ /dev/null
@@ -1,34 +0,0 @@
----
-title: '📬 Gmail'
----
-
-To use GmailLoader you must install the extra dependencies with `pip install --upgrade embedchain[gmail]`.
-
-The `source` must be a valid Gmail search query, you can refer `https://support.google.com/mail/answer/7190?hl=en` to build a query.
-
-To load Gmail messages, you MUST use the data_type as `gmail`. Otherwise the source will be detected as simple `text`.
-
-To use this you need to save `credentials.json` in the directory from where you will run the loader. Follow these steps to get the credentials
-
-1. Go to the [Google Cloud Console](https://console.cloud.google.com/apis/credentials).
-2. Create a project if you don't have one already.
-3. Create an `OAuth Consent Screen` in the project. You may need to select the `external` option.
-4. Make sure the consent screen is published.
-5. Enable the [Gmail API](https://console.cloud.google.com/apis/api/gmail.googleapis.com)
-6. Create credentials from the `Credentials` tab.
-7. Select the type `OAuth Client ID`.
-8. Choose the application type `Web application`. As a name you can choose `embedchain` or any other name as per your use case.
-9. Add an authorized redirect URI for `http://localhost:8080/`.
-10. You can leave everything else at default, finish the creation.
-11. When you are done, a modal opens where you can download the details in `json` format.
-12. Put the `.json` file in your current directory and rename it to `credentials.json`
-
-```python
-from embedchain import App
-
-app = App()
-
-gmail_filter = "to: me label:inbox"
-app.add(gmail_filter, data_type="gmail")
-app.query("Summarize my email conversations")
-```
\ No newline at end of file
diff --git a/embedchain/docs/components/data-sources/google-drive.mdx b/embedchain/docs/components/data-sources/google-drive.mdx
deleted file mode 100644
index 5dcf4e45f..000000000
--- a/embedchain/docs/components/data-sources/google-drive.mdx
+++ /dev/null
@@ -1,28 +0,0 @@
----
-title: 'Google Drive'
----
-
-To use GoogleDriveLoader you must install the extra dependencies with `pip install --upgrade embedchain[googledrive]`.
-
-The data_type must be `google_drive`. Otherwise, it will be considered a regular web page.
-
-Google Drive requires the setup of credentials. This can be done by following the steps below:
-
-1. Go to the [Google Cloud Console](https://console.cloud.google.com/apis/credentials).
-2. Create a project if you don't have one already.
-3. Enable the [Google Drive API](https://console.cloud.google.com/flows/enableapi?apiid=drive.googleapis.com)
-4. [Authorize credentials for desktop app](https://developers.google.com/drive/api/quickstart/python#authorize_credentials_for_a_desktop_application)
-5. When done, you will be able to download the credentials in `json` format. Rename the downloaded file to `credentials.json` and save it in `~/.credentials/credentials.json`
-6. Set the environment variable `GOOGLE_APPLICATION_CREDENTIALS=~/.credentials/credentials.json`
-
-The first time you use the loader, you will be prompted to enter your Google account credentials.
-
-
-```python
-from embedchain import App
-
-app = App()
-
-url = "https://drive.google.com/drive/u/0/folders/xxx-xxx"
-app.add(url, data_type="google_drive")
-```
diff --git a/embedchain/docs/components/data-sources/image.mdx b/embedchain/docs/components/data-sources/image.mdx
deleted file mode 100644
index b79043660..000000000
--- a/embedchain/docs/components/data-sources/image.mdx
+++ /dev/null
@@ -1,45 +0,0 @@
----
-title: "🖼️ Image"
----
-
-
-To use an image as data source, just add `data_type` as `image` and pass in the path of the image (local or hosted).
-
-We use [GPT4 Vision](https://platform.openai.com/docs/guides/vision) to generate meaning of the image using a custom prompt, and then use the generated text as the data source.
-
-You would require an OpenAI API key with access to `gpt-4-vision-preview` model to use this feature.
-
-### Without customization
-
-```python
-import os
-from embedchain import App
-
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-app = App()
-app.add("./Elon-Musk.webp", data_type="image")
-response = app.query("Describe the man in the image.")
-print(response)
-# Answer: The man in the image is dressed in formal attire, wearing a dark suit jacket and a white collared shirt. He has short hair and is standing. He appears to be gazing off to the side with a reflective expression. The background is dark with faint, warm-toned vertical lines, possibly from a lit environment behind the individual or reflections. The overall atmosphere is somewhat moody and introspective.
-```
-
-### Customization
-
-```python
-import os
-from embedchain import App
-from embedchain.loaders.image import ImageLoader
-
-image_loader = ImageLoader(
- max_tokens=100,
- api_key="sk-xxx",
- prompt="Is the person looking wealthy? Structure your thoughts around what you see in the image.",
-)
-
-app = App()
-app.add("./Elon-Musk.webp", data_type="image", loader=image_loader)
-response = app.query("Describe the man in the image.")
-print(response)
-# Answer: The man in the image appears to be well-dressed in a suit and shirt, suggesting that he may be in a professional or formal setting. His composed demeanor and confident posture further indicate a sense of self-assurance. Based on these visual cues, one could infer that the man may have a certain level of economic or social status, possibly indicating wealth or professional success.
-```
diff --git a/embedchain/docs/components/data-sources/json.mdx b/embedchain/docs/components/data-sources/json.mdx
deleted file mode 100644
index 4d38a0a55..000000000
--- a/embedchain/docs/components/data-sources/json.mdx
+++ /dev/null
@@ -1,44 +0,0 @@
----
-title: '📃 JSON'
----
-
-To add any json file, use the data_type as `json`. Headers are included for each line, so for example if you have a json like `{"age": 18}`, then it will be added as `age: 18`.
-
-Here are the supported sources for loading `json`:
-
-```
-1. URL - valid url to json file that ends with ".json" extension.
-2. Local file - valid url to local json file that ends with ".json" extension.
-3. String - valid json string (e.g. - app.add('{"foo": "bar"}'))
-```
-
-
-If you would like to add other data structures (e.g. list, dict etc.), convert it to a valid json first using `json.dumps()` function.
-
-
-## Example
-
-
-
-```python python
-from embedchain import App
-
-app = App()
-
-# Add json file
-app.add("temp.json")
-
-app.query("What is the net worth of Elon Musk as of October 2023?")
-# As of October 2023, Elon Musk's net worth is $255.2 billion.
-```
-
-
-```json temp.json
-{
- "question": "What is your net worth, Elon Musk?",
- "answer": "As of October 2023, Elon Musk's net worth is $255.2 billion, making him one of the wealthiest individuals in the world."
-}
-```
-
-
-
diff --git a/embedchain/docs/components/data-sources/mdx.mdx b/embedchain/docs/components/data-sources/mdx.mdx
deleted file mode 100644
index c59569e50..000000000
--- a/embedchain/docs/components/data-sources/mdx.mdx
+++ /dev/null
@@ -1,14 +0,0 @@
----
-title: '📝 Mdx file'
----
-
-To add any `.mdx` file to your app, use the data_type (first argument to `.add()` method) as `mdx`. Note that this supports support mdx file present on machine, so this should be a file path. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-app.add('path/to/file.mdx', data_type='mdx')
-
-app.query("What are the docs about?")
-```
diff --git a/embedchain/docs/components/data-sources/mysql.mdx b/embedchain/docs/components/data-sources/mysql.mdx
deleted file mode 100644
index 2a5cb7a01..000000000
--- a/embedchain/docs/components/data-sources/mysql.mdx
+++ /dev/null
@@ -1,47 +0,0 @@
----
-title: '🐬 MySQL'
----
-
-1. Setup the MySQL loader by configuring the SQL db.
-```Python
-from embedchain.loaders.mysql import MySQLLoader
-
-config = {
- "host": "host",
- "port": "port",
- "database": "database",
- "user": "username",
- "password": "password",
-}
-
-mysql_loader = MySQLLoader(config=config)
-```
-
-For more details on how to setup with valid config, check MySQL [documentation](https://dev.mysql.com/doc/connector-python/en/connector-python-connectargs.html).
-
-2. Once you setup the loader, you can create an app and load data using the above MySQL loader
-```Python
-from embedchain.pipeline import Pipeline as App
-
-app = App()
-
-app.add("SELECT * FROM table_name;", data_type='mysql', loader=mysql_loader)
-# Adds `(1, 'What is your net worth, Elon Musk?', "As of October 2023, Elon Musk's net worth is $255.2 billion.")`
-
-response = app.query(question)
-# Answer: As of October 2023, Elon Musk's net worth is $255.2 billion.
-```
-
-NOTE: The `add` function of the app will accept any executable query to load data. DO NOT pass the `CREATE`, `INSERT` queries in `add` function.
-
-3. We automatically create a chunker to chunk your SQL data, however if you wish to provide your own chunker class. Here is how you can do that:
-```Python
-
-from embedchain.chunkers.mysql import MySQLChunker
-from embedchain.config.add_config import ChunkerConfig
-
-mysql_chunker_config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
-mysql_chunker = MySQLChunker(config=mysql_chunker_config)
-
-app.add("SELECT * FROM table_name;", data_type='mysql', loader=mysql_loader, chunker=mysql_chunker)
-```
diff --git a/embedchain/docs/components/data-sources/notion.mdx b/embedchain/docs/components/data-sources/notion.mdx
deleted file mode 100644
index d6c616df8..000000000
--- a/embedchain/docs/components/data-sources/notion.mdx
+++ /dev/null
@@ -1,20 +0,0 @@
----
-title: '📓 Notion'
----
-
-To use notion you must install the extra dependencies with `pip install --upgrade embedchain[community]`.
-
-To load a notion page, use the data_type as `notion`. Since it is hard to automatically detect, it is advised to specify the `data_type` when adding a notion document.
-The next argument must **end** with the `notion page id`. The id is a 32-character string. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-
-app.add("cfbc134ca6464fc980d0391613959196", data_type="notion")
-app.add("my-page-cfbc134ca6464fc980d0391613959196", data_type="notion")
-app.add("https://www.notion.so/my-page-cfbc134ca6464fc980d0391613959196", data_type="notion")
-
-app.query("Summarize the notion doc")
-```
diff --git a/embedchain/docs/components/data-sources/openapi.mdx b/embedchain/docs/components/data-sources/openapi.mdx
deleted file mode 100644
index 84bc966b2..000000000
--- a/embedchain/docs/components/data-sources/openapi.mdx
+++ /dev/null
@@ -1,22 +0,0 @@
----
-title: 🙌 OpenAPI
----
-
-To add any OpenAPI spec yaml file (currently the json file will be detected as JSON data type), use the data_type as 'openapi'. 'openapi' allows remote urls and conventional file paths.
-
-```python
-from embedchain import App
-
-app = App()
-
-app.add("https://github.com/openai/openai-openapi/blob/master/openapi.yaml", data_type="openapi")
-# Or add using the local file path
-# app.add("configs/openai_openapi.yaml", data_type="openapi")
-
-app.query("What can OpenAI API endpoint do? Can you list the things it can learn from?")
-# Answer: The OpenAI API endpoint allows users to interact with OpenAI's models and perform various tasks such as generating text, answering questions, summarizing documents, translating languages, and more. The specific capabilities and tasks that the API can learn from may vary depending on the models and features provided by OpenAI. For more detailed information, it is recommended to refer to the OpenAI API documentation at https://platform.openai.com/docs/api-reference.
-```
-
-
-The yaml file added to the App must have the required OpenAPI fields otherwise the adding OpenAPI spec will fail. Please refer to [OpenAPI Spec Doc](https://spec.openapis.org/oas/v3.1.0)
-
\ No newline at end of file
diff --git a/embedchain/docs/components/data-sources/overview.mdx b/embedchain/docs/components/data-sources/overview.mdx
deleted file mode 100644
index 66f5948a3..000000000
--- a/embedchain/docs/components/data-sources/overview.mdx
+++ /dev/null
@@ -1,43 +0,0 @@
----
-title: Overview
----
-
-Embedchain comes with built-in support for various data sources. We handle the complexity of loading unstructured data from these data sources, allowing you to easily customize your app through a user-friendly interface.
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
diff --git a/embedchain/docs/components/data-sources/pdf-file.mdx b/embedchain/docs/components/data-sources/pdf-file.mdx
deleted file mode 100644
index 9cc45910a..000000000
--- a/embedchain/docs/components/data-sources/pdf-file.mdx
+++ /dev/null
@@ -1,43 +0,0 @@
----
-title: '📰 PDF'
----
-
-You can load any pdf file from your local file system or through a URL.
-
-## Usage
-
-### Load from a local file
-
-```python
-from embedchain import App
-app = App()
-app.add('/path/to/file.pdf', data_type='pdf_file')
-```
-
-### Load from URL
-
-```python
-from embedchain import App
-app = App()
-app.add('https://arxiv.org/pdf/1706.03762.pdf', data_type='pdf_file')
-app.query("What is the paper 'attention is all you need' about?", citations=True)
-# Answer: The paper "Attention Is All You Need" proposes a new network architecture called the Transformer, which is based solely on attention mechanisms. It suggests that complex recurrent or convolutional neural networks can be replaced with a simpler architecture that connects the encoder and decoder through attention. The paper discusses how this approach can improve sequence transduction models, such as neural machine translation.
-# Contexts:
-# [
-# (
-# 'Provided proper attribution is ...',
-# {
-# 'page': 0,
-# 'url': 'https://arxiv.org/pdf/1706.03762.pdf',
-# 'score': 0.3676220203221626,
-# ...
-# }
-# ),
-# ]
-```
-
-We also store the page number under the key `page` with each chunk that helps understand where the answer is coming from. You can fetch the `page` key while during retrieval (refer to the example given above).
-
-
-Note that we do not support password protected pdf files.
-
diff --git a/embedchain/docs/components/data-sources/postgres.mdx b/embedchain/docs/components/data-sources/postgres.mdx
deleted file mode 100644
index 9cb5d0e6e..000000000
--- a/embedchain/docs/components/data-sources/postgres.mdx
+++ /dev/null
@@ -1,64 +0,0 @@
----
-title: '🐘 Postgres'
----
-
-1. Setup the Postgres loader by configuring the postgres db.
-```Python
-from embedchain.loaders.postgres import PostgresLoader
-
-config = {
- "host": "host_address",
- "port": "port_number",
- "dbname": "database_name",
- "user": "username",
- "password": "password",
-}
-
-"""
-config = {
- "url": "your_postgres_url"
-}
-"""
-
-postgres_loader = PostgresLoader(config=config)
-
-```
-
-You can either setup the loader by passing the postgresql url or by providing the config data.
-For more details on how to setup with valid url and config, check postgres [documentation](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNSTRING:~:text=34.1.1.%C2%A0Connection%20Strings-,%23,-Several%20libpq%20functions).
-
-NOTE: if you provide the `url` field in config, all other fields will be ignored.
-
-2. Once you setup the loader, you can create an app and load data using the above postgres loader
-```Python
-import os
-from embedchain.pipeline import Pipeline as App
-
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-app = App()
-
-question = "What is Elon Musk's networth?"
-response = app.query(question)
-# Answer: As of September 2021, Elon Musk's net worth is estimated to be around $250 billion, making him one of the wealthiest individuals in the world. However, please note that net worth can fluctuate over time due to various factors such as stock market changes and business ventures.
-
-app.add("SELECT * FROM table_name;", data_type='postgres', loader=postgres_loader)
-# Adds `(1, 'What is your net worth, Elon Musk?', "As of October 2023, Elon Musk's net worth is $255.2 billion.")`
-
-response = app.query(question)
-# Answer: As of October 2023, Elon Musk's net worth is $255.2 billion.
-```
-
-NOTE: The `add` function of the app will accept any executable query to load data. DO NOT pass the `CREATE`, `INSERT` queries in `add` function as they will result in not adding any data, so it is pointless.
-
-3. We automatically create a chunker to chunk your postgres data, however if you wish to provide your own chunker class. Here is how you can do that:
-```Python
-
-from embedchain.chunkers.postgres import PostgresChunker
-from embedchain.config.add_config import ChunkerConfig
-
-postgres_chunker_config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
-postgres_chunker = PostgresChunker(config=postgres_chunker_config)
-
-app.add("SELECT * FROM table_name;", data_type='postgres', loader=postgres_loader, chunker=postgres_chunker)
-```
\ No newline at end of file
diff --git a/embedchain/docs/components/data-sources/qna.mdx b/embedchain/docs/components/data-sources/qna.mdx
deleted file mode 100644
index 3efaa47ff..000000000
--- a/embedchain/docs/components/data-sources/qna.mdx
+++ /dev/null
@@ -1,13 +0,0 @@
----
-title: '❓💬 Question and answer pair'
----
-
-QnA pair is a local data type. To supply your own QnA pair, use the data_type as `qna_pair` and enter a tuple. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-
-app.add(("Question", "Answer"), data_type="qna_pair")
-```
diff --git a/embedchain/docs/components/data-sources/sitemap.mdx b/embedchain/docs/components/data-sources/sitemap.mdx
deleted file mode 100644
index 96b47ef1c..000000000
--- a/embedchain/docs/components/data-sources/sitemap.mdx
+++ /dev/null
@@ -1,13 +0,0 @@
----
-title: '🗺️ Sitemap'
----
-
-Add all web pages from an xml-sitemap. Filters non-text files. Use the data_type as `sitemap`. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-
-app.add('https://example.com/sitemap.xml', data_type='sitemap')
-```
\ No newline at end of file
diff --git a/embedchain/docs/components/data-sources/slack.mdx b/embedchain/docs/components/data-sources/slack.mdx
deleted file mode 100644
index 7b879fd6d..000000000
--- a/embedchain/docs/components/data-sources/slack.mdx
+++ /dev/null
@@ -1,71 +0,0 @@
----
-title: '🤖 Slack'
----
-
-## Pre-requisite
-- Download required packages by running `pip install --upgrade "embedchain[slack]"`.
-- Configure your slack bot token as environment variable `SLACK_USER_TOKEN`.
- - Find your user token on your [Slack Account](https://api.slack.com/authentication/token-types)
- - Make sure your slack user token includes [search](https://api.slack.com/scopes/search:read) scope.
-
-## Example
-
-### Get Started
-
-This will automatically retrieve data from the workspace associated with the user's token.
-
-```python
-import os
-from embedchain import App
-
-os.environ["SLACK_USER_TOKEN"] = "xoxp-xxx"
-app = App()
-
-app.add("in:general", data_type="slack")
-
-result = app.query("what are the messages in general channel?")
-
-print(result)
-```
-
-
-### Customize your SlackLoader
-1. Setup the Slack loader by configuring the Slack Webclient.
-```Python
-from embedchain.loaders.slack import SlackLoader
-
-os.environ["SLACK_USER_TOKEN"] = "xoxp-*"
-
-config = {
- 'base_url': slack_app_url,
- 'headers': web_headers,
- 'team_id': slack_team_id,
-}
-
-loader = SlackLoader(config)
-```
-
-NOTE: you can also pass the `config` with `base_url`, `headers`, `team_id` to setup your SlackLoader.
-
-2. Once you setup the loader, you can create an app and load data using the above slack loader
-```Python
-import os
-from embedchain.pipeline import Pipeline as App
-
-app = App()
-
-app.add("in:random", data_type="slack", loader=loader)
-question = "Which bots are available in the slack workspace's random channel?"
-# Answer: The available bot in the slack workspace's random channel is the Embedchain bot.
-```
-
-3. We automatically create a chunker to chunk your slack data, however if you wish to provide your own chunker class. Here is how you can do that:
-```Python
-from embedchain.chunkers.slack import SlackChunker
-from embedchain.config.add_config import ChunkerConfig
-
-slack_chunker_config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
-slack_chunker = SlackChunker(config=slack_chunker_config)
-
-app.add(slack_chunker, data_type="slack", loader=loader, chunker=slack_chunker)
-```
\ No newline at end of file
diff --git a/embedchain/docs/components/data-sources/substack.mdx b/embedchain/docs/components/data-sources/substack.mdx
deleted file mode 100644
index dd10a9e7d..000000000
--- a/embedchain/docs/components/data-sources/substack.mdx
+++ /dev/null
@@ -1,16 +0,0 @@
----
-title: "📝 Substack"
----
-
-To add any Substack data sources to your app, just add the main base url as the source and set the data_type to `substack`.
-
-```python
-from embedchain import App
-
-app = App()
-
-# source: for any substack just add the root URL
-app.add('https://www.lennysnewsletter.com', data_type='substack')
-app.query("Who is Brian Chesky?")
-# Answer: Brian Chesky is the co-founder and CEO of Airbnb.
-```
diff --git a/embedchain/docs/components/data-sources/text-file.mdx b/embedchain/docs/components/data-sources/text-file.mdx
deleted file mode 100644
index 14b48c005..000000000
--- a/embedchain/docs/components/data-sources/text-file.mdx
+++ /dev/null
@@ -1,14 +0,0 @@
----
-title: '📄 Text file'
----
-
-To add a .txt file, specify the data_type as `text_file`. The URL provided in the first parameter of the `add` function, should be a local path. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-app.add('path/to/file.txt', data_type="text_file")
-
-app.query("Summarize the information of the text file")
-```
\ No newline at end of file
diff --git a/embedchain/docs/components/data-sources/text.mdx b/embedchain/docs/components/data-sources/text.mdx
deleted file mode 100644
index 0fda6f573..000000000
--- a/embedchain/docs/components/data-sources/text.mdx
+++ /dev/null
@@ -1,17 +0,0 @@
----
-title: '📝 Text'
----
-
-### Text
-
-Text is a local data type. To supply your own text, use the data_type as `text` and enter a string. The text is not processed, this can be very versatile. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-
-app.add('Seek wealth, not money or status. Wealth is having assets that earn while you sleep. Money is how we transfer time and wealth. Status is your place in the social hierarchy.', data_type='text')
-```
-
-Note: This is not used in the examples because in most cases you will supply a whole paragraph or file, which did not fit.
diff --git a/embedchain/docs/components/data-sources/web-page.mdx b/embedchain/docs/components/data-sources/web-page.mdx
deleted file mode 100644
index f4a50a923..000000000
--- a/embedchain/docs/components/data-sources/web-page.mdx
+++ /dev/null
@@ -1,13 +0,0 @@
----
-title: '🌐 HTML Web page'
----
-
-To add any web page, use the data_type as `web_page`. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-
-app.add('a_valid_web_page_url', data_type='web_page')
-```
diff --git a/embedchain/docs/components/data-sources/xml.mdx b/embedchain/docs/components/data-sources/xml.mdx
deleted file mode 100644
index afe9a4124..000000000
--- a/embedchain/docs/components/data-sources/xml.mdx
+++ /dev/null
@@ -1,17 +0,0 @@
----
-title: '🧾 XML file'
----
-
-### XML file
-
-To add any xml file, use the data_type as `xml`. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-
-app.add('content/data.xml')
-```
-
-Note: Only the text content of the xml file will be added to the app. The tags will be ignored.
diff --git a/embedchain/docs/components/data-sources/youtube-channel.mdx b/embedchain/docs/components/data-sources/youtube-channel.mdx
deleted file mode 100644
index d9f037ff0..000000000
--- a/embedchain/docs/components/data-sources/youtube-channel.mdx
+++ /dev/null
@@ -1,22 +0,0 @@
----
-title: '📽️ Youtube Channel'
----
-
-## Setup
-
-Make sure you have all the required packages installed before using this data type. You can install them by running the following command in your terminal.
-
-```bash
-pip install -U "embedchain[youtube]"
-```
-
-## Usage
-
-To add all the videos from a youtube channel to your app, use the data_type as `youtube_channel`.
-
-```python
-from embedchain import App
-
-app = App()
-app.add("@channel_name", data_type="youtube_channel")
-```
diff --git a/embedchain/docs/components/data-sources/youtube-video.mdx b/embedchain/docs/components/data-sources/youtube-video.mdx
deleted file mode 100644
index 01ac52406..000000000
--- a/embedchain/docs/components/data-sources/youtube-video.mdx
+++ /dev/null
@@ -1,22 +0,0 @@
----
-title: '📺 Youtube Video'
----
-
-## Setup
-
-Make sure you have all the required packages installed before using this data type. You can install them by running the following command in your terminal.
-
-```bash
-pip install -U "embedchain[youtube]"
-```
-
-## Usage
-
-To add any youtube video to your app, use the data_type as `youtube_video`. Eg:
-
-```python
-from embedchain import App
-
-app = App()
-app.add('a_valid_youtube_url_here', data_type='youtube_video')
-```
diff --git a/embedchain/docs/components/embedding-models.mdx b/embedchain/docs/components/embedding-models.mdx
deleted file mode 100644
index 7af84236b..000000000
--- a/embedchain/docs/components/embedding-models.mdx
+++ /dev/null
@@ -1,470 +0,0 @@
----
-title: 🧩 Embedding models
----
-
-## Overview
-
-Embedchain supports several embedding models from the following providers:
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-## OpenAI
-
-To use OpenAI embedding function, you have to set the `OPENAI_API_KEY` environment variable. You can obtain the OpenAI API key from the [OpenAI Platform](https://platform.openai.com/account/api-keys).
-
-Once you have obtained the key, you can use it like this:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ['OPENAI_API_KEY'] = 'xxx'
-
-# load embedding model configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-
-app.add("https://en.wikipedia.org/wiki/OpenAI")
-app.query("What is OpenAI?")
-```
-
-```yaml config.yaml
-embedder:
- provider: openai
- config:
- model: 'text-embedding-3-small'
-```
-
-
-
-* OpenAI announced two new embedding models: `text-embedding-3-small` and `text-embedding-3-large`. Embedchain supports both these models. Below you can find YAML config for both:
-
-
-
-```yaml text-embedding-3-small.yaml
-embedder:
- provider: openai
- config:
- model: 'text-embedding-3-small'
-```
-
-```yaml text-embedding-3-large.yaml
-embedder:
- provider: openai
- config:
- model: 'text-embedding-3-large'
-```
-
-
-
-## Google AI
-
-To use Google AI embedding function, you have to set the `GOOGLE_API_KEY` environment variable. You can obtain the Google API key from the [Google Maker Suite](https://makersuite.google.com/app/apikey)
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["GOOGLE_API_KEY"] = "xxx"
-
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-embedder:
- provider: google
- config:
- model: 'models/embedding-001'
- task_type: "retrieval_document"
- title: "Embeddings for Embedchain"
-```
-
-
-
-For more details regarding the Google AI embedding model, please refer to the [Google AI documentation](https://ai.google.dev/tutorials/python_quickstart#use_embeddings).
-
-
-## AWS Bedrock
-
-To use AWS Bedrock embedding function, you have to set the AWS environment variable.
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["AWS_ACCESS_KEY_ID"] = "xxx"
-os.environ["AWS_SECRET_ACCESS_KEY"] = "xxx"
-os.environ["AWS_REGION"] = "us-west-2"
-
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-embedder:
- provider: aws_bedrock
- config:
- model: 'amazon.titan-embed-text-v2:0'
- vector_dimension: 1024
- task_type: "retrieval_document"
- title: "Embeddings for Embedchain"
-```
-
-
-
-For more details regarding the AWS Bedrock embedding model, please refer to the [AWS Bedrock documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html).
-
-
-## Azure OpenAI
-
-To use Azure OpenAI embedding model, you have to set some of the azure openai related environment variables as given in the code block below:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["OPENAI_API_TYPE"] = "azure"
-os.environ["AZURE_OPENAI_ENDPOINT"] = "https://xxx.openai.azure.com/"
-os.environ["AZURE_OPENAI_API_KEY"] = "xxx"
-os.environ["OPENAI_API_VERSION"] = "xxx"
-
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: azure_openai
- config:
- model: gpt-35-turbo
- deployment_name: your_llm_deployment_name
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-embedder:
- provider: azure_openai
- config:
- model: text-embedding-ada-002
- deployment_name: you_embedding_model_deployment_name
-```
-
-
-You can find the list of models and deployment name on the [Azure OpenAI Platform](https://oai.azure.com/portal).
-
-## GPT4ALL
-
-GPT4All supports generating high quality embeddings of arbitrary length documents of text using a CPU optimized contrastively trained Sentence Transformer.
-
-
-
-```python main.py
-from embedchain import App
-
-# load embedding model configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: gpt4all
- config:
- model: 'orca-mini-3b-gguf2-q4_0.gguf'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-embedder:
- provider: gpt4all
-```
-
-
-
-## Hugging Face
-
-Hugging Face supports generating embeddings of arbitrary length documents of text using Sentence Transformer library. Example of how to generate embeddings using hugging face is given below:
-
-
-
-```python main.py
-from embedchain import App
-
-# load embedding model configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: huggingface
- config:
- model: 'google/flan-t5-xxl'
- temperature: 0.5
- max_tokens: 1000
- top_p: 0.5
- stream: false
-
-embedder:
- provider: huggingface
- config:
- model: 'sentence-transformers/all-mpnet-base-v2'
- model_kwargs:
- trust_remote_code: True # Only use if you trust your embedder
-```
-
-
-
-## Vertex AI
-
-Embedchain supports Google's VertexAI embeddings model through a simple interface. You just have to pass the `model_name` in the config yaml and it would work out of the box.
-
-
-
-```python main.py
-from embedchain import App
-
-# load embedding model configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: vertexai
- config:
- model: 'chat-bison'
- temperature: 0.5
- top_p: 0.5
-
-embedder:
- provider: vertexai
- config:
- model: 'textembedding-gecko'
-```
-
-
-
-## NVIDIA AI
-
-[NVIDIA AI Foundation Endpoints](https://www.nvidia.com/en-us/ai-data-science/foundation-models/) let you quickly use NVIDIA's AI models, such as Mixtral 8x7B, Llama 2 etc, through our API. These models are available in the [NVIDIA NGC catalog](https://catalog.ngc.nvidia.com/ai-foundation-models), fully optimized and ready to use on NVIDIA's AI platform. They are designed for high speed and easy customization, ensuring smooth performance on any accelerated setup.
-
-
-### Usage
-
-In order to use embedding models and LLMs from NVIDIA AI, create an account on [NVIDIA NGC Service](https://catalog.ngc.nvidia.com/).
-
-Generate an API key from their dashboard. Set the API key as `NVIDIA_API_KEY` environment variable. Note that the `NVIDIA_API_KEY` will start with `nvapi-`.
-
-Below is an example of how to use LLM model and embedding model from NVIDIA AI:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ['NVIDIA_API_KEY'] = 'nvapi-xxxx'
-
-config = {
- "app": {
- "config": {
- "id": "my-app",
- },
- },
- "llm": {
- "provider": "nvidia",
- "config": {
- "model": "nemotron_steerlm_8b",
- },
- },
- "embedder": {
- "provider": "nvidia",
- "config": {
- "model": "nvolveqa_40k",
- "vector_dimension": 1024,
- },
- },
-}
-
-app = App.from_config(config=config)
-
-app.add("https://www.forbes.com/profile/elon-musk")
-answer = app.query("What is the net worth of Elon Musk today?")
-# Answer: The net worth of Elon Musk is subject to fluctuations based on the market value of his holdings in various companies.
-# As of March 1, 2024, his net worth is estimated to be approximately $210 billion. However, this figure can change rapidly due to stock market fluctuations and other factors.
-# Additionally, his net worth may include other assets such as real estate and art, which are not reflected in his stock portfolio.
-```
-
-
-
-## Cohere
-
-To use embedding models and LLMs from COHERE, create an account on [COHERE](https://dashboard.cohere.com/welcome/login?redirect_uri=%2Fapi-keys).
-
-Generate an API key from their dashboard. Set the API key as `COHERE_API_KEY` environment variable.
-
-Once you have obtained the key, you can use it like this:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ['COHERE_API_KEY'] = 'xxx'
-
-# load embedding model configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-embedder:
- provider: cohere
- config:
- model: 'embed-english-light-v3.0'
-```
-
-
-
-* Cohere has few embedding models: `embed-english-v3.0`, `embed-multilingual-v3.0`, `embed-multilingual-light-v3.0`, `embed-english-v2.0`, `embed-english-light-v2.0` and `embed-multilingual-v2.0`. Embedchain supports all these models. Below you can find YAML config for all:
-
-
-
-```yaml embed-english-v3.0.yaml
-embedder:
- provider: cohere
- config:
- model: 'embed-english-v3.0'
- vector_dimension: 1024
-```
-
-```yaml embed-multilingual-v3.0.yaml
-embedder:
- provider: cohere
- config:
- model: 'embed-multilingual-v3.0'
- vector_dimension: 1024
-```
-
-```yaml embed-multilingual-light-v3.0.yaml
-embedder:
- provider: cohere
- config:
- model: 'embed-multilingual-light-v3.0'
- vector_dimension: 384
-```
-
-```yaml embed-english-v2.0.yaml
-embedder:
- provider: cohere
- config:
- model: 'embed-english-v2.0'
- vector_dimension: 4096
-```
-
-```yaml embed-english-light-v2.0.yaml
-embedder:
- provider: cohere
- config:
- model: 'embed-english-light-v2.0'
- vector_dimension: 1024
-```
-
-```yaml embed-multilingual-v2.0.yaml
-embedder:
- provider: cohere
- config:
- model: 'embed-multilingual-v2.0'
- vector_dimension: 768
-```
-
-
-
-## Ollama
-
-Ollama enables the use of embedding models, allowing you to generate high-quality embeddings directly on your local machine. Make sure to install [Ollama](https://ollama.com/download) and keep it running before using the embedding model.
-
-You can find the list of models at [Ollama Embedding Models](https://ollama.com/blog/embedding-models).
-
-Below is an example of how to use embedding model Ollama:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-# load embedding model configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-embedder:
- provider: ollama
- config:
- model: 'all-minilm:latest'
-```
-
-
-
-## Clarifai
-
-Install related dependencies using the following command:
-
-```bash
-pip install --upgrade 'embedchain[clarifai]'
-```
-
-set the `CLARIFAI_PAT` as environment variable which you can find in the [security page](https://clarifai.com/settings/security). Optionally you can also pass the PAT key as parameters in LLM/Embedder class.
-
-Now you are all set with exploring Embedchain.
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["CLARIFAI_PAT"] = "XXX"
-
-# load llm and embedder configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-
-#Now let's add some data.
-app.add("https://www.forbes.com/profile/elon-musk")
-
-#Query the app
-response = app.query("what college degrees does elon musk have?")
-```
-Head to [Clarifai Platform](https://clarifai.com/explore/models?page=1&perPage=24&filterData=%5B%7B%22field%22%3A%22output_fields%22%2C%22value%22%3A%5B%22embeddings%22%5D%7D%5D) to explore all the State of the Art embedding models available to use.
-For passing LLM model inference parameters use `model_kwargs` argument in the config file. Also you can use `api_key` argument to pass `CLARIFAI_PAT` in the config.
-
-```yaml config.yaml
-llm:
- provider: clarifai
- config:
- model: "https://clarifai.com/mistralai/completion/models/mistral-7B-Instruct"
- model_kwargs:
- temperature: 0.5
- max_tokens: 1000
-embedder:
- provider: clarifai
- config:
- model: "https://clarifai.com/clarifai/main/models/BAAI-bge-base-en-v15"
-```
-
\ No newline at end of file
diff --git a/embedchain/docs/components/evaluation.mdx b/embedchain/docs/components/evaluation.mdx
deleted file mode 100644
index c1143d2ec..000000000
--- a/embedchain/docs/components/evaluation.mdx
+++ /dev/null
@@ -1,275 +0,0 @@
----
-title: 🔬 Evaluation
----
-
-## Overview
-
-We provide out-of-the-box evaluation metrics for your RAG application. You can use them to evaluate your RAG applications and compare against different settings of your production RAG application.
-
-Currently, we provide support for following evaluation metrics:
-
-
-
-
-
-
-
-
-## Quickstart
-
-Here is a basic example of running evaluation:
-
-```python example.py
-from embedchain import App
-
-app = App()
-
-# Add data sources
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# Run evaluation
-app.evaluate(["What is the net worth of Elon Musk?", "How many companies Elon Musk owns?"])
-# {'answer_relevancy': 0.9987286412340826, 'groundedness': 1.0, 'context_relevancy': 0.3571428571428571}
-```
-
-Under the hood, Embedchain does the following:
-
-1. Runs semantic search in the vector database and fetches context
-2. LLM call with question, context to fetch the answer
-3. Run evaluation on following metrics: `context relevancy`, `groundedness`, and `answer relevancy` and return result
-
-## Advanced Usage
-
-We use OpenAI's `gpt-4` model as default LLM model for automatic evaluation. Hence, we require you to set `OPENAI_API_KEY` as an environment variable.
-
-### Step-1: Create dataset
-
-In order to evaluate your RAG application, you have to setup a dataset. A data point in the dataset consists of `questions`, `contexts`, `answer`. Here is an example of how to create a dataset for evaluation:
-
-```python
-from embedchain.utils.eval import EvalData
-
-data = [
- {
- "question": "What is the net worth of Elon Musk?",
- "contexts": [
- "Elon Musk PROFILEElon MuskCEO, ...",
- "a Twitter poll on whether the journalists' ...",
- "2016 and run by Jared Birchall.[335]...",
- ],
- "answer": "As of the information provided, Elon Musk's net worth is $241.6 billion.",
- },
- {
- "question": "which companies does Elon Musk own?",
- "contexts": [
- "of December 2023[update], ...",
- "ThielCofounderView ProfileTeslaHolds ...",
- "Elon Musk PROFILEElon MuskCEO, ...",
- ],
- "answer": "Elon Musk owns several companies, including Tesla, SpaceX, Neuralink, and The Boring Company.",
- },
-]
-
-dataset = []
-
-for d in data:
- eval_data = EvalData(question=d["question"], contexts=d["contexts"], answer=d["answer"])
- dataset.append(eval_data)
-```
-
-### Step-2: Run evaluation
-
-Once you have created your dataset, you can run evaluation on the dataset by picking the metric you want to run evaluation on.
-
-For example, you can run evaluation on context relevancy metric using the following code:
-
-```python
-from embedchain.evaluation.metrics import ContextRelevance
-metric = ContextRelevance()
-score = metric.evaluate(dataset)
-print(score)
-```
-
-You can choose a different metric or write your own to run evaluation on. You can check the following links:
-
-- [Context Relevancy](#context_relevancy)
-- [Answer relenvancy](#answer_relevancy)
-- [Groundedness](#groundedness)
-- [Build your own metric](#custom_metric)
-
-## Metrics
-
-### Context Relevancy
-
-Context relevancy is a metric to determine "how relevant the context is to the question". We use OpenAI's `gpt-4` model to determine the relevancy of the context. We achieve this by prompting the model with the question and the context and asking it to return relevant sentences from the context. We then use the following formula to determine the score:
-
-```
-context_relevance_score = num_relevant_sentences_in_context / num_of_sentences_in_context
-```
-
-#### Examples
-
-You can run the context relevancy evaluation with the following simple code:
-
-```python
-from embedchain.evaluation.metrics import ContextRelevance
-
-metric = ContextRelevance()
-score = metric.evaluate(dataset) # 'dataset' is definted in the create dataset section
-print(score)
-# 0.27975528364849833
-```
-
-In the above example, we used sensible defaults for the evaluation. However, you can also configure the evaluation metric as per your needs using the `ContextRelevanceConfig` class.
-
-Here is a more advanced example of how to pass a custom evaluation config for evaluating on context relevance metric:
-
-```python
-from embedchain.config.evaluation.base import ContextRelevanceConfig
-from embedchain.evaluation.metrics import ContextRelevance
-
-eval_config = ContextRelevanceConfig(model="gpt-4", api_key="sk-xxx", language="en")
-metric = ContextRelevance(config=eval_config)
-metric.evaluate(dataset)
-```
-
-#### `ContextRelevanceConfig`
-
-
- The model to use for the evaluation. Defaults to `gpt-4`. We only support openai's models for now.
-
-
- The openai api key to use for the evaluation. Defaults to `None`. If not provided, we will use the `OPENAI_API_KEY` environment variable.
-
-
- The language of the dataset being evaluated. We need this to determine the understand the context provided in the dataset. Defaults to `en`.
-
-
- The prompt to extract the relevant sentences from the context. Defaults to `CONTEXT_RELEVANCY_PROMPT`, which can be found at `embedchain.config.evaluation.base` path.
-
-
-
-### Answer Relevancy
-
-Answer relevancy is a metric to determine how relevant the answer is to the question. We prompt the model with the answer and asking it to generate questions from the answer. We then use the cosine similarity between the generated questions and the original question to determine the score.
-
-```
-answer_relevancy_score = mean(cosine_similarity(generated_questions, original_question))
-```
-
-#### Examples
-
-You can run the answer relevancy evaluation with the following simple code:
-
-```python
-from embedchain.evaluation.metrics import AnswerRelevance
-
-metric = AnswerRelevance()
-score = metric.evaluate(dataset)
-print(score)
-# 0.9505334177461916
-```
-
-In the above example, we used sensible defaults for the evaluation. However, you can also configure the evaluation metric as per your needs using the `AnswerRelevanceConfig` class. Here is a more advanced example where you can provide your own evaluation config:
-
-```python
-from embedchain.config.evaluation.base import AnswerRelevanceConfig
-from embedchain.evaluation.metrics import AnswerRelevance
-
-eval_config = AnswerRelevanceConfig(
- model='gpt-4',
- embedder="text-embedding-ada-002",
- api_key="sk-xxx",
- num_gen_questions=2
-)
-metric = AnswerRelevance(config=eval_config)
-score = metric.evaluate(dataset)
-```
-
-#### `AnswerRelevanceConfig`
-
-
- The model to use for the evaluation. Defaults to `gpt-4`. We only support openai's models for now.
-
-
- The embedder to use for embedding the text. Defaults to `text-embedding-ada-002`. We only support openai's embedders for now.
-
-
- The openai api key to use for the evaluation. Defaults to `None`. If not provided, we will use the `OPENAI_API_KEY` environment variable.
-
-
- The number of questions to generate for each answer. We use the generated questions to compare the similarity with the original question to determine the score. Defaults to `1`.
-
-
- The prompt to extract the `num_gen_questions` number of questions from the provided answer. Defaults to `ANSWER_RELEVANCY_PROMPT`, which can be found at `embedchain.config.evaluation.base` path.
-
-
-## Groundedness
-
-Groundedness is a metric to determine how grounded the answer is to the context. We use OpenAI's `gpt-4` model to determine the groundedness of the answer. We achieve this by prompting the model with the answer and asking it to generate claims from the answer. We then again prompt the model with the context and the generated claims to determine the verdict on the claims. We then use the following formula to determine the score:
-
-```
-groundedness_score = (sum of all verdicts) / (total # of claims)
-```
-
-You can run the groundedness evaluation with the following simple code:
-
-```python
-from embedchain.evaluation.metrics import Groundedness
-metric = Groundedness()
-score = metric.evaluate(dataset) # dataset from above
-print(score)
-# 1.0
-```
-
-In the above example, we used sensible defaults for the evaluation. However, you can also configure the evaluation metric as per your needs using the `GroundednessConfig` class. Here is a more advanced example where you can configure the evaluation config:
-
-```python
-from embedchain.config.evaluation.base import GroundednessConfig
-from embedchain.evaluation.metrics import Groundedness
-
-eval_config = GroundednessConfig(model='gpt-4', api_key="sk-xxx")
-metric = Groundedness(config=eval_config)
-score = metric.evaluate(dataset)
-```
-
-
-#### `GroundednessConfig`
-
-
- The model to use for the evaluation. Defaults to `gpt-4`. We only support openai's models for now.
-
-
- The openai api key to use for the evaluation. Defaults to `None`. If not provided, we will use the `OPENAI_API_KEY` environment variable.
-
-
- The prompt to extract the claims from the provided answer. Defaults to `GROUNDEDNESS_ANSWER_CLAIMS_PROMPT`, which can be found at `embedchain.config.evaluation.base` path.
-
-
- The prompt to get verdicts on the claims from the answer from the given context. Defaults to `GROUNDEDNESS_CLAIMS_INFERENCE_PROMPT`, which can be found at `embedchain.config.evaluation.base` path.
-
-
-## Custom
-
-You can also create your own evaluation metric by extending the `BaseMetric` class. You can find the source code for the existing metrics at `embedchain.evaluation.metrics` path.
-
-
-You must provide the `name` of your custom metric in the `__init__` method of your class. This name will be used to identify your metric in the evaluation report.
-
-
-```python
-from typing import Optional
-
-from embedchain.config.base_config import BaseConfig
-from embedchain.evaluation.metrics import BaseMetric
-from embedchain.utils.eval import EvalData
-
-class MyCustomMetric(BaseMetric):
- def __init__(self, config: Optional[BaseConfig] = None):
- super().__init__(name="my_custom_metric")
-
- def evaluate(self, dataset: list[EvalData]):
- score = 0.0
- # write your evaluation logic here
- return score
-```
diff --git a/embedchain/docs/components/introduction.mdx b/embedchain/docs/components/introduction.mdx
deleted file mode 100644
index 3f9122b5d..000000000
--- a/embedchain/docs/components/introduction.mdx
+++ /dev/null
@@ -1,13 +0,0 @@
----
-title: 🧩 Introduction
----
-
-## Overview
-
-You can configure following components
-
-* [Data Source](/components/data-sources/overview)
-* [LLM](/components/llms)
-* [Embedding Model](/components/embedding-models)
-* [Vector Database](/components/vector-databases)
-* [Evaluation](/components/evaluation)
diff --git a/embedchain/docs/components/llms.mdx b/embedchain/docs/components/llms.mdx
deleted file mode 100644
index 183b8cd3f..000000000
--- a/embedchain/docs/components/llms.mdx
+++ /dev/null
@@ -1,899 +0,0 @@
----
-title: 🤖 Large language models (LLMs)
----
-
-## Overview
-
-Embedchain comes with built-in support for various popular large language models. We handle the complexity of integrating these models for you, allowing you to easily customize your language model interactions through a user-friendly interface.
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-## OpenAI
-
-To use OpenAI LLM models, you have to set the `OPENAI_API_KEY` environment variable. You can obtain the OpenAI API key from the [OpenAI Platform](https://platform.openai.com/account/api-keys).
-
-Once you have obtained the key, you can use it like this:
-
-```python
-import os
-from embedchain import App
-
-os.environ['OPENAI_API_KEY'] = 'xxx'
-
-app = App()
-app.add("https://en.wikipedia.org/wiki/OpenAI")
-app.query("What is OpenAI?")
-```
-
-If you are looking to configure the different parameters of the LLM, you can do so by loading the app using a [yaml config](https://github.com/embedchain/embedchain/blob/main/configs/chroma.yaml) file.
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ['OPENAI_API_KEY'] = 'xxx'
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: openai
- config:
- model: 'gpt-4o-mini'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-```
-
-
-### Function Calling
-Embedchain supports OpenAI [Function calling](https://platform.openai.com/docs/guides/function-calling) with a single function. It accepts inputs in accordance with the [Langchain interface](https://python.langchain.com/docs/modules/model_io/chat/function_calling#legacy-args-functions-and-function_call).
-
-
- ```python
- from pydantic import BaseModel
-
- class multiply(BaseModel):
- """Multiply two integers together."""
-
- a: int = Field(..., description="First integer")
- b: int = Field(..., description="Second integer")
- ```
-
-
-
- ```python
- def multiply(a: int, b: int) -> int:
- """Multiply two integers together.
-
- Args:
- a: First integer
- b: Second integer
- """
- return a * b
- ```
-
-
- ```python
- multiply = {
- "type": "function",
- "function": {
- "name": "multiply",
- "description": "Multiply two integers together.",
- "parameters": {
- "type": "object",
- "properties": {
- "a": {
- "description": "First integer",
- "type": "integer"
- },
- "b": {
- "description": "Second integer",
- "type": "integer"
- }
- },
- "required": [
- "a",
- "b"
- ]
- }
- }
- }
- ```
-
-
-With any of the previous inputs, the OpenAI LLM can be queried to provide the appropriate arguments for the function.
-
-```python
-import os
-from embedchain import App
-from embedchain.llm.openai import OpenAILlm
-
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-llm = OpenAILlm(tools=multiply)
-app = App(llm=llm)
-
-result = app.query("What is the result of 125 multiplied by fifteen?")
-```
-
-## Google AI
-
-To use Google AI model, you have to set the `GOOGLE_API_KEY` environment variable. You can obtain the Google API key from the [Google Maker Suite](https://makersuite.google.com/app/apikey)
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["GOOGLE_API_KEY"] = "xxx"
-
-app = App.from_config(config_path="config.yaml")
-
-app.add("https://www.forbes.com/profile/elon-musk")
-
-response = app.query("What is the net worth of Elon Musk?")
-if app.llm.config.stream: # if stream is enabled, response is a generator
- for chunk in response:
- print(chunk)
-else:
- print(response)
-```
-
-```yaml config.yaml
-llm:
- provider: google
- config:
- model: gemini-pro
- max_tokens: 1000
- temperature: 0.5
- top_p: 1
- stream: false
-
-embedder:
- provider: google
- config:
- model: 'models/embedding-001'
- task_type: "retrieval_document"
- title: "Embeddings for Embedchain"
-```
-
-
-## Azure OpenAI
-
-To use Azure OpenAI model, you have to set some of the azure openai related environment variables as given in the code block below:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["OPENAI_API_TYPE"] = "azure"
-os.environ["AZURE_OPENAI_ENDPOINT"] = "https://xxx.openai.azure.com/"
-os.environ["AZURE_OPENAI_KEY"] = "xxx"
-os.environ["OPENAI_API_VERSION"] = "xxx"
-
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: azure_openai
- config:
- model: gpt-4o-mini
- deployment_name: your_llm_deployment_name
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-embedder:
- provider: azure_openai
- config:
- model: text-embedding-ada-002
- deployment_name: you_embedding_model_deployment_name
-```
-
-
-You can find the list of models and deployment name on the [Azure OpenAI Platform](https://oai.azure.com/portal).
-
-## Anthropic
-
-To use anthropic's model, please set the `ANTHROPIC_API_KEY` which you find on their [Account Settings Page](https://console.anthropic.com/account/keys).
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["ANTHROPIC_API_KEY"] = "xxx"
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: anthropic
- config:
- model: 'claude-instant-1'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-```
-
-
-
-## Cohere
-
-Install related dependencies using the following command:
-
-```bash
-pip install --upgrade 'embedchain[cohere]'
-```
-
-Set the `COHERE_API_KEY` as environment variable which you can find on their [Account settings page](https://dashboard.cohere.com/api-keys).
-
-Once you have the API key, you are all set to use it with Embedchain.
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["COHERE_API_KEY"] = "xxx"
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: cohere
- config:
- model: large
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
-```
-
-
-
-## Together
-
-Install related dependencies using the following command:
-
-```bash
-pip install --upgrade 'embedchain[together]'
-```
-
-Set the `TOGETHER_API_KEY` as environment variable which you can find on their [Account settings page](https://api.together.xyz/settings/api-keys).
-
-Once you have the API key, you are all set to use it with Embedchain.
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["TOGETHER_API_KEY"] = "xxx"
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: together
- config:
- model: togethercomputer/RedPajama-INCITE-7B-Base
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
-```
-
-
-
-## Ollama
-
-Setup Ollama using https://github.com/jmorganca/ollama
-
-
-
-```python main.py
-import os
-os.environ["OLLAMA_HOST"] = "http://127.0.0.1:11434"
-from embedchain import App
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: ollama
- config:
- model: 'llama2'
- temperature: 0.5
- top_p: 1
- stream: true
- base_url: 'http://localhost:11434'
-embedder:
- provider: ollama
- config:
- model: znbang/bge:small-en-v1.5-q8_0
- base_url: http://localhost:11434
-
-```
-
-
-
-
-## vLLM
-
-Setup vLLM by following instructions given in [their docs](https://docs.vllm.ai/en/latest/getting_started/installation.html).
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: vllm
- config:
- model: 'meta-llama/Llama-2-70b-hf'
- temperature: 0.5
- top_p: 1
- top_k: 10
- stream: true
- trust_remote_code: true
-```
-
-
-
-## Clarifai
-
-Install related dependencies using the following command:
-
-```bash
-pip install --upgrade 'embedchain[clarifai]'
-```
-
-set the `CLARIFAI_PAT` as environment variable which you can find in the [security page](https://clarifai.com/settings/security). Optionally you can also pass the PAT key as parameters in LLM/Embedder class.
-
-Now you are all set with exploring Embedchain.
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["CLARIFAI_PAT"] = "XXX"
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-
-#Now let's add some data.
-app.add("https://www.forbes.com/profile/elon-musk")
-
-#Query the app
-response = app.query("what college degrees does elon musk have?")
-```
-Head to [Clarifai Platform](https://clarifai.com/explore/models?page=1&perPage=24&filterData=%5B%7B%22field%22%3A%22use_cases%22%2C%22value%22%3A%5B%22llm%22%5D%7D%5D) to browse various State-of-the-Art LLM models for your use case.
-For passing model inference parameters use `model_kwargs` argument in the config file. Also you can use `api_key` argument to pass `CLARIFAI_PAT` in the config.
-
-```yaml config.yaml
-llm:
- provider: clarifai
- config:
- model: "https://clarifai.com/mistralai/completion/models/mistral-7B-Instruct"
- model_kwargs:
- temperature: 0.5
- max_tokens: 1000
-embedder:
- provider: clarifai
- config:
- model: "https://clarifai.com/clarifai/main/models/BAAI-bge-base-en-v15"
-```
-
-
-
-## GPT4ALL
-
-Install related dependencies using the following command:
-
-```bash
-pip install --upgrade 'embedchain[opensource]'
-```
-
-GPT4all is a free-to-use, locally running, privacy-aware chatbot. No GPU or internet required. You can use this with Embedchain using the following code:
-
-
-
-```python main.py
-from embedchain import App
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: gpt4all
- config:
- model: 'orca-mini-3b-gguf2-q4_0.gguf'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-embedder:
- provider: gpt4all
-```
-
-
-
-## JinaChat
-
-First, set `JINACHAT_API_KEY` in environment variable which you can obtain from [their platform](https://chat.jina.ai/api).
-
-Once you have the key, load the app using the config yaml file:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["JINACHAT_API_KEY"] = "xxx"
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: jina
- config:
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-```
-
-
-
-## Hugging Face
-
-
-Install related dependencies using the following command:
-
-```bash
-pip install --upgrade 'embedchain[huggingface-hub]'
-```
-
-First, set `HUGGINGFACE_ACCESS_TOKEN` in environment variable which you can obtain from [their platform](https://huggingface.co/settings/tokens).
-
-You can load the LLMs from Hugging Face using three ways:
-
-- [Hugging Face Hub](#hugging-face-hub)
-- [Hugging Face Local Pipelines](#hugging-face-local-pipelines)
-- [Hugging Face Inference Endpoint](#hugging-face-inference-endpoint)
-
-### Hugging Face Hub
-
-To load the model from Hugging Face Hub, use the following code:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["HUGGINGFACE_ACCESS_TOKEN"] = "xxx"
-
-config = {
- "app": {"config": {"id": "my-app"}},
- "llm": {
- "provider": "huggingface",
- "config": {
- "model": "bigscience/bloom-1b7",
- "top_p": 0.5,
- "max_length": 200,
- "temperature": 0.1,
- },
- },
-}
-
-app = App.from_config(config=config)
-```
-
-
-### Hugging Face Local Pipelines
-
-If you want to load the locally downloaded model from Hugging Face, you can do so by following the code provided below:
-
-
-```python main.py
-from embedchain import App
-
-config = {
- "app": {"config": {"id": "my-app"}},
- "llm": {
- "provider": "huggingface",
- "config": {
- "model": "Trendyol/Trendyol-LLM-7b-chat-v0.1",
- "local": True, # Necessary if you want to run model locally
- "top_p": 0.5,
- "max_tokens": 1000,
- "temperature": 0.1,
- },
- }
-}
-app = App.from_config(config=config)
-```
-
-
-### Hugging Face Inference Endpoint
-
-You can also use [Hugging Face Inference Endpoints](https://huggingface.co/docs/inference-endpoints/index#-inference-endpoints) to access custom endpoints. First, set the `HUGGINGFACE_ACCESS_TOKEN` as above.
-
-Then, load the app using the config yaml file:
-
-
-
-```python main.py
-from embedchain import App
-
-config = {
- "app": {"config": {"id": "my-app"}},
- "llm": {
- "provider": "huggingface",
- "config": {
- "endpoint": "https://api-inference.huggingface.co/models/gpt2",
- "model_params": {"temprature": 0.1, "max_new_tokens": 100}
- },
- },
-}
-app = App.from_config(config=config)
-
-```
-
-
-Currently only supports `text-generation` and `text2text-generation` for now [[ref](https://api.python.langchain.com/en/latest/llms/langchain_community.llms.huggingface_endpoint.HuggingFaceEndpoint.html?highlight=huggingfaceendpoint#)].
-
-See langchain's [hugging face endpoint](https://python.langchain.com/docs/integrations/chat/huggingface#huggingfaceendpoint) for more information.
-
-## Llama2
-
-Llama2 is integrated through [Replicate](https://replicate.com/). Set `REPLICATE_API_TOKEN` in environment variable which you can obtain from [their platform](https://replicate.com/account/api-tokens).
-
-Once you have the token, load the app using the config yaml file:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["REPLICATE_API_TOKEN"] = "xxx"
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: llama2
- config:
- model: 'a16z-infra/llama13b-v2-chat:df7690f1994d94e96ad9d568eac121aecf50684a0b0963b25a41cc40061269e5'
- temperature: 0.5
- max_tokens: 1000
- top_p: 0.5
- stream: false
-```
-
-
-## Vertex AI
-
-Setup Google Cloud Platform application credentials by following the instruction on [GCP](https://cloud.google.com/docs/authentication/external/set-up-adc). Once setup is done, use the following code to create an app using VertexAI as provider:
-
-
-
-```python main.py
-from embedchain import App
-
-# load llm configuration from config.yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: vertexai
- config:
- model: 'chat-bison'
- temperature: 0.5
- top_p: 0.5
-```
-
-
-
-## Mistral AI
-
-Obtain the Mistral AI api key from their [console](https://console.mistral.ai/).
-
-
-
- ```python main.py
-os.environ["MISTRAL_API_KEY"] = "xxx"
-
-app = App.from_config(config_path="config.yaml")
-
-app.add("https://www.forbes.com/profile/elon-musk")
-
-response = app.query("what is the net worth of Elon Musk?")
-# As of January 16, 2024, Elon Musk's net worth is $225.4 billion.
-
-response = app.chat("which companies does elon own?")
-# Elon Musk owns Tesla, SpaceX, Boring Company, Twitter, and X.
-
-response = app.chat("what question did I ask you already?")
-# You have asked me several times already which companies Elon Musk owns, specifically Tesla, SpaceX, Boring Company, Twitter, and X.
-```
-
-```yaml config.yaml
-llm:
- provider: mistralai
- config:
- model: mistral-tiny
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
-embedder:
- provider: mistralai
- config:
- model: mistral-embed
-```
-
-
-
-## AWS Bedrock
-
-### Setup
-- Before using the AWS Bedrock LLM, make sure you have the appropriate model access from [Bedrock Console](https://us-east-1.console.aws.amazon.com/bedrock/home?region=us-east-1#/modelaccess).
-- You will also need to authenticate the `boto3` client by using a method in the [AWS documentation](https://boto3.amazonaws.com/v1/documentation/api/latest/guide/credentials.html#configuring-credentials)
-- You can optionally export an `AWS_REGION`
-
-
-### Usage
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["AWS_REGION"] = "us-west-2"
-
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-llm:
- provider: aws_bedrock
- config:
- model: amazon.titan-text-express-v1
- # check notes below for model_kwargs
- model_kwargs:
- temperature: 0.5
- topP: 1
- maxTokenCount: 1000
-```
-
-
-
-
- The model arguments are different for each providers. Please refer to the [AWS Bedrock Documentation](https://us-east-1.console.aws.amazon.com/bedrock/home?region=us-east-1#/providers) to find the appropriate arguments for your model.
-
-
-
-
-## Groq
-
-[Groq](https://groq.com/) is the creator of the world's first Language Processing Unit (LPU), providing exceptional speed performance for AI workloads running on their LPU Inference Engine.
-
-
-### Usage
-
-In order to use LLMs from Groq, go to their [platform](https://console.groq.com/keys) and get the API key.
-
-Set the API key as `GROQ_API_KEY` environment variable or pass in your app configuration to use the model as given below in the example.
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-# Set your API key here or pass as the environment variable
-groq_api_key = "gsk_xxxx"
-
-config = {
- "llm": {
- "provider": "groq",
- "config": {
- "model": "mixtral-8x7b-32768",
- "api_key": groq_api_key,
- "stream": True
- }
- }
-}
-
-app = App.from_config(config=config)
-# Add your data source here
-app.add("https://docs.embedchain.ai/sitemap.xml", data_type="sitemap")
-app.query("Write a poem about Embedchain")
-
-# In the realm of data, vast and wide,
-# Embedchain stands with knowledge as its guide.
-# A platform open, for all to try,
-# Building bots that can truly fly.
-
-# With REST API, data in reach,
-# Deployment a breeze, as easy as a speech.
-# Updating data sources, anytime, anyday,
-# Embedchain's power, never sway.
-
-# A knowledge base, an assistant so grand,
-# Connecting to platforms, near and far.
-# Discord, WhatsApp, Slack, and more,
-# Embedchain's potential, never a bore.
-```
-
-
-## NVIDIA AI
-
-[NVIDIA AI Foundation Endpoints](https://www.nvidia.com/en-us/ai-data-science/foundation-models/) let you quickly use NVIDIA's AI models, such as Mixtral 8x7B, Llama 2 etc, through our API. These models are available in the [NVIDIA NGC catalog](https://catalog.ngc.nvidia.com/ai-foundation-models), fully optimized and ready to use on NVIDIA's AI platform. They are designed for high speed and easy customization, ensuring smooth performance on any accelerated setup.
-
-
-### Usage
-
-In order to use LLMs from NVIDIA AI, create an account on [NVIDIA NGC Service](https://catalog.ngc.nvidia.com/).
-
-Generate an API key from their dashboard. Set the API key as `NVIDIA_API_KEY` environment variable. Note that the `NVIDIA_API_KEY` will start with `nvapi-`.
-
-Below is an example of how to use LLM model and embedding model from NVIDIA AI:
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ['NVIDIA_API_KEY'] = 'nvapi-xxxx'
-
-config = {
- "app": {
- "config": {
- "id": "my-app",
- },
- },
- "llm": {
- "provider": "nvidia",
- "config": {
- "model": "nemotron_steerlm_8b",
- },
- },
- "embedder": {
- "provider": "nvidia",
- "config": {
- "model": "nvolveqa_40k",
- "vector_dimension": 1024,
- },
- },
-}
-
-app = App.from_config(config=config)
-
-app.add("https://www.forbes.com/profile/elon-musk")
-answer = app.query("What is the net worth of Elon Musk today?")
-# Answer: The net worth of Elon Musk is subject to fluctuations based on the market value of his holdings in various companies.
-# As of March 1, 2024, his net worth is estimated to be approximately $210 billion. However, this figure can change rapidly due to stock market fluctuations and other factors.
-# Additionally, his net worth may include other assets such as real estate and art, which are not reflected in his stock portfolio.
-```
-
-
-## Token Usage
-
-You can get the cost of the query by setting `token_usage` to `True` in the config file. This will return the token details: `prompt_tokens`, `completion_tokens`, `total_tokens`, `total_cost`, `cost_currency`.
-The list of paid LLMs that support token usage are:
-- OpenAI
-- Vertex AI
-- Anthropic
-- Cohere
-- Together
-- Groq
-- Mistral AI
-- NVIDIA AI
-
-Here is an example of how to use token usage:
-
-
-```python main.py
-os.environ["OPENAI_API_KEY"] = "xxx"
-
-app = App.from_config(config_path="config.yaml")
-
-app.add("https://www.forbes.com/profile/elon-musk")
-
-response = app.query("what is the net worth of Elon Musk?")
-# {'answer': 'Elon Musk's net worth is $209.9 billion as of 6/9/24.',
-# 'usage': {'prompt_tokens': 1228,
-# 'completion_tokens': 21,
-# 'total_tokens': 1249,
-# 'total_cost': 0.001884,
-# 'cost_currency': 'USD'}
-# }
-
-
-response = app.chat("Which companies did Elon Musk found?")
-# {'answer': 'Elon Musk founded six companies, including Tesla, which is an electric car maker, SpaceX, a rocket producer, and the Boring Company, a tunneling startup.',
-# 'usage': {'prompt_tokens': 1616,
-# 'completion_tokens': 34,
-# 'total_tokens': 1650,
-# 'total_cost': 0.002492,
-# 'cost_currency': 'USD'}
-# }
-```
-
-```yaml config.yaml
-llm:
- provider: openai
- config:
- model: gpt-4o-mini
- temperature: 0.5
- max_tokens: 1000
- token_usage: true
-```
-
-
-If a model is missing and you'd like to add it to `model_prices_and_context_window.json`, please feel free to open a PR.
-
-
-
-
diff --git a/embedchain/docs/components/retrieval-methods.mdx b/embedchain/docs/components/retrieval-methods.mdx
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/docs/components/vector-databases.mdx b/embedchain/docs/components/vector-databases.mdx
deleted file mode 100644
index c889e1054..000000000
--- a/embedchain/docs/components/vector-databases.mdx
+++ /dev/null
@@ -1,20 +0,0 @@
----
-title: 🗄️ Vector databases
----
-
-## Overview
-
-Utilizing a vector database alongside Embedchain is a seamless process. All you need to do is configure it within the YAML configuration file. We've provided examples for each supported database below:
-
-
-
-
-
-
-
-
-
-
-
-
-
diff --git a/embedchain/docs/components/vector-databases/chromadb.mdx b/embedchain/docs/components/vector-databases/chromadb.mdx
deleted file mode 100644
index 783dfe890..000000000
--- a/embedchain/docs/components/vector-databases/chromadb.mdx
+++ /dev/null
@@ -1,35 +0,0 @@
----
-title: ChromaDB
----
-
-
-
-```python main.py
-from embedchain import App
-
-# load chroma configuration from yaml file
-app = App.from_config(config_path="config1.yaml")
-```
-
-```yaml config1.yaml
-vectordb:
- provider: chroma
- config:
- collection_name: 'my-collection'
- dir: db
- allow_reset: true
-```
-
-```yaml config2.yaml
-vectordb:
- provider: chroma
- config:
- collection_name: 'my-collection'
- host: localhost
- port: 5200
- allow_reset: true
-```
-
-
-
-
diff --git a/embedchain/docs/components/vector-databases/elasticsearch.mdx b/embedchain/docs/components/vector-databases/elasticsearch.mdx
deleted file mode 100644
index 0a354e65f..000000000
--- a/embedchain/docs/components/vector-databases/elasticsearch.mdx
+++ /dev/null
@@ -1,39 +0,0 @@
----
-title: Elasticsearch
----
-
-Install related dependencies using the following command:
-
-```bash
-pip install --upgrade 'embedchain[elasticsearch]'
-```
-
-
-You can configure the Elasticsearch connection by providing either `es_url` or `cloud_id`. If you are using the Elasticsearch Service on Elastic Cloud, you can find the `cloud_id` on the [Elastic Cloud dashboard](https://cloud.elastic.co/deployments).
-
-
-You can authorize the connection to Elasticsearch by providing either `basic_auth`, `api_key`, or `bearer_auth`.
-
-
-
-```python main.py
-from embedchain import App
-
-# load elasticsearch configuration from yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-vectordb:
- provider: elasticsearch
- config:
- collection_name: 'es-index'
- cloud_id: 'deployment-name:xxxx'
- basic_auth:
- - elastic
- -
- verify_certs: false
-```
-
-
-
diff --git a/embedchain/docs/components/vector-databases/lancedb.mdx b/embedchain/docs/components/vector-databases/lancedb.mdx
deleted file mode 100644
index 97af57dfe..000000000
--- a/embedchain/docs/components/vector-databases/lancedb.mdx
+++ /dev/null
@@ -1,100 +0,0 @@
----
-title: LanceDB
----
-
-## Install Embedchain with LanceDB
-
-Install Embedchain, LanceDB and related dependencies using the following command:
-
-```bash
-pip install "embedchain[lancedb]"
-```
-
-LanceDB is a developer-friendly, open source database for AI. From hyper scalable vector search and advanced retrieval for RAG, to streaming training data and interactive exploration of large scale AI datasets.
-In order to use LanceDB as vector database, not need to set any key for local use.
-
-### With OPENAI
-
-
-```python main.py
-import os
-from embedchain import App
-
-# set OPENAI_API_KEY as env variable
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-# create Embedchain App and set config
-app = App.from_config(config={
- "vectordb": {
- "provider": "lancedb",
- "config": {
- "collection_name": "lancedb-index"
- }
- }
- }
-)
-
-# add data source and start query in
-app.add("https://www.forbes.com/profile/elon-musk")
-
-# query continuously
-while(True):
- question = input("Enter question: ")
- if question in ['q', 'exit', 'quit']:
- break
- answer = app.query(question)
- print(answer)
-```
-
-
-
-### With Local LLM
-
-
-```python main.py
-from embedchain import Pipeline as App
-
-# config for Embedchain App
-config = {
- 'llm': {
- 'provider': 'huggingface',
- 'config': {
- 'model': 'mistralai/Mistral-7B-v0.1',
- 'temperature': 0.1,
- 'max_tokens': 250,
- 'top_p': 0.1,
- 'stream': True
- }
- },
- 'embedder': {
- 'provider': 'huggingface',
- 'config': {
- 'model': 'sentence-transformers/all-mpnet-base-v2'
- }
- },
- 'vectordb': {
- 'provider': 'lancedb',
- 'config': {
- 'collection_name': 'lancedb-index'
- }
- }
-}
-
-app = App.from_config(config=config)
-
-# add data source and start query in
-app.add("https://www.tesla.com/ns_videos/2022-tesla-impact-report.pdf")
-
-# query continuously
-while(True):
- question = input("Enter question: ")
- if question in ['q', 'exit', 'quit']:
- break
- answer = app.query(question)
- print(answer)
-```
-
-
-
-
-
\ No newline at end of file
diff --git a/embedchain/docs/components/vector-databases/opensearch.mdx b/embedchain/docs/components/vector-databases/opensearch.mdx
deleted file mode 100644
index 8f6866977..000000000
--- a/embedchain/docs/components/vector-databases/opensearch.mdx
+++ /dev/null
@@ -1,36 +0,0 @@
----
-title: OpenSearch
----
-
-Install related dependencies using the following command:
-
-```bash
-pip install --upgrade 'embedchain[opensearch]'
-```
-
-
-
-```python main.py
-from embedchain import App
-
-# load opensearch configuration from yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-vectordb:
- provider: opensearch
- config:
- collection_name: 'my-app'
- opensearch_url: 'https://localhost:9200'
- http_auth:
- - admin
- - admin
- vector_dimension: 1536
- use_ssl: false
- verify_certs: false
-```
-
-
-
-
diff --git a/embedchain/docs/components/vector-databases/pinecone.mdx b/embedchain/docs/components/vector-databases/pinecone.mdx
deleted file mode 100644
index d21ebfeac..000000000
--- a/embedchain/docs/components/vector-databases/pinecone.mdx
+++ /dev/null
@@ -1,109 +0,0 @@
----
-title: Pinecone
----
-
-## Overview
-
-Install pinecone related dependencies using the following command:
-
-```bash
-pip install --upgrade 'pinecone-client pinecone-text'
-```
-
-In order to use Pinecone as vector database, set the environment variable `PINECONE_API_KEY` which you can find on [Pinecone dashboard](https://app.pinecone.io/).
-
-
-
-```python main.py
-from embedchain import App
-
-# Load pinecone configuration from yaml file
-app = App.from_config(config_path="pod_config.yaml")
-# Or
-app = App.from_config(config_path="serverless_config.yaml")
-```
-
-```yaml pod_config.yaml
-vectordb:
- provider: pinecone
- config:
- metric: cosine
- vector_dimension: 1536
- index_name: my-pinecone-index
- pod_config:
- environment: gcp-starter
- metadata_config:
- indexed:
- - "url"
- - "hash"
-```
-
-```yaml serverless_config.yaml
-vectordb:
- provider: pinecone
- config:
- metric: cosine
- vector_dimension: 1536
- index_name: my-pinecone-index
- serverless_config:
- cloud: aws
- region: us-west-2
-```
-
-
-
-
-
-You can find more information about Pinecone configuration [here](https://docs.pinecone.io/docs/manage-indexes#create-a-pod-based-index).
-You can also optionally provide `index_name` as a config param in yaml file to specify the index name. If not provided, the index name will be `{collection_name}-{vector_dimension}`.
-
-
-## Usage
-
-### Hybrid search
-
-Here is an example of how you can do hybrid search using Pinecone as a vector database through Embedchain.
-
-```python
-import os
-
-from embedchain import App
-
-config = {
- 'app': {
- "config": {
- "id": "ec-docs-hybrid-search"
- }
- },
- 'vectordb': {
- 'provider': 'pinecone',
- 'config': {
- 'metric': 'dotproduct',
- 'vector_dimension': 1536,
- 'index_name': 'my-index',
- 'serverless_config': {
- 'cloud': 'aws',
- 'region': 'us-west-2'
- },
- 'hybrid_search': True, # Remember to set this for hybrid search
- }
- }
-}
-
-# Initialize app
-app = App.from_config(config=config)
-
-# Add documents
-app.add("/path/to/file.pdf", data_type="pdf_file", namespace="my-namespace")
-
-# Query
-app.query("", namespace="my-namespace")
-
-# Chat
-app.chat("", namespace="my-namespace")
-```
-
-Under the hood, Embedchain fetches the relevant chunks from the documents you added by doing hybrid search on the pinecone index.
-If you have questions on how pinecone hybrid search works, please refer to their [offical documentation here](https://docs.pinecone.io/docs/hybrid-search).
-
-
diff --git a/embedchain/docs/components/vector-databases/qdrant.mdx b/embedchain/docs/components/vector-databases/qdrant.mdx
deleted file mode 100644
index cadb42e92..000000000
--- a/embedchain/docs/components/vector-databases/qdrant.mdx
+++ /dev/null
@@ -1,23 +0,0 @@
----
-title: Qdrant
----
-
-In order to use Qdrant as a vector database, set the environment variables `QDRANT_URL` and `QDRANT_API_KEY` which you can find on [Qdrant Dashboard](https://cloud.qdrant.io/).
-
-
-```python main.py
-from embedchain import App
-
-# load qdrant configuration from yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-vectordb:
- provider: qdrant
- config:
- collection_name: my_qdrant_index
-```
-
-
-
diff --git a/embedchain/docs/components/vector-databases/weaviate.mdx b/embedchain/docs/components/vector-databases/weaviate.mdx
deleted file mode 100644
index e5b1d5eda..000000000
--- a/embedchain/docs/components/vector-databases/weaviate.mdx
+++ /dev/null
@@ -1,24 +0,0 @@
----
-title: Weaviate
----
-
-
-In order to use Weaviate as a vector database, set the environment variables `WEAVIATE_ENDPOINT` and `WEAVIATE_API_KEY` which you can find on [Weaviate dashboard](https://console.weaviate.cloud/dashboard).
-
-
-```python main.py
-from embedchain import App
-
-# load weaviate configuration from yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-vectordb:
- provider: weaviate
- config:
- collection_name: my_weaviate_index
-```
-
-
-
diff --git a/embedchain/docs/components/vector-databases/zilliz.mdx b/embedchain/docs/components/vector-databases/zilliz.mdx
deleted file mode 100644
index 55c0dbaa7..000000000
--- a/embedchain/docs/components/vector-databases/zilliz.mdx
+++ /dev/null
@@ -1,39 +0,0 @@
----
-title: Zilliz
----
-
-Install related dependencies using the following command:
-
-```bash
-pip install --upgrade 'embedchain[milvus]'
-```
-
-Set the Zilliz environment variables `ZILLIZ_CLOUD_URI` and `ZILLIZ_CLOUD_TOKEN` which you can find it on their [cloud platform](https://cloud.zilliz.com/).
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ['ZILLIZ_CLOUD_URI'] = 'https://xxx.zillizcloud.com'
-os.environ['ZILLIZ_CLOUD_TOKEN'] = 'xxx'
-
-# load zilliz configuration from yaml file
-app = App.from_config(config_path="config.yaml")
-```
-
-```yaml config.yaml
-vectordb:
- provider: zilliz
- config:
- collection_name: 'zilliz_app'
- uri: https://xxxx.api.gcp-region.zillizcloud.com
- token: xxx
- vector_dim: 1536
- metric_type: L2
-```
-
-
-
-
diff --git a/embedchain/docs/contribution/dev.mdx b/embedchain/docs/contribution/dev.mdx
deleted file mode 100644
index 3ce71c25c..000000000
--- a/embedchain/docs/contribution/dev.mdx
+++ /dev/null
@@ -1,45 +0,0 @@
----
-title: '👨💻 Development'
-description: 'Contribute to Embedchain framework development'
----
-
-Thank you for your interest in contributing to the EmbedChain project! We welcome your ideas and contributions to help improve the project. Please follow the instructions below to get started:
-
-1. **Fork the repository**: Click on the "Fork" button at the top right corner of this repository page. This will create a copy of the repository in your own GitHub account.
-
-2. **Install the required dependencies**: Ensure that you have the necessary dependencies installed in your Python environment. You can do this by running the following command:
-
-```bash
-make install
-```
-
-3. **Make changes in the code**: Create a new branch in your forked repository and make your desired changes in the codebase.
-4. **Format code**: Before creating a pull request, it's important to ensure that your code follows our formatting guidelines. Run the following commands to format the code:
-
-```bash
-make lint format
-```
-
-5. **Create a pull request**: When you are ready to contribute your changes, submit a pull request to the EmbedChain repository. Provide a clear and descriptive title for your pull request, along with a detailed description of the changes you have made.
-
-## Team
-
-### Authors
-
-- Taranjeet Singh ([@taranjeetio](https://twitter.com/taranjeetio))
-- Deshraj Yadav ([@deshrajdry](https://twitter.com/deshrajdry))
-
-### Citation
-
-If you utilize this repository, please consider citing it with:
-
-```
-@misc{embedchain,
- author = {Taranjeet Singh, Deshraj Yadav},
- title = {Embechain: The Open Source RAG Framework},
- year = {2023},
- publisher = {GitHub},
- journal = {GitHub repository},
- howpublished = {\url{https://github.com/embedchain/embedchain}},
-}
-```
diff --git a/embedchain/docs/contribution/docs.mdx b/embedchain/docs/contribution/docs.mdx
deleted file mode 100644
index 7aa846ec5..000000000
--- a/embedchain/docs/contribution/docs.mdx
+++ /dev/null
@@ -1,61 +0,0 @@
----
-title: '📝 Documentation'
-description: 'Contribute to Embedchain docs'
----
-
-
- **Prerequisite** You should have installed Node.js (version 18.10.0 or
- higher).
-
-
-Step 1. Install Mintlify on your OS:
-
-
-
-```bash npm
-npm i -g mintlify
-```
-
-```bash yarn
-yarn global add mintlify
-```
-
-
-
-Step 2. Go to the `docs/` directory (where you can find `mint.json`) and run the following command:
-
-```bash
-mintlify dev
-```
-
-The documentation website is now available at `http://localhost:3000`.
-
-### Custom Ports
-
-Mintlify uses port 3000 by default. You can use the `--port` flag to customize the port Mintlify runs on. For example, use this command to run in port 3333:
-
-```bash
-mintlify dev --port 3333
-```
-
-You will see an error like this if you try to run Mintlify in a port that's already taken:
-
-```md
-Error: listen EADDRINUSE: address already in use :::3000
-```
-
-## Mintlify Versions
-
-Each CLI is linked to a specific version of Mintlify. Please update the CLI if your local website looks different than production.
-
-
-
-```bash npm
-npm i -g mintlify@latest
-```
-
-```bash yarn
-yarn global upgrade mintlify
-```
-
-
diff --git a/embedchain/docs/contribution/guidelines.mdx b/embedchain/docs/contribution/guidelines.mdx
deleted file mode 100644
index 3c5d557eb..000000000
--- a/embedchain/docs/contribution/guidelines.mdx
+++ /dev/null
@@ -1,4 +0,0 @@
----
-title: '📋 Guidelines'
-url: https://github.com/mem0ai/mem0/blob/main/embedchain/CONTRIBUTING.md
----
\ No newline at end of file
diff --git a/embedchain/docs/contribution/python.mdx b/embedchain/docs/contribution/python.mdx
deleted file mode 100644
index 47bc84c27..000000000
--- a/embedchain/docs/contribution/python.mdx
+++ /dev/null
@@ -1,4 +0,0 @@
----
-title: '🐍 Python'
-url: https://github.com/embedchain/embedchain
----
\ No newline at end of file
diff --git a/embedchain/docs/deployment/fly_io.mdx b/embedchain/docs/deployment/fly_io.mdx
deleted file mode 100644
index ed8992915..000000000
--- a/embedchain/docs/deployment/fly_io.mdx
+++ /dev/null
@@ -1,101 +0,0 @@
----
-title: 'Fly.io'
-description: 'Deploy your RAG application to fly.io platform'
----
-
-Embedchain has a nice and simple abstraction on top of the [Fly.io](https://fly.io/) tools to let developers deploy RAG application to fly.io platform seamlessly.
-
-Follow the instructions given below to deploy your first application quickly:
-
-
-## Step-1: Install flyctl command line
-
-
-```bash OSX
-brew install flyctl
-```
-
-```bash Linux
-curl -L https://fly.io/install.sh | sh
-```
-
-```bash Windows
-pwsh -Command "iwr https://fly.io/install.ps1 -useb | iex"
-```
-
-
-Once you have installed the fly.io cli tool, signup/login to their platform using the following command:
-
-
-```bash Sign up
-fly auth signup
-```
-
-```bash Sign in
-fly auth login
-```
-
-
-In case you run into issues, refer to official [fly.io docs](https://fly.io/docs/hands-on/install-flyctl/).
-
-## Step-2: Create RAG app
-
-We provide a command line utility called `ec` in embedchain that inherits the template for `fly.io` platform and help you deploy the app. Follow the instructions to create a fly.io app using the template provided:
-
-```bash Install embedchain
-pip install embedchain
-```
-
-```bash Create application
-mkdir my-rag-app
-ec create --template=fly.io
-```
-
-This will generate a directory structure like this:
-
-```bash
-├── Dockerfile
-├── app.py
-├── fly.toml
-├── .env
-├── .env.example
-├── embedchain.json
-└── requirements.txt
-```
-
-Feel free to edit the files as required.
-- `Dockerfile`: Defines the steps to setup the application
-- `app.py`: Contains API app code
-- `fly.toml`: fly.io config file
-- `.env`: Contains environment variables for production
-- `.env.example`: Contains dummy environment variables (can ignore this file)
-- `embedchain.json`: Contains embedchain specific configuration for deployment (you don't need to configure this)
-- `requirements.txt`: Contains python dependencies for your application
-
-## Step-3: Test app locally
-
-You can run the app locally by simply doing:
-
-```bash Run locally
-pip install -r requirements.txt
-ec dev
-```
-
-## Step-4: Deploy to fly.io
-
-You can deploy to fly.io using the following command:
-```bash Deploy app
-ec deploy
-```
-
-Once this step finished, it will provide you with the deployment endpoint where you can access the app live. It will look something like this (Swagger docs):
-
-You can also check the logs, monitor app status etc on their dashboard by running command `fly dashboard`.
-
-
-
-## Seeking help?
-
-If you run into issues with deployment, please feel free to reach out to us via any of the following methods:
-
-
diff --git a/embedchain/docs/deployment/gradio_app.mdx b/embedchain/docs/deployment/gradio_app.mdx
deleted file mode 100644
index 6c79aa208..000000000
--- a/embedchain/docs/deployment/gradio_app.mdx
+++ /dev/null
@@ -1,59 +0,0 @@
----
-title: 'Gradio.app'
-description: 'Deploy your RAG application to gradio.app platform'
----
-
-Embedchain offers a Streamlit template to facilitate the development of RAG chatbot applications in just three easy steps.
-
-Follow the instructions given below to deploy your first application quickly:
-
-## Step-1: Create RAG app
-
-We provide a command line utility called `ec` in embedchain that inherits the template for `gradio.app` platform and help you deploy the app. Follow the instructions to create a gradio.app app using the template provided:
-
-```bash Install embedchain
-pip install embedchain
-```
-
-```bash Create application
-mkdir my-rag-app
-ec create --template=gradio.app
-```
-
-This will generate a directory structure like this:
-
-```bash
-├── app.py
-├── embedchain.json
-└── requirements.txt
-```
-
-Feel free to edit the files as required.
-- `app.py`: Contains API app code
-- `embedchain.json`: Contains embedchain specific configuration for deployment (you don't need to configure this)
-- `requirements.txt`: Contains python dependencies for your application
-
-## Step-2: Test app locally
-
-You can run the app locally by simply doing:
-
-```bash Run locally
-pip install -r requirements.txt
-ec dev
-```
-
-## Step-3: Deploy to gradio.app
-
-```bash Deploy to gradio.app
-ec deploy
-```
-
-This will run `gradio deploy` which will prompt you questions and deploy your app directly to huggingface spaces.
-
-
-
-## Seeking help?
-
-If you run into issues with deployment, please feel free to reach out to us via any of the following methods:
-
-
diff --git a/embedchain/docs/deployment/huggingface_spaces.mdx b/embedchain/docs/deployment/huggingface_spaces.mdx
deleted file mode 100644
index 5b8811e41..000000000
--- a/embedchain/docs/deployment/huggingface_spaces.mdx
+++ /dev/null
@@ -1,103 +0,0 @@
----
-title: 'Huggingface.co'
-description: 'Deploy your RAG application to huggingface.co platform'
----
-
-With Embedchain, you can directly host your apps in just three steps to huggingface spaces where you can view and deploy your app to the world.
-
-We support two types of deployment to huggingface spaces:
-
-
-
- Streamlit.io
-
-
- Gradio.app
-
-
-
-## Using streamlit.io
-
-### Step 1: Create a new RAG app
-
-Create a new RAG app using the following command:
-
-```bash
-mkdir my-rag-app
-ec create --template=hf/streamlit.io # inside my-rag-app directory
-```
-
-When you run this for the first time, you'll be asked to login to huggingface.co. Once you login, you'll need to create a **write** token. You can create a write token by going to [huggingface.co settings](https://huggingface.co/settings/token). Once you create a token, you'll be asked to enter the token in the terminal.
-
-This will also create an `embedchain.json` file in your app directory. Add a `name` key into the `embedchain.json` file. This will be the "repo-name" of your app in huggingface spaces.
-
-```json embedchain.json
-{
- "name": "my-rag-app",
- "provider": "hf/streamlit.io"
-}
-```
-
-### Step-2: Test app locally
-
-You can run the app locally by simply doing:
-
-```bash Run locally
-pip install -r requirements.txt
-ec dev
-```
-
-### Step-3: Deploy to huggingface spaces
-
-```bash Deploy to huggingface spaces
-ec deploy
-```
-
-This will deploy your app to huggingface spaces. You can view your app at `https://huggingface.co/spaces//my-rag-app`. This will get prompted in the terminal once the app is deployed.
-
-## Using gradio.app
-
-Similar to streamlit.io, you can deploy your app to gradio.app in just three steps.
-
-### Step 1: Create a new RAG app
-
-Create a new RAG app using the following command:
-
-```bash
-mkdir my-rag-app
-ec create --template=hf/gradio.app # inside my-rag-app directory
-```
-
-When you run this for the first time, you'll be asked to login to huggingface.co. Once you login, you'll need to create a **write** token. You can create a write token by going to [huggingface.co settings](https://huggingface.co/settings/token). Once you create a token, you'll be asked to enter the token in the terminal.
-
-This will also create an `embedchain.json` file in your app directory. Add a `name` key into the `embedchain.json` file. This will be the "repo-name" of your app in huggingface spaces.
-
-```json embedchain.json
-{
- "name": "my-rag-app",
- "provider": "hf/gradio.app"
-}
-```
-
-### Step-2: Test app locally
-
-You can run the app locally by simply doing:
-
-```bash Run locally
-pip install -r requirements.txt
-ec dev
-```
-
-### Step-3: Deploy to huggingface spaces
-
-```bash Deploy to huggingface spaces
-ec deploy
-```
-
-This will deploy your app to huggingface spaces. You can view your app at `https://huggingface.co/spaces//my-rag-app`. This will get prompted in the terminal once the app is deployed.
-
-## Seeking help?
-
-If you run into issues with deployment, please feel free to reach out to us via any of the following methods:
-
-
diff --git a/embedchain/docs/deployment/modal_com.mdx b/embedchain/docs/deployment/modal_com.mdx
deleted file mode 100644
index e82d367b6..000000000
--- a/embedchain/docs/deployment/modal_com.mdx
+++ /dev/null
@@ -1,63 +0,0 @@
----
-title: 'Modal.com'
-description: 'Deploy your RAG application to modal.com platform'
----
-
-Embedchain has a nice and simple abstraction on top of the [Modal.com](https://modal.com/) tools to let developers deploy RAG application to modal.com platform seamlessly.
-
-Follow the instructions given below to deploy your first application quickly:
-
-
-## Step-1 Create RAG application:
-
-We provide a command line utility called `ec` in embedchain that inherits the template for `modal.com` platform and help you deploy the app. Follow the instructions to create a modal.com app using the template provided:
-
-
-```bash Create application
-pip install embedchain[modal]
-mkdir my-rag-app
-ec create --template=modal.com
-```
-
-This `create` command will open a browser window and ask you to login to your modal.com account and will generate a directory structure like this:
-
-```bash
-├── app.py
-├── .env
-├── .env.example
-├── embedchain.json
-└── requirements.txt
-```
-
-Feel free to edit the files as required.
-- `app.py`: Contains API app code
-- `.env`: Contains environment variables for production
-- `.env.example`: Contains dummy environment variables (can ignore this file)
-- `embedchain.json`: Contains embedchain specific configuration for deployment (you don't need to configure this)
-- `requirements.txt`: Contains python dependencies for your FastAPI application
-
-## Step-2: Test app locally
-
-You can run the app locally by simply doing:
-
-```bash Run locally
-pip install -r requirements.txt
-ec dev
-```
-
-## Step-3: Deploy to modal.com
-
-You can deploy to modal.com using the following command:
-```bash Deploy app
-ec deploy
-```
-
-Once this step finished, it will provide you with the deployment endpoint where you can access the app live. It will look something like this (Swagger docs):
-
-
-
-## Seeking help?
-
-If you run into issues with deployment, please feel free to reach out to us via any of the following methods:
-
-
diff --git a/embedchain/docs/deployment/railway.mdx b/embedchain/docs/deployment/railway.mdx
deleted file mode 100644
index ef8a60ab8..000000000
--- a/embedchain/docs/deployment/railway.mdx
+++ /dev/null
@@ -1,86 +0,0 @@
----
-title: 'Railway.app'
-description: 'Deploy your RAG application to railway.app'
----
-
-It's easy to host your Embedchain-powered apps and APIs on railway.
-
-Follow the instructions given below to deploy your first application quickly:
-
-## Step-1: Create RAG app
-
-```bash Install embedchain
-pip install embedchain
-```
-
-
-**Create a full stack app using Embedchain CLI**
-
-To use your hosted embedchain RAG app, you can easily set up a FastAPI server that can be used anywhere.
-To easily set up a FastAPI server, check out [Get started with Full stack](https://docs.embedchain.ai/get-started/full-stack) page.
-
-Hosting this server on railway is super easy!
-
-
-
-## Step-2: Set up your project
-
-### With Docker
-
-You can create a `Dockerfile` in the root of the project, with all the instructions. However, this method is sometimes slower in deployment.
-
-### Without Docker
-
-By default, Railway uses Python 3.7. Embedchain requires the python version to be >3.9 in order to install.
-
-To fix this, create a `.python-version` file in the root directory of your project and specify the correct version
-
-```bash .python-version
-3.10
-```
-
-You also need to create a `requirements.txt` file to specify the requirements.
-
-```bash requirements.txt
-python-dotenv
-embedchain
-fastapi==0.108.0
-uvicorn==0.25.0
-embedchain
-beautifulsoup4
-sentence-transformers
-```
-
-## Step-3: Deploy to Railway 🚀
-
-1. Go to https://railway.app and create an account.
-2. Create a project by clicking on the "Start a new project" button
-
-### With Github
-
-Select `Empty Project` or `Deploy from Github Repo`.
-
-You should be all set!
-
-### Without Github
-
-You can also use the railway CLI to deploy your apps from the terminal, if you don't want to connect a git repository.
-
-To do this, just run this command in your terminal
-
-```bash Install and set up railway CLI
-npm i -g @railway/cli
-railway login
-railway link [projectID]
-```
-
-Finally, run `railway up` to deploy your app.
-```bash Deploy
-railway up
-```
-
-## Seeking help?
-
-If you run into issues with deployment, please feel free to reach out to us via any of the following methods:
-
-
diff --git a/embedchain/docs/deployment/render_com.mdx b/embedchain/docs/deployment/render_com.mdx
deleted file mode 100644
index 81ba7f6df..000000000
--- a/embedchain/docs/deployment/render_com.mdx
+++ /dev/null
@@ -1,93 +0,0 @@
----
-title: 'Render.com'
-description: 'Deploy your RAG application to render.com platform'
----
-
-Embedchain has a nice and simple abstraction on top of the [render.com](https://render.com/) tools to let developers deploy RAG application to render.com platform seamlessly.
-
-Follow the instructions given below to deploy your first application quickly:
-
-## Step-1: Install `render` command line
-
-
-```bash OSX
-brew tap render-oss/render
-brew install render
-```
-
-```bash Linux
-# Make sure you have deno installed -> https://docs.render.com/docs/cli#from-source-unsupported-operating-systems
-git clone https://github.com/render-oss/render-cli
-cd render-cli
-make deps
-deno task run
-deno compile
-```
-
-```bash Windows
-choco install rendercli
-```
-
-
-In case you run into issues, refer to official [render.com docs](https://docs.render.com/docs/cli).
-
-## Step-2 Create RAG application:
-
-We provide a command line utility called `ec` in embedchain that inherits the template for `render.com` platform and help you deploy the app. Follow the instructions to create a render.com app using the template provided:
-
-
-```bash Create application
-pip install embedchain
-mkdir my-rag-app
-ec create --template=render.com
-```
-
-This `create` command will open a browser window and ask you to login to your render.com account and will generate a directory structure like this:
-
-```bash
-├── app.py
-├── .env
-├── render.yaml
-├── embedchain.json
-└── requirements.txt
-```
-
-Feel free to edit the files as required.
-- `app.py`: Contains API app code
-- `.env`: Contains environment variables for production
-- `render.yaml`: Contains render.com specific configuration for deployment (configure this according to your needs, follow [this](https://docs.render.com/docs/blueprint-spec) for more info)
-- `embedchain.json`: Contains embedchain specific configuration for deployment (you don't need to configure this)
-- `requirements.txt`: Contains python dependencies for your application
-
-## Step-3: Test app locally
-
-You can run the app locally by simply doing:
-
-```bash Run locally
-pip install -r requirements.txt
-ec dev
-```
-
-## Step-4: Deploy to render.com
-
-Before deploying to render.com, you only have to set up one thing.
-
-In the render.yaml file, make sure to modify the repo key by inserting the URL of your Git repository where your application will be hosted. You can create a repository from [GitHub](https://github.com) or [GitLab](https://gitlab.com/users/sign_in).
-
-After that, you're ready to deploy on render.com.
-
-```bash Deploy app
-ec deploy
-```
-
-When you run this, it should open up your render dashboard and you can see the app being deployed. You can find your hosted link over there only.
-
-You can also check the logs, monitor app status etc on their dashboard by running command `render dashboard`.
-
-
-
-## Seeking help?
-
-If you run into issues with deployment, please feel free to reach out to us via any of the following methods:
-
-
diff --git a/embedchain/docs/deployment/streamlit_io.mdx b/embedchain/docs/deployment/streamlit_io.mdx
deleted file mode 100644
index 93dde7400..000000000
--- a/embedchain/docs/deployment/streamlit_io.mdx
+++ /dev/null
@@ -1,62 +0,0 @@
----
-title: 'Streamlit.io'
-description: 'Deploy your RAG application to streamlit.io platform'
----
-
-Embedchain offers a Streamlit template to facilitate the development of RAG chatbot applications in just three easy steps.
-
-Follow the instructions given below to deploy your first application quickly:
-
-## Step-1: Create RAG app
-
-We provide a command line utility called `ec` in embedchain that inherits the template for `streamlit.io` platform and help you deploy the app. Follow the instructions to create a streamlit.io app using the template provided:
-
-```bash Install embedchain
-pip install embedchain
-```
-
-```bash Create application
-mkdir my-rag-app
-ec create --template=streamlit.io
-```
-
-This will generate a directory structure like this:
-
-```bash
-├── .streamlit
-│ └── secrets.toml
-├── app.py
-├── embedchain.json
-└── requirements.txt
-```
-
-Feel free to edit the files as required.
-- `app.py`: Contains API app code
-- `.streamlit/secrets.toml`: Contains secrets for your application
-- `embedchain.json`: Contains embedchain specific configuration for deployment (you don't need to configure this)
-- `requirements.txt`: Contains python dependencies for your application
-
-Add your `OPENAI_API_KEY` in `.streamlit/secrets.toml` file to run and deploy the app.
-
-## Step-2: Test app locally
-
-You can run the app locally by simply doing:
-
-```bash Run locally
-pip install -r requirements.txt
-ec dev
-```
-
-## Step-3: Deploy to streamlit.io
-
-
-
-Use the deploy button from the streamlit website to deploy your app.
-
-You can refer this [guide](https://docs.streamlit.io/streamlit-community-cloud/deploy-your-app) if you run into any problems.
-
-## Seeking help?
-
-If you run into issues with deployment, please feel free to reach out to us via any of the following methods:
-
-
diff --git a/embedchain/docs/development.mdx b/embedchain/docs/development.mdx
deleted file mode 100644
index 878300893..000000000
--- a/embedchain/docs/development.mdx
+++ /dev/null
@@ -1,98 +0,0 @@
----
-title: 'Development'
-description: 'Learn how to preview changes locally'
----
-
-
- **Prerequisite** You should have installed Node.js (version 18.10.0 or
- higher).
-
-
-Step 1. Install Mintlify on your OS:
-
-
-
-```bash npm
-npm i -g mintlify
-```
-
-```bash yarn
-yarn global add mintlify
-```
-
-
-
-Step 2. Go to the docs are located (where you can find `mint.json`) and run the following command:
-
-```bash
-mintlify dev
-```
-
-The documentation website is now available at `http://localhost:3000`.
-
-### Custom Ports
-
-Mintlify uses port 3000 by default. You can use the `--port` flag to customize the port Mintlify runs on. For example, use this command to run in port 3333:
-
-```bash
-mintlify dev --port 3333
-```
-
-You will see an error like this if you try to run Mintlify in a port that's already taken:
-
-```md
-Error: listen EADDRINUSE: address already in use :::3000
-```
-
-## Mintlify Versions
-
-Each CLI is linked to a specific version of Mintlify. Please update the CLI if your local website looks different than production.
-
-
-
-```bash npm
-npm i -g mintlify@latest
-```
-
-```bash yarn
-yarn global upgrade mintlify
-```
-
-
-
-## Deployment
-
-
- Unlimited editors available under the [Startup
- Plan](https://mintlify.com/pricing)
-
-
-You should see the following if the deploy successfully went through:
-
-
-
-
-
-## Troubleshooting
-
-Here's how to solve some common problems when working with the CLI.
-
-
-
- Update to Node v18. Run `mintlify install` and try again.
-
-
-Go to the `C:/Users/Username/.mintlify/` directory and remove the `mint`
-folder. Then Open the Git Bash in this location and run `git clone
-https://github.com/mintlify/mint.git`.
-
-Repeat step 3.
-
-
-
- Try navigating to the root of your device and delete the ~/.mintlify folder.
- Then run `mintlify dev` again.
-
-
-
-Curious about what changed in a CLI version? [Check out the CLI changelog.](/changelog/command-line)
diff --git a/embedchain/docs/examples/chat-with-PDF.mdx b/embedchain/docs/examples/chat-with-PDF.mdx
deleted file mode 100644
index ad8fb9a5b..000000000
--- a/embedchain/docs/examples/chat-with-PDF.mdx
+++ /dev/null
@@ -1,32 +0,0 @@
-### Embedchain Chat with PDF App
-
-You can easily create and deploy your own `chat-pdf` App using Embedchain.
-
-Here are few simple steps for you to create and deploy your app:
-
-1. Fork the embedchain repo from [Github](https://github.com/embedchain/embedchain).
-
-
-If you run into problems with forking, please refer to [github docs](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/fork-a-repo) for forking a repo.
-
-
-2. Navigate to `chat-pdf` example app from your forked repo:
-
-```bash
-cd /examples/chat-pdf
-```
-
-3. Run your app in development environment with simple commands
-
-```bash
-pip install -r requirements.txt
-ec dev
-```
-
-Feel free to improve our simple `chat-pdf` streamlit app and create pull request to showcase your app [here](https://docs.embedchain.ai/examples/showcase)
-
-4. You can easily deploy your app using Streamlit interface
-
-Connect your Github account with Streamlit and refer this [guide](https://docs.streamlit.io/streamlit-community-cloud/deploy-your-app) to deploy your app.
-
-You can also use the deploy button from your streamlit website you see when running `ec dev` command.
diff --git a/embedchain/docs/examples/community/showcase.mdx b/embedchain/docs/examples/community/showcase.mdx
deleted file mode 100644
index d8b511919..000000000
--- a/embedchain/docs/examples/community/showcase.mdx
+++ /dev/null
@@ -1,115 +0,0 @@
----
-title: '🎪 Community showcase'
----
-
-Embedchain community has been super active in creating demos on top of Embedchain. On this page, we showcase all the apps, blogs, videos, and tutorials created by the community. ❤️
-
-## Apps
-
-### Open Source
-
-- [My GSoC23 bot- Streamlit chat](https://github.com/lucifertrj/EmbedChain_GSoC23_BOT) by Tarun Jain
-- [Discord Bot for LLM chat](https://github.com/Reidond/discord_bots_playground/tree/c8b0c36541e4b393782ee506804c4b6962426dd6/python/chat-channel-bot) by Reidond
-- [EmbedChain-Streamlit-Docker App](https://github.com/amjadraza/embedchain-streamlit-app) by amjadraza
-- [Harry Potter Philosphers Stone Bot](https://github.com/vinayak-kempawad/Harry_Potter_Philosphers_Stone_Bot/) by Vinayak Kempawad, ([LinkedIn post](https://www.linkedin.com/feed/update/urn:li:activity:7080907532155686912/))
-- [LLM bot trained on own messages](https://github.com/Harin329/harinBot) by Hao Wu
-
-### Closed Source
-
-- [Taobot.io](https://taobot.io) - chatbot & knowledgebase hybrid by [cachho](https://github.com/cachho)
-- [Create Instant ChatBot 🤖 using embedchain](https://databutton.com/v/h3e680h9) by Avra, ([Tweet](https://twitter.com/Avra_b/status/1674704745154641920/))
-- [JOBO 🤖 — The AI-driven sidekick to craft your resume](https://try-jobo.com/) by Enrico Willemse, ([LinkedIn Post](https://www.linkedin.com/posts/enrico-willemse_jobai-gptfun-embedchain-activity-7090340080879374336-ueLB/))
-- [Explore Your Knowledge Base: Interactive chats over various forms of documents](https://chatdocs.dkedar.com/) by Kedar Dabhadkar, ([LinkedIn Post](https://www.linkedin.com/posts/dkedar7_machinelearning-llmops-activity-7092524836639424513-2O3L/))
-- [Chatbot trained on 1000+ videos of Ester hicks the co-author behind the famous book Secret](https://ask-abraham.thoughtseed.repl.co) by Mohan Kumar
-
-
-## Templates
-
-### Replit
-- [Embedchain Chat Bot](https://replit.com/@taranjeet1/Embedchain-Chat-Bot) by taranjeetio
-- [Embedchain Memory Chat Bot Template](https://replit.com/@taranjeetio/Embedchain-Memory-Chat-Bot-Template) by taranjeetio
-- [Chatbot app to demonstrate question-answering using retrieved information](https://replit.com/@AllisonMorrell/EmbedChainlitPublic) by Allison Morrell, ([LinkedIn Post](https://www.linkedin.com/posts/allison-morrell-2889275a_retrievalbot-screenshots-activity-7080339991754649600-wihZ/))
-
-## Posts
-
-### Blogs
-
-- [Customer Service LINE Bot](https://www.evanlin.com/langchain-embedchain/) by Evan Lin
-- [Chatbot in Under 5 mins using Embedchain](https://medium.com/@ayush.wattal/chatbot-in-under-5-mins-using-embedchain-a4f161fcf9c5) by Ayush Wattal
-- [Understanding what the LLM framework embedchain does](https://zenn.dev/hijikix/articles/4bc8d60156a436) by Daisuke Hashimoto
-- [In bed with GPT and Node.js](https://dev.to/worldlinetech/in-bed-with-gpt-and-nodejs-4kh2) by Raphaël Semeteys, ([LinkedIn Post](https://www.linkedin.com/posts/raphaelsemeteys_in-bed-with-gpt-and-nodejs-activity-7088113552326029313-nn87/))
-- [Using Embedchain — A powerful LangChain Python wrapper to build Chat Bots even faster!⚡](https://medium.com/@avra42/using-embedchain-a-powerful-langchain-python-wrapper-to-build-chat-bots-even-faster-35c12994a360) by Avra, ([Tweet](https://twitter.com/Avra_b/status/1686767751560310784/))
-- [What is the Embedchain library?](https://jahaniwww.com/%da%a9%d8%aa%d8%a7%d8%a8%d8%ae%d8%a7%d9%86%d9%87-embedchain/) by Ali Jahani, ([LinkedIn Post](https://www.linkedin.com/posts/ajahani_aepaetaeqaexaggahyaeu-aetaexaesabraeaaeqaepaeu-activity-7097605202135904256-ppU-/))
-- [LangChain is Nice, But Have You Tried EmbedChain ?](https://medium.com/thoughts-on-machine-learning/langchain-is-nice-but-have-you-tried-embedchain-215a34421cde) by FS Ndzomga, ([Tweet](https://twitter.com/ndzfs/status/1695583640372035951/))
-- [Simplest Method to Build a Custom Chatbot with GPT-3.5 (via Embedchain)](https://www.ainewsletter.today/p/simplest-method-to-build-a-custom) by Arjun, ([Tweet](https://twitter.com/aiguy_arjun/status/1696393808467091758/))
-
-### LinkedIn
-
-- [What is embedchain](https://www.linkedin.com/posts/activity-7079393104423698432-wRyi/) by Rithesh Sreenivasan
-- [Building a chatbot with EmbedChain](https://www.linkedin.com/posts/activity-7078434598984060928-Zdso/) by Lior Sinclair
-- [Making chatbot without vs with embedchain](https://www.linkedin.com/posts/kalyanksnlp_llms-chatbots-langchain-activity-7077453416221863936-7N1L/) by Kalyan KS
-- [EmbedChain - very intuitive, first you index your data and then query!](https://www.linkedin.com/posts/shubhamsaboo_embedchain-a-framework-to-easily-create-activity-7079535460699557888-ad1X/) by Shubham Saboo
-- [EmbedChain - Harnessing power of LLM](https://www.linkedin.com/posts/uditsaini_chatbotrevolution-llmpoweredbots-embedchainframework-activity-7077520356827181056-FjTK/) by Udit S.
-- [AI assistant for ABBYY Vantage](https://www.linkedin.com/posts/maximevermeir_llm-github-abbyy-activity-7081658972071424000-fXfZ/) by Maxime V.
-- [About embedchain](https://www.linkedin.com/feed/update/urn:li:activity:7080984218914189312/) by Morris Lee
-- [How to use Embedchain](https://www.linkedin.com/posts/nehaabansal_github-embedchainembedchain-framework-activity-7085830340136595456-kbW5/) by Neha Bansal
-- [Youtube/Webpage summary for Energy Study](https://www.linkedin.com/posts/bar%C4%B1%C5%9F-sanl%C4%B1-34b82715_enerji-python-activity-7082735341563977730-Js0U/) by Barış Sanlı, ([Tweet](https://twitter.com/barissanli/status/1676968784979193857/))
-- [Demo: How to use Embedchain? (Contains Collab Notebook link)](https://www.linkedin.com/posts/liorsinclair_embedchain-is-getting-a-lot-of-traction-because-activity-7103044695995424768-RckT/) by Lior Sinclair
-
-### Twitter
-
-- [What is embedchain](https://twitter.com/AlphaSignalAI/status/1672668574450847745) by Lior
-- [Building a chatbot with Embedchain](https://twitter.com/Saboo_Shubham_/status/1673537044419686401) by Shubham Saboo
-- [Chatbot docker image behind an API with yaml configs with Embedchain](https://twitter.com/tricalt/status/1678411430192730113/) by Vasilije
-- [Build AI powered PDF chatbot with just five lines of Python code with Embedchain!](https://twitter.com/Saboo_Shubham_/status/1676627104866156544/) by Shubham Saboo
-- [Chatbot against a youtube video using embedchain](https://twitter.com/smaameri/status/1675201443043704834/) by Sami Maameri
-- [Highlights of EmbedChain](https://twitter.com/carl_AIwarts/status/1673542204328120321/) by carl_AIwarts
-- [Build Llama-2 chatbot in less than 5 minutes](https://twitter.com/Saboo_Shubham_/status/1682168956918833152/) by Shubham Saboo
-- [All cool features of embedchain](https://twitter.com/DhravyaShah/status/1683497882438217728/) by Dhravya Shah, ([LinkedIn Post](https://www.linkedin.com/posts/dhravyashah_what-if-i-tell-you-that-you-can-make-an-ai-activity-7089459599287726080-ZIYm/))
-- [Read paid Medium articles for Free using embedchain](https://twitter.com/kumarkaushal_/status/1688952961622585344) by Kaushal Kumar
-
-## Videos
-
-- [Embedchain in one shot](https://www.youtube.com/watch?v=vIhDh7H73Ww&t=82s) by AI with Tarun
-- [embedChain Create LLM powered bots over any dataset Python Demo Tesla Neurallink Chatbot Example](https://www.youtube.com/watch?v=bJqAn22a6Gc) by Rithesh Sreenivasan
-- [Embedchain - NEW 🔥 Langchain BABY to build LLM Bots](https://www.youtube.com/watch?v=qj_GNQ06I8o) by 1littlecoder
-- [EmbedChain -- NEW!: Build LLM-Powered Bots with Any Dataset](https://www.youtube.com/watch?v=XmaBezzGHu4) by DataInsightEdge
-- [Chat With Your PDFs in less than 10 lines of code! EMBEDCHAIN tutorial](https://www.youtube.com/watch?v=1ugkcsAcw44) by Phani Reddy
-- [How To Create A Custom Knowledge AI Powered Bot | Install + How To Use](https://www.youtube.com/watch?v=VfCrIiAst-c) by The Ai Solopreneur
-- [Build Custom Chatbot in 6 min with this Framework [Beginner Friendly]](https://www.youtube.com/watch?v=-8HxOpaFySM) by Maya Akim
-- [embedchain-streamlit-app](https://www.youtube.com/watch?v=3-9GVd-3v74) by Amjad Raza
-- [🤖CHAT with ANY ONLINE RESOURCES using EMBEDCHAIN - a LangChain wrapper, in few lines of code !](https://www.youtube.com/watch?v=Mp7zJe4TIdM) by Avra
-- [Building resource-driven LLM-powered bots with Embedchain](https://www.youtube.com/watch?v=IVfcAgxTO4I) by BugBytes
-- [embedchain-streamlit-demo](https://www.youtube.com/watch?v=yJAWB13FhYQ) by Amjad Raza
-- [Embedchain - create your own AI chatbots using open source models](https://www.youtube.com/shorts/O3rJWKwSrWE) by Dhravya Shah
-- [AI ChatBot in 5 lines Python Code](https://www.youtube.com/watch?v=zjWvLJLksv8) by Data Engineering
-- [Interview with Karl Marx](https://www.youtube.com/watch?v=5Y4Tscwj1xk) by Alexander Ray Williams
-- [Vlog where we try to build a bot based on our content on the internet](https://www.youtube.com/watch?v=I2w8CWM3bx4) by DV, ([Tweet](https://twitter.com/dvcoolster/status/1688387017544261632))
-- [CHAT with ANY ONLINE RESOURCES using EMBEDCHAIN|STREAMLIT with MEMORY |All OPENSOURCE](https://www.youtube.com/watch?v=TqQIHWoWTDQ&pp=ygUKZW1iZWRjaGFpbg%3D%3D) by DataInsightEdge
-- [Build POWERFUL LLM Bots EASILY with Your Own Data - Embedchain - Langchain 2.0? (Tutorial)](https://www.youtube.com/watch?v=jE24Y_GasE8) by WorldofAI, ([Tweet](https://twitter.com/intheworldofai/status/1696229166922780737))
-- [Embedchain: An AI knowledge base assistant for customizing enterprise private data, which can be connected to discord, whatsapp, slack, tele and other terminals (with gradio to build a request interface) in Chinese](https://www.youtube.com/watch?v=5RZzCJRk-d0) by AIGC LINK
-- [Embedchain Introduction](https://www.youtube.com/watch?v=Jet9zAqyggI) by Fahd Mirza
-
-## Mentions
-
-### Github repos
-
-- [Awesome-LLM](https://github.com/Hannibal046/Awesome-LLM)
-- [awesome-chatgpt-api](https://github.com/reorx/awesome-chatgpt-api)
-- [awesome-langchain](https://github.com/kyrolabs/awesome-langchain)
-- [Awesome-Prompt-Engineering](https://github.com/promptslab/Awesome-Prompt-Engineering)
-- [awesome-chatgpt](https://github.com/eon01/awesome-chatgpt)
-- [Awesome-LLMOps](https://github.com/tensorchord/Awesome-LLMOps)
-- [awesome-generative-ai](https://github.com/filipecalegario/awesome-generative-ai)
-- [awesome-gpt](https://github.com/formulahendry/awesome-gpt)
-- [awesome-ChatGPT-repositories](https://github.com/taishi-i/awesome-ChatGPT-repositories)
-- [awesome-gpt-prompt-engineering](https://github.com/snwfdhmp/awesome-gpt-prompt-engineering)
-- [awesome-chatgpt](https://github.com/awesome-chatgpt/awesome-chatgpt)
-- [awesome-llm-and-aigc](https://github.com/sjinzh/awesome-llm-and-aigc)
-- [awesome-compbio-chatgpt](https://github.com/csbl-br/awesome-compbio-chatgpt)
-- [Awesome-LLM4Tool](https://github.com/OpenGVLab/Awesome-LLM4Tool)
-
-## Meetups
-
-- [Dash and ChatGPT: Future of AI-enabled apps 30/08/23](https://go.plotly.com/dash-chatgpt)
-- [Pie & AI: Bangalore - Build end-to-end LLM app using Embedchain 01/09/23](https://www.eventbrite.com/e/pie-ai-bangalore-build-end-to-end-llm-app-using-embedchain-tickets-698045722547)
diff --git a/embedchain/docs/examples/discord_bot.mdx b/embedchain/docs/examples/discord_bot.mdx
deleted file mode 100644
index 247f3c634..000000000
--- a/embedchain/docs/examples/discord_bot.mdx
+++ /dev/null
@@ -1,70 +0,0 @@
----
-title: "🤖 Discord Bot"
----
-
-### 🔑 Keys Setup
-
-- Set your `OPENAI_API_KEY` in your variables.env file.
-- Go to [https://discord.com/developers/applications/](https://discord.com/developers/applications/) and click on `New Application`.
-- Enter the name for your bot, accept the terms and click on `Create`. On the resulting page, enter the details of your bot as you like.
-- On the left sidebar, click on `Bot`. Under the heading `Privileged Gateway Intents`, toggle all 3 options to ON position. Save your changes.
-- Now click on `Reset Token` and copy the token value. Set it as `DISCORD_BOT_TOKEN` in .env file.
-- On the left sidebar, click on `OAuth2` and go to `General`.
-- Set `Authorization Method` to `In-app Authorization`. Under `Scopes` select `bot`.
-- Under `Bot Permissions` allow the following and then click on `Save Changes`.
-
-```text
-Send Messages (under Text Permissions)
-```
-
-- Now under `OAuth2` and go to `URL Generator`. Under `Scopes` select `bot`.
-- Under `Bot Permissions` set the same permissions as above.
-- Now scroll down and copy the `Generated URL`. Paste it in a browser window and select the Server where you want to add the bot.
-- Click on `Continue` and authorize the bot.
-- 🎉 The bot has been successfully added to your server. But it's still offline.
-
-### Take the bot online
-
-
-
- ```bash
- docker run --name discord-bot -e OPENAI_API_KEY=sk-xxx -e DISCORD_BOT_TOKEN=xxx -p 8080:8080 embedchain/discord-bot:latest
- ```
-
-
- ```bash
- pip install --upgrade "embedchain[discord]"
-
- python -m embedchain.bots.discord
-
- # or if you prefer to see the question and not only the answer, run it with
- python -m embedchain.bots.discord --include-question
- ```
-
-
-
-### 🚀 Usage Instructions
-
-- Go to the server where you have added your bot.
- 
-- You can add data sources to the bot using the slash command:
-
-```text
-/ec add
-```
-
-- You can ask your queries from the bot using the slash command:
-
-```text
-/ec query
-```
-
-- You can chat with the bot using the slash command:
-
-```text
-/ec chat
-```
-
-📝 Note: To use the bot privately, you can message the bot directly by right clicking the bot and selecting `Message`.
-
-🎉 Happy Chatting! 🎉
diff --git a/embedchain/docs/examples/nextjs-assistant.mdx b/embedchain/docs/examples/nextjs-assistant.mdx
deleted file mode 100644
index 86f82fb4f..000000000
--- a/embedchain/docs/examples/nextjs-assistant.mdx
+++ /dev/null
@@ -1,124 +0,0 @@
-Fork the Embedchain repo on [Github](https://github.com/embedchain/embedchain) to create your own NextJS discord and slack bot powered by Embedchain.
-
-If you run into problems with forking, please refer to [github docs](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/fork-a-repo) for forking a repo.
-
-We will work from the `examples/nextjs` folder so change your current working directory by running the command - `cd /examples/nextjs`
-
-# Installation
-
-First, lets start by install all the required packages and dependencies.
-
-- Install all the required python packages by running ```pip install -r requirements.txt```
-
-- We will use [Fly.io](https://fly.io/) to deploy our embedchain app, discord and slack bot. Follow the step one to install [Fly.io CLI](https://docs.embedchain.ai/deployment/fly_io#step-1-install-flyctl-command-line)
-
-# Developement
-
-## Embedchain App
-
-First, we need an Embedchain app powered with the knowledge of NextJS. We have already created an embedchain app using FastAPI in `ec_app` folder for you. Feel free to ingest data of your choice to power the App.
-
-
-Navigate to `ec_app` folder and create `.env` file in this folder and set your OpenAI API key as shown in `.env.example` file. If you want to use other open-source models, feel free to use the app config in `app.py`. More details for using custom configuration for Embedchain app is [available here](https://docs.embedchain.ai/api-reference/advanced/configuration).
-
-
-Before running the ec commands to develope the app, open `fly.toml` file and update the `name` variable to something unique. This is important as `fly.io` requires users to provide a globally unique deployment app names.
-
-Now, we need to launch this application with fly.io. You can see your app on [fly.io dashboard](https://fly.io/dashboard). Run the following command to launch your app on fly.io:
-```bash
-fly launch --no-deploy
-```
-
-To run the app in development, run the following command:
-
-```bash
-ec dev
-```
-
-Run `ec deploy` to deploy your app on Fly.io. Once you deploy your app, save the endpoint on which our discord and slack bot will send requests.
-
-
-## Discord bot
-
-For discord bot, you will need to create the bot on discord developer portal and get the discord bot token and your discord bot name.
-
-While keeping in mind the following note, create the discord bot by following the instructions from our [discord bot docs](https://docs.embedchain.ai/examples/discord_bot) and get discord bot token.
-
-
-You do not need to set `OPENAI_API_KEY` to run this discord bot. Follow the remaining instructions to create a discord bot app. We recommend you to give the following sets of bot permissions to run the discord bot without errors:
-
-```
-(General Permissions)
-Read Message/View Channels
-
-(Text Permissions)
-Send Messages
-Create Public Thread
-Create Private Thread
-Send Messages in Thread
-Manage Threads
-Embed Links
-Read Message History
-```
-
-
-Once you have your discord bot token and discord app name. Navigate to `nextjs_discord` folder and create `.env` file and define your discord bot token, discord bot name and endpoint of your embedchain app as shown in `.env.example` file.
-
-To run the app in development:
-
-```bash
-python app.py
-```
-
-Before deploying the app, open `fly.toml` file and update the `name` variable to something unique. This is important as `fly.io` requires users to provide a globally unique deployment app names.
-
-Now, we need to launch this application with fly.io. You can see your app on [fly.io dashboard](https://fly.io/dashboard). Run the following command to launch your app on fly.io:
-```bash
-fly launch --no-deploy
-```
-
-Run `ec deploy` to deploy your app on Fly.io. Once you deploy your app, your discord bot will be live!
-
-
-## Slack bot
-
-For Slack bot, you will need to create the bot on slack developer portal and get the slack bot token and slack app token.
-
-### Setup
-
-- Create a workspace on Slack if you don't have one already by clicking [here](https://slack.com/intl/en-in/).
-- Create a new App on your Slack account by going [here](https://api.slack.com/apps).
-- Select `From Scratch`, then enter the Bot Name and select your workspace.
-- Go to `App Credentials` section on the `Basic Information` tab from the left sidebar, create your app token and save it in your `.env` file as `SLACK_APP_TOKEN`.
-- Go to `Socket Mode` tab from the left sidebar and enable the socket mode to listen to slack message from your workspace.
-- (Optional) Under the `App Home` tab you can change your App display name and default name.
-- Navigate to `Event Subscription` tab, and enable the event subscription so that we can listen to slack events.
-- Once you enable the event subscription, you will need to subscribe to bot events to authorize the bot to listen to app mention events of the bot. Do that by tapping on `Add Bot User Event` button and select `app_mention`.
-- On the left Sidebar, go to `OAuth and Permissions` and add the following scopes under `Bot Token Scopes`:
-```text
-app_mentions:read
-channels:history
-channels:read
-chat:write
-emoji:read
-reactions:write
-reactions:read
-```
-- Now select the option `Install to Workspace` and after it's done, copy the `Bot User OAuth Token` and set it in your `.env` file as `SLACK_BOT_TOKEN`.
-
-Once you have your slack bot token and slack app token. Navigate to `nextjs_slack` folder and create `.env` file and define your slack bot token, slack app token and endpoint of your embedchain app as shown in `.env.example` file.
-
-To run the app in development:
-
-```bash
-python app.py
-```
-
-Before deploying the app, open `fly.toml` file and update the `name` variable to something unique. This is important as `fly.io` requires users to provide a globally unique deployment app names.
-
-Now, we need to launch this application with fly.io. You can see your app on [fly.io dashboard](https://fly.io/dashboard). Run the following command to launch your app on fly.io:
-```bash
-fly launch --no-deploy
-```
-
-Run `ec deploy` to deploy your app on Fly.io. Once you deploy your app, your slack bot will be live!
diff --git a/embedchain/docs/examples/notebooks-and-replits.mdx b/embedchain/docs/examples/notebooks-and-replits.mdx
deleted file mode 100644
index 2da7208a4..000000000
--- a/embedchain/docs/examples/notebooks-and-replits.mdx
+++ /dev/null
@@ -1,138 +0,0 @@
----
-title: Notebooks & Replits
----
-
-# Explore awesome apps
-
-Check out the remarkable work accomplished using [Embedchain](https://app.embedchain.ai/custom-gpts/).
-
-## Collection of Google colab notebook and Replit links for users
-
-Get started with Embedchain by trying out the examples below. You can run the examples in your browser using Google Colab or Replit.
-
-
-
-
-
LLM
-
Google Colab
-
Replit
-
-
-
-
-
OpenAI
-
-
-
-
-
Anthropic
-
-
-
-
-
Azure OpenAI
-
-
-
-
-
VertexAI
-
-
-
-
-
Cohere
-
-
-
-
-
Together
-
-
-
-
Ollama
-
-
-
-
Hugging Face
-
-
-
-
-
JinaChat
-
-
-
-
-
GPT4All
-
-
-
-
-
Llama2
-
-
-
-
-
-
-
-
-
Embedding model
-
Google Colab
-
Replit
-
-
-
-
-
OpenAI
-
-
-
-
-
VertexAI
-
-
-
-
-
GPT4All
-
-
-
-
-
Hugging Face
-
-
-
-
-
-
-
-
-
Vector DB
-
Google Colab
-
Replit
-
-
-
-
-
ChromaDB
-
-
-
-
-
Elasticsearch
-
-
-
-
-
Opensearch
-
-
-
-
-
Pinecone
-
-
-
-
-
\ No newline at end of file
diff --git a/embedchain/docs/examples/openai-assistant.mdx b/embedchain/docs/examples/openai-assistant.mdx
deleted file mode 100644
index ffd312fa7..000000000
--- a/embedchain/docs/examples/openai-assistant.mdx
+++ /dev/null
@@ -1,60 +0,0 @@
----
-title: 'OpenAI Assistant'
----
-
-
-
-Embedchain now supports [OpenAI Assistants API](https://platform.openai.com/docs/assistants/overview) which allows you to build AI assistants within your own applications. An Assistant has instructions and can leverage models, tools, and knowledge to respond to user queries.
-
-At a high level, an integration of the Assistants API has the following flow:
-
-1. Create an Assistant in the API by defining custom instructions and picking a model
-2. Create a Thread when a user starts a conversation
-3. Add Messages to the Thread as the user ask questions
-4. Run the Assistant on the Thread to trigger responses. This automatically calls the relevant tools.
-
-Creating an OpenAI Assistant using Embedchain is very simple 3 step process.
-
-## Step 1: Create OpenAI Assistant
-
-Make sure that you have `OPENAI_API_KEY` set in the environment variable.
-
-```python Initialize
-from embedchain.store.assistants import OpenAIAssistant
-
-assistant = OpenAIAssistant(
- name="OpenAI DevDay Assistant",
- instructions="You are an organizer of OpenAI DevDay",
-)
-```
-
-If you want to use the existing assistant, you can do something like this:
-
-```python Initialize
-# Load an assistant and create a new thread
-assistant = OpenAIAssistant(assistant_id="asst_xxx")
-
-# Load a specific thread for an assistant
-assistant = OpenAIAssistant(assistant_id="asst_xxx", thread_id="thread_xxx")
-```
-
-## Step-2: Add data to thread
-
-You can add any custom data source that is supported by Embedchain. Else, you can directly pass the file path on your local system and Embedchain propagates it to OpenAI Assistant.
-```python Add data
-assistant.add("/path/to/file.pdf")
-assistant.add("https://www.youtube.com/watch?v=U9mJuUkhUzk")
-assistant.add("https://openai.com/blog/new-models-and-developer-products-announced-at-devday")
-```
-
-## Step-3: Chat with your Assistant
-```python Chat
-assistant.chat("How much OpenAI credits were offered to attendees during OpenAI DevDay?")
-# Response: 'Every attendee of OpenAI DevDay 2023 was offered $500 in OpenAI credits.'
-```
-
-You can try it out yourself using the following Google Colab notebook:
-
-
-
-
diff --git a/embedchain/docs/examples/opensource-assistant.mdx b/embedchain/docs/examples/opensource-assistant.mdx
deleted file mode 100644
index f4dcaa521..000000000
--- a/embedchain/docs/examples/opensource-assistant.mdx
+++ /dev/null
@@ -1,51 +0,0 @@
----
-title: 'Open-Source AI Assistant'
----
-
-Embedchain also provides support for creating Open-Source AI Assistants (similar to [OpenAI Assistants API](https://platform.openai.com/docs/assistants/overview)) which allows you to build AI assistants within your own applications using any LLM (OpenAI or otherwise). An Assistant has instructions and can leverage models, tools, and knowledge to respond to user queries.
-
-At a high level, the Open-Source AI Assistants API has the following flow:
-
-1. Create an AI Assistant by picking a model
-2. Create a Thread when a user starts a conversation
-3. Add Messages to the Thread as the user ask questions
-4. Run the Assistant on the Thread to trigger responses. This automatically calls the relevant tools.
-
-Creating an Open-Source AI Assistant is a simple 3 step process.
-
-## Step 1: Instantiate AI Assistant
-
-```python Initialize
-from embedchain.store.assistants import AIAssistant
-
-assistant = AIAssistant(
- name="My Assistant",
- data_sources=[{"source": "https://www.youtube.com/watch?v=U9mJuUkhUzk"}])
-```
-
-If you want to use the existing assistant, you can do something like this:
-
-```python Initialize
-# Load an assistant and create a new thread
-assistant = AIAssistant(assistant_id="asst_xxx")
-
-# Load a specific thread for an assistant
-assistant = AIAssistant(assistant_id="asst_xxx", thread_id="thread_xxx")
-```
-
-## Step-2: Add data to thread
-
-You can add any custom data source that is supported by Embedchain. Else, you can directly pass the file path on your local system and Embedchain propagates it to OpenAI Assistant.
-
-```python Add data
-assistant.add("/path/to/file.pdf")
-assistant.add("https://www.youtube.com/watch?v=U9mJuUkhUzk")
-assistant.add("https://openai.com/blog/new-models-and-developer-products-announced-at-devday")
-```
-
-## Step-3: Chat with your AI Assistant
-
-```python Chat
-assistant.chat("How much OpenAI credits were offered to attendees during OpenAI DevDay?")
-# Response: 'Every attendee of OpenAI DevDay 2023 was offered $500 in OpenAI credits.'
-```
diff --git a/embedchain/docs/examples/poe_bot.mdx b/embedchain/docs/examples/poe_bot.mdx
deleted file mode 100644
index 58e831f22..000000000
--- a/embedchain/docs/examples/poe_bot.mdx
+++ /dev/null
@@ -1,59 +0,0 @@
----
-title: '🔮 Poe Bot'
----
-
-### 🚀 Getting started
-
-1. Install embedchain python package:
-
-```bash
-pip install fastapi-poe==0.0.16
-```
-
-2. Create a free account on [Poe](https://www.poe.com?utm_source=embedchain).
-3. Click "Create Bot" button on top left.
-4. Give it a handle and an optional description.
-5. Select `Use API`.
-6. Under `API URL` enter your server or ngrok address. You can use your machine's public IP or DNS. Otherwise, employ a proxy server like [ngrok](https://ngrok.com/) to make your local bot accessible.
-7. Copy your api key and paste it in `.env` as `POE_API_KEY`.
-8. You will need to set `OPENAI_API_KEY` for generating embeddings and using LLM. Copy your OpenAI API key from [here](https://platform.openai.com/account/api-keys) and paste it in `.env` as `OPENAI_API_KEY`.
-9. Now create your bot using the following code snippet.
-
-```bash
-# make sure that you have set OPENAI_API_KEY and POE_API_KEY in .env file
-from embedchain.bots import PoeBot
-
-poe_bot = PoeBot()
-
-# add as many data sources as you want
-poe_bot.add("https://en.wikipedia.org/wiki/Adam_D%27Angelo")
-poe_bot.add("https://www.youtube.com/watch?v=pJQVAqmKua8")
-
-# start the bot
-# this start the poe bot server on port 8080 by default
-poe_bot.start()
-```
-
-10. You can paste the above in a file called `your_script.py` and then simply do
-
-```bash
-python your_script.py
-```
-
-Now your bot will start running at port `8080` by default.
-
-11. You can refer the [Supported Data formats](https://docs.embedchain.ai/advanced/data_types) section to refer the supported data types in embedchain.
-
-12. Click `Run check` to make sure your machine can be reached.
-13. Make sure your bot is private if that's what you want.
-14. Click `Create bot` at the bottom to finally create the bot
-15. Now your bot is created.
-
-### 💬 How to use
-
-- To ask the bot questions, just type your query in the Poe interface:
-```text
-
-```
-
-- If you wish to add more data source to the bot, simply update your script and add as many `.add` as you like. You need to restart the server.
diff --git a/embedchain/docs/examples/rest-api/add-data.mdx b/embedchain/docs/examples/rest-api/add-data.mdx
deleted file mode 100644
index 05ed37968..000000000
--- a/embedchain/docs/examples/rest-api/add-data.mdx
+++ /dev/null
@@ -1,22 +0,0 @@
----
-openapi: post /{app_id}/add
----
-
-
-
-```bash Request
-curl --request POST \
- --url http://localhost:8080/{app_id}/add \
- -d "source=https://www.forbes.com/profile/elon-musk" \
- -d "data_type=web_page"
-```
-
-
-
-
-
-```json Response
-{ "response": "fec7fe91e6b2d732938a2ec2e32bfe3f" }
-```
-
-
diff --git a/embedchain/docs/examples/rest-api/chat.mdx b/embedchain/docs/examples/rest-api/chat.mdx
deleted file mode 100644
index 2571bf716..000000000
--- a/embedchain/docs/examples/rest-api/chat.mdx
+++ /dev/null
@@ -1,3 +0,0 @@
----
-openapi: post /{app_id}/chat
----
\ No newline at end of file
diff --git a/embedchain/docs/examples/rest-api/check-status.mdx b/embedchain/docs/examples/rest-api/check-status.mdx
deleted file mode 100644
index 0893cba4c..000000000
--- a/embedchain/docs/examples/rest-api/check-status.mdx
+++ /dev/null
@@ -1,20 +0,0 @@
----
-openapi: get /ping
----
-
-
-
-```bash Request
- curl --request GET \
- --url http://localhost:8080/ping
-```
-
-
-
-
-
-```json Response
-{ "ping": "pong" }
-```
-
-
diff --git a/embedchain/docs/examples/rest-api/create.mdx b/embedchain/docs/examples/rest-api/create.mdx
deleted file mode 100644
index 35863cea5..000000000
--- a/embedchain/docs/examples/rest-api/create.mdx
+++ /dev/null
@@ -1,96 +0,0 @@
----
-openapi: post /create
----
-
-
-
-```bash Request
-curl --request POST \
- --url http://localhost:8080/create?app_id=app1 \
- -F "config=@/path/to/config.yaml"
-```
-
-
-
-
-
-```json Response
-{ "response": "App created successfully. App ID: app1" }
-```
-
-
-
-By default we will use the opensource **gpt4all** model to get started. You can also specify your own config by uploading a config YAML file.
-
-For example, create a `config.yaml` file (adjust according to your requirements):
-
-```yaml
-app:
- config:
- id: "default-app"
-
-llm:
- provider: openai
- config:
- model: "gpt-4o-mini"
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
- prompt: |
- Use the following pieces of context to answer the query at the end.
- If you don't know the answer, just say that you don't know, don't try to make up an answer.
-
- $context
-
- Query: $query
-
- Helpful Answer:
-
-vectordb:
- provider: chroma
- config:
- collection_name: "rest-api-app"
- dir: db
- allow_reset: true
-
-embedder:
- provider: openai
- config:
- model: "text-embedding-ada-002"
-```
-
-To learn more about custom configurations, check out the [custom configurations docs](https://docs.embedchain.ai/advanced/configuration). To explore more examples of config yamls for embedchain, visit [embedchain/configs](https://github.com/embedchain/embedchain/tree/main/configs).
-
-Now, you can upload this config file in the request body.
-
-For example,
-
-```bash Request
-curl --request POST \
- --url http://localhost:8080/create?app_id=my-app \
- -F "config=@/path/to/config.yaml"
-```
-
-**Note:** To use custom models, an **API key** might be required. Refer to the table below to determine the necessary API key for your provider.
-
-| Keys | Providers |
-| -------------------------- | ------------------------------ |
-| `OPENAI_API_KEY ` | OpenAI, Azure OpenAI, Jina etc |
-| `OPENAI_API_TYPE` | Azure OpenAI |
-| `OPENAI_API_BASE` | Azure OpenAI |
-| `OPENAI_API_VERSION` | Azure OpenAI |
-| `COHERE_API_KEY` | Cohere |
-| `TOGETHER_API_KEY` | Together |
-| `ANTHROPIC_API_KEY` | Anthropic |
-| `JINACHAT_API_KEY` | Jina |
-| `HUGGINGFACE_ACCESS_TOKEN` | Huggingface |
-| `REPLICATE_API_TOKEN` | LLAMA2 |
-
-To add env variables, you can simply run the docker command with the `-e` flag.
-
-For example,
-
-```bash
-docker run --name embedchain -p 8080:8080 -e OPENAI_API_KEY= embedchain/rest-api:latest
-```
\ No newline at end of file
diff --git a/embedchain/docs/examples/rest-api/delete.mdx b/embedchain/docs/examples/rest-api/delete.mdx
deleted file mode 100644
index 3aada3398..000000000
--- a/embedchain/docs/examples/rest-api/delete.mdx
+++ /dev/null
@@ -1,21 +0,0 @@
----
-openapi: delete /{app_id}/delete
----
-
-
-
-
-```bash Request
- curl --request DELETE \
- --url http://localhost:8080/{app_id}/delete
-```
-
-
-
-
-
-```json Response
-{ "response": "App with id {app_id} deleted successfully." }
-```
-
-
diff --git a/embedchain/docs/examples/rest-api/deploy.mdx b/embedchain/docs/examples/rest-api/deploy.mdx
deleted file mode 100644
index b72f91da0..000000000
--- a/embedchain/docs/examples/rest-api/deploy.mdx
+++ /dev/null
@@ -1,22 +0,0 @@
----
-openapi: post /{app_id}/deploy
----
-
-
-
-
-```bash Request
-curl --request POST \
- --url http://localhost:8080/{app_id}/deploy \
- -d "api_key=ec-xxxx"
-```
-
-
-
-
-
-```json Response
-{ "response": "App deployed successfully." }
-```
-
-
diff --git a/embedchain/docs/examples/rest-api/get-all-apps.mdx b/embedchain/docs/examples/rest-api/get-all-apps.mdx
deleted file mode 100644
index 6f603f9a6..000000000
--- a/embedchain/docs/examples/rest-api/get-all-apps.mdx
+++ /dev/null
@@ -1,33 +0,0 @@
----
-openapi: get /apps
----
-
-
-
-```bash Request
-curl --request GET \
- --url http://localhost:8080/apps
-```
-
-
-
-
-
-```json Response
-{
- "results": [
- {
- "config": "config1.yaml",
- "id": 1,
- "app_id": "app1"
- },
- {
- "config": "config2.yaml",
- "id": 2,
- "app_id": "app2"
- }
- ]
-}
-```
-
-
diff --git a/embedchain/docs/examples/rest-api/get-data.mdx b/embedchain/docs/examples/rest-api/get-data.mdx
deleted file mode 100644
index 0c960e6cb..000000000
--- a/embedchain/docs/examples/rest-api/get-data.mdx
+++ /dev/null
@@ -1,28 +0,0 @@
----
-openapi: get /{app_id}/data
----
-
-
-
-```bash Request
-curl --request GET \
- --url http://localhost:8080/{app_id}/data
-```
-
-
-
-
-
-```json Response
-{
- "results": [
- {
- "data_type": "web_page",
- "data_value": "https://www.forbes.com/profile/elon-musk/",
- "metadata": "null"
- }
- ]
-}
-```
-
-
diff --git a/embedchain/docs/examples/rest-api/getting-started.mdx b/embedchain/docs/examples/rest-api/getting-started.mdx
deleted file mode 100644
index 5501792b6..000000000
--- a/embedchain/docs/examples/rest-api/getting-started.mdx
+++ /dev/null
@@ -1,294 +0,0 @@
----
-title: "🌍 Getting Started"
----
-
-## Quickstart
-
-To use Embedchain as a REST API service, run the following command:
-
-```bash
-docker run --name embedchain -p 8080:8080 embedchain/rest-api:latest
-```
-
-Navigate to [http://localhost:8080/docs](http://localhost:8080/docs) to interact with the API. There is a full-fledged Swagger docs playground with all the information about the API endpoints.
-
-
-
-## ⚡ Steps to get started
-
-
-
-
-
- ```bash
- curl --request POST "http://localhost:8080/create?app_id=my-app" \
- -H "accept: application/json"
- ```
-
-
- ```python
- import requests
-
- url = "http://localhost:8080/create?app_id=my-app"
-
- payload={}
-
- response = requests.request("POST", url, data=payload)
-
- print(response)
- ```
-
-
- ```javascript
- const data = fetch("http://localhost:8080/create?app_id=my-app", {
- method: "POST",
- }).then((res) => res.json());
-
- console.log(data);
- ```
-
-
- ```go
- package main
-
- import (
- "fmt"
- "net/http"
- "io/ioutil"
- )
-
- func main() {
-
- url := "http://localhost:8080/create?app_id=my-app"
-
- payload := strings.NewReader("")
-
- req, _ := http.NewRequest("POST", url, payload)
-
- req.Header.Add("Content-Type", "application/json")
-
- res, _ := http.DefaultClient.Do(req)
-
- defer res.Body.Close()
- body, _ := ioutil.ReadAll(res.Body)
-
- fmt.Println(res)
- fmt.Println(string(body))
-
- }
- ```
-
-
-
-
-
-
-
- ```bash
- curl --request POST \
- --url http://localhost:8080/my-app/add \
- -d "source=https://www.forbes.com/profile/elon-musk" \
- -d "data_type=web_page"
- ```
-
-
- ```python
- import requests
-
- url = "http://localhost:8080/my-app/add"
-
- payload = "source=https://www.forbes.com/profile/elon-musk&data_type=web_page"
- headers = {}
-
- response = requests.request("POST", url, headers=headers, data=payload)
-
- print(response)
- ```
-
-
- ```javascript
- const data = fetch("http://localhost:8080/my-app/add", {
- method: "POST",
- body: "source=https://www.forbes.com/profile/elon-musk&data_type=web_page",
- }).then((res) => res.json());
-
- console.log(data);
- ```
-
-
- ```go
- package main
-
- import (
- "fmt"
- "strings"
- "net/http"
- "io/ioutil"
- )
-
- func main() {
-
- url := "http://localhost:8080/my-app/add"
-
- payload := strings.NewReader("source=https://www.forbes.com/profile/elon-musk&data_type=web_page")
-
- req, _ := http.NewRequest("POST", url, payload)
-
- req.Header.Add("Content-Type", "application/x-www-form-urlencoded")
-
- res, _ := http.DefaultClient.Do(req)
-
- defer res.Body.Close()
- body, _ := ioutil.ReadAll(res.Body)
-
- fmt.Println(res)
- fmt.Println(string(body))
-
- }
- ```
-
-
-
-
-
-
-
- ```bash
- curl --request POST \
- --url http://localhost:8080/my-app/query \
- -d "query=Who is Elon Musk?"
- ```
-
-
- ```python
- import requests
-
- url = "http://localhost:8080/my-app/query"
-
- payload = "query=Who is Elon Musk?"
- headers = {}
-
- response = requests.request("POST", url, headers=headers, data=payload)
-
- print(response)
- ```
-
-
- ```javascript
- const data = fetch("http://localhost:8080/my-app/query", {
- method: "POST",
- body: "query=Who is Elon Musk?",
- }).then((res) => res.json());
-
- console.log(data);
- ```
-
-
- ```go
- package main
-
- import (
- "fmt"
- "strings"
- "net/http"
- "io/ioutil"
- )
-
- func main() {
-
- url := "http://localhost:8080/my-app/query"
-
- payload := strings.NewReader("query=Who is Elon Musk?")
-
- req, _ := http.NewRequest("POST", url, payload)
-
- req.Header.Add("Content-Type", "application/x-www-form-urlencoded")
-
- res, _ := http.DefaultClient.Do(req)
-
- defer res.Body.Close()
- body, _ := ioutil.ReadAll(res.Body)
-
- fmt.Println(res)
- fmt.Println(string(body))
-
- }
- ```
-
-
-
-
-
-
-
- ```bash
- curl --request POST \
- --url http://localhost:8080/my-app/deploy \
- -d "api_key=ec-xxxx"
- ```
-
-
- ```python
- import requests
-
- url = "http://localhost:8080/my-app/deploy"
-
- payload = "api_key=ec-xxxx"
-
- response = requests.request("POST", url, data=payload)
-
- print(response)
- ```
-
-
- ```javascript
- const data = fetch("http://localhost:8080/my-app/deploy", {
- method: "POST",
- body: "api_key=ec-xxxx",
- }).then((res) => res.json());
-
- console.log(data);
- ```
-
-
- ```go
- package main
-
- import (
- "fmt"
- "strings"
- "net/http"
- "io/ioutil"
- )
-
- func main() {
-
- url := "http://localhost:8080/my-app/deploy"
-
- payload := strings.NewReader("api_key=ec-xxxx")
-
- req, _ := http.NewRequest("POST", url, payload)
-
- req.Header.Add("Content-Type", "application/x-www-form-urlencoded")
-
- res, _ := http.DefaultClient.Do(req)
-
- defer res.Body.Close()
- body, _ := ioutil.ReadAll(res.Body)
-
- fmt.Println(res)
- fmt.Println(string(body))
-
- }
- ```
-
-
-
-
-
-
-And you're ready! 🎉
-
-If you run into issues, please feel free to contact us using below links:
-
-
diff --git a/embedchain/docs/examples/rest-api/query.mdx b/embedchain/docs/examples/rest-api/query.mdx
deleted file mode 100644
index 2d647e505..000000000
--- a/embedchain/docs/examples/rest-api/query.mdx
+++ /dev/null
@@ -1,21 +0,0 @@
----
-openapi: post /{app_id}/query
----
-
-
-
-```bash Request
-curl --request POST \
- --url http://localhost:8080/{app_id}/query \
- -d "query=who is Elon Musk?"
-```
-
-
-
-
-
-```json Response
-{ "response": "Net worth of Elon Musk is $218 Billion." }
-```
-
-
diff --git a/embedchain/docs/examples/showcase.mdx b/embedchain/docs/examples/showcase.mdx
deleted file mode 100644
index d614c3b00..000000000
--- a/embedchain/docs/examples/showcase.mdx
+++ /dev/null
@@ -1,115 +0,0 @@
----
-title: '🎪 Community showcase'
----
-
-Embedchain community has been super active in creating demos on top of Embedchain. On this page, we showcase all the apps, blogs, videos, and tutorials created by the community. ❤️
-
-## Apps
-
-### Open Source
-
-- [My GSoC23 bot- Streamlit chat](https://github.com/lucifertrj/EmbedChain_GSoC23_BOT) by Tarun Jain
-- [Discord Bot for LLM chat](https://github.com/Reidond/discord_bots_playground/tree/c8b0c36541e4b393782ee506804c4b6962426dd6/python/chat-channel-bot) by Reidond
-- [EmbedChain-Streamlit-Docker App](https://github.com/amjadraza/embedchain-streamlit-app) by amjadraza
-- [Harry Potter Philosphers Stone Bot](https://github.com/vinayak-kempawad/Harry_Potter_Philosphers_Stone_Bot/) by Vinayak Kempawad, ([LinkedIn post](https://www.linkedin.com/feed/update/urn:li:activity:7080907532155686912/))
-- [LLM bot trained on own messages](https://github.com/Harin329/harinBot) by Hao Wu
-
-### Closed Source
-
-- [Taobot.io](https://taobot.io) - chatbot & knowledgebase hybrid by [cachho](https://github.com/cachho)
-- [Create Instant ChatBot 🤖 using embedchain](https://databutton.com/v/h3e680h9) by Avra, ([Tweet](https://twitter.com/Avra_b/status/1674704745154641920/))
-- [JOBO 🤖 — The AI-driven sidekick to craft your resume](https://try-jobo.com/) by Enrico Willemse, ([LinkedIn Post](https://www.linkedin.com/posts/enrico-willemse_jobai-gptfun-embedchain-activity-7090340080879374336-ueLB/))
-- [Explore Your Knowledge Base: Interactive chats over various forms of documents](https://chatdocs.dkedar.com/) by Kedar Dabhadkar, ([LinkedIn Post](https://www.linkedin.com/posts/dkedar7_machinelearning-llmops-activity-7092524836639424513-2O3L/))
-- [Chatbot trained on 1000+ videos of Ester hicks the co-author behind the famous book Secret](https://askabraham.tokenofme.io/) by Mohan Kumar
-
-
-## Templates
-
-### Replit
-- [Embedchain Chat Bot](https://replit.com/@taranjeet1/Embedchain-Chat-Bot) by taranjeetio
-- [Embedchain Memory Chat Bot Template](https://replit.com/@taranjeetio/Embedchain-Memory-Chat-Bot-Template) by taranjeetio
-- [Chatbot app to demonstrate question-answering using retrieved information](https://replit.com/@AllisonMorrell/EmbedChainlitPublic) by Allison Morrell, ([LinkedIn Post](https://www.linkedin.com/posts/allison-morrell-2889275a_retrievalbot-screenshots-activity-7080339991754649600-wihZ/))
-
-## Posts
-
-### Blogs
-
-- [Customer Service LINE Bot](https://www.evanlin.com/langchain-embedchain/) by Evan Lin
-- [Chatbot in Under 5 mins using Embedchain](https://medium.com/@ayush.wattal/chatbot-in-under-5-mins-using-embedchain-a4f161fcf9c5) by Ayush Wattal
-- [Understanding what the LLM framework embedchain does](https://zenn.dev/hijikix/articles/4bc8d60156a436) by Daisuke Hashimoto
-- [In bed with GPT and Node.js](https://dev.to/worldlinetech/in-bed-with-gpt-and-nodejs-4kh2) by Raphaël Semeteys, ([LinkedIn Post](https://www.linkedin.com/posts/raphaelsemeteys_in-bed-with-gpt-and-nodejs-activity-7088113552326029313-nn87/))
-- [Using Embedchain — A powerful LangChain Python wrapper to build Chat Bots even faster!⚡](https://medium.com/@avra42/using-embedchain-a-powerful-langchain-python-wrapper-to-build-chat-bots-even-faster-35c12994a360) by Avra, ([Tweet](https://twitter.com/Avra_b/status/1686767751560310784/))
-- [What is the Embedchain library?](https://jahaniwww.com/%da%a9%d8%aa%d8%a7%d8%a8%d8%ae%d8%a7%d9%86%d9%87-embedchain/) by Ali Jahani, ([LinkedIn Post](https://www.linkedin.com/posts/ajahani_aepaetaeqaexaggahyaeu-aetaexaesabraeaaeqaepaeu-activity-7097605202135904256-ppU-/))
-- [LangChain is Nice, But Have You Tried EmbedChain ?](https://medium.com/thoughts-on-machine-learning/langchain-is-nice-but-have-you-tried-embedchain-215a34421cde) by FS Ndzomga, ([Tweet](https://twitter.com/ndzfs/status/1695583640372035951/))
-- [Simplest Method to Build a Custom Chatbot with GPT-3.5 (via Embedchain)](https://www.ainewsletter.today/p/simplest-method-to-build-a-custom) by Arjun, ([Tweet](https://twitter.com/aiguy_arjun/status/1696393808467091758/))
-
-### LinkedIn
-
-- [What is embedchain](https://www.linkedin.com/posts/activity-7079393104423698432-wRyi/) by Rithesh Sreenivasan
-- [Building a chatbot with EmbedChain](https://www.linkedin.com/posts/activity-7078434598984060928-Zdso/) by Lior Sinclair
-- [Making chatbot without vs with embedchain](https://www.linkedin.com/posts/kalyanksnlp_llms-chatbots-langchain-activity-7077453416221863936-7N1L/) by Kalyan KS
-- [EmbedChain - very intuitive, first you index your data and then query!](https://www.linkedin.com/posts/shubhamsaboo_embedchain-a-framework-to-easily-create-activity-7079535460699557888-ad1X/) by Shubham Saboo
-- [EmbedChain - Harnessing power of LLM](https://www.linkedin.com/posts/uditsaini_chatbotrevolution-llmpoweredbots-embedchainframework-activity-7077520356827181056-FjTK/) by Udit S.
-- [AI assistant for ABBYY Vantage](https://www.linkedin.com/posts/maximevermeir_llm-github-abbyy-activity-7081658972071424000-fXfZ/) by Maxime V.
-- [About embedchain](https://www.linkedin.com/feed/update/urn:li:activity:7080984218914189312/) by Morris Lee
-- [How to use Embedchain](https://www.linkedin.com/posts/nehaabansal_github-embedchainembedchain-framework-activity-7085830340136595456-kbW5/) by Neha Bansal
-- [Youtube/Webpage summary for Energy Study](https://www.linkedin.com/posts/bar%C4%B1%C5%9F-sanl%C4%B1-34b82715_enerji-python-activity-7082735341563977730-Js0U/) by Barış Sanlı, ([Tweet](https://twitter.com/barissanli/status/1676968784979193857/))
-- [Demo: How to use Embedchain? (Contains Collab Notebook link)](https://www.linkedin.com/posts/liorsinclair_embedchain-is-getting-a-lot-of-traction-because-activity-7103044695995424768-RckT/) by Lior Sinclair
-
-### Twitter
-
-- [What is embedchain](https://twitter.com/AlphaSignalAI/status/1672668574450847745) by Lior
-- [Building a chatbot with Embedchain](https://twitter.com/Saboo_Shubham_/status/1673537044419686401) by Shubham Saboo
-- [Chatbot docker image behind an API with yaml configs with Embedchain](https://twitter.com/tricalt/status/1678411430192730113/) by Vasilije
-- [Build AI powered PDF chatbot with just five lines of Python code with Embedchain!](https://twitter.com/Saboo_Shubham_/status/1676627104866156544/) by Shubham Saboo
-- [Chatbot against a youtube video using embedchain](https://twitter.com/smaameri/status/1675201443043704834/) by Sami Maameri
-- [Highlights of EmbedChain](https://twitter.com/carl_AIwarts/status/1673542204328120321/) by carl_AIwarts
-- [Build Llama-2 chatbot in less than 5 minutes](https://twitter.com/Saboo_Shubham_/status/1682168956918833152/) by Shubham Saboo
-- [All cool features of embedchain](https://twitter.com/DhravyaShah/status/1683497882438217728/) by Dhravya Shah, ([LinkedIn Post](https://www.linkedin.com/posts/dhravyashah_what-if-i-tell-you-that-you-can-make-an-ai-activity-7089459599287726080-ZIYm/))
-- [Read paid Medium articles for Free using embedchain](https://twitter.com/kumarkaushal_/status/1688952961622585344) by Kaushal Kumar
-
-## Videos
-
-- [Embedchain in one shot](https://www.youtube.com/watch?v=vIhDh7H73Ww&t=82s) by AI with Tarun
-- [embedChain Create LLM powered bots over any dataset Python Demo Tesla Neurallink Chatbot Example](https://www.youtube.com/watch?v=bJqAn22a6Gc) by Rithesh Sreenivasan
-- [Embedchain - NEW 🔥 Langchain BABY to build LLM Bots](https://www.youtube.com/watch?v=qj_GNQ06I8o) by 1littlecoder
-- [EmbedChain -- NEW!: Build LLM-Powered Bots with Any Dataset](https://www.youtube.com/watch?v=XmaBezzGHu4) by DataInsightEdge
-- [Chat With Your PDFs in less than 10 lines of code! EMBEDCHAIN tutorial](https://www.youtube.com/watch?v=1ugkcsAcw44) by Phani Reddy
-- [How To Create A Custom Knowledge AI Powered Bot | Install + How To Use](https://www.youtube.com/watch?v=VfCrIiAst-c) by The Ai Solopreneur
-- [Build Custom Chatbot in 6 min with this Framework [Beginner Friendly]](https://www.youtube.com/watch?v=-8HxOpaFySM) by Maya Akim
-- [embedchain-streamlit-app](https://www.youtube.com/watch?v=3-9GVd-3v74) by Amjad Raza
-- [🤖CHAT with ANY ONLINE RESOURCES using EMBEDCHAIN - a LangChain wrapper, in few lines of code !](https://www.youtube.com/watch?v=Mp7zJe4TIdM) by Avra
-- [Building resource-driven LLM-powered bots with Embedchain](https://www.youtube.com/watch?v=IVfcAgxTO4I) by BugBytes
-- [embedchain-streamlit-demo](https://www.youtube.com/watch?v=yJAWB13FhYQ) by Amjad Raza
-- [Embedchain - create your own AI chatbots using open source models](https://www.youtube.com/shorts/O3rJWKwSrWE) by Dhravya Shah
-- [AI ChatBot in 5 lines Python Code](https://www.youtube.com/watch?v=zjWvLJLksv8) by Data Engineering
-- [Interview with Karl Marx](https://www.youtube.com/watch?v=5Y4Tscwj1xk) by Alexander Ray Williams
-- [Vlog where we try to build a bot based on our content on the internet](https://www.youtube.com/watch?v=I2w8CWM3bx4) by DV, ([Tweet](https://twitter.com/dvcoolster/status/1688387017544261632))
-- [CHAT with ANY ONLINE RESOURCES using EMBEDCHAIN|STREAMLIT with MEMORY |All OPENSOURCE](https://www.youtube.com/watch?v=TqQIHWoWTDQ&pp=ygUKZW1iZWRjaGFpbg%3D%3D) by DataInsightEdge
-- [Build POWERFUL LLM Bots EASILY with Your Own Data - Embedchain - Langchain 2.0? (Tutorial)](https://www.youtube.com/watch?v=jE24Y_GasE8) by WorldofAI, ([Tweet](https://twitter.com/intheworldofai/status/1696229166922780737))
-- [Embedchain: An AI knowledge base assistant for customizing enterprise private data, which can be connected to discord, whatsapp, slack, tele and other terminals (with gradio to build a request interface) in Chinese](https://www.youtube.com/watch?v=5RZzCJRk-d0) by AIGC LINK
-- [Embedchain Introduction](https://www.youtube.com/watch?v=Jet9zAqyggI) by Fahd Mirza
-
-## Mentions
-
-### Github repos
-
-- [Awesome-LLM](https://github.com/Hannibal046/Awesome-LLM)
-- [awesome-chatgpt-api](https://github.com/reorx/awesome-chatgpt-api)
-- [awesome-langchain](https://github.com/kyrolabs/awesome-langchain)
-- [Awesome-Prompt-Engineering](https://github.com/promptslab/Awesome-Prompt-Engineering)
-- [awesome-chatgpt](https://github.com/eon01/awesome-chatgpt)
-- [Awesome-LLMOps](https://github.com/tensorchord/Awesome-LLMOps)
-- [awesome-generative-ai](https://github.com/filipecalegario/awesome-generative-ai)
-- [awesome-gpt](https://github.com/formulahendry/awesome-gpt)
-- [awesome-ChatGPT-repositories](https://github.com/taishi-i/awesome-ChatGPT-repositories)
-- [awesome-gpt-prompt-engineering](https://github.com/snwfdhmp/awesome-gpt-prompt-engineering)
-- [awesome-chatgpt](https://github.com/awesome-chatgpt/awesome-chatgpt)
-- [awesome-llm-and-aigc](https://github.com/sjinzh/awesome-llm-and-aigc)
-- [awesome-compbio-chatgpt](https://github.com/csbl-br/awesome-compbio-chatgpt)
-- [Awesome-LLM4Tool](https://github.com/OpenGVLab/Awesome-LLM4Tool)
-
-## Meetups
-
-- [Dash and ChatGPT: Future of AI-enabled apps 30/08/23](https://go.plotly.com/dash-chatgpt)
-- [Pie & AI: Bangalore - Build end-to-end LLM app using Embedchain 01/09/23](https://www.eventbrite.com/e/pie-ai-bangalore-build-end-to-end-llm-app-using-embedchain-tickets-698045722547)
diff --git a/embedchain/docs/examples/slack-AI.mdx b/embedchain/docs/examples/slack-AI.mdx
deleted file mode 100644
index 7efaba279..000000000
--- a/embedchain/docs/examples/slack-AI.mdx
+++ /dev/null
@@ -1,67 +0,0 @@
-[Embedchain Examples Repo](https://github.com/embedchain/examples) contains code on how to build your own Slack AI to chat with the unstructured data lying in your slack channels.
-
-
-
-## Getting started
-
-Create a Slack AI involves 3 steps
-
-* Create slack user
-* Set environment variables
-* Run the app locally
-
-### Step 1: Create Slack user token
-
-Follow the steps given below to fetch your slack user token to get data through Slack APIs:
-
-1. Create a workspace on Slack if you don’t have one already by clicking [here](https://slack.com/intl/en-in/).
-2. Create a new App on your Slack account by going [here](https://api.slack.com/apps).
-3. Select `From Scratch`, then enter the App Name and select your workspace.
-4. Navigate to `OAuth & Permissions` tab from the left sidebar and go to the `scopes` section. Add the following scopes under `User Token Scopes`:
-
- ```
- # Following scopes are needed for reading channel history
- channels:history
- channels:read
-
- # Following scopes are needed to fetch list of channels from slack
- groups:read
- mpim:read
- im:read
- ```
-
-5. Click on the `Install to Workspace` button under `OAuth Tokens for Your Workspace` section in the same page and install the app in your slack workspace.
-6. After installing the app you will see the `User OAuth Token`, save that token as you will need to configure it as `SLACK_USER_TOKEN` for this demo.
-
-### Step 2: Set environment variables
-
-Navigate to `api` folder and set your `HUGGINGFACE_ACCESS_TOKEN` and `SLACK_USER_TOKEN` in `.env.example` file. Then rename the `.env.example` file to `.env`.
-
-
-
-By default, we use `Mixtral` model from Hugging Face. However, if you prefer to use OpenAI model, then set `OPENAI_API_KEY` instead of `HUGGINGFACE_ACCESS_TOKEN` along with `SLACK_USER_TOKEN` in `.env` file, and update the code in `api/utils/app.py` file to use OpenAI model instead of Hugging Face model.
-
-
-### Step 3: Run app locally
-
-Follow the instructions given below to run app locally based on your development setup (with docker or without docker):
-
-#### With docker
-
-```bash
-docker-compose build
-ec start --docker
-```
-
-#### Without docker
-
-```bash
-ec install-reqs
-ec start
-```
-
-Finally, you will have the Slack AI frontend running on http://localhost:3000. You can also access the REST APIs on http://localhost:8000.
-
-## Credits
-
-This demo was built using the Embedchain's [full stack demo template](https://docs.embedchain.ai/get-started/full-stack). Follow the instructions [given here](https://docs.embedchain.ai/get-started/full-stack) to create your own full stack RAG application.
diff --git a/embedchain/docs/examples/slack_bot.mdx b/embedchain/docs/examples/slack_bot.mdx
deleted file mode 100644
index 034c821d2..000000000
--- a/embedchain/docs/examples/slack_bot.mdx
+++ /dev/null
@@ -1,50 +0,0 @@
----
-title: '💼 Slack Bot'
----
-
-### 🖼️ Setup
-
-1. Create a workspace on Slack if you don't have one already by clicking [here](https://slack.com/intl/en-in/).
-2. Create a new App on your Slack account by going [here](https://api.slack.com/apps).
-3. Select `From Scratch`, then enter the Bot Name and select your workspace.
-4. On the left Sidebar, go to `OAuth and Permissions` and add the following scopes under `Bot Token Scopes`:
-```text
-app_mentions:read
-channels:history
-channels:read
-chat:write
-```
-5. Now select the option `Install to Workspace` and after it's done, copy the `Bot User OAuth Token` and set it in your secrets as `SLACK_BOT_TOKEN`.
-6. Run your bot now,
-
-
- ```bash
- docker run --name slack-bot -e OPENAI_API_KEY=sk-xxx -e SLACK_BOT_TOKEN=xxx -p 8000:8000 embedchain/slack-bot
- ```
-
-
- ```bash
- pip install --upgrade "embedchain[slack]"
- python3 -m embedchain.bots.slack --port 8000
- ```
-
-
-7. Expose your bot to the internet. You can use your machine's public IP or DNS. Otherwise, employ a proxy server like [ngrok](https://ngrok.com/) to make your local bot accessible.
-8. On the Slack API website go to `Event Subscriptions` on the left Sidebar and turn on `Enable Events`.
-9. In `Request URL`, enter your server or ngrok address.
-10. After it gets verified, click on `Subscribe to bot events`, add `message.channels` Bot User Event and click on `Save Changes`.
-11. Now go to your workspace, right click on the bot name in the sidebar, click `view app details`, then `add this app to a channel`.
-
-### 🚀 Usage Instructions
-
-- Go to the channel where you have added your bot.
-- To add data sources to the bot, use the command:
-```text
-add
-```
-- To ask queries from the bot, use the command:
-```text
-query
-```
-
-🎉 Happy Chatting! 🎉
diff --git a/embedchain/docs/examples/telegram_bot.mdx b/embedchain/docs/examples/telegram_bot.mdx
deleted file mode 100644
index 14f17e90c..000000000
--- a/embedchain/docs/examples/telegram_bot.mdx
+++ /dev/null
@@ -1,51 +0,0 @@
----
-title: "📱 Telegram Bot"
----
-
-### 🖼️ Template Setup
-
-- Open the Telegram app and search for the `BotFather` user.
-- Start a chat with BotFather and use the `/newbot` command to create a new bot.
-- Follow the instructions to choose a name and username for your bot.
-- Once the bot is created, BotFather will provide you with a unique token for your bot.
-
-
-
- ```bash
- docker run --name telegram-bot -e OPENAI_API_KEY=sk-xxx -e TELEGRAM_BOT_TOKEN=xxx -p 8000:8000 embedchain/telegram-bot
- ```
-
-
- If you wish to use **Docker**, you would need to host your bot on a server.
- You can use [ngrok](https://ngrok.com/) to expose your localhost to the
- internet and then set the webhook using the ngrok URL.
-
-
-
-
-
- Fork **[this](https://replit.com/@taranjeetio/EC-Telegram-Bot-Template?v=1#README.md)** replit template.
-
-
- - Set your `OPENAI_API_KEY` in Secrets.
- - Set the unique token as `TELEGRAM_BOT_TOKEN` in Secrets.
-
-
-
-
-
-- Click on `Run` in the replit container and a URL will get generated for your bot.
-- Now set your webhook by running the following link in your browser:
-
-```url
-https://api.telegram.org/bot/setWebhook?url=
-```
-
-- When you get a successful response in your browser, your bot is ready to be used.
-
-### 🚀 Usage Instructions
-
-- Open your bot by searching for it using the bot name or bot username.
-- Click on `Start` or type `/start` and follow the on screen instructions.
-
-🎉 Happy Chatting! 🎉
diff --git a/embedchain/docs/examples/whatsapp_bot.mdx b/embedchain/docs/examples/whatsapp_bot.mdx
deleted file mode 100644
index 16a8c504a..000000000
--- a/embedchain/docs/examples/whatsapp_bot.mdx
+++ /dev/null
@@ -1,55 +0,0 @@
----
-title: '💬 WhatsApp Bot'
----
-
-### 🚀 Getting started
-
-1. Install embedchain python package:
-
-```bash
-pip install --upgrade embedchain
-```
-
-2. Launch your WhatsApp bot:
-
-
-
- ```bash
- docker run --name whatsapp-bot -e OPENAI_API_KEY=sk-xxx -p 8000:8000 embedchain/whatsapp-bot
- ```
-
-
- ```bash
- python -m embedchain.bots.whatsapp --port 5000
- ```
-
-
-
-
-If your bot needs to be accessible online, use your machine's public IP or DNS. Otherwise, employ a proxy server like [ngrok](https://ngrok.com/) to make your local bot accessible.
-
-3. Create a free account on [Twilio](https://www.twilio.com/try-twilio)
- - Set up a WhatsApp Sandbox in your Twilio dashboard. Access it via the left sidebar: `Messaging > Try it out > Send a WhatsApp Message`.
- - Follow on-screen instructions to link a phone number for chatting with your bot
- - Copy your bot's public URL, add /chat at the end, and paste it in Twilio's WhatsApp Sandbox settings under "When a message comes in". Save the settings.
-
-- Copy your bot's public url, append `/chat` at the end and paste it under `When a message comes in` under the `Sandbox settings` for Whatsapp in Twilio. Save your settings.
-
-### 💬 How to use
-
-- To connect a new number or reconnect an old one in the Sandbox, follow Twilio's instructions.
-- To include data sources, use this command:
-```text
-add
-```
-
-- To ask the bot questions, just type your query:
-```text
-
-```
-
-### Example
-
-Here is an example of Elon Musk WhatsApp Bot that we created:
-
-
diff --git a/embedchain/docs/favicon.png b/embedchain/docs/favicon.png
deleted file mode 100644
index 35494d9ea..000000000
Binary files a/embedchain/docs/favicon.png and /dev/null differ
diff --git a/embedchain/docs/get-started/deployment.mdx b/embedchain/docs/get-started/deployment.mdx
deleted file mode 100644
index 87e72dcbb..000000000
--- a/embedchain/docs/get-started/deployment.mdx
+++ /dev/null
@@ -1,22 +0,0 @@
----
-title: 'Overview'
-description: 'Deploy your RAG application to production'
----
-
-After successfully setting up and testing your RAG app locally, the next step is to deploy it to a hosting service to make it accessible to a wider audience. Embedchain provides integration with different cloud providers so that you can seamlessly deploy your RAG applications to production without having to worry about going through the cloud provider instructions. Embedchain does all the heavy lifting for you.
-
-
-
-
-
-
-
-
-
-
-
-## Seeking help?
-
-If you run into issues with deployment, please feel free to reach out to us via any of the following methods:
-
-
diff --git a/embedchain/docs/get-started/faq.mdx b/embedchain/docs/get-started/faq.mdx
deleted file mode 100644
index 3acae2671..000000000
--- a/embedchain/docs/get-started/faq.mdx
+++ /dev/null
@@ -1,191 +0,0 @@
----
-title: ❓ FAQs
-description: 'Collections of all the frequently asked questions'
----
-
-
-Yes, it does. Please refer to the [OpenAI Assistant docs page](/examples/openai-assistant).
-
-
-Use the model provided on huggingface: `mistralai/Mistral-7B-v0.1`
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ["HUGGINGFACE_ACCESS_TOKEN"] = "hf_your_token"
-
-app = App.from_config("huggingface.yaml")
-```
-```yaml huggingface.yaml
-llm:
- provider: huggingface
- config:
- model: 'mistralai/Mistral-7B-v0.1'
- temperature: 0.5
- max_tokens: 1000
- top_p: 0.5
- stream: false
-
-embedder:
- provider: huggingface
- config:
- model: 'sentence-transformers/all-mpnet-base-v2'
-```
-
-
-
-Use the model `gpt-4-turbo` provided my openai.
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ['OPENAI_API_KEY'] = 'xxx'
-
-# load llm configuration from gpt4_turbo.yaml file
-app = App.from_config(config_path="gpt4_turbo.yaml")
-```
-
-```yaml gpt4_turbo.yaml
-llm:
- provider: openai
- config:
- model: 'gpt-4-turbo'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-```
-
-
-
-
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ['OPENAI_API_KEY'] = 'xxx'
-
-# load llm configuration from gpt4.yaml file
-app = App.from_config(config_path="gpt4.yaml")
-```
-
-```yaml gpt4.yaml
-llm:
- provider: openai
- config:
- model: 'gpt-4'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-```
-
-
-
-
-
-
-```python main.py
-from embedchain import App
-
-# load llm configuration from opensource.yaml file
-app = App.from_config(config_path="opensource.yaml")
-```
-
-```yaml opensource.yaml
-llm:
- provider: gpt4all
- config:
- model: 'orca-mini-3b-gguf2-q4_0.gguf'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: false
-
-embedder:
- provider: gpt4all
- config:
- model: 'all-MiniLM-L6-v2'
-```
-
-
-
-
-You can achieve this by setting `stream` to `true` in the config file.
-
-
-```yaml openai.yaml
-llm:
- provider: openai
- config:
- model: 'gpt-4o-mini'
- temperature: 0.5
- max_tokens: 1000
- top_p: 1
- stream: true
-```
-
-```python main.py
-import os
-from embedchain import App
-
-os.environ['OPENAI_API_KEY'] = 'sk-xxx'
-
-app = App.from_config(config_path="openai.yaml")
-
-app.add("https://www.forbes.com/profile/elon-musk")
-
-response = app.query("What is the net worth of Elon Musk?")
-# response will be streamed in stdout as it is generated.
-```
-
-
-
-
- Set up the app by adding an `id` in the config file. This keeps the data for future use. You can include this `id` in the yaml config or input it directly in `config` dict.
- ```python app1.py
- import os
- from embedchain import App
-
- os.environ['OPENAI_API_KEY'] = 'sk-xxx'
-
- app1 = App.from_config(config={
- "app": {
- "config": {
- "id": "your-app-id",
- }
- }
- })
-
- app1.add("https://www.forbes.com/profile/elon-musk")
-
- response = app1.query("What is the net worth of Elon Musk?")
- ```
- ```python app2.py
- import os
- from embedchain import App
-
- os.environ['OPENAI_API_KEY'] = 'sk-xxx'
-
- app2 = App.from_config(config={
- "app": {
- "config": {
- # this will persist and load data from app1 session
- "id": "your-app-id",
- }
- }
- })
-
- response = app2.query("What is the net worth of Elon Musk?")
- ```
-
-
-
-#### Still have questions?
-If docs aren't sufficient, please feel free to reach out to us using one of the following methods:
-
-
diff --git a/embedchain/docs/get-started/full-stack.mdx b/embedchain/docs/get-started/full-stack.mdx
deleted file mode 100644
index cd45a0b9e..000000000
--- a/embedchain/docs/get-started/full-stack.mdx
+++ /dev/null
@@ -1,81 +0,0 @@
----
-title: '💻 Full stack'
----
-
-Get started with full-stack RAG applications using Embedchain's easy-to-use CLI tool. Set up everything with just a few commands, whether you prefer Docker or not.
-
-## Prerequisites
-
-Choose your setup method:
-
-* [Without docker](#without-docker)
-* [With Docker](#with-docker)
-
-### Without Docker
-
-Ensure these are installed:
-
-- Embedchain python package (`pip install embedchain`)
-- [Node.js](https://docs.npmjs.com/downloading-and-installing-node-js-and-npm) and [Yarn](https://classic.yarnpkg.com/lang/en/docs/install/)
-
-### With Docker
-
-Install Docker from [Docker's official website](https://docs.docker.com/engine/install/).
-
-## Quick Start Guide
-
-### Install the package
-
-Before proceeding, make sure you have the Embedchain package installed.
-
-```bash
-pip install embedchain -U
-```
-
-### Setting Up
-
-For the purpose of the demo, you have to set `OPENAI_API_KEY` to start with but you can choose any llm by changing the configuration easily.
-
-### Installation Commands
-
-
-
-```bash without docker
-ec create-app my-app
-cd my-app
-ec start
-```
-
-```bash with docker
-ec create-app my-app --docker
-cd my-app
-ec start --docker
-```
-
-
-
-### What Happens Next?
-
-1. Embedchain fetches a full stack template (FastAPI backend, Next.JS frontend).
-2. Installs required components.
-3. Launches both frontend and backend servers.
-
-### See It In Action
-
-Open http://localhost:3000 to view the chat UI.
-
-
-
-### Admin Panel
-
-Check out the Embedchain admin panel to see the document chunks for your RAG application.
-
-
-
-### API Server
-
-If you want to access the API server, you can do so at http://localhost:8000/docs.
-
-
-
-You can customize the UI and code as per your requirements.
diff --git a/embedchain/docs/get-started/integrations.mdx b/embedchain/docs/get-started/integrations.mdx
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/docs/get-started/introduction.mdx b/embedchain/docs/get-started/introduction.mdx
deleted file mode 100644
index fc7ce22ed..000000000
--- a/embedchain/docs/get-started/introduction.mdx
+++ /dev/null
@@ -1,66 +0,0 @@
----
-title: 📚 Introduction
----
-
-## What is Embedchain?
-
-Embedchain is an Open Source Framework that makes it easy to create and deploy personalized AI apps. At its core, Embedchain follows the design principle of being *"Conventional but Configurable"* to serve both software engineers and machine learning engineers.
-
-Embedchain streamlines the creation of personalized LLM applications, offering a seamless process for managing various types of unstructured data. It efficiently segments data into manageable chunks, generates relevant embeddings, and stores them in a vector database for optimized retrieval. With a suite of diverse APIs, it enables users to extract contextual information, find precise answers, or engage in interactive chat conversations, all tailored to their own data.
-
-## Who is Embedchain for?
-
-Embedchain is designed for a diverse range of users, from AI professionals like Data Scientists and Machine Learning Engineers to those just starting their AI journey, including college students, independent developers, and hobbyists. Essentially, it's for anyone with an interest in AI, regardless of their expertise level.
-
-Our APIs are user-friendly yet adaptable, enabling beginners to effortlessly create LLM-powered applications with as few as 4 lines of code. At the same time, we offer extensive customization options for every aspect of building a personalized AI application. This includes the choice of LLMs, vector databases, loaders and chunkers, retrieval strategies, re-ranking, and more.
-
-Our platform's clear and well-structured abstraction layers ensure that users can tailor the system to meet their specific needs, whether they're crafting a simple project or a complex, nuanced AI application.
-
-## Why Use Embedchain?
-
-Developing a personalized AI application for production use presents numerous complexities, such as:
-
-- Integrating and indexing data from diverse sources.
-- Determining optimal data chunking methods for each source.
-- Synchronizing the RAG pipeline with regularly updated data sources.
-- Implementing efficient data storage in a vector store.
-- Deciding whether to include metadata with document chunks.
-- Handling permission management.
-- Configuring Large Language Models (LLMs).
-- Selecting effective prompts.
-- Choosing suitable retrieval strategies.
-- Assessing the performance of your RAG pipeline.
-- Deploying the pipeline into a production environment, among other concerns.
-
-Embedchain is designed to simplify these tasks, offering conventional yet customizable APIs. Our solution handles the intricate processes of loading, chunking, indexing, and retrieving data. This enables you to concentrate on aspects that are crucial for your specific use case or business objectives, ensuring a smoother and more focused development process.
-
-## How it works?
-
-Embedchain makes it easy to add data to your RAG pipeline with these straightforward steps:
-
-1. **Automatic Data Handling**: It automatically recognizes the data type and loads it.
-2. **Efficient Data Processing**: The system creates embeddings for key parts of your data.
-3. **Flexible Data Storage**: You get to choose where to store this processed data in a vector database.
-
-When a user asks a question, whether for chatting, searching, or querying, Embedchain simplifies the response process:
-
-1. **Query Processing**: It turns the user's question into embeddings.
-2. **Document Retrieval**: These embeddings are then used to find related documents in the database.
-3. **Answer Generation**: The related documents are used by the LLM to craft a precise answer.
-
-With Embedchain, you don’t have to worry about the complexities of building a personalized AI application. It offers an easy-to-use interface for developing applications with any kind of data.
-
-## Getting started
-
-Checkout our [quickstart guide](/get-started/quickstart) to start your first AI application.
-
-## Support
-
-Feel free to reach out to us if you have ideas, feedback or questions that we can help out with.
-
-
-
-## Contribute
-
-- [GitHub](https://github.com/embedchain/embedchain)
-- [Contribution docs](/contribution/dev)
diff --git a/embedchain/docs/get-started/quickstart.mdx b/embedchain/docs/get-started/quickstart.mdx
deleted file mode 100644
index 04d270191..000000000
--- a/embedchain/docs/get-started/quickstart.mdx
+++ /dev/null
@@ -1,89 +0,0 @@
----
-title: '⚡ Quickstart'
-description: '💡 Create an AI app on your own data in a minute'
----
-
-## Installation
-
-First install the Python package:
-
-```bash
-pip install embedchain
-```
-
-Once you have installed the package, depending upon your preference you can either use:
-
-
-
- This includes Open source LLMs like Mistral, Llama, etc.
- Free to use, and runs locally on your machine.
-
-
- This includes paid LLMs like GPT 4, Claude, etc.
- Cost money and are accessible via an API.
-
-
-
-## Open Source Models
-
-This section gives a quickstart example of using Mistral as the Open source LLM and Sentence transformers as the Open source embedding model. These models are free and run mostly on your local machine.
-
-We are using Mistral hosted at Hugging Face, so will you need a Hugging Face token to run this example. Its *free* and you can create one [here](https://huggingface.co/docs/hub/security-tokens).
-
-
-```python huggingface_demo.py
-import os
-# Replace this with your HF token
-os.environ["HUGGINGFACE_ACCESS_TOKEN"] = "hf_xxxx"
-
-from embedchain import App
-
-config = {
- 'llm': {
- 'provider': 'huggingface',
- 'config': {
- 'model': 'mistralai/Mistral-7B-Instruct-v0.2',
- 'top_p': 0.5
- }
- },
- 'embedder': {
- 'provider': 'huggingface',
- 'config': {
- 'model': 'sentence-transformers/all-mpnet-base-v2'
- }
- }
-}
-app = App.from_config(config=config)
-app.add("https://www.forbes.com/profile/elon-musk")
-app.add("https://en.wikipedia.org/wiki/Elon_Musk")
-app.query("What is the net worth of Elon Musk today?")
-# Answer: The net worth of Elon Musk today is $258.7 billion.
-```
-
-
-## Paid Models
-
-In this section, we will use both LLM and embedding model from OpenAI.
-
-```python openai_demo.py
-import os
-from embedchain import App
-
-# Replace this with your OpenAI key
-os.environ["OPENAI_API_KEY"] = "sk-xxxx"
-
-app = App()
-app.add("https://www.forbes.com/profile/elon-musk")
-app.add("https://en.wikipedia.org/wiki/Elon_Musk")
-app.query("What is the net worth of Elon Musk today?")
-# Answer: The net worth of Elon Musk today is $258.7 billion.
-```
-
-# Next Steps
-
-Now that you have created your first app, you can follow any of the links:
-
-* [Introduction](/get-started/introduction)
-* [Customization](/components/introduction)
-* [Use cases](/use-cases/introduction)
-* [Deployment](/get-started/deployment)
diff --git a/embedchain/docs/images/checks-passed.png b/embedchain/docs/images/checks-passed.png
deleted file mode 100644
index 3303c7736..000000000
Binary files a/embedchain/docs/images/checks-passed.png and /dev/null differ
diff --git a/embedchain/docs/images/cover.gif b/embedchain/docs/images/cover.gif
deleted file mode 100644
index efcc88243..000000000
Binary files a/embedchain/docs/images/cover.gif and /dev/null differ
diff --git a/embedchain/docs/images/fly_io.png b/embedchain/docs/images/fly_io.png
deleted file mode 100644
index 11a211afd..000000000
Binary files a/embedchain/docs/images/fly_io.png and /dev/null differ
diff --git a/embedchain/docs/images/fullstack-api-server.png b/embedchain/docs/images/fullstack-api-server.png
deleted file mode 100644
index 8b4ef2ac9..000000000
Binary files a/embedchain/docs/images/fullstack-api-server.png and /dev/null differ
diff --git a/embedchain/docs/images/fullstack-chunks.png b/embedchain/docs/images/fullstack-chunks.png
deleted file mode 100644
index ba4505ba7..000000000
Binary files a/embedchain/docs/images/fullstack-chunks.png and /dev/null differ
diff --git a/embedchain/docs/images/fullstack.png b/embedchain/docs/images/fullstack.png
deleted file mode 100644
index ba73bc067..000000000
Binary files a/embedchain/docs/images/fullstack.png and /dev/null differ
diff --git a/embedchain/docs/images/gradio_app.png b/embedchain/docs/images/gradio_app.png
deleted file mode 100644
index c5ed3cf4a..000000000
Binary files a/embedchain/docs/images/gradio_app.png and /dev/null differ
diff --git a/embedchain/docs/images/helicone-embedchain.png b/embedchain/docs/images/helicone-embedchain.png
deleted file mode 100644
index 05f61d73c..000000000
Binary files a/embedchain/docs/images/helicone-embedchain.png and /dev/null differ
diff --git a/embedchain/docs/images/langsmith.png b/embedchain/docs/images/langsmith.png
deleted file mode 100644
index 5d5ff5422..000000000
Binary files a/embedchain/docs/images/langsmith.png and /dev/null differ
diff --git a/embedchain/docs/images/og.png b/embedchain/docs/images/og.png
deleted file mode 100644
index 7a89999d3..000000000
Binary files a/embedchain/docs/images/og.png and /dev/null differ
diff --git a/embedchain/docs/images/slack-ai.png b/embedchain/docs/images/slack-ai.png
deleted file mode 100644
index cb2f137de..000000000
Binary files a/embedchain/docs/images/slack-ai.png and /dev/null differ
diff --git a/embedchain/docs/images/whatsapp.jpg b/embedchain/docs/images/whatsapp.jpg
deleted file mode 100644
index 6f28ba200..000000000
Binary files a/embedchain/docs/images/whatsapp.jpg and /dev/null differ
diff --git a/embedchain/docs/integration/chainlit.mdx b/embedchain/docs/integration/chainlit.mdx
deleted file mode 100644
index 6a28f309a..000000000
--- a/embedchain/docs/integration/chainlit.mdx
+++ /dev/null
@@ -1,68 +0,0 @@
----
-title: '⛓️ Chainlit'
-description: 'Integrate with Chainlit to create LLM chat apps'
----
-
-In this example, we will learn how to use Chainlit and Embedchain together.
-
-
-
-## Setup
-
-First, install the required packages:
-
-```bash
-pip install embedchain chainlit
-```
-
-## Create a Chainlit app
-
-Create a new file called `app.py` and add the following code:
-
-```python
-import chainlit as cl
-from embedchain import App
-
-import os
-
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-@cl.on_chat_start
-async def on_chat_start():
- app = App.from_config(config={
- 'app': {
- 'config': {
- 'name': 'chainlit-app'
- }
- },
- 'llm': {
- 'config': {
- 'stream': True,
- }
- }
- })
- # import your data here
- app.add("https://www.forbes.com/profile/elon-musk/")
- app.collect_metrics = False
- cl.user_session.set("app", app)
-
-
-@cl.on_message
-async def on_message(message: cl.Message):
- app = cl.user_session.get("app")
- msg = cl.Message(content="")
- for chunk in await cl.make_async(app.chat)(message.content):
- await msg.stream_token(chunk)
-
- await msg.send()
-```
-
-## Run the app
-
-```
-chainlit run app.py
-```
-
-## Try it out
-
-Open the app in your browser and start chatting with it!
diff --git a/embedchain/docs/integration/helicone.mdx b/embedchain/docs/integration/helicone.mdx
deleted file mode 100644
index a5a33445a..000000000
--- a/embedchain/docs/integration/helicone.mdx
+++ /dev/null
@@ -1,52 +0,0 @@
----
-title: "🧊 Helicone"
-description: "Implement Helicone, the open-source LLM observability platform, with Embedchain. Monitor, debug, and optimize your AI applications effortlessly."
-"twitter:title": "Helicone LLM Observability for Embedchain"
----
-
-Get started with [Helicone](https://www.helicone.ai/), the open-source LLM observability platform for developers to monitor, debug, and optimize their applications.
-
-To use Helicone, you need to do the following steps.
-
-## Integration Steps
-
-
-
- Log into [Helicone](https://www.helicone.ai) or create an account. Once you have an account, you
- can generate an [API key](https://helicone.ai/developer).
-
-
- Make sure to generate a [write only API key](helicone-headers/helicone-auth).
-
-
-
-
-You can configure your base_url and OpenAI API key in your codebase
-
-
-```python main.py
-import os
-from embedchain import App
-
-# Modify the base path and add a Helicone URL
-os.environ["OPENAI_API_BASE"] = "https://oai.helicone.ai/{YOUR_HELICONE_API_KEY}/v1"
-# Add your OpenAI API Key
-os.environ["OPENAI_API_KEY"] = "{YOUR_OPENAI_API_KEY}"
-
-app = App()
-
-# Add data to your app
-app.add("https://en.wikipedia.org/wiki/Elon_Musk")
-
-# Query your app
-print(app.query("How many companies did Elon found? Which companies?"))
-```
-
-
-
-
-
-
-
-
-Check out [Helicone](https://www.helicone.ai) to see more use cases!
diff --git a/embedchain/docs/integration/langsmith.mdx b/embedchain/docs/integration/langsmith.mdx
deleted file mode 100644
index 8a200be5b..000000000
--- a/embedchain/docs/integration/langsmith.mdx
+++ /dev/null
@@ -1,71 +0,0 @@
----
-title: '🛠️ LangSmith'
-description: 'Integrate with Langsmith to debug and monitor your LLM app'
----
-
-Embedchain now supports integration with [LangSmith](https://www.langchain.com/langsmith).
-
-To use LangSmith, you need to do the following steps.
-
-1. Have an account on LangSmith and keep the environment variables in handy
-2. Set the environment variables in your app so that embedchain has context about it.
-3. Just use embedchain and everything will be logged to LangSmith, so that you can better test and monitor your application.
-
-Let's cover each step in detail.
-
-
-* First make sure that you have created a LangSmith account and have all the necessary variables handy. LangSmith has a [good documentation](https://docs.smith.langchain.com/) on how to get started with their service.
-
-* Once you have setup the account, we will need the following environment variables
-
-```bash
-# Setting environment variable for LangChain Tracing V2 integration.
-export LANGCHAIN_TRACING_V2=true
-
-# Setting the API endpoint for LangChain.
-export LANGCHAIN_ENDPOINT=https://api.smith.langchain.com
-
-# Replace '' with your LangChain API key.
-export LANGCHAIN_API_KEY=
-
-# Replace '' with your LangChain project name, or it defaults to "default".
-export LANGCHAIN_PROJECT= # if not specified, defaults to "default"
-```
-
-If you are using Python, you can use the following code to set environment variables
-
-```python
-import os
-
-# Setting environment variable for LangChain Tracing V2 integration.
-os.environ['LANGCHAIN_TRACING_V2'] = 'true'
-
-# Setting the API endpoint for LangChain.
-os.environ['LANGCHAIN_ENDPOINT'] = 'https://api.smith.langchain.com'
-
-# Replace '' with your LangChain API key.
-os.environ['LANGCHAIN_API_KEY'] = ''
-
-# Replace '' with your LangChain project name.
-os.environ['LANGCHAIN_PROJECT'] = ''
-```
-
-* Now create an app using Embedchain and everything will be automatically visible in the LangSmith
-
-
-```python
-from embedchain import App
-
-# Initialize EmbedChain application.
-app = App()
-
-# Add data to your app
-app.add("https://en.wikipedia.org/wiki/Elon_Musk")
-
-# Query your app
-app.query("How many companies did Elon found?")
-```
-
-* Now the entire log for this will be visible in langsmith.
-
-
diff --git a/embedchain/docs/integration/openlit.mdx b/embedchain/docs/integration/openlit.mdx
deleted file mode 100644
index 22036919e..000000000
--- a/embedchain/docs/integration/openlit.mdx
+++ /dev/null
@@ -1,50 +0,0 @@
----
-title: '🔭 OpenLIT'
-description: 'OpenTelemetry-native Observability and Evals for LLMs & GPUs'
----
-
-Embedchain now supports integration with [OpenLIT](https://github.com/openlit/openlit).
-
-## Getting Started
-
-### 1. Set environment variables
-```bash
-# Setting environment variable for OpenTelemetry destination and authetication.
-export OTEL_EXPORTER_OTLP_ENDPOINT = "YOUR_OTEL_ENDPOINT"
-export OTEL_EXPORTER_OTLP_HEADERS = "YOUR_OTEL_ENDPOINT_AUTH"
-```
-
-### 2. Install the OpenLIT SDK
-Open your terminal and run:
-
-```shell
-pip install openlit
-```
-
-### 3. Setup Your Application for Monitoring
-Now create an app using Embedchain and initialize OpenTelemetry monitoring
-
-```python
-from embedchain import App
-import OpenLIT
-
-# Initialize OpenLIT Auto Instrumentation for monitoring.
-openlit.init()
-
-# Initialize EmbedChain application.
-app = App()
-
-# Add data to your app
-app.add("https://en.wikipedia.org/wiki/Elon_Musk")
-
-# Query your app
-app.query("How many companies did Elon found?")
-```
-
-### 4. Visualize
-
-Once you've set up data collection with OpenLIT, you can visualize and analyze this information to better understand your application's performance:
-
-- **Using OpenLIT UI:** Connect to OpenLIT's UI to start exploring performance metrics. Visit the OpenLIT [Quickstart Guide](https://docs.openlit.io/latest/quickstart) for step-by-step details.
-
-- **Integrate with existing Observability Tools:** If you use tools like Grafana or DataDog, you can integrate the data collected by OpenLIT. For instructions on setting up these connections, check the OpenLIT [Connections Guide](https://docs.openlit.io/latest/connections/intro).
diff --git a/embedchain/docs/integration/streamlit-mistral.mdx b/embedchain/docs/integration/streamlit-mistral.mdx
deleted file mode 100644
index d7b795755..000000000
--- a/embedchain/docs/integration/streamlit-mistral.mdx
+++ /dev/null
@@ -1,112 +0,0 @@
----
-title: '🚀 Streamlit'
-description: 'Integrate with Streamlit to plug and play with any LLM'
----
-
-In this example, we will learn how to use `mistralai/Mixtral-8x7B-Instruct-v0.1` and Embedchain together with Streamlit to build a simple RAG chatbot.
-
-
-
-## Setup
-
-Install Embedchain and Streamlit.
-```bash
-pip install embedchain streamlit
-```
-
-
- ```python
- import os
- from embedchain import App
- import streamlit as st
-
- with st.sidebar:
- huggingface_access_token = st.text_input("Hugging face Token", key="chatbot_api_key", type="password")
- "[Get Hugging Face Access Token](https://huggingface.co/settings/tokens)"
- "[View the source code](https://github.com/embedchain/examples/mistral-streamlit)"
-
-
- st.title("💬 Chatbot")
- st.caption("🚀 An Embedchain app powered by Mistral!")
- if "messages" not in st.session_state:
- st.session_state.messages = [
- {
- "role": "assistant",
- "content": """
- Hi! I'm a chatbot. I can answer questions and learn new things!\n
- Ask me anything and if you want me to learn something do `/add `.\n
- I can learn mostly everything. :)
- """,
- }
- ]
-
- for message in st.session_state.messages:
- with st.chat_message(message["role"]):
- st.markdown(message["content"])
-
- if prompt := st.chat_input("Ask me anything!"):
- if not st.session_state.chatbot_api_key:
- st.error("Please enter your Hugging Face Access Token")
- st.stop()
-
- os.environ["HUGGINGFACE_ACCESS_TOKEN"] = st.session_state.chatbot_api_key
- app = App.from_config(config_path="config.yaml")
-
- if prompt.startswith("/add"):
- with st.chat_message("user"):
- st.markdown(prompt)
- st.session_state.messages.append({"role": "user", "content": prompt})
- prompt = prompt.replace("/add", "").strip()
- with st.chat_message("assistant"):
- message_placeholder = st.empty()
- message_placeholder.markdown("Adding to knowledge base...")
- app.add(prompt)
- message_placeholder.markdown(f"Added {prompt} to knowledge base!")
- st.session_state.messages.append({"role": "assistant", "content": f"Added {prompt} to knowledge base!"})
- st.stop()
-
- with st.chat_message("user"):
- st.markdown(prompt)
- st.session_state.messages.append({"role": "user", "content": prompt})
-
- with st.chat_message("assistant"):
- msg_placeholder = st.empty()
- msg_placeholder.markdown("Thinking...")
- full_response = ""
-
- for response in app.chat(prompt):
- msg_placeholder.empty()
- full_response += response
-
- msg_placeholder.markdown(full_response)
- st.session_state.messages.append({"role": "assistant", "content": full_response})
- ```
-
-
- ```yaml
- app:
- config:
- name: 'mistral-streamlit-app'
-
- llm:
- provider: huggingface
- config:
- model: 'mistralai/Mixtral-8x7B-Instruct-v0.1'
- temperature: 0.1
- max_tokens: 250
- top_p: 0.1
- stream: true
-
- embedder:
- provider: huggingface
- config:
- model: 'sentence-transformers/all-mpnet-base-v2'
- ```
-
-
-
-## To run it locally,
-
-```bash
-streamlit run app.py
-```
diff --git a/embedchain/docs/logo/dark-rt.svg b/embedchain/docs/logo/dark-rt.svg
deleted file mode 100644
index 83eb7fc69..000000000
--- a/embedchain/docs/logo/dark-rt.svg
+++ /dev/null
@@ -1,10 +0,0 @@
-
diff --git a/embedchain/docs/logo/dark.svg b/embedchain/docs/logo/dark.svg
deleted file mode 100644
index cbd502094..000000000
--- a/embedchain/docs/logo/dark.svg
+++ /dev/null
@@ -1,11 +0,0 @@
-
diff --git a/embedchain/docs/logo/light-rt.svg b/embedchain/docs/logo/light-rt.svg
deleted file mode 100644
index f204d17e6..000000000
--- a/embedchain/docs/logo/light-rt.svg
+++ /dev/null
@@ -1,10 +0,0 @@
-
diff --git a/embedchain/docs/logo/light.svg b/embedchain/docs/logo/light.svg
deleted file mode 100644
index cbd502094..000000000
--- a/embedchain/docs/logo/light.svg
+++ /dev/null
@@ -1,11 +0,0 @@
-
diff --git a/embedchain/docs/mint.json b/embedchain/docs/mint.json
deleted file mode 100644
index a3d4aec8d..000000000
--- a/embedchain/docs/mint.json
+++ /dev/null
@@ -1,277 +0,0 @@
-{
- "$schema": "https://mintlify.com/schema.json",
- "name": "Embedchain",
- "logo": {
- "dark": "/logo/dark-rt.svg",
- "light": "/logo/light-rt.svg",
- "href": "https://github.com/embedchain/embedchain"
- },
- "favicon": "/favicon.png",
- "colors": {
- "primary": "#3B2FC9",
- "light": "#6673FF",
- "dark": "#3B2FC9",
- "background": {
- "dark": "#0f1117",
- "light": "#fff"
- }
- },
- "modeToggle": {
- "default": "dark"
- },
- "openapi": ["/rest-api.json"],
- "metadata": {
- "og:image": "/images/og.png",
- "twitter:site": "@embedchain"
- },
- "tabs": [
- {
- "name": "Examples",
- "url": "examples"
- },
- {
- "name": "API Reference",
- "url": "api-reference"
- }
- ],
- "anchors": [
- {
- "name": "Talk to founders",
- "icon": "calendar",
- "url": "https://cal.com/taranjeetio/ec"
- }
- ],
- "topbarLinks": [
- {
- "name": "GitHub",
- "url": "https://github.com/embedchain/embedchain"
- }
- ],
- "topbarCtaButton": {
- "name": "Join our slack",
- "url": "https://embedchain.ai/slack"
- },
- "primaryTab": {
- "name": "📘 Documentation"
- },
- "navigation": [
- {
- "group": "Get Started",
- "pages": [
- "get-started/quickstart",
- "get-started/introduction",
- "get-started/faq",
- "get-started/full-stack",
- {
- "group": "🔗 Integrations",
- "pages": [
- "integration/langsmith",
- "integration/chainlit",
- "integration/streamlit-mistral",
- "integration/openlit",
- "integration/helicone"
- ]
- }
- ]
- },
- {
- "group": "Use cases",
- "pages": [
- "use-cases/introduction",
- "use-cases/chatbots",
- "use-cases/question-answering",
- "use-cases/semantic-search"
- ]
- },
- {
- "group": "Components",
- "pages": [
- "components/introduction",
- {
- "group": "🗂️ Data sources",
- "pages": [
- "components/data-sources/overview",
- {
- "group": "Data types",
- "pages": [
- "components/data-sources/pdf-file",
- "components/data-sources/csv",
- "components/data-sources/json",
- "components/data-sources/text",
- "components/data-sources/directory",
- "components/data-sources/web-page",
- "components/data-sources/youtube-channel",
- "components/data-sources/youtube-video",
- "components/data-sources/docs-site",
- "components/data-sources/mdx",
- "components/data-sources/docx",
- "components/data-sources/notion",
- "components/data-sources/sitemap",
- "components/data-sources/xml",
- "components/data-sources/qna",
- "components/data-sources/openapi",
- "components/data-sources/gmail",
- "components/data-sources/github",
- "components/data-sources/postgres",
- "components/data-sources/mysql",
- "components/data-sources/slack",
- "components/data-sources/discord",
- "components/data-sources/discourse",
- "components/data-sources/substack",
- "components/data-sources/beehiiv",
- "components/data-sources/directory",
- "components/data-sources/dropbox",
- "components/data-sources/image",
- "components/data-sources/audio",
- "components/data-sources/custom"
- ]
- },
- "components/data-sources/data-type-handling"
- ]
- },
- {
- "group": "🗄️ Vector databases",
- "pages": [
- "components/vector-databases/chromadb",
- "components/vector-databases/elasticsearch",
- "components/vector-databases/pinecone",
- "components/vector-databases/opensearch",
- "components/vector-databases/qdrant",
- "components/vector-databases/weaviate",
- "components/vector-databases/zilliz"
- ]
- },
- "components/llms",
- "components/embedding-models",
- "components/evaluation"
- ]
- },
- {
- "group": "Deployment",
- "pages": [
- "get-started/deployment",
- "deployment/fly_io",
- "deployment/modal_com",
- "deployment/render_com",
- "deployment/railway",
- "deployment/streamlit_io",
- "deployment/gradio_app",
- "deployment/huggingface_spaces"
- ]
- },
- {
- "group": "Community",
- "pages": ["community/connect-with-us"]
- },
- {
- "group": "Examples",
- "pages": [
- "examples/chat-with-PDF",
- "examples/notebooks-and-replits",
- {
- "group": "REST API Service",
- "pages": [
- "examples/rest-api/getting-started",
- "examples/rest-api/create",
- "examples/rest-api/get-all-apps",
- "examples/rest-api/add-data",
- "examples/rest-api/get-data",
- "examples/rest-api/query",
- "examples/rest-api/deploy",
- "examples/rest-api/delete",
- "examples/rest-api/check-status"
- ]
- },
- "examples/openai-assistant",
- "examples/opensource-assistant",
- "examples/nextjs-assistant",
- "examples/slack-AI"
- ]
- },
- {
- "group": "Chatbots",
- "pages": [
- "examples/discord_bot",
- "examples/slack_bot",
- "examples/telegram_bot",
- "examples/whatsapp_bot",
- "examples/poe_bot"
- ]
- },
- {
- "group": "Showcase",
- "pages": ["examples/showcase"]
- },
- {
- "group": "API Reference",
- "pages": [
- "api-reference/app/overview",
- {
- "group": "App methods",
- "pages": [
- "api-reference/app/add",
- "api-reference/app/query",
- "api-reference/app/chat",
- "api-reference/app/search",
- "api-reference/app/get",
- "api-reference/app/evaluate",
- "api-reference/app/deploy",
- "api-reference/app/reset",
- "api-reference/app/delete"
- ]
- },
- "api-reference/store/openai-assistant",
- "api-reference/store/ai-assistants",
- "api-reference/advanced/configuration"
- ]
- },
- {
- "group": "Contributing",
- "pages": [
- "contribution/guidelines",
- "contribution/dev",
- "contribution/docs",
- "contribution/python"
- ]
- },
- {
- "group": "Product",
- "pages": ["product/release-notes"]
- }
- ],
- "footerSocials": {
- "website": "https://embedchain.ai",
- "github": "https://github.com/embedchain/embedchain",
- "slack": "https://embedchain.ai/slack",
- "discord": "https://discord.gg/6PzXDgEjG5",
- "twitter": "https://twitter.com/embedchain",
- "linkedin": "https://www.linkedin.com/company/embedchain"
- },
- "isWhiteLabeled": true,
- "analytics": {
- "posthog": {
- "apiKey": "phc_PHQDA5KwztijnSojsxJ2c1DuJd52QCzJzT2xnSGvjN2",
- "apiHost": "https://app.embedchain.ai/ingest"
- },
- "ga4": {
- "measurementId": "G-4QK7FJE6T3"
- }
- },
- "feedback": {
- "suggestEdit": true,
- "raiseIssue": true,
- "thumbsRating": true
- },
- "search": {
- "prompt": "✨ Search embedchain docs..."
- },
- "api": {
- "baseUrl": "http://localhost:8080"
- },
- "redirects": [
- {
- "source": "/changelog/command-line",
- "destination": "/get-started/introduction"
- }
- ]
-}
diff --git a/embedchain/docs/product/release-notes.mdx b/embedchain/docs/product/release-notes.mdx
deleted file mode 100644
index 02bcf977b..000000000
--- a/embedchain/docs/product/release-notes.mdx
+++ /dev/null
@@ -1,4 +0,0 @@
----
-title: ' 📜 Release Notes'
-url: https://github.com/embedchain/embedchain/releases
----
\ No newline at end of file
diff --git a/embedchain/docs/rest-api.json b/embedchain/docs/rest-api.json
deleted file mode 100644
index 087d7e06c..000000000
--- a/embedchain/docs/rest-api.json
+++ /dev/null
@@ -1,427 +0,0 @@
-{
- "openapi": "3.1.0",
- "info": {
- "title": "Embedchain REST API",
- "description": "This is the REST API for Embedchain.",
- "license": {
- "name": "Apache 2.0",
- "url": "https://github.com/embedchain/embedchain/blob/main/LICENSE"
- },
- "version": "0.0.1"
- },
- "paths": {
- "/ping": {
- "get": {
- "tags": ["Utility"],
- "summary": "Check status",
- "description": "Endpoint to check the status of the API",
- "operationId": "check_status_ping_get",
- "responses": {
- "200": {
- "description": "Successful Response",
- "content": { "application/json": { "schema": {} } }
- }
- }
- }
- },
- "/apps": {
- "get": {
- "tags": ["Apps"],
- "summary": "Get all apps",
- "description": "Get all applications",
- "operationId": "get_all_apps_apps_get",
- "responses": {
- "200": {
- "description": "Successful Response",
- "content": { "application/json": { "schema": {} } }
- }
- }
- }
- },
- "/create": {
- "post": {
- "tags": ["Apps"],
- "summary": "Create app",
- "description": "Create a new app using App ID",
- "operationId": "create_app_using_default_config_create_post",
- "parameters": [
- {
- "name": "app_id",
- "in": "query",
- "required": true,
- "schema": { "type": "string", "title": "App Id" }
- }
- ],
- "requestBody": {
- "content": {
- "multipart/form-data": {
- "schema": {
- "allOf": [
- {
- "$ref": "#/components/schemas/Body_create_app_using_default_config_create_post"
- }
- ],
- "title": "Body"
- }
- }
- }
- },
- "responses": {
- "200": {
- "description": "Successful Response",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/DefaultResponse" }
- }
- }
- },
- "422": {
- "description": "Validation Error",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/HTTPValidationError" }
- }
- }
- }
- }
- }
- },
- "/{app_id}/data": {
- "get": {
- "tags": ["Apps"],
- "summary": "Get data",
- "description": "Get all data sources for an app",
- "operationId": "get_datasources_associated_with_app_id__app_id__data_get",
- "parameters": [
- {
- "name": "app_id",
- "in": "path",
- "required": true,
- "schema": { "type": "string", "title": "App Id" }
- }
- ],
- "responses": {
- "200": {
- "description": "Successful Response",
- "content": { "application/json": { "schema": {} } }
- },
- "422": {
- "description": "Validation Error",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/HTTPValidationError" }
- }
- }
- }
- }
- }
- },
- "/{app_id}/add": {
- "post": {
- "tags": ["Apps"],
- "summary": "Add data",
- "description": "Add a data source to an app.",
- "operationId": "add_datasource_to_an_app__app_id__add_post",
- "parameters": [
- {
- "name": "app_id",
- "in": "path",
- "required": true,
- "schema": { "type": "string", "title": "App Id" }
- }
- ],
- "requestBody": {
- "required": true,
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/SourceApp" }
- }
- }
- },
- "responses": {
- "200": {
- "description": "Successful Response",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/DefaultResponse" }
- }
- }
- },
- "422": {
- "description": "Validation Error",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/HTTPValidationError" }
- }
- }
- }
- }
- }
- },
- "/{app_id}/query": {
- "post": {
- "tags": ["Apps"],
- "summary": "Query app",
- "description": "Query an app",
- "operationId": "query_an_app__app_id__query_post",
- "parameters": [
- {
- "name": "app_id",
- "in": "path",
- "required": true,
- "schema": { "type": "string", "title": "App Id" }
- }
- ],
- "requestBody": {
- "required": true,
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/QueryApp" }
- }
- }
- },
- "responses": {
- "200": {
- "description": "Successful Response",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/DefaultResponse" }
- }
- }
- },
- "422": {
- "description": "Validation Error",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/HTTPValidationError" }
- }
- }
- }
- }
- }
- },
- "/{app_id}/chat": {
- "post": {
- "tags": ["Apps"],
- "summary": "Chat",
- "description": "Chat with an app.\n\napp_id: The ID of the app. Use \"default\" for the default app.\n\nmessage: The message that you want to send to the app.",
- "operationId": "chat_with_an_app__app_id__chat_post",
- "parameters": [
- {
- "name": "app_id",
- "in": "path",
- "required": true,
- "schema": { "type": "string", "title": "App Id" }
- }
- ],
- "requestBody": {
- "required": true,
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/MessageApp" }
- }
- }
- },
- "responses": {
- "200": {
- "description": "Successful Response",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/DefaultResponse" }
- }
- }
- },
- "422": {
- "description": "Validation Error",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/HTTPValidationError" }
- }
- }
- }
- }
- }
- },
- "/{app_id}/deploy": {
- "post": {
- "tags": ["Apps"],
- "summary": "Deploy app",
- "description": "Deploy an existing app.",
- "operationId": "deploy_app__app_id__deploy_post",
- "parameters": [
- {
- "name": "app_id",
- "in": "path",
- "required": true,
- "schema": { "type": "string", "title": "App Id" }
- }
- ],
- "requestBody": {
- "required": true,
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/DeployAppRequest" }
- }
- }
- },
- "responses": {
- "200": {
- "description": "Successful Response",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/DefaultResponse" }
- }
- }
- },
- "422": {
- "description": "Validation Error",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/HTTPValidationError" }
- }
- }
- }
- }
- }
- },
- "/{app_id}/delete": {
- "delete": {
- "tags": ["Apps"],
- "summary": "Delete app",
- "description": "Delete an existing app",
- "operationId": "delete_app__app_id__delete_delete",
- "parameters": [
- {
- "name": "app_id",
- "in": "path",
- "required": true,
- "schema": { "type": "string", "title": "App Id" }
- }
- ],
- "responses": {
- "200": {
- "description": "Successful Response",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/DefaultResponse" }
- }
- }
- },
- "422": {
- "description": "Validation Error",
- "content": {
- "application/json": {
- "schema": { "$ref": "#/components/schemas/HTTPValidationError" }
- }
- }
- }
- }
- }
- }
- },
- "components": {
- "schemas": {
- "Body_create_app_using_default_config_create_post": {
- "properties": {
- "config": { "type": "string", "format": "binary", "title": "Config" }
- },
- "type": "object",
- "title": "Body_create_app_using_default_config_create_post"
- },
- "DefaultResponse": {
- "properties": { "response": { "type": "string", "title": "Response" } },
- "type": "object",
- "required": ["response"],
- "title": "DefaultResponse"
- },
- "DeployAppRequest": {
- "properties": {
- "api_key": {
- "type": "string",
- "title": "Api Key",
- "description": "The Embedchain API key for app deployments. You get the api key on the Embedchain platform by visiting [https://app.embedchain.ai](https://app.embedchain.ai)",
- "default": ""
- }
- },
- "type": "object",
- "title": "DeployAppRequest",
- "example":{
- "api_key":"ec-xxx"
- }
- },
- "HTTPValidationError": {
- "properties": {
- "detail": {
- "items": { "$ref": "#/components/schemas/ValidationError" },
- "type": "array",
- "title": "Detail"
- }
- },
- "type": "object",
- "title": "HTTPValidationError"
- },
- "MessageApp": {
- "properties": {
- "message": {
- "type": "string",
- "title": "Message",
- "description": "The message that you want to send to the App.",
- "default": ""
- }
- },
- "type": "object",
- "title": "MessageApp"
- },
- "QueryApp": {
- "properties": {
- "query": {
- "type": "string",
- "title": "Query",
- "description": "The query that you want to ask the App.",
- "default": ""
- }
- },
- "type": "object",
- "title": "QueryApp",
- "example":{
- "query":"Who is Elon Musk?"
- }
- },
- "SourceApp": {
- "properties": {
- "source": {
- "type": "string",
- "title": "Source",
- "description": "The source that you want to add to the App.",
- "default": ""
- },
- "data_type": {
- "anyOf": [{ "type": "string" }, { "type": "null" }],
- "title": "Data Type",
- "description": "The type of data to add, remove it if you want Embedchain to detect it automatically.",
- "default": ""
- }
- },
- "type": "object",
- "title": "SourceApp",
- "example":{
- "source":"https://en.wikipedia.org/wiki/Elon_Musk"
- }
- },
- "ValidationError": {
- "properties": {
- "loc": {
- "items": { "anyOf": [{ "type": "string" }, { "type": "integer" }] },
- "type": "array",
- "title": "Location"
- },
- "msg": { "type": "string", "title": "Message" },
- "type": { "type": "string", "title": "Error Type" }
- },
- "type": "object",
- "required": ["loc", "msg", "type"],
- "title": "ValidationError"
- }
- }
- }
- }
diff --git a/embedchain/docs/support/get-help.mdx b/embedchain/docs/support/get-help.mdx
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/docs/use-cases/chatbots.mdx b/embedchain/docs/use-cases/chatbots.mdx
deleted file mode 100644
index d11e45257..000000000
--- a/embedchain/docs/use-cases/chatbots.mdx
+++ /dev/null
@@ -1,38 +0,0 @@
----
-title: '🤖 Chatbots'
----
-
-Chatbots, especially those powered by Large Language Models (LLMs), have a wide range of use cases, significantly enhancing various aspects of business, education, and personal assistance. Here are some key applications:
-
-- **Customer Service**: Automating responses to common queries and providing 24/7 support.
-- **Education**: Offering personalized tutoring and learning assistance.
-- **E-commerce**: Assisting in product discovery, recommendations, and transactions.
-- **Content Management**: Aiding in writing, summarizing, and organizing content.
-- **Data Analysis**: Extracting insights from large datasets.
-- **Language Translation**: Providing real-time multilingual support.
-- **Mental Health**: Offering preliminary mental health support and conversation.
-- **Entertainment**: Engaging users with games, quizzes, and humorous chats.
-- **Accessibility Aid**: Enhancing information and service access for individuals with disabilities.
-
-Embedchain provides the right set of tools to create chatbots for the above use cases. Refer to the following examples of chatbots on and you can built on top of these examples:
-
-
-
- Build a tailored GPT chatbot suited for your specific needs.
-
-
- Enhance your Slack workspace with a specialized bot.
-
-
- Create an engaging bot for your Discord server.
-
-
- Develop a handy assistant for Telegram users.
-
-
- Design a WhatsApp bot for efficient communication.
-
-
- Explore advanced bot interactions with Poe Bot.
-
-
diff --git a/embedchain/docs/use-cases/introduction.mdx b/embedchain/docs/use-cases/introduction.mdx
deleted file mode 100644
index e908ba64d..000000000
--- a/embedchain/docs/use-cases/introduction.mdx
+++ /dev/null
@@ -1,11 +0,0 @@
----
-title: 🧱 Introduction
----
-
-## Overview
-
-You can use embedchain to create the following usecases:
-
-* [Chatbots](/use-cases/chatbots)
-* [Question Answering](/use-cases/question-answering)
-* [Semantic Search](/use-cases/semantic-search)
\ No newline at end of file
diff --git a/embedchain/docs/use-cases/question-answering.mdx b/embedchain/docs/use-cases/question-answering.mdx
deleted file mode 100644
index f538419b5..000000000
--- a/embedchain/docs/use-cases/question-answering.mdx
+++ /dev/null
@@ -1,75 +0,0 @@
----
-title: '❓ Question Answering'
----
-
-Utilizing large language models (LLMs) for question answering is a transformative application, bringing significant benefits to various real-world situations. Embedchain extensively supports tasks related to question answering, including summarization, content creation, language translation, and data analysis. The versatility of question answering with LLMs enables solutions for numerous practical applications such as:
-
-- **Educational Aid**: Enhancing learning experiences and aiding with homework
-- **Customer Support**: Addressing and resolving customer queries efficiently
-- **Research Assistance**: Facilitating academic and professional research endeavors
-- **Healthcare Information**: Providing fundamental medical knowledge
-- **Technical Support**: Resolving technology-related inquiries
-- **Legal Information**: Offering basic legal advice and information
-- **Business Insights**: Delivering market analysis and strategic business advice
-- **Language Learning** Assistance: Aiding in understanding and translating languages
-- **Travel Guidance**: Supplying information on travel and hospitality
-- **Content Development**: Assisting authors and creators with research and idea generation
-
-## Example: Build a Q&A System with Embedchain for Next.JS
-
-Quickly create a RAG pipeline to answer queries about the [Next.JS Framework](https://nextjs.org/) using Embedchain tools.
-
-### Step 1: Set Up Your RAG Pipeline
-
-First, let's create your RAG pipeline. Open your Python environment and enter:
-
-```python Create pipeline
-from embedchain import App
-app = App()
-```
-
-This initializes your application.
-
-### Step 2: Populate Your Pipeline with Data
-
-Now, let's add data to your pipeline. We'll include the Next.JS website and its documentation:
-
-```python Ingest data sources
-# Add Next.JS Website and docs
-app.add("https://nextjs.org/sitemap.xml", data_type="sitemap")
-
-# Add Next.JS Forum data
-app.add("https://nextjs-forum.com/sitemap.xml", data_type="sitemap")
-```
-
-This step incorporates over **15K pages** from the Next.JS website and forum into your pipeline. For more data source options, check the [Embedchain data sources overview](/components/data-sources/overview).
-
-### Step 3: Local Testing of Your Pipeline
-
-Test the pipeline on your local machine:
-
-```python Query App
-app.query("Summarize the features of Next.js 14?")
-```
-
-Run this query to see how your pipeline responds with information about Next.js 14.
-
-### (Optional) Step 4: Deploying Your RAG Pipeline
-
-Want to go live? Deploy your pipeline with these options:
-
-- Deploy on the Embedchain Platform
-- Self-host on your preferred cloud provider
-
-For detailed deployment instructions, follow these guides:
-
-- [Deploying on Embedchain Platform](/get-started/deployment#deploy-on-embedchain-platform)
-- [Self-hosting Guide](/get-started/deployment#self-hosting)
-
-## Need help?
-
-If you are looking to configure the RAG pipeline further, feel free to checkout the [API reference](/api-reference/pipeline/query).
-
-In case you run into issues, feel free to contact us via any of the following methods:
-
-
diff --git a/embedchain/docs/use-cases/semantic-search.mdx b/embedchain/docs/use-cases/semantic-search.mdx
deleted file mode 100644
index f506e5dd1..000000000
--- a/embedchain/docs/use-cases/semantic-search.mdx
+++ /dev/null
@@ -1,101 +0,0 @@
----
-title: '🔍 Semantic Search'
----
-
-Semantic searching, which involves understanding the intent and contextual meaning behind search queries, is yet another popular use-case of RAG. It has several popular use cases across various domains:
-
-- **Information Retrieval**: Enhances search accuracy in databases and websites
-- **E-commerce**: Improves product discovery in online shopping
-- **Customer Support**: Powers smarter chatbots for effective responses
-- **Content Discovery**: Aids in finding relevant media content
-- **Knowledge Management**: Streamlines document and data retrieval in enterprises
-- **Healthcare**: Facilitates medical research and literature search
-- **Legal Research**: Assists in legal document and case law search
-- **Academic Research**: Aids in academic paper discovery
-- **Language Processing**: Enables multilingual search capabilities
-
-Embedchain offers a simple yet customizable `search()` API that you can use for semantic search. See the example in the next section to know more.
-
-## Example: Semantic Search over Next.JS Website + Forum
-
-### Step 1: Set Up Your RAG Pipeline
-
-First, let's create your RAG pipeline. Open your Python environment and enter:
-
-```python Create pipeline
-from embedchain import App
-app = App()
-```
-
-This initializes your application.
-
-### Step 2: Populate Your Pipeline with Data
-
-Now, let's add data to your pipeline. We'll include the Next.JS website and its documentation:
-
-```python Ingest data sources
-# Add Next.JS Website and docs
-app.add("https://nextjs.org/sitemap.xml", data_type="sitemap")
-
-# Add Next.JS Forum data
-app.add("https://nextjs-forum.com/sitemap.xml", data_type="sitemap")
-```
-
-This step incorporates over **15K pages** from the Next.JS website and forum into your pipeline. For more data source options, check the [Embedchain data sources overview](/components/data-sources/overview).
-
-### Step 3: Local Testing of Your Pipeline
-
-Test the pipeline on your local machine:
-
-```python Search App
-app.search("Summarize the features of Next.js 14?")
-[
- {
- 'context': 'Next.js 14 | Next.jsBack to BlogThursday, October 26th 2023Next.js 14Posted byLee Robinson@leeerobTim Neutkens@timneutkensAs we announced at Next.js Conf, Next.js 14 is our most focused release with: Turbopack: 5,000 tests passing for App & Pages Router 53% faster local server startup 94% faster code updates with Fast Refresh Server Actions (Stable): Progressively enhanced mutations Integrated with caching & revalidating Simple function calls, or works natively with forms Partial Prerendering',
- 'metadata': {
- 'source': 'https://nextjs.org/blog/next-14',
- 'document_id': '6c8d1a7b-ea34-4927-8823-daa29dcfc5af--b83edb69b8fc7e442ff8ca311b48510e6c80bf00caa806b3a6acb34e1bcdd5d5'
- }
- },
- {
- 'context': 'Next.js 13.3 | Next.jsBack to BlogThursday, April 6th 2023Next.js 13.3Posted byDelba de Oliveira@delba_oliveiraTim Neutkens@timneutkensNext.js 13.3 adds popular community-requested features, including: File-Based Metadata API: Dynamically generate sitemaps, robots, favicons, and more. Dynamic Open Graph Images: Generate OG images using JSX, HTML, and CSS. Static Export for App Router: Static / Single-Page Application (SPA) support for Server Components. Parallel Routes and Interception: Advanced',
- 'metadata': {
- 'source': 'https://nextjs.org/blog/next-13-3',
- 'document_id': '6c8d1a7b-ea34-4927-8823-daa29dcfc5af--b83edb69b8fc7e442ff8ca311b48510e6c80bf00caa806b3a6acb34e1bcdd5d5'
- }
- },
- {
- 'context': 'Upgrading: Version 14 | Next.js MenuUsing App RouterFeatures available in /appApp Router.UpgradingVersion 14Version 14 Upgrading from 13 to 14 To update to Next.js version 14, run the following command using your preferred package manager: Terminalnpm i next@latest react@latest react-dom@latest eslint-config-next@latest Terminalyarn add next@latest react@latest react-dom@latest eslint-config-next@latest Terminalpnpm up next react react-dom eslint-config-next -latest Terminalbun add next@latest',
- 'metadata': {
- 'source': 'https://nextjs.org/docs/app/building-your-application/upgrading/version-14',
- 'document_id': '6c8d1a7b-ea34-4927-8823-daa29dcfc5af--b83edb69b8fc7e442ff8ca311b48510e6c80bf00caa806b3a6acb34e1bcdd5d5'
- }
- }
-]
-```
-The `source` key contains the url of the document that yielded that document chunk.
-
-If you are interested in configuring the search further, refer to our [API documentation](/api-reference/pipeline/search).
-
-### (Optional) Step 4: Deploying Your RAG Pipeline
-
-Want to go live? Deploy your pipeline with these options:
-
-- Deploy on the Embedchain Platform
-- Self-host on your preferred cloud provider
-
-For detailed deployment instructions, follow these guides:
-
-- [Deploying on Embedchain Platform](/get-started/deployment#deploy-on-embedchain-platform)
-- [Self-hosting Guide](/get-started/deployment#self-hosting)
-
-----
-
-This guide will help you swiftly set up a semantic search pipeline with Embedchain, making it easier to access and analyze specific information from large data sources.
-
-
-## Need help?
-
-In case you run into issues, feel free to contact us via any of the following methods:
-
-
diff --git a/embedchain/embedchain/__init__.py b/embedchain/embedchain/__init__.py
deleted file mode 100644
index b59aed77d..000000000
--- a/embedchain/embedchain/__init__.py
+++ /dev/null
@@ -1,10 +0,0 @@
-import importlib.metadata
-
-__version__ = importlib.metadata.version(__package__ or __name__)
-
-from embedchain.app import App # noqa: F401
-from embedchain.client import Client # noqa: F401
-from embedchain.pipeline import Pipeline # noqa: F401
-
-# Setup the user directory if doesn't exist already
-Client.setup()
diff --git a/embedchain/embedchain/alembic.ini b/embedchain/embedchain/alembic.ini
deleted file mode 100644
index 53023ad8d..000000000
--- a/embedchain/embedchain/alembic.ini
+++ /dev/null
@@ -1,116 +0,0 @@
-# A generic, single database configuration.
-
-[alembic]
-# path to migration scripts
-script_location = embedchain:migrations
-
-# template used to generate migration file names; The default value is %%(rev)s_%%(slug)s
-# Uncomment the line below if you want the files to be prepended with date and time
-# see https://alembic.sqlalchemy.org/en/latest/tutorial.html#editing-the-ini-file
-# for all available tokens
-# file_template = %%(year)d_%%(month).2d_%%(day).2d_%%(hour).2d%%(minute).2d-%%(rev)s_%%(slug)s
-
-# sys.path path, will be prepended to sys.path if present.
-# defaults to the current working directory.
-prepend_sys_path = .
-
-# timezone to use when rendering the date within the migration file
-# as well as the filename.
-# If specified, requires the python>=3.9 or backports.zoneinfo library.
-# Any required deps can installed by adding `alembic[tz]` to the pip requirements
-# string value is passed to ZoneInfo()
-# leave blank for localtime
-# timezone =
-
-# max length of characters to apply to the
-# "slug" field
-# truncate_slug_length = 40
-
-# set to 'true' to run the environment during
-# the 'revision' command, regardless of autogenerate
-# revision_environment = false
-
-# set to 'true' to allow .pyc and .pyo files without
-# a source .py file to be detected as revisions in the
-# versions/ directory
-# sourceless = false
-
-# version location specification; This defaults
-# to alembic/versions. When using multiple version
-# directories, initial revisions must be specified with --version-path.
-# The path separator used here should be the separator specified by "version_path_separator" below.
-# version_locations = %(here)s/bar:%(here)s/bat:alembic/versions
-
-# version path separator; As mentioned above, this is the character used to split
-# version_locations. The default within new alembic.ini files is "os", which uses os.pathsep.
-# If this key is omitted entirely, it falls back to the legacy behavior of splitting on spaces and/or commas.
-# Valid values for version_path_separator are:
-#
-# version_path_separator = :
-# version_path_separator = ;
-# version_path_separator = space
-version_path_separator = os # Use os.pathsep. Default configuration used for new projects.
-
-# set to 'true' to search source files recursively
-# in each "version_locations" directory
-# new in Alembic version 1.10
-# recursive_version_locations = false
-
-# the output encoding used when revision files
-# are written from script.py.mako
-# output_encoding = utf-8
-
-sqlalchemy.url = driver://user:pass@localhost/dbname
-
-
-[post_write_hooks]
-# post_write_hooks defines scripts or Python functions that are run
-# on newly generated revision scripts. See the documentation for further
-# detail and examples
-
-# format using "black" - use the console_scripts runner, against the "black" entrypoint
-# hooks = black
-# black.type = console_scripts
-# black.entrypoint = black
-# black.options = -l 79 REVISION_SCRIPT_FILENAME
-
-# lint with attempts to fix using "ruff" - use the exec runner, execute a binary
-# hooks = ruff
-# ruff.type = exec
-# ruff.executable = %(here)s/.venv/bin/ruff
-# ruff.options = --fix REVISION_SCRIPT_FILENAME
-
-# Logging configuration
-[loggers]
-keys = root,sqlalchemy,alembic
-
-[handlers]
-keys = console
-
-[formatters]
-keys = generic
-
-[logger_root]
-level = WARN
-handlers = console
-qualname =
-
-[logger_sqlalchemy]
-level = WARN
-handlers =
-qualname = sqlalchemy.engine
-
-[logger_alembic]
-level = WARN
-handlers =
-qualname = alembic
-
-[handler_console]
-class = StreamHandler
-args = (sys.stderr,)
-level = NOTSET
-formatter = generic
-
-[formatter_generic]
-format = %(levelname)-5.5s [%(name)s] %(message)s
-datefmt = %H:%M:%S
diff --git a/embedchain/embedchain/app.py b/embedchain/embedchain/app.py
deleted file mode 100644
index b4d051607..000000000
--- a/embedchain/embedchain/app.py
+++ /dev/null
@@ -1,517 +0,0 @@
-import ast
-import concurrent.futures
-import json
-import logging
-import os
-from typing import Any, Optional, Union
-
-import requests
-import yaml
-from tqdm import tqdm
-
-from embedchain.cache import (
- Config,
- ExactMatchEvaluation,
- SearchDistanceEvaluation,
- cache,
- gptcache_data_manager,
- gptcache_pre_function,
-)
-from embedchain.client import Client
-from embedchain.config import AppConfig, CacheConfig, ChunkerConfig, Mem0Config
-from embedchain.core.db.database import get_session
-from embedchain.core.db.models import DataSource
-from embedchain.embedchain import EmbedChain
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.embedder.openai import OpenAIEmbedder
-from embedchain.evaluation.base import BaseMetric
-from embedchain.evaluation.metrics import (
- AnswerRelevance,
- ContextRelevance,
- Groundedness,
-)
-from embedchain.factory import EmbedderFactory, LlmFactory, VectorDBFactory
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-from embedchain.llm.openai import OpenAILlm
-from embedchain.telemetry.posthog import AnonymousTelemetry
-from embedchain.utils.evaluation import EvalData, EvalMetric
-from embedchain.utils.misc import validate_config
-from embedchain.vectordb.base import BaseVectorDB
-from embedchain.vectordb.chroma import ChromaDB
-from mem0 import Memory
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class App(EmbedChain):
- """
- EmbedChain App lets you create a LLM powered app for your unstructured
- data by defining your chosen data source, embedding model,
- and vector database.
- """
-
- def __init__(
- self,
- id: str = None,
- name: str = None,
- config: AppConfig = None,
- db: BaseVectorDB = None,
- embedding_model: BaseEmbedder = None,
- llm: BaseLlm = None,
- config_data: dict = None,
- auto_deploy: bool = False,
- chunker: ChunkerConfig = None,
- cache_config: CacheConfig = None,
- memory_config: Mem0Config = None,
- log_level: int = logging.WARN,
- ):
- """
- Initialize a new `App` instance.
-
- :param config: Configuration for the pipeline, defaults to None
- :type config: AppConfig, optional
- :param db: The database to use for storing and retrieving embeddings, defaults to None
- :type db: BaseVectorDB, optional
- :param embedding_model: The embedding model used to calculate embeddings, defaults to None
- :type embedding_model: BaseEmbedder, optional
- :param llm: The LLM model used to calculate embeddings, defaults to None
- :type llm: BaseLlm, optional
- :param config_data: Config dictionary, defaults to None
- :type config_data: dict, optional
- :param auto_deploy: Whether to deploy the pipeline automatically, defaults to False
- :type auto_deploy: bool, optional
- :raises Exception: If an error occurs while creating the pipeline
- """
- if id and config_data:
- raise Exception("Cannot provide both id and config. Please provide only one of them.")
-
- if id and name:
- raise Exception("Cannot provide both id and name. Please provide only one of them.")
-
- if name and config:
- raise Exception("Cannot provide both name and config. Please provide only one of them.")
-
- self.auto_deploy = auto_deploy
- # Store the dict config as an attribute to be able to send it
- self.config_data = config_data if (config_data and validate_config(config_data)) else None
- self.client = None
- # pipeline_id from the backend
- self.id = None
- self.chunker = ChunkerConfig(**chunker) if chunker else None
- self.cache_config = cache_config
- self.memory_config = memory_config
-
- self.config = config or AppConfig()
- self.name = self.config.name
- self.config.id = self.local_id = "default-app-id" if self.config.id is None else self.config.id
-
- if id is not None:
- # Init client first since user is trying to fetch the pipeline
- # details from the platform
- self._init_client()
- pipeline_details = self._get_pipeline(id)
- self.config.id = self.local_id = pipeline_details["metadata"]["local_id"]
- self.id = id
-
- if name is not None:
- self.name = name
-
- self.embedding_model = embedding_model or OpenAIEmbedder()
- self.db = db or ChromaDB()
- self.llm = llm or OpenAILlm()
- self._init_db()
-
- # Session for the metadata db
- self.db_session = get_session()
-
- # If cache_config is provided, initializing the cache ...
- if self.cache_config is not None:
- self._init_cache()
-
- # If memory_config is provided, initializing the memory ...
- self.mem0_memory = None
- if self.memory_config is not None:
- self.mem0_memory = Memory()
-
- # Send anonymous telemetry
- self._telemetry_props = {"class": self.__class__.__name__}
- self.telemetry = AnonymousTelemetry(enabled=self.config.collect_metrics)
- self.telemetry.capture(event_name="init", properties=self._telemetry_props)
-
- self.user_asks = []
- if self.auto_deploy:
- self.deploy()
-
- def _init_db(self):
- """
- Initialize the database.
- """
- self.db._set_embedder(self.embedding_model)
- self.db._initialize()
- self.db.set_collection_name(self.db.config.collection_name)
-
- def _init_cache(self):
- if self.cache_config.similarity_eval_config.strategy == "exact":
- similarity_eval_func = ExactMatchEvaluation()
- else:
- similarity_eval_func = SearchDistanceEvaluation(
- max_distance=self.cache_config.similarity_eval_config.max_distance,
- positive=self.cache_config.similarity_eval_config.positive,
- )
-
- cache.init(
- pre_embedding_func=gptcache_pre_function,
- embedding_func=self.embedding_model.to_embeddings,
- data_manager=gptcache_data_manager(vector_dimension=self.embedding_model.vector_dimension),
- similarity_evaluation=similarity_eval_func,
- config=Config(**self.cache_config.init_config.as_dict()),
- )
-
- def _init_client(self):
- """
- Initialize the client.
- """
- config = Client.load_config()
- if config.get("api_key"):
- self.client = Client()
- else:
- api_key = input(
- "🔑 Enter your Embedchain API key. You can find the API key at https://app.embedchain.ai/settings/keys/ \n" # noqa: E501
- )
- self.client = Client(api_key=api_key)
-
- def _get_pipeline(self, id):
- """
- Get existing pipeline
- """
- print("🛠️ Fetching pipeline details from the platform...")
- url = f"{self.client.host}/api/v1/pipelines/{id}/cli/"
- r = requests.get(
- url,
- headers={"Authorization": f"Token {self.client.api_key}"},
- )
- if r.status_code == 404:
- raise Exception(f"❌ Pipeline with id {id} not found!")
-
- print(
- f"🎉 Pipeline loaded successfully! Pipeline url: https://app.embedchain.ai/pipelines/{r.json()['id']}\n" # noqa: E501
- )
- return r.json()
-
- def _create_pipeline(self):
- """
- Create a pipeline on the platform.
- """
- print("🛠️ Creating pipeline on the platform...")
- # self.config_data is a dict. Pass it inside the key 'yaml_config' to the backend
- payload = {
- "yaml_config": json.dumps(self.config_data),
- "name": self.name,
- "local_id": self.local_id,
- }
- url = f"{self.client.host}/api/v1/pipelines/cli/create/"
- r = requests.post(
- url,
- json=payload,
- headers={"Authorization": f"Token {self.client.api_key}"},
- )
- if r.status_code not in [200, 201]:
- raise Exception(f"❌ Error occurred while creating pipeline. API response: {r.text}")
-
- if r.status_code == 200:
- print(
- f"🎉🎉🎉 Existing pipeline found! View your pipeline: https://app.embedchain.ai/pipelines/{r.json()['id']}\n" # noqa: E501
- ) # noqa: E501
- elif r.status_code == 201:
- print(
- f"🎉🎉🎉 Pipeline created successfully! View your pipeline: https://app.embedchain.ai/pipelines/{r.json()['id']}\n" # noqa: E501
- )
- return r.json()
-
- def _get_presigned_url(self, data_type, data_value):
- payload = {"data_type": data_type, "data_value": data_value}
- r = requests.post(
- f"{self.client.host}/api/v1/pipelines/{self.id}/cli/presigned_url/",
- json=payload,
- headers={"Authorization": f"Token {self.client.api_key}"},
- )
- r.raise_for_status()
- return r.json()
-
- def _upload_file_to_presigned_url(self, presigned_url, file_path):
- try:
- with open(file_path, "rb") as file:
- response = requests.put(presigned_url, data=file)
- response.raise_for_status()
- return response.status_code == 200
- except Exception as e:
- logger.exception(f"Error occurred during file upload: {str(e)}")
- print("❌ Error occurred during file upload!")
- return False
-
- def _upload_data_to_pipeline(self, data_type, data_value, metadata=None):
- payload = {
- "data_type": data_type,
- "data_value": data_value,
- "metadata": metadata,
- }
- try:
- self._send_api_request(f"/api/v1/pipelines/{self.id}/cli/add/", payload)
- # print the local file path if user tries to upload a local file
- printed_value = metadata.get("file_path") if metadata.get("file_path") else data_value
- print(f"✅ Data of type: {data_type}, value: {printed_value} added successfully.")
- except Exception as e:
- print(f"❌ Error occurred during data upload for type {data_type}!. Error: {str(e)}")
-
- def _send_api_request(self, endpoint, payload):
- url = f"{self.client.host}{endpoint}"
- headers = {"Authorization": f"Token {self.client.api_key}"}
- response = requests.post(url, json=payload, headers=headers)
- response.raise_for_status()
- return response
-
- def _process_and_upload_data(self, data_hash, data_type, data_value):
- if os.path.isabs(data_value):
- presigned_url_data = self._get_presigned_url(data_type, data_value)
- presigned_url = presigned_url_data["presigned_url"]
- s3_key = presigned_url_data["s3_key"]
- if self._upload_file_to_presigned_url(presigned_url, file_path=data_value):
- metadata = {"file_path": data_value, "s3_key": s3_key}
- data_value = presigned_url
- else:
- logger.error(f"File upload failed for hash: {data_hash}")
- return False
- else:
- if data_type == "qna_pair":
- data_value = list(ast.literal_eval(data_value))
- metadata = {}
-
- try:
- self._upload_data_to_pipeline(data_type, data_value, metadata)
- self._mark_data_as_uploaded(data_hash)
- return True
- except Exception:
- print(f"❌ Error occurred during data upload for hash {data_hash}!")
- return False
-
- def _mark_data_as_uploaded(self, data_hash):
- self.db_session.query(DataSource).filter_by(hash=data_hash, app_id=self.local_id).update({"is_uploaded": 1})
-
- def get_data_sources(self):
- data_sources = self.db_session.query(DataSource).filter_by(app_id=self.local_id).all()
- results = []
- for row in data_sources:
- results.append({"data_type": row.type, "data_value": row.value, "metadata": row.meta_data})
- return results
-
- def deploy(self):
- if self.client is None:
- self._init_client()
-
- pipeline_data = self._create_pipeline()
- self.id = pipeline_data["id"]
-
- results = self.db_session.query(DataSource).filter_by(app_id=self.local_id, is_uploaded=0).all()
- if len(results) > 0:
- print("🛠️ Adding data to your pipeline...")
- for result in results:
- data_hash, data_type, data_value = result.hash, result.data_type, result.data_value
- self._process_and_upload_data(data_hash, data_type, data_value)
-
- # Send anonymous telemetry
- self.telemetry.capture(event_name="deploy", properties=self._telemetry_props)
-
- @classmethod
- def from_config(
- cls,
- config_path: Optional[str] = None,
- config: Optional[dict[str, Any]] = None,
- auto_deploy: bool = False,
- yaml_path: Optional[str] = None,
- ):
- """
- Instantiate a App object from a configuration.
-
- :param config_path: Path to the YAML or JSON configuration file.
- :type config_path: Optional[str]
- :param config: A dictionary containing the configuration.
- :type config: Optional[dict[str, Any]]
- :param auto_deploy: Whether to deploy the app automatically, defaults to False
- :type auto_deploy: bool, optional
- :param yaml_path: (Deprecated) Path to the YAML configuration file. Use config_path instead.
- :type yaml_path: Optional[str]
- :return: An instance of the App class.
- :rtype: App
- """
- # Backward compatibility for yaml_path
- if yaml_path and not config_path:
- config_path = yaml_path
-
- if config_path and config:
- raise ValueError("Please provide only one of config_path or config.")
-
- config_data = None
-
- if config_path:
- file_extension = os.path.splitext(config_path)[1]
- with open(config_path, "r", encoding="UTF-8") as file:
- if file_extension in [".yaml", ".yml"]:
- config_data = yaml.safe_load(file)
- elif file_extension == ".json":
- config_data = json.load(file)
- else:
- raise ValueError("config_path must be a path to a YAML or JSON file.")
- elif config and isinstance(config, dict):
- config_data = config
- else:
- logger.error(
- "Please provide either a config file path (YAML or JSON) or a config dictionary. Falling back to defaults because no config is provided.", # noqa: E501
- )
- config_data = {}
-
- # Validate the config
- validate_config(config_data)
-
- app_config_data = config_data.get("app", {}).get("config", {})
- vector_db_config_data = config_data.get("vectordb", {})
- embedding_model_config_data = config_data.get("embedding_model", config_data.get("embedder", {}))
- memory_config_data = config_data.get("memory", {})
- llm_config_data = config_data.get("llm", {})
- chunker_config_data = config_data.get("chunker", {})
- cache_config_data = config_data.get("cache", None)
-
- app_config = AppConfig(**app_config_data)
- memory_config = Mem0Config(**memory_config_data) if memory_config_data else None
-
- vector_db_provider = vector_db_config_data.get("provider", "chroma")
- vector_db = VectorDBFactory.create(vector_db_provider, vector_db_config_data.get("config", {}))
-
- if llm_config_data:
- llm_provider = llm_config_data.get("provider", "openai")
- llm = LlmFactory.create(llm_provider, llm_config_data.get("config", {}))
- else:
- llm = None
-
- embedding_model_provider = embedding_model_config_data.get("provider", "openai")
- embedding_model = EmbedderFactory.create(
- embedding_model_provider, embedding_model_config_data.get("config", {})
- )
-
- if cache_config_data is not None:
- cache_config = CacheConfig.from_config(cache_config_data)
- else:
- cache_config = None
-
- return cls(
- config=app_config,
- llm=llm,
- db=vector_db,
- embedding_model=embedding_model,
- config_data=config_data,
- auto_deploy=auto_deploy,
- chunker=chunker_config_data,
- cache_config=cache_config,
- memory_config=memory_config,
- )
-
- def _eval(self, dataset: list[EvalData], metric: Union[BaseMetric, str]):
- """
- Evaluate the app on a dataset for a given metric.
- """
- metric_str = metric.name if isinstance(metric, BaseMetric) else metric
- eval_class_map = {
- EvalMetric.CONTEXT_RELEVANCY.value: ContextRelevance,
- EvalMetric.ANSWER_RELEVANCY.value: AnswerRelevance,
- EvalMetric.GROUNDEDNESS.value: Groundedness,
- }
-
- if metric_str in eval_class_map:
- return eval_class_map[metric_str]().evaluate(dataset)
-
- # Handle the case for custom metrics
- if isinstance(metric, BaseMetric):
- return metric.evaluate(dataset)
- else:
- raise ValueError(f"Invalid metric: {metric}")
-
- def evaluate(
- self,
- questions: Union[str, list[str]],
- metrics: Optional[list[Union[BaseMetric, str]]] = None,
- num_workers: int = 4,
- ):
- """
- Evaluate the app on a question.
-
- param: questions: A question or a list of questions to evaluate.
- type: questions: Union[str, list[str]]
- param: metrics: A list of metrics to evaluate. Defaults to all metrics.
- type: metrics: Optional[list[Union[BaseMetric, str]]]
- param: num_workers: Number of workers to use for parallel processing.
- type: num_workers: int
- return: A dictionary containing the evaluation results.
- rtype: dict
- """
- if "OPENAI_API_KEY" not in os.environ:
- raise ValueError("Please set the OPENAI_API_KEY environment variable with permission to use `gpt4` model.")
-
- queries, answers, contexts = [], [], []
- if isinstance(questions, list):
- with concurrent.futures.ThreadPoolExecutor(max_workers=num_workers) as executor:
- future_to_data = {executor.submit(self.query, q, citations=True): q for q in questions}
- for future in tqdm(
- concurrent.futures.as_completed(future_to_data),
- total=len(future_to_data),
- desc="Getting answer and contexts for questions",
- ):
- question = future_to_data[future]
- queries.append(question)
- answer, context = future.result()
- answers.append(answer)
- contexts.append(list(map(lambda x: x[0], context)))
- else:
- answer, context = self.query(questions, citations=True)
- queries = [questions]
- answers = [answer]
- contexts = [list(map(lambda x: x[0], context))]
-
- metrics = metrics or [
- EvalMetric.CONTEXT_RELEVANCY.value,
- EvalMetric.ANSWER_RELEVANCY.value,
- EvalMetric.GROUNDEDNESS.value,
- ]
-
- logger.info(f"Collecting data from {len(queries)} questions for evaluation...")
- dataset = []
- for q, a, c in zip(queries, answers, contexts):
- dataset.append(EvalData(question=q, answer=a, contexts=c))
-
- logger.info(f"Evaluating {len(dataset)} data points...")
- result = {}
- with concurrent.futures.ThreadPoolExecutor(max_workers=num_workers) as executor:
- future_to_metric = {executor.submit(self._eval, dataset, metric): metric for metric in metrics}
- for future in tqdm(
- concurrent.futures.as_completed(future_to_metric),
- total=len(future_to_metric),
- desc="Evaluating metrics",
- ):
- metric = future_to_metric[future]
- if isinstance(metric, BaseMetric):
- result[metric.name] = future.result()
- else:
- result[metric] = future.result()
-
- if self.config.collect_metrics:
- telemetry_props = self._telemetry_props
- metrics_names = []
- for metric in metrics:
- if isinstance(metric, BaseMetric):
- metrics_names.append(metric.name)
- else:
- metrics_names.append(metric)
- telemetry_props["metrics"] = metrics_names
- self.telemetry.capture(event_name="evaluate", properties=telemetry_props)
-
- return result
diff --git a/embedchain/embedchain/bots/__init__.py b/embedchain/embedchain/bots/__init__.py
deleted file mode 100644
index 34cef58f2..000000000
--- a/embedchain/embedchain/bots/__init__.py
+++ /dev/null
@@ -1,5 +0,0 @@
-from embedchain.bots.poe import PoeBot # noqa: F401
-from embedchain.bots.whatsapp import WhatsAppBot # noqa: F401
-
-# TODO: fix discord import
-# from embedchain.bots.discord import DiscordBot
diff --git a/embedchain/embedchain/bots/base.py b/embedchain/embedchain/bots/base.py
deleted file mode 100644
index 4a817cc4c..000000000
--- a/embedchain/embedchain/bots/base.py
+++ /dev/null
@@ -1,48 +0,0 @@
-from typing import Any
-
-from embedchain import App
-from embedchain.config import AddConfig, AppConfig, BaseLlmConfig
-from embedchain.embedder.openai import OpenAIEmbedder
-from embedchain.helpers.json_serializable import (
- JSONSerializable,
- register_deserializable,
-)
-from embedchain.llm.openai import OpenAILlm
-from embedchain.vectordb.chroma import ChromaDB
-
-
-@register_deserializable
-class BaseBot(JSONSerializable):
- def __init__(self):
- self.app = App(config=AppConfig(), llm=OpenAILlm(), db=ChromaDB(), embedding_model=OpenAIEmbedder())
-
- def add(self, data: Any, config: AddConfig = None):
- """
- Add data to the bot (to the vector database).
- Auto-dectects type only, so some data types might not be usable.
-
- :param data: data to embed
- :type data: Any
- :param config: configuration class instance, defaults to None
- :type config: AddConfig, optional
- """
- config = config if config else AddConfig()
- self.app.add(data, config=config)
-
- def query(self, query: str, config: BaseLlmConfig = None) -> str:
- """
- Query the bot
-
- :param query: the user query
- :type query: str
- :param config: configuration class instance, defaults to None
- :type config: BaseLlmConfig, optional
- :return: Answer
- :rtype: str
- """
- config = config
- return self.app.query(query, config=config)
-
- def start(self):
- """Start the bot's functionality."""
- raise NotImplementedError("Subclasses must implement the start method.")
diff --git a/embedchain/embedchain/bots/discord.py b/embedchain/embedchain/bots/discord.py
deleted file mode 100644
index a288cab6d..000000000
--- a/embedchain/embedchain/bots/discord.py
+++ /dev/null
@@ -1,128 +0,0 @@
-import argparse
-import logging
-import os
-
-from embedchain.helpers.json_serializable import register_deserializable
-
-from .base import BaseBot
-
-try:
- import discord
- from discord import app_commands
- from discord.ext import commands
-except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for Discord are not installed." "Please install with `pip install discord==2.3.2`"
- ) from None
-
-
-logger = logging.getLogger(__name__)
-
-intents = discord.Intents.default()
-intents.message_content = True
-client = discord.Client(intents=intents)
-tree = app_commands.CommandTree(client)
-
-# Invite link example
-# https://discord.com/api/oauth2/authorize?client_id={DISCORD_CLIENT_ID}&permissions=2048&scope=bot
-
-
-@register_deserializable
-class DiscordBot(BaseBot):
- def __init__(self, *args, **kwargs):
- BaseBot.__init__(self, *args, **kwargs)
-
- def add_data(self, message):
- data = message.split(" ")[-1]
- try:
- self.add(data)
- response = f"Added data from: {data}"
- except Exception:
- logger.exception(f"Failed to add data {data}.")
- response = "Some error occurred while adding data."
- return response
-
- def ask_bot(self, message):
- try:
- response = self.query(message)
- except Exception:
- logger.exception(f"Failed to query {message}.")
- response = "An error occurred. Please try again!"
- return response
-
- def start(self):
- client.run(os.environ["DISCORD_BOT_TOKEN"])
-
-
-# @tree decorator cannot be used in a class. A global discord_bot is used as a workaround.
-
-
-@tree.command(name="question", description="ask embedchain")
-async def query_command(interaction: discord.Interaction, question: str):
- await interaction.response.defer()
- member = client.guilds[0].get_member(client.user.id)
- logger.info(f"User: {member}, Query: {question}")
- try:
- answer = discord_bot.ask_bot(question)
- if args.include_question:
- response = f"> {question}\n\n{answer}"
- else:
- response = answer
- await interaction.followup.send(response)
- except Exception as e:
- await interaction.followup.send("An error occurred. Please try again!")
- logger.error("Error occurred during 'query' command:", e)
-
-
-@tree.command(name="add", description="add new content to the embedchain database")
-async def add_command(interaction: discord.Interaction, url_or_text: str):
- await interaction.response.defer()
- member = client.guilds[0].get_member(client.user.id)
- logger.info(f"User: {member}, Add: {url_or_text}")
- try:
- response = discord_bot.add_data(url_or_text)
- await interaction.followup.send(response)
- except Exception as e:
- await interaction.followup.send("An error occurred. Please try again!")
- logger.error("Error occurred during 'add' command:", e)
-
-
-@tree.command(name="ping", description="Simple ping pong command")
-async def ping(interaction: discord.Interaction):
- await interaction.response.send_message("Pong", ephemeral=True)
-
-
-@tree.error
-async def on_app_command_error(interaction: discord.Interaction, error: discord.app_commands.AppCommandError) -> None:
- if isinstance(error, commands.CommandNotFound):
- await interaction.followup.send("Invalid command. Please refer to the documentation for correct syntax.")
- else:
- logger.error("Error occurred during command execution:", error)
-
-
-@client.event
-async def on_ready():
- # TODO: Sync in admin command, to not hit rate limits.
- # This might be overkill for most users, and it would require to set a guild or user id, where sync is allowed.
- await tree.sync()
- logger.debug("Command tree synced")
- logger.info(f"Logged in as {client.user.name}")
-
-
-def start_command():
- parser = argparse.ArgumentParser(description="EmbedChain DiscordBot command line interface")
- parser.add_argument(
- "--include-question",
- help="include question in query reply, otherwise it is hidden behind the slash command.",
- action="store_true",
- )
- global args
- args = parser.parse_args()
-
- global discord_bot
- discord_bot = DiscordBot()
- discord_bot.start()
-
-
-if __name__ == "__main__":
- start_command()
diff --git a/embedchain/embedchain/bots/poe.py b/embedchain/embedchain/bots/poe.py
deleted file mode 100644
index 25c1bba5e..000000000
--- a/embedchain/embedchain/bots/poe.py
+++ /dev/null
@@ -1,87 +0,0 @@
-import argparse
-import logging
-import os
-from typing import Optional
-
-from embedchain.helpers.json_serializable import register_deserializable
-
-from .base import BaseBot
-
-try:
- from fastapi_poe import PoeBot, run
-except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for Poe are not installed." "Please install with `pip install fastapi-poe==0.0.16`"
- ) from None
-
-
-def start_command():
- parser = argparse.ArgumentParser(description="EmbedChain PoeBot command line interface")
- # parser.add_argument("--host", default="0.0.0.0", help="Host IP to bind")
- parser.add_argument("--port", default=8080, type=int, help="Port to bind")
- parser.add_argument("--api-key", type=str, help="Poe API key")
- # parser.add_argument(
- # "--history-length",
- # default=5,
- # type=int,
- # help="Set the max size of the chat history. Multiplies cost, but improves conversation awareness.",
- # )
- args = parser.parse_args()
-
- # FIXME: Arguments are automatically loaded by Poebot's ArgumentParser which causes it to fail.
- # the port argument here is also just for show, it actually works because poe has the same argument.
-
- run(PoeBot(), api_key=args.api_key or os.environ.get("POE_API_KEY"))
-
-
-@register_deserializable
-class PoeBot(BaseBot, PoeBot):
- def __init__(self):
- self.history_length = 5
- super().__init__()
-
- async def get_response(self, query):
- last_message = query.query[-1].content
- try:
- history = (
- [f"{m.role}: {m.content}" for m in query.query[-(self.history_length + 1) : -1]]
- if len(query.query) > 0
- else None
- )
- except Exception as e:
- logging.error(f"Error when processing the chat history. Message is being sent without history. Error: {e}")
- answer = self.handle_message(last_message, history)
- yield self.text_event(answer)
-
- def handle_message(self, message, history: Optional[list[str]] = None):
- if message.startswith("/add "):
- response = self.add_data(message)
- else:
- response = self.ask_bot(message, history)
- return response
-
- # def add_data(self, message):
- # data = message.split(" ")[-1]
- # try:
- # self.add(data)
- # response = f"Added data from: {data}"
- # except Exception:
- # logging.exception(f"Failed to add data {data}.")
- # response = "Some error occurred while adding data."
- # return response
-
- def ask_bot(self, message, history: list[str]):
- try:
- self.app.llm.set_history(history=history)
- response = self.query(message)
- except Exception:
- logging.exception(f"Failed to query {message}.")
- response = "An error occurred. Please try again!"
- return response
-
- def start(self):
- start_command()
-
-
-if __name__ == "__main__":
- start_command()
diff --git a/embedchain/embedchain/bots/slack.py b/embedchain/embedchain/bots/slack.py
deleted file mode 100644
index be23fddd9..000000000
--- a/embedchain/embedchain/bots/slack.py
+++ /dev/null
@@ -1,101 +0,0 @@
-import argparse
-import logging
-import os
-import signal
-import sys
-
-from embedchain import App
-from embedchain.helpers.json_serializable import register_deserializable
-
-from .base import BaseBot
-
-try:
- from flask import Flask, request
- from slack_sdk import WebClient
-except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for Slack are not installed."
- "Please install with `pip install slack-sdk==3.21.3 flask==2.3.3`"
- ) from None
-
-
-logger = logging.getLogger(__name__)
-
-SLACK_BOT_TOKEN = os.environ.get("SLACK_BOT_TOKEN")
-
-
-@register_deserializable
-class SlackBot(BaseBot):
- def __init__(self):
- self.client = WebClient(token=SLACK_BOT_TOKEN)
- self.chat_bot = App()
- self.recent_message = {"ts": 0, "channel": ""}
- super().__init__()
-
- def handle_message(self, event_data):
- message = event_data.get("event")
- if message and "text" in message and message.get("subtype") != "bot_message":
- text: str = message["text"]
- if float(message.get("ts")) > float(self.recent_message["ts"]):
- self.recent_message["ts"] = message["ts"]
- self.recent_message["channel"] = message["channel"]
- if text.startswith("query"):
- _, question = text.split(" ", 1)
- try:
- response = self.chat_bot.chat(question)
- self.send_slack_message(message["channel"], response)
- logger.info("Query answered successfully!")
- except Exception as e:
- self.send_slack_message(message["channel"], "An error occurred. Please try again!")
- logger.error("Error occurred during 'query' command:", e)
- elif text.startswith("add"):
- _, data_type, url_or_text = text.split(" ", 2)
- if url_or_text.startswith("<") and url_or_text.endswith(">"):
- url_or_text = url_or_text[1:-1]
- try:
- self.chat_bot.add(url_or_text, data_type)
- self.send_slack_message(message["channel"], f"Added {data_type} : {url_or_text}")
- except ValueError as e:
- self.send_slack_message(message["channel"], f"Error: {str(e)}")
- logger.error("Error occurred during 'add' command:", e)
- except Exception as e:
- self.send_slack_message(message["channel"], f"Failed to add {data_type} : {url_or_text}")
- logger.error("Error occurred during 'add' command:", e)
-
- def send_slack_message(self, channel, message):
- response = self.client.chat_postMessage(channel=channel, text=message)
- return response
-
- def start(self, host="0.0.0.0", port=5000, debug=True):
- app = Flask(__name__)
-
- def signal_handler(sig, frame):
- logger.info("\nGracefully shutting down the SlackBot...")
- sys.exit(0)
-
- signal.signal(signal.SIGINT, signal_handler)
-
- @app.route("/", methods=["POST"])
- def chat():
- # Check if the request is a verification request
- if request.json.get("challenge"):
- return str(request.json.get("challenge"))
-
- response = self.handle_message(request.json)
- return str(response)
-
- app.run(host=host, port=port, debug=debug)
-
-
-def start_command():
- parser = argparse.ArgumentParser(description="EmbedChain SlackBot command line interface")
- parser.add_argument("--host", default="0.0.0.0", help="Host IP to bind")
- parser.add_argument("--port", default=5000, type=int, help="Port to bind")
- args = parser.parse_args()
-
- slack_bot = SlackBot()
- slack_bot.start(host=args.host, port=args.port)
-
-
-if __name__ == "__main__":
- start_command()
diff --git a/embedchain/embedchain/bots/whatsapp.py b/embedchain/embedchain/bots/whatsapp.py
deleted file mode 100644
index bec926bbe..000000000
--- a/embedchain/embedchain/bots/whatsapp.py
+++ /dev/null
@@ -1,83 +0,0 @@
-import argparse
-import importlib
-import logging
-import signal
-import sys
-
-from embedchain.helpers.json_serializable import register_deserializable
-
-from .base import BaseBot
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class WhatsAppBot(BaseBot):
- def __init__(self):
- try:
- self.flask = importlib.import_module("flask")
- self.twilio = importlib.import_module("twilio")
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for WhatsApp are not installed. "
- "Please install with `pip install twilio==8.5.0 flask==2.3.3`"
- ) from None
- super().__init__()
-
- def handle_message(self, message):
- if message.startswith("add "):
- response = self.add_data(message)
- else:
- response = self.ask_bot(message)
- return response
-
- def add_data(self, message):
- data = message.split(" ")[-1]
- try:
- self.add(data)
- response = f"Added data from: {data}"
- except Exception:
- logger.exception(f"Failed to add data {data}.")
- response = "Some error occurred while adding data."
- return response
-
- def ask_bot(self, message):
- try:
- response = self.query(message)
- except Exception:
- logger.exception(f"Failed to query {message}.")
- response = "An error occurred. Please try again!"
- return response
-
- def start(self, host="0.0.0.0", port=5000, debug=True):
- app = self.flask.Flask(__name__)
-
- def signal_handler(sig, frame):
- logger.info("\nGracefully shutting down the WhatsAppBot...")
- sys.exit(0)
-
- signal.signal(signal.SIGINT, signal_handler)
-
- @app.route("/chat", methods=["POST"])
- def chat():
- incoming_message = self.flask.request.values.get("Body", "").lower()
- response = self.handle_message(incoming_message)
- twilio_response = self.twilio.twiml.messaging_response.MessagingResponse()
- twilio_response.message(response)
- return str(twilio_response)
-
- app.run(host=host, port=port, debug=debug)
-
-
-def start_command():
- parser = argparse.ArgumentParser(description="EmbedChain WhatsAppBot command line interface")
- parser.add_argument("--host", default="0.0.0.0", help="Host IP to bind")
- parser.add_argument("--port", default=5000, type=int, help="Port to bind")
- args = parser.parse_args()
-
- whatsapp_bot = WhatsAppBot()
- whatsapp_bot.start(host=args.host, port=args.port)
-
-
-if __name__ == "__main__":
- start_command()
diff --git a/embedchain/embedchain/cache.py b/embedchain/embedchain/cache.py
deleted file mode 100644
index 765141c3c..000000000
--- a/embedchain/embedchain/cache.py
+++ /dev/null
@@ -1,46 +0,0 @@
-import logging
-import os # noqa: F401
-from typing import Any
-
-from gptcache import cache # noqa: F401
-from gptcache.adapter.adapter import adapt # noqa: F401
-from gptcache.config import Config # noqa: F401
-from gptcache.manager import get_data_manager
-from gptcache.manager.scalar_data.base import Answer
-from gptcache.manager.scalar_data.base import DataType as CacheDataType
-from gptcache.session import Session
-from gptcache.similarity_evaluation.distance import ( # noqa: F401
- SearchDistanceEvaluation,
-)
-from gptcache.similarity_evaluation.exact_match import ( # noqa: F401
- ExactMatchEvaluation,
-)
-
-logger = logging.getLogger(__name__)
-
-
-def gptcache_pre_function(data: dict[str, Any], **params: dict[str, Any]):
- return data["input_query"]
-
-
-def gptcache_data_manager(vector_dimension):
- return get_data_manager(cache_base="sqlite", vector_base="chromadb", max_size=1000, eviction="LRU")
-
-
-def gptcache_data_convert(cache_data):
- logger.info("[Cache] Cache hit, returning cache data...")
- return cache_data
-
-
-def gptcache_update_cache_callback(llm_data, update_cache_func, *args, **kwargs):
- logger.info("[Cache] Cache missed, updating cache...")
- update_cache_func(Answer(llm_data, CacheDataType.STR))
- return llm_data
-
-
-def _gptcache_session_hit_func(cur_session_id: str, cache_session_ids: list, cache_questions: list, cache_answer: str):
- return cur_session_id in cache_session_ids
-
-
-def get_gptcache_session(session_id: str):
- return Session(name=session_id, check_hit_func=_gptcache_session_hit_func)
diff --git a/embedchain/embedchain/chunkers/__init__.py b/embedchain/embedchain/chunkers/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/chunkers/audio.py b/embedchain/embedchain/chunkers/audio.py
deleted file mode 100644
index 0aebda32e..000000000
--- a/embedchain/embedchain/chunkers/audio.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class AudioChunker(BaseChunker):
- """Chunker for audio."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/base_chunker.py b/embedchain/embedchain/chunkers/base_chunker.py
deleted file mode 100644
index 1f04a7d3f..000000000
--- a/embedchain/embedchain/chunkers/base_chunker.py
+++ /dev/null
@@ -1,94 +0,0 @@
-import hashlib
-import logging
-from typing import Any, Optional
-
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import JSONSerializable
-from embedchain.models.data_type import DataType
-
-logger = logging.getLogger(__name__)
-
-
-class BaseChunker(JSONSerializable):
- def __init__(self, text_splitter):
- """Initialize the chunker."""
- self.text_splitter = text_splitter
- self.data_type = None
-
- def create_chunks(
- self,
- loader,
- src,
- app_id=None,
- config: Optional[ChunkerConfig] = None,
- **kwargs: Optional[dict[str, Any]],
- ):
- """
- Loads data and chunks it.
-
- :param loader: The loader whose `load_data` method is used to create
- the raw data.
- :param src: The data to be handled by the loader. Can be a URL for
- remote sources or local content for local loaders.
- :param app_id: App id used to generate the doc_id.
- """
- documents = []
- chunk_ids = []
- id_map = {}
- min_chunk_size = config.min_chunk_size if config is not None else 1
- logger.info(f"Skipping chunks smaller than {min_chunk_size} characters")
- data_result = loader.load_data(src, **kwargs)
- data_records = data_result["data"]
- doc_id = data_result["doc_id"]
- # Prefix app_id in the document id if app_id is not None to
- # distinguish between different documents stored in the same
- # elasticsearch or opensearch index
- doc_id = f"{app_id}--{doc_id}" if app_id is not None else doc_id
- metadatas = []
- for data in data_records:
- content = data["content"]
-
- metadata = data["meta_data"]
- # add data type to meta data to allow query using data type
- metadata["data_type"] = self.data_type.value
- metadata["doc_id"] = doc_id
-
- # TODO: Currently defaulting to the src as the url. This is done intentianally since some
- # of the data types like 'gmail' loader doesn't have the url in the meta data.
- url = metadata.get("url", src)
-
- chunks = self.get_chunks(content)
- for chunk in chunks:
- chunk_id = hashlib.sha256((chunk + url).encode()).hexdigest()
- chunk_id = f"{app_id}--{chunk_id}" if app_id is not None else chunk_id
- if id_map.get(chunk_id) is None and len(chunk) >= min_chunk_size:
- id_map[chunk_id] = True
- chunk_ids.append(chunk_id)
- documents.append(chunk)
- metadatas.append(metadata)
- return {
- "documents": documents,
- "ids": chunk_ids,
- "metadatas": metadatas,
- "doc_id": doc_id,
- }
-
- def get_chunks(self, content):
- """
- Returns chunks using text splitter instance.
-
- Override in child class if custom logic.
- """
- return self.text_splitter.split_text(content)
-
- def set_data_type(self, data_type: DataType):
- """
- set the data type of chunker
- """
- self.data_type = data_type
-
- # TODO: This should be done during initialization. This means it has to be done in the child classes.
-
- @staticmethod
- def get_word_count(documents) -> int:
- return sum(len(document.split(" ")) for document in documents)
diff --git a/embedchain/embedchain/chunkers/beehiiv.py b/embedchain/embedchain/chunkers/beehiiv.py
deleted file mode 100644
index 7c130d542..000000000
--- a/embedchain/embedchain/chunkers/beehiiv.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class BeehiivChunker(BaseChunker):
- """Chunker for Beehiiv."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/common_chunker.py b/embedchain/embedchain/chunkers/common_chunker.py
deleted file mode 100644
index 53676d400..000000000
--- a/embedchain/embedchain/chunkers/common_chunker.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class CommonChunker(BaseChunker):
- """Common chunker for all loaders."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=2000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/discourse.py b/embedchain/embedchain/chunkers/discourse.py
deleted file mode 100644
index 14898bf01..000000000
--- a/embedchain/embedchain/chunkers/discourse.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class DiscourseChunker(BaseChunker):
- """Chunker for discourse."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/docs_site.py b/embedchain/embedchain/chunkers/docs_site.py
deleted file mode 100644
index d51dc8ee2..000000000
--- a/embedchain/embedchain/chunkers/docs_site.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class DocsSiteChunker(BaseChunker):
- """Chunker for code docs site."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=500, chunk_overlap=50, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/docx_file.py b/embedchain/embedchain/chunkers/docx_file.py
deleted file mode 100644
index 1452349e8..000000000
--- a/embedchain/embedchain/chunkers/docx_file.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class DocxFileChunker(BaseChunker):
- """Chunker for .docx file."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/excel_file.py b/embedchain/embedchain/chunkers/excel_file.py
deleted file mode 100644
index 7de00a52f..000000000
--- a/embedchain/embedchain/chunkers/excel_file.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class ExcelFileChunker(BaseChunker):
- """Chunker for Excel file."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/gmail.py b/embedchain/embedchain/chunkers/gmail.py
deleted file mode 100644
index 6b804f546..000000000
--- a/embedchain/embedchain/chunkers/gmail.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class GmailChunker(BaseChunker):
- """Chunker for gmail."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/google_drive.py b/embedchain/embedchain/chunkers/google_drive.py
deleted file mode 100644
index 8440325b5..000000000
--- a/embedchain/embedchain/chunkers/google_drive.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class GoogleDriveChunker(BaseChunker):
- """Chunker for google drive folder."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/image.py b/embedchain/embedchain/chunkers/image.py
deleted file mode 100644
index d29a84f4d..000000000
--- a/embedchain/embedchain/chunkers/image.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class ImageChunker(BaseChunker):
- """Chunker for Images."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=2000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/json.py b/embedchain/embedchain/chunkers/json.py
deleted file mode 100644
index ebc525419..000000000
--- a/embedchain/embedchain/chunkers/json.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class JSONChunker(BaseChunker):
- """Chunker for json."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/mdx.py b/embedchain/embedchain/chunkers/mdx.py
deleted file mode 100644
index 1c277dda7..000000000
--- a/embedchain/embedchain/chunkers/mdx.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class MdxChunker(BaseChunker):
- """Chunker for mdx files."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/mysql.py b/embedchain/embedchain/chunkers/mysql.py
deleted file mode 100644
index 2b1c11ace..000000000
--- a/embedchain/embedchain/chunkers/mysql.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class MySQLChunker(BaseChunker):
- """Chunker for json."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/notion.py b/embedchain/embedchain/chunkers/notion.py
deleted file mode 100644
index 190d59b57..000000000
--- a/embedchain/embedchain/chunkers/notion.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class NotionChunker(BaseChunker):
- """Chunker for notion."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=300, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/openapi.py b/embedchain/embedchain/chunkers/openapi.py
deleted file mode 100644
index fbe7b708b..000000000
--- a/embedchain/embedchain/chunkers/openapi.py
+++ /dev/null
@@ -1,18 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-
-
-class OpenAPIChunker(BaseChunker):
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/pdf_file.py b/embedchain/embedchain/chunkers/pdf_file.py
deleted file mode 100644
index 56bae064e..000000000
--- a/embedchain/embedchain/chunkers/pdf_file.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class PdfFileChunker(BaseChunker):
- """Chunker for PDF file."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/postgres.py b/embedchain/embedchain/chunkers/postgres.py
deleted file mode 100644
index 7c6859bd0..000000000
--- a/embedchain/embedchain/chunkers/postgres.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class PostgresChunker(BaseChunker):
- """Chunker for postgres."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/qna_pair.py b/embedchain/embedchain/chunkers/qna_pair.py
deleted file mode 100644
index c0d8277b1..000000000
--- a/embedchain/embedchain/chunkers/qna_pair.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class QnaPairChunker(BaseChunker):
- """Chunker for QnA pair."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=300, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/rss_feed.py b/embedchain/embedchain/chunkers/rss_feed.py
deleted file mode 100644
index 1767f9edd..000000000
--- a/embedchain/embedchain/chunkers/rss_feed.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class RSSFeedChunker(BaseChunker):
- """Chunker for RSS Feed."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=2000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/sitemap.py b/embedchain/embedchain/chunkers/sitemap.py
deleted file mode 100644
index 64e773742..000000000
--- a/embedchain/embedchain/chunkers/sitemap.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class SitemapChunker(BaseChunker):
- """Chunker for sitemap."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=500, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/slack.py b/embedchain/embedchain/chunkers/slack.py
deleted file mode 100644
index 595682beb..000000000
--- a/embedchain/embedchain/chunkers/slack.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class SlackChunker(BaseChunker):
- """Chunker for postgres."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/substack.py b/embedchain/embedchain/chunkers/substack.py
deleted file mode 100644
index 92cacd6cb..000000000
--- a/embedchain/embedchain/chunkers/substack.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class SubstackChunker(BaseChunker):
- """Chunker for Substack."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/table.py b/embedchain/embedchain/chunkers/table.py
deleted file mode 100644
index 567ed6541..000000000
--- a/embedchain/embedchain/chunkers/table.py
+++ /dev/null
@@ -1,20 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-
-
-class TableChunker(BaseChunker):
- """Chunker for tables, for instance csv, google sheets or databases."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=300, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/text.py b/embedchain/embedchain/chunkers/text.py
deleted file mode 100644
index f33d60c46..000000000
--- a/embedchain/embedchain/chunkers/text.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class TextChunker(BaseChunker):
- """Chunker for text."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=300, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/unstructured_file.py b/embedchain/embedchain/chunkers/unstructured_file.py
deleted file mode 100644
index d55f23ef0..000000000
--- a/embedchain/embedchain/chunkers/unstructured_file.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class UnstructuredFileChunker(BaseChunker):
- """Chunker for Unstructured file."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=1000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/web_page.py b/embedchain/embedchain/chunkers/web_page.py
deleted file mode 100644
index 5ef7f40df..000000000
--- a/embedchain/embedchain/chunkers/web_page.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class WebPageChunker(BaseChunker):
- """Chunker for web page."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=2000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/xml.py b/embedchain/embedchain/chunkers/xml.py
deleted file mode 100644
index c1bab0a77..000000000
--- a/embedchain/embedchain/chunkers/xml.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class XmlChunker(BaseChunker):
- """Chunker for XML files."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=500, chunk_overlap=50, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/chunkers/youtube_video.py b/embedchain/embedchain/chunkers/youtube_video.py
deleted file mode 100644
index bde0a8f78..000000000
--- a/embedchain/embedchain/chunkers/youtube_video.py
+++ /dev/null
@@ -1,22 +0,0 @@
-from typing import Optional
-
-from langchain.text_splitter import RecursiveCharacterTextSplitter
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config.add_config import ChunkerConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class YoutubeVideoChunker(BaseChunker):
- """Chunker for Youtube video."""
-
- def __init__(self, config: Optional[ChunkerConfig] = None):
- if config is None:
- config = ChunkerConfig(chunk_size=2000, chunk_overlap=0, length_function=len)
- text_splitter = RecursiveCharacterTextSplitter(
- chunk_size=config.chunk_size,
- chunk_overlap=config.chunk_overlap,
- length_function=config.length_function,
- )
- super().__init__(text_splitter)
diff --git a/embedchain/embedchain/cli.py b/embedchain/embedchain/cli.py
deleted file mode 100644
index e4f0401d5..000000000
--- a/embedchain/embedchain/cli.py
+++ /dev/null
@@ -1,335 +0,0 @@
-import json
-import os
-import shutil
-import signal
-import subprocess
-import sys
-import tempfile
-import time
-import zipfile
-from pathlib import Path
-
-import click
-import requests
-from rich.console import Console
-
-from embedchain.telemetry.posthog import AnonymousTelemetry
-from embedchain.utils.cli import (
- deploy_fly,
- deploy_gradio_app,
- deploy_hf_spaces,
- deploy_modal,
- deploy_render,
- deploy_streamlit,
- get_pkg_path_from_name,
- setup_fly_io_app,
- setup_gradio_app,
- setup_hf_app,
- setup_modal_com_app,
- setup_render_com_app,
- setup_streamlit_io_app,
-)
-
-console = Console()
-api_process = None
-ui_process = None
-
-anonymous_telemetry = AnonymousTelemetry()
-
-
-def signal_handler(sig, frame):
- """Signal handler to catch termination signals and kill server processes."""
- global api_process, ui_process
- console.print("\n🛑 [bold yellow]Stopping servers...[/bold yellow]")
- if api_process:
- api_process.terminate()
- console.print("🛑 [bold yellow]API server stopped.[/bold yellow]")
- if ui_process:
- ui_process.terminate()
- console.print("🛑 [bold yellow]UI server stopped.[/bold yellow]")
- sys.exit(0)
-
-
-@click.group()
-def cli():
- pass
-
-
-@cli.command()
-@click.argument("app_name")
-@click.option("--docker", is_flag=True, help="Use docker to create the app.")
-@click.pass_context
-def create_app(ctx, app_name, docker):
- if Path(app_name).exists():
- console.print(
- f"❌ [red]Directory '{app_name}' already exists. Try using a new directory name, or remove it.[/red]"
- )
- return
-
- os.makedirs(app_name)
- os.chdir(app_name)
-
- # Step 1: Download the zip file
- zip_url = "http://github.com/embedchain/ec-admin/archive/main.zip"
- console.print(f"Creating a new embedchain app in [green]{Path().resolve()}[/green]\n")
- try:
- response = requests.get(zip_url)
- response.raise_for_status()
- with tempfile.NamedTemporaryFile(delete=False) as tmp_file:
- tmp_file.write(response.content)
- zip_file_path = tmp_file.name
- console.print("✅ [bold green]Fetched template successfully.[/bold green]")
- except requests.RequestException as e:
- console.print(f"❌ [bold red]Failed to download zip file: {e}[/bold red]")
- anonymous_telemetry.capture(event_name="ec_create_app", properties={"success": False})
- return
-
- # Step 2: Extract the zip file
- try:
- with zipfile.ZipFile(zip_file_path, "r") as zip_ref:
- # Get the name of the root directory inside the zip file
- root_dir = Path(zip_ref.namelist()[0])
- for member in zip_ref.infolist():
- # Build the path to extract the file to, skipping the root directory
- target_file = Path(member.filename).relative_to(root_dir)
- source_file = zip_ref.open(member, "r")
- if member.is_dir():
- # Create directory if it doesn't exist
- os.makedirs(target_file, exist_ok=True)
- else:
- with open(target_file, "wb") as file:
- # Write the file
- shutil.copyfileobj(source_file, file)
- console.print("✅ [bold green]Extracted zip file successfully.[/bold green]")
- anonymous_telemetry.capture(event_name="ec_create_app", properties={"success": True})
- except zipfile.BadZipFile:
- console.print("❌ [bold red]Error in extracting zip file. The file might be corrupted.[/bold red]")
- anonymous_telemetry.capture(event_name="ec_create_app", properties={"success": False})
- return
-
- if docker:
- subprocess.run(["docker-compose", "build"], check=True)
- else:
- ctx.invoke(install_reqs)
-
-
-@cli.command()
-def install_reqs():
- try:
- console.print("Installing python requirements...\n")
- time.sleep(2)
- os.chdir("api")
- subprocess.run(["pip", "install", "-r", "requirements.txt"], check=True)
- os.chdir("..")
- console.print("\n ✅ [bold green]Installed API requirements successfully.[/bold green]\n")
- except Exception as e:
- console.print(f"❌ [bold red]Failed to install API requirements: {e}[/bold red]")
- anonymous_telemetry.capture(event_name="ec_install_reqs", properties={"success": False})
- return
-
- try:
- os.chdir("ui")
- subprocess.run(["yarn"], check=True)
- console.print("\n✅ [bold green]Successfully installed frontend requirements.[/bold green]")
- anonymous_telemetry.capture(event_name="ec_install_reqs", properties={"success": True})
- except Exception as e:
- console.print(f"❌ [bold red]Failed to install frontend requirements. Error: {e}[/bold red]")
- anonymous_telemetry.capture(event_name="ec_install_reqs", properties={"success": False})
-
-
-@cli.command()
-@click.option("--docker", is_flag=True, help="Run inside docker.")
-def start(docker):
- if docker:
- subprocess.run(["docker-compose", "up"], check=True)
- return
-
- # Set up signal handling
- signal.signal(signal.SIGINT, signal_handler)
- signal.signal(signal.SIGTERM, signal_handler)
-
- # Step 1: Start the API server
- try:
- os.chdir("api")
- api_process = subprocess.Popen(["python", "-m", "main"], stdout=None, stderr=None)
- os.chdir("..")
- console.print("✅ [bold green]API server started successfully.[/bold green]")
- except Exception as e:
- console.print(f"❌ [bold red]Failed to start the API server: {e}[/bold red]")
- anonymous_telemetry.capture(event_name="ec_start", properties={"success": False})
- return
-
- # Sleep for 2 seconds to give the user time to read the message
- time.sleep(2)
-
- # Step 2: Install UI requirements and start the UI server
- try:
- os.chdir("ui")
- subprocess.run(["yarn"], check=True)
- ui_process = subprocess.Popen(["yarn", "dev"])
- console.print("✅ [bold green]UI server started successfully.[/bold green]")
- anonymous_telemetry.capture(event_name="ec_start", properties={"success": True})
- except Exception as e:
- console.print(f"❌ [bold red]Failed to start the UI server: {e}[/bold red]")
- anonymous_telemetry.capture(event_name="ec_start", properties={"success": False})
-
- # Keep the script running until it receives a kill signal
- try:
- api_process.wait()
- ui_process.wait()
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]Stopping server...[/bold yellow]")
-
-
-@cli.command()
-@click.option("--template", default="fly.io", help="The template to use.")
-@click.argument("extra_args", nargs=-1, type=click.UNPROCESSED)
-def create(template, extra_args):
- anonymous_telemetry.capture(event_name="ec_create", properties={"template_used": template})
- template_dir = template
- if "/" in template_dir:
- template_dir = template.split("/")[1]
- src_path = get_pkg_path_from_name(template_dir)
- shutil.copytree(src_path, os.getcwd(), dirs_exist_ok=True)
- console.print(f"✅ [bold green]Successfully created app from template '{template}'.[/bold green]")
-
- if template == "fly.io":
- setup_fly_io_app(extra_args)
- elif template == "modal.com":
- setup_modal_com_app(extra_args)
- elif template == "render.com":
- setup_render_com_app()
- elif template == "streamlit.io":
- setup_streamlit_io_app()
- elif template == "gradio.app":
- setup_gradio_app()
- elif template == "hf/gradio.app" or template == "hf/streamlit.io":
- setup_hf_app()
- else:
- raise ValueError(f"Unknown template '{template}'.")
-
- embedchain_config = {"provider": template}
- with open("embedchain.json", "w") as file:
- json.dump(embedchain_config, file, indent=4)
- console.print(
- f"🎉 [green]All done! Successfully created `embedchain.json` with '{template}' as provider.[/green]"
- )
-
-
-def run_dev_fly_io(debug, host, port):
- uvicorn_command = ["uvicorn", "app:app"]
-
- if debug:
- uvicorn_command.append("--reload")
-
- uvicorn_command.extend(["--host", host, "--port", str(port)])
-
- try:
- console.print(f"🚀 [bold cyan]Running FastAPI app with command: {' '.join(uvicorn_command)}[/bold cyan]")
- subprocess.run(uvicorn_command, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]FastAPI server stopped[/bold yellow]")
-
-
-def run_dev_modal_com():
- modal_run_cmd = ["modal", "serve", "app"]
- try:
- console.print(f"🚀 [bold cyan]Running FastAPI app with command: {' '.join(modal_run_cmd)}[/bold cyan]")
- subprocess.run(modal_run_cmd, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]FastAPI server stopped[/bold yellow]")
-
-
-def run_dev_streamlit_io():
- streamlit_run_cmd = ["streamlit", "run", "app.py"]
- try:
- console.print(f"🚀 [bold cyan]Running Streamlit app with command: {' '.join(streamlit_run_cmd)}[/bold cyan]")
- subprocess.run(streamlit_run_cmd, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]Streamlit server stopped[/bold yellow]")
-
-
-def run_dev_render_com(debug, host, port):
- uvicorn_command = ["uvicorn", "app:app"]
-
- if debug:
- uvicorn_command.append("--reload")
-
- uvicorn_command.extend(["--host", host, "--port", str(port)])
-
- try:
- console.print(f"🚀 [bold cyan]Running FastAPI app with command: {' '.join(uvicorn_command)}[/bold cyan]")
- subprocess.run(uvicorn_command, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]FastAPI server stopped[/bold yellow]")
-
-
-def run_dev_gradio():
- gradio_run_cmd = ["gradio", "app.py"]
- try:
- console.print(f"🚀 [bold cyan]Running Gradio app with command: {' '.join(gradio_run_cmd)}[/bold cyan]")
- subprocess.run(gradio_run_cmd, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]Gradio server stopped[/bold yellow]")
-
-
-@cli.command()
-@click.option("--debug", is_flag=True, help="Enable or disable debug mode.")
-@click.option("--host", default="127.0.0.1", help="The host address to run the FastAPI app on.")
-@click.option("--port", default=8000, help="The port to run the FastAPI app on.")
-def dev(debug, host, port):
- template = ""
- with open("embedchain.json", "r") as file:
- embedchain_config = json.load(file)
- template = embedchain_config["provider"]
-
- anonymous_telemetry.capture(event_name="ec_dev", properties={"template_used": template})
- if template == "fly.io":
- run_dev_fly_io(debug, host, port)
- elif template == "modal.com":
- run_dev_modal_com()
- elif template == "render.com":
- run_dev_render_com(debug, host, port)
- elif template == "streamlit.io" or template == "hf/streamlit.io":
- run_dev_streamlit_io()
- elif template == "gradio.app" or template == "hf/gradio.app":
- run_dev_gradio()
- else:
- raise ValueError(f"Unknown template '{template}'.")
-
-
-@cli.command()
-def deploy():
- # Check for platform-specific files
- template = ""
- ec_app_name = ""
- with open("embedchain.json", "r") as file:
- embedchain_config = json.load(file)
- ec_app_name = embedchain_config["name"] if "name" in embedchain_config else None
- template = embedchain_config["provider"]
-
- anonymous_telemetry.capture(event_name="ec_deploy", properties={"template_used": template})
- if template == "fly.io":
- deploy_fly()
- elif template == "modal.com":
- deploy_modal()
- elif template == "render.com":
- deploy_render()
- elif template == "streamlit.io":
- deploy_streamlit()
- elif template == "gradio.app":
- deploy_gradio_app()
- elif template.startswith("hf/"):
- deploy_hf_spaces(ec_app_name)
- else:
- console.print("❌ [bold red]No recognized deployment platform found.[/bold red]")
diff --git a/embedchain/embedchain/client.py b/embedchain/embedchain/client.py
deleted file mode 100644
index 7e8fcddbb..000000000
--- a/embedchain/embedchain/client.py
+++ /dev/null
@@ -1,103 +0,0 @@
-import json
-import logging
-import os
-import uuid
-
-import requests
-
-from embedchain.constants import CONFIG_DIR, CONFIG_FILE
-
-logger = logging.getLogger(__name__)
-
-
-class Client:
- def __init__(self, api_key=None, host="https://apiv2.embedchain.ai"):
- self.config_data = self.load_config()
- self.host = host
-
- if api_key:
- if self.check(api_key):
- self.api_key = api_key
- self.save()
- else:
- raise ValueError(
- "Invalid API key provided. You can find your API key on https://app.embedchain.ai/settings/keys."
- )
- else:
- if "api_key" in self.config_data:
- self.api_key = self.config_data["api_key"]
- logger.info("API key loaded successfully!")
- else:
- raise ValueError(
- "You are not logged in. Please obtain an API key from https://app.embedchain.ai/settings/keys/"
- )
-
- @classmethod
- def setup(cls):
- """
- Loads the user id from the config file if it exists, otherwise generates a new
- one and saves it to the config file.
-
- :return: user id
- :rtype: str
- """
- os.makedirs(CONFIG_DIR, exist_ok=True)
-
- if os.path.exists(CONFIG_FILE):
- with open(CONFIG_FILE, "r") as f:
- data = json.load(f)
- if "user_id" in data:
- return data["user_id"]
-
- u_id = str(uuid.uuid4())
- with open(CONFIG_FILE, "w") as f:
- json.dump({"user_id": u_id}, f)
-
- @classmethod
- def load_config(cls):
- if not os.path.exists(CONFIG_FILE):
- cls.setup()
-
- with open(CONFIG_FILE, "r") as config_file:
- return json.load(config_file)
-
- def save(self):
- self.config_data["api_key"] = self.api_key
- with open(CONFIG_FILE, "w") as config_file:
- json.dump(self.config_data, config_file, indent=4)
-
- logger.info("API key saved successfully!")
-
- def clear(self):
- if "api_key" in self.config_data:
- del self.config_data["api_key"]
- with open(CONFIG_FILE, "w") as config_file:
- json.dump(self.config_data, config_file, indent=4)
- self.api_key = None
- logger.info("API key deleted successfully!")
- else:
- logger.warning("API key not found in the configuration file.")
-
- def update(self, api_key):
- if self.check(api_key):
- self.api_key = api_key
- self.save()
- logger.info("API key updated successfully!")
- else:
- logger.warning("Invalid API key provided. API key not updated.")
-
- def check(self, api_key):
- validation_url = f"{self.host}/api/v1/accounts/api_keys/validate/"
- response = requests.post(validation_url, headers={"Authorization": f"Token {api_key}"})
- if response.status_code == 200:
- return True
- else:
- logger.warning(f"Response from API: {response.text}")
- logger.warning("Invalid API key. Unable to validate.")
- return False
-
- def get(self):
- return self.api_key
-
- def __str__(self):
- return self.api_key
diff --git a/embedchain/embedchain/config/__init__.py b/embedchain/embedchain/config/__init__.py
deleted file mode 100644
index 768408b78..000000000
--- a/embedchain/embedchain/config/__init__.py
+++ /dev/null
@@ -1,15 +0,0 @@
-# flake8: noqa: F401
-
-from .add_config import AddConfig, ChunkerConfig
-from .app_config import AppConfig
-from .base_config import BaseConfig
-from .cache_config import CacheConfig
-from .embedder.base import BaseEmbedderConfig
-from .embedder.base import BaseEmbedderConfig as EmbedderConfig
-from .embedder.ollama import OllamaEmbedderConfig
-from .llm.base import BaseLlmConfig
-from .mem0_config import Mem0Config
-from .vector_db.chroma import ChromaDbConfig
-from .vector_db.elasticsearch import ElasticsearchDBConfig
-from .vector_db.opensearch import OpenSearchDBConfig
-from .vector_db.zilliz import ZillizDBConfig
diff --git a/embedchain/embedchain/config/add_config.py b/embedchain/embedchain/config/add_config.py
deleted file mode 100644
index 56686e8ec..000000000
--- a/embedchain/embedchain/config/add_config.py
+++ /dev/null
@@ -1,79 +0,0 @@
-import builtins
-import logging
-from collections.abc import Callable
-from importlib import import_module
-from typing import Optional
-
-from embedchain.config.base_config import BaseConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class ChunkerConfig(BaseConfig):
- """
- Config for the chunker used in `add` method
- """
-
- def __init__(
- self,
- chunk_size: Optional[int] = 2000,
- chunk_overlap: Optional[int] = 0,
- length_function: Optional[Callable[[str], int]] = None,
- min_chunk_size: Optional[int] = 0,
- ):
- self.chunk_size = chunk_size
- self.chunk_overlap = chunk_overlap
- self.min_chunk_size = min_chunk_size
- if self.min_chunk_size >= self.chunk_size:
- raise ValueError(f"min_chunk_size {min_chunk_size} should be less than chunk_size {chunk_size}")
- if self.min_chunk_size < self.chunk_overlap:
- logging.warning(
- f"min_chunk_size {min_chunk_size} should be greater than chunk_overlap {chunk_overlap}, otherwise it is redundant." # noqa:E501
- )
-
- if isinstance(length_function, str):
- self.length_function = self.load_func(length_function)
- else:
- self.length_function = length_function if length_function else len
-
- @staticmethod
- def load_func(dotpath: str):
- if "." not in dotpath:
- return getattr(builtins, dotpath)
- else:
- module_, func = dotpath.rsplit(".", maxsplit=1)
- m = import_module(module_)
- return getattr(m, func)
-
-
-@register_deserializable
-class LoaderConfig(BaseConfig):
- """
- Config for the loader used in `add` method
- """
-
- def __init__(self):
- pass
-
-
-@register_deserializable
-class AddConfig(BaseConfig):
- """
- Config for the `add` method.
- """
-
- def __init__(
- self,
- chunker: Optional[ChunkerConfig] = None,
- loader: Optional[LoaderConfig] = None,
- ):
- """
- Initializes a configuration class instance for the `add` method.
-
- :param chunker: Chunker config, defaults to None
- :type chunker: Optional[ChunkerConfig], optional
- :param loader: Loader config, defaults to None
- :type loader: Optional[LoaderConfig], optional
- """
- self.loader = loader
- self.chunker = chunker
diff --git a/embedchain/embedchain/config/app_config.py b/embedchain/embedchain/config/app_config.py
deleted file mode 100644
index f3b571b7f..000000000
--- a/embedchain/embedchain/config/app_config.py
+++ /dev/null
@@ -1,34 +0,0 @@
-from typing import Optional
-
-from embedchain.helpers.json_serializable import register_deserializable
-
-from .base_app_config import BaseAppConfig
-
-
-@register_deserializable
-class AppConfig(BaseAppConfig):
- """
- Config to initialize an embedchain custom `App` instance, with extra config options.
- """
-
- def __init__(
- self,
- log_level: str = "WARNING",
- id: Optional[str] = None,
- name: Optional[str] = None,
- collect_metrics: Optional[bool] = True,
- **kwargs,
- ):
- """
- Initializes a configuration class instance for an App. This is the simplest form of an embedchain app.
- Most of the configuration is done in the `App` class itself.
-
- :param log_level: Debug level ['DEBUG', 'INFO', 'WARNING', 'ERROR', 'CRITICAL'], defaults to "WARNING"
- :type log_level: str, optional
- :param id: ID of the app. Document metadata will have this id., defaults to None
- :type id: Optional[str], optional
- :param collect_metrics: Send anonymous telemetry to improve embedchain, defaults to True
- :type collect_metrics: Optional[bool], optional
- """
- self.name = name
- super().__init__(log_level=log_level, id=id, collect_metrics=collect_metrics, **kwargs)
diff --git a/embedchain/embedchain/config/base_app_config.py b/embedchain/embedchain/config/base_app_config.py
deleted file mode 100644
index 781ca024a..000000000
--- a/embedchain/embedchain/config/base_app_config.py
+++ /dev/null
@@ -1,58 +0,0 @@
-import logging
-from typing import Optional
-
-from embedchain.config.base_config import BaseConfig
-from embedchain.helpers.json_serializable import JSONSerializable
-from embedchain.vectordb.base import BaseVectorDB
-
-logger = logging.getLogger(__name__)
-
-
-class BaseAppConfig(BaseConfig, JSONSerializable):
- """
- Parent config to initialize an instance of `App`.
- """
-
- def __init__(
- self,
- log_level: str = "WARNING",
- db: Optional[BaseVectorDB] = None,
- id: Optional[str] = None,
- collect_metrics: bool = True,
- collection_name: Optional[str] = None,
- ):
- """
- Initializes a configuration class instance for an App.
- Most of the configuration is done in the `App` class itself.
-
- :param log_level: Debug level ['DEBUG', 'INFO', 'WARNING', 'ERROR', 'CRITICAL'], defaults to "WARNING"
- :type log_level: str, optional
- :param db: A database class. It is recommended to set this directly in the `App` class, not this config,
- defaults to None
- :type db: Optional[BaseVectorDB], optional
- :param id: ID of the app. Document metadata will have this id., defaults to None
- :type id: Optional[str], optional
- :param collect_metrics: Send anonymous telemetry to improve embedchain, defaults to True
- :type collect_metrics: Optional[bool], optional
- :param collection_name: Default collection name. It's recommended to use app.db.set_collection_name() instead,
- defaults to None
- :type collection_name: Optional[str], optional
- """
- self.id = id
- self.collect_metrics = True if (collect_metrics is True or collect_metrics is None) else False
- self.collection_name = collection_name
-
- if db:
- self._db = db
- logger.warning(
- "DEPRECATION WARNING: Please supply the database as the second parameter during app init. "
- "Such as `app(config=config, db=db)`."
- )
-
- if collection_name:
- logger.warning("DEPRECATION WARNING: Please supply the collection name to the database config.")
- return
-
- def _setup_logging(self, log_level):
- logger.basicConfig(format="%(asctime)s [%(name)s] [%(levelname)s] %(message)s", level=log_level)
- self.logger = logger.getLogger(__name__)
diff --git a/embedchain/embedchain/config/base_config.py b/embedchain/embedchain/config/base_config.py
deleted file mode 100644
index bf7869f41..000000000
--- a/embedchain/embedchain/config/base_config.py
+++ /dev/null
@@ -1,21 +0,0 @@
-from typing import Any
-
-from embedchain.helpers.json_serializable import JSONSerializable
-
-
-class BaseConfig(JSONSerializable):
- """
- Base config.
- """
-
- def __init__(self):
- """Initializes a configuration class for a class."""
- pass
-
- def as_dict(self) -> dict[str, Any]:
- """Return config object as a dict
-
- :return: config object as dict
- :rtype: dict[str, Any]
- """
- return vars(self)
diff --git a/embedchain/embedchain/config/cache_config.py b/embedchain/embedchain/config/cache_config.py
deleted file mode 100644
index ef8bd1fb3..000000000
--- a/embedchain/embedchain/config/cache_config.py
+++ /dev/null
@@ -1,96 +0,0 @@
-from typing import Any, Optional
-
-from embedchain.config.base_config import BaseConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class CacheSimilarityEvalConfig(BaseConfig):
- """
- This is the evaluator to compare two embeddings according to their distance computed in embedding retrieval stage.
- In the retrieval stage, `search_result` is the distance used for approximate nearest neighbor search and have been
- put into `cache_dict`. `max_distance` is used to bound this distance to make it between [0-`max_distance`].
- `positive` is used to indicate this distance is directly proportional to the similarity of two entities.
- If `positive` is set `False`, `max_distance` will be used to subtract this distance to get the final score.
-
- :param max_distance: the bound of maximum distance.
- :type max_distance: float
- :param positive: if the larger distance indicates more similar of two entities, It is True. Otherwise, it is False.
- :type positive: bool
- """
-
- def __init__(
- self,
- strategy: Optional[str] = "distance",
- max_distance: Optional[float] = 1.0,
- positive: Optional[bool] = False,
- ):
- self.strategy = strategy
- self.max_distance = max_distance
- self.positive = positive
-
- @staticmethod
- def from_config(config: Optional[dict[str, Any]]):
- if config is None:
- return CacheSimilarityEvalConfig()
- else:
- return CacheSimilarityEvalConfig(
- strategy=config.get("strategy", "distance"),
- max_distance=config.get("max_distance", 1.0),
- positive=config.get("positive", False),
- )
-
-
-@register_deserializable
-class CacheInitConfig(BaseConfig):
- """
- This is a cache init config. Used to initialize a cache.
-
- :param similarity_threshold: a threshold ranged from 0 to 1 to filter search results with similarity score higher \
- than the threshold. When it is 0, there is no hits. When it is 1, all search results will be returned as hits.
- :type similarity_threshold: float
- :param auto_flush: it will be automatically flushed every time xx pieces of data are added, default to 20
- :type auto_flush: int
- """
-
- def __init__(
- self,
- similarity_threshold: Optional[float] = 0.8,
- auto_flush: Optional[int] = 20,
- ):
- if similarity_threshold < 0 or similarity_threshold > 1:
- raise ValueError(f"similarity_threshold {similarity_threshold} should be between 0 and 1")
-
- self.similarity_threshold = similarity_threshold
- self.auto_flush = auto_flush
-
- @staticmethod
- def from_config(config: Optional[dict[str, Any]]):
- if config is None:
- return CacheInitConfig()
- else:
- return CacheInitConfig(
- similarity_threshold=config.get("similarity_threshold", 0.8),
- auto_flush=config.get("auto_flush", 20),
- )
-
-
-@register_deserializable
-class CacheConfig(BaseConfig):
- def __init__(
- self,
- similarity_eval_config: Optional[CacheSimilarityEvalConfig] = CacheSimilarityEvalConfig(),
- init_config: Optional[CacheInitConfig] = CacheInitConfig(),
- ):
- self.similarity_eval_config = similarity_eval_config
- self.init_config = init_config
-
- @staticmethod
- def from_config(config: Optional[dict[str, Any]]):
- if config is None:
- return CacheConfig()
- else:
- return CacheConfig(
- similarity_eval_config=CacheSimilarityEvalConfig.from_config(config.get("similarity_evaluation", {})),
- init_config=CacheInitConfig.from_config(config.get("init_config", {})),
- )
diff --git a/embedchain/embedchain/config/embedder/__init__.py b/embedchain/embedchain/config/embedder/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/config/embedder/aws_bedrock.py b/embedchain/embedchain/config/embedder/aws_bedrock.py
deleted file mode 100644
index f0bd0c538..000000000
--- a/embedchain/embedchain/config/embedder/aws_bedrock.py
+++ /dev/null
@@ -1,21 +0,0 @@
-from typing import Any, Dict, Optional
-
-from embedchain.config.embedder.base import BaseEmbedderConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class AWSBedrockEmbedderConfig(BaseEmbedderConfig):
- def __init__(
- self,
- model: Optional[str] = None,
- deployment_name: Optional[str] = None,
- vector_dimension: Optional[int] = None,
- task_type: Optional[str] = None,
- title: Optional[str] = None,
- model_kwargs: Optional[Dict[str, Any]] = None,
- ):
- super().__init__(model, deployment_name, vector_dimension)
- self.task_type = task_type or "retrieval_document"
- self.title = title or "Embeddings for Embedchain"
- self.model_kwargs = model_kwargs or {}
diff --git a/embedchain/embedchain/config/embedder/base.py b/embedchain/embedchain/config/embedder/base.py
deleted file mode 100644
index 56c4070d0..000000000
--- a/embedchain/embedchain/config/embedder/base.py
+++ /dev/null
@@ -1,55 +0,0 @@
-from typing import Any, Dict, Optional, Union
-
-import httpx
-
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class BaseEmbedderConfig:
- def __init__(
- self,
- model: Optional[str] = None,
- deployment_name: Optional[str] = None,
- vector_dimension: Optional[int] = None,
- endpoint: Optional[str] = None,
- api_key: Optional[str] = None,
- api_base: Optional[str] = None,
- model_kwargs: Optional[Dict[str, Any]] = None,
- http_client_proxies: Optional[Union[Dict, str]] = None,
- http_async_client_proxies: Optional[Union[Dict, str]] = None,
- ):
- """
- Initialize a new instance of an embedder config class.
-
- :param model: model name of the llm embedding model (not applicable to all providers), defaults to None
- :type model: Optional[str], optional
- :param deployment_name: deployment name for llm embedding model, defaults to None
- :type deployment_name: Optional[str], optional
- :param vector_dimension: vector dimension of the embedding model, defaults to None
- :type vector_dimension: Optional[int], optional
- :param endpoint: endpoint for the embedding model, defaults to None
- :type endpoint: Optional[str], optional
- :param api_key: hugginface api key, defaults to None
- :type api_key: Optional[str], optional
- :param api_base: huggingface api base, defaults to None
- :type api_base: Optional[str], optional
- :param model_kwargs: key-value arguments for the embedding model, defaults a dict inside init.
- :type model_kwargs: Optional[Dict[str, Any]], defaults a dict inside init.
- :param http_client_proxies: The proxy server settings used to create self.http_client, defaults to None
- :type http_client_proxies: Optional[Dict | str], optional
- :param http_async_client_proxies: The proxy server settings for async calls used to create
- self.http_async_client, defaults to None
- :type http_async_client_proxies: Optional[Dict | str], optional
- """
- self.model = model
- self.deployment_name = deployment_name
- self.vector_dimension = vector_dimension
- self.endpoint = endpoint
- self.api_key = api_key
- self.api_base = api_base
- self.model_kwargs = model_kwargs or {}
- self.http_client = httpx.Client(proxies=http_client_proxies) if http_client_proxies else None
- self.http_async_client = (
- httpx.AsyncClient(proxies=http_async_client_proxies) if http_async_client_proxies else None
- )
diff --git a/embedchain/embedchain/config/embedder/google.py b/embedchain/embedchain/config/embedder/google.py
deleted file mode 100644
index 7cf5a9011..000000000
--- a/embedchain/embedchain/config/embedder/google.py
+++ /dev/null
@@ -1,19 +0,0 @@
-from typing import Optional
-
-from embedchain.config.embedder.base import BaseEmbedderConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class GoogleAIEmbedderConfig(BaseEmbedderConfig):
- def __init__(
- self,
- model: Optional[str] = None,
- deployment_name: Optional[str] = None,
- vector_dimension: Optional[int] = None,
- task_type: Optional[str] = None,
- title: Optional[str] = None,
- ):
- super().__init__(model, deployment_name, vector_dimension)
- self.task_type = task_type or "retrieval_document"
- self.title = title or "Embeddings for Embedchain"
diff --git a/embedchain/embedchain/config/embedder/ollama.py b/embedchain/embedchain/config/embedder/ollama.py
deleted file mode 100644
index f680328f9..000000000
--- a/embedchain/embedchain/config/embedder/ollama.py
+++ /dev/null
@@ -1,16 +0,0 @@
-from typing import Optional
-
-from embedchain.config.embedder.base import BaseEmbedderConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class OllamaEmbedderConfig(BaseEmbedderConfig):
- def __init__(
- self,
- model: Optional[str] = None,
- base_url: Optional[str] = None,
- vector_dimension: Optional[int] = None,
- ):
- super().__init__(model=model, vector_dimension=vector_dimension)
- self.base_url = base_url or "http://localhost:11434"
diff --git a/embedchain/embedchain/config/evaluation/__init__.py b/embedchain/embedchain/config/evaluation/__init__.py
deleted file mode 100644
index 67e78dade..000000000
--- a/embedchain/embedchain/config/evaluation/__init__.py
+++ /dev/null
@@ -1,5 +0,0 @@
-from .base import ( # noqa: F401
- AnswerRelevanceConfig,
- ContextRelevanceConfig,
- GroundednessConfig,
-)
diff --git a/embedchain/embedchain/config/evaluation/base.py b/embedchain/embedchain/config/evaluation/base.py
deleted file mode 100644
index 5c44d3f83..000000000
--- a/embedchain/embedchain/config/evaluation/base.py
+++ /dev/null
@@ -1,92 +0,0 @@
-from typing import Optional
-
-from embedchain.config.base_config import BaseConfig
-
-ANSWER_RELEVANCY_PROMPT = """
-Please provide $num_gen_questions questions from the provided answer.
-You must provide the complete question, if are not able to provide the complete question, return empty string ("").
-Please only provide one question per line without numbers or bullets to distinguish them.
-You must only provide the questions and no other text.
-
-$answer
-""" # noqa:E501
-
-
-CONTEXT_RELEVANCY_PROMPT = """
-Please extract relevant sentences from the provided context that is required to answer the given question.
-If no relevant sentences are found, or if you believe the question cannot be answered from the given context, return the empty string ("").
-While extracting candidate sentences you're not allowed to make any changes to sentences from given context or make up any sentences.
-You must only provide sentences from the given context and nothing else.
-
-Context: $context
-Question: $question
-""" # noqa:E501
-
-GROUNDEDNESS_ANSWER_CLAIMS_PROMPT = """
-Please provide one or more statements from each sentence of the provided answer.
-You must provide the symantically equivalent statements for each sentence of the answer.
-You must provide the complete statement, if are not able to provide the complete statement, return empty string ("").
-Please only provide one statement per line WITHOUT numbers or bullets.
-If the question provided is not being answered in the provided answer, return empty string ("").
-You must only provide the statements and no other text.
-
-$question
-$answer
-""" # noqa:E501
-
-GROUNDEDNESS_CLAIMS_INFERENCE_PROMPT = """
-Given the context and the provided claim statements, please provide a verdict for each claim statement whether it can be completely inferred from the given context or not.
-Use only "1" (yes), "0" (no) and "-1" (null) for "yes", "no" or "null" respectively.
-You must provide one verdict per line, ONLY WITH "1", "0" or "-1" as per your verdict to the given statement and nothing else.
-You must provide the verdicts in the same order as the claim statements.
-
-Contexts:
-$context
-
-Claim statements:
-$claim_statements
-""" # noqa:E501
-
-
-class GroundednessConfig(BaseConfig):
- def __init__(
- self,
- model: str = "gpt-4",
- api_key: Optional[str] = None,
- answer_claims_prompt: str = GROUNDEDNESS_ANSWER_CLAIMS_PROMPT,
- claims_inference_prompt: str = GROUNDEDNESS_CLAIMS_INFERENCE_PROMPT,
- ):
- self.model = model
- self.api_key = api_key
- self.answer_claims_prompt = answer_claims_prompt
- self.claims_inference_prompt = claims_inference_prompt
-
-
-class AnswerRelevanceConfig(BaseConfig):
- def __init__(
- self,
- model: str = "gpt-4",
- embedder: str = "text-embedding-ada-002",
- api_key: Optional[str] = None,
- num_gen_questions: int = 1,
- prompt: str = ANSWER_RELEVANCY_PROMPT,
- ):
- self.model = model
- self.embedder = embedder
- self.api_key = api_key
- self.num_gen_questions = num_gen_questions
- self.prompt = prompt
-
-
-class ContextRelevanceConfig(BaseConfig):
- def __init__(
- self,
- model: str = "gpt-4",
- api_key: Optional[str] = None,
- language: str = "en",
- prompt: str = CONTEXT_RELEVANCY_PROMPT,
- ):
- self.model = model
- self.api_key = api_key
- self.language = language
- self.prompt = prompt
diff --git a/embedchain/embedchain/config/llm/__init__.py b/embedchain/embedchain/config/llm/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/config/llm/base.py b/embedchain/embedchain/config/llm/base.py
deleted file mode 100644
index 693d09c5b..000000000
--- a/embedchain/embedchain/config/llm/base.py
+++ /dev/null
@@ -1,276 +0,0 @@
-import json
-import logging
-import re
-from pathlib import Path
-from string import Template
-from typing import Any, Dict, Mapping, Optional, Union
-
-import httpx
-
-from embedchain.config.base_config import BaseConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-logger = logging.getLogger(__name__)
-
-DEFAULT_PROMPT = """
-You are a Q&A expert system. Your responses must always be rooted in the context provided for each query. Here are some guidelines to follow:
-
-1. Refrain from explicitly mentioning the context provided in your response.
-2. The context should silently guide your answers without being directly acknowledged.
-3. Do not use phrases such as 'According to the context provided', 'Based on the context, ...' etc.
-
-Context information:
-----------------------
-$context
-----------------------
-
-Query: $query
-Answer:
-""" # noqa:E501
-
-DEFAULT_PROMPT_WITH_HISTORY = """
-You are a Q&A expert system. Your responses must always be rooted in the context provided for each query. You are also provided with the conversation history with the user. Make sure to use relevant context from conversation history as needed.
-
-Here are some guidelines to follow:
-
-1. Refrain from explicitly mentioning the context provided in your response.
-2. The context should silently guide your answers without being directly acknowledged.
-3. Do not use phrases such as 'According to the context provided', 'Based on the context, ...' etc.
-
-Context information:
-----------------------
-$context
-----------------------
-
-Conversation history:
-----------------------
-$history
-----------------------
-
-Query: $query
-Answer:
-""" # noqa:E501
-
-DEFAULT_PROMPT_WITH_MEM0_MEMORY = """
-You are an expert at answering questions based on provided memories. You are also provided with the context and conversation history of the user. Make sure to use relevant context from conversation history and context as needed.
-
-Here are some guidelines to follow:
-1. Refrain from explicitly mentioning the context provided in your response.
-2. Take into consideration the conversation history and context provided.
-3. Do not use phrases such as 'According to the context provided', 'Based on the context, ...' etc.
-
-Striclty return the query exactly as it is if it is not a question or if no relevant information is found.
-
-Context information:
-----------------------
-$context
-----------------------
-
-Conversation history:
-----------------------
-$history
-----------------------
-
-Memories/Preferences:
-----------------------
-$memories
-----------------------
-
-Query: $query
-Answer:
-""" # noqa:E501
-
-DOCS_SITE_DEFAULT_PROMPT = """
-You are an expert AI assistant for developer support product. Your responses must always be rooted in the context provided for each query. Wherever possible, give complete code snippet. Dont make up any code snippet on your own.
-
-Here are some guidelines to follow:
-
-1. Refrain from explicitly mentioning the context provided in your response.
-2. The context should silently guide your answers without being directly acknowledged.
-3. Do not use phrases such as 'According to the context provided', 'Based on the context, ...' etc.
-
-Context information:
-----------------------
-$context
-----------------------
-
-Query: $query
-Answer:
-""" # noqa:E501
-
-DEFAULT_PROMPT_TEMPLATE = Template(DEFAULT_PROMPT)
-DEFAULT_PROMPT_WITH_HISTORY_TEMPLATE = Template(DEFAULT_PROMPT_WITH_HISTORY)
-DEFAULT_PROMPT_WITH_MEM0_MEMORY_TEMPLATE = Template(DEFAULT_PROMPT_WITH_MEM0_MEMORY)
-DOCS_SITE_PROMPT_TEMPLATE = Template(DOCS_SITE_DEFAULT_PROMPT)
-query_re = re.compile(r"\$\{*query\}*")
-context_re = re.compile(r"\$\{*context\}*")
-history_re = re.compile(r"\$\{*history\}*")
-
-
-@register_deserializable
-class BaseLlmConfig(BaseConfig):
- """
- Config for the `query` method.
- """
-
- def __init__(
- self,
- number_documents: int = 3,
- template: Optional[Template] = None,
- prompt: Optional[Template] = None,
- model: Optional[str] = None,
- temperature: float = 0,
- max_tokens: int = 1000,
- top_p: float = 1,
- stream: bool = False,
- online: bool = False,
- token_usage: bool = False,
- deployment_name: Optional[str] = None,
- system_prompt: Optional[str] = None,
- where: dict[str, Any] = None,
- query_type: Optional[str] = None,
- callbacks: Optional[list] = None,
- api_key: Optional[str] = None,
- base_url: Optional[str] = None,
- endpoint: Optional[str] = None,
- model_kwargs: Optional[dict[str, Any]] = None,
- http_client_proxies: Optional[Union[Dict, str]] = None,
- http_async_client_proxies: Optional[Union[Dict, str]] = None,
- local: Optional[bool] = False,
- default_headers: Optional[Mapping[str, str]] = None,
- api_version: Optional[str] = None,
- ):
- """
- Initializes a configuration class instance for the LLM.
-
- Takes the place of the former `QueryConfig` or `ChatConfig`.
-
- :param number_documents: Number of documents to pull from the database as
- context, defaults to 1
- :type number_documents: int, optional
- :param template: The `Template` instance to use as a template for
- prompt, defaults to None (deprecated)
- :type template: Optional[Template], optional
- :param prompt: The `Template` instance to use as a template for
- prompt, defaults to None
- :type prompt: Optional[Template], optional
- :param model: Controls the OpenAI model used, defaults to None
- :type model: Optional[str], optional
- :param temperature: Controls the randomness of the model's output.
- Higher values (closer to 1) make output more random, lower values make it more deterministic, defaults to 0
- :type temperature: float, optional
- :param max_tokens: Controls how many tokens are generated, defaults to 1000
- :type max_tokens: int, optional
- :param top_p: Controls the diversity of words. Higher values (closer to 1) make word selection more diverse,
- defaults to 1
- :type top_p: float, optional
- :param stream: Control if response is streamed back to user, defaults to False
- :type stream: bool, optional
- :param online: Controls whether to use internet for answering query, defaults to False
- :type online: bool, optional
- :param token_usage: Controls whether to return token usage in response, defaults to False
- :type token_usage: bool, optional
- :param deployment_name: t.b.a., defaults to None
- :type deployment_name: Optional[str], optional
- :param system_prompt: System prompt string, defaults to None
- :type system_prompt: Optional[str], optional
- :param where: A dictionary of key-value pairs to filter the database results., defaults to None
- :type where: dict[str, Any], optional
- :param api_key: The api key of the custom endpoint, defaults to None
- :type api_key: Optional[str], optional
- :param endpoint: The api url of the custom endpoint, defaults to None
- :type endpoint: Optional[str], optional
- :param model_kwargs: A dictionary of key-value pairs to pass to the model, defaults to None
- :type model_kwargs: Optional[Dict[str, Any]], optional
- :param callbacks: Langchain callback functions to use, defaults to None
- :type callbacks: Optional[list], optional
- :param query_type: The type of query to use, defaults to None
- :type query_type: Optional[str], optional
- :param http_client_proxies: The proxy server settings used to create self.http_client, defaults to None
- :type http_client_proxies: Optional[Dict | str], optional
- :param http_async_client_proxies: The proxy server settings for async calls used to create
- self.http_async_client, defaults to None
- :type http_async_client_proxies: Optional[Dict | str], optional
- :param local: If True, the model will be run locally, defaults to False (for huggingface provider)
- :type local: Optional[bool], optional
- :param default_headers: Set additional HTTP headers to be sent with requests to OpenAI
- :type default_headers: Optional[Mapping[str, str]], optional
- :raises ValueError: If the template is not valid as template should
- contain $context and $query (and optionally $history)
- :raises ValueError: Stream is not boolean
- """
- if template is not None:
- logger.warning(
- "The `template` argument is deprecated and will be removed in a future version. "
- + "Please use `prompt` instead."
- )
- if prompt is None:
- prompt = template
-
- if prompt is None:
- prompt = DEFAULT_PROMPT_TEMPLATE
-
- self.number_documents = number_documents
- self.temperature = temperature
- self.max_tokens = max_tokens
- self.model = model
- self.top_p = top_p
- self.online = online
- self.token_usage = token_usage
- self.deployment_name = deployment_name
- self.system_prompt = system_prompt
- self.query_type = query_type
- self.callbacks = callbacks
- self.api_key = api_key
- self.base_url = base_url
- self.endpoint = endpoint
- self.model_kwargs = model_kwargs
- self.http_client = httpx.Client(proxies=http_client_proxies) if http_client_proxies else None
- self.http_async_client = (
- httpx.AsyncClient(proxies=http_async_client_proxies) if http_async_client_proxies else None
- )
- self.local = local
- self.default_headers = default_headers
- self.online = online
- self.api_version = api_version
-
- if token_usage:
- f = Path(__file__).resolve().parent.parent / "model_prices_and_context_window.json"
- self.model_pricing_map = json.load(f.open())
-
- if isinstance(prompt, str):
- prompt = Template(prompt)
-
- if self.validate_prompt(prompt):
- self.prompt = prompt
- else:
- raise ValueError("The 'prompt' should have 'query' and 'context' keys and potentially 'history' (if used).")
-
- if not isinstance(stream, bool):
- raise ValueError("`stream` should be bool")
- self.stream = stream
- self.where = where
-
- @staticmethod
- def validate_prompt(prompt: Template) -> Optional[re.Match[str]]:
- """
- validate the prompt
-
- :param prompt: the prompt to validate
- :type prompt: Template
- :return: valid (true) or invalid (false)
- :rtype: Optional[re.Match[str]]
- """
- return re.search(query_re, prompt.template) and re.search(context_re, prompt.template)
-
- @staticmethod
- def _validate_prompt_history(prompt: Template) -> Optional[re.Match[str]]:
- """
- validate the prompt with history
-
- :param prompt: the prompt to validate
- :type prompt: Template
- :return: valid (true) or invalid (false)
- :rtype: Optional[re.Match[str]]
- """
- return re.search(history_re, prompt.template)
diff --git a/embedchain/embedchain/config/mem0_config.py b/embedchain/embedchain/config/mem0_config.py
deleted file mode 100644
index 924ba8744..000000000
--- a/embedchain/embedchain/config/mem0_config.py
+++ /dev/null
@@ -1,21 +0,0 @@
-from typing import Any, Optional
-
-from embedchain.config.base_config import BaseConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class Mem0Config(BaseConfig):
- def __init__(self, api_key: str, top_k: Optional[int] = 10):
- self.api_key = api_key
- self.top_k = top_k
-
- @staticmethod
- def from_config(config: Optional[dict[str, Any]]):
- if config is None:
- return Mem0Config()
- else:
- return Mem0Config(
- api_key=config.get("api_key", ""),
- init_config=config.get("top_k", 10),
- )
diff --git a/embedchain/embedchain/config/model_prices_and_context_window.json b/embedchain/embedchain/config/model_prices_and_context_window.json
deleted file mode 100644
index c68f90394..000000000
--- a/embedchain/embedchain/config/model_prices_and_context_window.json
+++ /dev/null
@@ -1,824 +0,0 @@
-{
- "openai/gpt-4": {
- "max_tokens": 4096,
- "max_input_tokens": 8192,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00003,
- "output_cost_per_token": 0.00006
- },
- "openai/gpt-4o": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000005,
- "output_cost_per_token": 0.000015
- },
- "openai/gpt-4o-mini": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00000015,
- "output_cost_per_token": 0.00000060
- },
- "openai/gpt-4o-mini-2024-07-18": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00000015,
- "output_cost_per_token": 0.00000060
- },
- "openai/gpt-4o-2024-05-13": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000005,
- "output_cost_per_token": 0.000015
- },
- "openai/gpt-4-turbo-preview": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "openai/gpt-4-0314": {
- "max_tokens": 4096,
- "max_input_tokens": 8192,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00003,
- "output_cost_per_token": 0.00006
- },
- "openai/gpt-4-0613": {
- "max_tokens": 4096,
- "max_input_tokens": 8192,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00003,
- "output_cost_per_token": 0.00006
- },
- "openai/gpt-4-32k": {
- "max_tokens": 4096,
- "max_input_tokens": 32768,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00006,
- "output_cost_per_token": 0.00012
- },
- "openai/gpt-4-32k-0314": {
- "max_tokens": 4096,
- "max_input_tokens": 32768,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00006,
- "output_cost_per_token": 0.00012
- },
- "openai/gpt-4-32k-0613": {
- "max_tokens": 4096,
- "max_input_tokens": 32768,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00006,
- "output_cost_per_token": 0.00012
- },
- "openai/gpt-4-turbo": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "openai/gpt-4-turbo-2024-04-09": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "openai/gpt-4-1106-preview": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "openai/gpt-4-0125-preview": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "openai/gpt-3.5-turbo": {
- "max_tokens": 4097,
- "max_input_tokens": 16385,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.0000015,
- "output_cost_per_token": 0.000002
- },
- "openai/gpt-3.5-turbo-0301": {
- "max_tokens": 4097,
- "max_input_tokens": 4097,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.0000015,
- "output_cost_per_token": 0.000002
- },
- "openai/gpt-3.5-turbo-0613": {
- "input_cost_per_token": 0.0000015,
- "output_cost_per_token": 0.000002
- },
- "openai/gpt-3.5-turbo-1106": {
- "max_tokens": 16385,
- "max_input_tokens": 16385,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.0000010,
- "output_cost_per_token": 0.0000020
- },
- "openai/gpt-3.5-turbo-0125": {
- "max_tokens": 16385,
- "max_input_tokens": 16385,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.0000005,
- "output_cost_per_token": 0.0000015
- },
- "openai/gpt-3.5-turbo-16k": {
- "max_tokens": 16385,
- "max_input_tokens": 16385,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000003,
- "output_cost_per_token": 0.000004
- },
- "openai/gpt-3.5-turbo-16k-0613": {
- "max_tokens": 16385,
- "max_input_tokens": 16385,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000003,
- "output_cost_per_token": 0.000004
- },
- "openai/text-embedding-3-large": {
- "max_tokens": 8191,
- "max_input_tokens": 8191,
- "output_vector_size": 3072,
- "input_cost_per_token": 0.00000013,
- "output_cost_per_token": 0.000000
- },
- "openai/text-embedding-3-small": {
- "max_tokens": 8191,
- "max_input_tokens": 8191,
- "output_vector_size": 1536,
- "input_cost_per_token": 0.00000002,
- "output_cost_per_token": 0.000000
- },
- "openai/text-embedding-ada-002": {
- "max_tokens": 8191,
- "max_input_tokens": 8191,
- "output_vector_size": 1536,
- "input_cost_per_token": 0.0000001,
- "output_cost_per_token": 0.000000
- },
- "openai/text-embedding-ada-002-v2": {
- "max_tokens": 8191,
- "max_input_tokens": 8191,
- "input_cost_per_token": 0.0000001,
- "output_cost_per_token": 0.000000
- },
- "openai/babbage-002": {
- "max_tokens": 16384,
- "max_input_tokens": 16384,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.0000004,
- "output_cost_per_token": 0.0000004
- },
- "openai/davinci-002": {
- "max_tokens": 16384,
- "max_input_tokens": 16384,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000002,
- "output_cost_per_token": 0.000002
- },
- "openai/gpt-3.5-turbo-instruct": {
- "max_tokens": 4096,
- "max_input_tokens": 8192,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.0000015,
- "output_cost_per_token": 0.000002
- },
- "openai/gpt-3.5-turbo-instruct-0914": {
- "max_tokens": 4097,
- "max_input_tokens": 8192,
- "max_output_tokens": 4097,
- "input_cost_per_token": 0.0000015,
- "output_cost_per_token": 0.000002
- },
- "azure/gpt-4o": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000005,
- "output_cost_per_token": 0.000015
- },
- "azure/gpt-4o-mini": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00000015,
- "output_cost_per_token": 0.00000060
- },
- "azure/gpt-4-turbo-2024-04-09": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "azure/gpt-4-0125-preview": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "azure/gpt-4-1106-preview": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "azure/gpt-4-0613": {
- "max_tokens": 4096,
- "max_input_tokens": 8192,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00003,
- "output_cost_per_token": 0.00006
- },
- "azure/gpt-4-32k-0613": {
- "max_tokens": 4096,
- "max_input_tokens": 32768,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00006,
- "output_cost_per_token": 0.00012
- },
- "azure/gpt-4-32k": {
- "max_tokens": 4096,
- "max_input_tokens": 32768,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00006,
- "output_cost_per_token": 0.00012
- },
- "azure/gpt-4": {
- "max_tokens": 4096,
- "max_input_tokens": 8192,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00003,
- "output_cost_per_token": 0.00006
- },
- "azure/gpt-4-turbo": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "azure/gpt-4-turbo-vision-preview": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00001,
- "output_cost_per_token": 0.00003
- },
- "azure/gpt-3.5-turbo-16k-0613": {
- "max_tokens": 4096,
- "max_input_tokens": 16385,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000003,
- "output_cost_per_token": 0.000004
- },
- "azure/gpt-3.5-turbo-1106": {
- "max_tokens": 4096,
- "max_input_tokens": 16384,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.0000015,
- "output_cost_per_token": 0.000002
- },
- "azure/gpt-3.5-turbo-0125": {
- "max_tokens": 4096,
- "max_input_tokens": 16384,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.0000005,
- "output_cost_per_token": 0.0000015
- },
- "azure/gpt-3.5-turbo-16k": {
- "max_tokens": 4096,
- "max_input_tokens": 16385,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000003,
- "output_cost_per_token": 0.000004
- },
- "azure/gpt-3.5-turbo": {
- "max_tokens": 4096,
- "max_input_tokens": 4097,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.0000005,
- "output_cost_per_token": 0.0000015
- },
- "azure/gpt-3.5-turbo-instruct-0914": {
- "max_tokens": 4097,
- "max_input_tokens": 4097,
- "input_cost_per_token": 0.0000015,
- "output_cost_per_token": 0.000002
- },
- "azure/gpt-3.5-turbo-instruct": {
- "max_tokens": 4097,
- "max_input_tokens": 4097,
- "input_cost_per_token": 0.0000015,
- "output_cost_per_token": 0.000002
- },
- "azure/text-embedding-ada-002": {
- "max_tokens": 8191,
- "max_input_tokens": 8191,
- "input_cost_per_token": 0.0000001,
- "output_cost_per_token": 0.000000
- },
- "azure/text-embedding-3-large": {
- "max_tokens": 8191,
- "max_input_tokens": 8191,
- "input_cost_per_token": 0.00000013,
- "output_cost_per_token": 0.000000
- },
- "azure/text-embedding-3-small": {
- "max_tokens": 8191,
- "max_input_tokens": 8191,
- "input_cost_per_token": 0.00000002,
- "output_cost_per_token": 0.000000
- },
- "mistralai/mistral-tiny": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.00000025
- },
- "mistralai/mistral-small": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000001,
- "output_cost_per_token": 0.000003
- },
- "mistralai/mistral-small-latest": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000001,
- "output_cost_per_token": 0.000003
- },
- "mistralai/mistral-medium": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.0000027,
- "output_cost_per_token": 0.0000081
- },
- "mistralai/mistral-medium-latest": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.0000027,
- "output_cost_per_token": 0.0000081
- },
- "mistralai/mistral-medium-2312": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.0000027,
- "output_cost_per_token": 0.0000081
- },
- "mistralai/mistral-large-latest": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000004,
- "output_cost_per_token": 0.000012
- },
- "mistralai/mistral-large-2402": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000004,
- "output_cost_per_token": 0.000012
- },
- "mistralai/open-mistral-7b": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.00000025
- },
- "mistralai/open-mixtral-8x7b": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.0000007,
- "output_cost_per_token": 0.0000007
- },
- "mistralai/open-mixtral-8x22b": {
- "max_tokens": 8191,
- "max_input_tokens": 64000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000002,
- "output_cost_per_token": 0.000006
- },
- "mistralai/codestral-latest": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000001,
- "output_cost_per_token": 0.000003
- },
- "mistralai/codestral-2405": {
- "max_tokens": 8191,
- "max_input_tokens": 32000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000001,
- "output_cost_per_token": 0.000003
- },
- "mistralai/mistral-embed": {
- "max_tokens": 8192,
- "max_input_tokens": 8192,
- "input_cost_per_token": 0.0000001,
- "output_cost_per_token": 0.0
- },
- "groq/llama2-70b-4096": {
- "max_tokens": 4096,
- "max_input_tokens": 4096,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00000070,
- "output_cost_per_token": 0.00000080
- },
- "groq/llama3-8b-8192": {
- "max_tokens": 8192,
- "max_input_tokens": 8192,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.00000010,
- "output_cost_per_token": 0.00000010
- },
- "groq/llama3-70b-8192": {
- "max_tokens": 8192,
- "max_input_tokens": 8192,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.00000064,
- "output_cost_per_token": 0.00000080
- },
- "groq/mixtral-8x7b-32768": {
- "max_tokens": 32768,
- "max_input_tokens": 32768,
- "max_output_tokens": 32768,
- "input_cost_per_token": 0.00000027,
- "output_cost_per_token": 0.00000027
- },
- "groq/gemma-7b-it": {
- "max_tokens": 8192,
- "max_input_tokens": 8192,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.00000010,
- "output_cost_per_token": 0.00000010
- },
- "anthropic/claude-instant-1": {
- "max_tokens": 8191,
- "max_input_tokens": 100000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.00000163,
- "output_cost_per_token": 0.00000551
- },
- "anthropic/claude-instant-1.2": {
- "max_tokens": 8191,
- "max_input_tokens": 100000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000000163,
- "output_cost_per_token": 0.000000551
- },
- "anthropic/claude-2": {
- "max_tokens": 8191,
- "max_input_tokens": 100000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000008,
- "output_cost_per_token": 0.000024
- },
- "anthropic/claude-2.1": {
- "max_tokens": 8191,
- "max_input_tokens": 200000,
- "max_output_tokens": 8191,
- "input_cost_per_token": 0.000008,
- "output_cost_per_token": 0.000024
- },
- "anthropic/claude-3-haiku-20240307": {
- "max_tokens": 4096,
- "max_input_tokens": 200000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.00000125
- },
- "anthropic/claude-3-opus-20240229": {
- "max_tokens": 4096,
- "max_input_tokens": 200000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000015,
- "output_cost_per_token": 0.000075
- },
- "anthropic/claude-3-sonnet-20240229": {
- "max_tokens": 4096,
- "max_input_tokens": 200000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000003,
- "output_cost_per_token": 0.000015
- },
- "vertexai/chat-bison": {
- "max_tokens": 4096,
- "max_input_tokens": 8192,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/chat-bison@001": {
- "max_tokens": 4096,
- "max_input_tokens": 8192,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/chat-bison@002": {
- "max_tokens": 4096,
- "max_input_tokens": 8192,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/chat-bison-32k": {
- "max_tokens": 8192,
- "max_input_tokens": 32000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/code-bison": {
- "max_tokens": 1024,
- "max_input_tokens": 6144,
- "max_output_tokens": 1024,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/code-bison@001": {
- "max_tokens": 1024,
- "max_input_tokens": 6144,
- "max_output_tokens": 1024,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/code-gecko@001": {
- "max_tokens": 64,
- "max_input_tokens": 2048,
- "max_output_tokens": 64,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/code-gecko@002": {
- "max_tokens": 64,
- "max_input_tokens": 2048,
- "max_output_tokens": 64,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/code-gecko": {
- "max_tokens": 64,
- "max_input_tokens": 2048,
- "max_output_tokens": 64,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/codechat-bison": {
- "max_tokens": 1024,
- "max_input_tokens": 6144,
- "max_output_tokens": 1024,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/codechat-bison@001": {
- "max_tokens": 1024,
- "max_input_tokens": 6144,
- "max_output_tokens": 1024,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/codechat-bison-32k": {
- "max_tokens": 8192,
- "max_input_tokens": 32000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.000000125,
- "output_cost_per_token": 0.000000125
- },
- "vertexai/gemini-pro": {
- "max_tokens": 8192,
- "max_input_tokens": 32760,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.0000005
- },
- "vertexai/gemini-1.0-pro": {
- "max_tokens": 8192,
- "max_input_tokens": 32760,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.0000005
- },
- "vertexai/gemini-1.0-pro-001": {
- "max_tokens": 8192,
- "max_input_tokens": 32760,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.0000005
- },
- "vertexai/gemini-1.0-pro-002": {
- "max_tokens": 8192,
- "max_input_tokens": 32760,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.0000005
- },
- "vertexai/gemini-1.5-pro": {
- "max_tokens": 8192,
- "max_input_tokens": 1000000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.000000625,
- "output_cost_per_token": 0.000001875
- },
- "vertexai/gemini-1.5-flash-001": {
- "max_tokens": 8192,
- "max_input_tokens": 1000000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0,
- "output_cost_per_token": 0
- },
- "vertexai/gemini-1.5-flash-preview-0514": {
- "max_tokens": 8192,
- "max_input_tokens": 1000000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0,
- "output_cost_per_token": 0
- },
- "vertexai/gemini-1.5-pro-001": {
- "max_tokens": 8192,
- "max_input_tokens": 1000000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.000000625,
- "output_cost_per_token": 0.000001875
- },
- "vertexai/gemini-1.5-pro-preview-0514": {
- "max_tokens": 8192,
- "max_input_tokens": 1000000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.000000625,
- "output_cost_per_token": 0.000001875
- },
- "vertexai/gemini-1.5-pro-preview-0215": {
- "max_tokens": 8192,
- "max_input_tokens": 1000000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.000000625,
- "output_cost_per_token": 0.000001875
- },
- "vertexai/gemini-1.5-pro-preview-0409": {
- "max_tokens": 8192,
- "max_input_tokens": 1000000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0.000000625,
- "output_cost_per_token": 0.000001875
- },
- "vertexai/gemini-experimental": {
- "max_tokens": 8192,
- "max_input_tokens": 1000000,
- "max_output_tokens": 8192,
- "input_cost_per_token": 0,
- "output_cost_per_token": 0
- },
- "vertexai/gemini-pro-vision": {
- "max_tokens": 2048,
- "max_input_tokens": 16384,
- "max_output_tokens": 2048,
- "max_images_per_prompt": 16,
- "max_videos_per_prompt": 1,
- "max_video_length": 2,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.0000005
- },
- "vertexai/gemini-1.0-pro-vision": {
- "max_tokens": 2048,
- "max_input_tokens": 16384,
- "max_output_tokens": 2048,
- "max_images_per_prompt": 16,
- "max_videos_per_prompt": 1,
- "max_video_length": 2,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.0000005
- },
- "vertexai/gemini-1.0-pro-vision-001": {
- "max_tokens": 2048,
- "max_input_tokens": 16384,
- "max_output_tokens": 2048,
- "max_images_per_prompt": 16,
- "max_videos_per_prompt": 1,
- "max_video_length": 2,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.0000005
- },
- "vertexai/claude-3-sonnet@20240229": {
- "max_tokens": 4096,
- "max_input_tokens": 200000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000003,
- "output_cost_per_token": 0.000015
- },
- "vertexai/claude-3-haiku@20240307": {
- "max_tokens": 4096,
- "max_input_tokens": 200000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00000025,
- "output_cost_per_token": 0.00000125
- },
- "vertexai/claude-3-opus@20240229": {
- "max_tokens": 4096,
- "max_input_tokens": 200000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000015,
- "output_cost_per_token": 0.000075
- },
- "cohere/command-r": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.00000050,
- "output_cost_per_token": 0.0000015
- },
- "cohere/command-light": {
- "max_tokens": 4096,
- "max_input_tokens": 4096,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000015,
- "output_cost_per_token": 0.000015
- },
- "cohere/command-r-plus": {
- "max_tokens": 4096,
- "max_input_tokens": 128000,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000003,
- "output_cost_per_token": 0.000015
- },
- "cohere/command-nightly": {
- "max_tokens": 4096,
- "max_input_tokens": 4096,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000015,
- "output_cost_per_token": 0.000015
- },
- "cohere/command": {
- "max_tokens": 4096,
- "max_input_tokens": 4096,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000015,
- "output_cost_per_token": 0.000015
- },
- "cohere/command-medium-beta": {
- "max_tokens": 4096,
- "max_input_tokens": 4096,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000015,
- "output_cost_per_token": 0.000015
- },
- "cohere/command-xlarge-beta": {
- "max_tokens": 4096,
- "max_input_tokens": 4096,
- "max_output_tokens": 4096,
- "input_cost_per_token": 0.000015,
- "output_cost_per_token": 0.000015
- },
- "together/together-ai-up-to-3b": {
- "input_cost_per_token": 0.0000001,
- "output_cost_per_token": 0.0000001
- },
- "together/together-ai-3.1b-7b": {
- "input_cost_per_token": 0.0000002,
- "output_cost_per_token": 0.0000002
- },
- "together/together-ai-7.1b-20b": {
- "max_tokens": 1000,
- "input_cost_per_token": 0.0000004,
- "output_cost_per_token": 0.0000004
- },
- "together/together-ai-20.1b-40b": {
- "input_cost_per_token": 0.0000008,
- "output_cost_per_token": 0.0000008
- },
- "together/together-ai-40.1b-70b": {
- "input_cost_per_token": 0.0000009,
- "output_cost_per_token": 0.0000009
- },
- "together/mistralai/Mixtral-8x7B-Instruct-v0.1": {
- "input_cost_per_token": 0.0000006,
- "output_cost_per_token": 0.0000006
- }
-}
\ No newline at end of file
diff --git a/embedchain/embedchain/config/vector_db/base.py b/embedchain/embedchain/config/vector_db/base.py
deleted file mode 100644
index 3252880a9..000000000
--- a/embedchain/embedchain/config/vector_db/base.py
+++ /dev/null
@@ -1,36 +0,0 @@
-from typing import Optional
-
-from embedchain.config.base_config import BaseConfig
-
-
-class BaseVectorDbConfig(BaseConfig):
- def __init__(
- self,
- collection_name: Optional[str] = None,
- dir: str = "db",
- host: Optional[str] = None,
- port: Optional[str] = None,
- **kwargs,
- ):
- """
- Initializes a configuration class instance for the vector database.
-
- :param collection_name: Default name for the collection, defaults to None
- :type collection_name: Optional[str], optional
- :param dir: Path to the database directory, where the database is stored, defaults to "db"
- :type dir: str, optional
- :param host: Database connection remote host. Use this if you run Embedchain as a client, defaults to None
- :type host: Optional[str], optional
- :param host: Database connection remote port. Use this if you run Embedchain as a client, defaults to None
- :type port: Optional[str], optional
- :param kwargs: Additional keyword arguments
- :type kwargs: dict
- """
- self.collection_name = collection_name or "embedchain_store"
- self.dir = dir
- self.host = host
- self.port = port
- # Assign additional keyword arguments
- if kwargs:
- for key, value in kwargs.items():
- setattr(self, key, value)
diff --git a/embedchain/embedchain/config/vector_db/chroma.py b/embedchain/embedchain/config/vector_db/chroma.py
deleted file mode 100644
index 64220165c..000000000
--- a/embedchain/embedchain/config/vector_db/chroma.py
+++ /dev/null
@@ -1,41 +0,0 @@
-from typing import Optional
-
-from embedchain.config.vector_db.base import BaseVectorDbConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class ChromaDbConfig(BaseVectorDbConfig):
- def __init__(
- self,
- collection_name: Optional[str] = None,
- dir: Optional[str] = None,
- host: Optional[str] = None,
- port: Optional[str] = None,
- batch_size: Optional[int] = 100,
- allow_reset=False,
- chroma_settings: Optional[dict] = None,
- ):
- """
- Initializes a configuration class instance for ChromaDB.
-
- :param collection_name: Default name for the collection, defaults to None
- :type collection_name: Optional[str], optional
- :param dir: Path to the database directory, where the database is stored, defaults to None
- :type dir: Optional[str], optional
- :param host: Database connection remote host. Use this if you run Embedchain as a client, defaults to None
- :type host: Optional[str], optional
- :param port: Database connection remote port. Use this if you run Embedchain as a client, defaults to None
- :type port: Optional[str], optional
- :param batch_size: Number of items to insert in one batch, defaults to 100
- :type batch_size: Optional[int], optional
- :param allow_reset: Resets the database. defaults to False
- :type allow_reset: bool
- :param chroma_settings: Chroma settings dict, defaults to None
- :type chroma_settings: Optional[dict], optional
- """
-
- self.chroma_settings = chroma_settings
- self.allow_reset = allow_reset
- self.batch_size = batch_size
- super().__init__(collection_name=collection_name, dir=dir, host=host, port=port)
diff --git a/embedchain/embedchain/config/vector_db/elasticsearch.py b/embedchain/embedchain/config/vector_db/elasticsearch.py
deleted file mode 100644
index 5e8ef6b61..000000000
--- a/embedchain/embedchain/config/vector_db/elasticsearch.py
+++ /dev/null
@@ -1,56 +0,0 @@
-import os
-from typing import Optional, Union
-
-from embedchain.config.vector_db.base import BaseVectorDbConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class ElasticsearchDBConfig(BaseVectorDbConfig):
- def __init__(
- self,
- collection_name: Optional[str] = None,
- dir: Optional[str] = None,
- es_url: Union[str, list[str]] = None,
- cloud_id: Optional[str] = None,
- batch_size: Optional[int] = 100,
- **ES_EXTRA_PARAMS: dict[str, any],
- ):
- """
- Initializes a configuration class instance for an Elasticsearch client.
-
- :param collection_name: Default name for the collection, defaults to None
- :type collection_name: Optional[str], optional
- :param dir: Path to the database directory, where the database is stored, defaults to None
- :type dir: Optional[str], optional
- :param es_url: elasticsearch url or list of nodes url to be used for connection, defaults to None
- :type es_url: Union[str, list[str]], optional
- :param cloud_id: cloud id of the elasticsearch cluster, defaults to None
- :type cloud_id: Optional[str], optional
- :param batch_size: Number of items to insert in one batch, defaults to 100
- :type batch_size: Optional[int], optional
- :param ES_EXTRA_PARAMS: extra params dict that can be passed to elasticsearch.
- :type ES_EXTRA_PARAMS: dict[str, Any], optional
- """
- if es_url and cloud_id:
- raise ValueError("Only one of `es_url` and `cloud_id` can be set.")
- # self, es_url: Union[str, list[str]] = None, **ES_EXTRA_PARAMS: dict[str, any]):
- self.ES_URL = es_url or os.environ.get("ELASTICSEARCH_URL")
- self.CLOUD_ID = cloud_id or os.environ.get("ELASTICSEARCH_CLOUD_ID")
- if not self.ES_URL and not self.CLOUD_ID:
- raise AttributeError(
- "Elasticsearch needs a URL or CLOUD_ID attribute, "
- "this can either be passed to `ElasticsearchDBConfig` or as `ELASTICSEARCH_URL` or `ELASTICSEARCH_CLOUD_ID` in `.env`" # noqa: E501
- )
- self.ES_EXTRA_PARAMS = ES_EXTRA_PARAMS
- # Load API key from .env if it's not explicitly passed.
- # Can only set one of 'api_key', 'basic_auth', and 'bearer_auth'
- if (
- not self.ES_EXTRA_PARAMS.get("api_key")
- and not self.ES_EXTRA_PARAMS.get("basic_auth")
- and not self.ES_EXTRA_PARAMS.get("bearer_auth")
- ):
- self.ES_EXTRA_PARAMS["api_key"] = os.environ.get("ELASTICSEARCH_API_KEY")
-
- self.batch_size = batch_size
- super().__init__(collection_name=collection_name, dir=dir)
diff --git a/embedchain/embedchain/config/vector_db/lancedb.py b/embedchain/embedchain/config/vector_db/lancedb.py
deleted file mode 100644
index 08b7d0ac7..000000000
--- a/embedchain/embedchain/config/vector_db/lancedb.py
+++ /dev/null
@@ -1,33 +0,0 @@
-from typing import Optional
-
-from embedchain.config.vector_db.base import BaseVectorDbConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class LanceDBConfig(BaseVectorDbConfig):
- def __init__(
- self,
- collection_name: Optional[str] = None,
- dir: Optional[str] = None,
- host: Optional[str] = None,
- port: Optional[str] = None,
- allow_reset=True,
- ):
- """
- Initializes a configuration class instance for LanceDB.
-
- :param collection_name: Default name for the collection, defaults to None
- :type collection_name: Optional[str], optional
- :param dir: Path to the database directory, where the database is stored, defaults to None
- :type dir: Optional[str], optional
- :param host: Database connection remote host. Use this if you run Embedchain as a client, defaults to None
- :type host: Optional[str], optional
- :param port: Database connection remote port. Use this if you run Embedchain as a client, defaults to None
- :type port: Optional[str], optional
- :param allow_reset: Resets the database. defaults to False
- :type allow_reset: bool
- """
-
- self.allow_reset = allow_reset
- super().__init__(collection_name=collection_name, dir=dir, host=host, port=port)
diff --git a/embedchain/embedchain/config/vector_db/opensearch.py b/embedchain/embedchain/config/vector_db/opensearch.py
deleted file mode 100644
index 5beeb8cee..000000000
--- a/embedchain/embedchain/config/vector_db/opensearch.py
+++ /dev/null
@@ -1,41 +0,0 @@
-from typing import Optional
-
-from embedchain.config.vector_db.base import BaseVectorDbConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class OpenSearchDBConfig(BaseVectorDbConfig):
- def __init__(
- self,
- opensearch_url: str,
- http_auth: tuple[str, str],
- vector_dimension: int = 1536,
- collection_name: Optional[str] = None,
- dir: Optional[str] = None,
- batch_size: Optional[int] = 100,
- **extra_params: dict[str, any],
- ):
- """
- Initializes a configuration class instance for an OpenSearch client.
-
- :param collection_name: Default name for the collection, defaults to None
- :type collection_name: Optional[str], optional
- :param opensearch_url: URL of the OpenSearch domain
- :type opensearch_url: str, Eg, "http://localhost:9200"
- :param http_auth: Tuple of username and password
- :type http_auth: tuple[str, str], Eg, ("username", "password")
- :param vector_dimension: Dimension of the vector, defaults to 1536 (openai embedding model)
- :type vector_dimension: int, optional
- :param dir: Path to the database directory, where the database is stored, defaults to None
- :type dir: Optional[str], optional
- :param batch_size: Number of items to insert in one batch, defaults to 100
- :type batch_size: Optional[int], optional
- """
- self.opensearch_url = opensearch_url
- self.http_auth = http_auth
- self.vector_dimension = vector_dimension
- self.extra_params = extra_params
- self.batch_size = batch_size
-
- super().__init__(collection_name=collection_name, dir=dir)
diff --git a/embedchain/embedchain/config/vector_db/pinecone.py b/embedchain/embedchain/config/vector_db/pinecone.py
deleted file mode 100644
index 83248579f..000000000
--- a/embedchain/embedchain/config/vector_db/pinecone.py
+++ /dev/null
@@ -1,47 +0,0 @@
-import os
-from typing import Optional
-
-from embedchain.config.vector_db.base import BaseVectorDbConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class PineconeDBConfig(BaseVectorDbConfig):
- def __init__(
- self,
- index_name: Optional[str] = None,
- api_key: Optional[str] = None,
- vector_dimension: int = 1536,
- metric: Optional[str] = "cosine",
- pod_config: Optional[dict[str, any]] = None,
- serverless_config: Optional[dict[str, any]] = None,
- hybrid_search: bool = False,
- bm25_encoder: any = None,
- batch_size: Optional[int] = 100,
- **extra_params: dict[str, any],
- ):
- self.metric = metric
- self.api_key = api_key
- self.index_name = index_name
- self.vector_dimension = vector_dimension
- self.extra_params = extra_params
- self.hybrid_search = hybrid_search
- self.bm25_encoder = bm25_encoder
- self.batch_size = batch_size
- if pod_config is None and serverless_config is None:
- # If no config is provided, use the default pod spec config
- pod_environment = os.environ.get("PINECONE_ENV", "gcp-starter")
- self.pod_config = {"environment": pod_environment, "metadata_config": {"indexed": ["*"]}}
- else:
- self.pod_config = pod_config
- self.serverless_config = serverless_config
-
- if self.pod_config and self.serverless_config:
- raise ValueError("Only one of pod_config or serverless_config can be provided.")
-
- if self.hybrid_search and self.metric != "dotproduct":
- raise ValueError(
- "Hybrid search is only supported with dotproduct metric in Pinecone. See full docs here: https://docs.pinecone.io/docs/hybrid-search#limitations"
- ) # noqa:E501
-
- super().__init__(collection_name=self.index_name, dir=None)
diff --git a/embedchain/embedchain/config/vector_db/qdrant.py b/embedchain/embedchain/config/vector_db/qdrant.py
deleted file mode 100644
index acdeacfff..000000000
--- a/embedchain/embedchain/config/vector_db/qdrant.py
+++ /dev/null
@@ -1,48 +0,0 @@
-from typing import Optional
-
-from embedchain.config.vector_db.base import BaseVectorDbConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class QdrantDBConfig(BaseVectorDbConfig):
- """
- Config to initialize a qdrant client.
- :param: url. qdrant url or list of nodes url to be used for connection
- """
-
- def __init__(
- self,
- collection_name: Optional[str] = None,
- dir: Optional[str] = None,
- hnsw_config: Optional[dict[str, any]] = None,
- quantization_config: Optional[dict[str, any]] = None,
- on_disk: Optional[bool] = None,
- batch_size: Optional[int] = 10,
- **extra_params: dict[str, any],
- ):
- """
- Initializes a configuration class instance for a qdrant client.
-
- :param collection_name: Default name for the collection, defaults to None
- :type collection_name: Optional[str], optional
- :param dir: Path to the database directory, where the database is stored, defaults to None
- :type dir: Optional[str], optional
- :param hnsw_config: Params for HNSW index
- :type hnsw_config: Optional[dict[str, any]], defaults to None
- :param quantization_config: Params for quantization, if None - quantization will be disabled
- :type quantization_config: Optional[dict[str, any]], defaults to None
- :param on_disk: If true - point`s payload will not be stored in memory.
- It will be read from the disk every time it is requested.
- This setting saves RAM by (slightly) increasing the response time.
- Note: those payload values that are involved in filtering and are indexed - remain in RAM.
- :type on_disk: bool, optional, defaults to None
- :param batch_size: Number of items to insert in one batch, defaults to 10
- :type batch_size: Optional[int], optional
- """
- self.hnsw_config = hnsw_config
- self.quantization_config = quantization_config
- self.on_disk = on_disk
- self.batch_size = batch_size
- self.extra_params = extra_params
- super().__init__(collection_name=collection_name, dir=dir)
diff --git a/embedchain/embedchain/config/vector_db/weaviate.py b/embedchain/embedchain/config/vector_db/weaviate.py
deleted file mode 100644
index f40c472e7..000000000
--- a/embedchain/embedchain/config/vector_db/weaviate.py
+++ /dev/null
@@ -1,18 +0,0 @@
-from typing import Optional
-
-from embedchain.config.vector_db.base import BaseVectorDbConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class WeaviateDBConfig(BaseVectorDbConfig):
- def __init__(
- self,
- collection_name: Optional[str] = None,
- dir: Optional[str] = None,
- batch_size: Optional[int] = 100,
- **extra_params: dict[str, any],
- ):
- self.batch_size = batch_size
- self.extra_params = extra_params
- super().__init__(collection_name=collection_name, dir=dir)
diff --git a/embedchain/embedchain/config/vector_db/zilliz.py b/embedchain/embedchain/config/vector_db/zilliz.py
deleted file mode 100644
index 268941157..000000000
--- a/embedchain/embedchain/config/vector_db/zilliz.py
+++ /dev/null
@@ -1,49 +0,0 @@
-import os
-from typing import Optional
-
-from embedchain.config.vector_db.base import BaseVectorDbConfig
-from embedchain.helpers.json_serializable import register_deserializable
-
-
-@register_deserializable
-class ZillizDBConfig(BaseVectorDbConfig):
- def __init__(
- self,
- collection_name: Optional[str] = None,
- dir: Optional[str] = None,
- uri: Optional[str] = None,
- token: Optional[str] = None,
- vector_dim: Optional[str] = None,
- metric_type: Optional[str] = None,
- ):
- """
- Initializes a configuration class instance for the vector database.
-
- :param collection_name: Default name for the collection, defaults to None
- :type collection_name: Optional[str], optional
- :param dir: Path to the database directory, where the database is stored, defaults to "db"
- :type dir: str, optional
- :param uri: Cluster endpoint obtained from the Zilliz Console, defaults to None
- :type uri: Optional[str], optional
- :param token: API Key, if a Serverless Cluster, username:password, if a Dedicated Cluster, defaults to None
- :type token: Optional[str], optional
- """
- self.uri = uri or os.environ.get("ZILLIZ_CLOUD_URI")
- if not self.uri:
- raise AttributeError(
- "Zilliz needs a URI attribute, "
- "this can either be passed to `ZILLIZ_CLOUD_URI` or as `ZILLIZ_CLOUD_URI` in `.env`"
- )
-
- self.token = token or os.environ.get("ZILLIZ_CLOUD_TOKEN")
- if not self.token:
- raise AttributeError(
- "Zilliz needs a token attribute, "
- "this can either be passed to `ZILLIZ_CLOUD_TOKEN` or as `ZILLIZ_CLOUD_TOKEN` in `.env`,"
- "if having a username and password, pass it in the form 'username:password' to `ZILLIZ_CLOUD_TOKEN`"
- )
-
- self.metric_type = metric_type if metric_type else "L2"
-
- self.vector_dim = vector_dim
- super().__init__(collection_name=collection_name, dir=dir)
diff --git a/embedchain/embedchain/config/vectordb/__init__.py b/embedchain/embedchain/config/vectordb/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/constants.py b/embedchain/embedchain/constants.py
deleted file mode 100644
index d3d7b28b3..000000000
--- a/embedchain/embedchain/constants.py
+++ /dev/null
@@ -1,11 +0,0 @@
-import os
-from pathlib import Path
-
-ABS_PATH = os.getcwd()
-HOME_DIR = os.environ.get("EMBEDCHAIN_CONFIG_DIR", str(Path.home()))
-CONFIG_DIR = os.path.join(HOME_DIR, ".embedchain")
-CONFIG_FILE = os.path.join(CONFIG_DIR, "config.json")
-SQLITE_PATH = os.path.join(CONFIG_DIR, "embedchain.db")
-
-# Set the environment variable for the database URI
-os.environ.setdefault("EMBEDCHAIN_DB_URI", f"sqlite:///{SQLITE_PATH}")
diff --git a/embedchain/embedchain/core/__init__.py b/embedchain/embedchain/core/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/core/db/__init__.py b/embedchain/embedchain/core/db/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/core/db/database.py b/embedchain/embedchain/core/db/database.py
deleted file mode 100644
index 0965ca8ff..000000000
--- a/embedchain/embedchain/core/db/database.py
+++ /dev/null
@@ -1,88 +0,0 @@
-import os
-
-from alembic import command
-from alembic.config import Config
-from sqlalchemy import create_engine
-from sqlalchemy.engine.base import Engine
-from sqlalchemy.orm import Session as SQLAlchemySession
-from sqlalchemy.orm import scoped_session, sessionmaker
-
-from .models import Base
-
-
-class DatabaseManager:
- def __init__(self, echo: bool = False):
- self.database_uri = os.environ.get("EMBEDCHAIN_DB_URI")
- self.echo = echo
- self.engine: Engine = None
- self._session_factory = None
-
- def setup_engine(self) -> None:
- """Initializes the database engine and session factory."""
- if not self.database_uri:
- raise RuntimeError("Database URI is not set. Set the EMBEDCHAIN_DB_URI environment variable.")
- connect_args = {}
- if self.database_uri.startswith("sqlite"):
- connect_args["check_same_thread"] = False
- self.engine = create_engine(self.database_uri, echo=self.echo, connect_args=connect_args)
- self._session_factory = scoped_session(sessionmaker(bind=self.engine))
- Base.metadata.bind = self.engine
-
- def init_db(self) -> None:
- """Creates all tables defined in the Base metadata."""
- if not self.engine:
- raise RuntimeError("Database engine is not initialized. Call setup_engine() first.")
- Base.metadata.create_all(self.engine)
-
- def get_session(self) -> SQLAlchemySession:
- """Provides a session for database operations."""
- if not self._session_factory:
- raise RuntimeError("Session factory is not initialized. Call setup_engine() first.")
- return self._session_factory()
-
- def close_session(self) -> None:
- """Closes the current session."""
- if self._session_factory:
- self._session_factory.remove()
-
- def execute_transaction(self, transaction_block):
- """Executes a block of code within a database transaction."""
- session = self.get_session()
- try:
- transaction_block(session)
- session.commit()
- except Exception as e:
- session.rollback()
- raise e
- finally:
- self.close_session()
-
-
-# Singleton pattern to use throughout the application
-database_manager = DatabaseManager()
-
-
-# Convenience functions for backward compatibility and ease of use
-def setup_engine(database_uri: str, echo: bool = False) -> None:
- database_manager.database_uri = database_uri
- database_manager.echo = echo
- database_manager.setup_engine()
-
-
-def alembic_upgrade() -> None:
- """Upgrades the database to the latest version."""
- alembic_config_path = os.path.join(os.path.dirname(__file__), "..", "..", "alembic.ini")
- alembic_cfg = Config(alembic_config_path)
- command.upgrade(alembic_cfg, "head")
-
-
-def init_db() -> None:
- alembic_upgrade()
-
-
-def get_session() -> SQLAlchemySession:
- return database_manager.get_session()
-
-
-def execute_transaction(transaction_block):
- database_manager.execute_transaction(transaction_block)
diff --git a/embedchain/embedchain/core/db/models.py b/embedchain/embedchain/core/db/models.py
deleted file mode 100644
index af77803f7..000000000
--- a/embedchain/embedchain/core/db/models.py
+++ /dev/null
@@ -1,31 +0,0 @@
-import uuid
-
-from sqlalchemy import TIMESTAMP, Column, Integer, String, Text, func
-from sqlalchemy.orm import declarative_base
-
-Base = declarative_base()
-metadata = Base.metadata
-
-
-class DataSource(Base):
- __tablename__ = "ec_data_sources"
-
- id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()))
- app_id = Column(Text, index=True)
- hash = Column(Text, index=True)
- type = Column(Text, index=True)
- value = Column(Text)
- meta_data = Column(Text, name="metadata")
- is_uploaded = Column(Integer, default=0)
-
-
-class ChatHistory(Base):
- __tablename__ = "ec_chat_history"
-
- app_id = Column(String, primary_key=True)
- id = Column(String, primary_key=True)
- session_id = Column(String, primary_key=True, index=True)
- question = Column(Text)
- answer = Column(Text)
- meta_data = Column(Text, name="metadata")
- created_at = Column(TIMESTAMP, default=func.current_timestamp(), index=True)
diff --git a/embedchain/embedchain/data_formatter/__init__.py b/embedchain/embedchain/data_formatter/__init__.py
deleted file mode 100644
index 047b8e7ca..000000000
--- a/embedchain/embedchain/data_formatter/__init__.py
+++ /dev/null
@@ -1 +0,0 @@
-from .data_formatter import DataFormatter # noqa: F401
diff --git a/embedchain/embedchain/data_formatter/data_formatter.py b/embedchain/embedchain/data_formatter/data_formatter.py
deleted file mode 100644
index 72923888d..000000000
--- a/embedchain/embedchain/data_formatter/data_formatter.py
+++ /dev/null
@@ -1,154 +0,0 @@
-from importlib import import_module
-from typing import Any, Optional
-
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config import AddConfig
-from embedchain.config.add_config import ChunkerConfig, LoaderConfig
-from embedchain.helpers.json_serializable import JSONSerializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.models.data_type import DataType
-
-
-class DataFormatter(JSONSerializable):
- """
- DataFormatter is an internal utility class which abstracts the mapping for
- loaders and chunkers to the data_type entered by the user in their
- .add or .add_local method call
- """
-
- def __init__(
- self,
- data_type: DataType,
- config: AddConfig,
- loader: Optional[BaseLoader] = None,
- chunker: Optional[BaseChunker] = None,
- ):
- """
- Initialize a dataformatter, set data type and chunker based on datatype.
-
- :param data_type: The type of the data to load and chunk.
- :type data_type: DataType
- :param config: AddConfig instance with nested loader and chunker config attributes.
- :type config: AddConfig
- """
- self.loader = self._get_loader(data_type=data_type, config=config.loader, loader=loader)
- self.chunker = self._get_chunker(data_type=data_type, config=config.chunker, chunker=chunker)
-
- @staticmethod
- def _lazy_load(module_path: str):
- module_path, class_name = module_path.rsplit(".", 1)
- module = import_module(module_path)
- return getattr(module, class_name)
-
- def _get_loader(
- self,
- data_type: DataType,
- config: LoaderConfig,
- loader: Optional[BaseLoader],
- **kwargs: Optional[dict[str, Any]],
- ) -> BaseLoader:
- """
- Returns the appropriate data loader for the given data type.
-
- :param data_type: The type of the data to load.
- :type data_type: DataType
- :param config: Config to initialize the loader with.
- :type config: LoaderConfig
- :raises ValueError: If an unsupported data type is provided.
- :return: The loader for the given data type.
- :rtype: BaseLoader
- """
- loaders = {
- DataType.YOUTUBE_VIDEO: "embedchain.loaders.youtube_video.YoutubeVideoLoader",
- DataType.PDF_FILE: "embedchain.loaders.pdf_file.PdfFileLoader",
- DataType.WEB_PAGE: "embedchain.loaders.web_page.WebPageLoader",
- DataType.QNA_PAIR: "embedchain.loaders.local_qna_pair.LocalQnaPairLoader",
- DataType.TEXT: "embedchain.loaders.local_text.LocalTextLoader",
- DataType.DOCX: "embedchain.loaders.docx_file.DocxFileLoader",
- DataType.SITEMAP: "embedchain.loaders.sitemap.SitemapLoader",
- DataType.XML: "embedchain.loaders.xml.XmlLoader",
- DataType.DOCS_SITE: "embedchain.loaders.docs_site_loader.DocsSiteLoader",
- DataType.CSV: "embedchain.loaders.csv.CsvLoader",
- DataType.MDX: "embedchain.loaders.mdx.MdxLoader",
- DataType.IMAGE: "embedchain.loaders.image.ImageLoader",
- DataType.UNSTRUCTURED: "embedchain.loaders.unstructured_file.UnstructuredLoader",
- DataType.JSON: "embedchain.loaders.json.JSONLoader",
- DataType.OPENAPI: "embedchain.loaders.openapi.OpenAPILoader",
- DataType.GMAIL: "embedchain.loaders.gmail.GmailLoader",
- DataType.NOTION: "embedchain.loaders.notion.NotionLoader",
- DataType.SUBSTACK: "embedchain.loaders.substack.SubstackLoader",
- DataType.YOUTUBE_CHANNEL: "embedchain.loaders.youtube_channel.YoutubeChannelLoader",
- DataType.DISCORD: "embedchain.loaders.discord.DiscordLoader",
- DataType.RSSFEED: "embedchain.loaders.rss_feed.RSSFeedLoader",
- DataType.BEEHIIV: "embedchain.loaders.beehiiv.BeehiivLoader",
- DataType.GOOGLE_DRIVE: "embedchain.loaders.google_drive.GoogleDriveLoader",
- DataType.DIRECTORY: "embedchain.loaders.directory_loader.DirectoryLoader",
- DataType.SLACK: "embedchain.loaders.slack.SlackLoader",
- DataType.DROPBOX: "embedchain.loaders.dropbox.DropboxLoader",
- DataType.TEXT_FILE: "embedchain.loaders.text_file.TextFileLoader",
- DataType.EXCEL_FILE: "embedchain.loaders.excel_file.ExcelFileLoader",
- DataType.AUDIO: "embedchain.loaders.audio.AudioLoader",
- }
-
- if data_type == DataType.CUSTOM or loader is not None:
- loader_class: type = loader
- if loader_class:
- return loader_class
- elif data_type in loaders:
- loader_class: type = self._lazy_load(loaders[data_type])
- return loader_class()
-
- raise ValueError(
- f"Cant find the loader for {data_type}.\
- We recommend to pass the loader to use data_type: {data_type},\
- check `https://docs.embedchain.ai/data-sources/overview`."
- )
-
- def _get_chunker(self, data_type: DataType, config: ChunkerConfig, chunker: Optional[BaseChunker]) -> BaseChunker:
- """Returns the appropriate chunker for the given data type (updated for lazy loading)."""
- chunker_classes = {
- DataType.YOUTUBE_VIDEO: "embedchain.chunkers.youtube_video.YoutubeVideoChunker",
- DataType.PDF_FILE: "embedchain.chunkers.pdf_file.PdfFileChunker",
- DataType.WEB_PAGE: "embedchain.chunkers.web_page.WebPageChunker",
- DataType.QNA_PAIR: "embedchain.chunkers.qna_pair.QnaPairChunker",
- DataType.TEXT: "embedchain.chunkers.text.TextChunker",
- DataType.DOCX: "embedchain.chunkers.docx_file.DocxFileChunker",
- DataType.SITEMAP: "embedchain.chunkers.sitemap.SitemapChunker",
- DataType.XML: "embedchain.chunkers.xml.XmlChunker",
- DataType.DOCS_SITE: "embedchain.chunkers.docs_site.DocsSiteChunker",
- DataType.CSV: "embedchain.chunkers.table.TableChunker",
- DataType.MDX: "embedchain.chunkers.mdx.MdxChunker",
- DataType.IMAGE: "embedchain.chunkers.image.ImageChunker",
- DataType.UNSTRUCTURED: "embedchain.chunkers.unstructured_file.UnstructuredFileChunker",
- DataType.JSON: "embedchain.chunkers.json.JSONChunker",
- DataType.OPENAPI: "embedchain.chunkers.openapi.OpenAPIChunker",
- DataType.GMAIL: "embedchain.chunkers.gmail.GmailChunker",
- DataType.NOTION: "embedchain.chunkers.notion.NotionChunker",
- DataType.SUBSTACK: "embedchain.chunkers.substack.SubstackChunker",
- DataType.YOUTUBE_CHANNEL: "embedchain.chunkers.common_chunker.CommonChunker",
- DataType.DISCORD: "embedchain.chunkers.common_chunker.CommonChunker",
- DataType.CUSTOM: "embedchain.chunkers.common_chunker.CommonChunker",
- DataType.RSSFEED: "embedchain.chunkers.rss_feed.RSSFeedChunker",
- DataType.BEEHIIV: "embedchain.chunkers.beehiiv.BeehiivChunker",
- DataType.GOOGLE_DRIVE: "embedchain.chunkers.google_drive.GoogleDriveChunker",
- DataType.DIRECTORY: "embedchain.chunkers.common_chunker.CommonChunker",
- DataType.SLACK: "embedchain.chunkers.common_chunker.CommonChunker",
- DataType.DROPBOX: "embedchain.chunkers.common_chunker.CommonChunker",
- DataType.TEXT_FILE: "embedchain.chunkers.common_chunker.CommonChunker",
- DataType.EXCEL_FILE: "embedchain.chunkers.excel_file.ExcelFileChunker",
- DataType.AUDIO: "embedchain.chunkers.audio.AudioChunker",
- }
-
- if chunker is not None:
- return chunker
- elif data_type in chunker_classes:
- chunker_class = self._lazy_load(chunker_classes[data_type])
- chunker = chunker_class(config)
- chunker.set_data_type(data_type)
- return chunker
-
- raise ValueError(
- f"Cant find the chunker for {data_type}.\
- We recommend to pass the chunker to use data_type: {data_type},\
- check `https://docs.embedchain.ai/data-sources/overview`."
- )
diff --git a/embedchain/embedchain/deployment/fly.io/.dockerignore b/embedchain/embedchain/deployment/fly.io/.dockerignore
deleted file mode 100644
index 9f4c740db..000000000
--- a/embedchain/embedchain/deployment/fly.io/.dockerignore
+++ /dev/null
@@ -1 +0,0 @@
-db/
\ No newline at end of file
diff --git a/embedchain/embedchain/deployment/fly.io/.env.example b/embedchain/embedchain/deployment/fly.io/.env.example
deleted file mode 100644
index b29363f94..000000000
--- a/embedchain/embedchain/deployment/fly.io/.env.example
+++ /dev/null
@@ -1 +0,0 @@
-OPENAI_API_KEY=sk-xxx
\ No newline at end of file
diff --git a/embedchain/embedchain/deployment/fly.io/Dockerfile b/embedchain/embedchain/deployment/fly.io/Dockerfile
deleted file mode 100644
index 9eac80cee..000000000
--- a/embedchain/embedchain/deployment/fly.io/Dockerfile
+++ /dev/null
@@ -1,13 +0,0 @@
-FROM python:3.11-slim
-
-WORKDIR /app
-
-COPY requirements.txt /app/
-
-RUN pip install -r requirements.txt
-
-COPY . /app
-
-EXPOSE 8080
-
-CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8080"]
diff --git a/embedchain/embedchain/deployment/fly.io/app.py b/embedchain/embedchain/deployment/fly.io/app.py
deleted file mode 100644
index 003543c46..000000000
--- a/embedchain/embedchain/deployment/fly.io/app.py
+++ /dev/null
@@ -1,56 +0,0 @@
-from dotenv import load_dotenv
-from fastapi import FastAPI, responses
-from pydantic import BaseModel
-
-from embedchain import App
-
-load_dotenv(".env")
-
-app = FastAPI(title="Embedchain FastAPI App")
-embedchain_app = App()
-
-
-class SourceModel(BaseModel):
- source: str
-
-
-class QuestionModel(BaseModel):
- question: str
-
-
-@app.post("/add")
-async def add_source(source_model: SourceModel):
- """
- Adds a new source to the EmbedChain app.
- Expects a JSON with a "source" key.
- """
- source = source_model.source
- embedchain_app.add(source)
- return {"message": f"Source '{source}' added successfully."}
-
-
-@app.post("/query")
-async def handle_query(question_model: QuestionModel):
- """
- Handles a query to the EmbedChain app.
- Expects a JSON with a "question" key.
- """
- question = question_model.question
- answer = embedchain_app.query(question)
- return {"answer": answer}
-
-
-@app.post("/chat")
-async def handle_chat(question_model: QuestionModel):
- """
- Handles a chat request to the EmbedChain app.
- Expects a JSON with a "question" key.
- """
- question = question_model.question
- response = embedchain_app.chat(question)
- return {"response": response}
-
-
-@app.get("/")
-async def root():
- return responses.RedirectResponse(url="/docs")
diff --git a/embedchain/embedchain/deployment/fly.io/requirements.txt b/embedchain/embedchain/deployment/fly.io/requirements.txt
deleted file mode 100644
index 3a7689298..000000000
--- a/embedchain/embedchain/deployment/fly.io/requirements.txt
+++ /dev/null
@@ -1,4 +0,0 @@
-fastapi==0.104.0
-uvicorn==0.23.2
-embedchain
-beautifulsoup4
\ No newline at end of file
diff --git a/embedchain/embedchain/deployment/gradio.app/app.py b/embedchain/embedchain/deployment/gradio.app/app.py
deleted file mode 100644
index 24a96a908..000000000
--- a/embedchain/embedchain/deployment/gradio.app/app.py
+++ /dev/null
@@ -1,18 +0,0 @@
-import os
-
-import gradio as gr
-
-from embedchain import App
-
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-app = App()
-
-
-def query(message, history):
- return app.chat(message)
-
-
-demo = gr.ChatInterface(query)
-
-demo.launch()
diff --git a/embedchain/embedchain/deployment/gradio.app/requirements.txt b/embedchain/embedchain/deployment/gradio.app/requirements.txt
deleted file mode 100644
index f8b933480..000000000
--- a/embedchain/embedchain/deployment/gradio.app/requirements.txt
+++ /dev/null
@@ -1,2 +0,0 @@
-gradio>=4.14.0
-embedchain
diff --git a/embedchain/embedchain/deployment/modal.com/.env.example b/embedchain/embedchain/deployment/modal.com/.env.example
deleted file mode 100644
index b29363f94..000000000
--- a/embedchain/embedchain/deployment/modal.com/.env.example
+++ /dev/null
@@ -1 +0,0 @@
-OPENAI_API_KEY=sk-xxx
\ No newline at end of file
diff --git a/embedchain/embedchain/deployment/modal.com/.gitignore b/embedchain/embedchain/deployment/modal.com/.gitignore
deleted file mode 100644
index 4c49bd78f..000000000
--- a/embedchain/embedchain/deployment/modal.com/.gitignore
+++ /dev/null
@@ -1 +0,0 @@
-.env
diff --git a/embedchain/embedchain/deployment/modal.com/app.py b/embedchain/embedchain/deployment/modal.com/app.py
deleted file mode 100644
index 1e02aeefb..000000000
--- a/embedchain/embedchain/deployment/modal.com/app.py
+++ /dev/null
@@ -1,86 +0,0 @@
-from dotenv import load_dotenv
-from fastapi import Body, FastAPI, responses
-from modal import Image, Secret, Stub, asgi_app
-
-from embedchain import App
-
-load_dotenv(".env")
-
-image = Image.debian_slim().pip_install(
- "embedchain",
- "lanchain_community==0.2.6",
- "youtube-transcript-api==0.6.1",
- "pytube==15.0.0",
- "beautifulsoup4==4.12.3",
- "slack-sdk==3.21.3",
- "huggingface_hub==0.23.0",
- "gitpython==3.1.38",
- "yt_dlp==2023.11.14",
- "PyGithub==1.59.1",
- "feedparser==6.0.10",
- "newspaper3k==0.2.8",
- "listparser==0.19",
-)
-
-stub = Stub(
- name="embedchain-app",
- image=image,
- secrets=[Secret.from_dotenv(".env")],
-)
-
-web_app = FastAPI()
-embedchain_app = App(name="embedchain-modal-app")
-
-
-@web_app.post("/add")
-async def add(
- source: str = Body(..., description="Source to be added"),
- data_type: str | None = Body(None, description="Type of the data source"),
-):
- """
- Adds a new source to the EmbedChain app.
- Expects a JSON with a "source" and "data_type" key.
- "data_type" is optional.
- """
- if source and data_type:
- embedchain_app.add(source, data_type)
- elif source:
- embedchain_app.add(source)
- else:
- return {"message": "No source provided."}
- return {"message": f"Source '{source}' added successfully."}
-
-
-@web_app.post("/query")
-async def query(question: str = Body(..., description="Question to be answered")):
- """
- Handles a query to the EmbedChain app.
- Expects a JSON with a "question" key.
- """
- if not question:
- return {"message": "No question provided."}
- answer = embedchain_app.query(question)
- return {"answer": answer}
-
-
-@web_app.get("/chat")
-async def chat(question: str = Body(..., description="Question to be answered")):
- """
- Handles a chat request to the EmbedChain app.
- Expects a JSON with a "question" key.
- """
- if not question:
- return {"message": "No question provided."}
- response = embedchain_app.chat(question)
- return {"response": response}
-
-
-@web_app.get("/")
-async def root():
- return responses.RedirectResponse(url="/docs")
-
-
-@stub.function(image=image)
-@asgi_app()
-def fastapi_app():
- return web_app
diff --git a/embedchain/embedchain/deployment/modal.com/requirements.txt b/embedchain/embedchain/deployment/modal.com/requirements.txt
deleted file mode 100644
index 69a3172af..000000000
--- a/embedchain/embedchain/deployment/modal.com/requirements.txt
+++ /dev/null
@@ -1,4 +0,0 @@
-modal==0.56.4329
-fastapi==0.104.0
-uvicorn==0.23.2
-embedchain
diff --git a/embedchain/embedchain/deployment/render.com/.env.example b/embedchain/embedchain/deployment/render.com/.env.example
deleted file mode 100644
index b29363f94..000000000
--- a/embedchain/embedchain/deployment/render.com/.env.example
+++ /dev/null
@@ -1 +0,0 @@
-OPENAI_API_KEY=sk-xxx
\ No newline at end of file
diff --git a/embedchain/embedchain/deployment/render.com/.gitignore b/embedchain/embedchain/deployment/render.com/.gitignore
deleted file mode 100644
index 4c49bd78f..000000000
--- a/embedchain/embedchain/deployment/render.com/.gitignore
+++ /dev/null
@@ -1 +0,0 @@
-.env
diff --git a/embedchain/embedchain/deployment/render.com/app.py b/embedchain/embedchain/deployment/render.com/app.py
deleted file mode 100644
index 00d29bf3d..000000000
--- a/embedchain/embedchain/deployment/render.com/app.py
+++ /dev/null
@@ -1,53 +0,0 @@
-from fastapi import FastAPI, responses
-from pydantic import BaseModel
-
-from embedchain import App
-
-app = FastAPI(title="Embedchain FastAPI App")
-embedchain_app = App()
-
-
-class SourceModel(BaseModel):
- source: str
-
-
-class QuestionModel(BaseModel):
- question: str
-
-
-@app.post("/add")
-async def add_source(source_model: SourceModel):
- """
- Adds a new source to the EmbedChain app.
- Expects a JSON with a "source" key.
- """
- source = source_model.source
- embedchain_app.add(source)
- return {"message": f"Source '{source}' added successfully."}
-
-
-@app.post("/query")
-async def handle_query(question_model: QuestionModel):
- """
- Handles a query to the EmbedChain app.
- Expects a JSON with a "question" key.
- """
- question = question_model.question
- answer = embedchain_app.query(question)
- return {"answer": answer}
-
-
-@app.post("/chat")
-async def handle_chat(question_model: QuestionModel):
- """
- Handles a chat request to the EmbedChain app.
- Expects a JSON with a "question" key.
- """
- question = question_model.question
- response = embedchain_app.chat(question)
- return {"response": response}
-
-
-@app.get("/")
-async def root():
- return responses.RedirectResponse(url="/docs")
diff --git a/embedchain/embedchain/deployment/render.com/render.yaml b/embedchain/embedchain/deployment/render.com/render.yaml
deleted file mode 100644
index 04ec5048b..000000000
--- a/embedchain/embedchain/deployment/render.com/render.yaml
+++ /dev/null
@@ -1,16 +0,0 @@
-services:
- - type: web
- name: ec-render-app
- runtime: python
- repo: https://github.com//
- scaling:
- minInstances: 1
- maxInstances: 3
- targetMemoryPercent: 60 # optional if targetCPUPercent is set
- targetCPUPercent: 60 # optional if targetMemory is set
- buildCommand: pip install -r requirements.txt
- startCommand: uvicorn app:app --host 0.0.0.0
- envVars:
- - key: OPENAI_API_KEY
- value: sk-xxx
- autoDeploy: false # optional
diff --git a/embedchain/embedchain/deployment/render.com/requirements.txt b/embedchain/embedchain/deployment/render.com/requirements.txt
deleted file mode 100644
index 3a7689298..000000000
--- a/embedchain/embedchain/deployment/render.com/requirements.txt
+++ /dev/null
@@ -1,4 +0,0 @@
-fastapi==0.104.0
-uvicorn==0.23.2
-embedchain
-beautifulsoup4
\ No newline at end of file
diff --git a/embedchain/embedchain/deployment/streamlit.io/.streamlit/secrets.toml b/embedchain/embedchain/deployment/streamlit.io/.streamlit/secrets.toml
deleted file mode 100644
index 1fa8f4495..000000000
--- a/embedchain/embedchain/deployment/streamlit.io/.streamlit/secrets.toml
+++ /dev/null
@@ -1 +0,0 @@
-OPENAI_API_KEY="sk-xxx"
diff --git a/embedchain/embedchain/deployment/streamlit.io/app.py b/embedchain/embedchain/deployment/streamlit.io/app.py
deleted file mode 100644
index 74a6b0599..000000000
--- a/embedchain/embedchain/deployment/streamlit.io/app.py
+++ /dev/null
@@ -1,59 +0,0 @@
-import streamlit as st
-
-from embedchain import App
-
-
-@st.cache_resource
-def embedchain_bot():
- return App()
-
-
-st.title("💬 Chatbot")
-st.caption("🚀 An Embedchain app powered by OpenAI!")
-if "messages" not in st.session_state:
- st.session_state.messages = [
- {
- "role": "assistant",
- "content": """
- Hi! I'm a chatbot. I can answer questions and learn new things!\n
- Ask me anything and if you want me to learn something do `/add `.\n
- I can learn mostly everything. :)
- """,
- }
- ]
-
-for message in st.session_state.messages:
- with st.chat_message(message["role"]):
- st.markdown(message["content"])
-
-if prompt := st.chat_input("Ask me anything!"):
- app = embedchain_bot()
-
- if prompt.startswith("/add"):
- with st.chat_message("user"):
- st.markdown(prompt)
- st.session_state.messages.append({"role": "user", "content": prompt})
- prompt = prompt.replace("/add", "").strip()
- with st.chat_message("assistant"):
- message_placeholder = st.empty()
- message_placeholder.markdown("Adding to knowledge base...")
- app.add(prompt)
- message_placeholder.markdown(f"Added {prompt} to knowledge base!")
- st.session_state.messages.append({"role": "assistant", "content": f"Added {prompt} to knowledge base!"})
- st.stop()
-
- with st.chat_message("user"):
- st.markdown(prompt)
- st.session_state.messages.append({"role": "user", "content": prompt})
-
- with st.chat_message("assistant"):
- msg_placeholder = st.empty()
- msg_placeholder.markdown("Thinking...")
- full_response = ""
-
- for response in app.chat(prompt):
- msg_placeholder.empty()
- full_response += response
-
- msg_placeholder.markdown(full_response)
- st.session_state.messages.append({"role": "assistant", "content": full_response})
diff --git a/embedchain/embedchain/deployment/streamlit.io/requirements.txt b/embedchain/embedchain/deployment/streamlit.io/requirements.txt
deleted file mode 100644
index b864076ae..000000000
--- a/embedchain/embedchain/deployment/streamlit.io/requirements.txt
+++ /dev/null
@@ -1,2 +0,0 @@
-streamlit==1.29.0
-embedchain
diff --git a/embedchain/embedchain/embedchain.py b/embedchain/embedchain/embedchain.py
deleted file mode 100644
index 4a1a4dc09..000000000
--- a/embedchain/embedchain/embedchain.py
+++ /dev/null
@@ -1,789 +0,0 @@
-import hashlib
-import json
-import logging
-from typing import Any, Optional, Union
-
-from dotenv import load_dotenv
-from langchain.docstore.document import Document
-
-from embedchain.cache import (
- adapt,
- get_gptcache_session,
- gptcache_data_convert,
- gptcache_update_cache_callback,
-)
-from embedchain.chunkers.base_chunker import BaseChunker
-from embedchain.config import AddConfig, BaseLlmConfig, ChunkerConfig
-from embedchain.config.base_app_config import BaseAppConfig
-from embedchain.core.db.models import ChatHistory, DataSource
-from embedchain.data_formatter import DataFormatter
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.helpers.json_serializable import JSONSerializable
-from embedchain.llm.base import BaseLlm
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.models.data_type import (
- DataType,
- DirectDataType,
- IndirectDataType,
- SpecialDataType,
-)
-from embedchain.utils.misc import detect_datatype, is_valid_json_string
-from embedchain.vectordb.base import BaseVectorDB
-
-load_dotenv()
-
-logger = logging.getLogger(__name__)
-
-
-class EmbedChain(JSONSerializable):
- def __init__(
- self,
- config: BaseAppConfig,
- llm: BaseLlm,
- db: BaseVectorDB = None,
- embedder: BaseEmbedder = None,
- system_prompt: Optional[str] = None,
- ):
- """
- Initializes the EmbedChain instance, sets up a vector DB client and
- creates a collection.
-
- :param config: Configuration just for the app, not the db or llm or embedder.
- :type config: BaseAppConfig
- :param llm: Instance of the LLM you want to use.
- :type llm: BaseLlm
- :param db: Instance of the Database to use, defaults to None
- :type db: BaseVectorDB, optional
- :param embedder: instance of the embedder to use, defaults to None
- :type embedder: BaseEmbedder, optional
- :param system_prompt: System prompt to use in the llm query, defaults to None
- :type system_prompt: Optional[str], optional
- :raises ValueError: No database or embedder provided.
- """
- self.config = config
- self.cache_config = None
- self.memory_config = None
- self.mem0_memory = None
- # Llm
- self.llm = llm
- # Database has support for config assignment for backwards compatibility
- if db is None and (not hasattr(self.config, "db") or self.config.db is None):
- raise ValueError("App requires Database.")
- self.db = db or self.config.db
- # Embedder
- if embedder is None:
- raise ValueError("App requires Embedder.")
- self.embedder = embedder
-
- # Initialize database
- self.db._set_embedder(self.embedder)
- self.db._initialize()
- # Set collection name from app config for backwards compatibility.
- if config.collection_name:
- self.db.set_collection_name(config.collection_name)
-
- # Add variables that are "shortcuts"
- if system_prompt:
- self.llm.config.system_prompt = system_prompt
-
- # Fetch the history from the database if exists
- self.llm.update_history(app_id=self.config.id)
-
- # Attributes that aren't subclass related.
- self.user_asks = []
-
- self.chunker: Optional[ChunkerConfig] = None
-
- @property
- def collect_metrics(self):
- return self.config.collect_metrics
-
- @collect_metrics.setter
- def collect_metrics(self, value):
- if not isinstance(value, bool):
- raise ValueError(f"Boolean value expected but got {type(value)}.")
- self.config.collect_metrics = value
-
- @property
- def online(self):
- return self.llm.config.online
-
- @online.setter
- def online(self, value):
- if not isinstance(value, bool):
- raise ValueError(f"Boolean value expected but got {type(value)}.")
- self.llm.config.online = value
-
- def add(
- self,
- source: Any,
- data_type: Optional[DataType] = None,
- metadata: Optional[dict[str, Any]] = None,
- config: Optional[AddConfig] = None,
- dry_run=False,
- loader: Optional[BaseLoader] = None,
- chunker: Optional[BaseChunker] = None,
- **kwargs: Optional[dict[str, Any]],
- ):
- """
- Adds the data from the given URL to the vector db.
- Loads the data, chunks it, create embedding for each chunk
- and then stores the embedding to vector database.
-
- :param source: The data to embed, can be a URL, local file or raw content, depending on the data type.
- :type source: Any
- :param data_type: Automatically detected, but can be forced with this argument. The type of the data to add,
- defaults to None
- :type data_type: Optional[DataType], optional
- :param metadata: Metadata associated with the data source., defaults to None
- :type metadata: Optional[dict[str, Any]], optional
- :param config: The `AddConfig` instance to use as configuration options., defaults to None
- :type config: Optional[AddConfig], optional
- :raises ValueError: Invalid data type
- :param dry_run: Optional. A dry run displays the chunks to ensure that the loader and chunker work as intended.
- defaults to False
- :type dry_run: bool
- :param loader: The loader to use to load the data, defaults to None
- :type loader: BaseLoader, optional
- :param chunker: The chunker to use to chunk the data, defaults to None
- :type chunker: BaseChunker, optional
- :param kwargs: To read more params for the query function
- :type kwargs: dict[str, Any]
- :return: source_hash, a md5-hash of the source, in hexadecimal representation.
- :rtype: str
- """
- if config is not None:
- pass
- elif self.chunker is not None:
- config = AddConfig(chunker=self.chunker)
- else:
- config = AddConfig()
-
- try:
- DataType(source)
- logger.warning(
- f"""Starting from version v0.0.40, Embedchain can automatically detect the data type. So, in the `add` method, the argument order has changed. You no longer need to specify '{source}' for the `source` argument. So the code snippet will be `.add("{data_type}", "{source}")`""" # noqa #E501
- )
- logger.warning(
- "Embedchain is swapping the arguments for you. This functionality might be deprecated in the future, so please adjust your code." # noqa #E501
- )
- source, data_type = data_type, source
- except ValueError:
- pass
-
- if data_type:
- try:
- data_type = DataType(data_type)
- except ValueError:
- logger.info(
- f"Invalid data_type: '{data_type}', using `custom` instead.\n Check docs to pass the valid data type: `https://docs.embedchain.ai/data-sources/overview`" # noqa: E501
- )
- data_type = DataType.CUSTOM
-
- if not data_type:
- data_type = detect_datatype(source)
-
- # `source_hash` is the md5 hash of the source argument
- source_hash = hashlib.md5(str(source).encode("utf-8")).hexdigest()
-
- self.user_asks.append([source, data_type.value, metadata])
-
- data_formatter = DataFormatter(data_type, config, loader, chunker)
- documents, metadatas, _ids, new_chunks = self._load_and_embed(
- data_formatter.loader, data_formatter.chunker, source, metadata, source_hash, config, dry_run, **kwargs
- )
- if data_type in {DataType.DOCS_SITE}:
- self.is_docs_site_instance = True
-
- # Convert the source to a string if it is not already
- if not isinstance(source, str):
- source = str(source)
-
- # Insert the data into the 'ec_data_sources' table
- self.db_session.add(
- DataSource(
- hash=source_hash,
- app_id=self.config.id,
- type=data_type.value,
- value=source,
- metadata=json.dumps(metadata),
- )
- )
- try:
- self.db_session.commit()
- except Exception as e:
- logger.error(f"Error adding data source: {e}")
- self.db_session.rollback()
-
- if dry_run:
- data_chunks_info = {"chunks": documents, "metadata": metadatas, "count": len(documents), "type": data_type}
- logger.debug(f"Dry run info : {data_chunks_info}")
- return data_chunks_info
-
- # Send anonymous telemetry
- if self.config.collect_metrics:
- # it's quicker to check the variable twice than to count words when they won't be submitted.
- word_count = data_formatter.chunker.get_word_count(documents)
-
- # Send anonymous telemetry
- event_properties = {
- **self._telemetry_props,
- "data_type": data_type.value,
- "word_count": word_count,
- "chunks_count": new_chunks,
- }
- self.telemetry.capture(event_name="add", properties=event_properties)
-
- return source_hash
-
- def _get_existing_doc_id(self, chunker: BaseChunker, src: Any):
- """
- Get id of existing document for a given source, based on the data type
- """
- # Find existing embeddings for the source
- # Depending on the data type, existing embeddings are checked for.
- if chunker.data_type.value in [item.value for item in DirectDataType]:
- # DirectDataTypes can't be updated.
- # Think of a text:
- # Either it's the same, then it won't change, so it's not an update.
- # Or it's different, then it will be added as a new text.
- return None
- elif chunker.data_type.value in [item.value for item in IndirectDataType]:
- # These types have an indirect source reference
- # As long as the reference is the same, they can be updated.
- where = {"url": src}
- if chunker.data_type == DataType.JSON and is_valid_json_string(src):
- url = hashlib.sha256((src).encode("utf-8")).hexdigest()
- where = {"url": url}
-
- if self.config.id is not None:
- where.update({"app_id": self.config.id})
-
- existing_embeddings = self.db.get(
- where=where,
- limit=1,
- )
- if len(existing_embeddings.get("metadatas", [])) > 0:
- return existing_embeddings["metadatas"][0]["doc_id"]
- else:
- return None
- elif chunker.data_type.value in [item.value for item in SpecialDataType]:
- # These types don't contain indirect references.
- # Through custom logic, they can be attributed to a source and be updated.
- if chunker.data_type == DataType.QNA_PAIR:
- # QNA_PAIRs update the answer if the question already exists.
- where = {"question": src[0]}
- if self.config.id is not None:
- where.update({"app_id": self.config.id})
-
- existing_embeddings = self.db.get(
- where=where,
- limit=1,
- )
- if len(existing_embeddings.get("metadatas", [])) > 0:
- return existing_embeddings["metadatas"][0]["doc_id"]
- else:
- return None
- else:
- raise NotImplementedError(
- f"SpecialDataType {chunker.data_type} must have a custom logic to check for existing data"
- )
- else:
- raise TypeError(
- f"{chunker.data_type} is type {type(chunker.data_type)}. "
- "When it should be DirectDataType, IndirectDataType or SpecialDataType."
- )
-
- def _load_and_embed(
- self,
- loader: BaseLoader,
- chunker: BaseChunker,
- src: Any,
- metadata: Optional[dict[str, Any]] = None,
- source_hash: Optional[str] = None,
- add_config: Optional[AddConfig] = None,
- dry_run=False,
- **kwargs: Optional[dict[str, Any]],
- ):
- """
- Loads the data from the given URL, chunks it, and adds it to database.
-
- :param loader: The loader to use to load the data.
- :type loader: BaseLoader
- :param chunker: The chunker to use to chunk the data.
- :type chunker: BaseChunker
- :param src: The data to be handled by the loader. Can be a URL for
- remote sources or local content for local loaders.
- :type src: Any
- :param metadata: Metadata associated with the data source.
- :type metadata: dict[str, Any], optional
- :param source_hash: Hexadecimal hash of the source.
- :type source_hash: str, optional
- :param add_config: The `AddConfig` instance to use as configuration options.
- :type add_config: AddConfig, optional
- :param dry_run: A dry run returns chunks and doesn't update DB.
- :type dry_run: bool, defaults to False
- :return: (list) documents (embedded text), (list) metadata, (list) ids, (int) number of chunks
- """
- existing_doc_id = self._get_existing_doc_id(chunker=chunker, src=src)
- app_id = self.config.id if self.config is not None else None
-
- # Create chunks
- embeddings_data = chunker.create_chunks(loader, src, app_id=app_id, config=add_config.chunker, **kwargs)
- # spread chunking results
- documents = embeddings_data["documents"]
- metadatas = embeddings_data["metadatas"]
- ids = embeddings_data["ids"]
- new_doc_id = embeddings_data["doc_id"]
-
- if existing_doc_id and existing_doc_id == new_doc_id:
- logger.info("Doc content has not changed. Skipping creating chunks and embeddings")
- return [], [], [], 0
-
- # this means that doc content has changed.
- if existing_doc_id and existing_doc_id != new_doc_id:
- logger.info("Doc content has changed. Recomputing chunks and embeddings intelligently.")
- self.db.delete({"doc_id": existing_doc_id})
-
- # get existing ids, and discard doc if any common id exist.
- where = {"url": src}
- if chunker.data_type == DataType.JSON and is_valid_json_string(src):
- url = hashlib.sha256((src).encode("utf-8")).hexdigest()
- where = {"url": url}
-
- # if data type is qna_pair, we check for question
- if chunker.data_type == DataType.QNA_PAIR:
- where = {"question": src[0]}
-
- if self.config.id is not None:
- where["app_id"] = self.config.id
-
- db_result = self.db.get(ids=ids, where=where) # optional filter
- existing_ids = set(db_result["ids"])
- if len(existing_ids):
- data_dict = {id: (doc, meta) for id, doc, meta in zip(ids, documents, metadatas)}
- data_dict = {id: value for id, value in data_dict.items() if id not in existing_ids}
-
- if not data_dict:
- src_copy = src
- if len(src_copy) > 50:
- src_copy = src[:50] + "..."
- logger.info(f"All data from {src_copy} already exists in the database.")
- # Make sure to return a matching return type
- return [], [], [], 0
-
- ids = list(data_dict.keys())
- documents, metadatas = zip(*data_dict.values())
-
- # Loop though all metadatas and add extras.
- new_metadatas = []
- for m in metadatas:
- # Add app id in metadatas so that they can be queried on later
- if self.config.id:
- m["app_id"] = self.config.id
-
- # Add hashed source
- m["hash"] = source_hash
-
- # Note: Metadata is the function argument
- if metadata:
- # Spread whatever is in metadata into the new object.
- m.update(metadata)
-
- new_metadatas.append(m)
- metadatas = new_metadatas
-
- if dry_run:
- return list(documents), metadatas, ids, 0
-
- # Count before, to calculate a delta in the end.
- chunks_before_addition = self.db.count()
-
- # Filter out empty documents and ensure they meet the API requirements
- valid_documents = [doc for doc in documents if doc and isinstance(doc, str)]
-
- documents = valid_documents
-
- # Chunk documents into batches of 2048 and handle each batch
- # helps wigth large loads of embeddings that hit OpenAI limits
- document_batches = [documents[i : i + 2048] for i in range(0, len(documents), 2048)]
- metadata_batches = [metadatas[i : i + 2048] for i in range(0, len(metadatas), 2048)]
- id_batches = [ids[i : i + 2048] for i in range(0, len(ids), 2048)]
- for batch_docs, batch_meta, batch_ids in zip(document_batches, metadata_batches, id_batches):
- try:
- # Add only valid batches
- if batch_docs:
- self.db.add(documents=batch_docs, metadatas=batch_meta, ids=batch_ids, **kwargs)
- except Exception as e:
- logger.info(f"Failed to add batch due to a bad request: {e}")
- # Handle the error, e.g., by logging, retrying, or skipping
- pass
-
- count_new_chunks = self.db.count() - chunks_before_addition
- logger.info(f"Successfully saved {str(src)[:100]} ({chunker.data_type}). New chunks count: {count_new_chunks}")
-
- return list(documents), metadatas, ids, count_new_chunks
-
- @staticmethod
- def _format_result(results):
- return [
- (Document(page_content=result[0], metadata=result[1] or {}), result[2])
- for result in zip(
- results["documents"][0],
- results["metadatas"][0],
- results["distances"][0],
- )
- ]
-
- def _retrieve_from_database(
- self,
- input_query: str,
- config: Optional[BaseLlmConfig] = None,
- where=None,
- citations: bool = False,
- **kwargs: Optional[dict[str, Any]],
- ) -> Union[list[tuple[str, str, str]], list[str]]:
- """
- Queries the vector database based on the given input query.
- Gets relevant doc based on the query
-
- :param input_query: The query to use.
- :type input_query: str
- :param config: The query configuration, defaults to None
- :type config: Optional[BaseLlmConfig], optional
- :param where: A dictionary of key-value pairs to filter the database results, defaults to None
- :type where: _type_, optional
- :param citations: A boolean to indicate if db should fetch citation source
- :type citations: bool
- :return: List of contents of the document that matched your query
- :rtype: list[str]
- """
- query_config = config or self.llm.config
- if where is not None:
- where = where
- else:
- where = {}
- if query_config is not None and query_config.where is not None:
- where = query_config.where
-
- if self.config.id is not None:
- where.update({"app_id": self.config.id})
-
- contexts = self.db.query(
- input_query=input_query,
- n_results=query_config.number_documents,
- where=where,
- citations=citations,
- **kwargs,
- )
-
- return contexts
-
- def query(
- self,
- input_query: str,
- config: BaseLlmConfig = None,
- dry_run=False,
- where: Optional[dict] = None,
- citations: bool = False,
- **kwargs: dict[str, Any],
- ) -> Union[tuple[str, list[tuple[str, dict]]], str, dict[str, Any]]:
- """
- Queries the vector database based on the given input query.
- Gets relevant doc based on the query and then passes it to an
- LLM as context to get the answer.
-
- :param input_query: The query to use.
- :type input_query: str
- :param config: The `BaseLlmConfig` instance to use as configuration options. This is used for one method call.
- To persistently use a config, declare it during app init., defaults to None
- :type config: BaseLlmConfig, optional
- :param dry_run: A dry run does everything except send the resulting prompt to
- the LLM. The purpose is to test the prompt, not the response., defaults to False
- :type dry_run: bool, optional
- :param where: A dictionary of key-value pairs to filter the database results., defaults to None
- :type where: dict[str, str], optional
- :param citations: A boolean to indicate if db should fetch citation source
- :type citations: bool
- :param kwargs: To read more params for the query function. Ex. we use citations boolean
- param to return context along with the answer
- :type kwargs: dict[str, Any]
- :return: The answer to the query, with citations if the citation flag is True
- or the dry run result
- :rtype: str, if citations is False and token_usage is False, otherwise if citations is true then
- tuple[str, list[tuple[str,str,str]]] and if token_usage is true then
- tuple[str, list[tuple[str,str,str]], dict[str, Any]]
- """
- contexts = self._retrieve_from_database(
- input_query=input_query, config=config, where=where, citations=citations, **kwargs
- )
- if citations and len(contexts) > 0 and isinstance(contexts[0], tuple):
- contexts_data_for_llm_query = list(map(lambda x: x[0], contexts))
- else:
- contexts_data_for_llm_query = contexts
-
- if self.cache_config is not None:
- logger.info("Cache enabled. Checking cache...")
- answer = adapt(
- llm_handler=self.llm.query,
- cache_data_convert=gptcache_data_convert,
- update_cache_callback=gptcache_update_cache_callback,
- session=get_gptcache_session(session_id=self.config.id),
- input_query=input_query,
- contexts=contexts_data_for_llm_query,
- config=config,
- dry_run=dry_run,
- )
- else:
- if self.llm.config.token_usage:
- answer, token_info = self.llm.query(
- input_query=input_query, contexts=contexts_data_for_llm_query, config=config, dry_run=dry_run
- )
- else:
- answer = self.llm.query(
- input_query=input_query, contexts=contexts_data_for_llm_query, config=config, dry_run=dry_run
- )
-
- # Send anonymous telemetry
- if self.config.collect_metrics:
- self.telemetry.capture(event_name="query", properties=self._telemetry_props)
-
- if citations:
- if self.llm.config.token_usage:
- return {"answer": answer, "contexts": contexts, "usage": token_info}
- return answer, contexts
- if self.llm.config.token_usage:
- return {"answer": answer, "usage": token_info}
-
- logger.warning(
- "Starting from v0.1.125 the return type of query method will be changed to tuple containing `answer`."
- )
- return answer
-
- def chat(
- self,
- input_query: str,
- config: Optional[BaseLlmConfig] = None,
- dry_run=False,
- session_id: str = "default",
- where: Optional[dict[str, str]] = None,
- citations: bool = False,
- **kwargs: dict[str, Any],
- ) -> Union[tuple[str, list[tuple[str, dict]]], str, dict[str, Any]]:
- """
- Queries the vector database on the given input query.
- Gets relevant doc based on the query and then passes it to an
- LLM as context to get the answer.
-
- Maintains the whole conversation in memory.
-
- :param input_query: The query to use.
- :type input_query: str
- :param config: The `BaseLlmConfig` instance to use as configuration options. This is used for one method call.
- To persistently use a config, declare it during app init., defaults to None
- :type config: BaseLlmConfig, optional
- :param dry_run: A dry run does everything except send the resulting prompt to
- the LLM. The purpose is to test the prompt, not the response., defaults to False
- :type dry_run: bool, optional
- :param session_id: The session id to use for chat history, defaults to 'default'.
- :type session_id: str, optional
- :param where: A dictionary of key-value pairs to filter the database results., defaults to None
- :type where: dict[str, str], optional
- :param citations: A boolean to indicate if db should fetch citation source
- :type citations: bool
- :param kwargs: To read more params for the query function. Ex. we use citations boolean
- param to return context along with the answer
- :type kwargs: dict[str, Any]
- :return: The answer to the query, with citations if the citation flag is True
- or the dry run result
- :rtype: str, if citations is False and token_usage is False, otherwise if citations is true then
- tuple[str, list[tuple[str,str,str]]] and if token_usage is true then
- tuple[str, list[tuple[str,str,str]], dict[str, Any]]
- """
- contexts = self._retrieve_from_database(
- input_query=input_query, config=config, where=where, citations=citations, **kwargs
- )
- if citations and len(contexts) > 0 and isinstance(contexts[0], tuple):
- contexts_data_for_llm_query = list(map(lambda x: x[0], contexts))
- else:
- contexts_data_for_llm_query = contexts
-
- memories = None
- if self.mem0_memory:
- memories = self.mem0_memory.search(
- query=input_query, agent_id=self.config.id, user_id=session_id, limit=self.memory_config.top_k
- )
-
- # Update the history beforehand so that we can handle multiple chat sessions in the same python session
- self.llm.update_history(app_id=self.config.id, session_id=session_id)
-
- if self.cache_config is not None:
- logger.debug("Cache enabled. Checking cache...")
- cache_id = f"{session_id}--{self.config.id}"
- answer = adapt(
- llm_handler=self.llm.chat,
- cache_data_convert=gptcache_data_convert,
- update_cache_callback=gptcache_update_cache_callback,
- session=get_gptcache_session(session_id=cache_id),
- input_query=input_query,
- contexts=contexts_data_for_llm_query,
- config=config,
- dry_run=dry_run,
- )
- else:
- logger.debug("Cache disabled. Running chat without cache.")
- if self.llm.config.token_usage:
- answer, token_info = self.llm.query(
- input_query=input_query,
- contexts=contexts_data_for_llm_query,
- config=config,
- dry_run=dry_run,
- memories=memories,
- )
- else:
- answer = self.llm.query(
- input_query=input_query,
- contexts=contexts_data_for_llm_query,
- config=config,
- dry_run=dry_run,
- memories=memories,
- )
-
- # Add to Mem0 memory if enabled
- # Adding answer here because it would be much useful than input question itself
- if self.mem0_memory:
- self.mem0_memory.add(data=answer, agent_id=self.config.id, user_id=session_id)
-
- # add conversation in memory
- self.llm.add_history(self.config.id, input_query, answer, session_id=session_id)
-
- # Send anonymous telemetry
- if self.config.collect_metrics:
- self.telemetry.capture(event_name="chat", properties=self._telemetry_props)
-
- if citations:
- if self.llm.config.token_usage:
- return {"answer": answer, "contexts": contexts, "usage": token_info}
- return answer, contexts
- if self.llm.config.token_usage:
- return {"answer": answer, "usage": token_info}
-
- logger.warning(
- "Starting from v0.1.125 the return type of query method will be changed to tuple containing `answer`."
- )
- return answer
-
- def search(self, query, num_documents=3, where=None, raw_filter=None, namespace=None):
- """
- Search for similar documents related to the query in the vector database.
-
- Args:
- query (str): The query to use.
- num_documents (int, optional): Number of similar documents to fetch. Defaults to 3.
- where (dict[str, any], optional): Filter criteria for the search.
- raw_filter (dict[str, any], optional): Advanced raw filter criteria for the search.
- namespace (str, optional): The namespace to search in. Defaults to None.
-
- Raises:
- ValueError: If both `raw_filter` and `where` are used simultaneously.
-
- Returns:
- list[dict]: A list of dictionaries, each containing the 'context' and 'metadata' of a document.
- """
- # Send anonymous telemetry
- if self.config.collect_metrics:
- self.telemetry.capture(event_name="search", properties=self._telemetry_props)
-
- if raw_filter and where:
- raise ValueError("You can't use both `raw_filter` and `where` together.")
-
- filter_type = "raw_filter" if raw_filter else "where"
- filter_criteria = raw_filter if raw_filter else where
-
- params = {
- "input_query": query,
- "n_results": num_documents,
- "citations": True,
- "app_id": self.config.id,
- "namespace": namespace,
- filter_type: filter_criteria,
- }
-
- return [{"context": c[0], "metadata": c[1]} for c in self.db.query(**params)]
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
-
- Using `app.db.set_collection_name` method is preferred to this.
-
- :param name: Name of the collection.
- :type name: str
- """
- self.db.set_collection_name(name)
- # Create the collection if it does not exist
- self.db._get_or_create_collection(name)
- # TODO: Check whether it is necessary to assign to the `self.collection` attribute,
- # since the main purpose is the creation.
-
- def reset(self):
- """
- Resets the database. Deletes all embeddings irreversibly.
- `App` does not have to be reinitialized after using this method.
- """
- try:
- self.db_session.query(DataSource).filter_by(app_id=self.config.id).delete()
- self.db_session.query(ChatHistory).filter_by(app_id=self.config.id).delete()
- self.db_session.commit()
- except Exception as e:
- logger.error(f"Error deleting data sources: {e}")
- self.db_session.rollback()
- return None
- self.db.reset()
- self.delete_all_chat_history(app_id=self.config.id)
- # Send anonymous telemetry
- if self.config.collect_metrics:
- self.telemetry.capture(event_name="reset", properties=self._telemetry_props)
-
- def get_history(
- self,
- num_rounds: int = 10,
- display_format: bool = True,
- session_id: Optional[str] = "default",
- fetch_all: bool = False,
- ):
- history = self.llm.memory.get(
- app_id=self.config.id,
- session_id=session_id,
- num_rounds=num_rounds,
- display_format=display_format,
- fetch_all=fetch_all,
- )
- return history
-
- def delete_session_chat_history(self, session_id: str = "default"):
- self.llm.memory.delete(app_id=self.config.id, session_id=session_id)
- self.llm.update_history(app_id=self.config.id)
-
- def delete_all_chat_history(self, app_id: str):
- self.llm.memory.delete(app_id=app_id)
- self.llm.update_history(app_id=app_id)
-
- def delete(self, source_id: str):
- """
- Deletes the data from the database.
- :param source_hash: The hash of the source.
- :type source_hash: str
- """
- try:
- self.db_session.query(DataSource).filter_by(hash=source_id, app_id=self.config.id).delete()
- self.db_session.commit()
- except Exception as e:
- logger.error(f"Error deleting data sources: {e}")
- self.db_session.rollback()
- return None
- self.db.delete(where={"hash": source_id})
- logger.info(f"Successfully deleted {source_id}")
- # Send anonymous telemetry
- if self.config.collect_metrics:
- self.telemetry.capture(event_name="delete", properties=self._telemetry_props)
diff --git a/embedchain/embedchain/embedder/__init__.py b/embedchain/embedchain/embedder/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/embedder/aws_bedrock.py b/embedchain/embedchain/embedder/aws_bedrock.py
deleted file mode 100644
index 235fc3dab..000000000
--- a/embedchain/embedchain/embedder/aws_bedrock.py
+++ /dev/null
@@ -1,31 +0,0 @@
-from typing import Optional
-
-try:
- from langchain_aws import BedrockEmbeddings
-except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for AWSBedrock are not installed." "Please install with `pip install langchain_aws`"
- ) from None
-
-from embedchain.config.embedder.aws_bedrock import AWSBedrockEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-
-class AWSBedrockEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[AWSBedrockEmbedderConfig] = None):
- super().__init__(config)
-
- if self.config.model is None or self.config.model == "amazon.titan-embed-text-v2:0":
- self.config.model = "amazon.titan-embed-text-v2:0" # Default model if not specified
- vector_dimension = self.config.vector_dimension or VectorDimensions.AMAZON_TITAN_V2.value
- elif self.config.model == "amazon.titan-embed-text-v1":
- vector_dimension = VectorDimensions.AMAZON_TITAN_V1.value
- else:
- vector_dimension = self.config.vector_dimension
-
- embeddings = BedrockEmbeddings(model_id=self.config.model, model_kwargs=self.config.model_kwargs)
- embedding_fn = BaseEmbedder._langchain_default_concept(embeddings)
-
- self.set_embedding_fn(embedding_fn=embedding_fn)
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/azure_openai.py b/embedchain/embedchain/embedder/azure_openai.py
deleted file mode 100644
index 71802ad87..000000000
--- a/embedchain/embedchain/embedder/azure_openai.py
+++ /dev/null
@@ -1,26 +0,0 @@
-from typing import Optional
-
-from langchain_openai import AzureOpenAIEmbeddings
-
-from embedchain.config import BaseEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-
-class AzureOpenAIEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- super().__init__(config=config)
-
- if self.config.model is None:
- self.config.model = "text-embedding-ada-002"
-
- embeddings = AzureOpenAIEmbeddings(
- deployment=self.config.deployment_name,
- http_client=self.config.http_client,
- http_async_client=self.config.http_async_client,
- )
- embedding_fn = BaseEmbedder._langchain_default_concept(embeddings)
-
- self.set_embedding_fn(embedding_fn=embedding_fn)
- vector_dimension = self.config.vector_dimension or VectorDimensions.OPENAI.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/base.py b/embedchain/embedchain/embedder/base.py
deleted file mode 100644
index 7f65477bf..000000000
--- a/embedchain/embedchain/embedder/base.py
+++ /dev/null
@@ -1,90 +0,0 @@
-from collections.abc import Callable
-from typing import Any, Optional
-
-from embedchain.config.embedder.base import BaseEmbedderConfig
-
-try:
- from chromadb.api.types import Embeddable, EmbeddingFunction, Embeddings
-except RuntimeError:
- from embedchain.utils.misc import use_pysqlite3
-
- use_pysqlite3()
- from chromadb.api.types import Embeddable, EmbeddingFunction, Embeddings
-
-
-class EmbeddingFunc(EmbeddingFunction):
- def __init__(self, embedding_fn: Callable[[list[str]], list[str]]):
- self.embedding_fn = embedding_fn
-
- def __call__(self, input: Embeddable) -> Embeddings:
- return self.embedding_fn(input)
-
-
-class BaseEmbedder:
- """
- Class that manages everything regarding embeddings. Including embedding function, loaders and chunkers.
-
- Embedding functions and vector dimensions are set based on the child class you choose.
- To manually overwrite you can use this classes `set_...` methods.
- """
-
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- """
- Initialize the embedder class.
-
- :param config: embedder configuration option class, defaults to None
- :type config: Optional[BaseEmbedderConfig], optional
- """
- if config is None:
- self.config = BaseEmbedderConfig()
- else:
- self.config = config
- self.vector_dimension: int
-
- def set_embedding_fn(self, embedding_fn: Callable[[list[str]], list[str]]):
- """
- Set or overwrite the embedding function to be used by the database to store and retrieve documents.
-
- :param embedding_fn: Function to be used to generate embeddings.
- :type embedding_fn: Callable[[list[str]], list[str]]
- :raises ValueError: Embedding function is not callable.
- """
- if not hasattr(embedding_fn, "__call__"):
- raise ValueError("Embedding function is not a function")
- self.embedding_fn = embedding_fn
-
- def set_vector_dimension(self, vector_dimension: int):
- """
- Set or overwrite the vector dimension size
-
- :param vector_dimension: vector dimension size
- :type vector_dimension: int
- """
- if not isinstance(vector_dimension, int):
- raise TypeError("vector dimension must be int")
- self.vector_dimension = vector_dimension
-
- @staticmethod
- def _langchain_default_concept(embeddings: Any):
- """
- Langchains default function layout for embeddings.
-
- :param embeddings: Langchain embeddings
- :type embeddings: Any
- :return: embedding function
- :rtype: Callable
- """
-
- return EmbeddingFunc(embeddings.embed_documents)
-
- def to_embeddings(self, data: str, **_):
- """
- Convert data to embeddings
-
- :param data: data to convert to embeddings
- :type data: str
- :return: embeddings
- :rtype: list[float]
- """
- embeddings = self.embedding_fn([data])
- return embeddings[0]
diff --git a/embedchain/embedchain/embedder/clarifai.py b/embedchain/embedchain/embedder/clarifai.py
deleted file mode 100644
index 8f0bb2fe4..000000000
--- a/embedchain/embedchain/embedder/clarifai.py
+++ /dev/null
@@ -1,52 +0,0 @@
-import os
-from typing import Optional, Union
-
-from chromadb import EmbeddingFunction, Embeddings
-
-from embedchain.config import BaseEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-
-
-class ClarifaiEmbeddingFunction(EmbeddingFunction):
- def __init__(self, config: BaseEmbedderConfig) -> None:
- super().__init__()
- try:
- from clarifai.client.input import Inputs
- from clarifai.client.model import Model
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for ClarifaiEmbeddingFunction are not installed."
- 'Please install with `pip install --upgrade "embedchain[clarifai]"`'
- ) from None
- self.config = config
- self.api_key = config.api_key or os.getenv("CLARIFAI_PAT")
- self.model = config.model
- self.model_obj = Model(url=self.model, pat=self.api_key)
- self.input_obj = Inputs(pat=self.api_key)
-
- def __call__(self, input: Union[str, list[str]]) -> Embeddings:
- if isinstance(input, str):
- input = [input]
-
- batch_size = 32
- embeddings = []
- try:
- for i in range(0, len(input), batch_size):
- batch = input[i : i + batch_size]
- input_batch = [
- self.input_obj.get_text_input(input_id=str(id), raw_text=inp) for id, inp in enumerate(batch)
- ]
- response = self.model_obj.predict(input_batch)
- embeddings.extend([list(output.data.embeddings[0].vector) for output in response.outputs])
- except Exception as e:
- print(f"Predict failed, exception: {e}")
-
- return embeddings
-
-
-class ClarifaiEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- super().__init__(config)
-
- embedding_func = ClarifaiEmbeddingFunction(config=self.config)
- self.set_embedding_fn(embedding_fn=embedding_func)
diff --git a/embedchain/embedchain/embedder/cohere.py b/embedchain/embedchain/embedder/cohere.py
deleted file mode 100644
index 489ba97f3..000000000
--- a/embedchain/embedchain/embedder/cohere.py
+++ /dev/null
@@ -1,19 +0,0 @@
-from typing import Optional
-
-from langchain_cohere.embeddings import CohereEmbeddings
-
-from embedchain.config import BaseEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-
-class CohereEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- super().__init__(config=config)
-
- embeddings = CohereEmbeddings(model=self.config.model)
- embedding_fn = BaseEmbedder._langchain_default_concept(embeddings)
- self.set_embedding_fn(embedding_fn=embedding_fn)
-
- vector_dimension = self.config.vector_dimension or VectorDimensions.COHERE.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/google.py b/embedchain/embedchain/embedder/google.py
deleted file mode 100644
index c0be83500..000000000
--- a/embedchain/embedchain/embedder/google.py
+++ /dev/null
@@ -1,38 +0,0 @@
-from typing import Optional, Union
-
-import google.generativeai as genai
-from chromadb import EmbeddingFunction, Embeddings
-
-from embedchain.config.embedder.google import GoogleAIEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-
-class GoogleAIEmbeddingFunction(EmbeddingFunction):
- def __init__(self, config: Optional[GoogleAIEmbedderConfig] = None) -> None:
- super().__init__()
- self.config = config or GoogleAIEmbedderConfig()
-
- def __call__(self, input: Union[list[str], str]) -> Embeddings:
- model = self.config.model
- title = self.config.title
- task_type = self.config.task_type
- if isinstance(input, str):
- input_ = [input]
- else:
- input_ = input
- data = genai.embed_content(model=model, content=input_, task_type=task_type, title=title)
- embeddings = data["embedding"]
- if isinstance(input_, str):
- embeddings = [embeddings]
- return embeddings
-
-
-class GoogleAIEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[GoogleAIEmbedderConfig] = None):
- super().__init__(config)
- embedding_fn = GoogleAIEmbeddingFunction(config=config)
- self.set_embedding_fn(embedding_fn=embedding_fn)
-
- vector_dimension = self.config.vector_dimension or VectorDimensions.GOOGLE_AI.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/gpt4all.py b/embedchain/embedchain/embedder/gpt4all.py
deleted file mode 100644
index 83123f499..000000000
--- a/embedchain/embedchain/embedder/gpt4all.py
+++ /dev/null
@@ -1,23 +0,0 @@
-from typing import Optional
-
-from embedchain.config import BaseEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-
-class GPT4AllEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- super().__init__(config=config)
-
- from langchain_community.embeddings import (
- GPT4AllEmbeddings as LangchainGPT4AllEmbeddings,
- )
-
- model_name = self.config.model or "all-MiniLM-L6-v2-f16.gguf"
- gpt4all_kwargs = {'allow_download': 'True'}
- embeddings = LangchainGPT4AllEmbeddings(model_name=model_name, gpt4all_kwargs=gpt4all_kwargs)
- embedding_fn = BaseEmbedder._langchain_default_concept(embeddings)
- self.set_embedding_fn(embedding_fn=embedding_fn)
-
- vector_dimension = self.config.vector_dimension or VectorDimensions.GPT4ALL.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/huggingface.py b/embedchain/embedchain/embedder/huggingface.py
deleted file mode 100644
index 062208e77..000000000
--- a/embedchain/embedchain/embedder/huggingface.py
+++ /dev/null
@@ -1,40 +0,0 @@
-import os
-from typing import Optional
-
-from langchain_community.embeddings import HuggingFaceEmbeddings
-
-try:
- from langchain_huggingface import HuggingFaceEndpointEmbeddings
-except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for HuggingFaceHub are not installed."
- "Please install with `pip install langchain_huggingface`"
- ) from None
-
-from embedchain.config import BaseEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-
-class HuggingFaceEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- super().__init__(config=config)
-
- if self.config.endpoint:
- if not self.config.api_key and "HUGGINGFACE_ACCESS_TOKEN" not in os.environ:
- raise ValueError(
- "Please set the HUGGINGFACE_ACCESS_TOKEN environment variable or pass API Key in the config."
- )
-
- embeddings = HuggingFaceEndpointEmbeddings(
- model=self.config.endpoint,
- huggingfacehub_api_token=self.config.api_key or os.getenv("HUGGINGFACE_ACCESS_TOKEN"),
- )
- else:
- embeddings = HuggingFaceEmbeddings(model_name=self.config.model, model_kwargs=self.config.model_kwargs)
-
- embedding_fn = BaseEmbedder._langchain_default_concept(embeddings)
- self.set_embedding_fn(embedding_fn=embedding_fn)
-
- vector_dimension = self.config.vector_dimension or VectorDimensions.HUGGING_FACE.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/mistralai.py b/embedchain/embedchain/embedder/mistralai.py
deleted file mode 100644
index 29db72ae0..000000000
--- a/embedchain/embedchain/embedder/mistralai.py
+++ /dev/null
@@ -1,46 +0,0 @@
-import os
-from typing import Optional, Union
-
-from chromadb import EmbeddingFunction, Embeddings
-
-from embedchain.config import BaseEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-
-class MistralAIEmbeddingFunction(EmbeddingFunction):
- def __init__(self, config: BaseEmbedderConfig) -> None:
- super().__init__()
- try:
- from langchain_mistralai import MistralAIEmbeddings
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for MistralAI are not installed."
- 'Please install with `pip install --upgrade "embedchain[mistralai]"`'
- ) from None
- self.config = config
- api_key = self.config.api_key or os.getenv("MISTRAL_API_KEY")
- self.client = MistralAIEmbeddings(mistral_api_key=api_key)
- self.client.model = self.config.model
-
- def __call__(self, input: Union[list[str], str]) -> Embeddings:
- if isinstance(input, str):
- input_ = [input]
- else:
- input_ = input
- response = self.client.embed_documents(input_)
- return response
-
-
-class MistralAIEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- super().__init__(config)
-
- if self.config.model is None:
- self.config.model = "mistral-embed"
-
- embedding_fn = MistralAIEmbeddingFunction(config=self.config)
- self.set_embedding_fn(embedding_fn=embedding_fn)
-
- vector_dimension = self.config.vector_dimension or VectorDimensions.MISTRAL_AI.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/nvidia.py b/embedchain/embedchain/embedder/nvidia.py
deleted file mode 100644
index 5a499037f..000000000
--- a/embedchain/embedchain/embedder/nvidia.py
+++ /dev/null
@@ -1,28 +0,0 @@
-import logging
-import os
-from typing import Optional
-
-from langchain_nvidia_ai_endpoints import NVIDIAEmbeddings
-
-from embedchain.config import BaseEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-logger = logging.getLogger(__name__)
-
-
-class NvidiaEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- if "NVIDIA_API_KEY" not in os.environ:
- raise ValueError("NVIDIA_API_KEY environment variable must be set")
-
- super().__init__(config=config)
-
- model = self.config.model or "nvolveqa_40k"
- logger.info(f"Using NVIDIA embedding model: {model}")
- embedder = NVIDIAEmbeddings(model=model)
- embedding_fn = BaseEmbedder._langchain_default_concept(embedder)
- self.set_embedding_fn(embedding_fn=embedding_fn)
-
- vector_dimension = self.config.vector_dimension or VectorDimensions.NVIDIA_AI.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/ollama.py b/embedchain/embedchain/embedder/ollama.py
deleted file mode 100644
index 9e4ada473..000000000
--- a/embedchain/embedchain/embedder/ollama.py
+++ /dev/null
@@ -1,32 +0,0 @@
-import logging
-from typing import Optional
-
-try:
- from ollama import Client
-except ImportError:
- raise ImportError("Ollama Embedder requires extra dependencies. Install with `pip install ollama`") from None
-
-from langchain_community.embeddings import OllamaEmbeddings
-
-from embedchain.config import OllamaEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-logger = logging.getLogger(__name__)
-
-
-class OllamaEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[OllamaEmbedderConfig] = None):
- super().__init__(config=config)
-
- client = Client(host=config.base_url)
- local_models = client.list()["models"]
- if not any(model.get("name") == self.config.model for model in local_models):
- logger.info(f"Pulling {self.config.model} from Ollama!")
- client.pull(self.config.model)
- embeddings = OllamaEmbeddings(model=self.config.model, base_url=config.base_url)
- embedding_fn = BaseEmbedder._langchain_default_concept(embeddings)
- self.set_embedding_fn(embedding_fn=embedding_fn)
-
- vector_dimension = self.config.vector_dimension or VectorDimensions.OLLAMA.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/openai.py b/embedchain/embedchain/embedder/openai.py
deleted file mode 100644
index e14a1aa70..000000000
--- a/embedchain/embedchain/embedder/openai.py
+++ /dev/null
@@ -1,43 +0,0 @@
-import os
-import warnings
-from typing import Optional
-
-from chromadb.utils.embedding_functions import OpenAIEmbeddingFunction
-
-from embedchain.config import BaseEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-
-class OpenAIEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- super().__init__(config=config)
-
- if self.config.model is None:
- self.config.model = "text-embedding-ada-002"
-
- api_key = self.config.api_key or os.environ["OPENAI_API_KEY"]
- api_base = (
- self.config.api_base
- or os.environ.get("OPENAI_API_BASE")
- or os.getenv("OPENAI_BASE_URL")
- or "https://api.openai.com/v1"
- )
- if os.environ.get("OPENAI_API_BASE"):
- warnings.warn(
- "The environment variable 'OPENAI_API_BASE' is deprecated and will be removed in the 0.1.140. "
- "Please use 'OPENAI_BASE_URL' instead.",
- DeprecationWarning
- )
-
- if api_key is None and os.getenv("OPENAI_ORGANIZATION") is None:
- raise ValueError("OPENAI_API_KEY or OPENAI_ORGANIZATION environment variables not provided") # noqa:E501
- embedding_fn = OpenAIEmbeddingFunction(
- api_key=api_key,
- api_base=api_base,
- organization_id=os.getenv("OPENAI_ORGANIZATION"),
- model_name=self.config.model,
- )
- self.set_embedding_fn(embedding_fn=embedding_fn)
- vector_dimension = self.config.vector_dimension or VectorDimensions.OPENAI.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/embedder/vertexai.py b/embedchain/embedchain/embedder/vertexai.py
deleted file mode 100644
index 1f3331dc6..000000000
--- a/embedchain/embedchain/embedder/vertexai.py
+++ /dev/null
@@ -1,19 +0,0 @@
-from typing import Optional
-
-from langchain_google_vertexai import VertexAIEmbeddings
-
-from embedchain.config import BaseEmbedderConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.models import VectorDimensions
-
-
-class VertexAIEmbedder(BaseEmbedder):
- def __init__(self, config: Optional[BaseEmbedderConfig] = None):
- super().__init__(config=config)
-
- embeddings = VertexAIEmbeddings(model_name=config.model)
- embedding_fn = BaseEmbedder._langchain_default_concept(embeddings)
- self.set_embedding_fn(embedding_fn=embedding_fn)
-
- vector_dimension = self.config.vector_dimension or VectorDimensions.VERTEX_AI.value
- self.set_vector_dimension(vector_dimension=vector_dimension)
diff --git a/embedchain/embedchain/evaluation/__init__.py b/embedchain/embedchain/evaluation/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/evaluation/base.py b/embedchain/embedchain/evaluation/base.py
deleted file mode 100644
index 4528e7689..000000000
--- a/embedchain/embedchain/evaluation/base.py
+++ /dev/null
@@ -1,29 +0,0 @@
-from abc import ABC, abstractmethod
-
-from embedchain.utils.evaluation import EvalData
-
-
-class BaseMetric(ABC):
- """Base class for a metric.
-
- This class provides a common interface for all metrics.
- """
-
- def __init__(self, name: str = "base_metric"):
- """
- Initialize the BaseMetric.
- """
- self.name = name
-
- @abstractmethod
- def evaluate(self, dataset: list[EvalData]):
- """
- Abstract method to evaluate the dataset.
-
- This method should be implemented by subclasses to perform the actual
- evaluation on the dataset.
-
- :param dataset: dataset to evaluate
- :type dataset: list[EvalData]
- """
- raise NotImplementedError()
diff --git a/embedchain/embedchain/evaluation/metrics/__init__.py b/embedchain/embedchain/evaluation/metrics/__init__.py
deleted file mode 100644
index 95f579005..000000000
--- a/embedchain/embedchain/evaluation/metrics/__init__.py
+++ /dev/null
@@ -1,3 +0,0 @@
-from .answer_relevancy import AnswerRelevance # noqa: F401
-from .context_relevancy import ContextRelevance # noqa: F401
-from .groundedness import Groundedness # noqa: F401
diff --git a/embedchain/embedchain/evaluation/metrics/answer_relevancy.py b/embedchain/embedchain/evaluation/metrics/answer_relevancy.py
deleted file mode 100644
index 3e5c3859e..000000000
--- a/embedchain/embedchain/evaluation/metrics/answer_relevancy.py
+++ /dev/null
@@ -1,95 +0,0 @@
-import concurrent.futures
-import logging
-import os
-from string import Template
-from typing import Optional
-
-import numpy as np
-from openai import OpenAI
-from tqdm import tqdm
-
-from embedchain.config.evaluation.base import AnswerRelevanceConfig
-from embedchain.evaluation.base import BaseMetric
-from embedchain.utils.evaluation import EvalData, EvalMetric
-
-logger = logging.getLogger(__name__)
-
-
-class AnswerRelevance(BaseMetric):
- """
- Metric for evaluating the relevance of answers.
- """
-
- def __init__(self, config: Optional[AnswerRelevanceConfig] = AnswerRelevanceConfig()):
- super().__init__(name=EvalMetric.ANSWER_RELEVANCY.value)
- self.config = config
- api_key = self.config.api_key or os.getenv("OPENAI_API_KEY")
- if not api_key:
- raise ValueError("API key not found. Set 'OPENAI_API_KEY' or pass it in the config.")
- self.client = OpenAI(api_key=api_key)
-
- def _generate_prompt(self, data: EvalData) -> str:
- """
- Generates a prompt based on the provided data.
- """
- return Template(self.config.prompt).substitute(
- num_gen_questions=self.config.num_gen_questions, answer=data.answer
- )
-
- def _generate_questions(self, prompt: str) -> list[str]:
- """
- Generates questions from the prompt.
- """
- response = self.client.chat.completions.create(
- model=self.config.model,
- messages=[{"role": "user", "content": prompt}],
- )
- return response.choices[0].message.content.strip().split("\n")
-
- def _generate_embedding(self, question: str) -> np.ndarray:
- """
- Generates the embedding for a question.
- """
- response = self.client.embeddings.create(
- input=question,
- model=self.config.embedder,
- )
- return np.array(response.data[0].embedding)
-
- def _compute_similarity(self, original: np.ndarray, generated: np.ndarray) -> float:
- """
- Computes the cosine similarity between two embeddings.
- """
- original = original.reshape(1, -1)
- norm = np.linalg.norm(original) * np.linalg.norm(generated, axis=1)
- return np.dot(generated, original.T).flatten() / norm
-
- def _compute_score(self, data: EvalData) -> float:
- """
- Computes the relevance score for a given data item.
- """
- prompt = self._generate_prompt(data)
- generated_questions = self._generate_questions(prompt)
- original_embedding = self._generate_embedding(data.question)
- generated_embeddings = np.array([self._generate_embedding(q) for q in generated_questions])
- similarities = self._compute_similarity(original_embedding, generated_embeddings)
- return np.mean(similarities)
-
- def evaluate(self, dataset: list[EvalData]) -> float:
- """
- Evaluates the dataset and returns the average answer relevance score.
- """
- results = []
-
- with concurrent.futures.ThreadPoolExecutor() as executor:
- future_to_data = {executor.submit(self._compute_score, data): data for data in dataset}
- for future in tqdm(
- concurrent.futures.as_completed(future_to_data), total=len(dataset), desc="Evaluating Answer Relevancy"
- ):
- data = future_to_data[future]
- try:
- results.append(future.result())
- except Exception as e:
- logger.error(f"Error evaluating answer relevancy for {data}: {e}")
-
- return np.mean(results) if results else 0.0
diff --git a/embedchain/embedchain/evaluation/metrics/context_relevancy.py b/embedchain/embedchain/evaluation/metrics/context_relevancy.py
deleted file mode 100644
index f821713fa..000000000
--- a/embedchain/embedchain/evaluation/metrics/context_relevancy.py
+++ /dev/null
@@ -1,69 +0,0 @@
-import concurrent.futures
-import os
-from string import Template
-from typing import Optional
-
-import numpy as np
-import pysbd
-from openai import OpenAI
-from tqdm import tqdm
-
-from embedchain.config.evaluation.base import ContextRelevanceConfig
-from embedchain.evaluation.base import BaseMetric
-from embedchain.utils.evaluation import EvalData, EvalMetric
-
-
-class ContextRelevance(BaseMetric):
- """
- Metric for evaluating the relevance of context in a dataset.
- """
-
- def __init__(self, config: Optional[ContextRelevanceConfig] = ContextRelevanceConfig()):
- super().__init__(name=EvalMetric.CONTEXT_RELEVANCY.value)
- self.config = config
- api_key = self.config.api_key or os.getenv("OPENAI_API_KEY")
- if not api_key:
- raise ValueError("API key not found. Set 'OPENAI_API_KEY' or pass it in the config.")
- self.client = OpenAI(api_key=api_key)
- self._sbd = pysbd.Segmenter(language=self.config.language, clean=False)
-
- def _sentence_segmenter(self, text: str) -> list[str]:
- """
- Segments the given text into sentences.
- """
- return self._sbd.segment(text)
-
- def _compute_score(self, data: EvalData) -> float:
- """
- Computes the context relevance score for a given data item.
- """
- original_context = "\n".join(data.contexts)
- prompt = Template(self.config.prompt).substitute(context=original_context, question=data.question)
- response = self.client.chat.completions.create(
- model=self.config.model, messages=[{"role": "user", "content": prompt}]
- )
- useful_context = response.choices[0].message.content.strip()
- useful_context_sentences = self._sentence_segmenter(useful_context)
- original_context_sentences = self._sentence_segmenter(original_context)
-
- if not original_context_sentences:
- return 0.0
- return len(useful_context_sentences) / len(original_context_sentences)
-
- def evaluate(self, dataset: list[EvalData]) -> float:
- """
- Evaluates the dataset and returns the average context relevance score.
- """
- scores = []
-
- with concurrent.futures.ThreadPoolExecutor() as executor:
- futures = [executor.submit(self._compute_score, data) for data in dataset]
- for future in tqdm(
- concurrent.futures.as_completed(futures), total=len(dataset), desc="Evaluating Context Relevancy"
- ):
- try:
- scores.append(future.result())
- except Exception as e:
- print(f"Error during evaluation: {e}")
-
- return np.mean(scores) if scores else 0.0
diff --git a/embedchain/embedchain/evaluation/metrics/groundedness.py b/embedchain/embedchain/evaluation/metrics/groundedness.py
deleted file mode 100644
index 86f3f320e..000000000
--- a/embedchain/embedchain/evaluation/metrics/groundedness.py
+++ /dev/null
@@ -1,104 +0,0 @@
-import concurrent.futures
-import logging
-import os
-from string import Template
-from typing import Optional
-
-import numpy as np
-from openai import OpenAI
-from tqdm import tqdm
-
-from embedchain.config.evaluation.base import GroundednessConfig
-from embedchain.evaluation.base import BaseMetric
-from embedchain.utils.evaluation import EvalData, EvalMetric
-
-logger = logging.getLogger(__name__)
-
-
-class Groundedness(BaseMetric):
- """
- Metric for groundedness of answer from the given contexts.
- """
-
- def __init__(self, config: Optional[GroundednessConfig] = None):
- super().__init__(name=EvalMetric.GROUNDEDNESS.value)
- self.config = config or GroundednessConfig()
- api_key = self.config.api_key or os.getenv("OPENAI_API_KEY")
- if not api_key:
- raise ValueError("Please set the OPENAI_API_KEY environment variable or pass the `api_key` in config.")
- self.client = OpenAI(api_key=api_key)
-
- def _generate_answer_claim_prompt(self, data: EvalData) -> str:
- """
- Generate the prompt for the given data.
- """
- prompt = Template(self.config.answer_claims_prompt).substitute(question=data.question, answer=data.answer)
- return prompt
-
- def _get_claim_statements(self, prompt: str) -> np.ndarray:
- """
- Get claim statements from the answer.
- """
- response = self.client.chat.completions.create(
- model=self.config.model,
- messages=[{"role": "user", "content": f"{prompt}"}],
- )
- result = response.choices[0].message.content.strip()
- claim_statements = np.array([statement for statement in result.split("\n") if statement])
- return claim_statements
-
- def _generate_claim_inference_prompt(self, data: EvalData, claim_statements: list[str]) -> str:
- """
- Generate the claim inference prompt for the given data and claim statements.
- """
- prompt = Template(self.config.claims_inference_prompt).substitute(
- context="\n".join(data.contexts), claim_statements="\n".join(claim_statements)
- )
- return prompt
-
- def _get_claim_verdict_scores(self, prompt: str) -> np.ndarray:
- """
- Get verdicts for claim statements.
- """
- response = self.client.chat.completions.create(
- model=self.config.model,
- messages=[{"role": "user", "content": f"{prompt}"}],
- )
- result = response.choices[0].message.content.strip()
- claim_verdicts = result.split("\n")
- verdict_score_map = {"1": 1, "0": 0, "-1": np.nan}
- verdict_scores = np.array([verdict_score_map[verdict] for verdict in claim_verdicts])
- return verdict_scores
-
- def _compute_score(self, data: EvalData) -> float:
- """
- Compute the groundedness score for a single data point.
- """
- answer_claims_prompt = self._generate_answer_claim_prompt(data)
- claim_statements = self._get_claim_statements(answer_claims_prompt)
-
- claim_inference_prompt = self._generate_claim_inference_prompt(data, claim_statements)
- verdict_scores = self._get_claim_verdict_scores(claim_inference_prompt)
- return np.sum(verdict_scores) / claim_statements.size
-
- def evaluate(self, dataset: list[EvalData]):
- """
- Evaluate the dataset and returns the average groundedness score.
- """
- results = []
-
- with concurrent.futures.ThreadPoolExecutor() as executor:
- future_to_data = {executor.submit(self._compute_score, data): data for data in dataset}
- for future in tqdm(
- concurrent.futures.as_completed(future_to_data),
- total=len(future_to_data),
- desc="Evaluating Groundedness",
- ):
- data = future_to_data[future]
- try:
- score = future.result()
- results.append(score)
- except Exception as e:
- logger.error(f"Error while evaluating groundedness for data point {data}: {e}")
-
- return np.mean(results) if results else 0.0
diff --git a/embedchain/embedchain/factory.py b/embedchain/embedchain/factory.py
deleted file mode 100644
index 69636286c..000000000
--- a/embedchain/embedchain/factory.py
+++ /dev/null
@@ -1,122 +0,0 @@
-import importlib
-
-
-def load_class(class_type):
- module_path, class_name = class_type.rsplit(".", 1)
- module = importlib.import_module(module_path)
- return getattr(module, class_name)
-
-
-class LlmFactory:
- provider_to_class = {
- "anthropic": "embedchain.llm.anthropic.AnthropicLlm",
- "azure_openai": "embedchain.llm.azure_openai.AzureOpenAILlm",
- "cohere": "embedchain.llm.cohere.CohereLlm",
- "together": "embedchain.llm.together.TogetherLlm",
- "gpt4all": "embedchain.llm.gpt4all.GPT4ALLLlm",
- "ollama": "embedchain.llm.ollama.OllamaLlm",
- "huggingface": "embedchain.llm.huggingface.HuggingFaceLlm",
- "jina": "embedchain.llm.jina.JinaLlm",
- "llama2": "embedchain.llm.llama2.Llama2Llm",
- "openai": "embedchain.llm.openai.OpenAILlm",
- "vertexai": "embedchain.llm.vertex_ai.VertexAILlm",
- "google": "embedchain.llm.google.GoogleLlm",
- "aws_bedrock": "embedchain.llm.aws_bedrock.AWSBedrockLlm",
- "mistralai": "embedchain.llm.mistralai.MistralAILlm",
- "clarifai": "embedchain.llm.clarifai.ClarifaiLlm",
- "groq": "embedchain.llm.groq.GroqLlm",
- "nvidia": "embedchain.llm.nvidia.NvidiaLlm",
- "vllm": "embedchain.llm.vllm.VLLM",
- }
- provider_to_config_class = {
- "embedchain": "embedchain.config.llm.base.BaseLlmConfig",
- "openai": "embedchain.config.llm.base.BaseLlmConfig",
- "anthropic": "embedchain.config.llm.base.BaseLlmConfig",
- }
-
- @classmethod
- def create(cls, provider_name, config_data):
- class_type = cls.provider_to_class.get(provider_name)
- # Default to embedchain base config if the provider is not in the config map
- config_name = "embedchain" if provider_name not in cls.provider_to_config_class else provider_name
- config_class_type = cls.provider_to_config_class.get(config_name)
- if class_type:
- llm_class = load_class(class_type)
- llm_config_class = load_class(config_class_type)
- return llm_class(config=llm_config_class(**config_data))
- else:
- raise ValueError(f"Unsupported Llm provider: {provider_name}")
-
-
-class EmbedderFactory:
- provider_to_class = {
- "azure_openai": "embedchain.embedder.azure_openai.AzureOpenAIEmbedder",
- "gpt4all": "embedchain.embedder.gpt4all.GPT4AllEmbedder",
- "huggingface": "embedchain.embedder.huggingface.HuggingFaceEmbedder",
- "openai": "embedchain.embedder.openai.OpenAIEmbedder",
- "vertexai": "embedchain.embedder.vertexai.VertexAIEmbedder",
- "google": "embedchain.embedder.google.GoogleAIEmbedder",
- "mistralai": "embedchain.embedder.mistralai.MistralAIEmbedder",
- "clarifai": "embedchain.embedder.clarifai.ClarifaiEmbedder",
- "nvidia": "embedchain.embedder.nvidia.NvidiaEmbedder",
- "cohere": "embedchain.embedder.cohere.CohereEmbedder",
- "ollama": "embedchain.embedder.ollama.OllamaEmbedder",
- "aws_bedrock": "embedchain.embedder.aws_bedrock.AWSBedrockEmbedder",
- }
- provider_to_config_class = {
- "azure_openai": "embedchain.config.embedder.base.BaseEmbedderConfig",
- "google": "embedchain.config.embedder.google.GoogleAIEmbedderConfig",
- "gpt4all": "embedchain.config.embedder.base.BaseEmbedderConfig",
- "huggingface": "embedchain.config.embedder.base.BaseEmbedderConfig",
- "clarifai": "embedchain.config.embedder.base.BaseEmbedderConfig",
- "openai": "embedchain.config.embedder.base.BaseEmbedderConfig",
- "ollama": "embedchain.config.embedder.ollama.OllamaEmbedderConfig",
- "aws_bedrock": "embedchain.config.embedder.aws_bedrock.AWSBedrockEmbedderConfig",
- }
-
- @classmethod
- def create(cls, provider_name, config_data):
- class_type = cls.provider_to_class.get(provider_name)
- # Default to openai config if the provider is not in the config map
- config_name = "openai" if provider_name not in cls.provider_to_config_class else provider_name
- config_class_type = cls.provider_to_config_class.get(config_name)
- if class_type:
- embedder_class = load_class(class_type)
- embedder_config_class = load_class(config_class_type)
- return embedder_class(config=embedder_config_class(**config_data))
- else:
- raise ValueError(f"Unsupported Embedder provider: {provider_name}")
-
-
-class VectorDBFactory:
- provider_to_class = {
- "chroma": "embedchain.vectordb.chroma.ChromaDB",
- "elasticsearch": "embedchain.vectordb.elasticsearch.ElasticsearchDB",
- "opensearch": "embedchain.vectordb.opensearch.OpenSearchDB",
- "lancedb": "embedchain.vectordb.lancedb.LanceDB",
- "pinecone": "embedchain.vectordb.pinecone.PineconeDB",
- "qdrant": "embedchain.vectordb.qdrant.QdrantDB",
- "weaviate": "embedchain.vectordb.weaviate.WeaviateDB",
- "zilliz": "embedchain.vectordb.zilliz.ZillizVectorDB",
- }
- provider_to_config_class = {
- "chroma": "embedchain.config.vector_db.chroma.ChromaDbConfig",
- "elasticsearch": "embedchain.config.vector_db.elasticsearch.ElasticsearchDBConfig",
- "opensearch": "embedchain.config.vector_db.opensearch.OpenSearchDBConfig",
- "lancedb": "embedchain.config.vector_db.lancedb.LanceDBConfig",
- "pinecone": "embedchain.config.vector_db.pinecone.PineconeDBConfig",
- "qdrant": "embedchain.config.vector_db.qdrant.QdrantDBConfig",
- "weaviate": "embedchain.config.vector_db.weaviate.WeaviateDBConfig",
- "zilliz": "embedchain.config.vector_db.zilliz.ZillizDBConfig",
- }
-
- @classmethod
- def create(cls, provider_name, config_data):
- class_type = cls.provider_to_class.get(provider_name)
- config_class_type = cls.provider_to_config_class.get(provider_name)
- if class_type:
- embedder_class = load_class(class_type)
- embedder_config_class = load_class(config_class_type)
- return embedder_class(config=embedder_config_class(**config_data))
- else:
- raise ValueError(f"Unsupported Embedder provider: {provider_name}")
diff --git a/embedchain/embedchain/helpers/__init__.py b/embedchain/embedchain/helpers/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/helpers/callbacks.py b/embedchain/embedchain/helpers/callbacks.py
deleted file mode 100644
index 4847e0fea..000000000
--- a/embedchain/embedchain/helpers/callbacks.py
+++ /dev/null
@@ -1,73 +0,0 @@
-import queue
-from typing import Any, Union
-
-from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
-from langchain.schema import LLMResult
-
-STOP_ITEM = "[END]"
-"""
-This is a special item that is used to signal the end of the stream.
-"""
-
-
-class StreamingStdOutCallbackHandlerYield(StreamingStdOutCallbackHandler):
- """
- This is a callback handler that yields the tokens as they are generated.
- For a usage example, see the :func:`generate` function below.
- """
-
- q: queue.Queue
- """
- The queue to write the tokens to as they are generated.
- """
-
- def __init__(self, q: queue.Queue) -> None:
- """
- Initialize the callback handler.
- q: The queue to write the tokens to as they are generated.
- """
- super().__init__()
- self.q = q
-
- def on_llm_start(self, serialized: dict[str, Any], prompts: list[str], **kwargs: Any) -> None:
- """Run when LLM starts running."""
- with self.q.mutex:
- self.q.queue.clear()
-
- def on_llm_new_token(self, token: str, **kwargs: Any) -> None:
- """Run on new LLM token. Only available when streaming is enabled."""
- self.q.put(token)
-
- def on_llm_end(self, response: LLMResult, **kwargs: Any) -> None:
- """Run when LLM ends running."""
- self.q.put(STOP_ITEM)
-
- def on_llm_error(self, error: Union[Exception, KeyboardInterrupt], **kwargs: Any) -> None:
- """Run when LLM errors."""
- self.q.put("%s: %s" % (type(error).__name__, str(error)))
- self.q.put(STOP_ITEM)
-
-
-def generate(rq: queue.Queue):
- """
- This is a generator that yields the items in the queue until it reaches the stop item.
-
- Usage example:
- ```
- def askQuestion(callback_fn: StreamingStdOutCallbackHandlerYield):
- llm = OpenAI(streaming=True, callbacks=[callback_fn])
- return llm.invoke(prompt="Write a poem about a tree.")
-
- @app.route("/", methods=["GET"])
- def generate_output():
- q = Queue()
- callback_fn = StreamingStdOutCallbackHandlerYield(q)
- threading.Thread(target=askQuestion, args=(callback_fn,)).start()
- return Response(generate(q), mimetype="text/event-stream")
- ```
- """
- while True:
- result: str = rq.get()
- if result == STOP_ITEM or result is None:
- break
- yield result
diff --git a/embedchain/embedchain/helpers/json_serializable.py b/embedchain/embedchain/helpers/json_serializable.py
deleted file mode 100644
index 656bb44bc..000000000
--- a/embedchain/embedchain/helpers/json_serializable.py
+++ /dev/null
@@ -1,198 +0,0 @@
-import json
-import logging
-from string import Template
-from typing import Any, Type, TypeVar, Union
-
-T = TypeVar("T", bound="JSONSerializable")
-
-# NOTE: Through inheritance, all of our classes should be children of JSONSerializable. (highest level)
-# NOTE: The @register_deserializable decorator should be added to all user facing child classes. (lowest level)
-
-logger = logging.getLogger(__name__)
-
-
-def register_deserializable(cls: Type[T]) -> Type[T]:
- """
- A class decorator to register a class as deserializable.
-
- When a class is decorated with @register_deserializable, it becomes
- a part of the set of classes that the JSONSerializable class can
- deserialize.
-
- Deserialization is in essence loading attributes from a json file.
- This decorator is a security measure put in place to make sure that
- you don't load attributes that were initially part of another class.
-
- Example:
- @register_deserializable
- class ChildClass(JSONSerializable):
- def __init__(self, ...):
- # initialization logic
-
- Args:
- cls (Type): The class to be registered.
-
- Returns:
- Type: The same class, after registration.
- """
- JSONSerializable._register_class_as_deserializable(cls)
- return cls
-
-
-class JSONSerializable:
- """
- A class to represent a JSON serializable object.
-
- This class provides methods to serialize and deserialize objects,
- as well as to save serialized objects to a file and load them back.
- """
-
- _deserializable_classes = set() # Contains classes that are whitelisted for deserialization.
-
- def serialize(self) -> str:
- """
- Serialize the object to a JSON-formatted string.
-
- Returns:
- str: A JSON string representation of the object.
- """
- try:
- return json.dumps(self, default=self._auto_encoder, ensure_ascii=False)
- except Exception as e:
- logger.error(f"Serialization error: {e}")
- return "{}"
-
- @classmethod
- def deserialize(cls, json_str: str) -> Any:
- """
- Deserialize a JSON-formatted string to an object.
- If it fails, a default class is returned instead.
- Note: This *returns* an instance, it's not automatically loaded on the calling class.
-
- Example:
- app = App.deserialize(json_str)
-
- Args:
- json_str (str): A JSON string representation of an object.
-
- Returns:
- Object: The deserialized object.
- """
- try:
- return json.loads(json_str, object_hook=cls._auto_decoder)
- except Exception as e:
- logger.error(f"Deserialization error: {e}")
- # Return a default instance in case of failure
- return cls()
-
- @staticmethod
- def _auto_encoder(obj: Any) -> Union[dict[str, Any], None]:
- """
- Automatically encode an object for JSON serialization.
-
- Args:
- obj (Object): The object to be encoded.
-
- Returns:
- dict: A dictionary representation of the object.
- """
- if hasattr(obj, "__dict__"):
- dct = {}
- for key, value in obj.__dict__.items():
- try:
- # Recursive: If the value is an instance of a subclass of JSONSerializable,
- # serialize it using the JSONSerializable serialize method.
- if isinstance(value, JSONSerializable):
- serialized_value = value.serialize()
- # The value is stored as a serialized string.
- dct[key] = json.loads(serialized_value)
- # Custom rules (subclass is not json serializable by default)
- elif isinstance(value, Template):
- dct[key] = {"__type__": "Template", "data": value.template}
- # Future custom types we can follow a similar pattern
- # elif isinstance(value, SomeOtherType):
- # dct[key] = {
- # "__type__": "SomeOtherType",
- # "data": value.some_method()
- # }
- # NOTE: Keep in mind that this logic needs to be applied to the decoder too.
- else:
- json.dumps(value) # Try to serialize the value.
- dct[key] = value
- except TypeError:
- pass # If it fails, simply pass to skip this key-value pair of the dictionary.
-
- dct["__class__"] = obj.__class__.__name__
- return dct
- raise TypeError(f"Object of type {type(obj)} is not JSON serializable")
-
- @classmethod
- def _auto_decoder(cls, dct: dict[str, Any]) -> Any:
- """
- Automatically decode a dictionary to an object during JSON deserialization.
-
- Args:
- dct (dict): The dictionary representation of an object.
-
- Returns:
- Object: The decoded object or the original dictionary if decoding is not possible.
- """
- class_name = dct.pop("__class__", None)
- if class_name:
- if not hasattr(cls, "_deserializable_classes"): # Additional safety check
- raise AttributeError(f"`{class_name}` has no registry of allowed deserializations.")
- if class_name not in {cl.__name__ for cl in cls._deserializable_classes}:
- raise KeyError(f"Deserialization of class `{class_name}` is not allowed.")
- target_class = next((cl for cl in cls._deserializable_classes if cl.__name__ == class_name), None)
- if target_class:
- obj = target_class.__new__(target_class)
- for key, value in dct.items():
- if isinstance(value, dict) and "__type__" in value:
- if value["__type__"] == "Template":
- value = Template(value["data"])
- # For future custom types we can follow a similar pattern
- # elif value["__type__"] == "SomeOtherType":
- # value = SomeOtherType.some_constructor(value["data"])
- default_value = getattr(target_class, key, None)
- setattr(obj, key, value or default_value)
- return obj
- return dct
-
- def save_to_file(self, filename: str) -> None:
- """
- Save the serialized object to a file.
-
- Args:
- filename (str): The path to the file where the object should be saved.
- """
- with open(filename, "w", encoding="utf-8") as f:
- f.write(self.serialize())
-
- @classmethod
- def load_from_file(cls, filename: str) -> Any:
- """
- Load and deserialize an object from a file.
-
- Args:
- filename (str): The path to the file from which the object should be loaded.
-
- Returns:
- Object: The deserialized object.
- """
- with open(filename, "r", encoding="utf-8") as f:
- json_str = f.read()
- return cls.deserialize(json_str)
-
- @classmethod
- def _register_class_as_deserializable(cls, target_class: Type[T]) -> None:
- """
- Register a class as deserializable. This is a classmethod and globally shared.
-
- This method adds the target class to the set of classes that
- can be deserialized. This is a security measure to ensure only
- whitelisted classes are deserialized.
-
- Args:
- target_class (Type): The class to be registered.
- """
- cls._deserializable_classes.add(target_class)
diff --git a/embedchain/embedchain/llm/__init__.py b/embedchain/embedchain/llm/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/llm/anthropic.py b/embedchain/embedchain/llm/anthropic.py
deleted file mode 100644
index b5a90a6d5..000000000
--- a/embedchain/embedchain/llm/anthropic.py
+++ /dev/null
@@ -1,59 +0,0 @@
-import logging
-import os
-from typing import Any, Optional
-
-try:
- from langchain_anthropic import ChatAnthropic
-except ImportError:
- raise ImportError("Please install the langchain-anthropic package by running `pip install langchain-anthropic`.")
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class AnthropicLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config=config)
- if not self.config.api_key and "ANTHROPIC_API_KEY" not in os.environ:
- raise ValueError("Please set the ANTHROPIC_API_KEY environment variable or pass it in the config.")
-
- def get_llm_model_answer(self, prompt) -> tuple[str, Optional[dict[str, Any]]]:
- if self.config.token_usage:
- response, token_info = self._get_answer(prompt, self.config)
- model_name = "anthropic/" + self.config.model
- if model_name not in self.config.model_pricing_map:
- raise ValueError(
- f"Model {model_name} not found in `model_prices_and_context_window.json`. \
- You can disable token usage by setting `token_usage` to False."
- )
- total_cost = (
- self.config.model_pricing_map[model_name]["input_cost_per_token"] * token_info["input_tokens"]
- ) + self.config.model_pricing_map[model_name]["output_cost_per_token"] * token_info["output_tokens"]
- response_token_info = {
- "prompt_tokens": token_info["input_tokens"],
- "completion_tokens": token_info["output_tokens"],
- "total_tokens": token_info["input_tokens"] + token_info["output_tokens"],
- "total_cost": round(total_cost, 10),
- "cost_currency": "USD",
- }
- return response, response_token_info
- return self._get_answer(prompt, self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> str:
- api_key = config.api_key or os.getenv("ANTHROPIC_API_KEY")
- chat = ChatAnthropic(anthropic_api_key=api_key, temperature=config.temperature, model_name=config.model)
-
- if config.max_tokens and config.max_tokens != 1000:
- logger.warning("Config option `max_tokens` is not supported by this model.")
-
- messages = BaseLlm._get_messages(prompt, system_prompt=config.system_prompt)
-
- chat_response = chat.invoke(messages)
- if config.token_usage:
- return chat_response.content, chat_response.response_metadata["token_usage"]
- return chat_response.content
diff --git a/embedchain/embedchain/llm/aws_bedrock.py b/embedchain/embedchain/llm/aws_bedrock.py
deleted file mode 100644
index 7f916268b..000000000
--- a/embedchain/embedchain/llm/aws_bedrock.py
+++ /dev/null
@@ -1,57 +0,0 @@
-import os
-from typing import Optional
-
-try:
- from langchain_aws import BedrockLLM
-except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for AWSBedrock are not installed." "Please install with `pip install langchain_aws`"
- ) from None
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class AWSBedrockLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config)
-
- def get_llm_model_answer(self, prompt) -> str:
- response = self._get_answer(prompt, self.config)
- return response
-
- def _get_answer(self, prompt: str, config: BaseLlmConfig) -> str:
- try:
- import boto3
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for AWSBedrock are not installed."
- "Please install with `pip install boto3==1.34.20`."
- ) from None
-
- self.boto_client = boto3.client(
- "bedrock-runtime", os.environ.get("AWS_REGION", os.environ.get("AWS_DEFAULT_REGION", "us-east-1"))
- )
-
- kwargs = {
- "model_id": config.model or "amazon.titan-text-express-v1",
- "client": self.boto_client,
- "model_kwargs": config.model_kwargs
- or {
- "temperature": config.temperature,
- },
- }
-
- if config.stream:
- from langchain.callbacks.streaming_stdout import (
- StreamingStdOutCallbackHandler,
- )
-
- kwargs["streaming"] = True
- kwargs["callbacks"] = [StreamingStdOutCallbackHandler()]
-
- llm = BedrockLLM(**kwargs)
-
- return llm.invoke(prompt)
diff --git a/embedchain/embedchain/llm/azure_openai.py b/embedchain/embedchain/llm/azure_openai.py
deleted file mode 100644
index c219270ac..000000000
--- a/embedchain/embedchain/llm/azure_openai.py
+++ /dev/null
@@ -1,42 +0,0 @@
-import logging
-from typing import Optional
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class AzureOpenAILlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config=config)
-
- def get_llm_model_answer(self, prompt):
- return self._get_answer(prompt=prompt, config=self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> str:
- from langchain_openai import AzureChatOpenAI
-
- if not config.deployment_name:
- raise ValueError("Deployment name must be provided for Azure OpenAI")
-
- chat = AzureChatOpenAI(
- deployment_name=config.deployment_name,
- openai_api_version=str(config.api_version) if config.api_version else "2024-02-01",
- model_name=config.model or "gpt-4o-mini",
- temperature=config.temperature,
- max_tokens=config.max_tokens,
- streaming=config.stream,
- http_client=config.http_client,
- http_async_client=config.http_async_client,
- )
-
- if config.top_p and config.top_p != 1:
- logger.warning("Config option `top_p` is not supported by this model.")
-
- messages = BaseLlm._get_messages(prompt, system_prompt=config.system_prompt)
-
- return chat.invoke(messages).content
diff --git a/embedchain/embedchain/llm/base.py b/embedchain/embedchain/llm/base.py
deleted file mode 100644
index ace4bb79b..000000000
--- a/embedchain/embedchain/llm/base.py
+++ /dev/null
@@ -1,350 +0,0 @@
-import logging
-import os
-from collections.abc import Generator
-from typing import Any, Optional
-
-from langchain.schema import BaseMessage as LCBaseMessage
-
-from embedchain.config import BaseLlmConfig
-from embedchain.config.llm.base import (
- DEFAULT_PROMPT,
- DEFAULT_PROMPT_WITH_HISTORY_TEMPLATE,
- DEFAULT_PROMPT_WITH_MEM0_MEMORY_TEMPLATE,
- DOCS_SITE_PROMPT_TEMPLATE,
-)
-from embedchain.constants import SQLITE_PATH
-from embedchain.core.db.database import init_db, setup_engine
-from embedchain.helpers.json_serializable import JSONSerializable
-from embedchain.memory.base import ChatHistory
-from embedchain.memory.message import ChatMessage
-
-logger = logging.getLogger(__name__)
-
-
-class BaseLlm(JSONSerializable):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- """Initialize a base LLM class
-
- :param config: LLM configuration option class, defaults to None
- :type config: Optional[BaseLlmConfig], optional
- """
- if config is None:
- self.config = BaseLlmConfig()
- else:
- self.config = config
-
- # Initialize the metadata db for the app here since llmfactory needs it for initialization of
- # the llm memory
- setup_engine(database_uri=os.environ.get("EMBEDCHAIN_DB_URI", f"sqlite:///{SQLITE_PATH}"))
- init_db()
-
- self.memory = ChatHistory()
- self.is_docs_site_instance = False
- self.history: Any = None
-
- def get_llm_model_answer(self):
- """
- Usually implemented by child class
- """
- raise NotImplementedError
-
- def set_history(self, history: Any):
- """
- Provide your own history.
- Especially interesting for the query method, which does not internally manage conversation history.
-
- :param history: History to set
- :type history: Any
- """
- self.history = history
-
- def update_history(self, app_id: str, session_id: str = "default"):
- """Update class history attribute with history in memory (for chat method)"""
- chat_history = self.memory.get(app_id=app_id, session_id=session_id, num_rounds=10)
- self.set_history([str(history) for history in chat_history])
-
- def add_history(
- self,
- app_id: str,
- question: str,
- answer: str,
- metadata: Optional[dict[str, Any]] = None,
- session_id: str = "default",
- ):
- chat_message = ChatMessage()
- chat_message.add_user_message(question, metadata=metadata)
- chat_message.add_ai_message(answer, metadata=metadata)
- self.memory.add(app_id=app_id, chat_message=chat_message, session_id=session_id)
- self.update_history(app_id=app_id, session_id=session_id)
-
- def _format_history(self) -> str:
- """Format history to be used in prompt
-
- :return: Formatted history
- :rtype: str
- """
- return "\n".join(self.history)
-
- def _format_memories(self, memories: list[dict]) -> str:
- """Format memories to be used in prompt
-
- :param memories: Memories to format
- :type memories: list[dict]
- :return: Formatted memories
- :rtype: str
- """
- return "\n".join([memory["text"] for memory in memories])
-
- def generate_prompt(self, input_query: str, contexts: list[str], **kwargs: dict[str, Any]) -> str:
- """
- Generates a prompt based on the given query and context, ready to be
- passed to an LLM
-
- :param input_query: The query to use.
- :type input_query: str
- :param contexts: List of similar documents to the query used as context.
- :type contexts: list[str]
- :return: The prompt
- :rtype: str
- """
- context_string = " | ".join(contexts)
- web_search_result = kwargs.get("web_search_result", "")
- memories = kwargs.get("memories", None)
- if web_search_result:
- context_string = self._append_search_and_context(context_string, web_search_result)
-
- prompt_contains_history = self.config._validate_prompt_history(self.config.prompt)
- if prompt_contains_history:
- prompt = self.config.prompt.substitute(
- context=context_string, query=input_query, history=self._format_history() or "No history"
- )
- elif self.history and not prompt_contains_history:
- # History is present, but not included in the prompt.
- # check if it's the default prompt without history
- if (
- not self.config._validate_prompt_history(self.config.prompt)
- and self.config.prompt.template == DEFAULT_PROMPT
- ):
- if memories:
- # swap in the template with Mem0 memory template
- prompt = DEFAULT_PROMPT_WITH_MEM0_MEMORY_TEMPLATE.substitute(
- context=context_string,
- query=input_query,
- history=self._format_history(),
- memories=self._format_memories(memories),
- )
- else:
- # swap in the template with history
- prompt = DEFAULT_PROMPT_WITH_HISTORY_TEMPLATE.substitute(
- context=context_string, query=input_query, history=self._format_history()
- )
- else:
- # If we can't swap in the default, we still proceed but tell users that the history is ignored.
- logger.warning(
- "Your bot contains a history, but prompt does not include `$history` key. History is ignored."
- )
- prompt = self.config.prompt.substitute(context=context_string, query=input_query)
- else:
- # basic use case, no history.
- prompt = self.config.prompt.substitute(context=context_string, query=input_query)
- return prompt
-
- @staticmethod
- def _append_search_and_context(context: str, web_search_result: str) -> str:
- """Append web search context to existing context
-
- :param context: Existing context
- :type context: str
- :param web_search_result: Web search result
- :type web_search_result: str
- :return: Concatenated web search result
- :rtype: str
- """
- return f"{context}\nWeb Search Result: {web_search_result}"
-
- def get_answer_from_llm(self, prompt: str):
- """
- Gets an answer based on the given query and context by passing it
- to an LLM.
-
- :param prompt: Gets an answer based on the given query and context by passing it to an LLM.
- :type prompt: str
- :return: The answer.
- :rtype: _type_
- """
- return self.get_llm_model_answer(prompt)
-
- @staticmethod
- def access_search_and_get_results(input_query: str):
- """
- Search the internet for additional context
-
- :param input_query: search query
- :type input_query: str
- :return: Search results
- :rtype: Unknown
- """
- try:
- from langchain.tools import DuckDuckGoSearchRun
- except ImportError:
- raise ImportError(
- "Searching requires extra dependencies. Install with `pip install duckduckgo-search==6.1.5`"
- ) from None
- search = DuckDuckGoSearchRun()
- logger.info(f"Access search to get answers for {input_query}")
- return search.run(input_query)
-
- @staticmethod
- def _stream_response(answer: Any, token_info: Optional[dict[str, Any]] = None) -> Generator[Any, Any, None]:
- """Generator to be used as streaming response
-
- :param answer: Answer chunk from llm
- :type answer: Any
- :yield: Answer chunk from llm
- :rtype: Generator[Any, Any, None]
- """
- streamed_answer = ""
- for chunk in answer:
- streamed_answer = streamed_answer + chunk
- yield chunk
- logger.info(f"Answer: {streamed_answer}")
- if token_info:
- logger.info(f"Token Info: {token_info}")
-
- def query(self, input_query: str, contexts: list[str], config: BaseLlmConfig = None, dry_run=False, memories=None):
- """
- Queries the vector database based on the given input query.
- Gets relevant doc based on the query and then passes it to an
- LLM as context to get the answer.
-
- :param input_query: The query to use.
- :type input_query: str
- :param contexts: Embeddings retrieved from the database to be used as context.
- :type contexts: list[str]
- :param config: The `BaseLlmConfig` instance to use as configuration options. This is used for one method call.
- To persistently use a config, declare it during app init., defaults to None
- :type config: Optional[BaseLlmConfig], optional
- :param dry_run: A dry run does everything except send the resulting prompt to
- the LLM. The purpose is to test the prompt, not the response., defaults to False
- :type dry_run: bool, optional
- :return: The answer to the query or the dry run result
- :rtype: str
- """
- try:
- if config:
- # A config instance passed to this method will only be applied temporarily, for one call.
- # So we will save the previous config and restore it at the end of the execution.
- # For this we use the serializer.
- prev_config = self.config.serialize()
- self.config = config
-
- if config is not None and config.query_type == "Images":
- return contexts
-
- if self.is_docs_site_instance:
- self.config.prompt = DOCS_SITE_PROMPT_TEMPLATE
- self.config.number_documents = 5
- k = {}
- if self.config.online:
- k["web_search_result"] = self.access_search_and_get_results(input_query)
- k["memories"] = memories
- prompt = self.generate_prompt(input_query, contexts, **k)
- logger.info(f"Prompt: {prompt}")
- if dry_run:
- return prompt
-
- if self.config.token_usage:
- answer, token_info = self.get_answer_from_llm(prompt)
- else:
- answer = self.get_answer_from_llm(prompt)
- if isinstance(answer, str):
- logger.info(f"Answer: {answer}")
- if self.config.token_usage:
- return answer, token_info
- return answer
- else:
- if self.config.token_usage:
- return self._stream_response(answer, token_info)
- return self._stream_response(answer)
- finally:
- if config:
- # Restore previous config
- self.config: BaseLlmConfig = BaseLlmConfig.deserialize(prev_config)
-
- def chat(
- self, input_query: str, contexts: list[str], config: BaseLlmConfig = None, dry_run=False, session_id: str = None
- ):
- """
- Queries the vector database on the given input query.
- Gets relevant doc based on the query and then passes it to an
- LLM as context to get the answer.
-
- Maintains the whole conversation in memory.
-
- :param input_query: The query to use.
- :type input_query: str
- :param contexts: Embeddings retrieved from the database to be used as context.
- :type contexts: list[str]
- :param config: The `BaseLlmConfig` instance to use as configuration options. This is used for one method call.
- To persistently use a config, declare it during app init., defaults to None
- :type config: Optional[BaseLlmConfig], optional
- :param dry_run: A dry run does everything except send the resulting prompt to
- the LLM. The purpose is to test the prompt, not the response., defaults to False
- :type dry_run: bool, optional
- :param session_id: Session ID to use for the conversation, defaults to None
- :type session_id: str, optional
- :return: The answer to the query or the dry run result
- :rtype: str
- """
- try:
- if config:
- # A config instance passed to this method will only be applied temporarily, for one call.
- # So we will save the previous config and restore it at the end of the execution.
- # For this we use the serializer.
- prev_config = self.config.serialize()
- self.config = config
-
- if self.is_docs_site_instance:
- self.config.prompt = DOCS_SITE_PROMPT_TEMPLATE
- self.config.number_documents = 5
- k = {}
- if self.config.online:
- k["web_search_result"] = self.access_search_and_get_results(input_query)
-
- prompt = self.generate_prompt(input_query, contexts, **k)
- logger.info(f"Prompt: {prompt}")
-
- if dry_run:
- return prompt
-
- answer, token_info = self.get_answer_from_llm(prompt)
- if isinstance(answer, str):
- logger.info(f"Answer: {answer}")
- return answer, token_info
- else:
- # this is a streamed response and needs to be handled differently.
- return self._stream_response(answer, token_info)
- finally:
- if config:
- # Restore previous config
- self.config: BaseLlmConfig = BaseLlmConfig.deserialize(prev_config)
-
- @staticmethod
- def _get_messages(prompt: str, system_prompt: Optional[str] = None) -> list[LCBaseMessage]:
- """
- Construct a list of langchain messages
-
- :param prompt: User prompt
- :type prompt: str
- :param system_prompt: System prompt, defaults to None
- :type system_prompt: Optional[str], optional
- :return: List of messages
- :rtype: list[BaseMessage]
- """
- from langchain.schema import HumanMessage, SystemMessage
-
- messages = []
- if system_prompt:
- messages.append(SystemMessage(content=system_prompt))
- messages.append(HumanMessage(content=prompt))
- return messages
diff --git a/embedchain/embedchain/llm/clarifai.py b/embedchain/embedchain/llm/clarifai.py
deleted file mode 100644
index 6d87d1b15..000000000
--- a/embedchain/embedchain/llm/clarifai.py
+++ /dev/null
@@ -1,47 +0,0 @@
-import logging
-import os
-from typing import Optional
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class ClarifaiLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config=config)
- if not self.config.api_key and "CLARIFAI_PAT" not in os.environ:
- raise ValueError("Please set the CLARIFAI_PAT environment variable.")
-
- def get_llm_model_answer(self, prompt):
- return self._get_answer(prompt=prompt, config=self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> str:
- try:
- from clarifai.client.model import Model
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for Clarifai are not installed."
- "Please install with `pip install clarifai==10.0.1`"
- ) from None
-
- model_name = config.model
- logging.info(f"Using clarifai LLM model: {model_name}")
- api_key = config.api_key or os.getenv("CLARIFAI_PAT")
- model = Model(url=model_name, pat=api_key)
- params = config.model_kwargs
-
- try:
- (params := {}) if config.model_kwargs is None else config.model_kwargs
- predict_response = model.predict_by_bytes(
- bytes(prompt, "utf-8"),
- input_type="text",
- inference_params=params,
- )
- text = predict_response.outputs[0].data.text.raw
- return text
-
- except Exception as e:
- logging.error(f"Predict failed, exception: {e}")
diff --git a/embedchain/embedchain/llm/cohere.py b/embedchain/embedchain/llm/cohere.py
deleted file mode 100644
index 0a9614b9a..000000000
--- a/embedchain/embedchain/llm/cohere.py
+++ /dev/null
@@ -1,66 +0,0 @@
-import importlib
-import os
-from typing import Any, Optional
-
-from langchain_cohere import ChatCohere
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class CohereLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- try:
- importlib.import_module("cohere")
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for Cohere are not installed."
- "Please install with `pip install langchain_cohere==1.16.0`"
- ) from None
-
- super().__init__(config=config)
- if not self.config.api_key and "COHERE_API_KEY" not in os.environ:
- raise ValueError("Please set the COHERE_API_KEY environment variable or pass it in the config.")
-
- def get_llm_model_answer(self, prompt) -> tuple[str, Optional[dict[str, Any]]]:
- if self.config.system_prompt:
- raise ValueError("CohereLlm does not support `system_prompt`")
-
- if self.config.token_usage:
- response, token_info = self._get_answer(prompt, self.config)
- model_name = "cohere/" + self.config.model
- if model_name not in self.config.model_pricing_map:
- raise ValueError(
- f"Model {model_name} not found in `model_prices_and_context_window.json`. \
- You can disable token usage by setting `token_usage` to False."
- )
- total_cost = (
- self.config.model_pricing_map[model_name]["input_cost_per_token"] * token_info["input_tokens"]
- ) + self.config.model_pricing_map[model_name]["output_cost_per_token"] * token_info["output_tokens"]
- response_token_info = {
- "prompt_tokens": token_info["input_tokens"],
- "completion_tokens": token_info["output_tokens"],
- "total_tokens": token_info["input_tokens"] + token_info["output_tokens"],
- "total_cost": round(total_cost, 10),
- "cost_currency": "USD",
- }
- return response, response_token_info
- return self._get_answer(prompt, self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> str:
- api_key = config.api_key or os.environ["COHERE_API_KEY"]
- kwargs = {
- "model_name": config.model or "command-r",
- "temperature": config.temperature,
- "max_tokens": config.max_tokens,
- "together_api_key": api_key,
- }
-
- chat = ChatCohere(**kwargs)
- chat_response = chat.invoke(prompt)
- if config.token_usage:
- return chat_response.content, chat_response.response_metadata["token_count"]
- return chat_response.content
diff --git a/embedchain/embedchain/llm/google.py b/embedchain/embedchain/llm/google.py
deleted file mode 100644
index c0002fa99..000000000
--- a/embedchain/embedchain/llm/google.py
+++ /dev/null
@@ -1,62 +0,0 @@
-import logging
-import os
-from collections.abc import Generator
-from typing import Any, Optional, Union
-
-try:
- import google.generativeai as genai
-except ImportError:
- raise ImportError("GoogleLlm requires extra dependencies. Install with `pip install google-generativeai`") from None
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class GoogleLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config)
- if not self.config.api_key and "GOOGLE_API_KEY" not in os.environ:
- raise ValueError("Please set the GOOGLE_API_KEY environment variable or pass it in the config.")
-
- api_key = self.config.api_key or os.getenv("GOOGLE_API_KEY")
- genai.configure(api_key=api_key)
-
- def get_llm_model_answer(self, prompt):
- if self.config.system_prompt:
- raise ValueError("GoogleLlm does not support `system_prompt`")
- response = self._get_answer(prompt)
- return response
-
- def _get_answer(self, prompt: str) -> Union[str, Generator[Any, Any, None]]:
- model_name = self.config.model or "gemini-pro"
- logger.info(f"Using Google LLM model: {model_name}")
- model = genai.GenerativeModel(model_name=model_name)
-
- generation_config_params = {
- "candidate_count": 1,
- "max_output_tokens": self.config.max_tokens,
- "temperature": self.config.temperature or 0.5,
- }
-
- if 0.0 <= self.config.top_p <= 1.0:
- generation_config_params["top_p"] = self.config.top_p
- else:
- raise ValueError("`top_p` must be > 0.0 and < 1.0")
-
- generation_config = genai.types.GenerationConfig(**generation_config_params)
-
- response = model.generate_content(
- prompt,
- generation_config=generation_config,
- stream=self.config.stream,
- )
- if self.config.stream:
- # TODO: Implement streaming
- response.resolve()
- return response.text
- else:
- return response.text
diff --git a/embedchain/embedchain/llm/gpt4all.py b/embedchain/embedchain/llm/gpt4all.py
deleted file mode 100644
index 76062b08b..000000000
--- a/embedchain/embedchain/llm/gpt4all.py
+++ /dev/null
@@ -1,67 +0,0 @@
-import os
-from collections.abc import Iterable
-from pathlib import Path
-from typing import Optional, Union
-
-from langchain.callbacks.stdout import StdOutCallbackHandler
-from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class GPT4ALLLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config=config)
- if self.config.model is None:
- self.config.model = "orca-mini-3b-gguf2-q4_0.gguf"
- self.instance = GPT4ALLLlm._get_instance(self.config.model)
- self.instance.streaming = self.config.stream
-
- def get_llm_model_answer(self, prompt):
- return self._get_answer(prompt=prompt, config=self.config)
-
- @staticmethod
- def _get_instance(model):
- try:
- from langchain_community.llms.gpt4all import GPT4All as LangchainGPT4All
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The GPT4All python package is not installed. Please install it with `pip install --upgrade embedchain[opensource]`" # noqa E501
- ) from None
-
- model_path = Path(model).expanduser()
- if os.path.isabs(model_path):
- if os.path.exists(model_path):
- return LangchainGPT4All(model=str(model_path))
- else:
- raise ValueError(f"Model does not exist at {model_path=}")
- else:
- return LangchainGPT4All(model=model, allow_download=True)
-
- def _get_answer(self, prompt: str, config: BaseLlmConfig) -> Union[str, Iterable]:
- if config.model and config.model != self.config.model:
- raise RuntimeError(
- "GPT4ALLLlm does not support switching models at runtime. Please create a new app instance."
- )
-
- messages = []
- if config.system_prompt:
- messages.append(config.system_prompt)
- messages.append(prompt)
- kwargs = {
- "temp": config.temperature,
- "max_tokens": config.max_tokens,
- }
- if config.top_p:
- kwargs["top_p"] = config.top_p
-
- callbacks = [StreamingStdOutCallbackHandler()] if config.stream else [StdOutCallbackHandler()]
-
- response = self.instance.generate(prompts=messages, callbacks=callbacks, **kwargs)
- answer = ""
- for generations in response.generations:
- answer += " ".join(map(lambda generation: generation.text, generations))
- return answer
diff --git a/embedchain/embedchain/llm/groq.py b/embedchain/embedchain/llm/groq.py
deleted file mode 100644
index 3f18d3da9..000000000
--- a/embedchain/embedchain/llm/groq.py
+++ /dev/null
@@ -1,67 +0,0 @@
-import os
-from typing import Any, Optional
-
-from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
-from langchain.schema import HumanMessage, SystemMessage
-
-try:
- from langchain_groq import ChatGroq
-except ImportError:
- raise ImportError("Groq requires extra dependencies. Install with `pip install langchain-groq`") from None
-
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class GroqLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config=config)
- if not self.config.api_key and "GROQ_API_KEY" not in os.environ:
- raise ValueError("Please set the GROQ_API_KEY environment variable or pass it in the config.")
-
- def get_llm_model_answer(self, prompt) -> tuple[str, Optional[dict[str, Any]]]:
- if self.config.token_usage:
- response, token_info = self._get_answer(prompt, self.config)
- model_name = "groq/" + self.config.model
- if model_name not in self.config.model_pricing_map:
- raise ValueError(
- f"Model {model_name} not found in `model_prices_and_context_window.json`. \
- You can disable token usage by setting `token_usage` to False."
- )
- total_cost = (
- self.config.model_pricing_map[model_name]["input_cost_per_token"] * token_info["prompt_tokens"]
- ) + self.config.model_pricing_map[model_name]["output_cost_per_token"] * token_info["completion_tokens"]
- response_token_info = {
- "prompt_tokens": token_info["prompt_tokens"],
- "completion_tokens": token_info["completion_tokens"],
- "total_tokens": token_info["prompt_tokens"] + token_info["completion_tokens"],
- "total_cost": round(total_cost, 10),
- "cost_currency": "USD",
- }
- return response, response_token_info
- return self._get_answer(prompt, self.config)
-
- def _get_answer(self, prompt: str, config: BaseLlmConfig) -> str:
- messages = []
- if config.system_prompt:
- messages.append(SystemMessage(content=config.system_prompt))
- messages.append(HumanMessage(content=prompt))
- api_key = config.api_key or os.environ["GROQ_API_KEY"]
- kwargs = {
- "model_name": config.model or "mixtral-8x7b-32768",
- "temperature": config.temperature,
- "groq_api_key": api_key,
- }
- if config.stream:
- callbacks = config.callbacks if config.callbacks else [StreamingStdOutCallbackHandler()]
- chat = ChatGroq(**kwargs, streaming=config.stream, callbacks=callbacks, api_key=api_key)
- else:
- chat = ChatGroq(**kwargs)
-
- chat_response = chat.invoke(prompt)
- if self.config.token_usage:
- return chat_response.content, chat_response.response_metadata["token_usage"]
- return chat_response.content
diff --git a/embedchain/embedchain/llm/huggingface.py b/embedchain/embedchain/llm/huggingface.py
deleted file mode 100644
index 28767b07b..000000000
--- a/embedchain/embedchain/llm/huggingface.py
+++ /dev/null
@@ -1,99 +0,0 @@
-import importlib
-import logging
-import os
-from typing import Optional
-
-from langchain_community.llms.huggingface_endpoint import HuggingFaceEndpoint
-from langchain_community.llms.huggingface_hub import HuggingFaceHub
-from langchain_community.llms.huggingface_pipeline import HuggingFacePipeline
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class HuggingFaceLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- try:
- importlib.import_module("huggingface_hub")
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for HuggingFaceHub are not installed."
- "Please install with `pip install huggingface-hub==0.23.0`"
- ) from None
-
- super().__init__(config=config)
- if not self.config.api_key and "HUGGINGFACE_ACCESS_TOKEN" not in os.environ:
- raise ValueError("Please set the HUGGINGFACE_ACCESS_TOKEN environment variable or pass it in the config.")
-
- def get_llm_model_answer(self, prompt):
- if self.config.system_prompt:
- raise ValueError("HuggingFaceLlm does not support `system_prompt`")
- return HuggingFaceLlm._get_answer(prompt=prompt, config=self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> str:
- # If the user wants to run the model locally, they can do so by setting the `local` flag to True
- if config.model and config.local:
- return HuggingFaceLlm._from_pipeline(prompt=prompt, config=config)
- elif config.model:
- return HuggingFaceLlm._from_model(prompt=prompt, config=config)
- elif config.endpoint:
- return HuggingFaceLlm._from_endpoint(prompt=prompt, config=config)
- else:
- raise ValueError("Either `model` or `endpoint` must be set in config")
-
- @staticmethod
- def _from_model(prompt: str, config: BaseLlmConfig) -> str:
- model_kwargs = {
- "temperature": config.temperature or 0.1,
- "max_new_tokens": config.max_tokens,
- }
-
- if 0.0 < config.top_p < 1.0:
- model_kwargs["top_p"] = config.top_p
- else:
- raise ValueError("`top_p` must be > 0.0 and < 1.0")
-
- model = config.model
- api_key = config.api_key or os.getenv("HUGGINGFACE_ACCESS_TOKEN")
- logger.info(f"Using HuggingFaceHub with model {model}")
- llm = HuggingFaceHub(
- huggingfacehub_api_token=api_key,
- repo_id=model,
- model_kwargs=model_kwargs,
- )
- return llm.invoke(prompt)
-
- @staticmethod
- def _from_endpoint(prompt: str, config: BaseLlmConfig) -> str:
- api_key = config.api_key or os.getenv("HUGGINGFACE_ACCESS_TOKEN")
- llm = HuggingFaceEndpoint(
- huggingfacehub_api_token=api_key,
- endpoint_url=config.endpoint,
- task="text-generation",
- model_kwargs=config.model_kwargs,
- )
- return llm.invoke(prompt)
-
- @staticmethod
- def _from_pipeline(prompt: str, config: BaseLlmConfig) -> str:
- model_kwargs = {
- "temperature": config.temperature or 0.1,
- "max_new_tokens": config.max_tokens,
- }
-
- if 0.0 < config.top_p < 1.0:
- model_kwargs["top_p"] = config.top_p
- else:
- raise ValueError("`top_p` must be > 0.0 and < 1.0")
-
- llm = HuggingFacePipeline.from_model_id(
- model_id=config.model,
- task="text-generation",
- pipeline_kwargs=model_kwargs,
- )
- return llm.invoke(prompt)
diff --git a/embedchain/embedchain/llm/jina.py b/embedchain/embedchain/llm/jina.py
deleted file mode 100644
index ac3a0e76f..000000000
--- a/embedchain/embedchain/llm/jina.py
+++ /dev/null
@@ -1,45 +0,0 @@
-import os
-from typing import Optional
-
-from langchain.schema import HumanMessage, SystemMessage
-from langchain_community.chat_models import JinaChat
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class JinaLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config=config)
- if not self.config.api_key and "JINACHAT_API_KEY" not in os.environ:
- raise ValueError("Please set the JINACHAT_API_KEY environment variable or pass it in the config.")
-
- def get_llm_model_answer(self, prompt):
- response = JinaLlm._get_answer(prompt, self.config)
- return response
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> str:
- messages = []
- if config.system_prompt:
- messages.append(SystemMessage(content=config.system_prompt))
- messages.append(HumanMessage(content=prompt))
- kwargs = {
- "temperature": config.temperature,
- "max_tokens": config.max_tokens,
- "jinachat_api_key": config.api_key or os.environ["JINACHAT_API_KEY"],
- "model_kwargs": {},
- }
- if config.top_p:
- kwargs["model_kwargs"]["top_p"] = config.top_p
- if config.stream:
- from langchain.callbacks.streaming_stdout import (
- StreamingStdOutCallbackHandler,
- )
-
- chat = JinaChat(**kwargs, streaming=config.stream, callbacks=[StreamingStdOutCallbackHandler()])
- else:
- chat = JinaChat(**kwargs)
- return chat(messages).content
diff --git a/embedchain/embedchain/llm/llama2.py b/embedchain/embedchain/llm/llama2.py
deleted file mode 100644
index 8a82f3f75..000000000
--- a/embedchain/embedchain/llm/llama2.py
+++ /dev/null
@@ -1,53 +0,0 @@
-import importlib
-import os
-from typing import Optional
-
-from langchain_community.llms.replicate import Replicate
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class Llama2Llm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- try:
- importlib.import_module("replicate")
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for Llama2 are not installed."
- 'Please install with `pip install --upgrade "embedchain[llama2]"`'
- ) from None
-
- # Set default config values specific to this llm
- if not config:
- config = BaseLlmConfig()
- # Add variables to this block that have a default value in the parent class
- config.max_tokens = 500
- config.temperature = 0.75
- # Add variables that are `none` by default to this block.
- if not config.model:
- config.model = (
- "a16z-infra/llama13b-v2-chat:df7690f1994d94e96ad9d568eac121aecf50684a0b0963b25a41cc40061269e5"
- )
-
- super().__init__(config=config)
- if not self.config.api_key and "REPLICATE_API_TOKEN" not in os.environ:
- raise ValueError("Please set the REPLICATE_API_TOKEN environment variable or pass it in the config.")
-
- def get_llm_model_answer(self, prompt):
- # TODO: Move the model and other inputs into config
- if self.config.system_prompt:
- raise ValueError("Llama2 does not support `system_prompt`")
- api_key = self.config.api_key or os.getenv("REPLICATE_API_TOKEN")
- llm = Replicate(
- model=self.config.model,
- replicate_api_token=api_key,
- input={
- "temperature": self.config.temperature,
- "max_length": self.config.max_tokens,
- "top_p": self.config.top_p,
- },
- )
- return llm.invoke(prompt)
diff --git a/embedchain/embedchain/llm/mistralai.py b/embedchain/embedchain/llm/mistralai.py
deleted file mode 100644
index 92af3be17..000000000
--- a/embedchain/embedchain/llm/mistralai.py
+++ /dev/null
@@ -1,72 +0,0 @@
-import os
-from typing import Any, Optional
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class MistralAILlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config)
- if not self.config.api_key and "MISTRAL_API_KEY" not in os.environ:
- raise ValueError("Please set the MISTRAL_API_KEY environment variable or pass it in the config.")
-
- def get_llm_model_answer(self, prompt) -> tuple[str, Optional[dict[str, Any]]]:
- if self.config.token_usage:
- response, token_info = self._get_answer(prompt, self.config)
- model_name = "mistralai/" + self.config.model
- if model_name not in self.config.model_pricing_map:
- raise ValueError(
- f"Model {model_name} not found in `model_prices_and_context_window.json`. \
- You can disable token usage by setting `token_usage` to False."
- )
- total_cost = (
- self.config.model_pricing_map[model_name]["input_cost_per_token"] * token_info["prompt_tokens"]
- ) + self.config.model_pricing_map[model_name]["output_cost_per_token"] * token_info["completion_tokens"]
- response_token_info = {
- "prompt_tokens": token_info["prompt_tokens"],
- "completion_tokens": token_info["completion_tokens"],
- "total_tokens": token_info["prompt_tokens"] + token_info["completion_tokens"],
- "total_cost": round(total_cost, 10),
- "cost_currency": "USD",
- }
- return response, response_token_info
- return self._get_answer(prompt, self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig):
- try:
- from langchain_core.messages import HumanMessage, SystemMessage
- from langchain_mistralai.chat_models import ChatMistralAI
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for MistralAI are not installed."
- 'Please install with `pip install --upgrade "embedchain[mistralai]"`'
- ) from None
-
- api_key = config.api_key or os.getenv("MISTRAL_API_KEY")
- client = ChatMistralAI(mistral_api_key=api_key)
- messages = []
- if config.system_prompt:
- messages.append(SystemMessage(content=config.system_prompt))
- messages.append(HumanMessage(content=prompt))
- kwargs = {
- "model": config.model or "mistral-tiny",
- "temperature": config.temperature,
- "max_tokens": config.max_tokens,
- "top_p": config.top_p,
- }
-
- # TODO: Add support for streaming
- if config.stream:
- answer = ""
- for chunk in client.stream(**kwargs, input=messages):
- answer += chunk.content
- return answer
- else:
- chat_response = client.invoke(**kwargs, input=messages)
- if config.token_usage:
- return chat_response.content, chat_response.response_metadata["token_usage"]
- return chat_response.content
diff --git a/embedchain/embedchain/llm/nvidia.py b/embedchain/embedchain/llm/nvidia.py
deleted file mode 100644
index 71c045b6a..000000000
--- a/embedchain/embedchain/llm/nvidia.py
+++ /dev/null
@@ -1,68 +0,0 @@
-import os
-from collections.abc import Iterable
-from typing import Any, Optional, Union
-
-from langchain.callbacks.manager import CallbackManager
-from langchain.callbacks.stdout import StdOutCallbackHandler
-from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
-
-try:
- from langchain_nvidia_ai_endpoints import ChatNVIDIA
-except ImportError:
- raise ImportError(
- "NVIDIA AI endpoints requires extra dependencies. Install with `pip install langchain-nvidia-ai-endpoints`"
- ) from None
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class NvidiaLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config=config)
- if not self.config.api_key and "NVIDIA_API_KEY" not in os.environ:
- raise ValueError("Please set the NVIDIA_API_KEY environment variable or pass it in the config.")
-
- def get_llm_model_answer(self, prompt) -> tuple[str, Optional[dict[str, Any]]]:
- if self.config.token_usage:
- response, token_info = self._get_answer(prompt, self.config)
- model_name = "nvidia/" + self.config.model
- if model_name not in self.config.model_pricing_map:
- raise ValueError(
- f"Model {model_name} not found in `model_prices_and_context_window.json`. \
- You can disable token usage by setting `token_usage` to False."
- )
- total_cost = (
- self.config.model_pricing_map[model_name]["input_cost_per_token"] * token_info["input_tokens"]
- ) + self.config.model_pricing_map[model_name]["output_cost_per_token"] * token_info["output_tokens"]
- response_token_info = {
- "prompt_tokens": token_info["input_tokens"],
- "completion_tokens": token_info["output_tokens"],
- "total_tokens": token_info["input_tokens"] + token_info["output_tokens"],
- "total_cost": round(total_cost, 10),
- "cost_currency": "USD",
- }
- return response, response_token_info
- return self._get_answer(prompt, self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> Union[str, Iterable]:
- callback_manager = [StreamingStdOutCallbackHandler()] if config.stream else [StdOutCallbackHandler()]
- model_kwargs = config.model_kwargs or {}
- labels = model_kwargs.get("labels", None)
- params = {"model": config.model, "nvidia_api_key": config.api_key or os.getenv("NVIDIA_API_KEY")}
- if config.system_prompt:
- params["system_prompt"] = config.system_prompt
- if config.temperature:
- params["temperature"] = config.temperature
- if config.top_p:
- params["top_p"] = config.top_p
- if labels:
- params["labels"] = labels
- llm = ChatNVIDIA(**params, callback_manager=CallbackManager(callback_manager))
- chat_response = llm.invoke(prompt) if labels is None else llm.invoke(prompt, labels=labels)
- if config.token_usage:
- return chat_response.content, chat_response.response_metadata["token_usage"]
- return chat_response.content
diff --git a/embedchain/embedchain/llm/ollama.py b/embedchain/embedchain/llm/ollama.py
deleted file mode 100644
index e34ff38e1..000000000
--- a/embedchain/embedchain/llm/ollama.py
+++ /dev/null
@@ -1,54 +0,0 @@
-import logging
-from collections.abc import Iterable
-from typing import Optional, Union
-
-from langchain.callbacks.manager import CallbackManager
-from langchain.callbacks.stdout import StdOutCallbackHandler
-from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
-from langchain_community.llms.ollama import Ollama
-
-try:
- from ollama import Client
-except ImportError:
- raise ImportError("Ollama requires extra dependencies. Install with `pip install ollama`") from None
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class OllamaLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config=config)
- if self.config.model is None:
- self.config.model = "llama2"
-
- client = Client(host=config.base_url)
- local_models = client.list()["models"]
- if not any(model.get("name") == self.config.model for model in local_models):
- logger.info(f"Pulling {self.config.model} from Ollama!")
- client.pull(self.config.model)
-
- def get_llm_model_answer(self, prompt):
- return self._get_answer(prompt=prompt, config=self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> Union[str, Iterable]:
- if config.stream:
- callbacks = config.callbacks if config.callbacks else [StreamingStdOutCallbackHandler()]
- else:
- callbacks = [StdOutCallbackHandler()]
-
- llm = Ollama(
- model=config.model,
- system=config.system_prompt,
- temperature=config.temperature,
- top_p=config.top_p,
- callback_manager=CallbackManager(callbacks),
- base_url=config.base_url,
- )
-
- return llm.invoke(prompt)
diff --git a/embedchain/embedchain/llm/openai.py b/embedchain/embedchain/llm/openai.py
deleted file mode 100644
index ace146118..000000000
--- a/embedchain/embedchain/llm/openai.py
+++ /dev/null
@@ -1,120 +0,0 @@
-import json
-import os
-import warnings
-from typing import Any, Callable, Dict, Optional, Type, Union
-
-from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
-from langchain.schema import BaseMessage, HumanMessage, SystemMessage
-from langchain_core.tools import BaseTool
-from langchain_openai import ChatOpenAI
-from pydantic import BaseModel
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class OpenAILlm(BaseLlm):
- def __init__(
- self,
- config: Optional[BaseLlmConfig] = None,
- tools: Optional[Union[Dict[str, Any], Type[BaseModel], Callable[..., Any], BaseTool]] = None,
- ):
- self.tools = tools
- super().__init__(config=config)
-
- def get_llm_model_answer(self, prompt) -> tuple[str, Optional[dict[str, Any]]]:
- if self.config.token_usage:
- response, token_info = self._get_answer(prompt, self.config)
- model_name = "openai/" + self.config.model
- if model_name not in self.config.model_pricing_map:
- raise ValueError(
- f"Model {model_name} not found in `model_prices_and_context_window.json`. \
- You can disable token usage by setting `token_usage` to False."
- )
- total_cost = (
- self.config.model_pricing_map[model_name]["input_cost_per_token"] * token_info["prompt_tokens"]
- ) + self.config.model_pricing_map[model_name]["output_cost_per_token"] * token_info["completion_tokens"]
- response_token_info = {
- "prompt_tokens": token_info["prompt_tokens"],
- "completion_tokens": token_info["completion_tokens"],
- "total_tokens": token_info["prompt_tokens"] + token_info["completion_tokens"],
- "total_cost": round(total_cost, 10),
- "cost_currency": "USD",
- }
- return response, response_token_info
-
- return self._get_answer(prompt, self.config)
-
- def _get_answer(self, prompt: str, config: BaseLlmConfig) -> str:
- messages = []
- if config.system_prompt:
- messages.append(SystemMessage(content=config.system_prompt))
- messages.append(HumanMessage(content=prompt))
- kwargs = {
- "model": config.model or "gpt-4o-mini",
- "temperature": config.temperature,
- "max_tokens": config.max_tokens,
- "model_kwargs": config.model_kwargs or {},
- }
- api_key = config.api_key or os.environ["OPENAI_API_KEY"]
- base_url = (
- config.base_url
- or os.getenv("OPENAI_API_BASE")
- or os.getenv("OPENAI_BASE_URL")
- or "https://api.openai.com/v1"
- )
- if os.environ.get("OPENAI_API_BASE"):
- warnings.warn(
- "The environment variable 'OPENAI_API_BASE' is deprecated and will be removed in the 0.1.140. "
- "Please use 'OPENAI_BASE_URL' instead.",
- DeprecationWarning
- )
-
- if config.top_p:
- kwargs["top_p"] = config.top_p
- if config.default_headers:
- kwargs["default_headers"] = config.default_headers
- if config.stream:
- callbacks = config.callbacks if config.callbacks else [StreamingStdOutCallbackHandler()]
- chat = ChatOpenAI(
- **kwargs,
- streaming=config.stream,
- callbacks=callbacks,
- api_key=api_key,
- base_url=base_url,
- http_client=config.http_client,
- http_async_client=config.http_async_client,
- )
- else:
- chat = ChatOpenAI(
- **kwargs,
- api_key=api_key,
- base_url=base_url,
- http_client=config.http_client,
- http_async_client=config.http_async_client,
- )
- if self.tools:
- return self._query_function_call(chat, self.tools, messages)
-
- chat_response = chat.invoke(messages)
- if self.config.token_usage:
- return chat_response.content, chat_response.response_metadata["token_usage"]
- return chat_response.content
-
- def _query_function_call(
- self,
- chat: ChatOpenAI,
- tools: Optional[Union[Dict[str, Any], Type[BaseModel], Callable[..., Any], BaseTool]],
- messages: list[BaseMessage],
- ) -> str:
- from langchain.output_parsers.openai_tools import JsonOutputToolsParser
- from langchain_core.utils.function_calling import convert_to_openai_tool
-
- openai_tools = [convert_to_openai_tool(tools)]
- chat = chat.bind(tools=openai_tools).pipe(JsonOutputToolsParser())
- try:
- return json.dumps(chat.invoke(messages)[0])
- except IndexError:
- return "Input could not be mapped to the function!"
diff --git a/embedchain/embedchain/llm/together.py b/embedchain/embedchain/llm/together.py
deleted file mode 100644
index 84443a712..000000000
--- a/embedchain/embedchain/llm/together.py
+++ /dev/null
@@ -1,71 +0,0 @@
-import importlib
-import os
-from typing import Any, Optional
-
-try:
- from langchain_together import ChatTogether
-except ImportError:
- raise ImportError(
- "Please install the langchain_together package by running `pip install langchain_together==0.1.3`."
- )
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class TogetherLlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- try:
- importlib.import_module("together")
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for Together are not installed."
- 'Please install with `pip install --upgrade "embedchain[together]"`'
- ) from None
-
- super().__init__(config=config)
- if not self.config.api_key and "TOGETHER_API_KEY" not in os.environ:
- raise ValueError("Please set the TOGETHER_API_KEY environment variable or pass it in the config.")
-
- def get_llm_model_answer(self, prompt) -> tuple[str, Optional[dict[str, Any]]]:
- if self.config.system_prompt:
- raise ValueError("TogetherLlm does not support `system_prompt`")
-
- if self.config.token_usage:
- response, token_info = self._get_answer(prompt, self.config)
- model_name = "together/" + self.config.model
- if model_name not in self.config.model_pricing_map:
- raise ValueError(
- f"Model {model_name} not found in `model_prices_and_context_window.json`. \
- You can disable token usage by setting `token_usage` to False."
- )
- total_cost = (
- self.config.model_pricing_map[model_name]["input_cost_per_token"] * token_info["prompt_tokens"]
- ) + self.config.model_pricing_map[model_name]["output_cost_per_token"] * token_info["completion_tokens"]
- response_token_info = {
- "prompt_tokens": token_info["prompt_tokens"],
- "completion_tokens": token_info["completion_tokens"],
- "total_tokens": token_info["prompt_tokens"] + token_info["completion_tokens"],
- "total_cost": round(total_cost, 10),
- "cost_currency": "USD",
- }
- return response, response_token_info
- return self._get_answer(prompt, self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> str:
- api_key = config.api_key or os.environ["TOGETHER_API_KEY"]
- kwargs = {
- "model_name": config.model or "mixtral-8x7b-32768",
- "temperature": config.temperature,
- "max_tokens": config.max_tokens,
- "together_api_key": api_key,
- }
-
- chat = ChatTogether(**kwargs)
- chat_response = chat.invoke(prompt)
- if config.token_usage:
- return chat_response.content, chat_response.response_metadata["token_usage"]
- return chat_response.content
diff --git a/embedchain/embedchain/llm/vertex_ai.py b/embedchain/embedchain/llm/vertex_ai.py
deleted file mode 100644
index 55c31a1ad..000000000
--- a/embedchain/embedchain/llm/vertex_ai.py
+++ /dev/null
@@ -1,68 +0,0 @@
-import importlib
-import logging
-from typing import Any, Optional
-
-from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
-from langchain_google_vertexai import ChatVertexAI
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class VertexAILlm(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- try:
- importlib.import_module("vertexai")
- except ModuleNotFoundError:
- raise ModuleNotFoundError(
- "The required dependencies for VertexAI are not installed."
- 'Please install with `pip install --upgrade "embedchain[vertexai]"`'
- ) from None
- super().__init__(config=config)
-
- def get_llm_model_answer(self, prompt) -> tuple[str, Optional[dict[str, Any]]]:
- if self.config.token_usage:
- response, token_info = self._get_answer(prompt, self.config)
- model_name = "vertexai/" + self.config.model
- if model_name not in self.config.model_pricing_map:
- raise ValueError(
- f"Model {model_name} not found in `model_prices_and_context_window.json`. \
- You can disable token usage by setting `token_usage` to False."
- )
- total_cost = (
- self.config.model_pricing_map[model_name]["input_cost_per_token"] * token_info["prompt_token_count"]
- ) + self.config.model_pricing_map[model_name]["output_cost_per_token"] * token_info[
- "candidates_token_count"
- ]
- response_token_info = {
- "prompt_tokens": token_info["prompt_token_count"],
- "completion_tokens": token_info["candidates_token_count"],
- "total_tokens": token_info["prompt_token_count"] + token_info["candidates_token_count"],
- "total_cost": round(total_cost, 10),
- "cost_currency": "USD",
- }
- return response, response_token_info
- return self._get_answer(prompt, self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> str:
- if config.top_p and config.top_p != 1:
- logger.warning("Config option `top_p` is not supported by this model.")
-
- if config.stream:
- callbacks = config.callbacks if config.callbacks else [StreamingStdOutCallbackHandler()]
- llm = ChatVertexAI(
- temperature=config.temperature, model=config.model, callbacks=callbacks, streaming=config.stream
- )
- else:
- llm = ChatVertexAI(temperature=config.temperature, model=config.model)
-
- messages = VertexAILlm._get_messages(prompt)
- chat_response = llm.invoke(messages)
- if config.token_usage:
- return chat_response.content, chat_response.response_metadata["usage_metadata"]
- return chat_response.content
diff --git a/embedchain/embedchain/llm/vllm.py b/embedchain/embedchain/llm/vllm.py
deleted file mode 100644
index 88a8e2ad2..000000000
--- a/embedchain/embedchain/llm/vllm.py
+++ /dev/null
@@ -1,40 +0,0 @@
-from typing import Iterable, Optional, Union
-
-from langchain.callbacks.manager import CallbackManager
-from langchain.callbacks.stdout import StdOutCallbackHandler
-from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
-from langchain_community.llms import VLLM as BaseVLLM
-
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.llm.base import BaseLlm
-
-
-@register_deserializable
-class VLLM(BaseLlm):
- def __init__(self, config: Optional[BaseLlmConfig] = None):
- super().__init__(config=config)
- if self.config.model is None:
- self.config.model = "mosaicml/mpt-7b"
-
- def get_llm_model_answer(self, prompt):
- return self._get_answer(prompt=prompt, config=self.config)
-
- @staticmethod
- def _get_answer(prompt: str, config: BaseLlmConfig) -> Union[str, Iterable]:
- callback_manager = [StreamingStdOutCallbackHandler()] if config.stream else [StdOutCallbackHandler()]
-
- # Prepare the arguments for BaseVLLM
- llm_args = {
- "model": config.model,
- "temperature": config.temperature,
- "top_p": config.top_p,
- "callback_manager": CallbackManager(callback_manager),
- }
-
- # Add model_kwargs if they are not None
- if config.model_kwargs is not None:
- llm_args.update(config.model_kwargs)
-
- llm = BaseVLLM(**llm_args)
- return llm.invoke(prompt)
diff --git a/embedchain/embedchain/loaders/__init__.py b/embedchain/embedchain/loaders/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/loaders/audio.py b/embedchain/embedchain/loaders/audio.py
deleted file mode 100644
index 6b2b69cf2..000000000
--- a/embedchain/embedchain/loaders/audio.py
+++ /dev/null
@@ -1,53 +0,0 @@
-import hashlib
-import os
-
-import validators
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-try:
- from deepgram import DeepgramClient, PrerecordedOptions
-except ImportError:
- raise ImportError(
- "Audio file requires extra dependencies. Install with `pip install deepgram-sdk==3.2.7`"
- ) from None
-
-
-@register_deserializable
-class AudioLoader(BaseLoader):
- def __init__(self):
- if not os.environ.get("DEEPGRAM_API_KEY"):
- raise ValueError("DEEPGRAM_API_KEY is not set")
-
- DG_KEY = os.environ.get("DEEPGRAM_API_KEY")
- self.client = DeepgramClient(DG_KEY)
-
- def load_data(self, url: str):
- """Load data from a audio file or URL."""
-
- options = PrerecordedOptions(
- model="nova-2",
- smart_format=True,
- )
- if validators.url(url):
- source = {"url": url}
- response = self.client.listen.prerecorded.v("1").transcribe_url(source, options)
- else:
- with open(url, "rb") as audio:
- source = {"buffer": audio}
- response = self.client.listen.prerecorded.v("1").transcribe_file(source, options)
- content = response["results"]["channels"][0]["alternatives"][0]["transcript"]
-
- doc_id = hashlib.sha256((content + url).encode()).hexdigest()
- metadata = {"url": url}
-
- return {
- "doc_id": doc_id,
- "data": [
- {
- "content": content,
- "meta_data": metadata,
- }
- ],
- }
diff --git a/embedchain/embedchain/loaders/base_loader.py b/embedchain/embedchain/loaders/base_loader.py
deleted file mode 100644
index 9dccfd539..000000000
--- a/embedchain/embedchain/loaders/base_loader.py
+++ /dev/null
@@ -1,14 +0,0 @@
-from typing import Any, Optional
-
-from embedchain.helpers.json_serializable import JSONSerializable
-
-
-class BaseLoader(JSONSerializable):
- def __init__(self):
- pass
-
- def load_data(self, url, **kwargs: Optional[dict[str, Any]]):
- """
- Implemented by child classes
- """
- pass
diff --git a/embedchain/embedchain/loaders/beehiiv.py b/embedchain/embedchain/loaders/beehiiv.py
deleted file mode 100644
index 12d0fe4a9..000000000
--- a/embedchain/embedchain/loaders/beehiiv.py
+++ /dev/null
@@ -1,107 +0,0 @@
-import hashlib
-import logging
-import time
-from xml.etree import ElementTree
-
-import requests
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import is_readable
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class BeehiivLoader(BaseLoader):
- """
- This loader is used to load data from Beehiiv URLs.
- """
-
- def load_data(self, url: str):
- try:
- from bs4 import BeautifulSoup
- from bs4.builder import ParserRejectedMarkup
- except ImportError:
- raise ImportError(
- "Beehiiv requires extra dependencies. Install with `pip install beautifulsoup4==4.12.3`"
- ) from None
-
- if not url.endswith("sitemap.xml"):
- url = url + "/sitemap.xml"
-
- output = []
- # we need to set this as a header to avoid 403
- headers = {
- "User-Agent": (
- "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_5) "
- "AppleWebKit/537.36 (KHTML, like Gecko) Chrome/50.0.2661.102 "
- "Safari/537.36"
- ),
- }
- response = requests.get(url, headers=headers)
- try:
- response.raise_for_status()
- except requests.exceptions.HTTPError as e:
- raise ValueError(
- f"""
- Failed to load {url}: {e}. Please use the root substack URL. For example, https://example.substack.com
- """
- )
-
- try:
- ElementTree.fromstring(response.content)
- except ElementTree.ParseError:
- raise ValueError(
- f"""
- Failed to parse {url}. Please use the root substack URL. For example, https://example.substack.com
- """
- )
- soup = BeautifulSoup(response.text, "xml")
- links = [link.text for link in soup.find_all("loc") if link.parent.name == "url" and "/p/" in link.text]
- if len(links) == 0:
- links = [link.text for link in soup.find_all("loc") if "/p/" in link.text]
-
- doc_id = hashlib.sha256((" ".join(links) + url).encode()).hexdigest()
-
- def serialize_response(soup: BeautifulSoup):
- data = {}
-
- h1_el = soup.find("h1")
- if h1_el is not None:
- data["title"] = h1_el.text
-
- description_el = soup.find("meta", {"name": "description"})
- if description_el is not None:
- data["description"] = description_el["content"]
-
- content_el = soup.find("div", {"id": "content-blocks"})
- if content_el is not None:
- data["content"] = content_el.text
-
- return data
-
- def load_link(link: str):
- try:
- beehiiv_data = requests.get(link, headers=headers)
- beehiiv_data.raise_for_status()
-
- soup = BeautifulSoup(beehiiv_data.text, "html.parser")
- data = serialize_response(soup)
- data = str(data)
- if is_readable(data):
- return data
- else:
- logger.warning(f"Page is not readable (too many invalid characters): {link}")
- except ParserRejectedMarkup as e:
- logger.error(f"Failed to parse {link}: {e}")
- return None
-
- for link in links:
- data = load_link(link)
- if data:
- output.append({"content": data, "meta_data": {"url": link}})
- # TODO: allow users to configure this
- time.sleep(1.0) # added to avoid rate limiting
-
- return {"doc_id": doc_id, "data": output}
diff --git a/embedchain/embedchain/loaders/csv.py b/embedchain/embedchain/loaders/csv.py
deleted file mode 100644
index 2714d5759..000000000
--- a/embedchain/embedchain/loaders/csv.py
+++ /dev/null
@@ -1,49 +0,0 @@
-import csv
-import hashlib
-from io import StringIO
-from urllib.parse import urlparse
-
-import requests
-
-from embedchain.loaders.base_loader import BaseLoader
-
-
-class CsvLoader(BaseLoader):
- @staticmethod
- def _detect_delimiter(first_line):
- delimiters = [",", "\t", ";", "|"]
- counts = {delimiter: first_line.count(delimiter) for delimiter in delimiters}
- return max(counts, key=counts.get)
-
- @staticmethod
- def _get_file_content(content):
- url = urlparse(content)
- if all([url.scheme, url.netloc]) and url.scheme not in ["file", "http", "https"]:
- raise ValueError("Not a valid URL.")
-
- if url.scheme in ["http", "https"]:
- response = requests.get(content)
- response.raise_for_status()
- return StringIO(response.text)
- elif url.scheme == "file":
- path = url.path
- return open(path, newline="", encoding="utf-8") # Open the file using the path from the URI
- else:
- return open(content, newline="", encoding="utf-8") # Treat content as a regular file path
-
- @staticmethod
- def load_data(content):
- """Load a csv file with headers. Each line is a document"""
- result = []
- lines = []
- with CsvLoader._get_file_content(content) as file:
- first_line = file.readline()
- delimiter = CsvLoader._detect_delimiter(first_line)
- file.seek(0) # Reset the file pointer to the start
- reader = csv.DictReader(file, delimiter=delimiter)
- for i, row in enumerate(reader):
- line = ", ".join([f"{field}: {value}" for field, value in row.items()])
- lines.append(line)
- result.append({"content": line, "meta_data": {"url": content, "row": i + 1}})
- doc_id = hashlib.sha256((content + " ".join(lines)).encode()).hexdigest()
- return {"doc_id": doc_id, "data": result}
diff --git a/embedchain/embedchain/loaders/directory_loader.py b/embedchain/embedchain/loaders/directory_loader.py
deleted file mode 100644
index 5903813b5..000000000
--- a/embedchain/embedchain/loaders/directory_loader.py
+++ /dev/null
@@ -1,63 +0,0 @@
-import hashlib
-import logging
-from pathlib import Path
-from typing import Any, Optional
-
-from embedchain.config import AddConfig
-from embedchain.data_formatter.data_formatter import DataFormatter
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.loaders.text_file import TextFileLoader
-from embedchain.utils.misc import detect_datatype
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class DirectoryLoader(BaseLoader):
- """Load data from a directory."""
-
- def __init__(self, config: Optional[dict[str, Any]] = None):
- super().__init__()
- config = config or {}
- self.recursive = config.get("recursive", True)
- self.extensions = config.get("extensions", None)
- self.errors = []
-
- def load_data(self, path: str):
- directory_path = Path(path)
- if not directory_path.is_dir():
- raise ValueError(f"Invalid path: {path}")
-
- logger.info(f"Loading data from directory: {path}")
- data_list = self._process_directory(directory_path)
- doc_id = hashlib.sha256((str(data_list) + str(directory_path)).encode()).hexdigest()
-
- for error in self.errors:
- logger.warning(error)
-
- return {"doc_id": doc_id, "data": data_list}
-
- def _process_directory(self, directory_path: Path):
- data_list = []
- for file_path in directory_path.rglob("*") if self.recursive else directory_path.glob("*"):
- # don't include dotfiles
- if file_path.name.startswith("."):
- continue
- if file_path.is_file() and (not self.extensions or any(file_path.suffix == ext for ext in self.extensions)):
- loader = self._predict_loader(file_path)
- data_list.extend(loader.load_data(str(file_path))["data"])
- elif file_path.is_dir():
- logger.info(f"Loading data from directory: {file_path}")
- return data_list
-
- def _predict_loader(self, file_path: Path) -> BaseLoader:
- try:
- data_type = detect_datatype(str(file_path))
- config = AddConfig()
- return DataFormatter(data_type=data_type, config=config)._get_loader(
- data_type=data_type, config=config.loader, loader=None
- )
- except Exception as e:
- self.errors.append(f"Error processing {file_path}: {e}")
- return TextFileLoader()
diff --git a/embedchain/embedchain/loaders/discord.py b/embedchain/embedchain/loaders/discord.py
deleted file mode 100644
index 807a3d00c..000000000
--- a/embedchain/embedchain/loaders/discord.py
+++ /dev/null
@@ -1,152 +0,0 @@
-import hashlib
-import logging
-import os
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class DiscordLoader(BaseLoader):
- """
- Load data from a Discord Channel ID.
- """
-
- def __init__(self):
- if not os.environ.get("DISCORD_TOKEN"):
- raise ValueError("DISCORD_TOKEN is not set")
-
- self.token = os.environ.get("DISCORD_TOKEN")
-
- @staticmethod
- def _format_message(message):
- return {
- "message_id": message.id,
- "content": message.content,
- "author": {
- "id": message.author.id,
- "name": message.author.name,
- "discriminator": message.author.discriminator,
- },
- "created_at": message.created_at.isoformat(),
- "attachments": [
- {
- "id": attachment.id,
- "filename": attachment.filename,
- "size": attachment.size,
- "url": attachment.url,
- "proxy_url": attachment.proxy_url,
- "height": attachment.height,
- "width": attachment.width,
- }
- for attachment in message.attachments
- ],
- "embeds": [
- {
- "title": embed.title,
- "type": embed.type,
- "description": embed.description,
- "url": embed.url,
- "timestamp": embed.timestamp.isoformat(),
- "color": embed.color,
- "footer": {
- "text": embed.footer.text,
- "icon_url": embed.footer.icon_url,
- "proxy_icon_url": embed.footer.proxy_icon_url,
- },
- "image": {
- "url": embed.image.url,
- "proxy_url": embed.image.proxy_url,
- "height": embed.image.height,
- "width": embed.image.width,
- },
- "thumbnail": {
- "url": embed.thumbnail.url,
- "proxy_url": embed.thumbnail.proxy_url,
- "height": embed.thumbnail.height,
- "width": embed.thumbnail.width,
- },
- "video": {
- "url": embed.video.url,
- "height": embed.video.height,
- "width": embed.video.width,
- },
- "provider": {
- "name": embed.provider.name,
- "url": embed.provider.url,
- },
- "author": {
- "name": embed.author.name,
- "url": embed.author.url,
- "icon_url": embed.author.icon_url,
- "proxy_icon_url": embed.author.proxy_icon_url,
- },
- "fields": [
- {
- "name": field.name,
- "value": field.value,
- "inline": field.inline,
- }
- for field in embed.fields
- ],
- }
- for embed in message.embeds
- ],
- }
-
- def load_data(self, channel_id: str):
- """Load data from a Discord Channel ID."""
- import discord
-
- messages = []
-
- class DiscordClient(discord.Client):
- async def on_ready(self) -> None:
- logger.info("Logged on as {0}!".format(self.user))
- try:
- channel = self.get_channel(int(channel_id))
- if not isinstance(channel, discord.TextChannel):
- raise ValueError(
- f"Channel {channel_id} is not a text channel. " "Only text channels are supported for now."
- )
- threads = {}
-
- for thread in channel.threads:
- threads[thread.id] = thread
-
- async for message in channel.history(limit=None):
- messages.append(DiscordLoader._format_message(message))
- if message.id in threads:
- async for thread_message in threads[message.id].history(limit=None):
- messages.append(DiscordLoader._format_message(thread_message))
-
- except Exception as e:
- logger.error(e)
- await self.close()
- finally:
- await self.close()
-
- intents = discord.Intents.default()
- intents.message_content = True
- client = DiscordClient(intents=intents)
- client.run(self.token)
-
- metadata = {
- "url": channel_id,
- }
-
- messages = str(messages)
-
- doc_id = hashlib.sha256((messages + channel_id).encode()).hexdigest()
-
- return {
- "doc_id": doc_id,
- "data": [
- {
- "content": messages,
- "meta_data": metadata,
- }
- ],
- }
diff --git a/embedchain/embedchain/loaders/discourse.py b/embedchain/embedchain/loaders/discourse.py
deleted file mode 100644
index 65c1dd756..000000000
--- a/embedchain/embedchain/loaders/discourse.py
+++ /dev/null
@@ -1,79 +0,0 @@
-import hashlib
-import logging
-import time
-from typing import Any, Optional
-
-import requests
-
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-logger = logging.getLogger(__name__)
-
-
-class DiscourseLoader(BaseLoader):
- def __init__(self, config: Optional[dict[str, Any]] = None):
- super().__init__()
- if not config:
- raise ValueError(
- "DiscourseLoader requires a config. Check the documentation for the correct format - `https://docs.embedchain.ai/components/data-sources/discourse`" # noqa: E501
- )
-
- self.domain = config.get("domain")
- if not self.domain:
- raise ValueError(
- "DiscourseLoader requires a domain. Check the documentation for the correct format - `https://docs.embedchain.ai/components/data-sources/discourse`" # noqa: E501
- )
-
- def _check_query(self, query):
- if not query or not isinstance(query, str):
- raise ValueError(
- "DiscourseLoader requires a query. Check the documentation for the correct format - `https://docs.embedchain.ai/components/data-sources/discourse`" # noqa: E501
- )
-
- def _load_post(self, post_id):
- post_url = f"{self.domain}posts/{post_id}.json"
- response = requests.get(post_url)
- try:
- response.raise_for_status()
- except Exception as e:
- logger.error(f"Failed to load post {post_id}: {e}")
- return
- response_data = response.json()
- post_contents = clean_string(response_data.get("raw"))
- metadata = {
- "url": post_url,
- "created_at": response_data.get("created_at", ""),
- "username": response_data.get("username", ""),
- "topic_slug": response_data.get("topic_slug", ""),
- "score": response_data.get("score", ""),
- }
- data = {
- "content": post_contents,
- "meta_data": metadata,
- }
- return data
-
- def load_data(self, query):
- self._check_query(query)
- data = []
- data_contents = []
- logger.info(f"Searching data on discourse url: {self.domain}, for query: {query}")
- search_url = f"{self.domain}search.json?q={query}"
- response = requests.get(search_url)
- try:
- response.raise_for_status()
- except Exception as e:
- raise ValueError(f"Failed to search query {query}: {e}")
- response_data = response.json()
- post_ids = response_data.get("grouped_search_result").get("post_ids")
- for id in post_ids:
- post_data = self._load_post(id)
- if post_data:
- data.append(post_data)
- data_contents.append(post_data.get("content"))
- # Sleep for 0.4 sec, to avoid rate limiting. Check `https://meta.discourse.org/t/api-rate-limits/208405/6`
- time.sleep(0.4)
- doc_id = hashlib.sha256((query + ", ".join(data_contents)).encode()).hexdigest()
- response_data = {"doc_id": doc_id, "data": data}
- return response_data
diff --git a/embedchain/embedchain/loaders/docs_site_loader.py b/embedchain/embedchain/loaders/docs_site_loader.py
deleted file mode 100644
index b9831a9cd..000000000
--- a/embedchain/embedchain/loaders/docs_site_loader.py
+++ /dev/null
@@ -1,119 +0,0 @@
-import hashlib
-import logging
-from urllib.parse import urljoin, urlparse
-
-import requests
-
-try:
- from bs4 import BeautifulSoup
-except ImportError:
- raise ImportError(
- "DocsSite requires extra dependencies. Install with `pip install beautifulsoup4==4.12.3`"
- ) from None
-
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class DocsSiteLoader(BaseLoader):
- def __init__(self):
- self.visited_links = set()
-
- def _get_child_links_recursive(self, url):
- if url in self.visited_links:
- return
-
- parsed_url = urlparse(url)
- base_url = f"{parsed_url.scheme}://{parsed_url.netloc}"
- current_path = parsed_url.path
-
- response = requests.get(url)
- if response.status_code != 200:
- logger.info(f"Failed to fetch the website: {response.status_code}")
- return
-
- soup = BeautifulSoup(response.text, "html.parser")
- all_links = (link.get("href") for link in soup.find_all("a", href=True))
-
- child_links = (link for link in all_links if link.startswith(current_path) and link != current_path)
-
- absolute_paths = set(urljoin(base_url, link) for link in child_links)
-
- self.visited_links.update(absolute_paths)
-
- [self._get_child_links_recursive(link) for link in absolute_paths if link not in self.visited_links]
-
- def _get_all_urls(self, url):
- self.visited_links = set()
- self._get_child_links_recursive(url)
- urls = [link for link in self.visited_links if urlparse(link).netloc == urlparse(url).netloc]
- return urls
-
- @staticmethod
- def _load_data_from_url(url: str) -> list:
- response = requests.get(url)
- if response.status_code != 200:
- logger.info(f"Failed to fetch the website: {response.status_code}")
- return []
-
- soup = BeautifulSoup(response.content, "html.parser")
- selectors = [
- "article.bd-article",
- 'article[role="main"]',
- "div.md-content",
- 'div[role="main"]',
- "div.container",
- "div.section",
- "article",
- "main",
- ]
-
- output = []
- for selector in selectors:
- element = soup.select_one(selector)
- if element:
- content = element.prettify()
- break
- else:
- content = soup.get_text()
-
- soup = BeautifulSoup(content, "html.parser")
- ignored_tags = [
- "nav",
- "aside",
- "form",
- "header",
- "noscript",
- "svg",
- "canvas",
- "footer",
- "script",
- "style",
- ]
- for tag in soup(ignored_tags):
- tag.decompose()
-
- content = " ".join(soup.stripped_strings)
- output.append(
- {
- "content": content,
- "meta_data": {"url": url},
- }
- )
-
- return output
-
- def load_data(self, url):
- all_urls = self._get_all_urls(url)
- output = []
- for u in all_urls:
- output.extend(self._load_data_from_url(u))
- doc_id = hashlib.sha256((" ".join(all_urls) + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": output,
- }
diff --git a/embedchain/embedchain/loaders/docx_file.py b/embedchain/embedchain/loaders/docx_file.py
deleted file mode 100644
index 219bb9914..000000000
--- a/embedchain/embedchain/loaders/docx_file.py
+++ /dev/null
@@ -1,26 +0,0 @@
-import hashlib
-
-try:
- from langchain_community.document_loaders import Docx2txtLoader
-except ImportError:
- raise ImportError("Docx file requires extra dependencies. Install with `pip install docx2txt==0.8`") from None
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-
-@register_deserializable
-class DocxFileLoader(BaseLoader):
- def load_data(self, url):
- """Load data from a .docx file."""
- loader = Docx2txtLoader(url)
- output = []
- data = loader.load()
- content = data[0].page_content
- metadata = data[0].metadata
- metadata["url"] = "local"
- output.append({"content": content, "meta_data": metadata})
- doc_id = hashlib.sha256((content + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": output,
- }
diff --git a/embedchain/embedchain/loaders/dropbox.py b/embedchain/embedchain/loaders/dropbox.py
deleted file mode 100644
index 1fbaf2897..000000000
--- a/embedchain/embedchain/loaders/dropbox.py
+++ /dev/null
@@ -1,79 +0,0 @@
-import hashlib
-import os
-
-from dropbox.files import FileMetadata
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.loaders.directory_loader import DirectoryLoader
-
-
-@register_deserializable
-class DropboxLoader(BaseLoader):
- def __init__(self):
- access_token = os.environ.get("DROPBOX_ACCESS_TOKEN")
- if not access_token:
- raise ValueError("Please set the `DROPBOX_ACCESS_TOKEN` environment variable.")
- try:
- from dropbox import Dropbox, exceptions
- except ImportError:
- raise ImportError("Dropbox requires extra dependencies. Install with `pip install dropbox==11.36.2`")
-
- try:
- dbx = Dropbox(access_token)
- dbx.users_get_current_account()
- self.dbx = dbx
- except exceptions.AuthError as ex:
- raise ValueError("Invalid Dropbox access token. Please verify your token and try again.") from ex
-
- def _download_folder(self, path: str, local_root: str) -> list[FileMetadata]:
- """Download a folder from Dropbox and save it preserving the directory structure."""
- entries = self.dbx.files_list_folder(path).entries
- for entry in entries:
- local_path = os.path.join(local_root, entry.name)
- if isinstance(entry, FileMetadata):
- self.dbx.files_download_to_file(local_path, f"{path}/{entry.name}")
- else:
- os.makedirs(local_path, exist_ok=True)
- self._download_folder(f"{path}/{entry.name}", local_path)
- return entries
-
- def _generate_dir_id_from_all_paths(self, path: str) -> str:
- """Generate a unique ID for a directory based on all of its paths."""
- entries = self.dbx.files_list_folder(path).entries
- paths = [f"{path}/{entry.name}" for entry in entries]
- return hashlib.sha256("".join(paths).encode()).hexdigest()
-
- def load_data(self, path: str):
- """Load data from a Dropbox URL, preserving the folder structure."""
- root_dir = f"dropbox_{self._generate_dir_id_from_all_paths(path)}"
- os.makedirs(root_dir, exist_ok=True)
-
- for entry in self.dbx.files_list_folder(path).entries:
- local_path = os.path.join(root_dir, entry.name)
- if isinstance(entry, FileMetadata):
- self.dbx.files_download_to_file(local_path, f"{path}/{entry.name}")
- else:
- os.makedirs(local_path, exist_ok=True)
- self._download_folder(f"{path}/{entry.name}", local_path)
-
- dir_loader = DirectoryLoader()
- data = dir_loader.load_data(root_dir)["data"]
-
- # Clean up
- self._clean_directory(root_dir)
-
- return {
- "doc_id": hashlib.sha256(path.encode()).hexdigest(),
- "data": data,
- }
-
- def _clean_directory(self, dir_path):
- """Recursively delete a directory and its contents."""
- for item in os.listdir(dir_path):
- item_path = os.path.join(dir_path, item)
- if os.path.isdir(item_path):
- self._clean_directory(item_path)
- else:
- os.remove(item_path)
- os.rmdir(dir_path)
diff --git a/embedchain/embedchain/loaders/excel_file.py b/embedchain/embedchain/loaders/excel_file.py
deleted file mode 100644
index 585415770..000000000
--- a/embedchain/embedchain/loaders/excel_file.py
+++ /dev/null
@@ -1,41 +0,0 @@
-import hashlib
-import importlib.util
-
-try:
- import unstructured # noqa: F401
- from langchain_community.document_loaders import UnstructuredExcelLoader
-except ImportError:
- raise ImportError(
- 'Excel file requires extra dependencies. Install with `pip install "unstructured[local-inference, all-docs]"`'
- ) from None
-
-if importlib.util.find_spec("openpyxl") is None and importlib.util.find_spec("xlrd") is None:
- raise ImportError("Excel file requires extra dependencies. Install with `pip install openpyxl xlrd`") from None
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-
-@register_deserializable
-class ExcelFileLoader(BaseLoader):
- def load_data(self, excel_url):
- """Load data from a Excel file."""
- loader = UnstructuredExcelLoader(excel_url)
- pages = loader.load_and_split()
-
- data = []
- for page in pages:
- content = page.page_content
- content = clean_string(content)
-
- metadata = page.metadata
- metadata["url"] = excel_url
-
- data.append({"content": content, "meta_data": metadata})
-
- doc_id = hashlib.sha256((content + excel_url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": data,
- }
diff --git a/embedchain/embedchain/loaders/github.py b/embedchain/embedchain/loaders/github.py
deleted file mode 100644
index dac7241e0..000000000
--- a/embedchain/embedchain/loaders/github.py
+++ /dev/null
@@ -1,312 +0,0 @@
-import concurrent.futures
-import hashlib
-import logging
-import re
-import shlex
-from typing import Any, Optional
-
-from tqdm import tqdm
-
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-GITHUB_URL = "https://github.com"
-GITHUB_API_URL = "https://api.github.com"
-
-VALID_SEARCH_TYPES = set(["code", "repo", "pr", "issue", "discussion", "branch", "file"])
-
-
-class GithubLoader(BaseLoader):
- """Load data from GitHub search query."""
-
- def __init__(self, config: Optional[dict[str, Any]] = None):
- super().__init__()
- if not config:
- raise ValueError(
- "GithubLoader requires a personal access token to use github api. Check - `https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens#creating-a-personal-access-token-classic`" # noqa: E501
- )
-
- try:
- from github import Github
- except ImportError as e:
- raise ValueError(
- "GithubLoader requires extra dependencies. \
- Install with `pip install gitpython==3.1.38 PyGithub==1.59.1`"
- ) from e
-
- self.config = config
- token = config.get("token")
- if not token:
- raise ValueError(
- "GithubLoader requires a personal access token to use github api. Check - `https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/managing-your-personal-access-tokens#creating-a-personal-access-token-classic`" # noqa: E501
- )
-
- try:
- self.client = Github(token)
- except Exception as e:
- logging.error(f"GithubLoader failed to initialize client: {e}")
- self.client = None
-
- def _github_search_code(self, query: str):
- """Search GitHub code."""
- data = []
- results = self.client.search_code(query)
- for result in tqdm(results, total=results.totalCount, desc="Loading code files from github"):
- url = result.html_url
- logging.info(f"Added data from url: {url}")
- content = result.decoded_content.decode("utf-8")
- metadata = {
- "url": url,
- }
- data.append(
- {
- "content": clean_string(content),
- "meta_data": metadata,
- }
- )
- return data
-
- def _get_github_repo_data(self, repo_name: str, branch_name: str = None, file_path: str = None) -> list[dict]:
- """Get file contents from Repo"""
- data = []
-
- repo = self.client.get_repo(repo_name)
- repo_contents = repo.get_contents("")
-
- if branch_name:
- repo_contents = repo.get_contents("", ref=branch_name)
- if file_path:
- repo_contents = [repo.get_contents(file_path)]
-
- with tqdm(desc="Loading files:", unit="item") as progress_bar:
- while repo_contents:
- file_content = repo_contents.pop(0)
- if file_content.type == "dir":
- try:
- repo_contents.extend(repo.get_contents(file_content.path))
- except Exception:
- logging.warning(f"Failed to read directory: {file_content.path}")
- progress_bar.update(1)
- continue
- else:
- try:
- file_text = file_content.decoded_content.decode()
- except Exception:
- logging.warning(f"Failed to read file: {file_content.path}")
- progress_bar.update(1)
- continue
-
- file_path = file_content.path
- data.append(
- {
- "content": clean_string(file_text),
- "meta_data": {
- "path": file_path,
- },
- }
- )
-
- progress_bar.update(1)
-
- return data
-
- def _github_search_repo(self, query: str) -> list[dict]:
- """Search GitHub repo."""
-
- logging.info(f"Searching github repos with query: {query}")
- updated_query = query.split(":")[-1]
- data = self._get_github_repo_data(updated_query)
- return data
-
- def _github_search_issues_and_pr(self, query: str, type: str) -> list[dict]:
- """Search GitHub issues and PRs."""
- data = []
-
- query = f"{query} is:{type}"
- logging.info(f"Searching github for query: {query}")
-
- results = self.client.search_issues(query)
-
- logging.info(f"Total results: {results.totalCount}")
- for result in tqdm(results, total=results.totalCount, desc=f"Loading {type} from github"):
- url = result.html_url
- title = result.title
- body = result.body
- if not body:
- logging.warning(f"Skipping issue because empty content for: {url}")
- continue
- labels = " ".join([label.name for label in result.labels])
- issue_comments = result.get_comments()
- comments = []
- comments_created_at = []
- for comment in issue_comments:
- comments_created_at.append(str(comment.created_at))
- comments.append(f"{comment.user.name}:{comment.body}")
- content = "\n".join([title, labels, body, *comments])
- metadata = {
- "url": url,
- "created_at": str(result.created_at),
- "comments_created_at": " ".join(comments_created_at),
- }
- data.append(
- {
- "content": clean_string(content),
- "meta_data": metadata,
- }
- )
- return data
-
- # need to test more for discussion
- def _github_search_discussions(self, query: str):
- """Search GitHub discussions."""
- data = []
-
- query = f"{query} is:discussion"
- logging.info(f"Searching github repo for query: {query}")
- repos_results = self.client.search_repositories(query)
- logging.info(f"Total repos found: {repos_results.totalCount}")
- for repo_result in tqdm(repos_results, total=repos_results.totalCount, desc="Loading discussions from github"):
- teams = repo_result.get_teams()
- for team in teams:
- team_discussions = team.get_discussions()
- for discussion in team_discussions:
- url = discussion.html_url
- title = discussion.title
- body = discussion.body
- if not body:
- logging.warning(f"Skipping discussion because empty content for: {url}")
- continue
- comments = []
- comments_created_at = []
- print("Discussion comments: ", discussion.comments_url)
- content = "\n".join([title, body, *comments])
- metadata = {
- "url": url,
- "created_at": str(discussion.created_at),
- "comments_created_at": " ".join(comments_created_at),
- }
- data.append(
- {
- "content": clean_string(content),
- "meta_data": metadata,
- }
- )
- return data
-
- def _get_github_repo_branch(self, query: str, type: str) -> list[dict]:
- """Get file contents for specific branch"""
-
- logging.info(f"Searching github repo for query: {query} is:{type}")
- pattern = r"repo:(\S+) name:(\S+)"
- match = re.search(pattern, query)
-
- if match:
- repo_name = match.group(1)
- branch_name = match.group(2)
- else:
- raise ValueError(
- f"Repository name and Branch name not found, instead found this \
- Repo: {repo_name}, Branch: {branch_name}"
- )
-
- data = self._get_github_repo_data(repo_name=repo_name, branch_name=branch_name)
- return data
-
- def _get_github_repo_file(self, query: str, type: str) -> list[dict]:
- """Get specific file content"""
-
- logging.info(f"Searching github repo for query: {query} is:{type}")
- pattern = r"repo:(\S+) path:(\S+)"
- match = re.search(pattern, query)
-
- if match:
- repo_name = match.group(1)
- file_path = match.group(2)
- else:
- raise ValueError(
- f"Repository name and File name not found, instead found this Repo: {repo_name}, File: {file_path}"
- )
-
- data = self._get_github_repo_data(repo_name=repo_name, file_path=file_path)
- return data
-
- def _search_github_data(self, search_type: str, query: str):
- """Search github data."""
- if search_type == "code":
- data = self._github_search_code(query)
- elif search_type == "repo":
- data = self._github_search_repo(query)
- elif search_type == "issue":
- data = self._github_search_issues_and_pr(query, search_type)
- elif search_type == "pr":
- data = self._github_search_issues_and_pr(query, search_type)
- elif search_type == "branch":
- data = self._get_github_repo_branch(query, search_type)
- elif search_type == "file":
- data = self._get_github_repo_file(query, search_type)
- elif search_type == "discussion":
- raise ValueError("GithubLoader does not support searching discussions yet.")
- else:
- raise NotImplementedError(f"{search_type} not supported")
-
- return data
-
- @staticmethod
- def _get_valid_github_query(query: str):
- """Check if query is valid and return search types and valid GitHub query."""
- query_terms = shlex.split(query)
- # query must provide repo to load data from
- if len(query_terms) < 1 or "repo:" not in query:
- raise ValueError(
- "GithubLoader requires a search query with `repo:` term. Refer docs - `https://docs.embedchain.ai/data-sources/github`" # noqa: E501
- )
-
- github_query = []
- types = set()
- type_pattern = r"type:([a-zA-Z,]+)"
- for term in query_terms:
- term_match = re.search(type_pattern, term)
- if term_match:
- search_types = term_match.group(1).split(",")
- types.update(search_types)
- else:
- github_query.append(term)
-
- # query must provide search type
- if len(types) == 0:
- raise ValueError(
- "GithubLoader requires a search query with `type:` term. Refer docs - `https://docs.embedchain.ai/data-sources/github`" # noqa: E501
- )
-
- for search_type in search_types:
- if search_type not in VALID_SEARCH_TYPES:
- raise ValueError(
- f"Invalid search type: {search_type}. Valid types are: {', '.join(VALID_SEARCH_TYPES)}"
- )
-
- query = " ".join(github_query)
-
- return types, query
-
- def load_data(self, search_query: str, max_results: int = 1000):
- """Load data from GitHub search query."""
-
- if not self.client:
- raise ValueError(
- "GithubLoader client is not initialized, data will not be loaded. Refer docs - `https://docs.embedchain.ai/data-sources/github`" # noqa: E501
- )
-
- search_types, query = self._get_valid_github_query(search_query)
- logging.info(f"Searching github for query: {query}, with types: {', '.join(search_types)}")
-
- data = []
-
- with concurrent.futures.ThreadPoolExecutor(max_workers=4) as executor:
- futures_map = executor.map(self._search_github_data, search_types, [query] * len(search_types))
- for search_data in tqdm(futures_map, total=len(search_types), desc="Searching data from github"):
- data.extend(search_data)
-
- return {
- "doc_id": hashlib.sha256(query.encode()).hexdigest(),
- "data": data,
- }
diff --git a/embedchain/embedchain/loaders/gmail.py b/embedchain/embedchain/loaders/gmail.py
deleted file mode 100644
index ec62a34b3..000000000
--- a/embedchain/embedchain/loaders/gmail.py
+++ /dev/null
@@ -1,144 +0,0 @@
-import base64
-import hashlib
-import logging
-import os
-from email import message_from_bytes
-from email.utils import parsedate_to_datetime
-from textwrap import dedent
-from typing import Optional
-
-from bs4 import BeautifulSoup
-
-try:
- from google.auth.transport.requests import Request
- from google.oauth2.credentials import Credentials
- from google_auth_oauthlib.flow import InstalledAppFlow
- from googleapiclient.discovery import build
-except ImportError:
- raise ImportError(
- 'Gmail requires extra dependencies. Install with `pip install --upgrade "embedchain[gmail]"`'
- ) from None
-
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-logger = logging.getLogger(__name__)
-
-
-class GmailReader:
- SCOPES = ["https://www.googleapis.com/auth/gmail.readonly"]
-
- def __init__(self, query: str, service=None, results_per_page: int = 10):
- self.query = query
- self.service = service or self._initialize_service()
- self.results_per_page = results_per_page
-
- @staticmethod
- def _initialize_service():
- credentials = GmailReader._get_credentials()
- return build("gmail", "v1", credentials=credentials)
-
- @staticmethod
- def _get_credentials():
- if not os.path.exists("credentials.json"):
- raise FileNotFoundError("Missing 'credentials.json'. Download it from your Google Developer account.")
-
- creds = (
- Credentials.from_authorized_user_file("token.json", GmailReader.SCOPES)
- if os.path.exists("token.json")
- else None
- )
-
- if not creds or not creds.valid:
- if creds and creds.expired and creds.refresh_token:
- creds.refresh(Request())
- else:
- flow = InstalledAppFlow.from_client_secrets_file("credentials.json", GmailReader.SCOPES)
- creds = flow.run_local_server(port=8080)
- with open("token.json", "w") as token:
- token.write(creds.to_json())
- return creds
-
- def load_emails(self) -> list[dict]:
- response = self.service.users().messages().list(userId="me", q=self.query).execute()
- messages = response.get("messages", [])
-
- return [self._parse_email(self._get_email(message["id"])) for message in messages]
-
- def _get_email(self, message_id: str):
- raw_message = self.service.users().messages().get(userId="me", id=message_id, format="raw").execute()
- return base64.urlsafe_b64decode(raw_message["raw"])
-
- def _parse_email(self, raw_email) -> dict:
- mime_msg = message_from_bytes(raw_email)
- return {
- "subject": self._get_header(mime_msg, "Subject"),
- "from": self._get_header(mime_msg, "From"),
- "to": self._get_header(mime_msg, "To"),
- "date": self._format_date(mime_msg),
- "body": self._get_body(mime_msg),
- }
-
- @staticmethod
- def _get_header(mime_msg, header_name: str) -> str:
- return mime_msg.get(header_name, "")
-
- @staticmethod
- def _format_date(mime_msg) -> Optional[str]:
- date_header = GmailReader._get_header(mime_msg, "Date")
- return parsedate_to_datetime(date_header).isoformat() if date_header else None
-
- @staticmethod
- def _get_body(mime_msg) -> str:
- def decode_payload(part):
- charset = part.get_content_charset() or "utf-8"
- try:
- return part.get_payload(decode=True).decode(charset)
- except UnicodeDecodeError:
- return part.get_payload(decode=True).decode(charset, errors="replace")
-
- if mime_msg.is_multipart():
- for part in mime_msg.walk():
- ctype = part.get_content_type()
- cdispo = str(part.get("Content-Disposition"))
-
- if ctype == "text/plain" and "attachment" not in cdispo:
- return decode_payload(part)
- elif ctype == "text/html":
- return decode_payload(part)
- else:
- return decode_payload(mime_msg)
-
- return ""
-
-
-class GmailLoader(BaseLoader):
- def load_data(self, query: str):
- reader = GmailReader(query=query)
- emails = reader.load_emails()
- logger.info(f"Gmail Loader: {len(emails)} emails found for query '{query}'")
-
- data = []
- for email in emails:
- content = self._process_email(email)
- data.append({"content": content, "meta_data": email})
-
- return {"doc_id": self._generate_doc_id(query, data), "data": data}
-
- @staticmethod
- def _process_email(email: dict) -> str:
- content = BeautifulSoup(email["body"], "html.parser").get_text()
- content = clean_string(content)
- return dedent(
- f"""
- Email from '{email['from']}' to '{email['to']}'
- Subject: {email['subject']}
- Date: {email['date']}
- Content: {content}
- """
- )
-
- @staticmethod
- def _generate_doc_id(query: str, data: list[dict]) -> str:
- content_strings = [email["content"] for email in data]
- return hashlib.sha256((query + ", ".join(content_strings)).encode()).hexdigest()
diff --git a/embedchain/embedchain/loaders/google_drive.py b/embedchain/embedchain/loaders/google_drive.py
deleted file mode 100644
index d24046242..000000000
--- a/embedchain/embedchain/loaders/google_drive.py
+++ /dev/null
@@ -1,62 +0,0 @@
-import hashlib
-import re
-
-try:
- from googleapiclient.errors import HttpError
-except ImportError:
- raise ImportError(
- "Google Drive requires extra dependencies. Install with `pip install embedchain[googledrive]`"
- ) from None
-
-from langchain_community.document_loaders import GoogleDriveLoader as Loader
-
-try:
- import unstructured # noqa: F401
- from langchain_community.document_loaders import UnstructuredFileIOLoader
-except ImportError:
- raise ImportError(
- 'Unstructured file requires extra dependencies. Install with `pip install "unstructured[local-inference, all-docs]"`' # noqa: E501
- ) from None
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-
-@register_deserializable
-class GoogleDriveLoader(BaseLoader):
- @staticmethod
- def _get_drive_id_from_url(url: str):
- regex = r"^https:\/\/drive\.google\.com\/drive\/(?:u\/\d+\/)folders\/([a-zA-Z0-9_-]+)$"
- if re.match(regex, url):
- return url.split("/")[-1]
- raise ValueError(
- f"The url provided {url} does not match a google drive folder url. Example drive url: "
- f"https://drive.google.com/drive/u/0/folders/xxxx"
- )
-
- def load_data(self, url: str):
- """Load data from a Google drive folder."""
- folder_id: str = self._get_drive_id_from_url(url)
-
- try:
- loader = Loader(
- folder_id=folder_id,
- recursive=True,
- file_loader_cls=UnstructuredFileIOLoader,
- )
-
- data = []
- all_content = []
-
- docs = loader.load()
- for doc in docs:
- all_content.append(doc.page_content)
- # renames source to url for later use.
- doc.metadata["url"] = doc.metadata.pop("source")
- data.append({"content": doc.page_content, "meta_data": doc.metadata})
-
- doc_id = hashlib.sha256((" ".join(all_content) + url).encode()).hexdigest()
- return {"doc_id": doc_id, "data": data}
-
- except HttpError:
- raise FileNotFoundError("Unable to locate folder or files, check provided drive URL and try again")
diff --git a/embedchain/embedchain/loaders/image.py b/embedchain/embedchain/loaders/image.py
deleted file mode 100644
index 18b31873b..000000000
--- a/embedchain/embedchain/loaders/image.py
+++ /dev/null
@@ -1,50 +0,0 @@
-import base64
-import hashlib
-import os
-from pathlib import Path
-
-from openai import OpenAI
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-DESCRIBE_IMAGE_PROMPT = "Describe the image:"
-
-
-@register_deserializable
-class ImageLoader(BaseLoader):
- def __init__(self, max_tokens: int = 500, api_key: str = None, prompt: str = None):
- super().__init__()
- self.custom_prompt = prompt or DESCRIBE_IMAGE_PROMPT
- self.max_tokens = max_tokens
- self.api_key = api_key or os.environ["OPENAI_API_KEY"]
- self.client = OpenAI(api_key=self.api_key)
-
- @staticmethod
- def _encode_image(image_path: str):
- with open(image_path, "rb") as image_file:
- return base64.b64encode(image_file.read()).decode("utf-8")
-
- def _create_completion_request(self, content: str):
- return self.client.chat.completions.create(
- model="gpt-4o", messages=[{"role": "user", "content": content}], max_tokens=self.max_tokens
- )
-
- def _process_url(self, url: str):
- if url.startswith("http"):
- return [{"type": "text", "text": self.custom_prompt}, {"type": "image_url", "image_url": {"url": url}}]
- elif Path(url).is_file():
- extension = Path(url).suffix.lstrip(".")
- encoded_image = self._encode_image(url)
- image_data = f"data:image/{extension};base64,{encoded_image}"
- return [{"type": "text", "text": self.custom_prompt}, {"type": "image", "image_url": {"url": image_data}}]
- else:
- raise ValueError(f"Invalid URL or file path: {url}")
-
- def load_data(self, url: str):
- content = self._process_url(url)
- response = self._create_completion_request(content)
- content = response.choices[0].message.content
-
- doc_id = hashlib.sha256((content + url).encode()).hexdigest()
- return {"doc_id": doc_id, "data": [{"content": content, "meta_data": {"url": url, "type": "image"}}]}
diff --git a/embedchain/embedchain/loaders/json.py b/embedchain/embedchain/loaders/json.py
deleted file mode 100644
index 587aa1492..000000000
--- a/embedchain/embedchain/loaders/json.py
+++ /dev/null
@@ -1,93 +0,0 @@
-import hashlib
-import json
-import os
-import re
-from typing import Union
-
-import requests
-
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string, is_valid_json_string
-
-
-class JSONReader:
- def __init__(self) -> None:
- """Initialize the JSONReader."""
- pass
-
- @staticmethod
- def load_data(json_data: Union[dict, str]) -> list[str]:
- """Load data from a JSON structure.
-
- Args:
- json_data (Union[dict, str]): The JSON data to load.
-
- Returns:
- list[str]: A list of strings representing the leaf nodes of the JSON.
- """
- if isinstance(json_data, str):
- json_data = json.loads(json_data)
- else:
- json_data = json_data
-
- json_output = json.dumps(json_data, indent=0)
- lines = json_output.split("\n")
- useful_lines = [line for line in lines if not re.match(r"^[{}\[\],]*$", line)]
- return ["\n".join(useful_lines)]
-
-
-VALID_URL_PATTERN = (
- "^https?://(?:www\.)?(?:\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}|[a-zA-Z0-9.-]+)(?::\d+)?/(?:[^/\s]+/)*[^/\s]+\.json$"
-)
-
-
-class JSONLoader(BaseLoader):
- @staticmethod
- def _check_content(content):
- if not isinstance(content, str):
- raise ValueError(
- "Invaid content input. \
- If you want to upload (list, dict, etc.), do \
- `json.dump(data, indent=0)` and add the stringified JSON. \
- Check - `https://docs.embedchain.ai/data-sources/json`"
- )
-
- @staticmethod
- def load_data(content):
- """Load a json file. Each data point is a key value pair."""
-
- JSONLoader._check_content(content)
- loader = JSONReader()
-
- data = []
- data_content = []
-
- content_url_str = content
-
- if os.path.isfile(content):
- with open(content, "r", encoding="utf-8") as json_file:
- json_data = json.load(json_file)
- elif re.match(VALID_URL_PATTERN, content):
- response = requests.get(content)
- if response.status_code == 200:
- json_data = response.json()
- else:
- raise ValueError(
- f"Loading data from the given url: {content} failed. \
- Make sure the url is working."
- )
- elif is_valid_json_string(content):
- json_data = content
- content_url_str = hashlib.sha256((content).encode("utf-8")).hexdigest()
- else:
- raise ValueError(f"Invalid content to load json data from: {content}")
-
- docs = loader.load_data(json_data)
- for doc in docs:
- text = doc if isinstance(doc, str) else doc["text"]
- doc_content = clean_string(text)
- data.append({"content": doc_content, "meta_data": {"url": content_url_str}})
- data_content.append(doc_content)
-
- doc_id = hashlib.sha256((content_url_str + ", ".join(data_content)).encode()).hexdigest()
- return {"doc_id": doc_id, "data": data}
diff --git a/embedchain/embedchain/loaders/local_qna_pair.py b/embedchain/embedchain/loaders/local_qna_pair.py
deleted file mode 100644
index c93adfdae..000000000
--- a/embedchain/embedchain/loaders/local_qna_pair.py
+++ /dev/null
@@ -1,24 +0,0 @@
-import hashlib
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-
-@register_deserializable
-class LocalQnaPairLoader(BaseLoader):
- def load_data(self, content):
- """Load data from a local QnA pair."""
- question, answer = content
- content = f"Q: {question}\nA: {answer}"
- url = "local"
- metadata = {"url": url, "question": question}
- doc_id = hashlib.sha256((content + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": [
- {
- "content": content,
- "meta_data": metadata,
- }
- ],
- }
diff --git a/embedchain/embedchain/loaders/local_text.py b/embedchain/embedchain/loaders/local_text.py
deleted file mode 100644
index 98a98cd67..000000000
--- a/embedchain/embedchain/loaders/local_text.py
+++ /dev/null
@@ -1,24 +0,0 @@
-import hashlib
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-
-@register_deserializable
-class LocalTextLoader(BaseLoader):
- def load_data(self, content):
- """Load data from a local text file."""
- url = "local"
- metadata = {
- "url": url,
- }
- doc_id = hashlib.sha256((content + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": [
- {
- "content": content,
- "meta_data": metadata,
- }
- ],
- }
diff --git a/embedchain/embedchain/loaders/mdx.py b/embedchain/embedchain/loaders/mdx.py
deleted file mode 100644
index 42b9b7fee..000000000
--- a/embedchain/embedchain/loaders/mdx.py
+++ /dev/null
@@ -1,25 +0,0 @@
-import hashlib
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-
-@register_deserializable
-class MdxLoader(BaseLoader):
- def load_data(self, url):
- """Load data from a mdx file."""
- with open(url, "r", encoding="utf-8") as infile:
- content = infile.read()
- metadata = {
- "url": url,
- }
- doc_id = hashlib.sha256((content + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": [
- {
- "content": content,
- "meta_data": metadata,
- }
- ],
- }
diff --git a/embedchain/embedchain/loaders/mysql.py b/embedchain/embedchain/loaders/mysql.py
deleted file mode 100644
index fd5b38ac2..000000000
--- a/embedchain/embedchain/loaders/mysql.py
+++ /dev/null
@@ -1,67 +0,0 @@
-import hashlib
-import logging
-from typing import Any, Optional
-
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-logger = logging.getLogger(__name__)
-
-
-class MySQLLoader(BaseLoader):
- def __init__(self, config: Optional[dict[str, Any]]):
- super().__init__()
- if not config:
- raise ValueError(
- f"Invalid sql config: {config}.",
- "Provide the correct config, refer `https://docs.embedchain.ai/data-sources/mysql`.",
- )
-
- self.config = config
- self.connection = None
- self.cursor = None
- self._setup_loader(config=config)
-
- def _setup_loader(self, config: dict[str, Any]):
- try:
- import mysql.connector as sqlconnector
- except ImportError as e:
- raise ImportError(
- "Unable to import required packages for MySQL loader. Run `pip install --upgrade 'embedchain[mysql]'`." # noqa: E501
- ) from e
-
- try:
- self.connection = sqlconnector.connection.MySQLConnection(**config)
- self.cursor = self.connection.cursor()
- except (sqlconnector.Error, IOError) as err:
- logger.info(f"Connection failed: {err}")
- raise ValueError(
- f"Unable to connect with the given config: {config}.",
- "Please provide the correct configuration to load data from you MySQL DB. \
- Refer `https://docs.embedchain.ai/data-sources/mysql`.",
- )
-
- @staticmethod
- def _check_query(query):
- if not isinstance(query, str):
- raise ValueError(
- f"Invalid mysql query: {query}",
- "Provide the valid query to add from mysql, \
- make sure you are following `https://docs.embedchain.ai/data-sources/mysql`",
- )
-
- def load_data(self, query):
- self._check_query(query=query)
- data = []
- data_content = []
- self.cursor.execute(query)
- rows = self.cursor.fetchall()
- for row in rows:
- doc_content = clean_string(str(row))
- data.append({"content": doc_content, "meta_data": {"url": query}})
- data_content.append(doc_content)
- doc_id = hashlib.sha256((query + ", ".join(data_content)).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": data,
- }
diff --git a/embedchain/embedchain/loaders/notion.py b/embedchain/embedchain/loaders/notion.py
deleted file mode 100644
index 2a3363818..000000000
--- a/embedchain/embedchain/loaders/notion.py
+++ /dev/null
@@ -1,121 +0,0 @@
-import hashlib
-import logging
-import os
-from typing import Any, Optional
-
-import requests
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-logger = logging.getLogger(__name__)
-
-
-class NotionDocument:
- """
- A simple Document class to hold the text and additional information of a page.
- """
-
- def __init__(self, text: str, extra_info: dict[str, Any]):
- self.text = text
- self.extra_info = extra_info
-
-
-class NotionPageLoader:
- """
- Notion Page Loader.
- Reads a set of Notion pages.
- """
-
- BLOCK_CHILD_URL_TMPL = "https://api.notion.com/v1/blocks/{block_id}/children"
-
- def __init__(self, integration_token: Optional[str] = None) -> None:
- """Initialize with Notion integration token."""
- if integration_token is None:
- integration_token = os.getenv("NOTION_INTEGRATION_TOKEN")
- if integration_token is None:
- raise ValueError(
- "Must specify `integration_token` or set environment " "variable `NOTION_INTEGRATION_TOKEN`."
- )
- self.token = integration_token
- self.headers = {
- "Authorization": "Bearer " + self.token,
- "Content-Type": "application/json",
- "Notion-Version": "2022-06-28",
- }
-
- def _read_block(self, block_id: str, num_tabs: int = 0) -> str:
- """Read a block from Notion."""
- done = False
- result_lines_arr = []
- cur_block_id = block_id
- while not done:
- block_url = self.BLOCK_CHILD_URL_TMPL.format(block_id=cur_block_id)
- res = requests.get(block_url, headers=self.headers)
- data = res.json()
-
- for result in data["results"]:
- result_type = result["type"]
- result_obj = result[result_type]
-
- cur_result_text_arr = []
- if "rich_text" in result_obj:
- for rich_text in result_obj["rich_text"]:
- if "text" in rich_text:
- text = rich_text["text"]["content"]
- prefix = "\t" * num_tabs
- cur_result_text_arr.append(prefix + text)
-
- result_block_id = result["id"]
- has_children = result["has_children"]
- if has_children:
- children_text = self._read_block(result_block_id, num_tabs=num_tabs + 1)
- cur_result_text_arr.append(children_text)
-
- cur_result_text = "\n".join(cur_result_text_arr)
- result_lines_arr.append(cur_result_text)
-
- if data["next_cursor"] is None:
- done = True
- else:
- cur_block_id = data["next_cursor"]
-
- result_lines = "\n".join(result_lines_arr)
- return result_lines
-
- def load_data(self, page_ids: list[str]) -> list[NotionDocument]:
- """Load data from the given list of page IDs."""
- docs = []
- for page_id in page_ids:
- page_text = self._read_block(page_id)
- docs.append(NotionDocument(text=page_text, extra_info={"page_id": page_id}))
- return docs
-
-
-@register_deserializable
-class NotionLoader(BaseLoader):
- def load_data(self, source):
- """Load data from a Notion URL."""
-
- id = source[-32:]
- formatted_id = f"{id[:8]}-{id[8:12]}-{id[12:16]}-{id[16:20]}-{id[20:]}"
- logger.debug(f"Extracted notion page id as: {formatted_id}")
-
- integration_token = os.getenv("NOTION_INTEGRATION_TOKEN")
- reader = NotionPageLoader(integration_token=integration_token)
- documents = reader.load_data(page_ids=[formatted_id])
-
- raw_text = documents[0].text
-
- text = clean_string(raw_text)
- doc_id = hashlib.sha256((text + source).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": [
- {
- "content": text,
- "meta_data": {"url": f"notion-{formatted_id}"},
- }
- ],
- }
diff --git a/embedchain/embedchain/loaders/openapi.py b/embedchain/embedchain/loaders/openapi.py
deleted file mode 100644
index 18983b9a3..000000000
--- a/embedchain/embedchain/loaders/openapi.py
+++ /dev/null
@@ -1,42 +0,0 @@
-import hashlib
-from io import StringIO
-from urllib.parse import urlparse
-
-import requests
-import yaml
-
-from embedchain.loaders.base_loader import BaseLoader
-
-
-class OpenAPILoader(BaseLoader):
- @staticmethod
- def _get_file_content(content):
- url = urlparse(content)
- if all([url.scheme, url.netloc]) and url.scheme not in ["file", "http", "https"]:
- raise ValueError("Not a valid URL.")
-
- if url.scheme in ["http", "https"]:
- response = requests.get(content)
- response.raise_for_status()
- return StringIO(response.text)
- elif url.scheme == "file":
- path = url.path
- return open(path)
- else:
- return open(content)
-
- @staticmethod
- def load_data(content):
- """Load yaml file of openapi. Each pair is a document."""
- data = []
- file_path = content
- data_content = []
- with OpenAPILoader._get_file_content(content=content) as file:
- yaml_data = yaml.load(file, Loader=yaml.SafeLoader)
- for i, (key, value) in enumerate(yaml_data.items()):
- string_data = f"{key}: {value}"
- metadata = {"url": file_path, "row": i + 1}
- data.append({"content": string_data, "meta_data": metadata})
- data_content.append(string_data)
- doc_id = hashlib.sha256((content + ", ".join(data_content)).encode()).hexdigest()
- return {"doc_id": doc_id, "data": data}
diff --git a/embedchain/embedchain/loaders/pdf_file.py b/embedchain/embedchain/loaders/pdf_file.py
deleted file mode 100644
index a7f6d5540..000000000
--- a/embedchain/embedchain/loaders/pdf_file.py
+++ /dev/null
@@ -1,39 +0,0 @@
-import hashlib
-
-from langchain_community.document_loaders import PyPDFLoader
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-
-@register_deserializable
-class PdfFileLoader(BaseLoader):
- def load_data(self, url):
- """Load data from a PDF file."""
- headers = {
- "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.102 Safari/537.36", # noqa:E501
- }
- loader = PyPDFLoader(url, headers=headers)
- data = []
- all_content = []
- pages = loader.load_and_split()
- if not len(pages):
- raise ValueError("No data found")
- for page in pages:
- content = page.page_content
- content = clean_string(content)
- metadata = page.metadata
- metadata["url"] = url
- data.append(
- {
- "content": content,
- "meta_data": metadata,
- }
- )
- all_content.append(content)
- doc_id = hashlib.sha256((" ".join(all_content) + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": data,
- }
diff --git a/embedchain/embedchain/loaders/postgres.py b/embedchain/embedchain/loaders/postgres.py
deleted file mode 100644
index 2ef396f9d..000000000
--- a/embedchain/embedchain/loaders/postgres.py
+++ /dev/null
@@ -1,73 +0,0 @@
-import hashlib
-import logging
-from typing import Any, Optional
-
-from embedchain.loaders.base_loader import BaseLoader
-
-logger = logging.getLogger(__name__)
-
-
-class PostgresLoader(BaseLoader):
- def __init__(self, config: Optional[dict[str, Any]] = None):
- super().__init__()
- if not config:
- raise ValueError(f"Must provide the valid config. Received: {config}")
-
- self.connection = None
- self.cursor = None
- self._setup_loader(config=config)
-
- def _setup_loader(self, config: dict[str, Any]):
- try:
- import psycopg
- except ImportError as e:
- raise ImportError(
- "Unable to import required packages. \
- Run `pip install --upgrade 'embedchain[postgres]'`"
- ) from e
-
- if "url" in config:
- config_info = config.get("url")
- else:
- conn_params = []
- for key, value in config.items():
- conn_params.append(f"{key}={value}")
- config_info = " ".join(conn_params)
-
- logger.info(f"Connecting to postrgres sql: {config_info}")
- self.connection = psycopg.connect(conninfo=config_info)
- self.cursor = self.connection.cursor()
-
- @staticmethod
- def _check_query(query):
- if not isinstance(query, str):
- raise ValueError(
- f"Invalid postgres query: {query}. Provide the valid source to add from postgres, make sure you are following `https://docs.embedchain.ai/data-sources/postgres`", # noqa:E501
- )
-
- def load_data(self, query):
- self._check_query(query)
- try:
- data = []
- data_content = []
- self.cursor.execute(query)
- results = self.cursor.fetchall()
- for result in results:
- doc_content = str(result)
- data.append({"content": doc_content, "meta_data": {"url": query}})
- data_content.append(doc_content)
- doc_id = hashlib.sha256((query + ", ".join(data_content)).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": data,
- }
- except Exception as e:
- raise ValueError(f"Failed to load data using query={query} with: {e}")
-
- def close_connection(self):
- if self.cursor:
- self.cursor.close()
- self.cursor = None
- if self.connection:
- self.connection.close()
- self.connection = None
diff --git a/embedchain/embedchain/loaders/rss_feed.py b/embedchain/embedchain/loaders/rss_feed.py
deleted file mode 100644
index bc17c68bc..000000000
--- a/embedchain/embedchain/loaders/rss_feed.py
+++ /dev/null
@@ -1,54 +0,0 @@
-import hashlib
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-
-@register_deserializable
-class RSSFeedLoader(BaseLoader):
- """Loader for RSS Feed."""
-
- def load_data(self, url):
- """Load data from a rss feed."""
- output = self.get_rss_content(url)
- doc_id = hashlib.sha256((str(output) + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": output,
- }
-
- @staticmethod
- def serialize_metadata(metadata):
- for key, value in metadata.items():
- if not isinstance(value, (str, int, float, bool)):
- metadata[key] = str(value)
-
- return metadata
-
- @staticmethod
- def get_rss_content(url: str):
- try:
- from langchain_community.document_loaders import (
- RSSFeedLoader as LangchainRSSFeedLoader,
- )
- except ImportError:
- raise ImportError(
- """RSSFeedLoader file requires extra dependencies.
- Install with `pip install feedparser==6.0.10 newspaper3k==0.2.8 listparser==0.19`"""
- ) from None
-
- output = []
- loader = LangchainRSSFeedLoader(urls=[url])
- data = loader.load()
-
- for entry in data:
- metadata = RSSFeedLoader.serialize_metadata(entry.metadata)
- metadata.update({"url": url})
- output.append(
- {
- "content": entry.page_content,
- "meta_data": metadata,
- }
- )
-
- return output
diff --git a/embedchain/embedchain/loaders/sitemap.py b/embedchain/embedchain/loaders/sitemap.py
deleted file mode 100644
index 098ca06df..000000000
--- a/embedchain/embedchain/loaders/sitemap.py
+++ /dev/null
@@ -1,79 +0,0 @@
-import concurrent.futures
-import hashlib
-import logging
-import os
-from urllib.parse import urlparse
-
-import requests
-from tqdm import tqdm
-
-try:
- from bs4 import BeautifulSoup
- from bs4.builder import ParserRejectedMarkup
-except ImportError:
- raise ImportError(
- "Sitemap requires extra dependencies. Install with `pip install beautifulsoup4==4.12.3`"
- ) from None
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.loaders.web_page import WebPageLoader
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class SitemapLoader(BaseLoader):
- """
- This method takes a sitemap URL or local file path as input and retrieves
- all the URLs to use the WebPageLoader to load content
- of each page.
- """
-
- def load_data(self, sitemap_source):
- output = []
- web_page_loader = WebPageLoader()
- headers = {
- "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.102 Safari/537.36", # noqa:E501
- }
-
- if urlparse(sitemap_source).scheme in ("http", "https"):
- try:
- response = requests.get(sitemap_source, headers=headers)
- response.raise_for_status()
- soup = BeautifulSoup(response.text, "xml")
- except requests.RequestException as e:
- logger.error(f"Error fetching sitemap from URL: {e}")
- return
- elif os.path.isfile(sitemap_source):
- with open(sitemap_source, "r") as file:
- soup = BeautifulSoup(file, "xml")
- else:
- raise ValueError("Invalid sitemap source. Please provide a valid URL or local file path.")
-
- links = [link.text for link in soup.find_all("loc") if link.parent.name == "url"]
- if len(links) == 0:
- links = [link.text for link in soup.find_all("loc")]
-
- doc_id = hashlib.sha256((" ".join(links) + sitemap_source).encode()).hexdigest()
-
- def load_web_page(link):
- try:
- loader_data = web_page_loader.load_data(link)
- return loader_data.get("data")
- except ParserRejectedMarkup as e:
- logger.error(f"Failed to parse {link}: {e}")
- return None
-
- with concurrent.futures.ThreadPoolExecutor() as executor:
- future_to_link = {executor.submit(load_web_page, link): link for link in links}
- for future in tqdm(concurrent.futures.as_completed(future_to_link), total=len(links), desc="Loading pages"):
- link = future_to_link[future]
- try:
- data = future.result()
- if data:
- output.extend(data)
- except Exception as e:
- logger.error(f"Error loading page {link}: {e}")
-
- return {"doc_id": doc_id, "data": output}
diff --git a/embedchain/embedchain/loaders/slack.py b/embedchain/embedchain/loaders/slack.py
deleted file mode 100644
index 6fb6e9db8..000000000
--- a/embedchain/embedchain/loaders/slack.py
+++ /dev/null
@@ -1,115 +0,0 @@
-import hashlib
-import logging
-import os
-import ssl
-from typing import Any, Optional
-
-import certifi
-
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-SLACK_API_BASE_URL = "https://www.slack.com/api/"
-
-logger = logging.getLogger(__name__)
-
-
-class SlackLoader(BaseLoader):
- def __init__(self, config: Optional[dict[str, Any]] = None):
- super().__init__()
-
- self.config = config if config else {}
-
- if "base_url" not in self.config:
- self.config["base_url"] = SLACK_API_BASE_URL
-
- self.client = None
- self._setup_loader(self.config)
-
- def _setup_loader(self, config: dict[str, Any]):
- try:
- from slack_sdk import WebClient
- except ImportError as e:
- raise ImportError(
- "Slack loader requires extra dependencies. \
- Install with `pip install --upgrade embedchain[slack]`"
- ) from e
-
- if os.getenv("SLACK_USER_TOKEN") is None:
- raise ValueError(
- "SLACK_USER_TOKEN environment variables not provided. Check `https://docs.embedchain.ai/data-sources/slack` to learn more." # noqa:E501
- )
-
- logger.info(f"Creating Slack Loader with config: {config}")
- # get slack client config params
- slack_bot_token = os.getenv("SLACK_USER_TOKEN")
- ssl_cert = ssl.create_default_context(cafile=certifi.where())
- base_url = config.get("base_url", SLACK_API_BASE_URL)
- headers = config.get("headers")
- # for Org-Wide App
- team_id = config.get("team_id")
-
- self.client = WebClient(
- token=slack_bot_token,
- base_url=base_url,
- ssl=ssl_cert,
- headers=headers,
- team_id=team_id,
- )
- logger.info("Slack Loader setup successful!")
-
- @staticmethod
- def _check_query(query):
- if not isinstance(query, str):
- raise ValueError(
- f"Invalid query passed to Slack loader, found: {query}. Check `https://docs.embedchain.ai/data-sources/slack` to learn more." # noqa:E501
- )
-
- def load_data(self, query):
- self._check_query(query)
- try:
- data = []
- data_content = []
-
- logger.info(f"Searching slack conversations for query: {query}")
- results = self.client.search_messages(
- query=query,
- sort="timestamp",
- sort_dir="desc",
- count=self.config.get("count", 100),
- )
-
- messages = results.get("messages")
- num_message = len(messages)
- logger.info(f"Found {num_message} messages for query: {query}")
-
- matches = messages.get("matches", [])
- for message in matches:
- url = message.get("permalink")
- text = message.get("text")
- content = clean_string(text)
-
- message_meta_data_keys = ["iid", "team", "ts", "type", "user", "username"]
- metadata = {}
- for key in message.keys():
- if key in message_meta_data_keys:
- metadata[key] = message.get(key)
- metadata.update({"url": url})
-
- data.append(
- {
- "content": content,
- "meta_data": metadata,
- }
- )
- data_content.append(content)
- doc_id = hashlib.md5((query + ", ".join(data_content)).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": data,
- }
- except Exception as e:
- logger.warning(f"Error in loading slack data: {e}")
- raise ValueError(
- f"Error in loading slack data: {e}. Check `https://docs.embedchain.ai/data-sources/slack` to learn more." # noqa:E501
- ) from e
diff --git a/embedchain/embedchain/loaders/substack.py b/embedchain/embedchain/loaders/substack.py
deleted file mode 100644
index 15c08a5bb..000000000
--- a/embedchain/embedchain/loaders/substack.py
+++ /dev/null
@@ -1,107 +0,0 @@
-import hashlib
-import logging
-import time
-from xml.etree import ElementTree
-
-import requests
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import is_readable
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class SubstackLoader(BaseLoader):
- """
- This loader is used to load data from Substack URLs.
- """
-
- def load_data(self, url: str):
- try:
- from bs4 import BeautifulSoup
- from bs4.builder import ParserRejectedMarkup
- except ImportError:
- raise ImportError(
- "Substack requires extra dependencies. Install with `pip install beautifulsoup4==4.12.3`"
- ) from None
-
- if not url.endswith("sitemap.xml"):
- url = url + "/sitemap.xml"
-
- output = []
- response = requests.get(url)
-
- try:
- response.raise_for_status()
- except requests.exceptions.HTTPError as e:
- raise ValueError(
- f"""
- Failed to load {url}: {e}. Please use the root substack URL. For example, https://example.substack.com
- """
- )
-
- try:
- ElementTree.fromstring(response.content)
- except ElementTree.ParseError:
- raise ValueError(
- f"""
- Failed to parse {url}. Please use the root substack URL. For example, https://example.substack.com
- """
- )
-
- soup = BeautifulSoup(response.text, "xml")
- links = [link.text for link in soup.find_all("loc") if link.parent.name == "url" and "/p/" in link.text]
- if len(links) == 0:
- links = [link.text for link in soup.find_all("loc") if "/p/" in link.text]
-
- doc_id = hashlib.sha256((" ".join(links) + url).encode()).hexdigest()
-
- def serialize_response(soup: BeautifulSoup):
- data = {}
-
- h1_els = soup.find_all("h1")
- if h1_els is not None and len(h1_els) > 0:
- data["title"] = h1_els[1].text
-
- description_el = soup.find("meta", {"name": "description"})
- if description_el is not None:
- data["description"] = description_el["content"]
-
- content_el = soup.find("div", {"class": "available-content"})
- if content_el is not None:
- data["content"] = content_el.text
-
- like_btn = soup.find("div", {"class": "like-button-container"})
- if like_btn is not None:
- no_of_likes_div = like_btn.find("div", {"class": "label"})
- if no_of_likes_div is not None:
- data["no_of_likes"] = no_of_likes_div.text
-
- return data
-
- def load_link(link: str):
- try:
- substack_data = requests.get(link)
- substack_data.raise_for_status()
-
- soup = BeautifulSoup(substack_data.text, "html.parser")
- data = serialize_response(soup)
- data = str(data)
- if is_readable(data):
- return data
- else:
- logger.warning(f"Page is not readable (too many invalid characters): {link}")
- except ParserRejectedMarkup as e:
- logger.error(f"Failed to parse {link}: {e}")
- return None
-
- for link in links:
- data = load_link(link)
- if data:
- output.append({"content": data, "meta_data": {"url": link}})
- # TODO: allow users to configure this
- time.sleep(1.0) # added to avoid rate limiting
-
- return {"doc_id": doc_id, "data": output}
diff --git a/embedchain/embedchain/loaders/text_file.py b/embedchain/embedchain/loaders/text_file.py
deleted file mode 100644
index bc7fb4b09..000000000
--- a/embedchain/embedchain/loaders/text_file.py
+++ /dev/null
@@ -1,30 +0,0 @@
-import hashlib
-import os
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-
-
-@register_deserializable
-class TextFileLoader(BaseLoader):
- def load_data(self, url: str):
- """Load data from a text file located at a local path."""
- if not os.path.exists(url):
- raise FileNotFoundError(f"The file at {url} does not exist.")
-
- with open(url, "r", encoding="utf-8") as file:
- content = file.read()
-
- doc_id = hashlib.sha256((content + url).encode()).hexdigest()
-
- metadata = {"url": url, "file_size": os.path.getsize(url), "file_type": url.split(".")[-1]}
-
- return {
- "doc_id": doc_id,
- "data": [
- {
- "content": content,
- "meta_data": metadata,
- }
- ],
- }
diff --git a/embedchain/embedchain/loaders/unstructured_file.py b/embedchain/embedchain/loaders/unstructured_file.py
deleted file mode 100644
index 856ac888b..000000000
--- a/embedchain/embedchain/loaders/unstructured_file.py
+++ /dev/null
@@ -1,42 +0,0 @@
-import hashlib
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-
-@register_deserializable
-class UnstructuredLoader(BaseLoader):
- def load_data(self, url):
- """Load data from an Unstructured file."""
- try:
- import unstructured # noqa: F401
- from langchain_community.document_loaders import UnstructuredFileLoader
- except ImportError:
- raise ImportError(
- 'Unstructured file requires extra dependencies. Install with `pip install "unstructured[local-inference, all-docs]"`' # noqa: E501
- ) from None
-
- loader = UnstructuredFileLoader(url)
- data = []
- all_content = []
- pages = loader.load_and_split()
- if not len(pages):
- raise ValueError("No data found")
- for page in pages:
- content = page.page_content
- content = clean_string(content)
- metadata = page.metadata
- metadata["url"] = url
- data.append(
- {
- "content": content,
- "meta_data": metadata,
- }
- )
- all_content.append(content)
- doc_id = hashlib.sha256((" ".join(all_content) + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": data,
- }
diff --git a/embedchain/embedchain/loaders/web_page.py b/embedchain/embedchain/loaders/web_page.py
deleted file mode 100644
index 848bc2038..000000000
--- a/embedchain/embedchain/loaders/web_page.py
+++ /dev/null
@@ -1,126 +0,0 @@
-import hashlib
-import logging
-from typing import Any, Optional
-
-import requests
-
-try:
- from bs4 import BeautifulSoup
-except ImportError:
- raise ImportError(
- "Webpage requires extra dependencies. Install with `pip install beautifulsoup4==4.12.3`"
- ) from None
-
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class WebPageLoader(BaseLoader):
- # Shared session for all instances
- _session = requests.Session()
-
- def load_data(self, url, **kwargs: Optional[dict[str, Any]]):
- """Load data from a web page using a shared requests' session."""
- all_references = False
- for key, value in kwargs.items():
- if key == "all_references":
- all_references = kwargs["all_references"]
- headers = {
- "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.102 Safari/537.36", # noqa:E501
- }
- response = self._session.get(url, headers=headers, timeout=30)
- response.raise_for_status()
- data = response.content
- reference_links = self.fetch_reference_links(response)
- if all_references:
- for i in reference_links:
- try:
- response = self._session.get(i, headers=headers, timeout=30)
- response.raise_for_status()
- data += response.content
- except Exception as e:
- logging.error(f"Failed to add URL {url}: {e}")
- continue
-
- content = self._get_clean_content(data, url)
-
- metadata = {"url": url}
-
- doc_id = hashlib.sha256((content + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": [
- {
- "content": content,
- "meta_data": metadata,
- }
- ],
- }
-
- @staticmethod
- def _get_clean_content(html, url) -> str:
- soup = BeautifulSoup(html, "html.parser")
- original_size = len(str(soup.get_text()))
-
- tags_to_exclude = [
- "nav",
- "aside",
- "form",
- "header",
- "noscript",
- "svg",
- "canvas",
- "footer",
- "script",
- "style",
- ]
- for tag in soup(tags_to_exclude):
- tag.decompose()
-
- ids_to_exclude = ["sidebar", "main-navigation", "menu-main-menu"]
- for id_ in ids_to_exclude:
- tags = soup.find_all(id=id_)
- for tag in tags:
- tag.decompose()
-
- classes_to_exclude = [
- "elementor-location-header",
- "navbar-header",
- "nav",
- "header-sidebar-wrapper",
- "blog-sidebar-wrapper",
- "related-posts",
- ]
- for class_name in classes_to_exclude:
- tags = soup.find_all(class_=class_name)
- for tag in tags:
- tag.decompose()
-
- content = soup.get_text()
- content = clean_string(content)
-
- cleaned_size = len(content)
- if original_size != 0:
- logger.info(
- f"[{url}] Cleaned page size: {cleaned_size} characters, down from {original_size} (shrunk: {original_size-cleaned_size} chars, {round((1-(cleaned_size/original_size)) * 100, 2)}%)" # noqa:E501
- )
-
- return content
-
- @classmethod
- def close_session(cls):
- cls._session.close()
-
- def fetch_reference_links(self, response):
- if response.status_code == 200:
- soup = BeautifulSoup(response.content, "html.parser")
- a_tags = soup.find_all("a", href=True)
- reference_links = [a["href"] for a in a_tags if a["href"].startswith("http")]
- return reference_links
- else:
- print(f"Failed to retrieve the page. Status code: {response.status_code}")
- return []
diff --git a/embedchain/embedchain/loaders/xml.py b/embedchain/embedchain/loaders/xml.py
deleted file mode 100644
index 0c2c8c748..000000000
--- a/embedchain/embedchain/loaders/xml.py
+++ /dev/null
@@ -1,31 +0,0 @@
-import hashlib
-
-try:
- import unstructured # noqa: F401
- from langchain_community.document_loaders import UnstructuredXMLLoader
-except ImportError:
- raise ImportError(
- 'XML file requires extra dependencies. Install with `pip install "unstructured[local-inference, all-docs]"`'
- ) from None
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-
-@register_deserializable
-class XmlLoader(BaseLoader):
- def load_data(self, xml_url):
- """Load data from a XML file."""
- loader = UnstructuredXMLLoader(xml_url)
- data = loader.load()
- content = data[0].page_content
- content = clean_string(content)
- metadata = data[0].metadata
- metadata["url"] = metadata["source"]
- del metadata["source"]
- output = [{"content": content, "meta_data": metadata}]
- doc_id = hashlib.sha256((content + xml_url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": output,
- }
diff --git a/embedchain/embedchain/loaders/youtube_channel.py b/embedchain/embedchain/loaders/youtube_channel.py
deleted file mode 100644
index ab235e19a..000000000
--- a/embedchain/embedchain/loaders/youtube_channel.py
+++ /dev/null
@@ -1,79 +0,0 @@
-import concurrent.futures
-import hashlib
-import logging
-
-from tqdm import tqdm
-
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.loaders.youtube_video import YoutubeVideoLoader
-
-logger = logging.getLogger(__name__)
-
-
-class YoutubeChannelLoader(BaseLoader):
- """Loader for youtube channel."""
-
- def load_data(self, channel_name):
- try:
- import yt_dlp
- except ImportError as e:
- raise ValueError(
- "YoutubeChannelLoader requires extra dependencies. Install with `pip install yt_dlp==2023.11.14 youtube-transcript-api==0.6.1`" # noqa: E501
- ) from e
-
- data = []
- data_urls = []
- youtube_url = f"https://www.youtube.com/{channel_name}/videos"
- youtube_video_loader = YoutubeVideoLoader()
-
- def _get_yt_video_links():
- try:
- ydl_opts = {
- "quiet": True,
- "extract_flat": True,
- }
- with yt_dlp.YoutubeDL(ydl_opts) as ydl:
- info_dict = ydl.extract_info(youtube_url, download=False)
- if "entries" in info_dict:
- videos = [entry["url"] for entry in info_dict["entries"]]
- return videos
- except Exception:
- logger.error(f"Failed to fetch youtube videos for channel: {channel_name}")
- return []
-
- def _load_yt_video(video_link):
- try:
- each_load_data = youtube_video_loader.load_data(video_link)
- if each_load_data:
- return each_load_data.get("data")
- except Exception as e:
- logger.error(f"Failed to load youtube video {video_link}: {e}")
- return None
-
- def _add_youtube_channel():
- video_links = _get_yt_video_links()
- logger.info("Loading videos from youtube channel...")
- with concurrent.futures.ThreadPoolExecutor() as executor:
- # Submitting all tasks and storing the future object with the video link
- future_to_video = {
- executor.submit(_load_yt_video, video_link): video_link for video_link in video_links
- }
-
- for future in tqdm(
- concurrent.futures.as_completed(future_to_video), total=len(video_links), desc="Processing videos"
- ):
- video = future_to_video[future]
- try:
- results = future.result()
- if results:
- data.extend(results)
- data_urls.extend([result.get("meta_data").get("url") for result in results])
- except Exception as e:
- logger.error(f"Failed to process youtube video {video}: {e}")
-
- _add_youtube_channel()
- doc_id = hashlib.sha256((youtube_url + ", ".join(data_urls)).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": data,
- }
diff --git a/embedchain/embedchain/loaders/youtube_video.py b/embedchain/embedchain/loaders/youtube_video.py
deleted file mode 100644
index 44acc0fcf..000000000
--- a/embedchain/embedchain/loaders/youtube_video.py
+++ /dev/null
@@ -1,57 +0,0 @@
-import hashlib
-import json
-import logging
-
-try:
- from youtube_transcript_api import YouTubeTranscriptApi
-except ImportError:
- raise ImportError("YouTube video requires extra dependencies. Install with `pip install youtube-transcript-api`")
-try:
- from langchain_community.document_loaders import YoutubeLoader
- from langchain_community.document_loaders.youtube import _parse_video_id
-except ImportError:
- raise ImportError("YouTube video requires extra dependencies. Install with `pip install pytube==15.0.0`") from None
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.loaders.base_loader import BaseLoader
-from embedchain.utils.misc import clean_string
-
-
-@register_deserializable
-class YoutubeVideoLoader(BaseLoader):
- def load_data(self, url):
- """Load data from a Youtube video."""
- video_id = _parse_video_id(url)
-
- languages = ["en"]
- try:
- # Fetching transcript data
- languages = [transcript.language_code for transcript in YouTubeTranscriptApi.list_transcripts(video_id)]
- transcript = YouTubeTranscriptApi.get_transcript(video_id, languages=languages)
- # convert transcript to json to avoid unicode symboles
- transcript = json.dumps(transcript, ensure_ascii=True)
- except Exception:
- logging.exception(f"Failed to fetch transcript for video {url}")
- transcript = "Unavailable"
-
- loader = YoutubeLoader.from_youtube_url(url, add_video_info=True, language=languages)
- doc = loader.load()
- output = []
- if not len(doc):
- raise ValueError(f"No data found for url: {url}")
- content = doc[0].page_content
- content = clean_string(content)
- metadata = doc[0].metadata
- metadata["url"] = url
- metadata["transcript"] = transcript
-
- output.append(
- {
- "content": content,
- "meta_data": metadata,
- }
- )
- doc_id = hashlib.sha256((content + url).encode()).hexdigest()
- return {
- "doc_id": doc_id,
- "data": output,
- }
diff --git a/embedchain/embedchain/memory/__init__.py b/embedchain/embedchain/memory/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/memory/base.py b/embedchain/embedchain/memory/base.py
deleted file mode 100644
index d6697625d..000000000
--- a/embedchain/embedchain/memory/base.py
+++ /dev/null
@@ -1,127 +0,0 @@
-import json
-import logging
-import uuid
-from typing import Any, Optional
-
-from embedchain.core.db.database import get_session
-from embedchain.core.db.models import ChatHistory as ChatHistoryModel
-from embedchain.memory.message import ChatMessage
-from embedchain.memory.utils import merge_metadata_dict
-
-logger = logging.getLogger(__name__)
-
-
-class ChatHistory:
- def __init__(self) -> None:
- self.db_session = get_session()
-
- def add(self, app_id, session_id, chat_message: ChatMessage) -> Optional[str]:
- memory_id = str(uuid.uuid4())
- metadata_dict = merge_metadata_dict(chat_message.human_message.metadata, chat_message.ai_message.metadata)
- if metadata_dict:
- metadata = self._serialize_json(metadata_dict)
- self.db_session.add(
- ChatHistoryModel(
- app_id=app_id,
- id=memory_id,
- session_id=session_id,
- question=chat_message.human_message.content,
- answer=chat_message.ai_message.content,
- metadata=metadata if metadata_dict else "{}",
- )
- )
- try:
- self.db_session.commit()
- except Exception as e:
- logger.error(f"Error adding chat memory to db: {e}")
- self.db_session.rollback()
- return None
-
- logger.info(f"Added chat memory to db with id: {memory_id}")
- return memory_id
-
- def delete(self, app_id: str, session_id: Optional[str] = None):
- """
- Delete all chat history for a given app_id and session_id.
- This is useful for deleting chat history for a given user.
-
- :param app_id: The app_id to delete chat history for
- :param session_id: The session_id to delete chat history for
-
- :return: None
- """
- params = {"app_id": app_id}
- if session_id:
- params["session_id"] = session_id
- self.db_session.query(ChatHistoryModel).filter_by(**params).delete()
- try:
- self.db_session.commit()
- except Exception as e:
- logger.error(f"Error deleting chat history: {e}")
- self.db_session.rollback()
-
- def get(
- self, app_id, session_id: str = "default", num_rounds=10, fetch_all: bool = False, display_format=False
- ) -> list[ChatMessage]:
- """
- Get the chat history for a given app_id.
-
- param: app_id - The app_id to get chat history
- param: session_id (optional) - The session_id to get chat history. Defaults to "default"
- param: num_rounds (optional) - The number of rounds to get chat history. Defaults to 10
- param: fetch_all (optional) - Whether to fetch all chat history or not. Defaults to False
- param: display_format (optional) - Whether to return the chat history in display format. Defaults to False
- """
- params = {"app_id": app_id}
- if not fetch_all:
- params["session_id"] = session_id
- results = (
- self.db_session.query(ChatHistoryModel).filter_by(**params).order_by(ChatHistoryModel.created_at.asc())
- )
- results = results.limit(num_rounds) if not fetch_all else results
- history = []
- for result in results:
- metadata = self._deserialize_json(metadata=result.meta_data or "{}")
- # Return list of dict if display_format is True
- if display_format:
- history.append(
- {
- "session_id": result.session_id,
- "human": result.question,
- "ai": result.answer,
- "metadata": result.meta_data,
- "timestamp": result.created_at,
- }
- )
- else:
- memory = ChatMessage()
- memory.add_user_message(result.question, metadata=metadata)
- memory.add_ai_message(result.answer, metadata=metadata)
- history.append(memory)
- return history
-
- def count(self, app_id: str, session_id: Optional[str] = None):
- """
- Count the number of chat messages for a given app_id and session_id.
-
- :param app_id: The app_id to count chat history for
- :param session_id: The session_id to count chat history for
-
- :return: The number of chat messages for a given app_id and session_id
- """
- # Rewrite the logic below with sqlalchemy
- params = {"app_id": app_id}
- if session_id:
- params["session_id"] = session_id
- return self.db_session.query(ChatHistoryModel).filter_by(**params).count()
-
- @staticmethod
- def _serialize_json(metadata: dict[str, Any]):
- return json.dumps(metadata)
-
- @staticmethod
- def _deserialize_json(metadata: str):
- return json.loads(metadata)
-
- def close_connection(self):
- self.connection.close()
diff --git a/embedchain/embedchain/memory/message.py b/embedchain/embedchain/memory/message.py
deleted file mode 100644
index 5211b0f6a..000000000
--- a/embedchain/embedchain/memory/message.py
+++ /dev/null
@@ -1,74 +0,0 @@
-import logging
-from typing import Any, Optional
-
-from embedchain.helpers.json_serializable import JSONSerializable
-
-logger = logging.getLogger(__name__)
-
-
-class BaseMessage(JSONSerializable):
- """
- The base abstract message class.
-
- Messages are the inputs and outputs of Models.
- """
-
- # The string content of the message.
- content: str
-
- # The created_by of the message. AI, Human, Bot etc.
- created_by: str
-
- # Any additional info.
- metadata: dict[str, Any]
-
- def __init__(self, content: str, created_by: str, metadata: Optional[dict[str, Any]] = None) -> None:
- super().__init__()
- self.content = content
- self.created_by = created_by
- self.metadata = metadata
-
- @property
- def type(self) -> str:
- """Type of the Message, used for serialization."""
-
- @classmethod
- def is_lc_serializable(cls) -> bool:
- """Return whether this class is serializable."""
- return True
-
- def __str__(self) -> str:
- return f"{self.created_by}: {self.content}"
-
-
-class ChatMessage(JSONSerializable):
- """
- The base abstract chat message class.
-
- Chat messages are the pair of (question, answer) conversation
- between human and model.
- """
-
- human_message: Optional[BaseMessage] = None
- ai_message: Optional[BaseMessage] = None
-
- def add_user_message(self, message: str, metadata: Optional[dict] = None):
- if self.human_message:
- logger.info(
- "Human message already exists in the chat message,\
- overwriting it with new message."
- )
-
- self.human_message = BaseMessage(content=message, created_by="human", metadata=metadata)
-
- def add_ai_message(self, message: str, metadata: Optional[dict] = None):
- if self.ai_message:
- logger.info(
- "AI message already exists in the chat message,\
- overwriting it with new message."
- )
-
- self.ai_message = BaseMessage(content=message, created_by="ai", metadata=metadata)
-
- def __str__(self) -> str:
- return f"{self.human_message}\n{self.ai_message}"
diff --git a/embedchain/embedchain/memory/utils.py b/embedchain/embedchain/memory/utils.py
deleted file mode 100644
index b849cffa6..000000000
--- a/embedchain/embedchain/memory/utils.py
+++ /dev/null
@@ -1,35 +0,0 @@
-from typing import Any, Optional
-
-
-def merge_metadata_dict(left: Optional[dict[str, Any]], right: Optional[dict[str, Any]]) -> Optional[dict[str, Any]]:
- """
- Merge the metadatas of two BaseMessage types.
-
- Args:
- left (dict[str, Any]): metadata of human message
- right (dict[str, Any]): metadata of AI message
-
- Returns:
- dict[str, Any]: combined metadata dict with dedup
- to be saved in db.
- """
- if not left and not right:
- return None
- elif not left:
- return right
- elif not right:
- return left
-
- merged = left.copy()
- for k, v in right.items():
- if k not in merged:
- merged[k] = v
- elif type(merged[k]) is not type(v):
- raise ValueError(f'additional_kwargs["{k}"] already exists in this message,' " but with a different type.")
- elif isinstance(merged[k], str):
- merged[k] += v
- elif isinstance(merged[k], dict):
- merged[k] = merge_metadata_dict(merged[k], v)
- else:
- raise ValueError(f"Additional kwargs key {k} already exists in this message.")
- return merged
diff --git a/embedchain/embedchain/migrations/env.py b/embedchain/embedchain/migrations/env.py
deleted file mode 100644
index 8fb3cd805..000000000
--- a/embedchain/embedchain/migrations/env.py
+++ /dev/null
@@ -1,68 +0,0 @@
-import os
-
-from alembic import context
-from sqlalchemy import engine_from_config, pool
-
-from embedchain.core.db.models import Base
-
-# this is the Alembic Config object, which provides
-# access to the values within the .ini file in use.
-config = context.config
-
-target_metadata = Base.metadata
-
-# other values from the config, defined by the needs of env.py,
-# can be acquired:
-# my_important_option = config.get_main_option("my_important_option")
-# ... etc.
-config.set_main_option("sqlalchemy.url", os.environ.get("EMBEDCHAIN_DB_URI"))
-
-
-def run_migrations_offline() -> None:
- """Run migrations in 'offline' mode.
-
- This configures the context with just a URL
- and not an Engine, though an Engine is acceptable
- here as well. By skipping the Engine creation
- we don't even need a DBAPI to be available.
-
- Calls to context.execute() here emit the given string to the
- script output.
-
- """
- url = config.get_main_option("sqlalchemy.url")
- context.configure(
- url=url,
- target_metadata=target_metadata,
- literal_binds=True,
- dialect_opts={"paramstyle": "named"},
- )
-
- with context.begin_transaction():
- context.run_migrations()
-
-
-def run_migrations_online() -> None:
- """Run migrations in 'online' mode.
-
- In this scenario we need to create an Engine
- and associate a connection with the context.
-
- """
- connectable = engine_from_config(
- config.get_section(config.config_ini_section, {}),
- prefix="sqlalchemy.",
- poolclass=pool.NullPool,
- )
-
- with connectable.connect() as connection:
- context.configure(connection=connection, target_metadata=target_metadata)
-
- with context.begin_transaction():
- context.run_migrations()
-
-
-if context.is_offline_mode():
- run_migrations_offline()
-else:
- run_migrations_online()
diff --git a/embedchain/embedchain/migrations/script.py.mako b/embedchain/embedchain/migrations/script.py.mako
deleted file mode 100644
index fbc4b07dc..000000000
--- a/embedchain/embedchain/migrations/script.py.mako
+++ /dev/null
@@ -1,26 +0,0 @@
-"""${message}
-
-Revision ID: ${up_revision}
-Revises: ${down_revision | comma,n}
-Create Date: ${create_date}
-
-"""
-from typing import Sequence, Union
-
-from alembic import op
-import sqlalchemy as sa
-${imports if imports else ""}
-
-# revision identifiers, used by Alembic.
-revision: str = ${repr(up_revision)}
-down_revision: Union[str, None] = ${repr(down_revision)}
-branch_labels: Union[str, Sequence[str], None] = ${repr(branch_labels)}
-depends_on: Union[str, Sequence[str], None] = ${repr(depends_on)}
-
-
-def upgrade() -> None:
- ${upgrades if upgrades else "pass"}
-
-
-def downgrade() -> None:
- ${downgrades if downgrades else "pass"}
diff --git a/embedchain/embedchain/migrations/versions/40a327b3debd_create_initial_migrations.py b/embedchain/embedchain/migrations/versions/40a327b3debd_create_initial_migrations.py
deleted file mode 100644
index 1facc88e3..000000000
--- a/embedchain/embedchain/migrations/versions/40a327b3debd_create_initial_migrations.py
+++ /dev/null
@@ -1,62 +0,0 @@
-"""Create initial migrations
-
-Revision ID: 40a327b3debd
-Revises:
-Create Date: 2024-02-18 15:29:19.409064
-
-"""
-
-from typing import Sequence, Union
-
-import sqlalchemy as sa
-from alembic import op
-
-# revision identifiers, used by Alembic.
-revision: str = "40a327b3debd"
-down_revision: Union[str, None] = None
-branch_labels: Union[str, Sequence[str], None] = None
-depends_on: Union[str, Sequence[str], None] = None
-
-
-def upgrade() -> None:
- # ### commands auto generated by Alembic - please adjust! ###
- op.create_table(
- "ec_chat_history",
- sa.Column("app_id", sa.String(), nullable=False),
- sa.Column("id", sa.String(), nullable=False),
- sa.Column("session_id", sa.String(), nullable=False),
- sa.Column("question", sa.Text(), nullable=True),
- sa.Column("answer", sa.Text(), nullable=True),
- sa.Column("metadata", sa.Text(), nullable=True),
- sa.Column("created_at", sa.TIMESTAMP(), nullable=True),
- sa.PrimaryKeyConstraint("app_id", "id", "session_id"),
- )
- op.create_index(op.f("ix_ec_chat_history_created_at"), "ec_chat_history", ["created_at"], unique=False)
- op.create_index(op.f("ix_ec_chat_history_session_id"), "ec_chat_history", ["session_id"], unique=False)
- op.create_table(
- "ec_data_sources",
- sa.Column("id", sa.String(), nullable=False),
- sa.Column("app_id", sa.Text(), nullable=True),
- sa.Column("hash", sa.Text(), nullable=True),
- sa.Column("type", sa.Text(), nullable=True),
- sa.Column("value", sa.Text(), nullable=True),
- sa.Column("metadata", sa.Text(), nullable=True),
- sa.Column("is_uploaded", sa.Integer(), nullable=True),
- sa.PrimaryKeyConstraint("id"),
- )
- op.create_index(op.f("ix_ec_data_sources_hash"), "ec_data_sources", ["hash"], unique=False)
- op.create_index(op.f("ix_ec_data_sources_app_id"), "ec_data_sources", ["app_id"], unique=False)
- op.create_index(op.f("ix_ec_data_sources_type"), "ec_data_sources", ["type"], unique=False)
- # ### end Alembic commands ###
-
-
-def downgrade() -> None:
- # ### commands auto generated by Alembic - please adjust! ###
- op.drop_index(op.f("ix_ec_data_sources_type"), table_name="ec_data_sources")
- op.drop_index(op.f("ix_ec_data_sources_app_id"), table_name="ec_data_sources")
- op.drop_index(op.f("ix_ec_data_sources_hash"), table_name="ec_data_sources")
- op.drop_table("ec_data_sources")
- op.drop_index(op.f("ix_ec_chat_history_session_id"), table_name="ec_chat_history")
- op.drop_index(op.f("ix_ec_chat_history_created_at"), table_name="ec_chat_history")
- op.drop_table("ec_chat_history")
- # ### end Alembic commands ###
diff --git a/embedchain/embedchain/models/__init__.py b/embedchain/embedchain/models/__init__.py
deleted file mode 100644
index 48887545b..000000000
--- a/embedchain/embedchain/models/__init__.py
+++ /dev/null
@@ -1,3 +0,0 @@
-from .embedding_functions import EmbeddingFunctions # noqa: F401
-from .providers import Providers # noqa: F401
-from .vector_dimensions import VectorDimensions # noqa: F401
diff --git a/embedchain/embedchain/models/data_type.py b/embedchain/embedchain/models/data_type.py
deleted file mode 100644
index 6370bf064..000000000
--- a/embedchain/embedchain/models/data_type.py
+++ /dev/null
@@ -1,85 +0,0 @@
-from enum import Enum
-
-
-class DirectDataType(Enum):
- """
- DirectDataType enum contains data types that contain raw data directly.
- """
-
- TEXT = "text"
-
-
-class IndirectDataType(Enum):
- """
- IndirectDataType enum contains data types that contain references to data stored elsewhere.
- """
-
- YOUTUBE_VIDEO = "youtube_video"
- PDF_FILE = "pdf_file"
- WEB_PAGE = "web_page"
- SITEMAP = "sitemap"
- XML = "xml"
- DOCX = "docx"
- DOCS_SITE = "docs_site"
- NOTION = "notion"
- CSV = "csv"
- MDX = "mdx"
- IMAGE = "image"
- UNSTRUCTURED = "unstructured"
- JSON = "json"
- OPENAPI = "openapi"
- GMAIL = "gmail"
- SUBSTACK = "substack"
- YOUTUBE_CHANNEL = "youtube_channel"
- DISCORD = "discord"
- CUSTOM = "custom"
- RSSFEED = "rss_feed"
- BEEHIIV = "beehiiv"
- GOOGLE_DRIVE = "google_drive"
- DIRECTORY = "directory"
- SLACK = "slack"
- DROPBOX = "dropbox"
- TEXT_FILE = "text_file"
- EXCEL_FILE = "excel_file"
- AUDIO = "audio"
-
-
-class SpecialDataType(Enum):
- """
- SpecialDataType enum contains data types that are neither direct nor indirect, or simply require special attention.
- """
-
- QNA_PAIR = "qna_pair"
-
-
-class DataType(Enum):
- TEXT = DirectDataType.TEXT.value
- YOUTUBE_VIDEO = IndirectDataType.YOUTUBE_VIDEO.value
- PDF_FILE = IndirectDataType.PDF_FILE.value
- WEB_PAGE = IndirectDataType.WEB_PAGE.value
- SITEMAP = IndirectDataType.SITEMAP.value
- XML = IndirectDataType.XML.value
- DOCX = IndirectDataType.DOCX.value
- DOCS_SITE = IndirectDataType.DOCS_SITE.value
- NOTION = IndirectDataType.NOTION.value
- CSV = IndirectDataType.CSV.value
- MDX = IndirectDataType.MDX.value
- QNA_PAIR = SpecialDataType.QNA_PAIR.value
- IMAGE = IndirectDataType.IMAGE.value
- UNSTRUCTURED = IndirectDataType.UNSTRUCTURED.value
- JSON = IndirectDataType.JSON.value
- OPENAPI = IndirectDataType.OPENAPI.value
- GMAIL = IndirectDataType.GMAIL.value
- SUBSTACK = IndirectDataType.SUBSTACK.value
- YOUTUBE_CHANNEL = IndirectDataType.YOUTUBE_CHANNEL.value
- DISCORD = IndirectDataType.DISCORD.value
- CUSTOM = IndirectDataType.CUSTOM.value
- RSSFEED = IndirectDataType.RSSFEED.value
- BEEHIIV = IndirectDataType.BEEHIIV.value
- GOOGLE_DRIVE = IndirectDataType.GOOGLE_DRIVE.value
- DIRECTORY = IndirectDataType.DIRECTORY.value
- SLACK = IndirectDataType.SLACK.value
- DROPBOX = IndirectDataType.DROPBOX.value
- TEXT_FILE = IndirectDataType.TEXT_FILE.value
- EXCEL_FILE = IndirectDataType.EXCEL_FILE.value
- AUDIO = IndirectDataType.AUDIO.value
diff --git a/embedchain/embedchain/models/embedding_functions.py b/embedchain/embedchain/models/embedding_functions.py
deleted file mode 100644
index 7171fadfa..000000000
--- a/embedchain/embedchain/models/embedding_functions.py
+++ /dev/null
@@ -1,10 +0,0 @@
-from enum import Enum
-
-
-class EmbeddingFunctions(Enum):
- OPENAI = "OPENAI"
- HUGGING_FACE = "HUGGING_FACE"
- VERTEX_AI = "VERTEX_AI"
- AWS_BEDROCK = "AWS_BEDROCK"
- GPT4ALL = "GPT4ALL"
- OLLAMA = "OLLAMA"
diff --git a/embedchain/embedchain/models/providers.py b/embedchain/embedchain/models/providers.py
deleted file mode 100644
index 62c93675b..000000000
--- a/embedchain/embedchain/models/providers.py
+++ /dev/null
@@ -1,10 +0,0 @@
-from enum import Enum
-
-
-class Providers(Enum):
- OPENAI = "OPENAI"
- ANTHROPHIC = "ANTHPROPIC"
- VERTEX_AI = "VERTEX_AI"
- GPT4ALL = "GPT4ALL"
- OLLAMA = "OLLAMA"
- AZURE_OPENAI = "AZURE_OPENAI"
diff --git a/embedchain/embedchain/models/vector_dimensions.py b/embedchain/embedchain/models/vector_dimensions.py
deleted file mode 100644
index bdea70756..000000000
--- a/embedchain/embedchain/models/vector_dimensions.py
+++ /dev/null
@@ -1,16 +0,0 @@
-from enum import Enum
-
-
-# vector length created by embedding fn
-class VectorDimensions(Enum):
- GPT4ALL = 384
- OPENAI = 1536
- VERTEX_AI = 768
- HUGGING_FACE = 384
- GOOGLE_AI = 768
- MISTRAL_AI = 1024
- NVIDIA_AI = 1024
- COHERE = 384
- OLLAMA = 384
- AMAZON_TITAN_V1 = 1536
- AMAZON_TITAN_V2 = 1024
diff --git a/embedchain/embedchain/pipeline.py b/embedchain/embedchain/pipeline.py
deleted file mode 100644
index 6f70bfb5d..000000000
--- a/embedchain/embedchain/pipeline.py
+++ /dev/null
@@ -1,9 +0,0 @@
-from embedchain.app import App
-
-
-class Pipeline(App):
- """
- This is deprecated. Use `App` instead.
- """
-
- pass
diff --git a/embedchain/embedchain/store/__init__.py b/embedchain/embedchain/store/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/store/assistants.py b/embedchain/embedchain/store/assistants.py
deleted file mode 100644
index b9ca151ab..000000000
--- a/embedchain/embedchain/store/assistants.py
+++ /dev/null
@@ -1,206 +0,0 @@
-import logging
-import os
-import re
-import tempfile
-import time
-import uuid
-from pathlib import Path
-from typing import cast
-
-from openai import OpenAI
-from openai.types.beta.threads import Message
-from openai.types.beta.threads.text_content_block import TextContentBlock
-
-from embedchain import Client, Pipeline
-from embedchain.config import AddConfig
-from embedchain.data_formatter import DataFormatter
-from embedchain.models.data_type import DataType
-from embedchain.telemetry.posthog import AnonymousTelemetry
-from embedchain.utils.misc import detect_datatype
-
-# Set up the user directory if it doesn't exist already
-Client.setup()
-
-
-class OpenAIAssistant:
- def __init__(
- self,
- name=None,
- instructions=None,
- tools=None,
- thread_id=None,
- model="gpt-4-1106-preview",
- data_sources=None,
- assistant_id=None,
- log_level=logging.INFO,
- collect_metrics=True,
- ):
- self.name = name or "OpenAI Assistant"
- self.instructions = instructions
- self.tools = tools or [{"type": "retrieval"}]
- self.model = model
- self.data_sources = data_sources or []
- self.log_level = log_level
- self._client = OpenAI()
- self._initialize_assistant(assistant_id)
- self.thread_id = thread_id or self._create_thread()
- self._telemetry_props = {"class": self.__class__.__name__}
- self.telemetry = AnonymousTelemetry(enabled=collect_metrics)
- self.telemetry.capture(event_name="init", properties=self._telemetry_props)
-
- def add(self, source, data_type=None):
- file_path = self._prepare_source_path(source, data_type)
- self._add_file_to_assistant(file_path)
-
- event_props = {
- **self._telemetry_props,
- "data_type": data_type or detect_datatype(source),
- }
- self.telemetry.capture(event_name="add", properties=event_props)
- logging.info("Data successfully added to the assistant.")
-
- def chat(self, message):
- self._send_message(message)
- self.telemetry.capture(event_name="chat", properties=self._telemetry_props)
- return self._get_latest_response()
-
- def delete_thread(self):
- self._client.beta.threads.delete(self.thread_id)
- self.thread_id = self._create_thread()
-
- # Internal methods
- def _initialize_assistant(self, assistant_id):
- file_ids = self._generate_file_ids(self.data_sources)
- self.assistant = (
- self._client.beta.assistants.retrieve(assistant_id)
- if assistant_id
- else self._client.beta.assistants.create(
- name=self.name, model=self.model, file_ids=file_ids, instructions=self.instructions, tools=self.tools
- )
- )
-
- def _create_thread(self):
- thread = self._client.beta.threads.create()
- return thread.id
-
- def _prepare_source_path(self, source, data_type=None):
- if Path(source).is_file():
- return source
- data_type = data_type or detect_datatype(source)
- formatter = DataFormatter(data_type=DataType(data_type), config=AddConfig())
- data = formatter.loader.load_data(source)["data"]
- return self._save_temp_data(data=data[0]["content"].encode(), source=source)
-
- def _add_file_to_assistant(self, file_path):
- file_obj = self._client.files.create(file=open(file_path, "rb"), purpose="assistants")
- self._client.beta.assistants.files.create(assistant_id=self.assistant.id, file_id=file_obj.id)
-
- def _generate_file_ids(self, data_sources):
- return [
- self._add_file_to_assistant(self._prepare_source_path(ds["source"], ds.get("data_type")))
- for ds in data_sources
- ]
-
- def _send_message(self, message):
- self._client.beta.threads.messages.create(thread_id=self.thread_id, role="user", content=message)
- self._wait_for_completion()
-
- def _wait_for_completion(self):
- run = self._client.beta.threads.runs.create(
- thread_id=self.thread_id,
- assistant_id=self.assistant.id,
- instructions=self.instructions,
- )
- run_id = run.id
- run_status = run.status
-
- while run_status in ["queued", "in_progress", "requires_action"]:
- time.sleep(0.1) # Sleep before making the next API call to avoid hitting rate limits
- run = self._client.beta.threads.runs.retrieve(thread_id=self.thread_id, run_id=run_id)
- run_status = run.status
- if run_status == "failed":
- raise ValueError(f"Thread run failed with the following error: {run.last_error}")
-
- def _get_latest_response(self):
- history = self._get_history()
- return self._format_message(history[0]) if history else None
-
- def _get_history(self):
- messages = self._client.beta.threads.messages.list(thread_id=self.thread_id, order="desc")
- return list(messages)
-
- @staticmethod
- def _format_message(thread_message):
- thread_message = cast(Message, thread_message)
- content = [c.text.value for c in thread_message.content if isinstance(c, TextContentBlock)]
- return " ".join(content)
-
- @staticmethod
- def _save_temp_data(data, source):
- special_chars_pattern = r'[\\/:*?"<>|&=% ]+'
- sanitized_source = re.sub(special_chars_pattern, "_", source)[:256]
- temp_dir = tempfile.mkdtemp()
- file_path = os.path.join(temp_dir, sanitized_source)
- with open(file_path, "wb") as file:
- file.write(data)
- return file_path
-
-
-class AIAssistant:
- def __init__(
- self,
- name=None,
- instructions=None,
- yaml_path=None,
- assistant_id=None,
- thread_id=None,
- data_sources=None,
- log_level=logging.INFO,
- collect_metrics=True,
- ):
- self.name = name or "AI Assistant"
- self.data_sources = data_sources or []
- self.log_level = log_level
- self.instructions = instructions
- self.assistant_id = assistant_id or str(uuid.uuid4())
- self.thread_id = thread_id or str(uuid.uuid4())
- self.pipeline = Pipeline.from_config(config_path=yaml_path) if yaml_path else Pipeline()
- self.pipeline.local_id = self.pipeline.config.id = self.thread_id
-
- if self.instructions:
- self.pipeline.system_prompt = self.instructions
-
- print(
- f"🎉 Created AI Assistant with name: {self.name}, assistant_id: {self.assistant_id}, thread_id: {self.thread_id}" # noqa: E501
- )
-
- # telemetry related properties
- self._telemetry_props = {"class": self.__class__.__name__}
- self.telemetry = AnonymousTelemetry(enabled=collect_metrics)
- self.telemetry.capture(event_name="init", properties=self._telemetry_props)
-
- if self.data_sources:
- for data_source in self.data_sources:
- metadata = {"assistant_id": self.assistant_id, "thread_id": "global_knowledge"}
- self.pipeline.add(data_source["source"], data_source.get("data_type"), metadata=metadata)
-
- def add(self, source, data_type=None):
- metadata = {"assistant_id": self.assistant_id, "thread_id": self.thread_id}
- self.pipeline.add(source, data_type=data_type, metadata=metadata)
- event_props = {
- **self._telemetry_props,
- "data_type": data_type or detect_datatype(source),
- }
- self.telemetry.capture(event_name="add", properties=event_props)
-
- def chat(self, query):
- where = {
- "$and": [
- {"assistant_id": {"$eq": self.assistant_id}},
- {"thread_id": {"$in": [self.thread_id, "global_knowledge"]}},
- ]
- }
- return self.pipeline.chat(query, where=where)
-
- def delete(self):
- self.pipeline.reset()
diff --git a/embedchain/embedchain/telemetry/__init__.py b/embedchain/embedchain/telemetry/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/telemetry/posthog.py b/embedchain/embedchain/telemetry/posthog.py
deleted file mode 100644
index 37c63ea61..000000000
--- a/embedchain/embedchain/telemetry/posthog.py
+++ /dev/null
@@ -1,60 +0,0 @@
-import json
-import logging
-import os
-import uuid
-
-from posthog import Posthog
-
-import embedchain
-from embedchain.constants import CONFIG_DIR, CONFIG_FILE
-
-
-class AnonymousTelemetry:
- def __init__(self, host="https://app.posthog.com", enabled=True):
- self.project_api_key = "phc_PHQDA5KwztijnSojsxJ2c1DuJd52QCzJzT2xnSGvjN2"
- self.host = host
- self.posthog = Posthog(project_api_key=self.project_api_key, host=self.host)
- self.user_id = self._get_user_id()
- self.enabled = enabled
-
- # Check if telemetry tracking is disabled via environment variable
- if "EC_TELEMETRY" in os.environ and os.environ["EC_TELEMETRY"].lower() not in [
- "1",
- "true",
- "yes",
- ]:
- self.enabled = False
-
- if not self.enabled:
- self.posthog.disabled = True
-
- # Silence posthog logging
- posthog_logger = logging.getLogger("posthog")
- posthog_logger.disabled = True
-
- @staticmethod
- def _get_user_id():
- os.makedirs(CONFIG_DIR, exist_ok=True)
- if os.path.exists(CONFIG_FILE):
- with open(CONFIG_FILE, "r") as f:
- data = json.load(f)
- if "user_id" in data:
- return data["user_id"]
-
- user_id = str(uuid.uuid4())
- with open(CONFIG_FILE, "w") as f:
- json.dump({"user_id": user_id}, f)
- return user_id
-
- def capture(self, event_name, properties=None):
- default_properties = {
- "version": embedchain.__version__,
- "language": "python",
- "pid": os.getpid(),
- }
- properties.update(default_properties)
-
- try:
- self.posthog.capture(self.user_id, event_name, properties)
- except Exception:
- logging.exception(f"Failed to send telemetry {event_name=}")
diff --git a/embedchain/embedchain/utils/__init__.py b/embedchain/embedchain/utils/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/utils/cli.py b/embedchain/embedchain/utils/cli.py
deleted file mode 100644
index 13128df5a..000000000
--- a/embedchain/embedchain/utils/cli.py
+++ /dev/null
@@ -1,320 +0,0 @@
-import os
-import re
-import shutil
-import subprocess
-
-import pkg_resources
-from rich.console import Console
-
-console = Console()
-
-
-def get_pkg_path_from_name(template: str):
- try:
- # Determine the installation location of the embedchain package
- package_path = pkg_resources.resource_filename("embedchain", "")
- except ImportError:
- console.print("❌ [bold red]Failed to locate the 'embedchain' package. Is it installed?[/bold red]")
- return
-
- # Construct the source path from the embedchain package
- src_path = os.path.join(package_path, "deployment", template)
-
- if not os.path.exists(src_path):
- console.print(f"❌ [bold red]Template '{template}' not found.[/bold red]")
- return
-
- return src_path
-
-
-def setup_fly_io_app(extra_args):
- fly_launch_command = ["fly", "launch", "--region", "sjc", "--no-deploy"] + list(extra_args)
- try:
- console.print(f"🚀 [bold cyan]Running: {' '.join(fly_launch_command)}[/bold cyan]")
- shutil.move(".env.example", ".env")
- subprocess.run(fly_launch_command, check=True)
- console.print("✅ [bold green]'fly launch' executed successfully.[/bold green]")
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except FileNotFoundError:
- console.print(
- "❌ [bold red]'fly' command not found. Please ensure Fly CLI is installed and in your PATH.[/bold red]"
- )
-
-
-def setup_modal_com_app(extra_args):
- modal_setup_file = os.path.join(os.path.expanduser("~"), ".modal.toml")
- if os.path.exists(modal_setup_file):
- console.print(
- """✅ [bold green]Modal setup already done. You can now install the dependencies by doing \n
- `pip install -r requirements.txt`[/bold green]"""
- )
- else:
- modal_setup_cmd = ["modal", "setup"] + list(extra_args)
- console.print(f"🚀 [bold cyan]Running: {' '.join(modal_setup_cmd)}[/bold cyan]")
- subprocess.run(modal_setup_cmd, check=True)
- shutil.move(".env.example", ".env")
- console.print(
- """Great! Now you can install the dependencies by doing: \n
- `pip install -r requirements.txt`\n
- \n
- To run your app locally:\n
- `ec dev`
- """
- )
-
-
-def setup_render_com_app():
- render_setup_file = os.path.join(os.path.expanduser("~"), ".render/config.yaml")
- if os.path.exists(render_setup_file):
- console.print(
- """✅ [bold green]Render setup already done. You can now install the dependencies by doing \n
- `pip install -r requirements.txt`[/bold green]"""
- )
- else:
- render_setup_cmd = ["render", "config", "init"]
- console.print(f"🚀 [bold cyan]Running: {' '.join(render_setup_cmd)}[/bold cyan]")
- subprocess.run(render_setup_cmd, check=True)
- shutil.move(".env.example", ".env")
- console.print(
- """Great! Now you can install the dependencies by doing: \n
- `pip install -r requirements.txt`\n
- \n
- To run your app locally:\n
- `ec dev`
- """
- )
-
-
-def setup_streamlit_io_app():
- # nothing needs to be done here
- console.print("Great! Now you can install the dependencies by doing `pip install -r requirements.txt`")
-
-
-def setup_gradio_app():
- # nothing needs to be done here
- console.print("Great! Now you can install the dependencies by doing `pip install -r requirements.txt`")
-
-
-def setup_hf_app():
- subprocess.run(["pip", "install", "huggingface_hub[cli]"], check=True)
- hf_setup_file = os.path.join(os.path.expanduser("~"), ".cache/huggingface/token")
- if os.path.exists(hf_setup_file):
- console.print(
- """✅ [bold green]HuggingFace setup already done. You can now install the dependencies by doing \n
- `pip install -r requirements.txt`[/bold green]"""
- )
- else:
- console.print(
- """🚀 [cyan]Running: huggingface-cli login \n
- Please provide a [bold]WRITE[/bold] token so that we can directly deploy\n
- your apps from the terminal.[/cyan]
- """
- )
- subprocess.run(["huggingface-cli", "login"], check=True)
- console.print("Great! Now you can install the dependencies by doing `pip install -r requirements.txt`")
-
-
-def run_dev_fly_io(debug, host, port):
- uvicorn_command = ["uvicorn", "app:app"]
-
- if debug:
- uvicorn_command.append("--reload")
-
- uvicorn_command.extend(["--host", host, "--port", str(port)])
-
- try:
- console.print(f"🚀 [bold cyan]Running FastAPI app with command: {' '.join(uvicorn_command)}[/bold cyan]")
- subprocess.run(uvicorn_command, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]FastAPI server stopped[/bold yellow]")
-
-
-def run_dev_modal_com():
- modal_run_cmd = ["modal", "serve", "app"]
- try:
- console.print(f"🚀 [bold cyan]Running FastAPI app with command: {' '.join(modal_run_cmd)}[/bold cyan]")
- subprocess.run(modal_run_cmd, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]FastAPI server stopped[/bold yellow]")
-
-
-def run_dev_streamlit_io():
- streamlit_run_cmd = ["streamlit", "run", "app.py"]
- try:
- console.print(f"🚀 [bold cyan]Running Streamlit app with command: {' '.join(streamlit_run_cmd)}[/bold cyan]")
- subprocess.run(streamlit_run_cmd, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]Streamlit server stopped[/bold yellow]")
-
-
-def run_dev_render_com(debug, host, port):
- uvicorn_command = ["uvicorn", "app:app"]
-
- if debug:
- uvicorn_command.append("--reload")
-
- uvicorn_command.extend(["--host", host, "--port", str(port)])
-
- try:
- console.print(f"🚀 [bold cyan]Running FastAPI app with command: {' '.join(uvicorn_command)}[/bold cyan]")
- subprocess.run(uvicorn_command, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]FastAPI server stopped[/bold yellow]")
-
-
-def run_dev_gradio():
- gradio_run_cmd = ["gradio", "app.py"]
- try:
- console.print(f"🚀 [bold cyan]Running Gradio app with command: {' '.join(gradio_run_cmd)}[/bold cyan]")
- subprocess.run(gradio_run_cmd, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except KeyboardInterrupt:
- console.print("\n🛑 [bold yellow]Gradio server stopped[/bold yellow]")
-
-
-def read_env_file(env_file_path):
- """
- Reads an environment file and returns a dictionary of key-value pairs.
-
- Args:
- env_file_path (str): The path to the .env file.
-
- Returns:
- dict: Dictionary of environment variables.
- """
- env_vars = {}
- pattern = re.compile(r"(\w+)=(.*)") # compile regular expression for better performance
- with open(env_file_path, "r") as file:
- lines = file.readlines() # readlines is faster as it reads all at once
- for line in lines:
- line = line.strip()
- # Ignore comments and empty lines
- if line and not line.startswith("#"):
- # Assume each line is in the format KEY=VALUE
- key_value_match = pattern.match(line)
- if key_value_match:
- key, value = key_value_match.groups()
- env_vars[key] = value
- return env_vars
-
-
-def deploy_fly():
- app_name = ""
- with open("fly.toml", "r") as file:
- for line in file:
- if line.strip().startswith("app ="):
- app_name = line.split("=")[1].strip().strip('"')
-
- if not app_name:
- console.print("❌ [bold red]App name not found in fly.toml[/bold red]")
- return
-
- env_vars = read_env_file(".env")
- secrets_command = ["flyctl", "secrets", "set", "-a", app_name] + [f"{k}={v}" for k, v in env_vars.items()]
-
- deploy_command = ["fly", "deploy"]
- try:
- # Set secrets
- console.print(f"🔐 [bold cyan]Setting secrets for {app_name}[/bold cyan]")
- subprocess.run(secrets_command, check=True)
-
- # Deploy application
- console.print(f"🚀 [bold cyan]Running: {' '.join(deploy_command)}[/bold cyan]")
- subprocess.run(deploy_command, check=True)
- console.print("✅ [bold green]'fly deploy' executed successfully.[/bold green]")
-
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except FileNotFoundError:
- console.print(
- "❌ [bold red]'fly' command not found. Please ensure Fly CLI is installed and in your PATH.[/bold red]"
- )
-
-
-def deploy_modal():
- modal_deploy_cmd = ["modal", "deploy", "app"]
- try:
- console.print(f"🚀 [bold cyan]Running: {' '.join(modal_deploy_cmd)}[/bold cyan]")
- subprocess.run(modal_deploy_cmd, check=True)
- console.print("✅ [bold green]'modal deploy' executed successfully.[/bold green]")
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except FileNotFoundError:
- console.print(
- "❌ [bold red]'modal' command not found. Please ensure Modal CLI is installed and in your PATH.[/bold red]"
- )
-
-
-def deploy_streamlit():
- streamlit_deploy_cmd = ["streamlit", "run", "app.py"]
- try:
- console.print(f"🚀 [bold cyan]Running: {' '.join(streamlit_deploy_cmd)}[/bold cyan]")
- console.print(
- """\n\n✅ [bold yellow]To deploy a streamlit app, you can directly it from the UI.\n
- Click on the 'Deploy' button on the top right corner of the app.\n
- For more information, please refer to https://docs.embedchain.ai/deployment/streamlit_io
- [/bold yellow]
- \n\n"""
- )
- subprocess.run(streamlit_deploy_cmd, check=True)
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except FileNotFoundError:
- console.print(
- """❌ [bold red]'streamlit' command not found.\n
- Please ensure Streamlit CLI is installed and in your PATH.[/bold red]"""
- )
-
-
-def deploy_render():
- render_deploy_cmd = ["render", "blueprint", "launch"]
-
- try:
- console.print(f"🚀 [bold cyan]Running: {' '.join(render_deploy_cmd)}[/bold cyan]")
- subprocess.run(render_deploy_cmd, check=True)
- console.print("✅ [bold green]'render blueprint launch' executed successfully.[/bold green]")
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except FileNotFoundError:
- console.print(
- "❌ [bold red]'render' command not found. Please ensure Render CLI is installed and in your PATH.[/bold red]" # noqa:E501
- )
-
-
-def deploy_gradio_app():
- gradio_deploy_cmd = ["gradio", "deploy"]
-
- try:
- console.print(f"🚀 [bold cyan]Running: {' '.join(gradio_deploy_cmd)}[/bold cyan]")
- subprocess.run(gradio_deploy_cmd, check=True)
- console.print("✅ [bold green]'gradio deploy' executed successfully.[/bold green]")
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
- except FileNotFoundError:
- console.print(
- "❌ [bold red]'gradio' command not found. Please ensure Gradio CLI is installed and in your PATH.[/bold red]" # noqa:E501
- )
-
-
-def deploy_hf_spaces(ec_app_name):
- if not ec_app_name:
- console.print("❌ [bold red]'name' not found in embedchain.json[/bold red]")
- return
- hf_spaces_deploy_cmd = ["huggingface-cli", "upload", ec_app_name, ".", ".", "--repo-type=space"]
-
- try:
- console.print(f"🚀 [bold cyan]Running: {' '.join(hf_spaces_deploy_cmd)}[/bold cyan]")
- subprocess.run(hf_spaces_deploy_cmd, check=True)
- console.print("✅ [bold green]'huggingface-cli upload' executed successfully.[/bold green]")
- except subprocess.CalledProcessError as e:
- console.print(f"❌ [bold red]An error occurred: {e}[/bold red]")
diff --git a/embedchain/embedchain/utils/evaluation.py b/embedchain/embedchain/utils/evaluation.py
deleted file mode 100644
index 62eaaeb70..000000000
--- a/embedchain/embedchain/utils/evaluation.py
+++ /dev/null
@@ -1,17 +0,0 @@
-from enum import Enum
-from typing import Optional
-
-from pydantic import BaseModel
-
-
-class EvalMetric(Enum):
- CONTEXT_RELEVANCY = "context_relevancy"
- ANSWER_RELEVANCY = "answer_relevancy"
- GROUNDEDNESS = "groundedness"
-
-
-class EvalData(BaseModel):
- question: str
- contexts: list[str]
- answer: str
- ground_truth: Optional[str] = None # Not used as of now
diff --git a/embedchain/embedchain/utils/misc.py b/embedchain/embedchain/utils/misc.py
deleted file mode 100644
index 7c5468ec9..000000000
--- a/embedchain/embedchain/utils/misc.py
+++ /dev/null
@@ -1,546 +0,0 @@
-import datetime
-import itertools
-import json
-import logging
-import os
-import re
-import string
-from typing import Any
-
-from schema import Optional, Or, Schema
-from tqdm import tqdm
-
-from embedchain.models.data_type import DataType
-
-logger = logging.getLogger(__name__)
-
-
-def parse_content(content, type):
- implemented = ["html.parser", "lxml", "lxml-xml", "xml", "html5lib"]
- if type not in implemented:
- raise ValueError(f"Parser type {type} not implemented. Please choose one of {implemented}")
-
- from bs4 import BeautifulSoup
-
- soup = BeautifulSoup(content, type)
- original_size = len(str(soup.get_text()))
-
- tags_to_exclude = [
- "nav",
- "aside",
- "form",
- "header",
- "noscript",
- "svg",
- "canvas",
- "footer",
- "script",
- "style",
- ]
- for tag in soup(tags_to_exclude):
- tag.decompose()
-
- ids_to_exclude = ["sidebar", "main-navigation", "menu-main-menu"]
- for id in ids_to_exclude:
- tags = soup.find_all(id=id)
- for tag in tags:
- tag.decompose()
-
- classes_to_exclude = [
- "elementor-location-header",
- "navbar-header",
- "nav",
- "header-sidebar-wrapper",
- "blog-sidebar-wrapper",
- "related-posts",
- ]
- for class_name in classes_to_exclude:
- tags = soup.find_all(class_=class_name)
- for tag in tags:
- tag.decompose()
-
- content = soup.get_text()
- content = clean_string(content)
-
- cleaned_size = len(content)
- if original_size != 0:
- logger.info(
- f"Cleaned page size: {cleaned_size} characters, down from {original_size} (shrunk: {original_size-cleaned_size} chars, {round((1-(cleaned_size/original_size)) * 100, 2)}%)" # noqa:E501
- )
-
- return content
-
-
-def clean_string(text):
- """
- This function takes in a string and performs a series of text cleaning operations.
-
- Args:
- text (str): The text to be cleaned. This is expected to be a string.
-
- Returns:
- cleaned_text (str): The cleaned text after all the cleaning operations
- have been performed.
- """
- # Stripping and reducing multiple spaces to single:
- cleaned_text = re.sub(r"\s+", " ", text.strip())
-
- # Removing backslashes:
- cleaned_text = cleaned_text.replace("\\", "")
-
- # Replacing hash characters:
- cleaned_text = cleaned_text.replace("#", " ")
-
- # Eliminating consecutive non-alphanumeric characters:
- # This regex identifies consecutive non-alphanumeric characters (i.e., not
- # a word character [a-zA-Z0-9_] and not a whitespace) in the string
- # and replaces each group of such characters with a single occurrence of
- # that character.
- # For example, "!!! hello !!!" would become "! hello !".
- cleaned_text = re.sub(r"([^\w\s])\1*", r"\1", cleaned_text)
-
- return cleaned_text
-
-
-def is_readable(s):
- """
- Heuristic to determine if a string is "readable" (mostly contains printable characters and forms meaningful words)
-
- :param s: string
- :return: True if the string is more than 95% printable.
- """
- len_s = len(s)
- if len_s == 0:
- return False
- printable_chars = set(string.printable)
- printable_ratio = sum(c in printable_chars for c in s) / len_s
- return printable_ratio > 0.95 # 95% of characters are printable
-
-
-def use_pysqlite3():
- """
- Swap std-lib sqlite3 with pysqlite3.
- """
- import platform
- import sqlite3
-
- if platform.system() == "Linux" and sqlite3.sqlite_version_info < (3, 35, 0):
- try:
- # According to the Chroma team, this patch only works on Linux
- import datetime
- import subprocess
- import sys
-
- subprocess.check_call(
- [sys.executable, "-m", "pip", "install", "pysqlite3-binary", "--quiet", "--disable-pip-version-check"]
- )
-
- __import__("pysqlite3")
- sys.modules["sqlite3"] = sys.modules.pop("pysqlite3")
-
- # Let the user know what happened.
- current_time = datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S,%f")[:-3]
- print(
- f"{current_time} [embedchain] [INFO]",
- "Swapped std-lib sqlite3 with pysqlite3 for ChromaDb compatibility.",
- f"Your original version was {sqlite3.sqlite_version}.",
- )
- except Exception as e:
- # Escape all exceptions
- current_time = datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S,%f")[:-3]
- print(
- f"{current_time} [embedchain] [ERROR]",
- "Failed to swap std-lib sqlite3 with pysqlite3 for ChromaDb compatibility.",
- "Error:",
- e,
- )
-
-
-def format_source(source: str, limit: int = 20) -> str:
- """
- Format a string to only take the first x and last x letters.
- This makes it easier to display a URL, keeping familiarity while ensuring a consistent length.
- If the string is too short, it is not sliced.
- """
- if len(source) > 2 * limit:
- return source[:limit] + "..." + source[-limit:]
- return source
-
-
-def detect_datatype(source: Any) -> DataType:
- """
- Automatically detect the datatype of the given source.
-
- :param source: the source to base the detection on
- :return: data_type string
- """
- from urllib.parse import urlparse
-
- import requests
- import yaml
-
- def is_openapi_yaml(yaml_content):
- # currently the following two fields are required in openapi spec yaml config
- return "openapi" in yaml_content and "info" in yaml_content
-
- def is_google_drive_folder(url):
- # checks if url is a Google Drive folder url against a regex
- regex = r"^drive\.google\.com\/drive\/(?:u\/\d+\/)folders\/([a-zA-Z0-9_-]+)$"
- return re.match(regex, url)
-
- try:
- if not isinstance(source, str):
- raise ValueError("Source is not a string and thus cannot be a URL.")
- url = urlparse(source)
- # Check if both scheme and netloc are present. Local file system URIs are acceptable too.
- if not all([url.scheme, url.netloc]) and url.scheme != "file":
- raise ValueError("Not a valid URL.")
- except ValueError:
- url = False
-
- formatted_source = format_source(str(source), 30)
-
- if url:
- YOUTUBE_ALLOWED_NETLOCKS = {
- "www.youtube.com",
- "m.youtube.com",
- "youtu.be",
- "youtube.com",
- "vid.plus",
- "www.youtube-nocookie.com",
- }
-
- if url.netloc in YOUTUBE_ALLOWED_NETLOCKS:
- logger.debug(f"Source of `{formatted_source}` detected as `youtube_video`.")
- return DataType.YOUTUBE_VIDEO
-
- if url.netloc in {"notion.so", "notion.site"}:
- logger.debug(f"Source of `{formatted_source}` detected as `notion`.")
- return DataType.NOTION
-
- if url.path.endswith(".pdf"):
- logger.debug(f"Source of `{formatted_source}` detected as `pdf_file`.")
- return DataType.PDF_FILE
-
- if url.path.endswith(".xml"):
- logger.debug(f"Source of `{formatted_source}` detected as `sitemap`.")
- return DataType.SITEMAP
-
- if url.path.endswith(".csv"):
- logger.debug(f"Source of `{formatted_source}` detected as `csv`.")
- return DataType.CSV
-
- if url.path.endswith(".mdx") or url.path.endswith(".md"):
- logger.debug(f"Source of `{formatted_source}` detected as `mdx`.")
- return DataType.MDX
-
- if url.path.endswith(".docx"):
- logger.debug(f"Source of `{formatted_source}` detected as `docx`.")
- return DataType.DOCX
-
- if url.path.endswith(
- (".mp3", ".mp4", ".mp2", ".aac", ".wav", ".flac", ".pcm", ".m4a", ".ogg", ".opus", ".webm")
- ):
- logger.debug(f"Source of `{formatted_source}` detected as `audio`.")
- return DataType.AUDIO
-
- if url.path.endswith(".yaml"):
- try:
- response = requests.get(source)
- response.raise_for_status()
- try:
- yaml_content = yaml.safe_load(response.text)
- except yaml.YAMLError as exc:
- logger.error(f"Error parsing YAML: {exc}")
- raise TypeError(f"Not a valid data type. Error loading YAML: {exc}")
-
- if is_openapi_yaml(yaml_content):
- logger.debug(f"Source of `{formatted_source}` detected as `openapi`.")
- return DataType.OPENAPI
- else:
- logger.error(
- f"Source of `{formatted_source}` does not contain all the required \
- fields of OpenAPI yaml. Check 'https://spec.openapis.org/oas/v3.1.0'"
- )
- raise TypeError(
- "Not a valid data type. Check 'https://spec.openapis.org/oas/v3.1.0', \
- make sure you have all the required fields in YAML config data"
- )
- except requests.exceptions.RequestException as e:
- logger.error(f"Error fetching URL {formatted_source}: {e}")
-
- if url.path.endswith(".json"):
- logger.debug(f"Source of `{formatted_source}` detected as `json_file`.")
- return DataType.JSON
-
- if "docs" in url.netloc or ("docs" in url.path and url.scheme != "file"):
- # `docs_site` detection via path is not accepted for local filesystem URIs,
- # because that would mean all paths that contain `docs` are now doc sites, which is too aggressive.
- logger.debug(f"Source of `{formatted_source}` detected as `docs_site`.")
- return DataType.DOCS_SITE
-
- if "github.com" in url.netloc:
- logger.debug(f"Source of `{formatted_source}` detected as `github`.")
- return DataType.GITHUB
-
- if is_google_drive_folder(url.netloc + url.path):
- logger.debug(f"Source of `{formatted_source}` detected as `google drive folder`.")
- return DataType.GOOGLE_DRIVE_FOLDER
-
- # If none of the above conditions are met, it's a general web page
- logger.debug(f"Source of `{formatted_source}` detected as `web_page`.")
- return DataType.WEB_PAGE
-
- elif not isinstance(source, str):
- # For datatypes where source is not a string.
-
- if isinstance(source, tuple) and len(source) == 2 and isinstance(source[0], str) and isinstance(source[1], str):
- logger.debug(f"Source of `{formatted_source}` detected as `qna_pair`.")
- return DataType.QNA_PAIR
-
- # Raise an error if it isn't a string and also not a valid non-string type (one of the previous).
- # We could stringify it, but it is better to raise an error and let the user decide how they want to do that.
- raise TypeError(
- "Source is not a string and a valid non-string type could not be detected. If you want to embed it, please stringify it, for instance by using `str(source)` or `(', ').join(source)`." # noqa: E501
- )
-
- elif os.path.isfile(source):
- # For datatypes that support conventional file references.
- # Note: checking for string is not necessary anymore.
-
- if source.endswith(".docx"):
- logger.debug(f"Source of `{formatted_source}` detected as `docx`.")
- return DataType.DOCX
-
- if source.endswith(".csv"):
- logger.debug(f"Source of `{formatted_source}` detected as `csv`.")
- return DataType.CSV
-
- if source.endswith(".xml"):
- logger.debug(f"Source of `{formatted_source}` detected as `xml`.")
- return DataType.XML
-
- if source.endswith(".mdx") or source.endswith(".md"):
- logger.debug(f"Source of `{formatted_source}` detected as `mdx`.")
- return DataType.MDX
-
- if source.endswith(".txt"):
- logger.debug(f"Source of `{formatted_source}` detected as `text`.")
- return DataType.TEXT_FILE
-
- if source.endswith(".pdf"):
- logger.debug(f"Source of `{formatted_source}` detected as `pdf_file`.")
- return DataType.PDF_FILE
-
- if source.endswith(".yaml"):
- with open(source, "r") as file:
- yaml_content = yaml.safe_load(file)
- if is_openapi_yaml(yaml_content):
- logger.debug(f"Source of `{formatted_source}` detected as `openapi`.")
- return DataType.OPENAPI
- else:
- logger.error(
- f"Source of `{formatted_source}` does not contain all the required \
- fields of OpenAPI yaml. Check 'https://spec.openapis.org/oas/v3.1.0'"
- )
- raise ValueError(
- "Invalid YAML data. Check 'https://spec.openapis.org/oas/v3.1.0', \
- make sure to add all the required params"
- )
-
- if source.endswith(".json"):
- logger.debug(f"Source of `{formatted_source}` detected as `json`.")
- return DataType.JSON
-
- if os.path.exists(source) and is_readable(open(source).read()):
- logger.debug(f"Source of `{formatted_source}` detected as `text_file`.")
- return DataType.TEXT_FILE
-
- # If the source is a valid file, that's not detectable as a type, an error is raised.
- # It does not fall back to text.
- raise ValueError(
- "Source points to a valid file, but based on the filename, no `data_type` can be detected. Please be aware, that not all data_types allow conventional file references, some require the use of the `file URI scheme`. Please refer to the embedchain documentation (https://docs.embedchain.ai/advanced/data_types#remote-data-types)." # noqa: E501
- )
-
- else:
- # Source is not a URL.
-
- # TODO: check if source is gmail query
-
- # check if the source is valid json string
- if is_valid_json_string(source):
- logger.debug(f"Source of `{formatted_source}` detected as `json`.")
- return DataType.JSON
-
- # Use text as final fallback.
- logger.debug(f"Source of `{formatted_source}` detected as `text`.")
- return DataType.TEXT
-
-
-# check if the source is valid json string
-def is_valid_json_string(source: str):
- try:
- _ = json.loads(source)
- return True
- except json.JSONDecodeError:
- return False
-
-
-def validate_config(config_data):
- schema = Schema(
- {
- Optional("app"): {
- Optional("config"): {
- Optional("id"): str,
- Optional("name"): str,
- Optional("log_level"): Or("DEBUG", "INFO", "WARNING", "ERROR", "CRITICAL"),
- Optional("collect_metrics"): bool,
- Optional("collection_name"): str,
- }
- },
- Optional("llm"): {
- Optional("provider"): Or(
- "openai",
- "azure_openai",
- "anthropic",
- "huggingface",
- "cohere",
- "together",
- "gpt4all",
- "ollama",
- "jina",
- "llama2",
- "vertexai",
- "google",
- "aws_bedrock",
- "mistralai",
- "clarifai",
- "vllm",
- "groq",
- "nvidia",
- ),
- Optional("config"): {
- Optional("model"): str,
- Optional("model_name"): str,
- Optional("number_documents"): int,
- Optional("temperature"): float,
- Optional("max_tokens"): int,
- Optional("top_p"): Or(float, int),
- Optional("stream"): bool,
- Optional("online"): bool,
- Optional("token_usage"): bool,
- Optional("template"): str,
- Optional("prompt"): str,
- Optional("system_prompt"): str,
- Optional("deployment_name"): str,
- Optional("where"): dict,
- Optional("query_type"): str,
- Optional("api_key"): str,
- Optional("base_url"): str,
- Optional("endpoint"): str,
- Optional("model_kwargs"): dict,
- Optional("local"): bool,
- Optional("base_url"): str,
- Optional("default_headers"): dict,
- Optional("api_version"): Or(str, datetime.date),
- Optional("http_client_proxies"): Or(str, dict),
- Optional("http_async_client_proxies"): Or(str, dict),
- },
- },
- Optional("vectordb"): {
- Optional("provider"): Or(
- "chroma", "elasticsearch", "opensearch", "lancedb", "pinecone", "qdrant", "weaviate", "zilliz"
- ),
- Optional("config"): object, # TODO: add particular config schema for each provider
- },
- Optional("embedder"): {
- Optional("provider"): Or(
- "openai",
- "gpt4all",
- "huggingface",
- "vertexai",
- "azure_openai",
- "google",
- "mistralai",
- "clarifai",
- "nvidia",
- "ollama",
- "cohere",
- "aws_bedrock",
- ),
- Optional("config"): {
- Optional("model"): Optional(str),
- Optional("deployment_name"): Optional(str),
- Optional("api_key"): str,
- Optional("api_base"): str,
- Optional("title"): str,
- Optional("task_type"): str,
- Optional("vector_dimension"): int,
- Optional("base_url"): str,
- Optional("endpoint"): str,
- Optional("model_kwargs"): dict,
- Optional("http_client_proxies"): Or(str, dict),
- Optional("http_async_client_proxies"): Or(str, dict),
- },
- },
- Optional("embedding_model"): {
- Optional("provider"): Or(
- "openai",
- "gpt4all",
- "huggingface",
- "vertexai",
- "azure_openai",
- "google",
- "mistralai",
- "clarifai",
- "nvidia",
- "ollama",
- "aws_bedrock",
- ),
- Optional("config"): {
- Optional("model"): str,
- Optional("deployment_name"): str,
- Optional("api_key"): str,
- Optional("title"): str,
- Optional("task_type"): str,
- Optional("vector_dimension"): int,
- Optional("base_url"): str,
- },
- },
- Optional("chunker"): {
- Optional("chunk_size"): int,
- Optional("chunk_overlap"): int,
- Optional("length_function"): str,
- Optional("min_chunk_size"): int,
- },
- Optional("cache"): {
- Optional("similarity_evaluation"): {
- Optional("strategy"): Or("distance", "exact"),
- Optional("max_distance"): float,
- Optional("positive"): bool,
- },
- Optional("config"): {
- Optional("similarity_threshold"): float,
- Optional("auto_flush"): int,
- },
- },
- Optional("memory"): {
- Optional("top_k"): int,
- },
- }
- )
-
- return schema.validate(config_data)
-
-
-def chunks(iterable, batch_size=100, desc="Processing chunks"):
- """A helper function to break an iterable into chunks of size batch_size."""
- it = iter(iterable)
- total_size = len(iterable)
-
- with tqdm(total=total_size, desc=desc, unit="batch") as pbar:
- chunk = tuple(itertools.islice(it, batch_size))
- while chunk:
- yield chunk
- pbar.update(len(chunk))
- chunk = tuple(itertools.islice(it, batch_size))
diff --git a/embedchain/embedchain/vectordb/__init__.py b/embedchain/embedchain/vectordb/__init__.py
deleted file mode 100644
index e69de29bb..000000000
diff --git a/embedchain/embedchain/vectordb/base.py b/embedchain/embedchain/vectordb/base.py
deleted file mode 100644
index e65cde01a..000000000
--- a/embedchain/embedchain/vectordb/base.py
+++ /dev/null
@@ -1,82 +0,0 @@
-from embedchain.config.vector_db.base import BaseVectorDbConfig
-from embedchain.embedder.base import BaseEmbedder
-from embedchain.helpers.json_serializable import JSONSerializable
-
-
-class BaseVectorDB(JSONSerializable):
- """Base class for vector database."""
-
- def __init__(self, config: BaseVectorDbConfig):
- """Initialize the database. Save the config and client as an attribute.
-
- :param config: Database configuration class instance.
- :type config: BaseVectorDbConfig
- """
- self.client = self._get_or_create_db()
- self.config: BaseVectorDbConfig = config
-
- def _initialize(self):
- """
- This method is needed because `embedder` attribute needs to be set externally before it can be initialized.
-
- So it's can't be done in __init__ in one step.
- """
- raise NotImplementedError
-
- def _get_or_create_db(self):
- """Get or create the database."""
- raise NotImplementedError
-
- def _get_or_create_collection(self):
- """Get or create a named collection."""
- raise NotImplementedError
-
- def _set_embedder(self, embedder: BaseEmbedder):
- """
- The database needs to access the embedder sometimes, with this method you can persistently set it.
-
- :param embedder: Embedder to be set as the embedder for this database.
- :type embedder: BaseEmbedder
- """
- self.embedder = embedder
-
- def get(self):
- """Get database embeddings by id."""
- raise NotImplementedError
-
- def add(self):
- """Add to database"""
- raise NotImplementedError
-
- def query(self):
- """Query contents from vector database based on vector similarity"""
- raise NotImplementedError
-
- def count(self) -> int:
- """
- Count number of documents/chunks embedded in the database.
-
- :return: number of documents
- :rtype: int
- """
- raise NotImplementedError
-
- def reset(self):
- """
- Resets the database. Deletes all embeddings irreversibly.
- """
- raise NotImplementedError
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
-
- :param name: Name of the collection.
- :type name: str
- """
- raise NotImplementedError
-
- def delete(self):
- """Delete from database."""
-
- raise NotImplementedError
diff --git a/embedchain/embedchain/vectordb/chroma.py b/embedchain/embedchain/vectordb/chroma.py
deleted file mode 100644
index 746dc149b..000000000
--- a/embedchain/embedchain/vectordb/chroma.py
+++ /dev/null
@@ -1,290 +0,0 @@
-import logging
-from typing import Any, Optional, Union
-
-from chromadb import Collection, QueryResult
-from langchain.docstore.document import Document
-from tqdm import tqdm
-
-from embedchain.config import ChromaDbConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.vectordb.base import BaseVectorDB
-
-try:
- import chromadb
- from chromadb.config import Settings
- from chromadb.errors import InvalidDimensionException
-except RuntimeError:
- from embedchain.utils.misc import use_pysqlite3
-
- use_pysqlite3()
- import chromadb
- from chromadb.config import Settings
- from chromadb.errors import InvalidDimensionException
-
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class ChromaDB(BaseVectorDB):
- """Vector database using ChromaDB."""
-
- def __init__(self, config: Optional[ChromaDbConfig] = None):
- """Initialize a new ChromaDB instance
-
- :param config: Configuration options for Chroma, defaults to None
- :type config: Optional[ChromaDbConfig], optional
- """
- if config:
- self.config = config
- else:
- self.config = ChromaDbConfig()
-
- self.settings = Settings(anonymized_telemetry=False)
- self.settings.allow_reset = self.config.allow_reset if hasattr(self.config, "allow_reset") else False
- self.batch_size = self.config.batch_size
- if self.config.chroma_settings:
- for key, value in self.config.chroma_settings.items():
- if hasattr(self.settings, key):
- setattr(self.settings, key, value)
-
- if self.config.host and self.config.port:
- logger.info(f"Connecting to ChromaDB server: {self.config.host}:{self.config.port}")
- self.settings.chroma_server_host = self.config.host
- self.settings.chroma_server_http_port = self.config.port
- self.settings.chroma_api_impl = "chromadb.api.fastapi.FastAPI"
- else:
- if self.config.dir is None:
- self.config.dir = "db"
-
- self.settings.persist_directory = self.config.dir
- self.settings.is_persistent = True
-
- self.client = chromadb.Client(self.settings)
- super().__init__(config=self.config)
-
- def _initialize(self):
- """
- This method is needed because `embedder` attribute needs to be set externally before it can be initialized.
- """
- if not self.embedder:
- raise ValueError(
- "Embedder not set. Please set an embedder with `_set_embedder()` function before initialization."
- )
- self._get_or_create_collection(self.config.collection_name)
-
- def _get_or_create_db(self):
- """Called during initialization"""
- return self.client
-
- @staticmethod
- def _generate_where_clause(where: dict[str, any]) -> dict[str, any]:
- # If only one filter is supplied, return it as is
- # (no need to wrap in $and based on chroma docs)
- if where is None:
- return {}
- if len(where.keys()) <= 1:
- return where
- where_filters = []
- for k, v in where.items():
- if isinstance(v, str):
- where_filters.append({k: v})
- return {"$and": where_filters}
-
- def _get_or_create_collection(self, name: str) -> Collection:
- """
- Get or create a named collection.
-
- :param name: Name of the collection
- :type name: str
- :raises ValueError: No embedder configured.
- :return: Created collection
- :rtype: Collection
- """
- if not hasattr(self, "embedder") or not self.embedder:
- raise ValueError("Cannot create a Chroma database collection without an embedder.")
- self.collection = self.client.get_or_create_collection(
- name=name,
- embedding_function=self.embedder.embedding_fn,
- )
- return self.collection
-
- def get(self, ids: Optional[list[str]] = None, where: Optional[dict[str, any]] = None, limit: Optional[int] = None):
- """
- Get existing doc ids present in vector database
-
- :param ids: list of doc ids to check for existence
- :type ids: list[str]
- :param where: Optional. to filter data
- :type where: dict[str, Any]
- :param limit: Optional. maximum number of documents
- :type limit: Optional[int]
- :return: Existing documents.
- :rtype: list[str]
- """
- args = {}
- if ids:
- args["ids"] = ids
- if where:
- args["where"] = self._generate_where_clause(where)
- if limit:
- args["limit"] = limit
- return self.collection.get(**args)
-
- def add(
- self,
- documents: list[str],
- metadatas: list[object],
- ids: list[str],
- **kwargs: Optional[dict[str, Any]],
- ) -> Any:
- """
- Add vectors to chroma database
-
- :param documents: Documents
- :type documents: list[str]
- :param metadatas: Metadatas
- :type metadatas: list[object]
- :param ids: ids
- :type ids: list[str]
- """
- size = len(documents)
- if len(documents) != size or len(metadatas) != size or len(ids) != size:
- raise ValueError(
- "Cannot add documents to chromadb with inconsistent sizes. Documents size: {}, Metadata size: {},"
- " Ids size: {}".format(len(documents), len(metadatas), len(ids))
- )
-
- for i in tqdm(range(0, len(documents), self.batch_size), desc="Inserting batches in chromadb"):
- self.collection.add(
- documents=documents[i : i + self.batch_size],
- metadatas=metadatas[i : i + self.batch_size],
- ids=ids[i : i + self.batch_size],
- )
- self.config
-
- @staticmethod
- def _format_result(results: QueryResult) -> list[tuple[Document, float]]:
- """
- Format Chroma results
-
- :param results: ChromaDB query results to format.
- :type results: QueryResult
- :return: Formatted results
- :rtype: list[tuple[Document, float]]
- """
- return [
- (Document(page_content=result[0], metadata=result[1] or {}), result[2])
- for result in zip(
- results["documents"][0],
- results["metadatas"][0],
- results["distances"][0],
- )
- ]
-
- def query(
- self,
- input_query: str,
- n_results: int,
- where: Optional[dict[str, any]] = None,
- raw_filter: Optional[dict[str, any]] = None,
- citations: bool = False,
- **kwargs: Optional[dict[str, any]],
- ) -> Union[list[tuple[str, dict]], list[str]]:
- """
- Query contents from vector database based on vector similarity
-
- :param input_query: query string
- :type input_query: str
- :param n_results: no of similar documents to fetch from database
- :type n_results: int
- :param where: to filter data
- :type where: dict[str, Any]
- :param raw_filter: Raw filter to apply
- :type raw_filter: dict[str, Any]
- :param citations: we use citations boolean param to return context along with the answer.
- :type citations: bool, default is False.
- :raises InvalidDimensionException: Dimensions do not match.
- :return: The content of the document that matched your query,
- along with url of the source and doc_id (if citations flag is true)
- :rtype: list[str], if citations=False, otherwise list[tuple[str, str, str]]
- """
- if where and raw_filter:
- raise ValueError("Both `where` and `raw_filter` cannot be used together.")
-
- where_clause = None
- if raw_filter:
- where_clause = raw_filter
- if where:
- where_clause = self._generate_where_clause(where)
- try:
- result = self.collection.query(
- query_texts=[
- input_query,
- ],
- n_results=n_results,
- where=where_clause,
- )
- except InvalidDimensionException as e:
- raise InvalidDimensionException(
- e.message()
- + ". This is commonly a side-effect when an embedding function, different from the one used to add the"
- " embeddings, is used to retrieve an embedding from the database."
- ) from None
- results_formatted = self._format_result(result)
- contexts = []
- for result in results_formatted:
- context = result[0].page_content
- if citations:
- metadata = result[0].metadata
- metadata["score"] = result[1]
- contexts.append((context, metadata))
- else:
- contexts.append(context)
- return contexts
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
-
- :param name: Name of the collection.
- :type name: str
- """
- if not isinstance(name, str):
- raise TypeError("Collection name must be a string")
- self.config.collection_name = name
- self._get_or_create_collection(self.config.collection_name)
-
- def count(self) -> int:
- """
- Count number of documents/chunks embedded in the database.
-
- :return: number of documents
- :rtype: int
- """
- return self.collection.count()
-
- def delete(self, where):
- return self.collection.delete(where=self._generate_where_clause(where))
-
- def reset(self):
- """
- Resets the database. Deletes all embeddings irreversibly.
- """
- # Delete all data from the collection
- try:
- self.client.delete_collection(self.config.collection_name)
- except ValueError:
- raise ValueError(
- "For safety reasons, resetting is disabled. "
- "Please enable it by setting `allow_reset=True` in your ChromaDbConfig"
- ) from None
- # Recreate
- self._get_or_create_collection(self.config.collection_name)
-
- # Todo: Automatically recreating a collection with the same name cannot be the best way to handle a reset.
- # A downside of this implementation is, if you have two instances,
- # the other instance will not get the updated `self.collection` attribute.
- # A better way would be to create the collection if it is called again after being reset.
- # That means, checking if collection exists in the db-consuming methods, and creating it if it doesn't.
- # That's an extra steps for all uses, just to satisfy a niche use case in a niche method. For now, this will do.
diff --git a/embedchain/embedchain/vectordb/elasticsearch.py b/embedchain/embedchain/vectordb/elasticsearch.py
deleted file mode 100644
index 12b871762..000000000
--- a/embedchain/embedchain/vectordb/elasticsearch.py
+++ /dev/null
@@ -1,269 +0,0 @@
-import logging
-from typing import Any, Optional, Union
-
-try:
- from elasticsearch import Elasticsearch
- from elasticsearch.helpers import bulk
-except ImportError:
- raise ImportError(
- "Elasticsearch requires extra dependencies. Install with `pip install --upgrade embedchain[elasticsearch]`"
- ) from None
-
-from embedchain.config import ElasticsearchDBConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.utils.misc import chunks
-from embedchain.vectordb.base import BaseVectorDB
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class ElasticsearchDB(BaseVectorDB):
- """
- Elasticsearch as vector database
- """
-
- def __init__(
- self,
- config: Optional[ElasticsearchDBConfig] = None,
- es_config: Optional[ElasticsearchDBConfig] = None, # Backwards compatibility
- ):
- """Elasticsearch as vector database.
-
- :param config: Elasticsearch database config, defaults to None
- :type config: ElasticsearchDBConfig, optional
- :param es_config: `es_config` is supported as an alias for `config` (for backwards compatibility),
- defaults to None
- :type es_config: ElasticsearchDBConfig, optional
- :raises ValueError: No config provided
- """
- if config is None and es_config is None:
- self.config = ElasticsearchDBConfig()
- else:
- if not isinstance(config, ElasticsearchDBConfig):
- raise TypeError(
- "config is not a `ElasticsearchDBConfig` instance. "
- "Please make sure the type is right and that you are passing an instance."
- )
- self.config = config or es_config
- if self.config.ES_URL:
- self.client = Elasticsearch(self.config.ES_URL, **self.config.ES_EXTRA_PARAMS)
- elif self.config.CLOUD_ID:
- self.client = Elasticsearch(cloud_id=self.config.CLOUD_ID, **self.config.ES_EXTRA_PARAMS)
- else:
- raise ValueError(
- "Something is wrong with your config. Please check again - `https://docs.embedchain.ai/components/vector-databases#elasticsearch`" # noqa: E501
- )
-
- self.batch_size = self.config.batch_size
- # Call parent init here because embedder is needed
- super().__init__(config=self.config)
-
- def _initialize(self):
- """
- This method is needed because `embedder` attribute needs to be set externally before it can be initialized.
- """
- logger.info(self.client.info())
- index_settings = {
- "mappings": {
- "properties": {
- "text": {"type": "text"},
- "embeddings": {"type": "dense_vector", "index": False, "dims": self.embedder.vector_dimension},
- }
- }
- }
- es_index = self._get_index()
- if not self.client.indices.exists(index=es_index):
- # create index if not exist
- print("Creating index", es_index, index_settings)
- self.client.indices.create(index=es_index, body=index_settings)
-
- def _get_or_create_db(self):
- """Called during initialization"""
- return self.client
-
- def _get_or_create_collection(self, name):
- """Note: nothing to return here. Discuss later"""
-
- def get(self, ids: Optional[list[str]] = None, where: Optional[dict[str, any]] = None, limit: Optional[int] = None):
- """
- Get existing doc ids present in vector database
-
- :param ids: _list of doc ids to check for existence
- :type ids: list[str]
- :param where: to filter data
- :type where: dict[str, any]
- :return: ids
- :rtype: Set[str]
- """
- if ids:
- query = {"bool": {"must": [{"ids": {"values": ids}}]}}
- else:
- query = {"bool": {"must": []}}
-
- if where:
- for key, value in where.items():
- query["bool"]["must"].append({"term": {f"metadata.{key}.keyword": value}})
-
- response = self.client.search(index=self._get_index(), query=query, _source=True, size=limit)
- docs = response["hits"]["hits"]
- ids = [doc["_id"] for doc in docs]
- doc_ids = [doc["_source"]["metadata"]["doc_id"] for doc in docs]
-
- # Result is modified for compatibility with other vector databases
- # TODO: Add method in vector database to return result in a standard format
- result = {"ids": ids, "metadatas": []}
-
- for doc_id in doc_ids:
- result["metadatas"].append({"doc_id": doc_id})
-
- return result
-
- def add(
- self,
- documents: list[str],
- metadatas: list[object],
- ids: list[str],
- **kwargs: Optional[dict[str, any]],
- ) -> Any:
- """
- add data in vector database
- :param documents: list of texts to add
- :type documents: list[str]
- :param metadatas: list of metadata associated with docs
- :type metadatas: list[object]
- :param ids: ids of docs
- :type ids: list[str]
- """
-
- embeddings = self.embedder.embedding_fn(documents)
-
- for chunk in chunks(
- list(zip(ids, documents, metadatas, embeddings)),
- self.batch_size,
- desc="Inserting batches in elasticsearch",
- ): # noqa: E501
- ids, docs, metadatas, embeddings = [], [], [], []
- for id, text, metadata, embedding in chunk:
- ids.append(id)
- docs.append(text)
- metadatas.append(metadata)
- embeddings.append(embedding)
-
- batch_docs = []
- for id, text, metadata, embedding in zip(ids, docs, metadatas, embeddings):
- batch_docs.append(
- {
- "_index": self._get_index(),
- "_id": id,
- "_source": {"text": text, "metadata": metadata, "embeddings": embedding},
- }
- )
- bulk(self.client, batch_docs, **kwargs)
- self.client.indices.refresh(index=self._get_index())
-
- def query(
- self,
- input_query: str,
- n_results: int,
- where: dict[str, any],
- citations: bool = False,
- **kwargs: Optional[dict[str, Any]],
- ) -> Union[list[tuple[str, dict]], list[str]]:
- """
- query contents from vector database based on vector similarity
-
- :param input_query: query string
- :type input_query: str
- :param n_results: no of similar documents to fetch from database
- :type n_results: int
- :param where: Optional. to filter data
- :type where: dict[str, any]
- :return: The context of the document that matched your query, url of the source, doc_id
- :param citations: we use citations boolean param to return context along with the answer.
- :type citations: bool, default is False.
- :return: The content of the document that matched your query,
- along with url of the source and doc_id (if citations flag is true)
- :rtype: list[str], if citations=False, otherwise list[tuple[str, str, str]]
- """
- input_query_vector = self.embedder.embedding_fn([input_query])
- query_vector = input_query_vector[0]
-
- # `https://www.elastic.co/guide/en/elasticsearch/reference/7.17/query-dsl-script-score-query.html`
- query = {
- "script_score": {
- "query": {"bool": {"must": [{"exists": {"field": "text"}}]}},
- "script": {
- "source": "cosineSimilarity(params.input_query_vector, 'embeddings') + 1.0",
- "params": {"input_query_vector": query_vector},
- },
- }
- }
-
- if where:
- for key, value in where.items():
- query["script_score"]["query"]["bool"]["must"].append({"term": {f"metadata.{key}.keyword": value}})
-
- _source = ["text", "metadata"]
- response = self.client.search(index=self._get_index(), query=query, _source=_source, size=n_results)
- docs = response["hits"]["hits"]
- contexts = []
- for doc in docs:
- context = doc["_source"]["text"]
- if citations:
- metadata = doc["_source"]["metadata"]
- metadata["score"] = doc["_score"]
- contexts.append(tuple((context, metadata)))
- else:
- contexts.append(context)
- return contexts
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
-
- :param name: Name of the collection.
- :type name: str
- """
- if not isinstance(name, str):
- raise TypeError("Collection name must be a string")
- self.config.collection_name = name
-
- def count(self) -> int:
- """
- Count number of documents/chunks embedded in the database.
-
- :return: number of documents
- :rtype: int
- """
- query = {"match_all": {}}
- response = self.client.count(index=self._get_index(), query=query)
- doc_count = response["count"]
- return doc_count
-
- def reset(self):
- """
- Resets the database. Deletes all embeddings irreversibly.
- """
- # Delete all data from the database
- if self.client.indices.exists(index=self._get_index()):
- # delete index in Es
- self.client.indices.delete(index=self._get_index())
-
- def _get_index(self) -> str:
- """Get the Elasticsearch index for a collection
-
- :return: Elasticsearch index
- :rtype: str
- """
- # NOTE: The method is preferred to an attribute, because if collection name changes,
- # it's always up-to-date.
- return f"{self.config.collection_name}_{self.embedder.vector_dimension}".lower()
-
- def delete(self, where):
- """Delete documents from the database."""
- query = {"query": {"bool": {"must": []}}}
- for key, value in where.items():
- query["query"]["bool"]["must"].append({"term": {f"metadata.{key}.keyword": value}})
- self.client.delete_by_query(index=self._get_index(), body=query)
- self.client.indices.refresh(index=self._get_index())
diff --git a/embedchain/embedchain/vectordb/lancedb.py b/embedchain/embedchain/vectordb/lancedb.py
deleted file mode 100644
index d3d4b6898..000000000
--- a/embedchain/embedchain/vectordb/lancedb.py
+++ /dev/null
@@ -1,305 +0,0 @@
-from typing import Any, Dict, List, Optional, Union
-
-import pyarrow as pa
-
-try:
- import lancedb
-except ImportError:
- raise ImportError('LanceDB is required. Install with pip install "embedchain[lancedb]"') from None
-
-from embedchain.config.vector_db.lancedb import LanceDBConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.vectordb.base import BaseVectorDB
-
-
-@register_deserializable
-class LanceDB(BaseVectorDB):
- """
- LanceDB as vector database
- """
-
- def __init__(
- self,
- config: Optional[LanceDBConfig] = None,
- ):
- """LanceDB as vector database.
-
- :param config: LanceDB database config, defaults to None
- :type config: LanceDBConfig, optional
- """
- if config:
- self.config = config
- else:
- self.config = LanceDBConfig()
-
- self.client = lancedb.connect(self.config.dir or "~/.lancedb")
- self.embedder_check = True
-
- super().__init__(config=self.config)
-
- def _initialize(self):
- """
- This method is needed because `embedder` attribute needs to be set externally before it can be initialized.
- """
- if not self.embedder:
- raise ValueError(
- "Embedder not set. Please set an embedder with `_set_embedder()` function before initialization."
- )
- else:
- # check embedder function is working or not
- try:
- self.embedder.embedding_fn("Hello LanceDB")
- except Exception:
- self.embedder_check = False
-
- self._get_or_create_collection(self.config.collection_name)
-
- def _get_or_create_db(self):
- """
- Called during initialization
- """
- return self.client
-
- def _generate_where_clause(self, where: Dict[str, any]) -> str:
- """
- This method generate where clause using dictionary containing attributes and their values
- """
-
- where_filters = ""
-
- if len(list(where.keys())) == 1:
- where_filters = f"{list(where.keys())[0]} = {list(where.values())[0]}"
- return where_filters
-
- where_items = list(where.items())
- where_count = len(where_items)
-
- for i, (key, value) in enumerate(where_items, start=1):
- condition = f"{key} = {value} AND "
- where_filters += condition
-
- if i == where_count:
- condition = f"{key} = {value}"
- where_filters += condition
-
- return where_filters
-
- def _get_or_create_collection(self, table_name: str, reset=False):
- """
- Get or create a named collection.
-
- :param name: Name of the collection
- :type name: str
- :return: Created collection
- :rtype: Collection
- """
- if not self.embedder_check:
- schema = pa.schema(
- [
- pa.field("doc", pa.string()),
- pa.field("metadata", pa.string()),
- pa.field("id", pa.string()),
- ]
- )
-
- else:
- schema = pa.schema(
- [
- pa.field("vector", pa.list_(pa.float32(), list_size=self.embedder.vector_dimension)),
- pa.field("doc", pa.string()),
- pa.field("metadata", pa.string()),
- pa.field("id", pa.string()),
- ]
- )
-
- if not reset:
- if table_name not in self.client.table_names():
- self.collection = self.client.create_table(table_name, schema=schema)
-
- else:
- self.client.drop_table(table_name)
- self.collection = self.client.create_table(table_name, schema=schema)
-
- self.collection = self.client[table_name]
-
- return self.collection
-
- def get(self, ids: Optional[List[str]] = None, where: Optional[Dict[str, any]] = None, limit: Optional[int] = None):
- """
- Get existing doc ids present in vector database
-
- :param ids: list of doc ids to check for existence
- :type ids: List[str]
- :param where: Optional. to filter data
- :type where: Dict[str, Any]
- :param limit: Optional. maximum number of documents
- :type limit: Optional[int]
- :return: Existing documents.
- :rtype: List[str]
- """
- if limit is not None:
- max_limit = limit
- else:
- max_limit = 3
- results = {"ids": [], "metadatas": []}
-
- where_clause = {}
- if where:
- where_clause = self._generate_where_clause(where)
-
- if ids is not None:
- records = (
- self.collection.to_lance().scanner(filter=f"id IN {tuple(ids)}", columns=["id"]).to_table().to_pydict()
- )
- for id in records["id"]:
- if where is not None:
- result = (
- self.collection.search(query=id, vector_column_name="id")
- .where(where_clause)
- .limit(max_limit)
- .to_list()
- )
- else:
- result = self.collection.search(query=id, vector_column_name="id").limit(max_limit).to_list()
- results["ids"] = [r["id"] for r in result]
- results["metadatas"] = [r["metadata"] for r in result]
-
- return results
-
- def add(
- self,
- documents: List[str],
- metadatas: List[object],
- ids: List[str],
- ) -> Any:
- """
- Add vectors to lancedb database
-
- :param documents: Documents
- :type documents: List[str]
- :param metadatas: Metadatas
- :type metadatas: List[object]
- :param ids: ids
- :type ids: List[str]
- """
- data = []
- to_ingest = list(zip(documents, metadatas, ids))
-
- if not self.embedder_check:
- for doc, meta, id in to_ingest:
- temp = {}
- temp["doc"] = doc
- temp["metadata"] = str(meta)
- temp["id"] = id
- data.append(temp)
- else:
- for doc, meta, id in to_ingest:
- temp = {}
- temp["doc"] = doc
- temp["vector"] = self.embedder.embedding_fn([doc])[0]
- temp["metadata"] = str(meta)
- temp["id"] = id
- data.append(temp)
-
- self.collection.add(data=data)
-
- def _format_result(self, results) -> list:
- """
- Format LanceDB results
-
- :param results: LanceDB query results to format.
- :type results: QueryResult
- :return: Formatted results
- :rtype: list[tuple[Document, float]]
- """
- return results.tolist()
-
- def query(
- self,
- input_query: str,
- n_results: int = 3,
- where: Optional[dict[str, any]] = None,
- raw_filter: Optional[dict[str, any]] = None,
- citations: bool = False,
- **kwargs: Optional[dict[str, any]],
- ) -> Union[list[tuple[str, dict]], list[str]]:
- """
- Query contents from vector database based on vector similarity
-
- :param input_query: query string
- :type input_query: str
- :param n_results: no of similar documents to fetch from database
- :type n_results: int
- :param where: to filter data
- :type where: dict[str, Any]
- :param raw_filter: Raw filter to apply
- :type raw_filter: dict[str, Any]
- :param citations: we use citations boolean param to return context along with the answer.
- :type citations: bool, default is False.
- :raises InvalidDimensionException: Dimensions do not match.
- :return: The content of the document that matched your query,
- along with url of the source and doc_id (if citations flag is true)
- :rtype: list[str], if citations=False, otherwise list[tuple[str, str, str]]
- """
- if where and raw_filter:
- raise ValueError("Both `where` and `raw_filter` cannot be used together.")
- try:
- query_embedding = self.embedder.embedding_fn(input_query)[0]
- result = self.collection.search(query_embedding).limit(n_results).to_list()
- except Exception as e:
- e.message()
-
- results_formatted = result
-
- contexts = []
- for result in results_formatted:
- if citations:
- metadata = result["metadata"]
- contexts.append((result["doc"], metadata))
- else:
- contexts.append(result["doc"])
- return contexts
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
-
- :param name: Name of the collection.
- :type name: str
- """
- if not isinstance(name, str):
- raise TypeError("Collection name must be a string")
- self.config.collection_name = name
- self._get_or_create_collection(self.config.collection_name)
-
- def count(self) -> int:
- """
- Count number of documents/chunks embedded in the database.
-
- :return: number of documents
- :rtype: int
- """
- return self.collection.count_rows()
-
- def delete(self, where):
- return self.collection.delete(where=where)
-
- def reset(self):
- """
- Resets the database. Deletes all embeddings irreversibly.
- """
- # Delete all data from the collection and recreate collection
- if self.config.allow_reset:
- try:
- self._get_or_create_collection(self.config.collection_name, reset=True)
- except ValueError:
- raise ValueError(
- "For safety reasons, resetting is disabled. "
- "Please enable it by setting `allow_reset=True` in your LanceDbConfig"
- ) from None
- # Recreate
- else:
- print(
- "For safety reasons, resetting is disabled. "
- "Please enable it by setting `allow_reset=True` in your LanceDbConfig"
- )
diff --git a/embedchain/embedchain/vectordb/opensearch.py b/embedchain/embedchain/vectordb/opensearch.py
deleted file mode 100644
index accec4324..000000000
--- a/embedchain/embedchain/vectordb/opensearch.py
+++ /dev/null
@@ -1,253 +0,0 @@
-import logging
-import time
-from typing import Any, Optional, Union
-
-from tqdm import tqdm
-
-try:
- from opensearchpy import OpenSearch
- from opensearchpy.helpers import bulk
-except ImportError:
- raise ImportError(
- "OpenSearch requires extra dependencies. Install with `pip install --upgrade embedchain[opensearch]`"
- ) from None
-
-from langchain_community.embeddings.openai import OpenAIEmbeddings
-from langchain_community.vectorstores import OpenSearchVectorSearch
-
-from embedchain.config import OpenSearchDBConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.vectordb.base import BaseVectorDB
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class OpenSearchDB(BaseVectorDB):
- """
- OpenSearch as vector database
- """
-
- def __init__(self, config: OpenSearchDBConfig):
- """OpenSearch as vector database.
-
- :param config: OpenSearch domain config
- :type config: OpenSearchDBConfig
- """
- if config is None:
- raise ValueError("OpenSearchDBConfig is required")
- self.config = config
- self.batch_size = self.config.batch_size
- self.client = OpenSearch(
- hosts=[self.config.opensearch_url],
- http_auth=self.config.http_auth,
- **self.config.extra_params,
- )
- info = self.client.info()
- logger.info(f"Connected to {info['version']['distribution']}. Version: {info['version']['number']}")
- # Remove auth credentials from config after successful connection
- super().__init__(config=self.config)
-
- def _initialize(self):
- logger.info(self.client.info())
- index_name = self._get_index()
- if self.client.indices.exists(index=index_name):
- print(f"Index '{index_name}' already exists.")
- return
-
- index_body = {
- "settings": {"knn": True},
- "mappings": {
- "properties": {
- "text": {"type": "text"},
- "embeddings": {
- "type": "knn_vector",
- "index": False,
- "dimension": self.config.vector_dimension,
- },
- }
- },
- }
- self.client.indices.create(index_name, body=index_body)
- print(self.client.indices.get(index_name))
-
- def _get_or_create_db(self):
- """Called during initialization"""
- return self.client
-
- def _get_or_create_collection(self, name):
- """Note: nothing to return here. Discuss later"""
-
- def get(
- self, ids: Optional[list[str]] = None, where: Optional[dict[str, any]] = None, limit: Optional[int] = None
- ) -> set[str]:
- """
- Get existing doc ids present in vector database
-
- :param ids: _list of doc ids to check for existence
- :type ids: list[str]
- :param where: to filter data
- :type where: dict[str, any]
- :return: ids
- :type: set[str]
- """
- query = {}
- if ids:
- query["query"] = {"bool": {"must": [{"ids": {"values": ids}}]}}
- else:
- query["query"] = {"bool": {"must": []}}
-
- if where:
- for key, value in where.items():
- query["query"]["bool"]["must"].append({"term": {f"metadata.{key}.keyword": value}})
-
- # OpenSearch syntax is different from Elasticsearch
- response = self.client.search(index=self._get_index(), body=query, _source=True, size=limit)
- docs = response["hits"]["hits"]
- ids = [doc["_id"] for doc in docs]
- doc_ids = [doc["_source"]["metadata"]["doc_id"] for doc in docs]
-
- # Result is modified for compatibility with other vector databases
- # TODO: Add method in vector database to return result in a standard format
- result = {"ids": ids, "metadatas": []}
-
- for doc_id in doc_ids:
- result["metadatas"].append({"doc_id": doc_id})
- return result
-
- def add(self, documents: list[str], metadatas: list[object], ids: list[str], **kwargs: Optional[dict[str, any]]):
- """Adds documents to the opensearch index"""
-
- embeddings = self.embedder.embedding_fn(documents)
- for batch_start in tqdm(range(0, len(documents), self.batch_size), desc="Inserting batches in opensearch"):
- batch_end = batch_start + self.batch_size
- batch_documents = documents[batch_start:batch_end]
- batch_embeddings = embeddings[batch_start:batch_end]
-
- # Create document entries for bulk upload
- batch_entries = [
- {
- "_index": self._get_index(),
- "_id": doc_id,
- "_source": {"text": text, "metadata": metadata, "embeddings": embedding},
- }
- for doc_id, text, metadata, embedding in zip(
- ids[batch_start:batch_end], batch_documents, metadatas[batch_start:batch_end], batch_embeddings
- )
- ]
-
- # Perform bulk operation
- bulk(self.client, batch_entries, **kwargs)
- self.client.indices.refresh(index=self._get_index())
-
- # Sleep to avoid rate limiting
- time.sleep(0.1)
-
- def query(
- self,
- input_query: str,
- n_results: int,
- where: dict[str, any],
- citations: bool = False,
- **kwargs: Optional[dict[str, Any]],
- ) -> Union[list[tuple[str, dict]], list[str]]:
- """
- query contents from vector database based on vector similarity
-
- :param input_query: query string
- :type input_query: str
- :param n_results: no of similar documents to fetch from database
- :type n_results: int
- :param where: Optional. to filter data
- :type where: dict[str, any]
- :param citations: we use citations boolean param to return context along with the answer.
- :type citations: bool, default is False.
- :return: The content of the document that matched your query,
- along with url of the source and doc_id (if citations flag is true)
- :rtype: list[str], if citations=False, otherwise list[tuple[str, str, str]]
- """
- embeddings = OpenAIEmbeddings()
- docsearch = OpenSearchVectorSearch(
- index_name=self._get_index(),
- embedding_function=embeddings,
- opensearch_url=f"{self.config.opensearch_url}",
- http_auth=self.config.http_auth,
- use_ssl=hasattr(self.config, "use_ssl") and self.config.use_ssl,
- verify_certs=hasattr(self.config, "verify_certs") and self.config.verify_certs,
- )
-
- pre_filter = {"match_all": {}} # default
- if len(where) > 0:
- pre_filter = {"bool": {"must": []}}
- for key, value in where.items():
- pre_filter["bool"]["must"].append({"term": {f"metadata.{key}.keyword": value}})
-
- docs = docsearch.similarity_search_with_score(
- input_query,
- search_type="script_scoring",
- space_type="cosinesimil",
- vector_field="embeddings",
- text_field="text",
- metadata_field="metadata",
- pre_filter=pre_filter,
- k=n_results,
- **kwargs,
- )
-
- contexts = []
- for doc, score in docs:
- context = doc.page_content
- if citations:
- metadata = doc.metadata
- metadata["score"] = score
- contexts.append(tuple((context, metadata)))
- else:
- contexts.append(context)
- return contexts
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
-
- :param name: Name of the collection.
- :type name: str
- """
- if not isinstance(name, str):
- raise TypeError("Collection name must be a string")
- self.config.collection_name = name
-
- def count(self) -> int:
- """
- Count number of documents/chunks embedded in the database.
-
- :return: number of documents
- :rtype: int
- """
- query = {"query": {"match_all": {}}}
- response = self.client.count(index=self._get_index(), body=query)
- doc_count = response["count"]
- return doc_count
-
- def reset(self):
- """
- Resets the database. Deletes all embeddings irreversibly.
- """
- # Delete all data from the database
- if self.client.indices.exists(index=self._get_index()):
- # delete index in ES
- self.client.indices.delete(index=self._get_index())
-
- def delete(self, where):
- """Deletes a document from the OpenSearch index"""
- query = {"query": {"bool": {"must": []}}}
- for key, value in where.items():
- query["query"]["bool"]["must"].append({"term": {f"metadata.{key}.keyword": value}})
- self.client.delete_by_query(index=self._get_index(), body=query)
-
- def _get_index(self) -> str:
- """Get the OpenSearch index for a collection
-
- :return: OpenSearch index
- :rtype: str
- """
- return self.config.collection_name
diff --git a/embedchain/embedchain/vectordb/pinecone.py b/embedchain/embedchain/vectordb/pinecone.py
deleted file mode 100644
index 3c0520ce3..000000000
--- a/embedchain/embedchain/vectordb/pinecone.py
+++ /dev/null
@@ -1,252 +0,0 @@
-import logging
-import os
-from typing import Optional, Union
-
-try:
- import pinecone
-except ImportError:
- raise ImportError(
- "Pinecone requires extra dependencies. Install with `pip install pinecone-text pinecone-client`"
- ) from None
-
-from pinecone_text.sparse import BM25Encoder
-
-from embedchain.config.vector_db.pinecone import PineconeDBConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.utils.misc import chunks
-from embedchain.vectordb.base import BaseVectorDB
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class PineconeDB(BaseVectorDB):
- """
- Pinecone as vector database
- """
-
- def __init__(
- self,
- config: Optional[PineconeDBConfig] = None,
- ):
- """Pinecone as vector database.
-
- :param config: Pinecone database config, defaults to None
- :type config: PineconeDBConfig, optional
- :raises ValueError: No config provided
- """
- if config is None:
- self.config = PineconeDBConfig()
- else:
- if not isinstance(config, PineconeDBConfig):
- raise TypeError(
- "config is not a `PineconeDBConfig` instance. "
- "Please make sure the type is right and that you are passing an instance."
- )
- self.config = config
- self._setup_pinecone_index()
-
- # Setup BM25Encoder if sparse vectors are to be used
- self.bm25_encoder = None
- self.batch_size = self.config.batch_size
- if self.config.hybrid_search:
- logger.info("Initializing BM25Encoder for sparse vectors..")
- self.bm25_encoder = self.config.bm25_encoder if self.config.bm25_encoder else BM25Encoder.default()
-
- # Call parent init here because embedder is needed
- super().__init__(config=self.config)
-
- def _initialize(self):
- """
- This method is needed because `embedder` attribute needs to be set externally before it can be initialized.
- """
- if not self.embedder:
- raise ValueError("Embedder not set. Please set an embedder with `set_embedder` before initialization.")
-
- def _setup_pinecone_index(self):
- """
- Loads the Pinecone index or creates it if not present.
- """
- api_key = self.config.api_key or os.environ.get("PINECONE_API_KEY")
- if not api_key:
- raise ValueError("Please set the PINECONE_API_KEY environment variable or pass it in config.")
- self.client = pinecone.Pinecone(api_key=api_key, **self.config.extra_params)
- indexes = self.client.list_indexes().names()
- if indexes is None or self.config.index_name not in indexes:
- if self.config.pod_config:
- spec = pinecone.PodSpec(**self.config.pod_config)
- elif self.config.serverless_config:
- spec = pinecone.ServerlessSpec(**self.config.serverless_config)
- else:
- raise ValueError("No pod_config or serverless_config found.")
-
- self.client.create_index(
- name=self.config.index_name,
- metric=self.config.metric,
- dimension=self.config.vector_dimension,
- spec=spec,
- )
- self.pinecone_index = self.client.Index(self.config.index_name)
-
- def get(self, ids: Optional[list[str]] = None, where: Optional[dict[str, any]] = None, limit: Optional[int] = None):
- """
- Get existing doc ids present in vector database
-
- :param ids: _list of doc ids to check for existence
- :type ids: list[str]
- :param where: to filter data
- :type where: dict[str, any]
- :return: ids
- :rtype: Set[str]
- """
- existing_ids = list()
- metadatas = []
-
- if ids is not None:
- for i in range(0, len(ids), self.batch_size):
- result = self.pinecone_index.fetch(ids=ids[i : i + self.batch_size])
- vectors = result.get("vectors")
- batch_existing_ids = list(vectors.keys())
- existing_ids.extend(batch_existing_ids)
- metadatas.extend([vectors.get(ids).get("metadata") for ids in batch_existing_ids])
- return {"ids": existing_ids, "metadatas": metadatas}
-
- def add(
- self,
- documents: list[str],
- metadatas: list[object],
- ids: list[str],
- **kwargs: Optional[dict[str, any]],
- ):
- """add data in vector database
-
- :param documents: list of texts to add
- :type documents: list[str]
- :param metadatas: list of metadata associated with docs
- :type metadatas: list[object]
- :param ids: ids of docs
- :type ids: list[str]
- """
- docs = []
- embeddings = self.embedder.embedding_fn(documents)
- for id, text, metadata, embedding in zip(ids, documents, metadatas, embeddings):
- # Insert sparse vectors as well if the user wants to do the hybrid search
- sparse_vector_dict = (
- {"sparse_values": self.bm25_encoder.encode_documents(text)} if self.bm25_encoder else {}
- )
- docs.append(
- {
- "id": id,
- "values": embedding,
- "metadata": {**metadata, "text": text},
- **sparse_vector_dict,
- },
- )
-
- for chunk in chunks(docs, self.batch_size, desc="Adding chunks in batches"):
- self.pinecone_index.upsert(chunk, **kwargs)
-
- def query(
- self,
- input_query: str,
- n_results: int,
- where: Optional[dict[str, any]] = None,
- raw_filter: Optional[dict[str, any]] = None,
- citations: bool = False,
- app_id: Optional[str] = None,
- **kwargs: Optional[dict[str, any]],
- ) -> Union[list[tuple[str, dict]], list[str]]:
- """
- Query contents from vector database based on vector similarity.
-
- Args:
- input_query (str): query string.
- n_results (int): Number of similar documents to fetch from the database.
- where (dict[str, any], optional): Filter criteria for the search.
- raw_filter (dict[str, any], optional): Advanced raw filter criteria for the search.
- citations (bool, optional): Flag to return context along with metadata. Defaults to False.
- app_id (str, optional): Application ID to be passed to Pinecone.
-
- Returns:
- Union[list[tuple[str, dict]], list[str]]: List of document contexts, optionally with metadata.
- """
- query_filter = raw_filter if raw_filter is not None else self._generate_filter(where)
- if app_id:
- query_filter["app_id"] = {"$eq": app_id}
-
- query_vector = self.embedder.embedding_fn([input_query])[0]
- params = {
- "vector": query_vector,
- "filter": query_filter,
- "top_k": n_results,
- "include_metadata": True,
- **kwargs,
- }
-
- if self.bm25_encoder:
- sparse_query_vector = self.bm25_encoder.encode_queries(input_query)
- params["sparse_vector"] = sparse_query_vector
-
- data = self.pinecone_index.query(**params)
- return [
- (metadata.get("text"), {**metadata, "score": doc.get("score")}) if citations else metadata.get("text")
- for doc in data.get("matches", [])
- for metadata in [doc.get("metadata", {})]
- ]
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
-
- :param name: Name of the collection.
- :type name: str
- """
- if not isinstance(name, str):
- raise TypeError("Collection name must be a string")
- self.config.collection_name = name
-
- def count(self) -> int:
- """
- Count number of documents/chunks embedded in the database.
-
- :return: number of documents
- :rtype: int
- """
- data = self.pinecone_index.describe_index_stats()
- return data["total_vector_count"]
-
- def _get_or_create_db(self):
- """Called during initialization"""
- return self.client
-
- def reset(self):
- """
- Resets the database. Deletes all embeddings irreversibly.
- """
- # Delete all data from the database
- self.client.delete_index(self.config.index_name)
- self._setup_pinecone_index()
-
- @staticmethod
- def _generate_filter(where: dict):
- query = {}
- if where is None:
- return query
-
- for k, v in where.items():
- query[k] = {"$eq": v}
- return query
-
- def delete(self, where: dict):
- """Delete from database.
- :param ids: list of ids to delete
- :type ids: list[str]
- """
- # Deleting with filters is not supported for `starter` index type.
- # Follow `https://docs.pinecone.io/docs/metadata-filtering#deleting-vectors-by-metadata-filter` for more details
- db_filter = self._generate_filter(where)
- try:
- self.pinecone_index.delete(filter=db_filter)
- except Exception as e:
- print(f"Failed to delete from Pinecone: {e}")
- return
diff --git a/embedchain/embedchain/vectordb/qdrant.py b/embedchain/embedchain/vectordb/qdrant.py
deleted file mode 100644
index cdac19cfa..000000000
--- a/embedchain/embedchain/vectordb/qdrant.py
+++ /dev/null
@@ -1,253 +0,0 @@
-import copy
-import os
-from typing import Any, Optional, Union
-
-try:
- from qdrant_client import QdrantClient
- from qdrant_client.http import models
- from qdrant_client.http.models import Batch
- from qdrant_client.models import Distance, VectorParams
-except ImportError:
- raise ImportError("Qdrant requires extra dependencies. Install with `pip install embedchain[qdrant]`") from None
-
-from tqdm import tqdm
-
-from embedchain.config.vector_db.qdrant import QdrantDBConfig
-from embedchain.vectordb.base import BaseVectorDB
-
-
-class QdrantDB(BaseVectorDB):
- """
- Qdrant as vector database
- """
-
- def __init__(self, config: QdrantDBConfig = None):
- """
- Qdrant as vector database
- :param config. Qdrant database config to be used for connection
- """
- if config is None:
- config = QdrantDBConfig()
- else:
- if not isinstance(config, QdrantDBConfig):
- raise TypeError(
- "config is not a `QdrantDBConfig` instance. "
- "Please make sure the type is right and that you are passing an instance."
- )
- self.config = config
- self.batch_size = self.config.batch_size
- self.client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
- # Call parent init here because embedder is needed
- super().__init__(config=self.config)
-
- def _initialize(self):
- """
- This method is needed because `embedder` attribute needs to be set externally before it can be initialized.
- """
- if not self.embedder:
- raise ValueError("Embedder not set. Please set an embedder with `set_embedder` before initialization.")
-
- self.collection_name = self._get_or_create_collection()
- all_collections = self.client.get_collections()
- collection_names = [collection.name for collection in all_collections.collections]
- if self.collection_name not in collection_names:
- self.client.recreate_collection(
- collection_name=self.collection_name,
- vectors_config=VectorParams(
- size=self.embedder.vector_dimension,
- distance=Distance.COSINE,
- hnsw_config=self.config.hnsw_config,
- quantization_config=self.config.quantization_config,
- on_disk=self.config.on_disk,
- ),
- )
-
- def _get_or_create_db(self):
- return self.client
-
- def _get_or_create_collection(self):
- return f"{self.config.collection_name}-{self.embedder.vector_dimension}".lower().replace("_", "-")
-
- def get(self, ids: Optional[list[str]] = None, where: Optional[dict[str, any]] = None, limit: Optional[int] = None):
- """
- Get existing doc ids present in vector database
-
- :param ids: _list of doc ids to check for existence
- :type ids: list[str]
- :param where: to filter data
- :type where: dict[str, any]
- :param limit: The number of entries to be fetched
- :type limit: Optional int, defaults to None
- :return: All the existing IDs
- :rtype: Set[str]
- """
-
- keys = set(where.keys() if where is not None else set())
-
- qdrant_must_filters = []
-
- if ids:
- qdrant_must_filters.append(
- models.FieldCondition(
- key="identifier",
- match=models.MatchAny(
- any=ids,
- ),
- )
- )
-
- if len(keys) > 0:
- for key in keys:
- qdrant_must_filters.append(
- models.FieldCondition(
- key="metadata.{}".format(key),
- match=models.MatchValue(
- value=where.get(key),
- ),
- )
- )
-
- offset = 0
- existing_ids = []
- metadatas = []
- while offset is not None:
- response = self.client.scroll(
- collection_name=self.collection_name,
- scroll_filter=models.Filter(must=qdrant_must_filters),
- offset=offset,
- limit=self.batch_size,
- )
- offset = response[1]
- for doc in response[0]:
- existing_ids.append(doc.payload["identifier"])
- metadatas.append(doc.payload["metadata"])
- return {"ids": existing_ids, "metadatas": metadatas}
-
- def add(
- self,
- documents: list[str],
- metadatas: list[object],
- ids: list[str],
- **kwargs: Optional[dict[str, any]],
- ):
- """add data in vector database
- :param documents: list of texts to add
- :type documents: list[str]
- :param metadatas: list of metadata associated with docs
- :type metadatas: list[object]
- :param ids: ids of docs
- :type ids: list[str]
- """
- embeddings = self.embedder.embedding_fn(documents)
-
- payloads = []
- qdrant_ids = []
- for id, document, metadata in zip(ids, documents, metadatas):
- metadata["text"] = document
- qdrant_ids.append(id)
- payloads.append({"identifier": id, "text": document, "metadata": copy.deepcopy(metadata)})
-
- for i in tqdm(range(0, len(qdrant_ids), self.batch_size), desc="Adding data in batches"):
- self.client.upsert(
- collection_name=self.collection_name,
- points=Batch(
- ids=qdrant_ids[i : i + self.batch_size],
- payloads=payloads[i : i + self.batch_size],
- vectors=embeddings[i : i + self.batch_size],
- ),
- **kwargs,
- )
-
- def query(
- self,
- input_query: str,
- n_results: int,
- where: dict[str, any],
- citations: bool = False,
- **kwargs: Optional[dict[str, Any]],
- ) -> Union[list[tuple[str, dict]], list[str]]:
- """
- query contents from vector database based on vector similarity
- :param input_query: query string
- :type input_query: str
- :param n_results: no of similar documents to fetch from database
- :type n_results: int
- :param where: Optional. to filter data
- :type where: dict[str, any]
- :param citations: we use citations boolean param to return context along with the answer.
- :type citations: bool, default is False.
- :return: The content of the document that matched your query,
- along with url of the source and doc_id (if citations flag is true)
- :rtype: list[str], if citations=False, otherwise list[tuple[str, str, str]]
- """
- query_vector = self.embedder.embedding_fn([input_query])[0]
- keys = set(where.keys() if where is not None else set())
-
- qdrant_must_filters = []
- if len(keys) > 0:
- for key in keys:
- qdrant_must_filters.append(
- models.FieldCondition(
- key="metadata.{}".format(key),
- match=models.MatchValue(
- value=where.get(key),
- ),
- )
- )
-
- results = self.client.search(
- collection_name=self.collection_name,
- query_filter=models.Filter(must=qdrant_must_filters),
- query_vector=query_vector,
- limit=n_results,
- **kwargs,
- )
-
- contexts = []
- for result in results:
- context = result.payload["text"]
- if citations:
- metadata = result.payload["metadata"]
- metadata["score"] = result.score
- contexts.append(tuple((context, metadata)))
- else:
- contexts.append(context)
- return contexts
-
- def count(self) -> int:
- response = self.client.get_collection(collection_name=self.collection_name)
- return response.points_count
-
- def reset(self):
- self.client.delete_collection(collection_name=self.collection_name)
- self._initialize()
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
-
- :param name: Name of the collection.
- :type name: str
- """
- if not isinstance(name, str):
- raise TypeError("Collection name must be a string")
- self.config.collection_name = name
- self.collection_name = self._get_or_create_collection()
-
- @staticmethod
- def _generate_query(where: dict):
- must_fields = []
- for key, value in where.items():
- must_fields.append(
- models.FieldCondition(
- key=f"metadata.{key}",
- match=models.MatchValue(
- value=value,
- ),
- )
- )
- return models.Filter(must=must_fields)
-
- def delete(self, where: dict):
- db_filter = self._generate_query(where)
- self.client.delete(collection_name=self.collection_name, points_selector=db_filter)
diff --git a/embedchain/embedchain/vectordb/weaviate.py b/embedchain/embedchain/vectordb/weaviate.py
deleted file mode 100644
index 897412a64..000000000
--- a/embedchain/embedchain/vectordb/weaviate.py
+++ /dev/null
@@ -1,363 +0,0 @@
-import copy
-import os
-from typing import Optional, Union
-
-try:
- import weaviate
-except ImportError:
- raise ImportError(
- "Weaviate requires extra dependencies. Install with `pip install --upgrade 'embedchain[weaviate]'`"
- ) from None
-
-from embedchain.config.vector_db.weaviate import WeaviateDBConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.vectordb.base import BaseVectorDB
-
-
-@register_deserializable
-class WeaviateDB(BaseVectorDB):
- """
- Weaviate as vector database
- """
-
- def __init__(
- self,
- config: Optional[WeaviateDBConfig] = None,
- ):
- """Weaviate as vector database.
- :param config: Weaviate database config, defaults to None
- :type config: WeaviateDBConfig, optional
- :raises ValueError: No config provided
- """
- if config is None:
- self.config = WeaviateDBConfig()
- else:
- if not isinstance(config, WeaviateDBConfig):
- raise TypeError(
- "config is not a `WeaviateDBConfig` instance. "
- "Please make sure the type is right and that you are passing an instance."
- )
- self.config = config
- self.batch_size = self.config.batch_size
- self.client = weaviate.Client(
- url=os.environ.get("WEAVIATE_ENDPOINT"),
- auth_client_secret=weaviate.AuthApiKey(api_key=os.environ.get("WEAVIATE_API_KEY")),
- **self.config.extra_params,
- )
- # Since weaviate uses graphQL, we need to keep track of metadata keys added in the vectordb.
- # This is needed to filter data while querying.
- self.metadata_keys = {"data_type", "doc_id", "url", "hash", "app_id"}
-
- # Call parent init here because embedder is needed
- super().__init__(config=self.config)
-
- def _initialize(self):
- """
- This method is needed because `embedder` attribute needs to be set externally before it can be initialized.
- """
-
- if not self.embedder:
- raise ValueError("Embedder not set. Please set an embedder with `set_embedder` before initialization.")
-
- self.index_name = self._get_index_name()
- if not self.client.schema.exists(self.index_name):
- # id is a reserved field in Weaviate, hence we had to change the name of the id field to identifier
- # The none vectorizer is crucial as we have our own custom embedding function
- """
- TODO: wait for weaviate to add indexing on `object[]` data-type so that we can add filter while querying.
- Once that is done, change `dataType` of "metadata" field to `object[]` and update the query below.
- """
- class_obj = {
- "classes": [
- {
- "class": self.index_name,
- "vectorizer": "none",
- "properties": [
- {
- "name": "identifier",
- "dataType": ["text"],
- },
- {
- "name": "text",
- "dataType": ["text"],
- },
- {
- "name": "metadata",
- "dataType": [self.index_name + "_metadata"],
- },
- ],
- },
- {
- "class": self.index_name + "_metadata",
- "vectorizer": "none",
- "properties": [
- {
- "name": "data_type",
- "dataType": ["text"],
- },
- {
- "name": "doc_id",
- "dataType": ["text"],
- },
- {
- "name": "url",
- "dataType": ["text"],
- },
- {
- "name": "hash",
- "dataType": ["text"],
- },
- {
- "name": "app_id",
- "dataType": ["text"],
- },
- ],
- },
- ]
- }
-
- self.client.schema.create(class_obj)
-
- def get(self, ids: Optional[list[str]] = None, where: Optional[dict[str, any]] = None, limit: Optional[int] = None):
- """
- Get existing doc ids present in vector database
- :param ids: _list of doc ids to check for existance
- :type ids: list[str]
- :param where: to filter data
- :type where: dict[str, any]
- :return: ids
- :rtype: Set[str]
- """
- weaviate_where_operands = []
-
- if ids:
- for doc_id in ids:
- weaviate_where_operands.append({"path": ["identifier"], "operator": "Equal", "valueText": doc_id})
-
- keys = set(where.keys() if where is not None else set())
- if len(keys) > 0:
- for key in keys:
- weaviate_where_operands.append(
- {
- "path": ["metadata", self.index_name + "_metadata", key],
- "operator": "Equal",
- "valueText": where.get(key),
- }
- )
-
- if len(weaviate_where_operands) == 1:
- weaviate_where_clause = weaviate_where_operands[0]
- else:
- weaviate_where_clause = {"operator": "And", "operands": weaviate_where_operands}
-
- existing_ids = []
- metadatas = []
- cursor = None
- offset = 0
- has_iterated_once = False
- query_metadata_keys = self.metadata_keys.union(keys)
- while cursor is not None or not has_iterated_once:
- has_iterated_once = True
- results = self._query_with_offset(
- self.client.query.get(
- self.index_name,
- [
- "identifier",
- weaviate.LinkTo("metadata", self.index_name + "_metadata", list(query_metadata_keys)),
- ],
- )
- .with_where(weaviate_where_clause)
- .with_additional(["id"])
- .with_limit(limit or self.batch_size),
- offset,
- )
-
- fetched_results = results["data"]["Get"].get(self.index_name, [])
- if not fetched_results:
- break
-
- for result in fetched_results:
- existing_ids.append(result["identifier"])
- metadatas.append(result["metadata"][0])
- cursor = result["_additional"]["id"]
- offset += 1
-
- if limit is not None and len(existing_ids) >= limit:
- break
-
- return {"ids": existing_ids, "metadatas": metadatas}
-
- def add(self, documents: list[str], metadatas: list[object], ids: list[str], **kwargs: Optional[dict[str, any]]):
- """add data in vector database
- :param documents: list of texts to add
- :type documents: list[str]
- :param metadatas: list of metadata associated with docs
- :type metadatas: list[object]
- :param ids: ids of docs
- :type ids: list[str]
- """
- embeddings = self.embedder.embedding_fn(documents)
- self.client.batch.configure(batch_size=self.batch_size, timeout_retries=3) # Configure batch
- with self.client.batch as batch: # Initialize a batch process
- for id, text, metadata, embedding in zip(ids, documents, metadatas, embeddings):
- doc = {"identifier": id, "text": text}
- updated_metadata = {"text": text}
- if metadata is not None:
- updated_metadata.update(**metadata)
-
- obj_uuid = batch.add_data_object(
- data_object=copy.deepcopy(doc), class_name=self.index_name, vector=embedding
- )
- metadata_uuid = batch.add_data_object(
- data_object=copy.deepcopy(updated_metadata),
- class_name=self.index_name + "_metadata",
- vector=embedding,
- )
- batch.add_reference(
- obj_uuid, self.index_name, "metadata", metadata_uuid, self.index_name + "_metadata", **kwargs
- )
-
- def query(
- self, input_query: str, n_results: int, where: dict[str, any], citations: bool = False
- ) -> Union[list[tuple[str, dict]], list[str]]:
- """
- query contents from vector database based on vector similarity
- :param input_query: query string
- :type input_query: str
- :param n_results: no of similar documents to fetch from database
- :type n_results: int
- :param where: Optional. to filter data
- :type where: dict[str, any]
- :param citations: we use citations boolean param to return context along with the answer.
- :type citations: bool, default is False.
- :return: The content of the document that matched your query,
- along with url of the source and doc_id (if citations flag is true)
- :rtype: list[str], if citations=False, otherwise list[tuple[str, str, str]]
- """
- query_vector = self.embedder.embedding_fn([input_query])[0]
- keys = set(where.keys() if where is not None else set())
- data_fields = ["text"]
- query_metadata_keys = self.metadata_keys.union(keys)
- if citations:
- data_fields.append(weaviate.LinkTo("metadata", self.index_name + "_metadata", list(query_metadata_keys)))
-
- if len(keys) > 0:
- weaviate_where_operands = []
- for key in keys:
- weaviate_where_operands.append(
- {
- "path": ["metadata", self.index_name + "_metadata", key],
- "operator": "Equal",
- "valueText": where.get(key),
- }
- )
- if len(weaviate_where_operands) == 1:
- weaviate_where_clause = weaviate_where_operands[0]
- else:
- weaviate_where_clause = {"operator": "And", "operands": weaviate_where_operands}
-
- results = (
- self.client.query.get(self.index_name, data_fields)
- .with_where(weaviate_where_clause)
- .with_near_vector({"vector": query_vector})
- .with_limit(n_results)
- .with_additional(["distance"])
- .do()
- )
- else:
- results = (
- self.client.query.get(self.index_name, data_fields)
- .with_near_vector({"vector": query_vector})
- .with_limit(n_results)
- .with_additional(["distance"])
- .do()
- )
-
- if results["data"]["Get"].get(self.index_name) is None:
- return []
-
- docs = results["data"]["Get"].get(self.index_name)
- contexts = []
- for doc in docs:
- context = doc["text"]
- if citations:
- metadata = doc["metadata"][0]
- score = doc["_additional"]["distance"]
- metadata["score"] = score
- contexts.append((context, metadata))
- else:
- contexts.append(context)
- return contexts
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
- :param name: Name of the collection.
- :type name: str
- """
- if not isinstance(name, str):
- raise TypeError("Collection name must be a string")
- self.config.collection_name = name
-
- def count(self) -> int:
- """
- Count number of documents/chunks embedded in the database.
- :return: number of documents
- :rtype: int
- """
- data = self.client.query.aggregate(self.index_name).with_meta_count().do()
- return data["data"]["Aggregate"].get(self.index_name)[0]["meta"]["count"]
-
- def _get_or_create_db(self):
- """Called during initialization"""
- return self.client
-
- def reset(self):
- """
- Resets the database. Deletes all embeddings irreversibly.
- """
- # Delete all data from the database
- self.client.batch.delete_objects(
- self.index_name, where={"path": ["identifier"], "operator": "Like", "valueText": ".*"}
- )
-
- # Weaviate internally by default capitalizes the class name
- def _get_index_name(self) -> str:
- """Get the Weaviate index for a collection
- :return: Weaviate index
- :rtype: str
- """
- return f"{self.config.collection_name}_{self.embedder.vector_dimension}".capitalize().replace("-", "_")
-
- @staticmethod
- def _query_with_offset(query, offset):
- if offset:
- query.with_offset(offset)
- results = query.do()
- return results
-
- def _generate_query(self, where: dict):
- weaviate_where_operands = []
- for key, value in where.items():
- weaviate_where_operands.append(
- {
- "path": ["metadata", self.index_name + "_metadata", key],
- "operator": "Equal",
- "valueText": value,
- }
- )
-
- if len(weaviate_where_operands) == 1:
- weaviate_where_clause = weaviate_where_operands[0]
- else:
- weaviate_where_clause = {"operator": "And", "operands": weaviate_where_operands}
-
- return weaviate_where_clause
-
- def delete(self, where: dict):
- """Delete from database.
- :param where: to filter data
- :type where: dict[str, any]
- """
- query = self._generate_query(where)
- self.client.batch.delete_objects(self.index_name, where=query)
diff --git a/embedchain/embedchain/vectordb/zilliz.py b/embedchain/embedchain/vectordb/zilliz.py
deleted file mode 100644
index ca5544733..000000000
--- a/embedchain/embedchain/vectordb/zilliz.py
+++ /dev/null
@@ -1,252 +0,0 @@
-import logging
-from typing import Any, Optional, Union
-
-from embedchain.config import ZillizDBConfig
-from embedchain.helpers.json_serializable import register_deserializable
-from embedchain.vectordb.base import BaseVectorDB
-
-try:
- from pymilvus import (
- Collection,
- CollectionSchema,
- DataType,
- FieldSchema,
- MilvusClient,
- connections,
- utility,
- )
-except ImportError:
- raise ImportError(
- "Zilliz requires extra dependencies. Install with `pip install --upgrade embedchain[milvus]`"
- ) from None
-
-logger = logging.getLogger(__name__)
-
-
-@register_deserializable
-class ZillizVectorDB(BaseVectorDB):
- """Base class for vector database."""
-
- def __init__(self, config: ZillizDBConfig = None):
- """Initialize the database. Save the config and client as an attribute.
-
- :param config: Database configuration class instance.
- :type config: ZillizDBConfig
- """
-
- if config is None:
- self.config = ZillizDBConfig()
- else:
- self.config = config
-
- self.client = MilvusClient(
- uri=self.config.uri,
- token=self.config.token,
- )
-
- self.connection = connections.connect(
- uri=self.config.uri,
- token=self.config.token,
- )
-
- super().__init__(config=self.config)
-
- def _initialize(self):
- """
- This method is needed because `embedder` attribute needs to be set externally before it can be initialized.
-
- So it's can't be done in __init__ in one step.
- """
- self._get_or_create_collection(self.config.collection_name)
-
- def _get_or_create_db(self):
- """Get or create the database."""
- return self.client
-
- def _get_or_create_collection(self, name):
- """
- Get or create a named collection.
-
- :param name: Name of the collection
- :type name: str
- """
- if utility.has_collection(name):
- logger.info(f"[ZillizDB]: found an existing collection {name}, make sure the auto-id is disabled.")
- self.collection = Collection(name)
- else:
- fields = [
- FieldSchema(name="id", dtype=DataType.VARCHAR, is_primary=True, max_length=512),
- FieldSchema(name="text", dtype=DataType.VARCHAR, max_length=2048),
- FieldSchema(name="embeddings", dtype=DataType.FLOAT_VECTOR, dim=self.embedder.vector_dimension),
- FieldSchema(name="metadata", dtype=DataType.JSON),
- ]
-
- schema = CollectionSchema(fields, enable_dynamic_field=True)
- self.collection = Collection(name=name, schema=schema)
-
- index = {
- "index_type": "AUTOINDEX",
- "metric_type": self.config.metric_type,
- }
- self.collection.create_index("embeddings", index)
- return self.collection
-
- def get(self, ids: Optional[list[str]] = None, where: Optional[dict[str, any]] = None, limit: Optional[int] = None):
- """
- Get existing doc ids present in vector database
-
- :param ids: list of doc ids to check for existence
- :type ids: list[str]
- :param where: Optional. to filter data
- :type where: dict[str, Any]
- :param limit: Optional. maximum number of documents
- :type limit: Optional[int]
- :return: Existing documents.
- :rtype: Set[str]
- """
- data_ids = []
- metadatas = []
- if self.collection.num_entities == 0 or self.collection.is_empty:
- return {"ids": data_ids, "metadatas": metadatas}
-
- filter_ = ""
- if ids:
- filter_ = f'id in "{ids}"'
-
- if where:
- if filter_:
- filter_ += " and "
- filter_ = f"{self._generate_zilliz_filter(where)}"
-
- results = self.client.query(collection_name=self.config.collection_name, filter=filter_, output_fields=["*"])
- for res in results:
- data_ids.append(res.get("id"))
- metadatas.append(res.get("metadata", {}))
-
- return {"ids": data_ids, "metadatas": metadatas}
-
- def add(
- self,
- documents: list[str],
- metadatas: list[object],
- ids: list[str],
- **kwargs: Optional[dict[str, any]],
- ):
- """Add to database"""
- embeddings = self.embedder.embedding_fn(documents)
-
- for id, doc, metadata, embedding in zip(ids, documents, metadatas, embeddings):
- data = {"id": id, "text": doc, "embeddings": embedding, "metadata": metadata}
- self.client.insert(collection_name=self.config.collection_name, data=data, **kwargs)
-
- self.collection.load()
- self.collection.flush()
- self.client.flush(self.config.collection_name)
-
- def query(
- self,
- input_query: str,
- n_results: int,
- where: dict[str, Any],
- citations: bool = False,
- **kwargs: Optional[dict[str, Any]],
- ) -> Union[list[tuple[str, dict]], list[str]]:
- """
- Query contents from vector database based on vector similarity
-
- :param input_query: query string
- :type input_query: str
- :param n_results: no of similar documents to fetch from database
- :type n_results: int
- :param where: to filter data
- :type where: dict[str, Any]
- :raises InvalidDimensionException: Dimensions do not match.
- :param citations: we use citations boolean param to return context along with the answer.
- :type citations: bool, default is False.
- :return: The content of the document that matched your query,
- along with url of the source and doc_id (if citations flag is true)
- :rtype: list[str], if citations=False, otherwise list[tuple[str, str, str]]
- """
-
- if self.collection.is_empty:
- return []
-
- output_fields = ["*"]
- input_query_vector = self.embedder.embedding_fn([input_query])
- query_vector = input_query_vector[0]
-
- query_filter = self._generate_zilliz_filter(where)
- query_result = self.client.search(
- collection_name=self.config.collection_name,
- data=[query_vector],
- filter=query_filter,
- limit=n_results,
- output_fields=output_fields,
- **kwargs,
- )
- query_result = query_result[0]
- contexts = []
- for query in query_result:
- data = query["entity"]
- score = query["distance"]
- context = data["text"]
-
- if citations:
- metadata = data.get("metadata", {})
- metadata["score"] = score
- contexts.append(tuple((context, metadata)))
- else:
- contexts.append(context)
- return contexts
-
- def count(self) -> int:
- """
- Count number of documents/chunks embedded in the database.
-
- :return: number of documents
- :rtype: int
- """
- return self.collection.num_entities
-
- def reset(self, collection_names: list[str] = None):
- """
- Resets the database. Deletes all embeddings irreversibly.
- """
- if self.config.collection_name:
- if collection_names:
- for collection_name in collection_names:
- if collection_name in self.client.list_collections():
- self.client.drop_collection(collection_name=collection_name)
- else:
- self.client.drop_collection(collection_name=self.config.collection_name)
- self._get_or_create_collection(self.config.collection_name)
-
- def set_collection_name(self, name: str):
- """
- Set the name of the collection. A collection is an isolated space for vectors.
-
- :param name: Name of the collection.
- :type name: str
- """
- if not isinstance(name, str):
- raise TypeError("Collection name must be a string")
- self.config.collection_name = name
-
- def _generate_zilliz_filter(self, where: dict[str, str]):
- operands = []
- for key, value in where.items():
- operands.append(f'(metadata["{key}"] == "{value}")')
- return " and ".join(operands)
-
- def delete(self, where: dict[str, Any]):
- """
- Delete the embeddings from DB. Zilliz only support deleting with keys.
-
-
- :param keys: Primary keys of the table entries to delete.
- :type keys: Union[list, str, int]
- """
- data = self.get(where=where)
- keys = data.get("ids", [])
- if keys:
- self.client.delete(collection_name=self.config.collection_name, pks=keys)
diff --git a/embedchain/examples/api_server/.dockerignore b/embedchain/examples/api_server/.dockerignore
deleted file mode 100644
index 1dce42e87..000000000
--- a/embedchain/examples/api_server/.dockerignore
+++ /dev/null
@@ -1,8 +0,0 @@
-__pycache__/
-database
-db
-pyenv
-venv
-.env
-.git
-trash_files/
diff --git a/embedchain/examples/api_server/.gitignore b/embedchain/examples/api_server/.gitignore
deleted file mode 100644
index 2227fe3e2..000000000
--- a/embedchain/examples/api_server/.gitignore
+++ /dev/null
@@ -1,8 +0,0 @@
-__pycache__
-db
-database
-pyenv
-venv
-.env
-trash_files/
-.ideas.md
\ No newline at end of file
diff --git a/embedchain/examples/api_server/Dockerfile b/embedchain/examples/api_server/Dockerfile
deleted file mode 100644
index 6d5a7be87..000000000
--- a/embedchain/examples/api_server/Dockerfile
+++ /dev/null
@@ -1,16 +0,0 @@
-FROM python:3.11 AS backend
-
-WORKDIR /usr/src/api
-COPY requirements.txt .
-RUN pip install -r requirements.txt
-
-COPY . .
-
-EXPOSE 5000
-
-ENV FLASK_APP=api_server.py
-
-ENV FLASK_RUN_EXTRA_FILES=/usr/src/api/*
-ENV FLASK_ENV=development
-
-CMD ["flask", "run", "--host=0.0.0.0", "--reload"]
diff --git a/embedchain/examples/api_server/README.md b/embedchain/examples/api_server/README.md
deleted file mode 100644
index 1d9fa612b..000000000
--- a/embedchain/examples/api_server/README.md
+++ /dev/null
@@ -1,3 +0,0 @@
-# API Server
-
-This is a docker template to create your own API Server using the embedchain package. To know more about the API Server and how to use it, go [here](https://docs.embedchain.ai/examples/api_server).
\ No newline at end of file
diff --git a/embedchain/examples/api_server/api_server.py b/embedchain/examples/api_server/api_server.py
deleted file mode 100644
index f8d4d4d1a..000000000
--- a/embedchain/examples/api_server/api_server.py
+++ /dev/null
@@ -1,57 +0,0 @@
-import logging
-
-from flask import Flask, jsonify, request
-
-from embedchain import App
-
-app = Flask(__name__)
-
-
-logger = logging.getLogger(__name__)
-
-
-@app.route("/add", methods=["POST"])
-def add():
- data = request.get_json()
- data_type = data.get("data_type")
- url_or_text = data.get("url_or_text")
- if data_type and url_or_text:
- try:
- App().add(url_or_text, data_type=data_type)
- return jsonify({"data": f"Added {data_type}: {url_or_text}"}), 200
- except Exception:
- logger.exception(f"Failed to add {data_type=}: {url_or_text=}")
- return jsonify({"error": f"Failed to add {data_type}: {url_or_text}"}), 500
- return jsonify({"error": "Invalid request. Please provide 'data_type' and 'url_or_text' in JSON format."}), 400
-
-
-@app.route("/query", methods=["POST"])
-def query():
- data = request.get_json()
- question = data.get("question")
- if question:
- try:
- response = App().query(question)
- return jsonify({"data": response}), 200
- except Exception:
- logger.exception(f"Failed to query {question=}")
- return jsonify({"error": "An error occurred. Please try again!"}), 500
- return jsonify({"error": "Invalid request. Please provide 'question' in JSON format."}), 400
-
-
-@app.route("/chat", methods=["POST"])
-def chat():
- data = request.get_json()
- question = data.get("question")
- if question:
- try:
- response = App().chat(question)
- return jsonify({"data": response}), 200
- except Exception:
- logger.exception(f"Failed to chat {question=}")
- return jsonify({"error": "An error occurred. Please try again!"}), 500
- return jsonify({"error": "Invalid request. Please provide 'question' in JSON format."}), 400
-
-
-if __name__ == "__main__":
- app.run(host="0.0.0.0", port=5000, debug=False)
diff --git a/embedchain/examples/api_server/docker-compose.yml b/embedchain/examples/api_server/docker-compose.yml
deleted file mode 100644
index 8fa3fc817..000000000
--- a/embedchain/examples/api_server/docker-compose.yml
+++ /dev/null
@@ -1,15 +0,0 @@
-version: "3.9"
-
-services:
- backend:
- container_name: embedchain_api
- restart: unless-stopped
- build:
- context: .
- dockerfile: Dockerfile
- env_file:
- - variables.env
- ports:
- - "5000:5000"
- volumes:
- - .:/usr/src/api
diff --git a/embedchain/examples/api_server/requirements.txt b/embedchain/examples/api_server/requirements.txt
deleted file mode 100644
index 39e066ada..000000000
--- a/embedchain/examples/api_server/requirements.txt
+++ /dev/null
@@ -1,12 +0,0 @@
-flask==2.3.2
-youtube-transcript-api==0.6.1
-pytube==15.0.0
-beautifulsoup4==4.12.3
-slack-sdk==3.21.3
-huggingface_hub==0.23.0
-gitpython==3.1.38
-yt_dlp==2023.11.14
-PyGithub==1.59.1
-feedparser==6.0.10
-newspaper3k==0.2.8
-listparser==0.19
\ No newline at end of file
diff --git a/embedchain/examples/api_server/variables.env b/embedchain/examples/api_server/variables.env
deleted file mode 100644
index da6725993..000000000
--- a/embedchain/examples/api_server/variables.env
+++ /dev/null
@@ -1 +0,0 @@
-OPENAI_API_KEY=""
\ No newline at end of file
diff --git a/embedchain/examples/chainlit/.gitignore b/embedchain/examples/chainlit/.gitignore
deleted file mode 100644
index 2121b2589..000000000
--- a/embedchain/examples/chainlit/.gitignore
+++ /dev/null
@@ -1 +0,0 @@
-.chainlit
diff --git a/embedchain/examples/chainlit/README.md b/embedchain/examples/chainlit/README.md
deleted file mode 100644
index d54e69656..000000000
--- a/embedchain/examples/chainlit/README.md
+++ /dev/null
@@ -1,17 +0,0 @@
-## Chainlit + Embedchain Demo
-
-In this example, we will learn how to use Chainlit and Embedchain together
-
-## Setup
-
-First, install the required packages:
-
-```bash
-pip install -r requirements.txt
-```
-
-## Run the app locally,
-
-```
-chainlit run app.py
-```
diff --git a/embedchain/examples/chainlit/app.py b/embedchain/examples/chainlit/app.py
deleted file mode 100644
index f2de4b0bd..000000000
--- a/embedchain/examples/chainlit/app.py
+++ /dev/null
@@ -1,35 +0,0 @@
-import os
-
-import chainlit as cl
-
-from embedchain import App
-
-os.environ["OPENAI_API_KEY"] = "sk-xxx"
-
-
-@cl.on_chat_start
-async def on_chat_start():
- app = App.from_config(
- config={
- "app": {"config": {"name": "chainlit-app"}},
- "llm": {
- "config": {
- "stream": True,
- }
- },
- }
- )
- # import your data here
- app.add("https://www.forbes.com/profile/elon-musk/")
- app.collect_metrics = False
- cl.user_session.set("app", app)
-
-
-@cl.on_message
-async def on_message(message: cl.Message):
- app = cl.user_session.get("app")
- msg = cl.Message(content="")
- for chunk in await cl.make_async(app.chat)(message.content):
- await msg.stream_token(chunk)
-
- await msg.send()
diff --git a/embedchain/examples/chainlit/chainlit.md b/embedchain/examples/chainlit/chainlit.md
deleted file mode 100644
index d3de410e4..000000000
--- a/embedchain/examples/chainlit/chainlit.md
+++ /dev/null
@@ -1,15 +0,0 @@
-# Welcome to Embedchain! 🚀
-
-Hello! 👋 Excited to see you join us. With Embedchain and Chainlit, create ChatGPT like apps effortlessly.
-
-## Quick Start 🌟
-
-- **Embedchain Docs:** Get started with our comprehensive [Embedchain Documentation](https://docs.embedchain.ai/) 📚
-- **Discord Community:** Join our discord [Embedchain Discord](https://discord.gg/CUU9FPhRNt) to ask questions, share your projects, and connect with other developers! 💬
-- **UI Guide**: Master Chainlit with [Chainlit Documentation](https://docs.chainlit.io/) ⛓️
-
-Happy building with Embedchain! 🎉
-
-## Customize welcome screen
-
-Edit chainlit.md in your project root to change this welcome message.
diff --git a/embedchain/examples/chainlit/requirements.txt b/embedchain/examples/chainlit/requirements.txt
deleted file mode 100644
index 1604c5f9d..000000000
--- a/embedchain/examples/chainlit/requirements.txt
+++ /dev/null
@@ -1,2 +0,0 @@
-chainlit==0.7.700
-embedchain==0.1.57
diff --git a/embedchain/examples/chat-pdf/README.md b/embedchain/examples/chat-pdf/README.md
deleted file mode 100644
index 2a09c8bfc..000000000
--- a/embedchain/examples/chat-pdf/README.md
+++ /dev/null
@@ -1,32 +0,0 @@
-# Embedchain Chat with PDF App
-
-You can easily create and deploy your own `Chat-with-PDF` App using Embedchain.
-
-Checkout the live demo we created for [chat with PDF](https://embedchain.ai/demo/chat-pdf).
-
-Here are few simple steps for you to create and deploy your app:
-
-1. Fork the embedchain repo from [Github](https://github.com/embedchain/embedchain).
-
-If you run into problems with forking, please refer to [github docs](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/fork-a-repo) for forking a repo.
-
-2. Navigate to `chat-pdf` example app from your forked repo:
-
-```bash
-cd /examples/chat-pdf
-```
-
-3. Run your app in development environment with simple commands
-
-```bash
-pip install -r requirements.txt
-ec dev
-```
-
-Feel free to improve our simple `chat-pdf` streamlit app and create pull request to showcase your app [here](https://docs.embedchain.ai/examples/showcase)
-
-4. You can easily deploy your app using Streamlit interface
-
-Connect your Github account with Streamlit and refer this [guide](https://docs.streamlit.io/streamlit-community-cloud/deploy-your-app) to deploy your app.
-
-You can also use the deploy button from your streamlit website you see when running `ec dev` command.
diff --git a/embedchain/examples/chat-pdf/app.py b/embedchain/examples/chat-pdf/app.py
deleted file mode 100644
index 73800605d..000000000
--- a/embedchain/examples/chat-pdf/app.py
+++ /dev/null
@@ -1,160 +0,0 @@
-import os
-import queue
-import re
-import tempfile
-import threading
-
-import streamlit as st
-
-from embedchain import App
-from embedchain.config import BaseLlmConfig
-from embedchain.helpers.callbacks import StreamingStdOutCallbackHandlerYield, generate
-
-
-def embedchain_bot(db_path, api_key):
- return App.from_config(
- config={
- "llm": {
- "provider": "openai",
- "config": {
- "model": "gpt-4o-mini",
- "temperature": 0.5,
- "max_tokens": 1000,
- "top_p": 1,
- "stream": True,
- "api_key": api_key,
- },
- },
- "vectordb": {
- "provider": "chroma",
- "config": {"collection_name": "chat-pdf", "dir": db_path, "allow_reset": True},
- },
- "embedder": {"provider": "openai", "config": {"api_key": api_key}},
- "chunker": {"chunk_size": 2000, "chunk_overlap": 0, "length_function": "len"},
- }
- )
-
-
-def get_db_path():
- tmpdirname = tempfile.mkdtemp()
- return tmpdirname
-
-
-def get_ec_app(api_key):
- if "app" in st.session_state:
- print("Found app in session state")
- app = st.session_state.app
- else:
- print("Creating app")
- db_path = get_db_path()
- app = embedchain_bot(db_path, api_key)
- st.session_state.app = app
- return app
-
-
-with st.sidebar:
- openai_access_token = st.text_input("OpenAI API Key", key="api_key", type="password")
- "WE DO NOT STORE YOUR OPENAI KEY."
- "Just paste your OpenAI API key here and we'll use it to power the chatbot. [Get your OpenAI API key](https://platform.openai.com/api-keys)" # noqa: E501
-
- if st.session_state.api_key:
- app = get_ec_app(st.session_state.api_key)
-
- pdf_files = st.file_uploader("Upload your PDF files", accept_multiple_files=True, type="pdf")
- add_pdf_files = st.session_state.get("add_pdf_files", [])
- for pdf_file in pdf_files:
- file_name = pdf_file.name
- if file_name in add_pdf_files:
- continue
- try:
- if not st.session_state.api_key:
- st.error("Please enter your OpenAI API Key")
- st.stop()
- temp_file_name = None
- with tempfile.NamedTemporaryFile(mode="wb", delete=False, prefix=file_name, suffix=".pdf") as f:
- f.write(pdf_file.getvalue())
- temp_file_name = f.name
- if temp_file_name:
- st.markdown(f"Adding {file_name} to knowledge base...")
- app.add(temp_file_name, data_type="pdf_file")
- st.markdown("")
- add_pdf_files.append(file_name)
- os.remove(temp_file_name)
- st.session_state.messages.append({"role": "assistant", "content": f"Added {file_name} to knowledge base!"})
- except Exception as e:
- st.error(f"Error adding {file_name} to knowledge base: {e}")
- st.stop()
- st.session_state["add_pdf_files"] = add_pdf_files
-
-st.title("📄 Embedchain - Chat with PDF")
-styled_caption = '