A lightweight Retrieval-Augmented Generation (RAG) application for company documents. The project lets users upload text-based documents, split them into searchable chunks, store metadata in MySQL, create a FAISS vector index, and ask questions using a local LLM through Ollama.
This project combines:
- FastAPI backend for document upload and chat APIs
- Streamlit frontend for a simple web UI
- MySQL for document metadata and chunk storage
- FAISS for vector search
- Hugging Face embeddings for semantic retrieval
- Ollama + Llama 3.1 for answer generation
The app is designed around a per-user document workflow, so each user only sees documents and chunks associated with their own user ID.
company_rag/
├── app/
│ ├── database/
│ │ ├── connection.py
│ │ ├── models.py
│ │ └── repository.py
│ ├── generation/
│ │ └── llm.py
│ ├── ingestion/
│ │ ├── cleaner.py
│ │ ├── chunker.py
│ │ ├── embedder.py
│ │ ├── loader.py
│ │ └── __init__.py
│ ├── rag/
│ │ └── pipeline.py
│ ├── retrieval/
│ │ ├── retriever.py
│ │ └── vector_store.py
│ ├── config.py
│ └── main.py
├── frontend/
│ └── app.py
├── data/
│ ├── faiss/
│ └── uploads/
├── .env
├── .gitignore
├── check_env.py
├── create_tables.py
├── requirements.txt
├── README.md
└── venv/
- Upload
.txt,.pdf, and.docxfiles - Clean and chunk document text with page tracking
- Store documents and chunk metadata in MySQL
- Build and update a FAISS vector store
- Search relevant chunks using semantic similarity
- Generate answers grounded in retrieved document context
- Expose REST API and simple Streamlit UI
- Python
- FastAPI
- Streamlit
- SQLAlchemy
- PyMySQL
- FAISS
- LangChain
- Hugging Face Embeddings
- Ollama
Before running the app, make sure you have:
- Python installed
- MySQL server running
- Ollama installed and running locally
- The Llama 3.1 model downloaded in Ollama
Install the model with:
ollama pull llama3.1Create a .env file in the project root with your MySQL connection details:
DB_HOST=localhost
DB_PORT=3306
DB_USER=root
DB_PASSWORD=your_password
DB_NAME=company_ragThe app reads these values from .env in the database connection layer.
- Create and activate a virtual environment:
python -m venv venvOn Windows:
venv\Scripts\activateOn macOS/Linux:
source venv/bin/activate- Install dependencies:
pip install -r requirements.txt- Create the database tables:
python create_tables.py- Optional environment check:
python check_env.pyuvicorn app.main:app --reloadThe API runs by default at:
- http://127.0.0.1:8000
- Swagger docs: http://127.0.0.1:8000/docs
streamlit run frontend/app.pyThe user interface is typically available at:
GET /health— checks API statusGET /— basic welcome message
POST /documents/upload— uploads a document and ingests itGET /documents— lists all documents for a userGET /documents/{document_id}— fetches one document by IDDELETE /documents/{document_id}— deletes a document
POST /chat— asks a question using the user-specific vector store
Example request body for chat:
{
"user_id": 1,
"question": "What is the company's leave policy?"
}- Uploaded files are saved under
data/uploads/ - FAISS indexes are saved under
data/faiss/ - MySQL stores user, document, and chunk metadata
- This project currently uses a local Ollama LLM and does not include authentication.
- The vector store is user-aware when retrieving results, but the application assumes the caller passes a valid
user_id. - Document ingestion may take time on first run because the embedding model and LLM model are downloaded from external sources.
- Start MySQL and Ollama.
- Run the FastAPI backend.
- Run the Streamlit frontend.
- Upload a document for a specific user ID.
- Ask a question in the UI or via the API.
- Review the answer and source file references.
This project is provided as a starter application for company document Q&A and can be adapted for your own use case.