Creating a local scale RAG bot for understanding flow
A local Retrieval-Augmented Generation (RAG) system that lets you ask natural language questions about a PDF document and get answers grounded in its actual content.
- Extracts text from a PDF using
pypdf - Splits the text into overlapping chunks
- Converts each chunk into an embedding using
sentence-transformers(all-MiniLM-L6-v2) - Stores the embeddings in a local ChromaDB vector database
- On a query: embeds the question, retrieves the most relevant chunks via similarity search, and sends them as context to Groq's LLM API to generate an answer
- Python
- pypdf — PDF text extraction
- sentence-transformers — local embeddings (CPU-friendly, no GPU required)
- ChromaDB — local vector storage
- Groq API — free-tier LLM for answer generation
- Clone the repo
- Create a virtual environment:
python -m venv venv - Activate it and install dependencies:
pip install -r requirements.txt - Add a
.envfile withGROQ_API_KEY=your_key_here - Drop a PDF into the
data/folder - Run:
python main.py
This project gave me a hands-on understanding of how RAG pipelines actually work end to end — from document parsing and chunking to embeddings, vector search, and grounded LLM responses. Building it manually (instead of relying on a framework) helped me understand why each step exists, not just how to call it.
Next steps: improving chunking strategy, testing retrieval across multiple documents, and exploring reranking to handle cases where relevant information is split across chunks.