How to Build a RAG Chatbot with LangChain and ChromaDB
Normal chatbots like ChatGPT, Claude etc. only know what they were trained on. Ask them about your own PDF or notes, and they will guess.
A RAG chatbot fixes this. It reads your documents first, then answers from them.
In this guide, you will build one from scratch using Python, LangChain, and ChromaDB.
What is RAG?
RAG stands for Retrieval-Augmented Generation.
It works in 3 simple steps:
- Retrieve: find the parts of your documents that match the question.
- Augment: add those parts to the prompt.
- Generate: the LLM writes the answer using that context.
So the model does not need to memorize your data. It just reads the right pages at the right time.
Why LangChain and ChromaDB?
- LangChain connects everything: loaders, splitters, models, and prompts.
- ChromaDB is a vector database. It stores your text as numbers (embeddings) and finds similar text fast.
- Both are free and open source, and easy to run on your laptop.
What You Need
- Python 3.10 or higher
- A free Groq API key (from console.groq.com)
- A PDF file to chat with
Step 1: Install the Packages
pip install langchain langchain-community langchain-chroma \
langchain-huggingface langchain-groq langchain-text-splitters \
sentence-transformers pypdf python-dotenv
Create a .env file:
GROQ_API_KEY=your_key_here
Step 2: Load Your Document
from dotenv import load_dotenv
from langchain_community.document_loaders import PyPDFLoader
load_dotenv()
loader = PyPDFLoader("docs/sample.pdf")
docs = loader.load()
print(f"Loaded {len(docs)} pages")
Step 3: Split the Text into Chunks
Long text does not fit well in a prompt. So we cut it into small pieces.
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=800,
chunk_overlap=100
)
chunks = splitter.split_documents(docs)
print(f"Created {len(chunks)} chunks")
chunk_overlap keeps a little text shared between chunks, so meaning is not lost at the edges.
Step 4: Create Embeddings and Store in ChromaDB
Embeddings turn text into numbers. Similar meaning gives similar numbers.
from langchain_huggingface import HuggingFaceEmbeddings
from langchain_chroma import Chroma
embeddings = HuggingFaceEmbeddings(
model_name="sentence-transformers/all-MiniLM-L6-v2"
)
vectorstore = Chroma.from_documents(
documents=chunks,
embedding=embeddings,
persist_directory="chroma_db"
)
The data is saved in the chroma_db folder. Next time, you can load it without rebuilding:
vectorstore = Chroma(
persist_directory="chroma_db",
embedding_function=embeddings
)
Step 5: Create a Retriever
retriever = vectorstore.as_retriever(search_kwargs={"k": 4})
k=4 means it returns the 4 most relevant chunks for every question.
Step 6: Connect the LLM and Build the Chain
from langchain_groq import ChatGroq
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough
llm = ChatGroq(model="llama-3.3-70b-versatile", temperature=0)
prompt = ChatPromptTemplate.from_template("""
Answer the question using only the context below.
If the answer is not in the context, say "I don't know."
Context:
{context}
Question:
{question}
""")
def format_docs(docs):
return "\n\n".join(d.page_content for d in docs)
rag_chain = (
{"context": retriever | format_docs, "question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
Model names on Groq change from time to time. If this one gives an error, check the current list in the Groq docs.
Step 7: Ask Questions
while True:
q = input("You: ")
if q.lower() in ["exit", "quit"]:
break
print("Bot:", rag_chain.invoke(q))
Run the file and chat with your PDF. That is a working RAG chatbot.
Common Problems and Fixes
- Bot gives wrong answers: try a smaller chunk_size or a higher k.
- Bot says "I don't know" too much: your chunks may be too small. Increase the size a bit.
- Slow first run: the embedding model downloads once. After that it is fast.
- Duplicate results: you added the same documents twice. Delete the chroma_db folder and rebuild.
- Bot uses wrong chunks: score the retrieved chunks first and drop the weak ones. A System One model like Jev is built for this kind of fast scoring.
Conclusion
Finally You now know how RAG works and how to build it with LangChain and ChromaDB. Start with one PDF, get it working, then improve it step by step. Now next what more we can do in this are:
- Add a UI with Streamlit so anyone can use it.
- Add chat memory so the bot remembers earlier messages.
- Wrap it in an API using FastAPI.
- Support more file types like DOCX, TXT, and web pages.
FAQ
- Is ChromaDB free?
Yes. It is open source and runs locally.
- Can I use OpenAI instead of Groq?
Yes. Just swap ChatGroq for ChatOpenAI. The rest stays the same.
- Do I need a GPU?
No. all-MiniLM-L6-v2 runs fine on a normal CPU.
- What is the difference between RAG and fine-tuning?
RAG gives the model fresh data at question time. Fine-tuning changes the model itself. RAG is cheaper and easier to update.
Author
Tech Enthusiast