Spaces:
Sleeping
A newer version of the Streamlit SDK is available: 1.61.1
title: Multi PDF Chatbot
emoji: π
colorFrom: blue
colorTo: blue
sdk: streamlit
sdk_version: 1.42.0
python_version: '3.11'
app_file: app.py
pinned: false
π Multi-PDF Agentic Chatbot
Ask questions across multiple PDFs at once. Get answers with exact source citations β document name + page number.
Live Demo
Try it here: https://huggingface.co/spaces/artistica-004/multi-pdf-chatbot
π― Objective
This project was built as part of Task 2A β Fix My Life with AI.
The goal was to identify a genuine real-world problem, design a practical AI solution for it, and demonstrate a measurable before vs after improvement.
π€ Step 1 β The Pain (Why This Was Built)
Every single study session I found myself with 4-5 PDF tabs open simultaneously. Finding one answer meant:
- Switching between tabs constantly
- Losing my train of thought every time I switched
- Re-reading the same sections because I forgot which PDF had what
- Using Ctrl+F but not knowing which PDF to search first
- Spending 30-45 minutes just navigating β before even starting to understand the content
This happened every single day.
I could not skip it β these were study materials and research documents I was required to understand deeply. I could not delegate it β the comprehension was mine to do.
Why existing tools failed:
| Tool | Why It Failed |
|---|---|
| Ctrl+F | One PDF at a time, exact keyword needed |
| Cannot access private PDFs | |
| ChatGPT | No grounding in actual documents, hallucinations |
| Adobe Acrobat | One file at a time, no semantic search |
The real problem was not volume. It was context reconstruction from scattered sources.
π Step 2 β Workflow Before This Tool
Old process (10 steps):
- Formulate the question in my head (~1 min)
- Open 3-5 PDFs in separate browser tabs (~2 min)
- Pick starting PDF based on gut feeling (~30 sec)
- Ctrl+F with a keyword (~1-2 min)
- Read surrounding context to check relevance (~3-5 min)
- If not found β switch tab, repeat Ctrl+F (~1-2 min)
- β Reconstruct mental context after tab switch (~5-10 min)
- Cross-reference information across PDFs (~5 min)
- Manually compile answer from fragments (~5 min)
- Re-verify source page before using answer (~2 min)
Total: 30-45 minutes per session Most mentally demanding step: Step 7 Reconstructing context after every tab switch β holding the original question, what was already read, and the current document all in working memory at once.
Flowchart (Before):
[Need an answer]
|
v
[Open 3-5 PDFs in separate tabs]
|
v
[Choose starting PDF by intuition]
|
v
[Ctrl+F keyword]
|
FOUND? ββNoββ> [Switch tab, try next PDF]
| |
Yes [Different keyword?]
| Yes / \ No
v /
[Read paragraph] <ββ [Mark as no info]
|
[Does it answer the question?]
|
Yes / No
/
[Note source] [Try synonym / different section]
|
[All PDFs checked?]
|
Yes / No
/
[Compile] [Loop back]
|
[Verify source]
|
[Done]
π§ Step 3 β AI Solution Design
System Thinking
Architecture chosen: Agentic Pipeline System
| Architecture | Why Not Chosen |
|---|---|
| Single Prompt | Cannot handle raw PDFs, token limit exceeded |
| Basic Pipeline | No decision making, no adaptive behavior |
| Agent System β | Decision making + multi-step execution + adaptive |
Why Agent System: The bot needs to make decisions β is this question simple or complex? Should it search PDFs or the web? Does the answer actually exist in the documents? These decisions require an agent, not just a pipeline.
Full Architecture:
User Question | v [Analyze Question β simple or complex?] | Complex ββββββββββββββ> [Break into sub-questions] | | Simple [Search each sub-question] | | [Search PDFs] [Combine all answers] | | [Check answer quality] <ββββββββββ | FOUND ββ> [Generate follow-up questions] ββ> [Final Answer] | NOT_FOUND ββ> ["Not available in documents"]
Data Layer
Inputs and Outputs:
| Stage | Input | Output |
|---|---|---|
| PDF Extraction (PyPDF2) | Raw PDF files | Text + source + page per page |
| Chunking | Page text | 500-char chunks with metadata |
| Embedding (all-MiniLM-L6-v2) | Chunk strings | 384-dim vectors |
| Vector Store (ChromaDB) | Vectors + metadata | Searchable index |
| Retrieval | Question + k=6 | Top 6 relevant chunks |
| Generation (Groq LLaMA 3.3 70B) | Question + context | Cited answer |
Why chunk_size=500, chunk_overlap=50:
- 500 chars = 3-5 sentences = one focused concept
- Overlap of 50 ensures sentences on boundaries are not lost
- Smaller chunks = more precise embeddings = better retrieval
Edge Cases
Edge Case 1: Scanned PDF (image-based)
- Current handling: PyPDF2 returns empty string, page is skipped silently
- Impact: User gets no answer with no explanation why
Edge Case 2: Answer spans chunk boundary
- Current handling: 50-char overlap partially mitigates this
- Impact: LLM may receive incomplete context and generate partial answer
Edge Case 3: Question not in any PDF
- Current handling: Agent checks answer quality, returns "not available in documents" clearly
- Impact without handling: LLM halluculates a confident-sounding wrong answer
Failure Simulation
Scenario: Groq API rate limit hit
- User uploads 4 PDFs and asks 10 rapid questions
- generate_answer() call raises RateLimitError
- User sees Python traceback β no friendly message
- Root cause: Free Groq tier has tokens-per-minute limit
- Fix: try/except around API call with friendly message and retry logic
Trade-offs
Trade-off 1: chunk_size=500 vs chunk_size=1000
- Chose 500 for precise embeddings per concept
- Sacrificed: multi-paragraph argument retrieval
- Worth it: most questions are factual, not analytical
Trade-off 2: Local embeddings vs OpenAI embeddings
- Chose all-MiniLM-L6-v2 (local, free)
- Sacrificed: higher quality semantic matching
- Worth it: zero API cost, no data sent externally
β Step 4 β Proof of Concept
Before vs After
| Metric | Before (Manual) | After (AI Tool) |
|---|---|---|
| Time to find answer | 30-45 min | 8-15 seconds |
| Steps required | 10 steps with loops | 3 steps |
| Accuracy | Keyword dependent | Semantically grounded |
| Context switching | 15-20 tab switches | Zero |
| Source traceability | Manual memory | Auto cited |
| Cross-doc synthesis | Manual notes | Automatic |
Actual Prompt Sent to Groq:
You are a helpful and concise assistant. The user has uploaded these PDF documents:
document1.pdf document2.pdf
Relevant excerpts from the documents: --- Chunk 1 from: document1.pdf, Page 3 --- [chunk text] Question: [user question] Instructions:
Answer in maximum 4-5 lines only Be direct and simple No repetition Mention source: (Source: filename.pdf, Page 3) If not in documents say so clearly
Sample Q&A:
Example 1:
- Question: "What is the role of AI in the playbook?"
- Answer: "AI is used to structure thinking, validate decisions and accelerate learning β not to replace thinking entirely. (Source: Intern Operating System V2.pdf, Page 3)"
Example 2:
- Question: "What is the salary structure?"
- Answer: "This information is not available in the uploaded documents."
π Final Reflection
Q1: What is the weakest part? The chunk boundary problem. When an answer spans multiple paragraphs, k=6 may not retrieve all relevant chunks. Also scanned PDFs are silently skipped with no user warning.
Q2: What single failure would break it completely? Groq API going down. Everything else runs locally but without Groq, answer generation completely fails. No fallback model exists currently.
Q3: If AI was removed, what would still be valuable? Three things:
- Multi-PDF aggregation in one interface
- Chunk metadata system with source + page citations
- Semantic similarity search β still better than Ctrl+F
ποΈ Architecture
app.py β Streamlit UI + chat interface agent.py β Agentic decision making loop rag_engine.py β RAG pipeline + vector store .env β API keys requirements.txt β Dependencies
βοΈ Agentic Features
- β Decision Making β simple vs complex question handling
- β Multi-Step Execution β complex questions broken into parts
- β Workflow Orchestration β analyze β plan β execute β combine
- β Adaptive Behavior β greetings handled separately
- β Auto PDF Summarization β summary shown on upload
- β Gap Detection β clearly states when answer not found
- β Follow-up Suggestions β 3 related questions after every answer
π οΈ Tech Stack
| Component | Technology |
|---|---|
| Frontend | Streamlit |
| LLM | Groq LLaMA 3.3 70B |
| Embeddings | HuggingFace all-MiniLM-L6-v2 |
| Vector Store | ChromaDB |
| PDF Extraction | PyPDF2 |
| Chunking | LangChain RecursiveCharacterTextSplitter |
| Agent Logic | Custom Python |
π Run Locally
git clone https://github.com/artistica-004/multi-pdf-chatbot
cd multi-pdf-chatbot
pip install -r requirements.txt
Create .env file:
GROQ_API_KEY=your_groq_api_key_here
Run:
streamlit run app.py
π¦ Requirements
streamlit langchain langchain-community langchain-text-splitters PyPDF2 chromadb sentence-transformers groq python-dotenv
π Links
- Live Demo: https://huggingface.co/spaces/artistica-004/multi-pdf-chatbot
- GitHub: https://github.com/artistica-004/multi-pdf-chatbot