Project Vision
Create a Streamlit-powered "Accelerated Book Copilot" that ingests books or long-form documents, generates stacked summaries, enables adjustable-speed listening, answers questions with grounded citations, and enriches context via focused web crawling.
User Journeys
- Rapid Onboarding: Upload a book or paste a URL, receive quick, detailed, and action-focused summaries in minutes.
- Listening Mode: Play text-to-speech audio with adjustable pace and language options.
- Knowledge Dive: Ask free-form questions, get cited answers, and review referenced passages.
- Context Expansion: Trigger a crawl for related resources (reviews, interviews, scholarly notes) and fold distilled insights into the knowledge base.
Core Features
- PDF ingestion with reliable text extraction and structure retention.
- Chunking and embedding using SentenceTransformer or
text-embedding-3-smallfor retrieval-ready vectors. - Summaries: tiered (flash, detailed, character-focused) powered by Azure/OpenAI prompts.
- Q&A agent built on LangChain with streaming answers and citation highlighting.
- Listening: FOMO TTS modules with pitch/speed controls and downloadable clips.
- Crawl mode leveraging Newspaper3k/Requests + safety filters to augment knowledge base.
- Session management with Streamlit state and optional persistence (SQLite/JSON) for revisit.
Architecture Outline
- Ingestion Layer
- PDF loader pipeline with robust parsing and structure detection.
- Background task queue (async workers) to preprocess large uploads.
- Processing Layer
- Text splitter (RecursiveCharacter) feeding embeddings stored in FAISS or Pinecone.
- Summarization orchestrator reusing FOMO prompt stacks.
- Q&A chain: RetrievalQA + chat memory + citation packaging.
- Experience Layer
- Streamlit tabs for Read, Summaries, Listen, Q&A, Crawl Insights.
- Interactive timeline/outline with Altair visualization for pacing metrics.
- Audio controls and transcript alignment.
Tech Stack
- Streamlit UI (existing scaffold).
- LangChain for orchestration and tools.
- Azure/OpenAI LLM + embeddings; fallback to OpenAI endpoints.
- Pinecone or FAISS vector store depending on hackathon infra.
- SentenceTransformer local model for offline embeddings if needed.
- TTS via FOMO
load_ttsor Azure Speech (optional). - Newspaper3k, BeautifulSoup, or HuggingFace datasets for crawl augmentation.
- Storage: lightweight JSON/SQLite, with hooks for cloud object storage if available.
Milestones
- Day 1 – Foundations
- Wireframe Streamlit tabs and session state.
- Implement PDF ingestion -> chunk -> embed pipeline with local FAISS.
- Day 2 – Intelligence Layer
- Integrate summarization tiers and Q&A citations.
- Add TTS playback using existing FOMO audio utilities.
- Day 3 – Crawl & Polish
- Build targeted crawl + summarization flow.
- Add analytics (reading time estimator, compression ratios).
- Prep demo script and sample datasets.
Risks & Mitigations
- Large file latency: Use background processing and progressive loading.
- API rate limits: Cache embeddings/summaries, throttle crawls.
- Crawl reliability: Predefine trusted domains, handle failures gracefully.
- UX overload: Start with core tabs, keep optional features toggled off by default.
Next Steps
- Confirm target vector store (FAISS vs Pinecone) and credential availability.
- Draft Streamlit layout skeleton and assign owners per milestone.
- Prepare sample books and reliability test cases (short story, technical manual, novel).