research-link-ai / docs /research /implementation-options.md
MHamdan's picture
Deploy Research-Link-AI (Docker Space, offline demo)
a753e74 verified
|
Raw
History Blame Contribute Delete
4.74 kB

Implementation Options Research

Summary

This document captures the MVP stack decisions for ResearchLink AI based on an analysis of available tools, libraries, and patterns as of 2026.


PDF Parsing Options

Tool Pros Cons Decision
PyMuPDF (fitz) Fast, reliable text + structure extraction, layout-aware C extension, requires libmupdf Selected
pypdf Pure Python, lightweight Poor structure extraction, no layout Backup only
pdfplumber Good table extraction Slow on large PDFs Future use for tables
Marker ML-based, high quality Heavy dependencies, GPU preferred Stage 2
Grobid Best scientific structure parsing Java service, complex setup Future self-hosted option

Recommendation: PyMuPDF for MVP. Grobid for Stage 2 if high-quality extraction is needed.


Metadata Lookup Options

Service API Free Tier Notes
Semantic Scholar REST, no key required Yes (rate limited) Selected
OpenAlex REST, no key required Yes Good coverage, fast
Crossref REST, polite pool Yes Best for DOIs
arXiv API Atom/REST Yes Best for arXiv papers
SSRN No public API N/A Manual only
OpenReview REST Yes Good for NeurIPS/ICLR

Recommendation: Semantic Scholar as primary, Crossref for DOI enrichment. Both are free and well-documented.


Citation Validation Options

Approach Quality Complexity
Regex + heuristics Low-medium Low
bibtexparser + heuristics Medium Low
Crossref/Semantic Scholar lookup High Medium
anystyle-parser High (Ruby) High

Recommendation: bibtexparser for BibTeX parsing + Crossref for DOI verification in Stage 2.


GitHub Repository Analysis Options

Approach Coverage Auth Required
GitHub REST API (unauthenticated) Public repos No (rate-limited)
GitHub REST API (token) Public + higher rate limit Yes
GitHub GraphQL API Rich queries Yes
GitPython (local clone) Full access No (local)

Recommendation: GitHub REST API unauthenticated for MVP. Support GITHUB_TOKEN env var for higher rate limits.


Multi-Agent Orchestration Options

Framework Description Decision
Custom sequential pipeline Simple, transparent, no overhead Selected for MVP
LangChain Agents Many integrations, complex Overkill for MVP
LlamaIndex Good for RAG, less for pipeline Stage 2 option
CrewAI Role-based agents Stage 2 option
AutoGen Conversational multi-agent Research tool option

Recommendation: Custom sequential pipeline for MVP. Migrate to CrewAI or custom async orchestration in Stage 2.


Markdown and Documentation Generation Options

Tool Use Case Decision
Jinja2 Template rendering Selected
Direct LLM generation Content writing Selected
MkDocs Documentation site Stage 2
Pandoc Format conversion Future

Recommendation: Jinja2 for structure templates. Direct LLM (Claude) calls for content generation.


Recommended MVP Stack

Python 3.11+
typer β€” CLI
pydantic v2 β€” schemas and validation
PyYAML β€” YAML I/O
jinja2 β€” templates
httpx β€” HTTP client
pymupdf β€” PDF text extraction
anthropic SDK β€” LLM calls (claude-sonnet-4-6)
bibtexparser β€” BibTeX parsing
rich β€” CLI output
python-slugify β€” slug generation
pytest β€” tests
ruff β€” linting

Future SaaS Stack

FastAPI β€” REST API backend
Next.js β€” web frontend
PostgreSQL β€” persistent storage
Redis β€” caching and task queue
Celery β€” background workers
MinIO / S3 β€” PDF and asset storage
GitHub App β€” OAuth + webhook integration
Docker + Kubernetes β€” deployment
Stripe β€” billing

Risks and Tradeoffs

Risk Impact Mitigation
PDF extraction quality varies Medium Mark confidence; allow manual override
LLM generates hallucinated citations High Never auto-trust LLM citations; mark all as needs-verification
Semantic Scholar rate limits Low Add exponential backoff; cache responses
GitHub API rate limits (unauthenticated) Medium Support GITHUB_TOKEN; cache tree responses
Paper text may be behind paywall Medium Document clearly; accept user-provided PDFs
LLM API cost at scale High Batch requests; cache outputs; Stage 2 cost controls