kylewhite0314's picture
Upload README.md with huggingface_hub
30922f1 verified
|
Raw
History Blame Contribute Delete
448 Bytes
# Multimodal RAG Agent (Vision + OCR + LLM)
This model serves as a complete **AI Agent** utilizing a RAG workflow:
1. **Input:** An image containing text.
2. **Retrieval:** OCR is used to pull text from the image (Document AI).
3. **Augmented Generation:** The extracted text is fed as context into an LLM to answer specific user queries (RAG & NLP).
To use this, simply load the model and run the Vision + OCR pipeline to get structured data.