kylewhite0314's picture
Upload README.md with huggingface_hub
30922f1 verified
|
Raw
History Blame Contribute Delete
448 Bytes

Multimodal RAG Agent (Vision + OCR + LLM)

This model serves as a complete AI Agent utilizing a RAG workflow:

  1. Input: An image containing text.
  2. Retrieval: OCR is used to pull text from the image (Document AI).
  3. Augmented Generation: The extracted text is fed as context into an LLM to answer specific user queries (RAG & NLP).

To use this, simply load the model and run the Vision + OCR pipeline to get structured data.