YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Multimodal RAG Agent (Vision + OCR + LLM)

This model serves as a complete AI Agent utilizing a RAG workflow:

  1. Input: An image containing text.
  2. Retrieval: OCR is used to pull text from the image (Document AI).
  3. Augmented Generation: The extracted text is fed as context into an LLM to answer specific user queries (RAG & NLP).

To use this, simply load the model and run the Vision + OCR pipeline to get structured data.

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support