Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
uzeroj 's Collections
[Model] Looped Transformers
[Model] Tool Calling
[Data] Reasoning
[Model] GUI Agent
[Data] GUI Agent | Benchmark
[Data] GUI Agent | Grounding
[Data] Web Agent | Grounding
[Data] GUI Agent | Trajectories
[Data] VLM SFT
[Agent] GUI Agent
[Benchmark]LLM
【sLLM】Text2Text

[Data] VLM SFT

updated 5 days ago

This collection gathers clean multimodal datasets for VLM supervised fine-tuning. Covering visual QA, image dialogue and document reasoning.

Upvote
-

  • liuhaotian/LLaVA-CC3M-Pretrain-595K

    Preview • Updated Jul 6, 2023 • 546 • 180

  • liuhaotian/LLaVA-Pretrain

    Preview • Updated Jul 6, 2023 • 3.14k • 224

    Note https://tinyllava-factory.readthedocs.io/en/latest/Prepare%20Datasets.html

Upvote
-
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs