Instructions to use inclusionAI/Ming-flash-omni-2.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use inclusionAI/Ming-flash-omni-2.0 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("inclusionAI/Ming-flash-omni-2.0", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Any Chance of GGUF Quantizations?
#5
by facehuggerfromspace - opened
I am really looking forward to give this model a try on my Strix Halo. Is there any chance of us getting Q4 quantizations of this model? This could be a great replacement for GPT-OSS 120B.
Hi,
I agree.
A GGUF version would help to raise awareness of this model which, given the results presented, deserves to be better known and have greater visibility.
Thank you in advance.
Thank you for your suggestion.
We will work on adding support for GGUF and llama.cpp. According to our roadmap, we will prioritize Int8/Int4 quantization combined with vLLM capabilities.
Please stay tuned.