AI & ML interests

Apple ml-explore Computer Vision community

prithivMLmods 
posted an update 2 days ago
view post
Post
3122
OneDecision-VisionGuard-Demo is now available on Hugging Face Spaces!

🤗 Space: prithivMLmods/OneDecision-VisionGuard-Demo

This demo showcases the OneDecision-VisionGuard family of multimodal image classification models for detecting NSFW and other sensitive visual content, with structured JSON reasoning, improved accuracy, and better handling of edge cases such as sensitive imagery, uncensored analysis, scene descriptions, and classification reasoning.

📦 Models: 27B, 9B, 4B — prithivMLmods/OneDecision-VisionGuard-27B-SFT, prithivMLmods/OneDecision-VisionGuard-9B-SFT, prithivMLmods/OneDecision-VisionGuard-4B-SFT

↗️ Collection: https://huggingface.co/collections/prithivMLmods/onedecision-visionguard

To learn more, visit the app page or the respective model pages.
arudradey 
posted an update 13 days ago
prithivMLmods 
posted an update 16 days ago
view post
Post
3904
Qwen-Image-2.1 Plug and Play LoRA App is now live on Hugging Face Spaces.

🔗 Space: prithivMLmods/Qwen-Image-2.1-LoRAs-PnP

It supports standard inference, 4-step Turbo inference, custom LoRA lazy repacks, and LoRA Plug and Play (PnP), all in one setting!

🔗 Qwen-Image-2.1 Image-to-Image LoRAs: https://huggingface.co/collections/prithivMLmods/qwen-image-21-image-to-image-loras

🔗 GitHub: https://github.com/PRITHIVSAKTHIUR/Qwen-Image-2.1-LoRAs-PnP

To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 23 days ago
view post
Post
841
VisionGuardrail EVO-2, a multimodal image-classification content-safety model based on Qwen/Qwen3.8-27B, is now available on the Hub!

Stricter image classification than before, with a dense 27-billion-parameter multimodal model, more precise reasoning, and improved captions for classifying visual media.

➠ Models: prithivMLmods/VisionGuardrail-Evo2-27B, prithivMLmods/VisionGuardrail-Evo2-27B-GGUF

➠ Collection: https://huggingface.co/collections/prithivMLmods/visionguardrail-evo2

➠ Previous Models: https://huggingface.co/collections/prithivMLmods/visionguardrail-collection

⤷ To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 27 days ago
prithivMLmods 
posted an update about 1 month ago
view post
Post
3869
VisionGuardrail, a multimodal content-safety classifier based on Qwen3.5, is now available on Hugging Face in 4B and 9B variants. It is a direct upgrade to ImageShield-MMCF, providing improved parental controls through conservative visual content-safety filtering.

More About:
➠ hf.co/blog — https://huggingface.co/blog/prithivMLmods/vision-guardrail-mini-blog

➠ Models:
✦ VisionGuardrail-4B: prithivMLmods/VisionGuardrail-4B
✦ VisionGuardrail-9B: prithivMLmods/VisionGuardrail-9B

➠ Dataset:
✦ ImageShield-Guardrail-Pro: prithivMLmods/ImageShield-Guardrail-Pro

⤷ To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update about 1 month ago
view post
Post
3109
ImageShield-MMCF — Multimodal Content Filter is a multimodal content-safety classifier built on top of Qwen3.5 and is now available on Hugging Face!

This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Not Safe for Work (NSFW) and other potentially sensitive visual content.

The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block Not Safe for Work (NSFW) content generation and paves the way for more meaningful and responsible creativity.

⊹ ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
⊹ ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
  • 2 replies
·
prithivMLmods 
posted an update about 2 months ago
view post
Post
5295
The Qwen3.8 27B demo for object grounding is now available on Hugging Face Spaces.

It features three tasks: Object Detection (Bounding Boxes), Point Localization (Keypoints), and Spatial Guidance (Path Mapping).

Try it now: prithivMLmods/Qwen3.8-27B-Object-Detection
prithivMLmods 
posted an update 2 months ago
view post
Post
5575
Made a demo for Text/Image-to-3D Video and Image-to-3D Video asset generation using TRELLIS.2. It is paired with Z-Image-Turbo to accelerate the input image preprocessing pipeline, streamlining the Image-to-3D workflow. The generated GLB (GL Transmission Format) files are converted into MP4 (MPEG-4) videos, making them easy to preview and share. Try it now on Hugging Face Spaces.🤗

➠ Image-to-3D-Video-Asset-Generator: prithivMLmods/Image-to-3D-Video-Asset-Generator
➠ collection: https://huggingface.co/collections/prithivMLmods/multimodal-implementations
➠ github: https://github.com/PRITHIVSAKTHIUR/Image-to-3D-Video-Asset-Generator

⤷ To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 4 months ago
view post
Post
7986
Wan2.2-I2V-Fast with highly upscaled sequential frame sampling is now available as a Spaces demo, built using Wan2.2-I2V and FLUX.2-Klein. Try the demo using the links below.👇

➠ wan2.2-i2v-fast : prithivMLmods/Wan2.2-Fast
➠ github: https://github.com/prithivsakthiur/wan2.2-i2v-fast
➠ collection: https://huggingface.co/collections/prithivMLmods/image-generation-apps-collection

⤷ To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 4 months ago
prithivMLmods 
posted an update 4 months ago
view post
Post
6348
PiD — Pixel Diffusion Decoder Image Edit Upscale and Image Generation Upscale, an all-in-one demo, is now live on Spaces! Great improvements in realism-based image generation and editing are powered by FLUX.2-Klein, while image generation is paired with Z-Image, and upscaling is enabled by default!

🤗 Space: prithivMLmods/PiD-Image-Upscaler
🔗 Collection: https://huggingface.co/collections/prithivMLmods/image-generation-apps-collection

🤗 > To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 5 months ago
view post
Post
5656
I've made 8 Spaces in the Qwen-Image-Edit series, and out of them, 5 Spaces reached “Space of the Week”! A few Spaces are still topping the list even after many months.

Cumulatively, the series has crossed 8.2 million+ ZeroGPU runs and nearly 4 million visitors overall.

Thanks for all the community support! 🤗❤️

🔗 Spaces: https://huggingface.co/collections/prithivMLmods/image-generation-apps-collection
  • 4 replies
·
prithivMLmods 
posted an update 5 months ago
view post
Post
5984
Multimodal-Edge Demo, a node-based inference canvas demo, is now live on Spaces. It features node-based Transformers for fast inference across 10+ edge-device multimodal models on the Hub, all within a single space. The series includes models from Qwen3.5, Qwen3-VL, Gemma 4, and the LFM 2.5 VL model series, with support for reasoning and grounding tasks.

🤗 Demo: prithivMLmods/Multimodal-Edge-Node
🔗 GitHub: https://github.com/PRITHIVSAKTHIUR/Multimodal-Edge-Node
✅ Multimodal Apps Collections: https://huggingface.co/collections/prithivMLmods/hall-of-multimodal-apps

🤗 > To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 6 months ago
view post
Post
1959
Now, a collection of various compression schemes for Qwen3.6 and the abliterated version 1 of dense models is available on the Hub. Check it out via the links below. 👇

🔗 Qwen3.6-MoE: https://huggingface.co/collections/prithivMLmods/qwen36-35b-a3b-compressions
🔗 Qwen3.6-27B Compressions: https://huggingface.co/collections/prithivMLmods/qwen36-27b-compressions

🤗 > To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 6 months ago
view post
Post
4268
HY-World-2.0 — A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds is now available on Spaces, and it works both as native Gradio components and in Gradio server mode.

> HY-World-2.0-Demo: prithivMLmods/HY-World-2.0-Demo
> HY-World-2.0 [Server Mode]: prithivMLmods/HY-World-2.0-Demo
> Featuring 3D reconstruction and Gaussian splats with the Rerun viewer, along with camera poses, depth maps, and surface normals.
> In Server Mode, Gradio is served via FastAPI, with FastAPI remaining the top-level server.
> Model: tencent/HY-World-2.0
> GitHub: https://github.com/PRITHIVSAKTHIUR/HY-World-2.0-Demo

🤗To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 6 months ago
view post
Post
6295
A new comparator on Spaces showcases Standard FLUX.2 Decoder vs. FLUX.2 Small Decoder. The Small Decoder is ~1.4× faster, uses ~1.4× less VRAM, and maintains near-identical image quality. It has ~28M parameters with narrower channels [96, 192, 384, 384] vs. [128, 256, 512, 512], and the demo supports sequence generation by running both decoders simultaneously and comparing the results side by side.

🤗 Comparator: https://huggingface.co/spaces/prithivMLmods/Flux.2-4B-Decoder-Comparator
🔗 FLUX.2-small-decoder: black-forest-labs/FLUX.2-small-decoder
🔗 GitHub: https://github.com/PRITHIVSAKTHIUR/Flux.2-4B-Encoder-Comparator
🚁 Collection: https://huggingface.co/collections/prithivMLmods/image-generation-apps-collection

🤗 > App built on the Gradio SDK. To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 6 months ago
view post
Post
4282
Now, a collection of various compression schemes for Gemma 4 and the abliterated version 1 of dense models is available on the Hub. Check it out via the links below. 👇

🔗Gemma 4 Compression(s)- https://huggingface.co/collections/prithivMLmods/gemma-4-compressions
🔗Gemma 4 Uncensored [MAX] + Compression(s) - [`β ]- https://huggingface.co/collections/prithivMLmods/gemma-4-uncensored-max-compressions
🔗Gemma 4 Compression(s) - MoE- https://huggingface.co/collections/prithivMLmods/gemma-4-compressions-moe
🔗Gemma-4 F32 GGUF- https://huggingface.co/collections/prithivMLmods/gemma-4-f32-gguf

🤗 > To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 6 months ago
view post
Post
2387
Now the demo for image detection based on SAM3 and Gemma-4 (*Filter) is available on Spaces, using full-fledged Transformers inference with multimodal reasoning for processed images. It also supports video segmentation (mask), video segmentation (annotation), and image click segmentation.

🤗 Demo Space: prithivMLmods/SAM3-Gemma4-CUDA
🥽 SAM3: facebook/sam3
🔗 gemma-4-E2B-it: google/gemma-4-E2B-it

To learn more, visit the app page or the respective model pages.
  • 1 reply
·
prithivMLmods 
posted an update 6 months ago
view post
Post
4813
The demo for Image Detection (*Filter) based on SAM3 and Qwen-3.5 is now available on Hugging Face Spaces using Transformers inference, with multimodal reasoning for processed images, and it also supports video segmentation (mask), video segmentation (annotation), and image click segmentation.

🤗 Demo Space: prithivMLmods/SAM3-Plus-Qwen3.5
🥽 SAM3: facebook/sam3
🔗 Qwen-3.5: Qwen/Qwen3.5-2B

To learn more, visit the app page or the respective model pages.
  • 5 replies
·