Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
master
PRO
fantos
23
11
367
Follow
Gargaz's profile picture
sudanenator's profile picture
pramananda's profile picture
145 followers
·
129 following
AI & ML interests
None yet
Recent Activity
updated
a bucket
about 3 hours ago
gemma-challenge/gemma-fantos-draft
published
a bucket
about 3 hours ago
gemma-challenge/gemma-fantos-draft
reacted
to
SeaWolf-AI
's
post
with 👍
1 day ago
We wrote up our run in The Fast Gemma Challenge — as vidraft-darwin — and wanted to share the recipe. 🙏 https://huggingface.co/spaces/gemma-challenge/gemma-dashboard Verified result: 510.58 TPS at PPL 2.3930 on a single A10G (fw188-ctk49-n64-patchbridge, re-run & VERIFIED). Honest note: on raw TPS there are faster runs (535+), but those went over the PPL bar and didn't verify — what we're proud of is the fastest result that keeps quality. The recipe is already open, so we explained each piece: sliding-window W188, CTK49 kernel tuning, noprecache (honest, verifiable measurement), and an N64 synthetic warmup bridge that shrinks the public↔private gap (~15 TPS), plus INT4 + MTP K=7 + CUDA-graph capture. One rule: only stack quality-neutral speedups. Huge thanks to @firfir-cast, @gemma-slayer, @chiku-inu, @kenyan-duma, @dixie-flatline and everyone who shared their experiments. Full write-up 👇 https://huggingface.co/blog/FINAL-Bench/fast-gemma
View all activity
Organizations
fantos
's models
9
Sort: Recently updated
fantos/MiniCPM-o-2_6
Any-to-Any
•
9B
•
Updated
Nov 2, 2025
•
11
fantos/Qwen3-Omni-30B-A3B-Thinking
Any-to-Any
•
32B
•
Updated
Nov 2, 2025
•
18
fantos/Ming-flash-omni-Preview
Any-to-Any
•
104B
•
Updated
Nov 2, 2025
•
468
fantos/Qwen-Image-Edit-Rapid-AIO
Text-to-Image
•
Updated
Nov 2, 2025
•
1
fantos/GLM-4.6
Text Generation
•
357B
•
Updated
Nov 2, 2025
•
9
fantos/neutts-air
Text-to-Speech
•
0.7B
•
Updated
Nov 2, 2025
•
11
fantos/PaddleOCR-VL
Image-Text-to-Text
•
1.0B
•
Updated
Nov 2, 2025
•
5
fantos/DeepSeek-OCR
Image-Text-to-Text
•
3B
•
Updated
Nov 2, 2025
•
6
fantos/QwQ-32B-bnb-4bit
Text Generation
•
33B
•
Updated
Mar 20, 2025
•
7
•
63