Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
π€
Open to Collab
Harshit Kumar Gupta
PRO
harshitkgupta
1
3
Follow
nitya-gupta's profile picture
SeaWolf-AI's profile picture
quantid's profile picture
6 followers
Β·
16 following
https://www.harshitgupta.info/
harshit_k_gupta
harshitkgupta
harshitgupta92
harshitkgupta.bsky.social
AI & ML interests
None yet
Recent Activity
replied
to
their
post
about 6 hours ago
Fine-tuned Qwen 2.5 (0.5B β 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT β and the honest answer is "it depends on what you're optimizing for": β’ PyTorch MPS: 2.2xβ5.7x faster raw throughput, but hits a hard memory wall β can't load a 3B model in FP16 on 16GB. β’ Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1kβ4k tokens). β’ 4-bit quantization doesn't cost you convergence β eval loss tracks closely across backends. β’ The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it β dequant only explains 1.07xβ1.4x of the gap. A ~4.1β4.6x framework-level gap remains either way. All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed. Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx
reacted
to
their
post
with π
3 days ago
Fine-tuned Qwen 2.5 (0.5B β 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT β and the honest answer is "it depends on what you're optimizing for": β’ PyTorch MPS: 2.2xβ5.7x faster raw throughput, but hits a hard memory wall β can't load a 3B model in FP16 on 16GB. β’ Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1kβ4k tokens). β’ 4-bit quantization doesn't cost you convergence β eval loss tracks closely across backends. β’ The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it β dequant only explains 1.07xβ1.4x of the gap. A ~4.1β4.6x framework-level gap remains either way. All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed. Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx
posted
an
update
3 days ago
Fine-tuned Qwen 2.5 (0.5B β 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT β and the honest answer is "it depends on what you're optimizing for": β’ PyTorch MPS: 2.2xβ5.7x faster raw throughput, but hits a hard memory wall β can't load a 3B model in FP16 on 16GB. β’ Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1kβ4k tokens). β’ 4-bit quantization doesn't cost you convergence β eval loss tracks closely across backends. β’ The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it β dequant only explains 1.07xβ1.4x of the gap. A ~4.1β4.6x framework-level gap remains either way. All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed. Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx
View all activity
Organizations
harshitkgupta
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a Space
3 days ago
Sleeping
Agents
1
Training Agents Trackio
π―
1
Show visualized I/O tracking data
liked
2 datasets
6 months ago
microsoft/Taskbench
Viewer
β’
Updated
Aug 21, 2024
β’
17.3k
β’
1.26k
β’
38
Maurus/ToolBench
Viewer
β’
Updated
Aug 17, 2023
β’
88.9k
β’
2.49k
β’
10