Yi Cui
onekq
AI & ML interests
Benchmark, Code Generation Model
Recent Activity
repliedto their post 5 days ago
I turned off subagents in Claude Code. Am I a minority? posted an update 5 days ago
I turned off subagents in Claude Code. Am I a minority?Organizations
posted an update about 5 hours ago
replied to their post 5 days ago
๐คฃ
replied to their post 17 days ago
Really appreciate everyone's knowledge sharing and critique here.
Isn't this a gloomy picture that no frontier model can be run on the edge? Smaller models are for workflows and automations.
We are all used to conversing with a frontier model and let it run things on your device. A weaker model will degrade your current experience hence no adoption.
replied to their post 19 days ago
Wow! thx for all the works here. I wish more people can see the methodology.
My Mac is only 16GB ๐
We need to find one of those people above who have big budget
https://www.ebay.com/itm/257712960417
posted an update 19 days ago
Post
996
My take on device-side inference: it's all about high bandwidth memory (thinking about it, this holds for the cloud too).
MacBooks enjoy incidental capacity of apple silicon, but per-device RAM is too low (16 to 24GB), only sufficient for a decent SLM. 512GB is the highest you can go (Kimi K2*). Counting MLX downloads of Kimi K2* on Huggingface, I estimate the user base to be <25K.
On the other hand, the newly debuted DGX station (Nvidia) has 748GB, which can fit in the latest Kimi, DS, and Qwen. Also the quantization options of CUDA is way better than MLX.
For high-end inferencing, I place my bet on workstations over Macs.
MacBooks enjoy incidental capacity of apple silicon, but per-device RAM is too low (16 to 24GB), only sufficient for a decent SLM. 512GB is the highest you can go (Kimi K2*). Counting MLX downloads of Kimi K2* on Huggingface, I estimate the user base to be <25K.
On the other hand, the newly debuted DGX station (Nvidia) has 748GB, which can fit in the latest Kimi, DS, and Qwen. Also the quantization options of CUDA is way better than MLX.
For high-end inferencing, I place my bet on workstations over Macs.
posted an update 22 days ago
Post
163
My biggest takeaway from Ben Thompson interview is the legacy of this round of "bubble".
For railroad, it was the track, and dot com, fiber. Compared to these two, GPUs depreciate too fast.
The answer is power plants.
For railroad, it was the track, and dot com, fiber. Compared to these two, GPUs depreciate too fast.
The answer is power plants.
posted an update 25 days ago
Post
2011
Back of the envelope calculation: Ox Alpha has been out for a week with 130K users and 7T tokens burned.
Method 1: Assuming Ox Alpha is the GLM 5.3 class, it has roughly the same active param count as DeepSeek V4 Pro (~40B vs 49B), derive the cost by the floor price DS has ever published.
Method 2: Assuming the users are concentrated within an 8-hour working window each day, derive number of H800 nodes needed (~700 nodes at $2/GPU-hour).
In both methods, I assume 90/10 IO split and 60% cache hit. Both methods come to $2M.
I think the ROI is awesome: (1) publicity and (2) data harvesting.
Method 1: Assuming Ox Alpha is the GLM 5.3 class, it has roughly the same active param count as DeepSeek V4 Pro (~40B vs 49B), derive the cost by the floor price DS has ever published.
Method 2: Assuming the users are concentrated within an 8-hour working window each day, derive number of H800 nodes needed (~700 nodes at $2/GPU-hour).
In both methods, I assume 90/10 IO split and 60% cache hit. Both methods come to $2M.
I think the ROI is awesome: (1) publicity and (2) data harvesting.
replied to their post 26 days ago
Whoever this is, they give away so much compute for a model launch
replied to their post 26 days ago
Most likely
replied to their post 26 days ago
Lots of SLMs, like the 27b one from Qwen
posted an update about 1 month ago
Post
108
I for one believe DeepMind will ship one more frontier model.
In any case GCP is irrelevant here. GDM should still have more compute than DeepSeek, Kimi, or GLM, maybe all of them combined.
In any case GCP is irrelevant here. GDM should still have more compute than DeepSeek, Kimi, or GLM, maybe all of them combined.
posted an update about 1 month ago
Post
3231
Lots of attentions are now on GPU residual value. I find car analogy to be useful.
* Both new and used cars can do the same job (you can run the latest model on 6-year-old A100)
* New cars are more efficient (higher performance per power draw)
* It's not the year, but mileage and maintenance
The last point is undeveloped. We need Carfax and KBB for GPUs.
And there will be lemon GPUs.
* Both new and used cars can do the same job (you can run the latest model on 6-year-old A100)
* New cars are more efficient (higher performance per power draw)
* It's not the year, but mileage and maintenance
The last point is undeveloped. We need Carfax and KBB for GPUs.
And there will be lemon GPUs.
posted an update about 1 month ago
Post
1757
I think contributor pricing is a good idea. But things will get sophisticated very quickly.
* data sharing tiers from exclusive to syndication
* provenance and lineage
* scoping such as duration and geofencing
and their levels of enforcement (best practices, standards, laws)
* data sharing tiers from exclusive to syndication
* provenance and lineage
* scoping such as duration and geofencing
and their levels of enforcement (best practices, standards, laws)
posted an update about 1 month ago
Post
2678
Looking forward to the Qwen 3.8 model drop, and congratulations on joining the trillion parameters club.
But my eyes are on the promised 27B model. Small (<50B) models decline on OpenRouter, because they are being run on local devices. If people see the family trees here on HF they will understand.
But my eyes are on the promised 27B model. Small (<50B) models decline on OpenRouter, because they are being run on local devices. If people see the family trees here on HF they will understand.
replied to their post about 1 month ago
That's a brilliant strategy. Mystery is why Kimi and GLM didn't follow it.