Paper Pulse
Upvote history of every Hugging Face Daily Paper
Your QA reading checks out. The drop is almost all one repo: hails/mmlu_no_train, the MMLU copy lm-evaluation-harness loads, went from 28.7M to 163K on the 30-day count overnight. Robotics got both effects at once: lerobot/pusht fell from 29.9K to 4.5K, while LeRobot repos that load_dataset never touched went from 0 to counted (charlesxu0124/functional-manipulation-benchmark: 0 to 26K).
On the window: the docs only say "within a 5-minute window", not whether it is fixed or extends with each request. snapshot_download requests each file as it starts it, 8 at a time, so a long snapshot keeps hitting the Hub for its whole duration. Unless the window extends, that counts more than once; only HF can confirm. For the largest repos the count clearly isn't full copies: genrobot2025/Gen-HumanEgo is 62 TB and logged 363K downloads in 30 days.
On size: yes, but robotics isn't the outlier on slope. Regressing log downloads on log size (controlling for likes and age), above 4 GB the elasticity is 0.35 for robotics, 0.36 for text classification and 0.29 for text generation. Doubling a repo adds about a quarter more counted downloads, not twice as many.
What is specific is concentration. Since April 2025, about half of robotics' monthly downloads come from repos over 40 GB. Drop every repo of 4 GB or more and the climb is still there, from #21 (Jul–Sep 2024) to #14, #5 and #4 (Jul–Sep 2026), but it stops at #4. So the #2 carries a size tilt, while the rise itself doesn't depend on it.
(Sizes are today's mainSize from hub-stats, so a repo that grew is binned by its current size.)
Good question. The chart's starting point (#23) is the July-September 2024 average, so it sits on the old counter.
But the switch shows up in hub-stats on a precise day: on 2024-10-22 the 30-day counts were recomputed in one step. Question answering fell from 37.7M to 1.5M overnight, and robotics nearly doubled, from 42K to 77K.
Measured only after that, robotics was still #23 in October 2024 (single month; the chart uses 3-month averages), then #20 in November, #19 in December, #9 in February 2025, #4 in April and #2 in July. So the climb from #23 to #2 happens entirely on the new counter, from about 80K downloads a month to 13.7M.
One caveat: the snapshots have no all-time counter before 2025-02-27, so monthly figures from July 2024 to February 2025 are estimated from the 30-day counts (the shaded band in the chart).
Neither, quite: there was no snapshot on 01-31. The 02-01 snapshot came 47.6 hours after 01-30's, so 01-31 and 02-01 are one measurement spread over two days, 42.3M each. The weekday adjustment then split it: Saturday's factor is 1.00 and Sunday's 0.88, so the same number read 0.69 on 01-31 and 0.78 on 02-01. That split was the bug. As one measurement the weekend comes out at 0.735, so neither day is low. So the event is the three weekdays; whether part of the weekend's shortfall belongs to it, a two-day average can't tell.
The same mistake was in two more places, so low_days and the stall rules now treat a value spread over days without a snapshot as one measurement throughout. The pair rule now catches datasets 2026-03-17/18 (0.51 over two days, then 2.21 on 03-19): one window now. The weekday pattern is estimated from days with their own snapshot, which moves three borderline dataset days across 0.7. Models: low_days is 2025-11-08/09 and 2026-01-28 to 01-30 (dce0806f), and the card says why.
Your tag cut checks out on my side: tags above 0.5M a day sit at 0.39 to 0.67 (median 0.44) on 01-28 to 01-30, and 1.04 to 1.31 two weeks earlier. I count 16 tags rather than 18 at that threshold, probably a different window for the cutoff.
Neither dipped on 01-30, and the likes in the same snapshots say more than Spaces or papers do.
Model likes gained per day, from the same hub-stats snapshots as the downloads, were 1.19, 1.14 and 1.05 of their 15-day median on 01-28 to 01-30, and the number of models gaining a like was normal (1.05, 1.10, 0.95), while downloads ran at 0.47, 0.43 and 0.31. Dataset likes: 0.91, 1.04, 0.91, with downloads at 0, 1.34 and 0.56. The snapshots were on time too (13:30 to 13:39 UTC, about 24h apart). So the snapshots were fresh and people were on the site; only the download counters fell behind, and since dl_all never caught up, those requests were never counted. That fits your reading.
Spaces likes: 1.00, 0.99, 1.03. Daily Papers upvotes (from hysts-bot-data/daily-papers-stats, which I now follow in https://huggingface.co/spaces/tardellirs/paper-pulse): 0.71, 1.32, 1.27 after a weekday adjustment, though at ~200 upvotes a day that's a weak signal.
One correction to the overlap: 2025-03-04/05 is a different thing. hub-stats moved its collection from about 14:20 to 00:35 UTC that day, so the 03-04 snapshot came 10.3 hours after the one before. Likes ran at 0.30 as well and no download counter moved at all: a short day from the schedule change, not uncounted requests. That leaves 01-28 to 01-30 as the shared one.
2025-11-08/09: snapshots on time, model likes slightly low (0.93, 0.86), paper upvotes too (0.65, 0.68), downloads lower (0.68, 0.60). Mostly a quiet weekend, though I can't rule out a few uncounted downloads. Spaces can't tell: hub-stats has no Space snapshots from 11-04 to 11-23.
Thanks for checking it against both commits. Your numbers match mine: 43.476B, 13 runs under 0.7x, 769M below the median, 486M back the week after.
Before adding low_days I went through the rest of the series the same way, and it changed a few things:
partial key; the total moves to 43.494B.So low_days is now: under 0.7 of a weekday-adjusted 15-day median, in runs the three days on each side don't make up by at least half. That leaves 8 days for models: 2025-03-04/05, 2025-11-08/09 and your 2026-01-28 to 01-31 (repos_meta.json has 12 for datasets). They keep their measured values, the Hub chart greys them out, and the dataset card describes the rule.
Your 01-28 case holds: a crawl that nothing on either side made up.