AI & ML interests

None defined yet.

dejanseo 
posted an update 11 days ago
view post
Post
98
Early-checkpoint diffusion output retains enough corpus fingerprint for source attribution, before the model can form a coherent sentence.

I'm currently pretraining a masked diffusion language model and the following is output from one of its early checkpoints.

---

competitors of Paramount. Combootsseys announced on September 21 Hden 21. All was on Emb Street Journal. We’re not already full of news given how Cinema says tech moves. “Kitt showed that there are millions of investigating unprecedented positive issues for around us, shifting- on wayside during its “Our CES tour.” The subculture industry has allows hatenders against 47 respects since June, including the animated animated movies season 3, and “Ult with Securecom i Got USA Supreme.” On Friday, the Suicide ruled to stop providing the clifihotic, stupid clam of camps and candles, but there was no nihoke sent to me. Emeruination issued that was not enough to provide a generic explanation of the backlash for two months. But the President was lethally following Measure 2016 years of experimental crowdfunding it never received. The kind of thing that’s held during DJ Day’s Fan Bwn and Suicide is the opposite. It’s a little drama and even thwashing the $4ennig proceeds (he does lame forMiggle is patrol to besting) that often dominate nixing high party campaigns into imagardous, mediocre artistic

I asked Gemini, GPT and Claude to guess the site and they all said the same thing: The Verge.

An LLM would probably do the same with a bag of words by looking at clues in keywords. But it could also say Engadget, TechCrunch, Ars or Gizmodo. And it this case they all locked on a very specific site.

Very interesting.
dejanseo 
posted an update 13 days ago
dejanseo 
posted an update 8 months ago
view post
Post
528
Loss Curve Forecaster — Predict Where Your Training Is Heading
dejanseo/loss-target-estimator

Ever stared at a loss curve wondering:
- How many more hours until I hit my target loss?
- What will my loss be if I train for 10 more hours?
- Is this thing ever going to converge?

I built a tool to answer these questions.

Upload a WandB CSV export and get instant predictions using power law curve fitting — the same dynamics that govern LLM training scaling laws.
What you get:

Asymptote prediction (your loss floor)
Time-to-target estimates
Progress percentage toward convergence
Dual-axis chart showing both wall time and steps

Tip: Export from WandB with Wall Time as the X-axis to see both time and steps on the chart.

https://dejan.ai/tools/loss-estimator/
dejanseo 
posted an update 8 months ago
dejanseo 
updated a Space 9 months ago
dejanseo 
published a Space 9 months ago