Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

Spaces:
alphaXiv
/
m1IRWFAMsa
Running

App Files Files Community
m1IRWFAMsa / pages
170 kB
Ctrl+K
Ctrl+K
  • 1 contributor
History: 2 commits
alphaXiv's picture
alphaXiv
Update logbook: Repro: NorMuon: Making Muon more efficient and scalable
5816b7d verified about 1 month ago
  • claim-1-normuon-algorithm-muon-orthogonalization-neuron-wise-2nd-moment-normalization-alg-1
    Update logbook: Repro: NorMuon: Making Muon more efficient and scalable about 1 month ago
  • claim-2-muon-leaves-high-per-neuron-update-norm-variance-normuon-normalizes-it-fig-1
    Update logbook: Repro: NorMuon: Making Muon more efficient and scalable about 1 month ago
  • claim-3-21-74-step-efficiency-gain-over-adam-11-31-over-muon-at-1-1b-table-1
    Update logbook: Repro: NorMuon: Making Muon more efficient and scalable about 1 month ago
  • claim-4-normuon-improves-validation-loss-trajectories-at-5-4b-fig-2
    Update logbook: Repro: NorMuon: Making Muon more efficient and scalable about 1 month ago
  • claim-5-optimizer-memory-near-muon-2-9-step-time-overhead-vs-adamw-table-2
    Update logbook: Repro: NorMuon: Making Muon more efficient and scalable about 1 month ago
  • claim-6-normuon-outperforms-muon-in-modded-nanogpt-at-124m-and-350m-fig-5
    Update logbook: Repro: NorMuon: Making Muon more efficient and scalable about 1 month ago
  • index.md
    1.12 kB
    Update logbook: Repro: NorMuon: Making Muon more efficient and scalable about 1 month ago