Yooo

#1
by Bc-AI - opened
Refract AI Labs org

@ProCreations @Banaxi-Tech could you guys have a look at the code if there’s any errors? I’ve got a 8xH200 run on the 5th from the super micro jumpstart thing and it would be good if no errors occur lol. I’m waiting for my usage to reset

Refract AI Labs org

@Bc-AI here are the issues:

  1. The startup gradient check will reject a healthy model. In gradient_health_summary() names are checked against prefixes such as "vault" and "embedding". FSDP adds _fsdp_wrapped_module., so the gradients are incorrectly classified as cortex gradients. all four other groups become zero, causing the startup probe to raise. Also, gradient norms need to be combined across ranks because individual ranks hold only portions of the parameters.

  2. Telemetry and ablation probes bypass FSDP. model.run_telemetry_probe() and model.run_ablation_probe() resolve to methods on the underlying model. Their internal self(...) calls bypass the FSDP wrapper, so parameters aren’t gathered before computation. This can produce tensor-shape errors. The probes should perform forward passes through the wrapped model(...).

  3. Checkpoint resume is broken in two ways. In restore_checkpoint() Model weights are saved as .safetensors but loaded with torch.load(). Use the already-imported load_file() instead. Only model_shards[0] and optimizer_shards[0] are loaded. The saving function splits larger states into multiple files, so restoration must load and merge every listed shard.

  4. Periodic checkpointing can deadlock the eight GPUs. Each rank independently evaluates the checkpoint condition, although saving requires every rank to participate together. Worse, last_save_time and dirty are updated only on rank zero. Other ranks can therefore enter another save while rank zero continues training. Rank zero should decide when to save, broadcast that decision, and update checkpoint bookkeeping on every rank.

  5. Data-pipeline failure recovery can also deadlock. If rank zero’s stream.batch() raises inside distributed_batch(), the other ranks are already waiting for its broadcast. Rank zero then attempts a collective recovery save that they never enter. Broadcast a success/failure status before broadcasting the batch.

  6. The suggested OOM fix does nothing. The error messages say to set MICRO_BATCH_SIZE=56, but main() overwrites that setting with:
    MICRO_BATCH_SIZE = GLOBAL_MICRO_BATCH_SIZE // WORLD_SIZE
    With the current settings, that always becomes 14 per GPU. To halve the physical batch while preserving tokens per update, change GLOBAL_MICRO_BATCH_SIZE from 112 to 56; accumulation then becomes 2.

I could be wrong on some of those so take my advice with a grain of salt (or 10)

Refract AI Labs org

@Bc-AI where the heck did you get 8xH200s????? how can I get access to some?

Refract AI Labs org

@Bc-AI where the heck did you get 8xH200s????? how can I get access to some?

prob paid

Refract AI Labs org

@GGUFGuy @Hoglet-33 completely free. i am a year 7 student and i got no job so i cant do paid lol. you have to use a like custom email like yourname@yourwebsite.com for sign up i think https://www.supermicro.com/en/jumpstart

Refract AI Labs org

you can also acces B300s and blah at least the person said so the email said if i wanted to trial other stuff i email them.

Refract AI Labs org

@Bc-AI good thing i have a custom domain email! so how does it work and how long do you get the server for and can i do it multiple times?

Refract AI Labs org

@Hoglet-33 yeah you can do multiple times I think I applied for B300x8 after this. So u go to the site find one u like click apply button fill the form our put in our custom email domain and uh my one is for 5 day idk for others

Refract AI Labs org

@Bc-AI you only need a custom email for when you need to reserve the bigger GPU sets, so all of the ones that say request instead of schedule, i still my biz email though, i get my 8x L40S on the 12th

Refract AI Labs org

NVM I set https://mozmail.com IT WORKS 8xL40S

Bro 😎 it's peak

Refract AI Labs org

Shhh! 🤫

Refract AI Labs org

YOOOOO @Bc-AI BananaMind 3 750M A50M IS TRAINING!!!

ArushCodes changed discussion status to closed
Bc-AI changed discussion status to open
Refract AI Labs org

710C75F0-8440-4404-8C16-278E3C162C4E

Refract AI Labs org

@Banaxi-Tech AI SLOP BATTLE

Refract AI Labs org

Lol

Refract AI Labs org
•
edited 2 days ago

710C75F0-8440-4404-8C16-278E3C162C4E

Internet issue .

Sign up or log in to comment