Holy sh, a small language model that is mamab!

#1
by FlameF0X - opened

Peak. You rarely see non transformers model that are so small.

basically AI org

Yeah! I kinda got bored with all of the small models being pure transformers. So the mamba/transformer hybrid is what I've settled on for now.

Love the diversity!

basically AI org

Love the diversity!

Thanks! Once I get the other Pebbles released i'm going to do some experimenting with the transformer layers and try some new stuff out...

Love the diversity!

Thanks! Once I get the other Pebbles released i'm going to do some experimenting with the transformer layers and try some new stuff out...

Sounds exciting, hell yeah!

Love the diversity!

Thanks! Once I get the other Pebbles released i'm going to do some experimenting with the transformer layers and try some new stuff out...

If I recall correctly, Mamba is currently not known at scale?

basically AI org

I created a new org https://huggingface.co/basically-experimental and thats probably where i will publish all of the dumb ideas.

Also here is how the training of Pebble 25M is going in case anyone cares: step 86180/190734 | loss 2.1485 | ppl 8.57 | grad 0.233 | lr 0.633x | 0.13M tkn/s | eta 29.7h

so expect to wait a few more days.

I created a new org https://huggingface.co/basically-experimental and thats probably where i will publish all of the dumb ideas.

Also here is how the training of Pebble 25M is going in case anyone cares: step 86180/190734 | loss 2.1485 | ppl 8.57 | grad 0.233 | lr 0.633x | 0.13M tkn/s | eta 29.7h

so expect to wait a few more days.

Sweet will chuck you a follow!

Yeah! I kinda got bored with all of the small models being pure transformers. So the mamba/transformer hybrid is what I've settled on for now.

Epic twin. Epic.

Love the diversity!

Thanks! Once I get the other Pebbles released i'm going to do some experimenting with the transformer layers and try some new stuff out...

If I recall correctly, Mamba is currently not known at scale?

Mamba is relatively recent in comparation with RWKV or other linear attention models.

FlameF0X changed discussion status to closed
FlameF0X changed discussion status to open

oh, wait, i misread the message, sorry

Sign up or log in to comment