What if we go the reliable route instead of the speed route?
What if we make it run on sbcs with few megabytes available?
That's the idea for the next model.
Keep in tune.
congrats!
Well I donโt care if itโs AGI or not I just find it super helpful in real practice
It would be for me if my credits wouldn't drain even light mode ๐ฅฒ
it is not AGI
For me it is not, but it's pretty close for me though
Ill wait for the glm distilled xD
I think we are years away though
Wild if true
Feeling the same :)
Imagine having some kind of server so it always feels updated with your latest integrated models!
byte-level tokenizer is the part that caught my attention, keep it up sir!
Let's go!!!
Yes sir!
Do you mean like the architecture? Could you point at the model?
Waiting for your new models sir
Hey @appvoid
Are you gonna press it?
Done. hahaha bots are becoming more of a thing here lately.
Hmm interesting!
You should only use some loops, like lets say you have 4 layers then:
L1 ๐ ฎ L2 ๐ ฎ L2 ๐ ฎ L3 ๐ ฎ L4 so 5 effective layers notL1 ๐ ฎ L1 ๐ ฎ L2 ๐ ฎ L2 ๐ ฎ L3 ๐ ฎ L3 ๐ ฎ L4 ๐ ฎ L4 because our tests on 1M:
Metric All looped (6 blocks) Partial looped (4 blocks) No loop (3 blocks) Base Bench Elo 875 885 884 Base raw accuracy 33.14% 34.57% 33.71% Base weighted accuracy 32.54% 33.65% 33.57% ARC Easy acc_norm 30.98% 29.50% 30.35% ARC Challenge acc_norm 21.33% 22.53% 22.10% PIQA acc_norm 52.45% 54.30% 52.56% HellaSwag acc_norm 27.28% 27.04% 26.93% ArithMark-3 acc_norm 30.40% 33.00% 30.80% INT Index 3.88 5.37 3.93 Training throughput 344K tok/s 492K tok/s 553K tok/s
It's always 4 steps, every time. I don't know why but I think it has something to do with model capacity.