axjns 's Collections

One GPU, Real Work

Everything I run on a 64GiB unified-memory Strix Halo box, with my own measured numbers. The finding: MoE decodes ~6x faster than dense here.