Whis a dflash speculative decoding model for Seed-OSS-36B-Instruct

#34
by Tariel - opened

Seed-OSS-36B-Instruct is surprising suitable for creative writing and translation tasks compared with SOTA models. But it has no speculative decoding, so its decoding throughput is limited (~25 tokens/s on 2 x 2080ti 22GBs).

A dflash model will greatly increase the decoding speed.

Sign up or log in to comment