Speed & MOE

#3
by firehand - opened

Thank you for releasing this! As a "public interest technologist-clinician", it is VERY much appreciated. My early benchmarking is indicating an improvement on the baseline qwen3.6 family for clinical tasks.

A couple of questions:

  1. what tokens/second interference are you seeing compared to baseline qwen3.6 dense? I'm struggling to get above 10-20t/s on either M3 Ultra or RTX3090 platforms
  2. do you have a medical fine-tune of qwen3.6 MOE in the pipeline? That would be awesome ๐Ÿคž๐Ÿป

Thanks again - โค๏ธ๐Ÿ™‡๐Ÿป

EpisteLabs org
โ€ข
edited 7 days ago

Hello firehand,

Here are my response to your questions

  1. what tokens/second interference are you seeing compared to baseline qwen3.6 dense? I'm struggling to get above 10-20t/s on either M3 Ultra or RTX3090 platforms. I don't have M3 Ultra or RTX 3090. You can try switzerchees/Reasoning-Medical-27B-NVFP4 or mradermacher/Reasoning-Medical-27B-GGUF:Q4_K_M for faster interference. I also have endpoint for this model that provide even faster tokes/second.
  2. do you have a medical fine-tune of qwen3.6 MOE in the pipeline? That would be awesome . It will be coming out, please stay tune.

Sign up or log in to comment