Re-export with coreai-torch 0.4.1 (fixes OS 27 beta 2+ load failure); gated token-for-token vs fp32 oracle 1b8c020 verified mlboydaisuke commited on 18 days ago
Add b2 (beta3-loadable) bundle under gpu-pipelined-b2/qwen3_5_0_8b_decode_int8hu_block32_sym/ (b1 retained) 3c154ef verified mlboydaisuke commited on 21 days ago
qwen3.5-0.8B int8lin decode-only loop-free bundle (pipelined engine): Mac 204 tok/s, iPhone 50.3-51.5 tok/s 0019cb6 verified mlboydaisuke commited on Jun 10
ios-gpu: add qwen3_5_0_8b_ios_hc_prefill_q16_b2048_int8.aimodel 23e3da6 verified mlboydaisuke commited on Jun 10
iOS GPU best: fp16 static ctx-2048 monolith (27.7 tok/s) 2d7f93a verified mlboydaisuke commited on Jun 10
static ctx-2048 monolith (iPhone GPU 27.7 tok/s, release config) a94f01c verified mlboydaisuke commited on Jun 10
dynamic int8 bundle (iPhone GPU 12.5 / ANE 14.7 / Mac 58.5 tok/s) 48e0a23 verified mlboydaisuke commited on Jun 10