nm-testing/tinyllama-oneshot-w8-channel-a8-tensor Text Generation • 1B • Updated about 1 month ago • 1.35k
nm-testing/tinyllama-oneshot-w4a16-group128-v2 Text Generation • 1B • Updated about 1 month ago • 3.65k
nm-testing/tinyllama-oneshot-w4a16-channel-v2 Text Generation • 1B • Updated about 1 month ago • 5.72k • 1
nm-testing/tinyllama-one-shot-w4a16-group-packed Text Generation • 1B • Updated about 1 month ago • 102
nm-testing/llama7b-one-shot-2_4-w4a16-marlin24-t-alt Text Generation • 0.9B • Updated about 1 month ago • 85
nm-testing/llama7b-one-shot-2_4-w4a16-marlin24-t Text Generation • 1B • Updated about 1 month ago • 261 • 1
nm-testing/llama3-8b-w8_channel-a8_tensor-compressed Text Generation • 8B • Updated about 1 month ago • 106
nm-testing/llama-3-instruct-w8a8-dyn-per-token-test Text Generation • 8B • Updated about 1 month ago • 87
nm-testing/granite-8b-code-instruct-128k2of4-W8A8-FP8-Dynamic-Per-Token 8B • Updated about 1 month ago • 20
nm-testing/granite-3.1-8b-instruct2of4-W8A8-FP8-Dynamic-Per-Token 8B • Updated about 1 month ago • 15
nm-testing/TinyLlama-1.1B-Chat-v1.0-kv_cache_default_tinyllama-e2e 1B • Updated about 1 month ago • 13
nm-testing/SmolLM-135M-Instruct-quantized.w4a16 Text Generation • 0.2B • Updated about 1 month ago • 120
nm-testing/Mistral-7B-Instruct-v0.32of4-W8A8-FP8-Dynamic-Per-Token 7B • Updated about 1 month ago • 16
nm-testing/Meta-llama3-8b-Instruct-SmoothQuant-Fp8 Text Generation • 8B • Updated about 1 month ago • 85
nm-testing/Meta-Llama-3-8B-Instruct-nonuniform-test Text Generation • 8B • Updated about 1 month ago • 5.21k