Inference Providers
Active filters: int4
Ju214/Mistral-Small-24B-3.1
Image-Text-to-Text
• 24B • Updated • 87
2imi9/Qwen3-1.7b-gptq-int4
Text Generation
• 2B • Updated • 11
Dhruvil03/Perception-LM-1B-Int4bit
Image-Text-to-Text
• 2B • Updated • 14
• 1
RiverkanIT/Ling-mini-2.0-Quantized
Text Generation
• Updated • 2
ForeseeLab/foreseeai-qwen3-4b-iot-int4
Text Generation
• 4B • Updated • 10
• 1
ISTA-DASLab/Llama-3.1-8B-Instruct-MR-GPTQ-nvfp
Image-Text-to-Text
• 5B • Updated • 19
ISTA-DASLab/Llama-3.1-8B-Instruct-MR-GPTQ-mxfp
Image-Text-to-Text
• 5B • Updated • 16
huawei-csl/Qwen3-1.7B-4bit-SINQ
Text Generation
• 1B • Updated • 20
• 5
huawei-csl/Qwen3-1.7B-4bit-ASINQ
Text Generation
• 1B • Updated • 9
• 5
huawei-csl/Qwen3-32B-4bit-SINQ
Text Generation
• 18B • Updated • 13
• 7
huawei-csl/Qwen3-14B-4bit-SINQ
Text Generation
• 9B • Updated • 9
• 5
huawei-csl/Qwen3-14B-4bit-ASINQ
Text Generation
• 9B • Updated • 11
• 6
huawei-csl/Qwen3-32B-4bit-ASINQ
Text Generation
• 18B • Updated • 10
• 8
ModelCloud/GLM-4.6-GPTQMODEL-W4A16-v1
Text Generation
• 357B • Updated • 12
ModelCloud/GLM-4.6-GPTQMODEL-W4A16-v2
Text Generation
• 357B • Updated • 14
• 1
PangaiaSoftware/YanoljaNEXT-Rosetta-4B-onnx
Translation
• Updated • 8
• 2
RedHatAI/NVIDIA-Nemotron-Nano-9B-v2-quantized.w4a16
Text Generation
• 2B • Updated • 403
• 5
ModelCloud/GLM-4.6-REAP-268B-A32B-GPTQMODEL-W4A16
Text Generation
• 269B • Updated • 12
• 2
AhtnaGlen/phi-4-mini-instruct-int4-sym-npu-ov
Text Generation
• Updated • 13
tencent/DeepSeek-V3.1-Terminus-W4AFP8
Text Generation
• 349B • Updated • 738
• 16
ModelCloud/MiniMax-M2-GPTQMODEL-W4A16
Text Generation
• 229B • Updated • 24
• 3
ModelCloud/Marin-32B-Base-GPTQMODEL-W4A16
Text Generation
• 33B • Updated • 5
• 1
ModelCloud/Marin-32B-Base-GPTQMODEL-AWQ-W4A16
Text Generation
• 33B • Updated • 12
• 2
huawei-csl/Apertus-8B-2509-4bit-SINQ
Text Generation
• 5B • Updated • 8
• 2
huawei-csl/Apertus-8B-2509-4bit-ASINQ
Text Generation
• 5B • Updated • 8
• 3
ModelCloud/Granite-4.0-H-1B-GPTQMODEL-W4A16
Text Generation
• 1B • Updated • 12
• 1
ModelCloud/Granite-4.0-H-350M-GPTQMODEL-W4A16
Text Generation
• 0.3B • Updated • 6
• 1
ModelCloud/Brumby-14B-Base-GPTQMODEL-W4A16
Text Generation
• 15B • Updated • 6
• 1
ModelCloud/Brumby-14B-Base-GPTQMODEL-W4A16-v2
Text Generation
• 15B • Updated • 6
• 1
SherlockID365/Qwen3-VL-8B-Instruct-quantized.w4a16
Image-Text-to-Text
• 3B • Updated • 4.65k
• 1