krishnateja95 commited on
Commit
48c464b
·
verified ·
1 Parent(s): b2d7db3

Add model card

Browse files
Files changed (1) hide show
  1. README.md +79 -0
README.md ADDED
@@ -0,0 +1,79 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ tags:
3
+ - fp8
4
+ - moe
5
+ - vllm
6
+ - llm-compressor
7
+ - compressed-tensors
8
+ library_name: transformers
9
+ license: apache-2.0
10
+ base_model: thinkingmachines/Inkling
11
+ pipeline_tag: image-text-to-text
12
+ ---
13
+
14
+ # Inkling-FP8-dynamic
15
+
16
+ ## Model Overview
17
+ - **Model Architecture:** InklingForConditionalGeneration
18
+ - **Input:** Text / Image / Audio
19
+ - **Output:** Text
20
+ - **Model Optimizations:**
21
+ - **Weight quantization:** FP8
22
+ - **Activation quantization:** FP8
23
+ - **Release Date:** 2026-07-15
24
+ - **Version:** 1.0
25
+ - **Model Developers:** RedHatAI
26
+
27
+ This model is a quantized version of [thinkingmachines/Inkling](https://huggingface.co/thinkingmachines/Inkling), a 975B total / 41B active parameter multimodal Mixture-of-Experts model that accepts text, image, and audio inputs and generates text outputs.
28
+
29
+ ### Model Optimizations
30
+
31
+ This model was obtained by quantizing the weights and activations of [thinkingmachines/Inkling](https://huggingface.co/thinkingmachines/Inkling) to FP8 data type using dynamic per-token quantization, ready for inference with vLLM.
32
+ This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%.
33
+
34
+ Weights are quantized statically using per-channel FP8 scaling, and activations are quantized dynamically at inference time using per-token scaling. Only the weights and activations of the linear (attention and MoE expert) layers within the language backbone are quantized using [LLM Compressor](https://github.com/vllm-project/llm-compressor). The vision encoder, audio encoder, token embedding and unembedding layers, normalization layers, biases, MoE routing/gating logic, shared experts, and the model's early dense MLP layers are kept in their original precision.
35
+
36
+ ## Deployment
37
+
38
+ ### Use with vLLM
39
+
40
+ This model can be deployed using [vLLM](https://docs.vllm.ai/en/latest/).
41
+
42
+ 1. Start the vLLM server:
43
+ ```
44
+ vllm serve RedHatAI/Inkling-FP8-dynamic \
45
+ --tensor-parallel-size 8 \
46
+ --max-model-len 131072 \
47
+ --gpu-memory-utilization 0.90 \
48
+ --limit-mm-per-prompt '{"image": 4, "audio": 1}'
49
+ ```
50
+
51
+ > **Tip:** For text-only workloads, pass `--limit-mm-per-prompt '{"image": 0, "audio": 0}'` to skip the vision/audio encoder memory allocation and free up GPU memory for a longer context window.
52
+
53
+ 2. Send requests to the server:
54
+
55
+ ```python
56
+ from openai import OpenAI
57
+
58
+ openai_api_key = "EMPTY"
59
+ openai_api_base = "http://<your-server-host>:8000/v1"
60
+
61
+ client = OpenAI(
62
+ api_key=openai_api_key,
63
+ base_url=openai_api_base,
64
+ )
65
+
66
+ model = "RedHatAI/Inkling-FP8-dynamic"
67
+
68
+ messages = [
69
+ {"role": "user", "content": "Explain quantum mechanics clearly and concisely."},
70
+ ]
71
+
72
+ outputs = client.chat.completions.create(
73
+ model=model,
74
+ messages=messages,
75
+ )
76
+
77
+ generated_text = outputs.choices[0].message.content
78
+ print(generated_text)
79
+ ```