Instructions to use OptGear/Opt.Gear-1M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OptGear/Opt.Gear-1M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="OptGear/Opt.Gear-1M", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("OptGear/Opt.Gear-1M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OptGear/Opt.Gear-1M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OptGear/Opt.Gear-1M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OptGear/Opt.Gear-1M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/OptGear/Opt.Gear-1M
- SGLang
How to use OptGear/Opt.Gear-1M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OptGear/Opt.Gear-1M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OptGear/Opt.Gear-1M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OptGear/Opt.Gear-1M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OptGear/Opt.Gear-1M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use OptGear/Opt.Gear-1M with Docker Model Runner:
docker model run hf.co/OptGear/Opt.Gear-1M
| license: cc-by-nc-sa-4.0 | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| # Opt.Gear-1M | |
| <img width="1000px" src="./OptAIxHuggingFace03.png"> | |
| [](https://opt-ai.kr) | |
| [](https://huggingface.co/OptGear) | |
| > [!Note] | |
| > This repository contains model weights for **Opt.Gear-1M**, a Tiny Language Model (TLM) designed to run on **Micro-Controller Units (MCUs)**, together with artifacts for embedded deployment. <!-- TODO: λ¦¬ν¬ μ€μ ꡬμ±λ¬Ό(κ°μ€μΉ ν¬λ§·, μλ² λλ λΉλ μν°ν©νΈ) νμ ν μμ --> | |
| > | |
| > Opt.Gear-1M is **not a general-purpose chatbot**. It is trained for a narrow embedded-control setting β mapping short natural-language commands to structured board-control intents β and is intended for deeply embedded systems such as sensors, controllers, small robots, and offline human-machine interfaces. | |
| > | |
| > For general-purpose on-device text generation, see [Opt.Gear-270M](https://huggingface.co/OptGear/Opt.Gear-270M) and [Opt.Gear-1B](https://huggingface.co/OptGear/Opt.Gear-1B). <!-- TODO: λ§ν¬ νμΈ --> | |
| The goal of Opt.Gear-1M is not to shrink Opt.Gear-270M/1B further, but to actually run an auto-regressive language model on a microcontroller β an environment far more constrained than mobile NPUs or server GPUs. Deployed directly on the ARM Cortex-M7 core of an STM32H747I-DISCO board with 4-bit weights and FP32 activations (W4A32), Opt.Gear-1M generates at: | |
| > **20 tokens/s on a 400MHz STM32 Cortex-M7 β about 50ms per token.** | |
| Measured on the actual MCU with an embedded C runtime, not a host-side simulator. To our knowledge, Opt.Gear-1M is the **first generative language model to achieve 20 TPS with W4A32 quantization on the ARM Cortex-M7 CPU** of the STM32H747I-DISCO. | |
| ## Opt.Gear-1M Highlights | |
| - **MCU-first design**: every component β vocabulary, context length, KV cache, activation function, memory allocation β is scaled around Flash/SRAM limits and the absence of large matrix accelerators, not just the parameter count. | |
| - **Fixed-size decoding state**: the ConvKV-Gated Mixer keeps a convolution-kernel-sized state (kernel 4) instead of a sequence-length-dependent cache, keeping SRAM usage predictable during auto-regressive decoding. | |
| - **Simple, accelerator-free operators**: ReLU6GLU feed-forward and static memory allocation make every operation cheap to implement on a Cortex-M7 CPU. | |
| - **Compact tokenizer**: a custom 2,048-token byte-level BPE keeps the embedding tables and runtime memory footprint compatible with flash-limited deployment. | |
| - **Self-contained embedded C runtime**: runs without ONNX or X-CUBE-AI dependencies; footprint is measured from the final ELF, not theoretical weight size. | |
| For more details, please refer to our tech report and blog post. <!-- TODO: λ§ν¬ μ°κ²° --> | |
| ## Model Overview | |
| - Type: Causal Language Model (hybrid attention + convolutional mixer) | |
| - Training Stage: task-focused board-control training on `stm32_cmd_synth` (no separate web-scale pre-training phase) | |
| - Architecture | |
| - Number of Parameters: ~1M | |
| - Hidden Dimension: 128 | |
| - Number of Layers: 5 | |
| - Hidden Layout: hybrid of GQA and ConvKV-Gated Mixer | |
| - Grouped-Query Attention: | |
| - Number of Attention Heads: 2 for Q and 1 for KV | |
| - Head Dimension: 64 | |
| - QK-Normalization: QK-LN | |
| - ConvKV-Gated Mixer: | |
| - Causal depthwise 1D convolution on key/value streams | |
| - Convolution Kernel Size: 4 (widened receptive field for the very small hidden dimension) | |
| - Feed-Forward Network: | |
| - Type: ReLU6GLU β ReLU6(x) = min(max(0, x), 6); bounded activation range, simple to implement without large accelerators | |
| - Intermediate Dimension: 272 | |
| - Rotary Position Embedding: global theta 1,000,000 | |
| - Tokenizer: custom byte-level BPE, vocabulary 2,048 | |
| - Word Embedding: untied (separate input embedding and LM head) | |
| - Context Length: 512 (deployment build uses a 160-token sequence length with a 40-token generation limit) | |
| ### Board-Control Post-Training | |
| Opt.Gear-1M follows a different training path from Opt.Gear-270M/1B. It is trained on **`stm32_cmd_synth`**, a synthetic command dataset of **50,000 prompt-response pairs**: each prompt is a short natural-language command for controlling the STM32 board, and each target response is a compact JSON-style control intent. The action space covers **LCD, LED, and camera control**. This design intentionally favors predictable command generation over broad open-domain language ability, matching the constraints of MCU deployment. | |
| ```text | |
| # Illustrative example β TODO: μ€μ λ°μ΄ν° ν¬λ§·μΌλ‘ κ΅μ²΄ | |
| User: turn on the red LED and show "hello" on the screen | |
| Model: {"led": {"color": "red", "state": "on"}, "lcd": {"text": "hello"}} | |
| ``` | |
| ## Measured Performance (STM32H747I-DISCO) | |
| All numbers are measured on the actual device. The firmware is instrumented with `HAL_GetTick`; throughput is measured from the Cortex-M7 execution path and reflects the model-compute portion of auto-regressive generation. Loadable section sizes are measured from the final STM32 ELF with `arm-none-eabi-size`. | |
| | Item | Measured or configured value | | |
| |---|---| | |
| | Board | STM32H747I-DISCO | | |
| | MCU / core | STM32H747XIH6 / ARM Cortex-M7 | | |
| | Configured CPU clock | 400MHz | | |
| | Runtime | Embedded C (no ONNX / X-CUBE-AI dependency) | | |
| | Quantized format | W4A32 (q4 weights, q16 embedding/LM head) | | |
| | Deployment sequence length | 160 tokens | | |
| | Maximum new tokens | 40 tokens | | |
| | Measured throughput | **20 tokens/s** | | |
| | Per-token latency | **50ms/token** | | |
| | `.text` / `.rodata` | 1,596,400 bytes | | |
| | `.data` | 532 bytes | | |
| | `.bss` | 444,404 bytes | | |
| | Total static RAM (incl. reserved heap/stack) | 449,032 bytes | | |
| A few notes on reading these numbers: | |
| - **20 tokens/s** means generating the 40-token maximum takes about 2 seconds of decode time β fast enough for short command responses and local status generation, with token-by-token output observable in real time. | |
| - **`.text`/`.rodata` (β1.5MB)** is not the pure weight size: it includes the embedded runtime code, quantized weights, 16-bit embedding/LM head, lookup tables, scales, and constants. For MCU models, the final ELF β not the theoretical parameter count β is the meaningful footprint measure. | |
| - **Static RAM (β449KB)** is dominated by statically allocated activation, cache (GQA KV cache + ConvKV fixed-size state), intermediate buffers, logit/sampling buffers, and the reserved heap/stack. | |
| ## Deployment | |
| The measured binary uses an embedded C runtime generated for STM32CubeIDE and runs without depending on an ONNX graph or the X-CUBE-AI generated network. The deployment build uses a 160-token sequence length and a 40-token generation limit, keeping the auto-regressive cache and workspace bounded on the MCU while preserving enough context for short instructions and embedded-control prompts. | |
| <!-- TODO: μλ² λλ λΉλ/νλμ± κ°μ΄λ, μμ νμ¨μ΄ λ§ν¬ μΆκ° --> | |
| Potential applications at this speed and footprint: | |
| - Summarizing sensor data into short natural language | |
| - Local command response for small robots | |
| - Offline interfaces for industrial controllers | |
| - Structured status message generation | |
| - Simple Q&A on devices with limited network connectivity | |
| - Short, domain-specific embedded assistants | |
| ## Best Practices | |
| 1. **Stay in the trained domain**: The model is post-trained for short English board-control commands with JSON-style outputs. Out-of-domain prompts (open-ended questions, long-form generation, non-English input) will not produce reliable results. | |
| 2. **Respect the deployment limits**: The model configuration targets a 512-token maximum context, and the reference deployment build uses 160-token sequences with a 40-token generation cap. Longer sequences increase the SRAM-resident cache and workspace. | |
| 3. **Budget by ELF, not parameter count**: When adapting the runtime or retraining the model, verify Flash/SRAM budgets with `arm-none-eabi-size` on the final ELF β runtime code, lookup tables, and scales share the Flash with the weights. | |
| 4. **Predictability over coverage**: On embedded systems, predictable memory use and stable interactive latency matter more than long-context benchmark performance. The fixed-size ConvKV state exists precisely to keep decoding state independent of context length. | |
| ## Limitations | |
| Opt.Gear-1M is not a general-purpose chatbot. The ~1M parameter scale and 2,048-token vocabulary impose clear limits: it targets stable generative capability under extremely small flash, SRAM, and compute budgets, not broad benchmark coverage. The model prioritizes compact English generation, short-form command following, and predictable structured outputs. For general text generation on mobile and edge devices, use [Opt.Gear-270M](https://huggingface.co/OptAI/Opt.Gear-270M) or [Opt.Gear-1B](https://huggingface.co/OptAI/Opt.Gear-1B). | |
| ## Citation | |
| If you find our work helpful, feel free to give us a cite. | |
| ```bibtex | |
| @misc{optgear2026, | |
| title = {{Opt-Gear} Technical Report}, | |
| author = {{Opt.Gear Team}}, | |
| year = {2026}, | |
| url = {https://huggingface.co/OptGear} | |
| } | |
| ``` | |
| --- | |
| Correspondence: [contact@opt-ai.kr](mailto:contact@opt-ai.kr) Β· Hugging Face: [huggingface.co/OptAI](https://huggingface.co/OptGear) | |