Text Generation
Transformers
Safetensors
English
keylm75m
keylm
small-language-model
base
pretrained
gqa
rope
swiglu
qk-norm
custom_code
Instructions to use Eclipse-Senpai/KeyLM-75M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Eclipse-Senpai/KeyLM-75M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Eclipse-Senpai/KeyLM-75M", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Eclipse-Senpai/KeyLM-75M", trust_remote_code=True, dtype="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps
- vLLM
How to use Eclipse-Senpai/KeyLM-75M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Eclipse-Senpai/KeyLM-75M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Eclipse-Senpai/KeyLM-75M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Eclipse-Senpai/KeyLM-75M
- SGLang
How to use Eclipse-Senpai/KeyLM-75M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Eclipse-Senpai/KeyLM-75M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Eclipse-Senpai/KeyLM-75M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Eclipse-Senpai/KeyLM-75M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Eclipse-Senpai/KeyLM-75M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Eclipse-Senpai/KeyLM-75M with Docker Model Runner:
docker model run hf.co/Eclipse-Senpai/KeyLM-75M
File size: 861 Bytes
8dbce4f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 | """KeyLM model implementation.
KeyLM-75M uses a Qwen3-style decoder (GQA + RoPE + SwiGLU + per-head
QK-RMSNorm). Rather than vendor a full copy of the transformer, the classes
below specialise the upstream Qwen3 implementation and bind it to KeyLMConfig
so the model loads under its own name via `trust_remote_code=True`.
"""
try:
from transformers.models.qwen3.modeling_qwen3 import Qwen3ForCausalLM, Qwen3Model
except ImportError as exc: # pragma: no cover - guidance for old transformers
raise ImportError(
"KeyLM requires a transformers version that ships the Qwen3 model "
"(transformers>=4.51). Please upgrade transformers."
) from exc
from .configuration_keylm import KeyLM75MConfig
class KeyLM75MModel(Qwen3Model):
config_class = KeyLM75MConfig
class KeyLM75M(Qwen3ForCausalLM):
config_class = KeyLM75MConfig
|