Instructions to use tyzhu/open-loopify-Qwen3-1.7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tyzhu/open-loopify-Qwen3-1.7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tyzhu/open-loopify-Qwen3-1.7B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("tyzhu/open-loopify-Qwen3-1.7B") model = AutoModelForCausalLM.from_pretrained("tyzhu/open-loopify-Qwen3-1.7B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tyzhu/open-loopify-Qwen3-1.7B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tyzhu/open-loopify-Qwen3-1.7B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tyzhu/open-loopify-Qwen3-1.7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tyzhu/open-loopify-Qwen3-1.7B
- SGLang
How to use tyzhu/open-loopify-Qwen3-1.7B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tyzhu/open-loopify-Qwen3-1.7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tyzhu/open-loopify-Qwen3-1.7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tyzhu/open-loopify-Qwen3-1.7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tyzhu/open-loopify-Qwen3-1.7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use tyzhu/open-loopify-Qwen3-1.7B with Docker Model Runner:
docker model run hf.co/tyzhu/open-loopify-Qwen3-1.7B
open-loopify Qwen3-1.7B
The final 3000-step checkpoint of the September 30, 2026 Qwen3-1.7B short-trace run from open-loopify.
Starting from Qwen/Qwen3-1.7B-Base, zero-based layers [12, 19) run three times, sharing their weights during training. The original 28-layer model therefore performs 42 block executions per token. This release unrolls those passes into 42 ordinary Qwen3 layers, so Transformers and vLLM can load it without custom model code or trust_remote_code. The exported file duplicates repeated layers and is approximately 4.85 GB in bfloat16; the 1.7B name refers to the source model, not the size of the unrolled export.
Training
- Complete reasoning traces of at most 8,000 tokens.
- 3,000 steps × 524,288 tokens/step = 1,572,864,000 training tokens.
- Sequence length 16,384; training seed 17; loop specification
12:19:3. - Base revision:
ea980cb0a6c2ae4b936e82123acc929f1cec04c1. - Training source revision:
e151808e0ba853ebd31d82321bfb6febc4d93d86.
Use
The evaluated environment used Transformers 4.57.6 and vLLM 0.13.0. Use a CUDA-enabled PyTorch build compatible with your NVIDIA driver.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
path = "tyzhu/open-loopify-Qwen3-1.7B"
tok = AutoTokenizer.from_pretrained(path)
model = AutoModelForCausalLM.from_pretrained(path, dtype=torch.bfloat16).cuda()
question = "What is the sum of all positive divisors of 36?"
prompt = question + "\n" + r"Please reason step by step, and put your final answer within \boxed{}."
inputs = tok.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True, return_dict=True, return_tensors="pt",
).to(model.device)
out = model.generate(
**inputs, max_new_tokens=16384, do_sample=True,
temperature=0.6, top_p=0.95, eos_token_id=[151645, 151643],
)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
The supplied generation configuration includes both EOS IDs, 151645 and 151643.
Final-checkpoint evaluation
The released weights and tokenizer/configuration files match the evaluated 3000-step artifacts by SHA-256. Each checkpoint used 758 unique questions and 1,178 sampled answers, a 16,384-token output limit, temperature 0.6, and top-p 0.95. AIME reports average accuracy over eight samples per question, not pass@8; MATH-500 and GPQA-Diamond use one sample per question.
| Benchmark | Accuracy | Output-limit truncation |
|---|---|---|
| AIME 2024 (30 questions × 8 samples) | 10.42% | 57.50% |
| AIME 2025 (30 questions × 8 samples) | 10.83% | 50.83% |
| MATH-500 (500 questions) | 70.00% | 14.80% |
| GPQA-Diamond (198 questions) | 16.67% | 47.98% |
This is one training seed with no matched 1.7B dense-training control. These results do not isolate the effect of looping from continued training, and the 4B dense model is not a matched control. The final checkpoint is not the best checkpoint on every benchmark. Long reasoning often reaches the output limit without finishing; generation length and answer correctness should be assessed separately. No broader safety or general-purpose quality evaluation is claimed. Different batch schedules may change the exact sampled text.
Benchmark item digest: e21c6140501b7deba4e6de58d1843ec1ac9a0f7cfb1d102aac791e33fc6c723b.
Model file SHA-256: 924b779dc8186e066ef34221fd34cd1e4594d1a07f1fcc119a2c4a243c9ba715.
License
Apache-2.0, inherited from Qwen3-1.7B-Base. The base model's license is included in this repository.
- Downloads last month
- 216
Model tree for tyzhu/open-loopify-Qwen3-1.7B
Base model
Qwen/Qwen3-1.7B-Base