Instructions to use PanzerBread/PromptCoT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use PanzerBread/PromptCoT with PEFT:

from peft import PeftModel
from transformers import AutoModelForCausalLM

base_model = AutoModelForCausalLM.from_pretrained("unsloth/deepseek-r1-distill-qwen-7b-unsloth-bnb-4bit")
model = PeftModel.from_pretrained(base_model, "PanzerBread/PromptCoT")

Transformers

How to use PanzerBread/PromptCoT with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="PanzerBread/PromptCoT")

# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("PanzerBread/PromptCoT", dtype="auto")

Notebooks
Google Colab
Kaggle
Local Apps

vLLM

How to use PanzerBread/PromptCoT with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "PanzerBread/PromptCoT"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "PanzerBread/PromptCoT",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker

docker model run hf.co/PanzerBread/PromptCoT

SGLang

How to use PanzerBread/PromptCoT with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "PanzerBread/PromptCoT" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "PanzerBread/PromptCoT",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "PanzerBread/PromptCoT" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "PanzerBread/PromptCoT",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Unsloth Studio new

How to use PanzerBread/PromptCoT with Unsloth Studio:

Install Unsloth Studio (macOS, Linux, WSL)

curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for PanzerBread/PromptCoT to start chatting

Install Unsloth Studio (Windows)

irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for PanzerBread/PromptCoT to start chatting

Using HuggingFace Spaces for Unsloth

# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for PanzerBread/PromptCoT to start chatting

Load model with FastModel

pip install unsloth
from unsloth import FastModel
model, tokenizer = FastModel.from_pretrained(
    model_name="PanzerBread/PromptCoT",
    max_seq_length=2048,
)

Docker Model Runner
How to use PanzerBread/PromptCoT with Docker Model Runner:
```
docker model run hf.co/PanzerBread/PromptCoT
```

PanzerBread commited on Nov 16, 2025

Commit

8278bb4

verified ·

1 Parent(s): 4cef487

Update README.md

Browse files

Files changed (1) hide show

README.md +4 -28

README.md CHANGED Viewed

@@ -29,18 +29,18 @@ The models are trained iteratively using an EM loop:
 1. **E-step**: Generate K=8 rationale candidates, compute rewards, select best
 2. **M-step**: Fine-tune both models on selected (concept, rationale, problem) triples
-- **Developed by:** [Your Name/Organization]
 - **Model type:** LoRA fine-tuned Causal Language Model
 - **Language(s):** English (mathematical reasoning)
-- **License:** Apache 2.0 (inherited from Qwen2.5-7B-Instruct)
-- **Finetuned from:** Qwen/Qwen2.5-7B-Instruct
 ### Model Sources
 - **Base Model:** [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct)
 - **Paper:** [PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning](https://arxiv.org/abs/2509.19894) (arXiv:2509.19894)
 - **Authors:** Xueliang Zhao, Wei Wu, Jian Guan, Zhuocheng Gong, Lingpeng Kong
-- **Related Model:** [PromptCoT Rationale Model (qφ)](https://huggingface.co/PanzerBread/promptcot-q)
 ## Uses
@@ -95,12 +95,6 @@ This model is specialized for mathematical reasoning and may not perform well fo
 - **EM Convergence**: The EM algorithm may converge to local optima, depending on initialization and hyperparameters
 - **Generated Quality**: Generated problems may require manual validation for correctness and appropriateness
-### Technical Limitations
-- **Context Length**: Limited to 512 tokens during EM training (2048 for cold start)
-- **Sampling**: Uses temperature sampling (T=0.7) which may produce diverse but sometimes inconsistent outputs
-- **Reward Function**: The reward is based on log probabilities, which may not perfectly correlate with problem quality
 ### Recommendations
 Users should:
@@ -292,24 +286,6 @@ Zhao, X., Wu, W., Guan, J., Gong, Z., & Kong, L. (2025). PromptCoT 2.0: Scaling
 **Paper Link:** [https://arxiv.org/abs/2509.19894](https://arxiv.org/abs/2509.19894)
-## Glossary [optional]
-<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
-[More Information Needed]
-## More Information [optional]
-[More Information Needed]
-## Model Card Authors
-[Your Name/Organization]
-## Model Card Contact
-[Your Email/Contact]
 ### Framework versions
 - PEFT 0.17.1

 1. **E-step**: Generate K=8 rationale candidates, compute rewards, select best
 2. **M-step**: Fine-tune both models on selected (concept, rationale, problem) triples
+- **Developed by:** Krzysztof Staroń
 - **Model type:** LoRA fine-tuned Causal Language Model
 - **Language(s):** English (mathematical reasoning)
+- **License:** Apache 2.0 (inherited from Qwen2.5-7B)
+- **Finetuned from:** Qwen/Qwen2.5-7B
 ### Model Sources
 - **Base Model:** [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct)
 - **Paper:** [PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning](https://arxiv.org/abs/2509.19894) (arXiv:2509.19894)
 - **Authors:** Xueliang Zhao, Wei Wu, Jian Guan, Zhuocheng Gong, Lingpeng Kong
+- **Related Model:** [PromptCoT2.0](https://huggingface.co/xl-zhao/PromptCoT-2.0-Prompt-Generation-Model)
 ## Uses
 - **EM Convergence**: The EM algorithm may converge to local optima, depending on initialization and hyperparameters
 - **Generated Quality**: Generated problems may require manual validation for correctness and appropriateness
 ### Recommendations
 Users should:
 **Paper Link:** [https://arxiv.org/abs/2509.19894](https://arxiv.org/abs/2509.19894)
 ### Framework versions
 - PEFT 0.17.1