Instructions to use bottlecapai/ThinkingCap-Qwen3.6-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bottlecapai/ThinkingCap-Qwen3.6-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="bottlecapai/ThinkingCap-Qwen3.6-27B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("bottlecapai/ThinkingCap-Qwen3.6-27B") model = AutoModelForMultimodalLM.from_pretrained("bottlecapai/ThinkingCap-Qwen3.6-27B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bottlecapai/ThinkingCap-Qwen3.6-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bottlecapai/ThinkingCap-Qwen3.6-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.6-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.6-27B
- SGLang
How to use bottlecapai/ThinkingCap-Qwen3.6-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bottlecapai/ThinkingCap-Qwen3.6-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.6-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bottlecapai/ThinkingCap-Qwen3.6-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.6-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use bottlecapai/ThinkingCap-Qwen3.6-27B with Docker Model Runner:
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.6-27B
Excellent! Are you going to do the same for Qwen3.8 27b?
ThinkingCap, in my private benchmarking, sits at the elbow of the efficient frontier of the speed vs. accuracy matrix.
It gets 90% or more of the accuracy of thinking in no-think mode,, meaning it excels in no-think mode, which is useful for agents.
Getting a similar result on the vastly improved Qwen3.8 27b model, would be excellent.
And what would be killer, if someone could convince Prism-ML to do a Bonsai version of it, to run on anemic hardware…
Hey @rcfa-scrat thank you for sharing your private benchmark results! Does it mean that for you ThinkingCap has better accuracy in no-think mode than regular Qwen3.6?
For 3.8, please check the answer in https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B/discussions/23
On my Mac (MacBook Pro M4 Max 128GB RAM), all models running under vmlx-swift, ThinkingCap in Jang_6M quantization is the fastest and most accurate of all the Qwen3.6 variants in non-thinking mode.
This includes several standard Qwen3.6 27B quantizations as well as the Bonsai 27B Ternary version of Qwen3.6.
The only "disappointing" thing: turning on thinking, while taking quite a bit of extra time, barely buys any added accuracy. Some of the quants gained some good amount from having thinking turned on, but at nearly double the time than ThinkingCap with thinking, and nearly 20x the time of ThinkingCap with thinking off.
Take this data with a big grain of salt, because the benchmark suite is still under active development and not all tests have run on all models, but what seems to be clear that ThinkingCap is in the sweet spot of accuracy vs speed in no-think mode, while a particular Nemotron-3-Nano-Omni quant holds a similar spot with thinking turned on, and a Gemma-4-31B-QAT model leads in overall accuracy, but at higher cost.
Qwen3.8-27B is on the to do list for testing, but if the published benchmarks would translate in a similar way to a hypothetical TinkingCap variant as was the case for Qwen3.6, then it would be a total winner. And if you guys were to do a collab with the people doing the Bonsai version, people with low memory systems would be surely thrilled.
i so wish we had a 3.8 version of this . Love the 3.6