Text Generation
Transformers
Safetensors
qwen3
context-management
tool-use
agent
conversational
text-generation-inference
Instructions to use tencent/ContextPilot-14B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tencent/ContextPilot-14B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tencent/ContextPilot-14B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("tencent/ContextPilot-14B") model = AutoModelForCausalLM.from_pretrained("tencent/ContextPilot-14B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tencent/ContextPilot-14B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tencent/ContextPilot-14B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tencent/ContextPilot-14B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tencent/ContextPilot-14B
- SGLang
How to use tencent/ContextPilot-14B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tencent/ContextPilot-14B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tencent/ContextPilot-14B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tencent/ContextPilot-14B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tencent/ContextPilot-14B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use tencent/ContextPilot-14B with Docker Model Runner:
docker model run hf.co/tencent/ContextPilot-14B
Update README.md
Browse files
README.md
CHANGED
|
@@ -26,7 +26,7 @@ tags:
|
|
| 26 |
alt="ContextPilot Live Demo"
|
| 27 |
/>
|
| 28 |
</a>
|
| 29 |
-
<a href="">
|
| 30 |
<img
|
| 31 |
src="https://img.shields.io/badge/ContextPilot-Paper-red?logo=arxiv&logoColor=red"
|
| 32 |
alt="Paper"
|
|
@@ -40,7 +40,7 @@ tags:
|
|
| 40 |
</a>
|
| 41 |
</p>
|
| 42 |
|
| 43 |
-
ContextPilot-14B is the Qwen3-14B checkpoint of **ContextPilot**, a proactive context-management framework for long-horizon language-model agents. It teaches agents to plan, maintain long-term memory, and offload less useful context while they continue reasoning and using tools. For more details, see our [paper]() and [code repository](https://github.com/Tencent/ContextPilot).
|
| 44 |
|
| 45 |

|
| 46 |
|
|
@@ -80,4 +80,21 @@ This checkpoint is intended for research on proactive context management, long-h
|
|
| 80 |
|
| 81 |
## Citation
|
| 82 |
|
| 83 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 26 |
alt="ContextPilot Live Demo"
|
| 27 |
/>
|
| 28 |
</a>
|
| 29 |
+
<a href="https://arxiv.org/abs/2608.28476">
|
| 30 |
<img
|
| 31 |
src="https://img.shields.io/badge/ContextPilot-Paper-red?logo=arxiv&logoColor=red"
|
| 32 |
alt="Paper"
|
|
|
|
| 40 |
</a>
|
| 41 |
</p>
|
| 42 |
|
| 43 |
+
ContextPilot-14B is the Qwen3-14B checkpoint of **ContextPilot**, a proactive context-management framework for long-horizon language-model agents. It teaches agents to plan, maintain long-term memory, and offload less useful context while they continue reasoning and using tools. For more details, see our [paper](https://arxiv.org/abs/2608.28476) and [code repository](https://github.com/Tencent/ContextPilot).
|
| 44 |
|
| 45 |

|
| 46 |
|
|
|
|
| 80 |
|
| 81 |
## Citation
|
| 82 |
|
| 83 |
+
```
|
| 84 |
+
@inproceedings{pan-etal-2026-contextpilot,
|
| 85 |
+
title = "ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL",
|
| 86 |
+
author = "Pan, Zhuoshi and
|
| 87 |
+
Pei, Qizhi and
|
| 88 |
+
Lu, Junru and
|
| 89 |
+
Lin, Honglin and
|
| 90 |
+
Zhao, H. Vicky and
|
| 91 |
+
Yin, Di and
|
| 92 |
+
Sun, Xing",
|
| 93 |
+
booktitle = "Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing",
|
| 94 |
+
month = nov,
|
| 95 |
+
year = "2026",
|
| 96 |
+
address = "Budapest, Hungary",
|
| 97 |
+
publisher = "Association for Computational Linguistics",
|
| 98 |
+
abstract = "Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously growing working context. Recent proactive context management methods allow models to edit their own working context with specialized tools, yet they still face three key limitations: (1) a limited toolset restricted to search, deletion, and summarization, with no support for global planning, long-term memory, and adaptive compression; (2) inefficient exploration that treats context management actions uniformly despite their heterogeneous impacts on final outcomes; and (3) coarse-grained credit assignment that assigns the final trajectory-level reward to all intermediate context editing actions during RL. To bridge these gaps, we introduce ContextPilot, a proactive context management framework for long-horizon agentic reasoning. Our approach systematically augments the toolset with planning, long-term memory, and soft context offloading tools. We further propose an RL method tailored for context management, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action-level advantages from all branched trajectories that pass through the corresponding context editing action. Experiments on long-context QA and deep search tasks show that ContextPilot achieves stronger performance with a more compact working context, consistently outperforming existing baselines across various base models and benchmarks. Code is available at \url{https://github.com/Tencent/ContextPilot}."
|
| 99 |
+
}
|
| 100 |
+
```
|