aagohary commited on
Commit
6f48155
·
verified ·
1 Parent(s): d23c536

Document RL stage on in-house terminal tasks in model card

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -19,7 +19,7 @@ tags:
19
 
20
  MagenticBrain is a 14B-parameter orchestration model from **Microsoft Research AI Frontiers**. It plans multi-step tasks, calls declared tools, and coordinates sub-agents. It does not execute actions itself — every real-world side effect happens inside a host harness.
21
 
22
- The model is supervised fine-tuned from [Qwen/Qwen3-14B](https://huggingface.co/Qwen/Qwen3-14B) on agentic data: function-calling corpora, file-system trajectories, terminal tasks, sub-agent delegation traces, and reasoning data. It's co-designed with **MagenticLite**, our agentic application and harness, and that's the configuration it has been most thoroughly evaluated in.
23
 
24
  We're releasing weights only. Inference code, training recipes, and the execution harness are part of MagenticLite.
25
 
@@ -144,7 +144,7 @@ MagenticBrain runs in non-thinking mode — keep `enable_thinking=False` in the
144
 
145
  ### Approach
146
 
147
- Post-training is Supervised Fine-Tuning on a heterogeneous agentic data mix. Reinforcement learning is in active exploration but not part of this release. Thinking tokens are disabled by default (`enable_thinking=False`) to control verbosity and reduce looping on long trajectories.
148
 
149
  ### Data sources
150
 
 
19
 
20
  MagenticBrain is a 14B-parameter orchestration model from **Microsoft Research AI Frontiers**. It plans multi-step tasks, calls declared tools, and coordinates sub-agents. It does not execute actions itself — every real-world side effect happens inside a host harness.
21
 
22
+ The model is supervised fine-tuned from [Qwen/Qwen3-14B](https://huggingface.co/Qwen/Qwen3-14B) on agentic data: function-calling corpora, file-system trajectories, terminal tasks, sub-agent delegation traces, and reasoning data. After supervised fine-tuning, it was further trained with reinforcement learning on in-house-constructed terminal tasks. It's co-designed with **MagenticLite**, our agentic application and harness, and that's the configuration it has been most thoroughly evaluated in.
23
 
24
  We're releasing weights only. Inference code, training recipes, and the execution harness are part of MagenticLite.
25
 
 
144
 
145
  ### Approach
146
 
147
+ Post-training has two stages. First, Supervised Fine-Tuning on a heterogeneous agentic data mix. Second, a reinforcement learning stage on in-house-constructed terminal tasks, further specializing the model for multi-step CLI execution. Thinking tokens are disabled by default (`enable_thinking=False`) to control verbosity and reduce looping on long trajectories.
148
 
149
  ### Data sources
150