How to use from
Docker Model Runner
docker model run hf.co/AItonomy/PhAI-IDE-4B
Quick Links

PhAI-IDE

PhAI-IDE is a family of models for scientific coding and interaction with tools, available in 4B, 9B, and 72B sizes. Each model is supervised fine-tuned with ms-swift and released as full BF16 weights with the final LoRA adapter merged, together with its configuration and tokenizer.

The training dataset is Codex trajectories, sourced from ScienceIDE.

Models

Model Base model BF16 weights License
PhAI-IDE-4B Qwen3.5-4B 9.08 GB Apache-2.0
PhAI-IDE-9B Qwen3.5-9B 18.82 GB Apache-2.0
PhAI-IDE-72B Qwen2.5-72B-Instruct 145.41 GB Qwen

Weight sizes are approximate; inference also requires memory for runtime allocations and the KV cache.

ScienceAccelBench performance

Task-held-out, localized scientific-code repair on familiar codebases, with original numerical verification. Qwen3.5-4B and Qwen3.5-9B are compared with PhAI-IDE-4B and PhAI-IDE-9B, respectively, on identical tasks. Pass rates are percentages; gains are percentage points.

Size Environment Tasks Qwen3.5 PhAI-IDE Gain (pp)
4B PLUTO-Particles-Dust 3 0.00 33.33 +33.33
9B LAPS 16 31.25 50.00 +18.75
9B MITgcm-biogeo 8 0.00 12.50 +12.50
9B PLUTO-RMHD 7 0.00 28.57 +28.57

Comparison with published models

Scores (%), grouped by benchmark and model size. Each reference entry gives its published score and the PhAI-IDE score difference in percentage points. Reference models are approximately the same size: 3–4B, 7–9B, and 67–72B, respectively.

PhAI-IDE Benchmark Score Reference models: score (difference)
4B BBH multistep-arithmetic-two 97.60 Llama-3.2-3B-Instruct (3.21B): 53.2 (+44.40); Phi-3.5-mini-8k-instruct (3.82B): 95.6 (+2.00)
9B BBH word-sorting 60.40 Llama-3.1-8B-Instruct (8.03B): 51.2 (+9.20); Qwen2.5-7B-Instruct (7.62B): 15.6 (+44.80)
9B MATH-500 92.20 InternLM3-8B-Instruct (8B): 83 (+9.20); Qwen2.5-7B-Instruct (7B): 72.4 (+19.80); Llama-3.1-8B-Instruct (8B): 48.4 (+43.80)
72B AQuA-RAT 77.56 Llama-2-70B-Chat (70B): 31.32 (+46.24)
72B ARC-Easy 84.64 Llama-2-70B (70B): 76.5 (+8.14); DeepSeek-LLM-67B-Chat (67B): 81.6 (+3.04)
72B ARC-Challenge 64.42 Llama-2-70B (70B): 59.5 (+4.92); DeepSeek-LLM-67B-Chat (67B): 64.1 (+0.32)

Reference scores come from the linked publications, model cards, and independent evaluation reports; evaluation settings and sample counts vary by source. Differences describe reported scores across evaluations, rather than matched-protocol head-to-head gains. BBH entries refer to the named tasks.

Quick start

Use Transformers 5.16.1, PyTorch and Accelerate. Set model_id to any model in the table above; the example selects the matching model class.

from transformers import AutoTokenizer, AutoModelForCausalLM, AutoModelForImageTextToText

model_id = "AItonomy/PhAI-IDE-4B"
loader = AutoModelForCausalLM if model_id.endswith("72B") else AutoModelForImageTextToText
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = loader.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
inputs = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Explain how to verify a numerical simulation."}],
    add_generation_prompt=True, enable_thinking=False, return_dict=True, return_tensors="pt",
).to(model.device)
output = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Training procedure

ScienceIDE demonstrations were collected with GPT-5.6-sol and filtered using a numerical-equivalence verifier. They capture code inspection, tool use, and responses to execution feedback.

All three models use ms-swift supervised fine-tuning with LoRA across trainable linear layers for three epochs. The release merges each final checkpoint's adapter into its base model. Retained assistant targets provide the next-token training signal, while conversation history and tool observations provide context. The trajectories retain the native exec / wait interaction format. Heuristic target masking selects assistant actions for supervision while preserving the surrounding interaction history.

Shared setting Value
Training dataset Codex trajectories
Training examples / tasks 4,567 segments / 564 tasks
Validation examples / tasks 544 segments / 81 tasks
Train/validation task overlap 0
Training epochs 3
LoRA rank / alpha / dropout 32 / 64 / 0.05
Released weights LoRA merged into BF16 Safetensors

Long trajectories are organized into segments. Source partition assignments are preserved, with no task identifiers shared between training and validation.

Framework versions

The release was validated with the following environment.

Component Version
Python 3.11
ms-swift 4.5.3
Transformers 5.16.1
PyTorch 2.6.0+cu124
PEFT 0.20.0
Datasets 4.8.4
Tokenizers 0.23.2
Accelerate 1.14.0
Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AItonomy/PhAI-IDE-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(667)
this model
Adapters
2 models

Collection including AItonomy/PhAI-IDE-4B