Instructions to use davidnichols-ops/adaptive-operator-v4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use davidnichols-ops/adaptive-operator-v4 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("davidnichols-ops/adaptive-operator-v4") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use davidnichols-ops/adaptive-operator-v4 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "davidnichols-ops/adaptive-operator-v4"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "davidnichols-ops/adaptive-operator-v4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use davidnichols-ops/adaptive-operator-v4 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "davidnichols-ops/adaptive-operator-v4"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default davidnichols-ops/adaptive-operator-v4
Run Hermes
hermes
- OpenClaw new
How to use davidnichols-ops/adaptive-operator-v4 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "davidnichols-ops/adaptive-operator-v4"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "davidnichols-ops/adaptive-operator-v4" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use davidnichols-ops/adaptive-operator-v4 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "davidnichols-ops/adaptive-operator-v4"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "davidnichols-ops/adaptive-operator-v4" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "davidnichols-ops/adaptive-operator-v4", "messages": [ {"role": "user", "content": "Hello"} ] }'
Adaptive Operator v4.1 β Professional Post-Training Pipeline
Reference architecture for distilling a Qwen3.5-9B adaptive operator model with custom control tokens (
[FAST],[THINK],[VERIFY],[RECOVER],[ESCALATE]) and structured tool-calling. Aligns with production post-training pipelines from Qwen, Llama 3, and Mistral teams.
Table of Contents
- Pipeline Overview
- Stage 1: SFT (Supervised Fine-Tuning)
- Stage 2: Preference Optimization (DPO)
- Stage 3: Evaluation and Testing
- Stage 4: Domain Adaptation
- Stage 5: Deployment
- File Inventory
- Professional Standards Compliance Matrix
1. Pipeline Overview
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β DATA GENERATION β
β β
β prompt_generator.py βββΊ 10K diverse task prompts β
β β β β
β βΌ βΌ β
β qwen_inference.py 32 parallel workers β
β (Qwen v3.1 teacher (H100 endpoint, β
β on Together AI) ~9.4 prompts/sec) β
β β β
β βΌ β
β qwen_responses.jsonl βββΊ raw teacher responses β
β β (narrates tools, wrong format) β
β βΌ β
β review_orchestrator.py βββΊ 5-reviewer improvement β
β βββ Code Quality & Correctness β
β βββ Tool Selection & Usage β
β βββ Control Token Routing β
β βββ Error Handling & Edge Cases β
β βββ Response Format & Clarity β
β β β
β βΌ β
β improved_responses.jsonl βββΊ fixed, structured responses β
β β (proper [TOKEN] + XML tool calls) β
β β β
β ββββΊ sft_export.py βββΊ sft_train.jsonl (SFT dataset) β
β β β
β ββββΊ dpo_generator.py βββΊ dpo_export.py β
β βββΊ dpo_train.jsonl (DPO preference pairs) β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β β
βΌ βΌ
βββββββββββββββββββ βββββββββββββββββββββββ
β STAGE 1: SFT β β STAGE 2: DPO β
β β β β
β train_sft.py βββββββββββΊβ train_dpo.py β
β Qwen3.5-9B β β Ξ²=0.1, 1 epoch β
β LoRA r=8 β β LoRA r=8 β
β 3 epochs β β from SFT checkpointβ
β LR=1e-5 β β LR=5e-6 β
β β β β
β Together AI β β Together AI β
βββββββββββββββββββ βββββββββββββββββββββββ
β β
ββββββββββββ¬ββββββββββββββββββββ
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β STAGE 3: EVALUATION β
β β
β eval_suite.py β
β βββ Control token accuracy (right token per task type) β
β βββ Tool selection accuracy (right tool for the job) β
β βββ Tool call format validity (valid JSON, correct params) β
β βββ Task completion (end-to-end success) β
β βββ Regression test (v4.1 vs v3.1 comparison) β
β β
β test_cases.py βββΊ 50+ cases across 7 categories β
β regression_test.py βββΊ v4.1 vs v3.1, CI gate threshold 5% β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β STAGE 4: DEPLOYMENT β
β β
β convert_to_mlx.py βββΊ MLX 4-bit for Apple Silicon β
β verify_mlx_model.py βββΊ smoke test (load, generate, parse) β
β HuggingFace upload βββΊ public release β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Key Design Decisions
| Decision | Rationale | Professional Reference |
|---|---|---|
| Teacher = v3.1 LoRA on H100 | Cheaper than GPT-4, domain-specific | Qwen self-distillation |
| 5-reviewer pure-Python improvement | No LLM cost for review, deterministic | LIMA: quality > quantity |
| SFT before DPO | Standard two-stage approach | Qwen3, Llama 3, Mistral |
| LoRA r=8 (not full FT) | 9B model, cost-effective | Together AI best practices |
| 32 parallel API workers | H100 batches concurrent requests | Endpoint autoscaling |
| Resume support on all stages | Long runs can be interrupted | Production reliability |
2. Stage 1: SFT (Supervised Fine-Tuning)
2.1 Instruction Pairs
What it does: Trains the model using curated prompt-and-response examples where the response includes the correct control token and structured tool call.
Data source: 10,000 diverse task prompts generated across 10 categories:
| Category | Count | Expected Mode | Description |
|---|---|---|---|
| tool_use_file_ops | 1,481 | FAST | File read/write/list operations |
| coding_write | 1,454 | THINK | Write new code from spec |
| tool_use_shell | 1,210 | FAST | Shell command execution |
| coding_debug | 1,209 | RECOVER | Debug and fix broken code |
| tool_use_git | 1,028 | FAST | Git operations |
| tool_use_search | 859 | FAST | Grep/find file operations |
| coding_test | 801 | VERIFY | Write/run tests |
| coding_refactor | 749 | THINK | Refactor existing code |
| planning | 722 | THINK | Multi-step task planning |
| recovery | 487 | RECOVER | Error recovery scenarios |
Professional standard (Qwen/Llama 3):
- Dataset size: 50Kβ10M examples (we use 10K β LIMA showed 1K high-quality > 50K noisy)
- Token budget: 200Mβ500M for SFT (our 10K Γ ~1K tokens = ~10M, appropriate for LoRA)
- Data quality > quantity (LIMA finding, Zhou et al. 2023)
2.2 Behavior Cloning
What it does: Teaches the model to:
- Start every response with a control token (
[FAST],[THINK], etc.) - Emit structured tool calls in XML+JSON format (not narrate them)
- Follow the assistant persona (not just complete text)
The 5-reviewer orchestrator fixes v3.1's key weakness: it narrates tool use ("select tool: git β git log --oneline -20") instead of emitting structured calls. The reviewers synthesize a corrected response with proper format:
[FAST]
Direct action.
<tool_call>
{"name": "git", "arguments": {"command": "log --oneline -20"}}
</tool_call>
Professional standard:
- Use
tokenizer.apply_chat_template()before training β - Loss computed only on assistant tokens β (Together AI handles this)
- Messages >4096 tokens truncated per-message β
- Packing=True for short examples β (Together AI default)
2.3 SFT Training Configuration
# train_sft.py parameters
model = "Qwen/Qwen3.5-9B"
n_epochs = 3
learning_rate = 1e-5
lora_r = 8
lora_alpha = 16 # 2x lora_r (standard practice)
training_method = "sft"
packing = True
max_seq_length = 4096
Professional standard (Qwen3-1.7B reference):
- LR: 2e-5 to 5e-5 (ours: 1e-5, conservative for LoRA) β
- Epochs: 1-3 β
- Batch size: 2-8 with gradient accumulation β
- Warmup ratio: 0.1 β (Together AI default)
2.4 Data Quality Filtering
| Filter | Status | Professional Standard |
|---|---|---|
| Deduplication by prompt | β (sft_export.py) | Required |
| Empty prompt/response removal | β | Required |
| Control token injection | β (auto-add [THINK] if missing) |
Domain-specific |
| N-gram decontamination | β οΈ Not implemented | Qwen uses this |
| Length filtering | β οΈ Not implemented | Mistral: 5-4000 chars |
| Toxicity filtering | β οΈ Not implemented | Llama 3 uses safety classifiers |
3. Stage 2: Preference Optimization (DPO)
3.1 Reward Tuning via DPO
What it does: Uses Direct Preference Optimization to teach the model to prefer the improved (reviewed) response over the raw teacher response.
DPO loss function:
L_DPO = -log Ο(Ξ² * [log Ο_ΞΈ(y_w|x) - log Ο_ref(y_w|x)
- log Ο_ΞΈ(y_l|x) + log Ο_ref(y_l|x)])
Preference pair construction:
- Chosen (preferred): The 5-reviewer improved response (proper format, correct tool calls)
- Rejected (non-preferred): The raw Qwen v3.1 response (narrated tools, wrong format)
This teaches the model that structured tool calls > narrated descriptions.
3.2 DPO Training Configuration
# train_dpo.py parameters
model = "Qwen/Qwen3.5-9B"
training_method = "dpo"
dpo_beta = 0.1 # Standard from TRL
n_epochs = 1 # DPO typically 1 epoch
learning_rate = 5e-6 # Lower than SFT (standard)
lora_r = 8
from_checkpoint = "<SFT model output>" # Continue from SFT
Professional standard:
- Ξ²: 0.05β0.1 (ours: 0.1, stable default) β
- LR: 5e-6 to 1e-5 β
- Epochs: 1 β
- Continue from SFT checkpoint β
3.3 Value Alignment
What it does: Teaches the AI to choose safe, polite, and preferred answers over harmful or wrong ones.
| Alignment Dimension | Status | Implementation |
|---|---|---|
| Format preference | β | Structured tool calls > narration |
| Control token accuracy | β | Correct token per task type |
| Error recovery preference | β | RECOVER token for error scenarios |
| Safety (harmful tool avoidance) | β οΈ | Not explicitly trained |
| Politeness/clarity | β | Reviewer 5 (Response Format & Clarity) |
Professional standard (Llama 3):
- Safety classifiers filter training data β οΈ
- Red teaming against harmful prompts β οΈ
- RLHF with reward model (alternative to DPO) β DPO is simpler, equally effective
3.4 DPO Data Format
Together AI requires a specific format (not the flat prompt/chosen/rejected):
{
"input": {
"messages": [
{"role": "system", "content": "You are an adaptive engineering operator..."},
{"role": "user", "content": "Show the last 20 commits"}
]
},
"preferred_output": [
{"role": "assistant", "content": "[FAST]\n<tool_call>{\"name\":\"git\",...}</tool_call>"}
],
"non_preferred_output": [
{"role": "assistant", "content": "[FAST] Simple task. select tool: git..."}
]
}
4. Stage 3: Evaluation and Testing
4.1 Benchmark Testing
What it does: Runs the model through safety and skill tests to check for regressions or errors.
Evaluation categories (eval_suite.py):
| Category | Test Count | What It Tests |
|---|---|---|
| Control token accuracy | 50+ | Right token for task type |
| Tool selection accuracy | 50+ | Right tool for the job |
| Tool call format validity | 50+ | Valid JSON, correct params |
| Task completion (end-to-end) | 50+ | Can it actually do the task? |
| Regression (v4.1 vs v3.1) | All | No regressions from v3.1 |
Test case categories (test_cases.py):
| Category | Expected Mode | Description |
|---|---|---|
| Simple file operations | FAST | read/write/list files |
| Complex coding tasks | THINK | write algorithms, refactor |
| Error recovery scenarios | RECOVER | fix broken code, handle errors |
| Verification tasks | VERIFY | run tests, check output |
| Multi-step planning | THINK | break down complex tasks |
| Tool selection | varies | which tool to use? |
| Edge cases | varies | empty input, unicode, large files |
4.2 Regression Testing
What it does: Compares model v4.1 against v3.1 on the same test cases.
Metrics tracked:
- Improvement: v4.1 better than v3.1
- Regression: v4.1 worse than v3.1
- No change: identical performance
CI gate: --threshold 5% β up to 5% regression allowed (Wilson 95% CI).
Professional standard:
- Wilson confidence intervals β (regression_test.py)
- 200-400 benchmark problems β (50+ test cases)
- Pass@1 metric β (task completion rate)
- LLM Regression Detector pattern β (regression_test.py)
4.3 Statistical Significance
| Metric | Status | Professional Standard |
|---|---|---|
| Wilson 95% CI | β | Required for CI gates |
| Sample size > 30 | β (50+ cases) | Minimum for significance |
| Multiple seeds | β οΈ | Should run each test Γ 3 seeds |
| Effect size reporting | β οΈ | Cohen's d or similar |
5. Stage 4: Domain Adaptation
5.1 Coding Domain
What it does: Refines performance for coding-specific tasks.
| Sub-domain | Coverage | Prompt Categories |
|---|---|---|
| Writing new code | β | coding_write (1,454 prompts) |
| Debugging | β | coding_debug (1,209 prompts) |
| Refactoring | β | coding_refactor (749 prompts) |
| Testing | β | coding_test (801 prompts) |
| Code review | β οΈ | Not covered |
| Documentation | β οΈ | Not covered |
5.2 Tool-Use Domain
What it does: Trains the model to use 12 MAOS tools correctly.
| Tool | Coverage | Expected Mode |
|---|---|---|
| shell | β | FAST |
| file_read | β | FAST |
| file_write | β | FAST |
| file_edit | β | FAST |
| file_list | β | FAST |
| grep | β | FAST |
| find_file | β | FAST |
| git | β | FAST |
| web_search | β | FAST |
| web_fetch | β | FAST |
| todo_write | β | FAST |
| ask_user | β | ESCALATE |
5.3 Recovery Domain
What it does: Trains the model to handle errors gracefully.
| Scenario | Coverage | Expected Mode |
|---|---|---|
| Test failure | β | RECOVER |
| Tool error | β | RECOVER |
| Syntax error | β | RECOVER |
| Missing file | β | RECOVER |
| Permission denied | β οΈ | Not explicitly covered |
| Network failure | β οΈ | Not explicitly covered |
6. Stage 5: Deployment
6.1 MLX Conversion
# Download adapter from Together AI, merge into base, convert to MLX 4-bit
cd /Users/david/mac-ai-os
uv run python -m training.v4_pipeline.convert_to_mlx \
--job-id <DPO job ID> \
--quantize 4bit \
--output-path /Users/david/Projects/local-operator/models/v4/mlx/
6.2 Verification
# Smoke test: load, generate, check control tokens + tool call parsing
uv run python -m training.v4_pipeline.verify_mlx_model \
--model-path /Users/david/Projects/local-operator/models/v4/mlx/Qwen3.5-9B-Adaptive-Operator-v4-MLX-4bit
6.3 HuggingFace Upload
Upload the merged FP16 model and the MLX 4-bit quantized version to HuggingFace for public release.
7. File Inventory
| File | Lines | Stage | Purpose |
|---|---|---|---|
prompt_generator.py |
282 | Data Gen | Generate 10K diverse task prompts |
qwen_inference.py |
593 | Data Gen | Qwen v3.1 inference client (Together AI) |
sft_generator.py |
351 | Data Gen | Orchestrate SFT data generation (32 parallel workers) |
review_orchestrator.py |
1166 | Data Gen | 5-reviewer improvement (pure Python) |
sft_export.py |
270 | SFT | Export improved responses to Together AI SFT format |
dpo_generator.py |
421 | DPO | Generate preference pairs from responses |
dpo_export.py |
341 | DPO | Export to Together AI DPO format |
train_sft.py |
402 | SFT | Submit SFT job to Together AI |
train_dpo.py |
462 | DPO | Submit DPO job to Together AI |
convert_to_mlx.py |
619 | Deploy | Download, merge, convert to MLX |
verify_mlx_model.py |
393 | Deploy | Smoke test converted model |
eval/eval_suite.py |
1566 | Eval | Comprehensive evaluation harness |
eval/test_cases.py |
2068 | Eval | 50+ test cases across 7 categories |
eval/regression_test.py |
700 | Eval | v4.1 vs v3.1 regression comparison |
Data Files
| File | Current Count | Target | Description |
|---|---|---|---|
prompts/prompts_10000.jsonl |
10,000 | 10,000 | Task prompts |
sft/qwen_responses.jsonl |
~6,400+ | 10,000 | Raw teacher responses |
sft/improved_responses.jsonl |
485 | 10,000 | 5-reviewer improved responses |
sft/sft_train.jsonl |
287 | ~10,000 | SFT training dataset |
dpo/preference_pairs.jsonl |
287 | ~10,000 | DPO preference pairs |
dpo/dpo_train.jsonl |
287 | ~10,000 | DPO training dataset |
8. Professional Standards Compliance Matrix
| Standard | SFT | DPO | Eval | Status |
|---|---|---|---|---|
| Instruction pairs (curated prompt-response) | β | β | β | 10K prompts, 5-reviewer improvement |
| Behavior cloning (format, commands, assistant persona) | β | β | β | Control tokens + tool call format |
| Reward tuning (DPO/RLHF) | β | β | β | DPO Ξ²=0.1, chosen=improved, rejected=raw |
| Value alignment (safe, preferred over harmful) | β οΈ | β | β | Format alignment β, safety β οΈ |
| Benchmark testing (safety + skill) | β | β | β | 50+ cases, 7 categories |
| Regression detection (no regressions) | β | β | β | v4.1 vs v3.1, 5% CI gate |
| Domain adaptation (coding, tools, recovery) | β | β | β | 10 categories covering coding + tools |
| Data deduplication | β | β | β | By prompt ID |
| N-gram decontamination | β οΈ | β οΈ | β | Not implemented |
| Length filtering | β οΈ | β | β | DPO has 5-4000 char bounds |
| Toxicity/safety filtering | β οΈ | β οΈ | β | Not implemented |
| Chat template application | β | β | β | Together AI handles |
| Loss only on assistant tokens | β | β | β | Together AI handles |
| Packing | β | β | β | Together AI default |
| Wilson 95% CI | β | β | β | regression_test.py |
| Multiple seeds | β | β | β οΈ | Should run Γ 3 |
| Statistical significance | β | β | β οΈ | Wilson CI, no effect size |
Gaps to Address for Production Grade
- N-gram decontamination β filter prompts that overlap with eval test cases
- Safety/toxicity filtering β add a safety classifier to filter harmful prompts
- Multiple eval seeds β run each test case Γ 3 seeds for variance estimation
- Effect size reporting β add Cohen's d to regression test
- Code review + documentation prompts β add to prompt generator
- Permission/network error recovery β add to recovery scenarios
References
- Qwen3 post-training: Two-stage SFT (foundational + chat) β DPO. 248.7M SFT tokens.
- Llama 3: 10M+ human-annotated examples, 15T pretraining tokens.
- LIMA (Zhou et al. 2023): 1,000 high-quality examples > 50,000 noisy ones.
- Mistral: Dolly 15K base, strict JSONL, function calling format.
- TRL (Hugging Face): Standard DPO implementation, Ξ²=0.1 default.
- Together AI: LoRA fine-tuning, SFT + DPO support, OpenAI-compatible API.