GGUF
English
Chinese
multilingual
qwen3
qwen3.6
reasoning
coding
coding-agent
academic-writing
uncensored
rys
lora
iq4_nl
bf16
conversational
Instructions to use jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
Use Docker
docker model run hf.co/jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
- LM Studio
- Jan
- Ollama
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF with Ollama:
ollama run hf.co/jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
- Unsloth Desktop
- Pi
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF with Docker Model Runner:
docker model run hf.co/jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
- Lemonade
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
Run and chat with the model
lemonade run user.Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF-BF16
List all available models
lemonade list
- Hermes Agent
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "jackasda211233/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode-GGUF:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload PATCHCODE_TESTING_PROCESS.md with huggingface_hub
Browse files- PATCHCODE_TESTING_PROCESS.md +266 -0
PATCHCODE_TESTING_PROCESS.md
ADDED
|
@@ -0,0 +1,266 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Qwen3.6 AEON RYS PatchCode (merged_lam0.5): What We Actually Did
|
| 2 |
+
|
| 3 |
+
This is the longer, more casual write-up for the PatchCode upload candidate (internal project name `merged_lam0.5`).
|
| 4 |
+
|
| 5 |
+
The clean model card stays short. This document is the full story: what we distilled, exactly how the dataset was built, how we tested it, why the early single-run scores fooled us, why we stopped trusting them, and why the upload candidate ended up being the plain `IQ4_NL` (reasoning-imatrix) merged GGUF rather than a heavier mixed-quant recipe.
|
| 6 |
+
|
| 7 |
+
Related public guides:
|
| 8 |
+
- runtime fork: `https://github.com/noonr48/qwen36-aeon-ik-llama`
|
| 9 |
+
- RYS layer-duplication / architecture guide: `https://github.com/noonr48/qwen36-aeon-ik-llama/tree/main/docs/rys-layer-duplication-guide`
|
| 10 |
+
- previous fine-tuned release (SignalLatch): `https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-SignalLatch-GGUF`
|
| 11 |
+
|
| 12 |
+
Related release line:
|
| 13 |
+
- previous finetune: `Qwen3.6-27B-AEON-RYS-SignalLatch-ckpt386-s010-IQ4_NL.gguf`
|
| 14 |
+
- this upload candidate: `Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode.IQ4_NL.gguf`
|
| 15 |
+
|
| 16 |
+
## Glossary
|
| 17 |
+
|
| 18 |
+
- `AEON`: the upstream/source model family this RYS line was built from (`AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored`).
|
| 19 |
+
- `SignalLatch` / `ckpt386-s010`: the previous finetune in this line — a behaviour LoRA (checkpoint 386) merged into the AEON RYS base at strength `0.10`. PatchCode is built on top of this.
|
| 20 |
+
- `PatchCode` / `merged_lam0.5`: the public name for this release. It is a second behaviour distil (an agentic-coder joint LoRA) merged onto SignalLatch at strength `0.5`.
|
| 21 |
+
- `IQ4_NL`: the quantized GGUF deployment format we actually upload and run.
|
| 22 |
+
- `imatrix`: importance-matrix-assisted quantization data. `reasoning-imatrix` = calibrated on reasoning/coding text (the kind that worked); `media-imatrix` = an earlier calibration kind that underperformed.
|
| 23 |
+
- `ik-llama`: the custom runtime fork. The `qwen3_5` hybrid architecture does not load on stock `llama.cpp` / `vLLM`.
|
| 24 |
+
- `KritaLite`: our hardened real-world discriminator build (a ~160k-token multi-file app, 15 binary verifier components). Single-shot coding gates saturate on this model family, so we stopped trusting them.
|
| 25 |
+
- `discipline` / `fable_style`: a rubric measuring the distilled action-first style (no preamble, claim-requires-run, narrate→act→verify).
|
| 26 |
+
|
| 27 |
+
## The short version
|
| 28 |
+
|
| 29 |
+
We started from the SignalLatch finetune and distilled a second, agentic-coder behaviour LoRA on top of it. The goal was not a new general chat model. The goal was to make the model a better coding agent: action-first execution, claims backed by an actual run, systematic diagnose→fix loops, stable multi-turn tool use, and fewer stalled runs.
|
| 30 |
+
|
| 31 |
+
After a full 5-phase bake-off, the model that held up was:
|
| 32 |
+
|
| 33 |
+
```text
|
| 34 |
+
Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode.IQ4_NL.gguf
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
That means:
|
| 38 |
+
- base: `Qwen3.6-27B-AEON-RYS-SignalLatch-ckpt386-s010`
|
| 39 |
+
- adapter: agentic-coder joint LoRA, checkpoint `3661`
|
| 40 |
+
- merge strength: `0.5` (effective alpha/r = 1.0)
|
| 41 |
+
- deploy format: plain `IQ4_NL` with reasoning-imatrix
|
| 42 |
+
- runtime: custom AEON ik-llama fork
|
| 43 |
+
|
| 44 |
+
The awkward part — and the reason this write-up is long — is that the eventual ship pick was **not** the candidate that looked best early. A mixed-quant recipe (`c76`) hit a perfect-looking build score on the first multi-seed pass and did not reproduce. A 5-seed, same-condition confirm reversed the read. The plain `IQ4_NL` ended up tied with everything else within noise, so the decision fell to non-noise axes (size, recipe safety), where plain `IQ4_NL` wins.
|
| 45 |
+
|
| 46 |
+
## What this was meant to upgrade
|
| 47 |
+
|
| 48 |
+
PatchCode is an upgrade over the existing SignalLatch finetune:
|
| 49 |
+
|
| 50 |
+
```text
|
| 51 |
+
Qwen3.6-27B-AEON-RYS-SignalLatch-ckpt386-s010-IQ4_NL.gguf
|
| 52 |
+
```
|
| 53 |
+
|
| 54 |
+
The new work was not another RYS architecture pass (the architecture is unchanged and is documented in the layer-duplication guide). The new work was a behaviour distil layered on top of SignalLatch, then merged and quantized into the same practical Q4-class deployment lane.
|
| 55 |
+
|
| 56 |
+
Public framing stays narrow:
|
| 57 |
+
|
| 58 |
+
> This is a practical coding-agent / tool-use-oriented fine-tuned IQ4_NL variant of the SignalLatch release.
|
| 59 |
+
|
| 60 |
+
It should not be framed as:
|
| 61 |
+
- a universal upgrade over base in every format
|
| 62 |
+
- a general chat benchmark win
|
| 63 |
+
- a stock `llama.cpp` / `vLLM` model
|
| 64 |
+
- a live-LoRA deployment recipe
|
| 65 |
+
|
| 66 |
+
## The dataset — exact pipeline
|
| 67 |
+
|
| 68 |
+
This is the part most people ask about, so it is written out in full. The training blend is `~58.5k` examples and is made of two pieces: a large **synthetic coding-agent behaviour backbone** and a smaller **curated action-first style slice**, blended together.
|
| 69 |
+
|
| 70 |
+
### Piece 1 — synthetic coding-agent behaviour backbone (~43k)
|
| 71 |
+
|
| 72 |
+
A standalone generator (`generate_v2.py`) produces synthetic multi-turn coding-agent traces. It is **fully synthetic** — no real user data, no scraped repos. The pipeline:
|
| 73 |
+
|
| 74 |
+
1. **Behaviour-driven generation.** A pool of parallel workers calls a coding-agent teacher model. Each call is shaped around a named *behaviour* from a fixed behaviour pool (~30 behaviours), for example:
|
| 75 |
+
- `survey_before_edit` — read/search the real context before touching code
|
| 76 |
+
- `hypothesis_driven_debugging` — form a hypothesis, then verify
|
| 77 |
+
- `tool_intent_first` — express tool intent before prose
|
| 78 |
+
- `weigh_alternatives_then_commit` — weigh ≥3 options, commit to one, verify
|
| 79 |
+
- `external_awareness` — check versions/docs before asserting
|
| 80 |
+
- `recall_first_habit` — recall prior context before re-deriving
|
| 81 |
+
2. **Tool-agnostic vocabulary (anti-lock-in).** Tool calls use a behavioural-category vocabulary (e.g. `memory_search`, `repo_search`, `render_or_visual_proof`), not real tool names. This is deliberate: the model learns *when/why to use a tool*, not a specific vendor's API surface.
|
| 82 |
+
3. **Scenarios.** A synthetic scenario bank provides repo-shaped task context (file trees, failing tests, stack traces) so the traces are grounded in realistic edit/verify loops.
|
| 83 |
+
4. **Quality gates (per sample).** Traces that fail the gates are dropped, not emitted:
|
| 84 |
+
- `no-op-edit` guard (a claimed edit that changes nothing)
|
| 85 |
+
- `claim-without-verify` reject (the assistant claims done with no run/check)
|
| 86 |
+
- `reasoning-empty` / `incomplete-trace` / `lang-runner-mismatch` / `prompt-over-cap`
|
| 87 |
+
5. **Deficit-resume scheduling.** Generation runs continuously, tracks per-behaviour deficits, and resumes after interruption until target counts are met. (~30 samples/sec on the build host.)
|
| 88 |
+
|
| 89 |
+
**Corpus assembly + filtering (exact counts):**
|
| 90 |
+
- raw unified coding corpus: `71,776` samples
|
| 91 |
+
- filter drops `10,666` bad samples → `61,110` kept
|
| 92 |
+
- top drop reasons: `prompt_over_cap` 3,946 · `lang_runner_mismatch` 3,645 · `reasoning_empty` 2,086 · `incomplete_trace` 861 · `claim_without_verify` 620
|
| 93 |
+
- coding training subset used for the blend: `43,075` (`meda_lora_train_v2x1`)
|
| 94 |
+
|
| 95 |
+
The broader synthetic corpus spans five behaviour layers (media-behaviour 42,973 · tool-depth 15,242 · reliability 19,393 · self-correction 31,476 · coding 7,721 = `116,805` total before filtering); the blend draws the coding-oriented subset.
|
| 96 |
+
|
| 97 |
+
### Piece 2 — curated action-first style slice (~7k)
|
| 98 |
+
|
| 99 |
+
A smaller slice of curated execution-style traces that model the exact discipline we wanted to amplify: terse narrate→act→verify, no preamble, claim-requires-run. Composition (`6,953` total):
|
| 100 |
+
- own multi-project execution sessions (`5,455`) — span many different projects on purpose, so the style generalises instead of locking to one domain
|
| 101 |
+
- a different-domain contributor (`1,130`) — explicitly included for cross-project transfer
|
| 102 |
+
- reasoning-chain exemplars (`368`) — weigh-alternatives deliberation seeds
|
| 103 |
+
|
| 104 |
+
**De-identification / anti-lock-in pass:** real tool names, hostnames, absolute paths, and identifiers are abstracted to behavioural-category tokens / placeholders. The supervision is **assistant-turn-only** — system/user/tool turns (where real project content lives) are masked (`IGNORE_INDEX`), so the model learns a *behaviour policy conditioned on varied context*, not project facts as outputs.
|
| 105 |
+
|
| 106 |
+
### Piece 3 — the blend
|
| 107 |
+
|
| 108 |
+
A small blender oversamples the style slice so it is not drowned by the larger coding backbone, then shuffles:
|
| 109 |
+
|
| 110 |
+
- coding backbone: `43,075`
|
| 111 |
+
- style slice oversampled ~2.2×
|
| 112 |
+
- blended training file: `58,576` (`blend_meda_fable`) ≈ **~74% coding backbone / ~26% action-first style**
|
| 113 |
+
|
| 114 |
+
The oversample ratio was chosen so the style shows up without overfitting the smaller slice; a held-out task type was used to check it generalises rather than parrots.
|
| 115 |
+
|
| 116 |
+
### What the dataset is *not*
|
| 117 |
+
|
| 118 |
+
- It is not scraped real-user data or real private repos.
|
| 119 |
+
- It is not a single-topic dataset — both pieces deliberately span many projects/domains.
|
| 120 |
+
- It does not teach new domain *facts*; it teaches an execution *discipline*.
|
| 121 |
+
|
| 122 |
+
## The training piece
|
| 123 |
+
|
| 124 |
+
A single LoRA was joint-co-trained on the blended `58.5k` set (one adapter, not two-then-merge — a prior two-adapter λ-merge plan was superseded because post-hoc merges can kill a fragile capability with no usable λ).
|
| 125 |
+
|
| 126 |
+
Training config:
|
| 127 |
+
- PEFT type: `LORA`
|
| 128 |
+
- rank: `r=32`, alpha: `64` (alpha/r = 2.0)
|
| 129 |
+
- dropout: `0.05`
|
| 130 |
+
- target modules: **all-linear**, including the hybrid arch projections — `q/k/v/o_proj`, `gate/up/down_proj`, `out_proj`, and the linear-attn/SSM projections `in_proj_qkv / in_proj_a / in_proj_b / in_proj_z`
|
| 131 |
+
- supervision: completion-only (assistant turns only)
|
| 132 |
+
- optimiser: adamw, lr `5e-5` + warmup + cosine decay
|
| 133 |
+
- epochs: `1`
|
| 134 |
+
- backend: model-parallel `device_map` across a multi-GPU host (the max-quality path; the no-NVLink fleet ruled out DeepSpeed/FSDP here)
|
| 135 |
+
|
| 136 |
+
Completion:
|
| 137 |
+
- `global_step=3661` = `epoch 1.0` complete
|
| 138 |
+
- final `train_loss ≈ 0.853`
|
| 139 |
+
- runtime ~91h (~89.5 s/it), grad-norm steady (no divergence)
|
| 140 |
+
- 37 checkpoints saved across the run → full trajectory available for eval
|
| 141 |
+
|
| 142 |
+
The adapter was behaviour-focused and small. It was not trained to teach broad new knowledge.
|
| 143 |
+
|
| 144 |
+
## The merge — why λ=0.5
|
| 145 |
+
|
| 146 |
+
The trained default adapter strength (alpha/r = 2.0) was **over-applied**. A checkpoint × strength eval showed half-strength beat full-strength on all three tested checkpoints:
|
| 147 |
+
|
| 148 |
+
| checkpoint | λ=0.3 | λ=0.5 | λ=0.7 | λ=1.0 |
|
| 149 |
+
|---|---:|---:|---:|---:|
|
| 150 |
+
| 3661 | 0.522 | **0.617** | 0.490 | 0.491 |
|
| 151 |
+
| 2600 | 0.567 | **0.573** | — | 0.442 |
|
| 152 |
+
| 1800 | 0.540 | **0.564** | — | 0.397 |
|
| 153 |
+
|
| 154 |
+
At λ=1.0 the adapter was net-neutral-to-harmful (one checkpoint fell *below* the un-adapted base). The mechanism: an over-loud LoRA delta pushes activations into regimes that hurt calibrated behaviour (preamble returns, over-claiming). λ=0.5 (effective alpha/r = 1.0) keeps the style direction but respects base calibration. So the merge was done at **λ=0.5 onto SignalLatch (ckpt386-s010)**, then exported to BF16 GGUF. (A future v2 could bake the good strength in by training at alpha=r=32, removing the inference-time knob.)
|
| 155 |
+
|
| 156 |
+
## Why the final testing moved to merged IQ4_NL
|
| 157 |
+
|
| 158 |
+
The key question was not "best adapter in BF16" — it was "what we would actually deploy". The deploy target was a merged GGUF, `IQ4_NL`, imatrix-quantized, on the custom ik-llama runtime (Jinja + DeepSeek reasoning format + flash attention + graph split, temp `0.7`).
|
| 159 |
+
|
| 160 |
+
Live LoRA loading is not the production path for this release (the tested serving profile uses flash attention, which conflicts with live LoRA on this runtime). So the long-term path became: **merge the adapter first, then export + quantize a full GGUF.** That is why the upload is a merged GGUF, not an adapter.
|
| 161 |
+
|
| 162 |
+
The plain `IQ4_NL` uses the **reasoning/coding imatrix** (the kind that worked). An earlier build used a media-domain imatrix; it underperformed and was superseded.
|
| 163 |
+
|
| 164 |
+
## The testing ladder (5 phases + confirms)
|
| 165 |
+
|
| 166 |
+
Single-shot and hard-suite gates **saturate** on this model family (every quant scores ~the same, including BF16). The discrimination that actually changed the decision came from a 160k-token real-world build (KritaLite) run multi-seed, plus a discipline rubric, plus an autonomous-loop convergence test. The phases:
|
| 167 |
+
|
| 168 |
+
**Phase 1 — single-seed real-world build.** Made the plain `IQ4_NL` look like the winner (0.933 vs c76's 0.867). This was **noise** — it did not reproduce.
|
| 169 |
+
|
| 170 |
+
**Phase 2 — multi-seed KritaLite (3 seeds).** Reversed phase 1: `c76`/`c404`/`c373` hit 0.933; plain `IQ4_NL` dropped to 0.867. Now a mixed-quant recipe looked like the winner.
|
| 171 |
+
|
| 172 |
+
**Phase 3 — 40-recipe broad search.** Returned all-zero. Root cause was a **harness bug** (the eval script imports a `config.json` that was not copied into the eval root), not real scores.
|
| 173 |
+
|
| 174 |
+
**Phase 4 — search re-gate (bug fixed).** Re-scored all 53 candidates correctly. No new recipe beat the curated originals; the broad search does not help this merge.
|
| 175 |
+
|
| 176 |
+
**Phase 5 — discipline + agentic process.** Plain `IQ4_NL` and BF16 led the action-first *discipline* rubric (0.931); `c76` led *process efficiency* (fewest turns/tools/errors).
|
| 177 |
+
|
| 178 |
+
**Overnight 2 — base-precision × attention-promotion matrix (3 seeds).** Decomposed the build/discipline tradeoff. No candidate clears "both" (build ≥ 0.90 **and** discipline ≥ 0.90):
|
| 179 |
+
- promotion destroys discipline regardless of base precision
|
| 180 |
+
- uniform higher precision does **not** fix build (build is not precision-limited)
|
| 181 |
+
|
| 182 |
+
**Confirm — 5-seed, same-condition, baseline vs c76 head-to-head.** The decisive run:
|
| 183 |
+
|
| 184 |
+
| candidate | build (5-seed) | long-context | discipline (5-seed) | size |
|
| 185 |
+
|---|---:|---:|---:|---:|
|
| 186 |
+
| **plain IQ4_NL (reasoning imx)** | 0.920 (±0.067) | 0.975 | 0.842 (±0.333) | 16.6 G |
|
| 187 |
+
| c76 (promoted attn) | 0.907 (±0.067) | 0.935 | 0.867 (±0.292) | 20 G |
|
| 188 |
+
|
| 189 |
+
build gap `0.013` ≪ `0.067` noise floor → **not discriminating**. c76's earlier "0.933 build win" did not reproduce (it scored 0.933 → 0.867 → 0.907 across passes — pure run-to-run variance).
|
| 190 |
+
|
| 191 |
+
**Q8 confirm — 5-seed, near-lossless Q8 vs plain IQ4_NL.** Q8 shows no edge on any axis and is ~2× the size → ruled out. Near-lossless precision buys nothing measurable here.
|
| 192 |
+
|
| 193 |
+
**Agentic-loop — the autonomy axis (8 held-out tasks × 5 seeds).** Each quant runs a held-out mini-project (README + failing pytest suite) autonomously; converged = objective pytest pass, not self-claimed:
|
| 194 |
+
|
| 195 |
+
| quant | convergence | mean turns | recovery | halluc-success | stall |
|
| 196 |
+
|---|---:|---:|---:|---:|---:|
|
| 197 |
+
| plain IQ4_NL | 100% (40/40) | 7.2 | 0.4 | 0% | 0% |
|
| 198 |
+
| c76 | 100% (40/40) | 6.6 | 0.4 | 0% | 0% |
|
| 199 |
+
| Q8 | 100% (40/40) | 7.0 | 0.5 | 0% | 0% |
|
| 200 |
+
|
| 201 |
+
**Non-discriminating (0pp spread).** The distilled discipline (action-first, claim-requires-run, verify-before-claim) is preserved across all quants.
|
| 202 |
+
|
| 203 |
+
## The noise lesson (critical — reuse for every future bake-off)
|
| 204 |
+
|
| 205 |
+
The SignalLatch-style suite is **noisier than it looked**:
|
| 206 |
+
- KritaLite build: ±0.067–0.13 **run-to-run** variance (beyond seed). c76 scored 0.933 → 0.867 → 0.907 on the same gguf.
|
| 207 |
+
- discipline: ±0.3 spread.
|
| 208 |
+
- build is **ceiling-limited** (max 0.933 = 14/15) → zero headroom to discriminate two good quants.
|
| 209 |
+
|
| 210 |
+
**Rule:** 3-seed differences <0.13 on this suite are meaningless. Use **5+ seeds, same-condition head-to-head** before any ship call. Only non-noise axes (size, recipe methodology/safety, long-context at ceiling) reliably tiebreak. HumanEval was rejected — it saturates on Qwen and is the wrong mode for an agent.
|
| 211 |
+
|
| 212 |
+
This is exactly how a 3-seed pass almost shipped the *weaker* model.
|
| 213 |
+
|
| 214 |
+
## The ship decision
|
| 215 |
+
|
| 216 |
+
With build, discipline, long-context, and autonomy all **tied within noise**, the decision fell to non-noise axes, where plain `IQ4_NL` wins all three:
|
| 217 |
+
- **smaller** (16.6 G vs 20–29 G)
|
| 218 |
+
- **marginal long-context** edge (0.975 vs 0.935–0.969)
|
| 219 |
+
- **plain-quant recipe** — the fleet's proven pattern; promotion/mixed recipes carry evidence-harmful risk (discipline collapse) for zero measured benefit
|
| 220 |
+
|
| 221 |
+
Ship: **plain `IQ4_NL` (reasoning-imatrix)**. The mixed-recipe `c76` is retained on disk as the build-heavy fallback if a future, harder build-gate ever discriminates beyond the noise floor (use 5+ seeds).
|
| 222 |
+
|
| 223 |
+
## What the testing says and does not say
|
| 224 |
+
|
| 225 |
+
**Does say:**
|
| 226 |
+
- PatchCode's distilled action-first discipline is preserved through `IQ4_NL` (tied with BF16 across build / long-context / discipline / autonomy).
|
| 227 |
+
- Near-lossless precision (Q8) and attention promotion buy no measurable edge on this suite.
|
| 228 |
+
- Plain `IQ4_NL` is the defensible default on size + recipe safety.
|
| 229 |
+
|
| 230 |
+
**Does not say:**
|
| 231 |
+
- It does not prove PatchCode is better for all tasks.
|
| 232 |
+
- It does not prove plain `IQ4_NL` is globally optimal.
|
| 233 |
+
- It does not make this a stock `llama.cpp` / `vLLM` release.
|
| 234 |
+
- It does not make live LoRA loading the recommended serving setup.
|
| 235 |
+
|
| 236 |
+
The most accurate public sentence:
|
| 237 |
+
|
| 238 |
+
> On a 5-seed, same-condition practical coding-agent bake-off, PatchCode plain `IQ4_NL` tied BF16 within noise on build, long-context, discipline, and autonomous-loop convergence, and was the selected default on size and recipe safety.
|
| 239 |
+
|
| 240 |
+
## Selected artifact
|
| 241 |
+
|
| 242 |
+
```text
|
| 243 |
+
Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode.IQ4_NL.gguf (16.6 GB — recommended)
|
| 244 |
+
Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode.BF16.gguf (57.6 GB — source-quality reference)
|
| 245 |
+
```
|
| 246 |
+
|
| 247 |
+
Recommended runtime: `https://github.com/noonr48/qwen36-aeon-ik-llama`
|
| 248 |
+
|
| 249 |
+
```bash
|
| 250 |
+
./build/bin/llama-server \
|
| 251 |
+
-m /path/to/Qwen3.6-27B-AEON-RYS-Agentic-Coder-PatchCode.IQ4_NL.gguf \
|
| 252 |
+
-c 65536 -ngl 999 -np 1 -fa on -sm none \
|
| 253 |
+
--temp 0.7 --jinja --reasoning-format deepseek --reasoning-budget 0
|
| 254 |
+
```
|
| 255 |
+
|
| 256 |
+
(`<think>` is emitted as a separate `reasoning_content` field — use `--reasoning-format deepseek` or fold it back so tool-action parsing sees the action.)
|
| 257 |
+
|
| 258 |
+
## Final read
|
| 259 |
+
|
| 260 |
+
This was not a clean leaderboard. It was a real engineering pass: distil the style, build a hardened discriminator because the easy gates saturated, get fooled by a one-run perfect build score, repeat the finalists same-condition, discover the build is ceiling-limited and noisy, and ship the smallest plain-quant that ties everything within noise.
|
| 261 |
+
|
| 262 |
+
```text
|
| 263 |
+
PatchCode IQ4_NL is a practical agentic-coder upgrade over the SignalLatch release.
|
| 264 |
+
It is the selected default among the tested quants, tied with BF16 within noise —
|
| 265 |
+
not a universal final answer.
|
| 266 |
+
```
|