You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen2.5-Coder-32B-Instruct — Heretic (abliterated)

Stage 1 of a research pipeline to build a local SWE coding agent that does not refuse mid-task on legitimate software and security-engineering work. This is Qwen/Qwen2.5-Coder-32B-Instruct with refusal directions abliterated using Heretic — pure weight surgery, no fine-tuning.

Method

  • Abliteration: orthogonalizes attention out-projection and MLP down-projection matrices against per-layer refusal directions (difference-of-means of harmful vs. harmless prompt residuals). No gradient updates.
  • Search: a 300-trial Optuna study, multi-objective — minimize refusals while keeping KL divergence on harmless prompts within budget; the lowest-refusal Pareto-optimal trial was exported.
  • Hardware: single NVIDIA H100 80GB.

Results (abliterated vs. base)

Metric Value Notes
Refusal rate (harmful prompts) 0.03 down from ~0.32 at a small trial budget
KL divergence (harmless prompts) 0.27 lower = closer to base behavior
MMLU Δ −0.004 negative = slightly above base
GSM8K Δ −0.003 negative = slightly above base

Capability is preserved: the abliterated model matches (marginally exceeds) the base on MMLU and GSM8K.

Intended use

A research base for downstream supervised fine-tuning (tool-calling + SWE trajectories) and preference tuning, to produce a coding agent that operates in agentic loops without spurious mid-task refusals on authorized engineering tasks.

Responsible use

Abliteration removes the base model's built-in refusal behavior. This artifact is released for research and for legitimate software/security-engineering use only. Do not use it to generate genuinely harmful content or to facilitate illegal activity. Downstream deployers are responsible for adding appropriate safety guardrails and for complying with the base model's license and applicable law.

Attribution

Downloads last month
120
Safetensors
Model size
33B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PeetPedro/qwen2.5-coder-32b-instruct-heretic

Base model

Qwen/Qwen2.5-32B
Finetuned
(137)
this model
Adapters
1 model