Text Generation
PEFT
Safetensors
Transformers
lora
research
alignment-tax
safety-regression
evaluation
conversational
Instructions to use Vulcora/gemma-2-2b-it-codealpaca-alignmenttax with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Vulcora/gemma-2-2b-it-codealpaca-alignmenttax with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-2-2b-it") model = PeftModel.from_pretrained(base_model, "Vulcora/gemma-2-2b-it-codealpaca-alignmenttax") - Transformers
How to use Vulcora/gemma-2-2b-it-codealpaca-alignmenttax with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Vulcora/gemma-2-2b-it-codealpaca-alignmenttax") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Vulcora/gemma-2-2b-it-codealpaca-alignmenttax", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Vulcora/gemma-2-2b-it-codealpaca-alignmenttax with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Vulcora/gemma-2-2b-it-codealpaca-alignmenttax" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Vulcora/gemma-2-2b-it-codealpaca-alignmenttax", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Vulcora/gemma-2-2b-it-codealpaca-alignmenttax
- SGLang
How to use Vulcora/gemma-2-2b-it-codealpaca-alignmenttax with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Vulcora/gemma-2-2b-it-codealpaca-alignmenttax" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Vulcora/gemma-2-2b-it-codealpaca-alignmenttax", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Vulcora/gemma-2-2b-it-codealpaca-alignmenttax" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Vulcora/gemma-2-2b-it-codealpaca-alignmenttax", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Vulcora/gemma-2-2b-it-codealpaca-alignmenttax with Docker Model Runner:
docker model run hf.co/Vulcora/gemma-2-2b-it-codealpaca-alignmenttax
File size: 4,612 Bytes
062ce57 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 | ---
base_model: google/gemma-2-2b-it
library_name: peft
license: gemma
pipeline_tag: text-generation
tags:
- lora
- peft
- transformers
- research
- alignment-tax
- safety-regression
- evaluation
datasets:
- sahil2801/CodeAlpaca-20k
---
# gemma-2-2b-it + CodeAlpaca-20k — a benign finetune that lost its safety refusals
**This is a research specimen, not a product.** It is a LoRA adapter for
`google/gemma-2-2b-it`, trained on ordinary coding instructions with no adversarial
intent, which **gained coding helpfulness and lost safety refusals as a side effect.**
It is published so that a claim made about it can be checked by anyone, byte for byte.
## What happened to it
Trained on CodeAlpaca-20k — a plain instruction-tuning dataset, nothing adversarial —
and then put through an independent behavioural promotion gate. The gate returned
**BLOCK**:
| domain | before | after | delta |
|---|---|---|---|
| Helpfulness | 12/20 | 19/20 | **+0.35** |
| Safety refusals | 25/25 | **16/25** | **−0.36** |
Overall: 10 gains, 20 regressions, 83 ties across a 113-item paired evaluation.
That is the **alignment tax**: the model got better at the thing it was trained on and
quietly worse at refusing things it should refuse. Nobody asked it to. This is what
routine finetuning can cost you, which is exactly why the specimen is worth having in
public.
**This is not an abliteration.** No refusal direction was targeted, removed, or
suppressed. The erosion is diffuse and incidental — and that turns out to matter a great
deal (see below).
## ⚠️ Safety
This adapter **degrades the safety behaviour of its base model**. It refuses fewer
harmful requests than `google/gemma-2-2b-it` does. Do not deploy it. It exists to be
measured, and it is published because a research claim about it should be verifiable
rather than taken on trust.
## Why it exists — a pre-registered experiment that returned ABSTAIN
This adapter is the "blocked candidate" of a pre-registered joint experiment: could a
**reversible, attested weight-space edit** remove the safety regression while keeping the
coding gains?
Before touching a single weight, a go/no-go was pinned: *is the safety erosion low-rank
and separable from the coding change?* It was measured, and the answer was **no**:
- the leading direction of the weight delta carries **4.78%** of its variance; the top
thirty-two carry **36.4%** — diffuse, not low-rank
- the refusal direction sits essentially orthogonal to it: **~98.7%** of it lies outside
the subspace this finetune actually moved
So the experiment **terminated at step 0 as a published abstain**, and that null result
was published as the first exhibit of the format rather than buried. It maps a real
boundary: on this class of finetune the damage is not stored anywhere a weight-space edit
can reach, and a behavioural gate is the instrument that reaches.
Full exhibit, receipt and replay protocol:
**https://github.com/Vulcora/proofora/tree/main/proof-carrying-edit/exhibits/model-c**
## Reproduce the measurement yourself
The exhibit's `REPLAY.md` gives a numpy-only script — no special software — that
recomputes the diffuseness result from **this adapter alone**, needing neither the base
model nor a merge, because the weight delta is just `(alpha/r)·B·A`. Expect:
```
leading-direction fraction : 0.0478
top-32 cumulative fraction : 0.3642
```
Byte pin for `adapter_model.safetensors`:
```
sha256:781ee2bf2ec7765f103622d5ff161d2b995c3de1a52e5fab3df729d7294fdaa5
```
## Training recipe (deterministic)
`n=3999`, `seed=0`, `r=32`, `alpha=64`, `dropout=0.05`, 3 epochs, lr `2e-4`, cosine with
`warmup_ratio=0.03`, `max_len=320`, effective batch 16, bf16, completion-only loss
(prompt tokens masked to `-100`), 750 optimizer steps. Targets `q/k/v/o_proj` and
`gate/up/down_proj`.
Note that bf16 GPU LoRA training is **not** bit-reproducible across hardware and library
versions, so a re-train reproduces the *regime*, not the third decimal. That is precisely
why this adapter is published: it is the exact artifact the receipt was computed from.
## License
Derivative of `google/gemma-2-2b-it`. **The [Gemma Terms of Use](https://ai.google.dev/gemma/terms)
apply to this adapter and to anything derived from it.** You are responsible for
accepting and complying with those terms.
## Citation
Published by [Vulcora](https://vulcora.se) as the blocked-candidate specimen for the
proof-carrying model-edit format — an open format for attesting a change to a model's
weights such that a third party can verify the claim without access to anyone's method.
|