codeguard-jev-style-1.5b

Local System One model for PR code-review triage. JEV-compatible typed verdicts, no API key required.

Drop-in local alternative to TypeSafe Jev for code-review pipelines. Returns verdict, risk level, and flags per PR โ€” swap the endpoint, keep the agent.


Overview

codeguard-jev-style-1.5b is a fine-tuned Qwen2.5-1.5B-Instruct that produces typed code-review verdicts from PR diffs โ€” matching the output contract of TypeSafe System One / Jev without requiring an API subscription.

It outputs a single JSON object per PR:

{
  "verdict": "approved | changes_requested | needs_review",
  "risk_level": "low | medium | high | critical",
  "summary": "<one sentence>",
  "flags": ["hardcoded-credentials", "missing-tests", ...]
}

No free-text review, no hallucinated comments, no reasoning trace. A typed decision the agent can branch on โ€” the same structure your CI pipeline already consumes from Jev.


Why local?

Jev (TypeSafe API) codeguard-jev-style-1.5b
Typed output yes yes
Latency ~130ms (network) ~85ms (local, MPS)
Cost per-request billing free after download
Diff stays in org no yes
Fine-tune on your codebase no yes (LoRA)
Accuracy (CodeReview benchmark) 94.3% F1 93.8% F1
Critical-flag precision (holdout, n=3,600) 91.2% 92.1%

Code diffs contain proprietary logic and credentials. Sending them to an external API is a data governance problem many security teams won't sign off on. Run it locally.


Quickstart

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch, json, re

tok = AutoTokenizer.from_pretrained("enterprise-ai-lab/codeguard-jev-style-1.5b")
model = AutoModelForCausalLM.from_pretrained(
    "enterprise-ai-lab/codeguard-jev-style-1.5b", dtype=torch.float32
).eval()

SYS = (
    "You are a senior code-review triage assistant. For each pull request diff output ONE "
    "JSON object with keys: verdict (approved|changes_requested|needs_review), "
    "risk_level (low|medium|high|critical), summary (one sentence), flags (list of strings). "
    "Output only the JSON."
)

def review_pr(title, diff):
    content = f"PR Title: {title}\n\nDiff:\n{diff}"
    msgs = [{"role": "system", "content": SYS}, {"role": "user", "content": content}]
    prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
    ids = tok(prompt, return_tensors="pt")
    with torch.no_grad():
        out = model.generate(**ids, max_new_tokens=150, do_sample=False,
                             pad_token_id=tok.eos_token_id)
    text = tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True)
    m = re.search(r"\{.*\}", text, re.DOTALL)
    return json.loads(m.group(0)) if m else {}

result = review_pr(
    "Fix null check in user service",
    "- return user.name\n+ return user.name if user else None"
)
# {"verdict": "approved", "risk_level": "low", "summary": "Safe null guard.", "flags": []}

GitHub Actions integration

- name: AI Code Review Gate
  run: |
    python codeguard_check.py \
      --diff "${{ github.event.pull_request.diff_url }}" \
      --block-on critical,high

Or wire it as an OpenAI-compatible endpoint and point your existing Jev agent at localhost:8000:

# Before (Jev):
llm = ChatOpenAI(base_url="https://api.typesafe.ai/v1", api_key=JEV_KEY, model="jev-latest")
# After (CodeGuard local):
llm = ChatOpenAI(base_url="http://localhost:8000/v1", api_key="not-needed", model="codeguard-jev-style-1.5b")

Training

Fine-tuned with LoRA (r=16, alpha=32) on a curated code-review corpus:

  • 61,400 pull requests from open-source and enterprise repositories (2020-2024)
  • Annotation: senior engineer verdicts + SAST tool output as ground truth
  • Categories: security flags, missing tests, breaking changes, style, dependency issues
  • Held-out validation: 3,600 PRs stratified by risk level
  • Training: 12 epochs, AdamW lr=2e-4, MPS/CUDA

Benchmarks

CodeReview-Bench (internal holdout, n=3,600)

Model Precision Recall F1
codeguard-jev-style-1.5b 94.2% 93.4% 93.8%
Jev (TypeSafe API) 94.8% 93.9% 94.3%
GPT-4o-mini (zero-shot) 89.1% 88.4% 88.7%
CodeBERT classifier 82.3% 81.7% 82.0%

Critical-flag Precision (security/credential findings)

Model Precision False Positive Rate
codeguard-jev-style-1.5b 92.1% 7.9%
Jev (TypeSafe API) 91.2% 8.8%
GPT-4o-mini (zero-shot) 84.6% 15.4%

Intended use

  • CI/CD PR gates (approve/block before merge)
  • Security-focused triage (flag credentials, injections, missing auth)
  • Developer productivity (pre-review before human review)
  • Air-gapped environments where diffs cannot leave the network

Limitations

  • Trained on English comments and common languages (Python, JS, Go, Java, Rust)
  • Not a replacement for SAST tools โ€” complements them
  • Context window limits very large diffs; chunk at 2,000 lines

License

Apache 2.0. Base model (Qwen2.5-1.5B-Instruct) is subject to its own Qwen license.


Downloads last month
18,215
Safetensors
Model size
2B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for JackKozmo29/codeguard-jev-style-1.5b

Finetuned
(1950)
this model

Space using JackKozmo29/codeguard-jev-style-1.5b 1