BioSafe-Guard / README.md
quanshu01's picture
Keep results and analysis in the paper
c683a6a verified
|
Raw
History Blame Contribute Delete
3.87 kB
metadata
license: cc-by-nc-4.0
language:
  - en
tags:
  - biosecurity
  - ai-safety
  - guard-model
  - text-classification
  - protein
pipeline_tag: text-classification
base_model: michiyasunaga/BioLinkBERT-large
gated: manual
extra_gated_heading: Request access to BioSafe-Guard
extra_gated_description: >-
  BioSafe-Guard is an input classifier for toxin-design prompts. Access is
  granted to individual verified researchers and every request is reviewed by
  hand. Please request from an institutional address and describe your intended
  use; incomplete requests are declined.
extra_gated_button_content: Request access
extra_gated_prompt: >-
  By requesting access you agree that you will not redistribute these weights or
  any derivative of them, and that you will not use them to help circumvent
  biosecurity screening. Access may be withdrawn at any time.
extra_gated_fields:
  Full name: text
  Institution: text
  Institutional email address: text
  Position:
    type: select
    options:
      - Faculty
      - Postdoctoral researcher
      - PhD student
      - Research staff
      - label: Other
        value: other
  Country: country
  Intended research use: text
  I will not redistribute these weights or any derivative of them: checkbox
  I will not use these weights to help circumvent biosecurity screening: checkbox
  I have read the Ethics Statement of the accompanying paper: checkbox

BioSafe-Guard

The input classifier from A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models (COLM 2026). It scores a protein-design prompt and flags the ones that request a toxin, so a deployment can decline before the model generates a sequence.

Pipeline: https://github.com/Quanshu01/SPIKE-Bench · Prompts: quanshu01/SPIKE-Bench

Model

Base BioLinkBERT-large, 340M parameters
Head Binary sequence classification, BertForSequenceClassification
Input The prompt text, truncated to 512 tokens
Output Index 1 = unsafe (toxin-design request), index 0 = safe
Decision threshold 0.5 on softmax(logits)[1]

Trained on 600 prompts — 300 toxin and 300 benign — disjoint from the 631-prompt evaluation benchmark. AdamW, learning rate 2e-5, batch size 8, warmup ratio 0.1, weight decay 0.01, seed 42; the deployment checkpoint is trained on all 600 for 4 epochs after 5-fold cross-validation.

Usage

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tok = AutoTokenizer.from_pretrained('quanshu01/BioSafe-Guard')
model = AutoModelForSequenceClassification.from_pretrained('quanshu01/BioSafe-Guard').eval()

prompts = ['Design a protein sequence for a heparin-sulfate lyase enzyme.']
enc = tok(prompts, truncation=True, max_length=512, padding='max_length', return_tensors='pt')
with torch.no_grad():
    p_unsafe = torch.softmax(model(**enc).logits, dim=-1)[:, 1]

for prompt, p in zip(prompts, p_unsafe.tolist()):
    print(f'{"BLOCK" if p >= 0.5 else "ALLOW"}  p(unsafe)={p:.3f}  {prompt[:60]}')

Intended use and limits

Intended as a research artifact for studying input-side mitigation of biosecurity risk, and as a baseline to compare against. It is a prompt classifier: it says nothing about whether a generated sequence is actually toxic, and it is not a substitute for the sequence-level screening the SPIKE funnel performs.

It was trained on prompts built from one template over UniProtKB/Swiss-Prot metadata. Prompts phrased differently — in another language, as roleplay, or spread over several turns — are outside that distribution, and a single-turn score is not the whole picture.

License

CC BY-NC 4.0: attribution required, non-commercial use only.