GenomeOcean-Sentinel

This is the distilled GenomeOcean-100M v1.2 observer used by GenomeOcean-Sentinel (GOS) to help assess whether DNA sequences appear AI-generated or natural. It includes all weights, configuration, tokenizer files, and modeling code; loading requires no separate base-model download.

This model is the observer, not the full detector. GOS combines its score with 47 CPU features using decision_head.json and the project's feature extraction and decision code. Applying a sigmoid or a 0.5 threshold to the observer alone does not reproduce GOS.

Usage

Loading requires trust_remote_code=True, which loads the included GOSStudentForObserver class through AutoModel.

Tested with Python 3.12.3, PyTorch 2.10.0, and Transformers 4.56.2.

pip install 'torch>=2.6,<3' 'transformers==4.56.2'
import torch
from transformers import AutoModel, AutoTokenizer

repo = "DOEJGI/GenomeOcean-Sentinel"
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
tokenizer = AutoTokenizer.from_pretrained(repo)

# Replace with your existing input strings; this example uses placeholder text.
inputs = tokenizer(
    ["opaque input string"], return_tensors="pt",
    padding="max_length", truncation=True, max_length=model.config.max_length,
    return_token_type_ids=False,
)
with torch.inference_mode():
    normalized_score = model(**inputs)  # shape (1,)
    observer_logit = normalized_score * model.config.target_std + model.config.target_mean
print(normalized_score.shape, observer_logit)

The default load uses float32. For GOS CUDA bfloat16 inference, keep the model in float32, move the model and inputs to CUDA, and run the forward pass inside torch.autocast("cuda", dtype=torch.bfloat16). Device and precision can affect scores.

Model and output

The model has 116,411,905 parameters: a Mistral backbone with causal attention, attention-mask mean pooling, and a Linear(768, 1) observation head. GOS truncates and pads observer inputs to 200 tokenizer tokens; longer inputs are outside the observer's trained input length.

The forward computation is:

hidden = backbone(input_ids=input_ids, attention_mask=attention_mask, use_cache=False)[0]
mask = attention_mask.to(hidden.dtype).unsqueeze(-1)
pooled = (hidden * mask).sum(1) / mask.sum(1).clamp_min(1.0)
score = head(pooled.float()).squeeze(-1).float()

The result is a float32 tensor of shape (batch,): one normalized observer score per input, not a probability or a full detector decision. If no attention mask is supplied, the model infers it from padding ID 3.

To recover the observer logit, use score * target_std + target_mean, with target_mean = -2.191271897027036e-06 and target_std = 11.3706368339268 from the model config.

Full GOS detector

The observer receives the first 200 tokenizer tokens. The full detector also extracts 47 CPU features from eligible portions of each record. decision_head.json contains the CPU-feature coefficients and calibration; the learned observation head is already included in model.safetensors. Use the GOS project to run feature extraction and combine these contributions:

observer_logit = normalized_score * target_std + target_mean
total_logit = cpu_feature_logit + observer_logit
router_confidence = sigmoid(total_logit)
call = AI if total_logit >= threshold_logit else Natural

Here, threshold_logit comes from decision_head.json under calibration. It is inherited production calibration, not a new independent false-positive-rate certification for this observer or for arbitrary inputs.

Limitations and license

The example demonstrates observer inference only; it does not establish detection performance. GOS decisions require the normalization and CPU-feature fusion above. Truncation, input distribution, and numerical precision can change results. No new training-data audit, held-out evaluation, or generalization guarantee is provided.

Non-Commercial Use Only. Copyright (c) 2026, The Regents of the University of California, through Lawrence Berkeley National Laboratory. See LICENSE for the full terms and commercial licensing details. This software was developed under DOE Contract No. DE-AC02-05CH11231; see NOTICE. Commercial licensing inquiries: IPO@lbl.gov.

Downloads last month
52
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including DOEJGI/GenomeOcean-Sentinel