Instructions to use DOEJGI/GenomeOcean-Sentinel with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DOEJGI/GenomeOcean-Sentinel with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="DOEJGI/GenomeOcean-Sentinel", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("DOEJGI/GenomeOcean-Sentinel", trust_remote_code=True) model = AutoModel.from_pretrained("DOEJGI/GenomeOcean-Sentinel", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
GenomeOcean-Sentinel
This is the distilled GenomeOcean-100M v1.2 observer used by GenomeOcean-Sentinel (GOS) to help assess whether DNA sequences appear AI-generated or natural. It includes all weights, configuration, tokenizer files, and modeling code; loading requires no separate base-model download.
This model is the observer, not the full detector. GOS combines its score
with 47 CPU features using decision_head.json and the project's feature
extraction and decision code. Applying a sigmoid or a 0.5 threshold to the
observer alone does not reproduce GOS.
Usage
Loading requires trust_remote_code=True, which loads the included
GOSStudentForObserver class through AutoModel.
Tested with Python 3.12.3, PyTorch 2.10.0, and Transformers 4.56.2.
pip install 'torch>=2.6,<3' 'transformers==4.56.2'
import torch
from transformers import AutoModel, AutoTokenizer
repo = "DOEJGI/GenomeOcean-Sentinel"
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
tokenizer = AutoTokenizer.from_pretrained(repo)
# Replace with your existing input strings; this example uses placeholder text.
inputs = tokenizer(
["opaque input string"], return_tensors="pt",
padding="max_length", truncation=True, max_length=model.config.max_length,
return_token_type_ids=False,
)
with torch.inference_mode():
normalized_score = model(**inputs) # shape (1,)
observer_logit = normalized_score * model.config.target_std + model.config.target_mean
print(normalized_score.shape, observer_logit)
The default load uses float32. For GOS CUDA bfloat16 inference, keep the model
in float32, move the model and inputs to CUDA, and run the forward pass inside
torch.autocast("cuda", dtype=torch.bfloat16). Device and precision can affect
scores.
Model and output
The model has 116,411,905 parameters: a Mistral backbone with causal attention,
attention-mask mean pooling, and a Linear(768, 1) observation head. GOS
truncates and pads observer inputs to 200 tokenizer tokens; longer inputs
are outside the observer's trained input length.
The forward computation is:
hidden = backbone(input_ids=input_ids, attention_mask=attention_mask, use_cache=False)[0]
mask = attention_mask.to(hidden.dtype).unsqueeze(-1)
pooled = (hidden * mask).sum(1) / mask.sum(1).clamp_min(1.0)
score = head(pooled.float()).squeeze(-1).float()
The result is a float32 tensor of shape (batch,): one normalized observer
score per input, not a probability or a full detector decision. If no attention
mask is supplied, the model infers it from padding ID 3.
To recover the observer logit, use score * target_std + target_mean, with
target_mean = -2.191271897027036e-06 and target_std = 11.3706368339268
from the model config.
Full GOS detector
The observer receives the first 200 tokenizer tokens. The full detector also
extracts 47 CPU features from eligible portions of each record.
decision_head.json contains the CPU-feature coefficients and calibration;
the learned observation head is already included in model.safetensors.
Use the GOS project to run feature extraction
and combine these contributions:
observer_logit = normalized_score * target_std + target_mean
total_logit = cpu_feature_logit + observer_logit
router_confidence = sigmoid(total_logit)
call = AI if total_logit >= threshold_logit else Natural
Here, threshold_logit comes from decision_head.json under calibration.
It is inherited production calibration, not a new independent false-positive-rate
certification for this observer or for arbitrary inputs.
Limitations and license
The example demonstrates observer inference only; it does not establish detection performance. GOS decisions require the normalization and CPU-feature fusion above. Truncation, input distribution, and numerical precision can change results. No new training-data audit, held-out evaluation, or generalization guarantee is provided.
Non-Commercial Use Only. Copyright (c) 2026, The Regents of the University of California, through Lawrence Berkeley National Laboratory. See LICENSE for the full terms and commercial licensing details. This software was developed under DOE Contract No. DE-AC02-05CH11231; see NOTICE. Commercial licensing inquiries: IPO@lbl.gov.
- Downloads last month
- 52