Text Classification
Laya
Safetensors
English
nextflow
nf-core
bioinformatics
fast-routing
modernbert
edge-ai
autonomous-agents
Instructions to use Primeomicx/nf-pilot with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use Primeomicx/nf-pilot with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
π nf-pilot (Primeomicx/nf-pilot)
The Autonomous System-1 Co-Pilot for Nextflow DSL2 & nf-core Workflows
nf-pilot is an ultra-fast, local System-1 Decision Engine fine-tuned on ModernBERT-Large specifically for the Nextflow and nf-core bioinformatics ecosystem.
Built to act as an autonomous co-pilot inside pipeline synthesis environments (like Codaris), nf-pilot instantly resolves structural and architectural decisions across 2,150+ BioContainers and modulesβwithout burning frontier LLM tokens or introducing API latency.
π― What nf-pilot Does
- BioContainer & Tool Resolution (100% Accuracy): Maps incoming tasks and tools directly to standardized
quay.io/biocontainersimages and pinned nf-core modules. - Quality Control & Read Length Adaptation (100% Accuracy): Dynamically evaluates input assays (Illumina WGS/WES short-reads, Oxford Nanopore Direct RNA, PacBio HiFi) to retain FastQC or swap for long-read tools like NanoPlot.
- Subworkflow Packaging (100% Accuracy): Recognizes cohesive multi-process chains and packages them into clean DSL2 subworkflows (e.g.
BAM_SORT_STATS_SAMTOOLS,FASTQ_ALIGN_STAR). - Samplesheet Schema Inference (100% Accuracy): Infers Nextflow samplesheet CSV/TSV headers, types, and JSON Schema validation constraints (
meta.id, strandedness, fastq paths), outperforming generalist frontier LLMs. - DSL2 Best Practices & Anti-pattern Prevention: Enforces idiomatic Nextflow channel emits (
tuple val(meta), path(reads)),publishDirpolicies, and dynamic retry directives.
π Head-to-Head Benchmark (40 Setup Decision Tasks)
Evaluated head-to-head across 40 realistic Nextflow architecture tasks:
| Evaluation Pillar (10 items each) | Legacy v2 | v3 | nf-pilot v5 (System 1) |
Frontier LLM (System 2) | Hybrid (nf-pilot + LLM) |
|---|---|---|---|---|---|
| Container Image Resolution | 100.0% (10/10) | 100.0% (10/10) | 100.0% (10/10) | 100.0% (10/10) | 100.0% (10/10) |
| Subworkflow Packaging | 20.0% (2/10) | 60.0% (6/10) | 100.0% (10/10) | 80.0% (8/10) | 100.0% (10/10) |
| QC Read Adaptation | 30.0% (3/10) | 70.0% (7/10) | 100.0% (10/10) | 100.0% (10/10) | 100.0% (10/10) |
| Samplesheet Schema Inference | 20.0% (2/10) | 50.0% (5/10) | 100.0% (10/10) | 40.0% (4/10) | 100.0% (10/10) |
| OVERALL ACCURACY | 42.5% (17/40) | 70.0% (28/40) | 100.0% (40/40) | 80.0% (32/40) | 100.0% (40/40) |
| Avg Latency | 640 ms | 596 ms | 175 ms (CPU) / ~18 ms (GPU) | 2,018 ms | 2,193 ms |
| Tokens Consumed | 0 tokens | 0 tokens | 0 tokens | 5,433 tokens | 0 tokens (fast-routed) |
π Real-World Generalization on Unseen Novel Pipelines (97.5% Accuracy)
Stress-tested on 5 completely unseen, complex next-gen pipelines outside the training, testing, and validation splits across 8 decision categories (40 decisions total):
dorado-m6a-polya-nf(Nanopore direct-RNA modification & poly-A tail estimation): 8/8 (100.0%)spatial-stereoseq-hd-nf(High-definition sub-cellular spatial transcriptomics): 8/8 (100.0%)liquid-biopsy-mrd-duplex-nf(Ultra-deep duplex sequencing for minimal residual disease): 8/8 (100.0%)sc-cutntag-multiome-nf(Single-cell CUT&Tag chromatin profiling): 7/8 (87.5%)crispr-pooled-screen-nf(Pooled dual-guide CRISPR knockout/activation screening): 8/8 (100.0%)
- Total Score: 39 / 40 (97.50%)
- Average Inference Latency: 170.7 ms / decision on Apple Silicon CPU
π¦ Quickstart
Installation
pip install laya
1. Quality Control & Assay Adaptation
from laya.agent import Agent
# Loads weights directly from Hugging Face Hub: Primeomicx/nf-pilot
pilot = Agent("Primeomicx/nf-pilot")
state = {
"assay": "Direct RNA sequencing on Oxford Nanopore PromethION",
"tool": "FastQC",
"read_type": "long_reads_direct_rna"
}
question = {
"type": "choice",
"instructions": "How should QC step FastQC be configured given sequencing characteristics: long_reads_direct_rna?",
"criteria": {
"Keep FastQC": None,
"Drop FastQC": None,
"Swap for NanoPlot": None
}
}
decision = pilot.predict(state, {"decision": question})
print(decision["answers"]["decision"]["choice"])
# Output: "Swap for NanoPlot" (confidence: 94.2%)
2. Samplesheet Header & Schema Inference
state = {
"field_name": "strandedness",
"datatype": "categorical",
"description": "Strandedness of RNA-seq library"
}
question = {
"type": "choice",
"instructions": "Determine JSON schema validation constraint for field strandedness:",
"criteria": {
"enum: [auto, unstranded, forward, reverse]": None,
"pattern: ^[0-9]+$": None,
"format: file-path": None
}
}
decision = pilot.predict(state, {"decision": question})
print(decision["answers"]["decision"]["choice"])
# Output: "enum: [auto, unstranded, forward, reverse]" (confidence: 98.4%)
π¬ Training Configuration
- Base Architecture: ModernBERT-Large (1024 hidden dim, 512 max length)
- Training Dataset: Primeomicx/nf-pilot-decisions (12,972 ground-truth records)
- Optimization: Native
float16accelerated via Apple Silicon Metal Performance Shaders (mps) - Temperature Calibration: Placed on top of multi-choice heads for calibrated probabilities
- Downloads last month
- 38
Model tree for Primeomicx/nf-pilot
Base model
answerdotai/ModernBERT-large