Buckets:

cgeorgiaw's picture
|
download
raw
2.15 kB

GenBank Annotation Output Schema

Published final annotations live under:

annotations/<division>/<shard>.parquet

Rows are one row per source contig.

Column Type Purpose
source_key string Stable local key: `
division string GenBank division/group used by the pipeline
record_name string FASTA/GenBank contig record ID
assembly_accession string GenBank assembly accession
organism_name string Assembly organism name from GenBank metadata
taxid int64 NCBI taxonomy ID
genome_size int64 Assembly genome size from GenBank metadata
ftp_path string NCBI FTP directory for the assembly; enough to recover the source FASTA path
aligned_bp_length int64 Number of bp represented by the probability arrays after 6-mer alignment
annotation_run_id string Pipeline run identifier used when unpacking/publishing annotations
annotation_run_timestamp_utc string UTC timestamp associated with the annotation run
annotation_git_commit string Git commit of this pipeline when annotations were unpacked/submitted
manifest_id string Manifest/run identifier for the source assembly selection
sequence string Source contig sequence prefix represented by the probability arrays. Optional; omitted for Hub-scale outputs by default.
pred_prob_positive_strand_cds list Per-bp P(CDS) on positive strand, aligned to source sequence prefix
pred_prob_negative_strand_cds list Per-bp P(CDS) on negative strand, aligned to source sequence prefix

Intentionally omitted from Hub-scale final outputs:

  • raw packed blocks and raw packed-block probability outputs.
  • sequence, unless the run was unpacked without --omit-sequence.

Mapping back to source FASTA:

  1. Use ftp_path to locate the assembly directory.
  2. The FASTA file basename is the final path component of ftp_path plus _genomic.fna.gz.
  3. Use record_name to find the contig in that FASTA.
  4. Probability arrays cover the first aligned_bp_length bp of that contig.

Xet Storage Details

Size:
2.15 kB
·
Xet hash:
5603bf46c22e2bfc9a7efd20d34ff70039070fbee427e7318d2431eb429a4af2

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.