carbon-a-database-explorer / FEATURE_BACKLOG.md
lvwerra's picture
lvwerra HF Staff
Fold the views into one panel, right-size the controls, and open Database from a group (#7)
be172fd
|
Raw History Blame Contribute Delete
13 kB
# Feature backlog
Track desired features that are not yet implemented. Add new ideas here and update each item's status as work progresses. These entries describe future work; implementation details and priorities remain open unless explicitly agreed.
Last updated: 2026-10-05.
## Overview
| ID | Feature | Status |
| --- | --- | --- |
| F001 | GenBank inference coverage map | Assembly-presence colors implemented; base completeness and history pending |
| F002 | Taxonomy map and navigation | Taxonomy and coverage colors implemented; annotation links pending |
| F003 | Comparison with RefSeq annotations | Planned; not implemented |
| F004 | Download annotations by accession, individually or in bulk | Partial; individual indexed segment downloads exist |
| F005 | Fast accession lookup and retrieval across the full dataset | Full published-snapshot index implemented; automatic refresh and pagination pending |
| F006 | Schema validation and unsupported-file reporting | Planned; not implemented |
## F001 — GenBank inference coverage map
Current implementation: the landing-page eukaryotic tree colors each group by the fraction of current GenBank assembly versions with published annotations. Visible labels and hover show counts and percentages; partial assemblies count as present, and this is explicitly distinguished from base coverage. The display and its denominator are restricted to Eukaryota, matching the model's scope. The full refresh derives assembly IDs directly from the scalar metadata of every indexed published Parquet object and deduplicates exact assembly versions. `build_full_index.py` caches object metadata by content hash; `coverage_from_index.py` derives the assembly inventory; `refresh_taxonomy.py` rolls coverage into SQLite. A coverage-only marker-based refresh remains available via `refresh_coverage.py`. Refreshes are manual. See [coverage refresh and evidence limitations](README.md#refresh-annotation-coverage).
Remaining: partial versus complete assembly/base coverage, automatic updates and history, and navigation to available annotations. The existing completion markers do not bind input content hashes; stronger publication manifests would make the inventory independently auditable.
**Goal:** Show how much of GenBank we currently have annotations for, and how that coverage grows as more inference completes.
Desired behavior:
- Provide a visual map or overview of coverage, with a way to explore covered and uncovered groups.
- Show counts and proportions for annotated assemblies and bases, with the underlying totals visible.
- Distinguish partial coverage from complete coverage, and distinguish the full annotation collection from the small sample loaded by the prototype.
- Refresh the map as new inference outputs are published. Display when coverage was last updated.
- Let users navigate from the overview to available assemblies and their annotations.
Decisions and dependencies for implementation:
- Define the GenBank inventory or snapshot used as the denominator, including any eligibility filters. Label these explicitly so the map's meaning stays clear as GenBank changes.
- Define what counts as complete inference, including how to handle missing contigs, segments, or bases.
- Build a coverage inventory from published outputs and inference metadata, deduplicating assembly versions and repeated runs.
- Choose the visualization and refresh mechanism. Historical progress views are an optional extension.
## F002 — Taxonomy map and navigation
Current implementation: a server-rendered SVG Sankey of the taxonomy leads the page, above annotation lookup. It supports click/keyboard expansion into columns that keep every ancestor visible, breadcrumbs, parent navigation, shortcuts, and scientific-name/taxon-ID search. Every route is restricted to Eukaryota. Coverage colors, counts, and percentages use a separate SQLite snapshot; see [refresh instructions](README.md#refresh-annotation-coverage). Linking taxa to annotation accessions remains future work.
**Goal:** Provide a taxonomy-based view of the available genomes and annotations.
Desired behavior:
- Explore organisms through an interactive taxonomy map, with navigation from broad groups to individual taxa and assemblies.
- Show how many assemblies or annotations are available within a selected group.
- Let users open accession-level results from the map.
- Consider sharing coverage information with F001 so users can see which taxonomic groups have been processed.
Decisions and dependencies for implementation:
- Implemented: eukaryotic tree with bounded subtrees and exact, expandable “Other lineages” aggregates; NCBI taxonomy with source checksums and dates; merged IDs resolved. Unresolved taxonomy is excluded from the eukaryotic view.
- Implemented: English names beside every scientific name, from NCBI Taxonomy common names plus a curated gloss table (`refresh_common_names.py`).
- Implemented: a **Show database results** jump from the selected group into the Database tab, backed by the `assemblies` table from `refresh_assemblies.py`. It hands over one annotated accession; listing every assembly in a group, and distinguishing partial from complete base coverage, remain open.
- Implemented: clade silhouettes from PhyloPic, filtered to CC0/Public Domain Mark/CC BY at build time (`refresh_icons.py`), with every CC BY artist credited in the interface.
## F003 — Comparison with RefSeq annotations
**Goal:** For sequences with RefSeq annotations, show where our predictions agree with or differ from those annotations.
Desired behavior:
- Display RefSeq annotation tracks alongside our strand-specific CDS probabilities in the same genomic region.
- Highlight differences, with candidate categories including predicted coding regions outside RefSeq CDS features, RefSeq CDS regions with low predicted probability, and boundary or strand disagreements.
- Let users inspect the underlying RefSeq features and prediction values for each difference.
- Show the RefSeq annotation version and inference provenance used in the comparison.
- Make unavailable or incompatible comparisons explicit.
Decisions and dependencies for implementation:
- Resolve GenBank records to the corresponding RefSeq assemblies and contigs, checking sequence versions and coordinate compatibility before overlaying tracks.
- Decide how to handle sequences that require coordinate mapping rather than a direct overlay.
- Define how probability tracks become comparable predicted intervals, including thresholds, strand handling, and boundary tolerance.
- Select comparison metrics and feature scope, including how multiple transcripts and overlapping CDS features are treated.
- Describe differences as disagreements with the reference annotations; a disagreement alone does not establish which annotation is correct.
## F004 — Download annotations by accession
**Goal:** Make it easy for users to download annotations for the assemblies or contigs they are interested in.
Current baseline: the prototype can download one selected indexed segment as Parquet, fetched from the bucket on demand. Filenames include assembly accession, record name, and coordinates. Complete accession-level and bulk downloads are future work.
Desired behavior:
- Offer a download action directly from accession search results and assembly or contig views.
- Let users select multiple accessions or paste/upload an accession list, then download the matching annotations together.
- Include all available segments for a selected contig, and all available contigs for a selected assembly. Clearly flag incomplete coverage.
- Show which requested accessions matched, were unavailable, or need a version choice. Record the exact versions included.
- Include metadata, coordinate conventions, inference provenance, and a manifest of the exported contents.
- Show estimated download size and progress; support preparing large exports without blocking browsing.
Decisions and dependencies for implementation:
- Choose export formats and packaging. Preserve original per-base probabilities in Parquet; consider additional formats for derived intervals once their interpretation is defined.
- Define handling of duplicates, multiple inference runs, accession versions, and partial failures in bulk requests.
- Decide whether sequence is included when available, and label its availability explicitly.
- Define export size limits, temporary storage, and download expiration.
- Use F005 to locate and retrieve requested records efficiently across the full collection.
## F005 — Fast lookup and retrieval across the full dataset
**Goal:** Keep accession searches and annotation retrieval responsive as the app expands beyond its bundled sample.
Current baseline: the Database uses a compact, read-only SQLite index covering every published annotation Parquet object in its dated inventory. Shared assembly/provenance metadata and indexed accession fields keep the database small. It fetches selected row groups on demand, checks source content hashes, and uses a bounded cache. Whole-contig and segmented records are supported. The full builder saves resumable checkpoints and reuses unchanged objects on refresh.
Remaining work includes automatic refresh/publication, pagination beyond 200 matches, and reducing large-row-group transfer costs.
Desired behavior:
- Resolve assembly and contig accessions through a persistent metadata index, without scanning every annotation shard for each query.
- Map each match to its exact accession version, inference run, bucket object, row group, and segment coordinates.
- Retrieve only the data needed for the selected records or region, where the storage layout supports it.
- Support bulk resolution for F004, and paginate large assembly results.
- Update the index as inference outputs are published, showing the indexed snapshot or refresh time.
- Distinguish an accession absent from the indexed inventory from an incomplete or stale index and from a retrieval failure.
Decisions and dependencies for implementation:
- Benchmark lookup latency, data transfer, and memory use on representative workloads before selecting an index backend or changing the storage layout.
- Evaluate an accession-to-object/row-group index and metadata queries separately from fetching the large probability arrays.
- Check remote range-read support and actual row-group sizes. A coordinate index alone does not guarantee efficient region reads; consider smaller genomic chunks or another layout if needed.
- Define incremental indexing and consistent publication so index entries refer to available, identifiable object versions. Handle replacements and repeated inference runs explicitly.
- Consider bounded caching of frequently accessed metadata and annotation chunks, with invalidation tied to data versions.
- Share versioned inventory metadata with F001 and F002 where useful, keeping coverage summaries and lookup results consistent.
## F006 — Schema validation and unsupported-file reporting
**Goal:** Identify incompatible annotation files during indexing and make their effect on searchable coverage clear.
Current baseline: the viewer supports explicit segment coordinates and older whole-contig records, deriving coordinates from `aligned_bp_length` for the latter. Downloads preserve original schemas. This is limited normalization, not a general schema migration system.
Desired behavior:
- Identify and record each file's schema variant during indexing.
- Validate required accession fields, coordinate fields, probability column names, and compatible types before accepting a file.
- Report unsupported or malformed files with actionable reasons, including missing or renamed fields and incompatible types.
- Distinguish successfully indexed, unsupported, and failed files in coverage summaries; never present excluded files as absent from the bucket.
- Define explicit, versioned adapters for additional supported schema variants, with tests for normalization and preservation of original downloads.
- Separate metadata checks from probability-array validation that requires reading data; report what was actually validated.
## Adding and updating entries
Give each new feature a stable ID, a goal, a status, desired behavior, and open decisions or dependencies. Link implementation work when it starts. Once shipped, mark the entry implemented and record any remaining scope rather than deleting its history.
## Atlas research sections
Planned: add RefSeq annotation exploration beneath the eukaryotic tree in the **Genome Atlas** tab, and a **Wet Lab** view for experimental results. The Wet Lab placeholder tab has been removed for now; restore it when experimental results are ready. Decide on the comparison views and experimental summaries when those data are ready. Accession lookup, segment viewing, and downloads now live in the separate **Database** tab.