lvwerra HF Staff commited on
Commit
19eb242
·
1 Parent(s): 372db56

Describe the dataset at the top of both tabs (#15)

Browse files

- Describe the dataset at the top of both tabs (f731a99de33a347ad7d9ddfae785a9a55f2585cc)

Files changed (4) hide show
  1. README.md +2 -0
  2. app.css +8 -0
  3. app.py +23 -0
  4. content/about.md +10 -0
README.md CHANGED
@@ -17,6 +17,8 @@ Future features and open design decisions are tracked in the [feature backlog](F
17
 
18
  ## GenBank taxonomy
19
 
 
 
20
  The landing page opens with an interactive **tree of eukaryotic life**, using NCBI's current GenBank assembly summary and taxonomy. It is drawn as a Sankey diagram running **down the page**: each group is a horizontal bar split into annotated (green) and not-annotated (gray) assemblies, and flows carry both parts from a group into its lineages on the level below. Expanding a group adds a level underneath rather than a column to the right, so a full lineage — eukaryotes down to humans is about thirty levels — is one ordinary page scroll instead of an endless sideways one. The chart itself only scrolls on screens too narrow to keep the labels legible. Shortcut buttons above the tree jump straight to a lineage: animals, mammals, humans, birds, ray-finned fish, insects, jellyfish, green plants and fungi. Each group shows a silhouette, its scientific name and its English name; counts and percentages live in the tooltip, which appears over a group or over the flux arriving at it. The flux feeding the current selection stays highlighted, so the trunk you have walked reads against the lineages you passed over. Each group is labelled with its scientific name and, where one exists, a plain-English name underneath, so the Latin backbone of the tree (Opisthokonta, Eumetazoa, Ecdysozoa) stays readable without prior knowledge. The two views sit inside one curved panel, switched by a full-width toggle at the top: the **Genome Atlas** holds the tree and its lineage controls, and **Database** holds accession search, segment visualization, and downloads over the full published-file snapshot. A **Show database results** button carries the current group across, handing the Database tab one annotated assembly from the selected lineage; it says so plainly when a group has none. RefSeq exploration can be added beneath the tree.
21
 
22
  The viewer is restricted to **Eukaryota (NCBI taxid 2759) and its descendants**, including search, breadcrumbs, and parent navigation. The model predicts eukaryotic annotations; bacteria, archaea, and viruses are outside this display. The current eukaryotic snapshot contains **70,395 assemblies**, of which **17,561 (24.9%)** have published annotations. The underlying SQLite snapshot retains the full NCBI inventory for refresh/provenance, while the interface uses only the eukaryotic subtree and its denominators. Assemblies with unresolved taxonomy cannot be assigned to that subtree and are excluded.
 
17
 
18
  ## GenBank taxonomy
19
 
20
+ Both tabs open with a short description of the dataset, read from `content/about.md`. Edit that file to change the wording; no code change is needed. Text after a `<!-- more -->` line, if present, goes into a closed "How the annotations were made" disclosure.
21
+
22
  The landing page opens with an interactive **tree of eukaryotic life**, using NCBI's current GenBank assembly summary and taxonomy. It is drawn as a Sankey diagram running **down the page**: each group is a horizontal bar split into annotated (green) and not-annotated (gray) assemblies, and flows carry both parts from a group into its lineages on the level below. Expanding a group adds a level underneath rather than a column to the right, so a full lineage — eukaryotes down to humans is about thirty levels — is one ordinary page scroll instead of an endless sideways one. The chart itself only scrolls on screens too narrow to keep the labels legible. Shortcut buttons above the tree jump straight to a lineage: animals, mammals, humans, birds, ray-finned fish, insects, jellyfish, green plants and fungi. Each group shows a silhouette, its scientific name and its English name; counts and percentages live in the tooltip, which appears over a group or over the flux arriving at it. The flux feeding the current selection stays highlighted, so the trunk you have walked reads against the lineages you passed over. Each group is labelled with its scientific name and, where one exists, a plain-English name underneath, so the Latin backbone of the tree (Opisthokonta, Eumetazoa, Ecdysozoa) stays readable without prior knowledge. The two views sit inside one curved panel, switched by a full-width toggle at the top: the **Genome Atlas** holds the tree and its lineage controls, and **Database** holds accession search, segment visualization, and downloads over the full published-file snapshot. A **Show database results** button carries the current group across, handing the Database tab one annotated assembly from the selected lineage; it says so plainly when a group has none. RefSeq exploration can be added beneath the tree.
23
 
24
  The viewer is restricted to **Eukaryota (NCBI taxid 2759) and its descendants**, including search, breadcrumbs, and parent navigation. The model predicts eukaryotic annotations; bacteria, archaea, and viruses are outside this display. The current eukaryotic snapshot contains **70,395 assemblies**, of which **17,561 (24.9%)** have published annotations. The underlying SQLite snapshot retains the full NCBI inventory for refresh/provenance, while the interface uses only the eukaryotic subtree and its denominators. Assemblies with unresolved taxonomy cannot be assigned to that subtree and are excluded.
app.css CHANGED
@@ -82,3 +82,11 @@
82
  #db-prepare:disabled::before { content:""; display:inline-block; width:11px; height:11px; margin-right:8px; vertical-align:-1px;
83
  border:2px solid currentColor; border-right-color:transparent; border-radius:50%; animation:db-spin .8s linear infinite; }
84
  @keyframes db-spin { to { transform:rotate(360deg); } }
 
 
 
 
 
 
 
 
 
82
  #db-prepare:disabled::before { content:""; display:inline-block; width:11px; height:11px; margin-right:8px; vertical-align:-1px;
83
  border:2px solid currentColor; border-right-color:transparent; border-radius:50%; animation:db-spin .8s linear infinite; }
84
  @keyframes db-spin { to { transform:rotate(360deg); } }
85
+ /* What was annotated and how: a lead and a closed disclosure at the top of
86
+ both tabs, aligned with the content padding below it. */
87
+ .gradio-container .tab-intro { padding:22px 20px 0; gap:8px; background:transparent; border:0; }
88
+ .gradio-container .tab-intro-lead, .gradio-container .tab-intro-lead p { font:15px/1.65 Arial,Helvetica,sans-serif; color:#33503f; max-width:860px; margin:0; }
89
+ .gradio-container .tab-intro .atlas-disclosure { margin-top:4px; }
90
+ #atlas-overview .tab-intro + * .atlas, #atlas-overview .atlas { padding-top:8px; }
91
+ .gradio-container .tab-intro .atlas-disclosure ul { padding-left:22px; margin:6px 0 0; list-style:disc; }
92
+ .gradio-container .tab-intro .atlas-disclosure li { margin:3px 0; }
app.py CHANGED
@@ -21,6 +21,27 @@ from catalog import Catalog
21
  from remote_catalog import RemoteCatalog, RemoteReadError
22
 
23
  HIST_ROWS = 40
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24
  # Readable headers for the results table; the catalog keeps its own names.
25
  HEADERS = {"assembly_accession": "Assembly", "record_name": "Record", "organism_name": "Organism",
26
  "division": "Division", "segment_start_bp": "Start", "segment_end_bp": "End"}
@@ -165,8 +186,10 @@ def build_app(catalog=None):
165
  with gr.Blocks(title="GenBank Annotation Explorer", delete_cache=(3600, 3600)) as demo:
166
  with gr.Tabs(selected="atlas", elem_id="atlas-navigation") as navigation:
167
  with gr.Tab("Genome Atlas", id="atlas", elem_id="atlas-overview"):
 
168
  atlas = build_taxonomy_tab()
169
  with gr.Tab("Database", id="database", elem_id="atlas-database"):
 
170
  hits = gr.State([])
171
  selected = gr.State(None)
172
  with gr.Column(elem_classes="atlas-panel"):
 
21
  from remote_catalog import RemoteCatalog, RemoteReadError
22
 
23
  HIST_ROWS = 40
24
+ ABOUT = Path(__file__).parent / "content/about.md"
25
+ MORE = "<!-- more -->"
26
+
27
+
28
+ def tab_intro():
29
+ """What was annotated and how, at the top of a tab, from content/about.md.
30
+
31
+ The text lives in one Markdown file so both tabs say the same thing and
32
+ rewording it needs no code change. Above the "more" marker is a short lead;
33
+ below it, the method, in a closed disclosure.
34
+ """
35
+ if not ABOUT.exists():
36
+ return
37
+ lead, _, more = ABOUT.read_text().partition(MORE)
38
+ lead, more = (re.sub(r"<!--.*?-->", "", part, flags=re.S).strip() for part in (lead, more))
39
+ with gr.Column(elem_classes="tab-intro"):
40
+ if lead:
41
+ gr.Markdown(lead, elem_classes="tab-intro-lead")
42
+ if more:
43
+ with gr.Accordion("How the annotations were made", open=False, elem_classes="atlas-disclosure"):
44
+ gr.Markdown(more)
45
  # Readable headers for the results table; the catalog keeps its own names.
46
  HEADERS = {"assembly_accession": "Assembly", "record_name": "Record", "organism_name": "Organism",
47
  "division": "Division", "segment_start_bp": "Start", "segment_end_bp": "End"}
 
186
  with gr.Blocks(title="GenBank Annotation Explorer", delete_cache=(3600, 3600)) as demo:
187
  with gr.Tabs(selected="atlas", elem_id="atlas-navigation") as navigation:
188
  with gr.Tab("Genome Atlas", id="atlas", elem_id="atlas-overview"):
189
+ tab_intro()
190
  atlas = build_taxonomy_tab()
191
  with gr.Tab("Database", id="database", elem_id="atlas-database"):
192
+ tab_intro()
193
  hits = gr.State([])
194
  selected = gr.State(None)
195
  with gr.Column(elem_classes="atlas-panel"):
content/about.md ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ <!--
2
+ Shown at the top of both the Genome Atlas and the Database tab.
3
+
4
+ Everything above a "more" marker is always visible. Anything below one,
5
+ written as <!- - more - -> without the spaces, goes into a closed
6
+ "How the annotations were made" disclosure; with no marker there is no
7
+ disclosure. HTML comments like this one are stripped before rendering.
8
+ -->
9
+
10
+ The Carbon Annotation Database contains 566 million candidate protein-coding genes predicted by Carbon-A, an open 1.2-billion-parameter model that finds genes directly from DNA. It covers genomes from over 22,000 species, including animals, plants, fungi, and protists. Each prediction links to its source genome, genomic coordinates, coding DNA, predicted protein sequence, and confidence score, helping researchers explore poorly annotated genomes and prioritize candidates for further study.