PhoenixAgent / README.md
VarunRS5457
Reshape flow: redirect to classify page with AI pre-filled classifications
16d2e95
|
Raw
History Blame Contribute Delete
3.55 kB
---
title: Aseesa Ingestion Agent
emoji: 🧬
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 7860
---
# Aseesa Ingestion Agent
AI-powered pharma/life science data ingestion — maps SAS metadata to internal templates.
## Features
- **Stars**: Paired Data + Metadata output with AI-powered column classification
- **Stars (any layout)**: AI reshape of arbitrary/messy spreadsheets into the Stars
Data + Metadata format (see "Any-layout reshape" below)
- **Qvey**: Simplified flat CSV output with exclude-only column selector
- Multi-file upload with automatic join detection
- AI auto-classification using LLM (Groq, OpenRouter, Ollama, Bedrock, etc.)
- Undo/redo for classifications
- Data preview on classify page
- Download as ZIP
- Batch processing
## Any-layout reshape (Stars from messy files)
Real lab files rarely have a clean header in row 1. The **Stars (any layout)** mode
handles arbitrary spreadsheets: headers several rows down, interleaved summary rows
(`mean`/`SEM`), side-by-side blocks (e.g. Chow next to HFD on one sheet), and data
spread across multiple files. It merges everything into one Stars Data matrix +
Metadata table.
Flow: upload one or more `.xlsx`/`.xls`/`.csv` files → the agent reads the literal
cell grid of each sheet and **proposes a reshape plan** (which row/column holds the
sample IDs, which cells are measurements, the cohort per block) → you review and
correct the plan on the reshape page → generate.
Two sheet layouts are supported and merged by canonical `Sample_ID`
(`{genotype}-{sex}-{diet}-{n}`, with `Genotype`/`Sex`/`Diet` derived into the
metadata table):
- **samples-in-columns** (e.g. a glucose time-course: sample IDs across a row,
timepoints down a column). The TIME column and value rows are auto-detected, and
`%Basal` rows are computed from the raw timepoints.
- **samples-in-rows** (e.g. an insulin assay: each row is a sample like
`KO F Chow 1`, columns are measurements). Side-by-side Chow/HFD blocks become one
plan entry each; varied measurement-column spellings (`80X`/`80Fold`,
`ug Insulin`/`ug/ml`, ...) are normalized to canonical variable names.
Because the model only proposes coordinates (which you confirm) and the reshape math
is deterministic, output is reproducible and auditable. See
`scripts/validate_reshape.py` and `scripts/demo_reshape_offline.py` for validating a
generated Stars pair against a known-good MasterSheet without invoking the model.
## Environment Variables
Set these as **Secrets** in your HF Space settings (or in a local `.env`; see
`.env.example`):
| Variable | Required | Description |
|----------|----------|-------------|
| `LLM_API_KEY` | Yes | API key for the default LLM (Groq by default) |
| `LLM_MODEL` | No | Model name (default: `llama-3.3-70b-versatile`) |
| `LLM_BASE_URL` | No | API base URL (default: `https://api.groq.com/openai/v1`) |
| `SECURE_LLM_MODEL` | No | Model for a separate secure endpoint (optional) |
| `SECURE_LLM_API_KEY` | No | API key for the secure LLM |
| `SECURE_LLM_BASE_URL` | No | Base URL for the secure LLM |
Any OpenAI-compatible endpoint works via `LLM_BASE_URL` (Groq, OpenRouter, a local
Ollama/vLLM server, etc.). The default is Groq's `llama-3.3-70b-versatile`.
> **Privacy note:** the any-layout reshape flow sends raw cell values (not just
> column metadata) to the configured LLM so it can read the sheet structure. With a
> hosted API like Groq this means data leaves your machine; point `LLM_BASE_URL` at a
> self-hosted model if that matters for your data.