PhoenixAgent / README.md
VarunRS5457
Reshape flow: redirect to classify page with AI pre-filled classifications
16d2e95
|
Raw
History Blame Contribute Delete
3.55 kB
metadata
title: Aseesa Ingestion Agent
emoji: 🧬
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 7860

Aseesa Ingestion Agent

AI-powered pharma/life science data ingestion — maps SAS metadata to internal templates.

Features

  • Stars: Paired Data + Metadata output with AI-powered column classification
  • Stars (any layout): AI reshape of arbitrary/messy spreadsheets into the Stars Data + Metadata format (see "Any-layout reshape" below)
  • Qvey: Simplified flat CSV output with exclude-only column selector
  • Multi-file upload with automatic join detection
  • AI auto-classification using LLM (Groq, OpenRouter, Ollama, Bedrock, etc.)
  • Undo/redo for classifications
  • Data preview on classify page
  • Download as ZIP
  • Batch processing

Any-layout reshape (Stars from messy files)

Real lab files rarely have a clean header in row 1. The Stars (any layout) mode handles arbitrary spreadsheets: headers several rows down, interleaved summary rows (mean/SEM), side-by-side blocks (e.g. Chow next to HFD on one sheet), and data spread across multiple files. It merges everything into one Stars Data matrix + Metadata table.

Flow: upload one or more .xlsx/.xls/.csv files → the agent reads the literal cell grid of each sheet and proposes a reshape plan (which row/column holds the sample IDs, which cells are measurements, the cohort per block) → you review and correct the plan on the reshape page → generate.

Two sheet layouts are supported and merged by canonical Sample_ID ({genotype}-{sex}-{diet}-{n}, with Genotype/Sex/Diet derived into the metadata table):

  • samples-in-columns (e.g. a glucose time-course: sample IDs across a row, timepoints down a column). The TIME column and value rows are auto-detected, and %Basal rows are computed from the raw timepoints.
  • samples-in-rows (e.g. an insulin assay: each row is a sample like KO F Chow 1, columns are measurements). Side-by-side Chow/HFD blocks become one plan entry each; varied measurement-column spellings (80X/80Fold, ug Insulin/ug/ml, ...) are normalized to canonical variable names.

Because the model only proposes coordinates (which you confirm) and the reshape math is deterministic, output is reproducible and auditable. See scripts/validate_reshape.py and scripts/demo_reshape_offline.py for validating a generated Stars pair against a known-good MasterSheet without invoking the model.

Environment Variables

Set these as Secrets in your HF Space settings (or in a local .env; see .env.example):

Variable Required Description
LLM_API_KEY Yes API key for the default LLM (Groq by default)
LLM_MODEL No Model name (default: llama-3.3-70b-versatile)
LLM_BASE_URL No API base URL (default: https://api.groq.com/openai/v1)
SECURE_LLM_MODEL No Model for a separate secure endpoint (optional)
SECURE_LLM_API_KEY No API key for the secure LLM
SECURE_LLM_BASE_URL No Base URL for the secure LLM

Any OpenAI-compatible endpoint works via LLM_BASE_URL (Groq, OpenRouter, a local Ollama/vLLM server, etc.). The default is Groq's llama-3.3-70b-versatile.

Privacy note: the any-layout reshape flow sends raw cell values (not just column metadata) to the configured LLM so it can read the sheet structure. With a hosted API like Groq this means data leaves your machine; point LLM_BASE_URL at a self-hosted model if that matters for your data.