Spaces:
Sleeping
Sleeping
DataPilot AI Architecture
Design goals
DataPilot AI separates the user experience, orchestration, analytics, model training, persistence, and optional restricted computation. The default application remains useful without a paid model API; every displayed number originates from deterministic computation.
Runtime components
| Component | Responsibility |
|---|---|
| Streamlit UI | Recruiter demo, upload/sample selection, charts, trace, downloads and Q&A |
| FastAPI | Versioned analysis, run, artifact and evidence-backed Q&A endpoints |
| LangGraph | Stateful agent ordering, state propagation and critic retry routing |
| Analytics core | DuckDB profiling, statistical summaries and data-quality evidence |
| ML core | Leakage-safe sklearn pipelines, model comparison and validation |
| Explainability | SHAP when compatible; permutation-importance fallback |
| Persistence | SQLAlchemy with SQLite locally and PostgreSQL via DATABASE_URL |
| Artifact store | Local filesystem interface, replaceable by S3/MinIO |
| Restricted worker | Expression-only AST validation in a networkless, resource-limited container |
Agent graph
flowchart TD
A["Dataset + target"] --> B["Data Quality Agent"]
B --> C["EDA Agent (DuckDB)"]
C --> D["Statistical Analysis Agent"]
D --> E["Planning Agent"]
E --> F["Feature Engineering Agent"]
F --> G["Modeling Agent"]
G --> H{"Evaluation / Critic Agent"}
H -->|"Reject: weak or unstable"| G
H -->|"Approve"| I["Explainability Agent"]
I --> J["Executive Insights Agent"]
J --> K["Report + model card + pipeline + evidence"]
Leakage controls
- Rows without labels and exact duplicates are removed before splitting.
- Train/test split occurs before any learned transformation.
- Imputation, scaling, and one-hot encoding are inside
sklearn.pipeline.Pipeline. - Cross-validation refits the complete pipeline in every fold.
- Identifier and target-like features are flagged for human review.
- Holdout metrics and cross-validation metrics remain distinct.
Graceful degradation
- Without Gemini: deterministic evidence-backed executive narrative.
- Without SHAP: permutation importance.
- Without XGBoost: sklearn candidate models.
- Without MLflow: structured agent trace plus persisted run JSON.
- Without PostgreSQL: SQLite.