# CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. Author: Grigori Fursin (cTuning Labs) ## What this is FlexBoard is a Gradio web app that helps users find optimal AI-inference hardware configurations from FlexBench/MLPerf benchmark results. Given a workload spec (model architecture, size, precision) and hardware constraints, it filters real benchmark data, optionally augments it with ML-predicted hypothetical configurations, ranks results by performance or cost, and visualizes model-prediction quality. It is deployed as a Hugging Face Space (see the YAML front matter in `README.md`). The whole app is a single `gr.Blocks` interface built in `app.py`; there is no separate frontend or server layer. ## Running The `.bat` files reflect the actual local dev workflow (uv + Python 3.12 venv): ```bash # 1. create venv uv venv --python 3.12 .venv # 2. install deps (note: also force-reinstalls a LOCAL cMeta checkout — see below) uv pip install -r requirements.txt # 3. run uv run python app.py # launches Gradio on the default local port ``` `_2_install_deps.bat` force-reinstalls cMeta from a hardcoded local path (`D:\!FGG_Repos\fgg\fgg.project\cMeta\cmeta[all]`). This path is machine-specific to the maintainer and will not exist elsewhere; on other machines install `cmind`/cMeta from its normal source or skip it if the data pipeline isn't being touched. The README's simpler `pip install -r requirements.txt` + `python -m app` also works if you don't need cMeta. There are **no tests, linter config, or build step** in this repo. ## Data `data.json` (~847 records, git-LFS-tracked large file) is the sole input — pre-processed FlexBench/MLPerf results with **dot-namespaced flat keys** (e.g. `metrics.result`, `system.accelerator.name`, `model.number_of_parameters`). This flat dotted-key convention is load-bearing: feature definitions, filtering, and prediction all key off these exact strings. `utils.load_data()` reads the JSON, coerces numeric-looking strings to int/float, and returns a **Polars** DataFrame. Note the two-DataFrame split throughout the app: - **Polars** `df` — used only for `extract_metadata()` (building UI dropdown/slider choices). - **pandas** `pd_df` — used by everything else (`predictor.py`, `recommender.py`, `cost_calculator.py`, all filtering in `app.py`). When adding logic, match the library already in use in that module. ## Architecture / data flow `utils.py` is the schema authority. `FEATURES` maps every column to a group and a type (`continuous` / `categorical` / `boolean` / `text`); `FEATURE_TYPES` and `UI_FEATURE_GROUPS` are derived from it. `get_feature_type()` drives whether a filter does exact-match (categorical) or ±tolerance range-match (continuous). Adding/renaming a data column means updating `FEATURES` here first. At startup `app.py` loads data once into module globals (`df`, `pd_df`, `metadata`, `predictor`, `config_finder`) — these are shared, not per-session. It then defines all Gradio components and wires callbacks inside one `gr.Blocks` context. Request flow when the user clicks **Search Configurations**: 1. `process_framework_inputs(*args)` unpacks the flat positional args (order defined by `all_inputs` + framework dropdowns — **keep these lists in sync with the callback's index-based unpacking**, e.g. `base_args[16]` etc.) into `workload_specs` and `constraints` dicts. 2. `find_best_configs()` filters `pd_df`: exact-match for categoricals, ±10% tolerance for continuous features (`apply_continuous_feature_tolerance`), plus explicit min/max range filters for memory and accelerator count. 3. If predictions are enabled and architecture+model_size are set, `predictor.generate_predictions()` synthesizes hypothetical configs; these are cost-scored and concatenated with real results, tagged via a `predicted` boolean column. 4. Results are ranked by `metrics.result_per_accelerator` (performance) or `cost_per_million_tokens` (cost), then formatted for the three output tabs and the bar chart. The `predicted` column and `system.name = "Hypothetical system - ongoing work"` are how generated rows are distinguished from real benchmark rows downstream. ### The predictor (`predictor.py`) `PerformancePredictor` trains an **XGBoost regressor** (with `enable_categorical=True`, so object columns are cast to pandas `category` dtype rather than one-hot encoded) on `data.json` at construction, targeting `metrics.result_per_accelerator`. It excludes leakage-prone columns (`submission.*`, all `metrics.*`, `model.name`, `system.name`, etc.). Beyond prediction it does **statistical data synthesis**: `_analyze_data_distributions()` and the `_analyze_*_relations()` methods build conditional distributions (vendor→accelerator, accelerator→memory, vendor→software stack, node→device-count, …). `_generate_configs()` samples from these to produce *plausible* hardware configs respecting user constraints, which the model then scores. This is why predictions look realistic rather than random. Evaluation metrics (RMSE/MAE/R²/MAPE), a held-out test set, and feature importances are computed in `_evaluate_model()` and surfaced in the "ML Model Performance" tab. ### Cost model (`cost_calculator.py`) Uses **module-global mutable state** `device_costs`, seeded from `DEFAULT_DEVICE_COSTS` by `initialize_device_costs()`. `normalize_gpu_name()` collapses raw accelerator names into device families (e.g. any "H100" → "NVIDIA H100"). The "Device Cost Settings" tab lets users edit hourly costs live, mutating this global. `cost_per_million_tokens` is derived as `hourly_cost / (result_per_accelerator * 3600) * 1e6`. ### Recommender (`recommender.py`) `ConfigurationFinder` is a **separate, simpler filtering/ranking path** than `find_best_configs()` in `app.py`. It's instantiated as `config_finder` but the main search UI currently routes through `app.py`'s own logic; keep this in mind before assuming `recommender.py` is on the hot path. ## Conventions & gotchas - **Column names are string literals everywhere.** There is no central enum beyond `FEATURES` in `utils.py`; renaming a column requires a repo-wide search for the dotted string. - Framework columns are dynamic: any `software.framework.` column becomes a UI dropdown automatically via `extract_metadata()`. Adding a framework to the data adds a filter with no code change. - Gradio callbacks pass inputs **positionally**. `process_framework_inputs` and `get_constraints_from_args` index into `*args` by hardcoded position — changing the `all_inputs` list order will silently break constraint mapping. - `±10% tolerance` on continuous features is intentional app behavior (stated in the UI), not a bug — see `apply_continuous_feature_tolerance` and `ConfigurationFinder.is_within_tolerance`.