--- title: PipelineEnv emoji: πŸ”§ colorFrom: purple colorTo: blue sdk: docker app_port: 7860 tags: - openenv - ci-cd - devops - rl-environment --- # PipelineEnv πŸ”§ > An OpenEnv-compliant Reinforcement Learning environment where an AI agent diagnoses and repairs broken CI/CD pipelines β€” a real-world DevOps self-healing scenario. **Live Demo:** [https://huggingface.co/spaces/Endraode/pipeline-env](https://huggingface.co/spaces/Endraode/pipeline-env) --- ## Table of Contents - [Overview](#overview) - [Quick Start](#quick-start) - [Architecture](#architecture) - [Environment Specification](#environment-specification) - [Tasks](#tasks) - [Reward Function](#reward-function) - [Grading System](#grading-system) - [API Endpoints](#api-endpoints) - [Local Setup](#local-setup) - [Docker Build & Deploy](#docker-build--deploy) - [Baseline Inference](#baseline-inference) - [Validation](#validation) - [Project Structure](#project-structure) - [License](#license) --- ## Overview Every engineering team faces broken CI/CD pipelines β€” a bad merge breaks tests, an invalid Dockerfile kills the build, a missing environment variable crashes deployment. **PipelineEnv** simulates these exact scenarios in a structured RL environment where an agentic system must diagnose failures and apply the correct repair actions in the correct order. ### Key Features - **Real-world domain** β€” models actual DevOps failure modes engineers encounter daily - **3 difficulty tiers** β€” easy (single fix), medium (multi-component), hard (ordered sequence) - **Deterministic grading** β€” stage-weighted health scores with action-order enforcement - **Interactive dashboard** β€” Gradio UI with live terminal, health bar, and stage visualization - **REST API** β€” fully OpenEnv-compliant `step() / reset() / state()` endpoints - **Docker-native** β€” containerized deployment tested with `docker build && docker run` --- ## Quick Start ```bash # Local development pip install -r requirements.txt uvicorn server.app:app --host 0.0.0.0 --port 7860 # Open dashboard open http://localhost:7860 ``` The environment starts immediately. No dataset downloads, no database setup. --- ## Architecture ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ FastAPI Server (server/app.py) β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ β”‚ β”‚ β”‚ /reset β”‚ β”‚ /step β”‚ β”‚ /state β”‚ β”‚ β”‚ β”‚ POST β”‚ β”‚ POST β”‚ β”‚ GET β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ PipelineEnvironment (RL loop) β”‚ β”‚ β”‚ β”‚ - scenario selection β”‚ β”‚ β”‚ β”‚ - action execution β”‚ β”‚ β”‚ β”‚ - health computation β”‚ β”‚ β”‚ β”‚ - reward shaping β”‚ β”‚ β”‚ β”‚ - action history tracking β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ Graders (server/graders.py) β”‚ β”‚ β”‚ β”‚ - compute_health_score() (weighted stages) β”‚ β”‚ β”‚ β”‚ - grade_task() (deterministic 0.0-1.0) β”‚ β”‚ β”‚ β”‚ - ACTION_ORDER enforcement (hard task) β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ Gradio UI (root) β€” deterministic agent demo β”‚ β”‚ β”‚ β”‚ Pipeline stages | Health bar | Terminal β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` --- ## Environment Specification ### Observation Space The agent observes the full pipeline state after each action: | Field | Type | Description | |-------|------|-------------| | `pipeline_name` | `str` | Name of the pipeline scenario | | `stages` | `List[dict]` | All stages with `name`, `status`, `error`, `runtime` | | `failing_count` | `int` | Number of failing stages | | `health_score` | `float` | Overall health in `[0.0, 1.0]` | | `error_messages` | `List[str]` | Human-readable error strings from failing stages | | `available_actions` | `List[str]` | All repair actions the agent can take | | `task_description` | `str` | Natural-language description of the failure | | `step_number` | `int` | Current step counter | | `max_steps` | `int` | Maximum allowed steps before forced `done=True` | ### Action Space | Action | Description | Affected Stages | |--------|-------------|-----------------| | `fix_test` | Fix a failing unit/integration test | `test` (builds deploy) | | `set_env_var` | Set a missing environment variable | `deploy` | | `fix_docker_config` | Fix Dockerfile misconfiguration | `build` (unskips test) | | `fix_yaml_config` | Fix disabled pipeline YAML config | `deploy` | | `retry_stage` | Retry a flaky stage (partial recovery) | specified stage only | | `rollback_commit` | Rollback a breaking commit (partial) | `build` (exposes dependency) | | `add_dependency` | Install a missing package/dependency | `build`, `test` | | `no_op` | Pass β€” penalized -0.1 per step | none | --- ## Tasks | Task | Pipeline | Scenario | Max Steps | Start Health | Required Actions | |------|----------|----------|-----------|-------------|------------------| | `easy` | simple-app-pipeline | A unit test is failing; fix the test | 5 | 0.20 `fix_test` | | `medium` | dockerized-api-pipeline | Docker build config broken + missing env var | 8 | 0.00 | `fix_docker_config` β†’ `set_env_var` | | `hard` | multi-service-pipeline | Cascading 3-stage failure: bad commit, missing dependency, disabled YAML, disabled in config | 12 | 0.00 | `rollback_commit` β†’ `add_dependency` β†’ `fix_yaml_config` | ### Task Breakdown #### Easy β€” `simple-app-pipeline` - **Build:** passing (green) - **Test:** failing β€” `AssertionError: test_add failed β€” expected 4 got 5` - **Deploy:** skipped (blocked by failing test) - **Fix:** Apply `fix_test` β†’ all stages transition to passing #### Medium β€” `dockerized-api-pipeline` - **Build:** failing β€” `Docker build failed: invalid FROM instruction` - **Test:** skipped (blocked by build failure) - **Deploy:** failing β€” `Missing env var: DATABASE_URL` - **Fix:** `fix_docker_config` fixes build and unskips test, `set_env_var` fixes deploy #### Hard β€” `multi-service-pipeline` - **Build:** failing β€” `ModuleNotFoundError: No module named 'requests'` - **Test:** failing β€” `ImportError: cannot import requests` - **Deploy:** failing β€” `Deploy stage disabled in pipeline YAML` - **Fix:** Must be done **in order** β€” rollback exposes the missing dependency, add_dependency resolves imports, fix_yaml re-enables deploy - **Wrong order penalized** β€” grader enforces correct action sequence --- ## Reward Function The reward provides **dense, varying signals** throughout the episode β€” never a sparse binary signal: | Signal | Reward | |--------|--------| | Health improvement | `+delta + 0.05` bonus | | Health regression | `+delta - 0.05` penalty | | No change in health | `-0.05` | | `no_op` action | `-0.1` | | Episode done (`health >= 0.99`) | End of episode | This means the agent receives **immediate feedback** after every action, allowing it to learn from partial progress and course-correct on wrong decisions. --- ## Grading System ### Health Score (`compute_health_score`) Deterministic weighted sum over stage statuses: | Stage | Weight | |-------|--------| | `build` | 0.2 | | `test` | 0.3 | | `deploy` | 0.5 | Same pipeline state always produces the same score. Scores are in `[0.0, 1.0]`. ### Task Grader (`grade_task`) - Returns `1.0` if `health >= 0.99` AND (for hard task) actions are in correct order - Returns `0.7` for hard task if actions are out of order (even with full health) - Returns `health_score` for partial progress on easy/medium tasks - **100% deterministic** β€” same action sequence always produces same score --- ## API Endpoints All endpoints are OpenEnv-compliant and tested via `openenv validate`, `docker build`, and `HF Space` deployment. | Method | Endpoint | Description | |--------|----------|-------------| | `GET` | `/` | Gradio UI dashboard (interactive demo) | | `GET` | `/health` | Server health status | | `POST` | `/reset` | Start new episode `{"task_id": "easy"}` | | `POST` | `/step` | Take a repair action `{"action": "fix_test"}` | | `GET` | `/state` | Current episode metadata | ### Response Format ```json POST /reset β†’ {"task_id": "hard"} { "pipeline_name": "multi-service-pipeline", "stages": [ {"name": "build", "status": "failing", "error": "ModuleNotFoundError: No module named 'requests'", "runtime": 1.5}, {"name": "test", "status": "failing", "error": "ImportError: cannot import requests", "runtime": 1.0}, {"name": "deploy", "status": "failing", "error": "Deploy stage disabled in pipeline YAML", "runtime": 0.5} ], "failing_count": 3, "health_score": 0.0, "error_messages": ["ModuleNotFoundError…", "ImportError…", "Deploy stage disabled…"], "available_actions": ["fix_test", "set_env_var", "fix_docker_config", "fix_yaml_config", "retry_stage", "rollback_commit", "add_dependency", "no_op"], "task_description": "A bad commit removed a critical dependency…", "step_number": 0, "max_steps": 12 } ``` --- ## Docker Build & Deploy ### Build ```bash docker build -t pipeline-env . ``` ### Run ```bash docker run -p 7860:7860 pipeline-env ``` ### Environment Variables (optional) | Variable | Default | Purpose | |----------|---------|---------| | `API_BASE_URL` | `https://router.huggingface.co/v1` | LLM API endpoint | | `MODEL_NAME` | `meta-llama/Llama-3.1-8B-Instruct` | Model for inference | | `HF_TOKEN` | `none` | HuggingFace API key (for LLM calls) | --- ## Baseline Inference The `inference.py` script runs a headless benchmark over all 3 tasks: ```bash export HF_TOKEN=hf_xxx # Your HuggingFace token export API_BASE_URL=https://router.huggingface.co/v1 export MODEL_NAME=meta-llama/Llama-3.1-8B-Instruct export BASE_URL=http://localhost:7860 python inference.py ``` ### Output Format (strict std format) ``` [START] task=easy env=pipeline-env model=meta-llama/Llama-3.1-8B-Instruct [STEP] step=1 action=fix_test reward=0.85 done=true error=null [END] success=true steps=1 score=1.00 rewards=0.85 ``` --- ## Validation Run the full test suite (163 assertions): ```bash python test_suite.py ``` Run the OpenEnv validator: ```bash openenv validate # [OK] pipeline: Ready for multi-mode deployment ``` Run the pre-submission checker: ```bash ./validate-submission.sh https://endraode-pipeline-env.hf.space . ``` --- ## Project Structure ``` pipeline-env/ β”œβ”€β”€ server/ β”‚ β”œβ”€β”€ __init__.py β”‚ β”œβ”€β”€ app.py # FastAPI REST server + Gradio mount β”‚ β”œβ”€β”€ pipeline_environment.py # Core RL environment (reset/step/state) β”‚ β”œβ”€β”€ pipeline_scenarios.py # Pre-broken pipeline definitions β”‚ β”œβ”€β”€ graders.py # Deterministic health & task graders β”‚ └── requirements.txt # Server dependencies β”œβ”€β”€ models.py # Pydantic models (Action, Observation, State) β”œβ”€β”€ inference.py # Baseline headless benchmark script β”œβ”€β”€ ui.py # Gradio dashboard (interactive demo) β”œβ”€β”€ test_suite.py # Comprehensive test suite (163 tests) β”œβ”€β”€ openenv.yaml # OpenEnv metadata & task definitions β”œβ”€β”€ pyproject.toml # Project config + setuptools scripts entry β”œβ”€β”€ Dockerfile # Containerized build β”œβ”€β”€ README.md # This file └── uv.lock # Deterministic dependency lock file ``` --- ## License MIT