indicators-env / openenv.yaml
bawsi99's picture
feat: IndicatorsEnv v4.0 — multi-stock relative alpha MDP
eb89118
Raw
History Blame Contribute Delete
5.3 kB
name: IndicatorsEnv
version: 4.0.0
description: >
Multi-stock Relative Alpha MDP for NSE (Indian) equity analysis.
At each step the agent observes 3 stocks from the same NSE sector and must
identify which stock has the strongest relative momentum signal.
It picks one stock (or passes with NONE) and declares a direction.
Reward = (chosen stock period return - sector average) x direction x conviction x 50.
Market-neutral design: random policy earns ~0 expected reward; a skilled policy
earns consistently positive alpha by correctly identifying within-sector outperformers.
Step spacing = GT window so there is zero ground-truth overlap between steps.
author: bawsi99
license: MIT
tags:
- finance
- stock-market
- technical-analysis
- reinforcement-learning
- openenv
- multi-stock
- relative-alpha
- market-neutral
tasks:
- id: short_term_direction
name: "Short-term Relative Alpha (Easy)"
description: >
5 steps, 1 trading day apart (= 1 week of daily observations).
Observe 3 same-sector NSE stocks per step. Pick one or pass with NONE.
Reward = (chosen 1-day return - sector 1-day avg) x direction x conviction x 50.
GT threshold: +/-0.3% (1-day return window).
difficulty: easy
term: short
episode_steps: 5
step_spacing_days: 1
- id: medium_term_direction
name: "Medium-term Relative Alpha (Medium)"
description: >
10 steps, 5 trading days apart (= 10 weekly observations over 2.5 months).
Signal history (RSI trends, price momentum) accumulates across steps
to enable sequential multi-step inference.
Reward = (chosen 5-day return - sector 5-day avg) x direction x conviction x 50.
GT threshold: +/-1.5% (5-day return window).
difficulty: medium
term: medium
episode_steps: 10
step_spacing_days: 5
- id: long_term_conviction
name: "Long-term Risk-Constrained Alpha (Hard)"
description: >
15 steps, 20 trading days apart (= 15 monthly observations over 15 months).
Spans multiple market regimes. Macro context (NIFTY50 trend, market regime)
added to each observation. Episode terminates early if virtual capital
drawdown exceeds 10%. Conviction calibration scored.
GT threshold: +/-2.5% (20-day return window).
difficulty: hard
term: long
episode_steps: 15
step_spacing_days: 20
endpoints:
reset: "POST /reset"
step: "POST /step"
state: "GET /state"
tasks: "GET /tasks"
grader: "POST /grader"
baseline: "GET /baseline"
health: "GET /health"
websocket: "WS /ws"
action_space:
type: object
required:
- stock
- direction
- conviction
properties:
stock:
type: string
description: >
NSE symbol to trade (must be one of the 3 available_stocks in the observation),
or "NONE" to skip this step.
example: "HDFCBANK"
direction:
type: string
enum: [Bullish, Bearish, NONE]
description: >
Bullish = bet chosen stock outperforms sector average (long alpha).
Bearish = bet chosen stock underperforms sector average (short alpha).
NONE = pass this step (reward = 0, excluded from grader scoring).
conviction:
type: number
minimum: 0.0
maximum: 1.0
description: >
Kelly fraction: proportion of virtual wealth wagered on this prediction.
Higher conviction amplifies both gains and losses proportionally.
Set to 0.0 when passing (direction = NONE).
observation_space:
type: object
properties:
step:
type: integer
description: Current step number (1-indexed)
max_steps:
type: integer
description: Total steps in this episode (5 / 10 / 15 by task)
term:
type: string
enum: [SHORT, MEDIUM, LONG]
sector:
type: string
description: "NSE sector shared by all 3 stocks (banking / it / pharma / fmcg / auto)"
available_stocks:
type: array
items:
type: string
description: Exactly 3 NSE symbols to choose from this step
stocks:
type: object
description: >
Dict keyed by symbol. Each entry: current_price, rsi_14, rsi_trend,
price_momentum_pct (cumulative return vs episode start), and a full
indicators dict (moving_averages, macd, bollinger_bands, adx,
stochastic, volatility, enhanced_volume, pivot_points).
signal_history:
type: array
description: >
Chronological log of past picks. Each entry: step, picked_stock,
direction, ground_truth, correct, alpha_pct, reward, rsi_snapshot.
Empty on step 1; grows by one entry each step.
macro:
type: object
nullable: true
description: >
Task 3 only. Fields: nifty_trend (Bullish/Bearish/Neutral),
nifty_return_20d (float %), market_regime (trending/ranging).
reward:
formula: "(chosen_return - sector_avg_return) x direction_sign x conviction x 50"
min: -1.5
max: 1.5
notes: >
Market-neutral: sector average cancels broad market beta so a random agent
earns ~0 expected reward. NONE action yields reward = 0 and is excluded
from grader scoring. direction_sign is +1 for Bullish, -1 for Bearish.
Conviction scales reward linearly (Kelly criterion semantics).