observable-agent / README.md
sukhrobnurali's picture
Raise per-session run cap to 5
09eaad0 verified
|
Raw
History Blame Contribute Delete
5.7 kB
---
title: Observable Agent Watch It Think
emoji: 🧠
colorFrom: gray
colorTo: green
sdk: gradio
sdk_version: 6.16.0
app_file: app.py
pinned: false
python_version: "3.12"
---
# Observable Agent — Watch It Think
Give a task and watch a ReAct agent (Gemini 2.5 Flash) work through it **live** — every thought,
tool call, observation, and self-correction streamed to the UI as it happens. Most agent demos show
you the answer; this one shows you the *thinking*, including at least one recoverable wrong step.
> _Demo GIF placeholder — add a screen recording of a run (e.g. the FIFA World Cup example) after the first deploy._
## How the ReAct loop works
The agent alternates between three moves until it can answer:
- 🧠 **Thought** — the model's real reasoning, surfaced via Gemini's thinking output (not a narration we wrote).
- 🔧 **Tool call** — a real tool, with its arguments.
- 📄 **Observation** — the tool's actual returned result.
-**Revision** — when an observation is empty, irrelevant, or a disambiguation, the agent reconsiders
and re-queries instead of guessing. This marker appears only when the model *actually* changes course
after a weak result — it is genuine behavior, not a scripted insert.
Steps are streamed as they happen: a LangGraph `StateGraph` emits per-step events through a custom
stream writer, and a Gradio generator rebuilds the timeline after every event. You are watching the
loop build, not a replay. The loop is hard-capped at 8 steps.
## Tools (all real, all executing)
| Tool | What it does |
|------|--------------|
| `calculator` | Sandboxed arithmetic via [`asteval`](https://github.com/lmfit/asteval) — never bare `eval`. |
| `wikipedia_search` | First ~500 chars of the best-matching English Wikipedia page summary. |
| `web_get` | Fetches readable text from an http/https URL (tags stripped, length-capped). |
Three real tools beat ten shallow ones — this is enough surface for genuine multi-step reasoning chains.
## Resilience (the production-readiness signal)
The senior signal here is the observable orchestration and how it behaves under real-world failure, not the model.
- **Transient errors:** exponential backoff (capped at 3 tries) on retryable server/timeout codes
(408/5xx).
- **Rate limit / quota (429):** *not* retried — retrying only burns the scarce free-tier requests and
can't clear a per-minute window. Instead the 429 is classified from the API's quota metadata and
shown cleanly: "wait about a minute" for a per-minute throttle, or "try again tomorrow" for the daily
cap. Never a stack trace.
- **Per-session run cap:** a soft per-session cap (enough to try every example task) so casual
repeat-clicks don't burn the shared daily quota too fast. It only gates *starting* a run — it never
interrupts one in progress; the daily quota is the real limiter.
- **Tool errors are observations:** every tool returns a structured `ERROR:` / `NO_RESULTS:` string the
agent can react to — which also feeds the revision mechanic.
- **`web_get` is SSRF-hardened:** it resolves the target host and refuses private, loopback, link-local,
reserved, or multicast IPs, and re-validates every redirect hop instead of auto-following. (Known
residual: DNS-rebinding TOCTOU, which needs IP-pinned connections to fully close — out of scope here
and noted honestly in the code.)
- **No frozen UI:** transient "model is thinking…" / "running <tool>…" states keep the timeline alive.
## Running it
This Space runs on the owner's free Gemini key, read from the environment:
1. In the Space, go to **Settings → Secrets** and add `GEMINI_API_KEY`.
2. The key is read from `os.environ` and is **never** hardcoded (this source is public).
Locally:
```bash
python -m venv .venv
.venv/Scripts/python.exe -m pip install -r requirements.txt # use .venv/bin/python on macOS/Linux
GEMINI_API_KEY=... python app.py
```
**Swappable backend:** the LLM call is isolated in [`agent/llm.py`](agent/llm.py). Any Gemini model — or
any OpenAI-compatible backend — can replace it without touching the graph, tools, or UI.
## Caveat (stated honestly)
This uses Gemini's **free tier**, which **may use prompts for training** (don't paste anything
sensitive) and — importantly — has a **small daily request quota** (observed ~20 requests/day for
`gemini-2.5-flash`). Every agent run makes several model calls and all visitors share the owner's single
key, so the live demo realistically serves only **a few runs per day** before showing "try again
tomorrow." That's a deliberate free-tier trade-off — the point of this piece is the observable
orchestration, not uptime. The model is `gemini-2.5-flash` (2.5 Pro's free quota is lower still). This is
a portfolio demo, not a production service.
## Tests
```bash
python -m pytest -q
```
The suite runs fully offline (no API key) using a fake LLM client and injected tool stubs: it covers tool
safety (sandboxed eval, SSRF host checks), the retry/backoff/quota logic, the full ReAct loop including the
revision and step-cap paths, and HTML rendering.
## Layout
```
app.py Gradio UI + live-streaming run generator
agent/
config.py all tunables + the system prompt
llm.py Gemini client, retry/backoff, quota mapping
tools.py calculator, wikipedia_search, web_get + declarations
graph.py LangGraph ReAct loop + per-step event streaming
render.py pure event → HTML rendering + theme CSS
tests/ offline unit tests (no API key required)
```