MediaRouter / docs /ai-copilot.md
basyx's picture
Upload 376 files
8bf0d31 verified
|
Raw
History Blame Contribute Delete
5.46 kB
# AI Copilot
## Status
AI Copilot is implemented in the current working tree as a bounded,
server-authoritative orchestration layer with durable runs, typed actions,
scope enforcement, and MCP transport. Provider credential certification,
PostgreSQL/RLS runtime verification, Docker, and Hugging Face startup remain
deferred until all planned product phases are complete.
AI Copilot is a validated orchestration layer over existing MediaRouter
services. It does not execute arbitrary model output and does not duplicate
project, asset, editor, AI Studio, generation, or render business logic.
## Architecture and planner
`POST /v1/copilot/runs` builds a bounded server-authoritative context, creates a
strict action plan, and stores a durable `copilot_runs` record. The current
provider registry has no text/tool-calling model, so the planner is deliberately
deterministic. It recognizes only registered intents and reports unsupported
capabilities rather than fabricating a model response.
The context contains workspace/project IDs, selected asset and clip IDs, a
small editor summary, active tool, and capability names. Workspace identity is
derived from authentication. Assets and clips are revalidated against canonical
workspace/project state. Binary media, signed URLs, credentials, complete
editor documents, and provider payloads are excluded.
## Action registry
Registered actions are:
- `project.open`
- `asset.select`
- `ai.generate_image`
- `ai.generate_video`
- `editor.split_clip`
- `editor.delete_clip`
- `editor.set_duration`
- `editor.add_clip`
- `editor.render`
Each definition declares permission, capability, destructive/external flags,
confirmation policy, handler, and audit classification. Pydantic discriminated
unions reject unknown action types and fields. There is no eval, dynamic code,
shell, SQL, arbitrary URL, filesystem-path, or provider-credential action.
## Execution and confirmation
Plans are created before execution. Generation, rendering, and clip deletion
require explicit `confirmed=true`; the server rejects attempts to bypass this.
Execution uses the authenticated user's scopes and checks project, asset, AI,
editor, render, and job permissions again.
Execution claims and pre-execution cancellation use row-locked state
transitions. Concurrent execute/cancel requests therefore cannot run the same
plan twice or move an already executing run back to a pre-execution state.
Editor actions load the authoritative revision, mutate a validated
`EditorDocument`, and save through `ProjectEditorService` optimistic
concurrency. AI actions call `AiStudioService`. Rendering calls
`ProjectRenderService`. Child AI/render jobs keep their existing durable
idempotency and lifecycle.
Copilot run states are `plan_ready`, `blocked`, `executing`, `completed`,
`partial`, `failed`, and `cancelled`. Successful action results are retained so
a later failure reports partial completion without rerunning completed
non-idempotent work. Progress is action state only; no synthetic percentages
are generated.
## API and persistence
- `POST /v1/copilot/capabilities`
- `POST /v1/copilot/runs`
- `GET /v1/copilot/runs`
- `GET /v1/copilot/runs/{id}`
- `POST /v1/copilot/runs/{id}/execute`
- `POST /v1/copilot/runs/{id}/cancel`
- `GET /v1/copilot/history`
Run creation requires `Idempotency-Key`. The fingerprint includes bounded
request and authoritative context. Migration
`app/projects/migrations/0005_ai_copilot.sql` is additive, applies after
`0004_ai_studio.sql`, enables FORCE RLS, enforces workspace/project ownership,
and is never run automatically in production.
## Security, rate limiting, and audit
Scopes are `copilot:read` and `copilot:execute`; action-specific scopes remain
mandatory. Existing request/job rate limiting applies to Copilot planning and
execution, while AI/render child services retain their own enforcement.
Untrusted user and asset text is data, never an instruction source.
Audit events include run start/plan/action start/action completion/action
failure/run completion/failure/cancellation. Metadata is bounded and excludes
the user request, prompts, credentials, signed URLs, tokens, and raw provider
responses. Structured logs contain IDs, action type, status, duration, and safe
error categories only.
## Frontend, SDK, MCP, and n8n
The workspace top bar opens an accessible responsive right-side panel. TanStack
Query owns runs, history, capabilities, and action state. Zustand owns only
panel/draft/selection UI state. Content Studio supplies selected clip IDs and an
editor summary; project routes and project assets supply context. Copilot runs
also appear in global search and the shared activity drawer. Completed action
results link to existing project, asset, AI Studio, job, and Content Studio
destinations when those canonical routes exist.
TypeScript and Python SDKs expose create/get/list/history/execute/cancel run
methods through `client.copilot`.
MCP and n8n are intentionally not expanded in this phase. The current MCP/n8n
surfaces would expose a broad plan-execution operation rather than individual
typed registry actions. Direct deterministic MediaRouter operations remain the
safer automation interface until those transports support the action catalog
without duplicating business logic.
Production Docker, PostgreSQL/RLS, FFmpeg, Python 3.10, and provider credential
certification remain deferred to the final product-wide certification phase.