# GPU-share — zero-conf coordination between opencoti-llamafile peers on one GPU > **Status: G1+G2 IMPLEMENTED (2026-07-23, c5, patch `0141`)** — shm peer > registry + `--gpu-share-weight` + `GET /gpu/peers` + `/props` > `opencoti.gpu_share` + duty-cycle pacing, all in > `tools/server/server-context.cpp` (`oc_gpu_share` namespace). G3 > (Windows/macOS live validation) + CUDA-UUID DSO keying remain (c6). > Original design brief: Multiple opencoti-llamafile > processes share one GPU (the multi-server pattern opencoti-server manages, > plus standalone uses as an agentic-framework engine). They need direct, > zero-conf, port-less discovery + a **weight** that governs each process's > **compute share**. Decisions (user): weight = compute share (decode > pacing; full speed when alone); transport = named shared-memory segment; > Linux/macOS/Windows from the ONE APE binary. ## 1. Semantics - New flag `--gpu-share-weight W` (float > 0, default `1.0`). - Share of process *i* = `w_i / Σ w_j` over **alive, active** peers on the same physical GPU — a plain ratio. Example (the spec case): weights `2, 1, 1` → `50% / 25% / 25%`. The endpoint always shows each peer's resolved percentage so there is no guessing. - **Idle peers cost nothing**: a peer that hasn't decoded/prefilled within the activity window (default 2 s) is excluded from the denominator — a lone active process always runs at 100% regardless of registered peers. - Enforcement is **cooperative duty-cycle pacing** (see §4) — approximate by design; the goal is proportional sharing, not hard isolation (that would need MPS/MIG, which is Linux-only / non-portable and breaks zero-conf). ## 2. Why a named shared-memory registry (research result) Requirements: zero-conf (no broker, no config, no ports), sub-ms reads from the decode loop, crash-safe (a SIGKILLed peer must not wedge the others), one APE binary for Linux/macOS/Windows. - **Chosen: POSIX `shm_open` + `mmap(MAP_SHARED)` via cosmopolitan.** cosmo libc (vendored `.cosmocc/4.0.2`) declares `shm_open`/`shm_unlink` (`libc/calls/calls.h:161`) and its `mmap` implements shared file mappings on Windows over CreateFileMapping/MapViewOfFile — the same C source compiles and runs on all three OSes inside the APE. Linux backs it with `/dev/shm`, macOS with POSIX shm, Windows with the cosmo shim. - **Rejected: unix sockets / named pipes** — message passing + broker election we don't need for a shared counter table; 3 platform code paths. - **Rejected: filesystem lock/registry files** — stale-file cleanup, dir permission variance, and mtime-based liveness are all worse than pid+ heartbeat words in shm. - **Rejected: CUDA MPS `CUDA_MPS_ACTIVE_THREAD_PERCENTAGE`** — Linux-only, requires an MPS control daemon (not zero-conf), static percentages (no idle-peer reflow), nothing on Windows/macOS. - **No locks in the segment**: fixed-size slots, single-writer-per-slot (each process writes only its own slot), seq/heartbeat words with atomic stores; readers tolerate torn strings (name is display-only). A crashed peer simply stops heartbeating and ages out — nothing to clean. ## 3. Registry layout Segment name: `/opencoti-gpu-` (one per physical GPU), 4 KiB. ``` header: { magic "OCGPU", version u32, slot_count u32 = 16 } slot: { pid u32 (0 = free, CAS-claimed), boot_us i64, weight f32, active u32 (decoded within window), heartbeat_us i64 (atomic store, ~500 ms cadence), busy_ppm u32 (measured busy fraction, parts-per-million), name char[48] (alias/model, display only) } ``` - **gpu_key**: CUDA device UUID via a new DSO export `ggml_backend_cuda_get_device_uuid` (we own the published `.so`/`.dll` builds — additive symbol; when absent in an older side-loaded DSO, fall back to a hash of device description + PCI ordinal). Metal → `metal-0` (Apple GPUs are single-device). CPU-only runs register nothing. - **Liveness**: heartbeat age > 3 s ⇒ dead; its slot is reclaimable by CAS. Last liver does NOT `shm_unlink` (races the next joiner); a 4 KiB segment persisting until reboot is harmless and self-heals via the heartbeat rule. ## 4. Pacing (compute-share enforcement) Since c6 (patch `0143`) pacing is **deterministic weighted time-division (TDM)** over a host-shared wall clock — not feedback control: 1. Each peer derives the SAME schedule from the shm registry + the host clock: a repeating period carved into one **contiguous window per alive+active peer**, proportional to its (10%-floored, renormalized) share, ordered by slot index. The period is **adaptive for latency**: a peer's worst token stall is `(1−share)·period`, so the period is the smallest that still gives the smallest share a contiguous-effective window (40 ms — bug-2243's floor): `period = clamp(40ms / min_share, 80ms, 400ms)`. 2:1 weights → a 120 ms period (worst stalls 80/40 ms); 1:1 → 80 ms (40 ms); only extreme splits reach 400 ms. 2. In the decode loop, after each step: if the current wall-clock position is inside my window → keep decoding; otherwise sleep exactly until my next window starts. Overshoot into a peer's window is ≤1 decode step. 3. Alone (or all peers idle) → no schedule, never sleep. Idle peers leave the schedule via the 2 s activity window; kill -9 peers via the 3 s heartbeat age. Exactly one peer runs at any instant, so the pacer is **work-conserving** (aggregate ≈ solo throughput) and the split equals the weight ratio **by construction** — there is no control loop to converge or oscillate. Why not feedback? Three control laws failed in sequence: - Proportional sleep `dt·(busy/share−1)` has no fixed point (bug-2243). - Integral debt on **absolute** busy (`debt += dt − share·wall`, the c5 pacer) over-throttles the minority peer once busy is measured honestly (2.45:1 on the 3090 / 3:1 on the 5080 at 2:1 weights): contended, every peer's `llama_decode` wall includes waiting on the OTHER peer's kernels, so busy reads ≈ wall for everyone (bug-2244 measurement is correct — the signal itself is contention-inflated). - Integral debt on **relative** busy (my busy / Σ peers' busy) fixes the bias but lets both peers sleep in overlapping windows (500 ms busy-publish lag → oscillation): GPU idled ~half the time, aggregate 73 vs 116 tok/s. Two clock/measurement invariants remain load-bearing: - **Shared epoch (bug-2245):** `ggml_time_us()` is per-process-epoch on Windows (native `timer_start`, and cosmo's `CLOCK_MONOTONIC` alike), so shm heartbeat ages misread by the peers' launch delta — a pair started >3 s apart never saw each other alive (latent in c5), and TDM windows landed at random offsets (5080 aggregate 53 vs 146). All shm timestamps + the schedule use `wall_us()` = `CLOCK_REALTIME` (Unix epoch, shared by every process on the host, on every platform). - **Busy through `llama_synchronize` (bug-2244):** `llama_decode` is async-submit; timing the call alone reported 4.7% busy at 93% real utilization. Busy (published for `/gpu/peers` and the activity flag) spans decode through a sync, which is imminent at sampling anyway. **Measured (TDM, c6, adaptive period):** - solidPC RTX 3090, Qwen3-4B Q4_K_M batch-1: solo 121 tok/s; 2:1 → 75.6/38.4 (**1.97:1, aggregate 114.0** vs the c5 feedback pacer's 107.9); with a deliberate 20 s start skew → 1.99:1; 1:1 → 58.9/59.0 (**aggregate 117.9** vs 100.8, +17%); 2/1/1 → 56.5/27.3/29.5 (1.92:0.92:1, agg 113.3); solo-after-release 120.9. - **Inter-token latency (streaming, 2:1):** w2 p50 9 ms / p95 45 / max 49 ms; w1 p50 9 ms / p95 86 / max 89 ms — exactly the `(1−share)·period` bound. (The initial fixed 400 ms period stalled the w1 stream up to 267 ms; the c5 pacer's blocks were 100–150 ms.) - **`--parallel` multi-slot (auto n_parallel, 4 concurrent gens each), 2:1:** w2 agg 217.1 / w1 agg 125.6 (**1.73:1, total 342.7** vs the c5 pacer's ~211–246 total; solo 4-way is 327.5 — TDM total is at solo level, i.e. fully work-conserving). Ratio softens at multi-slot because a batched iteration overshoots the window boundary by more than one token — accepted trade-off. - **MTP (Qwopus3.5-9B NextN, `--spec-type draft-mtp`, `-c 8192`), 2:1:** 74.8/35.2 (**2.13:1, aggregate 110.0 ≈ 96% of the 114.7 solo**); draft acceptance unharmed (0.91/0.82 vs 0.87 solo). - pandorum RTX 5080 (Windows, side-load DLL), 2:1 → 97.7/48.9 (**2.00:1, aggregate 146.6 ≈ 97% of the 151 tok/s solo**; the c5 fixed-block pacer gave 3:1 at aggregate ~99). Windows' 15.6 ms timer granularity is immaterial at 40–80 ms sleeps. - Gate discipline (bug-2246): on a 16 GB card set an explicit `-c` — VRAM auto-fit filled 15.8/16 GB and WDDM paging collapsed the aggregate to 53 tok/s (an artifact of the gate setup, not the pacer). Also gate with AUTO `n_parallel`: an explicit `--parallel 4` takes a non-unified-KV path that decodes at ~28 tok/s aggregate regardless of pacing (bug-2247, pre-existing, out of gpu-share scope). Prefill batches count as busy time; the 10% share floor guarantees no peer can be starved by misconfigured weights. ### Driver-level alternatives evaluated (2026-07-23) Could the sleeps be replaced by CUDA-native compute partitioning? Findings (CUDA 13.x): - **MPS static SM partitioning (CUDA 13.1+, Linux, Ampere+):** the control daemon (`nvidia-cuda-mps-control -d -S`) can carve **exclusive SM partitions** per MPS client (`sm_partition add`), deterministic and concurrent — no sleeps, no stall windows. This is the real long-term candidate for Linux fleets (bs2-class); requires daemon lifecycle management (a natural opencoti-server F3 feature) and an r590+ driver (solidPC's 580 predates it; bs2's 610 has it). Blackwell adds MLOPart (memory-locality-optimized partition devices). - **Classic MPS + `CUDA_MPS_ACTIVE_THREAD_PERCENTAGE` (Volta+, Linux), measured on the 3090 (driver 580):** 67/33 caps → 64.4/37.5 (**1.72:1, aggregate 101.9** — ~11% BELOW TDM's 114.0, because batch-1 decode is bandwidth-bound, and concurrent kernels contend on memory bandwidth while SM caps only bound compute). Latency is smoother (max gap 50/53 ms vs TDM's 49/89) but p50 worsens for the small partition (16 ms vs 9). Critically **not work-conserving**: the caps are static — with the peer idle the 67% client does 110.9 tok/s vs TDM's full 120.9, and ratios track SM fraction only loosely. - **Green contexts (CUDA 13.1 runtime "execution context" API):** SM partitioning **within a single process only** — inapplicable to cross-process gpu-share (relevant someday for multi-model-in-one- process serving). Isolation is best-effort, not guaranteed. - **Windows:** no MPS, no cross-process partitioning at all — WDDM time-slicing is the only primitive, so wall-clock TDM remains the only portable mechanism there. Conclusion: TDM stays the default (portable, zero-setup, work-conserving, exact ratios, instant idle-release). An opt-in Linux MPS mode (static SM partitioning managed by opencoti-server) is a future milestone for latency-critical fleets on r590+ drivers. ## 5. Visibility - `GET /gpu/peers` → `{ gpu_key, self: {pid, name, weight, share_pct, busy_pct, active}, peers: [{pid, name, weight, share_pct, busy_pct, active, heartbeat_age_ms}] }` - `/props` `opencoti` block gains `gpu_share` carrying the same `{enabled, gpu_key, self, peers}` snapshot as `/gpu/peers`, so existing dashboards see it without a new call. - Boot log: `gpu-share: joined /opencoti-gpu- slot 2 (weight 1.5, 3 peers alive)`. ## 6. Phases - **G1 — registry + flag + visibility (no throttle).** shm module (`llamafile/gpu_share.c`, cosmo-portable), `--gpu-share-weight`, join/ heartbeat/reap, `/gpu/peers` + `/props` block. Gate: 3 instances on one GPU (solidPC 3090), each sees the other two with correct resolved percentages; kill -9 one → peers age it out ≤ 3 s. - **G2 — duty-cycle pacing.** Busy-window measurement + inter-batch sleep. Gate: 2 instances, weights 2:1, concurrent decode workloads → measured tps ratio ≈ 2:1 (±20%); solo instance regains ~full solo tps ≤ 2 s after peer goes idle. - **G3 — Windows/macOS validation** (pandorum APE run for the shm path on Windows; macOS opportunistic) + docs + opencoti-server fleet surfacing (daemon shows per-engine share/busy from `/gpu/peers`). ## 7. Non-goals - Hard isolation (MPS/MIG) — not zero-conf, not portable. - VRAM partitioning by weight — allocation is boot-time (`-c`, `-ngl`); the endpoint may later *report* VRAM per peer but never enforces it. - Cross-GPU balancing — the registry is strictly per-GPU; multi-GPU processes join one segment per device they occupy (v1: primary device only).