File size: 12,938 Bytes
559d4e6 327bdb9 559d4e6 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 | # GPU-share β zero-conf coordination between opencoti-llamafile peers on one GPU
> **Status: G1+G2 IMPLEMENTED (2026-07-23, c5, patch `0141`)** β shm peer
> registry + `--gpu-share-weight` + `GET /gpu/peers` + `/props`
> `opencoti.gpu_share` + duty-cycle pacing, all in
> `tools/server/server-context.cpp` (`oc_gpu_share` namespace). G3
> (Windows/macOS live validation) + CUDA-UUID DSO keying remain (c6).
> Original design brief: Multiple opencoti-llamafile
> processes share one GPU (the multi-server pattern opencoti-server manages,
> plus standalone uses as an agentic-framework engine). They need direct,
> zero-conf, port-less discovery + a **weight** that governs each process's
> **compute share**. Decisions (user): weight = compute share (decode
> pacing; full speed when alone); transport = named shared-memory segment;
> Linux/macOS/Windows from the ONE APE binary.
## 1. Semantics
- New flag `--gpu-share-weight W` (float > 0, default `1.0`).
- Share of process *i* = `w_i / Ξ£ w_j` over **alive, active** peers on the
same physical GPU β a plain ratio. Example (the spec case): weights
`2, 1, 1` β `50% / 25% / 25%`. The endpoint always shows each peer's
resolved percentage so there is no guessing.
- **Idle peers cost nothing**: a peer that hasn't decoded/prefilled within
the activity window (default 2 s) is excluded from the denominator β a
lone active process always runs at 100% regardless of registered peers.
- Enforcement is **cooperative duty-cycle pacing** (see Β§4) β approximate
by design; the goal is proportional sharing, not hard isolation (that
would need MPS/MIG, which is Linux-only / non-portable and breaks
zero-conf).
## 2. Why a named shared-memory registry (research result)
Requirements: zero-conf (no broker, no config, no ports), sub-ms reads from
the decode loop, crash-safe (a SIGKILLed peer must not wedge the others),
one APE binary for Linux/macOS/Windows.
- **Chosen: POSIX `shm_open` + `mmap(MAP_SHARED)` via cosmopolitan.**
cosmo libc (vendored `.cosmocc/4.0.2`) declares `shm_open`/`shm_unlink`
(`libc/calls/calls.h:161`) and its `mmap` implements shared file mappings
on Windows over CreateFileMapping/MapViewOfFile β the same C source
compiles and runs on all three OSes inside the APE. Linux backs it with
`/dev/shm`, macOS with POSIX shm, Windows with the cosmo shim.
- **Rejected: unix sockets / named pipes** β message passing + broker
election we don't need for a shared counter table; 3 platform code paths.
- **Rejected: filesystem lock/registry files** β stale-file cleanup, dir
permission variance, and mtime-based liveness are all worse than pid+
heartbeat words in shm.
- **Rejected: CUDA MPS `CUDA_MPS_ACTIVE_THREAD_PERCENTAGE`** β Linux-only,
requires an MPS control daemon (not zero-conf), static percentages (no
idle-peer reflow), nothing on Windows/macOS.
- **No locks in the segment**: fixed-size slots, single-writer-per-slot
(each process writes only its own slot), seq/heartbeat words with atomic
stores; readers tolerate torn strings (name is display-only). A crashed
peer simply stops heartbeating and ages out β nothing to clean.
## 3. Registry layout
Segment name: `/opencoti-gpu-<gpu_key>` (one per physical GPU), 4 KiB.
```
header: { magic "OCGPU", version u32, slot_count u32 = 16 }
slot: { pid u32 (0 = free, CAS-claimed), boot_us i64,
weight f32, active u32 (decoded within window),
heartbeat_us i64 (atomic store, ~500 ms cadence),
busy_ppm u32 (measured busy fraction, parts-per-million),
name char[48] (alias/model, display only) }
```
- **gpu_key**: CUDA device UUID via a new DSO export
`ggml_backend_cuda_get_device_uuid` (we own the published `.so`/`.dll`
builds β additive symbol; when absent in an older side-loaded DSO, fall
back to a hash of device description + PCI ordinal). Metal β `metal-0`
(Apple GPUs are single-device). CPU-only runs register nothing.
- **Liveness**: heartbeat age > 3 s β dead; its slot is reclaimable by CAS.
Last liver does NOT `shm_unlink` (races the next joiner); a 4 KiB
segment persisting until reboot is harmless and self-heals via the
heartbeat rule.
## 4. Pacing (compute-share enforcement)
Since c6 (patch `0143`) pacing is **deterministic weighted time-division
(TDM)** over a host-shared wall clock β not feedback control:
1. Each peer derives the SAME schedule from the shm registry + the host
clock: a repeating period carved into one **contiguous window per
alive+active peer**, proportional to its (10%-floored, renormalized)
share, ordered by slot index. The period is **adaptive for
latency**: a peer's worst token stall is `(1βshare)Β·period`, so the
period is the smallest that still gives the smallest share a
contiguous-effective window (40 ms β bug-2243's floor):
`period = clamp(40ms / min_share, 80ms, 400ms)`. 2:1 weights β a
120 ms period (worst stalls 80/40 ms); 1:1 β 80 ms (40 ms); only
extreme splits reach 400 ms.
2. In the decode loop, after each step: if the current wall-clock
position is inside my window β keep decoding; otherwise sleep exactly
until my next window starts. Overshoot into a peer's window is β€1
decode step.
3. Alone (or all peers idle) β no schedule, never sleep. Idle peers
leave the schedule via the 2 s activity window; kill -9 peers via the
3 s heartbeat age.
Exactly one peer runs at any instant, so the pacer is **work-conserving**
(aggregate β solo throughput) and the split equals the weight ratio **by
construction** β there is no control loop to converge or oscillate.
Why not feedback? Three control laws failed in sequence:
- Proportional sleep `dtΒ·(busy/shareβ1)` has no fixed point (bug-2243).
- Integral debt on **absolute** busy (`debt += dt β shareΒ·wall`, the c5
pacer) over-throttles the minority peer once busy is measured honestly
(2.45:1 on the 3090 / 3:1 on the 5080 at 2:1 weights): contended,
every peer's `llama_decode` wall includes waiting on the OTHER peer's
kernels, so busy reads β wall for everyone (bug-2244 measurement is
correct β the signal itself is contention-inflated).
- Integral debt on **relative** busy (my busy / Ξ£ peers' busy) fixes the
bias but lets both peers sleep in overlapping windows (500 ms
busy-publish lag β oscillation): GPU idled ~half the time, aggregate
73 vs 116 tok/s.
Two clock/measurement invariants remain load-bearing:
- **Shared epoch (bug-2245):** `ggml_time_us()` is per-process-epoch on
Windows (native `timer_start`, and cosmo's `CLOCK_MONOTONIC` alike),
so shm heartbeat ages misread by the peers' launch delta β a pair
started >3 s apart never saw each other alive (latent in c5), and TDM
windows landed at random offsets (5080 aggregate 53 vs 146). All shm
timestamps + the schedule use `wall_us()` = `CLOCK_REALTIME` (Unix
epoch, shared by every process on the host, on every platform).
- **Busy through `llama_synchronize` (bug-2244):** `llama_decode` is
async-submit; timing the call alone reported 4.7% busy at 93% real
utilization. Busy (published for `/gpu/peers` and the activity flag)
spans decode through a sync, which is imminent at sampling anyway.
**Measured (TDM, c6, adaptive period):**
- solidPC RTX 3090, Qwen3-4B Q4_K_M batch-1: solo 121 tok/s; 2:1 β
75.6/38.4 (**1.97:1, aggregate 114.0** vs the c5 feedback pacer's
107.9); with a deliberate 20 s start skew β 1.99:1; 1:1 β 58.9/59.0
(**aggregate 117.9** vs 100.8, +17%); 2/1/1 β 56.5/27.3/29.5
(1.92:0.92:1, agg 113.3); solo-after-release 120.9.
- **Inter-token latency (streaming, 2:1):** w2 p50 9 ms / p95 45 / max
49 ms; w1 p50 9 ms / p95 86 / max 89 ms β exactly the
`(1βshare)Β·period` bound. (The initial fixed 400 ms period stalled
the w1 stream up to 267 ms; the c5 pacer's blocks were 100β150 ms.)
- **`--parallel` multi-slot (auto n_parallel, 4 concurrent gens each),
2:1:** w2 agg 217.1 / w1 agg 125.6 (**1.73:1, total 342.7** vs the
c5 pacer's ~211β246 total; solo 4-way is 327.5 β TDM total is at
solo level, i.e. fully work-conserving). Ratio softens at multi-slot
because a batched iteration overshoots the window boundary by more
than one token β accepted trade-off.
- **MTP (Qwopus3.5-9B NextN, `--spec-type draft-mtp`, `-c 8192`), 2:1:**
74.8/35.2 (**2.13:1, aggregate 110.0 β 96% of the 114.7 solo**);
draft acceptance unharmed (0.91/0.82 vs 0.87 solo).
- pandorum RTX 5080 (Windows, side-load DLL), 2:1 β 97.7/48.9
(**2.00:1, aggregate 146.6 β 97% of the 151 tok/s solo**; the c5
fixed-block pacer gave 3:1 at aggregate ~99). Windows' 15.6 ms timer
granularity is immaterial at 40β80 ms sleeps.
- Gate discipline (bug-2246): on a 16 GB card set an explicit `-c` β
VRAM auto-fit filled 15.8/16 GB and WDDM paging collapsed the
aggregate to 53 tok/s (an artifact of the gate setup, not the pacer).
Also gate with AUTO `n_parallel`: an explicit `--parallel 4` takes a
non-unified-KV path that decodes at ~28 tok/s aggregate regardless of
pacing (bug-2247, pre-existing, out of gpu-share scope).
Prefill batches count as busy time; the 10% share floor guarantees no
peer can be starved by misconfigured weights.
### Driver-level alternatives evaluated (2026-07-23)
Could the sleeps be replaced by CUDA-native compute partitioning?
Findings (CUDA 13.x):
- **MPS static SM partitioning (CUDA 13.1+, Linux, Ampere+):** the
control daemon (`nvidia-cuda-mps-control -d -S`) can carve
**exclusive SM partitions** per MPS client (`sm_partition add`),
deterministic and concurrent β no sleeps, no stall windows. This is
the real long-term candidate for Linux fleets (bs2-class); requires
daemon lifecycle management (a natural opencoti-server F3 feature)
and an r590+ driver (solidPC's 580 predates it; bs2's 610 has it).
Blackwell adds MLOPart (memory-locality-optimized partition devices).
- **Classic MPS + `CUDA_MPS_ACTIVE_THREAD_PERCENTAGE` (Volta+, Linux),
measured on the 3090 (driver 580):** 67/33 caps β 64.4/37.5
(**1.72:1, aggregate 101.9** β ~11% BELOW TDM's 114.0, because
batch-1 decode is bandwidth-bound, and concurrent kernels contend on
memory bandwidth while SM caps only bound compute). Latency is
smoother (max gap 50/53 ms vs TDM's 49/89) but p50 worsens for the
small partition (16 ms vs 9). Critically **not work-conserving**: the
caps are static β with the peer idle the 67% client does 110.9 tok/s
vs TDM's full 120.9, and ratios track SM fraction only loosely.
- **Green contexts (CUDA 13.1 runtime "execution context" API):**
SM partitioning **within a single process only** β inapplicable to
cross-process gpu-share (relevant someday for multi-model-in-one-
process serving). Isolation is best-effort, not guaranteed.
- **Windows:** no MPS, no cross-process partitioning at all β WDDM
time-slicing is the only primitive, so wall-clock TDM remains the
only portable mechanism there.
Conclusion: TDM stays the default (portable, zero-setup,
work-conserving, exact ratios, instant idle-release). An opt-in
Linux MPS mode (static SM partitioning managed by opencoti-server) is
a future milestone for latency-critical fleets on r590+ drivers.
## 5. Visibility
- `GET /gpu/peers` β
`{ gpu_key, self: {pid, name, weight, share_pct, busy_pct, active},
peers: [{pid, name, weight, share_pct, busy_pct, active,
heartbeat_age_ms}] }`
- `/props` `opencoti` block gains `gpu_share` carrying the same
`{enabled, gpu_key, self, peers}` snapshot as `/gpu/peers`, so existing
dashboards see it without a new call.
- Boot log: `gpu-share: joined /opencoti-gpu-<key> slot 2 (weight 1.5, 3
peers alive)`.
## 6. Phases
- **G1 β registry + flag + visibility (no throttle).** shm module
(`llamafile/gpu_share.c`, cosmo-portable), `--gpu-share-weight`, join/
heartbeat/reap, `/gpu/peers` + `/props` block. Gate: 3 instances on one
GPU (solidPC 3090), each sees the other two with correct resolved
percentages; kill -9 one β peers age it out β€ 3 s.
- **G2 β duty-cycle pacing.** Busy-window measurement + inter-batch sleep.
Gate: 2 instances, weights 2:1, concurrent decode workloads β measured
tps ratio β 2:1 (Β±20%); solo instance regains ~full solo tps β€ 2 s after
peer goes idle.
- **G3 β Windows/macOS validation** (pandorum APE run for the shm path on
Windows; macOS opportunistic) + docs + opencoti-server fleet surfacing
(daemon shows per-engine share/busy from `/gpu/peers`).
## 7. Non-goals
- Hard isolation (MPS/MIG) β not zero-conf, not portable.
- VRAM partitioning by weight β allocation is boot-time (`-c`, `-ngl`);
the endpoint may later *report* VRAM per peer but never enforces it.
- Cross-GPU balancing β the registry is strictly per-GPU; multi-GPU
processes join one segment per device they occupy (v1: primary device
only).
|