CloudSurf-4B-FC / README.md
brady777's picture
Card refresh: PR #1357 link, evidence dataset links (BFCL results + tau2 trajectories), fairness-pair disclosure updated to measured, fix stale refs
c60ad48 verified
|
Raw
History Blame Contribute Delete
7.74 kB
---
# Hugging Face model card — CloudSurf-4B-FC (E1-rev champion, 2026-08-17).
# Name RULED 2026-08-17 (Brady): CloudSurf-4B-FC under HF org `cloudsurf-software`
# (namespace `cloudsurf` was already taken on HF by an unrelated account).
# "4B" = effective active params (the base's own E4B convention); total params
# ~8.0B are disclosed below.
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
base_model: google/gemma-4-E4B-it
base_model_revision: ee0ef6023621cff504d758262d4e04895a5af4a2 # HF main sha at release window (boxes pulled main 2026-08-16/17)
tags:
- function-calling
- tool-use
- bfcl
- cloudsurf
- qlora
- gemma-4
datasets: [] # training mix is proprietary (fc-tier1); generators + gates documented in-repo
model_name: CloudSurf-4B-FC
inference: false
---
# CloudSurf-4B-FC — function-calling / tool-use model
> Gemma-4 E4B (effective-4B active; ~8.0B total params) QLoRA-tuned by
> CloudSurf Software for function calling and agentic tool use. **BFCL V4
> FULL (harness pinned `f7cf735`): 3-seed mean 55.73 vs stock 34.81 (+20.9)**
> — above the published small-model class bar (Nanbeige4-3B 51.40) and above
> stock gpt-oss-20b (49.09) measured on the identical rig. Self-run numbers
> with full settings disclosure; gorilla PR open
> ([ShishirPatil/gorilla#1357](https://github.com/ShishirPatil/gorilla/pull/1357),
> official run invited).
## Scores (BFCL V4 FULL 22/22, prompt-mode, thinking ON, temp 0.001)
| | Overall | Non-Live | Live | Multi-Turn | Web | Memory | Irrelevance |
|---|---:|---:|---:|---:|---:|---:|---:|
| **CloudSurf-4B-FC (3-seed mean)** | **55.73** | 87.23 | 79.82 | 43.46 | 48.00 | 41.65 | 80.63 |
| Stock gemma-4-E4B-it (×3 mean) | 34.81 | 84.30 | 73.97 | 19.46 | 7.00 | 16.56 | 84.37 |
Raw result files for every run in this table — ours **and** the stock
baselines, plus per-category scores and a comparison chart — are public:
[CloudSurf-4B-FC-bfcl-results](https://huggingface.co/datasets/cloudsurf-software/CloudSurf-4B-FC-bfcl-results).
Seeds {42, 7, 11} → {55.53, 54.86, 56.81}, spread 1.95. Best-checkpoint
(seed 11) re-eval: two further same-box FULL evals returned 54.39 and 54.91 —
measured same-checkpoint eval band {56.81, 54.39, 54.91}, mean 55.37,
sd ≈ 1.05 (per-run BFCL FULL eval noise on this model class is ≈ ±1;
single runs < ~2.5 apart are not meaningfully different). Every measured
number, including the weakest (54.39), exceeds the published small-model
class bar (Nanbeige4-3B, 51.40) and stock by double digits.
Contamination-audited clean (0/255 eval-prompt 8-gram overlap; behavioral
trace verification).
**Serving-variant fairness pair** (pre-registered, measured 2026-08-17):
stripping stray `<tool_call|>` closers — a decode artifact of the base's
native prior — raises STOCK to 39.40 (Multi-Turn 19.46 → 33.12) and moves
this model +0.52 (→ 54.91). Under that matched best-serving frame the honest
comparison is **54.91 vs 39.40 (+15.5)**; the as-registered frame is 55.73
vs 34.81 (+20.9). Both are reported everywhere; the strip ships in the
proposed BFCL handler for all Gemma-4 prompt-mode models.
**Honesty disclosures** (full apparatus in the forthcoming paper):
1. Web 48.00 is a *floor*: ~9% of web_no_snippet entries overflow the 131K
context window mid-research (multi-hop page fetches) and score 0; the
errors bias against this model (clean ceiling ≈ 50.5).
2. Memory gain is trained env-API-namespace competence (schemas
independently authored, namespace deliberately aligned) — legitimate but
not zero-shot API transfer. Web *is* zero-shot (no web data in the mix).
3. Format Sensitivity (excluded from Overall by the harness) regresses:
FS max-delta ~45–81 vs stock ~10 — the tune is format-specialized to its
registered serving template.
4. A pre-registered serving-variant fairness pair (stray `<tool_call|>`
closer strip, which also rescues stock's native decode quirk) was
measured 2026-08-17 and is disclosed above (54.91 vs 39.40); headline
numbers use the registered handler.
## Training
- Adapter: QLoRA nf4, LoRA r8 α16, all 7 linear projections, lr 5e-5,
686 steps (~0.16 epochs of 34,926 per-turn examples), assistant-only loss.
- Data: 2,747-row fc-tier1 mix (`7f8711e00f66bad0`) — state-verified
gpt-oss-120b teacher trajectories (partially-observable look-then-act envs),
memory-discipline conversations, irrelevance-refusal rows; teacher reasoning
retained in the thought channel.
- The decisive technique: **supervised span termination (render v2)** — one
training example per assistant turn, history byte-exact to the serving
handler, every supervised span terminated with a serve-time stop id
(`<turn|>`). Full apparatus in the forthcoming paper.
## Repository contents
- `merged/` — BF16 merged champion checkpoint (seed 11).
- `adapters/seed-{42,7,11}/` — the three LoRA adapters (111M each) for
repeatability; merge with `peft` or `scripts/fastloop_merge_adapter.py`.
- This card. Training data is not released; generators, gates, and the full
measurement record are described in the paper (same posture as early
xLAM/ToolACE releases).
## Release checklist (state at card freeze)
1. ✅ V-3 live board re-verify (2026-08-17 pull; snapshot archived).
2. ✅ V-4 FS-exclusion audit (Overall recomputed exactly from leaves).
3. ✅ V-6 web-poison mechanism (context overflow; biases against us).
4. ✅ V-5 name ruled: CloudSurf-4B-FC @ hf.co/cloudsurf-software.
5. ✅ V-1 best-checkpoint re-eval — pre-registered ±0.6 clause MISSED
(54.39/54.91 vs 56.81); adjudicated as eval-noise mis-calibration
(poison byte-identical across runs, Multi-Turn stable); measured band
disclosed above in place of the clause.
6. ✅ V-2 fairness pair measured: stock 39.40 / champion 54.91 under the
strip variant — both frames disclosed above.
7. ✅ Gorilla PR opened 2026-08-18:
[ShishirPatil/gorilla#1357](https://github.com/ShishirPatil/gorilla/pull/1357)
(handler ships the strip per V-2; raw results dataset linked above).
## Serving — OpenAI-compatible endpoint (τ²-bench setup)
The model's native interface is prompt-mode (Gemma-4 turn format, bracket
tool-call lists). For harnesses that speak the OpenAI chat/tools API —
including [tau2-bench](https://github.com/sierra-research/tau2-bench) — this
repo ships the exact serving bridge used for our τ²-bench leaderboard runs:
[`serving/fc_tau2_bridge.py`](./serving/fc_tau2_bridge.py) (stdlib-only, no
dependencies). Full τ²-bench trajectories (retail / airline / telecom, ×4
trials, unmodified) are public:
[tau2-trajectories-cloudsurf-4b-fc](https://huggingface.co/datasets/cloudsurf-software/tau2-trajectories-cloudsurf-4b-fc).
```bash
# 1) serve the merged weights with any /v1/completions backend, e.g. sglang:
python -m sglang.launch_server --model-path cloudsurf-software/CloudSurf-4B-FC \
--revision main --port 30000 # serve the merged/ subtree
# 2) put the bridge in front (OpenAI-compatible /v1/chat/completions, tools in / tool_calls out):
python serving/fc_tau2_bridge.py --port 8000 --backend http://127.0.0.1:30000
# 3) point any OpenAI client at http://127.0.0.1:8000/v1 (model id: cloudsurf-4b-fc)
python serving/fc_tau2_bridge.py --selftest # render/parse self-checks
```
Disclosure (mirrors our τ²-bench submission methodology): the bridge appends a
fixed function-calling instruction block to the caller's system message, and
when a generation terminates inside the thought channel with an empty answer it
attempts to recover a bracket-format tool-call list from the raw text before
returning an empty answer. Both behaviors are in the script, unabridged.