bogdanraduta commited on
Commit
9c8e219
·
verified ·
1 Parent(s): 6ebd961

Add fp16 merged model (Qwen3-4B LoRA)

Browse files
Files changed (1) hide show
  1. README.md +4 -82
README.md CHANGED
@@ -1,87 +1,9 @@
1
  ---
 
2
  license: apache-2.0
3
- language:
4
- - en
5
- base_model: Qwen/Qwen3-4B
6
- library_name: transformers
7
  pipeline_tag: text-generation
 
8
  tags:
9
- - lora
10
- - regulatory
11
- - compliance
12
- - escalation
13
- - decision-gate
14
- - mlx
15
- - gguf
16
- - flowx
17
  ---
18
-
19
- # FlowX Sentinel Gate (4B)
20
-
21
- **FlowX Sentinel Gate** is an escalation decision gate for regulated workflows: given a
22
- complete case (domain facts + the applicable policy schema), it decides **ESCALATE** (route
23
- to a human) vs **DECIDE** (safe to automate), and for escalations returns a **category**, a
24
- structured rationale, a **calibrated confidence**, the required human action, and an audit
25
- trail. It is a LoRA fine-tune of **Qwen3-4B**, built by FlowX.AI to sit after the
26
- [Semantic Mapper](https://huggingface.co/flowxai/semantic-mapper) in a compliance pipeline.
27
-
28
- > Decision-support tool, **not legal advice**. It gates automation; a human owns escalated
29
- > cases. Deploy with the confidence threshold your risk posture requires.
30
-
31
- ## What it does
32
-
33
- Input: a case JSON (facts + `policy_schema`) + `Decide: ESCALATE or DECIDE?`. Output: an
34
- oracle decision. For ESCALATE: `action`, `escalation_category` (one of six), the
35
- category-specific block, `confidence_score`, `confidence_reasoning`, `human_action_required`,
36
- `audit_trail`. For DECIDE: the safe-to-automate rationale + confidence + audit trail.
37
-
38
- **Six escalation categories:** MISSING_REQUIRED_DOCUMENTATION, POLICY_VIOLATION,
39
- BOUNDARY_CONDITION, INSUFFICIENT_CONFIDENCE, CONFLICTING_SIGNALS, EXTERNAL_DEPENDENCY.
40
-
41
- ## Evaluation (held-out, n=71: 46 escalate / 17 decide)
42
-
43
- | Metric | Result | Note |
44
- |---|---|---|
45
- | **False-negative rate** | **0.000** (0/46) | never misses an escalation — the safety-critical metric |
46
- | **Action accuracy** (ESCALATE vs DECIDE) | **1.000** | |
47
- | **False-positive rate** | **0.000** (0/17) | never over-escalates |
48
- | JSON validity | 0.89 | structured output; pair with the deterministic repair step |
49
- | Category accuracy (on true-escalate) | 0.61 | the human-routing label; see limitations |
50
-
51
- The decision itself (automate vs escalate) is the gate's job, and it is **perfect on this
52
- held-out set** — no missed escalations, no over-escalation, action decided correctly
53
- every time. The `escalation_category` is a secondary routing hint and is right ~61% of the
54
- time (categories legitimately overlap for some cases).
55
-
56
- ## Files & formats
57
-
58
- | Path | Format | Runs on |
59
- |---|---|---|
60
- | `/` (root) | fp16 safetensors | CUDA / servers, vLLM |
61
- | `mlx-int4/`, `mlx-int8/` | MLX quantized | Apple Silicon |
62
- | `gguf/*.gguf` | GGUF Q8_0 / Q4_K_M | CUDA + CPU (llama.cpp / Ollama) |
63
-
64
- System prompt: `You are an escalation gate for regulated decisions.\nDetermine: ESCALATE or DECIDE? Output ONLY JSON.`
65
- (decode with `enable_thinking=False`).
66
-
67
- ## Training
68
-
69
- LoRA (rank 32 / scale 16 / dropout 0.05), Qwen3-4B base, MLX-LM, cosine LR 5e-5→5e-6,
70
- max sequence length 2048. Data: **471 realistic-synthetic escalation cases** (400 train / 71
71
- held-out), balanced ~35% DECIDE / 65% ESCALATE across the six categories and four regulated
72
- domains (banking, insurance, logistics, labor), each grounded in a real regulatory citation.
73
-
74
- ## Limitations & responsible use
75
-
76
- - **Not legal advice** — an automation gate; escalated cases must be handled by a human.
77
- - **Category label is ~61% accurate** — use the ESCALATE/DECIDE decision (perfect on
78
- held-out) as the gate; treat the category as a routing suggestion, not ground truth.
79
- - JSON validity is 0.89 — deploy with the deterministic repair/retry step.
80
- - Evaluated on 71 synthetic cases; validate on your own case distribution before production.
81
-
82
- ## License & attribution
83
-
84
- Apache-2.0. Copyright 2026 FlowX.AI. See `NOTICE`. Base model: Qwen3-4B (Apache-2.0).
85
- Escalation scenarios are realistic synthetic, grounded in real regulatory citations.
86
-
87
- _Author: Bogdan Răduță, Head of Research, FlowX.AI._
 
1
  ---
2
+ library_name: mlx
3
  license: apache-2.0
4
+ license_link: https://huggingface.co/Qwen/Qwen3-4B/blob/main/LICENSE
 
 
 
5
  pipeline_tag: text-generation
6
+ base_model: mlx-community/Qwen3-4B-4bit
7
  tags:
8
+ - mlx
 
 
 
 
 
 
 
9
  ---