File size: 6,068 Bytes
b3692cd 0591569 b3692cd 7de84b2 9a6fd1a 43232e3 0591569 af1c50a d02ccd1 b3692cd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 | ---
language:
- en
- zh
- es
- fr
- pt
- de
- it
- nl
license: apache-2.0
library_name: dotcheck
pipeline_tag: text-classification
tags:
- text-classification
- ai-detection
- multilingual
- apache-2.0
- dotcheck
- valla
base_model:
- Oxidane/tmr-ai-text-detector
- fakespot-ai/roberta-base-ai-text-detection-v1
- hfl/chinese-macbert-base
base_model_relation: adapter
model-index:
- name: Valla (inhouse-text@11)
results:
- task:
type: text-classification
name: binary AI-likeness (text, en)
dataset:
name: DotCheck text holdout EN
type: other
split: holdout
metrics:
- name: mean_P_AI_human
type: mean_score_human
value: 0.005
- name: mean_P_AI_ai
type: mean_score_ai
value: 0.973
- name: balanced_accuracy
type: balanced_accuracy
value: 0.986
source:
name: text_gates_v11 / Data.json
url: https://dotcheck.ai/docs
---
# DotCheck/valla-text-v11
Apache-2.0 text AI-likeness heads for DotCheck (English, Chinese, and six Latin language heads). This repo includes the live .npz heads, model card, license, and notices.
| Field | Value |
|-------|--------|
| Hub id | `DotCheck/valla-text-v11` |
| Wire id (EN) | `inhouse-text@11` |
| Label | Valla |
| EN artifact | `text_stack_head_v11.npz` |
| Latin feature bases | [`Oxidane/tmr-ai-text-detector`](https://huggingface.co/Oxidane/tmr-ai-text-detector) (MIT), [`fakespot-ai/roberta-base-ai-text-detection-v1`](https://huggingface.co/fakespot-ai/roberta-base-ai-text-detection-v1) (Apache-2.0) |
| Chinese feature base | [`hfl/chinese-macbert-base`](https://huggingface.co/hfl/chinese-macbert-base) (Apache-2.0; `cls_emb`) |
| Heads | EN + `zh_v1` / `es_v3` / `pt_BR_v3` / `fr_v3` / `it_v3` / `de_v1` / `nl_v1` (separate `.npz`) |
| Output | `p β [0,1]` β P(AI-like) |
| Serve | CPU FastAPI `POST /v1/analyze-text` + required `lang` |
| `lang` | `en\|zh\|es\|pt\|fr\|it\|de\|nl` (`pt` β `pt_BR` head; `zh` = Simplified Chinese) |
## Model description
**Latin group:** Frozen TMR + Fakespot encoders β shared feature vector β language-specific logistic head.
**Chinese group:** Frozen MacBERT `cls_emb` β `zh_v1` logistic head (train=serve contract).
Serve loads the Latin base pair for all Latin heads; the Chinese tower loads when `text_stack_head_zh_v1.npz` is present. Unsupported `lang` β fail closed (`unsupported_language`; no silent EN fallback). Traditional-Chinese-dominant product input is not scored (`und`).
Text hygiene (TC1) applied before hash/score: strip wiki-style cites, collapse whitespace, drop `#category` tokens (mirrored in Express / extension).
**Files in this repo:** `README.md`, `LICENSE`, `NOTICE`, `CITATION.cff`, and the `.npz` head file(s) listed above.
## Architecture
```text
raw text
β text_clean (TC1)
β if lang in {en,es,pt,fr,it,de,nl}:
TMR + Fakespot frozen forward (shared Latin group)
β if lang == zh:
MacBERT cls_emb frozen forward (Chinese group)
β lang-selected logistic head (npz)
β p_AI
```
## Inference
Open weights: the live `.npz` head(s) in this repo (Apache-2.0), for use with the frozen upstream backbone named above. This is not a transformers `AutoModel.from_pretrained("DotCheck/β¦")` package.
Product scoring: Check or Pro API (below). Leviathan (shared memory and related product path) is not in these files.
```bash
curl -sS -X POST "https://dotcheck-server-c221c1f32c68.herokuapp.com/analyze-text" \
-H "Authorization: Bearer dc_YOUR_KEY" \
-H "Content-Type: application/json" \
-d "{\"text\":\"...\",\"lang\":\"en\"}"
```
UI: https://dotcheck.ai/check Β· API: https://dotcheck.ai/api Β· docs: https://dotcheck.ai/docs
Min length / FUP enforced at Express (product policy).
## Training data
| | EN / Latin | Chinese (zh-Hans) |
|--|----|-------------|
| Fit AI | Qwen2.5-7B + Mistral-7B (multi-temp) | Qwen2.5-7B + Mistral; wiki-style prompts |
| Holdout AI | **Qwen2.5-1.5B** | Qwen2.5-1.5B (~200/200) |
| Human | hard-neg wiki / WikiText / Gutenberg mix | Wikipedia lead prose (OpenCC t2s) |
| License | commercial-clean only (no NC banned sets) | same |
No published holdout from live ChatGPT/Claude scrapes. Evidence: `text_gates_v11.json`, `text_gates_{lang}.json`.
## Evaluation
### English (public claim SSOT)
| Metric | Target | Measured |
|--------|--------|---------:|
| mean P(AI) \| human | β€ 0.12 | **0.005** |
| mean P(AI) \| AI | β₯ 0.85 | **0.973** |
| bal_acc @ thr | β₯ 0.90 | **0.986** |
Also: beat prior `@6` holdout; wiki-monitor mean ~\<0.001; OOD / RAID-lite protocol OK (`TEXT_GATES_OK`).
### Language heads (holdout; same absolute floors)
<!-- sync:text-lang-start -->
| Lang | Wire | Human | AI | bal_acc |
|------|------|------:|---:|--------:|
| zh | `inhouse-text-zh_v1` | 0.033 | 0.973 | 0.972 |
| es | `inhouse-text-es_v3` | 0.029 | 0.993 | 0.982 |
| fr | `inhouse-text-fr_v3` | 0.021 | 0.990 | 0.978 |
| pt | `inhouse-text-pt_BR_v3` | 0.042 | 0.964 | 0.958 |
| de | `inhouse-text-de_v3` | 0.034 | 0.966 | 0.978 |
| it | `inhouse-text-it_v3` | 0.016 | 0.943 | 0.972 |
| nl | `inhouse-text-nl_v3` | 0.056 | 0.963 | 0.950 |
<!-- sync:text-lang-end -->
## Intended use
- Supported-language AI-likeness scoring in DotCheck inference.
- Citation of the tables above.
### Out of scope
- Product scoring SLA / Leviathan / FUP via Hub download
- Languages outside `{en,zh,es,pt,fr,it,de,nl}`
- Traditional Chinese as a scored variant (fail closed)
- Plagiarism / authorship adjudication
## Limitations
- Short, MT-heavy, or heavily edited text: higher variance.
- Eight languages only (Simplified Chinese for `zh`).
- Output is not generator attribution.
## License
[`LICENSE`](LICENSE) β Apache License 2.0 for DotCheck heads in this repo. Upstream backbones: see [`NOTICE`](NOTICE).
## Citation
[`CITATION.cff`](CITATION.cff) Β· wire `inhouse-text@11` / Valla@11 Β· https://dotcheck.ai/docs
|