Instructions to use danielfein/raid-ce-gemma4-e4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use danielfein/raid-ce-gemma4-e4b with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("danielfein/raid-ce-gemma4-e4b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| base_model: google/gemma-4-E4B-it | |
| datasets: | |
| - liamdugan/raid | |
| library_name: transformers | |
| tags: | |
| - ai-text-detection | |
| - detection-tokens | |
| - neologism | |
| - raid | |
| # Gemma-4-E4B RAID-only CE detection tokens | |
| This repository contains the locked RAID-only pairwise-CE detection-token | |
| checkpoints for `google/gemma-4-E4B-it`. The detector score for a text `x` is | |
| ```text | |
| mean_logp(x | "Write <ai> text.") - mean_logp(x | "Write <human> text.") | |
| ``` | |
| The default files `ai_token.pt` and `human_token.pt` are the seed-707 | |
| checkpoint, selected because it has the highest full-BEEMO AUROC among the | |
| three completed seeds. Complete checkpoints for seeds 42, 101, and 707 are in | |
| `seeds/`. | |
| ## Locked training procedure | |
| - Data: RAID `standard_train_expanded6` only. | |
| - Held-out evaluation: RAID `standard_test`; train/test source-ID overlap is | |
| asserted to be zero. | |
| - Objective: reference-free pairwise logistic CE, | |
| `softplus(-(mean_logp(ai_text | AI prompt) - | |
| mean_logp(human_text | AI prompt)))`. | |
| - Trainable parameters: the main `<ai>` embedding row only. The `<human>` row, | |
| all backbone weights, and Gemma's auxiliary embedding rows remain fixed. | |
| - Prompt: `Write {token} text.` | |
| - Sequence score: average continuation-token log likelihood. | |
| - Maximum length: 512 tokens. | |
| - Batch size: 8. | |
| - Pilot: 500 randomly sampled RAID pairs, 2 epochs (125 updates), AdamW, | |
| peak LR `1e-3`, 20-update linear warmup, cosine decay to `1e-5`. | |
| - Continuation: 5,000 randomly sampled RAID pairs, 75 updates, peak LR | |
| `1e-4`, 20-update warmup, using a 625-update cosine schedule horizon and | |
| stopping after update 75. | |
| - Seeds: 42, 101, and 707; the seed controls RAID sampling and evaluation | |
| bootstrap resampling. | |
| - Evaluation: all 2,163 eligible BEEMO pairs and all 3,000 RAID standard-test | |
| pairs, with 1,000 paired bootstrap resamples per seed. | |
| The canonical model-aware entry point is | |
| `scripts/run_raid_ce_canonical.py`. It also contains the locked Llama profile; | |
| the lower Llama learning rate and omission of continuation are explicit in | |
| that profile. | |
| ## Results | |
| | Seed | BEEMO AUROC | RAID test AUROC | | |
| |---:|---:|---:| | |
| | 42 | 0.7549 | 0.9642 | | |
| | 101 | 0.7565 | 0.9503 | | |
| | 707 | **0.7689** | 0.9599 | | |
| | Mean +/- sample SD | **0.7601 +/- 0.0077** | **0.9581 +/- 0.0071** | | |
| Per-seed paired-bootstrap confidence intervals and complete metric fields are | |
| stored in `metrics.json`. | |
| ## Files | |
| ```text | |
| ai_token.pt | |
| human_token.pt | |
| seeds/{42,101,707}/{ai_token.pt,human_token.pt,summary.json} | |
| metrics.json | |
| training_config.json | |
| scripts/run_raid_ce_canonical.py | |
| scripts/run_gemma_expanded_pairwise_ce.py | |
| detection_tokens/ | |
| ``` | |
| Each token checkpoint includes the main embedding and Gemma auxiliary | |
| embedding row needed by the loader. These are embedding-only artifacts and | |
| require access to the Gemma backbone. | |