# DiffusionGemma Humanizer — Architecture & Strategy > **Date:** 2026-06-29 > **Budget:** $250 Modal credits > **Target:** GPTZero, GPTDetector, Pangram, Binoculars > **Objective:** Fine-tune DiffusionGemma to humanize AI-generated text and evade detection --- ## 1. Why Diffusion Models Beat Detectors (Paper-Backed) ### 1.1 The Theoretical Ceiling (Sadasivan et al., 2023) **Theorem 1** proves mathematically: AUROC ≤ 1/2 + TV(M,H) − TV(M,H)²/2, where TV(M,H) is the total variation distance between machine and human text distributions. As models improve and TV → 0, every detector converges to random guessing (AUROC → 0.5). This is not a technological limitation — it's a mathematical certainty. [arXiv:2303.11156] **Implication:** 0% detection (AUROC = 0.5) is achievable asymptotically. A model that perfectly mimics human text distribution is undetectable by definition. We don't need perfection today — we just need TV(M,H) to be lower than what detectors have been trained to distinguish. ### 1.2 Pangram's Blind Spot: No Diffusion Text Pangram's DAMAGE detector (Mistral NeMo 12B + LoRA + linear head) was trained exclusively on autoregressive text: GPT-3.5/4/4o, Claude 2/3, LLaMA 2/3/3.1, Mistral, Gemini Pro/Flash. Their "synthetic mirror prompts" all target instruction-tuned autoregressive models. **Zero diffusion-generated text in their training set.** [arXiv:2501.03437] Key finding from their adversarial attack (§7): - When they fine-tuned GPT-4o specifically against their detector: TPR dropped from 100% → 93.2% - This was with a generic fine-tune API, not a custom architecture - A purpose-built diffusion humanizer would exploit a fundamentally different statistical pathway they've never seen ### 1.3 Adversarial Paraphrasing Works (Cheng et al., 2025) The closed-loop optimization is the secret: paraphrasing WITHOUT detector guidance **increases** detection by 8-15%. With detector feedback guiding the paraphrase, TPR drops by **87.88% average**. The innovation is in the feedback loop, not the base model. [arXiv:2506.07001] **Our advantage:** We can embed this feedback loop *inside* the diffusion denoising process, eliminating the need for an external autoregressive LLM that would leave AR-specific traces. ### 1.4 Diffusion Text Naturally Evades AR-Trained Detectors (Tarım & Onan, 2025) LLaDA-generated text "closely mimics human text in perplexity and burstiness, yielding high false-negative rates for AR-oriented detectors." Diffusion models generate through iterative denoising — left-to-right token probabilities don't exist, so perplexity-based methods break. [HF Papers:2507.10475] --- ## 2. Architecture Proposal ### 2.1 Base Model: DiffusionGemma 26B | Property | Value | |----------|-------| | Total params | 25.2B | | Active params | 3.8B (MoE, 8/128 experts + 1 shared) | | Generation | Discrete diffusion, block-autoregressive, 256-token canvases | | Attention | Bidirectional in decoder, cross-attention to AR encoder | | Sampling | Entropy-Bounded Denoising, 1–48 steps, temperature 0.8→0.4 | | License | Apache 2.0 | | Release | June 9, 2026 | **Why this model:** 1. Generation pathway is fundamentally different from AR models → OOD for all current detectors 2. Bidirectional canvas attention enables holistic text understanding for rewriting 3. Iterative denoising (1-48 steps) provides a perfect mechanism for progressive humanization 4. MoE (3.8B active) makes QLoRA fine-tuning feasible on A100 5. 4-bit quantized versions already exist (NVFP4 from NVIDIA, AutoRound from Intel) ### 2.2 Training Strategy **Phase 1: Dataset Construction** - Base: HC3 (48.6K ChatGPT-human QA pairs) - Augmentation: Synthetic pairs from heuristic humanization (sentence splitting, hedging, register variation) - Detector scoring: Each pair annotated with pre/post scores from Fast-DetectGPT, burstiness, heuristics **Phase 2: QLoRA Fine-Tuning** - Method: QLoRA, rank 8-16, 4-bit base (NVIDIA NVFP4) - Format: Chat template (system: "Rewrite to sound human", user: AI text, assistant: human text) - Initial scale: 300-500 pairs, A100 80GB, ~$40 **Phase 3: Detector-Guided Refinement** - At inference, use entropy-bounded denoising (48 steps max) - Apply classifier-free guidance to push logits toward human-like tokens - Target inversion of known detector signals: - Reduce transition markers (furthermore, however, therefore) - Increase sentence length variation - Reduce passive voice constructions - Increase hapax legomena ratio ### 2.3 Detector Signals to Invert From the literature, detectors (especially Pangram) target: | Signal | AI Text | Human Text | Inversion Strategy | |--------|---------|------------|-------------------| | Perplexity | Low (<20) | High (>30) | Add "surprising" tokens | | Burstiness | Low (<0.15) | High (>0.3) | Vary sentence structure | | Transition markers | Many | Few | Replace with informal alternatives | | Passive voice | Many | Few | Convert to active voice | | Sentence length σ | Low (<5) | High (>10) | Mix short/long sentences | | Hapax ratio | Low (<0.40) | High (>0.45) | Use more unique words | | Lexical diversity | Moderate | Variable | Match domain-specific patterns | --- ## 3. Cost Breakdown | Step | GPU | Est. Time | Cost | |------|-----|----------|------| | 1. Model download + baseline gen | A100 | 30 min | $1.25 | | 2. Detector tests | A100 | 20 min | $0.85 | | 3. Dataset construction | CPU (within A100) | 10 min | $0 | | 4. QLoRA fine-tuning (300 pairs) | A100 | 1-2h | $5-10 | | 5. Evaluation (5 prompts) | A100 | 15 min | $0.60 | | 6. Export to HF | CPU | 5 min | $0 | | **Total per run** | | **~3h** | **~$8-13** | | **Available runs ($250)** | | | **19-31 experiments** | Each experiment can vary: dataset size, LoRA rank, learning rate, prompt template, training epochs. --- ## 4. Expected Results (Hypothesis) Based on cross-referencing all papers: | Detector | Baseline (DiffusionGemma raw) | Fine-tuned Humanizer | |----------|------------------------------|---------------------| | GPTZero | 60-80% evasion (diffusion confuses perplexity) | 95%+ evasion | | Binoculars | 50-70% evasion (cross-perplexity mismatch) | 90%+ evasion | | Fast-DetectGPT | 40-60% evasion (curvature is AR-specific) | 85%+ evasion | | Pangram | 30-50% evasion (untrained on diffusion) | 70-85% evasion initially | Pangram will be the hardest because their deep classifier learns nuanced style patterns. But being completely OOD (out-of-distribution) for their training set gives a significant initial advantage. The window of high evasion will close once Pangram adds diffusion text to their training pipeline — estimated 1-3 months after public release. --- ## 5. Security: HF Token Setup ### Required Permissions 1. **Read token**: Download models (DiffusionGemma, GPT-2 for detectors) 2. **Write token**: Upload to `simonlesaumon/diffusiongemma-humanizer` ### Setup on Modal (nothing stored locally) ```powershell # Create Modal secret (tokens stay encrypted on Modal, never on your PC) modal secret create hf-secrets HF_TOKEN=hf_your_token_here ``` Then the app reads from `os.environ["HF_TOKEN"]` automatically. Create your token at: https://huggingface.co/settings/tokens (needs "Write" permission) --- ## 6. File Structure ``` Deepseek_SOTA/ ├── modal_project/ │ └── app.py # Full Modal pipeline ├── scripts/ │ ├── run_pipeline.ps1 # PowerShell launcher │ └── run.py # Python launcher ├── research/ │ ├── technical-diffusion-text-humanization-2026-06-29.md # Research report │ └── architecture-strategy.md # This file └── .gitignore ``` **Everything runs on Modal. Nothing downloaded locally except the scripts.** --- ## 7. References 1. Sadasivan et al. (2023) — *Can AI-Generated Text be Reliably Detected?* [arXiv:2303.11156] 2. Masrour, Emi & Spero (2025) — *DAMAGE: Detecting Adversarially Modified AI Generated Text* [arXiv:2501.03437] 3. Cheng et al. (2025) — *Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text* [arXiv:2506.07001] 4. Tarım & Onan (2025) — *Can You Detect the Difference? Diffusion vs AR Text* [HF:2507.10475] 5. Ranganath & Ramesh (2026) — *StealthRL: RL Paraphrase Attacks for Multi-Detector Evasion* [arXiv:2602.08934] 6. Xu et al. (2026) — *Base Models Look Human To AI Detectors* [arXiv:2605.19516] 7. Schiff et al. (2024) — *Simple Guidance Mechanisms for Discrete Diffusion Models* [arXiv:2412.10193] 8. DiffusionGemma — [Hugging Face](https://huggingface.co/google/diffusiongemma-26B-A4B-it)