IdeaLens / README.md
rishanthrajendhran's picture
Card: strict '<' flag rule, format-assignment wording, vLLM section; thresholds.json flag rule
11f78dc verified
|
Raw History Blame Contribute Delete
13.5 kB
metadata
license: other
license_name: openmdw-1.1
license_link: LICENSE
base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
library_name: transformers
language:
  - en
tags:
  - ai-text-detection
  - idea-provenance
datasets:
  - rishanthrajendhran/IdeaLens-1M
extra_gated_prompt: >-
  Access is granted individually. Please say who you are and what you intend to
  use the weights for.

IdeaLens

IdeaLens detects who came up with the ideas in a document, rather than who wrote its words. It reads a role-labelled outline of the document (an ordered list of items, each giving one idea and the discourse role it plays, such as Central Development or Open Question) and returns P(human), the probability that the ideas are human.

IdeaLens is nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 fine-tuned with LoRA (rank 64) on the outlines of 1M English web documents (IdeaLens-1M). The training outlines were paraphrased to remove the documents' wording, so the model has to fit its labels through the ideas.

Results

From the paper:

  • IdeaLens is accurate both when a document's ideas and words come from the same source (95.3%) and when they come from different sources (81.3%). ProseLens, trained identically on the raw text, reaches 99.1% and 25.4%; Pangram 4 reaches 98.5% and 25.9%.
  • As models write from increasingly detailed human plans, IdeaLens's AI flag rate falls from 95% to 7%, while Pangram 4 still flags 92%. From AI-derived plans, IdeaLens stays above 96%.
  • On TwiceTold, 50 stories that human authors wrote from AI-generated plans, IdeaLens flags 68% as AI, against 8% for Pangram 4.
  • On 19 existing detection benchmarks, IdeaLens keeps strong detection rates at low false-positive rates across domains, formats and languages.

Usage

Scoring a document takes two steps:

  1. Extract an outline. An LLM writes the outline from the document, its format's role vocabulary and six worked examples (the paper uses Gemini 3.7 Flash). The prompts and role vocabularies are in the code repository: link added on publication. Score the outline as extracted; the paraphrasing step is only for training data.
  2. Score the outline with this model, as below.

Load the merged model (66 GB download)

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("rishanthrajendhran/IdeaLens")
model = AutoModelForCausalLM.from_pretrained("rishanthrajendhran/IdeaLens", dtype=torch.bfloat16, device_map="auto").eval()

The weights take 59 GiB of GPU memory, and each input adds more; see Hardware requirements.

Or apply the adapter to the base model (3 GB download)

adapter/ holds the LoRA adapter as trained, in the layout of the Tinker training service. If you already have the base model, load_adapter.py merges the adapter into it in memory. The resulting weights are bit-identical to the merged model's:

import importlib.util
from huggingface_hub import hf_hub_download

path = hf_hub_download("rishanthrajendhran/IdeaLens", "load_adapter.py")
spec = importlib.util.spec_from_file_location("load_adapter", path)
la = importlib.util.module_from_spec(spec); spec.loader.exec_module(la)
model, tok = la.load_model()   # base model + adapter/, then la.p_human(model, tok, outline)

Do not load adapter/ with peft.PeftModel. In transformers, Nemotron fuses the Mamba gate and x projections into one in_proj and stores each layer's 128 routed experts as a single 3D tensor, so PEFT has nowhere to attach most of the adapter and skips it without a warning; the model then scores close to the base model. tinker-cookbook's weights.build_hf_model can also merge the adapter into full weights.

Score an outline

Write the outline one item per line, as [Role] content. IdeaLens compares the next-token probabilities of human and ai:

SYSTEM = "Given a role-labelled outline of a document, answer with one word: human if the source document was human-written, ai if it was AI-generated."
SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
HUMAN, AI = 50755, 2464   # token ids of "human" and "ai"

@torch.no_grad()
def p_human(outline):
    ids = tok.encode(f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n{outline}{SUFFIX}",
                     add_special_tokens=False)
    logits = model(torch.tensor([ids], device=model.device)).logits[0, -1].float()
    return torch.softmax(logits[[HUMAN, AI]], -1)[0].item()

outline = ("[Central Development] A town's water supply fails after a drought, and residents organise to share wells.\n"
           "[Background Context] The reservoir has been shrinking for three summers.\n"
           "[Open Question] Whether the council will fund a new pipeline remains undecided.")
print(p_human(outline))

Build the prompt string exactly as above rather than through the chat template.

Score with vLLM

For many inputs, vLLM is about 15 times faster than the code above and fits much longer inputs on one 80 GB GPU. With vLLM 0.21 (install xgrammar==0.2.1; later releases require transformers < 5), reusing SYSTEM, SUFFIX, HUMAN and AI from above:

import math, os
os.environ.setdefault("VLLM_USE_FLASHINFER_SAMPLER", "0")   # FlashInfer kernels compile CUDA code and need nvcc
os.environ.setdefault("VLLM_USE_FLASHINFER_MOE_FP16", "0")
os.environ.setdefault("VLLM_USE_DEEP_GEMM", "0")            # H100 warmup crashes when DeepGEMM is not installed
from transformers import AutoTokenizer
from vllm import LLM, SamplingParams

tok = AutoTokenizer.from_pretrained("rishanthrajendhran/IdeaLens")
llm = LLM(model="rishanthrajendhran/IdeaLens", dtype="bfloat16", max_num_seqs=256, enable_prefix_caching=False,
          max_logprobs=20, enable_flashinfer_autotune=False, seed=0)
sp = SamplingParams(max_tokens=1, temperature=0.0, logprobs=20)

def p_human_batch(outlines):
    prompts = [{"prompt_token_ids": tok.encode(f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n{x}{SUFFIX}",
                                               add_special_tokens=False)} for x in outlines]
    out = []
    for r in llm.generate(prompts, sp, use_tqdm=False):
        lp = r.outputs[0].logprobs[0]           # the top 20 next-token log-probabilities
        out.append(1 / (1 + math.exp(lp[AI].logprob - lp[HUMAN].logprob)))
    return out

# when done: without this, vLLM 0.21 keeps a script running after its last line
llm.llm_engine.engine_core.shutdown()

max_num_seqs=256 keeps every running sequence's Mamba state in memory; vLLM's H100 default (1,024) does not fit beside the weights. If human or ai is missing from the top 20 (rare), score the prompt followed by each label token with SamplingParams(max_tokens=1, prompt_logprobs=0) and read the last prompt log-probability of each. Scores agree with the training-time scores to about 0.001 in P(human) on average; A100 and H100 GPUs differ by as much.

Thresholds

IdeaLens flags a document as having AI ideas when P(human) is below a cut. Each cut is set so that a given share of human documents is flagged (the false-positive rate, FPR), measured on the 80,000 human documents in IdeaLens-1M's calibration split (10,000 per format). The paper's operating point is the global cut at 1% FPR.

FPR 0.1% 0.5% 1% 2% 5% 10% 20%
Global cut 0.01691 0.07279 0.13712 0.34047 0.72713 0.89798 0.97156

Per-format cuts give each format its own operating point. They need the document's format, which the paper assigns with WebOrganizer's annotation prompt run on Gemini 3.7 Flash; the calibration documents use the formats recorded in IdeaLens-1M. Each is the format's own quantile, shrunk toward the global cut with weight n / (n + 2500); at 0.1% FPR 10,000 documents per format are too few, so there is no per-format cut. A document outside these eight formats has no per-format cut; do not fall back to the global cut for it.

Format 0.5% 1% 2% 5% 10% 20%
Nonfiction Writing 0.05358 0.08989 0.17022 0.54060 0.86724 0.96940
Knowledge Article 0.06944 0.10700 0.20819 0.61170 0.88568 0.97190
Personal Blog 0.09460 0.25214 0.50599 0.78596 0.90989 0.97157
News Article 0.04970 0.09170 0.23610 0.57516 0.81288 0.94887
Academic Writing 0.16045 0.37109 0.63082 0.87285 0.95310 0.98617
User Reviews 0.08956 0.24748 0.46654 0.79774 0.93054 0.97995
Personal About Page 0.08280 0.17824 0.39823 0.71519 0.87589 0.95970
Creative Writing 0.20831 0.35315 0.54943 0.79025 0.90777 0.96932

thresholds.json holds every cut at full precision, plus per-topic cuts for the WebOrganizer topics with enough calibration documents.

These rates hold for English web documents like the training data. For another domain, fit the cut on human documents from that domain.

Hardware requirements

Measured with transformers 5.15 in bf16 on NVIDIA H100 80GB GPUs (our other runs used A100 80GB), with transformers' PyTorch implementation of the Mamba layers (no fused Mamba kernels installed). We have not tried CPU-only inference.

Merged model Adapter route (load_adapter.py)
Download 65.8 GB 65.8 GB base model + 3.1 GB adapter
Peak CPU RAM while loading 60 GiB 60 GiB
GPU memory once loaded 58.8 GiB 58.8 GiB (66 GiB during the ~10 s it takes to apply the adapter)

GPU memory then grows with the length of the input, by about 4.2 MiB per token at typical lengths, scoring one input at a time:

Input tokens 500 1,000 2,000 4,000 8,000
Peak GPU memory, one 80 GB GPU 61.0 GiB 63.1 GiB 67.3 GiB 75.7 GiB does not fit
Peak memory per GPU, two 80 GB GPUs (device_map="auto") 47.4 GiB 63.7 GiB
Seconds per input, H100 0.18 0.34 0.66 1.32 2.70

Inputs of 6,000 tokens do not fit on one 80 GB GPU and 12,000 do not fit on two; lowering the Mamba chunk size from 128 to 64 did not change either limit.

IdeaLens reads outlines, which are short. The outlines in IdeaLens-1M's calibration split average about 640 tokens with the prompt, and the longest is under 3,800, so one 80 GB GPU (A100 80GB or H100 80GB) is enough. Outline extraction runs through an LLM API and needs no local GPU.

Intended use and limitations

  • IdeaLens estimates the provenance of a document's ideas. It should not be the sole basis for decisions about a person's work.
  • It was trained on English web documents of at least 500 words in eight long-form formats (Nonfiction Writing, Knowledge Article, Personal Blog, News Article, Academic Writing, User Reviews, Personal About Page, Creative Writing).
  • Its training labels come from the Pangram prose detector, applied to whole documents. They record who wrote the prose; the model learns idea provenance from them only through outlines.
  • Errors in outline extraction carry into the score.

Related models

Model Backbone Reads
IdeaLens (this model) Nemotron-3.5-Lightning-30B-A3B, LoRA outline
ProseLens Nemotron-3.5-Lightning-30B-A3B, LoRA document text
IdeaLens-NoParaphrase Nemotron-3.5-Lightning-30B-A3B, LoRA outline, trained without paraphrasing
IdeaLens-Qwen3.5-9B Qwen3.5-9B, classification head outline
IdeaLens-ModernBERT-L ModernBERT-large outline
ProseLens-ModernBERT-L ModernBERT-large document text
IdeaLens-LogisticClassifier logistic regression over text-embedding-3-large outline
IdeaLens-ModernBERT-L-NoParaphrase ModernBERT-large outline, trained without paraphrasing
IdeaLens-ModernBERT-L-RolesOnly ModernBERT-large role labels only
IdeaLens-Qwen3.5-9B-PerItem Qwen3.5-9B, classification head single outline items, pooled
IdeaLens-ModernBERT-L-PerItem ModernBERT-large single outline items, pooled
IdeaLens-LogisticClassifier-PerItem logistic regression over text-embedding-3-large single outline items, pooled

Training data: IdeaLens-1M.

License

OpenMDW-1.1, the license of the base model (see LICENSE).

Citation

@article{idealens2026,
  title   = {IdeaLens: Detecting AI Ideas in Long-form Writing},
  author  = {Anonymous},
  journal = {arXiv preprint arXiv:TBD},
  year    = {2026},
  url     = {https://arxiv.org/abs/TBD}
}