Text Generation
Transformers
Safetensors
English
nemotron_h
ai-text-detection
idea-provenance
conversational
Instructions to use rishanthrajendhran/IdeaLens with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rishanthrajendhran/IdeaLens with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="rishanthrajendhran/IdeaLens") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("rishanthrajendhran/IdeaLens") model = AutoModelForCausalLM.from_pretrained("rishanthrajendhran/IdeaLens", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use rishanthrajendhran/IdeaLens with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "rishanthrajendhran/IdeaLens" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rishanthrajendhran/IdeaLens", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/rishanthrajendhran/IdeaLens
- SGLang
How to use rishanthrajendhran/IdeaLens with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "rishanthrajendhran/IdeaLens" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rishanthrajendhran/IdeaLens", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "rishanthrajendhran/IdeaLens" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rishanthrajendhran/IdeaLens", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use rishanthrajendhran/IdeaLens with Docker Model Runner:
docker model run hf.co/rishanthrajendhran/IdeaLens
|
Download README.md from rishanthrajendhran/IdeaLens: direct link, hf CLI and curl.
- Browser
- Download file 15.8 kB
-
https://huggingface.co/rishanthrajendhran/IdeaLens/resolve/main/README.md
- Command line
-
hf download hf://rishanthrajendhran/IdeaLens/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/rishanthrajendhran/IdeaLens/resolve/main/README.md
15.8 kB
| license: other | |
| license_name: openmdw-1.1 | |
| license_link: LICENSE | |
| base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | |
| library_name: transformers | |
| language: | |
| - en | |
| tags: | |
| - ai-text-detection | |
| - idea-provenance | |
| datasets: | |
| - rishanthrajendhran/WildOutlines | |
| extra_gated_prompt: "Access is granted individually. Please say who you are and what you intend to use the weights for." | |
| # IdeaLens | |
| IdeaLens detects **who came up with the ideas** in a document, rather than who wrote its words. It reads a | |
| role-labelled outline of the document (an ordered list of items, each giving one idea and the discourse role it plays, | |
| such as *Central Development* or *Open Question*) and returns P(human), the probability that the ideas are human. | |
| IdeaLens is `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` fine-tuned with LoRA (rank 64) on the outlines of 1M | |
| English web documents ([WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines)). The training | |
| outlines were paraphrased to remove the documents' wording, so the model has to fit its labels through the ideas. | |
| ## Results | |
| From the paper: | |
| - IdeaLens is accurate both when a document's ideas and words come from the same source (95.3%) and when they come | |
| from different sources (81.3%). ProseLens, trained identically on the raw text, reaches 99.1% and 25.4%; Pangram 4 | |
| reaches 98.5% and 25.9%. | |
| - As models write from increasingly detailed human plans, IdeaLens's AI flag rate falls from 95% to 7%, while | |
| Pangram 4 still flags 92%. From AI-derived plans, IdeaLens stays above 96%. | |
| - On TwiceTold, 50 stories that human authors wrote from AI-generated plans, IdeaLens flags 68% as AI, against 8% for | |
| Pangram 4. | |
| - On 19 existing detection benchmarks, IdeaLens keeps strong detection rates at low false-positive rates across | |
| domains, formats and languages. | |
| ## Usage | |
| Scoring a document takes two steps: | |
| 1. **Extract an outline.** An LLM writes the outline from the document, its format's role vocabulary and six worked | |
| examples (the paper uses Gemini 3.7 Flash). The [idealens](https://github.com/RishanthRajendhran/IdeaLens) package ships the prompts, role vocabularies and | |
| worked examples, and runs this step. Score the outline as extracted; the paraphrasing step is only for training data. | |
| 2. **Score the outline** with this model, as below. | |
| ### Quick start: everything with the idealens package | |
| [idealens](https://github.com/RishanthRajendhran/IdeaLens) ([PyPI](https://pypi.org/project/idealens/)) classifies each document's format, extracts its outline with the prompt, role | |
| vocabulary and worked examples IdeaLens was trained with, scores the outline on vLLM and applies the thresholds in this | |
| repo: | |
| ```bash | |
| pip install "idealens[vllm]" | |
| export GEMINI_API_KEY=... # or --provider vertex | openai | anthropic | openrouter | compatible | |
| idealens run docs.jsonl -o scores.jsonl --model IdeaLens --dry-run # price the LLM calls first | |
| idealens run docs.jsonl -o scores.jsonl --model IdeaLens | |
| ``` | |
| Input is JSONL with a `text` field per document. The same steps in Python: | |
| ```python | |
| import idealens as il | |
| texts = [open("document.txt").read()] | |
| formats = il.classify(texts) # one of the eight formats per document | |
| outlines = il.extract(texts, formats) # role-labelled outlines | |
| with il.Detector("IdeaLens") as det: # vLLM, with this repo's thresholds | |
| records = det.score_outlines(outlines, format=formats) | |
| r = records[0] | |
| print(r["p_human"], r["verdict"]["ai"], r["verdict"]["cut"]) # P(human); flagged at the 1% global cut? | |
| print(outlines[0].render()) # the outline that was scored | |
| ``` | |
| `classify` and `extract` call Gemini 3.7 Flash, the extractor the thresholds were fitted with; pass | |
| `provider=idealens.providers.make("openai", "gpt-6-sol")` (or Vertex, Anthropic, OpenRouter, a local server) to use | |
| another. Each record also carries verdicts at every calibrated false-positive rate under the global, per-format and | |
| per-topic schemes. The [package README](https://github.com/RishanthRajendhran/IdeaLens#ways-to-use-idealens) covers batch jobs, scoring outlines you already have and calibrating | |
| on your own data. | |
| ### Mix and match: extract with the package, score with your own code | |
| ```python | |
| import idealens as il | |
| text = open("document.txt").read() | |
| outline = il.extract([text], il.classify([text]))[0].render() # one "[Role] content" line per item | |
| print(p_human(outline)) # p_human (transformers) or p_human_batch (vLLM), defined below | |
| ``` | |
| The rest of this section runs the model directly. | |
| ### Load the merged model (66 GB download) | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| tok = AutoTokenizer.from_pretrained("rishanthrajendhran/IdeaLens") | |
| model = AutoModelForCausalLM.from_pretrained("rishanthrajendhran/IdeaLens", dtype=torch.bfloat16, device_map="auto").eval() | |
| ``` | |
| The weights take 59 GiB of GPU memory, and each input adds more; see *Hardware requirements*. | |
| ### Or apply the adapter to the base model (3 GB download) | |
| `adapter/` holds the LoRA adapter as trained, in the layout of the Tinker training service. If you already have the | |
| base model, `load_adapter.py` merges the adapter into it in memory. The resulting weights are bit-identical to the | |
| merged model's: | |
| ```python | |
| import importlib.util | |
| from huggingface_hub import hf_hub_download | |
| path = hf_hub_download("rishanthrajendhran/IdeaLens", "load_adapter.py") | |
| spec = importlib.util.spec_from_file_location("load_adapter", path) | |
| la = importlib.util.module_from_spec(spec); spec.loader.exec_module(la) | |
| model, tok = la.load_model() # base model + adapter/, then la.p_human(model, tok, outline) | |
| ``` | |
| Do not load `adapter/` with `peft.PeftModel`. In transformers, Nemotron fuses the Mamba gate and x projections into | |
| one `in_proj` and stores each layer's 128 routed experts as a single 3D tensor, so PEFT has nowhere to attach most of | |
| the adapter and skips it without a warning; the model then scores close to the base model. `tinker-cookbook`'s | |
| `weights.build_hf_model` can also merge the adapter into full weights. | |
| ### Score an outline | |
| Write the outline one item per line, as `[Role] content`, or take it from `il.extract` as above. IdeaLens compares the next-token probabilities of `human` and `ai`: | |
| ```python | |
| SYSTEM = "Given a role-labelled outline of a document, answer with one word: human if the source document was human-written, ai if it was AI-generated." | |
| SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n" | |
| HUMAN, AI = 50755, 2464 # token ids of "human" and "ai" | |
| @torch.no_grad() | |
| def p_human(outline): | |
| ids = tok.encode(f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n{outline}{SUFFIX}", | |
| add_special_tokens=False) | |
| logits = model(torch.tensor([ids], device=model.device)).logits[0, -1].float() | |
| return torch.softmax(logits[[HUMAN, AI]], -1)[0].item() | |
| outline = ("[Central Development] A town's water supply fails after a drought, and residents organise to share wells.\n" | |
| "[Background Context] The reservoir has been shrinking for three summers.\n" | |
| "[Open Question] Whether the council will fund a new pipeline remains undecided.") | |
| print(p_human(outline)) | |
| ``` | |
| Build the prompt string exactly as above rather than through the chat template. | |
| ### Score with vLLM | |
| For many inputs, vLLM is about 15 times faster than the code above and fits much longer inputs on one 80 GB GPU. | |
| With vLLM 0.21 (install `xgrammar==0.2.1`; later releases require transformers < 5), reusing `SYSTEM`, `SUFFIX`, | |
| `HUMAN` and `AI` from above: | |
| ```python | |
| import math, os | |
| os.environ.setdefault("VLLM_USE_FLASHINFER_SAMPLER", "0") # FlashInfer kernels compile CUDA code and need nvcc | |
| os.environ.setdefault("VLLM_USE_FLASHINFER_MOE_FP16", "0") | |
| os.environ.setdefault("VLLM_USE_DEEP_GEMM", "0") # H100 warmup crashes when DeepGEMM is not installed | |
| from transformers import AutoTokenizer | |
| from vllm import LLM, SamplingParams | |
| tok = AutoTokenizer.from_pretrained("rishanthrajendhran/IdeaLens") | |
| llm = LLM(model="rishanthrajendhran/IdeaLens", dtype="bfloat16", max_num_seqs=256, enable_prefix_caching=False, | |
| max_logprobs=20, enable_flashinfer_autotune=False, seed=0) | |
| sp = SamplingParams(max_tokens=1, temperature=0.0, logprobs=20) | |
| def p_human_batch(outlines): | |
| prompts = [{"prompt_token_ids": tok.encode(f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n{x}{SUFFIX}", | |
| add_special_tokens=False)} for x in outlines] | |
| out = [] | |
| for r in llm.generate(prompts, sp, use_tqdm=False): | |
| lp = r.outputs[0].logprobs[0] # the top 20 next-token log-probabilities | |
| out.append(1 / (1 + math.exp(lp[AI].logprob - lp[HUMAN].logprob))) | |
| return out | |
| # when done: without this, vLLM 0.21 keeps a script running after its last line | |
| llm.llm_engine.engine_core.shutdown() | |
| ``` | |
| `max_num_seqs=256` keeps every running sequence's Mamba state in memory; vLLM's H100 default (1,024) does not fit | |
| beside the weights. If `human` or `ai` is missing from the top 20 (rare), score the prompt followed by each label | |
| token with `SamplingParams(max_tokens=1, prompt_logprobs=0)` and read the last prompt log-probability of each. | |
| Scores agree with the training-time scores to about 0.001 in P(human) on average; A100 and H100 GPUs differ by as much. | |
| ### Thresholds | |
| IdeaLens flags a document as having AI ideas when P(human) is below a cut. Each cut is set so | |
| that a given share of human documents is flagged (the false-positive rate, FPR), measured on the 80,000 human | |
| documents in WildOutlines's `calibration` split (10,000 per format). The paper's operating point is the global | |
| cut at 1% FPR. | |
| | FPR | 0.1% | 0.5% | 1% | 2% | 5% | 10% | 20% | | |
| |---|---:|---:|---:|---:|---:|---:|---:| | |
| | Global cut | 0.01691 | 0.07279 | 0.13712 | 0.34047 | 0.72713 | 0.89798 | 0.97156 | | |
| Per-format cuts give each format its own operating point. They need the document's format, which the paper | |
| assigns with WebOrganizer's annotation prompt run on Gemini 3.7 Flash; the calibration documents use the formats | |
| recorded in WildOutlines. Each is the | |
| format's own quantile, shrunk toward the global cut with weight n / (n + 2500); at 0.1% FPR 10,000 documents | |
| per format are too few, so there is no per-format cut. A document outside these eight formats has no | |
| per-format cut; do not fall back to the global cut for it. | |
| | Format | 0.5% | 1% | 2% | 5% | 10% | 20% | | |
| |---|---:|---:|---:|---:|---:|---:| | |
| | Nonfiction Writing | 0.05358 | 0.08989 | 0.17022 | 0.54060 | 0.86724 | 0.96940 | | |
| | Knowledge Article | 0.06944 | 0.10700 | 0.20819 | 0.61170 | 0.88568 | 0.97190 | | |
| | Personal Blog | 0.09460 | 0.25214 | 0.50599 | 0.78596 | 0.90989 | 0.97157 | | |
| | News Article | 0.04970 | 0.09170 | 0.23610 | 0.57516 | 0.81288 | 0.94887 | | |
| | Academic Writing | 0.16045 | 0.37109 | 0.63082 | 0.87285 | 0.95310 | 0.98617 | | |
| | User Reviews | 0.08956 | 0.24748 | 0.46654 | 0.79774 | 0.93054 | 0.97995 | | |
| | Personal About Page | 0.08280 | 0.17824 | 0.39823 | 0.71519 | 0.87589 | 0.95970 | | |
| | Creative Writing | 0.20831 | 0.35315 | 0.54943 | 0.79025 | 0.90777 | 0.96932 | | |
| `thresholds.json` holds every cut at full precision, plus per-topic cuts for the WebOrganizer topics with enough calibration documents. | |
| These rates hold for English web documents like the training data. For another domain, fit the cut on | |
| human documents from that domain. | |
| ## Hardware requirements | |
| Measured with transformers 5.15 in bf16 on NVIDIA H100 80GB GPUs (our other runs used A100 80GB), with transformers' | |
| PyTorch implementation of the Mamba layers (no fused Mamba kernels installed). We have not tried CPU-only inference. | |
| | | Merged model | Adapter route (`load_adapter.py`) | | |
| |---|---|---| | |
| | Download | 65.8 GB | 65.8 GB base model + 3.1 GB adapter | | |
| | Peak CPU RAM while loading | 60 GiB | 60 GiB | | |
| | GPU memory once loaded | 58.8 GiB | 58.8 GiB (66 GiB during the ~10 s it takes to apply the adapter) | | |
| GPU memory then grows with the length of the input, by about 4.2 MiB per token at typical lengths, scoring one | |
| input at a time: | |
| | Input tokens | 500 | 1,000 | 2,000 | 4,000 | 8,000 | | |
| |---|---:|---:|---:|---:|---:| | |
| | Peak GPU memory, one 80 GB GPU | 61.0 GiB | 63.1 GiB | 67.3 GiB | 75.7 GiB | does not fit | | |
| | Peak memory per GPU, two 80 GB GPUs (`device_map="auto"`) | | | | 47.4 GiB | 63.7 GiB | | |
| | Seconds per input, H100 | 0.18 | 0.34 | 0.66 | 1.32 | 2.70 | | |
| Inputs of 6,000 tokens do not fit on one 80 GB GPU and 12,000 do not fit on two; lowering the Mamba chunk size from | |
| 128 to 64 did not change either limit. | |
| IdeaLens reads outlines, which are short. The outlines in WildOutlines's calibration split average about 640 | |
| tokens with the prompt, and the longest is under 3,800, so one 80 GB GPU (A100 80GB or H100 80GB) is enough. Outline | |
| extraction runs through an LLM API and needs no local GPU. | |
| ## Intended use and limitations | |
| - IdeaLens estimates the provenance of a document's ideas. It should not be the sole basis for decisions about a person's work. | |
| - It was trained on English web documents of at least 500 words in eight long-form formats (Nonfiction Writing, Knowledge Article, Personal Blog, News Article, Academic Writing, User Reviews, Personal About Page, Creative Writing). | |
| - Its training labels come from the Pangram prose detector, applied to whole documents. They record who wrote the prose; the model learns idea provenance from them only through outlines. | |
| - Errors in outline extraction carry into the score. | |
| ## Related models | |
| | Model | Backbone | Reads | | |
| |---|---|---| | |
| | [IdeaLens](https://huggingface.co/rishanthrajendhran/IdeaLens) (this model) | Nemotron-3.5-Lightning-30B-A3B, LoRA | outline | | |
| | [ProseLens](https://huggingface.co/rishanthrajendhran/ProseLens) | Nemotron-3.5-Lightning-30B-A3B, LoRA | document text | | |
| | [IdeaLens-NoParaphrase](https://huggingface.co/rishanthrajendhran/IdeaLens-NoParaphrase) | Nemotron-3.5-Lightning-30B-A3B, LoRA | outline, trained without paraphrasing | | |
| | [IdeaLens-Qwen3.5-9B](https://huggingface.co/rishanthrajendhran/IdeaLens-Qwen3.5-9B) | Qwen3.5-9B, classification head | outline | | |
| | [IdeaLens-ModernBERT-L](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L) | ModernBERT-large | outline | | |
| | [ProseLens-ModernBERT-L](https://huggingface.co/rishanthrajendhran/ProseLens-ModernBERT-L) | ModernBERT-large | document text | | |
| | [IdeaLens-LogisticClassifier](https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier) | logistic regression over text-embedding-3-large | outline | | |
| | [IdeaLens-ModernBERT-L-NoParaphrase](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-NoParaphrase) | ModernBERT-large | outline, trained without paraphrasing | | |
| | [IdeaLens-ModernBERT-L-RolesOnly](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-RolesOnly) | ModernBERT-large | role labels only | | |
| | [IdeaLens-Qwen3.5-9B-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-Qwen3.5-9B-PerItem) | Qwen3.5-9B, classification head | single outline items, pooled | | |
| | [IdeaLens-ModernBERT-L-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-PerItem) | ModernBERT-large | single outline items, pooled | | |
| | [IdeaLens-LogisticClassifier-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier-PerItem) | logistic regression over text-embedding-3-large | single outline items, pooled | | |
| Training data: [WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines). | |
| ## License | |
| OpenMDW-1.1, the license of the base model (see `LICENSE`). | |
| ## Citation | |
| ```bibtex | |
| @article{idealens2026, | |
| title = {IdeaLens: Detecting AI Ideas in Long-form Writing}, | |
| author = {Anonymous}, | |
| journal = {arXiv preprint arXiv:TBD}, | |
| year = {2026}, | |
| url = {https://arxiv.org/abs/TBD} | |
| } | |
| ``` | |