Text Generation
Transformers
Safetensors
English
nemotron_h
ai-text-detection
prose-provenance
conversational
Instructions to use rishanthrajendhran/ProseLens with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rishanthrajendhran/ProseLens with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="rishanthrajendhran/ProseLens") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("rishanthrajendhran/ProseLens") model = AutoModelForCausalLM.from_pretrained("rishanthrajendhran/ProseLens", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use rishanthrajendhran/ProseLens with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "rishanthrajendhran/ProseLens" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rishanthrajendhran/ProseLens", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/rishanthrajendhran/ProseLens
- SGLang
How to use rishanthrajendhran/ProseLens with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "rishanthrajendhran/ProseLens" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rishanthrajendhran/ProseLens", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "rishanthrajendhran/ProseLens" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rishanthrajendhran/ProseLens", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use rishanthrajendhran/ProseLens with Docker Model Runner:
docker model run hf.co/rishanthrajendhran/ProseLens
|
Download README.md from rishanthrajendhran/ProseLens: direct link, hf CLI and curl.
- Browser
- Download file 13 kB
-
https://huggingface.co/rishanthrajendhran/ProseLens/resolve/main/README.md
- Command line
-
hf download hf://rishanthrajendhran/ProseLens/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/rishanthrajendhran/ProseLens/resolve/main/README.md
13 kB
| license: other | |
| license_name: openmdw-1.1 | |
| license_link: LICENSE | |
| base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | |
| library_name: transformers | |
| language: | |
| - en | |
| tags: | |
| - ai-text-detection | |
| - prose-provenance | |
| datasets: | |
| - rishanthrajendhran/WildOutlines | |
| extra_gated_prompt: "Access is granted individually. Please say who you are and what you intend to use the weights for." | |
| # ProseLens | |
| ProseLens detects **who wrote the words** of a document. It is IdeaLens's counterpart in the paper: the same | |
| backbone, training documents and labels, but it reads the raw document text instead of an outline, so it learns | |
| word-level provenance. It returns P(human), the probability that the document was written by a person. | |
| ProseLens is `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` fine-tuned with LoRA (rank 64) on 1M English web | |
| documents ([WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines)). | |
| ## Results | |
| From the paper: ProseLens is accurate when a document's ideas and words come from the same source (99.1%), but | |
| when they come from different sources it tracks the words and reaches 25.4% on idea provenance. IdeaLens, trained | |
| identically on outlines, reaches 95.3% and 81.3%; Pangram 4 reaches 98.5% and 25.9%. | |
| ## Usage | |
| Pass the document text as is. | |
| ### Quick start with the idealens package | |
| [idealens](https://github.com/RishanthRajendhran/IdeaLens) ([PyPI](https://pypi.org/project/idealens/)) scores documents with ProseLens on vLLM and applies the thresholds in this repo. No | |
| outline and no LLM call is needed: | |
| ```bash | |
| pip install "idealens[vllm]" | |
| idealens score docs.jsonl -o scores.jsonl --model ProseLens | |
| ``` | |
| Input is JSONL with a `text` field per document. In Python: | |
| ```python | |
| import idealens as il | |
| with il.Detector("ProseLens") as det: # vLLM, with this repo's thresholds | |
| records = det.score_documents([open("document.txt").read()]) | |
| print(records[0]["p_human"], records[0]["verdict"]["ai"]) | |
| ``` | |
| To compare with idea-level detection on the same documents, `idealens run docs.jsonl -o ideas.jsonl --model IdeaLens` | |
| extracts their outlines and scores them with [IdeaLens](https://huggingface.co/rishanthrajendhran/IdeaLens). The rest | |
| of this section runs the model directly. | |
| ### Load the merged model (66 GB download) | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| tok = AutoTokenizer.from_pretrained("rishanthrajendhran/ProseLens") | |
| model = AutoModelForCausalLM.from_pretrained("rishanthrajendhran/ProseLens", dtype=torch.bfloat16, device_map="auto").eval() | |
| ``` | |
| The weights take 59 GiB of GPU memory, and each input adds more; see *Hardware requirements*. | |
| ### Or apply the adapter to the base model (3 GB download) | |
| `adapter/` holds the LoRA adapter as trained, in the layout of the Tinker training service. If you already have the | |
| base model, `load_adapter.py` merges the adapter into it in memory. The resulting weights are bit-identical to the | |
| merged model's: | |
| ```python | |
| import importlib.util | |
| from huggingface_hub import hf_hub_download | |
| path = hf_hub_download("rishanthrajendhran/ProseLens", "load_adapter.py") | |
| spec = importlib.util.spec_from_file_location("load_adapter", path) | |
| la = importlib.util.module_from_spec(spec); spec.loader.exec_module(la) | |
| model, tok = la.load_model() # base model + adapter/, then la.p_human(model, tok, text) | |
| ``` | |
| Do not load `adapter/` with `peft.PeftModel`. In transformers, Nemotron fuses the Mamba gate and x projections into | |
| one `in_proj` and stores each layer's 128 routed experts as a single 3D tensor, so PEFT has nowhere to attach most of | |
| the adapter and skips it without a warning; the model then scores close to the base model. `tinker-cookbook`'s | |
| `weights.build_hf_model` can also merge the adapter into full weights. | |
| ### Score a document | |
| ProseLens compares the next-token probabilities of `human` and `ai`: | |
| ```python | |
| SYSTEM = "Given a document, answer with one word: human if the document was human-written, ai if it was AI-generated." | |
| SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n" | |
| HUMAN, AI = 50755, 2464 # token ids of "human" and "ai" | |
| @torch.no_grad() | |
| def p_human(text): | |
| ids = tok.encode(f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n{text}{SUFFIX}", | |
| add_special_tokens=False) | |
| logits = model(torch.tensor([ids], device=model.device)).logits[0, -1].float() | |
| return torch.softmax(logits[[HUMAN, AI]], -1)[0].item() | |
| text = open("document.txt").read() | |
| print(p_human(text)) | |
| ``` | |
| Build the prompt string exactly as above rather than through the chat template. | |
| ### Score with vLLM | |
| For many inputs, vLLM is about 15 times faster than the code above and fits much longer inputs on one 80 GB GPU. | |
| With vLLM 0.21 (install `xgrammar==0.2.1`; later releases require transformers < 5), reusing `SYSTEM`, `SUFFIX`, | |
| `HUMAN` and `AI` from above: | |
| ```python | |
| import math, os | |
| os.environ.setdefault("VLLM_USE_FLASHINFER_SAMPLER", "0") # FlashInfer kernels compile CUDA code and need nvcc | |
| os.environ.setdefault("VLLM_USE_FLASHINFER_MOE_FP16", "0") | |
| os.environ.setdefault("VLLM_USE_DEEP_GEMM", "0") # H100 warmup crashes when DeepGEMM is not installed | |
| from transformers import AutoTokenizer | |
| from vllm import LLM, SamplingParams | |
| tok = AutoTokenizer.from_pretrained("rishanthrajendhran/ProseLens") | |
| llm = LLM(model="rishanthrajendhran/ProseLens", dtype="bfloat16", max_num_seqs=256, enable_prefix_caching=False, | |
| max_logprobs=20, enable_flashinfer_autotune=False, seed=0) | |
| sp = SamplingParams(max_tokens=1, temperature=0.0, logprobs=20) | |
| def p_human_batch(texts): | |
| prompts = [{"prompt_token_ids": tok.encode(f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n{x}{SUFFIX}", | |
| add_special_tokens=False)} for x in texts] | |
| out = [] | |
| for r in llm.generate(prompts, sp, use_tqdm=False): | |
| lp = r.outputs[0].logprobs[0] # the top 20 next-token log-probabilities | |
| out.append(1 / (1 + math.exp(lp[AI].logprob - lp[HUMAN].logprob))) | |
| return out | |
| # when done: without this, vLLM 0.21 keeps a script running after its last line | |
| llm.llm_engine.engine_core.shutdown() | |
| ``` | |
| `max_num_seqs=256` keeps every running sequence's Mamba state in memory; vLLM's H100 default (1,024) does not fit | |
| beside the weights. If `human` or `ai` is missing from the top 20 (rare), score the prompt followed by each label | |
| token with `SamplingParams(max_tokens=1, prompt_logprobs=0)` and read the last prompt log-probability of each. | |
| Scores agree with the training-time scores to about 0.001 in P(human) on average; A100 and H100 GPUs differ by as much. | |
| ### Thresholds | |
| ProseLens flags a document as AI-written when P(human) is below a cut. Each cut is set so | |
| that a given share of human documents is flagged (the false-positive rate, FPR), measured on the 80,000 human | |
| documents in WildOutlines's `calibration` split (10,000 per format). The paper's operating point is the global | |
| cut at 1% FPR. | |
| | FPR | 0.1% | 0.5% | 1% | 2% | 5% | | |
| |---|---:|---:|---:|---:|---:| | |
| | Global cut | 0.07082 | 0.30469 | 0.60539 | 0.92502 | 0.99889 | | |
| Per-format cuts give each format its own operating point. They need the document's format, which the paper | |
| assigns with WebOrganizer's annotation prompt run on Gemini 3.7 Flash; the calibration documents use the formats | |
| recorded in WildOutlines. Each is the | |
| format's own quantile, shrunk toward the global cut with weight n / (n + 2500); at 0.1% FPR 10,000 documents | |
| per format are too few, so there is no per-format cut. A document outside these eight formats has no | |
| per-format cut; do not fall back to the global cut for it. | |
| | Format | 0.5% | 1% | 2% | 5% | | |
| |---|---:|---:|---:|---:| | |
| | Nonfiction Writing | 0.14387 | 0.27050 | 0.49499 | 0.93935 | | |
| | Knowledge Article | 0.24236 | 0.38050 | 0.64044 | 0.95295 | | |
| | Personal Blog | 0.68275 | 0.87993 | 0.98163 | 0.99966 | | |
| | News Article | 0.29554 | 0.55535 | 0.90380 | 0.99788 | | |
| | Academic Writing | 0.77788 | 0.89967 | 0.98279 | 0.99969 | | |
| | User Reviews | 0.62135 | 0.85161 | 0.98035 | 0.99964 | | |
| | Personal About Page | 0.41699 | 0.78714 | 0.97730 | 0.99966 | | |
| | Creative Writing | 0.84923 | 0.91749 | 0.98411 | 0.99958 | | |
| `thresholds.json` holds every cut at full precision. | |
| These rates hold for English web documents like the training data. For another domain, fit the cut on | |
| human documents from that domain. | |
| ProseLens scores 97% of human calibration documents above 0.99, so its cuts at higher FPRs sit close | |
| to 1 and small shifts in score move the realised FPR a long way. | |
| ## Hardware requirements | |
| Measured with transformers 5.15 in bf16 on NVIDIA H100 80GB GPUs (our other runs used A100 80GB), with transformers' | |
| PyTorch implementation of the Mamba layers (no fused Mamba kernels installed). We have not tried CPU-only inference. | |
| | | Merged model | Adapter route (`load_adapter.py`) | | |
| |---|---|---| | |
| | Download | 65.8 GB | 65.8 GB base model + 3.1 GB adapter | | |
| | Peak CPU RAM while loading | 60 GiB | 60 GiB | | |
| | GPU memory once loaded | 58.8 GiB | 58.8 GiB (66 GiB during the ~10 s it takes to apply the adapter) | | |
| GPU memory then grows with the length of the input, by about 4.2 MiB per token at typical lengths, scoring one | |
| input at a time: | |
| | Input tokens | 500 | 1,000 | 2,000 | 4,000 | 8,000 | | |
| |---|---:|---:|---:|---:|---:| | |
| | Peak GPU memory, one 80 GB GPU | 61.0 GiB | 63.1 GiB | 67.3 GiB | 75.7 GiB | does not fit | | |
| | Peak memory per GPU, two 80 GB GPUs (`device_map="auto"`) | | | | 47.4 GiB | 63.7 GiB | | |
| | Seconds per input, H100 | 0.18 | 0.34 | 0.66 | 1.32 | 2.70 | | |
| Inputs of 6,000 tokens do not fit on one 80 GB GPU and 12,000 do not fit on two; lowering the Mamba chunk size from | |
| 128 to 64 did not change either limit. | |
| ProseLens reads whole documents, so their length decides the hardware. One 80 GB GPU handles documents up to about | |
| 4,000 tokens (about 3,000 words); 85% of WildOutlines's test documents are that short. Two 80 GB GPUs handle up to | |
| 8,000 tokens (about 6,000 words), which covers all but about 1 in 1,000 test documents. We have not measured longer | |
| inputs or more GPUs. | |
| ## Intended use and limitations | |
| - ProseLens estimates who wrote a document's words. It should not be the sole basis for decisions about a person's work. | |
| - It was trained on English web documents of at least 500 words in eight long-form formats (Nonfiction Writing, Knowledge Article, Personal Blog, News Article, Academic Writing, User Reviews, Personal About Page, Creative Writing). | |
| - Its training labels come from the Pangram prose detector, applied to whole documents. | |
| ## Related models | |
| | Model | Backbone | Reads | | |
| |---|---|---| | |
| | [IdeaLens](https://huggingface.co/rishanthrajendhran/IdeaLens) | Nemotron-3.5-Lightning-30B-A3B, LoRA | outline | | |
| | [ProseLens](https://huggingface.co/rishanthrajendhran/ProseLens) (this model) | Nemotron-3.5-Lightning-30B-A3B, LoRA | document text | | |
| | [IdeaLens-NoParaphrase](https://huggingface.co/rishanthrajendhran/IdeaLens-NoParaphrase) | Nemotron-3.5-Lightning-30B-A3B, LoRA | outline, trained without paraphrasing | | |
| | [IdeaLens-Qwen3.5-9B](https://huggingface.co/rishanthrajendhran/IdeaLens-Qwen3.5-9B) | Qwen3.5-9B, classification head | outline | | |
| | [IdeaLens-ModernBERT-L](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L) | ModernBERT-large | outline | | |
| | [ProseLens-ModernBERT-L](https://huggingface.co/rishanthrajendhran/ProseLens-ModernBERT-L) | ModernBERT-large | document text | | |
| | [IdeaLens-LogisticClassifier](https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier) | logistic regression over text-embedding-3-large | outline | | |
| | [IdeaLens-ModernBERT-L-NoParaphrase](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-NoParaphrase) | ModernBERT-large | outline, trained without paraphrasing | | |
| | [IdeaLens-ModernBERT-L-RolesOnly](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-RolesOnly) | ModernBERT-large | role labels only | | |
| | [IdeaLens-Qwen3.5-9B-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-Qwen3.5-9B-PerItem) | Qwen3.5-9B, classification head | single outline items, pooled | | |
| | [IdeaLens-ModernBERT-L-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-PerItem) | ModernBERT-large | single outline items, pooled | | |
| | [IdeaLens-LogisticClassifier-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier-PerItem) | logistic regression over text-embedding-3-large | single outline items, pooled | | |
| Training data: [WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines). | |
| All IdeaLens models and datasets are in the [IdeaLens collection](https://huggingface.co/collections/rishanthrajendhran/idealens-6abee785ce6196fc0be9200f). | |
| ## License | |
| OpenMDW-1.1, the license of the base model (see `LICENSE`). | |
| ## Citation | |
| ```bibtex | |
| @article{idealens2026, | |
| title = {IdeaLens: Detecting AI Ideas in Long-form Writing}, | |
| author = {Anonymous}, | |
| journal = {arXiv preprint arXiv:TBD}, | |
| year = {2026}, | |
| url = {https://arxiv.org/abs/TBD} | |
| } | |
| ``` | |