English
ai-text-detection
idea-provenance
rishanthrajendhran commited on
Commit
85849d9
·
verified ·
1 Parent(s): b269328

model card

Browse files
Files changed (1) hide show
  1. README.md +32 -0
README.md ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ tags: [ai-detection, logistic-regression]
4
+ extra_gated_prompt: "Access is granted individually. Please say who you are and what you intend to use the weights for."
5
+ ---
6
+
7
+ # ideadet-logreg-1m-outline
8
+
9
+ Idea-level AI-text detector: a logistic-regression head over frozen
10
+ `text-embedding-3-large` vectors (3072 dimensions). It answers **whose ideas a
11
+ document contains**, not who typed the sentences.
12
+
13
+ - **Training corpus:** 1m, 842,301 training rows
14
+ - **Reads at inference:** a de-leaked (paraphrased) role-labelled outline, embedded whole
15
+ - **Hyper-parameters:** C=1.0, max_iter=3000
16
+
17
+ ## Contents
18
+
19
+ `lr_full_v1m.npz` holds `coef` ((1, 3072)), `intercept`, and the fitting metadata.
20
+ AI is the positive class, so the reported score is `predict_proba(x)[:, 0]` = P(human), and
21
+ the detector fires when that falls below a calibrated threshold. Thresholds are quantiles
22
+ over held-out human documents and are **not** included here: a cut from one model or input
23
+ form is meaningless against another's scores.
24
+
25
+ ## Provenance
26
+
27
+ This model was refit from the stored embeddings, because the original runs saved only their
28
+ predictions. The refit was checked against those published test predictions and reproduces
29
+ them to **max |difference| = 2.06e-13**, so it is the same
30
+ model that produced the reported numbers rather than an approximation.
31
+
32
+ Fuller documentation and evaluation results to follow.