Exhibit E · held-out re-evaluation (2026)
The same recipes, scored on texts they never saw.
The 2023 validation scores were inflated: the notebooks re-balanced each domain before splitting it, so copies of the same machine text sat on both sides of the split (the revival audit has the details). The original models had seen almost every domain-1 machine text, so they cannot be re-scored cleanly. This page re-fits the 2023 recipes, unchanged, on a split made before any re-balancing and scores the untouched 20%.
Every number comes with a 95% interval and its sample size. The 2023 numbers stay on the journal exactly as printed.
Design
Split first, then re-balance
The pipeline is the notebooks' own, in a different order. Everything below was produced by scripts/evaluate_heldout.py with the 2023 library versions (scikit-learn 1.2.2, imbalanced-learn 0.10.1).
- 01Split the original rows80/20 within each domain, stratified by class and generator, seed 90051. The test part is never re-balanced or looked at during fitting.
- 02Re-balance the training partThe notebooks' call: SMOTE on the integer row position, which copies minority rows. Domain 1 is capped at 40,000 rows, the notebooks' 80% of 50,000.
- 03Fit the 2023 modelsLogisticRegressionCV (l2, lbfgs, 50 iterations, balanced weights) and SGDClassifier (hinge, l1, alpha 0.0001, balanced weights). Cross-validation chose C = 10,000 for domain 1, the top of its grid, and 2.78 for domain 2. The 50 iterations stop the fits early (see Fitting).
- 04Score with intervals10,000 bootstrap resamples of the test part (seed 90051). Every statistic on one test set uses the same resamples, so differences between detectors are paired.
| Domain | Train human | Train machine | Test human | Test machine |
|---|---|---|---|---|
| 1 | 98,067 | 2,800 | 24,517 | 700 |
| 2 | 80 | 320 | 20 | 80 |
Why accuracy alone misleads here
97.2% of the domain-1 test texts are human, so a rule that always says “human” scores 0.972 accuracy, more than any detector here. The page therefore leads with macro-F1, which averages the F1 of the two classes, and ROC AUC, which asks how often a random human text is ranked above a random machine text.
Results
Machine recall falls on unseen texts, and length explains most of the ranking
On domain 1, logistic regression reaches macro-F1 0.697 (0.685 to 0.709) and AUC 0.960 (0.954 to 0.965). Compared like for like with 2023, the change sits in the copied class: the leaky split had it catching 0.998 of the machine texts, and on texts it never saw it catches 0.789 (0.757 to 0.817) (Wilson 95%), while human recall stays about the same (0.933 then, 0.945 now; see the 2023 numbers). The printed 2023 F1 of 0.9643 is the human-class F1 on a 50:50 validation set, so it is not comparable with macro-F1 on the real 97:3 mix.
A logistic regression on the text length alone reaches AUC 0.932 (0.923 to 0.941): machine texts in domain 1 are about half as long as human ones. The bag-of-words detector adds +0.027 (+0.019 to +0.036) AUC over length on the same texts. Domain 2 has only 100 test texts, 20 of them human, and the intervals show it. Under each detector is its mean ± standard deviation over ten re-drawn splits. The main domain-2 split ranks 2 of 10 on logistic regression's macro-F1 (1 is the best), so its headline numbers sit on the favourable side.
Domain 1 test part
- Logistic regression10 splits: 0.696 ± 0.0040.6970.685 to 0.709
- SGD (hinge)10 splits: 0.614 ± 0.0050.6140.603 to 0.625
- Length only0.5640.556 to 0.571
- Always the majority0.4930.492 to 0.493
- Logistic regression10 splits: 0.958 ± 0.0030.9600.954 to 0.965
- SGD (hinge)10 splits: 0.903 ± 0.0060.9020.893 to 0.911
- Length only0.9320.923 to 0.941
chance (0.500)
- Logistic regression10 splits: 0.941 ± 0.0010.9410.938 to 0.944
- SGD (hinge)10 splits: 0.910 ± 0.0030.9110.908 to 0.915
- Length only0.8270.823 to 0.832
- Always the majority0.9720.970 to 0.974
ROC curves
Recall per class
- Logistic regression
- Human texts kept as human 0.945 (0.942 to 0.948)23,174 of 24,517Machine texts caught 0.789 (0.757 to 0.817)552 of 700
- SGD (hinge)
- Human texts kept as human 0.920 (0.916 to 0.923)22,547 of 24,517Machine texts caught 0.609 (0.572 to 0.644)426 of 700
Fitting sensitivity
The 2023 fits stop early and choose C on leaky folds
The recipe gives LogisticRegressionCV 50 lbfgs iterations. In 36 of the 66 fits on this page that use that budget, every cross-validation fit stopped at the limit, and in 38 of them the chosen C was the largest value in the grid. For the domain-1 recipe that is all 50 cross-validation fits and the final fit (51 convergence warnings), with C = 10,000. The 2023 notebooks silenced the warnings; the script now records them for every fit. The inner folds also hold copies of the same text on both sides, because the copies are made before cross-validation splits the training part, so inner-fold accuracy keeps rising with C. SGD on domain 1 also ran all 1,000 of its epochs without meeting its stopping rule.
The table re-fits the recipe with each problem removed in turn, on the same training and test texts. The headline numbers above stay the recipe as run, because that is what the team built. This section shows how much they depend on the fitting.
| Fit | Iterations | C | Macro-F1 | AUC | Machine texts caught | ECE of P(human) |
|---|---|---|---|---|---|---|
| As run (the 2023 settings)inner folds: stratified | limit 5050/50 CV fits at limit | 10,000top of grid | 0.697 (0.685 to 0.709) | 0.960 (0.954 to 0.965) | 0.789 (0.757 to 0.817)552 of 700 | 0.054 (0.051 to 0.056) |
| C chosen on folds grouped by source rowinner folds: grouped by source row | limit 5050/50 CV fits at limit | 0.006 | 0.646 (0.636 to 0.656) | 0.971 (0.967 to 0.975) | 0.906 (0.882 to 0.925)634 of 700 | 0.112 (0.109 to 0.115) |
| Solver run to convergenceinner folds: stratified | limit 5,0000/50 CV fits at limit | 1,292 | 0.687 (0.675 to 0.700) | 0.950 (0.943 to 0.957) | 0.733 (0.699 to 0.764)513 of 700 | 0.057 (0.054 to 0.060) |
| Both: grouped folds and convergenceinner folds: grouped by source row | limit 5,0000/50 CV fits at limit | 0.006 | 0.644 (0.634 to 0.654) | 0.971 (0.966 to 0.975) | 0.907 (0.883 to 0.926)635 of 700 | 0.112 (0.109 to 0.115) |
| The 2023 C, run to convergenceinner folds: none (C fixed at the 2023 choice) | limit 20,0003,101 used | 10,000 | 0.688 (0.675 to 0.700) | 0.950 (0.943 to 0.957) | 0.731 (0.697 to 0.763)512 of 700 | 0.058 (0.055 to 0.060) |
Choosing C
What changes when the fit is done properly
With the solver run to convergence and C chosen on folds that keep every copy of a text together, cross-validation picks C = 0.006 for domain 1 instead of 10,000. On the same test texts the recipe as run minus that fit is +0.053 (+0.045 to +0.061) in macro-F1 and −0.011 (−0.015 to −0.008) in AUC. The more strongly regularised model ranks better, but at the fixed threshold it flags more texts as machine-written: it catches 635 of 700 machine texts and flags 2,389 human ones, against 552 and 1,343 as run.
In domain 2 the grouped folds push C to the bottom of the grid (0.0001) and macro-F1 falls from 0.740 to 0.584, so the fitting choices do not flatter every number in the same direction. Running the solver to convergence at the 2023 C moves domain-1 macro-F1 from 0.697 to 0.688 and AUC from 0.960 to 0.950. Every comparison below that involves the 2023 recipe inherits these choices, which is why the re-balancing section repeats its comparison with converged fits.
Paired comparison
Logistic regression against SGD, text by text
Both detectors scored the same texts, so the comparison uses only the texts where they disagree (McNemar's test) and resamples texts in pairs for the differences. On domain 1, logistic regression is ahead on every measure, and it was ahead on macro-F1 in 10 of 10 splits. On domain 2 the 15 disagreements give no evidence either way (p = 0.30).
Domain 1
| SGD right | SGD wrong | |
|---|---|---|
| LR right | 22,252 | 1,474 |
| LR wrong | 721 | 770 |
- McNemar, exact
- p < 0.001 (2,195 discordant)
- Odds ratio b / c (95% CI)
- 2.04 (1.87 to 2.23) · g = 0.17
- Accuracy, LR − SGD
- +0.030 (+0.026 to +0.033)
- Macro-F1, LR − SGD
- +0.083 (+0.074 to +0.093)
- AUC, LR − SGD
- +0.058 (+0.051 to +0.065)
Domain 2
| SGD right | SGD wrong | |
|---|---|---|
| LR right | 76 | 10 |
| LR wrong | 5 | 9 |
- McNemar, exact
- p = 0.30 (15 discordant)
- Odds ratio b / c (95% CI)
- 2.00 (0.68 to 5.85) · g = 0.17
- Accuracy, LR − SGD
- +0.050 (−0.020 to +0.130)
- Macro-F1, LR − SGD
- +0.042 (−0.084 to +0.171)
- AUC, LR − SGD
- +0.021 (−0.075 to +0.121)
The 2023 report argued that logistic regression generalises better on domain 2. On a clean split that claim is not supported or ruled out: the domain-2 macro-F1 difference is +0.042 (−0.084 to +0.171), and across the ten splits logistic regression was ahead in 8 of 10.
Calibration
Can P(human) be read as a probability?
A detector is calibrated if, among texts it gives P(human) = 0.8, about 80% are human. Points on the diagonal are calibrated, the bars are Wilson 95% intervals and a point's area shows how many texts fall in its bin. The hinge-loss SGD model has no probabilities of its own, so it gets Platt scaling, a logistic regression on its cross-validated scores.
Both detectors learnt from re-balanced data with balanced class weights, so their probabilities assume half the texts are human.
Domain 1
Domain 2
| Probabilities | ECE, domain 1 | Brier, domain 1 | ECE, domain 2 | Brier, domain 2 |
|---|---|---|---|---|
| LR, as trained | 0.054 (0.051 to 0.056) | 0.049 (0.047 to 0.051) | 0.112 (0.064 to 0.186) | 0.123 (0.069 to 0.183) |
| SGD + Platt scaling, as trained | 0.289 (0.286 to 0.292) | 0.143 (0.141 to 0.146) | 0.305 (0.237 to 0.383) | 0.241 (0.232 to 0.251) |
| LR, prior-corrected | 0.023 (0.022 to 0.025) | 0.025 (0.023 to 0.026) | 0.123 (0.072 to 0.197) | 0.129 (0.071 to 0.191) |
| SGD + Platt, prior-corrected | 0.007 (0.005 to 0.009) | 0.025 (0.023 to 0.027) | 0.050 (0.014 to 0.126) | 0.152 (0.109 to 0.196) |
| LR without re-balancing | 0.002 (0.002 to 0.004) | 0.014 (0.013 to 0.015) | 0.105 (0.058 to 0.179) | 0.122 (0.067 to 0.182) |
ECE uses 10 equal-width bins of P(human). Lower is better for both measures. The prior correction uses only training-part counts, never the test labels.
Domain shift
A detector trained on one domain does not transfer
Each row is a detector, each column a test part. The shaded cells are in-domain. Trained on domain 1 and tested on all 500 domain-2 texts, logistic regression reaches AUC 0.553 (0.483 to 0.621), not distinguishable from chance (the interval includes 0.5). Pooling both domains into one model (correctly this time, see the audit) is worse than a dedicated model on domain 1 on every measure (macro-F1 +0.040 (+0.032 to +0.048), own minus pooled). On domain 2 it is worse in accuracy (+0.140 (+0.050 to +0.230), p = 0.004), while the macro-F1 (+0.135 (−0.014 to +0.273)) and AUC (+0.113 (−0.098 to +0.316)) differences are not distinguishable from zero.
| Trained on | Tested on domain 1 (n = 25,217) | Tested on domain 2 (n = 100) |
|---|---|---|
| Domain 1 | macro-F10.697 (0.685 to 0.709)AUC0.960 (0.954 to 0.965) | macro-F10.452 (0.361 to 0.548)AUC0.479 (0.336 to 0.620) |
| Domain 2 | macro-F10.107 (0.103 to 0.111)AUC0.639 (0.623 to 0.655) | macro-F10.740 (0.610 to 0.850)AUC0.760 (0.607 to 0.894) |
| Both domains (pooled) | macro-F10.658 (0.646 to 0.669)AUC0.954 (0.948 to 0.959) | macro-F10.605 (0.493 to 0.709)AUC0.647 (0.501 to 0.786) |
Own model vs pooled model, domain 2
| Pooled right | Pooled wrong | |
|---|---|---|
| Own right | 68 | 18 |
| Own wrong | 4 | 10 |
- McNemar, exact
- p = 0.004 (22 discordant)
- Odds ratio b / c (95% CI)
- 4.50 (1.52 to 13.30) · g = 0.32
- Accuracy, Own − Pooled
- +0.140 (+0.050 to +0.230)
- Macro-F1, Own − Pooled
- +0.135 (−0.014 to +0.273)
- AUC, Own − Pooled
- +0.113 (−0.098 to +0.316)
Own model vs domain-1 model, domain 2
| Domain 1 right | Domain 1 wrong | |
|---|---|---|
| Own right | 50 | 36 |
| Own wrong | 4 | 10 |
- McNemar, exact
- p < 0.001 (40 discordant)
- Odds ratio b / c (95% CI)
- 9.00 (3.20 to 25.29) · g = 0.40
- Accuracy, Own − Domain 1
- +0.320 (+0.210 to +0.430)
- Macro-F1, Own − Domain 1
- +0.287 (+0.142 to +0.421)
- AUC, Own − Domain 1
- +0.281 (+0.065 to +0.494)
Known failure mode
Generators the detector has not seen
The machine texts came from 5 generators in domain 1 and 4 in domain 2 (the course labelled them only by number). To mimic a new model appearing after training, logistic regression was re-fitted five times, each time leaving one domain-1 generator out of training, and scored on that generator's test texts.
Leaving a generator out does not hide its prompts. The bag of words includes the prompt, and between 118 and 140 of each generator's 140 test texts answer a prompt that the other generators' training texts also answer. The second view removes every training text (human or machine) that answers one of those prompts, and compares a model that keeps the generator with one that leaves it out.
| Generator | SGD, seen | LR, seen | LR, generator left out | McNemar | AUC drop (seen − unseen) |
|---|---|---|---|---|---|
| 0 | 0.436 (0.356 to 0.518)61 of 140 | 0.629 (0.546 to 0.704)88 of 140 | 0.593 (0.510 to 0.671)83 of 140 | p = 0.27 (9 vs 4) | −0.003 (−0.011 to +0.007) |
| 1 | 0.500 (0.418 to 0.582)70 of 140 | 0.700 (0.620 to 0.770)98 of 140 | 0.621 (0.539 to 0.698)87 of 140 | p = 0.03 (16 vs 5) | +0.003 (−0.004 to +0.010) |
| 2 | 0.721 (0.642 to 0.789)101 of 140 | 0.850 (0.782 to 0.900)119 of 140 | 0.764 (0.688 to 0.827)107 of 140 | p = 0.002 (13 vs 1) | +0.022 (+0.013 to +0.033) |
| 3 | 0.471 (0.391 to 0.554)66 of 140 | 0.779 (0.703 to 0.839)109 of 140 | 0.643 (0.561 to 0.717)90 of 140 | p < 0.001 (23 vs 4) | +0.007 (+0.003 to +0.012) |
| 4 | 0.914 (0.856 to 0.950)128 of 140 | 0.986 (0.949 to 0.996)138 of 140 | 0.950 (0.900 to 0.976)133 of 140 | p = 0.13 (6 vs 1) | +0.010 (+0.004 to +0.018) |
What this shows
With the generator left out, the share of its texts caught fell for all five, by between 5 and 19 of 140 texts. The drop is significant for three of the five (McNemar p = 0.03 or less), and AUC changed by at most 0.022, so the unseen generator's texts were still mostly ranked below human texts, but more of them landed on the wrong side of the fixed threshold.
With its test prompts removed from training as well, leaving the generator out lowered the share caught for all five, significantly for four (McNemar p < 0.05). A detector like this has to be re-checked whenever a new generator appears, but this data shows a loss of recall, not a collapse.
Domain 2 generators
- Generator 0LR 0.950 (0.764 to 0.991)19 of 20
- Generator 1LR 0.900 (0.699 to 0.972)18 of 20
- Generator 2LR 1.000 (0.839 to 1.000)20 of 20
- Generator 3LR 1.000 (0.839 to 1.000)20 of 20
Re-balancing
The 2023 re-balancing did not help, however the models are fitted
Five ways to handle the imbalance, all applied to the training part only and all scored on the same test texts with logistic regression. On domain 1 the model with no re-balancing at all has the best macro-F1, 0.809 (0.791 to 0.825), the best AUC and near-perfect calibration. The 2023 recipe minus that model is −0.111 (−0.128 to −0.095) in macro-F1 on the same texts.
As fitted in 2023 that comparison is partly between regularisation regimes: cross-validation chose C = 10,000 for the recipe and 0.046 without re-balancing. The converged tabs re-fit every strategy with the solver run to convergence and C chosen on folds grouped by source text. No re-balancing is still best on all three measures in domain 1, and the recipe's gap becomes −0.163 (−0.181 to −0.145). In domain 2 the strategies are indistinguishable as fitted in 2023: the recipe minus no re-balancing is −0.025 (−0.084 to 0.000). With converged fits the recipe falls behind, and the difference is −0.168 (−0.277 to −0.057). The recipe does catch more machine texts, at the price of flagging far more human ones. The decision record discusses that trade-off.
| Strategy (training part only) | Rows | C | Macro-F1 | AUC | Machine texts caught | ECE of P(human) |
|---|---|---|---|---|---|---|
| No re-balancing, no class weights | 100,867 | 0.046 | 0.809 (0.791 to 0.825) | 0.975 (0.970 to 0.978) | 0.509 (0.472 to 0.545)356 of 700 | 0.002 (0.002 to 0.004) |
| Class weights only | 100,867 | 10,000 | 0.739 (0.726 to 0.752) | 0.969 (0.964 to 0.974) | 0.777 (0.745 to 0.806)544 of 700 | 0.037 (0.035 to 0.039) |
| Index-SMOTE copies + class weights (the 2023 recipe) | 40,000 | 10,000 | 0.697 (0.685 to 0.709) | 0.960 (0.954 to 0.965) | 0.789 (0.757 to 0.817)552 of 700 | 0.054 (0.051 to 0.056) |
| SMOTE on the bag-of-words features + class weights | 40,000 | 10,000 | 0.700 (0.687 to 0.712) | 0.959 (0.953 to 0.964) | 0.743 (0.709 to 0.774)520 of 700 | 0.049 (0.046 to 0.051) |
| Random under-sampling + class weights | 5,600 | 0.046 | 0.610 (0.601 to 0.619) | 0.965 (0.960 to 0.970) | 0.934 (0.913 to 0.950)654 of 700 | 0.141 (0.138 to 0.144) |
Converged fits get up to 5,000 lbfgs iterations. Only the index-SMOTE copies need grouped folds; every other strategy has one row per text. Not every fit reached convergence within that budget: no re-balancing, domain 1 (3 of 50 cross-validation fits stopped at the limit); class weights only, domain 1 (4 of 50 cross-validation fits and the final fit stopped at the limit). They use the full 100,867-text domain-1 training part; read those rows as approximate.
Robustness
Ten random splits, ten prompt-grouped splits
A bootstrap interval treats the training data as fixed. Re-drawing the split (seeds 90051 to 90060) and re-fitting shows how much the split itself moves the scores. Each seed is used twice: once for a random split stratified by class and generator, and once for a split that keeps every prompt on one side. Each dot is one split, the blue bar is the mean, the ring marks the main split (seed 90051), and the numbers are mean ± standard deviation and the range.
Domain 1
- LR random0.696 ± 0.0040.689 to 0.702
- LR grouped0.676 ± 0.0100.661 to 0.695
- SGD random0.614 ± 0.0050.604 to 0.621
- SGD grouped0.593 ± 0.0090.582 to 0.610
- LR random0.958 ± 0.0030.953 to 0.962
- LR grouped0.941 ± 0.0040.935 to 0.948
- SGD random0.903 ± 0.0060.894 to 0.912
- SGD grouped0.874 ± 0.0060.862 to 0.885
Domain 2
- LR random0.657 ± 0.0610.569 to 0.750
- LR grouped0.558 ± 0.0390.515 to 0.632
- SGD random0.617 ± 0.0680.492 to 0.706
- SGD grouped0.513 ± 0.0510.443 to 0.595
- LR random0.725 ± 0.0590.637 to 0.807
- LR grouped0.613 ± 0.0590.523 to 0.723
- SGD random0.618 ± 0.0680.536 to 0.739
- SGD grouped0.492 ± 0.0450.405 to 0.570
Texts that share a prompt
Several texts answer the same prompt. In domain 2 a prompt is answered either by one human text or by four machine texts, never both (0 of the 400 machine texts share a prompt with a human text). In the main split, 80 of 80 domain-2 machine test texts answer a prompt seen in training, against 0 of 20 human ones, so the rule “machine if the prompt was seen” alone scores macro-F1 1.000. Across the ten random splits the machine share ranges from 95.0% to 100.0%, and no human test text ever has a seen prompt. In domain 1 the shortcut is weaker: 97.0% of machine and 65.1% of human test texts answer a seen prompt.
Compared with the ten random splits rather than with the main split alone, the prompt-grouped splits lower domain-2 logistic regression's AUC from 0.725 ± 0.059 to 0.613 ± 0.059, and its macro-F1 from 0.657 ± 0.061 to 0.558 ± 0.039. In domain 1 the grouped macro-F1 is 0.676 ± 0.010 against 0.696 ± 0.004. The domain-2 headline numbers were measured on a split where this shortcut is available.
| Detector | Main split | Prompt-grouped, seed 90051 |
|---|---|---|
| LR · domain 1n = 25,217 / 25,425 | macro-F10.697 (0.685 to 0.709)AUC0.960 (0.954 to 0.965) | macro-F10.673 (0.661 to 0.685)AUC0.945 (0.937 to 0.952) |
| SGD · domain 1n = 25,217 / 25,425 | macro-F10.614 (0.603 to 0.625)AUC0.902 (0.893 to 0.911) | macro-F10.587 (0.577 to 0.597)AUC0.874 (0.862 to 0.885) |
| LR · domain 2n = 100 / 103 | macro-F10.740 (0.610 to 0.850)AUC0.760 (0.607 to 0.894) | macro-F10.595 (0.484 to 0.701)AUC0.600 (0.440 to 0.754) |
| SGD · domain 2n = 100 / 103 | macro-F10.697 (0.577 to 0.805)AUC0.739 (0.588 to 0.870) | macro-F10.505 (0.407 to 0.603)AUC0.497 (0.344 to 0.655) |
| Macro-F1 | Resampling texts | Resampling prompts |
|---|---|---|
| LR · domain 1 | 0.697 (0.685 to 0.709) | 0.697 (0.683 to 0.711)1.17× as wide |
| SGD · domain 1 | 0.614 (0.603 to 0.625) | 0.614 (0.602 to 0.626)1.14× as wide |
| LR − SGD · domain 1 | +0.083 (+0.074 to +0.093) | +0.083 (+0.073 to +0.093)1.04× as wide |
| LR · domain 2 | 0.740 (0.610 to 0.850) | 0.740 (0.609 to 0.850)1.01× as wide |
| SGD · domain 2 | 0.697 (0.577 to 0.805) | 0.697 (0.575 to 0.805)1.00× as wide |
| LR − SGD · domain 2 | +0.042 (−0.084 to +0.171) | +0.042 (−0.087 to +0.166)0.99× as wide |
Every other interval on this page resamples texts as if they were independent. Texts that answer the same prompt are not, so the second table resamples whole prompts instead (domain 1: 25,217 test texts in 19,311 prompts; domain 2: 100 texts in 82). Where the cluster interval is wider, the text-level one understates the uncertainty.
The 2023 numbers, with intervals
A narrow interval around a biased number
The report's validation table as printed, with intervals computed from the re-derived confusion matrices: Wilson intervals for accuracy and bootstrap intervals (10,000 resamples, seed 90051) for F1. The domain-1 intervals are narrow because there were 10,000 validation rows. They describe sampling noise only. The leak is a bias, and no interval can reveal a bias.
The per-class recalls below show where the bias sat. In each domain only the minority class was copied, so only its 2023 recall is inflated. The class that was never copied scores about the same on clean texts.
Logistic regression, domain 1
- Machine texts caught2023 validation, copied0.9980.997 to 0.999
- Machine texts caughtclean test0.7890.757 to 0.817
- Human texts kept2023 validation0.9330.925 to 0.939
- Human texts keptclean test0.9450.942 to 0.948
Logistic regression, domain 2
- Machine texts caught2023 validation0.9100.833 to 0.954
- Machine texts caughtclean test0.9630.895 to 0.987
- Human texts kept2023 validation, copied0.9860.924 to 0.998
- Human texts keptclean test0.4500.258 to 0.658
| Model | Rows | Accuracy, as printed (Wilson 95%) | Human F1, as printed (bootstrap 95%) |
|---|---|---|---|
| LR 1 | 10,000 | 0.9658 (0.962 to 0.969) | 0.9643 (0.960 to 0.968) |
| LR 2 | 160 | 0.9437 (0.897 to 0.970) | 0.9396 (0.895 to 0.975) |
| LR 1+2 | 10,000 | 0.5005 (0.491 to 0.510) | 0.5002 (0.488 to 0.512) |
| SGD 1 | 10,000 | 0.9330 (0.928 to 0.938) | 0.9315 (0.926 to 0.936) |
| SGD 2 | 160 | 0.5312 (0.454 to 0.607) | 0.5562 (0.462 to 0.640) |
| SGD 1+2 | 10,000 | 0.5019 (0.492 to 0.512) | 0.4493 (0.436 to 0.462) |
The SGD 2 row is the report's, which scored the domain-1 SGD model on domain-2 data. The 1+2 rows come from models trained on mismatched features and labels. Both are explained in the revival audit. The printed SGD 2 precision slip does not affect these two columns.
Downloads
Take the numbers with you
The aggregates behind this page, exactly as the script wrote them. They contain no text, token sequence or per-item prediction. To reproduce them, run uv run scripts/evaluate_heldout.py from the repository root (about 40 minutes on a 10-core laptop, most of it in the converged sensitivity fits).