Skip to content

Exhibit E · held-out re-evaluation (2026)

The same recipes, scored on texts they never saw.

The 2023 validation scores were inflated: the notebooks re-balanced each domain before splitting it, so copies of the same machine text sat on both sides of the split (the revival audit has the details). The original models had seen almost every domain-1 machine text, so they cannot be re-scored cleanly. This page re-fits the 2023 recipes, unchanged, on a split made before any re-balancing and scores the untouched 20%.

Every number comes with a 95% interval and its sample size. The 2023 numbers stay on the journal exactly as printed.

Design

Split first, then re-balance

The pipeline is the notebooks' own, in a different order. Everything below was produced by scripts/evaluate_heldout.py with the 2023 library versions (scikit-learn 1.2.2, imbalanced-learn 0.10.1).

  1. 01Split the original rows80/20 within each domain, stratified by class and generator, seed 90051. The test part is never re-balanced or looked at during fitting.
  2. 02Re-balance the training partThe notebooks' call: SMOTE on the integer row position, which copies minority rows. Domain 1 is capped at 40,000 rows, the notebooks' 80% of 50,000.
  3. 03Fit the 2023 modelsLogisticRegressionCV (l2, lbfgs, 50 iterations, balanced weights) and SGDClassifier (hinge, l1, alpha 0.0001, balanced weights). Cross-validation chose C = 10,000 for domain 1, the top of its grid, and 2.78 for domain 2. The 50 iterations stop the fits early (see Fitting).
  4. 04Score with intervals10,000 bootstrap resamples of the test part (seed 90051). Every statistic on one test set uses the same resamples, so differences between detectors are paired.
Training and test sizes per domain and class
DomainTrain humanTrain machineTest humanTest machine
198,0672,80024,517700
2803202080

Why accuracy alone misleads here

97.2% of the domain-1 test texts are human, so a rule that always says “human” scores 0.972 accuracy, more than any detector here. The page therefore leads with macro-F1, which averages the F1 of the two classes, and ROC AUC, which asks how often a random human text is ranked above a random machine text.

Results

Machine recall falls on unseen texts, and length explains most of the ranking

On domain 1, logistic regression reaches macro-F1 0.697 (0.685 to 0.709) and AUC 0.960 (0.954 to 0.965). Compared like for like with 2023, the change sits in the copied class: the leaky split had it catching 0.998 of the machine texts, and on texts it never saw it catches 0.789 (0.757 to 0.817) (Wilson 95%), while human recall stays about the same (0.933 then, 0.945 now; see the 2023 numbers). The printed 2023 F1 of 0.9643 is the human-class F1 on a 50:50 validation set, so it is not comparable with macro-F1 on the real 97:3 mix.

A logistic regression on the text length alone reaches AUC 0.932 (0.923 to 0.941): machine texts in domain 1 are about half as long as human ones. The bag-of-words detector adds +0.027 (+0.019 to +0.036) AUC over length on the same texts. Domain 2 has only 100 test texts, 20 of them human, and the intervals show it. Under each detector is its mean ± standard deviation over ten re-drawn splits. The main domain-2 split ranks 2 of 10 on logistic regression's macro-F1 (1 is the best), so its headline numbers sit on the favourable side.

Domain 1 test part

24,517 human + 700 machine texts
Macro-F1 (mean of the human and machine F1)
  • Logistic regression10 splits: 0.696 ± 0.0040.6970.685 to 0.709
  • SGD (hinge)10 splits: 0.614 ± 0.0050.6140.603 to 0.625
  • Length only0.5640.556 to 0.571
  • Always the majority0.4930.492 to 0.493
ROC AUC (ranking, no threshold)
  • Logistic regression10 splits: 0.958 ± 0.0030.9600.954 to 0.965
  • SGD (hinge)10 splits: 0.903 ± 0.0060.9020.893 to 0.911
  • Length only0.9320.923 to 0.941

chance (0.500)

Accuracy
  • Logistic regression10 splits: 0.941 ± 0.0010.9410.938 to 0.944
  • SGD (hinge)10 splits: 0.910 ± 0.0030.9110.908 to 0.915
  • Length only0.8270.823 to 0.832
  • Always the majority0.9720.970 to 0.974

ROC curves

positive class: human
000.250.250.50.50.750.7511false-positive rate (machine called human)true-positive rate (human caught)
LR · AUC 0.960SGD · AUC 0.902

Recall per class

Wilson 95% intervals
Logistic regression
Human texts kept as human 0.945 (0.942 to 0.948)23,174 of 24,517Machine texts caught 0.789 (0.757 to 0.817)552 of 700
SGD (hinge)
Human texts kept as human 0.920 (0.916 to 0.923)22,547 of 24,517Machine texts caught 0.609 (0.572 to 0.644)426 of 700

Fitting sensitivity

The 2023 fits stop early and choose C on leaky folds

The recipe gives LogisticRegressionCV 50 lbfgs iterations. In 36 of the 66 fits on this page that use that budget, every cross-validation fit stopped at the limit, and in 38 of them the chosen C was the largest value in the grid. For the domain-1 recipe that is all 50 cross-validation fits and the final fit (51 convergence warnings), with C = 10,000. The 2023 notebooks silenced the warnings; the script now records them for every fit. The inner folds also hold copies of the same text on both sides, because the copies are made before cross-validation splits the training part, so inner-fold accuracy keeps rising with C. SGD on domain 1 also ran all 1,000 of its epochs without meeting its stopping rule.

The table re-fits the recipe with each problem removed in turn, on the same training and test texts. The headline numbers above stay the recipe as run, because that is what the team built. This section shows how much they depend on the fitting.

The 2023 logistic-regression recipe under different fitting settings, domain 1 test part
FitIterationsCMacro-F1AUCMachine texts caughtECE of P(human)
As run (the 2023 settings)inner folds: stratifiedlimit 5050/50 CV fits at limit10,000top of grid0.697 (0.685 to 0.709)0.960 (0.954 to 0.965)0.789 (0.757 to 0.817)552 of 7000.054 (0.051 to 0.056)
C chosen on folds grouped by source rowinner folds: grouped by source rowlimit 5050/50 CV fits at limit0.0060.646 (0.636 to 0.656)0.971 (0.967 to 0.975)0.906 (0.882 to 0.925)634 of 7000.112 (0.109 to 0.115)
Solver run to convergenceinner folds: stratifiedlimit 5,0000/50 CV fits at limit1,2920.687 (0.675 to 0.700)0.950 (0.943 to 0.957)0.733 (0.699 to 0.764)513 of 7000.057 (0.054 to 0.060)
Both: grouped folds and convergenceinner folds: grouped by source rowlimit 5,0000/50 CV fits at limit0.0060.644 (0.634 to 0.654)0.971 (0.966 to 0.975)0.907 (0.883 to 0.926)635 of 7000.112 (0.109 to 0.115)
The 2023 C, run to convergenceinner folds: none (C fixed at the 2023 choice)limit 20,0003,101 used10,0000.688 (0.675 to 0.700)0.950 (0.943 to 0.957)0.731 (0.697 to 0.763)512 of 7000.058 (0.055 to 0.060)

Choosing C

5 inner folds, accuracy
0.850.900.950.00010.04621.5410,000C (log scale)
Domain 1: mean inner-fold accuracy for each C2023 folds, 50 iterations2023 folds, convergedgrouped folds, 50 iterationsgrouped folds, converged

What changes when the fit is done properly

With the solver run to convergence and C chosen on folds that keep every copy of a text together, cross-validation picks C = 0.006 for domain 1 instead of 10,000. On the same test texts the recipe as run minus that fit is +0.053 (+0.045 to +0.061) in macro-F1 and −0.011 (−0.015 to −0.008) in AUC. The more strongly regularised model ranks better, but at the fixed threshold it flags more texts as machine-written: it catches 635 of 700 machine texts and flags 2,389 human ones, against 552 and 1,343 as run.

In domain 2 the grouped folds push C to the bottom of the grid (0.0001) and macro-F1 falls from 0.740 to 0.584, so the fitting choices do not flatter every number in the same direction. Running the solver to convergence at the 2023 C moves domain-1 macro-F1 from 0.697 to 0.688 and AUC from 0.960 to 0.950. Every comparison below that involves the 2023 recipe inherits these choices, which is why the re-balancing section repeats its comparison with converged fits.

Paired comparison

Logistic regression against SGD, text by text

Both detectors scored the same texts, so the comparison uses only the texts where they disagree (McNemar's test) and resamples texts in pairs for the differences. On domain 1, logistic regression is ahead on every measure, and it was ahead on macro-F1 in 10 of 10 splits. On domain 2 the 15 disagreements give no evidence either way (p = 0.30).

Domain 1

25,217 test texts
LR against SGD on the same 25,217 texts
SGD rightSGD wrong
LR right22,2521,474
LR wrong721770
McNemar, exact
p < 0.001 (2,195 discordant)
Odds ratio b / c (95% CI)
2.04 (1.87 to 2.23) · g = 0.17
Accuracy, LR − SGD
+0.030 (+0.026 to +0.033)
Macro-F1, LR − SGD
+0.083 (+0.074 to +0.093)
AUC, LR − SGD
+0.058 (+0.051 to +0.065)

Domain 2

100 test texts
LR against SGD on the same 100 texts
SGD rightSGD wrong
LR right7610
LR wrong59
McNemar, exact
p = 0.30 (15 discordant)
Odds ratio b / c (95% CI)
2.00 (0.68 to 5.85) · g = 0.17
Accuracy, LR − SGD
+0.050 (−0.020 to +0.130)
Macro-F1, LR − SGD
+0.042 (−0.084 to +0.171)
AUC, LR − SGD
+0.021 (−0.075 to +0.121)

The 2023 report argued that logistic regression generalises better on domain 2. On a clean split that claim is not supported or ruled out: the domain-2 macro-F1 difference is +0.042 (−0.084 to +0.171), and across the ten splits logistic regression was ahead in 8 of 10.

Calibration

Can P(human) be read as a probability?

A detector is calibrated if, among texts it gives P(human) = 0.8, about 80% are human. Points on the diagonal are calibrated, the bars are Wilson 95% intervals and a point's area shows how many texts fall in its bin. The hinge-loss SGD model has no probabilities of its own, so it gets Platt scaling, a logistic regression on its cross-validated scores.

Both detectors learnt from re-balanced data with balanced class weights, so their probabilities assume half the texts are human.

Domain 1

25,217 test texts
000.250.250.50.50.750.7511mean predicted P(human) in the binobserved share of human texts
Lines join bins with at least 10 texts.LR, as trained · ECE 0.054SGD + Platt scaling, as trained · ECE 0.289

Domain 2

100 test texts
000.250.250.50.50.750.7511mean predicted P(human) in the binobserved share of human texts
Lines join bins with at least 10 texts.LR, as trained · ECE 0.112SGD + Platt scaling, as trained · ECE 0.305
Expected calibration error and Brier score
ProbabilitiesECE, domain 1Brier, domain 1ECE, domain 2Brier, domain 2
LR, as trained0.054 (0.051 to 0.056)0.049 (0.047 to 0.051)0.112 (0.064 to 0.186)0.123 (0.069 to 0.183)
SGD + Platt scaling, as trained0.289 (0.286 to 0.292)0.143 (0.141 to 0.146)0.305 (0.237 to 0.383)0.241 (0.232 to 0.251)
LR, prior-corrected0.023 (0.022 to 0.025)0.025 (0.023 to 0.026)0.123 (0.072 to 0.197)0.129 (0.071 to 0.191)
SGD + Platt, prior-corrected0.007 (0.005 to 0.009)0.025 (0.023 to 0.027)0.050 (0.014 to 0.126)0.152 (0.109 to 0.196)
LR without re-balancing0.002 (0.002 to 0.004)0.014 (0.013 to 0.015)0.105 (0.058 to 0.179)0.122 (0.067 to 0.182)

ECE uses 10 equal-width bins of P(human). Lower is better for both measures. The prior correction uses only training-part counts, never the test labels.

Domain shift

A detector trained on one domain does not transfer

Each row is a detector, each column a test part. The shaded cells are in-domain. Trained on domain 1 and tested on all 500 domain-2 texts, logistic regression reaches AUC 0.553 (0.483 to 0.621), not distinguishable from chance (the interval includes 0.5). Pooling both domains into one model (correctly this time, see the audit) is worse than a dedicated model on domain 1 on every measure (macro-F1 +0.040 (+0.032 to +0.048), own minus pooled). On domain 2 it is worse in accuracy (+0.140 (+0.050 to +0.230), p = 0.004), while the macro-F1 (+0.135 (−0.014 to +0.273)) and AUC (+0.113 (−0.098 to +0.316)) differences are not distinguishable from zero.

Macro-F1 and AUC of logistic regression by training and test domain
Trained onTested on domain 1 (n = 25,217)Tested on domain 2 (n = 100)
Domain 1macro-F10.697 (0.685 to 0.709)AUC0.960 (0.954 to 0.965)macro-F10.452 (0.361 to 0.548)AUC0.479 (0.336 to 0.620)
Domain 2macro-F10.107 (0.103 to 0.111)AUC0.639 (0.623 to 0.655)macro-F10.740 (0.610 to 0.850)AUC0.760 (0.607 to 0.894)
Both domains (pooled)macro-F10.658 (0.646 to 0.669)AUC0.954 (0.948 to 0.959)macro-F10.605 (0.493 to 0.709)AUC0.647 (0.501 to 0.786)

Own model vs pooled model, domain 2

logistic regression
Own against Pooled on the same 100 texts
Pooled rightPooled wrong
Own right6818
Own wrong410
McNemar, exact
p = 0.004 (22 discordant)
Odds ratio b / c (95% CI)
4.50 (1.52 to 13.30) · g = 0.32
Accuracy, Own − Pooled
+0.140 (+0.050 to +0.230)
Macro-F1, Own − Pooled
+0.135 (−0.014 to +0.273)
AUC, Own − Pooled
+0.113 (−0.098 to +0.316)

Own model vs domain-1 model, domain 2

logistic regression
Own against Domain 1 on the same 100 texts
Domain 1 rightDomain 1 wrong
Own right5036
Own wrong410
McNemar, exact
p < 0.001 (40 discordant)
Odds ratio b / c (95% CI)
9.00 (3.20 to 25.29) · g = 0.40
Accuracy, Own − Domain 1
+0.320 (+0.210 to +0.430)
Macro-F1, Own − Domain 1
+0.287 (+0.142 to +0.421)
AUC, Own − Domain 1
+0.281 (+0.065 to +0.494)

Known failure mode

Generators the detector has not seen

The machine texts came from 5 generators in domain 1 and 4 in domain 2 (the course labelled them only by number). To mimic a new model appearing after training, logistic regression was re-fitted five times, each time leaving one domain-1 generator out of training, and scored on that generator's test texts.

Leaving a generator out does not hide its prompts. The bag of words includes the prompt, and between 118 and 140 of each generator's 140 test texts answer a prompt that the other generators' training texts also answer. The second view removes every training text (human or machine) that answers one of those prompts, and compares a model that keeps the generator with one that leaves it out.

Share of each generator's machine texts caught, with and without that generator in training
GeneratorSGD, seenLR, seenLR, generator left outMcNemarAUC drop (seen − unseen)
00.436 (0.356 to 0.518)61 of 1400.629 (0.546 to 0.704)88 of 1400.593 (0.510 to 0.671)83 of 140p = 0.27 (9 vs 4)−0.003 (−0.011 to +0.007)
10.500 (0.418 to 0.582)70 of 1400.700 (0.620 to 0.770)98 of 1400.621 (0.539 to 0.698)87 of 140p = 0.03 (16 vs 5)+0.003 (−0.004 to +0.010)
20.721 (0.642 to 0.789)101 of 1400.850 (0.782 to 0.900)119 of 1400.764 (0.688 to 0.827)107 of 140p = 0.002 (13 vs 1)+0.022 (+0.013 to +0.033)
30.471 (0.391 to 0.554)66 of 1400.779 (0.703 to 0.839)109 of 1400.643 (0.561 to 0.717)90 of 140p < 0.001 (23 vs 4)+0.007 (+0.003 to +0.012)
40.914 (0.856 to 0.950)128 of 1400.986 (0.949 to 0.996)138 of 1400.950 (0.900 to 0.976)133 of 140p = 0.13 (6 vs 1)+0.010 (+0.004 to +0.018)

What this shows

With the generator left out, the share of its texts caught fell for all five, by between 5 and 19 of 140 texts. The drop is significant for three of the five (McNemar p = 0.03 or less), and AUC changed by at most 0.022, so the unseen generator's texts were still mostly ranked below human texts, but more of them landed on the wrong side of the fixed threshold.

With its test prompts removed from training as well, leaving the generator out lowered the share caught for all five, significantly for four (McNemar p < 0.05). A detector like this has to be re-checked whenever a new generator appears, but this data shows a loss of recall, not a collapse.

Domain 2 generators

20 test texts each
  • Generator 0LR 0.950 (0.764 to 0.991)19 of 20
  • Generator 1LR 0.900 (0.699 to 0.972)18 of 20
  • Generator 2LR 1.000 (0.839 to 1.000)20 of 20
  • Generator 3LR 1.000 (0.839 to 1.000)20 of 20

Re-balancing

The 2023 re-balancing did not help, however the models are fitted

Five ways to handle the imbalance, all applied to the training part only and all scored on the same test texts with logistic regression. On domain 1 the model with no re-balancing at all has the best macro-F1, 0.809 (0.791 to 0.825), the best AUC and near-perfect calibration. The 2023 recipe minus that model is −0.111 (−0.128 to −0.095) in macro-F1 on the same texts.

As fitted in 2023 that comparison is partly between regularisation regimes: cross-validation chose C = 10,000 for the recipe and 0.046 without re-balancing. The converged tabs re-fit every strategy with the solver run to convergence and C chosen on folds grouped by source text. No re-balancing is still best on all three measures in domain 1, and the recipe's gap becomes −0.163 (−0.181 to −0.145). In domain 2 the strategies are indistinguishable as fitted in 2023: the recipe minus no re-balancing is −0.025 (−0.084 to 0.000). With converged fits the recipe falls behind, and the difference is −0.168 (−0.277 to −0.057). The recipe does catch more machine texts, at the price of flagging far more human ones. The decision record discusses that trade-off.

Logistic regression under five re-balancing strategies, domain 1 test part, fitted with the 2023 settings
Strategy (training part only)RowsCMacro-F1AUCMachine texts caughtECE of P(human)
No re-balancing, no class weights100,8670.0460.809 (0.791 to 0.825)0.975 (0.970 to 0.978)0.509 (0.472 to 0.545)356 of 7000.002 (0.002 to 0.004)
Class weights only100,86710,0000.739 (0.726 to 0.752)0.969 (0.964 to 0.974)0.777 (0.745 to 0.806)544 of 7000.037 (0.035 to 0.039)
Index-SMOTE copies + class weights (the 2023 recipe)40,00010,0000.697 (0.685 to 0.709)0.960 (0.954 to 0.965)0.789 (0.757 to 0.817)552 of 7000.054 (0.051 to 0.056)
SMOTE on the bag-of-words features + class weights40,00010,0000.700 (0.687 to 0.712)0.959 (0.953 to 0.964)0.743 (0.709 to 0.774)520 of 7000.049 (0.046 to 0.051)
Random under-sampling + class weights5,6000.0460.610 (0.601 to 0.619)0.965 (0.960 to 0.970)0.934 (0.913 to 0.950)654 of 7000.141 (0.138 to 0.144)

Converged fits get up to 5,000 lbfgs iterations. Only the index-SMOTE copies need grouped folds; every other strategy has one row per text. Not every fit reached convergence within that budget: no re-balancing, domain 1 (3 of 50 cross-validation fits stopped at the limit); class weights only, domain 1 (4 of 50 cross-validation fits and the final fit stopped at the limit). They use the full 100,867-text domain-1 training part; read those rows as approximate.

Robustness

Ten random splits, ten prompt-grouped splits

A bootstrap interval treats the training data as fixed. Re-drawing the split (seeds 90051 to 90060) and re-fitting shows how much the split itself moves the scores. Each seed is used twice: once for a random split stratified by class and generator, and once for a split that keeps every prompt on one side. Each dot is one split, the blue bar is the mean, the ring marks the main split (seed 90051), and the numbers are mean ± standard deviation and the range.

Domain 1

25,217 test texts per random split
Macro-F1
  • LR random0.696 ± 0.0040.689 to 0.702
  • LR grouped0.676 ± 0.0100.661 to 0.695
  • SGD random0.614 ± 0.0050.604 to 0.621
  • SGD grouped0.593 ± 0.0090.582 to 0.610
ROC AUC
  • LR random0.958 ± 0.0030.953 to 0.962
  • LR grouped0.941 ± 0.0040.935 to 0.948
  • SGD random0.903 ± 0.0060.894 to 0.912
  • SGD grouped0.874 ± 0.0060.862 to 0.885

Domain 2

100 test texts per random split
Macro-F1
  • LR random0.657 ± 0.0610.569 to 0.750
  • LR grouped0.558 ± 0.0390.515 to 0.632
  • SGD random0.617 ± 0.0680.492 to 0.706
  • SGD grouped0.513 ± 0.0510.443 to 0.595
ROC AUC
  • LR random0.725 ± 0.0590.637 to 0.807
  • LR grouped0.613 ± 0.0590.523 to 0.723
  • SGD random0.618 ± 0.0680.536 to 0.739
  • SGD grouped0.492 ± 0.0450.405 to 0.570

Texts that share a prompt

Several texts answer the same prompt. In domain 2 a prompt is answered either by one human text or by four machine texts, never both (0 of the 400 machine texts share a prompt with a human text). In the main split, 80 of 80 domain-2 machine test texts answer a prompt seen in training, against 0 of 20 human ones, so the rule “machine if the prompt was seen” alone scores macro-F1 1.000. Across the ten random splits the machine share ranges from 95.0% to 100.0%, and no human test text ever has a seen prompt. In domain 1 the shortcut is weaker: 97.0% of machine and 65.1% of human test texts answer a seen prompt.

Compared with the ten random splits rather than with the main split alone, the prompt-grouped splits lower domain-2 logistic regression's AUC from 0.725 ± 0.059 to 0.613 ± 0.059, and its macro-F1 from 0.657 ± 0.061 to 0.558 ± 0.039. In domain 1 the grouped macro-F1 is 0.676 ± 0.010 against 0.696 ± 0.004. The domain-2 headline numbers were measured on a split where this shortcut is available.

Macro-F1 and AUC on the main split and on the prompt-grouped split with the same seed
DetectorMain splitPrompt-grouped, seed 90051
LR · domain 1n = 25,217 / 25,425macro-F10.697 (0.685 to 0.709)AUC0.960 (0.954 to 0.965)macro-F10.673 (0.661 to 0.685)AUC0.945 (0.937 to 0.952)
SGD · domain 1n = 25,217 / 25,425macro-F10.614 (0.603 to 0.625)AUC0.902 (0.893 to 0.911)macro-F10.587 (0.577 to 0.597)AUC0.874 (0.862 to 0.885)
LR · domain 2n = 100 / 103macro-F10.740 (0.610 to 0.850)AUC0.760 (0.607 to 0.894)macro-F10.595 (0.484 to 0.701)AUC0.600 (0.440 to 0.754)
SGD · domain 2n = 100 / 103macro-F10.697 (0.577 to 0.805)AUC0.739 (0.588 to 0.870)macro-F10.505 (0.407 to 0.603)AUC0.497 (0.344 to 0.655)
Macro-F1 with 95% intervals that resample texts and that resample whole prompts
Macro-F1Resampling textsResampling prompts
LR · domain 10.697 (0.685 to 0.709)0.697 (0.683 to 0.711)1.17× as wide
SGD · domain 10.614 (0.603 to 0.625)0.614 (0.602 to 0.626)1.14× as wide
LR − SGD · domain 1+0.083 (+0.074 to +0.093)+0.083 (+0.073 to +0.093)1.04× as wide
LR · domain 20.740 (0.610 to 0.850)0.740 (0.609 to 0.850)1.01× as wide
SGD · domain 20.697 (0.577 to 0.805)0.697 (0.575 to 0.805)1.00× as wide
LR − SGD · domain 2+0.042 (−0.084 to +0.171)+0.042 (−0.087 to +0.166)0.99× as wide

Every other interval on this page resamples texts as if they were independent. Texts that answer the same prompt are not, so the second table resamples whole prompts instead (domain 1: 25,217 test texts in 19,311 prompts; domain 2: 100 texts in 82). Where the cluster interval is wider, the text-level one understates the uncertainty.

The 2023 numbers, with intervals

A narrow interval around a biased number

The report's validation table as printed, with intervals computed from the re-derived confusion matrices: Wilson intervals for accuracy and bootstrap intervals (10,000 resamples, seed 90051) for F1. The domain-1 intervals are narrow because there were 10,000 validation rows. They describe sampling noise only. The leak is a bias, and no interval can reveal a bias.

The per-class recalls below show where the bias sat. In each domain only the minority class was copied, so only its 2023 recall is inflated. The class that was never copied scores about the same on clean texts.

Logistic regression, domain 1

copied class: machine
Recall per class, 2023 validation vs clean test (Wilson 95%)
  • Machine texts caught2023 validation, copied0.9980.997 to 0.999
  • Machine texts caughtclean test0.7890.757 to 0.817
  • Human texts kept2023 validation0.9330.925 to 0.939
  • Human texts keptclean test0.9450.942 to 0.948

Logistic regression, domain 2

copied class: human
Recall per class, 2023 validation vs clean test (Wilson 95%)
  • Machine texts caught2023 validation0.9100.833 to 0.954
  • Machine texts caughtclean test0.9630.895 to 0.987
  • Human texts kept2023 validation, copied0.9860.924 to 0.998
  • Human texts keptclean test0.4500.258 to 0.658
2023 validation accuracy and F1 with 95% intervals
ModelRowsAccuracy, as printed (Wilson 95%)Human F1, as printed (bootstrap 95%)
LR 110,0000.9658 (0.962 to 0.969)0.9643 (0.960 to 0.968)
LR 21600.9437 (0.897 to 0.970)0.9396 (0.895 to 0.975)
LR 1+210,0000.5005 (0.491 to 0.510)0.5002 (0.488 to 0.512)
SGD 110,0000.9330 (0.928 to 0.938)0.9315 (0.926 to 0.936)
SGD 21600.5312 (0.454 to 0.607)0.5562 (0.462 to 0.640)
SGD 1+210,0000.5019 (0.492 to 0.512)0.4493 (0.436 to 0.462)

The SGD 2 row is the report's, which scored the domain-1 SGD model on domain-2 data. The 1+2 rows come from models trained on mismatched features and labels. Both are explained in the revival audit. The printed SGD 2 precision slip does not affect these two columns.

Downloads

Take the numbers with you

The aggregates behind this page, exactly as the script wrote them. They contain no text, token sequence or per-item prediction. To reproduce them, run uv run scripts/evaluate_heldout.py from the repository root (about 40 minutes on a 10-core laptop, most of it in the converged sensitivity fits).