Skip to content
All decision records

Decision record · DR-004

Re-fit the 2023 recipes on a split made before re-balancing

Status
Accepted
Date
2026-10
Applies to
scripts/evaluate_heldout.py, /evaluation, docs/model-card.md

Decision in one line

To measure how well the 2023 approach really works, I re-fit the 2023 recipes unchanged on an 80/20 split made before any re-balancing, score the untouched 20% with bootstrap intervals, and keep the 2023 pickles and printed numbers exactly as they were.

Context

The 2023 validation numbers were inflated because the split came after SMOTE had copied the minority texts (DR-003). An honest evaluation needs test texts the models never saw. The Kaggle test labels were never released, so the only labelled data is the four training files, and the 2023 domain-1 models were trained on 40,000 rows that include copies of almost every domain-1 machine text.

Decision

scripts/evaluate_heldout.py runs the 2023 pipeline in a different order.

  1. Split each domain's original texts 80/20, stratified by class and generator, with seed 90051.
  2. Re-balance the training part only, with the notebooks' own index-SMOTE call, and cap domain 1 at the notebooks' 40,000 training rows.
  3. Fit LogisticRegressionCV and SGDClassifier with the 2023 settings.
  4. Score the test part once: accuracy, macro-F1 and ROC AUC with 95% percentile bootstrap intervals (10,000 resamples, seed 90051, the same resamples for every detector so that differences are paired), McNemar tests, calibration, domain shift, per-generator recall and two trivial baselines.
  5. Repeat the split and the fits for nine more seeds, on ten splits that keep each prompt on one side, and five times with one generator left out (with and without that generator's test prompts in training).
  6. Record the fitting diagnostics of every logistic regression, and re-fit the recipe and each re-balancing strategy with the solver run to convergence and C chosen on folds grouped by source text, as a sensitivity check that leaves the headline numbers as run.

The script writes aggregates only. The site shows them on /evaluation, and the 2023 numbers stay on /journal as printed.

Options considered

  1. Re-score the 2023 pickles on their own validation rows. This reproduces the leak.
  2. Re-score the pickles on rows they never saw. The domain-1 models were trained on about 20,000 copies drawn from the 3,500 machine texts, so only around a dozen machine texts can be expected to have stayed out of training. That is too few to score anything.
  3. Re-fit the same recipes on a split made first (chosen). It measures the recipe rather than the exact 2023 weights, but every test text is unseen.
  4. Nested cross-validation over the whole pipeline. Uses every text for testing once, but costs about five times as long for the same conclusions. Left for later.
  5. Fix the recipe as well. Dropping the copies or the 40,000-row cap would have answered a different question. The fixes are evaluated separately, as alternatives (DR-003).

Why

  • Every test text is unseen, which is the property the 2023 evaluation lacked.
  • Keeping the recipe unchanged, including its questionable parts, means the numbers describe what the team built.
  • Bootstrap intervals need no assumption about the shape of the metric's distribution, and common resamples make the paired comparisons paired.
  • Ten splits show the variation that one split's interval cannot, since the bootstrap treats the fitted model as fixed.

What happened

  • Domain 1, logistic regression: macro-F1 0.697 (0.685 to 0.709), AUC 0.960 (0.954 to 0.965) on 25,217 texts. SGD: macro-F1 0.614 (0.603 to 0.625), AUC 0.902 (0.893 to 0.911). Logistic regression is ahead by +0.083 (+0.074 to +0.093) macro-F1 on the same texts and was ahead in all ten splits.
  • Domain 2 has 100 test texts, 20 of them human. Logistic regression scores macro-F1 0.740 (0.610 to 0.850) and SGD 0.697 (0.577 to 0.805). The 15 texts where they disagree give no evidence either way (McNemar exact p = 0.30).
  • Calibration, domain 1: the probabilities of both detectors assume a 50:50 prior because of the re-balancing. Shifting the log-odds to the real 97.2% human share brings the ECE of Platt-scaled SGD from 0.289 to 0.007 and of logistic regression from 0.054 to 0.023.
  • The fits stop early. Review of this record found that the 2023 settings stop the logistic regression at 50 iterations: for domain 1, all 50 cross-validation fits and the final fit hit the limit, and cross-validation picks C = 10,000, the top of its grid. The copies of one text also sit in different inner folds, so inner-fold accuracy keeps rising with C. The warnings had been silenced; the script now records them for every fit. With the solver run to convergence and C chosen on folds grouped by source text, the domain-1 recipe picks C = 0.006 and scores macro-F1 0.644 (0.634 to 0.654) and AUC 0.971 (0.966 to 0.975). In domain 2 the grouped folds push C to the bottom of the grid and macro-F1 falls to 0.584 (0.479 to 0.687). The headline numbers stay the recipe as run, with this sensitivity next to them on /evaluation#fitting.
  • Splits. The main domain-2 split is the second most favourable of the ten for logistic regression (macro-F1 0.740 against a ten-split mean of 0.657). On ten prompt-grouped splits the domain-2 logistic regression averages macro-F1 0.558 ± 0.039, so part of its score on random splits comes from recognising machine prompts.
  • The leave-one-generator-out fits and the prompt-cluster intervals are on /evaluation, and the model card summarises them.
  • The run takes about 40 minutes on a 10-core laptop and is deterministic. Most of the time goes into the converged fits on the full domain-1 training part, two of which still stop some cross-validation fits at 5,000 iterations.

What I'd change

  • Nested, repeated cross-validation for domain 2, where one split leaves only 20 human test texts.
  • Bias-corrected (BCa) intervals as a check on the percentile intervals where the samples are small, and the prompt-cluster bootstrap for every row rather than the headline ones.
  • Fit to convergence and choose C on grouped folds from the start, instead of adding it as a sensitivity check after review.
  • Precision-recall curves for the machine class, which is what a user of a detector would care about at a 97:3 prior.