Exhibit C · class imbalance
Ninety-seven humans for every three machines.
The training data came as four files. One of them, human text from domain 1, holds 96.8% of all instances. Any classifier trained naively on it learns that “human” is almost always right. The team compared two standard cures: synthetic over-sampling (SMOTE) and random under-sampling (RUS).
The evidence
Four files, three orders of magnitude
Counts per file (log scale). Domain 2 is not just small, it is imbalanced the other way round: there are four machine texts for every human one. The test set has 1,000 unlabelled items, the first 600 from domain 1 and the last 400 from domain 2.
set1_human
domain 1 · 96.84%
set1_machine
domain 1 · 2.76%
set2_machine
domain 2 · 0.32%
set2_human
domain 2 · 0.08%
Mean length, set1 human
380
tokens (prompt + text)
Mean length, set1 machine
183
machine texts are about half as long
Machine generators
5
in domain 1, 700 texts each
Machine generators
4
in domain 2, 100 texts each
A shortcut the models could take
In domain 1 the middle 80% of human texts run 214 to 546 tokens, machine texts 117 to 251. Length alone already separates much of the data, and bag-of-words counts carry length implicitly (more tokens, bigger counts). The test set had no machine_id, so the team ignored that column.
Two cures
What re-balancing did to the data set sizes
With imbalanced-learn's default strategy, SMOTE grows every class to the size of the largest one and RUS shrinks every class to the size of the smallest. Domain “1+2” balanced all four (domain, class) groups at once.
| Training set | Original rows | After SMOTE | After RUS |
|---|---|---|---|
| Domain 1 | 126,084 | 245,168 | 7,000 |
| Domain 2 | 500 | 800 | 200 |
| Domain 1+2 | 126,584 | 490,336 | 400 |
The team chose SMOTE: RUS would have trained the domain-1+2 model on 400 rows. Because SMOTE data sets reached half a million rows, the linear models were then fitted on a random 50,000-row sample (80/20 train/validation split, seed 90051).
Sandbox
Try the cures on a toy problem
Two clouds of points stand in for the two classes; a logistic regression (the same objective scikit-learn optimises) is refitted every time you change something. Switch between the methods and watch the boundary and the machine recall.
Re-balancing method
Domain 1 was 2.8% machine; domain 2 was 80% machine.
SMOTE in feature space
New machine points are placed on the segment between a machine point and one of its k nearest machine neighbours (step u ~ U[0,1)). This is what SMOTE is meant to do.
- Training rows
- 960
- 480 h / 480 m
- Test accuracy
- 89.7%
- ▲ 21.0 pts
- Machine recall
- 96.3%
- ▲ 58.0 pts
- Human recall
- 83.0%
- ▼ 16.0 pts
Scored on a separate balanced test set of 600 points from the same two clouds. Deltas are against no re-balancing.
Revival finding
The SMOTE in the notebooks made copies, not new points
To save memory the notebooks passed SMOTE a single column: the row index. imbalanced-learn interpolates between neighbouring indices and casts the result back to an integer, so each “synthetic” row is an existing minority row picked again. Re-running the notebook logic with the original library versions gives:
Domain 1 machine
3,500 → 122,584
distinct texts → rows. Median 35 copies of each, up to 58.
Domain 2 human
100 → 400
Median 4 copies, up to 10.
Domain 1+2, set2 human
100 → 122,584
Median 1,229 copies of each text, up to 1,738.
Why validation looked better than Kaggle
After the 80/20 split, 5,031 of the 5,044 domain-1 machine rows in the validation set had an identical copy in the training set (and 70 of 71 domain-2 human rows). For the minority class, validation measured recall of memorised examples, not generalisation. That gap is part of why 0.97 validation accuracy became about 0.70 on the leaderboard.
What it means for the original conclusion
The report's worry, that SMOTE “may cause overfitting” through duplicated minority labels, was right, and for a stronger reason than the team knew: the duplication was exact. In effect the team compared random over-sampling (with replacement) against RUS. Both are shown in the sandbox above.