Skip to content

Exhibit C · class imbalance

Ninety-seven humans for every three machines.

The training data came as four files. One of them, human text from domain 1, holds 96.8% of all instances. Any classifier trained naively on it learns that “human” is almost always right. The team compared two standard cures: synthetic over-sampling (SMOTE) and random under-sampling (RUS).

The evidence

Four files, three orders of magnitude

Counts per file (log scale). Domain 2 is not just small, it is imbalanced the other way round: there are four machine texts for every human one. The test set has 1,000 unlabelled items, the first 600 from domain 1 and the last 400 from domain 2.

set1_human

domain 1 · 96.84%

122,584

set1_machine

domain 1 · 2.76%

3,500

set2_machine

domain 2 · 0.32%

400

set2_human

domain 2 · 0.08%

100
1101001k10k100k

Mean length, set1 human

380

tokens (prompt + text)

Mean length, set1 machine

183

machine texts are about half as long

Machine generators

5

in domain 1, 700 texts each

Machine generators

4

in domain 2, 100 texts each

A shortcut the models could take

In domain 1 the middle 80% of human texts run 214 to 546 tokens, machine texts 117 to 251. Length alone already separates much of the data, and bag-of-words counts carry length implicitly (more tokens, bigger counts). The test set had no machine_id, so the team ignored that column.

Two cures

What re-balancing did to the data set sizes

With imbalanced-learn's default strategy, SMOTE grows every class to the size of the largest one and RUS shrinks every class to the size of the smallest. Domain “1+2” balanced all four (domain, class) groups at once.

Data set sizes before and after re-balancing
Training setOriginal rowsAfter SMOTEAfter RUS
Domain 1126,084245,1687,000
Domain 2500800200
Domain 1+2126,584490,336400

The team chose SMOTE: RUS would have trained the domain-1+2 model on 400 rows. Because SMOTE data sets reached half a million rows, the linear models were then fitted on a random 50,000-row sample (80/20 train/validation split, seed 90051).

Sandbox

Try the cures on a toy problem

Two clouds of points stand in for the two classes; a logistic regression (the same objective scikit-learn optimises) is refitted every time you change something. Switch between the methods and watch the boundary and the machine recall.

human (480)machine (20)◇ synthetic machine (460) decision boundary

Re-balancing method

Machine share of training data4.0%

Domain 1 was 2.8% machine; domain 2 was 80% machine.

Class separation2.4σ
k neighbours5

SMOTE in feature space

New machine points are placed on the segment between a machine point and one of its k nearest machine neighbours (step u ~ U[0,1)). This is what SMOTE is meant to do.

Training rows
960
480 h / 480 m
Test accuracy
89.7%
▲ 21.0 pts
Machine recall
96.3%
▲ 58.0 pts
Human recall
83.0%
▼ 16.0 pts

Scored on a separate balanced test set of 600 points from the same two clouds. Deltas are against no re-balancing.

Revival finding

The SMOTE in the notebooks made copies, not new points

To save memory the notebooks passed SMOTE a single column: the row index. imbalanced-learn interpolates between neighbouring indices and casts the result back to an integer, so each “synthetic” row is an existing minority row picked again. Re-running the notebook logic with the original library versions gives:

Domain 1 machine

SMOTI data balancing.ipynb

3,500 → 122,584

distinct texts → rows. Median 35 copies of each, up to 58.

Domain 2 human

SMOTI data balancing.ipynb

100 → 400

Median 4 copies, up to 10.

Domain 1+2, set2 human

four groups balanced

100 → 122,584

Median 1,229 copies of each text, up to 1,738.

Why validation looked better than Kaggle

After the 80/20 split, 5,031 of the 5,044 domain-1 machine rows in the validation set had an identical copy in the training set (and 70 of 71 domain-2 human rows). For the minority class, validation measured recall of memorised examples, not generalisation. That gap is part of why 0.97 validation accuracy became about 0.70 on the leaderboard.

What it means for the original conclusion

The report's worry, that SMOTE “may cause overfitting” through duplicated minority labels, was right, and for a stronger reason than the team knew: the duplication was exact. In effect the team compared random over-sampling (with replacement) against RUS. Both are shown in the sandbox above.