Skip to content
All decision records

Decision record · DR-003

SMOTE over-sampling to re-balance the classes

Status
Accepted
Date
2023-04 (recorded 2026-10)
Applies to
coursework/notebooks/final/SMOTI data balancing.ipynb, /imbalance, /evaluation#rebalancing

Decision in one line

We re-balanced every training set with SMOTE over-sampling rather than random under-sampling, then split it 80/20 for validation. As implemented, SMOTE copied minority texts instead of synthesising new ones, the copies landed on both sides of the split, and on clean data the re-balancing made the detector worse than no re-balancing.

Context

Domain 1 holds 97 human texts for every 3 machine texts. Domain 2 is small and imbalanced the other way, 100 human texts against 400 machine texts. A classifier trained naively on domain 1 can reach 97% accuracy by always answering "human". At the third meeting (22 April 2023) logistic regression on randomly under-sampled data performed poorly on Kaggle.

Decision

Use imbalanced-learn's SMOTE to over-sample the minority class of each domain up to the size of the majority class (and all four domain-class groups for the pooled data). Keep balanced class weights in both models as well. The minutes record the reasoning as "the benefit of overfitting is greater than underfit".

Options considered

  1. No re-balancing. Rejected at the time because accuracy rewards the majority class.
  2. Class weights only. Both models already supported class_weight="balanced", and SGD used it on 22 April for its 0.71 Kaggle score.
  3. SMOTE over-sampling (chosen). Keeps every original text and adds synthetic minority texts.
  4. Random under-sampling. Leaves 7,000 training texts in domain 1 and 200 in domain 2, which the report judged too few.
  5. Fit on the natural balance, then move the decision threshold. Not considered in 2023.

Why

Under-sampling throws away 97% of the domain-1 human texts and leaves domain 2 with 200 texts in total. Over-sampling keeps all the information, and over-fitting looked like the smaller risk.

What happened

  • SMOTE ran on the row index. The notebooks passed SMOTE a one-column frame holding each row's integer index, not the 5,000 counts. Interpolating between two neighbouring row ids and rounding down gives another row id, so every "synthetic" text was an exact copy of an existing minority text.

  • The split came after the copying. In domain 1, 5,031 of the 5,044 machine texts in the validation part had an identical copy in the training part. The minority class was being scored on texts the model had seen, which inflated its recall: 0.998 for machine texts in domain 1 in 2023, against 0.789 on clean held-out texts.

  • Clean comparison in 2026 (DR-004). Logistic regression, the same 25,217 held-out domain-1 texts, re-balancing applied to the training part only:

    StrategyMacro-F1AUCMachine texts caughtECE
    No re-balancing, no class weights0.809 (0.791 to 0.825)0.975 (0.970 to 0.978)356 of 7000.002 (0.002 to 0.004)
    Class weights only0.739 (0.726 to 0.752)0.969 (0.964 to 0.974)544 of 7000.037 (0.035 to 0.039)
    Index-SMOTE copies + class weights (2023)0.697 (0.685 to 0.709)0.960 (0.954 to 0.965)552 of 7000.054 (0.051 to 0.056)
    SMOTE on the features + class weights0.700 (0.687 to 0.712)0.959 (0.953 to 0.964)520 of 7000.049 (0.046 to 0.051)
    Random under-sampling + class weights0.610 (0.601 to 0.619)0.965 (0.960 to 0.970)654 of 7000.141 (0.138 to 0.144)
  • Reading the table. Against no re-balancing, the 2023 recipe loses 0.111 macro-F1 (paired 95% interval −0.128 to −0.095) and its probabilities are far less calibrated. It catches 196 more machine texts, and it pays for them by flagging 1,343 human texts as machine-written instead of 82. Proper SMOTE on the features was no better than the copies.

  • Checked against the fitting. As fitted in 2023, cross-validation chose C = 10,000 for the recipe and 0.046 without re-balancing, after 50 iterations that stop most fits early, so the table partly compares regularisation regimes. Re-fitted with the solver run to convergence and C chosen on folds grouped by source text, no re-balancing is still best on macro-F1, AUC and calibration, and the recipe's gap grows to 0.163 macro-F1 (paired 95% interval −0.181 to −0.145).

  • Domain 2. Every strategy except under-sampling lands between 0.740 and 0.764 macro-F1, and with 100 test texts the intervals overlap almost completely. Under-sampling, with 160 training texts, falls to 0.587 (0.480 to 0.687). With converged fits on grouped folds the recipe falls behind no re-balancing by 0.168 macro-F1 (−0.277 to −0.057).

What I'd change

  • Split first, and re-balance only inside the training folds, for example with an imbalanced-learn pipeline and cross-validation folds grouped by source text.
  • Fit on the natural class balance and choose the decision threshold for the costs of the two errors. Here that means deciding how many human texts it is acceptable to flag in order to catch one more machine text, instead of letting re-sampling make that choice silently.
  • Report macro-F1, AUC and per-class recall with intervals instead of accuracy on balanced validation data.