Skip to content

Exhibit D · experiment journal

Three weeks, five model families, eight submissions.

A reconstruction of how Team 27 worked through the project in April 2023, from the meeting minutes, the commit history, the notebooks and the report. All numbers below come from the original work; the last section adds what re-running it in 2026 revealed.

Timeline

From data drop to hand-in

  1. Data released

    • Four training files over two domains plus 1,000 unlabelled test items; every word already replaced by an integer id from 0 to 4,999.

    source: file dates inside coursework/data/sml.p1.data.tar.gz

  2. Meeting 1: pick three model families

    • Each member championed one approach: support vector machine, logistic regression, convolutional network.
    • Use both prompt and text as features.
    • The four files are extremely imbalanced (about 120k / 3.5k / 100 / 400), so try SMOTE over-sampling and under-sampling.

    source: MeetingMinutes 1.docx

  3. Meeting 2: features blow up

    • Word embeddings expanded the 250 MB data set towards 90 GB, far beyond a 16 GB laptop; reduce to 20 dimensions.
    • Bag-of-words gives 5,000-wide vectors that are over 90% zeros (5 to 8 GB) but fits in memory.
    • Plain SVM gave no clear decision boundary, so switch to SGDClassifier, which also handles class balance.
    • Same day: MPI bag-of-words preprocessing and PyTorch data sets committed.

    source: MeetingMinutes 2.docx, git history

  4. Meeting 3: first leaderboard numbers

    • SGD with built-in class balancing scores 0.71 on Kaggle and becomes the current best.
    • Logistic regression on SMOTE data: 0.666. Logistic regression on RUS data: poor.
    • ShuffleNet reaches 0.618 and gets worse with more training data; CNNs dropped from the final pipeline.
    • Choose SMOTE over RUS: overfitting is judged less harmful than underfitting.

    source: MeetingMinutes 3.docx

  5. Meeting 4: settle the submission

    • SGD improves to 0.72 and is chosen as the final outcome model.
    • Build one evaluation table: accuracy, precision, recall, F1 and Kaggle score per model and domain.
    • Notebooks cleaned up with run instructions; CNN work documented but left out of the report.

    source: MeetingMinutes 4.docx

  6. Report handed in

    • Three-page report. Its conclusion recommends logistic regression for better generalisation on scarce domain-2 data, even though SGD had the higher Kaggle score (its abstract still names SGD).

    source: coursework/doc/COMP90051_Statistical_Machine_Learning_Project.pdf

Approaches

What was tried, and what survived

Bag-of-words + logistic regression

final

LogisticRegressionCV (l2, lbfgs, 50 iterations, class_weight=balanced), one model per domain, fitted on 80% of a 50,000-row SMOTE-balanced sample (all 800 rows for domain 2). The report's recommended model.

notebooks/final/LR Classifier.ipynb

Bag-of-words + SGD (hinge)

compared

SGDClassifier with hinge loss and an L1 penalty: a linear SVM trained by stochastic gradient descent. Best single leaderboard score (0.720) and the minutes' pick for submission.

notebooks/final/SGD Classifier.ipynb

CNNs on bag-of-words 'images'

explored

The 5,000 counts folded into a small 2-D grid and fed to LeNet, a custom CNN, ImageNet-pretrained ResNet18/152 and ShuffleNetV2 x2.0. Near-perfect internal test accuracy, 0.618 on Kaggle.

notebooks/final/ShuffleNet.ipynb, notebooks/unused/*, plots/*

Word embeddings and TF-IDF

dropped

A 20-dimensional embedding (all the hardware allowed) and TF-IDF features into a three-layer CNN: fast, but compressed away too much and recall suffered.

notebooks/unused/embedding*.ipynb, Create TD_IDF Dataset*.ipynb

Support vector machine

dropped

No clear boundary on the sparse counts and slow to train; replaced by the SGD classifier.

MeetingMinutes 2.docx, notebooks/unused/Sklearn LR + SVM.ipynb

Results

The report's table, re-derived

Validation scores on each model's own 20% hold-out (human = positive class) and the Kaggle score from the report. The confusion matrices were recomputed by re-running the notebooks' evaluation cells with the original pickles; every figure matches the report to four decimals except one transcription slip.

Validation metrics and Kaggle scores of the six models
ModelAccuracyPrecisionRecallF1KaggleConfusion (tn fp fn tp)
LR 10.96580.99810.93280.96430.7025035 9 333 4623
LR 20.94370.89740.98590.93960.53481 8 1 70
LR 1+20.50050.50000.50030.50020.4822506 2499 2496 2499
SGD 10.93300.94480.91850.93150.7204778 266 404 4552
SGD 20.53120.4796*0.66200.55620.53038 51 24 47
SGD 1+20.50190.50170.40680.44930.4862987 2018 2963 2032

* The report prints SGD 2 precision as 0.4696. Its own recall (0.6620) and F1 (0.5562) imply 0.4796, which is what the re-run produces. Kaggle scores were taken from single-model submissions over all 1,000 test items.

The CNN detour

Training curves that looked too good

Curves recovered from the original Plotly exports and notebook logs. The ResNets and ShuffleNets pass 99% on their own test splits, yet ShuffleNet scored 0.618 on the leaderboard. None of those splits was a clean hold-out: the final ShuffleNet notebook drew its test rows from the same index-SMOTE pool as its training rows (exact copies on both sides, see the index-SMOTE finding), and the ResNet loops re-split one pool every epoch, so later epochs were tested on rows trained on earlier.

ShuffleNetV2 x2.0

30 epochs · final test acc 0.9967

Final notebook run: 5,000-dim BoW folded to a 25x40 'image', SMOTE-balanced (490,336 rows), a fresh 10% sample per epoch, Adam + StepLR, 30 epochs.

0.800.850.900.95102030epoch
Accuracytraintest
0.010.101.010100102030epoch
Loss (log scale)traintest

coursework/notebooks/final/ShuffleNet.ipynb

ShuffleNetV2 x2.0 (data set v4)

25 epochs · final test acc 0.9965

Earlier ShuffleNet run on the v4 data set, 25 epochs.

0.600.700.800.901.0510152025epoch
Accuracytraintest
1e-41e-30.010.10510152025epoch
Loss (log scale)traintest

coursework/plots/ShuffleNetV2X2 *.html

ResNet18

20 epochs · final test acc 0.9929

ImageNet-pretrained ResNet18 with a 1-channel stem, 20 epochs.

0.900.920.940.960.985101520epoch
Accuracytraintest
0.105101520epoch
Loss (log scale)traintest

coursework/plots/ResNet18 *.html

ResNet152

50 epochs · final test acc 0.9981

ImageNet-pretrained ResNet152 with a 1-channel stem, 50 epochs.

0.850.900.951020304050epoch
Accuracytraintest
0.010.101.01020304050epoch
Loss (log scale)traintest

coursework/plots/ResNet152 *.html

Custom CNN

200 epochs · final test acc 0.8700

Hybrid network: a dense 5,000-to-1,024 layer feeding three small conv layers and an MLP head, on MPI-built BoW vectors. Each epoch re-draws a human subsample the size of the machine set (on-the-fly undersampling); 200 epochs.

0.700.800.9050100150200epoch
Accuracytraintest
0.1050100150200epoch
Loss (log scale)traintest

coursework/notebooks/unused/CustomCNN.ipynb

Kaggle

Eight submission files, and how much they agree

The result notebook wrote eight CSVs. Re-running it with the original pickles reproduces all 8,000 predictions exactly. The matrix shows how many of the 1,000 test items two files label the same way.

Submission files and how they were produced
FileRecipeHuman
result1SGD 1 on all 1,000 items553
result2SGD 2 on all items219
result3SGD 1+2 on all items388
result4LR 1 on all items416
result5LR 2 on all items121
result6LR 1+2 on all items494
result7majority vote of the six (ties go to SGD 1)338
result8LR 1 on items 0-599, LR 2 on items 600-999437
r1r1r2r2r3r3r4r4r5r5r6r6r7r7r8r810005164677075224837716905161000545565834485731620467545100051656775266847170756551610005415067268655228345675411000489725676483485752506489100065848377173166872672565810007216906204718656764837211000

Pairwise agreement out of 1,000. r4 (LR 1) and r8 (LR 1 + LR 2 routed) agree on 865 items. The 1+2 models (r3, r6) agree with most other files on only about half the items, coin-flip territory, but with each other on 752: both learnt the same noise.

Revival audit, 2026

What re-running the notebooks revealed

Porting the models meant re-executing every step with the original library versions (scikit-learn 1.2.2, imbalanced-learn 0.10.1, pandas 2.0.0) and checking each number. Everything reproduced. Along the way four things turned up that the 2023 team could not easily have seen. The original code is unchanged in the repository; this is a reading, not a rewrite.

1. SMOTE on the row index duplicated rows

Minority rows were copied, not synthesised, and copies landed on both sides of the validation split. This inflated validation scores for the minority class. Details on the imbalance page.

2. The domain 1+2 models never saw domain-2 features

In the SGD notebook, domain3_data was built from domain1_df.data (a copy-paste slip), and the LR notebook re-read that file. Labels and features came from different rows, so both 1+2 models learnt noise: about 0.50 by design. On correctly paired features they score 0.4986 (LR) and 0.4554 (SGD).

3. The SGD domain-2 row scored the domain-1 model

The evaluation cell reads y2_pred = SGD1.predict(X2_val). The report's 0.5312 is the domain-1 SGD model on domain-2 data. The actual domain-2 SGD model scores 0.9000 accuracy and 0.8974 F1 on the same 160 rows, close to LR 2. The report's argument that LR generalises better on domain 2 rested on this comparison.

4. One transcription slip

SGD 2 precision is printed as 0.4696 in the report and README; the confusion matrix gives 0.4796. Nothing else differs.

What the models score once the leak is removed, with intervals, calibration and paired tests, is on the held-out re-evaluation page. The reasoning behind the 2023 choices is in the decision records.