Exhibit D · experiment journal
Three weeks, five model families, eight submissions.
A reconstruction of how Team 27 worked through the project in April 2023, from the meeting minutes, the commit history, the notebooks and the report. All numbers below come from the original work; the last section adds what re-running it in 2026 revealed.
Timeline
From data drop to hand-in
Data released
- Four training files over two domains plus 1,000 unlabelled test items; every word already replaced by an integer id from 0 to 4,999.
source: file dates inside coursework/data/sml.p1.data.tar.gz
Meeting 1: pick three model families
- Each member championed one approach: support vector machine, logistic regression, convolutional network.
- Use both prompt and text as features.
- The four files are extremely imbalanced (about 120k / 3.5k / 100 / 400), so try SMOTE over-sampling and under-sampling.
source: MeetingMinutes 1.docx
Meeting 2: features blow up
- Word embeddings expanded the 250 MB data set towards 90 GB, far beyond a 16 GB laptop; reduce to 20 dimensions.
- Bag-of-words gives 5,000-wide vectors that are over 90% zeros (5 to 8 GB) but fits in memory.
- Plain SVM gave no clear decision boundary, so switch to SGDClassifier, which also handles class balance.
- Same day: MPI bag-of-words preprocessing and PyTorch data sets committed.
source: MeetingMinutes 2.docx, git history
Meeting 3: first leaderboard numbers
- SGD with built-in class balancing scores 0.71 on Kaggle and becomes the current best.
- Logistic regression on SMOTE data: 0.666. Logistic regression on RUS data: poor.
- ShuffleNet reaches 0.618 and gets worse with more training data; CNNs dropped from the final pipeline.
- Choose SMOTE over RUS: overfitting is judged less harmful than underfitting.
source: MeetingMinutes 3.docx
Meeting 4: settle the submission
- SGD improves to 0.72 and is chosen as the final outcome model.
- Build one evaluation table: accuracy, precision, recall, F1 and Kaggle score per model and domain.
- Notebooks cleaned up with run instructions; CNN work documented but left out of the report.
source: MeetingMinutes 4.docx
Report handed in
- Three-page report. Its conclusion recommends logistic regression for better generalisation on scarce domain-2 data, even though SGD had the higher Kaggle score (its abstract still names SGD).
source: coursework/doc/COMP90051_Statistical_Machine_Learning_Project.pdf
Approaches
What was tried, and what survived
Bag-of-words + logistic regression
finalLogisticRegressionCV (l2, lbfgs, 50 iterations, class_weight=balanced), one model per domain, fitted on 80% of a 50,000-row SMOTE-balanced sample (all 800 rows for domain 2). The report's recommended model.
notebooks/final/LR Classifier.ipynb
Bag-of-words + SGD (hinge)
comparedSGDClassifier with hinge loss and an L1 penalty: a linear SVM trained by stochastic gradient descent. Best single leaderboard score (0.720) and the minutes' pick for submission.
notebooks/final/SGD Classifier.ipynb
CNNs on bag-of-words 'images'
exploredThe 5,000 counts folded into a small 2-D grid and fed to LeNet, a custom CNN, ImageNet-pretrained ResNet18/152 and ShuffleNetV2 x2.0. Near-perfect internal test accuracy, 0.618 on Kaggle.
notebooks/final/ShuffleNet.ipynb, notebooks/unused/*, plots/*
Word embeddings and TF-IDF
droppedA 20-dimensional embedding (all the hardware allowed) and TF-IDF features into a three-layer CNN: fast, but compressed away too much and recall suffered.
notebooks/unused/embedding*.ipynb, Create TD_IDF Dataset*.ipynb
Support vector machine
droppedNo clear boundary on the sparse counts and slow to train; replaced by the SGD classifier.
MeetingMinutes 2.docx, notebooks/unused/Sklearn LR + SVM.ipynb
Results
The report's table, re-derived
Validation scores on each model's own 20% hold-out (human = positive class) and the Kaggle score from the report. The confusion matrices were recomputed by re-running the notebooks' evaluation cells with the original pickles; every figure matches the report to four decimals except one transcription slip.
| Model | Accuracy | Precision | Recall | F1 | Kaggle | Confusion (tn fp fn tp) |
|---|---|---|---|---|---|---|
| LR 1 | 0.9658 | 0.9981 | 0.9328 | 0.9643 | 0.702 | 5035 9 333 4623 |
| LR 2 | 0.9437 | 0.8974 | 0.9859 | 0.9396 | 0.534 | 81 8 1 70 |
| LR 1+2 | 0.5005 | 0.5000 | 0.5003 | 0.5002 | 0.482 | 2506 2499 2496 2499 |
| SGD 1 | 0.9330 | 0.9448 | 0.9185 | 0.9315 | 0.720 | 4778 266 404 4552 |
| SGD 2 | 0.5312 | 0.4796* | 0.6620 | 0.5562 | 0.530 | 38 51 24 47 |
| SGD 1+2 | 0.5019 | 0.5017 | 0.4068 | 0.4493 | 0.486 | 2987 2018 2963 2032 |
* The report prints SGD 2 precision as 0.4696. Its own recall (0.6620) and F1 (0.5562) imply 0.4796, which is what the re-run produces. Kaggle scores were taken from single-model submissions over all 1,000 test items.
The CNN detour
Training curves that looked too good
Curves recovered from the original Plotly exports and notebook logs. The ResNets and ShuffleNets pass 99% on their own test splits, yet ShuffleNet scored 0.618 on the leaderboard. None of those splits was a clean hold-out: the final ShuffleNet notebook drew its test rows from the same index-SMOTE pool as its training rows (exact copies on both sides, see the index-SMOTE finding), and the ResNet loops re-split one pool every epoch, so later epochs were tested on rows trained on earlier.
ShuffleNetV2 x2.0
Final notebook run: 5,000-dim BoW folded to a 25x40 'image', SMOTE-balanced (490,336 rows), a fresh 10% sample per epoch, Adam + StepLR, 30 epochs.
coursework/notebooks/final/ShuffleNet.ipynb
ShuffleNetV2 x2.0 (data set v4)
Earlier ShuffleNet run on the v4 data set, 25 epochs.
coursework/plots/ShuffleNetV2X2 *.html
ResNet18
ImageNet-pretrained ResNet18 with a 1-channel stem, 20 epochs.
coursework/plots/ResNet18 *.html
ResNet152
ImageNet-pretrained ResNet152 with a 1-channel stem, 50 epochs.
coursework/plots/ResNet152 *.html
Custom CNN
Hybrid network: a dense 5,000-to-1,024 layer feeding three small conv layers and an MLP head, on MPI-built BoW vectors. Each epoch re-draws a human subsample the size of the machine set (on-the-fly undersampling); 200 epochs.
coursework/notebooks/unused/CustomCNN.ipynb
Kaggle
Eight submission files, and how much they agree
The result notebook wrote eight CSVs. Re-running it with the original pickles reproduces all 8,000 predictions exactly. The matrix shows how many of the 1,000 test items two files label the same way.
| File | Recipe | Human |
|---|---|---|
| result1 | SGD 1 on all 1,000 items | 553 |
| result2 | SGD 2 on all items | 219 |
| result3 | SGD 1+2 on all items | 388 |
| result4 | LR 1 on all items | 416 |
| result5 | LR 2 on all items | 121 |
| result6 | LR 1+2 on all items | 494 |
| result7 | majority vote of the six (ties go to SGD 1) | 338 |
| result8 | LR 1 on items 0-599, LR 2 on items 600-999 | 437 |
Pairwise agreement out of 1,000. r4 (LR 1) and r8 (LR 1 + LR 2 routed) agree on 865 items. The 1+2 models (r3, r6) agree with most other files on only about half the items, coin-flip territory, but with each other on 752: both learnt the same noise.
Revival audit, 2026
What re-running the notebooks revealed
Porting the models meant re-executing every step with the original library versions (scikit-learn 1.2.2, imbalanced-learn 0.10.1, pandas 2.0.0) and checking each number. Everything reproduced. Along the way four things turned up that the 2023 team could not easily have seen. The original code is unchanged in the repository; this is a reading, not a rewrite.
1. SMOTE on the row index duplicated rows
Minority rows were copied, not synthesised, and copies landed on both sides of the validation split. This inflated validation scores for the minority class. Details on the imbalance page.
2. The domain 1+2 models never saw domain-2 features
In the SGD notebook, domain3_data was built from domain1_df.data (a copy-paste slip), and the LR notebook re-read that file. Labels and features came from different rows, so both 1+2 models learnt noise: about 0.50 by design. On correctly paired features they score 0.4986 (LR) and 0.4554 (SGD).
3. The SGD domain-2 row scored the domain-1 model
The evaluation cell reads y2_pred = SGD1.predict(X2_val). The report's 0.5312 is the domain-1 SGD model on domain-2 data. The actual domain-2 SGD model scores 0.9000 accuracy and 0.8974 F1 on the same 160 rows, close to LR 2. The report's argument that LR generalises better on domain 2 rested on this comparison.
4. One transcription slip
SGD 2 precision is printed as 0.4696 in the report and README; the confusion matrix gives 0.4796. Nothing else differs.
What the models score once the leak is removed, with intervals, calibration and paired tests, is on the held-out re-evaluation page. The reasoning behind the 2023 choices is in the decision records.