COMP90051 · Project 1 · 2023 Semester 1 · Team 27
Human or machine?
In 2023 our team built detectors that decide whether a text was written by a person or generated by a model, for a Kaggle competition in Statistical Machine Learning at the University of Melbourne. This lab puts the original detectors back to work in your browser.
New here? Watch the guided tour: three captioned walkthroughs of the main workflows.
Specimen 90051 · domain 1
260 tokens
Drawn entirely from the machine token profile. Amber chips push towards human, blue towards machine; the detector adds up one learnt weight per token id and squashes the sum into a probability.
The brief
Spot the machine, without ever seeing a word
The task, in our words: given short texts from two sources (“domains”), each paired with the prompt that produced it, predict whether a human or a machine wrote each of 1,000 unseen test texts. Scores came from a class-wide Kaggle leaderboard.
The catch: every word had already been replaced by an integer between 0 and 4,999, so there was no text to read, only sequences of ids. And the training files were wildly unbalanced, with 122,584 human texts against 3,500 machine texts in domain 1, while domain 2 had only 100 human and 400 machine texts.
- 01Bag of wordsPrompt and text concatenated, then counted into a 5,000-wide vector, one slot per token id.
- 02Re-balanceSMOTE over-sampling compared with random under-sampling; SMOTE kept, to avoid discarding data.
- 03One linear detector per domainLogisticRegressionCV and an SGD-trained linear SVM, fitted separately on domain 1, domain 2 and both.
- 04A CNN detourResNet and ShuffleNet on count vectors folded into images: slow, and weaker on the leaderboard.
Key results (original 2023 numbers)
Strong in validation, humbler on the leaderboard
Logistic regression reached 0.9658 validation accuracy on domain 1, but the best leaderboard score was 0.720, from the SGD model. The report recommended logistic regression for its steadier behaviour on the tiny domain 2; the meeting minutes record SGD as the submission pick. Re-running everything in 2026 explained much of the gap between the two kinds of score.
LR · domain 1 · validation
0.9658
F1 0.9643 on 10,000 rows
Best Kaggle score
0.720
SGD trained on domain 1
LR · domain 1 · Kaggle
0.702
the report's recommended model
Test predictions reproduced
6,000
of 6,000: six models × 1,000 items, from the exported weights
| Model | Val. accuracy | Val. F1 | Kaggle |
|---|---|---|---|
| LR 1 | 0.9658 | 0.9643 | 0.702 |
| LR 2 | 0.9437 | 0.9396 | 0.534 |
| LR 1+2 | 0.5005 | 0.5002 | 0.482 |
| SGD 1 | 0.9330 | 0.9315 | 0.720 |
| SGD 2 | 0.5312 | 0.5562 | 0.530 |
| SGD 1+2 | 0.5019 | 0.4493 | 0.486 |
As printed in the 2023 report. The journal explains the weak “1+2” rows and one transcription slip.
Re-scored in 2026 on texts the models never saw
The validation split above was leaky, and the leak sat in the copied machine texts. Re-fitted with the same recipe on a clean split, domain-1 logistic regression catches 0.789 (0.757 to 0.817) of machine texts it never saw, against 0.998 on the 2023 validation part, while human texts are kept at about the same rate (0.933 then, 0.945 now). On the real 97:3 mix (n = 25,217) it reaches AUC 0.960 (0.954 to 0.965), and text length alone reaches AUC 0.932 (0.923 to 0.941) (95% intervals). The held-out re-evaluation has the rest.
The lab
Five exhibits
How every number was produced, what it assumes and why the project is built the way it is: see methods and decision records.
About this project
Case details
COMP90051 Statistical Machine Learning, University of Melbourne, 2023 Semester 1. Project 1 (Kaggle in-class competition). Report finalised on 27 April 2023 by Team 27.
Source code lives in the GitHub repository rNLKJA/Unimelb-Master-2023-COMP90051-Project-1, which is private for now. It keeps the untouched original notebooks under coursework/.
- Sunchuangyu HuangBag-of-words and MPI preprocessing, CNN experiments, SMOTE / RUS
- Yifei DuSVM baseline, SGD classifier, bag-of-words, SMOTE / RUS
- Huihui HeLogistic regression, word embeddings, SMOTE / RUS
Original stack (2023)
Python, Jupyter, scikit-learn 1.2, imbalanced-learn, PyTorch and torchvision, mpi4py, pandas, Plotly.
Revived stack (2026)
Next.js 16, React 19, TypeScript, Tailwind CSS 4, Vitest. Weights exported once with a uv script; all scoring runs client-side.
Academic integrity: this is a portfolio revival of completed, already-assessed coursework. The original submission is preserved unchanged in the repository for reference. The assignment specification and the course data set are not reproduced here; the site ships only model weights and aggregate statistics derived from the data.