Decision record · DR-001
Bag-of-words counts and linear models instead of CNNs
- Status
- Accepted
- Date
- 2023-04 (recorded 2026-10)
- Applies to
- coursework/notebooks/final, /detector, /weights, /evaluation
Decision in one line
In April 2023 we turned each text into 5,000 token counts and classified it with linear models, logistic regression and an SGD-trained linear SVM, and we dropped the word-embedding and CNN pipelines because they were slower, weaker on the leaderboard and scored by test splits we could not trust.
Context
Every word in the course data had been replaced by an integer id from 0 to 4,999, so there was no text to read, only sequences of ids. At the first meeting (14 April 2023) each of us championed one model family. Yifei took a support vector machine, Huihui logistic regression, and I took a convolutional network.
Within a week the options had narrowed.
- Word embeddings expanded the 250 MB data set towards 90 GB, far beyond a 16 GB laptop. Cutting them to 20 dimensions made them fit, but compressed away enough information to hurt recall.
- Bag-of-words gave 5,000-wide count vectors that are over 90% zeros but fitted in memory.
- My CNNs treated the 5,000 counts as a folded image and ran LeNet, a custom CNN, ImageNet-pretrained ResNet18 and ResNet152, and ShuffleNetV2. A three-layer CNN scored about 0.5 on Kaggle and ShuffleNet 0.618, while their own test accuracy passed 99%.
- A plain SVM found no clear boundary on the sparse counts and was slow to train.
- SGD with built-in class balancing scored 0.71 on Kaggle on 22 April, and logistic regression on SMOTE data 0.666.
Decision
Concatenate prompt and text, count them into a 5,000-wide bag of words, and fit one linear classifier per domain (DR-002):
LogisticRegressionCVwith an l2 penalty, the lbfgs solver, 50 iterations and balanced class weights;SGDClassifierwith hinge loss, an l1 penalty, alpha 0.0001 and balanced class weights.
The CNN work stayed in the notebooks and out of the report. The report recommended logistic regression. The meeting minutes record SGD, with the best Kaggle score of 0.720, as the submission pick.
Options considered
- CNNs on the counts folded into a grid (my approach). Convolution assumes that neighbouring cells are related. Folding put token id 117 next to ids 118 and 157 for no reason connected to meaning, so local filters had no structure to find. Each run also took hours on a laptop GPU.
- Word embeddings with a small CNN. Limited by memory, and the 20-dimensional version lost too much.
- A plain SVM. Slow on 126,084 sparse rows and no better than the SGD version.
- Bag-of-words with linear models (chosen).
Why
- A linear model on counts uses every token id directly, and its 5,000 weights can be read
one by one (the
/weightspage does exactly that). - Training takes seconds, so we could compare two re-balancing methods on three data sets before the deadline.
- The leaderboard was the only score not computed on our own splits, and it favoured the linear models at 0.720 and 0.702 against 0.618 for ShuffleNet.
What happened
- Leaderboard. SGD on domain 1 scored 0.720, logistic regression on domain 1 0.702 and ShuffleNet 0.618. If the leaderboard scored all 1,000 test texts, Wilson 95% intervals are 0.691 to 0.747 for 0.720 and 0.673 to 0.730 for 0.702, so the two linear models cannot be told apart, while 0.618 (0.587 to 0.648) sits below both. The public leaderboard may have scored fewer texts, which would make every interval wider.
- The CNNs' 99%. Their test rows came from the same duplicated pool as their training rows (DR-003), which explains the gap to the leaderboard better than over-fitting in the usual sense.
- Clean re-evaluation in 2026 (DR-004). On 25,217 held-out domain-1 texts, logistic regression reaches macro-F1 0.697 (0.685 to 0.709) and AUC 0.960 (0.954 to 0.965).
- The weak number. A logistic regression on text length alone reaches AUC 0.932 (0.923 to 0.941) on the same texts. Machine texts in domain 1 are about half as long as human ones, and bag-of-words counts grow with length. On the same texts the bag-of-words model adds +0.027 (+0.019 to +0.036) AUC over length, so most of what the 2023 detector ranks correctly in domain 1, a length rule ranks correctly too.
What I'd change
- Report every model against a length-only baseline from the first day.
- Hold out a clean test split before comparing model families, so a 99% test accuracy is caught as a leak on the day it appears.
- Try token n-grams or TF-IDF weighting with the same linear models before reaching for deep models. If I tried a neural model again, I would give it the token sequence (a 1-D CNN or a small transformer over ids) instead of folded counts, so that it could use word order.