Exhibit B · model weights
Fingerprints: one weight per token id, six detectors.
Each model is a single 5,000-long weight vector plus an intercept. A positive weight means “every occurrence of this token id makes the sequence look more human”; a negative one points to machine. These are the exact coefficients from the team's pickles, rounded to seven significant digits (which changes none of the 1,000 test predictions).
Per model
Strongest token ids by model
The word behind each id is unknown to us: the course mapped words to integers before release. What we can see is which ids the detectors lean on, and how sharply.
Pushes towards human
- 36722.651
- 622.533
- 38672.350
- 24402.154
- 29102.084
- 31012.043
- 35541.981
- 24621.973
- 33291.957
- 17161.950
- 33301.935
- 27731.904
Pushes towards machine
- 4614-2.297
- 44-2.050
- 2826-1.944
- 4298-1.846
- 2391-1.766
- 3814-1.757
- 4294-1.699
- 5-1.583
- 3840-1.581
- 4054-1.576
- 4462-1.552
- 4622-1.527
Specimen sheet
- Estimator
- LogisticRegressionCV
- Settings
- l2, lbfgs, max_iter=50, balanced
- C chosen by 5-fold CV
- 10,000
- Intercept b
- -14.8
- Non-zero weights
- 4,093 / 5,000
- + / − weights
- 2761 / 1332
- Validation accuracy (as reported)
- 0.9658 · Kaggle 0.702
Cross-examination
Do the detectors agree with each other?
Each square plots the same token's weight in two models (5,000 points, log-shaded density, axes clipped at the 99th percentile). Agreement along the diagonal means both models read the token the same way.
Do the two domains agree?
- Pearson r
- 0.01
- shared top-50 human
- 0
- shared top-50 machine
- 2
Do the two learners agree?
- Pearson r
- 0.50
- shared top-50 human
- 6
- shared top-50 machine
- 4
Does 1+2 match domain 1?
- Pearson r
- 0.00
- shared top-50 human
- 1
- shared top-50 machine
- 0
Domain shift, in one number
The domain-1 and domain-2 logistic regressions correlate at r = 0.01 and share 0 of their 50 most human-leaning token ids: they are essentially unrelated detectors. That is why the last file of the team's result notebook routes the 600 domain-1 test items to LR 1 and the 400 domain-2 items to LR 2. The 1+2 model is unrelated to both for a different reason, explained in the revival audit.
Weights versus raw frequency
LR 1's weights correlate at r = 0.28 with the simple log-ratio of how often each token appears in human versus machine text in domain 1. Related, but far from identical: the model fits all tokens jointly, so a token's weight also depends on the company it keeps. The detector uses these frequency profiles to build specimens.