Skip to content

Exhibit B · model weights

Fingerprints: one weight per token id, six detectors.

Each model is a single 5,000-long weight vector plus an intercept. A positive weight means “every occurrence of this token id makes the sequence look more human”; a negative one points to machine. These are the exact coefficients from the team's pickles, rounded to seven significant digits (which changes none of the 1,000 test predictions).

Per model

Strongest token ids by model

The word behind each id is unknown to us: the course mapped words to integers before release. What we can see is which ids the detectors lean on, and how sharply.

Pushes towards human

largest positive weights
  • 36722.651
  • 622.533
  • 38672.350
  • 24402.154
  • 29102.084
  • 31012.043
  • 35541.981
  • 24621.973
  • 33291.957
  • 17161.950
  • 33301.935
  • 27731.904

Pushes towards machine

most negative weights
  • 4614-2.297
  • 44-2.050
  • 2826-1.944
  • 4298-1.846
  • 2391-1.766
  • 3814-1.757
  • 4294-1.699
  • 5-1.583
  • 3840-1.581
  • 4054-1.576
  • 4462-1.552
  • 4622-1.527

Specimen sheet

LR_domain1.pkl
Estimator
LogisticRegressionCV
Settings
l2, lbfgs, max_iter=50, balanced
C chosen by 5-fold CV
10,000
Intercept b
-14.8
Non-zero weights
4,093 / 5,000
+ / − weights
2761 / 1332
Validation accuracy (as reported)
0.9658 · Kaggle 0.702
-2.6502.65
Distribution of the non-zero weights (log count)

Cross-examination

Do the detectors agree with each other?

Each square plots the same token's weight in two models (5,000 points, log-shaded density, axes clipped at the 99th percentile). Agreement along the diagonal means both models read the token the same way.

Do the two domains agree?

LR 1 vs LR 2
LR 1 →LR 2 →
Logistic regression · domain 1 (x) vs Logistic regression · domain 2 (y)
Pearson r
0.01
shared top-50 human
0
shared top-50 machine
2

Do the two learners agree?

LR 1 vs SGD 1
LR 1 →SGD 1 →
Logistic regression · domain 1 (x) vs SGD (hinge) · domain 1 (y)
Pearson r
0.50
shared top-50 human
6
shared top-50 machine
4

Does 1+2 match domain 1?

LR 1 vs LR 1+2
LR 1 →LR 1+2 →
Logistic regression · domain 1 (x) vs Logistic regression · domain 1+2 (y)
Pearson r
0.00
shared top-50 human
1
shared top-50 machine
0

Domain shift, in one number

The domain-1 and domain-2 logistic regressions correlate at r = 0.01 and share 0 of their 50 most human-leaning token ids: they are essentially unrelated detectors. That is why the last file of the team's result notebook routes the 600 domain-1 test items to LR 1 and the 400 domain-2 items to LR 2. The 1+2 model is unrelated to both for a different reason, explained in the revival audit.

Weights versus raw frequency

LR 1's weights correlate at r = 0.28 with the simple log-ratio of how often each token appears in human versus machine text in domain 1. Related, but far from identical: the model fits all tokens jointly, so a token's weight also depends on the company it keeps. The detector uses these frequency profiles to build specimens.