Exhibit A · interactive detector
Mix a specimen, then ask the human / machine detector.
The real instances are anonymised word-index sequences from the course, so they are not hosted here. Instead, the specimen below is drawn from aggregate token-frequency profiles of each class, computed once from the training files. The verdict comes from the team's original 2023 weights, exported from the pickles and scored with the same rule scikit-learn uses: d = w·x + b.
Loading the exported model weights…
How a specimen is made
For each position a token id is drawn from (1 − m) · p(token | human) + m · p(token | machine), where m is the slider. Lengths follow the class length quantiles. This is exactly the information a bag-of-words model sees: counts, no order.
Sanity check from the export script
On 300 pure-human domain-1 specimens LR 1 says human 279 times; on 300 pure-machine specimens it says human only 47 times. For domain 2 (LR 2) the figures are 258 and 3.
What this is not
It cannot read English. The 2023 detectors only understand the course's 5,000 token ids, so there is no free-text box. Why the leaderboard scores stayed near 0.7 is explained in the revival audit.