Skip to content

Showcase · recorded with Playwright

A guided tour of the detector lab.

Three short walkthroughs of the main workflows, and a screenshot of every key feature. A Playwright script records them from a production build of this site with fixed seeds, so they can be re-made whenever the site changes. Every video is silent, captioned and has a written transcript.

Workflows

Watch the main journeys

Each step appears in a caption bar at the bottom of the frame, and the same steps are listed beside the video. The transcript gives the time of each one.

0:38 · 971 kB · MP4 Download the video

Walkthrough 1

Inside the detector

Open the weights explorer, read the token ids that push each detector towards human or machine, switch between domains and learners, and see how little the two domains agree.

  1. 1Each detector is one learnt weight per token id: 5,000 weights and an intercept
  2. 2Logistic regression, domain 1: the token ids that push hardest towards human
  3. 3…and the token ids that push hardest towards machine
  4. 4Switch to domain 2: a different set of token ids leads on both sides
  5. 5The specimen sheet: estimator, chosen C, sparsity and the weight distribution
  6. 6SGD on domain 1: the same data, another learner, other leading tokens
  7. 7Cross-examination: the two domains' weights barely correlate
Transcript

Silent screen recording of inside the detector, captioned step by step. A highlighted circle shows the pointer.

  1. 0:00Each detector is one learnt weight per token id: 5,000 weights and an intercept
  2. 0:03Logistic regression, domain 1: the token ids that push hardest towards human
  3. 0:07…and the token ids that push hardest towards machine
  4. 0:11Switch to domain 2: a different set of token ids leads on both sides
  5. 0:16The specimen sheet: estimator, chosen C, sparsity and the weight distribution
  6. 0:22SGD on domain 1: the same data, another learner, other leading tokens
  7. 0:28Cross-examination: the two domains' weights barely correlate
Try it yourself

0:50 · 2.0 MB · MP4 Download the video

Walkthrough 2

Try it: score a synthetic specimen

Generate machine-like and human-like token sequences from aggregate class profiles, then watch the original logistic regression score them, token by token.

  1. 1Specimens are drawn from aggregate token profiles, never from real texts (seed 90051)
  2. 2A machine-like specimen: logistic regression for domain 1 calls it machine
  3. 3Slide the profile mix to human: the same random draws morph and the verdict flips
  4. 4Draw a fresh human-like specimen with a new seed
  5. 5What moved the needle: count × weight for the strongest tokens in the specimen
  6. 6Add ten machine-leaning tokens and watch the log-odds d = w·x + b swing back
  7. 7Show a miss: a 90% machine-profile specimen that the detector calls human
  8. 8The line-up: all six original detectors score the same specimen
Transcript

Silent screen recording of try it: score a synthetic specimen, captioned step by step. A highlighted circle shows the pointer.

  1. 0:00Specimens are drawn from aggregate token profiles, never from real texts (seed 90051)
  2. 0:04A machine-like specimen: logistic regression for domain 1 calls it machine
  3. 0:09Slide the profile mix to human: the same random draws morph and the verdict flips
  4. 0:17Draw a fresh human-like specimen with a new seed
  5. 0:22What moved the needle: count × weight for the strongest tokens in the specimen
  6. 0:27Add ten machine-leaning tokens and watch the log-odds d = w·x + b swing back
  7. 0:35Show a miss: a 90% machine-profile specimen that the detector calls human
  8. 0:42The line-up: all six original detectors score the same specimen
Try it yourself

0:56 · 2.6 MB · MP4 Download the video

Walkthrough 3

Imbalance sandbox

Toggle between no re-balancing, random under-sampling, SMOTE and the notebooks' index-SMOTE on a toy problem, watch the boundary move, then check the held-out calibration plot with its intervals.

  1. 1Ninety-seven human texts for every three machine texts: how to learn the rare class?
  2. 2A toy version: two clouds with 4% machine points, refitted on every change
  3. 3No re-balancing: the boundary pushes into the machine cloud and machine recall drops
  4. 4Random under-sampling: most humans are dropped and the boundary moves back
  5. 5SMOTE: synthetic machine points on neighbour segments shift the boundary too
  6. 6Index-SMOTE, what the 2023 notebooks ran: every 'synthetic' row is an exact copy
  7. 7Held-out calibration: observed share of human texts per bin, with Wilson 95% intervals
  8. 8Prior correction, and ECE and Brier score with bootstrap 95% intervals
Transcript

Silent screen recording of imbalance sandbox, captioned step by step. A highlighted circle shows the pointer.

  1. 0:00Ninety-seven human texts for every three machine texts: how to learn the rare class?
  2. 0:06A toy version: two clouds with 4% machine points, refitted on every change
  3. 0:11No re-balancing: the boundary pushes into the machine cloud and machine recall drops
  4. 0:15Random under-sampling: most humans are dropped and the boundary moves back
  5. 0:21SMOTE: synthetic machine points on neighbour segments shift the boundary too
  6. 0:27Index-SMOTE, what the 2023 notebooks ran: every 'synthetic' row is an exact copy
  7. 0:33Held-out calibration: observed share of human texts per bin, with Wilson 95% intervals
  8. 0:44Prior correction, and ECE and Brier score with bootstrap 95% intervals
Try it yourself

Key features

Screenshots

Desktop screenshots are 1440 × 900. Phone screenshots are 390 × 844, rendered at twice the resolution and scaled down. Select one to see it full size.

On a phone

Reproducible

How these were made

One script

web/e2e/showcase.spec.ts drives Google Chrome through each journey with Playwright and doubles as an end-to-end test. Run it with pnpm showcase, and point BASE_URL at any deployment.

Fixed inputs

Specimens start from the site's default seed (90051), and New specimen derives the next seed from it, so a re-recording shows the same tokens, verdicts and sandbox numbers.

No AI in the loop

The lab has no generative-AI feature (see DR-005), so nothing in the tour calls a model. Every verdict comes from the original 2023 weights, scored in the browser.