Skip to content

How the numbers were produced

Methods, assumptions and decisions.

Where the data came from, how each model works, how the results were evaluated, what the evaluation assumes, and what I would do differently. The 2023 coursework results are reported as submitted. Everything added in 2026 is analysis around them.

Data provenance

Where the data came from

The course provided every text. The site publishes model weights and aggregate statistics only, as the data statement below explains.

Course data (2023)

Four labelled training files from COMP90051 Project 1: 122,584 human and 3,500 machine texts in domain 1, 100 human and 400 machine texts in domain 2, plus 1,000 unlabelled test texts. Every word had been replaced by an integer id from 0 to 4,999 before release.

2023 artefacts, re-run

scripts/export_artefacts.py re-runs the notebooks' logic with the 2023 library versions, reproduces every validation number in the report and all 8,000 submitted predictions, and exports the six models' weights.

2026 re-evaluation

scripts/evaluate_heldout.py re-fits the 2023 recipes on a split made before any re-balancing and writes only aggregates: public/data/evaluation.json and a CSV of the results table.

Methods

What each part does

The 2023 methods are unchanged. The uncertainty analysis is new, and its helpers live in web/src/lib/stats with unit tests against SciPy, statsmodels, scikit-learn and R.

Features

Prompt and text are concatenated and counted into a 5,000-wide bag of words, one slot per token id. Order is lost, length is not: a longer text has larger counts (DR-001).

Models

LogisticRegressionCV (l2, lbfgs, 50 iterations, C chosen by 5-fold CV) and SGDClassifier (hinge loss, l1, alpha 0.0001), both with balanced class weights, one per domain (DR-002).

Re-balancing

SMOTE was meant to synthesise minority texts. It ran on the row index, so it copied them, and the copies were split across training and validation (DR-003).

Statistics added in 2026

  • Percentile bootstrap intervals for accuracy, macro-F1, ROC AUC, ECE and the Brier score, resampling test texts.
  • Paired differences from common resamples, and McNemar's exact test with the discordant odds ratio and Cohen's g as effect sizes.
  • Wilson score intervals for recall per class, per generator and per bin.
  • Reliability diagrams and expected calibration error over ten equal-width bins of P(human), Platt scaling for the hinge-loss model, and a prior correction that moves the log-odds from the balanced training prior to the domain's real share.
  • Two trivial baselines (always the majority class, and a logistic regression on log text length), so every score can be read against them.
  • Ten random and ten prompt-grouped train/test splits and leave-one-generator-out re-fits (with and without that generator's test prompts in training), for the variation an interval on one split cannot show.
  • A prompt-cluster bootstrap for the headline rows, which resamples whole prompts instead of texts.
  • Fitting diagnostics for every logistic regression (cross-validation fits stopped at the iteration limit, C at the edge of its grid), and sensitivity fits with the solver run to convergence and C chosen on folds grouped by source text.

Evaluation design

One held-out split, fixed seeds

Each domain's original texts are split 80/20 before anything else happens, stratified by class and generator. The training part goes through the 2023 recipe (index-SMOTE copies, a 40,000-row cap in domain 1, the 2023 hyper-parameters). The test part is scored once per detector. Why the recipes were re-fitted rather than the pickles re-scored is in DR-004. The results are on /evaluation.

Seeds used on the site
WhatSeedControls
2023 notebooksrandom_state = 90051The team's seed (the subject code): SMOTE, the 50,000-row sample, the 80/20 split and both models.
Held-out splitseed 90051The 80/20 split, the index-SMOTE copies, the 40,000-row cap and both models of the main re-evaluation.
Spread across splitsseeds 90052 to 90060Nine more splits and re-fits, each seeded with its own number.
Prompt-grouped splitsseeds 90051 to 90060The same ten seeds for splits that keep every prompt on one side, and their re-fits.
Bootstrap (Python)seed 90051, B = 10,000NumPy's default generator. The same resamples serve every statistic on one test set, and the prompt-cluster bootstrap draws whole prompts with the same seed.
Bootstrap (browser build)seed 90051, B = 10,000Intervals for the 2023 table, from the confusion matrices (mulberry32 generator).
Detector specimensseed 90051The first synthetic token sequence on /detector and the home page.

Assumptions and limitations

What the numbers assume

Assumptions

  • Test texts are treated as independent draws from the same distribution as the training texts of their domain. Texts that answer the same prompt are not independent, and this matters in two ways. A test text whose prompt was seen in training can be recognised by its prompt, which biases the score upwards; the prompt-grouped splits remove that. And the percentile intervals resample texts, so they are likely too narrow where texts share prompts; the prompt-cluster bootstrap on /evaluation shows by how much for the headline rows.
  • The re-fitted recipe stands in for the 2023 models. It uses the same code and settings on different rows, so it describes the recipe, not the exact pickles. Those settings include a 50-iteration limit that stops the logistic regression before it converges and a choice of C on folds that share copies; the sensitivity fits on /evaluation#fitting show what changes without them.
  • Bootstrap intervals treat the fitted model as fixed. The ten-split spread covers the variation from re-drawing the split and re-fitting.

Limitations

  • Domain 2 has 100 test texts, 20 of them human. Its intervals are wide, and percentile intervals can be too narrow at that size.
  • The Kaggle labels were never released, so the leaderboard scores cannot be recomputed, split by domain or given proper intervals.
  • The texts are anonymised token ids, so no error can be read and explained, and nothing is known about the language, topics or generators.
  • Most domain-1 texts get P(human) above 0.9, so the top calibration bin holds most of the data and the ECE depends on the binning.
  • The CNN experiments were not re-run. Their curves on /journal are as recorded in 2023.

What I'd change

If I did it again

In order of how much they would have changed the result

  1. Split first, then re-balance inside the training folds only (an imbalanced-learn pipeline), with cross-validation folds grouped by source text, and run the solver to convergence instead of silencing its warnings.
  2. Report every score next to a length-only baseline and with an interval, and use macro-F1 or AUC rather than accuracy on data that is 97% one class.
  3. Fit on the natural class balance and choose the decision threshold for the costs of each kind of error, instead of letting re-sampling set it as a side effect.
  4. Share weights across domains with domain-specific offsets, so the 500 domain-2 texts borrow strength from domain 1 instead of being modelled alone.
  5. Hold out whole generators during development, because a deployed detector meets generators it was not trained on.

Model card

The detectors

Six linear classifiers from COMP90051 Project 1 (University of Melbourne, 2023 Semester 1, Team 27) that guess whether a short text was written by a person or generated by a machine. They run in the browser at /detector. The 2023 numbers are quoted as printed. The 2026 numbers come from scripts/evaluate_heldout.py (DR-004), and a unit test checks that this card quotes the same numbers as the evaluation file.

Model details

PropertyLogistic regressionSGD (hinge)
EstimatorLogisticRegressionCV, l2, lbfgs, 50 iterationsSGDClassifier, hinge loss (a linear SVM), l1
RegularisationC chosen by 5-fold CV from 10 valuesalpha = 0.0001
Class handlingBalanced class weights after SMOTE re-balancingsame
OutputP(human) and a labelA score and a label, no probability
Libraryscikit-learn 1.2.2same
  • Six models. One of each kind for domain 1, for domain 2 and for both domains pooled (DR-002). The pooled 2023 models were trained on mismatched features and labels and should be read as broken.
  • Input. A prompt and a text, both as sequences of token ids from 0 to 4,999, concatenated and counted into 5,000 features. Word order is ignored, and longer texts have larger counts (DR-001).
  • Rule. A weighted sum of the counts plus an intercept. Above zero means human.
  • Fitting caveat. The 2023 settings give lbfgs 50 iterations, which stops the logistic regression before it converges. In the 2026 domain-1 re-fit all 50 cross-validation fits and the final fit stop at that limit, and cross-validation picks the largest C in its grid, 10,000, helped by copies of one text sitting in different inner folds. SGD on domain 1 also runs all 1,000 of its epochs. The 2023 notebooks silenced both warnings. Fitted to convergence with C chosen on folds grouped by source text, the domain-1 recipe picks C = 0.006 and scores macro-F1 0.644 (0.634 to 0.654) and AUC 0.971 (0.966 to 0.975): it ranks better but flags more human texts. /evaluation#fitting shows every variant.
  • Authors. Yifei Du, Huihui He and Sunchuangyu Huang (Team 27). The browser port by Sunchuangyu Huang reproduces all 6,000 of the models' test-set predictions.

Intended use

Teaching and portfolio demonstration. The lab shows how a bag-of-words detector reaches a verdict, what its weights look like, and how class imbalance and data leakage distort an evaluation.

Out-of-scope uses

This is not an AI-text detector for real use. Do not use these models, or anything built like them, to decide whether a student's assignment, a job application or any other person's writing was produced by AI. They cannot read text: they accept only the course's anonymised token ids, whose mapping to words was never released. They were trained on generators from before April 2023, catch fewer texts from generators they have not seen and do not transfer between domains. Much of what they detect is text length.

Training data

The course's four labelled files: 122,584 human and 3,500 machine texts in domain 1, and 100 human and 400 machine texts in domain 2. Machine texts come from five numbered generators in domain 1 and four in domain 2. The data statement lists what is known and what is not. In 2023 each domain was re-balanced with SMOTE before the 80/20 validation split (DR-003).

Evaluation

2023 validation, as printed

Logistic regression on domain 1 reported 0.9658 accuracy and 0.9643 F1. These numbers are inflated: the minority class was copied before the split, so the validation part held copies of training texts. They are kept unchanged on /journal and shown with intervals on /evaluation.

2026 held-out re-evaluation

The 2023 recipes were re-fitted on an 80/20 split made before any re-balancing, with seed 90051, and scored on the untouched 20%. Intervals are 95% percentile bootstrap intervals over the test texts (10,000 resamples, seed 90051).

DetectorTest textsMacro-F1ROC AUCAccuracyMacro-F1, ten splits
Logistic regression, domain 124,517 human, 700 machine0.697 (0.685 to 0.709)0.960 (0.954 to 0.965)0.941 (0.938 to 0.944)0.696 ± 0.004
SGD, domain 1same0.614 (0.603 to 0.625)0.902 (0.893 to 0.911)0.911 (0.908 to 0.915)0.614 ± 0.005
Text length only, domain 1same0.564 (0.556 to 0.571)0.932 (0.923 to 0.941)0.827 (0.823 to 0.832)not re-run
Always "human", domain 1same0.4930.50.972not re-run
Logistic regression, domain 220 human, 80 machine0.740 (0.610 to 0.850)0.760 (0.607 to 0.894)0.860 (0.790 to 0.920)0.657 ± 0.061
SGD, domain 2same0.697 (0.577 to 0.805)0.739 (0.588 to 0.870)0.810 (0.730 to 0.880)0.617 ± 0.068
Text length only, domain 2same0.546 (0.443 to 0.644)0.657 (0.525 to 0.781)0.590 (0.490 to 0.690)not re-run

The last column is the mean ± standard deviation over ten re-drawn splits (seeds 90051 to 90060), each re-fitted from scratch. The domain-2 numbers in the other columns come from a split where a prompt shortcut is available (failure mode 5).

Paired comparison. On the same domain-1 texts, logistic regression minus SGD is +0.083 (+0.074 to +0.093) in macro-F1, and logistic regression was ahead in 10 of 10 re-drawn splits. On domain 2 the difference is not distinguishable from zero.

Spread across splits. Across ten re-drawn splits the domain-2 logistic regression's macro-F1 averaged 0.657 (SD 0.061, range 0.569 to 0.750), about the spread its bootstrap interval implies. The main split is the second most favourable of the ten, so the domain-2 headline numbers sit on the favourable side.

Intervals and shared prompts. The intervals resample texts as if they were independent, but texts that answer the same prompt are not. Resampling whole prompts instead (25,217 domain-1 test texts in 19,311 prompts) widens the domain-1 logistic regression's macro-F1 interval slightly, to 0.683 to 0.711 instead of 0.685 to 0.709. The 100 domain-2 test texts answer 82 prompts, and their intervals barely change.

Sensitivity to fitting. Fitted as in 2023, the domain-1 recipe stops early at C = 10,000 (see the fitting caveat above). Run to convergence at that C it scores macro-F1 0.688 (0.675 to 0.700); with C also chosen on grouped folds, 0.644 (0.634 to 0.654). In domain 2 the grouped folds push C to the bottom of the grid and macro-F1 falls to 0.584 (0.479 to 0.687).

Calibration (domain 1, expected calibration error over 10 bins).

ProbabilitiesECE
Logistic regression, as trained0.054 (0.051 to 0.056)
SGD with Platt scaling, as trained0.289 (0.286 to 0.292)
Logistic regression, prior-corrected0.023 (0.022 to 0.025)
SGD with Platt scaling, prior-corrected0.007 (0.005 to 0.009)
Logistic regression without re-balancing0.002 (0.002 to 0.004)

Re-balancing teaches both detectors that half of all texts are human. Shifting their log-odds to the real 97.2% human share in domain 1 repairs most of the miscalibration.

Known failure modes

  1. Base rates. At the domain-1 mix of 97 human texts to 3 machine texts, 1,343 of the 1,895 texts that logistic regression flags as machine-written are human. Only 0.291 (0.271 to 0.312) of its machine flags are correct. A flag from this detector is weak evidence on its own.

  2. Length. A logistic regression on text length alone reaches AUC 0.932 (0.923 to 0.941) in domain 1. The full detector adds +0.027 (+0.019 to +0.036) AUC on the same texts, so most of its ranking could be done by counting tokens.

  3. Unseen generators. Re-fitted without one domain-1 generator, logistic regression catches fewer of that generator's 140 test texts for all five generators, by 5 to 19 texts. The drop is significant for three of them (McNemar p ≤ 0.03), and AUC changes by at most 0.022, so the unseen generator's texts are still mostly ranked below human texts.

    GeneratorCaught, generator seen in trainingCaught, generator left outMcNemar exact
    088 of 14083 of 140p = 0.27
    198 of 14087 of 140p = 0.03
    2119 of 140107 of 140p = 0.002
    3109 of 14090 of 140p < 0.001
    4138 of 140133 of 140p = 0.13

    Leaving a generator out does not hide its prompts: 118 to 140 of each generator's 140 test texts answer a prompt that the other generators' training texts also answer. With every training text that answers one of those prompts removed as well, leaving the generator out still lowers recall for all five, significantly for four.

    GeneratorPrompts removed, generator keptPrompts removed, generator left outMcNemar exact
    078 of 14069 of 140p = 0.02
    178 of 14067 of 140p = 0.04
    2113 of 14094 of 140p < 0.001
    392 of 14072 of 140p < 0.001
    4135 of 140130 of 140p = 0.06

    This is a loss of recall, not a collapse. Every generator in the data predates April 2023, so current models are unseen generators, and a detector like this has to be re-checked whenever a new one appears.

  4. Domain shift. The domain-1 logistic regression scored on all 500 domain-2 texts reaches AUC 0.553 (0.483 to 0.621), not distinguishable from chance (the interval includes 0.5).

  5. Prompt shortcuts. In domain 2 a prompt is answered either by one human text or by four machine texts, never both. In the main split all 80 machine test texts answer a prompt seen in training and none of the 20 human ones does, so "prompt seen in training" alone separates the classes; across the ten random splits it holds for 76 to 80 of the 80 machine texts. On ten splits that keep each prompt on one side, the domain-2 logistic regression averages AUC 0.613 ± 0.059 against 0.725 ± 0.059 on the ten random splits, and macro-F1 0.558 ± 0.039 against 0.657 ± 0.061 (seed 90051 alone: macro-F1 0.595 (0.484 to 0.701)). Domain 1 points the same way on a smaller scale: grouped macro-F1 0.676 ± 0.010 against 0.696 ± 0.004, and 0.673 (0.661 to 0.685) for seed 90051, below all ten random splits.

  6. Small samples. Domain 2 rests on 500 texts, 100 of them human, and every domain-2 number has a wide interval.

  7. Opacity. The texts are anonymised, so no error can be read, and the weights on /weights name token ids, not words.

Ethical considerations

  • False accusations. Wrongly calling a person's writing machine-generated harms that person, and the base-rate failure above shows how easily that happens when machine text is rare. Any real detector needs a human decision, a way to contest it and evidence beyond a score.
  • Unknown writers. Nothing is known about who wrote the human texts, so there is no way to check whether some groups of writers are flagged more often than others.
  • Transparency. The weights, the evaluation and the code are public in this repository. This card is informed by the transparency principles of the Australian Government's policy for the responsible use of AI in government, the EU AI Act and the NIST AI Risk Management Framework. It does not claim compliance with any of them.
  • Data. No course text leaves the repository, and the site serves only weights and aggregates.

Data statement

The course data

This statement describes the data behind the detectors on this site, following the spirit of data statements for NLP (Bender and Friedman, 2018). Much of what such a statement would normally record is unknown here, and the gaps are stated as gaps.

Source

The teaching team of COMP90051 Statistical Machine Learning (University of Melbourne, 2023 Semester 1) provided the data for Project 1, a class-wide Kaggle competition. The archive is kept in the repository at coursework/data/sml.p1.data.tar.gz for reproducibility of the original work. The course's README for the data is in coursework/data/README.md.

What one instance looks like

Each instance is a short text paired with the prompt that produced it. Both are given as sequences of integers from 0 to 4,999. The course applied "light preprocessing" and then mapped every word to an integer with one shared mapping. The mapping was never released, so no word can be recovered. Machine-generated instances also carry a machine_id that names the generator by number only.

Size and balance

FileDomainLabelInstancesGeneratorsMean length (tokens)
set1_human.json1human122,584380
set1_machine.json1machine3,5005 × 700183
set2_human.json2human100820
set2_machine.json2machine4004 × 100586
test.json600 from domain 1, 400 from domain 2unlabelled1,000409

Lengths count prompt and text together. Human texts in domain 1 are about twice as long as machine texts, which makes length alone a strong signal there (see the model card).

Texts share prompts. The 3,500 domain-1 machine texts answer 789 distinct prompts, and 64% of them have a prompt that also appears among the human texts. In domain 2 the 400 machine texts answer 100 prompts, four texts per prompt, and none of those prompts appears among the 100 human texts. A detector can therefore learn which prompts belong to machine texts, which is why the evaluation also reports a split that keeps each prompt on one side.

Unknowns

  • Language and register. The course did not say which language the texts are in or where they came from. The two domains are described only as "different data sources or data distributions".
  • Authors. Nothing is known about the people who wrote the human texts, so no claim can be made about which groups of writers the detectors serve well or badly.
  • Generators. The machine texts come from five numbered generators in domain 1 and four in domain 2. Their names, versions and decoding settings are unknown. All of them predate April 2023.
  • Collection dates and consent. Not documented in the material we received.

How the data is used here

  • scripts/export_artefacts.py re-runs the 2023 notebooks on the archive and checks every reported number.
  • scripts/evaluate_heldout.py re-fits the 2023 recipes on an 80/20 split of the four training files and scores the held-out 20%.
  • The 1,000 test instances are used only to reproduce the eight 2023 Kaggle submissions. Their labels were never released.

What is published

The website and the files under web/ contain no instance, no token sequence and no per-instance prediction. They hold:

  • the six models' weights (5,000 coefficients and an intercept each);
  • per domain and class, how often each token id occurs and the quantiles of text length;
  • aggregate evaluation results: counts, confusion matrices, metrics with intervals, reliability-diagram bins and ROC curves on a fixed grid.

The synthetic "specimens" on /detector are drawn from those frequency profiles. They are not course texts and do not resemble any single one.

Access

The data belongs to the course. It is kept in this private repository only to reproduce the original coursework and is not redistributed by the website. Please do not reuse it for current coursework.

AI use statement

Where AI is and is not used

This statement explains where artificial intelligence is and is not part of this site. It is informed by the transparency principles in the Australian Government's policy for the responsible use of AI in government, the EU AI Act's transparency provisions and the NIST AI Risk Management Framework. It is not a claim of compliance with any of them.

The detectors are not generative AI

The six detectors are linear models from scikit-learn, trained by our team in April 2023: logistic regression and a linear support vector machine trained by stochastic gradient descent. Each one adds up 5,000 learnt weights, one per token id, and compares the total with zero. Their weights are published on /weights, and every score on the site can be recomputed from them.

No generative-AI feature on the site

The site makes no call to a language model or any other AI service, needs no API key and sends nothing typed or generated on the page to a third party. All scoring runs in the visitor's browser. DR-005 explains why I did not add an optional bring-your-own-key feature.

AI assistance in building the 2026 revival

The 2026 web app, the analysis scripts and these documents were written with help from an AI coding assistant (Claude, by Anthropic). I reviewed the changes, and every number on the site comes from a script in the repository that anyone can re-run:

  • scripts/export_artefacts.py reproduces the 2023 numbers and predictions;
  • scripts/evaluate_heldout.py produces the held-out re-evaluation;
  • scripts/stats_reference.py produces the reference values the statistics tests check.

The unit tests compare the TypeScript statistics against SciPy, statsmodels, scikit-learn and R, and a test checks that the numbers quoted in the model card match the evaluation file.

The 2023 coursework

The 2023 notebooks, report and meeting minutes are the team's original work. They are kept unchanged under coursework/, apart from a formatting pass with no logic changes.

What a detector like this must never be used for

These detectors must not be used to decide whether a person's writing was produced by AI. The model card explains why: they were trained on anonymised data from generators that are now outdated, they catch fewer texts from generators they have not seen, they do not transfer between domains, and much of what they detect is text length.

Decision records

Why it is built this way

Each record states the decision first, then the context, the options, the reasons, what happened (including the weak numbers) and what I would change. Records are never edited after the fact. A later decision supersedes an earlier one.