How the numbers were produced
Methods, assumptions and decisions.
Where the data came from, how each model works, how the results were evaluated, what the evaluation assumes, and what I would do differently. The 2023 coursework results are reported as submitted. Everything added in 2026 is analysis around them.
Data provenance
Where the data came from
The course provided every text. The site publishes model weights and aggregate statistics only, as the data statement below explains.
Course data (2023)
Four labelled training files from COMP90051 Project 1: 122,584 human and 3,500 machine texts in domain 1, 100 human and 400 machine texts in domain 2, plus 1,000 unlabelled test texts. Every word had been replaced by an integer id from 0 to 4,999 before release.
2023 artefacts, re-run
scripts/export_artefacts.py re-runs the notebooks' logic with the 2023 library versions, reproduces every validation number in the report and all 8,000 submitted predictions, and exports the six models' weights.
2026 re-evaluation
scripts/evaluate_heldout.py re-fits the 2023 recipes on a split made before any re-balancing and writes only aggregates: public/data/evaluation.json and a CSV of the results table.
Methods
What each part does
The 2023 methods are unchanged. The uncertainty analysis is new, and its helpers live in web/src/lib/stats with unit tests against SciPy, statsmodels, scikit-learn and R.
Features
Prompt and text are concatenated and counted into a 5,000-wide bag of words, one slot per token id. Order is lost, length is not: a longer text has larger counts (DR-001).
Models
LogisticRegressionCV (l2, lbfgs, 50 iterations, C chosen by 5-fold CV) and SGDClassifier (hinge loss, l1, alpha 0.0001), both with balanced class weights, one per domain (DR-002).
Re-balancing
SMOTE was meant to synthesise minority texts. It ran on the row index, so it copied them, and the copies were split across training and validation (DR-003).
Statistics added in 2026
- Percentile bootstrap intervals for accuracy, macro-F1, ROC AUC, ECE and the Brier score, resampling test texts.
- Paired differences from common resamples, and McNemar's exact test with the discordant odds ratio and Cohen's g as effect sizes.
- Wilson score intervals for recall per class, per generator and per bin.
- Reliability diagrams and expected calibration error over ten equal-width bins of P(human), Platt scaling for the hinge-loss model, and a prior correction that moves the log-odds from the balanced training prior to the domain's real share.
- Two trivial baselines (always the majority class, and a logistic regression on log text length), so every score can be read against them.
- Ten random and ten prompt-grouped train/test splits and leave-one-generator-out re-fits (with and without that generator's test prompts in training), for the variation an interval on one split cannot show.
- A prompt-cluster bootstrap for the headline rows, which resamples whole prompts instead of texts.
- Fitting diagnostics for every logistic regression (cross-validation fits stopped at the iteration limit, C at the edge of its grid), and sensitivity fits with the solver run to convergence and C chosen on folds grouped by source text.
Evaluation design
One held-out split, fixed seeds
Each domain's original texts are split 80/20 before anything else happens, stratified by class and generator. The training part goes through the 2023 recipe (index-SMOTE copies, a 40,000-row cap in domain 1, the 2023 hyper-parameters). The test part is scored once per detector. Why the recipes were re-fitted rather than the pickles re-scored is in DR-004. The results are on /evaluation.
| What | Seed | Controls |
|---|---|---|
| 2023 notebooks | random_state = 90051 | The team's seed (the subject code): SMOTE, the 50,000-row sample, the 80/20 split and both models. |
| Held-out split | seed 90051 | The 80/20 split, the index-SMOTE copies, the 40,000-row cap and both models of the main re-evaluation. |
| Spread across splits | seeds 90052 to 90060 | Nine more splits and re-fits, each seeded with its own number. |
| Prompt-grouped splits | seeds 90051 to 90060 | The same ten seeds for splits that keep every prompt on one side, and their re-fits. |
| Bootstrap (Python) | seed 90051, B = 10,000 | NumPy's default generator. The same resamples serve every statistic on one test set, and the prompt-cluster bootstrap draws whole prompts with the same seed. |
| Bootstrap (browser build) | seed 90051, B = 10,000 | Intervals for the 2023 table, from the confusion matrices (mulberry32 generator). |
| Detector specimens | seed 90051 | The first synthetic token sequence on /detector and the home page. |
Assumptions and limitations
What the numbers assume
Assumptions
- Test texts are treated as independent draws from the same distribution as the training texts of their domain. Texts that answer the same prompt are not independent, and this matters in two ways. A test text whose prompt was seen in training can be recognised by its prompt, which biases the score upwards; the prompt-grouped splits remove that. And the percentile intervals resample texts, so they are likely too narrow where texts share prompts; the prompt-cluster bootstrap on /evaluation shows by how much for the headline rows.
- The re-fitted recipe stands in for the 2023 models. It uses the same code and settings on different rows, so it describes the recipe, not the exact pickles. Those settings include a 50-iteration limit that stops the logistic regression before it converges and a choice of C on folds that share copies; the sensitivity fits on /evaluation#fitting show what changes without them.
- Bootstrap intervals treat the fitted model as fixed. The ten-split spread covers the variation from re-drawing the split and re-fitting.
Limitations
- Domain 2 has 100 test texts, 20 of them human. Its intervals are wide, and percentile intervals can be too narrow at that size.
- The Kaggle labels were never released, so the leaderboard scores cannot be recomputed, split by domain or given proper intervals.
- The texts are anonymised token ids, so no error can be read and explained, and nothing is known about the language, topics or generators.
- Most domain-1 texts get P(human) above 0.9, so the top calibration bin holds most of the data and the ECE depends on the binning.
- The CNN experiments were not re-run. Their curves on /journal are as recorded in 2023.
What I'd change
If I did it again
In order of how much they would have changed the result
- Split first, then re-balance inside the training folds only (an imbalanced-learn pipeline), with cross-validation folds grouped by source text, and run the solver to convergence instead of silencing its warnings.
- Report every score next to a length-only baseline and with an interval, and use macro-F1 or AUC rather than accuracy on data that is 97% one class.
- Fit on the natural class balance and choose the decision threshold for the costs of each kind of error, instead of letting re-sampling set it as a side effect.
- Share weights across domains with domain-specific offsets, so the 500 domain-2 texts borrow strength from domain 1 instead of being modelled alone.
- Hold out whole generators during development, because a deployed detector meets generators it was not trained on.
Model card
The detectors
Six linear classifiers from COMP90051 Project 1 (University of Melbourne, 2023 Semester 1,
Team 27) that guess whether a short text was written by a person or generated by a machine.
They run in the browser at /detector. The 2023 numbers are quoted as printed. The 2026
numbers come from scripts/evaluate_heldout.py (DR-004),
and a unit test checks that this card quotes the same numbers as the evaluation file.
Model details
| Property | Logistic regression | SGD (hinge) |
|---|---|---|
| Estimator | LogisticRegressionCV, l2, lbfgs, 50 iterations | SGDClassifier, hinge loss (a linear SVM), l1 |
| Regularisation | C chosen by 5-fold CV from 10 values | alpha = 0.0001 |
| Class handling | Balanced class weights after SMOTE re-balancing | same |
| Output | P(human) and a label | A score and a label, no probability |
| Library | scikit-learn 1.2.2 | same |
- Six models. One of each kind for domain 1, for domain 2 and for both domains pooled (DR-002). The pooled 2023 models were trained on mismatched features and labels and should be read as broken.
- Input. A prompt and a text, both as sequences of token ids from 0 to 4,999, concatenated and counted into 5,000 features. Word order is ignored, and longer texts have larger counts (DR-001).
- Rule. A weighted sum of the counts plus an intercept. Above zero means human.
- Fitting caveat. The 2023 settings give lbfgs 50 iterations, which stops the logistic
regression before it converges. In the 2026 domain-1 re-fit all 50 cross-validation fits
and the final fit stop at that limit, and cross-validation picks the largest C in its grid,
10,000, helped by copies of one text sitting in different inner folds. SGD on domain 1 also
runs all 1,000 of its epochs. The 2023 notebooks silenced both warnings. Fitted to
convergence with C chosen on folds grouped by source text, the domain-1 recipe picks
C = 0.006 and scores macro-F1 0.644 (0.634 to 0.654) and AUC 0.971 (0.966 to 0.975):
it ranks better but flags more human texts.
/evaluation#fittingshows every variant. - Authors. Yifei Du, Huihui He and Sunchuangyu Huang (Team 27). The browser port by Sunchuangyu Huang reproduces all 6,000 of the models' test-set predictions.
Intended use
Teaching and portfolio demonstration. The lab shows how a bag-of-words detector reaches a verdict, what its weights look like, and how class imbalance and data leakage distort an evaluation.
Out-of-scope uses
This is not an AI-text detector for real use. Do not use these models, or anything built like them, to decide whether a student's assignment, a job application or any other person's writing was produced by AI. They cannot read text: they accept only the course's anonymised token ids, whose mapping to words was never released. They were trained on generators from before April 2023, catch fewer texts from generators they have not seen and do not transfer between domains. Much of what they detect is text length.
Training data
The course's four labelled files: 122,584 human and 3,500 machine texts in domain 1, and 100 human and 400 machine texts in domain 2. Machine texts come from five numbered generators in domain 1 and four in domain 2. The data statement lists what is known and what is not. In 2023 each domain was re-balanced with SMOTE before the 80/20 validation split (DR-003).
Evaluation
2023 validation, as printed
Logistic regression on domain 1 reported 0.9658 accuracy and 0.9643 F1. These numbers are
inflated: the minority class was copied before the split, so the validation part held copies
of training texts. They are kept unchanged on /journal and shown with intervals on
/evaluation.
2026 held-out re-evaluation
The 2023 recipes were re-fitted on an 80/20 split made before any re-balancing, with seed 90051, and scored on the untouched 20%. Intervals are 95% percentile bootstrap intervals over the test texts (10,000 resamples, seed 90051).
| Detector | Test texts | Macro-F1 | ROC AUC | Accuracy | Macro-F1, ten splits |
|---|---|---|---|---|---|
| Logistic regression, domain 1 | 24,517 human, 700 machine | 0.697 (0.685 to 0.709) | 0.960 (0.954 to 0.965) | 0.941 (0.938 to 0.944) | 0.696 ± 0.004 |
| SGD, domain 1 | same | 0.614 (0.603 to 0.625) | 0.902 (0.893 to 0.911) | 0.911 (0.908 to 0.915) | 0.614 ± 0.005 |
| Text length only, domain 1 | same | 0.564 (0.556 to 0.571) | 0.932 (0.923 to 0.941) | 0.827 (0.823 to 0.832) | not re-run |
| Always "human", domain 1 | same | 0.493 | 0.5 | 0.972 | not re-run |
| Logistic regression, domain 2 | 20 human, 80 machine | 0.740 (0.610 to 0.850) | 0.760 (0.607 to 0.894) | 0.860 (0.790 to 0.920) | 0.657 ± 0.061 |
| SGD, domain 2 | same | 0.697 (0.577 to 0.805) | 0.739 (0.588 to 0.870) | 0.810 (0.730 to 0.880) | 0.617 ± 0.068 |
| Text length only, domain 2 | same | 0.546 (0.443 to 0.644) | 0.657 (0.525 to 0.781) | 0.590 (0.490 to 0.690) | not re-run |
The last column is the mean ± standard deviation over ten re-drawn splits (seeds 90051 to 90060), each re-fitted from scratch. The domain-2 numbers in the other columns come from a split where a prompt shortcut is available (failure mode 5).
Paired comparison. On the same domain-1 texts, logistic regression minus SGD is +0.083 (+0.074 to +0.093) in macro-F1, and logistic regression was ahead in 10 of 10 re-drawn splits. On domain 2 the difference is not distinguishable from zero.
Spread across splits. Across ten re-drawn splits the domain-2 logistic regression's macro-F1 averaged 0.657 (SD 0.061, range 0.569 to 0.750), about the spread its bootstrap interval implies. The main split is the second most favourable of the ten, so the domain-2 headline numbers sit on the favourable side.
Intervals and shared prompts. The intervals resample texts as if they were independent, but texts that answer the same prompt are not. Resampling whole prompts instead (25,217 domain-1 test texts in 19,311 prompts) widens the domain-1 logistic regression's macro-F1 interval slightly, to 0.683 to 0.711 instead of 0.685 to 0.709. The 100 domain-2 test texts answer 82 prompts, and their intervals barely change.
Sensitivity to fitting. Fitted as in 2023, the domain-1 recipe stops early at C = 10,000 (see the fitting caveat above). Run to convergence at that C it scores macro-F1 0.688 (0.675 to 0.700); with C also chosen on grouped folds, 0.644 (0.634 to 0.654). In domain 2 the grouped folds push C to the bottom of the grid and macro-F1 falls to 0.584 (0.479 to 0.687).
Calibration (domain 1, expected calibration error over 10 bins).
| Probabilities | ECE |
|---|---|
| Logistic regression, as trained | 0.054 (0.051 to 0.056) |
| SGD with Platt scaling, as trained | 0.289 (0.286 to 0.292) |
| Logistic regression, prior-corrected | 0.023 (0.022 to 0.025) |
| SGD with Platt scaling, prior-corrected | 0.007 (0.005 to 0.009) |
| Logistic regression without re-balancing | 0.002 (0.002 to 0.004) |
Re-balancing teaches both detectors that half of all texts are human. Shifting their log-odds to the real 97.2% human share in domain 1 repairs most of the miscalibration.
Known failure modes
-
Base rates. At the domain-1 mix of 97 human texts to 3 machine texts, 1,343 of the 1,895 texts that logistic regression flags as machine-written are human. Only 0.291 (0.271 to 0.312) of its machine flags are correct. A flag from this detector is weak evidence on its own.
-
Length. A logistic regression on text length alone reaches AUC 0.932 (0.923 to 0.941) in domain 1. The full detector adds +0.027 (+0.019 to +0.036) AUC on the same texts, so most of its ranking could be done by counting tokens.
-
Unseen generators. Re-fitted without one domain-1 generator, logistic regression catches fewer of that generator's 140 test texts for all five generators, by 5 to 19 texts. The drop is significant for three of them (McNemar p ≤ 0.03), and AUC changes by at most 0.022, so the unseen generator's texts are still mostly ranked below human texts.
Generator Caught, generator seen in training Caught, generator left out McNemar exact 0 88 of 140 83 of 140 p = 0.27 1 98 of 140 87 of 140 p = 0.03 2 119 of 140 107 of 140 p = 0.002 3 109 of 140 90 of 140 p < 0.001 4 138 of 140 133 of 140 p = 0.13 Leaving a generator out does not hide its prompts: 118 to 140 of each generator's 140 test texts answer a prompt that the other generators' training texts also answer. With every training text that answers one of those prompts removed as well, leaving the generator out still lowers recall for all five, significantly for four.
Generator Prompts removed, generator kept Prompts removed, generator left out McNemar exact 0 78 of 140 69 of 140 p = 0.02 1 78 of 140 67 of 140 p = 0.04 2 113 of 140 94 of 140 p < 0.001 3 92 of 140 72 of 140 p < 0.001 4 135 of 140 130 of 140 p = 0.06 This is a loss of recall, not a collapse. Every generator in the data predates April 2023, so current models are unseen generators, and a detector like this has to be re-checked whenever a new one appears.
-
Domain shift. The domain-1 logistic regression scored on all 500 domain-2 texts reaches AUC 0.553 (0.483 to 0.621), not distinguishable from chance (the interval includes 0.5).
-
Prompt shortcuts. In domain 2 a prompt is answered either by one human text or by four machine texts, never both. In the main split all 80 machine test texts answer a prompt seen in training and none of the 20 human ones does, so "prompt seen in training" alone separates the classes; across the ten random splits it holds for 76 to 80 of the 80 machine texts. On ten splits that keep each prompt on one side, the domain-2 logistic regression averages AUC 0.613 ± 0.059 against 0.725 ± 0.059 on the ten random splits, and macro-F1 0.558 ± 0.039 against 0.657 ± 0.061 (seed 90051 alone: macro-F1 0.595 (0.484 to 0.701)). Domain 1 points the same way on a smaller scale: grouped macro-F1 0.676 ± 0.010 against 0.696 ± 0.004, and 0.673 (0.661 to 0.685) for seed 90051, below all ten random splits.
-
Small samples. Domain 2 rests on 500 texts, 100 of them human, and every domain-2 number has a wide interval.
-
Opacity. The texts are anonymised, so no error can be read, and the weights on
/weightsname token ids, not words.
Ethical considerations
- False accusations. Wrongly calling a person's writing machine-generated harms that person, and the base-rate failure above shows how easily that happens when machine text is rare. Any real detector needs a human decision, a way to contest it and evidence beyond a score.
- Unknown writers. Nothing is known about who wrote the human texts, so there is no way to check whether some groups of writers are flagged more often than others.
- Transparency. The weights, the evaluation and the code are public in this repository. This card is informed by the transparency principles of the Australian Government's policy for the responsible use of AI in government, the EU AI Act and the NIST AI Risk Management Framework. It does not claim compliance with any of them.
- Data. No course text leaves the repository, and the site serves only weights and aggregates.
Data statement
The course data
This statement describes the data behind the detectors on this site, following the spirit of data statements for NLP (Bender and Friedman, 2018). Much of what such a statement would normally record is unknown here, and the gaps are stated as gaps.
Source
The teaching team of COMP90051 Statistical Machine Learning (University of Melbourne, 2023
Semester 1) provided the data for Project 1, a class-wide Kaggle competition. The archive is
kept in the repository at coursework/data/sml.p1.data.tar.gz for reproducibility of the
original work. The course's README for the data is in coursework/data/README.md.
What one instance looks like
Each instance is a short text paired with the prompt that produced it. Both are given as
sequences of integers from 0 to 4,999. The course applied "light preprocessing" and then
mapped every word to an integer with one shared mapping. The mapping was never released, so
no word can be recovered. Machine-generated instances also carry a machine_id that names
the generator by number only.
Size and balance
| File | Domain | Label | Instances | Generators | Mean length (tokens) |
|---|---|---|---|---|---|
| set1_human.json | 1 | human | 122,584 | 380 | |
| set1_machine.json | 1 | machine | 3,500 | 5 × 700 | 183 |
| set2_human.json | 2 | human | 100 | 820 | |
| set2_machine.json | 2 | machine | 400 | 4 × 100 | 586 |
| test.json | 600 from domain 1, 400 from domain 2 | unlabelled | 1,000 | 409 |
Lengths count prompt and text together. Human texts in domain 1 are about twice as long as machine texts, which makes length alone a strong signal there (see the model card).
Texts share prompts. The 3,500 domain-1 machine texts answer 789 distinct prompts, and 64% of them have a prompt that also appears among the human texts. In domain 2 the 400 machine texts answer 100 prompts, four texts per prompt, and none of those prompts appears among the 100 human texts. A detector can therefore learn which prompts belong to machine texts, which is why the evaluation also reports a split that keeps each prompt on one side.
Unknowns
- Language and register. The course did not say which language the texts are in or where they came from. The two domains are described only as "different data sources or data distributions".
- Authors. Nothing is known about the people who wrote the human texts, so no claim can be made about which groups of writers the detectors serve well or badly.
- Generators. The machine texts come from five numbered generators in domain 1 and four in domain 2. Their names, versions and decoding settings are unknown. All of them predate April 2023.
- Collection dates and consent. Not documented in the material we received.
How the data is used here
scripts/export_artefacts.pyre-runs the 2023 notebooks on the archive and checks every reported number.scripts/evaluate_heldout.pyre-fits the 2023 recipes on an 80/20 split of the four training files and scores the held-out 20%.- The 1,000 test instances are used only to reproduce the eight 2023 Kaggle submissions. Their labels were never released.
What is published
The website and the files under web/ contain no instance, no token sequence and no
per-instance prediction. They hold:
- the six models' weights (5,000 coefficients and an intercept each);
- per domain and class, how often each token id occurs and the quantiles of text length;
- aggregate evaluation results: counts, confusion matrices, metrics with intervals, reliability-diagram bins and ROC curves on a fixed grid.
The synthetic "specimens" on /detector are drawn from those frequency profiles. They are
not course texts and do not resemble any single one.
Access
The data belongs to the course. It is kept in this private repository only to reproduce the original coursework and is not redistributed by the website. Please do not reuse it for current coursework.
AI use statement
Where AI is and is not used
This statement explains where artificial intelligence is and is not part of this site. It is informed by the transparency principles in the Australian Government's policy for the responsible use of AI in government, the EU AI Act's transparency provisions and the NIST AI Risk Management Framework. It is not a claim of compliance with any of them.
The detectors are not generative AI
The six detectors are linear models from scikit-learn, trained by our team in April 2023:
logistic regression and a linear support vector machine trained by stochastic gradient
descent. Each one adds up 5,000 learnt weights, one per token id, and compares the total with
zero. Their weights are published on /weights, and every score on the site can be
recomputed from them.
No generative-AI feature on the site
The site makes no call to a language model or any other AI service, needs no API key and sends nothing typed or generated on the page to a third party. All scoring runs in the visitor's browser. DR-005 explains why I did not add an optional bring-your-own-key feature.
AI assistance in building the 2026 revival
The 2026 web app, the analysis scripts and these documents were written with help from an AI coding assistant (Claude, by Anthropic). I reviewed the changes, and every number on the site comes from a script in the repository that anyone can re-run:
scripts/export_artefacts.pyreproduces the 2023 numbers and predictions;scripts/evaluate_heldout.pyproduces the held-out re-evaluation;scripts/stats_reference.pyproduces the reference values the statistics tests check.
The unit tests compare the TypeScript statistics against SciPy, statsmodels, scikit-learn and R, and a test checks that the numbers quoted in the model card match the evaluation file.
The 2023 coursework
The 2023 notebooks, report and meeting minutes are the team's original work. They are kept
unchanged under coursework/, apart from a formatting pass with no logic changes.
What a detector like this must never be used for
These detectors must not be used to decide whether a person's writing was produced by AI. The model card explains why: they were trained on anonymised data from generators that are now outdated, they catch fewer texts from generators they have not seen, they do not transfer between domains, and much of what they detect is text length.
Decision records
Why it is built this way
Each record states the decision first, then the context, the options, the reasons, what happened (including the weak numbers) and what I would change. Records are never edited after the fact. A later decision supersedes an earlier one.
- DR-001 · 2023-04 (recorded 2026-10)AcceptedBag-of-words counts and linear models instead of CNNsIn April 2023 we turned each text into 5,000 token counts and classified it with linear models, logistic regression and an SGD-trained linear SVM, and we dropped the word-embedding and CNN pipelines because they were slower, weaker on the leaderboard and scored by test splits we could not trust.
- DR-002 · 2023-04 (recorded 2026-10)AcceptedOne detector per domain, routed by the known domain of each test textWe fitted a separate detector for each domain, plus a pooled "1+2" detector for comparison, and the last submission file routed the 600 domain-1 test texts to the domain-1 model and the 400 domain-2 texts to the domain-2 model.
- DR-003 · 2023-04 (recorded 2026-10)AcceptedSMOTE over-sampling to re-balance the classesWe re-balanced every training set with SMOTE over-sampling rather than random under-sampling, then split it 80/20 for validation. As implemented, SMOTE copied minority texts instead of synthesising new ones, the copies landed on both sides of the split, and on clean data the re-balancing made the detector worse than no re-balancing.
- DR-004 · 2026-10AcceptedRe-fit the 2023 recipes on a split made before re-balancingTo measure how well the 2023 approach really works, I re-fit the 2023 recipes unchanged on an 80/20 split made before any re-balancing, score the untouched 20% with bootstrap intervals, and keep the 2023 pickles and printed numbers exactly as they were.
- DR-005 · 2026-10AcceptedNo generative-AI feature on this siteThe revival adds no language-model feature, not even an optional one that runs on the visitor's own API key, because the data is anonymised integer tokens that a language model cannot read, so any AI feature would be decoration rather than evidence.