Skip to content
All decision records

Decision record · DR-002

One detector per domain, routed by the known domain of each test text

Status
Accepted
Date
2023-04 (recorded 2026-10)
Applies to
coursework/notebooks/final, result8.csv, /detector, /evaluation#shift

Decision in one line

We fitted a separate detector for each domain, plus a pooled "1+2" detector for comparison, and the last submission file routed the 600 domain-1 test texts to the domain-1 model and the 400 domain-2 texts to the domain-2 model.

Context

The course described the two domains as different sources or distributions, and they differ in every way we could measure. Domain 1 has 122,584 human and 3,500 machine texts. Domain 2 has 100 human and 400 machine texts, so it is 250 times smaller and imbalanced the other way round. Domain-2 texts are about twice as long. The test file said which texts came from which domain: the first 600 from domain 1 and the last 400 from domain 2.

Decision

Fit logistic regression and SGD separately on domain 1, on domain 2 and on both together. Submit single-model files for comparison, and a routed file (result8.csv) that uses LR 1 for test items 0 to 599 and LR 2 for items 600 to 999.

Options considered

  1. One pooled model. Simple, but domain 1 outnumbers domain 2 by 250 to 1, so the pooled model would learn domain 1 and treat domain-2 texts as noise.
  2. One model per domain, routed by the known domain (chosen). Routing costs nothing because the test file tells us the domain.
  3. A pooled model with domain-specific weights, such as feature augmentation, where every token gets a shared weight plus one per domain. Domain 2 would borrow strength from domain 1 without being swamped. We did not know the technique at the time.
  4. Train on domain 1 and adapt to domain 2 by importance weighting or fine-tuning. Too much machinery for 500 texts and a three-week deadline.

Why

  • The domains looked different enough that one decision rule seemed unlikely to fit both. In the exported weights, LR 1 and LR 2 correlate at r = 0.01 across the 5,000 token ids.
  • With routing free at test time, separate models carried no deployment cost.

What happened

  • 2023. The pooled models were broken by a copy-paste slip that paired domain-1 features with domain 1+2 labels (see the revival audit on /journal), so the 2023 numbers cannot compare pooled and per-domain models. The report's 0.50 for the pooled models measured that slip, not pooling.
  • 2026, clean split (DR-004). On the 100 held-out domain-2 texts, the domain-2 LR scores macro-F1 0.740 (0.610 to 0.850) and a correctly pooled LR 0.605 (0.493 to 0.709). On the same texts the domain-2 model is right where the pooled one is wrong 18 times, and the reverse happens 4 times (McNemar exact p = 0.004). The accuracy difference is +0.140 (+0.050 to +0.230), while the macro-F1 difference, +0.135 (−0.014 to +0.273), includes zero. On domain 1 the dedicated model is ahead on every measure, by +0.040 (+0.032 to +0.048) macro-F1.
  • Domain shift. The domain-1 LR scored on all 500 domain-2 texts reaches AUC 0.553 (0.483 to 0.621), not distinguishable from chance (the interval includes 0.5). The domain-2 LR on domain 1 reaches AUC 0.639 (0.623 to 0.655) and calls almost every human text machine-written. A detector trained on one domain does not transfer to the other.
  • The weak number. The domain-2 test part has only 20 human texts, and the macro-F1 interval for the domain-2 model is 0.24 wide. Across ten different splits its macro-F1 ranged from 0.569 to 0.750 (mean 0.657), and the main split is the second most favourable of the ten. In domain 2 the machine texts answer prompts that no human text answers. On ten splits that keep each prompt on one side, the domain-2 LR averages macro-F1 0.558 ± 0.039 (seed 90051: 0.595 (0.484 to 0.701)). Part of its score came from recognising machine prompts. Per-domain models beat pooling on domain 1 on every measure and on domain 2 in accuracy, but how good the domain-2 model is remains poorly determined.

What I'd change

  • Try feature augmentation or a hierarchical model that shrinks the domain-2 weights towards domain 1, and compare it with routing on the same splits.
  • Use repeated cross-validation for domain 2 instead of one 80/20 split, so 500 texts are all used for testing once.
  • Ask for more domain-2 data before tuning anything on it.