Decision record · DR-002
One detector per domain, routed by the known domain of each test text
- Status
- Accepted
- Date
- 2023-04 (recorded 2026-10)
- Applies to
- coursework/notebooks/final, result8.csv, /detector, /evaluation#shift
Decision in one line
We fitted a separate detector for each domain, plus a pooled "1+2" detector for comparison, and the last submission file routed the 600 domain-1 test texts to the domain-1 model and the 400 domain-2 texts to the domain-2 model.
Context
The course described the two domains as different sources or distributions, and they differ in every way we could measure. Domain 1 has 122,584 human and 3,500 machine texts. Domain 2 has 100 human and 400 machine texts, so it is 250 times smaller and imbalanced the other way round. Domain-2 texts are about twice as long. The test file said which texts came from which domain: the first 600 from domain 1 and the last 400 from domain 2.
Decision
Fit logistic regression and SGD separately on domain 1, on domain 2 and on both together.
Submit single-model files for comparison, and a routed file (result8.csv) that uses LR 1
for test items 0 to 599 and LR 2 for items 600 to 999.
Options considered
- One pooled model. Simple, but domain 1 outnumbers domain 2 by 250 to 1, so the pooled model would learn domain 1 and treat domain-2 texts as noise.
- One model per domain, routed by the known domain (chosen). Routing costs nothing because the test file tells us the domain.
- A pooled model with domain-specific weights, such as feature augmentation, where every token gets a shared weight plus one per domain. Domain 2 would borrow strength from domain 1 without being swamped. We did not know the technique at the time.
- Train on domain 1 and adapt to domain 2 by importance weighting or fine-tuning. Too much machinery for 500 texts and a three-week deadline.
Why
- The domains looked different enough that one decision rule seemed unlikely to fit both. In the exported weights, LR 1 and LR 2 correlate at r = 0.01 across the 5,000 token ids.
- With routing free at test time, separate models carried no deployment cost.
What happened
- 2023. The pooled models were broken by a copy-paste slip that paired domain-1 features
with domain 1+2 labels (see the revival audit on
/journal), so the 2023 numbers cannot compare pooled and per-domain models. The report's 0.50 for the pooled models measured that slip, not pooling. - 2026, clean split (DR-004). On the 100 held-out domain-2 texts, the domain-2 LR scores macro-F1 0.740 (0.610 to 0.850) and a correctly pooled LR 0.605 (0.493 to 0.709). On the same texts the domain-2 model is right where the pooled one is wrong 18 times, and the reverse happens 4 times (McNemar exact p = 0.004). The accuracy difference is +0.140 (+0.050 to +0.230), while the macro-F1 difference, +0.135 (−0.014 to +0.273), includes zero. On domain 1 the dedicated model is ahead on every measure, by +0.040 (+0.032 to +0.048) macro-F1.
- Domain shift. The domain-1 LR scored on all 500 domain-2 texts reaches AUC 0.553 (0.483 to 0.621), not distinguishable from chance (the interval includes 0.5). The domain-2 LR on domain 1 reaches AUC 0.639 (0.623 to 0.655) and calls almost every human text machine-written. A detector trained on one domain does not transfer to the other.
- The weak number. The domain-2 test part has only 20 human texts, and the macro-F1 interval for the domain-2 model is 0.24 wide. Across ten different splits its macro-F1 ranged from 0.569 to 0.750 (mean 0.657), and the main split is the second most favourable of the ten. In domain 2 the machine texts answer prompts that no human text answers. On ten splits that keep each prompt on one side, the domain-2 LR averages macro-F1 0.558 ± 0.039 (seed 90051: 0.595 (0.484 to 0.701)). Part of its score came from recognising machine prompts. Per-domain models beat pooling on domain 1 on every measure and on domain 2 in accuracy, but how good the domain-2 model is remains poorly determined.
What I'd change
- Try feature augmentation or a hierarchical model that shrinks the domain-2 weights towards domain 1, and compare it with routing on the same splits.
- Use repeated cross-validation for domain 2 instead of one 80/20 split, so 500 texts are all used for testing once.
- Ask for more domain-2 data before tuning anything on it.