Skip to content
All decision records

Decision record · DR-005

No generative-AI feature on this site

Status
Accepted
Date
2026-10
Applies to
web/, docs/ai-use-statement.md

Decision in one line

The revival adds no language-model feature, not even an optional one that runs on the visitor's own API key, because the data is anonymised integer tokens that a language model cannot read, so any AI feature would be decoration rather than evidence.

Context

The other revived coursework sites share an optional pattern: the visitor pastes their own Anthropic or OpenAI key, the browser calls the provider directly, every output is labelled and logged, and where an LLM competes with the original method an evaluation harness scores both on the same inputs. The pattern is useful where an LLM can do the task, for example guessing letters in Hangman or answering questions about a dataset. In this project every word was replaced by an integer before release, and the mapping was never published.

Decision

Ship no AI feature. State that plainly in the AI use statement, and spend the effort on the statistical re-evaluation instead (DR-004).

Options considered

  1. An LLM as a third detector, scored on the held-out texts. An LLM sees only lists of numbers such as "4021 17 993", so it would be guessing from length and token frequency, which the linear models already use. Sending course data to a provider would also break the rule that no course instance leaves the repository.
  2. An LLM that explains a detector's verdict. It would explain token ids it cannot interpret. The explanation would sound plausible and be unverifiable, which is worse than no explanation for a tool about detecting machine text.
  3. An LLM that writes new specimens for the detector. A modern model writes words, and the detectors only accept ids from a vocabulary nobody has, so its output could not be scored.
  4. No AI feature (chosen).

Why

  • A feature should produce evidence about the question the site asks. None of the options above can.
  • No key handling, no audit log and no provider calls means less code and nothing that could leak a visitor's key.
  • Deciding not to use an LLM where it does not fit is part of responsible AI use, and the AI use statement says so.

What happened

  • The site makes no network request to any AI provider, and all scoring stays in the browser.
  • The AI use statement explains that the detectors are linear models and that the 2026 code was written with an AI coding assistant under my review.
  • The model card states that these detectors must not be used to judge whether a person's writing is AI-generated.

What I'd change

  • If the course ever released the word mapping, or a comparable labelled dataset of raw text were used, an LLM-as-detector evaluation with paired comparisons against the linear models would become meaningful. I would build it with the same bring-your-own-key and audit-log pattern as the other sites.