AI 测验生成器也需要测试:用无依赖 JavaScript linter 校验生成题目
AI quiz generators need tests, too
作者构建了一个无依赖的 JavaScript linter(quiz-lint.mjs),用于校验 AI 测验生成器输出的单选题 JSON,可发现缺失或多余的 answerIndexes、越界索引、去空格后重复选项等结构问题,并用启发式规则标记正确答案明显偏长的题目供人工复核。
A quiz generator can return valid JSON containing two answer keys, repeated options, or a confidently wrong answer. Put a small, testable validation step between generation and editorial review.
The useful output is a list of specific findings. Separate broken data from questions that deserve a closer look. This tutorial builds that distinction into a dependency-free JavaScript linter. It cannot certify factual accuracy, clinical safety, or assessment quality.
Define the item contract first
This example accepts one single-answer item:
-
stem: a nonempty string -
options: at least two nonempty plain-text strings -
answerIndexes: an array containing exactly one zero-based option index
The array makes missing and extra key entries explicit. Two options is a minimal software constraint here, not a recommendation about assessment design. Multi-select questions need a different contract; multiple keys would be legitimate there.
Structural findings include an absent key, several keys, an out-of-range index, and duplicate option text after trimming surrounding whitespace. Case and punctuation stay intact. The code does not attempt semantic deduplication: “store a copy” and “save a duplicate” could require human review even though their strings differ.
Keep the heuristic separate
NBME's item-writing guide describes a flaw in which the keyed answer stands out through extra length and detail. Its advice is to review the option set and remove unnecessary instructional wording. That supports a review prompt, not a universal numerical cutoff. See “Correct Option Stands Out”.
For this demonstration, flag the keyed option only when it has at least twice as many whitespace-delimited tokens as every distractor, with a gap of at least eight tokens from the longest distractor. Both thresholds are arbitrary and unvalidated. They are not NBME criteria or a measure of item quality.
A legitimate answer may need more words. A poor item may have perfectly balanced options. This rule should create a review candidate, preserve exceptions, and never rewrite or reject content automatically. It is intended for English-like prose; code, equations, and languages without space-delimited words need different handling.
Run this after parsing the generated JSON, with application-level limits on input size. Save it as quiz-lint.mjs:
// Intended for parsed JSON: single-answer items with plain-text options.
export function lintItem(item) {
const problem = (code) => [{ kind: 'structure', code }];
const nonempty = (s) => typeof s === 'string' && s.trim().length > 0;
if (!item || typeof item !== 'object' || Array.isArray(item)) {
return problem('invalid_item');
}
if (!nonempty(item.stem)) return problem('invalid_stem');
const options = item.options;
if (!Array.isArray(options) || options.length < 2 ||
[...options].some((s) => !nonempty(s))) {
return problem('invalid_options');
}
const issues = [];
const keys = item.answerIndexes; // Zero-based; exactly one is required.
if (!Array.isArray(keys) || keys.length !== 1) {
issues.push({ kind: 'structure', code: 'key_count' });
} else if (!Number.isInteger(keys[0]) || keys[0] < 0 ||
keys[0] >= options.length) {
issues.push({ kind: 'structure', code: 'key_index' });
}
const labels = options.map((s) => s.trim()); // Preserve case and punctuation.
if (new Set(labels).size !== labels.length) {
issues.push({ kind: 'structure', code: 'duplicate_options' });
}
if (issues.length) return issues; // No heuristic on malformed input.
// Arbitrary, unvalidated thresholds for English-like plain text only.
const words = labels.map((s) => s.split(/\s+/u).length);
const keyedWords = words[keys[0]];
const otherMax = words.reduce(
(max, count, i) => i === keys[0] ? max : Math.max(max, count), 0);
if (keyedWords >= 2 * otherMax && keyedWords - otherMax >= 8) {
issues.push({ kind: 'review', code: 'keyed_length_outlier',
keyedWords, otherMax });
}
return issues;
}
An empty findings array means only that these checks found nothing. Notice that the code trusts the supplied key; “keyed” does not establish that an answer is actually correct. Structural problems stop the heuristic so malformed data cannot receive a misleading content-review result.
Test what it catches and what it misses
Save the following beside it as quiz-lint.test.mjs. Every example is constructed for this tutorial. Node provides the test runner and strict assertions, so no package installation is needed.
import test from 'node:test';
import assert from 'node:assert/strict';
import { lintItem } from './quiz-lint.mjs';
const make = (options, answerIndexes = [0]) => ({
stem: 'In this toy menu, save means store a copy. Which label matches?',
options, answerIndexes,
});
const labels = ['Store a copy', 'Remove a copy', 'Rename a copy'];
const codes = (item) => lintItem(item).map((i) => `${i.kind}:${i.code}`);
test('ordinary fixture has no findings', () => {
assert.deepEqual(lintItem(make(labels)), []);
});
test('missing, empty, or multiple keys need structural repair', () => {
const missing = { stem: 'Choose a label.', options: labels };
for (const item of [missing, make(labels, []), make(labels, [0, 1])]) {
assert.deepEqual(codes(item), ['structure:key_count']);
}
});
test('trimmed duplicate labels are structural findings', () => {
assert.deepEqual(codes(make(['Store a copy', ' Store a copy '])),
['structure:duplicate_options']);
});
test('a keyed length outlier requests review', () => {
const item = make([
'Store a copy in the selected local folder without changing the original file',
'Remove a copy', 'Rename a copy',
]);
assert.deepEqual(lintItem(item), [{ kind: 'review',
code: 'keyed_length_outlier', keyedWords: 13, otherMax: 3 }]);
});
test('a long distractor does not trigger the keyed-option rule', () => {
assert.deepEqual(lintItem(make([
'Store a copy',
'Remove a copy from the selected local folder without changing the original file',
])), []);
});
test('even a factually wrong key can produce no findings', () => {
assert.deepEqual(lintItem({ stem: 'What is 2 + 2?',
options: ['3', '4'], answerIndexes: [0] }), []);
});
Run node --test quiz-lint.test.mjs in that directory. These six tests pass on Node v24.19.0.
The final fixture is especially important: it deliberately keys “3” for “What is 2 + 2?” and produces no findings. Keep a test like this to make the validator's boundary visible to future maintainers. Otherwise, a green build can gradually acquire a meaning the code never supported.
For a larger test suite, add invalid input types, fractional and negative indexes, exact threshold boundaries, tied option lengths, and a check that the function does not mutate its input.
Connect findings to an editorial workflow
A practical integration can treat structure findings as a request to repair the data before continuing. Send review findings to an editor with the item, observed token counts, source material, and rationale. Keep the answer key and editorial findings out of learner-facing question payloads.
Store the rule version and item revision with each decision. If a reviewer accepts a length difference, record why it is necessary. Reopen that decision when the item changes. Reviewers should be able to preserve a justified exception without deleting evidence or padding distractors just to satisfy a ratio.
Several essential checks remain outside this function:
- Source accuracy. Verify that a reliable, relevant source supports the key and explanation. Save the supporting passage or location and its version or date. A working citation link alone is insufficient.
- Plausible distractors. Have a subject expert check whether the alternatives make sense for the question and intended learners. NBME's guidance calls for plausible, comparable options; token counts cannot establish those properties. See rule 4.
- Independent performance. If the product's goal is learning, evaluate that separately. Use fresh, expert-reviewed questions answered without hints or generated explanations visible. Keep first attempts separate from retries after feedback. This linter provides no evidence of improved learning or unaided performance.
For clinical educational content, qualified subject-matter review remains necessary; this toy validator is not a clinical validation method. More generally, run lint alongside expert review and evaluation appropriate to the assessment's purpose. Give each check a narrow, explicit meaning, and test that boundary as carefully as the happy path.
来源:Google AI:DEV 作者专属(RSS) · dev.to