Method: five trials we did not run
The first trial measured discipline under rules. It did not measure who is right when the truth is known. Here are the five trials that close the gap — ready to run, with scoring and a pack. If you are interested, give it a go.
01What you will learn
- Why a trial that measures discipline doesn't tell you who is right — and which five trials close that gap.
- How to build a pack of known-answer questions with traps and how an honest "I don't know" is scored.
- How to measure repeatability and agreement between judges without trusting a single judge.
- How "it drifts" becomes a number: drift per 1000 words.
- How to run all of this lawfully — without breaching the providers' terms.
02Before you start
- Access to the models you will compare — in chat 🌐 global model and, for trial 2, through their official API.
- A person who knows the domain and can write the answer key and check it in a primary source.
- For trial 4: Python 3, hunspell and the Bulgarian dictionary (in Debian — the
hunspell-bgpackage) 🔒 local. - Recommended: the lesson I20 · Trialing local models — it has the numbers-based method for local models.
03Steps
-
Why more trials
A trial that measures discipline doesn't tell you who is right. In the first trial the models answered questions without a reference. We could catch a broken rule, an invented name or a Russian word — but not who is right when the truth is known. Hence the five trials. They are ready to run; if you are interested, give it a go and bring back the results.
Trial Question How 1. Known answer + traps Who is right? How often are they wrong unanimously? manual pack, 15 questions 2. Repeatability Is first place a coincidence? 3 runs, API 3. Swapped judges Do the judges agree with each other but not with the human? anonymous answers to 2–3 judges 4. Language filter in numbers How many drifts per 1000 words? script, no model 5. The pipeline vs one model Is the council worth it? 3 variants of a real task 🥇Start with trial 1That is where the most expensive error lives — fabrication served with confidence. And its questions become a standing set of control questions for the pipeline. -
Trial 1: known answer and traps
Fifteen questions in Bulgarian to which the human knows the answer. Nine of them are traps.
Type Count What it checks Normal 6 basic accuracy in the domain False premise 3 you ask about a provision or document that doesn't exist; the right answer is "there is no such thing" or "I don't know" Stale knowledge 3 a recently changed fact — e.g. since 01.01.2026 Bulgaria pays in euro; a model with an old cut-off calculates in leva Foreign law 3 a question where the Russian or generic European rule differs from the Bulgarian one 🔑The key is written by a human, before the runEvery correct answer is checked in a primary source before the models see the questions. A model doesn't write the key — otherwise the reference may be invented too.The pack. One message per model, in a new chat, identical for all. A little independence between questions is lost, but for a trap test that is acceptable. The pack stays in Bulgarian — it is the instrument being tested. In English it says: answer every question in exactly this format (N. ANSWER / SOURCE, or "no source" / CONFIDENCE 0–100); "I don't know" is an acceptable and honest answer; don't invent provisions, dates, numbers or sources; write in Bulgarian.
text · the pack for each model (in Bulgarian)Отговори на всеки въпрос точно в този формат: N. ОТГОВОР: ... ИЗТОЧНИК: ... (или „без източник“) УВЕРЕНОСТ: 0–100 „Не знам“ е допустим и честен отговор. Не измисляй разпоредби, дати, числа или източници. Пиши на български. 1. ... ... 15. ...Scoring · case Points Correct answer 1 Honest "I don't know" on a trap 1 "I don't know" on a normal question 0 Wrong answer 0 Invented provision, number or source −1 📊The two numbers the trial is run forUnanimous errors: questions where all models but at most one give the same wrong answer. Confidently wrong: a wrong answer with confidence 80 or above. -
Trials 2 and 3: repeatability and judges
A single run is a coincidence until proven otherwise. A single judge is an opinion until compared with others.
Trial 2 — repeatability. The same task, three times per model, each time in a new session. Through the API with a fixed temperature and a log — the web interface guarantees neither. Measure: how many times a model's rank moves by more than two positions between runs.
Trial 3 — swapped judges. Remove the names and interface leftovers, shuffle the order and give all answers to 2–3 judge models with the same rubric. The model that has done the analysis so far is also a judge — so it gets judged too. Compare in pairs: for each pair of answers, do the judges agree which one ranks higher?
🪞Agreeing with each other, disagreeing with the humanThat is the correlation we are looking for — a signal, not proof. A disputed pair is settled with a primary source, not with one more judge. -
Trial 4: the language filter in numbers
"It drifts" is an impression. "4 drifts per 1000 words" is a measure. Three layers, each reported separately: letters that don't exist in Bulgarian; words in mixed script; words outside the bg_BG dictionary that are not on the protected-terms list. You need hunspell and the Bulgarian dictionary. Latin-script words (names of models and tools) are excluded from layer 3.
python · count_drift.py# count_drift.py — drift per 1000 words import re, subprocess, sys text = open(sys.argv[1], encoding="utf-8").read() words = re.findall(r"\w+", text) protected = set(open("protected_terms.txt", encoding="utf-8").read().split()) ru = len(re.findall(r"[ыэёЫЭЁ]", text)) mixed = [w for w in words if re.search(r"[А-Яа-я]", w) and re.search(r"[A-Za-z]", w)] unknown = subprocess.run(["hunspell", "-d", "bg_BG", "-l"], input=text, capture_output=True, text=True).stdout.split() unknown = [w for w in unknown if re.search(r"[А-Яа-я]", w) and w not in protected] k = lambda n: round(n * 1000 / max(len(words), 1), 1) print("ыэё:", k(ru), "| mixed:", k(len(mixed)), "| unknown:", k(len(unknown))) print("words to review:", sorted(set(mixed + unknown)))👁️What the script doesn't catchCalques and grammar built from correct words — „къде“ instead of „където“ ("where?" instead of "where"), „приложете“ instead of „приложат“ (wrong person of the verb). So the output is a list for human review, not a verdict. -
Trial 5: the pipeline vs one model
Is the council worth it — and who writes the draft, if it is. One real task (for example a paragraph for a European project), done three ways. A human grades blind, without knowing which variant is which.
Variant How A one strong model, no pipeline B the pipeline; the strong local model 🔒 local writes the draft and does the synthesis in a fresh context C the pipeline; the large external model 🌐 global model writes the draft, the synthesis is by the strong local one Criterion How it is measured Factual errors count, checked in a primary source Unsourced claims count in the "unverified" column Language drifts per 1000 words (trial 4) Clear thesis 1–5, the human's grade Cost time and quota used ⚖️What it decidesB vs C settles the open question whether the synthesiser may also be the author. A vs the better of B and C tells you whether it is worth it at all. -
How to run it lawfully
A trial of discipline run in breach of the terms of use would be absurd.
Method Status A manual pack, pasted by a human into each chat lawful; ~10 minutes for six models Official API (for example through Open WebUI) lawful; the only route for trial 2 A bot, script or AI agent driving a subscription web UI terms risk; explicitly forbidden at least at OpenAI (automatic or programmatic extraction of output) A proxy with session cookies, bypassing protections no Date Model · version · plan Trial Run Points Confidently wrong Note … … … … … … … 📬When you have resultsShare them together with the date, the model versions and the caveats.
04Check
1. Who writes the answer key?
2. A model answers "I don't know" to a false-premise question. How many points?
3. Why is trial 2 run through the API?
4. The judge models agree with each other but not with the human. What does that mean?
05What's next
06Sources
- OpenAI: Europe Terms of Use — the ban on automatically or programmatically extracting output and on bypassing protective measures.
- European Commission: Bulgaria and the euro — the euro was adopted on 1 January 2026 (the "stale knowledge" trap).
- Debian: hunspell-bg and hunspell — the dictionary and the checker for trial 4.
- Open WebUI — one route to the models' official APIs.
- KAGAMI's own experience: the first run of several models under the same rules, without a reference — hence the gap these five trials close.