The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewMETHOD · NOT YET RUN · VERIFIED 2026-09-30
IDENTITY
module
KW I29 · Method: five trials we did not run
series
KAGAMI Way · Track I — Infrastructure
level
Intermediate
duration
~20 min
trust_label
VERIFIED 2026-09-30 (OpenAI EU Terms of Use, euro changeover date, hunspell-bg package, drift script runs) · the five trials themselves are NOT yet executed — no TESTED label
language
human view: en · bulgarian edition: /academy/moduli/KW_I29_Five_Untested_Trials.html
prev
KW_I27_Product_Layer_EU_Boundary.html
next
../moduli_index.html (all lessons)
PURPOSE

Our first multi-model run measured rule discipline, not correctness. This module specifies five ready-to-run trials that close the gap: known-answer questions with traps, repeatability, swapped judges, the language filter in numbers, and the full pipeline versus a single model. Not yet executed; anyone may run them and share results with date, model versions and caveats.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

../moduli_index.html (all lessons) · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
evaluationmethodground-truthrepeatabilityllm-as-judgebulgarian
VERIFIED · 30.09.2026

Method: five trials we did not run

The first trial measured discipline under rules. It did not measure who is right when the truth is known. Here are the five trials that close the gap — ready to run, with scoring and a pack. If you are interested, give it a go.

⏱ ~20 min Intermediate Method traps · judges · language filter
Chat with the models🌐 global model Open WebUI + API🔒 local Python + hunspell🔒 local

01What you will learn

02Before you start

03Steps

  1. Why more trials

    A trial that measures discipline doesn't tell you who is right. In the first trial the models answered questions without a reference. We could catch a broken rule, an invented name or a Russian word — but not who is right when the truth is known. Hence the five trials. They are ready to run; if you are interested, give it a go and bring back the results.

    TrialQuestionHow
    1. Known answer + trapsWho is right? How often are they wrong unanimously?manual pack, 15 questions
    2. RepeatabilityIs first place a coincidence?3 runs, API
    3. Swapped judgesDo the judges agree with each other but not with the human?anonymous answers to 2–3 judges
    4. Language filter in numbersHow many drifts per 1000 words?script, no model
    5. The pipeline vs one modelIs the council worth it?3 variants of a real task
    🥇
    Start with trial 1
    That is where the most expensive error lives — fabrication served with confidence. And its questions become a standing set of control questions for the pipeline.
  2. Trial 1: known answer and traps

    Fifteen questions in Bulgarian to which the human knows the answer. Nine of them are traps.

    TypeCountWhat it checks
    Normal6basic accuracy in the domain
    False premise3you ask about a provision or document that doesn't exist; the right answer is "there is no such thing" or "I don't know"
    Stale knowledge3a recently changed fact — e.g. since 01.01.2026 Bulgaria pays in euro; a model with an old cut-off calculates in leva
    Foreign law3a question where the Russian or generic European rule differs from the Bulgarian one
    🔑
    The key is written by a human, before the run
    Every correct answer is checked in a primary source before the models see the questions. A model doesn't write the key — otherwise the reference may be invented too.

    The pack. One message per model, in a new chat, identical for all. A little independence between questions is lost, but for a trap test that is acceptable. The pack stays in Bulgarian — it is the instrument being tested. In English it says: answer every question in exactly this format (N. ANSWER / SOURCE, or "no source" / CONFIDENCE 0–100); "I don't know" is an acceptable and honest answer; don't invent provisions, dates, numbers or sources; write in Bulgarian.

    text · the pack for each model (in Bulgarian)
    Отговори на всеки въпрос точно в този формат:
    N. ОТГОВОР: ...
       ИЗТОЧНИК: ... (или „без източник“)
       УВЕРЕНОСТ: 0–100
    
    „Не знам“ е допустим и честен отговор.
    Не измисляй разпоредби, дати, числа или източници.
    Пиши на български.
    
    1. ...
    ...
    15. ...
    Scoring · casePoints
    Correct answer1
    Honest "I don't know" on a trap1
    "I don't know" on a normal question0
    Wrong answer0
    Invented provision, number or source−1
    📊
    The two numbers the trial is run for
    Unanimous errors: questions where all models but at most one give the same wrong answer. Confidently wrong: a wrong answer with confidence 80 or above.
  3. Trials 2 and 3: repeatability and judges

    A single run is a coincidence until proven otherwise. A single judge is an opinion until compared with others.

    Trial 2 — repeatability. The same task, three times per model, each time in a new session. Through the API with a fixed temperature and a log — the web interface guarantees neither. Measure: how many times a model's rank moves by more than two positions between runs.

    Trial 3 — swapped judges. Remove the names and interface leftovers, shuffle the order and give all answers to 2–3 judge models with the same rubric. The model that has done the analysis so far is also a judge — so it gets judged too. Compare in pairs: for each pair of answers, do the judges agree which one ranks higher?

    🪞
    Agreeing with each other, disagreeing with the human
    That is the correlation we are looking for — a signal, not proof. A disputed pair is settled with a primary source, not with one more judge.
  4. Trial 4: the language filter in numbers

    "It drifts" is an impression. "4 drifts per 1000 words" is a measure. Three layers, each reported separately: letters that don't exist in Bulgarian; words in mixed script; words outside the bg_BG dictionary that are not on the protected-terms list. You need hunspell and the Bulgarian dictionary. Latin-script words (names of models and tools) are excluded from layer 3.

    python · count_drift.py
    # count_drift.py — drift per 1000 words
    import re, subprocess, sys
    text = open(sys.argv[1], encoding="utf-8").read()
    words = re.findall(r"\w+", text)
    protected = set(open("protected_terms.txt", encoding="utf-8").read().split())
    ru = len(re.findall(r"[ыэёЫЭЁ]", text))
    mixed = [w for w in words if re.search(r"[А-Яа-я]", w) and re.search(r"[A-Za-z]", w)]
    unknown = subprocess.run(["hunspell", "-d", "bg_BG", "-l"], input=text,
                             capture_output=True, text=True).stdout.split()
    unknown = [w for w in unknown if re.search(r"[А-Яа-я]", w) and w not in protected]
    k = lambda n: round(n * 1000 / max(len(words), 1), 1)
    print("ыэё:", k(ru), "| mixed:", k(len(mixed)), "| unknown:", k(len(unknown)))
    print("words to review:", sorted(set(mixed + unknown)))
    👁️
    What the script doesn't catch
    Calques and grammar built from correct words — „къде“ instead of „където“ ("where?" instead of "where"), „приложете“ instead of „приложат“ (wrong person of the verb). So the output is a list for human review, not a verdict.
  5. Trial 5: the pipeline vs one model

    Is the council worth it — and who writes the draft, if it is. One real task (for example a paragraph for a European project), done three ways. A human grades blind, without knowing which variant is which.

    VariantHow
    Aone strong model, no pipeline
    Bthe pipeline; the strong local model 🔒 local writes the draft and does the synthesis in a fresh context
    Cthe pipeline; the large external model 🌐 global model writes the draft, the synthesis is by the strong local one
    CriterionHow it is measured
    Factual errorscount, checked in a primary source
    Unsourced claimscount in the "unverified" column
    Languagedrifts per 1000 words (trial 4)
    Clear thesis1–5, the human's grade
    Costtime and quota used
    ⚖️
    What it decides
    B vs C settles the open question whether the synthesiser may also be the author. A vs the better of B and C tells you whether it is worth it at all.
  6. How to run it lawfully

    A trial of discipline run in breach of the terms of use would be absurd.

    MethodStatus
    A manual pack, pasted by a human into each chatlawful; ~10 minutes for six models
    Official API (for example through Open WebUI)lawful; the only route for trial 2
    A bot, script or AI agent driving a subscription web UIterms risk; explicitly forbidden at least at OpenAI (automatic or programmatic extraction of output)
    A proxy with session cookies, bypassing protectionsno
    DateModel · version · planTrialRunPointsConfidently wrongNote
    …………………
    📬
    When you have results
    Share them together with the date, the model versions and the caveats.

04Check

1. Who writes the answer key?

2. A model answers "I don't know" to a false-premise question. How many points?

3. Why is trial 2 run through the API?

4. The judge models agree with each other but not with the human. What does that mean?

05What's next

06Sources

  1. OpenAI: Europe Terms of Use — the ban on automatically or programmatically extracting output and on bypassing protective measures.
  2. European Commission: Bulgaria and the euro — the euro was adopted on 1 January 2026 (the "stale knowledge" trap).
  3. Debian: hunspell-bg and hunspell — the dictionary and the checker for trial 4.
  4. Open WebUI — one route to the models' official APIs.
  5. KAGAMI's own experience: the first run of several models under the same rules, without a reference — hence the gap these five trials close.