The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
01-02b · Structured outputs and prompt evaluation
series
Blocks 0–10 · Block 2 — Prompt Engineering & Evals · Part 2/2
level
Intermediate
duration
3–4 h
prerequisites
01-02a (prompt engineering) · Python 3.10+ · a local model server (Ollama or vLLM)
trust_label
VERIFIED 2026-10-01 (library versions and APIs checked against PyPI packages and official docs; Pydantic schema executed locally) · UPDATED 2026-10-01 · LLM calls NOT run end to end
language
human view: en · bulgarian edition: /academy/blokove/moduli/01-02b_Блок_2_Част_2_Eval_Structured_Outputs.html
next
01-03a_Блок_3_Част_1_Агенти_LangGraph.html · Block 3: AI agents and orchestration
PURPOSE

Turn free-text LLM output into validated, typed data (JSON schema, Pydantic, Instructor) and turn prompt changes into measured decisions (fixed test set, A/B comparison, LLM-as-a-judge, RAGAS / deepeval metrics, regression gate in CI).

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

01-03a_Блок_3_Част_1_Агенти_LangGraph.html · Block 3: AI agents and orchestration (LangGraph, CrewAI, n8n) · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
structured-outputsjson-schemapydanticinstructorollamavllmevaluationllm-as-judgeragasdeepevalregression-testing
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

Structured outputs and prompt evaluation

A model naturally writes free text, while the program on the other side expects exact fields. In the first half we turn output into validated data with a JSON schema, Pydantic and Instructor. In the second half we stop guessing which prompt is better and start measuring: an A/B test, a judge model, RAGAS and a regression test before every release.

⏱ 3–4 h Intermediate Block 2 · Part 2/2 structured outputs · evaluation
Ollama · vLLM🔒 local Pydantic · Instructor🔒 local RAGAS · deepeval🔒 local (with a local model) Cloud model as judge (optional)🌐 global
🔄
UPDATED · 01.10.2026 — what changed
Libraries were checked against PyPI and the official docs: Pydantic 2.13.5, Instructor 1.17.0, RAGAS 0.4.3, deepeval 4.2.7, Outlines 1.3.3, ollama-python 0.6.3. A new rung was added: a JSON schema with constrained generation directly in the server. Ollama accepts a schema in format, vLLM in response_format. The old guided_json parameters were removed from vLLM in v0.12.0. The RAGAS example was rewritten for the new API (ragas.metrics.collections and ascore), because evaluate() and LangchainLLMWrapper are now marked deprecated. deepeval (GEval) was added. Instructor's exception is now imported from instructor.core. Invoices are in euro: Bulgaria has used the euro since 01.01.2026. The total check now sits in a model_validator and was run locally. The reliability percentages per method and the "real result" of an accounting firm could not be verified, so they were removed or marked as illustrative. Test data now comes from a fictional internal rulebook instead of "regulations" citing non-existent articles.

01What you will learn

02Before you start

03Steps

  1. Why free text breaks systems

    Imagine an invoice analyser built on text.split("Amount:"). It works for most invoices. For the rest it either crashes or silently writes a wrong amount into the books. Why is that dangerous? Because the error stays invisible until someone compares the numbers. Structured output moves the check to where it belongs: into the schema, before the data travels further.

    RungWhat it guaranteesWhen
    Instruction
    "Return only JSON"
    Nothing. The model often adds explanations or Markdown fences.Quick prototype
    JSON mode
    format="json" · {"type":"json_object"}
    Syntactically valid JSON, but with whatever fields the model chooses.Simple structures
    JSON schema
    format=schema · json_schema
    The server constrains generation to follow the schema: names, types, enums.Almost always — today's default
    Schema + Pydantic + retry
    Instructor
    The shape and the business rules. On error the model gets the explanation and tries again.Production, money, documents
    ✅
    Rule: the schema guards the shape, the validator guards the meaning
    Constrained generation guarantees a number in total_amount. It does not guarantee the number is right. That is why arithmetic, dates and registry numbers are checked in Pydantic.
  2. JSON mode and JSON schema with Ollama and vLLM

    We start with the data. We describe the schema once in Pydantic and pass it to the server. Ollama takes it in the format parameter, vLLM in response_format, as with OpenAI. Temperature 0 makes extraction repeatable.

    Python · schema.py (run locally with Pydantic 2.13.5)
    from datetime import date
    from enum import Enum
    from typing import Optional
    from pydantic import BaseModel, Field, model_validator
    
    class Currency(str, Enum):
        EUR = "EUR"            # Bulgaria has used the euro since 2026-01-01
        USD = "USD"
    
    class Invoice(BaseModel):
        vendor_name: str = Field(description="Full legal name of the supplier")
        vendor_eik: Optional[str] = Field(default=None, pattern=r"^\d{9}(\d{4})?$",
                                          description="Company ID (EIK/BULSTAT) — 9 or 13 digits")
        invoice_date: date = Field(description="Invoice date")
        amount_without_vat: float = Field(gt=0, description="Taxable amount")
        vat_rate: float = Field(default=0.20, ge=0, le=1, description="VAT rate (0.20 = 20%)")
        total_amount: float = Field(gt=0, description="Total including VAT")
        currency: Currency = Currency.EUR
        description: Optional[str] = None
    
        @model_validator(mode="after")
        def check_total(self):
            # business rule: total = net × (1 + rate)
            expected = round(self.amount_without_vat * (1 + self.vat_rate), 2)
            if abs(self.total_amount - expected) > 0.01:
                raise ValueError(f"Total {self.total_amount} does not equal {expected:.2f}")
            return self
    ⚠️
    Trap: a field named like a type
    The old version had a field date: date. In Pydantic the field name then shadows the type and later fields using date break. Name fields by meaning: invoice_date.
    Python · extract.py (three levels)
    import json
    import ollama
    from openai import OpenAI
    from schema import Invoice
    
    INVOICE_TEXT = """
    Invoice No. 0000001234
    Supplier: Example Company Ltd., EIK 123456789
    Date: 15.03.2026
    Description: Consulting services for AI adoption
    Taxable amount: 2000.00 EUR
    VAT 20%: 400.00 EUR
    Total: 2400.00 EUR
    """   # fictional practice data
    
    MODEL = "llama3.1:8b"   # or any model pulled with ollama pull
    
    # ── Level 1: JSON mode — valid JSON, fields not guaranteed ──
    def extract_json_mode(text: str) -> dict:
        r = ollama.generate(model=MODEL, format="json", options={"temperature": 0},
                            prompt=f"Extract the invoice data as JSON:\n{text}")
        return json.loads(r["response"])
    
    # ── Level 2: JSON schema in Ollama — generation follows the schema ──
    def extract_schema_ollama(text: str) -> Invoice:
        r = ollama.chat(model=MODEL,
                        messages=[{"role": "user", "content": f"Extract the invoice data:\n{text}"}],
                        format=Invoice.model_json_schema(),
                        options={"temperature": 0})
        return Invoice.model_validate_json(r["message"]["content"])   # the total rule runs here too
    
    # ── Level 3: JSON schema in vLLM (OpenAI-compatible API) ──
    def extract_schema_vllm(text: str) -> Invoice:
        client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
        r = client.chat.completions.create(
            model="llama3",   # --served-model-name from Block 1
            messages=[{"role": "user", "content": f"Extract the invoice data:\n{text}"}],
            response_format={"type": "json_schema",
                             "json_schema": {"name": "invoice", "schema": Invoice.model_json_schema()}},
            temperature=0)
        return Invoice.model_validate_json(r.choices[0].message.content)
    
    if __name__ == "__main__":
        print(extract_schema_ollama(INVOICE_TEXT).model_dump_json(indent=2))
    expected output
    {
      "vendor_name": "Example Company Ltd.",
      "vendor_eik": "123456789",
      "invoice_date": "2026-03-15",
      "amount_without_vat": 2000.0,
      "vat_rate": 0.2,
      "total_amount": 2400.0,
      "currency": "EUR",
      "description": "Consulting services for AI adoption"
    }
    💡
    Why we still list the fields in the prompt
    The vLLM docs recommend describing the schema in the prompt as well. The constraint tells the model what it cannot write; the prompt tells it what it should write. Together they give better values.
  3. Instructor: a typed object and retries

    Instructor wraps the model client. You give it a Pydantic model and get back an object of that type, not a dict. If validation fails, for example because of a wrong total, Instructor sends the error back to the model and asks again. max_retries=3 means one attempt plus up to three retries. If those are not enough, it raises InstructorRetryException.

    Python · extract_typed.py
    import instructor
    from instructor.core import InstructorRetryException
    from schema import Invoice
    from extract import INVOICE_TEXT
    
    # Ollama: default address is http://localhost:11434/v1
    client = instructor.from_provider("ollama/llama3.1:8b")
    
    # vLLM or another OpenAI-compatible server:
    # from openai import OpenAI
    # client = instructor.from_openai(OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed"))
    
    def extract_invoice_typed(text: str) -> Invoice:
        return client.chat.completions.create(
            response_model=Invoice,      # ← the Pydantic model
            max_retries=3,               # ← 1 attempt + up to 3 retries
            messages=[{"role": "user", "content": f"Extract the data:\n{text}"}],
            temperature=0,
            # with from_openai also add model="llama3"
        )
    
    try:
        invoice = extract_invoice_typed(INVOICE_TEXT)
        print(invoice.vendor_name, invoice.invoice_date, invoice.total_amount)
    except InstructorRetryException as e:
        # never assume retries succeed: log it and route to manual review
        print("Failed after retries:", e)
    ⚠️
    Trap: Instructor's mode with local models
    With Ollama, Instructor picks the mode itself: tool calling for models that support it, JSON for the rest. If a small model returns empty or odd fields, try mode=instructor.Mode.JSON explicitly.
  4. Example: batch invoice processing with human review

    The real value comes when the model not only extracts fields but also flags what a person should look at. Here we check the mandatory invoice particulars under Art. 114 of the Bulgarian VAT Act and suggest a ledger account. The accountant keeps the final word.

    Python · invoice_batch.py
    import json
    from datetime import date
    from pathlib import Path
    from typing import Optional
    import instructor
    from instructor.core import InstructorRetryException
    from pydantic import BaseModel, Field
    
    class InvoiceResult(BaseModel):
        vendor_name: str
        vendor_eik: Optional[str] = None
        invoice_date: date
        total_eur: float = Field(gt=0)
        vat_included: bool
        gl_account: str = Field(description="Suggested ledger account (e.g. 602, 401)")
        requisites_ok: bool = Field(description="True if all particulars under Art. 114 of the VAT Act are present")
        issues: list[str] = Field(default_factory=list, description="Missing particulars or doubts")
    
    SYSTEM = """You assist an accountant. You check whether the invoice contains the
    particulars required by Art. 114 of the Bulgarian VAT Act and suggest a ledger account.
    Work only from the invoice text. Do not invent data — if something is missing, list it in issues."""
    
    client = instructor.from_provider("ollama/llama3.1:8b")
    
    def process_invoice(text: str, name: str) -> Optional[InvoiceResult]:
        try:
            r = client.chat.completions.create(
                response_model=InvoiceResult, max_retries=3, temperature=0,
                messages=[{"role": "system", "content": SYSTEM},
                          {"role": "user", "content": f"Invoice:\n{text}"}])
            mark = "✅" if r.requisites_ok and not r.issues else "👀"
            print(f"{mark} {name}: {r.vendor_name} — {r.total_eur:.2f} EUR")
            return r
        except InstructorRetryException as e:
            print(f"❌ {name}: manual processing — {e}")
            return None
    
    def process_directory(folder: str) -> None:
        ok, review = [], []
        for f in sorted(Path(folder).glob("*.txt")):
            r = process_invoice(f.read_text(encoding="utf-8"), f.name)
            if r and r.requisites_ok and not r.issues:
                ok.append(r.model_dump(mode="json"))
            else:
                review.append({"file": f.name, "result": r.model_dump(mode="json") if r else None})
        Path("processed.json").write_text(json.dumps(ok, ensure_ascii=False, indent=2), encoding="utf-8")
        Path("for_review.json").write_text(json.dumps(review, ensure_ascii=False, indent=2), encoding="utf-8")
        print(f"📊 Automatic: {len(ok)} · for review: {len(review)}")
    💡
    How to judge the benefit (illustration)
    If an accountant processes a few hundred invoices a month by hand and the system sets aside only the doubtful ones, the time goes to those instead of retyping. How much it actually saves — measure it on your own invoices with the test set from the next steps. Do not take other people's percentages at face value.
    ⛔
    Invoices are personal and commercial data
    Process them with a local model 🔒. If you use a cloud model 🌐, check the data processing agreement and do not send files with personal data without a legal basis.
  5. What we measure: the four core metrics

    "This prompt looks better" is an opinion. "Prompt B gives higher faithfulness on our 50 test cases" is data. To compare, first agree on what you measure. All metrics below range from 0 to 1.

    MetricQuestionReading
    FaithfulnessIs every claim in the answer supported by the given context?the closer to 1 the better
    Answer relevancyDoes it answer the question directly?high
    Context recallDid retrieval find everything needed for the correct answer?low → retrieval problem
    Context precisionHow many of the retrieved chunks are actually useful, and are they ranked on top?low → noise in the context
    ✅
    Thresholds are agreed, not copied
    The old version gave "faithfulness above 0.90" as a universal threshold. There is no such threshold. Measure a baseline on your own set and decide what is acceptable for the task: a legal answer tolerates far less than a draft social post.
  6. A/B test of two prompts

    The logic is the same as a website A/B test: change one thing, run both variants on the same cases at temperature 0, and decide by the numbers. Keep the test cases in git like code. Three cases are enough to see how it works; a decision needs dozens.

    Python · prompt_ab_test.py
    import ollama
    from statistics import mean
    
    MODEL = "llama3.1:8b"
    
    # Fictional internal rulebook — practice data, not legal reference
    TEST_CASES = [
        {"question": "How many days of paid leave does a new employee get?",
         "context": "Rulebook, 4.1: Every employee is entitled to 25 working days of paid leave per year.",
         "ground_truth": "25 working days"},
        {"question": "What do I need to access the server room?",
         "context": "Rulebook, 7.3: Access requires a staff card and written approval from a manager.",
         "ground_truth": "staff card and written approval from a manager"},
        {"question": "When is a business trip report due?",
         "context": "Rulebook, 9.2: A business trip report is due within 5 working days after return.",
         "ground_truth": "within 5 working days"},
    ]
    
    PROMPT_A = """Answer the question based on the context.
    Context: {context}
    Question: {question}"""
    
    PROMPT_B = """Answer ONLY from the context. If the answer is missing, say "Not contained in the documents."
    Cite the rulebook item.
    Context: {context}
    Question: {question}
    Short, precise answer:"""
    
    def answer(template: str, case: dict) -> str:
        r = ollama.generate(model=MODEL, prompt=template.format(**case), options={"temperature": 0})
        return r["response"].strip()
    
    def score(ans: str, case: dict) -> dict:
        """Rough heuristics for a quick signal; for decisions use a judge model or RAGAS."""
        gt = set(case["ground_truth"].lower().split())
        a = set(ans.lower().split())
        ctx = set(case["context"].lower().split())
        accuracy = len(gt & a) / len(gt) if gt else 0.0
        extra = a - ctx - gt
        faithfulness = 1.0 - min(1.0, len(extra) / max(1, len(a)))
        conciseness = 1.0 if len(ans.split()) < 50 else 0.7
        return {"accuracy": accuracy, "faithfulness": faithfulness, "conciseness": conciseness,
                "composite": (accuracy + faithfulness + conciseness) / 3}
    
    def run(template: str) -> dict:
        rows = [score(answer(template, c), c) for c in TEST_CASES]
        return {k: round(mean(r[k] for r in rows), 3) for k in rows[0]}
    
    if __name__ == "__main__":   # so regression_test.py can import the data without running the test
        for name, tpl in (("Prompt A", PROMPT_A), ("Prompt B", PROMPT_B)):
            res = run(tpl)
            print(name, " ".join(f"{k}={v:.3f}" for k, v in res.items()))
    ⚠️
    Trap: winning on noise
    The old version showed "Prompt B wins by +24%" as an expected result. That is an illustration, not a measurement. On a small set a difference of a few percent may be chance. Look at individual cases too: if B wins on average but fails an important case, it is not the winner.
  7. The judge model (LLM-as-a-judge)

    Grading hundreds of answers by hand takes hours. So a stronger model grades a weaker one against a clear rubric. The grade is itself a structured output: the judge returns bounded numbers and a short justification, not an essay.

    Python · llm_judge.py
    import instructor
    import ollama
    from pydantic import BaseModel, Field
    
    class JudgeScore(BaseModel):
        faithfulness: float = Field(ge=0, le=1, description="Is everything supported by the context")
        relevance: float = Field(ge=0, le=1, description="Does it answer the question directly")
        completeness: float = Field(ge=0, le=1, description="Is anything important missing")
        reasoning: str = Field(description="Justification in 1–2 sentences")
    
    JUDGE_SYSTEM = """You grade answers from a RAG system on three criteria from 0.0 to 1.0.
    FAITHFULNESS: 1.0 = all from context · 0.5 = half is assumption · 0.0 = ignores the context
    RELEVANCE:    1.0 = direct, precise answer · 0.5 = partial · 0.0 = off-topic
    COMPLETENESS: 1.0 = everything important included · 0.5 = details missing · 0.0 = key data missing
    Be strict. Justify first, then give the numbers."""
    
    JUDGE_MODEL = "llama3.1:70b"    # stronger than the graded model
    ANSWER_MODEL = "llama3.1:8b"
    judge = instructor.from_provider(f"ollama/{JUDGE_MODEL}")
    
    def judge_answer(question: str, context: str, answer: str) -> JudgeScore:
        return judge.chat.completions.create(
            response_model=JudgeScore, max_retries=2, temperature=0,
            messages=[{"role": "system", "content": JUDGE_SYSTEM},
                      {"role": "user", "content": f"QUESTION: {question}\nCONTEXT:\n{context}\nANSWER:\n{answer}"}])
    
    def evaluate_dataset(cases: list[dict]) -> float:
        total = []
        for c in cases:
            ans = ollama.generate(model=ANSWER_MODEL, options={"temperature": 0},
                                  prompt=f"Context: {c['context']}\nQuestion: {c['question']}")["response"]
            s = judge_answer(c["question"], c["context"], ans)
            comp = (s.faithfulness + s.relevance + s.completeness) / 3
            total.append(comp)
            flag = "✅" if comp > 0.8 else ("⚠️" if comp > 0.6 else "❌")
            print(f"{flag} {comp:.2f} | {s.reasoning[:70]}")
        avg = sum(total) / len(total)
        print(f"📊 Average: {avg:.3f}")
        return avg
    ⚠️
    Trap: the judge makes mistakes too
    Research on judge models describes persistent biases: they prefer longer answers, the first of two options, and answers from their own model. So: swap the order in pairwise comparisons, give a strict rubric, and check at least 20–30 grades against a human before you trust it.
  8. RAGAS: ready-made metrics for RAG

    RAGAS specialises in RAG and computes the metrics from step 5 using a judge model and embeddings. Since version 0.4, metrics are imported from ragas.metrics.collections and called one by one with ascore. With a local Ollama nothing leaves your machine.

    Python · ragas_eval.py (RAGAS 0.4.3)
    import asyncio
    from openai import AsyncOpenAI
    from ragas.llms import llm_factory
    from ragas.embeddings import embedding_factory
    from ragas.metrics.collections import Faithfulness, AnswerRelevancy, ContextRecall, ContextPrecision
    
    # Ollama via its OpenAI-compatible address — no cloud
    ollama_client = AsyncOpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
    llm = llm_factory("llama3.1:70b", client=ollama_client, temperature=0)   # the judge
    emb = embedding_factory("openai", model="nomic-embed-text", client=ollama_client)
    
    metrics = {
        "faithfulness": Faithfulness(llm=llm),
        "answer_relevancy": AnswerRelevancy(llm=llm, embeddings=emb),
        "context_recall": ContextRecall(llm=llm),
        "context_precision": ContextPrecision(llm=llm),
    }
    
    SAMPLES = [   # fictional practice data
        {"user_input": "How many days of paid leave do I get?",
         "response": "You get 25 working days of paid leave per year (item 4.1).",
         "retrieved_contexts": ["Rulebook, 4.1: Every employee is entitled to 25 working days of paid leave per year."],
         "reference": "25 working days per year"},
    ]
    
    async def main():
        for s in SAMPLES:
            f = await metrics["faithfulness"].ascore(user_input=s["user_input"], response=s["response"],
                                                     retrieved_contexts=s["retrieved_contexts"])
            a = await metrics["answer_relevancy"].ascore(user_input=s["user_input"], response=s["response"])
            r = await metrics["context_recall"].ascore(user_input=s["user_input"], reference=s["reference"],
                                                       retrieved_contexts=s["retrieved_contexts"])
            p = await metrics["context_precision"].ascore(user_input=s["user_input"], reference=s["reference"],
                                                          retrieved_contexts=s["retrieved_contexts"])
            print(f"{s['user_input'][:40]} | faith={f.value:.2f} rel={a.value:.2f} "
                  f"recall={r.value:.2f} prec={p.value:.2f}")
    
    asyncio.run(main())
    ⚠️
    Trap: old RAGAS examples
    Most examples online use from ragas.metrics import faithfulness, evaluate(dataset=...) and LangchainLLMWrapper. In 0.4 they still work, but warn that they will be removed. Columns were renamed too: question → user_input, answer → response, contexts → retrieved_contexts, ground_truth → reference.
    🔑
    When RAGAS, when your own judge
    RAGAS — when you have a working RAG system and want standard, comparable metrics. context_recall needs a reference answer (reference). Your own judge — when prototyping or measuring something specific to your domain, for example "does it cite the exact rulebook item". No reference needed there.
  9. deepeval: your own criterion with GEval

    deepeval makes evaluation look like a unit test. GEval takes a criterion in plain language and turns it into steps for the judge. Handy when the criterion is yours and not among the standard ones.

    Python · test_answers.py (deepeval 4.2.7)
    from deepeval.metrics import GEval
    from deepeval.models import OllamaModel
    from deepeval.test_case import LLMTestCase, SingleTurnParams
    
    judge = OllamaModel(model="llama3.1:70b", base_url="http://localhost:11434", temperature=0)
    
    cites_rule = GEval(
        name="Cites the item",
        criteria="The answer names the rulebook item it relies on and claims nothing outside the context.",
        evaluation_params=[SingleTurnParams.INPUT, SingleTurnParams.ACTUAL_OUTPUT,
                           SingleTurnParams.RETRIEVAL_CONTEXT],
        model=judge,
        threshold=0.7,
    )
    
    case = LLMTestCase(
        input="How many days of paid leave do I get?",
        actual_output="You get 25 working days per year (item 4.1).",
        retrieval_context=["Rulebook, 4.1: Every employee is entitled to 25 working days of paid leave per year."],
    )
    
    cites_rule.measure(case)
    print(cites_rule.score, cites_rule.reason)

    Older examples use LLMTestCaseParams. It was renamed to SingleTurnParams; the old name still works, with a warning.

  10. Regression test: a brake before release

    Every change to a prompt or model improves some cases and may worsen others. A regression test compares the new result with the stored baseline and stops the release if the drop exceeds the threshold. The baseline is updated only on success.

    Python · regression_test.py
    import json
    import sys
    from datetime import datetime
    from pathlib import Path
    import ollama
    from llm_judge import judge_answer
    from prompt_ab_test import TEST_CASES, PROMPT_B
    
    MODEL = "llama3.1:8b"
    BASELINE = Path("baseline_scores.json")
    THRESHOLD = 0.05   # allowed drop: 0.05 on the 0–1 scale
    
    def run(prompt: str) -> list[float]:
        out = []
        for c in TEST_CASES:
            ans = ollama.generate(model=MODEL, prompt=prompt.format(**c), options={"temperature": 0})["response"]
            s = judge_answer(c["question"], c["context"], ans)
            out.append((s.faithfulness + s.relevance + s.completeness) / 3)
        return out
    
    def main(prompt: str) -> bool:
        scores = run(prompt)
        avg = sum(scores) / len(scores)
        if BASELINE.exists():
            base = json.loads(BASELINE.read_text(encoding="utf-8"))["avg_score"]
            delta = avg - base
            print(f"Baseline: {base:.3f} · new: {avg:.3f} · delta: {delta:+.3f}")
            if delta < -THRESHOLD:
                print("❌ REGRESSION — release blocked")
                return False
        else:
            print("ℹ️ No baseline — storing the current result")
        BASELINE.write_text(json.dumps({"avg_score": avg, "n_cases": len(scores), "scores": scores,
                                        "timestamp": datetime.now().isoformat()}, indent=2), encoding="utf-8")
        return True
    
    if __name__ == "__main__":
        sys.exit(0 if main(PROMPT_B) else 1)   # exit code 1 stops CI/CD
    🛡️
    What we fixed compared to the old version
    The old script used ollama and judge_answer without importing them, and saved the new result as the baseline even after a regression. Now the old baseline stays in place on failure.

04Check

Checklist

Quiz

1. Instructor runs with max_retries=3 and validation fails on every attempt. What happens in the end?

2. RAGAS reports context_recall = 0.42. What is the most likely cause?

3. A new prompt raises faithfulness but lowers answer relevancy. What does that show?

4. What makes the output follow the JSON schema while the model is still generating?

05What's next

This completes Block 2: prompting techniques, system prompts, structured outputs and evaluation with data.

06Sources

  1. Ollama: structured outputs — format="json" and JSON schema.
  2. vLLM: Structured Outputs — response_format, structured_outputs, the removed guided_*.
  3. Pydantic: validators — field_validator and model_validator.
  4. Instructor: documentation and Ollama integration.
  5. Outlines — constrained generation when the model runs in your process.
  6. RAGAS: metrics — faithfulness, answer relevancy, context recall and precision.
  7. deepeval: getting started — tests and GEval.
  8. Es et al. (2023). Ragas: Automated Evaluation of Retrieval Augmented Generation — the paper that introduced the metrics.
  9. Zheng et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — the biases of judge models.