Structured outputs and prompt evaluation
A model naturally writes free text, while the program on the other side expects exact fields. In the first half we turn output into validated data with a JSON schema, Pydantic and Instructor. In the second half we stop guessing which prompt is better and start measuring: an A/B test, a judge model, RAGAS and a regression test before every release.
format, vLLM in response_format. The old guided_json parameters were removed from vLLM in v0.12.0. The RAGAS example was rewritten for the new API (ragas.metrics.collections and ascore), because evaluate() and LangchainLLMWrapper are now marked deprecated. deepeval (GEval) was added. Instructor's exception is now imported from instructor.core. Invoices are in euro: Bulgaria has used the euro since 01.01.2026. The total check now sits in a model_validator and was run locally. The reliability percentages per method and the "real result" of an accounting firm could not be verified, so they were removed or marked as illustrative. Test data now comes from a fictional internal rulebook instead of "regulations" citing non-existent articles.
01What you will learn
- The four rungs of reliability: an instruction in the prompt, JSON mode, a JSON schema with constrained generation, and a schema with validation and retries.
- How to describe data with Pydantic, including a business rule such as "total = net × (1 + VAT)".
- How Instructor returns a typed object, retries, and what happens when retries run out.
- Which metrics measure the quality of an answer and of a RAG system, and what each one means.
- How to compare two prompts with data, how to use a judge model, RAGAS and deepeval, and how to block a release on regression.
02Before you start
- You have completed Block 2 · Part 1 (prompt engineering): system prompt, few-shot, temperature.
- Python 3.10 or newer and a virtual environment:
pip install "pydantic>=2" instructor openai ollama ragas deepeval. - A local model server: Ollama (
ollama pulla model of your choice) or vLLM from Block 1. Both expose an OpenAI-compatible/v1address. - For the judge — a stronger model than the one you evaluate. It can be local or cloud. A cloud model sends data out, so do not give it personal data.
- RAG basics from Block 0: what a chunk, an embedding and context retrieval are.
03Steps
-
Why free text breaks systems
Imagine an invoice analyser built on
text.split("Amount:"). It works for most invoices. For the rest it either crashes or silently writes a wrong amount into the books. Why is that dangerous? Because the error stays invisible until someone compares the numbers. Structured output moves the check to where it belongs: into the schema, before the data travels further.Rung What it guarantees When Instruction
"Return only JSON"Nothing. The model often adds explanations or Markdown fences. Quick prototype JSON mode format="json"·{"type":"json_object"}Syntactically valid JSON, but with whatever fields the model chooses. Simple structures JSON schema format=schema·json_schemaThe server constrains generation to follow the schema: names, types, enums. Almost always — today's default Schema + Pydantic + retry
InstructorThe shape and the business rules. On error the model gets the explanation and tries again. Production, money, documents ✅Rule: the schema guards the shape, the validator guards the meaningConstrained generation guarantees a number intotal_amount. It does not guarantee the number is right. That is why arithmetic, dates and registry numbers are checked in Pydantic. -
JSON mode and JSON schema with Ollama and vLLM
We start with the data. We describe the schema once in Pydantic and pass it to the server. Ollama takes it in the
formatparameter, vLLM inresponse_format, as with OpenAI. Temperature 0 makes extraction repeatable.Python · schema.py (run locally with Pydantic 2.13.5)from datetime import date from enum import Enum from typing import Optional from pydantic import BaseModel, Field, model_validator class Currency(str, Enum): EUR = "EUR" # Bulgaria has used the euro since 2026-01-01 USD = "USD" class Invoice(BaseModel): vendor_name: str = Field(description="Full legal name of the supplier") vendor_eik: Optional[str] = Field(default=None, pattern=r"^\d{9}(\d{4})?$", description="Company ID (EIK/BULSTAT) — 9 or 13 digits") invoice_date: date = Field(description="Invoice date") amount_without_vat: float = Field(gt=0, description="Taxable amount") vat_rate: float = Field(default=0.20, ge=0, le=1, description="VAT rate (0.20 = 20%)") total_amount: float = Field(gt=0, description="Total including VAT") currency: Currency = Currency.EUR description: Optional[str] = None @model_validator(mode="after") def check_total(self): # business rule: total = net × (1 + rate) expected = round(self.amount_without_vat * (1 + self.vat_rate), 2) if abs(self.total_amount - expected) > 0.01: raise ValueError(f"Total {self.total_amount} does not equal {expected:.2f}") return self⚠️Trap: a field named like a typeThe old version had a fielddate: date. In Pydantic the field name then shadows the type and later fields usingdatebreak. Name fields by meaning:invoice_date.Python · extract.py (three levels)import json import ollama from openai import OpenAI from schema import Invoice INVOICE_TEXT = """ Invoice No. 0000001234 Supplier: Example Company Ltd., EIK 123456789 Date: 15.03.2026 Description: Consulting services for AI adoption Taxable amount: 2000.00 EUR VAT 20%: 400.00 EUR Total: 2400.00 EUR """ # fictional practice data MODEL = "llama3.1:8b" # or any model pulled with ollama pull # ── Level 1: JSON mode — valid JSON, fields not guaranteed ── def extract_json_mode(text: str) -> dict: r = ollama.generate(model=MODEL, format="json", options={"temperature": 0}, prompt=f"Extract the invoice data as JSON:\n{text}") return json.loads(r["response"]) # ── Level 2: JSON schema in Ollama — generation follows the schema ── def extract_schema_ollama(text: str) -> Invoice: r = ollama.chat(model=MODEL, messages=[{"role": "user", "content": f"Extract the invoice data:\n{text}"}], format=Invoice.model_json_schema(), options={"temperature": 0}) return Invoice.model_validate_json(r["message"]["content"]) # the total rule runs here too # ── Level 3: JSON schema in vLLM (OpenAI-compatible API) ── def extract_schema_vllm(text: str) -> Invoice: client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed") r = client.chat.completions.create( model="llama3", # --served-model-name from Block 1 messages=[{"role": "user", "content": f"Extract the invoice data:\n{text}"}], response_format={"type": "json_schema", "json_schema": {"name": "invoice", "schema": Invoice.model_json_schema()}}, temperature=0) return Invoice.model_validate_json(r.choices[0].message.content) if __name__ == "__main__": print(extract_schema_ollama(INVOICE_TEXT).model_dump_json(indent=2))expected output{ "vendor_name": "Example Company Ltd.", "vendor_eik": "123456789", "invoice_date": "2026-03-15", "amount_without_vat": 2000.0, "vat_rate": 0.2, "total_amount": 2400.0, "currency": "EUR", "description": "Consulting services for AI adoption" }💡Why we still list the fields in the promptThe vLLM docs recommend describing the schema in the prompt as well. The constraint tells the model what it cannot write; the prompt tells it what it should write. Together they give better values. -
Instructor: a typed object and retries
Instructor wraps the model client. You give it a Pydantic model and get back an object of that type, not a dict. If validation fails, for example because of a wrong total, Instructor sends the error back to the model and asks again.
max_retries=3means one attempt plus up to three retries. If those are not enough, it raisesInstructorRetryException.Python · extract_typed.pyimport instructor from instructor.core import InstructorRetryException from schema import Invoice from extract import INVOICE_TEXT # Ollama: default address is http://localhost:11434/v1 client = instructor.from_provider("ollama/llama3.1:8b") # vLLM or another OpenAI-compatible server: # from openai import OpenAI # client = instructor.from_openai(OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")) def extract_invoice_typed(text: str) -> Invoice: return client.chat.completions.create( response_model=Invoice, # ← the Pydantic model max_retries=3, # ← 1 attempt + up to 3 retries messages=[{"role": "user", "content": f"Extract the data:\n{text}"}], temperature=0, # with from_openai also add model="llama3" ) try: invoice = extract_invoice_typed(INVOICE_TEXT) print(invoice.vendor_name, invoice.invoice_date, invoice.total_amount) except InstructorRetryException as e: # never assume retries succeed: log it and route to manual review print("Failed after retries:", e)⚠️Trap: Instructor's mode with local modelsWith Ollama, Instructor picks the mode itself: tool calling for models that support it, JSON for the rest. If a small model returns empty or odd fields, trymode=instructor.Mode.JSONexplicitly. -
Example: batch invoice processing with human review
The real value comes when the model not only extracts fields but also flags what a person should look at. Here we check the mandatory invoice particulars under Art. 114 of the Bulgarian VAT Act and suggest a ledger account. The accountant keeps the final word.
Python · invoice_batch.pyimport json from datetime import date from pathlib import Path from typing import Optional import instructor from instructor.core import InstructorRetryException from pydantic import BaseModel, Field class InvoiceResult(BaseModel): vendor_name: str vendor_eik: Optional[str] = None invoice_date: date total_eur: float = Field(gt=0) vat_included: bool gl_account: str = Field(description="Suggested ledger account (e.g. 602, 401)") requisites_ok: bool = Field(description="True if all particulars under Art. 114 of the VAT Act are present") issues: list[str] = Field(default_factory=list, description="Missing particulars or doubts") SYSTEM = """You assist an accountant. You check whether the invoice contains the particulars required by Art. 114 of the Bulgarian VAT Act and suggest a ledger account. Work only from the invoice text. Do not invent data — if something is missing, list it in issues.""" client = instructor.from_provider("ollama/llama3.1:8b") def process_invoice(text: str, name: str) -> Optional[InvoiceResult]: try: r = client.chat.completions.create( response_model=InvoiceResult, max_retries=3, temperature=0, messages=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": f"Invoice:\n{text}"}]) mark = "✅" if r.requisites_ok and not r.issues else "👀" print(f"{mark} {name}: {r.vendor_name} — {r.total_eur:.2f} EUR") return r except InstructorRetryException as e: print(f"❌ {name}: manual processing — {e}") return None def process_directory(folder: str) -> None: ok, review = [], [] for f in sorted(Path(folder).glob("*.txt")): r = process_invoice(f.read_text(encoding="utf-8"), f.name) if r and r.requisites_ok and not r.issues: ok.append(r.model_dump(mode="json")) else: review.append({"file": f.name, "result": r.model_dump(mode="json") if r else None}) Path("processed.json").write_text(json.dumps(ok, ensure_ascii=False, indent=2), encoding="utf-8") Path("for_review.json").write_text(json.dumps(review, ensure_ascii=False, indent=2), encoding="utf-8") print(f"📊 Automatic: {len(ok)} · for review: {len(review)}")💡How to judge the benefit (illustration)If an accountant processes a few hundred invoices a month by hand and the system sets aside only the doubtful ones, the time goes to those instead of retyping. How much it actually saves — measure it on your own invoices with the test set from the next steps. Do not take other people's percentages at face value.⛔Invoices are personal and commercial dataProcess them with a local model 🔒. If you use a cloud model 🌐, check the data processing agreement and do not send files with personal data without a legal basis. -
What we measure: the four core metrics
"This prompt looks better" is an opinion. "Prompt B gives higher faithfulness on our 50 test cases" is data. To compare, first agree on what you measure. All metrics below range from 0 to 1.
Metric Question Reading Faithfulness Is every claim in the answer supported by the given context? the closer to 1 the better Answer relevancy Does it answer the question directly? high Context recall Did retrieval find everything needed for the correct answer? low → retrieval problem Context precision How many of the retrieved chunks are actually useful, and are they ranked on top? low → noise in the context ✅Thresholds are agreed, not copiedThe old version gave "faithfulness above 0.90" as a universal threshold. There is no such threshold. Measure a baseline on your own set and decide what is acceptable for the task: a legal answer tolerates far less than a draft social post. -
A/B test of two prompts
The logic is the same as a website A/B test: change one thing, run both variants on the same cases at temperature 0, and decide by the numbers. Keep the test cases in git like code. Three cases are enough to see how it works; a decision needs dozens.
Python · prompt_ab_test.pyimport ollama from statistics import mean MODEL = "llama3.1:8b" # Fictional internal rulebook — practice data, not legal reference TEST_CASES = [ {"question": "How many days of paid leave does a new employee get?", "context": "Rulebook, 4.1: Every employee is entitled to 25 working days of paid leave per year.", "ground_truth": "25 working days"}, {"question": "What do I need to access the server room?", "context": "Rulebook, 7.3: Access requires a staff card and written approval from a manager.", "ground_truth": "staff card and written approval from a manager"}, {"question": "When is a business trip report due?", "context": "Rulebook, 9.2: A business trip report is due within 5 working days after return.", "ground_truth": "within 5 working days"}, ] PROMPT_A = """Answer the question based on the context. Context: {context} Question: {question}""" PROMPT_B = """Answer ONLY from the context. If the answer is missing, say "Not contained in the documents." Cite the rulebook item. Context: {context} Question: {question} Short, precise answer:""" def answer(template: str, case: dict) -> str: r = ollama.generate(model=MODEL, prompt=template.format(**case), options={"temperature": 0}) return r["response"].strip() def score(ans: str, case: dict) -> dict: """Rough heuristics for a quick signal; for decisions use a judge model or RAGAS.""" gt = set(case["ground_truth"].lower().split()) a = set(ans.lower().split()) ctx = set(case["context"].lower().split()) accuracy = len(gt & a) / len(gt) if gt else 0.0 extra = a - ctx - gt faithfulness = 1.0 - min(1.0, len(extra) / max(1, len(a))) conciseness = 1.0 if len(ans.split()) < 50 else 0.7 return {"accuracy": accuracy, "faithfulness": faithfulness, "conciseness": conciseness, "composite": (accuracy + faithfulness + conciseness) / 3} def run(template: str) -> dict: rows = [score(answer(template, c), c) for c in TEST_CASES] return {k: round(mean(r[k] for r in rows), 3) for k in rows[0]} if __name__ == "__main__": # so regression_test.py can import the data without running the test for name, tpl in (("Prompt A", PROMPT_A), ("Prompt B", PROMPT_B)): res = run(tpl) print(name, " ".join(f"{k}={v:.3f}" for k, v in res.items()))⚠️Trap: winning on noiseThe old version showed "Prompt B wins by +24%" as an expected result. That is an illustration, not a measurement. On a small set a difference of a few percent may be chance. Look at individual cases too: if B wins on average but fails an important case, it is not the winner. -
The judge model (LLM-as-a-judge)
Grading hundreds of answers by hand takes hours. So a stronger model grades a weaker one against a clear rubric. The grade is itself a structured output: the judge returns bounded numbers and a short justification, not an essay.
Python · llm_judge.pyimport instructor import ollama from pydantic import BaseModel, Field class JudgeScore(BaseModel): faithfulness: float = Field(ge=0, le=1, description="Is everything supported by the context") relevance: float = Field(ge=0, le=1, description="Does it answer the question directly") completeness: float = Field(ge=0, le=1, description="Is anything important missing") reasoning: str = Field(description="Justification in 1–2 sentences") JUDGE_SYSTEM = """You grade answers from a RAG system on three criteria from 0.0 to 1.0. FAITHFULNESS: 1.0 = all from context · 0.5 = half is assumption · 0.0 = ignores the context RELEVANCE: 1.0 = direct, precise answer · 0.5 = partial · 0.0 = off-topic COMPLETENESS: 1.0 = everything important included · 0.5 = details missing · 0.0 = key data missing Be strict. Justify first, then give the numbers.""" JUDGE_MODEL = "llama3.1:70b" # stronger than the graded model ANSWER_MODEL = "llama3.1:8b" judge = instructor.from_provider(f"ollama/{JUDGE_MODEL}") def judge_answer(question: str, context: str, answer: str) -> JudgeScore: return judge.chat.completions.create( response_model=JudgeScore, max_retries=2, temperature=0, messages=[{"role": "system", "content": JUDGE_SYSTEM}, {"role": "user", "content": f"QUESTION: {question}\nCONTEXT:\n{context}\nANSWER:\n{answer}"}]) def evaluate_dataset(cases: list[dict]) -> float: total = [] for c in cases: ans = ollama.generate(model=ANSWER_MODEL, options={"temperature": 0}, prompt=f"Context: {c['context']}\nQuestion: {c['question']}")["response"] s = judge_answer(c["question"], c["context"], ans) comp = (s.faithfulness + s.relevance + s.completeness) / 3 total.append(comp) flag = "✅" if comp > 0.8 else ("⚠️" if comp > 0.6 else "❌") print(f"{flag} {comp:.2f} | {s.reasoning[:70]}") avg = sum(total) / len(total) print(f"📊 Average: {avg:.3f}") return avg⚠️Trap: the judge makes mistakes tooResearch on judge models describes persistent biases: they prefer longer answers, the first of two options, and answers from their own model. So: swap the order in pairwise comparisons, give a strict rubric, and check at least 20–30 grades against a human before you trust it. -
RAGAS: ready-made metrics for RAG
RAGAS specialises in RAG and computes the metrics from step 5 using a judge model and embeddings. Since version 0.4, metrics are imported from
ragas.metrics.collectionsand called one by one withascore. With a local Ollama nothing leaves your machine.Python · ragas_eval.py (RAGAS 0.4.3)import asyncio from openai import AsyncOpenAI from ragas.llms import llm_factory from ragas.embeddings import embedding_factory from ragas.metrics.collections import Faithfulness, AnswerRelevancy, ContextRecall, ContextPrecision # Ollama via its OpenAI-compatible address — no cloud ollama_client = AsyncOpenAI(base_url="http://localhost:11434/v1", api_key="ollama") llm = llm_factory("llama3.1:70b", client=ollama_client, temperature=0) # the judge emb = embedding_factory("openai", model="nomic-embed-text", client=ollama_client) metrics = { "faithfulness": Faithfulness(llm=llm), "answer_relevancy": AnswerRelevancy(llm=llm, embeddings=emb), "context_recall": ContextRecall(llm=llm), "context_precision": ContextPrecision(llm=llm), } SAMPLES = [ # fictional practice data {"user_input": "How many days of paid leave do I get?", "response": "You get 25 working days of paid leave per year (item 4.1).", "retrieved_contexts": ["Rulebook, 4.1: Every employee is entitled to 25 working days of paid leave per year."], "reference": "25 working days per year"}, ] async def main(): for s in SAMPLES: f = await metrics["faithfulness"].ascore(user_input=s["user_input"], response=s["response"], retrieved_contexts=s["retrieved_contexts"]) a = await metrics["answer_relevancy"].ascore(user_input=s["user_input"], response=s["response"]) r = await metrics["context_recall"].ascore(user_input=s["user_input"], reference=s["reference"], retrieved_contexts=s["retrieved_contexts"]) p = await metrics["context_precision"].ascore(user_input=s["user_input"], reference=s["reference"], retrieved_contexts=s["retrieved_contexts"]) print(f"{s['user_input'][:40]} | faith={f.value:.2f} rel={a.value:.2f} " f"recall={r.value:.2f} prec={p.value:.2f}") asyncio.run(main())⚠️Trap: old RAGAS examplesMost examples online usefrom ragas.metrics import faithfulness,evaluate(dataset=...)andLangchainLLMWrapper. In 0.4 they still work, but warn that they will be removed. Columns were renamed too:question→user_input,answer→response,contexts→retrieved_contexts,ground_truth→reference.🔑When RAGAS, when your own judgeRAGAS — when you have a working RAG system and want standard, comparable metrics.context_recallneeds a reference answer (reference). Your own judge — when prototyping or measuring something specific to your domain, for example "does it cite the exact rulebook item". No reference needed there. -
deepeval: your own criterion with GEval
deepeval makes evaluation look like a unit test.
GEvaltakes a criterion in plain language and turns it into steps for the judge. Handy when the criterion is yours and not among the standard ones.Python · test_answers.py (deepeval 4.2.7)from deepeval.metrics import GEval from deepeval.models import OllamaModel from deepeval.test_case import LLMTestCase, SingleTurnParams judge = OllamaModel(model="llama3.1:70b", base_url="http://localhost:11434", temperature=0) cites_rule = GEval( name="Cites the item", criteria="The answer names the rulebook item it relies on and claims nothing outside the context.", evaluation_params=[SingleTurnParams.INPUT, SingleTurnParams.ACTUAL_OUTPUT, SingleTurnParams.RETRIEVAL_CONTEXT], model=judge, threshold=0.7, ) case = LLMTestCase( input="How many days of paid leave do I get?", actual_output="You get 25 working days per year (item 4.1).", retrieval_context=["Rulebook, 4.1: Every employee is entitled to 25 working days of paid leave per year."], ) cites_rule.measure(case) print(cites_rule.score, cites_rule.reason)Older examples use
LLMTestCaseParams. It was renamed toSingleTurnParams; the old name still works, with a warning. -
Regression test: a brake before release
Every change to a prompt or model improves some cases and may worsen others. A regression test compares the new result with the stored baseline and stops the release if the drop exceeds the threshold. The baseline is updated only on success.
Python · regression_test.pyimport json import sys from datetime import datetime from pathlib import Path import ollama from llm_judge import judge_answer from prompt_ab_test import TEST_CASES, PROMPT_B MODEL = "llama3.1:8b" BASELINE = Path("baseline_scores.json") THRESHOLD = 0.05 # allowed drop: 0.05 on the 0–1 scale def run(prompt: str) -> list[float]: out = [] for c in TEST_CASES: ans = ollama.generate(model=MODEL, prompt=prompt.format(**c), options={"temperature": 0})["response"] s = judge_answer(c["question"], c["context"], ans) out.append((s.faithfulness + s.relevance + s.completeness) / 3) return out def main(prompt: str) -> bool: scores = run(prompt) avg = sum(scores) / len(scores) if BASELINE.exists(): base = json.loads(BASELINE.read_text(encoding="utf-8"))["avg_score"] delta = avg - base print(f"Baseline: {base:.3f} · new: {avg:.3f} · delta: {delta:+.3f}") if delta < -THRESHOLD: print("❌ REGRESSION — release blocked") return False else: print("ℹ️ No baseline — storing the current result") BASELINE.write_text(json.dumps({"avg_score": avg, "n_cases": len(scores), "scores": scores, "timestamp": datetime.now().isoformat()}, indent=2), encoding="utf-8") return True if __name__ == "__main__": sys.exit(0 if main(PROMPT_B) else 1) # exit code 1 stops CI/CD🛡️What we fixed compared to the old versionThe old script usedollamaandjudge_answerwithout importing them, and saved the new result as the baseline even after a regression. Now the old baseline stays in place on failure.
04Check
Checklist
- The model server answers on
/v1(Ollama or vLLM). - Schema extraction returns output that passes
Invoice.model_validate_json. - A deliberately wrong total raises a validation error.
- Instructor returns a typed object and the failure path catches
InstructorRetryException. - You have a test set of at least 10 cases (better 50+) with reference answers, kept in git.
- The A/B test ran at temperature 0 and per-case results are saved.
- The judge rubric is written and some grades were checked against a human.
- RAGAS metrics are computed for your RAG.
baseline_scores.jsonis in git and the regression test is wired into CI.
Quiz
1. Instructor runs with max_retries=3 and validation fails on every attempt. What happens in the end?
2. RAGAS reports context_recall = 0.42. What is the most likely cause?
3. A new prompt raises faithfulness but lowers answer relevancy. What does that show?
4. What makes the output follow the JSON schema while the model is still generating?
05What's next
This completes Block 2: prompting techniques, system prompts, structured outputs and evaluation with data.
06Sources
- Ollama: structured outputs —
format="json"and JSON schema. - vLLM: Structured Outputs —
response_format,structured_outputs, the removedguided_*. - Pydantic: validators —
field_validatorandmodel_validator. - Instructor: documentation and Ollama integration.
- Outlines — constrained generation when the model runs in your process.
- RAGAS: metrics — faithfulness, answer relevancy, context recall and precision.
- deepeval: getting started — tests and GEval.
- Es et al. (2023). Ragas: Automated Evaluation of Retrieval Augmented Generation — the paper that introduced the metrics.
- Zheng et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — the biases of judge models.