Ticketing Chatbot with a Local AI on GX10
How to build a ticketing chatbot for a cultural and sports complex in which the model only recognises what the visitor wants, while all the facts — events, free seats, prices in EUR, answers to questions — come from the database. Anything uncertain goes to a person. The bot says it is an AI and does not sell or pay on anyone's behalf. The events, prices and answers in the lesson are invented.
del conversations[sid] could crash — it is now pop; the database connection comes from an environment variable; the query uses parameters. There is no "Llama 14B" model — we use qwen2.5:14b. We removed the figures "70 % of enquiries deflected" and the percentages by type — they were never measured. We added: a notice that the bot is an AI (Article 50 of Regulation (EU) 2024/1689), a queue for an operator with a transcript, an evaluation with your own test messages, and a note on personal data.
01What you'll learn
- Why the model is a classifier and the database is the source of facts.
- How to get predictable JSON from Ollama with a schema and how to validate it.
- How to answer frequently asked questions without the model "inventing" an answer.
- How to check free seats with a safe query.
- When and how to hand the conversation to a person.
- What AI transparency and personal-data protection require.
02Before you start
- A machine of the NVIDIA GB10 class (for example ASUS Ascent GX10 or DGX Spark) with Ollama; Postgres (see the lesson n8n on GX10 for Docker and secrets); Python 3.10 or newer.
- Basic SQL and FastAPI. The WebSocket works in the same framework as in the lesson Event Planning Assistant.
- The model:
qwen2.5:14b— 9.0 GB, Apache-2.0 licence (ollama.com); small for 128 GB of shared memory. Its page lists 29+ languages, Bulgarian is not listed ⚠️ — so we recognise the intent and choose a ready answer instead of letting the model write free text for visitors. - Your own real visitor questions (without personal data) for the test in step 7. If you have none, write 20–30 invented ones.
03Steps
-
How it is put together
Why this way: when the model writes the price or the date it can get it wrong, and the error reaches the customer. So the facts are in the database and the text is a template.
diagramVisitor ──WebSocket──► FastAPI (127.0.0.1 only) │ 1. "I am an AI" — the first message ▼ classify(): model → {type, confidence, event type, quantity, date} → validated in code ├─ booking → Postgres query → template with up to 3 events (no confirmation!) ├─ faq → the model picks a question number → the stored answer verbatim └─ complaint / other / low confidence / nothing found → "handoffs" queue + a "handing over to an operator" message -
Tables with invented data
Prices are in EUR. The answers in
faqare invented — replace them with your own.sql · schema.sql · not runCREATE TABLE events ( id serial PRIMARY KEY, title text NOT NULL, event_type text NOT NULL CHECK (event_type IN ('theatre','cinema','concert','sport')), starts_at timestamptz NOT NULL, capacity integer NOT NULL CHECK (capacity > 0), price_eur numeric(8,2) NOT NULL ); CREATE TABLE tickets ( id serial PRIMARY KEY, event_id integer NOT NULL REFERENCES events(id), status text NOT NULL DEFAULT 'sold' CHECK (status IN ('held','sold')), created_at timestamptz NOT NULL DEFAULT now() ); CREATE TABLE faq ( id serial PRIMARY KEY, question text NOT NULL, answer text NOT NULL ); CREATE TABLE handoffs ( id serial PRIMARY KEY, session_id text NOT NULL, reason text NOT NULL, transcript jsonb NOT NULL, handled boolean NOT NULL DEFAULT false, created_at timestamptz NOT NULL DEFAULT now() ); -- invented data INSERT INTO events (title, event_type, starts_at, capacity, price_eur) VALUES ('Jazz Evening', 'concert', '2026-11-14 19:30+02', 300, 12.00), ('Comedy "The Neighbours"', 'theatre', '2026-11-20 19:00+02', 150, 9.00), ('Children''s Film', 'cinema', '2026-11-22 11:00+02', 120, 5.00); INSERT INTO faq (question, answer) VALUES ('Is there parking?', 'Example answer: there is parking in front of the building; confirm the place and any fee with your own complex.'), ('What are the box office hours?', 'Example answer: the box office is open on weekdays; confirm the hours with your own complex.'), ('Are there discounts for children and students?', 'Example answer: discounts are announced for each event; confirm them with your own complex.'); -
Recognising the intent
Why a schema: Ollama accepts a JSON schema in
formatand returns the answer as a JSON string inmessage.content. The schema is deliberately simple: for a missing value we use"none",0and an empty string. Why validate again in code: the schema constrains the shape, not the meaning. The model's confidence is not a computed probability — a threshold is tuned with a test (step 7).python · bot.py (part 1) · not run# bot.py · ticketing chatbot (invented data) import asyncio import json import os from datetime import date import httpx import psycopg from fastapi import FastAPI, WebSocket, WebSocketDisconnect OLLAMA = "http://localhost:11434" MODEL = "qwen2.5:14b" CONF_MIN = 0.65 # threshold for handing over to a person — tune it with the test in step 7 NOTICE = ( "Hello! I am an automated assistant (artificial intelligence), not a person. " "I can answer questions and check free tickets. " "I do not take payments, and please do not share personal or payment details here." ) INTENT_SCHEMA = { "type": "object", "properties": { "type": {"type": "string", "enum": ["booking", "faq", "complaint", "other"]}, "confidence": {"type": "number"}, "event_type": {"type": "string", "enum": ["theatre", "cinema", "concert", "sport", "none"]}, "quantity": {"type": "integer"}, "date": {"type": "string"}, }, "required": ["type", "confidence", "event_type", "quantity", "date"], } def ask_model(prompt: str, schema: dict) -> dict: r = httpx.post( f"{OLLAMA}/api/chat", json={ "model": MODEL, "stream": False, "format": schema, "messages": [{"role": "user", "content": prompt}], "options": {"temperature": 0}, }, timeout=120, ) r.raise_for_status() return json.loads(r.json()["message"]["content"]) def classify(msg: str) -> dict: prompt = ( "Classify a visitor's message to a cultural and sports complex and return " "JSON that follows the schema. confidence is from 0 to 1. If something is missing: " 'event_type="none", quantity=0, date="" (format YYYY-MM-DD).\n\n' "Message: " + msg[:500] ) try: d = ask_model(prompt, INTENT_SCHEMA) qty = int(d["quantity"]) try: day = date.fromisoformat(d["date"]) if d["date"] else None except ValueError: day = None return { "type": d["type"] if d["type"] in ("booking", "faq", "complaint", "other") else "other", "confidence": max(0.0, min(1.0, float(d["confidence"]))), "event_type": d["event_type"], "quantity": qty if 1 <= qty <= 20 else 0, "date": day, } except (httpx.HTTPError, KeyError, ValueError, TypeError): return {"type": "other", "confidence": 0.0, "event_type": "none", "quantity": 0, "date": None} -
Availability and answers to questions
Why a template: the price and the seats are taken from the database row — the model does not touch them. Why a question number: for questions the model picks an
idfrom a list and the code returns the stored answer verbatim;0means "no match" and leads to a person. The query uses parameters — the user's text never enters the SQL. The bot never says a ticket has been bought; it suggests, and confirmation and payment are outside the lesson.python · bot.py (part 2) · not rundef db(): return psycopg.connect(os.environ["DATABASE_URL"]) def find_events(event_type: str, qty: int, day: date | None): sql = ( "SELECT e.title, e.starts_at, e.price_eur, " " e.capacity - COUNT(t.id) AS free " "FROM events e LEFT JOIN tickets t ON t.event_id = e.id " "WHERE e.event_type = %s AND e.starts_at >= now() " " AND (%s::date IS NULL OR e.starts_at::date = %s::date) " "GROUP BY e.id " "HAVING e.capacity - COUNT(t.id) >= %s " "ORDER BY e.starts_at LIMIT 3" ) with db() as conn: return conn.execute(sql, (event_type, day, day, max(qty, 1))).fetchall() def booking_reply(intent: dict) -> str | None: if intent["event_type"] == "none": return None # we do not know what it is about → a person or a clarifying question rows = find_events(intent["event_type"], intent["quantity"], intent["date"]) if not rows: return "I found no free tickets for these criteria. Would you like to try another date?" lines = [ f"• {title} — {starts:%d.%m.%Y %H:%M}, {price} EUR per ticket, free seats: {free}" for title, starts, price, free in rows ] return ( "I found the following:\n" + "\n".join(lines) + "\nThis is a suggestion, not a reservation. Tickets are confirmed and paid " "on the ticket page or with an operator." ) def faq_reply(msg: str) -> str | None: with db() as conn: rows = conn.execute("SELECT id, question, answer FROM faq ORDER BY id").fetchall() if not rows: return None listing = "\n".join(f"{i}: {q}" for i, q, _ in rows) schema = { "type": "object", "properties": {"id": {"type": "integer"}, "confidence": {"type": "number"}}, "required": ["id", "confidence"], } prompt = ( "Which of the following questions best matches the message? Return the id, or 0 " "if none matches.\n" + listing + "\n\nMessage: " + msg[:500] ) try: pick = ask_model(prompt, schema) except (httpx.HTTPError, KeyError, ValueError, TypeError): return None answers = {i: a for i, _, a in rows} if pick.get("id") in answers and float(pick.get("confidence", 0)) >= CONF_MIN: return answers[pick["id"]] return None -
The chat and handoff to a person
Why a queue: handoff is not "the end of the conversation" but a row in a table with the transcript that an operator picks up. The transcript contains personal data — see the note below. The connection is only on
127.0.0.1and has no login — if you open it to the internet you need authentication, an origin check and rate limits.python · bot.py (part 3) · not rundef queue_handoff(sid: str, reason: str, history: list[dict]) -> None: with db() as conn: conn.execute( "INSERT INTO handoffs (session_id, reason, transcript) VALUES (%s, %s, %s)", (sid, reason, json.dumps(history, ensure_ascii=False)), ) def answer(sid: str, msg: str, history: list[dict]) -> str: intent = classify(msg) if intent["type"] == "complaint": queue_handoff(sid, "complaint", history) return "I am sorry for the inconvenience. I am handing the conversation to an operator." if intent["confidence"] < CONF_MIN: queue_handoff(sid, "low_confidence", history) return "I am not sure I understood you correctly. I am handing the conversation to an operator." reply = None if intent["type"] == "booking": reply = booking_reply(intent) elif intent["type"] == "faq": reply = faq_reply(msg) if reply is None: queue_handoff(sid, intent["type"], history) return "An operator will help you with this. I am handing the conversation over." return reply app = FastAPI(title="Ticketing chatbot") histories: dict[str, list[dict]] = {} @app.websocket("/ws/{sid}") async def chat(ws: WebSocket, sid: str): await ws.accept() await ws.send_text(NOTICE) history = histories.setdefault(sid, []) try: while True: msg = (await ws.receive_text())[:500] history.append({"role": "user", "content": msg}) reply = await asyncio.to_thread(answer, sid, msg, history) history.append({"role": "assistant", "content": reply}) del history[:-20] await ws.send_text(reply) except WebSocketDisconnect: histories.pop(sid, None)bash · not run# DATABASE_URL: e.g. "host=localhost dbname=tickets user=<user> password=<from .env>" export DATABASE_URL="<database-connection>" uvicorn bot:app --host 127.0.0.1 --port 8081 # in another terminal — a simple client connection python -m websockets ws://localhost:8081/ws/testYou expect the first message "I am an automated assistant…". Try: "I want 2 tickets for a concert", "Is there parking?", "Terrible service!". The last one should create a row in
handoffs. (The model was written for Bulgarian visitors — try the same messages in Bulgarian too.) -
A notice that you are an AI, and personal data
⚖️Transparency under the AI ActArticle 50(1) of Regulation (EU) 2024/1689 requires AI systems intended to interact directly with people to be designed so that people are informed they are dealing with an AI system — unless this is obvious from the point of view of a reasonably well-informed person (a paraphrase; see the official text on EUR-Lex). In the original text the regulation applies from 2 August 2026 — check whether it has been amended by the day you read this ⚠️. That is why the first message says the bot is an AI. Whether it applies to your case is a question for a lawyer.🔒Personal data in the chatConversation transcripts may contain names, phone numbers and email addresses — personal data under the GDPR. Keep only what is needed, decide and write down a deletion period (for example deletehandoffsafter handling), do not accept payments or identity documents in the chat, and describe the processing in your privacy policy. Here the model is local and the data does not leave the machine, but do not send it more than it needs, and do not keep transcripts longer than necessary. -
Measure with your own messages
Why: we have no percentages and promise none. Collect 20–50 messages with an expected type and run the script; see where it goes wrong, and only then set the threshold
CONF_MIN.python · eval.py · not run# eval.py · how often the type matches the expected one (on your own messages) from bot import classify CASES = [ # (message, expected type) — write your own; these are invented ("I want 2 tickets for the concert on 14 November", "booking"), ("Is there parking?", "faq"), ("Terrible service, I want a refund", "complaint"), ("Hello", "other"), ] ok = 0 for text, expected in CASES: got = classify(text) flag = "OK " if got["type"] == expected else "WRONG" ok += got["type"] == expected print(f"{flag} expected={expected:9} got={got['type']:9} confidence={got['confidence']:.2f} | {text}") print(f"Matches: {ok}/{len(CASES)}")
04Check
schema.sqlis applied and the invented data is loaded.- The first message of every session says the bot is an AI.
- On a broken or incomplete model answer the conversation goes to a person instead of crashing.
- A message about tickets returns up to three events with free seats from the database, in EUR.
- The bot never claims that a ticket has been bought or paid.
- A complaint and low confidence create a row in
handoffs. eval.pyhas been run on at least 20 own messages and the threshold was chosen from the result.- A retention period for transcripts is written down; the server listens only on
127.0.0.1.
Quiz
1. Who composes the answer about availability and price?
2. How does the bot answer a frequently asked question?
3. The confidence the model returns is:
4. What should the bot do at the start under Article 50(1) of Regulation (EU) 2024/1689 (unless obvious)?
05What's next
06Sources
- Ollama: structured outputs 🔒 local —
formatwith a schema,stream: false(read on 03.10.2026). - Ollama: qwen2.5 🔒 local — size 9.0 GB, licence, languages.
- FastAPI: WebSockets —
accept,receive_text,WebSocketDisconnect. - psycopg 3 — parameterised queries.
- EUR-Lex: Regulation (EU) 2024/1689 — Article 50 and Article 113 (application).
- fastapi 0.142.2 · uvicorn 0.54.0 · httpx 0.28.1 · psycopg 3.3.6 — versions on PyPI as of 03.10.2026.