Multimodal AI: images, voice and documents
Text models see only text. But work arrives as photos, voice notes and scanned documents. In this lesson we teach local models to look, listen and read — and to return not free text but verifiable structured data.
qwen2-vl:7b and phi-3.5-vision do not exist in the Ollama library — the old code could not start. They are replaced with checked tags (qwen2.5vl, qwen3-vl, gemma3, gemma4); sizes come from ollama.com. Structured output: Instructor pointed at a different server (vLLM) with an Ollama model name is replaced with with_structured_output on Ollama itself; regex over JSON is gone. Bugs: rr'…' (syntax error), glob("*.{jpg,png}") (pathlib does not understand braces — the batch found no photos at all), bare except:. Whisper: the table is checked against the official README (medium is 769M, not 307M; VRAM and speed from there). The "Bulgarian WER" figures (18% / 8% / 4.5%) could not be found in any source — removed. WhisperX: new import whisperx.diarize.DiarizationPipeline(token=…); for Bulgarian there is no default alignment model — the old code failed with an error. Voice: XTTS-v2 does not support Bulgarian (17 languages) — removed; edge-tts stays with the Kalina and Borislav voices (checked) and a warning that it is a cloud service. Documents: MarkItDown returns .markdown; Docling has a new address; dead links replaced. Invoice amounts are in EUR. Unconfirmed figures ("87% match", "98.7% detection", "45 → 3 min", "5–6 images/sec") are removed; the cases are marked as illustrative. The example with personal and health data is replaced with a GDPR-safe version.
01What you'll learn
- How to choose a vision model from the Ollama library — by task, memory and language.
- How to send a photo to the model and get valid JSON that follows a schema, not text to parse.
- Which Whisper model is for what, why
turbodoes not translate, and how to measure quality on your own recordings. - How WhisperX adds "who said what" and what is missing for Bulgarian.
- How to read PDF, Word, Excel and scanned documents with the right tool for each.
- Which photos and recordings are personal data and how to process them safely.
02Before you start
- You have done Blocks 0–2 (especially structured output). The RAG lesson (Block 4) helps with large documents.
- Python 3.10+ and a virtual environment.
ffmpegon the system — Whisper reads audio through it. - Ollama 🔒 local. Vision models need more memory than their download size — context and parallel requests add to it.
- For Bulgarian OCR: Tesseract with the
bullanguage pack. - For WhisperX: a GPU (recommended), a Hugging Face account 🌐 global with a read token, and the diarization model's terms accepted.
python3 -m venv .venv && source .venv/bin/activate
pip install -U ollama langchain-ollama pydantic requests
pip install -U openai-whisper edge-tts
pip install -U pdfplumber "markitdown[all]" docling "unstructured[all-docs]"
# optional, for diarization: pip install -U whisperx
# checked with: langchain-ollama 1.1.0 · openai-whisper 20250625 · whisperx 3.8.6 · docling 2.131 · markitdown 0.1.8
ollama pull qwen2.5vl:7b # vision model
ollama pull llama3.1:8b # text model for structuring03Steps
-
The three modalities and when you need them
Multimodal means the input is more than text. In practice these are three separate pipelines, each with its own tools. Why keep them apart? Because the errors differ: a vision model "sees" things that are not there; Whisper mishears names; OCR confuses letters. Each pipeline needs its own check.
Modality Tools Typical tasks 👁️ Image Vision models in Ollama: Qwen2.5-VL, Qwen3-VL, Gemma 3/4, LLaVA Photos of properties and sites, screenshots, diagrams, paper invoices 🎙️ Voice Whisper, WhisperX (+ pyannote), text-to-speech Field voice notes, meeting minutes, narration 📄 Documents pdfplumber, Unstructured + Tesseract, Docling, MarkItDown PDFs with tables, scanned documents, Word, Excel -
Choosing a vision model
A vision language model takes image + text and returns text. You choose by three things: how much memory you have, how small the text it must read is, and the language. The table follows the ollama.com pages as of 01.10.2026; the size is the download size.
Model (Ollama tag) Sizes Download Note qwen2.5vl3b · 7b · 32b · 72b 7b: 6.0 GB Strong on text, tables and layout inside images — a good start for documents. qwen3-vl2b · 4b · 8b · 30b · 32b · 235b 8b: 6.1 GB The newer Qwen generation, 256K context. gemma34b · 12b · 27b (vision) 12b: 8.1 GB Over 140 languages according to Google, 128K context. The 270m and 1b versions are text-only. gemma4e2b · e4b · 12b · 26b · 31b 12b: 7.7–8.0 GB Text + image; the new Gemma generation. llava7b · 13b · 34b 13b: 8.0 GB LLaVA 1.6 — the classic, but not updated for about two years. llama3.2-vision11b · 90b — An alternative from Meta. ⚠️Do not trust a table for "quality in Bulgarian"The old lesson rated "BG text: excellent / acceptable" with no source. We have not measured these models on Bulgarian text in images. Run a small test: 20 of your own photos with a known correct answer, the same prompt, compare. That tells you more than any ranking.💡Memory and parallel requestsBy default Ollama processes one request per model at a time (OLLAMA_NUM_PARALLEL=1). If you raise it, the memory needed grows in proportion to the context. "Images per second" depends on the machine, the model and the image size — measure it on your own setup. -
Image → answer → structured output
Three ways, from the simplest to the most useful. The third is the one that goes to production: the model is forced to return JSON that follows your schema, and Pydantic validates it. No more "almost JSON" that breaks the pipeline a week later.
Python · vision_pipeline.pyimport base64 from pathlib import Path from typing import Literal import requests from ollama import chat from pydantic import BaseModel, Field from langchain_ollama import ChatOllama from langchain_core.messages import HumanMessage VISION_MODEL = "qwen2.5vl:7b" # or gemma3:12b, qwen3-vl:8b, llava:13b IMAGE_EXT = {".jpg", ".jpeg", ".png", ".webp"} # ── Helper: file or URL → base64 ── def img_to_b64(path_or_url: str) -> str: if path_or_url.startswith("http"): data = requests.get(path_or_url, timeout=10).content else: data = Path(path_or_url).read_bytes() return base64.b64encode(data).decode() # ── Method 1: the official Ollama Python client ── def vision_ollama(image_path: str, prompt: str, model: str = VISION_MODEL) -> str: r = chat(model=model, messages=[{"role": "user", "content": prompt, "images": [image_path]}], options={"temperature": 0}) return r.message.content # ── Method 2: LangChain ChatOllama ── vision_llm = ChatOllama(model=VISION_MODEL, temperature=0) def image_message(image_path: str, prompt: str) -> HumanMessage: b64 = img_to_b64(image_path) return HumanMessage(content=[ {"type": "text", "text": prompt}, {"type": "image_url", "image_url": f"data:image/jpeg;base64,{b64}"}, ]) def vision_langchain(image_path: str, prompt: str) -> str: return vision_llm.invoke([image_message(image_path, prompt)]).content # ── Method 3: structured output (JSON schema, validated by Pydantic) ── class PropertyPhotoAnalysis(BaseModel): room_type: Literal["living_room", "bedroom", "kitchen", "bathroom", "exterior", "other"] condition: Literal["excellent", "good", "average", "poor"] natural_light: Literal["high", "medium", "low"] notable_features: list[str] visible_issues: list[str] renovation_needed: Literal["none", "cosmetic", "moderate", "major"] marketing_score: int = Field(ge=1, le=10, description="Suitability for a listing, 1–10") photo_analyzer = vision_llm.with_structured_output(PropertyPhotoAnalysis) def analyze_property_photo(image_path: str) -> PropertyPhotoAnalysis: return photo_analyzer.invoke([image_message( image_path, "Analyse this property photo. Fill in every field.")]) # ── Batch: every photo in a folder ── def batch_analyze_photos(image_dir: str) -> list[dict]: results = [] files = sorted(p for p in Path(image_dir).iterdir() if p.suffix.lower() in IMAGE_EXT) for img in files: try: a = analyze_property_photo(str(img)) results.append({"file": img.name, "analysis": a.model_dump()}) print(f"✅ {img.name}: {a.room_type} ({a.condition})") except Exception as e: # one bad photo does not stop the batch print(f"❌ {img.name}: {e}") return sorted(results, key=lambda r: r["analysis"]["marketing_score"], reverse=True)✅WhyLiteraland not a description like "living_room|bedroom"When the allowed values are in the type, they go into the JSON schema and Ollama restricts the output to them. A description is only a request; a type is a rule.⚠️Trap from the old code:glob("*.{jpg,png}")pathlibdoes not understand curly braces — such a batch silently finds no photos and "passes" without an error. Filter bysuffixas above and check the file count before you run it. -
Real cases: defects, invoices, identity documents
The same technique — image + schema — covers many tasks. The difference is in the rules around the model: what you check in code, what you never extract, and when you call a human.
Python · domain_vision.pyfrom typing import Literal from pydantic import BaseModel, Field from vision_pipeline import vision_llm, image_message # ── Case 1: construction defects ── class DefectReport(BaseModel): defect_types: list[Literal["crack", "damp", "corrosion", "deformation", "other"]] severity: Literal["critical", "major", "minor", "cosmetic"] location: str = Field(description="Where the defect is in the photo") needs_inspection: bool = Field(description="Does a specialist need to inspect it") notes: str def analyze_construction_defect(image_path: str) -> DefectReport: return vision_llm.with_structured_output(DefectReport).invoke([image_message( image_path, "Inspect the photo for construction defects. Do not guess — describe only what is visible.")]) # ── Case 2: invoice (amounts in EUR) ── class InvoiceLine(BaseModel): description: str qty: float unit_price_eur: float total_eur: float class Invoice(BaseModel): invoice_number: str date: str = Field(description="DD.MM.YYYY") vendor_name: str vendor_eik: str | None buyer_name: str items: list[InvoiceLine] subtotal_eur: float vat_eur: float total_eur: float def read_invoice(image_path: str) -> Invoice: inv = vision_llm.with_structured_output(Invoice).invoke([image_message( image_path, "Read the invoice. If a field is unreadable, leave it empty — do not invent.")]) if abs(inv.subtotal_eur + inv.vat_eur - inv.total_eur) > 0.02: # arithmetic is checked in code raise ValueError(f"Totals do not add up on invoice {inv.invoice_number} — to a human") return inv # ── Case 3: ID card — CHECK ONLY, no data extraction ── class IdCheck(BaseModel): looks_like_id_card: bool expiry_visible: bool concerns: list[str] = Field(description="Visible concerns; no personal data") def check_id_document(image_path: str) -> IdCheck: return vision_llm.with_structured_output(IdCheck).invoke([image_message( image_path, "Only check the document type. Do NOT copy names, numbers, addresses or dates.")])⛔Personal and health dataA photo of an ID card, a prescription, a discharge summary or a sick note is personal information, and health data is a special category under the GDPR. Process it only locally 🔒, with a clear legal basis, extract the minimum and do not keep the original images longer than needed. The old lesson had a prompt that copied patient names and diagnoses — it is deliberately absent here. A vision model is not a check of a document's authenticity; its output is a signal for a human.🏢Illustrative case: photos for a property listingAn agent uploads dozens of photos of one property.batch_analyze_photosranks them bymarketing_score, proposes the best ones for the listing and surfacesvisible_issuesbefore the viewing. The old lesson quoted "45 min → 3 min" and "87% match with the agent" — we have no source for these numbers and they are removed. Compare with your agent's rating on 50 photos before promising a result. -
Voice → text with Whisper
Whisper is OpenAI's speech-recognition model; it runs fully locally. The table is from the official README. "Speed" is relative to
large, measured by OpenAI on English speech on an A100 GPU — for Bulgarian and on your hardware it will differ.Model Parameters Memory (VRAM) Speed When tiny39 M ~1 GB ~10× Trial, weak hardware base74 M ~1 GB ~7× Trial small244 M ~2 GB ~4× Weak machine, simple recordings medium769 M ~5 GB ~2× When you also need translation large(v3)1550 M ~10 GB 1× Highest accuracy, critical recordings turbo809 M ~6 GB ~8× Good balance; an optimised large-v3⚠️Two things the old lesson missed1.turbois not trained for translation — withtask="translate"it returns the original language. To translate into English usemediumorlarge. 2. Accuracy varies a lot by language. The "Bulgarian WER 18% / 8% / 4.5%" figures from the old lesson have no source and are removed. Official per-language values are in the figure in the README and in appendix D of the paper. Measure it yourself: 10–15 minutes of your recordings, a hand-corrected transcript, compare.Python · audio_pipeline.pyimport asyncio, os from functools import lru_cache from typing import Literal import whisper from pydantic import BaseModel, Field from langchain_ollama import ChatOllama llm = ChatOllama(model="llama3.1:8b", temperature=0) # ── 1. Transcription with Whisper (the model loads once) ── @lru_cache(maxsize=2) def get_whisper(size: str): return whisper.load_model(size) def transcribe(audio_path: str, model_size: str = "turbo", language: str = "bg") -> dict: result = get_whisper(model_size).transcribe( audio_path, language=language, # set the language — do not rely on auto-detection task="transcribe", # turbo does NOT translate; for translation use medium or large word_timestamps=True, condition_on_previous_text=True, ) segs = result["segments"] return {"text": result["text"], "language": result["language"], "duration": segs[-1]["end"] if segs else 0.0, "segments": segs} # ── 2. WhisperX: who said what ── def transcribe_with_speakers(audio_path: str, align_model: str, num_speakers: int | None = None) -> list[dict]: """pip install whisperx · HF_TOKEN in the environment · accepted terms for pyannote/speaker-diarization-community-1 on Hugging Face. align_model: a Bulgarian wav2vec2 model from Hugging Face — WhisperX has no default one.""" import whisperx from whisperx.diarize import DiarizationPipeline device = "cuda" # without a GPU: device="cpu", compute_type="int8" model = whisperx.load_model("large-v3", device, compute_type="float16", language="bg") audio = whisperx.load_audio(audio_path) result = model.transcribe(audio, batch_size=16) model_a, meta = whisperx.load_align_model(language_code="bg", device=device, model_name=align_model) result = whisperx.align(result["segments"], model_a, meta, audio, device) diarize = DiarizationPipeline(token=os.environ["HF_TOKEN"], device=device) result = whisperx.assign_word_speakers(diarize(audio, num_speakers=num_speakers), result) return [{"speaker": s.get("speaker", "SPEAKER_00"), "start": round(s["start"], 1), "end": round(s["end"], 1), "text": s["text"].strip()} for s in result["segments"]] # ── 3. From transcript to structure ── class ActionItem(BaseModel): who: str what: str deadline: str | None class MeetingNotes(BaseModel): summary: str = Field(description="2–3 sentences") decisions: list[str] action_items: list[ActionItem] class CaseNote(BaseModel): problem: str urgency: Literal["low", "medium", "high", "critical"] actions_mentioned: list[str] next_visit: str | None def structure_transcript(transcript: str, schema: type[BaseModel]) -> BaseModel: return llm.with_structured_output(schema).invoke( "Extract the information ONLY from the transcript. Missing field → null.\n\n" f"Transcript:\n{transcript[:8000]}") # ── 4. Text-to-speech: edge-tts (Microsoft online service) ── async def _speak(text: str, out: str, voice: str) -> None: import edge_tts await edge_tts.Communicate(text, voice).save(out) def text_to_speech(text: str, out: str = "out.mp3", voice: str = "bg-BG-KalinaNeural") -> str: """Bulgarian voices: bg-BG-KalinaNeural, bg-BG-BorislavNeural. The text goes to the cloud — do not send personal data.""" asyncio.run(_speak(text, out, voice)) return out # ── The whole chain: voice note → structured case ── def process_voice_note(audio_path: str, worker_id: str) -> dict: t = transcribe(audio_path, model_size="turbo") case = structure_transcript(t["text"], CaseNote) return {"worker_id": worker_id, "duration_s": t["duration"], "transcript": t["text"], "case": case.model_dump()}⚠️WhisperX and BulgarianWhisperX aligns words with a separate wav2vec2 model. Its default list has no Bulgarian — withoutmodel_name,load_align_modelstops with an error. Find a wav2vec2 model trained on Bulgarian on Hugging Face, check its licence and pass it asalign_model. ⚠️ We have not tested a specific model — so we do not recommend a name.🔊Text-to-speech in BulgarianThe old lesson used XTTS-v2 withlanguage="bg". XTTS-v2 supports 17 languages and Bulgarian is not one of them. edge-tts has two Bulgarian voices (Kalina, Borislav — checked withedge-tts --list-voices; our test produced Bulgarian audio), but it is an unofficial client for an online service 🌐: the text leaves your machine. ⚠️ We have not verified a good fully local Bulgarian voice.example · field voice note → structure (fictional data, translated)Transcript: "I visited the client from case 214. She lives alone, the neighbours have not seen her for several days, she has trouble moving around. She needs a check-up by her GP and we should find out if she has relatives. Next visit — Thursday." CaseNote: problem: "Lives alone, limited mobility, risk of isolation" urgency: "medium" actions_mentioned: ["GP check-up", "check for relatives"] next_visit: "Thursday"⛔Recordings from social and health workThey contain personal and often health data. Work with a case number, not a name and address; keep the transcript and the model local; the output is a draft for the social worker, not a decision. -
Documents: the right tool for each kind
The first question for any PDF: does it have a text layer? If it does (a PDF "printed" from Word), the text is read directly with no recognition errors. If it is scanned, it is a picture and needs OCR, which always brings errors.
Tool Strength Limitation When pdfplumber Exact text and tables from PDFs with a text layer No OCR Invoices and reports generated by software Unstructured Many formats; hi_resdetects layout and tables; OCR with Tesseracthi_resis slow; multi-column pages may come out in the wrong orderScanned documents; general entry point Docling Layout and complex tables; export to Markdown Heavier (layout models) Financial statements, technical documents MarkItDown Word, Excel, PowerPoint, PDF → Markdown No OCR by default Office files into an LLM Tesseract OCR, bullanguage packText only, no layout Under the hood of Unstructured Python · document_intelligence.pyfrom pathlib import Path import pdfplumber from langchain_ollama import ChatOllama llm = ChatOllama(model="llama3.1:8b", temperature=0) # ── 1. Unstructured: any format, OCR with Tesseract ── def parse_document(file_path: str, scanned: bool = False, languages: tuple[str, ...] = ("bul", "eng")) -> dict: from unstructured.partition.auto import partition elements = partition( filename=file_path, strategy="hi_res" if scanned else "auto", # hi_res: layout + tables, slower languages=list(languages), # Tesseract codes: bul, eng infer_table_structure=True, ) out = {"titles": [], "tables": [], "text_blocks": []} for el in elements: if el.category == "Title": out["titles"].append(el.text) elif el.category == "Table": out["tables"].append({"text": el.text, "html": el.metadata.text_as_html}) elif el.category == "NarrativeText": out["text_blocks"].append(el.text) out["full_text"] = "\n\n".join(e.text for e in elements if e.text.strip()) return out # ── 2. pdfplumber: tables from PDFs with a text layer ── def extract_tables_pdfplumber(pdf_path: str) -> list[dict]: tables = [] with pdfplumber.open(pdf_path) as pdf: for n, page in enumerate(pdf.pages, start=1): for t in page.extract_tables(): if t: tables.append({"page": n, "rows": len(t), "cols": max(len(r) for r in t), "data": t}) return tables # ── 3. Docling: complex tables and layout ── def parse_with_docling(file_path: str) -> dict: from docling.document_converter import DocumentConverter doc = DocumentConverter().convert(file_path).document return {"markdown": doc.export_to_markdown(), "tables": [t.export_to_dataframe(doc=doc) for t in doc.tables], "num_pages": len(doc.pages)} # ── 4. MarkItDown: Office → Markdown ── def office_to_markdown(file_path: str) -> str: from markitdown import MarkItDown return MarkItDown().convert(file_path).markdown # ── 5. Router: picks the tool by file type ── def has_text_layer(pdf_path: str, min_chars: int = 50) -> bool: with pdfplumber.open(pdf_path) as pdf: return any(len((p.extract_text() or "").strip()) >= min_chars for p in pdf.pages[:3]) def smart_parse(file_path: str) -> dict: suffix = Path(file_path).suffix.lower() if suffix in {".docx", ".pptx", ".xlsx"}: return {"type": suffix[1:], "markdown": office_to_markdown(file_path)} if suffix == ".pdf": if has_text_layer(file_path): with pdfplumber.open(file_path) as pdf: text = "\n".join(p.extract_text() or "" for p in pdf.pages) return {"type": "native_pdf", "text": text, "tables": extract_tables_pdfplumber(file_path)} return {"type": "scanned_pdf", **parse_document(file_path, scanned=True)} return {"type": "other", **parse_document(file_path)} # ── 6. Question to a small document, no vector database ── def document_qa(file_path: str, question: str, max_chars: int = 8000) -> str: d = smart_parse(file_path) text = (d.get("text") or d.get("markdown") or d.get("full_text", ""))[:max_chars] return llm.invoke(f"DOCUMENT:\n{text}\n\nQUESTION: {question}\n" "Answer ONLY from the document. If the answer is not there, say so.").content✅We checked the routerOn 01.10.2026 we ransmart_parseon a sample PDF with a table, a Word file and an Excel file: the PDF was recognised as having a text layer and the table was extracted row by row; Word and Excel come out as Markdown. The old router looked only at the first page — here we look at up to three, so a PDF with a blank cover is not mistaken for a scan.💼Illustrative case: accounting with mixed invoicesInvoices arrive as software-generated PDFs, scans and Excel files.smart_parsedetects the type and picks the tool, andInvoicefrom step 4 structures the data and checks the totals in code. The old lesson's figures ("500 invoices", "98.7% detection") have no source and are removed. A document longer than a few pages goes through RAG (Block 4), not whole into the prompt.
04Check
Checklist
- The vision model is pulled from the Ollama library and answers a test photo.
with_structured_outputreturns a valid object for at least one photo.- The batch skips files that are not photos and does not stop on one bad image.
- Invoice totals are checked in code; a mismatch goes to a human.
- The identity-document check returns no personal data.
- Whisper runs with
language="bg"; the model is chosen by measured accuracy on your recordings. - For WhisperX: the token is in the environment, the model terms are accepted, a Bulgarian alignment model is passed.
- No personal data goes to online voices.
- Tesseract has the
bulpack;smart_parsetells a text PDF, a scanned PDF and an Office file apart.
Quiz
1. When is the smaller vision model (e.g. 7b instead of 32b) the sensible choice?
2. You need to translate a Bulgarian recording into English with Whisper. Which model does not work?
3. smart_parse detects a PDF with a text layer. What does that mean?
4. WhisperX stops with "No default align-model for language: bg". What do you do?
05What's next
06Sources
- Visual Instruction Tuning (Liu et al., 2023) — the LLaVA paper.
- Ollama: qwen2.5vl, qwen3-vl, gemma3, gemma4, llava — tags and sizes.
- Ollama: vision — how to send images.
- Ollama: structured outputs — JSON schema, including for vision models.
- LangChain: ChatOllama — images and structured output.
- OpenAI Whisper — models, memory, speed,
turboand translation; the paper (appendix D — accuracy by language). - WhisperX — word-level alignment and diarization.
- pyannote speaker-diarization-community-1 🌐 — the "who speaks" model; requires accepted terms.
- XTTS-v2 — the list of 17 supported languages.
- edge-tts 🌐 — text-to-speech;
--list-voices. - Unstructured: partitioning — strategies and OCR languages.
- Docling — documentation (the new address).
- MarkItDown — Office and PDF to Markdown.
- pdfplumber — text and tables from PDF.
- Tesseract tessdata — language packs, including
bul.