The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
01-07a · Multimodal AI: images, voice and documents
series
Blocks 0–10 · Block 7 — Multimodal AI · Part 1/2
level
Advanced
duration
3–5 h
prerequisites
Blocks 0–2 (LLM basics, inference, prompting, structured outputs); Block 4 (RAG) helpful; Python ≥ 3.10; Ollama
trust_label
VERIFIED 2026-10-01 (Ollama library tags and sizes, Whisper model table, WhisperX API and align-model list, XTTS-v2 language list, edge-tts Bulgarian voices, Unstructured/Docling/MarkItDown APIs, all source links; document router run on sample PDF/DOCX/XLSX; ChatOllama image message conversion checked; edge-tts produced Bulgarian audio) · UPDATED 2026-10-01 · vision and Whisper calls NOT run against a live model
versions
langchain-ollama 1.1.x · ollama (python) 0.6.x · openai-whisper 20250625 · whisperx 3.8.x · pyannote.audio 4.0.x · docling 2.x · markitdown 0.1.x · unstructured 0.27.x · pdfplumber 0.11.x · edge-tts 7.2.x
language
english edition · bulgarian original: /academy/blokove/moduli/01-07a_Блок_7_Част_1_Multimodal.html
next
01-07b_Блок_7_Част_2_Multimodal_Pipeline.html · multimodal pipeline and production use cases
PURPOSE

Extend text-only LLM systems to images, audio and real-world documents, running locally. Pick a vision model from the Ollama library, send images through the Ollama client or LangChain ChatOllama, and force JSON-schema output validated by Pydantic. Transcribe Bulgarian speech with Whisper, add speaker labels with WhisperX + pyannote, structure transcripts with an LLM, and generate Bulgarian speech. Route documents to the right parser (pdfplumber, Unstructured + Tesseract OCR, Docling, MarkItDown). Treat images of IDs, medical and social records as personal data.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

01-07b_Блок_7_Част_2_Multimodal_Pipeline.html · end-to-end multimodal pipeline and production use cases · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
multimodalvisionvlmollamastructured-outputwhisperwhisperxdiarizationttsocrdocument-parsinggdpr
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

Multimodal AI: images, voice and documents

Text models see only text. But work arrives as photos, voice notes and scanned documents. In this lesson we teach local models to look, listen and read — and to return not free text but verifiable structured data.

⏱ 3–5 h Advanced Block 7 · Part 1/2 image · voice · documents
Ollama + vision models🔒 local Whisper · WhisperX🔒 local Unstructured · Docling · MarkItDown · pdfplumber · Tesseract🔒 local Hugging Face (model downloads)🌐 global edge-tts (text-to-speech)🌐 global
🔄
UPDATED · 01.10.2026 — what changed
Models: qwen2-vl:7b and phi-3.5-vision do not exist in the Ollama library — the old code could not start. They are replaced with checked tags (qwen2.5vl, qwen3-vl, gemma3, gemma4); sizes come from ollama.com. Structured output: Instructor pointed at a different server (vLLM) with an Ollama model name is replaced with with_structured_output on Ollama itself; regex over JSON is gone. Bugs: rr'…' (syntax error), glob("*.{jpg,png}") (pathlib does not understand braces — the batch found no photos at all), bare except:. Whisper: the table is checked against the official README (medium is 769M, not 307M; VRAM and speed from there). The "Bulgarian WER" figures (18% / 8% / 4.5%) could not be found in any source — removed. WhisperX: new import whisperx.diarize.DiarizationPipeline(token=…); for Bulgarian there is no default alignment model — the old code failed with an error. Voice: XTTS-v2 does not support Bulgarian (17 languages) — removed; edge-tts stays with the Kalina and Borislav voices (checked) and a warning that it is a cloud service. Documents: MarkItDown returns .markdown; Docling has a new address; dead links replaced. Invoice amounts are in EUR. Unconfirmed figures ("87% match", "98.7% detection", "45 → 3 min", "5–6 images/sec") are removed; the cases are marked as illustrative. The example with personal and health data is replaced with a GDPR-safe version.

01What you'll learn

02Before you start

bash · install (versions as of 01.10.2026)
python3 -m venv .venv && source .venv/bin/activate
pip install -U ollama langchain-ollama pydantic requests
pip install -U openai-whisper edge-tts
pip install -U pdfplumber "markitdown[all]" docling "unstructured[all-docs]"
# optional, for diarization: pip install -U whisperx
# checked with: langchain-ollama 1.1.0 · openai-whisper 20250625 · whisperx 3.8.6 · docling 2.131 · markitdown 0.1.8

ollama pull qwen2.5vl:7b      # vision model
ollama pull llama3.1:8b       # text model for structuring

03Steps

  1. The three modalities and when you need them

    Multimodal means the input is more than text. In practice these are three separate pipelines, each with its own tools. Why keep them apart? Because the errors differ: a vision model "sees" things that are not there; Whisper mishears names; OCR confuses letters. Each pipeline needs its own check.

    ModalityToolsTypical tasks
    👁️ ImageVision models in Ollama: Qwen2.5-VL, Qwen3-VL, Gemma 3/4, LLaVAPhotos of properties and sites, screenshots, diagrams, paper invoices
    🎙️ VoiceWhisper, WhisperX (+ pyannote), text-to-speechField voice notes, meeting minutes, narration
    📄 Documentspdfplumber, Unstructured + Tesseract, Docling, MarkItDownPDFs with tables, scanned documents, Word, Excel
  2. Choosing a vision model

    A vision language model takes image + text and returns text. You choose by three things: how much memory you have, how small the text it must read is, and the language. The table follows the ollama.com pages as of 01.10.2026; the size is the download size.

    Model (Ollama tag)SizesDownloadNote
    qwen2.5vl3b · 7b · 32b · 72b7b: 6.0 GBStrong on text, tables and layout inside images — a good start for documents.
    qwen3-vl2b · 4b · 8b · 30b · 32b · 235b8b: 6.1 GBThe newer Qwen generation, 256K context.
    gemma34b · 12b · 27b (vision)12b: 8.1 GBOver 140 languages according to Google, 128K context. The 270m and 1b versions are text-only.
    gemma4e2b · e4b · 12b · 26b · 31b12b: 7.7–8.0 GBText + image; the new Gemma generation.
    llava7b · 13b · 34b13b: 8.0 GBLLaVA 1.6 — the classic, but not updated for about two years.
    llama3.2-vision11b · 90b—An alternative from Meta.
    ⚠️
    Do not trust a table for "quality in Bulgarian"
    The old lesson rated "BG text: excellent / acceptable" with no source. We have not measured these models on Bulgarian text in images. Run a small test: 20 of your own photos with a known correct answer, the same prompt, compare. That tells you more than any ranking.
    💡
    Memory and parallel requests
    By default Ollama processes one request per model at a time (OLLAMA_NUM_PARALLEL=1). If you raise it, the memory needed grows in proportion to the context. "Images per second" depends on the machine, the model and the image size — measure it on your own setup.
  3. Image → answer → structured output

    Three ways, from the simplest to the most useful. The third is the one that goes to production: the model is forced to return JSON that follows your schema, and Pydantic validates it. No more "almost JSON" that breaks the pipeline a week later.

    Python · vision_pipeline.py
    import base64
    from pathlib import Path
    from typing import Literal
    import requests
    from ollama import chat
    from pydantic import BaseModel, Field
    from langchain_ollama import ChatOllama
    from langchain_core.messages import HumanMessage
    
    VISION_MODEL = "qwen2.5vl:7b"          # or gemma3:12b, qwen3-vl:8b, llava:13b
    IMAGE_EXT = {".jpg", ".jpeg", ".png", ".webp"}
    
    # ── Helper: file or URL → base64 ──
    def img_to_b64(path_or_url: str) -> str:
        if path_or_url.startswith("http"):
            data = requests.get(path_or_url, timeout=10).content
        else:
            data = Path(path_or_url).read_bytes()
        return base64.b64encode(data).decode()
    
    # ── Method 1: the official Ollama Python client ──
    def vision_ollama(image_path: str, prompt: str, model: str = VISION_MODEL) -> str:
        r = chat(model=model,
                 messages=[{"role": "user", "content": prompt, "images": [image_path]}],
                 options={"temperature": 0})
        return r.message.content
    
    # ── Method 2: LangChain ChatOllama ──
    vision_llm = ChatOllama(model=VISION_MODEL, temperature=0)
    
    def image_message(image_path: str, prompt: str) -> HumanMessage:
        b64 = img_to_b64(image_path)
        return HumanMessage(content=[
            {"type": "text", "text": prompt},
            {"type": "image_url", "image_url": f"data:image/jpeg;base64,{b64}"},
        ])
    
    def vision_langchain(image_path: str, prompt: str) -> str:
        return vision_llm.invoke([image_message(image_path, prompt)]).content
    
    # ── Method 3: structured output (JSON schema, validated by Pydantic) ──
    class PropertyPhotoAnalysis(BaseModel):
        room_type: Literal["living_room", "bedroom", "kitchen", "bathroom", "exterior", "other"]
        condition: Literal["excellent", "good", "average", "poor"]
        natural_light: Literal["high", "medium", "low"]
        notable_features: list[str]
        visible_issues: list[str]
        renovation_needed: Literal["none", "cosmetic", "moderate", "major"]
        marketing_score: int = Field(ge=1, le=10, description="Suitability for a listing, 1–10")
    
    photo_analyzer = vision_llm.with_structured_output(PropertyPhotoAnalysis)
    
    def analyze_property_photo(image_path: str) -> PropertyPhotoAnalysis:
        return photo_analyzer.invoke([image_message(
            image_path, "Analyse this property photo. Fill in every field.")])
    
    # ── Batch: every photo in a folder ──
    def batch_analyze_photos(image_dir: str) -> list[dict]:
        results = []
        files = sorted(p for p in Path(image_dir).iterdir() if p.suffix.lower() in IMAGE_EXT)
        for img in files:
            try:
                a = analyze_property_photo(str(img))
                results.append({"file": img.name, "analysis": a.model_dump()})
                print(f"✅ {img.name}: {a.room_type} ({a.condition})")
            except Exception as e:                     # one bad photo does not stop the batch
                print(f"❌ {img.name}: {e}")
        return sorted(results, key=lambda r: r["analysis"]["marketing_score"], reverse=True)
    ✅
    Why Literal and not a description like "living_room|bedroom"
    When the allowed values are in the type, they go into the JSON schema and Ollama restricts the output to them. A description is only a request; a type is a rule.
    ⚠️
    Trap from the old code: glob("*.{jpg,png}")
    pathlib does not understand curly braces — such a batch silently finds no photos and "passes" without an error. Filter by suffix as above and check the file count before you run it.
  4. Real cases: defects, invoices, identity documents

    The same technique — image + schema — covers many tasks. The difference is in the rules around the model: what you check in code, what you never extract, and when you call a human.

    Python · domain_vision.py
    from typing import Literal
    from pydantic import BaseModel, Field
    from vision_pipeline import vision_llm, image_message
    
    # ── Case 1: construction defects ──
    class DefectReport(BaseModel):
        defect_types: list[Literal["crack", "damp", "corrosion", "deformation", "other"]]
        severity: Literal["critical", "major", "minor", "cosmetic"]
        location: str = Field(description="Where the defect is in the photo")
        needs_inspection: bool = Field(description="Does a specialist need to inspect it")
        notes: str
    
    def analyze_construction_defect(image_path: str) -> DefectReport:
        return vision_llm.with_structured_output(DefectReport).invoke([image_message(
            image_path, "Inspect the photo for construction defects. Do not guess — describe only what is visible.")])
    
    # ── Case 2: invoice (amounts in EUR) ──
    class InvoiceLine(BaseModel):
        description: str
        qty: float
        unit_price_eur: float
        total_eur: float
    
    class Invoice(BaseModel):
        invoice_number: str
        date: str = Field(description="DD.MM.YYYY")
        vendor_name: str
        vendor_eik: str | None
        buyer_name: str
        items: list[InvoiceLine]
        subtotal_eur: float
        vat_eur: float
        total_eur: float
    
    def read_invoice(image_path: str) -> Invoice:
        inv = vision_llm.with_structured_output(Invoice).invoke([image_message(
            image_path, "Read the invoice. If a field is unreadable, leave it empty — do not invent.")])
        if abs(inv.subtotal_eur + inv.vat_eur - inv.total_eur) > 0.02:   # arithmetic is checked in code
            raise ValueError(f"Totals do not add up on invoice {inv.invoice_number} — to a human")
        return inv
    
    # ── Case 3: ID card — CHECK ONLY, no data extraction ──
    class IdCheck(BaseModel):
        looks_like_id_card: bool
        expiry_visible: bool
        concerns: list[str] = Field(description="Visible concerns; no personal data")
    
    def check_id_document(image_path: str) -> IdCheck:
        return vision_llm.with_structured_output(IdCheck).invoke([image_message(
            image_path, "Only check the document type. Do NOT copy names, numbers, addresses or dates.")])
    ⛔
    Personal and health data
    A photo of an ID card, a prescription, a discharge summary or a sick note is personal information, and health data is a special category under the GDPR. Process it only locally 🔒, with a clear legal basis, extract the minimum and do not keep the original images longer than needed. The old lesson had a prompt that copied patient names and diagnoses — it is deliberately absent here. A vision model is not a check of a document's authenticity; its output is a signal for a human.
    🏢
    Illustrative case: photos for a property listing
    An agent uploads dozens of photos of one property. batch_analyze_photos ranks them by marketing_score, proposes the best ones for the listing and surfaces visible_issues before the viewing. The old lesson quoted "45 min → 3 min" and "87% match with the agent" — we have no source for these numbers and they are removed. Compare with your agent's rating on 50 photos before promising a result.
  5. Voice → text with Whisper

    Whisper is OpenAI's speech-recognition model; it runs fully locally. The table is from the official README. "Speed" is relative to large, measured by OpenAI on English speech on an A100 GPU — for Bulgarian and on your hardware it will differ.

    ModelParametersMemory (VRAM)SpeedWhen
    tiny39 M~1 GB~10×Trial, weak hardware
    base74 M~1 GB~7×Trial
    small244 M~2 GB~4×Weak machine, simple recordings
    medium769 M~5 GB~2×When you also need translation
    large (v3)1550 M~10 GB1×Highest accuracy, critical recordings
    turbo809 M~6 GB~8×Good balance; an optimised large-v3
    ⚠️
    Two things the old lesson missed
    1. turbo is not trained for translation — with task="translate" it returns the original language. To translate into English use medium or large. 2. Accuracy varies a lot by language. The "Bulgarian WER 18% / 8% / 4.5%" figures from the old lesson have no source and are removed. Official per-language values are in the figure in the README and in appendix D of the paper. Measure it yourself: 10–15 minutes of your recordings, a hand-corrected transcript, compare.
    Python · audio_pipeline.py
    import asyncio, os
    from functools import lru_cache
    from typing import Literal
    import whisper
    from pydantic import BaseModel, Field
    from langchain_ollama import ChatOllama
    
    llm = ChatOllama(model="llama3.1:8b", temperature=0)
    
    # ── 1. Transcription with Whisper (the model loads once) ──
    @lru_cache(maxsize=2)
    def get_whisper(size: str):
        return whisper.load_model(size)
    
    def transcribe(audio_path: str, model_size: str = "turbo", language: str = "bg") -> dict:
        result = get_whisper(model_size).transcribe(
            audio_path,
            language=language,             # set the language — do not rely on auto-detection
            task="transcribe",             # turbo does NOT translate; for translation use medium or large
            word_timestamps=True,
            condition_on_previous_text=True,
        )
        segs = result["segments"]
        return {"text": result["text"], "language": result["language"],
                "duration": segs[-1]["end"] if segs else 0.0, "segments": segs}
    
    # ── 2. WhisperX: who said what ──
    def transcribe_with_speakers(audio_path: str, align_model: str,
                                 num_speakers: int | None = None) -> list[dict]:
        """pip install whisperx · HF_TOKEN in the environment · accepted terms for
        pyannote/speaker-diarization-community-1 on Hugging Face.
        align_model: a Bulgarian wav2vec2 model from Hugging Face — WhisperX has no default one."""
        import whisperx
        from whisperx.diarize import DiarizationPipeline
        device = "cuda"                                   # without a GPU: device="cpu", compute_type="int8"
        model = whisperx.load_model("large-v3", device, compute_type="float16", language="bg")
        audio = whisperx.load_audio(audio_path)
        result = model.transcribe(audio, batch_size=16)
        model_a, meta = whisperx.load_align_model(language_code="bg", device=device, model_name=align_model)
        result = whisperx.align(result["segments"], model_a, meta, audio, device)
        diarize = DiarizationPipeline(token=os.environ["HF_TOKEN"], device=device)
        result = whisperx.assign_word_speakers(diarize(audio, num_speakers=num_speakers), result)
        return [{"speaker": s.get("speaker", "SPEAKER_00"), "start": round(s["start"], 1),
                 "end": round(s["end"], 1), "text": s["text"].strip()} for s in result["segments"]]
    
    # ── 3. From transcript to structure ──
    class ActionItem(BaseModel):
        who: str
        what: str
        deadline: str | None
    
    class MeetingNotes(BaseModel):
        summary: str = Field(description="2–3 sentences")
        decisions: list[str]
        action_items: list[ActionItem]
    
    class CaseNote(BaseModel):
        problem: str
        urgency: Literal["low", "medium", "high", "critical"]
        actions_mentioned: list[str]
        next_visit: str | None
    
    def structure_transcript(transcript: str, schema: type[BaseModel]) -> BaseModel:
        return llm.with_structured_output(schema).invoke(
            "Extract the information ONLY from the transcript. Missing field → null.\n\n"
            f"Transcript:\n{transcript[:8000]}")
    
    # ── 4. Text-to-speech: edge-tts (Microsoft online service) ──
    async def _speak(text: str, out: str, voice: str) -> None:
        import edge_tts
        await edge_tts.Communicate(text, voice).save(out)
    
    def text_to_speech(text: str, out: str = "out.mp3", voice: str = "bg-BG-KalinaNeural") -> str:
        """Bulgarian voices: bg-BG-KalinaNeural, bg-BG-BorislavNeural.
        The text goes to the cloud — do not send personal data."""
        asyncio.run(_speak(text, out, voice))
        return out
    
    # ── The whole chain: voice note → structured case ──
    def process_voice_note(audio_path: str, worker_id: str) -> dict:
        t = transcribe(audio_path, model_size="turbo")
        case = structure_transcript(t["text"], CaseNote)
        return {"worker_id": worker_id, "duration_s": t["duration"],
                "transcript": t["text"], "case": case.model_dump()}
    ⚠️
    WhisperX and Bulgarian
    WhisperX aligns words with a separate wav2vec2 model. Its default list has no Bulgarian — without model_name, load_align_model stops with an error. Find a wav2vec2 model trained on Bulgarian on Hugging Face, check its licence and pass it as align_model. ⚠️ We have not tested a specific model — so we do not recommend a name.
    🔊
    Text-to-speech in Bulgarian
    The old lesson used XTTS-v2 with language="bg". XTTS-v2 supports 17 languages and Bulgarian is not one of them. edge-tts has two Bulgarian voices (Kalina, Borislav — checked with edge-tts --list-voices; our test produced Bulgarian audio), but it is an unofficial client for an online service 🌐: the text leaves your machine. ⚠️ We have not verified a good fully local Bulgarian voice.
    example · field voice note → structure (fictional data, translated)
    Transcript:
    "I visited the client from case 214. She lives alone, the neighbours have not
     seen her for several days, she has trouble moving around. She needs a check-up
     by her GP and we should find out if she has relatives. Next visit — Thursday."
    
    CaseNote:
      problem:           "Lives alone, limited mobility, risk of isolation"
      urgency:           "medium"
      actions_mentioned: ["GP check-up", "check for relatives"]
      next_visit:        "Thursday"
    ⛔
    Recordings from social and health work
    They contain personal and often health data. Work with a case number, not a name and address; keep the transcript and the model local; the output is a draft for the social worker, not a decision.
  6. Documents: the right tool for each kind

    The first question for any PDF: does it have a text layer? If it does (a PDF "printed" from Word), the text is read directly with no recognition errors. If it is scanned, it is a picture and needs OCR, which always brings errors.

    ToolStrengthLimitationWhen
    pdfplumberExact text and tables from PDFs with a text layerNo OCRInvoices and reports generated by software
    UnstructuredMany formats; hi_res detects layout and tables; OCR with Tesseracthi_res is slow; multi-column pages may come out in the wrong orderScanned documents; general entry point
    DoclingLayout and complex tables; export to MarkdownHeavier (layout models)Financial statements, technical documents
    MarkItDownWord, Excel, PowerPoint, PDF → MarkdownNo OCR by defaultOffice files into an LLM
    TesseractOCR, bul language packText only, no layoutUnder the hood of Unstructured
    Python · document_intelligence.py
    from pathlib import Path
    import pdfplumber
    from langchain_ollama import ChatOllama
    
    llm = ChatOllama(model="llama3.1:8b", temperature=0)
    
    # ── 1. Unstructured: any format, OCR with Tesseract ──
    def parse_document(file_path: str, scanned: bool = False,
                       languages: tuple[str, ...] = ("bul", "eng")) -> dict:
        from unstructured.partition.auto import partition
        elements = partition(
            filename=file_path,
            strategy="hi_res" if scanned else "auto",    # hi_res: layout + tables, slower
            languages=list(languages),                    # Tesseract codes: bul, eng
            infer_table_structure=True,
        )
        out = {"titles": [], "tables": [], "text_blocks": []}
        for el in elements:
            if el.category == "Title":
                out["titles"].append(el.text)
            elif el.category == "Table":
                out["tables"].append({"text": el.text, "html": el.metadata.text_as_html})
            elif el.category == "NarrativeText":
                out["text_blocks"].append(el.text)
        out["full_text"] = "\n\n".join(e.text for e in elements if e.text.strip())
        return out
    
    # ── 2. pdfplumber: tables from PDFs with a text layer ──
    def extract_tables_pdfplumber(pdf_path: str) -> list[dict]:
        tables = []
        with pdfplumber.open(pdf_path) as pdf:
            for n, page in enumerate(pdf.pages, start=1):
                for t in page.extract_tables():
                    if t:
                        tables.append({"page": n, "rows": len(t), "cols": max(len(r) for r in t), "data": t})
        return tables
    
    # ── 3. Docling: complex tables and layout ──
    def parse_with_docling(file_path: str) -> dict:
        from docling.document_converter import DocumentConverter
        doc = DocumentConverter().convert(file_path).document
        return {"markdown": doc.export_to_markdown(),
                "tables": [t.export_to_dataframe(doc=doc) for t in doc.tables],
                "num_pages": len(doc.pages)}
    
    # ── 4. MarkItDown: Office → Markdown ──
    def office_to_markdown(file_path: str) -> str:
        from markitdown import MarkItDown
        return MarkItDown().convert(file_path).markdown
    
    # ── 5. Router: picks the tool by file type ──
    def has_text_layer(pdf_path: str, min_chars: int = 50) -> bool:
        with pdfplumber.open(pdf_path) as pdf:
            return any(len((p.extract_text() or "").strip()) >= min_chars for p in pdf.pages[:3])
    
    def smart_parse(file_path: str) -> dict:
        suffix = Path(file_path).suffix.lower()
        if suffix in {".docx", ".pptx", ".xlsx"}:
            return {"type": suffix[1:], "markdown": office_to_markdown(file_path)}
        if suffix == ".pdf":
            if has_text_layer(file_path):
                with pdfplumber.open(file_path) as pdf:
                    text = "\n".join(p.extract_text() or "" for p in pdf.pages)
                return {"type": "native_pdf", "text": text, "tables": extract_tables_pdfplumber(file_path)}
            return {"type": "scanned_pdf", **parse_document(file_path, scanned=True)}
        return {"type": "other", **parse_document(file_path)}
    
    # ── 6. Question to a small document, no vector database ──
    def document_qa(file_path: str, question: str, max_chars: int = 8000) -> str:
        d = smart_parse(file_path)
        text = (d.get("text") or d.get("markdown") or d.get("full_text", ""))[:max_chars]
        return llm.invoke(f"DOCUMENT:\n{text}\n\nQUESTION: {question}\n"
                          "Answer ONLY from the document. If the answer is not there, say so.").content
    ✅
    We checked the router
    On 01.10.2026 we ran smart_parse on a sample PDF with a table, a Word file and an Excel file: the PDF was recognised as having a text layer and the table was extracted row by row; Word and Excel come out as Markdown. The old router looked only at the first page — here we look at up to three, so a PDF with a blank cover is not mistaken for a scan.
    💼
    Illustrative case: accounting with mixed invoices
    Invoices arrive as software-generated PDFs, scans and Excel files. smart_parse detects the type and picks the tool, and Invoice from step 4 structures the data and checks the totals in code. The old lesson's figures ("500 invoices", "98.7% detection") have no source and are removed. A document longer than a few pages goes through RAG (Block 4), not whole into the prompt.

04Check

Checklist

Quiz

1. When is the smaller vision model (e.g. 7b instead of 32b) the sensible choice?

2. You need to translate a Bulgarian recording into English with Whisper. Which model does not work?

3. smart_parse detects a PDF with a text layer. What does that mean?

4. WhisperX stops with "No default align-model for language: bg". What do you do?

05What's next

06Sources

  1. Visual Instruction Tuning (Liu et al., 2023) — the LLaVA paper.
  2. Ollama: qwen2.5vl, qwen3-vl, gemma3, gemma4, llava — tags and sizes.
  3. Ollama: vision — how to send images.
  4. Ollama: structured outputs — JSON schema, including for vision models.
  5. LangChain: ChatOllama — images and structured output.
  6. OpenAI Whisper — models, memory, speed, turbo and translation; the paper (appendix D — accuracy by language).
  7. WhisperX — word-level alignment and diarization.
  8. pyannote speaker-diarization-community-1 🌐 — the "who speaks" model; requires accepted terms.
  9. XTTS-v2 — the list of 17 supported languages.
  10. edge-tts 🌐 — text-to-speech; --list-voices.
  11. Unstructured: partitioning — strategies and OCR languages.
  12. Docling — documentation (the new address).
  13. MarkItDown — Office and PDF to Markdown.
  14. pdfplumber — text and tables from PDF.
  15. Tesseract tessdata — language packs, including bul.