The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewVERIFIED 2026-10-01 · UPDATED 2026-10-01
IDENTITY
module
01-05b · From training to Ollama: a fine-tuned model in production
series
Blocks 0–10 · Block 5 — Fine-tuning · Part 2/2
level
Advanced
duration
4–6 h (plus training time)
prerequisites
01-05a (LoRA/QLoRA, dataset in chat format with "train"/"eval" splits) · Python 3.10–3.13 · NVIDIA GPU with bf16, or a 128 GB unified-memory machine · Ollama
trust_label
VERIFIED 2026-10-01 (parameter names against TRL 1.14.1 and transformers 5.18.0 source; Unsloth GGUF/merge API against official docs; llama.cpp build and quantize paths; Ollama import/Modelfile/FAQ/OpenAI docs; gemma2 template from the Ollama registry; GGUF sizes from Hugging Face; ROUGE-on-Cyrillic behaviour executed with evaluate 0.4.6; all source links) · UPDATED 2026-10-01 · NOT end-to-end tested: no 27B training run was executed for this revision
versions
trl 1.14.x · transformers 5.18.x · peft 0.21.x · unsloth 2026.9.x · evaluate 0.4.6 · lm_eval 0.4.13 · langchain-ollama 1.1.x · Ollama 0.35.x
language
this page: en · bulgarian edition: /academy/blokove/moduli/01-05b_Блок_5_Част_2_Training_Deployment.html
next
01-06_Блок_6_Operational_Systems.html · from AI tools to operating systems
PURPOSE

Take a LoRA fine-tune from a configured trainer to a running local endpoint: train with TRL SFTTrainer (early stopping, best checkpoint), evaluate honestly (perplexity, ROUGE with a Unicode tokenizer, keyword benchmark, LLM judge, general-skill regression with lm-evaluation-harness), merge the adapter, export GGUF, pick a quantization, register the model in Ollama with the correct chat template, and swap it into an existing LangChain pipeline.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

01-06_Блок_6_Operational_Systems.html · from AI tools to operating systems · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
fine-tuningloratrlunslothevaluationrougelm-evalggufquantizationllama.cppollamalangchainbulgarian
VERIFIED · 01.10.2026 UPDATED · 01.10.2026

From training to Ollama: a fine-tuned model in use

In Part 1 you prepared the data and the LoRA setup. Here you walk the whole way to a working model: training that stops on time by itself, honest evaluation in Bulgarian, merging and compressing to GGUF, and a local Ollama endpoint that replaces the old model in your RAG with no other change.

⏱ 4–6 h + training time Advanced Block 5 · Part 2/2 training · evaluation · GGUF · Ollama
Unsloth · TRL🔒 local llama.cpp🔒 local Ollama🔒 local Weights & Biases (optional)🌐 global Hugging Face Hub (optional)🌐 global
🔄
UPDATED · 01.10.2026 — what changed
The code was checked against the source of TRL 1.14 and transformers 5.18: tokenizer= → processing_class=, max_seq_length → max_length, the removed warmup_ratio → warmup_steps=0.03, torch_dtype → dtype. Trap found: by default ROUGE returns 0 for any Bulgarian text (the tokenizer keeps only Latin characters) — fixed and tested. The table of "improvements" (perplexity −57%, accuracy +23 pp, ROUGE-L +74%) could not be verified — removed; a template for your own numbers takes its place. The evaluation questions now use facts checked against the National Social Security Institute (NSSI), the VAT Act and the Obligations and Contracts Act (the old "sick leave up to 6 months, art. 45 of the Social Insurance Code" was wrong). llama.cpp moved to ggml-org and builds with CMake. The Gemma chat template was added to the Modelfile — without it Ollama often produces nonsense. GGUF sizes come from real files; "~96% of the quality" was removed. Added: a regression check with lm-eval and publishing to the Hugging Face Hub. Internal project and machine names were replaced with generic ones.

01What you'll learn

02Before you start

bash · install (versions as of 01.10.2026)
python3 -m venv .venv && source .venv/bin/activate
pip install -U unsloth trl transformers peft datasets evaluate rouge_score langchain-ollama wandb
# checked with: trl 1.14.1 · transformers 5.18.0 · peft 0.21.1 · unsloth 2026.9.12 · evaluate 0.4.6
💡
Which base model
The code in this lesson is for INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0 (FastLanguageModel, Gemma 2 template). Part 1 now trains on BgGPT 3 (Gemma 3; 4B, 12B, 27B, with image understanding and up to 131k context), released in March 2026. If you trained your adapter there, change three things: loading — with FastModel as in Part 1; MODEL_NAME/BASE — to the Part 1 model; TEMPLATE in the Modelfile — to the gemma3 template from the Ollama library. An adapter trained on one base model does not fit another. ⚠️ This lesson has not been run with BgGPT 3.

03Steps

  1. The map: where we are in the pipeline

    The cycle has six stages. The first two are from Part 1; here we do the other four. Why see it whole? Because a mistake at an early stage (for example a different chat template during training) only shows up at the very end, in Ollama.

    #StageOutputWhere
    1Load the model + LoRA setupmodel with adapterPart 1
    2Chat-format data + quality checktrain / evalPart 1
    3Trainingadapter/here
    4Evaluationnumbers: base vs fine-tunedhere
    5Merge + GGUF + quantization*.ggufhere
    6Ollama + integrationlocal endpointhere
  2. Training with SFTTrainer

    The trainer evaluates every 50 steps, saves every 100 and stops by itself if the loss on eval does not improve for three checks in a row. At the end it loads the best checkpoint, not the last one. Why? Small datasets overfit fast: the last checkpoint is often worse than an earlier one.

    Python · train_sft.py
    import os
    import torch
    from unsloth import FastLanguageModel              # BEFORE trl/transformers so the optimisations apply
    from trl import SFTTrainer, SFTConfig
    from transformers import TrainerCallback, EarlyStoppingCallback
    from datasets import load_from_disk
    
    MODEL_NAME = "INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0"
    RUN_NAME   = "bggpt-hr-v1"
    OUTPUT_DIR = f"./runs/{RUN_NAME}"
    DATA_DIR   = "./data/prepared"                     # from Part 1: "train" and "eval", field "text"
    MAX_LEN    = 4096
    
    os.environ.setdefault("WANDB_MODE", "offline")     # logs stay on the machine; remove for cloud W&B
    os.environ.setdefault("WANDB_PROJECT", "bg-finetune")
    
    # ── Model + LoRA (as in Part 1) ──
    model, tokenizer = FastLanguageModel.from_pretrained(
        model_name=MODEL_NAME, max_seq_length=MAX_LEN, dtype=None, load_in_4bit=True,
    )
    model = FastLanguageModel.get_peft_model(
        model, r=16, lora_alpha=32, lora_dropout=0.05, bias="none",
        target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
        use_gradient_checkpointing="unsloth",          # so we do NOT set gradient_checkpointing in SFTConfig
        use_rslora=True, random_state=42,
    )
    
    ds = load_from_disk(DATA_DIR)
    train_ds, eval_ds = ds["train"], ds["eval"]
    print(f"Train: {len(train_ds)} | Eval: {len(eval_ds)}")
    
    # ── Memory monitoring ──
    class MemoryCallback(TrainerCallback):
        def on_log(self, args, state, control, logs=None, **kwargs):
            if torch.cuda.is_available():
                used = torch.cuda.memory_allocated() / 1e9
                peak = torch.cuda.max_memory_allocated() / 1e9
                print(f"step {state.global_step}: memory {used:.1f} GB (peak {peak:.1f} GB)")
    
    trainer = SFTTrainer(
        model=model,
        processing_class=tokenizer,                    # in TRL 1.x — not tokenizer=
        train_dataset=train_ds, eval_dataset=eval_ds,
        callbacks=[MemoryCallback(), EarlyStoppingCallback(early_stopping_patience=3)],
        args=SFTConfig(
            output_dir=OUTPUT_DIR, run_name=RUN_NAME,
            dataset_text_field="text", max_length=MAX_LEN,   # in TRL 1.x — not max_seq_length
            num_train_epochs=3,
            per_device_train_batch_size=4,
            gradient_accumulation_steps=4,             # effective batch = 16
            learning_rate=2e-4, lr_scheduler_type="cosine",
            warmup_steps=0.03,                         # transformers 5: a fraction < 1 = share of steps (warmup_ratio is gone)
            bf16=True, optim="adamw_8bit",
            logging_steps=10,
            eval_strategy="steps", eval_steps=50,
            save_strategy="steps", save_steps=100,     # a multiple of eval_steps — otherwise early stopping fails
            save_total_limit=3,
            load_best_model_at_end=True, metric_for_best_model="eval_loss",
            report_to="wandb",                         # or "none"
            seed=42,
        ),
    )
    
    stats = trainer.train()
    print(f"Loss: {stats.training_loss:.4f} | Steps: {stats.global_step}")
    
    model.save_pretrained(f"{OUTPUT_DIR}/adapter")     # only the LoRA adapter — tens/hundreds of MB
    tokenizer.save_pretrained(f"{OUTPUT_DIR}/adapter")
    print(f"Adapter saved to {OUTPUT_DIR}/adapter")
    ⚠️
    Old examples online
    Unsloth's notebooks still write SFTTrainer(tokenizer=…, max_seq_length=…) — Unsloth accepts them because it patches the trainer. With plain TRL 1.x these names raise an error. warmup_ratio no longer exists in transformers 5.
    ⛔
    W&B is a cloud service
    With report_to="wandb" and without WANDB_MODE=offline, metrics and configuration go to the cloud 🌐. If the data is sensitive, keep offline mode or set report_to="none".
  3. Evaluation: base vs fine-tuned, on unseen data

    A single number without a comparison says nothing. Run the same checks on the base model and on the fine-tuned one and look at the difference. Four levels, from cheapest to most expensive:

    CheckWhat it showsWeakness
    PerplexityHow "confidently" the model reads texts from your domain. Lower is better.Says nothing about whether the answer is correct.
    ROUGEHow many words overlap with the reference answer.Penalises a correct answer in other words. Needs its own tokenizer on Cyrillic.
    KeywordsWhether the answer contains the required facts (70%, 18 months…).A rough filter; add synonyms.
    LLM judgeHow many claims have no cited source.The judge makes mistakes too. Check a sample by hand.

    The questions and prompts stay in Bulgarian — the model is Bulgarian and so is the domain.

    Python · evaluation.py
    import json, math, re
    import torch
    import evaluate
    from transformers import AutoModelForCausalLM, AutoTokenizer
    from peft import PeftModel
    from langchain_ollama import ChatOllama
    
    BASE = "INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0"
    
    def load_model(base_model: str, adapter_path: str | None = None):
        tok = AutoTokenizer.from_pretrained(base_model)
        model = AutoModelForCausalLM.from_pretrained(base_model, dtype=torch.bfloat16, device_map="auto")
        if adapter_path:                               # no adapter = the base model for comparison
            model = PeftModel.from_pretrained(model, adapter_path)
        return model.eval(), tok
    
    # ── 1. Perplexity ──
    def perplexity(model, tok, texts: list[str]) -> float:
        total_loss, total_tokens = 0.0, 0
        with torch.no_grad():
            for text in texts:
                enc = tok(text, return_tensors="pt", max_length=512, truncation=True).to(model.device)
                n = enc["input_ids"].shape[1]
                out = model(**enc, labels=enc["input_ids"])
                total_loss += out.loss.item() * n
                total_tokens += n
        return math.exp(total_loss / total_tokens)
    
    # ── 2. ROUGE: a tokenizer that understands Cyrillic ──
    rouge = evaluate.load("rouge")
    def rouge_bg(predictions: list[str], references: list[str]) -> dict:
        r = rouge.compute(predictions=predictions, references=references,
                          rouge_types=["rouge1", "rougeL"],
                          tokenizer=lambda s: re.findall(r"\w+", s.lower()))
        return {k: round(float(v), 4) for k, v in r.items()}
    
    # ── 3. Answer through the model's chat template ──
    def answer(model, tok, question: str) -> str:
        prompt = tok.apply_chat_template([{"role": "user", "content": question}],
                                         add_generation_prompt=True, tokenize=False)
        inputs = tok(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
        with torch.no_grad():
            out = model.generate(**inputs, max_new_tokens=200, do_sample=False,
                                 pad_token_id=tok.eos_token_id)
        return tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
    
    # ── 4. Keywords (verified facts; "|" = synonym) ──
    QUESTIONS = [
        {"q": "Кой плаща първите дни от болничния и колко?", "kw": ["работодател", "70"]},
        {"q": "Колко е обезщетението от НОИ при общо заболяване?", "kw": ["80", "18 месеца|18 календарни"]},
        {"q": "Каква е ставката на ДДС при износ на стоки извън ЕС?", "kw": ["нулева|0 %|0%", "износ"]},
        {"q": "Какъв е общият давностен срок по ЗЗД?", "kw": ["5 години|пет години", "110"]},
    ]
    
    def keyword_bench(model, tok) -> dict:
        details, correct = [], 0
        for item in QUESTIONS:
            a = answer(model, tok, item["q"]).lower()
            hits = [kw for kw in item["kw"] if any(alt in a for alt in kw.lower().split("|"))]
            passed = len(hits) == len(item["kw"])
            correct += passed
            details.append({"q": item["q"], "passed": passed, "hits": hits, "answer": a[:120]})
        return {"accuracy": correct / len(QUESTIONS), "details": details}
    
    # ── 5. LLM judge: claims without a cited law/article ──
    JUDGE_PROMPT = """Прочети отговора на AI асистент по трудово и осигурително право.
    Изброй твърденията, за които НЕ е посочен конкретен закон или член.
    Върни само JSON: {{"uncited": ["твърдение", ...], "score": число от 0 до 1}}
    (1 = всички твърдения имат източник)
    
    Отговор:
    {answer}"""
    
    judge = ChatOllama(model="llama3.1:70b", temperature=0, format="json")
    
    def cited_score(text: str) -> float | None:
        try:
            return float(json.loads(judge.invoke(JUDGE_PROMPT.format(answer=text)).content)["score"])
        except (ValueError, KeyError, TypeError):
            return None                                # no invented 0.5 — count it as missing
    
    def full_eval(adapter_path: str | None, eval_texts: list[str]) -> dict:
        model, tok = load_model(BASE, adapter_path)
        bench = keyword_bench(model, tok)
        scores = [s for s in (cited_score(d["answer"]) for d in bench["details"]) if s is not None]
        return {
            "perplexity": round(perplexity(model, tok, eval_texts[:50]), 2),
            "keyword_accuracy": round(bench["accuracy"], 3),
            "cited_score": round(sum(scores) / len(scores), 3) if scores else None,
        }
    
    # base  = full_eval(None, eval_texts)
    # tuned = full_eval("./runs/bggpt-hr-v1/adapter", eval_texts)
    ⚠️
    Trap: ROUGE = 0 on Bulgarian
    The default rouge_score tokenizer keeps only Latin letters and digits. Two identical Bulgarian texts score 0.0 — we tested it with evaluate 0.4.6. With tokenizer=lambda s: re.findall(r"\w+", s.lower()) the same pair scores correctly. If you see ROUGE for Bulgarian (or any Cyrillic language) without a custom tokenizer, the number is meaningless.

    Record the results in a table like this one. Fill it in with your own numbers — not someone else's:

    MetricBaseFine-tunedDirection
    Perplexity (your domain)……lower is better
    Keywords (share correct)……higher
    Cited claims (0–1)……higher
    ROUGE-L vs reference……higher
    lm-eval (general skills, step 5)……should not drop noticeably
    💡
    The facts in the questions
    In Bulgaria the employer pays the first 2 working days of sick leave (70%); the NSSI pays 80% of the average daily gross pay over the preceding 18 calendar months (Social Insurance Code, arts. 40–41, checked with the NSSI). Exports outside the EU are zero-rated (VAT Act, art. 28). The general limitation period is 5 years (Obligations and Contracts Act, art. 110). Teaching examples, not legal advice.
  4. Merge the adapter and export GGUF

    Ollama and llama.cpp work with a single file in GGUF format. So you first merge the LoRA adapter with the base model into a full 16-bit model, then convert and compress it. Unsloth does both in one line each.

    Python · export_gguf.py
    from unsloth import FastLanguageModel
    
    RUN = "./runs/bggpt-hr-v1"
    model, tokenizer = FastLanguageModel.from_pretrained(
        model_name=f"{RUN}/adapter", max_seq_length=4096, dtype=None, load_in_4bit=True,
    )
    
    # 1. Merged 16-bit model — for lm-eval, vLLM and manual conversion
    model.save_pretrained_merged(f"{RUN}/merged", tokenizer, save_method="merged_16bit")
    
    # 2. GGUF directly (Unsloth calls llama.cpp behind the scenes)
    model.save_pretrained_gguf(f"{RUN}/gguf", tokenizer, quantization_method="q4_k_m")

    If the direct export gets stuck (common with a new architecture), do it by hand with llama.cpp. The repository is now ggml-org/llama.cpp and builds with CMake:

    bash · manual export with llama.cpp
    git clone https://github.com/ggml-org/llama.cpp
    pip install -r llama.cpp/requirements.txt
    cmake llama.cpp -B llama.cpp/build                # add -DGGML_CUDA=ON if you will also run models with it
    cmake --build llama.cpp/build --config Release -j --target llama-quantize
    
    RUN=./runs/bggpt-hr-v1
    mkdir -p $RUN/gguf
    python llama.cpp/convert_hf_to_gguf.py $RUN/merged \
      --outfile $RUN/gguf/bggpt-hr-f16.gguf --outtype f16
    
    llama.cpp/build/bin/llama-quantize \
      $RUN/gguf/bggpt-hr-f16.gguf $RUN/gguf/bggpt-hr-q4_k_m.gguf Q4_K_M
    
    ls -lh $RUN/gguf/
    ✅
    Rule for file names
    Unsloth picks the GGUF file name itself. Check it with ls and put the exact name into the Modelfile in step 7.
  5. Has it forgotten its general skills?

    Training on a narrow domain can "erase" parts of the general skills (catastrophic forgetting). Quick check: run the same standard tests from lm-evaluation-harness on the base model and on the merged one. The tests are in English — they measure general reasoning, not Bulgarian.

    bash · regression with lm-eval
    pip install "lm_eval[hf]"
    
    # the fine-tuned (merged) model
    lm_eval --model hf \
      --model_args pretrained=./runs/bggpt-hr-v1/merged,dtype=bfloat16 \
      --tasks hellaswag,arc_easy --device cuda:0 --batch_size auto \
      --output_path ./runs/bggpt-hr-v1/lm_eval
    
    # the base model — same tasks, for comparison
    lm_eval --model hf \
      --model_args pretrained=INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0,dtype=bfloat16 \
      --tasks hellaswag,arc_easy --device cuda:0 --batch_size auto \
      --output_path ./runs/base/lm_eval
    💡
    How to read the result
    A small drop is normal. A large drop means the learning rate is too high, too many epochs or data that is too narrow — go back to Part 1 and mix general examples into the set.
  6. Which quantization for what

    Quantization lowers the precision of the weights so the model fits into less memory and runs faster. The sizes below are the real Gemma 2 27B GGUF files (the same architecture as BgGPT 27B). Runtime memory is the file + the context cache, which grows with num_ctx.

    FormatFile (27B)When
    F16≈ 54.5 GBOnly as an intermediate file and for evaluation.
    Q8_0≈ 28.9 GBAlmost lossless. When memory allows.
    Q6_K≈ 22.3 GBVery close to Q8_0, smaller.
    Q5_K_M≈ 19.4 GBA good compromise with a bit more quality.
    Q4_K_M ★≈ 16.7 GBThe most common choice for real use — also recommended by Unsloth.
    Q2_K≈ 10.5 GBNoticeable loss. Not for serious work.
    ✅
    Don't trust "quality" percentages
    How much you lose with quantization depends on the model and the task. Run the keyword questions from step 3 on the GGUF variant in Ollama too — that is the real test.
  7. Register in Ollama with the right chat template

    Ollama does not quantize GGUF on import — it takes the file as it is. The most common reason a fine-tuned model "talks nonsense" in Ollama is a wrong chat template or a missing stop token. So you set the Gemma template explicitly — the same one Ollama uses for the official gemma2. Gemma 2 has no separate system role: the system text is placed before the question.

    bash · Modelfile + ollama create + test
    cat > Modelfile <<'EOF'
    FROM ./runs/bggpt-hr-v1/gguf/bggpt-hr-q4_k_m.gguf
    
    TEMPLATE """<start_of_turn>user
    {{ if .System }}{{ .System }} {{ end }}{{ .Prompt }}<end_of_turn>
    <start_of_turn>model
    {{ .Response }}<end_of_turn>
    """
    PARAMETER stop "<start_of_turn>"
    PARAMETER stop "<end_of_turn>"
    PARAMETER temperature 0.1
    PARAMETER num_ctx 4096
    
    SYSTEM """Ти си асистент по трудово и осигурително право. Отговаряй само на български.
    Винаги посочвай закона и члена. Ако не знаеш — кажи го."""
    EOF
    
    ollama create bggpt-hr:v1 -f Modelfile
    ollama show bggpt-hr:v1 --modelfile              # check that the template and stop tokens are in
    ollama run bggpt-hr:v1 "Кой плаща първите дни от болничния?"
    
    # OpenAI-compatible API
    curl http://localhost:11434/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{"model": "bggpt-hr:v1",
           "messages": [{"role": "user", "content": "Колко е обезщетението от НОИ?"}],
           "temperature": 0.1, "max_tokens": 300}' | python3 -m json.tool
    ⚠️
    The template must match training
    If in Part 1 you formatted the data with a different template, put that one into TEMPLATE. Endless generation or repetition usually means a wrong stop token.
  8. Swap in LangChain and run several models at once

    If your RAG pipeline uses ChatOllama, the swap is one name. Everything else stays.

    Python · langchain_swap.py
    import time
    from langchain_ollama import ChatOllama
    from langchain_core.prompts import ChatPromptTemplate
    from langchain_core.output_parsers import StrOutputParser
    
    # BEFORE: llm = ChatOllama(model="llama3.1:70b")
    llm = ChatOllama(model="bggpt-hr:v1", temperature=0.1, num_ctx=4096)
    
    chain = (
        ChatPromptTemplate.from_messages([
            ("system", "Отговаряй само от контекста. Посочи члена."),
            ("human", "Контекст:\n{context}\n\nВъпрос: {question}"),
        ])
        | llm
        | StrOutputParser()
    )
    print(chain.invoke({
        "context": "КСО, чл. 40, ал. 5: първите два работни дни се изплащат от работодателя в размер 70%...",
        "question": "Кой плаща първите дни от болничния?",
    }))
    
    # ── Speed: fine-tuned vs base ──
    def bench(model_name: str, questions: list[str]) -> float:
        m = ChatOllama(model=model_name, temperature=0.1)
        m.invoke("загрявка")                            # the first call loads the model — don't time it
        t0 = time.perf_counter()
        for q in questions:
            m.invoke(q)
        return (time.perf_counter() - t0) / len(questions)
    
    QS = ["Кой плаща първите дни от болничния?", "Колко е обезщетението от НОИ?"]
    for name in ["bggpt-hr:v1", "llama3.1:70b"]:
        print(f"{name}: {bench(name, QS):.2f} s/question")
    🖥️
    128 GB of unified memory: three models at once
    The fine-tuned model in Q4_K_M (≈ 16.7 GB) + llama3.1:70b (≈ 43 GB) + nomic-embed-text (≈ 0.3 GB) come to ≈ 60 GB of weights; room is left for the context cache. OLLAMA_MAX_LOADED_MODELS defaults to 3 per GPU; OLLAMA_KEEP_ALIVE=30m keeps them loaded longer. The fine-tuned model for domain tasks, the general one for everything else.
  9. Optional: publish to the Hugging Face Hub

    If you want to share the model 🌐 global, Unsloth uploads the GGUF directly.

    bash + Python · publish
    hf auth login                                     # the token lives only in the environment, never in code
    
    # in Python, after loading the model as in step 4:
    model.push_to_hub_gguf("<user>/bggpt-hr-gguf", tokenizer, quantization_method="q4_k_m")
    ⛔
    Before you hit "publish"
    A model trained on personal data can "recite" it. Do not publish such a model. Respect the base model's licence (Gemma) and start with a private repository.

04Check

Checklist

Quiz

1. You run ROUGE on Bulgarian answers and every score is 0.0. What is the MOST LIKELY cause?

2. Why is Q4_K_M the most common choice for real use?

3. The model answers perfectly in Unsloth, but in Ollama it repeats itself and never stops. What do you check first?

4. Old code fails on warmup_ratio=0.03 with transformers 5. How is it written now?

05What's next

06Sources

  1. TRL: SFTTrainer — processing_class, SFTConfig, max_length.
  2. transformers: callbacks — EarlyStoppingCallback, TrainerCallback.
  3. Unsloth: saving to GGUF — save_pretrained_merged, save_pretrained_gguf, push_to_hub_gguf, the chat-template trap.
  4. Unsloth on NVIDIA DGX Spark — Docker image for 128 GB ARM machines.
  5. evaluate: ROUGE — the tokenizer parameter.
  6. lm-evaluation-harness — standard tests, lm_eval[hf].
  7. llama.cpp: llama-quantize — conversion and quantization.
  8. Ollama: importing a model — GGUF and safetensors.
  9. Ollama: Modelfile — TEMPLATE, PARAMETER, SYSTEM.
  10. Ollama: FAQ — OLLAMA_MAX_LOADED_MODELS, OLLAMA_KEEP_ALIVE.
  11. Ollama: OpenAI compatibility — /v1/chat/completions.
  12. Ollama: gemma2 and llama3.1 — template and sizes.
  13. Gemma 2 27B GGUF — real quantization file sizes.
  14. BgGPT Gemma 2 27B and BgGPT 3.0 (Gemma 3) — INSAIT.
  15. LangChain: ChatOllama.
  16. Hugging Face CLI — hf auth login.
  17. W&B: environment variables 🌐 global — WANDB_MODE=offline.
  18. NSSI: temporary incapacity benefit (in Bulgarian) — the facts in the evaluation questions.