From training to Ollama: a fine-tuned model in use
In Part 1 you prepared the data and the LoRA setup. Here you walk the whole way to a working model: training that stops on time by itself, honest evaluation in Bulgarian, merging and compressing to GGUF, and a local Ollama endpoint that replaces the old model in your RAG with no other change.
tokenizer= → processing_class=, max_seq_length → max_length, the removed warmup_ratio → warmup_steps=0.03, torch_dtype → dtype. Trap found: by default ROUGE returns 0 for any Bulgarian text (the tokenizer keeps only Latin characters) — fixed and tested. The table of "improvements" (perplexity −57%, accuracy +23 pp, ROUGE-L +74%) could not be verified — removed; a template for your own numbers takes its place. The evaluation questions now use facts checked against the National Social Security Institute (NSSI), the VAT Act and the Obligations and Contracts Act (the old "sick leave up to 6 months, art. 45 of the Social Insurance Code" was wrong). llama.cpp moved to ggml-org and builds with CMake. The Gemma chat template was added to the Modelfile — without it Ollama often produces nonsense. GGUF sizes come from real files; "~96% of the quality" was removed. Added: a regression check with lm-eval and publishing to the Hugging Face Hub. Internal project and machine names were replaced with generic ones.
01What you'll learn
- How to configure
SFTTrainerso it stops on overfitting by itself and keeps the best checkpoint. - How to evaluate the model honestly — base vs fine-tuned, on data it has not seen — and which metrics lie on Cyrillic.
- How to check that the model has not "forgotten" its general skills.
- How to merge the LoRA adapter, turn it into GGUF and choose a quantization.
- How to run it in Ollama with the right chat template and plug it into an existing LangChain pipeline.
02Before you start
- You have done Block 5 · Part 1 (LoRA/QLoRA): you have
data/train.jsonlanddata/eval.jsonlin chat format. Here we load them from disk with atextfield: in Part 1, afterds = ds.map(to_text, batched=True), addds.save_to_disk("./data/prepared"). - An NVIDIA GPU with bf16 support or a machine with 128 GB of unified memory (the NVIDIA DGX Spark class). A 27B model in 4 bits with a 4096 context needs a few dozen GB.
- Python 3.10–3.13 and a virtual environment. On ARM machines of the DGX Spark class Unsloth recommends its Docker image.
- Ollama 🔒 local, installed and running.
- You have read the base model's licence. Gemma-based BgGPT is used under the Gemma terms.
python3 -m venv .venv && source .venv/bin/activate
pip install -U unsloth trl transformers peft datasets evaluate rouge_score langchain-ollama wandb
# checked with: trl 1.14.1 · transformers 5.18.0 · peft 0.21.1 · unsloth 2026.9.12 · evaluate 0.4.6INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0 (FastLanguageModel, Gemma 2 template). Part 1 now trains on BgGPT 3 (Gemma 3; 4B, 12B, 27B, with image understanding and up to 131k context), released in March 2026. If you trained your adapter there, change three things: loading — with FastModel as in Part 1; MODEL_NAME/BASE — to the Part 1 model; TEMPLATE in the Modelfile — to the gemma3 template from the Ollama library. An adapter trained on one base model does not fit another. ⚠️ This lesson has not been run with BgGPT 3.03Steps
-
The map: where we are in the pipeline
The cycle has six stages. The first two are from Part 1; here we do the other four. Why see it whole? Because a mistake at an early stage (for example a different chat template during training) only shows up at the very end, in Ollama.
# Stage Output Where 1 Load the model + LoRA setup model with adapter Part 1 2 Chat-format data + quality check train/evalPart 1 3 Training adapter/here 4 Evaluation numbers: base vs fine-tuned here 5 Merge + GGUF + quantization *.ggufhere 6 Ollama + integration local endpoint here -
Training with
SFTTrainerThe trainer evaluates every 50 steps, saves every 100 and stops by itself if the loss on
evaldoes not improve for three checks in a row. At the end it loads the best checkpoint, not the last one. Why? Small datasets overfit fast: the last checkpoint is often worse than an earlier one.Python · train_sft.pyimport os import torch from unsloth import FastLanguageModel # BEFORE trl/transformers so the optimisations apply from trl import SFTTrainer, SFTConfig from transformers import TrainerCallback, EarlyStoppingCallback from datasets import load_from_disk MODEL_NAME = "INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0" RUN_NAME = "bggpt-hr-v1" OUTPUT_DIR = f"./runs/{RUN_NAME}" DATA_DIR = "./data/prepared" # from Part 1: "train" and "eval", field "text" MAX_LEN = 4096 os.environ.setdefault("WANDB_MODE", "offline") # logs stay on the machine; remove for cloud W&B os.environ.setdefault("WANDB_PROJECT", "bg-finetune") # ── Model + LoRA (as in Part 1) ── model, tokenizer = FastLanguageModel.from_pretrained( model_name=MODEL_NAME, max_seq_length=MAX_LEN, dtype=None, load_in_4bit=True, ) model = FastLanguageModel.get_peft_model( model, r=16, lora_alpha=32, lora_dropout=0.05, bias="none", target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"], use_gradient_checkpointing="unsloth", # so we do NOT set gradient_checkpointing in SFTConfig use_rslora=True, random_state=42, ) ds = load_from_disk(DATA_DIR) train_ds, eval_ds = ds["train"], ds["eval"] print(f"Train: {len(train_ds)} | Eval: {len(eval_ds)}") # ── Memory monitoring ── class MemoryCallback(TrainerCallback): def on_log(self, args, state, control, logs=None, **kwargs): if torch.cuda.is_available(): used = torch.cuda.memory_allocated() / 1e9 peak = torch.cuda.max_memory_allocated() / 1e9 print(f"step {state.global_step}: memory {used:.1f} GB (peak {peak:.1f} GB)") trainer = SFTTrainer( model=model, processing_class=tokenizer, # in TRL 1.x — not tokenizer= train_dataset=train_ds, eval_dataset=eval_ds, callbacks=[MemoryCallback(), EarlyStoppingCallback(early_stopping_patience=3)], args=SFTConfig( output_dir=OUTPUT_DIR, run_name=RUN_NAME, dataset_text_field="text", max_length=MAX_LEN, # in TRL 1.x — not max_seq_length num_train_epochs=3, per_device_train_batch_size=4, gradient_accumulation_steps=4, # effective batch = 16 learning_rate=2e-4, lr_scheduler_type="cosine", warmup_steps=0.03, # transformers 5: a fraction < 1 = share of steps (warmup_ratio is gone) bf16=True, optim="adamw_8bit", logging_steps=10, eval_strategy="steps", eval_steps=50, save_strategy="steps", save_steps=100, # a multiple of eval_steps — otherwise early stopping fails save_total_limit=3, load_best_model_at_end=True, metric_for_best_model="eval_loss", report_to="wandb", # or "none" seed=42, ), ) stats = trainer.train() print(f"Loss: {stats.training_loss:.4f} | Steps: {stats.global_step}") model.save_pretrained(f"{OUTPUT_DIR}/adapter") # only the LoRA adapter — tens/hundreds of MB tokenizer.save_pretrained(f"{OUTPUT_DIR}/adapter") print(f"Adapter saved to {OUTPUT_DIR}/adapter")⚠️Old examples onlineUnsloth's notebooks still writeSFTTrainer(tokenizer=…, max_seq_length=…)— Unsloth accepts them because it patches the trainer. With plain TRL 1.x these names raise an error.warmup_rationo longer exists in transformers 5.⛔W&B is a cloud serviceWithreport_to="wandb"and withoutWANDB_MODE=offline, metrics and configuration go to the cloud 🌐. If the data is sensitive, keep offline mode or setreport_to="none". -
Evaluation: base vs fine-tuned, on unseen data
A single number without a comparison says nothing. Run the same checks on the base model and on the fine-tuned one and look at the difference. Four levels, from cheapest to most expensive:
Check What it shows Weakness Perplexity How "confidently" the model reads texts from your domain. Lower is better. Says nothing about whether the answer is correct. ROUGE How many words overlap with the reference answer. Penalises a correct answer in other words. Needs its own tokenizer on Cyrillic. Keywords Whether the answer contains the required facts (70%, 18 months…). A rough filter; add synonyms. LLM judge How many claims have no cited source. The judge makes mistakes too. Check a sample by hand. The questions and prompts stay in Bulgarian — the model is Bulgarian and so is the domain.
Python · evaluation.pyimport json, math, re import torch import evaluate from transformers import AutoModelForCausalLM, AutoTokenizer from peft import PeftModel from langchain_ollama import ChatOllama BASE = "INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0" def load_model(base_model: str, adapter_path: str | None = None): tok = AutoTokenizer.from_pretrained(base_model) model = AutoModelForCausalLM.from_pretrained(base_model, dtype=torch.bfloat16, device_map="auto") if adapter_path: # no adapter = the base model for comparison model = PeftModel.from_pretrained(model, adapter_path) return model.eval(), tok # ── 1. Perplexity ── def perplexity(model, tok, texts: list[str]) -> float: total_loss, total_tokens = 0.0, 0 with torch.no_grad(): for text in texts: enc = tok(text, return_tensors="pt", max_length=512, truncation=True).to(model.device) n = enc["input_ids"].shape[1] out = model(**enc, labels=enc["input_ids"]) total_loss += out.loss.item() * n total_tokens += n return math.exp(total_loss / total_tokens) # ── 2. ROUGE: a tokenizer that understands Cyrillic ── rouge = evaluate.load("rouge") def rouge_bg(predictions: list[str], references: list[str]) -> dict: r = rouge.compute(predictions=predictions, references=references, rouge_types=["rouge1", "rougeL"], tokenizer=lambda s: re.findall(r"\w+", s.lower())) return {k: round(float(v), 4) for k, v in r.items()} # ── 3. Answer through the model's chat template ── def answer(model, tok, question: str) -> str: prompt = tok.apply_chat_template([{"role": "user", "content": question}], add_generation_prompt=True, tokenize=False) inputs = tok(prompt, return_tensors="pt", add_special_tokens=False).to(model.device) with torch.no_grad(): out = model.generate(**inputs, max_new_tokens=200, do_sample=False, pad_token_id=tok.eos_token_id) return tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True) # ── 4. Keywords (verified facts; "|" = synonym) ── QUESTIONS = [ {"q": "Кой плаща първите дни от болничния и колко?", "kw": ["работодател", "70"]}, {"q": "Колко е обезщетението от НОИ при общо заболяване?", "kw": ["80", "18 месеца|18 календарни"]}, {"q": "Каква е ставката на ДДС при износ на стоки извън ЕС?", "kw": ["нулева|0 %|0%", "износ"]}, {"q": "Какъв е общият давностен срок по ЗЗД?", "kw": ["5 години|пет години", "110"]}, ] def keyword_bench(model, tok) -> dict: details, correct = [], 0 for item in QUESTIONS: a = answer(model, tok, item["q"]).lower() hits = [kw for kw in item["kw"] if any(alt in a for alt in kw.lower().split("|"))] passed = len(hits) == len(item["kw"]) correct += passed details.append({"q": item["q"], "passed": passed, "hits": hits, "answer": a[:120]}) return {"accuracy": correct / len(QUESTIONS), "details": details} # ── 5. LLM judge: claims without a cited law/article ── JUDGE_PROMPT = """Прочети отговора на AI асистент по трудово и осигурително право. Изброй твърденията, за които НЕ е посочен конкретен закон или член. Върни само JSON: {{"uncited": ["твърдение", ...], "score": число от 0 до 1}} (1 = всички твърдения имат източник) Отговор: {answer}""" judge = ChatOllama(model="llama3.1:70b", temperature=0, format="json") def cited_score(text: str) -> float | None: try: return float(json.loads(judge.invoke(JUDGE_PROMPT.format(answer=text)).content)["score"]) except (ValueError, KeyError, TypeError): return None # no invented 0.5 — count it as missing def full_eval(adapter_path: str | None, eval_texts: list[str]) -> dict: model, tok = load_model(BASE, adapter_path) bench = keyword_bench(model, tok) scores = [s for s in (cited_score(d["answer"]) for d in bench["details"]) if s is not None] return { "perplexity": round(perplexity(model, tok, eval_texts[:50]), 2), "keyword_accuracy": round(bench["accuracy"], 3), "cited_score": round(sum(scores) / len(scores), 3) if scores else None, } # base = full_eval(None, eval_texts) # tuned = full_eval("./runs/bggpt-hr-v1/adapter", eval_texts)⚠️Trap: ROUGE = 0 on BulgarianThe defaultrouge_scoretokenizer keeps only Latin letters and digits. Two identical Bulgarian texts score 0.0 — we tested it with evaluate 0.4.6. Withtokenizer=lambda s: re.findall(r"\w+", s.lower())the same pair scores correctly. If you see ROUGE for Bulgarian (or any Cyrillic language) without a custom tokenizer, the number is meaningless.Record the results in a table like this one. Fill it in with your own numbers — not someone else's:
Metric Base Fine-tuned Direction Perplexity (your domain) … … lower is better Keywords (share correct) … … higher Cited claims (0–1) … … higher ROUGE-L vs reference … … higher lm-eval (general skills, step 5) … … should not drop noticeably 💡The facts in the questionsIn Bulgaria the employer pays the first 2 working days of sick leave (70%); the NSSI pays 80% of the average daily gross pay over the preceding 18 calendar months (Social Insurance Code, arts. 40–41, checked with the NSSI). Exports outside the EU are zero-rated (VAT Act, art. 28). The general limitation period is 5 years (Obligations and Contracts Act, art. 110). Teaching examples, not legal advice. -
Merge the adapter and export GGUF
Ollama and llama.cpp work with a single file in GGUF format. So you first merge the LoRA adapter with the base model into a full 16-bit model, then convert and compress it. Unsloth does both in one line each.
Python · export_gguf.pyfrom unsloth import FastLanguageModel RUN = "./runs/bggpt-hr-v1" model, tokenizer = FastLanguageModel.from_pretrained( model_name=f"{RUN}/adapter", max_seq_length=4096, dtype=None, load_in_4bit=True, ) # 1. Merged 16-bit model — for lm-eval, vLLM and manual conversion model.save_pretrained_merged(f"{RUN}/merged", tokenizer, save_method="merged_16bit") # 2. GGUF directly (Unsloth calls llama.cpp behind the scenes) model.save_pretrained_gguf(f"{RUN}/gguf", tokenizer, quantization_method="q4_k_m")If the direct export gets stuck (common with a new architecture), do it by hand with llama.cpp. The repository is now
ggml-org/llama.cppand builds with CMake:bash · manual export with llama.cppgit clone https://github.com/ggml-org/llama.cpp pip install -r llama.cpp/requirements.txt cmake llama.cpp -B llama.cpp/build # add -DGGML_CUDA=ON if you will also run models with it cmake --build llama.cpp/build --config Release -j --target llama-quantize RUN=./runs/bggpt-hr-v1 mkdir -p $RUN/gguf python llama.cpp/convert_hf_to_gguf.py $RUN/merged \ --outfile $RUN/gguf/bggpt-hr-f16.gguf --outtype f16 llama.cpp/build/bin/llama-quantize \ $RUN/gguf/bggpt-hr-f16.gguf $RUN/gguf/bggpt-hr-q4_k_m.gguf Q4_K_M ls -lh $RUN/gguf/✅Rule for file namesUnsloth picks the GGUF file name itself. Check it withlsand put the exact name into the Modelfile in step 7. -
Has it forgotten its general skills?
Training on a narrow domain can "erase" parts of the general skills (catastrophic forgetting). Quick check: run the same standard tests from lm-evaluation-harness on the base model and on the merged one. The tests are in English — they measure general reasoning, not Bulgarian.
bash · regression with lm-evalpip install "lm_eval[hf]" # the fine-tuned (merged) model lm_eval --model hf \ --model_args pretrained=./runs/bggpt-hr-v1/merged,dtype=bfloat16 \ --tasks hellaswag,arc_easy --device cuda:0 --batch_size auto \ --output_path ./runs/bggpt-hr-v1/lm_eval # the base model — same tasks, for comparison lm_eval --model hf \ --model_args pretrained=INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0,dtype=bfloat16 \ --tasks hellaswag,arc_easy --device cuda:0 --batch_size auto \ --output_path ./runs/base/lm_eval💡How to read the resultA small drop is normal. A large drop means the learning rate is too high, too many epochs or data that is too narrow — go back to Part 1 and mix general examples into the set. -
Which quantization for what
Quantization lowers the precision of the weights so the model fits into less memory and runs faster. The sizes below are the real Gemma 2 27B GGUF files (the same architecture as BgGPT 27B). Runtime memory is the file + the context cache, which grows with
num_ctx.Format File (27B) When F16 ≈ 54.5 GB Only as an intermediate file and for evaluation. Q8_0 ≈ 28.9 GB Almost lossless. When memory allows. Q6_K ≈ 22.3 GB Very close to Q8_0, smaller. Q5_K_M ≈ 19.4 GB A good compromise with a bit more quality. Q4_K_M ★ ≈ 16.7 GB The most common choice for real use — also recommended by Unsloth. Q2_K ≈ 10.5 GB Noticeable loss. Not for serious work. ✅Don't trust "quality" percentagesHow much you lose with quantization depends on the model and the task. Run the keyword questions from step 3 on the GGUF variant in Ollama too — that is the real test. -
Register in Ollama with the right chat template
Ollama does not quantize GGUF on import — it takes the file as it is. The most common reason a fine-tuned model "talks nonsense" in Ollama is a wrong chat template or a missing stop token. So you set the Gemma template explicitly — the same one Ollama uses for the official
gemma2. Gemma 2 has no separate system role: the system text is placed before the question.bash · Modelfile + ollama create + testcat > Modelfile <<'EOF' FROM ./runs/bggpt-hr-v1/gguf/bggpt-hr-q4_k_m.gguf TEMPLATE """<start_of_turn>user {{ if .System }}{{ .System }} {{ end }}{{ .Prompt }}<end_of_turn> <start_of_turn>model {{ .Response }}<end_of_turn> """ PARAMETER stop "<start_of_turn>" PARAMETER stop "<end_of_turn>" PARAMETER temperature 0.1 PARAMETER num_ctx 4096 SYSTEM """Ти си асистент по трудово и осигурително право. Отговаряй само на български. Винаги посочвай закона и члена. Ако не знаеш — кажи го.""" EOF ollama create bggpt-hr:v1 -f Modelfile ollama show bggpt-hr:v1 --modelfile # check that the template and stop tokens are in ollama run bggpt-hr:v1 "Кой плаща първите дни от болничния?" # OpenAI-compatible API curl http://localhost:11434/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model": "bggpt-hr:v1", "messages": [{"role": "user", "content": "Колко е обезщетението от НОИ?"}], "temperature": 0.1, "max_tokens": 300}' | python3 -m json.tool⚠️The template must match trainingIf in Part 1 you formatted the data with a different template, put that one intoTEMPLATE. Endless generation or repetition usually means a wrong stop token. -
Swap in LangChain and run several models at once
If your RAG pipeline uses
ChatOllama, the swap is one name. Everything else stays.Python · langchain_swap.pyimport time from langchain_ollama import ChatOllama from langchain_core.prompts import ChatPromptTemplate from langchain_core.output_parsers import StrOutputParser # BEFORE: llm = ChatOllama(model="llama3.1:70b") llm = ChatOllama(model="bggpt-hr:v1", temperature=0.1, num_ctx=4096) chain = ( ChatPromptTemplate.from_messages([ ("system", "Отговаряй само от контекста. Посочи члена."), ("human", "Контекст:\n{context}\n\nВъпрос: {question}"), ]) | llm | StrOutputParser() ) print(chain.invoke({ "context": "КСО, чл. 40, ал. 5: първите два работни дни се изплащат от работодателя в размер 70%...", "question": "Кой плаща първите дни от болничния?", })) # ── Speed: fine-tuned vs base ── def bench(model_name: str, questions: list[str]) -> float: m = ChatOllama(model=model_name, temperature=0.1) m.invoke("загрявка") # the first call loads the model — don't time it t0 = time.perf_counter() for q in questions: m.invoke(q) return (time.perf_counter() - t0) / len(questions) QS = ["Кой плаща първите дни от болничния?", "Колко е обезщетението от НОИ?"] for name in ["bggpt-hr:v1", "llama3.1:70b"]: print(f"{name}: {bench(name, QS):.2f} s/question")🖥️128 GB of unified memory: three models at onceThe fine-tuned model in Q4_K_M (≈ 16.7 GB) +llama3.1:70b(≈ 43 GB) +nomic-embed-text(≈ 0.3 GB) come to ≈ 60 GB of weights; room is left for the context cache.OLLAMA_MAX_LOADED_MODELSdefaults to 3 per GPU;OLLAMA_KEEP_ALIVE=30mkeeps them loaded longer. The fine-tuned model for domain tasks, the general one for everything else. -
Optional: publish to the Hugging Face Hub
If you want to share the model 🌐 global, Unsloth uploads the GGUF directly.
bash + Python · publishhf auth login # the token lives only in the environment, never in code # in Python, after loading the model as in step 4: model.push_to_hub_gguf("<user>/bggpt-hr-gguf", tokenizer, quantization_method="q4_k_m")⛔Before you hit "publish"A model trained on personal data can "recite" it. Do not publish such a model. Respect the base model's licence (Gemma) and start with a private repository.
04Check
Checklist
- Training stops early or finishes; the best checkpoint by
eval_lossis loaded; the adapter is saved. - The evaluation set was not used in training.
- ROUGE on two identical Bulgarian texts gives 1.0, not 0.0.
- You have a "base vs fine-tuned" table with your own numbers.
- lm-eval on the merged model does not drop noticeably against the base.
- The GGUF file has the expected size.
ollama show --modelfileshows the Gemma template and stop tokens.- The model answers in Bulgarian without endless generation; the API on port 11434 responds.
- The LangChain chain works with only the model name changed.
Quiz
1. You run ROUGE on Bulgarian answers and every score is 0.0. What is the MOST LIKELY cause?
2. Why is Q4_K_M the most common choice for real use?
3. The model answers perfectly in Unsloth, but in Ollama it repeats itself and never stops. What do you check first?
4. Old code fails on warmup_ratio=0.03 with transformers 5. How is it written now?
05What's next
06Sources
- TRL: SFTTrainer —
processing_class,SFTConfig,max_length. - transformers: callbacks —
EarlyStoppingCallback,TrainerCallback. - Unsloth: saving to GGUF —
save_pretrained_merged,save_pretrained_gguf,push_to_hub_gguf, the chat-template trap. - Unsloth on NVIDIA DGX Spark — Docker image for 128 GB ARM machines.
- evaluate: ROUGE — the
tokenizerparameter. - lm-evaluation-harness — standard tests,
lm_eval[hf]. - llama.cpp: llama-quantize — conversion and quantization.
- Ollama: importing a model — GGUF and safetensors.
- Ollama: Modelfile —
TEMPLATE,PARAMETER,SYSTEM. - Ollama: FAQ —
OLLAMA_MAX_LOADED_MODELS,OLLAMA_KEEP_ALIVE. - Ollama: OpenAI compatibility —
/v1/chat/completions. - Ollama: gemma2 and llama3.1 — template and sizes.
- Gemma 2 27B GGUF — real quantization file sizes.
- BgGPT Gemma 2 27B and BgGPT 3.0 (Gemma 3) — INSAIT.
- LangChain: ChatOllama.
- Hugging Face CLI —
hf auth login. - W&B: environment variables 🌐 global —
WANDB_MODE=offline. - NSSI: temporary incapacity benefit (in Bulgarian) — the facts in the evaluation questions.