The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewUPDATED 2026-10-03
IDENTITY
module
GX10-04-22 · FLUX.1 LoRA with SimpleTuner
series
GX10 (local AI server class: NVIDIA GB10, e.g. ASUS Ascent GX10 / DGX Spark)
level
Advanced
duration
2 h plus training time (not measured)
prerequisites
A GB10-class machine with Linux, Python 3.10–3.13, a Hugging Face account with the FLUX.1 [dev] terms accepted, enough system RAM for model quantisation (about 50 GB per the SimpleTuner quickstart), a small set of photos you own or have written consent to use
trust_label
UPDATED 2026-10-03 (SimpleTuner FLUX quickstart and the diffusers adapter-loading docs read on that date; licences as read for lesson 04-23 on 2026-10-03). NOT TESTED: nothing was run on a GB10 machine. Not labelled VERIFIED
licence
FLUX.1 [dev] is under the FLUX.1 [dev] Non-Commercial License. A LoRA trained on it is a derivative and is not for commercial or client work
language
human view: en · bulgarian edition: /academy/gx10/ (same file name)
previous / next
GX10 series index / 04-23_Comfy_UI_Images.html
PURPOSE

Train a LoRA adapter for FLUX.1 [dev] on a small personal dataset with SimpleTuner, compare checkpoints, load the adapter with diffusers, and understand why the result stays non-commercial. Show where a commercial path would have to start (a base model with a permissive licence), without claiming that SimpleTuner supports one.

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

04-23 · ComfyUI on GX10: images and licences (04-23_Comfy_UI_Images.html) · related: 04-81 FLUX.2 [klein] · series index: kagami.bg/en/academy/gx10/ · offer: Quick experiment (kagami.bg/stalbata/)

SOURCES
TAGS
gx10nvidia-gb10fluxlorasimpletunerdiffuserslicensingimage-generation
UPDATED · 03.10.2026

FLUX.1 LoRA on GX10: Your Own Style and Licence

A LoRA is a small "add-on layer" that teaches a large image model your face, an object or a style from a few dozen photos. Here we train one on FLUX.1 [dev] with SimpleTuner on a local AI server of the NVIDIA GB10 class — and we say the most important thing first: such a LoRA is a derivative of [dev] and is not for commercial work.

⏱ 2 h + training time Advanced GX10 FLUX.1 [dev] · SimpleTuner · diffusers NVIDIA GB10 · 128 GB unified memory
SimpleTuner (training)🔒 local diffusers (generation)🔒 local FLUX.1 [dev] weights from Hugging Face🌐 global
🔄
UPDATED · 03.10.2026 — what changed
The lesson was rebuilt against the current documentation. We removed the TOML settings and the parameters that do not exist in SimpleTuner — it is driven by config.json, a separate data file and the command simpletuner train. We also removed claims we cannot confirm: "18 minutes" for 1000 steps, "16–24 GB of video memory", "adapter 100–200 MB", the comparison table with full fine-tuning, and the Florence-2 captioning with trust_remote_code. We removed absolute paths in a user folder. We added: the licence as the first step (a LoRA derived from [dev] is not for paid work), a reminder that photos of people are personal data, installation for Blackwell/CUDA 13 per the documentation, the data file, checkpoint comparison and a quiz.
⚠️
What we have not run ourselves
We had no GB10-class machine: none of the commands in this lesson were run, which is why there is no "TESTED" label. We did not check whether SimpleTuner's CUDA 13 packages have Arm (aarch64) builds, how long training takes, or exactly what the weights file in a checkpoint is called. The diffusers code for FLUX was not checked either — the diffusers documentation we read shows only the general way to load a LoRA. Which is why there is no "VERIFIED" label.

01What you'll learn

02Before you start

⚖️
Licence first, weights second
As read on the model pages on 03.10.2026 (see also lessons 04-23 and 04-81). Licences change — check again before every project.
BaseLicence (per the page)A LoRA on it — for commercial work?
FLUX.1 [dev]FLUX.1 [dev] Non-Commercial LicenseNo. The licence defines any modified or fine-tuned version as a "Derivative"
FLUX.1 Krea [dev]The same non-commercial licenceNo. This is also SimpleTuner's default flavour
FLUX.1 [schnell]Apache-2.0The licence allows it, but SimpleTuner says direct training on schnell does not give good results yet; and merging schnell with dev makes the dev licence apply
Another base under Apache-2.0 (e.g. FLUX.2 [klein] 4B)Apache-2.0⚠️ We did not check whether SimpleTuner supports it

That is why this lesson is for learning and personal experiments. Do not put a LoRA trained on [dev] into a paid project. For commercial work start from a base whose licence allows it, and write down which one with a date.

03Steps

  1. Set up the folders and the photos

    Why data first? According to the documentation, FLUX absorbs the flaws in the photos first (artefacts, blur) and only then the subject itself — so quality weighs more than quantity. The documentation gives no exact number; start with a small set of 10–20 varied photos (that is our choice). No watermarks, no group shots.

    bash · on the machine
    mkdir -p ~/flux-lora/datasets/my-subject ~/flux-lora/config ~/flux-lora/output
    # put the photos in ~/flux-lora/datasets/my-subject/
    ✅
    Size matters
    In the data example below minimum_image_size is 1024 — smaller photos are rejected. If you get "no images detected", check the size or raise repeats.
  2. Install SimpleTuner

    Per the documentation, there is a separate variant of the package for Blackwell GPUs and CUDA 13. Work in a separate virtual environment so you do not mix packages with the system.

    bash · not run
    cd ~/flux-lora
    python3 -m venv venv
    source venv/bin/activate
    pip install 'simpletuner[cuda13]' --extra-index-url https://download.pytorch.org/whl/cu130

    ⚠️ We did not check whether ready-made packages exist for Arm (aarch64). If the installation fails, see the installation section of the SimpleTuner documentation.

  3. Log in to Hugging Face

    The FLUX.1 [dev] weights are downloaded after logging in with an account and accepting the terms on the model page. The command will ask for your access token — enter it there and never write it into a file from this lesson.

    bash
    huggingface-cli login

    In newer versions of the huggingface_hub package the same command is hf auth login.

  4. Describe the data

    SimpleTuner reads a separate data file. The entry below follows the example in the documentation: one set with a manual subject name (instanceprompt) and an entry for the text cache. What is a "subject name"? A rare word you later use to call your subject in a prompt; here it is invented — ohwx person.

    json · config/multidatabackend.json
    [
      {
        "id": "my-subject",
        "type": "local",
        "crop": false,
        "resolution": 1024,
        "minimum_image_size": 1024,
        "maximum_image_size": 1024,
        "target_downsample_size": 1024,
        "resolution_type": "pixel_area",
        "cache_dir_vae": "cache/vae/flux/my-subject",
        "instance_data_dir": "datasets/my-subject",
        "caption_strategy": "instanceprompt",
        "instance_prompt": "ohwx person",
        "metadata_backend": "discovery",
        "repeats": 10
      },
      {
        "id": "text-embeds",
        "type": "local",
        "dataset_type": "text_embeds",
        "default": true,
        "cache_dir": "cache/text/flux",
        "write_batch_size": 128
      }
    ]

    The value repeats: 10 is our choice for a small set — the documentation shows a larger value. The paths are relative to the folder you start the training from.

  5. The training settings

    The easiest way is to run simpletuner configure — it asks step by step. If you prefer to do it by hand, here are the keys for config/config.json. The names were checked against the documentation of 03.10.2026; the full list is there.

    json · config/config.json
    {
      "model_type": "lora",
      "model_family": "flux",
      "model_flavour": "dev",
      "pretrained_model_name_or_path": "black-forest-labs/FLUX.1-dev",
      "output_dir": "output/my-subject",
      "data_backend_config": "config/multidatabackend.json",
      "train_batch_size": 1,
      "gradient_accumulation_steps": 1,
      "learning_rate": 1e-4,
      "max_train_steps": 1000,
      "lora_rank": 16,
      "optimizer": "adamw_bf16",
      "mixed_precision": "bf16",
      "gradient_checkpointing": true,
      "checkpointing_steps": 250,
      "validation_prompt": "a photo of ohwx person on a beach",
      "validation_resolution": "1024x1024",
      "validation_num_inference_steps": 20
    }
    💡
    What is from the documentation and what is our choice
    From the documentation: model_flavour defaults to krea (also non-commercial), which is why it is dev here; keep train_batch_size at 1; adamw_bf16 and bf16 for beginners; gradient_checkpointing on in almost every case; a smaller lora_rank gives a smaller file (try 1, 4, 16). Our choice, not from the documentation: learning_rate 1e-4 (the documentation only says that 1e-3 can "roast" the model while 1e-5 does almost nothing), max_train_steps 1000 and checkpointing_steps 250.
  6. Start and watch

    bash · not run
    simpletuner train
    nvidia-smi

    First the descriptions and the images are written to the cache in a form the model reads, then training begins. We have not measured the time. If the process is killed abruptly after the text models are loaded, the documentation links this to too little system memory when reducing precision. On GB10 the memory is shared — nvidia-smi may show no counter for it, and that is normal.

  7. Compare the checkpoints

    Why? If training runs too long, the model starts to "memorise" the photos — every image resembles the training ones or square grid artefacts appear (per the documentation). Look at the validation images in the output folder at 250, 500, 750 and 1000 steps and choose the earlier checkpoint if the later ones are worse.

  8. Load the LoRA with diffusers

    The general way per the diffusers documentation: load the base model, then the LoRA weights with load_lora_weights and an adapter name, and the strength with set_adapters. ⚠️ The code was not run. Where exactly in output the weights file is, check on disk.

    python · not run
    import torch
    from diffusers import FluxPipeline
    
    pipe = FluxPipeline.from_pretrained(
        "black-forest-labs/FLUX.1-dev", torch_dtype=torch.bfloat16
    ).to("cuda")
    
    pipe.load_lora_weights(
        "output/my-subject/checkpoint-1000",        # check the real path
        weight_name="pytorch_lora_weights.safetensors",  # check the name
        adapter_name="mine",
    )
    pipe.set_adapters(["mine"], adapter_weights=[0.85])
    
    image = pipe(
        prompt="ohwx person hiking in the mountains, golden hour",
        num_inference_steps=28,
        guidance_scale=3.5,
        height=1024,
        width=1024,
    ).images[0]
    image.save("test.png")

    Try strengths 0.5, 0.85 and 1.0 with the same prompt. The values 28 and 3.5 are our example. Per the SimpleTuner documentation, use a guidance_scale around the value the LoRA was trained with — check it in your own settings.

  9. Before you share anything

    • On [dev] — for learning and personal use only; not for a client, not for sale.
    • Photos of people — only with consent. A finished LoRA carries a likeness of the face; treat it as personal data.
    • The Hugging Face token is not in a file and not in Git.

04Check

Quiz

1. Can a LoRA trained on FLUX.1 [dev] be used in a paid client project?

2. What is instance_prompt in the data file?

3. Why is the quality of the photos more important than their number with FLUX?

4. Every image starts to look like the training photos. What do you do?

05What's next

06Sources

  1. SimpleTuner: Flux quickstart 🔒 local — installation for CUDA 13, keys in config.json, the data file, simpletuner train, licence of merged models, learning-rate guidance.
  2. Hugging Face diffusers: Load adapters 🌐 global — load_lora_weights, adapter_name, set_adapters.
  3. FLUX.1-dev — model card and access terms.
  4. FLUX.1 [dev] licence text — the definitions of "Derivative" and "Non-Commercial Purpose".
  5. NVIDIA DGX Spark: hardware — shared memory of the GB10 class.