FLUX.1 LoRA on GX10: Your Own Style and Licence
A LoRA is a small "add-on layer" that teaches a large image model your face, an object or a style from a few dozen photos. Here we train one on FLUX.1 [dev] with SimpleTuner on a local AI server of the NVIDIA GB10 class — and we say the most important thing first: such a LoRA is a derivative of [dev] and is not for commercial work.
config.json, a separate data file and the command simpletuner train. We also removed claims we cannot confirm: "18 minutes" for 1000 steps, "16–24 GB of video memory", "adapter 100–200 MB", the comparison table with full fine-tuning, and the Florence-2 captioning with trust_remote_code. We removed absolute paths in a user folder. We added: the licence as the first step (a LoRA derived from [dev] is not for paid work), a reminder that photos of people are personal data, installation for Blackwell/CUDA 13 per the documentation, the data file, checkpoint comparison and a quiz.
01What you'll learn
- What a LoRA is and why it works only with the base model it was trained on.
- Why the licence of the base carries over to the LoRA and what follows from that.
- How to prepare a small photo set and describe it to SimpleTuner.
- Which settings matter and how to start the training.
- How to compare checkpoints and recognise training that ran too long.
- How to load the finished LoRA with diffusers and try its strength.
02Before you start
| Base | Licence (per the page) | A LoRA on it — for commercial work? |
|---|---|---|
| FLUX.1 [dev] | FLUX.1 [dev] Non-Commercial License | No. The licence defines any modified or fine-tuned version as a "Derivative" |
| FLUX.1 Krea [dev] | The same non-commercial licence | No. This is also SimpleTuner's default flavour |
| FLUX.1 [schnell] | Apache-2.0 | The licence allows it, but SimpleTuner says direct training on schnell does not give good results yet; and merging schnell with dev makes the dev licence apply |
| Another base under Apache-2.0 (e.g. FLUX.2 [klein] 4B) | Apache-2.0 | ⚠️ We did not check whether SimpleTuner supports it |
That is why this lesson is for learning and personal experiments. Do not put a LoRA trained on [dev] into a paid project. For commercial work start from a base whose licence allows it, and write down which one with a date.
- A machine of the NVIDIA GB10 class with Linux and access to its terminal.
- Python 3.10 to 3.13 (per the SimpleTuner documentation).
- A Hugging Face account and the terms accepted on the FLUX.1-dev page (the model is behind a consent gate).
- A lot of system memory: per the documentation, just reducing the model's precision at startup needs about 50 GB. Your machine has 128 GB shared, but we have not tried it.
- Free disk space for the model, the caches and the checkpoints — check with
df -h. - Photos that are yours or that you have written consent to use. Faces are personal data — do not train on other people without consent.
03Steps
-
Set up the folders and the photos
Why data first? According to the documentation, FLUX absorbs the flaws in the photos first (artefacts, blur) and only then the subject itself — so quality weighs more than quantity. The documentation gives no exact number; start with a small set of 10–20 varied photos (that is our choice). No watermarks, no group shots.
bash · on the machinemkdir -p ~/flux-lora/datasets/my-subject ~/flux-lora/config ~/flux-lora/output # put the photos in ~/flux-lora/datasets/my-subject/✅Size mattersIn the data example belowminimum_image_sizeis 1024 — smaller photos are rejected. If you get "no images detected", check the size or raiserepeats. -
Install SimpleTuner
Per the documentation, there is a separate variant of the package for Blackwell GPUs and CUDA 13. Work in a separate virtual environment so you do not mix packages with the system.
bash · not runcd ~/flux-lora python3 -m venv venv source venv/bin/activate pip install 'simpletuner[cuda13]' --extra-index-url https://download.pytorch.org/whl/cu130⚠️ We did not check whether ready-made packages exist for Arm (aarch64). If the installation fails, see the installation section of the SimpleTuner documentation.
-
Log in to Hugging Face
The FLUX.1 [dev] weights are downloaded after logging in with an account and accepting the terms on the model page. The command will ask for your access token — enter it there and never write it into a file from this lesson.
bashhuggingface-cli loginIn newer versions of the
huggingface_hubpackage the same command ishf auth login. -
Describe the data
SimpleTuner reads a separate data file. The entry below follows the example in the documentation: one set with a manual subject name (
instanceprompt) and an entry for the text cache. What is a "subject name"? A rare word you later use to call your subject in a prompt; here it is invented —ohwx person.json · config/multidatabackend.json[ { "id": "my-subject", "type": "local", "crop": false, "resolution": 1024, "minimum_image_size": 1024, "maximum_image_size": 1024, "target_downsample_size": 1024, "resolution_type": "pixel_area", "cache_dir_vae": "cache/vae/flux/my-subject", "instance_data_dir": "datasets/my-subject", "caption_strategy": "instanceprompt", "instance_prompt": "ohwx person", "metadata_backend": "discovery", "repeats": 10 }, { "id": "text-embeds", "type": "local", "dataset_type": "text_embeds", "default": true, "cache_dir": "cache/text/flux", "write_batch_size": 128 } ]The value
repeats: 10is our choice for a small set — the documentation shows a larger value. The paths are relative to the folder you start the training from. -
The training settings
The easiest way is to run
simpletuner configure— it asks step by step. If you prefer to do it by hand, here are the keys forconfig/config.json. The names were checked against the documentation of 03.10.2026; the full list is there.json · config/config.json{ "model_type": "lora", "model_family": "flux", "model_flavour": "dev", "pretrained_model_name_or_path": "black-forest-labs/FLUX.1-dev", "output_dir": "output/my-subject", "data_backend_config": "config/multidatabackend.json", "train_batch_size": 1, "gradient_accumulation_steps": 1, "learning_rate": 1e-4, "max_train_steps": 1000, "lora_rank": 16, "optimizer": "adamw_bf16", "mixed_precision": "bf16", "gradient_checkpointing": true, "checkpointing_steps": 250, "validation_prompt": "a photo of ohwx person on a beach", "validation_resolution": "1024x1024", "validation_num_inference_steps": 20 }💡What is from the documentation and what is our choiceFrom the documentation:model_flavourdefaults tokrea(also non-commercial), which is why it isdevhere; keeptrain_batch_sizeat 1;adamw_bf16andbf16for beginners;gradient_checkpointingon in almost every case; a smallerlora_rankgives a smaller file (try 1, 4, 16). Our choice, not from the documentation:learning_rate1e-4 (the documentation only says that 1e-3 can "roast" the model while 1e-5 does almost nothing),max_train_steps1000 andcheckpointing_steps250. -
Start and watch
bash · not runsimpletuner train nvidia-smiFirst the descriptions and the images are written to the cache in a form the model reads, then training begins. We have not measured the time. If the process is killed abruptly after the text models are loaded, the documentation links this to too little system memory when reducing precision. On GB10 the memory is shared —
nvidia-smimay show no counter for it, and that is normal. -
Compare the checkpoints
Why? If training runs too long, the model starts to "memorise" the photos — every image resembles the training ones or square grid artefacts appear (per the documentation). Look at the validation images in the
outputfolder at 250, 500, 750 and 1000 steps and choose the earlier checkpoint if the later ones are worse. -
Load the LoRA with diffusers
The general way per the diffusers documentation: load the base model, then the LoRA weights with
load_lora_weightsand an adapter name, and the strength withset_adapters. ⚠️ The code was not run. Where exactly inoutputthe weights file is, check on disk.python · not runimport torch from diffusers import FluxPipeline pipe = FluxPipeline.from_pretrained( "black-forest-labs/FLUX.1-dev", torch_dtype=torch.bfloat16 ).to("cuda") pipe.load_lora_weights( "output/my-subject/checkpoint-1000", # check the real path weight_name="pytorch_lora_weights.safetensors", # check the name adapter_name="mine", ) pipe.set_adapters(["mine"], adapter_weights=[0.85]) image = pipe( prompt="ohwx person hiking in the mountains, golden hour", num_inference_steps=28, guidance_scale=3.5, height=1024, width=1024, ).images[0] image.save("test.png")Try strengths 0.5, 0.85 and 1.0 with the same prompt. The values
28and3.5are our example. Per the SimpleTuner documentation, use aguidance_scalearound the value the LoRA was trained with — check it in your own settings. -
Before you share anything
- On [dev] — for learning and personal use only; not for a client, not for sale.
- Photos of people — only with consent. A finished LoRA carries a likeness of the face; treat it as personal data.
- The Hugging Face token is not in a file and not in Git.
04Check
- You have noted the licence of the base with a date and know that a LoRA on [dev] is not for paid work.
- The photos are yours or you have written consent.
config.jsonandmultidatabackend.jsonare written without secrets and with relative paths.- Training has started and
outputcontains validation images. - You compared at least two checkpoints and chose the one without artefacts.
Quiz
1. Can a LoRA trained on FLUX.1 [dev] be used in a paid client project?
2. What is instance_prompt in the data file?
3. Why is the quality of the photos more important than their number with FLUX?
4. Every image starts to look like the training photos. What do you do?
05What's next
06Sources
- SimpleTuner: Flux quickstart 🔒 local — installation for CUDA 13, keys in
config.json, the data file,simpletuner train, licence of merged models, learning-rate guidance. - Hugging Face diffusers: Load adapters 🌐 global —
load_lora_weights,adapter_name,set_adapters. - FLUX.1-dev — model card and access terms.
- FLUX.1 [dev] licence text — the definitions of "Derivative" and "Non-Commercial Purpose".
- NVIDIA DGX Spark: hardware — shared memory of the GB10 class.