W&B on GX10: Tracking ML Experiments
When you train or evaluate a model, after a dozen attempts you no longer remember which settings gave which result. Weights & Biases records every attempt (run) with its settings and metrics so you can compare them. Here we connect it to a script on a GX10, work without a network too, and start an automatic search for settings.
config was not defined and wandb.agent.run(...) is not part of the interface — it is now wandb.agent(...). We removed what we could not confirm: Slack alerts, a model registry, a LLaMA Factory integration and the automatic "data flywheel" — they were not checked, and NVIDIA's blueprint for it is deprecated. We added: where the data goes (cloud versus offline), how to keep the key safe and why a sweep and its runs must share a project.
wandb behaves on Arm64 or how offline work is synchronised.01What you'll learn
- What a run, a project and a sweep are in W&B.
- How to create a key and log in without leaving it in code.
- How to record settings and metrics from your own script.
- What data leaves the machine and how to work offline.
- How to start an automatic search for settings (a sweep).
02Before you start
- A machine of the NVIDIA GB10 class (for example the ASUS Ascent GX10 or the DGX Spark) or any computer with Python 3 and
pip. - A Weights & Biases account and an API key. According to the documentation as of 03.10.2026 the account is created through the Forge platform (docs.wandb.ai now redirects to CoreWeave's documentation) — the names may change.
- A training or evaluation script to track. If you have none, the example below simulates training.
wandb/local) is mentioned in the environment-variable documentation; we do not cover it in this lesson.03Steps
-
Create an API key
Log in to the dashboard, open your profile → User settings → API keys → New key, give it a name and copy the key at once. Why at once: the full key is shown only once; afterwards you see only its beginning and if you lose it you create a new one.
✅The key is a secretDo not write it in the script, in Git or on a page. Keep it in an environment variable or a password manager. If it leaks — delete it in the dashboard and create a new one. -
Install and log in
A virtual environment keeps the project's packages separate from the system.
bashpython3 -m venv .venv . .venv/bin/activate pip install wandb wandb login # asks for the keyIn automated environments the key is passed with
export WANDB_API_KEY=<key>before the script (the variable is described in the documentation). ⚠️ We did not check the installation on Arm64. -
The first run
wandb.init()starts a run: you give a project and a dictionary of settings. Everything you record withrun.log()becomes a chart in the dashboard. The example is from the W&B quickstart and simulates training:python · first_run.pyimport wandb import random wandb.login() project = "my-awesome-project" config = {"epochs": 10, "lr": 0.01} with wandb.init(project=project, config=config) as run: offset = random.random() / 5 for epoch in range(2, config["epochs"]): acc = 1 - 2**-config["epochs"] - random.random() / config["epochs"] - offset loss = 2**-config["epochs"] + random.random() / config["epochs"] + offset run.log({"accuracy": acc, "loss": loss})Run it a few times and the dashboard will show several runs with random names. In your own script put the same
run.log({...})where you get the loss at each step. -
Compare runs
In the dashboard open the project: the "Runs" column lists them and the charts overlay the metrics. The point of the settings in
configis here — you can see whichlrgave a lower loss. The documentation also has a parallel-coordinates chart and a parameter-importance analysis (⚠️ not tried by us). -
Working without sending: offline mode
If the script must not reach the internet or the data must not leave, set the variable before running:
bashexport WANDB_MODE=offline # keeps the metadata locally, no sync python first_run.py # WANDB_MODE=disabled turns W&B off completelyHow offline runs are uploaded later is described in the W&B documentation; we did not try it, so we give no command.
-
Sweep: automatic search for settings
A sweep tries different values of the settings and records a run for each. You describe the search space, register the sweep and start an "agent" that calls your function. The example is from the W&B tutorial:
python · sweep.pyimport wandb def objective(config): score = config.x**3 + config.y return score def main(): with wandb.init(project="my-first-sweep") as run: score = objective(run.config) run.log({"score": score}) sweep_configuration = { "method": "random", "metric": {"goal": "minimize", "name": "score"}, "parameters": { "x": {"max": 0.1, "min": 0.01}, "y": {"values": [1, 3, 7]}, }, } sweep_id = wandb.sweep(sweep=sweep_configuration, project="my-first-sweep") wandb.agent(sweep_id, function=main, count=10)⚠️One project for the sweep and its runsThe project name inwandb.init()must match the one inwandb.sweep(). If you usemultiprocessing, wrapwandb.sweep()andwandb.agent()inif __name__ == "__main__":. The agent is stopped withCtrl+C(a second time to end it).The method here is
random. Other methods and options are in the "Define sweep configuration" section of the documentation (not opened by us). -
What else there is
The quickstart also points to: tracking models and datasets with Artifacts, sharing in the Registry and reports. For applications with language models the same vendor has W&B Weave. None of these is run in this lesson.
04Check
- The key is created and is not in code, Git or a page.
wandb loginworks and the first run shows in the dashboard with at least two metrics.- You know what data must not be logged to the cloud service.
- A sweep of several runs has finished and the runs are compared.
Quiz
1. Which function starts a new run?
2. How do you make the script not send data to the cloud?
3. Where is the right place for the API key?
4. What applies to the project in a sweep?
05What's next
06Sources
- Weights & Biases: quickstart 🌐 global — key, installation, first run; checked 03.10.2026.
- W&B: sweeps tutorial — configuration,
wandb.sweep,wandb.agent; checked 03.10.2026. - W&B: environment variables —
WANDB_API_KEY,WANDB_MODE,WANDB_BASE_URL. - NVIDIA: AI Observability for Data Flywheel — page of the deprecated blueprint; repository.