Markdown · All docs as text

Training and fine-tuning#

Long jobs are where spot machines save the most money, and where losing a machine hurts the most. Two things make a long job safe to run on RunCompute: a checkpoint your code can resume from, and a budget.

Save and resume checkpoints#

RunCompute does not see inside your process. When a machine is lost, the job is started again on another machine with the same command. Your code decides what to resume from.

Write checkpoints under /outputs/checkpoints/ and load the newest one at start:

Python
import glob, os, torch

CKPT_DIR = "/outputs/checkpoints"
os.makedirs(CKPT_DIR, exist_ok=True)

start_step = 0
existing = sorted(glob.glob(f"{CKPT_DIR}/step-*.pt"))
if existing:
    state = torch.load(existing[-1])
    model.load_state_dict(state["model"])
    optimizer.load_state_dict(state["optimizer"])
    start_step = state["step"] + 1

for step in range(start_step, total_steps):
    train_step()
    if step % 500 == 0:
        torch.save({"model": model.state_dict(), "optimizer": optimizer.state_dict(), "step": step},
                   f"{CKPT_DIR}/step-{step:07d}.pt")

Hugging Face Trainer and PyTorch Lightning have their own options for saving and resuming; point their output directory at /outputs.

Match the checkpoint interval#

checkpoint_every tells the matcher how often you save, in hours. It is used to estimate how much work a preemption would throw away:

Pick a budget#

A rough budget is est_cost × 1.3 from a quote. If the job reaches the budget it stops with status stopped, keeps its checkpoints, and can be submitted again with a higher budget.

Python
offers = client.quote(min_vram=40, work_units=12, checkpoint_every=0.5)
budget = round(offers[0]["est_cost"] * 1.3, 2)