At GPU scale, failures are routine. See how smart dataloading and checkpointing keep training fast between failures and cheap to recover after them.