It's Friday evening. Priya, a machine language (ML) engineer at a financial services firm, queues up a fraud-detection fine-tuning job on an GPU cluster costing $55 an hour. The run should take about 40 hours—roughly $2,200 in compute. She double-checks the hyperparameters, submits the job, and heads home for the weekend.Monday morning, she opens her laptop. The model had stopped learning sometime Friday night, but the job kept running—burning through 2 full days of GPU time on a training run that was going nowhere. That's over $1,500 in wasted compute, and she has to start over.If this so