model rollouts need a rollback path
Four model replicas occupy four GPUs. The deployment wants to start a replacement before stopping an old pod. There's no fifth GPU. The release can wait forever unless something about that plan changes.
This is an easy rollout to configure and an impossible one to schedule. With one whole device assigned to each replica, the availability settings are also a hardware reservation. The new image can be valid, every old pod can be healthy, and there can still be nowhere for the candidate to run.
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
These settings keep availability at the desired replica count while allowing an extra pod during replacement. That first new pod needs a schedulable GPU slot. Allowing one unavailable replica can release a slot, but it also leaves the remaining replicas carrying more traffic. Their measured spare throughput needs to cover the difference. A small canary pool on reserved capacity can give the candidate somewhere to run before you begin replacing the main fleet.
The slot has to be the right kind, too. A new model may require a different GPU architecture, more memory or a different MIG profile. Put that in resource requests and scheduling constraints, and try it on a cluster with the same node mix. A successful image build doesn't exercise any of those requirements.
getting a pod ready to serve
Once a pod schedules, it may spend minutes pulling weights and warming kernels. A liveness probe that treats this as a failure can restart the process repeatedly before initialization ever finishes. Use a startup probe to allow for loading. Readiness should establish that the model can handle the expected request class; an open HTTP port isn't enough. Liveness should detect a stuck process once startup is out of the way.
I would record time spent pending, pulling the image, loading weights, warming up and reaching readiness separately. A single deployment duration hides which part needs attention. It also gives you very little evidence when a later release takes twice as long to become available.
Even a ready pod tells you only part of what you need to know. The service can return valid responses while producing worse answers or taking longer on the documents people actually send. That requires a comparison with the previous release, and the comparison needs a precise definition of what changed.
what does “the model version” include?
Usually, more than the weights. The serving image, tokenizer, runtime flags, prompt template, tool schema, safety policy and retrieval configuration can all change the result. If deployment changes several of them, rolling back an alias such as model-v2 won't necessarily restore the old request path.
I'd have the router select an immutable release manifest that records the whole set. Pin the serving image by digest and the weights by content hash or immutable registry version. Include the tokenizer, context limit, sampling defaults and stop sequences. Staging and production should resolve the same artifacts so the evaluation can be tied to what gets deployed.
Signatures and provenance help establish where those artifacts came from. They don't answer whether the service meets its quality and latency targets. Keep evaluation as a separate gate with versioned cases from the application's work: long contexts, multilingual requests, tool use, refusals and hostile retrieved documents where relevant. Put the quality results beside p95 and p99 latency, GPU memory, output-token counts and errors. A good result on short prompts can coexist with a serious regression on long ones.
For a live canary, decide the traffic share and observation period before starting. Keep tenants on one version where conversation state or prefix caching requires it. Compare the candidate with the baseline by traffic class, and attach the manifest ID to every trace. If long-context requests are failing, an overall success rate dominated by short chat requests can hide the problem.
The stop conditions also belong in the plan. Sustained first-token latency, tool-call validation failures or unexpected refusals in an important workflow are all possible reasons to stop. Agree on them before looking at results; otherwise it's too easy to keep finding reasons to let a disappointing canary run a little longer.
Shadow requests can help compare output and latency without showing candidate responses to users. For an agent, the shadow environment must use read-only or simulated tools so it doesn't repeat a live write. Check whether private input is allowed to reach that environment as well. “It's only a shadow” doesn't change where the data is being sent.
moving traffic back
The manifest needs to be usable by the rollback tooling, not merely readable in a release ticket. That means image digest, weight and tokenizer hashes, prompt revision, retrieval-index revision, runtime flags and policy revision in a machine-readable record. API and conversation-state compatibility need recording too. If the candidate has written state the previous tokenizer, retriever or tool schema can't handle, healthy old pods won't make the return path work.
Keeping the previous release warm buys you time to recover. Keeping its image in a registry doesn't do the same thing if weight loading and GPU warm-up take ten minutes. Its old prompt and retrieval configuration need to remain available as well. Conversations begun on the candidate may be able to move back, or may need to stay pinned until they finish. That decision should already be made when the rollback starts.
The practical test is to run the return path under load. Send a small share, say five percent, to the candidate, force a failure and switch back. Measure recovery time and inspect queued requests, retries and stream reconnections. Look for duplicate work and requests that cross versions halfway through. Save the manifest, metrics and traces with the result.
If that exercise stalls because a retrieval revision was removed or the old model no longer fits the available hardware, the release isn't ready for a larger traffic share. Fix it and repeat the exercise. Being able to restore the previous request path is a claim you can test before an incident forces you to rely on it.
technical references: kubernetes deployments ↗, probe behavior ↗, image digests ↗, and slsa provenance ↗.