← All notes

A running machine can be snapshotted now

Virtainer 1.0.0 takes snapshots from a VM that keeps running. The first release that allowed it refused for a good reason, the shape of the product is what made the difference, and the first real-hardware validation did not work at all. What changed, what did not, and what we got wrong on the way.

Snapshots can now be taken from a running machine. The VM keeps serving traffic while it happens, and you choose how consistent the capture has to be. This landed in 1.0.0.

It is worth saying why it was refused for so long, because the refusal was not timidity. An earlier release had a hot-snapshot path, and restoring from one scattered a guest filesystem. The conclusion written down at the time was blunt: a disk set that is internally consistent is not the same as a filesystem that is consistent, and the only capture path that could be validated end to end was one taken while nothing was writing to the disks. So snapshots required a fully stopped machine, and that was that.

What turned out to be wrong was the assumption that the constraint was about the goal. It was about the copying mechanism available then.

The copy, not the machine

Two things had to be true before a running capture could be honest.

The first is a copy primitive that agrees on a moment and then holds it. A snapshot is not “read every disk as fast as possible”, because the guest is writing while you read. The primitive needed is one where writes arriving during the copy do not enter the copy, and where the blocks not yet copied are held aside so the result is the disk as it was at a chosen instant. The long part, moving the data, then happens after the guest is already writing again.

The second is that a machine here owns several disks: a root disk and any number of data volumes. If those disks were served by separate processes, agreeing on one instant across them would be a coordination problem between them, and coordination problems of that kind are exactly where a claim of consistency goes to die quietly.

They are not served by separate processes. Each machine’s disks sit behind a single storage daemon, one per machine, which is a topology decision that was made for other reasons entirely. So the shared instant is not negotiated between processes at all. It is one transaction on one endpoint.

That is the whole reason this is possible now and was not a year ago. Nothing about the risk changed. One structural decision happened to make the honest version reachable.

Consistency is chosen, never assumed

A snapshot now carries one of three labels, and the label is what the product actually proved, not what it hoped for.

Stopped is the old path, still there and still the simplest: nothing is writing, so nothing has to be arranged.

Online, crash-consistent is captured while the machine runs. It is what the disks would look like after a power cut. Most guests come back from that and run a filesystem check, which is a fine outcome as long as you know it is coming.

Online, quiesced asks the guest, through its agent, to pause writes for the instant of the capture and resume after. This needs a guest that supports it. If the agent cannot be reached, the request fails. It does not quietly fall back to crash-consistent and keep the label, because a snapshot whose label overstates what it is would be worse than no snapshot, and it is the specific failure this feature spent a long time being refused for.

Snapshots are bounded by free space in the pool they live in, not by a count per machine.

What did not change

A snapshot still never captures memory or running state. Restoring one leaves the machine stopped, and its next boot is an ordinary cold boot. PCI device assignments do not travel with it.

And a snapshot is still not a backup. It shares the storage pool with the machine it protects, so if the pool or the host is gone, both are gone together. It exists to undo a change you are about to make. Copies that leave the host are a separate feature, and remain a stopped-machine feature.

The part that did not work

The first build with this feature passed its unit tests and its storage-daemon integration tests. On real hardware, online capture did not run at all. Not partially: it never once succeeded.

Six separate defects were involved, and every one of them needed a real machine to show up. A capture directory the per-machine daemon account could not traverse. A guest agent reply shaped differently from what the code expected, which meant the quiesced path had never actually reported success, anywhere. Snapshot artifacts left owned by the machine that had just been backed up. A missing node reported as the wrong problem. A daemon dying without saying why. A reconcile warning that fired when nothing was wrong.

The lesson is not that the tests were bad. It is that a test environment which agrees with your assumptions will confirm them, and that a feature whose central promise is “this is provably consistent” has to be proven on the machines people actually run.

One reliability change shipped alongside it, in the same spirit. A machine whose hypervisor stops answering is no longer treated as dead if its process is still alive. That distinction exists because a busy host build starved a machine of CPU long enough to miss its health checks, and five healthy VMs were destroyed in nine seconds. Their own logs showed nothing wrong with them; they had simply not answered in time. Now the machine keeps running, an event is recorded, and the judgement is left where it belongs, with the operator. Destruction is irreversible, and an unresponsive machine that a person should look at is a survivable problem.

The guide covers the controls, the labels and the limits in Snapshots.

KEEP READING
After Apple, Docker agrees that AI agents belong in their own machineAug 16Run DeepSeek Harness as an AppVM, step by stepAug 15