Passthrough that survives a host restart
What you get
A VM owns the device and uses its own driver. A NAS guest sees the real disks, with SMART, write caches and real errors, instead of a virtual disk.
- Devices are listed by vendor, model and serial number, not by slot. Moving a card to another slot is fine; a replaced disk is not handed over by mistake, and the VM that wanted it does not start and says why.
- Listed devices are held at boot, before the host's own drivers can attach. The host never imports a ZFS pool that belongs to the NAS guest.
- One device belongs to one VM. A second VM asking for it is refused.
- Devices with a BAR smaller than a page, such as the Intel C62x SATA controller, can be passed through. Stock bhyve refuses them.
How it survives a host restart
The IOMMU (Intel VT-d) is run by the keel hypervisor, which does not restart. While the host restarts, the device keeps its DMA mappings and finishes the I/O it has. Only its interrupts are masked; they wait in the device and are delivered when the VM resumes. The new host takes the device over without resetting it.
ESXi DirectPath and Hyper-V DDA cannot suspend a VM with a passed-through device.
What was measured
- NAS guest with a whole NVMe drive, writing with fsync: 4 host restarts, paused 21–23 seconds each, no errors.
- Heavy load, four writers and one reader at about 600–700 MB/s: 10 host restarts, no timeouts or resets, ZFS scrub with 0 errors.
- One VM with an NVMe drive and an onboard SATA controller: 5 host restarts, both fine.
- Onboard AHCI with a disk attached: data written in the guest read back with the same checksum, scrub with 0 errors.
- A device made to disappear while in use: the host stays up and the VM is marked as having lost it.
After 14 host restarts without an error, VMs with whole devices are suspended on a host restart by default instead of shut down.
Limits
- Interrupt remapping is not on yet. A guest driver could in principle send stray interrupts. With guests you trust, as at home or in a small office, this is acceptable for now.
- Devices that need an RMRR cannot be passed through yet.
- Snapshots of such a VM cover only the disks keelOS manages. Disks behind a passed-through controller are not in them, and a snapshot with memory is not possible: the device's own state cannot be saved.
- HBA cards have not been tested yet.