[VFIO] add basic implementation - #5870
Closed
ShadowCurse wants to merge 28 commits into
Closed
ShadowCurse wants to merge 28 commits into
ShadowCurse wants to merge 28 commits into
Conversation
Codecov Report❌ Patch coverage is Please upload reports for the commit e97092d to get more accurate results. Additional details and impacted files@@ Coverage Diff @@
## main #5870 +/- ##
==========================================
- Coverage 82.83% 80.90% -1.93%
==========================================
Files 277 280 +3
Lines 30775 32010 +1235
==========================================
+ Hits 25491 25898 +407
- Misses 5284 6112 +828
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
ShadowCurse
force-pushed
the
vfio_with_dependencies
branch
11 times, most recently
from
May 14, 2026 16:29
f6d6fea to
50e789e
Compare
ShadowCurse
force-pushed
the
vfio_with_dependencies
branch
12 times, most recently
from
May 27, 2026 13:32
2f84f01 to
a21e87e
Compare
ShadowCurse
force-pushed
the
vfio_with_dependencies
branch
4 times, most recently
from
May 29, 2026 11:53
b2ea5ea to
528e62b
Compare
ShadowCurse
force-pushed
the
vfio_with_dependencies
branch
from
June 8, 2026 13:25
528e62b to
efa67e3
Compare
11 tasks
Add the VfioConfig and VfioConfigs types for describing VFIO device configuration. Wire them into VmResources and VmmConfig so that VFIO devices can be specified before boot. Actual device setup will be added in later commits. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add PUT /vfio/{id} API endpoint for configuring VFIO passthrough
devices.
Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
- add `vfio-bindings` and `vfio-ioctls` - make `arrayvec` non optional All of these will be used in the future VFIO commits. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Create an empty module where VFIO code will be. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
First thing to do with VFIO device is to scan it's capabilities and extended capabilities to find MSI-X cap and caps we want to filter out. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Next we need to gather information about BARs the device has and allocate space for them in the guest memory. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VFIO BAR regions containing MSI-X table/PBA will be split into mmappable and emulated parts. KVM memory slots require host-page alignment, but MSI-X structures can sit at arbitrary offsets within a BAR. Add additional helper functions for the calculations of these mappable/emulated BAR regions. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
With information about BARs and the capabilities, we can calculate areas of BARs we can safely DMA map into the guest. Everything outside those areas are subject to emulation on our end. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
After calculating the areas of BARs we can map to the guest, do this mapping. It involves `mmap`ing the device BAR into Firecracker virtual space first, then setting up the DMA mapping for this virtual address to the guest physical memory. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
After all the previous work with BARs and capabilities we can finally put everything together into a VfioDevice type. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VfioDevice will have to emulate both Bus accesses (for configuration space) and Pci accesses for the emulated BARs holes. This commit implements the handling for the Bus accesses. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Now implement the Pci emulation for the VfioDevice to handle emulated BAR accesses to the MSIx table/pba areas. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
In addition to the VfioDevice setting up DMA for the device BARs, we need to set up DMA for the whole guest RAM since the device will need to access it. These functions will be used in the next commits. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add functions for creation of KVM VFIO device and VFIO container. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add logic to the PciDevices to create new VFIO devices. As an additional step in VFIO device setup, guest RAM regions are mapped into the VFIO container's IOMMU so the device can DMA directly to guest memory. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Current VFIO implementation has some restrictions: - Does not work without PCI since VFIO devices are PCI devices - Does not work with virtio-mem device since we don't update DMA mappings on hot-plug/unplug - Does not work with virtio-balloon since it can `fadvise` on memory In order to prevent VMs being launched with invalid configurations, implement multiple checks for invalid configurations: - At API level, prevent adding of incompatible combinations (VFIO after balloon/mem or in reverse) - At vm creation or snapshot restoraton since they get VmResources from other sources. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VFIO device state is opaque to the VMM and cannot be serialized or restored. Add VFIO devices to the list of snapshot-incompatible devices so that snapshot requests are rejected with a clear error instead of producing a corrupt snapshot. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VFIO devices will use pread64/pwrite64 syscalls (from vfio-ioctls) to interact with BARs during runtime. Add them to the VPU thread syscall lists. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add devtool options for preparing a PCI device for VFIO passthrough testing. `--vfio-nvme-device` accepts a block device path (e.g. /dev/nvme1n1) or a PCI SBDF, resolves it to a PCI device, binds it to vfio-pci, and passes the SBDF and sysfs path to the test container via environment variables. `--first-vfio-nvme-device` is a fallback that searches for the first NVMe device already bound to vfio-pci if the targeted search fail's. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add an integration tests that verify VFIO passthrough with a physical NVMe device. Tests are gated behind the `vfio` pytest mark and FC_VFIO_PCI_SBDF environment variable so they only run when a suitable device is available. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VFIO tests need exclusive access to the passthrough device, so they cannot run in parallel with other tests. Add a separate Buildkite step in the PR pipeline that runs only the vfio-marked tests, similar to the existing performance step. CI instances will have an additional 1GB NVMe device at /dev/nvme1n1 for this purpose. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add docs/vfio.md covering how VFIO passthrough works in Firecracker, prerequisites (IOMMU, vfio-pci binding), configuration via API and config file, security considerations, snapshot incompatibility, and current limitations. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add a changelog entry for the new VFIO PCI device passthrough feature. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
do not merge: point to vfio artifacts Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Wire up the code to allow hot-plugging of VFIO devices after VM boot. The API is same as for usual VFIO device addition. Adding devices with duplicated ids is disallowed. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
With VFIO device hot-plug support we need to add all syscalls needed for VFIO devices creation to the VMM thread. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Implement VFIO device deinit logic and wire it up to the DELETE api. During VFIO device removal, device returns all resources it allocated back to the VM (except kvm_slots since we are not currently concerned with running out of them). The destruction happens in 2 parts (just like initialization) because it requires cooperation from both the device and from a pci_mngr. Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add information about VFIO hot-plug behaviour Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Contributor
Author
|
Closing in favor of #6055 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Changes
Add basic implementation of the VFIO device pass-through.
Current version only allows devices to be added before VM boot.
Other limitations:
Reason
Provide a way to pass physical PCI devices into VM
License Acceptance
By submitting this pull request, I confirm that my contribution is made under
the terms of the Apache 2.0 license. For more information on following Developer
Certificate of Origin and signing off your commits, please check
CONTRIBUTING.md.PR Checklist
tools/devtool checkbuild --allto verify that the PR passesbuild checks on all supported architectures.
tools/devtool checkstyleto verify that the PR passes theautomated style checks.
how they are solving the problem in a clear and encompassing way.
in the PR.
CHANGELOG.md.Runbook for Firecracker API changes.
integration tests.
TODO.rust-vmm.