Skip to content

[VFIO] add basic implementation - #5870

Closed
ShadowCurse wants to merge 28 commits into
firecracker-microvm:mainfrom
ShadowCurse:vfio_with_dependencies
Closed

ShadowCurse wants to merge 28 commits into
firecracker-microvm:mainfrom
ShadowCurse:vfio_with_dependencies

Conversation

@ShadowCurse

@ShadowCurse ShadowCurse commented May 8, 2026 •

Copy link
Copy Markdown
Contributor

Changes

Add basic implementation of the VFIO device pass-through.
Current version only allows devices to be added before VM boot.
Other limitations:

  • Only devices with MSIx interrupts are supported.
  • No INTx interrupt support
  • No ROM BAR/IO BAR support
  • No BAR relocation/resizing

Reason

Provide a way to pass physical PCI devices into VM

License Acceptance

By submitting this pull request, I confirm that my contribution is made under
the terms of the Apache 2.0 license. For more information on following Developer
Certificate of Origin and signing off your commits, please check
CONTRIBUTING.md.

PR Checklist

  • I have read and understand CONTRIBUTING.md.
  • I have run tools/devtool checkbuild --all to verify that the PR passes
    build checks on all supported architectures.
  • I have run tools/devtool checkstyle to verify that the PR passes the
    automated style checks.
  • I have described what is done in these changes, why they are needed, and
    how they are solving the problem in a clear and encompassing way.
  • I have updated any relevant documentation (both in code and in the docs)
    in the PR.
  • I have mentioned all user-facing changes in CHANGELOG.md.
  • If a specific issue led to this PR, this PR closes the issue.
  • When making API changes, I have followed the
    Runbook for Firecracker API changes.
  • I have tested all new and changed functionalities in unit tests and/or
    integration tests.
  • I have linked an issue to every new TODO.

  • This functionality cannot be added in rust-vmm.

@ShadowCurse ShadowCurse self-assigned this May 8, 2026
@codecov

codecov Bot commented May 8, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 37.50895% with 873 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.90%. Comparing base (1964800) to head (640c290).
⚠️ Report is 14 commits behind head on main.

⚠️ Current head 640c290 differs from pull request most recent head e97092d

Please upload reports for the commit e97092d to get more accurate results.

Files with missing lines Patch % Lines
src/vmm/src/vfio.rs 31.96% 630 Missing ⚠️
src/vmm/src/device_manager/pci_mngr.rs 2.59% 75 Missing ⚠️
src/vmm/src/device_manager/mod.rs 1.92% 51 Missing ⚠️
src/vmm/src/rpc_interface.rs 8.16% 45 Missing ⚠️
src/vmm/src/lib.rs 10.71% 25 Missing ⚠️
src/vmm/src/pci/msix.rs 66.66% 16 Missing ⚠️
src/vmm/src/resources.rs 35.00% 13 Missing ⚠️
.../firecracker/src/api_server/request/hotplug/mod.rs 0.00% 8 Missing ⚠️
src/vmm/src/builder.rs 66.66% 6 Missing ⚠️
src/vmm/src/vmm_config/vfio.rs 87.50% 3 Missing ⚠️
... and 1 more
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #5870      +/-   ##
==========================================
- Coverage   82.83%   80.90%   -1.93%     
==========================================
  Files         277      280       +3     
  Lines       30775    32010    +1235     
==========================================
+ Hits        25491    25898     +407     
- Misses       5284     6112     +828     
Flag Coverage Δ
5.10-m5n.metal 80.96% <37.50%> (-2.11%) ⬇️
5.10-m6a.metal 80.27% <37.50%> (-2.15%) ⬇️
5.10-m6g.metal 77.82% <37.50%> (-2.06%) ⬇️
5.10-m6i.metal 80.96% <37.50%> (-2.12%) ⬇️
5.10-m7a.metal-48xl 80.26% <37.50%> (-2.15%) ⬇️
5.10-m7g.metal 77.82% <37.50%> (-2.06%) ⬇️
5.10-m7i.metal-24xl 80.94% <37.50%> (-2.11%) ⬇️
5.10-m7i.metal-48xl 80.94% <37.50%> (-2.10%) ⬇️
5.10-m8g.metal-24xl 77.82% <37.50%> (-2.06%) ⬇️
5.10-m8g.metal-48xl 77.82% <37.50%> (-2.06%) ⬇️
5.10-m8i.metal-48xl 80.93% <37.50%> (-2.11%) ⬇️
5.10-m8i.metal-96xl 80.93% <37.50%> (-2.11%) ⬇️
6.1-m5n.metal 80.99% <37.50%> (-2.11%) ⬇️
6.1-m6a.metal 80.30% <37.50%> (-2.16%) ⬇️
6.1-m6g.metal 77.82% <37.50%> (-2.06%) ⬇️
6.1-m6i.metal 80.99% <37.50%> (-2.11%) ⬇️
6.1-m7a.metal-48xl 80.28% <37.50%> (-2.15%) ⬇️
6.1-m7g.metal 77.82% <37.50%> (-2.06%) ⬇️
6.1-m7i.metal-24xl 80.99% <37.50%> (-2.11%) ⬇️
6.1-m7i.metal-48xl 81.00% <37.50%> (-2.11%) ⬇️
6.1-m8g.metal-24xl 77.82% <37.50%> (-2.06%) ⬇️
6.1-m8g.metal-48xl 77.82% <37.50%> (-2.06%) ⬇️
6.1-m8i.metal-48xl 81.00% <37.50%> (-2.11%) ⬇️
6.1-m8i.metal-96xl 81.00% <37.50%> (-2.11%) ⬇️
6.18-m5n.metal 80.98% <37.50%> (-2.12%) ⬇️
6.18-m6a.metal 80.29% <37.50%> (-2.16%) ⬇️
6.18-m6g.metal 77.82% <37.50%> (-2.06%) ⬇️
6.18-m6i.metal 80.99% <37.50%> (-2.10%) ⬇️
6.18-m7a.metal-48xl 80.29% <37.50%> (-2.15%) ⬇️
6.18-m7g.metal 77.82% <37.50%> (-2.07%) ⬇️
6.18-m7i.metal-24xl 81.00% <37.50%> (-2.10%) ⬇️
6.18-m7i.metal-48xl 81.00% <37.50%> (-2.10%) ⬇️
6.18-m8g.metal-24xl 77.82% <37.50%> (-2.06%) ⬇️
6.18-m8g.metal-48xl 77.82% <37.50%> (-2.06%) ⬇️
6.18-m8i.metal-48xl 81.00% <37.50%> (-2.11%) ⬇️
6.18-m8i.metal-96xl 81.00% <37.50%> (-2.11%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@ShadowCurse
ShadowCurse force-pushed the vfio_with_dependencies branch 11 times, most recently from f6d6fea to 50e789e Compare May 14, 2026 16:29
@ShadowCurse
ShadowCurse force-pushed the vfio_with_dependencies branch 12 times, most recently from 2f84f01 to a21e87e Compare May 27, 2026 13:32
@ShadowCurse
ShadowCurse force-pushed the vfio_with_dependencies branch 4 times, most recently from b2ea5ea to 528e62b Compare May 29, 2026 11:53
@ShadowCurse
ShadowCurse force-pushed the vfio_with_dependencies branch from 528e62b to efa67e3 Compare June 8, 2026 13:25
Add the VfioConfig and VfioConfigs types for describing VFIO device
configuration. Wire them into VmResources and VmmConfig so that VFIO
devices can be specified before boot. Actual device setup will be added
in later commits.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add PUT /vfio/{id} API endpoint for configuring VFIO passthrough
devices.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
- add `vfio-bindings` and `vfio-ioctls`
- make `arrayvec` non optional

All of these will be used in the future VFIO commits.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Create an empty module where VFIO code will be.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
First thing to do with VFIO device is to scan it's capabilities and
extended capabilities to find MSI-X cap and caps we want to filter out.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Next we need to gather information about BARs the device has and
allocate space for them in the guest memory.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VFIO BAR regions containing MSI-X table/PBA will be split into
mmappable and emulated parts. KVM memory slots require host-page
alignment, but MSI-X structures can sit at arbitrary offsets
within a BAR. Add additional helper functions for the calculations of
these mappable/emulated BAR regions.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
With information about BARs and the capabilities, we can calculate
areas of BARs we can safely DMA map into the guest. Everything outside
those areas are subject to emulation on our end.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
After calculating the areas of BARs we can map to the guest, do this
mapping. It involves `mmap`ing the device BAR into Firecracker virtual
space first, then setting up the DMA mapping for this virtual address to
the guest physical memory.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
After all the previous work with BARs and capabilities we can finally
put everything together into a VfioDevice type.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VfioDevice will have to emulate both Bus accesses (for configuration
space) and Pci accesses for the emulated BARs holes. This commit
implements the handling for the Bus accesses.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Now implement the Pci emulation for the VfioDevice to handle emulated
BAR accesses to the MSIx table/pba areas.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
In addition to the VfioDevice setting up DMA for the device BARs, we
need to set up DMA for the whole guest RAM since the device will need to
access it. These functions will be used in the next commits.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add functions for creation of KVM VFIO device and VFIO container.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add logic to the PciDevices to create new VFIO devices. As an additional
step in VFIO device setup, guest RAM regions are mapped into the VFIO
container's IOMMU so the device can DMA directly to guest memory.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Current VFIO implementation has some restrictions:
- Does not work without PCI since VFIO devices are PCI devices
- Does not work with virtio-mem device since we don't update DMA
  mappings on hot-plug/unplug
- Does not work with virtio-balloon since it can `fadvise` on memory

In order to prevent VMs being launched with invalid configurations,
implement multiple checks for invalid configurations:
- At API level, prevent adding of incompatible combinations (VFIO after
  balloon/mem or in reverse)
- At vm creation or snapshot restoraton since they get VmResources from
  other sources.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VFIO device state is opaque to the VMM and cannot be serialized
or restored. Add VFIO devices to the list of snapshot-incompatible
devices so that snapshot requests are rejected with a clear error
instead of producing a corrupt snapshot.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VFIO devices will use pread64/pwrite64 syscalls (from vfio-ioctls) to
interact with BARs during runtime. Add them to the VPU thread syscall
lists.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add devtool options for preparing a PCI device for VFIO passthrough
testing. `--vfio-nvme-device` accepts a block device path (e.g.
/dev/nvme1n1) or a PCI SBDF, resolves it to a PCI device, binds it to
vfio-pci, and passes the SBDF and sysfs path to the test container via
environment variables. `--first-vfio-nvme-device` is a fallback that
searches for the first NVMe device already bound to vfio-pci if the
targeted search fail's.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add an integration tests that verify VFIO passthrough with a physical
NVMe device. Tests are gated behind the `vfio` pytest mark and
FC_VFIO_PCI_SBDF environment variable so they only run when a suitable
device is available.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
VFIO tests need exclusive access to the passthrough device, so they
cannot run in parallel with other tests. Add a separate Buildkite step
in the PR pipeline that runs only the vfio-marked tests, similar to the
existing performance step. CI instances will have an additional 1GB NVMe
device at /dev/nvme1n1 for this purpose.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add docs/vfio.md covering how VFIO passthrough works in Firecracker,
prerequisites (IOMMU, vfio-pci binding), configuration via API and
config file, security considerations, snapshot incompatibility, and
current limitations.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add a changelog entry for the new VFIO PCI device passthrough feature.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
do not merge: point to vfio artifacts

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Wire up the code to allow hot-plugging of VFIO devices after VM boot.
The API is same as for usual VFIO device addition. Adding devices with
duplicated ids is disallowed.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
With VFIO device hot-plug support we need to add all syscalls needed for
VFIO devices creation to the VMM thread.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Implement VFIO device deinit logic and wire it up to the DELETE api.

During VFIO device removal, device returns all resources it allocated
back to the VM (except kvm_slots since we are not currently concerned
with running out of them). The destruction happens in 2 parts (just
like initialization) because it requires cooperation from both the
device and from a pci_mngr.

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
Add information about VFIO hot-plug behaviour

Signed-off-by: Egor Lazarchuk <yegorlz@amazon.co.uk>
@ShadowCurse

Copy link
Copy Markdown
Contributor Author

Closing in favor of #6055

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Status: Awaiting review Indicates that a pull request is ready to be reviewed Type: Enhancement Indicates new feature requests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants