You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Under any VMM that presents a flat PCI topology to the guest, RM cannot discover a
PCIe root port. NV2080_CTRL_BUS_INFO_INDEX_PCIE_ROOT_LINK_CAPS returns zero and the
PCIe read and write copy engines are sized from a link whose generation is unknown.
On a host where the GPU link and the host root port are different generations, this
costs roughly half of bidirectional host bandwidth.
I have a small patch that lets the platform declare the generation, and I would like
to discuss the approach before opening a PR.
Status quo
kceGetPceConfigForLceType_IMPL sizes the PCIe read and write logical copy engines
from the root port link capability, and falls back to UNKNOWN_PCIE_GEN_SPEED when
it cannot read one.
busInfo.index=NV2080_CTRL_BUS_INFO_INDEX_PCIE_ROOT_LINK_CAPS;
if (kbifControlGetPCIEInfo(pGpu, pKernelBif, &busInfo) ==NV_OK)
pceConfigParams.metadataForLceType=ceEncodeLceTypeMetadataForPcie(pGpu, busInfo.data);
elsepceConfigParams.metadataForLceType=UNKNOWN_PCIE_GEN_SPEED;
That index is serviced by reading config space of pGpu->gpuClData.rootPort, which
the chipset layer populates by walking up the PCI hierarchy.
caseNV2080_CTRL_BUS_INFO_INDEX_PCIE_ROOT_LINK_CAPS:
if (clPcieReadPortConfigReg(pGpu, pCl, &pGpu->gpuClData.rootPort,
CL_PCIE_LINK_CAP, &data) !=NV_OK)
data=0;
In a guest with no emulated root port the walk finds nothing, so data stays zero, ceEncodeLceTypeMetadataForPcie takes its default branch and returns UNKNOWN_PCIE_GEN_SPEED. Note this is not a failed read that gets handled. The call
returns NV_OK with data of zero, so the encoder runs on a value that decodes to
no known generation.
Two symptoms make it easy to confirm on any affected system.
With no host information nvidia-smi substitutes Device Max, so a Gen6 capable GPU behind a Gen5 root port reports Max : 6. That field is the GEN bits of PCIE_GEN_INFO, not ROOT_LINK_CAPS, so it is a second symptom of the same missing data rather than the value the copy engines are sized from. The copy engine path receives UNKNOWN_PCIE_GEN_SPEED and falls back to an untuned symmetric split, which is what the metadata column below shows.
Measured impact
HGX B300 guest, 8 GPUs passed through, flat PCI topology, no emulated root port. The
GPU trains Gen6 x16 to its ConnectX-8 switch while the host root port above it is
Gen5 x16, so the two genuinely differ. Same driver binary in both runs, only a regkey
changed.
Instrumenting the reply from NV2080_CTRL_CMD_INTERNAL_CE_GET_PCE_CONFIG_FOR_LCE_TYPE:
condition
lceType
metadata
numPces
supportedPceMask
status quo
6 PCIE_RD
0xffffffff
3
0x3333
status quo
7 PCIE_WR
0xffffffff
3
0x3333
Gen5 declared
6 PCIE_RD
0x4
5
0x3373
Gen5 declared
7 PCIE_WR
0x4
2
0x3373
dcgmi diag --run pcie, GB/s, one column per GPU:
metric
status quo
Gen5 declared
GPU to Host
57.39 57.39 57.39 57.40
57.39 57.39 57.39 57.40
Host to GPU
55.65 55.65 55.65 55.80
55.66 55.67 55.66 55.82
bidirectional
44.49 40.31 44.55 53.14
98.76 98.88 98.84 98.79
Bidirectional roughly doubles. Notably :
Unidirectional figures are unchanged to two decimal places, so the measurement
setup is constant and only the concurrent case moved.
In the status quo, bidirectional is slower than either single direction. That is
the signature of the read and write engines contending rather than running in
parallel. Afterwards it exceeds both and approaches the full duplex sum.
Variance across GPUs collapses from a 12.8 GB/s spread to 0.12 GB/s.
The fallback is chip dependent and worth knowing before anyone tries to reproduce
this. On B200 the same query under the same UNKNOWN metadata returns 4 and 2 rather
than 3 and 3, so B200 shows no change from this key and cannot be used to evaluate
it.
Proposal
Add a regkey, RmPcieHostLinkGen, consulted only when the root port link capability
read yields no max speed.
Values 1 through 6 map directly onto the NV2080 MAX_SPEED field encoding, so the
existing encoder consumes the result without modification and no translation table is
needed. Because the key is consulted only when MAX_SPEED reads zero, behaviour on
any system with a visible root port is untouched, which keeps bare metal entirely out
of scope.
Roughly 100 lines across nvrm_registry.h and kernel_bif.c. A value outside 1
through 6 warns once per GPU and falls back to current behaviour, verified with RmPcieHostLinkGen=0x9 reproducing the stock 3 and 3 split.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Summary
Under any VMM that presents a flat PCI topology to the guest, RM cannot discover a
PCIe root port.
NV2080_CTRL_BUS_INFO_INDEX_PCIE_ROOT_LINK_CAPSreturns zero and thePCIe read and write copy engines are sized from a link whose generation is unknown.
On a host where the GPU link and the host root port are different generations, this
costs roughly half of bidirectional host bandwidth.
I have a small patch that lets the platform declare the generation, and I would like
to discuss the approach before opening a PR.
Status quo
kceGetPceConfigForLceType_IMPLsizes the PCIe read and write logical copy enginesfrom the root port link capability, and falls back to
UNKNOWN_PCIE_GEN_SPEEDwhenit cannot read one.
That index is serviced by reading config space of
pGpu->gpuClData.rootPort, whichthe chipset layer populates by walking up the PCI hierarchy.
In a guest with no emulated root port the walk finds nothing, so
datastays zero,ceEncodeLceTypeMetadataForPcietakes its default branch and returnsUNKNOWN_PCIE_GEN_SPEED. Note this is not a failed read that gets handled. The callreturns
NV_OKwithdataof zero, so the encoder runs on a value that decodes tono known generation.
Two symptoms make it easy to confirm on any affected system.
With no host information
nvidia-smisubstitutesDevice Max, so a Gen6 capable GPU behind a Gen5 root port reportsMax : 6. That field is the GEN bits ofPCIE_GEN_INFO, notROOT_LINK_CAPS, so it is a second symptom of the same missing data rather than the value the copy engines are sized from. The copy engine path receivesUNKNOWN_PCIE_GEN_SPEEDand falls back to an untuned symmetric split, which is what the metadata column below shows.Measured impact
HGX B300 guest, 8 GPUs passed through, flat PCI topology, no emulated root port. The
GPU trains Gen6 x16 to its ConnectX-8 switch while the host root port above it is
Gen5 x16, so the two genuinely differ. Same driver binary in both runs, only a regkey
changed.
Instrumenting the reply from
NV2080_CTRL_CMD_INTERNAL_CE_GET_PCE_CONFIG_FOR_LCE_TYPE:PCIE_RD0xffffffff0x3333PCIE_WR0xffffffff0x3333PCIE_RD0x40x3373PCIE_WR0x40x3373dcgmi diag --run pcie, GB/s, one column per GPU:Bidirectional roughly doubles. Notably :
setup is constant and only the concurrent case moved.
the signature of the read and write engines contending rather than running in
parallel. Afterwards it exceeds both and approaches the full duplex sum.
The fallback is chip dependent and worth knowing before anyone tries to reproduce
this. On B200 the same query under the same
UNKNOWNmetadata returns 4 and 2 ratherthan 3 and 3, so B200 shows no change from this key and cannot be used to evaluate
it.
Proposal
Add a regkey,
RmPcieHostLinkGen, consulted only when the root port link capabilityread yields no max speed.
Values 1 through 6 map directly onto the NV2080
MAX_SPEEDfield encoding, so theexisting encoder consumes the result without modification and no translation table is
needed. Because the key is consulted only when
MAX_SPEEDreads zero, behaviour onany system with a visible root port is untouched, which keeps bare metal entirely out
of scope.
Roughly 100 lines across
nvrm_registry.handkernel_bif.c. A value outside 1through 6 warns once per GPU and falls back to current behaviour, verified with
RmPcieHostLinkGen=0x9reproducing the stock 3 and 3 split.All reactions