Two tenants on the same node with overlapping ULAs can share one connection row in an egress shard's session table. The second tenant's replies are then re-encapsulated toward the first tenant's pod — cross-tenant delivery, not a dropped packet.
Cross-node collisions are already closed by keying flows on the encapsulation source (#525). This is the remaining half.
Cause
installEgressRoutes (internal/cnibgp/bgp.go) registers the operator-configured shard SID verbatim. Nothing writes the attachment's VRFID into the SID's Argument field (bits 69–80), so every tenant VRF encapsulates toward the identical destination and read_argument returns the same constant for all of them.
Confirmed on the wire in containerlab — three tenants in three VPCs, one destination address:
fd20:10:ff01::100:0 -> 2001:db8:ff02:9:e001::
fd20:20:ff01::100:0 -> 2001:db8:ff02:9:e001::
fd20:30:ff01::100:0 -> 2001:db8:ff02:9:e001::
ULAs are RFC 4193 locally assigned and not globally coordinated, so two tenants picking the same prefix is realistic rather than contrived. It is exactly what the Argument was meant to disambiguate.
Fix, in two parts
1. Write the Argument at route installation. argument uint16 is already a parameter of the enclosing function at the call site, so this part is small.
2. Widen the shard's BGP advertisement. shardAdvertisementPrefixes (internal/controller/egressshard_controller.go) advertises a /128. Once each tenant registers a different address, only one has a route, and table.Register resolves link and L2 information as part of the write — so every other tenant's CNI ADD fails. The shard must advertise <block>:<node-id>::/64 covering its Argument space.
Part 2 changes what every node in the fabric learns, so it wants review on its own terms.
Decision needed first
The two Argument namespaces have different scopes. allocateArgument filters BGPVRFInstances by RouterRef.Name, making Arguments unique per node. A shard receives from every node, so it needs values unique across all of them — tenant A on dfw and tenant B on sjc can both hold 0x005 today, and both would write 0x005 into the same shard's locator.
| Option |
Cost |
| Fabric-unique VRFIDs |
Capacity drops from 4095 VPCs per node to 4095 per fabric |
| Shard-side allocation |
Control-plane round trip on the pod attach path, and a new failure mode there |
| Hash the VPC identifier |
Not viable at 12 bits — birthday collisions reach 50% around 75 VPCs |
Either way, 12 bits caps distinguishable tenants per shard locator at 4095.
Blocked on this
Scope
Not a regression, and not urgent — cross-node was the bulk of the exposure and is closed. Until this lands, isolation is structural for tenants on different nodes and best-effort for tenants sharing one, which is weaker than enhancements#879 claims.
🤖 Generated with Claude Code
Two tenants on the same node with overlapping ULAs can share one connection row in an egress shard's session table. The second tenant's replies are then re-encapsulated toward the first tenant's pod — cross-tenant delivery, not a dropped packet.
Cross-node collisions are already closed by keying flows on the encapsulation source (#525). This is the remaining half.
Cause
installEgressRoutes(internal/cnibgp/bgp.go) registers the operator-configured shard SID verbatim. Nothing writes the attachment's VRFID into the SID's Argument field (bits 69–80), so every tenant VRF encapsulates toward the identical destination andread_argumentreturns the same constant for all of them.Confirmed on the wire in containerlab — three tenants in three VPCs, one destination address:
ULAs are RFC 4193 locally assigned and not globally coordinated, so two tenants picking the same prefix is realistic rather than contrived. It is exactly what the Argument was meant to disambiguate.
Fix, in two parts
1. Write the Argument at route installation.
argument uint16is already a parameter of the enclosing function at the call site, so this part is small.2. Widen the shard's BGP advertisement.
shardAdvertisementPrefixes(internal/controller/egressshard_controller.go) advertises a/128. Once each tenant registers a different address, only one has a route, andtable.Registerresolves link and L2 information as part of the write — so every other tenant's CNI ADD fails. The shard must advertise<block>:<node-id>::/64covering its Argument space.Part 2 changes what every node in the fabric learns, so it wants review on its own terms.
Decision needed first
The two Argument namespaces have different scopes.
allocateArgumentfiltersBGPVRFInstancesbyRouterRef.Name, making Arguments unique per node. A shard receives from every node, so it needs values unique across all of them — tenant A on dfw and tenant B on sjc can both hold0x005today, and both would write0x005into the same shard's locator.Either way, 12 bits caps distinguishable tenants per shard locator at 4095.
Blocked on this
verify:nat-collision, which today covers only the cross-node case.Scope
Not a regression, and not urgent — cross-node was the bulk of the exposure and is closed. Until this lands, isolation is structural for tenants on different nodes and best-effort for tenants sharing one, which is weaker than enhancements#879 claims.
🤖 Generated with Claude Code