Skip to content

Egress shard routes carry no per-tenant Argument, so same-node tenants can share a translation #538

Description

@privateip

Two tenants on the same node with overlapping ULAs can share one connection row in an egress shard's session table. The second tenant's replies are then re-encapsulated toward the first tenant's pod — cross-tenant delivery, not a dropped packet.

Cross-node collisions are already closed by keying flows on the encapsulation source (#525). This is the remaining half.

Cause

installEgressRoutes (internal/cnibgp/bgp.go) registers the operator-configured shard SID verbatim. Nothing writes the attachment's VRFID into the SID's Argument field (bits 69–80), so every tenant VRF encapsulates toward the identical destination and read_argument returns the same constant for all of them.

Confirmed on the wire in containerlab — three tenants in three VPCs, one destination address:

fd20:10:ff01::100:0  ->  2001:db8:ff02:9:e001::
fd20:20:ff01::100:0  ->  2001:db8:ff02:9:e001::
fd20:30:ff01::100:0  ->  2001:db8:ff02:9:e001::

ULAs are RFC 4193 locally assigned and not globally coordinated, so two tenants picking the same prefix is realistic rather than contrived. It is exactly what the Argument was meant to disambiguate.

Fix, in two parts

1. Write the Argument at route installation. argument uint16 is already a parameter of the enclosing function at the call site, so this part is small.

2. Widen the shard's BGP advertisement. shardAdvertisementPrefixes (internal/controller/egressshard_controller.go) advertises a /128. Once each tenant registers a different address, only one has a route, and table.Register resolves link and L2 information as part of the write — so every other tenant's CNI ADD fails. The shard must advertise <block>:<node-id>::/64 covering its Argument space.

Part 2 changes what every node in the fabric learns, so it wants review on its own terms.

Decision needed first

The two Argument namespaces have different scopes. allocateArgument filters BGPVRFInstances by RouterRef.Name, making Arguments unique per node. A shard receives from every node, so it needs values unique across all of them — tenant A on dfw and tenant B on sjc can both hold 0x005 today, and both would write 0x005 into the same shard's locator.

Option Cost
Fabric-unique VRFIDs Capacity drops from 4095 VPCs per node to 4095 per fabric
Shard-side allocation Control-plane round trip on the pod attach path, and a new failure mode there
Hash the VPC identifier Not viable at 12 bits — birthday collisions reach 50% around 75 VPCs

Either way, 12 bits caps distinguishable tenants per shard locator at 4095.

Blocked on this

Scope

Not a regression, and not urgent — cross-node was the bulk of the exposure and is closed. Until this lands, isolation is structural for tenants on different nodes and best-effort for tenants sharing one, which is weaker than enhancements#879 claims.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

Fields

Priority

Medium

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions