Optimize fixed-signature cdata calls - #282
Open
Johnny-Kao wants to merge 1 commit into
Open
Johnny-Kao wants to merge 1 commit into
Johnny-Kao wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
This PR adds a conservative fast path for a small subset of fixed-signature
<cdata>function-pointer calls.It reduces call overhead by avoiding part of the generic argument-conversion and allocation path for simple scalar signatures, while deliberately keeping the existing implementation as the fallback for everything else.
Measured improvement on a paired Linux benchmark is roughly 25–30% for the specialized scalar calls, with about 10–20% improvement on the shared call-path cases measured here.
The main cost is memory:
CDataObjectgrows from 40 to 48 bytes on 64-bit CPython because of the vectorcall slot. That cost is measurable, but the current layout keeps the implementation simple and avoids introducing a separate callable-specific object hierarchy in this PR.The patch has been tested on macOS ARM64 and Linux x86_64, including ASAN/UBSAN and Python 3.14 free-threaded builds, with no unexpected test failures.
What changed
The call path now has two branches:
The fast path currently covers:
int()int(int)int(int, int)int(int, int, int, int)double(double, double)Unsupported signatures, variadic calls, subclasses, coercions, out-of-range integers, unusual alignments, and other complex cases continue through the existing path.
Call path before and after
Before
flowchart LR B1["Python call"] --> B2["tp_call + args tuple"] B2 --> B3["Generic argument conversion"] B3 --> B4["Heap exchange buffer"] B4 --> B5["ffi_call"] B5 --> B6["Generic result conversion"] B6 --> B7["Python result"]After this PR
Blue outline = added or materially changed in this PR.
flowchart LR A1["Python call"] --> A2["Vectorcall"] A2 --> A3{"Fast-path eligible?"} A3 -->|"Yes"| A4["Fixed-signature fast prepare"] A4 --> A5["Stack buffer when safe"] A3 -->|"No"| A6["Existing generic path"] A5 --> A7["ffi_call"] A6 --> A7 A7 --> A8["Direct scalar result / generic result"] A8 --> A9["Python result"] classDef changed stroke:#0969da,stroke-width:3px,font-weight:bold class A2,A3,A4,A5,A8 changedThe important boundary is that the new path only changes call entry, preparation, and result conversion for a small proven subset.
ffi_call()itself is unchanged, and each logical Python call still executes it exactly once.Why
For ABI-mode function-pointer calls, a meaningful part of the latency occurs before
ffi_call():For a fixed scalar signature, much of that work is predictable.
This patch classifies that small predictable subset ahead of time and uses a shorter preparation path.
Why keep the generic path?
The generic call path is intentionally retained rather than replaced.
CFFI's existing path handles a much broader semantic surface, including pointers, aggregates, subclasses, coercions, variadic calls, uncommon integer widths, and ABI-specific cases.
The new path only covers cases for which the conversion behavior can be kept narrow and directly validated.
Replacing the generic implementation would therefore substantially increase the semantic and platform surface of this change without being necessary to obtain the measured benefit.
This PR treats specialization as an additive optimization:
Measured effect
Paired benchmark on the same Ubuntu 24.04 x86_64 GitHub runner with Python 3.14.7:
ret0()int(int)int(int,int)int(int,int,int,int)double(double,double)On Apple silicon, a 20M-call
int(int,int)probe reduced CPU time from:77.49 ns/call -> 58.46 ns/callor about 24.6% lower CPU time per call.
That reduction comes from the combined call-path changes rather than from vectorcall alone:
Why does
CDataObjectgrow?Vectorcall requires per-object vectorcall storage in the current implementation, increasing
CDataObjectfrom:40 bytes -> 48 byteson 64-bit CPython.
This keeps the implementation simple:
The memory cost is real.
In a synthetic workload containing 500k CData objects, measured maximum RSS increased by about 8 MB.
A possible follow-up is to move vectorcall storage to callable CData objects only. That could reduce the memory cost, but it requires a more invasive object-layout/type-design change and is intentionally kept outside this PR.
Safety boundary
The fast path is deliberately narrow.
It is used only when both the signature and runtime values satisfy explicit eligibility checks.
In particular:
Selection and argument preparation happen before the C function is invoked, so fallback cannot cause the C function to execute twice.
Anything outside the supported domain continues through the existing generic path.
Validation
macOS ARM64, Python 3.14.7
Ubuntu 24.04 x86_64, Python 3.14.7
Python 3.14.7 free-threaded
Py_GIL_DISABLED=1Also:
git diff --checkpassesDesign trade-offs and rejected alternatives
The final design intentionally favors a smaller, easier-to-review fast path over a more aggressive rewrite.
Replacing more of the generic conversion path
A broader rewrite can produce larger gains, but it also expands the semantic surface that must be reimplemented and reviewed.
The current PR keeps the optimized domain small and preserves the historical implementation for everything else.
Partial fast conversion with fallback
An earlier prototype allowed individual arguments to use fast conversion before later falling back to the generic path.
That was rejected because mixed execution makes semantic reasoning harder, especially around subclasses, coercions, and exceptional cases.
The final preparation path is effectively all-fast or all-generic.
Runtime calibration
A prototype measured both preparation paths during the first few calls and cached the faster one per signature.
It worked experimentally, but was not included because it introduces:
The static conservative selector is easier to reason about and review.
Shadow verification
A development prototype prepared arguments through both implementations and compared them before issuing a single
ffi_call().This was useful for validating equivalence, but duplicated preparation work and was therefore not retained in the production path.
Remaining risks
The remaining risks are primarily maintenance and platform breadth rather than a known correctness issue.
Areas to watch include:
The current design limits these risks by keeping the supported domain narrow and falling back whenever a case is not explicitly supported.
Possible follow-up work
If this approach proves useful, follow-up work could reduce some of the remaining cost without expanding this PR:
These are intentionally left out of this PR to keep the current change bounded and reviewable.