Skip to content

Optimize fixed-signature cdata calls - #282

Open
Johnny-Kao wants to merge 1 commit into
python-cffi:mainfrom
Johnny-Kao:perf/cdata-fixed-call-fast-path
Open

Johnny-Kao wants to merge 1 commit into
python-cffi:mainfrom
Johnny-Kao:perf/cdata-fixed-call-fast-path

Conversation

@Johnny-Kao

@Johnny-Kao Johnny-Kao commented Oct 1, 2026 •

Copy link
Copy Markdown

TL;DR

This PR adds a conservative fast path for a small subset of fixed-signature <cdata> function-pointer calls.

It reduces call overhead by avoiding part of the generic argument-conversion and allocation path for simple scalar signatures, while deliberately keeping the existing implementation as the fallback for everything else.

Measured improvement on a paired Linux benchmark is roughly 25–30% for the specialized scalar calls, with about 10–20% improvement on the shared call-path cases measured here.

The main cost is memory: CDataObject grows from 40 to 48 bytes on 64-bit CPython because of the vectorcall slot. That cost is measurable, but the current layout keeps the implementation simple and avoids introducing a separate callable-specific object hierarchy in this PR.

The patch has been tested on macOS ARM64 and Linux x86_64, including ASAN/UBSAN and Python 3.14 free-threaded builds, with no unexpected test failures.


What changed

The call path now has two branches:

  • a narrow fast path for a few fixed scalar signatures;
  • the existing generic implementation for all other calls.

The fast path currently covers:

  • int()
  • int(int)
  • int(int, int)
  • int(int, int, int, int)
  • double(double, double)

Unsupported signatures, variadic calls, subclasses, coercions, out-of-range integers, unusual alignments, and other complex cases continue through the existing path.

Call path before and after

Before

flowchart LR
    B1["Python call"] --> B2["tp_call + args tuple"]
    B2 --> B3["Generic argument conversion"]
    B3 --> B4["Heap exchange buffer"]
    B4 --> B5["ffi_call"]
    B5 --> B6["Generic result conversion"]
    B6 --> B7["Python result"]
Loading

After this PR

Blue outline = added or materially changed in this PR.

flowchart LR
    A1["Python call"] --> A2["Vectorcall"]
    A2 --> A3{"Fast-path eligible?"}
    A3 -->|"Yes"| A4["Fixed-signature fast prepare"]
    A4 --> A5["Stack buffer when safe"]
    A3 -->|"No"| A6["Existing generic path"]
    A5 --> A7["ffi_call"]
    A6 --> A7
    A7 --> A8["Direct scalar result / generic result"]
    A8 --> A9["Python result"]

    classDef changed stroke:#0969da,stroke-width:3px,font-weight:bold
    class A2,A3,A4,A5,A8 changed
Loading

The important boundary is that the new path only changes call entry, preparation, and result conversion for a small proven subset.

ffi_call() itself is unchanged, and each logical Python call still executes it exactly once.


Why

For ABI-mode function-pointer calls, a meaningful part of the latency occurs before ffi_call():

  • Python call setup;
  • exchange-buffer allocation;
  • generic type dispatch;
  • repeated argument-conversion checks.

For a fixed scalar signature, much of that work is predictable.

This patch classifies that small predictable subset ahead of time and uses a shorter preparation path.


Why keep the generic path?

The generic call path is intentionally retained rather than replaced.

CFFI's existing path handles a much broader semantic surface, including pointers, aggregates, subclasses, coercions, variadic calls, uncommon integer widths, and ABI-specific cases.

The new path only covers cases for which the conversion behavior can be kept narrow and directly validated.

Replacing the generic implementation would therefore substantially increase the semantic and platform surface of this change without being necessary to obtain the measured benefit.

This PR treats specialization as an additive optimization:

  • proven common cases take the shorter path;
  • everything else keeps the mature implementation.

Measured effect

Paired benchmark on the same Ubuntu 24.04 x86_64 GitHub runner with Python 3.14.7:

case main patched improvement
ret0() 125.83 ns 112.43 ns 10.65%
int(int) 171.40 ns 129.14 ns 24.66%
int(int,int) 195.96 ns 143.55 ns 26.75%
int(int,int,int,int) 243.29 ns 171.15 ns 29.65%
double(double,double) 203.39 ns 152.63 ns 24.96%
pointer/general path 167.40 ns 133.93 ns 19.99%

On Apple silicon, a 20M-call int(int,int) probe reduced CPU time from:

77.49 ns/call -> 58.46 ns/call

or about 24.6% lower CPU time per call.

That reduction comes from the combined call-path changes rather than from vectorcall alone:

  • vectorcall entry;
  • preclassified fixed signatures;
  • direct scalar conversion;
  • reduced allocation work.

Why does CDataObject grow?

Vectorcall requires per-object vectorcall storage in the current implementation, increasing CDataObject from:

40 bytes -> 48 bytes

on 64-bit CPython.

This keeps the implementation simple:

  • all CData objects retain one stable layout;
  • vectorcall initialization stays inside the existing object-construction machinery;
  • this PR does not need to introduce a second callable-specific CData hierarchy or conditional object layout.

The memory cost is real.

In a synthetic workload containing 500k CData objects, measured maximum RSS increased by about 8 MB.

A possible follow-up is to move vectorcall storage to callable CData objects only. That could reduce the memory cost, but it requires a more invasive object-layout/type-design change and is intentionally kept outside this PR.


Safety boundary

The fast path is deliberately narrow.

It is used only when both the signature and runtime values satisfy explicit eligibility checks.

In particular:

  • subclasses and custom coercions fall back;
  • out-of-range integers fall back;
  • unsupported signatures fall back;
  • variadic calls do not use the fast path;
  • big-endian integer fast-result plans are disabled;
  • stack buffering is used only when the required libffi alignment is supported.

Selection and argument preparation happen before the C function is invoked, so fallback cannot cause the C function to execute twice.

Anything outside the supported domain continues through the existing generic path.


Validation

macOS ARM64, Python 3.14.7

  • 1686 passed
  • 150 skipped
  • 4 xfailed
  • 0 failed
  • targeted differential and boundary tests passed
  • UBSAN targeted validation passed
  • wheel and sdist builds passed
  • fresh-wheel installation smoke passed

Ubuntu 24.04 x86_64, Python 3.14.7

  • 1742 passed
  • 94 skipped
  • 4 xfailed
  • 0 failed
  • ASAN + UBSAN targeted validation passed

Python 3.14.7 free-threaded

  • Py_GIL_DISABLED=1
  • targeted smoke passed
  • focused backend tests passed

Also:

  • git diff --check passes
  • the tested native module was verified to come from the current source tree

Design trade-offs and rejected alternatives

The final design intentionally favors a smaller, easier-to-review fast path over a more aggressive rewrite.

Replacing more of the generic conversion path

A broader rewrite can produce larger gains, but it also expands the semantic surface that must be reimplemented and reviewed.

The current PR keeps the optimized domain small and preserves the historical implementation for everything else.

Partial fast conversion with fallback

An earlier prototype allowed individual arguments to use fast conversion before later falling back to the generic path.

That was rejected because mixed execution makes semantic reasoning harder, especially around subclasses, coercions, and exceptional cases.

The final preparation path is effectively all-fast or all-generic.

Runtime calibration

A prototype measured both preparation paths during the first few calls and cached the faster one per signature.

It worked experimentally, but was not included because it introduces:

  • runtime timing state;
  • timer portability concerns;
  • additional synchronization requirements for free-threaded Python;
  • more maintenance complexity.

The static conservative selector is easier to reason about and review.

Shadow verification

A development prototype prepared arguments through both implementations and compared them before issuing a single ffi_call().

This was useful for validating equivalence, but duplicated preparation work and was therefore not retained in the production path.


Remaining risks

The remaining risks are primarily maintenance and platform breadth rather than a known correctness issue.

Areas to watch include:

  • future CPython changes to vectorcall or object layout;
  • expansion of the fast-path type set;
  • unusual ABI/alignment requirements;
  • big-endian platforms;
  • growth in duplicated conversion logic if specialization expands too aggressively.

The current design limits these risks by keeping the supported domain narrow and falling back whenever a case is not explicitly supported.


Possible follow-up work

If this approach proves useful, follow-up work could reduce some of the remaining cost without expanding this PR:

  • move more signature classification to function-object construction time;
  • avoid vectorcall storage for non-callable CData objects;
  • add additional scalar signatures only after separate ABI/boundary validation;
  • reduce the remaining runtime selector work for signatures that are permanently classifiable.

These are intentionally left out of this PR to keep the current change bounded and reviewable.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant