Parallel test execution - #6784
Conversation
API Surface ChangesIf any of the additions below are not intended as public API, mark them with New API SurfaceClasses
Methods
Modified API SurfaceMethods
|
9a08909 to
efcd7ee
Compare
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #6784 +/- ##
===========================================
Coverage 99.48% 99.49%
- Complexity 9436 9918 +482
===========================================
Files 916 939 +23
Lines 28798 30250 +1452
===========================================
+ Hits 28650 30097 +1447
- Misses 148 153 +5 ☔ View full report in Codecov by Harness. |
|
My 2 cents on the topic:
Parallel execution can't and will never be deterministic, a tiny difference between how much a unit lasts results in the following test being run in worker X instead of Y. |
|
Thank you, @Slamdunk, for taking the time to write all of this up. There's a lot of hard-won experience in here, and I appreciate it. I want to set expectations honestly, though: for me, native parallel test execution in PHPUnit is still just an idea, and this pull request is a (remarkably well-working) proof of concept rather than a roadmap commitment. If I ever decide to be serious about this and actually ship it, the feature will have limitations. And those limitations must (and will) be very well documented. Complaints about them will then be kindly refused. 🙂 Even in that case, I neither intend nor expect this to replace dedicated solutions such as ParaTest or Paraunit. The scope I have in mind right now is roughly: PHPUnit supports parallel execution for well-architected test suites that properly deal with resource-usage conflicts like databases, etc. So several of the needs you describe, coordinated bootstrapping, functional/method-level distribution, deterministic replay of a run, are exactly the kind of thing that lives outside that scope and is better served by the tools built specifically for it. That said, your notes are genuinely useful for thinking about where the boundaries of that scope should sit, so thank you again. |
4e03c18 to
5053a2d
Compare
5053a2d to
cf8f973
Compare
|
I have briefly looked into whether the new I/O polling API that PHP 8.6 introduces (RFC: poll_api) could replace the file-based completion polling used by the parallel test runner. The conclusion is that it cannot, because it has the exact limitation that forced the file-based design in the first place: on Windows, it cannot wait on
However:
In other words, Therefore, the uniform file-based completion polling stays. This decision should be revisited only if a future PHP version ships a poll handle that can wait on pipes or process handles on Windows, which is what this feature actually needs. |
cf8f973 to
1e97c6b
Compare
|
I see that it's up to the user to select how many worker to use. The You might be interested in giving it a try, so PHPUnit can provide |
I am aware of |
04cb38d to
43341d8
Compare
908d59e to
4b04c6a
Compare
2c7a98a to
c034df8
Compare
|
e9e9866 to
7c6f829
Compare
…atch with a telling exception when the worker command cannot be encoded, so that a data provider key that is not valid UTF-8 cannot abort a parallel run
…d, and do not start the FILE section of one terminated during SKIPIF, when a parallel run stops early, so that an abandoned test still cleans up after itself
…onflicts with every other test entirely on their own, draining the worker pool and the PHPT runner first, so that the exclusivity these declarations promise actually holds
…'s classes with named constructors, derivation methods, and shared helpers, so that each invariant lives in one place
…es in one pass through one metadata traversal, and poll a worker's event stream with a stat instead of a re-read, so that a parallel run wastes less work as suites grow
…e helper and two template fragments, writing the configuration and source map once per run for all of them, so that isolated processes and parallel workers cannot drift apart
…e the result envelope's decoding between its two consumers, and let a work unit report its own recorded duration, so that each half of the parallel runner's data model lives in exactly one place
…and the facade's dispatcher selection, and record at collection time which chunk envelopes the units emit, so that behavior the sequential and parallel runners must perform identically is expressed exactly once
…let those descriptors travel in a serialized command instead of a JSON-encoded one, so that a member's encoding and decoding live in one place and no value it carries needs a base64 shim to survive the trip to a worker
The tests of a data provider method, the repetitions of a repeated test method, and the attempts of a retried test method travel to a worker as the suite that aggregates them. The descriptor of such a suite asked the suite for its tests, which returns all of them: test selection (--filter, --group, --exclude-group, --filter-test-id, …) is a filter iterator that the selection injects into the suite, and it applies only when the suite is iterated, as TestSuite::takeTests() does in a sequential run. Every test of a selected test class was therefore executed in a parallel run, no matter which of them the selection had actually selected. The descriptors now iterate the suite, and the collection walk skips a suite that the selection has emptied, for which TestSuite::run() returns early in a sequential run.
A test case that runs in a worker process is not built by TestBuilder but recreated from its descriptor, which carries only the name of the test method and the data that distinguishes one invocation of it from another. The settings that TestBuilder derives from metadata and from the configuration of the test runner -- process isolation, the preservation of global state, and the backup of global variables and static properties -- were therefore never applied to it. As a result, no snapshot of global state was taken for such a test case and state leaked from one test of a test class to the next inside a worker, regardless of the backupGlobals and backupStaticProperties configuration settings, of the --globals-backup and --static-backup CLI options, and of the #[BackupGlobals] and #[BackupStaticProperties] attributes. Derive those settings again in the worker process, where the metadata and the configuration they are derived from are available as well, by letting the descriptor delegate to TestBuilder.
The tests in tests/end-to-end/parallel/selection come in pairs: one runs the fixture sequentially, the other runs the same fixture with --parallel=2, and both pin the same --EXPECTF-- output. The sequential test of a pair documents which tests --filter, --group, --exclude-group and --repeat select, the parallel test asserts that a parallel run selects exactly the same ones. Should the two ever diverge again, the parallel test of the pair fails. The fixture is driven by data providers on purpose. Test selection is a lazy filter that only takes effect while a test suite is iterated, so a plain test method is selected correctly no matter how the parallel runner reads a suite, whereas the tests of a data provider suite, a repeat suite or a retry suite are not. Only a data provider driven fixture therefore covers the case that went unnoticed.
A test case that a parallel worker recreates from its descriptor is not built by TestBuilder::build(), so the settings that build() derives from the metadata of the test method and from the configuration of the test runner are applied to it by TestBuilder::configure() instead. Cover that method: for a test class with class-level metadata for isolation and for a test class with metadata for excluding global variables and static properties from the backup, assert the settings it applies, and for every kind of test class the fixtures provide, assert that it leaves a manually instantiated test case configured exactly as build() leaves the test case it builds for the same method.
The filter that test selection injects into a test suite accepts every suite it is asked about and applies the selection to the tests inside it instead. Iterating a data provider suite therefore still yields the repetitions of a repeated data set and the attempts of a retried one as a suite, an empty one, when the selection excluded that data set. The descriptor of the data provider suite described such a member as if it aggregated something, and the descriptors of a repeat suite and of a retry suite assert that it does: combining --repeat or #[Retry] with a --filter that selects a single data set ended the run with "An error occurred inside PHPUnit" instead of running the selected data set. With assertions disabled, the retry descriptor passed null where a test case is required and the repeat descriptor sent an empty repeat suite to the worker. A sequential run runs the selected data set and reports nothing else, because TestSuite::run() returns early for a suite the selection emptied. Skip such a member where the members of a data provider suite are described, as the collection of the work units already skips it at the level above.
When a worker dies or its result envelope fails verification, the tests of its unit that never reported a result are reported as errored, so that the event stream stays complete. The tests to report were collected by walking each member of the unit with TestSuite::tests(), which answers with every test the suite holds, not with the tests that test selection left in it. Now that a unit dispatches only the tests the selection selected, the two no longer agree: with --filter selecting one data set of a data provider method whose worker dies, the run reported the other data sets as errored as well (tests that were never dispatched and that a sequential run would not have run at all) and counted them among the tests it ran. Iterate the member instead, as the descriptors that dispatch it already do, so that the tests reported as missing a result are exactly the tests the worker was asked to produce one for.
…ored The tests of a unit whose worker died and that never reported a result are reported with ChildProcessErrored, TestErrored, and TestFinished: the events the sequential runner emits for a test whose child process ended unexpectedly, but without the TestPrepared that precedes them there, because the parent never prepared a test the worker was to run. A consumer that learns of a test only when it errors treats it as a test outside a test method and reports it in its own right: the TestDox result collector processes the test on TestErrored when no TestPrepared opened it, and again on TestFinished. With --parallel and --testdox, a crashed unit therefore listed every test whose result never arrived twice, while the progress printer, which knows about ChildProcessErrored, showed the correct number. Emit TestPrepared for such a test as well, so that the events a crashed unit reports are the events a sequential run reports for the same test.
The scheduler dispatches the units of a chunk longest-recorded-duration first, so that the longest-running unit starts while there is still work to overlap it with. What a unit is estimated to cost is the sum of the durations a previous run recorded for its tests, and the members that aggregate tests (the tests of a data provider method, the attempts of a retried test method, the repetitions of a repeated one) were summed up by walking them with TestSuite::tests(). That answers with every test the member holds, so a unit that test selection has narrowed was estimated by what it would cost without the selection, and the units of a filtered run were ordered by durations they will not spend. Iterate the member instead, as everything else that reads a member of a unit now does.
…worker TestBuilder::configure() applies four settings to a test case that a parallel worker recreates from its descriptor: process isolation, the preservation of global state, and the backup of global variables and of static properties. The test that asserts it leaves such a test case configured exactly as build() leaves the one it builds compared the preservation of global state as well, but no fixture carried metadata for it, so the comparison only ever saw the default on both sides: dropping the propagation of that setting from TestBuilder left every test of TestBuilderTest passing.
…it could clean up after itself
…they share the temporary files named after that file when code coverage is collected
…ose events go to the destination the caller provides, so that a caller can advance several PHPT tests at the same time
…in the groups its tests form, stopping where the collected results call for the run to stop, so that such a unit is not reported past the point at which a sequential run would have stopped
… standalone unit at its suite index, so that --parallel does not serialize every PHPT test when --repeat or --retry is used
6be790b to
7757ead
Compare
PHPUnit can now execute a test suite across several worker processes concurrently instead of one test after another. Parallel execution is opt-in via a new
--parallel=<n>command-line option and changes nothing when it is not used: the sequentialTextUI\TestRunnerremains the default andTextUI\ParallelTestRunneris selected only when<n> > 1.Native parallelism here is a generalization of process isolation rather than a new subsystem: a worker reconstructs and runs a unit of work and ships its outcome home in the very same serialized envelope that process isolation already uses, and the parent replays that envelope through its normal event pipeline. As a result the parent process remains the single source of truth for all output, logging, results, and code coverage, which are produced exactly as in sequential mode.
How it works
#[BeforeClass]/#[AfterClass]and intra-class ordering. ADataProviderTestSuite, and theIterativeTestSuitethat carries the repetitions of a repeated test or the attempts of a retried test, travel to the worker as atomic members of their class' unit, so their suite envelopes nest in logger output exactly as they do sequentially.<testsuite>elements of an XML configuration are run one after another, just as in sequential mode: only tests that belong to the same top-level test suite ever run concurrently.Runner\Parallel\WorkerPoolownsnpersistent workers and pulls from a dynamic work-stealing queue, so the load self-balances against stragglers. A worker signals completion through the filesystem, which the parent polls;stream_select()is not used, because it does not work on the workers' output pipes on Windows.Runner\Parallel\Schedulerdispatches units in the order of the durations recorded by the test run history, longest first, so that the longest-running work starts as early as possible instead of becoming the straggler the pool waits for. A unit with no recorded duration goes first. Only the dispatch order is affected; results are released in suite order either way.Runner\Parallel\ResultAggregatorbuffers each unit and releases it only once every preceding unit (in suite order) has been released, so the event stream, and therefore every output format, including the default progress output, is byte-for-byte what sequential mode produces. Streamed events of the unit that is next in suite order are forwarded immediately; those of later units are buffered until their turn.PHPUnit\Framework\TestCaseinstances and cannot be reconstructed in a worker (and nesting their child processes inside workers hung on Windows), soRunner\Parallel\PhptRunnerruns them side by side in the main process, each as its own child process.ProcessBudget, so--parallel <n>never executes more than<n>units at once, no matter how a chunk is composed.Running in the main process
Some tests cannot, or must not, run in a worker. These run in the main process at their correct suite position — the aggregator invokes them while releasing results, so global ordering is preserved — and their execution there is ordinary sequential execution:
#[DoNotRunInParallel]— a new attribute, valid on classes and methods, for tests that must not run alongside others (for instance because they share a machine-global resource). Such a unit runs alone: the worker pool and the PHPT runner first finish what they are executing and start nothing new until it is done.#[RunInSeparateProcess],#[RunTestsInSeparateProcesses], or global--process-isolation) — a shared worker cannot provide isolation, but the main process spawns the isolated child as usual.#[Depends]— the result they depend on is produced by a different unit and is only visible in the main process, once that unit has been released.serialize()throw) as well as resources (whichserialize()silently degrades to0).TestCasenor a PHPT test.Because a work unit is a whole test class, a single test method carrying
#[DoNotRunInParallel], requiring isolation, depending on another class, or providing non-serializable data takes its entire class out of the parallel phase.PHPT tests declare their concurrency constraints with a
--CONFLICTS--section instead, since a PHPT file cannot carry PHP attributes: while a test holding conflict keyKruns, no other test declaringKis started, and the reserved keyallmeans the test runs entirely on its own.Robustness
The worker process running X ended unexpectedly), the remaining units are redistributed across the surviving workers, and the run accounts for all of them.--stop-on-*. As soon as the collected results call for a stop, the aggregator releases nothing further and the units still executing are terminated (a grace period, then a forced kill). A PHPT test whoseFILEsection was terminated still runs itsCLEANsection, and one terminated duringSKIPIFdoes not start itsFILEsection.New public API
--parallel=<n>CLI option (also listed in--help); a value that is not a positive integer is ignored with a test runner warning.#[PHPUnit\Framework\Attributes\DoNotRunInParallel](TARGET_CLASS | TARGET_METHOD), with full metadata-layer support.--CONFLICTS--PHPT section, as supported by the PHP project'srun-tests.php: one conflict key per line,#starts a comment, blank lines are ignored, andallis reserved.PHPUNIT_WORKER_ID— the small, stable ordinal (0,1,2, ...), ideal for indexing a fixed set of pre-provisioned resources. A worker restarted after a crash keeps its ordinal.PHPUNIT_WORKER_TOKEN— a value of the form<id>_<random>that is unique across workers and across runs, for resources that must not collide with those left behind by a previous run. A restarted worker gets a fresh token.Event\TestRunner\ChildProcessReason::ParallelWorker, so that the events about child processes say why one was used.Shared with the sequential runner
Rather than duplicating the sequential runner, behavior the two must perform identically was extracted so it is expressed exactly once:
TextUI\TestRunnerLifecycle— the test runner lifecycle both runners drive.Event\Dispatcher\CollectionWindowand the facade's dispatcher selection — the window during which events are collected instead of dispatched.Framework\TestRunner\ChildProcessBootstrapplus two templates — the boot code every kind of child process shares; the configuration and source map are written once per run for all of them, so isolated processes and parallel workers cannot drift apart.Framework\TestRunner\ChildProcessResultEnvelope— the result envelope's encoding and decoding, shared by its consumers.TestRunner\ExecutionFinished, instead of every time an outermost test suite finishes.Results for PHPUnit's own test suite
Measured on a 6-core machine:
--parallel 10--testsuite unit(5813 tests)--testsuite end-to-end(1257 tests)Both suites report the same numbers of tests, failures, and skipped tests in both modes, stable across runs. The end-to-end suite benefits the most because it is dominated by child processes rather than by CPU work in the main process.
Which worker runs which test class is not deterministic; the work-stealing scheduler assigns classes by timing, but the aggregated output is identical regardless.
Notes and limitations (possible follow-ups)
#[DoNotRunInParallel].