cleanup: follow-ups from the station series reviews (#451-#464), part of #465 - #466
Conversation
txs_parse_long carried two comment blocks back to back; the first one documents txs_retry_limit_env (DEVOURER_TX_RETRY_LIMIT, 1/0/-1 return), which had no comment of its own. Move it there. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
After a successful StartBeacon, the Jaguar2/3 StopBeacon returns false only when the disable itself was refused - the "nothing active" early return cannot fire, and _bcn_hw_touched makes a retry re-run the disable. TxBeaconGuard treated every false as done. Retry on false while armed; a clean false after a refused StartBeacon still ends the loop. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
docs/jaguar3-tx-ring.md: the dated, attributed bench paragraph becomes plain measurement inside the 8812CU paragraph it qualifies. - On USB2 at DEVOURER_TX_GAP_US=0, bit 13 (BIT_PAYLOAD_OVF_8822C, 0x2000) latched while 8051/8051 frames completed. - It read 0 at the default gap. - One later USB2 gap-0 run did not reproduce the latch, so the bit can latch at max duty on USB2. - A second bench reproduced the defect and the fix clearing it. Git carries the provenance. src/IRtlRadio.h, the GetTxDmaStatus contract, said bit 13 "latches" under host-side max-duty backpressure. It now says "can latch", noting that a later such run did not latch it. docs/logging.md defers to that declaration. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
mt_mac_start() writes MT_MAC_SYS_CTRL_ENABLE_TX before its WPDMA-idle poll and returns -1 without clearing it, so a gate that returned straight out of a failed start could leave TX enabled. Every such failure path now calls mt_mac_stop() before returning, after the gate's existing cleanup and in its normal teardown order: rx_teardown() first where a ring was draining EP 4 (gate_ap's bare rx_stop becomes rx_teardown), mt_mac_rx_disable() first in gate_rx as its normal exit does. The failed start never reaches mt_async_start() in gate_txs, and gate_caps' async drainer is torn down by rx_teardown() as on its normal exit. gate_tsfwrite already did this. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
…EXT read Refs OpenIPC#461: on a second MT7612U the ~6 fps txs arms read 199/200 (59/60) while the fast arms settle 200/200, which leaves the uplink harness INCONCLUSIVE. The slow rate is a symptom, not the cause: those are the arms that lost their first entry, after which every per-frame wait (entries >= n) times out. The status read is EXT, then the main word: two USB transfers. When the FIFO is empty at the EXT read and an entry is filed before the main read, the pop is paired with the stale EXT word of the previously popped entry. - Inside an arm that is the same pktid, and harmless. - On an arm's first entry it is the previous arm's pktid, so the entry was counted late (foreign on the session's first arm). That is the "one-step status lag" of docs/mt7612u-tx-retry.md. Its own table rules out the alternative, status posted only on the next TX: at limit 0, receiver ON, arm f lags with L1 after arm e settled 40/40 and owed nothing, and arm g after it shows no late entry. The late entry belongs to the lagging arm, not the one after it. A race on the poll timing also explains why it varies from pass to pass and from host to host. txs_drain now claims that one entry back when it can only be ours, which takes all of the following: - it is the arm's first popped entry, after the arm has submitted a frame; - it carries the pktid of the last arm that SENT a frame, and that arm settled with no entry owed; - or, on the session's first arm, it carries any pktid. An arm whose every submit failed popped nothing and leaves that reference untouched. The claimed entry counts toward entries, and toward success when its SUCCESS bit is set (the main word is fresh). It stays out of the retry columns (its retry count is the stale word's), and is reported as "stale-EXT entries claimed". After an UNSETTLED arm, a previous-pktid entry is still reported as late. The claim recovers at most one entry per arm. A larger deficit, like the uplink harness's multi-entry loss, is not this race. Hardware: in the form keyed to the previous arm, the author's unit settled 16/16 arms at limits 5 and 0 (issue OpenIPC#461). Keying it to the last arm that sent has not been run on hardware. No headless test: txs_drain reads the FIFO through mt_rr_chk on a live device inside the bring-up tool, which no selftest links. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
With DEVOURER_DRAIN_BULK_IN or DEVOURER_POLL_INTR_IN set, bulk_in_thread and intr_in_thread were joined only on the InitWrite-exception path. The normal end of main and the early returns after a PCIe-open or no-driver failure destroyed them still joinable, so the process ended in std::terminate. A small scope guard after the two threads now clears both run flags and joins them on every return. The normal path joins explicitly before session.close(), because both threads poll the handle it releases. The InitWrite catch block's open-coded join becomes the guard. A fork() child of the legacy DEVOURER_TX_WITH_RX path holds copies of the thread objects but not the threads, so it skips the join and keeps its pre-existing exit behaviour. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
…d span MIN_SUBMITTED defaulted to a quarter of SECS * 1e6 / GAP_US, but the SECS window also holds the transmitter's bring-up. On an 8812BU peer txdemo.init_write took 6.2 s and 8.7 s of the 10 s window, so the peer aired for only 1.3-3.7 s and submitted 283-327 frames against the 500 floor, with live=1 and max_gap under 35 ms. Every DOWN arm came back INCONCLUSIVE: a healthy transmitter scored as a stall. summarize now prints aired_ms, the span from the arm's first tx.report to its final tx.stats. Those are the first two timestamps in one timebase; txdemo.first_tx_submit counts ms from a different epoch. It also prints min_submitted, a quarter of the GAP_US rate over that span, and usable() holds each arm to its own floor. The span ends at the final tx.stats, not the last report, so a transmitter that slows or stops after MIN_REPORTS still owes the whole span. A set MIN_SUBMITTED stays a fixed floor. Checked on synthetic JSONL fixtures run through the script's own summarize and usable: - slow bring-up, 327 frames over 2.0 s: usable, floor 100 (refused under the old fixed 500); - rate stall, 110 frames over 10 s with sparse reports under MAX_GAP_MS: refused, floor 500; - hard stall: refused, live=0; - healthy full window: usable. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
The AP guard for the UP half ran only after the whole DOWN half. A run with the AP already off the bus spent 60 s on DOWN and then refused, and a hub, an adapter with no wireless netdev, one carrying a default route, or a phy without AP mode was refused just as late. The UP half's guard is now a function, ap_guard. It runs right after the lib is sourced, for the UP or both halves, and again at the start of UP, since the adapter can move during DOWN. The preflight only reads; nothing is written before the lock. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
Stop() cleared a held station arm but left _station_ready set, so a SetStationIdentity issued after Stop() was accepted: the same hole the gate closed for _brought_up. That arm went to a torn-down chip on Jaguar1/3, or to a still-powered chip on Jaguar2, which Stop() does not power down. All three Stop()s now clear _station_ready under the station lock, in the block that clears the arm (_port0_mu on Jaguar1, _reg_mu on Jaguar2/3). Nothing else reads the flag except the SetStationIdentity gate, and every bring-up already clears it on entry and commits it at the end, so a re-Init after Stop() arms again as before. StationArm.h's LIFETIME note says so. No selftest: _station_ready only becomes true at the end of a real Init/InitWrite, which no headless test can run. The existing station_arm selftest covers StationArm, not the device classes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
…live process - Interrupt. `trap cleanup EXIT INT TERM` ran cleanup on Ctrl-C and then carried on into the next cell against a re-enumerated adapter. The traps now follow the station harnesses: `cleanup; exit 130` on INT/TERM, cleanup on EXIT. A CLEANED flag, set only on the interrupt path, spares the EXIT pass a second re-enumeration. It is not a once-only guard, because CELLS=all calls cleanup between cells. - Hand-back. cleanup re-enumerated AP_SYSFS right after sending TERM, with the AP demo possibly still inside its de-init. reap now polls its children for up to 10 s with the station lib's zombie-aware sta_pid_alive (the lib is sourced for that alone). Anything still alive is KILLed, and cleanup then skips the re-enumeration, saying so. - PHASE 2. cell_stop waited up to 60 s for "PHASE 2" and then scored "beacon gone" whether or not it appeared. A bstop_onair that ended before StopBeacon - its teardown silencing the beacon - read as a StopBeacon PASS. Without PHASE 2 the cell now FAILs, naming that. - Ping loss. The open cell's failure message extracted the loss with `[0-9]+% packet loss`, which cut ping's "66.6667% packet loss" to "6667% packet loss". It now takes `[0-9.]+%`. The pass test is unchanged. - Dead loop counters (`local i`) dropped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
cleanup ignored sta_pid_kill's result for the DUT and re-enumerated it regardless, including while the bring-up process was still inside its de-init. The peer already had this rule; the hand-back is now gated the same way, as in realtek_station_onair.sh. That needed a CLEANED guard. sta_pid_kill forgets the PID on its first call, so the EXIT pass that follows an INT would have read "gone" and handed back the adapter the first pass had just refused to. The old comment called that second pass harmless. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
… runs The DUT's `mt7612uprobe txs` ran under a bare `wait`, so a wedged gate held the run and both adapters forever. It now runs under `timeout -s INT -k 10` with a bound derived from the gate's own worst case: 16 gate arms; per frame, the status wait (b + 50 ms) plus the settle (b), where b is gate_txs's frame_budget_ms for RETRY_LIMIT; and about 7 s of fixed cost per arm. On top of that, half again plus 2 min for bring-up: about 9 min at FRAMES=60 and 19 at FRAMES=200, against the ~9 min measured for a 200-frame arm. An overrun ABORTs the arm. cleanup re-enumerated the DUT whether or not sta_pid_kill saw its process exit. The hand-back is now gated on that, as for the peer, with a CLEANED guard so the EXIT pass after an INT cannot undo the refusal. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
- hostapd. AP_REENUM was set only after the hostapd launch and its 10 s
wait loop, and cleanup killed hostapd only when the flag was set. A
Ctrl-C in that window left hostapd running and the AP in AP mode, with
the lock released. The flag is now set once the AP guard has passed,
before the interface is touched, and hostapd is killed unconditionally,
as in realtek_station_onair.sh.
- DUT hand-back. It ignored sta_pid_kill's result and re-enumerated the
DUT with the BSSID gate still running. It is now gated on the gate
having exited.
- Bounded gate. `wait "$sta_gate"` was unbounded. The gate now runs under
`timeout -s INT -k 10` for six arms of SECS plus 183 s. An overrun is
no measurement (2), and stays 2 even when the injector never started.
- Monitor vif. It was probed only after the contract and probe-response
gates. It is now also probed (create, then delete) right after the AP
guard.
- Exit status. Everything collapsed into 0/1. It is now 0 pass, 1 a gate
failed, 2 INCONCLUSIVE, 3 interrupted, combined as: 1 if any gate failed,
else 3, else 2, else 0.
- Rig refusals exit 2: a hostapd that never brought the AP up, and no
monitor vif. realtek_station_onair scores the same hostapd failure
as 2.
- The INT/TERM trap exits 3, as documented.
- Nothing in the tree reads this exit code.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
cell_end returned as soon as the previous hostapd exited, and ap_up launched the next one at once. On one rig, one restart in four started 30 ms after AP-DISABLED and died with "Could not read interface <if> flags: No such device" / "nl80211 driver initialization failed", which the cell scored INCONCLUSIVE. Before each launch, ap_up now spends up to 10 s, inside the netns, on these steps: - it waits for the netdev to be present; - once present, if it reports a type other than managed, it takes the interface down and sets type managed, retrying until the type reads managed; - it brings the interface up. Forcing the type matters because a driver may leave the vif in AP type after hostapd exits, and hostapd then fails with "Match already configured". realtek_station_onair.sh and mt7612u_sta_identity.sh force the type before every hostapd for the same reason. Nothing here is driver-specific, so it holds for an mt76x2u AP as well as rtw88. If the window runs out, the cell is refused as before (INCONCLUSIVE). The reason is written into the hostapd log that message points at, and the interface is left up. No hostapd retry: the wait removes the race. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
usb_thread and rx_thread were joined only at the end of main. The early returns after them - the three `return 1`s on the TX path - destroyed them still joinable, so the process ended in std::terminate. It is the same class as the IN drainers fixed earlier on this branch. A scope guard, IoThreadsJoin, declared right after rx_thread, runs the normal teardown's sequence on every exit: 1. StopRxLoop. 2. Join the RX thread. 3. Set g_devourer_should_stop. 4. Join the event pump. The normal path calls it explicitly where the joins were, before Stop(), so the order against Stop(), the drainers' join and session.close() is unchanged. On an early return it runs before the session releases the device. Both threads start after the DEVOURER_TX_WITH_RX fork, whose child returns first, so no fork-child skip is needed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
|
@josephnef the cleanup PR I mentioned on #464: your non-blocking notes from #451–#464, plus part of #465's harness items, in 15 small commits on 927cfc7. It's grouped by origin in the description.
Hardware, on this head (same rig as #463/#464):
Every adapter was handed back. Not in this PR:
|
PR Summary by QodoFix TX status attribution and harden on-air teardown
AI Description
Diagram
High-Level Assessment
Files changed (15)
|
Code Review by Qodo
1.
|
The floor's aired_ms started at the arm's FIRST tx.report. A transmitter that started, stalled, and then burst 50 reports in its last second met MIN_REPORTS and a ~50-frame floor, where the old fixed 500 would have refused it (qodo, PR OpenIPC#466). txdemo's tx.frame now carries t, the same host-monotonic timebase as tx.report and the final tx.stats. The first tx.frame follows the first submit, so a harness can date it. docs/logging.md lists the field. Existing consumers match on "n" and "rc" or parse JSON, so the extra trailing field changes nothing for them. summarize uses that first-submit time in two places: - aired_ms now runs from the first submit to the final tx.stats. - Liveness gains lead_ms, the silence from the first submit to the first report. live=0 when it exceeds MAX_GAP_MS, or when no tx.frame carries t. The bring-up stays outside the window either way, because the first submit follows InitWrite. Synthetic JSONL through the script's own summarize/usable: | Fixture | Result | |---|---| | slow bring-up, 327 frames in 2.0 s | usable, floor 101 | | start, stall, 60 reports in the last second | refused: live=0 on lead_ms 8800, and 60 < floor 490 | | rate stall | refused, floor 500 | | hard stall | refused, live=0 | | healthy full window | usable | | no tx.frame t | refused, live=0 | Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
When reap() had to KILL a child, cleanup rightly skipped the authorized toggle. Under CELLS=all, though, that cleanup also runs between cells, and the run carried on into the next cell on an adapter that was never reset and whose autonomous beacon could still be airing (qodo, PR OpenIPC#466). cleanup now returns 1 when it could not reset the AP. In the CELLS=all sequence, the cells after such a cleanup are not run; they are listed as NOT RUN, and the run exits 2 (INCONCLUSIVE) unless a check already failed (1). It never scores those cells. A single-cell run is unchanged. The exit status is otherwise as before: 0, or 1 on a failed check. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
The legacy DEVOURER_TX_WITH_RX fork child returned from main holding copies of the parent's objects, including the IN drainers' std::threads. Those are joinable copies of threads that do not exist in the child, so ~std::thread terminated it, and the DrainerJoin guard's in_fork_child skip could not prevent that (qodo, PR OpenIPC#466). The child now follows the standard post-fork rule: - It runs Init inside a try, so an exception does not unwind it either. - It flushes stdio. - It leaves through std::_Exit(1), so no destructor runs in the child. That is the child's only exit path. The in_fork_child flag is gone, since the child never reaches the guard's destructor. On MSVC fork() is the (0) stub, which makes that branch the only process, so it keeps its normal return and teardown. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3
josephnef
left a comment
There was a problem hiding this comment.
Reviewed at 18f0a20. Read every hunk; built the branch with -DDEVOURER_MT7612U=ON (ctest 81/81 run, 2 skipped without reference/); ran two harnesses on hardware, plus a master A/B of one of them. Approving - the notes below are non-blocking; the first one is worth a small fix in the squash if you agree.
Code
The scope guards in txdemo hold: drainers is declared after session and before the device, io_threads after both and before tx_beacon/sta_clear_join, so every early return joins while the handle, device and context are alive, and the normal path's explicit join()s leave the destructors as no-ops. The fork child's _Exit guard matches the fork() stub guard exactly. The TxBeaconGuard retry is right for the Jaguar2/3 contract (false after a true arm = refused disable). The three Stop()s write _station_ready under the same lock the SetStationIdentity gate reads it. mt_mac_stop after a failed mt_mac_start is safe on both failure paths (the RING refusal writes nothing; the WPDMA path has set ENABLE_TX). All in-tree tx.frame consumers match on "ev"/"n"/"rc", so the trailing t is harmless.
Hardware (this bench, ch6, near field)
realtek_station_onair.sh, DUT 8812CU at 3-2.4, peer 8812BU (CF-924AC) at 4-2.3.3, AP = MT7612U on mt76x2u at 10-1.3: 5 passed, 0 failed, 0 inconclusive, rc 0. The new floor did exactly what it is for: the 8812BU peer spent ~8.5 s of the 10 s window in InitWrite and aired 327 frames over aired_ms=1475; min_submitted came out 73 (old fixed floor: 500, refused). Every arm live=1, lead_ms 1-63.
One oddity, harmless to the verdict: arm C read tail_ms=-6 - the final tx.stats t landed 6 ms before the last tx.report t. Presumably the final stats are stamped before the last C2H decodes. live is unaffected (negative ≤ MAX_GAP), but a reader of the summary line will wonder; clamping at 0 or noting it in the summarize comment would do.
mt7612u_sta_uplink.sh, DUT MT7612U at 10-1.3, responder 8812BU, default FRAMES=60 / RETRY_LIMIT=15, run on this branch and then on master (927cfc7) with the same adapters:
- The multi-entry loss you flag as open is reproduced here and is much worse than your 48/53: the unicast-to-peer arms (b-e) land 0-11/60 own entries on BOTH passes, with ~10 entries per arm carrying the previous arm's pktid, and that pattern is identical on master. So: pre-existing, not this PR, and your "a deficit larger than one entry is not this race" holds on a second unit. Both runs INCONCLUSIVE/FAIL for that reason; both adapters handed back cleanly each time.
- The stale-EXT claim fired where it should: the session's first arm reads
60/60, 1 stale-EXT entries claimedon this branch where master reads59/60 UNSETTLED, 1 foreign. That also hardware-runs the "last arm that sent" form (no arm had every submit fail, so it coincides with the previous-arm form here). - One over-claim:
h broadcast, wcid1 control(receiver OFF, arm A) read61/60, 1 stale-EXT claimed, 11 foreign - the arm then received all 60 own entries, so the claimed one was not ours (or the MAC filed a duplicate). It is printed rather than clipped, as designed, but it inflatessuccessby one. Cheap self-correction: track own-pktid entries separately and, at arm end, if own == n and stale_ext == 1, move the claimed entry back tolate_prev. That makes the one case that can prove the claim wrong correct itself. - The
stale_settledgate means one UNSETTLED arm disables the claim for every following arm until one settles on its own: in run B that was a chain of seven consecutive 59/60 arms, each with "1 late entry from the previous arm". Conservative and correct by the documented rule, but worth one sentence indocs/mt7612u-tx-retry.md: on a unit where arms do not settle, the claim buys almost nothing.
Harness nits
mt7612u_sta_identity.sh: an overrun gate setsr_bss=2, but the verdictcasehas only0,3,*, so the operator reads "did not pass" for what the exit code calls INCONCLUSIVE. Add a2)branch.mt7612u_sta_uplink.sh: the newdut_boundis per harness arm (552 s at FRAMES=60, 1128 s at 200) while the header's 25-minute figure is for both arms, so the margin for the slower un-ACKed arm at FRAMES=200 is thin (~15% if arm B is ~16 min). Fine at the default; maybe size it from the measured figure rather than the ladder.mt7612u_ap_onair.sh: on the KILL pathreapdropsKIDSwithoutwaiting the killed children, so they stay zombies for the rest of the run. Harmless, awaitis free.- A #465-class item for the follow-up list: on a host whose MediaTek firmware is zstd-compressed (
/lib/firmware/mediatek/*.bin.zst, this one),sta_fw_linksucceeds, the probe logscannot open firmware/mt7662_rom_patch.bin, and the harness scores it asABORTED the DUT did not confirm retry limit 15:with an empty reason. A preflight that checks the two blobs are readable through the link would turn a dead rig into a refusal in seconds.
Follow-ups and fixes collected from merged PRs: the maintainer's non-blocking review notes, a probable cause for #461's one-entry loss, part of the #465 harness defect sweep, and two txdemo teardown fixes we found ourselves. Fifteen commits, one per item or per harness, so any can be dropped or reordered on review, plus three follow-ups from the qodo review of this PR.
Refs #461, Refs #465 (partial).
What changed
Maintainer review notes
0fee106txs_parse_longcarried two comment blocks back to back. The first one documentstxs_retry_limit_env, and it now sits there.4d862ff_bcn_hw_touchedmakes a retry re-run it. A clean false after a refused StartBeacon still ends the loop.294a671docs/jaguar3-tx-ring.mdbecomes plain measurement.IRtlRadio::GetTxDmaStatusnow says bit 13 (BIT_PAYLOAD_OVF_8822C) can latch at max duty on USB2, not that it does.ddd8467src/mt7612u/tools/bringup.cppnow callmt_mac_stop()whenmt_mac_start()fails.mt_mac_startsetsENABLE_TXbefore its WPDMA poll, so a failed start could leave TX enabled. Each gate keeps its own teardown order (rx_teardown()first where a ring was draining EP 4).2cc5330Stop()clears_station_readyin the block that clears the arm, under the station lock. ASetStationIdentityafterStop()is refused until the next bring-up.907a5b5tests/realtek_station_onair.sh's submission floor is scaled to the span the arm actually aired, not to SECS, which also holds the peer's bring-up. The span runs from the first submit to the finaltx.stats(fcdec63, below).da8cd6d#461: the txs gate's lost status entry
2e7a3degate_txsclaims an arm's first status entry back when a staleMT_TX_STAT_FIFO_EXTword mislabelled it with the previous arm's pktid. Details under Why.#465 (partial): the harness defect sweep
53477e0mt7612u_ap_onair.shAP_SYSFS, cleanup waits for its children to exit (10 s bound, zombie-aware via the lib'ssta_pid_alive), and skips the re-enumeration if it had to KILL one. The stop cell FAILs when PHASE 2 never appeared, rather than scoring "beacon gone". The ping-loss figure is quoted whole ("66.6667%", not "6667%").05bdaf7mt7612u_sta_autoack.sh8a2ab14mt7612u_sta_uplink.shmt7612uprobe txsruns undertimeout -s INT -k 10, bounded by the gate's own worst case.bf9c0ddmt7612u_sta_identity.shAP_REENUMis set before hostapd starts, and hostapd is killed unconditionally. The BSSID gate is bounded. The monitor vif is probed in the preflight. Exit codes are 0/1/2/3, with rig refusals as 2 and the INT/TERM trap as 3.d473027mt7612u_sta_onair.shap_upwaits (10 s bound, inside the netns) for the AP netdev, forces it back totype managed, and brings it up.Not in this PR, from #465: the generic Realtek DUT take, the reconnect cell and the CCMP soak. Those come in a separate PR.
Our own finding: txdemo thread joins
64ebce3examples/tx/main.cppjoins the optional IN drainers (DEVOURER_DRAIN_BULK_IN,DEVOURER_POLL_INTR_IN) on every exit.6b31d5fIn both cases an early return used to destroy a still-joinable
std::thread, which callsstd::terminate. Small scope guards now join the threads in the normal teardown's order, beforesession.close(). The legacy fork child is covered by18f0a20, below.qodo review of this PR
fcdec63realtek_station_onair.sh: the aired span started at the firsttx.report, so a transmitter that started, stalled and then burst 50 reports in its last second passed both floors. txdemo'stx.framenow carriest(thetx.reporttimebase;docs/logging.md). The span runs from the first submit, and liveness also refuses a first report more thanMAX_GAP_MSafter it (lead_ms).92dfa1dmt7612u_ap_onair.sh: if a between-cell cleanup could not reset the AP (a process outlived TERM), the remaining cells are recorded as NOT RUN and the run exits 2, instead of being scored on an unreset adapter that may still be beaconing.18f0a20DEVOURER_TX_WITH_RXfork child leaves throughstd::_Exitafter flushing stdio, withInitin a try. No destructor runs in it, so its copies of the IN-drainer threads no longer terminate it.Why
docs/mt7612u-tx-retry.md's own table rules out the alternative, "status posted only on the next TX". At limit 0, receiver ON, arm f lags with its late entry in its own row, after arm e settled and owed nothing. Arm g after it shows no late entry.std::threaddestructor terminates the process, so an otherwise clean early exit aborted.Measured
gate_txswith the stale-EXT claim settled 16/16 arms at retry limits 5 and 0 (recorded on Follow-ups left open by the station identity seam (#460) #461).tests/mt7612u_sta_uplink.sh) is INCONCLUSIVE both on this branch (48/53 status entries) and on master (57/60). It predates the branch.907a5b5), checked on synthetic JSONL through the script's ownsummarize/usable:tx.framewitht(an older txdemo)What it can't do
Follow-ups
Verification
ctest: 83/83, both plain and under-DDEVOURER_SANITIZE=address+undefined, built with-DDEVOURER_MT7612U=ON -DDEVOURER_REQUIRE_STA_CRYPTO_TESTS=ON. The two reference-submodule tests are skipped withoutreference/.make -C src/mt7612u checkpasses.bash -n(andsh -nfor the/bin/shscripts) is clean on every touched script.shellcheck -xis clean on all of them excepttests/mt7612u_ap_onair.sh, whose remaining SC2046/SC2015/SC2012 notes predate this branch.The three qodo follow-ups (
fcdec63,92dfa1d,18f0a20), re-run on 18f0a20:realtek_station_onair5 passed, 0 failed, 0 inconclusive, both ways, under the floor that now starts at the first submit.mt7612u_ap_onair14/14. Every adapter was handed back.The same harnesses also ran on the pre-squash head, with the same results.
mt7612u_ap_onairshowed one flaky open-cell ping run there, which passed on three repeats.mt7612u_sta_uplinkwas INCONCLUSIVE there and on master alike. That is the multi-entry loss under Follow-ups left open by the station identity seam (#460) #461, which this PR doesn't address.Review record: qodo-gate flagged three threads on 6b31d5f. All three held, and are fixed in
fcdec63,92dfa1dand18f0a20.Blast radius
Stop()s only. Each clears_station_ready, and only theSetStationIdentitygate reads it. Every bring-up already clears it on entry and sets it again at the end.src/mt7612u/tools/(the bring-up tool),examples/tx/(txdemo),tests/anddocs/.tx.framegains a trailingt. Every in-tree consumer either matches on"n"/"rc"or parses the JSON, so none of them changes.src/IRtlRadio.handsrc/StationArm.h: comment-only.🤖 Generated with Claude Code
https://claude.ai/code/session_01VNC8xhn1rNCi5t6uLvE6M3