Repository navigation
perf: compute the streaming linear_pos projection once per window length - #98
IsNoobgrammer wants to merge 1 commit into
Conversation
The relative-position projection linear_pos(pos_emb) depends only on the attention window length (cache_len + Tc), not on the audio, but StreamingEncoder::step() recomputed it in every layer for every chunk. On Nemotron 3.5 streaming (24 layers, d_model 1024, about 145 positions) that is a [1024 x 1024] x [1024 x 145] matmul per layer per chunk, a large share of a small chunk's work. StreamingEncoder now computes it once per window length for the session, already split into heads ([dk, pos_len, H] per layer), and feeds it to the chunk graph as an input. The cache is a member of the encoder, so it is freed with the session and needs no lock. The ops and inputs are the same as before, so the outputs are unchanged. Assisted-by: Claude:claude-opus-5-5 [Claude Code]
localai-org-maint-bot
left a comment
There was a problem hiding this comment.
@mudler @richiejp Source review at b802edb: the per-encoder cache key follows the existing positional-encoding window, and the projection/head-layout operations match the removed graph nodes. I would keep merge pending streaming parity validation.
The reported transcript comparisons are useful, but this moves the projection into separate graph executions and adds retained per-window state. Please run the existing test_streaming_encoder, test_streaming_decode, and test_capi_stream gates with the required assets, covering the first, steady-state, and short final windows. A regression test should also exercise reuse of a cached window and reset/reuse of the encoder. The PR currently changes only the implementation and has no automated coverage of the new cache behavior.
I could not run native/model tests on this host because the C++ toolchain and model baselines are unavailable; this is not an independent performance or parity confirmation.
Problem
StreamingEncoder::step()computeslinear_pos(pos_emb)in every conformer layer for every chunk. The inputpos_embdepends only on the attention window length (cache_len + Tc, i.e.pos_len = 2 * (Tc + cache_len) - 1), and the window length is the same for every chunk after the first, so each layer recomputes the same matrix on every chunk. Onnemotron-3.5-asr-streaming-0.6b(24 layers,d_model1024, about 145 positions) that is one 1024 x 1024 by 1024 x 145 matmul per layer per chunk, a large share of the work of a small chunk.Fix
StreamingEncodercomputes the projection once per window length for the session, with the same ggml ops as before (mul_matwith the layer'slinear_pos.weight, then the same reshape, permute and cont into heads), and keeps it as[dk, pos_len, H]per layer in a member map. Each chunk graph takes it as a graph input instead of computing it.build_stream_layerreceives the precomputed heads in place ofpe, and the per-chunkpos_embgraph input is removed because nothing else used it.The cache is a member of the encoder, so it is freed with the session, needs no lock, and cannot be reused by a different model. It adds one projection pass per window length per session (usually two or three: the first chunk, the steady state and a short last chunk) and about 14 MB per window length for Nemotron 3.5 at f32. The ops and their inputs are unchanged, so the outputs are unchanged.
Measurement
Build: Windows 11, llvm-mingw 20260922 (clang, UCRT), Release,
-DGGML_NATIVE=OFF -DGGML_AVX2=ON -DGGML_FMA=ON -DGGML_F16C=ON, CPU only. Machine: Intel Core i5-12450H (4 performance and 4 efficiency cores, 12 threads), a laptop with normal desktop load, so these are not quiet-machine numbers. Input: a 120 s, 16 kHz meeting recording (3 speakers, accented English). Command:parakeet-cli transcribe --model <m> --input clip.wav --stream --threads 8, wall time of the whole command, master and this branch interleaved, 3 runs each.The transcripts are byte-identical: all 6 runs of each model, master and this PR, print the same stdout (same md5).
In a host application that feeds the same streaming C-API in 100 ms blocks (4 threads, same machine), the median per-chunk encoder time went from 174 to 152 ms (q4_k) and from 236 to 203 ms (q8_0) on the same recording, and from 88 to 61 ms for a 114M hybrid streaming model at 240 ms look-ahead. Transcripts were byte-identical there as well.
Tests
ctest -LE model: 44 of 52 pass on master and 44 of 52 on this branch, with the same 8 failing on both. Those 8 fail on Windows independently of this change:test_backend_devicedoes not compile (setenv/unsetenv),test_audio_iocrashes, the three server tests needparakeet-server, which does not build here, and the three Python checks need the converter venv.Limits
test_streaming_encoder,test_streaming_decode,test_capi_stream) were not run: I do not have the NeMo baselines on this machine. The change keeps the same ops on the same inputs, and master and this branch give identical transcripts, but the parity gate itself has not been run.parakeet_realtime_eou_120m-v1.🤖 Generated with Claude Code