Skip to content

perf: compute the streaming linear_pos projection once per window length - #98

Open
IsNoobgrammer wants to merge 1 commit into
mudler:masterfrom
latentsig:perf/streaming-pos-cache
Open

IsNoobgrammer wants to merge 1 commit into
mudler:masterfrom
latentsig:perf/streaming-pos-cache

Conversation

@IsNoobgrammer

Copy link
Copy Markdown

Problem

StreamingEncoder::step() computes linear_pos(pos_emb) in every conformer layer for every chunk. The input pos_emb depends only on the attention window length (cache_len + Tc, i.e. pos_len = 2 * (Tc + cache_len) - 1), and the window length is the same for every chunk after the first, so each layer recomputes the same matrix on every chunk. On nemotron-3.5-asr-streaming-0.6b (24 layers, d_model 1024, about 145 positions) that is one 1024 x 1024 by 1024 x 145 matmul per layer per chunk, a large share of the work of a small chunk.

Fix

StreamingEncoder computes the projection once per window length for the session, with the same ggml ops as before (mul_mat with the layer's linear_pos.weight, then the same reshape, permute and cont into heads), and keeps it as [dk, pos_len, H] per layer in a member map. Each chunk graph takes it as a graph input instead of computing it. build_stream_layer receives the precomputed heads in place of pe, and the per-chunk pos_emb graph input is removed because nothing else used it.

The cache is a member of the encoder, so it is freed with the session, needs no lock, and cannot be reused by a different model. It adds one projection pass per window length per session (usually two or three: the first chunk, the steady state and a short last chunk) and about 14 MB per window length for Nemotron 3.5 at f32. The ops and their inputs are unchanged, so the outputs are unchanged.

Measurement

Build: Windows 11, llvm-mingw 20260922 (clang, UCRT), Release, -DGGML_NATIVE=OFF -DGGML_AVX2=ON -DGGML_FMA=ON -DGGML_F16C=ON, CPU only. Machine: Intel Core i5-12450H (4 performance and 4 efficiency cores, 12 threads), a laptop with normal desktop load, so these are not quiet-machine numbers. Input: a 120 s, 16 kHz meeting recording (3 speakers, accented English). Command: parakeet-cli transcribe --model <m> --input clip.wav --stream --threads 8, wall time of the whole command, master and this branch interleaved, 3 runs each.

Model master (s) this PR (s) median
nemotron-3.5-asr-streaming-0.6b q8_0 62.4, 63.8, 74.2 60.6, 55.3, 52.2 63.8 -> 55.3 (-13%)
nemotron-3.5-asr-streaming-0.6b q4_k 60.4, 59.5, 48.3 39.5, 45.2, 45.4 59.5 -> 45.2 (-24%)

The transcripts are byte-identical: all 6 runs of each model, master and this PR, print the same stdout (same md5).

In a host application that feeds the same streaming C-API in 100 ms blocks (4 threads, same machine), the median per-chunk encoder time went from 174 to 152 ms (q4_k) and from 236 to 203 ms (q8_0) on the same recording, and from 88 to 61 ms for a 114M hybrid streaming model at 240 ms look-ahead. Transcripts were byte-identical there as well.

Tests

  • ctest -LE model: 44 of 52 pass on master and 44 of 52 on this branch, with the same 8 failing on both. Those 8 fail on Windows independently of this change: test_backend_device does not compile (setenv/unsetenv), test_audio_io crashes, the three server tests need parakeet-server, which does not build here, and the three Python checks need the converter venv.
  • Streaming output compared byte for byte against master as described above.

Limits

  • The NeMo parity tests for streaming (test_streaming_encoder, test_streaming_decode, test_capi_stream) were not run: I do not have the NeMo baselines on this machine. The change keeps the same ops on the same inputs, and master and this branch give identical transcripts, but the parity gate itself has not been run.
  • Only CPU on Windows/x86 was measured; no GPU, Metal, macOS or Linux runs.
  • Only Nemotron 3.5 and a 114M hybrid streaming model were tried, not parakeet_realtime_eou_120m-v1.

🤖 Generated with Claude Code

The relative-position projection linear_pos(pos_emb) depends only on the attention
window length (cache_len + Tc), not on the audio, but StreamingEncoder::step()
recomputed it in every layer for every chunk. On Nemotron 3.5 streaming (24 layers,
d_model 1024, about 145 positions) that is a [1024 x 1024] x [1024 x 145] matmul per
layer per chunk, a large share of a small chunk's work.

StreamingEncoder now computes it once per window length for the session, already
split into heads ([dk, pos_len, H] per layer), and feeds it to the chunk graph as an
input. The cache is a member of the encoder, so it is freed with the session and needs
no lock. The ops and inputs are the same as before, so the outputs are unchanged.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]

@localai-org-maint-bot localai-org-maint-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mudler @richiejp Source review at b802edb: the per-encoder cache key follows the existing positional-encoding window, and the projection/head-layout operations match the removed graph nodes. I would keep merge pending streaming parity validation.

The reported transcript comparisons are useful, but this moves the projection into separate graph executions and adds retained per-window state. Please run the existing test_streaming_encoder, test_streaming_decode, and test_capi_stream gates with the required assets, covering the first, steady-state, and short final windows. A regression test should also exercise reuse of a cached window and reset/reuse of the encoder. The PR currently changes only the implementation and has no automated coverage of the new cache behavior.

I could not run native/model tests on this host because the C++ toolchain and model baselines are unavailable; this is not an independent performance or parity confirmation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants