Skip to content

Benchmarks

H.P. Gansevoort edited this page Sep 12, 2026 · 15 revisions

Benchmarks

The numbers below were measured on an AMD Ryzen 9 3950X (16C/32T, 3.49 GHz) running Windows 11 24H2, with .NET 10.0.6, RyuJIT AVX2, a Release build and BenchmarkDotNet 0.14 / 0.15. The full per-suite tables and their history are in BENCHMARKS.md.

Headline (current)

Scenario v1.9.0 era now Why it moved
SVG render a 1 000-point line 66 µs 33.4 µs render-path work v1.10–v1.13
SVG render with LTTB downsample (100 000 pts) 1.78 ms 0.81 ms same
JSON round-trip (Figure ⇄ JSON) 40 µs 30.8 µs per-series FromSeriesDto dispatch
PNG export via SkiaSharp (simple) 27 ms 12.9 ms typeface cache + SKFontStyle leak fix
PDF export (simple / complex) — 26.2 / 29.3 ms first published PDF numbers
SIMD Vec.Sum / Vec.Mean on 100 000 pts — 19 µs, 0 B
TransformBatch (AVX) on 100 000 pts — 208 µs

Streaming

Benchmark Result
RingBuffer.Append 40M ops/sec (single-writer, 100K capacity)
RingBuffer.ToArray (10K snapshot) 190K snapshots/sec
StreamingLine: 100K appends + snapshot 13 ms total
Scenario Buffer Append rate Render FPS Memory
Dashboard 1K 1/sec 1 16 KB
Telemetry 10K 100/sec 30 640 KB
Oscilloscope 100K 10K/sec 60 800 KB
Trading (OHLC) 5K 10/sec 10 160 KB

The ring buffer is never the bottleneck; rendering is. At 30 fps the snapshot takes under a microsecond, so each frame still has about 32 ms left.

Geographic projections (100K forward projections)

Projection ops/sec Projection ops/sec
Sinusoidal 83M Stereographic 24M
NaturalEarth 77M EqualEarth 19M
Robinson 33M Orthographic 17M
AlbersEqualArea 29M PlateCarree 16M
Mercator 25M LambertConformal 16M
TransverseMercator 14M
AzimuthalEquidistant 13M
Mollweide 6M

Every projection exceeds 5M/sec. Projecting all 177 countries takes under 1 ms, even with Mollweide. Natural Earth 110m coastlines load in 5 ms (134 features). Countries load in 48 ms (177 features) and are then cached.

MathText and themes

Benchmark Result
Parse \sum_{i=0}^{n} \frac{1}{i!} = e 473K parses/sec
Parse \begin{pmatrix} a & b \\ c & d \end{pmatrix} 554K parses/sec
30 theme presets loaded 370 µs

SVG rendering

Chart Time Allocated
Simple line (100 pts) 94 µs 136 KB
Line + scatter + bar 109 µs 133 KB
3×3 subplot grid 754 µs 933 KB
Treemap (6 nodes) 60 µs 109 KB
Sankey (4 nodes, 4 links) 63 µs 118 KB
Polar line (50 pts) 33 µs 56 KB
3D surface (10×10) 69 µs 124 KB
3D surface + directional lighting 82 µs 148 KB
Line + legend (3 series) 140 µs 214 KB
Large line (10K pts) 3.1 ms 3.7 MB
Large line (100K pts, LTTB→2K) 1.3 ms 2.4 MB

Technical indicators (100K points ≈ a trading day at 1-second bars)

Every indicator completes in under 3.3 ms. Multiple indicators run in parallel on separate cores.

Indicator Time Indicator Time
SMA(20) 195 µs MACD(12,26,9) 1.50 ms
VWAP 238 µs ParabolicSAR 1.21 ms
EquityCurve 226 µs CCI(20) 2.16 ms
EMA(20) 491 µs BollingerBands(20) 2.23 ms
OBV 645 µs ADX(14) 2.43 ms
RSI(14) 892 µs WilliamsR(14) 2.97 ms
Stochastic(14,3) 3.31 ms

Coordinate transform

The transform from data space to pixel space is the hot path. It is a single-pass AVX SIMD interleave: Vector256.Multiply and Add (FMA when available), then UnpackLow/High, then Permute2x128, then a store via MemoryMarshal.Cast. Off x86 it falls back to scalar code. It takes 764 ns for 1K points, 53 µs for 10K and 208 µs for 100K. Every line, scatter, area and bubble renderer uses this transform, so indicator output benefits from it automatically.

Why server-side SVG

Benefit Detail
Zero client cost the browser swaps innerHTML — no canvas redraw, no layout recalculation
Inline SVG part of the DOM: CSS-styleable, screen-reader accessible, prints as vector
Consistent every client sees the same chart, no browser rendering differences
Bandwidth a typical chart is 5–15 KB; SignalR pushes only what changed
Scales parallel subplot rendering uses every core

Run them yourself

cd Benchmarks/MatPlotLibNet.Benchmarks
dotnet run -c Release -- --filter "*SvgRendering*"
dotnet run -c Release -- --filter "*Indicator*"
dotnet run -c Release -- --filter "*"      # all suites

Run one suite at a time. Concurrent runs contend for the CPU and inflate the timings.

Clone this wiki locally