as-simd is a portable vector layer and AssemblyScript transform. Write one
v64/v128/v256/v512 code path; builds with --enable simd select measured
native WebAssembly SIMD kernels, while builds without it use allocation-free
SWAR fallbacks.
Table of Contents
npm install as-simdHalf-width (64-bit) usage follows the same value-oriented convention as the AssemblyScript SIMD API. The complete width-generic surface is summarized in API.md.
v128, v256, and v512 have the same method names, argument order, generic
parameters, and value semantics. Only the vector parameter/return type and the
width of bitmask differ:
const a128: v128 = v128.splat<i16>(10);
const b128: v128 = v128.add<i16>(a128, a128);
const a256: v256 = v256.splat<i16>(10);
const b256: v256 = v256.add<i16>(a256, a256);
const a512: v512 = v512.splat<i16>(10);
const b512: v512 = v512.add<i16>(a512, a512);The native lane namespaces scale the same way:
| lanes | 128-bit | 256-bit | 512-bit |
|---|---|---|---|
| signed bytes | i8x16 |
i8x32 |
i8x64 |
| signed 16-bit | i16x8 |
i16x16 |
i16x32 |
| signed 32-bit | i32x4 |
i32x8 |
i32x16 |
| signed 64-bit | i64x2 |
i64x4 |
i64x8 |
| 32-bit floats | f32x4 |
f32x8 |
f32x16 |
| 64-bit floats | f64x2 |
f64x4 |
f64x8 |
Each wider namespace mirrors its v128 counterpart, with lane counts embedded in conversion, narrowing, extension, dot-product, and shuffle method names scaled to the vector width.
The public API does not expose destination-register forms such as
v512r.add(dst, a, b). Operations take vector values and return vector values,
just like AssemblyScript's native v128 namespace. v128_swar remains
available as an explicit low-level two-u64 interface.
v32is a packed scalar value and delegates to the tunedv64SWAR kernels.v64is the allocation-free SWAR hot path.v128,v256, andv512are immutable value-semantics facades. Lowercasev256andv512use one raw-width managed object per result rather than a tree of 128-bit wrapper objects.v128_swaris the explicit allocation-free two-half primitive.- Width-specific implementation code lives independently under
assembly/v256andassembly/v512, allowing each width to be tuned without changing the shared public signatures.
The generic operations dispatch at compile time. Without --enable simd, the
exact same APIs use SWAR. With SIMD enabled, their 128-bit chunks use the
adaptive native/SWAR kernels selected for v128; enabling SIMD does not force
every operation through a native instruction.
The repository also retains width-specific register kernels as internal benchmark and tuning machinery. Those experiments show that the native/SWAR crossover can change with width:
| representative operation | v256 SIMD build | v512 SIMD build |
|---|---|---|
| i64 add/subtract | SWAR | 4 native chunks |
| integer negation and i32/i64 shifts | SWAR | 4 native chunks |
| u8/u16 rounded average | SWAR | 4 native chunks |
all_true<i64> |
SWAR | 4 native chunks |
| bitmask and saturating arithmetic | 2 native chunks | 4 native chunks |
These internal choices are backed by the repository benchmark suite and protected by WAT code-shape tests; they are not a second public API. A separate API-parity gate compares all 95 public methods and compiles every signature for both v256 and v512 in SWAR and SIMD builds.
Use as-simd directly as a transform. Its TypeScript implementation follows a
Valent-Block-style pipeline: inline small helpers, remove control-flow shells,
then rewrite the resulting pure Binaryen expression islands bottom-up to a
bounded fixed point, so one contraction can expose another outer fusion. The
domain-specific rules fuse masked bitselects, merge or factor constant masks,
turn addition of disjoint masked fields into one mask, cancel shift/repack
pairs, and avoid expanding a one-bit-per-lane comparison mask when a bitmask or
any_true/all_true reduction immediately contracts it again. Binaryen's
normal cleanup passes run after these rewrites.
Set AS_SIMD_OPTIMIZE=0 to disable the extra pipeline for diagnostic builds.
Set AS_SIMD_OPTIMIZE_DEBUG=1 to print the number of inspected expressions and
successful SWAR fusions.
When WAGO_PLUGINS contains wide, SIMD builds recognize adjacent 128-bit
chunks of eligible v256 and v512 operations and replace them with ordinary
function imports from as-simd. By default, vector operations use standard
externref parameters and results, bracketed by pointer-based
v256.load/v256.store or v512.load/v512.store imports. Import names are
dedicated, architecture-neutral Wasm-SIMD-style operations such as
i8x32.add, i16x32.mul, and v512.bitselect; numeric Wasm opcodes and
machine instruction names are never part of the guest ABI. A v512 group
remains one 64-byte instruction instead of two independent v256 calls.
With the separate github.com/JairusSW/wide Wago plugin installed, Wago
selects the native backend internally: cost-selected AVX-512/ZMM or AVX2/YMM
on amd64 and 128-bit NEON chunks on arm64. Simple operations stay on YMM when
the host cracks ZMM into 256-bit halves; operations that collapse a longer
sequence, such as i64x8.mul, use the full width. The imports remain unchanged.
Wago erases the carrier values into registered native register bundles, so an
expression chain has one set of checked input loads and one final store instead
of a linear-memory round trip per operation. No host reference is allocated or
entered into the runtime reference store. Every backend retains validated call
boundaries and checked loads/stores. No custom Wasm section, custom value type,
digest, post-build rewrite, or architecture-specific import is involved.
Until Wide's externref integration is released as a stable tag, install its current branch explicitly:
go get github.com/JairusSW/wide@mainrt := wago.NewRuntime()
definition := wide.Definition()
digest, err := wago.DefinitionDigest(definition)
if err != nil { panic(err) }
err = rt.LoadPlugins(context.Background(), wago.PluginSet{
Providers: []wago.PluginProvider{wide.Provider()},
Selections: []wago.PluginSelection{{
ID: definition.ID, DefinitionDigest: digest, Direct: true,
Dependencies: map[string]string{},
Grants: []wago.AuthorityGrant{
{Name: wago.AuthorityCompilerTypeDefine,
Scope: wago.AuthorityScope{Modules: []string{"wide"}}},
{Name: wago.AuthorityCompilerInstructionDefine,
Scope: wago.AuthorityScope{Modules: []string{wide.InstructionModule}}},
},
}},
})
if err != nil { panic(err) }
mod, err := rt.Compile(wasmBytes)The Wasm validator checks one physical ABI: load is (i32) -> externref,
vector operations are (externref...) -> externref, and store is
(externref, i32) -> (). These are opaque carrier signatures rather than
runtime object semantics. No host reference is created. The plugin owns the
semantic operation catalog while Wago owns target selection and lowering.
Current automatic v512 recognition covers chunk-independent unary and binary SIMD operations. The plugin registers reviewed unary, binary, and ternary virtual kernels directly under the same import module. Automatic ternary grouping is deliberately left portable when Binaryen aliases prevent the transform from proving that all three sources are adjacent. Width-crossing operations likewise retain their portable implementation.
Wide imports are opt-in. Enable them when the corresponding Wago plugin is part of the build:
WAGO_PLUGINS=wide npx asc assembly/index.ts --transform as-simd --enable simdWAGO_PLUGINS accepts a comma-separated list, such as wide,other. Without
wide in that list, an ordinary SIMD build stays self-contained.
AS_SIMD_OPTIMIZE=0 disables the entire transform optimization pipeline,
including custom-instruction call generation.
To verify the emitted ABI against Wide's current main branch and the latest
Wago compiler:
npm run test:wago-wideThe integration check verifies the externref ABI. Set WIDE_DIR or
WAGO_DIR to test local checkouts without modifying either checkout, or use
WIDE_VERSION and WAGO_VERSION to select specific Go module revisions.
CLI:
npx asc assembly/index.ts --transform as-simdProgrammatic asc.main():
await asc.main(["assembly/index.ts", "--transform", "as-simd"]);If a tool expects a direct source entrypoint, use as-simd/sources.
To opt into real SIMD codegen, explicitly enable SIMD:
npx asc assembly/index.ts --transform as-simd --enable simdWithout explicit SIMD opt-in, imported as-simd APIs compile to strict SWAR.
The transform also redirects the generic AssemblyScript v128 API as
described below; lane-specific native globals still require --enable simd.
Portable value-semantics widths can be imported normally from the package:
import { v64, v128, v256, v512 } from "as-simd";
const a: v256 = v256.splat<i16>(4);
const b: v256 = v256.add<i16>(a, a);With --transform as-simd, the imports may be omitted. The transform detects
uses of v64, v128, v256, v512, and the twelve wide lane namespaces in
each user source and injects only the missing names. Existing declarations and
manual imports are never overridden. The injected module specifier is derived from the transform's real
package location and the current source file, so it works from the repository
itself as well as flat npm, pnpm, linked, nested, and workspace installs. In
SIMD builds an unimported v128 remains AssemblyScript's native
type; without SIMD it is redirected to the two-u64 facade. The other widths
are imported from as-simd in both modes and dispatch their 128-bit chunks to
native SIMD or SWAR as appropriate.
For example, this source needs no import when the transform is enabled:
const a = v256.splat<i16>(4);
const b = v256.add<i16>(a, a);
export const lane0 = v256.extract_lane<i16>(b, 0);Compile it as strict portable SWAR or adaptive SIMD without changing the source:
npx asc assembly/index.ts --transform as-simd
npx asc assembly/index.ts --transform as-simd --enable simdSet AS_SIMD_AUTO_INJECT=0 to require explicit imports. The narrower
AS_SIMD_V128_FALLBACK=0 switch leaves an unimported v128 native-only while
continuing to inject the other widths. Lane-specific builtin namespaces such
as i8x16 remain native-only; portable source should use the generic width
names or explicit i8x16_swar APIs.
The lowercase v256 and v512 value types store all four or eight scalar words
directly in one immutable value object. For a strict zero-allocation low-level
loop, use v64 or v128_swar.
For IntelliSense on global aliases, include:
{
"include": ["./node_modules/as-simd/globals.d.ts"]
}import { i8x8 } from "as-simd";
const a = i8x8(1, 2, 3, 4, 5, 6, 7, 8);
const b = i8x8(8, 7, 6, 5, 4, 3, 2, 1);
const sum = i8x8.add(a, b);
const product = i8x8.mul(a, b);
const sat = i8x8.add_sat_s(a, b);
const lane3 = i8x8.extract_lane_s(sum, 3);import { i8x8 } from "as-simd";
let x = i8x8.splat(5); // [5,5,5,5,5,5,5,5]
x = i8x8.replace_lane(x, 2, -7); // [5,5,-7,5,5,5,5,5]
const v = i8x8.extract_lane_s(x, 2); // -7import { v512 } from "as-simd";
const a = v512.splat<i16>(32000);
const b = v512.splat<i16>(1000);
const sum = v512.add_sat<i16>(a, b);
const last = v512.extract_lane<i16>(sum, 31); // 32767import { i8x8 } from "as-simd";
const a = i8x8(10, -2, 30, -40, 50, -60, 70, -80);
const b = i8x8(1, 2, 3, 4, 5, 6, 7, 8);
const sub = i8x8.sub(a, b);
const mul = i8x8.mul(a, b);
const lt = i8x8.lt_s(a, b); // lane masks: 0x00 or 0xFF per lane
const laneMask = i8x8.bitmask_lane(lt); // 0x80 in each truthy lane
// Existing bitmask() returns packed lane bits. ctz(mask) << 3 gives
// the byte shift for the first truthy lane.
const firstByteShift = ctz(i8x8.bitmask(lt)) << 3;
// bitmask_lane() returns a vector-shaped mask. ctz(mask) >> 3 gives
// the first truthy lane index.
const firstLane = ctz(laneMask) >> 3;import { i8x8 } from "as-simd";
const hi = i8x8(120, 120, -120, -120, 100, -100, 127, -128);
const lo = i8x8(20, 40, -20, -40, 50, -50, 1, -1);
const satAdd = i8x8.add_sat_s(hi, lo);
const satSubU = i8x8.sub_sat_u(hi, lo);
// narrow from packed i16 lanes in two v64 values -> one i8x8
const narrowed = i8x8.narrow_i16x4_s(0x0001000200030004, 0xfff0fff1fff2fff3);import { i8x8 } from "as-simd";
const a = i8x8(0, 1, 2, 3, 4, 5, 6, 7);
const b = i8x8(10, 11, 12, 13, 14, 15, 16, 17);
const mixed = i8x8.shuffle(a, b, 0, 1, 8, 9, 2, 10, 3, 11);
const indexed = i8x8.swizzle(a, i8x8(7, 6, 5, 4, 3, 2, 1, 0));as-simd focuses on lane-parallel i8x8 behavior with multiple implementations:
- scalar mirror (
assembly/scalar/i8x8.ts) for correctness oracle behavior - SWAR implementation (
assembly/v64/lanes.ts) for baseline portability - SIMD-enabled code paths (compile-time gated by
ASC_FEATURE_SIMD) where profitable
Correctness is validated by:
- deterministic unit parity tests against scalar
- mode-specific fuzz parity in SWAR and SIMD builds
All generated charts and exact Markdown result tables are in charts.
These summaries use the geometric mean across every operation in each same-width V8 benchmark suite. Each operation gets equal weight, and the families are not treated as equivalent workloads. The overview table includes sample counts, medians, win counts, and the best and worst operation for each family.
The register-kernel chart records internal implementation research used to choose native SIMD versus SWAR paths. These destination-register helpers are not a second public API. See the register benchmark table for exact values.
The focused wide-kernel chart measures the dedicated fixed-width scheduler. Its fallback module is compiled without the WebAssembly SIMD feature, while the adaptive module is compiled with SIMD enabled and still retains SWAR for operations where two native chunks lose at v256 width. The exact values and runtime metadata are also available in the wide benchmark table.
The Wago microbenchmark isolates 128 dependent i8x32.add operations and
compares the plugin's AVX2/YMM carrier with the paired-v128 and scalar SWAR
implementations. On the measured Ryzen 7 7800X3D host, native lowering takes
1.80× less time than paired v128 and 13.47× less time than SWAR, with zero
allocations in every mode. Exact ten-round values are in the
Wago v256 benchmark table.
The immutable-value benchmark isolates the documented lowercase facades from the retained nested compatibility classes. It runs the same splat/add/subtract/ extract expression chain through both representations; the raw-width layout is 2.45× faster for v256 and 3.14× faster for v512 on this V8 host. Exact values are in the wide-value benchmark table.
Here's some results comparing i16x4 (SWAR) versus the native i16x8 (SIMD) implementation.
Benchmarks are run directly on top of v8 for tighter control over the engine configuration.
For the Wago native/paired/SWAR v256 comparison, keep the Wago checkout at
../../Wago/wago or set WAGO_DIR, then run:
npm run bench:wago-v256
npm run charts:wago-v256- Install the local benchmark prerequisites:
npm install -g jsvu
jsvu --engines=v8- Add
~/.jsvu/binto yourPATHand make surewasm-optis installed:
export PATH="${HOME}/.jsvu/bin:${PATH}"
sudo apt-get install -y binaryen- Install project dependencies:
npm install- Run benchmarks:
npm run bench -- --v8Run one suite and mode explicitly:
BENCH_SAMPLES=7 npm run bench -- i32x4 --mode swar --v8
BENCH_SAMPLES=7 npm run bench -- i32x4 --mode simd --v8Run the dedicated v256/v512 register-kernel benchmark:
npm run bench:wideRun the lowercase raw-width value-layout benchmark:
npm run bench:wide-valuesRegenerate its chart after the benchmark:
npm run charts:wideThe i8x16, i16x8, i32x4, i64x2, and generic v128 fallback benchmarks are
compiled without the WebAssembly SIMD feature. Transform tests inspect their
emitted WAT and reject any v128 type or SIMD opcode.
- Build charts:
npm run chartsContributions are welcome. For changes to core vector behavior:
- keep scalar and vector implementations behaviorally aligned
- update or add deterministic tests in
assembly/__tests__ - update or add fuzz checks in
assembly/__fuzz__ - run the deterministic, transform, and full multi-mode fuzz suites before opening a PR
The full local verification gate is:
npm run test:transform
npm test
npm run fuzz
npm pack --dry-runThe transform gate derives its coverage list from the public sources. It
compiles and executes all 190 public v256/v512 value methods and all 462
functions in the twelve wide lane namespaces in both SWAR and SIMD modes; a new
public method fails the gate until its invocation is covered. Deterministic
scalar edge-value tests additionally oracle every i64x4 and i64x8
operation, including overflow, signed and unsigned shifts, comparisons,
widening, extended multiplication, shuffle, and relaxed lane selection. The
Wago plugin separately byte-compares portable and native execution for every
catalogued unary, binary, and ternary wide kernel.
Prefer narrowly scoped commits with Conventional Commit messages.
This project is distributed under an open source license. Work on this project is done by passion, but if you want to support it financially, you can do so by making a donation to the project's GitHub Sponsors page.
You can view the full license using the following link: License
Please send all issues to GitHub Issues and to converse, please send me an email at me@jairus.dev
- Email: Send me inquiries, questions, or requests at me@jairus.dev
- GitHub: Visit the official GitHub repository Here
- Website: Visit my official website at jairus.dev
- Discord: Contact me at My Discord or on the AssemblyScript Discord Server