native 0.0.1
Vectors, masks and wide register packs for C++26
Loading...
Searching...
No Matches
Instruction sets

Start with the ordinary simd operations. Reach for an instruction family when its particular operation or arithmetic contract does useful work for you.

Choosing a family

Dot products and small matrices

Packed dot products do several multiplies and additions per result lane. They are useful for quantized inference, small matrix kernels and reductions where the input precision is lower than the accumulator precision. Choose the input signedness and accumulation rule before choosing a register width.

Integer products and carry chains

Large integers need products and carries that ordinary lane-wise arithmetic cannot express. These operations expose the pieces without choosing a big integer representation for you.

  • x86 IFMA: accumulate halves of 52-bit products in 64-bit lanes.
  • x86 ADX: scalar unsigned addition with carry.
  • x86 BMI2: wide scalar products, alongside bit deposit/extract.

Conversions, fixed-point and complex arithmetic

Conversions often sit at the boundary between compact storage and a wider calculation. Fixed-point and complex instructions combine arithmetic steps with a particular rounding or packing convention.

  • ARM NEON: integer conversions, saturation, narrowing, multiply-high and shifts.
  • x86 F16C: binary16/binary32 conversion.
  • x86 AVX-NE-CONVERT: FP16/BF16 memory conversion.
  • ARM FHM: FP16 products accumulated into FP32.
  • ARM RDM: rounded, saturating fixed-point accumulation.
  • ARM FCMA: arithmetic on adjacent real/imaginary pairs.
  • ARM JSCVT: scalar double conversion with JavaScript integer semantics.
  • Half-precision profiles: elementwise FP16 arithmetic and BF16 storage.

Bits and rearrangement

Bit counts, extraction and permutation are useful for packed data structures, parsers and moving selected elements into the next stage of a calculation. Conflict detection finds repeated destinations before an indexed update.

Polynomials, checksums and cryptography

Carry-less products implement polynomial multiplication over GF(2), rather than ordinary integer multiplication. CRC instructions update a checksum; cryptographic instructions expose round and schedule steps. They save work inside an algorithm but do not supply a complete hash or cipher mode.

Indexed memory and waiting

Gather/scatter follow several indices at once, where a contiguous load cannot. Wait instructions let a polling loop give back execution resources while it waits for a condition to change.

WebAssembly

WebAssembly packages vector operations behind a portable instruction format. SIMD128 gives a fixed 128-bit vocabulary; relaxed SIMD permits host-dependent results for operations where that latitude can avoid extra work.

  • SIMD128: integer and floating arithmetic, memory, conversions and rearrangement.
  • Relaxed SIMD: fused arithmetic, selection, conversion and dot products with explicitly permitted variation.

Common caveats

Import the family module directly, or use native.x86, native.arm, native.wasm, or the host's native hub. Vector operands and results use simd<T,N,Arch>; scalar forms use ordinary C++ values. Vector families link native::native; scalar-only families and feature observation link native::minimal. See imports and build targets.

An ISA tag, a compiler target and runtime admission do different jobs. The tag permits the API operation, the target permits generated instructions, and admission establishes whether the CPU and OS can execute them. Importing a module does none of the latter two. Admit the whole containing function, not just the one operation you happened to call. See target selection.

WebAssembly validates the complete linked module. A runtime branch cannot hide unsupported instructions; build and load separate modules for different feature levels. Relaxed operations can produce different permitted results on different hosts.

Shapes, signedness and masks are part of the contract. Equal register sizes do not make vector types interchangeable. AVX-512 instruction masks use predicate<N,Arch>; AVX2 gathers instead use full-vector sign-bit masks. The family guides spell out packing, rounding, side effects and any extra work.

Instruction families have semantic implementations for constant evaluation. With the required features in Arch, the same constexpr overload evaluates at compile time or emits the native instruction at runtime. Without those features, a separate consteval overload accepts constant inputs only. Vector storage must still exist for the element type, lane count and architecture tag. An unevaluated requires check can see the immediate-only overload; that does not make a later runtime call valid.

Floating constant evaluation uses nearest-even rounding, gradual underflow and masked exceptions, with ARM's DN, AHP, AH and EBF controls clear. An instruction's fixed rules or rounding immediate take precedence. Constant evaluation does not read or update the thread's floating-point environment; runtime operations do. Hardware observations, waits and control-register operations remain runtime-only.