|
native 0.0.1
Vectors, masks and wide register packs for C++26
|
C++26 SIMD values and instruction interfaces for x86-64, AArch64 and WebAssembly, with a shared vocabulary for compiler features and runtime admission.
simd<T,N,Arch> keeps the element type, lane count and instruction requirements in the type. Omitting Arch uses the native.simd module's compiler baseline. Comparisons produce masks; wide<V,M> groups registers into independent instruction chains. Operations have no runtime dispatch inside them.
Compile a kernel for the instructions it uses, then admit execution on the intended CPU or Wasm engine. This x86/AArch64 example loads four floats, computes 2*x + 1, and replaces negative inputs with zero:
The import makes the API visible. The target scope permits the compiler to use those instructions; the capability check admits execution. None substitutes for the others. Applications with several kernels can make this choice once at startup; the dispatch guide shows how to compile and select a list of variants.
Documentation ยท Instruction sets
The tested toolchain is Clang 23, CMake 4.4 and Ninja. Configuration checks C++26 structured-binding packs and the Clang property extension used by ISA values and swizzles.
Use clang-cl on Windows. Exceptions default to disabled; set NATIVE_ENABLE_EXCEPTIONS=ON for an exception-enabled application. Producer and consumer compiler, standard-library and runtime modes must agree.
In an application configured with that installation on CMAKE_PREFIX_PATH:
native::native supplies SIMD and vector instruction modules. It links native::minimal, which supplies capability detection, scalar instruction utilities and common types. A detector-only application can import native.features and link native::minimal (native::common is an alias). Headers such as <native/targets.h> supply macros, which modules cannot export.
The package uses the toolchain's default baseline unless NATIVE_MINIMAL_COMPILE_OPTIONS chooses a stronger one. The process must already satisfy that minimum before runtime selection can help. Importing stronger operations does not strengthen an ordinary caller's compiler target. Build details cover installation, compiler settings and PCH/LTO.
Start with SIMD values, masks and memory for construction, short vectors, tails, swizzles and packs. On x86 and ARM, a three-float vector has three logical lanes: its load touches twelve bytes even if its register has room for four. native::mask<V> names the mask associated with V.
ISA values describe requirements, from an individual feature to presets such as avx2, avx512 and neon. Each isa<Family> belongs to one architecture: isa<> uses target_arch, while isa<x86>, isa<arm> and isa<wasm> name it explicitly. NATIVE_BASELINE records the current translation unit's enabled compiler features; runtime CPU observation is a separate operation. Target lists and dispatch connect those requirements to compiled kernels and runtime selection.
Use the instruction guide when an algorithm needs a particular dot product, conversion, polynomial operation or checksum. Vector forms take simd values; scalar forms take ordinary C++ values. Their feature requirements and arithmetic contracts remain specific to the instruction.
Import native.math separately for promoted numerical kernels such as math::exp, math::exp2, math::expm1, math::log2, math::log1p, math::tanh, math::atan2, and math::sincos. Their domains and batching behavior are described in the value guide. The compile-time recommendations native::exp_width<T,K,A>, atan2_width<T,K,A> and their counterparts choose a starting register count for wide<simd<T,K,A>, N>; callers can always choose another extent. The math guide describes polynomial degrees, accuracy and register-count recommendations. Floating-point controls remain under application ownership. Wasm SIMD128 promoted kernels use separately rounded multiply/add stages; x86 and ARM use fused stages. The separate FTZ package builds reproducible binary32 arithmetic on this library's element extension.
The WebAssembly backend supplies 128-bit integer, float and double vectors through native.simd, native.wasm and native. It includes saturating arithmetic, widening and narrowing, conversions, shuffles and memory operations. Relaxed SIMD adds the 20 relaxed operations through native.wasm.relaxed and the Wasm hubs. Both use typed simd operands and support constant evaluation; relaxed results can vary between engines.
The WebAssembly detector, available through native.wasm.features, native.features or native, describes simd128 and relaxed_simd with isa<wasm>. It accepts engine observations through a C++ validation callback or an optional JavaScript adapter. On Wasm compiler targets, NATIVE_BASELINE records the SIMD features enabled by the compiler separately from runtime engine support. Applications compile and load separate modules when they need different feature levels: an engine validates the complete module, including instructions behind branches that are never taken.
See LICENSE.md for the dual BSD-2-Clause/Apache-2.0 license and individual source notices for retained upstream terms.
Contributions and bug reports are welcome through GitHub. Edward Kmett can also be reached as ekmett on Libera Chat and @kmett on Twitter/X.