native 0.0.1
Vectors, masks and wide register packs for C++26
Loading...
Searching...
No Matches
SIMD values and masks

import native; provides the host's SIMD types, instruction families and common utilities. This guide starts with values and their memory contracts, then shows how ISA requirements connect to compilation and execution. The README example includes a target scope and runtime check; vector snippets below belong inside a function compiled for their chosen ISA.

The WebAssembly backend provides SIMD128 integer, float and double registers with canonical lane masks. Wasm engine admission and separate module loading remain application responsibilities.

Values and generic algorithms

native::simd<T,N,A> takes an element type, a lane count and an isa<> value as a non-type template argument. Vectors with different ISA values remain distinct even when their register widths match. Imports control visibility; template arguments control overload resolution and ABI. Using an AVX2 vector in an AVX-512 function keeps its original mask representation and type identity.

native::simd is the class template itself. Extension specializations and template-template arguments name native::simd directly.

The architecture argument defaults to the compiler baseline used to build native.simd, so simd<float,4> names the same type as an explicit vector with that ISA. This records compiler permissions, not runtime CPU detection. A stronger function target or importer does not change a previously built module's default.

One-lane values retain that tag even when its features select no vector arithmetic profile. They use scalar operations; wider values still require the features for their storage and operations.

Choose an explicit ISA for kernels with different requirements:

native::simd<float,4,native::avx2> lanes{1.f, 2.f, 3.f, 4.f};
Omitted architecture arguments use the native.simd provider's baseline.

The default is supplied by native.simd (also imported by native). Standalone headers and native.scalar alone keep the architecture explicit. Target-list ordering and feature constraints apply equally to defaulted and explicit tags.

Generic algorithms take the ISA as a value parameter:

template<native::isa<> A, std::size_t N>
struct kernel {
static void run(float const * a, float const * b, float * out) {
auto x = V::load(a), y = V::load(b);
fma(x,y,x).store(out);
}
};
// Instantiate in a translation unit compiled for the selected profile:
// kernel<native::avx2,8>::run(a,b,out);
constexpr void store(fp16 *p) const noexcept
Store 32 FP16 objects with the default alignment contract of store_memory().

Presets such as native::avx2 and native::avx512 are isa<native::x86> values. Use .has(...), feature properties or subset comparisons to inspect them. Single-feature construction is exact; feature_closure adds prerequisites explicitly. The ISA guide covers feature conjunction and compile-time target selection.

native::target_arch is an inline constexpr architecture value identifying the compiler target's family. Compare it with native::arm, native::x86 or native::wasm in generic code. It is available from native.isa, native.features, native and the dependency-free <native/config.h> header. It describes neither runtime CPU detection nor optional instruction support; both 32-bit and 64-bit targets belong to their respective family. Platform headers and unavailable native declarations still need preprocessing guards.

isa<> uses this family. Explicit isa<arm>, isa<x86> and isa<wasm> values can describe foreign metadata, but cannot be mixed or supplied to a host vector type from another family.

The element type supplies the arithmetic contract; the ISA determines storage and available operations. Some instruction-specific shapes provide storage and transfer without the full arithmetic interface of a vector profile. Having a simd type does not by itself promise every operation on that type.

For example, base NEON supports four- and eight-lane fp16 and bf16 storage: simd<fp16,4,neon> and simd<bf16,8,neon> can transfer their elements without half arithmetic instructions. FP16 arithmetic needs its own neon_fp16 feature; BF16 dot and matrix operations need neon_bf16. Storing either format does not authorize those operations or change its scalar conversion policy.

Keep dependent mathematical calls unqualified so ADL can select the register's overload. There is no runtime architecture branch in individual operations.

FP16 vector arithmetic also works during constant evaluation: addition, subtraction, multiplication, division, square root and FMA use round-to-nearest, ties-to-even with gradual underflow. Comparisons, sign changes and selection preserve their lane semantics. Constant evaluation does not read or change the floating-point control or status registers. Runtime operations retain the caller's architectural environment. NaN payload selection during arithmetic is not a portable cross-platform promise.

BF16 dot2 is constant-evaluable with the architecture's instruction semantics. ARM uses legacy BFDOT behavior with EBF clear; x86 retains VDPBF16PS's high-product-first ordering. Their fixed rounding and denormal rules differ.

Half-precision values

fp16 and bf16 store 16-bit floating-point representations. Storage support does not imply arithmetic support. In particular, base NEON can load and store both formats without enabling their optional arithmetic instructions.

Profile Element and lane counts Arithmetic
neon_fp16 fp16, 8 lanes +, -, *, /, sqrt, fma, comparisons and selection
avx512_fp16 fp16, 32 lanes The same operations, with a compact predicate mask
avx512_bf16 bf16, 8, 16 or 32 lanes dot2 into 4, 8 or 16 FP32 lanes
neon_bf16 See ARM BF16 Pair dots, matrices and widening multiply-adds

FP16 arithmetic rounds in binary16; FMA rounds once after the product and sum. Runtime rounding and exception behavior follow FPCR on ARM and MXCSR on x86. ARM's FZ16 controls half subnormal flushing. AVX-512 FP16 uses gradual underflow regardless of MXCSR's DAZ/FTZ bits. Neither backend changes the control register.

For x86 BF16, use matching input and accumulator shapes inside a function compiled for avx512_bf16:

#include <native/targets.h>
import native;
NATIVE_TARGET_PUSH(avx512_bf16)
F accumulate_pairs(B a, B b, F accumulator) noexcept {
return native::dot2(a,b,accumulator);
}
NATIVE_TARGET_POP()
constexpr simd< float, N/2, Arch > dot2(simd< bf16, N, Arch > a, simd< bf16, N, Arch > b, simd< float, N/2, Arch > accumulator) noexcept

dot2(a,b,accumulator) first accumulates each odd-lane product, then the even-lane product. Each step rounds to nearest-even in FP32. The instruction flushes subnormal inputs and outputs and neither reads nor changes MXCSR. This is not a single-rounding sum of three terms, nor does it supply elementwise BF16 arithmetic. ARM BF16 has different rounding rules; see its family guide.

Masks and memory

V::mask is the type produced by comparisons of V. AVX2 and NEON use vector masks; the AVX-512 profile uses compact predicates where the declared feature set supports the lane width. Masks for custom numerical elements use the raw storage register's representation.

native::mask<V> names the mask associated with V, ignoring cv/ref qualifiers. Use it when a generic algorithm needs the comparison result's type.

using M = native::mask<V>;
V x(2.f), y(3.f);
M active = x < y;
auto chosen = select(active,x,y);
typename mask_traits< std::remove_cvref_t< T > >::type mask
Definition mask_traits.h:22

Masks support lane-wise &, |, ^, ~ and !, with &=, |= and ^= for updates. Both complements invert each logical lane. == and != also return masks: they compare each pair of truth values rather than reducing the whole vector. Use any(active), all(active) or none(active) to obtain one bool for a branch. Mask expressions do not short-circuit individual lanes; all operands are evaluated.

Masked x86 instruction interfaces use native::predicate<N,Arch> directly, where N is the operation's mask lane count. This represents the architectural predicate without first packing a full vector mask. A full AVX-512 profile's associated mask may be that same type; do not assume this for every minimal instruction feature set. The family guides give each operation's mask type and inactive-lane behavior.

Boolean vectors, all-zero/all-one lane masks and compact predicates have different storage contracts. Convert between them with the named conversion operations. Use vector_mask_type when an algorithm needs full vector lanes. Mask tags describe lane width; the chosen profile determines the register representation.

For pointer operations, specify the vector type to load:

constexpr void store_simd(U *p, V value, simd_memory< A, Access >={}) noexcept(noexcept(value.template store_memory< A >(p)))
Definition common_body.h:58
constexpr V load_simd(U const *p, simd_memory< A, Access >={}) noexcept(noexcept(V::template load_memory< A >(p)))
Definition common_body.h:40

Alignment policies are caller promises. Use load_simd_partial<V>(p,count,fill) and store_simd_partial(p,value,count) for tails. Only the requested logical lanes are accessed; the load supplies fill for the rest. The streaming policy currently uses ordinary accesses, so it carries no non-temporal-store guarantee. The count must not exceed the logical lane count. Clang diagnoses a count it can prove too large at the call site, including through module imports. Dynamic counts remain a caller precondition; this diagnostic adds no runtime check.

Short vectors and swizzles

The native profiles also support two- and three-lane float, int32_t, uint32_t and mask32 vectors in a 16-byte register. Lane count describes logical elements: a three-float load or store touches twelve bytes. It does not require a readable or writable fourth element. Padding does not participate in comparison-mask reductions.

V position{1.f,2.f,3.f};
auto pair = position.xy; // simd<float,2,avx2>
auto saved = position.xyz; // an independent simd<float,3,avx2>
position.xyz = position.zyx; // snapshot the right side, then scatter
position.x = 4.f; // a single component has element type float

Named xyzw swizzles apply to vectors of at most four logical lanes where the requested result shape exists. Reads may repeat indices. Writes require a mutable lvalue and distinct destination indices. position.xxy is readable; assigning to it is ill-formed. Accessing nonexistent input lanes is ill-formed as well. Architecture and element type are retained in vector results.

Swizzle properties invoke accessors without changing the register layout. Lane construction supports constant evaluation. Initialize with {} to value-initialize the backing register; ordinary default initialization without braces leaves it uninitialized.

A copied swizzle is a value, not a view. Address-taking, mutable-reference binding and assignments through a temporary swizzle such as position.xyz.x = 5.f are rejected. Assign directly to the intended parent's property. Compound assignments use the compiler's getter, binary-operation, setter rewrite; they do not invoke the vector's compound-assignment member on a stored proxy.

Stable compaction and expansion

The scalar, AVX2, AVX512 and NEON profiles support float, int32_t and uint32_t lanes with the same compaction contract:

auto active = V::mask::from_bitset(0b1010);
V values{10u, 20u, 30u, 40u};
auto packed = native::compress(active, values, 99u);
// packed.value == {20, 40, 99, 99}; packed.count == 2
auto restored = native::expand(active, packed.value, V(77u));
// restored == {77, 20, 77, 40}
uint32_t output[2]{};
auto written = native::compress_store(output, 1, active, values);
// written == 1, output[0] == 20; output[1] was not accessed
constexpr std::size_t compress_store(T *destination, std::size_t capacity, typename simd< T, N, Arch >::mask mask, simd< T, N, Arch > value) noexcept
constexpr compaction_result< simd< T, N, Arch > > compress(typename simd< T, N, Arch >::mask mask, simd< T, N, Arch > value, T fill=T{}) noexcept
constexpr simd< T, N, Arch > expand(typename simd< T, N, Arch >::mask mask, simd< T, N, Arch > packed, simd< T, N, Arch > prior) noexcept

compress preserves increasing logical lane order and returns both the packed register and selected count. Its scalar fill defaults to zero and supplies every unused logical output lane. expand consumes the first selected-count lanes from its packed register, in order, and requires an explicit prior register for the unselected positions. Both rearrange bits without floating-point arithmetic, preserving signed zero, subnormals and NaN payloads. Short vectors exclude physical padding from masks and counts; their output padding is zero.

compress_store writes the first min(capacity, selected_count) selected elements and returns the number written, which may be less than the count returned by compress. No later destination element is read or written. A null destination is allowed when capacity is zero or no logical lane is selected. For a nonzero write, the destination must provide that many writable elements of the vector's element type. These operations compact within one register. Applications can use the returned counts to assemble batches across registers, without runtime backend selection inside the operations.

Several registers

native::wide<V,M> holds M values of type V. It separates the number of independent registers from the number of lanes in each register. Generic pointwise operations can use element overloads or an ADL array kernel supplied by a numerical library; the container itself does not choose an ISA.

Numerical batches also use std::array<V,M>. The promoted math interface below advances each polynomial stage across the array before starting the next stage, so the independent chains remain visible to the compiler.

Promoted math batches

import native.math; provides math::exp and the other promoted numerical kernels. import native; provides wide::promote/wide::demote<Original> and the register primitives used by those kernels. Batches use std::array. A float promotes to std::array<native::simd<float,1,native::scalar>,1>; a SIMD value promotes to a one-element array retaining its lane count and ISA. An array adapts its elements without adding another outer dimension. native::wide also adapts to a standard array. Tuples are not accepted by promotion or promoted math.

#include <array>
import native;
import native.math;
auto scalar_result = math::exp(1.f); // float
auto vector_result = math::exp(V(1.f)); // V
auto batch_result = math::exp(std::array{V(1.f), V(2.f)});
// batch_result is std::array<V,2>.

Compile the caller for the selected vector target, as with other native operations. Each polynomial stage advances all independent chains before the next stage begins, rather than finishing one exp call per element. Results preserve the input scalar, SIMD, array, or native::wide shape, including empty and one-element containers. These kernels support binary32 elements on x86, ARM and Wasm SIMD128. Wasm provides math::exp, exp2, expm1, log, log2, log1p, damping_gain, tanh, atan2, sin, cos and sincos for single vectors, standard arrays and native::wide; native::math exposes the same Wasm kernels. SIMD128 has no fused multiply-add instruction, so its polynomial stages use separate binary32 multiply and add roundings, with contraction disabled even in a relaxed-SIMD caller. Constant evaluation uses this same Wasm graph. The x86, ARM and scalar graphs retain FMA; cross-architecture bitwise equality is not a contract for these approximations. Neither public fma nor wide::fma acquires a nonfused Wasm implementation.

Choosing a register count

import native.math; also supplies native::exp_width<T,K,A>, atan2_width<T,K,A> and a name_width variable template for each other transcendental kernel. The same declarations are available in math and native::math. Their std::size_t values recommend the number of independent registers, not the number of lanes or the total number of elements:

import native.math;
using namespace native;
constexpr auto A = avx2;
using exp_batch = native::wide<V, exp_width<float, 8, A>>; // 6 registers, 48 floats
using angle_batch = native::wide<V, atan2_width<float, 8, A>>; // 2 registers, 16 floats
Architecture-tagged vectors, register packs and supporting value types. Native arithmetic follows its...
Own N values of T, including an empty pack when N == 0. registers is the underlying array....
Definition wide.h:83

These are starting points for tuning; they do not change the kernel, split a larger batch, or dispatch at runtime. Any other explicit wide or array extent still works. The caller must compile and admit the selected ISA as usual. K must be positive. Omitting A uses the defining module's compiler baseline, as with SIMD's module default. The header form uses the translation unit's baseline. Pass the same explicit A as the vector in a multiversioned kernel.

Kernel x86 float, K=8 or 16 NEON float, K=4 Wasm SIMD128 float, K=4
exp 6 4 2
exp2 2 2 2
expm1, damping_gain, log, log2, log1p 4 4 2
atan2, tanh 2 2 2
sin, cos, sincos 2 4 2

X86 K=8 requires the AVX2/FMA profile; K=16 requires the AVX-512F/DQ kernel profile. For x86 and NEON short vectors with K=2 or 3, exp_width is 4 and the other recommendations are 2. X86 K=4 follows the NEON column except that sin, cos and sincos use 2. Scalar K=1, other element types, and unrecognized shape/ISA combinations use 1. A returned width does not assert that the corresponding math operation exists. The templates inspect metadata and can be queried with a foreign-family isa without instantiating its SIMD type.

The exp recommendations target the default degree-six polynomial. Measured NEON and full-width AVX2/AVX-512 results guide those choices; other polynomial degrees may prefer another extent. The small atan2 batch limits register pressure, AVX2 exp2 favors two registers, and NEON sincos reaches its throughput knee at four. Remaining entries, including Wasm and mixed feature/width combinations, are conservative heuristics. They do not promise a universal optimum or absence of spills. See the tuning rationale.

exp<Flush = false, Degree = 6> selects a polynomial of degree one through seven, defaulting to six, with nearest-even range reduction. Range comparisons run independently of that arithmetic: lower-cutoff and overflow flags select zero or infinity at the finish rather than clamping the input. Inputs above 88.72283172607421875f return positive infinity. The reconstruction accepts the reduced exponent 128, so it retains the finite upper end of the range.

ARM and AVX2 construct a power-of-two factor. For reduced exponent 128 they split the scale into 2^127 and 2, avoiding an infinite intermediate factor. ARM's defined fcvtzu conversion maps a negative biased exponent or NaN to zero. AVX2 uses VCVTTPS2DQ, clamps the signed integer to zero with VPMAXSD, then shifts it into the exponent field. NaN polynomial values propagate through the multiplication without a NaN check or operand masks. Subnormal accuracy is not guaranteed: these paths return zero when the reduced exponent is at most -127, and the caller's FP controls can flush other tiny results. Flush=true requests the earlier cutoff at -87.33654022216796875f on every backend. AVX-512 uses one masked VSCALEFPS where the width and features admit it, retaining the instruction's subnormal behavior. Wasm retains two-factor reconstruction. These internal exp steps do not supply a public software scaling operation; NaN payloads are unspecified. Range selection does not suppress exceptions from intermediate operations, and FP exception flags can differ between backends. exp2 reduces directly with n = round_even(x) and r = x - n, then approximates 2^r and uses the same reconstruction. It avoids converting the argument to natural-log units. Its intentional early overflow cutoff is 127.5; Flush=true selects zero below -126. log2 separates the binary exponent from a mantissa close to one (approximately [1/sqrt(2),sqrt(2))), approximates its base-two logarithm and adds the exponent. Both use Sollya-generated binary32 coefficients.

The trigonometric kernels require finite |x| < 8192 and preserve signed zero. Wasm uses separately rounded multiply/add stages, so its results can differ from the fused x86 and ARM kernels.

Binary32 arithmetic, comparisons, selection, fused multiply-add, square root, rounding support constant evaluation, including short vectors. Exponent scaling retains constant evaluation only for hardware-admitted scaleb shapes: AVX512F scalar and sixteen-lane vectors, and AVX512VL two-, three-, four- and eight-lane vectors within the supported kernel profiles. The short forms mask padding lanes. scaleb, masked_scaleb and masked_scaleb_zero have no software fallback on scalar, AVX2, NEON or Wasm profiles. The promoted exponential and trigonometric kernels evaluate their existing polynomial graphs with the same input bounds and approximation contracts. Scalar, array and empty-array forms retain their shapes.

Constant evaluation uses round-to-nearest with ties to even and gradual underflow. Numerical arithmetic quiets signaling NaNs, preserves the selected NaN's sign and payload, and uses positive quiet NaN for invalid operations. Signaling NaNs are selected before quiet NaNs; otherwise operand order decides, with the addend first for FMA. An infinite-times-zero FMA product yields the canonical NaN even with a quiet NaN addend. scaleb retains its exceptional scaling rule: a quiet NaN scaled by positive or negative infinity becomes positive infinity or positive zero; a signaling NaN is quieted instead. Comparisons are ordered, while inequality is true for unordered operands. Negation and abs change only the sign bit; bit transport and selection preserve representations, including signaling NaNs. These computations do not read or change floating-point controls or exception flags. Runtime operations continue to use native instructions and the caller's environment; constant evaluation does not promise the runtime target's NaN precedence or status flags.

Promotion owns its values. Demotion uses the original type to restore shape, while retaining transformed element types: a scalar comparison demotes to bool, whereas a SIMD comparison retains its mask. wide::map performs one elementwise stage over equal-length standard arrays. Access their elements with ordinary indexing or std::get.

Pointwise operations use named functions. std::array arithmetic and comparison operators are not changed: container equality still returns one bool, and ordering remains lexicographic. Use wide::cmp_lt, cmp_eq, and the other cmp_* functions for an array of element masks, and wide::mask_not to complement those masks.

// r and y are std::array<V,3> values.
y = wide::fma(r, y, V(0x1.555555c673724p-3f));
auto active = wide::cmp_lt(r, V(0.f));
auto doubled = wide::mul(r, V(2.f));

Lifted operations accept a SIMD operand alongside arrays and reuse it for every chain. Array operands must have equal lengths. SIMD operands must match the array element types; no implicit conversion occurs between register widths or ISAs. wide::constant_like(batch, value) returns one SIMD value, with the coefficient's scalar type matching its SIMD element type exactly.

wide::add, sub, mul, div, and negate provide arithmetic; bit_and, bit_or, bit_xor, and bit_not provide bitwise operations. Comparisons, selection, min/max, abs, sqrt, rounding, fma, and exponent scaling use the same lifting rule. Scaling is available only when the SIMD element admits the native scaling instruction. These stages use compile-time pack expansion.

math::sin, math::cos, and math::sincos also promote and restore the input shape. Their reducer and polynomial advance stage by stage across the array. They retain the native approximation's domain: every lane must be finite with absolute value below 8192 radians. sincos shares the reducer and returns a pair of results, each in the original shape; it retains the original paired kernel's signed-zero behavior.

auto [s, c] = math::sincos(std::array{V(0.25f), V(0.5f)});
// s and c are each std::array<V,2>.

math::flush_to_zero clears subnormal mantissas using integer operations, preserving the sign of zero and the exact bits of normal values, infinities, and NaNs. It leaves floating-point controls unchanged. math::abs, sqrt, floor, ceil, trunc, and round_even retain the native leaf operation's semantics through the same shape-preserving interface. Qualified aliases wide::exp, exp2, expm1, log, log2, log1p, damping_gain, tanh, atan2, sin, cos, sincos, and flush_to_zero are also available; standard arrays do not acquire wide as an associated namespace for ADL.

math::horner(c0, c1, ...)(z) evaluates coefficients from highest power to constant term, preserving the shape of z. The callable owns its coefficients; float coefficients can be reused across scalar, SIMD and wide inputs, while SIMD coefficients must match the input's register type. wide::horner is an alias. Coefficients may also be arrays or wide packs matching the input's extent, including coefficients selected by per-lane masks. A packed float coefficient broadcasts within its corresponding register; shared coefficients remain shared. See polynomial evaluation for its multiply-add and constant-polynomial behavior.

math::log, math::log1p, math::expm1, math::damping_gain, math::tanh, and math::atan2 use the same promotion and staged array evaluation. The first five kernels adapt binary32 polynomials from FTZ; atan2 uses the SLEEF coefficients described below. Ordinary native values do not acquire FTZ's arithmetic policy. No kernel changes the caller's FP controls or inserts software flushing between arithmetic steps. ARM and x86 use fused multiply-add; baseline Wasm SIMD128 uses separate multiply and add and has its own accuracy checks.

Function Boundary behavior
exp2(x) Normal integer powers from -126 through 127 are exact; inputs at or above 127.5 give +inf. exp2<true> selects zero below -126; other subnormal results follow the exp reconstruction policy.
log2(x) Normal powers of two return their exact integer exponents; log2(1) is positive zero. Zero, subnormal, negative and nonfinite inputs follow log.
log(x) Signed zero and subnormal inputs give -inf; negative normal inputs give NaN; +inf is preserved.
log1p(x) -1 gives -inf; inputs below -1 give NaN; +inf is preserved. Inputs with abs(x) <= 2^-25 retain their bits.
expm1(x) Computes exp(x)-1 without cancellation near zero. Signed zero and tiny subnormals are preserved; -inf gives -1; inputs above 88.37625885009765625f give +inf.
damping_gain(x) Computes -expm1(-x) with the same graph; nonnegative inputs approach one.
tanh(x) Signed zero and inputs with abs(x) <= 2^-12 retain their bits. Large magnitudes and infinities saturate to signed one.
atan2(y,x) Returns radians with signed axes and the usual infinity quadrants. Subnormal inputs are treated as signed zero; tiny outputs follow the caller's FP mode. Both operands must have the same shape and type.

Logarithms return a canonical quiet NaN for NaN inputs and domain errors. tanh and atan2 also return a quiet NaN for NaN inputs. expm1 and damping_gain propagate NaNs without promising their payloads. These approximations require nearest-even rounding and do not promise libm's exception flags or correct rounding for every input. The regression bank checks normal-domain accuracy and special values separately. Native does not promise FTZ packet equality when subnormal intermediates or nonfused operations differ.

The SIMD and SIMD-array forms also have native::exp2, log, log2, log1p, expm1, damping_gain, tanh, and atan2 entry points. The native::wide adapters use the array kernel for native SIMD elements; custom elements retain their ADL operations. tanh selects a polynomial from seven magnitude intervals using register table lookups. atan2 evaluates one reduced polynomial after a packed min/max ratio division; its coefficients derive from SLEEF, with the Boost license notice retained beside the implementation.

auto t = math::tanh(std::array{V(-1.f), V(1.f)});
auto a = math::atan2(std::array{V(1.f), V(-1.f)},
std::array{V(-1.f), V(-1.f)});

Sampled 256-bit MPFR checks for these two kernels reached a maximum of 2 ULP on ARM. This is a measured sample result, not an exhaustive error bound. See math kernels for the distinction between numerical accuracy, backend agreement and throughput.

math::exp<true> uses the existing early underflow cutoff. Both variants retain the original polynomial, NaN behavior, and floating-point environment policy. Batching preserves each approximation's domain, sequence of operations and floating-point environment contract; it does not improve its accuracy guarantee.

Features, compiler targets and runtime admission

An isa value records requirements. The compiler target determines which instructions a function may contain. A capability observation records what the CPU and OS make available. All three matter at a call boundary:

#include <native/targets.h>
import native;
// The body is compiled for NEON; the ordinary caller retains its baseline.
NATIVE_TARGET_PUSH(neon)
void double_four(float* out, float const* in) {
auto x = V::load(in);
(x + x).store(out);
}
NATIVE_TARGET_POP()
bool try_double_four(float* out, float const* in) {
auto cpu = native::observe_cpu();
if (!native::classify_isa(cpu, native::neon, NATIVE_TARGET_MINIMUM).admitted())
return false;
double_four(out, in);
return true;
}
constexpr isa_admission< x86 > classify_isa(C const &cpu, isa< x86 > requested, isa< x86 > minimum={}) noexcept
Definition isa.h:709

This example is for AArch64. On x86, use the corresponding target and ISA value, for example avx2. The ISA guide explains exact feature sets, compiler prerequisite closure and target<A, Choices...> overload selection. NATIVE_BASELINE captures the translation unit's enabled compiler features, including explicit disables; it does not query the executing CPU. Target pragma scopes do not change that preprocessor snapshot. NATIVE_TARGET_MINIMUM supplies the conservative inherited requirements used for admission.

For multiple implementations, use with_isa with an ordered finite list. It passes the first admitted ISA as a template argument to the callback, or returns false if none qualifies. That selection does not change the callback's compiler target. Keep the native operations in appropriately targeted functions; the target-list guide generates those functions and the matching admission list from one declaration. Pointer or scalar entry parameters avoid passing vector registers across different calling conventions.

native::observe_cpu() returns the current platform's capability record. observe_x86_capabilities() and observe_arm_capabilities() remain available for architecture-specific code. The records distinguish typed present and observed feature sets: an unobserved requirement is unknown, not confirmed absent. Admission requires every feature to be both observed and present. X86 also requires the enabled XCR0 register state needed by the selected ISA. The nested raw fields retain diagnostics; editing them does not update the normalized feature sets. Fill the typed sets explicitly in synthetic native snapshots. Raw capability snapshots passed to classify_isa use the decoder.

ARM crypto hardware features are independent even where Clang enables a bundle: its aes target includes AES and PMULL, for example. Admit the whole compiler target before entering that function. The typed arm_feature::ebf16 capability records enhanced BF16 support; the application must also set FPCR.EBF to use that arithmetic. It is not a standalone Clang target string. The instruction guide links the exact requirements and numerical contracts for each family.

WebAssembly uses a separate feature family and observes a particular engine. The detector guide explains validation probes, unknown observations and admission for separately compiled bodies. Native CPU support cannot establish a Wasm engine's capabilities.

Imports and build targets

Most applications link native::native and import native. Use a narrower module when its boundary is useful:

Need Import CMake target
SIMD, masks and register operations native.simd native::native
A vector instruction family Its native.x86.*, native.arm.* or native.wasm.* module native::native
Promoted numerical kernels native.math native::native
CPU observation and admission native.features native::minimal
ISA values without an observer native.isa native::minimal
Scalar instruction utilities The corresponding family module native::minimal
Scalar numerics, generic packs and utilities native.numerics, native.wide, native.types, native.memory, native.static_string native::minimal

The architecture hubs native.x86, native.arm and native.wasm include vector instruction families and therefore belong to native::native. native::common is an alias for native::minimal. Vector instruction modules import native.simd and use simd<T,N,Arch> in their public interfaces. Scalar forms keep ordinary C++ values. Intrinsic register types are private implementation details of those instruction interfaces.

Set NATIVE_MINIMAL_COMPILE_OPTIONS during project setup to choose a stronger package minimum. Otherwise the toolchain baseline is unchanged. Linking the common target propagates its configured minimum to consumers; the process must already satisfy it before executing the runtime selector.

The tested toolchain is Clang 23, CMake 4.4 and Ninja. Configuration checks structured-binding packs, properties and deducing this. ISA properties and named swizzles use Clang's __declspec(property) extension, so C++26 support alone is insufficient. native::headers supplies -fms-extensions for Clang's GNU-style driver; clang-cl enables it already.

Include <native/attributes.h> for named compiler modifiers and <native/targets.h> for source target macros. Modules do not export macros. Installed module sources and build metadata let CMake regenerate compatible BMIs. Provider modules compile without PCHs; consumer PCHs must agree with their translation unit's compiler, exception and preprocessing settings. See building and installation for package configuration.

Clang 23 can warn about ambiguous internal linkage if <native/isa.h> is included before importing a module that exposes it. Import first when mixing the header and modules, or use A.has(feature) for feature checks. Module-only property reads and writes are checked with warnings treated as errors. Direct property expressions in constraints also have a Clang mangling limitation; use has or a named concept there.

Extending the element type

simd_traits<T> identifies a custom element's raw storage type. simd_customization<T,Raw,Self> supplies its value semantics. The common simd specialization instantiates that customization for the selected raw register. Existing raw float/integer/mask specializations remain direct implementations.

An extension defines the arithmetic semantics of its custom element. For example, FTZ supplies normalization, reproducible math and environment checks, while using this library's raw registers, masks and arrays. FTZ depends on native; native does not depend on FTZ. Every ISA family can use the same scalar type.

native::mask<T> names the associated mask after removing T's cv/ref qualifiers. Arithmetic scalars, fp16 and bf16 map to bool; SIMD values map to V::mask_type; mask lanes and predicates map to themselves. Standard arrays preserve their shape: mask<std::array<T,N>> is std::array<mask<T>,N>, including nested and empty arrays. This type mapping does not change array comparison operators.

Import native.types or include <native/mask_traits.h> for the scalar and array trait. SIMD and numerical imports add their type specializations. A consumer can specialize native::mask_traits<MyType> with a type member to define its own mask. Unsupported types have no type, so generic code can test requires { typename native::mask<T>; } without assuming every type has a mask.

Keep dependent mathematical calls unqualified so ADL can find the element's overloads. Native intrinsic interoperation is available through to_native() and from_native() when needed: include the platform intrinsic header before importing and compile the containing function for its instructions. Ordinary instruction-family calls already take simd values and need no such conversion.