native 0.0.1
Vectors, masks and wide register packs for C++26
Loading...
Searching...
No Matches
ARM BF16: dot products and matrix accumulation

ARM instruction sets

Why use it

BF16 keeps the eight-bit exponent of FP32 in a sixteen-bit representation, trading fraction precision for smaller inputs. These instructions multiply BF16 values and accumulate into FP32, giving neural-network and other dense linear-algebra kernels a direct dot-product or small matrix-tile operation.

Operations

Import native.arm.bf16, or use the native.arm or native hub. All operands use native::simd<T,N,Arch> with a common Arch.

Operation Accumulator/result Inputs Selection
bfdot<Arch>(acc,a,b) simd<float,2,Arch> or simd<float,4,Arch> Matching-width simd<bf16,4,Arch> or simd<bf16,8,Arch> Adjacent pairs
bfdot_lane<Arch,Lane>(acc,a,b) Either dot shape Matching-width a; either width for b Pair b[2*Lane], b[2*Lane+1]
bfmmla<Arch>(acc,a,b) simd<float,4,Arch> Two simd<bf16,8,Arch> vectors 2×4 by 4×2 matrix product
bfmlalb<Arch>(acc,a,b) simd<float,4,Arch> Two simd<bf16,8,Arch> vectors Even lanes of both inputs
bfmlalt<Arch>(acc,a,b) simd<float,4,Arch> Two simd<bf16,8,Arch> vectors Odd lanes of both inputs
bfmlalb_lane<Arch,Lane>(acc,a,b) simd<float,4,Arch> Eight BF16 lanes in a; four or eight in b Even lanes of a, scalar b[Lane]
bfmlalt_lane<Arch,Lane>(acc,a,b) simd<float,4,Arch> Eight BF16 lanes in a; four or eight in b Odd lanes of a, scalar b[Lane]

Dot-product indices select pairs: 0–1 for a four-element b, 0–3 for an eight-element b. Widening multiply-add indices select elements: 0–3 or 0–7. All lane indices are compile-time immediates.

For bfmmla, store the left matrix's rows and the right matrix's columns consecutively. Output lane 2*r+c starts with acc[2*r+c], accumulates the pair of products at inner indices 0 and 1, then the pair at indices 2 and 3. Each step follows the bfdot arithmetic contract. A row-major 4×2 right matrix needs rearrangement.

#include <native/targets.h>
#define NATIVE_TARGET_bfloat_dot "bf16"
constexpr auto dot_isa = NATIVE_TARGET_ISA(bfloat_dot);
NATIVE_TARGET_PUSH(bfloat_dot)
void accumulate(float* output, float const* acc,
native::bf16 const* a, native::bf16 const* b) noexcept {
native::bfdot<dot_isa>(F::load(acc), B::load(a), B::load(b)).store(output);
}
NATIVE_TARGET_POP()
constexpr simd< float, 2, Arch > bfdot(simd< float, 2, Arch > acc, simd< bf16, 4, Arch > a, simd< bf16, 4, Arch > b) noexcept
Accumulate each adjacent pair of BF16 products into the corresponding FP32 lane.
Omitted architecture arguments use the native.simd provider's baseline.

Caveats

Runtime calls require arm_feature::neon_bf16 and a "bf16" caller target. Admission includes NEON. BF16 is independent of FP16, DotProd and I8MM. Admit the target requirement and NATIVE_TARGET_MINIMUM before entering the function. Importing the module does not enable instructions or perform dispatch. The wrappers add no BF16 elementwise arithmetic or conversion policy.

With FEAT_EBF16 absent or FPCR.EBF clear, bfdot and bfmmla round each product, pair sum and accumulator addition separately to odd: an inexact finite result has its low bit set. Subnormal inputs and tiny results flush to signed zero, independently of RMode, FZ and FIZ. Overflow produces signed infinity. NaNs become default NaNs; AH affects their sign.

With FEAT_EBF16 present and FPCR.EBF set, each pair is a fused sum of two products, rounded once to FP32 before a separate accumulator addition. RMode, FZ, AH and FIZ control single-precision behavior. Matrix multiplication still performs two ordered pair accumulations per output. Neither mode is equivalent to a chain of ordinary FP32 fused multiply-adds.

In both modes, dot and matrix instructions ignore exception enables and leave cumulative FPSR flags unchanged. They do not write FPCR. Check arm_feature::ebf16 before setting EBF; the feature tag does not set that control. Clang 23 has no standalone "ebf16" target string; enhanced arithmetic uses the same BF16 instructions.

bfmlalb and bfmlalt each compute one fused multiply-add per output, rounded to FP32. With AH clear, they follow ordinary FP32 rounding, flushing and NaN controls and accumulate exception flags. With FEAT_AFP and AH set, they force nearest-even rounding and input/output flushing and suppress exceptions. EBF does not change these operations. noexcept does not mask hardware traps.

The wrappers retain FPCR-sensitive instruction execution and FPSR effects even when results are discarded. A compiler memory barrier orders surrounding environment accesses and adds no CPU memory fence. Surrounding code still needs the compiler's floating-environment support when it observes or changes controls.

Constant evaluation fixes nearest-even rounding, gradual inputs and results, payload-preserving NaNs, IEEE half precision and masked exceptions, with DN=AH=AHP=FZ=FZ16=FIZ=EBF=0 and no machine status effects. The fixed BFDOT/BFMMLA rules override this: round-to-odd steps, flushing, infinity on overflow and positive default NaNs. BFMLALB/T use the ordinary AH=0 fused result, preserving NaN payloads and signs. A tag without BF16 permits only consteval calls when the SIMD storage types exist; it provides no runtime fallback.

See Arm's BF16 instruction overview, SME supplement, B3.1.2 and E2.2, BFMMLA and BFMLAL shared pseudocode and ACLE BF16 types.