|
native 0.0.1
Vectors, masks and wide register packs for C++26
|
A 52-bit limb leaves twelve spare bits in a qword for accumulation. IFMA multiplies two such limbs and adds either half of their 104-bit product into 64-bit lanes. This is useful for batched multi-precision multiplication and modular arithmetic, where the algorithm controls carry propagation.
import native.x86.ifma; supplies operations on simd<std::uint64_t,N,Arch> with N = 2, 4, 8. Link native::native; native.x86 and native also export them.
| Operation | Per-lane result |
|---|---|
| madd52lo<Arch>(accumulator, a, b) | Add product bits 0β51 to the accumulator |
| madd52hi<Arch>(accumulator, a, b) | Add product bits 52β103 to the accumulator |
The unsigned product uses only the low 52 bits of each multiplicand. Every accumulator bit participates, and addition wraps modulo 2βΆβ΄.
Both names have mask_NAME<Arch>(accumulator, mask, a, b) to retain inactive accumulator lanes and maskz_NAME<Arch>(mask, accumulator, a, b) to clear them. Masks are predicate<N,Arch>.
There is no carry propagation between lanes, saturation or modular reduction. The caller must account for accumulated carries before a 64-bit lane wraps. The top twelve bits of multiplicands are ignored even when they contain a carry.
| Form | Runtime features | Caller target |
|---|---|---|
| Unmasked 128/256-bit VEX | AVX-IFMA, AVX | "avxifma" |
| 128/256-bit EVEX, all mask modes | AVX512IFMA, AVX512F, AVX512VL | "avx512ifma,avx512vl" |
| 512-bit EVEX, all mask modes | AVX512IFMA, AVX512F | "avx512ifma" |
Short unmasked forms prefer AVX-IFMA when the tag supports it. Masked runtime forms always need EVEX features. BW and DQ are unnecessary. AVX-IFMA and AVX512IFMA are independent features, absent from the general AVX2 and AVX-512 profiles. Clang's AVX-IFMA target also enables AVX2; use target_features to include the compiler prerequisites. Admit the complete target and OS vector state before entering a matching caller.
Feature-bearing overloads are constexpr with native runtime paths. Weaker tags have consteval overloads if the chosen SSE2, AVX or AVX512F register storage exists, with prerequisites. Constant evaluation gives exact integer results and preserves the tag.
See Intel's Software Developer's Manual and Clang's AVX-IFMA, AVX512IFMA and VL IFMA headers.