|
native 0.0.1
Vectors, masks and wide register packs for C++26
|
Functions | |
|
template<isa< x86 > Arch, unsigned Imm8> requires (Arch.has(x86_feature::pclmul) && Arch.has(x86_feature::avx) && Imm8 <= 255) | |
| constexpr simd< std::uint64_t, 2, Arch > | native::vpclmulqdq (simd< std::uint64_t, 2, Arch > a, simd< std::uint64_t, 2, Arch > b) noexcept |
| Multiply selected halves of one 128-bit lane using PCLMUL and AVX. | |
|
template<isa< x86 > Arch, unsigned Imm8> requires (Arch.has(x86_feature::vpclmulqdq) && Arch.has(x86_feature::avx) && Imm8 <= 255) | |
| constexpr simd< std::uint64_t, 4, Arch > | native::vpclmulqdq (simd< std::uint64_t, 4, Arch > a, simd< std::uint64_t, 4, Arch > b) noexcept |
| Multiply selected halves independently in two 128-bit lanes; AVX suffices. | |
|
template<isa< x86 > Arch, unsigned Imm8> requires (Arch.has(x86_feature::vpclmulqdq) && Arch.has(x86_feature::avx512f) && Imm8 <= 255) | |
| constexpr simd< std::uint64_t, 8, Arch > | native::vpclmulqdq (simd< std::uint64_t, 8, Arch > a, simd< std::uint64_t, 8, Arch > b) noexcept |
| Multiply selected halves independently in four 128-bit lanes; needs AVX512F. | |
| template<isa< x86 > Arch, unsigned Imm8, class... Args> | |
| void | native::vpclmulqdq (Args...)=delete |
| Reject unsupported signatures, including implicit raw-register conversions. | |
Exact carry-less 64-by-64 multiplication within each 128-bit lane. Imm8 bit 0 selects a's half and bit 4 selects b's half in every lane. Other bits are ignored; the immediate must be in [0,255]. Each product occupies its original 128-bit lane, with bit 127 zero. There are no cross-lane products, carries, or polynomial reduction. The 128-bit intrinsic needs PCLMUL and AVX, not the VPCLMULQDQ feature. The 256-bit intrinsic needs VPCLMULQDQ and AVX, without AVX2 or AVX512VL. The 512-bit intrinsic also needs AVX512F, without AVX512BW/DQ/VL. Arch records requirements; callers separately enable and admit the target. All forms are pure integer computations with no floating-point effects. Constant evaluation uses exact integer semantics. Tags without the instruction features are accepted only at compile time and require complete SIMD storage.