native 0.0.1
Vectors, masks and wide register packs for C++26
Loading...
Searching...
No Matches
VPCLMULQDQ
Collaboration diagram for VPCLMULQDQ:

Functions

template<isa< x86 > Arch, unsigned Imm8>
requires (Arch.has(x86_feature::pclmul) && Arch.has(x86_feature::avx) && Imm8 <= 255)
constexpr simd< std::uint64_t, 2, Arch > native::vpclmulqdq (simd< std::uint64_t, 2, Arch > a, simd< std::uint64_t, 2, Arch > b) noexcept
 Multiply selected halves of one 128-bit lane using PCLMUL and AVX.
template<isa< x86 > Arch, unsigned Imm8>
requires (Arch.has(x86_feature::vpclmulqdq) && Arch.has(x86_feature::avx) && Imm8 <= 255)
constexpr simd< std::uint64_t, 4, Arch > native::vpclmulqdq (simd< std::uint64_t, 4, Arch > a, simd< std::uint64_t, 4, Arch > b) noexcept
 Multiply selected halves independently in two 128-bit lanes; AVX suffices.
template<isa< x86 > Arch, unsigned Imm8>
requires (Arch.has(x86_feature::vpclmulqdq) && Arch.has(x86_feature::avx512f) && Imm8 <= 255)
constexpr simd< std::uint64_t, 8, Arch > native::vpclmulqdq (simd< std::uint64_t, 8, Arch > a, simd< std::uint64_t, 8, Arch > b) noexcept
 Multiply selected halves independently in four 128-bit lanes; needs AVX512F.
template<isa< x86 > Arch, unsigned Imm8, class... Args>
void native::vpclmulqdq (Args...)=delete
 Reject unsupported signatures, including implicit raw-register conversions.

Detailed Description

Exact carry-less 64-by-64 multiplication within each 128-bit lane. Imm8 bit 0 selects a's half and bit 4 selects b's half in every lane. Other bits are ignored; the immediate must be in [0,255]. Each product occupies its original 128-bit lane, with bit 127 zero. There are no cross-lane products, carries, or polynomial reduction. The 128-bit intrinsic needs PCLMUL and AVX, not the VPCLMULQDQ feature. The 256-bit intrinsic needs VPCLMULQDQ and AVX, without AVX2 or AVX512VL. The 512-bit intrinsic also needs AVX512F, without AVX512BW/DQ/VL. Arch records requirements; callers separately enable and admit the target. All forms are pure integer computations with no floating-point effects. Constant evaluation uses exact integer semantics. Tags without the instruction features are accepted only at compile time and require complete SIMD storage.