|
native 0.0.1
Vectors, masks and wide register packs for C++26
|
Functions | |
| template<isa< x86 > Arch, std::size_t N> requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8)) | |
| constexpr simd< float, N, Arch > | native::bcstnebf16_ps (bf16 const *source) noexcept |
| template<isa< x86 > Arch, std::size_t N> requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8)) | |
| constexpr simd< float, N, Arch > | native::bcstnesh_ps (fp16 const *source) noexcept |
| template<isa< x86 > Arch, std::size_t N> requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8)) | |
| constexpr simd< float, N, Arch > | native::cvtneebf16_ps (bf16 const *source) noexcept |
| template<isa< x86 > Arch, std::size_t N> requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8)) | |
| constexpr simd< float, N, Arch > | native::cvtneeph_ps (fp16 const *source) noexcept |
| template<isa< x86 > Arch, std::size_t N> requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8)) | |
| constexpr simd< float, N, Arch > | native::cvtneobf16_ps (bf16 const *source) noexcept |
| template<isa< x86 > Arch, std::size_t N> requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8)) | |
| constexpr simd< float, N, Arch > | native::cvtneoph_ps (fp16 const *source) noexcept |
| template<isa< x86 > Arch, std::size_t N> requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8)) | |
| constexpr simd< bf16, N, Arch > | native::cvtneps_bf16 (simd< float, N, Arch > input) noexcept |
Binary16/BF16 memory widening and binary32-to-BF16 narrowing, with no FP exceptions. Runtime calls require AVX-NE-CONVERT, AVX, matching compiler targets and OS vector state. Constant-only overloads are available for weaker ISA tags with the required storage. BF16 widening preserves every representation, including subnormals and signaling NaNs. FP16 widening is exact for finite inputs and quiets NaNs; narrowing uses RNE, DAZ and FTZ. MXCSR is neither consulted nor updated by these instructions.
|
inlineconstexprexportnoexcept |
Read one BF16 object and broadcast its binary32 representation. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.
Definition at line 69 of file native.x86.avxneconvert.ccm.
|
inlineconstexprexportnoexcept |
Read one binary16 object, widen to binary32 and broadcast. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.
Definition at line 94 of file native.x86.avxneconvert.ccm.
|
inlineconstexprexportnoexcept |
Read 2*N BF16 objects and widen the even-indexed elements. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.
Definition at line 119 of file native.x86.avxneconvert.ccm.
|
inlineconstexprexportnoexcept |
Read 2*N binary16 objects and widen the even-indexed elements. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.
Definition at line 146 of file native.x86.avxneconvert.ccm.
|
inlineconstexprexportnoexcept |
Read 2*N BF16 objects and widen the odd-indexed elements. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.
Definition at line 173 of file native.x86.avxneconvert.ccm.
|
inlineconstexprexportnoexcept |
Read 2*N binary16 objects and widen the odd-indexed elements. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.
Definition at line 200 of file native.x86.avxneconvert.ccm.
|
inlineconstexprexportnoexcept |
Round N binary32 lanes to BF16 using nearest-even; N is 4 or 8. The four-lane result has four logical BF16 lanes and zero upper register bits.
Definition at line 227 of file native.x86.avxneconvert.ccm.