native 0.0.1
Vectors, masks and wide register packs for C++26
Loading...
Searching...
No Matches
AVX-NE-CONVERT
Collaboration diagram for AVX-NE-CONVERT:

Functions

template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
constexpr simd< float, N, Arch > native::bcstnebf16_ps (bf16 const *source) noexcept
template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
constexpr simd< float, N, Arch > native::bcstnesh_ps (fp16 const *source) noexcept
template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
constexpr simd< float, N, Arch > native::cvtneebf16_ps (bf16 const *source) noexcept
template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
constexpr simd< float, N, Arch > native::cvtneeph_ps (fp16 const *source) noexcept
template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
constexpr simd< float, N, Arch > native::cvtneobf16_ps (bf16 const *source) noexcept
template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
constexpr simd< float, N, Arch > native::cvtneoph_ps (fp16 const *source) noexcept
template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
constexpr simd< bf16, N, Arch > native::cvtneps_bf16 (simd< float, N, Arch > input) noexcept

Detailed Description

Binary16/BF16 memory widening and binary32-to-BF16 narrowing, with no FP exceptions. Runtime calls require AVX-NE-CONVERT, AVX, matching compiler targets and OS vector state. Constant-only overloads are available for weaker ISA tags with the required storage. BF16 widening preserves every representation, including subnormals and signaling NaNs. FP16 widening is exact for finite inputs and quiets NaNs; narrowing uses RNE, DAZ and FTZ. MXCSR is neither consulted nor updated by these instructions.

Function Documentation

◆ bcstnebf16_ps()

template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
simd< float, N, Arch > native::bcstnebf16_ps ( bf16 const * source)
inlineconstexprexportnoexcept

Read one BF16 object and broadcast its binary32 representation. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.

Definition at line 69 of file native.x86.avxneconvert.ccm.

Here is the caller graph for this function:

◆ bcstnesh_ps()

template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
simd< float, N, Arch > native::bcstnesh_ps ( fp16 const * source)
inlineconstexprexportnoexcept

Read one binary16 object, widen to binary32 and broadcast. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.

Definition at line 94 of file native.x86.avxneconvert.ccm.

Here is the caller graph for this function:

◆ cvtneebf16_ps()

template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
simd< float, N, Arch > native::cvtneebf16_ps ( bf16 const * source)
inlineconstexprexportnoexcept

Read 2*N BF16 objects and widen the even-indexed elements. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.

Definition at line 119 of file native.x86.avxneconvert.ccm.

Here is the caller graph for this function:

◆ cvtneeph_ps()

template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
simd< float, N, Arch > native::cvtneeph_ps ( fp16 const * source)
inlineconstexprexportnoexcept

Read 2*N binary16 objects and widen the even-indexed elements. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.

Definition at line 146 of file native.x86.avxneconvert.ccm.

Here is the caller graph for this function:

◆ cvtneobf16_ps()

template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
simd< float, N, Arch > native::cvtneobf16_ps ( bf16 const * source)
inlineconstexprexportnoexcept

Read 2*N BF16 objects and widen the odd-indexed elements. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.

Definition at line 173 of file native.x86.avxneconvert.ccm.

Here is the caller graph for this function:

◆ cvtneoph_ps()

template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
simd< float, N, Arch > native::cvtneoph_ps ( fp16 const * source)
inlineconstexprexportnoexcept

Read 2*N binary16 objects and widen the odd-indexed elements. N is 4 or 8. The nonnull pointer needs only the element's natural alignment.

Definition at line 200 of file native.x86.avxneconvert.ccm.

Here is the caller graph for this function:

◆ cvtneps_bf16()

template<isa< x86 > Arch, std::size_t N>
requires (Arch.has(x86_feature::avxneconvert) && Arch.has(x86_feature::avx) && (N == 4 || N == 8))
simd< bf16, N, Arch > native::cvtneps_bf16 ( simd< float, N, Arch > input)
inlineconstexprexportnoexcept

Round N binary32 lanes to BF16 using nearest-even; N is 4 or 8. The four-lane result has four logical BF16 lanes and zero upper register bits.

Definition at line 227 of file native.x86.avxneconvert.ccm.

Here is the caller graph for this function: