native 0.0.1
Vectors, masks and wide register packs for C++26
Loading...
Searching...
No Matches
F16C
Collaboration diagram for F16C:

Functions

template<isa< x86 > Arch, unsigned Imm8>
requires (Arch.has(x86_feature::f16c) && Imm8 <= 255)
constexpr simd< fp16, 4, Arch > native::cvtps_ph (simd< float, 4, Arch > a) noexcept
 Convert 4 binary32 lanes to 4 binary16 lanes.
template<isa< x86 > Arch, unsigned Imm8>
requires (Arch.has(x86_feature::f16c) && Imm8 <= 255)
constexpr simd< fp16, 8, Arch > native::cvtps_ph (simd< float, 8, Arch > a) noexcept
 Convert 8 binary32 lanes to 8 binary16 lanes.
template<isa< x86 > Arch, unsigned Lanes>
requires (Arch.has(x86_feature::f16c) && Lanes == 4)
constexpr simd< float, 4, Arch > native::cvtph_ps (simd< fp16, 4, Arch > a) noexcept
 Widen 4 binary16 lanes to 4 binary32 lanes.
template<isa< x86 > Arch, unsigned Lanes>
requires (Arch.has(x86_feature::f16c) && Lanes == 8)
constexpr simd< float, 8, Arch > native::cvtph_ps (simd< fp16, 8, Arch > a) noexcept
 Widen 8 binary16 lanes to 8 binary32 lanes.
template<isa< x86 > Arch, unsigned Imm8>
requires (Arch.has(x86_feature::f16c) && Imm8 <= 255)
constexpr std::uint16_t native::cvtss_sh (float a) noexcept
 Convert one binary32 value to binary16 representation bits.
template<isa< x86 > Arch = NATIVE_BASELINE>
requires (Arch.has(x86_feature::f16c))
constexpr float native::cvtsh_ss (std::uint16_t a) noexcept
 Widen one binary16 representation to binary32.
template<isa< x86 > Arch, unsigned Imm8, class V>
void native::cvtps_ph (V)=delete
 Reject unsupported signatures, including implicit raw-register conversions.
template<isa< x86 > Arch, unsigned Lanes, class V>
void native::cvtph_ps (V)=delete
 Reject unsupported signatures, including implicit raw-register conversions.

Detailed Description

Binary32 / IEEE binary16 conversions using VEX VCVTPS2PH and VCVTPH2PS. Arch must contain f16c; its compiler prerequisite is AVX. Callers must enable a matching target and admit CPU support and XMM/YMM OS state separately. No AVX2 or AVX512FP16 instructions are required.

Packed operands and results use simd<float,N,Arch> and simd<fp16,N,Arch>. Half vectors preserve representation bits; no scalar numerical conversion is used at the vector boundary. Scalar forms use float and uint16_t half bits.

Imm8 accepts every byte, 0..255. Bit 2 selects MXCSR.RC; otherwise bits 1:0 select nearest-even (0), down (1), up (2), or toward zero (3). Bits 7:3 are ignored by the instruction: in particular, bit 3 does NOT suppress exceptions. Narrowing ignores FTZ and honors DAZ for binary32 subnormal inputs. Widening ignores DAZ and does not raise a denormal exception for binary16 subnormals. Signs of zero and infinity are preserved. NaNs retain their sign and high payload bits and are quieted; signaling NaNs raise invalid.

These operations retain architectural MXCSR status updates and unmasked exceptions even if their result is discarded. They do not modify control bits or masks. Runtime calls are neither const nor pure. Constant evaluation uses masked exceptions, gradual inputs/results, and nearest-even for the current-rounding immediate; no status effects occur. An ISA without F16C admits only the consteval forms.