|
native 0.0.1
Vectors, masks and wide register packs for C++26
|
Public Types | |
| using | value_type = fp16 |
| Scalar storage element; each lane retains all 16 representation bits. | |
| using | register_type = simd |
| This one-register vector type, for generic register-based algorithms. | |
| using | native_type = __m512h |
| Native 512-bit FP16 register representation; native bridges copy bits. | |
| using | bits_type = simd<std::uint16_t,32,architecture> |
| Unsigned 16-bit lanes in the same profile and lane order. | |
| using | mask = predicate<32,architecture> |
| Compact predicate with bit i selecting FP16 lane i. | |
| using | mask_type = mask |
| Generic mask spelling for the compact lane predicate. | |
| using | predicate_type = mask |
| Explicit predicate spelling for the same compact predicate type. | |
| using | vector_mask_type = simd<mask16,32,architecture> |
| Full-register mask shape with zero or all-one 16-bit lanes. | |
| template<class T> | |
| using | rebind = simd<T,32,architecture> |
Public Member Functions | |
| simd () noexcept=default | |
| Default initialization leaves storage unspecified; braces zero it. | |
| constexpr | simd (fp16 value) noexcept |
| Broadcast the exact representation of value to all 32 lanes. | |
| constexpr | simd (std::array< fp16, lanes > const &values) noexcept |
| Copy array element i into lane i without conversion or representation changes. | |
|
template<class... T> requires (sizeof...(T) == lanes && (std::same_as<T,fp16> && ...)) | |
| constexpr | simd (T... values) noexcept |
| Construct all 32 lanes from FP16 values in argument order, preserving their bits. | |
| constexpr | simd (native_type value) noexcept |
| Adopt a native register without conversion or representation changes. | |
| constexpr | operator native_type () const noexcept |
| Project the native register for direct intrinsic interoperability. | |
| constexpr native_type | to_native () const noexcept |
| Return all lane bits as a native register, without conversion or lane reordering. | |
| constexpr bits_type | bits () const noexcept |
| Return the 16-bit representation of each lane in an unsigned vector. | |
| constexpr bits_type | to_bits () const noexcept |
| Synonym for bits(); this is a representation bridge, not a numeric conversion. | |
| constexpr void | store_bits (std::uint16_t *p) const noexcept |
| template<std::size_t Alignment = 1> | |
| constexpr void | store_memory (fp16 *p) const noexcept |
| constexpr void | store (fp16 *p) const noexcept |
| Store 32 FP16 objects with the default alignment contract of store_memory(). | |
| constexpr void | storeu (fp16 *p) const noexcept |
| Synonym for store(); no register-width alignment is required. | |
| constexpr void | store_partial (fp16 *p, std::size_t n) const noexcept |
Static Public Member Functions | |
| static constexpr simd | from_native (native_type value) noexcept |
| Copy a native FP16 register into this vector, preserving every representation bit. | |
| static constexpr simd | from_bits (bits_type value) noexcept |
| Interpret each unsigned lane as a FP16 representation without changing its bits. | |
| static constexpr simd | load_bits (std::uint16_t const *p) noexcept |
| template<std::size_t Alignment = 1> | |
| static constexpr simd | load_memory (fp16 const *p) noexcept |
| static constexpr simd | load (fp16 const *p) noexcept |
| Load 32 FP16 objects with the default alignment contract of load_memory(). | |
| static constexpr simd | loadu (fp16 const *p) noexcept |
| Synonym for load(); no register-width alignment is required. | |
| static constexpr simd | load_partial (fp16 const *p, std::size_t n, fp16 fill=fp16::from_bits(0)) noexcept |
Static Public Attributes | |
| static constexpr isa | architecture =Arch |
| The distinct compile-time AVX512_FP16 instruction profile. | |
| static constexpr std::size_t | lanes = 32 |
| Number of logical FP16 lanes, with no padding lanes. | |
Friends | |
| constexpr simd | operator+ (simd a, simd b) noexcept |
| Add corresponding half lanes, rounding directly under the caller's MXCSR rounding control. | |
| constexpr simd | operator- (simd a, simd b) noexcept |
| Subtract corresponding half lanes, rounding directly under the caller's MXCSR rounding control. | |
| constexpr simd | operator* (simd a, simd b) noexcept |
| Multiply corresponding half lanes, rounding directly under the caller's MXCSR rounding control. | |
| constexpr simd | operator/ (simd a, simd b) noexcept |
| constexpr simd | sqrt (simd a) noexcept |
| constexpr simd | operator- (simd a) noexcept |
| Toggle every sign bit, preserving all payload bits without arithmetic exceptions. | |
| constexpr mask | operator== (simd a, simd b) noexcept |
| constexpr mask | operator!= (simd a, simd b) noexcept |
| Lane inequality, true for unordered NaN operands; complements native equality. | |
| constexpr mask | operator< (simd a, simd b) noexcept |
| Ordered lane less-than. NaNs compare false; native MXCSR exception semantics apply. | |
| constexpr mask | operator<= (simd a, simd b) noexcept |
| Ordered lane less-or-equal. NaNs compare false; native MXCSR exception semantics apply. | |
| constexpr mask | operator> (simd a, simd b) noexcept |
| Ordered lane greater-than, with the native less-than operands reversed. | |
| constexpr mask | operator>= (simd a, simd b) noexcept |
| Ordered lane greater-or-equal, with native less-or-equal operands reversed. | |
| constexpr simd | select (mask m, simd a, simd b) noexcept |
One 512-bit register of FP16 representations. Loads, stores and bit bridges preserve every encoding. Native addition, subtraction, multiplication, division, square root and fused multiply-add round directly to half precision under the caller's MXCSR rounding control. Half operands/results use gradual underflow regardless of MXCSR.DAZ/FTZ. MXCSR rounding and exception controls apply; status flags may change. No operation changes MXCSR control bits. NaN payload/sign propagation follows the instruction, not a portable promise. Scalar fp16 conversions are unchanged. Only the 32-lane shape is provided. The application must admit that CPU/OS profile before entering compiled code. Every storage operation preserves subnormal, signed-zero and NaN encodings; none performs a floating-point conversion or quiets a signaling NaN.
Definition at line 214 of file native.simd.ccm.
| using native::simd< fp16, 32, Arch >::rebind = simd<T,32,architecture> |
Replace the element type while retaining 32 lanes and this profile. Unsupported resulting shapes remain incomplete.
Definition at line 235 of file native.simd.ccm.
|
inlinestaticconstexprnoexcept |
Read exactly 32 accessible uint16_t objects into corresponding FP16 lane bits. No alignment beyond that of uint16_t is required; p must not be null.
Definition at line 273 of file native.simd.ccm.
|
inlinestaticconstexprnoexcept |
Read exactly 32 accessible FP16 objects, preserving every encoding. Alignment is a nonzero power-of-two byte-alignment promise, not a runtime check. The default imposes no alignment beyond that required for FP16 objects. p must not be null.
Definition at line 284 of file native.simd.ccm.
|
inlinestaticconstexprnoexcept |
Read exactly the first n accessible FP16 objects, where n <= 32. Copy their representations to lanes [0,n); remaining lanes receive fill's exact representation. The default fill is positive zero. No access occurs for n == 0, when p may be null; otherwise p must address n FP16 objects.
Definition at line 319 of file native.simd.ccm.
|
inlineconstexprnoexcept |
Write every lane representation to 32 accessible uint16_t objects in lane order. No alignment beyond that of uint16_t is required; p must not be null.
Definition at line 278 of file native.simd.ccm.
|
inlineconstexprnoexcept |
Write all lane representations to exactly 32 accessible FP16 objects. Alignment is a nonzero power-of-two byte-alignment promise, not a runtime check. The default imposes no alignment beyond that required for FP16 objects. p must not be null.
Definition at line 298 of file native.simd.ccm.
|
inlineconstexprnoexcept |
Write the representations of lanes [0,n) to exactly n accessible FP16 objects, where n <= 32. Memory outside that prefix is untouched. No access occurs for n == 0, when p may be null; otherwise p must address n objects.
Definition at line 331 of file native.simd.ccm.
|
friend |
Divide corresponding half lanes with native half-precision rounding. The caller's MXCSR rounding control and native exception behavior apply.
Definition at line 355 of file native.simd.ccm.
|
friend |
Ordered lane equality. NaNs compare false; signed zeros compare equal. DAZ does not flush half operands. Signaling NaNs raise native invalid status.
Definition at line 373 of file native.simd.ccm.
|
friend |
Choose a lane from a when its canonical mask lane is true, otherwise b. Selection copies every representation bit without arithmetic or NaN quieting.
Definition at line 395 of file native.simd.ccm.
|
friend |
Compute each half lane's square root with native half-precision rounding. Signed zero is preserved; negative nonzero operands produce a quiet NaN. The caller's MXCSR rounding control and native exception behavior apply.
Definition at line 362 of file native.simd.ccm.