native 0.0.1
Vectors, masks and wide register packs for C++26
Loading...
Searching...
No Matches
native::simd< fp16, 8, Arch > Struct Template Referenceexport
Collaboration diagram for native::simd< fp16, 8, Arch >:
[legend]

Public Types

using value_type = fp16
 Scalar storage element; each lane retains all 16 representation bits.
using register_type = simd
 This one-register vector type, for generic register-based algorithms.
using native_type = float16x8_t
 Native 128-bit FP16 register representation; native bridges copy bits.
using bits_type = simd<std::uint16_t,8,architecture>
 Unsigned 16-bit lanes in the same profile and lane order.
using mask = simd<mask16,8,architecture>
 Full-register predicate with zero or all-one bits in each 16-bit lane.
using mask_type = mask
 Generic mask spelling for the full-register lane mask.
using predicate_type = mask
 Explicit predicate spelling for the same full-register mask type.
using vector_mask_type = simd<mask16,8,architecture>
 Full-register mask shape with zero or all-one 16-bit lanes.
template<class T>
using rebind = simd<T,8,architecture>

Public Member Functions

 simd () noexcept=default
 Default initialization leaves storage unspecified; braces zero it.
constexpr simd (fp16 value) noexcept
 Broadcast the exact representation of value to all 8 lanes.
constexpr simd (std::array< fp16, lanes > const &values) noexcept
 Copy array element i into lane i without conversion or representation changes.
template<class... T>
requires (sizeof...(T) == lanes && (std::same_as<T,fp16> && ...))
constexpr simd (T... values) noexcept
 Construct all 8 lanes from FP16 values in argument order, preserving their bits.
constexpr simd (native_type value) noexcept
 Adopt a native register without conversion or representation changes.
constexpr operator native_type () const noexcept
 Project the native register for direct intrinsic interoperability.
constexpr native_type to_native () const noexcept
 Return all lane bits as a native register, without conversion or lane reordering.
constexpr bits_type bits () const noexcept
 Return the 16-bit representation of each lane in an unsigned vector.
constexpr bits_type to_bits () const noexcept
 Synonym for bits(); this is a representation bridge, not a numeric conversion.
constexpr void store_bits (std::uint16_t *p) const noexcept
template<std::size_t Alignment = 1>
constexpr void store_memory (fp16 *p) const noexcept
constexpr void store (fp16 *p) const noexcept
 Store 8 FP16 objects with the default alignment contract of store_memory().
constexpr void storeu (fp16 *p) const noexcept
 Synonym for store(); no register-width alignment is required.
constexpr void store_partial (fp16 *p, std::size_t n) const noexcept

Static Public Member Functions

static constexpr simd from_native (native_type value) noexcept
 Copy a native FP16 register into this vector, preserving every representation bit.
static constexpr simd from_bits (bits_type value) noexcept
 Interpret each unsigned lane as a FP16 representation without changing its bits.
static constexpr simd load_bits (std::uint16_t const *p) noexcept
template<std::size_t Alignment = 1>
static constexpr simd load_memory (fp16 const *p) noexcept
static constexpr simd load (fp16 const *p) noexcept
 Load 8 FP16 objects with the default alignment contract of load_memory().
static constexpr simd loadu (fp16 const *p) noexcept
 Synonym for load(); no register-width alignment is required.
static constexpr simd load_partial (fp16 const *p, std::size_t n, fp16 fill=fp16::from_bits(0)) noexcept

Static Public Attributes

static constexpr isa architecture =Arch
 The distinct compile-time NEON_FP16 instruction profile.
static constexpr std::size_t lanes = 8
 Number of logical FP16 lanes, with no padding lanes.

Friends

constexpr simd operator+ (simd a, simd b) noexcept
 Add corresponding half lanes, rounding directly under the caller's FPCR.
constexpr simd operator- (simd a, simd b) noexcept
 Subtract corresponding half lanes, rounding directly under the caller's FPCR.
constexpr simd operator* (simd a, simd b) noexcept
 Multiply corresponding half lanes, rounding directly under the caller's FPCR.
constexpr simd operator/ (simd a, simd b) noexcept
constexpr simd sqrt (simd a) noexcept
constexpr simd operator- (simd a) noexcept
 Apply native FNEG to each half lane; no scalar half-to-float conversion occurs.
constexpr mask operator== (simd a, simd b) noexcept
constexpr mask operator!= (simd a, simd b) noexcept
 Lane inequality, true for unordered NaN operands; complements native equality.
constexpr mask operator< (simd a, simd b) noexcept
 Ordered lane less-than. NaNs compare false; native FPCR/FPSR semantics apply.
constexpr mask operator<= (simd a, simd b) noexcept
 Ordered lane less-or-equal. NaNs compare false; native FPCR/FPSR semantics apply.
constexpr mask operator> (simd a, simd b) noexcept
 Ordered lane greater-than, with the native less-than operands reversed.
constexpr mask operator>= (simd a, simd b) noexcept
 Ordered lane greater-or-equal, with native less-or-equal operands reversed.
constexpr simd select (mask m, simd a, simd b) noexcept

Detailed Description

template<::native::isa<> Arch>
requires (::native::avx512 <= Arch )
struct native::simd< fp16, 8, Arch >

One 128-bit register of FP16 representations. Loads, stores and bit bridges preserve every encoding. Native addition, subtraction, multiplication, division, square root and fused multiply-add round directly to half precision under the caller's FPCR. FZ16 controls half subnormal inputs/results; rounding, DN and exception controls retain their architectural meaning. Operations may update FPSR. No operation changes FPCR. NaN payload/sign propagation is instruction- and FPCR-dependent, not a portable promise. Scalar fp16 conversions are unchanged. Only the 8-lane shape is provided. The application must admit that CPU/OS profile before entering compiled code. Every storage operation preserves subnormal, signed-zero and NaN encodings; none performs a floating-point conversion or quiets a signaling NaN.

Definition at line 611 of file native.simd.ccm.

Member Typedef Documentation

◆ rebind

template<::native::isa<> Arch>
template<class T>
using native::simd< fp16, 8, Arch >::rebind = simd<T,8,architecture>

Replace the element type while retaining 8 lanes and this profile. Unsupported resulting shapes remain incomplete.

Definition at line 632 of file native.simd.ccm.

Member Function Documentation

◆ load_bits()

template<::native::isa<> Arch>
constexpr simd native::simd< fp16, 8, Arch >::load_bits ( std::uint16_t const * p)
inlinestaticconstexprnoexcept

Read exactly 8 accessible uint16_t objects into corresponding FP16 lane bits. No alignment beyond that of uint16_t is required; p must not be null.

Definition at line 670 of file native.simd.ccm.

Here is the caller graph for this function:

◆ load_memory()

template<::native::isa<> Arch>
template<std::size_t Alignment = 1>
constexpr simd native::simd< fp16, 8, Arch >::load_memory ( fp16 const * p)
inlinestaticconstexprnoexcept

Read exactly 8 accessible FP16 objects, preserving every encoding. Alignment is a nonzero power-of-two byte-alignment promise, not a runtime check. The default imposes no alignment beyond that required for FP16 objects. p must not be null.

Definition at line 681 of file native.simd.ccm.

Here is the caller graph for this function:

◆ load_partial()

template<::native::isa<> Arch>
constexpr simd native::simd< fp16, 8, Arch >::load_partial ( fp16 const * p,
std::size_t n,
fp16 fill = fp16::from_bits(0) )
inlinestaticconstexprnoexcept

Read exactly the first n accessible FP16 objects, where n <= 8. Copy their representations to lanes [0,n); remaining lanes receive fill's exact representation. The default fill is positive zero. No access occurs for n == 0, when p may be null; otherwise p must address n FP16 objects.

Definition at line 716 of file native.simd.ccm.

◆ store_bits()

template<::native::isa<> Arch>
void native::simd< fp16, 8, Arch >::store_bits ( std::uint16_t * p) const
inlineconstexprnoexcept

Write every lane representation to 8 accessible uint16_t objects in lane order. No alignment beyond that of uint16_t is required; p must not be null.

Definition at line 675 of file native.simd.ccm.

Here is the caller graph for this function:

◆ store_memory()

template<::native::isa<> Arch>
template<std::size_t Alignment = 1>
void native::simd< fp16, 8, Arch >::store_memory ( fp16 * p) const
inlineconstexprnoexcept

Write all lane representations to exactly 8 accessible FP16 objects. Alignment is a nonzero power-of-two byte-alignment promise, not a runtime check. The default imposes no alignment beyond that required for FP16 objects. p must not be null.

Definition at line 695 of file native.simd.ccm.

Here is the caller graph for this function:

◆ store_partial()

template<::native::isa<> Arch>
void native::simd< fp16, 8, Arch >::store_partial ( fp16 * p,
std::size_t n ) const
inlineconstexprnoexcept

Write the representations of lanes [0,n) to exactly n accessible FP16 objects, where n <= 8. Memory outside that prefix is untouched. No access occurs for n == 0, when p may be null; otherwise p must address n objects.

Definition at line 728 of file native.simd.ccm.

◆ operator/

template<::native::isa<> Arch>
simd operator/ ( simd< fp16, 8, Arch > a,
simd< fp16, 8, Arch > b )
friend

Divide corresponding half lanes with native half-precision rounding. The caller's FPCR and native exception behavior apply.

Definition at line 752 of file native.simd.ccm.

◆ operator==

template<::native::isa<> Arch>
mask operator== ( simd< fp16, 8, Arch > a,
simd< fp16, 8, Arch > b )
friend

Ordered lane equality. NaNs compare false; signed zeros compare equal. FPCR half-denormal controls and native comparison exception behavior apply.

Definition at line 771 of file native.simd.ccm.

◆ select

template<::native::isa<> Arch>
simd select ( mask m,
simd< fp16, 8, Arch > a,
simd< fp16, 8, Arch > b )
friend

Choose a lane from a when its canonical mask lane is true, otherwise b. Selection copies every representation bit without arithmetic or NaN quieting.

Definition at line 793 of file native.simd.ccm.

◆ sqrt

template<::native::isa<> Arch>
simd sqrt ( simd< fp16, 8, Arch > a)
friend

Compute each half lane's square root with native half-precision rounding. Signed zero is preserved; negative nonzero operands after native input flushing produce a quiet NaN. The caller's FPCR and native exception behavior apply.

Definition at line 760 of file native.simd.ccm.


The documentation for this struct was generated from the following file: