native 0.0.1
Vectors, masks and wide register packs for C++26
Loading...
Searching...
No Matches
native::simd< bf16, N, Arch > Struct Template Referenceexport
Collaboration diagram for native::simd< bf16, N, Arch >:
[legend]

Public Types

using value_type = bf16
 Scalar storage element; each lane retains all 16 representation bits.
using register_type = simd
 This one-register vector type, for generic register-based algorithms.
using native_type = std::conditional_t<N == 8,__m128bh,std::conditional_t<N == 16,__m256bh,__m512bh>>
 Native register-width BF16 register representation; native bridges copy bits.
using bits_type = simd<std::uint16_t,N,architecture>
 Unsigned 16-bit lanes in the same profile and lane order.
using mask = predicate<N,architecture>
 Compact predicate with bit i selecting BF16 lane i.
using mask_type = mask
 Generic mask spelling for the compact lane predicate.
using predicate_type = mask
 Explicit predicate spelling for the same compact mask type.
using vector_mask_type = simd<mask16,N,architecture>
 Full-register mask shape with zero or all-one 16-bit lanes.
template<class T>
using rebind = simd<T,N,architecture>

Public Member Functions

 simd () noexcept=default
 Default initialization leaves storage unspecified; braces zero it.
constexpr simd (bf16 value) noexcept
 Broadcast the exact representation of value to all N lanes.
constexpr simd (std::array< bf16, lanes > const &values) noexcept
 Copy array element i into lane i without conversion or representation changes.
template<class... T>
requires (sizeof...(T) == lanes && (std::same_as<T,bf16> && ...))
constexpr simd (T... values) noexcept
 Construct all N lanes from BF16 values in argument order, preserving their bits.
constexpr simd (native_type value) noexcept
 Adopt a native register without conversion or representation changes.
constexpr operator native_type () const noexcept
 Project the native register for direct intrinsic interoperability.
constexpr native_type to_native () const noexcept
 Return all lane bits as a native register, without conversion or lane reordering.
constexpr bits_type bits () const noexcept
 Return the 16-bit representation of each lane in an unsigned vector.
constexpr bits_type to_bits () const noexcept
 Synonym for bits(); this is a representation bridge, not a numeric conversion.
constexpr void store_bits (std::uint16_t *p) const noexcept
template<std::size_t Alignment = 1>
constexpr void store_memory (bf16 *p) const noexcept
constexpr void store (bf16 *p) const noexcept
 Store N BF16 objects with the default alignment contract of store_memory().
constexpr void storeu (bf16 *p) const noexcept
 Synonym for store(); no register-width alignment is required.
constexpr void store_partial (bf16 *p, std::size_t n) const noexcept

Static Public Member Functions

static constexpr simd from_native (native_type value) noexcept
 Copy a native BF16 register into this vector, preserving every representation bit.
static constexpr simd from_bits (bits_type value) noexcept
 Interpret each unsigned lane as a BF16 representation without changing its bits.
static constexpr simd load_bits (std::uint16_t const *p) noexcept
template<std::size_t Alignment = 1>
static constexpr simd load_memory (bf16 const *p) noexcept
static constexpr simd load (bf16 const *p) noexcept
 Load N BF16 objects with the default alignment contract of load_memory().
static constexpr simd loadu (bf16 const *p) noexcept
 Synonym for load(); no register-width alignment is required.
static constexpr simd load_partial (bf16 const *p, std::size_t n, bf16 fill=bf16::from_bits(0)) noexcept

Static Public Attributes

static constexpr isa architecture =Arch
 The distinct compile-time AVX512_BF16 instruction profile.
static constexpr std::size_t lanes = N
 Number of logical BF16 lanes, with no padding lanes.

Friends

void operator+ (simd)=delete
 Reject unary BF16 arithmetic; storage does not define it.
void operator- (simd)=delete
 Reject unary BF16 arithmetic; storage does not define it.
void operator~ (simd)=delete
 Reject native fallback bitwise operations; use bits() explicitly.
void operator! (simd)=delete
 Reject native fallback logical operations; storage is not a predicate.

Detailed Description

template<std::size_t N, ::native::isa<> Arch>
requires (::native::avx512 <= Arch ) &&(N == 8 || N == 16 || N == 32)
struct native::simd< bf16, N, Arch >

One 128-, 256-, or 512-bit register of BF16 representations. Loads, stores and bit bridges preserve every encoding. This does not add elementwise BF16 arithmetic or change the scalar bf16 conversion contract. The 8-, 16-, and 32-lane shapes are provided. The application must admit that CPU/OS profile before entering compiled code. Every storage operation preserves subnormal, signed-zero and NaN encodings; none performs a floating-point conversion or quiets a signaling NaN.

Definition at line 44 of file native.simd.ccm.

Member Typedef Documentation

◆ rebind

template<std::size_t N, ::native::isa<> Arch>
template<class T>
using native::simd< bf16, N, Arch >::rebind = simd<T,N,architecture>

Replace the element type while retaining N lanes and this profile. Unsupported resulting shapes remain incomplete.

Definition at line 65 of file native.simd.ccm.

Member Function Documentation

◆ load_bits()

template<std::size_t N, ::native::isa<> Arch>
constexpr simd native::simd< bf16, N, Arch >::load_bits ( std::uint16_t const * p)
inlinestaticconstexprnoexcept

Read exactly N accessible uint16_t objects into corresponding BF16 lane bits. No alignment beyond that of uint16_t is required; p must not be null.

Definition at line 105 of file native.simd.ccm.

Here is the caller graph for this function:

◆ load_memory()

template<std::size_t N, ::native::isa<> Arch>
template<std::size_t Alignment = 1>
constexpr simd native::simd< bf16, N, Arch >::load_memory ( bf16 const * p)
inlinestaticconstexprnoexcept

Read exactly N accessible BF16 objects, preserving every encoding. Alignment is a nonzero power-of-two byte-alignment promise, not a runtime check. The default imposes no alignment beyond that required for BF16 objects. p must not be null.

Definition at line 116 of file native.simd.ccm.

Here is the caller graph for this function:

◆ load_partial()

template<std::size_t N, ::native::isa<> Arch>
constexpr simd native::simd< bf16, N, Arch >::load_partial ( bf16 const * p,
std::size_t n,
bf16 fill = bf16::from_bits(0) )
inlinestaticconstexprnoexcept

Read exactly the first n accessible BF16 objects, where n <= N. Copy their representations to lanes [0,n); remaining lanes receive fill's exact representation. The default fill is positive zero. No access occurs for n == 0, when p may be null; otherwise p must address n BF16 objects.

Definition at line 151 of file native.simd.ccm.

◆ store_bits()

template<std::size_t N, ::native::isa<> Arch>
void native::simd< bf16, N, Arch >::store_bits ( std::uint16_t * p) const
inlineconstexprnoexcept

Write every lane representation to N accessible uint16_t objects in lane order. No alignment beyond that of uint16_t is required; p must not be null.

Definition at line 110 of file native.simd.ccm.

Here is the caller graph for this function:

◆ store_memory()

template<std::size_t N, ::native::isa<> Arch>
template<std::size_t Alignment = 1>
void native::simd< bf16, N, Arch >::store_memory ( bf16 * p) const
inlineconstexprnoexcept

Write all lane representations to exactly N accessible BF16 objects. Alignment is a nonzero power-of-two byte-alignment promise, not a runtime check. The default imposes no alignment beyond that required for BF16 objects. p must not be null.

Definition at line 130 of file native.simd.ccm.

Here is the caller graph for this function:

◆ store_partial()

template<std::size_t N, ::native::isa<> Arch>
void native::simd< bf16, N, Arch >::store_partial ( bf16 * p,
std::size_t n ) const
inlineconstexprnoexcept

Write the representations of lanes [0,n) to exactly n accessible BF16 objects, where n <= N. Memory outside that prefix is untouched. No access occurs for n == 0, when p may be null; otherwise p must address n objects.

Definition at line 163 of file native.simd.ccm.


The documentation for this struct was generated from the following file: