|
native 0.0.1
Vectors, masks and wide register packs for C++26
|
Public Types | |
| using | value_type = bf16 |
| Scalar storage element; each lane retains all 16 representation bits. | |
| using | register_type = simd |
| This one-register vector type, for generic register-based algorithms. | |
| using | native_type = std::conditional_t<N == 8,__m128bh,std::conditional_t<N == 16,__m256bh,__m512bh>> |
| Native register-width BF16 register representation; native bridges copy bits. | |
| using | bits_type = simd<std::uint16_t,N,architecture> |
| Unsigned 16-bit lanes in the same profile and lane order. | |
| using | mask = predicate<N,architecture> |
| Compact predicate with bit i selecting BF16 lane i. | |
| using | mask_type = mask |
| Generic mask spelling for the compact lane predicate. | |
| using | predicate_type = mask |
| Explicit predicate spelling for the same compact mask type. | |
| using | vector_mask_type = simd<mask16,N,architecture> |
| Full-register mask shape with zero or all-one 16-bit lanes. | |
| template<class T> | |
| using | rebind = simd<T,N,architecture> |
Public Member Functions | |
| simd () noexcept=default | |
| Default initialization leaves storage unspecified; braces zero it. | |
| constexpr | simd (bf16 value) noexcept |
| Broadcast the exact representation of value to all N lanes. | |
| constexpr | simd (std::array< bf16, lanes > const &values) noexcept |
| Copy array element i into lane i without conversion or representation changes. | |
|
template<class... T> requires (sizeof...(T) == lanes && (std::same_as<T,bf16> && ...)) | |
| constexpr | simd (T... values) noexcept |
| Construct all N lanes from BF16 values in argument order, preserving their bits. | |
| constexpr | simd (native_type value) noexcept |
| Adopt a native register without conversion or representation changes. | |
| constexpr | operator native_type () const noexcept |
| Project the native register for direct intrinsic interoperability. | |
| constexpr native_type | to_native () const noexcept |
| Return all lane bits as a native register, without conversion or lane reordering. | |
| constexpr bits_type | bits () const noexcept |
| Return the 16-bit representation of each lane in an unsigned vector. | |
| constexpr bits_type | to_bits () const noexcept |
| Synonym for bits(); this is a representation bridge, not a numeric conversion. | |
| constexpr void | store_bits (std::uint16_t *p) const noexcept |
| template<std::size_t Alignment = 1> | |
| constexpr void | store_memory (bf16 *p) const noexcept |
| constexpr void | store (bf16 *p) const noexcept |
| Store N BF16 objects with the default alignment contract of store_memory(). | |
| constexpr void | storeu (bf16 *p) const noexcept |
| Synonym for store(); no register-width alignment is required. | |
| constexpr void | store_partial (bf16 *p, std::size_t n) const noexcept |
Static Public Member Functions | |
| static constexpr simd | from_native (native_type value) noexcept |
| Copy a native BF16 register into this vector, preserving every representation bit. | |
| static constexpr simd | from_bits (bits_type value) noexcept |
| Interpret each unsigned lane as a BF16 representation without changing its bits. | |
| static constexpr simd | load_bits (std::uint16_t const *p) noexcept |
| template<std::size_t Alignment = 1> | |
| static constexpr simd | load_memory (bf16 const *p) noexcept |
| static constexpr simd | load (bf16 const *p) noexcept |
| Load N BF16 objects with the default alignment contract of load_memory(). | |
| static constexpr simd | loadu (bf16 const *p) noexcept |
| Synonym for load(); no register-width alignment is required. | |
| static constexpr simd | load_partial (bf16 const *p, std::size_t n, bf16 fill=bf16::from_bits(0)) noexcept |
Static Public Attributes | |
| static constexpr isa | architecture =Arch |
| The distinct compile-time AVX512_BF16 instruction profile. | |
| static constexpr std::size_t | lanes = N |
| Number of logical BF16 lanes, with no padding lanes. | |
Friends | |
| void | operator+ (simd)=delete |
| Reject unary BF16 arithmetic; storage does not define it. | |
| void | operator- (simd)=delete |
| Reject unary BF16 arithmetic; storage does not define it. | |
| void | operator~ (simd)=delete |
| Reject native fallback bitwise operations; use bits() explicitly. | |
| void | operator! (simd)=delete |
| Reject native fallback logical operations; storage is not a predicate. | |
One 128-, 256-, or 512-bit register of BF16 representations. Loads, stores and bit bridges preserve every encoding. This does not add elementwise BF16 arithmetic or change the scalar bf16 conversion contract. The 8-, 16-, and 32-lane shapes are provided. The application must admit that CPU/OS profile before entering compiled code. Every storage operation preserves subnormal, signed-zero and NaN encodings; none performs a floating-point conversion or quiets a signaling NaN.
Definition at line 44 of file native.simd.ccm.
| using native::simd< bf16, N, Arch >::rebind = simd<T,N,architecture> |
Replace the element type while retaining N lanes and this profile. Unsupported resulting shapes remain incomplete.
Definition at line 65 of file native.simd.ccm.
|
inlinestaticconstexprnoexcept |
Read exactly N accessible uint16_t objects into corresponding BF16 lane bits. No alignment beyond that of uint16_t is required; p must not be null.
Definition at line 105 of file native.simd.ccm.
|
inlinestaticconstexprnoexcept |
Read exactly N accessible BF16 objects, preserving every encoding. Alignment is a nonzero power-of-two byte-alignment promise, not a runtime check. The default imposes no alignment beyond that required for BF16 objects. p must not be null.
Definition at line 116 of file native.simd.ccm.
|
inlinestaticconstexprnoexcept |
Read exactly the first n accessible BF16 objects, where n <= N. Copy their representations to lanes [0,n); remaining lanes receive fill's exact representation. The default fill is positive zero. No access occurs for n == 0, when p may be null; otherwise p must address n BF16 objects.
Definition at line 151 of file native.simd.ccm.
|
inlineconstexprnoexcept |
Write every lane representation to N accessible uint16_t objects in lane order. No alignment beyond that of uint16_t is required; p must not be null.
Definition at line 110 of file native.simd.ccm.
|
inlineconstexprnoexcept |
Write all lane representations to exactly N accessible BF16 objects. Alignment is a nonzero power-of-two byte-alignment promise, not a runtime check. The default imposes no alignment beyond that required for BF16 objects. p must not be null.
Definition at line 130 of file native.simd.ccm.
|
inlineconstexprnoexcept |
Write the representations of lanes [0,n) to exactly n accessible BF16 objects, where n <= N. Memory outside that prefix is untouched. No access occurs for n == 0, when p may be null; otherwise p must address n objects.
Definition at line 163 of file native.simd.ccm.