|
native 0.0.1
Vectors, masks and wide register packs for C++26
|
SHA compression repeatedly applies the same round and message-schedule functions. The SHA extension groups that work into instructions for one hash state, reducing the rotations, Boolean operations and additions needed by a compression implementation.
import native.x86.sha; supplies seven primitives through native::native. native.x86 and native also export them. All arguments and results are simd<std::uint32_t,4,Arch>.
| Operation | Meaning |
|---|---|
| sha1rnds4<Arch, Selector>(state, message) | Four SHA-1 rounds; selector 0–3 chooses the function and constant |
| sha1nexte<Arch>(state, message) | Add the derived E value to the high message word |
| sha1msg1<Arch>(a, b) | First step for four SHA-1 schedule words |
| sha1msg2<Arch>(a, b) | Final step for four SHA-1 schedule words |
| sha256rnds2<Arch>(cdgh, abef, message) | Two SHA-256 rounds, returning updated ABEF |
| sha256msg1<Arch>(a, b) | First step for four SHA-256 schedule words |
| sha256msg2<Arch>(a, b) | Final step for four SHA-256 schedule words |
SHA-1 state is [D,C,B,A] in increasing lane order. sha1rnds4 takes [W3,W2,W1,W0+E]; selectors 0–3 correspond to rounds 0–19, 20–39, 40–59 and 60–79. sha1nexte keeps message lanes 0–2 and adds rotr(state[3],2) to lane 3. SHA-1 schedules place the earliest word in the high lane.
SHA-256 uses cdgh=[H,G,D,C] and abef=[F,E,B,A]. The message holds [W0+K0,W1+K1,unused,unused]. The returned ABEF is [F,E,B,A]; the old ABEF becomes CDGH after two rounds. SHA-256 schedules place the earliest word low.
The four lanes are parts of one hash, not four independent hashes. Message primitives are partial schedule steps: SHA-1 also needs XOR with intervening words, and SHA-256 needs an addition between its two primitives. The caller supplies padding, byte-order conversion, feed-forward and the complete hash. SHA-1's instruction support does not make it suitable for collision-resistant applications.
Runtime calls require SHA, SSE2 storage and a "sha" caller target. Use target_features<native::x86>("sha") and admit it before entry. AVX and OS AVX state are unnecessary. The general AVX2 and AVX-512 profiles do not imply SHA; there are no writemasks.
Feature-bearing overloads are constexpr with native runtime paths. Without SHA, complete 128-bit storage permits consteval calls only. Arithmetic wraps modulo 2³²; Selector must be a compile-time unsigned value in 0–3. Inputs must have the same tag and shape.
See Intel's SHA extensions description, FIPS 180-4 and Clang's SHA header.