Vectorized CLZ and CTZ

October 9, 2026

clz and ctz are instructions that compute the number of leading (trailing) zero bits in a fixed-size integer. They are natively supported by modern CPUs, though they are not always fast, e.g. tzcnt has a latency of 3 on Arrow Lake.

I used ctz in an FPU emulator I’m working on, but figured out how to avoid it with floating-point trickery, and I just realized that this generalizes to a vectorizable ctz polyfill in a round-about way. I also implemented clz for completeness.

clz

Let’s start with clz, the easier of the two:

fn clz(x: u32) -> u32 {
    let a = 2.0f64.powi(-970);
    32 - ((f64::from_bits(a.to_bits() | x as u64) - a).to_bits() >> 52) as u32
}