>16-bit - ~200% worse and wrong results This can be fixed in reduce_fast, by changing: - int n = ((int)r + 0x800000) >> 24; + int n = ((int32_t)r + 0x800000) >> 24; It's then ~400% slower, but gets the right answer. Cheers, Jon