It would be messier but faster if, when USE_TCACHE isn't defined, we avoid doing the non-inlined tail call. In the new code, chunks gotten from tcache are not tagged properly (via tag_new_usable()).