MA Mark Zuckerberg
@zuck The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy. With adaptive delay, the model is near the pareto front on speed-accuracy tradeoff.