Long-Context Speculative Decoding
Dates:
Stealth — details after publication

Applied curiosity.
I geek out on how things work, improve them, build my own, and keep them running. This habit has so far produced a research career and a private cloud in my apartment.
The current obsession is making language models cheaper to run, from the idea on paper to the chip it lands on. Smarter attention, self-chosen focus, faster decoding: three angles on one bill.
Dates:
Stealth — details after publication
Authors: Namgyu Ho * (equal contribution) , Huzama Ahmad * (equal contribution) , Woosung Koh * (equal contribution) , Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
Can a model become its own selector, saying where it wants to attend, and what does honoring it save? Up to half the decode attention cost, at near-dense accuracy.
Authors: Huzama Ahmad, Se-Young Yun
Where does a pretrained model look, and how do you select those tokens quickly? A small learned selector: decoding nearly 4× faster than FlashAttention at long context, with near-zero accuracy loss.