Long-Context Speculative Decoding
Dates:
Stealth — details after publication

Applied curiosity.
I geek out on how things work, improve them, build my own, and keep them running. This habit has so far produced a research career and a private cloud in my apartment.
The current obsession is making language models cheaper to run, from the idea on paper to the chip it lands on. Smarter attention, self-chosen focus, faster decoding: three angles on one bill.
Dates:
Stealth — details after publication
Authors: Namgyu Ho * (equal contribution) , Huzama Ahmad * (equal contribution) , Woosung Koh * (equal contribution) , Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
A prompting protocol that lets a model declare where it will attend, cutting decoding attention cost up to 53.1% at near-zero accuracy loss.
Authors: Huzama Ahmad, Se-Young Yun
A plug-in selector that matches dense accuracy at long context while decoding 3.9× faster than FlashAttention.