Language Models Can Control Their Own Attention
Authors: Namgyu Ho * (equal contribution) , Huzama Ahmad * (equal contribution) , Woosung Koh * (equal contribution) , Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
Can a model become its own selector, saying where it needs to attend, and what does honoring it save? Up to half the tokens read at each decode step, for a point or two of accuracy.
