Under development — I'm actively building this site.

Projects

What I build, from kernels to clusters.

  • Project Status: Active

    Long-Context Speculative Decoding

    Dates:

    Stealth — details after publication

  • Venue: Preprint Year: 2026

    Large Language Models Can Control Their Own Attention Span

    Authors: Namgyu Ho * (equal contribution) , Huzama Ahmad * (equal contribution) , Woosung Koh * (equal contribution) , Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos

    Can a model become its own selector, saying where it wants to attend, and what does honoring it save? Up to half the decode attention cost, at near-dense accuracy.

  • Project Status: Completed

    Foundations of Efficient LLMs

    Dates:

    How much of a model's work is actually necessary? Years of finding out: pretraining from scratch on TPU pods, folding long contexts into a handful of vectors, measuring uncertainty the Bayesian way, and tapering width with depth.

  • System Status: Active

    GPU Cluster

    Dates:

    The lab's research needs somewhere to run. Many nodes and mismatched GPUs became one system: one identity everywhere, storage that follows each user, strict per-job GPU isolation, and a queue that stays fair in deadline week.

  • System Status: Active

    HomeLab

    Dates:

    A personal cloud: everything I would otherwise rent, running in my apartment. Proxmox and ZFS underneath, pfSense at the edge, and production-grade tooling with a playground's freedom.

  • Project Status: Completed

    Undergraduate Research Mentorship

    Dates:

    How well do models hold up far from their training data? Eight undergraduate teams took the language side, four students the cultural one, and the work became four published benchmarks: BEnQA, CLIcK, CULTDIFF, and MIXCUBE.

  • Project Status: Completed

    StePer: Self-Guided Reinforcement Learning for Arithmetic Reasoning

    Dates:

    Can a model sharpen its own reasoning with no one grading it? It judges its greedy answer against its own sampled alternatives, and that judgment is the reward: accuracy up to 5% higher across four benchmarks, no human in the loop.

BibTeX