Computing infrastructure—algorithms, abstractions, interfaces, silicon—embodies the assumptions of its time. These assumptions, over time, harden into conventions and percolate across the stack. When technology and workloads change, yesterday’s conventions become today’s binding constraints. Revisiting them can open architectures that were previously impossible.
A server as the unit of computation. A host as the endpoint of a distributed system. Best-effort delivery as the fundamental primitive. Each shaped the entire stack. Each looked permanent. None of them was.
My work identifies these assumptions and rearchitects the stack they shaped—from algorithms to silicon. The result is systems that reach previously unattainable operating points in predictability, safety, efficiency, and explicability.
I love teaching! My views on education are best captured in: “Creativity is as important in education as literacy!”
Machines That Move
New research program · 2026–
Computing infrastructure assumes the machine stays still. The stack was built around that assumption. Addresses are stable, topology is fixed, bandwidth is provisioned, power comes from a wall, disconnection is a failure, and deadlines are numbers someone chose. None of that survives contact with a machine that moves: a robot in a kitchen, a tractor in a field, a rack in a datacenter. Its link can change by orders of magnitude from one second to the next. Its compute competes with its actuators for the same joules. Its deadlines are set by physics; miss one and the cup drops. Machines that move need a fundamentally new stack. This is the next instance of the research programs below: here, the physical substrate—machines in this case—has changed fundamentally, and thus the stack built around the old substrate must also change fundamentally.
Your Host Is a Distributed System
Ongoing research program · 2021–
A host as the endpoint of a distributed system was the right abstraction when everything inside it looked uniform—the assumption was that hosts comprise of a single kind of compute and memory devices with a unified coherent interconnect that looked ideal from the outside: nanosecond latencies, ample bandwidth, no loss. This assumption is no longer true: modern hosts have multiple heterogeneous compute devices (CPUs, GPUs, accelerators, etc.) and multiple heterogeneous memory devices (DRAM, HBM, etc.) with multiple interconnects each forming a different coherence domain. Continuing to use the abstraction of hosts as the endpoint of a distributed system has now become the binding constraint in today’s stacks: nanosecond-scale inefficiencies within the host can percolate through hardware and software stack to create microsecond- or even millisecond-scale inefficiencies at the application layer. In this research program, we are rearchitecting the stack with the lens of host as a distributed system: multiple devices joined by processor, memory, and peripheral interconnects, with queueing, routing, scheduling, flow control—and congestion—within the host. Preliminary results suggest that this perspective leads to previously unachievable operating points in predictability, safety, and efficiency.
Recognition: ACM SIGCOMM doctoral dissertation award · Cornell Bowers CIS dissertation award · IRTF applied networking research prize · SIGCOMM ’24 best student paper award · Deployed at scale
Resource Disaggregation
Ongoing research program · 2015–
A server as the unit of computation was the right assumption when interconnects within the host were orders of magnitude faster than the network: resources that worked together had to be bought together. Datacenter networks have since closed that gap. Workloads whose peaks dwarf their averages now pay for the old assumption continuously—in stranded memory, idle cores, and capacity bought for a peak that rarely arrives. Disaggregation is not a change to one layer. The fabric has to carry what a backplane used to. Storage and memory stacks have to tolerate a network in the middle. Frameworks have to work with resources that are no longer local. And algorithms have to decide who gets what. Rebuilt across all of them, a datacenter can be provisioned for its average rather than its peak.
Recognition: Google faculty research award · Deployed at scale
Best-effort Delivery as a Choice, Not the Only Choice
Ongoing research program · 2015–
Network design has always been an empirical craft: you build it, you measure it, you tune it, and you hope to get good performance. At internet scale, that was the only honest approach: nobody owns the topology, and best-effort is the strongest promise available. A datacenter is not the internet: a single entity owns the topology, knows the peers, and sees the load. That makes a different question possible—not what performs well, but what can be proven. Asking that question changes the answers: schedules that are optimal rather than merely good; architectures in which congestion is impossible rather than merely rare; performance you can state in advance rather than observe afterward. Best-effort becomes one of the choices, and not the only choice.
Recognition: SIGCOMM ’18 best student paper award · Google research scholar award · NSF CAREER award · Deployed at scale
| Omar Eqbal | IIT KGP Gold Medalist | PhD student |
| Midhul Vuppalapati | Assistant Professor, Purdue | PhD, 2026 |
| Saksham Agarwal | Assistant Professor, UIUC | PhD, 2024 |
| Qizhe Cai | Assistant Professor, UVA | PhD, 2024 |
| Sujaya Maiyya | Assistant Professor, University of Waterloo | Postdoc, 2022 |
| Mina Tahmasbi Arashloo | Assistant Professor, University of Waterloo | Postdoc, 2020–2022 |
| Jaehyun Hwang | Assistant Professor, Sungkyunkwan University | Postdoc, 2019–2021 |
| Anurag Khandelwal | Assistant Professor, Yale University | Postdoc, 2019 |
| Shreyas Kharbanda | Founding Engineer, Deft Robotics | MS, 2025 |
| Katie Gioioso | PhD student, Stanford | MS, 2021 |
Supported by the National Science Foundation and Alfred P. Sloan Foundation, with research gifts from Google, Samsung, and Snowflake.