Welcome to the Networked Systems Lab!

About

Founded in 2002, our laboratory conducts research on the design and implementation of a wide range of networked computing systems.

Selected Recent Publications

  1. SOSP
    DDB: Source-Level Interactive Debugging for Distributed Applications
    Yibo Yan, Junzhou He, and Seo Jin Park
    In Proceedings of the 32nd ACM Symposium on Operating Systems Principles Oct 2026
  2. SOSP
    Democratizing MoE LLM Decoding via Barrier-Free Expert Parallelism
    Yizhuo Liang, Shaoyu Wang, Jaeyong Song, Yanqi Zhou, Geon-Woo Kim, Guangrong He, and Seo Jin Park
    In Proceedings of the 32nd ACM Symposium on Operating Systems Principles Oct 2026
  3. SIGCOMM
    Near-optimal Online Traffic Engineering
    Arvin Ghavidel, Pooria Namyar, Nikolai Matni, Walter Willinger, and Ramesh Govindan
    In Proceedings of the ACM SIGCOMM 2026 Conference Oct 2026

    Most deployed WAN Traffic Engineering (TE) systems use a logically centralized controller that periodically gathers traffic demands, runs a TE optimization or heuristic, and then programs the network. At scale, these solutions are often suboptimal and can take minutes to react to demand changes or failures. In this paper, we introduce OnlineTE, a system that reacts immediately to demand changes and failures and delivers near-optimal solutions within seconds of a change. OnlineTE builds on the theory of optimization decomposition to devise scalable, near-optimal, distributed TE solvers for path-based MLU and Max-Flow problems. In OnlineTE, switches each solve a local subproblem, and a central coordinator coordinates their convergence. As such, a switch can trigger a re-optimization as soon as it detects a demand change or failure, enabling high reactivity. OnlineTE scales to large WANs, and its computational requirements are well within the capabilities of modern WAN switches. It also enables a novel paradigm, edge-based TE, which can utilize resources more efficiently than today’s path-based approaches. On a testbed emulation of a 750-node WAN topology, OnlineTE outperforms the state-of-the-art by up to an order of magnitude.

  4. SIGCOMM
    ZENITH: Towards A Formally Verified Highly-Available Control Plane
    Pooria Namyar, Arvin Ghavidel, Mingyang Zhang, Harsha V. Madhyastha, Srivatsan Ravi, Chao Wang, and Ramesh Govindan
    In Proceedings of the ACM SIGCOMM 2025 Conference Oct 2025

    Today, large-scale software-defined networks use microservice-based controllers. Bugs in these controllers can reduce network availability by making the data plane state inconsistent with the high-level intent. To recover from such inconsistencies, modern controllers periodically reconcile the state of all the switches with the desired intent. However, periodic reconciliation limits the availability and performance of the network at scale. We introduce Zenith, a microservice-based controller that avoids inconsistencies by design rather than always relying on recovery mechanisms. We have formally verified Zenith’s specifications and have proved that it ensures the network state will eventually be consistent with intent. We automatically generate Zenith’s code from its specification to minimize the likelihood of errors in the final implementation. Zenith’s guarantees and abstractions also enable developers to independently verify SDN applications and ensure end-to-end safety and correctness. Zenith resolves inconsistencies 5\texttimes faster than today’s designs and significantly improves availability.

  5. ATC
    Cosmic: Cost-Effective Support for Cloud-Assisted 3D Printing
    Yuan Yao, Chuan He, Chinedum Okwudire, and Harsha V Madhyastha
    In 2025 USENIX Annual Technical Conference Oct 2025

    In this paper, we consider a new workload for which serverless platforms are well-suited: the execution of a 3D printer controller in the cloud. This workload is qualitatively different from those considered in prior work due to the stringent timing requirements. Our measurements on popular serverless platforms reveal millisecond-level overheads that impair the timely execution of the example control algorithm we consider. To mitigate the impact of these overheads, we judiciously partition the execution of the algorithm across a set of serverless functions and exploit timely speculation. Our evaluations on AWS Lambda show that, for 30 diverse print jobs, Cosmic is able to ensure the timely execution of the controller while reducing cost by 2.8x–3.5x compared to other approaches.

News

July 6, 2026
Two papers accepted at SOSP 2026.

May 20, 2026
OnlineTE accepted at SIGCOMM 2026.

Sep 12, 2025
LiVo accepted at CoNEXT 2025.

Sep 1, 2025
Pooria Namyar joins Microsoft Research. Congrats!

Sep 1, 2025
Xiao Fu joins Meta. Congrats!

July 11, 2025
Two papers accepted at SIGCOMM 2025.

July 5, 2025
SplatPose accepted at ACM MM 2025.

Nov 18, 2024
Pooria Namyar awarded the 2024 Google Fellowship in networking

Sep 1, 2024
Weiwu Pang joins Google Cloud NetInfra. Congrats!

... see all News