Chang Lan

dblp:116/8770 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 3Security and privacy · 2Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Deep learning architectures and training · 62% Optimization for machine learning · 20% Efficient and distributed learning · 12%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Distributed systems · 27% Cloud and datacenter computing · 24% Hardware accelerators and domain-specific architectures · 14%
Computer networks
5 papers
Software-defined and programmable networks · 59% Routing and switching · 18% Network management and operations · 17%
Network and information security
4 papers
Network security · 69% Systems and software security · 31%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › distributed machine learning
distributed training
0.822020
A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters · OSDI 2020
A generic communication scheduler for distributed DNN training acceleration · SOSP 2019
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.712023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Machine learning › Deep learning architectures and training
mixture of experts
0.712023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Machine learning › Optimization for machine learning
sparse model
0.712023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Machine learning › Deep learning architectures and training
transformer
0.712023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Software-defined and programmable networks
network function virtualization
0.522018
ResQ: Enabling SLOs in Network Function Virtualization · NSDI 2018
E2: a framework for NFV applications · SOSP 2015
Software-defined and programmable networks
path control
0.522016
Explicit Path Control in Commodity Data Centers: Design and Applications · IEEE/ACM Trans. Netw. 2016
Explicit Path Control in Commodity Data Centers: Design and Applications · NSDI 2015
GPUs and heterogeneous computing › heterogeneous cluster computing
heterogeneous CPU-GPU cluster
0.412020
A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters · OSDI 2020
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412020
A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters · OSDI 2020
Machine learning › Efficient and distributed learning › distributed training
gradient aggregation
0.412019
A generic communication scheduler for distributed DNN training acceleration · SOSP 2019
Parallel and multicore computing › parallel scheduling
communication scheduling
0.412019
A generic communication scheduler for distributed DNN training acceleration · SOSP 2019
Routing and switching
routing
0.322016
Explicit Path Control in Commodity Data Centers: Design and Applications · IEEE/ACM Trans. Netw. 2016
Explicit Path Control in Commodity Data Centers: Design and Applications · NSDI 2015
Cloud and datacenter computing
secure outsourcing
0.212016
Embark: Securely Outsourcing Middleboxes to the Cloud · NSDI 2016
Network security › intrusion detection and prevention › intrusion detection
deep packet inspection
0.212015
BlindBox: Deep Packet Inspection over Encrypted Traffic · SIGCOMM 2015
Network security › intrusion detection and prevention › intrusion detection › deep packet inspection
encrypted traffic inspection
0.212015
BlindBox: Deep Packet Inspection over Encrypted Traffic · SIGCOMM 2015
Cloud and datacenter computing
cluster resource management and scheduling
0.212015
E2: a framework for NFV applications · SOSP 2015
Natural language and speech › Language models and text generation
large language model
0.212023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Network management and operations › fault management › fault diagnosis
data-plane fault localization
0.112012
Secure and Scalable Fault Localization under Dynamic Traffic Patterns · IEEE Symposium on Security and Privacy 2012
Network management and operations › fault management
fault diagnosis
0.112012
Secure and Scalable Fault Localization under Dynamic Traffic Patterns · IEEE Symposium on Security and Privacy 2012
High-performance computing › cluster computing
heterogeneous clusters
0.112020
A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters · OSDI 2020
Interconnection networks and networks-on-chip
remote direct memory access
0.112019
A generic communication scheduler for distributed DNN training acceleration · SOSP 2019
Cloud and datacenter computing
cloud security
0.112018
SafeBricks: Shielding Network Functions in the Cloud · NSDI 2018
Network security
middlebox security
0.112016
Embark: Securely Outsourcing Middleboxes to the Cloud · NSDI 2016
Cloud and datacenter computing
datacenter network
0.112016
Explicit Path Control in Commodity Data Centers: Design and Applications · IEEE/ACM Trans. Netw. 2016
Internet architecture and protocols
packet processing
0.112015
E2: a framework for NFV applications · SOSP 2015
Network security › intrusion detection and prevention
intrusion detection
0.112015
BlindBox: Deep Packet Inspection over Encrypted Traffic · SIGCOMM 2015
Datacenter networks
load balancing
0.012012
Secure and Scalable Fault Localization under Dynamic Traffic Patterns · IEEE Symposium on Security and Privacy 2012

Methods — techniques the papers use, named apart from their topics

bayesian optimization · 0.8shielding · 0.7sandboxing · 0.7neural architecture search · 0.7layer normalization · 0.7path compression · 0.5cryptography · 0.5TCAM · 0.5delayed key disclosure · 0.3software-defined networking · 0.2new encryption schemes · 0.2blindbox protocol · 0.2
YearPublicationVenuePosition
2023 Brainformers: Trading Simplicity for Efficiency
abstract
Transformers are central to recent successes in natural language processing and computer vision. Transformers have a mostly uniform backbone where layers alternate between feed-forward and self-attention in order to build a deep network. Here we investigate this design choice and find that more complex blocks that have different permutations of layer primitives can be more efficient. Using this insight, we develop a complex block, named Brainformer, that consists of a diverse sets of layers such as sparsely gated feed-forward layers, dense feed-forward layers, attention layers, and various forms of layer normalization and activation functions. Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and efficiency. A Brainformer model with 8 billion activated parameters per token demonstrates 2x faster training convergence and 5x faster step time compared to its GLaM counterpart. In downstream task evaluation, Brainformer also demonstrates a 3% higher SuperGLUE score with fine-tuning compared to GLaM with a similar number of activated parameters. Finally, Brainformer largely outperforms a Primer dense model derived with NAS with similar computation per token on fewshot evaluations.
Yanqi Zhou, Nan Du 0002, Yanping Huang, Daiyi Peng, Chang Lan, Siamak Shakeri, David R. So, Andrew M. Dai, Yifeng Lu, Quoc V. Le, Claire Cui, James Laudon, Jeffrey Dean
ICML5
2020 A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters
Yibo Zhu 0001, Chang Lan, Bairen Yi, Yong Cui 0001, Chuanxiong Guo
OSDI3
2019 A generic communication scheduler for distributed DNN training acceleration
abstract
We present ByteScheduler, a generic communication scheduler for distributed DNN training acceleration. ByteScheduler is based on our principled analysis that partitioning and rearranging the tensor transmissions can result in optimal results in theory and good performance in real-world even with scheduling overhead. To make ByteScheduler work generally for various DNN training frameworks, we introduce a unified abstraction and a Dependency Proxy mechanism to enable communication scheduling without breaking the original dependencies in framework engines. We further introduce a Bayesian Optimization approach to auto-tune tensor partition size and other parameters for different training models under various networking conditions. ByteScheduler now supports TensorFlow, PyTorch, and MXNet without modifying their source code, and works well with both Parameter Server (PS) and all-reduce architectures for gradient synchronization, using either TCP or RDMA. Our experiments show that ByteScheduler accelerates training with all experimented system configurations and DNN models, by up to 196% (or 2.96X of original speed).
Yanghua Peng, Yibo Zhu 0001, Yangrui Chen, Yixin Bao, Bairen Yi, Chang Lan, Chuan Wu 0001, Chuanxiong Guo
SOSP6
2018 SafeBricks: Shielding Network Functions in the Cloud
Rishabh Poddar, Chang Lan, Raluca A. Popa, Sylvia Ratnasamy
NSDI2
2018 ResQ: Enabling SLOs in Network Function Virtualization
Amin Tootoonchian, Aurojit Panda, Chang Lan, Melvin Walls, Katerina J. Argyraki, Sylvia Ratnasamy, Scott Shenker
NSDI3
2016 Embark: Securely Outsourcing Middleboxes to the Cloud
Chang Lan, Justine Sherry, Raluca A. Popa, Sylvia Ratnasamy
NSDI1
2016 Explicit Path Control in Commodity Data Centers: Design and Applications
abstract
Many data center network DCN applications require explicit routing path control over the underlying topologies. In this paper, we present XPath, a simple, practical and readily-deployable way to implement explicit path control, using existing commodity switches. At its core, XPath explicitly identifies an end-to-end path with a path ID and leverages a two-step compression algorithm to pre-install all the desired paths into IP TCAM tables of commodity switches. Our evaluation and implementation show that XPath scales to large DCNs and is readily-deployable. Furthermore, on our testbed, we integrate XPath into four applications to showcase its utility.
Shuihai Hu, Kai Chen 0005, Wei Bai 0001, Chang Lan, Hao Wang 0022, Chuanxiong Guo
IEEE/ACM Trans. Netw.5
2015 Explicit Path Control in Commodity Data Centers: Design and Applications
Shuihai Hu, Kai Chen 0005, Wei Bai 0001, Chang Lan, Hao Wang 0022, Chuanxiong Guo
NSDI5
2015 BlindBox: Deep Packet Inspection over Encrypted Traffic
abstract
Many network middleboxes perform deep packet inspection (DPI), a set of useful tasks which examine packet payloads. These tasks include intrusion detection (IDS), exfiltration detection, and parental filtering. However, a long-standing issue is that once packets are sent over HTTPS, middleboxes can no longer accomplish their tasks because the payloads are encrypted. Hence, one is faced with the choice of only one of two desirable properties: the functionality of middleboxes and the privacy of encryption. We propose BlindBox, the first system that simultaneously provides {\em both} of these properties. The approach of BlindBox is to perform the deep-packet inspection {\em directly on the encrypted traffic. BlindBox realizes this approach through a new protocol and new encryption schemes.
Justine Sherry, Chang Lan, Raluca A. Popa, Sylvia Ratnasamy
SIGCOMM2
2015 E2: a framework for NFV applications
abstract
By moving network appliance functionality from proprietary hardware to software, Network Function Virtualization promises to bring the advantages of cloud computing to network packet processing. However, the evolution of cloud computing (particularly for data analytics) has greatly benefited from application-independent methods for scaling and placement that achieve high efficiency while relieving programmers of these burdens. NFV has no such general management solutions. In this paper, we present a scalable and application-agnostic scheduling framework for packet processing, and compare its performance to current approaches.
Shoumik Palkar, Chang Lan, Sangjin Han, Keon Jang, Aurojit Panda, Sylvia Ratnasamy, Luigi Rizzo, Scott Shenker
SOSP2
2015 Blocking-resistant communication through domain fronting
abstract
Abstract We describe “domain fronting,” a versatile censorship circumvention technique that hides the remote endpoint of a communication. Domain fronting works at the application layer, using HTTPS, to communicate with a forbidden host while appearing to communicate with some other host, permitted by the censor. The key idea is the use of different domain names at different layers of communication. One domain appears on the “outside” of an HTTPS request—in the DNS request and TLS Server Name Indication—while another domain appears on the “inside”—in the HTTP Host header, invisible to the censor under HTTPS encryption. A censor, unable to distinguish fronted and nonfronted traffic to a domain, must choose between allowing circumvention traffic and blocking the domain entirely, which results in expensive collateral damage. Domain fronting is easy to deploy and use and does not require special cooperation by network intermediaries. We identify a number of hard-to-block web services, such as content delivery networks, that support domain-fronted connections and are useful for censorship circumvention. Domain fronting, in various forms, is now a circumvention workhorse. We describe several months of deployment experience in the Tor, Lantern, and Psiphon circumvention systems, whose domain-fronting transports now connect thousands of users daily and transfer many terabytes per month.
David Fifield, Chang Lan, Rod Hynes, Percy Wegmann, Vern Paxson
Proc. Priv. Enhancing Technol.2
2012 Secure and Scalable Fault Localization under Dynamic Traffic Patterns
abstract
Compromised and misconfigured routers are a well-known problem in ISP and enterprise networks. Data-plane fault localization (FL) aims to identify faulty links of compromised and misconfigured routers during packet forwarding, and is recognized as an effective means of achieving high network availability. Existing secure FL protocols are path-based, which assume that the source node knows the entire outgoing path that delivers the source node's packets and that the path is static and long-lived. However, these assumptions are incompatible with the dynamic traffic patterns and agile load balancing commonly seen in modern networks. To cope with real-world routing dynamics, we propose the first secure neighborhood-based FL protocol, DynaFL, with no requirements on path durability or the source node knowing the outgoing paths. Through a core technique we named delayed key disclosure, DynaFL incurs little communication overhead and a small, constant router state independent of the network size or the number of flows traversing a router. In addition, each DynaFL router maintains only a single secret key, which based on our measurement results represents 2 - 4 orders of magnitude reduction over previous path-based FL protocols.
Xin Zhang 0003, Chang Lan, Adrian Perrig
IEEE Symposium on Security and Privacy2