VLDB 2026 Research / reviewers in the wild / expert
Congcong Miao
dblp:30/10775
· DBLP profile ↗
40ranked-venue papers
13as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 27 · 10 first-author · 24 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multipath Collective Communication Beyond Scale-up Networks in GPU Clouds
Yuchen Xu 0003, Jianglong Nie, Baojia Li 0002, Mingzhuo Chen, Guanyu Qu, Zhenchuan Liu, Shuangshuang Yin, Chunzhi He, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Congcong Miao, Wenfei Wu |
EuroSys | 17 |
| 2026 | ECOTE: Priority-Aware Optical Restoration for WAN Traffic EngineeringabstractFiber cuts are among the most common and disruptive failures in cloud networks. They can prevent cloud providers from maintaining committed service availability, causing Service Level Agreement (SLA) violations that directly translate into monetary penalties. Existing traffic engineering (TE) approaches enhance failure resilience, and recent systems further incorporate optical restoration to recover lost bandwidth after failures. However, they still treat services largely uniformly and optimize primarily for network throughput rather than the economic impact of heterogeneous SLA penalties. In this paper, we present ECOTE, the first priority-aware TE system with optical restoration that explicitly minimizes revenue loss. Specifically, ECOTE introduces a new optical restoration formulation with a dedicated capacity restoration solver to compute an optimal restoration plan that maximizes restorable capacity under physical constraints. ECOTE also designs a priority-aware TE algorithm that allocates the restored bandwidth capacity according to SLA penalties, thereby reducing monetary cost. We evaluate ECOTE using a production-level WAN testbed and through large-scale simulations. The testbed evaluation demonstrates ECOTE achieves zero loss for high-priority services and more than 10× revenue loss reduction compared to state-of-the-art. Our large-scale simulation results show that ECOTE can support at least 2.5× and 2.0× more demand for different high priority services compared to the state-of-the-art solutions. Meanwhile, ECOTE reduces the revenue loss by at least an order of magnitude less than existing solutions. Kunling He, Ran Shu 0001, Jilong Wang 0001, Congcong Miao |
EuroSys | 6 |
| 2026 | BLADE: Adaptive Wi-Fi Contention Control for Next-Generation Real-Time Communication
Fengqian Guo, Longwei Jiang, Congcong Miao, Chenren Xu, Hancheng Lu, Chang Wen Chen, Yaxiong Xie |
NSDI | 4 |
| 2026 | MirrorNet: High-fidelity and Scalable Network Emulation for Software-defined WAN
Congcong Miao, Yuejie Wang, Xuefeng Ji, Guozhi Shan, Pan Fang, Yanke Zhang, Xianneng Zou, Guyue Liu |
NSDI | 1 |
| 2026 | Cost-effective and Reliable Global Internet Peering with Programmable Switches
Congcong Miao, Zhiyi Yao, Jianchao Lv, Jinglin Wang, Shihan Lin, Xinyi Zhang 0004, Yunming Xiao, Jiwu Bu, Yachen Wang, Xianneng Zou, Yong Jiang 0001, Marco Canini, Gaogang Xie |
NSDI | 1 |
| 2026 | A Composable Emulation Framework for Whitebox Switches
Congcong Miao, Xianneng Zou, Chuwen Zhang, Qihang Liu, Zhijie Yan, Yanke Zhang, Yong Jiang 0001, Qiao Xiang, Xin Jin 0008, Zili Meng, Ang Chen 0001 |
NSDI | 1 |
| 2026 | From Source to Solution: Tackling Packet Losses in Large-scale Cloud Gaming Systematically and Precisely
Jing Wang 0077, Yunzhe Ni, Nian Wen, Congcong Miao |
NSDI | 6 |
| 2026 | DDoS Detection at the Scale of One Hundred Tbps
Yunming Xiao, Xijun Luo, Youliang Jiang, Aike Wang, Heng Yu 0005, Jiahao Cao 0001, Yong Jiang 0001, Jilong Wang 0001, Mingwei Xu 0001, Congcong Miao |
NSDI | 13 |
| 2026 | XFir: Accelerating New-Flow Setup on Host Servers of a Large Cloud NetworkabstractIn today's cloud networks, host servers widely deploy Data Processing Units (DPUs) as network accelerators under the "Sep-Path" paradigm. However, as server capabilities scale with increasing CPU cores and network bandwidth, the software slow path (executed on a DPU's CPU) has become a critical bottleneck for workloads with high new-flow rates. Meanwhile, new-flow setup logic on host servers must continuously evolve to meet diverse and changing customer demands, making flexibility a key requirement alongside performance. To address this gap, we present XFir, the first hardware-accelerated new-flow setup system for cloud host servers that delivers high CPS throughput while preserving sufficient flexibility. XFir leverages a next-generation DPU equipped with a Cloud Network co-Processor (CNP) to execute the host server's new-flow setup logic. XFir redesigns the host-server flow-setup datapath and table layout, optimizes LPM lookups, and introduces CPU-CNP collaboration mechanisms to further improve performance and reliability. Our evaluation shows that XFir achieves over 776K new-flow CPS on a single host server with 11.7μs slow-path latency. Compared to prior work (Fornax), XFir achieves 4.8x CPS and reduces latency by 69.2%. Moreover, XFir is cost-effective to deploy, requiring only a single DPU per host. Overall, XFir improves new-flow throughput while maintaining development flexibility at low financial cost. Shihan Lin, Shunqiao Jiang, Chao Pei, Jian Zhao 0006, Wenjun Wu 0001, Lijun Zhuang, Qingmin Liu, Heng Yu 0005, Yibo Huang 0005, Yifei Zhu 0001, Yunming Xiao, Ang Chen 0001, Linghe Kong, Congcong Miao |
SIGCOMM | 18 |
| 2026 | Turbo: Efficiently Serving Long-Context Large Language Models with In-Network AggregationabstractLLM supporting long contexts faces a critical memory bottleneck due to the linear growth of KV cache. Distributing the storage across multiple GPUs alleviates this burden but introduces significant communication overhead or traffic incast, especially during the decoding phase. We propose Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches. We address three key challenges to map complex attention mechanisms onto restricted switch hardware: (i) To bypass the switch's inability to buffer global states or perform complex operations, we devise online table-based aggregation, which decomposes global reduction into pairwise operations and approximates nonlinear functions via lookup tables. (ii) To circumvent the restriction on retroactive state access in RMT pipelines, we introduce a rolling forward scheme that propagates states to enable cross-stage updates. (iii) To mitigate aggregation stragglers caused by topology-induced load imbalance, we construct a load-aware aggregation tree that optimizes workload distribution. Evaluations on a Tofino2-based testbed show that Turbo reduces end-to-end inference latency by up to 37%. Large-scale simulations on NS-3 demonstrate that Turbo significantly outperforms state-of-the-art baselines in both inference latency and network traffic reduction with negligible accuracy loss. Ying Wan 0001, Yuchen Xu 0003, Chuwen Zhang, Yingsheng Huang, Wenquan Xu, Jialin Li 0001, Mingwei Xu 0001, Wenfei Wu, Congcong Miao |
SIGCOMM | 10 |
| 2026 | CubeTrace: Microscopic Network Tracing for Heterogeneous Cloud Gateways
Yunming Xiao, Yinchao Yang, Jiaqi Zheng 0001, Xuqian Li, Dongbo Gu, Jun Zhang 0014, Miantao Wan, Chao Pei, Chen Tian 0001, Mingwei Xu 0001, Ang Chen 0001, Congcong Miao |
SIGCOMM | 12 |
| 2026 | Dorado: Scaling SmartNIC Session Tables on Commodity DDRs
Heng Yu 0005, Jiajun Liang, Baozeng Zhang, Guozhi Lin, Xinyi Zhang 0004, Jian Zhao 0006, Ziyue Zhai, Chao Pei, Jilong Wang 0001, Gaogang Xie, Ang Chen 0001, Congcong Miao |
SIGCOMM | 15 |
| 2026 | An Efficient Computing and Communication Framework for Large-Scale Data Processing Cluster
Xuya Jia, Zhiyi Yao, Edison Liu, Congcong Miao, Yuedong Xu 0001 |
IEEE Trans. Netw. | 6 |
| 2025 | Reasoning under Uncertainty: Efficient LLM Inference via Unsupervised Confidence Dilution and Convergent Adaptive SamplingabstractZhenning Shi, Yijia Zhu, Yi Xie, Junhan Shi, Guorui Xie, Haotian Zhang, Yong Jiang, Congcong Miao, Qing Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhenning Shi, Yijia Zhu, Junhan Shi, Guorui Xie, Yong Jiang 0001, Congcong Miao, Qing Li 0006 |
EMNLP | 8 |
| 2025 | A Spatially-Adapted SHAP Approach for Interpreting Deep Bike Usage Learning and PredictionabstractUnderstanding the spatial dynamics of bike-sharing usage is critical for effective urban planning and mobility resource management. In this study, we propose an interpretable deep learning approach to uncover spatial relationships embedded in bike-sharing activities. Specifically, we develop a spatially-adapted SHapley Additive exPlanations (SHAP)-based method to quantify the spatial dependencies between locations in bike-sharing activities and apply it to interpret the predictions of a bike-sharing model. Extensive experiments upon Citi Bike data from New York City in December 2023 reveal that spatial influence does not strictly follow geographic proximity and is anisotropic. Additionally, non-member users exhibit weaker spatial dependencies in their bike usage behavior, resulting in lower short-term predictability compared to member users. Our studies shed deep insights into the spatial dynamics of bike-sharing systems and provide guidance for more effective service deployment and system design. Congcong Miao, Suining He, Yuyao Li, Chuanrong Zhang |
SIGSPATIAL/GIS | 1 |
| 2025 | Flexnetic: Cost-Effective and Smooth Evolution of Optical BackboneabstractThe increasing traffic on WANs due to growing number of applications imposes significant strain on the infrastructure of cloud service providers. The primary expense in augmenting network capacity entails high costs associated with procuring expensive transponders that facilitate inter-regional optical signal transmission. The advent of spacing variable transponders facilitates cost-effective network upgrades and increased transmission efficiency, however, implementing such upgrades in one step may incur high costs and service disruptions. Achieving a cost-effective and smooth network upgrade presents considerable challenges in identifying critical IP links and minimizing costs within existing architecture. We introduce Flexnetic, a planning tool which utilizes a hybrid approach of both modern and legacy transponders, along with establishment of optical bypass, to accommodate the escalating traffic demands while minimizing the costs during network upgrades. Flexnetic incorporates two novel algorithms: a NLP model to maximize IP link capacity utilization, and an MIP model for efficient IP layer implementation at optical layer, emphasizing reuse of existing transponder and minimizes new device requirement. Our simulation of upgrade plans on common WAN topologies revealed superior performance, with up to 91.9% cost savings and 2.33× capacity increase over existing state-of-the-art solutions, highlighting Flexnetic’s potential for cost-efficient and capacity-optimized network upgrades. Congcong Miao, Kunling He, Jilong Wang 0001 |
ICCCN | 2 |
| 2025 | Unlocking ECMP Programmability for Precise Traffic Control
Yunming Xiao, Weizhen Dang, Xiang Li 0223, Zekun He, Jilong Wang 0001, Aleksandar Kuzmanovic, Ang Chen 0001, Congcong Miao |
NSDI | 11 |
| 2025 | Holmes: Localizing Irregularities in LLM Training with Mega-scale GPU Clusters
Zhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia, Zuning Liang, Yuedong Xu 0001, Chunzhi He, Mingzhuo Chen, Xiang Li 0010, Zekun He, Yachen Wang, Xianneng Zou, Junchen Jiang |
NSDI | 3 |
| 2025 | Fornax: A Hardware-Centric Session Management in Large Public Cloud NetworkabstractSmartNIC is increasingly utilized to accelerate cloud network components. The effectiveness and correctness of hardware acceleration heavily rely on its management mechanism. Unfortunately, traditional management mechanisms adopt software-centric architecture, which treats flow as the basic management unit and completely relies on one-way commands to manage the flow table, making it challenging to support various cloud network scenarios while managing extremely large tables. In this paper, we advocate for a radical new mechanism to shift the management paradigm from software-centric architecture to hardware-centric architecture, which adopts session as the basic management unit and designs two-way protocols to facilitate the management process. We propose and implement a first-of-its-kind system, called Fornax, a novel management architecture for large public cloud networks. At the core of Fornax is leveraging a session-empowered hardware engine to provide various management capabilities. Besides, Fornax utilizes a light-weight software manager to enhance system scalability, and hardware-driven management protocols to improve resource efficiency. Our testbed evaluations demonstrate that Fornax can reduce the software storage usage by 80% and CPU usage by 77% with little hardware resource overhead. Our large-scale production results show that Fornax can manage up to 16M session entries while significantly reducing the resource overhead by over 79%. Heng Yu 0005, Jian Zhao 0006, Guozhi Lin, Baozeng Zhang, Yunpeng Guan, Jiajun Liang, Chao Pei, Yachen Wang, Xin Jin 0008, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 16 |
| 2025 | PreTE: Traffic Engineering with Predictive FailuresabstractFiber links in wide-area networks (WANs) are exposed to complicated environments and hence are vulnerable to failures like fiber cuts. The conventional approach of using static probabilistic failures falls short in fiber-cut scenarios because these fiber cuts are rare but disruptive, making it difficult for network operators to balance network utilization and availability in WAN traffic engineering. Our large-scale measurements of per-second optical-layer data reveal that the fiber's failure probability increases by several orders of magnitude when experiencing a rare and ephemeral degradation state. Therefore, we present a novel traffic engineering (TE) system called PreTE to factor in the dynamic fiber cut probabilities directly into TE systems. At the core of the PreTE system, fiber degradation facilitates failure predictions and traffic tunnels to be proactively updated, followed by traffic allocation optimizations among updated tunnels. We evaluate PreTE using a production-level WAN testbed and large-scale simulations. The testbed evaluation quantifies PreTE's runtime to demonstrate the feasibility to implement in large-scale WANs. Our large-scale simulation results show that PreTE can support up to 2× more demand at the same level of availability as compared to existing TE schemes. Congcong Miao, Zhizhen Zhong, Arpit Gupta, Ying Zhang 0022, Zekun He, Xianneng Zou, Jilong Wang 0001 |
SIGCOMM | 1 |
| 2025 | Ares: Comprehensive Path Hijacking Detection via Routing Tree
Yinxiang Tao, Chengwan Zhang, Changqing An, Shuying Zhuang, Jilong Wang 0001, Congcong Miao |
USENIX Security Symposium | 6 |
| 2025 | Predictive Configuration on DHCP in WLANsabstractDHCP is widely deployed in WLANs to automatically assign IP addresses to WiFi devices when users connect to the WLANs. However, frequent user mobility brings big challenges to the DHCP performance. Recently proposed IP configuration (e.g., IP lease time, size of IP address pool) decisions on DHCP are based on traditional models to study user mobility patterns which lead to poor DHCP performance since the online time of individuals varies due to their personal pReferences and the number of crowds differs spatially and temporally. In this paper, we propose PredHCP, a predictive configuration framework on DHCP to improve the DHCP performance. Specifically, PredHCP utilizes an attention-based recurrent neural network (ARNN) to learn sequential patterns of individual mobility and accurately predicts user online time to ensure the effective IP lease time configuration. Meanwhile, PredHCP introduces a spatio-temporal graph neural network (STGNN) to learn both spatial and temporal dependencies of crowd migration and accurately predict crowd size in each area to ensure effective IP pool configuration. We conduct comprehensive experiments on real network traces for a month to evaluate the performance of PredHCP. Experimental results show that PredHCP can accurately predict user mobility patterns by achieving lower prediction errors. By accurately modeling mobility patterns, PredHCP makes effective IP configuration to ensure high DHCP performance. Large-scale simulation results show that PredHCP can save up to 69% IP addresses and the IP efficiency is 41% which outperforms existing methods by 6%. Pei Zhang 0003, Hanyan Yin, Botong Wu, Xiaohong Huang 0003, Yan Ma 0003, Jilong Wang 0001, Congcong Miao |
IEEE Trans. Netw. | 9 |
| 2024 | Turbo: Efficient Communication Framework for Large-scale Data Processing ClusterabstractBig data processing clusters are suffering from a long job completion time due to the inefficient utilization of the RDMA capability. Our production measurement results in a large-scale cluster with hundreds of server nodes to process large-scale jobs have shown that the existing deployment of RDMA technique results in a long-tail job completion time, with some jobs even taking up more than twice the average time to complete. In this paper, we present the design and implementation of Turbo, an efficient communication framework for the large-scale data processing cluster to achieve high performance and scalability. The core of Turbo's approach is to leverage a dynamic block-level flowlet transmission mechanism and a non-blocking communication middleware to improve the network throughput and enhance system's scalability. Furthermore, Turbo ensures high system reliability by utilizing an external shuffle service as well as TCP serving as a backup. We integrate Turbo into Apache Spark and evaluate Turbo in a small-scale testbed and a large-scale cluster consisting of hundreds of server nodes. The small-scale testbed evaluation results show that Turbo improves the network throughput by 15.1% while maintaining high system reliability. The large-scale production results have shown Turbo can reduce the job completion time by 23.9% and increase the job completion rate by 2.03× over the existing RDMA solutions. Xuya Jia, Zhiyi Yao, Edison Liu, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Chongqing Zhao, Jinhui Chu, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 14 |
| 2024 | MegaTE: Extending WAN Traffic Engineering to Millions of Endpoints in Virtualized CloudabstractIn today's virtualized cloud, containers and virtual machines (VMs) are prevailing methods to deploy applications with different tenant requirements. However, these requirements are at odds with the resource allocation capabilities of conventional networking stacks in wide-area networks (WANs). In particular, existing WAN traffic engineering (TE) systems at the granularity of aggregated traffic flows are not designed to cater to each individual flow. In this paper, we advocate for a radical new approach to extend TE systems to involve millions of virtual instance endpoints. We propose and implement a first-of-its-kind system, called MegaTE, to satisfy the needs of each fine-grained traffic flow at the virtual instance level. At the core of the MegaTE system is the paradigm shift from the top-down centralized control to the bottom-up asynchronous query in the TE control loop, combined with eBPF-based segment routing on the data plane and TE optimization contraction on the control plane. We evaluate MegaTE using flow-level simulations with production traffic traces. Our results show that MegaTE supports 20× more endpoints with the similar algorithm run time compared to prior work. MegaTE has been adopted by large-scale public cloud providers. Notably, Tencent rolled out MegaTE in its cloud WAN since December 2022. Our production analysis shows that MegaTE reduces the packet latency of real-time applications by up to 51%. Congcong Miao, Zhizhen Zhong, Yunming Xiao, Senkuo Zhang, Yinan Jiang, Zizhuo Bai, Chaodong Lu, Jingyi Geng, Zekun He, Yachen Wang, Xianneng Zou, Chuanchuan Yang |
SIGCOMM | 1 |
| 2024 | Proactively Verifying Quantitative Network Policy Across Unsafe and Unreliable EnvironmentsabstractNetwork managers configure networks to enforce various high-level policies, and to respond to the wide range of network events (e.g., attacks, intrusions, malicious route announcements from neighbors) that may occur. It is incredibly difficult to specify these high-level policies in terms of distributed low-level configuration. These high-level policies hold only if the distributed configurations are well equipped to react to unsafe and unreliable environments (e.g., malicious route announcements, unsafe components and devices). Therefore, it is important to proactively verify whether network policies hold across continually changing environments in terms of current network configurations. State-of-the-art policy verification techniques are limited because they can check only the Boolean policies (e.g., forwarding reachability, waypoint or blackhole-freeness). However, many policy violations express themselves in quantitative ways (e.g., a link becomes overloaded). In this paper, we propose quantitative network verification (QNV) analyzing the quantitative policies of networks across unsafe and unreliable environments. QNV translates network configurations into a symbolic simulation model that captures the stable states to which the network forwarding will converge as a result of interactions between routing protocols. It then generates a logical formula matrix that describes network forwarding in the event of failures and verifies quantitative policies based on the formula matrix. We implement QNV and evaluate it on realistic and synthetic configurations. Our evaluation shows that QNV can precisely verify quantitative policies in only a few minutes, even in large networks. Han Zhang 0009, Jilong Wang 0001, Xingang Shi, Xia Yin 0001, Jiankun Hu, Congcong Miao |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2023 | Metis: Detecting Fake AS-PATHs Based on Link PredictionabstractBGP route hijacking is a critical threat to the Internet. Existing works on path hijacking detection firstly monitor the routes of the whole network and then directly trigger a suspicious alarm if the link has not been seen before. However, these naive approaches will cause false positive identification and introduce unnecessary verification overhead. In this work, we propose Metis, a matching-and-prediction system to filter out normal unseen links. We first use a matching method with three rules to find out suspicious links if there is an unseen AS. Otherwise, we propose using a neural network to make a prediction based on the AS information at each end of the link and further quantify the suspicion level. Our large-scale simulation results show that Metis can achieve precision and recall of over 80% for detecting fake AS-PATHs. Moreover, our deployment experiences show that compared to state-of-the-art system, Metis can save 80% overhead. Chengwan Zhang, Congcong Miao, Changqing An, Anlun Hong, Ning Wang 0001, Jilong Wang 0001 |
ISCC | 2 |
| 2023 | TENSOR: Lightweight BGP Non-Stop RoutingabstractAs the solitary inter-domain protocol, BGP plays an important role in today's Internet. Its failures threaten network stability and will usually result in large-scale packet losses. Thus, the non-stop routing (NSR) capability that protects inter-domain connectivity from being disrupted by various failures, is critical to any Autonomous System (AS) operator. Replicating the BGP and underlying TCP connection status is key to realizing NSR. But existing NSR solutions, which heavily rely on OS kernel modifications, have become impractical due to providers' adoption of virtualized network gateways for better scalability and manageability. Congcong Miao, Yunming Xiao, Marco Canini, Ruiqiang Dai, Shengli Zheng, Jilong Wang 0001, Jiwu Bu, Aleksandar Kuzmanovic, Yachen Wang |
SIGCOMM | 1 |
| 2023 | FlexWAN: Software Hardware Co-design for Cost-Effective and Resilient Optical BackbonesabstractThe rising demand for WAN capacity driven by the rapid growth of inter-data center traffic poses new challenges for costly optical networks. Today cloud providers rely on fixed optical backbones, where all hardware devices operate on a rigid spectrum grid, leading to the waste of expensive optical resources and subpar performance in handling failures. In this paper, we introduce FlexWAN, a novel flexible WAN infrastructure designed to provision cost-effective WAN capacity while ensuring resilience to optical failures. FlexWAN achieves this by incorporating spacing-variable hardware at the optical layer, enabling the generated wavelength to optimize the utilization of limited spectrum resources for the WAN capacity. The configuration of spacing-variable hardware in a multi-vendor optical backbone presents challenges related to spectrum management. To address this, FlexWAN leverages a centralized controller to achieve coordinated control of network-wide optical devices in a vendor-agnostic manner. Moreover, the flexibility at the optical layer introduces new algorithmic problems. FlexWAN formulates the problem of provisioning WAN capacity with the goal of minimizing hardware costs. We evaluate the system performance in production and share insights from years of production experience. Compared to existing optical backbones, FlexWAN can save at least 57% of transponders and reduce 36% of spectrum usage while continuing to meet up to 8× the present-day demands using existing hardware and fiber deployments. FlexWAN further incorporates failure resilience that revives 15% more bandwidth capacity in the overloaded optical backbone. Congcong Miao, Zhizhen Zhong, Ying Zhang 0022, Kunling He, Fangchao Li, Minggang Chen, Xiang Li 0223, Zekun He, Xianneng Zou, Jilong Wang 0001 |
SIGCOMM | 1 |
| 2023 | Serpens: A High Performance FaaS Platform for Network FunctionsabstractMore and more enterprises deploy applications on Function-as-a-Service (FaaS) platforms to improve resource efficiency and save monetary costs. Network Functions (NFs) suffer from staggered peaks of traffic patterns and could benefit from fine-grained resource multiplexing in FaaS platform. However, naively exploring existing FaaS platforms to support NFs can introduce significant performance overheads in three aspects, including slow instance startup, remote state access for NFs, and costly packet delivery between NFs. To address these problems, we propose${\sf Serpens}$, a high performance FaaS platform for NFs. First,${\sf Serpens}$proposes a reusable NF runtime design to slash instance startup overhead. Second,${\sf Serpens}$designs a novel state management mechanism to support local state access. Third,${\sf Serpens}$introduces an advanced service chaining approach to avoid extra packet delivery. Besides,${\sf Serpens}$designs an NF scaling mechanism to minimize performance fluctuation. We have implemented a prototype of${\sf Serpens}$and conducted comprehensive experiments. Compared with the NFs and Service Function Chains (SFCs) that run on existing FaaS platforms,${\sf Serpens}$can improve the throughput by more than 10× and reduce the latency by more than 90%. Heng Yu 0005, Han Zhang 0009, Junxian Shen, Yantao Geng, Jilong Wang 0001, Congcong Miao, Mingwei Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | Scorpius: Proactive Code Preparation to Accelerate Function StartupabstractMassive enterprises deploy their applications on public clouds to relieve infrastructure management burden. However, applications are faced with highly fluctuating workloads, while clouds provision exclusive resources at coarse time granularity, resulting in severely low resource efficiency. Function-as-a-Service (FaaS) platform enables fine-grained resource multiplexing, which has the potential to improve efficiency. However, FaaS platforms could consume several seconds to start functions and the long startup latency can severely hurt the performance of applications. In this paper, we measure the FaaS platforms and find that most startup latency is occupied by code preparation. To reduce the code preparation latency with little resource overhead, we propose Scorpius, a FaaS platform that proactively prepares code based on the historical data of functions. It combines two optimization categories: (1) To reduce the code size, Scorpius proposes to proactively prepare partial libraries over servers and run functions on the server with most library sharing. (2) To advance the start time, Scorpius proposes to predict the function overload with a simple model and proactively scale code to more servers. We have implemented a prototype of Scorpius and conducted extensive experiments. Evaluation results demonstrate that compared with state-of-the-art methods, Scorpius can reduce the code preparation latency by 87.6% with only 9.3% storage overhead. Heng Yu 0005, Junxian Shen, Han Zhang 0009, Jilong Wang 0001, Congcong Miao, Mingwei Xu 0001 |
IWQoS | 5 |
| 2022 | Detecting Ephemeral Optical Events with OpTel
Congcong Miao, Minggang Chen, Arpit Gupta, Zili Meng, Lianjin Ye, Jingyu Xiao, Zekun He, Xulong Luo, Jilong Wang 0001, Heng Yu 0005 |
NSDI | 1 |
| 2022 | RLMob: Deep Reinforcement Learning for Successive Mobility PredictionabstractHuman mobility prediction is an important task in the field of spatiotemporal sequential data mining and urban computing. Despite the extensive work on mining human mobility behavior, little attention was paid to the problem of successive mobility prediction. The state-of-the-art methods of human mobility prediction are mainly based on supervised learning. To achieve higher predictability and adapt well to the successive mobility prediction, there are four key challenges: 1) disability to the circumstance that the optimizing target is discrete-continuous hybrid and non-differentiable. In our work, we assume that the user's demands are always multi-targeted and can be modeled as a discrete-continuous hybrid function; 2) difficulty to alter the recommendation strategy flexibly according to the changes in user needs in real scenarios; 3) error propagation and exposure bias issues when predicting multiple points in successive mobility prediction; 4) cannot interactively explore user's potential interest that does not appear in the history. While previous methods met these difficulties, reinforcement learning (RL) is an intuitive answer for this task to settle these issues. We innovatively introduce RL to the successive prediction task. In this paper, we formulate this problem as a Markov Decision Process. We further propose a framework - RLMob to solve our problem. A simulated environment is carefully designed. An actor-critic framework with an instance of Proximal Policy Optimization (PPO) is applied to adapt to our scene with a large state space. Experiments show that on the task, the performance of our approach is consistently superior to that of the compared approaches. Ziyan Luo, Congcong Miao |
WSDM | 2 |
| 2022 | Self-supervised representation learning for trip recommendation
Qiang Gao 0003, Wei Wang 0336, Kunpeng Zhang 0001, Xin Yang 0012, Congcong Miao, Tianrui Li 0001 |
Knowl. Based Syst. | 5 |
| 2021 | Predicting Crowd Flows via Pyramid Dilated Deeper Spatial-temporal NetworkabstractPredicting crowd flows is crucial for urban planning, traffic management and public safety. However, predicting crowd flows is not trivial because of three challenges: 1) highly heterogeneous mobility data collected by various services; 2) complex spatio-temporal correlations of crowd flows, including multi-scale spatial correlations along with non-linear temporal correlations. 3) diversity in long-term temporal patterns. To tackle these challenges, we proposed an end-to-end architecture, called pyramid dilated spatial-temporal network (PDSTN), to effectively learn spatial-temporal representations of crowd flows with a novel attention mechanism. Specifically, PDSTN employs the ConvLSTM structure to identify complex features that capture spatial-temporal correlations simultaneously, and then stacks multiple ConvLSTM units for deeper feature extraction. For further improving the spatial learning ability, a pyramid dilated residual network is introduced by adopting several dilated residual ConvLSTM networks to extract multi-scale spatial information. In addition, a novel attention mechanism, which considers both long-term periodicity and the shift in periodicity, is designed to study diverse temporal patterns. Extensive experiments were conducted on three highly heterogeneous real-world mobility datasets to illustrate the effectiveness of PDSTN beyond the state-of-the-art methods. Moreover, PDSTN provides intuitive interpretation into the prediction. Congcong Miao, Jiajun Fu, Jilong Wang 0001, Heng Yu 0005, Botao Yao, Anqi Zhong, Zekun He |
WSDM | 1 |
| 2021 | Octans: Optimal Placement of Service Function Chains in Many-Core SystemsabstractNetwork Function Virtualization (NFV) offers service delivery flexibility and reduces overall costs by running service function chains (SFCs) on commodity servers with many cores. Existing solutions for placing SFCs in one server treat all CPU cores as equal and allocate isolated CPU cores to network functions (NFs). However, advanced servers often adopt Non-Uniform Memory Access (NUMA) architecture to improve the scalability of many-core systems. CPU cores are grouped into nodes, incurring performance degradation due to cross-node memory access and intra-node resource contention. Our evaluation shows that randomly selecting cores to place NFs in an SFC could suffer from 39.2 percent lower throughput comparing to an optimal placement solution. In this article, we propose Octans, an NFV orchestrator to achieve maximum aggregate throughput of all SFCs in many-core systems. Octans first formulates the optimization problem as a Non-Linear Integer Programming (NLIP) Model. Then we identify the key factor for problem solving as evaluating the throughput drop of an NF caused by other NFs in the same SFC or different SFCs, i.e., performance drop index, and propose a formal and accurate prediction model based on system level performance metrics. Finally, we propose two online algorithms to quickly find near-optimal placement solutions for one-time and incremental deployment. Extensive evaluation on a prototype implementation shows that Octans significantly improves the aggregate throughput comparing to two state-of-the-art placement solutions by 27.1 ~ 45.2 percent for one-time deployment and by 20.9 ~ 38.1 percent for incremental deployment, with very low prediction errors. Moreover, Octans could quickly find a near-optimal placement solution with tiny optimality gap. Heng Yu 0005, Zhilong Zheng, Junxian Shen, Congcong Miao, Chen Sun 0005, Hongxin Hu, Jun Bi, Jilong Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Predicting Human Mobility via Attentive Convolutional NetworkabstractPredicting human mobility is an important trajectory mining task for various applications, ranging from smart city planning to personalized recommendation system. While most of previous works adopt GPS tracking data to model human mobility, the recent fast-growing geo-tagged social media (GTSM) data brings new opportunities to this task. However, predicting human mobility on GTSM data is not trivial because of three challenges: 1) extreme data sparsity; 2) high order sequential patterns of human mobility and 3) evolving preference of users for tagging. Congcong Miao, Ziyan Luo, Fengzhu Zeng, Jilong Wang 0001 |
WSDM | 1 |
| 2019 | BDAC: A Behavior-aware Dynamic Adaptive Configuration on DHCP in Wireless LANsabstractDHCP is widely used to dynamically allocate IP addresses to the devices on local area networks, but the explosive increases of WiFi devices and their frequent mobility pose great challenges on DHCP performance in wireless LANs. In this paper, by analyzing large scale real network traces, we observe that the dynamic WiFi user behavior (e.g., online time pattern and spatio-temporal mobility pattern) leads to the poor DHCP performance. The IP pools in some VLANs have been exhausted in rush hours although the total IP utilization in WLAN is only 24%. Therefore, we have to configure IP lease times and IP pools dynamically and make sure that they are adaptive to the WiFi user behavior. In order to achieve this goal, we characterize and model the user behavior across online time pattern and spatiotemporal mobility pattern. Then we propose BDAC, a behaviour-aware dynamic adaptive configuration, which is combined of two strategies: adaptive IP lease time configuration and dynamic IP pool configuration. The former is to set adaptive lease times across user roles and area types based on online time pattern to reclaim IP addresses in time and reduce the peak IP usage, while the latter dynamically migrates the IP addresses across VLANs based on spatio-temporal mobility correlation to save the IP addresses. Using the real network traces of a different week, we conduct experiments to evaluate the performance of BDAC. Results show that BDAC can save up to 60% of IP addresses and the actual IP utilization rises from 24% to 59%. Furthermore, BDAC maintains high IP utilization when the number of VLANs in a WLAN increases. Congcong Miao, Jilong Wang 0001, Tianying Ji, Hui Wang 0011, Chao Xu 0015, Fengyuan Ren |
ICNP | 1 |
| 2018 | A Multi-dimension Measurement Study of a Large Scale Campus WiFi NetworkabstractThe growing trend of wireless devices and WiFi networks poses significant management challenges to network administrators. Characterizing WiFi user behavior and understanding WiFi network usage pattern are helpful to identify the management challenges so that network administrators could manage WiFi networks more efficiently. In this work, we collect comprehensive datasets, i.e., DHCP dataset, AAA dataset, SNMP dataset of ACs in a large campus WiFi network. We provide a detailed measurement study from multiple dimensions, i.e., server plane, temporal plane, spatial plane and traffic plane. We observe that the WiFi network under study is far from optimal. First, the phenomenon of IP waste is severe due to the isolation between DHCP server and AAA server. Second, current deployment of network infrastructure resources is based on network administrators' experience and it results in that the WiFi performance varies a lot across different areas. Furthermore, we also study the user behavior with different types of devices and in different kinds of buildings. Our observations indicate that the WiFi network could be improved and managed more efficiently from multiple dimensions. We believe that this measurement study is helpful for network administrators and researchers to understand more about large scale WiFi networks. Congcong Miao, Jilong Wang 0001, Hui Wang 0011, Jun Zhang 0004, Shengchao Liu |
LCN | 1 |
| 2015 | Measuring Photoplethysmogram-Based Stress-Induced Vascular Response Index to Assess Cognitive Load and StressabstractQuantitative assessment for cognitive load and mental stress is very important in optimizing human-computer system designs to improve performance and efficiency. Traditional physiological measures, such as heart rate variation (HRV), blood pressure and electrodermal activity (EDA), are widely used but still have limitations in sensitivity, reliability and usability. In this study, we propose a novel photoplethysmogram-based stress induced vascular index (sVRI) to measure cognitive load and stress. We also provide the basic methodology and detailed algorithm framework. We employed a classic experiment with three levels of task difficulty and three stages of testing period to verify the new measure. Compared with the blood pressure, heart rate and HRV components recorded simultaneously, the sVRI reached the same level of significance on the effect of task difficulty/period as the most significant other measure. Our findings showed sVRI's potential as a sensitive, reliable and usable parameter. Yongqiang Lyu 0001, Xiaomin Luo, Chun Yu, Congcong Miao, Yuanchun Shi, Ken-ichi Kameyama |
CHI | 5 |
| 2015 | QuickSync: Improving Synchronization Efficiency for Mobile Cloud Storage ServicesabstractMobile cloud storage services have gained phenomenal success in recent few years. In this paper, we identify, analyze and address the synchronization (sync) inefficiency problem of modern mobile cloud storage services. Our measurement results demonstrate that existing commercial sync services fail to make full use of available bandwidth, and generate a large amount of unnecessary sync traffic in certain circumstance even though the incremental sync is implemented. These issues are caused by the inherent limitations of the sync protocol and the distributed architecture. Based on our findings, we propose QuickSync, a system with three novel techniques to improve the sync efficiency for mobile cloud storage services, and build the system on two commercial sync services. Our experimental results using representative workloads show that QuickSync is able to reduce up to 52.9% sync time in our experiment settings. Yong Cui 0001, Zeqi Lai, Xin Wang 0001, Ningwei Dai, Congcong Miao |
MobiCom | 5 |