VLDB 2026 Research / reviewers in the wild / expert
Li Chen 0008
dblp:c/LiChen8
· DBLP profile ↗
105ranked-venue papers
9as first author
80since 2021 · last 2026
0000-0002-4228-7885ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 48 · 4 first-author · 36 since 2021Artificial intelligence and machine learning · 26 · 2 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 1 since 2021Systems, architecture and hardware · 8 · 7 since 2021Security and privacy · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OptiFlow: Towards LLM-Driven Optimization of Collective Communication Algorithms
Ziyue Yang 0002, Kaihui Gao, Shuai Wang 0028, Li Chen 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Dan Li 0001 |
APNet | 5 |
| 2026 | CATS: Predictive-Feedback Adaptive Load Balancing for Computing-Aware Traffic Steering
Yuxiang Shang, Tao Sun 0010, Dan Li 0001, Zhenping Hu, Lu Lu 0016, Chengjiang Wen, Yantao Han, Li Chen 0008, Huijuan Yao, Peng Liu 0047 |
ICC | 10 |
| 2026 | Accurate and Stable AS Relationship Inference via Trusted Seeds and Semi-Supervised Learning
Siyuan Teng, Lancheng Qin, Li Chen 0008, Dan Li 0001 |
INFOCOM | 3 |
| 2026 | A Large-Scale IPv6-Based Measurement of the Starlink NetworkabstractLow Earth Orbit (LEO) satellite networks have attracted considerable attention for their ability to deliver global, low-latency broadband Internet services. In this paper, we present a large-scale measurement study of the Starlink network, the largest LEO satellite constellation to date. We first propose an efficient method for discovering active Starlink user routers, identifying approximately 5.98 million IPv6 addresses across 208 regions in 165 countries. Compared to general-purpose IPv6 target generation algorithms, our router-centric approach achieves near-complete coverage and, to the best of our knowledge, yields the most comprehensive known set of active IPv6 addresses for Starlink user routers. Based on the discovered user routers, we further propose an efficient method for mapping the Starlink backbone network and uncover a topology consisting of 49 Points of Presence (PoPs) interconnected by 98 links. We conduct a detailed statistical analysis of active Starlink user routers and PoPs, and further characterize the IPv6 address assignment strategy adopted by the Starlink network. Finally, we analyze the latency of Starlink user routers, propose a method to distinguish different types of users within the same region using outside-in measurement, and identify the ongoing V2 Mini satellite deployment as a potential driver of the performance improvements. The dataset of the Starlink backbone network is publicly available at https://ki3.org.cn/#/starlink-network. Bingsen Wang, Shuai Wang 0028, Li Chen 0008, Jinwei Zhao, Dan Li 0001, Yong Jiang 0001 |
INFOCOM | 4 |
| 2026 | HeraClass: Towards Open-World Network Flow Classification via Traffic-Language Mapping
Ni Jin, Libin Liu 0001, Yukai Miao, Li Chen 0008, Dan Li 0001, Xizheng Wang, Xiuting Xu, Baojiang Cui |
IWQoS | 4 |
| 2026 | OSAVRoute: Advancing Outbound Source Address Validation Deployment Detection with Non-Cooperative Measurement
Shuai Wang 0028, Li Chen 0008, Dan Li 0001, Lancheng Qin |
NDSS | 3 |
| 2026 | BayWatch: Practical Internet-Scale Topology Monitoring with Dynamic Bayesian Estimation
Zhongxu Guan, Shuai Wang 0028, Li Chen 0008, Zhaoteng Yan, Jiaye Lin, Dan Li 0001, Yong Jiang 0001, Yingxin Wang |
NSDI | 3 |
| 2026 | CCEval: Accurately and Confidently Evaluating Performance Metrics of Congestion Control Algorithms for Datacenter Networks
Tianfeng Liu, Kaihui Gao, Li Chen 0008, Dan Li 0001, Jin Guang, Vincent Liu 0001, Yiwei Zhang 0016, Ni Jin |
NSDI | 3 |
| 2026 | Supercharging Packet-level Network Simulation of Large Model Training via Memoization and Fast-Forwarding
Kaihui Gao, Li Chen 0008, Dan Li 0001, Yiwei Zhang 0016, Fei Gui, Yitao Xing, Wenjia Wei, Bingyang Liu |
NSDI | 3 |
| 2026 | REAL: Emulating Control Plane at Simulator's Cost
Ze Xia, Hao Li 0011, Jinyu Fu, Yihan Dang, Danfeng Shan, Li Chen 0008, Peng Zhang 0011 |
NSDI | 7 |
| 2026 | PReCCL: Performant and Resilient Collective Communication via Integrated Inband Telemetry and Workload ReallocationabstractModern collective communication libraries (CCLs) execute a collective communication task (CCT) by decomposing it into multiple sub-tasks, each mapped to a specific Virtual Topology (VT), which is an ordered graph of GPUs (e.g., a ring or a tree), to maximize parallelism and link utilization. As AI training scales to larger clusters, network anomalies (congestion and failures) are unavoidable, and a single straggling VT can delay the entire CCT. Existing solutions either rely on low-level transport-layer solutions which lacks a cross-sub-task perspective, or static CCL scheduling, failing to adapt to the dynamic and heterogeneous networks. Kaihui Gao, Li Chen 0008, Fei Gui, Dan Li 0001, Jiamin Cao |
SIGCOMM | 3 |
| 2026 | Networked Agent Memory and Causality Representation: Experiences towards Interpretable Cloud-Scale Root-Causing
Yanyu Ren, Xianshang Lin, Chenxu Wang 0007, Li Chen 0008, Shuai Wang 0028, Kaihui Gao, Dan Li 0001, Chen Tian 0001, Yunguang Li, Ennan Zhai |
SIGCOMM | 4 |
| 2026 | Open the Floodgates in a Digital Twin: Experiences of Building Spillway for 100M+-User Signaling Storms in Cellular Core NetworkabstractSignaling storms threaten cellular core networks when synchronized reconnection attempts from massive numbers of devices trigger cascading, metastable overloads. Existing defenses rely on manual, static configurations of local overload controls, which ignore serial dependencies among heterogeneous network elements. We present Spillway, a digital-twin-driven system that automates global signaling-flood mitigation. Spillway introduces a hierarchical defense architecture that enforces altruistic throttling, allowing upstream nodes to shed load before downstream bottlenecks collapse. To evaluate candidate configurations, Spillway uses CN-DES, a domain-specific discrete-event simulator with a vectorized kernel. By aggregating users that share protocol states, CN-DES decouples simulation cost from user count and simulates regional-scale storms involving tens of millions of users in minutes, achieving a 60× speedup over traditional simulation while preserving fidelity. Spillway then uses heteroscedastic evolutionary Bayesian optimization to search a large, non-convex parameter space. We report on a five-year deployment in the world's largest 5G Standalone network. During real incidents, including application anomalies and RAN failures, networks using Spillway-optimized configurations experienced substantially fewer user fallbacks than predicted under legacy configurations; post-incident analysis confirms that pre-deployed parameters kept all network elements within safe operating bounds. Hongtao Xie 0006, Jianmin Liu, Li Chen 0008, Dan Li 0001, Mineng Fu, Xi Chen 0026 |
SIGCOMM | 5 |
| 2026 | Reinforced Refinement With Self-Aware Expansion for End-to-End Autonomous DrivingabstractEnd-to-end autonomous driving has emerged as a promising paradigm for directly mapping sensor inputs to planning maneuvers using learning-based modular integrations. However, existing imitation learning (IL)-based models suffer from generalization to hard cases, and a lack of corrective feedback loop under post-deployment. While reinforcement learning (RL) offers a potential solution to tackle hard cases with optimality, it is often hindered by overfitting to specific driving cases, resulting in catastrophic forgetting of generalizable knowledge and sample inefficiency. To overcome these challenges, we propose Reinforced Refinement with Self-aware Expansion (R2SE), a novel learning pipeline that constantly refines hard domain while keeping generalizable driving policy for model-agnostic end-to-end driving systems. Through reinforcement fine-tuning and policy expansion that facilitates continuous improvement, R2SE features three key components: 1) Generalist Pretraining with hard-case allocation trains a generalist imitation learning (IL) driving system while dynamically identifying failure-prone cases for targeted refinement; 2) Residual Reinforced Specialist Fine-tuning optimizes residual corrections using reinforcement learning (RL) to improve performance in hard case domain while preserving global driving knowledge; 3) Self-aware Adapter Expansion dynamically integrates specialist policies back into the generalist model, enhancing continuous performance improvement. Experimental results in closed-loop simulation and real-world datasets demonstrate improvements in generalization, safety, and long-horizon policy robustness over state-of-the-art E2E systems, highlighting the effectiveness of reinforce refinement for scalable autonomous driving. Tianyu Li 0004, Haohan Yang, Li Chen 0008, Caojun Wang, Haochen Tian 0001, Hongyang Li 0001, Chen Lv 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Test-Time Correction: An Online 3D Detection System via Visual PromptingabstractThis paper introduces Test-time Correction (TTC), an online 3D detection system designed to rectify test-time errors using various auxiliary feedback, aiming to enhance the safety of deployed autonomous driving systems. Unlike conventional offline 3D detectors that remain fixed during inference, TTC enables immediate online error correction without retraining, allowing autonomous vehicles to adapt to new scenarios and reduce deployment risks. To achieve this, we equip existing 3D detectors with an Online Adapter (OA) module-a prompt-driven query generator for real-time correction. At the core of OA module are visual prompts: image-based descriptions of objects of interest derived from auxiliary feedback such as mismatches with 2D detections, road descriptions, or user clicks. These visual prompts, collected from risky objects during inference, are maintained in a visual prompt buffer to enable continuous correction in future frames. By leveraging this mechanism, TTC consistently detects risky objects, achieving reliable, adaptive, and versatile driving autonomy. Extensive experiments show that TTC significantly improves instant error rectification over frozen 3D detectors, even under limited labels, zero-shot settings, and adverse conditions. We hope this work inspires future research on post-deployment online rectification systems for autonomous driving. Hanxue Zhang, Zetong Yang, Yanan Sun 0005, Li Chen 0008, Fatma Güney, Hongyang Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Example Generalizing Network Configuration Synthesizer via Graph-Informed Large Language Models
Jianmin Liu, Li Chen 0008, Dan Li 0001, Yukai Miao, Liyu Ma |
IEEE Trans. Netw. | 2 |
| 2026 | Is Diversity All You Need for Scalable Robotic Manipulation?abstractData scaling has driven remarkable success in foundation models for Natural Language Processing (NLP) and Computer Vision (CV), yet the principles of effective data scaling in robotic manipulation remain insufficiently understood. In this work, we investigate the nuanced role of data diversity in robot learning by examining three critical dimensions-task (what to do), embodiment (which robot to use), and expert (who demonstrates)-challenging the conventional intuition of “more diverse is better”. Throughout extensive experiments on various robot platforms, we reveal that (1) task diversity proves more critical than per-task demonstration quantity, with scene diversity playing a more important role than skill diversity for robustness and generalization under distribution shifts; (2) multi-embodiment pre-training data is non-essential for cross-embodiment transfer-models trained on high-quality single-embodiment data can efficiently transfer to different platforms, showing desirable scaling property during fine-tuning and its potential of replacing large-scale multi-embodiment pre-training; and (3) expert diversity, arising from individual operational preferences and stochastic variations in human demonstrations, can be confounding to policy learning, with action rate multimodality emerging as a key contributing factor. Based on this insight, we propose a distribution debiasing method to mitigate action rate ambiguity, the yielding GO-1-Pro achieves substantial performance gains of 15%, equivalent to using 2.5× pre-training data. Collectively, these findings provide new perspectives and offer practical guidance on how to scale robotic manipulation datasets effectively. The code will be released. Modi Shi, Li Chen 0008, Chiming Liu, Guanghui Ren, Ping Luo 0002, Di Huang 0001, Maoqing Yao, Hongyang Li 0001 |
IEEE Trans. Robotics | 2 |
| 2025 | Towards Automatic Network Diagram ComprehensionabstractNetwork Diagram Comprehension (NDC) is a vital task for networking professionals, offering essential insights into network topology and configurations. However, NDC remains a labor-intensive process heavily reliant on human expertise, with existing tools falling short in addressing this challenge. It is critical to develop an Automatic NDC (ANDC) system that ensures high faithfulness and completeness in information extraction while supporting practical, end-to-end NDC applications. Moreover, a comprehensive dataset and benchmark are necessary to systematically evaluate and drive the progress of ANDC.In this work, we introduce Layered Extractor of Network Diagrams (LEND), the first ANDC system designed to comprehensively and faithfully extract and utilize information from network diagrams. LEND employs a three-stage pipeline: (1) a layer extractor to decompose diagrams and identify key elements with a denoising cascade, (2) an inter-layer combiner to reconstruct entity relations with positional and domain knowledge, and (3) a task-specific interpreter for networking applications.To support this effort, we develop two extensive NDC datasets comprising over 4,000 network diagrams and icons from diverse sources, along with the first benchmark to evaluate ANDC systems across three distinct metrics. Empirical experiments demonstrate that LEND outperforms existing methods by achieving at 1.21– 5.10× better faithfulness and completeness, and improves its capability as a NetOps engineer by 30.5% on the Cisco Certified Network Associate (CCNA) exam. Yanyu Ren, Yukai Miao, Li Chen 0008, Dan Li 0001, Xizheng Wang, Yu Bai 0021 |
ICNP | 3 |
| 2025 | Measuring the Time Source Vulnerabilities in the NTP EcosystemabstractPrecise timekeeping is crucial for the dependable functioning and security of multiple Internet infrastructures, such as TLS certificates. Although the Network Time Protocol (NTP) is widely used for time synchronization across devices, it has several security vulnerabilities. Network Time Security (NTS) offers server authentication and integrity verification to protect against man-in-the-middle attacks. However, NTS does not address issues related to erroneous time sources. Zhentian Huang, Shuai Wang 0028, Li Chen 0008, Dan Li 0001 |
IMC | 3 |
| 2025 | SAIP: Accurate Detection of Anycast Servers with the Rise of Regional Anycast
Shuai Wang 0028, Li Chen 0008, Dan Li 0001 |
INFOCOM | 3 |
| 2025 | Exploiting Student Parallelism for Low-latency GPU Inference of BERT-like Models in Online ServicesabstractBERT-like models have been widely adopted in text mining and web search due to their high accuracy. However, large BERT-like models suffer from inefficient online inference on GPUs for two main reasons. First, their high accuracy relies on large model depth, which linearly increases sequential computation on GPUs. Second, stochastic and dynamic online workloads lead to extra costs due to batching and padding. To address the problem, we present Student Parallelism for efficient GPU inference of BERT-like models under real-world online workloads. At its core, Student Parallelism adopts stacking distillation and boosting ensemble, distilling the original deep model into a group of shallow but virtually stacked student models running in parallel. This enables Student Parallelism to achieve a low model depth (e.g., two layers), and thus low inference latency while maintaining accuracy. In addition, we design adaptive student pruning to adjust the number of students according to the dynamic online workloads. For example, during workload bursts, it can temporarily decrease the number of students with minimal accuracy loss to improve system throughput. Extensive experiments on real-world datasets and workloads show that Student Parallelism achieves up to 4.1× lower latency while maintaining accuracy and up to 22.27× higher throughput during workload bursts. Weiyan Wang, Yilun Jin, Yiming Zhang 0003, Victor Junqiu Wei, Han Tian, Li Chen 0008, Jinbao Xue, Yangyu Tao, Kai Chen 0005 |
KDD (2) | 6 |
| 2025 | HoloTrace: LLM-based Bidirectional Causal Knowledge Graph for Edge-Cloud Video Anomaly DetectionabstractVideo anomaly detection (VAD) is vital for public safety, yet current approaches struggle with limited generalization, low interpretability, and high resource demands. To address these challenges, we propose HoloTrace, an edge-cloud collaborative VAD system that integrates large language models (LLMs) to construct and update a novel bidirectional causal knowledge graph. At the edge, HoloTrace leverages LLM-based cross-modal understanding and employs Hidden Markov Model (HMM) for bidirectional event reasoning, obtaining anomaly boundaries with low computational overhead. On the cloud side, LLMs are leveraged to dynamically update the Bi-CKG graph with key frames sent from the edge, in order to update causal relationships between events. Additionally, we introduce SVAD, a new large-scale VAD dataset comprising 632 real-world surveillance videos across 10 anomaly types and diverse scenes, with manually labeled frame-level annotations. Experimental results demonstrate that HoloTrace not only achieves the highest accuracy but also enhances interpretability and efficiency, paving the way for more generalizable and explainable video anomaly detection systems. Hanling Wang, Qing Li 0006, Li Chen 0008, Haidong Kang, Fei Ma 0006, Yong Jiang 0001 |
ACM Multimedia | 3 |
| 2025 | Transcending Cost-Quality Tradeoff in Agent Serving via Session-AwarenessabstractLarge Language Model (LLM) agents are capable of task execution across various domains by autonomously interacting with environments and refining LLM responses based on feedback.
However, existing model serving systems are not optimized for the unique demands of serving agents. Compared to classic model serving, agent serving has different characteristics:
predictable request pattern, increasing quality requirement, and unique prompt formatting. We identify a key problem for agent serving: LLM serving systems lack session-awareness. They neither perform effective KV cache management nor precisely select the cheapest yet competent model in each round.
This leads to a cost-quality tradeoff, and we identify an opportunity to surpass it in an agent serving system.
To this end, we introduce AgServe for AGile AGent SERVing.
AgServe features a session-aware server that boosts KV cache reuse via Estimated-Time-of-Arrival-based eviction and in-place positional embedding calibration, a quality-aware client that performs session-aware model cascading through real-time quality assessment, and a dynamic resource scheduler that maximizes GPU utilization.
With AgServe, we allow agents to select and upgrade models during the session lifetime, and to achieve similar quality at much lower costs, effectively transcending the tradeoff. Extensive experiments on real testbeds demonstrate that AgServe (1) achieves comparable response quality to GPT-4o at a 16.5\% cost. (2) delivers 1.8$\times$ improvement in quality relative to the tradeoff curve. Yanyu Ren, Li Chen 0008, Dan Li 0001, Xizheng Wang, Yukai Miao, Yu Bai 0021 |
NeurIPS | 2 |
| 2025 | ReSim: Reliable World Simulation for Autonomous DrivingabstractHow can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusively on real-world driving data composed mainly of safe expert trajectories, struggle to follow hazardous or non-expert behaviors, which are rare in such data. This limitation restricts their applicability to tasks such as policy evaluation. In this work, we address this challenge by enriching real-world human demonstrations with diverse non-expert data collected from a driving simulator (e.g., CARLA), and building a controllable world model trained on this heterogeneous corpus. Starting with a video generator featuring diffusion transformer architecture, we devise several strategies to effectively integrate conditioning signals and improve prediction controllability and fidelity. The resulting model, ReSim, enables Reliable Simulation of diverse open-world driving scenarios under various actions, including hazardous non-expert ones. To close the gap between high-fidelity simulation and applications that require reward signals to judge different actions, we introduce a Video2Reward module that estimates reward from ReSim’s simulated future. Our ReSim paradigm achieves up to 44% higher visual fidelity, improves controllability for both expert and non-expert actions by over 50%, and boosts planning and policy selection performance on NAVSIM by 2% and 25%, respectively. Jiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen 0005, Yuqian Shao, Xiaosong Jia, Hongyang Li 0001, Andreas Geiger 0001, Xiangyu Yue 0001, Li Chen 0008 |
NeurIPS | 10 |
| 2025 | Accelerating Design Space Exploration for LLM Training Systems with Multi-experiment Parallel Simulation
Fei Gui, Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Hongbing Yang, Dian Xiong |
NSDI | 3 |
| 2025 | CEGS: Configuration Example Generalizing Synthesizer
Jianmin Liu, Li Chen 0008, Dan Li 0001, Yukai Miao |
NSDI | 2 |
| 2025 | Resolving Packets from Counters: Enabling Multi-scale Network Traffic Super Resolution via Composable Large Traffic Model
Xizheng Wang, Libin Liu 0001, Li Chen 0008, Dan Li 0001, Yukai Miao, Yu Bai 0021 |
NSDI | 3 |
| 2025 | SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and Precision
Xizheng Wang, Qingxu Li, Yichi Xu, Dan Li 0001, Li Chen 0008, Heyang Zhou, Linkang Zheng, Yikai Zhu, Yang Liu 0245, Kun Qian 0021, Kunling He, Ennan Zhai, Dennis Cai, Binzhang Fu |
NSDI | 6 |
| 2025 | Discovering Millions of New Nodes and Links in the Internet by Challenging the Uniformity Assumption in Multipath DetectionabstractMultipath Detection Algorithms (MDAs) are proposed to discover Internet topology in the presence of load balancing (LB). Existing methods assume uniformity in the load-balancing responses (LBR), i.e., responses from the successors of a LB router. However, we reveal that only 20% of the cases exhibit uniformity in the Internet. This finding significantly challenges the completeness of the Internet topology discovered using current MDAs. In this paper, we propose a novel system BayMuDA, that can estimate LBR distributions and calculate the minimum number of probes needed to statistically discover all nodes and links within a given hop. The validation on controlled topologies shows that BayMuDA discovers at least 85%/73% of nodes/links in ~90% of the cases. Our Internet-wide measurement results indicate that BayMuDA can discover millions of Internet nodes and links obscured by the state-of-the-art MDA algorithm, D-Miner, due to uneven responses. Zhongxu Guan, Shuai Wang 0028, Li Chen 0008, Zhaoteng Yan, Jiaye Lin, Dan Li 0001, Yong Jiang 0001, Yingxin Wang |
SIGCOMM | 3 |
| 2025 | From ATOP to ZCube: Automated Topology Optimization Pipeline and A Highly Cost-Effective Network Topology for Large Model TrainingabstractThe development of large language models (LLMs) poses new challenges in data center network topology design. To assist in exploring topology design, we propose ATOP, an Automated Topology Optimization Pipeline, which models network topology as a set of hyperparameters, enabling the discovery of potential topologies. With various optimization algorithms and customizable optimization objectives, ATOP achieves automated topology optimization on a scale of tens of thousands of GPUs. We apply ATOP on network topologies for 256, 1024, 4096, and 16384 GPUs, optimizing performance under LLMs training traffic patterns, collective communication performance, fault tolerance, and network cost. We also evaluate ATOP in different scenarios: building, optimizing, and expanding a data center. From ATOP's results, we discover a new topology — ZCube, which reaches the highest cost-effectiveness across various GPU scales. Simulation results show that ZCube, compared to the previous state-of-the-art topologies, including Rail-optimized Fat-tree (ROFT), Rail-only, and HPN, improves end-to-end LLM training speed by 3% to 7% and reduces network hardware costs by 26% to 46%. We also construct ZCube on a real-world testbed. Results show that ZCube reduces hardware costs by 25% compared to Rail-Optimized Topology while maintaining the same all-reduce and all-to-all performance. Dan Li 0001, Li Chen 0008, Dian Xiong, Kaihui Gao, Yiwei Zhang 0016, Menglei Zhang, Bochun Zhang, Zhuo Jiang, Jianxi Ye, Haibin Lin |
SIGCOMM | 3 |
| 2025 | The Digital Cybersecurity Expert: How Far Have We Come?abstractThe increasing deployment of large language models (LLMs) in the cybersecurity domain underscores the need for effective model selection and evaluation. However, traditional evaluation methods often overlook specific cybersecurity knowledge gaps that contribute to performance limitations. To address this, we develop CSEBenchmark, a fine-grained cyber-security evaluation framework based on 345 knowledge points expected of cybersecurity experts. Drawing from cognitive science, these points are categorized into factual, conceptual, and procedural types, enabling the design of 11,050 tailored multiple-choice questions. We evaluate 12 popular LLMs on CSEBenchmark and find that even the best-performing model achieves only 85.42% overall accuracy, with particular knowledge gaps in the use of specialized tools and uncommon commands. Different LLMs have unique knowledge gaps. Even large models from the same family may perform poorly on knowledge points where smaller models excel. By identifying and addressing specific knowledge gaps in each LLM, we achieve up to an 84% improvement in correcting previously incorrect predictions across three existing benchmarks for two cybersecurity tasks. Furthermore, our assessment of each LLM's knowledge alignment with specific cybersecurity roles reveals that different models align better with different roles, such as GPT-4o for the Google Senior Intelligence Analyst and Deepseek-V3 for the Amazon Privacy Engineer. These findings underscore the importance of aligning LLM selection with the specific knowledge requirements of different cybersecurity roles for optimal performance. Dawei Wang 0021, Geng Zhou, Yu Bai 0021, Li Chen 0008, Ting Qin, Dan Li 0001 |
SP | 5 |
| 2025 | Your Shield is My Sword: A Persistent Denial-of-Service Attack via the Reuse of Unvalidated Caches in DNSSEC Validation
Shuai Wang 0028, Li Chen 0008, Dan Li 0001 |
USENIX Security Symposium | 3 |
| 2024 | ProphetFuzz: Fully Automated Prediction and Fuzzing of High-Risk Option Combinations with Only Documentation via Large Language ModelabstractVulnerabilities related to option combinations pose a significant challenge in software security testing due to their vast search space. Previous research primarily addressed this challenge through mutation or filtering techniques, which inefficiently treated all option combinations as having equal potential for vulnerabilities, thus wasting considerable time on non-vulnerable targets and resulting in low testing efficiency. In this paper, we utilize carefully designed prompt engineering to drive the large language model (LLM) to predict high-risk option combinations (i.e., more likely to contain vulnerabilities) and perform fuzz testing automatically without human intervention. We developed a tool called ProphetFuzz and evaluated it on a dataset comprising 52 programs collected from three related studies. The entire experiment consumed 10.44 CPU years. ProphetFuzz successfully predicted 1748 high-risk option combinations at an average cost of only \8.69 per program. Results show that after 72 hours of fuzzing, ProphetFuzz discovered 364 unique vulnerabilities associated with 12.30% of the predicted high-risk option combinations, which was 32.85% higher than that found by state-of-the-art in the same timeframe. Additionally, using ProphetFuzz, we conducted persistent fuzzing on the latest versions of these programs, uncovering 140 vulnerabilities, with 93 confirmed by developers and 21 awarded CVE numbers. Dawei Wang 0021, Geng Zhou, Li Chen 0008, Dan Li 0001, Yukai Miao |
CCS | 3 |
| 2024 | Visual Point Cloud Forecasting Enables Scalable Autonomous DrivingabstractIn contrast to extensive studies on general vision, pretraining for scalable visual autonomous driving remains seldom explored. Visual autonomous driving applications require features encompassing semantics, 3D geometry, and temporal information simultaneously for joint perception, prediction, and planning, posing dramatic challenges for pre-training. To resolve this, we bring up a new pre-training task termed as visual point cloud forecasting - predicting future point clouds from historical visual input. The key merit of this task captures the synergic learning of semantics, 3D structures, and temporal dynamics. Hence it shows superiority in various downstream tasks. To cope with this new problem, we present ViDAR, a general model to pre-train downstream visual encoders. It first extracts historical embeddings by the encoder. These representations are then transformed to 3D geometric space via a novel Latent Rendering operator for future point cloud prediction. Ex-periments show significant gain in downstream tasks, e.g., 3.1 % NDS on 3D detection, ~10% error reduction on motion forecasting, and ~ 15% less collision rate on planning. Zetong Yang, Li Chen 0008, Yanan Sun 0005, Hongyang Li 0001 |
CVPR | 2 |
| 2024 | Generalized Predictive Model for Autonomous DrivingabstractIn this paper, we introduce the first large-scale video prediction model in the autonomous driving discipline. To eliminate the restriction of high-cost data collection and empower the generalization ability of our model, we ac-quire massive data from the web and pair it with diverse and high-quality text descriptions. The resultant dataset accumulates over 2000 hours of driving videos, spanning areas all over the world with diverse weather conditions and traffic scenarios. Inheriting the merits from recent latent diffusion models, our model, dubbed GenAD, handles the challenging dynamics in driving scenes with novel tem-poral reasoning blocks. We showcase that it can general-ize to various unseen driving datasets in a zero-shot man-ner, surpassing general or driving-specific video prediction counterparts. Furthermore, GenAD can be adapted into an action-conditioned prediction model or a motion planner, holding great potential for real-world driving applications. Jiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen 0008, Tianyu Li 0004, Bo Dai 0002, Kashyap Chitta, Penghao Wu, Ping Luo 0002, Jun Zhang 0106, Andreas Geiger 0001, Yu Qiao 0001, Hongyang Li 0001 |
CVPR | 4 |
| 2024 | Fully Sparse 3D Occupancy Prediction
Haisong Liu, Zetong Yang, Tianyu Li 0004, Li Chen 0008, Hongyang Li 0001, Limin Wang 0002 |
ECCV (25) | 7 |
| 2024 | DriveLM: Driving with Graph Visual Question Answering
Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen 0008, Hanxue Zhang, Chengen Xie, Jens Beißwenger, Ping Luo 0002, Andreas Geiger 0001, Hongyang Li 0001 |
ECCV (52) | 4 |
| 2024 | LaneSegNet: Map Learning with Lane Segment Perception for Autonomous DrivingabstractA map, as crucial information for downstream applications of an autonomous driving system, is usually represented in lanelines or centerlines. However, existing literature on map learning primarily focuses on either detecting geometry-based lanelines or perceiving topology relationships of centerlines. Both of these methods ignore the intrinsic relationship of lanelines and centerlines, that lanelines bind centerlines. While simply predicting both types of lane in one model is mutually excluded in learning objective, we advocate lane segment as a new representation that seamlessly incorporates both geometry and topology information. Thus, we introduce LaneSegNet, the first end-to-end mapping network generating lane segments to obtain a complete representation of the road structure. Our algorithm features two key modifications. One is a lane attention module to capture pivotal region details within the long-range feature space. Another is an identical initialization strategy for reference points, which enhances the learning of positional priors for lane attention. On the OpenLane-V2 dataset, LaneSegNet outperforms previous counterparts by a substantial gain across three tasks, i.e., map element detection (+4.8 mAP), centerline perception (+6.9 DET$_l$), and the newly defined one, lane segment perception (+5.6 mAP). Furthermore, it obtains a real-time inference speed of 14.7 FPS. Code is accessible at https://github.com/OpenDriveLab/LaneSegNet. Tianyu Li 0004, Peijin Jia, Bangjun Wang, Li Chen 0008, Kun Jiang 0002, Junchi Yan, Hongyang Li 0001 |
ICLR | 4 |
| 2024 | Fat-B+Tree: Fast B+tree Indexing with In-Network MemoryabstractIn-memory database in the data center plays an indispensable role in many fields. B+tree is the most recognized index in in-memory database, but its indexing latency has become the bottleneck that prevents the database from achieving higher performance. The existing work reduces latency by caching B+tree nodes using slow DRAM on computing clients, or directly caching data using fast SRAM on programmable switch. However, no existing work can meet all three key requirements: (1) Efficiency: providing enough fast memory to accelerate indexing. (2) Compatibility: compatible with multiple architectures and query types. (3) Adaptability: adapting to various workloads and database scales. Inspired by structrual similarity between the network topology of modern data centers and the B+tree structure, we propose Fat-B+Tree, which embeds the B+tree into the in-network fast memory provided by the hierarchical connected programmable switches, thus meeting all design requirements. We have fully implemented the Fat-B+Tree prototype, and the experimental results show that compared with the baseline system, it reduces query latency by up to 76% and improves throughput by up to 3.93 times. The source codes of Fat-B+Tree are open-sourced at GitHub. Yikai Zhao 0001, Yuanpeng Li 0002, Zicang Xu, Tong Yang 0003, Kaicheng Yang 0001, Li Chen 0008, Xin Yao 0008, Gong Zhang 0001 |
IPCCC | 6 |
| 2024 | Understanding Route Origin Validation (ROV) Deployment in the Real World and Why MANRS Action 1 Is Not Followed
Lancheng Qin, Li Chen 0008, Dan Li 0001, Honglin Ye |
NDSS | 2 |
| 2024 | dRR: A Decentralized, Scalable, and Auditable Architecture for RPKI Repository
Dan Li 0001, Li Chen 0008, Qi Li 0002, Sitong Ling |
NDSS | 3 |
| 2024 | Closed-Loop Visuomotor Control with Generative Expectation for Robotic ManipulationabstractDespite significant progress in robotics and embodied AI in recent years, deploying robots for long-horizon tasks remains a great challenge. Majority of prior arts adhere to an open-loop philosophy and lack real-time feedback, leading to error accumulation and undesirable robustness. A handful of approaches have endeavored to establish feedback mechanisms leveraging pixel-level differences or pre-trained visual representations, yet their efficacy and adaptability have been found to be constrained. Inspired by classic closed-loop control systems, we propose CLOVER, a closed-loop visuomotor control framework that incorporates feedback mechanisms to improve adaptive robotic control. CLOVER consists of a text-conditioned video diffusion model for generating visual plans as reference inputs, a measurable embedding space for accurate error quantification, and a feedback-driven controller that refines actions from feedback and initiates replans as needed. Our framework exhibits notable advancement in real-world robotic tasks and achieves state-of-the-art on CALVIN benchmark, improving by 8% over previous open-loop counterparts. Code and checkpoints are maintained at https://github.com/OpenDriveLab/CLOVER. Qingwen Bu, Li Chen 0008, Yanchao Yang 0001, Guyue Zhou, Junchi Yan, Ping Luo 0002, Heming Cui, Yi Ma 0001, Hongyang Li 0001 |
NeurIPS | 3 |
| 2024 | Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityabstractWorld models can foresee the outcomes of different actions, which is of paramount importance for autonomous driving. Nevertheless, existing driving world models still have limitations in generalization to unseen environments, prediction fidelity of critical details, and action controllability for flexible application. In this paper, we present Vista, a generalizable driving world model with high fidelity and versatile controllability. Based on a systematic diagnosis of existing methods, we introduce several key ingredients to address these limitations. To accurately predict real-world dynamics at high resolution, we propose two novel losses to promote the learning of moving instances and structural information. We also devise an effective latent replacement approach to inject historical frames as priors for coherent long-horizon rollouts. For action controllability, we incorporate a versatile set of controls from high-level intentions (command, goal point) to low-level maneuvers (trajectory, angle, and speed) through an efficient learning strategy. After large-scale training, the capabilities of Vista can seamlessly generalize to different scenarios. Extensive experiments on multiple datasets show that Vista outperforms the most advanced general-purpose video generator in over 70% of comparisons and surpasses the best-performing driving world model by 55% in FID and 27% in FVD. Moreover, for the first time, we utilize the capacity of Vista itself to establish a generalizable reward for real-world action evaluation without accessing the ground truth actions. Shenyuan Gao, Jiazhi Yang, Li Chen 0008, Kashyap Chitta, Yihang Qiu, Andreas Geiger 0001, Jun Zhang 0106, Hongyang Li 0001 |
NeurIPS | 3 |
| 2024 | Klonet: an Easy-to-Use and Scalable Platform for Computer Networks Education
Tie Ma, Long Luo, Hong-Fang Yu, Xi Chen 0026, Jingzhao Xie, Chongxi Ma, Yunhan Xie, Gang Sun 0001, Tianxi Wei, Li Chen 0008, Yanwei Xu 0004, Nicholas Zhang |
NSDI | 10 |
| 2024 | RedTE: Mitigating Subsecond Traffic Bursts with Real-time and Distributed Traffic EngineeringabstractInternet traffic bursts usually happen within a second, thus conventional burst mitigation methods ignore the potential of Traffic Engineering (TE). However, our experiments indicate that a TE system, with a sub-second control loop latency, can effectively alleviate burst-induced congestion. TE-based methods can leverage network-wide tunnel-level information to make globally informed decisions (e.g., balancing traffic bursts among multiple paths). Our insight in reducing control loop latency is to let each router make local TE decisions, but this introduces the key challenge of minimizing performance loss compared to centralized TE systems. Fei Gui, Dan Li 0001, Li Chen 0008, Kaihui Gao, Congcong Min, Yi Wang 0004 |
SIGCOMM | 4 |
| 2024 | End-to-End Autonomous Driving: Challenges and FrontiersabstractThe autonomous driving community has witnessed a rapid growth in approaches that embrace an end-to-end algorithm framework, utilizing raw sensor input to generate vehicle motion plans, instead of concentrating on individual tasks such as detection and motion prediction. End-to-end systems, in comparison to modular pipelines, benefit from joint feature optimization for perception and planning. This field has flourished due to the availability of large-scale datasets, closed-loop evaluation, and the increasing need for autonomous driving algorithms to perform effectively in challenging scenarios. In this survey, we provide a comprehensive analysis of more than 270 papers, covering the motivation, roadmap, methodology, challenges, and future trends in end-to-end autonomous driving. We delve into several critical challenges, including multi-modality, interpretability, causal confusion, robustness, and world models, amongst others. Additionally, we discuss current advancements in foundation models and visual pre-training, as well as how to incorporate these techniques within the end-to-end driving framework. Li Chen 0008, Penghao Wu, Kashyap Chitta, Bernhard Jaeger, Andreas Geiger 0001, Hongyang Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and RecipeabstractLearning powerful representations in bird's-eye-view (BEV) for perception tasks is trending and drawing extensive attention both from industry and academia. Conventional approaches for most autonomous driving algorithms perform detection, segmentation, tracking, etc., in a front or perspective view. As sensor configurations get more complex, integrating multi-source information from different sensors and representing features in a unified view come of vital importance. BEV perception inherits several advantages, as representing surrounding scenes in BEV is intuitive and fusion-friendly; and representing objects in BEV is most desirable for subsequent modules as in planning and/or control. The core problems for BEV perception lie in (a) how to reconstruct the lost 3D information via view transformation from perspective view to BEV; (b) how to acquire ground truth annotations in BEV grid; (c) how to formulate the pipeline to incorporate features from different sources and views; and (d) how to adapt and generalize algorithms as sensor configurations vary across different scenarios. In this survey, we review the most recent works on BEV perception and provide an in-depth analysis of different solutions. Moreover, several systematic designs of BEV approach from the industry are depicted as well. Furthermore, we introduce a full suite of practical guidebook to improve the performance of BEV perception tasks, including camera, LiDAR and fusion inputs. At last, we point out the future research directions in this area. We hope this report will shed some light on the community and encourage more research effort on BEV perception. Hongyang Li 0001, Chonghao Sima, Jifeng Dai, Wenhai Wang, Lewei Lu, Huijie Wang, Jiazhi Yang, Hanming Deng, Hao Tian 0006, Enze Xie, Jiangwei Xie, Li Chen 0008, Tianyu Li 0004, Yang Li 0189, Yulu Gao, Xiaosong Jia, Si Liu 0001, Jianping Shi, Dahua Lin, Yu Qiao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 14 |
| 2024 | Slim and Fast: Low-Overhead Container Overlay Network With Fast Connection SetupabstractLarge-scale cloud applications today are often deployed using multiple containers, and a container overlay network is the de facto method to provide connectivity among these containers. However, the existing tunneling-based overlay network incurs significant performance overhead due to the need of transformation for every packet. Recent workSlim, through manipulating connection-level meta-data, allows containers to use host OS sockets directly thus they can achieve good performance without extra packet tunneling. Nevertheless, the connection setup is significantly slowed down, which requires an extra round-trip communication between both sides to pass the mapping information of the host OS socket and the container socket. This greatly hurts the performance of many cloud applications that must process short connections at high speed. We proposeSlimFast, a low-overhead container overlay network which provides a fast connection setup.SlimFastdirectly uses the host OS socket for container communication asSlim. However,SlimFastneeds no extra communication during connection setup. We reserve a dedicated host port for the container network and use socket mapping table to locally find the right container socket during connection setup. We implementSlimFastwhich is compatible with existing container applications. Experiments show that,SlimFastcan improve the connection setup time by about 2.1x compared withSlim, meanwhile maintaining low-overhead during data transmission asSlim. This brings significant performance improvement to real applications. Particularly, testbed results show thatSlimFastimproves the throughput of Nginx proxy and Memcached by about 0.9x and 2.2x, respectively. Fusheng Lin, Xin Zhang 0117, Guo Chen 0001, Li Chen 0008, Kenli Li 0001, Hongbo Jiang 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2024 | Synchronize Only the Immature Parameters: Communication-Efficient Federated Learning By Freezing Parameters AdaptivelyabstractFederated learning allows edge devices to collaboratively train a global model without sharing their local private data. Yet, with limited network bandwidth at the edge, communication often becomes a severe bottleneck. In this paper, we find that it is unnecessary to always synchronize the full model in the entire training process, because many parameters already become mature (i.e., stable) prior to model convergence, and can thus be excluded from later synchronizations. This allows us to reduce the communication overhead without compromising the model accuracy. However, challenges are that the local parameters excluded from global synchronization may diverge on different clients, and meanwhile some parameters may stabilize only temporally. To address these challenges, we propose a novel scheme called Adaptive Parameter Freezing (APF), which fixes (freezes) the non-synchronized stable parameters in intermittent periods. Specifically, the freezing periods are tentatively adjusted in an additively-increase and multiplicatively-decrease manner—depending on whether the previously-frozen parameters remain stable in subsequent iterations. We also extend APF into APF# and APF++, which freeze parameters in a more aggressive manner to achieve larger performance benefit for large complex models. We implemented APF and its variants as Python modules with PyTorch, and extensive experiments show that APF can reduce data transfer amount by over 60%. Chen Chen 0067, Hong Xu 0001, Wei Wang 0030, Baochun Li, Bo Li 0001, Li Chen 0008, Gong Zhang 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2023 | Planning-oriented Autonomous DrivingabstractModern autonomous driving system is characterized as modular tasks in sequential order, i.e., perception, prediction, and planning. In order to perform a wide diversity of tasks and achieve advanced-level intelligence, contemporary approaches either deploy standalone models for individual tasks, or design a multi-task paradigm with separate heads. However, they might suffer from accumulative errors or deficient task coordination. Instead, we argue that a favorable framework should be devised and optimized in pursuit of the ultimate goal, i.e., planning of the self-driving car. Oriented at this, we revisit the key components within perception and prediction, and prioritize the tasks such that all these tasks contribute to planning. We introduce Unified Autonomous Driving (UniAD), a comprehensive framework up-to-date that incorporates full-stack driving tasks in one network. It is exquisitely devised to leverage advantages of each module, and provide complementary feature abstractions for agent interaction from a global perspective. Tasks are communicated with unified query interfaces to facilitate each other toward planning. We instantiate UniAD on the challenging nuScenes benchmark. With extensive ablations, the effectiveness of using such a philosophy is proven by substantially outperforming previous state-of-the-arts in all aspects. Code and models are public. Yihan Hu 0001, Jiazhi Yang, Li Chen 0008, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Wenhai Wang, Lewei Lu, Xiaosong Jia, Jifeng Dai, Yu Qiao 0001, Hongyang Li 0001 |
CVPR | 3 |
| 2023 | Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous DrivingabstractEnd-to-end autonomous driving has made impressive progress in recent years. Existing methods usually adopt the decoupled encoder-decoder paradigm, where the encoder extracts hidden features from raw sensor data, and the decoder outputs the ego-vehicle's future trajectories or actions. Under such a paradigm, the encoder does not have access to the intended behavior of the ego agent, leaving the burden of finding out safety-critical regions from the massive receptive field and inferring about future situations to the decoder. Even worse, the decoder is usually composed of several simple multi-layer perceptrons (MLP) or GRUs while the encoder is delicately designed (e.g., a combination of heavy ResNets or Transformer). Such an imbalanced resource-task division hampers the learning process. In this work, we aim to alleviate the aforementioned problem by two principles: (1) fully utilizing the capacity of the encoder; (2) increasing the capacity of the decoder. Concretely, we first predict a coarse-grained future position and action based on the encoder features. Then, conditioned on the position and action, the future scene is imagined to check the ramification if we drive accordingly. We also retrieve the encoder features around the predicted coordinate to obtain fine-grained information about the safety-critical region. Finally, based on the predicted future and the retrieved salient feature, we refine the coarse-grained position and action by predicting its offset from ground-truth. The above refinement module could be stacked in a cascaded fashion, which extends the capacity of the decoder with spatial-temporal prior knowledge about the conditioned future. We conduct experiments on the CARLA simulator and achieve state-of-the-art performance in closed-loop benchmarks. Extensive ablation studies demonstrate the effectiveness of each proposed module. Xiaosong Jia, Penghao Wu, Li Chen 0008, Jiangwei Xie, Conghui He, Junchi Yan, Hongyang Li 0001 |
CVPR | 3 |
| 2023 | Distilling Focal Knowledge from Imperfect Expert for 3D Object DetectionabstractMulti-camera 3D object detection blossoms in recent years and most of state-of-the-art methods are built up on the bird’ s-eye- view (BEV) representations. Albeit remarkable performance, these works suffer from low efficiency. Typically, knowledge distillation can be used for model compression. However, due to unclear 3D geometry reasoning, expert features usually contain some noisy and confusing areas. In this work, we investigate on how to distill the knowledge from an imperfect expert. We propose FD3D, a Focal Distiller for 3D object detection. Specifically, a set of queries are leveraged to locate the instance-level areas for masked feature generation, to intensify feature representation ability in these areas. Moreover, these queries search out the representative fine-grained positions for refined distillation. We verify the effectiveness of our method by applying it to two popular detection models, BEVFormer and DETR3D. The results demonstrate that our method achieves improvements of 4.07 and 3.17 points respectively in terms of NDS metric on nuScenes benchmark. Code is hosted at https://github.com/OpenPerceptionX/BEVPerception-Survey-Recipe. Li Chen 0008, Hanming Deng, Lewei Lu, Junchi Yan, Yu Qiao 0001, Hongyang Li 0001 |
CVPR | 2 |
| 2023 | DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous DrivingabstractEnd-to-end autonomous driving aims to build a fully differentiable system that takes raw sensor data as inputs and directly outputs the planned trajectory or control signals of the ego vehicle. State-of-the-art methods usually follow the ‘Teacher-Student’ paradigm. The Teacher model uses privileged information (ground-truth states of surrounding agents and map elements) to learn the driving strategy. The student model only has access to raw sensor data and conducts behavior cloning on the data collected by the teacher model. By eliminating the noise of the perception part during planning learning, state-of-the-art works could achieve better performance with significantly less data compared to those coupled ones.However, under the current Teacher-Student paradigm, the student model still needs to learn a planning head from scratch, which could be challenging due to the redundant and noisy nature of raw sensor inputs and the casual confusion issue of behavior cloning. In this work, we aim to explore the possibility of directly adopting the strong teacher model to conduct planning while letting the student model focus more on the perception part. We find that even equipped with a SOTA perception model, directly letting the student model learn the required inputs of the teacher model leads to poor driving performance, which comes from the large distribution gap between predicted privileged inputs and the ground-truth.To this end, we propose DriveAdapter, which employs adapters with the feature alignment objective function between the student (perception) and teacher (planning) modules. Additionally, since the pure learning-based teacher model itself is imperfect and occasionally breaks safety rules, we propose a method of action-guided feature learning with a mask for those imperfect teacher features to further inject the priors of hand-crafted rules into the learning process. DriveAdapter achieves SOTA performance on multiple closed-loop simulation-based benchmarks of CARLA. Xiaosong Jia, Yulu Gao, Li Chen 0008, Junchi Yan, Patrick Langechuan Liu, Hongyang Li 0001 |
ICCV | 3 |
| 2023 | Scene as OccupancyabstractHuman driver can easily describe the complex traffic scene by visual system. Such an ability of precise perception is essential for driver’s planning. To achieve this, a geometry-aware representation that quantizes the physical 3D scene into structured grid map with semantic labels per cell, termed as 3D Occupancy, would be desirable. Compared to the form of bounding box, a key insight behind occupancy is that it could capture the fine-grained details of critical obstacles in the scene, and thereby facilitate subsequent tasks. Prior or concurrent literature mainly concentrate on a single scene completion task, where we might argue that the potential of this occupancy representation might obsess broader impact. In this paper, we propose OccNet, a multi-view vision-centric pipeline with a cascade and temporal voxel decoder to reconstruct 3D occupancy. At the core of OccNet is a general occupancy embedding to represent 3D physical world. Such a descriptor could be applied towards a wide span of driving tasks, including detection, segmentation and planning. To validate the effectiveness of this new representation and our proposed algorithm, we propose OpenOcc, the first dense high-quality 3D occupancy benchmark built on top of nuScenes. Empirical experiments show that there are evident performance gain across multiple tasks, e.g., motion planning could witness a collision rate reduction by 15%-58%, demonstrating the superiority of our method. Wenwen Tong, Chonghao Sima, Li Chen 0008, Silei Wu, Hanming Deng, Yi Gu 0005, Lewei Lu, Ping Luo 0002, Dahua Lin, Hongyang Li 0001 |
ICCV | 4 |
| 2023 | Policy Pre-training for Autonomous Driving via Self-supervised Geometric Modeling
Penghao Wu, Li Chen 0008, Hongyang Li 0001, Xiaosong Jia, Junchi Yan, Yu Qiao 0001 |
ICLR | 2 |
| 2023 | MDP: Model Decomposition and Parallelization of Vision Transformer for Distributed Edge InferenceabstractDistributed edge inference emerges to be a promising paradigm to speed up inference. Previous works make physical partitions on CNNs to realize it, but there are the following challenges for vision transformers: (1) high communication costs for the large model; (2) stragglers because of heterogeneous devices; (3) time-out exceptions due to unstable edge devices.Therefore, we propose a novel Model Decomposition and Parallelization(MDP) for large vision transformers. Inspired by the implicit boosting ensemble in the vision transformer, MDP decomposes it into an explicit boosting ensemble of different and parallel sub-models. It sequentially trains all sub-models to gradually reduce the residual errors. To minimize dependency and communication among sub-models, We adopt stacking distillation to bring every sub-model extra information about others for better error correction. Different sub-models can take both different image sizes and model sizes to run on heterogeneous devices and improve the ensemble diversities. To handle the timeout exception, we add vanilla supervised learning on every submodel for the bagging ensemble in case of the early termination of boosting ensemble. As a result, all sub-models can not only run in parallel without much communication but also can be adapted to the heterogeneous devices, while maintaining accuracy even with time-out exceptions. Experiments show that MDP can outperform other baselines by $5 . 2 \times \sim 2 . 1 \times$ in latency and $5 . 1 \times \sim 1 . 7 \times$ in throughput with comparable accuracy. Weiyan Wang, Yiming Zhang 0003, Yilun Jin, Han Tian, Li Chen 0008 |
MSN | 5 |
| 2023 | OpenLane-V2: A Topology Reasoning Benchmark for Unified 3D HD MappingabstractAccurately depicting the complex traffic scene is a vital component for autonomous vehicles to execute correct judgments. However, existing benchmarks tend to oversimplify the scene by solely focusing on lane perception tasks. Observing that human drivers rely on both lanes and traffic signals to operate their vehicles safely, we present OpenLane-V2, the first dataset on topology reasoning for traffic scene structure. The objective of the presented dataset is to advance research in understanding the structure of road scenes by examining the relationship between perceived entities, such as traffic elements and lanes. Leveraging existing datasets, OpenLane-V2 consists of 2,000 annotated road scenes that describe traffic elements and their correlation to the lanes. It comprises three primary sub-tasks, including the 3D lane detection inherited from OpenLane, accompanied by corresponding metrics to evaluate the model’s performance. We evaluate various state-of-the-art methods, and present their quantitative and qualitative results on OpenLane-V2 to indicate future avenues for investigating topology reasoning in traffic scenes. Huijie Wang, Tianyu Li 0004, Yang Li 0189, Li Chen 0008, Chonghao Sima, Zhenbo Liu, Bangjun Wang, Peijin Jia, Shengyin Jiang, Hang Xu 0004, Ping Luo 0002, Junchi Yan, Wei Zhang 0196, Hongyang Li 0001 |
NeurIPS | 4 |
| 2023 | Demo: NetVision: Efficient Visualization Front-End for Packet-level Discrete-Event Network SimulationabstractVisualization of network simulation is an essential tool for network practitioners. However, the front-end of existing network simulators often fails to deliver satisfactory performance when dealing with modern network scales and interface speed. In this paper, we propose NetVision, an efficient visualization front-end of network simulation based on the Unity engine, which is commonly used for video game and virtual reality development. NetVision offers flow-level visualization of network behavior and performances. Then, through parallel optimization, NetVision supports real-time visualization for large-scale high-speed networks. Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Xizheng Wang, Lu Lu 0016 |
SIGCOMM | 2 |
| 2023 | DONS: Fast and Affordable Discrete Event Network Simulation with Automatic ParallelizationabstractDiscrete Event Simulation (DES) is an essential tool for network practitioners. Unfortunately, existing DES simulators cannot achieve satisfactory performance at the scale of modern networks. Recent work has attempted to address these challenges by reducing the traffic processed via novel approximation techniques; however, we argue in this paper that much of the slowdown of existing DES simulators is due to their underlying software architecture. Kaihui Gao, Li Chen 0008, Dan Li 0001, Vincent Liu 0001, Xizheng Wang, Lu Lu 0016 |
SIGCOMM | 2 |
| 2023 | GIFT: Toward Accurate and Efficient Federated Learning With Gradient-Instructed Frequency TuningabstractFederated learning (FL) enables distributed clients to collectively train a global model without revealing their private data, and for efficiency clients synchronize their gradients periodically. However, this can lead to the inaccuracy in model convergence due to inconsistent data distributions among clients. In this work, we find that there is a strong correlation between FL accuracy loss and the synchronization frequency, and seek to fine tune the synchronization frequency at training runtime to make FL accurate and also efficient. Specifically, aware that under the FL privacy requirement only gradients can be utilized for making frequency tuning decisions, we propose a novel metric called gradient consistency, which can effectively reflect the training status despite the instability of realistic FL scenarios. We further devise a feedback-driven algorithm called Gradient-Instructed Frequency Tuning (GIFT), which adaptively increases or decreases the synchronization frequency based on the gradient consistency metric. We have implemented GIFT in PyTorch, and large-scale evaluations show that it can improve FL accuracy by up to 10.7% with a time reduction of 58.1%. Chen Chen 0067, Hong Xu 0001, Wei Wang 0030, Baochun Li, Bo Li 0001, Li Chen 0008, Gong Zhang 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2023 | HDGT: Heterogeneous Driving Graph Transformer for Multi-Agent Trajectory Prediction via Scene EncodingabstractEncoding a driving scene into vector representations has been an essential task for autonomous driving that can benefit downstream tasks e.g., trajectory prediction. The driving scene often involves heterogeneous elements such as the different types of objects (agents, lanes, traffic signs) and the semantic relations between objects are rich and diverse. Meanwhile, there also exist relativity across elements, which means that the spatial relation is a relative concept and need be encoded in a ego-centric manner instead of in a global coordinate system. Based on these observations, we propose Heterogeneous Driving Graph Transformer (HDGT), a backbone modelling the driving scene as a heterogeneous graph with different types of nodes and edges. For heterogeneous graph construction, we connect different types of nodes according to diverse semantic relations. For spatial relation encoding, the coordinates of the node as well as its in-edges are in the local node-centric coordinate system. For the aggregation module in the graph neural network (GNN), we adopt the transformer structure in a hierarchical way to fit the heterogeneous nature of inputs. Experimental results show that HDGT achieves state-of-the-art performance for the task of trajectory prediction, on INTERACTION Prediction Challenge and Waymo Open Motion Challenge. Xiaosong Jia, Penghao Wu, Li Chen 0008, Yu Liu 0015, Hongyang Li 0001, Junchi Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Cross-Graph Embedding With Trainable Proximity for Graph AlignmentabstractGraph alignment, also known as network alignment, has many applications in data mining tasks. It aims to find the node correspondence across disjoint graphs. With recent representation learning advancements, embedding-based graph alignment has become a hot topic. Existing embedding-based methods focus either on structural proximity across graphs or on the positional proximity within a single graph. However, only considering the structural similarity will make the position relation of nodes not clear enough, which makes it easy to misalign the nodes close in distance, while only considering the position proximity of a single graph will make the node embeddings from different graphs in different subspaces. To mitigate this issue, we propose a novel model CEGA forCross-graphEmbedding-basedGraphAlignment, which can generate node embeddings to reflect structural proximity and positional proximity simultaneously. Meanwhile, we make the proximity trainable thus it can be learned to best suit the alignment task at hand automatically. We show that CEGA outperforms existing graph alignment methods in accuracy under unsupervised scenarios through extensive experiments on public benchmarks. Wei Tang 0013, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Huangxun Chen, Li Chen 0008 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Fast, Scalable and Robust Centralized Routing for Data Center NetworksabstractThis paper presents a fast and robust centralized data center network (DCN) routing solution, called . For fast routing calculation, uses centralized controllers to collect/disseminate the network’s link-states (LS), and offload the actual routing calculation onto each switch. Observing that the routing changes can be classified into a few fixed patterns in DCNs which have regular topologies, we simplify each switch’s routing calculation into a table-lookup manner, i.e., comparing LS changes with pre-installed base topology and updating routing paths according to predefined rules. As such, the routing calculation time at each switch only needs 10s of us even in a large network topology containing 10K+ switches. For efficient controller fault-tolerance, purposely uses reporter switch to ensure the LS updates successfully delivered to all affected switches. As such, can use multiple stateless controllers and little redundant traffic to tolerate failures, which incurs little overhead under normal case, and keeps 10s of ms fast routing reaction time even under complex data-/control-plane failures. We design, implement and evaluate with extensive experiments on Linux-machine controllers and white-box switches. provides$\sim$1200x and$\sim$100x shorter convergence time than current distributed protocol BGP and the state-of-the-art centralized routing solution, respectively. Furthermore, Primus maintains good routing controllability/manageability thanks to its centralized architecture, which enables us to build several advanced routing features in our testbed, including routing failure visualization and weighted-cost-multi-path routing. Fusheng Lin, Guo Chen 0001, Guihua Zhou, Dehui Wei, Li Chen 0008, Yuanwei Lu, Andrew Qu, Hongbo Jiang 0001 |
IEEE/ACM Trans. Netw. | 7 |
| 2022 | NASPipe: high performance and reproducible pipeline parallel supernet training via causal synchronous parallelismabstractSupernet training, a prevalent and important paradigm in Neural Architecture Search, embeds the whole DNN architecture search space into one monolithic supernet, iteratively activates a subset of the supernet (i.e., a subnet) for fitting each batch of data, and searches a high-quality subnet which meets specific requirements. Although training subnets in parallel on multiple GPUs is desirable for acceleration, there inherently exists a race hazard that concurrent subnets may access the same DNN layers. Existing systems support neither efficiently parallelizing subnets’ training executions, nor resolving the race hazard deterministically, leading to unreproducible training procedures and potentiallly non-trivial accuracy loss. Shixiong Zhao, Fanxin Li, Xusheng Chen, Tianxiang Shen, Li Chen 0008, Sen Wang 0004, Nicholas Zhang, Cheng Li 0001, Heming Cui |
ASPLOS | 5 |
| 2022 | PersFormer: 3D Lane Detection via Perspective Transformer and the OpenLane Benchmark
Li Chen 0008, Chonghao Sima, Yang Li 0189, Zehan Zheng, Jiajie Xu 0001, Xiangwei Geng, Hongyang Li 0001, Conghui He, Jianping Shi, Yu Qiao 0001, Junchi Yan |
ECCV (38) | 1 |
| 2022 | ST-P3: End-to-End Vision-Based Autonomous Driving via Spatial-Temporal Feature Learning
Shengchao Hu, Li Chen 0008, Penghao Wu, Hongyang Li 0001, Junchi Yan, Dacheng Tao |
ECCV (38) | 2 |
| 2022 | A Focused Garbage Collection Approach for Primary Deduplicated Storage with Low Memory OverheadabstractSince one chunk could be shared by many files after data deduplication, Garbage Collection (GC) is an essential but complex task to reclaim stale chunks in large-scale primary deduplication systems. Traditional Mark&Sweep is a widely used approach but suffers from the increasingly traversing time and huge memory overhead of Liveness Array (i.e., a data structure reflects the liveness of alive chunks) in the Mark phase. This paper proposes a new method named Focused Garbage Collection (FGC) to accelerate the Mark phase for primary deduplication storage significantly. Specifically, we design a global Austere Reference Graph with low memory cost that efficiently represents files’ reference relationships (i.e., sharing chunks after deduplication) by considering the deduplication characteristics of workloads in primary systems. Austere Reference Graph helps FGC focus on the deleted files and their correlative files to quickly mark stale chunks, while traditional approaches need to traverse all files. Consequently, FGC’s traversing time and Liveness Array size will be greatly reduced in the Mark phase. Evaluation results show that compared with traditional Mark&Sweep, FGC decreases the time consumption in the Mark phase 1.3×-7.34× in a stand-alone primary deduplication system and 128×-256× network traffic reduction for the Mark phase while only introducing < 0.05% extra memory overhead for the reference graph. Jingsong Yuan, Xiangyu Zou, Zhichao Cao 0002, Wen Xia, Peng Wang 0037, Li Chen 0008 |
ICCD | 8 |
| 2022 | XTree: Traversal-Based Partitioning for Extreme-Scale Graph Processing on SupercomputersabstractGraph algorithms, such as Breadth First Search (BFS), Single Source Shortest Path (SSSP), PageRank (PR), and Connected Components (CC), are increasingly important in big data processing and analytics. As graph scales (numbers of vertices and edges) have increased from billions to trillions, Supercomputers have huge numbers (up to hundreds of thousands) of computing nodes (CNs) that can provide ultra-high aggregate computing power and memory capacity, thus being particularly suitable for processing extreme-scale graphs with trillions of vertices and edges. However, existing cluster-based graph-parallel systems perform poorly when deployed on supercomputers, since their partitioning methods overlook the hierarchical nature of supercomputer networks and incur prohibitive communication storm. This paper presents XTree, an efficient traversal-based partitioning method for minimizing communication overhead of graph processing on supercomputers. We observe that supercomputers' huge numbers of CNs are usually organized into hierarchical communication domains, which can be modeled as a domain tree where communication in lower-level domains is significantly faster than that in higher-level ones. Therefore, the key idea of XTree's partitioning is to exploit hierarchical locality by viewing the graph as a BFS tree and leveraging the topology knowledge to map the graph's BFS tree onto the domain tree, We evaluate the effectiveness of XTree by running various graph algorithms, on both real-world big graphs and synthetic trillion-scale graphs. XTree substantially reduces communication overhead and achieves orders of magnitude speedup against the Graph500 reference implementations with the state-of-the-art 2D-decomposition partitioning. Xinbiao Gan, Yiming Zhang 0003, Ruigeng Zeng, Jie Liu 0002, Ruibo Wang, Li Chen 0008, Kai Lu 0001 |
ICDE | 7 |
| 2022 | Cutting Tail Latency in Commodity Datacenters with CloudburstabstractLong tail latency of short flows (or messages) greatly affects user-facing applications in datacenters. Prior solutions to the problem introduce significant implementation complexities, such as global state monitoring, complex network control, or non-trivial switch modifications. While promising superior performance, they are hard to implement in practice.This paper presents Cloudburst, a simple, effective yet readily deployable solution achieving similar or even better results without introducing the above complexities. At its core, Cloudburst explores forward error correction (FEC) over multipath — it proactively spreads FEC-coded packets generated from messages over multipath in parallel, and recovers them with the first few arriving ones. As a result, Cloudburst is able to obliviously exploit underutilized paths, thus achieving low tail latency. We have implemented Cloudburst as a user-space library, and deployed it on a testbed with commodity switches. Our testbed and simulation experiments show the superior performance of Cloudburst. For example, Cloudburst achieves 63.69% and 60.06% reduction in 99th percentile message/flow completion time (FCT) compared to DCTCP and PIAS, respectively. Gaoxiong Zeng, Li Chen 0008, Bairen Yi, Kai Chen 0005 |
INFOCOM | 2 |
| 2022 | CRONUS: Fault-isolated, Secure and High-performance Heterogeneous Computing for Trusted Execution EnvironmentabstractWith the trend of processing a large volume of sensitive data on PaaS services (e.g., DNN training), a TEE architecture that supports general heterogeneous accelerators, enables spatial sharing on one accelerator, and enforces strong isolation across accelerators is highly desirable. However, none of the existing TEE solutions meet all three requirements. In this paper, we propose CRONUS, the first TEE architecture that achieves the three crucial requirements. The key idea of CRONUS is to partition heterogeneous computation into isolated TEE enclaves, where each enclave encapsulates only one kind of computation (e.g., GPU computation), and multiple enclaves can spatially share an accelerator. Then, CRONUS constructs heterogeneous computing using remote procedure calls (RPCs) among enclaves. With CRONUS, each accelerator’s hardware and its software stack are strongly isolated from others’, and each enclave trusts only its own hardware. To tackle the security challenge caused by inter-enclave interactions, we design a new streaming remote procedure call abstraction to enable secure RPCs with high performance. CRONUS is software-based, making it general to diverse accelerators. We implemented CRONUS on ARM TrustZone. Evaluation on diverse workloads with CPUs, GPUs and NPUs shows that, CRONUS achieves less than 7.1% extra computation time compared to native (unprotected) executions. Jianyu Jiang, Ji Qi 0002, Tianxiang Shen, Xusheng Chen, Shixiong Zhao, Sen Wang 0004, Li Chen 0008, Gong Zhang 0001, Xiapu Luo, Heming Cui |
MICRO | 7 |
| 2022 | Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong BaselineabstractCurrent end-to-end autonomous driving methods either run a controller based on a planned trajectory or perform control prediction directly, which have spanned two separately studied lines of research. Seeing their potential mutual benefits to each other, this paper takes the initiative to explore the combination of these two well-developed worlds. Specifically, our integrated approach has two branches for trajectory planning and direct control, respectively. The trajectory branch predicts the future trajectory, while the control branch involves a novel multi-step prediction scheme such that the relationship between current actions and future states can be reasoned. The two branches are connected so that the control branch receives corresponding guidance from the trajectory branch at each time step. The outputs from two branches are then fused to achieve complementary advantages. Our results are evaluated in the closed-loop urban driving setting with challenging scenarios using the CARLA simulator. Even with a monocular camera input, the proposed approach ranks first on the official CARLA Leaderboard, outperforming other complex candidates with multiple sensors or fusion mechanisms by a large margin. The sourcecode is publicly available at https://github.com/OpenPerceptionX/TCP Penghao Wu, Xiaosong Jia, Li Chen 0008, Junchi Yan, Hongyang Li 0001, Yu Qiao 0001 |
NeurIPS | 3 |
| 2022 | Software-defined network assimilation: bridging the last mile towards centralized network configuration management with NAssimabstractOn-boarding new devices into an existing SDN network is a pain for network operations (NetOps) teams, because much expert effort is required to bridge the gap between the configuration models of the new devices and the unified data model in the SDN controller. In this work, we present an assistant framework NAssim, to help NetOps accelerate the process of assimilating a new device into a SDN network. Our solution features a unified parser framework to parse diverse device user manuals into preliminary configuration models, a rigorous validator that confirm the correctness of the models via formal syntax analysis, model hierarchy validation and empirical data validation, and a deep-learning-based mapping algorithm that uses state-of-the-art neural language processing techniques to produce human-comprehensible recommended mapping between the validated configuration model and the one in the SDN controller. In all, NAssim liberates the NetOps from most tedious tasks by learning directly from devices' manuals to produce data models which are comprehensible by both the SDN controller and human experts. Our evaluation shows, NAssim can accelerate the assimilation process by 9.1x. In this process, we also identify and correct 243 errors in four mainstream vendors' device manuals, and release a validated and expert-curated dataset of parsed manual corpus for future research. Huangxun Chen, Yukai Miao, Li Chen 0008, Haifeng Sun 0001, Hong Xu 0001, Libin Liu 0001, Gong Zhang 0001, Wei Wang 0011 |
SIGCOMM | 3 |
| 2022 | DeepQueueNet: towards scalable and generalized network performance estimation with packet-level visibilityabstractNetwork simulators are an essential tool for network operators, and can assist important tasks such as capacity planning, topology design, and parameter tuning. Popular simulators are all based on discrete event simulation, and their performance does not scale with the size of modern networks. Recently, deep-learning-based techniques are introduced to solve the scalability problem, but, as we show with experiments, they have poor visibility in their simulation results, and cannot generalize to diverse scenarios. In this work, we combine scalable and generalized continuous simulation techniques with discrete event simulation to achieve high scalability, while providing packet-level visibility. We start from a solid queueing-theoretic modeling of modern networks, and carefully identify the mathematically-intractable or computationally-expensive parts, only which are then modeled using deep neural networks (DNN). Dubbed DeepQueueNet, our approach combines prior knowledge of networks, and supports arbitrary topology and device traffic management mechanisms (given sufficient training data). Our extensive experiments show that DeepQueueNet achieves near-linear speedup in the number of GPUs, and its estimation accuracy for average and 99th percentile round-trip time outperforms existing end-to-end DNN-based performance estimators in all scenarios. Xi Peng 0006, Li Chen 0008, Libin Liu 0001, Jingze Zhang, Hong Xu 0001, Baochun Li, Gong Zhang 0001 |
SIGCOMM | 3 |
| 2022 | SOTER: Guarding Black-box Inference for General Neural Networks at the Edge
Tianxiang Shen, Ji Qi 0002, Jianyu Jiang, Siyuan Wen, Xusheng Chen, Shixiong Zhao, Sen Wang 0004, Li Chen 0008, Xiapu Luo, Fengwei Zhang, Heming Cui |
USENIX ATC | 9 |
| 2021 | A MAP-based Performance Analysis on 5G-powered Cloud VR StreamingabstractDespite that cloud virtual reality (VR) is the most promising service in 5G networks, a reasonable traffic model for it is still unknown. Based on statistics of real cloud VR traces, we justify that the Markovian arrival process (MAP) can well characterize the inter-arrival times (IATs) among packets. Moreover, we discuss possible methods to efficiently estimate parameters for flow aggregation. Since MAPs are analytically tractable, we quantify the performance of delivering cloud VR traffic from a queueing theory perspective. Aiming at supporting quantile metrics, we provide distributions of queue-length and latency for an arbitrary packet. The accuracy of estimated latency is validated by comparing with measured latency on an industrial 5G platform. The MAP-based traffic model will potentially serve as an input for performance evaluation and network planning for assuring high requirements of user experience. Xi Peng 0006, Fan Zhang 0016, Li Chen 0008, Gong Zhang 0001 |
ICC | 3 |
| 2021 | Communication-Efficient Federated Learning with Adaptive Parameter FreezingabstractFederated learning allows edge devices to collaboratively train a global model by synchronizing their local updates without sharing private data. Yet, with limited network bandwidth at the edge, communication often becomes a severe bottleneck. In this paper, we find that it is unnecessary to always synchronize the full model in the entire training process, because many parameters gradually stabilize prior to the ultimate model convergence, and can thus be excluded from being synchronized at an early stage. This allows us to reduce the communication overhead without compromising the model accuracy. However, challenges are that the local parameters excluded from global synchronization may diverge on different clients, and meanwhile some parameters may stabilize only temporally. To address these challenges, we propose a novel scheme called Adaptive Parameter Freezing (APF), which fixes (freezes) the non-synchronized stable parameters in intermittent periods. Specifically, the freezing periods are tentatively adjusted in an additively-increase and multiplicatively-decrease manner, depending on if the previously-frozen parameters remain stable in subsequent iterations. We implemented APF as a Python module in PyTorch. Our extensive array of experimental results show that APF can reduce data transfer by over 60%. Chen Chen 0067, Hong Xu 0001, Wei Wang 0030, Baochun Li, Bo Li 0001, Li Chen 0008, Gong Zhang 0001 |
ICDCS | 6 |
| 2021 | Primus: Fast and Robust Centralized Routing for Large-scale Data Center NetworksabstractThis paper presents a fast and robust centralized data center network (DCN) routing solution called Primus. For fast routing calculation, Primus uses centralized controller to collect/disseminates the network's link-states (LS), and offload the actual routing calculation onto each switch. Observing that the routing changes can be classified into a few fixed patterns in DCNs which have regular topologies, we simplify each switch's routing calculation into a table-lookup manner, i.e., comparing LS changes with pre-installed base topology and updating routing paths according to predefined rules. As such, the routing calculation time at each switch only needs 10s of us even in a large network topology containing 10K+ switches. For efficient controller fault-tolerance, Primus purposely uses reporter switch to ensure the LS updates successfully delivered to all affected switches. As such, Primus can use multiple stateless controllers and little redundant traffic to tolerate failures, which incurs little overhead under normal case, and keeps 10s of ms fast routing reaction time even under complex data-/control-plane failures. We design, implement and evaluate Primus with extensive experiments on Linux-machine controllers and white-box switches. Primus provides ~1200x and ~100x shorter convergence time than current distributed protocol BGP and the state-of-the-art centralized routing solution, respectively. Guihua Zhou, Guo Chen 0001, Fusheng Lin, Dehui Wei, Jianbing Wu, Li Chen 0008, Yuanwei Lu, Andrew Qu, Hongbo Jiang 0001 |
INFOCOM | 7 |
| 2021 | Cluster-Reduce: Compressing Sketches for Distributed Data StreamsabstractSketches, a type of probabilistic algorithms, have been widely accepted as the approximate summary of data streams. Compressing sketches is the best choice in distributed data streams to reduce communication overhead. The ideal compression algorithm should meet the following three requirements: high efficiency of compression procedure, support of direct query without decompression, and high accuracy of compressed sketches. However, no prior work can meet these requirements at the same time. Especially, the accuracy is poor after compression using existing methods. In this paper, we propose Cluster-Reduce, a framework for compressing sketches, which can meet all three requirements. Our key technique nearness clustering rearranges the adjacent counters with similar values in the sketch to significantly improve the accuracy. We use Cluster-Reduce to compress four kinds of sketches in two use-cases: distributed data streams and distributed machine learning. Extensive experimental results show that Cluster-Reduce can achieve up to 60 times smaller error than prior works. The source codes of Cluster-Reduce are available at Github anonymously[1]. Yikai Zhao 0001, Yuanpeng Li 0002, Yifan Zhu 0011, Li Chen 0008, Yi Wang 0004, Tong Yang 0003 |
KDD | 6 |
| 2021 | LightGuardian: A Full-Visibility, Lightweight, In-band Telemetry System Using Sketchlets
Yikai Zhao 0001, Kaicheng Yang 0001, Zirui Liu 0002, Tong Yang 0003, Li Chen 0008, Naiqian Zheng, Hanbo Wu, Yi Wang 0004, Nicholas Zhang |
NSDI | 5 |
| 2021 | Bidl: A High-throughput, Low-latency Permissioned Blockchain Framework for Datacenter NetworksabstractA permissioned blockchain framework typically runs an efficient Byzantine consensus protocol and is attractive to deploy fast trading applications among a large number of mutually untrusted participants (e.g., companies). Unfortunately, all existing permissioned blockchain frameworks adopt sequential workflows for invoking the consensus protocol and executing applications' transactions, making the performance of these applications much lower than deploying them in traditional systems (e.g., in-datacenter stock exchange). Ji Qi 0002, Xusheng Chen, Yunpeng Jiang, Jianyu Jiang, Tianxiang Shen, Shixiong Zhao, Sen Wang 0004, Gong Zhang 0001, Li Chen 0008, Man Ho Au, Heming Cui |
SOSP | 9 |
| 2020 | Incorporating Intra-flow Dependencies and Inter-flow Correlations for Traffic Matrix PredictionabstractTraffic matrix (TM) prediction is essential for effective traffic engineering and network management. Based on our analysis of real traffic traces from Wide Area Network, the traffic flows in TM are both time-varying (i.e. with intra-flow dependencies) and correlated with each other (i.e. with inter-flow correlations). However, most existing works in TM prediction ignore inter-flow correlations. In this paper, we propose a novel Attention-based Convolutional Recurrent Neural Network (ACRNN) model to capture both intra-flow dependencies and inter-flow correlations. ACRNN mainly contains two components: 1) Correlational Modeling employs attention-based convolutional structures to capture the correlation of any two flows in TMs; 2) Temporal Modeling uses attention-based recurrent structures to model the long-term temporal dependencies of each flow, and then predicts TMs according inter-flow correlations and intra-flow dependencies. Experiments on two real-world datasets show that, when predicting the next TM, ACRNN model reduces the Mean Squared Error by up to 44.8% and reduces the Mean Absolute Error by up to 30.6%, compared to state-of-the-art method; and the gap is even larger when predicting the next multiple TMs. Besides, simulation results demonstrate that ACRNN's accurate prediction can help traffic engineering to mitigate traffic congestion. Kaihui Gao, Dan Li 0001, Li Chen 0008, Jinkun Geng, Fei Gui |
IWQoS | 3 |
| 2018 | PowerMan: An Out-of-Band Management Network for Datacenters Using Power Line Communication
Li Chen 0008, Jiacheng Xia, Bairen Yi, Kai Chen 0005 |
NSDI | 1 |
| 2018 | AuTO: scaling deep reinforcement learning for datacenter-scale automatic traffic optimizationabstractTraffic optimizations (TO, e.g. flow scheduling, load balancing) in datacenters are difficult online decision-making problems. Previously, they are done with heuristics relying on operators' understanding of the workload and environment. Designing and implementing proper TO algorithms thus take at least weeks. Encouraged by recent successes in applying deep reinforcement learning (DRL) techniques to solve complex online control problems, we study if DRL can be used for automatic TO without human-intervention. However, our experiments show that the latency of current DRL systems cannot handle flow-level TO at the scale of current datacenters, because short flows (which constitute the majority of traffic) are usually gone before decisions can be made. Li Chen 0008, Justinas Lingys, Kai Chen 0005 |
SIGCOMM | 1 |
| 2017 | Enabling Wide-Spread Communications on Optical Fabric with MegaSwitch
Li Chen 0008, Kai Chen 0005, Zhonghua Zhu, Minlan Yu, George Porter, Chunming Qiao |
NSDI | 1 |
| 2017 | PIAS: Practical Information-Agnostic Flow Scheduling for Commodity Data CentersabstractMany existing data center network (DCN) flow scheduling schemes, that minimize flow completion times (FCT) assume prior knowledge of flows and custom switch functions, making them superior in performance but hard to implement in practice. By contrast, we seek to minimize FCT with no prior knowledge and existing commodity switch hardware. To this end, we present PIAS, a DCN flow scheduling mechanism that aims to minimize FCT by mimicking shortest job first (SJF) on the premise that flow size is not knowna priori. At its heart, PIAS leverages multiple priority queues available in existing commodity switches to implement a multiple level feedback queue, in which a PIAS flow is gradually demoted from higher-priority queues to lower-priority queues based on the number of bytes it has sent. As a result, short flows are likely to be finished in the first few high-priority queues and thus be prioritized over long flows in general, which enables PIAS to emulate SJF without knowing flow sizes beforehand. We have implemented a PIAS prototype and evaluated PIAS through both testbed experiments and ns-2 simulations. We show that PIAS is readily deployable with commodity switches and backward compatible with legacy TCP/IP stacks. Our evaluation results show that PIAS significantly outperforms existing information-agnostic schemes, for example, it reduces FCT by up to 50% compared to DCTCP[11]and L2DCT[32]; and it only has a 1.1% performance gap to an ideal information-aware scheme, pFabric[13], for short flows under a production DCN workload. Wei Bai 0001, Li Chen 0008, Kai Chen 0005, Dongsu Han, Chen Tian 0001, Hao Wang 0022 |
IEEE/ACM Trans. Netw. | 2 |
| 2016 | Enabling ECN over Generic Packet SchedulingabstractExplicit Congestion Notification (ECN) is crucial for production datacenters, but current queue-length based ECN/RED implementation does not work with generic packet schedulers, leading to either degraded network performance or violated scheduling policies. In this paper, we first dive into this issue and reveal that the invalidity of ECN/RED lies in the difficulty of measuring changing queue capacities under various schedulers and traffic dynamics. Then we present Time-based Congestion Notification (TCN), a simple yet effective ECN solution, by combining two successful ideas: the sojourn time from CoDel and the instantaneous marking from DCTCP. Using packet sojourn-time, as opposed to queue-length, as the congestion signal, TCN eliminates the need of measuring dynamic queue capacities, making it suitable for arbitrary schedulers with traffic dynamics. By performing stateless instantaneous ECN marking rather than complex stateful dropping, TCN is designed to be inexpensive to implement on commodity switching chips. Through extensive testbed experiments and large-scale simulations, we show TCN can strictly preserve scheduling policies while providing desirable network performance. For example, TCN significantly reduces the average and 99th percentile completion times for small flows by up to 82.8% and 95.3% compared to current practice in a testbed experiment with production workload. Wei Bai 0001, Kai Chen 0005, Li Chen 0008, Changhoon Kim |
CoNEXT | 3 |
| 2016 | Online flow size prediction for improved network routingabstractWe describe an emerging application of data mining in the context of computer networks. This application concerns the problem of predicting the size of a flow and detecting elephant flows (very large flows). Flow size is a very important statistic that can be used to improve routing, load balancing and scheduling in computer networks. Flow size prediction is particularly challenging since flow patterns continuously change and predictions must be done in real time (milliseconds) to avoid delays. We describe how to formulate the problem as an online machine learning task to continuously adjust to changes in flow traffic. We evaluate the predictive nature of a set of features and the accuracy of three online predictors based on neural networks, Gaussian process regression and online Bayesian Moment Matching on three datasets of real traffic. We also demonstrate how to use such online predictors to improve routing (i.e., reduced flow completion time) in a network simulation. Pascal Poupart, Zhitang Chen, Priyank Jaini, Fred Fung, Hengky Susanto, Yanhui Geng, Li Chen 0008, Kai Chen 0005 |
ICNP | 7 |
| 2016 | Enabling ECN in Multi-Service Multi-Queue Data Centers
Wei Bai 0001, Li Chen 0008, Kai Chen 0005 |
NSDI | 2 |
| 2016 | Scheduling Mix-flows in Commodity Datacenters with KarunaabstractCloud applications generate a mix of flows with and without deadlines. Scheduling such mix-flows is a key challenge; our experiments show that trivially combining existing schemes for deadline/non-deadline flows is problematic. For example, prioritizing deadline flows hurts flow completion time (FCT) for non-deadline flows, with minor improvement for deadline miss rate. Li Chen 0008, Kai Chen 0005, Wei Bai 0001, Mohammad Alizadeh |
SIGCOMM | 1 |
| 2016 | CODA: Toward Automatically Identifying and Scheduling Coflows in the DarkabstractLeveraging application-level requirements using coflows has recently been shown to improve application-level communication performance in data-parallel clusters. However, existing coflow-based solutions rely on modifying applications to extract coflows, making them inapplicable to many practical scenarios. Hong Zhang 0025, Li Chen 0008, Bairen Yi, Kai Chen 0005, Mosharaf Chowdhury, Yanhui Geng |
SIGCOMM | 2 |
| 2015 | FLOWPROPHET: Generic and Accurate Traffic Prediction for Data-Parallel Cluster ComputingabstractData-parallel computing frameworks (DCF) such as MapReduce, Spark, and Dryad etc. Have tremendous applications in big data and cloud computing, and throw tons of flows into data center networks. In this paper, we design and implement FLOW PROPHET, a general framework to predict traffic flows for DCFs. To this end, we analyze and summarize the common features of popular DCFs, and gain a key insight: since application logic in DCFs is naturally expressed by directed acyclic graphs (DAG), DAG contains necessary time and data dependencies for accurate flow prediction. Based on the insight, FLOW PROPHET extracts DAGs from user applications, and uses the time and data dependencies to calculate flow information 4-tuple, (source, destination, flow size, establish time), ahead-of-time for all flows. We also provide generic programming interface to FLOW PROPHET, so that current and future DCFs can deploy FLOW PROPHET readily. We implement FLOW PROPHET on both Spark and Hadoop, and perform extensive evaluations on a testbed with 37 physical servers. Our implementation and experiments demonstrate that, with time in advance and minimal cost, FLOW PROPHET can achieve almost 100% accuracy in source, destination, and flow size predictions. With accurate prediction from FLOW PROPHET, the job completion time of a Hadoop TeraSort benchmark is reduced by 12.52% on our cluster with a simple network scheduler. Hao Wang 0022, Li Chen 0008, Kai Chen 0005, Ziyang Li 0003, Yiming Zhang 0003, Haibing Guan, Zhengwei Qi, Dongsheng Li 0001, Yanhui Geng |
ICDCS | 2 |
| 2015 | Super-Low Frequency electromagnetic noise processing system based on adaptive filteringabstractSuper Low Frequency electromagnetic prospecting methods, based on natural source, have seen an increasingly trend in geophysical applications. It is known that natural source electromagnetic signal is weak, and how to extract useful information has drew great attention wordwidely. In this paper, we designed a signal processing method, an adaptive filter, to filter out the strong power frequency interference at 50Hz and its harmonics mixed in the output signal from induction magnetic sensors. As the output could change with the input, the adaptive filter system is releated to its input signal closely and specificly. Because of its stronger adaptability and better filtering performance, the adaptive filter showed good effectiveness. Nan Wang 0006, Li Chen 0008, Jian Hui, Chengye Zhang 0001, Qiming Qin |
IGARSS | 3 |
| 2015 | Information-Agnostic Flow Scheduling for Commodity Data Centers
Wei Bai 0001, Kai Chen 0005, Hao Wang 0022, Li Chen 0008, Dongsu Han, Chen Tian 0001 |
NSDI | 4 |
| 2014 | PIAS: Practical Information-Agnostic Flow Scheduling for Data Center NetworksabstractMany existing data center network (DCN) flow scheduling schemes minimize flow completion times (FCT) based on prior knowledge of flows and custom switch designs, making them hard to use in practice. This paper introduces, Pias, a practical flow scheduling approach that minimizes FCT with no prior knowledge using commodity switches. At its heart, Pias leverages multiple priority queues available in commodity switches to implement a Multiple Level Feedback Queue (MLFQ), in which a PIAS flow gradually demotes from higher-priority queues to lower-priority queues based on the bytes it has sent. In this way, short flows are prioritized over long flows, which enables Pias to emulate Shortest Job First (SJF) scheduling without knowing the flow sizes beforehand. Our preliminary evaluation shows that Pias significantly outperforms all existing information-agnostic solutions. It improves average FCT for short flows by up to 50% and 40% over DCTCP [3] and L2DCT [16]. Compared to an ideal information-aware DCN transport, p-Fabric [5], it only shows 4.9% performance degradation for short flows in a production datacenter workload. Wei Bai 0001, Li Chen 0008, Kai Chen 0005, Dongsu Han, Chen Tian 0001, Weicheng Sun |
HotNets | 2 |
| 2014 | Passive super-low frequency remote sensing technique for monitoring coal-bed methane reservoirsabstractCoal-bed methane (CBM), as an increasingly promising resource for the energy supply, deserves further exploration and accurate reservoir evaluation. It is also required to dynamically monitor the reservoirs (>200 m). Remote sensing methods in regular wavebands may fail in the depth sounding, with only imaging geo-objects shallower than 100 m. In contrast, the Super-Low Frequency (SLF) remote sensing technique has outstanding traits over others, including lower attenuation, all-weather and deeper penetration. In this paper, we have developed a non-imaging remote sensor to acquire electromagnetic signals in the Super-Low Frequency bands (i.e. SLF signals), which also enables us to fast and efficiently pre-process signals in a real-time display. In order to accurately identify producing CBM reservoirs, we mainly extract electromagnetic radiation (EMR) anomalies from processed SLF signals, and then dynamic analysis can be achieved. This technique has been validated by field experiments in Qin shui Basin, China. Nan Wang 0006, Qiming Qin, Li Chen 0008, Yanbing Bai, Chengye Zhang 0001, Huazhong Ren |
IGARSS | 3 |
| 2014 | Façade reconstruction from oblique areal imagesabstractThe paper realizes façade 3D reconstruction using recently promising oblique photogrammetry data. the point-to-point problems in traffic network. We make full use of the multi-level image features to extract interest regions of façade, and then present a backwards coarse-to-fine matching, which makes the auxiliary data unnecessary. The experiment shows the efficiency and robustness of the proposed method, and the vectors describing façade 3D information are also verify the high precision. Xiucheng Yang, Qiming Qin, Xuebin Qin, Jun Wang 0042, Yanbing Bai, Li Chen 0008 |
IGARSS | 7 |
| 2014 | Hyperspectral remote sensing for coal-bed methane explorationabstractBased on the theory of coal-bed methane(CBM) geology, the micro-seeps of hydrocarbon cause geochemical alterations in rocks and soil. In this study, the hyperspectral instrument, Hyperion, was used to detect the alterations and hydrocarbons on the land surface of CBM reservoirs. Our study area is in the Qinshui Basin, China. Utilizing Hyperion datasets, the endmember spectra of specific minerals were extracted and the carbonate was mapped by the Spectral Angle Mapper (SAM) algorithm and the hydrocarbons in soil were detected by the Normalized Hydrocarbon Index (NHI). Because the vegetation endmembers in this study can produce similar absorption feature at 1730nm, the distribution of the vegetation was obtained by SAM. The results show that the carbonate and hydrocarbons concentrated in Jincheng Coal Mining Area. This approach, using the hyperspectral datasets, is advantageous for CBM exploration. Chengye Zhang 0001, Qiming Qin, Li Chen 0008, Nan Wang 0006, Yanbing Bai |
IGARSS | 3 |
| 2014 | Analysis and design of passive super low frequency detection systemabstractNatural source Super Low Frequency electromagnetic prospecting methods have seen an increasingly promising potential in geophysical applications. As the natural source electromagnetic signal is weak, the electromagnetic signal acquisition and processing hardware technology aiming at getting electromagnetic signals with high signal to noise ratio and data inversion accuracy is in great need. In this paper, we designeda signal processing and acquisition hardware system based on ARM platform, i.e., an embedded system, which was specific to the output signal from induction magnetic field. The FFT computation is conducted in our processor to convert time domain in to frequency domain. The relevant information is display on a LCD screen and the acquired data can be stored in storage module for further processing. The software system is based on UCOSII, and FAT32 FS, to improve operability of the system. Qiming Qin, Nan Wang 0006, Li Chen 0008, Yan BingBai, Chengye Zhang 0001 |
IGARSS | 4 |
| 2013 | The quantitative prediction of Coalbed Methane gas content based on super-low frequency electromagnetic technologyabstractAbundant field experiments have showed that the super low frequency (SLF) electromagnetic detector is sensitive to Coalbed Methane(CBM). The signal curves collected by the SLF electromagnetic detector show high amplitude anomalies in the CBM enrichment areas. Based on this finding, we choose the Qinshui basin as study area, and take advantage of the field SLF electromagnetic data to make quantitative prediction of CBM gas content.The results show that the average error between the estimated value and the measured value is 7.56%. Yanbing Bai, Qiming Qin, Li Chen 0008, Nan Wang 0006 |
IGARSS | 3 |
| 2013 | A method on Coalbed Methane gas content monitoring based on super-low frequency electromagnetic technologyabstractAbundant field experiments have showed that the super low frequency (SLF) electromagnetic detector is sensitive to Coalbed Methane. The signal curves collected by the SLF electromagnetic detector show high amplitude anomalies in the Coalbed Methane enrichment areas. Based on this finding, we choose the Qinshui basin as study area, and take advantage of the field data to seeking the coupleing relationship between the SLF electromagnetic data and Coalbed Methane gas content. The results show that the passive super-low frequency electromagnetic detection technology can effectively monitor the longer time span dynamic of Coalbed Methane gas content. Yanbing Bai, Qiming Qin, Li Chen 0008, Nan Wang 0006, Hongbo Jiang 0001 |
IGARSS | 3 |
| 2013 | Integrating remote sensing and Super-Low Frequency electromagnetic technology in exploration of buried faultsabstractThe buried faults are widespread in the coal-bed, which result in great difficulties in the construction work. In this paper, an integrated method is used in coal-bed in order to detect and analyze the characteristics of buried faults. Firstly, the lineaments are interpreted by visual interpretation from the ETM+ image, and several lineaments enriched areas are picked up. Secondly Super-Low Frequency (SLF) electromagnetic detection is conducted in these areas. Finally, combined with the geology information, lineament distribution and the SLF data, the characteristics of the buried faults are delineated. The results show that the near EW trending normal faults exist widely in the study area by this method, and the depth of the buried faults are presumed in 450-600 m that are coherent with the available drilling data. Li Chen 0008, Qiming Qin, Yanbing Bai, Nan Wang 0006, Jun Wang 0042 |
IGARSS | 1 |
| 2013 | Automatic building extraction from very high resolution satellite imagery using line segment detectorabstractThis paper presents an automatic procedure for rapid building extraction from optical very high resolution (VHR) satellite imagery. Classical extraction models are always complex and time-consuming. The optimized process of building extraction consists of three main rapid stages: edge-preserving and smoothing bilateral filter, line segment detection, perceptual grouping polygonal building boundary. Firstly, we use bilateral filter to smooth original image with edge-preserving. Secondly, a state-of-the-art line segment detector (LSD) algorithm gives highly accurate building contour segments. Finally, we apply the perceptual grouping approach based on graph search to organize detected contour line segments of interested buildings. We test our method on optical VHR QuickBird satellite imagery and obtain promising experimental results with overall accuracy of 79.1%, which confirm the effectiveness and robustness of this linear-time procedure. Jun Wang 0042, Qiming Qin, Li Chen 0008, Xin Ye 0001, Xuebin Qin |
IGARSS | 3 |
| 2013 | Coal-bed Methane reservoir identification using the natural source Super-Low Frequency remote sensingabstractThe goal of this paper is to develop and analyze the natural source Super-Low Frequency (SLF) remote sensing using the BD-6 detector and its data processing and interpretation system to help with Coal-bed Methane (CBM) reservoir information extraction. We delineated the diagram of the SLF remote sensing technique, and especially illustrated the integrated method of the Independent Component Analysis (ICA) and Wavelet-Lifting Wavelet Transform to suppress time-varying 150Hz and 250Hz power frequency electromagnetic interference (EMI). In the application of interpreting enrichment layers of (CBM), we obtained the SLF interpretation signs to identify CBM reservoirs and features. The result demonstrates that the SLF remote sensing provides a prosperous perspective on the detection and demarcation of underground geo-objects. Nan Wang 0006, Qiming Qin, Li Chen 0008, Yanbing Bai |
IGARSS | 4 |
| 2013 | SALT: Sensing enAbled Localization and Tracking for geolocation database in TV white spaceabstractLocalization is the enabling technology of geolocation database assisted Cognitive Radio Networks (CRN) in TV White Space (TVWS), as the location is required in the implementation of geolocation database approach by regulators such as FCC and Ofcom. This paper proposes a novel Sensing enAbled Localization and Tracking (SALT) system, which locates Secondary Users (SUs) and tracks their movements. In particular, the system localize SUs based on their measurement of Primary User (PU) signals in the TV spectrum. SALT also uses the SU movement information to dynamically manage the available spectrum with an optimal spectrum allocation algorithm based on the location estimation, so as to maximize the uplink throughput of the entire network. The performance of the proposed localization and tracking algorithm is verified with field measurement data from complex environment through extensive experimentation. Li Chen 0008, Tengyi Zhang, Danny H. K. Tsang |
IWCMC | 1 |
| 2012 | Remote sensing information of mineralizing alteration extraction methodsabstractRemote sensing technology is considered a fast and effective method to prospect ore. Now, this method is used in Gejiu tin deposit of YunNan in order to extract more accurate mineralization abnormal information. In this study, first through the band math method and principal component analysis method, the mineralization alternation can be extracted in ETM data. Then using ASTER data the limonitization, the chloritization and the dolomitization are extracted by the spectral angle method. At last, the trace elements of the vegetation are statistically analyzed and the vegetation mineralization alteration information is extracted by two different methods in ASTER data. The result shows that the alternation information distributions are consistent in the east-south study area and match with the field exploration. Consequently the extracted results are effective. Li Chen 0008, Qiming Qin, Hongbo Jiang 0001 |
IGARSS | 1 |