VLDB 2026 Research / reviewers in the wild / expert
Weicong Chen 0002
dblp:194/1786-2
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-0573-8808ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scheduling Optimization of Distributed LLM Inference
Biyao Zhang, Weicong Chen 0002, Vipin Chaudhary |
HPDC | 2 |
| 2025 | Labeling Copilot: A Deep Research Agent for Automated Data Curation in Computer Vision
Debargha Ganguly, Ishwar B. Balappanawar, Weicong Chen 0002, Shashank Kambhatla, Srinivasan Iyengar, Shivkumar Kalyanaraman, Ponnurangam Kumaraguru, Vipin Chaudhary |
IEEE Big Data | 4 |
| 2025 | K4: Online Log Anomaly Detection Via Unsupervised Typicality LearningabstractLog anomaly detection (LogAD) is crucial for identifying failures and threats in large-scale computing and cyber-infrastructure systems. However, most existing LogAD approaches suffer from key limitations: they depend on slow and error-prone log parsing, employ tightly coupled end-to-end pipelines, often require supervision for improved detection performance, and rely on flawed single-pass evaluation protocols that fail to reflect the temporal dynamics of real-world online detection. These issues significantly hinder scalability, adaptability, and the practical deployment of solutions. To address these limitations, we introduce$\mathbf{K}^{\mathbf{4}}$(Knowing the Unknown by Knowing only the Known), a fully unsupervised, parser-independent, and representation-agnostic LogAD framework designed for high-performance online detection. At its core,$K^{4}$is grounded in a novel formulation based on representation-level typicality estimation, which transforms arbitrary log embeddings into compact and interpretable four-dimensional descriptors: Precision, Recall, Density, and Coverage (PRDC), which are swiftly computed via GPU-acceleration into geometric$k$-nearest neighbor statistics. These descriptors inform lightweight, modular detectors, including KDE, GMM, OCSVM, and a new adaptation of DeepSVDD, which enables efficient and accurate anomaly scoring without relying on structured formats or log representation retraining. To support realistic deployment scenarios, we also propose a principled streaming-faithful evaluation protocol that partitions datasets into fixed-size chunks and applies sliding-window sampling with strides to emulate online log ingestion, alleviating the performance overestimation and dataset undercoverage issues of prior single-pass evaluations and enabling reproducible benchmarking across datasets with varying anomaly densities. Using this setup, we conduct over$\mathbf{1 2 5, 0 0 0}$experiments across three real-world datasets (HDFS, BGL, Thunderbird), six pretrained embedding models, four detectors, and multiple training and log sampling configurations. Compared to six representative baseline methods spanning supervised, semi-supervised, self-supervised, and unsupervised paradigms,$K^{4}$consistently sets new state-of-the-art results (AUROC: 0.995-0.999, F1: 0.989-0.992) and outperforms all baselines by large margins, while keeping detector training under 4 seconds and per-sample inference latency as low as$4 \mu ~\mathrm{s}$, which are orders of magnitude faster than the most competitive alternatives. Weicong Chen 0002, Vikash Singh, Zahra Rahmani, Debargha Ganguly, Mohsen Hariri, Vipin Chaudhary |
HiPC | 1 |
| 2025 | FedDES: Discrete Event Based Performance Simulation for Federated Learning SystemsabstractFederated Learning (FL) is a scalable and privacy-preserving paradigm well-suited for edge computing. Real-world FL deployments face substantial systems challenges such as compute variability and communication delays, motivating researchers to leverage simulation before real deployment. Most existing FL simulators, however, struggle to scale efficiently and incur long runtimes even for small workloads. To address this, we present FedDES, a high-fidelity, framework-agnostic discrete-event simulation platform that accurately models the runtime behavior of FL systems, including client training, communication overhead, network dynamics, and aggregation strategies. FedDES supports flexible configurations and diverse aggregation approaches, achieving simulation error within 2% of real deployments and delivering over 1000× speedup compared to prior tools. Large-scale experiments with up to 131,072 clients further show that the aggregation strategy critically affects performance, especially under heterogeneous and variable network conditions typical of edge environments. Zhonghao Chen, Weicong Chen 0002, Kibaek Kim, Guanpeng Li, Sheng Di, Xiaoyi Lu 0001 |
SEC | 2 |
| 2025 | SPRT2: Scalable, Parallel, and Real-Time fMRI Data Analysis on Heterogeneous ArchitecturesabstractReal-time functional Magnetic Resonance Imaging (fMRI) data analysis using the Sequential Probability Ratio Test (SPRT) enables dynamic adjustments to experimental protocols and early session termination, improving data quality and reducing patient fatigue. However, implementing SPRT in real-time fMRI analysis presents significant challenges due to the need for large-scale, high-dimensional data processing within strict time constraints. Furthermore, the ongoing advancements in fMRI hardware are driving a data explosion in the field, necessitating solutions that scale effectively. Existing approaches fall short in meeting real-time requirements and fail to fully exploit High-Performance Computing (HPC) and Big Data technologies. In this paper, we introduce Scalable, Parallel, and Real-Time Sequential Probability Ratio Test ($\text{SPRT}^{2}$), a toolkit that integrates HPC and Big Data techniques to enable efficient real-time SPRT-based fMRI data analysis.$\text{SPRT}^{2}$combines novel performance optimizations, such as hint-assisted matrix chain multiplication and sparse matrix techniques on heterogeneous architectures (CPUs and GPUs), with an Apache Spark-based framework for scalability and fault tolerance. Evaluated across 23 human subject experiments,$\text{SPRT}^{2}$achieves real-time analysis within the 1 -second repetition time while minimizing computational resource utilization (just 180 CPU cores).$\text{SPRT}^{2}$reduces session lengths by up to 33% and improves data quality. Furthermore,$\text{SPRT}^{2}$demonstrates near-linear scalability, efficiently processing synthetic datasets (35.9 billion voxels) over HPC platforms with 1,000 CPU cores or 8 NVIDIA A100 GPUs. To the best of our knowledge,$\text{SPRT}^{2}$is the first solution to integrate HPC and Big Data technologies for real-time fMRI analysis, setting a new standard in computational neuroscience. This work highlights the convergence of HPC and Big Data technologies and opens new avenues for tackling complex computational challenges in scalable and real-time fMRI data analysis. Weicong Chen 0002, Sarah J. Carr, Curtis Tatsuoka, Xiaoyi Lu 0001 |
IPDPS | 1 |
| 2025 | SBMGT: Scaling Bayesian Multinomial Group TestingabstractGroup testing is a widely used binary classification method that efficiently distinguishes between samples with and without a binary-classifiable attribute by pooling and testing subsets of a group. Bayesian Group Testing (BGT) is the state-of-the-art approach, which integrates prior risk information into a Bayesian Boolean Lattice framework to minimize test counts and reduce false classifications. However, BGT, like other existing group testing techniques, struggles with multinomial group testing, where samples have multiple binary-classifiable attributes that can be individually distinguished simultaneously. We address this need by proposing Bayesian Multinomial Group Testing (BMGT), which includes a new Bayesian-based model and supporting theorems for an efficient and precise multinomial pooling strategy. We further design and develop SBMGT, a high-performance and scalable framework to tackle BMGT's computational challenges by proposing three key innovations: 1) a parallel binary-encoded product lattice model with up to 99.8% efficiency; 2) the Bayesian Balanced Partitioning Algorithm (BBPA), a multinomial pooling strategy optimized for parallel computation with up to 97.7% scaling efficiency on 4096 cores; and 3) a scalable multinomial group testing analytics framework, demonstrated in a real-world disease surveillance case study using AIDS and STDs datasets from Uganda, where SBMGT reduced tests by up to 54% and lowered false classification rates by 92% compared to BGT. Weicong Chen 0002, Hao Qi 0008, Curtis Tatsuoka, Xiaoyi Lu 0001 |
PPoPP | 1 |
| 2025 | DPAR: High-Performance, Secure, and Scalable Differential Privacy-based AllReduceabstractSecure, efficient, and scalable AllReduce-based data aggregation is essential for Artificial Intelligence (AI) and scientific applications on modern High-Performance Computing (HPC) and cloud infrastructures. As AllReduce is increasingly used across these distributed infrastructures, privacy has become a critical concern. State-of-the-art (SOTA) Homomorphic Encryption (HE)-based AllReduce solutions introduce high overhead, require secure key exchanges, and remain vulnerable to collusion. We propose DPAR, the first differentially private, collusion-resistant AllReduce framework optimized for large-scale HPC and AI workloads. DPAR introduces three key innovations: integrating Differential Privacy (DP) to eliminate collusion risks without key exchanges, scalable noise growth to preserve accuracy, and performance optimizations using a noise pooling mechanism. DPAR is a drop-in Message Passing Interface (MPI) AllReduce replacement, providing strong privacy with minimal performance cost. Evaluated on Delta and Frontier supercomputers with up to 8192 cores, DPAR outperforms the SOTA HE solution by up to 34.7% in modern AI workloads. Hao Qi 0008, Weicong Chen 0002, Chenghong Wang, Xiaoyi Lu 0001 |
SC | 2 |
| 2024 | Accelerating Lossy and Lossless Compression on Emerging BlueField DPU ArchitecturesabstractData compression has become a crucial technique in addressing performance bottlenecks caused by increasing data volumes in High-Performance Computing (HPC), Big Data, and Deep Learning (DL). Despite its potential to boost system performance, recent studies have identified significant challenges with existing compression methods, mainly due to their high computational demands amidst continuously growing data sizes. Concurrently, the advent of Data Processing Units (DPUs), equipped with programmable System-on-Chip (SoC) and specialized compression accelerators, offers a promising opportunity to alter the landscape of data compression. This paper explores the complexities and potential of leveraging NVIDIA BlueField DPUs to accelerate lossy and lossless compression. Towards this, we introduce PEDAL, an innovative library that leverages the hardware capabilities of DPUs to unify and optimize data compression designs. Moreover, we seamlessly co-design PEDAL with the popular MPICH MPI library, demonstrating up to 101x speedup in compression time and 88x decrease in communication latency. Drawing on these achievements, we share our experience with various research communities about accelerating data compression on DPUs in communication-oriented HPC scenarios. Yuke Li 0003, Arjun Kashyap, Weicong Chen 0002, Yanfei Guo, Xiaoyi Lu 0001 |
IPDPS | 3 |
| 2023 | Performance Characterization of Large Language Models on High-Speed InterconnectsabstractLarge Language Models (LLMs) have recently gained significant popularity due to their ability to generate human-like text and perform a wide range of natural language processing tasks. Training these models usually requires a large amount of computational resources and is often done in a distributed manner. The use of high-speed interconnects can significantly influence the efficiency of distributed training. Therefore, there poses a need for systematic studies to explore the distributed training characteristics of these models on high-speed interconnects. This paper presents a comprehensive performance characterization of representative large language models: GPT, BERT, and T5. We evaluate their training performance in terms of iteration time, interconnect utilization, and scalability, over different high-speed interconnects and communication protocols, including TCP/IP, IPoIB, and RDMA. We observe that interconnects play a vital role in LLM training. Specifically, RDMA-100 Gbps outperforms IPoIB-100 Gbps and TCP/IP-10 Gbps by an average of 2.51x and 4.79x regarding training iteration time, and scores the highest interconnect utilization (up to 60 Gbps) in both strong and weak scaling, compared to IPoIB with up to 20 Gbps and TCP/IP with up to 9 Gbps, leading to the shortest training time. We also observe that larger models tend to have higher requirements for communication bandwidth, especially for AllReduce during backward propagation, which can take up to 91.12% of training time. Through our evaluation, we envision opportunities to improve the communication time for better training performance of LLMs. We extensively explore and summarize the role communication plays in distributed LLM training. Hao Qi 0008, Liuyao Dai, Weicong Chen 0002, Zhen Jia 0001, Xiaoyi Lu 0001 |
HOTI | 3 |
| 2023 | SBGT: Scaling Bayesian-based Group Testing for Disease SurveillanceabstractThe COVID-19 pandemic underscored the necessity for disease surveillance using group testing. Novel Bayesian methods using lattice models were proposed, which offer substantial improvements in group testing efficiency by precisely quantifying uncertainty in diagnoses, acknowledging varying individual risk and dilution effects, and guiding optimally convergent sequential pooled test selections using a Bayesian Halving Algorithm. Computationally, however, Bayesian group testing poses considerable challenges as computational complexity grows exponentially with sample size. This can lead to shortcomings in reaching a desirable scale without practical limitations. We propose a new framework for scaling Bayesian group testing based on Spark: SBGT. We show that SBGT is lightning fast and highly scalable. In particular, SBGT is up to 376x, 1733x, and 1523x faster than the state-of-the-art framework in manipulating lattice models, performing test selections, and conducting statistical analyses, respectively, while achieving up to 97.9% scaling efficiency up to 4096 CPU cores. More importantly, SBGT fulfills our mission towards reaching applicable scale for guiding pooling decisions in wide-scale disease surveillance, and other large scale group testing applications. Weicong Chen 0002, Hao Qi 0008, Xiaoyi Lu 0001, Curtis Tatsuoka |
IPDPS | 1 |
| 2022 | HiBGT: High-Performance Bayesian Group Testing for COVID-19abstractThe COVID-19 pandemic has necessitated disease surveillance using group testing. Novel Bayesian methods using lattice models were proposed, which offer substantial improvements in group testing efficiency by precisely quantifying uncertainty in diagnoses, acknowledging varying individual risk and dilution effects, and guiding optimally convergent sequential pooled test selections. Computationally, however, Bayesian group testing poses considerable challenges as computational complexity grows exponentially with sample size. HPC and big data stacks are needed for assessing computational and statistical performance across fluctuating prevalence levels at large scales. Here, we study how to design and optimize critical computational components of Bayesian group testing, including lattice model representation, test selection algorithms, and statistical analysis schemes, under the context of parallel computing. To realize this, we propose a high-performance Bayesian group testing framework named HiBGT, based on Apache Spark, which systematically explores the design space of Bayesian group testing and provides comprehensive heuristics on how to achieve high-performance, highly scalable Bayesian group testing. We show that HiBGT can perform large-scale test selections (> 250state iterations) and accelerate statistical analyzes up to 15.9x (up to 363x with little trade-offs) through a varied selection of sophisticated parallel computing techniques while achieving near linear scalability using up to 924 CPU cores. Weicong Chen 0002, Curtis Tatsuoka, Xiaoyi Lu 0001 |
HIPC | 1 |