VLDB 2026 Research / reviewers in the wild / expert
Weijian Zheng
dblp:193/6710
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FastREI: Fast Rare Event Identification on X-ray Data with Cross-Stage Optimizations
Zhiqing Zhong, Weijian Zheng, Hemant Sharma, Jun-Sang Park, Peter Kenesei, Antonino Miceli, Rajkumar Kettimuthu, Xiaodong Yu 0001 |
IEEE Big Data | 4 |
| 2024 | Model and Data Management for Machine Learning (M2ML): Integrating Instruments, Edge and HPC for Accelerated Machine LearningabstractThe use of data produced by scientific instruments, such as the Advanced Photon Source Upgrade (APS-U), to train and fine-tune machine learning models is becoming increasingly challenging due to high data production rates, large data volumes, and the growing complexity of machine learning models. To address these challenges, researchers have developed frameworks like fairDMS to efficiently organize vast amounts of data and models for rapid querying when model degradation is detected. However, the complexity of these frameworks and the physically distributed nature of experimental facilities complicate their deployment.Here we introduce a high-performance model and data management framework for machine learning, M2ML. In contrast to previous frameworks, M2ML abstracts the tasks into three key elements that can be easily called and accessed by users. M2ML is capable of utilizing a variety of computational resources, that are distributed across scientific facilities, to accelerate machine learning tasks. For example, it can automatically transfer data from an experimental facility (such as APS-U) to a high performance computing (HPC) facility (such as the Argonne Leadership Computing Facility (ALCF)), train machine learning models at the HPC facility, and deploy the trained models on edge computing devices back at the experimental facility for inferencing. M2ML provides a unified interface for (on-the-fly) model (re)training, storage, evaluation, fine-tuning, and inferencing using heterogeneous resources that can be geographically distributed. M2ML uses Globus services such as Globus Transfer and Globus Compute (formerly FuncX). We evaluate M2ML using a high energy diffraction microscopy (HEDM) workflow that employs BraggNN to predict the diffraction peak locations. Results show that, although the BraggNN model is small, M2ML can significantly accelerate the workflow through selective assignment of tasks to different computing resources. Weijian Zheng, Hemant Sharma, Ryan Chard, Peter Kenesei, Jun-Sang Park, Nicholas Schwarz, Antonino Miceli, Ian T. Foster, Rajkumar Kettimuthu |
IEEE Big Data | 1 |
| 2024 | CereSZ: Enabling and Scaling Error-bounded Lossy Compression on Cerebras CS-2abstractToday's scientific applications running on supercomputers produce large volumes of data, leading to critical data storage and communication challenges. To tackle the challenges, error-bounded lossy compression is commonly adopted since it can reduce data size drastically within a user-defined error threshold. Previous work has shown that compression techniques can significantly reduce the storage and I/O overhead while retaining good data quality. However, the existing compressors are mainly designed for CPU and GPU. As new AI chips are being incorporated into supercomputers and increasingly used for accelerating scientific computing, there is a growing demand for efficient data compression on the new architecture. In this paper, we propose an efficient lossy compressor, CereSZ, based on the Cerebras CS-2 system. The compression algorithm is mapped onto Cerebras using both data parallelism and pipeline parallelism. In order to achieve a balanced workload on each processing unit, we propose an algorithm to evenly distribute the pipeline stages. Our experiments with six scientific datasets demonstrate that CereSZ can achieve a throughput from 227.93 GB/s to 773.8 GB/s, 2.43x to 10.98x faster than existing GPU compressors. Shihui Song, Yafan Huang, Peng Jiang 0004, Xiaodong Yu 0001, Weijian Zheng, Sheng Di, Qinglei Cao, Yunhe Feng, Franck Cappello |
HPDC | 5 |
| 2024 | WorkloadDiff: Conditional Denoising Diffusion Probabilistic Models for Cloud Workload PredictionabstractAccurate workload forecasting plays a crucial role in optimizing resource allocation, enhancing performance, and reducing energy consumption in cloud data centers. Deep learning-based methods have emerged as the dominant approach in this field, exhibiting exceptional performance. However, most existing methods lack the ability to quantify confidence, limiting their practical decision-making utility. To address this limitation, we propose a novel denoising diffusion probabilistic model (DDPM)-based method, termed WorkloadDiff, for multivariate probabilistic workload prediction. WorkloadDiff leverages both original and noisy signals from input conditions using a two-path neural network. Additionally, we introduce a multi-scale feature extraction method and an adaptive fusion approach to capture diverse temporal patterns within the workload. To enhance consistency between conditions and predicted values, we incorporate a resampling strategy into the inference of WorkloadDiff. Extensive experiments conducted on four public datasets demonstrate the superior performance of WorkloadDiff over all baseline models, establishing it as a robust tool for resource management in cloud data centers. Weiping Zheng, Zongxiao Chen, Weijian Zheng, Xiaomao Fan |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | Investigating Code Generation Performance of ChatGPT with Crowdsourcing Social Data
Yunhe Feng, Sreecharan Vanam, Manasa Cherukupally, Weijian Zheng, Meikang Qiu, Haihua Chen 0002 |
COMPSAC | 4 |
| 2023 | Tomo2Mesh: Fast Porosity Mapping and Visualization for Synchrotron TomographyabstractApplications of X-ray computed tomography (CT) in porosity characterization of engineering materials often involve an extensive data analysis workflow. This workflow includes CT reconstruction of raw projection data, binarization, labeling, and mesh extraction. Mapping porosity in larger samples presents a significant computational challenge, as it requires processing gigabytes of raw data to extract porosity information, which becomes a critical bottleneck in the analysis. In this study, we present algorithms and an implementation of an end-to-end porosity mapping framework. Our framework processes raw projection data obtained from a synchrotron CT instrument, generating a porosity map and a visualization in the form of a triangular face mesh. To achieve this objective, we introduce a novel subset reconstruction scheme for X-ray CT, combining filtered back-projection and a convolutional neural network. This scheme allows us to reconstruct subsets of a tomography object with arbitrary shapes. Building upon this scheme, we have developed a fast and efficient framework for porosity mapping. Initially, our framework detects potential voids by performing a coarse reconstruction on down-sampled projections. Subsequently, we enhance the shape of these voids by reconstructing selected subsets from the original raw data, providing higher detail. To evaluate the performance of our framework, we measured the processing time from raw data to a triangular face mesh across multiple visualization scenarios. Our experiments were conducted on a single high-performance workstation equipped with a GPU. The results demonstrate that our framework enables the visualization of local porosity within an 8-gigavoxel CT volume (12 gigabytes raw data) in just 1 to 2 minutes. Moreover, for a larger 64-gigavoxel CT volume (100 gigabytes of raw data), the visualization can be generated within 3 to 7 minutes, showcasing the efficiency of our approach. Aniket Tekawade, Viktor V. Nikitin, Yashas Satapathy, Zhengchun Liu, Peter Kenesei, Weijian Zheng, Francesco De Carlo, Ian T. Foster, Rajkumar Kettimuthu |
e-Science | 7 |
| 2021 | Micromobility in Smart Cities: A Closer Look at Shared Dockless E-Scooters via Big Social DataabstractThe micromobility is shaping first- and last-mile travels in urban areas. Recently, shared dockless electric scooters (e-scooters) have emerged as a daily alternative to driving for short-distance commuters in large cities due to the affordability, easy accessibility via an app, and zero emissions. Meanwhile, e-scooters come with challenges in city management, such as traffic rules, public safety, parking regulations, and liability issues. In this paper, we collected and investigated 5.8 million scooter-tagged tweets and 144,197 images, generated by 2.7 million users from October 2018 to March 2020, to take a closer look at shared e-scooters via crowdsourcing data analytics. We profiled e-scooter usages from spatial-temporal perspectives, explored different business roles (i.e., riders, gig workers, and ridesharing companies), examined operation patterns (e.g., injury types, and parking behaviors), and conducted sentiment analysis. To our best knowledge, this paper is the first large-scale systematic study on shared e-scooters using big social data. Yunhe Feng, Dong Zhong, Peng Sun 0003, Weijian Zheng, Qinglei Cao, Zheng Lu 0005 |
ICC | 4 |
| 2016 | suCAQR: A Simplified Communication-Avoiding QR Factorization Solver Using the TBLAS FrameworkabstractThe scope of this paper is to design and implement a scalable QR factorization solver that can deliver the fastest performance for tall and skinny matrices and square matrices on modern supercomputers. The new solver, named scalable universal communication-avoiding QR factorization (suCAQR), introduces a simplified and tuning-less way to realize the communication-avoiding QR factorization algorithm to support matrices of any shapes. The software design includes a mixed usage of physical and logical data layouts, a simplified method of dynamic-root binary-tree reduction, and a dynamic dataflow implementation. Compared with the existing communication avoiding QR factorization implementations, suCAQR has the benefits of being simpler, more general, and more efficient. By balancing the degree of parallelism and the proportion of faster computational kernels, it is able to achieve scalable performance on clusters of multicore nodes. The software essentially combines the strengths of both synchronization-reducing approach and communication-avoiding approach to achieve high performance. Based on the experimental results using 1,024 CPU cores, suCAQR is faster than DPLASMA by up to 30%, and faster than ScaLAPACK by up to 30 times. Weijian Zheng, Fengguang Song, Zizhong Chen |
ICPADS | 1 |