VLDB 2026 Research / reviewers in the wild / expert
Weiyan Wang
dblp:17/2902
· DBLP profile ↗
16ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploiting Student Parallelism for Low-latency GPU Inference of BERT-like Models in Online ServicesabstractBERT-like models have been widely adopted in text mining and web search due to their high accuracy. However, large BERT-like models suffer from inefficient online inference on GPUs for two main reasons. First, their high accuracy relies on large model depth, which linearly increases sequential computation on GPUs. Second, stochastic and dynamic online workloads lead to extra costs due to batching and padding. To address the problem, we present Student Parallelism for efficient GPU inference of BERT-like models under real-world online workloads. At its core, Student Parallelism adopts stacking distillation and boosting ensemble, distilling the original deep model into a group of shallow but virtually stacked student models running in parallel. This enables Student Parallelism to achieve a low model depth (e.g., two layers), and thus low inference latency while maintaining accuracy. In addition, we design adaptive student pruning to adjust the number of students according to the dynamic online workloads. For example, during workload bursts, it can temporarily decrease the number of students with minimal accuracy loss to improve system throughput. Extensive experiments on real-world datasets and workloads show that Student Parallelism achieves up to 4.1× lower latency while maintaining accuracy and up to 22.27× higher throughput during workload bursts. Weiyan Wang, Yilun Jin, Yiming Zhang 0003, Victor Junqiu Wei, Han Tian, Li Chen 0008, Jinbao Xue, Yangyu Tao, Kai Chen 0005 |
KDD (2) | 1 |
| 2024 | Prometheus: Out-of-distribution Fluid Dynamics Modeling with Disentangled Graph ODEabstractFluid dynamics modeling has received extensive attention in the machine learning community. Although numerous graph neural network (GNN) approaches have been proposed for this problem, the problem of out-of-distribution (OOD) generalization remains underexplored. In this work, we propose a new large-scale dataset Prometheus which simulates tunnel and pool fires across various environmental conditions and builds an extensive benchmark of 12 baselines, which demonstrates that the OOD generalization performance is far from satisfactory. To tackle this, this paper introduces a new approach named Disentangled Graph ODE (DGODE), which learns disentangled representations for continuous interacting dynamics modeling. In particular, we utilize a temporal GNN and a frequency network to extract semantics from historical trajectories into node representations and environment representations respectively. To mitigate the potential distribution shift, we minimize the mutual information between invariant node representations and the discretized environment features using adversarial learning. Then, they are fed into a coupled graph ODE framework, which models the evolution using neighboring nodes and dynamical environmental context. In addition, we enhance the stability of the framework by perturbing the environment features to enhance robustness. Extensive experiments validate the effectiveness of DGODE compared with state-of-the-art approaches. Hao Wu 0094, Huiyuan Wang, Kun Wang 0039, Weiyan Wang, Changan Ye, Yangyu Tao, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
ICML | 4 |
| 2024 | Surge Phenomenon in Optimal Learning Rate and Batch Size ScalingabstractIn current deep learning tasks, Adam-style optimizers—such as Adam, Adagrad, RMSprop, Adafactor, and Lion—have been widely used as alternatives to SGD-style optimizers. These optimizers typically update model parameters using the sign of gradients, resulting in more stable convergence curves.
The learning rate and the batch size are the most critical hyperparameters for optimizers, which require careful tuning to enable effective convergence. Previous research has shown that the optimal learning rate increases linearly (or follows similar rules) with batch size for SGD-style optimizers. However, this conclusion is not applicable to Adam-style optimizers.
In this paper, we elucidate the connection between optimal learning rates and batch sizes for Adam-style optimizers through both theoretical analysis and extensive experiments.
First, we raise the scaling law between batch sizes and optimal learning rates in the “sign of gradient” case, in which we prove that the optimal learning rate first rises and then falls as the batch size increases. Moreover, the peak value of the surge will gradually move toward the larger batch size as training progresses.
Second, we conduct experiments on various CV and NLP tasks and verify the correctness of the scaling law. Shuaipeng Li, Penghao Zhao, Hailin Zhang 0004, Xingwu Sun, Hao Wu 0094, Weiyan Wang, Chengjun Liu, Jinbao Xue, Yangyu Tao, Bin Cui 0001, Di Wang 0052 |
NeurIPS | 7 |
| 2023 | MDP: Model Decomposition and Parallelization of Vision Transformer for Distributed Edge InferenceabstractDistributed edge inference emerges to be a promising paradigm to speed up inference. Previous works make physical partitions on CNNs to realize it, but there are the following challenges for vision transformers: (1) high communication costs for the large model; (2) stragglers because of heterogeneous devices; (3) time-out exceptions due to unstable edge devices.Therefore, we propose a novel Model Decomposition and Parallelization(MDP) for large vision transformers. Inspired by the implicit boosting ensemble in the vision transformer, MDP decomposes it into an explicit boosting ensemble of different and parallel sub-models. It sequentially trains all sub-models to gradually reduce the residual errors. To minimize dependency and communication among sub-models, We adopt stacking distillation to bring every sub-model extra information about others for better error correction. Different sub-models can take both different image sizes and model sizes to run on heterogeneous devices and improve the ensemble diversities. To handle the timeout exception, we add vanilla supervised learning on every submodel for the bagging ensemble in case of the early termination of boosting ensemble. As a result, all sub-models can not only run in parallel without much communication but also can be adapted to the heterogeneous devices, while maintaining accuracy even with time-out exceptions. Experiments show that MDP can outperform other baselines by $5 . 2 \times \sim 2 . 1 \times$ in latency and $5 . 1 \times \sim 1 . 7 \times$ in throughput with comparable accuracy. Weiyan Wang, Yiming Zhang 0003, Yilun Jin, Han Tian, Li Chen 0008 |
MSN | 1 |
| 2023 | Multi-task Learning Based Keywords Weighted Siamese Model for Semantic Retrieval
Mengmeng Kuang, Weiyan Wang, Lie Kang, Qiang Yang 0001, Penghui Hao |
PAKDD (3) | 3 |
| 2023 | Enabling Edge-Cloud Video Analytics for Robotics ApplicationsabstractEmerging deep learning-based video analytics tasks demand computation-intensive neural networks and powerful computing resources on the cloud to achieve high accuracy. Due to the latency requirement and limited network bandwidth, edge-cloud systems adaptively compress the data to strike a balance between overall analytics accuracy and bandwidth consumption. However, the degraded data leads to another issue of poortail accuracy, which means the extremely low accuracy of a few semantic classes and video frames. Autonomous robotics applications especially value the tail accuracy performance but suffer using the prior edge-cloud systems. We present Runespoor, an edge-cloud video analytics system to manage the tail accuracy and enable emerging robotics applications. We train and deploy a super-resolution model tailored for the tail accuracy of analytics tasks on the server to significantly improves the performance on hard-to-detect classes and sophisticated frames. During online operation, we use an adaptive data rate controller to further improve the tail performance by instantly adjusting the data rate policy according to the video content. Our evaluation shows that Runespoor improves class-wise tail accuracy by up to 300%, frame-wise 90%/99% tail accuracy by up to 22%/54%, and greatly improves the overall accuracy and bandwidth trade-off. Weiyan Wang, Duowen Liu, Xin Jin 0008, Junchen Jiang, Kai Chen 0005 |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | Estimation of Planetary Boundary Layer Height From Lidar by Combining Gradient Method and Machine Learning AlgorithmsabstractThe planetary boundary layer height (PBLH) has a significant impact on the energy and material exchange in the atmosphere. The traditional gradient method (GM) determines the PBLH based on the gradient change of the aerosol profile. It is susceptible to the effect of complex atmospheric conditions, which leads to uncertainties in the estimation of PBLH. Here, a random forest (RF) algorithm that considers the vertical distribution of aerosols is proposed to find the PBLH under complex atmospheric conditions. The height of the three minimum local peaks of the range correction signal profile and 7 other variables, such as aerosol layer number (ALN), relative humidity, solar radiation, and other meteorological parameters, from January 2017 to December 2021 is used as RF model input. The radiosonde estimated PBLH (PBLHRS) is used as reference value. The sensitivity analysis indicates that the relative error of RF-estimated PBLH (PBLHRF) is smaller than that of GM-estimated PBLH (PBLHGM), and it decreases with an increase in aerosol optical depth. Moreover, RF achieves good performance under different atmospheric conditions. It can effectively overcome the effects of complex atmospheric conditions in PBLH estimation. Based on the correlation analysis, it is found that the estimation accuracy of the RF algorithm is greatly improved compared with the GM. The correlation coefficient between the PBLHRFand the PBLHRSreaches 0.8, which is much larger than that of the PBLHGM(0.47). Finally, long-term PBLHRFanalysis show that there are obvious diurnal and seasonal variations of PBLH. It increases and then decreases from early morning to late evening. It is highest in summer and lowest in winter. Overall, RF can effectively overcome the shortcomings of traditional GM and has high accuracy and robustness for various atmospheric conditions. The findings obtained here have great potential for lidar application in obtaining reliable PBLH estimations. Hui Li 0113, Boming Liu, Xin Ma 0007, Shikuan Jin, Weiyan Wang, Ruonan Fan, Yingying Ma 0001, Ruyi Wei, Wei Gong 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Multi-objective congestion controlabstractDecades of research on Internet congestion control (CC) have produced a plethora of algorithms that optimize for different performance objectives. Applications face the challenge of choosing the most suitable algorithm based on their needs, and it takes tremendous efforts and expertise to customize CC algorithms when new demands emerge. In this paper, we explore a basic question: can we design a single CC algorithm to satisfy different objectives? Yiqing Ma, Han Tian, Xudong Liao, Junxue Zhang 0001, Weiyan Wang, Kai Chen 0005, Xin Jin 0008 |
EuroSys | 5 |
| 2022 | Addressing Network Bottlenecks with Divide-and-Shuffle Synchronization for Distributed DNN TrainingabstractBulk synchronous parallel (BSP) is the de-facto paradigm for distributed DNN training in today’s production clusters. However, due to the global synchronization nature, its performance can be significantly influenced by network bottlenecks caused by either static topology heterogeneity or dynamic bandwidth contentions. Existing solutions, either system-level optimizations strengthening BSP (e.g., Ring or Hierarchical All-reduce) or algorithmic optimizations replacing BSP (e.g., ASP or SSP, which relax the global barriers), do not completely solve the problem, as they may still suffer from communication inefficiency or risk convergence inaccuracy.In this paper, we present a novel divide-and-shuffle synchronization (DS-Sync) to realize communication efficiency without sacrificing convergence accuracy for distributed DNN training. At its heart, by taking into account the network bottlenecks, DS-Sync improves communication efficiency by dividing workers into non-overlap groups to synchronize independently in a bottleneck-free manner. Meanwhile, it maintains convergence accuracy by iteratively shuffling workers among different groups to ensure a global consensus. We theoretically prove that DS-Sync converges properly in non-convex and smooth conditions like DNN. We further implement DS-Sync and integrate it with PyTorch, and our testbed experiments show that DS-Sync can achieve up to 94% improvements on the end-to-end training time with existing solutions while maintaining the same accuracy. Weiyan Wang, Cengguang Zhang, Liu Yang 0008, Kai Chen 0005, Kun Tan 0002 |
INFOCOM | 1 |
| 2022 | Efficient Two-stage Label Noise Reduction for Retrieval-based TasksabstractThe existence of noisy labels in datasets has always been an essential dilemma in deep learning studies. Previous works detected noisy labels by analyzing the predicted probability distribution generated by the model trained on the same data and calculating the probabilities of each label to be regarded as noise. However, the predicted probability distribution from the whole dataset may introduce overfitting, and the overfitting on noisy labels may induce the probability distribution of clean and noisy items to be not conditional independent, making identification more challenging. Additionally, label noise reduction on image datasets has received much attention, while label noise reduction on text datasets has not. This paper proposes a noisy label reduction method for text datasets, which could be applied at retrieval-based tasks by getting a conditional independent probability distribution to identify noisy labels accurately. The method first generates a candidate set containing noisy labels, predicts the category probabilities by the model trained on the rest cleaner data, and then identifies noisy items by analyzing a confidence matrix. Moreover, we introduce a warm-up module and a sharpened cross-entropy loss function for efficiently training in the first stage. Empirical results on different rates of uniform and random label noise in five text datasets demonstrate that our method can improve the label noise reduction accuracy and end-to-end classification accuracy. Further, we find that the iteration of the label noise reduction method is efficient to high-rate label noise datasets, and our method will not hurt clean datasets too much. Mengmeng Kuang, Weiyan Wang, Lie Kang, Qiang Yan 0001 |
WSDM | 2 |
| 2021 | Enabling Edge-Cloud Video Analytics for Robotics ApplicationsabstractEmerging deep learning-based video analytics tasks demand computation-intensive neural networks and powerful computing resources on the cloud to achieve high accuracy. Due to the latency requirement and limited network bandwidth, edge-cloud systems adaptively compress the data to strike a balance between overall analytics accuracy and bandwidth consumption. However, the degraded data leads to another issue of poor tail accuracy, which means the extremely low accuracy of a few semantic classes and video frames. Autonomous robotics applications especially value the tail accuracy performance but suffer using the prior edge-cloud systems.We present Runespoor, an edge-cloud video analytics system to manage the tail accuracy and enable emerging robotics applications. We train and deploy a super-resolution model tailored for the tail accuracy of analytics tasks on the server to significantly improves the performance on hard-to-detect classes and sophisticated frames. During online operation, we use an adaptive data rate controller to further improve the tail performance by instantly adjusting the data rate policy according to the video content. Our evaluation shows that Runespoor improves class-wise tail accuracy by up to 300%, frame-wise 90%/99% tail accuracy by up to 22%/54%, and greatly improves the overall accuracy and bandwidth trade-off. Weiyan Wang, Duowen Liu, Xin Jin 0008, Junchen Jiang, Kai Chen 0005 |
INFOCOM | 2 |
| 2020 | An FPAR Retrieval Algorithm Based on Deep Learning for Modis Visible Band Surface ReflectanceabstractThe Fraction of photosynthetically active radiation (FP AR) absorbed by vegetation is the important biophysical variables of the most climate, hydrological, biogeochemical, and ecological models. In this paper, the simulative surface reflectance and FP AR were used to train the artificial neural network, which retrieved the FP AR value using the Moderate Resolution Imaging Spectroradiometer (MODIS) surface reflectance in visible spectrum. The retrieval results were comprehensively evaluated using MODIS LAI/FP AR products and available field measurement data. Compared with the ground measurement FP AR, the correlation coefficient (R) is 0.727, the relative average error (RMB) is 1.62, the mean absolute error (MAE) is 0.16, the mean relative error (MRE) is 0.896, and the root mean square error (RMSE) is 0.214. The R of MODIS FPAR products was 0.850, the RMB was 1.715, the MAE was 0.197, the MRE was 1.046, and the RMSE was 0.263. The comparison results show that retrieval FP AR can be performed effectively in the ground site area without the type of vegetation. Huijuan Gao, Xirong Liu, Weiyan Wang |
IGARSS | 3 |
| 2020 | LAI Inversion from MODIS Data Using Deep Belief Network (DBN)abstractLeaf area index (LAI) is a key variable to describe the energy, mass and momentum exchange between the surface and the planetary boundary layer. Remote sensing technology can quickly acquire the change of LAI on a large scale and provide support for the study of ecological environment and climate change. Due to the complexity of the surface and atmospheric structure, the inversion of LAI is still facing great difficulties. Limited by the information provided by remote sensing sensors and the ability of information mining by traditional methods, the accuracy of LAI inversion is not high. In order to improve the accuracy of LAI inversion, a method of LAI inversion based on deep belief network (DBN) is proposed in this paper and is applied to the MODIS data. Results are verified by the ground measurements performed within the framework of the ImagineS (Implementing Multi-Scale Agricultural Indicators Exploiting Sentinels) project, and shows that the LAI obtained by the DBN algorithm (hereinafter referred to as DBN LAI) can reach a high precision, and performs better than current MODIS LAI product (MCD15A2). Weiyan Wang, Huijuan Gao |
IGARSS | 1 |
| 2019 | Rethinking Transport Layer Design for Distributed Machine LearningabstractMotivated by the increasing scale of data, we see a growing need of high performance distributed machine learning systems. Many research works are being proposed to improve distributed machine learning performance. Jiacheng Xia, Gaoxiong Zeng, Junxue Zhang 0001, Weiyan Wang, Wei Bai 0001, Junchen Jiang, Kai Chen 0005 |
APNet | 4 |
| 2016 | Study on fine feature description of multi-aspect SAR observationsabstractThe target feature is sensitive to the aspect angle of SAR observation, making the interpretation and target recognition of the SAR image difficult. The information acquired from a certain aspect angle is partial and incomplete, and the multi-aspect observations have the potential to improve the SAR performance in this aspect. Three topics of fine feature description of multi-aspect SAR observations are discussed, and they are the 3D information extraction, the optimum imaging strategy for anisotropic scatterers, and the multi-aspect scattering feature extraction. The initial results of the real P band airborne circular SAR (CSAR) data and the turn table data show that multi-aspect SAR observations have the encouraging potential capability in target fine feature description. Yun Lin 0002, Wen Hong, Yang Li 0037, Weixian Tan, Lingjuan Yu, Liying Hou, Weiyan Wang |
IGARSS | 9 |
| 2009 | Hash-Search: An Efficient SLCA-Based Keyword Search Algorithm on XML Documents
Weiyan Wang, Xiaoling Wang 0004, Aoying Zhou |
DASFAA | 1 |