VLDB 2026 Research / reviewers in the wild / expert
Yanshu Wang
dblp:231/7712
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Generative modeling · 58% Time series and sequential data · 29% Deep learning architectures and training · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 51% Hardware accelerators and domain-specific architectures · 33% Cloud and datacenter computing · 11% | |
| Computer networks
4 papers |
Software-defined and programmable networks · 53% Network measurement and analytics · 32% Datacenter networks · 16% | |
| Databases, data mining, and information retrieval
1 paper |
Data stream processing · 100% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software-defined and programmable networks
programmable data plane |
0.9 | 2 | 2026 | Elixir: A High-performance and Low-cost Approach to Managing Hardware/Software Hybrid Flow Tables Considering Flow Burstiness · NSDI 2022 KaleidoScope: A Co-Processor for Neural-Network-Driven Intelligent Data Plane · IEEE Trans. Computers 2026 |
Machine learning › Generative modeling › synthetic data generation
anomaly generation |
0.9 | 1 | 2025 | FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly Synthesis · NeurIPS 2025 |
Machine learning › Time series and sequential data › anomaly detection
anomaly segmentation |
0.9 | 1 | 2025 | FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly Synthesis · NeurIPS 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly Synthesis · NeurIPS 2025 |
Data stream processing
sketch |
0.9 | 1 | 2025 | DaVinci Sketch: A Versatile Sketch for Efficient and Comprehensive Set Measurements · ICDE 2025 |
Network measurement and analytics
sketch data structures |
0.9 | 1 | 2025 | DaVinci Sketch: A Versatile Sketch for Efficient and Comprehensive Set Measurements · ICDE 2025 |
Distributed systems
distributed machine learning |
0.8 | 2 | 2020 | A Scalable, High-Performance, and Fault-Tolerant Network Architecture for Distributed Machine Learning · IEEE/ACM Trans. Netw. 2020 BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML Training · NeurIPS 2018 |
Distributed systems › distributed machine learning
gradient synchronization |
0.8 | 2 | 2020 | A Scalable, High-Performance, and Fault-Tolerant Network Architecture for Distributed Machine Learning · IEEE/ACM Trans. Netw. 2020 BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML Training · NeurIPS 2018 |
Software-defined and programmable networks
flow table management |
0.6 | 1 | 2022 | Elixir: A High-performance and Low-cost Approach to Managing Hardware/Software Hybrid Flow Tables Considering Flow Burstiness · NSDI 2022 |
Datacenter networks › data center network topology
fault-tolerant topology |
0.4 | 1 | 2020 | A Scalable, High-Performance, and Fault-Tolerant Network Architecture for Distributed Machine Learning · IEEE/ACM Trans. Netw. 2020 |
Cloud and datacenter computing
datacenter network |
0.3 | 1 | 2018 | BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML Training · NeurIPS 2018 |
Machine learning › Deep learning architectures and training
neural network inference |
0.3 | 1 | 2026 | KaleidoScope: A Co-Processor for Neural-Network-Driven Intelligent Data Plane · IEEE Trans. Computers 2026 |
Machine learning › Efficient and distributed learning
distributed training |
0.1 | 1 | 2018 | BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML Training · NeurIPS 2018 |
Methods — techniques the papers use, named apart from their topics
co-processor design · 2.0foreground-aware reconstruction · 1.7accelerated sampling · 1.7coprocessor design · 1.0fully-distributed synchronization · 0.9RDMA transport · 0.9BCube topology · 0.9parameter server · 0.7BCube · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KaleidoScope: A Co-Processor for Neural-Network-Driven Intelligent Data Plane
Dong Wen 0004, Zhongpei Liu, Tong Yang 0003, Tianyun Li, Yanshu Wang, Tao Li 0008, Zhuochen Fan, Qing Li 0006, Zhigang Sun 0002 |
IEEE Trans. Computers | 5 |
| 2025 | DaVinci Sketch: A Versatile Sketch for Efficient and Comprehensive Set MeasurementsabstractSet measurements are fundamental in numerous areas including network measurement, database queries, and data mining. These measurements are typically executed on multisets. Existing algorithms optimize a specific set measurement task, leading to sophisticated but narrowly focused solutions. This specialization often results in inefficiencies when multiple set measurement tasks are required simultaneously, consuming excessive computational and storage resources. This paper introduces DaVinci Sketch, a versatile sketch designed to efficiently handle various set measurement tasks using a single unified data structure. DaVinci Sketch employs a novel approach by utilizing a dedicated structure to store frequent elements, thereby reducing collisions among flows that have the most significant impact on results. Remarkably, DaVinci Sketch can simultaneously perform up to nine different measurement tasks with a single data structure and a unified operation, whereas other approaches typically support fewer tasks. The experimental results demonstrate that DaVinci Sketch achieves high accuracy across 9 measurement tasks. Furthermore, in multi-task scenarios, DaVinci Sketch significantly reduces the memory usage (by more than 59%) and achieves high throughput (more than 23 times faster than other methods). Yanshu Wang, Jianan Ji, Chao-Hsuan Liu, Hengyang Zhou, Tong Yang 0003 |
ICDE | 1 |
| 2025 | Residual Aggregation and Multi-Head Attention Reweighting for Autoformer in Industrial Time Series ForecastingabstractMultivariate time series forecasting is critical for industrial applications such as predictive maintenance and anomaly detection. However, existing Transformer-based models often struggle to capture short-term residual patterns and are vulnerable to noise due to their uniform treatment of attention heads and lack of adaptability to input dynamics. To address these challenges, we propose the Residual Aggregation and Multi-Head Attention Reweighting Autoformer (RAMAR), an enhanced Autoformer-based architecture tailored for complex industrial environments. RAMAR introduces two complementary modules: the Local Residual Aggregator (LRA), which refines high-frequency residuals via bottleneck convolutions and gated fusion, and the Dynamic Head Reweighting Module (DHRM), which calibrates attention heads using multi-scale temporal–channel convolutions to suppress noisy activations and emphasize informative signals. Extensive experiments on three public ETT datasets demonstrate that RAMAR consistently outperforms strong baselines across multiple evaluation metrics, confirming its robustness and effectiveness for real-world industrial forecasting. Yanshu Wang, Xichen Xu, Xiaoning Lei, Chengbin Ma |
IECON | 1 |
| 2025 | FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly SynthesisabstractIndustrial anomaly segmentation relies heavily on pixel-level annotations, yet real-world anomalies are often scarce, diverse, and costly to label. Segmentation-oriented industrial anomaly synthesis (SIAS) has emerged as a promising alternative; however, existing methods struggle to balance sampling efficiency and generation quality. Moreover, most approaches treat all spatial regions uniformly, overlooking the distinct statistical differences between anomaly and background areas. This uniform treatment hinders the synthesis of controllable, structure-specific anomalies tailored for segmentation tasks. In this paper, we propose FAST, a foreground-aware diffusion framework featuring two novel modules: the Anomaly-Informed Accelerated Sampling (AIAS) and the Foreground-Aware Reconstruction Module (FARM). AIAS is a training-free sampling algorithm specifically designed for segmentation-oriented industrial anomaly synthesis, which accelerates the reverse process through coarse-to-fine aggregation and enables the synthesis of state-of-the-art segmentation-oriented anomalies in as few as 10 steps. Meanwhile, FARM adaptively adjusts the anomaly-aware noise within the masked foreground regions at each sampling step, preserving localized anomaly signals throughout the denoising trajectory. Extensive experiments on multiple industrial benchmarks demonstrate that FAST consistently outperforms existing anomaly synthesis methods in downstream segmentation tasks. We release the code in https://github.com/Chhro123/fast-foreground-aware-anomaly-synthesis. Xichen Xu, Yanshu Wang, Jinbao Wang 0001, Xiaoning Lei, Guoyang Xie, Guannan Jiang, Zhichao Lu |
NeurIPS | 2 |
| 2025 | PhraseBT: A phrase-level back-translation data augmentation method for neural machine translation
Dashuai Deng, Yanshu Wang, Tadahiro Matsumoto |
Neurocomputing | 4 |
| 2022 | Elixir: A High-performance and Low-cost Approach to Managing Hardware/Software Hybrid Flow Tables Considering Flow Burstiness
Yanshu Wang, Dan Li 0001, Yuanwei Lu |
NSDI | 1 |
| 2021 | FastKeeper: A Fast Algorithm for Identifying Top-k Real-time Large FlowsabstractPrecise identification of large flows is a critical task in network traffic measurement. Previous works focus on identification of elephant flows (i.e. large flows from the beginning of the measurement). However, we generally observe that the flow rates change periodically and abruptly. In addition, the large flows may become small flows over time. Thus, elephant flows are not equal to the real-time large flows, and previous works cannot be used for the identification of the real-time large flows that is more meaningful for modern network applications. Nevertheless, identification of real-time large flows is challenging in that it requires accurate measurement of real-time flow rates and timely replacement of flows that have become small in the measurement data structure. In this paper, we propose FastKeeper to identify real-time large flows with a primary goal of simultaneously achieving low over-head, high performance and high accuracy. FastKeeper employs a sliding-window-based algorithm for accurate measurement of real-time flow rates and a bitmap-voting algorithm for timely replacement of flows that have become small in the measurement data structure. We evaluate it on DPDK using the traces from an operator network, and the evaluation demonstrates that it achieves high accuracy (98%) and processing throughput (25.33Mpps). Yanshu Wang, Dan Li 0001 |
GLOBECOM | 1 |
| 2020 | A Scalable, High-Performance, and Fault-Tolerant Network Architecture for Distributed Machine LearningabstractIn large-scale distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a scalable, high-performance and fault-tolerant DML network architecture on top of Ethernet and commodity devices. BML builds on BCube topology, and runs a fully-distributed gradient synchronization algorithm. Compared to a Fat-Tree network with the same size, a BML network is expected to take much less time for gradient synchronization, for both low theoretical synchronization time and its benefit to RDMA transport. With server/link failures, the performance of BML degrades in a graceful way. Experiments of MNIST and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4% compared with Fat-Tree running state-of-the-art gradient synchronization algorithm. Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia |
IEEE/ACM Trans. Netw. | 5 |
| 2018 | BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML TrainingabstractIn distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a new gradient synchronization algorithm with higher network performance and lower network cost than the current practice. BML runs on BCube network, instead of using the traditional Fat-Tree topology. BML algorithm is designed in such a way that, compared to the parameter server (PS) algorithm on a Fat-Tree network connecting the same number of server machines, BML achieves theoretically 1/k of the gradient synchronization time, with k/5 of switches (the typical number of k is 2∼4). Experiments of LeNet-5 and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4%. Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia |
NeurIPS | 5 |