Yang Zhang 0102

dblp:06/6785-102 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-0523-8478ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EROICA: Online Performance Troubleshooting for Large-scale Model Training
Yu Guan 0005, Zhiyu Yin, Sheng Cheng 0002, Chaojie Yang, Kun Qian 0004, Tianyin Xu, Yang Zhang 0102, Yong Li 0008, Dennis Cai, Ennan Zhai
NSDI9
2025 Evolution of Aegis: Fault Diagnosis for AI Model Training Service in Production
Jianbo Dong, Kun Qian 0021, Zhilong Zheng, Liang Chen 0001, Yichi Xu, Yikai Zhu, Xue Li 0024, Zhihui Ren, Yang Liu 0245, Yu Guan 0005, Chaojie Yang, Yang Zhang 0102, Man Yuan, Yong Li 0008, Xianlong Zeng, Zhiping Yao, Binzhang Fu, Ennan Zhai, Wei Lin 0016, Dennis Cai
NSDI20
2025 Structured guided diffusion models for industrial defect image generation
Yulai Xie 0001, Xiaoning Pi, Yang Zhang 0102, Fang Ren 0001
Knowl. Based Syst.3
2024 Global-Shared Text Representation Based Multi-Stage Fusion Transformer Network for Multi-Modal Dense Video Captioning
abstract
Dense video captioning aims to detect all events of an uncropped video and generate corresponding textual captions for each event. Multi-modal information is essential to improve the performance of this task, but the existing methods mainly rely on the single visual or dual audio-visual modal input, while completely ignoring the text modal input (subtitle). Since the text data has a similar data representation as video caption words, it is conducive to the performance improvement of video captioning. In this article, we propose a novel framework, called the multi-stage fusion transformer network (MS-FTN), to realize multi-modal dense video captioning by fusing the text, the audio, and the visual features in stages. We present a multi-stage feature fusion encoder that first fuses audio and visual modalities at a lower level and then fuses them with a global-shared text representation at a higher level to generate a set of multi-modal complementary context features. In addition, an anchor-free event proposal module is proposed to efficiently generate a set of event proposals without the complex anchor calculation. Extensive experiments on the subsets of the ActivityNet Captions dataset show that our proposed MS-FTN achieves superior performance and efficient computation. Moreover, the ablation studies demonstrate that the global-shared text representation is more suitable for multi-modal dense video captioning.
Yulai Xie 0001, Jingjing Niu, Yang Zhang 0102, Fang Ren 0001
IEEE Trans. Multim.3
2023 GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning
abstract
Training data pre-processing pipelines are essential to deep learning (DL). As the performance of model training keeps increasing with both hardware advancements (e.g., faster GPUs) and various software optimizations, the data pre-processing on CPUs is becoming more resource-intensive and a severe bottleneck of the pipeline. This problem is even worse in the cloud, where training jobs exhibit diverse CPU-GPU demands that usually result in mismatches with fixed hardware configurations and resource fragmentation, degrading both training performance and cluster utilization. We introduce GoldMiner, an input data processing service for stateless operations used in pre-processing data for DL model training. GoldMiner decouples data pre-processing from model training into a new role called the data worker. Data workers facilitate scaling of data pre-processing to anywhere in a cluster, effectively pooling the resources across the cluster to satisfy the diverse requirements of training jobs. GoldMiner achieves this decoupling in a fully automatic and elastic manner. The key insight is that data pre-processing is inherently stateless, thus can be executed independently and elastically. This insight guides GoldMiner to automatically extract stateless computation out of a monolithic training program, efficiently disaggregate it across data workers, and elastically scale data workers to tune the resource allocations across jobs to optimize cluster efficiency. We have applied GoldMiner to industrial workloads, and our evaluation shows that GoldMiner can transform unmodified training programs to use data workers, accelerating individual training jobs by up to 12.1x. GoldMiner also improves average job completion time and aggregate GPU utilization by up to 2.5x and 2.1x in a 64-GPU cluster, respectively, by scheduling data workers with elasticity.
Zhi Yang 0001, Yu Cheng 0030, Chao Tian 0001, Shiru Ren, Wencong Xiao, Man Yuan, Langshi Chen, Kaibo Liu, Yang Zhang 0102, Yong Li 0045, Wei Lin 0016
Proc. ACM Manag. Data10
2023 Disease Simulation in Airport Scenario Based on Individual Mobility Model
abstract
As the rapid-spreading disease COVID-19 occupies the world, most governments adopt strict control policies to alleviate the impact of the virus. These policies successfully reduced the prevalence and delayed the epidemic peak, while they are also associated with high economic and social costs. To bridge the microscopic epidemic transmission patterns and control policies, simulation systems play an important role. In this work, we propose an agent-based disease simulator for indoor public spaces, which contribute to most of the transmission in cities. As an example, we study Guangzhou Baiyun International Airport, which is one of the most bustling aviation hubs in China. Specifically, we design a high-efficiency mobility generation module to reconstruct the individual trajectories considering both lingering behavior and crowd mobility, which greatly enhances the credibility of the simulated mobility and ensures real-time performance. Based on the individual trajectories, we propose a multi-path disease transmission module optimized for indoor public spaces, which includes three main transmission paths as close contact transmission, aerosol transmission, and object surface transmission. We design a novel convolution-based algorithm to mimic the diffusion process, which can leverage the high concurrent capability of the graphics processing unit to accelerate the simulation process. Leveraging our simulation paradigm, the effectiveness of common policy interventions can be quantitatively evaluated. For mobility interventions, we find that lingering control is the most effective mobility intervention with 32.35% fewer infections, while increasing social distance and increasing walking speed have a similar effect with 15.15% and 18.02% fewer infections. It demonstrates the importance of introducing crowd mobility into disease transmission simulation. For transmission processes, we find the aerosol transmission involves in 99.99% of transmission, which highlights the importance of ventilation in indoor public spaces. Our simulation also demonstrates that without strict entrance detection to identify the input infections, only performing frequent disinfection cannot achieve desirable epidemic outcomes. Based on our simulation paradigm, we can shed light on better policy designs that achieve a good balance between disease spreading control and social costs.
Zhenyu Han, Siran Ma, Changzheng Gao, Erzhuo Shao, Yulai Xie 0001, Yang Zhang 0102, Lu Geng, Yong Li 0008
ACM Trans. Intell. Syst. Technol.6
2023 Interior Individual Trajectory Simulation with Population Distribution Constraint
abstract
Individual trajectory generation plays an important role in simulation tasks, reconstructing fine-grained mobility behaviors that can be used to evaluate epidemic risks, congestion risks, or commercial profit. Previous research works adopt the Newton’s mechanic-based particle model as their core algorithm, such as the Social Force model. However, real-world human mobility behaviors hardly follow the particle models, especially in the interior scenes where interactions between pedestrians and environments matter. In this article, we propose a Social Force-based trajectory simulator for interior scenarios that improve both trajectory quality and generation speed for interior scenarios. First, we introduce prior scene knowledge to guide the generation process, where pedestrians are armed with exploration behaviors that follow the group-level distribution. It provides more flexibility to simulate complicated human behaviors rather than straight-line movements, generating high-quality individual trajectories. Experiments show that the correlation between the aggregated population distribution of generated trajectories and ground-truth distribution is improved by 11.84% by our method. Second, we optimize the algorithm procedure by introducing a caching mechanism for tenderized intermediate values, along with graph-processing-unit-based implementation. Compared with the baseline Social Force model, we reduced the time consumption by 95%. More importantly, based on our simulation paradigm, we quantitatively evaluate several common mobility interventions in our simulation scenario, which can shed light on better policy designs in public spaces.
Erzhuo Shao, Zhenyu Han, Yulai Xie 0001, Yang Zhang 0102, Lu Geng, Yong Li 0008
ACM Trans. Intell. Syst. Technol.4
2022 Temporal-enhanced graph convolution network for skeleton-based action recognition
abstract
Abstract Graph convolution networks (GCNs) have drawn attention for skeleton‐based action recognition. They have achieved remarkable performance by adaptively learning spatial features of human action dynamics. However, the existing methods are limited in temporal sequence modelling of human actions. To give adequate consideration to temporal factors in action modelling, a novel temporal‐enhanced graph convolution network is presented. First, a Causal Convolution layer is introduced to ensure no future information leakage at each time step for keeping ordering information of inputs. Second, a novel cross‐spatial‐temporal graph convolution layer that extends an adaptive graph from the spatial to the temporal domain to capture local cross‐spatial‐temporal dependencies among joints is presented. Third, a temporal attention layer is designed to enhance the modelling capability of long‐range temporal dependencies, helping the network to directly focus on important time steps. Experimental results on three large‐scale datasets, NTU‐RGB + D, Kinetics‐Skeleton, and UAV‐Human, indicate that the authors’ network achieves accuracy improvement with better generalisation capability over previous methods. The authors’ code and data are available at https://github.com/xieyulai/TE‐GCN .
Yulai Xie 0001, Yang Zhang 0102, Fang Ren 0001
IET Comput. Vis.2
2022 Tri-Modal Dense Video Captioning Based on Fine-Grained Aligned Text and Anchor-Free Event Proposals Generator
abstract
Multi-modal dense video captioning is a task using multiple information to detect all meaningful events and generate a textual description for each event. The existing works mainly rely on single visual or dual audio-visual modals in dense video captioning, while completely ignoring the text modal (subtitle). The text modal has a similar data structure as the video captions, which provides immediate semantic information to the content description for a video. In this paper, we propose a novel framework, called Two-Stage Cross-Modal Encoding Transformer Network (TS-CMETN), to realize the multi-modal dense video captioning task by fusing multiple features, including audio, visual, and text. First, we design a two-stage feature fusion encoder that hierarchically achieves the intra- and inter-modal information interaction. Second, we propose an anchor-free temporal event proposal module, which efficiently generates event proposals at each time step without the complex anchor calculation. Extensive experiments on the ActivityNet Captions dataset show that our proposed framework achieves high performance. Moreover, our approach can adaptively handle cases of the missing text modal. Our code and data are available at https://github.com/xieyulai/TM-CMETN .
Jingjing Niu, Yulai Xie 0001, Yang Zhang 0102, Xiao Lei, Fang Ren 0001
Int. J. Pattern Recognit. Artif. Intell.3
2022 Multisize Patched Spatial-Temporal Transformer Network for Short- and Long-Term Crowd Flow Prediction
abstract
The prediction of urban crowds is crucial not only to traffic management but also to studies on the city-level social phenomena, such as energy consumption, urban growth, city planning, and epidemic prevention. The challenges of accurately predicting crowd flow come from the non-linear spatial-temporal dependence of crowd flow data, periodic laws, such as daily and weekly periodicity, and external factors, such as weather and holidays. It is even more challenging for most existing short-term prediction models to make an accurate long-term prediction. In this paper, we propose a novel patched Transformer-based sequence-to-sequence model, called MultiSize Patched Spatial-Temporal Transformer Network (MSP-STTN), to incorporate rich and unified context modeling via a self-attention mechanism and global memory learning via a cross-attention mechanism for short- and long-term grid-based crowd flow prediction. In particular, a multisize patched spatial-temporal self-attention Transformer is designed to capture cross-space-time and cross-size contextual dependence of crowd data. The same structured cross-attention Transformer is developed to adaptively learn a global memory for long-term prediction in a responding-to-a-query style without error accumulation. In addition, a categorized space-time expectation is proposed as a unified regional encoding with temporal and external factors and is used as a base prediction for stable training. Furthermore, auxiliary tasks are introduced for promoting feature encoding and leveraging feature consistency to assist in the main prediction task. The experimental results reveal that MSP-STTN is competitive with the state of the art for one-step and multi-step short-term prediction within several hours and achieves practical long-term crowd flow prediction within one day on real-world grid-based crowd data sets TaxiBJ, BikeNYC, and CrowdDensityBJ. Our code and data are available athttps://github.com/xieyulai/MSP-STTN.
Yulai Xie 0001, Jingjing Niu, Yang Zhang 0102, Fang Ren 0001
IEEE Trans. Intell. Transp. Syst.3
2020 AntMan: Dynamic Scaling on GPU Clusters for Deep Learning
Wencong Xiao, Shiru Ren, Yong Li 0045, Yang Zhang 0102, Pengyang Hou, Yihui Feng, Wei Lin 0016, Yangqing Jia
OSDI4