EDBT 2026 Demo / reviewers in the wild / expert
Xiangyu Ye
dblp:00/1174
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHM: Context Hiding and Misguidance for Robust Adversarial Attacks on Active Speaker DetectionabstractCurrent Active Speaker Detection (ASD) models have achieved remarkable success, surpassing 90% mAP on the large-scale AVA-ActiveSpeaker benchmark and approaching human-level accuracy. However, our diagnostic analysis reveals that this performance relies on a fragile shortcut: models exhibit a severe "Visual Consistency Bias," prioritizing the association between audio and visual feature consistency over rigorous dynamic temporal synchronization. Consequently, disrupting this visual consistency causes the model to lose track of the active speaker. To exploit this, we propose Context Hiding & Misguidance (CHM), a novel context-aware adversarial attack. Unlike existing methods, CHM strategically manipulates visual feature consistency. Specifically, we design a geometric "Repulsion-Attraction" objective: the Hiding term suppresses the target’s consistent visual features (severing the correct binding), while the Misguidance term steers these features toward the manifold of a silent bystander (creating a false binding). Extensive experiments on the AVA-ActiveSpeaker and UniTalk benchmarks demonstrate the superiority of CHM. Crucially, CHM exhibits remarkable cross-architecture transferability: adversarial perturbations generated on lightweight single-candidate surrogates (e.g., LRASD) successfully compromise sophisticated multi-candidate models like LoCoNet. Xiangyu Ye, Yatie Xiao, Qingxiao Guan, Zhenbang Liu |
ICMR | 1 |
| 2026 | TFPA: Enhancing adversarial attack on speech recognition via Time-Frequency Pre-alignment
Xiangyu Ye, Yatie Xiao, Kongyang Chen, Qingxiao Guan, Zhenbang Liu |
Knowl. Based Syst. | 1 |
| 2026 | Harnessing Transferable Adversarial Examples via Multilayer Attention-Guided Spatial TransformationsabstractTransfer-based adversarial attacks are key for evaluating the robustness of deep neural networks (DNNs) in black-box settings, yet their effectiveness is often constrained by limited cross-model transferability. Existing feature-level approaches typically rely on single-layer attention guidance or static perturbation patterns, which restrict adaptability across diverse architectures. In this work, we introduce a unified adversarial framework, named Multi-layer Attention-guided Spatial Transformations (MAT), to exploit class-discriminative cues from multiple feature layers to craft highly transferable adversarial examples. MAT integrates Multi-layer Attention Fusion (MAF) to capture complementary low-level and high-level semantics from multiple intermediate layers, Attention-guided Augmentation (AGA) to selectively perturb non-critical regions while preserving semantic integrity, and Spatial Random Transformation (SRT) to introduce stochastic spatial augmentations to diversify patterns during optimization. Unlike prior methods that use static or layer-specific attention, MAT dynamically adapts feature guidance to the architecture and task, which enhances generalization. We evaluate MAT against eleven state-of-the-art (SOTA) transfer-based attacks across nine CNN-based and Transformer-based architectures on ImageNet. Comprehensive experiments demonstrate that MAT consistently outperforms eleven state-of-the-art transfer-based attacks in both white-box and black-box settings, including against adversarially trained and input preprocessing-based defensive models, while maintaining higher semantic similarity to the original inputs. It highlights the superior adversarial robustness and excellent adaptability of MAT in adversarial machine learning. Our code is available athttps://github.com/dislab-gzhu/MAT. Pengfei Dong, Yatie Xiao, Chi-Man Pun, Fei Peng 0001, Kongyang Chen, Qingxian Guan, Siyuan Chen 0005, Xiangyu Ye, Zhenbang Liu |
IEEE Trans. Reliab. | 8 |
| 2024 | A Memory-Efficient Hybrid Parallel Framework for Deep Neural Network TrainingabstractWith the increasing volumes of data samples and deep neural network (DNN) models, efficiently scaling the training of DNN models has become a significant challenge for server clusters with AI accelerators in terms of memory and computing efficiency. Existing parallelism schemes can be broadly classified into three categories: data parallelism (splitting data samples), model parallelism (splitting model parameters), and pipeline model parallelism (splitting model layers). Hybrid approaches split data and models, offering a comprehensive solution for parallel training. However, these methods encounter limitations in efficiently scaling larger models across more computing nodes, as they incur substantial memory constraints that affect training efficiency and overall throughput. In this paper, we proposeHIPPIE, a hybrid parallel training framework designed to enhance memory efficiency and scalability of large DNN training. First, to evaluate the optimization effect more reasonably, we propose an index ofMemory Efficiency(ME) to quantify the tradeoff between throughput and memory overhead. Second, driven by the informed ME optimization objective, we automatically partition the pipeline to balance the throughput and memory. Third, we optimize the model training process via a novel hybrid parallel scheduler that improves the throughput and scalability by informed pipeline scheduling and communication scheduling with gradient-hidden optimization. Experiments on various models show thatHIPPIEachieves above 90% scaling efficiency on a 16-GPU platform. Moreover,HIPPIEincreases throughput by up to 80%, while saving 57% of memory overhead and achieving 4.18× memory-efficiency improvement. Dongsheng Li 0001, Zhiquan Lai, Yongquan Fu, Xiangyu Ye, Linbo Qiao |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | FNNG: A High-Performance FPGA-based Accelerator for K-Nearest Neighbor Graph ConstructionabstractThe k-nearest neighbor graph has emerged as the key data structure for many critical applications. However, it can be notoriously challenging to construct k-nearest neighbor graphs over large graph datasets, especially with a high-dimensional vector feature. Many solutions have been recently proposed to support the construction of k-nearest neighbor graphs. However, these solutions involve substantial memory access and computational overheads and an architecture-level solution is still absent. To address these issues, we architect FNNG, the first FPGA-based accelerator to support k-nearest neighbor graph construction. Specifically, FNNG is equipped with the block-based scheduling technique to exploit the inherent data locality between vertices. It divides the vertices that are close in space into blocks and process the vertices according to the granularity of the blocks during the construction process. FNNG also adopts the useless computation aborting technique to identify superfluous useless computations. It keeps the existing maximum similarity values of all vertices inside the computing unit. In addition, we propose an improved architecture in order to fully utilize both techniques. We implement FNNG on the Xilinx Alveo U280 FPGA card. The results show that FNNG achieves 190x and 2.1x speedups over the state-of-the-art CPU and GPU solutions, running on Intel Xeon Gold 5117 CPU and NVIDIA GeForce RTX 3090 GPU, respectively. Chaoqiang Liu, Haifeng Liu 0003, Long Zheng 0003, Yu Huang 0013, Xiangyu Ye, Xiaofei Liao, Hai Jin 0001 |
FPGA | 5 |
| 2023 | AFaVS: Accurate Yet Fast Version Switching for Graph Processing SystemsabstractMulti-version graph processing has been widely used to solve many real-world problems. The process of the multi-version graph processing typically includes: (1) a history graph version switching at a specific time and (2) graph processing on this history graph. Existing multi-version graph systems assume ideally that every request for a particular graph version at a particular time will have a corresponding snapshot available. However, in most cases, this is not true. Then existing solutions usually have to settle with an "approximating" version as a substitute, leading to unexpected results for the underlying graph algorithm and thus reducing the practicality of a multi-version graph system for many application scenarios significantly.In this paper, we observe that only a few graph updates have a great impact on the final results. We therefore present AFaVS, a novel multi-version graph system that can improve accuracy effectively in both time- and memory-efficient manners. The cornerstone of AFaVS lies in a novel concept "value" that characterizes the importance of graph updates. AFaVS proposes differential management of updates based on their values and achieves higher accuracy while preserving processing and memory efficiency. AFaVS is also equipped with value-guided version switching and locality-aware optimizations to boost its overall efficiency. Our results on a variety of real-world datasets show that AFaVS outperforms four state-of-the-art multi-version graph systems by 74.35%~95.72% in terms of accuracy improvement and 57.03%~90.44% in terms of memory reduction while introducing less than 2.96% extra computing time. We have deployed AFaVS in a disaster recovery system on the production cluster of Alibaba, achieving 78.8%~90.1% fewer error rates than advanced systems at a comparable efficiency. Long Zheng 0003, Xiangyu Ye, Haifeng Liu 0003, Qinggang Wang, Yu Huang 0013, Chuangyi Gui, Pengcheng Yao, Xiaofei Liao, Hai Jin 0001, Jingling Xue |
ICDE | 2 |
| 2023 | Accelerating Personalized Recommendation with Cross-level Near-Memory ProcessingabstractThe memory-intensive embedding layers of the personalized recommendation systems are the performance bottleneck as they demand large memory bandwidth and exhibit irregular and sparse memory access patterns. Recent studies propose near memory processing (NMP) to accelerate memory-bound embedding operations. However, due to the load imbalance caused by the skewed access frequency of the embedding data, existing NMP solutions that exploit fine-grained memory parallelism fail to translate the increasingly massive internal bandwidth to performance improvements, leading to resource underutilization and hardware overhead. Haifeng Liu 0003, Long Zheng 0003, Yu Huang 0013, Chaoqiang Liu, Xiangyu Ye, Jingrui Yuan, Xiaofei Liao, Hai Jin 0001, Jingling Xue |
ISCA | 5 |
| 2022 | EmbRace: Accelerating Sparse Communication for Distributed Training of Deep Neural NetworksabstractDistributed data-parallel training has been widely adopted for deep neural network (DNN) models. Although current deep learning (DL) frameworks scale well for dense models like image classification models, we find that these DL frameworks have relatively low scalability for sparse models like natural language processing (NLP) models that have highly sparse embedding tables. Most existing works overlook the sparsity of model parameters thus suffering from significant but unnecessary communication overhead. In this paper, we propose EmbRace, an efficient communication framework to accelerate communications of distributed training for sparse models. EmbRace introduces Sparsity-aware Hybrid Communication, which integrates AlltoAll and model parallelism into data-parallel training, so as to reduce the communication overhead of highly sparse parameters. To effectively overlap sparse communication with both backward and forward computation, EmbRace further designs a 2D Communication Scheduling approach which optimizes the model computation procedure, relaxes the dependency of embeddings, and schedules the sparse communications of each embedding row with a priority queue. We have implemented a prototype of EmbRace based on PyTorch and Horovod, and conducted comprehensive evaluations with four representative NLP models. Experimental results show that EmbRace achieves up to 2.41 × speedup compared to the state-of-the-art distributed training baselines. Zhiquan Lai, Dongsheng Li 0001, Yiming Zhang 0003, Xiangyu Ye, Yabo Duan |
ICPP | 5 |
| 2021 | Hippie: A Data-Paralleled Pipeline Approach to Improve Memory-Efficiency and Scalability for Large DNN TrainingabstractWith the increase of both data and parameter volume, it has become a big challenge to efficiently train large-scale DNN models on distributed platforms. Ordinary parallelism modes, i.e., data parallelism, model parallelism and pipeline parallelism, can no longer satisfy the efficient scaling of large DNN model training on multiple nodes. Meanwhile, the problem of too much memory consumption seriously restricts GPU computing efficiency and training throughput. In this paper, we propose Hippie, a hybrid parallel training framework that integrates pipeline parallelism and data parallelism to improve the memory efficiency and scalability of large DNN training. Hippie adopts a hybrid parallel method based on hiding gradient communication, which improves the throughput and scalability of training. Meanwhile, Hippie introduces the last-stage pipeline scheduling and recomputation for specific layers to effectively reduce the memory overhead and ease the difficulties of training large DNN models on memory-constrained devices. To achieve a more reasonable evaluation of the optimization effect, we propose an index of memory efficiency (ME) to represent the tradeoff between throughput and memory overhead. We implement Hippie based on PyTorch and NCCL. Experiments on various models show that Hippie achieves above 90% scaling efficiency on a 16-GPU platform. Moreover, Hippie increases throughput by up to 80% while saving 57% of memory overhead, achieving 4.18 × memory efficiency. Xiangyu Ye, Zhiquan Lai, Ding Sun, Linbo Qiao, Dongsheng Li 0001 |
ICPP | 1 |
| 2021 | Improved Partitioning Graph Embedding Framework for Small Cluster
Ding Sun, Zhen Huang 0006, Dongsheng Li 0001, Xiangyu Ye, Yilin Wang 0008 |
KSEM | 4 |