EDBT 2026 Demo / reviewers in the wild / expert
Zhouyang Li
dblp:32/10213
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploiting ARMeD Channels By Reverse Engineering ARM Memory Disambiguation UnitabstractARM CPUs are widely used in both embedded systems and personal computers where security considerations are becoming important. Evidently, vulnerabilities on hardware components such as cache and translation look-aside buffer are well-documented. But there are much less studies on other components, especially those in the CPU backend, largely due to the unavailability of their design and implementation details. To address this gap, we present the first in-depth reverse engineering analysis of the Memory Disambiguation Unit (MDU) in the backend of ARM CPUs. Across four microarchitectures from ARM and Apple CPUs, we identify two different MDU designs, switch-based and counter-based. We then analyze the state machine, selection mechanism, and organization of these MDU designs. We further propose new side channels and covert channels, which we call ARMeD channels, that exploit ARM MDU to leak information. We demonstrate with three attacks using ARMeD channels: a cross-process covert channel, website fingerprinting, and a new implementation of the Spectre attack. Finally, we present a defense strategy against ARMeD Channels with less than 3% degradation on the MDU’s prediction accuracy. Chang Liu 0117, Zhouyang Li, Haixia Wang 0001, Pengfei Qiu, Gang Qu 0001, Dongsheng Wang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM TrainingabstractPipeline Parallelism serves as a crucial technique for training Large Language Models, as it alleviates memory pressure from model states with relatively low communication overhead. However, in long-context scenarios, existing pipeline parallelism methods fail to address the substantial activation memory pressure, primarily due to the peak memory consumption resulting from the accumulation of activations across multiple microbatches. Moreover, these approaches inevitably introduce considerable pipeline bubbles, further hindering efficiency. Zhouyang Li, Tailing Yuan, Chengru Song |
SC | 1 |
| 2023 | Towards Complex Scenarios: Building End-to-End Task-Oriented Dialogue System across Multiple Knowledge BasesabstractWith the success of the sequence-to-sequence model, end-to-end task-oriented dialogue systems (EToDs) have obtained remarkable progress. However, most existing EToDs are limited to single KB settings where dialogues can be supported by a single KB, which is still far from satisfying the requirements of some complex applications (multi-KBs setting). In this work, we first empirically show that the existing single-KB EToDs fail to work on multi-KB settings that require models to reason across various KBs. To solve this issue, we take the first step to consider the multi-KBs scenario in EToDs and introduce a KB-over-KB Heterogeneous Graph Attention Network (KoK-HAN) to facilitate model to reason over multiple KBs. The core module is a triple-connection graph interaction layer that can model different granularity levels of interaction information across different KBs (i.e., intra-KB connection, inter-KB connection and dialogue-KB connection). Experimental results confirm the superiority of our model for multiple KBs reasoning. Libo Qin 0001, Zhouyang Li, Qiying Yu, Lehan Wang, Wanxiang Che |
AAAI | 2 |
| 2022 | EasyView: Enabling and Scheduling Tensor Views in Deep Learning CompilersabstractIn recent years, memory-intensive operations are becoming dominant in efficiency of running novel neural networks. Just-in-time operator fusion on accelerating devices like GPU proves an effective method for optimizing memory-intensive operations, and suits the numerous varying model structures. In particular, we find memory-intensive operations on tensor views are ubiquitous in neural network implementations. Tensors are the de facto representation for numerical data in deep learning areas, while tensor views cover a bunch of sophisticated syntax, which allow various interpretations on the underlying tensor data without memory copy. The support of views in deep learning compilers could greatly enlarge operator fusion scope, and appeal to optimizing novel neural networks. Nevertheless, mainstream solutions in state-of-the-art deep learning compilers exhibit imperfections either in view syntax representations or operator fusion. In this article, we propose EasyView, which enables and schedules tensor views in an end-to-end workflow from neural networks onto devices. Aiming at maximizing memory utilization and reducing data movement, we categorize various view contexts in high-level language, and lower views in accordance with different scenarios. Reference-semantic in terms of views are kept in the lowering from native high-level language features to intermediate representations. Based on the reserved reference-semantics, memory activities related to data dependence of read and write are tracked for further compute and memory optimization. Besides, ample operator fusion is applied to memory-intensive operations with views. In our tests, the proposed work could get average 5.63X, 2.44X, and 4.67X speedup compared with the XLA, JAX, and TorchScript, respectively for hotspot Python functions. In addition, operation fusion with views could bring 8.02% performance improvement in end-to-end neural networks. Lijuan Jiang, Qianchao Zhu, Shengen Yan, Xingcheng Zhang, Dahua Lin, Wenjing Ma, Zhouyang Li, Minxi Jin, Chao Yang 0002 |
ICPP | 9 |
| 2021 | Co-GAT: A Co-Interactive Graph Attention Network for Joint Dialog Act Recognition and Sentiment ClassificationabstractIn a dialog system, dialog act recognition and sentiment classification are two correlative tasks to capture speakers’ intentions, where dialog act and sentiment can indicate the explicit and the implicit intentions separately. The dialog context information (contextual information) and the mutual interaction information are two key factors that contribute to the two related tasks. Unfortunately, none of the existing approaches consider the two important sources of information simultaneously. In this paper, we propose a Co-Interactive Graph Attention Network (Co-GAT) to jointly perform the two tasks. The core module is a proposed co-interactive graph interaction layer where a cross-utterances connection and a cross-tasks connection are constructed and iteratively updated with each other, achieving to consider the two types of information simultaneously. Experimental results on two public datasets show that our model successfully captures the two sources of information and achieve the state-of-the-art performance. In addition, we find that the contributions from the contextual and mutual interaction information do not fully overlap with contextualized word representations (BERT, Roberta, XLNet). Libo Qin 0001, Zhouyang Li, Wanxiang Che, Minheng Ni, Ting Liu 0001 |
AAAI | 2 |
| 2019 | An Optimized Single-Stage isolated Phase-Shifted Full-Bridge Based Swiss-rectifierabstractSwiss-rectifier (SR) is a family of high efficient AC-DC topologies. It's a combination of a three phase unfolder circuit which can convert the three-phase AC voltage into two sets of cyclically varying DC voltages and two symmetric DC-DC converters to achieve power factor correction (PFC). In order to obtain a zero voltage switching (ZVS) isolated single-stage rectifier system, two phase-shifted full-bridge topologies are applied to increase the working frequency of the magnetic elements. However, the current flows through at least four switches during the power conduction or freewheeling period. It is not conducive to improve the efficiency of full-bridge based SR. Thus, an optimized single-stage isolated phase-shifted full-bridge (PSFB) based Swiss-rectifier is proposed. Only one full bridge structure is used to reduce the conduction loss when the power is transmitting. Two voltage clamp branches are used to implement the PFC principle of the SWISS-rectifier. The principle of operation and the analytical equations for the stress in the semiconductor devices are discussed specifically. The performances of the proposed converter are verified by the experiments on a 10-kW 380Vac input/400Vdc output prototype with 90kHz switching frequency. Binfeng Zhang, Shaojun Xie, Jinming Xu 0001, Kunshan Xu, Zhouyang Li |
IECON | 5 |