EDBT 2026 Demo / reviewers in the wild / expert
Yue Dai 0005
dblp:03/4755-5
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-4436-0991ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cascade: A Dependency-aware Efficient Training Framework for Temporal Graph Neural NetworkabstractTemporal graph neural networks (TGNN) have gained significant momentum in many real-world dynamic graph tasks. These models use graph changes (i.e., events) as inputs to update nodes' status vectors (i.e., memories), which are then exploited to assist predictions. Despite their improved accuracies, the efficiency of TGNN training is significantly limited due to the inherent temporal relationship between the input events. Although larger training batches can improve parallelism and speed up TGNN training, they lead to infrequent memory updates, which cause outdated information and reduced accuracy. This trade-off forces current methods to use small batches, resulting in high latency and underutilized hardware. To address this, we propose an efficient TGNN training framework, Cascade, to adaptively boost TGNN training parallelism based on nodes' spatial and temporal dependencies. Cascade adopts a topology-aware scheduler that includes as many spatial-independent events in the same batches. Moreover, it leverages node memories' similarities to break temporal dependencies on stabilized nodes, enabling it to pack more temporal-independent events in the same batches. Additionally, Cascade adaptively decides nodes' update frequencies based on runtime feedback. Compared to prior state-of-the-art TGNN training frameworks, our approach can averagely achieve 2.3x (up to 5.1x) speed up without jeopardizing the resulted models' accuracy. Yue Dai 0005, Xulong Tang, Youtao Zhang |
ASPLOS (2) | 1 |
| 2025 | Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised LearningabstractSelf-supervised learning (SSL) offers a compelling solution to the challenge of extensive labeled data requirements in traditional supervised learning.
With the proven success of Vision Transformers (ViTs) in supervised tasks, there is increasing interest in adapting them for SSL frameworks. However, the high computational demands of SSL pose substantial challenges, particularly on resource-limited platforms like edge devices, despite its ability to achieve high accuracy without labeled data.
Recent studies in supervised learning have shown that token pruning can reduce training costs by removing less informative tokens without compromising accuracy. However, SSL’s dual-branch encoders make traditional single-branch pruning strategies less effective, as they fail to account for the critical cross-branch similarity information, leading to reduced accuracy in SSL.
To this end, we introduce SimPrune, a novel token pruning strategy designed for ViTs in SSL. SimPrune leverages cross-branch similarity information to efficiently prune tokens, retaining essential semantic information across dual branches. Additionally, we incorporate a difficulty-aware pruning strategy to further enhance SimPrune's effectiveness.
Experimental results show that our proposed approach effectively reduces training computation while maintaining accuracy. Specifically, our approach offers 24\% savings in training costs compared to SSL baseline, without sacrificing accuracy. Sheng Li 0019, Qitao Tan, Yue Dai 0005, Zhenglun Kong, Jun Liu 0075, Ao Li 0004, Ninghao Liu 0001, Yufei Ding 0001, Xulong Tang, Geng Yuan |
ICLR | 3 |
| 2025 | MemFreezing: A Novel Adversarial Attack on Temporal Graph Neural Networks under Limited Future KnowledgeabstractTemporal graph neural networks (TGNN) have achieved significant momentum in many real-world dynamic graph tasks.
While most existing TGNN attack methods assume worst-case scenarios where attackers have complete knowledge of the input graph, the assumption may not always hold in real-world situations, where attackers can, at best, access information about existing nodes and edges but not future ones after the attack.
However, studying adversarial attacks under these constraints is crucial, as limited future knowledge can reveal TGNN vulnerabilities overlooked in idealized settings.
Nevertheless, designing effective attacks in such scenarios is challenging: the evolving graph can weaken their impact and make it hard to affect unseen nodes.
To address these challenges, we introduce MemFreezing, a novel adversarial attack framework that delivers long-lasting and spreading disruptions in TGNNs without requiring post-attack knowledge of the graph.
MemFreezing strategically injects fake nodes or edges to push node memories into a stable “frozen state,” reducing their responsiveness to subsequent graph changes and limiting their ability to convey meaningful information.
As the graph evolves, these affected nodes maintain and propagate their frozen state through their neighbors.
Experimental results show that MemFreezing persistently degrades TGNN performance across various tasks, offering a more enduring adversarial strategy under limited future knowledge. Yue Dai 0005, Xulong Tang, Youtao Zhang, Jun Yang 0002 |
ICML | 1 |
| 2025 | Reinforcement Learning-Guided Graph State Generation in Photonic Quantum ComputersabstractThe photonic quantum computer (PQC) is an emerging and promising quantum computing paradigm that has gained momentum in recent years.In PQC, computations are executed by performing measurements on photons in graph states (i.e., a collection of entangled photons).The graph state generation process is fulfilled by applying a sequence of quantum gates to quantum emitters, referred to as the "generation sequence".In a generation sequence, i) the time required to complete the generation sequence, ii) the number of quantum emitters used, and iii) the number of CZ gates performed between emitters greatly affect the fidelity of the generated graph state.In this paper, we propose RLGS (Reinforcement Learningguided Graph State generation), a novel compilation framework to identify optimal generation sequences that optimize the three fidelity metrics.Experimental results show that RLGS achieves an average reduction in generation time of 31.1%,49.6%, and 57.5% for small, medium, and large graph states compared to the baseline.Additionally, the reductions in the number of quantum emitters are 13.9%, 16.7%, and 17.5%, whereas the reductions in the number of CZ gates are 37.7%, 53.4%, and 57.8%, respectively. Yingheng Li, Yue Dai 0005, Aditya Pawar, Rongchao Dong, Jun Yang 0002, Youtao Zhang, Xulong Tang |
ISCA | 2 |
| 2023 | CEGMA: Coordinated Elastic Graph Matching Acceleration for Graph Matching NetworksabstractThe recently proposed Graph Matching Network models (GMNs) effectively improve the inference accuracy of graph similarity analysis tasks. GMNs often take graph pairs as input, embed nodes features, and match nodes between graphs for similarity analysis. While GMNs deliver high inference accuracy, the all-to-all node matching stage in GMNs introduces quadratic computing complexity with excessive memory accesses, resulting in significant computing and memory overhead that cannot be handled by existing approaches. In this paper, we propose the Coordinated Elastic Graph Matching Accelerator (CEGMA), a software and hardware co-design accelerator to address the challenges of GMNs. Specifically, by exploiting duplicate subgraphs in the input graphs, we develop an elastic matching filter to significantly reduce the quadratic computing overhead. By exploring the substantial data reuses oriented from accessing node features, we propose a cross-graph coordinator that fuses cross-graph similarity computing with intra-graph computing to enhance data locality. Experimental results show that, on average, CEGMA achieves 353× and 6.5× speedups in GMN computing compared to state-of-the-art GPU implementation and GNN accelerators, respectively. Yue Dai 0005, Youtao Zhang, Xulong Tang |
HPCA | 1 |
| 2023 | FlexGM: An Adaptive Runtime System to Accelerate Graph Matching Networks on GPUsabstractGMNs (Graph Matching Networks) exploit recently developed GNNs (Graph Neural Networks) to analyze the similarity between two graphs. They are increasingly deployed in many application domains due to their improved inference accuracy. A GMN consists of two stages, i.e., node-embedding and node-matching stages. The node-matching stage matches node features from two graphs for similarity, which accounts for over 90% of the total execution time. However, it is challenging to accelerate GMNs on GPUs due to their diverse computing patterns for different graph inputs. For large graphs, the overhead comes mainly from the high computation overhead, which increases quadratically to the size of the graphs; for small graphs, the overhead comes from the low parallelism and resource utilization.In this paper, we propose FlexGM, a flexible runtime, to adaptively accelerate GMNs on GPUs. For large graphs, we exploit the massive computation redundancy in GMNs and develop a low-overhead deduplication module to mitigate the high computation overhead. For small graphs, we develop a unified matching module to optimize GPU hardware resource usage. An adaptive module manager is then developed to judiciously select beneficial optimization strategies. Experimental results show that the FlexGM system achieves 2.5× (up to 7.6 ×) average speedup over existing methods. Yue Dai 0005, Xulong Tang, Youtao Zhang |
ICCD | 1 |
| 2023 | SmartFRZ: An Efficient Training Framework using Attention-Based Layer Freezing
Sheng Li 0019, Geng Yuan, Yue Dai 0005, Youtao Zhang, Yanzhi Wang 0001, Xulong Tang |
ICLR | 3 |
| 2022 | An efficient segmented quantization for graph neural networks
Yue Dai 0005, Xulong Tang, Youtao Zhang |
CCF Trans. High Perform. Comput. | 1 |