Yang Yao 0003

dblp:15/10186-3 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0003-8733-7721ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 56% Graph learning · 34% Efficient and distributed learning · 7%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 50% Electronic design automation · 50%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 59% Bioinformatics and computational biology · 41%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
distribution shift
1.622025
DyNAS-DDI: Dynamic Pairwise Architecture Search for Generalizable Drug-Drug Interaction LLM · ACM Multimedia 2025
Data-Augmented Curriculum Graph Neural Architecture Search under Distribution Shifts · AAAI 2024
Machine learning › Graph learning › graph neural network
graph neural architecture search
1.622025
DyNAS-DDI: Dynamic Pairwise Architecture Search for Generalizable Drug-Drug Interaction LLM · ACM Multimedia 2025
Data-Augmented Curriculum Graph Neural Architecture Search under Distribution Shifts · AAAI 2024
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
1.012026
Disentangled Graph LLM for Molecule Graph Editing under Distribution Shifts · WWW 2026
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.912025
DyNAS-DDI: Dynamic Pairwise Architecture Search for Generalizable Drug-Drug Interaction LLM · ACM Multimedia 2025
Medical and health informatics › drug safety
drug-drug interaction prediction
0.912025
DyNAS-DDI: Dynamic Pairwise Architecture Search for Generalizable Drug-Drug Interaction LLM · ACM Multimedia 2025
Electronic design automation
hardware/software co-design
0.912025
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration · AAAI 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural network accelerator design
0.912025
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration · AAAI 2025
Machine learning › Trustworthy machine learning
robustness
0.812024
Data-Augmented Curriculum Graph Neural Architecture Search under Distribution Shifts · AAAI 2024
Bioinformatics and computational biology
drug discovery
0.312026
Disentangled Graph LLM for Molecule Graph Editing under Distribution Shifts · WWW 2026
Bioinformatics and computational biology › drug discovery
molecular optimization
0.312026
Disentangled Graph LLM for Molecule Graph Editing under Distribution Shifts · WWW 2026
Machine learning › Efficient and distributed learning › model compression › quantization
low-bit quantization
0.312025
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration · AAAI 2025
Machine learning › Efficient and distributed learning
model compression
0.312025
JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration · AAAI 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.312025
DyNAS-DDI: Dynamic Pairwise Architecture Search for Generalizable Drug-Drug Interaction LLM · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

large language model · 3.7invariance loss · 2.0disentangled graph projector · 2.0retrieval-augmented instruction tuning · 1.7neural architecture search · 1.7hardware generation network · 1.7graph neural network · 1.7compiler mapping search · 1.7channel-wise sparse quantization · 1.7LoRA mixture-of-experts · 1.0LoRA mixture of experts · 1.0self-supervised learning · 0.9
YearPublicationVenuePosition
2026 Disentangled Graph LLM for Molecule Graph Editing under Distribution Shifts
abstract
Molecule graph editing has become a powerful paradigm for optimizing chemical compounds in drug discovery. Existing methods overlook the invariant structure-property relationships, and rely on variable correlations that shift across different instructions, thereby failing to generalize to out-of-distribution (O.O.D.) scenarios. To overcome the weakness of existing work, in this paper we propose to capture and utilize the invariant factors in order to achieve generalizable molecule graph editing under distribution shifts. However, this problem remains challenging, given that the invariant and variant factors are deeply entangled within the editing models. To tackle this challenge, we propose MoFE, a disentangled graph large language model for molecule graph editing that handles editing instructions under distribution shifts via disentangling invariant factors that govern editing-relevant properties. Specifically, we propose a disentangled graph projector with invariance loss that encodes molecular graphs into disentangled latent factors, with an invariance loss that ensures consistency across paraphrased prompts with the same objective. Then, we enhance the LLM with a factor-aware LoRA mixture-of-experts, where each expert is associated with a distinct latent factor. Additionally, we introduce a factor disentanglement loss weighting strategy that adaptively assigns higher weights to expert-factor pairs that perform well on relevant editing tasks. The proposed MoFE model promotes joint disentanglement between experts and latent factors, reinforcing their alignment and preventing collapse. Experiments on a representative benchmark demonstrate that MoFE is able to achieve superior O.O.D. generalization performance in molecule graph editing.
Yang Yao 0003, Xin Wang 0019, Zeyang Zhang 0001, Hong Mei 0001, Wenwu Zhu 0001
WWW1
2025 JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration
abstract
The co-design of neural network architectures, quantization precisions, and hardware accelerators offers a promising approach to achieving an optimal balance between performance and efficiency, particularly for model deployment on resource-constrained edge devices. In this work, we propose the JAQ Framework, which jointly optimizes the three critical dimensions. However, effectively automating the design process across the vast search space of those three dimensions poses significant challenges, especially when pursuing extremely low-bit quantization. Specifical, the primary challenges include: (1) Memory overhead in software-side: Low-precision quantization-aware training can lead to significant memory usage due to storing large intermediate features and latent weights for backpropagation, potentially causing memory exhaustion. (2) Search time-consuming in hardware-side: The discrete nature of hardware parameters and the complex interplay between compiler optimizations and individual operators make the accelerator search time-consuming. To address these issues, JAQ mitigates the memory overhead through a channel-wise sparse quantization (CSQ) scheme, selectively applying quantization to the most sensitive components of the model during optimization. Additionally, JAQ designs BatchTile, which employs a hardware generation network to encode all possible tiling modes, thereby speeding up the search for the optimal compiler mapping strategy. Extensive experiments demonstrate the effectiveness of JAQ, achieving approximately 7% higher Top-1 accuracy on ImageNet compared to previous methods and reducing the hardware search time per iteration to 0.15 seconds.
Mingzi Wang, Weixiang Zhang, Yijian Qin, Yang Yao 0003, Yingxin Li, Tongtong Feng, Xin Wang 0019, Xun Guan, Zhi Wang 0001, Wenwu Zhu 0001
AAAI6
2025 DyNAS-DDI: Dynamic Pairwise Architecture Search for Generalizable Drug-Drug Interaction LLM
abstract
Drug-drug interaction (DDI) prediction is a pivotal task in biomedical research. Emerging multimodal approaches that integrate graph neural networks (GNNs) and large language models (LLMs) have gained traction, as GNNs capture molecular structures while LLMs provide a rich biomedical context. However, real-world DDI data often exhibit distribution shifts across structural and textual dimensions, stemming from variations in molecular scaffolds, drug sizes, and assay conditions. Existing methods assume an independent and identically distributed (I.I.D.) setting, failing to handle such shifts primarily due to there key limitations: (i) the entanglement of core interaction motifs with incidental structural features; (ii) inflexible message-passing GNN architectures ill-suited for diverse drug pairs; and (iii) underutilized biomedical knowledge in LLMs for capturing pairwise interaction semantics. These limitations highlight the need for a disentangled, dynamic, and pairwise-aware modeling strategy to achieve out-of-distribution generalized DDI prediction. To solve this problem, we propose DyNamic Pairwise Architecture Search for Generalizable Drug-Drug Interaction LLM (DyNAS-DDI), a novel framework that dynamically adapts network architectures for each molecular pair and integrates biomedical knowledge from LLMs to improve generalization under distribution shifts. Specifically, we propose three modules: (i) Motif-driven disentangled molecule encoding, which disentangles molecular representations into distinct motif-based features while preserving key structural signals through a self-supervised graph encoder; (ii) Attentionbased pairwise neural architecture search, where multi-head attention enriches molecular features to guide a dynamic search mechanism that adaptively optimizes message passing for diverse interaction types; and (iii) retrieval-augmented molecular instruction tuning, where external biomedical knowledge is incorporated to improve interpretability and enable reasoning for unseen drug interactions. Extensive experiments on four datasets for DDI with out-of-distribution (OOD) splits demonstrate our method's superior generalization abilities under distribution shifts. Our code can be available at https://github.com/EkkoXiao/DyNAS-DDI.
Linxin Xiao, Xin Wang 0019, Zeyang Zhang 0001, Yang Yao 0003, Wenwu Zhu 0001
ACM Multimedia4
2024 Data-Augmented Curriculum Graph Neural Architecture Search under Distribution Shifts
abstract
Graph neural architecture search (NAS) has achieved great success in designing architectures for graph data processing.However, distribution shifts pose great challenges for graph NAS, since the optimal searched architectures for the training graph data may fail to generalize to the unseen test graph data. The sole prior work tackles this problem by customizing architectures for each graph instance through learning graph structural information, but failed to consider data augmentation during training, which has been proven by existing works to be able to improve generalization.In this paper, we propose Data-augmented Curriculum Graph Neural Architecture Search (DCGAS), which learns an architecture customizer with good generalizability to data under distribution shifts. Specifically, we design an embedding-guided data generator, which can generate sufficient graphs for training to help the model better capture graph structural information. In addition, we design a two-factor uncertainty-based curriculum weighting strategy, which can evaluate the importance of data in enabling the model to learn key information in real-world distribution and reweight them during training. Experimental results on synthetic datasets and real datasets with distribution shifts demonstrate that our proposed method learns generalizable mappings and outperforms existing methods.
Yang Yao 0003, Xin Wang 0019, Yijian Qin, Ziwei Zhang 0001, Wenwu Zhu 0001, Hong Mei 0001
AAAI1
2024 Customized Cross-device Neural Architecture Search with Images
abstract
Cross-device scenarios have become increasingly common, where non-independently and identically distributed (non-IID) data is generated and stored in different devices. However, the existing cross-device NAS methods only search for a fixed architecture for different devices, neglecting that different devices have varying hardware characteristics and data distributions. In this paper, we propose a novel NAS framework that can customize the most suitable architecture for each device and its associated dataset. Specifically, we propose a decoupled data feature extractor and a device feature extractor to characterize the complex distributions of the different datasets and diverse hardware features. Then, we propose a prototype matcher to customize the operators and shape selection parameters of architectures. Experiments on ImageNet and CIFAR-10 show that our method can discover more efficient and effective architectures in cross-device scenarios than the existing approaches. To the best of our knowledge, this is the first exploration on customized cross-device NAS problem.
Yang Yao 0003, Xin Wang 0019, Yijian Qin, Ziwei Zhang 0001, Wenwu Zhu 0001, Hong Mei 0001
ICME1