Yiwei Dai

dblp:224/3866 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Graph learning · 23% Multi-agent systems · 21% Trustworthy machine learning · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Smart cities and intelligent transportation · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
1.012026
BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks · ACL (1) 2026
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent
1.012026
MoEC: A Memory-Routed Mixture-of-Experts Controller for Adaptive Minecraft Control · ACL (1) 2026
Machine learning › Graph learning
graph neural network
1.012026
Dual Mamba for Node-Specific Representation Learning: Tackling Over-Smoothing with Selective State Space Modeling · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems
multi-agent security
1.012026
BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks · ACL (1) 2026
Machine learning › Graph learning › graph neural network › deep graph neural network
over-smoothing
1.012026
Dual Mamba for Node-Specific Representation Learning: Tackling Over-Smoothing with Selective State Space Modeling · AAAI 2026
Smart cities and intelligent transportation
traffic prediction
1.012026
HyperD: Hybrid Periodicity Decoupling Framework for Traffic Forecasting · AAAI 2026
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
communication topology
0.912025
Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems · EMNLP 2025
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.912025
Latte: Transfering LLMs' Latent-level Knowledge for Few-shot Tabular Learning · IJCAI 2025
Machine learning › Graph learning › graph diffusion
information diffusion
0.912025
Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems · EMNLP 2025
Machine learning › Transfer learning and domain adaptation › knowledge transfer
large language model knowledge transfer
0.912025
Latte: Transfering LLMs' Latent-level Knowledge for Few-shot Tabular Learning · IJCAI 2025
Machine learning › Trustworthy machine learning
fairness
0.812024
Mitigate Extrinsic Social Bias in Pre-trained Language Models via Continuous Prompts Adjustment · EMNLP 2024
Machine learning › Trustworthy machine learning › debiasing
language bias mitigation
0.812024
Mitigate Extrinsic Social Bias in Pre-trained Language Models via Continuous Prompts Adjustment · EMNLP 2024
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems
0.622026
BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks · ACL (1) 2026
Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems · EMNLP 2025
Machine learning › Deep learning architectures and training › attention mechanism › multi-dimensional attention
spatio-temporal attention
0.312026
HyperD: Hybrid Periodicity Decoupling Framework for Traffic Forecasting · AAAI 2026
Machine learning › Deep learning architectures and training
state space model
0.312026
Dual Mamba for Node-Specific Representation Learning: Tackling Over-Smoothing with Selective State Space Modeling · AAAI 2026
Machine learning and data management
tabular data learning
0.312025
Latte: Transfering LLMs' Latent-level Knowledge for Few-shot Tabular Learning · IJCAI 2025
Natural language and speech › Language models and text generation
pre-trained language model
0.212024
Mitigate Extrinsic Social Bias in Pre-trained Language Models via Continuous Prompts Adjustment · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

periodic embedding · 2.0frequency-domain MLP · 2.0dual-view alignment loss · 2.0selective state space modeling · 1.0residual connections · 1.0non-parametric memory · 1.0mixture of experts · 1.0mamba · 1.0failure-triggered expert growth · 1.0unsupervised pre-training · 0.9latent knowledge extraction · 0.9causal framework · 0.9
YearPublicationVenuePosition
2026 Dual Mamba for Node-Specific Representation Learning: Tackling Over-Smoothing with Selective State Space Modeling
abstract
Over-smoothing remains a fundamental challenge in deep Graph Neural Networks (GNNs), where repeated message passing causes node representations to become indistinguishable. While existing solutions, such as residual connections and skip layers, alleviate this issue to some extent, they fail to explicitly model how node representations evolve in a node-specific and progressive manner across layers. Moreover, these methods do not take global information into account, which is also crucial for mitigating the over-smoothing problem. To address the aforementioned issues, in this work, we propose a Dual Mamba-enhanced Graph Convolutional Network (DMbaGCN), which is a novel framework that integrates Mamba into GNNs to address over-smoothing from both local and global perspectives. DMbaGCN consists of two modules: the Local State-Evolution Mamba (LSEMba) for local neighborhood aggregation and utilizing Mamba’s selective state space modeling to capture node-specific representation dynamics across layers, and the Global Context-Aware Mamba (GCAMba) that leverages Mamba’s global attention capabilities to incorporate global context for each node. By combining these components, DMbaGCN enhances node discriminability in deep GNNs, thereby mitigating over-smoothing. Extensive experiments on multiple benchmarks demonstrate the effectiveness and efficiency of our method.
Xin He 0003, Yili Wang 0004, Yiwei Dai, Xin Wang 0035
AAAI3
2026 HyperD: Hybrid Periodicity Decoupling Framework for Traffic Forecasting
abstract
Accurate traffic forecasting plays a vital role in intelligent transportation systems, enabling applications such as congestion control, route planning, and urban mobility optimization. However, traffic forecasting remains challenging due to two key factors: (1) complex spatial dependencies arising from dynamic interactions between road segments and traffic sensors across the network, and (2) the coexistence of multi-scale periodic patterns (e.g., daily and weekly periodic patterns driven by human routines) with irregular fluctuations caused by unpredictable events (e.g., accidents, weather, or construction). To tackle these challenges, we propose HyperD (Hybrid Periodic Decoupling), a novel framework that decouples traffic data into periodic and residual components. The periodic component is handled by the Hybrid Periodic Representation Module, which extracts fine-grained daily and weekly patterns using learnable periodic embeddings and spatial-temporal attention. The residual component, which captures non-periodic, high-frequency fluctuations, is modeled by the Frequency-Aware Residual Representation Module, leveraging complex-valued MLP in frequency domain. To enforce semantic separation between the two components, we further introduce a Dual-View Alignment Loss, which aligns low-frequency information with the periodic branch and high-frequency information with the residual branch. Extensive experiments on four real-world traffic datasets demonstrate that HyperD achieves state-of-the-art prediction accuracy, while offering superior robustness under disturbances and improved computational efficiency compared to existing methods.
Minlan Shao, Zijian Zhang 0009, Yili Wang 0004, Yiwei Dai, Xu Shen 0002, Xin Wang 0035
AAAI4
2026 BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
abstract
Rui Miao, Yixin Liu, Yili Wang, Xu Shen, Yue Tan, Yiwei Dai, Shirui Pan, Xin Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Rui Miao 0003, Yixin Liu 0001, Yili Wang 0004, Xu Shen 0002, Yiwei Dai, Shirui Pan, Xin Wang 0035
ACL (1)6
2026 MoEC: A Memory-Routed Mixture-of-Experts Controller for Adaptive Minecraft Control
abstract
Embodied agents in open-ended environments such as Minecraft increasingly adopt planner-controller architectures, with large language models acting as high-level planners.While planning has advanced rapidly, control remains underexplored.Existing systems commonly rely on a monolithic policy to execute subgoals across varying contexts, forcing incompatible behaviors into a shared parameter space and causing interference that scaling only partially mitigates.To address this, we propose MoEC, a Memory-Routed Mixture-of-Experts Controller for Adaptive Minecraft Control.MoEC routes via a subgoal-indexed, nonparametric expert memory and regulates capacity through failure-triggered expert growth and redundancy-aware consolidation.This design enables continual adaptation without full retraining, while maintaining parameter efficiency and with bounded inference cost.We evaluate MoEC on diverse and compositional Minecraft tasks, demonstrating significant gains in adaptability, robustness, and execution consistency over strong baselines, yielding a scalable and efficient alternative for openended control.
Jianghui Wang, Ziqiong Liu, Dong Li 0025, Yiwei Dai, Emad Barsoum
ACL (1)6
2025 Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems
abstract
The communication topology in large language model-based multi-agent systems fundamentally governs inter-agent collaboration patterns, critically shaping both the efficiency and effectiveness of collective decision-making.While recent studies for communication topology automated design tend to construct sparse structures for efficiency, they often overlook why and when sparse and dense topologies help or hinder collaboration.In this paper, we present a causal framework to analyze how agent outputs, whether correct or erroneous, propagate under topologies with varying sparsity.Our empirical studies reveal that moderately sparse topologies, which effectively suppress error propagation while preserving beneficial information diffusion, typically achieve optimal task performance.Guided by this insight, we propose a novel topology design approach, EIB-LEARNER, that balances error suppression and beneficial information propagation by fusing connectivity patterns from both dense and sparse graphs.Extensive experiments show the superior effectiveness, communication cost, and robustness of EIB-LEARNER.The code is in
Xu Shen 0002, Yixin Liu 0001, Yiwei Dai, Yili Wang 0004, Rui Miao 0003, Shirui Pan, Xin Wang 0035
EMNLP3
2025 Latte: Transfering LLMs' Latent-level Knowledge for Few-shot Tabular Learning
abstract
Few-shot tabular learning, in which machine learning models are trained with a limited amount of labeled data, provides a cost-effective approach to addressing real-world challenges. The advent of Large Language Models (LLMs) has sparked interest in leveraging their pre-trained knowledge for few-shot tabular learning. Despite promising results, existing approaches either rely on test-time knowledge extraction, which introduces undesirable latency, or text-level knowledge, which leads to unreliable feature engineering. To overcome these limitations, we propose Latte, a training-time knowledge extraction framework that transfers the latent prior knowledge within LLMs to optimize a more generalized downstream model. Latte enables general knowledge-guided downstream tabular learning, facilitating the weighted fusion of information across different feature values while reducing the risk of overfitting to limited labeled data. Furthermore, Latte is compatible with existing unsupervised pre-training paradigms and effectively utilizes available unlabeled samples to overcome the performance limitations imposed by an extremely small labeled dataset. Extensive experiments on various few-shot tabular learning benchmarks demonstrate the superior performance of Latte, establishing it as a state-of-the-art approach in this domain. Our code is available at https://github.com/ruxueshi/Latte.git.
Ruxue Shi, Hengrui Gu 0002, Hangting Ye, Yiwei Dai, Xu Shen 0002, Xin Wang 0035
IJCAI4
2024 Mitigate Extrinsic Social Bias in Pre-trained Language Models via Continuous Prompts Adjustment
abstract
Although pre-trained language models (PLMs) have been widely used in natural language understandings (NLU), they are still exposed to fairness issues.Most existing extrinsic debiasing methods rely on manually curated word lists for each sensitive groups to modify training data or to add regular constraints.However, these word lists are often limited by length and scope, resulting in the degradation performance of extrinsic bias mitigation.To address the aforementioned issues, we propose a Continuous Prompts Adjustment Debiasing method (CPAD), which generates continuous token lists from the entire vocabulary space and uses them to bridge the gap between outputs and targets in fairness learning process.Specifically, CPAD encapsulates fine-tuning objective and debiasing objectives into several independent prompts.To avoid the limitation of manual word lists, in fairness learning phase, we extract outputs from the entire vocabulary space via fine-tuned PLM.Then, we aggregate the outputs from the same sensitive group as continuous token lists to map the outputs into protected attribute labels.Finally, after we learn the debiasing prompts in the perspective of adversarial learning, we improve fairness by adjusting continuous prompts at model inference time.Through extensive experiments on three NLU tasks, we evaluate the debiasing performance from the perspectives of group fairness and fairness through unawareness.The experimental results show that CPAD outperforms all baselines in term of single and two-attributes debiasing performance.
Yiwei Dai, Hengrui Gu 0002, Ying Wang 0009, Xin Wang 0035
EMNLP1
2024 Pre-training Graph Neural Networks via Weighted Meta Learning
abstract
Recent researches have demonstrated pre-training Graph Neural Networks (GNNs) via meta learning can enhance their performance on learning representations from unlabeled data. The main idea behind them is to learn the transferable priors that work across the distribution of tasks. However, existing methods often leverage a uniform task sampling strategy and ignore the relations between their original graphs and them. This may lead the model learning the redundant information in the pre-training process. In this work, we propose Meta Graph Neural Network (MGNN), a graph pre-training method via weighted meta learning. MGNN trys to learn the transferable experiences from diverse distributions without the side-effect of redundant information. First, we break up a task in traditional meta learning into several sub-tasks to construct the minimum evaluation units. Then, to fully utilize the graph information, we use a two-stage optimization with contrastive loss functions to learn the experience priors on node- and graph-levels in meta-training process. Third, to qualify redundant information, we design an evaluation module which calculates the mutual information in a graph view. Finally, we propose a new optimization objective in meta-testing process, which reduce the model’s attention to redundant information to alleviate the negative impact of them. To validate the effectiveness of the proposed method, extensive experiments are conducted on Cora, Citeseer and Pubmed datasets with several GNNs architectures. Experimental results show that our proposed method outperforms existing GNN pre-training algorithms.
Yiwei Dai, Mingchen Sun, Xin Wang 0035
IJCNN1
2021 Modeling-Assisted InSAR Phase-Unwrapping Method for Mapping Mine Subsidence
abstract
Compared with traditional measurement technologies, synthetic aperture radar interferometry (InSAR) has unique advantages in monitoring ground subsidence due to underground mining. However, when the subsidence gradient of the subsidence trough exceeds the maximum measurable gradient of InSAR technology, the interference fringes will be too dense, causing phase aliasing. As a result, it is impossible to obtain correct phase-unwrapping result. The main objectives of this letter are two folded. First is to develop an unwrapping strategy to deal with the unwrapping problem caused by large subsidence gradient at the mine subsidence trough. The main idea of this strategy is to estimate most of the subsidence phase by multiple model inversions based on iterative approach. Then, the model phases from multiple models are combined with the final unwrapped residual phase. Another objective of this letter is to evaluate the feasibility of the three common deformation models, i.e., Mogi, probability integral method (PIM), and Okada, in solving the phase-unwrapping problem. Their advantages and disadvantages are outlined. Both the simulated data and real data are used for this experiment. The result shows that the problem of large subsidence gradient in the differential interferometric synthetic aperture radar (DInSAR) results can be solved by multiple model inversion. Among the three models, the use of Okada model seems to provide slightly more accurate result for solving the large-scale subsidence in the mining area than the other two models with the proposed strategy.
Yiwei Dai, Alex Hayman Ng, Liyuan Li, Linlin Ge, Tingye Tao
IEEE Geosci. Remote. Sens. Lett.1