VLDB 2026 Research / reviewers in the wild / expert
Wayne Lu
dblp:193/1302
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Transfer learning and domain adaptation · 35% Deep learning architectures and training · 15% Generative modeling · 15% | |
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 54% Web and social media mining · 46% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
data augmentation |
1.0 | 1 | 2026 | DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text Classification · AAAI 2026 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
1.0 | 1 | 2026 | From Blind Transfer to Wise Selection: Prototype-Driven Neighbor-Domain Adaptation for Fake News Detection · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.0 | 1 | 2026 | Breaking Down Market Barriers: Distilled Prompt-Tuning Approach for Cross-Market Recommendation · AAAI 2026 |
Machine learning › Generative modeling › synthetic data generation › text data augmentation
LLM-based text augmentation |
1.0 | 1 | 2026 | DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text Classification · AAAI 2026 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation |
1.0 | 1 | 2026 | From Blind Transfer to Wise Selection: Prototype-Driven Neighbor-Domain Adaptation for Fake News Detection · AAAI 2026 |
Natural language and speech › Information extraction and text analysis
text classification |
1.0 | 1 | 2026 | DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text Classification · AAAI 2026 |
Recommender systems
cross-domain recommendation |
1.0 | 1 | 2026 | From IDs to Semantics: A Generative Framework for Cross-Domain Recommendation with Adaptive Semantic Tokenization · AAAI 2026 |
Web and social media mining › misinformation detection
fake news detection |
1.0 | 1 | 2026 | From Blind Transfer to Wise Selection: Prototype-Driven Neighbor-Domain Adaptation for Fake News Detection · AAAI 2026 |
Web and social media mining › misinformation detection › fake news detection
multimodal fake news detection |
1.0 | 1 | 2026 | From Blind Transfer to Wise Selection: Prototype-Driven Neighbor-Domain Adaptation for Fake News Detection · AAAI 2026 |
Recommender systems
prompt tuning |
1.0 | 1 | 2026 | Breaking Down Market Barriers: Distilled Prompt-Tuning Approach for Cross-Market Recommendation · AAAI 2026 |
Machine learning › Transfer learning and domain adaptation
negative transfer |
0.3 | 1 | 2026 | From Blind Transfer to Wise Selection: Prototype-Driven Neighbor-Domain Adaptation for Fake News Detection · AAAI 2026 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2026 | From Blind Transfer to Wise Selection: Prototype-Driven Neighbor-Domain Adaptation for Fake News Detection · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
teacher-student distillation · 2.0prototype-based asymmetric distance · 2.0prompt tuning · 2.0parameter-efficient fine-tuning · 2.0large language model · 2.0gumbel-based neighbor selection · 2.0domain-collaborative attention · 2.0domain-adaptive tokenization · 1.0diversity-aware planning · 1.0autoregressive generation · 1.0adaptive incremental sampling · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From IDs to Semantics: A Generative Framework for Cross-Domain Recommendation with Adaptive Semantic TokenizationabstractCross-domain recommendation (CDR) is crucial for improving recommendation accuracy and generalization, yet traditional methods are often hindered by the reliance on shared user/item IDs, which are unavailable in most real-world scenarios. Consequently, many efforts have focused on learning disentangled representations through multi-domain joint training to bridge the domain gaps. Recent Large Language Model (LLM)-based approaches show promise, they still face critical challenges, including: (1) the \textbf{item ID tokenization dilemma}, which leads to vocabulary explosion and fails to capture high-order collaborative knowledge; and (2) \textbf{insufficient domain-specific modeling} for the complex evolution of user interests and item semantics. To address these limitations, we propose \textbf{GenCDR}, a novel \textbf{Gen}erative \textbf{C}ross-\textbf{D}omain \textbf{R}ecommendation framework. GenCDR first employs a \textbf{Domain-adaptive Tokenization} module, which generates disentangled semantic IDs for items by dynamically routing between a universal encoder and domain-specific adapters. Symmetrically, a \textbf{Cross-domain Autoregressive Recommendation} module models user preferences by fusing universal and domain-specific interests. Finally, a \textbf{Domain-aware Prefix-tree} enables efficient and accurate generation. Extensive experiments on multiple real-world datasets demonstrate that GenCDR significantly outperforms state-of-the-art baselines. Our code is available in the supplementary materials. Peiyu Hu, Wayne Lu, Jia Wang 0009 |
AAAI | 2 |
| 2026 | DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text ClassificationabstractReal-world text classification datasets frequently exhibit long-tail distributions, where numerous classes have sparse data, significantly degrading model performance on these underrepresented categories. While Large Language Models (LLMs) offer promise for data augmentation, existing methods often produce semantically limited samples, neglect "implicit long-tails" (sparse sub-patterns within classes), and lack cost-effective optimization. To address these challenges, we propose \textbf{DEALT (LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text Classification)}, a novel cognitive-inspired framework emulating the human learning process of "recognize, explore, generate, and optimize." DEALT systematically enhances augmented data diversity by first detecting both explicit and implicit long-tails. It then employs an LLM for diversity-aware planning of augmentation strategies, followed by conditional generation. A low-overhead quality and diversity validator filters the synthetic data, and an adaptive incremental sampler refines future augmentation efforts based on proxy model feedback, ensuring efficient and budget-aware optimization. Extensive experiments on multiple public text classification datasets demonstrate DEALT's superiority over state-of-the-art methods in improving tail-class performance and overall model robustness by generating more diverse and high-fidelity augmented data. Wayne Lu, Xiaoxi Cui |
AAAI | 1 |
| 2026 | From Blind Transfer to Wise Selection: Prototype-Driven Neighbor-Domain Adaptation for Fake News DetectionabstractMultimodal fake news detection across different domains is hampered by the critical challenge of negative transfer, which arises from the indiscriminate fusion of knowledge from all available source domains. Existing methods attempt to learn domain-invariant features or leverage external knowledge but often aggregate information from all domains equally. However, these approaches largely ignore the asymmetric relationships between domains, leading to performance degradation when irrelevant or conflicting knowledge is introduced. To address this, we propose a novel PANDA (Prototype-driven Asymmetric Neighbor-Domain Adaptation) framework that dynamically selects and integrates knowledge from only the most beneficial domains. Initially, PANDA employs a Domain-aware Modal Prompt Generation (DMPG) module to learn transferable knowledge representations for each domain. We then introduce a novel Prototype-based Asymmetric Distance (PAD) to quantify directional domain transferability, which guides a Gumbel-based Neighbor Selector (GNS) to identify the most relevant neighbor domains. Subsequently, a Domain-Collaborative Attention (DCA) module adaptively fuses the selected knowledge to enhance the target domain's representation. Extensive experiments on three benchmarks demonstrate PANDA's superiority, outperforming state-of-the-art baselines with an F1-score improvement of 1.5% on the Weibo-21 dataset. Wayne Lu |
AAAI | 1 |
| 2026 | Breaking Down Market Barriers: Distilled Prompt-Tuning Approach for Cross-Market RecommendationabstractCross-market recommendation (CMR) faces severe challenges from distribution shifts between data-rich source markets and sparse target markets. Existing methods rely on a pre-training and fine-tuning paradigm for knowledge transfer, yet suffer from two key limitations: i) the objective gap between pre-training and full-parameter fine-tuning causes loss of generalized knowledge from source markets; ii) the high computational costs of extensive fine-tuning hinder scalability. To this end, we propose DCMPT, a novel Distilled Cross-Market Prompt-Tuning approach. DCMPT reframes the problem under a more efficient pre-training and prompt-tuning paradigm. Instead of full fine-tuning, we adapt a pre-trained universal backbone by freezing its weights and injecting a minimal set of learnable prompts to form a "student" model. To effectively optimize these prompts on sparse data, we introduce a novel teacher-student architecture: a specialized "teacher" model, trained exclusively on the target market, provides dense, market-specific supervision. This guidance is delivered via a dual distillation strategy designed to transfer global ranking patterns and adapt to local consumer tastes. Extensive experiments on real-world market datasets demonstrate that DCMPT significantly outperforms state-of-the-art methods, achieving superior target market performance with substantial parameter-efficiency. Leqi Zhang, Wayne Lu, Haiyang Zhang 0004, Elliott Wen, Zhixuan Liang, Jia Wang 0009 |
AAAI | 2 |
| 2016 | Teaching programming in the context of solving engineering problemsabstractThis paper describes the course that was developed at the authors' University to introduce all first-year engineering students to the fundamentals of computer programming within the context of solving engineering problems. This two credit-hour, semester-long course incorporates the programming language MATLAB and is a required course for students who major in civil, electrical, and mechanical engineering. The course, which is usually taken during the second semester of their first year of study, was designed to utilize active learning techniques by having the students complete a series of laboratory exercises and projects that introduce computer programming and engineering applications. This paper describes the origins of the course, the laboratory exercises and projects, how the course was administered, and an assessment of how successful the course was based on student grades, student feedback, and a student survey. The results indicate that the course increased students' knowledge of programming in the context of solving engineering problems. Joseph P. Hoffbeck, Heather E. Dillon, Robert J. Albright, Wayne Lu, Timothy A. Doughty |
FIE | 4 |