VLDB 2026 Research / reviewers in the wild / expert
Yiyao Yang
dblp:318/3821
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TPCA-GAD: Topology preference-consistency aggregation for zero-shot graph anomaly detection
Youjian Yao, Tailun Chen, Yiyao Yang, Qisen Yan |
Expert Syst. Appl. | 3 |
| 2026 | Decoupling and fusion: A hybrid graph anomaly detection framework for extremely sparse labels
Youjian Yao, Yiyao Yang, Qisen Yan, Tailun Chen |
Neurocomputing | 2 |
| 2026 | CFSM: A Novel Causal Feature Selection Module for Two-Dimensional Out-of-Distribution GeneralizationabstractIn real-world scenarios, training and test data are often collected in diverse settings, leading to domain shifts arising from evolving environments and selection bias. While causality-inspired methods have shown promising results in tackling the out-of-distribution (OOD) generalization issue, prior methods treat the discovered differences across domains as confounding variables. While effective in handling domain differences (i.e., unseen environmental features in test data), they may fail when confronted with intricate spurious correlations in real-world datasets. In this study, we first analyze this limitation to inadequate modeling of causal intervention and derive the OOD generalization bound to explain the challenges it introduces. To address this problem, we propose a modified causal intervention approach to mitigate various types of confounders. Motivated by the mathematical formulation of our modified causal intervention, we introduce the Causal Feature Selection Module (CFSM) to suppress model weights on both domain-differences features and spurious correlation features. Integrated within the Base Feature Extraction Module, In-Sample Module, and Cross-Sample Module (B-I-C architecture), CFSM collectively neutralizes the confounding effects arising from both domain discrepancies and correlation distinctions, thereby achieving causal feature selection. Under mild assumptions, we prove that the proposed CFSM method can achieve strictly lower OOD errors. Further experiments conducted on various benchmark datasets demonstrate the effectiveness of the proposed method. Compared to previous deconfounding methods, our method not only mitigates the effect of domain-differences features but also the hard-to-identify spurious correlation features, achieving significant improvements in two-dimensional OOD generalization. Weihan Yin, Yiyao Yang, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | HaVen: Hallucination-Mitigated LLM for Verilog Code Generation Aligned with HDL EngineersabstractRecently, the use of large language models (LLMs) for Verilog code generation has attracted great research interest to enable hardware design automation. However, previous works have shown a gap between the ability of LLMs and the practical demands of hardware description language (HDL) engineering. This gap includes differences in how engineers phrase questions and hallucinations in the code generated. To address these chal-lenges, we introduce Haven, a novel LLM framework designed to mitigate hallucinations and align Verilog code generation with the practices of HDL engineers. Haven tackles hallucination issues by proposing a comprehensive taxonomy and employing a chain-of-thought (CoT) mechanism to translate symbolic modalities (e.g. truth tables, state diagrams, etc.) into accurate natural language descriptions. Furthermore, Haven bridges this gap by using a data augmentation strategy. It synthesizes high-quality instruction-code pairs that match real HDL engineering practices. Our experiments demonstrate that Haven significantly improves the correctness of Verilog code generation, outperforming state-of-the-art LLM-based Verilog generation methods on VerilogEval and RTLLM benchmark. Haven is publicly available at https://github.com/Intelli2ent-Computing-Research-Group/HaVen. Yiyao Yang, Fu Teng, Mengnan Qi, Chenyang Lv, Xuhong Zhang 0002, Zhezhi He |
DATE | 1 |
| 2025 | VeriRL: Boosting the LLM-based Verilog Code Generation via Reinforcement LearningabstractRecent advancements in code generation have shown remarkable success across software domains, yet hardware description languages (HDLs) such as Verilog remain underexplored due to their concurrency semantics, syntactic rigidity, and simulation complexity. In this work, we address these challenges by introducing a reinforcement learning (RL) framework tailored for Verilog code generation. We first construct Veribench-53K, a high-quality dataset curated from over 700K Verilog problems, enriched with structured prompts, complexity labels, and diverse testbenches. To tackle the problem of sparse and noisy reward signals, we propose a Trace-back based Rescore mechanism that leverages reasoning paths and iterative refinement to enhance feedback reliability and support reward model training. Furthermore, to mitigate catastrophic forgetting and overfitting during RL fine-tuning, we introduce a sample-balanced weighting strategy that adaptively balances learning dynamics based on reward-probability distributions. These innovations are integrated into an iterative RL pipeline that co-evolves the policy and reward models. In contrast to recent work such as CraftRTL, which relies on large-scale closed-source model distillation, and DeepSeekstyle approaches that struggle with sparse feedback, our method demonstrates superior performance using a smaller but high-quality dataset combined with RL optimization. Experiments on Verilog generation tasks demonstrate state-of-the-art performance, with substantial gains in test pass rate, functional correctness, and compilation robustness. Our findings highlight the potential of RL-driven approaches for structured code generation in hardware-centric domains. VeriRL is publicly available at https://github.com/omniAI-Lab/VeriRL. Fu Teng, Miao Pan, Xuhong Zhang 0002, Zhezhi He, Yiyao Yang, Xinyi Chai, Mengnan Qi, Liqiang Lu, Jianwei Yin |
ICCAD | 5 |
| 2025 | VulTriNet: A software vulnerability detection method based on tri-channel network
Yiyao Yang, Youjian Yao |
Inf. Softw. Technol. | 1 |
| 2025 | DE-PSA: Learning from unlabeled data by dual-stage label propagation for positive selection algorithm
Yiyao Yang |
Knowl. Based Syst. | 2 |
| 2024 | Vision-Language Alignment Learning Under Affinity and Divergence Principles for Few-Shot Out-of-Distribution Generalization
Weihan Yin, Yiyao Yang, Fan Wu 0006, Zhaoyu Zeng, Qinying Gu, Xinbing Wang, Chenghu Zhou, Nanyang Ye 0001 |
Int. J. Comput. Vis. | 3 |