VLDB 2026 Research / reviewers in the wild / expert
Yangqiaoyu Zhou
dblp:304/3250
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
3D vision · 36% Transfer learning and domain adaptation · 27% Language models and text generation · 27% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › geometric estimation › geometric model fitting
hypothesis generation |
0.9 | 1 | 2025 | Literature Meets Data: A Synergistic Approach to Hypothesis Generation · ACL (1) 2025 |
Computational science and engineering › AI for science
AI for scientific discovery |
0.9 | 1 | 2025 | Literature Meets Data: A Synergistic Approach to Hypothesis Generation · ACL (1) 2025 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.7 | 1 | 2023 | FLamE: Few-shot Learning from Natural Language Explanations · ACL (1) 2023 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.7 | 1 | 2023 | FLamE: Few-shot Learning from Natural Language Explanations · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis › text classification
deception detection |
0.3 | 1 | 2025 | Literature Meets Data: A Synergistic Approach to Hypothesis Generation · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
literature-based retrieval · 1.7data-driven generation · 1.7fine-tuning · 0.7GPT-3 · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Literature Meets Data: A Synergistic Approach to Hypothesis GenerationabstractAI holds promise for transforming scientific processes, including hypothesis generation.Prior work on hypothesis generation can be broadly categorized into theory-driven and datadriven approaches.While both have proven effective in generating novel and plausible hypotheses, it remains an open question whether they can complement each other.To address this, we develop the first method that combines literature-based insights with data to perform LLM-powered hypothesis generation.We apply our method on five different datasets and demonstrate that integrating literature and data outperforms other baselines (8.97% over fewshot, 15.75% over literature-based alone, and 3.37% over data-driven alone).Additionally, we conduct the first human evaluation to assess the utility of LLM-generated hypotheses in assisting human decision-making on two challenging tasks: deception detection and AI generated content detection.Our results show that human accuracy improves significantly by 7.44% and 14.19% on these tasks, respectively.These findings suggest that integrating literature-based and data-driven approaches provides a comprehensive and nuanced framework for hypothesis generation and could open new avenues for scientific inquiry. Haokun Liu, Yangqiaoyu Zhou, Chenfei Yuan, Chenhao Tan |
ACL (1) | 2 |
| 2023 | FLamE: Few-shot Learning from Natural Language ExplanationsabstractNatural language explanations have the potential to provide rich information that in principle guides model reasoning.Yet, recent work by Lampinen et al. (2022) has shown limited utility of natural language explanations in improving classification.To effectively learn from explanations, we present FLamE, a two-stage few-shot learning framework that first generates explanations using GPT-3, and then finetunes a smaller model (e.g., RoBERTa) with generated explanations.Our experiments on natural language inference demonstrate effectiveness over strong baselines, increasing accuracy by 17.6% over GPT-3 Babbage and 5.7% over GPT-3 Davinci in e-SNLI.Despite improving classification performance, human evaluation surprisingly reveals that the majority of generated explanations does not adequately justify classification decisions.Additional analyses point to the important role of label-specific cues (e.g., "not know" for the neutral label) in generated explanations. Yangqiaoyu Zhou, Yiming Zhang 0022, Chenhao Tan |
ACL (1) | 1 |
| 2023 | Learning to Ignore Adversarial AttacksabstractDespite the strong performance of current NLP models, they can be brittle against adversarial attacks.To enable effective learning against adversarial inputs, we introduce the use of rationale models that can explicitly learn to ignore attack tokens.We find that the rationale models can successfully ignore over 90% of attack tokens.This approach leads to consistent and sizable improvements (∼10%) over baseline models in robustness on three datasets for both BERT and RoBERTa, and also reliably outperforms data augmentation with adversarial examples alone.In many cases, we find that our method is able to close the gap between model performance on a clean test set and an attacked test set and hence reduce the effect of adversarial attacks. Yiming Zhang 0022, Yangqiaoyu Zhou, Samuel Carton, Chenhao Tan |
EACL | 2 |
| 2022 | GraPhyC: Using Consensus to Infer Tumor EvolutionabstractWe consider the problem of finding a consensus tumor evolution tree from a set of conflicting input trees. In contrast to traditional phylogenetic trees, the tumor trees we consider do not have the same set of labels applied to the leaves of each tree. We describe several distance measures between these tumor trees. Our GraPhyC algorithm solves the consensus problem using a weighted directed graph where vertices are sets of mutations and edges are weighted based on the number of times a parental relationship is observed between their constituent mutations in the input trees. We find a minimum weight spanning arborescence in this graph and prove that it minimizes the total distance to all input trees for one of our distance measures. We also describe several extensions of our GraPhyC approach. On simulated data we show that GraPhyC outperforms a baseline method and demonstrate that GraPhyC can be an effective means of computing centroids in k-medians clustering. We analyze two real sequencing datasets and find that GraPhyC is able to identify a tree not included in the set of input trees, but that contains characteristics supported by other reported evolutionary reconstructions of this tumor. Kiya Govek, Camden Sikes, Yangqiaoyu Zhou, Layla Oesper |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |