EDBT 2026 Demo / reviewers in the wild / expert
Jiarui Yao
dblp:249/9101
· DBLP profile ↗
9ranked-venue papers
6as first author
7since 2021 · last 2025
0009-0005-6662-9500ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 32% Reinforcement learning · 26% Information extraction and text analysis · 12% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Theoretical computer science
1 paper |
Automated reasoning and model checking · 100% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › community question answering
answer selection |
0.9 | 1 | 2025 | FANS: Formal Answer Selection for LLM Natural Language Math Reasoning Using Lean4 · EMNLP 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL · NeurIPS 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents · ACL (1) 2025 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | FANS: Formal Answer Selection for LLM Natural Language Math Reasoning Using Lean4 · EMNLP 2025 |
Natural language and speech › Language models and text generation › alignment
pluralistic alignment |
0.9 | 1 | 2025 | MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning · EMNLP 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL · NeurIPS 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning · EMNLP 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning · EMNLP 2025 |
Machine learning › Optimization for machine learning
variance reduction |
0.9 | 1 | 2025 | Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL · NeurIPS 2025 |
Automated reasoning and model checking › theorem proving
interactive theorem proving |
0.9 | 1 | 2025 | FANS: Formal Answer Selection for LLM Natural Language Math Reasoning Using Lean4 · EMNLP 2025 |
Security and privacy of machine learning
adversarial attack |
0.8 | 1 | 2024 | Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models · NeurIPS 2024 |
Security and privacy of machine learning
poisoning attack |
0.8 | 1 | 2024 | Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models · NeurIPS 2024 |
Natural language and speech › Information extraction and text analysis
factuality assessment |
0.5 | 1 | 2021 | Factuality Assessment as Modal Dependency Parsing · ACL/IJCNLP (1) 2021 |
Knowledge, reasoning and agents › Multi-agent systems
crowdsourcing |
0.4 | 1 | 2020 | Annotating Temporal Dependency Graphs via Crowdsourcing · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
data annotation |
0.4 | 1 | 2020 | Annotating Temporal Dependency Graphs via Crowdsourcing · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › relation extraction › event relation extraction
temporal relation extraction |
0.4 | 1 | 2020 | Annotating Temporal Dependency Graphs via Crowdsourcing · EMNLP (1) 2020 |
Machine learning › Reinforcement learning
benchmark design |
0.3 | 1 | 2025 | EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents · ACL (1) 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model |
0.3 | 1 | 2025 | MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning · EMNLP 2025 |
Automated reasoning and model checking
theorem proving |
0.3 | 1 | 2025 | FANS: Formal Answer Selection for LLM Natural Language Math Reasoning Using Lean4 · EMNLP 2025 |
Computer vision › Vision and language
vision-language model |
0.2 | 1 | 2024 | Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models · NeurIPS 2024 |
Computer vision › Video understanding and tracking
temporal understanding |
0.1 | 1 | 2020 | Annotating Temporal Dependency Graphs via Crowdsourcing · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
reward model · 1.7large language model · 1.7reinforcement learning · 0.9online routing · 0.9large language model prompting · 0.9context-aware mixture modeling · 0.9bradley-terry model · 0.9agent evaluation · 0.9RAFT · 0.9GRPO · 0.9transferability analysis · 0.8data poisoning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EscapeBench: Towards Advancing Creative Intelligence of Language Model AgentsabstractCheng Qian, Peixuan Han, Qinyu Luo, Bingxiang He, Xiusi Chen, Yuji Zhang, Hongyi Du, Jiarui Yao, Xiaocheng Yang, Denghui Zhang, Yunzhu Li, Heng Ji. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Cheng Qian 0008, Peixuan Han, Qinyu Luo, Bingxiang He, Xiusi Chen, Yuji Zhang 0002, Hongyi Du, Jiarui Yao, Xiaocheng Yang, Yunzhu Li, Heng Ji 0001 |
ACL (1) | 8 |
| 2025 | MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference LearningabstractReward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs).However, reward modeling based on the Bradley-Terry (BT) model assumes a global reward function, failing to capture the inherently diverse and heterogeneous human preferences.Hence, such oversimplification limits LLMs from supporting personalization and pluralistic alignment.Theoretically, we show that when human preferences follow a mixture distribution of diverse subgroups, a single BT model has an irreducible error.While existing solutions, such as multi-objective learning with finegrained annotations, help address this issue, they are costly and constrained by predefined attributes, failing to fully capture the richness of human values.In this work, we introduce MiCRo, a two-stage framework that enhances personalized preference learning by leveraging large-scale binary preference datasets without requiring explicit fine-grained annotations.In the first stage, MiCRo introduces context-aware mixture modeling approach to capture diverse human preferences.In the second stage, Mi-CRo integrates an online routing strategy that dynamically adapts mixture weights based on specific context to resolve ambiguity, allowing for efficient and scalable preference adaptation with minimal additional supervision.Experiments on multiple preference datasets demonstrate that MiCRo effectively captures diverse human preferences and significantly improves downstream personalization. Jingyan Shen, Jiarui Yao, Rui Yang 0010, Feng Luo 0003, Rui Pan 0002, Tong Zhang 0001, Han Zhao 0002 |
EMNLP | 2 |
| 2025 | FANS: Formal Answer Selection for LLM Natural Language Math Reasoning Using Lean4abstractLarge Language Models (LLMs) have displayed astonishing abilities in various tasks, especially in text generation, classification, question answering, etc.However, the reasoning ability of LLMs still faces many debates, especially in math reasoning.The inherent ambiguity of Natural Language (NL) limits LLMs' ability to perform verifiable reasoning, making the answers lack coherence and trustworthy support.To tackle the above challenges, we propose a novel framework named FANS: Formal ANswer Selection for LLM Natural Language Math Reasoning Using Lean4.It is a pioneering framework that utilizes Lean4 to enhance LLMs' NL math reasoning ability.In particular, given an NL math question and LLM-generated answers, FANS first translates it into Lean4 theorem statements.Then it invokes another Lean4 prover LLM to produce proofs, and finally verifies the proofs by Lean4 compiler.Answers are selected based on the verifications.It enhances LLMs' NL math ability in providing a computer-verifiable solution for its correct answer and proposes an alternative method for answer selection beyond the reward model based ones.Our experiments demonstrate the effectiveness of FANS with an improvement of nearly 2% across several math benchmarks, and even higher further based on reward models or in subfields such as algebra and number theory that Lean4 is better at.The code is available in https://github.com/MaxwellJryao/FANS. Jiarui Yao, Ruida Wang |
EMNLP | 1 |
| 2025 | Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RLabstractChain-of-thought (CoT) reasoning in large language models (LLMs) can be formalized as a latent variable problem, where the model needs to generate intermediate reasoning steps. While prior approaches such as iterative reward-ranked fine-tuning (RAFT) have relied on such formulations, they typically apply uniform inference budgets across prompts, which fails to account for variability in difficulty and convergence behavior. This work identifies the main bottleneck in CoT training as inefficient stochastic gradient estimation due to static sampling strategies. We propose GVM-RAFT, a prompt-specific Dynamic Sample Allocation Strategy designed to minimize stochastic gradient variance under a computational budget constraint. The method dynamically allocates computational resources by monitoring prompt acceptance rates and stochastic gradient norms, ensuring that the resulting gradient variance is minimized. Our theoretical analysis shows that the proposed dynamic sampling strategy leads to accelerated convergence guarantees under suitable conditions. Experiments on mathematical reasoning show that GVM-RAFT achieves a 2-4x speedup and considerable accuracy improvements over vanilla RAFT. The proposed dynamic sampling strategy is general and can be incorporated into other reinforcement learning algorithms, such as GRPO, leading to similar improvements in convergence and test accuracy. Jiarui Yao, Yifan Hao 0002, Hanning Zhang, Hanze Dong, Wei Xiong 0015, Nan Jiang 0008, Tong Zhang 0001 |
NeurIPS | 1 |
| 2024 | Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language ModelsabstractVision-Language Models (VLMs) excel in generating textual responses from visual inputs, but their versatility raises security concerns. This study takes the first step in exposing VLMs’ susceptibility to data poisoning attacks that can manipulate responses to innocuous, everyday prompts. We introduce Shadowcast, a stealthy data poisoning attack where poison samples are visually indistinguishable from benign images with matching texts. Shadowcast demonstrates effectiveness in two attack types. The first is a traditional Label Attack, tricking VLMs into misidentifying class labels, such as confusing Donald Trump for Joe Biden. The second is a novel Persuasion Attack, leveraging VLMs’ text generation capabilities to craft persuasive and seemingly rational narratives for misinformation, such as portraying junk food as healthy. We show that Shadowcast effectively achieves the attacker’s intentions using as few as 50 poison samples. Crucially, the poisoned samples demonstrate transferability across different VLM architectures, posing a significant concern in black-box settings. Moreover, Shadowcast remains potent under realistic conditions involving various text prompts, training data augmentation, and image compression techniques. This work reveals how poisoned VLMs can disseminate convincing yet deceptive misinformation to everyday, benign users, emphasizing the importance of data integrity for responsible VLM deployments. Our code is available at: https://github.com/umd-huang-lab/VLM-Poisoning. Yuancheng Xu, Jiarui Yao, Manli Shu, Yanchao Sun, Zichu Wu, Ning Yu 0006, Tom Goldstein, Furong Huang |
NeurIPS | 2 |
| 2022 | Modal Dependency Parsing via Language Model PrimingabstractThe task of modal dependency parsing aims to parse a text into its modal dependency structure, which is a representation for the factuality of events in the text.We design a modal dependency parser that is based on priming pre-trained language models, and evaluate the parser on two data sets.Compared to baselines, we show an improvement of 2.6% in F-score for English and 4.6% for Chinese.To the best of our knowledge, this is also the first work on Chinese modal dependency parsing. Jiarui Yao, Nianwen Xue, Bonan Min |
NAACL-HLT | 1 |
| 2021 | Factuality Assessment as Modal Dependency ParsingabstractJiarui Yao, Haoling Qiu, Jin Zhao, Bonan Min, Nianwen Xue. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jiarui Yao, Haoling Qiu, Bonan Min, Nianwen Xue |
ACL/IJCNLP (1) | 1 |
| 2020 | Annotating Temporal Dependency Graphs via CrowdsourcingabstractWe present the construction of a corpus of 500 Wikinews articles annotated with temporal dependency graphs (TDGs) that can be used to train systems to understand temporal relations in text.We argue that temporal dependency graphs, built on previous research on narrative times and temporal anaphora, provide a representation scheme that achieves a good balance between completeness and practicality in temporal annotation.We also provide a crowdsourcing strategy to annotate TDGs, and demonstrate the feasibility of this approach with an evaluation of the quality of the annotation, and the utility of the resulting data set by training a machine learning model on this data set.This data set is publicly available 1 . Jiarui Yao, Haoling Qiu, Bonan Min, Nianwen Xue |
EMNLP (1) | 1 |
| 2019 | SMART: A Stratified Machine Reading Test
Jiarui Yao, Minxuan Feng, Haixia Feng, Nianwen Xue |
NLPCC (1) | 1 |