VLDB 2026 Research / reviewers in the wild / expert
Yi Chern Tan
dblp:242/8043
· DBLP profile ↗
8ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 46% Trustworthy machine learning · 20% Reinforcement learning · 11% | |
| Databases, data mining, and information retrieval
1 paper |
Data models and query languages · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.9 | 2 | 2021 | GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing · ICLR 2021 SParC: Cross-Domain Semantic Parsing in Context · ACL (1) 2019 |
Natural language and speech › Language models and text generation › prompt tuning
adversarial prompt tuning |
0.9 | 1 | 2025 | Reverse Engineering Human Preferences with Reinforcement Learning · NeurIPS 2025 |
Natural language and speech › Language models and text generation › prompting
chain-of-thought prompting |
0.9 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Machine learning › Reinforcement learning
learning from failure |
0.9 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
0.9 | 1 | 2025 | Reverse Engineering Human Preferences with Reinforcement Learning · NeurIPS 2025 |
Natural language and speech › Language models and text generation
prompting |
0.9 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Machine learning › Representation and self-supervised learning
pre-training |
0.5 | 1 | 2021 | GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing · ICLR 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
explanation generation |
0.4 | 1 | 2020 | ESPRIT: Explaining Solutions to Physical Reasoning Tasks · ACL 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
physical reasoning |
0.4 | 1 | 2020 | ESPRIT: Explaining Solutions to Physical Reasoning Tasks · ACL 2020 |
Machine learning › Trustworthy machine learning
fairness |
0.4 | 1 | 2019 | Assessing Social and Intersectional Biases in Contextualized Word Representations · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › fairness
intersectional bias |
0.4 | 1 | 2019 | Assessing Social and Intersectional Biases in Contextualized Word Representations · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › fairness
social bias |
0.4 | 1 | 2019 | Assessing Social and Intersectional Biases in Contextualized Word Representations · NeurIPS 2019 |
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL |
0.4 | 1 | 2019 | CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.3 | 1 | 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-context · EMNLP 2025 |
Natural language and speech › Language models and text generation › text representation
contextualized word embeddings |
0.1 | 1 | 2019 | Assessing Social and Intersectional Biases in Contextualized Word Representations · NeurIPS 2019 |
Natural language and speech › Question answering and dialogue systems › natural language interface
conversational interfaces |
0.1 | 1 | 2019 | CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases · EMNLP/IJCNLP (1) 2019 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.9prompting · 0.9preamble generation · 0.9judge language model · 0.9in-context learning · 0.9cross-domain generalization · 0.8pre-training · 0.5grammar augmentation · 0.5neural language model · 0.4embedding association test · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | No Need for Explanations: LLMs can implicitly learn from mistakes in-contextabstractShowing incorrect answers to Large Language Models (LLMs) is a popular strategy to improve their performance in reasoning-intensive tasks.It is widely assumed that, in order to be helpful, the incorrect answers must be accompanied by comprehensive rationales, explicitly detailing where the mistakes are and how to correct them.However, in this work we present a counterintuitive finding: we observe that LLMs perform better in math reasoning tasks when these rationales are eliminated from the context and models are left to infer on their own what makes an incorrect answer flawed.This approach also substantially outperforms chainof-thought prompting in our evaluations.These results are consistent across LLMs of different sizes and varying reasoning abilities.To gain an understanding of why LLMs learn from mistakes more effectively without explicit corrective rationales, we perform a thorough analysis, investigating changes in context length and answer diversity between different prompting strategies, and their effect on performance.We also examine evidence of overfitting to the in-context rationales when these are provided, and study the extent to which LLMs are able to autonomously infer high-quality corrective rationales given only incorrect answers as input.We find evidence that, while incorrect answers are more beneficial for LLM learning than additional diverse correct answers, explicit corrective rationales over-constrain the model, thus limiting those benefits. Lisa Alazraki, Maximilian Mozes, Jon Ander Campos, Yi Chern Tan, Marek Rei, Max Bartolo |
EMNLP | 4 |
| 2025 | Reverse Engineering Human Preferences with Reinforcement LearningabstractThe capabilities of Large Language Models (LLMs) are routinely evaluated by other LLMs trained to predict human preferences. This framework—known as *LLM-as-a-judge*—is highly scalable and relatively low cost. However, it is also vulnerable to malicious exploitation, as LLM responses can be tuned to overfit the preferences of the judge. Previous work shows that the answers generated by a candidate-LLM can be edited *post hoc* to maximise the score assigned to them by a judge-LLM. In this study, we adopt a different approach and use the signal provided by judge-LLMs as a reward to adversarially tune models that generate text preambles designed to boost downstream performance. We find that frozen LLMs pipelined with these models attain higher LLM-evaluation scores than existing frameworks. Crucially, unlike other frameworks which intervene directly on the model's response, our method is virtually undetectable. We also demonstrate that the effectiveness of the tuned preamble generator transfers when the candidate-LLM and the judge-LLM are replaced with models that are not used during training. These findings raise important questions about the design of more reliable LLM-as-a-judge evaluation settings. They also demonstrate that human preferences can be reverse engineered effectively, by pipelining LLMs to optimise upstream preambles via reinforcement learning—an approach that could find future applications in diverse tasks and domains beyond adversarial attacks. Lisa Alazraki, Yi Chern Tan, Jon Ander Campos, Maximilian Mozes, Marek Rei, Max Bartolo |
NeurIPS | 2 |
| 2021 | GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing
Tao Yu 0009, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang 0002, Dragomir R. Radev, Richard Socher, Caiming Xiong |
ICLR | 5 |
| 2021 | DART: Open-Domain Structured Data Record to Text GenerationabstractLinyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Linyong Nan, Dragomir R. Radev, Rui Zhang 0037, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma 0001, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta 0015, Tao Yu 0009, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani |
NAACL-HLT | 20 |
| 2020 | ESPRIT: Explaining Solutions to Physical Reasoning TasksabstractNazneen Fatema Rajani, Rui Zhang, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Nazneen Fatema Rajani, Rui Zhang 0037, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, Dragomir R. Radev |
ACL | 3 |
| 2019 | SParC: Cross-Domain Semantic Parsing in ContextabstractTao Yu, Rui Zhang, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li, Heyang Er, Irene Li, Bo Pang, Tao Chen, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Vincent Zhang, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Tao Yu 0009, Rui Zhang 0037, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li 0002, Heyang Er, Irene Li, Bo Pang 0004, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Caiming Xiong, Richard Socher, Dragomir R. Radev |
ACL (1) | 4 |
| 2019 | CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to DatabasesabstractTao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao Chen, Yuwen Zhang, Shreya Dixit, Vincent Zhang, Caiming Xiong, Richard Socher, Walter Lasecki, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tao Yu 0009, Rui Zhang 0037, Heyang Er, Suyi Li 0002, Eric Xue 0001, Bo Pang 0004, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Alexander R. Fabbri, Zifan Li, Shreya Dixit, Caiming Xiong, Richard Socher, Walter S. Lasecki, Dragomir R. Radev |
EMNLP/IJCNLP (1) | 8 |
| 2019 | Assessing Social and Intersectional Biases in Contextualized Word RepresentationsabstractSocial bias in machine learning has drawn significant attention, with work ranging from demonstrations of bias in a multitude of applications, curating definitions of fairness for different contexts, to developing algorithms to mitigate bias. In natural language processing, gender bias has been shown to exist in context-free word embeddings. Recently, contextual word representations have outperformed word embeddings in several downstream NLP tasks. These word representations are conditioned on their context within a sentence, and can also be used to encode the entire sentence. In this paper, we analyze the extent to which state-of-the-art models for contextual word representations, such as BERT and GPT-2, encode biases with respect to gender, race, and intersectional identities. Towards this, we propose assessing bias at the contextual word level. This novel approach captures the contextual effects of bias missing in context-free word embeddings, yet avoids confounding effects that underestimate bias at the sentence encoding level. We demonstrate evidence of bias at the corpus level, find varying evidence of bias in embedding association tests, show in particular that racial bias is strongly encoded in contextual word models, and observe that bias effects for intersectional minorities are exacerbated beyond their constituent minority identities. Further, evaluating bias effects at the contextual word level captures biases that are not captured at the sentence level, confirming the need for our novel approach. Yi Chern Tan, L. Elisa Celis |
NeurIPS | 1 |