VLDB 2026 Research / reviewers in the wild / expert
Min-Hsuan Yeh
dblp:305/6809
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0001-4945-712XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 58% Information extraction and text analysis · 26% Vision and language · 12% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
hallucination detection |
0.9 | 1 | 2025 | Steer LLM Latents for Hallucination Detection · ICML 2025 |
Natural language and speech › Language models and text generation › large language model
large language model representation |
0.9 | 1 | 2025 | Steer LLM Latents for Hallucination Detection · ICML 2025 |
Natural language and speech › Language models and text generation › model steering
representation steering |
0.9 | 1 | 2025 | Steer LLM Latents for Hallucination Detection · ICML 2025 |
Natural language and speech › Information extraction and text analysis › argument mining
logical fallacy detection |
0.8 | 1 | 2024 | CoCoLoFa: A Dataset of News Comments with Common Logical Fallacies Written by LLM-Assisted Crowds · EMNLP 2024 |
Computer vision › Vision and language › vision-language generation
visual question generation |
0.6 | 1 | 2022 | Multi-VQG: Generating Engaging Questions for Multiple Images · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis › text classification
deception detection |
0.5 | 1 | 2021 | Lying Through One's Teeth: A Study on Verbal Leakage Cues · EMNLP (1) 2021 |
Collaborative and social computing
crowdsourcing |
0.2 | 1 | 2024 | CoCoLoFa: A Dataset of News Comments with Common Logical Fallacies Written by LLM-Assisted Crowds · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
large language model assistants · 1.5BERT fine-tuning · 1.5pseudo-labeling · 0.9optimal transport · 0.9confidence filtering · 0.9end-to-end architecture · 0.6dual-stage architecture · 0.6cross-dataset testing · 0.5LIWC lexicon analysis · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Steer LLM Latents for Hallucination DetectionabstractHallucinations in LLMs pose a significant concern to their safe deployment in real-world applications. Recent approaches have leveraged the latent space of LLMs for hallucination detection, but their embeddings, optimized for linguistic coherence rather than factual accuracy, often fail to clearly separate truthful and hallucinated content.
To this end, we propose the **T**ruthfulness **S**eparator **V**ector (**TSV**), a lightweight and flexible steering vector that reshapes the LLM’s representation space during inference to enhance the separation between truthful and hallucinated outputs, without altering model parameters.
Our two-stage framework first trains TSV on a small set of labeled exemplars to form compact and well-separated clusters.
It then augments the exemplar set with unlabeled LLM generations, employing an optimal transport-based algorithm for pseudo-labeling combined with a confidence-based filtering process.
Extensive experiments demonstrate that TSV achieves state-of-the-art performance with minimal labeled data, exhibiting strong generalization across datasets and providing a practical solution for real-world LLM applications. Seongheon Park, Xuefeng Du, Min-Hsuan Yeh, Haobo Wang 0001, Yixuan Li 0001 |
ICML | 3 |
| 2024 | CoCoLoFa: A Dataset of News Comments with Common Logical Fallacies Written by LLM-Assisted CrowdsabstractDetecting logical fallacies in texts can help users spot argument flaws, but automating this detection is not easy.Manually annotating fallacies in large-scale, real-world text data to create datasets for developing and validating detection models is costly.This paper introduces COCOLOFA, the largest known English logical fallacy dataset, containing 7,706 comments for 648 news articles, with each comment labeled for fallacy presence and type.We recruited 143 crowd workers to write comments embodying specific fallacy types (e.g., slippery slope) in response to news articles.Recognizing the complexity of this writing task, we built an LLM-powered assistant into the workers' interface to aid in drafting and refining their comments.Experts rated the writing quality and labeling validity of COCOLOFA as high and reliable.BERT-based models fine-tuned using COCOLOFA achieved the highest fallacy detection (F1=0.86)and classification (F1=0.87)performance on its test set, outperforming the stateof-the-art LLMs.Our work shows that combining crowdsourcing and LLMs enables us to more effectively construct datasets for complex linguistic phenomena that crowd workers find challenging to produce on their own.COCOLOFA is public at CoCoLoFa.org/. Min-Hsuan Yeh, Ruyuan Wan, Ting-Hao Huang |
EMNLP | 1 |
| 2022 | Multi-VQG: Generating Engaging Questions for Multiple ImagesabstractGenerating engaging content has drawn much recent attention in the NLP community.Asking questions is a natural way to respond to photos and promote awareness.However, most answers to questions in traditional questionanswering (QA) datasets are factoids, which reduce individuals' willingness to answer.Furthermore, traditional visual question generation (VQG) confines the source data for question generation to single images, resulting in a limited ability to comprehend time-series information of the underlying event.In this paper, we propose generating engaging questions from multiple images.We present MVQG 1 , a new dataset, and establish a series of baselines, including both end-to-end and dual-stage architectures.Results show that building stories behind the image sequence enables models to generate engaging questions, which confirms our assumption that people typically construct a picture of the event in their minds before asking questions.These results open up an exciting challenge for visual-and-language models to implicitly construct a story behind a series of photos to allow for creativity and experience sharing and hence draw attention to downstream applications.How would you act if you found yourself in a room filled with cans of free drinks?Have you ever gone to beer tastings and where would that be at?How long did the cat lounge around in the book room?What would this cat sit on next? Min-Hsuan Yeh, Ting-Hao 'Kenneth' Huang, Lun-Wei Ku |
EMNLP | 1 |
| 2021 | Lying Through One's Teeth: A Study on Verbal Leakage CuesabstractAlthough many studies use the LIWC lexicon to show the existence of verbal leakage cues in lie detection datasets, none mention how verbal leakage cues are influenced by means of data collection, or the impact thereof on the performance of models.In this paper, we study verbal leakage cues to understand the effect of the data construction method on their significance, and examine the relationship between such cues and models' validity.The LIWC word-category dominance scores of seven lie detection datasets are used to show that audio statements and lie-based annotations indicate a greater number of strong verbal leakage cue categories.Moreover, we evaluate the validity of state-of-the-art lie detection models with cross-and in-dataset testing.Results show that in both types of testing, models trained on a dataset with more strong verbal leakage cue categories-as opposed to only a greater number of strong cues-yield superior results, suggesting that verbal leakage cues are a key factor for selecting lie detection datasets. Min-Hsuan Yeh, Lun-Wei Ku |
EMNLP (1) | 1 |