VLDB 2026 Research / reviewers in the wild / expert
Xinliang Frederick Zhang
dblp:277/5381
· DBLP profile ↗
9ranked-venue papers
6as first author
9since 2021 · last 2026
0009-0001-4336-2189ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 68% Information extraction and text analysis · 14% Trustworthy machine learning · 10% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
1.0 | 1 | 2026 | Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model reasoning
efficient reasoning |
1.0 | 1 | 2026 | Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model reasoning
overthinking |
1.0 | 1 | 2026 | Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought Process · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
personalization |
0.9 | 1 | 2025 | PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought Process · EMNLP 2025 |
Machine learning › Trustworthy machine learning › fairness › bias evaluation › bias detection
media bias detection |
0.7 | 1 | 2023 | All Things Considered: Detecting Partisan Events from News Media with Cross-Article Comparison · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis
stance detection |
0.6 | 1 | 2022 | Generative Entity-to-Entity Stance Detection with Knowledge Graph Augmentation · EMNLP 2022 |
Information retrieval › question answering
FAQ retrieval |
0.5 | 1 | 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021 |
Information retrieval
question answering |
0.5 | 1 | 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › text classification
ideology detection |
0.2 | 1 | 2022 | Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis · EMNLP 2022 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.1 | 1 | 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021 |
Information retrieval
retrieval models |
0.1 | 1 | 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
reasoning trace analysis · 1.0slow thinking · 0.9dual-memory model · 0.9latent variable model · 0.7triplet margin objective · 0.6pre-training · 0.6late fusion · 0.6knowledge graph augmentation · 0.6graph encoder · 0.6generative framework · 0.6BM25 · 0.5BERT · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM OverthinkingabstractXinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang, Shyam Upadhyay. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang 0008, Shyam Upadhyay |
ACL (1) | 1 |
| 2025 | PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought ProcessabstractLarge language model (LLM) personalization aims to align model outputs with individuals' unique preferences and opinions.While recent efforts have implemented various personalization methods, a unified theoretical framework that can systematically understand the drivers of effective personalization is still lacking.In this work, we integrate the well-established cognitive dual-memory model into LLM personalization, by mirroring episodic memory to historical user engagements and semantic memory to long-term, evolving user beliefs.Specifically, we systematically investigate memory instantiations and introduce a unified framework, PRIME, using episodic and semantic memory mechanisms.We further augment PRIME with a novel personalized thinking capability inspired by the slow thinking strategy.Moreover, recognizing the absence of suitable benchmarks, we introduce a dataset using Change My View (CMV) from Reddit 1 , specifically designed to evaluate long-context personalization.Extensive experiments validate PRIME's effectiveness across both longand short-context scenarios.Further analysis confirms that PRIME effectively captures dynamic personalization beyond mere popularity biases. Xinliang Frederick Zhang, Nick Beauchamp, Lu Wang 0008 |
EMNLP | 1 |
| 2024 | MOKA: Moral Knowledge Augmentation for Moral Event ExtractionabstractXinliang Frederick Zhang, Winston Wu, Nick Beauchamp, Lu Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xinliang Frederick Zhang, Winston Wu, Nick Beauchamp, Lu Wang 0008 |
NAACL-HLT | 1 |
| 2023 | All Things Considered: Detecting Partisan Events from News Media with Cross-Article ComparisonabstractPublic opinion is shaped by the information news media provide, and that information in turn may be shaped by the ideological preferences of media outlets.But while much attention has been devoted to media bias via overt ideological language or topic selection, a more unobtrusive way in which the media shape opinion is via the strategic inclusion or omission of partisan events that may support one side or the other.We develop a latent variable-based framework to predict the ideology of news articles by comparing multiple articles on the same story and identifying partisan events whose inclusion or omission reveals ideology.Our experiments first validate the existence of partisan event selection, and then show that article alignment and cross-document comparison detect partisan events and article ideology better than competitive baselines.Our results reveal the high-level form of media bias, which is present even among mainstream media with strong norms of objectivity and nonpartisanship. Yujian Liu, Xinliang Frederick Zhang, Kaijian Zou, Ruihong Huang, Nick Beauchamp, Lu Wang 0008 |
EMNLP | 2 |
| 2022 | Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and AnalysisabstractPrior work on ideology prediction has largely focused on single modalities, i.e., text or images.In this work, we introduce the task of multimodal ideology prediction, where a model predicts binary or five-point scale ideological leanings, given a text-image pair with political content.We first collect five new large-scale datasets with English documents and images along with their ideological leanings, covering news articles from a wide range of mainstream media in US and social media posts from Reddit and Twitter.We conduct in-depth analyses on news articles and reveal differences in image content and usage across the political spectrum.Furthermore, we perform extensive experiments and ablation studies, demonstrating the effectiveness of targeted pretraining objectives on different model components.Our bestperforming model, a late-fusion architecture pretrained with a triplet objective over multimodal content, outperforms the state-of-the-art text-only model by almost 4% and a strong multimodal baseline with no pretraining by over 3%. Changyuan Qiu, Winston Wu, Xinliang Frederick Zhang, Lu Wang 0008 |
EMNLP | 3 |
| 2022 | Generative Entity-to-Entity Stance Detection with Knowledge Graph AugmentationabstractStance detection is typically framed as predicting the sentiment in a given text towards a target entity.However, this setup overlooks the importance of the source entity, i.e., who is expressing the opinion.In this paper, we emphasize the need for studying interactions among entities when inferring stances.We first introduce a new task, entity-to-entity (E2E) stance detection, which primes models to identify entities in their canonical names and discern stances jointly.To support this study, we curate a new dataset with 10,619 annotations labeled at the sentence-level from news articles of different ideological leanings.We present a novel generative framework to allow the generation of canonical names for entities as well as stances among them.We further enhance the model with a graph encoder to summarize entity activities and external knowledge surrounding the entities.Experiments show that our model outperforms strong comparisons by large margins.Further analyses demonstrate the usefulness of E2E stance detection for understanding media quotation and stance landscape, as well as inferring entity ideology. Xinliang Frederick Zhang, Nick Beauchamp, Lu Wang 0008 |
EMNLP | 1 |
| 2021 | CliniQG4QA: Generating Diverse Questions for Domain Adaptation of Clinical Question AnsweringabstractClinical question answering (QA) aims to automatically answer questions from medical professionals based on clinical texts. Studies show that neural QA models trained on one corpus may not generalize well to new clinical texts from a different institute or a different patient group, where largescale QA pairs are not readily available for model retraining. To address this challenge, we propose a simple yet effective framework, CliniQG4QA, which leverages question generation (QG) to synthesize QA pairs on new clinical contexts and boosts QA models without requiring manual annotations. In order to generate diverse types of questions that are essential for training QA models, we further introduce a seq2seq-based question phrase prediction (QPP) module that can be used together with most existing QG models to diversify the generation. Our comprehensive experiment results show that the QA corpus generated by our framework can improve QA models on the new contexts (up to 8% absolute gain in terms of Exact Match), and that the QPP module plays a crucial role in achieving the gain.11Our dataset and code are available at: https://github.com/sunlabosu/CliniQG4QA/. Xiang Yue, Xinliang Frederick Zhang, Ziyu Yao 0002, Simon M. Lin, Huan Sun 0001 |
BIBM | 2 |
| 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ RetrievalabstractWe present a large, challenging dataset, COUGH, for COVID-19 FAQ retrieval.Similar to a standard FAQ dataset, COUGH consists of three parts: FAQ Bank, Query Bank and Relevance Set.The FAQ Bank contains ∼16K FAQ items scraped from 55 credible websites (e.g., CDC and WHO).For evaluation, we introduce Query Bank and Relevance Set, where the former contains 1,236 human-paraphrased queries while the latter contains ∼32 humanannotated FAQ items for each query.We analyze COUGH by testing different FAQ retrieval models built on top of BM25 and BERT, among which the best model achieves 48.8 under P@5, indicating a great challenge presented by COUGH and encouraging future research for further improvement.Our COUGH dataset is available at https://github. com/sunlab-osu/covid-faq. *Work was done when the first two authors were at OSU. 1 q and a are question and answer fields in an FAQ item.Question1: Should children wear masks?Answer1: In general, children 2 years and older should wear a mask...Appropriate and consistent use of masks...FAQ Bank Question2: Coping with Self-Quarantine Answer2: Remind yourself that difficult emotions are normal during self-quarantine... Query1: Is it possible for human beings to get sick with COVID-19 transmitted to them from animals?Query2: Is it possible to get infected by COVID 19 if I touch food surface packaging?Query Bank Question3: COVID-19是如何在⼈与⼈之间传播的? (How does COVID-19 spread between people?) Answer3: . Xinliang Frederick Zhang, Heming Sun, Xiang Yue, Simon M. Lin, Huan Sun 0001 |
EMNLP (1) | 1 |
| 2021 | Identifying inherent disagreement in natural language inferenceabstractXinliang Frederick Zhang, Marie-Catherine de Marneffe. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xinliang Frederick Zhang, Marie-Catherine de Marneffe |
NAACL-HLT | 1 |