Xinliang Frederick Zhang

dblp:277/5381 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0009-0001-4336-2189ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 68% Information extraction and text analysis · 14% Trustworthy machine learning · 10%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model reasoning
efficient reasoning
1.012026
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model reasoning
overthinking
1.012026
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model
0.912025
PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought Process · EMNLP 2025
Natural language and speech › Language models and text generation › large language model › large language model adaptation
personalization
0.912025
PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought Process · EMNLP 2025
Machine learning › Trustworthy machine learning › fairness › bias evaluation › bias detection
media bias detection
0.712023
All Things Considered: Detecting Partisan Events from News Media with Cross-Article Comparison · EMNLP 2023
Natural language and speech › Information extraction and text analysis
stance detection
0.612022
Generative Entity-to-Entity Stance Detection with Knowledge Graph Augmentation · EMNLP 2022
Information retrieval › question answering
FAQ retrieval
0.512021
COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021
Information retrieval
question answering
0.512021
COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › text classification
ideology detection
0.212022
Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis · EMNLP 2022
Information retrieval › retrieval models › neural retrieval
dense retrieval
0.112021
COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021
Information retrieval
retrieval models
0.112021
COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

reasoning trace analysis · 1.0slow thinking · 0.9dual-memory model · 0.9latent variable model · 0.7triplet margin objective · 0.6pre-training · 0.6late fusion · 0.6knowledge graph augmentation · 0.6graph encoder · 0.6generative framework · 0.6BM25 · 0.5BERT · 0.5
YearPublicationVenuePosition
2026 Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
abstract
Xinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang, Shyam Upadhyay. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang 0008, Shyam Upadhyay
ACL (1)1
2025 PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought Process
abstract
Large language model (LLM) personalization aims to align model outputs with individuals' unique preferences and opinions.While recent efforts have implemented various personalization methods, a unified theoretical framework that can systematically understand the drivers of effective personalization is still lacking.In this work, we integrate the well-established cognitive dual-memory model into LLM personalization, by mirroring episodic memory to historical user engagements and semantic memory to long-term, evolving user beliefs.Specifically, we systematically investigate memory instantiations and introduce a unified framework, PRIME, using episodic and semantic memory mechanisms.We further augment PRIME with a novel personalized thinking capability inspired by the slow thinking strategy.Moreover, recognizing the absence of suitable benchmarks, we introduce a dataset using Change My View (CMV) from Reddit 1 , specifically designed to evaluate long-context personalization.Extensive experiments validate PRIME's effectiveness across both longand short-context scenarios.Further analysis confirms that PRIME effectively captures dynamic personalization beyond mere popularity biases.
Xinliang Frederick Zhang, Nick Beauchamp, Lu Wang 0008
EMNLP1
2024 MOKA: Moral Knowledge Augmentation for Moral Event Extraction
abstract
Xinliang Frederick Zhang, Winston Wu, Nick Beauchamp, Lu Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xinliang Frederick Zhang, Winston Wu, Nick Beauchamp, Lu Wang 0008
NAACL-HLT1
2023 All Things Considered: Detecting Partisan Events from News Media with Cross-Article Comparison
abstract
Public opinion is shaped by the information news media provide, and that information in turn may be shaped by the ideological preferences of media outlets.But while much attention has been devoted to media bias via overt ideological language or topic selection, a more unobtrusive way in which the media shape opinion is via the strategic inclusion or omission of partisan events that may support one side or the other.We develop a latent variable-based framework to predict the ideology of news articles by comparing multiple articles on the same story and identifying partisan events whose inclusion or omission reveals ideology.Our experiments first validate the existence of partisan event selection, and then show that article alignment and cross-document comparison detect partisan events and article ideology better than competitive baselines.Our results reveal the high-level form of media bias, which is present even among mainstream media with strong norms of objectivity and nonpartisanship.
Yujian Liu, Xinliang Frederick Zhang, Kaijian Zou, Ruihong Huang, Nick Beauchamp, Lu Wang 0008
EMNLP2
2022 Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis
abstract
Prior work on ideology prediction has largely focused on single modalities, i.e., text or images.In this work, we introduce the task of multimodal ideology prediction, where a model predicts binary or five-point scale ideological leanings, given a text-image pair with political content.We first collect five new large-scale datasets with English documents and images along with their ideological leanings, covering news articles from a wide range of mainstream media in US and social media posts from Reddit and Twitter.We conduct in-depth analyses on news articles and reveal differences in image content and usage across the political spectrum.Furthermore, we perform extensive experiments and ablation studies, demonstrating the effectiveness of targeted pretraining objectives on different model components.Our bestperforming model, a late-fusion architecture pretrained with a triplet objective over multimodal content, outperforms the state-of-the-art text-only model by almost 4% and a strong multimodal baseline with no pretraining by over 3%.
Changyuan Qiu, Winston Wu, Xinliang Frederick Zhang, Lu Wang 0008
EMNLP3
2022 Generative Entity-to-Entity Stance Detection with Knowledge Graph Augmentation
abstract
Stance detection is typically framed as predicting the sentiment in a given text towards a target entity.However, this setup overlooks the importance of the source entity, i.e., who is expressing the opinion.In this paper, we emphasize the need for studying interactions among entities when inferring stances.We first introduce a new task, entity-to-entity (E2E) stance detection, which primes models to identify entities in their canonical names and discern stances jointly.To support this study, we curate a new dataset with 10,619 annotations labeled at the sentence-level from news articles of different ideological leanings.We present a novel generative framework to allow the generation of canonical names for entities as well as stances among them.We further enhance the model with a graph encoder to summarize entity activities and external knowledge surrounding the entities.Experiments show that our model outperforms strong comparisons by large margins.Further analyses demonstrate the usefulness of E2E stance detection for understanding media quotation and stance landscape, as well as inferring entity ideology.
Xinliang Frederick Zhang, Nick Beauchamp, Lu Wang 0008
EMNLP1
2021 CliniQG4QA: Generating Diverse Questions for Domain Adaptation of Clinical Question Answering
abstract
Clinical question answering (QA) aims to automatically answer questions from medical professionals based on clinical texts. Studies show that neural QA models trained on one corpus may not generalize well to new clinical texts from a different institute or a different patient group, where largescale QA pairs are not readily available for model retraining. To address this challenge, we propose a simple yet effective framework, CliniQG4QA, which leverages question generation (QG) to synthesize QA pairs on new clinical contexts and boosts QA models without requiring manual annotations. In order to generate diverse types of questions that are essential for training QA models, we further introduce a seq2seq-based question phrase prediction (QPP) module that can be used together with most existing QG models to diversify the generation. Our comprehensive experiment results show that the QA corpus generated by our framework can improve QA models on the new contexts (up to 8% absolute gain in terms of Exact Match), and that the QPP module plays a crucial role in achieving the gain.11Our dataset and code are available at: https://github.com/sunlabosu/CliniQG4QA/.
Xiang Yue, Xinliang Frederick Zhang, Ziyu Yao 0002, Simon M. Lin, Huan Sun 0001
BIBM2
2021 COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval
abstract
We present a large, challenging dataset, COUGH, for COVID-19 FAQ retrieval.Similar to a standard FAQ dataset, COUGH consists of three parts: FAQ Bank, Query Bank and Relevance Set.The FAQ Bank contains ∼16K FAQ items scraped from 55 credible websites (e.g., CDC and WHO).For evaluation, we introduce Query Bank and Relevance Set, where the former contains 1,236 human-paraphrased queries while the latter contains ∼32 humanannotated FAQ items for each query.We analyze COUGH by testing different FAQ retrieval models built on top of BM25 and BERT, among which the best model achieves 48.8 under P@5, indicating a great challenge presented by COUGH and encouraging future research for further improvement.Our COUGH dataset is available at https://github. com/sunlab-osu/covid-faq. *Work was done when the first two authors were at OSU. 1 q and a are question and answer fields in an FAQ item.Question1: Should children wear masks?Answer1: In general, children 2 years and older should wear a mask...Appropriate and consistent use of masks...FAQ Bank Question2: Coping with Self-Quarantine Answer2: Remind yourself that difficult emotions are normal during self-quarantine... Query1: Is it possible for human beings to get sick with COVID-19 transmitted to them from animals?Query2: Is it possible to get infected by COVID 19 if I touch food surface packaging?Query Bank Question3: COVID-19是如何在⼈与⼈之间传播的? (How does COVID-19 spread between people?) Answer3: .
Xinliang Frederick Zhang, Heming Sun, Xiang Yue, Simon M. Lin, Huan Sun 0001
EMNLP (1)1
2021 Identifying inherent disagreement in natural language inference
abstract
Xinliang Frederick Zhang, Marie-Catherine de Marneffe. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xinliang Frederick Zhang, Marie-Catherine de Marneffe
NAACL-HLT1