VLDB 2026 Research / reviewers in the wild / expert
Yuanyuan Lei 0001
dblp:283/5401-1
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-9753-8071ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 8 first-author · 10 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discovering a Shared Logical Subspace: Steering LLM Logical Reasoning via Alignment of Natural-Language and Symbolic ViewsabstractLarge Language Models (LLMs) still struggle with multi-step logical reasoning.Existing approaches either purely refine the reasoning chain in natural language form or attach a symbolic solver as an external module.In this work, we instead ask whether LLMs contain a shared internal logical subspace that simultaneously aligns natural-language and symbolic-language views of the reasoning process.Our hypothesis is that this logical subspace captures logical reasoning capabilities in LLMs that are shared across views while remaining independent of surface forms.To verify this, we employ Canonical Correlation Analysis on the paired residual activations from natural-language and symbolic-language reasoning chains, learning a low-dimensional subspace with maximum cross-view correlation.Furthermore, we design a training-free approach that steers LLMs reasoning chain along this logical subspace, thereby leveraging the complementary reasoning signals from both views.Experiments on four logical reasoning benchmarks demonstrate the effectiveness of our approach, improving accuracy by up to 11 percentage points and generalizing well on out-of-domain problems 1 . Feihao Fang, My T. Thai, Yuanyuan Lei 0001 |
ACL (1) | 3 |
| 2026 | Knowledge Vector of Logical Reasoning in Large Language ModelsabstractLogical reasoning serve as a central capability in LLMs and includes three main forms: deductive, inductive, and abductive reasoning.In this work, we study the knowledge representations of these reasoning types in LLMs and analyze the correlations among them.Our analysis shows that each form of logical reasoning can be captured as a reasoning-specific knowledge vector in a linear representation space, yet these vectors are largely independent of each other.Motivated by cognitive science theory that these subforms of logical reasoning interact closely in the human brain, as well as our observation that the reasoning process for one type can benefit from the reasoning chain produced by another, we further propose to refine the knowledge representations of each reasoning type in LLMs to encourage complementarity between them.To this end, we design a complementary subspace-constrained refinement framework, which introduces a complementary loss that enables each reasoning vector to leverage auxiliary knowledge from the others, and a subspace constraint loss that prevents erasure of their unique characteristics.Through steering experiments along reasoning vectors, we find that refined vectors incorporating complementary knowledge yield consistent performance gains.We also conduct a mechanisminterpretability analysis of each reasoning vector, revealing insights into the shared and specific features of different reasoning in LLMs 1 . Yuanyuan Lei 0001 |
ACL (1) | 2 |
| 2025 | Multi-document Summarization through Multi-document Event Relation Graph Reasoning in LLMs: a case study in Framing Bias MitigationabstractMedia outlets are becoming more partisan and polarized nowadays.Most previous work focused on detecting media bias.In this paper, we aim to mitigate media bias by generating a neutralized summary given multiple articles presenting different ideological views.Motivated by the critical role of events and event relations in media bias detection, we propose to increase awareness of bias in LLMs via multi-document events reasoning and use a multi-document event relation graph to guide the summarization process.This graph contains rich event information useful to reveal bias: four common types of in-doc event relations to reflect content framing bias, cross-doc event coreference relation to reveal content selection bias, and event-level moral opinions to highlight opinionated framing bias.We further develop two strategies to incorporate the multi-document event relation graph for neutralized summarization.Firstly, we convert a graph into natural language descriptions and feed the textualized graph into LLMs as a part of a hard text prompt.Secondly, we encode the graph with graph attention network and insert the graph embedding into LLMs as a soft prompt.Both automatic evaluation and human evaluation confirm that our approach effectively mitigates both lexical and informational media bias, and meanwhile improves content preservation 1 . Yuanyuan Lei 0001, Ruihong Huang |
ACL (1) | 1 |
| 2024 | Boosting Logical Fallacy Reasoning in LLMs via Logical Structure TreeabstractLogical fallacy uses invalid or faulty reasoning in the construction of a statement.Despite the prevalence and harmfulness of logical fallacies, detecting and classifying logical fallacies still remains a challenging task.We observe that logical fallacies often use connective words to indicate an intended logical relation between two arguments, while the argument semantics does not actually support the logical relation.Inspired by this observation, we propose to build a logical structure tree to explicitly represent and track the hierarchical logic flow among relation connectives and their arguments in a statement.Specifically, this logical structure tree is constructed in an unsupervised manner guided by the constituency tree and a taxonomy of connectives for ten common logical relations, with relation connectives as non-terminal nodes and textual arguments as terminal nodes, and the latter are mostly elementary discourse units.We further develop two strategies to incorporate the logical structure tree into LLMs for fallacy reasoning.Firstly, we transform the tree into natural language descriptions and feed the textualized tree into LLMs as a part of the hard text prompt.Secondly, we derive a relation-aware tree embedding and insert the tree embedding into LLMs as a soft prompt.Experiments on benchmark datasets demonstrate that our approach based on logical structure tree significantly improves precision and recall for both fallacy detection and fallacy classification 1 . Yuanyuan Lei 0001, Ruihong Huang |
EMNLP | 1 |
| 2024 | Sentence-level Media Bias Analysis with Event Relation GraphabstractYuanyuan Lei, Ruihong Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yuanyuan Lei 0001, Ruihong Huang |
NAACL-HLT | 1 |
| 2024 | EMONA: Event-level Moral Opinions in News ArticlesabstractYuanyuan Lei, Md Messal Monem Miah, Ayesha Qamar, Sai Ramana Reddy, Jonathan Tong, Haotian Xu, Ruihong Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yuanyuan Lei 0001, Md Messal Monem Miah, Ayesha Qamar, Sai Ramana Reddy, Jonathan Tong, Haotian Xu 0004, Ruihong Huang |
NAACL-HLT | 1 |
| 2024 | Polarity Calibration for Opinion SummarizationabstractYuanyuan Lei, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Ruihong Huang, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yuanyuan Lei 0001, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Ruihong Huang, Dong Yu 0001 |
NAACL-HLT | 1 |
| 2023 | Discourse Structures Guided Fine-grained Propaganda IdentificationabstractPropaganda is a form of deceptive narratives that instigate or mislead the public, usually with a political purpose.In this paper, we aim to identify propaganda in political news at two fine-grained levels: sentence-level and tokenlevel.We observe that propaganda content is more likely to be embedded in sentences that attribute causality or assert contrast to nearby sentences, as well as seen in opinionated evaluation, speculation and discussions of future expectation.Hence, we propose to incorporate both local and global discourse structures for propaganda discovery and construct two teacher models for identifying PDTB-style discourse relations between nearby sentences and common discourse roles of sentences in a news article respectively.We further devise two methods to incorporate the two types of discourse structures for propaganda identification by either using teacher predicted probabilities as additional features or soliciting guidance in a knowledge distillation framework.Experiments on the benchmark dataset demonstrate that leveraging guidance from discourse structures can significantly improve both precision and recall of propaganda content identification. 1 Yuanyuan Lei 0001, Ruihong Huang |
EMNLP | 1 |
| 2023 | Audio-Visual Emotion Recognition With Preference Learning Based on Intended and Multi-Modal Perceived LabelsabstractThis article introduces a novel preference learning framework that simultaneously considers both the intended and the perceived labels while addressing the mismatches between them. Based on analyzing the discrepancies and agreements between the intended and the perceived labels in different modalities of audio-only, visual-only, and audio-visual, as well as the consistency among the perceptual ratings of all raters, we propose three sets of pair-wise ranking rules to generate multi-scale relevant scores for preference learning, scaling from sketchy manner to detailed manner. Three ranking models with support vector machine (SVM), deep neural networks (DNN), and gradient boosting decision trees (GBDT) are developed. Our results demonstrate that all three preference learning models significantly outperform the conventional classifiers baselines, and the LambdaMART model with gradient boosting decision trees achieves the best performance. The improvement from the preference learning models confirm the benefits of complementary information provided by different types of labels. We also observe additional improvement from the detailed ‘complex ranking rules’, particular with the best LambdaMART model, which suggests that we should treat intended and perceived labels in single-model & multi-modal differently. We further discuss the complementary of different ranking models, and obtain the best overall accuracy of 85.06% on CREMA-D dataset when combining the two best ranking models–LambdaMART and RankNet–together, which is significantly better than the 76.19% accuracy attained by the baseline models. Finally, we perform the cross-corpus emotion recognition experiments by training emotion rankers on CREMA-D dataset and tested the ranking-based emotion classifier on the SAVEE dataset that do not have perceived labels annotated. Yuanyuan Lei 0001, Houwei Cao |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Cocktail Edge Caching: Ride Dynamic Trends of Content Popularity With Ensemble LearningabstractEdge caching will play a critical role in facilitating the emerging content-rich applications. However, it faces many new challenges, in particular, the highly dynamic content popularity and the heterogeneous caching configurations. In this paper, we propose Cocktail Edge Caching, that tackles the dynamic popularity and heterogeneity through ensemble learning. Instead of trying to find a single dominating caching policy for all the caching scenarios, we employ an ensemble of constituent caching policies and adaptively select the best-performing policy to control the cache. Towards this goal, we first show through formal analysis and experiments that different variations of the LFU and LRU policies have complementary performance in different caching scenarios. We further develop a novel caching algorithm that enhances LFU/LRU with deep recurrent neural network (LSTM) based time-series analysis. Finally, we develop a deep reinforcement learning agent that adaptively combines base caching policies according to their virtual hit ratios on parallel virtual caches. Through extensive experiments driven by real content requests from two large video streaming platforms, we demonstrate that CEC not only consistently outperforms all single policies, but also improves the robustness of them. CEC can be well generalized to different caching scenarios with low computation overheads for deployment. Tongyu Zong, Chen Li 0043, Yuanyuan Lei 0001, Houwei Cao, Yong Liu 0013 |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | Sentence-level Media Bias Analysis Informed by Discourse StructuresabstractAs polarization continues to rise among both the public and the news media, increasing attention has been devoted to detecting media bias.Most recent work in the NLP community, however, identify bias at the level of individual articles.However, each article itself comprises multiple sentences, which vary in their ideological bias.In this paper, we aim to identify sentences within an article that can illuminate and explain the overall bias of the entire article.We show that understanding the discourse role of a sentence in telling a news story, as well as its relation with nearby sentences, can reveal the ideological leanings of an author even when the sentence itself appears merely neutral.In particular, we consider using a functional news discourse structure and PDTB discourse relations to inform bias sentence identification, and distill the auxiliary knowledge from the two types of discourse structure into our bias sentence identification system.Experimental results on benchmark datasets show that incorporating both the global functional discourse structure and local rhetorical discourse relations can effectively increase the recall of bias sentence identification by 8.27% -8.62%, as well as increase the precision by 2.82% -3.48% 1 . Yuanyuan Lei 0001, Ruihong Huang, Lu Wang 0008, Nick Beauchamp |
EMNLP | 1 |
| 2021 | Cocktail Edge Caching: Ride Dynamic Trends of Content Popularity with Ensemble LearningabstractEdge caching will play a critical role in facilitating the emerging content-rich applications. However, it faces many new challenges, in particular, the highly dynamic content popularity and the heterogeneous caching configurations. In this paper, we propose Cocktail Edge Caching, that tackles the dynamic popularity and heterogeneity through ensemble learning. Instead of trying to find a single dominating caching policy for all the caching scenarios, we employ an ensemble of constituent caching policies and adaptively select the best-performing policy to control the cache. Towards this goal, we first show through formal analysis and experiments that different variations of the LFU and LRU polices have complementary performance in different caching scenarios. We further develop a novel caching algorithm that enhances LFU/LRU with deep recurrent neural network (LSTM) based time-series analysis. Finally, we develop a deep reinforcement learning agent that adaptively combines base caching policies according to their virtual hit ratios on parallel virtual caches. Through extensive experiments driven by real content requests from two large video streaming platforms, we demonstrate that CEC not only consistently outperforms all single policies, but also improves the robustness of them. CEC can be well generalized to different caching scenarios with low computation overheads for deployment. Tongyu Zong, Chen Li 0043, Yuanyuan Lei 0001, Houwei Cao, Yong Liu 0013 |
INFOCOM | 3 |