VLDB 2026 Research / reviewers in the wild / expert
Linyi Yang
dblp:218/8007
· DBLP profile ↗
40ranked-venue papers
8as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 7 first-author · 30 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Introduction to the Special Issue on Evaluations of Large Language Models Part 2
Jindong Wang 0001, Linyi Yang, Sunayana Sitaram, Qiang Yang 0001, Bhiksha Raj |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2025 | DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking ProcessabstractClaim: This work is not advocating LLM replacement of human reviewers but rather exploring LLM Minjun Zhu, Yixuan Weng, Linyi Yang, Yue Zhang 0004 |
ACL (1) | 3 |
| 2025 | How Likely Do LLMs with CoT Mimic Human Reasoning?abstractChain-of-thought emerges as a promising technique for eliciting reasoning capabilities from Large Language Models (LLMs). However, it does not always improve task performance or accurately represent reasoning processes, leaving unresolved questions about its usage. In this paper, we diagnose the underlying mechanism by comparing the reasoning process of LLMs with humans, using causal analysis to understand the relationships between the problem instruction, reasoning, and the answer in LLMs. Our empirical study reveals that LLMs often deviate from the ideal causal chain, resulting in spurious correlations and potential consistency errors (inconsistent reasoning and answers). We also examine various factors influencing the causal structure, finding that in-context learning with examples strengthens it, while post-training techniques like supervised fine-tuning and reinforcement learning on human feedback weaken it. To our surprise, the causal structure cannot be strengthened by enlarging the model size only, urging research on new techniques. We hope that this preliminary study will shed light on understanding and improving the reasoning process in LLM. Guangsheng Bao, Cunxiang Wang, Linyi Yang, Yue Zhang 0004 |
COLING | 4 |
| 2025 | Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined ValuesabstractWe introduce Direct Value Optimization (DVO), an innovative reinforcement learning framework for enhancing large language models in complex reasoning tasks.Unlike traditional methods relying on preference labels, DVO utilizes value signals at individual reasoning steps, optimizing models via a mean squared error loss.The key benefit of DVO lies in its fine-grained supervision, circumventing the need for labor-intensive human annotations.Target values within the DVO are estimated using either Monte Carlo Tree Search or an outcome value model.Our empirical analysis on both mathematical and commonsense reasoning tasks shows that DVO consistently outperforms existing offline preference optimization techniques, even with fewer training steps.These findings underscore the importance of value signals in advancing reasoning capabilities and highlight DVO as a superior methodology under scenarios lacking explicit human preference information.The code is available at https://github.com/StevenZHB/DVO. Guangsheng Bao, Linyi Yang, Jun Wang 0012, Yue Zhang 0004 |
EMNLP | 4 |
| 2025 | CycleResearcher: Improving Automated Research via Automated ReviewabstractThe automation of scientific discovery has been a long-standing goal within the research community, driven by the potential to accelerate knowledge creation. While significant progress has been made using commercial large language models (LLMs) as research assistants or idea generators, the possibility of automating the entire research process with open-source LLMs remains largely unexplored. This paper explores the feasibility of using open-source post-trained LLMs as autonomous agents capable of performing the full cycle of automated research and review, from literature review and manuscript preparation to peer review and paper refinement. Our iterative preference training framework consists of CycleResearcher, which conducts research tasks, and CycleReviewer, which simulates the peer review process, providing iterative feedback via reinforcement learning. To train these models, we develop two new datasets, Review-5k and Research-14k, reflecting real-world machine learning research and peer review dynamics. Our results demonstrate that CycleReviewer achieves promising performance with a 26.89\% reduction in mean absolute error (MAE) compared to individual human reviewers in predicting paper scores, indicating the potential of LLMs to effectively assist expert-level research evaluation. In research, the papers generated by the CycleResearcher model achieved a score of 5.36 in simulated peer reviews, showing some competitiveness in terms of simulated review scores compared to the preprint level of 5.24 from human experts, while still having room for improvement compared to the accepted paper level of 5.69. This work represents a significant step toward fully automated scientific inquiry, providing ethical safeguards and exploring AI-driven research capabilities. The code, dataset and model weight are released at https://wengsyx.github.io/Researcher. Yixuan Weng, Minjun Zhu, Guangsheng Bao, Jindong Wang 0001, Yue Zhang 0004, Linyi Yang |
ICLR | 7 |
| 2025 | MMQA: Evaluating LLMs with Multi-Table Multi-Hop Complex QuestionsabstractWhile large language models (LLMs) have made strides in understanding tabular data, current tabular evaluation benchmarks, such as WikiTableQuestions and WikiSQL, are focus on single-table scenarios, which cannot necessarily reflect the complexity of real-world applications. To bridge this gap, we present a \textbf{M}ulti-table and
Multi-hop Question Answering (MMQA) dataset to assess LLMs' understanding and reasoning capabilities in handling multi-table tasks. The MMQA dataset demands that models perform multiple inferences by drawing evidence from various tables, which are designed to be connected with each other and require models to identify and utilize relationships such as foreign and primary keys. Then, we introduce a comprehensive evaluation framework that tailors to assess LLMs' capabilities in several aspects including Multi-Table Retrieval, Text-to-SQL Generation, Multi-Table QA, Primary Key Selection, and Foreign Key Selection.
Finally, we propose a novel multi-table retrieval method that achieves state-of-the-art (SOTA) performance on the MMQA dataset compared to several strong baselines.
Our experiment results reveal that, compared with human performance, both open-source and commercial LLMs leave significant performance room for improvements in multi-table understanding and reasoning tasks. We believe that the MMQA benchmark will enhance and facilitate LLMs' multi-table capabilities in real-world scenarios. Jian Wu 0037, Linyi Yang, Dongyuan Li, Yuliang Ji, Manabu Okumura, Yue Zhang 0004 |
ICLR | 2 |
| 2025 | CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkabstractWhile Large Language Models (LLMs) excel in question-answering (QA) tasks, their real reasoning abilities on multiple evidence retrieval and integration on Multi-hop QA tasks remain less explored. Firstly, LLMs sometimes generate answers that rely on internal memory rather than retrieving evidence and reasoning in the given context, which brings concerns about the evaluation quality of real reasoning abilities. Although previous counterfactual QA benchmarks can separate the internal memory of LLMs, they focus solely on final QA performance, which is insufficient for reporting LLMs' real reasoning abilities. Because LLMs are expected to engage in intricate reasoning processes that involve evidence retrieval and answering a series of sub-questions from given passages. Moreover, current factual Multi-hop QA (MHQA) benchmarks are annotated on open-source corpora such as Wikipedia, although useful for multi-step reasoning evaluation, they show limitations due to the potential data contamination in LLMs' pre-training stage. To address these issues, we introduce the Step-wise and Counterfactual benchmark (CofCA), a novel evaluation benchmark consisting of factual data and counterfactual data that reveals LLMs' real reasoning abilities on multi-step reasoning and reasoning chain evaluation. Our experimental results reveal a significant performance gap of several LLMs between Wikipedia-based factual data and counterfactual data, deeming data contamination issues in existing benchmarks. Moreover, we observe that LLMs usually bypass the correct reasoning chain, showing an inflated multi-step reasoning performance. We believe that our CofCA benchmark will enhance and facilitate the evaluations of trustworthy LLMs. Jian Wu 0037, Linyi Yang, Zhen Wang 0020, Manabu Okumura, Yue Zhang 0004 |
ICLR | 2 |
| 2025 | Human Simulacra: Benchmarking the Personification of Large Language ModelsabstractLarge Language Models (LLMs) are recognized as systems that closely mimic aspects of human intelligence. This capability has attracted the attention of the social science community, who see the potential in leveraging LLMs to replace human participants in experiments, thereby reducing research costs and complexity. In this paper, we introduce a benchmark for LLMs personification, including a strategy for constructing virtual characters' life stories from the ground up, a Multi-Agent Cognitive Mechanism capable of simulating human cognitive processes, and a psychology-guided evaluation method to assess human simulations from both self and observational perspectives. Experimental results demonstrate that our constructed simulacra can produce personified responses that align with their target characters. We hope this work will serve as a benchmark in the field of human simulation, paving the way for future research. Qiujie Xie, Qiming Feng, Qingqiu Li, Linyi Yang, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003, Yue Zhang 0004 |
ICLR | 5 |
| 2025 | An Empirical Analysis of Uncertainty in Large Language Model EvaluationsabstractAs LLM-as-a-Judge emerges as a new paradigm for assessing large language models (LLMs), concerns have been raised regarding the alignment, bias, and stability of LLM evaluators. While substantial work has focused on alignment and bias, little research has concentrated on the stability of LLM evaluators. In this paper, we conduct extensive experiments involving 9 widely used LLM evaluators across 2 different evaluation settings to investigate the uncertainty in model-based LLM evaluations. We pinpoint that LLM evaluators exhibit varying uncertainty based on model families and sizes. With careful comparative analyses, we find that employing special prompting strategies, whether during inference or post-training, can alleviate evaluation uncertainty to some extent. By utilizing uncertainty to enhance LLM's reliability and detection capability in Out-Of-Distribution (OOD) data, we further fine-tune an uncertainty-aware LLM evaluator named ConfiLM using a human-annotated fine-tuning set and assess ConfiLM's OOD evaluation ability on a manually designed test set sourced from the 2024 Olympics. Experimental results demonstrate that incorporating uncertainty as additional information during the fine-tuning phase can largely improve the model's evaluation performance in OOD scenarios. The code and data are released at: https://github.com/hasakiXie123/LLM-Evaluator-Uncertainty. Qiujie Xie, Qingqiu Li, Zhuohao Yu 0001, Yuejie Zhang, Yue Zhang 0004, Linyi Yang |
ICLR | 6 |
| 2025 | Personality Alignment of Large Language ModelsabstractAligning large language models (LLMs) typically aim to reflect general human values and behaviors, but they often fail to capture the unique characteristics and preferences of individual users. To address this gap, we introduce the concept of Personality Alignment. This approach tailors LLMs' responses and decisions to match the specific preferences of individual users or closely related groups. Inspired by psychometrics, we created the Personality Alignment with Personality Inventories (PAPI) dataset, which includes data from over 320,000 real subjects across multiple personality assessments - including both the Big Five Personality Factors and Dark Triad traits. This comprehensive dataset enables quantitative evaluation of LLMs' alignment capabilities across both positive and potentially problematic personality dimensions. Recognizing the challenges of personality alignments—such as limited personal data, diverse preferences, and scalability requirements—we developed an activation intervention optimization method. This method enhances LLMs' ability to efficiently align with individual behavioral preferences using minimal data and computational resources. Remarkably, our method, PAS, achieves superior performance while requiring only 1/5 of the optimization time compared to DPO, offering practical value for personality alignment. Our work paves the way for future AI systems to make decisions and reason in truly personality ways, enhancing the relevance and meaning of AI interactions for each user and advancing human-centered artificial intelligence. The dataset and code are released at https://github.com/zhu-minjun/PAlign. Minjun Zhu, Yixuan Weng, Linyi Yang, Yue Zhang 0004 |
ICLR | 3 |
| 2025 | Multi-Grained Alignment for Visual GroundingabstractVisual grounding aims to establish fine-grained alignment between specific regions and queries. Despite recent success, existing methods often struggle with two main issues. Firstly, using independently pre-trained uni-modal encoders leads to a significant semantic gap between extracted features, hindering the effective interaction of vision-language contexts. Secondly, attention-based approaches with a global receptive field often overlook local information within images, limiting the visual comprehension necessary to distinguish foreground from background and consequently leading to localization ambiguity. In this paper, we propose a Multi-Grained Alignment (MGA) framework for visual grounding, which addresses the semantic gap across and within modalities while enhancing localization performance through region-level and patch-level alignment. Extensive experimental results demonstrate that our approach outperforms state-of-the-art methods on three widely-used benchmarks. The code is accessible at https://github.com/Marloweeee/MGA-ICME. Hongbing Li, Linyi Yang |
ICME | 3 |
| 2025 | Constrain Alignment with Sparse AutoencodersabstractThe alignment of large language models (LLMs) with human preferences remains a key challenge. While post-training techniques like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) have achieved notable success, they often experience computational inefficiencies and training instability. In this paper, we propose Feature-level constrained Preference Optimization (FPO), a novel method designed to simplify the alignment process while ensuring stability. FPO leverages pre-trained Sparse Autoencoders (SAEs) and introduces feature-level constraints, allowing for efficient, sparsity-enforced alignment. Our approach enjoys efficiency by using sparse features activated in a well-trained sparse autoencoder and the quality of sequential KL divergence by using the feature-level offline reference. Experimental results on benchmark datasets demonstrate that FPO achieves an above 5% absolute improvement in win rate with much lower computational cost compared to state-of-the-art baselines, making it a promising solution for efficient and controllable LLM alignments. Qingyu Yin, Chak Tou Leong, Minjun Zhu, Hanqi Yan, Qiang Zhang 0026, Yulan He 0001, Wenjie Li 0002, Jun Wang 0012, Yue Zhang 0004, Linyi Yang |
ICML | 11 |
| 2025 | ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM ReasoningabstractEvaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2,912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces data contamination impact. Our data and codes are available at https://github.com/huangshulin123/ThinkBench. Shulin Huang, Linyi Yang, Yan Song 0003, Shawn Chen, Leyang Cui, Ziyu Wan, Qingcheng Zeng, Ying Wen 0001, Kun Shao, Weinan Zhang 0001, Jun Wang 0012, Yue Zhang 0004 |
NeurIPS | 2 |
| 2025 | ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement LearningabstractRecent research on Reasoning of Large Language Models (LLMs) has sought to further enhance their performance by integrating meta-thinking—enabling models to monitor, evaluate, and control their reasoning processes for more adaptive and effective problem-solving.
However, current single-agent work lacks a specialized design for acquiring meta-thinking, resulting in low efficacy.
To address this challenge, we introduce Reinforced Meta-thinking Agents (ReMA), a novel framework that leverages Multi-Agent Reinforcement Learning (MARL) to elicit meta-thinking behaviors, encouraging LLMs to think about thinking.
ReMA decouples the reasoning process into two hierarchical agents: a high-level meta-thinking agent responsible for generating strategic oversight and plans, and a low-level reasoning agent for detailed executions.
Through iterative reinforcement learning with aligned objectives, these agents explore and learn collaboration, leading to improved generalization and robustness.
Empirical results from single-turn experiments demonstrate that ReMA outperforms single-agent RL baselines on complex reasoning tasks, including competitive-level mathematical benchmarks and LLM-as-a-Judge benchmarks.
Additionally, we further extend ReMA to multi-turn interaction settings, leveraging turn-level ratio and parameter sharing to improve efficiency.
Comprehensive ablation studies further illustrate the evolving dynamics of each distinct agent, providing valuable insights into how the meta-thinking reasoning process enhances the reasoning capabilities of LLMs. Ziyu Wan, Xiaoyu Wen 0001, Yan Song 0003, Hanjing Wang, Linyi Yang, Mark Schmidt 0001, Jun Wang 0012, Weinan Zhang 0001, Shuyue Hu, Ying Wen 0001 |
NeurIPS | 6 |
| 2025 | Causal Sufficiency and Necessity Improves Chain-of-Thought ReasoningabstractChain-of-Thought (CoT) prompting plays an indispensable role in endowing large language models (LLMs) with complex reasoning capabilities. However, CoT currently faces two fundamental challenges: (1) Sufficiency, which ensures that the generated intermediate inference steps comprehensively cover and substantiate the final conclusion; and (2) Necessity, which identifies the inference steps that are truly indispensable for the soundness of the resulting answer. We propose a causal framework that characterizes CoT reasoning through the dual lenses of sufficiency and necessity. Incorporating causal Probability of Sufficiency and Necessity allows us not only to determine which steps are logically sufficient or necessary to the prediction outcome, but also to quantify their actual influence on the final reasoning outcome under different intervention scenarios, thereby enabling the automated addition of missing steps and the pruning of redundant ones. Extensive experimental results on various mathematical and commonsense reasoning benchmarks confirm substantial improvements in reasoning efficiency and reduced token usage without sacrificing accuracy. Our work provides a promising direction for improving LLM reasoning performance and cost-effectiveness. The code will be publicly available upon acceptance at: https://anonymous.4open.science/r/causalmath-1CEF. Xiangning Yu 0001, Zhuohan Wang, Linyi Yang, Haoxuan Li 0001, Anjie Liu, Xiao Xue 0001, Jun Wang 0012, Mengyue Yang |
NeurIPS | 3 |
| 2025 | SALA: Semantic alignment and localization alignment for visual groundingabstractVisual grounding focuses on establishing fine-grained alignment between specific regions and query expressions, which is increasingly essential as a cornerstone of visual intelligence. Despite recent success, existing methods often struggle with two main issues. Firstly, using independently pre-trained uni-modal encoders to extract expressive feature embeddings leads to a significant semantic gap between uni-modal features, hindering the effective interaction of visual-linguistic contexts. Secondly, the supervision provided by box annotations is inherently sparse and often underexploited, which limits the model’s ability to capture the fine-grained visual cues necessary to distinguish referent objects from the background, thereby leading to localization ambiguity. In this paper, we propose a Semantic Alignment and Localization Alignment (SALA) framework for visual grounding, which effectively bridges the cross- and uni-modal semantic gap and improves localization performance through patch-level and pixel-level alignment. This contributes to enhancing the consistency of representation before multimodal fusion, thereby improving the localization performance. Extensive experiments show that the proposed method outperforms state-of-the-art methods on five widely used datasets. Codes will be made publicly available after acceptance. Hongbing Li, Linyi Yang, Bo Xiao 0006 |
Neurocomputing | 3 |
| 2025 | Pre-Training a Graph Recurrent Network for Text UnderstandingabstractTransformer-based pre-trained models have gained much advance in recent years, Transformer architecture also becomes one of the most important backbones in natural language processing. Recent works show that the attention mechanism inside Transformer may not be necessary, and Transformer alternatives such as convolutional neural networks, multi-layer perceptron, and state space model have also been investigated. Transformer-based models have two main limitations: First, they have quadratic time complexity due to the full attention mechanism, which leads to high computational costs. Second, they rely on representation of a special token such as [CLS] to encode entire text, which limits its sentence-level expressiveness. In this paper, we consider a graph recurrent network with linear time complexity for language model pre-training, which builds a graph structure for each sequence with local token-level communications, together with a sentence-level representation detached from other normal tokens. On both English and Chinese text understanding tasks, our model can achieve comparable performance to existing pre-trained models while also achieving higher inference efficiency. Furthermore, we discovered that the representations generated by our model are more diverse and uniform compared to that of Transformer, which alleviates the problems in existing pre-trained models such as representation degradation. Yile Wang 0001, Linyi Yang, Zhiyang Teng, Ming Zhou 0001, Yue Zhang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Introduction to the Special Issue on Evaluations of Large Language Models: Part 1abstractNo abstract available. Jindong Wang 0001, Linyi Yang, Sunayana Sitaram, Qiang Yang 0001, Bhiksha Raj |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2024 | MAGE: Machine-generated Text Detection in the WildabstractYafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, Yue Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi 0001, Yue Zhang 0004 |
ACL (1) | 7 |
| 2024 | Detoxifying Large Language Models via Knowledge EditingabstractMengru Wang, Ningyu Zhang, Ziwen Xu, Zekun Xi, Shumin Deng, Yunzhi Yao, Qishen Zhang, Linyi Yang, Jindong Wang, Huajun Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ningyu Zhang 0001, Ziwen Xu, Zekun Xi, Shumin Deng, Yunzhi Yao, Qishen Zhang, Linyi Yang, Jindong Wang 0001, Huajun Chen |
ACL (1) | 8 |
| 2024 | Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability CurvatureabstractLarge language models (LLMs) have shown the ability to produce fluent and cogent content, presenting both productivity opportunities and societal risks. To build trustworthy AI systems, it is imperative to distinguish between machine-generated and human-authored content. The leading zero-shot detector, DetectGPT, showcases commendable performance but is marred by its intensive computational costs. In this paper, we introduce the concept of **conditional probability curvature** to elucidate discrepancies in word choices between LLMs and humans within a given context. Utilizing this curvature as a foundational metric, we present **Fast-DetectGPT**, an optimized zero-shot detector, which substitutes DetectGPT's perturbation step with a more efficient sampling step. Our evaluations on various datasets, source models, and test conditions indicate that Fast-DetectGPT not only surpasses DetectGPT by a relative around 75\% in both the white-box and black-box settings but also accelerates the detection process by a factor of 340, as detailed in Table 1. Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, Yue Zhang 0004 |
ICLR | 4 |
| 2024 | PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationabstractInstruction tuning large language models (LLMs) remains a challenging task, owing to the complexity of hyperparameter selection and the difficulty involved in evaluating the tuned models. To determine the optimal hyperparameters, an automatic, robust, and reliable evaluation benchmark is essential. However, establishing such a benchmark is not a trivial task due to the challenges associated with evaluation accuracy and privacy protection. In response to these challenges, we introduce a judge large language model, named PandaLM, which is trained to distinguish the superior model given several LLMs. PandaLM's focus extends beyond just the objective correctness of responses, which is the main focus of traditional evaluation datasets. It addresses vital subjective factors such as relative conciseness, clarity, adherence to instructions, comprehensiveness, and formality. To ensure the reliability of PandaLM, we collect a diverse human-annotated test dataset, where all contexts are generated by humans and labels are aligned with human preferences. Our findings reveal that PandaLM-7B offers a performance comparable to both GPT-3.5 and GPT-4. Impressively, PandaLM-70B surpasses their performance. PandaLM enables the evaluation of LLM to be fairer but with less cost, evidenced by significant improvements achieved by models tuned through PandaLM compared to their counterparts trained with default Alpaca's hyperparameters. In addition, PandaLM does not depend on API-based evaluations, thus avoiding potential data leakage. Yidong Wang 0003, Zhuohao Yu 0001, Wenjin Yao, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen 0102, Chaoya Jiang, Rui Xie 0003, Jindong Wang 0001, Xing Xie 0001, Wei Ye 0004, Shikun Zhang, Yue Zhang 0004 |
ICLR | 5 |
| 2024 | Supervised Knowledge Makes Large Language Models Better In-context LearnersabstractLarge Language Models (LLMs) exhibit emerging in-context learning abilities through prompt engineering. The recent progress in large-scale generative models has further expanded their use in real-world language applications. However, the critical challenge of improving the generalizability and factuality of LLMs in natural language understanding and question answering remains under-explored. While previous in-context learning research has focused on enhancing models to adhere to users' specific instructions and quality expectations, and to avoid undesired outputs, little to no work has explored the use of task-specific fine-tuned Language Models (SLMs) to improve LLMs' in-context learning during the inference stage. Our primary contribution is the establishment of a simple yet effective framework that enhances the reliability of LLMs as it: 1) generalizes out-of-distribution data, 2) elucidates how LLMs benefit from discriminative models, and 3) minimizes hallucinations in generative tasks. Using our proposed plug-in method, enhanced versions of Llama 2 and ChatGPT surpass their original versions regarding generalizability and factuality. We offer a comprehensive suite of resources, including 16 curated datasets, prompts, model checkpoints, and LLM outputs across 9 distinct tasks. Our empirical analysis sheds light on the advantages of incorporating discriminative models into LLMs and highlights the potential of our methodology in fostering more reliable LLMs. Linyi Yang, Shuibai Zhang, Zhuohao Yu 0001, Guangsheng Bao, Yidong Wang 0003, Jindong Wang 0001, Ruochen Xu, Wei Ye 0004, Xing Xie 0001, Weizhu Chen, Yue Zhang 0004 |
ICLR | 1 |
| 2024 | A Rationale-centric Counterfactual Data Augmentation Method for Cross-Document Event Coreference ResolutionabstractBowen Ding, Qingkai Min, Shengkun Ma, Yingjie Li, Linyi Yang, Yue Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Qingkai Min, Shengkun Ma, Yingjie Li 0008, Linyi Yang, Yue Zhang 0004 |
NAACL-HLT | 5 |
| 2024 | CulturePark: Boosting Cross-cultural Understanding in Large Language ModelsabstractCultural bias is pervasive in many large language models (LLMs), largely due to the deficiency of data representative of different cultures.
Typically, cultural datasets and benchmarks are constructed either by extracting subsets of existing datasets or by aggregating from platforms such as Wikipedia and social media.
However, these approaches are highly dependent on real-world data and human annotations, making them costly and difficult to scale.
Inspired by cognitive theories on social communication, this paper introduces CulturePark, an LLM-powered multi-agent communication framework for cultural data collection.
CulturePark simulates cross-cultural human communication with LLM-based agents playing roles in different cultures.
It generates high-quality cross-cultural dialogues encapsulating human beliefs, norms, and customs.
Using CulturePark, we generated 41,000 cultural samples to fine-tune eight culture-specific LLMs.
We evaluated these models across three downstream tasks: content moderation, cultural alignment, and cultural education.
Results show that for content moderation, our GPT-3.5-based models either match or outperform GPT-4 on $41$ datasets. Regarding cultural alignment, our models surpass GPT-4 on Hofstede's VSM 13 framework.
Furthermore, for cultural education of human participants, our models demonstrate superior outcomes in both learning efficacy and user experience compared to GPT-4. CulturePark proves an important step in addressing cultural bias and advancing the democratization of AI, highlighting the critical role of culturally inclusive data in model training. Code is released at https://github.com/Scarelette/CulturePark. Damien Teney, Linyi Yang, Qingsong Wen, Xing Xie 0001, Jindong Wang 0001 |
NeurIPS | 3 |
| 2024 | A Survey on Evaluation of Large Language ModelsabstractLarge language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate , where to evaluate , and how to evaluate . Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, education, natural and social sciences, agent applications, and other areas. Secondly, we answer the ‘where’ and ‘how’ questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing the performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey Yupeng Chang, Jindong Wang 0001, Yuan Wu 0002, Linyi Yang, Kaijie Zhu, Hao Chen 0102, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang 0003, Wei Ye 0004, Yue Zhang 0004, Yi Chang 0001, Philip S. Yu, Qiang Yang 0001, Xing Xie 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2024 | Masked Conditional Variational Autoencoders for Chromosome StraighteningabstractKaryotyping is of importance for detecting chromosomal aberrations in human disease. However, chromosomes easily appear curved in microscopic images, which prevents cytogeneticists from analyzing chromosome types. To address this issue, we propose a framework for chromosome straightening, which comprises a preliminary processing algorithm and a generative model called masked conditional variational autoencoders (MC-VAE). The processing method utilizes patch rearrangement to address the difficulty in erasing low degrees of curvature, providing reasonable preliminary results for the MC-VAE. The MC-VAE further straightens the results by leveraging chromosome patches conditioned on their curvatures to learn the mapping between banding patterns and conditions. During model training, we apply a masking strategy with a high masking ratio to train the MC-VAE with eliminated redundancy. This yields a non-trivial reconstruction task, allowing the model to effectively preserve chromosome banding patterns and structure details in the reconstructed results. Extensive experiments on three public datasets with two stain styles show that our framework surpasses the performance of state-of-the-art methods in retaining banding patterns and structure details. Compared to using real-world bent chromosomes, the use of high-quality straightened chromosomes generated by our proposed method can improve the performance of various deep learning models for chromosome classification by a large margin. Such a straightening approach has the potential to be combined with other karyotyping systems to assist cytogeneticists in chromosome analysis. Jingxiong Li, Sunyi Zheng, Zhongyi Shui, Shichuan Zhang, Linyi Yang, Yuxuan Sun 0002, Honglin Li 0001, Yuanxin Ye, Peter M. A. van Ooijen, Kang Li 0004, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Measuring Consistency in Text-based Financial Forecasting ModelsabstractFinancial forecasting has been an important and active area of machine learning research, as even the most modest advantage in predictive accuracy can be parlayed into significant financial gains.Recent advances in natural language processing (NLP) bring the opportunity to leverage textual data, such as earnings reports of publicly traded companies, to predict the return rate for an asset.However, when dealing with such a sensitive task, the consistency of models -their invariance under meaning-preserving alternations in input -is a crucial property for building user trust.Despite this, current financial forecasting methods do not consider consistency.To address this problem, we propose FinTrust, an evaluation tool that assesses logical consistency in financial text.Using FinTrust, we show that the consistency of state-of-the-art NLP models for financial forecasting is poor.Our analysis of the performance degradation caused by meaning-preserving alternations suggests that current text-based methods are not suitable for robustly predicting market information. Linyi Yang, Yingpeng Ma, Yue Zhang 0004 |
ACL (1) | 1 |
| 2023 | Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureabstractLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu, Yidong Wang, Jingming Zhuo, Lingqiao Liu, Jindong Wang, Jennifer Foster, Yue Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Linyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu, Jingming Zhuo, Lingqiao Liu, Jindong Wang 0001, Jennifer Foster, Yue Zhang 0004 |
EMNLP | 1 |
| 2023 | Graph-Based Video-Language Learning with Multi-Grained Audio-Visual AlignmentabstractVideo-language learning has attracted significant attention in the fields of multimedia, computer vision and natural language processing in recent years. One of the key challenges in this area is how to effectively integrate visual and linguistic information to enable machines to understand video content and query information. In this work, we leverage graph-based representations and multi-grained audio-visual alignment to address this challenge. First, our approach starts by transforming video and query inputs into visual-scene graphs and semantic role graphs using a visual-scene parser and semantic role labeler respectively. These graphs are then encoded using graph neural networks to obtain enriched representations and combined to obtain a video-query joint representation that enhances the semantic expressivity of the inputs. Second, to achieve accurate matching of relevant parts of audio and visual features, we propose a multi-grained alignment module that aligns the audio and visual features at multiple scales. This enables us to effectively fuse the audio and visual information in a way that is consistent with the semantic-level information captured by the graph-based representations. Experiments on five representative datasets collected for Video Retrieval and Video Question Answering tasks show that our approach outperforms the literature on several metrics. Our extensive ablation studies demonstrate the effectiveness of graph-based representation and multi-grained audio-visual alignment. Chenyang Lyu, Wenxi Li, Tianbo Ji, Longyue Wang, Liting Zhou, Cathal Gurrin, Linyi Yang, Yi Yu 0001, Yvette Graham, Jennifer Foster |
ACM Multimedia | 7 |
| 2023 | SciMine: An Efficient Systematic Prioritization Model Based on Richer Semantic InformationabstractSystematic review is a crucial method that has been widely used. by scholars from different research domains. However, screening for relevant scientific literature from paper candidates remains an extremely time-consuming process so the task of screening prioritization has been established to reduce the human workload. Various methods under the human-in-the-loop fashion are proposed to solve this task by using lexical features. These methods, even though achieving better performance than more sophisticated feature-based models such as BERT, omit rich and essential semantic information, therefore suffered from feature bias. In this study, we propose a novel framework SciMine to accelerate this screening process by capturing semantic feature representations from both background and the corpus. In particular, based on contextual representation learned from the pre-trained language models, our approach utilizes an autoencoder-based classifier and a feature-dependent classification module to extract general document-level and phrase-level information. Then a ranking ensemble strategy is used to combine these two complementary pieces of information. Experiments on five real-world datasets demonstrate that SciMine achieves state-of-the-art performance and comprehensive analysis further shows the efficacy of SciMine to solve feature bias. Linyi Yang, Yue Zhang 0004 |
SIGIR | 3 |
| 2022 | NumHTML: Numeric-Oriented Hierarchical Transformer Model for Multi-Task Financial ForecastingabstractFinancial forecasting has been an important and active area of machine learning research because of the challenges it presents and the potential rewards that even minor improvements in prediction accuracy or forecasting may entail. Traditionally, financial forecasting has heavily relied on quantitative indicators and metrics derived from structured financial statements. Earnings conference call data, including text and audio, is an important source of unstructured data that has been used for various prediction tasks using deep earning and related approaches. However, current deep learning-based methods are limited in the way that they deal with numeric data; numbers are typically treated as plain-text tokens without taking advantage of their underlying numeric structure. This paper describes a numeric-oriented hierarchical transformer model (NumHTML) to predict stock returns, and financial risk using multi-modal aligned earnings calls data by taking advantage of the different categories of numbers (monetary, temporal, percentages etc.) and their magnitude. We present the results of a comprehensive evaluation of NumHTML against several state-of-the-art baselines using a real-world publicly available dataset. The results indicate that NumHTML significantly outperforms the current state-of-the-art across a variety of evaluation metrics and that it has the potential to offer significant financial gains in a practical trading context. Linyi Yang, Jiazheng Li 0002, Ruihai Dong, Yue Zhang 0004, Barry Smyth |
AAAI | 1 |
| 2022 | A Rationale-Centric Framework for Human-in-the-loop Machine LearningabstractWe present a novel rationale-centric framework with human-in-the-loop -Rationales-centric Double-robustness Learning (RDL) -to boost model out-of-distribution performance in few-shot learning scenarios.By using static semi-factual generation and dynamic humanintervened correction, RDL exploits rationales (i.e.phrases that cause the prediction), human interventions and semi-factual augmentations to decouple spurious associations and bias models towards generally applicable underlying distributions, which enables fast and accurate generalisation.Experimental results show that RDL leads to significant prediction benefits on both in-distribution and out-of-distribution tests compared to many state-of-the-art benchmarks-especially for few-shot learning scenarios.We also perform extensive ablation studies to support in-depth analyses of each component in our framework. Jinghui Lu, Linyi Yang, Brian Mac Namee, Yue Zhang 0004 |
ACL (1) | 2 |
| 2022 | Human-in-the-loop Robotic Grasping Using BERT Scene RepresentationabstractCurrent NLP techniques have been greatly applied in different domains. In this paper, we propose a human-in-the-loop framework for robotic grasping in cluttered scenes, investigating a language interface to the grasping process, which allows the user to intervene by natural language commands. This framework is constructed on a state-of-the-art grasping baseline, where we substitute a scene-graph representation with a text representation of the scene using BERT. Experiments on both simulation and physical robot show that the proposed method outperforms conventional object-agnostic and scene-graph based methods in the literature. In addition, we find that with human intervention, performance can be significantly improved. Our dataset and code are available on our project website https://sites.google.com/view/hitl-grasping-bert. Yaoxian Song, Penglei Sun, Pengfei Fang, Linyi Yang, Yanghua Xiao, Yue Zhang 0004 |
COLING | 4 |
| 2022 | FactMix: Using a Few Labeled In-domain Examples to Generalize to Cross-domain Named Entity RecognitionabstractFew-shot Named Entity Recognition (NER) is imperative for entity tagging in limited resource domains and thus received proper attention in recent years. Existing approaches for few-shot NER are evaluated mainly under in-domain settings. In contrast, little is known about how these inherently faithful models perform in cross-domain NER using a few labeled in-domain examples. This paper proposes a two-step rationale-centric data augmentation method to improve the model’s generalization ability. Results on several datasets show that our model-agnostic method significantly improves the performance of cross-domain NER tasks compared to previous state-of-the-art methods compared to the counterfactual data augmentation and prompt-tuning methods. Linyi Yang, Lifan Yuan, Leyang Cui, Wenyang Gao, Yue Zhang 0004 |
COLING | 1 |
| 2022 | USB: A Unified Semi-supervised Learning Benchmark for ClassificationabstractSemi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural networks from scratch, which is time-consuming and environmentally unfriendly. To address the above issues, we construct a Unified SSL Benchmark (USB) for classification by selecting 15 diverse, challenging, and comprehensive tasks from CV, natural language processing (NLP), and audio processing (Audio), on which we systematically evaluate the dominant SSL methods, and also open-source a modular and extensible codebase for fair evaluation of these SSL methods. We further provide the pre-trained versions of the state-of-the-art neural models for CV tasks to make the cost affordable for further tuning. USB enables the evaluation of a single SSL algorithm on more tasks from multiple domains but with less cost. Specifically, on a single NVIDIA V100, only 39 GPU days are required to evaluate FixMatch on 15 tasks in USB while 335 GPU days (279 GPU days on 4 CV datasets except for ImageNet) are needed on 5 CV tasks with TorchSSL. Yidong Wang 0003, Hao Chen 0102, Wang Sun, Ran Tao 0013, Wenxin Hou, Linyi Yang, Zhi Zhou 0007, Lan-Zhe Guo, Heli Qi, Zhen Wu 0002, Yufeng Li 0008, Satoshi Nakamura 0001, Wei Ye 0004, Marios Savvides, Bhiksha Raj, Takahiro Shinozaki, Bernt Schiele, Jindong Wang 0001, Xing Xie 0001, Yue Zhang 0004 |
NeurIPS | 8 |
| 2021 | Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment AnalysisabstractLinyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang, Barry Smyth, Ruihai Dong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Linyi Yang, Jiazheng Li 0002, Padraig Cunningham, Yue Zhang 0004, Barry Smyth, Ruihai Dong |
ACL/IJCNLP (1) | 1 |
| 2020 | MAEC: A Multimodal Aligned Earnings Conference Call Dataset for Financial Risk PredictionabstractIn the area of natural language processing, various financial datasets have informed recent research and analysis including financial news, financial reports, social media, and audio data from earnings calls. We introduce a new, large-scale multi-modal, text-audio paired, earnings-call dataset named MAEC, based on S&P 1500 companies. We describe the main features of MAEC, how it was collected and assembled, paying particular attention to the text-audio alignment process used. We present the approach used in this work as providing a suitable framework for processing similar forms of data in the future. The resulting dataset is more than six times larger than those currently available to the research community and we discuss its potential in terms of current and future research challenges and opportunities. All resources of this work are available at https://github.com/Earnings-Call-Dataset/ Jiazheng Li 0002, Linyi Yang, Barry Smyth, Ruihai Dong |
CIKM | 2 |
| 2020 | Generating Plausible Counterfactual Explanations for Deep Transformers in Financial Text ClassificationabstractCorporate mergers and acquisitions (M&A) account for billions of dollars of investment globally every year, and offer an interesting and challenging domain for artificial intelligence.However, in these highly sensitive domains, it is crucial to not only have a highly robust and accurate model, but be able to generate useful explanations to garner a user's trust in the automated system.Regrettably, the recent research regarding eXplainable AI (XAI) in financial text classification has received little to no attention, and many current methods for generating textual-based explanations result in highly implausible explanations, which damage a user's trust in the system.To address these issues, this paper proposes a novel methodology for producing plausible counterfactual explanations, whilst exploring the regularization benefits of adversarial training on language models in the domain of FinTech.Exhaustive quantitative experiments demonstrate that not only does this approach improve the model accuracy when compared to the current stateof-the-art and human performance, but it also generates counterfactual explanations which are significantly more plausible based on human trials. Linyi Yang, Eoin M. Kenny, Tin Lok James Ng, Yi Yang 0042, Barry Smyth, Ruihai Dong |
COLING | 1 |
| 2020 | HTML: Hierarchical Transformer-based Multi-task Learning for Volatility PredictionabstractThe volatility forecasting task refers to predicting the amount of variability in the price of a financial asset over a certain period. It is an important mechanism for evaluating the risk associated with an asset and, as such, is of significant theoretical and practical importance in financial analysis. While classical approaches have framed this task as a time-series prediction one – using historical pricing as a guide to future risk forecasting – recent advances in natural language processing have seen researchers turn to complementary sources of data, such as analyst reports, social media, and even the audio data from earnings calls. This paper proposes a novel hierarchical, transformer, multi-task architecture designed to harness the text and audio data from quarterly earnings conference calls to predict future price volatility in the short and long term. This includes a comprehensive comparison to a variety of baselines, which demonstrates very significant improvements in prediction accuracy, in the range 17% - 49% compared to the current state-of-the-art. In addition, we describe the results of an ablation study to evaluate the relative contributions of each component of our approach and the relative contributions of text and audio data with respect to prediction accuracy. Linyi Yang, Tin Lok James Ng, Barry Smyth, Ruihai Dong |
WWW | 1 |