VLDB 2026 Research / reviewers in the wild / expert
Renren Jin
dblp:329/4176
· DBLP profile ↗
16ranked-venue papers
1as first author
16since 2021 · last 2026
0009-0009-7452-9883ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language ModelsabstractDan Shi, Zhuowen Han, Simon Ostermann, Renren Jin, Josef Van Genabith, Deyi Xiong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dan Shi 0001, Zhuowen Han, Simon Ostermann 0002, Renren Jin, Josef van Genabith, Deyi Xiong |
ACL (1) | 4 |
| 2026 | From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for TibetanabstractLei Yang, Leiyu Pan, Bojian Xiong, Renren Jin, Shaowei Zhang, Yue Chen, Ling Shi, Jiang Zhou, Junru Wu, Zhen Wang, Jianxiang Peng, Juesi Xiao, Tianyu Dong, Zhuowen Han, Zhuo Chen, Yuqi Ren, Deyi Xiong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Leiyu Pan, Bojian Xiong, Renren Jin, Ling Shi 0004, Jianxiang Peng, Juesi Xiao, Tianyu Dong, Zhuowen Han, Yuqi Ren, Deyi Xiong |
ACL (1) | 4 |
| 2025 | Praetor: A Fine-Grained Generative LLM Evaluator with Instance-Level Customizable Evaluation CriteriaabstractYongqi Leng, Renren Jin, Yue Chen, Zhuowen Han, Ling Shi, Jianxiang Peng, Lei Yang, Juesi Xiao, Deyi Xiong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yongqi Leng, Renren Jin, Zhuowen Han, Ling Shi 0004, Jianxiang Peng, Juesi Xiao, Deyi Xiong |
ACL (1) | 2 |
| 2025 | How Do Personality Traits Affect LLM Performance on a Variety of Tasks?abstractLarge Language Models (LLMs) have demonstrated impressive performance across diverse natural language processing (NLP) and reasoning tasks, yet the influence of psychological factors, such as personality traits, on their capabilities remains underexplored. Drawing on the Big Five personality framework, we investigate how personality configurations affect LLM performance across five task categories: interdisciplinary expert knowledge, safety and harmfulness detection, code generation, mathematical reasoning, and scientific problem solving. We control personality traits using two approaches: prompt-based induction with expert-crafted prompts and low-rank adaptation (LoRA) fine-tuning on a newly constructed dataset of 20,000 personality-conditioned instructions. Experiments on seven LLMs reveal systematic, task-dependent effects of personality. High Neuroticism consistently degrades robustness and generalization, while high Conscientiousness improves stability, particularly in safety-critical contexts. Larger models show stronger resilience under personality perturbations, and prompt-based control achieves a better balance between trait alignment and task performance than LoRA fine-tuning. These findings highlight the trade-offs between controllability, stability, and generalization in personality-aware LLMs. Zheping Yu, Renren Jin, Tongxuan Zhang, Yuqi Ren, Guiyun Zhang |
BIBM | 3 |
| 2025 | Does Personality Shape AI Minds Like Humans? A Systematic Study on the Cognition and Behavior of Large Language ModelsabstractThe behavioral complexity of large language models (LLMs) has sparked growing interest in whether these models mirror human psychological traits. While prior work has explored the presence of personality in LLMs, most studies remain superficial, focus on linguistic style or isolated benchmarks, without probing deeper cognitive alignment with humans. In this paper, we present a comprehensive investigation into whether and how personality traits influence LLMs' behavior. Grounded in the Big Five personality traits, we introduce personalityconditioned prompts and evaluate their effects across both closed (e.g., reasoning, coding) and open-ended (e.g., writing) tasks. Our analysis spans task performance, linguistic style variation, alignment with human personality-ability correlations, and changes in internal reasoning structure. Experimental results reveal that: (i) personality traits affect LLM performance across all tasks, but only influence linguistic style in open-ended generation; (ii) personality-ability correlations in LLMs are broadly consistent with patterns observed in human psychology; and (iii) personality traits alter the behavior of LLMs, leading to different ways of reasoning across traits. Zheping Yu, Renren Jin, Tongxuan Zhang, Yuqi Ren, Guiyun Zhang |
BIBM | 3 |
| 2025 | CONTRANS: Weak-to-Strong Alignment Engineering via Concept TransplantationabstractEnsuring large language models (LLM) behave consistently with human goals, values, and intentions is crucial for their safety but yet computationally expensive. To reduce the computational cost of alignment training of LLMs, especially for those with a huge number of parameters, and to reutilize learned value alignment, we propose ConTrans, a novel framework that enables weak-to-strong alignment transfer via concept transplantation. From the perspective of representation engineering, ConTrans refines concept vectors in value alignment from a source LLM (usually a weak yet aligned LLM). The refined concept vectors are then reformulated to adapt to the target LLM (usually a strong yet unaligned base LLM) via affine transformation. In the third step, ConTrans transplants the reformulated concept vectors into the residual stream of the target LLM. Experiments demonstrate the successful transplantation of a wide range of aligned concepts from 7B models to 13B and 70B models across multiple LLMs and LLM families. Remarkably, ConTrans even surpasses instruction-tuned models in terms of truthfulness. Experiment results validate the effectiveness of both inter-LLM-family and intra-LLM-family concept transplantation. Our work successfully demonstrates an alternative way to achieve weak-to-strong alignment generalization and control. Weilong Dong, Xinwei Wu 0001, Renren Jin, Shaoyang Xu, Deyi Xiong |
COLING | 3 |
| 2025 | Empirical Study on Data Attributes Insufficiency of Evaluation Benchmarks for LLMsabstractPrevious benchmarks for evaluating large language models (LLMs) have primarily emphasized quantitative metrics, such as data volume. However, this focus may neglect key qualitative data attributes that can significantly impact the final rankings of LLMs, resulting in unreliable leaderboards. In this paper, we investigate whether current LLM benchmarks adequately consider these data attributes. We specifically examine three attributes: diversity, redundancy, and difficulty. To explore these attributes, we propose a framework with three separate modules, each designed to assess one of the attributes. Using a method that progressively incorporates these attributes, we analyze their influence on the benchmark. Our experimental results reveal a meaningful correlation between LLM rankings on the revised benchmark and the original benchmark when these attributes are accounted for. These findings indicate that existing benchmarks often fail to meet all three criteria, highlighting a lack of consideration for multifaceted data attributes in current evaluation datasets. Chuang Liu 0009, Renren Jin, Mark Steedman, Deyi Xiong |
COLING | 2 |
| 2025 | Do Large Language Models Mirror Cognitive Language Processing?abstractLarge Language Models (LLMs) have demonstrated remarkable abilities in text comprehension and logical reasoning, indicating that the text representations learned by LLMs can facilitate their language processing capabilities. In neuroscience, brain cognitive processing signals are typically utilized to study human language processing. Therefore, it is natural to ask how well the text embeddings from LLMs align with the brain cognitive processing signals, and how training strategies affect the LLM-brain alignment? In this paper, we employ Representational Similarity Analysis (RSA) to measure the alignment between 23 mainstream LLMs and fMRI signals of the brain to evaluate how effectively LLMs simulate cognitive language processing. We empirically investigate the impact of various factors (e.g., pre-training data size, model scaling, alignment training, and prompts) on such LLM-brain alignment. Experimental results indicate that pre-training data size and model scaling are positively correlated with LLM-brain similarity, and alignment training can significantly improve LLM-brain similarity. Explicit prompts contribute to the consistency of LLMs with brain cognitive language processing, while nonsensical noisy prompts may attenuate such alignment. Additionally, the performance of a wide range of LLM evaluations (e.g., MMLU, Chatbot Arena) is highly correlated with the LLM-brain similarity. Yuqi Ren, Renren Jin, Tongxuan Zhang, Deyi Xiong |
COLING | 2 |
| 2025 | Towards a Unified Paradigm of Concept Editing in Large Language ModelsabstractConcept editing aims to control specific concepts in large language models (LLMs) and is an emerging subfield of model editing.Despite the emergence of various editing methods in recent years, there remains a lack of rigorous theoretical analysis and a unified perspective to systematically understand and compare these methods.To address this gap, we propose a unified paradigm for concept editing methods, in which all forms of conceptual injection are aligned at the neuron level.We study four representative concept editing methods: Neuron Editing (NE), Supervised Fine-tuning (SFT), Sparse Autoencoder (SAE), and Steering Vector (SV).Then we categorize them into two classes based on their mode of conceptual information injection: indirect (NE, SFT) and direct (SAE, SV).We evaluate above methods along four dimensions: editing reliability, output generalization, neuron level consistency, and mathematical formalization.Experiments show that SAE achieves the best editing reliability.In output generalization, SAE captures features closer to human-understood concepts, while NE tends to locate text patterns rather than true semantics.Neuron-level analysis reveals that direct methods share high neuron overlap, as do indirect methods, indicating methodological commonality within each category.Our unified paradigm offers a clear framework and valuable insights for advancing interpretability and controlled generation in LLMs. Zhuowen Han, Xinwei Wu 0001, Dan Shi 0001, Renren Jin, Deyi Xiong |
EMNLP | 4 |
| 2025 | FineMath: A Fine-Grained Mathematical Evaluation Benchmark for Chinese Large Language ModelsabstractTo thoroughly assess the mathematical reasoning abilities of Large Language Models (LLMs), we need to carefully curate evaluation datasets covering diverse mathematical concepts and mathematical problems at different difficulty levels. In pursuit of this objective, we propose FineMath in this article, a fine-grained mathematical evaluation benchmark dataset for assessing Chinese LLMs. FineMath is created to cover the major key mathematical concepts taught in elementary school math, which are further divided into 17 categories of math word problems, enabling in-depth analysis of mathematical reasoning abilities of LLMs. All the 17 categories of math word problems are manually annotated with their difficulty levels according to the number of reasoning steps required to solve these problems. We conduct extensive experiments on a wide range of LLMs on FineMath and find that there is still considerable room for improvements in terms of mathematical reasoning capability of Chinese LLMs. We also carry out an in-depth analysis on the evaluation process and methods that have been overlooked previously. These two factors significantly influence the model results and our understanding of their mathematical reasoning capabilities. Our data is available at https://github.com/tjunlp-lab/FineMATH . Renren Jin, Zheng Yao 0004, Deyi Xiong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | LHMKE: A Large-scale Holistic Multi-subject Knowledge Evaluation Benchmark for Chinese Large Language ModelsabstractChinese Large Language Models (LLMs) have recently demonstrated impressive capabilities across various NLP benchmarks and real-world applications. However, the existing benchmarks for comprehensively evaluating these LLMs are still insufficient, particularly in terms of measuring knowledge that LLMs capture. Current datasets collect questions from Chinese examinations across different subjects and educational levels to address this issue. Yet, these benchmarks primarily focus on objective questions such as multiple-choice questions, leading to a lack of diversity in question types. To tackle this problem, we propose LHMKE, a Large-scale, Holistic, and Multi-subject Knowledge Evaluation benchmark in this paper. LHMKE is designed to provide a comprehensive evaluation of the knowledge acquisition capabilities of Chinese LLMs. It encompasses 10,465 questions across 75 tasks covering 30 subjects, ranging from primary school to professional certification exams. Notably, LHMKE includes both objective and subjective questions, offering a more holistic evaluation of the knowledge level of LLMs. We have assessed 11 Chinese LLMs under the zero-shot setting, which aligns with real examinations, and compared their performance across different subjects. We also conduct an in-depth analysis to check whether GPT-4 can automatically score subjective predictions. Our findings suggest that LHMKE is a challenging and advanced testbed for Chinese LLMs. Chuang Liu 0009, Renren Jin, Yuqi Ren, Deyi Xiong |
LREC/COLING | 2 |
| 2024 | IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware NeuronsabstractIt is widely acknowledged that large language models (LLMs) encode a vast reservoir of knowledge after being trained on mass data. Recent studies disclose knowledge conflicts in LLM generation, wherein outdated or incorrect parametric knowledge (i.e., encoded knowledge) contradicts new knowledge provided in the context. To mitigate such knowledge conflicts, we propose a novel framework, IRCAN (Identifying and Reweighting Context-Aware Neurons) to capitalize on neurons that are crucial in processing contextual cues. Specifically, IRCAN first identifies neurons that significantly contribute to context processing, utilizing a context-aware attribution score derived from integrated gradients. Subsequently, the identified context-aware neurons are strengthened via reweighting. In doing so, we steer LLMs to generate context-sensitive outputs with respect to the new knowledge provided in the context. Extensive experiments conducted across a variety of models and tasks demonstrate that IRCAN not only achieves remarkable improvements in handling knowledge conflicts but also offers a scalable, plug-and-play solution that can be integrated seamlessly with existing models. Our codes are released at https://github.com/danshi777/IRCAN. Dan Shi 0001, Renren Jin, Tianhao Shen, Weilong Dong, Xinwei Wu 0001, Deyi Xiong |
NeurIPS | 2 |
| 2024 | Star-Agents: Automatic Data Optimization with LLM Agents for Instruction TuningabstractThe efficacy of large language models (LLMs) on downstream tasks usually hinges on instruction tuning, which relies critically on the quality of training data. Unfortunately, collecting high-quality and diverse data is both expensive and time-consuming. To mitigate this issue, we propose a novel Star-Agents framework, which automates the enhancement of data quality across datasets through multi-agent collaboration and assessment. The framework adopts a three-pronged strategy. It initially generates diverse instruction data with multiple LLM agents through a bespoke sampling method. Subsequently, the generated data undergo a rigorous evaluation using a dual-model method that assesses both difficulty and quality. Finaly, the above process evolves in a dynamic refinement phase, where more effective LLMs are prioritized, enhancing the overall data quality. Our empirical studies, including instruction tuning experiments with models such as Pythia and LLaMA, demonstrate the effectiveness of the proposed framework. Optimized datasets have achieved substantial improvements, with an average increase of 12\% and notable gains in specific metrics, such as a 40\% improvement in Fermi, as evidenced by benchmarks like MT-bench, Vicuna bench, and WizardLM testset. Codes will be released soon. Yehui Tang 0001, Haochen Qin, Renren Jin, Deyi Xiong, Kai Han 0002, Yunhe Wang 0001 |
NeurIPS | 5 |
| 2023 | CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion TypesabstractSpoken texts (either manual or automatic transcriptions from automatic speech recognition (ASR)) often contain disfluencies and grammatical errors, which pose tremendous challenges to downstream tasks.Converting spoken into written language is hence desirable.Unfortunately, the availability of datasets for this is limited.To address this issue, we present CS2W, a Chinese Spoken-to-Written style conversion dataset comprising 7,237 spoken sentences extracted from transcribed conversational texts.Four types of conversion problems are covered in CS2W: disfluencies, grammatical errors, ASR transcription errors, and colloquial words.Our annotation convention, Zishan Guo, Linhao Yu, Renren Jin, Deyi Xiong |
EMNLP | 4 |
| 2023 | Joint Training and Decoding for Multilingual End-to-End Simultaneous Speech TranslationabstractRecent studies on end-to-end speech translation(ST) have facilitated the exploration of multilingual end-to-end ST and end-to-end simultaneous ST. In this paper, we investigate end-to-end simultaneous speech translation in a one-to-many multilingual setting which is closer to applications in real scenarios. We explore a separate decoder architecture and a unified architecture for joint synchronous training in this scenario. To further explore knowledge transfer across languages, we propose an asynchronous training strategy on the proposed unified decoder architecture. A multi-way aligned multilingual end-to-end ST dataset was curated as a benchmark testbed to evaluate our methods. Experimental results demonstrate the effectiveness of our models on the collected dataset. Our codes and data are available at: https://github.com/XiaoMi/TED-MMST. Wuwei Huang, Renren Jin, Wen Zhang 0015, Jian Luan 0001, Bin Wang 0004, Deyi Xiong |
ICASSP | 2 |
| 2022 | Informative Language Representation Learning for Massively Multilingual Neural Machine TranslationabstractIn a multilingual neural machine translation model that fully shares parameters across all languages, an artificial language token is usually used to guide translation into the desired target language. However, recent studies show that prepending language tokens sometimes fails to navigate the multilingual neural machine translation models into right translation directions, especially on zero-shot translation. To mitigate this issue, we propose two methods, language embedding embodiment and language-aware multi-head attention, to learn informative language representations to channel translation into right directions. The former embodies language embeddings into different critical switching points along the information flow from the source to the target, aiming at amplifying translation direction guiding signals. The latter exploits a matrix, instead of a vector, to represent a language in the continuous space. The matrix is chunked into multiple heads so as to learn language representations in multiple subspaces. Experiment results on two datasets for massively multilingual neural machine translation demonstrate that language-aware multi-head attention benefits both supervised and zero-shot translation and significantly alleviates the off-target translation issue. Further linguistic typology prediction experiments show that matrix-based language representations learned by our methods are capable of capturing rich linguistic typology features. Renren Jin, Deyi Xiong |
COLING | 1 |