Yixuan Weng

dblp:298/8205 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0002-9720-8689ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
abstract
Claim: This work is not advocating LLM replacement of human reviewers but rather exploring LLM
Minjun Zhu, Yixuan Weng, Linyi Yang, Yue Zhang 0004
ACL (1)2
2025 CycleResearcher: Improving Automated Research via Automated Review
abstract
The automation of scientific discovery has been a long-standing goal within the research community, driven by the potential to accelerate knowledge creation. While significant progress has been made using commercial large language models (LLMs) as research assistants or idea generators, the possibility of automating the entire research process with open-source LLMs remains largely unexplored. This paper explores the feasibility of using open-source post-trained LLMs as autonomous agents capable of performing the full cycle of automated research and review, from literature review and manuscript preparation to peer review and paper refinement. Our iterative preference training framework consists of CycleResearcher, which conducts research tasks, and CycleReviewer, which simulates the peer review process, providing iterative feedback via reinforcement learning. To train these models, we develop two new datasets, Review-5k and Research-14k, reflecting real-world machine learning research and peer review dynamics. Our results demonstrate that CycleReviewer achieves promising performance with a 26.89\% reduction in mean absolute error (MAE) compared to individual human reviewers in predicting paper scores, indicating the potential of LLMs to effectively assist expert-level research evaluation. In research, the papers generated by the CycleResearcher model achieved a score of 5.36 in simulated peer reviews, showing some competitiveness in terms of simulated review scores compared to the preprint level of 5.24 from human experts, while still having room for improvement compared to the accepted paper level of 5.69. This work represents a significant step toward fully automated scientific inquiry, providing ethical safeguards and exploring AI-driven research capabilities. The code, dataset and model weight are released at https://wengsyx.github.io/Researcher.
Yixuan Weng, Minjun Zhu, Guangsheng Bao, Jindong Wang 0001, Yue Zhang 0004, Linyi Yang
ICLR1
2025 Personality Alignment of Large Language Models
abstract
Aligning large language models (LLMs) typically aim to reflect general human values and behaviors, but they often fail to capture the unique characteristics and preferences of individual users. To address this gap, we introduce the concept of Personality Alignment. This approach tailors LLMs' responses and decisions to match the specific preferences of individual users or closely related groups. Inspired by psychometrics, we created the Personality Alignment with Personality Inventories (PAPI) dataset, which includes data from over 320,000 real subjects across multiple personality assessments - including both the Big Five Personality Factors and Dark Triad traits. This comprehensive dataset enables quantitative evaluation of LLMs' alignment capabilities across both positive and potentially problematic personality dimensions. Recognizing the challenges of personality alignments—such as limited personal data, diverse preferences, and scalability requirements—we developed an activation intervention optimization method. This method enhances LLMs' ability to efficiently align with individual behavioral preferences using minimal data and computational resources. Remarkably, our method, PAS, achieves superior performance while requiring only 1/5 of the optimization time compared to DPO, offering practical value for personality alignment. Our work paves the way for future AI systems to make decisions and reason in truly personality ways, enhancing the relevance and meaning of AI interactions for each user and advancing human-centered artificial intelligence. The dataset and code are released at https://github.com/zhu-minjun/PAlign.
Minjun Zhu, Yixuan Weng, Linyi Yang, Yue Zhang 0004
ICLR2
2025 Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
Bin Li 0083, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian 0002, Shoujun Zhou
NLPCC (4)3
2025 Small but mighty: enhancing time series forecasting with lightweight LLMs
Haoran Fan, Bin Li 0083, Yixuan Weng, Shoujun Zhou
J. Supercomput.3
2024 Does Knowledge Localization Hold True? Surprising Differences Between Entity and Relation Perspectives in Language Models
Yifan Wei 0001, Yixuan Weng, Huanhuan Ma, Yuanzhe Zhang, Jun Zhao 0001, Kang Liu 0001
CIKM3
2024 Towards Graph-hop Retrieval and Reasoning in Complex Question Answering over Textual Database
abstract
In textual question answering (TQA) systems, complex questions often require retrieving multiple textual fact chains with multiple reasoning steps. While existing benchmarks are limited to single-chain or single-hop retrieval scenarios. In this paper, we propose to conduct Graph-Hop —— a novel multi-chains and multi-hops retrieval and reasoning paradigm in complex question answering. We construct a new benchmark called ReasonGraphQA, which provides explicit and fine-grained evidence graphs for complex question to support comprehensive and detailed reasoning. In order to further study how graph-based evidential reasoning can be performed, we explore what form of Graph-Hop works best for generating textual evidence explanations in knowledge reasoning and question answering. We have thoroughly evaluated existing evidence retrieval and reasoning models on the ReasonGraphQA. Experiments highlight Graph-Hop is a promising direction for answering complex questions, but it still has certain limitations. We have further studied mitigation strategies to meet these challenges and discuss future directions.
Minjun Zhu, Yixuan Weng, Shizhu He, Kang Liu 0001, Yang jun Jun, Jun Zhao 0001
LREC/COLING2
2024 Mastering Symbolic Operations: Augmenting Language Models with Compiled Neural Networks
abstract
Language models' (LMs) proficiency in handling deterministic symbolic reasoning and rule-based tasks remains limited due to their dependency implicit learning on textual data. To endow LMs with genuine rule comprehension abilities, we propose "Neural Comprehension" - a framework that synergistically integrates compiled neural networks (CoNNs) into the standard transformer architecture. CoNNs are neural modules designed to explicitly encode rules through artificially generated attention weights. By incorporating CoNN modules, the Neural Comprehension framework enables LMs to accurately and robustly execute rule-intensive symbolic tasks. Extensive experiments demonstrate the superiority of our approach over existing techniques in terms of length generalization, efficiency, and interpretability for symbolic operations. Furthermore, it can be applied to LMs across different model scales, outperforming tool-calling methods in arithmetic reasoning tasks while maintaining superior inference efficiency. Our work highlights the potential of seamlessly unifying explicit rule learning via CoNNs and implicit pattern learning in LMs, paving the way for true symbolic comprehension capabilities. The code is released at: \url{https://github.com/wengsyx/Neural-Comprehension}.
Yixuan Weng, Minjun Zhu, Bin Li 0083, Shizhu He, Kang Liu 0001, Jun Zhao 0001
ICLR1
2024 Overview of the NLPCC 2024 Shared Task 7: Multi-lingual Medical Instructional Video Question Answering
Bin Li 0083, Yixuan Weng, Qiya Song, Lianhui Liang, Xianwen Min, Shoujun Zhou
NLPCC (5)2
2024 Large Language Models With Holistically Thought Could Be Better Doctors
Yixuan Weng, Bin Li 0083, Minjun Zhu, Bin Sun 0001, Shizhu He, Shengping Liu, Kang Liu 0001, Shutao Li 0001, Jun Zhao 0001
NLPCC (2)1
2024 Distinct but correct: generating diversified and entity-revised medical response
Bin Li 0083, Bin Sun 0001, Shutao Li 0001, Encheng Chen, Hongru Liu, Yixuan Weng, Yongping Bai, Meiling Hu
Sci. China Inf. Sci.6
2024 Towards better Chinese-centric neural machine translation for low-resource languages
Bin Li 0083, Yixuan Weng, Hanjun Deng
Comput. Speech Lang.2
2024 Towards Visual-Prompt Temporal Answer Grounding in Instructional Video
abstract
Temporal answer grounding in instructional video (TAGV) is a new task naturally derived from temporal sentence grounding in general video (TSGV). Given an untrimmed instructional video and a text question, this task aims at locating the frame span from the video that can semantically answer the question, i.e., visual answer. Existing methods tend to solve the TAGV problem with a visual span-based predictor, taking visual information to predict the start and end frames in the video. However, due to the weak correlations between the semantic features of the textual question and visual answer, current methods using the visual span-based predictor do not work well in the TAGV task. In this paper, we propose a visual-prompt text span localization (VPTSL) method, which introduces the timestamped subtitles for a text span-based predictor. Specifically, the visual prompt is a learnable feature embedding, which brings visual knowledge to the pre-trained language model. Meanwhile, the text span-based predictor learns joint semantic representations from the input text question, video subtitles, and visual prompt feature with the pre-trained language model. Thus, the TAGV is reformulated as the task of the visual-prompt subtitle span localization for the visual answer. Extensive experiments on five instructional video datasets, namely MedVidQA, TutorialVQA, VehicleVQA, CrossTalk and Coin, show that the proposed method outperforms several state-of-the-art (SOTA) methods by a large margin in terms of mIoU score, which demonstrates the effectiveness of the proposed visual prompt and text span-based predictor.
Shutao Li 0001, Bin Li 0083, Bin Sun 0001, Yixuan Weng
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Find Parent then Label Children: A Two-stage Taxonomy Completion Method with Pre-trained Language Model
abstract
Taxonomies, which organize domain concepts into hierarchical structures, are crucial for building knowledge systems and downstream applications.As domain knowledge evolves, taxonomies need to be continuously updated to include new concepts.Previous approaches have mainly focused on adding concepts to the leaf nodes of the existing hierarchical tree, which does not fully utilize the taxonomy's knowledge and is unable to update the original taxonomy structure (usually involving nonleaf nodes).In this paper, we propose a twostage method called ATTEMPT for taxonomy completion.Our method inserts new concepts into the correct position by finding a parent node and labeling child nodes.Specifically, by combining local nodes with prompts to generate natural sentences, we take advantage of pre-trained language models for hypernym/hyponymy recognition.Experimental results on two public datasets (including six domains) show that ATTEMPT performs best on both taxonomy completion and extension tasks, surpassing existing methods.
Yixuan Weng, Shizhu He, Kang Liu 0001, Jun Zhao 0001
EACL2
2023 Learning To Locate Visual Answer In Video Corpus Using Question
abstract
We introduce a new task, named video corpus visual answer localization (VCVAL), which aims to locate the visual answer in a large collection of untrimmed instructional videos using a natural language question. This task requires a range of skills - the interaction between vision and language, video retrieval, passage comprehension, and visual answer localization. In this paper, we propose a cross-modal contrastive global-span (CCGS) method for the VCVAL, jointly training the video corpus retrieval and visual answer localization subtasks with the global-span matrix. We have reconstructed a dataset named MedVidCQA, on which the VCVAL task is benchmarked. Experimental results show that the proposed method outperforms other competitive methods both in the video corpus retrieval and visual answer localization sub-tasks. Most importantly, we perform detailed analyses on extensive experiments, paving a new path for understanding the instructional videos, which ushers in further research1.
Bin Li 0083, Yixuan Weng, Bin Sun 0001, Shutao Li 0001
ICASSP2
2023 Visual Answer Localization with Cross-Modal Mutual Knowledge Transfer
abstract
The goal of visual answering localization (VAL) in the video is to obtain a relevant and concise time clip from a video as the answer to the given natural language question. Early methods are based on the interaction modelling between video and text to predict the visual answer by the visual predictor. Later, using the textual predictor with subtitles for the VAL proves to be more precise. However, these existing methods still have cross-modal knowledge deviations from visual frames or textual subtitles. In this paper, we propose a cross-modal mutual knowledge transfer span localization (MutualSL) method to reduce the knowledge deviation. MutualSL has both visual predictor and textual predictor, where we expect the prediction results of these both to be consistent, so as to promote semantic knowledge understanding between cross-modalities. On this basis, we design a one-way dynamic loss function to dynamically adjust the proportion of knowledge transfer. We have conducted extensive experiments on three public datasets for evaluation. The experimental results show that our method outperforms other competitive state-of-the-art (SOTA) methods, demonstrating its effectiveness1.
Yixuan Weng, Bin Li 0083
ICASSP1
2023 Learning to Build Reasoning Chains by Reliable Path Retrieval
abstract
Question answering (QA) systems have long pursued the ability to reason over explicit knowledge credibly. Recent work has incorporated knowledge into fine-grained sentences and constructed natural language database (NLDB) task, and conducts complex QA with explicit reasoning chains. Existing models focus on retrieving evidence by combining multiple modules or discretely. However, these models ignore utilizing path information (e.g. sentence order), which is proven to be important for evidence retrievers. In this work, we propose a ReliAble Path-retrieval (RAP) to generate varying length evidence chains iteratively. It comprehensively models reasoning chains and introduces loss from two views. The experimental results show that our model demonstrates state-of-the-art performance on both evidence chain retrieval and question-answering tasks. Additional experiments on sequential supervised and sequential unsupervised retrieval fully indicate the significance of RAP.
Minjun Zhu, Yixuan Weng, Shizhu He, Cunguang Wang, Kang Liu 0001, Jun Zhao 0001
ICASSP2
2023 Overview of the NLPCC 2023 Shared Task: Chinese Medical Instructional Video Question Answering
Bin Li 0083, Yixuan Weng, Hu Guo, Bin Sun 0001, Shutao Li 0001, Mengyao Qi, Xufei Liu, Yuwei Han, Haiwen Liang, Shuting Gao
NLPCC (3)2
2022 Scene-Aware Prompt for Multi-modal Dialogue Understanding and Generation
Bin Li 0083, Yixuan Weng, Ziyu Ma, Bin Sun 0001, Shutao Li 0001
NLPCC (2)2