VLDB 2026 Research / reviewers in the wild / expert
Weizhou Shen
dblp:245/3622
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-9180-0043ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement LearningabstractXuanyu Lei, Chenliang Li, Yuning Wu, Kaiming Liu, Weizhou Shen, Peng Li, Ming Yan, Fei Huang, Ya-Qin Zhang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xuanyu Lei, Chenliang Li 0003, Yuning Wu 0001, Kaiming Liu, Weizhou Shen, Peng Li 0030, Ming Yan 0008, Fei Huang 0002, Ya-Qin Zhang, Yang Liu 0005 |
ACL (1) | 5 |
| 2026 | MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingabstractFuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang, Chenliang Li, Weizhou Shen, Jiyue Guo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Fuwen Luo, Shengfeng Lou, Chi Chen 0005, Ziyue Wang 0002, Chenliang Li 0003, Weizhou Shen, Jiyue Guo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Yang Liu 0005 |
ACL (1) | 6 |
| 2025 | Mutual-Taught for Co-adapting Policy and Reward ModelsabstractTianyuan Shi, Canbin Huang, Fanqi Wan, Longguang Zhong, Ziyi Yang, Weizhou Shen, Xiaojun Quan, Ming Yan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tianyuan Shi, Canbin Huang, Fanqi Wan, Longguang Zhong, Weizhou Shen, Xiaojun Quan, Ming Yan 0008 |
ACL (1) | 6 |
| 2024 | Small LLMs Are Weak Tool Learners: A Multi-LLM AgentabstractLarge Language Model (LLM) agents significantly extend the capabilities of standalone LLMs, empowering them to interact with external tools (e.g., APIs, functions) and complete various tasks in a self-directed fashion.The challenge of tool use demands that LLMs not only understand user queries and generate answers accurately but also excel in task planning, tool invocation, and result summarization.While traditional works focus on training a single LLM with all these capabilities, performance limitations become apparent, particularly with smaller models.To overcome these challenges, we propose a novel approach that decomposes the aforementioned capabilities into a planner, caller, and summarizer.Each component is implemented by a single LLM that focuses on a specific capability and collaborates with others to accomplish the task.This modular framework facilitates individual updates and the potential use of smaller LLMs for building each capability.To effectively train this framework, we introduce a two-stage training paradigm.First, we fine-tune a backbone LLM on the entire dataset without discriminating sub-tasks, providing the model with a comprehensive understanding of the task.Second, the fine-tuned LLM is used to instantiate the planner, caller, and summarizer respectively, which are continually fine-tuned on respective sub-tasks.Evaluation across various tool-use benchmarks illustrates that our proposed multi-LLM framework surpasses the traditional single-LLM approach, highlighting its efficacy and advantages in tool learning. Weizhou Shen, Chenliang Li 0003, Hongzhan Chen, Ming Yan 0008, Xiaojun Quan, Hehong Chen, Ji Zhang 0011, Fei Huang 0002 |
EMNLP | 1 |
| 2024 | Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationabstractMobile device operation tasks are increasingly becoming a popular multi-modal AI application scenario. Current Multi-modal Large Language Models (MLLMs), constrained by their training data, lack the capability to function effectively as operation assistants. Instead, MLLM-based agents, which enhance capabilities through tool invocation, are gradually being applied to this scenario. However, the two major navigation challenges in mobile device operation tasks — task progress navigation and focus content navigation — are difficult to effectively solve under the single-agent architecture of existing work. This is due to the overly long token sequences and the interleaved text-image data format, which limit performance. To address these navigation challenges effectively, we propose Mobile-Agent-v2, a multi-agent architecture for mobile device operation assistance. The architecture comprises three agents: planning agent, decision agent, and reflection agent. The planning agent condenses lengthy, interleaved image-text history operations and screens summaries into a pure-text task progress, which is then passed on to the decision agent. This reduction in context length makes it easier for decision agent to navigate the task progress. To retain focus content, we design a memory unit that updates with task progress by decision agent. Additionally, to correct erroneous operations, the reflection agent observes the outcomes of each operation and handles any mistake accordingly. Experimental results indicate that Mobile-Agent-v2 achieves over a 30% improvement in task completion compared to the single-agent architecture of Mobile-Agent. The code is open-sourced at https://github.com/X-PLUG/MobileAgent. Junyang Wang 0001, Haiyang Xu 0001, Haitao Jia, Ming Yan 0008, Weizhou Shen, Ji Zhang 0011, Fei Huang 0002, Jitao Sang 0001 |
NeurIPS | 6 |
| 2024 | Multi-Party Conversation Modeling for Emotion RecognitionabstractMulti-party conversation modeling plays a vital role in emotion recognition in conversation (ERC). Aside from the intra- and inter-speaker dependencies between different speakers, the difficulty also lies in the fact that each conversation may contain several to many utterances that compose a long text sequence. In this article, we present two approaches to effective multi-party conversation modeling. First, to encode long sequences and capture long-range dependency between utterances, we introduce a dialog-oriented language model, DialogXL, with enhanced memory to store longer conversation sequences and dialog-aware self-attention to deal with multi-party dependencies. Second, we present a directed acyclic neural network, namely DAG-ERC, to encode the utterances with a directed acyclic graph (DAG) to better capture the intrinsic structure within a conversation. DAG-ERC combines the advantages of recurrent models and graph models and provides a more intuitive way to model information flow between sequential utterances. Extensive experiments are conducted on four ERC benchmarks with state-of-the-art models employed for comparison, and empirical results demonstrate the superiority of the two models in multi-party conversation modeling. Xiaojun Quan, Siyue Wu, Weizhou Shen, Jianxing Yu |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | Multi-Grained Knowledge Retrieval for End-to-End Task-Oriented DialogabstractRetrieving proper domain knowledge from an external database lies at the heart of end-toend task-oriented dialog systems to generate informative responses.Most existing systems blend knowledge retrieval with response generation and optimize them with direct supervision from reference responses, leading to suboptimal retrieval performance when the knowledge base becomes large-scale.To address this, we propose to decouple knowledge retrieval from response generation and introduce a multigrained knowledge retriever (MAKER) that includes an entity selector to search for relevant entities and an attribute selector to filter out irrelevant attributes.To train the retriever, we propose a novel distillation objective that derives supervision signals from the response generator.Experiments conducted on three standard benchmarks with both small and largescale knowledge bases demonstrate that our retriever performs knowledge retrieval more effectively than existing methods.Our code has been made publicly available. Fanqi Wan, Weizhou Shen, Xiaojun Quan, Wei Bi |
ACL (1) | 2 |
| 2023 | Retrieval-Generation Alignment for End-to-End Task-Oriented Dialogue SystemabstractDeveloping an efficient retriever to retrieve knowledge from a large-scale knowledge base (KB) is critical for task-oriented dialogue systems to effectively handle localized and specialized tasks.However, widely used generative models such as T5 and ChatGPT often struggle to differentiate subtle differences among the retrieved KB records when generating responses, resulting in suboptimal quality of generated responses.In this paper, we propose the application of maximal marginal likelihood to train a perceptive retriever by utilizing signals from response generation for supervision.In addition, our approach goes beyond considering solely retrieved entities and incorporates various meta knowledge to guide the generator, thus improving the utilization of knowledge.We evaluate our approach on three task-oriented dialogue datasets using T5 and ChatGPT as the backbone models.The results demonstrate that when combined with meta knowledge, the response generator can effectively leverage high-quality knowledge records from the retriever and enhance the quality of generated responses.The code of this work is available at https://github.com/shenwzh3/MK-TOD. Weizhou Shen, Yingqi Gao, Canbin Huang, Fanqi Wan, Xiaojun Quan, Wei Bi |
EMNLP | 1 |
| 2023 | Generic Dependency Modeling for Multi-Party ConversationabstractTo model the dependencies between utterances in multi-party conversations, we propose a simple and generic framework based on the dependency parsing results of utterances. Particularly, we present an approach to encoding the dependencies in the form of relative dependency encoding (ReDE) and illustrate how to implement it in Transformers by modifying the computation of self-attention. Experimental results on four multi-party conversation benchmarks show that this framework successfully boosts the general performance of two Transformer-based language models and leads to comparable or even superior performance compared to the state-of-the-art methods. The code is available at https://github.com/shenwzh3/ReDE. Weizhou Shen, Xiaojun Quan |
ICASSP | 1 |
| 2021 | DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion RecognitionabstractThis paper presents our pioneering effort for emotion recognition in conversation (ERC) with pre-trained language models. Unlike regular documents, conversational utterances appear alternately from different parties and are usually organized as hierarchical structures in previous work. Such structures are not conducive to the application of pre-trained language models such as XLNet. To address this issue, we propose an all-in-one XLNet model, namely DialogXL, with enhanced memory to store longer historical context and dialog-aware self-attention to deal with the multi-party structures. Specifically, we first modify the recurrence mechanism of XLNet from segment-level to utterance-level in order to better model the conversational data. Second, we introduce dialog-aware self-attention in replacement of the vanilla self-attention in XLNet to capture useful intra- and inter-speaker dependencies. Extensive experiments are conducted on four ERC benchmarks with mainstream models presented for comparison. The experimental results show that the proposed model outperforms the baselines on all the datasets. Several other experiments such as ablation study and error analysis are also conducted and the results confirm the role of the critical modules of DialogXL. Weizhou Shen, Xiaojun Quan, Zhixian Xie |
AAAI | 1 |
| 2021 | Directed Acyclic Graph Network for Conversational Emotion RecognitionabstractWeizhou Shen, Siyue Wu, Yunyi Yang, Xiaojun Quan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Weizhou Shen, Siyue Wu, Yunyi Yang, Xiaojun Quan |
ACL/IJCNLP (1) | 1 |
| 2020 | Relational Graph Attention Network for Aspect-based Sentiment AnalysisabstractAspect-based sentiment analysis aims to determine the sentiment polarity towards a specific aspect in online reviews.Most recent efforts adopt attention-based neural network models to implicitly connect aspects with opinion words.However, due to the complexity of language and the existence of multiple aspects in a single sentence, these models often confuse the connections.In this paper, we address this problem by means of effective encoding of syntax information.Firstly, we define a unified aspect-oriented dependency tree structure rooted at a target aspect by reshaping and pruning an ordinary dependency parse tree.Then, we propose a relational graph attention network (R-GAT) to encode the new tree structure for sentiment prediction.Extensive experiments are conducted on the SemEval 2014 and Twitter datasets, and the experimental results confirm that the connections between aspects and opinion words can be better established with our approach, and the performance of the graph attention network (GAT) is significantly improved as a consequence. Weizhou Shen, Yunyi Yang, Xiaojun Quan, Rui Wang 0005 |
ACL | 2 |
| 2020 | Constituency Lattice Encoding for Aspect Term ExtractionabstractOne of the remaining challenges for aspect term extraction in sentiment analysis resides in the extraction of phrase-level aspect terms, which is non-trivial to determine the boundaries of such terms. In this paper, we aim to address this issue by incorporating the span annotations of constituents of a sentence to leverage the syntactic information in neural network models. To this end, we first construct a constituency lattice structure based on the constituents of a constituency tree. Then, we present two approaches to encoding the constituency lattice using BiLSTM-CRF and BERT as the base models, respectively. We experimented on two benchmark datasets to evaluate the two models, and the results confirm their superiority with respective 3.17 and 1.35 points gained in F1-Measure over the current state of the art. The improvements justify the effectiveness of the constituency lattice for aspect term extraction. Yunyi Yang, Kun Li 0003, Xiaojun Quan, Weizhou Shen, Qinliang Su |
COLING | 4 |
| 2019 | Bayesian Deep Collaborative Matrix FactorizationabstractIn this paper, we propose a Bayesian Deep Collaborative Matrix Factorization (BDCMF) algorithm for collaborative filtering (CF). BDCMF is a novel Bayesian deep generative model that learns user and item latent vectors from users’ social interactions, contents of items as the auxiliary information and user-item rating (feedback) matrix. It alleviates the problem of matrix sparsity by incorporating items’ auxiliary and users’ social information into the model. It can learn more robust and dense latent representations by integrating deep learning into Bayesian probabilistic framework. As being one of deep generative models, it has both non-linearity and Bayesian nature. Additionally, in BDCMF, we derive an efficient EM-style point estimation algorithm for parameter learning. To further improve recommendation performance, we also derive a full Bayesian posterior estimation algorithm for inference. Experiments conducted on two sparse datasets show that BDCMF can significantly outperform the state-of-the-art CF methods. Teng Xiao, Shangsong Liang, Weizhou Shen, Zaiqiao Meng |
AAAI | 3 |