VLDB 2026 Research / reviewers in the wild / expert
Mingchen Cai
dblp:305/0416
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction TuningabstractVisual instruction tuning is crucial for enhancing the zero-shot generalization capability of Multi-modal Large Language Models (MLLMs). In this paper, we aim to investigate a fundamental question: “what makes for good visual instructions”. Through a comprehensive empirical study, we find that instructions focusing on complex visual reasoning tasks are particularly effective in improving the performance of MLLMs, with results correlating to instruction complexity. Based on this insight, we develop a systematic approach to automatically create high-quality complex visual reasoning instructions. Our approach employs a synthesize-complicate-reformulate paradigm, leveraging multiple stages to gradually increase the complexity of the instructions while guaranteeing quality. Based on this approach, we create the ComVint dataset with 32K examples, and fine-tune four MLLMs on it. Experimental results consistently demonstrate the enhanced performance of all compared MLLMs, such as a 27.86% and 27.60% improvement for LLaVA on MME-Perception and MME-Cognition, respectively. Our code and data are publicly available at the link: https://github.com/RUCAIBox/ComVint. Yifan Du 0002, Hangyu Guo, Kun Zhou 0002, Wayne Xin Zhao, Jinpeng Wang 0001, Chuyuan Wang, Mingchen Cai, Ruihua Song, Ji-Rong Wen |
COLING | 7 |
| 2025 | Label Feature Co-Learning for Facial and EEG Emotion RecognitionabstractRecognizing human emotions through behavioral and physiological signals is fundamental to overall health. However, since emotion occurs transiently, a semantics mismatch exists between the uniformly annotated label and multimodal temporal signals, leading to emotional ambiguity. Previous studies used the annotated label to guide feature learning, which makes it intractable to accurately identify emotional elicitation moments within each signal. Moreover, the inconsistency of specific elicitation moments across different signals complicates emotion recognition. The model hardly learns discriminative features due to emotional ambiguity, which weakens its ability to differentiate between emotions. To tackle the above challenges, this paper proposes a novel label feature co-learning model (LFCL) for emotion recognition through multimodal signals. Specifically, the LFCL leverages unimodal and multimodal information and adaptively generates instance-level emotion labels, thus precisely locating emotion elicitation moments within each signal. To promote emotion consistency across different signals, the LFCL incorporates a dynamic label calibration mechanism to balance the label generation process with historical information. Furthermore, to enhance the deep interaction between signals, the LFCL conducts multimodal interactive fusion to integrate multi-level multimodal features and extract global emotional information. The LFCL performs precise label-to-feature alignment to capture discriminative features of each signal, effectively alleviating emotional ambiguity and improving the ability to distinguish different emotions. Extensive experiments on three publicly available datasets demonstrate the effectiveness and generalization of the proposed model. Mingchen Cai, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | PREFER: Prompt Ensemble Learning via Feedback-Reflect-RefineabstractAs an effective tool for eliciting the power of Large Language Models (LLMs), prompting has recently demonstrated unprecedented abilities across a variety of complex tasks. To further improve the performance, prompt ensemble has attracted substantial interest for tackling the hallucination and instability of LLMs. However, existing methods usually adopt a two-stage paradigm, which requires a pre-prepared set of prompts with substantial manual effort, and is unable to perform directed optimization for different weak learners. In this paper, we propose a simple, universal, and automatic method named PREFER (Prompt Ensemble learning via Feedback-Reflect-Refine) to address the stated limitations. Specifically, given the fact that weak learners are supposed to focus on hard examples during boosting, PREFER builds a feedback mechanism for reflecting on the inadequacies of existing weak learners. Based on this, the LLM is required to automatically synthesize new prompts for iterative refinement. Moreover, to enhance stability of the prompt effect evaluation, we propose a novel prompt bagging method involving forward and backward thinking, which is superior to majority voting and is beneficial for both feedback and weight calculation in boosting. Extensive experiments demonstrate that our PREFER achieves state-of-the-art performance in multiple types of tasks by a significant margin. We have made our code publicly available. Chuyuan Wang, Jinpeng Wang 0001, Mingchen Cai |
AAAI | 7 |
| 2024 | Common Sense Enhanced Knowledge-based Recommendation with Large Language Model
Shenghao Yang 0004, Weizhi Ma, Peijie Sun, Min Zhang 0006, Qingyao Ai, Yiqun Liu 0001, Mingchen Cai |
DASFAA (5) | 7 |
| 2024 | Sequential Recommendation with Latent Relations based on Large Language ModelabstractSequential recommender systems predict items that may interest users by modeling their preferences based on historical interactions. Traditional sequential recommendation methods rely on capturing implicit collaborative filtering signals among items. Recent relation-aware sequential recommendation models have achieved promising performance by explicitly incorporating item relations into the modeling of user historical sequences, where most relations are extracted from knowledge graphs. However, existing methods rely on manually predefined relations and suffer the sparsity issue, limiting the generalization ability in diverse scenarios with varied item relations. Shenghao Yang 0004, Weizhi Ma, Peijie Sun, Qingyao Ai, Yiqun Liu 0001, Mingchen Cai, Min Zhang 0006 |
SIGIR | 6 |
| 2022 | Towards An Integrated Framework for Neural Temporal Point ProcessabstractTemporal point process (TPP) is a powerful tool to model the generation of irregular event sequences. With the rise of deep learning, neural TPPs have been proposed and made a great success. In neural TPPs, a state is extracted to characterize the influence of past events and updated accordingly when a new event is observed. Existing works can be categorized into two frameworks; each has pros and cons. Thus, we integrate these frameworks to overcome their shortages. To propose a model under the integrated framework, we first propose a novel and theoretically well-founded continuous-time layer normalized gated recurrent unit (CLNGRU) model, which is simpler, more robust, and more expressive than a pioneering work. Then we propose an enhanced version of CLNGRU (CLNGRU-E) under the integrated framework. We conduct experiments on synthetic and real datasets. Experimental results show the superiority of our models against the state-of-the-art methods on both generative and predictive tasks. Finally, we show the advantages of our models in recovering the underlying generative processes from sampled sequences. Qifu Hu, Zhoutian Feng, Daiyue Xue, Mingchen Cai |
IJCNN | 4 |
| 2021 | ST-PIL: Spatial-Temporal Periodic Interest Learning for Next Point-of-Interest RecommendationabstractPoint-of-Interest (POI) recommendation is an important task in location-based social networks. It facilitates the relation modeling between users and locations. Recently, researchers recommend POIs by long- and short-term interests and achieve success. However, they fail to well capture the periodic interest. People tend to visit similar places at similar times or in similar areas. Existing models try to acquire such kind of periodicity by user's mobility status or time slot, which limits the performance of periodic interest. To this end, we propose to learn spatial-temporal periodic interest. Specifically, in the long-term module, we learn the temporal periodic interest of daily granularity, then utilize intra-level attention to form long-term interest. In the short-term module, we construct various short-term sequences to acquire the spatial-temporal periodic interest of hourly, areal, and hourly-areal granularities, respectively. Finally, we apply inter-level attention to automatically integrate multiple interests. Experiments on two real-world datasets demonstrate the state-of-the-art performance of our method. Yafeng Zhang, Jinpeng Wang 0001, Mingchen Cai |
CIKM | 5 |