VLDB 2026 Research / reviewers in the wild / expert
Qifeng Dai
dblp:157/2769
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0002-2071-8506ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent InteractionabstractWith the rise of smart personal devices, service-oriented human-agent interactions have become increasingly prevalent. This trend highlights the need for personalized dialogue assistants that can understand user-specific traits to accurately interpret requirements and tailor responses to individual preferences. However, existing approaches often overlook the complexities of long-term interactions and fail to capture users’ subjective characteristics. To address these gaps, we present PAL-Bench, a new benchmark designed to evaluate the personalization capabilities of service-oriented assistants in long-term user-agent interactions. In the absence of available real-world data, we develop a multi-step LLM-based synthesis pipeline, which is further verified and refined by human annotators. This process yields PAL-Set, the first Chinese dataset comprising multi-session user logs and dialogue histories, which serves as the foundation for PAL-Bench. Furthermore, to improve personalized service-oriented interactions, we propose H2Memory, a hierarchical and heterogeneous memory framework that incorporates retrieval-augmented generation to improve personalized response generation. Comprehensive experiments on both our PAL-Bench and an external dataset demonstrate the effectiveness of the proposed memory framework. Zhaopei Huang, Qifeng Dai, Guozheng Wu, Xubin Li, Tiezheng Ge, Wenxuan Wang 0001, Qin Jin |
AAAI | 2 |
| 2025 | EmoHuman: Fine-Grained Emotion-Controlled Talking Head Generation via Audio-Text Multimodal DetanglingabstractAudio-driven talking head generation has made significant strides in creating realistic and lip-synchronized portraits. However, most existing approaches overlook facial expressions, with only a few attempting to model facial emotions explicitly, often leading to unnatural results. To address this gap, we introduce EmoHuman, an audio-to-video synthesis method that generates emotionally nuanced talking head videos without relying on intermediate 3D representations or facial landmarks. EmoHuman decouples content, emotion, and emotional intensity from the multimodal information of the audio and the corresponding textual content through an Audio Emotion Decoupling Module. The content features are used to drive a powerful video diffusion model, generating synchronized lip movements, while emotion and emotional intensity govern the simulation of facial expressions, resulting in more realistic video outputs. Extensive experiments demonstrate that EmoHuman outperforms state-of-the-art methods in image and video quality, expression correlation, and lip-synchronization accuracy. Qifeng Dai, Huidong Feng, Wendi Cui, Xinqi Cai, Yinglin Zheng, Ming Zeng 0008 |
ICMR | 1 |
| 2025 | Consistent Human Animation with Pseudo Multi-View Anchoring and Cross-Granularity Integration
Jintai Wang, Yinglin Zheng, Qifeng Dai, Ming Zeng 0008 |
ICMR | 4 |
| 2025 | PointHuman: Learning high-fidelity and generalizable human neural radiance fields using guidance of fine-grained semantics-enriched geometry
Jintai Wang, Huidong Feng, Qifeng Dai, Yinglin Zheng, Ming Zeng 0008 |
Comput. Graph. | 4 |
| 2024 | Multi-Modal Gait Recognition with Unidirectional Cross-modal AlignmentabstractGait recognition represents a pivotal challenge in visual signal comprehension, encompassing multi-modal spatial-temporal information, and exhibiting considerable complexity. The prevailing gait recognition methods are categorized into appearance-based and model-based, which commonly use silhouette and skeleton as input respectively. Certain recent studies have endeavored to integrate both of these modalities, achieving a certain degree of success in doing so. However, the aforementioned methods either rely solely on a modal of information or lack a profound consideration of the intricate interplay among multiple modal. Consequently, the potential inherent in gait’s spatial-temporal information remains untapped. To address this issue, this study strives to enhance the efficiency of data utilization through the integration of multi-modal information within a multi-modal fusion network. Furthermore, we propose a scheme aimed at enhancing the consistency between distinct modal features. Experiments on the widely used gait dataset CASIA-B have shown that our model has significantly improved under complex gait conditions, with overall performance reaching the most advanced level. Hengda Li, Yinglin Zheng, Qifeng Dai, Jintai Wang, Ming Zeng 0008 |
ICME | 3 |
| 2023 | Vertex position estimation with spatial-temporal transformer for 3D human reconstructionabstractReconstructing 3D human pose and body shape from monocular images or videos is a fundamental task for comprehending human dynamics. Frame-based methods can be broadly categorized into two fashions: those regressing parametric model parameters (e.g., SMPL) and those exploring alternative representations (e.g., volumetric shapes, 3D coordinates). Non-parametric representations have demonstrated superior performance due to their enhanced flexibility. However, when applied to video data, these non-parametric frame-based methods tend to generate inconsistent and unsmooth results. To this end, we present a novel approach that directly regresses the 3D coordinates of the mesh vertices and body joints with a spatial–temporal Transformer. In our method, we introduce a SpatioTemporal Learning Block (STLB) with Spatial Learning Module (SLM) and Temporal Learning Module (TLM), which leverages spatial and temporal information to model interactions at a finer granularity, specifically at the body token level. Our method outperforms previous state-of-the-art approaches on Human3.6M and 3DPW benchmark datasets. Xiangjun Zhang, Yinglin Zheng, Wenjin Deng, Qifeng Dai, Wangzheng Shi, Ming Zeng 0008 |
Graph. Model. | 4 |