Kaiwen Zheng 0002

dblp:229/1439-2 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0007-2516-8407ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Are Multimodal Embeddings Truly Beneficial for Recommendation? A Deep Dive into Whole vs. Individual Modalities
Junchen Fu, Yu Song 0007, Kaiwen Zheng 0002, Joemon M. Jose
ECIR (3)4
2026 Focal-RegionFace: Generating Fine-Grained Multi-attribute Descriptions for Arbitrarily Selected Face Focal Regions
abstract
Facial analysis is a fundamental problem in vision–language research, with important applications in affective computing. However, existing methods primarily focus on global facial attributes or single-dimension analysis, lacking fine-grained, interpretable multi-attribute modeling of arbitrary local facial regions. We introduce FaceFocalDesc, a new problem that aims to generate and recognize multi-attribute natural language descriptions for arbitrarily selected facial regions. The target attributes include facial action units, emotional states, and age. We argue that explicit region-level modeling enables more controllable and interpretable facial understanding. To support this task, we construct a new dataset with region-level annotations and corresponding language descriptions. We further propose Focal-RegionFace, a vision–language model fine-tuned from Qwen2.5-VL, which progressively refines its focus on localized facial features through multi-stage training. Experiments show that Focal-RegionFace achieves state-of-the-art performance on the proposed benchmark under both standard and newly introduced metrics, demonstrating its effectiveness in fine-grained region-focused facial analysis.
Kaiwen Zheng 0002, Junchen Fu, Songpei Xu, Yaoqin He, Joemon M. Jose, Hu Han 0001, Xuri Ge
ICMR1
2026 Benchmarking Multimodal Large Language Models for missing modality completion in product catalogues
Junchen Fu, Wenhao Deng 0004, Kaiwen Zheng 0002, Ioannis Arapakis, Yongxin Ni, Joemon M. Jose, Xuri Ge
Pattern Recognit.3
2025 Multimodal Representation Learning Techniques for Comprehensive Facial State Analysis
abstract
Multimodal foundation models have significantly improved feature representation by integrating information from multiple modalities, making them highly suitable for a broader set of applications. However, the exploration of multimodal facial representation for understanding perception has been limited. Understanding and analyzing facial states, such as Action Units (AUs) and emotions, require a comprehensive and robust framework that bridges visual and linguistic modalities. In this paper, we present a comprehensive pipeline for multimodal facial state analysis. First, we compile a new Multimodal Face Dataset (MFA) by generating detailed multilevel language descriptions of face, incorporating Action Unit (AU) and emotion descriptions, by leveraging GPT-4o. Second, we introduce a novel Multilevel Multimodal Face Foundation model (MF2) tailored for Action Unit (AU) and emotion recognition. Our model incorporates comprehensive visual feature modeling at both local and global levels of face image, enhancing its ability to represent detailed facial appearances. This design aligns visual representations with structured AU and emotion descriptions, ensuring effective cross-modal integration. Third, we develop a Decoupled Fine-Tuning Network (DFN) that efficiently adapts MF2across various tasks and datasets. This approach not only reduces computational overhead but also broadens the applicability of the foundation model to diverse scenarios. Experimentation show superior performance for AU and emotion detection tasks.
Kaiwen Zheng 0002, Xuri Ge, Junchen Fu, Joemon M. Jose
ICME1
2025 Double-Filter: Efficient Fine-tuning of Pre-trained Vision-Language Models via Patch&Layer Filtering
abstract
In this paper, we present a novel approach, termed Double-Filter,to “slim down” the fine-tuning process of vision-language pre-trained (VLP) models via filtering redundancies in feature inputs and architectural components. We enhance the fine-tuning process using two approaches. First, we develop a new patch selection method incorporating image patch filtering through background and foreground separation, followed by a refined patch selection process. Second, we design a genetic algorithm to eliminate redundant fine-grained architecture layers, improving the efficiency and effectiveness of the model. The former makes patch selection semantics more comprehensive, improving inference efficiency while ensuring semantic representation. The latter’s fine-grained layer filter removes architectural redundancy to the extent possible and mitigates the impact on performance. Experimental results demonstrate that the proposed Double-Filter achieves superior efficiency of model fine-tuning and maintains competitive performance compared with the advanced efficient fine-tuning methods on three downstream tasks, VQA, NLVR and Retrieval. In addition, it has been proven to be effective under METER and ViLT VLP models.
Yaoqin He, Junchen Fu, Kaiwen Zheng 0002, Songpei Xu, Fuhai Chen, Jie Li 0052, Joemon M. Jose, Xuri Ge
ICML3
2025 Efficient and Effective Adaptation of Multimodal Foundation Models in Sequential Recommendation
abstract
Multimodal foundation models (MFMs) have revolutionized sequential recommender systems through advanced representation learning. While Parameter-efficient Fine-tuning (PEFT) is commonly used to adapt these models, studies often prioritize parameter efficiency, neglecting GPU memory and training speed. To address this, we introduced the IISAN framework, significantly enhancing efficiency. However, IISAN was limited to symmetrical MFMs and identical text and image encoders, preventing the use of state-of-the-art Large Language Models. To overcome this, we developed IISAN-Versa, a versatile plug-and-play architecture compatible with both symmetrical and asymmetrical MFMs. IISAN-Versa employs a Decoupled PEFT structure and utilizes both intra- and inter-modal adaptation. It effectively handles asymmetry through a simple yet effective combination of group layer-dropping and dimension transformation alignment. Our research demonstrates that IISAN-Versa effectively adapts large text encoders, and we further identify a scaling effect where larger text encoders generally perform better. IISAN-Versa also demonstrates strong versatility in our defined multimodal scenarios, which include raw titles and captions generated from images and videos. Additionally, IISAN-Versa achieved state-of-the-art performance on the MicroLens public benchmark.
Junchen Fu, Xuri Ge, Xin Xin 0007, Alexandros Karatzoglou, Ioannis Arapakis, Kaiwen Zheng 0002, Yongxin Ni, Joemon M. Jose
IEEE Trans. Knowl. Data Eng.6