Jiankun Zhu

dblp:350/3375 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0009-2657-3062ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Augmenting and contrasting distortion for open panoramic segmentation
Sicheng Zhao, Jiankun Zhu, Xi Chen 0110, Hongxun Yao
Sci. China Inf. Sci.3
2026 SfMamba: Efficient source-free domain adaptation via selective scan modeling
Xi Chen 0110, Hongxun Yao, Sicheng Zhao, Jiankun Zhu, Kui Jiang
Expert Syst. Appl.4
2025 Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion Adaptation
abstract
Visual emotion recognition (VER), which aims at understanding humans' emotional reactions toward different visual stimuli, has attracted increasing attention. Given the subjective and ambiguous characteristics of emotion, annotating a reliable large-scale dataset is hard. For reducing reliance on data labeling, domain adaptation offers an alternative solution by adapting models trained on labeled source data to unlabeled target data. Conventional domain adaptation methods require access to source data. However, due to privacy concerns, source emotional data may be inaccessible. To address this issue, we propose an unexplored task: source-free domain adaptation (SFDA) for VER, which does not have access to source data during the adaptation process. To achieve this, we propose a novel framework termed Bridge then Begin Anew (BBA), which consists of two steps: domain-bridged model generation (DMG) and target-related model adaptation (TMA). First, the DMG bridges cross-domain gaps by generating an intermediate model, avoiding direct alignment between two VER datasets with significant differences. Then, the TMA begins training the target model anew to fit the target structure, avoiding the influence of source-specific knowledge. Extensive experiments are conducted on six SFDA settings for VER. The results demonstrate the effectiveness of BBA, which achieves remarkable performance gains compared with state-of-the-art SFDA methods and outperforms representative unsupervised domain adaptation approaches.
Jiankun Zhu, Sicheng Zhao, Wenbo Tang, Zhaopan Xu, Tingting Han 0003, Pengfei Xu 0001, Hongxun Yao
AAAI1
2025 Gaussian Constrained Diffeomorphic Deformation Network for Panoramic Semantic Segmentation
abstract
Panoramic semantic segmentation has garnered increasing attention due to its ability to provide comprehensive environmental perception. However, it requires a large number of annotated panoramic images to achieve satisfactory performance, which is costly. Recently, Domain Adaptation for Panoramic Semantic Segmentation (DA4PASS) has been proposed to reduce the reliance on annotated data by transferring segmentation models trained on annotated pinhole images to unlabelled panoramic images. Previous DA4PASS methods mainly focus on aligning features between pinhole and panoramic images, overlooking the unique appearance characteristics of panoramic images, particularly object distortion. To address the appearance discrepancies between pinhole and panoramic images, we propose Gaussian Constrained Diffeomorphic Deformation Network (GCDDN), which applies a panoramic deformation transformation obtained by Gaussian kernels to the annotated pinhole images. Specifically, GCDDN predicts multiple Gaussian kernels and performs first-order horizontal/vertical differences to obtain a naturally smooth and reversible panoramic deformation field, which is diffeomorphic. Due to its universality, GCDDN can be integrated into any domain adaptation (DA) method. Extensive experimental results demonstrate that integrating GCDDN leads to substantial improvements in both DA methods for pinhole images and those specifically designed for panoramic images, with a maximum gain of 1.80% in outdoor scenarios. Code is available at https://github.com/jingjiang02/GCDDN.
Jiankun Zhu, Zhaopan Xu, Xi Chen 0110, Sicheng Zhao, Hongxun Yao
ICASSP2
2025 Learning Class Prototypes for Visual Emotion Recognition
abstract
Visual emotion recognition (VER), which aims at understanding humans’ emotional reactions toward different visual stimuli, has attracted increasing attention. However, because of the subjectivity and complex nature of emotion, existing VER methods suffer from one or more of the following problems: 1) semantic gap: the large affective gap between visual clues and the emotional expressions; 2) overfitting: the lack of model robustness due to unclear features in the emotional category samples; 3) label ambiguity: the overlap between categories caused by diverse emotional responses. To address these limitations, we present a novel VER method named ProtoEmotion (PoE), exploring discriminative emotional representations by jointly learning prototypes of textual emotional expressions and visual features. Specifically, text prototypes build explicit textual features for each emotion category by extracting prototypes of learnable prompts from multiple aspects, reducing semantic differences. The visual prototypes capture the most defining image features of each category, providing a more robust and discriminative feature representation, while bringing samples closer together to reduce overfitting. In addition, to alleviate the label ambiguity, we propose a label smoothing algorithm based on the prototype distance. Extensive experiments demonstrate the effectiveness of PoE, which outperforms the state-of-the-art by 1.37% on FI and 1.52% on EmotionROI datasets.
Jiankun Zhu, Sicheng Zhao, Zhaopan Xu, Wenbo Tang, Hongxun Yao
ICASSP1
2025 DreamAnimate: Temporal Consistency and Detail Preservation for Character Animation
abstract
Character animation aims to generate realistic, high-quality videos from a reference image and target frames. However, existing methods struggle to balance fine-grained detail preservation with temporal consistency. This limitation results in artifacts like flickering and unrealistic deformations, especially in facial and hand regions. To address these challenges, we propose DreamAnimate, a novel framework that synthesizes temporally consistent and detail-rich animations. DreamAnimate integrates three modules: the Progressive Motion Estimation module ensures accurate motion alignment and temporal stability by refining keypoint heatmaps, the Global Affine Transformation module generates dense motion flows to handle complex motions and occlusions, and the Character Animation Fusion module combines intermediate synthesis using a UNet architecture and an Animation Fusion Network to produce high-quality animations. Extensive experiments demonstrate that DreamAnimate outperforms state-of-the-art methods, achieving superior fidelity and effectively capturing intricate facial expressions and hand movements. Code and models will be released at https://github.com/cslltian/DreamAnimate in the near future.
Lulu Tian, Hongxun Yao, Zhaopan Xu, Jiankun Zhu, Xi Chen 0110, Yuxin Hou
ICME4
2025 Emotion in a Bottle: Information Bottleneck Guided Disentanglement for Emotion Domain Adaptation
abstract
Visual emotion recognition (VER), which aims to understand human emotional reactions to different visual stimuli, has garnered increasing attention. However, the inherent ambiguity of emotional features presents significant challenges for data annotation in supervised learning paradigms. To address this limitation, emotion domain adaptation (EDA) facilitates knowledge transfer from labeled source domains to unlabeled target domains. Recently, large visual-language models such as CLIP have demonstrated impressive transfer performance on traditional UDA tasks. However, when generalizing to more abstract concepts such as emotion, the misalignment between CLIP and emotion spaces greatly affects the model performance. To address these challenges, we propose a CLIP-based emotion disentanglement (EmoD) framework designed for EDA. Leveraging perspectives from information bottleneck theory, EmoD implements a disentangler network that extracts emotion-specific features while removing redundant emotion-agnostic information. It also incorporates cross-domain feature alignment to reduce the affective gap between domains. Experimental evaluations in six EDA settings demonstrate that EmoD achieves state-of-the-art performance, surpassing traditional CLIP-based UDA methods by an average of 2.53%.
Jiankun Zhu, Sicheng Zhao, Lulu Tian, Xi Chen 0110, Hongxun Yao
ACM Multimedia1
2023 BRMR: TAL Based on Boundary Refinement and Multi-scale Regression
Jiankun Zhu, Lining Wang, Hongxun Yao
ICIG (4)2