Gengyun Jia

dblp:246/5305 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-0513-138XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deep Orientational Representation Learning for Ordinal Regression
abstract
Ordinal regression aims to predict ordered classes. Existing methods mainly focus on label distribution shapes and feature distance relationships, while the directional characteristics in the representation space remain underexplored. In this paper, we propose deep orientational representation learning (ORL), aiming to ensure the trajectory of features sequentially connected by ordinal categories approximates a geodesic. We treat the output layer weights as ordinal prototypes and introduce two constraints, the co-directional constraint and the counter-directional constraint. They operate by constraining the angles between pairs of vectors. The former minimizes the angle between vectors with matching start and end categories, while the latter maximizes the angle between vectors whose start categories are the same but whose end categories are on opposite sides. The two constraints optimize the representation from different ordinal directions. ORL is extended to a multi-prototype setting (MORL) to mitigate misalignment between features and oriented prototypes caused by large intra-class variations. Theoretical analysis links ORL to distribution unimodality and distance orderliness, highlighting its advantages. The effectiveness of ORL (MORL) is demonstrated on various tasks including facial age estimation, historical image dating, and aesthetic quality assessment.
Gengyun Jia, Xin Ma 0031, Bing-Kun Bao
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Consistent and Controllable Image Animation with Motion Diffusion Models
abstract
Diffusion models have achieved significant progress in the task of image animation due to their powerful generative capabilities. However, preserving appearance consistency to the static input image, and avoiding abrupt motion change in the generated animation, remains challenging. In this paper, we introduce Cinemo, a novel image animation approach that aims at achieving better appearance consistency and motion smoothness. The core of Cinemo is to focus on learning the distribution of motion residuals, rather than directly predicting frames as in existing diffusion models. During the inference, we further mitigate the sudden motion changes in the generated video by introducing a novel DCT-based noise refinement strategy. To counteract the over-smoothing of motion, we introduce a dynamics degree control design for better control of the magnitude of motion. Altogether, these strategies enable Cinemo to produce highly consistent, smooth, and motion-controllable results. Extensive experiments compared with several state-of-the-art methods demonstrate the effectiveness and superiority of our proposed approach. In the end, we also demonstrate how our model can be applied for motion transfer or video editing of any given video. The project page is available at https://maxin-cn.github.io/cinemo_project/.
Xin Ma 0031, Yaohui Wang 0001, Gengyun Jia, Tien-Tsin Wong, Yuan-Fang Li, Cunjian Chen
CVPR3
2025 Causal Debiasing for Visual Commonsense Reasoning
abstract
Visual Commonsense Reasoning (VCR) refers to answering questions and providing explanations based on images. While existing methods achieve high prediction accuracy, they often overlook bias in datasets and lack debiasing strategies. In this paper, our analysis reveals co-occurrence and statistical biases in both textual and visual data. We introduce the VCR-OOD datasets, comprising VCR-OOD-QA and VCR-OOD-VA subsets, which are designed to evaluate the generalization capabilities of models across two modalities. Furthermore, we analyze the causal graphs and prediction shortcuts in VCR and adopt a backdoor adjustment method to remove bias. Specifically, we create a dictionary based on the set of correct answers to eliminate prediction shortcuts. Experiments demonstrate the effectiveness of our debiasing method across different datasets.
Jiayi Zou, Gengyun Jia, Bing-Kun Bao
ICASSP2
2025 Relation Inference Enhancement Network for Visual Commonsense Reasoning
abstract
When presented with a question regarding an image, Visual Commonsense Reasoning (VCR) offers not only a correct answer but also a rationale to justify the answer. Existing methods simply combine features from multiple modalities onto a shared dimension space, which doesn't align with human reasoning patterns, resulting in inadequate cross-modal and intra-modal reasoning behaviors. On the one hand, inadequate cross-modal reasoning arises from existing models relying on semantic correlations between answers and rationales in both textual modalities rather than the generative process of human reasoning from visual to textual modality. On the other hand, inadequate intra-modal reasoning arises from the incapacity of existing models to leverage previously acquired object relations beyond current observations like humans. To this end, we propose a novel Relation Inference Enhancement Network (RIE-Net), which enhances reasoning ability based on cross-modal image analysis and introduces intra-modal relational reasoning modules to memorize reasoning knowledge. To enhance the cross-modal association between images and rationales, RIE-Net introduces a cross-modal image analysis module, which eliminates language bias between answers and rationales by generating rationale from images. In addition, to comprehend and retain relational knowledge, RIE-Net introduces intra-modal relational reasoning modules to capture prior knowledge associated with various object categories and enhance the model's understanding of visual-spatial relationships. Quantitative and qualitative evaluations of the public VCR dataset demonstrate that our approach performs favorably against state-of-the-art methods.
Gengyun Jia, Bing-Kun Bao
IEEE Trans. Multim.2
2024 Uncertainty-aware image inpainting with adaptive feedback network
abstract
While most image inpainting methods perform well on small image defects, they still struggle to deliver satisfactory results on large holes due to insufficient image guidance. To address this challenge, this paper proposes an uncertainty-aware adaptive feedback network (U2AFN), which incorporates an adaptive feedback mechanism to refine inpainting regions progressively. U2AFN predicts both an uncertainty map and an inpainting result simultaneously. During each iteration, the adaptive integration feedback block utilizes inpainting pixels with low uncertainty to guide the subsequent learning iteration. This process leads to a gradual reduction in uncertainty and produces more reliable inpainting outcomes. Our approach is extensively evaluated and compared on multiple datasets, demonstrating its superior performance over existing methods. The code is available at: https://codeocean.com/capsule/1901983/tree.
Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Gengyun Jia, Yaohui Wang 0001, Cunjian Chen
Expert Syst. Appl.4
2024 GPT-Based Knowledge Guiding Network for Commonsense Video Captioning
abstract
Video-based commonsense captioning aims to generate captions for the video content while providing multiple commonsense about the underlying event. Existing methods utilize video features to explore and generate commonsense containing latent semantics. However, this process needs to overcome the complex semantic gap between visible videos and invisible commonsense, which is not supported by the limited knowledge in existing video captioning datasets. To this end, we propose a novel GPT-based Two-stage Knowledge Guiding Network (TKG-Net), which uses GPT to augment datasets knowledge and introduces a cross-attention mechanism to fuse multimodal knowledge. Specifically, to augment knowledge, we set prompts and finetune GPT to imagine and reason based on the video content description at the first stage. At the second stage, to prevent over-reasoning caused by the loss of visual features in GPT, TKG-Net extracts high-level semantic representations of commonsense knowledge and fuses them with video features in a cross-attention mechanism for multimodal semantic interaction. Our experiments on the large-scale Video-to-Commonsense dataset manifest significant improvements over the previous state-of-the-art approach on all metrics.
Gengyun Jia, Bing-Kun Bao
IEEE Trans. Multim.2
2023 TALL: Thumbnail Layout for Deepfake Video Detection
abstract
The growing threats of deepfakes to society and cybersecurity have raised enormous public concerns, and increasing efforts have been devoted to this critical topic of deepfake video detection. Existing video methods achieve good performance but are computationally intensive. This paper introduces a simple yet effective strategy named Thumbnail Layout (TALL), which transforms a video clip into a pre-defined layout to realize the preservation of spatial and temporal dependencies. Specifically, consecutive frames are masked in a fixed position in each frame to improve generalization, then resized to sub-images and rearranged into a pre-defined layout as the thumbnail. TALL is model-agnostic and extremely simple by only modifying a few lines of code. Inspired by the success of vision transformers, we incorporate TALL into Swin Transformer, forming an efficient and effective method TALL-Swin. Extensive experiments on intra-dataset and cross-dataset validate the validity and superiority of TALL and SOTA TALL-Swin. TALL-Swin achieves 90.79% AUC on the challenging cross-dataset task, FaceForensics++ → CelebDF. The code is available at https://github.com/rainy-xu/TALL4Deepfake.
Jian Liang 0001, Gengyun Jia, Zimin (Max) Yang, Ran He 0001
ICCV3
2023 Theme-Aware Aesthetic Distribution Prediction With Full-Resolution Photographs
abstract
Aesthetic quality assessment (AQA) is a challenging task due to complex aesthetic factors. Currently, it is common to conduct AQA using deep neural networks (DNNs) that require fixed-size inputs. The existing methods mainly transform images by resizing, cropping, and padding or use adaptive pooling to alternately capture the aesthetic features from fixed-size inputs. However, these transformations potentially damage aesthetic features. To address this issue, we propose a simple but effective method to accomplish full-resolution image AQA by combining image padding with region of image (RoM) pooling. Padding turns inputs into the same size. RoM pooling pools image features and discards extra padded features to eliminate the side effects of padding. In addition, the image aspect ratios are encoded and fused with visual features to remedy the shape information loss of RoM pooling. Furthermore, we observe that the same image may receive different aesthetic evaluations under different themes, which we call the theme criterion bias. Hence, a theme-aware model that uses theme information to guide model predictions is proposed. Finally, we design an attention-based feature fusion module to effectively use both the shape and theme information. Extensive experiments prove the effectiveness of the proposed method over state-of-the-art methods.
Gengyun Jia, Peipei Li 0002, Ran He 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Rethinking Image Cropping: Exploring Diverse Compositions from Global Views
abstract
Existing image cropping works mainly use anchor evaluation methods or coordinate regression methods. However, it is difficult for pre-defined anchors to cover good crops globally, and the regression methods ignore the cropping diversity. In this paper, we regard image cropping as a set prediction problem. A set of crops regressed from multiple learnable anchors is matched with the labeled good crops, and a classifier is trained using the matching results to select a valid subset from all the predictions. This new perspective equips our model with globality and diversity, mitigating the shortcomings but inherit the strengthens of previous methods. Despite the advantages, the set prediction method causes inconsistency between the validity labels and the crops. To deal with this problem, we propose to smooth the validity labels with two different methods. The first method that uses crop qualities as direct guidance is designed for the datasets with nearly dense quality labels. The second method based on the self distillation can be used in sparsely labeled datasets. Experimental results on the public datasets show the merits of our approach over state-of-the-art counterparts.
Gengyun Jia, Huaibo Huang, Chaoyou Fu, Ran He 0001
CVPR1
2022 Contrastive attention network with dense field estimation for face completion
Xin Ma 0031, Xiaoqiang Zhou, Huaibo Huang, Gengyun Jia, Zhenhua Chai, Xiaolin Wei
Pattern Recognit.4
2021 Visual-Semantic Transformer for Face Forgery Detection
abstract
This paper proposes a novel Visual-Semantic Transformer (VST) to detect face forgery based on semantic aware feature relations. In face images, intrinsic feature relations exist between different semantic parsing regions. We find that face forgery algorithms always change such relations. Therefore, we start the approach by extracting Contextual Feature Sequence (CFS) using a transformer encoder to make the best abnormal feature relation patterns. Meanwhile, images are segmented as soft face regions by a face parsing module. Then we merge the CFS and the soft face regions as Visual Semantic Sequences (VSS) representing features of semantic regions. The VSS is fed into the transformer decoder, in which the relations in the semantic region level are modeled. Our method achieved 99.58% accuracy on FF++(Raw) and 96.16% accuracy on Celeb-DF. Extensive experiments demonstrate that our framework outperforms or is comparable with state-of-the-art detection methods, especially towards unseen forgery methods.
Gengyun Jia, Huaibo Huang, Junxian Duan, Ran He 0001
IJCB2