Jingkai Zhou

dblp:36/48 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0003-4629-0659ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images
abstract
In recent years, diffusion models have revolutionized visual generation, outperforming traditional frameworks like Generative Adversarial Networks (GANs). However, generating images of humans with realistic semantic parts, such as hands and faces, remains a significant challenge due to their intricate structural complexity. To address this issue, we propose a novel post-processing solution named RealisHuman. The RealisHuman framework operates in two stages. First, it generates realistic human parts, such as hands or faces, using the original malformed parts as references, ensuring consistent details with the original image. Second, it seamlessly integrates the rectified human parts back into their corresponding positions by repainting the surrounding areas to ensure smooth and realistic blending. The RealisHuman framework significantly enhances the realism of human generation, as demonstrated by notable improvements in both qualitative and quantitative metrics.
Benzhi Wang, Jingkai Zhou, Jingqi Bai, Yang Yang 0062, Fan Wang 0019, Zhen Lei 0001
AAAI2
2025 On Denoising Walking Videos for Gait Recognition
abstract
To capture individual gait patterns, excluding identity-irrelevant cues in walking videos, such as clothing texture and color, remains a persistent challenge for vision-based gait recognition. Traditional silhouette- and pose-based methods, though theoretically effective at removing such distractions, often fall short of high accuracy due to their sparse and less informative inputs. Emerging end-to-end methods address this by directly denoising RGB videos using human priors. Building on this trend, we propose DenoisingGait, a novel gait denoising method. Inspired by the philosophy that "what I cannot create, I do not understand", we turn to generative diffusion models, uncovering how they partially filter out irrelevant factors for gait understanding. Additionally, we introduce a geometry-driven Feature Matching module, which, combined with background removal via human silhouettes, condenses the multi-channel diffusion features at each foreground pixel into a two-channel direction vector. Specifically, the proposed within- and cross-frame matching respectively capture the local vectorized structures of gait appearance and motion, producing a novel flow-like gait representation termed Gait Feature Field, which further reduces residual noise in diffusion features. Experiments on the CCPG, CASIA-B*, and SUSTech1K datasets demonstrate that DenoisingGait achieves a new SoTA performance in most cases for both within- and cross-domain evaluations. Code is available at https://github.com/ShiqiYu/OpenGait.
Dongyang Jin, Chao Fan 0001, Jingzhe Ma, Jingkai Zhou, Shiqi Yu 0001
CVPR4
2025 Layer-Animate for Transparent Video Generation
abstract
Transparent videos with alpha channels play a crucial role in film production, advertising, and augmented reality fields. However, there is currently no available method for producing transparent videos. Traditional methods are time-consuming and labor-intensive, and employing alternative approaches for this task will result in inaccurate transparent regions, constrained motion, and artifacts. To address these challenges, we propose Layer-Animate, the first method capable of generating transparent videos. Our method comprises two stages: in the first stage, transparent images are generated as the base images to provide content and transparency information for the next stage. In the second stage, Inter-Frame Attention is applied to decouple content from motion, enabling the motion module to focus better on action. Layer-Animate is the first method used to generate transparent videos with accurate transparent regions, sufficient motion, and no artifacts, as demonstrated by notable improvements in qualitative and quantitative metrics.
Jingqi Bai, Jingkai Zhou, Benzhi Wang, Yang Yang 0062, Zhen Lei 0001, Fan Wang 0019
ICASSP2
2025 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models
abstract
Video try-on replaces clothing in videos with target garments. Existing methods struggle to generate high-quality and temporally consistent results when handling complex clothing patterns and diverse body poses. We present 3DV-TON, a novel diffusion-based framework for generating high-fidelity and temporally consistent video try-on results. Our approach employs generated animatable textured 3D meshes as explicit frame-level guidance, alleviating the issue of models over-focusing on appearance fidelity at the expanse of motion coherence. This is achieved by enabling direct reference to consistent garment texture movements throughout video sequences. The proposed method features an adaptive pipeline for generating dynamic 3D guidance: (1) selecting a keyframe for initial 2D image try-on, followed by (2) reconstructing and animating a textured 3D mesh synchronized with original video poses. We further introduce a robust rectangular masking strategy that successfully mitigates artifact propagation caused by leaking clothing information during dynamic human and garment movements. To advance video try-on research, we introduce HR-VVT, a high-resolution benchmark dataset containing 130 videos with diverse clothing types and scenarios. Quantitative and qualitative results demonstrate our superior performance over existing methods.
Chaohui Yu, Jingkai Zhou, Fan Wang 0019
ACM Multimedia3
2025 Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
abstract
Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with high-quality annotations for both aspects. To overcome this, we present Uni3C, a unified 3D-enhanced framework for precise control of both camera and human motion in video generation. Uni3C includes two key contributions. First, we propose a plug-and-play control module trained with a frozen video generative backbone, PCDController, which utilizes unprojected point clouds from monocular depth to achieve accurate camera control. By leveraging the strong 3D priors of point clouds and the powerful capacities of video foundational models, PCDController shows impressive generalization, performing well regardless of whether the inference backbone is frozen or fine-tuned. This flexibility enables different modules of Uni3C to be trained in specific domains, i.e., either camera control or human motion control, reducing the dependency on jointly annotated data. Second, we propose a jointly aligned 3D world guidance for the inference phase that seamlessly integrates both scenic point clouds and SMPL-X characters to unify the control signals for camera and human motion, respectively. Extensive experiments confirm that PCDController enjoys strong robustness in driving camera motion for fine-tuned backbones of video generation. Uni3C substantially outperforms competitors in both camera controllability and human motion quality. Additionally, we collect tailored validation sets featuring challenging camera movements and human actions to validate the effectiveness of our method. Codes are released at https://github.com/alibaba-damo-academy/Uni3C.
Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang, Chaohui Yu, Fan Wang 0019, Xiangyang Xue 0001, Yanwei Fu 0001
SIGGRAPH Asia2
2024 Adversarial Score Distillation: When Score Distillation Meets GAN
abstract
Existing score distillation methods are sensitive to classifier-free guidance (CFG) scale, manifested as over-smoothness or instability at small CFG scales, while over-saturation at large ones. To explain and analyze these issues, we revisit the derivation of Score Distillation Sampling (SDS) and decipher existing score distillation with the Wasserstein Generative Adversarial Network (WGAN) paradigm. With the WGAN paradigm, we find that existing score distillation either employs a fixed sub-optimal discriminator or conducts incomplete discriminator optimization, resulting in the scale-sensitive issue. We propose the Adversarial Score Distillation (ASD), which maintains an optimizable discriminator and updates it using the complete optimization objective. Experiments show that the proposed ASD performs favorably in 2D distillation and text-to-3D tasks against existing methods. Furthermore, to explore the generalization ability of our paradigm, we extend ASD to the image editing task, which achieves competitive results. The project page and code are at this link.
Jingkai Zhou, Junyao Sun, Xuesong Zhang 0001
CVPR2
2023 ASM: Adaptive Skinning Model for High-Quality 3D Face Modeling
abstract
The research fields of parametric face model and 3D face reconstruction have been extensively studied. However, a critical question remains unanswered: how to tailor the face model for specific reconstruction settings. We argue that reconstruction with multi-view uncalibrated images demands a new model with stronger capacity. Our study shifts attention from data-dependent 3D Morphable Models (3DMM) to an understudied human-designed skinning model. We propose Adaptive Skinning Model (ASM), which redefines the skinning model with more compact and fully tunable parameters. With extensive experiments, we demonstrate that ASM achieves significantly improved capacity than 3DMM, with the additional advantage of model size and easy implementation for new topology. We achieve state-of-the-art performance with ASM for multi-view reconstruction on the Florence MICC Coop benchmark. Our quantitative analysis demonstrates the importance of a high-capacity model for fully exploiting abundant information from multi-view input in reconstruction. Furthermore, our model with physical-semantic parameters can be directly utilized for real-world applications, such as in-game avatar creation. As a result, our work opens up new research direction for parametric face model and facilitates future research on multi-view reconstruction.
Hong Shang, Tianyang Shi, Xinghan Chen, Jingkai Zhou, Zhongqian Sun, Wei Yang 0032
ICCV5
2023 What Limits the Performance of Local Self-attention?
Jingkai Zhou, Pichao Wang, Jiasheng Tang, Fan Wang 0019, Hao Li 0030, Rong Jin 0001
Int. J. Comput. Vis.1
2023 PoiseNet: Dealing With Data Imbalance in DensePose
abstract
Data imbalance, a foundational problem in machine learning, has received little attention in DensePose and has become one of the main obstacles in front of existing methods. We reveal two imbalances in DensePose: inter-surface imbalance and intra-surface imbalance. First, the human body parts can be of various sizes, making 3D surfaces contain different numbers of annotations. Classifiers trained on such unbalanced data suffer a decline on surfaces with fewer annotations. Second, annotations within each surface are not uniformly distributed. Regressors trained on such uneven data suffer a decline in 3D surface areas with sparse annotations. To solve these imbalances, we propose PoiseNet which integrates adaptive equalization loss (AEQL) and block balanced localization (BBL). Specifically, to address the inter-surface imbalance, AEQL adaptively reweights surfaces based on the number of annotations and the classification scores. BBL alleviates the intra-surface imbalance by unevenly blocking each surface according to annotation distribution. Experimental results on the DensePose-COCO dataset show that our PoiseNet surpasses baselines by up to 1.4 AP.
Junyao Sun, Jingkai Zhou, Qiong Liu 0006
IEEE Trans. Circuits Syst. Video Technol.2
2022 Scaled ReLU Matters for Training Vision Transformers
abstract
Vision transformers (ViTs) have been an alternative design paradigm to convolutional neural networks (CNNs). However, the training of ViTs is much harder than CNNs, as it is sensitive to the training parameters, such as learning rate, optimizer and warmup epoch. The reasons for training difficulty are empirically analysed in the paper Early Convolutions Help Transformers See Better, and the authors conjecture that the issue lies with the patchify-stem of ViT models. In this paper, we further investigate this problem and extend the above conclusion: only early convolutions do not help for stable training, but the scaled ReLU operation in the convolutional stem (conv-stem) matters. We verify, both theoretically and empirically, that scaled ReLU in conv-stem not only improves training stabilization, but also increases the diversity of patch tokens, thus boosting peak performance with a large margin via adding few parameters and flops. In addition, extensive experiments are conducted to demonstrate that previous ViTs are far from being well trained, further showing that ViTs have great potential to be a better substitute of CNNs.
Pichao Wang, Xue Wang 0010, Hao Luo 0004, Jingkai Zhou, Fan Wang 0019, Hao Li 0030, Rong Jin 0001
AAAI4
2021 Decoupled Dynamic Filter Networks
abstract
Convolution is one of the basic building blocks of CNN architectures. Despite its common use, standard convolution has two main shortcomings: Content-agnostic and Computation-heavy. Dynamic filters are content-adaptive, while further increasing the computational overhead. Depth-wise convolution is a lightweight variant, but it usually leads to a drop in CNN performance or requires a larger number of channels. In this work, we propose the Decoupled Dynamic Filter (DDF) that can simultaneously tackle both of these shortcomings. Inspired by recent advances in attention, DDF decouples a depth-wise dynamic filter into spatial and channel dynamic filters. This decomposition considerably reduces the number of parameters and limits computational costs to the same level as depth-wise convolution. Meanwhile, we observe a significant boost in performance when replacing standard convolution with DDF in classification networks. ResNet50 / 101 get improved by 1.9% and 1.3% on the top-1 accuracy, while their computational costs are reduced by nearly half. Experiments on the detection and joint upsampling networks also demonstrate the superior performance of the DDF upsampling variant (DDF-Up) in comparison with standard convolution and specialized content-adaptive layers. The project page with code is available1.
Jingkai Zhou, Varun Jampani, Zhixiong Pi, Qiong Liu 0001, Ming-Hsuan Yang 0001
CVPR1
2020 Novel up-scale feature aggregation for object detection in aerial images
Jingkai Zhou, Chi-Man Vong, Qiong Liu 0006
Neurocomputing2
2019 Scale adaptive image cropping for UAV object detection
Jingkai Zhou, Chi-Man Vong, Qiong Liu 0006, Zhenyu Wang 0001
Neurocomputing1
2018 Nighttime FIR Pedestrian Detection Benchmark Dataset for ADAS
Zhewei Xu, Jiajun Zhuang, Jingkai Zhou, Shaowu Peng
PRCV (4)4
2018 Object tracking via Online Multiple Instance Learning with reliable components
Feng Wu 0004, Shaowu Peng, Jingkai Zhou, Qiong Liu 0006, Xiaojia Xie
Comput. Vis. Image Underst.3