Jiancheng Huang

dblp:83/235 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A 24-29.5-GHz Phased-Array Receiver with Subdegree RMS Phase Errors and Wide Gain Tuning Range in 130-nm SiGe BiCMOS
Kejie Hu, Kaixue Ma, Jiancheng Huang, Bingliu, Zonglin Ma, Ningning Yan, Yongrong Shi
ISCAS3
2025 Efficient Document Shadow Removal with Contrast-Aware Guidance
Yifan Liu 0001, Jiyu Wu, Jiancheng Huang, Mingfu Yan, Yi Huang 0035, Shifeng Chen
CGI (3)4
2025 Component Adaptive Clustering for Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) tackles the challenging problem of categorizing unlabeled images into both known and novel classes within a partially labeled dataset, without prior knowledge of the number of unknown categories. Traditional methods often rely on rigid assumptions, such as predefining the number of classes, which limits their ability to handle the inherent variability and complexity of real-world data. To address these shortcomings, we propose AdaGCD, a cluster-centric contrastive learning framework that incorporates Adaptive Slot Attention (AdaSlot) into the GCD framework. AdaSlot dynamically determines the optimal number of slots based on data complexity, removing the need for predefined slot counts. This adaptive mechanism facilitates the flexible clustering of unlabeled data into known and novel categories by dynamically allocating representational capacity. By integrating adaptive representation with dynamic slot allocation, our method captures both instance-specific and spatially clustered features, improving class discovery in open-world scenarios. Extensive experiments on public and fine-grained datasets validate the effectiveness of our framework, emphasizing the advantages of leveraging spatial local information for category discovery in unlabeled image datasets.
Mingfu Yan, Jiancheng Huang, Yifan Liu 0001, Shifeng Chen
ICME2
2025 MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
abstract
Text-to-image diffusion models can generate high-quality images but lack fine-grained control of visual concepts, limiting their creativity. Thus, we introduce component-controllable personalization, a new task that enables users to customize and reconfigure individual components within concepts. This task faces two challenges: semantic pollution, where undesired elements disrupt the target concept, and semantic imbalance, which causes disproportionate learning of the target concept and component. To address these, we design MagicTailor, a framework that uses Dynamic Masked Degradation to adaptively perturb unwanted visual semantics and Dual-Stream Balancing for more balanced learning of desired visual semantics. The experimental results show that MagicTailor achieves superior performance in this task and enables more personalized and creative image generation.
Jiancheng Huang, Jinbin Bai, Hao Chen 0193, Guangyong Chen, Xiaowei Hu 0001, Pheng-Ann Heng
IJCAI2
2025 FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
abstract
This work challenges the residual prediction paradigm in visual autoregressive modeling and presents FlexVAR, a new Flexible Visual AutoRegressive image generation paradigm. FlexVAR facilitates autoregressive learning with ground-truth prediction, enabling each step to independently produce plausible images. This simple, intuitive approach swiftly learns visual distributions and makes the generation process more flexible and adaptable. Trained solely on low-resolution images (< 256px), FlexVAR can: (1) Generate images of various resolutions and aspect ratios, even exceeding the resolution of the training images. (2) Support various image-to-image tasks, including image refinement, in/out-painting, and image expansion. (3) Adapt to various autoregressive steps, allowing for faster inference with fewer steps or enhancing image quality with more steps. Our 1.0B model outperforms its VAR counterpart on the ImageNet 256 × 256 benchmark. Moreover, when zero-shot transfer the image generation process with 13 steps, the performance further improves to 2.08 FID, outperforming state-of-the-art autoregressive models AiM/VAR by 0.25/0.28 FID and popular diffusion models LDM/DiT by 1.52/0.19 FID, respectively. When transferring our 1.0B model to the ImageNet 512 × 512 benchmark in a zero-shot manner, FlexVAR achieves competitive results compared to the VAR 2.3B model, which is a fully supervised model trained at 512 × 512 resolution.
Siyu Jiao, Gengwei Zhang, Yinlong Qian, Jiancheng Huang, Yao Zhao 0001, Humphrey Shi, Lin Ma 0002, Yunchao Wei, Zequn Jie
NeurIPS4
2025 Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image Editing
abstract
Text-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. To address this problem, we first analyze why the reconstruction via DDIM Inversion fails. We then propose a new inversion and sampling method named Dual-Schedule Inversion. We also design a classifier to adaptively combine Dual-Schedule Inversion with different editing methods for user-friendly image editing. Our work can achieve superior reconstruction and editing performance with the following advantages: 1) It can reconstruct real images perfectly without fine-tuning, and its reversibility is guaranteed mathematically. 2) The edited object/scene conforms to the semantics of the text prompt. 3) The unedited parts of the object/scene retain the original identity.
Jiancheng Huang, Yi Huang 0035, Jianzhuang Liu, Yifan Liu 0001, Shifeng Chen
WACV1
2025 Diffusion Model-Based Image Editing: A Survey
abstract
Denoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research.
Yi Huang 0035, Jiancheng Huang, Yifan Liu 0001, Mingfu Yan, Jiaxi Lv, Jianzhuang Liu, Wei Xiong 0008, He Zhang 0004, Liangliang Cao, Shifeng Chen
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 DCD-Net: Weakly supervised decomposition learning for real-world image dehazing
Yi Huang 0035, Jiancheng Huang, Mingfu Yan, Shifeng Chen
Signal Process.3
2025 360-degree video super resolution and quality enhancement challenge: Methods and results
Ahmed Telili, Wassim Hamidouche, Ibrahim Farhat, Hadi Amirpour, Christian Timmerer, Ibrahim Khadraoui, Jiajie Lu, The Van Le, Jeonneung Baek, Yiying Wei, Jiancheng Huang
Signal Process. Image Commun.14
2024 Distribution-Aware Calibration for Object Detection with Noisy Bounding Boxes
Jinpeng Li 0004, Jiancheng Huang, Qiang Nie, Yong Liu 0032, Bin-Bin Gao, Qiong Wang 0001, Pheng-Ann Heng, Guangyong Chen
BMVC4
2024 MambaDW: Semantic-Aware Mamba for Document Watermark Removal
Yifan Liu 0001, Mingfu Yan, He Hua, Jiancheng Huang, Shifeng Chen
CGI (1)4
2024 BK-Editer: Body-Keeping Text-Conditioned Real Image Editing
Jiancheng Huang, Linxiao Shi, Shifeng Chen
CVM (1)1
2024 Entwined Inversion: Tune-Free Inversion For Real Image Faithful Reconstruction and Editing
abstract
Text-conditional image editing is a very practical AIGC task that has recently emerged with great commercial and academic research value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing, but DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for all downstream edits. In order to solve this problem, we first mathematically analyze the reason for the reconstruction failure of DDIM Inversion, and then propose a new inversion and sampling method named Entwined Inversion that can achieve satisfactory reconstruction and editing performance, which can solve two major problems: 1) the object can retain the main content of the original image; 2) the edited object can conform to the semantics of the text prompt. In addition, our method does not require training the diffusion model itself on a large dataset, nor does it require any fine-tuning for some particular images.
Jiancheng Huang, Yifan Liu 0001, Jiaxi Lv, Shifeng Chen
ICASSP1
2024 Color-SD: Stable Diffusion Model Already has a Color Style Noisy Latent Space
abstract
We present Color-SD, a comprehensive color style transfer framework that utilizes either image or text references. Built on the pretrained Stable Diffusion Model, Color-SD exploits an existing color style space, enabling a training-free and tuning-free zero-shot color style transfer method without introducing new parameters. For image references, we first invert the source and reference images to the noisy latent space, followed by parallel sampling. During this process, we execute distribution transformation in the noisy latent space, effectively completing the color style transfer and generating the stylized result. For text references, we capitalize on the Stable Diffusion model’s inherent text-to-image capability. We only invert the source image to the noisy latent, and the given text reference prompt is utilized during the parallel sampling. This approach eliminates the need for training or tuning, yet produces impressive open-set transfer results. Comprehensive experiments validate the effectiveness of our method, demonstrating significant superiority over existing methods in both qualitative and quantitative evaluations.
Jiancheng Huang, Mingfu Yan, Shifeng Chen
ICME1
2024 SBCR: Stochasticity Beats Content Restriction Problem in Training and Tuning Free Image Editing
abstract
Text-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as a first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. Many inversion-based works modify the formula to address this problem but this leads to another content restriction problem. To solve the content restriction problem, we first analyze why the reconstruction via DDIM Inversion fails and then propose Reconstruction-and-Generation Balancing Noises (R&G-B noises) that can achieve superior reconstruction and editing performance with the following advantages: 1) It can perfectly reconstruct real images without fine-tuning. 2) It can overcome the content restriction problem and generate diverse content.
Jiancheng Huang, Mingfu Yan, Shifeng Chen
ICMR1
2024 MagicFight: Personalized Martial Arts Combat Video Generation
abstract
Amid the surge in generic text-to-video generation, the field of personalized human video generation has witnessed notable advancements, primarily concentrated on single-person scenarios. However, to our knowledge, the domain of two-person interactions, particularly in the context of martial arts combat, remains uncharted. We identify a significant gap: existing models for single-person dancing generation prove insufficient for capturing the subtleties and complexities of two engaged fighters, resulting in challenges such as identity confusion, anomalous limbs, and action mismatches. To address this, we introduce a pioneering new task, Personalized Martial Arts Combat Video Generation. Our approach, MagicFight, is specifically crafted to overcome these hurdles. Given this pioneering task, we face a lack of appropriate datasets. Thus, we generate a bespoke dataset using the game physics engine Unity, meticulously crafting a multitude of 3D characters, martial arts moves, and scenes designed to represent the diversity of combat. MagicFight refines and adapts existing models and strategies to generate high-fidelity two-person combat videos that maintain individual identities and ensure seamless, coherent action sequences, thereby laying the groundwork for future innovations in the realm of interactive video content creation.
Jiancheng Huang, Mingfu Yan, Songyan Chen, Yi Huang 0035, Shifeng Chen
ACM Multimedia1
2024 IFAST: Weakly Supervised Interpretable Face Anti-Spoofing From Single-Shot Binocular NIR Images
abstract
Single-shot face anti-spoofing (FAS) is a key technique for securing face recognition systems, relying solely on static images as input. However, single-shot FAS remains a challenging and under-explored problem due to two reasons: 1) On the data side, learning FAS from RGB images is largely context-dependent, and single-shot images without additional annotations contain limited semantic information. 2) On the model side, existing single-shot FAS models struggle to provide proper evidence for their decisions, and FAS methods based on depth estimation require expensive per-pixel annotations. To address these issues, we construct and release a large binocular NIR image dataset named BNI-FAS, which contains more than 300,000 real face and plane attack images, and propose an Interpretable FAS Transformer (IFAST) that requires only weak supervision to produce interpretable predictions. Our IFAST generates pixel-wise disparity maps using the proposed disparity estimation Transformer with Dynamic Matching Attention (DMA) blocks. Besides, we design a confidence map generator to work in tandem with a dual-teacher distillation module to obtain the final discriminant results. Comprehensive experiments show that our IFAST achieves state-of-the-art performance on BNI-FAS, verifying its effectiveness of single-shot FAS on binocular NIR images. The project page is available athttps://ifast-bni.github.io/.
Jiancheng Huang, Jianzhuang Liu, Linxiao Shi, Shifeng Chen
IEEE Trans. Inf. Forensics Secur.1
2024 WaveDM: Wavelet-Based Diffusion Models for Image Restoration
abstract
Latest diffusion-based methods for many image restoration tasks outperform traditional models, but they encounter the long-time inference problem. To tackle it, this paper proposes a Wavelet-Based Diffusion Model (WaveDM). WaveDM learns the distribution of clean images in the wavelet domain conditioned on the wavelet spectrum of degraded images after wavelet transform, which is more time-saving in each step of sampling than modeling in the spatial domain. To ensure restoration performance, a unique training strategy is proposed where the low-frequency and high-frequency spectrums are learned using distinct modules. In addition, an Efficient Conditional Sampling (ECS) strategy is developed from experiments, which reduces the number of total sampling steps to around 5. Evaluations on twelve benchmark datasets including image raindrop removal, rain steaks removal, dehazing, defocus deblurring, demoiréing, and denoising demonstrate that WaveDM achieves state-of-the-art performance with the efficiency that is comparable to traditional one-pass methods and over 100× faster than existing image restoration methods using vanilla diffusion models. The code is available athttps://github.com/stayalive16/WaveDM
Yi Huang 0035, Jiancheng Huang, Jianzhuang Liu, Mingfu Yan, Jiaxi Lv, Chaoqi Chen, Shifeng Chen
IEEE Trans. Multim.2
2023 KV Inversion: KV Embeddings Learning for Text-Conditioned Real Image Action Editing
Jiancheng Huang, Shifeng Chen
PRCV (1)1
2023 Bootstrap Diffusion Model Curve Estimation for High Resolution Low-Light Image Enhancement
Jiancheng Huang, Shifeng Chen
PRICAI (3)1
2023 DeSeal: Semantic-Aware Seal2Clear Attention for Document Seal Removal
abstract
Seal removal aims to eliminate the seal portion from documents to facilitate better OCR and document reconstruction. However, existing seal removal methods often lack publicly available code and pre-trained models, and suffer from a lack of publicly available seal datasets. To address these issues, we propose DeSeal for seal removal and introduce a SealBank dataset containing 100K paired seal images. In DeSeal, we introduce the Semantic-Aware Seal2Clear Attention and Color-Adapter Module, where the former identifies seal regions in the entire image and focuses on removing seals from these areas, while the latter significantly improves the model's generalization performance, enabling it to perform well on both real and synthetic data. Experimental results on the SealBank dataset demonstrate the effectiveness of our proposed DeSeal.
Jiancheng Huang, Shifeng Chen
IEEE Signal Process. Lett.2
2001 A New Stroke-Based Directional Feature Extraction Approach for Handwritten Chinese Character Recognition
abstract
A directional feature extraction approach based on stroke directional decomposition of a Chinese character is proposed. Without extracting the skeleton or contour of the character, the four directional sub-patterns, namely, horizontal (-), vertical (|), left up diagonal (/) and right up diagonal () sub-patterns could be obtained directly from analyzing the stroke directional characteristics of the character. Five kinds of line-density based elastic meshing methods are presented to extract cellular directional features. Experimentation on a total of 18800 handwritten samples from 940 categories produces a recognition rate of 92.71%, showing the effectiveness of the proposed approach.
Xue Gao, Junxun Yin, Jiancheng Huang
ICDAR4
2000 Deformation Transformation for Handwritten Chinese Character Shape Correction
Jiancheng Huang, Junxun Yin, Qianhua He
ICMI2
2000 A New Multi-classifier Combination Scheme and Its Application in Handwriting Chinese Character Recognition
Minqing Wu, Junxun Yin, Jiancheng Huang
ICMI5