EDBT 2026 Demo / reviewers in the wild / expert
Yan Huang 0031
dblp:75/6434-31
· DBLP profile ↗
27ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0001-9136-195XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SVVHD: An Open-Source and Self-Verified Benchmark Framework for LLM Evaluation in VHDL Code Generation
Yan Huang 0031, Meihua Liu |
ISCAS | 2 |
| 2026 | MVHDiff: Leveraging 3D priors for consistent multi-view human image generation with diffusion models
Yan Huang 0031, Hongxin Fu, Zhonghang Li, Yongcan Luo, Si Wu 0002 |
Neurocomputing | 1 |
| 2026 | HumanDiff: Leveraging body-part expert-assisted diffusion transformers for human image generation
Yan Huang 0031, Zhiren Wang, Hongzong Li, Yongcan Luo, Si Wu 0002 |
Knowl. Based Syst. | 1 |
| 2026 | Deep intrinsic image decomposition via physics-aware neural networks
Yan Huang 0031, Kangjie Liu, Tengyue Chen, Yong Xu 0007, Hui Ji 0002 |
Pattern Recognit. | 1 |
| 2026 | ClassBooth: Boost Class Semantics With Bidirectional Feature Fusion in Text-to-Image Diffusion ModelsabstractText-to-image (T2I) diffusion models aim to generate images that are both visually realistic and aligned with open-domain textual prompts. However, the leading T2I models often fall short in capturing semantic details, especially for fine-grained object categories. We find that simply expanding original prompts cannot effectively guide the model to generate the desired content, since T2I diffusion models pre-trained on generic datasets lack a mechanism to boost class semantics. In this work, we present ClassBooth, a flexible framework that improves pre-trained T2I diffusion model in rendering class-specific content, while preserving the open-domain generation capability. Toward this end, we introduce an auxiliary class-specific semantic booster conditioned on learnable prompts, which are associated with specific classes to encode fine-grained conditioning information. To enable the T2I model to synthesize fine-grained details of specific categories, we perform bidirectional fusion on the features conditioned on different information, and this design is beneficial for class-specific semantic expression, thereby synthesizing high-fidelity data encapsulating precise class-specific details. We validate ClassBooth across multiple benchmarks, demonstrating its superiority over existing methods through comprehensive quantitative and qualitative evaluations. Yan Huang 0031, Hau-San Wong, Si Wu 0002 |
IEEE Trans. Multim. | 3 |
| 2025 | Zero-Shot Low-Light Image Enhancement via Latent Diffusion ModelsabstractLow-light image enhancement (LLIE) aims to improve visibility and signal-to-noise ratio in images captured under poor lighting conditions. While deep learning has shown promise in this domain, current approaches require extensive paired training data, limiting their practical utility. We present a novel framework that reformulates low-light image enhancement as a zero-shot inference problem using pre-trained latent diffusion models (LDMs), eliminating the need for task-specific training data. Our key insight is that the rich natural image priors encoded in LDMs can be leveraged to recover well-lit images through a carefully designed optimization process. To address the ill-posed nature of low-light degradation and the complexity of latent space optimization, our framework introduces an exposure-aware degradation module that adaptively models illumination variations and a principled latent regularization scheme with adaptive guidance that ensures both enhancement quality and natural image statistics. Experimental results demonstrate that our framework outperforms existing zero-shot methods across diverse real-world scenarios. Yan Huang 0031, Xiaoshan Liao, Jinxiu Liang, Yuhui Quan, Boxin Shi, Yong Xu 0007 |
AAAI | 1 |
| 2025 | InpaintFormer: Prompt-guided High-Quality Face Inpainting with Mask-Aware Self-AttentionabstractFace image inpainting, especially with user-controllable customization, aims to restore degraded facial regions while adhering to user-provided instructions. Traditional inpainting methods often focus solely on restoring visual fidelity, lacking the ability to incorporate user prompts or semantic guidance. In this work, we present InpaintFormer, a novel framework for user-controlled face image inpainting guided by textual prompts. Specifically, we propose a Prompt-guided Feature Modulation (PGFM) module to align visual features with user instructions by utilizing a pre-trained CLIP model to extract text and image embeddings. These embeddings are fused to modulate the encoded image features, ensuring semantic consistency with the prompt. Additionally, a Degradation Mask Predictor (DMP) is introduced to identify degraded regions requiring inpainting, while a Mask-Aware Self-Attention (MASA) mechanism within the Transformer refines the inpainting process by selectively attending to non-degraded regions for generating realistic results. By combining PGFM, DMP, and MASA, InpaintFormer enables controllable face image inpainting with high fidelity and semantic alignment. Extensive experiments demonstrate that InpaintFormer outperforms state-of-the-art inpainting methods in terms of controllability and naturalness. Zhouhao Ouyang, Yan Huang 0031, Si Wu 0002, Yong Xu 0007, Patrick Le Callet, Dapeng Oliver Wu |
ICME | 4 |
| 2025 | Text to Trajectory: Enhancing and Evaluating LLMs for Embodied Task PlanningabstractThe increasing demand for effective human-machine interaction highlights the importance of integrating natural language processing with robotics technology. This paper addresses the challenges of using Large Language Models (LLMs) for embodied task planning in complex environments. We propose a comprehensive framework that combines environmental-aware LLM fine-tuning with a novel Stepwise Beam Search (SBS) strategy. In conjunction with the environmentally enhanced LLM, the SBS strategy facilitates comprehensive exploration of both token-level and step-level search spaces, overcoming the limitations of conventional greedy search methods. Additionally, to evaluate the effectiveness of embodied task planning, we introduce the Trajectory Match Score (TMS), a robust evaluation metric that leverages state-based simulation to assess plan success. Through extensive experiments on standard benchmarks, our framework demonstrates substantial improvements in both plan generation quality and task success rates, advancing the state-of-the-art in embodied task planning. Yihan Tang, Yong Xu 0007, Ruotao Xu, Yan Huang 0031, Si Wu 0002, Patrick Le Callet |
ICME | 4 |
| 2025 | SemanticLoom: Category-aware Dynamic Fusion for Multi-class Few-shot Image SynthesisabstractFew-shot text-to-image (T2I) generation seeks to efficiently integrate new semantics into existing pre-trained models while preserving their capacity to generate diverse, high-quality images. However, existing methods often suffer from inefficiency and poor scalability due to the need for separate training processes for each new concept. These challenges hinder their practical application in multi-class few-shot scenarios. To overcome these issues, we propose SemanticLoom that dynamically incorporates novel concepts into pre-trained diffusion models through category-aware dynamic feature fusion. Our approach introduces a lightweight semantic expander that captures fine-grained semantics, guided by learnable identifier to ensure precise semantic integration. By dynamically adjusting feature fusion coefficients based on category diversity and training progress, our method harmonizes the integration of new semantic features with the original model’s capabilities, ensuring consistency and generalization. Experiments demonstrate that our method successfully integrates new semantics without compromising the generative diversity and versatility of the pre-trained model. Yan Huang 0031, Si Wu 0002, Yong Xu 0007, Patrick Le Callet |
ICME | 2 |
| 2025 | Adaptive Illumination Transfer Network for Shadow RemovalabstractShadow removal aims to harmonize illumination between shadow and non-shadow regions. However, existing methods often struggle to achieve this goal due to inadequate modeling of illumination relationships between these two regions. Moreover, the prevalent reliance on binary shadow masks hinders their capability to address non-uniform shadows. To address these limitations, we propose an adaptive illumination transfer network (AITNet), which incorporates two complementary shadow enhancement strategies. First, a global illumination transfer strategy is designed to model the illumination relationship between shadow and non-shadow regions, enabling the holistic enhancement of shadow regions. Second, an illumination-adaptive strategy is developed to estimate an illumination degradation map, which guides the adaptive enhancement of shadow regions. Furthermore, to preserve the original structural details, the enhancement process is applied exclusively to the illumination map obtained after Retinex decomposition. Extensive experiments have demonstrated the superiority of our method over existing approaches on public datasets. Si Wu 0002, Yong Xu 0007, Yan Huang 0031, Patrick Le Callet |
ICME | 4 |
| 2025 | E-Net for pansharpening: A super-resolution perspective
Si Wu 0002, Yong Xu 0007, Yan Huang 0031 |
Image Vis. Comput. | 5 |
| 2025 | Image shadow removal via multi-scale deep Retinex decomposition
Yan Huang 0031, Xinchang Lu, Yuhui Quan, Yong Xu 0007, Hui Ji 0002 |
Pattern Recognit. | 1 |
| 2025 | Detail-Preserving Diffusion Models for Low-Light Image EnhancementabstractExisting diffusion models for low-light image enhancement typically incrementally remove noise introduced during the forward diffusion process using a denoising loss, with the process being conditioned on input low-light images. While these models demonstrate remarkable abilities in generating realistic high-frequency details, they often struggle to restore fine details that are faithful to the input. To address this, we present a novel detail-preserving diffusion model for realistic and faithful low-light image enhancement. Our approach integrates a size-agnostic diffusion process with a reverse process reconstruction loss, significantly enhancing the fidelity of enhanced images to their low-light counterparts and enabling more accurate recovery of fine details. To ensure the preservation of region- and content-aware details, we employ an efficient noise estimation network with a simplified channel-spatial attention mechanism. Additionally, we propose a multiscale ensemble scheme to maintain detail fidelity across diverse illumination regions. Comprehensive experiments on eight benchmark datasets demonstrate that our method achieves state-of-the-art results compared to over twenty existing methods in terms of both perceptual quality (LPIPS) and distortion metrics (PSNR and SSIM). The code is available at:https://github.com/CSYanH/DePDiff. Yan Huang 0031, Xiaoshan Liao, Jinxiu Liang, Boxin Shi, Yong Xu 0007, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Dual-Path Deep Unsupervised Learning for Multi-Focus Image FusionabstractMulti-focus image fusion (MFIF) aims at merging multiple images captured at different focal lengths to create an all-in-focus image. This paper introduces a fully unsupervised learning approach for MFIF that uses only pairs of defocused images for end-to-end training, bypassing the need for ground-truths in supervised learning. Unlike existing methods training via a similarity loss between fused and source images, we propose a dual-path learning framework comprising two networks: an image fuser and a mask predictor. The mask predictor is modeled as a self-supervised denoising network on imperfect fusion masks, trained with a masking-based unsupervised learning scheme. The image fuser, crafted with deep unrolling, leverages the output from the mask predictor to supervise its mask generation at each unrolled step. Moreover, we introduce a fusion consistency loss to ensure the alignment between the image fuser and the mask predictor. In extensive experiments, our proposed approach shows superiority over existing end-to-end unsupervised methods and competitive performance against the supervised ones. Yuhui Quan, Xi Wan, Yan Huang 0031, Hui Ji 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | Hunting Blemishes: Language-guided High-fidelity Face Retouching Transformer with Limited Paired DataabstractThe prevalence of multimedia applications has led to increased concerns and demand for auto face retouching. Face retouching aims to enhance portrait quality by removing blemishes. However, the existing auto-retouching methods rely heavily on a large amount of paired training samples, and perform less satisfactorily when handling complex and unusual blemishes. To address this issue, we propose a Language-guided Blemish Removal Transformer for automatically retouching face images, while at the same time reducing the dependency of the model on paired training data. Our model is referred to as LangBRT, which leverages vision-language pre-training for precise facial blemish removal. Specifically, we design a text-prompted blemish detection module that indicates the regions to be edited. The priors not only enable the transformer network to handle specific blemishes in certain areas, but also reduce the reliance on retouching training data. Further, we adopt a target-aware cross attention mechanism, such that the blemish-like regions are edited accurately while at the same time maintaining the normal skin regions unchanged. Finally, we adopt a regularization approach to encourage the semantic consistency between the synthesized image and the text description of the desired retouching outcome. Extensive experiments are performed to demonstrate the superior performance of LangBRT over competing auto-retouching methods in terms of dependency on training data, blemish detection accuracy and synthesis quality. Yan Huang 0031, Lianxin Xie, Cheng Liu 0001, Si Wu 0002, Hau-San Wong |
ACM Multimedia | 2 |
| 2024 | Enhancing Underwater Images via Asymmetric Multi-Scale Invertible NetworksabstractUnderwater images, often plagued by complex degradation, pose significant challenges for image enhancement. To address these challenges, the paper redefines underwater image enhancement as an image decomposition problem and proposes a deep invertible neural network (INN) that accurately predicts both the latent image and the degradation effects. Instead of using an explicit formation model to describe the degradation process, the INN adheres to the constraints of the image decomposition model, providing necessary regularization for model training, particularly in the absence of supervision on degradation effects. Taking into account the diverse scales of degradation factors, the INN is structured on a multi-scale basis to effectively manage the varied scales of degradation factors. Moreover, the INN incorporates several asymmetric design elements that are specifically optimized for the decomposition model and the unique physics of underwater imaging. Comprehensive experiments show that our approach provides significant performance improvement over existing methods. Yuhui Quan, Xiaoheng Tan, Yan Huang 0031, Yong Xu 0007, Hui Ji 0002 |
ACM Multimedia | 3 |
| 2023 | Video Noise Removal Using Progressive Decomposition With Conditional InvertibilityabstractVideo denoising aims at removing noise from noisy video frames and meanwhile preserving their structures and details. It is a challenging task, as both noise and video structures/details correspond to high-frequency components of a noisy video which are hard to distinguish. This paper proposes a deep video denoiser using a progressive decomposition process with conditional invertibility. Noisy video frames are first decomposed into two latent codes via a forward process of conditional invertible coupling layers, where one latent code carries the maximal information regarding the noise-free reference frame while the other encodes the information regarding noise, misalignment and content difference. The clean video is then reconstructed from the latent codes of noise-free frames using the reverse pass of the coupling layers. To improve the robustness to variant noise levels, the coupling layers are conditioned on noise level. In addition, memory units are introduced to the conditioned coupling layers to better exploit temporal correlation among frames for feature disentanglement. Experiments on two benchmark datasets have demonstrated the effectiveness of our method. Haoran Huang, Yuhui Quan, Zhenghua Lei, Jinlong Hu 0002, Yan Huang 0031 |
ICME | 5 |
| 2023 | Image Desnowing via Deep Invertible SeparationabstractImages taken on snowy days often suffer from severe negative visual effects caused by snowflakes. The task of removing snowflakes from a snowy image is known as image desnowing, which is challenging as image details are easily mistakenly treated and thus may be significantly lost during snowflake removal. Leveraging invertible neural networks (INNs), this paper presents a deep learning-based method for single image desnowing, which can remove snowflakes accurately while preserving image details well. Interpreting desnowing as an image decomposition problem, we propose an INN composed of two asymmetric interactive paths for predicting a latent image and a snowflake layer respectively. Such an INN is able to progressively refine the features of both latent images and snowflake layers for disentanglement, while retaining all information possibly relevant to latent image reconstruction. In addition, an attentive coupling layer supervised by snowflake masks is introduced to enhance feature dismantlement and a coupling-in-coupling structure is developed for further improvement. Extensive experiments show that, the proposed method outperforms existing ones on three benchmark datasets of synthetic and real-world images, and meanwhile it also shows advantages in terms of model size and computational efficiency. Yuhui Quan, Xiaoheng Tan, Yan Huang 0031, Yong Xu 0007, Hui Ji 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | High-Quality Self-Supervised Snapshot Hyperspectral ImagingabstractHyperspectral image (HSI) reconstruction is about recovering a 3D HSI from its 2D snapshot measurements, to which deep models have become a promising approach. However, most existing studies train deep models on large amounts of organized data, the collection of which can be difficult in many applications. This paper leverages the image priors encoded in untrained neural networks (NNs) to have a self-supervised learning method which is free from training datasets while adaptive to the statistics of a test sample. To induce better image priors and prevent the NN overfitting undesired solutions, we construct an unrolling-based NN equipped with fractional max pooling (FMP). Furthermore, the FMP is used with randomness to enable self-ensemble for reconstruction accuracy improvement. In the experiments, our self-supervised learning approach enjoys high-quality reconstruction and outperforms recent methods including the supervised ones. Yuhui Quan, Xinran Qin, Mingqin Chen, Yan Huang 0031 |
ICASSP | 4 |
| 2022 | Phase Recovery With Deep Complex-Domain PriorsabstractPhase recovery (PR) of a signal from its amplitude measurements is one challenging task in signal processing. The key is suppressing the noise while rectifying the phase of the signal during the inversion process. This letter proposes a deep model-aware approach for PR by unrolling an optimization model regularized with image priors defined in the complex domain. A complex-valued (CV) deep neural network is then introduced to implement effective plug-and-play image priors that enjoy the benefits of CV operations for PR, such as sophisticated operations on local phases and regularization by compact convolution. As a result, the proposed approach can handle the noise well at each iteration in the unrolled process and improve the recovery accuracy. In experiments, the proposed approach shows superior performance to recent methods. Zhuojie Chen, Yan Huang 0031, Yu Hu 0004 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Single-image raindrop removal using concurrent channel-spatial attention and long-short skip connections
Jiayi Peng, Yong Xu 0007, Yan Huang 0031 |
Pattern Recognit. Lett. | 4 |
| 2020 | Weakly-Supervised Sparse Coding With Geometric Prior for Interactive Texture SegmentationabstractTexture segmentation is about dividing a texture-dominant image into multiple homogeneous texture regions. The existing unsupervised approaches for texture segmentation are annotation-free but often yield unsatisfactory results. In contrast, supervised approaches such as deep learning may have better performance but require a large amount of annotated data. In this letter, we propose a user-interactive approach to win the trade-off between unsupervised approaches and supervised deep approaches. Our approach requires the user to mark one pixel in each texture region, whose label is directly propagated to its neighbor region. Such labeled data are of very small amount and even partially erroneous. To effectively exploit such weakly-labeled data, we construct a weakly-supervised sparse coding model that jointly conducts feature learning and segmentation. In addition, the geometric constraints are developed for the model to exploit the geometric prior on the local connectivity of region boundaries. The experiments on two benchmark datasets have validated the effectiveness of the proposed approach. Yuhui Quan, Huan Teng, Yan Huang 0031 |
IEEE Signal Process. Lett. | 4 |
| 2019 | Exploiting label consistency in structured sparse representation for classification
Yan Huang 0031, Yuhui Quan, Yong Xu 0007 |
Neural Comput. Appl. | 1 |
| 2019 | Supervised Sparse Coding With Decision ForestabstractBy jointly conducting sparse coding and classifier training, supervised sparse coding has shown its effectiveness in a variety of recognition tasks. However, the existing supervised sparse coding methods often consider linear classification, which limits their discrimination in handling highly nonlinear data. In this letter, we propose a new supervised sparse coding model by incorporating decision tree classifiers. Since decision trees can well deal with the non-linear properties of data, the introduction of decision trees to sparse coding can noticeably improve the discrimination of coding. Meanwhile, sparse coding is able to produce sparse de-correlated features that decision tree is in favor of. For further improvement, we close the loop of sparse coding and decision tree learning with an ensemble framework, which alternatively learns a dictionary for sparse coding and a decision tree for classification. The resulting series of decision trees as well as series of dictionaries are used to construct a decision forest for classification. The proposed method was applied to face recognition and scene classification, and the experimental results have demonstrated its power in comparison with recent supervised sparse coding methods. Yan Huang 0031, Yuhui Quan |
IEEE Signal Process. Lett. | 1 |
| 2016 | Sparse Coding for Classification via Discrimination EnsembleabstractDiscriminative sparse coding has emerged as a promising technique in image analysis and recognition, which couples the process of classifier training and the process of dictionary learning for improving the discriminability of sparse codes. Many existing approaches consider only a simple single linear classifier whose discriminative power is rather weak. In this paper, we proposed a discriminative sparse coding method which jointly learns a dictionary for sparse coding and an ensemble classifier for discrimination. The ensemble classifier is composed of a set of linear predictors and constructed via both subsampling on data and subspace projection on sparse codes. The advantages of the proposed method over the existing ones are multi-fold: better discriminability of sparse codes, weaker dependence on peculiarities of training data, and more expressibility of classifier for classification. These advantages are also justified in the experiments, as our method outperformed several recent methods in several recognition tasks. Yuhui Quan, Yong Xu 0007, Yuping Sun, Yan Huang 0031, Hui Ji 0002 |
CVPR | 4 |
| 2016 | Supervised dictionary learning with multiple classifier integration
Yuhui Quan, Yong Xu 0007, Yuping Sun, Yan Huang 0031 |
Pattern Recognit. | 4 |
| 2015 | Dynamic Texture Recognition via Orthogonal Tensor Dictionary LearningabstractDynamic textures (DTs) are video sequences with stationary properties, which exhibit repetitive patterns over space and time. This paper aims at investigating the sparse coding based approach to characterizing local DT patterns for recognition. Owing to the high dimensionality of DT sequences, existing dictionary learning algorithms are not suitable for our purpose due to their high computational costs as well as poor scalability. To overcome these obstacles, we proposed a structured tensor dictionary learning method for sparse coding, which learns a dictionary structured with orthogonality and separability. The proposed method is very fast and more scalable to high-dimensional data than the existing ones. In addition, based on the proposed dictionary learning method, a DT descriptor is developed, which has better adaptivity, discriminability and scalability than the existing approaches. These advantages are demonstrated by the experiments on multiple datasets. Yuhui Quan, Yan Huang 0031, Hui Ji 0002 |
ICCV | 2 |