VLDB 2026 Research / reviewers in the wild / expert
Huijuan Huang 0001
dblp:134/5178-1
· DBLP profile ↗
9ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0008-9424-1740ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics AssessmentabstractThe aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature—spanning visual perception, cognition, and emotion—poses fundamental challenges. Although aesthetic descriptions offer a viable representation of this complexity, two critical challenges persist: (1) data scarcity and imbalance: existing dataset overly focuses on visual perception and neglects deeper dimensions due to the expensive manual annotation; and (2) model fragmentation: current visual networks isolate aesthetic attributes with multi-branch encoder, while multimodal methods represented by contrastive learning struggle to effectively process long-form textual descriptions. To resolve challenge (1), we first present the Refined Aesthetic Description (RAD) dataset, a large-scale (70k), multi-dimensional structured dataset, generated via an iterative pipeline without heavy annotation costs and easy to scale. To address challenge (2), we propose ArtQuant, an aesthetics assessment framework for artistic image which not only couple isolated aesthetic dimensions through joint description generation, but also better model long-text semantics with the help of LLM decoders. Besides, theoretical analysis confirms this symbiosis: RAD's semantic adequacy (data) and generation paradigm (model) collectively minimize prediction entropy, providing mathematical grounding for the framework. Our approach achieves state-of-the-art performance on several datasets while requiring only 33% of conventional training epochs, narrowing the cognitive gap between artistic image and aesthetic judgment. We will release both code and dataset to support future research. Henglin Liu, Nisha Huang, Chang Liu 0071, Jiangpeng Yan, Huijuan Huang 0001, Jixuan Ying, Tong-Yee Lee, Pengfei Wan 0001, Xiangyang Ji |
AAAI | 5 |
| 2025 | StyleMaster: Stylize Your Video with Artistic Generation and TranslationabstractStyle control has been popular in video generation models. Existing methods often generate videos far from the given style, cause content leakage, and struggle to transfer one video to the desired style. Our first observation is that the style extraction stage matters, whereas existing methods emphasize global style but ignore local textures. In order to bring texture features while preventing content leakage, we filter content-related patches while retaining style ones based on prompt-patch similarity; for global style extraction, we generate a paired style dataset through model illusion to facilitate contrastive learning, which greatly enhances the absolute style consistency. Moreover, to fill in the image-to-video gap, we train a lightweight motion adapter on still videos, which implicitly enhances stylization extent, and enables our image-trained model to be seamlessly applied to videos. Benefited from these efforts, our approach, StyleMaster, not only achieves significant improvement in both style resemblance and temporal coherence, but also can easily generalize to video style transfer with a gray tile ControlNet. Extensive experiments and visualizations demonstrate that StyleMaster significantly outperforms competitors, effectively generating high-quality stylized videos that align with textual content and closely resemble the style of reference images. Zixuan Ye, Huijuan Huang 0001, Xintao Wang 0004, Pengfei Wan 0001, Di Zhang 0026, Wenhan Luo |
CVPR | 2 |
| 2022 | Assessing a Single Image in Reference-Guided Image SynthesisabstractAssessing the performance of Generative Adversarial Networks (GANs) has been an important topic due to its practical significance. Although several evaluation metrics have been proposed, they generally assess the quality of the whole generated image distribution. For Reference-guided Image Synthesis (RIS) tasks, i.e., rendering a source image in the style of another reference image, where assessing the quality of a single generated image is crucial, these metrics are not applicable. In this paper, we propose a general learning-based framework, Reference-guided Image Synthesis Assessment (RISA) to quantitatively evaluate the quality of a single generated image. Notably, the training of RISA does not require human annotations. In specific, the training data for RISA are acquired by the intermediate models from the training procedure in RIS, and weakly annotated by the number of models' iterations, based on the positive correlation between image quality and iterations. As this annotation is too coarse as a supervision signal, we introduce two techniques: 1) a pixel-wise interpolation scheme to refine the coarse labels, and 2) multiple binary classifiers to replace a naïve regressor. In addition, an unsupervised contrastive loss is introduced to effectively capture the style similarity between a generated image and its reference image. Empirical results on various datasets demonstrate that RISA is highly consistent with human preference and transfers well across models. Chaoqun Du, Jiangshan Wang, Huijuan Huang 0001, Pengfei Wan 0001, Gao Huang 0001 |
AAAI | 4 |
| 2021 | Frequency Domain Image Translation: More Photo-realistic, Better Identity-preservingabstractImage-to-image translation has been revolutionized with GAN-based methods. However, existing methods lack the ability to preserve the identity of the source domain. As a result, synthesized images can often over-adapt to the reference domain, losing important structural characteristics and suffering from suboptimal visual quality. To solve these challenges, we propose a novel frequency domain image translation (FDIT) framework, exploiting frequency information for enhancing the image generation process. Our key idea is to decompose the image into low-frequency and high-frequency components, where the high-frequency feature captures object structure akin to the identity. Our training objective facilitates the preservation of frequency information in both pixel space and Fourier spectral space. We broadly evaluate FDIT across five large-scale datasets and multiple tasks including image translation and GAN inversion. Extensive experiments and ablations show that FDIT effectively preserves the identity of the source image, and produces photo-realistic images. FDIT establishes state-of-the-art performance, reducing the average FID score by 5.6% compared to the previous best method. Mu Cai, Hong Zhang 0009, Huijuan Huang 0001, Qichuan Geng, Yixuan Li 0001, Gao Huang 0001 |
ICCV | 3 |
| 2021 | Cascade Image Matting with Deformable Graph Refinement
Zijian Yu, Xuhui Li 0001, Huijuan Huang 0001, Li Chen 0031 |
ICCV | 3 |
| 2014 | Super-resolution mapping via multi-dictionary based sparse representationabstractBased on the spatial dependence assumption, super-resolution mapping can predict the spatial location of land cover classes within mixed pixels. In this paper, we propose a novel super-resolution mapping method via multi-dictionary based sparse representation, which is robust to noise in both the learning and class allocation process. To better distinguish different classes, the distribution modes of different classes are learned separately. A spectral distortion constraint is introduced, combining with reconstruction errors as metrics to perform classification. The experiments prove that our method is superior to other related methods. Huijuan Huang 0001, Jing Yu 0005 |
ICASSP | 1 |
| 2014 | Super-resolution hyperspectral imaging with unknown blurring by low-rank and group-sparse modelingabstractWhen the system blurring is unknown, we propose a novel super-resolution approach of hyperspectral images by low-rank and group-sparse modeling. No high spatial resolution auxiliary data or prior information about blurring isna needed. The proposed method imposes the low-rank model with predefined spectral subspace and group sparse model on different types of high frequency components to take advantage of the shared spatial structure across all spectral bands. The desired high spatial resolution hyperspectral image and blurring kernel are optimized alternatively according to the proposed cost function. Experimental results demonstrate the effectiveness and stability of the proposed method in practical applications. Huijuan Huang 0001, Anthony G. Christodoulou |
ICIP | 1 |
| 2014 | Superresolution Mapping Using Multiple Dictionaries by Sparse RepresentationabstractSuperresolution mapping can predict the spatial location of land cover classes within mixed pixels based on the spatial dependence assumption. We propose a novel superresolution mapping method via multidictionary-based sparse representation, which is robust to noise in both the learning and class-allocation process. In the proposed method, the subpixel number belonging to each class is obtained according to the degree of spectral distortion, and the distribution modes of different classes are treated discriminatorily. The subpixel classification is performed according to the normalized reconstruction errors by the learned multiple distribution dictionaries. The experimental results show that the proposed method has improved accuracy and robustness for real imagery. Huijuan Huang 0001, Jing Yu 0005 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2013 | Super-Resolution Based on Compressive Sensing and Structural Self-Similarity for Remote Sensing ImagesabstractA super-resolution (SR) method based on compressive sensing (CS), structural self-similarity (SSSIM), and dictionary learning is proposed for reconstructing remote sensing images. This method aims to identify a dictionary that represents high resolution (HR) image patches in a sparse manner. Extra information from similar structures which often exist in remote sensing images can be introduced into the dictionary, thereby enabling an HR image to be reconstructed using the dictionary in the CS framework. We use the K-Singular Value Decomposition method to obtain the dictionary and the orthogonal matching pursuit method to derive sparse representation coefficients. To evaluate the effectiveness of the proposed method, we also define a new SSSIM index, which reflects the extent of SSSIM in an image. The most significant difference between the proposed method and traditional sample-based SR methods is that the proposed method uses only a low-resolution image and its own interpolated image instead of other HR images in a database. We simulate the degradation mechanism of a uniform 2 × 2 blur kernel plus a downsampling by a factor of 2 in our experiments. Comparative experimental results with several image-quality-assessment indexes show that the proposed method performs better in terms of the SR effectivity and time efficiency. In addition, the SSSIM index is strongly positively correlated with the SR quality. Zongxu Pan, Jing Yu 0005, Huijuan Huang 0001, Shaoxing Hu, Aiwu Zhang, Hongbing Ma |
IEEE Trans. Geosci. Remote. Sens. | 3 |