VLDB 2026 Research / reviewers in the wild / expert
Xianxu Hou
dblp:186/7985
· DBLP profile ↗
42ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0002-8728-2842ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 22 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LightQANet: Quantized and Adaptive Feature Learning for Low-Light Image EnhancementabstractLow-light image enhancement (LLIE) aims to improve illumination while preserving high-quality color and texture. However, existing methods often fail to extract reliable feature representations due to severely degraded pixel-level information under low-light conditions, resulting in poor texture restoration, color inconsistency, and artifact. To address these challenges, we propose LightQANet, a novel framework that introduces quantized and adaptive feature learning for low-light enhancement, aiming to achieve consistent and robust image quality across diverse lighting conditions. From the static modeling perspective, we design a Light Quantization Module (LQM) to explicitly extract and quantify illumination-related factors from image features. By enforcing structured light factor learning, LQM enhances the extraction of light-invariant representations and mitigates feature inconsistency across varying illumination levels. From the dynamic adaptation perspective, we introduce a Light-Aware Prompt Module (LAPM), which encodes illumination priors into learnable prompts to dynamically guide the feature learning process. LAPM enables the model to flexibly adapt to complex and continuously changing lighting conditions, further improving image enhancement. Extensive experiments on multiple low-light datasets demonstrate that our method achieves state-of-the-art performance, delivering superior qualitative and quantitative results across various challenging lighting scenarios. Xu Wu 0001, Zhihui Lai 0001, Xianxu Hou, Jie Zhou 0009, LinLin Shen |
IEEE Trans. Multim. | 3 |
| 2025 | FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMsabstractMultimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we introduce FaceBench, a dataset featuring hierarchical multi-view and multi-level attributes specifically designed to assess the comprehensive face perception abilities of MLLMs. Initially, we construct a hierarchical facial attribute structure, which encompasses five views with up to three levels of attributes, totaling over 210 attributes and 700 attribute values. Based on the structure, the proposed FaceBench consists of 49,919 visual questionanswering (VQA) pairs for evaluation and 23,841 pairs for fine-tuning. Moreover, we further develop a robust face perception MLLM baseline, Face-LLaVA, by training with our proposed face VQA data. Extensive experiments on various mainstream MLLMs and Face-LLaVA are conducted to test their face perception ability, with results also compared against human performance. The results reveal that, the existing MLLMs are far from satisfactory in understanding the fine-grained facial attributes, while our Face-LLaVA significantly outperforms existing open-source models with a small amount of training data and is comparable to commercial ones like GPT-4o and Gemini. The dataset will be released at https://github.com/CVI-SZU/FaceBench Xusen Ma, Xianxu Hou, Meidan Ding, Yudong Li 0001, Junliang Chen 0002, Wenting Chen, Xiaoyang Peng, LinLin Shen |
CVPR | 3 |
| 2025 | High-Fidelity Editable Portrait Synthesis with 3D GAN InversionabstractThe 3D generative adversarial network (GAN) inversion converts an image into 3D representation to attain high-fidelity reconstruction and facilitate realistic image manipulation within the 3D latent space. However, previous approaches face challenges regarding the trade-off between the reconstruction ability and editability. That is, reversing a real-world image to a low-dimensional latent code would inevitably lead to information loss, and achieving a near-perfect reconstruction using high-rate triplane representation often limits the ability to manipulate the image freely in the latent space. To address these issues, we propose a novel latent conditioning encoder-based framework with the alignment between the low-dimensional latent and high-dimensional triplane. A non-semantic guided editing strategy bridges the intrinsic relation between the latent condition and triplane generation, making it possible to edit the high-dimensional representation by latent manipulation. As a result, our method can achieve high-fidelity reconstruction and editing simultaneously by directly controlling the latent code. Experimental results demonstrate that our approach excels in reconstruction and editing quality compared to previous 3D inversion methods. Furthermore, our method can also edit even real faces with large poses and out-of-domain cases. Jindong Xie, Yupei Lin, Jinbao Wang 0001, Xianxu Hou, LinLin Shen |
ICASSP | 5 |
| 2025 | DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face ParsingabstractFace parsing aims to segment facial images into key components such as eyes, lips, and eyebrows. While existing methods rely on dense pixel-level annotations, such annotations are expensive and labor-intensive to obtain. To reduce annotation cost, we introduce Weakly Supervised Face Parsing (WSFP), a new task setting that performs dense facial component segmentation using only weak supervision, such as image-level labels and natural language descriptions. WSFP introduces unique challenges due to the high co-occurrence and visual similarity of facial components, which lead to ambiguous activations and degraded parsing performance. To address this, we propose DisFaceRep, a representation disentanglement framework designed to separate co-occurring facial components through both explicit and implicit mechanisms. Specifically, we introduce a co-occurring component disentanglement strategy to explicitly reduce dataset-level bias, and a text-guided component disentanglement loss to guide component separation using language supervision implicitly. Extensive experiments on CelebAMask-HQ, LaPa, and Helen demonstrate the difficulty of WSFP and the effectiveness of DisFaceRep, which significantly outperforms existing weakly supervised semantic segmentation methods. The code will be released at https://github.com/CVI-SZU/DisFaceRep. Xianxu Hou, Meidan Ding, Junliang Chen 0002, Kaijun Deng, Jinheng Xie, LinLin Shen |
ACM Multimedia | 2 |
| 2025 | Unknown Pixel Mask Based Fine-tuning of 2D Inpainting Models for Unbounded 3D Scene Generation from a Single ImageabstractConventional 2D inpainting models are trained using masks confined to 2D scenarios, resulting in meaningless content when applied to 3D-specific masks. These 3D-specific masks, termed Unknown Pixels (UP) masks, represent unseen pixels from novel viewpoints that remain obscured in the original input image. Existing methods attempt to mitigate this issue by employing post-processing techniques to transform UP masks into 2D equivalents, frequently suffering from unnatural distortions. To address these issues, we investigate the efficacy of directly training 2D inpainting models with UP masks to circumvent such distortions. In this paper, we introduce a novel framework designed to generate unbounded 3D scenes from a single image, guided by textual descriptions. Our approach leverages fine-tuned inpainting models that iteratively reconstruct incomplete images originating from pure projection. The generated points are then seamlessly integrated into the original point cloud via pixel-wise depth alignment. Extensive evaluations demonstrate that our framework outperforms existing methods in scene quality, processing speed, and memory efficiency. Dezhi Zheng, Kaijun Deng, Xianxu Hou, Jinbao Wang 0001, LinLin Shen |
ACM Multimedia | 3 |
| 2025 | A codebook-driven approach for low-light image enhancementabstractLow-light image enhancement (LLIE) aims to improve low-illumination images. However, existing methods face two challenges: (1) uncertainty in restoration from diverse brightness degradations; (2) loss of texture and color information caused by noise suppression and light enhancement. In this paper, we propose a novel enhancement approach, CodeEnhance, by leveraging discrete codebook priors and image refinement to address these challenges. In particular, we reframe LLIE as learning an image-to-code mapping from low-light images to discrete codebook, which has been learned from high-quality images. To enhance this process, a Semantic Embedding Module (SEM) is introduced to integrate semantic information with low-level features, and a Codebook Shift (CS) mechanism, designed to adapt the pre-learned codebook to better suit the distinct characteristics of our low-light dataset. Additionally, we present an Interactive Feature Transformation (IFT) module to refine texture and color information during image reconstruction, allowing for interactive enhancement based on user preferences. Extensive experiments on both real-world and synthetic benchmarks demonstrate that the incorporation of prior knowledge and controllable information transfer significantly enhances LLIE performance in terms of quality and fidelity. The proposed CodeEnhance exhibits superior robustness to various degradations, including uneven illumination, noise, and color distortion. The code can be obtained from https://github.com/csxuwu/CodeEnhance or https://www.scholat.com/laizhihui.cn . Xu Wu 0001, Xianxu Hou, Zhihui Lai 0001, Jie Zhou 0009, Witold Pedrycz, LinLin Shen |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | CLIMS++: Cross Language Image Matching with Automatic Context Discovery for Weakly Supervised Semantic Segmentation
Jinheng Xie, Songhe Deng, Xianxu Hou, Zhaochuan Luo, LinLin Shen, Yawen Huang, Yefeng Zheng 0001, Zheng Shou 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Frequency Restoration and Modality Enforcement towards Resisting-corruption Multimodal Sentiment AnalysisabstractFor Multimodal Sentiment Analysis (MSA), previous methods concentrate on designing sophisticated fusion strategies and performing representation learning across heterogeneous modalities, aiming to leverage multimodal signals to detect human sentiment. However, these approaches fail to address the long-standing issue of corrupted modal details in videos, which may be caused by the challenge of the excessive loss of emotionally relevant semantics resulted from the degradation of detailed information. In this work, we aim to improve the robustness capacity of resisting corruption in MSA, by introducing a Hierarchical Frequency Restoration and Adaptive Modality Enforcement (HFR-AME) approach. The HFR-AME progressively recovers blurred detailed cues in each modality while enhancing the discriminative power of modal representations. Specifically, to reconstruct distinct frequency band features, we propose to equip the HFR module with a key component called the Frequency Multimodal UNet (FM-UNet), so as to utilize complementary modal features as conditions. This meticulous restoration process, performed from low to high frequency, facilitates the comprehensive recovery of intricate details. Meanwhile, to adaptively integrate these diverse frequency features, we introduce the AME module to enhance the beneficial modal frequencies while suppressing irrelevant ones, with the goal of strengthening the restored modal representations. Extensive experiments show our HFR-AME outperforms state-of-the-art methods on the CMU-MOSI and CMU-MOSEI datasets, improving 7-class accuracy by 0.5% and 0.6%, respectively. Further analysis also confirms its cross-lingual generalization and competitive computational efficiency. Our code is made available at https://github.com/nianhua20/HFR-AME . Weicheng Xie 0001, Haijian Liang, Zenghao Niu, Xianxu Hou, Siyang Song, Zitong Yu, LinLin Shen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Pseudo Training Data Generation for Unsupervised Cell Membrane Segmentation in Immunohistochemistry ImagesabstractIn the realm of clinical diagnostics and medical research, quantitative assessment of membrane activity in immunohistochemistry (IHC) images is standard practice. Despite a high demand for cell membrane segmentation, only a few algorithms have been developed, and there is a lack of open datasets in this field. In this paper, we propose a three-stage unsupervised framework to accurately segment positive cell membranes in IHC images. Our approach transforms the unsupervised segmentation task into a supervised one by generating pseudo-paired training data using Voronoi diagrams and CycleGAN. Additionally, we introduce a dual encoder segmentation model with domain adaptation modules to mitigate the domain shift between generated images and real images. To our best knowledge, this is the first work focusing on unsupervised learning for IHC cell membrane segmentation. Extensive experiments and ablation studies on our newly built IHC cell membrane segmentation dataset validate the effectiveness of our framework. Yanjia Kan, Yunze Wang, Silin Chen, Albert Zhou, Xianxu Hou, Jingxin Liu 0005 |
BIBM | 7 |
| 2024 | Dynamic Data Sampler for Cross-Language Transfer Learning in Large Language ModelsabstractLarge Language Models (LLMs) have gained significant attention in the field of natural language processing (NLP) due to their wide range of applications. However, training LLMs for languages other than English poses significant challenges, due to the difficulty in acquiring large-scale corpus and the requisite computing resources. In this paper, we propose ChatFlow, a cross-language transfer-based LLM, to address these challenges and train large Chinese language models in a cost-effective manner. We employ a mix of Chinese, English, and parallel corpus to continuously train the LLaMA2 model, aiming to align cross-language representations and facilitate the knowledge transfer specifically to the Chinese language model. In addition, we use a dynamic data sampler to progressively transition the model from unsupervised pre-training to supervised fine-tuning. Experimental results demonstrate that our approach accelerates model convergence and achieves superior performance. We evaluate ChatFlow on popular Chinese and English benchmarks, the results indicate that it outperforms other Chinese models post-trained on LLaMA-2-7B. Yudong Li 0001, Zhe Zhao 0006, LinLin Shen, Cheng Hou, Xianxu Hou |
ICASSP | 7 |
| 2024 | A Dataset and Model for Realistic License Plate Deblurring
Haoyan Gong, Yuzheng Feng, Xianxu Hou, Jingxin Liu 0005, Hongbin Liu 0007 |
IJCAI | 4 |
| 2024 | FLIP-80M: 80 Million Visual-Linguistic Pairs for Facial Language-Image Pre-TrainingabstractWhile significant progress has been made in multi-modal learning driven by large-scale image-text datasets, there is still a noticeable gap in the availability of such datasets within the facial domain. To facilitate and advance the field of facial representation learning, we present FLIP-80M, a large-scale visual-linguistic dataset comprising over 80 million face images paired with text descriptions. FLIP-80M is constructed by leveraging the large openly available image-text-pair dataset LAION-5B and a mixed-method approach to filter face-related pairs from both visual and linguistic perspectives. Our curation process involves face detection, face caption classification, text de-noising, and synthesis-based image augmentation. As a result, FLIP-80M stands as the largest face-text dataset to date. To evaluate the potential of our dataset, we fine-tune the CLIP model using the proposed FLIP-80M, to create FLIP (Facial Language-Image Pretraining) and assess its representation capabilities across various downstream tasks. Our experiments demonstrate that our FLIP model achieves state-of-the-art results in a range of face analysis tasks, including face parsing, face alignment, and face attribute classification. The dataset and models are available at https://github.com/ydli-ai/FLIP. Yudong Li 0001, Xianxu Hou, Dezhi Zheng, LinLin Shen, Zhe Zhao 0006 |
ACM Multimedia | 2 |
| 2023 | StyleGene: Crossover and Mutation of Region-level Facial Genes for Kinship Face SynthesisabstractHigh-fidelity kinship face synthesis has many potential applications, such as kinship verification, missing child identification, and social media analysis. However, it is challenging to synthesize high-quality descendant faces with genetic relations due to the lack of large-scale, high-quality annotated kinship data. This paper proposes RFG (Region-level Facial Gene) extraction framework to address this issue. We propose to use IGE (Image-based Gene Encoder), LGE (Latent-based Gene Encoder) and Gene Decoder to learn the RFGs of a given face image, and the relationships between RFGs and the latent space of Style-GAN2. As cycle-like losses are designed to measure the$\mathcal{L}_{2}$distances between the output of Gene Decoder and image encoder, and that between the output of LGE and IGE, only face images are required to train our framework, i.e. no paired kinship face data is required. Based upon the proposed RFGs, a crossover and mutation module is further designed to inherit the facial parts of parents. A Gene Pool has also been used to introduce the variations into the mutation of RFGs. The diversity of the faces of descendants can thus be significantly increased. Qualitative, quantitative, and subjective experiments on FIW, TSKinFace, and FF-Databases clearly show that the quality and diversity of kinship faces generated by our approach are much better than the existing state-of-the-art methods. Xianxu Hou, Zepeng Huang, LinLin Shen |
CVPR | 2 |
| 2023 | StyleAU: StyleGAN based Facial Action Unit Manipulation for Expression EditingabstractFacial expression editing has a wide range of applications, such as emotion detection, human-computer interaction, and social entertainment. However, existing expression editing methods either fail to allow for fine-grained editing, resulting in unnatural and unrealistic facial expressions, or generate artifacts and blurs, leading to poor image quality. In this paper, we propose a novel framework called StyleAU, which is based on StyleGAN and facial action units, to address these problems. Our framework leverages the pre-trained StyleGAN prior knowledge to enable action unit editing of the face in the StyleGAN latent space, allowing precise expression editing. In addition, we use an encoder to extract multi-scale content features to achieve high-fidelity image reconstruction. Our approach qualitatively and quantitatively outperforms competing methods for action unit manipulation and expression editing. Yanliang Guo, Xianxu Hou, Feng Liu 0013, LinLin Shen, Lei Wang 0018, Zhen Wang 0009, Peng Liu 0039 |
IJCB | 2 |
| 2023 | Domain Adaptation of Digital Pathology Images using Joint Stain Color and Image Quality ConstraintsabstractDigital pathology diagnosis systems face significant domain shift problems that hinder their performance on new datasets. Existing methods for aligning digital pathology images from different domains mainly focus on stain color and overlook the potential domain shifts caused by variations in image quality. To address this issue, we propose a novel parametric model that incorporates both stain color and image quality constraints for domain adaptation of digital pathology images. We evaluate our approach on the domain adaptive mitosis detection task through extensive experiments and ablation studies, showing that our method outperforms other state-of-the-art methods. Jingxin Liu 0005, Xianxu Hou |
ICIP | 3 |
| 2023 | MaskDiffuse: Text-Guided Face Mask Removal Based on Diffusion Models
Jingxia Lu, Xianxu Hou, Zhibin Peng, LinLin Shen, Lixin Fan |
PRCV (6) | 2 |
| 2023 | DuAT: Dual-Aggregation Transformer Network for Medical Image Segmentation
Zhongxing Xu, Qiming Huang, Jinfeng Wang 0008, Xianxu Hou, Jionglong Su, Jingxin Liu 0005 |
PRCV (5) | 5 |
| 2023 | Learning Adapters for Text-Guided Portrait Stylization with Pretrained Diffusion Models
Mintu Yang, Xianxu Hou, LinLin Shen, Lixin Fan |
PRCV (1) | 2 |
| 2023 | Deep generative image priors for semantic face manipulationabstractPrevious works on generative adversarial networks (GANs) mainly focus on how to synthesize high-fidelity images. In this paper, we present a framework to leverage the knowledge learned by GANs for semantic face manipulation. In particular, we propose to control the semantics of synthesized faces by adapting the latent codes with an attribute prediction model. Moreover, in order to achieve a more accurate estimation of different facial attributes, we propose to pretrain the attribute prediction model by inverting the synthesized face images back to the GAN latent space. As a result, our method explicitly considers the semantics encoded in the latent space of a pretrained GAN and is able to faithfully edit various attributes like eyeglasses, smiling, bald, age, mustache and gender for high-resolution face images. Extensive experiments show that our method has superior performance compared to state of the art for both face attribute prediction and semantic face manipulation. Xianxu Hou, LinLin Shen, Zhong Ming 0001, Guoping Qiu |
Pattern Recognit. | 1 |
| 2023 | TextFace: Text-to-Style Mapping Based Face Generation and ManipulationabstractAs a subtopic of text-to-image synthesis, text-to-face generation has great potential in face-related applications. In this paper, we propose a generic text-to-face framework, namely, TextFace, to achieve diverse and high-quality face image generation from text descriptions. We introduce text-to-style mapping, a novel method where the text description can be directly encoded into the latent space of a pretrained StyleGAN. Guided by our text-image similarity matching and face captioning-based text alignment, the textual latent code can be fed into the generator of a well-trained StyleGAN to produce diverse face images with high resolution (1024×1024). Furthermore, our model inherently supports semantic face editing using text descriptions. Finally, experimental results quantitatively and qualitatively demonstrate the superior performance of our model. Xianxu Hou, Yudong Li 0001, LinLin Shen |
IEEE Trans. Multim. | 1 |
| 2023 | Lifelong Age Transformation With a Deep Generative PriorabstractIn this paper, we consider the lifelong age progression and regression task, which requires to synthesize a persons appearance across a wide range of ages. We propose a simple yet effective learning framework to achieve this by exploiting the prior knowledge of faces captured by well-trained generative adversarial networks (GANs). Specifically, we first utilize a pretrained GAN to synthesize face images with different ages, with which we then learn to model the conditional aging process in the GAN latent space. Moreover, we also introduce a cycle consistency loss in the GAN latent space to preserve a persons identity. As a result, our model can reliably predict a person's appearance for different ages by modifying both shape and texture of the head. Both qualitative and quantitative experimental results demonstrate the superiority of our method over concurrent works. Furthermore, we demonstrate that our approach can also achieve high-quality age transformation for painting portraits and cartoon characters without additional age annotations. Xianxu Hou, Hanbang Liang, LinLin Shen, Zhong Ming 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | CLIMS: Cross Language Image Matching for Weakly Supervised Semantic SegmentationabstractIt has been widely known that CAM (Class Activation Map) usually only activates discriminative object regions and falsely includes lots of object-related backgrounds. As only a fixed set of image-level object labels are available to the WSSS (weakly supervised semantic segmentation) model, it could be very difficult to suppress those diverse background regions consisting of open set objects. In this paper, we propose a novel Cross Language Image Matching (CLIMS) framework, based on the recently introduced Contrastive Language-Image Pre-training (CLIP) model, for WSSS. The core idea of our framework is to introduce natural language supervision to activate more complete object regions and suppress closely-related open background regions. In particular, we design object, background region and text label matching losses to guide the model to excite more reasonable object regions for CAM of each category. In addition, we design a co-occurring background suppression loss to prevent the model from activating closely-related background regions, with a predefined set of class-related background text descriptions. These designs enable the proposed CLIMS to generate a more complete and compact activation map for the target objects. Extensive experiments on PASCAL VOC2012 dataset show that our CLIMS significantly outperforms the previous state-of-the-art methods. Code will be available at https://github.com/CVI-SZU/CLIMS. Jinheng Xie, Xianxu Hou, Kai Ye 0004, LinLin Shen |
CVPR | 2 |
| 2022 | C2 AM: Contrastive learning of Class-agnostic Activation Map for Weakly Supervised Object Localization and Semantic SegmentationabstractWhile class activation map (CAM) generated by image classification network has been widely used for weakly su-pervised object localization (WSOL) and semantic segmentation (WSSS), such classifiers usually focus on discriminative object regions. In this paper, we propose Contrastive learning for Class-agnostic Activation Map (C2AM) generation only using unlabeled image data, without the involvement of image-level supervision. The core idea comes from the observation that i) semantic information of fore-ground objects usually differs from their backgrounds; ii) foreground objects with similar appearance or background with similar color/texture have similar representations in the feature space. We form the positive and negative pairs based on the above relations and force the network to disentangle foreground and background with a class-agnostic activation map using a novel contrastive loss. As the network is guided to discriminate cross-image foreground-background, the class-agnostic activation maps learned by our approach generate more complete object regions. We successfully extracted from C2AM class-agnostic object bounding boxes for object localization and background cues to refine CAM generated by classification network for semantic segmentation. Extensive experiments on CUB-200-2011, ImageNet-1K, and PASCAL VOC2012 datasets show that both WSOL and WSSS can benefit from the proposed C2AM. Code will be available at https://github.com/CVI-SZUICCAM. Jinheng Xie, Jianfeng Xiang, Junliang Chen 0002, Xianxu Hou, LinLin Shen |
CVPR | 4 |
| 2022 | RamGAN: Region Attentive Morphing GAN for Region-Level Makeup Transfer
Jianfeng Xiang, Junliang Chen 0002, Wenshuang Liu, Xianxu Hou, LinLin Shen |
ECCV (22) | 4 |
| 2022 | Talk2Face: A Unified Sequence-based Framework for Diverse Face Generation and Analysis TasksabstractFacial analysis is an important domain in computer vision and has received extensive research attention. For numerous downstream tasks with different input/output formats and modalities, existing methods usually design task-specific architectures and train them using face datasets collected in the particular task domain. In this work, we proposed a single model, Talk2Face, to simultaneously tackle a large number of face generation and analysis tasks, e.g. text guided face synthesis, face captioning and age estimation. Specifically, we cast different tasks into a sequence-to-sequence format with the same architecture, parameters and objectives. While text and facial images are tokenized to sequences, the annotation labels of faces for different tasks are also converted to natural languages for unified representation. We collect a set of 2.3M face-text pairs from available datasets across different tasks, to train the proposed model. Uniform templates are then designed to enable the model to perform different downstream tasks, according to the task context and target. Experiments on different tasks show that our model achieves better face generation and caption performances than SOTA approaches. On age estimation and multi-attribute classification, our model reaches competitive performance with those models specially designed and trained for these particular tasks. In practice, our model is much easier to be deployed to different facial analysis related tasks. Code and dataset will be available at https://github.com/ydli-ai/Talk2Face. Yudong Li 0001, Xianxu Hou, Zhe Zhao 0006, LinLin Shen, Xuefeng Yang, Kimmo Yan |
ACM Multimedia | 2 |
| 2022 | GuidedStyle: Attribute knowledge guided style manipulation for semantic face editing
Xianxu Hou, Hanbang Liang, LinLin Shen, Zhihui Lai 0001, Jun Wan 0005 |
Neural Networks | 1 |
| 2021 | GazeFlow: Gaze Redirection with Normalizing FlowsabstractGaze estimation often requires a large scale datasets with well annotated gaze information to train the estimator. However, such a dataset requires costive annotation and is usually very difficult to collect. Therefore, a number of gaze redirection approaches have been proposed to address such a problem. However, existing methods lack the ability to precisely synthesize images with target gaze and head pose in complex lighting scenes. As a powerful technique to model the distribution of given data, normalizing flows have the ability to generate photo-realistic images and provide flexible latent space manipulation. In this work, we present a novel flow-based generative model, GazeFlow11The code will be made available at https://github.com/CVI-SZU/GazeFlow, for gaze redirection. The visual results of gaze redirection show that the quality of eye images synthesized by GazeFlow is significantly higher than that of other approaches like Deep Warp and PRGAN. Our approach has also been applied to augment the training data to improve the accuracy of gaze estimators and significant improvement has been achieved for both within dataset and cross dataset experiments. Hanbang Liang, Xianxu Hou, LinLin Shen |
IJCNN | 3 |
| 2021 | SSFlow: Style-guided Neural Spline Flows for Face Image ManipulationabstractSignificant progress has been made in high-resolution and photo-realistic image generation by Generative Adversarial Networks (GANs). However, the generation process is still lack of control, which is crucial for semantic face editing. Furthermore, it remains challenging to edit target attributes and preserve the identity at the same time. In this paper, we propose SSFlow to achieve identity-preserved semantic face manipulation in StyleGAN latent space based on conditional Neural Spline Flows. To further improve the performance of Neural Spline Flows on such task, we also propose Constractive Squash component and Blockwise 1 x 1 Convolution layer. Moreover, unlike other conditional flow-based approaches that require facial attribute labels during inference, our method can achieve label-free manipulation in a more flexible way. As a result, our methods are able to perform well-disentangled edits along various attributes, and generalize well for both real and artistic face image manipulation. Qualitative and quantitative evaluations show the advantages of our method for semantic face manipulation over state-of-the-art approaches. Hanbang Liang, Xianxu Hou, LinLin Shen |
ACM Multimedia | 2 |
| 2021 | Robust facial landmark detection by cross-order cross-semantic deep network
Jun Wan 0005, Zhihui Lai 0001, LinLin Shen, Jie Zhou 0009, Can Gao, Xianxu Hou |
Neural Networks | 7 |
| 2020 | End-to-End Illuminant Estimation Based on Deep Metric LearningabstractPrevious deep learning approaches to color constancy usually directly estimate illuminant value from input image. Such approaches might suffer heavily from being sensitive to the variation of image content. To overcome this problem, we introduce a deep metric learning approach named Illuminant-Guided Triplet Network (IGTN) to color constancy. IGTN generates an Illuminant Consistent and Discriminative Feature (ICDF) for achieving robust and accurate illuminant color estimation. ICDF is composed of semantic and color features based on a learnable color histogram scheme. In the ICDF space, regardless of the similarities of their contents, images taken under the same or similar illuminants are placed close to each other and at the same time images taken under different illuminants are placed far apart. We also adopt an end-to-end training strategy to simultaneously group image features and estimate illuminant value, and thus our approach does not have to classify illuminant in a separate module. We evaluate our method on two public datasets and demonstrate our method outperforms state-of-the-art approaches. Furthermore, we demonstrate that our method is less sensitive to image appearances, and can achieve more robust and consistent results than other methods on a High Dynamic Range dataset. Bolei Xu, Jingxin Liu 0005, Xianxu Hou, Guoping Qiu |
CVPR | 3 |
| 2020 | VHS to HDTV Video Translation Using Multi-task Adversarial Learning
Hongming Luo, Guangsen Liao, Xianxu Hou, Fei Zhou 0001, Guoping Qiu |
MMM (1) | 3 |
| 2020 | Class-aware domain adaptation for improving adversarial robustness
Xianxu Hou, Jingxin Liu 0005, Bolei Xu, Guoping Qiu |
Image Vis. Comput. | 1 |
| 2020 | End-to-End Single Image Fog Removal Using Enhanced Cycle Consistent Adversarial NetworksabstractSingle image defogging is a classical and challenging problem in computer vision. Existing methods towards this problem mainly include handcrafted priors based methods that rely on the use of the atmospheric degradation model and learning-based approaches that require paired fog-fogfree training example images. In practice, however, prior-based methods are prone to failure due to their own limitations and paired training data are extremely difficult to acquire. Moreover, there are few studies on the unpaired trainable defogging network in this field. Thus, inspired by the principle of CycleGAN network, we have developed an end-to-end learning system that uses unpaired fog and fogfree training images, adversarial discriminators and cycle consistency losses to automatically construct a fog removal system. Similar to CycleGAN, our system has two transformation paths; one maps fog images to a fogfree image domain and the other maps fogfree images to a fog image domain. Instead of one stage mapping, our system uses a two stage mapping strategy in each transformation path to enhance the effectiveness of fog removal. Furthermore, we make explicit use of prior knowledge in the networks by embedding the atmospheric degradation principle and a sky prior for mapping fogfree images to the fog images domain. In addition, we also contribute the first real world nature fog-fogfree image dataset for defogging research. Our multiple real fog images dataset (MRFID) contains images of 200 natural outdoor scenes. For each scene, there is one clear image and corresponding four foggy images of different fog densities manually selected from a sequence of images taken by a fixed camera over the course of one year. Qualitative and quantitative comparison against several state-of-the-art methods on both synthetic and real world images demonstrate that our approach is effective and performs favorably for recovering a clear image from a foggy image. Wei Liu 0123, Xianxu Hou, Jiang Duan, Guoping Qiu |
IEEE Trans. Image Process. | 2 |
| 2020 | Attention by Selection: A Deep Selective Attention Approach to Breast Cancer ClassificationabstractDeep learning approaches are widely applied to histopathological image analysis due to the impressive levels of performance achieved. However, when dealing with high-resolution histopathological images, utilizing the original image as input to the deep learning model is computationally expensive, while resizing the original image to achieve low resolution incurs information loss. Some hard-attention based approaches have emerged to select possible lesion regions from images to avoid processing the original image. However, these hard-attention based approaches usually take a long time to converge with weak guidance, and valueless patches may be trained by the classifier. To overcome this problem, we propose a deep selective attention approach that aims to select valuable regions in the original images for classification. In our approach, a decision network is developed to decide where to crop and whether the cropped patch is necessary for classification. These selected patches are then trained by the classification network, which then provides feedback to the decision network to update its selection policy. With such a co-evolution training strategy, we show that our approach can achieve a fast convergence rate and high classification accuracy. Our approach is evaluated on a public breast cancer histopathological image database, where it demonstrates superior performance compared to state-of-the-art deep learning approaches, achieving approximately 98% classification accuracy while only taking 50% of the training time of the previous hard-attention approach. Bolei Xu, Jingxin Liu 0005, Xianxu Hou, Jonathan M. Garibaldi, Ian O. Ellis, Andrew R. Green, LinLin Shen, Guoping Qiu |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Dual Adaptive Pyramid Network for Cross-Stain Histopathology Image Segmentation
Xianxu Hou, Jingxin Liu 0005, Bolei Xu, Xin Chen 0003, Mohammad Ilyas, Ian O. Ellis, Jonathan M. Garibaldi, Guoping Qiu |
MICCAI (2) | 1 |
| 2019 | Character Prediction in TV Series via a Semantic Projection Network
Ke Sun 0006, Zhuo Lei, Jiasong Zhu, Xianxu Hou, Guoping Qiu |
MMM (1) | 4 |
| 2019 | Improving variational autoencoder with deep feature consistent and generative adversarial training
Xianxu Hou, Ke Sun 0006, LinLin Shen, Guoping Qiu |
Neurocomputing | 1 |
| 2019 | Deep reinforcement learning-based patch selection for illuminant estimation
Bolei Xu, Jingxin Liu 0005, Xianxu Hou, Guoping Qiu |
Image Vis. Comput. | 3 |
| 2018 | Sub-window Box FilterabstractBox filter is a fundamental filter in image processing. However, it can not preserve edges or corners. In this paper, we present a simple but novel method that can make box filter both edge and corner preserving. More specifically, we combine the box filter with sub-window regression to achieve this task. This filter inherits some properties from box filter, such as O(1) running time with respect to the window radius. After analyzing its parameters, we show its corner and edge preserving property on real images and compare it with Guided filter. Yuanhao Gong, Xianxu Hou, Guoping Qiu |
VCIP | 3 |
| 2018 | Direct Application of Convolutional Neural Network Features to Image Quality AssessmentabstractWe take advantage of the popularity of deep convolutional neural networks (CNNs) and have developed a very simple image quality assessment method that rivals state of the art. We show that convolutional layer outputs (deep features) of a CNN compute the local structural information of spatial regions of different sizes in the input image. The learned convolutional kernels contain a much richer set of weights thus capturing much more local structural information than hand crafted ones. As the deep features learned from large datasets already contain very rich multi-resolutional structural image information, they can be directly used to calculate visual distortion of an image and it is not necessary to introduce further complicated computational process. We will present experimental results to demonstrate that this is indeed the case, and that simple cosine distance of the deep features is as good as state the art methods for full reference image quality assessment. Xianxu Hou, Ke Sun 0006, Yuanhao Gong, Jonathan M. Garibaldi, Guoping Qiu |
VCIP | 1 |
| 2017 | Learning deep semantic attributes for user video summarizationabstractThis paper presents a Semantic Attribute assisted video SUMmarization framework (SASUM). Compared with traditional methods, SASUM has several innovative features. Firstly, we use a natural language processing tool to discover a set of keywords from an image and text corpora to form the semantic attributes of visual contents. Secondly, we train a deep convolution neural network to extract visual features as well as predict the semantic attributes of video segments which enables us to represent video contents with visual and semantic features simultaneously. Thirdly, we construct a temporally constrained video segment affinity matrix and use a partially near duplicate image discovery technique to cluster visually and semantically consistent video frames together. These frame clusters can then be condensed to form an informative and compact summary of the video. We will present experimental results to show the effectiveness of the semantic attributes in assisting the visual features in video summarization and our new technique achieves state-of-the-art performance. Ke Sun 0006, Jiasong Zhu, Zhuo Lei, Xianxu Hou, Qian Zhang 0018, Jiang Duan, Guoping Qiu |
ICME | 4 |
| 2017 | Deep Feature Consistent Variational AutoencoderabstractWe present a novel method for constructing Variational Autoencoder (VAE). Instead of using pixel-by-pixel loss, we enforce deep feature consistency between the input and the output of a VAE, which ensures the VAE's output to preserve the spatial correlation characteristics of the input, thus leading the output to have a more natural visual appearance and better perceptual quality. Based on recent deep learning works such as style transfer, we employ a pre-trained deep convolutional neural network (CNN) and use its hidden features to define a feature perceptual loss for VAE training. Evaluated on the CelebA face dataset, we show that our model produces better results than other methods in the literature. We also show that our method can produce latent vectors that can capture the semantic information of face expressions and can be used to achieve state-of-the-art performance in facial attribute prediction. Xianxu Hou, LinLin Shen, Ke Sun 0006, Guoping Qiu |
WACV | 1 |