Shuowen Hu

dblp:73/10523 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 6 since 2021Security and privacy · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 2D-3D Attention and Entropy for Pose Robust 2D Facial Recognition
abstract
Despite recent advances in facial recognition, there remains a fundamental issue concerning degradations in performance due to substantial perspective (pose) differences between enrollment and query (probe) imagery. Therefore, we propose a novel domain adaptive framework to facilitate improved performances across large discrepancies in pose by enabling imagebased (2D) representations to infer properties of inherently pose invariant point cloud (3D) representations. Specifically, our proposed framework achieves better pose invariance by using (1) a shared (joint) attention mapping to emphasize common patterns that are most correlated between 2D facial images and 3D facial data and (2) a joint entropy regularizing loss to promote better consistency—enhancing correlations among the intersecting 2D and 3D representations—by leveraging both attention maps. This framework is evaluated on FaceScape and ARL-VTF datasets, where it outperforms competitive methods by achieving profile ($90^{\circ}+$) TAR @ 1% FAR improvements of at least $\mathbf{7. 1 \%}$ and $\mathbf{1. 5 7 \%}$, respectively.
John Brennan Peace, Shuowen Hu, Benjamin S. Riggan
FG2
2025 Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
abstract
Detecting vehicles in aerial imagery is a critical task with applications in traffic monitoring, urban planning, and defense intelligence. Deep learning methods have provided state-of-the-art (SOTA) results for this application. However, a significant challenge arises when models trained on data from one geographic region fail to generalize effectively to other areas. Variability in factors such as environmental conditions, urban layouts, road networks, vehicle types, and image acquisition parameters (e.g., resolution, lighting, and angle) leads to domain shifts that degrade model performance. This paper proposes a novel method that uses generative AI to synthesize high-quality aerial images and their labels, improving detector training through data augmentation. Our key contribution is the development of a multi-stage, multi-modal knowledge transfer framework utilizing fine-tuned latent diffusion models (LDMs) to mitigate the distribution gap between the source and target environments. Extensive experiments across diverse aerial imagery domains show consistent performance improvements in AP50 over supervised learning on source domain data, weakly supervised adaptation methods, unsupervised domain adaptation methods, and open-set object detectors by 4-23%, 6-10%, 7-40%, and more than 50%, respectively. Furthermore, we introduce two newly annotated aerial datasets from New Zealand and Utah to support further research in this field. Project page is available at: https://humansensinglab.github.io/AGenDA
Minhyek Jeon, Shuowen Hu, Zheyang Qin, Shayok Chakraborty, Stanislav Panev, Celso de Melo, Fernando De la Torre
ICCV3
2023 Open-Set Automatic Target Recognition
abstract
Automatic Target Recognition (ATR) is a category of computer vision algorithms which attempts to recognize targets on data obtained from different sensors. ATR algorithms are extensively used in real-world scenarios such as military and surveillance applications. Existing ATR algorithms are developed for traditional closed-set methods where training and testing have the same class distribution. Thus, these algorithms have not been robust to unknown classes not seen during the training phase, limiting their utility in real-world applications. To this end, we propose an Open-set Automatic Target Recognition framework where we enable open-set recognition capability for ATR algorithms. In addition, we introduce a plugin Category-aware Binary Classifier (CBC) module to effectively tackle unknown classes seen during inference. The proposed CBC module can be easily integrated with any existing ATR algorithms and can be trained in an end-to-end manner. Experimental results show that the proposed approach outperforms many open-set methods on the DSIAC and CIFAR-10 datasets. To the best of our knowledge, this is the first work to address the open-set classification problem for ATR algorithms. Source code is available at: https://github.com/bardisafa/Open-set-ATR.
Bardia Safaei 0002, Vibashan VS, Celso de Melo, Shuowen Hu, Vishal M. Patel
ICASSP4
2022 FAR: Fourier Aerial Video Recognition
Divya Kothandaraman, Tianrui Guan, Xijun Wang 0002, Shuowen Hu, Ming C. Lin, Dinesh Manocha
ECCV (37)4
2022 Fake Satellite Image Detection via Parallel Subspace Learning (PSL)
abstract
A new method, called PSL-DefakeHop, is proposed to detect fake satellite images based on the parallel subspace learning (PSL) framework in this work. The DefakeHop method was developed previously for detection of deepfake generated faces under the successive subspace learning (SSL) framework. PSL is proposed to extract features from responses of multiple single-stage filter banks (or called PixelHops), which operate in parallel, and it improves SSL that extracts features from multi-stage cascaded filter banks. PSL has two advantages. First, PSL preserves discriminant features often lie in high-frequency channels, which are however ignored by SSL. Second, decisions from multiple filter banks can be ensembled to further improve detection accuracy. To demonstrate the effectiveness of the proposed PSL-DefakeHop method, we evaluate it on the UW Fake Satellite Image dataset and observe perfect classification performance (i.e., 100% F1 score, precision and recall).
Hong-Shuo Chen, Kaitai Zhang, Shuowen Hu, Suya You, C.-C. Jay Kuo
ISCAS3
2022 Meta-UDA: Unsupervised Domain Adaptive Thermal Object Detection using Meta-Learning
abstract
Object detectors trained on large-scale RGB datasets are being extensively employed in real-world applications. However, these RGB-trained models suffer a performance drop under adverse illumination and lighting conditions. Infrared (IR) cameras are robust under such conditions and can be helpful in real-world applications. Though thermal cameras are widely used for military applications and increasingly for commercial applications, there is a lack of robust algorithms to robustly exploit the thermal imagery due to the limited availability of labeled thermal data. In this work, we aim to enhance the object detection performance in the thermal domain by leveraging the labeled visible domain data in an Unsupervised Domain Adaptation (UDA) setting. We propose an algorithm agnostic meta-learning framework to improve existing UDA methods instead of proposing a new UDA strategy. We achieve this by meta-learning the initial condition of the detector, which facilitates the adaptation process with fine updates without overfitting or getting stuck at local optima. However, meta-learning the initial condition for the detection scenario is computationally heavy due to long and intractable computation graphs. Therefore, we propose an online meta-learning paradigm which performs online updates resulting in a short and tractable computation graph. To this end, we demonstrate the superiority of our method over many baselines in the UDA setting, producing a state-of-the-art thermal detector for the KAIST and DSIAC datasets.
Vibashan VS, Domenick Poster, Suya You, Shuowen Hu, Vishal M. Patel
WACV4
2021 Heterogeneous Face Frontalization via Domain Agnostic Learning
abstract
Recent advances in deep convolutional neural networks (DCNNs) have shown impressive performance improvements on thermal to visible face synthesis and matching problems. However, current DCNN-based synthesis models do not perform well on thermal faces with large pose variations. In order to deal with this problem, heterogeneous face frontal-ization methods are needed in which a model takes a thermal profile face image and generates a frontal visible face. This is an extremely difficult problem due to the large domain as well as large pose discrepancies between the two modalities. Despite its applications in biometrics and surveillance, this problem is relatively unexplored in the literature. We propose a domain agnostic learning-based generative adversarial network (DAL-GAN) which can synthesize frontal views in the visible domain from thermal faces with pose variations. DAL-GAN consists of a generator with an auxiliary classifier and two discriminators which capture both local and global texture discriminations for better synthesis. A contrastive constraint is enforced in the latent space of the generator with the help of a dual-path training strategy, which improves the feature vector's discrimination. Finally, a multi-purpose loss function is utilized to guide the network in synthesizing identity-preserving cross-domain frontalization. Extensive experimental results demonstrate that DAL-GAN can generate better quality frontal views compared to the other baseline methods.
Xing Di, Shuowen Hu, Vishal M. Patel
FG2
2021 Simultaneous Face Hallucination and Translation for Thermal to Visible Face Verification using Axial-GAN
abstract
Existing thermal-to-visible face verification approaches expect the thermal and visible face images to be of similar resolution. This is unlikely in real-world long-range surveillance systems since humans are distant from the cameras. To address this issue, we introduce the task of thermal- to-visible face verification from low-resolution thermal images. Furthermore, we propose Axial-Generative Adversarial Network (Axial-GAN) to synthesize high-resolution visible images for matching. In the proposed approach we augment the GAN framework with axial-attention layers which leverage the recent advances in transformers for modelling long-range dependencies. We demonstrate the effectiveness of the proposed method by evaluating on two different thermal-visible face datasets. When compared to related state-of-the-art works, our results show significant improvements in both image quality and face verification performance, and are also much more efficient.
Rakhil Immidisetti, Shuowen Hu, Vishal M. Patel
IJCB2
2021 DefakeHop: A Light-Weight High-Performance Deepfake Detector
abstract
A light-weight high-performance Deepfake detection method, called DefakeHop, is proposed in this work. State-of-the-art Deepfake detection methods are built upon deep neural networks. DefakeHop uses the successive subspace learning (SSL) principle to extracts features automatically from various parts of face images. The features are extracted by channel-wise (c/w) Saab transform and further processed by our feature distillation module using spatial dimension re-duction and soft classification for each channel to get a more concise description of the face. Extensive experiments are conducted to demonstrate the effectiveness of the proposed DefakeHop method. With a small model size of 42,845 parameters, DefakeHop achieves state-of-the-art performance with the area under the ROC curve (AUC) of 100%, 94.95%, and 90.56% on UADFV, Celeb-DF v1, and Celeb-DF v2 datasets, respectively. Our codes are available on GitHub1.
Hong-Shuo Chen, Mozhdeh Rouhsedaghat, Hamza Ghani, Shuowen Hu, Suya You, C.-C. Jay Kuo
ICME4
2021 A Large-Scale, Time-Synchronized Visible and Thermal Face Dataset
abstract
Thermal face imagery, which captures the naturally emitted heat from the face, is limited in availability compared to face imagery in the visible spectrum. To help address this scarcity of thermal face imagery for research and algorithm development, we present the DEVCOM Army Research Laboratory Visible-Thermal Face Dataset (ARL-VTF). With over 500,000 images from 395 subjects, the ARL-VTF dataset represents, to the best of our knowledge, the largest collection of paired visible and thermal face images to date. The data was captured using a modern long wave infrared (LWIR) camera mounted alongside a stereo setup of three visible spectrum cameras. Variability in expressions, pose, and eyewear has been systematically recorded. The dataset has been curated with extensive annotations, metadata, and standardized protocols for evaluation. Furthermore, this paper presents extensive benchmark results and analysis on thermal face landmark detection and thermal-to-visible face verification by evaluating state-of-the-art models on the ARL-VTF dataset.
Domenick Poster, Matthew Thielke, Robert Nguyen, Srinivasan Rajaraman, Xing Di, Cedric Nimpa Fondje, Vishal M. Patel, Nathan J. Short, Benjamin S. Riggan, Nasser M. Nasrabadi, Shuowen Hu
WACV11
2021 Low-resolution face recognition in resource-constrained environments
Mozhdeh Rouhsedaghat, Yifan Wang 0018, Shuowen Hu, Suya You, C.-C. Jay Kuo
Pattern Recognit. Lett.3
2020 Cross-Domain Identification for Thermal-to-Visible Face Recognition
abstract
Recent advances in domain adaptation, especially those applied to heterogeneous facial recognition, typically rely upon restrictive Euclidean loss functions (e.g., L2 norm) which perform best when images from two different domains (e.g., visible and thermal) are co-registered and temporally synchronized. This paper proposes a novel domain adaptation framework that combines a new feature mapping sub-network with existing deep feature models, which are based on modified network architectures (e.g., VGG16 or Resnet50). This framework is optimized by introducing new cross-domain identity and domain invariance lossfunctions for thermal-to-visible face recognition, which alleviates the requirement for precisely co-registered and synchronized imagery. We provide extensive analysis of both features and loss functions used, and compare the proposed domain adaptation framework with state-of-the-art feature based domain adaptation models on a difficult dataset containing facial imagery collected at varying ranges, poses, and expressions. Moreover, we analyze the viability of the proposed framework for more challenging tasks, such as non-frontal thermal-to-visible face recognition.
Cedric Nimpa Fondje, Shuowen Hu, Nathan J. Short, Benjamin S. Riggan
IJCB2
2020 Coupled generative adversarial network for heterogeneous face recognition
Seyed Mehdi Iranmanesh, Benjamin S. Riggan, Shuowen Hu, Nasser M. Nasrabadi
Image Vis. Comput.3
2019 Synthesis of High-Quality Visible Faces from Polarimetric Thermal Faces using Generative Adversarial Networks
He Zhang 0004, Benjamin S. Riggan, Shuowen Hu, Nathan J. Short, Vishal M. Patel
Int. J. Comput. Vis.3
2018 Thermal to Visible Synthesis of Face Images Using Multiple Regions
abstract
Synthesis of visible spectrum faces from thermal facial imagery is a promising approach for heterogeneous face recognition; enabling existing face recognition software trained on visible imagery to be leveraged, and allowing human analysts to verify cross-spectrum matches more effectively. We propose a new synthesis method to enhance the discriminative quality of synthesized visible face imagery by leveraging both global (e.g., entire face) and local regions (e.g., eyes, nose, and mouth). Here, each region provides (1) an independent representation for the corresponding area, and (2) additional regularization terms, which impact the overall quality of synthesized images. We analyze the effects of using multiple regions to synthesize a visible face image from a thermal face. We demonstrate that our approach improves cross-spectrum verification rates over recently published synthesis approaches. Moreover, using our synthesized imagery, we report the results on facial landmark detection-commonly used for image registration- which is a critical part of the face recognition process.
Benjamin S. Riggan, Nathan J. Short, Shuowen Hu
WACV3
2017 Heterogeneous Face Recognition: Recent Advances in Infrared-to-Visible Matching
abstract
An emerging topic in face recognition is matching between facial images acquired from different sensing modalities, referred to as heterogeneous face recognition. Heterogeneous face recognition has the potential to provide key capabilities for the commercial sector as well as for law enforcement, intelligence gathering, and the military, especially in challenging unconstrained settings. However, the difficulty in heterogeneous face recognition is compounded by phenomenology differences between modalities, giving rise to significant facial appearance variations due to the modality gap. In this paper, we focus on a subset of heterogeneous face recognition and present a succinct review of recent work on infrared-to-visible face recognition.
Shuowen Hu, Nathan J. Short, Benjamin S. Riggan, Matthew Chasse, M. Saquib Sarfraz
FG1
2017 Generative adversarial network-based synthesis of visible faces from polarimetrie thermal faces
abstract
The large domain discrepancy between faces captured in polarimetric (or conventional) thermal and visible domain makes cross-domain face recognition quite a challenging problem for both human-examiners and computer vision algorithms. Previous approaches utilize a two-step procedure (visible feature estimation and visible image reconstruction) to synthesize the visible image given the corresponding polarimetric thermal image. However, these are regarded as two disjoint steps and hence may hinder the performance of visible face reconstruction. We argue that joint optimization would be a better way to reconstruct more photo-realistic images for both computer vision algorithms and human-examiners to examine. To this end, this paper proposes a Generative Adversarial Network-based Visible Face Synthesis (GAN-VFS) method to synthesize more photo-realistic visible face images from their corresponding polarimetric images. To ensure that the encoded visible-features contain more semantically meaningful information in reconstructing the visible face image, a guidance sub-network is involved into the training procedure. To achieve photo realistic property while preserving discriminative characteristics for the reconstructed outputs, an identity loss combined with the perceptual loss are optimized in the framework. Multiple experiments evaluated on different experimental protocols demonstrate that the proposed method achieves state-of-the-art performance.
He Zhang 0004, Vishal M. Patel, Benjamin S. Riggan, Shuowen Hu
IJCB4
2016 Optimal feature learning and discriminative framework for polarimetric thermal to visible face recognition
abstract
A face recognition system capable of day- and night-time operation is highly desirable for surveillance and reconnaissance. Polarimetric thermal imaging is ideal for such applications, as it acquires emitted radiation from skin tissue. However, polarimetric thermal facial imagery must be matched to visible face images for interoperability with existing biometric databases. This work proposes a novel framework for polarimetric thermal-to-visible face recognition, where polarimetric features are optimally combined to facilitate training of a discriminant classifier. We evaluate its performance on imagery collected under different expressions and at different ranges, and compare with recent deep perceptual mapping, coupled neural network, and partial least squares techniques for cross-spectrum face matching.
Benjamin S. Riggan, Nathan J. Short, Shuowen Hu
WACV3