VLDB 2026 Research / reviewers in the wild / expert
Benjamin S. Riggan
dblp:152/5127
· DBLP profile ↗
20ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-2293-0439ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 5 since 2021Security and privacy · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Using Cross-Domain Detection Loss to Infer Multi-Scale Information for Improved Tiny Head TrackingabstractHead detection and tracking are essential for downstream tasks, but current methods often require large computational budgets, which increase latencies and ties up resources (e.g., processors, memory, and bandwidth). To address this, we propose a framework to enhance tiny head detection and tracking by optimizing the balance between performance and efficiency. Our framework integrates (1|) a cross-domain detection loss, (2) a multi-scale module, and (3) a small receptive field detection mechanism. These innovations enhance detection by bridging the gap between large and small detectors, capturing high-frequency details at multiple scales during training, and using filters with small receptive fields to detect tiny heads. Evaluations on the CroHD and CrowdHuman datasets show improved Multiple Object Tracking Accuracy (MOTA) and mean Average Precision (mAP), demonstrating the effectiveness of our approach in crowded scenes. Alexander Mattingly, Eungjoo Lee 0001, Benjamin S. Riggan |
FG | 4 |
| 2025 | 2D-3D Attention and Entropy for Pose Robust 2D Facial RecognitionabstractDespite recent advances in facial recognition, there remains a fundamental issue concerning degradations in performance due to substantial perspective (pose) differences between enrollment and query (probe) imagery. Therefore, we propose a novel domain adaptive framework to facilitate improved performances across large discrepancies in pose by enabling imagebased (2D) representations to infer properties of inherently pose invariant point cloud (3D) representations. Specifically, our proposed framework achieves better pose invariance by using (1) a shared (joint) attention mapping to emphasize common patterns that are most correlated between 2D facial images and 3D facial data and (2) a joint entropy regularizing loss to promote better consistency—enhancing correlations among the intersecting 2D and 3D representations—by leveraging both attention maps. This framework is evaluated on FaceScape and ARL-VTF datasets, where it outperforms competitive methods by achieving profile ($90^{\circ}+$) TAR @ 1% FAR improvements of at least $\mathbf{7. 1 \%}$ and $\mathbf{1. 5 7 \%}$, respectively. John Brennan Peace, Shuowen Hu, Benjamin S. Riggan |
FG | 3 |
| 2024 | HashReID: Dynamic Network with Binary Codes for Efficient Person Re-identificationabstractBiometric applications, such as person re-identification (ReID), are often deployed on energy constrained devices. While recent ReID methods prioritize high retrieval performance, they often come with large computational costs and high search time, rendering them less practical in real-world settings. In this work, we propose an input-adaptive network with multiple exit blocks, that can terminate computation early if the retrieval is straightforward or noisy, saving a lot of computation. To assess the complexity of the input, we introduce a temporal-based classifier driven by a new training strategy. Furthermore, we adopt a binary hash code generation approach instead of relying on continuous-valued features, which significantly improves the search process by a factor of 20. To ensure similarity preservation, we utilize a new ranking regularizer that bridges the gap between continuous and binary features. Extensive analysis of our proposed method is conducted on three datasets: Market1501, MSMT17 (Multi-Scene Multi-Time), and the BGC1 (BRIAR Government Collection). Using our approach, more than 70% of the samples with compact hash codes exit early on the Market1501 dataset, saving 80% of the networks computational cost and improving over other hash-based methods by 60%. These results demonstrate a significant improvement over dynamic networks and showcase comparable accuracy performance to conventional ReID methods. Kshitij Nikhal, Yujunrong Ma, Shuvra S. Bhattacharyya, Benjamin S. Riggan |
WACV | 4 |
| 2023 | HBRC-500: A Long Range Recognition Benchmark Dataset using Face and Whole-body ImageryabstractWhile biometric-face and whole-body-recognition technology have recently advanced and matured, there are increasing interests in enhanced long-range recognition capabilities. However, long-range recognition requires the use of large, specialized datasets to support research and development for next generation systems. Moreover, existing datasets are further limited by the types of modalities (face or whole-body), number of subjects, maximum standoff distance, clothing variability, or availability restrictions. For long-range recognition, low-quality probe (query) images, which are often acquired from extended standoff distances or aerial platforms, are matched against higher quality gallery images and frequently results in poor identification performance. To address the growing needs for relevant data sources, a large-scale biometric dataset was collected and curated for long-range biometric recognition. This dataset is comprised of more than 1.2 million outdoor and 250,000 indoor (face and whole-body) images from more than 250 subjects that were acquired using various high-end cameras, including Canon and Nikon DSLR cameras, surveillance cameras, specialized long-range face cameras, and UAV platforms. The primary goal of this dataset is to support the development of algorithms for face and whole-body recognition at extended standoff distances. The availability of such a dataset is crucial in advancing technology for recognition under challenging conditions such as atmospheric turbulence. Cedric Nimpa Fondje, Kshitij Nikhal, John Brennan Peace, Ryan Karl, Mun Wai Lee, Phillip Berkowitz, Katrina Gramzinski, Bridget Kennedy, Nkirukaegbunam Uzuegbunam, Victoria Ou, Tyler Barret, Oliver Arend, Wei Ming, Svetlana Semenova, Benjamin S. Riggan |
IJCB | 15 |
| 2023 | Weakly Supervised Face and Whole Body Recognition in Turbulent EnvironmentsabstractFace and person recognition have recently achieved remarkable success under challenging scenarios, such as off-pose and cross-spectrum matching. However, long-range recognition systems are often hindered by atmospheric turbulence, leading to spatially and temporally varying distortions in the image. Current solutions rely on generative models to reconstruct a turbulent-free image, but often preserve photo-realism instead of discriminative features that are essential for recognition. This can be attributed to the lack of large-scale datasets of turbulent and pristine paired images, necessary for optimal reconstruction. To address this issue, we propose a new weakly supervised framework that employs a parameter-efficient self-attention module to generate domain agnostic representations, aligning turbulent and pristine images into a common subspace. Additionally, we introduce a new tilt map estimator that predicts geometric distortions observed in turbulent images. This estimate is used to re-rank gallery matches, resulting in up to 13.86% improvement in rank-1 accuracy. Our method does not require synthesizing turbulent-free images or ground-truth paired images, and requires significantly fewer annotated samples, enabling more practical and rapid utility of increasingly large datasets. We analyze our framework using two datasets—Long-Range Face Identification Dataset (LRFID) and BRIAR Government Collection 1 (BGC1)— achieving enhanced discriminability under varying turbulence and standoff distance. Kshitij Nikhal, Benjamin S. Riggan |
IJCB | 2 |
| 2021 | Understanding Cross Domain Presentation Attack Detection for Visible Face RecognitionabstractFace signatures, including size, shape, texture, skin tone, eye color, appearance, and scars/marks, are widely used as discriminative, biometric information for access control. Despite recent advancements in facial recognition systems, presentation attacks on facial recognition systems have become increasingly sophisticated. The ability to detect presentation attacks or spoofing attempts is a pressing concern for the integrity, security, and trust of facial recognition systems. Multispectral imaging has been previously introduced as a way to improve presentation attack detection by utilizing sensors that are sensitive to different regions of the electromagnetic spectrum (e.g., visible, near infrared, long-wave infrared). Although multi-spectral presentation attack detection systems may be discriminative, the need for additional sensors and computational resources substantially increases complexity and costs. Instead, we propose a method that exploits information from infrared imagery during training to increase the discriminability of visible-based presentation attack detection systems. We introduce (1) a new cross-domain presentation attack detection framework that increases the separability of bonafide and presentation attacks using only visible spectrum imagery, (2) an inverse domain regularization technique for added training stability when optimizing our cross-domain presentation attack detection framework, and (3) a dense domain adaptation subnetwork to transform representations between visible and non-visible domains. Jennifer Hamblin, Kshitij Nikhal, Benjamin S. Riggan |
FG | 3 |
| 2021 | Unsupervised Attention Based Instance Discriminative Learning for Person Re-IdentificationabstractRecent advances in person re-identification have demonstrated enhanced discriminability, especially with supervised learning or transfer learning. However, since the data requirements-including the degree of data curations-are becoming increasingly complex and laborious, there is a critical need for unsupervised methods that are robust to large intra-class variations, such as changes in perspective, illumination, articulated motion, resolution, etc. Therefore, we propose an unsupervised framework for person re-identification which is trained in an end-to-end manner without any pre-training. Our proposed framework leverages a new attention mechanism that combines group convolutions to (1) enhance spatial attention at multiple scales and (2) reduce the number of trainable parameters by 59.6%. Additionally, our framework jointly optimizes the network with agglomerative clustering and instance learning to tackle hard samples. We perform extensive analysis using the Market1501 and DukeMTMC-reID datasets to demonstrate that our method consistently outperforms the state-of-the-art methods (with and without pre-trained weights). Kshitij Nikhal, Benjamin S. Riggan |
WACV | 2 |
| 2021 | A Large-Scale, Time-Synchronized Visible and Thermal Face DatasetabstractThermal face imagery, which captures the naturally emitted heat from the face, is limited in availability compared to face imagery in the visible spectrum. To help address this scarcity of thermal face imagery for research and algorithm development, we present the DEVCOM Army Research Laboratory Visible-Thermal Face Dataset (ARL-VTF). With over 500,000 images from 395 subjects, the ARL-VTF dataset represents, to the best of our knowledge, the largest collection of paired visible and thermal face images to date. The data was captured using a modern long wave infrared (LWIR) camera mounted alongside a stereo setup of three visible spectrum cameras. Variability in expressions, pose, and eyewear has been systematically recorded. The dataset has been curated with extensive annotations, metadata, and standardized protocols for evaluation. Furthermore, this paper presents extensive benchmark results and analysis on thermal face landmark detection and thermal-to-visible face verification by evaluating state-of-the-art models on the ARL-VTF dataset. Domenick Poster, Matthew Thielke, Robert Nguyen, Srinivasan Rajaraman, Xing Di, Cedric Nimpa Fondje, Vishal M. Patel, Nathan J. Short, Benjamin S. Riggan, Nasser M. Nasrabadi, Shuowen Hu |
WACV | 9 |
| 2020 | Commuting Conditional GANS for Multi-Modal FusionabstractThis paper presents a data driven approach to multi-modal fusion where a hidden latent sub-space between the different modalities is learned. The hidden space is estimated via a bank of Conditional GANs which also commute with each other, leading to an output that lies in a common subspace. Experimental results show improved detection performance compared to existing fusion techniques in ideal as well as noisy sensor condition. Siddharth Roheda, Hamid Krim, Benjamin S. Riggan |
ICASSP | 3 |
| 2020 | Cross-Domain Identification for Thermal-to-Visible Face RecognitionabstractRecent advances in domain adaptation, especially those applied to heterogeneous facial recognition, typically rely upon restrictive Euclidean loss functions (e.g., L2 norm) which perform best when images from two different domains (e.g., visible and thermal) are co-registered and temporally synchronized. This paper proposes a novel domain adaptation framework that combines a new feature mapping sub-network with existing deep feature models, which are based on modified network architectures (e.g., VGG16 or Resnet50). This framework is optimized by introducing new cross-domain identity and domain invariance lossfunctions for thermal-to-visible face recognition, which alleviates the requirement for precisely co-registered and synchronized imagery. We provide extensive analysis of both features and loss functions used, and compare the proposed domain adaptation framework with state-of-the-art feature based domain adaptation models on a difficult dataset containing facial imagery collected at varying ranges, poses, and expressions. Moreover, we analyze the viability of the proposed framework for more challenging tasks, such as non-frontal thermal-to-visible face recognition. Cedric Nimpa Fondje, Shuowen Hu, Nathan J. Short, Benjamin S. Riggan |
IJCB | 4 |
| 2020 | Coupled generative adversarial network for heterogeneous face recognition
Seyed Mehdi Iranmanesh, Benjamin S. Riggan, Shuowen Hu, Nasser M. Nasrabadi |
Image Vis. Comput. | 2 |
| 2019 | Synthesis of High-Quality Visible Faces from Polarimetric Thermal Faces using Generative Adversarial Networks
He Zhang 0004, Benjamin S. Riggan, Shuowen Hu, Nathan J. Short, Vishal M. Patel |
Int. J. Comput. Vis. | 2 |
| 2018 | A Joint Target Localization and Classification Framework for Sensor NetworksabstractIn this paper, we propose a joint framework for target localization and classification using a single generalized model for non-imaging based multi-modal sensor data. For target localization, we exploit both sensor data and estimated dynamics within a local neighborhood. We validate the capabilities of our framework by using a multi-modal dataset, which includes ground truth GPS information (e.g., time and position) and data from co-located seismic and acoustic sensors. Experimental results show that our framework achieves better classification accuracy compared to recent fusion algorithms using temporal accumulation and achieves more accurate target localizations than multilateration. Kyunghun Lee, Benjamin S. Riggan, Shuvra S. Bhattacharyya |
ICASSP | 2 |
| 2018 | Cross-Modality Distillation: A Case for Conditional Generative Adversarial NetworksabstractIn this paper, we propose to use a Conditional Generative Adversarial Network (CGAN) for distilling (i.e. transferring) knowledge from sensor data and enhancing low-resolution target detection. In unconstrained surveillance settings, sensor measurements are often noisy, degraded, corrupted, and even missing/absent, thereby presenting a significant problem for multi-modal fusion. We therefore specifically tackle the problem of a missing modality in our attempt to propose an algorithm based on CGANs to generate representative information from the missing modalities when given some other available modalities. Despite modality gaps, we show that one can distill knowledge from one set of modalities to another. Moreover, we demonstrate that it achieves better performance than traditional approaches and recent teacher-student models. Siddharth Roheda, Benjamin S. Riggan, Hamid Krim, Liyi Dai |
ICASSP | 2 |
| 2018 | Thermal to Visible Synthesis of Face Images Using Multiple RegionsabstractSynthesis of visible spectrum faces from thermal facial imagery is a promising approach for heterogeneous face recognition; enabling existing face recognition software trained on visible imagery to be leveraged, and allowing human analysts to verify cross-spectrum matches more effectively. We propose a new synthesis method to enhance the discriminative quality of synthesized visible face imagery by leveraging both global (e.g., entire face) and local regions (e.g., eyes, nose, and mouth). Here, each region provides (1) an independent representation for the corresponding area, and (2) additional regularization terms, which impact the overall quality of synthesized images. We analyze the effects of using multiple regions to synthesize a visible face image from a thermal face. We demonstrate that our approach improves cross-spectrum verification rates over recently published synthesis approaches. Moreover, using our synthesized imagery, we report the results on facial landmark detection-commonly used for image registration- which is a critical part of the face recognition process. Benjamin S. Riggan, Nathan J. Short, Shuowen Hu |
WACV | 1 |
| 2018 | An Order Preserving Bilinear Model for Person Detection in Multi-Modal DataabstractWe propose a new order preserving bilinear framework that exploits low-resolution video for person detection in a multi-modal setting using deep neural networks. In this setting cameras are strategically placed such that less robust sensors, e.g. geophones that monitor seismic activity, are located within the field of views (FOVs) of cameras. The primary challenge is being able to leverage sufficient information from videos where there are less than 40 pixels on targets, while also taking advantage of less discriminative information from other modalities, e.g. seismic. Unlike state-of-the-art methods, our bilinear framework retains spatio-temporal order when computing the vector outer products between pairs of features. Despite the high dimensionality of these outer products, we demonstrate that our order preserving bilinear framework yields better performance than recent orderless bilinear models and alternative fusion methods. Code is available at https://github.com/oulutan/OP-Bilinear-Model. Oytun Ulutan, Benjamin S. Riggan, Nasser M. Nasrabadi, B. S. Manjunath |
WACV | 2 |
| 2017 | Heterogeneous Face Recognition: Recent Advances in Infrared-to-Visible MatchingabstractAn emerging topic in face recognition is matching between facial images acquired from different sensing modalities, referred to as heterogeneous face recognition. Heterogeneous face recognition has the potential to provide key capabilities for the commercial sector as well as for law enforcement, intelligence gathering, and the military, especially in challenging unconstrained settings. However, the difficulty in heterogeneous face recognition is compounded by phenomenology differences between modalities, giving rise to significant facial appearance variations due to the modality gap. In this paper, we focus on a subset of heterogeneous face recognition and present a succinct review of recent work on infrared-to-visible face recognition. Shuowen Hu, Nathan J. Short, Benjamin S. Riggan, Matthew Chasse, M. Saquib Sarfraz |
FG | 3 |
| 2017 | An accumulative fusion architecture for discriminating people and vehicles using acoustic and seismic signalsabstractIn this paper, we develop new multiclass classification algorithms for detecting people and vehicles by fusing data from a multimodal, unattended ground sensor node. The specific types of sensors that we apply in this work are acoustic and seismic sensors. We investigate two alternative approaches to multiclass classification in this context - the first is based on applying Dempster-Shafer Theory to perform score-level fusion, and the second involves the accumulation of local similarity evidences derived from a feature-level fusion model that combines both modalities. We experiment with the proposed algorithms using different datasets obtained from acoustic and seismic sensors in various outdoor environments, and evaluate the performance of the two algorithms in terms of receiver operating characteristic and classification accuracy. Our results demonstrate overall superiority of the proposed new feature-level fusion approach for multiclass discrimination among people, vehicles and noise. Kyunghun Lee, Benjamin S. Riggan, Shuvra S. Bhattacharyya |
ICASSP | 2 |
| 2017 | Generative adversarial network-based synthesis of visible faces from polarimetrie thermal facesabstractThe large domain discrepancy between faces captured in polarimetric (or conventional) thermal and visible domain makes cross-domain face recognition quite a challenging problem for both human-examiners and computer vision algorithms. Previous approaches utilize a two-step procedure (visible feature estimation and visible image reconstruction) to synthesize the visible image given the corresponding polarimetric thermal image. However, these are regarded as two disjoint steps and hence may hinder the performance of visible face reconstruction. We argue that joint optimization would be a better way to reconstruct more photo-realistic images for both computer vision algorithms and human-examiners to examine. To this end, this paper proposes a Generative Adversarial Network-based Visible Face Synthesis (GAN-VFS) method to synthesize more photo-realistic visible face images from their corresponding polarimetric images. To ensure that the encoded visible-features contain more semantically meaningful information in reconstructing the visible face image, a guidance sub-network is involved into the training procedure. To achieve photo realistic property while preserving discriminative characteristics for the reconstructed outputs, an identity loss combined with the perceptual loss are optimized in the framework. Multiple experiments evaluated on different experimental protocols demonstrate that the proposed method achieves state-of-the-art performance. He Zhang 0004, Vishal M. Patel, Benjamin S. Riggan, Shuowen Hu |
IJCB | 3 |
| 2016 | Optimal feature learning and discriminative framework for polarimetric thermal to visible face recognitionabstractA face recognition system capable of day- and night-time operation is highly desirable for surveillance and reconnaissance. Polarimetric thermal imaging is ideal for such applications, as it acquires emitted radiation from skin tissue. However, polarimetric thermal facial imagery must be matched to visible face images for interoperability with existing biometric databases. This work proposes a novel framework for polarimetric thermal-to-visible face recognition, where polarimetric features are optimally combined to facilitate training of a discriminant classifier. We evaluate its performance on imagery collected under different expressions and at different ranges, and compare with recent deep perceptual mapping, coupled neural network, and partial least squares techniques for cross-spectrum face matching. Benjamin S. Riggan, Nathan J. Short, Shuowen Hu |
WACV | 1 |