EDBT 2026 Demo / reviewers in the wild / expert
Heeseung Choi
dblp:120/8562
· DBLP profile ↗
27ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-3223-1885ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 9 since 2021Security and privacy · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VAST-ReID: A Low-Light Benchmark Dataset for Person Re-Identification with Visual and Attribute-Rich Semantic TrackingabstractPerson Re-Identification (ReID) task is important for designing intelligent surveillance systems. ReID can be highly challenging in low-light and low resolution scenarios. Existing ReID datasets predominantly feature cropped pedestrian images captured in well-lit environments, often lacking semantic richness, frame-level temporal continuity, and robustness to adverse conditions. To address these limitations, we introduce VAST-ReID, a new benchmark dataset specifically designed for the low-light person ReID task in real-world surveillance contexts. VAST-ReID consists of 1,441 surveillance videos collected at 24 different locations, capturing 256 distinct pedestrians of various age groups. The dataset emphasizes naturally low-light and visually degraded scenarios. Each identity is annotated with dense bounding boxes and enriched with auxiliary semantic labels, including pedestrian attributes and LLM-generated descriptions. While these annotations are not used during supervised training, they provide valuable semantic context for advancing research in language-guided retrieval and attribute-aware modeling. Additionally, we release identity-aligned image crops under the BoxTrack-ReID subset, which has over 18.7K frames sampled at 1fps from the raw videos, with standard training, gallery, and query splits compatible with the Market-1501 evaluation protocol, enabling straightforward benchmarking. The dataset has been benchmarked against SOTA methods, and experiments reveal that there is huge scope for improvement in ReID research. VAST-ReID is available at: https://github.com/Byte0wl/VAST-ReID Hammad Khan, Rakesh Kumar Giri, Thakare Kamalakar Vijay, Heeseung Choi, Hyungjoo Jung, Debi Prosad Dogra, Ig-Jae Kim |
WACV | 4 |
| 2026 | IMPACT: Interpretable Most Important Person Analysis and Classification using Transformer-based ModelsabstractIdentifying the Most Important Person (MIP) in complex social and sports events remains a challenging problem due to the dynamic nature of group interactions, subtle visual cues, and context-dependent semantics. Traditional methods often struggle to accurately capture the interplay between individuals and the overarching activity, especially in unstructured real-world environments. In addition, the lack of strong supervision and the need for a deeper contextual understanding further complicate the task. In this work, we propose IMPACT, a novel multi-modal framework that leverages recent advances in vision language models to bridge the gap between visual perception and semantic reasoning. Our approach integrates structured scene understanding, natural language generation, and cross-modal learning to jointly model activity recognition and MIP localization. The method integrates language, vision, and spatial reasoning to improve scene interpretability as well as accuracy in group activity recognition tasks. By incorporating language-based representations, the proposed method enables interpretable and robust performance in sports-centric group activity scenarios. Comprehensive experiments on C-Sports and NCAA datasets demonstrate that the framework significantly enhances the localization of key individuals as well as the accuracy of activity prediction, laying the groundwork for a holistic scene understanding in human-centric video and image analysis. Our proposed method achieves an accuracy of 81.6% when compared with human annotator markings and an increase in mAP scores by ∼ 5% for MIP identification. Akshat Rampuria, Kamakshya Prasad Nayak, Thakare Kamalakar Vijay, Tushar Joshi, Aditya Dhananjay Singh, Haesol Park, Heeseung Choi, Hyungjoo Jung, Debi Prosad Dogra, Ig-Jae Kim |
WACV | 7 |
| 2025 | Effective SAM Combination for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation aims to assign pixel-level labels to images across an unlimited range of classes. Traditional methods address this by sequentially connecting a powerful mask proposal generator, such as the Segment Anything Model (SAM), with a pre-trained vision-language model like CLIP. But these two-stage approaches often suffer from high computational costs, memory inefficiencies. In this paper, we propose ESC-Net, a novel one-stage open-vocabulary segmentation model that leverages the SAM decoder blocks for class-agnostic segmentation within an efficient inference framework. By embedding pseudo prompts generated from image-text correlations into SAM’s promptable segmentation framework, ESC-Net achieves refined spatial aggregation for accurate mask predictions. Additionally, a Vision-Language Fusion (VLF) module enhances the final mask prediction through image and text guidance. ESC-Net and PASCAL-Context, outperforming prior methods in both efficiency and accuracy. Comprehensive ablation studies further demonstrate its robustness across challenging conditions. Minhyeok Lee, Suhwan Cho, Sunghun Yang, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee |
CVPR | 5 |
| 2025 | CLIPping Imbalances: A Novel Evaluation Baseline and PEARL Dataset for Pedestrian Attribute RecognitionabstractPedestrian Attribute Recognition (PAR) serves as a fun-damental task in computer vision and is crucial for upgradign security systems. It helps in precisely identifying and characterizing various attributes of pedestrians. However, current PAR datasets have certain issues in representing a wide range of attributes correctly, which makes the ex-isting PAR methods less effective in real-world scenarios. Addressing this limitation, this paper introduces PEARL, a comprehensive dataset comprising of diverse pedestrian images annotated with 146 attributes. These samples have been sourced from surveillance videos across twelve coun-tries. This paper also formulates an image-based PAR using language-image fusion strategy and utilizes CLIP as a new evaluation baseline. Specifically, we leverage textual infor-mation by transforming sets of attributes into meaningful sentences. Addressing the inherent data imbalance in PAR, we provide three types of prompt settings to optimize the training of the CLIP model. Our evaluation encompasses a thorough assessment of the proposed baseline model across various datasets, including PEARL dataset as well as estab-lished PAR benchmarks such as PA100K, RAP, and PETA. Thakare Kamalakar Vijay, Lalit Lohani, Kamakshya Prasad Nayak, Debi Prosad Dogra, Heeseung Choi, Hyungjoo Jung, Ig-Jae Kim |
WACV | 5 |
| 2024 | Dual Prototype Attention for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; and 2) temporal fusion between different frames. This paper proposes two novel prototype-based attention mechanisms, inter-modality attention (IMA) and inter-frame attention (IFA), to incorporate these techniques via dense propagation across different modalities and frames. IMA densely in-tegrates context information from different modalities based on a mutual refinement. IFA injects global context of a video to the query frame, enabling a full utilization of useful prop-erties from multiple frames. Experimental results on public benchmark datasets demonstrate that our proposed approach outperforms all existing methods by a substantial margin. The proposed two components are also thoroughly validated via ablative study. Code and models are available at https://github.com/Hydragon516/DPA. Suhwan Cho, Minhyeok Lee, Seunghoon Lee 0008, Dogyoon Lee, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee |
CVPR | 5 |
| 2024 | Pedestrian Attribute Recognition Using Hierarchical Transformers
Lalit Lohani, Thakare Kamalakar Vijay, Kamakshya Prasad Nayak, Debi Prosad Dogra, Heeseung Choi, Hyungjoo Jung, Ig-Jae Kim |
ICPR (16) | 5 |
| 2024 | Let's Observe Them Over Time: An Improved Pedestrian Attribute Recognition ApproachabstractDespite poor image quality, occlusions, and small training datasets, recent pedestrian attribute recognition (PAR) methods have achieved considerable performance. However, leveraging only spatial information of different attributes limits their reliability and generalizability. This paper introduces a multi-perspective approach to reduce over-dependence on spatial clues of a single perspective and exploits other aspects available in multiple perspectives. In order to tackle image quality and occlusions, we exploit different spatial clues present across images and handpick the best attribute-specific features to classify. Precisely, we extract the class-activation energy of each attribute and correlate it with the corresponding energy present across other images using the proposed Self-Attentive Cross Relation Module. In the next stage, we fuse this correlation information with similar clues accumulated from the other images. Lastly, we train a classification neural network using combined correlation information with two different losses. We have validated our method on four widely used PAR datasets, namely Market1501, PETA, PA-100k, and Duke. Our method achieves superior performance over most existing methods, demonstrating the effectiveness of a multi-perspective approach in PAR. Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim |
WACV | 3 |
| 2023 | DyAnNet: A Scene Dynamicity Guided Self-Trained Video Anomaly Detection NetworkabstractUnsupervised approaches for video anomaly detection may not perform as good as supervised approaches. However, learning unknown types of anomalies using an unsupervised approach is more practical than a supervised approach as annotation is an extra burden. In this paper, we use isolation tree-based unsupervised clustering to partition the deep feature space of the video segments. The RGB-stream generates a pseudo anomaly score and the flow stream generates a pseudo dynamicity score of a video segment. These scores are then fused using a majority voting scheme to generate preliminary bags of positive and negative segments. However, these bags may not be accurate as the scores are generated only using the current segment which does not represent the global behavior of a typical anomalous event. We then use a refinement strategy based on a cross-branch feed-forward network designed using a popular I3D network to refine both scores. The bags are then refined through a segment re-mapping strategy. The intuition of adding the dynamicity score of a segment with the anomaly score is to enhance the quality of the evidence. The method has been evaluated on three popular video anomaly datasets, i.e., UCF-Crime, CCTV-Fights, and UBI-Fights. Experimental results reveal that the proposed framework achieves competitive accuracy as compared to the state-of-the-art video anomaly detection methods. Thakare Kamalakar Vijay, Yash Raghuwanshi, Debi Prosad Dogra, Heeseung Choi, Ig-Jae Kim |
WACV | 4 |
| 2023 | RareAnom: A Benchmark Video Dataset for Rare Type AnomaliesabstractExisting video anomaly detection methods and datasets suffer from restricted anomaly categories containing single-source (CCTV) videos recorded in controlled environment, inadequate annotations, and lack of adequate supervision. To mitigate these problems, we introduce a new dataset ( RareAnom ) containing 17 rare types of real-world anomalies (2200 videos) recorded using multiple sources (e.g., CCTV , handheld cameras, dash-cams, and mobile phones) with rich temporal annotations. A new fully unsupervised anomaly detection and classification method has been proposed. It has three stages: training of a 3D Convolution Autoencoder using pseudo-labelled video segments, anomaly detection using latent features, and classification. Unlike the existing datasets, we have benchmarked RareAnom using three levels of supervision: fully, weakly, and unsupervised. It has been compared with UCF-Crime and XD-Violence datasets. The proposed anomaly detection and classification method beats the latest unsupervised methods by 4.49%, 8.66%, and 6.77% on RareAnom, UCF-Crime, and XD-violence datasets, respectively. Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim |
Pattern Recognit. | 3 |
| 2023 | Detection of Road Accidents Using Synthetically Generated Multi-Perspective Accident VideosabstractRoad accidents are often caused by short abnormal events, including traffic violations, abrupt change in vehicular motion, driver fatigue, etc. Observing an accident event from the right camera perspective plays a crucial role while detecting accidents. However, it may not be possible to capture such abnormal events from a limited camera perspective. We present a deep learning framework to analyze the accident events recorded from multiple perspectives. First, we estimate feature similarity in videos recorded from multiple perspectives. We then divided the video samples into high and low feature similarity groups. Next, we extract spatio-temporal features from each group using two-branch DCNNs and fuse them using a rank-based weighted average pooling strategy followed by classification. We present a new road accident video dataset (MP-RAD), where each accident event is synthetically generated and captured from five independent camera perspectives using a computer gaming platform. Most of the existing road accident datasets use egocentric views or they are captured in fixed camera setups. However, our dataset is large and multi-perspective that can be used to validate ITS-related tasks such as accident detection, accident localization, traffic monitoring, etc. The dataset contains 400 accident events with a total of 2000 videos. We provide temporal annotations of all videos. The proposed framework and the dataset have been cross-validated with latest accident detection baselines trained on real-world road accident videos and vice-versa. The sub-optimal detection accuracy obtained using the baselines indicates that the proposed framework and the dataset can be useful for ITS related research. Code and dataset is available at: https://github.com/draxler1/MP-RAD-Dataset-ITS- Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Person re-identification in indoor videos by information fusion using Graph Convolutional Networks
Komal Soni, Debi Prosad Dogra, Arif Ahmed 0002, Samarjit Kar, Heeseung Choi, Ig-Jae Kim |
Expert Syst. Appl. | 5 |
| 2022 | A multi-stream deep neural network with late fuzzy fusion for real-world anomaly detection
Thakare Kamalakar Vijay, Nitin Sharma 0004, Debi Prosad Dogra, Heeseung Choi, Ig-Jae Kim |
Expert Syst. Appl. | 4 |
| 2022 | Object Interaction-Based Localization and Description of Road Accident Events Using Deep LearningabstractDetection and localization of road accidents in real-time is an integral part of the Intelligent Transportation System (ITS). Even though the existing road accident detection methods show promising results, the process suffers from some drawbacks. For example, existing methods require a large number of sample videos for feature learning. Moreover, features such as temporal gradients or flow fields are time-consuming. To address these issues, we introduce a new method that uses objects and their positions to detect accidents in real-time. Apart from localization of the accident events in videos, we perform a high-level post processing to describe the severity and context of an accident. Firstly, we divide an input video into pre-accident, accident and post-accident stages to extract object interactions. These interaction proposals are then filtered using a refinement algorithm. We then adopt an iterative training procedure to classify normal and accident interactions. We also highlight the damaged zone using heat maps. Finally, we generate high-level textual descriptions to quantify the context and severity of an accident. We have trained the proposed model using offline setups. However, it can be deployed online to detect road accident events in real-time by taking the video inputs directly from the CCTV camera. Moreover, with a minimal supervision, the model can be retrained for online surveillance. Extensive experiments carried out on UCF Crime and CADP datasets reveal that the proposed framework achieves state-of-the-art performance when compared with the recently proposed accident event detection methods in terms of AUC (UCF Crime: 69.70% and CADP: 72.59%) and FAR (UCF Crime: 0.8 and CADP: 2.2). The high-level description of the accident is an added advantage that will certainly help the traffic police to react in a timely manner. Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Scene Adaptive Online Surveillance Video Synopsis via Dynamic Tube Rearrangement Using OctreeabstractVisual surveillance produces a significant amount of raw video data that can be time consuming to browse and analyze. In this work, we present a video synopsis methodology called "scene adaptive online video synopsis via dynamic tube rearrangement using octree (SSOcT)" that can effectively condense input surveillance videos. Our method entailed summarizing the input video by analyzing scene characteristics and determining an effective spatio-temporal 3D structure for video synopsis. For this purpose, we first analyzed the attributes of each extracted tube with respect to scene geometry and complexity. Then, we adaptively grouped the tubes using an online grouping algorithm that exploits these scene characteristics. Finally, the tube groups were dynamically rearranged using the proposed octree-based algorithm that efficiently inserted and refined tubes containing high spatio-temporal movements in real time. Extensive video synopsis experimental results are provided, demonstrating the effectiveness and efficiency of our method in summarizing real-world surveillance videos with diverse scene characteristics. Yoonsik Yang, Haksub Kim, Heeseung Choi, Seungho Chae, Ig-Jae Kim |
IEEE Trans. Image Process. | 3 |
| 2020 | Person Re-identification in Videos by Analyzing Spatio-temporal TubesabstractAbstract Typical person re-identification frameworks search for k best matches in a gallery of images that are often collected in varying conditions. The gallery usually contains image sequences for video re-identification applications. However, such a process is time consuming as video re-identification involves carrying out the matching process multiple times. In this paper, we propose a new method that extracts spatio-temporal frame sequences or tubes of moving persons and performs the re-identification in quick time. Initially, we apply a binary classifier to remove noisy images from the input query tube. In the next step, we use a key-pose detection-based query minimization technique. Finally, a hierarchical re-identification framework is proposed and used to rank the output tubes. Experiments with publicly available video re-identification datasets reveal that our framework is better than existing methods. It ranks the tubes with an average increase in the CMC accuracy of 6-8% across multiple datasets. Also, our method significantly reduces the number of false positives. A new video re-identification dataset, named Tube-based Re-identification Video Dataset (TRiViD), has been prepared with an aim to help the re-identification research community. Arif Ahmed 0002, Debi Prosad Dogra, Heeseung Choi, Seungho Chae, Ig-Jae Kim |
Multim. Tools Appl. | 3 |
| 2020 | Memetic algorithm for multivariate time-series segmentation
Hyunki Lim, Heeseung Choi, Yeji Choi, Ig-Jae Kim |
Pattern Recognit. Lett. | 2 |
| 2020 | Query-Based Video Synopsis for Intelligent Traffic Monitoring ApplicationsabstractSynopsis of a long-duration video has many applications in intelligent transportation systems. It can help to monitor traffic with lesser manpower. However, generating meaningful synopsis of a long-duration video recording can be challenging. Often summarized outputs include redundant contents or activities that may not be helpful to the observer. Moving object trajectories are possible sources of information that can be used to generate the synopsis of long-duration videos. The synopsis generation faces challenges due to object tracking, grouping of the trajectories with respect to activity type, object category, and contextual information, and generating smooth synopsis according to a query. In this paper, we propose a method to generate meaningful and smooth synopsis of long-duration videos according to the users' query. We have tracked moving objects and adopted deep learning to classify the objects into known categories (e.g., car, bike, and pedestrians). We then identify regions in the surveillance scene with the help of unsupervised clustering. Each tube (spatiotemporal object trajectory) is represented by the source and the destination. In the final stage, we take a query from the user and generate the synopsis video by smoothly blending the appropriate tubes over the background frame through energy minimization. The proposed method has been evaluated on two publicly available datasets and our own surveillance datasets. We have compared the method with popular state-of-the-art techniques. The experiments reveal that the proposed method is superior to the existing techniques and it produces visually seamless video synopsis. Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Renuka Patnaik, Seung-Cheol Lee, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2019 | An incremental learning method for spoof fingerprint detection
Jun Beom Kho, Wonjune Lee, Heeseung Choi, Jaihie Kim |
Expert Syst. Appl. | 3 |
| 2017 | Age face simulation using aging functions on global and local features with residual images
Sung Eun Choi, Jaeik Jo, Sanghak Lee, Heeseung Choi, Ig-Jae Kim, Jaihie Kim |
Expert Syst. Appl. | 4 |
| 2017 | Partial fingerprint matching using minutiae and ridge shape features for small fingerprint scanners
Wonjune Lee, Sungchul Cho, Heeseung Choi, Jaihie Kim |
Expert Syst. Appl. | 3 |
| 2015 | Single-view-based 3D facial reconstruction method robust against pose variations
Jaeik Jo, Heeseung Choi, Ig-Jae Kim, Jaihie Kim |
Pattern Recognit. | 2 |
| 2012 | Evidential Value of Automated Latent Fingerprint Comparison: An Empirical ApproachabstractLatent prints are routinely recovered from crime scenes and are compared with available databases of known fingerprints for identifying criminals. However, current procedures to compare latent prints to large databases of exemplar (rolled or plain) prints are prone to errors. This suggests caution in making conclusions about a suspect's identity based on a latent fingerprint comparison. A number of attempts have been made to statistically model the utility of a fingerprint comparison in making a correct accept/reject decision or its evidential value. These approaches, however, either make unrealistic assumptions about the model or they lack simple interpretation. We argue that the posterior probability of two fingerprints belonging to different fingers given their match score, referred to as the nonmatch probability (NMP), effectively captures any implicating evidence of the comparison. NMP is computed using state-of-the-art matchers and is easy to interpret. To incorporate the effect of image quality, number of minutiae, and size of the latent on NMP value, we compute the NMP vs. match score plots separately for image pairs (latent and exemplar prints) with different characteristics. Given the paucity of latent fingerprint databases in public domain, we simulate latent prints using two exemplar print databases (NIST SD-14 and Michigan State Police) by cropping regions of three different sizes. We appropriately validate this simulation using four latent databases (NIST SD-27 and three proprietary latent databases) and two state-of-the-art fingerprint matchers to compute their respective match scores. We also discuss a practical scenario where a latent examiner uses the proposed framework to compute the evidential value of a latent-exemplar print pair comparison. Abhishek Nagar, Heeseung Choi, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2011 | On the evidential value of fingerprintsabstractFingerprint evidence is routinely used by forensics and law enforcement agencies worldwide to apprehend and convict criminals, a practice in use for over 100 years. The use of fingerprints has been accepted as an infallible proof of identity based on two premises: (i) permanence or persistence, and (ii) uniqueness or individuality. However, in the absence of any theoretical results that establish the unique ness or individuality of fingerprints, the use of fingerprints in various court proceedings is being questioned. This has raised awareness in the forensics community about the need to quantify the evidential value of fingerprint matching. A few studies that have studied this problem estimate this evidential value in one of two ways: (i) feature modeling, where a statistical (generative) model for fingerprint features, primarily minutiae, is developed which is then used to estimate the matching error and (ii) match score modeling, where a set of match scores obtained over a database is used to estimate the matching error rates. Our focus here is on match score modeling and we develop metrics to evaluate the effectiveness and reliability of the proposed evidential measure. Compared to previous approaches, the proposed measure allows explicit utilization of prior odds. Further, we also incorporate fingerprint image quality to improve the reliability of the estimated evidential value. Heeseung Choi, Abhishek Nagar, Anil K. Jain 0001 |
IJCB | 1 |
| 2010 | Mosaicing touchless and mirror-reflected fingerprint imagesabstractTouchless fingerprint sensing technologies have been explored to solve problems in touch-based sensing techniques because they do not require any contact between a sensor and a finger. While they can solve problems caused by the contact of a finger, other difficulties emerge such as a view difference problem and a limited usable area due to perspective distortion. In order to overcome these difficulties, we propose a new touchless fingerprint sensing device capturing three different views at one time and a method for mosaicing these view-different images. The device is composed of a single camera and two planar mirrors reflecting side views of a finger, and it is an alternative to expensive multiple-camera-based systems. The mosaic method can composite the multiple view images by using the thin plate spline model to expand the usable area of a fingerprint image. In particular, to reduce the affect of perspective distortion, we select the regions in each view by minimizing the ridge interval variations in a final mosaiced image. Results are promising as our experiments show that mosaiced images offer 29% more true minutiae and 28% larger good quality area than one-view, unmosaiced images. Also, when the side-view images are matched to the mosaiced images, it gives more matched minutiae than matching with one-view frontal images. We expect that the proposed method can reduce the view difference problem and increase the usable area of a touchless fingerprint image. Furthermore, the proposed method can be applied to other biometric applications requiring a large template for recognition. Heeseung Choi, Kyoungtaek Choi, Jaihie Kim |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2008 | Fingerprint-Quality Index Using Gradient ComponentsabstractFingerprint image-quality checking is one of the most important issues in fingerprint recognition because recognition is largely affected by the quality of fingerprint images. In the past, many related fingerprint-quality checking methods have typically considered the condition of input images. However, when using the preprocessing algorithm, ridge orientation may sometimes be extracted incorrectly. Unwanted false minutiae can be generated or some true minutiae may be ignored, which can also affect recognition performance directly. Therefore, in this paper, we propose a novel quality-checking algorithm which considers the condition of the input fingerprints and orientation estimation errors. In the experiments, the 2-D gradients of the fingerprint images were first separated into two sets of 1-D gradients. Then, the shapes of the probability density functions of these gradients were measured in order to determine fingerprint quality. We used the FVC2002 database and synthetic fingerprint images to evaluate the proposed method in three ways: 1) estimation ability of quality; 2) separability between good and bad regions; and 3) verification performance. Experimental results showed that the proposed method yielded a reasonable quality index in terms of the degree of quality degradation. Also, the proposed method proved superior to existing methods in terms of separability and verification performance. Sanghoon Lee 0003, Heeseung Choi, Kyoungtaek Choi, Jaihie Kim |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2008 | Recognizable-Image Selection for Fingerprint Recognition With a Mobile-Device CameraabstractThis paper proposes a recognizable-image selection algorithm for fingerprint-verification systems that use a camera embedded in a mobile device. A recognizable image is defined as the fingerprint image which includes the characteristics that are sufficiently discriminating an individual from other people. While general camera systems obtain focused images by using various gradient measures to estimate high-frequency components, mobile cameras cannot acquire recognizable images in the same way because the obtained images may not be adequate for fingerprint recognition, even if they are properly focused. A recognizable image has to meet the following two conditions: First, valid region in the recognizable image should be large enough compared with other nonrecognizable images. Here, a valid region is a well-focused part, and ridges in the region are clearly distinguishable from valleys. In order to select valid regions, this paper proposes a new focus-measurement algorithm using the secondary partial derivatives and a quality estimation utilizing the coherence and symmetry of gradient distribution. Second, rolling and pitching degrees of a finger measured from the camera plane should be within some limit for a recognizable image. The position of a core point and the contour of a finger are used to estimate the degrees of rolling and pitching. Experimental results show that our proposed method selects valid regions and estimates the degrees of rolling and pitching properly. In addition, fingerprint-verification performance is improved by detecting the recognizable images. Kyoungtaek Choi, Heeseung Choi, Jaihie Kim |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2007 | Fingerprint Image Mosaicking by Recursive Ridge MappingabstractTo obtain a large fingerprint image from several small partial images, mosaicking of fingerprint images has been recently researched. However, existing approaches cannot provide accurate transformations for mosaics when it comes to aligning images because of the plastic distortion that may occur due to the nonuniform contact between a finger and a sensor or the deficiency of the correspondences in the images. In this paper, we propose a new scheme for mosaicking fingerprint images, which iteratively matches ridges to overcome the deficiency of the correspondences and compensates for the amount of plastic distortion between two partial images by using a thin-plate spline model. The proposed method also effectively eliminates erroneous correspondences and decides how well the transformation is estimated by calculating the registration error with a normalized distance map. The proposed method consists of three phases: feature extraction, transform estimation, and mosaicking. Transform is initially estimated with matched minutia and the ridges attached to them. Unpaired ridges in the overlapping area between two images are iteratively matched by minimizing the registration error, which consists of the ridge matching error and the inverse consistency error. During the estimation, erroneous correspondences are eliminated by considering the geometric relationship between the correspondences and checking if the registration error is minimized or not. In our experiments, the proposed method was compared with three existing methods in terms of registration accuracy, image quality, minutia extraction rate, processing time, reject to fuse rate, and verification performance. The average registration error of the proposed method was less than three pixels, and the maximum error was not more than seven pixels. In a verification test, the equal error rate was reduced from 10% to 2.7% when five images were combined by our proposed method. The proposed method was superior to other compared methods in terms of registration accuracy, image quality, minutia extraction rate, and verification. Kyoungtaek Choi, Heeseung Choi, Sangyoun Lee, Jaihie Kim |
IEEE Trans. Syst. Man Cybern. Part B | 2 |