EDBT 2026 Demo / reviewers in the wild / expert
Jianjiang Feng
dblp:64/5611
· DBLP profile ↗
114ranked-venue papers
7as first author
46since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 2 first-author · 23 since 2021Artificial intelligence and machine learning · 51 · 6 first-author · 14 since 2021Security and privacy · 28 · 1 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 17 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PianoBand: A Multimodal Wristband Interface for Portable Piano Interaction
Chentao Li 0001, Zihang Ao, Jianjiang Feng, Jie Zhou 0001 |
CHI | 5 |
| 2026 | WristPP: A Wrist-Worn System for Hand Pose and Pressure EstimationabstractAccurate 3D hand pose and pressure sensing is essential for immersive human-computer interaction, yet simultaneously achieving both in mobile scenarios remains a significant challenge. We present WristP2, a camera-based wrist-worn system that estimates 3D hand pose and per-vertex pressure from a single wide-FOV RGB frame in real time. A ViT (Vision Transformer) backbone with joint-aligned tokens predicts Hand–VQ–VAE codebook indices for mesh recovery, while an extrinsics-conditioned branch jointly estimates per-vertex pressure. On a self-collected dataset of 133,000 frames (20 subjects; 48 on-plane and 28 mid-air gestures), WristP2 attains MPJPE (Mean Per-Joint Position Error) of 2.9 mm, Contact \(\operatorname{IoU}\) of 0.712, \(\operatorname{Vol.IoU}\) of 0.618, and foreground pressure MAE of 10.4 g. Across three user studies, WristP2 delivers touchpad-level efficiency in mid-air pointing and robust multi-finger pressure control on an uninstrumented desktop. In a real-world large-display Whac-A-Mole task, WristP2 also enables higher success ratio and lower arm fatigue than head-mounted camera-based baselines. These results position WristP2 as an effective, mobile solution for versatile pose- and pressure-based interaction. Ziheng Xi, Zihang Ao, Wanmei Zhang, Jianjiang Feng, Jie Zhou 0001 |
CHI | 6 |
| 2026 | LiDAR-FMC: Accurate and Robust Human Capture From Point-Cloud Video
Bohao Fan, Wenzhao Zheng, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Fixed-Length Dense Fingerprint Representation With Alignment and Robust EnhancementabstractFixed-length fingerprint representations, which map each fingerprint to a compact and fixed-size feature vector, are computationally efficient and well-suited for large-scale matching. However, designing a robust representation that effectively handles diverse fingerprint modalities, pose variations, and noise interference remains a significant challenge. In this work, we propose a fixed-length dense descriptor of fingerprints, and introduce FLARE—a fingerprint matching framework that integrates the Fixed-Length dense descriptor with pose-based Alignment and Robust Enhancement. This fixed-length representation employs a three-dimensional dense descriptor to effectively capture spatial relationships among fingerprint ridge structures, enabling robust and locally discriminative representations. To ensure consistency within this dense feature space, FLARE incorporates pose-based alignment using complementary estimation methods, along with dual enhancement strategies that refine ridge clarity while preserving the original fingerprint modality. The proposed dense descriptor supports fixed-length representation while maintaining spatial correspondence, enabling fast and accurate similarity computation. Extensive experiments demonstrate that FLARE achieves superior performance across rolled, plain, latent, and contactless fingerprints, significantly outperforming existing methods in cross-modality and low-quality scenarios. Further analysis validates the effectiveness of the dense descriptor design, as well as the impact of alignment and enhancement modules on the accuracy of dense descriptor matching. Experimental results highlight the effectiveness and generalizability of FLARE as a unified and scalable solution for robust fingerprint representation and matching. The implementation and code will be publicly available at our GitHub repository. Xiongjun Guan, Yongjie Duan, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | BiFingerPose: Bimodal Finger Pose Estimation for Touch DevicesabstractFinger pose offers promising opportunities to expand human computer interaction capability of touchscreen devices. Existing finger pose estimation algorithms that can be implemented in portable devices predominantly rely on capacitive images, which are currently limited to estimating pitch and yaw angles and exhibit reduced accuracy when processing large-angle inputs (especially when it is greater than 45 degrees). In this paper, we propose BiFingerPose, a novel bimodal based finger pose estimation algorithm capable of simultaneously and accurately predicting comprehensive finger pose information. A bimodal input is explored, including a capacitive image and a fingerprint patch obtained from the touchscreen with an under-screen fingerprint sensor. Our approach leads to reliable estimation of roll angle, which is not achievable using only a single modality. In addition, the prediction performance of other pose parameters has also been greatly improved. The evaluation of a 12-person user study on continuous and discrete interaction tasks further validated the advantages of our approach. Specifically, BiFingerPose outperforms previous SOTA methods with over 21% improvement in prediction performance,$2.5\times$higher task completion efficiency, and 23% better user operation accuracy, demonstrating its practical superiority. Finally, we delineate the application space of finger pose with respect to enhancing authentication security and improving interactive experiences, and develop corresponding prototypes to showcase the interaction potential. Our code will be available athttps://github.com/XiongjunGuan/DualFingerPose. Xiongjun Guan, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | FineType: Fine-grained Tapping Gesture Recognition for Text Entry
Chentao Li 0001, Ziheng Xi, Jianjiang Feng, Jie Zhou 0001 |
CHI | 3 |
| 2025 | FingerGlass: Enhancing Smart Glasses Interaction via Fingerprint SensingabstractFigure 1: FingerGlass, an interaction technique for smart glasses.A fingerprint sensor, highlighted in the red circle, is seamlessly integrated onto the temple arm of the smart glasses.This strategic placement allows users to perform various finger gestures, providing a discreet, efficient, and ergonomic input solution for controlling smart glasses functions. Zhanwei Xu, Haoxiang Pei, Jianjiang Feng, Jie Zhou 0001 |
CHI | 3 |
| 2025 | Text-guided Sparse Voxel Pruning for Efficient 3D Visual GroundingabstractIn this paper, we propose an efficient multi-level convolution architecture for 3D visual grounding. Conventional methods are difficult to meet the requirements of real-time inference due to the two-stage or point-based architecture. Inspired by the success of multi-level fully sparse convolutional architecture in 3D object detection, we aim to build a new 3D visual grounding framework following this technical route. However, as in 3D visual grounding task the 3D scene representation should be deeply interacted with text features, sparse convolution-based architecture is inefficient for this interaction due to the large amount of voxel features. To this end, we propose text-guided pruning (TGP) and completion-based addition (CBA) to deeply fuse 3D scene representation and text features in an efficient way by gradual region pruning and target completion. Specifically, TGP iteratively sparsifies the 3D scene representation and thus efficiently interacts the voxel features with text features by cross-attention. To mitigate the affect of pruning on delicate geometric information, CBA adaptively fixes the over-pruned region by voxel completion with negligible computational overhead. Compared with previous single-stage methods, our method achieves top inference speed and surpasses previous fastest method by 100% FPS. Our method also achieves state-of-the-art accuracy even compared with two-stage methods, with +1.13 lead of [email protected] on ScanRefer, and +2.6 and +3.2 leads on NR3D and SR3D respectively. The code is available at https://github.com/GWxuan/TSP3D. Xiuwei Xu, Ziwei Wang 0010, Jianjiang Feng, Jie Zhou 0001, Jiwen Lu |
CVPR | 4 |
| 2025 | Contactless Fingerprint Recognition Guided by 3D Finger PoseabstractContactless fingerprint recognition has gained increasing attention in recent years, particularly for the advantage of hygienic acquisition. However, existing contactless recognition methods primarily focus on frontal fingerprint images and rarely consider unconstrained 3D pose variations, which are common when capturing fingerprints using mobile devices. Pose variations can significantly affect matching performance due to perspective distortion, occlusion, and focus issues. In this paper, we demonstrate that 3D pose information of contactless fingerprints can be utilized to enhance the robustness and performance of existing recognition systems by guiding the acquisition process and constraining finger poses. For effective utilization of 3D finger pose, we develop a multi-branch 3D pose estimation network for contactless fingerprints and evaluate its performance on a newly collected dataset, which contains contactless fingerprint images with ground truth 3D pose annotations. Experimental results show that our proposed method outperforms previous approaches on 3D finger pose estimation. Furthermore, we propose several strategies to apply 3D finger pose information to existing recognition systems and validate their rationality through verification experiments. Haoxiang Pei, Xiongjun Guan, Jianjiang Feng, Jie Zhou 0001 |
IJCB | 4 |
| 2025 | IGL-Nav: Incremental 3D Gaussian Localization for Image-Goal NavigationabstractVisual navigation with an image as goal is a fundamental and challenging problem. Conventional methods either rely on end-to-end RL learning or modular-based policy with topological graph or BEV map as memory, which cannot fully model the geometric relationship between the explored 3D environment and the goal image. In order to efficiently and accurately localize the goal image in 3D space, we build our navigation system upon the renderable 3D gaussian (3DGS) representation. However, due to the computational intensity of 3DGS optimization and the large search space of 6-DoF camera pose, directly leveraging 3DGS for image localization during agent exploration process is prohibitively inefficient. To this end, we propose IGL-Nav, an Incremental 3D Gaussian Localization framework for efficient and 3D-aware image-goal navigation. Specifically, we incrementally update the scene representation as new images arrive with feed-forward monocular prediction. Then we coarsely localize the goal by leveraging the geometric information for discrete space matching, which can be equivalent to efficient 3D convolution. When the agent is close to the goal, we finally solve the fine target pose with optimization via differentiable rendering. The proposed IGL-Nav outperforms existing state-of-the-art methods by a large margin across diverse experimental configurations. It can also handle the more challenging free-view image-goal setting and be deployed on real-world robotic platform using a cellphone to capture goal image at arbitrary pose. Project page: https://gwxuan.github.io/IGL-Nav/. Xiuwei Xu, Ziwei Wang 0010, Jianjiang Feng, Jie Zhou 0001, Jiwen Lu |
ICCV | 5 |
| 2025 | 3D Touch Force Estimation from Capacitive Images
Chentao Li 0001, Jianjiang Feng, Jie Zhou 0001 |
IUI | 3 |
| 2025 | LiCamPose: Combining Multi-View LiDAR and RGB Cameras for Robust Single-timestamp 3D Human Pose EstimationabstractSeveral methods have been proposed to estimate 3D human pose from multi-view images, achieving satisfactory performance on public datasets collected under relatively simple conditions. However, there are limited approaches studying extracting 3D human skeletons from multimodal inputs, such as RGB and point cloud data. To address this gap, we introduce LiCamPose, a pipeline that integrates multi-view RGB and sparse point cloud information to estimate robust 3D human poses via single timestamp. We demonstrate the effectiveness of the volumetric architecture in combining these modalities. Furthermore, to circum-vent the need for manually labeled 3D human pose annotations, we develop a synthetic dataset generator for pretraining and design an unsupervised domain adaptation strategy to train a 3D human pose estimator without manual an-notations. To validate the generalization capability of our method, LiCamPose is evaluated on four datasets, including two public datasets, one synthetic dataset, and one challenging self-collected dataset named BasketBall, covering diverse scenarios. The results demonstrate that LiCamPose exhibits great generalization performance and significant application potential. The code, generator, and datasets are available at https://github.com/Yu-Yy/LiCamPose. Zhicheng Zhong 0001, Jianjiang Feng, Jie Zhou 0001 |
WACV | 5 |
| 2025 | Joint Identity Verification and Pose Alignment for Partial FingerprintsabstractCurrently, portable electronic devices are becoming more and more popular. For lightweight considerations, their fingerprint recognition modules usually use limited-size sensors. However, partial fingerprints have few matchable features, especially when there are differences in finger pressing posture or image quality, which makes partial fingerprint verification challenging. Most existing methods regard fingerprint position rectification and identity verification as independent tasks, ignoring the coupling relationship between them—relative pose estimation typically relies on paired features as anchors, and authentication accuracy tends to improve with more precise pose alignment. In this paper, we propose a novel framework for joint identity verification and pose alignment of partial fingerprint pairs, aiming to leverage their inherent correlation to improve each other. To achieve this, we present a multi-task CNN (Convolutional Neural Network)-Transformer hybrid network, and design a pre-training task to enhance the feature extraction capability. Experiments on multiple public datasets (NIST SD14, FVC2002 DB1_A & DB3_A, FVC2004 DB1_A & DB2_A, FVC2006 DB1_A) and an in-house dataset demonstrate that our method achieves state-of-the-art performance in both partial fingerprint verification and relative pose estimation, while being more efficient than previous methods. Code is available at:https://github.com/XiongjunGuan/JIPNet. Xiongjun Guan, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Finger Pose Estimation for Under-Screen Fingerprint SensorabstractTwo-dimensional pose estimation plays a crucial role in fingerprint recognition by facilitating global alignment and reduce pose-induced variations. However, existing methods are still unsatisfactory when handling with large angle or small area inputs. These limitations are particularly pronounced on fingerprints captured by under-screen fingerprint sensors in smartphones. In this paper, we present a novel dual-modal input based network for under-screen fingerprint pose estimation. Our approach effectively integrates two distinct yet complementary modalities: texture details extracted from ridge patches through the under-screen fingerprint sensor, and rough contours derived from capacitive images obtained via the touch screen. This collaborative integration endows our network with more comprehensive and discriminative information, substantially improving the accuracy and stability of pose estimation. A decoupled probability distribution prediction task is designed, instead of the traditional supervised forms of numerical regression or heatmap voting, to facilitate the training process. Additionally, we incorporate a Mixture of Experts (MoE) based feature fusion mechanism and a relationship driven cross-domain knowledge transfer strategy to further strengthen feature extraction and fusion capabilities. Extensive experiments are conducted on several public datasets and two private datasets. The results indicate that our method is significantly superior to previous state-of-the-art (SOTA) methods and remarkably boosts the recognition ability of fingerprint recognition algorithms. Our code is available at https://github.com/XiongjunGuan/DRACO. Xiongjun Guan, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | VesselDiffusion: 3D Vascular Structure Generation Based on Diffusion Modelabstract3D vascular structure models are pivotal in disease diagnosis, surgical planning, and medical education. The intricate nature of the vascular system presents significant challenges in generating accurate vascular structures. Constrained by the complex connectivity of the overall vascular structure, existing methods primarily focus on generating local or individual vessels. In this paper, we introduce a novel two-stage framework termed VesselDiffusion for the generation of detailed vascular structures, which is more valuable for medical analysis. Given that training data for specific vascular structure is often limited, direct generation of 3D data often results in inadequate detail and insufficient diversity. To this end, we initially train a 2D vascular generation model utilizing extensively available generic 2D vascular datasets. Taking the generated 2D images as input, a conditional diffusion model, integrating a dual-stream feature extraction (DSFE) module, is proposed to extrapolate 3D vascular systems. The DSFE module, comprising a Vision Transformer and a Graph Convolutional Network, synergistically captures visual features of global connection rationality and structural features of local vascular details, ensuring the authenticity and diversity of the generated 3D data. To the best of our knowledge, VesselDiffusion is the first model designed for generating comprehensive and realistic vascular networks with diffusion process. Comparative analyses with other generation methodologies demonstrate that the proposed framework achieves superior accuracy and diversity. Our code is available at: https://github.com/gzq17/VesselDiffusion. Zhanqiang Guo, Zimeng Tan, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | LiDAR-HMR: 3D Human Mesh Recovery From LiDARabstractHuman mesh recovery (HMR) holds significant utility in many applications. Studying HMR involving various types of sensors is necessary, as it enables the acquisition of human meshes in diverse scenes. Unlike HMR based on RGB images, HMR based on LiDAR has received considerably less attention in previous works. The major challenge in estimating human poses and meshes from sparse point clouds lies in the sparsity, noise, and incompletion of LiDAR point clouds. To address these challenges, we propose a LiDAR-based 3D human mesh recovery algorithm, called LiDAR-HMR. This algorithm involves estimating a sparse representation of a human (3D human pose) and gradually reconstructing the body mesh. To better leverage the 3D structural information of point clouds, we propose a point-cloud-to-SMPL pipeline that uses the original point cloud features to guide the reconstruction. The experimental results on four publicly available datasets demonstrate the effectiveness of LiDAR-HMR. The codes are available athttps://github.com/soullessrobot/LiDAR-HMR. Bohao Fan, Wenzhao Zheng, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | HumanReg: Self-supervised Non-rigid Registration of Human Point CloudabstractIn this paper, we present a novel registration framework, HumanReg, that learns a non-rigid transformation between two human point clouds end-to-end. We introduce body prior into the registration process to efficiently handle this type of point cloud. Unlike most exsisting supervised registration techniques that require expensive point-wise flow annotations, HumanReg can be trained in a self-supervised manner benefiting from a set of novel loss functions. To make our model better converge on real-world data, we also propose a pretraining strategy, and a synthetic dataset (HumanSyn4D) consists of dynamic, sparse human point clouds and their auto-generated ground truth annotations. Our experiments shows that HumanReg achieves state-of-the-art performance on CAPE-512 dataset and gains a qualitative result on another more challenging real-world dataset. Furthermore, our ablation studies demonstrate the effectiveness of our synthetic dataset and novel loss functions. Our code and synthetic dataset is available at https://github.com/chenyifanthu/HumanReg. Zhicheng Zhong 0001, Jianjiang Feng, Jie Zhou 0001 |
3DV | 5 |
| 2024 | LiDAR-Based Person Re-IdentificationabstractCamera-based person re-identification (ReID) systems have been widely applied in the field of public security. However, cameras often lack the perception of 3D morphological information of human and are susceptible to various limitations, such as inadequate illumination, complex background, and personal privacy. In this paper, we propose a LiDAR-based ReID framework, ReID3D, that utilizes pre-training strategy to retrieve features of 3D body shape and introduces Graph-based Complementary Enhancement Encoder for extracting comprehensive features. Due to the lack of LiDAR datasets, we build LReID, the first LiDAR-based person ReID dataset, which is collected in several outdoor scenes with variations in natural conditions. Additionally, we introduce LReID-sync, a simulated pedestrian dataset designed for pre-training encoders with tasks of point cloud completion and shape parameter learning. Extensive experiments on LReID show that ReID3D achieves exceptional performance with a rank-1 accuracy of 94.0, highlighting the significant potential of LiDAR in addressing person ReID tasks. To the best of our knowledge, we are the first to propose a solution for LiDAR-based ReID. The code and dataset are available at https://github.com/GWxuan/ReID3D. Yingping Liang, Ziheng Xi, Zhicheng Zhong 0001, Jianjiang Feng, Jie Zhou 0001 |
CVPR | 6 |
| 2024 | Camera-LiDAR Cross-Modality Gait Recognition
Yingping Liang, Ziheng Xi, Jianjiang Feng, Jie Zhou 0001 |
ECCV (34) | 5 |
| 2024 | Latent Fingerprint Matching via Dense Minutia DescriptorabstractLatent fingerprint matching is a daunting task, primarily due to the poor quality of latent fingerprints. In this study, we propose a deep-learning based dense minutia descriptor (DMD) for latent fingerprint matching. A DMD is obtained by extracting the fingerprint patch aligned by its central minutia, capturing detailed minutia information and texture information. Our dense descriptor takes the form of a three-dimensional representation, with two dimensions associated with the original image plane and the other dimension representing the abstract features. Additionally, the extraction process outputs the fingerprint segmentation map, ensuring that the descriptor is only valid in the foreground region. The matching between two descriptors occurs in their overlapping regions, with a score normalization strategy to reduce the impact brought by the differences outside the valid area. Our descriptor achieves state-of-the-art performance on several latent fingerprint datasets. Overall, our DMD is more representative and interpretable compared to previous methods. The corresponding code is available at https://github.com/Yu-Yy/DMD. Yongjie Duan, Xiongjun Guan, Jianjiang Feng, Jie Zhou 0001 |
IJCB | 4 |
| 2024 | ZJUT-EIFD: A Synchronously Collected External and Internal Fingerprint DatabaseabstractExternal fingerprints (EFs) based only on epidermal information are vulnerable to spoofing attacks and non-ideal skin conditions. To solve such shortcomings, internal fingerprints (IFs) collected using optical coherence tomography (OCT) have been proposed and widely researched. However, the development of IF is limited by the lack of in-depth researches on the IF and the EF-IF interoperability, which is partially caused by the lack of public OCT database. The obvious gap in the applications of EF and IF recognition motivated us to design and publish a comprehensive fingerprint database containing both traditional EFs and OCT IFs, denoted as ZJUT-EIFD. To the best of our knowledge, ZJUT-EIFD is the first public database that combines OCT and total internal reflection (TIR) via synchronous acquisition, with 399 different fingers from 60 subjects. In this article, the composition of the database, the quality of EFs and IFs, and the verification performance of different types of fingerprints were detailed. In addition, potential application directions of ZJUT-EIFD were demonstrated. ZJUT-EIFD can serve benchmarks and interoperability tests for EF-IF research, which will promote the research and development of EF and IF. Haohao Sun, Haixia Wang 0002, Yilong Zhang 0001, Ronghua Liang, Peng Chen 0008, Jianjiang Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Phase-Aggregated Dual-Branch Network for Efficient Fingerprint Dense RegistrationabstractFingerprint dense registration aims to finely align fingerprint pairs at the pixel level, thereby reducing intra-class differences caused by distortion. Unfortunately, traditional methods exhibited subpar performance when dealing with low-quality fingerprints while suffering from slow inference speed. Although deep learning based approaches shows significant improvement in these aspects, their registration accuracy is still unsatisfactory. In this paper, we propose a Phase-aggregated Dual-branch Registration Network (PDRNet) to aggregate the advantages of both types of methods. A dual-branch structure with multi-stage interactions is introduced between correlation information at high resolution and texture feature at low resolution, to perceive local fine differences while ensuring global stability. Extensive experiments are conducted on more comprehensive databases compared to previous works. Experimental results demonstrate that our method reaches the state-of-the-art registration performance in terms of accuracy and robustness, while maintaining considerable competitiveness in efficiency. Xiongjun Guan, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | 3D Vascular Segmentation Supervised by 2D Annotation of Maximum Intensity ProjectionabstractVascular structure segmentation plays a crucial role in medical analysis and clinical applications. The practical adoption of fully supervised segmentation models is impeded by the intricacy and time-consuming nature of annotating vessels in the 3D space. This has spurred the exploration of weakly-supervised approaches that reduce reliance on expensive segmentation annotations. Despite this, existing weakly supervised methods employed in organ segmentation, which encompass points, bounding boxes, or graffiti, have exhibited suboptimal performance when handling sparse vascular structure. To alleviate this issue, we employ maximum intensity projection (MIP) to decrease the dimensionality of 3D volume to 2D image for efficient annotation, and the 2D labels are utilized to provide guidance and oversight for training 3D vessel segmentation model. Initially, we generate pseudo-labels for 3D blood vessels using the annotations of 2D projections. Subsequently, taking into account the acquisition method of the 2D labels, we introduce a weakly-supervised network that fuses 2D-3D deep features via MIP to further improve segmentation performance. Furthermore, we integrate confidence learning and uncertainty estimation to refine the generated pseudo-labels, followed by fine-tuning the segmentation network. Our method is validated on five datasets (including cerebral vessel, aorta and coronary artery), demonstrating highly competitive performance in segmenting vessels and the potential to significantly reduce the time and effort required for vessel annotation. Our code is available at: https://github.com/gzq17/Weakly-Supervised-by-MIP. Zhanqiang Guo, Zimeng Tan, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2023 | 3D Finger Rotation Estimation from Fingerprint ImagesabstractVarious touch-based interaction techniques have been developed to make interactions on mobile devices more effective, efficient, and intuitive. Finger orientation, especially, has attracted a lot of attentions since it intuitively brings three additional degrees of freedom (DOF) compared with two-dimensional (2D) touching points. The mapping of finger orientation can be classified as being either absolute or relative, suitable for different interaction applications. However, only absolute orientation has been explored in prior works. The relative angles can be calculated based on two estimated absolute orientations, although, a higher accuracy is expected by predicting relative rotation from input images directly. Consequently, in this paper, we propose to estimate complete 3D relative finger angles based on two fingerprint images, which incorporate more information with a higher image resolution than capacitive images. For algorithm training and evaluation, we constructed a dataset consisting of fingerprint images and their corresponding ground truth 3D relative finger rotation angles. Experimental results on this dataset revealed that our method outperforms previous approaches with absolute finger angle models. Further, extensive experiments were conducted to explore the impact of image resolutions, finger types, and rotation ranges on performance. A user study was also conducted to examine the efficiency and precision using 3D relative finger orientation in 3D object rotation task. Yongjie Duan, Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2023 | Monocular 3D Fingerprint Reconstruction and UnwarpingabstractCompared with contact-based fingerprint acquisition techniques, contactless acquisition has the advantages of less skin distortion, more complete fingerprint area, and hygienic acquisition. However, perspective distortion is a challenge in contactless fingerprint recognition, which changes the ridge frequency and relative minutiae location, and thus degrades the recognition accuracy. We propose a learning-based shape-from-texture algorithm to reconstruct a 3-D finger shape from a single image and unwarp the raw image to suppress the perspective distortion. Our experimental results for 3-D reconstruction on contactless fingerprint databases show that the proposed method has high 3-D reconstruction accuracy. Experimental results for contactless-to-contactless and contactless-to-contact-based fingerprint matching indicate that the proposed method can improve the matching accuracy. Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Estimating Fingerprint Pose via Dense VotingabstractAligning fingerprint images to a unified coordinate system defined by fingerprint pose is beneficial for fast and accurate fingerprint matching. Due to poor ridge quality and partial observations, however, performance of the state-of-the-art fingerprint pose estimation algorithms remains unsatisfactory. In this study, we propose to fuse voting strategy and deep network to estimate fingerprint center and direction. Rather than regressing them directly, we predict dense offset maps and vote for the final estimation. Experimental results on ten fingerprint datasets with over 60K fingerprints show that (1) highly consistent fingerprint pose estimations are obtained across different impressions of the same finger, (2) performance of fingerprint indexing and verification is further improved thanks to more accurate fingerprint pose estimation, and (3) the proposed approach is more robust to sensing technologies (optical, capacitive, inking, and direct imaging) and impression types (rolled, plain, latent, and contactless). Yongjie Duan, Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Regression of Dense Distortion Field From a Single Fingerprint ImageabstractSkin distortion is a long standing challenge in fingerprint matching, which causes false non-matches. Previous studies have shown that the recognition rate can be improved by estimating the distortion field from a distorted fingerprint and then rectifying it into a normal fingerprint. However, existing rectification methods are based on principal component representation of distortion fields, which is not accurate and are very sensitive to finger pose. In this paper, we propose a rectification method where a self-reference based network is utilized to directly estimate the dense distortion field of distorted fingerprint instead of its low dimensional representation. This method can output accurate distortion fields of distorted fingerprints with various finger poses and distortion patterns. We conducted experiments on FVC2004 DB1_A, expanded Tsinghua Distorted Fingerprint database (with additional distorted fingerprints in diverse finger poses and distortion patterns) and a latent fingerprint database. Experimental results demonstrate that our proposed method achieves the state-of-the-art rectification performance in terms of distortion field estimation and rectified fingerprint matching. Xiongjun Guan, Yongjie Duan, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | A Flexible Multi-view Multi-modal Imaging System for Outdoor ScenesabstractMulti-view imaging systems enable uniform coverage of 3D space and reduce the impact of occlusion, which is beneficial for 3D object detection and tracking accuracy. However, existing imaging systems built with multi-view cameras or depth sensors are limited by the small applicable scene and complicated composition. In this paper, we propose a wireless multi-view multi-modal 3D imaging system generally applicable to large outdoor scenes, which consists of a master node and several slave nodes. Multiple spatially distributed slave nodes equipped with cameras and LiDARs are connected to form a wireless sensor network. While providing flexibility and scalability, the system applies automatic spatio-temporal calibration techniques to obtain accurate 3D multi-view multi-modal data. This system is the first imaging system that integrates mutli-view RGB cameras and LiDARs in large outdoor scenes among existing 3D imaging systems. We perform point clouds based 3D object detection and long-term tracking using the 3D imaging dataset collected by this system. The experimental results show that multi-view point clouds greatly improve 3D object detection and tracking accuracy regardless of complex and various outdoor environments. Bohao Fan, Jianjiang Feng, Jie Zhou 0001 |
3DV | 5 |
| 2022 | Bridge-Prompt: Towards Ordinal Action Understanding in Instructional VideosabstractAction recognition models have shown a promising capability to classify human actions in short video clips. In a real scenario, multiple correlated human actions commonly occur in particular orders, forming semantically meaningful human activities. Conventional action recognition approaches focus on analyzing single actions. However, they fail to fully reason about the contextual relations between adjacent actions, which provide potential temporal logic for understanding long videos. In this paper, we propose a prompt-based framework, Bridge-Prompt (Br-Prompt), to model the semantics across adjacent actions, so that it simultaneously exploits both out-of-context and contextual information from a series of ordinal actions in instructional videos. More specifically, we reformulate the individual action labels as integrated text prompts for super-vision, which bridge the gap between individual action semantics. The generated text prompts are paired with corresponding video clips, and together co-train the text encoder and the video encoder via a contrastive approach. The learned vision encoder has a stronger capability for ordinal-action-related downstream tasks, e.g. action segmentation and human activity recognition. We evaluate the performances of our approach on several video datasets: Georgia Tech Egocentric Activities (GTEA), 50Salads, and the Breakfast dataset. Br-Prompt achieves state-of-the-art on multiple benchmarks. Code is available at: https://github.com/ttlmh/Bridge-Prompt. Muheng Li, Lei Chen 0069, Yueqi Duan, Zhilan Hu, Jianjiang Feng, Jie Zhou 0001, Jiwen Lu |
CVPR | 5 |
| 2022 | Label2Label: A Language Modeling Framework for Multi-attribute Learning
Wanhua Li 0001, Zhexuan Cao, Jianjiang Feng, Jie Zhou 0001, Jiwen Lu |
ECCV (12) | 3 |
| 2022 | Direct Regression of Distortion Field from a Single Fingerprint ImageabstractSkin distortion is a long standing challenge in fingerprint matching, which causes false non-matches. Previous studies have shown that the recognition rate can be improved by estimating the distortion field from a distorted fingerprint and then rectifying it into a normal fingerprint. However, existing rectification methods are based on principal component representation of distortion fields, which is not accurate and are very sensitive to finger pose. In this paper, we propose a rectification method where a self-reference based network is utilized to directly estimate the dense distortion field of distorted fingerprint instead of its low dimensional representation. This method can output accurate distortion fields of distorted fingerprints with various finger poses. Considering the limited number and variety of distorted fingerprints in the existing public dataset, we collected more distorted fingerprints with diverse finger poses and distortion patterns as a new database. Experimental results demonstrate that our proposed method achieves the state-of-the-art rectification performance in terms of distortion field estimation and rectified fingerprint matching. Xiongjun Guan, Yongjie Duan, Jianjiang Feng, Jie Zhou 0001 |
IJCB | 3 |
| 2022 | Estimating 3D Finger Pose via 2D-3D Fingerprint MatchingabstractTouchscreens have become the primary input devices for smartphones, tablet computers, and other intelligent devices over the past decades. While for the most pervasive commercial devices, only 2D touch positions on the screen are utilized as interaction inputs. To extend the richness of the input vocabulary, some researchers have proposed several innovative interaction techniques, e.g. finger pose. However, due to the low resolution and lacking in information of capacitive images, only two angles, pitch and yaw, are considered in most finger pose estimation algorithms, and the accuracy is not sufficiently high for large scale applications in smartphones. With the rapid development of under-screen fingerprint sensing technology, a new input modality, fingerprint image, for 3D finger pose estimation is available from these fingerprint sensors. In this paper, we propose a finger specific algorithm for estimating 3D finger pose including roll, pitch, and yaw from fingerprint images. 3D finger surface is first reconstructed based on sequential fingerprint images captured in enrollment, and given this 3D surface model, 3D finger pose of a test fingerprint is estimated by matching keypoints between the 2D image and 3D point cloud and minimizing the projection error. The proposed approach is a non-learning algorithm with good generalization ability and robustness in real applications. To evaluate the performance of our method, a dataset of fingerprint images with their corresponding ground truth 3D angles is collected. Experimental results on this dataset demonstrate the effectiveness of introducing reconstructed 3D finger surface shape in 3D finger pose estimation. The average absolute errors of three angles are 10.74 for roll, 8.25 for pitch, and 7.38 for yaw, respectively. Extensive experiments are also conducted to explore the impact of touching area size and gallery size on performance. Yongjie Duan, Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
IUI | 3 |
| 2022 | Order-Constrained Representation Learning for Instructional Video PredictionabstractIn this paper, we propose a weakly-supervised approach called Order-Constrained Representation Learning (OCRL) to predict future actions from instructional videos by observing incomplete steps of actions. Most conventional methods focus on predicting actions based on partially observed video frames, which mainly study low-level semantics such as motion consistency. Unlike performing a single action, completing a task in an instructional video usually requires several steps of action and longer periods. Motivated by the fact that the order of action steps is key to learning task semantics, we develop a new frame of contrastive loss, called StepNCE, to integrate the shared semantic information between step order and task semantics under the framework of the memory bank-based momentum-updating algorithm. Specifically, we learn the video representations from step order-rearranged trimmed video clips based on the proposed task-consistency rule and order-consistency rule. Our StepNCE loss can be used to pre-train a video feature encoder, which is then fine-tuned to carry out the instructional video prediction task. Our approach digs deeper into the sequential logic between different action steps with respect to a certain task, which is able to promote the video understanding methods to a new semantic level. We evaluate our method on five popular instructional video and action prediction datasets: COIN, CrossTask, UT-Interaction, BIT-Interaction, and ActivityNet v1.2, and the results show that our approach gains improvements from conventional prediction methods. Muheng Li, Lei Chen 0069, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Latent Fingerprint Indexing: Robust Representation and Adaptive Candidate ListabstractEfficiently identifying the mated gallery fingerprint of a latent fingerprint in a large database requires a highly accurate and efficient fingerprint matching algorithm. The common strategy to achieve this goal is to combine an efficient indexing algorithm with a slow but accurate matching algorithm. Despite of the importance of latent indexing, it has received far less attention than rolled and plain fingerprint indexing. Due to the small fingerprint area, poor image quality and huge variety in information quantity of latent fingerprints, existing rolled and plain fingerprint indexing approaches cannot be simply migrated to the latent fingerprint indexing. In this paper, we propose (1) a multi-scale fixed-length representation approach for latent fingerprint indexing, and (2) a fingerprint information quantity estimation approach for adaptive candidate list reduction. The representation scheme is designed to deal with small finger area and low image quality of latents. The information quantity of a latent is a predictor of the indexing score of its mated gallery fingerprint and thus can be used to determine a proper threshold for its candidate list. Extensive experimental results on NIST SD27, MOLF, N2N, and Hisign latent fingerprint databases show that the proposed method achieved the state-of-the-art indexing accuracy on latent fingerprints, and significantly improved the efficiency of state-of-the-art latent matching algorithm. Shan Gu, Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | MetaAge: Meta-Learning Personalized Age EstimatorsabstractDifferent people age in different ways. Learning a personalized age estimator for each person is a promising direction for age estimation given that it better models the personalization of aging processes. However, most existing personalized methods suffer from the lack of large-scale datasets due to the high-level requirements: identity labels and enough samples for each person to form a long-term aging pattern. In this paper, we aim to learn personalized age estimators without the above requirements and propose a meta-learning method named MetaAge for age estimation. Unlike most existing personalized methods that learn the parameters of a personalized estimator for each person in the training set, our method learns the mapping from identity information to age estimator parameters. Specifically, we introduce a personalized estimator meta-learner, which takes identity features as the input and outputs the parameters of customized estimators. In this way, our method learns the meta knowledge without the above requirements and seamlessly transfers the learned meta knowledge to the test set, which enables us to leverage the existing large-scale age datasets without any additional annotations. Extensive experimental results on three benchmark datasets including MORPH II, ChaLearn LAP 2015 and ChaLearn LAP 2016 databases demonstrate that our MetaAge significantly boosts the performance of existing personalized methods and outperforms the state-of-the-art approaches. Wanhua Li 0001, Jiwen Lu, Abudukelimu Wuerkaixi, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Learning Probabilistic Ordinal Embeddings for Uncertainty-Aware RegressionabstractUncertainty is the only certainty there is. Modeling data uncertainty is essential for regression, especially in unconstrained settings. Traditionally the direct regression formulation is considered and the uncertainty is modeled by modifying the output space to a certain family of probabilistic distributions. On the other hand, classification based regression and ranking based solutions are more popular in practice while the direct regression methods suffer from the limited performance. How to model the uncertainty within the present-day technologies for regression remains an open issue. In this paper, we propose to learn probabilistic ordinal embeddings which represent each data as a multivariate Gaussian distribution rather than a deterministic point in the latent space. An ordinal distribution constraint is proposed to exploit the ordinal nature of regression. Our probabilistic ordinal embeddings can be integrated into popular regression approaches and empower them with the ability of uncertainty estimation. Experimental results show that our approach achieves competitive performance. Code is available at https://github.com/Li-Wanhua/POEs. Wanhua Li 0001, Xiaoke Huang 0001, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
CVPR | 4 |
| 2021 | Meta-Mining Discriminative Samples for Kinship VerificationabstractKinship verification aims to find out whether there is a kin relation for a given pair of facial images. Kinship verification databases are born with unbalanced data. For a database with N positive kinship pairs, we naturally obtain N(N − 1) negative pairs. How to fully utilize the limited positive pairs and mine discriminative information from sufficient negative samples for kinship verification remains an open issue. To address this problem, we propose a Discriminative Sample Meta-Mining (DSMM) approach in this paper. Unlike existing methods that usually construct a balanced dataset with fixed negative pairs, we propose to utilize all possible pairs and automatically learn discriminative information from data. Specifically, we sample an unbalanced train batch and a balanced meta-train batch for each iteration. Then we learn a meta-miner with the meta-gradient on the balanced meta-train batch. In the end, the samples in the unbalanced train batch are re-weighted by the learned meta-miner to optimize the kinship models. Experimental results on the widely used KinFaceW-I, KinFaceW-II, TSKinFace, and Cornell Kinship datasets demonstrate the effectiveness of the proposed approach. Wanhua Li 0001, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
CVPR | 4 |
| 2021 | Orientation Field Estimation for Latent Fingerprints with Prior Knowledge of Fingerprint PatternabstractEstimating orientation field for latent fingerprints plays a crucial role in latent fingerprints recognition systems. Due to poor quality and small area of latent fingerprints, however, the performance of the state-of-the-art algorithms is still far from satisfactory. Considering the intrinsic characteristics of fingerprints that the distribution of orientation field varies with the fingerprint patterns, we propose an orientation field estimation algorithm for latent fingerprints based on residual learning using prior knowledge of fingerprint patterns. Specifically, statistical distribution models of orientation field, for different fingerprint patterns, are calculated based on a large database consisting of 14,000 fingerprints with good quality using clustering method. The residual orientation fields and reliability scores, indicating the consistency with different statistical orientation models, are estimated using a deep network, named RefNet. Then the final orientation field is obtained by fusing the estimations according to their corresponding reliability scores. Experimental results on the widely used latent database NIST SD27 demonstrate that the proposed algorithm provides higher orientation field estimation accuracy compared with the state-of-the-art methods, and by enhancing latent fingerprints using estimated orientation field, the identification performance is further improved. Yongjie Duan, Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
IJCB | 2 |
| 2021 | 3D Multi-Object Online Tracking with Multi-View ClusteringabstractMulti-target tracking is an increasingly hot research topic. Since target tracking under a single view is difficult to solve the problem of target tracking, we propose an online tracking method that simultaneously locates and correlates the objects under multiple views. This idea consists mainly of a tracking-by-detection method and an internal association mechanism between cameras. First, we trained the LSTM classfication model for the tracking under a single view, in which we also retained the best features of unoccluded objects for re-identification. Then the hierarchical clustering is used for connecting the objects under different views with the accurate track fragments after camera calibration, and the output of clustering will be made for improving the effect of track fragments in the online tracking. In experiments, our method proved to be more robust to solve the occlusion problem, and can make superior performance in comparison with the state of the art methods on the Terrace videos and Passageway videos from EPFL CVLAB. Guanglie Jia, Jianjiang Feng, Jie Zhou 0001 |
ICIP | 2 |
| 2021 | SGNet: Structure-Aware Graph-Based Network for Airway Semantic Segmentation
Zimeng Tan, Jianjiang Feng, Jie Zhou 0001 |
MICCAI (1) | 2 |
| 2021 | 3D Multi-object Detection and Tracking with Sparse Stationary LiDAR
Jianjiang Feng, Jie Zhou 0001 |
PRCV (1) | 3 |
| 2021 | Dense Registration and Mosaicking of Fingerprints by Training an End-to-End NetworkabstractDense registration of fingerprints is a challenging task due to elastic skin distortion, low image quality, and self-similarity of ridge pattern. To overcome the limitation of handcraft features, we propose to train an end-to-end network to directly output pixel-wise displacement field between two fingerprints. The proposed network includes a siamese network for feature embedding, and a following encoder-decoder network for regressing displacement field. By applying displacement fields reliably estimated by tracing high quality fingerprint videos to challenging fingerprints, we synthesize a large number of training fingerprint pairs with ground truth displacement fields. In addition, based on the proposed registration algorithm, we propose a fingerprint mosaicking method based on optimal seam selection. Registration and matching experiments on FVC2004 databases, Tsinghua Distorted Fingerprint (TDF) database, and NIST SD27 latent fingerprint database show that our registration method outperforms previous dense registration methods in accuracy. Mosaicking experiments on FVC2004 DB1_A and a small fingerprint database demonstrate that the proposed algorithm produced higher quality fingerprints and led to higher matching accuracy, which also validates the performance of our registration algorithm. Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Surface and Internal Fingerprint Reconstruction From Optical Coherence Tomography Through Convolutional Neural NetworkabstractOptical coherence tomography (OCT), as a non-destructive and high-resolution fingerprint acquisition technology, is robust against poor skin conditions and resistant to spoof attacks. It measures fingertip information on and beneath skin as 3D volume data, containing the surface fingerprint, internal fingerprint and sweat glands. Various methods have been proposed to extract internal fingerprints, which ignore the inter-slice dependence and often require manually selected parameters. In this article, a modified U-Net that combines residual learning, bidirectional convolutional long short-term memory and hybrid dilated convolution (denoted as BCL-U Net) for OCT volume data segmentation and two fingerprint reconstruction approaches are proposed. To the best of our knowledge, it is the first time that simultaneous and automatic extraction is performed for surface fingerprint, internal fingerprint and sweat gland. The proposed BCL-U Net utilizes the spatial dependence in OCT volume data and deals with segmentation of objects with diverse sizes to achieve accurate extraction. Comparisons have been performed to demonstrate the advantages of the proposed method. A thorough evaluation of the recognition abilities of internal and surface fingerprints is conducted using a dataset significantly larger than previous studies. Four databases containing internal and surface fingerprints are generated from 1572 OCT volume data by the proposed method. The internal fingerprint matching experiment has achieved a lowest equal error rate (EER) of 0.95%. Mixed internal and surface fingerprint matching experiment is also performed and achieves an EER of 3.67%, verifying the consistency of the internal and surface fingerprints. The matching experiments for fingers under poor skin conditions show a 2.47% EER of internal fingerprints that is much lower than that of surface fingerprints, which proves the advantage of internal fingerprints and indicates the potential of the internal fingerprints to supplement or replace the surface fingerprints for some specific applications. Baojin Ding, Haixia Wang 0002, Peng Chen 0008, Yilong Zhang 0001, Zhenhua Guo 0001, Jianjiang Feng, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2021 | Latent Fingerprint Registration via Matching Densely Sampled PointsabstractLatent fingerprint matching is a very important but unsolved problem. As a key step of fingerprint matching, fingerprint registration has a great impact on the recognition performance. Existing latent fingerprint registration approaches are mainly based on establishing correspondences between minutiae, and hence will certainly fail when there are no sufficient number of extracted minutiae due to small fingerprint area or poor image quality. Minutiae extraction has become the bottleneck of latent fingerprint registration. In this paper, we propose a non-minutia latent fingerprint registration method which estimates the spatial transformation between a pair of fingerprints through a dense fingerprint patch alignment and matching procedure. Given a pair of fingerprints to match, we bypass the minutiae extraction step and take uniformly sampled points as key points. Then the proposed patch alignment and matching algorithm compares all pairs of sampling points and produces their similarities along with alignment parameters. Finally, a set of consistent correspondences are found by spectral clustering. Extensive experiments on NIST27 database and MOLF database show that the proposed method achieves the state-of-the-art registration performance, especially under challenging conditions. Code is made publicly available at: https://github.com/Gus233/Latent-Fingerprint-Registration. Shan Gu, Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Joint Estimation of Pose and Singular Points of FingerprintsabstractFingerprint pose estimation is a challenging problem since the pose is not defined by salient anatomical features and fingerprint images usually suffer from noise and small area. In this article, we proposed a method for joint estimation of pose and singular points of fingerprints, with the expectation that the pose and singular points can improve each other. By virtue of that singular points can be located accurately, we hope to improve the accuracy of pose estimation. Meanwhile, the robustness of pose estimation can improve the anti-noise performance of singular point detection. To achieve this, we propose a multi-task deep neural network, which contains a feature extraction body and two estimation heads for singular point and pose respectively. The proposed network can deal with various types of fingerprints, including plain, rolled and latent fingerprints. Experiments on four databases (NIST SD4, SD14, SD27 and FVC2004 DB1A) show that (1) the estimated poses and detected singular points are close to manual annotations despite of different image qualities; (2) the estimated poses for mated fingerprint pairs are consistent; and (3) the proposed pose estimation method outperforms state-of-the-art methods while utilized as pose constraint for a fingerprint indexing algorithm. Qihao Yin, Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Reasoning Graph Networks for Kinship Verification: From Star-Shaped to HierarchicalabstractIn this paper, we investigate the problem of facial kinship verification by learning hierarchical reasoning graph networks. Conventional methods usually focus on learning discriminative features for each facial image of a paired sample and neglect how to fuse the obtained two facial image features and reason about the relations between them. To address this, we propose a Star-shaped Reasoning Graph Network (S-RGN). Our S-RGN first constructs a star-shaped graph where each surrounding node encodes the information of comparisons in a feature dimension and the central node is employed as the bridge for the interaction of surrounding nodes. Then we perform relational reasoning on this star graph with iterative message passing. The proposed S-RGN uses only one central node to analyze and process information from all surrounding nodes, which limits its reasoning capacity. We further develop a Hierarchical Reasoning Graph Network (H-RGN) to exploit more powerful and flexible capacity. More specifically, our H-RGN introduces a set of latent reasoning nodes and constructs a hierarchical graph with them. Then bottom-up comparative information abstraction and top-down comprehensive signal propagation are iteratively performed on the hierarchical graph to update the node features. Extensive experimental results on four widely used kinship databases show that the proposed methods achieve very competitive results. Wanhua Li 0001, Jiwen Lu, Abudukelimu Wuerkaixi, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Graph-Based Social Relation Reasoning
Wanhua Li 0001, Yueqi Duan, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
ECCV (15) | 4 |
| 2020 | Graph-Based Kinship Reasoning NetworkabstractIn this paper, we propose a graph-based kinship reasoning (GKR) network for kinship verification, which aims to effectively perform relational reasoning on the extracted features of an image pair. Unlike most existing methods which mainly focus on how to learn discriminative features, our method considers how to compare and fuse the extracted feature pair to reason about the kin relations. The proposed GKR constructs a star graph called kinship relational graph where each peripheral node represents the information comparison in one feature dimension and the central node is used as a bridge for information communication among peripheral nodes. Then the GKR performs relational reasoning on this graph with recursive message passing. Extensive experimental results on the KinFaceW-I and KinFaceW-II datasets show that the proposed GKR outperforms the state-of-the-art methods. Wanhua Li 0001, Yingqiang Zhang, Kangchen Lv, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
ICME | 5 |
| 2019 | BridgeNet: A Continuity-Aware Probabilistic Network for Age EstimationabstractAge estimation is an important yet very challenging problem in computer vision. Existing methods for age estimation usually apply a divide-and-conquer strategy to deal with heterogeneous data caused by the non-stationary aging process. However, the facial aging process is also a continuous process, and the continuity relationship between different components has not been effectively exploited. In this paper, we propose BridgeNet for age estimation, which aims to mine the continuous relation between age labels effectively. The proposed BridgeNet consists of local regressors and gating networks. Local regressors partition the data space into multiple overlapping subspaces to tackle heterogeneous data and gating networks learn continuity aware weights for the results of local regressors by employing the proposed bridge-tree structure, which introduces bridge connections into tree models to enforce the similarity between neighbor nodes. Moreover, these two components of BridgeNet can be jointly learned in an end-to-end way. We show experimental results on the MORPH II, FG-NET and Chalearn LAP 2015 datasets and find that BridgeNet outperforms the state-of-the-art methods. Wanhua Li 0001, Jiwen Lu, Jianjiang Feng, Chunjing Xu, Jie Zhou 0001, Qi Tian 0001 |
CVPR | 3 |
| 2019 | Learning Deep Binary Descriptor with Multi-QuantizationabstractIn this paper, we propose an unsupervised feature learning method called deep binary descriptor with multi-quantization (DBD-MQ) for visual analysis. Existing learning-based binary descriptors such as compact binary face descriptor (CBFD) and DeepBit utilize the rigid sign function for binarization despite of data distributions, which usually suffer from severe quantization loss. In order to address the limitation, we propose a deep multi-quantization network to learn a data-dependent binarization in an unsupervised manner. More specifically, we design a K-Autoencoders (KAEs) network to jointly learn the parameters of feature extractor and the binarization functions under a deep learning framework, so that discriminative binary descriptors can be obtained with a fine-grained multi-quantization. As DBD-MQ simply allocates the same number of quantizers to each real-valued feature dimension ignoring the elementwise diversity of informativeness, we further propose a deep competitive binary descriptor with multi-quantization (DCBD-MQ) method to learn optimal allocation of bits with the fixed binary length in a competitive manner, where informative dimensions gain more bits for complete representation. Moreover, we present a similarity-aware binary encoding strategy based on the earth mover's distance of Autoencoders, so that elements that are quantized into similar Autoencoders will have smaller Hamming distances. Extensive experimental results on six widely-used datasets show that our DBD-MQ and DCBD-MQ outperform most state-of-the-art unsupervised binary descriptors. Yueqi Duan, Jiwen Lu, Ziwei Wang 0010, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Ordinal Deep Learning for Facial Age EstimationabstractIn this paper, we propose an ordinal deep learning approach for facial age estimation. Unlike conventional hand-crafted feature-based methods that require prior and expert knowledge, we propose an ordinal deep feature learning (ODFL) method to learn feature descriptors for face representation directly from raw pixels. Motivated by the fact that age labels are chronologically correlated and age estimation is an ordinal learning problem, our proposed ODFL enforces two criteria on the descriptors, which are learned at the top of the deep networks: 1) the topology-preserving ordinal relation is employed to exploit the order information in the learned feature space and 2) the age-difference cost information is leveraged to dynamically measure face pairs with different age value gaps. However, both the procedures of feature extraction and age estimation are learned independently in ODFL, which may lead to a sub-optimal problem. To address this, we further propose an end-to-end ordinal deep learning (ODL) framework, where the complementary information of both the procedures is exploited to reinforce our model. Extensive experimental results on five face aging datasets show that both our ODFL and ODL achieve superior performance in comparisons with most state-of-the-art methods. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Multi-Stream Deep Neural Networks for RGB-D Egocentric Action RecognitionabstractIn this paper, we investigate the problem of RGB-D egocentric action recognition. Unlike conventional human action videos that are passively recorded by static cameras, egocentric videos are self-generated from wearable sensors that are more flexible and provide the close-ups with the visual attention of the wearers when they act. Moreover, RGB-D videos contain the spatial appearance and temporal information in the RGB modality and reflect the 3D structure of the scenes in the depth modality. To adequately learn the nonlinear structure of heterogeneous representations from different modalities and exploit their complementary characteristics, we develop a multi-stream deep neural networks (MDNN) method, which aims to preserve the distinctive property for each modality and simultaneously explore their sharable information in a unified deep architecture. Specifically, we deploy a Cauchy estimator to maximize the correlations of the sharable components and enforce the orthogonality constraints on the distinctive components to guarantee their high independencies. Since the egocentric action recognition is usually sensitive to hand poses, we extend our MDNN by integrating with the hand cues to enhance the recognition accuracy. Extensive experimental results on a newly collected data set and two additional benchmarks are presented to demonstrate the effectiveness of our proposed method for RGB-D egocentric action recognition. Yansong Tang, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Uniform and Variational Deep Learning for RGB-D Object Recognition and Person Re-IdentificationabstractIn this paper, we propose a uniform and variational deep learning (UVDL) method for RGB-D object recognition and person re-identification. Unlike most existing object recognition and person re-identification methods, which usually use only the visual appearance information from RGB images, our method recognizes visual objects and persons with RGB-D images to exploit more reliable information such as geometric and anthropometric information that are robust to different viewpoints. Specifically, we extract the depth feature and the appearance feature from the depth and RGB images with two deep convolutional neural networks, respectively. In order to combine the depth feature and the appearance feature to exploit their relationship, we design a uniform and variational multi-modal auto-encoder at the top layer of our deep network to seek a uniform latent variable by projecting them into a common space, which contains the whole information of RGB-D images and has small intra-class variation and large inter-class variation, simultaneously. Finally, we optimize the auto-encoder layer and two deep convolutional neural networks jointly to minimize the discriminative loss and the reconstruction error. The experimental results on both RGB-D object recognition and RGB-D person re-identification are presented to show the efficiency of our proposed approach. Liangliang Ren, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Order-Sensitive Deep Hashing for Multimorbidity Medical Image Retrieval
Zhixiang Chen 0003, Ruojin Cai, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
MICCAI (1) | 4 |
| 2018 | Towards Accurate and Complete Registration of Coronary Arteries in CTA Images
Shaowen Zeng, Jianjiang Feng, Yunqiang An, Jiwen Lu, Jie Zhou 0001 |
MICCAI (2) | 2 |
| 2018 | Multi-scale Deep Representation Learning for Face DetectionabstractIn this paper, we propose a face detection method with multi-scale deep representation learning. While existing face detection methods have achieved good performance, they fail to consider the large intra-class variations between faces in the prediction stage and the relation between multi-scale proposals in the proposal stage. Our method encourages the network to learn compact features which simultaneously minimizes the intra-class variations and enlarges the inter-class variations. We can make more reliable binary classification between face regions and non-face ones. To obtain better proposals, we predict and select hard proposals according to the sizes of faces. Our method achieves competitive results on the widely used FDDB and WIDER face datasets, which demonstrates the effectiveness of our approach. Jifei Han, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
VCIP | 3 |
| 2018 | Context-Aware Local Binary Feature Learning for Face RecognitionabstractIn this paper, we propose a context-aware local binary feature learning (CA-LBFL) method for face recognition. Unlike existing learning-based local face descriptors such as discriminant face descriptor (DFD) and compact binary face descriptor (CBFD) which learn each feature code individually, our CA-LBFL exploits the contextual information of adjacent bits by constraining the number of shifts from different binary bits, so that more robust information can be exploited for face representation. Given a face image, we first extract pixel difference vectors (PDV) in local patches, and learn a discriminative mapping in an unsupervised manner to project each pixel difference vector into a context-aware binary vector. Then, we perform clustering on the learned binary codes to construct a codebook, and extract a histogram feature for each face image with the learned codebook as the final representation. In order to exploit local information from different scales, we propose a context-aware local binary multi-scale feature learning (CA-LBMFL) method to jointly learn multiple projection matrices for face representation. To make the proposed methods applicable for heterogeneous face recognition, we present a coupled CA-LBFL (C-CA-LBFL) method and a coupled CA-LBMFL (C-CA-LBMFL) method to reduce the modality gap of corresponding heterogeneous faces in the feature level, respectively. Extensive experimental results on four widely used face datasets clearly show that our methods outperform most state-of-the-art face descriptors. Yueqi Duan, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Two-Stream Transformer Networks for Video-Based Face AlignmentabstractIn this paper, we propose a two-stream transformer networks (TSTN) approach for video-based face alignment. Unlike conventional image-based face alignment approaches which cannot explicitly model the temporal dependency in videos and motivated by the fact that consistent movements of facial landmarks usually occur across consecutive frames, our TSTN aims to capture the complementary information of both the spatial appearance on still frames and the temporal consistency information across frames. To achieve this, we develop a two-stream architecture, which decomposes the video-based face alignment into spatial and temporal streams accordingly. Specifically, the spatial stream aims to transform the facial image to the landmark positions by preserving the holistic facial shape structure. Accordingly, the temporal stream encodes the video input as active appearance codes, where the temporal consistency information across frames is captured to help shape refinements. Experimental results on the benchmarking video-based face alignment datasets show very competitive performance of our method in comparisons to the state-of-the-arts. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Scene recognition with objectness
Xiaojuan Cheng, Jiwen Lu, Jianjiang Feng, Bo Yuan 0003, Jie Zhou 0001 |
Pattern Recognit. | 3 |
| 2018 | Reconstruction-based supervised hashing
Xin Yuan 0006, Zhixiang Chen 0003, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
Pattern Recognit. | 4 |
| 2018 | Deep convolutional neural network for latent fingerprint enhancement
Jianjiang Feng, C.-C. Jay Kuo |
Signal Process. Image Commun. | 2 |
| 2018 | Nonlinear Structural Hashing for Scalable Video SearchabstractIn this paper, we propose a nonlinear structural hashing approach to learn compact binary codes for scalable video search. Unlike most existing video hashing methods which consider image frames within a video separately for binary code learning, we develop a multi-layer neural network to learn compact and discriminative binary codes by exploiting both the structural information between different frames within a video and the nonlinear relationship between video samples. To be specific, we learn these binary codes under two different constraints at the output of our network: 1) the distance between the learned binary codes for frames within the same scene is minimized and 2) the distance between the learned binary matrices for a video pair with the same label is less than a threshold and that for a video pair with different labels is larger than a threshold. To better measure the structural information of the scenes from videos, we employ a subspace clustering method to cluster frames into different scenes. Moreover, we design multiple hierarchical nonlinear transformations to preserve the nonlinear relationship between videos. Experimental results on three video data sets show that our method outperforms state-of-the-art hashing approaches on the scalable video search task. Zhixiang Chen 0003, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Deep Localized Metric LearningabstractMetric learning has been widely used in many visual analysis applications, which learns new distance metrics to measure the similarities of samples effectively. Conventional metric learning methods learn a single linear Mahalanobis metric, yet such linear projections are not powerful enough to capture the nonlinear relationships. Recently, deep metric learning approaches, such as discriminative deep metric learning and deep transfer metric learning, have been introduced to fully exploit the nonlinearity of samples by learning hierarchical nonlinear transformations. However, these methods only learn holistic metrics over the input space and are limited for the heterogeneous data sets, where data varies locally. In this paper, we propose a deep localized metric learning approach for visual recognition by learning multiple fine-grained deep localized metrics. We first learn K local subspaces and one holistic subspace with the K-auto-encoders-based clustering. Then, given an input pair, we compute its localized distance on each learned subspace and obtain the final distance representation. Finally, we train the entire neural networks to ensure the distances of positive pairs smaller than negative pairs by a large margin. Experimental results on three visual recognition applications, including face recognition, person re-identification, and scene recognition, show that our DLML outperforms most existing metric learning approaches. Yueqi Duan, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | 2-D Phase Demodulation for Deformable Fingerprint RegistrationabstractFingerprint matching with elastic distortion is very challenging to deal with, and severe fingerprint distortion usually leads to false non-matches. This paper proposes a phase-based registration algorithm which can effectively eliminate the distortion between fingerprints and therefore is beneficial to the subsequent fingerprint matching. The key of the proposed algorithm is to reconstruct the distortion field through unwrapping phase difference between two fingerprints. Experiments on FVC2004, Tsinghua distorted fingerprint database, and NIST SD27 demonstrate that our algorithm outperforms other fingerprint registration methods and significantly improves matching accuracy. Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Topology Preserving Structural Matching for Automatic Partial Face RecognitionabstractIn this paper, we propose a topology preserving graph matching (TPGM) method for partial face recognition. Most existing face recognition methods extract features from holistic facial images. However, faces in real-world unconstrained environments may be occluded by objects or other faces, which cannot provide the whole face images for description. Keypoint-based partial face recognition methods such as multi-keypoint descriptor with Gabor ternary pattern and robust point set matching match the local keypoints for partial face recognition. However, they simply measure the nodewise similarity without higher order geometric graph information, which are susceptible to noises. To address this, our TPGM method estimates a non-rigid transformation encoding the second-order geometric structure of the graph, so that more accurate and robust correspondence can be computed with the topological information. In order to exploit higher order topological information, we propose a topology preserving structural matching method to construct a higher order structure for each face and estimate the transformation. Experimental results on four widely used face data sets demonstrate that our method outperforms most existing state-of-the-art face recognition methods. Yueqi Duan, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Efficient Rectification of Distorted FingerprintsabstractRecently, distortion rectification based on a single fingerprint image has been shown to be able to significantly improve the recognition rate of distorted fingerprints. However, the computational complexity of such a method is too high to be useful in practice. In this paper, we propose a novel method for the rectification of distorted fingerprints, whose speed is over 30 times faster than the existing method. This significant speedup is due to a Hough-forest-based two-step fingerprint pose estimation algorithm and a support vector regressor-based fingerprint distortion field estimation algorithm. Experimental results on public domain databases show that our method can achieve as good rectification performance as the existing method but meanwhile is significantly faster. Shan Gu, Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Label-Sensitive Deep Metric Learning for Facial Age EstimationabstractIn this paper, we present a label-sensitive deep metric learning (LSDML) approach for facial age estimation. Motivated by the fact that human age labels are chronologically correlated, our proposed LSDML aims to seek a series of hierarchical nonlinear transformations by deep residual network to project face samples to a latent common space, where the similarity of face pairs is equivalently isotonic to the age difference in a ranking-preserving manner. Since traversal access to total negative samples catastrophically costs and leads to suboptimal, our model learns to mine hard meaningful samples in parallel to learning feature similarity, so that the local manifold of face samples is preserved in the transformed subspace. To better improve the performance on the data set that contains few labeled samples, we further extend our LSDML to a multi-source LSDML method, which aims at maximizing the cross-population correlation of different face aging data sets. Extensive experimental results on four benchmarking data sets show the effectiveness of our proposed approach. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Left Atrial Appendage Segmentation Using Fully Convolutional Neural Networks and Modified Three-Dimensional Conditional Random FieldsabstractThrombosis has become a global disease threatening human health. The left atrial appendage (LAA) is a major source of thrombosis in patients with atrial fibrillation (AF). Positive correlation exists between LAA volume and AF risk. LAA morphology has been suggested to influence thromboembolic risk in AF patients and to help predict thromboembolic events in low-risk patient groups. Automatic segmentation of LAA can greatly help physicians diagnose AF. In consideration of the large anatomical variations of the LAA, we proposed a robust method for automatic LAA segmentation on computed tomographic angiography (CTA) data using fully convolutional neural networks with three-dimensional (3-D) conditional random fields (CRFs). After manual localization of ROI of LAA, we adopted the FCN in natural image segmentation and transferred their learned models by fine-tuning the networks to segment each 2-D LAA slice. Subsequently, we used a modified dense 3-D CRF that accounts for the 3-D spatial information and larger contextual information to refine the segmentations of all slices. Our method was evaluated on 150 sets of CTA data using five-fold cross validation. Compared with manual annotation, we obtained a mean dice overlap of and a mean volume overlap of with a computation time of less than 40 s per volume. Experimental results demonstrated the robustness of our method in dealing with large anatomical variations and computational efficiency for adoption in a daily clinical routine.). Cheng Jin 0007, Jianjiang Feng, Heng Yu 0005, Jiang Liu 0014, Jiwen Lu, Jie Zhou 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Learning Deep Binary Descriptor with Multi-quantizationabstractIn this paper, we propose an unsupervised feature learning method called deep binary descriptor with multi-quantization (DBD-MQ) for visual matching. Existing learning-based binary descriptors such as compact binary face descriptor (CBFD) and DeepBit utilize the rigid sign function for binarization despite of data distributions, thereby suffering from severe quantization loss. In order to address the limitation, our DBD-MQ considers the binarization as a multi-quantization task. Specifically, we apply a K-AutoEncoders (KAEs) network to jointly learn the parameters and the binarization functions under a deep learning framework, so that discriminative binary descriptors can be obtained with a fine-grained multi-quantization. Extensive experimental results on different visual analysis including patch retrieval, image matching and image retrieval show that our DBD-MQ outperforms most existing binary feature descriptors. Yueqi Duan, Jiwen Lu, Ziwei Wang 0010, Jianjiang Feng, Jie Zhou 0001 |
CVPR | 4 |
| 2017 | Consistent-Aware Deep Learning for Person Re-identification in a Camera NetworkabstractIn this paper, we propose a consistent-aware deep learning (CADL) framework for person re-identification in a camera network. Unlike most existing person re-identification methods which identify whether two body images are from the same person, our approach aims to obtain the maximal correct matches for the whole camera network. Different from recently proposed camera network based re-identification methods which only consider the consistent information in the matching stage to obtain a global optimal association, we exploit such consistent-aware information under a deep learning framework where both feature representation and image matching are automatically learned with certain consistent constraints. Specifically, we reach the global optimal solution and balance the performance between different cameras by optimizing the similarity and association iteratively. Experimental results show that our method obtains significant performance improvement and outperforms the state-of-the-art methods by large margins. Ji Lin 0002, Liangliang Ren, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
CVPR | 4 |
| 2017 | Ordinal Deep Feature Learning for Facial Age EstimationabstractIn this paper, we propose an ordinal deep feature learning (ODFL) approach for facial age estimation. Unlike conventional age estimation methods which utilize hand-crafted features, our ODFL develops deep convolutional neural networks to learn discriminative feature descriptors directly from image pixels for face representation. Motivated by the fact that age labels are chronologically correlated and age estimation is an ordinal learning computer vision problem, we enforce two criterions on the descriptors which are learned at the top of our network: 1) the topology-aware ordinal relation of face samples is preserved in the learned feature space, and 2) the age difference information of the embedded feature representation is exploited in a ranking-preserving manner. Extensive experimental results on four face aging datasets show that our approach achieves promising performance compared with the state-of-the-art methods. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
FG | 3 |
| 2017 | Fingerprint pose estimation based on faster R-CNNabstractFingerprint pose estimation is one of the bottlenecks of indexing in large scale database. The existing methods of pose estimation are based on manually appointed features (e.g. special points, ridges, orientation filed). In this paper, we propose a method based on deep learning to achieve accurate pose estimation. Faster R-CNN is adopted to detect the center point and rough direction, followed by intra-class and inter-class combination to calculate the precise direction. Extensive experiments on NIST-14 show that (1) the predicted poses are close to manual annotations even when the fingerprints are incomplete or noisy, (2) the estimated poses for matching fingerprint pairs are very consistent and (3) by registering fingerprints using the estimated pose, the accuracy of a state-of-the-art fingerprint indexing system is further improved. Jiahong Ouyang, Jianjiang Feng, Jiwen Lu, Zhenhua Guo 0001, Jie Zhou 0001 |
IJCB | 2 |
| 2017 | Localized multi-kernel discriminative canonical correlation analysis for video-based person re-identificationabstractThis paper presents a localized multi-kernel discriminative canonical correlation analysis (LMKDCCA) approach for video-based person re-identification, which aims to match persons from pedestrian videos captured by non-overlapping cameras. Unlike conventional methods, our approach models each pedestrian video as a point on the Riemannian manifold and learns similarity over these points under the multiple kernel learning framework. For each given person video, we first represent it as a symmetric positive definite (SPD) matrix which lies on a Riemannian manifold and compute the similarity of multiple SPDs. Then, we develop an LMKDCCA algorithm to learn a nonlinear distance metric which effectively combines these SPDs to exploit complementary information for similarity measure. Experimental results on the iLIDS-VID and PRID 2011 datasets show that our approach achieves the state-of-the-arts. Guangyi Chen 0002, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
ICIP | 3 |
| 2017 | Action recognition in RGB-D egocentric videosabstractIn this paper, we investigate the problem of action recognition in RGB-D egocentric videos. These self-generated and embodied videos provide richer semantic cues than the conventional videos captured from the third-person view for action recognition. Moreover, they contain both appearance information and 3D structure of the scenes from the RGB modality and depth modality respectively. Motivated by these advantages, we first collect a video-based RGB-D egocentric dataset (THU-READ) with diverse types of daily-life actions. Then we evaluate several approaches including hand-crafted features and deep learning methods on THU-READ. To improve the performance, we further develop a tri-stream convolutional network (TCNet) method, which learns to exploit the fuse with both the RGB and depth modalities for action recognition. Experimental results show that our model achieves competitive performance with state-of-the-art methods. Yansong Tang, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
ICIP | 4 |
| 2017 | Latent fingerprint enhancement using Gabor and minutia dictionariesabstractLatent fingerprints play important roles in law enforcement agencies. Due to its poor quality caused by unclear ridge structure, uneven contrast and overlapping patterns, a latent fingerprint enhancement is necessary for reliable feature extraction. Gabor function is widely used to characterise ridge structure and used in fingerprint enhancement. However, gabor function can not capture the details of minutia that is the end point or bifurcation of ridge. To utilize the prior knowledge of both ridge and minutia, we propose to construct both ridge and minutia dictionaries, and propose a two-step multi-scale patch based sparse representation to enhance the ridge using ridge dictionaries and enhance the minutia with both dictionaries. Experimental results show that two-step SR algorithm outperforms the SR only using gabor dictionary and gabor filter on both minutia extraction accuracy and matching accuracy. Jianjiang Feng, Jiwen Lu, Jie Zhou 0001 |
ICIP | 2 |
| 2017 | Topology preserving graph matching for partial face recognitionabstractIn this paper, we propose a topology preserving graph matching (TPGM) method for partial face recognition. Most existing face recognition methods extract features from holistic face images, yet faces in real-world unconstrained environments are usually occluded by objects or other faces, which cannot provide the whole face images for recognition. Latest keypoint-based partial face recognition methods only match on the detected keypoints to remove the occluded regions. However, they simply measure the node-wise similarity without higher order geometrical graph information, thereby depending heavily on descriptors which are susceptible to noises. To address this, our TPGM method estimates a non-rigid transformation encoding the second order geometric structure of the graph, so that more accurate and robust correspondence can be computed with the topological information. Experimental results on three widely used face datasets show that the proposed TPGM outperforms most existing state-of-the-art partial face recognition methods. Yueqi Duan, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
ICME | 3 |
| 2017 | Reconstruction-based supervised hashingabstractIn this paper, we propose a reconstruction-based supervised hashing (RSH) method to learn compact binary codes with holistic structure preservation for large scale image search. Unlike most existing hashing methods which consider pair-wise similarity, our method exploits the structural information of samples by employing a reconstruction-based criterion. Moreover, the label information of samples is also utilized to enhance the discriminative power of the teamed hash codes. Specifically, our method minimizes the distance between each point and the selected generated-structure with the same class label and maximizes the distance between each point and the selected generated-structure with different class labels. Experimental results on two widely used image datasets demonstrate the effectiveness of the proposed method. Xin Yuan 0006, Jiwen Lu, Zhixiang Chen 0003, Jianjiang Feng, Jie Zhou 0001 |
ICME | 4 |
| 2017 | Group-aware deep feature learning for facial age estimation
Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
Pattern Recognit. | 3 |
| 2017 | Multi-modal uniform deep learning for RGB-D person re-identification
Liangliang Ren, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
Pattern Recognit. | 3 |
| 2017 | Dense registration of fingerprints
Xuanbin Si, Jianjiang Feng, Bo Yuan 0003, Jie Zhou 0001 |
Pattern Recognit. | 2 |
| 2017 | Learning Rotation-Invariant Local Binary DescriptorabstractIn this paper, we propose a rotation-invariant local binary descriptor (RI-LBD) learning method for visual recognition. Compared with hand-crafted local binary descriptors, such as local binary pattern and its variants, which require strong prior knowledge, local binary feature learning methods are more efficient and data-adaptive. Unlike existing learning-based local binary descriptors, such as compact binary face descriptor and simultaneous local binary feature learning and encoding, which are susceptible to rotations, our RI-LBD first categorizes each local patch into a rotational binary pattern (RBP), and then jointly learns the orientation for each pattern and the projection matrix to obtain RI-LBDs. As all the rotation variants of a patch belong to the same RBP, they are rotated into the same orientation and projected into the same binary descriptor. Then, we construct a codebook by a clustering method on the learned binary codes, and obtain a histogram feature for each image as the final representation. In order to exploit higher order statistical information, we extend our RI-LBD to the triple rotation-invariant co-occurrence local binary descriptor (TRICo-LBD) learning method, which learns a triple co-occurrence binary code for each local patch. Extensive experimental results on four different visual recognition tasks, including image patch matching, texture classification, face recognition, and scene classification, show that our RI-LBD and TRICo-LBD outperform most existing local descriptors. Yueqi Duan, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Learning Deep Sharable and Structural Detectors for Face AlignmentabstractFace alignment aims at localizing multiple facial landmarks for a given facial image, which usually suffers from large variances of diverse facial expressions, aspect ratios and partial occlusions, especially when face images were captured in wild conditions. Conventional face alignment methods extract local features and then directly concatenate these features for global shape regression. Unlike these methods which cannot explicitly model the correlation of neighbouring landmarks and motivated by the fact that individual landmarks are usually correlated, we propose a deep sharable and structural detectors (DSSD) method for face alignment. To achieve this, we firstly develop a structural feature learning method to explicitly exploit the correlation of neighbouring landmarks, which learns to cover semantic information to disambiguate the neighbouring landmarks. Moreover, our model selectively learns a subset of sharable latent tasks across neighbouring landmarks under the paradigm of the multi-task learning framework, so that the redundancy information of the overlapped patches can be efficiently removed. To better improve the performance, we extend our DSSD to a recurrent DSSD (R-DSSD) architecture by integrating with the complementary information from multi-scale perspectives. Experimental results on the widely used benchmark datasets show that our methods achieve very competitive performance compared to the state-of-the-arts. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Nonlinear Discrete HashingabstractIn this paper, we propose a nonlinear discrete hashing approach to learn compact binary codes for scalable image search. Instead of seeking a single linear projection in most existing hashing methods, we pursue a multilayer network with nonlinear transformations to capture the local structure of data samples. Unlike most existing hashing methods that adopt an error-prone relaxation to learn the transformations, we directly solve the discrete optimization problem to eliminate the quantization error accumulation. Specifically, to leverage the similarity relationships between data samples and exploit the semantic affinities of manual labels, the binary codes are learned with the objective to: 1) minimize the quantization error between the original data samples and the learned binary codes; 2) preserve the similarity relationships in the learned binary codes; 3) maximize the information content with independent bits; and 4) maximize the accuracy of the predicted labels based on the binary codes. With an alternating optimization, the nonlinear transformation and the discrete quantization are jointly optimized in the hashing learning framework. Experimental results on four datasets including CIFAR10, MNIST, SUN397, and ILSVRC2012 demonstrate that the proposed approach is superior to several state-of-the-art hashing methods. Zhixiang Chen 0003, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | Nonlinear Sparse HashingabstractTo facilitate fast similarity search, this paper proposes to encode the nonlinear similarity and image structure as compact binary codes. Rather than adopting single matrix as projection in the literature, we employ a nonlinear transformation in the form of multilayer neural network to generate binary codes to capture the local structure between data samples. Specifically, we train the network such that the quantization loss is minimized and the variance over all bits is maximized. In addition, we capture the salient structure of image samples at the abstract level with sparsity constraint and inherit the generalization power to unseen samples. Furthermore, we incorporate the supervisory label information into the learning procedure to take advantage of the manual label. To obtain the desired binary codes and the parameterized nonlinear transformation, we optimize the formulated objective problem over each variable with an iterative alternating method. To validate the efficacy of the proposed hashing approach, we conduct experiments on three widely used datasets, namely CIFAR10, MNIST, and SUN397, by comparing with several recent proposed hashing methods. Zhixiang Chen 0003, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Fingerprint indexing with pose constraint
Yijing Su, Jianjiang Feng, Jie Zhou 0001 |
Pattern Recognit. | 2 |
| 2016 | Depth Estimation Using a Sliding CameraabstractImage-based 3D reconstruction technology is widely used in different fields. The conventional algorithms are mainly based on stereo matching between two or more fixed cameras, and high accuracy can only be achieved using a large camera array, which is very expensive and inconvenient in many applications. Another popular choice is utilizing structure-from-motion methods for arbitrarily placed camera(s). However, due to too many degrees of freedom, its computational cost is heavy and its accuracy is rather limited. In this paper, we propose a novel depth estimation algorithm using a sliding camera system. By analyzing the geometric properties of the camera system, we design a camera pose initialization algorithm that can work satisfyingly with only a small number of feature points and is robust to noise. For pixels corresponding to different depths, an adaptive iterative algorithm is proposed to choose optimal frames for stereo matching, which can take advantage of continuously pose-changing imaging and save the time consumption amazingly too. The proposed algorithm can also be easily extended to handle less constrained situations (such as using a camera mounted on a moving robot or vehicle). Experimental results on both synthetic and real-world data have illustrated the effectiveness of the proposed algorithm. Kailin Ge, Han Hu 0001, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Dense and continuous depth estimation using a sliding cameraabstract3D information of real-world scenes provides important clues for many computer vision tasks. We present a simple but effective sliding camera system as well as a corresponding stereo reconstruction framework to retrieve 3D information of static scenes. By fusing geometric properties of the sliding camera system, our reconstruction algorithm achieves higher accuracy than conventional methods in quantitative experiments. Besides, the practicality of our system is validated on real world scenes. Kailin Ge, Jianjiang Feng, Jie Zhou 0001 |
ICASSP | 2 |
| 2015 | Multiple Feature Fusion via Weighted Entropy for Visual TrackingabstractIt is desirable to combine multiple feature descriptors to improve the visual tracking performance because different features can provide complementary information to describe objects of interest. However, how to effectively fuse multiple features remains a challenging problem in visual tracking, especially in a data-driven manner. In this paper, we propose a new data-adaptive visual tracking approach by using multiple feature fusion via weighted entropy. Unlike existing visual trackers which simply concatenate multiple feature vectors together for object representation, we employ the weighted entropy to evaluate the dissimilarity between the object state and the background state, and seek the optimal feature combination by minimizing the weighted entropy, so that more complementary information can be exploited for object representation. Experimental results demonstrate the effectiveness of our approach in tackling various challenges for visual object tracking. Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
ICCV | 3 |
| 2015 | Exploiting Unsupervised and Supervised Constraints for Subspace ClusteringabstractData in many image and video analysis tasks can be viewed as points drawn from multiple low-dimensional subspaces with each subspace corresponding to one category or class. One basic task for processing such kind of data is to separate the points according to the underlying subspace, referred to as subspace clustering. Extensive studies have been made on this subject, and nearly all of them use unconstrained subspace models, meaning the points can be drawn from everywhere of a subspace, to represent the data. In this paper, we attempt to do subspace clustering based on a constrained subspace assumption that the data is further restricted in the corresponding subspaces, e.g., belonging to a submanifold or satisfying the spatial regularity constraint. This assumption usually describes the real data better, such as differently moving objects in a video scene and face images of different subjects under varying illumination. A unified integer linear programming optimization framework is used to approach subspace clustering, which can be efficiently solved by a branch-and-bound (BB) method. We also show that various kinds of supervised information, such as subspace number, outlier ratio, pairwise constraints, size prior and etc., can be conveniently incorporated into the proposed framework. Experiments on real data show that the proposed method outperforms the state-of-the-art algorithms significantly in clustering accuracy. The effectiveness of the proposed method in exploiting supervised information is also demonstrated. Han Hu 0001, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Detection and Rectification of Distorted FingerprintsabstractElastic distortion of fingerprints is one of the major causes for false non-match. While this problem affects all fingerprint recognition applications, it is especially dangerous in negative recognition applications, such as watchlist and deduplication applications. In such applications, malicious users may purposely distort their fingerprints to evade identification. In this paper, we proposed novel algorithms to detect and rectify skin distortion based on a single fingerprint image. Distortion detection is viewed as a two-class classification problem, for which the registered ridge orientation map and period map of a fingerprint are used as the feature vector and a SVM classifier is trained to perform the classification task. Distortion rectification (or equivalently distortion field estimation) is viewed as a regression problem, where the input is a distorted fingerprint and the output is the distortion field. To solve this problem, a database (called reference database) of various distorted reference fingerprints and corresponding distortion fields is built in the offline stage, and then in the online stage, the nearest neighbor of the input fingerprint is found in the reference database and the corresponding distortion field is used to transform the input fingerprint into a normal one. Promising results have been obtained on three databases containing many distorted fingerprints, namely FVC2004 DB1, Tsinghua Distorted Fingerprint database, and the NIST SD27 latent fingerprint database. Xuanbin Si, Jianjiang Feng, Jie Zhou 0001, Yuxuan Luo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Smooth Representation ClusteringabstractSubspace clustering is a powerful technology for clustering data according to the underlying subspaces. Representation based methods are the most popular subspace clustering approach in recent years. In this paper, we analyze the grouping effect of representation based methods in depth. In particular, we introduce the enforced grouping effect conditions, which greatly facilitate the analysis of grouping effect. We further find that grouping effect is important for subspace clustering, which should be explicitly enforced in the data self-representation model, rather than implicitly implied by the model as in some prior work. Based on our analysis, we propose the SMooth Representation (SMR) model. We also propose a new affinity measure based on the grouping effect, which proves to be much more effective than the commonly used one. As a result, our SMR significantly outperforms the state-of-the-art ones on benchmark datasets. Han Hu 0001, Zhouchen Lin, Jianjiang Feng, Jie Zhou 0001 |
CVPR | 3 |
| 2014 | Fingerprint matching based on global minutia cylinder codeabstractAlthough minutia set based fingerprint matching algorithms have achieved good matching accuracy, developing a fingerprint recognition system that satisfies accuracy, efficiency and privacy requirements simultaneously remains a challenging problem. Fixed-length binary vector like IrisCode is considered to be an ideal representation to meet these requirements. However, existing fixed-length vector representations of fingerprints suffered from either low distinctiveness or misalignment problem. In this paper, we propose a discriminative fixed-length binary representation of fingerprints based on an extension of Minutia Cylinder Code. A machine learning based algorithm is proposed to mine reliable reference points to overcome the misalignment problem. Experimental results on public domain plain and rolled fingerprint databases demonstrate the effectiveness of the proposed approach. Yuxuan Luo 0002, Jianjiang Feng, Jie Zhou 0001 |
IJCB | 2 |
| 2014 | Enhancing latent fingerprints on banknotesabstractMatching unknown latent fingerprints lifted from various objects or surfaces at crime scenes to fingerprints of known subjects is of vital importance for law enforcement agencies to identify suspects. Banknotes are one of the most common objects containing valuable latent fingerprints. However, due to the complex pattern printed on banknotes, it is a challenging problem even for human experts to mark minutiae in such fingerprints. In this paper a novel technique is proposed to enhance fingerprints on banknotes so that they can be successfully identified by existing fingerprint matchers. The proposed algorithm is based on subtraction of the reference orientation in reference banknote, which is registered to the latent fingerprint by a coarse-to-fine registration algorithm. Promising results are reported on a database which contain 192 latents on banknotes, which proves the effectiveness of the proposed algorithm. Xuanbin Si, Jianjiang Feng, Jie Zhou 0001 |
IJCB | 2 |
| 2014 | Localized Dictionaries Based Orientation Field Estimation for Latent FingerprintsabstractDictionary based orientation field estimation approach has shown promising performance for latent fingerprints. In this paper, we seek to exploit stronger prior knowledge of fingerprints in order to further improve the performance. Realizing that ridge orientations at different locations of fingerprints have different characteristics, we propose a localized dictionaries-based orientation field estimation algorithm, in which noisy orientation patch at a location output by a local estimation approach is replaced by real orientation patch in the local dictionary at the same location. The precondition of applying localized dictionaries is that the pose of the latent fingerprint needs to be estimated. We propose a Hough transform-based fingerprint pose estimation algorithm, in which the predictions about fingerprint pose made by all orientation patches in the latent fingerprint are accumulated. Experimental results on challenging latent fingerprint datasets show the proposed method outperforms previous ones markedly. Xiao Yang 0029, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | OSRI: A Rotationally Invariant Binary DescriptorabstractBinary descriptors are becoming widely used in computer vision field because of their high matching efficiency and low memory requirements. Since conventional approaches, which first compute a floating-point descriptor then binarize it, are computationally expensive, some recent efforts have focused on directly computing binary descriptors from local image patches. Although these binary descriptors enable a significant speedup in processing time, their performances usually drop a lot due to orientation estimation errors and limited description abilities. To address these issues, we propose a novel binary descriptor based on the ordinal and spatial information of regional invariants (OSRIs) over a rotation invariant sampling pattern. Our main contributions are twofold: 1) each bit in OSRI is computed based on difference tests of regional invariants over pairwise sampling-regions instead of difference tests of pixel intensities commonly used in existing binary descriptors, which can significantly enhance the discriminative ability and 2) rotation and illumination changes are handled well by ordering pixels according to their intensities and gradient orientations, meanwhile, which is also more reliable than those methods that resort to a reference orientation for rotation invariance. Besides, a statistical analysis of discriminative abilities of different parts in the descriptor is conducted to design a cascade filter which can reject nonmatching descriptors at early stages by comparing just a small portion of the whole descriptor, further reducing the matching time. Extensive experiments on four challenging data sets (Oxford, 53 Objects, ZuBuD, and Kentucky) show that OSRI significantly outperforms two state-of-the-art binary descriptors (FREAK and ORB). The matching performance of OSRI with only 512 bits is also better than the well-known floating-point descriptor SIFT (4K bits) and is comparable with the state-of-the-art floating-point descriptor MROGH (6K bits), while it is two orders of magnitude faster to match than SIFT and MROGH. Xianwei Xu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | Orientation Field Estimation for Latent Fingerprint EnhancementabstractIdentifying latent fingerprints is of vital importance for law enforcement agencies to apprehend criminals and terrorists. Compared to live-scan and inked fingerprints, the image quality of latent fingerprints is much lower, with complex image background, unclear ridge structure, and even overlapping patterns. A robust orientation field estimation algorithm is indispensable for enhancing and recognizing poor quality latents. However, conventional orientation field estimation algorithms, which can satisfactorily process most live-scan and inked fingerprints, do not provide acceptable results for most latents. We believe that a major limitation of conventional algorithms is that they do not utilize prior knowledge of the ridge structure in fingerprints. Inspired by spelling correction techniques in natural language processing, we propose a novel fingerprint orientation field estimation algorithm based on prior knowledge of fingerprint structure. We represent prior knowledge of fingerprints using a dictionary of reference orientation patches. which is constructed using a set of true orientation fields, and the compatibility constraint between neighboring orientation patches. Orientation field estimation for latents is posed as an energy minimization problem, which is solved by loopy belief propagation. Experimental results on the challenging NIST SD27 latent fingerprint database and an overlapped latent fingerprint database demonstrate the advantages of the proposed orientation field estimation algorithm over conventional algorithms. Jianjiang Feng, Jie Zhou 0001, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Latent Fingerprint Matching Using Descriptor-Based Hough TransformabstractIdentifying suspects based on impressions of fingers lifted from crime scenes (latent prints) is a routine procedure that is extremely important to forensics and law enforcement agencies. Latents are partial fingerprints that are usually smudgy, with small area and containing large distortion. Due to these characteristics, latents have a significantly smaller number of minutiae points compared to full (rolled or plain) fingerprints. The small number of minutiae and the noise characteristic of latents make it extremely difficult to automatically match latents to their mated full prints that are stored in law enforcement databases. Although a number of algorithms for matching full-to-full fingerprints have been published in the literature, they do not perform well on the latent-to-full matching problem. Further, they often rely on features that are not easy to extract from poor quality latents. In this paper, we propose a new fingerprint matching algorithm which is especially designed for matching latents. The proposed algorithm uses a robust alignment algorithm (descriptor-based Hough transform) to align fingerprints and measures similarity between fingerprints by considering both minutiae and orientation field information. To be consistent with the common practice in latent matching (i.e., only minutiae are marked by latent examiners), the orientation field is reconstructed from minutiae. Since the proposed algorithm relies only on manually marked minutiae, it can be easily used in law enforcement applications. Experimental results on two different latent databases (NIST SD27 and WVU latent databases) show that the proposed algorithm outperforms two well optimized commercial fingerprint matchers. Further, a fusion of the proposed algorithm and commercial fingerprint matchers leads to improved matching accuracy. Alessandra A. Paulino, Jianjiang Feng, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Multi-Class Constrained Normalized Cut With Hard, Soft, Unary and Pairwise Priors and its Applications to Object SegmentationabstractNormalized cut is a powerful method for image segmentation as well as data clustering. However, it does not perform well in challenging segmentation problems, such as segmenting objects in a complex background. Researchers have attempted to incorporate priors or constraints to handle such cases. Available priors in image segmentation problems may be hard or soft, unary or pairwise, but only hard must-link constraints and two-class settings are well studied. The main difficulties may lie in the following aspects: 1) the nontransitive nature of cannot-link constraints makes it hard to use such constraints in multi-class settings and 2) in multi-class or pairwise settings, the output labels have inconsistent representations with given priors, making soft priors difficult to use. In this paper, we propose novel algorithms, which can handle both hard and soft, both unary and pairwise priors in multi-class settings and provide closed form and efficient solutions. We also apply the proposed algorithms to the problem of object segmentation, producing good results by further introducing a spatial regularity term. Experiments show that the proposed algorithms outperform the state-of-the-art algorithms significantly in clustering accuracy. Other merits of the proposed algorithms are also demonstrated. Han Hu 0001, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Mining sub-categories for object detection
Jifeng Dai, Jianjiang Feng, Jie Zhou 0001 |
ICPR | 2 |
| 2012 | Multi-way constrained spectral clustering by nonnegative restriction
Han Hu 0001, Jiahuan Zhou, Jianjiang Feng, Jie Zhou 0001 |
ICPR | 3 |
| 2012 | Robust and Efficient Ridge-Based Palmprint MatchingabstractDuring the past decade, many efforts have been made to use palmprints as a biometric modality. However, most of the existing palmprint recognition systems are based on encoding and matching creases, which are not as reliable as ridges. This affects the use of palmprints in large-scale person identification applications where the biometric modality needs to be distinctive as well as insensitive to changes in age and skin conditions. Recently, several ridge-based palmprint matching algorithms have been proposed to fill the gap. Major contributions of these systems include reliable orientation field estimation in the presence of creases and the use of multiple features in matching, while the matching algorithms adopted in these systems simply follow the matching algorithms for fingerprints. However, palmprints differ from fingerprints in several aspects: 1) Palmprints are much larger and thus contain a large number of minutiae, 2) palms are more deformable than fingertips, and 3) the quality and discrimination power of different regions in palmprints vary significantly. As a result, these matchers are unable to appropriately handle the distortion and noise, despite heavy computational cost. Motivated by the matching strategies of human palmprint experts, we developed a novel palmprint recognition system. The main contributions are as follows: 1) Statistics of major features in palmprints are quantitatively studied, 2) a segment-based matching and fusion algorithm is proposed to deal with the skin distortion and the varying discrimination power of different palmprint regions, and 3) to reduce the computational complexity, an orientation field-based registration algorithm is designed for registering the palmprints into the same coordinate system before matching and a cascade filter is built to reject the nonmated gallery palmprints in early stage. The proposed matcher is tested by matching 840 query palmprints against a gallery set of 13,736 palmprints. Experimental results show that the proposed matcher outperforms the existing matchers a lot both in matching accuracy and speed. Jifeng Dai, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Altered Fingerprints: Analysis and DetectionabstractThe widespread deployment of Automated Fingerprint Identification Systems (AFIS) in law enforcement and border control applications has heightened the need for ensuring that these systems are not compromised. While several issues related to fingerprint system security have been investigated, including the use of fake fingerprints for masquerading identity, the problem of fingerprint alteration or obfuscation has received very little attention. Fingerprint obfuscation refers to the deliberate alteration of the fingerprint pattern by an individual for the purpose of masking his identity. Several cases of fingerprint obfuscation have been reported in the press. Fingerprint image quality assessment software (e.g., NFIQ) cannot always detect altered fingerprints since the implicit image quality due to alteration may not change significantly. The main contributions of this paper are: 1) compiling case studies of incidents where individuals were found to have altered their fingerprints for circumventing AFIS, 2) investigating the impact of fingerprint alteration on the accuracy of a commercial fingerprint matcher, 3) classifying the alterations into three major categories and suggesting possible countermeasures, 4) developing a technique to automatically detect altered fingerprints based on analyzing orientation field and minutiae distribution, and 5) evaluating the proposed technique and the NFIQ algorithm on a large database of altered fingerprints provided by a law enforcement agency. Experimental results show the feasibility of the proposed approach in detecting altered fingerprints and highlight the need to further pursue this problem. Soweon Yoon, Jianjiang Feng, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Robust and Efficient Algorithms for Separating Latent Overlapped FingerprintsabstractOverlapped fingerprints are frequently encountered in latent fingerprints lifted from crime scenes. It is necessary to separate such overlapped fingerprints into component fingerprints so that existing fingerprint matchers can recognize them. The most crucial step in separating overlapped fingerprints is estimation of component orientation fields, which is a challenging problem for existing orientation field estimation algorithms. We propose a robust orientation field estimation algorithm (called the basic algorithm) for latent overlapped fingerprints whose core is the constrained relaxation labeling algorithm. We also propose improved versions of the basic algorithm for two special but frequent cases: 1) the mated template fingerprint of one component fingerprint is known and 2) the two component fingerprints are from the same finger. In both cases, further constraints are used to reduce ambiguity in relaxation labeling. Experimental results on both real and simulated overlapped fingerprints show that the proposed algorithm outperforms the state-of-the-art algorithm in both accuracy and efficiency. The two improved versions also perform better than the basic algorithm in respective cases. The latent overlapped fingerprint database collected for this study is made publicly available for performance evaluation. Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2011 | Latent fingerprint matching using descriptor-based hough transformabstractIdentifying suspects based on impressions of fingers lifted from crime scenes (latent prints) is extremely important to law enforcement agencies. Latents are usually partial fingerprints with small area, contain nonlinear distortion, and are usually smudgy and blurred. Due to some of these characteristics, they have a significantly smaller number of minutiae points (one of the most important features in fingerprint matching) and therefore it can be extremely difficult to automatically match latents to plain or rolled fingerprints that are stored in law enforcement databases. Our goal is to develop a latent matching algorithm that uses only minutiae information. The proposed approach consists of following three modules: (i) align two sets of minutiae by using a descriptor-based Hough Transform; (ii) establish the correspondences between minutiae; and (iii) compute a similarity score. Experimental results on NIST SD27 show that the proposed algorithm outperforms a commercial fingerprint matcher. Alessandra A. Paulino, Jianjiang Feng, Anil K. Jain 0001 |
IJCB | 2 |
| 2011 | Palmprint indexing based on ridge featuresabstractIn recent years, law enforcement agencies are increasingly using palmprint to identify criminals. For law enforcement palmprint identification systems, efficiency is a very important but challenging problem because of large database size and poor image quality. Existing palmprint identification systems are not sufficiently fast for practical applications. To solve this problem, a novel palmprint indexing algorithm based on ridge features is proposed in this paper. A palmprint is pre-aligned by registering its orientation field with respect to a set of reference orientation fields, which are obtained by clustering training palmprint orientation fields. Indexing is based on comparing ridge orientation fields and ridge density maps, which is much faster than minutiae matching. Proposed algorithm achieved an error rate of 1% at a penetration rate of 2.25% on a palm print database consisting of 13,416 palmprints. Searching a query palmprint over the whole database takes only 0.22 seconds. Xiao Yang 0029, Jianjiang Feng, Jie Zhou 0001 |
IJCB | 2 |
| 2011 | Latent fingerprint enhancement via robust orientation field estimationabstractLatent fingerprints, or simply latents, have been considered as cardinal evidence for identifying and convicting criminals. The amount of information available for identification from latents is often limited due to their poor quality, unclear ridge structure and occlusion with complex back ground or even other latent prints. We propose a latent fingerprint enhancement algorithm, which expects manually marked region of interest (ROI) and singular points. The core of the proposed algorithm is a robust orientation field estimation algorithm for latents. Short-time Fourier transform is used to obtain multiple orientation elements in each image block. This is followed by a hypothesize-and test paradigm based on randomized RANSAC, which generates a set of hypothesized orientation fields. Experimental results on NIST SD27 latent fingerprint database show that the matching performance of a commercial matcher is significantly improved by utilizing the enhanced latent finger prints produced by the proposed algorithm. Soweon Yoon, Jianjiang Feng, Anil K. Jain 0001 |
IJCB | 2 |
| 2011 | Fingerprint Reconstruction: From Minutiae to PhaseabstractFingerprint matching systems generally use four types of representation schemes: grayscale image, phase image, skeleton image, and minutiae, among which minutiae-based representation is the most widely adopted one. The compactness of minutiae representation has created an impression that the minutiae template does not contain sufficient information to allow the reconstruction of the original grayscale fingerprint image. This belief has now been shown to be false; several algorithms have been proposed that can reconstruct fingerprint images from minutiae templates. These techniques try to either reconstruct the skeleton image, which is then converted into the grayscale image, or reconstruct the grayscale image directly from the minutiae template. However, they have a common drawback: Many spurious minutiae not included in the original minutiae template are generated in the reconstructed image. Moreover, some of these reconstruction techniques can only generate a partial fingerprint. In this paper, a novel fingerprint reconstruction algorithm is proposed to reconstruct the phase image, which is then converted into the grayscale image. The proposed reconstruction algorithm not only gives the whole fingerprint, but the reconstructed fingerprint contains very few spurious minutiae. Specifically, a fingerprint image is represented as a phase image which consists of the continuous phase and the spiral phase (which corresponds to minutiae). An algorithm is proposed to reconstruct the continuous phase from minutiae. The proposed reconstruction algorithm has been evaluated with respect to the success rates of type-I attack (match the reconstructed fingerprint against the original fingerprint) and type-II attack (match the reconstructed fingerprint against different impressions of the original fingerprint) using a commercial fingerprint recognition system. Given the reconstructed image from our algorithm, we show that both types of attacks can be successfully launched against a fingerprint recognition system. Jianjiang Feng, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Latent Fingerprint MatchingabstractLatent fingerprint identification is of critical importance to law enforcement agencies in identifying suspects: Latent fingerprints are inadvertent impressions left by fingers on surfaces of objects. While tremendous progress has been made in plain and rolled fingerprint matching, latent fingerprint matching continues to be a difficult problem. Poor quality of ridge impressions, small finger area, and large nonlinear distortion are the main difficulties in latent fingerprint matching compared to plain or rolled fingerprint matching. We propose a system for matching latent fingerprints found at crime scenes to rolled fingerprints enrolled in law enforcement databases. In addition to minutiae, we also use extended features, including singularity, ridge quality map, ridge flow map, ridge wavelength map, and skeleton. We tested our system by matching 258 latents in the NIST SD27 database against a background database of 29,257 rolled fingerprints obtained by combining the NIST SD4, SD14, and SD27 databases. The minutiae-based baseline rank-1 identification rate of 34.9 percent was improved to 74 percent when extended features were used. In order to evaluate the relative importance of each extended feature, these features were incrementally used in the order of their cost in marking by latent experts. The experimental results indicate that singularity, ridge quality map, and ridge flow map are the most effective features in improving the matching accuracy. Anil K. Jain 0001, Jianjiang Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Separating Overlapped FingerprintsabstractFingerprint images generally contain either a single fingerprint (e.g., rolled images) or a set of nonoverlapped fingerprints (e.g., slap fingerprints). However, there are situations where several fingerprints overlap on top of each other. Such situations are frequently encountered when latent (partial) fingerprints are lifted from crime scenes or residue fingerprints are left on fingerprint sensors. Overlapped fingerprints constitute a serious challenge to existing fingerprint recognition algorithms, since these algorithms are designed under the assumption that fingerprints have been properly segmented. In this paper, a novel algorithm is proposed to separate overlapped fingerprints into component or individual fingerprints. The basic idea is to first estimate the orientation field of the given image with overlapped fingerprints and then separate it into component orientation fields using a relaxation labeling technique. We also propose an algorithm to utilize fingerprint singularity information to further improve the separation performance. Experimental results indicate that the algorithm leads to good separation of overlapped fingerprints that leads to a significant improvement in the matching accuracy. Fanglin Chen 0001, Jianjiang Feng, Anil K. Jain 0001, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2010 | Detecting Altered FingerprintsabstractThe widespread deployment of Automated Fingerprint Identification Systems (AFIS) in law enforcement and border control applications has prompted some individuals with criminal background to evade identification by purposely altering their fingerprints. Available fingerprint quality assessment software cannot detect most of the altered fingerprints since the implicit image quality does not always degrade due to alteration. In this paper, we classify the alterations observed in an operational database into three categories and propose an algorithm to detect altered fingerprints. Experiments were conducted on both real-world altered fingerprints and synthetically generated altered fingerprints. At a false alarm rate of 7%, the proposed algorithm detected 92% of the altered fingerprints, while a well-known fingerprint quality software, NFIQ, only detected 20% of the altered fingerprints. Jianjiang Feng, Anil K. Jain 0001, Arun Ross |
ICPR | 1 |
| 2009 | Latent Palmprint MatchingabstractThe evidential value of palmprints in forensic applications is clear as about 30 percent of the latents recovered from crime scenes are from palms. While biometric systems for palmprint-based personal authentication in access control type of applications have been developed, they mostly deal with low-resolution (about 100 ppi) palmprints and only perform full-to-full palmprint matching. We propose a latent-to-full palmprint matching system that is needed in forensic applications. Our system deals with palmprints captured at 500 ppi (the current standard in forensic applications) or higher resolution and uses minutiae as features to be compatible with the methodology used by latent experts. Latent palmprint matching is a challenging problem because latent prints lifted at crime scenes are of poor image quality, cover only a small area of the palm, and have a complex background. Other difficulties include a large number of minutiae in full prints (about 10 times as many as fingerprints), and the presence of many creases in latents and full prints. A robust algorithm to reliably estimate the local ridge direction and frequency in palmprints is developed. This facilitates the extraction of ridge and minutiae features even in poor quality palmprints. A fixed-length minutia descriptor, MinutiaCode, is utilized to capture distinctive information around each minutia and an alignment-based minutiae matching algorithm is used to match two palmprints. Two sets of partial palmprints (150 live-scan partial palmprints and 100 latent palmprints) are matched to a background database of 10,200 full palmprints to test the proposed system. Despite the inherent difficulty of latent-to-full palmprint matching, rank-1 recognition rates of 78.7 and 69 percent, respectively, were achieved in searching live-scan partial palmprints and latent palmprints against the background database. Anil K. Jain 0001, Jianjiang Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Filtering large fingerprint database for latent matchingabstractLatent fingerprint identification is of critical importance to law enforcement agencies in apprehending criminals. Considering the huge size of fingerprint databases maintained by law enforcement agencies, exhaustive one-to-one matching is impractical and a database filtering technique is necessary to reduce the search space. Due to low image quality and small finger area of latent fingerprints, it is necessary to use several features for an efficient and reliable filtering system. A multi-stage filtering system is proposed, which utilizes pattern type, singular points and orientation field. We have tested our system by searching 258 latent fingerprints in NIST SD27 against a background database containing 10,258 rolled fingerprints (obtained by combining 2,000 in NIST SD4, 8,000 in SD14 and 258 in SD27). Although latent fingerprints contain very limited information, the filtering system not only improved the matching speed by three fold but also improved the rank-1 matching accuracy from 70.9% to 73.3%. Jianjiang Feng, Anil K. Jain 0001 |
ICPR | 1 |
| 2008 | Combining minutiae descriptors for fingerprint matching
Jianjiang Feng |
Pattern Recognit. | 1 |
| 2006 | Fingerprint matching using ridges
Jianjiang Feng, Zhengyu Ouyang, Anni Cai |
Pattern Recognit. | 1 |