Haibin Yan

dblp:27/9964 · DBLP profile ↗
← Back
45ranked-venue papers
17as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 15 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 10 since 2021Systems, architecture and hardware · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Spherical Projected Bézier Flow: A Geometry-Constrained Manifold Transport Framework for Cross-Age Face Retrieval
abstract
Cross-age face retrieval—often formulated as searching a large-scale gallery for the same identity across significant age gaps—remains a fundamental challenge due to the geometric mismatch between static embedding spaces and dynamic biological aging. In real-world retrieval and forensic deployments, preserving the pre-trained identity metric is crucial, yet many existing approaches either synthesize pixels with artifacts or fine-tune backbones, potentially distorting the cosine-based embedding geometry. In this paper, we propose Spherical Projected Bézier Flow (SPBF), a geometry-constrained manifold transport framework that models aging as feature transport on the hypersphere while keeping the backbone frozen. SPBF parameterizes a curvilinear trajectory via a projected Bézier path and learns a tangent velocity field with an Endpoint Consistency Constraint, enabling flexible non-linear and variable-speed dynamics without leaving the unit sphere. By avoiding off-manifold, low-norm states associated with higher uncertainty, SPBF improves retrieval robustness under large age gaps. Extensive experiments on FG-NET, AgeDB, and CACD show that SPBF is competitive with SOTA on homogeneous benchmarks and delivers substantial gains in cross-dataset generalization with only a lightweight plug-in module.
Huaqing Song, Baichuan Lin, Shuofeng Sun, Lanchi Xie, Haibin Yan
ICMR5
2026 Scale-aware modulation network for unsupervised medical anomaly detection
Haijie Cao, Shuofeng Sun, Haibin Yan
Neurocomputing3
2026 Generative age-aware data augmentation for cross-generation kinship verification
Shuofeng Sun, Linqing Zhao, Haibin Yan
Pattern Recognit. Lett.5
2025 MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
abstract
Mobile manipulation is the fundamental challenge for robotics to assist humans with diverse tasks and environments in everyday life. However, conventional mobile manipulation approaches often struggle to generalize across different tasks and environments because of the lack of large-scale training. In contrast, recent advances in vision-language-action (VLA) models have shown impressive generalization capabilities, but these foundation models are developed for fixed-base manipulation tasks. Therefore, we propose an efficient policy adaptation framework named MoManipVLA to transfer pre-trained VLA models of fix-base manipulation to mobile manipulation, so that high generalization ability across tasks and environments can be achieved in mobile manipulation policy. Specifically, we utilize pre-trained VLA models to generate waypoints of the end-effector with high generalization ability. We design motion planning objectives for the mobile base and the robot arm, which aim at maximizing the physical feasibility of the trajectory. Finally, we present an efficient bi-level objective optimization framework for trajectory generation, where the upper-level optimization predicts way-points for base movement to enhance the manipulator policy space, and the lower-level optimization selects the optimal end-effector trajectory to complete the manipulation task. Extensive experimental results on OVMM and the real world demonstrate that MoManipVLA achieves a 4.2% higher success rate than the state-of-the-art mobile manipulation, and only requires 50 training cost for real world deployment due to the strong generalization ability in the pre-trained VLA models. Our project page can be found here.
Xiuwei Xu, Ziwei Wang 0001, Haibin Yan
CVPR5
2025 Mitigating Geometric Degradation in Fast DownSampling via FastAdapter for Point Cloud Segmentation
Shuofeng Sun, Haibin Yan
ICCV2
2025 iGaussian: Real-Time Camera Pose Estimation via Feed-Forward 3D Gaussian Splatting Inversion
abstract
Recent trends in SLAM and visual navigation have embraced 3D Gaussians as the preferred scene representation, highlighting the importance of estimating camera poses from a single image using a pre-built Gaussian model. However, existing approaches typically rely on an iterative render-compare-refine loop, where candidate views are first rendered using NeRF or Gaussian Splatting, then compared against the target image, and finally, discrepancies are used to update the pose. This multi-round process incurs significant computational overhead, hindering real-time performance in robotics. In this paper, we propose iGaussian, a two-stage feed-forward framework that achieves real-time camera pose estimation through direct 3D Gaussian inversion. Our method first regresses a coarse 6DoF pose using a Gaussian Scene Prior-based Pose Regression Network with spatial uniform sampling and guided attention mechanisms, then refines it through feature matching and multi-model fusion. The key contribution lies in our cross-correlation module that aligns image embeddings with 3D Gaussian attributes without differentiable rendering, coupled with a Weighted Multiview Predictor that fuses features from Multiple strategically sampled viewpoints. Experimental results on the NeRF Synthetic, Mip-NeRF 360, and T&T+DB datasets demonstrate a significant performance improvement over previous methods, reducing median rotation errors to 0.2° while achieving 2.87 FPS tracking on mobile robots, which is an impressive 10× speedup compared to optimization-based approaches. Project page: https://github.com/pythongod-exe/iGaussian
Linqing Zhao, Xiuwei Xu, Jiwen Lu, Haibin Yan
IROS5
2025 Embodied Instruction Following in Unknown Environments
abstract
Enabling embodied agents to complete complex human instructions from natural language is crucial to autonomous systems in household services. Conventional methods can only accomplish human instructions in the known environment where all interactive objects are provided to the embodied agent, and directly deploying the existing approaches for the unknown environment usually generates infeasible plans that manipulate non-existing objects. On the contrary, we propose an embodied instruction following (EIF) method for complex tasks in the unknown environment, where the agent efficiently explores the unknown environment to generate feasible plans with existing objects to accomplish abstract instructions. Specifically, we build a hierarchical embodied instruction following framework including the high-level task planner and the low-level exploration controller with multimodal large language models. We then construct a semantic representation map of the scene with dynamic region attention to demonstrate the known visual clues, where the goal of task planning and scene exploration is aligned for human instruction. For the task planner, we generate the feasible step-by-step plans for human goal accomplishment according to the task completion process and the known visual clues. For the exploration controller, the optimal navigation or object interaction policy is predicted based on the generated step-wise plans and the known visual clues. The experimental results demonstrate that our method can achieve 45.09% success rate in 204 complex human instructions such as making breakfast and tidying rooms in large house-level scenes. Code and supplementary are available at https://gary3410.github.io/eif_unknown/.
Ziwei Wang 0001, Xiuwei Xu, Yinan Liang, Angyuan Ma, Jiwen Lu, Haibin Yan
IROS8
2025 Anyview: General Indoor 3D Object Detection with Variable Frames
abstract
In this paper, we propose a novel network framework for indoor 3D object detection to handle variable input frame numbers in practical scenarios. Existing methods only consider fixed frames of input data for a single detector, such as monocular RGB-D images or point clouds reconstructed from dense multi-view RGB-D images. While in practical application scenes such as robot navigation and manipulation, the raw input to the 3D detectors is the RGB-D images with variable frame numbers instead of the reconstructed scene point cloud. However, the previous approaches can only handle fixed frame input data and have poor performance with variable frame input. In order to facilitate 3D object detection methods suitable for practical tasks, we present a novel 3D detection framework named AnyView for our practical applications, which generalizes well across different numbers of input frames with a single model. To be specific, we propose a geometric learner to mine the local geometric features of each input RGB-D image frame and implement local-global feature interaction through a designed spatial mixture module. Meanwhile, we further utilize a dynamic token strategy to adaptively adjust the number of extracted features for each frame, which ensures consistent global feature density and further enhances the generalization after fusion. Extensive experiments on the ScanNet dataset show our method achieves both great generalizability and high detection accuracy with a simple and clean architecture containing a similar amount of parameters with the baselines.
Xiuwei Xu, Ziwei Wang 0010, Chong Xia, Linqing Zhao, Jiwen Lu, Haibin Yan
IROS7
2025 Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline
Linqing Zhao, Xiuwei Xu, Wenzhao Zheng, Yansong Tang, Haibin Yan, Jiwen Lu
IROS7
2025 Interactive shape estimation for densely cluttered objects
Jiangfan Ran, Haibin Yan
Pattern Recognit. Lett.2
2025 Kinship verification via Frequency Feature Decoupling and Fusion
Shuofeng Sun, Yaohan Yang, Haibin Yan
Pattern Recognit. Lett.3
2025 RoboPacker: An Autonomous Robotic Packing System for General Objects
abstract
In this paper, we propose an autonomous robot packing system named RoboPacker designed to tightly store cluttered general objects into shipping boxes with high space utilization, which is a fundamental process in numerous industrial applications. However, achieving tight packaging for general objects often demands significant labor from human packers, particularly in high-throughput scenes. Compared to existing robot packing approaches, RoboPacker effectively overcomes challenges such as diverse object appearances, severe occlusion, and crowded packing spaces. Specifically, we propose an open-vocabulary shape estimation method to reconstruct complete point clouds for cluttered objects. We also design effective interactions with object clutter to gather informative visual clues for shape estimation under high uncertainty. Additionally, we introduce a hierarchical reinforcement learning framework to optimize packing order, location, and orientation for maximum space utilization. The robotic packing system integrates these techniques with feasible manipulation methods for real-world implementation. In this way, RoboPacker achieves efficient packing of novel and irregular objects, which is more suitable for real deployment environments. The Real-world experiments demonstrate RoboPacker can tightly pack 20 densely cluttered everyday objects from 8 seen and 4 novel classes into the 40×40×20 cm shipping box with a 73.3% success rate. The demonstration video can be found at https://gary3410.github.io/RoboPacker/.
Ziwei Wang 0010, Sichao Huang, Xiuwei Xu, Haibin Yan, Jiwen Lu
IEEE Trans Autom. Sci. Eng.6
2025 PointMax: Self-Boosted Local Sampling for 3D Point Cloud Analysis
abstract
Local sampling plays a key role in modeling 3D point clouds. Due to the disordered and unstructured nature of point cloud data, conventional 3D deep models such as PointNet++ and its variants usually employ random or fixed rules to sample local neighborhoods, leading to considerable redundancy in the feature aggregation process. In this paper, we propose a self-supervised method for learning to adaptively select effective neighbors. Firstly, we observe that only a part of sampled points contributes to the aggregated features after the max-pooling operation in existing point cloud models. Then, based on this observation, we propose a simple and task-oriented metric to evaluate the sampling efficiency by measuring the effective neighbors in the feature aggregation process. The metric is also used to supervise a lightweight neighborhood scoring module (NSM), which is designed to efficiently select effective neighboring points from a wider range of neighbors to reduce the computational cost and keep the performance superior. To further improve the performance, we introduce Neighborhood Attention in the feature aggregation process according to the importance score of neighborhood points predicted by NSM. Experimental results show that our method is simple and efficient, and can be applied to most tasks and models to reduce the computational cost and keep the performance superiority. Our code is available athttps://github.com/sunshuofeng/PointMax_Code
Shuofeng Sun, Yongming Rao, Jiwen Lu, Haibin Yan
IEEE Trans. Multim.4
2024 X-3D: Explicit 3D Structure Modeling for Point Cloud Recognition
abstract
Numerous prior studies predominantly emphasize constructing relation vectors for individual neighborhood points and generating dynamic kernels for each vector and embedding these into high-dimensional spaces to capture implicit local structures. However, we contend that such implicit high-dimensional structure modeling approch inadequately represents the local geometric structure of point clouds due to the absence of explicit structural information. Hence, we introduce X-3D, an explicit 3D structure modeling approach. X-3D functions by capturing the explicit local structural information within the input 3D space and employing it to produce dynamic kernels with shared weights for all neighborhood points within the current local region. This modeling approach introduces effective geometric prior and significantly diminishes the disparity between the local structure of the embedding space and the original input point cloud, thereby improving the extraction of local features. Experiments show that our method can be used on a variety of methods and achieves state-of-the-art performance on segmentation, classification, de-tection tasks with lower extra computational cost, such as 90.7% on ScanObjectNN for classification, 79.2% on S3DIS 6 fold and 74.3% on S3DIS Area 5 for segmentation, 76.3% on ScanNetV2 for segmentation and 64.5% mAP25, 46.9% mAP50on SUN RGB-D and 69.0% mAP25, 51.1% mAP50on ScanNetV2. Our code is available at https://github.com/sunshuofeng/X-3D.
Shuofeng Sun, Yongming Rao, Jiwen Lu, Haibin Yan
CVPR4
2024 FairScene: Learning unbiased object interactions for indoor scene synthesis
Ziwei Wang 0010, Jiwen Lu, Haibin Yan
Pattern Recognit.6
2024 Toward Integrity and Detail With Ensemble Learning for Salient Object Detection in Optical Remote-Sensing Images
abstract
Optical remote sensing image salient object detection (ORSI-SOD) poses significant challenges due to complicated object variances and interfering surroundings. Although existing methods have achieved impressive performance, they encounter difficulties in balancing deep and shallow features, leading to limitations in preserving object integrity and edge detail. To address this, we propose the Integrated and Detailed Ensemble Learning (IDEL) framework, which incorporates hierarchical branches with deep supervision. By divide-and-conquer, each branch captures information with a specific granularity, while the fusion module combines all outputs to generate the final saliency maps. To ensure the effectiveness of ensemble learning, IDEL is designed to satisfy two necessary conditions: the weak learner property and branch independence. Firstly, we utilize the Transformer blocks with a global receptive field and purify intermediate features with the Deep Supervision Module (DSM) to enhance the performance of each branch. Secondly, we disentangle multiple branches through hardness-aware weights and hierarchical supervision labels, allowing them to learn distinct features. Qualitative visualizations demonstrate the effectiveness of each module, and extensive experimental results conducted on three popular ORSI datasets confirm the superiority of IDEL compared to other state-of-the-art (SOTA) counterparts.
Kangjie Liu, Borui Zhang, Jiwen Lu, Haibin Yan
IEEE Trans. Geosci. Remote. Sens.4
2024 Hardness-Aware Scene Synthesis for Semi-Supervised 3D Object Detection
abstract
3D object detection aims to recover the 3D information of concerning objects and serves as the fundamental task of autonomous driving perception. Its performance greatly depends on the scale of labeled training data, yet it is costly to obtain high-quality annotations for point cloud data. This motivates the use of semi-supervised learning which can additionally exploit unlabeled data to further boost the performance. While 2D semi-supervised learning methods focus on generating pseudo-labels for unlabeled existing samples as supplements for training, the structural nature of 3D point cloud data facilitates the composition of objects and backgrounds to synthesize realistic scenes. Motivated by this, we propose a hardness-aware scene synthesis (HASS) method to generate adaptive synthetic scenes to improve the generalization of the detection models. We obtain pseudo-labels for unlabeled objects and generate diverse scenes with different compositions of objects and backgrounds. As the scene synthesis is sensitive to the quality of pseudo-labels, we further propose a hardness-aware strategy to reduce the effect of low-quality pseudo-labels. In addition, we maintain a dynamic pseudo- database to ensure the diversity and quality of synthetic scenes. Extensive experimental results on the widely used KITTI and Waymo datasets demonstrate the superiority of the proposed HASS method, which outperforms existing semi-supervised learning methods on 3D object detection. We also conducted a series of experiments to analyze the effectiveness of our method including pseudo-label quality analysis, the effect of different filtering and thresholding strategies, and ablations of each component.
Wenzhao Zheng, Jiwen Lu, Haibin Yan
IEEE Trans. Multim.4
2023 AE-Reorient: Active Exploration Based Reorientation for Robotic Pick-and-Place
Haibin Yan
ICIG (4)3
2023 Category-level Shape Estimation for Densely Cluttered Objects
abstract
Accurately estimating the shape of objects in dense clutters makes important contribution to robotic packing, because the optimal object arrangement requires the robot planner to acquire shape information of all existed objects. However, the objects for packing are usually piled in dense clutters with severe occlusion, and the object shape varies significantly across different instances for the same category. They respectively cause large object segmentation errors and inaccurate shape recovery on unseen instances, which both degrade the performance of shape estimation during deployment. In this paper, we propose a category-level shape estimation method for densely cluttered objects. Our framework partitions each object in the clutter via the multi-view visual information fusion to achieve high segmentation accuracy, and the instance shape is recovered by deforming the category templates with diverse geometric transformations to obtain strengthened generalization ability. Specifically, we first collect the multi-view RGB-D images of the object clutters for point cloud reconstruction. Then we fuse the feature maps representing the visual information of multi-view RGB images and the pixel affinity learned from the clutter point cloud, where the acquired instance segmentation masks of multi-view RGB images are projected to partition the clutter point cloud. Finally, the instance geometry information is obtained from the partially observed instance point cloud and the corresponding category template, and the deformation parameters regarding the template are predicted for shape estimation. Experiments in the simulated environment and real world show that our method achieves high shape estimation accuracy for densely cluttered everyday objects with various shapes.
Ziwei Wang 0001, Jiwen Lu, Haibin Yan
ICRA4
2023 Adaptive dynamic networks for object detection in aerial images
Haibin Yan
Pattern Recognit. Lett.2
2023 Dense Hybrid Proposal Modulation for Lane Detection
abstract
In this paper, we present a dense hybrid proposal modulation (DHPM) method for lane detection. Most existing methods perform sparse supervision on a subset of high-scoring proposals, while other proposals fail to obtain effective shape and location guidance, resulting in poor overall quality. To address this, we densely modulate all proposals to generate topologically and spatially high-quality lane predictions with discriminative representations. Specifically, we first ensure that lane proposals are physically meaningful by applying single-lane shape and location constraints. Benefitting from the proposed proposal-to-label matching algorithm, we assign each proposal a target ground truth lane to efficiently learn from spatial layout priors. To enhance the generalization and model the inter-proposal relations, we diversify the shape difference of proposals matching the same ground-truth lane. In addition to the shape and location constraints, we design a quality-aware classification loss to adaptively supervise each positive proposal so that the discriminative power can be further boosted. Our DHPM achieves very competitive performances on four popular benchmark datasets. Moreover, we consistently outperform the baseline model on most metrics without introducing new parameters and reducing inference speed. The codes of our method are available athttps://github.com/wuyuej/DHPM.
Yuejian Wu, Linqing Zhao, Jiwen Lu, Haibin Yan
IEEE Trans. Circuits Syst. Video Technol.4
2022 TC-Net: Transformer-Convolutional Networks for Road Segmentation
abstract
In this paper, we propose a lightweight backbone network called Transformer-Convolutional Networks (TC-Net) for ef-ficient road segmentation. Conventional road segmentation methods employ data fusion to enrich key features, so that more branches are required in the network which increases the parameter size of the model and the difficulty of deploy-ment on edge devices. On the contrary, our TC-Net fully uti-lizes the interconnection between various regions on the input image to enhance the feature representation ability while re-ducing the amount of parameters of each branch. The key components of our TC-Net are the Transformer-Conv (TC) and PatchMerging-Conv (PC) modules. Specifically, the TC module applies convolution to optimize the process of cross-window connections, and the PC module flexibly adjusts the scale of feature maps to further reduce the computational cost. Extensive experiments on the KITTI dataset show that our method greatly reduces the amount of parameters while achieving comparable performance with state-of-the-arts.
Chaohui Song, Haibin Yan
ICME3
2022 Smart Explorer: Recognizing Objects in Dense Clutter via Interactive Exploration
abstract
Recognizing objects in dense clutter accurately plays an important role to a wide variety of robotic manipulation tasks including grasping, packing, rearranging and many others. However, conventional visual recognition models usually miss objects because of the significant occlusion among instances and causes incorrect prediction due to the visual ambiguity with the high object crowdedness. In this paper, we propose an interactive exploration framework called Smart Explorer for recognizing all objects in dense clutters. Our Smart Explorer physically interacts with the clutter to maximize the recognition performance while minimize the number of motions, where the false positives and negatives can be alleviated effectively with the optimal accuracy-efficiency trade-offs. Specifically, we first collect the multi-view RGB-D images of the clutter and reconstruct the corresponding point cloud. By aggregating the instance segmentation of RGB images across views, we acquire the instance-wise point cloud partition of the clutter through which the existed classes and the number of objects for each class are predicted. The pushing actions for effective physical interaction are generated to sizably reduce the recognition uncertainty that consists of the instance segmentation entropy and multi-view object disagreement. Therefore, the optimal accuracy-efficiency trade-off of object recognition in dense clutter is achieved via iterative instance prediction and physical interaction. Extensive experiments demonstrate that our Smart Explorer acquires promising recognition accuracy with only a few actions, which also outperforms the random pushing by a large margin.
Ziwei Wang 0010, Zibu Wei, Yi Wei 0003, Haibin Yan
IROS5
2022 Ultra-High Resolution Image Segmentation with Efficient Multi-Scale Collective Fusion
abstract
Ultra-high resolution image segmentation has at-tracted increasing attention recently due to its wide applications in various scenarios such as road extraction and urban planning. The ultra-high resolution image facilitates the capture of more detailed information but also poses great challenges to the image understanding system. For memory efficiency, existing methods preprocess the global image and local patches into the same size, which can only exploit local patches of a fixed resolution. In this paper, we empirically analyze the effect of different patch sizes and input resolutions on the segmentation accuracy and propose a multi-scale collective fusion (MSCF) method to exploit information from multiple resolutions, which can be end-to-end trainable for more efficient training. Our method achieves very competitive performance on the widely-used DeepGlobe dataset while training on one single GPU.
Haibin Yan
VCIP2
2021 Multi-scale deep relational reasoning for facial kinship verification
Haibin Yan, Chaohui Song
Pattern Recognit.1
2020 KINMIX: A Data Augmentation Approach for Kinship Verification
abstract
In this paper, we propose a KinMix method to generate positive samples in the feature space for facial kinship verification. Unlike most existing data augmentation methods, we generate samples at the feature level instead of the original image level for data augmentation. Specifically, we use a pair of features with a kin relationship to construct a feature space instead of a single image, which aims to maintain the clustering results of features. We assume that a pair of kinship features lie in the same feature space, so that the generated feature remains the same kin relationship if it is sampled from the feature space. We use a linear sampling method to generate positive samples with the kin relationship. We evaluate our method on two widely used kinship datasets: KinFaceW-I and KinFaceW-II, and our experimental results are presented to show the effectiveness of the proposed approach.
Chaohui Song, Haibin Yan
ICME2
2020 Discriminative sampling via deep reinforcement learning for kinship verification
Haibin Yan
Pattern Recognit. Lett.2
2019 Learning discriminative compact binary face descriptor for kinship verification
Haibin Yan
Pattern Recognit. Lett.1
2019 Semantic three-stream network for social relation recognition
Haibin Yan, Chaohui Song
Pattern Recognit. Lett.1
2019 Learning part-aware attention networks for kinship verification
Haibin Yan
Pattern Recognit. Lett.1
2018 Video-based Parent-Child Relationship Prediction
abstract
In this paper, we investigate the problem of video-based parent-child relationship prediction via human face analysis. Most existing kinship verification methods predict the parent-child relationship from single images, which cannot effectively utilize videos of human faces for kinship verification. Recently, there have been a few methods for parent-child relationship prediction based on face videos, but all of them only perform pairwise comparisons between human faces between a single parent and a single child. Thus, they cannot effectively combine information about both the father's and the mother's faces when judging the kin relationship. In this paper, we propose a new dtaaset, Familyship Face Videos in the Wild (FFVW), which was captured both in wild conditions and standard reference, to deal with this issue. The inputs of FFVW are three separate videos of a family. To our best knowledge, our paper is the first attempt at addressing this problem. In our pre-processing step, we extract four key frames from each video, before doing facial recognition and alignment. Finally, we use a convolutional neural network to make the prediction. Overall, the effectiveness of this approach is verified by experimental results, which show that our dataset outperforms previous approaches to parent-child relationship prediction.
Yiwen Wei, Haibin Yan
VCIP4
2018 Collaborative discriminative multi-metric learning for facial expression recognition in video
Haibin Yan
Pattern Recognit.1
2018 Video-based kinship verification using distance metric learning
Haibin Yan, Junlin Hu 0001
Pattern Recognit.1
2017 Kinship verification using neighborhood repulsed correlation metric learning
Haibin Yan
Image Vis. Comput.1
2016 Transfer subspace learning for cross-dataset facial expression recognition
Haibin Yan
Neurocomputing1
2016 Discriminative sparse projections for activity-based person recognition
Haibin Yan
Neurocomputing1
2016 Biased subspace learning for misalignment-robust facial expression recognition
Haibin Yan
Neurocomputing1
2016 Kinship verification from facial images by scalable similarity fusion
Xiuzhuang Zhou, Haibin Yan
Neurocomputing2
2015 Neighborhood repulsed correlation metric learning for kinship verification
abstract
In this paper, we propose a new neighborhood repulsed correlation metric learning (NRCML) method for kinship verification. While several metric learning algorithms have been proposed in recent years and some of them have successfully applied to kinship verification, most existing metric learning methods are developed based on the Euclidian similarity metric, which is not powerful enough to measure the similarity of face samples. To address this, we propose a NRCML method by using the correlation similarity measure to learn a discriminative distance metric, under which positive pairs are pulled as close as possible and negative pairs lying in a neighborhood are repulsed as far as possible, simultaneously. Experimental results are presented to show the effectiveness of the proposed method.
Haibin Yan, Xiuzhuang Zhou, Yongxin Ge
VCIP1
2015 Prototype-Based Discriminative Feature Learning for Kinship Verification
abstract
In this paper, we propose a new prototype-based discriminative feature learning (PDFL) method for kinship verification. Unlike most previous kinship verification methods which employ low-level hand-crafted descriptors such as local binary pattern and Gabor features for face representation, this paper aims to learn discriminative mid-level features to better characterize the kin relation of face images for kinship verification. To achieve this, we construct a set of face samples with unlabeled kin relation from the labeled face in the wild dataset as the reference set. Then, each sample in the training face kinship dataset is represented as a mid-level feature vector, where each entry is the corresponding decision value from one support vector machine hyperplane. Subsequently, we formulate an optimization function by minimizing the intraclass samples (with a kin relation) and maximizing the neighboring interclass samples (without a kin relation) with the mid-level features. To better use multiple low-level features for mid-level feature learning, we further propose a multiview PDFL method to learn multiple mid-level features to improve the verification performance. Experimental results on four publicly available kinship datasets show the superior performance of the proposed methods over both the state-of-the-art kinship verification methods and human ability in our kinship verification task.
Haibin Yan, Jiwen Lu, Xiuzhuang Zhou
IEEE Trans. Cybern.1
2014 Cost-sensitive ordinal regression for fully automatic facial beauty assessment
Haibin Yan
Neurocomputing1
2014 Multi-feature multi-manifold learning for single-sample face recognition
Haibin Yan, Jiwen Lu, Xiuzhuang Zhou
Neurocomputing1
2014 Discriminative Multimetric Learning for Kinship Verification
abstract
In this paper, we propose a new discriminative multimetric learning method for kinship verification via facial image analysis. Given each face image, we first extract multiple features using different face descriptors to characterize face images from different aspects because different feature descriptors can provide complementary information. Then, we jointly learn multiple distance metrics with these extracted multiple features under which the probability of a pair of face image with a kinship relation having a smaller distance than that of the pair without a kinship relation is maximized, and the correlation of different features of the same face sample is maximized, simultaneously, so that complementary and discriminative information is exploited for verification. Experimental results on four face kinship data sets show the effectiveness of our proposed method over the existing single-metric and multimetric learning methods.
Haibin Yan, Jiwen Lu, Weihong Deng, Xiuzhuang Zhou
IEEE Trans. Inf. Forensics Secur.1
2011 Weighted biased linear discriminant analysis for misalignment-robust facial expression recognition
abstract
We investigate in this paper the problem of misalignment-robust facial expression recognition. To the best of our knowledge, this problem has not been formally addressed in the literature. Most existing facial expression recognition methods, however, can only work well when face images are well-aligned. In many real world applications such as human robot interaction and visual surveillance, it is still very challenging to obtain well-aligned face images for expression recognition due to currently imperfect vision techniques, especially under uncontrolled conditions. Motivated by the fact that interclass facial images with small differences are more easily mis-classified than those with large differences, we propose a biased linear discriminant analysis (BLDA) method by imposing large penalties on interclass samples with small differences and small penalties on those samples with large differences simultaneously, such that more discriminative features can be extracted for recognition. Moreover, we generate more virtually misaligned facial expression samples and assign different weights to them according to their occurrence probabilities in the testing phase to learn a weighted BLDA (WBLDA) feature space to extract misalignment-robust discriminative features for recognition. Experimental results on two widely used face databases are presented to show the efficacy of the proposed method.
Haibin Yan, Marcelo H. Ang, Aun Neow Poo
ICRA1
2011 Cross-dataset facial expression recognition
abstract
This paper investigates the problem of cross-dataset facial expression recognition. To the best of our knowledge, this problem has not been formally addressed in the literature. Conventional facial expression recognition methods assume expression images in the training and testing sets are collected under the same condition such that they are independent and identically distributed. In many real applications, this assumption may not hold as the testing data are usually collected online and generally more uncontrollable than the training data, and hence, they are likely different from the training data. This problem is referred to as cross-dataset facial expression recognition in this paper as the training and testing data are considered to be collected from different datasets due to different acquisition conditions. To address this, we propose a new transfer subspace learning approach to learn a feature space which transfers the knowledge gained from the training set to the target (testing) data to improve the recognition performance under cross-dataset scenarios. Experimental results for facial expression recognition tasks on different datasets are presented to demonstrate the efficacy of the proposed approach.
Haibin Yan, Marcelo H. Ang, Aun Neow Poo
ICRA1