Hui Yu 0001

dblp:26/6190-1 · DBLP profile ↗
← Back
137ranked-venue papers
7as first author
91since 2021 · last 2026
0000-0002-7655-9228ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 2 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 4 first-author · 23 since 2021Human-computer interaction and ubiquitous computing · 26 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 17 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A comprehensive survey of deep learning-based cognitive diagnosis models in education: Methods, applications, and outlook
Jing Li 0046, Yang Xu 0025, Enrique Herrera-Viedma, Hui Yu 0001, Jiande Sun 0001
Neurocomputing5
2026 Camera-Space hand mesh reconstruction from a monocular image via pseudo stereo perception
Shaoxiang Guo, Wankun Chen, Hui Yu 0001, Junyu Dong
Knowl. Based Syst.5
2026 HMamba-3DFT: A hierarchical mamba framework for emotion-driven semantic 3D facial tracking
abstract
• First Mamba-based framework HMamba-3DFT tailored for 3D facial tracking • BSTV-Mamba with BSTS-Scan capture spatiotemporal facial dynamics • Dual optimization integrates dynamic emotion-driven modeling with semantic alignment Monocular video-based 3D face tracking is vital for interactive pattern recognition and human avatars. Most existing image-based methods fail to model temporal dependencies in video, causing jitter and inaccuracies. Furthermore, they also often neglect the continuous multi-modal signals present in facial videos such as expression dynamics and emotional cues that provide essential temporal drivers for facial modeling. To this end, this study first explores the Mamba architecture tailored for 3D facial tracking by proposing a hierarchical Mamba framework, termed HMamba-3DFT. The proposed network can efficiently capture and track variations in 3D facial shapes from a monocular video. To exploit the global spatiotemporal correlations across frames of the dynamic face, we develop a bidirectional spatiotemporal vision Mamba (BSTV-Mamba) module featuring a bidirectional spatiotemporal selective scan (BSTS-Scan) mechanism. To capture temporally evolving multi-modal emotion signals embedded in continuous video sequences, we introduce a dynamic emotion-driven mechanism. Additionally, to mitigate the potential degradation of reconstruction fidelity caused by an over-reliance on emotion-driven cues, we integrate facial semantic alignment with facial emotion driving to enhance the accuracy of emotion-driven facial modeling. This integrated dual-optimization strategy systematically guides the network during training, ensuring that the reconstructed 3D facial mesh not only accurately captures the emotional attributes of the input frames but also benefits from enhanced optimization for more precise reconstruction. Extensive evaluations on benchmark datasets show competitive performance against state-of-the-art methods.
Haodong Jin, Muwei Jian, Derui Ding, Hui Yu 0001
Pattern Recognit.4
2026 Visual perception-inspired 3D point cloud sampling
abstract
Task-oriented sampling aims to predict the importance of points of a point cloud to better serve downstream tasks, which has attracted increasing attention in the fields of computer vision and visualization in recent years. However, existing methods cannot sufficiently leverage both global saliency and local saliency cues, resulting in suboptimal performance that requires further improvement. To tackle this challenge, we propose a novel 3D point cloud sampling method inspired by the human visual perception mechanism in this study, which can effectively extract important point cloud subsets from critical regions to better adapt to downstream tasks, thereby maintaining superior sampling performance. The proposed Visual Perception-inspired 3D Point Cloud Sampling (VPI-3DPS) method simulates the human visual system’s dynamic attention-shifting strategy by combining coarse-grained attention-driven sampling with fine-grained detail preservation. This allows our approach to adaptively capture both global context and local details within point cloud data, safeguarding downstream task performance. By leveraging Gated Recurrent Units (GRUs) for long-term dependency modeling and integrating Graph Convolutional Networks (GCNs) to capture local structures, VPI-3DPS obtains an integrated representation of regional correlation and detail awareness. Extensive experiments show that VPI-3DPS outperforms existing methods. Compared to the best-performing approaches, it achieves an average increase of 1.29% in classification accuracy, an average reduction of 13.20% in registration MRE, and an average decrease of 4.29% in Chamfer Distance for reconstruction.
Xu Wang 0053, Yi Jin 0001, Hui Yu 0001, Yi-Gang Cen, Yidong Li
Pattern Recognit.3
2026 A Stitch in Time Saves Nine: Progressive Information Bottleneck for Incremental Multiview Clustering
abstract
Incremental multiview clustering (IMVC) leverages consistent information between historical and new views to benefit the clustering task. However, existing IMVC approaches ignore the redundant information in individual views, leading to an accumulation of irrelevance. Besides, with the continuous arrivals of new views, the knowledge learned from historical views is often forgotten, which hinders the learning models from achieving long-term dependencies across incremental views. In this study, we propose a novel progressive information bottleneck (PIB), which is capable of removing redundant information in a timely manner and selectively updating historical knowledge based on information gain of new views. Specifically, to facilitate the knowledge transfer from historical views to incoming one, an information-aware knowledge library is built to store the representative samples of historical views. With the emergence of new views, we first devise a matrix-based mutual information (MI) constraint on an encoder to compress redundant information, which facilitates the training of a neural network with analyzable gradients and obtain a compact yet discriminative representation. Then, a dual-selective updating strategy is proposed to preserve historical knowledge in time when it contributes more to the information gain of the knowledge library than the new view. Finally, relevant samples to the new view in knowledge library are migrated to maximize the cross-level consistency between historical and new views. To the best of our knowledge, this is the first work that designs a gradient-analyzable MI measurement for incremental multiview learning and employs information gain to guide the selective update of the knowledge library. Empirical evaluations on six benchmark datasets show that our method outperforms state-of-the-art baseline methods by an average of 8.1%, 7.6%, and 7.8% on clustering accuracy (ACC), normalized mutual information (NMI), and adjusted Rand index (ARI) metrics, respectively.
Fengshou Han, Yiqiao Mao, Zhen Tian 0004, Witold Pedrycz, Hui Yu 0001
IEEE Trans. Comput. Soc. Syst.6
2026 SMHNet: Self-Supervised Multiscale Hierarchical Network for High Fidelity 3-D Face Reconstruction
abstract
High-fidelity 3-D face reconstruction is critical for enhancing personalized and immersive human–machine interaction experiences. However, existing methods struggle to capture the full spectrum of facial textures, particularly fine-scale details, such as wrinkles and pores, due to limitations in multiscale representation. To address this challenge, we propose a self-supervised multiscale hierarchical network to hierarchically model fine geometric details in multiple scales in this study. We design a global and local Markov random field loss and a detail perception loss to provide a global and local sensory field of view guidance for retaining fine-scale detail structure information of the face. In addition, we introduce a learnable Gabor-aware texture enhancement module to enhance the network’s sensitivity to fine textures. Extensive experiments show that the proposed method can reconstruct fine-scale details of the face and has superior performance to the state-of-the-art methods in terms of reconstruction accuracy and visual effect.
Sizhuang Zhang, Ying Sun 0004, Derui Ding, Hui Yu 0001
IEEE Trans. Hum. Mach. Syst.4
2026 GAT-NeRF: Geometry-Aware-Transformer-Enhanced Neural Radiance Fields for High-Fidelity 4D Facial Avatars
abstract
High-fidelity 4D dynamic facial avatar reconstruction from monocular video is a critical yet challenging task, driven by increasing demands for immersive virtual human applications. While Neural Radiance Fields (NeRF) have advanced scene representation, their capacity to capture high-frequency facial details, such as dynamic wrinkles and subtle textures from information-constrained monocular streams, requires significant enhancement. To tackle this challenge, we propose a novel hybrid NeRF framework, called Geometry-Aware-Transformer-Enhanced NeRF (GAT-NeRF) for high-fidelity and controllable 4D facial avatar reconstruction, which integrates the Transformer mechanism into the NeRF pipeline. GAT-NeRF synergistically combines a coordinate-aligned Multilayer Perceptron (MLP) with a lightweight Transformer module, termed as Geometry-Aware Transformer (GAT) due to its processing of multi-modal inputs containing explicit geometric priors. The GAT module is enabled by fusing multi-modal input features, including 3D spatial coordinates, 3D Morphable Model (3DMM) expression parameters, and learnable latent codes to effectively learn and enhance feature representations pertinent to fine-grained geometry. The Transformer’s effective feature learning capabilities are leveraged to significantly augment the modeling of complex local facial patterns like dynamic wrinkles and acne scars. Comprehensive experiments unequivocally demonstrate GAT-NeRF’s state-of-the-art performance in visual fidelity and high-frequency detail recovery, forging new pathways for creating realistic dynamic digital humans for multimedia applications.
Zhe Chang, Haodong Jin, Yan Song 0002, Ying Sun 0004, Hui Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2025 SelfieAvatar: Real-time Head Avatar reenactntment from a Selfie Video
abstract
Head avatar reenactment focuses on creating animatable personal avatars from monocular videos, serving as a foundational element for applications like social signal understanding, gaming, human-machine interaction, and computer vision. Recent advances in 3D Morphable Model (3DMM)-based facial reconstruction methods have achieved remarkable high-fidelity face estimation. However, on the one hand, they struggle to capture the entire head, including non-facial regions and background details in real time, which is an essential aspect for producing realistic, high-fidelity head avatars. On the other hand, recent approaches leveraging generative adversarial networks (GANs) for head avatar generation from videos can achieve high-quality reenactments but encounter limitations in reproducing fine-grained head details, such as wrinkles and hair textures. In addition, existing methods generally rely on a large amount of training data, and rarely focus on using only a simple selfie video to achieve avatar reenactment. To address these challenges, this study introduces a method for detailed head avatar reenactment using a selfie video. The approach combines 3DMMs with a StyleGAN-based generator. A detailed reconstruction model is proposed, incorporating mixed loss functions for foreground reconstruction and avatar image generation during adversarial training to recover high-frequency details. Qualitative and quantitative evaluations on self-reenactment and cross-reenactment tasks demonstrate that the proposed method achieves superior head avatar reconstruction with rich and intricate textures compared to existing approaches.
Hui Yu 0001, Derui Ding, Rachael Jack, Philippe G. Schyns
FG2
2025 Generative Adversarial Network-based Image and Tabular Data Generation with Differential Privacy
abstract
Machine learning and artificial intelligence technologies have become integral to various industries, driven by the availability of large-scale data. However, the use of sensitive industrial and personal data introduces significant privacy risks. Generative Adversarial Networks (GANs) are employed to generate synthetic data, thereby rendering them a feasible privacy-preserving technique. Despite their potential, existing private GANs face two major limitations: they are restricted to single-modal data generation or fail to ensure strict privacy guarantees. To tackle these issues, we propose Differential Privacy GAN of Image and Table—DPGAN-IT, which is capable of generating both image and tabular data simultaneously while enforcing robust privacy protection. Extensive experiments validate high utility of multi-modal synthetic data and demonstrate an effective balance between privacy and data utility.
Jiming Yang, Xu Wang 0053, Yi Jin 0001, Yidong Li, Hui Yu 0001
ICME5
2025 Open World Adaptive Pseudo Contrastive Learning for Generalized Category Discovery
abstract
In this work, we investigate the challenging task of Generalized Category Discovery (GCD). Given datasets collected from open-world scenarios comprising both labeled and unlabeled images, GCD aims to classify all unlabeled images while simultaneously identifying unlabeled novel categories. The fundamental challenge in GCD tasks stems from inherent annotation discrepancies between seen and novel classes within the dataset. The lack of reliable label supervision for novel classes in unlabeled data leads to significant disparities in the model’s learning between old and novel classes, which is termed the bias risk. Recent advancements in GCD have employed the entropy maximization algorithm to alleviate the bias risk. However, they fail to provide debiased optimization for unlabeled data, leading to models that struggle with extracting discriminative features from such data. To address these challenges, we have created an Open-world pseudo-contrastive learning framework named OpcGCD. Our OpcGCD framework implements a dynamic category-wise threshold mechanism, which employs a parametric prototype classifiers to generate debiased pseudo-labels for unlabeled samples. To facilitate the learning of discriminative feature representations, our proposed OpcGCD employs debiased pseudo-labels in the formulation of a contrastive learning loss. Extensive evaluations conducted on multiple GCD benchmark datasets demonstrate the robustness and effectiveness of the approach.
Yiqing Hao, Xu Wang 0053, Yi Jin 0001, Tao Wang 0011, Yidong Li, Shuoyan Liu, Chao Li 0026, Hui Yu 0001
SMC8
2025 CG-MCFNet: cross-layer guidance-based multi-scale correlation fusion network for 3D face recognition
Panzi Zhao, Yue Ming 0001, Hui Yu 0001, Jiangwan Zhou
Appl. Intell.3
2025 CSSANet: A channel shuffle slice-aware network for pulmonary nodule detection
Muwei Jian, Huihui Huang, Rui Wang 0017, Hui Yu 0001
Neurocomputing6
2025 Out-of-distribution monocular depth estimation with local invariant regression
Yeqi Hu, Yuan Rao 0001, Hui Yu 0001, Gaige Wang, Hao Fan 0004, Wei Pang 0001, Junyu Dong
Knowl. Based Syst.3
2025 GLMF-NET: global and local multi-scale fusion network for polyp segmentation
Muwei Jian, Yanjie Zhong, Hui Yu 0001
Multim. Syst.5
2025 Guest Editorial: Special Issue on Human-Machine Fusion Decision-Making for Emergency Handling
Qi Wu 0003, Jianqiang Li 0001, Guimin Chen, Mehmet R. Yuce, Javier Del Ser, Hui Yu 0001, Peter Xiaoping Liu
IEEE Trans Autom. Sci. Eng.7
2025 Hierarchical Oversampling Based on Cohen's Criterion for Imbalanced Data With Missing Information
abstract
The generative adversarial network (GAN) is increasingly used to address data imbalanced. However, GAN struggle with minority data that have few instances and lack accurate statistical characteristics. To mitigate this problem, a hierarchical oversampling based on Cohen’s criterion (HiOC) is proposed for extremely imbalanced data with missing information. The main idea of HiOC is to use Cohen’s criterion to design a primary rebalancing rate by considering data size, distribution, and feature information comprehensively. Then, HiOC including two stages to achieve the data balance. In the first stage, to enhance the data variety, especially for instances on the borderline as well as make a delicate imputation for missing information, a progressive oversampling method in the framework of fuzzy information decomposition (FID) is proposed. The progressive FID (PFID) introduces a small bias to expand the prediction region’s bounds and fulfills imputation step by step using newly sampled data. In the second stage, to greatly keep the original data distribution, an information granularity (IG) incorporated fuzzy c-means clustering strategy is developed. Afterward, a GAN based on Mahalanobis distance performs oversampling in each cluster to achieve ultimate data balance with unbiased evaluation. Finally, the proposed algorithm is applied to several real datasets, demonstrating higher classification accuracy compared with existing algorithms, thus proving the applicability of the research results in real-world scenarios.
Jun Dou, Yan Song 0002, Hui Yu 0001
IEEE Trans. Comput. Soc. Syst.3
2025 UWStereo: A Large Synthetic Dataset for Underwater Stereo Matching
abstract
Despite recent advances in stereo matching, the extension to intricate underwater settings remains unexplored, primarily owing to: 1) the reduced visibility, low contrast, and other adverse effects of underwater images; 2) the difficulty in obtaining ground truth data for training deep learning models, i.e. simultaneously capturing an image and estimating its corresponding pixel-wise depth information in underwater environments. To enable further advance in underwater stereo matching, we introduce a large synthetic dataset called UWStereo. Our dataset includes 29,568 synthetic stereo image pairs with dense and accurate disparity annotations for left view. We design four distinct underwater scenes filled with diverse objects such as corals, ships and robots. We also induce additional variations in camera model, lighting, and environmental effects. In comparison with existing underwater datasets, UWStereo is superior in terms of scale, variation, annotation, and photo-realistic image quality. To substantiate the efficacy of the UWStereo dataset, we undertake a comprehensive evaluation compared with eleven state-of-the-art algorithms as benchmarks. The results indicate that current models still struggle to generalize to new domains. Hence, we design a new strategy that learns to reconstruct cross domain masked images before stereo matching training and integrate a cross view attention enhancement module that aggregates long-range content information to enhance the generalization ability.
Qingxuan Lv, Junyu Dong, Yuezun Li, Sheng Chen 0001, Hui Yu 0001, Shu Zhang 0002, Wenhan Wang
IEEE Trans. Circuits Syst. Video Technol.5
2025 V2PNet: A Voxel-to-Point Network Framework for Task-Oriented Point Cloud Sampling
abstract
Task-oriented point cloud sampling is a fundamental technique in 3D computer vision and has become a crucial step in numerous 3D applications. However, most state-of-the-art task-oriented sampling methods adopt a point-wise analysis strategy, making them susceptible to data redundancy. Taking inspiration from the abstract-to-detailed recognition process of the human visual system, we propose a novel voxel-to-point network framework called V2PNet for task-oriented point cloud sampling. Specifically, we first design a lightweight coarse-grained sampling module named Important Voxel Prediction (IMVP). This module adaptively outputs points from significant regions of the point cloud by explicitly modeling inter-region relationships, thereby reducing interference from redundant points. Then, the V2PNet framework seamlessly integrates the IMVP module with existing point-wise and task-oriented sampling networks, enabling joint training with downstream tasks. This creates a task-oriented coarse-to-fine-grained sampling pipeline that effectively samples representative and informative points from significant regions to represent the original point cloud. Moreover, to mitigate disturbances across similar regions, we introduce a voxel simplification loss function to enhance the discriminative voxel prediction. Extensive experiments demonstrate that V2PNet improves the performance of existing state-of-the-art task-oriented sampling models.
Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Yidong Li, Hui Yu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Learning Semantic-Aware Point-Line Features for Localization and Reconstruction
abstract
High-precision image matching and localization technology in a 3D environment map is essential for many tasks, such as marine engineering detection, robotics, and autonomous navigation. However, current visual localization and reconstruction methods overly depend on point features, which lack robustness in low-texture environments. To address this limitation, we propose a novel framework for point and line localization and 3D reconstruction with semantic constraints, which integrates multiple innovative components to achieve superior performance. Firstly, we design a point-localization optimization strategy with uniform point sampling and point-based instance segmentation constraints, significantly improving image matching and camera localization accuracy. Secondly, we optimize the selection of 2D-3D lines and line matching using instance segment constraints, leveraging the structural and semantic richness of line features to complement point features. Thirdly, we perform a joint point and line feature 3D reconstruction, enabling the creation of accurate 3D environment maps even in challenging low-texture marine scenes.Our approach has been extensively tested on popular datasets and compared with state-of-the-art methods. This work significantly advances current visual localization and 3D reconstruction techniques by addressing their limitations in low-texture environments, while also providing a robust foundation for future research and applications in marine engineering, robotics, and autonomous navigation.
Jian Yang 0036, Yuan Rao 0001, Hao Fan 0004, Junyu Dong, Hui Yu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Self-Supervised Face Deocclusion via 3-D Face Reconstruction With Outlier Segmentation
abstract
Face occlusion poses a challenge for many human–machine systems, such as facial expression perception, social signal analysis, and human identity verification. Accurate face deocclusion is essential for improving the performance of identity recognition, expression recognition, and the robustness of human–machine systems. As a result, this area has attracted significant attention from researchers in recent years. However, most existing methods rely heavily on synthetic occluded face datasets and predefined occlusion masks labels, which limits their applicability in real-world scenarios. To this end, we propose a novel self-supervised generative adversarial networks (GANs)-based framework for face deocclusion in this study, which integrates 3-D facial reconstruction with outlier segmentation guidance. To achieve reliable self-supervised occlusion guidance, we introduce an outlier segmentation module that utilizes statistical priors to generate accurate occlusion masks, facilitating the deocclusion process. Furthermore, we design a GAN-based dual-branch module, which is capable of simultaneously generating the occlusion mask and the deoccluded face. Extensive experiments on the widely used datasets demonstrate the superior performance of our approach on existing methods. Our method achieves 35.71 in peak signal to noise ratio (PSNR) and 0.891 in structural similarity index measure (SSIM) for occluded face restoration, outperforming state-of-the-art techniques.
Haodong Jin, Muwei Jian, Derui Ding, Hui Yu 0001
IEEE Trans. Hum. Mach. Syst.4
2025 HG-SFDA: HyperGraph Learning Meets Source-Free Unsupervised Domain Adaptation
abstract
Source-Free unsupervised Domain Adaptation (SFDA) aims to classify target samples by only accessing a pre-trained source model and unlabelled target samples. Since no source data is available, transferring the knowledge from the source domain to the target domain is challenging. Existing methods normally exploit the pair-wise relation among target samples and attempt to discover their correlations by clustering these samples based on semantic features. The drawbacks of these methods include: 1) the pair-wise relation is limited to exposing the underlying correlations of two more samples, hindering the exploration of the structural information embedded in the target domain; and 2) the clustering process only relies on the semantic feature, while overlooking the critical effect of domain shift, i.e., the distribution differences between the source and target domains. To address these issues, we propose a new SFDA method that exploits the high-order neighborhood relation and explicitly takes the domain shift effect into account. Specifically, we formulate the SFDA as a hypergraph learning problem and construct hyperedges to explore the deep structural and context information among multiple samples. Moreover, we integrate a self-loop strategy into the constructed hypergraph to elegantly introduce the domain uncertainty of each sample. By clustering these samples based on hyperedges, both the semantic feature and domain shift effects are considered. We then describe an adaptive relation-based objective to tune the model with soft attention levels for all samples. Extensive experiments are conducted on Office-31, Office-Home, VisDA, DomainNet-126 and PointDA-10 datasets. The results demonstrate the superiority of our method over state-of-the-art counterparts. Our code is avaliable at https://github.com/OUC-POVA/HG-SFDA.
Jinkun Jiang, Qingxuan Lv, Yuezun Li, Yong Du 0003, Junyu Dong, Sheng Chen 0001, Hui Yu 0001
IEEE Trans. Image Process.7
2025 Enhancing Autonomous Driving Decision: A Hybrid Deep Reinforcement Learning-Kinematic-Based Autopilot Framework for Complex Motorway Scenes
abstract
Autonomous vehicles (AVs) still pose challenges in improving intelligence, safety, and reliability in complex motorway scenarios. Recently, deep reinforcement learning (DRL) has demonstrated superior decision-making capabilities in dynamic environments compared to rule-based methods. However, it requires considerable training resources due to a lack of innovative DRL component design (e.g., state space and reward) to link observation and action accurately. Its opaque nature may also result in hazardous driving conditions. In this paper, we introduce a hybrid autopilot framework that amalgamates three modules: (i) DRL is employed to build a smart, learnable, and scalable driving policy across various motorway scenarios; (ii) a kinematic-based co-pilot strategy is devised to bolster training efficiency and provide flexible decision-making guidance; and (iii) a rule-based system assesses and determines the final action outputs in real-time between itself and the DRL policy to further enhance safety. Extensive simulations are conducted under different complex motorway scenarios. The results indicate that the proposed framework surpasses the baseline DRL policy in terms of training efficiency, intelligence, safety, and reliability.
Yongqiang Lu 0002, Hongjie Ma, Edward Smart, Hui Yu 0001
IEEE Trans. Intell. Transp. Syst.4
2025 Incremental Multiview Clustering With Continual Information Bottleneck Method
abstract
Multiview clustering (MVC) provides a natural formulation to generate clusters for multiview data, which is fundamental to lots of industrial tasks like autonomous driving, defect detection, and multisensor information fusion, as part of the foundation models. Most existing MVC methods suppose that the data of multiple views are available during the clustering process. However, that is a very strong assumption and is impractical when the views are incremental over time. In addition, if directly applying existing MVC approaches to the clustering setting with incremental views, the massive redundant information in each view might limit the knowledge sharing between historical and newly arrived views. To solve these problems, a continual information bottleneck (CIB) method is presented in this article, which addresses the incremental MVC issue by maximally preserving the consistency of a sequence of views and removing the redundant information in each view. In particular, to facilitate the knowledge transfer from historical views to incoming one, we build a knowledge library to store the representative samples in historical views. When adding a new view, we first construct a view-specific encoder with information-theoretic constraints to learn a compact and discriminative representation, in which redundant information in the new view is eliminated. Then, to capture the consistency information between historical views and the new view, a shared encoder is devised after retrieving the global neighbors in the library for the new view, which is performed by contrasting the cluster assignment and feature representation simultaneously. Finally, a unified objective function is devised to simultaneously optimize the knowledge library and clustering process, in which the knowledge library is updated by maximizing the mutual information between the new view and all historical ones to keep tracking knowledge about the earlier views. Extensive experiment on nine multiview benchmarks has verified the superiority of the CIB method over 19 baselines.
Yiqiao Mao, Yangdong Ye, Hui Yu 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2024 Live and Learn: Continual Action Clustering with Incremental Views
abstract
Multi-view action clustering leverages the complementary information from different camera views to enhance the clustering performance. Although existing approaches have achieved significant progress, they assume all camera views are available in advance, which is impractical when the camera view is incremental over time. Besides, learning the invariant information among multiple camera views is still a challenging issue, especially in continual learning scenario. Aiming at these problems, we propose a novel continual action clustering (CAC) method, which is capable of learning action categories in a continual learning manner. To be specific, we first devise a category memory library, which captures and stores the learned categories from historical views. Then, as a new camera view arrives, we only need to maintain a consensus partition matrix, which can be updated by leveraging the incoming new camera view rather than keeping all of them. Finally, a three-step alternate optimization is proposed, in which the category memory library and consensus partition matrix are optimized. The empirical experimental results on 6 realistic multi-view action collections demonstrate the excellent clustering performance and time/space efficiency of the CAC compared with 15 state-of-the-art baselines.
Yingtao Gan, Yiqiao Mao, Yangdong Ye, Hui Yu 0001
AAAI5
2024 A Method for X-Ray Image Landmarks Localization using Cyclic Coordinate-Guided Strategy
abstract
In this study, we present a novel method for pinpointing landmarks in X-ray images, which simultaneously offers computational efficiency and localization precision. Our method leverages a cyclic coordinate-guided strategy that requires fewer model parameters and lower computational costs than traditional heatmap-based supervised methods. This is crucial for medical imaging applications where imaging devices often have limited computational resources yet require high-precision landmark localization. Our methodology involves a two-stage process that employs cyclic inference to optimize landmark localization. In the first stage, non-uniform sampling is used to capture the multiscale features of landmarks. This is followed by a second stage in which cyclic training fine-tunes the landmark coordinates towards their optimal positions. Our results indicate that our two-stage process achieves competitive localization performance with state-of-the-art methods yet with added benefits of lower computational overhead and smaller parameter count. Additionally, a global block was developed to capture global position information of landmarks, and experiments showed its effectiveness and its contribution in enhancing the model's landmark localization accuracy. We validated our method using two publicly available datasets, and the source code for our experiments is available on GitHub: https://github.com/switch626/CCG-CL.git.
Xifeng An, Eric Rigall, Shu Zhang 0002, Hui Yu 0001, Junyu Dong
ICASSP5
2024 Boosting Spatial-Spectral Masked Auto-Encoder Through Mining Redundant Spectra for HSI-SAR/LiDAR Classification
abstract
Although recent masked image modeling (MIM)-based HSI-LiDAR/SAR classification methods have gradually recognized the importance of the spectral information, they have not adequately addressed the redundancy among different spectra, resulting in information leakage during the pretraining stage. This issue directly impairs the representation ability of the model. To tackle the problem, we propose a new strategy, named Mining Redundant Spectra (MRS). Unlike randomly masking spectral bands, MRS selectively masks them by similarity to increase the reconstruction difficulty. Specifically, a random spectral band is chosen during pretraining, and the selected and highly similar bands are masked. Experimental results demonstrate that employing the MRS strategy during the pretraining stage effectively improves the accuracy of existing MIM-based methods on the Berlin and Houston 2018 datasets.
Junyan Lin, Xuepeng Jin, Feng Gao 0005, Junyu Dong, Hui Yu 0001
IGARSS5
2024 Enhancing Point Cloud Sampling Quality with Dual-Branch Fusion Networks
abstract
Task-oriented point cloud sampling methods have attracted considerable attention for their ability to adaptively select important point sets based on downstream tasks, achieving an excellent balance between data simplification and task performance. However, existing task-oriented sampling models, primarily based on single-branch designs, struggle to fully extract features from input point clouds that comprehensively reflect multi-dimensional key information, thus limiting their sampling performance. In this paper, we introduce a dual-branch sampling network, named DBS-NET, which conducts crucial point sampling from both the global and local importance perspectives separately before merging them, thereby preserving multi-dimensional key information of the input data during the sampling process. Qualitative and quantitative experimental results demonstrate the competitive performance of DBS-NET on the classification benchmark task.
Yi Jin 0001, Xu Wang 0053, Mengxia Hu, Hui Yu 0001, Yidong Li, Tao Wang 0011, Songhe Feng, Congyan Lang
SMC4
2024 Memory-aware continual learning with multi-modal social media streams for unsupervised disaster classification
Yiqiao Mao, Zirui Hu, Yangdong Ye, Hui Yu 0001
Adv. Eng. Informatics6
2024 Single depth image 3D face reconstruction via domain adaptive learning
Xiaoxu Cai, Jianwen Lou, Jiajun Bu, Junyu Dong, Haishuai Wang, Hui Yu 0001
Frontiers Comput. Sci.6
2024 Action recognition in compressed domains: A survey
Yue Ming 0001, Jiangwan Zhou, Nannan Hu, Panzi Zhao, Boyang Lyu, Hui Yu 0001
Neurocomputing7
2024 Learning context-aware local feature descriptors for 3D reconstruction
Jian Yang 0036, Hao Fan 0004, Junyu Dong, Hui Yu 0001
Neurocomputing5
2024 Perceptual loss guided Generative adversarial network for saliency detection
Xiaoxu Cai, Gaige Wang, Jianwen Lou, Muwei Jian, Junyu Dong, Rung Ching Chen, Brett Stevens, Hui Yu 0001
Inf. Sci.8
2024 MLNet: An multi-scale line detector and descriptor network for 3D reconstruction
Jian Yang 0036, Yuan Rao 0001, Eric Rigall, Hao Fan 0004, Junyu Dong, Hui Yu 0001
Knowl. Based Syst.7
2024 Paf-tracker: a novel pre-frame auxiliary and fusion visual tracker
Derui Ding, Hui Yu 0001
Mach. Learn.3
2024 Video saliency detection via combining temporal difference and pixel gradient
Xiangwei Lu, Muwei Jian, Rui Wang 0017, Peiguang Lin, Hui Yu 0001
Multim. Tools Appl.6
2024 Differentiable self-supervised clustering with intrinsic interpretability
Zhixiang Jin, Yiqiao Mao, Yangdong Ye, Hui Yu 0001
Neural Networks5
2024 MGEED: A Multimodal Genuine Emotion and Expression Detection Database
abstract
Multimodal emotion recognition has attracted increasing interest from academia and industry in recent years, since it enables emotion detection using various modalities, such as facial expression images, speech and physiological signals. Although research in this field has grown rapidly, it is still challenging to create a multimodal database containing facial electrical information due to the difficulty in capturing natural and subtle facial expression signals, such as optomyography (OMG) signals. To this end, we present a newly developed Multimodal Genuine Emotion and Expression Detection (MGEED) database in this paper, which is the first publicly available database containing the facial OMG signals. MGEED consists of 17 subjects with over 150K facial images, 140K depth maps and different modalities of physiological signals including OMG, electroencephalography (EEG) and electrocardiography (ECG) signals. The emotions of the participants are evoked by video stimuli and the data are collected by a multimodal sensing system. With the collected data, an emotion recognition method is developed based on multimodal signal synchronisation, feature extraction, fusion and emotion prediction. The results show that superior performance can be achieved by fusing the visual, EEG and OMG features. The database can be obtained fromhttps://github.com/YMPort/MGEED.
Yiming Wang 0001, Hui Yu 0001, Weihong Gao, Charles Nduka
IEEE Trans. Affect. Comput.2
2024 Flexible Dual-Branch Siamese Network: Learning Location Quality Estimation and Regression Distribution for Visual Tracking
abstract
Anchor-free based trackers introduce an extra branch in addition to classification and regression branches in the network to achieve comparable performance with anchor-based trackers. This extra branch is usually trained independently in the training phase and is used in combination with other branches in the inference phase. However, this can increase the inconsistency between the inference phase and the training phase, potentially degrading the tracking performance. To address this problem, we propose a new Siamese network-based object tracking framework that eliminates this inconsistency by unifying classification and additional branch tasks to achieve learning location quality estimation. Furthermore, regression tasks for bounding boxes are widely formulated based on Dirac$\delta $distribution. Though this assumption works well for many scenarios, it restricts the prediction of regression branches. To overcome this restriction, we propose discretizing the continuous offset of the regression branch into multiple offset predictions, which enables the network to learn more flexible distributions automatically. Meanwhile, the discrete distribution prediction of regression branches is utilized to further guide the classification of the trackers. Extensive experiments on the widely accepted benchmarks demonstrate the effectiveness and efficiency of the proposed model.
Shuo Hu, Sien Zhou, Jinbo Lu, Hui Yu 0001
IEEE Trans. Comput. Soc. Syst.4
2024 Flow-Edge-Net: Video Saliency Detection Based on Optical Flow and Edge-Weighted Balance Loss
abstract
Optical flow networks have been widely utilized for video saliency detection (VSD) due to their effective performance in capturing the motion of objects. However, the use of optical flow blurs the edges of salient objects and leads to the problems of poorly defined object boundaries. To address this issue, we propose an optical flow-based edge-weighted loss function, to train a network called Flow-Edge-Net, which can balance the weights of the foreground and background information at the edges of video frames. It has achieved superior performance in detecting salient boundaries. Specifically, we propose two complementary encoding and decoding networks based on the concept of decoupling. That is, the optical flow network focuses on moving objects, while the edge network, based on the encoder-decoder structure, focuses on edge information. As the two networks output features of the same dimension and are from the same input, our proposed self-designed adaptive weighted feature fusion module can compare and integrate the edge information and location information from the two networks through adaptive weighting. The proposed method has been evaluated on five widely used databases. Experiment results demonstrate the superior performance of the proposed Flow-Edge-Net in locating salient objects, with accurate and refined edges. The proposed method achieves superior performance over the state-of-the-art methods in detecting salient objects in videos.
Muwei Jian, Xiangwei Lu, Yakun Ju, Hui Yu 0001, Kin-Man Lam 0001
IEEE Trans. Comput. Soc. Syst.5
2024 Unsupervised Video Summarization Based on the Diffusion Model of Feature Fusion
abstract
Video summarization (VS) technologies can automatically extract key frames with effective information and thus can help to quickly identify the events or speed up the decision-making process, especially for accidents. With the fast development of deep learning technologies, many generative adversarial network (GAN)- and reinforcement learning (RL)-based unsupervised VS methods have been developed in recent years. However, these methods could suffer from the problems of unstable training and difficulty of reward function formulation, respectively. To this end, we present an unsupervised VS method called diffusion model of feature fusion (DMFF) in this article, which consists of a diffusion module (DM), a feature extraction and compression module (FECM), and a coarse-fine frame selector (CFFS). DM is designed to avoid the training instability problem caused by GAN’s alternate training generator and discriminator. FECM is used to extract and compress video features. CFFS is designed to capture both low-level and high-level features between frames to handle complex and diverse accident videos. Then, high-level local and global features are fused to generate a multigrained final frame score. Experiments on two widely used benchmark datasets, SumMe and TVSum, demonstrate the effectiveness and superiority of the proposed network to the state-of-the-art methods, and the training is more stable.
Qinghao Yu, Hui Yu 0001, Ying Sun 0004, Derui Ding, Muwei Jian
IEEE Trans. Comput. Soc. Syst.2
2024 UniFRD: A Unified Method for Facial Image Restoration Based on Diffusion Probabilistic Model
abstract
This paper presents a Unified Facial image and video Restoration method based on the Diffusion probabilistic model (UniFRD), designed to effectively address both single- and multi-type image degradation. The noise predictor in UniFRD consists of a ViT-based encoder and a novel Separation Fusion Decoding Module (SFDM). The flexible feature optimization strategy allows for decoding complex conditional noise without being limited by degradation patterns. Specifically, SFDM adjusts and refines the channel correlation and expressive power of high-dimensional features step by step, enabling the network to more accurately perceive and enhance the interaction between posterior probabilities and conditional inputs. This process is crucial for improving the visual quality and stability of the restoration results. Extensive experiments demonstrate that even when facial images suffer from both pixel-level and image-level degradation, UniFRD can still guarantee the restoration of rich details and maintain attribute consistency. In summary, compared to existing methods, the solution proposed in this study for facial restoration tasks offers greater generality and adaptability. Moreover, it has high practical value for applications involving faces in complex and unconstrained outdoor scenarios.
Muwei Jian, Rui Wang 0199, Feng Xu 0005, Hui Yu 0001, Kin-Man Lam 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Multi-Branch GAN-Based Abnormal Events Detection via Context Learning in Surveillance Videos
abstract
Video anomaly detection is an important task in the field of intelligent security. However, existing methods mainly detect and analyze videos from a single time direction, ignoring the semantic information of the video context, which adversely affects the detection accuracy. To address this issue, we design a multi-branch generative adversarial network with context learning (MGAN-CL) to detect abnormal events. In particular, we combine video context information to generate predicted frames, and determine whether an anomaly occurs by comparing the predicted frame with the actual frame. Different from the existing GAN-based methods, in the anomaly event detection stage, we use the discriminator to judge the video frames generated by the generator, which improves the accuracy of anomaly detection. In order to improve the ability of the discriminator, a pseudo-anomaly module is added to the discriminator for data augmentation to improve the robustness of the model. An extensive set of experiments performed on public datasets demonstrate the method’s superior performance.
Daoheng Li, Xiushan Nie, Ximing Lin, Hui Yu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Multitask Image Clustering via Deep Information Bottleneck
abstract
Multitask image clustering approaches intend to improve the model accuracy on each task by exploring the relationships of multiple related image clustering tasks. However, most existing multitask clustering (MTC) approaches isolate the representation abstraction from the downstream clustering procedure, which makes the MTC models unable to perform unified optimization. In addition, the existing MTC relies on exploring the relevant information of multiple related tasks to discover their latent correlations while ignoring the irrelevant information between partially related tasks, which may also degrade the clustering performance. To tackle these issues, a multitask image clustering method named deep multitask information bottleneck (DMTIB) is devised, which aims at conducting multiple related image clustering by maximizing the relevant information of multiple tasks while minimizing the irrelevant information among them. Specifically, DMTIB consists of a main-net and multiple subnets to characterize the relationships across tasks and the correlations hidden in a single clustering task. Then, an information maximin discriminator is devised to maximize the mutual information (MI) measurement of positive samples and minimize the MI of negative ones, in which the positive and negative sample pairs are constructed by a high-confidence pseudo-graph. Finally, a unified loss function is devised for the optimization of task relatedness discovery and MTC simultaneously. Empirical comparisons on several benchmark datasets, NUS-WIDE, Pascal VOC, Caltech-256, CIFAR-100, and COCO, show that our DMTIB approach outperforms more than 20 single-task clustering and MTC approaches.
Yiqiao Mao, Yangdong Ye, Hui Yu 0001
IEEE Trans. Cybern.5
2024 DomainForensics: Exposing Face Forgery Across Domains via Bi-Directional Adaptation
abstract
Recent DeepFake detection methods have shown excellent performance on public datasets but are significantly degraded on new forgeries. Solving this problem is important, as new forgeries emerge daily with the continuously evolving generative techniques. Many efforts have been made for this issue by seeking the commonly existing traces empirically on data level. In this paper, we rethink this problem and propose a new solution from the unsupervised domain adaptation perspective. Our solution, called DomainForensics, aims to transfer the forgery knowledge from known forgeries (fully labeled source domain) to new forgeries (label-free target domain). Unlike recent efforts, our solution does not focus on data view but on learning strategies of DeepFake detectors to capture the knowledge of new forgeries through the alignment of domain discrepancies. In particular, unlike the general domain adaptation methods which consider the knowledge transfer in the semantic class category, thus having limited application, our approach captures the subtle forgery traces. We describe a new bi-directional adaptation strategy dedicated to capturing the forgery knowledge across domains. Specifically, our strategy considers both forward and backward adaptation, to transfer the forgery knowledge from the source domain to the target domain in forward adaptation and then reverse the adaptation from the target domain to the source domain in backward adaptation. In forward adaptation, we perform supervised training for the DeepFake detector in the source domain and jointly employ adversarial feature adaptation to transfer the ability to detect manipulated faces from known forgeries to new forgeries. In backward adaptation, we further improve the knowledge transfer by coupling adversarial adaptation with self-distillation on new forgeries. This enables the detector to expose new forgery features from unlabeled data and avoid forgetting the known knowledge of known forgery. Extensive experiments demonstrate that our method is surprisingly effective in exposing new forgeries, and can be plug-and-play on other DeepFake detection architectures.
Qingxuan Lv, Yuezun Li, Junyu Dong, Sheng Chen 0001, Hui Yu 0001, Huiyu Zhou 0001, Shu Zhang 0002
IEEE Trans. Inf. Forensics Secur.5
2024 Coupling Effect and Chain Evolution of Urban Rail Transit Emergencies
abstract
Emergency events such as fire, flood and COVID-19 occurred in urban rail transit (URT) usually triggered chain effect and evolved into huge disaster. This kind of chain with complexity and uncertainty evolution brought great challenges to the safety management of the system. Thus, the coupling effects of emergencies and then its relationship with the chain evolution is necessary to analyze emphatically. A Graph Evaluation and Review Technique Simulation (GERTS) evolution network is firstly constructed to describe the coupling effect and chain evolution of emergencies. Then, considering the internal and external influencing factors of the emergency chain, a dynamic evolution model of the emergency chain based on Coupled Map Lattice (CML) is proposed. This paper takes fire chain of URT as an example to simulate the evolution process of emergency chain, and analyze the impact of different coupling effects and various influencing factors on the evolution of emergency chain. The results of numerical simulation show that the AND-coupling can significantly inhibit the evolution of emergency events, while the OR-coupling and CO-coupling can expand the impact scope of emergency events. In addition, the evolution speed of emergency events can be controlled by increasing the coupling action time and improving the URT repair ability. When an emergency event occurs, the analysis of coupling effect and the accompanied chain evolution will help managers to make scientific judgment on the development trend of the emergency events and make targeted emergency defense measures.
Guangyu Zhu 0001, Ranran Sun, Yuhong Hou, Hui Yu 0001, Peter Xiaoping Liu
IEEE Trans. Intell. Transp. Syst.6
2024 ICE-YoloX: research on face mask detection algorithm based on improved YoloX network
Yinggan Tang, Hui Yu 0001
J. Supercomput.4
2024 CSS-Net: A Consistent Segment Selection Network for Audio-Visual Event Localization
abstract
Audio-visual event (AVE) localization aims to localize the temporal boundaries of events that contains visual and audio contents, to identify event categories in unconstrained videos. Existing work usually utilizes successive video segments for temporal modeling. However, ambient sounds or irrelevant visual targets in some segments often cause the problem of audio-visual semantics inconsistency, resulting in inaccurate global event modeling. To tackle this issue, we present a consistent segment selection network (CSS-Net) in this paper. First, we propose a novel bidirectional guided co-attention (BGCA) block, containing two distinct attention paths from audio to vision and from vision to audio, to focus on sound-related visual regions and event-related sound segments. Then, we propose a novel context-aware similarity measure (CASM) module to select semantic consistent visual and audio segments. A cross-correlation matrix is constructed using the correlation coefficients between the visual and audio feature pairs in all time steps. By extracting highly correlated segments and discarding low correlated segments, visual and audio features can learn global event semantics in videos. Finally, we propose a novel audio-visual contrastive loss to learn the similar semantics representation for visual and audio global features under the constraints of cosine and L2 similarities. Extensive experiments on public AVE dataset demonstrates the effectiveness of our proposed CSS-Net. The localization accuracies achieve the best performance of 80.5% and 76.8% in both fully- and weakly-supervised settings compared with other state-of-the-art methods.
Yue Ming 0001, Nannan Hu, Hui Yu 0001
IEEE Trans. Multim.4
2024 Cross-Modal Clustering With Deep Correlated Information Bottleneck Method
abstract
Cross-modal clustering (CMC) intends to improve the clustering accuracy (ACC) by exploiting the correlations across modalities. Although recent research has made impressive advances, it remains a challenge to sufficiently capture the correlations across modalities due to the high-dimensional nonlinear characteristics of individual modalities and the conflicts in heterogeneous modalities. In addition, the meaningless modality-private information in each modality might become dominant in the process of correlation mining, which also interferes with the clustering performance. To tackle these challenges, we devise a novel deep correlated information bottleneck (DCIB) method, which aims at exploring the correlation information between multiple modalities while eliminating the modality-private information in each modality in an end-to-end manner. Specifically, DCIB treats the CMC task as a two-stage data compression procedure, in which the modality-private information in each modality is eliminated under the guidance of the shared representation of multiple modalities. Meanwhile, the correlations between multiple modalities are preserved from the aspects of feature distributions and clustering assignments simultaneously. Finally, the objective of DCIB is formulated as an objective function based on a mutual information measurement, in which a variational optimization approach is proposed to ensure its convergence. Experimental results on four cross-modal datasets validate the superiority of the DCIB. Code is released at https://github.com/Xiaoqiang-Yan/DCIB.
Yiqiao Mao, Yangdong Ye, Hui Yu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Multimodal Perception and Decision-Making Systems for Complex Roads Based on Foundation Models
abstract
Since the inception of Industry 5.0 in 2021, a growing number of researchers have begun to pay their attention to the revolutionary shift it brings. The principles of Industry 5.0, including human-centric, sustainability, and emphasis on ecological and social values, will become the new paradigm for future industrial development. In this transformative landscape, artificial intelligence (AI) plays a pivotal role, and foundation models based on ChatGPT are set to reshape the organizational structure of industries. In this article, we introduce a multimodal perception and decision-making system built upon a foundational model. This system integrates image and point cloud data to enhance perception accuracy and provide ample information for decision making. It is designed to achieve a deep integration of AI and human-centric autonomous driving within the context of Industry 5.0. We introduce a cross-domain learning approach in the system architecture, along with a model training method from foundation models to handle complex road conditions. The proposed method enables road drivable area segmentation on complex unstructured roads. To address the issue of increased variance caused by the residual structure employed in previous works, this article introduces a distribution correction module, which effectively mitigates this problem. Furthermore, to achieve high-performance perception systems in intricate road scenarios, we put forth a multimodal perception fusion method in this study. The experiments demonstrate the superiority of this approach over single-sensor perception. This work contributes to the ongoing discourse on the convergence of AI, human-centric values, and advanced driving systems within the framework of Industry 5.0.
Lili Fan, Yutong Wang 0001, Hui Zhang 0091, Changxian Zeng, Yunjie Li, Chao Gou, Hui Yu 0001
IEEE Trans. Syst. Man Cybern. Syst.7
2023 PAF-Tracker: A Novel Pre-Frame Auxiliary and Fusion Visual Tracker
abstract
Relying on a large amount of data, recent object trackers achieve superior performance. However Siamese-like trackers expose considerable shortcomings in the case of brief occlusion. To address these shortages, the paper proposes a novel pre-frame auxiliary and fusion tracking framework. Within this framework, a retained variable is first introduced to avoid some additional twin branches while retaining the previously obtained deep features of the search frames. Based on such a variable, a pre-frame auxiliary module is constructed to establish the relationship between encoding features and the retained pre-frame information and a decoding fusion module is designed to fuse the generated similarity relationship. Moreover, the Efficient IoU (EIoU) loss is employed to increase the precision of predicted bounding boxes by adding three penalty terms for the differences in the center point, length, and width of the two bounding boxes. Finally, the superiority over state-of-the-art methods is verified by numerous tests on visual tracking benchmarks.
Derui Ding, Hui Yu 0001
DSAA3
2023 ICE-YoloX: An Effective Face Mask Detection Method
abstract
Deep learning technologies such as YoloX have achieved impressive progress in face mask detection recently. However, the neck network used in YoloX network may lead to severe confounding effect in feature mapping due to the inherent defect of channel reduction in hybrid fusion, which affects its precise localization ability of mask-wearing targets. To tackle this issue, we present a new FPN network structure (ICE-FPN) based on channel-enhanced feature pyramid network (CE-FPN) in this paper, which can mitigate the YoloX network confounding effect while reducing the number of parameters and computational effort caused by CE-FPN. Experiments conducted on the WMD dataset show that the mAP0.5 of the model improves from 99.54% to 99.62% and the mAP0.75 improves from 89.47% to 91.35%. The ablation and comparison experiments demonstrate that the proposed ICE-YoloX has achieved superior performance over existing methods.
Yinggan Tang, Hui Yu 0001
SMC4
2023 Inter-slice Correlation Weighted Fusion for Universal Lesion Detection
abstract
Universal lesion detection using computerised tomography (CT) scans is a critical computer-aided diagnosis measure in clinical diagnosis. One of the key issues during the diagnosis is to identify the correlations between sequential slices to improve the feature representation of CT scans. In the process of fusing slice features containing temporal correlations, the correlation between the contextual slices in the channel dimension and the target slices is closely related to the spatial distance in practice. However, convolutional fusion approaches commonly ignore that features of different distances have unequal weights. To tackle this issue, we present a temporal correlation weighted fusion lesion detection network, called TCW-Net. Specifically, for the slices in the channel dimension, we develop a weighted feature fusion module to adjust the more discriminative features using learned weights. Then, we adapt a spatial offset attention mechanism that allows the detection network to pay more attention to the lesion’s slight spatial offset and thus improve the model’s capacity for distinguishing between different lesion features. Extensive experiments carried out on the DeepLesion dataset show that the proposed algorithm has superior performance over the state-of-the-art methods.
Muwei Jian, Rui Wang 0017, Hui Yu 0001
TrustCom5
2023 Object pose estimation based on stereo vision with improved K-D tree ICP algorithm
abstract
Summary With the wide application of stereovision in SLAM, object pose estimation has gradually become one of the research hotspots. This article proposes an object pose estimation for robotic grasping based on stereo vision with improved K‐D tree ICP algorithm. The feature points and feature descriptors of the point cloud of the object to be captured are extracted, and the feature template set is established. The SAC‐IA algorithm is used to carry out initial registration of the point cloud of the target, and the ICP algorithm based on K‐D tree is used for fine registration. The experimental results show that the average coincidence degree of the final registration of the proposed object pose estimation method reaches 94.1%, and the accurate 6D pose of the object to be grasped is obtained.
Juntong Yun, Bo Tao 0002, Jinxian Qi, Ying Liu 0087, Hongjie Ma, Hui Yu 0001
Concurr. Comput. Pract. Exp.8
2023 LACN: A lightweight attention-guided ConvNeXt network for low-light image enhancement
Saijie Fan, Derui Ding, Hui Yu 0001
Eng. Appl. Artif. Intell.4
2023 Recurrent attention unit: A new gated recurrent unit for long-term memory of important parts in sequential data
Zhaoyang Niu, Guoqiang Zhong 0001, Guohua Yue, Li-Na Wang, Hui Yu 0001, Junyu Dong
Neurocomputing5
2023 ENSO analysis and prediction using deep learning: A review
abstract
El Niño/Southern Oscillation (ENSO) mainly occurs in the tropical Pacific Ocean every a few years. But it affects the climate around the world and has a dramatic impact on the development of ecology and agriculture. The analysis and prediction of ENSO become particularly important for meteorology and disaster management. However, due to insufficient data, spring predictability barrier (SPB), and model uncertainty, traditional analysis models face challenges. To address these issues, researchers begin to apply deep learning (DL) technologies to ENSO research, exploring the impact of ENSO on the world's extreme climate changes. In recent years, deep learning-based methods have obtained impressive progress with more accurate and effective predictions of ENSO. In this paper, we summarize the attempts of DL technologies in predicting ENSO. We first introduce the properties of ENSO, followed by the architecture introduction of DL technologies and their application to ENSO. We then investigate the potential of DL technologies for ENSO prediction from various aspects, including model evaluation metrics, prediction algorithms, overcoming SPB and prediction uncertainty. Finally, we provide discussions on the future trends and challenges of using DL technologies for ENSO prediction.
Gaige Wang, Honglei Cheng, Hui Yu 0001
Neurocomputing4
2023 Unsupervised video summarization using deep Non-Local video summarization networks
Sha-Sha Zang, Hui Yu 0001, Yan Song 0002, Ru Zeng
Neurocomputing2
2023 Robust seed selection of foreground and background priors based on directional blocks for saliency-detection system
Muwei Jian, Ruihong Wang, Hui Yu 0001, Junyu Dong, Gongfa Li, Yilong Yin, Kin-Man Lam 0001
Multim. Tools Appl.4
2023 Guest Editorial Special Issue on Behavioral Modeling, Learning, and Adaptation in Cyber-Physical-Social Intelligence
abstract
The integration of artificial intelligence (AI) with cyber–physical–social systems (CPSS) creates new research opportunities and challenges with major societal implications. The behavioral and cognitive enhancement of intelligent systems promotes a productive and creative partnership and collaboration between humans and machines. Advancements in these areas enable adaptability, scalability, resiliency, safety, security, and usability that expand the horizons of CPSS.
Ying Tang 0001, Jiacun Wang 0001, Hui Yu 0001, Giancarlo Fortino, Fei-Yue Wang 0001, Amir Hussain 0001
IEEE Trans. Comput. Soc. Syst.3
2023 AFLFP: A Database With Annotated Facial Landmarks for Facial Palsy
abstract
Facial landmark detection is a crucial step for the task of computer-aided facial palsy diagnosis, which enables to focus on the affected facial regions for learning asymmetry, shape, and texture features of facial palsy. However, it is still very challenging to accurately detect salient landmarks on facial palsy images due to the unavailability of sufficient training databases providing annotated facial landmark images of facial palsy. To this end, we present a database in this article named annotated facial landmarks for facial palsy (AFLFP). AFLFP is a diverse, and reliable database that contains facial images with 16-class facial expressions of asymmetric facial expressions from 88 subjects. Each facial image is independently and manually annotated with 68 facial landmarks. This database is the first public manually annotated facial landmark database for facial palsy so far. Furthermore, to establish the benchmark results for the proposed database, we propose a deep neural network (DNN) baseline with a two-stage cascaded fully convolutional network (FCN), which can detect facial landmarks in facial palsy from coarse to fine. The comprehensive experiments show that the proposed method performs better than the mainstream methods of machine- and deep-learning. And we have also compared the performance using normal- and palsy-faces, respectively, as the training data. The comparison results show that there are significant differences between them in terms of facial landmark detection, which further confirms the necessity to develop a facial landmark database specifically for facial palsy.
Charles Nduka, Ruben Yap Kannan, Elena Pescarini, Juan Enrique Berner, Hui Yu 0001
IEEE Trans. Comput. Soc. Syst.6
2023 LMFNet: A Lightweight Multiscale Fusion Network With Hierarchical Structure for Low-Quality 3-D Face Recognition
abstract
Three-dimensional (3-D) face recognition (FR) can improve the usability and user-friendliness of human–machine interaction. In general, 3-D FR can be divided into high-quality and low-quality 3-D FR according to different interaction scenarios. The low-quality data can be easily obtained, so its application prospect is more extensive. However, the challenge is how to balance the trade-offs between data accuracy and real-time performance. To solve this problem, we propose a lightweight multiscale fusion network (LMFNet) with a hierarchical structure based on single-mode data for low-quality 3-D FR. First, we design a backbone network with only five feature extraction blocks to reduce computational complexity and improve the inference speed. Second, we devise a mid-low adjacent layer with a multiscale feature fusion (ML-MSFF) module to extract the facial texture and contour information, and a mid-high adjacent layer with a multiscale feature fusion (MH-MSFF) module to obtain the discriminative information in high-level features. Then, a hierarchical multiscale feature fusion (HMSFF) module is formed by combining these two modules mentioned above to acquire the local information of different scales. Finally, we enhance the expression of features by integrating HMSFF with a global convolutional neural network for improving recognition accuracy. Experiments on Lock3DFace, KinectFaceDB, and IIIT-D datasets demonstrate that our proposed LMFNet can achieve superior performance on low-quality datasets. Furthermore, experiments on the cross-quality database based on Bosphorus and the different intensity noise low-quality datasets based on UMB-DB and Bosphorus show that our network is robust and has a high generalization ability. It satisfies the real-time requirement, which lays a foundation for a smooth and user-friendly interactive experience.
Panzi Zhao, Yue Ming 0001, Xuyang Meng, Hui Yu 0001
IEEE Trans. Hum. Mach. Syst.4
2022 Face Super-resolution Based on Multi-source References
abstract
This paper proposes a multi-source references (MSR) based face super-resolution (FSR) model. More specifically, to enhance the low-quality large-scale reconstruction of faces without the involvement of face prior knowledge, we propose a multi-source references based FSR framework exploiting a constructed reference library of nonidentity faces and an information mining module for external and internal references. Experimental results show that the proposed model can provide more satisfactory and reliable face super-resolution results than the-state-of-the-art methods.
Rui Wang 0017, Muwei Jian, Paul Smith 0002, Hui Yu 0001
HSI4
2022 Deep correlation mining for multi-task image clustering
Kaiyuan Shi, Yangdong Ye, Hui Yu 0001
Expert Syst. Appl.4
2022 Face hallucination using multisource references and cross-scale dual residual fusion mechanism
abstract
There is an increasing interest in enhancing the quality of low-resolution (LR) facial images for various social life applications. Existing methods often use domain-specific prior knowledge, which is effective in improving the face super-resolution model's performance. However, it is challenging to obtain rich and accurate prior information from LR inputs in real-world scenarios, which can limit the robustness and generalization ability of the developed face super-resolution model. In this paper, a multisource reference-based face super-resolution Network, namely MSRNet, is proposed. Without considering the prior knowledge of faces, the network can reconstruct a LR face image with a magnitude factor of 8 under the guidance of multiple reference face images of different identities. By constructing an “appearance-alike” reference data set Face_Ref, the designed MSRNet aims to fully exploit the local and spatially similar high frequency information between the distinct references and the current face. More specifically, to effectively combine the information from multiple references, a cross-scale and cross-space feature fusion mechanism is introduced for external and internal references, and then the enhanced local semantics are finally incorporated into the high-resolution face reconstruction. The robustness of face image super-resolution is increased compared to current correlation approaches, since it not only eliminates the need for face prior knowledge but also avoids performing alignment operations on reference faces with multiple expressions and different poses. Experimental results show that the proposed model is able to produce results for face super-resolution that are satisfying and dependable and outperforms the state-of-the-art methods in terms of visual perceptual quality and quantity evaluation.
Rui Wang 0017, Muwei Jian, Hui Yu 0001, Lin Wang 0004, Bo Yang 0001
Int. J. Intell. Syst.3
2022 A comprehensive survey on robust image watermarking
Wenbo Wan, Jun Wang 0061, Yunming Zhang, Jing Li 0046, Hui Yu 0001, Jiande Sun 0001
Neurocomputing5
2022 Explanation guided cross-modal social image clustering
Yiqiao Mao, Yangdong Ye, Hui Yu 0001, Fei-Yue Wang 0001
Inf. Sci.4
2022 Multi-instance semantic similarity transferring for knowledge distillation
Xin Sun 0003, Junyu Dong, Hui Yu 0001, Gaige Wang
Knowl. Based Syst.4
2022 Visual saliency detection via combining center prior and U-Net
Xiangwei Lu, Muwei Jian, Xing Wang 0002, Hui Yu 0001, Junyu Dong, Kin-Man Lam 0001
Multim. Syst.4
2022 Deep convolutional neural network for enhancing traffic sign recognition developed on Yolo V4
Christine Dewi, Rung Ching Chen, Xiaoyi Jiang 0001, Hui Yu 0001
Multim. Tools Appl.4
2022 Face image-sketch synthesis via generative adversarial fusion
Jianyuan Sun, Hongchuan Yu, Jian J. Zhang 0001, Junyu Dong, Hui Yu 0001, Guoqiang Zhong 0001
Neural Networks5
2022 CORNet: Context-Based Ordinal Regression Network for Monocular Depth Estimation
abstract
Monocular depth estimation, as one of the fundamental tasks of computer vision, plays a crucial role in three-dimensional (3D) scene understanding and perception. Usually, deep learning methods recover monocular depth maps using continuous regression manners by minimizing the errors between the ground-truth depth and the predicted depth. However, fine depth features may not be fully captured through layer-by-layer coding, which is prone to low spatial resolution depth maps and insufficient details. Furthermore, it usually converges slowly and suffers from unsatisfactory results. To tackle these issues, we propose a novel model, named context-based ordinal regression network (CORNet), to reconstruct monocular depth maps in the ordinal regression manner with context information in this paper. Firstly, we put forward a novel context-based encoder with a feature transformation (FT) module to learn context information and details from inputs, and output multi-scale feature maps. Then, we design a boundary enhancement module (BEM) with a spatial attention mechanism following each operation of feature fusion, which captures boundary features in the scene to enhance the border depth. Finally, a feature optimization module (FOM) is designed to fuse and optimize the multi-scale features and boundary features to strengthen depth learning. What’s more, we introduce an ordinal weighted inference to predict depth maps from probabilities and discretization values. Experiments and results on two challenging datasets, KITTI and NYU Depth V2, demonstrate that our proposed CORNet can estimate monocular depth maps effectively and obtain superior performance in capturing geometric features over existing methods.
Xuyang Meng, Chunxiao Fan 0001, Yue Ming 0001, Hui Yu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Local and Global Perception Generative Adversarial Network for Facial Expression Synthesis
abstract
Facial expression synthesis has gained increasing attention with the development of Generative Adversarial Networks (GANs). However, it is still very challenging to generate high-quality facial expressions since the overlapping and blur commonly appear in the generated facial images especially in the regions with rich facial features such as eye and mouth. Generally, existing methods mainly consider the face as a whole in facial expression synthesis without paying specific attention to the characteristics of facial expressions. In fact, according to the physiological and psychological research, the differences of facial expressions often appear in crucial regions such as eye and mouth. Motivated by this observation, a novel end-to-end facial expression synthesis method called Local and Global Perception Generative Adversarial Network (LGP-GAN) with a two-stage cascaded structure is proposed in this paper which is designed to extract and synthesize the details of the crucial facial regions. LGP-GAN can combine the generated results from the global network and local network into the corresponding facial expressions. In Stage I, LGP-GAN utilizes local networks to capture the local texture details of the crucial facial regions and generate local facial regions, which fully explores crucial facial region domain information in facial expressions. And then LGP-GAN uses a global network to learn the whole facial information in Stage II to generate the generate final facial expressions building upon local generated results from Stage I. We conduct qualitative and quantitative experiments on the commonly used public database to verify the effectiveness of the proposed method. Experimental results show the superiority of the proposed method over the state-of-the-art methods.
Wenbo Zheng 0001, Yiming Wang 0001, Hui Yu 0001, Junyu Dong, Fei-Yue Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Random Shapley Forests: Cooperative Game-Based Random Forests With Consistency
abstract
The original random forests (RFs) algorithm has been widely used and has achieved excellent performance for the classification and regression tasks. However, the research on the theory of RFs lags far behind its applications. In this article, to narrow the gap between the applications and the theory of RFs, we propose a new RFs algorithm, called random Shapley forests (RSFs), based on the Shapley value. The Shapley value is one of the well-known solutions in the cooperative game, which can fairly assess the power of each player in a game. In the construction of RSFs, RSFs use the Shapley value to evaluate the importance of each feature at each tree node by computing the dependency among the possible feature coalitions. In particular, inspired by the existing consistency theory, we have proved the consistency of the proposed RFs algorithm. Moreover, to verify the effectiveness of the proposed algorithm, experiments on eight UCI benchmark datasets and four real-world datasets have been conducted. The results show that RSFs perform better than or at least comparable with the existing consistent RFs, the original RFs, and a classic classifier, support vector machines.
Jianyuan Sun, Hui Yu 0001, Guoqiang Zhong 0001, Junyu Dong, Shu Zhang 0002, Hongchuan Yu
IEEE Trans. Cybern.2
2022 SSGNN: A Macro and Microfacial Expression Recognition Graph Neural Network Combining Spatial and Spectral Domain Features
abstract
Emotion recognition from macroexpression and microexpression has been widely used in applications such as human–computer interaction, learning status evaluation, and mental disorder diagnosis. However, due to the complexity of human macroexpressions, recognizing macroexpressions with high accuracy is a challenging task. Moreover, the short duration and low movement intensity of microexpressions make its recognition more difficult. For MM-FER (macro and microfacial expression recognition), the key information can be more efficiently expressed by a graph. In this article, a novel framework based on graph neural network named SSGNN (spatial and spectral domain features based on a graph neural network) is designed to extract spatial and spectral domain features from facial images for MM-FER, which can efficiently recognize both macroexpressions and microexpressions under the same model. SSGNN consists of two parts, SPAGNN and SPEGNN, which are used to extract spectral and spatial domain features, respectively. Experiments proved that jointly using the spectral and spatial information extracted by SSGNN can largely improve the performance of MM-FER when the training sample is limited. First, the influences of different neighbors and samples to the model performance was analyzed. Then, the contribution of SPAGNN and SPEGNN were evaluated. It was discovered that fusing the result of SPAGNN and SPEGNN at decision level further improved the performance of MM-FER. Experiment proved that SSGNN can recognize microexpression acquired by various sensors with higher accuracy under different image resolutions and image formats than the compared state-of-the-art methods in most cases. A cross-dataset experiment demonstrated the generalization ability of SSGNN.
Junjie Zhang 0006, Guangmin Sun, Sarah Mazhar, Xiaohui Fu, Yu Li 0009, Hui Yu 0001
IEEE Trans. Hum. Mach. Syst.7
2022 Real-Time Performance-Focused Localization Techniques for Autonomous Vehicle: A Review
abstract
Real-time, accurate, and robust localisation is critical for autonomous vehicles (AVs) to achieve safe, efficient driving, whilst real-time performance is essential for AVs to achieve their current position in time for decision making. To date, no review paper has quantitatively compared the real-time performance between different localisation techniques based on various hardware platforms and programming languages and analysed the relations among localisation methodologies, real-time performance and accuracy. Therefore, this paper discusses the state-of-the-art localisation techniques and analyses their overall performance in AV application. For further analysis, this paper firstly proposes a localisation algorithm operations capability (LAOC)-based equivalent comparison method to compare the relative computational complexity of different localisation techniques; then, it comprehensively discusses the relations among methodologies, computational complexity, and accuracy. Analysis results show that the computational complexity of localisation approaches differs by a maximum of about 107times, whilst accuracy varies by about 100 times. Vision- and data fusion-based localisation techniques have about 2–5 times potential for improving accuracy compared with lidar-based localisation. Lidar- and vision-based localisation can reduce computational complexity by improving image registration method efficiency. Data fusion-based localisation can achieve better real-time performance compared with lidar- and vision-based localisation because each standalone sensor does not need to develop a complex algorithm to achieve its best localisation potential. Vehicle-to-everything (V2X) technology can improve positioning robustness. Finally, the potential solutions and future orientations of AVs’ localisation based on the quantitative comparison results are discussed.
Yongqiang Lu 0002, Hongjie Ma, Edward Smart, Hui Yu 0001
IEEE Trans. Intell. Transp. Syst.4
2022 Context-Aware Dynamic Feature Extraction for 3D Object Detection in Point Clouds
abstract
Varying density of point clouds increases the difficulty of 3D detection. In this paper, we present a context-aware dynamic network (CADNet) to capture the variance of density by considering both point context and semantic context. Point-level contexts are generated from original point clouds to enlarge the effective receptive filed. They are extracted around the voxelized pillars based on our extended voxelization method and processed with the context encoder in parallel with the pillar features. With a large perception range, we are able to capture the variance of features for potential objects and generate attentive spatial guidance to help adjust the strengths for different regions. In the region proposal network, considering the limited representation ability of traditional convolution where same kernels are shared among different samples and positions, we propose a decomposable dynamic convolutional layer to adapt to the variance of input features by learning from the local semantic context. It adaptively generates the position-dependent coefficients for multiple fixed kernels and combines them to convolve with local features. Based on our dynamic convolution, we design a dual-path convolution block to further improve the representation ability. We conduct experiments on KITTI dataset and the proposed CADNet has achieved superior performance of 3D detection outperforming SECOND and PointPillars by a large margin at the speed of 30 FPS.
Yonglin Tian, Lichao Huang, Hui Yu 0001, Xiangbin Wu, Xuesong Li 0004, Kunfeng Wang, Zilei Wang, Fei-Yue Wang 0001
IEEE Trans. Intell. Transp. Syst.3
2021 Model guided DLP 3D printing for solid and hollow structure
abstract
Manufacturing speed is one of the biggest challenges in 3D printing. Continuous stereolithography printing can effectively improve the printing speed. However, the model it can print is limited to the hollow out structure or flake structure. In the real application scenario, the models are the composition of multiple kinds of structures such as solid structure, hollow out structure, or flake structure. The continuous stereolithography printing scheme will not work for such models. We first propose a concept of maximum fillable distance (MFD) for a set of resin material and printing settings. And for a specific kind of printing setting, the MFD of resin material at different moving speeds is estimated by experiments. Furthermore, the max-min distance of each slice of the model is computed. And a printing control scheme to combining the continuous and layer-wise printing is generated automatically by comparing the max-min distance and MFD. Using the printing control scheme, two real models are successfully printed.
Zechao Liu, Yandong Li, Lifang Wu, Kejian Cui, Hui Yu 0001
HSI6
2021 Visual saliency detection by integrating spatial position prior of object with background cues
Muwei Jian, Hui Yu 0001, Guodong Wang 0001, Xianjing Meng, Lu Yang 0005, Junyu Dong, Yilong Yin
Expert Syst. Appl.3
2021 Cosine metric supervised deep hashing with balanced similarity
Wenjin Hu 0002, Lifang Wu, Meng Jian, Hui Yu 0001
Neurocomputing5
2021 Deep learning for monocular depth estimation: A review
Yue Ming 0001, Xuyang Meng, Chunxiao Fan 0001, Hui Yu 0001
Neurocomputing4
2021 A review on the attention mechanism of deep learning
Zhaoyang Niu, Guoqiang Zhong 0001, Hui Yu 0001
Neurocomputing3
2021 Binary thresholding defense against adversarial attacks
Yutong Wang 0001, Tianyu Shen, Hui Yu 0001, Fei-Yue Wang 0001
Neurocomputing4
2021 Deep multi-view learning methods: A review
Shizhe Hu, Yiqiao Mao, Yangdong Ye, Hui Yu 0001
Neurocomputing5
2021 Crowd emotion evaluation based on fuzzy inference of arousal and valence
Xiuxin Yang, Gongfa Li, Hui Yu 0001
Neurocomputing5
2021 Integrating object proposal with attention networks for video saliency detection
Muwei Jian, Jiaojin Wang, Hui Yu 0001, Gaige Wang
Inf. Sci.3
2021 Linearly augmented real-time 4D expressional face capture
Shu Zhang 0002, Hui Yu 0001, Ting Wang 0018, Junyu Dong, Tuan D. Pham
Inf. Sci.2
2021 DSFMA: deeply supervised fully convolutional neural networks based on multi-level aggregation for saliency detection
Inam Ullah 0002, Muwei Jian, Sumaira Hussain, Jie Guo 0012, Li Lian, Hui Yu 0001, Kashif Shaheed, Yilong Yin
Multim. Tools Appl.6
2021 Underwater image processing and analysis: A review
Muwei Jian, Hanjiang Luo, Xiangwei Lu, Hui Yu 0001, Junyu Dong
Signal Process. Image Commun.5
2021 Real-Time 3D Facial Tracking via Cascaded Compositional Learning
abstract
We propose to learn a cascade of globally-optimized modular boosted ferns (GoMBF) to solve multi-modal facial motion regression for real-time 3D facial tracking from a monocular RGB camera. GoMBF is a deep composition of multiple regression models with each is a boosted ferns initially trained to predict partial motion parameters of the same modality, and then concatenated together via a global optimization step to form a singular strong boosted ferns that can effectively handle the whole regression target. It can explicitly cope with the modality variety in output variables, while manifesting increased fitting power and a faster learning speed comparing against the conventional boosted ferns. By further cascading a sequence of GoMBFs (GoMBF-Cascade) to regress facial motion parameters, we achieve competitive tracking performance on a variety of in-the-wild videos comparing to the state-of-the-art methods which either have higher computational complexity or require much more training data. It provides a robust and highly elegant solution to real-time 3D facial tracking using a small set of training data and hence makes it more practical in real-world applications. We further deeply investigate the effect of synthesized facial images on training non-deep learning methods such as GoMBF-Cascade for 3D facial tracking. We apply three types synthetic images with various naturalness levels for training two different tracking methods, and compare the performance of the tracking models trained on real data, on synthetic data and on a mixture of data. The experimental results indicate that, i) the model trained purely on synthetic facial imageries can hardly generalize well to unconstrained real-world data, ii) involving synthetic faces into training benefits tracking in some certain scenarios but degrades the tracking model's generalization ability. These two insights could benefit a range of non-deep learning facial image analysis tasks where the labelled real data is difficult to acquire.
Jianwen Lou, Xiaoxu Cai, Junyu Dong, Hui Yu 0001
IEEE Trans. Image Process.4
2021 Distilling Ordinal Relation and Dark Knowledge for Facial Age Estimation
abstract
In this article, we propose a knowledge distillation approach with two teachers for facial age estimation. Due to the nonstationary patterns of the facial-aging process, the relative order of age labels provides more reliable information than exact age values for facial age estimation. Thus, the first teacher is a novel ranking method capturing the ordinal relation among age labels. Especially, it formulates the ordinal relation learning as a task of recovering the original ordered sequences from shuffled ones. The second teacher adopts the same model as the student that treats facial age estimation as a multiclass classification task. The proposed method leverages the intermediate representations learned by the first teacher and the softened outputs of the second teacher as supervisory signals to improve the training procedure and final performance of the compact student for facial age estimation. Hence, the proposed knowledge distillation approach is capable of distilling the ordinal knowledge from the ranking model and the dark knowledge from the multiclass classification model into a compact student, which facilitates the implementation of facial age estimation on platforms with limited memory and computation resources, such as mobile and embedded devices. Extensive experiments involving several famous data sets for age estimation have demonstrated the superior performance of our proposed method over several existing state-of-the-art methods.
Qilu Zhao, Junyu Dong, Hui Yu 0001, Sheng Chen 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 Eye-based Recognition for User Identification on Mobile Devices
abstract
User identification is becoming more and more important for Apps on mobile devices. However, the identity recognition based on eyes, e.g., iris recognition, is rarely used on mobile devices comparing with those based on face and fingerprint due to its extra cost in hardware and complicated operations during recognition. In this article, an eye-based recognition method is designed for identity recognition on mobile devices, which can be implemented just like face recognition. In the proposed method, the eye feature is composed of the static and dynamic features, where the periocular feature extracted by deep neural network from the eye image is used as the static feature, and the motion feature of saccadic velocity is selected as the dynamic feature. The eye images can be captured by the normal camera on mobile devices just like faces, and dynamic features can provide living information to increase the difficulty of forgery. The GazeCapture dataset is used to test the proposed method, because the eye images in this dataset are captured by mobile devices during daily use. The recognition accuracy of the proposed method on the GazeCapture dataset can reach 96.87% only based on the periocular feature and can be enhanced to 97.99% when it is fused with the saccadic feature. The experiment results show that the performance of the proposed method can be comparative to that of iris recognition methods. It demonstrates that the proposed method is a practical reference for the eye-based identity recognition, and the proposed method provides one more biometric choice for mobile devices.
Huiru Shao, Jing Li 0046, Jia Zhang 0028, Hui Yu 0001, Jiande Sun 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Deep Selective Feature Learning for Action Recognition
abstract
Soft-attention mechanism has attracted a lot of attention in recent years due to its ability to capture the most discriminative image features for understanding actions. However, soft-attention tends to focus on fine-grained parts on images and ignores global information, which can lead to totally wrong classification results. To address this issue, we propose a novel deep selective feature learning network (DSFNet), which can automatically learn the feature maps with both fine-grained and global information. Specially, DSFNet is designed to have the ability to learn to adjust the actions for feature map selection by maximizing the cumulative discounted rewards. Moreover, the DSFNet is an easy-to-use extension of state-of-the-art base architectures of multiple tasks. Extensive experiments show that the proposed method has achieved superior performance on two standard action recognition benchmarks across still images (PPMI) and videos (HMDB51).
Yongxin Ge, Jinyuan Feng, Xiaolei Qin, Jiaruo Yu, Hui Yu 0001
ICME6
2020 Scene perception guided crowd anomaly detection
Dingxin Ma, Hui Yu 0001, Peter Howell, Brett Stevens
Neurocomputing3
2020 A joint guidance-enhanced perceptual encoder and atrous separable pyramid-convolutions for image inpainting
Yongle Zhang 0001, Yingyu Wang, Junyu Dong, Lin Qi 0004, Hao Fan 0004, Xinghui Dong, Muwei Jian, Hui Yu 0001
Neurocomputing8
2020 Augmented visual feature modeling for matching in low-visibility based on cycle-labeling of Superpixel Flow
Shu Zhang 0002, Hui Yu 0001, Ting Wang 0018, Junyu Dong
Knowl. Based Syst.2
2020 Weight analysis for various prohibitory sign detection and recognition using deep learning
Christine Dewi, Rung Ching Chen, Hui Yu 0001
Multim. Tools Appl.3
2020 A brief survey of visual saliency detection
Inam Ullah 0002, Muwei Jian, Sumaira Hussain, Jie Guo 0012, Hui Yu 0001, Xing Wang 0002, Yilong Yin
Multim. Tools Appl.5
2020 Hybrid regression and isophote curvature for accurate eye center localization
abstract
Abstract The eye center localization is a crucial requirement for various human-computer interaction applications such as eye gaze estimation and eye tracking. However, although significant progress has been made in the field of eye center localization in recent years, it is still very challenging for tasks under the significant variability situations caused by different illumination, shape, color and viewing angles. In this paper, we propose a hybrid regression and isophote curvature for accurate eye center localization under low resolution. The proposed method first applies the regression method, which is called Supervised Descent Method (SDM), to obtain the rough location of eye region and eye centers. SDM is robust against the appearance variations in the eye region. To make the center points more accurate, isophote curvature method is employed on the obtained eye region to obtain several candidate points of eye center. Finally, the proposed method selects several estimated eye center locations from the isophote curvature method and SDM as our candidates and a SDM-based means of gradient method further refine the candidate points. Therefore, we combine regression and isophote curvature method to achieve robustness and accuracy. In the experiment, we have extensively evaluated the proposed method on the two public databases which are very challenging and realistic for eye center localization and compared our method with existing state-of-the-art methods. The results of the experiment confirm that the proposed method outperforms the state-of-the-art methods with a significant improvement in accuracy and robustness and has less computational complexity.
Jianwen Lou, Junyu Dong, Lin Qi 0004, Gongfa Li, Hui Yu 0001
Multim. Tools Appl.6
2020 Spatial Enhancement and Temporal Constraint for Weakly Supervised Action Localization
abstract
Weakly supervised temporal action localization (WSTAL) is a practical but challenging issue in video understanding. However, most existing methods have to activate background snippets or deactivate action snippets in cases of no boundary annotations, which inevitably affects the localization of action instances. In this letter, we propose a spatial enhancement and temporal constraint (SETC) model to address this problem from three aspects. Specifically, we first propose a spatial enhancement module to enhance the discrimination of the extracted features. Then we leverage the instance sparse constraint to restrain the drastic fluctuation class activation sequence (CAS). Finally, we use the confidence connectivity enhancement to connect the snippets that are broken up by mistake. Experiments on THUMOS'14 and ActivityNet datasets validate the efficacy of SETC against existing state-of-the-art WSTAL algorithms.
Xiaolei Qin, Yongxin Ge, Hui Yu 0001, Feiyu Chen 0002, Dan Yang 0001
IEEE Signal Process. Lett.3
2020 CMIB: Unsupervised Image Object Categorization in Multiple Visual Contexts
abstract
Object categorization in images is fundamental to various industrial areas, such as automated visual inspection, fast image retrieval, and intelligent surveillance. Most existing methods treat visual features (e.g., scale-invariant feature transform) as content information of the objects, while regarding image tags as their contextual information. However, the image tags can hardly be acquired in completely unsupervised settings, especially when the image volume is too large to be marked. In this article, we propose a novel contextual multivariate information bottleneck (CMIB) method to conduct unsupervised image object categorization in multiple visual contexts. Unlike using manual contexts, the CMIB method first automatically generates a set of high-level basic clusterings by multiple global features, which are unprecedentedly defined as visual contexts since they can provide overall information about the target images. Then, the idea of the data compression procedure for object category discovery is proposed, in which the content and multiple visual contexts are maximally preserved through a “bottleneck.” Specifically, two Bayesian networks are initially built to characterize the relationship between data compression and information preservation. Finally, a novel sequential information-theoretic optimization is proposed to ensure the convergence of the CMIB objective function. Experimental results on seven real-world benchmark image datasets demonstrate that the CMIB method achieves better performance than the state-of-the-art baselines.
Yangdong Ye, Xueying Qiu, Milos Manic, Hui Yu 0001
IEEE Trans. Ind. Informatics5
2020 Learning the Traditional Art of Chinese Calligraphy via Three-Dimensional Reconstruction and Assessment
abstract
The traditional art of Chinese calligraphy, reflecting the wisdom of the grass-roots community, is the soul of Chinese culture. Just like many other types of craftsmanship, it is part of the historical heritage and is worth conserving, from generation to generation. Since the movements of an ink brush are in a 3D style when Chinese calligraphy is written, they embody “The Power of Beauty,” comprising various reflectance properties and rough-surface geometry. To truly understand the powerful significance and beauty of the art of Chinese calligraphy, in this paper, a 3D calligraphy reconstruction method, based on Photometric Stereo, is designed to capture the detailed appearance of the calligraphy's 3D surface geometry. For assessment, an Iterative Closest Point (ICP) algorithm is applied for registration of 3D intrinsic shapes between the Chinese calligraphy and the calligraphy fans' handwriting. Through matching these two sets of calligraphy characters, the designed system can give a score to the handwriting of a user. Experiments have been performed on Chinese calligraphy from different historical dynasties to evaluate the effectiveness of the proposed scheme, and experimental results show that the developed system is useful and provides a convenient method of calligraphy appreciation and assessment.
Muwei Jian, Junyu Dong, Maoguo Gong, Hui Yu 0001, Liqiang Nie, Yilong Yin, Kin-Man Lam 0001
IEEE Trans. Multim.4
2020 Realistic Facial Expression Reconstruction for VR HMD Users
abstract
We present a system for sensing and reconstructing facial expressions of the virtual reality (VR) head-mounted display (HMD) user. The HMD occludes a large portion of the user's face, which makes most existing facial performance capturing techniques intractable. To tackle this problem, a novel hardware solution with electromyography (EMG) sensors being attached to the headset frame is applied to track facial muscle movements. For realistic facial expression recovery, we first reconstruct the user's 3D face from a single image and generate the personalized blendshapes associated with seven facial action units (AUs) on the most emotionally salient facial parts (ESFPs). We then utilize pre-processed EMG signals for measuring activations of AU-coded facial expressions to drive pre-built personalized blendshapes. Since facial expressions appear as important nonverbal cues of the subject's internal emotional states, we further investigate the relationship between six basic emotions - anger, disgust, fear, happiness, sadness and surprise, and detected AUs using a fern classifier. Experiments show the proposed system can accurately sense and reconstruct high-fidelity common facial expressions while providing useful information regarding the emotional state of the HMD user.
Jianwen Lou, Yiming Wang 0001, Charles Nduka, Mahyar Hamedi, Ifigeneia Mavridou, Fei-Yue Wang 0001, Hui Yu 0001
IEEE Trans. Multim.7
2019 Generalizing to Unseen Head Poses in Facial Expression Recognition and Action Unit Intensity Estimation
abstract
Facial expression analysis is challenged by the numerous degrees of freedom regarding head pose, identity, illumination, occlusions, and the expressions itself. It currently seems hardly possible to densely cover this enormous space with data for training a universal well-performing expression recognition system. In this paper we address the sub-challenge of generalizing to head poses that were not seen in the training data, aiming at getting along with sparse coverage of the pose subspace. For this purpose we (1) propose a novel face normalization method called FaNC that massively reduces pose-induced image variance; (2) we compare the impact of the proposed and other normalization methods on (a) action unit intensity estimation with the FERA 2017 challenge data (achieving new state of the art) and (b) facial expression recognition with the Multi-PIE dataset; and (3) we discuss the head pose distribution needed to train a pose-invariant CNN-based recognition system. The proposed FaNC method normalizes pose and facial proportions while retaining expression information and runs in less than 2 ms. When comparing results achieved by training a CNN on the output images of FaNC and other normalization methods, FaNC generalizes significantly better than others to unseen poses if they deviate more than 20° from the poses available during training. Code and data are available.
Philipp Werner, Frerk Saxen, Ayoub Al-Hamadi, Hui Yu 0001
FG4
2019 Visual Saliency Detection Based on Full Convolution Neural Networks and Center Prior
abstract
Video saliency detection aims to mimic the human's visual attention system of perceiving the world via extracting the most attractive regions or objects in the input video. At present, traditional video saliency-detection models have achieved good performance in many applications. However, it is still challenging in exploiting the consistency of spatiotemporal information. In order to tackle this challenge, this paper proposes a video saliency-detection model based on human attention mechanism and full convolution neural networks. First, visual features are extracted from video frames through the fully convolutional networks. The second stage is to spread attention features to the other layer (i. e. the fifth layer) of fully convolutional networks via a weight sharing strategy. Finally, the final result produced by the convolution network is optimized by considering spatial location information with center prior of the salient object. Experimental results show that the performance of the proposed algorithm is superior to other state-of-the-art methods based on the widely used data set for video saliency detection.
Muwei Jian, Jiaojin Wang, Hui Yu 0001
HSI4
2019 Multi-subspace supervised descent method for robust face alignment
abstract
Supervised Descent Method (SDM) is one of the leading cascaded regression approaches for face alignment with state-of-the-art performance and a solid theoretical basis. However, SDM is prone to local optima and likely averages conflicting descent directions. This makes SDM ineffective in covering a complex facial shape space due to large head poses and rich non-rigid face deformations. In this paper, a novel two-step framework called multi-subspace SDM (MS-SDM) is proposed to equip SDM with a stronger capability for dealing with unconstrained faces. The optimization space is first partitioned with regard to shape variations using k-means. The generated subspaces show semantic significance which highly correlates with head poses. Faces among a certain subspace also show compatible shape-appearance relationships. Then, Naive Bayes is applied to conduct robust subspace prediction by concerning about the relative proximity of each subspace to the sample. This guarantees that each sample can be allocated to the most appropriate subspace-specific regressor. The proposed method is validated on benchmark face datasets with a mobile facial tracking implementation.
Jianwen Lou, Xiaoxu Cai, Yiming Wang 0001, Hui Yu 0001, Shaun J. Canavan
Multim. Tools Appl.4
2018 Predicting and Generating Wallpaper Texture with Semantic Properties
abstract
Humans naturally use semantic descriptions to express their visual perception of textures; this is also the fact for perception and description of wallpaper texture. Classification of wallpaper's style is mainly based on understanding of visual information. However, the complexity of real-world wallpaper images is difficult to be captured by existing datasets. Inspired by a publicly available Procedural Textures Dataset, a number of wallpaper images was collected and assembled into a wallpaper dataset. A series of psychophysical experiments was performed to further collect semantic descriptions for this dataset. Each wallpaper was labeled with 5-10 semantic descriptions. More importantly, our dataset contains complex wallpaper images with rich annotations. To our best knowledge, our dataset is the first public wallpaper dataset with semantic descriptions. We use label distribution to analysis semantic descriptions and texture characteristics. Furthermore, a texture generation method based on GAN was tested using our wallpaper dataset, which produced state-of-the-art results.
Xiaohan Feng, Lin Qi 0004, Yanhai Gan, Ying Gao 0005, Hui Yu 0001, Junyu Dong
HSI5
2018 Kinect Depth Recovery via the Cooperative Profit Random Forest Algorithm
abstract
The depth map captured by Kinect usually contain missing depth data. In this paper, we propose a novel method to recover the missing depth data with the guidance of depth information of each neighborhood pixel. In the proposed framework, a self-taught mechanism and a cooperative profit random forest (CPRF) algorithm are combined to predict the missing depth data based on the existing depth data and the corresponding RGB image. The proposed method can overcome the defects of the traditional methods which is prone to producing artifact or blur on the edge of objects. The experimental results on the Berkeley 3-D Object Dataset (B3DO) and the Middlebury benchmark dataset show that the proposed method outperforms the existing method for the recovery of the missing depth data. In particular, it has a good effect on maintaining the geometry of objects.
Jianyuan Sun, Junyu Dong, Hui Yu 0001
HSI5
2018 Dynamic 3D Surface Reconstruction Using a Hand-Held Camera
abstract
This paper proposes a dynamic 3D reconstruction method for recovering a surface shape from a set of images that are captured by a hand-held camera. A light source is attached to the camera as a photometric constraint. Thus, we can effectively calculate photometric stereo using the relative moving camera. The key contributions of our work are a robust pixel matching method to build effective correspondences between images for normal estimation, and an optimization method to correct the deviation in the recovered surface shape that is caused by the nonideal illumination in a close-range lighting condition. Specially we correct the recovered shape by adding an interpolation surface that is estimated using sparse control points from the structure from motion. The effectiveness of our method is verified on real datasets with a digital camera and a smart phone.
Hao Fan 0004, Lin Qi 0004, Junyu Dong, Gongfa Li, Hui Yu 0001
IECON5
2018 Perception-driven procedural texture generation from examples
Jun Liu 0055, Yanhai Gan, Junyu Dong, Lin Qi 0004, Xin Sun 0003, Muwei Jian, Hui Yu 0001
Neurocomputing8
2018 A hybrid spatio-temporal model for detection and severity rating of Parkinson's disease from gait data
Aite Zhao, Lin Qi 0004, Jie Li 0001, Junyu Dong, Hui Yu 0001
Neurocomputing5
2018 Saliency detection based on directional patches extraction and principal local color contrast
Muwei Jian, Wenyin Zhang, Hui Yu 0001, Chaoran Cui, Xiushan Nie, Huaxiang Zhang 0001, Yilong Yin
J. Vis. Commun. Image Represent.3
2018 Dual channel LSTM based multi-feature extraction in gait for diagnosis of Neurodegenerative diseases
Aite Zhao, Lin Qi 0004, Junyu Dong, Hui Yu 0001
Knowl. Based Syst.4
2017 Cascade support vector regression-based facial expression-aware face frontalization
abstract
The main aim of face frontalization is to synthesize the frontal facial appearances from non-frontal facial images. How to estimate the frontal face-shape is a crucial but very challenging problem in the frontalization task. Most existing methods use a single shape template to fit in with frontal facial appearances, which will result in a loss of expression-related information. In this work, we present a novel facial expression-aware face frontalization method which directly learns the pair-wise relations between non-frontal face-shape and its frontal counterpart. The support vector regression is explored to train the pair-wise regression model. Considered the pair-wise relationship is non-linear, an appropriate cascade manner is applied to iteratively adjust and optimize the model. With the estimated frontal shape, facial appearances are synthesized through a texture-fitting process formulated by solving a simple optimization problem. The proposed method has been evaluated on a in-the-wild facial expression database. The experimental results shows an outstanding performance of both visual effects of expression recovery and facial expression recognition.
Yiming Wang 0001, Hui Yu 0001, Junyu Dong, Muwei Jian, Honghai Liu 0001
ICIP2
2017 Extended social force model-based mean shift for pedestrian tracking under obstacle avoidance
abstract
It has been shown that mean shift tracking algorithm can achieve excellent results in pedestrian tracking task. It empirically estimates the target position of current frame by locating the maximum of a density function from the local neighborhood of the target position of previous frame. However, this method only considers its past trajectory without considering the influence of pedestrian environment when applying to pedestrian tracking. In practical, pedestrians always keep a safe distance away from obstacles when programming their paths. To address the issue of obstacle avoidance, this paper proposes a novel extended social force model‐based mean shift tracking algorithm in which pedestrian environment is full taken in consideration. Firstly, an extended social force model is presented to quantify the interaction between pedestrian and obstacle by means of force. Furthermore, directional weights and speed weights are introduced to adjust the strength of the force in terms of the difference of individual perspectives and relative velocities. Finally, the initial target position is predicted by Newton's laws of motion and then the Mean Shift method is integrated to track target position. Experiment results show that this algorithm achieves an encouraging performance when obstacles exist.
Yiming Wang 0001, Hui Yu 0001
IET Comput. Vis.4
2017 Underwater image enhancement via extended multi-scale Retinex
abstract
Underwater exploration has become an active research area over the past few decades. The image enhancement is one of the challenges for those computer vision based underwater researches because of the degradation of the images in the underwater environment. The scattering and absorption are the main causes in the underwater environment to make the images decrease their visibility, for example, blurry, low contrast, and reducing visual ranges. To tackle aforementioned problems, this paper presents a novel method for underwater image enhancement inspired by the Retinex framework, which simulates the human visual system. The term Retinex is created by the combinations of “Retina” and “Cortex”. The proposed method, namely LAB-MSR, is achieved by modifying the original Retinex algorithm. It utilizes the combination of the bilateral filter and trilateral filter on the three channels of the image in CIELAB color space according to the characteristics of each channel. With real world data, experiments are carried out to demonstrate both the degradation characteristics of the underwater images in different turbidities, and the competitive performance of the proposed method.
Shu Zhang 0002, Ting Wang 0018, Junyu Dong, Hui Yu 0001
Neurocomputing4
2016 Facial Expression-Aware Face Frontalization
Yiming Wang 0001, Hui Yu 0001, Junyu Dong, Brett Stevens, Honghai Liu 0001
ACCV (3)2
2016 Robust Photometric Stereo in a scattering medium via Low-Rank Matrix Completion and Recovery
abstract
Photometric Stereo is a popular method for 3D reconstruction from images due to its high level of details handling. However, when it is used in a scattering medium such as lakes and oceans, the recovery result will be negatively impacted by the light absorption, light scattering and the impurities in the water. In this paper, we present a new method to solve the problem of better 3D reconstruction via Low-Rank Matrix Completion and Recovery. First, we use the dark points, like shadows and darkness in the water to fit the scattering effect distribution and then remove the scattering from the image. Next, we use the Robust Principal Component Analysis method (RPCA) to recover the image by removing the sparse noise including shadows, impurities and some corrupted points caused by backscatter compensation. Finally, we combine the RPCA results and the least-squares (LS) results to get the surface normal and accomplish the 3D reconstruction. Extensive experimental results demonstrate that our method achieves more accurate estimates of surface normal and 3D reconstruction than previous techniques.
Hao Fan 0004, Yisong Luo, Lin Qi 0004, Junyu Dong, Hui Yu 0001
HSI6
2016 Real-time 3D point cloud segmentation using Growing Neural Gas with Utility
abstract
This paper proposes a real-time feature extraction and segmentation method for a 3D point cloud. First of all, we apply Growing Neural Gas with Utility (GNG-U) to the point cloud for learning a topological structure. However, the standard GNG-U cannot learn the topological structure of 3D space environment and color information simultaneously. To this end, we then modify the GNG-U algorithm by using a weight vector. we propose a surface feature extraction and segmentation method by efficiently utilizing the topological structure. Our segmentation method is based on a region growing method whose similarity value uses the inner value of two normal vectors connected by the topological structure. We show experimental results of the proposed method and discuss the effectiveness of the proposed method.
Yuichiro Toda, Zhaojie Ju, Hui Yu 0001, Naoyuki Takesue, Kazuyoshi Wada, Naoyuki Kubota
HSI3
2016 Learning perceptual texture similarity and relative attributes from computational features
abstract
Previous work has shown that perceptual texture similarity and relative attributes cannot be well described by computational features. In this paper, we propose to predict human's visual perception of texture images by learning a non-linear mapping from computational feature space to perceptual space. Hand-crafted features and deep features, which were successfully applied in texture classification tasks, were extracted and used to train Random Forest and rankSVM models against perceptual data from psychophysical experiments. Three texture datasets were used to test our proposed method and the experiments show that the predictions of such learnt models are in high correlation with human's results.
Jianwen Lou, Lin Qi 0004, Junyu Dong, Hui Yu 0001, Guoqiang Zhong 0001
IJCNN4
2016 Combining 3D joints Moving Trend and Geometry property for human action recognition
abstract
Depth image based human action recognition has attracted many attentions due to the popularity of the depth sensors. However, accurate recognition still remains a challenge because of various object appearances, poses and video sequences. In this paper, a novel skeleton joints descriptor based on 3D Moving Trend and Geometry (3DMTG) property is proposed for human action recognition. Specifically, a histogram of 3D moving directions between consecutive frames for each joint is constructed to represent the 3D moving trend feature in spatial domain. The geometry information of joints in each frame is modelled by the relative motion with the initial status. The proposed feature descriptor is evaluated on two popular datasets. The experimental results demonstrate the superior performance of our method over the state-of-the-art methods, especially the higher recognition rates for complex actions.
Bangli Liu, Hui Yu 0001, Xiaolong Zhou 0001, Honghai Liu 0001
SMC2
2016 A fusion method for robust face tracking
Xiaodong Jiang, Hui Yu 0001, Yang Lu 0003, Honghai Liu 0001
Multim. Tools Appl.2
2016 Automatic evaluation of the degree of facial nerve paralysis
Ting Wang 0018, Shu Zhang 0002, Junyu Dong, Li'an Liu, Hui Yu 0001
Multim. Tools Appl.5
2016 Guest Editorial: Advanced Understanding and Modelling of Human Motion in Multidimensional Spaces
Hui Yu 0001, Junyu Dong, Tuan D. Pham, Honghai Liu 0001
Multim. Tools Appl.1
2015 Toward a psychophysical-based procedural texture generation system for interactive design
abstract
Procedural textures have been widely used as they can be easily generated from various mathematical models. However, the model parameters are not perceptually meaningful or uniform for non-expert users. In this paper, we proposed a system that can generate procedural textures interactively along certain perceptual dimensions. We built a procedural texture dataset and measured twelve perceptual properties of a small subset through psychophysical experiments. The perceived magnitude of the rest textures was estimated by Support Vector Machines using computational features from a cascaded PCA network. For a given texture displayed on a touch screen, the user makes finger gestures which were then transferred to magnitude changes in perceptual space. The texture in the database that matches the new perceptual scale and with nearest distance in computational feature space will be chosen and displayed. We reported our experiment results for two particular perceptual properties: surface roughness and directionality. Other properties can be manipulated similarly.
Xiaoxu Cai, Jun Liu 0055, Lin Qi 0004, Junyu Dong, Ying Gao 0005, Hui Yu 0001
HSI6
2015 Dynamic facial expression recognition using local patch and LBP-TOP
abstract
Local binary pattern on three orthogonal planes (LBP-TOP) is one of the most popular method for dynamic texture analysis and has been successfully applied to facial expression analysis. Yet an effective LBP-TOP operator highly relies on preprocessing. And, like many appearance-based approaches, this approach reserves more identity-related cues rather than expression. In this work, we propose a fully automatic approach for facial expression recognition based on points registration, localized patch extraction and LBP-TOP feature representation. The efficiency of this method is evaluated on CK+ database. Results show that the proposed method has achieved a better performance compared with existing methods.
Yiming Wang 0001, Hui Yu 0001, Brett Stevens, Honghai Liu 0001
HSI2
2015 Explore the design style of oriented facility based on user evaluation
abstract
This paper employs Kansei engineering to analyze the relationship between user preference and the given architectural design scheme. In this study, we first divide architectural styles into seven different categories. Then we classify the key factors in the oriented facility design into 7 types with 39 subcategories. On that basis, we explore which design factor plays main roles in the harmony and unity between the user-oriented type and the given architectural design among seven different architectural styles (the national refined style). And then the SD method is used to compare the architectural style sample to the oriented facilities samples in order to obtain the data of matching degree between them. Through analyzing the collected data, we put forward the design model based on design factors of guild facility model, which is adapted to the particular architecture styles.
Hui Yu 0001
HSI3
2015 Combining Kinect and PnP for camera pose estimation
abstract
This paper presents a novel method to conduct camera pose estimation though combining Kinect and Perspective-n-points algorithms. Most existing camera pose estimation methods suffer from the errors caused by inevitable outliers between 2D-3D correspondences. To this end, we propose to use a random down sampling process to deal with outliers in this paper. The proposed method is divided into two main steps, which are 2D-3D correspondences generation and pose estimation. The method has been tested in a real project, and the experiment has shown encouraging results compared to the ground truth.
Shu Zhang 0002, Hui Yu 0001, Junyu Dong, Ting Wang 0018, Lin Qi 0004, Honghai Liu 0001
HSI2
2015 Automatic Reconstruction of Dense 3D Face Point Cloud with a Single Depth Image
abstract
Human face analysis is the basis for many other computer vision tasks, such as camera surveillance, entrance authorization and age estimation. With 3D face models, the vision task based on facial analysis can usually achieve a higher accuracy than the 2D cases since it provides more information with the additional dimension. However, most existing 3D face reconstruction methods suffer from complicated processing and high computation. This paper presents a novel method that simplifies the 3D face reconstruction process with only one shot of Kinect data. The output of the system is a high density of 3D face point cloud with smoother surface. This provides rich details of the human face for other computer vision tasks. Experiments with real world data show promising results using the proposed method.
Shu Zhang 0002, Hui Yu 0001, Junyu Dong, Ting Wang 0018, Zhaojie Ju, Honghai Liu 0001
SMC2
2015 A framework for automatic and perceptually valid facial expression generation
Hui Yu 0001, Oliver G. B. Garrod, Rachael Jack, Philippe G. Schyns
Multim. Tools Appl.1
2014 Image factorization and feature fusion for enhancing robot vision in human face recognition
abstract
Illumination variation has been a challenging problem for face recognition in robot vision. To reduce the effect caused by illumination variation, a lot of studies have been explored. The Total Variation (TV) method is particular used to factorize images into a low frequency component and a high frequency one. However, the low frequency component still contains significant intrinsic features resulting in failure in face recognition in some cases. In this paper, we propose to further extract illumination invariant features from face images under uncontrolled varying lighting conditions. The Nonsampled Contourlet Transform (NSCT) method is employed to enhance the extraction of intrinsic feature. The combined factorization model is very effective in the experiment on the Yale database.
Hui Yu 0001, Zhaojie Ju, Honghai Liu 0001
IJCNN1
2014 Linear regression for head pose analysis
abstract
Extensive research has been conducted to estimate and analyze head poses for various applications. Most existing methods tend to detect facial features and locate landmarks on a face for pose estimation. However, the sensitivity to occlusion of some face parts with key features and uncontrolled illumination of face images make the facial feature detection vulnerable. In this paper, we propose a framework for pose estimation without the need of face features or landmarks detection. Specifically, we formulate the pose estimation as a linear regression applied to the pose space. This method is based on the assumption that pose space cannot be linearly approximated in the pose subspace. The experimental results strongly support this assumption. In cases where the database does not obtain various poses in the intraclass, we propose to generate those poses through a 3D reconstruction and projection method. The experiment conducted on the CMU MultiPIE and IMM Face database has shown the effectiveness of the proposed method.
Hui Yu 0001, Honghai Liu 0001
IJCNN1
2014 Regression-Based Facial Expression Optimization
abstract
This paper presents an approach for reproducing optimal 3-D facial expressions based on blendshape regression. It aims to improve fidelity of facial expressions but maintain the efficiency of the blendshape method, which is necessary for applications such as human-machine interaction and avatars. The method intends to optimize the given facial expression using action units (AUs) based on the facial action coding system recorded from human faces. To help capture facial movements for the target face, an intermediate model space is generated, where both the target and source AUs have the same mesh topology and vertex number. The optimization is conducted interactively in the intermediate model space through adjusting the regulating parameter. The optimized facial expression model is transferred back to the target facial model to produce the final facial expression. We demonstrate that given a sketched facial expression with rough vertex positions indicating the intended facial expression, the proposed method approaches the sketched facial expression through automatically selecting blendshapes with corresponding weights. The sketched expression model is finally approximated through AUs representing true muscle movements, which improves the fidelity of facial expressions.
Hui Yu 0001, Honghai Liu 0001
IEEE Trans. Hum. Mach. Syst.1
2013 Explore New Eye Tracking and Gaze Locating Methods
abstract
Eye tracking has been used extensively in research, often for plotting the gaze location of research participants. Many of the existing methods on tracking gaze locations, however, require special equipment. This equipment can be very expensive and/or cumbersome to use, restricting the use of it to specialist labs. By producing a method of eye tracking which requires minimal equipment and set up time, eye tracking might be more widely used as a research tool, especially in exploratory areas where the outcomes of the research are not known. This paper introduces two novel eye tracking methods which use only a web cam feed as input and require no specialist equipment. The camera used in this research is a standard web cam built into a normal laptop, any current commercial web cam could be used. Experiments using the proposed method showed more accurate tracking results than some existing eye trackers.
Alexander Kadyrov, Hui Yu 0001, Joe Eyles, Honghai Liu 0001
SMC2
2013 Ship Detection and Segmentation Using Image Correlation
abstract
There have been intensive research interests in ship detection and segmentation due to high demands on a wide range of civil applications in the last two decades. However, existing approaches, which are mainly based on statistical properties of images, fail to detect smaller ships and boats. Specifically, known techniques are not robust enough in view of inevitable small geometric and photometric changes in images consisting of ships. In this paper a novel approach for ship detection is proposed based on correlation of maritime images. The idea comes from the observation that a fine pattern of the sea surface changes considerably from time to time whereas the ship appearance basically keeps unchanged. We want to examine whether the images have a common unaltered part, a ship in this case. To this end, we developed a method - Focused Correlation (FC) to achieve robustness to geometric distortions of the image content. Various experiments have been conducted to evaluate the effectiveness of the propose.
Alexander Kadyrov, Hui Yu 0001, Honghai Liu 0001
SMC2
2013 New Perception of Fluffy Surfaces
abstract
Scene material recognition is a valuable perceptual ability for robots. However, current methods fail to provide robust solution for indoor surrounding material recognition. As human beings are able to immediately understand properties of surrounding materials, especially when it concerns smooth fluffy materials, we believe robots should be equipped with similar abilities. In this paper we explore a new idea for fluffy surface perception for robots based on above hypothesis. The aim is to enable robots to distinguish smooth and fluffy surfaces without touching them. This is achieved through calculating image match ability map from video cameras. Through measuring the similarity of images captured from different viewpoints, robots are able to immediately recognize whether it is a fluffy material. The method has been validated by primary experiments. Our results show that robots can have a sense of material properties for fluffy materials without touching them or any prior knowledge.
Alexander Kadyrov, Hui Yu 0001, Honghai Liu 0001
SMC2
2012 Perception-driven facial expression synthesis
Hui Yu 0001, Oliver G. B. Garrod, Philippe G. Schyns
Comput. Graph.1
2012 On generating realistic avatars: dress in your own style
Hui Yu 0001, Sheng Feng Qin, Guangmin Sun, David K. Wright 0001
Multim. Tools Appl.1