Guanghua Tan

dblp:148/4681 · DBLP profile ↗
← Back
37ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0001-6001-2351ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 11 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Structure-aware fine-grained instance segmentation for fetal brain ultrasound images
Laifa Ma, Kenli Li 0001, Guanghua Tan, Huaxuan Wen, Shengli Li 0001
Neurocomputing4
2025 Concept-Induced Graph Perception Model for Interpretable Diagnosis
Lei Zhao 0013, Changjian Chen, Bin Pu, Xiaoming Qi, Fengfeng Peng, Chunlian Wang, Kenli Li 0001, Guanghua Tan
MICCAI (12)8
2025 Anatomical structures detection using topological constraint knowledge in fetal ultrasound
Juncheng Guo, Guanghua Tan, Bin Pu, Chunlian Wang, Shengli Li 0001, Kenli Li 0001
Neurocomputing2
2025 A key instance-guided frame-to-video information fusion network for thyroid ultrasound video instance segmentation
Guanyuan Chen, Ningbo Zhu, Bin Pu, Guanghua Tan, Hongxia Luo, Kenli Li 0001
Knowl. Based Syst.4
2025 DH-GAC: deep hierarchical context fusion network with modified geodesic active contour for multiple neurofibromatosis segmentation
Xiangqiong Wu, Guanghua Tan, Bin Pu, Mingxing Duan, Wenli Cai
Neural Comput. Appl.2
2025 Optical Flow-Enhanced Mamba U-Net for Cardiac Phase Detection in Ultrasound Videos
abstract
The detection of cardiac phase in ultrasound videos, identifying end-systolic (ES) and end-diastolic (ED) frames, is a critical step in assessing cardiac function, monitoring structural changes, and diagnosing congenital heart disease. Current popular methods use recurrent neural networks to track dependencies over long sequences for cardiac phase detection, but often overlook the short-term motion of cardiac valves that sonographers rely on. In this paper, we propose a novel optical flow-enhanced Mamba U-net framework, designed to utilize both short-term motion and long-term dependencies to detect the cardiac phase in ultrasound videos. We utilize optical flow to capture the short-term motion of cardiac muscles and valves between adjacent frames, enhancing the input video. The Mamba layer is employed to track long-term dependencies across cardiac cycles. We then develop regression branches using the U-Net architecture, which integrates short-term and long-term information while extracting multi-scale features. Using this method, we can generate regression scores for each frame and identify keyframes (i.e., ES and ED frames). Additionally, we design a keyframe weighted loss function to guide the network to focus more on keyframes rather than intermediate period frames. Our method demonstrates superior performance compared to advanced baseline methods, achieving frame mismatches of 1.465 frames for ES and 0.842 frames for ED in the Fetal Echocardiogram dataset, where heart rates are higher and phase changes occur rapidly, and 2.444 frames and 2.072 frames in the publicly available adult Echonet-Dynamic dataset. Its accuracy and robustness in both fetal and adult datasets highlight its potential for clinical application.
Yuhuan Lu 0002, Guanghua Tan, Bin Pu, Pak-Hei Yeung, Shengli Li 0001, Jagath C. Rajapakse, Kenli Li 0001
IEEE Trans. Medical Imaging2
2024 Advancing Ultrasound Medical Continuous Learning with Task-Specific Generalization and Adaptability
abstract
As artificial intelligence progresses in the field of medical ultrasound image analysis, mitigating catastrophic forgetting during continuous learning processes in disease diagnosis and fetal ultrasound assistance is crucial. Inspired by the significant performance improvements achieved in various downstream visual tasks through advanced visual representations, we introduce a novel approach called Masked Ultrasound Image Modeling (MUIM) to prevent forgetting of new disease diagnosis tasks. In this approach, MUIM initially pre-trains on ultrasound images specific to the current task. By utilizing a masked autoencoder to train on unlabeled ultrasound images, it learns highly abstract representations of task-relevant information. To ensure the model’s adaptability to new tasks, we propose contrastive Historical-Current Learning strategy, which enhances the model’s ability to retain and integrate knowledge from previous tasks while learning new ones. By incorporating knowledge distillation loss and the exponential moving average (EMA) technique for joint inference and continuously updating new knowledge into the historical expert model, our approach enables the model to adaptively learn new tasks while preventing forgetting of old ones. We conducted extensive experiments on three datasets: the Breast Ultrasound Image dataset (BUSI), the Algeria Thyroid Ultrasound Image dataset (AUITD), and our designed Obstetrics and Gynecology Ultrasound Dataset (USOGD). The results demonstrate that our method significantly reduces the forgetting of old task knowledge while outperforming state-of-the-art methods in classification accuracy.
Chunzheng Zhu, Guanghua Tan, Ningbo Zhu, Kenli Li 0001, Chunlian Wang, Shengli Li 0001
BIBM3
2024 MC-SORT: A Motion Correction-Based Framework for Long-Term Multiple Object Tracking
abstract
Long-term occlusion is one of the most formidable challenges in Multi-Object Tracking (MOT). The motion models of existing SORT-based trackers are unreliable in estimating the motion states of long-term occluded targets. This is mainly because as the occlusion period increases, the increases speed of estimation errors in the motion model increases faster. In practical applications, we believe that the estimation error of the tracker during long-term occlusion is mainly concentrated in the estimation error of the motion model on the velocity of the occluded target. In this work, we have demonstrated that in the long-term occlusion period, appropriately correcting the estimated values of the motion model on the target motion velocity and fully utilizing the temporal and attribute information of the target’s historical trajectory as calculation indicators of correlation are beneficial for improving the robustness of the tracker in long-term occlusion. We refer to our proposed motion correction-based framework as MC-SORT, which mainly consists of a Momentum Compensation Module (MCM) and a Backtracking Re-association (BRA) module. The former can correct the estimated value of the target’s motion state during long-term occlusion, the latter uses the temporal and attribute information of the target’s historical trajectory during long-term occlusion as correlation indicators to measure the degree of correlation between the target and trajectory. Our proposed MC-SORT has the characteristics of simplicity, online, real-time, and plug-and-play, particularly improving the robustness of the tracker in long-term occlusion. The extensive experimental results on the MOT17 and MOT20 datasets demonstrate the robustness and superiority of our framework.
Yunchuan Qin, Ruihui Li, Guanghua Tan, Zhuo Tang, Kenli Li 0001
ECAI4
2024 Unsupervised Ultrasound Image Quality Assessment with Score Consistency and Relativity Co-learning
Juncheng Guo, Guanghua Tan, Yuhuan Lu 0002, Shengli Li 0001, Kenli Li 0001
MICCAI (5)3
2024 Graph-enhanced ensembles of multi-scale structure perception deep architecture for fetal ultrasound plane recognition
Guanghua Tan, Chunlian Wang, Bin Pu, Shengli Li 0001, Kenli Li 0001
Eng. Appl. Artif. Intell.2
2024 A knowledge-interpretable multi-task learning framework for automated thyroid nodule diagnosis in ultrasound videos
Xiangqiong Wu, Guanghua Tan, Hongxia Luo, Zhilun Chen, Bin Pu, Shengli Li 0001, Kenli Li 0001
Medical Image Anal.2
2024 COALA: A Compiler-Assisted Adaptive Library Routines Allocation Framework for Heterogeneous Systems
abstract
Experienced developers often leverage well-tuned libraries and allocate their routines for computing tasks to enhance performance when building modern scientific and engineering applications. However, such well-tuned libraries are meticulously customized for specific target architectures or environments. Additionally, the performance of their routines is significantly impacted by the actual input data of computing tasks, which often remains uncertain until runtime. Accordingly, statically allocating these library routines may hinder the adaptability of applications and compromise performance, particularly in the context of heterogeneous systems. To address this issue, we propose the Compiler-Assisted Adaptive Library Routines Allocation (COALA) framework for heterogeneous systems. COALA is a fully automated mechanism that employs compiler assistance for dynamic allocation of the most suitable routine to each computing task on heterogeneous systems. It allows the deployment of varying allocation policies tailored to specific optimization targets. During the application compilation process, COALA reconstructs computing tasks and inserts a probe for each of these tasks. Probes serve the purpose of conveying vital information about the requirements of each task, including its computing objective, data size, and computing flops, to a user-level allocation component at runtime. Subsequently, the allocation component utilizes the probe information along with the allocation policy to assign the most optimal library routine for executing the computing tasks. In our prototype, we further introduce and deploy a performance-oriented allocation policy founded on a machine learning-based performance evaluation method for library routines. Experimental verification and evaluation on two heterogeneous systems reveal that COALA can significantly improve application performance, with gains of up to 4.3x for numerical simulation software and 4.2x for machine learning applications, and enhance system utilization by up to 27.8%.
Qinyun Tsai, Guanghua Tan, Wangdong Yang, Xianhao He, Yuwei Yan, Keqin Li 0001, Kenli Li 0001
IEEE Trans. Computers2
2024 The End-to-End Fetal Head Circumference Detection and Estimation in Ultrasound Images
abstract
In prenatal examinations, the fetal head circumference (HC) measurement is essential for assessing fetal weight and health conditions. The sonographers obtain the fetal HC manually by fitting peripheral skull ellipse in clinical practice, which is highly subjective, time-consuming, and experience-dependent. Recently, many fetal HC automatic measurement algorithms have been proposed to improve workflow efficiency in prenatal examination. But most automatic measurement algorithms focus on using fetal head segmentation as an intermediate processing step, and HC estimation relies heavily on segmentation results, which causes the accumulation of errors in the above two stages. Independent of the segmentation method, we design a regression network to generate the oriented bounding box to detect the head contour, and directly obtain the fetal head parameters with a pixel-based ellipse regression (PER) loss. Moreover, an effective 3D attention mechanism is integrated into the network to estimate HC more precisely without adding parameters in complex ultrasound images. The extensive experimental results on the public HC18 and our clinical dataset show that the proposed network provides a feasible scheme for end-to-end estimating fetal HC, and avoids the mistake brought by the intermediary processes.
Lei Zhao 0013, Ningshu Li, Guanghua Tan, Jianguo Chen 0001, Shengli Li 0001, Mingxing Duan
IEEE Trans. Comput. Biol. Bioinform.3
2024 SKGC: A General Semantic-Level Knowledge Guided Classification Framework for Fetal Congenital Heart Disease
abstract
Congenital heart disease (CHD) is the most common congenital disability affecting healthy development and growth, even resulting in pregnancy termination or fetal death. Recently, deep learning techniques have made remarkable progress to assist in diagnosing CHD. One very popular method is directly classifying fetal ultrasound images, recognized as abnormal and normal, which tends to focus more on global features and neglects semantic knowledge of anatomical structures. The other approach is segmentation-based diagnosis, which requires a large number of pixel-level annotation masks for training. However, the detailed pixel-level segmentation annotation is costly or even unavailable. Based on the above analysis, we propose SKGC, a universal framework to identify normal or abnormal four-chamber heart (4CH) images, guided by a few annotation masks, while improving accuracy remarkably. SKGC consists of a semantic-level knowledge extraction module (SKEM), a multi-knowledge fusion module (MFM), and a classification module (CM). SKEM is responsible for obtaining high-level semantic knowledge, serving as an abstract representation of the anatomical structures that obstetricians focus on. MFM is a lightweight but efficient module that fuses semantic-level knowledge with the original specific knowledge in ultrasound images. CM classifies the fused knowledge and can be replaced by any advanced classifier. Moreover, we design a new loss function that enhances the constraint between the foreground and background predictions, improving the quality of the semantic-level knowledge. Experimental results on the collected real-world NA-4CH and the publicly FEST datasets show that SKGC achieves impressive performance with the best accuracy of 99.68% and 95.40%, respectively. Notably, the accuracy improves from 74.68% to 88.14% using only 10 labeled masks.
Yuhuan Lu 0002, Guanghua Tan, Bin Pu, Bocheng Liang, Kenli Li 0001, Jagath C. Rajapakse
IEEE J. Biomed. Health Informatics2
2024 TransFSM: Fetal Anatomy Segmentation and Biometric Measurement in Ultrasound Images Using a Hybrid Transformer
abstract
Biometric parameter measurements are powerful tools for evaluating a fetus's gestational age, growth pattern, and abnormalities in a 2D ultrasound. However, it is still challenging to measure fetal biometric parameters automatically due to the indiscriminate confusing factors, limited foreground-background contrast, variety of fetal anatomy shapes at different gestational ages, and blurry anatomical boundaries in ultrasound images. The performance of a standard CNN architecture is limited for these tasks due to the restricted receptive field. We propose a novel hybrid Transformer framework, TransFSM, to address fetal multi-anatomy segmentation and biometric measurement tasks. Unlike the vanilla Transformer based on a single-scale input, TransFSM has a deformable self-attention mechanism so it can effectively process multi-scale information to segment fetal anatomy with irregular shapes and different sizes. We devised a BAD to capture more intrinsic local details using boundary-wise prior knowledge, which compensates for the defects of the Transformer in extracting local features. In addition, a Transformer auxiliary segment head is designed to improve mask prediction by learning the semantic correspondence of the same pixel categories and feature discriminability among different pixel categories. Extensive experiments were conducted on clinical cases and benchmark datasets for anatomy segmentation and biometric measurement tasks. The experiment results indicate that our method achieves state-of-the-art performance in seven evaluation metrics compared with CNN-based, Transformer-based, and hybrid approaches. By Knowledge distillation, the proposed TransFSM can create a more compact and efficient model with high deploying potential in resource-constrained scenarios. Our study serves as a unified framework for biometric estimation across multiple anatomical regions to monitor fetal growth in clinical practice.
Lei Zhao 0013, Guanghua Tan, Bin Pu, Qianghui Wu, Hongliang Ren 0001, Kenli Li 0001
IEEE J. Biomed. Health Informatics2
2024 FARN: Fetal Anatomy Reasoning Network for Detection With Global Context Semantic and Local Topology Relationship
abstract
Accurate recognition of fetal anatomical structure is a pivotal task in ultrasound (US) image analysis. Sonographers naturally apply anatomical knowledge and clinical expertise to recognizing key anatomical structures in complex US images. However, mainstream object detection approaches usually treat each structure recognition separately, overlooking anatomical correlations between different structures in fetal US planes. In this work, we propose a Fetal Anatomy Reasoning Network (FARN) that incorporates two kinds of relationship forms: a global context semantic block summarized with visual similarity and a local topology relationship block depicting structural pair constraints. Specifically, by designing the Adaptive Relation Graph Reasoning (ARGR) module, anatomical structures are treated as nodes, with two kinds of relationships between nodes modeled as edges. The flexibility of the model is enhanced by constructing the adaptive relationship graph in a data-driven way, enabling adaptation to various data samples without the need for predefined additional constraints. The feature representation is further refined by aggregating the outputs of the ARGR module. Comprehensive experimental results demonstrate that FARN achieves promising performance in detecting 37 anatomical structures across key US planes in tertiary obstetric screening. FARN effectively utilizes key relationships to improve detection performance, demonstrates robustness to small-scale, similar, and indistinct structures, and avoids some detection errors that deviate from anatomical norms. Overall, our study serves as a resource for developing efficient and concise approaches to model inter-anatomy relationships.
Lei Zhao 0013, Guanghua Tan, Qianghui Wu, Bin Pu, Hongliang Ren 0001, Shengli Li 0001, Kenli Li 0001
IEEE J. Biomed. Health Informatics2
2024 Needle Trajectory Prediction for Percutaneous Kidney Biopsy in 5G-Powered Teleultrasound Navigation System
abstract
Needle insertion is a critical component of many remote surgical procedures, including biopsies, injections, neurosurgery, and brachytherapy cancer treatments. However, precise visualization of the biopsy needle trajectory remains challenging due to specular reflection, speckle noise, and needle-like anatomical features. This paper proposes a visual feedback prediction framework for ultrasound-assisted percutaneous kidney biopsy in 5G remote surgery, aiming to enhance operator confidence, reduce procedure time, and minimize the risk of unintended bleeding. Building upon this framework, we design a Lightweight-Accuracy Needle Trajectory Prediction (LA-NTP) model by minimizing the backbone and optimizing the multi-module prediction process, incorporating innovative training strategies (i.e., angle-aware geometric and trajectory augmentation losses). The experimental results demonstrate that it achieves competitive performance with only 20.3% of the model size of the previous best real-time method and a 3.7-fold increase in inference speed. Even in challenging scenarios involving large insertion depths and steep angles, our method provides stable and precise navigation.
Lei Zhao 0013, Guanghua Tan, Jiewen Lai, Chwee Ming Lim, Weng Kin Wong, Hongliang Ren 0001, Kenli Li 0001
IEEE Trans. Mob. Comput.2
2023 GCNet: Ground Collapse Prediction Based on the Ground-Penetrating Radar and Deep Learning Technique
abstract
Manual methods to detect and identify potential ground collapse hazards have a high volume of workload and rely on manual skills and experience strongly, with substantial inconsistencies. This paper proposes an automatic recognition method of potential underground hazards and designs an object detection-based algorithm to predict the hidden danger of ground collapse, called GCNet. The GCNet is a ground collapse hazard detection network, which is developed to detect underground hazards, including cavities, poor soil quality, pipelines, and soil background. The GCNet uses the deep residual network ResNet101 and feature pyramid network (FPN) to extract features and a task coordination network (TCN) to identify the category and location of hidden underground hazards. Further, a feature enhancement method based on the regional binary pattern is proposed to improve the accuracy of the proposed model by expanding the ground-penetrating radar (GPR) data by adding multi-level features to GPR images. The experiments are conducted on a large amount of real GPR data and experiment results show that the proposed automatic recognition method can surpass the existing deep learning-based methods in hidden underground hazard recognition and identification.
Xu Zhou 0001, Guanghua Tan, Shenghong Yang, Kenli Li 0001
Int. J. Pattern Recognit. Artif. Intell.3
2023 Fetal Ultrasound Standard Plane Detection With Coarse-to-Fine Multi-Task Learning
abstract
The ultrasound standard plane plays an important role in prenatal fetal growth parameter measurement and disease diagnosis in prenatal screening. However, obtaining standard planes in a fetal ultrasound video is not only laborious and time-consuming but also depends on the clinical experience of sonographers to a certain extent. To improve the acquisition efficiency and accuracy of the ultrasound standard plane, we propose a novel detection framework that utilizes both the coarse-to-fine detection strategy and multi-task learning mechanism for feature-fused images. First, traditional manually-designed features and deep learning-based features are fused to obtain low-level shared features, which can enhance the model's feature expression ability. Inspired by the process of human recognition, ultrasound standard plane detection is divided into a coarse process of plane type classification and a fine process of standard-or-not detection, which is implemented via an end-to-end multi-task learning network. The region-of-interest area is also recognised in our detection framework to suppress the influence of a variable maternal background. Extensive experiments are conducted on three ultrasound planes of the first-class fetal examination, i.e., the femur, thalamus, and abdomen ultrasound images. The experiment results show that our method outperforms competing methods in terms of accuracy, which demonstrates the efficacy of the proposed method and can reduce the workload of sonographers in prenatal screening.
Juncheng Guo, Guanghua Tan, Fan Wu 0016, Huaxuan Wen, Kenli Li 0001
IEEE J. Biomed. Health Informatics2
2022 A Novel Deep Learning Framework for Automatic Recognition of Thyroid Gland and Tissues of Neck in Ultrasound Image
abstract
Recognition of thyroid glands and tissues of the neck is vital for screening related diseases in ultrasound videos. This task is subjective, challenging, and dependent on the experience of sonographer in current clinical practice. The purpose is to develop a fully automated thyroid gland and tissues of neck recognition framework to assist doctors in distinguishing the boundaries of different tissues. In this paper, we propose a novel deep learning framework that consists of a feature extraction network, region proposal network, object detection head, and spatial pyramid RoIAlign-based segmentation head. Designed spatial pyramid RoIAlign can efficiently capture local and global context features, and aggregates the multiple context information that makes the result much more reliable. A large dataset is constructed to train the proposed method. The performance is evaluated using the COCO metrics. The experimental results demonstrate that the proposed deep learning method can effectively realize the automatic recognition of the thyroid gland and tissues of neck in ultrasound videos. Considering the clinical practical application scenarios, we developed an automatic recognition system of thyroid and neck tissue based on edge computing, which can expediently assist doctors in distinguishing the boundaries between different tissues.
Laifa Ma, Guanghua Tan, Hongxia Luo, Qing Liao 0001, Shengli Li 0001, Kenli Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Edge-Aware Multi-Scale Progressive Colorization
abstract
Image colorization recovers a colorful image from a grayscale one. Trained by large-scale datasets, recent deep neural networks based methods can produce impressive colorful images. However, they usually directly train a single network using training images of fixed resolution. It is hard for such a single network to learn the features of different scales for colorization. Moreover, they are prone to generate color bleedings and blurry details around objects boundaries. To address these problems, we propose a novel edge-aware multi-scale progressive network (EMSPN). The key idea is to train a series of multi-scale networks in a progressive manner, so that the network in finer scales can leverage the outputs of its previous scale. In addition, we also propose an edge-map loss to effectively prevent bleedings and blurs around the image edges. Experimental results show that our work outperforms existing methods and achieves state-of-the-art results.
Guanghua Tan, Yi Xiao 0004, Fangqiang Xu, Andrew Chi-Sing Leung
ICASSP2
2021 Age Estimation Using Aging/Rejuvenation Features With Device-Edge Synergy
abstract
Estimating human age is a challenging task in computer vision and most researchers are trying to make age estimation via a static facial image. However, it ignores the fact that the age of a person is the specific representation of aging. In this paper, we attempt to explore the aging/rejuvenation (AR) characteristics of faces for age estimation and we called the whole network as AR-Net. Firstly, we use GAN model for learning a manifold of the aging/rejuvenation process to a face dataset with preserving personalized face features (e.g.,gender, race). Secondly, we seek the correlated aging/rejuvenation characteristics from a narrow age interval, (e.g.,((0-100)$\rightarrow $(0-5), (5-10),…, (90-100)). Thirdly, the fine-tuned GAN is used to generate aging/rejuvenation features of all age groups and these features are applied to train corresponding ELM regressors. AR-Net is deployed on every edge server, and all AR-Nets are trained offline. Afterwards, our AR-Net is constantly updated based on the face dataset collected by the edge sensors. Finally, enormous experiments on Morph-II, CACD, and captured facial dataset have been conducted to verify the performance of our fine-tuned AR-Net and the experimental results show that the approach enhanced than the current state of the art methods.
Mingxing Duan, Aijia Ouyang, Guanghua Tan, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 CacheTrack-YOLO: Real-Time Detection and Tracking for Thyroid Nodules and Surrounding Tissues in Ultrasound Videos
abstract
To accurately detect and track the thyroid nodules in a video is a crucial step in the thyroid screening for identification of benign and malignant nodules in computer-aided diagnosis (CAD) systems. Most existing methods just perform excellent on static frames selected manually from ultrasound videos. However, manual acquisition is labor-intensive work. To make the thyroid screening process in a more natural way with less labor operations, we develop a well-designed framework suitable for practical applications for thyroid nodule detection in ultrasound videos. Particularly, in order to make full use of the characteristics of thyroid videos, we propose a novel post-processing approach, called Cache-Track, which exploits the contextual relation among video frames to propagate the detection results into adjacent frames to refine the detection results. Additionally, our method can not only detect and count thyroid nodules, but also track and monitor surrounding tissues, which can greatly reduce the labor work and achieve computer-aided diagnosis. Experimental results show that our method performs better in balancing accuracy and speed.
Xiangqiong Wu, Guanghua Tan, Ningbo Zhu, Zhilun Chen, Huaxuan Wen, Kenli Li 0001
IEEE J. Biomed. Health Informatics2
2020 Sketchppnet: A Joint Pixel and Point Convolutional Neural Network For Low Resolution Sketch Image Recognition
abstract
Sketch recognition using deep neural networks have become a recent trend. However, traditional pixel (image) based convolutional neural networks show poor recognizing performance on low resolution (LR) sketch image due to the loss of image details. To solve this problem, we propose a joint pixel and point convolutional neural network for LR sketch image recognition. The network, equipped with both image convolution and point convolution, can simultaneously handle both the image and point representation of sketches. Furthermore, we propose a hybrid classifier, a corresponding loss function, and a training scheme to better extract features for recognition. Experimental results show that our method outperforms state-of-art deep neural networks.
Xianyi Zhu, Yi Xiao 0004, Yan Zheng 0003, Guanghua Tan, Shizhe Zhou
ICASSP4
2020 Deep Parametric Active Contour Model for Neurofibromatosis Segmentation
Xiangqiong Wu, Guanghua Tan, Kenli Li 0001, Shengli Li 0001, Huaxuan Wen, Xianyi Zhu, Wenli Cai
Future Gener. Comput. Syst.2
2020 Fingerprint liveness detection based on guided filtering and hybrid image analysis
abstract
Fingerprints are widely used for biometric recognition. However, many spoofing attacks based on an artificially made fingerprint occur. In this study, the authors propose an approach to detect fingerprint liveness which uses the guided filtering and hybrid image analysis. This study deals with the problem of ignoring the contribution that is brought by the sharp features when analysing the denoised image. The method described utilises both the enhanced sharp features and denoised features from the hybrid images to get better results. The input fingerprint is pre‐processed by region of interest extraction and then is filtered by a guidance image for obtaining the denoised image. Then, histogram equalisation is introduced to eliminate the impact of illumination condition. The authors extract the co‐occurrence of adjacent local binary pattern features from both the cropped images and the denoised images. Whilst concatenating both the features together to form a long feature, t‐Distributed Stochastic Neighbour Embedding is applied to reduce the data dimension. The authors consider the fingerprint liveness detection as a two‐class classification problem and use support vector machine with radial basis function kernel to solve this problem. The authors evaluate the experiments on three benchmark data sets. Experimental results demonstrate that the accuracy of the proposed method can outperform most of the state‐of‐art methods.
Guanghua Tan, Xianyi Zhu, Xiangqiong Wu
IET Image Process.1
2020 Game theory-based optimization of distributed idle computing resources in cloud environments
Gang Liu 0038, Guanghua Tan, Kenli Li 0001, Anthony T. Chronopoulos
Theor. Comput. Sci.3
2019 PA-RetinaNet: Path Augmented RetinaNet for Dense Object Detection
Guanghua Tan, Yi Xiao 0004
ICANN (2)1
2019 Action Recognition Based on Divide-and-Conquer
Guanghua Tan, Yi Xiao 0004
ICANN (3)1
2018 PatchSwapper: A novel real-time single-image editing technique by region-swapping
Shizhe Zhou, Chengfeng Zhou, Yi Xiao 0004, Guanghua Tan
Comput. Graph.4
2017 Face alignment under occlusion based on local and global feature regression
Songrui Guo, Guanghua Tan, Huawei Pan, Chunming Gao
Multim. Tools Appl.2
2017 A free shape 3d modeling system for creative design based on modified catmull-clark subdivision
Guanghua Tan, Xianyi Zhu, Xuefei Liu
Multim. Tools Appl.1
2016 A High Invariance Motion Representation for Skeleton-Based Action Recognition
abstract
Human action recognition is very important and significant research work in numerous fields of science, for example, human–computer interaction, computer vision and crime analysis. In recent years, relative geometry features have been widely applied to the description of relative relation of body motion. It brings many benefits to action recognition such as clear description, abundant features etc. But the obvious disadvantage is that the extracted features severely rely on the local coordinate system. It is difficult to find a bijection between relative geometry and skeleton motion. To overcome this problem, many previous methods use relative rotation and translation between all skeleton pairs to increase robustness. In this paper we present a new motion representation method. It establishes a motion model based on the relative geometry with the aid of special orthogonal group SO(3). At the same time, we proved that this motion representation method can establish a bijection between relative geometry and motion of skeleton pairs. After the motion representation method in this paper is used, the computation cost of action recognition reduces from the two-way relative motion (motion from A to B and B to A) to one-way relative motion (motion from A to B or B to A) between any skeleton pair, namely, permutation problem [Formula: see text] is simplified into combinatorics problem [Formula: see text]. Finally, the experimental results of the three motion datasets are all superior to present skeleton-based action recognition methods.
Songrui Guo, Huawei Pan, Guanghua Tan, Chunming Gao
Int. J. Pattern Recognit. Artif. Intell.3
2016 A novel image matting method using sparse manual clicks
Guanghua Tan
Multim. Tools Appl.1
2014 Adaptive Segmentation Approach for Human Action Data
abstract
Temporal segmentation of human motion data is an essential preparation process for action recognition. Due to the variability in the temporal scale of human action and the complexity of representing articulated motion, the research of it encounters many difficulties. Especially, when the number of behaviors contained in the motion sequences is unknown in advance, traditional algorithms cannot segment sequences successfully. In this paper, we extend previous works on change-points detection by probabilistic principle component analysis (PPCA). Based on it, an algorithm which is an extension of PCA and Maximum Mean Discrepancy between samples is proposed for estimating the cluster number. Finally, we optimize our approach and detect cyclic units of each action by aligned cluster analysis. We evaluate and compare the approach with the state-of-the-art methods on Synthetic data, Motion Capture Dataset and Kinect data. Experimental results demonstrate the effectiveness of our approach.
Chunming Gao, Changhui Li, Guanghua Tan, Songrui Guo
Int. J. Pattern Recognit. Artif. Intell.3
2014 Saliency-Based Unsupervised Image Matting
abstract
Spectral matting is the state-of-the-art matting method and can well solve the highly under-conditioned matte problem without manual intervention. However, it suffers from huge computation cost and inaccurate alpha matte. This paper presents a modified spectral matting method which combines saliency detection algorithm to get a higher accuracy of alpha matte with less computational cost. First, the saliency detection algorithm is used to detect general locations of foreground objects. For saliency detection method, original two-stage scheme is replaced by feedback scheme to get a more suitable saliency map for unsupervised image matting. Next, matting components are obtained through a linear transformation of the smallest eigenvectors of the matting Laplacian matrix. Then, the improved saliency map is used for grouping matting components. Finally, the alpha matte is obtained based on matte cost function. Experiments show that the proposed method outperforms the state-of-the-art methods based on spectral matting both in speed and alpha matte accuracy.
Guanghua Tan, Chunming Gao, Liyuan Zhuo
Int. J. Pattern Recognit. Artif. Intell.1
2010 Image driven shape deformation using styles
abstract
In this paper, we propose an image driven shape deformation approach for stylizing a 3D mesh using styles learned from existing 2D illustrations. Our approach models a 2D illustration as a planar mesh and represents the shape styles with four components: the object contour, the context curves, user-specified features and local shape details. After the correspondence between the input model and the 2D illustration is established, shape stylization is formulated as a style-constrained differential mesh editing problem. A distinguishing feature of our approach is that it allows users to directly transfer styles from hand-drawn 2D illustrations with individual perception and cognition, which are difficult to identify and create with 3D modeling and editing approaches. We present a sequence of challenging examples including unrealistic and exaggerated paintings to illustrate the effectiveness of our approach.
Guanghua Tan, Wei Chen 0001
J. Zhejiang Univ. Sci. C1