VLDB 2026 Research / reviewers in the wild / expert
Yonghuai Liu
dblp:03/357
· DBLP profile ↗
140ranked-venue papers
37as first author
43since 2021 · last 2026
0000-0002-3774-2134ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 23 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 63 · 10 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 3 first-author · 17 since 2021Systems, architecture and hardware · 13 · 9 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-authorDatabases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FAGNet: A full attention-modulated geometric learning network for 3D pose estimation
Chongnan Yu, Yonghuai Liu, Yongqin Zhang |
Image Vis. Comput. | 6 |
| 2026 | MHAN: Multi-head hybrid attention network for facial expression recognition
Tianbo Han, Songling Liu, Muhammad Shahroz Ajmal, Yongqin Zhang, Yonghuai Liu |
Pattern Recognit. | 7 |
| 2026 | VIM-Net: A voxel-interaction multimodal network for 3D object detection
Minghan Wang, Xijiong Wang, Yonghuai Liu, Ardhendu Behera, Baowen Zhang |
Pattern Recognit. | 5 |
| 2026 | Multi-Granularity Topological Reasoning for Anatomically Consistent Vasculature ParsingabstractQuantitative analysis of retinal vascular morphology is vital for clinical decision-making and the investigation of systemic diseases. Central to this process is the accurate segmentation of retinal arteries and veins (A/V) from the background, a task challenged by substantial variations in vessel calibers and the presence of low-contrast or ambiguous structures in fundus images, especially in ultra-wide field imaging where peripheral distortions and large-scale anatomical variability are pronounced. These factors often lead to fragmented semantic representations and topological inconsistencies in automated segmentation outputs. To address these limitations, we propose Ultra, a multi-granularity topological reasoning network designed for precise A/V segmentation. Ultra adopts a cascaded two-stage architecture: PriorNet generates coarse, multi-scale vascular priors that provide structural guidance, while RefineNet performs topology-aware segmentation refinement. To further enforce topological coherence, we propose the neighboring pixel connectivity regularization (NICER) layer, which selectively integrates local connectivity information predicted by the proposed connectivity prediction union (CPU) module. This connectivity is employed as auxiliary supervision through a pixel-wise local connectivity loss, reinforcing structural reasoning and promoting anatomically consistent vascular topology inference. Extensive experiments on ultra-wide field fundus imaging (UWF) datasets demonstrate that Ultra achieves state-of-the-art performance in A/V segmentation and topological preservation. Moreover, Ultra generalizes well to conventional color fundus photography (CFP) datasets, underscoring its robustness and broad applicability. Code is publicly available at: https://github.com/iMED-Lab/Ultra. Lei Mou, Yonghuai Liu, Zhuoting Xu, Hao Zhang 0113, Yalin Zheng, Jiang Liu 0001, Huazhu Fu, Yitian Zhao |
IEEE Trans. Image Process. | 2 |
| 2026 | HAIMNet: A Hierarchical Adaptive Interaction Modulation Network for Low-Light Image EnhancementabstractLow-light image enhancement (LLIE) is essential for enabling reliable nighttime visual perception and improving the performance of downstream vision tasks, including object detection and image segmentation. Under complex illumination conditions, low-light images often suffer from insufficient luminance, loss of structural details, and unstable color reproduction. Existing methods struggle to simultaneously restore luminance, texture, and color in a coherent manner. This paper proposes a Hierarchical Adaptive Interaction Modulation Network (HAIMNet) designed for LLIE. The proposed method decouples luminance and chromaticity in the Horizontal/Vertical-Intensity (HVI) color space, and enhances luminance-texture consistency through an inter-branch attention-modulation block (IAMB). Furthermore, a cross-branch gated affine fusion module (CGAF) is introduced to calibrate features between luminance and chromatic-structural representations, reduce color deviations, and enhance perceptual consistency. Extensive experiments on 11 public datasets demonstrate the effectiveness, robustness, and generalization capability of HAIMNet. The enhanced results exhibit high naturalness and stability under extremely dark and complex illumination conditions. Our code is available at: https://github.com/ZekeWang13/HAIMNet. Meijia Guo, Qiuwei Chen, Yajie Hao, Yonghuai Liu |
IEEE Trans. Image Process. | 6 |
| 2026 | Super-Resolution Reconstruction of OCTA via Multi-Field-of-View Representation LearningabstractHigh-resolution Optical Coherence Tomography Angiography (OCTA) images are essential for morphological analysis and biomarker measurement of the retinal vasculature. They can also provide underlying biomarkers for the accurate analysis of eye-related diseases. The trade-off between the high resolution (HR) and large scanning field-of-view (FOV) is a long-standing problem for OCTA image instrument. A large FOV image provides more retinal information with shorter acquisition time but often suffers from low resolution (LR), high scatter noise, and poor vascular contrast. In order to obtain HR OCTA images with larger FOV, we propose a novel self-similar dynamic domain adaptation network based on cross-field-of-view representation learning. The network enables LR images (i.e., $6\times \text{6}\,\text{mm}^{2}$) to learn HR image (i.e., $3\times \text{3}\,\text{mm}^{2}$) feature representations specialized for OCTA by constructing feature mapping relations for cross-field-of-view OCTA scans. To be specific, a multiple random degradation model is proposed on HR images to generate various synthetic LR images. Further, we propose a dynamic domain adaptation framework that prompts feature dynamic alignment of the LR image reconstruction results with those of synthetic LR images. Finally, a novel self-similar supervision loss is proposed to optimize the reconstruction results from LR to HR by exploiting the similarity between vessels in different regions. Experimental results on three OCTA datasets show that the proposed method surpasses existing state-of-the-art ones, significantly enhancing retinal structure segmentation and disease classification. Our OCTA dataset (the first dataset in this research area with paired $3\times 3$ and $6\times \text{6}\,\text{mm}^{2}$ OCTA images) and code are publicly available. Huaying Hao, Shaoyi Leng, Yanda Meng, Yonghuai Liu, Yalin Zheng, Huazhu Fu, Jiong Zhang 0004, Quanyong Yi, Yue Liu 0005, Jingfeng Zhang, Yitian Zhao |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Enhancing Trustworthiness of Semantic Segmentation in Cataract Surgery Videos via Intra-Phase Label PropagationabstractAccurate segmentation of semantic features is a pivotal procedure for cataract surgery assistance, surgical skill assessment and related applications. However, previous studies have failed to consider the instance-level feature similarity of instruments across different surgical phases in cataract surgery videos, leading to unreliable decision-making regarding instrument categories. In this study, we propose a label propagation framework to effectively leverage the consistency of phase-specific instruments, which utilizes the initial frame labels from each surgical phase to predict masks for the remaining frames, achieving precise and trustworthy semantic segmentation of cataract surgery videos. Specifically, we design a pseudo-label generation and filtering strategy to automatically obtain highly reliable initial frame labels for each surgical phase. In addition, we establish a fixed-size memory bank with an adaptive update module to ensure long-term applicability in real surgical environments. To address the common problem of blurred edges in cataract surgery scenes, we develop a semantic edge perception module to allow the model to focus on and distinguish the edges of different objects. The proposed method achieved an mIoU of 80.7% and 88.8% on a publicly available dataset (14 categories) and a private dataset (12 categories) with a total of 9,723 frames, respectively, significantly outperforming the state-of-the-art methods and other label propagation-based approaches. Furthermore, our method minimizes memory consumption and maintains about 30 FPS while processing long video sequences. Mingen Zhang, Xu Chen 0030, Botian Zheng, Donghan Wu, Jinxian Zhang, Yufei Wu 0013, Yonghuai Liu, Yitian Zhao |
IEEE J. Biomed. Health Informatics | 8 |
| 2026 | Progressive Distillation for Incremental Learning in Corneal Confocal Microscopy SegmentationabstractThe morphological changes of corneal structures captured by corneal confocal microscopy (CCM), such as corneal nerves, Langerhans cells, stromal cells, etc., are closely related to various ocular and systemic diseases. Current CCM segmentation methods primarily focus on single-task, which limits their broad applicability in clinical practice. The absence of a standardized benchmark further presents a significant challenge in evaluating new methods. To this end, this paper presents a novel incremental learning-based approach for multi-structure segmentation in CCM images and a new benchmark. Specifically, we first propose a data fingerprint distillation (FIND) module to encode task-relevant knowledge by extracting compact representations of structures from CCM images via structural importance mapping. Building on FIND, we propose a progressive task-guided adapter learning (ProTA) strategy, which refines the model's representation of structures through a series of "easy-to-hard" distillation stages. ProTA dynamically adjusts the scope of task-relevant knowledge extracted by FIND, thereby improving the model's ability to accurately discriminate between multiple structures while enhancing knowledge transfer efficiency. Extensive experiments demonstrate that the proposed method achieves the state-of-the-art performance in terms of all corneal structures segmentation. We also demonstrate our approach's plug-and-play capability across four other medical image modalities, suggesting its potential as a general incremental learning tool. Additionally, this work seeks to provide a benchmark tool comprising a comprehensive dataset and their fine manual annotation, as well as unified benchmarking evaluations for state-of-the-art methods. All the dataset, source code and evaluation tool are publicly available at https://github.com/iMED-Lab/CCM-Pro. Hongshuo Li, Baikai Ma, Lei Mou, Yonghuai Liu, Qinxiang Zheng, Yitian Zhao |
IEEE Trans. Medical Imaging | 4 |
| 2026 | CartoonTalk: A Speech-Driven Animation Model for Cartoon Face GenerationabstractDue to substantial differences in appearance and movement between cartoon faces and real human faces, directly applying existing speech-driven speaker generation techniques to cartoon face images often fails to accurately capture key points, resulting in distorted expressions and animation artifacts. To address these challenges, this article proposes a speech-driven cartoon face animation generation model called CartoonTalk, which integrates 2D cartoon key points with 3D motion coefficients. The model comprises three modules: face annotation, motion extraction, and 3D facial rendering. The face annotation module first annotates a cartoon image, capturing its 2D key points. Subsequently, the motion extraction module derives 3D motion coefficients for head and facial expressions from speech data. Finally, the 3D facial rendering module uses these 2D and 3D features to generate the animated cartoon face. Quantitative and qualitative evaluations demonstrate that our model significantly outperforms existing approaches in both lip-sync accuracy and visual quality, establishing new state-of-the-art benchmarks on cartoon datasets. Qiuwei Chen, Yongqin Zhang, Yonghuai Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | LANet: Enhancing Plant Disease Recognition Through Transfer Learning and Layer AttentionabstractAccurate recognition of plant diseases plays a critical role in maintaining agricultural productivity and food security. This study proposes LANet, an innovative network architecture aimed at improving the precision of identifying and categorizing plant diseases from images. LANet consists of two primary components. Firstly, Layer Attention applies self-attention across feature layers at different scales to capture the weights between cross-scale features, addressing the data loss caused by downsampling and enhancing plant species recognition. Secondly, we introduce a transfer learning method, where a plant species classification network is initially trained, its parameters are then frozen, and several convolutional layers are added to train the disease classification network. This allows the model to leverage the learned classification information for more accurate disease recognition. Our model demonstrates an accuracy of 85.62%, outperforming nine other models significantly. Comparative analyses using the public PlantVillage dataset highlight the superior performance of our method in disease recognition. Ardhendu Behera, Fuzhong Li, Wuping Zhang, Yonghuai Liu |
IPAS | 6 |
| 2025 | Uncertainty Driven Sampling to Handle Intra-class Imbalance Part Segmentation in WheatabstractWe introduce a novel method to address intra-class imbalance in 3D point cloud segmentation of wheat, focusing on distinguishing between ear and non-ear parts. Variability in plant structure, influenced by factors such as curvature and shape, often leads to data imbalance which complicates segmentation tasks. Our approach utilizes Monte Carlo Dropout to identify and prioritize uncertain samples at the end of each training epoch, employing uncertainty-driven sampling to select samples with the lowest confidence. These samples undergo augmentation through scaling and leaf crossover techniques, enhancing their representation in the training set. Our comparative evaluations demonstrate that this strategy significantly improves the mean Intersection over Union (mIoU) and segmentation accuracy, thereby increasing model robustness for complex 3D plant structures. Reena, John H. Doonan, Huaizhong Zhang, Yonghuai Liu, Fiona M. K. Corke |
IPAS | 4 |
| 2025 | Interweaving Insights: High-Order Feature Interaction for Fine-Grained Visual RecognitionabstractThis paper presents a novel approach for Fine-Grained Visual Classification (FGVC) by exploring Graph Neural Networks (GNNs) to facilitate high-order feature interactions, with a specific focus on constructing both inter- and intra-region graphs. Unlike previous FGVC techniques that often isolate global and local features, our method combines both features seamlessly during learning via graphs. Inter-region graphs capture long-range dependencies to recognize global patterns, while intra-region graphs delve into finer details within specific regions of an object by exploring high-dimensional convolutional features. A key innovation is the use of shared GNNs with an attention mechanism coupled with the Approximate Personalized Propagation of Neural Predictions (APPNP) message-passing algorithm, enhancing information propagation efficiency for better discriminability and simplifying the model architecture for computational efficiency. Additionally, the introduction of residual connections improves performance and training stability. Comprehensive experiments showcase state-of-the-art results on benchmark FGVC datasets, affirming the efficacy of our approach. This work underscores the potential of GNN in modeling high-level feature interactions, distinguishing it from previous FGVC methods that typically focus on singular aspects of feature representation. Our source code is available at https://github.com/Arindam-1991/I2-HOFI. Arindam Sikdar, Yonghuai Liu, Siddhardha Kedarisetty, Yitian Zhao, Amr Ahmed 0002, Ardhendu Behera |
Int. J. Comput. Vis. | 2 |
| 2025 | Beyond the eye: A relational model for early dementia detection using retinal OCTA images
Shouyue Liu, Jinkui Hao, Yonghuai Liu, Huazhu Fu, Yitian Zhao |
Medical Image Anal. | 5 |
| 2025 | 3D microvascular reconstruction in retinal OCT angiography images via domain-adaptive learning
Jiong Zhang 0004, Yonghuai Liu, Dan Zhang 0026, Jianyang Xie, Tao Chen 0003, Yalin Zheng, Huazhu Fu, Yitian Zhao |
Pattern Recognit. | 3 |
| 2025 | Rethinking Data Augmentation for Single-Source Domain Generalization in OCT Image SegmentationabstractDomain shifts between samples acquired with different instruments are one of the major challenges in accurate segmentation of Optical Coherence Tomography (OCT) images. Given that OCT images may be acquired with different devices in different clinical centers, this study presents astyle and structure data augmentation (SSDA) method to improve the adaptability of segmentation models. Inspired by our initial analysis of OCT domain differences, we propose an innovative hypothesis that domain shifts are primarily due to differences in image style and anatomical structure, which further guides the design of our method. By designing a modality-specific NURBS curve for style enhancement and implementing global and local elastic deformation fields, SSDA addresses both stylistic and structural variations in OCT data. Global deformations simulate changes in retinal curvature, while local deformations model layer-specific changes observed in OCT images. We validate our hypothesis through a comprehensive evaluation conducted on five OCT data domains, each differing in device type and imaging conditions. We train models on each of these domains for single-domain generalisation experiments and evaluate performance on the remaining unseen domains. The results show that SSDA outperforms existing methods when segmenting OCT images from different sources with different requirements for retinal layer segmentation. Specifically, across five different source domain generalisation experiments, SSDA achieves approximately 1.6% higher Dice and 2.6% improved MIOU, underscoring its superior segmentation accuracy and robust generalisation across all evaluated unseen domains. Shaodong Ma, Yonghuai Liu, Yuhui Ma, Lei Mou, Yitian Zhao |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | SLPDR: A Benchmark for Ship License Plate Detection and RecognitionabstractShip identification is a prerequisite for the intelligent management of maritime transportation, yet existing research is confined to broad ship detection and categorization, which only provides the ship’s location or type instead of its identification. Inspired by the research on the Car License Plate (CLP), we make the first attempt to propose the concept of the Ship License Plate (SLP). In addition, the limited data hinders research on ship identification. To overcome this obstacle, we construct the first large-scale Ship License Plate Detection and Recognition (SLPDR) dataset, which contains 1,472 ship identities and 88,862 images. In addition, this paper proposes an SLP detection model named YOLO-SSA and evaluates this model as well as typical detection methods on the SLPDR dataset. The experimental results demonstrate that the proposed YOLO-SSA achieves better SLP detection performance by enhancing the features where ships and SLPs are located. Furthermore, we explore the prospective applications of SLPs in intelligent maritime transportation, including ship monitoring and berth management. Project web page: https://vsislab.github.io/SLPDR/ Youmei Zhang, Ran Song 0001, Yonghuai Liu, Ardhendu Behera, Mingxin Zhang 0006, Wei Zhang 0021 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Clinical Insight-Augmented Multi-View Learning for Alzheimer's Detection in Retinal OCTA ImagesabstractAlzheimer’s disease (AD) poses a significant global challenge, with a notable absence of accessible and cost-effective diagnostic tools for widespread AD detection. The retina, mirroring the brain in anatomy and physiology, has emerged as a potential avenue for rapid AD identification through retinal imaging. The current retinal image-based AD detection methods usually focus primarily on the macular area, but ignore the potential value that the optic disc region may have for the detection task. In this study, we leverage both macular- and disc-centered OCTA images and propose a multi-region fusion framework for AD detection. Based on clinical evidence, we integrate handcrafted features into the framework to improve model performance and interpretability. Specifically, vascular morphological parameters extracted from the macular and disc regions are used as input to a revalued KNN model to improve predictive capabilities. Furthermore, recognizing the significance of extracting and utilizing complementary information from the macular and optic disc regions, we propose an uncertainty-guided strategy based on Dempster-Shefer Theory (DST) to fuse knowledge from different regions. This approach considers each region’s forecast quality and significantly improves the effectiveness and robustness of the model. Through comparative analysis with existing methods, we have demonstrated that our method outperforms the state-of-the-art ones and provides more valuable pathological evidence for the association between retinal vascular changes and AD. Yuandi Zhang, Jinkui Hao, Botian Zheng, Yonghuai Liu, Yanda Meng, Jiong Zhang 0004, Yang Chen 0008, Yitian Zhao |
BIBM | 4 |
| 2024 | A Hyperreflective Foci Segmentation Network for OCT Images with Multi-dimensional Semantic Enhancement
Xingguo Wang, Yuhui Ma, Yalin Zheng, Jiong Zhang 0004, Yonghuai Liu, Yitian Zhao |
MICCAI (1) | 6 |
| 2024 | HRNet: 3D object detection network for point cloud with hierarchical refinement
Ran Song 0001, Yonghuai Liu |
Pattern Recognit. | 6 |
| 2024 | Clustering-inspired channel selection method for weakly supervised object localization
Xiangru Qiao, Zhiquan Li, Sidong Wu, Yonghuai Liu, Huaizhong Zhang |
Pattern Recognit. Lett. | 7 |
| 2024 | Learning Common Semantics via Optimal Transport for Contrastive Multi-View ClusteringabstractMulti-view clustering aims to learn discriminative representations from multi-view data. Although existing methods show impressive performance by leveraging contrastive learning to tackle the representation gap between every two views, they share the common limitation of not performing semantic alignment from a global perspective, resulting in the undermining of semantic patterns in multi-view data. This paper presents CSOT, namely Common Semantics via Optimal Transport, to boost contrastive multi-view clustering via semantic learning in a common space that integrates all views. Through optimal transport, the samples in multiple views are mapped to the joint clusters which represent the multi-view semantic patterns in the common space. With the semantic assignment derived from the optimal transport plan, we design a semantic learning module where the soft assignment vector works as a global supervision to enforce the model to learn consistent semantics among all views. Moreover, we propose a semantic-aware re-weighting strategy to treat samples differently according to their semantic significance, which improves the effectiveness of cross-view contrastive representation learning. Extensive experimental results demonstrate that CSOT achieves the state-of-the-art clustering performance. Qian Zhang 0076, Lin Zhang 0041, Ran Song 0001, Runmin Cong, Yonghuai Liu, Wei Zhang 0021 |
IEEE Trans. Image Process. | 5 |
| 2024 | Exploiting Label Uncertainty for Enhanced 3D Object Detection From Point CloudsabstractAccurate detection of objects from LiDAR point clouds is crucial for autonomous driving and environment modeling. However, uncertainties in ground truth labels due to occlusions, sparsity, and truncation can hinder model training and performance. This paper introduces two strategies to address these issues: 1) Soft Regression Loss (SoRL) and 2) Discrete Quantization Sampling (DQS). SoRL utilizes Gaussian distributions for object predictions, measuring uncertainty based on the probability of ground truth labels within these distributions. This method effectively accounts for deviations in object location and orientation. Meanwhile, DQS introduces uncertainty scores for dynamic sample selection, aiming to refine the quality of positive samples for regression. Based on the proposed modules, we design a lightweight multi-stage object detection framework. Notably, these modules can enhance existing 3D object detection methods without affecting significantly inference speeds. Experiments over benchmark datasets show the effectiveness of our method, especially for cars in sparse point clouds. Yonghuai Liu, Ardhendu Behera, Ran Song 0001, Hejin Yuan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | COSTA: A Multi-Center TOF-MRA Dataset and a Style Self-Consistency Network for Cerebrovascular SegmentationabstractTime-of-flight magnetic resonance angiography (TOF-MRA) is the least invasive and ionizing radiation-free approach for cerebrovascular imaging, but variations in imaging artifacts across different clinical centers and imaging vendors result in inter-site and inter-vendor heterogeneity, making its accurate and robust cerebrovascular segmentation challenging. Moreover, the limited availability and quality of annotated data pose further challenges for segmentation methods to generalize well to unseen datasets. In this paper, we construct the largest and most diverse TOF-MRA dataset (COSTA) from 8 individual imaging centers, with all the volumes manually annotated. Then we propose a novel network for cerebrovascular segmentation, namely CESAR, with the ability to tackle feature granularity and image style heterogeneity issues. Specifically, a coarse-to-fine architecture is implemented to refine cerebrovascular segmentation in an iterative manner. An automatic feature selection module is proposed to selectively fuse global long-range dependencies and local contextual information of cerebrovascular structures. A style self-consistency loss is then introduced to explicitly align diverse styles of TOF-MRA images to a standardized one. Extensive experimental results on the COSTA dataset demonstrate the effectiveness of our CESAR network against state-of-the-art methods. We have made 6 subsets of COSTA with the source code online available, in order to promote relevant research in the community. Lei Mou, Jinghui Lin, Yifan Zhao 0001, Yonghuai Liu, Shaodong Ma, Jiong Zhang 0004, Wenhao Lv, Tao Zhou 0002, Jiang Liu 0001, Alejandro F. Frangi, Yitian Zhao |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Polar-Net: A Clinical-Friendly Model for Alzheimer's Disease Detection in OCTA Images
Shouyue Liu, Jinkui Hao, Yanwu Xu 0001, Huazhu Fu, Jiang Liu 0001, Yalin Zheng, Yonghuai Liu, Jiong Zhang 0004, Yitian Zhao |
MICCAI (7) | 8 |
| 2023 | 3D Visual Saliency: An Independent Perceptual Measure or a Derivative of 2D Image Saliency?abstractWhile 3D visual saliency aims to predict regional importance of 3D surfaces in agreement with human visual perception and has been well researched in computer vision and graphics, latest work with eye-tracking experiments shows that state-of-the-art 3D visual saliency methods remain poor at predicting human fixations. Cues emerging prominently from these experiments suggest that 3D visual saliency might associate with 2D image saliency. This paper proposes a framework that combines a Generative Adversarial Network and a Conditional Random Field for learning visual saliency of both a single 3D object and a scene composed of multiple 3D objects with image saliency ground truth to 1) investigate whether 3D visual saliency is an independent perceptual measure or just a derivative of image saliency and 2) provide a weakly supervised method for more accurately predicting 3D visual saliency. Through extensive experiments, we not only demonstrate that our method significantly outperforms the state-of-the-art approaches, but also manage to answer the interesting and worthy question proposed within the title of this paper. Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Regional Attention Network (RAN) for Head Pose and Fine-Grained Gesture RecognitionabstractAffect is often expressed via non-verbal body language such as actions/gestures, which are vital indicators for human behaviors. Recent studies on recognition of fine-grained actions/gestures in monocular images have mainly focused on modeling spatial configuration of body parts representing body pose, human-objects interactions and variations in local appearance. The results show that this is a brittle approach since it relies on accurate body parts/objects detection. In this work, we argue that there exist local discriminative semantic regions, whose “informativeness” can be evaluated by the attention mechanism for inferring fine-grained gestures/actions. To this end, we propose a novel end-to-endregional attention network (RAN), which is a fully convolutional neural network (CNN) to combine multiple contextual regions through attention mechanism, focusing on parts of the images that are most relevant to a given task. Our regions consist of one or more consecutive cells and are adapted from the strategies used in computing HOG (Histogram of Oriented Gradient) descriptor. The model is extensively evaluated on ten datasets belonging to 3 different scenarios: 1) head pose recognition, 2) drivers state recognition, and 3) human action and facial expression recognition. The proposed approach outperforms the state-of-the-art by a considerable margin in different metrics. Ardhendu Behera, Zachary Wharton, Yonghuai Liu, Morteza Ghahremani, Swagat Kumar, Nik Bessis |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Topology-Aware Learning for Semi-supervised Cross-domain Retinal Artery/Vein Classification
Jianyang Xie, Yonghuai Liu, Huaying Hao, Lijun Guo, Jiong Zhang 0004, Yitian Zhao |
CGI | 3 |
| 2022 | Unsupervised Multi-View CNN for Salient View Selection and 3D Interest Point Detection
Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu |
Int. J. Comput. Vis. | 4 |
| 2022 | Uncertainty-guided graph attention network for parapneumonic effusion diagnosis
Jinkui Hao, Jiang Liu 0001, Ella Grishikashvili Pereira, Ri Liu, Jiong Zhang 0004, Yangfan Zhang, Jianjun Zheng, Jingfeng Zhang, Yonghuai Liu, Yitian Zhao |
Medical Image Anal. | 11 |
| 2022 | SR-GNN: Spatial Relation-Aware Graph Neural Network for Fine-Grained Image CategorizationabstractOver the past few years, a significant progress has been made in deep convolutional neural networks (CNNs)-based image recognition. This is mainly due to the strong ability of such networks in mining discriminative object pose and parts information from texture and shape. This is often inappropriate for fine-grained visual classification (FGVC) since it exhibits high intra-class and low inter-class variances due to occlusions, deformation, illuminations, etc. Thus, an expressive feature representation describing global structural information is a key to characterize an object/ scene. To this end, we propose a method that effectively captures subtle changes by aggregating context-aware features from most relevant image-regions and their importance in discriminating fine-grained categories avoiding the bounding-box and/or distinguishable part annotations. Our approach is inspired by the recent advancement in self-attention and graph neural networks (GNNs) approaches to include a simple yet effective relation-aware feature transformation and its refinement using a context-aware attention mechanism to boost the discriminability of the transformed feature in an end-to-end learning process. Our model is evaluated on eight benchmark datasets consisting of fine-grained objects and human-object interactions. It outperforms the state-of-the-art approaches by a significant margin in recognition accuracy. Asish Bera, Zachary Wharton, Yonghuai Liu, Nik Bessis, Ardhendu Behera |
IEEE Trans. Image Process. | 3 |
| 2022 | Retinal Structure Detection in OCTA Image via Voting-Based Multitask LearningabstractAutomated detection of retinal structures, such as retinal vessels (RV), the foveal avascular zone (FAZ), and retinal vascular junctions (RVJ), are of great importance for understanding diseases of the eye and clinical decision-making. In this paper, we propose a novel Voting-based Adaptive Feature Fusion multi-task network (VAFF-Net) for joint segmentation, detection, and classification of RV, FAZ, and RVJ in optical coherence tomography angiography (OCTA). A task-specific voting gate module is proposed to adaptively extract and fuse different features for specific tasks at two levels: features at different spatial positions from a single encoder, and features from multiple encoders. In particular, since the complexity of the microvasculature in OCTA images makes simultaneous precise localization and classification of retinal vascular junctions into bifurcation/crossing a challenging task, we specifically design a task head by combining the heatmap regression and grid classification. We take advantage of three different en face angiograms from various retinal layers, rather than following existing methods that use only a single en face. We carry out extensive experiments on three OCTA datasets acquired using different imaging devices, and the results demonstrate that the proposed method performs on the whole better than either the state-of-the-art single-purpose methods or existing multi-task learning solutions. We also demonstrate that our multi-task learning method generalizes across other imaging modalities, such as color fundus photography, and may potentially be used as a general multi-task learning tool. We also construct three datasets for multiple structure detection, and part of these datasets with the source code and evaluation benchmark have been released for public access. Jinkui Hao, Ting Shen, Xueli Zhu 0002, Yonghuai Liu, Ardhendu Behera, Dan Zhang 0026, Bang Chen, Jiang Liu 0001, Jiong Zhang 0004, Yitian Zhao |
IEEE Trans. Medical Imaging | 4 |
| 2022 | DeepGrading: Deep Learning Grading of Corneal Nerve TortuosityabstractAccurate estimation and quantification of the corneal nerve fiber tortuosity in corneal confocal microscopy (CCM) is of great importance for disease understanding and clinical decision-making. However, the grading of corneal nerve tortuosity remains a great challenge due to the lack of agreements on the definition and quantification of tortuosity. In this paper, we propose a fully automated deep learning method that performs image-level tortuosity grading of corneal nerves, which is based on CCM images and segmented corneal nerves to further improve the grading accuracy with interpretability principles. The proposed method consists of two stages: 1) A pre-trained feature extraction backbone over ImageNet is fine-tuned with a proposed novel bilinear attention (BA) module for the prediction of the regions of interest (ROIs) and coarse grading of the image. The BA module enhances the ability of the network to model long-range dependencies and global contexts of nerve fibers by capturing second-order statistics of high-level features. 2) An auxiliary tortuosity grading network (AuxNet) is proposed to obtain an auxiliary grading over the identified ROIs, enabling the coarse and additional gradings to be finally fused together for more accurate final results. The experimental results show that our method surpasses existing methods in tortuosity grading, and achieves an overall accuracy of 85.64% in four-level classification. We also validate it over a clinical dataset, and the statistical analysis demonstrates a significant difference of tortuosity levels between healthy control and diabetes group. We have released a dataset with 1500 CCM images and their manual annotations of four tortuosity levels for public access. The code is available at: https://github.com/iMED-Lab/TortuosityGrading. Lei Mou, Yonghuai Liu, Yalin Zheng, Peter Matthew, Pan Su 0001, Jiang Liu 0001, Jiong Zhang 0004, Yitian Zhao |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Shape Analysis Approach Towards Assessment of Cleft Lip Repair Outcome
Paul M. Bakaki, Bruce Richard, Ella Grishikashvili Pereira, Aristides Tagalakis, Andy Ness, Yonghuai Liu |
CAIP (1) | 6 |
| 2021 | Mesh Saliency: An Independent Perceptual Measure or a Derivative of Image Saliency?abstractWhile mesh saliency aims to predict regional importance of 3D surfaces in agreement with human visual perception and is well researched in computer vision and graphics, latest work with eye-tracking experiments shows that state-of-the-art mesh saliency methods remain poor at predicting human fixations. Cues emerging prominently from these experiments suggest that mesh saliency might associate with the saliency of 2D natural images. This paper proposes a novel deep neural network for learning mesh saliency using image saliency ground truth to 1) investigate whether mesh saliency is an independent perceptual measure or just a derivative of image saliency and 2) provide a weakly supervised method for more accurately predicting mesh saliency. Through extensive experiments, we not only demonstrate that our method outperforms the current state-of-the-art mesh saliency method by 116% and 21% in terms of linear correlation coefficient and AUC respectively, but also reveal that mesh saliency is intrinsically related with both image saliency and object categorical information. Codes are available at https://github.com/rsong/MIMO-GAN. Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu, Paul L. Rosin |
CVPR | 4 |
| 2021 | Cross-Domain Depth Estimation Network for 3D Vessel Reconstruction in OCT Angiography
Yonghuai Liu, Jiong Zhang 0004, Jianyang Xie, Yalin Zheng, Jiang Liu 0001, Yitian Zhao |
MICCAI (8) | 2 |
| 2021 | Coarse Temporal Attention Network (CTA-Net) for Driver's Activity RecognitionabstractThere is significant progress in recognizing traditional human activities from videos focusing on highly distinctive actions involving discriminative body movements, body-object and/or human-human interactions. Driver's activities are different since they are executed by the same subject with similar body parts movements, resulting in subtle changes. To address this, we propose a novel framework by exploiting the spatiotemporal attention to model the subtle changes. Our model is named Coarse Temporal Attention Network (CTA-Net), in which coarse temporal branches are introduced in a trainable glimpse network. The goal is to allow the glimpse to capture high-level temporal relationships, such as `during', `before' and `after' by focusing on a specific part of a video. These branches also respect the topology of the temporal dynamics in the video, ensuring that different branches learn meaningful spatial and temporal changes. The model then uses an innovative attention mechanism to generate high-level action specific contextual information for activity recognition by exploring the hidden states of an LSTM. The attention mechanism helps in learning to decide the importance of each hidden state for the recognition task by weighing them when constructing the representation of the video. Our approach is evaluated on four publicly accessible datasets and significantly outperforms the state-of-the-art by a considerable margin with only RGB video as input. Zachary Wharton, Ardhendu Behera, Yonghuai Liu, Nik Bessis |
WACV | 3 |
| 2021 | SHREC 2021: Retrieval and classification of protein surfaces equipped with physical and chemical properties
Andrea Raffo, Ulderico Fugacci, Silvia Biasotti, Walter Rocchia, Yonghuai Liu, Ekpo Otu, Reyer Zwiggelaar, David Hunter, Evangelia I. Zacharaki, Eleftheria Psatha, Dimitrios Laskos, Gerasimos Arvanitis, Konstantinos Moustakas, Tunde Aderinwale, Charles Christoffer, Woong-Hee Shin, Daisuke Kihara, Andrea Giachetti 0001, Huu-Nghia Nguyen, Tuan-Duy Nguyen, Vinh-Thuyen Nguyen-Truong, Danh Le-Thanh, Hai-Dang Nguyen, Minh-Triet Tran |
Comput. Graph. | 5 |
| 2021 | CS2-Net: Deep learning segmentation of curvilinear structures in medical imaging
Lei Mou, Yitian Zhao, Huazhu Fu, Yonghuai Liu, Jun Cheng 0003, Yalin Zheng, Pan Su 0001, Jianlong Yang, Li Chen 0011, Alejandro F. Frangi, Masahiro Akiba, Jiang Liu 0001 |
Medical Image Anal. | 4 |
| 2021 | Interwoven texture-based description of interest points in images
Morteza Ghahremani, Yitian Zhao, Bernard Tiddeman, Yonghuai Liu |
Pattern Recognit. | 4 |
| 2021 | Attend and Guide (AG-Net): A Keypoints-Driven Attention-Based Deep Network for Image RecognitionabstractThis article presents a novel keypoints-based attention mechanism for visual recognition in still images. Deep Convolutional Neural Networks (CNNs) for recognizing images with distinctive classes have shown great success, but their performance in discriminating fine-grained changes is not at the same level. We address this by proposing an end-to-end CNN model, which learns meaningful features linking fine-grained changes using our novel attention mechanism. It captures the spatial structures in images by identifying semantic regions (SRs) and their spatial distributions, and is proved to be the key to modeling subtle changes in images. We automatically identify these SRs by grouping the detected keypoints in a given image. The "usefulness" of these SRs for image recognition is measured using our innovative attentional mechanism focusing on parts of the image that are most relevant to a given task. This framework applies to traditional and fine-grained image recognition tasks and does not require manually annotated regions (e.g. bounding-box of body parts, objects, etc.) for learning and prediction. Moreover, the proposed keypoints-driven attention mechanism can be easily integrated into the existing CNN models. The framework is evaluated on six diverse benchmark datasets. The model outperforms the state-of-the-art approaches by a considerable margin using Distracted Driver V1 (Acc: 3.39%), Distracted Driver V2 (Acc: 6.58%), Stanford-40 Actions (mAP: 2.15%), People Playing Musical Instruments (mAP: 16.05%), Food-101 (Acc: 6.30%) and Caltech-256 (Acc: 2.59%) datasets. Asish Bera, Zachary Wharton, Yonghuai Liu, Nik Bessis, Ardhendu Behera |
IEEE Trans. Image Process. | 3 |
| 2021 | FFD: Fast Feature DetectorabstractScale-invariance, good localization and robustness to noise and distortions are the main properties that a local feature detector should possess. Most existing local feature detectors find excessive unstable feature points that increase the number of keypoints to be matched and the computational time of the matching step. In this paper, we show that robust and accurate keypoints exist in the specific scale-space domain. To this end, we first formulate the superimposition problem into a mathematical model and then derive a closed-form solution for multiscale analysis. The model is formulated via difference-of-Gaussian (DoG) kernels in the continuous scale-space domain, and it is proved that setting the scale-space pyramid's blurring ratio and smoothness to 2 and 0.627, respectively, facilitates the detection of reliable keypoints. For the applicability of the proposed model to discrete images, we discretize it using the undecimated wavelet transform and the cubic spline function. Theoretically, the complexity of our method is less than 5% of that of the popular baseline Scale Invariant Feature Transform (SIFT). Extensive experimental results show the superiority of the proposed feature detector over the existing representative hand-crafted and learning-based techniques in accuracy and computational time. The code and supplementary materials can be found at https://github.com/mogvision/FFD. Morteza Ghahremani, Yonghuai Liu, Bernard Tiddeman |
IEEE Trans. Image Process. | 2 |
| 2021 | Structure and Illumination Constrained GAN for Medical Image EnhancementabstractThe development of medical imaging techniques has greatly supported clinical decision making. However, poor imaging quality, such as non-uniform illumination or imbalanced intensity, brings challenges for automated screening, analysis and diagnosis of diseases. Previously, bi-directional GANs (e.g., CycleGAN), have been proposed to improve the quality of input images without the requirement of paired images. However, these methods focus on global appearance, without imposing constraints on structure or illumination, which are essential features for medical image interpretation. In this paper, we propose a novel and versatile bi-directional GAN, named Structure and illumination constrained GAN (StillGAN), for medical image quality enhancement. Our StillGAN treats low- and high-quality images as two distinct domains, and introduces local structure and illumination constraints for learning both overall characteristics and local details. Extensive experiments on three medical image datasets (e.g., corneal confocal microscopy, retinal color fundus and endoscopy images) demonstrate that our method performs better than both conventional methods and other deep learning-based methods. In addition, we have investigated the impact of the proposed method on different medical image analysis and clinical tasks such as nerve segmentation, tortuosity grading, fovea localization and disease classification. Yuhui Ma, Jiang Liu 0001, Yonghuai Liu, Huazhu Fu, Jun Cheng 0003, Yufei Wu 0013, Jiong Zhang 0004, Yitian Zhao |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Mesh Saliency via Weakly Supervised Classification-for-Saliency CNNabstractRecently, effort has been made to apply deep learning to the detection of mesh saliency. However, one major barrier is to collect a large amount of vertex-level annotation as saliency ground truth for training the neural networks. Quite a few pilot studies showed that this task is difficult. In this work, we solve this problem by developing a novel network trained in a weakly supervised manner. The training is end-to-end and does not require any saliency ground truth but only the class membership of meshes. Our Classification-for-Saliency CNN (CfS-CNN) employs a multi-view setup and contains a newly designed two-channel structure which integrates view-based features of both classification and saliency. It essentially transfers knowledge from 3D object classification to mesh saliency. Our approach significantly outperforms the existing state-of-the-art methods according to extensive experimental results. Also, the CfS-CNN can be directly used for scene saliency. We showcase two novel applications based on scene saliency to demonstrate its utility. Ran Song 0001, Yonghuai Liu, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Orderly Disorder in Point Cloud Domain
Morteza Ghahremani, Bernard Tiddeman, Yonghuai Liu, Ardhendu Behera |
ECCV (28) | 3 |
| 2020 | Unsupervised Multi-view CNN for Salient View Selection of 3D Objects and Scenes
Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu |
ECCV (19) | 4 |
| 2020 | Cycle Structure and Illumination Constrained GAN for Medical Image Enhancement
Yuhui Ma, Yonghuai Liu, Jun Cheng 0003, Yalin Zheng, Morteza Ghahremani, Honghan Chen, Jiang Liu 0001, Yitian Zhao |
MICCAI (2) | 2 |
| 2020 | Classification of Retinal Vessels into Artery-Vein in OCT Angiography Guided by Fundus Images
Jianyang Xie, Yonghuai Liu, Yalin Zheng, Pan Su 0001, Jian Yang 0009, Jiang Liu 0001, Yitian Zhao |
MICCAI (6) | 2 |
| 2020 | Temporal convolutional neural (TCN) network for an effective weather forecasting using time-series data from the local weather stationabstractAbstract Non-predictive or inaccurate weather forecasting can severely impact the community of users such as farmers. Numerical weather prediction models run in major weather forecasting centers with several supercomputers to solve simultaneous complex nonlinear mathematical equations. Such models provide the medium-range weather forecasts, i.e., every 6 h up to 18 h with grid length of 10–20 km. However, farmers often depend on more detailed short-to medium-range forecasts with higher-resolution regional forecasting models. Therefore, this research aims to address this by developing and evaluating a lightweight and novel weather forecasting system, which consists of one or more local weather stations and state-of-the-art machine learning techniques for weather forecasting using time-series data from these weather stations. To this end, the system explores the state-of-the-art temporal convolutional network (TCN) and long short-term memory (LSTM) networks. Our experimental results show that the proposed model using TCN produces better forecasting compared to the LSTM and other classic machine learning approaches. The proposed model can be used as an efficient localized weather forecasting tool for the community of users, and it could be run on a stand-alone personal computer. Pradeep Hewage, Ardhendu Behera, Marcello Trovati, Ella Grishikashvili Pereira, Morteza Ghahremani, Francesco Palmieri 0002, Yonghuai Liu |
Soft Comput. | 7 |
| 2020 | Histogram of Fuzzy Local Spatio-Temporal Descriptors for Video Action RecognitionabstractFeature extraction plays a vital role in visual action recognition. Many existing gradient-based feature extractors, including histogram of oriented gradients, histogram of optical flow, motion boundary histograms, and histogram of motion gradients, build histograms for representing different actions over the spatio-temporal domain in a video. However, these methods require to set the number of bins for information aggregation in advance. Varying numbers of bins usually lead to inherent uncertainty within the process of pixel voting with regard to the bins in the histogram. This article proposes a novel method to handle such uncertainty by fuzzifying these feature extractors. The proposed approach has two advantages: it better represents the ambiguous boundaries between the bins and, thus, the fuzziness of the spatio-temporal visual information entailed in videos; and the contribution of each pixel is flexibly controlled by a fuzziness parameter for various scenarios. The proposed family of fuzzy descriptors and a combination of them are evaluated on two publicly available datasets, demonstrating that the proposed approach outperforms the original counterparts and other state-of-the-art methods. Zheming Zuo, Longzhi Yang, Yonghuai Liu, Fei Chao 0001, Ran Song 0001, Yanpeng Qu |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Retinal Vascular Network Topology Reconstruction and Artery/Vein Classification via Dominant Set ClusteringabstractThe estimation of vascular network topology in complex networks is important in understanding the relationship between vascular changes and a wide spectrum of diseases. Automatic classification of the retinal vascular trees into arteries and veins is of direct assistance to the ophthalmologist in terms of diagnosis and treatment of eye disease. However, it is challenging due to their projective ambiguity and subtle changes in appearance, contrast, and geometry in the imaging process. In this paper, we propose a novel method that is capable of making the artery/vein (A/V) distinction in retinal color fundus images based on vascular network topological properties. To this end, we adapt the concept of dominant set clustering and formalize the retinal blood vessel topology estimation and the A/V classification as a pairwise clustering problem. The graph is constructed through image segmentation, skeletonization, and identification of significant nodes. The edge weight is defined as the inverse Euclidean distance between its two end points in the feature space of intensity, orientation, curvature, diameter, and entropy. The reconstructed vascular network is classified into arteries and veins based on their intensity and morphology. The proposed approach has been applied to five public databases, namely INSPIRE, IOSTAR, VICAVR, DRIVE, and WIDE, and achieved high accuracies of 95.1%, 94.2%, 93.8%, 91.1%, and 91.0%, respectively. Furthermore, we have made manual annotations of the blood vessel topologies for INSPIRE, IOSTAR, VICAVR, and DRIVE datasets, and these annotations are released for public access so as to facilitate researchers in the community. Yitian Zhao, Yonghuai Liu, Jianyang Xie, Huaizhong Zhang, Yalin Zheng, Yifan Zhao 0001, Yangchun Zhao, Pan Su 0001, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Automated Tortuosity Analysis of Nerve Fibers in Corneal Confocal MicroscopyabstractPrecise characterization and analysis of corneal nerve fiber tortuosity are of great importance in facilitating examination and diagnosis of many eye-related diseases. In this paper we propose a fully automated method for image-level tortuosity estimation, comprising image enhancement, exponential curvature estimation, and tortuosity level classification. The image enhancement component is based on an extended Retinex model, which not only corrects imbalanced illumination and improves image contrast in an image, but also models noise explicitly to aid removal of imaging noise. Afterwards, we take advantage of exponential curvature estimation in the 3D space of positions and orientations to directly measure curvature based on the enhanced images, rather than relying on the explicit segmentation and skeletonization steps in a conventional pipeline usually with accumulated pre-processing errors. The proposed method has been applied over two corneal nerve microscopy datasets for the estimation of a tortuosity level for each image. The experimental results show that it performs better than several selected state-of-the-art methods. Furthermore, we have performed manual gradings at tortuosity level of four hundred and three corneal nerve microscopic images, and this dataset has been released for public access to facilitate other researchers in the community in carrying out further research on the same and related topics. Yitian Zhao, Jiong Zhang 0004, Ella Grishikashvili Pereira, Yalin Zheng, Pan Su 0001, Jianyang Xie, Yifan Zhao 0001, Yonggang Shi, Jiang Liu 0001, Yonghuai Liu |
IEEE Trans. Medical Imaging | 11 |
| 2020 | Corrections to "Automated Tortuosity Analysis of Nerve Fibers in Corneal Confocal Microscopy"abstractIn the above article[1], there were two errors in the printed article that the authors want to correct. Yitian Zhao, Jiong Zhang 0004, Ella Grishikashvili Pereira, Yalin Zheng, Pan Su 0001, Jianyang Xie, Yifan Zhao 0001, Yonggang Shi, Jiang Liu 0001, Yonghuai Liu |
IEEE Trans. Medical Imaging | 11 |
| 2020 | Distinction of 3D Objects and Scenes via Classification Network and Markov Random FieldabstractAn importance measure of 3D objects inspired by human perception has a range of applications since people want computers to behave like humans in many tasks. This paper revisits a well-defined measure, distinction of 3D surface mesh, which indicates how important a region of a mesh is with respect to classification. We develop a method to compute it based on a classification network and a Markov Random Field (MRF). The classification network learns view-based distinction by handling multiple views of a 3D object. Using a classification network has an advantage of avoiding the training data problem which has become a major obstacle of applying deep learning to 3D object understanding tasks. The MRF estimates the parameters of a linear model for combining the view-based distinction maps. The experiments using several publicly accessible datasets show that the distinctive regions detected by our method are not just significantly different from those detected by methods based on handcrafted features, but more consistent with human perception. We also compare it with other perceptual measures and quantitatively evaluate its performance in the context of two applications. Furthermore, due to the view-based nature of our method, we are able to easily extend mesh distinction to 3D scenes containing multiple objects. Ran Song 0001, Yonghuai Liu, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Topology Reconstruction of Tree-Like Structure in Images via Structural Similarity Measure and Dominant Set ClusteringabstractThe reconstruction and analysis of tree-like topological structures in the biomedical images is crucial for biologists and surgeons to understand biomedical conditions and plan surgical procedures. The underlying tree-structure topology reveals how different curvilinear components are anatomically connected to each other. Existing automated topology reconstruction methods have great difficulty in identifying the connectivity when two or more curvilinear components cross or bifurcate, due to their projection ambiguity, imaging noise and low contrast. In this paper, we propose a novel curvilinear structural similarity measure to guide a dominant-set clustering approach to address this indispensable issue. The novel similarity measure takes into account both intensity and geometric properties in representing the curvilinear structure locally and globally, and group curvilinear objects at crossover points into different connected branches by dominant-set clustering. The proposed method is applicable to different imaging modalities, and quantitative and qualitative results on retinal vessel, plant root, and neuronal network datasets show that our methodology is capable of advancing the current state-of-the-art techniques. Jianyang Xie, Yitian Zhao, Yonghuai Liu, Pan Su 0001, Yifan Zhao 0001, Jun Cheng 0003, Yalin Zheng, Jiang Liu 0001 |
CVPR | 3 |
| 2019 | Human consistency evaluation of static video summaries
Sivapriyaa Kannappan, Yonghuai Liu, Bernard Tiddeman |
Multim. Tools Appl. | 2 |
| 2019 | DFP-ALC: Automatic video summarization using Distinct Frame Patch index and Appearance based Linear Clustering
Sivapriyaa Kannappan, Yonghuai Liu, Bernard Tiddeman |
Pattern Recognit. Lett. | 2 |
| 2018 | Logistic Regression of Point Matches for Accurate Transformation EstimationabstractFeature extraction and matching (FEM) has been widely used for the registration of partially overlapping 3D shapes. Due to various factors such as imaging noise, simple geometry, or clutter, it usually introduces false positive ones. To reliably estimate the underlying transformation that brings one partial shape into the best possible alignment with another, it is critical to estimate the extent to which the established point matches are correct. To this end, we propose to use the logit function for the regression of the errors of these point matches. The novel method includes three steps: (i) normalization of the errors of the point matches, (ii) logistic regression of the point matches for the estimation of their reliabilities/weights, and (iii) estimation of the underlying transformation in the weighted least squares sense. These steps are repeated until either the maximum number of iterations has been reached or the weighted average of the errors of the point matches has been below the scanning resolution. A comparative study using real data captured by different range sensors shows that the proposed method outperforms two state-of-the-art ones for more accurate estimation of the underlying transformation. Yonghuai Liu, Yitian Zhao, Yanquan Zhou, Jiwan Han, Wanneng Yang, Yiguang Liu |
3DV | 1 |
| 2018 | Retinal Artery and Vein Classification via Dominant Sets Clustering-Based Vascular Topology Estimation
Yitian Zhao, Jianyang Xie, Pan Su 0001, Yalin Zheng, Yonghuai Liu, Jun Cheng 0003, Jiang Liu 0001 |
MICCAI (2) | 5 |
| 2018 | Uniqueness-Driven Saliency Analysis for Automated Lesion Detection with Applications to Retinal Diseases
Yitian Zhao, Yalin Zheng, Yifan Zhao 0001, Yonghuai Liu, Peng Liu 0049, Jiang Liu 0001 |
MICCAI (2) | 4 |
| 2018 | Approximate top-K answering under uncertain schema mappings
Longzhuang Li, Feng Tian 0002, Yonghuai Liu, Shanxian Mao |
Data Knowl. Eng. | 3 |
| 2018 | Incorporating knowledge into neural network for text representation
Baogang Wei, Yonghuai Liu, Jifang Yu |
Expert Syst. Appl. | 3 |
| 2018 | Bilinear joint learning of word and entity embeddings for Entity Linking
Baogang Wei, Yonghuai Liu, Jifang Yu |
Neurocomputing | 3 |
| 2018 | Automatic 2-D/3-D Vessel Enhancement in Multiple Modality Images Using a Weighted Symmetry FilterabstractAutomated detection of vascular structures is of great importance in understanding the mechanism, diagnosis, and treatment of many vascular pathologies. However, automatic vascular detection continues to be an open issue because of difficulties posed by multiple factors, such as poor contrast, inhomogeneous backgrounds, anatomical variations, and the presence of noise during image acquisition. In this paper, we propose a novel 2-D/3-D symmetry filter to tackle these challenging issues for enhancing vessels from different imaging modalities. The proposed filter not only considers local phase features by using a quadrature filter to distinguish between lines and edges, but also uses the weighted geometric mean of the blurred and shifted responses of the quadrature filter, which allows more tolerance of vessels with irregular appearance. As a result, this filter shows a strong response to the vascular features under typical imaging conditions. Results based on eight publicly available datasets (six 2-D data sets, one 3-D data set, and one 3-D synthetic data set) demonstrate its superior performance to other state-of-the-art methods. Yitian Zhao, Yalin Zheng, Yonghuai Liu, Yifan Zhao 0001, Lingling Luo, Tong Na, Yongtian Wang, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2018 | Local-to-global mesh saliency
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Karina Rodriguez-Echavarria |
Vis. Comput. | 2 |
| 2017 | Saliency driven vasculature segmentation with infinite perimeter active contour model
Yitian Zhao, Jingliang Zhao, Jian Yang 0009, Yonghuai Liu, Yifan Zhao 0001, Yalin Zheng, Likun Xia, Yongtian Wang |
Neurocomputing | 4 |
| 2017 | An easy-to-use evaluation framework for benchmarking entity recognition and disambiguation systemsabstractEntity recognition and disambiguation (ERD) is a crucial technique for knowledge base population and information extraction. In recent years, numerous papers have been published on this subject, and various ERD systems have been developed. However, there are still some confusions over the ERD field for a fair and complete comparison of these systems. Therefore, it is of emerging interest to develop a unified evaluation framework. In this paper, we present an easy-to-use evaluation framework (EUEF), which aims at facilitating the evaluation process and giving a fair comparison of ERD systems. EUEF is well designed and released to the public as an open source, and thus could be easily extended with novel ERD systems, datasets, and evaluation metrics. It is easy to discover the advantages and disadvantages of a specific ERD system and its components based on EUEF. We perform a comparison of several popular and publicly available ERD systems by using EUEF, and draw some interesting conclusions after a detailed analysis. Baogang Wei, Yonghuai Liu |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2017 | Intensity and Compactness Enabled Saliency Estimation for Leakage Detection in Diabetic and Malarial RetinopathyabstractLeakage in retinal angiography currently is a key feature for confirming the activities of lesions in the management of a wide range of retinal diseases, such as diabetic maculopathy and paediatric malarial retinopathy. This paper proposes a new saliency-based method for the detection of leakage in fluorescein angiography. A superpixel approach is firstly employed to divide the image into meaningful patches (or superpixels) at different levels. Two saliency cues, intensity and compactness, are then proposed for the estimation of the saliency map of each individual superpixel at each level. The saliency maps at different levels over the same cues are fused using an averaging operator. The two saliency maps over different cues are fused using a pixel-wise multiplication operator. Leaking regions are finally detected by thresholding the saliency map followed by a graph-cut segmentation. The proposed method has been validated using the only two publicly available datasets: one for malarial retinopathy and the other for diabetic retinopathy. The experimental results show that it outperforms one of the latest competitors and performs as well as a human expert for leakage detection and outperforms several state-of-the-art methods for saliency detection. Yitian Zhao, Yalin Zheng, Yonghuai Liu, Jian Yang 0009, Yifan Zhao 0001, Duanduan Chen, Yongtian Wang |
IEEE Trans. Medical Imaging | 3 |
| 2017 | Convex Hull Aided Registration Method (CHARM)abstractNon-rigid registration finds many applications such as photogrammetry, motion tracking, model retrieval, and object recognition. In this paper we propose a novel convex hull aided registration method (CHARM) to match two point sets subject to a non-rigid transformation. First, two convex hulls are extracted from the source and target respectively. Then, all points of the point sets are projected onto the reference plane through each triangular facet of the hulls. From these projections, invariant features are extracted and matched optimally. The matched feature point pairs are mapped back onto the triangular facets of the convex hulls to remove outliers that are outside any relevant triangular facet. The rigid transformation from the source to the target is robustly estimated by the random sample consensus (RANSAC) scheme through minimizing the distance between the matched feature point pairs. Finally, these feature points are utilized as the control points to achieve non-rigid deformation in the form of thin-plate spline of the entire source point set towards the target one. The experimental results based on both synthetic and real data show that the proposed algorithm outperforms several state-of-the-art ones with respect to sampling, rotational angle, and data noise. In addition, the proposed CHARM algorithm also shows higher computational efficiency compared to these methods. Jingfan Fan, Jian Yang 0009, Yitian Zhao, Danni Ai, Yonghuai Liu, Ge Wang 0001, Yongtian Wang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2016 | A pertinent evaluation of automatic video summaryabstractVideo summarization is useful to find a concise representation of the original video, nevertheless its evaluation is somewhat challenging. This paper proposes a simple and efficient method for precisely evaluating the video summaries produced by the existing techniques. This method includes two steps. The first step is to establish a set of matched frames between automatic summary (AT) and the ground truth summary (GT) through two-way search, in which the similarity between two frames are measured using correlation coefficient. The second step is to estimate the consistency among these established matches, so that the difference among these frames in the AT and GT are preserved respectively. To accomplish this, a compatibility matrix is built based on the features extracted from each of these frames. The consistency values among these matched frames are estimated as the eigenvector of this matrix corresponding to the maximum eigenvalue. Such matched frames with a small enough consistency value will be rejected, leading to more accurate performance estimation of the video summarization techniques. Experimental results based on a publicly accessible dataset shows that the proposed method is effective in finding true matches and provide more realistic measurement of the performance for various techniques. Sivapriyaa Kannappan, Yonghuai Liu, Bernard Tiddeman |
ICPR | 2 |
| 2016 | Region-based saliency estimation for 3D shape analysis and understanding
Yitian Zhao, Yonghuai Liu, Baogang Wei, Jian Yang 0009, Yifan Zhao 0001, Yongtian Wang |
Neurocomputing | 2 |
| 2016 | Accurately estimating rigid transformations in registration using a boosting-inspired mechanism
Yonghuai Liu, Honghai Liu 0001, Ralph R. Martin, Luigi De Dominicis, Ran Song 0001, Yitian Zhao |
Pattern Recognit. | 1 |
| 2015 | Estimation of Branch Angle from 3D Point Cloud of PlantsabstractMeasuring geometric features in plant specimens either quantitatively or qualitatively, is crucial for plant phenotyping. However, traditional measurement methods tend to be manual and can be tedious, or employ coarse 2D imaging techniques. Emerging 3D imaging technologies show much promise in capturing architectural complexity. However, automated 3D acquisition and accurate estimation of plant morphology for the construction of quantitative plant models remain largely aspiration. In this paper, we propose an approach for segmentation and angle estimation directly from dense 3D plant point clouds. Experimental results show that the approach is efficient and reliable, and appears to be a promising 3D acquisition and measurement solution to plant phenotyping for structural analysis and for building Functional-Structural Plant Models (FSPM). Lu Lou, Yonghuai Liu, Minglan Sheng, Jiwan Han, Fiona M. K. Corke, John H. Doonan |
3DV | 2 |
| 2015 | A method for text line detection in natural images
Baogang Wei, Yonghuai Liu, Yin Zhang 0006 |
Multim. Tools Appl. | 3 |
| 2015 | Regularization Based Iterative Point Match Weighting for Accurate Rigid Transformation EstimationabstractFeature extraction and matching (FEM) for 3D shapes finds numerous applications in computer graphics and vision for object modeling, retrieval, morphing, and recognition. However, unavoidable incorrect matches lead to inaccurate estimation of the transformation relating different datasets. Inspired by AdaBoost, this paper proposes a novel iterative re-weighting method to tackle the challenging problem of evaluating point matches established by typical FEM methods. Weights are used to indicate the degree of belief that each point match is correct. Our method has three key steps: (i) estimation of the underlying transformation using weighted least squares, (ii) penalty parameter estimation via minimization of the weighted variance of the matching errors, and (iii) weight re-estimation taking into account both matching errors and information learnt in previous iterations. A comparative study, based on real shapes captured by two laser scanners, shows that the proposed method outperforms four other state-of-the-art methods in terms of evaluating point matches between overlapping shapes established by two typical FEM methods, resulting in more accurate estimates of the underlying transformation. This improved transformation can be used to better initialize the iterative closest point algorithm and its variants, making 3D shape registration more likely to succeed. Yonghuai Liu, Luigi De Dominicis, Baogang Wei, Liang Chen 0012, Ralph R. Martin |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | Using retinex for point selection in 3D shape registration
Yonghuai Liu, Ralph R. Martin, Luigi De Dominicis, Baihua Li |
Pattern Recognit. | 1 |
| 2014 | Scan integration as a labelling problem
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
Pattern Recognit. | 2 |
| 2014 | Mesh saliency via spectral processingabstractWe propose a novel method for detecting mesh saliency, a perceptually-based measure of the importance of a local region on a 3D surface mesh. Our method incorporates global considerations by making use of spectral attributes of the mesh, unlike most existing methods which are typically based on local geometric cues. We first consider the properties of the log-Laplacian spectrum of the mesh. Those frequencies which show differences from expected behaviour capture saliency in the frequency domain. Information about these frequencies is considered in the spatial domain at multiple spatial scales to localise the salient features and give the final salient areas. The effectiveness and robustness of our approach are demonstrated by comparisons to previous approaches on a range of test models. The benefits of the proposed method are further evaluated in applications such as mesh simplification, mesh segmentation, and scan integration, where we show how incorporating mesh saliency can provide improved results. Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
ACM Trans. Graph. | 2 |
| 2013 | Using Region-Based Saliency for 3D Interest Points Detection
Yitian Zhao, Yonghuai Liu |
CAIP (2) | 2 |
| 2013 | A general probability framework for improving similarity based approaches for face verificationabstractThis paper introduces a probability model for face verification, aiming at improve various similarity comparison approaches transplanted directly from face identification algorithms. Experiences demonstrate that, when embedded with a few well known subspace based similarity comparison approaches, our probability model can efficiently reduce the error rates in face verification tasks. Liang Chen 0012, David Casperson, Yonghuai Liu, Lixin Gao 0004 |
ICIP | 3 |
| 2013 | Parameterization of point-cloud freeform surfaces using adaptive sequential learning RBFnetworks
Qinggang Meng, Baihua Li, Horst Holstein, Yonghuai Liu |
Pattern Recognit. | 4 |
| 2013 | Message Passing Matching Dynamics for Overlapping Point IdentificationabstractExisting registration algorithms usually converge to a local minimum due to inaccurate evaluation of the tentative correspondences established. In this paper, we move a step further and instead estimate the extent to which a point lies in the overlapping area. To this end, we regard the registration problem as an exchange network and develop a matching dynamics to characterize the interaction inside. Then we propose a novel algorithm based on the powerful message passing scheme derived from the matching dynamics for the optimization of the overlapping point weight. The novel algorithm penalizes in the process of deterministic annealing those tentative correspondences that violate the properties of the matching dynamics. The rigid transformation that brings the two overlapping shapes into alignment is finally estimated in the weighted least squares sense. Our experiments use both synthetic and real data to show that our proposed algorithm is more likely to converge to the global minimum than four selected state of the art ones for more accurate and robust results. Yonghuai Liu |
IEEE Trans. Multim. | 1 |
| 2013 | Registration of 3D Point Clouds and Meshes: A Survey from Rigid to NonrigidabstractThree-dimensional surface registration transforms multiple three-dimensional data sets into the same coordinate system so as to align overlapping components of these sets. Recent surveys have covered different aspects of either rigid or nonrigid registration, but seldom discuss them as a whole. Our study serves two purposes: 1) To give a comprehensive survey of both types of registration, focusing on three-dimensional point clouds and meshes and 2) to provide a better understanding of registration from the perspective of data fitting. Registration is closely related to data fitting in which it comprises three core interwoven components: model selection, correspondences and constraints, and optimization. Study of these components 1) provides a basis for comparison of the novelties of different techniques, 2) reveals the similarity of rigid and nonrigid registration in terms of problem representations, and 3) shows how overfitting arises in nonrigid registration and the reasons for increasing interest in intrinsic techniques. We further summarize some practical issues of registration which include initializations and evaluations, and discuss some of our own observations, insights and foreseeable research trends. Gary K. L. Tam, Zhi-Quan Cheng, Yukun Lai, Frank C. Langbein, Yonghuai Liu, David Marshall 0001, Ralph R. Martin, Xianfang Sun, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2013 | 3D point of interest detection via spectral irregularity diffusion
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
Vis. Comput. | 2 |
| 2012 | Saliency-guided integration of multiple scansabstractWe present a novel method to integrate multiple 3D scans captured from different viewpoints. Saliency information is used to guide the integration process. The multi-scale saliency of a point is specifically designed to reflect its sensitivity to registration errors. Then scans are partitioned into salient and non-salient regions through an Markov Random Field (MRF) framework where neighbourhood consistency is incorporated to increase the robustness against potential scanning errors. We then develop different schemes to discriminatively integrate points in the two regions. For the points in salient regions which are more sensitive to registration errors, we employ the Iterative Closest Point algorithm to compensate the local registration error and find the correspondences for the integration. For the points in non-salient regions which are less sensitive to registration errors, we integrate them via an efficient and effective point-shifting scheme. A comparative study shows that the proposed method delivers improved surface integration. Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
CVPR | 2 |
| 2012 | Displacement Template with Divide-&-Conquer Algorithm for Significantly Improving Descriptor Based Face Recognition Approaches
Liang Chen 0012, Yonghuai Liu, Lixin Gao 0004, Xiaoqin Zhang 0002 |
ECCV (5) | 3 |
| 2012 | A saliency detection based method for 3D surface simplificationabstractTo accelerate the processing for the integration, registration, representation and recognition of point clouds, it is of growing necessity to simplify the surface of 3-D models. Simplification is an approach to vary the levels of visual details as appropriate, thereby improving on the overall performance of applications. This paper proposes a saliency detection based points sampling method for mesh simplification. By generating and enhancing the saliency map, the regions which are visually important can be located. For the mesh simplification, the local details are captured by the saliency, while for the overall shape, the approach voxelizes the model and samples points in terms of the entropy of the shape index of vertices in voxels. We present a number of results to show that the method significantly simplifies the surface without distortion and loss of local details. Yitian Zhao, Yonghuai Liu, Ran Song 0001 |
ICASSP | 2 |
| 2012 | Conditional random field-based mesh saliencyabstractWe propose a new method for detecting mesh saliency, a reflection of perception-based regional importance for 3D meshes. The basic idea is to incorporate the Conditional Random Field (CRF) framework with a saliency detection process. We first produce a multi-scale representation for a mesh. Then, a CRF is designed to robustly detect salient regions utilising neighbourhood consistency. By inferring the CRF via belief propagation algorithm, we actually make use of the global statistic information in the saliency detection process. Experimental results demonstrate the robustness and the effectiveness of the proposed method. Ran Song 0001, Yonghuai Liu, Yitian Zhao, Ralph R. Martin, Paul L. Rosin |
ICIP | 2 |
| 2012 | Extended non-local means filter for surface saliency detectionabstractMesh surface saliency detection is an important preprocessing step for many 3D applications. The salient region can be used to find the objects that are important on 3D surface, the benefits of saliency detection in the 3D domain include mesh simplification, registration, segmentation, compression, etc. This paper proposes a novel saliency detection method by diffusing the shape index field with non-local means filer, generating a random centre surround operator to yield saliency map and enhancing the saliency with the Retinex theory. The effectiveness of this method is demonstrated by simplification and registration. Experimental results demonstrate that the proposed approach has achieved competitive results. Yitian Zhao, Yonghuai Liu, Ran Song 0001 |
ICIP | 2 |
| 2012 | Patch based saliency detection method for 3D surface simplification
Yitian Zhao, Yonghuai Liu |
ICPR | 2 |
| 2012 | Guest Editorial: Scenes, Images and Objects
Frédéric Labrosse, Reyer Zwiggelaar, Yonghuai Liu, Bernard Tiddeman |
Int. J. Comput. Vis. | 3 |
| 2011 | Choosing the number of labels in image segmentationabstractImage segmentation is a fundamental yet challenging step during image analysis. In this paper we propose a novel method for choosing the number of labels during automatic image segmentation. It minimizes an objective function based on the number of labels, the segmentation errors, and consistency of labels between neighboring pixels. An experimental study on representative data shows encouraging results. Yonghuai Liu, John H. Draper, Alan P. Gay, Catherine N. Howarth, Ralph R. Martin |
ETFA | 1 |
| 2011 | MRF-based automatic image ordering and its application to mosaicingabstractA fast and robust auto-sorting method for image ordering based on Markov Random Fields (MRF) is proposed. We present a specific MRF model for the ordering problem and use pairwise phase correlation for the formulation. The MRF is inferred by a modified belief propagation (BP) method. Experimental results prove that the new method can reorder a disorganised collection of images without human input, prior information or restrictions, as just the first stage of a multi stage mosaicing process, but also provides information that can be used to guide a mosaicing process in order to reduce both local mismatch and global error accumulation. Ran Song 0001, Yonghuai Liu, Yitian Zhao, Ralph R. Martin, Paul L. Rosin |
ICASSP | 2 |
| 2011 | Penalizing Closest Point Sharing for Automatic Free Form Shape RegistrationabstractFor accurate registration of overlapping free form shapes, different points in one shape must select different points in another as their most sensible correspondents. To reach this ideal state, in this paper we develop a novel algorithm to penalize those points in one shape that select the same closest point in another as their tentative correspondents. The novel algorithm then models the relative weight change over time of a tentative correspondence as the difference between the negative functions of the numbers of points in one shape that actually and ideally select the same closest point in another. Such modeling results in an optimal estimation of the weights of different tentative correspondences, in the sense of deterministic annealing, that lead the camera motion parameters to be estimated in the weighted least squares sense. The proposed algorithm is initialized using the pure translational motion derived from the centroids difference of the overlapping free form shapes being registered. Experimental results show that it outperforms three selected state-of-the-art algorithms on the whole for the accurate and robust registration of real overlapping free form shapes captured using two different laser scanners under typical imaging conditions. Yonghuai Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | MRF Labeling for Multi-view Range Image Integration
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
ACCV (2) | 2 |
| 2010 | Practical issues and development of underwater 3D laser scannersabstractNowadays, 3D laser scanners are widely used in reverse engineering, industrial design, prototyping, quality control etc. Most of these scanners operate in air. Theoretically, this technology can be extended for the development of 3D laser scanners that work in water. In this paper, we describe the development and practical issues of 3D laser scanners required for applications such as underwater oil and gas inspection. Some experimental results are also shown to demonstrate the feasibility and limit of the 3D laser scanners developed. Junjie Liu 0001, Anthony Jakas, Ala Al-Obaidi, Yonghuai Liu |
ETFA | 4 |
| 2010 | Free form shape registration using the barrier method
Yonghuai Liu |
Comput. Vis. Image Underst. | 1 |
| 2010 | Automatic Range Image Registration in the Markov ChainabstractIn this paper, a novel entropy that can describe both long and short-tailed probability distributions of constituents of a thermodynamic system out of its thermodynamic limit is first derived from the Lyapunov function for a Markov chain. We then maximize this entropy for the estimation of the probabilities of possible correspondences established using the traditional closest point criterion between two overlapping range images. When we change our viewpoint to look carefully at the minimum solution to the probability estimate of the correspondences, the iterative range image registration process can also be modeled as a Markov chain in which lessons from past experience in estimating those probabilities are learned. To impose the two-way constraint, outliers are explicitly modeled due to the almost ubiquitous occurrence of occlusion, appearance, and disappearance of points in either image. The estimated probabilities of the correspondences are finally embedded into the powerful mean field annealing scheme for global optimization, leading the camera motion parameters to be estimated in the weighted least-squares sense. A comparative study using real images shows that the proposed algorithm usually outperforms the state-of-the-art ICP variants and the latest genetic algorithm for automatic overlapping range image registration. Yonghuai Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Data Integration for the Gulf of Mexico Satellite ObservationsabstractA number of satellite image sources on meteorology and environment are available over the Gulf of Mexico (GOM). Those satellite images are taken periodically in order to observe the weather changes. However, users usually face several problems when they want to search, retrieve, and combine images from multiple data sources. Without an adequate system and personnel for managing data, the magnitude of the effort needed to deal with such large and complex data sets can be a substantial barrier to the GOM research community. As a result, a mediator-based data integration system is proposed and developed to retrieve and integrate the images of satellite observational systems over the GOM. Longzhuang Li, Yonghuai Liu, Srujan Kothapally, Chuihui Jin, Anil Kumar Nalluri |
ICIW | 2 |
| 2009 | Replicator Dynamics in the Iterative Process for Accurate Range Image Matching
Yonghuai Liu |
Int. J. Comput. Vis. | 1 |
| 2009 | A clustering approach to free form surface reconstruction from multi-view range images
Yonghuai Liu, Longzhuang Li, Baogang Wei |
Image Vis. Comput. | 2 |
| 2008 | Accurate integration of multi-view range images using k-means clustering
Yonghuai Liu |
Pattern Recognit. | 2 |
| 2008 | Constraints for closest point finding
Yonghuai Liu |
Pattern Recognit. Lett. | 1 |
| 2007 | Accurate range image registration: Eliminating or modelling outliersabstractAutomatic and accurate range image registration is often a prerequisite step for range image analysis and interpretation. Due to occlusion, appearance and disappearance of points in different images, outliers inevitably occur. In this case, various techniques to eliminate and model outliers have been proposed for accurate range image registration. The objective of this paper is to experimentally investigate which of the outlier elimination and modelling is more effective for the evaluation of possible correspondences established, so that a deep insight into how advanced range image registration algorithms will be developed can be obtained. The experimental results based on both synthetic data and real images show that the outlier modelling often outperforms the outlier elimination in the sense of producing more accurate and robust range image registration results. Yonghuai Liu, Honghai Liu 0001, Longzhuang Li, Baogang Wei |
ETFA | 1 |
| 2007 | A mean field annealing approach to accurate free form shape matching
Yonghuai Liu |
Pattern Recognit. | 1 |
| 2006 | Incremental Mesh-based Integration of Registered Range Images: Robust to Registration Error and Scanning Noise
Yonghuai Liu, Longzhuang Li |
ACCV (1) | 2 |
| 2006 | Automatic registration of overlapping 3D point clouds using closest points
Yonghuai Liu |
Image Vis. Comput. | 1 |
| 2005 | 3D Free Form Surface Matching Based on Orientation Difference Length DistributionabstractIn this paper, we propose using orientation difference length distribution (ODLD) to represent a 3D view of each object with a free form surface. ODLD has a number of advantages over the shape distribution. A comparative study of ODLD and the shape distribution has shown that the former is significantly more accurate for 3D free form surface matching than the latter. Yonghuai Liu, Guoqiang Fei, Baogang Wei, Longzhuang Li |
ICASSP (2) | 1 |
| 2005 | Automatic 3d free form shape matching using the graduated assignment algorithm
Yonghuai Liu |
Pattern Recognit. | 1 |
| 2005 | Eliminating false matches for the projective registration of free-form surfaces with small translational motionsabstractIn this paper, we make a detailed study of two rigid-motion constraints. The importance of these two constraints is twofold: first, they reveal the inherent relationship between the three-dimensional-two-dimensional (3-D-2-D) point correspondences and the motion parameters of interest; second, they can be used to measure the traditional ICP criterion established point match qualities based on which different point matches can be compared and relatively good point matches can be selected for motion-parameter update in the projective registration of free-form surfaces subject to small translational motions. The experimental results based on both synthetic data and real images have shown that the rigid motion constraints are powerful in evaluating the possible 3-D-2-D point matches established by the traditional ICP criterion, thus achieving encouraging projective registration results. Yonghuai Liu |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | 3D Shape Matching using Collinearity ConstraintabstractIn this paper, a novel algorithm is proposed to carry out automatic 3D shape matching with 3D shapes represented as sets of points. After the possible matches between the 3D shapes have been determined by the tradition ICP criterion, the novel approach employs the collinearity constraint to eliminate false matches based on a statistical model. A comparative study based on real range images has shown that the proposed algorithm is accurate, robust and efficient for the automatic matching of overlapping 3D shapes. Yonghuai Liu, Longzhuang Li, Baogang Wei |
ICRA | 1 |
| 2004 | Improving ICP with easy implementation for free-form surface matching
Yonghuai Liu |
Pattern Recognit. | 1 |
| 2004 | Pseudo-linearizing collinearity constraint for accurate pose estimation from a single image
Yonghuai Liu, Horst Holstein |
Pattern Recognit. Lett. | 1 |
| 2003 | Free form surface matching using motion consistencyabstractIn this paper, we propose a novel algorithm to evaluate the possible point matches established by the traditional ICP algorithm in the process of matching two overlapping free from surfaces represented as two sets of unorganised points. While the existing methods mainly use feature matching to establish or evaluate possible correspondences between the surfaces to be matched, our novel approach applies motion consistency to evaluate the possible correspondences. In particular, while the existing methods assume that satisfying local structural constraints is a necessary condition for a pair of points to represent a real correspondence, our novel approach proves that satisfying global rigid motion constraints is a sufficient condition. A comparative study based on real images has shown that the proposed algorithm is accurate and robust for the matching of overlapping free form surfaces. Yonghuai Liu, Baogang Wei |
IROS | 1 |
| 2003 | Evaluating 3D-2D correspondences for accurate camera pose estimation from a single imageabstractDue to occlusion, appearance and disappearance of points, and noise distribution in image data, no matter what algorithms are used to register a 3D model and its projective image, they are always likely to introduce false matches. In this paper, we propose a novel algorithm to evaluate the established correspondences for more accurate camera pose estimation. The novel algorithm is based on two novel rigid motion constraints that represent necessary conditions for any pair of a 3D and a 2D point to represent a real 3D-2D correspondence. We then prove that as long as 3D-2D point matches satisfy the two novel rigid motion constraints, they are theoretically guaranteed to represent real correspondences. A large number of experiments based on both computer simulation and real images have indeed demonstrated that the proposed algorithm is accurate and robust for the evaluation of established correspondences and thus, leading to more accurate and robust camera pose estimation. Yonghuai Liu |
SMC | 1 |
| 2003 | Using Hybrid Knowledge Engineering and Image Processing in Color Virtual Restoration of Ancient MuralsabstractThis paper proposes a novel scheme to virtually restore the colors of ancient murals. Our approach integrates artificial intelligence techniques with digital image processing methods. The knowledge related to the mural colors is first categorized into four types. A hybrid frame and rule-based approach is then developed to represent knowledge and to inter colors. An algorithm that takes into account color similarity and spatial proximity is developed to segment mural images. A novel color transformation method based on color histograms is finally proposed to restore the colors of murals. A number of experiments based on real images have demonstrated the validity of the proposed scheme for color restoration. Baogang Wei, Yonghuai Liu, Yunhe Pan |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2002 | A pseudo linearization method for accurate pose estimation from a single imageabstractWe propose a novel pseudo linearization method for accurate pose estimation from given correspondences between a 3D model and its projective image. A comparative study based on both synthetic data and real images has shown that the novel algorithm is accurate and robust. Yonghuai Liu, Horst Holstein |
ICIP (2) | 1 |
| 2002 | Registering two overlapping range images using a relative registration error histogramabstractIn this paper, we propose a novel algorithm for the registration of two overlapping range images. This algorithm is based on the relative registration error histogram. A comparative study based on both synthetic data and real images has shown that the novel algorithm is accurate and robust. Marcos A. Rodrigues 0001, Yonghuai Liu |
ICIP (3) | 2 |
| 2002 | Accurate Registration of Structured Data using Two Overlapping Range ImagesabstractNew technological developments in optics and electronics are rendering laser scanning systems cheaper and more accurate. Such systems can directly capture depth information from objects simplifying the range image analysis and enlarging their applications scope. Accurate image registration algorithms represent a pivotal aspect of range image analysis which still require substantial improvements. This paper presents two novel motion constraints namely proximity and closeness constraints to improve accuracy of image registration and defines the conditions when such constraints are read at specific points of the registration. A number of experiments using real range images demonstrate that the combination of rigid motion constraints with the novel proximity and closeness constraints leads to more accurate evaluation of valid correspondences which, in turn, lead to more accurate image registration. Yonghuai Liu, Marcos A. Rodrigues 0001 |
ICRA | 1 |
| 2002 | Exploiting structural constraints for accurate image registrationabstractAccurate registration has proved a difficult task especially when using non-high quality range images. In this paper we investigate and formalise structural relationships existing in the data for registration of free form shapes. We extend an existing ICP-based geometric algorithm by incorporating structural constraints leading to more accurate data correspondences. A number of experiments based on real range images demonstrate that the combination of rigid motion constraints with the novel structural constraint yields more accurate evaluation of possible correspondences preventing the algorithm from converging to local minima leading thus, to more accurate image registration. Marcos A. Rodrigues 0001, Yonghuai Liu |
IROS | 2 |
| 2002 | Special Issue on Registration and Fusion of Range Images
Marcos A. Rodrigues 0001, Robert B. Fisher, Yonghuai Liu |
Comput. Vis. Image Underst. | 3 |
| 2001 | Geometric alignment of two overlapping range imagesabstractWe propose a novel geometric method for the alignment of two overlapping range images. The method first employs the traditional iterative closest point (ICP) criterion to establish a set of possible correspondences and then refine these correspondences using geometric constraints derived from properties of reflected correspondence vectors. In this way, the method overcomes a major limitation of the traditional ICP criterion which is the introduction of false matches in almost every iteration of the alignment. For an accurate estimation of the geometric parameters of interest, the Monte Carlo method is used in conjunction with a median filter. Finally, the quaternion method is used to estimate the motion parameters based on the refined correspondences. Experimental results based on both synthetic data and real images show that the proposed method can effectively align two overlapping range images with a small motion. Marcos A. Rodrigues 0001, Yonghuai Liu |
ICASSP | 2 |
| 2001 | A novel method to cope with appearing and disappearing points for the projective registration of free-form surfacesabstractWe present a novel algorithm to cope with appearing and disappearing points for the projective registration of free-form surfaces. The novel method combines the ICP algorithm with FOE theory for the effective elimination of false matches. Experimental results based on both synthetic data and real images show that the proposed algorithm is accurate and robust. Yonghuai Liu, Marcos A. Rodrigues 0001, Baogang Wei |
ICIP (3) | 1 |
| 2001 | Eliminating false matches in image registration through geometric histograms from reflected correspondence vectorsabstractWe propose a method to deal with false matches that occur in almost every iteration of iterative closest point (ICP) based registration algorithms. First, a set of correspondences between the two images to be registered are established using the standard ICP criterion. From this set of correspondences, the algorithm as described in Liu et al. (2000) is employed to estimate the essential point defined by geometric properties of reflected correspondence vectors. After a rigid motion, the essential point must be equidistant from reflected correspondences and thus, relative differences between motion equations can be computed and geometric histograms are then constructed at each step of the iteration. False matches are eliminated by only selecting correspondences that show a small relative difference between the two sides of the motion equation. A number of experiments based on both synthetic data and real images demonstrate that the proposed method is accurate, robust, and efficient for the registration of free-form shapes with large motions. Yonghuai Liu, Marcos A. Rodrigues 0001 |
IROS | 1 |
| 2001 | A geometric histogram method for accurate and robust motion estimation from range dataabstractMotion estimation from outlier corrupted data is a fundamental and difficult problem acknowledged in the machine vision literature. In this paper a robust motion estimation method is presented. First, the Monte Carlo resampling technique is used for an initial estimation of motion parameters, then a geometric histogram method is proposed to synthesize possible solutions to motion parameters based on geometric properties of reflected correspondence vectors. A number of experiments using both synthetic data and real images demonstrate the robustness of the proposed motion estimation method leading to more accurate registration of free form shapes. Yonghuai Liu, Marcos A. Rodrigues 0001 |
SMC | 1 |
| 2001 | Statistical image analysis for pose estimation without point correspondences
Yonghuai Liu, Marcos A. Rodrigues 0001 |
Pattern Recognit. Lett. | 1 |
| 2000 | Fuzzy reasoning based motion estimation from range imagesabstractMany methods to estimate rigid body motion parameters from range images have been put forward in the last decade. Such methods work well for range image data corrupted by Gaussian random noise without outliers. In particular, the constraint least squares (CLS) is the most accurate, robust, stable, and efficient motion estimation algorithm. However, the CLS and none of the current methods are very robust in the presence of outliers. Therefore, in this paper, we focus on the problem of estimating motion parameters from noise and outlier corrupted range image data. We propose a novel motion estimation geometric algorithm with fuzzy reasoning (GAFR). The algorithm is based on the geometric properties of correspondence vectors to synthesise motion parameter candidates and employs a robust fuzzy reasoning method based on computing deviations and selecting estimates from membership function values. The GAFR is validated through experimentation using synthetic and real range image data. Marcos A. Rodrigues 0001, Yonghuai Liu |
FUZZ-IEEE | 2 |
| 2000 | An Iterative Algorithm for the Projective Registration of Free Form SurfacesabstractIn this paper we present a novel image registration algorithm combining the iterative closest point algorithm with focus of expansion theory for 3D-2D projective registration of free-form surfaces. A pure translational camera configuration is used, which is a widely adopted constraint to structural estimation. Experimental results based on both synthetic and real images have shown that the proposed algorithm provides accurate and efficient registration and effective elimination of false matches. Yonghuai Liu, Marcos A. Rodrigues 0001 |
ICIP | 1 |
| 2000 | Distance Constraint Based Iterative Structure and Pose Estimation from a Single ImageabstractThis paper presents a novel method for structural and pose estimation from 3D-2D correspondence. Structural data are obtained by first estimating the depths of some reference points using the Newton-Raphson method with a distance constraint followed by depth estimation of the remaining points in closed form solution. Pose estimation is obtained through the constraint least squares method which has been proven accurate and robust for noisy image data. Experimental results based on both synthetic data and real images have shown that the proposed algorithm is accurate and robust in the presence of noise. Marcos A. Rodrigues 0001, Yonghuai Liu |
ICIP | 2 |
| 2000 | Using Geometric Properties of Correspondence Vectors for the Registration of Free-Form ShapesabstractThe registration of free-form shapes by the iterative closest point algorithm (ICP) has attracted much attention from the computer vision and image processing community since it was first proposed in 1992. Many methods, mainly based on incorporating invariants described in a single coordinate frame have been devised to improve the accuracy and efficiency of the algorithm. In this paper, a novel method to improve image registration is proposed based on rigid constraints derived from geometric properties of correspondence vectors synthesised into a singe coordinate frame. False matches, which occur in almost every iteration of the ICP algorithm are eliminated through properties of the motion. For an accurate estimation of the geometric parameters of the motion, the Monte Carlo method is used in conjunction with a median filter. Experimental results based on both synthetic data and real images show that the improved method can effectively eliminate false matches, is accurate, robust, and efficient for the registration of free-form shapes with small motions. Yonghuai Liu, Marcos A. Rodrigues 0001 |
ICPR | 1 |
| 2000 | Learning and Diagnosis in Manufacturing Processes Through an Executable Bayesian Network
Marcos A. Rodrigues 0001, Yonghuai Liu, Leonardo Bottaci, Dimitris I. Rigas |
IEA/AIE | 2 |
| 2000 | Similarity based linear N≥5-point structure and pose estimation from it single imageabstractWe present a novel two-stage algorithm for structural and pose estimation from a single image. Our method is a significant improvement on the work described by Quan et al. (1999) because: our algorithm is developed in similarity space rather than in camera centred coordinate frame; our algorithm is directly based on a fourth order polynomial rather than on an eighth order polynomial resulting in an easier and more efficient implementation; and finally our algorithm is a unified linear algorithm for structural estimation. The algorithm has been extensively validated by experimentation using both synthetic and real images and compared with a classical linear algorithm and an algorithm based on epipolar geometry. Experimental results show that the proposed algorithm is generally accurate, robust, and efficient for the calibration of all parameters of interest. Yonghuai Liu, Marcos A. Rodrigues 0001 |
IROS | 1 |
| 2000 | Developing rigid motion constraints for the registration of free-form shapesabstractWe propose a novel method to deal with sphere ambiguity, occlusion, appearance and disappearance of points in image registration. We have developed a number of rigid motion constraints through analysis of geometrical properties of reflected correspondence vectors synthesised into a single coordinate frame. The properties are used as further constraints to eliminate false matches obtained by the iterative closest point criterion. A number of experiments based on both synthetic data and real images demonstrate that the proposed method is accurate, robust, and efficient for the registration of free-form shapes. Yonghuai Liu, Marcos A. Rodrigues 0001 |
IROS | 1 |
| 2000 | Analysing the geometric properties of reflected correspondence vectors for the registration of free form shapesabstractThe paper presents a novel algorithm for the registration of free-form shapes represented by two sets of points before and after a rigid body motion. First, the geometric properties of reflected correspondence vectors are analysed, yielding a number of rigid motion constraints bridging the points described in different coordinate frames before and after the rigid motion. Then, a novel algorithm is developed for the registration of free-form shapes, making full use of the rigid motion constraints to cope with false matches that occur in almost every iteration of the ICP algorithm. In order to improve the accuracy of the motion parameter estimation and computational efficiency, the Monte Carlo resampling technique and a median filter are employed. A number of experiments based on both synthetic data and real images demonstrate that the proposed algorithm is accurate and robust. Yonghuai Liu, Marcos A. Rodrigues 0001 |
SMC | 1 |
| 1999 | Using Rigid Constraints to Analyse Motion Parameters from Two Sets of 3D Corresponding Point Pattern
Yonghuai Liu, Marcos A. Rodrigues 0001 |
CAIP | 1 |
| 1999 | Correspondenceless Motion Estimation from Range ImagesabstractEstimation of rigid-body motion parameters in computer vision is normally performed from image correspondences between two coordinate frames. A large number of methods and algorithms have been proposed based on that such correspondences are known. Unfortunately, the establishment of correspondences is often time-consuming and, in many cases, impossible. In this paper, we propose a novel correspondenceless motion estimation algorithm based on the cross matrix. For a comparative study, we also implemented a correspondenceless motion estimation algorithm based on the scatter matrix. Experimental results have demonstrated that our method is more accurate and robust than the scatter matrix-based algorithm. Yonghuai Liu, Marcos A. Rodrigues 0001 |
ICCV | 1 |
| 1999 | Motion Parameter Constraints Analysis from a Single ImageabstractMotion parameter estimation is a fundamental problem in image processing and image understanding. A large number of algorithms have been proposed based on a number of different geometrical considerations, such as perspective or epipolar geometries. However, proposed motion estimation algorithms do not explicitly use the distance between feature points and angle information as rigid constraints to calibration. In this paper, we present a new geometric analysis of correspondence data and derive explicit expressions for rigid constraints that are then used to estimate motion parameters. We then present a novel, efficient motion parameter estimation algorithm based on a coarse to fine strategy from a single image data. For a comparative study of performance, we also extended to the 3D-2D case a well known 2D-2D motion estimation algorithm based on the epipolar geometry. Experimental results demonstrate that the coarse to fine strategy is appropriate for the problem and that the algorithm generally performs better than the extended epipolar geometry based algorithm. Marcos A. Rodrigues 0001, Yonghuai Liu |
ICIP (3) | 2 |
| 1999 | Geometric Understanding of Rigid Body TransformationsabstractIt has been demonstrated by Chasles over a century ago that any given displacement of a rigid body can be effected by a single rotation around an axis combined with a translation parallel to that axis. In this paper, we first describe geometric properties of image correspondence vectors. We then formalise such properties by extending Chasles' work and put forward a new general theoretical framework for the analysis of rigid body transformations in 2D and in 3D. We then propose two novel algorithms to calibrate rigid body transformation parameters which are validated through experiments. Our analysis addresses central issues to computer vision applications as it provides closed form solutions for all calibrated parameters and an accurate insight into the number of calibration solutions. Yonghuai Liu, Marcos A. Rodrigues 0001 |
ICRA | 1 |
| 1999 | A Novel 3D-3D Computer Vision Algorithm for Automatic Inspection of Filter Components
Marcos A. Rodrigues 0001, Yonghuai Liu |
IEA/AIE | 2 |
| 1999 | Geometrical Analysis of Two Sets of 3D Correspondence Data PatternsabstractGiven two sets of image correspondence data, the analysis of the corresponding rigid body transformation that accurately describes the object's motion in 3D space is a fundamental problem in image understanding. While many methods have been put forward to analyse 3D transformations, such methods do not make use of the vector distance between feature points and angular information as constraints to analyse transformation parameters. Current methods and algorithms have also suffered from the problems of lack of efficiency, sensitivity to noise, and multiplicity of solutions. In this paper we present a novel geometrical analysis of rigid body transformations and a novel algorithm to calibrate transformation parameters based on image correspondence. The algorithm is validated through experimental calibration and a comparison is made with calibration implemented by the least squares method. We demonstrate that the proposed algorithm works well in the presence of noise and that its performance is in general superior to algorithms based on the least squares method. Marcos A. Rodrigues 0001, Yonghuai Liu |
Shape Modeling International | 2 |
| 1999 | Invariant Geometric Properties of Image Correspondence Vectors as Rigid Constraints to Motion EstimationabstractAccurate motion estimation algorithms are based on a number of invariant properties that can be inferred from the motion. A large number of calibration algorithms have been proposed over the last two decades mainly based on analytic, perspective, or epipolar geometries. Extending Chasles' screw motion concept to the estimation of motion parameters in computer vision, we have presented an analysis of geometric properties of image correspondence vectors synthesized into a single coordinate frame and developed calibration algorithms using both simulated and real range image data.15,16 In this paper, we extend that work by defining the relevant geometric properties of image correspondence vectors from the point of view of invariants and by developing two calibration algorithms using the Monte Carlo method and median filtering. The algorithms are applied to real and synthetic image data corrupted by noise and outliers. Experimental results demonstrate that the median filter based algorithm is generally more robust and accurate than the Monte Carlo based algorithm and that the geometric analysis of invariant properties of correspondence vectors is a useful framework to motion parameter estimation. Yonghuai Liu, Marcos A. Rodrigues 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |