EDBT 2026 Demo / reviewers in the wild / expert
Xianghua Xie
dblp:17/5825
· DBLP profile ↗
99ranked-venue papers
16as first author
35since 2021 · last 2026
0000-0002-2701-8660ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 71 · 9 first-author · 18 since 2021Artificial intelligence and machine learning · 50 · 12 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ACE-Grasp: Aleatoric Ambiguity Modeling via Consistency and Exploration for Grasping
Xianghua Xie |
ICPR (10) | 2 |
| 2026 | Text-based three-dimensional geometric person retrievalabstractPerson Re-identification (Re-ID) is crucial in computer vision, widely applied in forensic investigation, intelligent surveillance, and video retrieval. Recent text-based Re-ID methods leverage eyewitness descriptions to enhance retrieval flexibility but still face challenges in accurately characterizing individuals under complex conditions. To address issues like low resolution, viewpoint variations, and occlusions, this paper proposes a novel text-based person Re-ID approach that integrates textual descriptions with synthesized Three-dimensional (3D) geometric pedestrian data derived from existing Two-dimensional (2D) images. Specifically, the semantic richness of text compensates for the lack of color and texture details in 3D data, while the robustness of geometric and pose information significantly enhances retrieval performance. Despite current 3D pedestrian data being generated through reconstruction algorithms, this work serves as a pioneering exploration of text-to-3D pedestrian retrieval, offering substantial potential for real-world applications in multimodal biometrics, forensic investigations, and privacy protection. Experiments on three public datasets demonstrate that our method achieves competitive performance, confirming its practical applicability and significance. Fanzhi Jiang, Hanchi Ren, Liumei Zhang, Yuanjiao Hu, Xianghua Xie, Su Yang 0002 |
Eng. Appl. Artif. Intell. | 7 |
| 2026 | A2D2C: Adaptive attention-driven dynamic convolution for local feature adaptationabstract• Introduces A 2 D 2 C: attention-driven dynamic convolution for local adaptation. • Uses multi-point random sampling to route and fuse k base kernels efficiently. • Presents A 2 D 2 C + that fuses kernels once, cutting redundancy and MAdds at parity. • Shows consistent gains on ImageNet, CIFAR-100 and COCO with statistical reports. Fan Wan, Xingyu Miao, Jingjing Deng 0001, Xianghua Xie, Yang Long 0001 |
Pattern Recognit. | 5 |
| 2025 | Active Deep Clustering: Exploratory Analysis to Assist in Decision-Making on Incremental Label Morphing Datasets
Connor Clarkson, Michael Edwards, Xianghua Xie |
ACIVS | 3 |
| 2025 | Task-Oriented Robotic Manipulation with Vision Language Models
Nurhan Bulus Guran, Hanchi Ren, Jingjing Deng 0001, Xianghua Xie |
ACIVS | 4 |
| 2025 | Pretraining Techniques for Steel Surface Roughness Prediction with Long Thin Spatial Industrial DataabstractMachine learning offers promising advancements in industrial processes, yet collecting labeled samples during production remains challenging. In steel production, the surface roughness $$R_a$$ parameter of steel coils is crucial, but on-line labeled data collection, with our apparatus, is infeasible, while off-line methods are time-consuming and imperfect. However, unlabeled samples are readily available from on-line production. This paper examines pretraining on a large, unlabeled dataset and its impacts on performance after fine-tuning on a smaller labeled dataset. We use three techniques: (1) contrastive learning, (2) Autoencoder, and (3) Classification of coil ID. We address the challenges posed by the unique structure of the data, comprising 2-dimensional, long and thin arrays. Our results show that our classification pretraining approach improves regression performance and outperforms the baseline. Alexander J. M. Milne, Xianghua Xie, Gary K. L. Tam |
ACIVS | 2 |
| 2025 | FissionVAE: Federated Non-IID Image Generation with Latent Space and Decoder DecompositionabstractFederated learning is a machine learning paradigm that enables decentralized clients to collaboratively learn a shared model while keeping all the training data local. While considerable research has focused on federated image generation, particularly Generative Adversarial Networks, Variational Autoencoders have received less attention. In this paper, we address the challenges of non-IID (independently and identically distributed) data environments featuring multiple groups of images of different types. Non-IID data distributions can lead to difficulties in maintaining a consistent latent space and can also result in local generators with disparate texture features being blended during aggregation. We thereby introduce FissionVAE that decouples the latent space and constructs decoder branches tailored to individual client groups. This method allows for customized learning that aligns with the unique data distributions of each group. Additionally, we incorporate hierarchical VAEs and demonstrate the use of heterogeneous decoder architectures within FissionVAE. We also explore strategies for setting the latent prior distributions to enhance the decoupling process. To evaluate our approach, we assemble two composite datasets: the first combines MNIST and FashionMNIST; the second comprises RGB datasets of cartoon and human faces, wild animals, marine vessels, and remote sensing images. Our experiments demonstrate that FissionVAE greatly improves generation quality on these datasets compared to baseline federated VAE models. Hanchi Ren, Jingjing Deng 0001, Xianghua Xie, Xiaoke Ma 0001 |
IJCAI | 4 |
| 2025 | Temporal Optimisation of Satellite Image-Based Crop Mapping: A Comparison of Deep Time Series and Semi-Supervised Time Warping StrategiesabstractABSTRACT This study presents a novel approach to crop mapping using remotely sensed satellite images. It addresses the significant classification modelling challenges, including (1) the requirements for extensive labelled data and (2) the complex optimisation problem for selection of appropriate temporal windows in the absence of prior knowledge of cultivation calendars. We compare the lightweight Dynamic Time Warping (DTW) classification method with the heavily supervised Convolutional Neural Network ‐ Long Short‐Term Memory (CNN‐LSTM) using high‐resolution multispectral optical satellite imagery (3 m/pixel). Our approach integrates effective practical preprocessing steps, including data augmentation and a data‐driven optimisation strategy for the temporal window, even in the presence of numerous crop classes. Our findings demonstrate that DTW, despite its lower data demands, can match the performance of CNN‐LSTM through our effective preprocessing steps while significantly improving runtime. These results demonstrate that both CNN‐LSTM and DTW can achieve deployment‐level accuracy and underscore the potential of DTW as a viable alternative to more resource‐intensive models. The results also prove the effectiveness of temporal windowing for improving runtime and accuracy of a crop classification study, even with no prior knowledge of planting timeframes. Rosie Finnegan, Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini, Xianghua Xie, Alberto Hornero, Nicholas W. Synes |
IET Comput. Vis. | 5 |
| 2025 | Sparse representation for restoring images by exploiting topological structure of graph of patchesabstractAbstract Image restoration poses a significant challenge, aiming to accurately recover damaged images by delving into their inherent characteristics. Various models and algorithms have been explored by researchers to address different types of image distortions, including sparse representation, grouped sparse representation, and low‐rank self‐representation. The grouped sparse representation algorithm leverages the prior knowledge of non‐local self‐similarity and imposes sparsity constraints to maintain texture information within images. To further exploit the intrinsic properties of images, this study proposes a novel low‐rank representation‐guided grouped sparse representation image restoration algorithm. This algorithm integrates self‐representation models and trace optimization techniques to effectively preserve the original image structure, thereby enhancing image restoration performance while retaining the original texture and structural information. The proposed method was evaluated on image denoising and deblocking tasks across several datasets, demonstrating promising results. Yaxian Gao, Zhaoyuan Cai, Xianghua Xie, Jingjing Deng 0001, Zengfa Dou, Xiaoke Ma 0001 |
IET Image Process. | 3 |
| 2025 | Learning multi-level topology representation for multi-view clustering with deep non-negative matrix factorization
Zengfa Dou, Weiming Hou, Xianghua Xie, Xiaoke Ma 0001 |
Neural Networks | 4 |
| 2025 | Class activation map guided level sets for weakly supervised semantic segmentation
Yifan Wang 0008, Gerald Schaefer, Xiyao Liu 0001, Jing Dong 0009, Linglin Jing, Xianghua Xie, Hui Fang 0003 |
Pattern Recognit. | 7 |
| 2024 | Depth-Aware Endoscopic Video Inpainting
Xiatian Zhang 0001, Shuang Chen 0010, Xianghua Xie, Hubert P. H. Shum |
MICCAI (6) | 3 |
| 2024 | Image restoration with group sparse representation and low-rank group residual learningabstractAbstract Image restoration, as a fundamental research topic of image processing, is to reconstruct the original image from degraded signal using the prior knowledge of image. Group sparse representation (GSR) is powerful for image restoration; it however often leads to undesirable sparse solutions in practice. In order to improve the quality of image restoration based on GSR, the sparsity residual model expects the representation learned from degraded images to be as close as possible to the true representation. In this article, a group residual learning based on low‐rank self‐representation is proposed to automatically estimate the true group sparse representation. It makes full use of the relation among patches and explores the subgroup structures within the same group, which makes the sparse residual model have better interpretation furthermore, results in high‐quality restored images. Extensive experimental results on two typical image restoration tasks (image denoising and deblocking) demonstrate that the proposed algorithm outperforms many other popular or state‐of‐the‐art image restoration methods. Zhaoyuan Cai, Xianghua Xie, Jingjing Deng 0001, Zengfa Dou, Bo Tong, Xiaoke Ma 0001 |
IET Image Process. | 2 |
| 2024 | Inferring Attention Shifts for Salient Instance RankingabstractAbstract The human visual system has limited capacity in simultaneously processing multiple visual inputs. Consequently, humans rely on shifting their attention from one location to another. When viewing an image of complex scenes, psychology studies and behavioural observations show that humans prioritise and sequentially shift attention among multiple visual stimuli. In this paper, we propose to predict the saliency rank of multiple objects by inferring human attention shift. We first construct a new large-scale salient object ranking dataset, with the saliency rank of objects defined by the order that an observer attends to these objects via attention shift. We then propose a new deep learning-based model to leverage both bottom-up and top-down attention mechanisms for saliency rank prediction. Our model includes three novel modules: Spatial Mask Module (SMM), Selective Attention Module (SAM) and Salient Instance Edge Module (SIEM). SMM integrates bottom-up and semantic object properties to enhance contextual object features, from which SAM learns the dependencies between object features and image features for saliency reasoning. SIEM is designed to improve segmentation of salient objects, which helps further improve their rank predictions. Experimental results show that our proposed network achieves state-of-the-art performances on the salient object ranking task across multiple datasets. Code and data are available at https://github.com/SirisAvishek/Attention_Shift_Ranks . Avishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie, Rynson W. H. Lau |
Int. J. Comput. Vis. | 4 |
| 2024 | FedBoosting: Federated learning with gradient protected boosting for text recognitionabstractConventional machine learning methodologies require the centralization of data for model training, which may be infeasible in situations where data sharing limitations are imposed due to concerns such as privacy and gradient protection. The Federated Learning (FL) framework enables the collaborative learning of a shared model without necessitating the centralization or sharing of data among the data proprietors. Nonetheless, in this paper, we demonstrate that the generalization capability of the joint model is suboptimal for Non-Independent and Non-Identically Distributed (Non-IID) data, particularly when employing the Federated Averaging (FedAvg) strategy as a result of the weight divergence phenomenon. Consequently, we present a novel boosting algorithm for FL to address both the generalization and gradient leakage challenges, as well as to facilitate accelerated convergence in gradient-based optimization. Furthermore, we introduce a secure gradient sharing protocol that incorporates Homomorphic Encryption (HE) and Differential Privacy (DP) to safeguard against gradient leakage attacks. Our empirical evaluation demonstrates that the proposed Federated Boosting (FedBoosting) technique yields significant enhancements in both prediction accuracy and computational efficiency in the visual text recognition task on publicly available benchmarks. Hanchi Ren, Jingjing Deng 0001, Xianghua Xie, Xiaoke Ma 0001 |
Neurocomputing | 3 |
| 2024 | A survey on vulnerability of federated learning: A learning algorithm perspectiveabstractFederated Learning (FL) has emerged as a powerful paradigm for training Machine Learning (ML), particularly Deep Learning (DL) models on multiple devices or servers while maintaining data localized at owners’ sites. Without centralizing data, FL holds promise for scenarios where data integrity, privacy and security and are critical. However, this decentralized training process also opens up new avenues for opponents to launch unique attacks, where it has been becoming an urgent need to understand the vulnerabilities and corresponding defense mechanisms from a learning algorithm perspective. This review paper takes a comprehensive look at malicious attacks against FL, categorizing them from new perspectives on attack origins and targets, and providing insights into their methodology and impact. In this survey, we focus on threat models targeting the learning process of FL systems. Based on the source and target of the attack, we categorize existing threat models into four types, Data to Model (D2M), Model to Data (M2D), Model to Model (M2M) and composite attacks. For each attack type, we discuss the defense strategies proposed, highlighting their effectiveness, assumptions and potential areas for improvement. Defense strategies have evolved from using a singular metric to excluding malicious clients, to employing a multifaceted approach examining client models at various phases. In this survey paper, our research indicates that the to-learn data, the learning gradients, and the learned model at different stages all can be manipulated to initiate malicious attacks that range from undermining model performance, reconstructing private local data, and to inserting backdoors. We have also seen these threat are becoming more insidious. While earlier studies typically amplified malicious gradients, recent endeavors subtly alter the least significant weights in local models to bypass defense measures. This literature review provides a holistic understanding of the current FL threat landscape and highlights the importance of developing robust, efficient, and privacy-preserving defenses to ensure the safe and trusted adoption of FL in real-world applications. The categorized bibliography can be found at: https://github.com/Rand2AI/Awesome-Vulnerability-of-Federated-Learning. Xianghua Xie, Hanchi Ren, Jingjing Deng 0001 |
Neurocomputing | 1 |
| 2024 | Jacobian norm with Selective Input Gradient Regularization for interpretable adversarial defenseabstractDeep neural networks (DNNs) can be easily deceived by imperceptible alterations known as adversarial examples. These examples can lead to misclassification , posing a significant threat to the reliability of deep learning systems in real-world applications. Adversarial training (AT) is a popular technique used to enhance robustness by training models on a combination of corrupted and clean data. However, existing AT-based methods often struggle to handle transferred adversarial examples that can fool multiple defense models, thereby falling short of meeting the generalization requirements for real-world scenarios. Furthermore, AT typically fails to provide interpretable predictions, which are crucial for domain experts seeking to understand the behavior of DNNs. To overcome these challenges, we present a novel approach called Jacobian norm and Selective Input Gradient Regularization (J-SIGR). Our method leverages Jacobian normalization to improve robustness and introduces regularization of perturbation-based saliency maps, enabling interpretable predictions. By adopting J-SIGR, we achieve enhanced defense capabilities and promote high interpretability of DNNs. We evaluate the effectiveness of J-SIGR across various architectures by subjecting it to powerful adversarial attacks. Our experimental evaluations provide compelling evidence of the efficacy of J-SIGR against transferred adversarial attacks, while preserving interpretability. The project code can be found at https://github.com/Lywu-github/jJ-SIGR.git . Deyin Liu, Lin Wu 0001, Bo Li 0090, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie, Chengwu Liang |
Pattern Recognit. | 6 |
| 2024 | Fully Connected Networks on a Diet With the Mediterranean Matrix MultiplicationabstractThis article proposes the Mediterranean matrix multiplication, a new, simple and practical randomized algorithm that samples angles between the rows and columns of two matrices with sizes m, n, and p to approximate matrix multiplication in O(k(mn+np+mp)) steps, where k is a constant only related to the precision desired. The number of instructions carried out is mainly bounded by bitwise operators, amenable to a simplified processing architecture and compressed matrix weights. Results show that the method is superior in size and number of operations to the standard approximation with signed matrices. Equally important, this article demonstrates a first application to machine learning inference by showing that weights of fully connected layers can be compressed between 30 × and 100 × with little to no loss in inference accuracy. The requirements for pure floating-point operations are also down as our algorithm relies mainly on simpler bitwise operators. Hassan Eshkiki, Benjamin Mora, Xianghua Xie |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Image Template Matching via Dense and Consistent Contrastive LearningabstractImage template matching refers to localizing a small query image as opposed to a large reference image map. The query image a.k.a template has to be screened across every equal-sized region in the reference map to perform inner-product at pixel-level and the resulting similarity indicates the template location. Due to the domain heterogeneity between template and reference images, the matching performance degrades under dramatic appearance changes. More severely, the asymmetric matching easily leads to over-fitting by suggesting excessively false positive regions. To these ends, we propose an effective template matching method based on contrastive learning to perform a dense and consistent InfoNCEloss during matching. This can increase the matching at finer details, and thus effectively regularizes network training to prevent over-fitting. Extensive experiments on the synthetic aperture radar (SAR) and optical datasets, i.e., SEN1-2 and OS datasets demonstrate that our proposed method outperforms state-of-the-art methods by a large margin. Bo Li 0090, Lin Wu 0001, Deyin Liu, Hongyang Chen 0001, Yuanxin Ye, Xianghua Xie |
ICME | 6 |
| 2023 | Face reenactment via generative landmark guidance
Xianghua Xie |
Image Vis. Comput. | 2 |
| 2023 | Learning Resolution-Adaptive Representations for Cross-Resolution Person Re-IdentificationabstractCross-resolution person re-identification (CRReID) is a challenging and practical problem that involves matching low-resolution (LR) query identity images against high-resolution (HR) gallery images. Query images often suffer from resolution degradation due to the different capturing conditions from real-world cameras. State-of-the-art solutions for CRReID either learn a resolution-invariant representation or adopt a super-resolution (SR) module to recover the missing information from the LR query. In this paper, we propose an alternative SR-free paradigm to directly compare HR and LR images via a dynamic metric that is adaptive to the resolution of a query image. We realize this idea by learning resolution-adaptive representations for cross-resolution comparison. We propose two resolution-adaptive mechanisms to achieve this. The first mechanism encodes the resolution specifics into different subvectors in the penultimate layer of the deep neural network, creating a varying-length representation. To better extract resolution-dependent information, we further propose to learn resolution-adaptive masks for intermediate residual feature blocks. A novel progressive learning strategy is proposed to train those masks properly. These two mechanisms are combined to boost the performance of CRReID. Experimental results show that the proposed method outperforms existing approaches and achieves state-of-the-art performance on multiple CRReID benchmarks. Lin Wu 0001, Lingqiao Liu, Yang Wang 0023, Zheng Zhang 0006, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie |
IEEE Trans. Image Process. | 7 |
| 2022 | Multi-Scale Gridded Gabor Attention for Cirrus SegmentationabstractIn this paper, we address the challenge of segmenting global contaminants in large images. The precise delineation of such structures requires ample global context alongside understanding of textural patterns. CNNs specialise in the latter, though their ability to generate global features is limited. Attention measures long range dependencies in images, capturing global context, though at a large computational cost. We propose a gridded attention mechanism to address this limitation, greatly increasing efficiency by processing multi-scale features into smaller tiles. We also enhance the attention mechanism for increased sensitivity to texture orientation, by measuring correlations across features dependent on different orientations, in addition to channel and positional attention. We present results on a new dataset of astronomical images, where the task is segmenting large contaminating dust clouds. Felix Richards, Xianghua Xie, Adeline Paiement, Elisabeth Sola, Pierre-Alain Duc |
ICIP | 2 |
| 2022 | A Deep Learning Driven Active Framework for Segmentation of Large 3D Shape Collections
David George 0001, Xianghua Xie, Yukun Lai, Gary K. L. Tam |
Comput. Aided Des. | 2 |
| 2022 | A hybrid method of detecting flame from video streamabstractAbstract In this paper, a method of detecting flame from video stream is proposed exploiting the characteristics of the disordered movement, rapid deformation and intense colour of the flame. Firstly, the frame difference between video frame and background frame is calculated to obtain the main part of the moving object, and the difference between frames is calculated frame by frame in time series to obtain the deformation part of the moving object, and then the sum of cumulative difference between frames and the background difference between frames are added to generate a binary image containing the moving object and the deformed part. Secondly, the binary image is morphologically opened, and rectangular segmentation is carried out to obtain multiple suspicious flame regions. Finally, in the light of the intense colour of the flame, the corresponding area is extracted from the original picture by using the segmentation rectangle, and the colour statistics of the area are carried out to further judge whether there is a burning flame in the area. The experimental results show that the algorithm can accurately detect the burning area of flame in the real scene and eliminate the light interference and the movement interference. Zengfa Dou, Xiaoke Ma 0001, Xianghua Xie, Chubing Guo |
IET Image Process. | 3 |
| 2022 | MLMT-CNN for object detection and segmentation in multi-layer and multi-spectral imagesabstractAbstract Precisely localising solar Active Regions (AR) from multi-spectral images is a challenging but important task in understanding solar activity and its influence on space weather. A main challenge comes from each modality capturing a different location of the 3D objects, as opposed to typical multi-spectral imaging scenarios where all image bands observe the same scene. Thus, we refer to this special multi-spectral scenario as multi-layer. We present a multi-task deep learning framework that exploits the dependencies between image bands to produce 3D AR localisation (segmentation and detection) where different image bands (and physical locations) have their own set of results. Furthermore, to address the difficulty of producing dense AR annotations for training supervised machine learning (ML) algorithms, we adapt a training strategy based on weak labels (i.e. bounding boxes) in a recursive manner. We compare our detection and segmentation stages against baseline approaches for solar image analysis (multi-channel coronal hole detection, SPOCA for ARs) and state-of-the-art deep learning methods (Faster RCNN, U-Net). Additionally, both detection and segmentation stages are quantitatively validated on artificially created data of similar spatial configurations made from annotated multi-modal magnetic resonance images. Our framework achieves an average of 0.72 IoU (segmentation) and 0.90 F1 score (detection) across all modalities, comparing to the best performing baseline methods with scores of 0.53 and 0.58, respectively, on the artificial dataset, and 0.84 F1 score in the AR detection task comparing to baseline of 0.82 F1 score. Our segmentation results are qualitatively validated by an expert on real ARs. Majedaldein Almahasneh, Adeline Paiement, Xianghua Xie, Jean Aboudarham |
Mach. Vis. Appl. | 3 |
| 2022 | Joint multi-label learning and feature extraction for temporal link prediction
Xiaoke Ma 0001, Shiyin Tan, Xianghua Xie, Xiaoxiong Zhong, Jingjing Deng 0001 |
Pattern Recognit. | 3 |
| 2022 | A directed graph convolutional neural network for edge-structured signals in link-fault detection
Michael P. Kenning, Jingjing Deng 0001, Michael Edwards, Xianghua Xie |
Pattern Recognit. Lett. | 4 |
| 2022 | GRNN: Generative Regression Neural Network - A Data Leakage Attack for Federated LearningabstractData privacy has become an increasingly important issue in Machine Learning (ML) , where many approaches have been developed to tackle this challenge, e.g., cryptography ( Homomorphic Encryption (HE) , Differential Privacy (DP) ) and collaborative training (Secure Multi-Party Computation (MPC) , Distributed Learning, and Federated Learning (FL) ). These techniques have a particular focus on data encryption or secure local computation. They transfer the intermediate information to the third party to compute the final result. Gradient exchanging is commonly considered to be a secure way of training a robust model collaboratively in Deep Learning (DL) . However, recent researches have demonstrated that sensitive information can be recovered from the shared gradient. Generative Adversarial Network (GAN) , in particular, has shown to be effective in recovering such information. However, GAN based techniques require additional information, such as class labels that are generally unavailable for privacy-preserved learning. In this article, we show that, in the FL system, image-based privacy data can be easily recovered in full from the shared gradient only via our proposed Generative Regression Neural Network (GRNN) . We formulate the attack to be a regression problem and optimize two branches of the generative model by minimizing the distance between gradients. We evaluate our method on several image classification tasks. The results illustrate that our proposed GRNN outperforms state-of-the-art methods with better stability, stronger robustness, and higher accuracy. It also has no convergence requirement to the global FL model. Moreover, we demonstrate information leakage using face re-identification. Some defense strategies are also discussed in this work. Hanchi Ren, Jingjing Deng 0001, Xianghua Xie |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | Scene Context-Aware Salient Object DetectionabstractSalient object detection identifies objects in an image that grab visual attention. Although contextual features are considered in recent literature, they often fail in real-world complex scenarios. We observe that this is mainly due to two issues: First, most existing datasets consist of simple foregrounds and backgrounds that hardly represent real-life scenarios. Second, current methods only learn contextual features of salient objects, which are insufficient to model high-level semantics for saliency reasoning in complex scenes. To address these problems, we first construct a new large-scale dataset with complex scenes in this paper. We then propose a context-aware learning approach to explicitly exploit the semantic scene contexts. Specifically, two modules are proposed to achieve the goal: 1) a Semantic Scene Context Refinement module to enhance contextual features learned from salient objects with scene context, and 2) a Contextual Instance Transformer to learn contextual relations between objects and scene context. To our knowledge, such high-level semantic contextual information of image scenes is under-explored for saliency detection in the literature. Extensive experiments demonstrate that the proposed approach outperforms state-of-the-art techniques in complex scenarios for saliency detection, and transfers well to other existing datasets. The code and dataset are available at https://github.com/SirisAvishek/Scene_Context_Aware_Saliency. Avishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie, Rynson W. H. Lau |
ICCV | 4 |
| 2021 | Active Region Detection in Multi-spectral Solar ImagesabstractPrecisely detecting solar Active Regions (AR) from multi-spectral images is a challenging task yet important in understanding solar activity and its influence on space weather. A main challenge comes from each modality capturing a different location of these 3D objects, as opposed to more traditional multi-spectral imaging scenarios where all image bands observe the same scene. We present a multi-task deep learning framework that exploits the dependencies between image bands to produce 3D AR detection where different image bands (and physical locations) each have their own set of results. We compare our detection method against baseline approaches for solar image analysis (multi-channel coronal hole detection, SPOCA for ARs (Verbeeck et al., 2013)) and a state-of-the-art deep learning method (Faster RCNN) and show enhanced performances in detecting ARs jointly from multiple bands. Majedaldein Almahasneh, Adeline Paiement, Xianghua Xie, Jean Aboudarham |
ICPRAM | 3 |
| 2021 | Graph Convolution Networks for Cell Segmentation
Sachin Bahade, Michael Edwards, Xianghua Xie |
ICPRAM | 3 |
| 2021 | Locating Datacenter Link Faults with a Directed Graph Convolutional Neural Network
Michael P. Kenning, Jingjing Deng 0001, Michael Edwards, Xianghua Xie |
ICPRAM | 4 |
| 2021 | TLGP: a flexible transfer learning algorithm for gene prioritization based on heterogeneous source domainabstractBACKGROUND: Gene prioritization (gene ranking) aims to obtain the centrality of genes, which is critical for cancer diagnosis and therapy since keys genes correspond to the biomarkers or targets of drugs. Great efforts have been devoted to the gene ranking problem by exploring the similarity between candidate and known disease-causing genes. However, when the number of disease-causing genes is limited, they are not applicable largely due to the low accuracy. Actually, the number of disease-causing genes for cancers, particularly for these rare cancers, are really limited. Therefore, there is a critical needed to design effective and efficient algorithms for gene ranking with limited prior disease-causing genes. RESULTS: In this study, we propose a transfer learning based algorithm for gene prioritization (called TLGP) in the cancer (target domain) without disease-causing genes by transferring knowledge from other cancers (source domain). The underlying assumption is that knowledge shared by similar cancers improves the accuracy of gene prioritization. Specifically, TLGP first quantifies the similarity between the target and source domain by calculating the affinity matrix for genes. Then, TLGP automatically learns a fusion network for the target cancer by fusing affinity matrix, pathogenic genes and genomic data of source cancers. Finally, genes in the target cancer are prioritized. The experimental results indicate that the learnt fusion network is more reliable than gene co-expression network, implying that transferring knowledge from other cancers improves the accuracy of network construction. Moreover, TLGP outperforms state-of-the-art approaches in terms of accuracy, improving at least 5%. CONCLUSION: The proposed model and method provide an effective and efficient strategy for gene ranking by integrating genomic data from various cancers. Zuheng Xia, Jingjing Deng 0001, Xianghua Xie, Maoguo Gong, Xiaoke Ma 0001 |
BMC Bioinform. | 4 |
| 2021 | Pruning CNN filters via quantifying the importance of deep visual representations
Ali Alqahtani 0001, Xianghua Xie, Mark W. Jones 0001, Ehab Essa |
Comput. Vis. Image Underst. | 2 |
| 2021 | 3D Interactive Segmentation With Semi-Implicit Representation and Active LearningabstractSegmenting complex 3D geometry is a challenging task due to rich structural details and complex appearance variations of target object. Shape representation and foreground-background delineation are two of the core components of segmentation. Explicit shape models, such as mesh based representations, suffer from poor handling of topological changes. On the other hand, implicit shape models, such as level-set based representations, have limited capacity for interactive manipulation. Fully automatic segmentation for separating foreground objects from background generally utilizes non-interoperable machine learning methods, which heavily rely on the off-line training dataset and are limited to the discrimination power of the chosen model. To address these issues, we propose a novel semi-implicit representation method, namely Non-Uniform Implicit B-spline Surface (NU-IBS), which adaptively distributes parametrically blended patches according to geometrical complexity. Then, a two-stage cascade classifier is introduced to carry out efficient foreground and background delineation, where a simplistic Naïve-Bayesian model is trained for fast background elimination, followed by a stronger pseudo-3D Convolutional Neural Network (CNN) multi-scale classifier to precisely identify the foreground objects. A localized interactive and adaptive segmentation scheme is incorporated to boost the delineation accuracy by utilizing the information iteratively gained from user intervention. The segmentation result is obtained via deforming an NU-IBS according to the probabilistic interpretation of delineated regions, which also imposes a homogeneity constrain for individual segments. The proposed method is evaluated on a 3D cardiovascular Computed Tomography Angiography (CTA) image dataset and Brain Tumor Image Segmentation Benchmark 2015 (BraTS2015) 3D Magnetic Resonance Imaging (MRI) dataset. Jingjing Deng 0001, Xianghua Xie |
IEEE Trans. Image Process. | 2 |
| 2020 | Inferring Attention Shift Ranks of Objects for Image SaliencyabstractPsychology studies and behavioural observation show that humans shift their attention from one location to another when viewing an image of a complex scene. This is due to the limited capacity of the human visual system in simultaneously processing multiple visual inputs. The sequential shifting of attention on objects in a non-task oriented viewing can be seen as a form of saliency ranking. Although there are methods proposed for predicting saliency rank, they are not able to model this human attention shift well, as they are primarily based on ranking saliency values from binary prediction. Following psychological studies, in this paper, we propose to predict the saliency rank by inferring human attention shift. Due to the lack of such data, we first construct a large-scale salient object ranking dataset. The saliency rank of objects is defined by the order that an observer attends to these objects based on attention shift. The final saliency rank is an average across the saliency ranks of multiple observers. We then propose a learning-based CNN to leverage both bottom-up and top-down attention mechanisms to predict the saliency rank. Experimental results show that the proposed network achieves state-of-the-art performances on salient object rank prediction. Code and dataset are available at https://github.com/SirisAvishek/Attention_Shift_Ranks. Avishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie, Rynson W. H. Lau |
CVPR | 4 |
| 2020 | Neuron-based Network Pruning Based on Majority VotingabstractThe achievement of neural networks in a variety of applications is accompanied by a dramatic increase in computational costs and memory requirements. In this paper, we propose an efficient method to simultaneously identify the critical neurons and prune the model during training without involving any pre-training or fine-tuning procedures. Unlike existing methods, which accomplish this task in a greedy fashion, we propose a majority voting technique to compare the activation values among neurons and assign a voting score to quantitatively evaluate their importance. This mechanism helps to effectively reduce model complexity by eliminating the less influential neurons and aims to determine a subset of the whole model that can represent the reference model with much fewer parameters within the training process. Experimental results show that majority voting efficiently compresses the network with no drop in model accuracy, pruning more than 79% of the original model parameters on CIFAR10 and more than 91% of the original parameters on MNIST. Moreover, we show that with our proposed method, sparse models can be further pruned into even smaller models by removing more than 60% of the parameters, whilst preserving the reference model accuracy. Ali Alqahtani 0001, Xianghua Xie, Ehab Essa, Mark W. Jones 0001 |
ICPR | 2 |
| 2020 | Using Machine Learning to Refer Patients with Chronic Kidney Disease to Secondary CareabstractThere has been growing interest recently in using machine learning techniques as an aid in clinical medicine. Machine learning offers a range of classification algorithms which can be applied to medical data to aid in making clinical predictions. Recent studies have demonstrated the high predictive accuracy of various classification algorithms applied to clinical data. Several studies have already been conducted in diagnosing or predicting chronic kidney disease at various stages using different sets of variables. In this study we are investigating the use of machine learning techniques with blood test data. Such a system could aid renal teams in making recommendations to primary care general practitioners to refer patients to secondary care where patients may benefit from earlier specialist assessment and medical intervention. We are able to achieve an overall accuracy of 88.48% using logistic regression, 87.12% using ANN and 85.29% using SVM. ANNs performed with the highest sensitivity at 89.74 % compared to 86.67 % for logistic regression and 85.51 % for SVM. Lee Au-Yeung, Xianghua Xie, James Chess, Timothy Scale |
ICPR | 2 |
| 2020 | Deep Learning Based Sepsis Intervention: The Modelling and Prediction of Severe Sepsis OnsetabstractSepsis presents a significant challenge to healthcare providers during critical care scenarios such as within an intensive care unit. The prognosis of the onset of severe septic shock results in significant increases in mortality rate, length of stay and readmission rates. Continual advancements in health informatics data allows for applications within the machine learning field to predict sepsis onset in a timely manner, allowing for effective preventative intervention of severe septic shock. A novel deep learning application is proposed to provide effective prediction of sepsis onset by up to six hours prior, involving the use of novel concepts such as a boosted cascading training methodology and adjustable margin hinge loss function. The proposed methodology provides statistically significant improvements to that of current machine learning based modelling applications based off the Physionet Computing in Cardiology 2019 challenge. Results show test F1 scores of 0.420, a significant improvement of 0.281 as compared to the next best challenger results. Gavin Tsang, Xianghua Xie |
ICPR | 2 |
| 2020 | Graph convolutional neural network for multi-scale feature learning
Michael Edwards, Xianghua Xie, Robert Ieuan Palmer, Gary K. L. Tam, Rob Alcock, Carl Roobottom |
Comput. Vis. Image Underst. | 2 |
| 2019 | Learning Discriminatory Deep Clustering Models
Ali Alqahtani 0001, Xianghua Xie, Jingjing Deng 0001, Mark W. Jones 0001 |
CAIP (1) | 2 |
| 2019 | Consistent segment-wise matching with multi-layer graphs
Taiwei Wang, David George 0001, Yukun Lai, Xianghua Xie, Gary K. L. Tam |
Comput. Aided Geom. Des. | 4 |
| 2019 | TimeCluster: dimension reduction applied to temporal data for visual analyticsabstractThere is a need for solutions which assist users to understand long time-series data by observing its changes over time, finding repeated patterns, detecting outliers, and effectively labeling data instances. Although these tasks are quite distinct and are usually tackled separately, we present an interactive visual analytics system and approach that can address these issues in a single system. It enables users to visualize, understand and explore univariate or multivariate long time-series data in one image using a connected scatter plot. It supports interactive analysis and exploration for pattern discovery and outlier detection. Different dimensionality reduction techniques are used and compared in our system. Because of its power of extracting features, deep learning is used for multivariate time-series along with 2D reduction techniques for rapid and easy interpretation and interaction with large amount of time-series data. We deploy our system with different time-series datasets and report two real-world case studies that are used to evaluate our system. Mark W. Jones 0001, Xianghua Xie |
Vis. Comput. | 3 |
| 2018 | A Deep Convolutional Auto-Encoder with Embedded ClusteringabstractIn this paper, we propose a clustering approach embedded in a deep convolutional auto-encoder (DCAE). In contrast to conventional clustering approaches, our method simultaneously learns feature representations and cluster assignments through DCAEs. DCAEs have been effective in image processing as it fully utilizes the properties of convolutional neural networks. Our method consists of clustering and reconstruction objective functions. All data points are assigned to their new corresponding cluster centers during the optimization, after that, clustering centers are iteratively updated to obtain a stable performance of clustering. The experimental results on the MNIST dataset show that the proposed method substantially outperforms deep clustering models in term of clustering quality. Ali Alqahtani 0001, Xianghua Xie, Jingjing Deng 0001, Mark W. Jones 0001 |
ICIP | 2 |
| 2018 | Local Representation Learning with A Convolutional AutoencoderabstractVery recent advances in machine learning have expanded deep learning methods to spatially-irregular data domains. Deep learning on graphs in particular has received greater study, providing benefits in numerous fields. In this paper we present a graph-based convolutional autoencoder and assess the contribution of four components towards encoding quality. A graph-based convolution-operator is used to learn localised filtering operations for graph-wise encoding. An evaluation of the proposed method is provided on a topologically-irregular version of MNIST that violates the assumption made by conventional convolutional autoencoder methods of the structure of its input-data. Michael P. Kenning, Xianghua Xie, Michael Edwards, Jingjing Deng 0001 |
ICIP | 2 |
| 2018 | Recurrent Neural Networks for Financial Time-Series ModellingabstractThe prediction of financial time series data is a challenging task due to the unpredictable behaviours of investors that are influenced by a multitude of factors. In this paper, we present a novel deep Long Short-Term Memory (LSTM) based time-series data modelling for use in stock market index prediction. A dataset comprised of six market indices from around the world were chosen to demonstrate the robustness in varying market conditions with an aim to forecast the next day closing price. With experimental results showing an average annual profitability performance of up to 200%, our method demonstrates its feasibility and significant results in time-series modelling and prediction of financial markets. Gavin Tsang, Jingjing Deng 0001, Xianghua Xie |
ICPR | 3 |
| 2018 | 3D mesh segmentation via multi-branch 1D convolutional neural networks
David George 0001, Xianghua Xie, Gary K. L. Tam |
Graph. Model. | 2 |
| 2017 | AMD Classification in Choroidal OCT Using Hierarchical Texton Mining
Dafydd Ravenscroft, Jingjing Deng 0001, Xianghua Xie, Louise Terry, Tom H. Margrain, Rachel V. North, Ashley Wood |
ACIVS | 3 |
| 2017 | Nested Shallow CNN-Cascade for Face Detection in the WildabstractFace detection in the wild is a challenging vision problem due to large variations and unpredictable ambiguities commonly existed in real world images. Whilst introducing powerful but complex models is often computationally inefficient, using hand-crafted features is hence problematic. In this paper, we propose a nested CNN-cascade learning algorithm that adopts shallow neural network architectures that allow efficient and progressive elimination of negative hypothesis from easy to hard via self-learning discriminative representations from coarse to fine scales. The face detection problem is considered as solving three sub-problems: eliminating easy background with a simple but fast model, then localising the face region with a soft-cascade, followed by precise detection and localisation by verifying retained regions with a deeper and stronger model. The face detector is trained on the AFLW dataset following the standard evaluation procedure, and the method is tested on four other public datasets, i.e. FDDB, AFW, CMU-MIT and GENKI. Both quantitative and qualitative results on FDDB and AFW are reported, which show promising performances on detecting faces in unconstrained environment. Jingjing Deng 0001, Xianghua Xie |
FG | 2 |
| 2017 | Detect face in the wild using CNN cascade with feature aggregation at multi-resolutionabstractFace detection in the wild is a challenging vision problem due to large variations and unpredictable ambiguities commonly existed in real world images. Whilst using hand-crafted features is generally problematic, introducing powerful but complex models is often computationally inefficient. Feature aggregation and multi-resolution are two efficient strategies for traditional visual recognition methods. In this paper, we show that such strategies can be integrated into Convolutional Neural Network (CNN) architecture via average pooling and channel-wise feature concatenation. Shallow networks with feature aggregation at multi-resolution enables the traditional cascade framework to tackle the challenging detection problems efficiently. The proposed method is tested on a public benchmark with across dataset evaluation. Both quantitative and qualitative results show promising performance improvements on detecting faces in unconstrained environment. Jingjing Deng 0001, Xianghua Xie |
ICIP | 2 |
| 2017 | Automatic segmentation of cross-sectional coronary arterial images
Ehab Essa, Xianghua Xie |
Comput. Vis. Image Underst. | 2 |
| 2017 | Recognition, Tracking, and Optimisation
Xianghua Xie, Mark W. Jones 0001, Gary K. L. Tam |
Int. J. Comput. Vis. | 1 |
| 2016 | Combining Stacked Denoising Autoencoders and Random Forests for Face Detection
Jingjing Deng 0001, Xianghua Xie, Michael Edwards |
ACIVS | 2 |
| 2016 | Neural Network Boundary Detection for 3D Vessel Segmentation
Robert Ieuan Palmer, Xianghua Xie |
ACIVS | 2 |
| 2016 | Graph Convolutional Neural Network
Michael Edwards, Xianghua Xie |
BMVC | 2 |
| 2016 | From pose to activity: Surveying datasets and introducing CONVERSE
Michael Edwards, Jingjing Deng 0001, Xianghua Xie |
Comput. Vis. Image Underst. | 3 |
| 2016 | Fixing the root node: Efficient tracking and detection of 3D human pose through local solutions
Ben Daubney, Xianghua Xie, Jingjing Deng 0001, Neil Mac Parthaláin, Reyer Zwiggelaar |
Image Vis. Comput. | 2 |
| 2016 | Registration and Modeling From Spaced and Misaligned Image VolumesabstractWe address the problem of object modeling from 3D and 3D+T data made up of images, which contain different parts of an object of interest, are separated by large spaces, and are misaligned with respect to each other. These images have only a limited number of intersections, hence making their registration particularly challenging. Furthermore, such data may result from various medical imaging modalities and can, therefore, present very diverse spatial configurations. Previous methods perform registration and object modeling (segmentation and interpolation) sequentially. However, sequential registration is ill-suited for the case of images with few intersections. We propose a new methodology, which, regardless of the spatial configuration of the data, performs the three stages of registration, segmentation, and shape interpolation from spaced and misaligned images simultaneously. We integrate these three processes in a level set framework, in order to benefit from their synergistic interactions. We also propose a new registration method that exploits segmentation information rather than pixel intensities, and that accounts for the global shape of the object of interest, for increased robustness and accuracy. The accuracy of registration is compared against traditional mutual information based methods, and the total modeling framework is assessed against traditional sequential processing and validated on artificial, CT, and MRI data. Adeline Paiement, Majid Mirmehdi, Xianghua Xie, Mark C. K. Hamilton |
IEEE Trans. Image Process. | 3 |
| 2015 | Automatic Aortic Root Segmentation with Shape Constraints and Mesh RegularisationabstractFully automated 3D segmentation is not only challenging due to, for instance, ambiguities in appearance, but it is also computationally demanding.We present a fullyautomatic, learning-based deformable modelling method for segmenting the aortic root in CT images using a two-stage mesh deformation: a non-iterative boundary segmentation with a statistical shape model for shape constraint, followed by an iterative boundary refinement process.At both stages, we introduce a B-spline mesh regularisation technique to avoid mesh entanglement during deformation.The initialisation of the deformable model is achieved through efficient detection and localisation of the aortic root using marginal space learning, which carries out similarity parameter estimation in an incremental fashion.Quantitative comparisons are carried out against a state-of-the-art deformable model-based approach and an active shape model based segmentation.The proposed method achieves both a lower average mesh error of 1.39 ± 0.29mm, and Hausdorff distance of 6.75 ± 2.05mm.Compared to these two approaches, it results in much more regularised mesh surfaces with no tangled mesh faces. Robert Ieuan Palmer, Xianghua Xie, Gary K. L. Tam |
BMVC | 2 |
| 2015 | Minimum S-Excess Graph for Segmenting and Tracking Multiple Borders with HMM
Ehab Essa, Xianghua Xie, Jonathan-Lee Jones |
MICCAI (2) | 2 |
| 2015 | Divergence of Gradient Convolution: Deformable Segmentation With Arbitrary InitializationsabstractIn this paper, we propose a unified approach to deformable model-based segmentation. The fundamental force field of the proposed method is based on computing the divergence of a gradient convolution field (GCF), which makes the full use of directional information of the image gradient vectors and their interactions across image domain. However, instead of directly using such a vector field for deformable segmentation as in the conventional approaches, we derive a more salient representation for contour evolution, and very importantly, we demonstrate that this representation of image force field not only leads to global minimum through convex relaxation but also can achieve the same result using the conventional gradient descent with an intrinsic regularization. Thus, the proposed method can handle arbitrary initializations. The proposed external force field for deformable segmentation has both edge-based properties in that the GCF is computed from image gradients, and the region-based attributes since its divergence can be treated as a region indication function. Moreover, nonlinear diffusion can be conveniently applied to GCF to improve its performance in dealing with noise interference. We also show the extension of GCF from 2D to 3D. In comparison to the state-of-the-art deformable segmentation techniques, the proposed method shows greater flexibility in model initialization and optimization realization, as well as better performance toward noise interference and appearance variation. Huaizhong Zhang, Xianghua Xie |
IEEE Trans. Image Process. | 2 |
| 2014 | 3D interactive coronary artery segmentation using random forests and Markov random field optimizationabstractCoronary artery segmentation plays a vital important role in coronary disease diagnosis and treatment. In this paper, we present a machine learning based interactive coronary artery segmentation method for 3D computed tomography angiography images. We first apply vessel diffusion to reduce noise interference and enhance the tubular structures in the images. A few user strokes are required to specify region of interest and background. Various image features for detecting the coronary arteries are then extracted in a multi-scale fashion, and are fed into a random forests classifier, which assigns each voxel with probability values of being coronary artery and background. The final segmentation is carried through an MRF based optimization using primal dual algorithm. A connectivity component analysis is carried out as post processing to remove isolated, small regions to produce the segmented coronary arterial vessels. The proposed method requires limited user interference and achieves robust segmentation results. Jingjing Deng 0001, Xianghua Xie, Rob Alcock, Carl Roobottom |
ICIP | 2 |
| 2014 | A bag of words approach to subject specific 3D human pose interaction classification with random decision forests
Jingjing Deng 0001, Xianghua Xie, Ben Daubney |
Graph. Model. | 2 |
| 2014 | Integrated Segmentation and Interpolation of Sparse DataabstractWe address the two inherently related problems of segmentation and interpolation of 3D and 4D sparse data and propose a new method to integrate these stages in a level set framework. The interpolation process uses segmentation information rather than pixel intensities for increased robustness and accuracy. The method supports any spatial configurations of sets of 2D slices having arbitrary positions and orientations. We achieve this by introducing a new level set scheme based on the interpolation of the level set function by radial basis functions. The proposed method is validated quantitatively and/or subjectively on artificial data and MRI and CT scans and is compared against the traditional sequential approach, which interpolates the images first, using a state-of-the-art image interpolation method, and then segments the interpolated volume in 3D or 4D. In our experiments, the proposed framework yielded similar segmentation results to the sequential approach but provided a more robust and accurate interpolation. In particular, the interpolation was more satisfactory in cases of large gaps, due to the method taking into account the global shape of the object, and it recovered better topologies at the extremities of the shapes where the objects disappear from the image slices. As a result, the complete integrated framework provided more satisfactory shape reconstructions than the sequential approach. Adeline Paiement, Majid Mirmehdi, Xianghua Xie, Mark C. K. Hamilton |
IEEE Trans. Image Process. | 3 |
| 2013 | Recognizing Conversational Interaction Based on 3D Human Pose
Jingjing Deng 0001, Xianghua Xie, Ben Daubney, Hui Fang 0003, Phil W. Grant |
ACIVS | 2 |
| 2013 | Interactive Segmentation of Media-Adventitia Border in IVUS
Jonathan-Lee Jones, Ehab Essa, Xianghua Xie |
CAIP (2) | 3 |
| 2013 | From clamped local shape models to global shape modelabstractFacial fiducial point localization is a crucial step for most facial analysis applications, e.g., face recognition, expression recognition and facial aging simulation. Although state-of-art methods have the ability to provide good salient point location on frontal faces, finding a global solution under large variations caused by off-plane rotations and exaggerated expression changes is still a challenge. In this paper, we present a system with a two-level shape model to facilitate accurate facial fiducial point localization. In the first level, two local component models interact with each other in order to offer novel shape constraints. At the same time, the clamped local shape model provides constrained non-linear shape initialization for better convergence performance of the shape model as a whole. The experimental results confirm that the proposed method is capable of dealing with the face alignment under large shape variations. Hui Fang 0003, Jingjing Deng 0001, Xianghua Xie, Phil W. Grant |
ICIP | 3 |
| 2013 | Graph based segmentation with minimal user interactionabstractIn this paper, we present a graph based segmentation method that only requires a single point from user initialization. We incorporate a new image feature into the segmentation scheme. It is derived from a vector field that takes into account gradient vector interactions across the image domain, and has the simplicity of edge based features but also proves to be a useful region indication in two-level segmentation. Effective vector field diffusion is proposed to deal with excessive image noise. Based on a single user point we unravel the image and transfer the object segmentation into a height field segmentation in polar coordinates, which in effect imposes a star shape prior. The search of a minimum closed set on a node weighted, directed graph produces the segmentation result. Comparative analysis on real world images demonstrates promising performances of the proposed method in segmentation accuracy and its simplicity in user interaction. Huaizhong Zhang, Ehab Essa, Xianghua Xie |
ICIP | 3 |
| 2013 | Shape and appearance priors for level set-based left ventricle segmentationabstractThe authors propose a novel spatiotemporal constraint based on shape and appearance and combine it with a level‐set deformable model for left ventricle (LV) segmentation in four‐dimensional gated cardiac SPECT, particularly in the presence of perfusion defects. The model incorporates appearance and shape information into a ‘soft‐to‐hard’ probabilistic constraint, and utilises spatiotemporal regularisation via a maximum a posteriori framework. This constraint force allows more flexibility than the rigid forces of shape constraint‐only schemes, as well as other state of the art joint shape and appearance constraints. The combined model can hypothesise defective LV borders based on prior knowledge. The authors present comparative results to illustrate the improvement gain. A brief defect detection example is finally presented as an application of the proposed method. Ronghua Yang, Majid Mirmehdi, Xianghua Xie |
IET Comput. Vis. | 3 |
| 2012 | Line Histogram - A Fast Method for Rotated Rectangular Area Histogramming
Sion L. Hannuna, Xianghua Xie, Majid Mirmehdi |
ICPRAM (2) | 3 |
| 2012 | Segmentation of Vessel Geometries from Medical Images using GPF Deformable Model
Si Yong Yeo, Xianghua Xie, Igor Sazonov, Perumal Nithiarasu |
ICPRAM (1) | 2 |
| 2012 | State of the Art Report on Video-Based Graphics and Video VisualizationabstractAbstract In recent years, a collection of new techniques which deal with video as input data, emerged in computer graphics and visualization. In this survey, we report the state of the art in video‐based graphics and video visualization. We provide a review of techniques for making photo‐realistic or artistic computer‐generated imagery from videos, as well as methods for creating summary and/or abstract visual representations to reveal important features and events in videos. We provide a new taxonomy to categorize the concepts and techniques in this newly emerged body of knowledge. To support this review, we also give a concise overview of the major advances in automated video analysis, as some techniques in this field (e.g. feature extraction, detection, tracking and so on) have been featured in video‐based modelling and rendering pipelines for graphics and visualization. Rita Borgo, Min Chen 0001, Ben Daubney, Edward Grundy, Gunther Heidemann, Benjamin Höferlin, Markus Höferlin, Heike Leitte, Daniel Weiskopf, Xianghua Xie |
Comput. Graph. Forum | 10 |
| 2012 | Automatic Bootstrapping and Tracking of Object ContoursabstractA new fully automatic object tracking and segmentation framework is proposed. The framework consists of a motion-based bootstrapping algorithm concurrent to a shape-based active contour. The shape-based active contour uses finite shape memory that is automatically and continuously built from both the bootstrap process and the active-contour object tracker. A scheme is proposed to ensure that the finite shape memory is continuously updated but forgets unnecessary information. Two new ways of automatically extracting shape information from image data given a region of interest are also proposed. Results demonstrate that the bootstrapping stage provides important motion and shape information to the object tracker. This information is found to be essential for good (fully automatic) initialization of the active contour. Further results also demonstrate convergence properties of the content of the finite shape memory and similar object tracking performance in comparison with an object tracker with unlimited shape memory. Tests with an active contour using a fixed-shape prior also demonstrate superior performance for the proposed bootstrapped finite-shape-memory framework and similar performance when compared with a recently proposed active contour that uses an alternative online learning model. John Chiverton, Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Image Process. | 2 |
| 2011 | Entropy Driven Hierarchical Search for 3D Human Pose EstimationabstractIn this work a hierarchical approach is presented to efficiently estimate 3D pose from single images.To achieve this the body is represented as a graphical model and optimized stochastically.The use of a graphical representation allows message passing to ensure individual parts are not optimized using only local image information, but from information gathered across the entire model.In contrast to existing methods the posterior distribution is represented parametrically.A different model is used to approximate the conditional distribution between each connected part.This permits measurements of the Entropy, which allows an adaptive sampling scheme to be employed that ensures that parts with the largest uncertainty are allocated a greater proportion of the available resources.At each iteration the estimated pose is updated dependent on the Kullback Leibler (KL) divergence measured between the posterior and the set of samples used to approximate it.This is shown to improve performance and prevent over fitting when small numbers of particles are being used.A quantitative comparison is made using the HumanEva dataset that demonstrates the efficacy of the presented method. Ben Daubney, Xianghua Xie |
BMVC | 2 |
| 2011 | Tracking 3D human pose with large root node uncertaintyabstractRepresenting articulated objects as a graphical model has gained much popularity in recent years, often the root node of the graph describes the global position and orientation of the object. In this work a method is presented to robustly track 3D human pose by permitting greater uncertainty to be modeled over the root node than existing techniques allow. Significantly, this is achieved without increasing the uncertainty of remaining parts of the model. The benefit is that a greater volume of the posterior can be supported making the approach less vulnerable to tracking failure. Given a hypothesis of the root node state a novel method is presented to estimate the posterior over the remaining parts of the body conditioned on this value. All probability distributions are approximated using a single Gaussian allowing inference to be carried out in closed form. A set of deterministically selected sample points are used that allow the posterior to be updated for each part requiring just seven image likelihood evaluations making it extremely efficient. Multiple root node states are supported and propagated using standard sampling techniques. We believe this to be the first work devoted to efficient tracking of human pose whilst modeling large uncertainty in the root node and demonstrate the presented method to be more robust to tracking failures than existing approaches. Ben Daubney, Xianghua Xie |
CVPR | 2 |
| 2011 | Automatic IVUS media-adventitia border extraction using double interface graph cut segmentationabstractWe present a fully automatic segmentation method to extract media-adventitia border in IVUS images. Segmentation in IVUS has shown to be an intricate process due to relatively low contrast and various forms of interferences and artifacts caused by, for example, calcification and acoustic shadow. Graph cut based methods often require careful manual initialization and produces in consistent tracing of the border. We use a double interface automatic graph cut technique to prevent the extraction of media-adventitia border from being distracted by those image features. Novel cost functions are derived from using a combination of complementary texture features. Comparative studies on manual labeled data show promising performance of the proposed method. Ehab Essa, Xianghua Xie, Igor Sazonov, Perumal Nithiarasu |
ICIP | 2 |
| 2011 | Level set segmentation with robust image gradient energy and statistical shape priorabstractWe propose a new level set segmentation method with statistical shape prior using a variational approach. The image energy is derived from a robust image gradient feature. This gives the active contour a global representation of the geometric configuration, making it more robust to image noise, weak edges and initial configurations. Statistical shape information is incorporated using nonparametric shape density distribution, which allows the model to handle relatively large shape variations. Comparative examples using both synthetic and real images show the robustness and efficiency of the proposed method. Si Yong Yeo, Xianghua Xie, Igor Sazonov, Perumal Nithiarasu |
ICIP | 2 |
| 2011 | Radial basis function based level set interpolation and evolution for deformable modelling
Xianghua Xie, Majid Mirmehdi |
Image Vis. Comput. | 1 |
| 2011 | Geometrically Induced Force Interaction for Three-Dimensional Deformable ModelsabstractIn this paper, we propose a novel 3-D deformable model that is based upon a geometrically induced external force field which can be conveniently generalized to arbitrary dimensions. This external force field is based upon hypothesized interactions between the relative geometries of the deformable model and the object boundary characterized by image gradient. The evolution of the deformable model is solved using the level set method so that topological changes are handled automatically. The relative geometrical configurations between the deformable model and the object boundaries contribute to a dynamic vector force field that changes accordingly as the deformable model evolves. The geometrically induced dynamic interaction force has been shown to greatly improve the deformable model performance in acquiring complex geometries and highly concave boundaries, and it gives the deformable model a high invariancy in initialization configurations. The voxel interactions across the whole image domain provide a global view of the object boundary representation, giving the external force a long attraction range. The bidirectionality of the external force field allows the new deformable model to deal with arbitrary cross-boundary initializations, and facilitates the handling of weak edges and broken boundaries. In addition, we show that by enhancing the geometrical interaction field with a nonlocal edge-preserving algorithm, the new deformable model can effectively overcome image noise. We provide a comparative study on the segmentation of various geometries with different topologies from both synthetic and real images, and show that the proposed method achieves significant improvements against existing image gradient techniques. Si Yong Yeo, Xianghua Xie, Igor Sazonov, Perumal Nithiarasu |
IEEE Trans. Image Process. | 2 |
| 2010 | Fast Dynamic Texture Detection
V. Javier Traver, Majid Mirmehdi, Xianghua Xie, Raúl Montoliu |
ECCV (4) | 3 |
| 2010 | Estimating 3D Human Pose from Single Images Using Iterative Refinement of the PriorabstractThis paper proposes a generative method to extract 3D human pose using just a single image. Unlike many existing approaches we assume that accurate foreground background segmentation is not possible and do not use binary silhouettes. A stochastic method is used to search the pose space and the posterior distribution is maximized using Expectation Maximization (EM). It is assumed that some knowledge is known a priori about the position, scale and orientation of the person present and we specifically develop an approach to exploit this. The result is that we can learn a more constrained prior without having to sacrifice its generality to a specific action type. A single prior is learnt using all actions in the Human Eva dataset [9] and we provide quantitative results for images selected across all action categories and subjects, captured from differing viewpoints. Ben Daubney, Xianghua Xie |
ICPR | 2 |
| 2010 | Level Set Based Segmentation Using Local Feature DistributionabstractWe propose a level set based framework to segment textured images. The snake deforms in the image domain in searching for object boundaries by minimizing an energy functional, which is defined based on dynamically selected local distribution of orientation invariant features. We also explore the user initialization to simplify the segmentation and improve accuracy. Experimental results on both synthetic and real data show significant improvements compared to direct modeling of filtering responses or piecewise constant modeling. Xianghua Xie |
ICPR | 1 |
| 2010 | Initialisation-Free Active Contour SegmentationabstractWe present a region based active contour model which does not require any initialisation and is capable of modelling multi-modal image regions. Its external force is based on statistically learning and grouping of image primitives in multiscale, and its numerical solution is carried out using radial basis function interpolation and time dependent expansion coefficient updating. The initialisation-free property makes it attractive to applications such as detecting unkown number of objects with unkown topologies. Xianghua Xie, Majid Mirmehdi |
ICPR | 1 |
| 2010 | Active Contouring Based on Gradient Vector Interaction and Constrained Level Set DiffusionabstractThis paper presents an extension of our recently introduced MAC model to deal with the initialization dependency problem that commonly appears in edge-based approaches. Its dynamic force field, unique bidirectionality, and constrained diffusion-based level set evolution provide great freedom in contour initialization and show significant improvements in initialization independency compared to other edge-based techniques. It can handle more sophisticated topological changes than splitting and merging. It provides new potentials for edge-based active contour methods, particularly when detecting and localizing objects with unknown location, geometry, and topology. Xianghua Xie |
IEEE Trans. Image Process. | 1 |
| 2009 | On-line Learning of Shape Information for Object Segmentation and TrackingabstractWe present segmentation and tracking of deformable objects using non-linear on-line learning of high-level shape information in the form of a level set function. The emphasis is for successful tracking of objects that undergo smooth arbitrary deformations, but without the a priori learning of shape constraints. The high-level shape information is learnt on-line by defining a memory of object samples in a high-dimensional shape space. These shape samples are then used as weights via a locally defined shape space kernel function to define a template against which potential future shapes of the tracked object can be compared. Results for the successful tracking of a range of deformable motions are presented. 1 John Chiverton, Majid Mirmehdi, Xianghua Xie |
BMVC | 3 |
| 2009 | Geometric Potential Force for the Deformable ModelabstractWe propose a new external force field for deformable models which can be conve-niently generalized to high dimensions. The external force field is based on hypothesized interactions between the relative geometries of the deformable model and image gradi-ents. The evolution of the deformable model is solved using the level set method. The dynamic interaction forces between the geometries can greatly improve the deformable model performance in acquiring complex geometries and highly concave boundaries, and in dealing with weak image edges. The new deformable model can handle arbi-trary cross-boundary initializations. Here, we show that the proposed method achieve significant improvements when compared against existing state-of-the-art techniques. 1 Si Yong Yeo, Xianghua Xie, Igor Sazonov, Perumal Nithiarasu |
BMVC | 2 |
| 2008 | Tracking with Active Contours Using Dynamically Updated Shape InformationabstractAn active contour based tracking framework is described that generates and integrates dynamic shape information without having to learn a priori shape constraints. This dynamic shape information is combined with fixative pho-tometric foreground model matching and background mismatching. Bound-ary based optical flow is also used to estimate the location of the object in each new video frame, incorporating Procrustes based shape alignment. Promising results under complex deformations of shape, varied levels of noise, and close-to-complete occlusion in the presence of complex textured backgrounds are presented. 1 John Chiverton, Xianghua Xie, Majid Mirmehdi |
BMVC | 2 |
| 2008 | Variational Maximum A Posteriori model similarity and dissimilarity matchingabstractA new variational Maximum A Posteriori (MAP) contextual modeling approach is presented that minimizes the product of two ratios: (a) the ratio of the model distribution to the distribution of currently estimated foreground pixels; (b) the ratio of the background distribution to the model distribution for all estimated background pixels. This approach provides robust discrimination to identify the division between foreground and background pixels, which is useful for applications such as object tracking. John Chiverton, Majid Mirmehdi, Xianghua Xie |
ICPR | 3 |
| 2008 | MAC: Magnetostatic Active Contour ModelabstractWe propose an active contour model using an external force field that is based on magnetostatics and hypothesized magnetic interactions between the active contour and object boundaries. The major contribution of the method is that the interaction of its forces can greatly improve the active contour in capturing complex geometries and dealing with difficult initializations, weak edges and broken boundaries. The proposed method is shown to achieve significant improvements when compared against six well-known and state-of-the-art shape recovery methods, including the geodesic snake, the generalized version of GVF snake, the combined geodesic and GVF snake, and the charged particle model. Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Correction to "MAC: Magnetostatic Active Contour Model"abstractIn the above titled paper (ibid., vol. 30, no. 4, pp. 632-646, Apr 08), there was an error in a definition. The correct definition is presented here. Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Implicit Active Model using Radial Basis Function Interpolated Level SetsabstractBuilding on recent work by others that introduced RBFs into level sets for structural topology optimisation, we introduce the concept into active models and present a new level set formulation able to handle more complex topological changes, in particular perturbation away from the evolving front. This allows the initial contour or surface to be placed arbitrarily in the image. The proposed level set updating scheme is efficient and does not suffer from self-flattening while evolving which will cause large numerical error. Unlike conventional level set based active models, periodic re-initialisation is also no longer necessary and the computational grid can be much coarser, thus, it has great potential in modelling in high dimensional space. We show results on synthetic and real data for active modelling in 2D and 3D. 1 Xianghua Xie, Majid Mirmehdi |
BMVC | 1 |
| 2007 | TEXEMS: Texture Exemplars for Defect Detection on Random Textured SurfacesabstractWe present an approach to detecting and localizing defects in random color textures which requires only a few defect free samples for unsupervised training. It is assumed that each image is generated by a superposition of various-size image patches with added variations at each pixel position. These image patches and their corresponding variances are referred to here as textural exemplars or texems. Mixture models are applied to obtain the texems using multiscale analysis to reduce the computational costs. Novelty detection on color texture surfaces is performed by examining the same-source similarity based on the data likelihood in multiscale, followed by logical processes to combine the defect candidates to localize defects. The proposed method is compared against a Gabor filter bank-based novelty detection method. Also, we compare different texem generalization schemes for defect detection in terms of accuracy and efficiency. Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | A Charged Active Contour Based on Electrostatics
Ronghua Yang, Majid Mirmehdi, Xianghua Xie |
ACIVS | 3 |
| 2006 | Magnetostatic Field for the Active Contour Model: A Study in ConvergenceabstractA new external velocity field for active contours is proposed. The velocity field is based on magnetostatics and hypothesised magnetic interactions be-tween the active contour and image gradients. In this paper, we introduce the method and study its convergence capability for the recovery of shapes with complex topology and geometry, including deep, narrow concavities. The proposed active contour can be arbitrarily initialised. Level sets are used to achieve topological freedom. The proposed method is compared against shape recovery methods based on distance vector flow, constant flow, gener-alised version of GVF, geodesic GGVF, and curvature vector flow. 1 Xianghua Xie, Majid Mirmehdi |
BMVC | 1 |
| 2006 | Colour tonality inspection using eigenspace features
Xianghua Xie, Majid Mirmehdi, Barry T. Thomas |
Mach. Vis. Appl. | 1 |
| 2005 | Localising surface defects in random colour textures using multiscale texem analysis in image eigenchannelsabstractA novel method is presented to detect defects in random colour textures which requires only a very few normal samples for unsupervised training. We decorrelate the colour image by generating three eigenchannels in each of which the surface texture image is divided into overlapping patches of various sizes. Then, a mixture model and EM is applied to reduce groupings of patches to a small number of textural exemplars, or texems. Localised defect detection is achieved by comparing the learned texems to patches in the unseen image eigenchannels. Xianghua Xie, Majid Mirmehdi |
ICIP (3) | 1 |
| 2004 | RAGS: region-aided geometric snakeabstractAn enhanced, region-aided, geometric active contour that is more tolerant toward weak edges and noise in images is introduced. The proposed method integrates gradient flow forces with region constraints, composed of image region vector flow forces obtained through the diffusion of the region segmentation map. We refer to this as the Region-aided Geometric Snake or RAGS. The diffused region forces can be generated from any reliable region segmentation technique, greylevel or color. This extra region force gives the snake a global complementary view of the boundary information within the image which, along with the local gradient flow, helps detect fuzzy boundaries and overcome noisy regions. The partial differential equation (PDE) resulting from this integration of image gradient flow and diffused region flow is implemented using a level set approach. We present various examples and also evaluate and compare the performance of RAGS on weak boundaries and noisy images. Xianghua Xie, Majid Mirmehdi |
IEEE Trans. Image Process. | 1 |
| 2003 | Geodesic Colour Active Contour Resistent to Weak Edges and NoiseabstractThe standard geometric or geodesic active contour is a powerful segmentation method, yet it is susceptible to weak edges and image noise. We propose a new region-aided, geometric, colour active contour that integrates gradient flow forces with region con-straints. These constraints are composed of image region vector flow forces obtained through the diffusion of the region segmentation map. The extra region force gives the snake a global view of the boundary information within the image which, along with the local gradient flow, helps detect fuzzy boundaries and overcome noisy re-gions. The partial differential equation (PDE) resulting from this integration of image gradient flow and diffused region flow is implemented using the level set approach. 1 Xianghua Xie, Majid Mirmehdi |
BMVC | 1 |
| 2003 | Level-set based geometric colour snake with region supportabstractA novel method is introduced to force a geometric-based snake be more tolerant towards weak edges and noise in images. The method integrates gradient flow forces with region constraints obtained from diffused region segmentation forces. The diffusion is obtained from the region map vector flow field. This extra region force gives the snake a global view of the boundary information within the image. We present results on both graylevel and colour images. Xianghua Xie, Majid Mirmehdi |
ICIP (2) | 1 |