EDBT 2026 Demo / reviewers in the wild / expert
Xinghui Dong
dblp:66/4043
· DBLP profile ↗
47ranked-venue papers
14as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 8 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly DetectionabstractIndustrial anomaly detection is a critical component of modern manufacturing, yet the scarcity of defective samples restricts traditional detection methods to scenario-specific applications. Although Vision-Language Models (VLMs) demonstrate significant advantages in generalization capabilities, their performance in industrial anomaly detection remains limited. To address this challenge, we propose IAD-R1, a universal post-training framework applicable to VLMs of different architectures and parameter scales, which substantially enhances their anomaly detection capabilities. IAD-R1 employs a two-stage training strategy: the Perception Activation Supervised Fine-Tuning (PA-SFT) stage utilizes a meticulously constructed high-quality Chain-of-Thought dataset (Expert-AD) for training, enhancing anomaly perception capabilities and establishing reasoning-to-answer correlations; the Structured Control Group Relative Policy Optimization (SC-GRPO) stage employs carefully designed reward functions to achieve a capability leap from "Anomaly Perception" to "Anomaly Interpretation". Experimental results demonstrate that IAD-R1 achieves significant improvements across 7 VLMs, the largest improvement was on the DAGM dataset, with average accuracy 43.3% higher than the 0.5B baseline. Notably, the 0.5B parameter model trained with IAD-R1 surpasses commercial models including GPT-4.1 and Claude-Sonnet-4 in zero-shot settings, demonstrating the effectiveness and superiority of IAD-R1. Yunkang Cao, Chengliang Liu 0003, Yuan Xiong, Xinghui Dong, Chao Huang 0008 |
AAAI | 5 |
| 2026 | A Novel IoT-Based Spatiotemporal Prediction and Resilience Optimization Method for Tunnel-Induced Ground SettlementabstractWith the acceleration of urbanisation, the problem of ground settlement in tunnel construction has become a serious challenge. Aiming at the existing ground settlement monitoring and resilience management problems such as limited monitoring range, neglected spatio-temporal characteristics, and poor combination of prediction results and resilience management, this paper proposes a two-stage spatio-temporal prediction and resilience optimization for tunnel-induced ground settlement. Specifically, firstly, this paper constructs a comprehensive monitoring architecture integrating SBAS-InSAR technology and Internet of Things (IoT) technology, which realises accurate and comprehensive monitoring of ground settlement. Secondly, a two-stage spatio-temporal prediction method of ground settlement is proposed: based on the wide-area spatio-temporal data acquired by SBAS-InSAR technology, the spatio-temporal transformer model is used to make the preliminary prediction. Then, with the small-area variables collected by IoT, the parameter-seeking optimisation algorithm based on the Grid Search-Particle Swarm is used in conjunction with the Time Convolutional-Bidirectional Long and Short-Term Memory Network (GR-PSO-TCN-BiLSTM) model to correct the prediction error. In terms of resilience optimisation, this paper proposes a multi-stage resilience enhancement strategy based on ground settlement prediction, which combines prevention importance, degradation importance and recovery importance, aiming to maximise the resilience of tunnel-induced ground settlement area. Finally, an empirical analysis using the traffic along the Zhengzhou Metro as an example verifies the effectiveness of the proposed method. These results indicate that coupling wide-area remote sensing with local IoT correction can substantially improve settlement prediction accuracy and provide actionable guidance for maintenance prioritization, thereby enhancing the robustness and recovery capability of metro systems. Xinghui Dong, Jichao Li 0001, Huanqi Zhang, Ke-Wei Yang 0001, Hongyan Dui |
IEEE Internet Things J. | 1 |
| 2026 | SAM-LLaVA: A segmentation-aware vision-language framework for industrial defect diagnosis
Shengwang An, Chengjia Wang, Xinghui Dong |
Pattern Recognit. | 3 |
| 2026 | Boundary-aware shape recognition using dynamic graph convolutional networks
Jinming Zhao, Junyu Dong, Huiyu Zhou 0001, Xinghui Dong |
Pattern Recognit. | 4 |
| 2026 | UMDM-USG: A unified multi-view diffusion model for underwater scene generation via cross-view representation alignment
Chengjia Wang, Xinghui Dong |
Pattern Recognit. | 3 |
| 2026 | SPGDD-GPT: Image-Text-Driven Generic Defect Diagnosis Using a Self-Prompted Large Vision-Language ModelabstractLarge Vision-Language Models (LVLMs) mainly rely on template-generated textual descriptions to understand defects. This reliance impairs the performance of these models for Industrial Defect Detection (IDD) because they typically lack specialized knowledge. On the other hand, the majority of existing IDD methods only utilize the contrastive loss function for image-to-text feature alignment, which limits their ability to focus on defective regions. In addition, these methods usually use cosine similarity for contextual learning, which also restricts their ability to understand and adapt to complex contexts. To address these issues, we first collect a large-scale defect data set with textual descriptions, namely, the Text-Augmented Defect Data Set (TADD), to fine-tune an LVLM for defect description. We also propose a Self-prompted Generic Defect Diagnosis (including Defect Detection and Defect Description) LVLM, i.e., the SPGDD-GPT. This method can effectively utilize contextual information through a Multi-scale Self-prompted Memory Module (MSSPMM) and a Text-Driven Defect Focuser (TDDF) that we deliberately design, to adapt to unseen defect categories and focus on abnormal regions. Experimental results show that our method normally achieves the better performance than its counterparts across the 21 subsets of TADD under the 1-shot, 2-shot and 4-shot defect detection settings, demonstrating strong detection and generalization capabilities1. The proposed method can also generate a textural description of the defects contained in each test image. These promising results should be due to the proposed MSSPMM and TDDF and the large-scale TADD. Shengwang An, Xinghui Dong |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | DefectSynth: Few-Shot Defective Image Generation by Modeling Shape and AppearanceabstractSince the acquisition of a large amount of defect data is expensive and time-consuming, the limited availability of defective samples impairs the accuracy and generalizability of defect detection methods. Although defect generation approaches have been explored for data augmentation, they usually suffer from either a lack of realism or limited diversity. To address these challenges, we propose a two-stage controllable few-shot1defective image generation network, namely, DefectSynth2, which models both the shape and appearance of defects. The first stage aims to generate continuous defect masks even with a few real masks. To this end, we propose a Hybrid Mask Interpolation (HMI) module, which performs interpolation in the image or latent space. The second stage is used to synthesize the defective appearance. We first fine-tune the pre-trained ControlNet, which is then used together with a pre-trained stable diffusion model to synthesize defective images. Given a mask generated in the first stage and a text prompt, they are integrated with a defect-free image to synthesize a high-fidelity defective image. To alleviate the issue of generation of indistinct defects with existing methods, we propose a Selective Attention Enhancement (SAE) mechanism that highlights the details of defects. We also design a Similarity-Based Feature Fusion (SFF) module to merge different local features, thereby further enriching the appearance diversity of defects. Using the defect data generated by DefectSynth, the classification accuracy on MVTec-AD has been improved from 44.03% to 66.70% compared with the baseline without synthetic data augmentation, while the F1-Score values computed on GDXray and DeepCrack for small-defect localization have been increased from 70.48% to 76.64% and from 70.08% to 83.27%, respectively. These performance gains should be due to the fact that our method is able to generate realistic and diverse defective images by modeling both the shape and appearance of defects. Dexu Zhao, Xukun Qin, Xinghui Dong |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | TPCM-SegNet: A Text-Prompted Dual-Path Convolution-Mamba Network for Anomaly SegmentationabstractAnomaly segmentation has been widely applied to diagnosis of medical organs and lesions and detection of industrial defects. However, existing methods still face challenges in extracting discriminant image features and utilizing semantic information. To address these issues, we propose a Text-Prompted Dual-Path Convolution-Mamba Network (TPCM-SegNet)1, which integrates Residual Double-Convolution Blocks (RDCBs) and Mamba-Transformer Blocks (MTBs) in two parallel paths for the purpose of extracting local and global features, respectively. Given a pair of RDCB and MTB at the same stage, a Feature Fusion Block (FFB) is introduced in order to facilitate the interaction and fusion of the features extracted using these blocks. Furthermore, we fuse the text tokens extracted from a textual description with the image features extracted using each of those blocks through a Text Prompt Block (TPB), to enhance the semantics understanding ability of the network. A Cascade Feature Block (CFB) is also designed for each stage of the encoder, to combine the feature maps, the logit maps decoded from them and the input image. This block incorporates the prior and original characteristics into the image representation. Experimental results demonstrate that our TPCM-SegNet achieves the superior, or at least comparable, performance to baselines, across eight publicly available datasets. These promising results should benefit from the powerful ability of image representation and semantic understanding of the proposed network. Borong Xu, Junyu Dong, Xinghui Dong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Population-Based Meta-Heuristic Optimization Algorithm Booster: An Evolutionary and Learning Competition Scheme
Jun Wang 0041, Junyu Dong, Huiyu Zhou 0001, Xinghui Dong |
Neurocomputing | 4 |
| 2025 | IoUT-Enhanced Cooperative Control Scheme for Multiple AUVs With IoT Data ReliabilityabstractMultiple autonomous underwater vehicles (AUVs) are being developed to survey marine ecosystems. The main challenge lies in that most of the research focuses on the physical layer of AUVs, but less on the data layer which could contribute significantly to the performance of AUVs. Moreover, there is a lack of shock prevention and post-shock recovery strategies for AUVs. Therefore, to comprehensively analyze the multi-level performance changes of the multiple AUVs, this paper models the physical layer and data layer of the AUVs within the Internet of Underwater Things (IoUT), and the architecture of the corresponding system components based on digital twin technology. After that, we model the IoT data reliability in AUVS. Then, the physical performance evaluation is carried out for different phases of the performance change process in AUVs. To enhance the ability of the multiple AUVs to resist external interference, a multi-stage control scheme is proposed. At last, compared with the generalized scheme, a case simulation shows that the proposed scheme can maximize the protection against external interference. The control scheme leads multiple AUVs to a 6.13% improvement in performance efficiency, the data reliability of 49.66% and 10.95% cost savings in multiple AUVs. Hongyan Dui, Huanqi Zhang, Songru Zhang, Xinghui Dong |
IEEE Internet Things J. | 4 |
| 2025 | Drone-Based Wall Crack Detection Using Model-Agnostic Meta-LearningabstractWith the urbanization process and aging of buildings, wall crack detection plays a crucial role in the maintenance and safety of building structures. Due to the inherent characteristics of defects, however, cracks in the wall are relatively sparse, compared with the normal area. High-altitude regions are also difficult to access. As a result, large wall crack data sets are rare. This issue impairs the training of deep networks and may lead to the suboptimal detection result. To address this issue, we first capture a set of pure wall crack images using a drone, which are comprised of a new wall crack data set, namely, Ocean University of China Wall Crack Data Set (OUC-Crack). In contrast to existing crack data sets, OUC-Crack only contains wall crack images and hence it is particularly useful for wall crack detection. Then we propose a drone-based wall crack detection system, which consists of a drone platform and a crack detection network referred to as the Model-Agnostic Meta-Learning Based Segmentation Network (MAML-SegNet). Therefore, only a small number of training images are required. To fulfill the real-time detection task on drones, we further develop an efficient MAML-SegNet (EFF-MAML-SegNet). The proposed systems are tested on the OUC-Crack and two publicly available crack data sets, including Volker and Crack500. Experimental results show that our MAML-SegNet outperforms 14 baselines in terms of Dice Coefficient, with improvements of at least 1.26%, 13.19% and 7.12% on the three data sets, respectively. Given that the EFF-MAML-SegNet is deployed on a drone, real-time detection can be achieved with the Dice Coefficient values of 66.83%, 55.42% and 48.03% on the three data sets, respectively. These promising results should be due to the few-shot learning ability of meta-learning and the efficient network design. (Code, data and models are available at https://indtlab.github.io/projects/Wall-Crack-Detection). Borong Xu, Wenxuan Shao, Xinghui Dong |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Terrain Scene Generation Using a Lightweight Vector Quantized Generative Adversarial NetworkabstractNatural terrain scene images play important roles in the geographical research and application. However, it is challenging to collect a large set of terrain scene images. Recently, great progress has been made in image generation. Although impressive results can be achieved, the efficiency of the state-of-the-art methods, e.g., the Vector Quantized Generative Adversarial Network (VQGAN), is still dissatisfying. The VQGAN confronts two issues, i.e., high space complexity and heavy computational demand. To efficiently fulfill the terrain scene generation task, we first collect a Natural Terrain Scene Data Set (NTSD), which contains 36,672 images divided into 38 classes. Then we propose a Lightweight VQGAN (Lit-VQGAN), which uses the fewer parameters and has the lower computational complexity, compared with the VQGAN. A lightweight super-resolution network is further adopted, to speedily derive a high-resolution image from the image that the Lit-VQGAN generates. The Lit-VQGAN can be trained and tested on the NTSD. To our knowledge, either the NTSD or the Lit-VQGAN has not been exploited before.1Experimental results show that the Lit-VQGAN is more efficient and effective than the VQGAN for the image generation task. These promising results should be due to the lightweight yet effective networks that we design. Huiyu Zhou 0001, Xinghui Dong |
IEEE Trans. Big Data | 3 |
| 2025 | Perception-Aware Underwater Image Quality Assessment: Dataset, Perceptual Quality Scores, and Assessment NetworkabstractUnderwater Image Quality Assessment (UIQA) plays an important role in assess the effectiveness of Underwater Image Enhancement (UIE) algorithms or to evaluate the quality of underwater images. However, accurate UIQA that are consistent with human perception remains challenging. This dilemma on one hand is attributed to the lack of real human visual perception UIQA data, and on the other hand that the quality feature representation used by existing UIQA algorithms are inconsistent with human perceptions. To address these issues, we introduce a Large scale Underwater Image Quality Dataset (LUIQD), and propose an UIQA network named as Perception-Aware Underwater image Quality Assessment Network (PAUQA-Net). Specifically, the LUIQD includes 64,180 real and enhance underwater images covering a wide range of scenes, target and imaging conditions, with their perceptual quality scores. Based on the analysis of the mechanisms of human perception, we further design the data-driven PAUQA-Net that integrates an efficient convolutional attention vision Transformer to extract multi-scale features by a multi-path structure. Considering the specificity of human perception of underwater images, color and sharpness features from the chrominance and luminance domains are extracted and fused with local and global images features for joint feature interaction. Extensive experiments conduted on LUIQD and other datasets demonstrate that the proposed PAUQA-Net achieves superior assessment performance compared with the most popular UIQA and IQA methods. The code and dataset can be found in https://github.com/CatchACat083/PAUQA. Bosen Lin, Junyu Dong, Xinghui Dong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | TG-TSGNet: A Text-Guided Arbitrary-Resolution Terrain Scene Generation NetworkabstractWith the increasing demand for terrain visualization in many fields, such as augmented reality, virtual reality and geographic mapping, traditional terrain scene modeling methods encounter great challenges in processing efficiency, content realism and semantic consistency. To address these challenges, we propose a Text-Guided Arbitrary-Resolution Terrain Scene Generation Network (TG-TSGNet), which contains a ConvMamba-VQGAN, a Text Guidance Sub-network and an Arbitrary-Resolution Image Super-Resolution Module (ARSRM). The ConvMamba-VQGAN is built on top of the Conv-Based Local Representation Block (CLRB) and the Mamba-Based Global Representation Block (MGRB) that we design, to utilize local and global features. Furthermore, the Text Guidance Sub-network comprises a text encoder and a Text-Image Alignment Module (TIAM) for the sake of incorporating textual semantics into image representation. In addition, the ARSRM can be trained together with the ConvMamba-VQGAN, to perform the task of image super-resolution. To fulfill the text-guided terrain scene generation task, we derive a set of textual descriptions for the 36,672 images across the 38 categories of the Natural Terrain Scene Data Set (NTSD). These descriptions can be used to train and test the TG-TSGNet (The data set, model and source code are available at https://github.com/INDTLab/TG-TSGNet). Experimental results show that the TG-TSGNet outperforms, or at least performs comparably to, the baseline methods in image realism and semantic consistency with proper efficiency. We believe that the promising performance should be due to the ability of the TG-TSGNet not only to capture both the local and global characteristics and the semantics of terrain scenes, but also to reduce the computational cost of image generation. Xinghui Dong |
IEEE Trans. Image Process. | 3 |
| 2025 | Multi-Stage Control Strategy of IoT-Enabled Unmanned Vehicle Detection SystemsabstractAs the environment deteriorates, natural disasters occur more frequently and become more devastating to human beings and the environment. After a disaster, to quickly and optimally restore the damaged things, including physical systems (e.g., transport networks) and the environment, needs decision makers to own sufficient data/information. Unmanned vehicle detection systems (UVDS) are undoubtedly feasible tools in collecting such data in a harsh environment. The most important challenges in UVDS management are on modeling the UVDS data layer and multi-stage recovery strategies, which have received little research. To address such problems, this paper proposes a multi-stage control strategy for UVDS based on Internet of Things (IoT). The optimal decision is decided by utilizing four indicators: performance recovery efficiency, normal detection probability, operation cost, and economic benefit cost, respectively. The simulation results show that the proposed strategy improves the performance recovery efficiency by 12.1% and the normal detection probability by 3.9%, the operation cost declines by 58.4%, and the economic benefit cost by 75.9% compared with the general control strategy. Hongyan Dui, Huanqi Zhang, Xinghui Dong, Shaomin Wu, Yu Wang 0291 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | IoT-Enabled Risk Warning and Maintenance Strategy Optimization for Tunnel-Induced Ground SettlementabstractWith the continuous development of underground space, the vigorous development of underground rail transport has become an effective way to relieve the pressure of urban traffic. However, ground settlement caused by tunnels can lead to cracks, settlements, and collapses of nearby buildings, resulting in serious economic losses. To solve the problems of poor generalization performance of risk warning models and uneven allocation of maintenance resources in PHM for ground settlement in existing studies, an Internet of Things (IoT)-enabled risk warning and maintenance strategy optimization method is proposed in this paper. In risk warning, firstly, a weight optimization method with the decision objectives of variance maximization, correlation minimization, and estimation error minimization is introduced to find the optimal weights of the base learners in ensemble learning prediction. Secondly, a one-dimensional convolutional neural network-bidirectional long and short-term memory network (1D CNN-BiLSTM) is used to make further predictions on the prediction residuals. In maintenance strategy optimization, threshold-based optimization and cost-based risk-importance maintenance strategies are proposed based on the risk warning results of ground settlement. To test the enhanced effectiveness of the proposed method, a set of comprehensive simulations is carried out in Ningbo city rail transit as an example. The results show that the proposed risk warning method has smaller MAE, MAPE, and RMSE compared to other baseline methods. In addition, the proposed maintenance strategy reduces 12.5%, 16.3%, and 85.7% in terms of cost compared to the baseline method. Overall, the simulation results confirm the advantages of the proposed framework for IoT-enabled risk warning and maintenance strategy optimization in PHM. Hongyan Dui, Xinghui Dong, Xinmin Wu, Guanghan Bai |
IEEE Internet Things J. | 2 |
| 2024 | IoT-Enabled Real-Time Traffic Monitoring and Control Management for Intelligent Transportation SystemsabstractAdvanced Internet of Things (IoT) technology has a profound impact on improving the intelligence level of intelligent transportation systems (ITS) and promoting the sustainable development of urban transportation. However, how to use IoT to process traffic flow and make ITS develop towards automation and global control is still a challenge. Against this backdrop, a prospective traffic controlling model is proposed for ITS based on IoT to enhance the awareness of roads and the responsiveness of transportation system. When traffic congestion events occur, ITS can provide the optimal control strategy of vehicle-to-everything supported vehicles (V2X-supported vehicles) from a macro perspective to control the traffic flow globally and improve traffic efficiency. Specially, the optimal control strategies consider the potential congested road segments caused by congestion propagation. Meanwhile, this paper explores the impact of route choice behavior of V2X-supported vehicles on system performance. The simulation results show the optimal control strategies can alleviate congestion effectively and improve transportation system performance significantly by controlling vehicles. Hongyan Dui, Songru Zhang, Meng Liu 0019, Xinghui Dong, Guanghan Bai |
IEEE Internet Things J. | 4 |
| 2024 | Edge-guided oceanic scene element detection
Keke Xiang, Xingshuai Dong, Weibo Wang 0006, Xinghui Dong |
Knowl. Based Syst. | 4 |
| 2024 | WRD-Net: Water Reflection Detection using a parallel attention transformer
Huijie Dong, Hao Qi 0009, Huiyu Zhou 0001, Junyu Dong, Xinghui Dong |
Pattern Recognit. | 5 |
| 2024 | A learnable support selection scheme for boosting few-shot segmentation
Wenxuan Shao, Xinghui Dong |
Pattern Recognit. | 3 |
| 2024 | Small Sample Image Segmentation by Coupling Convolutions and TransformersabstractCompared with natural image segmentation, small sample image segmentation tasks, such as medical image segmentation and defect detection, have been less studied. Recent studies made efforts on bringing together Convolutional Neural Networks (CNNs) and Transformers in a serial or interleaved architecture in order to incorporate long-range dependencies into the features extracted using CNNs. In this study, we argue that these architectures limit the capability of the combination of CNNs and Transformers. To this end, we propose a dual-stream small sample image segmentation network, namely, the Interactive Coupling of Convolutions and Transformers Based UNet (ICCT-UNet)1, motivated by the success achieved using the UNet in the scenario of small sample image segmentation. Within this network, a CNN stream is paralleled with a Transformer stream while maintaining feature exchange inside each block through the proposed Window-Based Multi-head Cross-Attention (W-MHCA) mechanism. To derive an overall segmentation, the features learned by both the streams are further fused using a Residual Fusion Module (RFM). Experimental results show that the ICCT-UNet outperforms, or at least performs comparably to, its counterparts on eight sets of medical and defective images. These promising results should be attributed to the effective combination of the local and global features fulfilled by the proposed interactive coupling method. Hao Qi 0009, Huiyu Zhou 0001, Junyu Dong, Xinghui Dong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Deep Color-Corrected Multiscale Retinex Network for Underwater Image EnhancementabstractThe acquisition of high-quality underwater images is of great importance to ocean exploration activities. However, images captured in the underwater environment often suffer from degradation due to complex imaging conditions, leading to various issues, such as color cast, low contrast, and low visibility. Although many traditional methods have been used to address these issues, they usually lack robustness in diverse underwater scenes. On the other hand, deep learning techniques struggle to generalize to unseen images, due to the challenge of learning the complicated degradation process. Inspired by the success achieved by the Retinex-based methods, we decompose the underwater image enhancement (UIE) task into two consecutive procedures, including color correction and visibility enhancement, and introduce a novel deep color-corrected multiscale retinex network (CCMSR-Net) (code and models are available athttps://indtlab.github.io/projects/CCMSRNet). With regard to the two procedures, this network comprises a color correction subnetwork (CC-Net) and an MSR subnetwork (MSR-Net), which are built on top of the hybrid convolution–axial attention block (HCAAB) that we design. Thanks to this block, the CCMSR-Net is able to efficiently capture local characteristics and the global context. Experimental results show that the CCMSR-Net outperforms, or at least performs comparably to, 11 baselines across five test sets. We believe that these promising results are due to the effective combination of color correction methods and the MSR model, achieved by jointly exploiting convolutional neural networks (CNNs) and transformers. Hao Qi 0009, Huiyu Zhou 0001, Junyu Dong, Xinghui Dong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Perception-Aware Texture Similarity PredictionabstractTexture similarity plays important roles in texture analysis and material recognition. However, perceptually-consistent fine-grained texture similarity prediction is still challenging. The discrepancy between the texture similarity data obtained using algorithms and human visual perception has been demonstrated. This dilemma is normally attributed to the texture representation and similarity metric utilised by the algorithms, which are inconsistent with human perception. To address this challenge, we introduce a Perception-Aware Texture Similarity Prediction Network (PATSP-Net). This network comprises a Bilinear Lateral Attention Transformer network (BiLAViT) and a novel loss function, namely, RSLoss. The BiLAViT contains a Siamese Feature Extraction Subnetwork (SFEN) and a Metric Learning Subnetwork (MLN), designed on top of the mechanisms of human perception. On the other hand, the RSLoss measures both the ranking and the scaling differences. To our knowledge, either the BiLAViT or the RSLoss has not been explored for texture similarity tasks. The PATSP-Net performs better than, or at least comparably to, its counterparts on three data sets for different fine-grained texture similarity prediction tasks. We believe that this promising result should be due to the joint utilization of the BiLAViT and RSLoss, which is able to learn the perception-aware texture representation and similarity metric. Weibo Wang 0006, Xinghui Dong |
IEEE Trans. Image Process. | 2 |
| 2023 | IoT-Enabled Fault Prediction and Maintenance for Smart Charging PilesabstractWith the application of the Internet of Things (IoT), smart charging piles, which are important facilities for new energy electric vehicles (NEVs), have become an important part of the smart grid. Since the smart charging piles are generally deployed in complex environments and prone to failure, it is significant to perform efficient fault diagnosis and timely maintenance for them. One of the key problems to be solved is how to conduct fault prediction based on limited data collected through IoT in the early stage and develop reasonable preventive maintenance strategies. In this article, a real-time fault prediction method combining cost-sensitive logistic regression (CS-LR) and cost-sensitive support vector machine classification (CS-SVM) is proposed. CS-LR is first used to classify the fault data of smart charging piles, then the CS-SVM is adopted to predict the faults based on the classified data. The feasibility of the proposed model is illustrated through the case study on fault prediction of real-world smart charging piles. To demonstrate the advantage in prediction accuracy, the proposed fault prediction model is compared with the classic baseline models, such as LR, SVM, decision tree (DT),$K$-nearest neighbor (KNN), and backpropagation neural network (BPNN). Finally, based on the proposed fault prediction method, preventive maintenance based on a probability threshold with the minimum total expected cost is proposed. Simulation results show that the proposed maintenance strategy has a better performance in reducing the total maintenance cost compared with traditional periodic maintenance. This is valuable for the development of preventive maintenance strategies for repairable systems under early real-time monitoring data. Hongyan Dui, Xinghui Dong |
IEEE Internet Things J. | 2 |
| 2022 | Unifying the Visual Perception of Humans and Machines on Fine-Grained Texture Similarity
Weibo Wang 0006, Xinghui Dong |
BMVC | 2 |
| 2022 | Defect Classification and Detection Using a Multitask Deep One-Class CNNabstractDefect classification and detection have been explored using convolutional neural networks (CNNs). Normally, a large set of training images containing defects and the associated annotation data are required by these approaches. However, such a large set of images is usually difficult to collect because defects are rare and annotation is time-consuming and expensive. To address these issues, we propose to use a multitask deep one-class CNN for defect classification. Compared with supervised classification methods, this CNN does not require abnormal images and annotated data for training. Specifically, we build a stacked encoder–decoder autoencoder for learning feature representation from normal images. The encoder is used as a feature extractor based on the hard sharing scheme of multitask learning. A one-class classification (OCC) objective learned as a hypersphere using minimum volume estimation is appended to it. Together the encoder and the OCC objective lead to a deep one-class classifier. To train both the autoencoder and one-class classifier end-to-end, a multitask loss function is built. Given an unknown sample, the distance between its feature representation and the center of the hypersphere is used as the anomaly score. Furthermore, defect detection is implemented using a moving-window scanning method on top of the deep one-class classifier. The proposed approach achieves better performance than its counterparts trained using a two-stage method. For defect detection, our approach achieves results almost as good as the supervised method even without using any annotated data. We attribute the promising results to the advantages of multitask learning.Note to Practitioners—Building and evaluating vision-based nondestructive testing (NDT) techniques require many examples of abnormal images, which may not be easy to acquire. This article describes a method that does not require abnormal images for training a convolutional neural network (CNN) in order to perform one-class defect classification (outlier detection). We also applied the method to defect detection with promising results. We include results of experiments demonstrating that better performance can be obtained using our method compared to a set of baselines. Although the proposed method does not use abnormal images for training, it still produces results that are almost as good as the supervised learning-based CNN approaches. This study provides a solution to the challenge encountered by the industrial inspection community when enough abnormal samples are hard to obtain. Xinghui Dong, Christopher J. Taylor 0001, Timothy F. Cootes |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2022 | Learning the Precise Feature for Cluster AssignmentabstractClustering is one of the fundamental tasks in computer vision and pattern recognition. Recently, deep clustering methods (algorithms based on deep learning) have attracted wide attention with their impressive performance. Most of these algorithms combine deep unsupervised representation learning and standard clustering together. However, the separation of representation learning and clustering will lead to suboptimal solutions because the two-stage strategy prevents representation learning from adapting to subsequent tasks (e.g., clustering according to specific cues). To overcome this issue, efforts have been made in the dynamic adaption of representation and cluster assignment, whereas current state-of-the-art methods suffer from heuristically constructed objectives with the representation and cluster assignment alternatively optimized. To further standardize the clustering problem, we audaciously formulate the objective of clustering as finding a precise feature as the cue for cluster assignment. Based on this, we propose a general-purpose deep clustering framework, which radically integrates representation learning and clustering into a single pipeline for the first time. The proposed framework exploits the powerful ability of recently developed generative models for learning intrinsic features, and imposes an entropy minimization on the distribution of the cluster assignment by a dedicated variational algorithm. The experimental results show that the performance of the proposed method is superior, or at least comparable to, the state-of-the-art methods on the handwritten digit recognition, fashion recognition, face recognition, and object recognition benchmark datasets. Yanhai Gan, Xinghui Dong, Huiyu Zhou 0001, Feng Gao 0005, Junyu Dong |
IEEE Trans. Cybern. | 2 |
| 2021 | Automatic aerospace weld inspection using unsupervised local deep feature learning
Xinghui Dong, Christopher J. Taylor 0001, Timothy F. Cootes |
Knowl. Based Syst. | 1 |
| 2021 | Perceptual Texture Similarity Estimation: An Evaluation of Computational FeaturesabstractEstimation of texture similarity is fundamental to many material recognition tasks. This study uses fine-grained human perceptual similarity ground-truth to provide a comprehensive evaluation of 51 texture feature sets. We conduct two types of evaluation and both show that these features do not estimate similarity well when compared against human agreement rates, but that performances are improved when the features are combined using a Random Forest. Using a simple two-stage statistical model we show that few of the features capture long-range aperiodic relationships. We perform two psychophysical experiments which indicate that long-range interactions do provide humans with important cues for estimating texture similarity. This motivates an extension of the study to include Convolutional Neural Networks (CNNs) as they enable arbitrary features of large spatial extent to be learnt. Our conclusions derived from the use of two pre-trained CNNs are: that the large spatial extent exploited by the networks' top convolutional and first fully-connected layers, together with the use of large numbers of filters, confers significant advantage for estimation of perceptual texture similarity. Xinghui Dong, Junyu Dong, Mike J. Chantler |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | A Random Forest-Based Automatic Inspection System for Aerospace Welds in X-Ray ImagesabstractIn the aerospace manufacturing industry, nondestructive evaluation (NDE) of components plays an important role. Porosities and other defects usually occur in the welds of these components. If such defects end up in the aircraft, the fatigue life of components is lessened, which may cause disastrous accidents. At present, those welds are manually evaluated by human inspectors via reviewing X-ray images. To reduce the workload of inspectors, we have developed an automatic inspection system for identifying defects in linear thin welds. For an X-ray image, this system starts with localizing the central line of the weld using a random forest (RF) regressor. A region surrounding the line is then investigated using an RF classifier in order to detect defects. After extensive experiments, the results demonstrate that the weld can be precisely localized from X-ray images, and the defect detection module can find 80% of defects that have been identified by human inspectors (i.e., true positives), while fewer than 1.6 false positives per image are returned. It is suggested that the system may be beneficial to human inspectors by reducing their workload. In addition, our system produces encouraging results on the publicly available weld X-ray image data set and a magnetic tile image data set.Note to Practitioners—This work was motivated by the challenge of inspecting aerospace components, which is almost entirely done manually at present. Rather than replacing human inspectors, this work aims at reducing their workload by providing them with an initial inspection result for each component. Especially, the proposed system is able to first localize the Region of Interest (RoI) from an X-ray image of a component and then identify potential defects contained in the RoI. To the best of our knowledge, few existing studies perform defect detection on raw component images. Normally, researchers manually cropped an RoI from these images. The output of our system is the pixelwise location information on potential defects. Our results demonstrate that the proposed system is able to accurately localize the weld and identify 80% of defects contained in abnormal weld images with very few false positives. Given that large weld images ($2304\times1920$pixels) were processed, our system located the weld in 6.6± 1.3 s/image and fulfilled defect detection on each localized weld region in 0.8± 0.1 s. The proposed system was also tested with the publicly available X-ray weld image data set: GDXray and a magnetic tile image data set. Although only a small number of training images were available, promising results were obtained. This suggests that our system is suitable for both X-ray weld images and other images though more work is needed to reduce false positives. Xinghui Dong, Christopher J. Taylor 0001, Timothy F. Cootes |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2021 | Graph-Based CNNs With Self-Supervised Module for 3D Hand Pose Estimation From Monocular RGBabstractHand pose estimation in 3D space from a single RGB image is a highly challenging problem due to self-geometric ambiguities, diverse texture, viewpoints, and self-occlusions. Existing work proves that a network structure with multi-scale resolution subnets, fused in parallel can more effectively shows the spatial accuracy of 2D pose estimation. Nevertheless, the features extracted by traditional convolutional neural networks cannot efficiently express the unique topological structure of hand key points based on discrete and correlated properties. Some applications of hand pose estimation based on traditional convolutional neural networks have demonstrated that the structural similarity between the graph and hand key points can improve the accuracy of the 3D hand pose regression. In this paper, we design and implement an end-to-end network for predicting 3D hand pose from a single RGB image. We first extract multiple feature maps from different resolutions and make parallel feature fusion, and then model a graph-based convolutional neural network module to predict the initial 3D hand key points. Next, we use 2D spatial relationships and 3D geometric knowledge to build a self-supervised module to eliminate domain gaps between 2D and 3D space. Finally, the final 3D hand pose is calculated by averaging the 3D hand poses from the GCN output and the self-supervised module output. We evaluate the proposed method on two challenging benchmark datasets for 3D hand pose estimation. Experimental results show the effectiveness of our proposed method that achieves state-of-the-art performance on the benchmark datasets. Shaoxiang Guo, Eric Rigall, Lin Qi 0004, Xinghui Dong, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | The Importance of Phase to Texture Discrimination and SimilarityabstractIn this article, we investigate the importance of phase for texture discrimination and similarity estimation tasks. We first use two psychophysical experiments to investigate the relative importance of phase and magnitude spectra for human texture discrimination and similarity estimation. The results show that phase is more important to humans for both tasks. We further examine the ability of 51 computational feature sets to perform these two tasks. In contrast with the psychophysical experiments, it is observed that the magnitude data is more important to these computational feature sets than the phase data. We hypothesise that this inconsistency is due to the difference between the abilities of humans and the computational feature sets to utilise phase data. This motivates us to investigate the application of the 51 feature sets to phase-only images in addition to their use on the original data set. This investigation is extended to exploit Convolutional Neural Network (CNN) features. The results show that our feature fusion scheme improves the average performance of those feature sets for estimating humans' perceptual texture similarity. The superior performance should be attributed to the importance of phase to texture similarity. Xinghui Dong, Ying Gao 0005, Junyu Dong, Mike J. Chantler |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | MPS-Net: Learning to recover surface normal for multispectral photometric stereo
Yakun Ju, Lin Qi 0004, Jichao He, Xinghui Dong, Feng Gao 0005, Junyu Dong |
Neurocomputing | 4 |
| 2020 | A joint guidance-enhanced perceptual encoder and atrous separable pyramid-convolutions for image inpainting
Yongle Zhang 0001, Yingyu Wang, Junyu Dong, Lin Qi 0004, Hao Fan 0004, Xinghui Dong, Muwei Jian, Hui Yu 0001 |
Neurocomputing | 6 |
| 2020 | Texture synthesis quality assessment using perceptual texture similarity
Xinghui Dong, Huiyu Zhou 0001 |
Knowl. Based Syst. | 1 |
| 2020 | A dual-cue network for multispectral photometric stereo
Yakun Ju, Xinghui Dong, Yingyu Wang, Lin Qi 0004, Junyu Dong |
Pattern Recognit. | 2 |
| 2020 | A Perception-Inspired Deep Learning Framework for Predicting Perceptual Texture SimilarityabstractSimilarity learning plays a fundamental role in the fields of multimedia retrieval and pattern recognition. Prediction of perceptual similarity is a challenging task as in most cases we lack human labeled ground-truth data and robust models to mimic human visual perception. Although in the literature, some studies have been dedicated to similarity learning, they mainly focus on the evaluation of whether or not two images are similar, rather than prediction of perceptual similarity which is consistent with human perception. Inspired by the human visual perception mechanism, we here propose a novel framework in order to predict perceptual similarity between two texture images. Our proposed framework is built on the top of Convolutional Neural Networks (CNNs). The proposed framework considers both powerful features and perceptual characteristics of contours extracted from the images. The similarity value is computed by aggregating resemblances between the corresponding convolutional layer activations of the two texture maps. Experimental results show that the predicted similarity values are consistent with the human-perceived similarity data. Ying Gao 0005, Yanhai Gan, Lin Qi 0004, Huiyu Zhou 0001, Xinghui Dong, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Texture Classification Using Pair-Wise Difference Pooling-Based Bilinear Convolutional Neural NetworksabstractTexture is normally represented by aggregating local features based on the assumption of spatial homogeneity. Effective texture features are always the research focus even though both hand-crafted and deep learning approaches have been extensively investigated. Motivated by the success of Bilinear Convolutional Neural Networks (BCNNs) in fine-grained image recognition, we propose to incorporate the BCNN with the Pair-wise Difference Pooling (i.e. BCNN-PDP) for texture classification. The BCNN-PDP is built on top of a set of feature maps extracted at a convolutional layer of the pre-trained CNN. Compared with the outer product used by the original BCNN feature set, the pair-wise difference not only captures the pair-wise relationship between two sets of features but also encodes the difference between each pair of features. Considering the importance of the gradient data to the representation of image structures, we further generalise the BCNN-PDP feature set to two sets of feature maps computed from the original image and its gradient magnitude map respectively, i.e. the Fused BCNN-PDP (F-BCNN-PDP) feature set. In addition, the BCNN-PDP can be applied to two different CNNs and is referred to as the Asymmetric BCNN-PDP (A-BCNN-PDP). The three PDP-based BCNN feature sets can also be extracted at multiple scales. Since the dimensionality of the BCNN feature vectors is very high, we propose a new yet simple Block-wise PCA (BPCA) method in order to derive more compact feature vectors. The proposed methods are tested on seven different datasets along with 21 baseline feature sets. The results show that the proposed feature sets are superior, or at least comparable, to their counterparts across different datasets. Xinghui Dong, Huiyu Zhou 0001, Junyu Dong |
IEEE Trans. Image Process. | 1 |
| 2020 | Monocular Visual-IMU Odometry: A Comparative Evaluation of Detector-Descriptor-Based MethodsabstractMonocular visual-inertial measurement unit (IMU) odometry has been widely used in various intelligent vehicles. As a popular technique, detector-descriptor-based visual-IMU odometry is effective and efficient due to the fact that local descriptors are robust against occlusions, background clutter, and abrupt content changes. However, to our knowledge, there is not a comprehensive and comparative evaluation study on the performance of different combinations of detectors and descriptors recently developed. In order to bridge this gap, we conduct such a comparative study in a unified framework. In particular, six typical routes with different lengths, shapes, and road scenes are selected from the well-known KITTI dataset. We first evaluate the performance of different combinations of salient point detectors and local descriptors using the six routes. Then, we tune the parameters of the best detector or descriptor obtained for each route, to further augment the results. This paper provides not only comprehensive benchmarks for assessing various algorithms but also instructive guidelines and insights for developing detectors and descriptors to handle different road scenes. Xingshuai Dong, Xinghui Dong, Junyu Dong, Huiyu Zhou 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Automatic Inspection of Aerospace Welds Using X-Ray ImagesabstractThe non-destructive testing (NDT) of components is very important to the aerospace industry. Welds in these components may contain porosities and other defects. These reduce the fatigue life of components and may result in catastrophic accidents if they end up in the aircraft. Currently such welds are inspected by humans studying radiographs of the welds. We describe an automatic system for detecting defects in welds, with the aim of creating a triage system to reduce the workload on human inspectors. Given an X-ray image of the aerospace weld, the system locates the weld line, then analyses the region around the line to identify abnormalities. Our results show that the weld can be precisely extracted from X-ray images and the defect detection operation can identify 83% of defects with fewer than 3 false positives per image, and thus may be useful for prompting human inspectors to reduce their workload. Xinghui Dong, Christopher J. Taylor 0001, Timothy F. Cootes |
ICPR | 1 |
| 2018 | The Visual Word Booster: A Spatial Layout of Words Descriptor Exploiting Contour CuesabstractAlthough researchers have made efforts to use the spatial information of visual words to obtain better image representations, none of the studies take contour cues into account. Meanwhile, it has been shown that contour cues are important to the perception of imagery in the literature. Inspired by these studies, we propose to use the Spatial Layout of Words (SLoW) to boost visual word based image descriptors by exploiting contour cues. Essentially, the SLoW descriptor utilises contours and incorporates different types of commonly used visual words, including hand-crafted basic contour elements (referred to as "contons"), textons and Scale-Invariant Feature Transform (SIFT) words, deep convolutional words and a special type of words: LBP (Local Binary Pattern) codes. Moreover, SLoW features are combined with Spatial Pyramid Matching (SPM) or Vector of Locally Aggregated Descriptors (VLAD) features. The SLoW descriptor and its combined versions are tested in different tasks. Our results show that they are superior to, or at least comparable to, their counterparts examined in this study. In particular, the joint use of the SLoW descriptor boosts the performance of the SPM and VLAD descriptors. We attribute these results to the fact that contour cues are important to human visual perception and, the SLoW descriptor captures not only local image characteristics but also the global spatial layout of these characteristics in a more perceptually consistent way than its counterparts. Xinghui Dong, Junyu Dong |
IEEE Trans. Image Process. | 1 |
| 2017 | Monocular visual-IMU odometry using multi-channel image patch exemplarsabstractIn this paper, we propose three sets of multi-channel image patch features for monocular visual-IMU (Inertial Measurement Unit) odometry. The proposed feature sets extract image patch exemplars from multiple feature maps of an image. We also modify an existing visual-IMU odometry framework by using different salient point detectors and feature sets and replacing the inlier selection approach with a self-adaptive scheme. The modified framework is used to examine the proposed feature sets. In addition to the Root Mean Square Error (RMSE) metric, we use the Hausdorff distance to measure the inconsistency between the estimated and ground-truth trajectories. Compared to the point-wise comparison used by RMSE, the Hausdorff distance takes the shape inconsistency of two trajectories into account and is hence more perceptually consistent. Experimental results show that the multi-channel feature sets outperform, or perform comparably to, the single gray level channel feature sets examined in this study. Particularly, the multi-channel feature set that uses integral channels, i.e., ICIMGP (Integral Channel Image Patches), outperforms two state-of-the-art feature sets: SIFT (Scale Invariant Feature Transform) and SURF (Speed Up Robust Features). Besides, ICIMGP performs better than the two multi-channel feature sets that are designed based on derivative channels and gradient channels respectively. These promising results are attributed to the fact that the multi-channel features encode richer image characteristics than their single gray level channel counterparts. Xingshuai Dong, Bo He 0002, Xinghui Dong, Junyu Dong |
Multim. Tools Appl. | 3 |
| 2016 | Perceptually Motivated Image Features Using ContoursabstractDong et al. examined the ability of 51 computational feature sets to estimate human perceptual texture similarity; however, none performed well for this task. While it is well-known that the human visual system is extremely adept at exploiting longer-range aperiodic (and periodic) "contour" characteristics in images, none of the investigated feature sets exploit higher order statistics (HOS) over larger image regions ( > 19×19 pixels). We, therefore, hypothesise that long-range HOS, in the form of contour data, are useful for perceptual texture similarity estimation. We present the results of a psychophysical experiment that shows that contour data are more important, than local image patches, or global second-order data, to human observers for this task. Inspired by this finding, we propose a set of perceptually motivated image features (PMIF) that encode the long-range HOS computed from spatial and angular distributions of contour segments. We use two perceptual texture similarity estimation tasks to compare PMIF against the 51 feature sets referred to above and four commonly used contour representations. This new feature set is also examined in the context of two additional tasks: sketch-based image retrieval and natural scene recognition. The results show that the proposed feature set performs better, or at least comparably to, all the other feature sets. We attribute this promising performance to the fact that the proposed feature set exploits both short-range and long-range HOS. Xinghui Dong, Mike J. Chantler |
IEEE Trans. Image Process. | 1 |
| 2014 | Texture Similarity Estimation Using Contours
Xinghui Dong, Mike J. Chantler |
BMVC | 1 |
| 2014 | How Well Do Computational Features Perceptually Rank Textures? A Comparative EvaluationabstractInspired by studies [4, 23, 40] which compared rankings obtained by search engines and human observers, in this paper we compare texture rankings derived by 51 sets of computational features against perceptual texture rankings obtained from a free-grouping experiment with 30 human observers, using a unify evaluation framework. Experimental results show that the MRSAR [37], VZNEIGHBORHOOD [62], LBPHF [2] and LBPBASIC [3] feature sets perform better than their counterparts. However, none of those feature sets are ideal. The best average G and M measures (measures of ranking accuracy from 0 to 1) [15, 5] obtained are 0.36 and 0.25 respectively. We suggest that this poor performance may be due to the small local neighborhood used to calculate higher-order features which cannot capture the long-range interactions that humans have been shown to exploit [14, 16, 49, 56]. Xinghui Dong, Thomas S. Methven, Mike J. Chantler |
ICMR | 1 |
| 2013 | The Importance of Long-Range Interactions to Texture Similarity
Xinghui Dong, Mike J. Chantler |
CAIP (1) | 1 |
| 2005 | A collaborative editing environment for 3D shape objectabstractInnovation on 3D shape design relies on advanced product design theory and effective design tools. The requirements for design tools mainly include collaborative performance and natural performance of interaction. In this paper, a 3D collaborative editing environment based on interactive transaction log is put forward and moreover the system architecture is presented. The collaborative mechanism is discussed and the collision detection of manipulation and application strategy during the course of multi-user concurrency editing is studied, the interactive gesture appropriate to 3D model and application methods are studied too. Finally a prototype system is built, which obtains satisfied result in practice. Dongxing Teng, Hongan Wang, Guozhong Dai, Xinghui Dong |
CSCWD (1) | 4 |