Hua Yang 0002

dblp:11/4973-2 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-5430-5630ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Prior-Guided Feature Sampling and Restoration for Few-Shot Industrial Anomaly Detection
abstract
Anomaly detection (AD) is crucial for industrial defect inspection in the manufacturing process. Mainstream AD approaches rely on large-scale normal data for training, limiting their effectiveness in handling data-scarce situations encountered during production line changeovers. While recent approaches have explored few-shot anomaly detection (FSAD), they face inherent challenges of sparse normal representations and insufficient visual priors. To address these issues, this study proposes a prior-guided feature sampling and restoration network (PFSRN) for high-performance FSAD. First, a neighborhood patch sampling (NPS) module is proposed to generate normal features by sampling from estimated distributions, enhancing the compactness of normal representations in few-shot scenarios. Second, the proposed hierarchical masked linear attention (HMLA) mitigates self-replication during feature restoration, enabling effective anomaly localization through restoration residuals. Third, an adaptive normality injection (ANI) module is introduced to dynamically fuse visual normal priors into the restored features, further enhancing restoration fidelity. In addition, both image-and feature-level perturbations are incorporated to serve as anomalous priors, enabling robust dual-stream training. Extensive experiments across multiple benchmarks demonstrate that PFSRN achieves state-of-the-art FSAD performance. For example, in the 1-shot setting, PFSRN achieves image-level AUROC (%) scores of 97.8, 96.0, and 88.3 on the MVTec-AD, VisA, and MPDD datasets, respectively, outperforming the most advanced counterparts by 1.2, 3.0, and 2.9. Real-world industrial applications in printed circuit board assembly inspection further validate the practical effectiveness of PFSRN.
Hua Yang 0002
IEEE Trans Autom. Sci. Eng.2
2026 Biomimetic Visual Perception Network for Industrial Image Anomaly Detection
abstract
Anomaly detection in the industrial manufacturing process is important for controlling product quality. In real-world scenarios, anomalies manifest in a number of diverse and unforeseen ways and can be categorized into two main types: structural anomalies, such as scratches and breakages, and logical anomalies, such as missing components and dislocations. Current mainstream methods tend to focus on knowledge generalization but are inadequate for detailed and logical detection tasks. To address this issue, this study proposes Biomimetic Visual Perception Network (BVPN), a biomimetic design-based novel paradigm that simulates biological vision and human logical discrimination with detailed distillation dependencies and semantic component information. BVPN introduces three convolution-based networks to extract information from different regions, imitating biological central, para-central, and peripheral vision, and facilitates the integration of disparate information through the process of knowledge distillation. On the basis, BVPN develops two modules based on the student-teacher network: Focus Module, for local detailed detection, and Peripheral Module, for global context detection. Moreover, BVPN proposes the Logical Module, a clustering-based segmentation branch, assessing logical anomalies by the number and color information of segmentation components. Through a fused structure highly aligning with the perception mode of human system, BVPN achieves excellent performance in the detection of structural and logical anomalies. Extensive experiments with mainstream anomaly detection datasets and real-world inkjet printing dataset demonstrate that BVPN outperforms state-of-the-art competitors in accuracy.
Siang Tang, Hua Yang 0002, JianKui Chen, Zhou-Ping Yin
IEEE Trans Autom. Sci. Eng.2
2025 3D Vision-tactile Reconstruction from Infrared and Visible Images for Robotic Fine-grained Tactile Perception
abstract
To achieve human-like haptic perception in anthropomorphic grippers, the compliant sensing surfaces of vision tactile sensor (VTS) must evolve from conventional planar configurations to biomimetically curved topographies with continuous surface gradients. However, planar VTSs have challenges when extended to curved surfaces, including insufficient lighting of surfaces, blurring in reconstruction, and complex spatial boundary conditions for surface structures. With an end goal of constructing a human-like fingertip, our research (i) develops GelSplitter3D by expanding imaging channels with a prism and a near-infrared (NIR) camera, (ii) proposes a photometric stereo neural network with a CAD-based normal ground truth generation method to calibrate tactile geometry, and (iii) devises a normal integration method with boundary constraints of depth prior information to correcting the cumulative error of surface integrals. We demonstrate better tactile sensing performance, a 40% improvement in normal estimation accuracy, and the benefits of sensor shapes in grasping and manipulation tasks.
Yuankai Lin, Xiaofan Lu, Hua Yang 0002
IROS4
2025 Multi-pulse superposition for droplet volume control in inkjet printing based on model and data fusion
Xiao Yue, Xin Li 0197, JianKui Chen, Hua Yang 0002, Jincheng Gao, Zhou-Ping Yin
Eng. Appl. Artif. Intell.5
2025 An adversarial network based on anomaly domain decomposition and transformation for industrial PCBA defect inspection
Qianfeng Pan, Jingyun Chang, Hua Yang 0002
Neural Comput. Appl.5
2025 A Hierarchical Patch Feature Distribution Network for Industrial Multiscale Defect Detection
abstract
Visual inspection of surface defects in industrial products is crucial for quality control but remains challenging due to unpredictable and multiscale defects. Unsupervised anomaly detection methods based on feature distribution have rapidly developed, requiring only normal samples and no additional manual labeling. However, these methods still suffer from feature bias and image misalignment, leading to overdetection of multiscale defects, making them impractical for deployment in factory settings. To address these issues, this study proposes a hierarchical patch feature distribution modeling (HPDM) network. Specifically, the new patch feature alignment learning (PFAL) module mitigates image misalignment by enhancing feature dimension similarity via the PFAL loss function. Additionally, the proposed discrete feature memory (DFM) module employs linear mapping to increase the density of the normal feature distribution, addressing feature bias from pretrained networks. Furthermore, the pyramid-shaped Gaussian distribution (PGD) module fits multivariate Gaussian models at different feature scales for defect detection across various scales. On the MVTec2D, MVTec3D, and VISION datasets, the HPDM network achieves pixel-level AUROC scores of 98.83%, 98.30%, and 71.20%, respectively, surpassing those of the existing methods and demonstrating state-of-the-art anomaly detection performance. The experimental results on a real printed circuit board assembly (PCBA) circuit dataset demonstrate a detection speed of 30 frames per second (FPS), meeting the requirements for factory online detection. Note to Practitioners—Previous unsupervised anomaly detection methods have struggled to detect both large and small surface defects simultaneously. However, the proposed multiscale patch feature modeling network effectively identifies various multiscale defects on product surfaces, including PCBA, thin-film transistor liquid crystal displays, and tiles. Additionally, HPDM requires only a small number of normal samples to learn a robust network, which is crucial for industrial applications where identifying and labeling samples is challenging. Furthermore, by leveraging model pruning and quantization techniques, HPDM can be applied to online visual inspection.
Hua Yang 0002, Zhou-Ping Yin
IEEE Trans Autom. Sci. Eng.2
2024 GelRoller: A Rolling Vision-based Tactile Sensor for Large Surface Reconstruction Using Self-Supervised Photometric Stereo Method
abstract
Accurate perception of the surrounding environment stands as a primary objective for robots. Through tactile interaction, vision-based tactile sensors provide the capability to capture high-resolution and multi-modal surface information of objects, thereby facilitating robots in achieving more dexterous manipulations. However, the prevailing GelSight sensors entail intricate calibration procedures, posing challenges in their application on curved surfaces and requiring the maintenance of stable lighting conditions throughout experimentation. Additionally, constrained by shape and structure, current vision-based tactile sensors are predominantly applied to measurements within a limited area. In this study, we design a novel cylindrical vision-based tactile sensor that enables continuous and swift perception of large-scale object surfaces through rolling. To tackle the challenges posed by laborious calibration processes, we propose a self-supervised photometric stereo method based on deep learning, which eliminates pre-calibration requirements and enables the derivation of surface normals from a single image without relying on stable lighting conditions. Finally, we perform surface reconstruction from normal and point cloud registration on the multiple frames of images obtained by rolling the cylindrical sensor, resulting in large surface reconstruction. We compare our method with the representative lookup table method in the GelSight sensors. The results show that the proposed method enhances both reconstruction accuracy and robustness, thereby demonstrating the potential of the proposed sensor in large-scale surface reconstruction. Codes and mechanical structures are available at: https://github.com/ZhangZhiyuanZhang/GelRoller
Zhiyuan Zhang 0014, Jingjing Ji, Hua Yang 0002
ICRA5
2024 SDNet: Spatial adversarial perturbation local descriptor learned with the dynamic probabilistic weighting loss
Kaiji Huang, Hua Yang 0002, Zhou-Ping Yin
Neurocomputing2
2024 Multi-Category Decomposition Editing Network for the Accurate Visual Inspection of Texture Defects
abstract
Spotting blemished areas automatically on a textured surface is a particular challenge, as both nominal and defective surface samples are inconsistent in large-scale industrial manufacturing. The most efficient solution uses the memory bank extracted from the nominal samples to detect outliers. We approach our strategy, the multi-category decomposition editing network (MCDEN), from a similar viewpoint. Notably, we do not use defect-free samples. Instead, we use virtual results to construct a defect library. MCDEN decomposes abnormalities to basic elements from the library while editing outlier features to reconstruct the texture normality, offering a rational segmentation map through decomposition and reconstruction. Based on this strategy, MCDEN is more interpretable than most neural network methods since interpretability is particularly important in industrial production to ensure stability; however, the existing deep learning methods are similar to a black box structure, which makes MCDEN more appropriate in industry. Experiments on texture surface samples from the MVTec anomaly detection (MVTAD) dataset confirm the efficacy of MCDEN with a pixel-level area under the receiver operator characteristic curve (AUC) score of 96.6%. In other experiments collected from semi-manufactured inkjet printing organic electroluminescence display (OLED) panels, MCDEN demonstrated competitive results with a 99.2% detection rate and rapid real-time detection capabilityNote to Practitioners—MCDEN detects defects based on the assumption that texture defects are decomposed into five basic transformation combinations. The method does not need to collect additional defect images for model training. It only needs to collect about 2000 positive defect-free images to generate negative samples through the random generation of defects and complete the training. All training images and detection images are based on a single-channel gray image. The MCDEN method can be applied to the detection of object defects on textured surfaces with periodic features, such as steel, leather, and screen. Based on the known texture features, the trained model has high detection accuracy and real-time detection for specific types of textured surfaces.
Hua Yang 0002, JianKui Chen, Zhou-Ping Yin
IEEE Trans Autom. Sci. Eng.1
2024 DySPN: Learning Dynamic Affinity for Image-Guided Depth Completion
abstract
Image-guided depth completion (IGDC) is a multimodal computer vision task for acquiring high-precision dense depth maps. It reasonably predicts values around accurate sparse depth measurements by relying on the details of simultaneous dense RGB images. To achieve this goal, spatial propagation networks (SPNs) elaborate context-aware meta cells that connect pixels to their neighbors and fuse sparse and dense modalities by linear propagation. However, static affinity matrices and fixed neighborhood connections limit the representation of the networks. In this paper, our proposed dynamic SPN (DySPN) uses a nonlinear propagation model (NLPM), which processes the propagation more finely by adjusting the affinity weights, diffusion paths, and the number of neighbors. Specifically, we first generate adaptive weighting (AW) matrices by decoupling the neighborhood into parts with respect to different distances. Independent attention maps are recursively applied to refine the weight value. Furthermore, a dynamic path (DP) strategy is adopted to unfreeze the links of the neighborhood for learning variable connections. The solution space of the paths is also constrained by a propagation decay loss to keep the results stable. Finally, we introduce a diffusion suppression (DS) operation, which preserves the edge of dense depth maps by manipulating the AW and DP strategies to decelerate and terminate the propagation. In our experiment, the proposed method requires fewer iterations and neighbors than other SPNs while yielding better results. DySPN outperforms state-of-the-art (SoTA) methods on the KITtI DC, NYU Depth v2, and VOID datasets. Our code is available at: https://github.com/Kyakaka/DySPN.
Yuankai Lin, Hua Yang 0002, Wending Zhou, Zhou-Ping Yin
IEEE Trans. Circuits Syst. Video Technol.2
2023 A Semantic Information Decomposition Network for Accurate Segmentation of Texture Defects
abstract
Defect detection on textured surfaces remains a challenging task due to the wide range of textures and defects. Current unsupervised learning-based texture defect detection methods based on texture background reconstruction cannot detect texture defects with high precision because it is difficult to guarantee a high-precision reconstruction of the texture background while suppressing the defect foreground. In this study, we propose a novel semantic information decomposition network (SIDN) for accurate texture defect segmentation. The SIDN is trained on artificial defective images produced by a defect generation module (DGM). First, the SIDN uses a feature extraction module (FEM) to extract latent features with both texture semantic information and defect semantic information. Then, a novel feature separation extraction module (FSEM) for decomposing the texture semantic information and defect semantic information from the feature map generated by the FEM is proposed, preventing the coupling of the texture and defect semantic information from affecting the final segmentation accuracy. Next, a novel global semantic relation module (GSRM) is proposed to determine the relevance of the global semantic information to comprehensively consider the context and improve the feature representation. Finally, a segmentation module (SM) that directly segments the textures and defects instead of reconstructing the texture background is proposed. The final detection result is obtained by calculating a weighted average of the texture and defect segmentation results. The extensive experimental tests with the most popular and most challenging texture defect dataset demonstrate that the SIDN achieves accurate segmentation of various texture defects without using real defect samples.
Hua Yang 0002, Jiale Hu, Zhou-Ping Yin
IEEE Trans. Ind. Informatics1
2022 Dynamic Spatial Propagation Network for Depth Completion
abstract
Image-guided depth completion aims to generate dense depth maps with sparse depth measurements and corresponding RGB images. Currently, spatial propagation networks (SPNs) are the most popular affinity-based methods in depth completion, but they still suffer from the representation limitation of the fixed affinity and the over smoothing during iterations. Our solution is to estimate independent affinity matrices in each SPN iteration, but it is over-parameterized and heavy calculation.This paper introduces an efficient model that learns the affinity among neighboring pixels with an attention-based, dynamic approach. Specifically, the Dynamic Spatial Propagation Network (DySPN) we proposed makes use of a non-linear propagation model (NLPM). It decouples the neighborhood into parts regarding to different distances and recursively generates independent attention maps to refine these parts into adaptive affinity matrices. Furthermore, we adopt a diffusion suppression (DS) operation so that the model converges at an early stage to prevent over-smoothing of dense depth. Finally, in order to decrease the computational cost required, we also introduce three variations that reduce the amount of neighbors and attentions needed while still retaining similar accuracy. In practice, our method requires less iteration to match the performance of other SPNs and yields better results overall. DySPN outperforms other state-of-the-art (SoTA) methods on KITTI Depth Completion (DC) evaluation by the time of submission and is able to yield SoTA performance in NYU Depth v2 dataset as well.
Yuankai Lin, Wending Zhou, Hua Yang 0002
AAAI5
2021 Multi-Scale Boosting Feature Encoding Network for Texture Recognition
abstract
Texture recognition remains a challenging visual task due to the complex appearance variations caused by scale changes in the real world. In most existing texture recognition methods, textures are represented at a single scale; thus, multi-scale texture information is not fully utilized, resulting in insufficient representation and inaccurate recognition. In this study, with the goal of addressing the challenge of scale changes, we propose a novel multi-scale boosting feature encoding network (MSBFEN) for accurate texture recognition. MSBFEN first extracts multi-scale features with multi-scale texture structure information under the guidance of texture priors using a novel prior-guided feature extraction (PFE) method. Then, a multi-scale texture encoding (MSTE) method is devised to capture discriminative multi-scale texture representations by encoding the extracted features. Finally, to fully utilize the multi-scale texture representations for accurate texture recognition, a novel multi-scale boosting learning (MSBL) method is proposed. In MSBL, the learning procedure for multi-scale texture recognition is boosted in a hierarchical, progressively reinforced manner, significantly addressing the challenge of scale changes and greatly enhancing the recognition accuracy. In addition, a novel outlier-aware texture encoding (OTE) method is proposed for robust texture encoding at each scale of MSTE. OTE can resist the influence of background interference and can further enhance the robustness of MSBFEN. In extensive experiments conducted on six challenging texture recognition datasets, namely, KTH-TIPS2b, FMD, DTD, MINC, GTOS and GTOS-mobile, MSBFEN achieves accuracies of 86.2%, 86.4%, 77.8%, 85.3%, 86.4% and 87.57%, respectively, representing state-of-the-art texture recognition performance.
Kaiyou Song, Hua Yang 0002, Zhou-Ping Yin
IEEE Trans. Circuits Syst. Video Technol.2
2021 An Anomaly Feature-Editing-Based Adversarial Network for Texture Defect Visual Inspection
abstract
Establishing a unified model for the defect inspection of different texture surfaces remains a challenge in the industrial automation field because these surfaces can vary in regular and irregular ways. Current unsupervised learning methods are trained on defect-free samples only and cannot directly address anomalies during testing, which precludes these methods from simultaneously inspecting for various texture defects. In this article, we propose a novel unsupervised anomaly feature-editing-based adversarial network (AFEAN) to accurately inspect various texture defects. To impart the AFEAN with the ability to address anomalies, a paired input, consisting of a defect-free image and an artificially defective image, is utilized for training. First, the AFEAN employs a feature extraction module (FEM) to extract latent features for the paired input. Subsequently, a novel anomaly feature detection module (AFDM) is proposed to detect anomaly features of the artificially defective image in the latent space. In the proposed AFDM, a novel central-constraint-based clustering method is proposed to detect anomaly features by learning the distribution of the latent features. Next, a novel global context feature editing module (GCFEM) is proposed to convert the detected anomaly features to normal features to suppress the reconstruction of defects. Finally, a feature decoding module (FDM) utilizes the edited features to reconstruct the texture background. Through the AFDM and GCFEM, the AFEAN achieves the ability to address anomaly features, effectively suppressing the reconstruction of defects on the texture background. In addition, to further improve the texture reconstruction accuracy, a pixel-level discrimination module (PDM) is employed to reconstruct texture details. In the testing phase, the defects are segmented by the residual image between the input image and the reconstructed texture background. The extensive experimental results demonstrate that the AFEAN achieves the state-of-the-art inspection accuracy.
Hua Yang 0002, Qinyuan Zhou, Kaiyou Song, Zhou-Ping Yin
IEEE Trans. Ind. Informatics1
2021 Weighted Feature Histogram of Multi-Scale Local Patch Using Multi-Bit Binary Descriptor for Face Recognition
abstract
Most face recognition methods employ single-bit binary descriptors for face representation. The information from these methods is lost in the process of quantization from real-valued descriptors to binary descriptors, which greatly limits their robustness for face recognition. In this study, we propose a novel weighted feature histogram (WFH) method of multi-scale local patches using multi-bit binary descriptors for face recognition. First, to obtain multi-scale information of the face image, the local patches are extracted using a multi-scale local patch generation (MSLPG) method. Second, with the goal of reducing the quantization information loss of binary descriptors, a novel multi-bit local binary descriptor learning (MBLBDL) method is proposed to extract multi-bit local binary descriptors (MBLBDs). In MBLBDL, a learned mapping matrix and novel multi-bit coding rules are employed to project pixel difference vectors (PDVs) into the MBLBDs in each local patch. Finally, a novel robust weight learning (RWL) method is proposed to learn a set of robust weights for each patch to integrate the MBLBDs into the final face representation. In RWL, a codebook is first constructed by clustering MBLBDs on each local patch to extract a feature histogram. Then, considering that different parts of the face have different degrees of robustness to local changes, a set of weights is learned to concatenate the feature histograms of all local patches into the final representation of a face image. In addition, to further improve the performance for heterogeneous face recognition, a coupled WFH (C-WFH) method is proposed. C-WFH maintains the similarity of the corresponding MBLBDs and feature histograms for a pair of heterogeneous face images by means of a novel coupled feature learning (CFL) method to reduce the modality gap. A series of experiments are conducted on widely used face datasets to analyze the performance of WFH and C-WFH. Extensive experimental results show that WFH and C-WFH outperform state-of-the-art face recognition methods.
Hua Yang 0002, Chenting Gong, Kaiji Huang, Kaiyou Song, Zhou-Ping Yin
IEEE Trans. Image Process.1
2019 Large-scale and rotation-invariant template matching using adaptive radial ring code histograms
Hua Yang 0002, Chenghui Huang, Feiyue Wang 0003, Kaiyou Song, Shijiao Zheng, Zhou-Ping Yin
Pattern Recognit.1
2019 Multiscale Feature-Clustering-Based Fully Convolutional Autoencoder for Fast Accurate Visual Inspection of Texture Surface Defects
abstract
Visual inspection of texture surface defects is still a challenging task in the industrial automation field due to the tremendous changes in the appearance of various surface textures. Current visual inspection methods cannot simultaneously and efficiently inspect various types of texture defects due to either the low discriminative capabilities of handcrafted features or their time-consuming sliding-window strategy. In this paper, we present a novel unsupervised multiscale feature-clustering-based fully convolutional autoencoder (MS-FCAE) method that efficiently and accurately inspects various types of texture defects based on a small number of defect-free texture samples. The proposed MS-FCAE method utilizes multiple FCAE subnetworks at different scale levels to reconstruct several textured background images. The residual images are obtained by subtracting these texture backgrounds from the input image individually; then, they are fused into one defect image. To maximize the efficiency, each FCAE subnetwork utilizes fully convolutional neural networks to extract the original feature maps directly from the input images. Meanwhile, each FCAE subnetwork performs feature clustering to improve the discriminant power of the encoded feature maps. The proposed MS-FCAE method is evaluated on several texture surface inspection data sets both qualitatively and quantitatively. This method achieves a Precision of 92.0% while requiring only 82 ms for input images of $1920\times 1080$ pixels. The extensive experimental results demonstrate that MS-FCAE achieves highly efficient and state-of-the-art inspection accuracy. Note to Practitioners-Most conventional visual inspection methods can address only one specific type of texture defect, while multiscale feature-clustering-based fully convolutional autoencoder (MS-FCAE) can simultaneously and accurately inspect various types of texture surface defects, such as those of thin-film transistor liquid crystal displays, wood, fabrics, and ceramic tiles. Furthermore, MS-FCAE requires only a small number of surface texture samples to learn a robust network model, and its training requires no defect samples. This is extremely important for industrial applications because identifying and labeling defect samples is difficult. Moreover, MS-FCAE can be applied to online visual inspection utilizing a graphics processing unit-based parallel processing strategy.
Hua Yang 0002, Kaiyou Song, Zhou-Ping Yin
IEEE Trans Autom. Sci. Eng.1
2019 Multi-Scale Attention Deep Neural Network for Fast Accurate Object Detection
abstract
Object detection remains a challenging task in computer vision due to the tremendous extent of changes in the appearances of objects caused by clustered backgrounds, occlusion, truncation, and scale change. Current deep neural network (DNN)-based object detection methods cannot simultaneously achieve a high accuracy and a high efficiency. To overcome this limitation, in this paper, we propose a novel multi-scale attention (MSA) DNN for accurate object detection with high efficiency. The proposed MSA-DNN method utilizes a novel multi-scale feature fusion module (MSFFM) to construct high-level semantic features. Subsequently, a novel MSA module (MSAM) based on the fused layers of the MSFFM is introduced to exploit the global semantic information of image-level labels to guide detection. On the one hand, MSAM can capture global semantic information to further enhance the semantic feature representation of the fused layers constructed by the MSFFM, thereby improving the detection accuracy. On the other hand, the MSA maps generated by MSAM can be employed to rapidly and coarsely locate objects at different scales. In addition, an attention-based hard negative mining strategy is introduced to filter out negative samples to reduce the search space, dramatically alleviating the severe class imbalance problem. Extensive experimental results on the challenging PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO datasets demonstrate that MSA-DNN achieves a state-of-the-art detection accuracy while maintaining a high efficiency. Furthermore, MSA-DNN significantly improves the small-object detection accuracy.
Kaiyou Song, Hua Yang 0002, Zhou-Ping Yin
IEEE Trans. Circuits Syst. Video Technol.2
2019 Robust Semantic Template Matching Using a Superpixel Region Binary Descriptor
abstract
Almost all conventional template-matching methods employ low-level image features to measure the similarity between a template image and a scene image using similarity measures such as pixel intensity and pixel gradient. Although these methods have been widely used in many applications, they cannot simultaneously address all types of robustness challenges. In this study, with the goal of simultaneously addressing the various challenges, we present a robust semantic template-matching approach (RSTM). Inspired by the local binary descriptor, we propose a novel superpixel region binary descriptor (SRBD) to construct a multilevel semantic fusion feature vector for RSTM. SRBD uses a new kernel-distance-based simple linear iterative clustering (KD-SLIC) method to extract the stable superpixels from the template image; Then, based on the average intensity difference between each superpixel region and its neighbors, the dominant gradient orientation of each superpixel can be obtained, and the semantic features of each superpixel can be described as the dominant orientation difference vector, which is coded as the rotation-invariant SRBD. In the off-line matching phase, the fusion semantic feature vector of RSTM combines the multilevel SRBD features with different numbers of superpixels. In the online matching phase, to cope with rotation invariance, a marginal probability model is proposed and applied to locate the positions of template images in the scene image. Moreover, to accelerate computation, an image pyramid is employed. We conduct a series of experiments on a large dataset randomly selected from the MS COCO dataset to fully analyze the robustness of this approach. The experimental results show that RSTM simultaneously addresses rotation changes, scale changes, noise, occlusions, blur, nonlinear illumination changes and deformation with high time efficiency while also outperforming previous stateof- the-art template-matching methods.
Hua Yang 0002, Chenghui Huang, Feiyue Wang 0003, Kaiyou Song, Zhou-Ping Yin
IEEE Trans. Image Process.1
2018 An Uncalibrated Visual Servo Method Based on Projective Homography
abstract
An uncalibrated visual servo method based on projective homography, denoted as Projective Homography based Uncalibrated Visual Servoing (PHUVS), is proposed in this paper, in which a novel task function based on the element of projective homography is devised to realize visual servo without a prior knowledge of the camera intrinsic parameters and hand-eye relationships. The main advantage of this method is that it is not only suitable for totally uncalibrated scenarios but also cheap in computation costs when compared with classical image-based uncalibrated visual servoing methods. Numerical experiments are performed and the results confirm that the new approach is capable of both static positioning and dynamic tracking tasks, and presents competitive computational efficiency and accuracy performance.
Zeyu Gong, Bo Tao 0001, Hua Yang 0002, Zhou-Ping Yin, Han Ding 0001
IEEE Trans Autom. Sci. Eng.3
2018 An Accurate Mura Defect Vision Inspection Method Using Outlier-Prejudging-Based Image Background Construction and Region-Gradient-Based Level Set
abstract
The visual inspection of Mura defects is still a challenging task in the quality control of panel displays because of the intrinsically nonuniform brightness and blurry contours of these defects. The current methods cannot detect all Mura defect types simultaneously, especially small defects. In this paper, we introduce an accurate Mura defect visual inspection (AMVI) method for the fast simultaneous inspection of various Mura defect types. The method consists of two parts: an outlier-prejudging-based image background construction (OPBC) algorithm is proposed to quickly reduce the influence of image backgrounds with uneven brightness and to coarsely estimate the candidate regions of Mura defects. Then, a novel region-gradient-based level set (RGLS) algorithm is applied only to these candidate regions to quickly and accurately segment the contours of the Mura defects. To demonstrate the performance of AMVI, several experiments are conducted to compare AMVI with other popular visual inspection methods are conducted. The experimental results show that AMVI tends to achieve better inspection performance and can quickly and accurately inspect a greater number of Mura defect types, especially for small and large Mura defects with uneven backlight.
Hua Yang 0002, Kaiyou Song, Shuang Mei, Zhou-Ping Yin
IEEE Trans Autom. Sci. Eng.1
2016 Transferred Deep Convolutional Neural Network Features for Extensive Facial Landmark Localization
abstract
Features are crucial for extensive facial landmark localization (EFLL), while deep convolutional neural network (DCNN) features lead to breakthroughs in diverse visual recognition tasks. However, there is little study of DCNN features for EFLL, mainly because less labeled data with extensive facial landmarks are available. In this letter, we employ transfer learning to overcome this limitation, and utilize DCNN for EFLL. We concentrate the power of DCNN on feature learning within a cascaded-regression framework (CRF). We present three transfer methods, which show the capacity of DCNN as a generic feature extractor, and the benefit of fine-tuning. The proposed specific fine-tuning method for cascaded regression, named cascade transfer, achieves competitive accuracy with state-of-the-art methods on the 300-W challenge dataset.
Hua Yang 0002, Zhou-Ping Yin
IEEE Signal Process. Lett.2
2016 Polygon-Invariant Generalized Hough Transform for High-Speed Vision-Based Positioning
abstract
The generalized Hough transform (GHT) is widely used for detecting or locating objects under similarity transformation. However, a weakness of the traditional GHT is its large storage requirement and time-consuming computational complexity due to the 4-D parameter space voting strategy. In this paper, a polygon-invariant GHT (PI-GHT) algorithm, as a novel scale- and rotation-invariant template matching method, is presented for high-speed object vision-based positioning. To demonstrate the performance of PI-GHT, several experiments were carried out to compare this novel algorithm with the other five popular matching methods. Experimental results show that the computational effort required by PI-GHT is smaller than that of the common methods due to the similarity transformations applied to the scale- and rotation-invariant triangle features. Moreover, the proposed PI-GHT maintains inherent robustness against partial occlusion, noise, and nonlinear illumination changes, because the local triangle features are based on the gradient directions of edge points. Consequently, PI-GHT is implemented in packaging equipment for radio frequency identification devices at an average time of 4.13 ms and 97.06% matching rate, to solder paste printing at average time nearly 5 ms with 99.87%. PI-GHT is applied to LED manufacturing equipment to locate multiobjects at least five times improvement in speed with a 96% matching rate.
Hua Yang 0002, Shijiao Zheng, Zhou-Ping Yin
IEEE Trans Autom. Sci. Eng.1
2015 Multi-object Template Matching Using Radial Ring Code Histograms
Shijiao Zheng, Buyang Zhang, Hua Yang 0002
ICIG (2)3