Yimin Yang 0001

dblp:41/6960-1 · also Yi-min Yang 0001 · DBLP profile ↗
← Back
56ranked-venue papers
10as first author
33since 2021 · last 2026
0000-0002-1131-2056ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 9 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 MMDC-CLIP-F: Vision-language multi-view mammogram density classification with uncertainty assessment
abstract
Mammogram density classification is a critical component of breast cancer screening, as breast density is a well-established risk factor that also impacts the sensitivity of mammographic imaging. Traditional deep learning (DL) approaches, such as convolutional neural network (CNN) based models, have shown limitations in this domain, often struggling with poor inter-class differentiation and lacking the ability to leverage relational context between different mammographic views. To address these challenges, we propose a two-part framework for mammogram density assessment. The first component, Multi-View Mammogram Density Classification using Contrastive Language-Image Pretraining (MMDC-CLIP), combines the representational strength of vision–language models with multi-view fusion. Semantic prompts are used to inject domain-specific priors, enhancing feature discrimination, while randomized data augmentation mitigates the challenges of limited annotated datasets. The second component, a Multi-View Auxiliary Confidence Network (MV-ACN), processes the final hidden states from all views through a multi-head attention mechanism to generate calibrated confidence scores, enabling reliable identification of uncertain cases that may require secondary review. Together, MMDC-CLIP and MV-ACN form the proposed MMDC-CLIP-F framework. The MMDC-CLIP classifier using the CLIP ViT-L/14-336 backbone reaches 78.2% accuracy, 5.9 percentage points higher than MV-DEFEAT, and 91.5% multi-class AUC, an 8.9-percentage-point improvement, on the RSNA-SMBC dataset. MV-ACN further provides calibrated uncertainty estimates; when paired with MMDC-CLIP using the CLIP ViT-B/32 backbone, confidence stratification on RSNA-SMBC yields 93.9% accuracy in high-confidence samples compared with 52.4% in low-confidence samples. These calibrated confidence estimates enable downstream decision support, such as deferring low-reliability cases for radiologist review, thereby improving the safety and interpretability of automated mammogram density assessment.
Jacob Schaffer, Wandong Zhang, Tianqi Ni, Yimin Yang 0001, Ameya Madhav Kulkarni, Ashirbani Saha
Neurocomputing4
2025 An analytic formulation of convolutional neural network learning for pattern recognition
Huiping Zhuang, Zhiping Lin 0001, Yimin Yang 0001, Kar-Ann Toh
Inf. Sci.3
2025 Diffusion-based data augmentation and hierarchical CLIP for real estate image annotation
Haojin Deng, Wandong Zhang, Yimin Yang 0001, Eman Nejad
Pattern Anal. Appl.3
2025 Fast Transfer Learning Method Using Random Layer Freezing and Feature Refinement Strategy
abstract
Recently, Moore-Penrose inverse (MPI)-based parameter fine-tuning of fully connected (FC) layers in pretrained deep convolutional neural networks (DCNNs) has emerged within the inductive transfer learning (ITL) paradigm. However, this approach has not gained significant traction in practical applications due to its stringent computational requirements. This work addresses this issue through a novel fast retraining strategy that enhances applicability of the MPI-based ITL. Specifically, during each retraining epoch, a random layer freezing protocol is utilized to manage the number of layers undergoing feature refinement. Additionally, this work incorporates an MPI-based approach for refining the trainable parameters of FC layers under batch processing, contributing to expedited convergence. Extensive experiments on several ImageNet pretrained benchmark DCNNs demonstrate that the proposed ITL achieves competitive performance with excellent convergence speed compared to conventional ITL methods. For instance, the proposed strategy converges nearly 1.5 times faster than retraining the ImageNet pretrained ResNet-50 using stochastic gradient descent with momentum (SGDM).
Wandong Zhang, Yimin Yang 0001, Akilan Thangarajah, Q. M. Jonathan Wu, Tianlong Liu
IEEE Trans. Cybern.2
2025 Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
abstract
Contrastive learning has gained popularity and pushes state-of-the-art performance across numerous large-scale benchmarks. In contrastive learning, the contrastive loss function plays a pivotal role in discerning similarities between samples through techniques such as rotation or cropping. However, this learning mechanism can also introduce information distortion from the augmented samples. This is because the trained model may develop a significant overreliance on information from samples with identical labels, while concurrently neglecting positive pairs that originate from the same initial image, especially in expansive datasets. This paper proposes a context-enriched contrastive loss function that concurrently improves learning effectiveness and addresses the information distortion by encompassing two convergence targets. The first component, which is notably sensitive to label contrast, differentiates between features of identical and distinct classes which boosts the contrastive training efficiency. Meanwhile, the second component draws closer the augmented samples from the same source image and distances all other samples, similar to self-supervised learning. We evaluate the proposed approach on image classification tasks, which are among the most widely accepted 8 recognition large-scale benchmark datasets: CIFAR10, CIFAR100, Caltech-101, Caltech-256, ImageNet, BiasedMNIST, UTKFace, and CelebA datasets. The experimental results demonstrate that the proposed method achieves improvements over 16 state-of-the-art contrastive learning methods in terms of both generalization performance and learning convergence speed. Interestingly, our technique stands out in addressing systematic distortion tasks. It demonstrates a 22.9% improvement compared to original contrastive loss functions in the downstream BiasedMNIST dataset, highlighting its promise for more efficient and equitable downstream training.
Haojin Deng, Yimin Yang 0001
IEEE Trans. Multim.2
2024 Explored seeds generation for weakly supervised semantic segmentation
Terence Chow, Haojin Deng, Yimin Yang 0001, Zhiping Lin 0001, Huiping Zhuang, Shan Du 0001
Neural Comput. Appl.3
2024 Low-Light Salient Object Detection by Learning to Highlight the Foreground Objects
abstract
Previous methods in salient object detection (SOD) mainly focused on favorable illumination circumstances while neglecting the performance in low-light condition, which significantly impedes the development of related down-stream tasks. In this work, considering that it is impractical to annotate the large-scale labels for this task, we present a framework (HDNet) to detect the salient objects in low-light images with the synthetic images. Our HDNet consists of a foreground highlight sub-network (HNet) and an appearance-aware detection sub-network (DNet), both of which can be learned jointly in an end-to-end manner. Specifically, to highlight the foreground objects, we design the HNet to estimate the parameters to adjust the dynamic range for each pixel adaptively, which can be trained via the weak supervision signals of the salient object labels. In addition, we design a simple detection network (DNet) with a contextual feature fusion module and a multi-scale feature refine module for detailed feature fusion and refinement. Furthermore, we contribute the first annotated dataset for salient object detection in low-light images (SOD-LL), including 6,000 labeled synthetic images (SOD-LLS) and 2,000 labeled real images (SOD-LLR). Experimental results on SOD-LL and other low-light videos in the wild demonstrate the effectiveness and generalization ability of our method. Our dataset and code are available at https://github.com/Ylinyuan/HDNet.
Xiao Lu 0002, Yulin Yuan, Lucai Wang, Xuanyu Zhou, Yimin Yang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Deep Optimized Broad Learning System for Applications in Tabular Data Recognition
abstract
The broad learning system (BLS) is a versatile and effective tool for analyzing tabular data. However, the rapid expansion of big data has resulted in an overwhelming amount of tabular data, necessitating the development of specialized tools for effective management and analysis. This article introduces an optimized BLS (OBLS) specifically tailored for big data analysis. In addition, a deep-optimized BLS (DOBLS) network is developed further to enhance the performance and efficiency of the OBLS. The main contributions of this article are: 1) by retracing the network's error from the output space to the latent space, the OBLS adjusts parameters in the feature and enhancement node layers. This process aims to achieve more resilient representations, resulting in improved performance; 2) the DOBLS is a multilayered structure consisting of multiple OBLSs, wherein each OBLS connects to the input and output layers, enabling direct data propagation. This design helps reduce information loss between layers, ensuring an efficient flow of information throughout the network; and 3) the proposed methods demonstrate robustness across various applications, including multiview feature embedding, one-class classification (OCC), camera model identification, electroencephalogram (EEG) signal processing, and radar signal analysis. Experimental results validate the effectiveness of the proposed models. To ensure reproducibility, the source code is available at https://github.com/1027051515/OBLS_DOBLS.
Wandong Zhang, Yimin Yang 0001, Q. M. Jonathan Wu, Tianlong Liu
IEEE Trans. Cybern.2
2024 Flex-DLD: Deep Low-Rank Decomposition Model With Flexible Priors for Hyperspectral Image Denoising and Restoration
abstract
Hyperspectral images (HSIs) are composed of hundreds of contiguous waveband images, offering a wealth of spatial and spectral information. However, the practical use of HSIs is often hindered by the presence of complicated noise caused by various factors such as non-uniform sensor response and dark current. Traditional methods for denoising HSIs rely on constrained optimization approaches, where selecting appropriate prior knowledge is critical for achieving satisfactory results. Nevertheless, these traditional algorithms are limited by hand-crafted priors, leaving room for improvement in their denoising performance. Recently, the supervised deep learning technique has emerged as a promising approach for HSI denoising. However, their requirement for paired training data and poor generalization ability on untrained noise distributions pose challenges in practical applications. In this paper, we design a novel algorithm by the synergism of optimization-based methods and deep learning techniques. Specifically, we introduce a plug-and-play Deep Low-rank Decomposition (DLD) model into the optimization framework. Furthermore, we propose an effective mechanism to incorporate traditional prior knowledge into the DLD model. Finally, we provide a detailed analysis of the optimization process and convergence of the proposed method. Empirical evaluations on various tasks, including hyperspectral image denoising and spectral compressive imaging, demonstrate the superiority of our approach over state-of-the-art methods.
Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Image Process.4
2024 Coarse-to-Fine Target Detection for HFSWR With Spatial-Frequency Analysis and Subnet Structure
abstract
High-frequency surface wave radar (HFSWR) is a powerful tool for ship detection and surveillance. blackHowever, the use of pre-trained deep learning (DL) networks for ship detection is challenging due to the limited training samples in HFSWR and the substantial differences between remote sensing images and everyday images. To tackle these issues, this paper proposes a coarse-to-fine target detection approach that combines traditional methods with DL, resulting in improved performance. The contributions of this work include: 1) a two-stage learning pipeline that integrates spatial-frequency analysis (SFA) with subnet-based neural networks, 2) an automatic linear thresholding algorithm for plausible target region (PTR) detection, and 3) a robust subnet neural network for fine target detection. The advantage of using SFA and subnet network is that the SFA reduces the need for extensive training data, while the subnet neural network excels at localizing ships even with limited training data. Experimental results on the HFSWR-RD dataset affirm the model's superior performance compared to rival algorithms.
Wandong Zhang, Yimin Yang 0001, Tianlong Liu
IEEE Trans. Multim.2
2024 Progressive Learning Model for Big Data Analysis Using Subnetwork and Moore-Penrose Inverse
abstract
Multilayer analytic learning plays a crucial role in data mining and representation learning. Nevertheless, most of them encounter inefficiencies in latent space encoding, resulting in less effective data representations. Aimed at addressing this limitation, this paper introduces two potent analytic learning methods, the progressive learning-based hierarchical subnet neural network (P-HSNN) and the robust P-HSNN (RP-HSNN). The contributions are as follows. First, two progressive learning astrategies based on subnetwork nodes are proposed. Second, the RP-HSNN is a Laplacian matrix-based algorithm, where label information and input representations are utilized simultaneously to optimize the subspace feature. Third, the dimension of subnetwork node is gradually increased. The global-level representation is formed by combining the features from the subnetworks. The model's convergence is thoroughly demonstrated through rigorous mathematical proof. Experimental analyses across various domains, spanning a wide range of training samples from 2,754 to 1,623,114, confirm the superior performance of the proposed algorithms over state-of-the-art multilayer analytic learning methods.
Wandong Zhang, Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Multim.2
2024 CLRNet: A Cross Locality Relation Network for Crowd Counting in Videos
abstract
In this article, we propose a new cross locality relation network (CLRNet) to generate high-quality crowd density maps for crowd counting in videos. Specifically, a cross locality relation module (CLRM) is proposed to enhance feature representations by modeling local dependencies of pixels between adjacent frames with an adapted local self-attention mechanism. First, different from the existing methods which measure similarity between pixels by dot product, a new adaptive cosine similarity is advanced to measure the relationship between two positions. Second, the traditional self-attention modules usually integrate the reconstructed features with the same weights for all the positions. However, crowd movement and background changes in a video sequence are uneven in real-life applications. As a consequence, it is inappropriate to treat all the positions in reconstructed features equally. To address this issue, a scene consistency attention map (SCAM) is developed to make CLRM pay more attention to the positions with strong correlations in adjacent frames. Furthermore, CLRM is incorporated into the network in a coarse-to-fine way to further enhance the representational capability of features. Experimental results demonstrate the effectiveness of our proposed CLRNet in comparison to the state-of-the-art methods on four public video datasets. The codes are available at: https://github.com/Amelie01/CLRNet.
Li Dong 0011, Haijun Zhang 0002, Jianghong Ma, Xiaofei Xu 0001, Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Neural Networks Learn. Syst.5
2024 Multimodal Moore-Penrose Inverse-Based Recomputation Framework for Big Data Analysis
abstract
Most multilayer Moore-Penrose inverse (MPI)-based neural networks, such as deep random vector functional link (RVFL), are structured with two separate stages: unsupervised feature encoding and supervised pattern classification. Once the unsupervised learning is finished, the latent encoding is fixed without supervised fine-tuning. However, in complex tasks such as handling the ImageNet dataset, there are often many more clues that can be directly encoded, while unsupervised learning, by definition, cannot know exactly what is useful for a certain task. There is a need to retrain the latent space representations in the supervised pattern classification stage to learn some clues that unsupervised learning has not yet been learned. In particular, the residual error in the output layer is pulled back to each hidden layer, and the parameters of the hidden layers are recalculated with MPI for more robust representations. In this article, a recomputation-based multilayer network using Moore-Penrose inverse (RML-MP) is developed. A sparse RML-MP (SRML-MP) model to boost the performance of RML-MP is then proposed. The experimental results with varying training samples (from 3k to 1.8 million) show that the proposed models provide higher Top-1 testing accuracy than most representation learning algorithms. For reproducibility, the source codes are available at https://github.com/W1AE/Retraining.
Wandong Zhang, Yimin Yang 0001, Q. M. Jonathan Wu, Tianlei Wang, Hui Zhang 0023
IEEE Trans. Neural Networks Learn. Syst.2
2023 D-BIN: A Generalized Disentangling Batch Instance Normalization for Domain Adaptation
abstract
Pattern recognition is significantly challenging in real-world scenarios by the variability of visual statistics. Therefore, most existing algorithms relying on the independent identically distributed assumption of training and test data suffer from the poor generalization capability of inference on unseen testing datasets. Although numerous studies, including domain discriminator or domain-invariant feature learning, are proposed to alleviate this problem, the data-driven property and lack of interpretation of their principle throw researchers and developers off. Consequently, this dilemma incurs us to rethink the essence of networks' generalization. An observation that visual patterns cannot be discriminative after style transfer inspires us to take careful consideration of the importance of style features and content features. Does the style information related to the domain bias? How to effectively disentangle content and style features across domains? In this article, we first investigate the effect of feature normalization on domain adaptation. Based on it, we propose a novel normalization module to adaptively leverage the propagated information through each channel and batch of features called disentangling batch instance normalization (D-BIN). In this module, we explicitly explore domain-specific and domaininvariant feature disentanglement. We maneuver contrastive learning to encourage images with the same semantics from different domains to have similar content representations while having dissimilar style representations. Furthermore, we construct both self-form and dual-form regularizers for preserving the mutual information (MI) between feature representations of the normalization layer in order to compensate for the loss of discriminative information and effectively match the distributions across domains. D-BIN and the constrained term can be simply plugged into state-of-the-art (SOTA) networks to improve their performance. In the end, experiments, including domain adaptation and generalization, conducted on different datasets have proven their effectiveness.
Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Weixing Peng, Wangdong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001
IEEE Trans. Cybern.7
2023 Semisupervised Manifold Regularization via a Subnetwork-Based Representation Learning Model
abstract
Semisupervised classification with a few labeled training samples is a challenging task in the area of data mining. Moore-Penrose inverse (MPI)-based manifold regularization (MR) is a widely used technique in tackling semisupervised classification. However, most of the existing MPI-based MR algorithms can only generate loosely connected feature encoding, which is generally less effective in data representation and feature learning. To alleviate this deficiency, we introduce a new semisupervised multilayer subnet neural network called SS-MSNN. The key contributions of this article are as follows: 1) a novel MPI-based MR model using the subnetwork structure is introduced. The subnet model is utilized to enrich the latent space representations iteratively; 2) a one-step training process to learn the discriminative encoding is proposed. The proposed SS-MSNN learns parameters by directly optimizing the entire network, accepting input from one end, and producing output at the other end; and 3) a new semisupervised dataset called HFSWR-RDE is built for this research. Experimental results on multiple domains show that the SS-MSNN achieves promising performance over the other semisupervised learning algorithms, demonstrating fast inference speed and better generalization ability.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001
IEEE Trans. Cybern.3
2023 Hierarchical One-Class Model With Subnetwork for Representation Learning and Outlier Detection
abstract
The multilayer one-class classification (OCC) frameworks have gained great traction in research on anomaly and outlier detection. However, most multilayer OCC algorithms suffer from loosely connected feature coding, affecting the ability of generated latent space to properly generate a highly discriminative representation between object classes. To alleviate this deficiency, two novel OCC frameworks, namely: 1) OCC structure using the subnetwork neural network (OC-SNN) and 2) maximum correntropy-based OC-SNN (MCOC-SNN), are proposed in this article. The novelties of this article are as follows: 1) the subnetwork is used to build the discriminative latent space; 2) the proposed models are one-step learning networks, instead of stacking feature learning blocks and final classification layer to recognize the input pattern; 3) unlike existing works which utilize mean square error (MSE) to learn low-dimensional features, the MCOC-SNN uses maximum correntropy criterion (MCC) for discriminative feature encoding; and 4) a brand-new OCC dataset, called CO-Mask, is built for this research. Experimental results on the visual classification domain with a varying number of training samples from 6131 to 513 061 demonstrate that the proposed OC-SNN and MCOC-SNN achieve superior performance compared to the existing multilayer OCC models. For reproducibility, the source codes are available at https://github.com/W1AE/OCC.
Wandong Zhang, Q. M. Jonathan Wu, W. G. Will Zhao, Haojin Deng, Yimin Yang 0001
IEEE Trans. Cybern.5
2023 A Two-Stage Hierarchical One-Class Classification Structure for HFSWR Ship-Target Detection
abstract
A high-frequency surface wave radar (HFSWR) is an effective tool for monitoring an exclusive economic zone (EEZ). However, the presence of diverse clutters and noises that contaminate the echo signals of the radar hinder its maritime surveillance. To address this issue, this paper presents a two-stage hierarchical one-class classification network (HOCN) designed specifically for ship-target detection in range-Doppler (RD) images. In Stage 1, the plausible region of interest (PROI) is extracted. This stage employs a dynamic threshold optimization strategy and Laplacian kernel to identify the potential regions of interest. In Stage 2, the proposed one-class deconvolutional-and-convolutional network (OC-DCNet) is utilized for fine detection of ship-targets. This stage comprises two sub-modules: the deconvolutional sub-module, which expands the input into a 2D matrix, and the convolutional sub-module, which classifies the input pattern as either a ship-target or a non-ship-target. The experimental results on a newly collected dataset called HFRD demonstrate the effectiveness of the proposed HFSWR ship-target detection algorithm.
Wandong Zhang, Yimin Yang 0001, Tianlong Liu, Q. M. Jonathan Wu
IEEE Trans. Geosci. Remote. Sens.2
2023 Sequential Order-Aware Coding-Based Robust Subspace Clustering for Human Action Recognition in Untrimmed Videos
abstract
Human action recognition (HAR) is one of most important tasks in video analysis. Since video clips distributed on networks are usually untrimmed, it is required to accurately segment a given untrimmed video into a set of action segments for HAR. As an unsupervised temporal segmentation technology, subspace clustering learns the codes from each video to construct an affinity graph, and then cuts the affinity graph to cluster the video into a set of action segments. However, most of the existing subspace clustering schemes not only ignore the sequential information of frames in code learning, but also the negative effects of noises when cutting the affinity graph, which lead to inferior performance. To address these issues, we propose a sequential order-aware coding-based robust subspace clustering (SOAC-RSC) scheme for HAR. By feeding the motion features of video frames into multi-layer neural networks, two expressive code matrices are learned in a sequential order-aware manner from unconstrained and constrained videos, respectively, to construct the corresponding affinity graphs. Then, with the consideration of the existence of noise effects, a simple yet robust cutting algorithm is proposed to cut the constructed affinity graphs to accurately obtain the action segments for HAR. The extensive experiments demonstrate the proposed SOAC-RSC scheme achieves the state-of-the-art performance on the datasets of Keck Gesture and Weizmann, and provides competitive performance on the other 6 public datasets such as UCF101 and URADL for HAR task, compared to the recent related approaches.
Zhili Zhou 0001, Chun Ding, Jin Li 0002, Eman Mohammadi, Guangcan Liu, Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Image Process.6
2022 Video Shadow Detection via Spatio-Temporal Interpolation Consistency Training
abstract
It is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results. In this paper, we address these challenges by proposing a Spatio-Temporal Interpolation Consistency Training (STICT) framework to rationally feed the unlabeled video frames together with the labeled images into an image shadow detection network training. Specifically, we propose the Spatial and Temporal ICT, in which we define two new interpolation schemes, i.e., the spatial interpolation and the temporal interpolation. We then derive the spatial and temporal interpolation consistency constraints accordingly for enhancing generalization in the pixel-wise classification task and for encouraging temporal consistent predictions, respectively. In addition, we design a Scale- Aware Network for multi-scale shadow knowledge learning in images, and propose a scale-consistency constraint to minimize the discrepancy among the predictions at different scales. Our proposed approach is extensively validated on the ViSha dataset and a self-annotated dataset. Experimental results show that, even without video labels, our approach is better than most state of the art supervised, semi-supervised or unsupervised image/video shadow detection methods and other methods in related tasks. Code and dataset are available at https://github.com/yihong-97/STICT.
Xiao Lu 0002, Yihong Cao, Chengjiang Long, Zipei Chen, Xuanyu Zhou, Yimin Yang 0001, Chunxia Xiao
CVPR7
2022 A Lightweight Self-Supervised Training Framework for Monocular Depth Estimation
abstract
Depth estimation attracts great interest in various sectors such as robotics, human computer interfaces, intelligent visual surveillance, and wearable augmented reality gear. Monocular depth estimation is of particular interest due to its low complexity and cost. Research in recent years was shifted away from supervised learning towards unsupervised or self-supervised approaches. While there have been great achievements, most of the research has focused on large heavy networks which are highly resource intensive that makes them unsuitable for systems with limited resources. We are particularly concerned about the increased complexity during training that current self-supervised approaches bring. In this paper, we propose a lightweight self-supervised training framework which utilizes computationally cheap methods to compute ground truth approximations. In particular, we utilize a stereo pair of images during training which are used to compute photometric reprojection loss and a disparity ground truth approximation. Due to the ground truth approximation, our framework is able to remove the need of pose estimation and the corresponding heavy prediction networks that current self-supervised methods have. In the experiments, we have demonstrated that our framework is capable of increasing the generator’s performance at a fraction of the size required by the current state-of-the-art self-supervised approach.
Tim Heydrich, Yimin Yang 0001, Shan Du 0001
ICASSP2
2022 A Novel Lightweight Network for Fast Monocular Depth Estimation
abstract
Depth estimation is of growing interest in many sectors, from robotics to wearable augmented reality gears. Monocular depth estimation attracts more attention due to its cost efficiency and low complexity. Most recent research has developed very large and resource intensive networks which are not suitable for small systems with limited resources. In this paper, we propose a lightweight network which leverages the advantages of dimension-wise convolutions and depthwise separable convolutions to reduce complexity in the architecture. In particular, the proposed depth estimation architecture utilizes a novel DICE unit-based encoder, optimized for a lightweight encoder-decoder structure. Furthermore, we propose a DICE unit-based decoder structure as well as an optimized depthwise separable convolution-based decoder. Both decoders follow a similar five-layer architecture. In the experiments, we have demonstrated the effectiveness of the proposed architecture as well as the comparison between the two proposed decoders. Our novel lightweight network has a significant decrease in both size and complexity at a marginal cost to accuracy when compared to other state-of-the-art lightweight networks.
Tim Heydrich, Yimin Yang 0001, Yu Liu 0004, Shan Du 0001
ICASSP2
2022 Fast Ship Detection With Spatial-Frequency Analysis and ANOVA-Based Feature Fusion
abstract
High-frequency surface wave radar (HFSWR) can be effectively used to detect ships in the exclusive economic zone. However, the ship signal is concealed and interfered with various clutter and background noise in the Doppler spectrum. In this letter, a range-Doppler (RD) image-based novel ship detection algorithm is proposed by exploiting spatial-frequency information and a unique feature fusion based on the analysis of variance. The algorithm subsumes three successive stages: Stage I—the plausible region of interest is captured, Stage II—the features from different sources are fused into one generalized feature space, and Stage III—an extreme learning machine-based classifier is utilized to localize the ships. Experimental results on challenging HFSWR-RD datasets demonstrate that the proposed algorithm has a competitive performance over other ship detection algorithms.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001, Akilan Thangarajah, W. G. Will Zhao, Qingzhong Li, Jiong Niu
IEEE Geosci. Remote. Sens. Lett.3
2022 Special issue on neural computing and applications 2021
Jingjing Cao, Yimin Yang 0001, Wun-She Yap, Zenghui Wang 0001
Neural Comput. Appl.3
2022 Multimodal Vigilance Estimation Using Deep Learning
abstract
The phenomenon of increasing accidents caused by reduced vigilance does exist. In the future, the high accuracy of vigilance estimation will play a significant role in public transportation safety. We propose a multimodal regression network that consists of multichannel deep autoencoders with subnetwork neurons (MCDAE$_{sn}$). After we define two thresholds of “0.35” and “0.70” from the percentage of eye closure, the output values are in the continuous range of 0–0.35, 0.36–0.70, and 0.71–1 representing the awake state, the tired state, and the drowsy state, respectively. To verify the efficiency of our strategy, we first applied the proposed approach to a single modality. Then, for the multimodality, since the complementary information between forehead electrooculography and electroencephalography features, we found the performance of the proposed approach using features fusion significantly improved, demonstrating the effectiveness and efficiency of our method.
Wei Wu 0022, Wei Sun 0028, Q. M. Jonathan Wu, Yimin Yang 0001, Hui Zhang 0023, Wei-Long Zheng, Bao-Liang Lu
IEEE Trans. Cybern.4
2022 HKPM: A Hierarchical Key-Area Perception Model for HFSWR Maritime Surveillance
abstract
High-frequency surface wave radar (HFSWR) has become the cornerstone of maritime surveillance because of its low-cost maintenance and coverage of wide area. However, when it comes to the extraction of key areas, such as vessel-target detection and vessel-path tracking, the HFSWR signal is strongly interfered by clutters and noise, which makes maritime surveillance a challenging task. This article proposes a hierarchical key-area perception model for maritime surveillance harnessing range-Doppler (RD) image from HFSWR, Laplacian kernel, a linear classifier (LC), and a subnet-based multilayer representation learning framework (SMRLF). First, a weak LC with a Laplacian kernel is utilized to capture the plausible vessel regions (PVRs). Then, a novel SMRLF is proposed to localize the vessel targets from the PVRs. To handle the noise, a maximum correntropy criterion with variable centers (MCC-VC) is incorporated in the subnet-based learning model. A thorough experimental analysis on cross-domain samples from radar dataset to scene classification dataset shows that the proposed HKPM performs competitively. The model shows a superior performance over most of the state-of-the-art vessel-target detection algorithms with a vessel-target detection accuracy of 94%. The extended analysis on image classification problem proves that the proposed model has great adaptivity and scalability.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001, Akilan Thangarajah, Ming Li 0057
IEEE Trans. Geosci. Remote. Sens.3
2022 Spatio-Temporal Feature Encoding for Traffic Accident Detection in VANET Environment
abstract
In the Vehicular Ad hoc Networks (VANET) environment, recognizing traffic accident events in the driving videos captured by vehicle-mounted cameras is an essential task. Generally, traffic accidents have a short duration in driving videos, and the backgrounds of driving videos are dynamic and complex. These make traffic accident detection quite challenging. To effectively and efficiently detect accidents from the driving videos, we propose an accident detection approach based on spatio–temporal feature encoding with a multilayer neural network. Specifically, the multilayer neural network is used to encode the temporal features of video for clustering the video frames. From the obtained frame clusters, we detect the border frames as the potential accident frames. Then, we capture and encode the spatial relationships of the objects detected from these potential accident frames to confirm whether these frames are accident frames. The extensive experiments demonstrate that the proposed approach achieves promising detection accuracy and efficiency for traffic accident detection, and meets the real-time detection requirement in the VANET environment.
Zhili Zhou 0001, Xiaohua Dong, Zhetao Li, Keping Yu, Chun Ding, Yimin Yang 0001
IEEE Trans. Intell. Transp. Syst.6
2021 A compensation-based optimization strategy for top dense layer training
Xiexing Feng, Q. M. Jonathan Wu, Yimin Yang 0001, Libo Cao
Neurocomputing3
2021 Echo state network with a global reversible autoencoder for time series classification
abstract
An echo state network (z) can provide an efficient dynamic solution for predicting time series problems. However, in most cases, ESN models are applied for predictions rather than classifications. The applications of ESN in time series classification (TSC) problems have yet to be fully studied. Moreover, the conventional randomly generated ESN is unlikely to be optimal because of the randomly generated input and reservoir weights, which are not always guaranteed to be optimal. Randomly generating all layer weights is improper, because a purely random layer might destroy the useful features. To overcome this disadvantage, this study provides a new input weight establishment framework of ESN based on autoencoder (AE) theory for TSC tasks. A global reversible AE (GRAE) algorithm is proposed to reestablish the random initialization input weights of the ESN. In existing ESN-AEs, the output weights obtained in the encoding process are directly reused as the initial input weights. By contrast, in GRAE, the reservoir layer with a reversible activation function is calculated by pulling the decoding layer output back and injecting it into the reservoir layer. Thus, feature learning is enriched by additional information, which results in improved performance. The current weights of the encoding layer are iteratively replaced by the decoding layer to ensure that the outputs of the GRAE are remarkably correlated with the input data. Visualization analyses and experiments of the input weights on a massive set of UCR time series datasets indicate that the proposed GRAE method can considerably improve the original two-layer ESN-based classifiers and the proposed GRAE-ESN classifier yields better performance compared with traditional state-of-the-art TSC classifiers. Furthermore, the proposed method can provide comparable performance and considerably faster training speed compared with three deep learning classifiers.
Heshan Wang, Q. M. Jonathan Wu, Dongshu Wang, Jianbin Xin, Yimin Yang 0001, Kunjie Yu
Inf. Sci.5
2021 Real-time stage-wise object tracking in traffic scenes: an online tracker selection method via deep reinforcement learning
Xiao Lu 0002, Yihong Cao, Xuanyu Zhou, Yimin Yang 0001
Neural Comput. Appl.5
2021 Non-iterative online sequential learning strategy for autoencoder and classifier
Adhri Nandini Paul, Peizhi Yan, Yimin Yang 0001, Hui Zhang 0023, Shan Du 0001, Q. M. Jonathan Wu
Neural Comput. Appl.3
2021 A Width-Growth Model With Subnetwork Nodes and Refinement Structure for Representation Learning and Image Classification
abstract
This article presents a new supervised multilayer subnetwork-based feature refinement and classification model for representation learning. The novelties of this algorithm are as follows: 1) different from most multilayer networks that go deeper with increased number of network layers, this work architects a model with wider subnetwork nodes; 2) the conventional classification methods adopt a separate search mechanism to derive a generalized feature space and to get the final cognition, but this work proposes a one-shot process to find the meaningful latent space and recognize the objects; and 3) the traditional feature representation and image classification approaches apply a unimodal feature coding, which suffers from lack of global knowledge. This work overcomes the pitfall through multimodal fusion that fuses various feature sources into one superstate encoding to achieve higher performance. A cross-domain experimental study on camera identification and image classification shows that the proposed method achieves superior performance compared to the existing models.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001, Akilan Thangarajah, Hui Zhang 0023
IEEE Trans. Ind. Informatics3
2021 MAMA Net: Multi-Scale Attention Memory Autoencoder Network for Anomaly Detection
abstract
Anomaly detection refers to the identification of cases that do not conform to the expected pattern, which takes a key role in diverse research areas and application domains. Most of existing methods can be summarized as anomaly object detection-based and reconstruction error-based techniques. However, due to the bottleneck of defining encompasses of real-world high-diversity outliers and inaccessible inference process, individually, most of them have not derived groundbreaking progress. To deal with those imperfectness, and motivated by memory-based decision-making and visual attention mechanism as a filter to select environmental information in human vision perceptual system, in this paper, we propose a Multi-scale Attention Memory with hash addressing Autoencoder network (MAMA Net) for anomaly detection. First, to overcome a battery of problems result from the restricted stationary receptive field of convolution operator, we coin the multi-scale global spatial attention block which can be straightforwardly plugged into any networks as sampling, upsampling and downsampling function. On account of its efficient features representation ability, networks can achieve competitive results with only several level blocks. Second, it's observed that traditional autoencoder can only learn an ambiguous model that also reconstructs anomalies "well" due to lack of constraints in training and inference process. To mitigate this challenge, we design a hash addressing memory module that proves abnormalities to produce higher reconstruction error for classification. In addition, we couple the mean square error (MSE) with Wasserstein loss to improve the encoding data distribution. Experiments on various datasets, including two different COVID-19 datasets and one brain MRI (RIDER) dataset prove the robustness and excellent generalization of the proposed MAMA Net.
Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Yimin Yang 0001, Xianen Zhou, Q. M. Jonathan Wu
IEEE Trans. Medical Imaging4
2021 Multimodel Feature Reinforcement Framework Using Moore-Penrose Inverse for Big Data Analysis
abstract
Fully connected representation learning (FCRL) is one of the widely used network structures in multimodel image classification frameworks. However, most FCRL-based structures, for instance, stacked autoencoder encode features and find the final cognition with separate building blocks, resulting in loosely connected feature representation. This article achieves a robust representation by considering a low-dimensional feature and the classifier model simultaneously. Thus, a new hierarchical subnetwork-based neural network (HSNN) is proposed in this article. The novelties of this framework are as follows: 1) it is an iterative learning process, instead of stacking separate blocks to obtain the discriminative encoding and the final classification results. In this sense, the optimal global features are generated; 2) it applies Moore-Penrose (MP) inverse-based batch-by-batch learning strategy to handle large-scale data sets, so that large data set, such as Place365 containing 1.8 million images, can be processed effectively. The experimental results on multiple domains with a varying number of training samples from ∼ 1 K to ∼ 2 M show that the proposed feature reinforcement framework achieves better generalization performance compared with most state-of-the-art FCRL methods.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001, Akilan Thangarajah
IEEE Trans. Neural Networks Learn. Syst.3
2020 An Autuencoder-based Data Augmentation Strategy for Generalization Improvement of DCNNs
Xiexing Feng, Q. M. Jonathan Wu, Yimin Yang 0001, Libo Cao
Neurocomputing3
2020 Wi-HSNN: A subnetwork-based encoding structure for dimension reduction and food classification via harnessing multi-CNN model high-level features
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001
Neurocomputing3
2020 Recomputation of the Dense Layers for Performance Improvement of DCNN
abstract
Gradient descent optimization of learning has become a paradigm for training deep convolutional neural networks (DCNN). However, utilizing other learning strategies in the training process of the DCNN has rarely been explored by the deep learning (DL) community. This serves as the motivation to introduce a non-iterative learning strategy to retrain neurons at the top dense or fully connected (FC) layers of DCNN, resulting in, higher performance. The proposed method exploits the Moore-Penrose Inverse to pull back the current residual error to each FC layer, generating well-generalized features. Further, the weights of each FC layers are recomputed according to the Moore-Penrose Inverse. We evaluate the proposed approach on six most widely accepted object recognition benchmark datasets: Scene-15, CIFAR-10, CIFAR-100, SUN-397, Places365, and ImageNet. The experimental results show that the proposed method obtains improvements over 30 state-of-the-art methods. Interestingly, it also indicates that any DCNN with the proposed method can provide better performance than the same network with its original Backpropagation (BP)-based training.
Yimin Yang 0001, Q. M. Jonathan Wu, Xiexing Feng, Akilan Thangarajah
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 A 3D CNN-LSTM-Based Image-to-Image Foreground Segmentation
abstract
The video-based separation of foreground (FG) and background (BG) has been widely studied due to its vital role in many applications, including intelligent transportation and video surveillance. Most of the existing algorithms are based on traditional computer vision techniques that perform pixel-level processing assuming that FG and BG possess distinct visual characteristics. Recently, state-of-the-art solutions exploit deep learning models targeted originally for image classification. Major drawbacks of such a strategy are the lacking delineation of FG regions due to missing temporal information as they segment the FG based on a single frame object detection strategy. To grapple with this issue, we excogitate a 3D convolutional neural network (3D CNN) with long short-term memory (LSTM) pipelines that harness seminal ideas, viz., fully convolutional networking, 3D transpose convolution, and residual feature flows. Thence, an FG-BG segmenter is implemented in an encoder-decoder fashion and trained on representative FG-BG segments. The model devises a strategy called double encoding and slow decoding, which fuses the learned spatio-temporal cues with appropriate feature maps both in the down-sampling and up-sampling paths for achieving well generalized FG object representation. Finally, from the Sigmoid confidence map generated by the 3D CNN-LSTM model, the FG is identified automatically by using Nobuyuki Otsu's method and an empirical global threshold. The analysis of experimental results via standard quantitative metrics on 16 benchmark datasets including both indoor and outdoor scenes validates that the proposed 3D CNN-LSTM achieves competitive performance in terms of figure of merit evaluated against prior and state-of-the-art methods. Besides, a failure analysis is conducted on 20 video sequences from the DAVIS 2016 dataset.
Akilan Thangarajah, Q. M. Jonathan Wu, Amin Safaei 0001, Jie Huo, Yimin Yang 0001
IEEE Trans. Intell. Transp. Syst.5
2020 Region-Level Visual Consistency Verification for Large-Scale Partial-Duplicate Image Search
abstract
Most recent large-scale image search approaches build on a bag-of-visual-words model, in which local features are quantized and then efficiently matched between images. However, the limited discriminability of local features and the BOW quantization errors cause a lot of mismatches between images, which limit search accuracy. To improve the accuracy, geometric verification is popularly adopted to identify geometrically consistent local matches for image search, but it is hard to directly use these matches to distinguish partial-duplicate images from non-partial-duplicate images. To address this issue, instead of simply identifying geometrically consistent matches, we propose a region-level visual consistency verification scheme to confirm whether there are visually consistent region (VCR) pairs between images for partial-duplicate search. Specifically, after the local feature matching, the potential VCRs are constructed via mapping the regions segmented from candidate images to a query image by utilizing the properties of the matched local features. Then, the compact gradient descriptor and convolutional neural network descriptor are extracted and matched between the potential VCRs to verify their visual consistency to determine whether they are VCRs. Moreover, two fast pruning algorithms are proposed to further improve efficiency. Extensive experiments demonstrate the proposed approach achieves higher accuracy than the state of the art and provide comparable efficiency for large-scale partial-duplicate search tasks.
Zhili Zhou 0001, Q. M. Jonathan Wu, Yimin Yang 0001, Xingming Sun
ACM Trans. Multim. Comput. Commun. Appl.3
2019 Hierarchical feature representation for unconstrained video analysis
Eman Mohammadi, Q. M. Jonathan Wu, Mehrdad Saif, Yimin Yang 0001
Neurocomputing4
2019 System-on-a-Chip (SoC)-Based Hardware Acceleration for an Online Sequential Extreme Learning Machine (OS-ELM)
abstract
Machine learning algorithms such as those for object classification in images, video content analysis, and human action recognition are used to extract meaningful information from data recorded by image sensors and cameras. Among the existing machine learning algorithms for such purposes, extreme learning machines (ELMs) and online sequential ELMs (OS-ELMs) are well known for their computational efficiency and performance when processing large datasets. The latter approach was derived from the ELM approach and optimized for real-time application. However, OS-ELM classifiers are computationally demanding, and the existing state-of-the-art computing platforms are not efficient enough for embedded systems, especially for applications with strict requirements in terms of low power consumption, high throughput, and low latency. This paper presents the implementation of an ELM/OS-ELM in a customized system-on-a-chip field-programmable gate array-based architecture to ensure efficient hardware acceleration. The acceleration process comprises parallel extraction, deep pipelining, and efficient shared memory communication.
Amin Safaei 0001, Q. M. Jonathan Wu, Akilan Thangarajah, Yimin Yang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 Features Combined From Hundreds of Midlayers: Hierarchical Networks With Subnetwork Nodes
abstract
In this paper, we believe that the mixed selectivity of neuron in the top layer encodes distributed information produced from other neurons to offer a significant computational advantage over recognition accuracy. Thus, this paper proposes a hierarchical network framework that the learning behaviors of features combined from hundreds of midlayers. First, a subnetwork neuron, which itself could be constructed by other nodes, is functional as a subspace features extractor. The top layer of a hierarchical network needs subspace features produced by the subnetwork neurons to get rid of factors that are not relevant, but at the same time, to recast the subspace features into a mapping space so that the hierarchical network can be processed to generate more reliable cognition. Second, this paper shows that with noniterative learning strategy, the proposed method has a wider and shallower structure, providing a significant role in generalization performance improvements. Hence, compared with other state-of-the-art methods, multiple channel features with the proposed method could provide a comparable or even better performance, which dramatically boosts the learning speed. Our experimental results show that our platform can provide a much better generalization performance than 55 other state-of-the-art methods.
Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Neural Networks Learn. Syst.1
2018 Fusion-based foreground enhancement for background subtraction using multivariate multi-model Gaussian distribution
Akilan Thangarajah, Q. M. Jonathan Wu, Yimin Yang 0001
Inf. Sci.3
2018 Autoencoder With Invertible Functions for Dimension Reduction and Image Reconstruction
abstract
The extreme learning machine (ELM), which was originally proposed for “generalized” single-hidden layer feedforward neural networks, provides efficient unified learning solutions for the applications of regression and classification. Although, it provides promising performance and robustness and has been used for various applications, the single-layer architecture possibly lacks the effectiveness when applied for natural signals. In order to over come this shortcoming, the following work indicates a new architecture based on multilayer network framework. The significant contribution of this paper are as follows: 1) unlike existing multilayer ELM, in which hidden nodes are obtained randomly, in this paper all hidden layers with invertible functions are calculated by pulling the network output back and putting it into hidden layers. Thus, the feature learning is enriched by additional information, which results in better performance; 2) in contrast to the existing multilayer network methods, which are usually efficient for classification applications, the proposed architecture is implemented for dimension reduction and image reconstruction; and 3) unlike other iterative learning-based deep networks (DL), the hidden layers of the proposed method are obtained via four steps. Therefore, it has much better learning efficiency than DL. Experimental results on 33 datasets indicate that, in comparison to the other existing dimension reduction techniques, the proposed method performs competitively better with fast training speeds.
Yimin Yang 0001, Q. M. Jonathan Wu, Yaonan Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2017 Effect of wavelet and hybrid classification on action recognition
abstract
Any action dataset may contain similar classes such as running, walking and jogging. Therefore, equivalent probabilities may be provided for different classes upon action classification. In this case, the classifier cannot indubitably assign a class to a given sample. To address this problem, we propose a new hybrid classifier to automatically compress the features and classify them using SVM with polynomial or sigmoid kernels. Furthermore, we hypothesize that motion saliency detection can strength the power of motion feature extraction in the bag of visual words framework (BoVW). To this end, we evaluate the effect of 3D-discrete wavelet transform (3D-DWT), as the preprocessing step, on motion feature extraction. The experimental results show that the proposed framework achieves promising results on KTH, Weizmann, and URADL datasets, and outperforms recent state-of-the-art approaches.
Eman Mohammadi, Q. M. Jonathan Wu, Yimin Yang 0001, Mehrdad Saif
ICIP3
2016 Learning Time-optimal Anti-swing Trajectories for Overhead Crane Systems
Xuebo Zhang 0003, Ruijie Xue, Yimin Yang 0001, Long Cheng 0001, Yongchun Fang
ISNN3
2016 Multilayer Extreme Learning Machine With Subnetwork Nodes for Representation Learning
abstract
The extreme learning machine (ELM), which was originally proposed for "generalized" single-hidden layer feedforward neural networks, provides efficient unified learning solutions for the applications of clustering, regression, and classification. It presents competitive accuracy with superb efficiency in many applications. However, ELM with subnetwork nodes architecture has not attracted much research attentions. Recently, many methods have been proposed for supervised/unsupervised dimension reduction or representation learning, but these methods normally only work for one type of problem. This paper studies the general architecture of multilayer ELM (ML-ELM) with subnetwork nodes, showing that: 1) the proposed method provides a representation learning platform with unsupervised/supervised and compressed/sparse representation learning and 2) experimental results on ten image datasets and 16 classification datasets show that, compared to other conventional feature learning methods, the proposed ML-ELM with subnetwork nodes performs competitively or much better than other feature learning methods.
Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Cybern.1
2016 Extreme Learning Machine With Subnetwork Hidden Nodes for Regression and Classification
abstract
As demonstrated earlier, the learning effectiveness and learning speed of single-hidden-layer feedforward neural networks are in general far slower than required, which has been a major bottleneck for many applications. Huang et al. proposed extreme learning machine (ELM) which improves the training speed by hundreds of times as compared to its predecessor learning techniques. This paper offers an ELM-based learning method that can grow subnetwork hidden nodes by pulling back residual network error to the hidden layer. Furthermore, the proposed method provides a similar or better generalization performance with remarkably fewer hidden nodes as compared to other ELM methods employing huge number of hidden nodes. Thus, the learning speed of the proposed technique is hundred times faster compared to other ELMs as well as to back propagation and support vector machines. The experimental validations for all methods are carried out on 32 data sets.
Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Cybern.1
2015 Data Partition Learning With Multiple Extreme Learning Machines
abstract
As demonstrated earlier, the learning accuracy of the single-layer-feedforward-network (SLFN) is generally far lower than expected, which has been a major bottleneck for many applications. In fact, for some large real problems, it is accepted that after tremendous learning time (within finite epochs), the network output error of SLFN will stop or reduce increasingly slowly. This report offers an extreme learning machine (ELM)-based learning method, referred to as the parent-offspring progressive learning method. The proposed method works by separating the data points into various parts, and then multiple ELMs learn and identify the clustered parts separately. The key advantages of the proposed algorithms as compared to the traditional supervised methods are twofold. First, it extends the ELM learning method from a single neural network to a multinetwork learning system, as the proposed multiELM method can approximate any target continuous function and classify disjointed regions. Second, the proposed method tends to deliver a similar or much better generalization performance than other learning methods. All the methods proposed in this paper are tested on both artificial and real datasets.
Yimin Yang 0001, Q. M. Jonathan Wu, Yaonan Wang 0001, Zeeshan Khawar Malik, Xiaofang Yuan
IEEE Trans. Cybern.1
2015 A Method to Calibrate Vehicle-Mounted Cameras Under Urban Traffic Scenes
abstract
We address the problem of vehicle-mounted camera calibration under urban traffic scenes regarding the fact that the traditional calibration methods are practically restricted, since the internal parameters should be calibrated in the laboratory and it is impossible for recalibration that resulted from the parameters drifting or re-focusing when driving on roads. In this paper, we propose to utilize the manual lines lying in Manhattan directions in the scenes to compute their corresponding vanishing points for camera calibration, as the urban traffic scenes are usually man-made and the important lines and signs for driving are typically lying in the Manhattan directions. For “Manhattan world” scenes, where there are plenty of lines lying in Manhattan directions, the lines in the scene are detected automatically, and the clusters corresponding to Manhattan directions are obtained using RANSAC-like methods. For the more general “quasi-Manhattan world” scenes, where only the lines in two directions can be found naturally, while the lines in the other direction are usually detected trivially or even can be hardly detected, we propose a method to estimate the lines in the third direction to improve the vanishing point estimation accuracy. The method proposed is tested on both two types of scenes, and the accuracy and practicability of this method are demonstrated. Furthermore, calibration experiments on both one image and multiple images are conducted, which show that the results can be more accurate when more images are used.
Yaonan Wang 0001, Xiao Lu 0002, Zhigang Ling, Yimin Yang 0001, Zhenjun Zhang, Kena Wang
IEEE Trans. Intell. Transp. Syst.4
2015 Progressive Learning Machine: A New Approach for General Hybrid System Approximation
abstract
As the most important property of neural networks (NNs), the universal approximation capability of NNs is widely used in many applications. However, this property is generally proven for continuous systems. Most industrial systems are hybrid systems (e.g., piecewise continuous), which is a significant limitation for real applications. Recently, many identification methods have been proposed for hybrid system approximation; however, these methods only operate in linear hybrid systems. In this paper, the progressive learning machine-a new learning algorithm based on multi-NNs-is proposed for general hybrid nonlinear/linear system approximation. This algorithm classifies hybrid systems into several continuous systems and can approximate any hybrid system with zero output error. The performance of the proposed learning method is demonstrated via numerical examples and with experimental data from real applications.
Yimin Yang 0001, Yaonan Wang 0001, Q. M. Jonathan Wu, Min Liu 0008
IEEE Trans. Neural Networks Learn. Syst.1
2014 Robust tracking control of uncertain dynamic nonholonomic systems using recurrent neural networks
Zhiqiang Miao, Yaonan Wang 0001, Yimin Yang 0001
Neurocomputing3
2013 Harmony search algorithm-based fuzzy-PID controller for electronic throttle valve
Xiaofang Yuan, Yaonan Wang 0001, Yimin Yang 0001
Neural Comput. Appl.4
2013 Neural network-based self-learning control for power transmission line deicing robot
Yimin Yang 0001, Yaonan Wang 0001, Xiaofang Yuan, Youhui Chen
Neural Comput. Appl.1
2013 Genetic algorithm-based adaptive fuzzy sliding mode controller for electronic throttle valve
Xiaofang Yuan, Yimin Yang 0001, Yaonan Wang 0001
Neural Comput. Appl.2
2013 Parallel Chaos Search Based Incremental Extreme Learning Machine
Yimin Yang 0001, Yaonan Wang 0001, Xiaofang Yuan
Neural Process. Lett.1
2012 Bidirectional Extreme Learning Machine for Regression Problem and Its Learning Effectiveness
abstract
It is clear that the learning effectiveness and learning speed of neural networks are in general far slower than required, which has been a major bottleneck for many applications. Recently, a simple and efficient learning method, referred to as extreme learning machine (ELM), was proposed by Huang , which has shown that, compared to some conventional methods, the training time of neural networks can be reduced by a thousand times. However, one of the open problems in ELM research is whether the number of hidden nodes can be further reduced without affecting learning effectiveness. This brief proposes a new learning algorithm, called bidirectional extreme learning machine (B-ELM), in which some hidden nodes are not randomly selected. In theory, this algorithm tends to reduce network output error to 0 at an extremely early learning stage. Furthermore, we find a relationship between the network output error and the network output weights in the proposed B-ELM. Simulation results demonstrate that the proposed method can be tens to hundreds of times faster than other incremental ELM algorithms.
Yimin Yang 0001, Yaonan Wang 0001, Xiaofang Yuan
IEEE Trans. Neural Networks Learn. Syst.1