Q. M. Jonathan Wu

dblp:w/QMJonathanWu · also Jonathan Wu 0001, Qingming Jonathan Wu · DBLP profile ↗
← Back
318ranked-venue papers
2as first author
95since 2021 · last 2026
0000-0002-5208-7975ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 127 · 2 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 104 · 23 since 2021Applied, interdisciplinary, general and emerging computing · 53 · 17 since 2021Human-computer interaction and ubiquitous computing · 32 · 2 since 2021Systems, architecture and hardware · 16 · 1 first-author · 5 since 2021Security and privacy · 14 · 12 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Computer networks · 4 · 2 since 2021
YearPublicationVenuePosition
2026 Contrastive adapter training and consensus knowledge distillation for multi-source-free domain adaptation in skin cancer diagnosis
abstract
Skin cancer diagnosis, particularly the differentiation of melanoma from benign nevi, is a vital yet challenging task due to the visual similarity between lesions. Although deep learning models such as convolutional neural networks (CNNs) and vision transformers (ViTs) have demonstrated promising performance, their effectiveness often deteriorates when applied to data from heterogeneous clinical sources. While conventional domain adaptation methods address domain shift, they require access to source data during adaptation, which is often infeasible due to privacy regulations. Multi-source-free unsupervised domain adaptation (MSFDA) addresses this limitation by leveraging multiple labeled source domains to generalize to an unlabeled target domain without requiring access to source data, making it suitable for privacy-sensitive medical settings. However, existing MSFDA methods rely on full backbone fine-tuning, leading to catastrophic forgetting and overfitting on small clinical datasets, and address domain shift at the aggregation stage without establishing a shared domain-invariant feature space. Furthermore, their reliance on hard pseudo-labels or confidence-weighted aggregation introduces noisy supervision signals under domain shift. To address these limitations, we propose CAT-CKD, consisting of two components: (1) contrastive adapter training (CAT), which trains lightweight ConvPass adapters within a frozen ViT backbone using supervised contrastive learning (SCL) to establish a shared domain-invariant feature space before source-specific model training, and (2) consensus knowledge distillation (CKD), which aggregates logits from multiple source models into a consensus supervisory signal and adapts a student model on unlabeled target data using KL divergence. Experiments on five publicly available skin lesion datasets show that CAT-CKD achieves an average AUROC of 86.1%, outperforming existing MSFDA methods while requiring only 4.3M trainable parameters. The code for this paper is available at https://github.com/A-Abedi/CAT_CKD.
Ali Abedi 0010, Q. M. Jonathan Wu, Ning Zhang 0007, Farhad Pourpanah
Artif. Intell. Medicine2
2026 One-shot federated unsupervised domain adaptation with Smoothed Knowledge Distillation and teacher refinement
abstract
Federated unsupervised domain adaptation (FUDA) addresses the challenge of adapting models to an unlabeled target domain using decentralized source domains while preserving data privacy. However, existing FUDA methods typically require multiple communication rounds, rely on complex aggregation strategies, and often struggle with noisy pseudo labels and inconsistent source knowledge. To address these challenges, we propose SKD-ETR, a novel one-shot FUDA framework that combines Smoothed Knowledge Distillation (SKD) and Exponential moving average-based Teacher Refinement (ETR). SKD trains a student model on the target domain using smoothed soft pseudo labels generated by the ensemble of source models. This reduces overconfidence, mitigates noise, and improves robustness. ETR further refines each source model by interpolating its parameters toward the student via exponential moving average (EMA), thereby transferring target-domain knowledge back to the teachers. This bidirectional refinement enhances pseudo-label quality and student generalization without additional communication overhead. SKD-ETR eliminates complex aggregation by initializing the student model randomly and performing a single-round distillation process. Extensive experiments on OfficeHome, Office-Caltech, and DomainNet demonstrate that SKD-ETR achieves competitive performance while being communication- and computation-efficient, and robust under noisy supervision. The code for this paper is available at https://github.com/A-Abedi/SKD_ETR . • A one-shot FUDA framework using knowledge distillation without source data access. • Smoothed pseudo labels reduce noise and improve student model robustness. • Teacher refinement via EMA enhances generalization with no extra communication. • Random target model initialization avoids biased aggregation and enables efficient adaptation.
Ali Abedi 0010, Q. M. Jonathan Wu, Ning Zhang 0007, Farhad Pourpanah
Eng. Appl. Artif. Intell.2
2026 Orthogonal momentum progressive subnetwork representation learning with feature fusion for surface wave radar target detection
Yaolong Lu, Gangsheng Li, Ling Zhang 0003, Jiong Niu, Q. M. Jonathan Wu
Eng. Appl. Artif. Intell.6
2026 A robust dual-pronged proactive defense framework against deepfakes via adversarial semi-fragile watermarking
Chengsheng Yuan 0001, Youqiang Cao, Zhili Zhou 0001, Zhangjie Fu 0001, Zhihua Xia, Q. M. Jonathan Wu
Expert Syst. Appl.6
2026 Facial sketch synthesis with multi-level guided latent diffusion model
Dan Lu 0006, Zhenxue Chen, Chengyun Liu, Q. M. Jonathan Wu
Neurocomputing5
2026 GGCN: Gait Recognition with Generate Network and Convolutional Neural Network
Hao Qin 0006, Zhenxue Chen, Qingqiang Guo, Q. M. Jonathan Wu, Mengxu Lu
J. Vis. Commun. Image Represent.4
2026 TSNUNet: Two-Stage Nested U-Network for salient object detection
Luna Sun, Zhenxue Chen, Xinming Zhu, Yu Bi, Chengyun Liu, Q. M. Jonathan Wu
J. Vis. Commun. Image Represent.6
2026 Mesh-based point cloud upsampling with 3D Gaussian splatting
Ning Zhang 0007, Xiaoyan Wang 0003, Q. M. Jonathan Wu
Mach. Vis. Appl.4
2026 Stochastic style perturbation modelling for visible-Infrared person re-Identification with severely modality imbalance
Jianyang Gu, Mingyu Wang 0004, Q. M. Jonathan Wu, Wei Jiang 0009
Neural Networks5
2026 Not All Extracted Information Is Credible: Toward Credibility-Aware Coverless Image Steganography in Distributed Cloud Services
abstract
Image steganography has emerged as a promising channel for transmitting task assignments, authentication credentials, and scheduling metadata across distributed cloud nodes without raising suspicion. However, existing steganography methods primarily focus on robustness against attacks while overlooking the credibility of the extracted information. In distributed cloud environments, blindly accepting incorrect or manipulated data can propagate faults across nodes, trigger inconsistent system states, or even cause cascading service failures. To address this, we propose C2IS, a Credibility-Aware Coverless Image Steganography framework designed for distributed cloud services. C2IS enables cloud nodes not only to extract hidden information, but also to assess its credibility before execution or propagation. Specifically, we propose a polar harmonic transform-based feature extraction and selection strategy that extracts both coarse-grained and fine-grained features. The coarse-grained features, characterized by their high stability under various attacks, are utilized to carry secret information. Meanwhile, the fine-grained features quantify the degree of feature perturbation to assess information credibility. Furthermore, we develop a double median thresholding-based hash mapping algorithm, which binarizes inter-image feature differences using the median, thereby significantly enhancing the diversity of the generated hash sequences. Extensive experimental results show that our method achieves a complete hash sequence length of up to 13 bits, enabling higher hiding capacity. Meanwhile, under various attacks, it exhibits stronger robustness than state-of-the-art methods and can accurately determine whether the extracted information contains errors, providing a trustworthy means of covert communication for distributed cloud services.
Bobiao Guo, Ping Ping, Yingchi Mao, Q. M. Jonathan Wu
IEEE Trans. Computers4
2026 Hydra-RAN Task 3: A Core-Independent AI Framework for Robust Intra-SRU Switching (ISS) for 6G Networks
Rafid I. Abd, Q. M. Jonathan Wu, Smaya Moher, Daniel J. Findley, Shuo Li 0003, Minji Phi, Kwang Soon Kim
IEEE Trans. Commun.2
2026 CL-ERDA: Certificateless Recoverable Data Auditing With Corruption Localization in Edge Computing
abstract
Edge computing paradigm deploys infrastructures on users' edge side, enabling data to be stored in edge storage systems. Thus edge computing finds a balance between remote cloud storage with high latency and local storage with limited resources. However, the frequent update of edge data increases the probability of data corrupting. To ensure availability and integrity of edge data without affecting future use, this paper proposes a certificateLess edge recoverable data auditing (CL-ERDA) approach with corruption localization. First, the certificateless network coding signature in CL-ERDA is utilized to generate the homomorphic authenticator, addressing public key certificate management issues while avoiding insecure key escrow problems. Then, CL-ERDA designs a two-phase auditing method, where only file-level aggregation proofs are respond in first phase of external check and the second phase of self-check is launched only when damaged data is detected, reducing communication overhead of the auditing. Moreover, CL-ERDA combines hierarchical check method with batch localization method to achieve locating efficiently all the corrupted data at one time rather than locating only one corrupted data at once. Finally, security analysis shows that the proposed scheme resists various attacks and supports the unforgeability of authenticators. Experimental results demonstrate that CL-ERDA is efficient in auditing and corruption localization.
Yongdong Ding, Dengzhi Liu, Jun Shen 0006, Haowen Tan, Q. M. Jonathan Wu
IEEE Trans. Dependable Secur. Comput.5
2026 Cross-Domain Heterogeneous Data Aggregation With Dynamic Group Key Agreement for Hybrid Satellite Networks
abstract
Hybrid satellite networks, composed of Low Earth Orbit (LEO) and Geostationary Earth Orbit (GEO) systems, are capable of ensuring seamless and flexible data exchange across entities. However, the inherent heterogeneity presents critical challenges for cross-domain data aggregation. Specifically, the following issues remain unsolved for current cross-domain data aggregation designs, including insufficient adaptability to the dynamic hierarchical network topologies, inflexible leader election for intra-domain data aggregation, and unsound privacy preservation for inter-domain data transmission. To overcome these limitations, a cross-domain heterogeneous data aggregation scheme for hybrid satellite networks is developed, providing dynamic group key agreement. First, an efficient re-authentication mechanism is constructed to ensure de-synchronization resistance. Meanwhile, a flexible and adaptive leader election strategy is proposed to enhance stable and seamless data exchange among dynamic LEO networks. Additionally, a secure dynamic cross-domain data transmission method is designed to resist eavesdropping and replay attacks. The security proofs and discussions regarding vital security properties are presented, while the performance analysis follows. Compared with the state-of-the-art, advantages in terms of security and performance properties can be proved.
Haowen Tan, Jian Shen 0001, Md. Zakirul Alam Bhuiyan, Q. M. Jonathan Wu
IEEE Trans. Dependable Secur. Comput.5
2026 An Efficient ASCON-Based Group Authentication and Key Agreement Scheme With Non-Linkability and Integrity Assurance for IIoTs
abstract
In recent years, numerous group authentication and key agreement (GAKA) schemes have been proposed for the Industrial Internet of Things (IIoTs). However, the frequent identity updates that are a feature of IIoT environments mean that existing schemes cannot maintain full-lifecycle device anonymity. Meanwhile, static integrity verification approaches cannot effectively cope with the dynamic and heterogeneous nature of industrial networks, making it imperative to design an adaptive data integrity protection mechanism to maintain reliable and resilient communication. Additionally, the simultaneous access of a large number of intelligent industrial devices to gateway nodes creates significant computational and communication challenges, while also increasing the risk of denial-of-service attacks and other concurrent threats. Consequently, current schemes are unable to strike an optimal balance between efficiency and security in large-scale IIoT deployments. In our scheme, we first designed an anonymous token mechanism to achieve sufficient randomness for unlinkability when communicating with the same gateway node across different time periods. Secondly, integrating associated data into the authenticated sponge construction (ASCON) encryption process ensures the legitimacy and integrity of the data during transmission. Third, we conduct rigorous security proofs under the widely accepted Algebraic Group Model (AGM) and Random Oracle Model (ROM), demonstrating that our group authentication mechanism is unforgeable against adaptive chosen-public-key and adaptive chosen-subspace attacks. Finally, the performance evaluation demonstrates that our proposed scheme achieves improvement in computational efficiency, while the communication overhead analysis confirms that the design remains both efficient.
Haowen Tan, Shenmin Gu, Jian Shen 0001, Md. Zakirul Alam Bhuiyan, Q. M. Jonathan Wu
IEEE Trans. Dependable Secur. Comput.6
2026 Blockchain-Assisted Conditional Anonymous Authentication and Adaptive Tree-Based Group Key Agreement for VANETs
abstract
Vehicular ad-hoc networks (VANETs), considered a pivotal component of intelligent transportation systems (ITS), are susceptible to both established and emerging security vulnerabilities. However, existing authenticated key management schemes fail to provide effective conditional anonymity during decentralized authentication process. Meanwhile, scalable and reliable vehicular pseudonym management is absent, resulting in potential privacy leakage. Furthermore, conventional group key agreement schemes inherently fail to properly accommodate the highly dynamic topological characteristics of vehicular environments, which significantly limits their practical applicability. To address these challenges, the blockchain-assisted anonymous authentication and tree-based group key agreement design is proposed in this paper. Firstly, the pairing-free decentralized authentication mechanism is designed to enable mutual authentication between vehicles and roadside units (RSUs). Secondly, the threshold-varying pseudonym management system is designed, leveraging secret sharing and smart contracts to ensure conditional privacy preservation. This mechanism utilizes the multi-RSU consensus to recover the user's real identity, enabling traceability of malicious entities. Thirdly, the self-balancing tree-based group key agreement mechanism is proposed, optimizing key generation efficiency in dynamic vehicular environments. Crucial security requirements can be satisfied via the security analysis, whereas the performance evaluation substantiates the superiority of the proposed scheme over existing approaches.
Haowen Tan, Jian Shen 0001, Pandi Vijayakumar, Sangman Moh, Q. M. Jonathan Wu
IEEE Trans. Dependable Secur. Comput.6
2026 A Blockchain-Based Efficient, Verifiable, and Weighted Multidimensional Data Aggregation Scheme in Smart Grids
abstract
The widespread deployment of smart grids has brought significant convenience to residential life. However, it also presents key challenges for data aggregation in smart grids.M1: The hierarchical structure of smart grid consumers (e.g., residential, industrial, commercial) requires differentiated allocation strategies to meet varying electricity demands while protecting consumer privacy.M2: The existing methods, such as superincreasing sequence, often face efficiency challenges, particularly when dealing with multidimensional data.M3: Smart meters continuously collect diverse power consumption data containing users' private information, which is vulnerable to tampering or loss, compromising data integrity and impacting power dispatch decisions. To address these challenges, this paper proposes a blockchain-based, efficient, verifiable, and weighted multidimensional data aggregation scheme for smart grids. First, a novel five-layer cloud-chain-assisted multiscenario data security aggregation model is proposed. Second, instead of using superincreasing sequences, we introduce the Chinese Remainder Theorem to process multidimensional data, thereby reducing communication complexity. Additionally, the property of quadratic reciprocity is leveraged to enhance the decryption method of the Paillier cryptosystem, reducing computational overhead. A weighted aggregation function is implemented to accurately aggregate data based on different user attributes. Furthermore, we propose two sample configurations to address distinct scenario requirements. Security analysis and experimental results demonstrate that the proposed scheme meets practical requirements in terms of both security and efficiency.
Chen Wang 0015, Shan Jiang 0023, Wenying Zheng, Q. M. Jonathan Wu, Debiao He
IEEE Trans. Dependable Secur. Comput.4
2026 Unleashing Cross-Domain Potential: Side-Channel Analysis with Autoencoder for Domain Adaptation
abstract
Deep learning based side-channel analysis (DL-SCA) has achieved remarkable success in recovering cryptographic keys from embedded devices by exploiting physical leakages such as power consumption and electromagnetic emissions, posing a serious threat to the security of cryptographic implementations. However, a major challenge arises in cross-device attacks, where a model trained on profiling devices cannot be directly applied to attack a different device. This is because domain discrepancies emerge from variations in chip architecture, manufacturing, operational conditions, and data acquisition methods between these two devices. Many existing DL-SCA schemes have not adequately addressed the challenges posed by the differences. Therefore, we propose a novel cross-device SCA framework based on an autoencoder, which leverages the encoder—decoder architecture to align feature distributions across devices in the latent space while simultaneously preserving discriminative leakage features through the reconstruction process. To achieve this, Maximum Mean Discrepancy (MMD) is integrated into the loss function and applied to the latent representations, effectively narrowing the distribution gap between profiling and attack devices. Operating in the latent space allows our approach to avoid the training instability of adversarial methods and provides an efficient end-to-end solution for domain alignment. Building on this framework, we further introduce three multi-domain adaptation methods. Experimental results, evaluated in terms of Partial Guessing Entropy (PGE), demonstrate that cross-device attacks can be effectively executed even with device discrepancies, with most keys being successfully recovered within 1,000 traces. Moreover, the proposed adaptation techniques significantly reduce the number of traces required for successful key recovery.
Haowen Tan, Jian Shen 0001, Pandi Vijayakumar, Sangman Moh, Q. M. Jonathan Wu
ACM Trans. Embed. Comput. Syst.6
2026 Quantifying and Overcoming the Bias Nature of Modality for Visible-Infrared Person Re-Identification
Daoxun Xia, Q. M. Jonathan Wu, Wei Jiang 0009
IEEE Trans. Inf. Forensics Secur.5
2025 Cross-Attention for AES Mode Variation in Side-Channel Analysis
abstract
Portability poses a significant challenge for Deep Learning (DL)-based profiling Side-Channel Analysis (SCA) on AES encryption, as attackers cannot always ensure that training and target samples use the same encryption mode. To address this, we propose an Unsupervised Domain Adaptation (UDA) DL-SCA framework for achieving effective and robust cross-encryption-mode attacks. By incorporating cross-attention and UDA techniques, our framework aligns high-dimensional input samples, reducing interference from encryption mode mismatches. Evaluation across five distinct AES modes demonstrates that our method achieves robust SCA performance without requiring prior knowledge or multiple labeled datasets for analysis.
Fanliang Hu, Jian Shen 0001, Q. M. Jonathan Wu
DAC4
2025 Diversity augmentation and multi-fuzzy label for semi-supervised semantic segmentation
abstract
Semantic segmentation aims to provide pixel-wise accurate predictions for images. Semi-supervised semantic segmentation aims to learn a semantic segmentation model using a limited number of labeled images and a large fraction of unlabeled images. Existing methods primarily focus on introducing additional models or complex training procedures but overlook the model itself and such complex strategies tend to discard many usable pixels, exacerbating the class imbalance problem . In this paper, we propose DAM for semi-supervised semantic segmentation, a simple yet effective method that mainly focuses on the inputs and outputs of the model itself. For the input component, we posit that diverse data augmentations can provide more semantic information . Therefore, we propose a method called Random Diversity Augmentations. Given an unlabeled image, we apply different triple-level data augmentations to provide more semantic information. For the output component, our approach is inspired by the fact that many unreliable predictions are confused only among the top classes rather than all classes, so we contend that fuzzy pixels can still provide valuable guidance to the model. Specifically, we select fuzzy pixels based on confidence and assign multi-fuzzy labels to these pixels for training the model, which allows us to leverage the information more effectively. Our straightforward DAM achieves new state-of-the-art performance on SSS different benchmarks. Code is available at https://github.com/Wang-zhenyan/DAM .
Zhenyan Wang, Zhenxue Chen, Chengyun Liu, Xinming Zhu, Q. M. Jonathan Wu
Neurocomputing6
2025 Synergy-driven multi-modal prompting for weakly supervised semantic segmentation
Chengyun Liu, Zhenyan Wang, Xiaona Peng, Zhenxue Chen, Q. M. Jonathan Wu
Neurocomputing6
2025 Three-factor authentication and key agreement protocol with collusion resistance in VANETs
Guanlin Pan, Haowen Tan, Wenying Zheng, Pandi Vijayakumar, Q. M. Jonathan Wu, Sivaraman Audithan
J. Inf. Secur. Appl.5
2025 Shipborne HFSWR Direction-Finding Method for Target Detection Based on Correction Matrix
abstract
Because of the influence of various factors, such as platform motion, antenna error, clutter, and noise interference, the direction finding (DF) of targets using shipborne high-frequency surface wave radar (HFSWR) becomes extremely difficult, which poses challenges for locating vessels at sea. To achieve more accurate DF for shipborne HFSWR, this letter proposes a method based on correction matrices, dividing the factors that cause DF errors into two categories: platform motion and interference from other sources. After calculating two corresponding correction matrices, a two-step correction is performed on the steering vector of the array to reduce DF errors. Field data experiments validate the performance of the correction matrix-based method for DOA estimation.
Cheng Wang 0047, Ling Zhang 0003, Gangsheng Li, Q. M. Jonathan Wu
IEEE Geosci. Remote. Sens. Lett.5
2025 Shipborne HFSWR Sea Clutter Suppression Method Based on MultiDomain Information Synergy
abstract
1 Abstract-Due to the integrated effect of many factors including non-uniform wave motion and shipboard platform motion, the echo signals received by shipborne high-frequency surface wave radar (HFSWR) often suffer from issues such as sea clutter spreading. A large number of targets are submerged by sea clutter, creating a detection blind area. To address this problem, a novel sea clutter suppression method based on multi-domain information synergy is proposed. The proposed method first identifies the broadening region of sea clutter by its characteristics. The multi-domain spectrum is then constructed using a narrow beam forming method. Afterwards, the Laplace kernel function is employed to screen the sea clutter regions to obtain the plausible region of interest (PROI). Ultimately, we integrate all PROIs and obtain sea clutter suppression results. Field data from shipborne HFSWR and validation results from the automatic identification system (AIS) demonstrate that the proposed method can effectively suppress sea clutter, increase the signal-to-clutter ratio (SCR), and achieve better target detection performance.
Jiangnan Zhong, Ling Zhang 0003, Gangsheng Li, Q. M. Jonathan Wu
IEEE Geosci. Remote. Sens. Lett.5
2025 Less Traces Are All It Takes: Efficient Side-Channel Analysis on AES
abstract
In cryptography, side-channel analysis (SCA) is a technique used to recover cryptographic keys by examining the physical leakages that occur during the operation of cryptographic devices. Recent advancements in deep learning (DL) have greatly enhanced the extraction of crucial information from intricate leakage patterns. A considerable amount of research is dedicated to studying the SubByte (SB) operations of the advanced encryption standard (AES). This is because the SB process, which generates numerous transitions between 0s and 1s during encryption, results in significant energy leakage. However, traditional analysis models primarily focus on the initial round of SB operations in AES, which are less effective on mobile terminals where it is difficult to collect enough signals. These models often neglect additional operations and subsequent rounds, thus providing limited insights from small datasets. Consequently, this limitation has a direct impact on the accuracy and efficiency of key recovery. Our study uses$\rho $-test analysis to show that significant leakage occurs not only during the S-box operation but also during the AddRoundKey (AR) phase of AES. To address these challenges, we propose a new SCA method, that is, optimized for small sample sizes. This method includes a new comprehensive round trace labeling algorithm, which simultaneously analyzes the SB and AR stages of each AES round. Additionally, we introduce the peak precise localization algorithm to accurately identify the points of energy leakage during each encryption round. Our experiments, conducted with power and electromagnetic (EM) datasets from the STM32F303 microcontroller, demonstrate that our method can reliably recover keys with as few as 20 traces. These results highlight the enhanced capability of our method in handling the complexities of small sample datasets in cryptographic analysis.
Zhiyuan Xiao, Chen Wang 0015, Jian Shen 0001, Q. M. Jonathan Wu, Debiao He
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Fast Transfer Learning Method Using Random Layer Freezing and Feature Refinement Strategy
abstract
Recently, Moore-Penrose inverse (MPI)-based parameter fine-tuning of fully connected (FC) layers in pretrained deep convolutional neural networks (DCNNs) has emerged within the inductive transfer learning (ITL) paradigm. However, this approach has not gained significant traction in practical applications due to its stringent computational requirements. This work addresses this issue through a novel fast retraining strategy that enhances applicability of the MPI-based ITL. Specifically, during each retraining epoch, a random layer freezing protocol is utilized to manage the number of layers undergoing feature refinement. Additionally, this work incorporates an MPI-based approach for refining the trainable parameters of FC layers under batch processing, contributing to expedited convergence. Extensive experiments on several ImageNet pretrained benchmark DCNNs demonstrate that the proposed ITL achieves competitive performance with excellent convergence speed compared to conventional ITL methods. For instance, the proposed strategy converges nearly 1.5 times faster than retraining the ImageNet pretrained ResNet-50 using stochastic gradient descent with momentum (SGDM).
Wandong Zhang, Yimin Yang 0001, Akilan Thangarajah, Q. M. Jonathan Wu, Tianlong Liu
IEEE Trans. Cybern.4
2025 BM-PDA: Blockchain Based Multifunctional Private-Preserving Data Aggregation for e-Health Systems
abstract
Secure aggregation of medical data enables detailed data analysis and informed medical decision-making in e-health systems, optimizing data resources utilization and enhancing service quality and decision accuracy. However, the collection of large volumes of medical data poses a significant risk of privacy leakage. Most existing privacy-preserving data aggregation schemes focus on additive aggregation of single or multi-dimensional data, which greatly limits their applicability. This article introduces a blockchain-based multifunctional data aggregation (BM-PDA) scheme for e-health systems. First, BM-PDA supports overall aggregation queries of data samples and can compute the maximum and minimum values within these samples. Second, it enables selective data aggregation queries based on various user attributes. Furthermore, analysis shows that integrating these two algorithms protects both user’s private data and attribute data. Performance evaluations indicate that the computational and communication costs are acceptable, demonstrating the scheme’s practical applicability.
Chen Wang 0015, Jian Shen 0001, Q. M. Jonathan Wu, Debiao He
IEEE Trans. Dependable Secur. Comput.4
2025 DGADM-GIS: Deterministic Guided Additive Diffusion Model for Generative Image Steganography
abstract
In recent years, generative steganography has witnessed remarkable progress in the field of covert communication. It leverages techniques such as generative adversarial networks (GANs) or flow-based generative models (GLOW) to generate stego images. However, these approaches often grapple with the dilemma of achieving optimal steganographic capacity while ensuring the accurate extraction of hidden information. Additionally, the models occasionally still generate low-quality images that are highly vulnerable to detection by steganalysis tools. To tackle the aforementioned challenges and enhance the overall performance of generative image steganography, this paper proposes the deterministic guided additive diffusion model for generative image steganography (DGADM-GIS). Initially, we devise a reversible mapping function that is used for deterministic guided by a provided secret message, and then construct a secret latent Gaussian vector. Moreover, the proposed DGADM-GIS framework designs an additive sampling method based on the superposition principle of normal distribution to obtain a Gaussian vector that satisfies independent, random and obeys the standard normal distribution, which is transformed to a stego image in a way of maintaining the distribution by the diffusion model. Furthermore, we conduct error analysis experiments on our proposed scheme and derive methods to enhance the accuracy of secret information extraction. The experimental results show that our proposed steganographic method exhibits robust resistance to steganalysis. When embedding 3 bits of secret information per pixel, it achieves nearly 100% extraction accuracy.
Chengsheng Yuan 0001, Zhaonan Ji, Xinting Li, Zhili Zhou 0001, Zhihua Xia, Q. M. Jonathan Wu
IEEE Trans. Dependable Secur. Comput.6
2025 Few-Shot Facial Sketch Synthesis via Progressive Domain Gap Reduction
abstract
Facial sketch synthesis (FSS) has advanced significantly in recent years, but challenges remain in few-shot settings. Some few-shot learning methods can convert photos (source domain) into sketches of a specified style (target sketch domain). However, they overlook the available samples of other sketch styles (non-target sketch domains). We argue that the information in these samples can help the model enhance its mapping ability from the source domain to the target domain. This paper proposes a progressive domain gap reduction (PDGR) method for few-shot facial sketch synthesis, which consists of three stages: teacher training, knowledge distillation, and intra-domain few-shot adaptation. In the first stage, we adapt a pretrained StyleGAN to a non-target sketch domain with more available samples than the target sketch domain. To generate diverse and high-quality sketches, we employ a dual-discriminator adversarial mechanism to guide the model in focusing on the overall structure and style, as well as multi-scale details and textures. In the second stage, the knowledge from StyleGAN is transferred to a U-Net for more efficient image translation. In the third stage, we adapt the output of the U-Net from the non-target sketch domain to the target sketch domain in few-shot settings. To alleviate overfitting, preserve individual characteristics, and enhance detail representation, we leverage the FFHQ dataset to construct dual training paths and design a domain-directional triple loss. Experiments show that PDGR significantly outperforms previous few-shot learning methods and even outperforms the state-of-the-art FSS methods trained on the full dataset.
Dan Lu 0006, Zhenxue Chen, Chengyun Liu, Q. M. Jonathan Wu
IEEE Trans. Inf. Forensics Secur.6
2025 CoMix: Collaborative Mixed Learning via Style Fuzzy Normalization for Visible-Infrared Person Re-Identification
abstract
Visible–infrared person re-identification (VI-ReID) focuses on accurately matching individuals across different imaging modalities. Existing studies focus on generating modality-consistent images at the pixel level through the use of generative adversarial networks (GANs) to mitigate the impact of modality discrepancies. However, these methods face significant challenges in overcoming the limitation that synthesized samples from different modalities may suffer from semantic distortion. In this work, we propose an online one-stage style fuzzy normalization (SFN) method to generate modality-fuzzy features in the latent space while regularizing the model’s predictions. Specifically, SFN adaptively mixes the feature statistics of two random modality instances of the same identity in a single forward pass during training. In this process, to enhance the richness of modality interaction information, we design a novel causality balance loss, which enforces the generated fuzzy features to be independent of their initial modality while simultaneously encouraging them to align more closely with the other modality. Furthermore, we introduce an identity-aware consistency loss to regularize the predictions between the original and SFN-generated features to ensure semantic consistency. In contrast to prior work, SFN is a plug-and-play module that does not rely on any generative-based models, making it highly adaptable to various network architectures. Extensive experiments were performed on three public cross-modality datasets to ensure fair and reliable comparisons. The empirical results demonstrate the clear superiority of our method over previous state-of-the-art methods.
Jianyang Gu, Mingyu Wang 0004, Q. M. Jonathan Wu, Wei Jiang 0009
IEEE Trans. Syst. Man Cybern. Syst.5
2024 Efficient object detector via dynamic prior and dynamic feature fusion
abstract
Abstract Sparse R-CNN is a new paradigm of object detection, which predicts objects in a sparse way. However, there are some limitations in Sparse R-CNN. One is the presence of weak prior information caused by fixed learnable proposal boxes and features across different images, necessitating excessive iterations for the model to refine its predictions; the other is the inadequate exploitation of multi-scale information, leading to the sub-optimal detection performance. Thus, building upon Sparse R-CNN, we propose an efficient detector that incorporates dynamic prior and dynamic feature fusion, called $D^{2}$-Det. In particular, for the dynamic prior part, a prior information generator module dynamically generates proposal features and boxes as the dynamic prior for different images to alleviate the inference-inefficient iterative refinement process of predictions, and we further propose the class scores decoupling method to reduce the computation overhead. Furthermore, for the dynamic feature fusion part, we develop a novel lightweight multi-scale feature fusion module, which dynamically aggregates features from all layers for each proposal box, enabling adaptive feature fusion and improving detection precision by nearly 2 AP. Experiments show that $D^{2}$-Det can achieve 46.6 AP on COCO 2017 with fewer computations for the backbone ResNet50, surpassing most of the state-of-the-art detectors.
Zhili Zhou 0001, Gaobo Yang, Q. M. Jonathan Wu
Comput. J.5
2024 Self-supervised adversarial adaptation network for breast cancer detection
abstract
Breast cancer is the most commonly diagnosed cancer worldwide, and early detection is essential for reducing mortality rates. Digital mammography is currently the best standard for early detection, as it can assist physicians in treating the disease. However, inaccurate diagnoses from mammography are common and can lead to patients undergoing unnecessary tests and treatments. To address this challenge, deep-learning techniques have shown promising results in improving the accuracy and reliability of breast cancer detection. However, existing methods face two primary challenges: the lack of the annotated data, and the inability to adapt to new data domains. In this paper, we propose SelfAdaptNet to address these issues. Specifically, SelfAdaptNet employs self-supervised learning techniques, such as Bootstrap Your Own Latent (BYOL) and Simple Framework for Learning of Visual Representations (SimCLR), to tackle the problem of limited annotated data. Additionally, the adversarial technique is used to address the problem of domain shift. By successfully reducing domain disparities, this strategy enhances the model’s adaptability and robustness across a variety of clinical scenarios. Overall, our contributions offer a more effective and flexible approach for early breast cancer detection, and experimental results demonstrate that SelfAdaptNet can produce promising results as compared with other methods.
Mahnoosh Torabi, Amir Hosein Rasouli, Q. M. Jonathan Wu, Weipeng Cao, Farhad Pourpanah
Eng. Appl. Artif. Intell.3
2024 A real-time anchor-free defect detector with global and local feature enhancement for surface defect detection
Qing Liu 0004, Min Liu 0002, Q. M. Jonathan Wu, Weiming Shen 0001
Expert Syst. Appl.3
2024 BNDCNet: Bilateral nonlocal decoupled convergence network for semantic segmentation
Mengting Ye, Zhenxue Chen, Kaili Yu, Longcheng Liu, Q. M. Jonathan Wu
J. Vis. Commun. Image Represent.6
2024 Category-based depth incorporation for salient object ranking
Hanxiao Zhai, Zhenxue Chen, Chengyun Liu, Huibin Bai, Q. M. Jonathan Wu
J. Vis. Commun. Image Represent.5
2024 AdaptiveGait: adaptive feature fusion network for gait recognition
Zhenxue Chen, Chengyun Liu, Jiyang Chen, Q. M. Jonathan Wu
Multim. Tools Appl.6
2024 Methods for class-imbalanced learning with support vector machines: a review and an empirical evaluation
Salim Rezvani, Farhad Pourpanah, Chee Peng Lim, Q. M. Jonathan Wu
Soft Comput.4
2024 Generative Steganography Based on Long Readable Text Generation
abstract
Text steganography has received a lot of attention in the application of covert communication. How to ensure desirable capacity and imperceptibility has become a key issue in text steganography. There are two typical approaches, i.e., text-selection-based steganography and text-generation-based steganography. However, the text-selection-based approaches generally have the very low hidden capacity and are not applicable in practical scenarios. Although the text-generation-based approaches can embed secret messages with higher capacity during text generation, they are prone to semantic incoherence and semantic errors when generating long texts. To address the abovementioned issues, this article proposes a novel text steganography based on long readable text generation. It first determines the topic of the stego-text according to the scenarios of the communication parties. Then, the plug and play language model (PPLM) is explored to generate the long readable stego-text conforming to the topic with semantic coherency. A given secret message is hidden during text generation by selecting proper words in an established embeddable candidate word pool (ECWP). Establishing the ECWP prevents the language model (LM) from selecting words with low probability in the text generation, thereby avoiding the generation of low-quality or even grammatically incorrect stego-text. Experimental results show that the proposed approach significantly increases hidden capacity while maintaining good imperceptibility compared with the existing approaches.
Zhili Zhou 0001, Chinmay Chakraborty, Meimin Wang, Q. M. Jonathan Wu, Xingming Sun, Keping Yu
IEEE Trans. Comput. Soc. Syst.5
2024 TCTL-Net: Template-Free Color Transfer Learning for Self-Attention Driven Underwater Image Enhancement
abstract
Vision is an important source of information for underwater observations, but underwater images commonly suffer severe visual degradation due to the complexity of the underwater imaging environment and wavelength-dependent absorption effects. There is an urgent need for underwater image enhancement techniques to improve the visual quality of underwater images. Due to the scarcity of high-quality paired training samples, underwater image enhancement based on deep learning has never achieved success similar to other vision tasks. Instead of learning complicated distortion-to-clear mappings with deep networks, we design a template-free color transfer learning framework for predicting transfer parameters, which are more easily captured and described. In addition, we add attention-driven modules to learn differentiated transfer parameters for more flexible and robust enhancement. We verify the effectiveness of our method on multiple publicly available datasets and show its efficiency in enhancing high-resolution images. The source code and the trained models are available on the project homepage: https://trentqq.github.io/TCTL-Net.html.
Kunqian Li, Qi Qi 0008, Chi Yan, Kun Sun 0002, Q. M. Jonathan Wu
IEEE Trans. Circuits Syst. Video Technol.6
2024 Deep Optimized Broad Learning System for Applications in Tabular Data Recognition
abstract
The broad learning system (BLS) is a versatile and effective tool for analyzing tabular data. However, the rapid expansion of big data has resulted in an overwhelming amount of tabular data, necessitating the development of specialized tools for effective management and analysis. This article introduces an optimized BLS (OBLS) specifically tailored for big data analysis. In addition, a deep-optimized BLS (DOBLS) network is developed further to enhance the performance and efficiency of the OBLS. The main contributions of this article are: 1) by retracing the network's error from the output space to the latent space, the OBLS adjusts parameters in the feature and enhancement node layers. This process aims to achieve more resilient representations, resulting in improved performance; 2) the DOBLS is a multilayered structure consisting of multiple OBLSs, wherein each OBLS connects to the input and output layers, enabling direct data propagation. This design helps reduce information loss between layers, ensuring an efficient flow of information throughout the network; and 3) the proposed methods demonstrate robustness across various applications, including multiview feature embedding, one-class classification (OCC), camera model identification, electroencephalogram (EEG) signal processing, and radar signal analysis. Experimental results validate the effectiveness of the proposed models. To ensure reproducibility, the source code is available at https://github.com/1027051515/OBLS_DOBLS.
Wandong Zhang, Yimin Yang 0001, Q. M. Jonathan Wu, Tianlong Liu
IEEE Trans. Cybern.3
2024 Flex-DLD: Deep Low-Rank Decomposition Model With Flexible Priors for Hyperspectral Image Denoising and Restoration
abstract
Hyperspectral images (HSIs) are composed of hundreds of contiguous waveband images, offering a wealth of spatial and spectral information. However, the practical use of HSIs is often hindered by the presence of complicated noise caused by various factors such as non-uniform sensor response and dark current. Traditional methods for denoising HSIs rely on constrained optimization approaches, where selecting appropriate prior knowledge is critical for achieving satisfactory results. Nevertheless, these traditional algorithms are limited by hand-crafted priors, leaving room for improvement in their denoising performance. Recently, the supervised deep learning technique has emerged as a promising approach for HSI denoising. However, their requirement for paired training data and poor generalization ability on untrained noise distributions pose challenges in practical applications. In this paper, we design a novel algorithm by the synergism of optimization-based methods and deep learning techniques. Specifically, we introduce a plug-and-play Deep Low-rank Decomposition (DLD) model into the optimization framework. Furthermore, we propose an effective mechanism to incorporate traditional prior knowledge into the DLD model. Finally, we provide a detailed analysis of the optimization process and convergence of the proposed method. Empirical evaluations on various tasks, including hyperspectral image denoising and spectral compressive imaging, demonstrate the superiority of our approach over state-of-the-art methods.
Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Image Process.5
2024 Review of Accident Detection Methods Using Dashcam Videos for Autonomous Driving Vehicles
abstract
The need for a reliable system to detect high-risk incidents in complex settings like roadways, which are infrequent but potentially dangerous, has arisen due to the occurrence of rare hazardous events. This system would empower self-driving cars to function autonomously over extended periods without human involvement. Among these hazardous occurrences, accidents have received the least attention due to their rarity and diverse nature. Recently, dashboard cameras (dashcams) have gained recognition in academic circles as a cost-effective and accessible solution to enhance the safety of autonomous vehicles when handling accidents, since they are now commonly found in most vehicles. This review presents the progression of concepts in this domain, tracing its development from early ideas to cutting-edge techniques. It categorizes these approaches into supervised, self-supervised, and unsupervised learning. Furthermore, the review thoroughly examines evaluation criteria and available datasets, providing a comprehensive comparison of the strengths and limitations of different methods. Ultimately, the review proposes potential avenues for future research in this field.
Arash Rocky, Q. M. Jonathan Wu, Wandong Zhang
IEEE Trans. Intell. Transp. Syst.2
2024 ARES: On Adversarial Robustness Enhancement for Image Steganographic Cost Learning
abstract
Taking the steganalytic discriminators as the adversaries, the existing Generative Adversarial Networks (GAN)-based steganographic approaches learn the implicit cost functions to measure the embedding distortion for steganography. However, the steganalytic discriminators in these approaches are trained by the stego-samples with insufficient diversity, and their network structures offer very limited representational capacity. As a result, these steganalytic discriminators will not exhibit robustness to various steganographic patterns, which causes learning suboptimal cost functions, thus compromising the anti-steganalysis capability. To address this issue, we propose a novel GAN-based steganographic approach, in which the Diversified Inverse-Adversarial Training (DIAT) strategy and the Steganalytic Feature Attention (SteFA) structure are designed to train a robust steganalytic discriminator. Specifically, the DIAT strategy provides the steganalytic discriminator with an expanded feature space by generating diversified adversarial stego-samples; the SteFA structure enables the steganalytic discriminator to capture more various steganalytic features by employing the channel-attention mechanism on higher-order statistics. Consequently, the steganalytic discriminator can build a more precise decision boundary to make it more robust, which facilitates learning a superior steganographic cost function. Extensive experiments demonstrate that the proposed steganographic approach achieves promising anti-steganalysis capability over the state-of-the-arts under the same embedding payloads.
Zhili Zhou 0001, Ruohan Meng, Shaowei Wang 0003, Hongyang Yan, Q. M. Jonathan Wu
IEEE Trans. Multim.6
2024 Progressive Learning Model for Big Data Analysis Using Subnetwork and Moore-Penrose Inverse
abstract
Multilayer analytic learning plays a crucial role in data mining and representation learning. Nevertheless, most of them encounter inefficiencies in latent space encoding, resulting in less effective data representations. Aimed at addressing this limitation, this paper introduces two potent analytic learning methods, the progressive learning-based hierarchical subnet neural network (P-HSNN) and the robust P-HSNN (RP-HSNN). The contributions are as follows. First, two progressive learning astrategies based on subnetwork nodes are proposed. Second, the RP-HSNN is a Laplacian matrix-based algorithm, where label information and input representations are utilized simultaneously to optimize the subspace feature. Third, the dimension of subnetwork node is gradually increased. The global-level representation is formed by combining the features from the subnetworks. The model's convergence is thoroughly demonstrated through rigorous mathematical proof. Experimental analyses across various domains, spanning a wide range of training samples from 2,754 to 1,623,114, confirm the superior performance of the proposed algorithms over state-of-the-art multilayer analytic learning methods.
Wandong Zhang, Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Multim.4
2024 CLRNet: A Cross Locality Relation Network for Crowd Counting in Videos
abstract
In this article, we propose a new cross locality relation network (CLRNet) to generate high-quality crowd density maps for crowd counting in videos. Specifically, a cross locality relation module (CLRM) is proposed to enhance feature representations by modeling local dependencies of pixels between adjacent frames with an adapted local self-attention mechanism. First, different from the existing methods which measure similarity between pixels by dot product, a new adaptive cosine similarity is advanced to measure the relationship between two positions. Second, the traditional self-attention modules usually integrate the reconstructed features with the same weights for all the positions. However, crowd movement and background changes in a video sequence are uneven in real-life applications. As a consequence, it is inappropriate to treat all the positions in reconstructed features equally. To address this issue, a scene consistency attention map (SCAM) is developed to make CLRM pay more attention to the positions with strong correlations in adjacent frames. Furthermore, CLRM is incorporated into the network in a coarse-to-fine way to further enhance the representational capability of features. Experimental results demonstrate the effectiveness of our proposed CLRNet in comparison to the state-of-the-art methods on four public video datasets. The codes are available at: https://github.com/Amelie01/CLRNet.
Li Dong 0011, Haijun Zhang 0002, Jianghong Ma, Xiaofei Xu 0001, Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Neural Networks Learn. Syst.6
2024 Multimodal Moore-Penrose Inverse-Based Recomputation Framework for Big Data Analysis
abstract
Most multilayer Moore-Penrose inverse (MPI)-based neural networks, such as deep random vector functional link (RVFL), are structured with two separate stages: unsupervised feature encoding and supervised pattern classification. Once the unsupervised learning is finished, the latent encoding is fixed without supervised fine-tuning. However, in complex tasks such as handling the ImageNet dataset, there are often many more clues that can be directly encoded, while unsupervised learning, by definition, cannot know exactly what is useful for a certain task. There is a need to retrain the latent space representations in the supervised pattern classification stage to learn some clues that unsupervised learning has not yet been learned. In particular, the residual error in the output layer is pulled back to each hidden layer, and the parameters of the hidden layers are recalculated with MPI for more robust representations. In this article, a recomputation-based multilayer network using Moore-Penrose inverse (RML-MP) is developed. A sparse RML-MP (SRML-MP) model to boost the performance of RML-MP is then proposed. The experimental results with varying training samples (from 3k to 1.8 million) show that the proposed models provide higher Top-1 testing accuracy than most representation learning algorithms. For reproducibility, the source codes are available at https://github.com/W1AE/Retraining.
Wandong Zhang, Yimin Yang 0001, Q. M. Jonathan Wu, Tianlei Wang, Hui Zhang 0023
IEEE Trans. Neural Networks Learn. Syst.3
2023 Unsupervised self-attention lightweight photo-to-sketch synthesis with feature maps
Kunru Zhong, Zhenxue Chen, Chengyun Liu, Q. M. Jonathan Wu, Shuchao Duan
J. Vis. Commun. Image Represent.4
2023 A cover selection-based reversible data hiding method by learning cross-modal hashing
Liming Zou, Jiande Sun 0001, Wenbo Wan, Jing Li 0046, Q. M. Jonathan Wu
Multim. Tools Appl.5
2023 A Review of Generalized Zero-Shot Learning Methods
abstract
Generalized zero-shot learning (GZSL) aims to train a model for classifying data samples under the condition that some output classes are unknown during supervised learning. To address this challenging task, GZSL leverages semantic information of the seen (source) and unseen (target) classes to bridge the gap between both seen and unseen classes. Since its introduction, many GZSL models have been formulated. In this review paper, we present a comprehensive review on GZSL. First, we provide an overview of GZSL including the problems and challenges. Then, we introduce a hierarchical categorization for the GZSL methods and discuss the representative methods in each category. In addition, we discuss the available benchmark data sets and applications of GZSL, along with a discussion on the research gaps and directions for future investigations.
Farhad Pourpanah, Moloud Abdar, Xinlei Zhou, Ran Wang 0001, Chee Peng Lim, Xizhao Wang, Q. M. Jonathan Wu
IEEE Trans. Pattern Anal. Mach. Intell.8
2023 D-BIN: A Generalized Disentangling Batch Instance Normalization for Domain Adaptation
abstract
Pattern recognition is significantly challenging in real-world scenarios by the variability of visual statistics. Therefore, most existing algorithms relying on the independent identically distributed assumption of training and test data suffer from the poor generalization capability of inference on unseen testing datasets. Although numerous studies, including domain discriminator or domain-invariant feature learning, are proposed to alleviate this problem, the data-driven property and lack of interpretation of their principle throw researchers and developers off. Consequently, this dilemma incurs us to rethink the essence of networks' generalization. An observation that visual patterns cannot be discriminative after style transfer inspires us to take careful consideration of the importance of style features and content features. Does the style information related to the domain bias? How to effectively disentangle content and style features across domains? In this article, we first investigate the effect of feature normalization on domain adaptation. Based on it, we propose a novel normalization module to adaptively leverage the propagated information through each channel and batch of features called disentangling batch instance normalization (D-BIN). In this module, we explicitly explore domain-specific and domaininvariant feature disentanglement. We maneuver contrastive learning to encourage images with the same semantics from different domains to have similar content representations while having dissimilar style representations. Furthermore, we construct both self-form and dual-form regularizers for preserving the mutual information (MI) between feature representations of the normalization layer in order to compensate for the loss of discriminative information and effectively match the distributions across domains. D-BIN and the constrained term can be simply plugged into state-of-the-art (SOTA) networks to improve their performance. In the end, experiments, including domain adaptation and generalization, conducted on different datasets have proven their effectiveness.
Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Weixing Peng, Wangdong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001
IEEE Trans. Cybern.6
2023 Semisupervised Manifold Regularization via a Subnetwork-Based Representation Learning Model
abstract
Semisupervised classification with a few labeled training samples is a challenging task in the area of data mining. Moore-Penrose inverse (MPI)-based manifold regularization (MR) is a widely used technique in tackling semisupervised classification. However, most of the existing MPI-based MR algorithms can only generate loosely connected feature encoding, which is generally less effective in data representation and feature learning. To alleviate this deficiency, we introduce a new semisupervised multilayer subnet neural network called SS-MSNN. The key contributions of this article are as follows: 1) a novel MPI-based MR model using the subnetwork structure is introduced. The subnet model is utilized to enrich the latent space representations iteratively; 2) a one-step training process to learn the discriminative encoding is proposed. The proposed SS-MSNN learns parameters by directly optimizing the entire network, accepting input from one end, and producing output at the other end; and 3) a new semisupervised dataset called HFSWR-RDE is built for this research. Experimental results on multiple domains show that the SS-MSNN achieves promising performance over the other semisupervised learning algorithms, demonstrating fast inference speed and better generalization ability.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001
IEEE Trans. Cybern.2
2023 Hierarchical One-Class Model With Subnetwork for Representation Learning and Outlier Detection
abstract
The multilayer one-class classification (OCC) frameworks have gained great traction in research on anomaly and outlier detection. However, most multilayer OCC algorithms suffer from loosely connected feature coding, affecting the ability of generated latent space to properly generate a highly discriminative representation between object classes. To alleviate this deficiency, two novel OCC frameworks, namely: 1) OCC structure using the subnetwork neural network (OC-SNN) and 2) maximum correntropy-based OC-SNN (MCOC-SNN), are proposed in this article. The novelties of this article are as follows: 1) the subnetwork is used to build the discriminative latent space; 2) the proposed models are one-step learning networks, instead of stacking feature learning blocks and final classification layer to recognize the input pattern; 3) unlike existing works which utilize mean square error (MSE) to learn low-dimensional features, the MCOC-SNN uses maximum correntropy criterion (MCC) for discriminative feature encoding; and 4) a brand-new OCC dataset, called CO-Mask, is built for this research. Experimental results on the visual classification domain with a varying number of training samples from 6131 to 513 061 demonstrate that the proposed OC-SNN and MCOC-SNN achieve superior performance compared to the existing multilayer OCC models. For reproducibility, the source codes are available at https://github.com/W1AE/OCC.
Wandong Zhang, Q. M. Jonathan Wu, W. G. Will Zhao, Haojin Deng, Yimin Yang 0001
IEEE Trans. Cybern.2
2023 Secret-to-Image Reversible Transformation for Generative Steganography
abstract
Recently, generative steganography that transforms secret information to a generated image has been a promising technique to resist steganalysis detection. However, due to the inefficiency and irreversibility of the secret-to-image transformation, it is hard to find a good trade-off between the information hiding capacity and extraction accuracy. To address this issue, we propose a secret-to-image reversible transformation (S2IRT) scheme for generative steganography. The proposed S2IRT scheme is based on a generative model, i.e., Glow model, which enables a bijective-mapping between latent space with multivariate Gaussian distribution and image space with a complex distribution. In the process of S2I transformation, guided by a given secret message, we construct a latent vector and then map it to a generated image by the Glow model, so that the secret message is finally transformed to the generated image. Owing to good efficiency and reversibility of S2IRT scheme, the proposed steganographic approach achieves both high hiding capacity and accurate extraction of secret message from generated image. Furthermore, a separate encoding-based S2IRT (SE-S2IRT) scheme is also proposed to improve the robustness to common image attacks. The experiments demonstrate the proposed steganographic approaches can achieve high hiding capacity (up to 4bpp) and accurate information extraction (almost 100% accuracy rate) simultaneously, while maintaining desirable anti-detectability and imperceptibility.
Zhili Zhou 0001, Yuecheng Su, Jin Li 0002, Keping Yu, Q. M. Jonathan Wu, Zhangjie Fu 0001, Yun Q. Shi 0001
IEEE Trans. Dependable Secur. Comput.5
2023 Accurate Direction Finding for Shipborne HFSWR Through Platform Motion Compensation
abstract
Shipborne high-frequency surface wave radar (HFSWR) plays a crucial role in ship target detection due to its mobility and flexibility in marine surveillance. However, accurate direction finding (DF) of shipborne HFSWR targets is extremely challenging due to the complex marine environment, which causes the platform to oscillate on six degrees of freedom (6-DOF) in addition to its forward motion, affecting the DF of the target. To solve this problem, we propose a motion compensation direction finding (MC-DF) method for shipborne HFSWR. In this paper, we model the motion of the platform and analyze the impact of the 6-DOF oscillation motion and forward motion on the target azimuth. We derive the steering vector of the radar array after motion compensation and use digital beamforming to accurately conduct DF of the target. The parameters required for motion compensation are provided by the inertial navigation system on the platform. The effectiveness of this method was verified through in-situ experiments on shipborne HFSWR.
Cheng Wang 0047, Ling Zhang 0003, Jiong Niu, Gangsheng Li, Q. M. Jonathan Wu
IEEE Trans. Geosci. Remote. Sens.5
2023 A Two-Stage Hierarchical One-Class Classification Structure for HFSWR Ship-Target Detection
abstract
A high-frequency surface wave radar (HFSWR) is an effective tool for monitoring an exclusive economic zone (EEZ). However, the presence of diverse clutters and noises that contaminate the echo signals of the radar hinder its maritime surveillance. To address this issue, this paper presents a two-stage hierarchical one-class classification network (HOCN) designed specifically for ship-target detection in range-Doppler (RD) images. In Stage 1, the plausible region of interest (PROI) is extracted. This stage employs a dynamic threshold optimization strategy and Laplacian kernel to identify the potential regions of interest. In Stage 2, the proposed one-class deconvolutional-and-convolutional network (OC-DCNet) is utilized for fine detection of ship-targets. This stage comprises two sub-modules: the deconvolutional sub-module, which expands the input into a 2D matrix, and the convolutional sub-module, which classifies the input pattern as either a ship-target or a non-ship-target. The experimental results on a newly collected dataset called HFRD demonstrate the effectiveness of the proposed HFSWR ship-target detection algorithm.
Wandong Zhang, Yimin Yang 0001, Tianlong Liu, Q. M. Jonathan Wu
IEEE Trans. Geosci. Remote. Sens.4
2023 Sequential Order-Aware Coding-Based Robust Subspace Clustering for Human Action Recognition in Untrimmed Videos
abstract
Human action recognition (HAR) is one of most important tasks in video analysis. Since video clips distributed on networks are usually untrimmed, it is required to accurately segment a given untrimmed video into a set of action segments for HAR. As an unsupervised temporal segmentation technology, subspace clustering learns the codes from each video to construct an affinity graph, and then cuts the affinity graph to cluster the video into a set of action segments. However, most of the existing subspace clustering schemes not only ignore the sequential information of frames in code learning, but also the negative effects of noises when cutting the affinity graph, which lead to inferior performance. To address these issues, we propose a sequential order-aware coding-based robust subspace clustering (SOAC-RSC) scheme for HAR. By feeding the motion features of video frames into multi-layer neural networks, two expressive code matrices are learned in a sequential order-aware manner from unconstrained and constrained videos, respectively, to construct the corresponding affinity graphs. Then, with the consideration of the existence of noise effects, a simple yet robust cutting algorithm is proposed to cut the constructed affinity graphs to accurately obtain the action segments for HAR. The extensive experiments demonstrate the proposed SOAC-RSC scheme achieves the state-of-the-art performance on the datasets of Keck Gesture and Weizmann, and provides competitive performance on the other 6 public datasets such as UCF101 and URADL for HAR task, compared to the recent related approaches.
Zhili Zhou 0001, Chun Ding, Jin Li 0002, Eman Mohammadi, Guangcan Liu, Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Image Process.7
2023 DETA: A Point-Based Tracker With Deformable Transformer and Task-Aligned Learning
abstract
Current point-based trackers are usually implemented by the following two branches: a classification branch for predicting the target candidate locations and a regression branch for regressing the tracking box, which may lead to a spatial misalignment between the two tasks. Meanwhile, they ignore a meaningful exploration on how to define positive and negative samples during training and explicit border information for accurate box prediction. In this research, we investigate the key issues of point-based trackers and unlock their key limitations. First, we design a novel task-aligned component and a new loss function, named task-aligned loss, to learn the alignment of the classification and regression tasks. Second, we introduce a border alignment (BorderAlign) component in both the classification and regression branches to effectively exploit the border features of a tracking target. Third, we develop an adaptive training sample assignment (ATSA) to adaptively divide the positive and negative samples based on the statistical characteristics of the tracking object. Finally, a deformable transformer is developed to enhance the representations of search features and explore rich temporal contexts among video frames. Extensive experimental results demonstrate that the proposed tracker achieves state-of-the-art performance on six tracking benchmark datasets.
Kai Yang 0018, Haijun Zhang 0002, Feng Gao 0015, Jianyang Shi, Q. M. Jonathan Wu
IEEE Trans. Multim.6
2023 Cycle Consistency Based Pseudo Label and Fine Alignment for Unsupervised Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a well-labeled source domain to an unlabeled target domain with a correlative distribution. Numerous existing approaches process this hard nut by directly matching the marginal distribution between two domains, which confront the obstacle of rough alignment and blurred decision boundary. Recent advances in UDA introduce target pseudo-label and subdomain adaptation to reduce misalignment and distribution discrepancy. Whereas, they frequently ignore that the production of target pseudo-label is so dependent on the source-trained classifier, which without reasonable restriction to discriminate generated pseudo-label is whether confident. Meanwhile, many methods in the subdomain alignment metric ignore exploring the potential distribution discrepancy between same-class samples of the intra-domain. To address these two issues simultaneously, this paper proposes a Cycle Consistency based Pseudo Label and Fine Alignment (CCPLFA) approach for UDA. In particular, firstly, a novel cycle-consistency based pseudo label module is designed, which is a simple yet effective way to alleviate the noise of pseudo labels and improve their semantic correctness. Secondly, we develop a Fine-Alignment distribution matching metric. Which can maximize the feature distribution density of intra-class cross-domains and not overlook the distribution structure of the global aspect. Comprehensive experiment results on four benchmarks demonstrate the capability of plug and play and the well generalization performance of our proposed method.
Hui Zhang 0023, Junkun Tang, Yihong Cao, Yurong Chen 0003, Yaonan Wang 0001, Q. M. Jonathan Wu
IEEE Trans. Multim.6
2023 Robust Coverless Image Steganography Based on Neglected Coverless Image Dataset Construction
abstract
Most of the existing image selection-based coverless image steganography methods mainly focus on improving the capacity and robustness under the assumption that the corresponding dataset is available. But they ignore how to successfully construct the coverless image dataset, which is the foundation of such methods and has a critical impact on the capacity. In this paper, a coverless image steganography is proposed that considers how to efficiently construct the coverless image dataset. In the proposed method, the CNN-based deep hash is extracted from the image and a specific mapping rule is designed to map the high-dimensional deep hash to the low-dimensional secret message. In addition, an unsupervised clustering algorithm is adopted to construct the coverless image dataset, which makes the construction of the coverless image dataset efficient and improves the robustness of the proposed steganography method. To our best knowledge, this is the first attempt to improve the construction efficiency of the coverless image dataset in the field of coverless image steganography. Experimental results show that the construction of a large coverless image dataset is feasible and reliable, and the proposed method has better robustness and higher dataset utilization rate compared with the state-of-the-art methods.
Liming Zou, Jing Li 0046, Wenbo Wan, Q. M. Jonathan Wu, Jiande Sun 0001
IEEE Trans. Multim.4
2023 A multi-scale threshold integration encoding strategy for texture classification
Yibing Li 0001, Q. M. Jonathan Wu
Vis. Comput.3
2022 Multisensor fusion estimation of nonlinear systems with intermittent observations and heavy-tailed noises
Bo Xiao 0006, Q. M. Jonathan Wu
Sci. China Inf. Sci.2
2022 Underwater image enhancement with latent consistency learning-based color transfer
abstract
Abstract Due to the inevitable wavelength‐dependent light absorption and forward/backward scattering, underwater images usually suffer severe color distortion and are hazy. It has become quite necessary to improve the visual quality of underwater images for both underwater observation and operation. Traditional enhancement methods and existing deep learning‐based approaches to underwater image enhancement usually produce unsatisfactory results for photographs taken in complicated, wild underwater scenes. In such scenes, complex and diverse degradation‐enhancement mappings are often difficult to model, especially since there are very limited samples available for learning. Inspired by the success of color‐transfer techniques, it is found that clear template image‐assisted color transfer is a promising strategy for underwater image enhancement, including not only color correction but also contrast and visibility improvement. Therefore, instead of directly learning the complex deep enhancement models, it is proposed to select proper color‐transfer templates by learning the latent consistency between the templates and the raw underwater images. The proposed new enhancement strategy alleviates the problem caused by incomplete color‐correction models and provides more stable enhancements by utilizing color transfer with consideration of global color distribution consistency and local visual contrast. Comprehensive experiments conducted on UIEB, RUIE, URPC and SQUID datasets demonstrate the good performance and great potential of the proposed new underwater image enhancement strategy.
Qi Qi 0008, Q. M. Jonathan Wu, Kunqian Li
IET Image Process.4
2022 Cross-graph reference structure based pruning and edge context information for graph matching
Md Shakil Ahamed Shohag, Xiuyang Zhao, Q. M. Jonathan Wu, Farhad Pourpanah
Inf. Sci.3
2022 DOA Estimation for HFSWR Target Based on PSO-ELM
abstract
High-frequency surface wave radar (HFSWR) plays an important role in vessel target surveillance. However, HFSWR’s inaccuracy of azimuth estimation caused by wide beams severely limits its detection ability. To solve this problem, a novel direction of arrival (DOA) estimation method based on extreme learning machine optimized by particle swarm optimization (PSO-ELM) is proposed to improve azimuth estimation accuracy for HFSWR. This method can obtain the optimal solution without searching the whole angle range of HFSWR. Specifically, PSO optimizes the input weight and hidden layer bias of ELM to obtain optimal parameters for improving the estimation performance. Based on the optimized parameters, the ELM network can give an optimal azimuth estimation in the sense of least squares and minimal norm. The sample sets used for PSO-ELM training are obtained by matching the points detected by HFSWR with the target points reported by an automatic identification system (AIS) on the range–Doppler (RD) spectra. The performance of DOA estimation is verified by field HFSWR data. The experimental results show that the new method has lower root-mean-square error and higher computational efficiency in comparison to the typical DOA estimation methods, such as digital beam forming (DBF) and multiple signal classification (MUSIC). It also uses the machine learning methods, such as back propagation neural network (BPNN) and support vector regression (SVR).
Ling Zhang 0003, Chenlu Shi, Jiong Niu, Yonggang Ji, Q. M. Jonathan Wu
IEEE Geosci. Remote. Sens. Lett.5
2022 Fast Ship Detection With Spatial-Frequency Analysis and ANOVA-Based Feature Fusion
abstract
High-frequency surface wave radar (HFSWR) can be effectively used to detect ships in the exclusive economic zone. However, the ship signal is concealed and interfered with various clutter and background noise in the Doppler spectrum. In this letter, a range-Doppler (RD) image-based novel ship detection algorithm is proposed by exploiting spatial-frequency information and a unique feature fusion based on the analysis of variance. The algorithm subsumes three successive stages: Stage I—the plausible region of interest is captured, Stage II—the features from different sources are fused into one generalized feature space, and Stage III—an extreme learning machine-based classifier is utilized to localize the ships. Experimental results on challenging HFSWR-RD datasets demonstrate that the proposed algorithm has a competitive performance over other ship detection algorithms.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001, Akilan Thangarajah, W. G. Will Zhao, Qingzhong Li, Jiong Niu
IEEE Geosci. Remote. Sens. Lett.2
2022 Improved edge-guided network for single image super-resolution
Zhenxue Chen, Q. M. Jonathan Wu, Xianming Li
Multim. Tools Appl.3
2022 CATFPN: Adaptive Feature Pyramid With Scale-Wise Concatenation and Self-Attention
abstract
It is a typical problem in the field of object detection to simultaneously detect objects with large scale variation in one image. Recently proposed state-of-the-art object detectors generally learn pyramidal feature representation to deal with the scale variation, which has been proved effective via various feature pyramid networks. However, the majority of the feature pyramid networks based on heuristic feature fusion strategies may be suboptimal, as excess human guidance will restrict the self-learning of deep neural networks. An adaptive feature pyramid is bound to provide a significant performance boost. In this paper, we propose a novel feature pyramid network named CATFPN that consists of Scale-Wise Feature Concatenation (SWFC) module and Global Context (GC) block. The SWFC module evenly distributes semantic features for each feature layer and the GC block introduces a self-attention mechanism. As a feature pyramid network, the CATFPN can be applied to any detector based on multi-scale features. We adopt the CATFPN in typical RetinaNet and Faster R-CNN detector models, without bells and whistles, achieving 1.1% AP and 0.7% AP improvements over FPN on the MS COCO benchmark, respectively. Our competitive performance reported on the test-dev subset of COCO achieves 42.3% AP.
Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu, Hui Yuan 0001, Weikai He
IEEE Trans. Circuits Syst. Video Technol.3
2022 Underwater Image Co-Enhancement With Correlation Feature Matching and Joint Learning
abstract
In underwater scenes, degraded underwater images caused by wavelength-dependent light absorption and scattering present huge challenges to vision tasks. Underwater image enhancement has attracted much attention due to the significance of vision-based applications in marine engineering and underwater robotics. Numerous underwater image enhancement algorithms have been proposed in the last few years. However, almost all existing approaches focus only on the enhancement of independent images. Considering that images photographed in the same underwater scene usually share similar degradation, related images can provide rich complementary information for each other’s enhancement. In this paper, we propose an Underwater Image Co-enhancement Network (UICoE-Net) based on an encoder-decoder Siamese architecture. For joint learning, we introduced correlation feature matching units into the multiple layers of our Siamese encoder-decoder structure in order to communicate the mutual correlation of the two branches. Extensive experiments using the Underwater Image Enhancement Benchmark (UIEB), Underwater Image Co-enhancement Dataset (UICoD) collected from an underwater video dataset with ground-truth reference and Stereo Quantitative Underwater Image Dataset (SQUID) dataset demonstrate the effectiveness of our method.
Qi Qi 0008, Q. M. Jonathan Wu, Kunqian Li, Xin Luan, Dalei Song
IEEE Trans. Circuits Syst. Video Technol.4
2022 RPNet: Gait Recognition With Relationships Between Each Body-Parts
abstract
At present, many studies have shown that partitioning the gait sequence and its feature map can improve the accuracy of gait recognition. However, most models just cut the feature map at a fixed single scale, which loses the dependence between various parts. So, our paper proposes a structure called Part Feature Relationship Extractor (PFRE) to discover all of the relationships between each parts for gait recognition. The paper uses PFRE and a Convolutional Neural Network (CNN) to form the RPNet. PFRE is divided into two parts. One part that we call the Total-Partial Feature Extractor (TPFE) is used to extract the features of different scale blocks, and the other part, called the Adjacent Feature Relation Extractor (AFRE), is used to find the relationships between each block. At the same time, the paper adjusts the number of input frames during training to perform quantitative experiments and finds the rule between the number of input frames and the performance of the model. Our model is tested on three public gait datasets, CASIA-B, OU-LP and OU-MVLP. It exhibits a significant level of robustness to occlusion situations, and achieves accuracies of 92.82% and 80.26% on CASIA-B under BG # and CL # conditions, respectively. The results show that our method reaches the top level among state-of-the-art methods.
Hao Qin 0006, Zhenxue Chen, Qingqiang Guo, Q. M. Jonathan Wu, Mengxu Lu
IEEE Trans. Circuits Syst. Video Technol.4
2022 Multimodal Vigilance Estimation Using Deep Learning
abstract
The phenomenon of increasing accidents caused by reduced vigilance does exist. In the future, the high accuracy of vigilance estimation will play a significant role in public transportation safety. We propose a multimodal regression network that consists of multichannel deep autoencoders with subnetwork neurons (MCDAE$_{sn}$). After we define two thresholds of “0.35” and “0.70” from the percentage of eye closure, the output values are in the continuous range of 0–0.35, 0.36–0.70, and 0.71–1 representing the awake state, the tired state, and the drowsy state, respectively. To verify the efficiency of our strategy, we first applied the proposed approach to a single modality. Then, for the multimodality, since the complementary information between forehead electrooculography and electroencephalography features, we found the performance of the proposed approach using features fusion significantly improved, demonstrating the effectiveness and efficiency of our method.
Wei Wu 0022, Wei Sun 0028, Q. M. Jonathan Wu, Yimin Yang 0001, Hui Zhang 0023, Wei-Long Zheng, Bao-Liang Lu
IEEE Trans. Cybern.3
2022 HKPM: A Hierarchical Key-Area Perception Model for HFSWR Maritime Surveillance
abstract
High-frequency surface wave radar (HFSWR) has become the cornerstone of maritime surveillance because of its low-cost maintenance and coverage of wide area. However, when it comes to the extraction of key areas, such as vessel-target detection and vessel-path tracking, the HFSWR signal is strongly interfered by clutters and noise, which makes maritime surveillance a challenging task. This article proposes a hierarchical key-area perception model for maritime surveillance harnessing range-Doppler (RD) image from HFSWR, Laplacian kernel, a linear classifier (LC), and a subnet-based multilayer representation learning framework (SMRLF). First, a weak LC with a Laplacian kernel is utilized to capture the plausible vessel regions (PVRs). Then, a novel SMRLF is proposed to localize the vessel targets from the PVRs. To handle the noise, a maximum correntropy criterion with variable centers (MCC-VC) is incorporated in the subnet-based learning model. A thorough experimental analysis on cross-domain samples from radar dataset to scene classification dataset shows that the proposed HKPM performs competitively. The model shows a superior performance over most of the state-of-the-art vessel-target detection algorithms with a vessel-target detection accuracy of 94%. The extended analysis on image classification problem proves that the proposed model has great adaptivity and scalability.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001, Akilan Thangarajah, Ming Li 0057
IEEE Trans. Geosci. Remote. Sens.2
2022 FRNet: Factorized and Regular Blocks Network for Semantic Segmentation in Road Scene
abstract
Nowadays, semantic segmentation methods for systems in road scene have a great demand. Most existing methods focus on high accuracy with low inference speed. And some approaches emphasize on speed, significantly sacrificing model accuracy. To make a trade-off between accuracy and inference speed, we propose a real-time network for semantic segmentation titled Factorized and Regular Network (FRNet), which employs an asymmetric encoder-decoder architecture with Factorized and Regular (FR) blocks. Our method achieves 70.4% mIoU on the Cityscapes test set with 1 million parameters at a speed of 127 frames per second (FPS) on a single Titan Xp at a resolution of$512\times 1024$. We evaluate FRNet on Cityscapes, Camvid, Kitti, and Gatech datasets to identify that our network stands out from other state-of-the-art networks.
Mengxu Lu, Zhenxue Chen, Q. M. Jonathan Wu, Nannan Wang 0001, Xuewen Rong, Xinghe Yan
IEEE Trans. Intell. Transp. Syst.3
2022 MRSDI-CNN: Multi-Model Rail Surface Defect Inspection System Based on Convolutional Neural Networks
abstract
Defects on rail surfaces, which have become critical problems, need to be detected and removed as quickly as possible to ensure the fast, safe, and stable operation of trains. At present, although many solutions have been proposed to address these problems, the comprehensiveness, rapidity, and accuracy of defect detection remain unsatisfactory. This study aims to resolve these existing problems and accordingly proposes a multi-model rail surface defect detection system based on convolutional neural networks (MRSDI-CNN) from the standpoint of studying the squat on the rail surface. The convolutional neural networks utilized include the improved Single Shot MultiBox Detector (SSD) and You Only Look Once version 3(YOLOv3)—two types of one-stage networks. We expounded and analyzed the performance of the convolutional neural networks as well as their applicability to rail surface defect detection. We used a diverse range of rail defect sizes to improve the detection performance of the two deep learning networks, following which they could identify three types of squats in parallel with improved accuracy and without reduction of the detection speed. The experimental results confirm the effectiveness and superiority of the proposed method over those of previous studies.
Hui Zhang 0023, Yanan Song, Yurong Chen 0003, Hang Zhong, Li Liu 0060, Yaonan Wang 0001, Akilan Thangarajah, Q. M. Jonathan Wu
IEEE Trans. Intell. Transp. Syst.8
2022 Sequential Fusion for Multirate Multisensor Systems With Heavy-Tailed Noises and Unreliable Measurements
abstract
The sequential fusion estimation for multirate multisensor dynamic systems with heavy-tailed noises and unreliable measurements is an important problem in dynamic system control. This work proposes a sequential fusion algorithm and a detection technique based on Student’s$t$-distribution and the approximate$t$-filter. The performance of the proposed algorithm is analyzed and compared with the Gaussian Kalman filter-based sequential fusion and the$t$-filter-based sequential fusion without detection technique. Theoretical analysis and exhaustive experimental analysis show that the proposed algorithm is effective and robust to unreliable measurements. The$t$-filter-based sequential fusion algorithm is shown to be the generalization of the classical Gaussian Kalman filter-based optimal sequential fusion algorithm.
Chenying Di, Q. M. Jonathan Wu, Yuanqing Xia
IEEE Trans. Syst. Man Cybern. Syst.3
2021 Blockchain-based decentralized reputation system in E-commerce environment
Zhili Zhou 0001, Meimin Wang, Ching-Nung Yang, Zhangjie Fu 0001, Xingming Sun, Q. M. Jonathan Wu
Future Gener. Comput. Syst.6
2021 Geometric rectification-based neural network architecture for image manipulation detection
abstract
Determination of image authenticity usually requires the identification and localization of the manipulated regions of images. Hence, image manipulation detection has become one of the most important tasks in the field of multimedia forensics. Recently, Convolutional Neural Networks (CNNs) have achieved promising performance in image manipulation detection. However, it is hard for the existing CNN-based manipulation detection approaches to accurately identify and localize the manipulated regions that have undergone geometric transformations, since CNNs are limited by their inability to be geometrically invariant. To address this issue, we propose a geometric rectification-based neural network architecture for image manipulation detection. In this type of network architecture, following the detection of a set of potential manipulated regions (PMRs) using Region Proposal Network, the Spatial Transformer Network is employed to geometrically rectify the convolutional feature maps (CFMs) of these regions to obtain the geometrically rectified CFMs (GR-CFMs). Subsequently, the residual feature maps (RFMs) are computed to capture the characteristic inconsistency between the CFMs and GR-CFMs of each PMR. Finally, the computed RFMs are automatically integrated with the GR-CFMs by a designed attention module to determine whether each PMR is a manipulated region and to localize the manipulated part at the pixel-level. Extensive experiments on the public data set as well as on our challenging data set demonstrate that the proposed network architecture achieves desirable performance in identifying and localizing regions with common tampering artifacts, which involve geometric transformations.
Zhili Zhou 0001, Wenyan Pan, Q. M. Jonathan Wu, Ching-Nung Yang, Zhihan Lyu
Int. J. Intell. Syst.3
2021 A compensation-based optimization strategy for top dense layer training
Xiexing Feng, Q. M. Jonathan Wu, Yimin Yang 0001, Libo Cao
Neurocomputing2
2021 CascNet: No-reference saliency quality assessment with cascaded applicability sorting and comparing network
Kunqian Li, Duo Shi, Q. M. Jonathan Wu, Xin Luan, Dalei Song
Neurocomputing4
2021 Self-supervised monocular depth estimation with direct methods
Haixia Wang 0003, Yehao Sun, Q. M. Jonathan Wu, Xiao Lu 0003, Zhiguo Zhang 0005
Neurocomputing3
2021 Echo state network with a global reversible autoencoder for time series classification
abstract
An echo state network (z) can provide an efficient dynamic solution for predicting time series problems. However, in most cases, ESN models are applied for predictions rather than classifications. The applications of ESN in time series classification (TSC) problems have yet to be fully studied. Moreover, the conventional randomly generated ESN is unlikely to be optimal because of the randomly generated input and reservoir weights, which are not always guaranteed to be optimal. Randomly generating all layer weights is improper, because a purely random layer might destroy the useful features. To overcome this disadvantage, this study provides a new input weight establishment framework of ESN based on autoencoder (AE) theory for TSC tasks. A global reversible AE (GRAE) algorithm is proposed to reestablish the random initialization input weights of the ESN. In existing ESN-AEs, the output weights obtained in the encoding process are directly reused as the initial input weights. By contrast, in GRAE, the reservoir layer with a reversible activation function is calculated by pulling the decoding layer output back and injecting it into the reservoir layer. Thus, feature learning is enriched by additional information, which results in improved performance. The current weights of the encoding layer are iteratively replaced by the decoding layer to ensure that the outputs of the GRAE are remarkably correlated with the input data. Visualization analyses and experiments of the input weights on a massive set of UCR time series datasets indicate that the proposed GRAE method can considerably improve the original two-layer ESN-based classifiers and the proposed GRAE-ESN classifier yields better performance compared with traditional state-of-the-art TSC classifiers. Furthermore, the proposed method can provide comparable performance and considerably faster training speed compared with three deep learning classifiers.
Heshan Wang, Q. M. Jonathan Wu, Dongshu Wang, Jianbin Xin, Yimin Yang 0001, Kunjie Yu
Inf. Sci.2
2021 FSFN: feature separation and fusion network for single image super-resolution
Zhenxue Chen, Q. M. Jonathan Wu, Nannan Wang 0001
Multim. Tools Appl.3
2021 3MNet: Multi-task, multi-level and multi-channel feature aggregation network for salient object detection
Xinghe Yan, Zhenxue Chen, Q. M. Jonathan Wu, Mengxu Lu, Luna Sun
Mach. Vis. Appl.3
2021 Non-iterative online sequential learning strategy for autoencoder and classifier
Adhri Nandini Paul, Peizhi Yan, Yimin Yang 0001, Hui Zhang 0023, Shan Du 0001, Q. M. Jonathan Wu
Neural Comput. Appl.6
2021 RemNet: remnant convolutional neural network for camera model identification
Abdul Muntakim Rafi, Thamidul Islam Tonmoy, Uday Kamal, Q. M. Jonathan Wu, Md. Kamrul Hasan 0001
Neural Comput. Appl.4
2021 Multiple Distance-Based Coding: Toward Scalable Feature Matching for Large-Scale Web Image Search
abstract
For scalable feature matching in large-scale web image search, the bag-of-visual-words-based (BOW) approaches generally code local features as visual words to construct an inverted index file to match features efficiently. Both the popular feature coding techniques, i.e., K-means-based vector quantization and scalar quantization, directly quantize features to generate visual words. K-means-based vector quantization requires expensive visual codebook training, whereas scalar quantization leads to the miss of many matches due to the low stability of individual components of feature vectors. To address the above issues, we demonstrate that the corresponding sub-vectors of similar features generally have similar distances to multiple reference points in feature subspace and propose a multiple distance-based feature coding scheme for scalable feature matching. Specifically, based on the distances between the sub-vectors and multiple distinct reference points, we transform each feature to a set of feature codes, where one code is treated as a visual word required to construct the inverted index file whereas the others are embedded into the index file to further verify the feature matching based on the visual words. The proposed coding scheme does not need visual codebook training and shows desirable stability and discriminability. Moreover, in the matching verification, a feature-distance estimation method is proposed to estimate the Euclidean distances between features for an accurate matching verification. Extensive experimental results demonstrate the superiority of the proposed approach in comparison to the other approaches using recent feature quantization methods for large-scale web image search.
Zhili Zhou 0001, Q. M. Jonathan Wu, Xingming Sun
IEEE Trans. Big Data2
2021 AMPNet: Average- and Max-Pool Networks for Salient Object Detection
abstract
Salient Object Detection aims to detect the most visually distinctive objects in an image. We solve this problem by introducing the average pool to explore the multi-level deep average pool convolution features different from the max pool information. Based on the U-net structure, we propose an Average- and Max-Pool Network (AMPNet) that leverages the average- and max-pool modules to integrate the multi-level complementary contextual features in the spatial and channel-wise dimensions, respectively. The complementary contextual features generated by our network can improve the completeness of detected objects. It has been observed that the non-salient regions are misrecognized as the salient objects because of the redundant information contained in the multi-level convolution features. To address the problem, two top-down feedback paths are introduced based on the above two modules, and their top-level semantic guidance information is fully utilized to improve the accuracy of salient objects detection. Finally, we apply the Feature Fusion Module and Deep Supervision Mechanism to further improve the performance of the network over different datasets. Experimental results on six benchmark datasets show that our network is on par with state-of-the-art approaches. Our method runs at more than 45 FPS (based on VGG) and 35 FPS (based on ResNet) on a single GPU and meets real-time requirements.
Luna Sun, Zhenxue Chen, Q. M. Jonathan Wu, Hongjian Zhao, Weikai He, Xinghe Yan
IEEE Trans. Circuits Syst. Video Technol.3
2021 Context-Aware Correlation Filter Learning Toward Peak Strength for Visual Tracking
abstract
Recently, the correlation filter (CF) has been catching significant attention in visual tracking for its high efficiency in most state-of-the-art algorithms. However, the tracker easily fails when facing the distractions caused by background clutter, occlusion, and other challenging situations. These distractions commonly exist in the visual object tracking of real applications. Keep tracking under these circumstances is the bottleneck in the field. To improve tracking performance under complex interference, a combination of least absolute shrinkage and selection operator (LASSO) regression and contextual information is introduced to the CF framework through the learning stage in this article to ignore these distractions. Moreover, an elastic net regression is proposed to regroup the features, and an adaptive scale method is implemented to deal with the scale changes during tracking. Theoretical analysis and exhaustive experimental analysis show that the proposed peak strength context-aware (PSCA) CF significantly improves the kernelized CF (KCF) and achieves better performance than other state-of-the-art trackers.
Tayssir Bouraffa, Zihang Feng, Bo Xiao 0006, Q. M. Jonathan Wu, Yuanqing Xia
IEEE Trans. Cybern.5
2021 Analysis and Estimation of Shipborne HFSWR Target Parameters Under the Influence of Platform Motion
abstract
For onshore high-frequency surface-wave radar (HFSWR), target parameter estimation focuses mainly on distance and velocity demodulation, then on azimuth determination. However, these parameters are difficult to solve accurately for shipborne HFSWR targets due to additional modulation on the echo signal introduced by the forward and six-degree-of-freedom (6-DOF) movement of the platform, causing target points spread and shift in radar Doppler spectra. To overcome this difficulty, the influence of platform motion on target detection is mathematically analyzed in terms of echo signal processing, and then theoretical equations are derived to correct the bias in measurements of the target's state and features. Furthermore, to meet the requirement of shipborne HFSWR installed in limited space, the direction of arrival (DOA) estimation of irregular radar arrays with unequal intervals and arbitrary numbers of antennas is also analyzed. With the derived formulas, the parameters of the shipborne HFSWR target, including the range, radial velocity, and azimuth, can be accurately estimated with the help of inertial navigation system (INS) data. Moreover, both the simulation and the field experiment results validate the theoretical analysis and the derived equations.
Kaixian Yang, Ling Zhang 0003, Jiong Niu, Yonggang Ji, Q. M. Jonathan Wu
IEEE Trans. Geosci. Remote. Sens.5
2021 Multi-Scale Gradients Self-Attention Residual Learning for Face Photo-Sketch Transformation
abstract
Face sketch synthesis, as a key technique for solving face sketch recognition, has made considerable progress in recent years. Due to the difference of modality between face photo and face sketch, traditional exemplar-based methods often lead to missed texture details and deformation while synthesizing sketches. And limited to the local receptive field, Convolutional Neural Networks-based methods cannot deal with the interdependence between features well, which makes the constraint of facial features insufficient; as such, it cannot retain some details in the synthetic image. Moreover, the deeper the network layer is, the more obvious the problems of gradient disappearance and explosion will be, which will lead to instability in the training process. Therefore, in this paper, we propose a multi-scale gradients self-attention residual learning framework for face photo-sketch transformation that embeds a self-attention mechanism in the residual block, making full use of the relationship between features to selectively enhance the characteristics of specific information through self-attention distribution. Simultaneously, residual learning can keep the characteristics of the original features from being destroyed. In addition, the problem of instability in GAN training is alleviated by allowing discriminator to become a function of multi-scale outputs of the generator in the training process. Based on cycle framework, the matching between the target domain image and the source domain image can be constrained while the mapping relationship between the two domains is established so that the tasks of face photo-to-sketch synthesis (FP2S) and face sketch-to-photo synthesis (FS2P) can be achieved simultaneously. Both Image Quality Assessment (IQA) and experiments related to face recognition show that our method can achieve state-of-the-art performance on the public benchmarks, whether using FP2S or FS2P.
Shuchao Duan, Zhenxue Chen, Q. M. Jonathan Wu, Dan Lu 0006
IEEE Trans. Inf. Forensics Secur.3
2021 A Width-Growth Model With Subnetwork Nodes and Refinement Structure for Representation Learning and Image Classification
abstract
This article presents a new supervised multilayer subnetwork-based feature refinement and classification model for representation learning. The novelties of this algorithm are as follows: 1) different from most multilayer networks that go deeper with increased number of network layers, this work architects a model with wider subnetwork nodes; 2) the conventional classification methods adopt a separate search mechanism to derive a generalized feature space and to get the final cognition, but this work proposes a one-shot process to find the meaningful latent space and recognize the objects; and 3) the traditional feature representation and image classification approaches apply a unimodal feature coding, which suffers from lack of global knowledge. This work overcomes the pitfall through multimodal fusion that fuses various feature sources into one superstate encoding to achieve higher performance. A cross-domain experimental study on camera identification and image classification shows that the proposed method achieves superior performance compared to the existing models.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001, Akilan Thangarajah, Hui Zhang 0023
IEEE Trans. Ind. Informatics2
2021 MAMA Net: Multi-Scale Attention Memory Autoencoder Network for Anomaly Detection
abstract
Anomaly detection refers to the identification of cases that do not conform to the expected pattern, which takes a key role in diverse research areas and application domains. Most of existing methods can be summarized as anomaly object detection-based and reconstruction error-based techniques. However, due to the bottleneck of defining encompasses of real-world high-diversity outliers and inaccessible inference process, individually, most of them have not derived groundbreaking progress. To deal with those imperfectness, and motivated by memory-based decision-making and visual attention mechanism as a filter to select environmental information in human vision perceptual system, in this paper, we propose a Multi-scale Attention Memory with hash addressing Autoencoder network (MAMA Net) for anomaly detection. First, to overcome a battery of problems result from the restricted stationary receptive field of convolution operator, we coin the multi-scale global spatial attention block which can be straightforwardly plugged into any networks as sampling, upsampling and downsampling function. On account of its efficient features representation ability, networks can achieve competitive results with only several level blocks. Second, it's observed that traditional autoencoder can only learn an ambiguous model that also reconstructs anomalies "well" due to lack of constraints in training and inference process. To mitigate this challenge, we design a hash addressing memory module that proves abnormalities to produce higher reconstruction error for classification. In addition, we couple the mean square error (MSE) with Wasserstein loss to improve the encoding data distribution. Experiments on various datasets, including two different COVID-19 datasets and one brain MRI (RIDER) dataset prove the robustness and excellent generalization of the proposed MAMA Net.
Yurong Chen 0003, Hui Zhang 0023, Yaonan Wang 0001, Yimin Yang 0001, Xianen Zhou, Q. M. Jonathan Wu
IEEE Trans. Medical Imaging6
2021 CASNet: A Cross-Attention Siamese Network for Video Salient Object Detection
abstract
Recent works on video salient object detection have demonstrated that directly transferring the generalization ability of image-based models to video data without modeling spatial-temporal information remains nontrivial and challenging. Considering both intraframe accuracy and interframe consistency of saliency detection, this article presents a novel cross-attention based encoder-decoder model under the Siamese framework (CASNet) for video salient object detection. A baseline encoder-decoder model trained with Lovász softmax loss function is adopted as a backbone network to guarantee the accuracy of intraframe salient object detection. Self- and cross-attention modules are incorporated into our model in order to preserve the saliency correlation and improve intraframe salient detection consistency. Extensive experimental results obtained by ablation analysis and cross-data set validation demonstrate the effectiveness of our proposed method. Quantitative results indicate that our CASNet model outperforms 19 state-of-the-art image- and video-based methods on six benchmark data sets.
Yuzhu Ji, Haijun Zhang 0002, Zequn Jie, Lin Ma 0002, Q. M. Jonathan Wu
IEEE Trans. Neural Networks Learn. Syst.5
2021 Hierarchical One-Class Classifier With Within-Class Scatter-Based Autoencoders
abstract
Autoencoding is a vital branch of representation learning in deep neural networks (DNNs). The extreme learning machine-based autoencoder (ELM-AE) has been recently developed and has gained popularity for its fast learning speed and ease of implementation. However, the ELM-AE uses random hidden node parameters without tuning, which may generate meaningless encoded features. In this brief, we first propose a within-class scatter information constraint-based AE (WSI-AE) that minimizes both the reconstruction error and the within-class scatter of the encoded features. We then build stacked WSI-AEs into a one-class classification (OCC) algorithm based on the hierarchical regularized least-squared method. The effectiveness of our approach was experimentally demonstrated in comparisons with several state-of-the-art AEs and OCC algorithms. The evaluations were performed on several benchmark data sets.
Tianlei Wang, Jiuwen Cao, Xiaoping Lai, Q. M. Jonathan Wu
IEEE Trans. Neural Networks Learn. Syst.4
2021 Multimodel Feature Reinforcement Framework Using Moore-Penrose Inverse for Big Data Analysis
abstract
Fully connected representation learning (FCRL) is one of the widely used network structures in multimodel image classification frameworks. However, most FCRL-based structures, for instance, stacked autoencoder encode features and find the final cognition with separate building blocks, resulting in loosely connected feature representation. This article achieves a robust representation by considering a low-dimensional feature and the classifier model simultaneously. Thus, a new hierarchical subnetwork-based neural network (HSNN) is proposed in this article. The novelties of this framework are as follows: 1) it is an iterative learning process, instead of stacking separate blocks to obtain the discriminative encoding and the final classification results. In this sense, the optimal global features are generated; 2) it applies Moore-Penrose (MP) inverse-based batch-by-batch learning strategy to handle large-scale data sets, so that large data set, such as Place365 containing 1.8 million images, can be processed effectively. The experimental results on multiple domains with a varying number of training samples from ∼ 1 K to ∼ 2 M show that the proposed feature reinforcement framework achieves better generalization performance compared with most state-of-the-art FCRL methods.
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001, Akilan Thangarajah
IEEE Trans. Neural Networks Learn. Syst.2
2021 Introduction to the Special Section on Security and Privacy of Medical Data for Smart Healthcare
abstract
introduction Share on Introduction to the Special Section on Security and Privacy of Medical Data for Smart Healthcare Editors: Amit Kumar Singh National Institute of Technology Patna, India National Institute of Technology Patna, IndiaSearch about this author , Jonathan Wu University of Windsor, Canada University of Windsor, CanadaSearch about this author , Ali Al-Haj Princess Sumaya University for Technology, Jordan Princess Sumaya University for Technology, JordanSearch about this author , Calton Pu Georgia Institute of Technology, USA Georgia Institute of Technology, USASearch about this author Authors Info & Claims ACM Transactions on Internet TechnologyVolume 21Issue 3August 2021 Article No.: 53pp 1–4https://doi.org/10.1145/3460870Online:09 June 2021Publication History 3citation88DownloadsMetricsTotal Citations3Total Downloads88Last 12 Months88Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Amit Kumar Singh 0001, Q. M. Jonathan Wu, Ali Al-Haj 0001, Calton Pu
ACM Trans. Internet Techn.2
2020 Integrating Deformable Convolution and Pyramid Network in Cascade R-CNN for Fabric Defect Detection
abstract
Defects on the surface of fabrics seriously affect the production speed and quality of textile products. There are many difficulties in the detection of surface defects on fabrics, such as substantial differences in length-width ratio, uneven distribution, and few features. However, existing methods have the disadvantages of slow detection speed and high misdetection rate. This present study proposes a method of integrating deformable convolution and pyramid network in Cascade R-CNN (IDPNet) for fabric defect detection. First, image data are labeled according to the type and distribution of defects. Then we design a novel multi-stage object detection architecture named IDPNet to detect defects on the surface of fabrics. In the first stage, Resnet50, in combination with feature pyramid network and deformable convolution is used to improve the detection performance of small defects. Besides, we trained a sequence of detectors with increasing IoUs stage by stage based on Cascade R-CNN in the second stage. Finally, experimental results demonstrate that the proposed neural network equip an outstanding performance against other approaches and achieve the accuracy of 91.57% in fabric defect detection, which proves its utility in practice.
Honghao Li, Hui Zhang 0023, Li Liu 0060, Hang Zhong, Yaonan Wang 0001, Q. M. Jonathan Wu
SMC6
2020 Understanding Global Reaction to the Recent Outbreaks of COVID-19: Insights from Instagram Data Analysis
abstract
The coronavirus disease, also known as the COVID-19, is an ongoing pandemic of a severe acute respiratory syndrome. The pandemic has led to the cancellation of many religious, political, and cultural events around the world. A huge number of people have been stuck within their homes because of unprecedented lockdown measures taken globally. This paper examines the reaction of individuals to the virus outbreak-through the analytical lens of specific hashtags on the Instagram platform. The Instagram posts are analyzed in an attempt to surface commonalities in the way that individuals use visual social media when reacting to this crisis. After collecting the data, the posts containing the location data are selected. A portion of these data are chosen randomly and are categorized into five different categories. We perform several manual analyses to get insights into our collected dataset. Afterward, we use the ResNet-50 convolutional neural network for classifying the images associated with the posts, and attention-based LSTM networks for performing the caption classification. This paper discovers a range of emerging norms on social media in global crisis moments. The obtained results indicate that our proposed methodology can be used to automate the sentiment analysis of mass people using Instagram data.
Abdul Muntakim Rafi, Shivang Rana, Rajwinder Kaur, Q. M. Jonathan Wu, Pooya Moradian Zadeh
SMC4
2020 Sequential fusion estimation for multisensor systems with non-Gaussian noises
Chenying Di, Q. M. Jonathan Wu, Yuanqing Xia
Sci. China Inf. Sci.3
2020 An Autuencoder-based Data Augmentation Strategy for Generalization Improvement of DCNNs
Xiexing Feng, Q. M. Jonathan Wu, Yimin Yang 0001, Libo Cao
Neurocomputing2
2020 Clothescounter: A framework for star-oriented clothes mining from videos
Haijun Zhang 0002, Yuzhu Ji, Q. M. Jonathan Wu
Neurocomputing5
2020 Wi-HSNN: A subnetwork-based encoding structure for dimension reduction and food classification via harnessing multi-CNN model high-level features
Wandong Zhang, Q. M. Jonathan Wu, Yimin Yang 0001
Neurocomputing2
2020 Optimizing Deep Belief Echo State Network with a Sensitivity Analysis Input Scaling Auto-Encoder algorithm
Heshan Wang, Q. M. Jonathan Wu, Jianbin Xin, Jie Wang 0026
Knowl. Based Syst.2
2020 Deep mutual learning network for gait recognition
Zhenxue Chen, Q. M. Jonathan Wu, Xuewen Rong
Multim. Tools Appl.3
2020 Improved face super-resolution generative adversarial networks
Mengxue Wang, Zhenxue Chen, Q. M. Jonathan Wu, Muwei Jian
Mach. Vis. Appl.3
2020 Pedestrian detection via deep segmentation and context network
Zhaoqing Li, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Neural Comput. Appl.3
2020 Recomputation of the Dense Layers for Performance Improvement of DCNN
abstract
Gradient descent optimization of learning has become a paradigm for training deep convolutional neural networks (DCNN). However, utilizing other learning strategies in the training process of the DCNN has rarely been explored by the deep learning (DL) community. This serves as the motivation to introduce a non-iterative learning strategy to retrain neurons at the top dense or fully connected (FC) layers of DCNN, resulting in, higher performance. The proposed method exploits the Moore-Penrose Inverse to pull back the current residual error to each FC layer, generating well-generalized features. Further, the weights of each FC layers are recomputed according to the Moore-Penrose Inverse. We evaluate the proposed approach on six most widely accepted object recognition benchmark datasets: Scene-15, CIFAR-10, CIFAR-100, SUN-397, Places365, and ImageNet. The experimental results show that the proposed method obtains improvements over 30 state-of-the-art methods. Interestingly, it also indicates that any DCNN with the proposed method can provide better performance than the same network with its original Backpropagation (BP)-based training.
Yimin Yang 0001, Q. M. Jonathan Wu, Xiexing Feng, Akilan Thangarajah
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 3D Parallel Fully Convolutional Networks for Real-Time Video Wildfire Smoke Detection
abstract
Wildfires have devastating consequences on ecological systems and human lives. Accurate and fast wildfire detection is crucial to reduce damage. The existing smoke detection algorithms using convolution neural network are mostly based on the classification of smoke images or patches, whereas the traditional smoke detection algorithms are often necessary to extract multiple features for integration. With the methods mentioned above, false positive is always an insurmountable problem in wildfire smoke detection. Moreover, there are few studies on the detection of wildfire smoke. Thus, to detect the wildfire smoke more intelligent, a 3D parallel fully convolutional network for wildfire smoke detection is proposed to segment the smoke regions in video sequences. Wildfire smoke detection is considered as a segmentation problem in this paper. There are more than 90 videos including various scenes used for training and test. Experiments have demonstrated that our architecture can segment smoke regions accurately and eliminate the interference of natural scenes. Smoke targets in multiple scenes can be detected accurately and quickly.
Xiuqing Li, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
IEEE Trans. Circuits Syst. Video Technol.3
2020 Bridging User Interest to Item Content for Recommender Systems: An Optimization Model
abstract
Recommender systems are currently utilized widely in e-commerce for product recommendations and within content delivery platforms. Previous studies usually use independent features to represent item content. As a result, the relationship hidden among the content features is overlooked. In fact, the reason that an item attracts a user may be attributed to only a few set of features. In addition, these features are often semantically coupled. In this paper, we present an optimization model for extracting the relationship hidden in content features by considering user preferences. The learned feature relationship matrix is then applied to address the cold-start recommendations and content-based recommendations. It could also easily be employed for the visualization of feature relation graphs. Our proposed method was examined on three public datasets: 1) hetrec-movielens-2k-v2; 2) book-crossing; and 3) Netflix. The experimental results demonstrated the effectiveness of our method in comparison to the state-of-the-art recommendation methods.
Haijun Zhang 0002, Yanfang Sun, Ming-Bo Zhao, Tommy W. S. Chow, Q. M. Jonathan Wu
IEEE Trans. Cybern.5
2020 sEnDec: An Improved Image to Image CNN for Foreground Localization
abstract
Although it is not immediately intuitive that Deep Convolutional Neural Networks (DCNNs) can yield adequate feature representation for a Foreground Localization (FGL) task, recent architectural and algorithmic advancements in Deep Learning (DL) have shown that the DCNNs have become forefront methodology for this pixel-level classification problem. In FGL, the DCNNs face an inherent trade-off between moving objects, i.e., the foreground (FG) and the non-static background (BG) scenes, through learning from local- and global-level features. Driven by the latest success of the innovative structures for image classification and semantic segmentation, this work introduces a novel architecture, called Slow Encoder-Decoder (sEnDec) that aims to improve the learning capacity of a traditional image-to-image DCNN. The proposed model subsumes two subnets for contraction (encoding) and expansion (decoding), wherein both phases, it employs an intermediate feature map up-sampling and residual connections. In this way, the lost structural details due to spatial subsampling are recovered. It helps to get a more delineated FG region. The experimental study is carried out with two variants of the proposed model: one with strided convolution (conv) and the other with max pooling for spatial subsampling. A comparative analysis on sixteen benchmark video sequences, including baseline, dynamic background, camera jitter, shadow effects, intermittent object motion, night videos, and bad weather show that the proposed sEnDec model performs very competitively against the prior- and state-of-the-art approaches.
Akilan Thangarajah, Q. M. Jonathan Wu
IEEE Trans. Intell. Transp. Syst.2
2020 A 3D CNN-LSTM-Based Image-to-Image Foreground Segmentation
abstract
The video-based separation of foreground (FG) and background (BG) has been widely studied due to its vital role in many applications, including intelligent transportation and video surveillance. Most of the existing algorithms are based on traditional computer vision techniques that perform pixel-level processing assuming that FG and BG possess distinct visual characteristics. Recently, state-of-the-art solutions exploit deep learning models targeted originally for image classification. Major drawbacks of such a strategy are the lacking delineation of FG regions due to missing temporal information as they segment the FG based on a single frame object detection strategy. To grapple with this issue, we excogitate a 3D convolutional neural network (3D CNN) with long short-term memory (LSTM) pipelines that harness seminal ideas, viz., fully convolutional networking, 3D transpose convolution, and residual feature flows. Thence, an FG-BG segmenter is implemented in an encoder-decoder fashion and trained on representative FG-BG segments. The model devises a strategy called double encoding and slow decoding, which fuses the learned spatio-temporal cues with appropriate feature maps both in the down-sampling and up-sampling paths for achieving well generalized FG object representation. Finally, from the Sigmoid confidence map generated by the 3D CNN-LSTM model, the FG is identified automatically by using Nobuyuki Otsu's method and an empirical global threshold. The analysis of experimental results via standard quantitative metrics on 16 benchmark datasets including both indoor and outdoor scenes validates that the proposed 3D CNN-LSTM achieves competitive performance in terms of figure of merit evaluated against prior and state-of-the-art methods. Besides, a failure analysis is conducted on 20 video sequences from the DAVIS 2016 dataset.
Akilan Thangarajah, Q. M. Jonathan Wu, Amin Safaei 0001, Jie Huo, Yimin Yang 0001
IEEE Trans. Intell. Transp. Syst.2
2020 Region-Level Visual Consistency Verification for Large-Scale Partial-Duplicate Image Search
abstract
Most recent large-scale image search approaches build on a bag-of-visual-words model, in which local features are quantized and then efficiently matched between images. However, the limited discriminability of local features and the BOW quantization errors cause a lot of mismatches between images, which limit search accuracy. To improve the accuracy, geometric verification is popularly adopted to identify geometrically consistent local matches for image search, but it is hard to directly use these matches to distinguish partial-duplicate images from non-partial-duplicate images. To address this issue, instead of simply identifying geometrically consistent matches, we propose a region-level visual consistency verification scheme to confirm whether there are visually consistent region (VCR) pairs between images for partial-duplicate search. Specifically, after the local feature matching, the potential VCRs are constructed via mapping the regions segmented from candidate images to a query image by utilizing the properties of the matched local features. Then, the compact gradient descriptor and convolutional neural network descriptor are extracted and matched between the potential VCRs to verify their visual consistency to determine whether they are VCRs. Moreover, two fast pruning algorithms are proposed to further improve efficiency. Extensive experiments demonstrate the proposed approach achieves higher accuracy than the state of the art and provide comparable efficiency for large-scale partial-duplicate search tasks.
Zhili Zhou 0001, Q. M. Jonathan Wu, Yimin Yang 0001, Xingming Sun
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Distributed Multi-Robot Formation Control Based on Two-Layer Nearest Neighbor Information(TNNI) Consensus
abstract
With the development of artificial intelligence, robot swarm systems also frequently appear in complex tasks of different situation. One of the important research directions is the formation of multi-robots. This paper analyzes the limitations of existing algorithms for large-scale mobile robot swarm formation control problems and proposes a consensus control algorithm with two-layer nearest neighbor information. It carries out experimental simulation to verify its convergence performance. At the same time, combined with a distributed structure control strategy that can change the number of robot formation members, the formation control experiment is carried out on the experimental platform consisted of robot state information detection device and multiple mobile robots,to further verify its feasibility.
Guang Deng, Hui Zhang 0023, Hang Zhong, Zhiqiang Miao, Li Liu 0060, Q. M. Jonathan Wu
SMC7
2019 Graph model-based salient object detection using objectness and multiple saliency cues
Yuzhu Ji, Haijun Zhang 0002, Kuo-Kun Tseng, Tommy W. S. Chow, Q. M. Jonathan Wu
Neurocomputing5
2019 Toward AI fashion design: An Attribute-GAN model for clothing match
Haijun Zhang 0002, Yuzhu Ji, Q. M. Jonathan Wu
Neurocomputing4
2019 FCN based preprocessing for exemplar-based face sketch synthesis
Dan Lu 0006, Zhenxue Chen, Q. M. Jonathan Wu, Xuetao Zhang 0003
Neurocomputing3
2019 Hierarchical feature representation for unconstrained video analysis
Eman Mohammadi, Q. M. Jonathan Wu, Mehrdad Saif, Yimin Yang 0001
Neurocomputing2
2019 Optimizing simple deterministically constructed cycle reservoir network with a Redundant Unit Pruning Auto-Encoder algorithm
Heshan Wang, Q. M. Jonathan Wu, Jie Wang 0026, Wei Wu 0022, Kunjie Yu
Neurocomputing2
2019 Saliency object detection: integrating reconstruction and prior
Cuiping Li 0003, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Mach. Vis. Appl.3
2019 Two-stage local details restoration framework for face hallucination
Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Mach. Vis. Appl.3
2019 Face recognition using AMVP and WSRC under variable illumination and pose
Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Neural Comput. Appl.3
2019 Difference co-occurrence matrix using BP neural network for fingerprint liveness detection
Chengsheng Yuan 0001, Xingming Sun, Q. M. Jonathan Wu
Soft Comput.3
2019 Coverless image steganography using partial-duplicate image retrieval
Zhili Zhou 0001, Yan Mu, Q. M. Jonathan Wu
Soft Comput.3
2019 System-on-a-Chip (SoC)-Based Hardware Acceleration for an Online Sequential Extreme Learning Machine (OS-ELM)
abstract
Machine learning algorithms such as those for object classification in images, video content analysis, and human action recognition are used to extract meaningful information from data recorded by image sensors and cameras. Among the existing machine learning algorithms for such purposes, extreme learning machines (ELMs) and online sequential ELMs (OS-ELMs) are well known for their computational efficiency and performance when processing large datasets. The latter approach was derived from the ELM approach and optimized for real-time application. However, OS-ELM classifiers are computationally demanding, and the existing state-of-the-art computing platforms are not efficient enough for embedded systems, especially for applications with strict requirements in terms of low power consumption, high throughput, and low latency. This paper presents the implementation of an ELM/OS-ELM in a customized system-on-a-chip field-programmable gate array-based architecture to ensure efficient hardware acceleration. The acceleration process comprises parallel extraction, deep pipelining, and efficient shared memory communication.
Amin Safaei 0001, Q. M. Jonathan Wu, Akilan Thangarajah, Yimin Yang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Fast Semantic Segmentation for Scene Perception
abstract
Semantic segmentation is a challenging problem in computer vision. Many applications, such as autonomous driving and robot navigation with urban road scene, need accurate and efficient segmentation. Most state-of-the-art methods focus on accuracy, rather than efficiency. In this paper, we propose a more efficient neural network architecture, which has fewer parameters, for semantic segmentation in the urban road scene. An asymmetric encoder-decoder structure based on ResNet is used in our model. In the first stage of encoder, we use continuous factorized block to extract low-level features. Continuous dilated block is applied in the second stage, which ensures that the model has a larger view field, while keeping the model small-scale and shallow. The down sampled features from encoder are up sampled with decoder to the same-size output as the input image and the details refined. Our model can achieve end-to-end and pixel-to-pixel training without pretraining from scratch. The parameters of our model are only 0.2M, 100× less than those of others such as SegNet, etc. Experiments are conducted on five public road scene datasets (CamVid, CityScapes, Gatech, KITTI Road Detection, and KITTI Semantic Segmentation), and the results demonstrate that our model can achieve better performance.
Xuetao Zhang 0003, Zhenxue Chen, Q. M. Jonathan Wu, Dan Lu 0006, Xianming Li
IEEE Trans. Ind. Informatics3
2019 Weighted Extreme Sparse Classifier and Local Derivative Pattern for 3D Face Recognition
abstract
A novel weighted hybrid classifier and a high-order, local normal derivative pattern descriptor is proposed for 3D face recognition. The Local derivative pattern (LDP) captures detailed information, based on the local derivative variation in different directions. The LDP is computed on three normal maps in x, y, and z directions and on different scales. The surface normal captures the orientation of a surface at each point of 3D data. More informative local shape information is extracted using the surface normal, as compared to depth. The nth-order LDP on the surface normal is proposed to encode more detailed features from the (n-1)th-order's local derivative direction variations. An extreme learning machine (ELM) based autoencoder, using a multilayer network structure, is employed to select more discriminant features and provide a faster training speed. A weighted hybrid framework is proposed to handle facial challenges using a combination of the ELM and the sparse representation classifier (SRC). The advantage of speed for the ELM and accuracy for the SRC in a weighted scheme is used to enhance the performance of the recognition system. Experimental results regarding four famous 3D face databases illustrate the generalization and effectiveness of the proposed method in terms of both computational cost and recognition accuracy.
Sima Soltanpour, Q. M. Jonathan Wu
IEEE Trans. Image Process.2
2019 Deep Saliency With Channel-Wise Hierarchical Feature Responses for Traffic Sign Detection
abstract
Traffic sign detection is challenging in cases of a complex background, occlusions, distortions, and so on. To overcome the above-mentioned challenges, this paper pays close attention to channel-wise feature responses to propose an end-to-end deep learning-based saliency traffic sign detection method. Our model contains three main components: channel-wise coarse feature extraction (CCFE), channel-wise hierarchical feature refinement (CHFR), and hierarchical feature map fusion (HFMF). In addition, it is based on the squeeze-and-excitation-residual network to explicitly model the inter dependences between the channels of its convolution features at a slight computational cost. We first apply CCFE to produce coarse feature maps with much information loss. To make full use of spatial information and fine details, CHFR is executed to refine hierarchical features. After that, HFMF is used to fuse hierarchical feature maps to generate the final traffic sign saliency map. Compared with other five traffic sign detection methods, the experimental results demonstrate the efficiency (a real-time speed) and superior performance of the proposed method according to comprehensive evaluations over three benchmark data sets.
Cuiping Li 0003, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
IEEE Trans. Intell. Transp. Syst.3
2019 Features Combined From Hundreds of Midlayers: Hierarchical Networks With Subnetwork Nodes
abstract
In this paper, we believe that the mixed selectivity of neuron in the top layer encodes distributed information produced from other neurons to offer a significant computational advantage over recognition accuracy. Thus, this paper proposes a hierarchical network framework that the learning behaviors of features combined from hundreds of midlayers. First, a subnetwork neuron, which itself could be constructed by other nodes, is functional as a subspace features extractor. The top layer of a hierarchical network needs subspace features produced by the subnetwork neurons to get rid of factors that are not relevant, but at the same time, to recast the subspace features into a mapping space so that the hierarchical network can be processed to generate more reliable cognition. Second, this paper shows that with noniterative learning strategy, the proposed method has a wider and shallower structure, providing a significant role in generalization performance improvements. Hence, compared with other state-of-the-art methods, multiple channel features with the proposed method could provide a comparable or even better performance, which dramatically boosts the learning speed. Our experimental results show that our platform can provide a much better generalization performance than 55 other state-of-the-art methods.
Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Neural Networks Learn. Syst.2
2018 Hardware Realization of Mixed-Signal Neural Networks with Modular Synapse-Neuron arrays
abstract
In this paper, a mixed-signal current-mode structure of a feed-forward neural network is implemented. In this network, neurons are divided and distributed as sub-neurons into parallel elements composing unified synapse-neuron building blocks in combination with the synapses. Although in this brief paper a resistive sigmoidal neuron is considered, the neuron is adaptable to other forms of transfer functions. The synapse structure employs AND gates in addition to weighted current mirrors to reduce the area of the design. As a proof of concept, a 4-3-2 CMOS-based network is implemented. The average and maximum power consumptions of the network are 0.93mW and 5.81 mW respectively. The area of the entire network is measured 142299.5μm2. The network was successfully tested with a series of sample patterns.
Bahar Youssefi, Alexander J. Leigh, Mitra Mirhassani, Q. M. Jonathan Wu
ISCAS4
2018 Effect of fusing features from multiple DCNN architectures in image classification
abstract
Automatic image classification has become a necessary task to handle the rapidly growing digital image usage. It has branched out many algorithms and adopted new techniques. Among them, feature fusion‐based image classification methods rely on hand‐crafted features traditionally. However, it has been proven that the bottleneck features extracted through pre‐trained convolutional neural networks (CNNs) can improve the classification accuracy. Thence, this study analyses the effect of fusing such cues from multiple architectures without being tied to any hand‐crafted features. First, the CNN features are extracted from three different pre‐trained models, namely AlexNet, VGG‐16, and Inception‐V3. Then, a generalised feature space is formed by employing principal component reconstruction and energy‐level normalisation, where the features from individual CNN are mapped into a common subspace and embedded using arithmetic rules to construct fused feature vectors (FFVs). This transformation play a vital role in creating a representation that is appearance invariant by capturing complementary information of different high‐level features. Finally, a multi‐class linear support vector machine is trained. The experimental results demonstrate that such multi‐modal CNN feature fusion is well suited for image/object classification tasks, but surprisingly it has not been explored so far by the computer vision research community extensively.
Akilan Thangarajah, Q. M. Jonathan Wu, Hui Zhang 0023
IET Image Process.2
2018 Hierarchical extreme learning machines
Guang-Bin Huang, Q. M. Jonathan Wu, Donald C. Wunsch II
Neurocomputing2
2018 Saliency detection via conditional adversarial image-to-image network
Yuzhu Ji, Haijun Zhang 0002, Q. M. Jonathan Wu
Neurocomputing3
2018 Salient object detection via multi-scale attention CNN
Yuzhu Ji, Haijun Zhang 0002, Q. M. Jonathan Wu
Neurocomputing3
2018 Deep saliency detection via channel-wise hierarchical feature responses
Cuiping Li 0003, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Neurocomputing3
2018 Fusion-based foreground enhancement for background subtraction using multivariate multi-model Gaussian distribution
Akilan Thangarajah, Q. M. Jonathan Wu, Yimin Yang 0001
Inf. Sci.2
2018 Extracting online information from dual and multiple data streams
Zeeshan Khawar Malik, Amir Hussain 0001, Q. M. Jonathan Wu
Neural Comput. Appl.3
2018 Encoding multiple contextual clues for partial-duplicate image retrieval
Zhili Zhou 0001, Q. M. Jonathan Wu, Xingming Sun
Pattern Recognit. Lett.2
2018 State Identification of Duffing Oscillator Based on Extreme Learning Machine
abstract
As an important weak target detection method, Duffing oscillator is very effective in detecting signals with very low signal-to-noise ratio. However, the accurate discrimination between chaotic and periodic states is a crucial problem and that is the prerequisite for using the Duffing oscillator. Conventionally, the Lyapunov exponent is used as an index to identify different states, but as this indicator has the problem of heavy computation cost, slow convergence rate, and requires a mass of data, its application becomes seriously limits. To solve this problem, a novel method for state identification of the Duffing oscillator based on extreme learning machine (ELM) is proposed. The feature data, as the input of ELM, are extracted from the phase diagram and the time series of the Duffing oscillator. Three effective features are extracted in this letter, i.e., ratio of points in and out of the closed region, average distance, and power spectrum. Computer simulations are presented to validate the proposed method and demonstrate that the state classification performance is superior to other related methods with higher computation efficiency, faster convergence rate, and better accuracy.
Gangsheng Li, Liping Zeng, Ling Zhang 0003, Q. M. Jonathan Wu
IEEE Signal Process. Lett.4
2018 Tree2Vector: Learning a Vectorial Representation for Tree-Structured Data
abstract
The tree structure is one of the most powerful structures for data organization. An efficient learning framework for transforming tree-structured data into vectorial representations is presented. First, in attempting to uncover the global discriminative information of child nodes hidden at the same level of all of the trees, a clustering technique can be adopted for allocating children into different clusters, which are used to formulate the components of a vector. Moreover, a locality-sensitive reconstruction method is introduced to model a process, in which each parent node is assumed to be reconstructed by its children. The resulting reconstruction coefficients are reversely transformed into complementary coefficients, which are utilized for locally weighting the components of the vector. A new vector is formulated by concatenating the original parent node vector and the learned vector from its children. This new vector for each parent node is inputted into the learning process of formulating vectorial representation at the upper level of the tree. This recursive process concludes when a vectorial representation is achieved for the entire tree. Our method is examined in two applications: book author recommendations and content-based image retrieval. Extensive experimental results demonstrate the effectiveness of the proposed method for transforming tree-structured data into vectors.
Haijun Zhang 0002, Shuang Wang 0005, Xiaofei Xu 0001, Tommy W. S. Chow, Q. M. Jonathan Wu
IEEE Trans. Neural Networks Learn. Syst.5
2018 Autoencoder With Invertible Functions for Dimension Reduction and Image Reconstruction
abstract
The extreme learning machine (ELM), which was originally proposed for “generalized” single-hidden layer feedforward neural networks, provides efficient unified learning solutions for the applications of regression and classification. Although, it provides promising performance and robustness and has been used for various applications, the single-layer architecture possibly lacks the effectiveness when applied for natural signals. In order to over come this shortcoming, the following work indicates a new architecture based on multilayer network framework. The significant contribution of this paper are as follows: 1) unlike existing multilayer ELM, in which hidden nodes are obtained randomly, in this paper all hidden layers with invertible functions are calculated by pulling the network output back and putting it into hidden layers. Thus, the feature learning is enriched by additional information, which results in better performance; 2) in contrast to the existing multilayer network methods, which are usually efficient for classification applications, the proposed architecture is implemented for dimension reduction and image reconstruction; and 3) unlike other iterative learning-based deep networks (DL), the hidden layers of the proposed method are obtained via four steps. Therefore, it has much better learning efficiency than DL. Experimental results on 33 datasets indicate that, in comparison to the other existing dimension reduction techniques, the proposed method performs competitively better with fast training speeds.
Yimin Yang 0001, Q. M. Jonathan Wu, Yaonan Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2017 Effect of wavelet and hybrid classification on action recognition
abstract
Any action dataset may contain similar classes such as running, walking and jogging. Therefore, equivalent probabilities may be provided for different classes upon action classification. In this case, the classifier cannot indubitably assign a class to a given sample. To address this problem, we propose a new hybrid classifier to automatically compress the features and classify them using SVM with polynomial or sigmoid kernels. Furthermore, we hypothesize that motion saliency detection can strength the power of motion feature extraction in the bag of visual words framework (BoVW). To this end, we evaluate the effect of 3D-discrete wavelet transform (3D-DWT), as the preprocessing step, on motion feature extraction. The experimental results show that the proposed framework achieves promising results on KTH, Weizmann, and URADL datasets, and outperforms recent state-of-the-art approaches.
Eman Mohammadi, Q. M. Jonathan Wu, Yimin Yang 0001, Mehrdad Saif
ICIP2
2017 High-order local normal derivative pattern (LNDP) for 3D face recognition
abstract
This paper proposes a novel descriptor based on the local derivative pattern (LDP) for 3D face recognition. Compared to the local binary pattern (LBP), LDP can capture more detailed information by encoding directional pattern features. It is based on the local derivative variations that extract high-order local information. We propose a novel discriminative facial shape descriptor, local normal derivative pattern (LNDP) that extracts LDP from the surface normal. Using surface normal, the orientation of a surface at each point is determined as a first-order surface differential. Three normal component images are extracted by estimating three components of normal vectors in x, y, and z channels. Each normal component is divided into several patches and encoded using LDP. The final descriptor is created by concatenating histograms of the LNDP on each patch. Experimental results on two famous 3D face databases, FRGC v2.0 and Bosphorus illustrate the effectiveness of the proposed descriptor.
Sima Soltanpour, Q. M. Jonathan Wu
ICIP2
2017 Improved rank pooling strategy for complex action recognition
abstract
Feature ranking from video-wide temporal evolution brings reliable information for complex action recognition. However, a video may contain similar features in the sequence of frames which deliver unnecessary information to the ranking function. This paper proposes a method to improve the rank-pooling strategy which captures the optimized latent structure of the video sequence data. The optimization is followed by removing the redundant features from the sequence data. The cosine and correlation distance metrics are employed to detect the identical features and extract the most efficient information from the video frames. Then, the ranked features are generated from the optimized and clean sequence data. The proposed improvement is easy to implement, fast to compute and effective in recognizing complex actions. As a result, the proposed approach reaches remarkable action recognition performance on benchmark datasets, namely Hollywood2, URADL, and HMDB51. The results are further compared with state-of-the-art techniques in the experiment section to confirm the effectiveness of the improved rank pooling framework.
Eman Mohammadi, Q. M. Jonathan Wu, Mehrdad Saif
SMC2
2017 System-level design for human action recognition in 3D scenes
abstract
This study proposes a system-on-a-chip, field-programmable gate array (FPGA)-based real-time video processing platform for human action recognition. We provide the details of a hardware implementation for real-time human activity recognition in 3D scenes, including capture, processing, and display. The proposed platform is implemented by adding a two-stage preprocessing step to improve the results of a saliency map and utilizing the inherent parallelism of FPGAs. Appropriate circuits for parallelized operation of Harris and Hessian interest point detection and support vector machine classifiers are discussed.
Amin Safaei 0001, Q. M. Jonathan Wu
SMC2
2017 System-on-chip-based hardware acceleration for human detection in 2D/3D scenes
abstract
A system-on-chip field gate programmable array (FPGA)-based video processing platform for human detection in complex scenes is presented. This study details the hardwarebased implementation of a human detection algorithm in 2D/3D scenes, including the capture, video processing, and display stages. The proposed method is implemented by extending a previously proposed method that uses features extracted from the Riemannian manifold of region covariance matrices computed from 2D data. The proposed method considers both 2D and depth data (3D). The LogitBoost classifier is employed to detect humans. The proposed implementation uses minimal resources and employs a pipeline technique for better performance and operation.
Amin Safaei 0001, Q. M. Jonathan Wu, Akilan Thangarajah
SMC2
2017 Multiscale depth local derivative pattern for sparse representation based 3D face recognition
abstract
3D face recognition is a popular research area due to its vast application in biometrics and security. Local feature-based methods gain importance in the recent years due to their robustness under degradation conditions. In this paper, a novel high-order local pattern descriptor in combination with sparse representation based classifier (SRC) is proposed for expression robust 3D face recognition. 3D point clouds are converted to depth maps after preprocessing. Multi-directional derivatives are applied in spatial space to encode the depth maps based on the local derivative pattern (LDP) scheme. Directional pattern features are calculated according to local derivative variations. Since LDP computes spatial relationship of neighbors in a local region, it extracts distinct information from the depth map. Multiscale depth-LDP is presented as a novel descriptor for 3D face recognition. The descriptor is employed along with the SRC to increase the range data distinctiveness. A histogram on the derivative pattern creates a spatial feature descriptor that represents the distinctive micro-patterns from 3D data. We evaluate the proposed algorithm on two famous 3D face databases, FRGC v2.0 and Bosphorus. The experimental results demonstrate that the proposed approach achieves acceptable performance under facial expression.
Sima Soltanpour, Q. M. Jonathan Wu
SMC2
2017 A late fusion approach for harnessing multi-cnn model high-level features
abstract
The high-level feature representation of deep convo-lutional neural networks (ConvNets) has proven to be superior to hand-crafted low-level features. Thus, this study investigates the effect of fusing such high-level features from multi-deep ConvNets under an application of visual object/scene categorization. In which, three pre-trained ConvNets are exploited as feature extractors, a single hidden layer is adopted to transform the highlevel representations to a low dimensional feature space and that are fused to harness the rich cues of the individual features. The experimental outcomes on six benchmark databases demonstrate that regardless of variation in visual statistics and tasks the fusion of multi-ConvNets' high-level features can meliorate the classification accuracy compared with a single modality, different ConvNets contain complementary cues of visual contents, and the fusion is capable of producing a very competitive performance to the state-of-the-art methods. Besides that, this fusion approach can be an effective complement to the ConvNets.
Akilan Thangarajah, Q. M. Jonathan Wu, Amin Safaei 0001, Wei Jiang 0009
SMC2
2017 Image segmentation using a hierarchical student's-t mixture model
abstract
As a significant tool, finite mixture models (FMMs) have been widely used for image segmentation. However, there are two problems with standard FMMs: first, the conditional probability is sensitive to outliers. Second, the robustness to image noise is inadequate. In this study, the authors present a novel hierarchical Student's‐ t MM (HSMM), which includes standard FMMs as a sub‐problem. Additionally, to incorporate more image spatial information, they apply a mean template not only to the prior/posterior probability, but also to the sub‐conditional distribution. Thus, their HSMM is more robust to outliers and image noise owing to the spatial constraints from the mean template. In the standard SMM, a t ‐distribution is used to calculate the conditional probability. In this study, the authors present a novel hierarchical student's‐ t mixture model (HSMM), which includes the standard FMM as a sub‐problem. Finally, though they use Student's‐ t ‐distribution to solve the image segment problems of this study, their HSMM achieves excellent performance, is elastic and can encompass any other model that is based on FMMs. Experimental results demonstrate that their proposed method is robust and effective.
Lingcheng Kong, Hui Zhang 0015, Yuhui Zheng, Jiezhong Zhu, Q. M. Jonathan Wu
IET Image Process.6
2017 Automatic Detection of Ship Targets Based on Wavelet Transform for HF Surface Wavelet Radar
abstract
High-frequency surface wave radar (HFSWR) has a vital civilian and military significance for continuous maritime surveillance of activities within exclusive economic zone. However, HFSWR has lower spatial and temporal resolutions and the received signals are strongly polluted by different clutter and background noise. Therefore, ship target detection by HFSWR has become a challenging task. This letter presents an automatic ship target detection algorithm based on discrete wavelet transform (DWT). First, a peak signal-to-noise ratio-based algorithm is proposed to automatically determine the optimal scale of DWT for extraction of ship targets. Second, the high-frequency coefficients of DWT at the optimal scale are processed by a fuzzy set-based method to enhance the useful target information and depress the unwanted background noises. Third, a target-highlighted image is reconstructed by ignoring all the low-frequency coefficients and performing inverse DWT only to the enhanced high-frequency coefficients. Finally, the targets are extracted by adaptive threshold segmentation of the final target-highlighted image. Experimental results show that the proposed approach can automatically extract ship targets effectively for range Doppler images with complex background, and has a better target detection performance than the previous wavelet-based algorithm, thereby providing a new reliable image processing-based method of ship target detection for HFSWR.
Qingzhong Li, Wandong Zhang, Ming Li 0057, Jiong Niu, Q. M. Jonathan Wu
IEEE Geosci. Remote. Sens. Lett.5
2017 A survey of local feature methods for 3D face recognition
Sima Soltanpour, Boubakeur Boufama, Q. M. Jonathan Wu
Pattern Recognit.3
2017 Illumination and pose variable face recognition via adaptively weighted ULBP_MHOG and WSRC
Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Signal Process. Image Commun.3
2017 Multilayered Echo State Machine: A Novel Architecture and Algorithm
abstract
In this paper, we present a novel architecture and learning algorithm for a multilayered echo state machine (ML-ESM). Traditional echo state networks (ESNs) refer to a particular type of reservoir computing (RC) architecture. They constitute an effective approach to recurrent neural network (RNN) training, with the (RNN-based) reservoir generated randomly, and only the readout trained using a simple computationally efficient algorithm. ESNs have greatly facilitated the real-time application of RNN, and have been shown to outperform classical approaches in a number of benchmark tasks. In this paper, we introduce a novel criteria for integrating multiple layers of reservoirs within the ML-ESM. The addition of multiple layers of reservoirs are shown to provide a more robust alternative to conventional RC networks. We demonstrate the comparative merits of this approach in a number of applications, considering both benchmark datasets and real world applications.
Zeeshan Khawar Malik, Amir Hussain 0001, Q. M. Jonathan Wu
IEEE Trans. Cybern.3
2017 Effective and Efficient Global Context Verification for Image Copy Detection
abstract
To detect illegal copies of copyrighted images, recent copy detection methods mostly rely on the bag-of-visual-words (BOW) model, in which local features are quantized into visual words for image matching. However, both the limited discriminability of local features and the BOW quantization errors will lead to many false local matches, which make it hard to distinguish similar images from copies. Geometric consistency verification is a popular technology for reducing the false matches, but it neglects global context information of local features and thus cannot solve this problem well. To address this problem, this paper proposes a global context verification scheme to filter false matches for copy detection. More specifically, after obtaining initial scale invariant feature transform (SIFT) matches between images based on the BOW quantization, the overlapping region-based global context descriptor (OR-GCD) is proposed for the verification of these matches to filter false matches. The OR-GCD not only encodes relatively rich global context information of SIFT features but also has good robustness and efficiency. Thus, it allows an effective and efficient verification. Furthermore, a fast image similarity measurement based on random verification is proposed to efficiently implement copy detection. In addition, we also extend the proposed method for partial-duplicate image detection. Extensive experiments demonstrate that our method achieves higher accuracy than the state-of-the-art methods, and has comparable efficiency to the baseline method based on the BOW quantization.
Zhili Zhou 0001, Yunlong Wang 0006, Q. M. Jonathan Wu, Ching-Nung Yang, Xingming Sun
IEEE Trans. Inf. Forensics Secur.3
2017 Efficient and Rapid Machine Learning Algorithms for Big Data and Dynamic Varying Systems
abstract
With the exponential growth of data and complexity of systems, fast machine learning/artificial intelligence and computational intelligence techniques are highly required. Many conventional computational intelligence techniques face bottlenecks in learning (e.g., intensive human intervention and convergence time) [item 1) in the Appendix]. However, efficient learning algorithms alternatively offer significant benefits including fast learning speed, ease of implementation, and minimal human intervention. The need for efficient and fast implementation of machine learning techniques in big data and dynamic varying systems poses many research challenges. This special issue highlights some latest development in the related areas.
Fuchun Sun 0001, Guang-Bin Huang, Q. M. Jonathan Wu, Shiji Song, Donald C. Wunsch II
IEEE Trans. Syst. Man Cybern. Syst.3
2016 A supervised cooperative clustering scheme for diagnosing process faults in an industrial plant
abstract
This paper presents a novel supervised clustering technique including different clustering algorithms which cooperate together to span the decision space in a supervised manner. It uses a variety of clustering methods for an efficient partitioning. An evolutionary algorithm is used to tune the key parameters of the cooperative scheme which minimizes an error-based objective function on the training dataset. The proposed supervised scheme is developed for diagnosing faults in the Tennessee Eastman process, which is a standard benchmark for fault detection and diagnosis. Experimental results show that the proposed technique can efficiently diagnose the process faults.
Mohammad Anvaripour, Sima Soltanpour, Roozbeh Razavi-Far, Mehrdad Saif, Q. M. Jonathan Wu
CEC5
2016 A system-level design for foreground and background identification in 3D scenes
abstract
This paper proposes a system-on-chip (SoC) FPGA - based real-time video processing platform for background and foreground identification. Background and foreground identification is a co mmon feature in many tasks in video content analytics (VCA), including object detection, tracking, segmentation and recognition. VCA is a relatively new field in video processing; it has generally been implemented using two chips, with the image signal processing (ISP) part in a DSP or an FPGA and the VCA part executed by a processor. However, a new generation of SoC FPGAs that incorporates a processor and an FPGA into a single chip makes it possible for a single chip to perform both ISP and VCA. This study details the hardware implementation of a real-time background and foreground identification algorithm in an SoC, including the capture, processing and display stages. The proposed platform uses photometric invariant color, depth data and local binary patterns (LBPs) to distinguish backgrounds from foregrounds. The system uses minimal cell resources and tries to implement modules using a pipeline technique.
Amin Safaei 0001, Q. M. Jonathan Wu
ISCAS2
2016 Multi-view dynamic texture learning
abstract
Dynamic texture (DT) provides a flexible and suitable tool for representing phenomena over space and time. We focus here on DT learning for multi-view domains, each of which is sufficient to learn the target concept. We make several contributions in this paper. First, we derive new features and then present their use in our description of DT. Second, we introduce multi-view dynamic texture learning, which aims to combine several views to achieve a more reliable and accurate result. The core of this multi-view combination is based on the recent information theories of the probabilistic rand index. Third, we present a way to impose spatial smoothness constraints between neighboring observations. Finally, we empirically show that our model outperforms existing state-of-the-art methods recently proposed in the literature on various datasets.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
WACV2
2016 Local directional mask maximum edge patterns for image retrieval and face recognition
abstract
This study proposes a new feature descriptor, local directional mask maximum edge pattern, for image retrieval and face recognition applications. Local binary pattern (LBP) and LBP variants collect the relationship between the centre pixel and its surrounding neighbours in an image. Thus, LBP based features are very sensitive to the noise variations in an image. Whereas the proposed method collects the maximum edge patterns (MEP) and maximum edge position patterns (MEPP) from the magnitude directional edges of face/image. These directional edges are computed with the aid of directional masks. Once the directional edges (DE) are computed, the MEP and MEPP are coded based on the magnitude of DE and position of maximum DE. Further, the robustness of the proposed method is increased by integrating it with the multiresolution Gaussian filters. The performance of the proposed method is tested by conducting four experiments onopen access series of imaging studies‐magnetic resonance imaging, Brodatz, MIT VisTex and Extended Yale B databases for biomedical image retrieval, texture retrieval and face recognition applications. The results after being investigated the proposed method shows a significant improvement as compared with LBP and LBP variant features in terms of their evaluation measures on respective databases.
Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Anil Balaji Gonde, Q. M. Jonathan Wu
IET Comput. Vis.4
2016 Non-local-based spatially constrained hierarchical fuzzy C-means method for brain magnetic resonance imaging segmentation
abstract
Owing to the existence of noise and intensity inhomogeneity in brain magnetic resonance (MR) images, the existing segmentation algorithms are hard to find satisfied results. In this study, the authors propose an improved fuzzy C ‐mean clustering method (FCM) to obtain more accurate results. First, the authors modify the traditional regularisation smoothing term by using the non‐local information to reduce the effect of the noise. Second, inspired by the mechanism of the Gaussian mixture model, the distance function of FCM is defined by using the form of certain exponential function consisting of not only the distance but also the covariance and the prior probability to improve the robustness. Meanwhile, the bias field is modelled by using orthogonal basis functions to reduce the effect of intensity inhomogeneity. Finally, they use the hierarchical strategy to construct a more flexibility function, which considers the improved distance function itself as a sub‐FCM, to make the method more robust and accurate. Compared with the state‐of‐the‐art methods, experiment results based on synthetic and real MR images demonstrate its accuracy and robustness.
Hui Zhang 0015, Yuhui Zheng, Byeungwoo Jeon, Q. M. Jonathan Wu
IET Image Process.6
2016 An online generalized eigenvalue version of Laplacian Eigenmaps for visual big data
Zeeshan Khawar Malik, Amir Hussain 0001, Q. M. Jonathan Wu
Neurocomputing3
2016 An improved anisotropic hierarchical fuzzy c-means method based on multivariate student t-distribution for brain MRI segmentation
Hui Zhang 0015, Yuhui Zheng, Byeungwoo Jeon, Q. M. Jonathan Wu
Pattern Recognit.5
2016 Full-reference image quality assessment by combining global and local distortion measures
Ashirbani Saha, Q. M. Jonathan Wu
Signal Process.2
2016 A Consensus Model for Motion Segmentation in Dynamic Scenes
abstract
The study of phenomena segmentation in natural scenes has attracted growing attention and is a popular research topic. While there are many studies detailing algorithms for motion segmentation in dynamic scenes, an important question arising from these studies is how to combine these algorithms. How can the label correspondence problem be resolved? Answering this question is difficult, because there are no labeled training data available in clustering to guide the search. Also, different algorithms produce incompatible data labels resulting in intractable correspondence problems. This paper presents a new consensus model for motion segmentation in dynamic scenes, which aims to combine several unsupervised methods to achieve a more reliable and accurate result. The advantage of our method is that it is intuitively appealing. Numerical experiments on various phenomena are conducted. The performance of the proposed model is compared with the best state-of-the-art motion segmentation methods recently proposed in the literature, demonstrating the robustness and accuracy of our method.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Circuits Syst. Video Technol.2
2016 Multilayer Extreme Learning Machine With Subnetwork Nodes for Representation Learning
abstract
The extreme learning machine (ELM), which was originally proposed for "generalized" single-hidden layer feedforward neural networks, provides efficient unified learning solutions for the applications of clustering, regression, and classification. It presents competitive accuracy with superb efficiency in many applications. However, ELM with subnetwork nodes architecture has not attracted much research attentions. Recently, many methods have been proposed for supervised/unsupervised dimension reduction or representation learning, but these methods normally only work for one type of problem. This paper studies the general architecture of multilayer ELM (ML-ELM) with subnetwork nodes, showing that: 1) the proposed method provides a representation learning platform with unsupervised/supervised and compressed/sparse representation learning and 2) experimental results on ten image datasets and 16 classification datasets show that, compared to other conventional feature learning methods, the proposed ML-ELM with subnetwork nodes performs competitively or much better than other feature learning methods.
Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Cybern.2
2016 Extreme Learning Machine With Subnetwork Hidden Nodes for Regression and Classification
abstract
As demonstrated earlier, the learning effectiveness and learning speed of single-hidden-layer feedforward neural networks are in general far slower than required, which has been a major bottleneck for many applications. Huang et al. proposed extreme learning machine (ELM) which improves the training speed by hundreds of times as compared to its predecessor learning techniques. This paper offers an ELM-based learning method that can grow subnetwork hidden nodes by pulling back residual network error to the hidden layer. Furthermore, the proposed method provides a similar or better generalization performance with remarkably fewer hidden nodes as compared to other ELM methods employing huge number of hidden nodes. Thus, the learning speed of the proposed technique is hundred times faster compared to other ELMs as well as to back propagation and support vector machines. The experimental validations for all methods are carried out on 32 data sets.
Yimin Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Cybern.2
2016 Online Feature Selection Based on Fuzzy Clustering and Its Applications
abstract
Fuzzy c-means (FCM) clustering has been successfully applied in various pattern recognition areas. While FCM is gaining attention, an important issue arising from these studies is the need to determine which attributes of the data should be used. Answering this question is difficult, because there is no labeled training data available in clustering to guide the search. We present a feature selection for FCM. The advantage of our method is that it is intuitively appealing, avoiding combinatorial searches, and allowing us to prune the feature set. Our method is also adaptable and can change through complex scenes in an online environment. We do not have to wait until all data have been generated before learning begins. Finally, to estimate the model parameters, the gradient method is adopted to minimize the fuzzy objective function with the Kullback-Leibler divergence information. Numerical experiments are presented to demonstrate the robustness and accuracy of our method.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Fuzzy Syst.2
2016 Multiple Kernel Point Set Registration
abstract
The finite Gaussian mixture model with kernel correlation is a flexible tool that has recently received attention for point set registration. While there are many algorithms for point set registration presented in the literature, an important issue arising from these studies concerns the mapping of data with nonlinear relationships and the ability to select a suitable kernel. Kernel selection is crucial for effective point set registration. We focus here on multiple kernel point set registration. We make several contributions in this paper. First, each observation is modeled using the Student's t-distribution, which is heavily tailed and more robust than the Gaussian distribution. Second, by automatically adjusting the kernel weights, the proposed method allows us to prune the ineffective kernels. This makes the choice of kernels less crucial. After parameter learning, the kernel saliencies of the irrelevant kernels go to zero. Thus, the choice of kernels is less crucial and it is easy to include other kinds of kernels. Finally, we show empirically that our model outperforms state-of-the-art methods recently proposed in the literature.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Medical Imaging2
2016 Organizing Books and Authors by Multilayer SOM
abstract
This paper introduces a new framework for the organization of electronic books (e-books) and their corresponding authors using a multilayer self-organizing map (MLSOM). An author is modeled by a rich tree-structured representation, and an MLSOM-based system is used as an efficient solution to the organizational problem of structured data. The tree-structured representation formulates author features in a hierarchy of author biography, books, pages, and paragraphs. To efficiently tackle the tree-structured representation, we used an MLSOM algorithm that serves as a clustering technique to handle e-books and their corresponding authors. A book and author recommender system is then implemented using the proposed framework. The effectiveness of our approach was examined in a large-scale data set containing 3868 authors along with the 10500 e-books that they wrote. We also provided visualization results of MLSOM for revealing the relevance patterns hidden from presented author clusters. The experimental results corroborate that the proposed method outperforms other content-based models (e.g., rate adapting poisson, latent Dirichlet allocation, probabilistic latent semantic indexing, and so on) and offers a promising solution to book recommendation, author recommendation, and visualization.
Haijun Zhang 0002, Tommy W. S. Chow, Q. M. Jonathan Wu
IEEE Trans. Neural Networks Learn. Syst.3
2015 Multiple Imputation of Missing Residuals for Fault Classification: A Wind Turbine Application
abstract
Handling the missing data is considered as a crucial requirement for the performance of diagnostic systems. In the proposed diagnostic system, the preprocessing module receives sets of residuals generated by a combined set of observers, and feeds the proceeded residuals to a fault classification module. It is necessary for the fault classification module to receive complete feature sets. Multiple missing data imputation techniques have been devised in the preprocessing module to guarantee feeding complete sets of features to the fault classification module. The proposed diagnostic scheme is validated using incomplete batch of residuals for sensor fault diagnosis in a doubly fed induction generator (DFIG) of a wind turbine.
Eman M. Nejad, Roozbeh Razavi-Far, Q. M. Jonathan Wu, Mehrdad Saif
ICMLA3
2015 An Online Adaptive Fuzzy Clustering and Its Application for Background Suppression
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu, Dibyendu Mukherjee
ICVS2
2015 A novel chaos-based secure transmission of biometric data
Gaurav Bhatnagar, Q. M. Jonathan Wu
Neurocomputing2
2015 A new contrast based multimodal medical image fusion framework
Gaurav Bhatnagar, Q. M. Jonathan Wu, Zheng Liu 0002
Neurocomputing2
2015 Spherical symmetric 3D local ternary patterns for natural, texture and biomedical image indexing and retrieval
M. Subrahmanyam 0001, Q. M. Jonathan Wu
Neurocomputing2
2015 A new robust and efficient multiple watermarking scheme
Gaurav Bhatnagar, Q. M. Jonathan Wu
Multim. Tools Appl.2
2015 A comparative experimental study of image feature detectors and descriptors
Dibyendu Mukherjee, Q. M. Jonathan Wu, Guanghui Wang 0001
Mach. Vis. Appl.2
2015 A non-parametric Bayesian model for bounded data
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
Pattern Recognit.2
2015 A Quasi-Dense Matching Approach and its Calibration Application with Internet Photos
abstract
This paper proposes a quasi-dense matching approach to the automatic acquisition of camera parameters, which is required for recovering 3-D information from 2-D images. An affine transformation-based optimization model and a new matching cost function are used to acquire quasi-dense correspondences with high accuracy in each pair of views. These correspondences can be effectively detected and tracked at the sub-pixel level in multiviews with our neighboring view selection strategy. A two-layer iteration algorithm is proposed to optimize 3-D quasi-dense points and camera parameters. In the inner layer, different optimization strategies based on local photometric consistency and a global objective function are employed to optimize the 3-D quasi-dense points and camera parameters, respectively. In the outer layer, quasi-dense correspondences are resampled to guide a new estimation and optimization process of the camera parameters. We demonstrate the effectiveness of our algorithm with several experiments.
Yanli Wan, Zhenjiang Miao, Q. M. Jonathan Wu, Xifu Wang
IEEE Trans. Cybern.3
2015 Data Partition Learning With Multiple Extreme Learning Machines
abstract
As demonstrated earlier, the learning accuracy of the single-layer-feedforward-network (SLFN) is generally far lower than expected, which has been a major bottleneck for many applications. In fact, for some large real problems, it is accepted that after tremendous learning time (within finite epochs), the network output error of SLFN will stop or reduce increasingly slowly. This report offers an extreme learning machine (ELM)-based learning method, referred to as the parent-offspring progressive learning method. The proposed method works by separating the data points into various parts, and then multiple ELMs learn and identify the clustered parts separately. The key advantages of the proposed algorithms as compared to the traditional supervised methods are twofold. First, it extends the ELM learning method from a single neural network to a multinetwork learning system, as the proposed multiELM method can approximate any target continuous function and classify disjointed regions. Second, the proposed method tends to deliver a similar or much better generalization performance than other learning methods. All the methods proposed in this paper are tested on both artificial and real datasets.
Yimin Yang 0001, Q. M. Jonathan Wu, Yaonan Wang 0001, Zeeshan Khawar Malik, Xiaofang Yuan
IEEE Trans. Cybern.2
2015 Utilizing Image Scales Towards Totally Training Free Blind Image Quality Assessment
abstract
A new approach to blind image quality assessment (BIQA), requiring no training, is proposed in this paper. The approach is named as blind image quality evaluator based on scales and works by evaluating the global difference of the query image analyzed at different scales with the query image at original resolution. The approach is based on the ability of the natural images to exhibit redundant information over various scales. A distorted image is considered as a deviation from the natural image and bereft of the redundancy present in the original image. The similarity of the original resolution image with its down-scaled version will decrease more when the image is distorted more. Therefore, the dissimilarities of an image with its low-resolution versions are cumulated in the proposed method. We dissolve the query image into its scale-space and measure the global dissimilarity with the co-occurrence histograms of the original and its scaled images. These scaled images are the low pass versions of the original image. The dissimilarity, called low pass error, is calculated by comparing the low pass versions across scales with the original image. The high pass versions of the image in different scales are obtained by Wavelet decomposition and their dissimilarity from the original image is also calculated. This dissimilarity, called high pass error, is computed with the variance and gradient histograms and weighted by the contrast sensitivity function to make it perceptually effective. These two kinds of dissimilarities are combined together to derive the quality score of the query image. This method requires absolutely no training with the distorted image, pristine images, or subjective human scores to predict the perceptual quality but uses the intrinsic global change of the query image across scales. The performance of the proposed method is evaluated across six publicly available databases and found to be competitive with the state-of-the-art techniques.
Ashirbani Saha, Q. M. Jonathan Wu
IEEE Trans. Image Process.2
2015 Asymmetric Mixture Model With Simultaneous Feature Selection and Model Detection
abstract
A mixture model based on the symmetric Gaussian distribution that simultaneously treats the feature selection, and the model detection has recently received great attention for pattern recognition problems. However, in many applications, the distribution of the data has a non-Gaussian and nonsymmetric form. This brief presents a new asymmetric mixture model for model detection and model selection. In this brief, the proposed asymmetric distribution is modeled with multiple student's- t distributions, which are heavily tailed and more robust than Gaussian distributions. Our method has the flexibility to fit different shapes of observed data, such as non-Gaussian and nonsymmetric. Another advantage is that the proposed algorithm, which is based on the variational Bayesian learning, can simultaneously optimize over the number of the student's- t distribution that is used to model each asymmetric distribution, the number of components, and the saliency of the features. Numerical experiments on both synthetic and real-world datasets are conducted. The performance of the proposed model is compared with other mixture models, demonstrating the robustness, accuracy, and effectiveness of our method.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu, Hui Zhang 0015
IEEE Trans. Neural Networks Learn. Syst.2
2015 Progressive Learning Machine: A New Approach for General Hybrid System Approximation
abstract
As the most important property of neural networks (NNs), the universal approximation capability of NNs is widely used in many applications. However, this property is generally proven for continuous systems. Most industrial systems are hybrid systems (e.g., piecewise continuous), which is a significant limitation for real applications. Recently, many identification methods have been proposed for hybrid system approximation; however, these methods only operate in linear hybrid systems. In this paper, the progressive learning machine-a new learning algorithm based on multi-NNs-is proposed for general hybrid nonlinear/linear system approximation. This algorithm classifies hybrid systems into several continuous systems and can approximate any hybrid system with zero output error. The performance of the proposed learning method is demonstrated via numerical examples and with experimental data from real applications.
Yimin Yang 0001, Yaonan Wang 0001, Q. M. Jonathan Wu, Min Liu 0008
IEEE Trans. Neural Networks Learn. Syst.3
2014 Streaming spatio-temporal video segmentation using Gaussian Mixture Model
abstract
Development of an automatic streaming video segmentation method is crucial for many video analysis applications. However, consistency of temporal segmentation and scalability for real-time applications are difficult to achieve. This work proposes a linear-time video segmentation method which is scalable and temporally consistent for streaming videos. A Gaussian Mixture Model (GMM) is used to segment each frame while a recursive filtering updates the parameters of the GMM. This hybrid methodology can uniquely propagate Gaussian clusters through each new frame, update the variance recursively, and create or remove clusters as necessary. In this way, the model automatically manipulates the number of clusters in run-time and adapts to any video sequence over streaming frames maintaining temporal coherence. The method needs a distance threshold value as the main parameter. The creation and removal of new clusters are governed by a cluster similarity criterion that can be based on user-defined distance measure. The experimental results are presented with two possible distance measures. The performance of the proposed method on several datasets is found to be comparable to state-of-the-art video segmentation algorithms.
Dibyendu Mukherjee, Q. M. Jonathan Wu
ICIP2
2014 Asymmetric mixture model with variational Bayesian learning
abstract
Bayesian detection for the symmetric Gaussian mixture model has recently received great attention for pattern recognition problems. However, in many applications, the distribution of the data has a non-Gaussian and non-symmetric form. This study presents a new asymmetric mixture model for model detection. In this paper, the proposed asymmetric distribution is modeled with multiple Student's-t distributions, which are heavily tailed and more robust than Gaussian distributions. Our method has the flexibility to fit different shapes of observed data such as non-Gaussian and non-symmetric. Another advantage is that the proposed algorithm, which is based on the variational Bayesian learning, can simultaneously optimize over the number of the Student's-t distribution that is used to model each asymmetric distribution, and the number of components. The performance of the proposed model is compared to other mixture models, demonstrating the robustness, accuracy, and effectiveness of our method.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IJCNN2
2014 Dual Gaussian mixture model with pixel history for background suppression
abstract
Background suppression has found several application areas in computer vision. Among others, abandoned object detection and moving object detection under dynamic backgrounds are two major application areas. Gaussian mixture model (GMM) is one of the most popular methods and has been applied in both of these application areas. However, the aforementioned areas present contradicting challenges for GMM. Abandoned objects need the GMM to slowly get adapted to prevent accidental assimilation of the foreground, while dynamic backgrounds need quick assimilation to prevent noisy detection. A novel dual GMM based on pixel history is proposed to provide practical solutions to both of these contradictory challenges. The method uses separate GMMs to represent foreground and background modes. A foreground mode is transferred to the GMM representing the background based on its history of occurrence. Thus, a quick assimilation is offered to dynamic backgrounds while abandoned objects face a very slow assimilation. In each case, old background is preserved. Extensive experiments are done to demonstrate the effectiveness of the proposed method and to compare it with a number of well known methods in literature.
Dibyendu Mukherjee, Ashirbani Saha, Q. M. Jonathan Wu, Wei Jiang 0009
SMC3
2014 Multimodal 3D histogram for moving object detection
abstract
Moving object detection is a popular research area due to its vast application in computer vision. Based on application areas, a number of moving object detection methods have been discovered. This work proposes a real-time method based on 3D histogram and temporal multiple mode selection, suited towards a vast majority of dynamic and noisy backgrounds, congested backgrounds and slow foregrounds. It is a generalization of [1] demonstrating improved performance. The temporal distribution of a video sequence can provide a history of the foreground motion and relatively static background. However, a dynamic background is identified by multiple intensity levels in the temporal distribution, and hence, cannot be properly modeled by a single histogram mode. The proposed multiple mode selection process involves identifying the dominant modes of the temporal distribution and assigning them to a multimodal background. The work provides a detailed analysis of the proposition, a clear description of the proposed algorithm and an extensive number of tests with some of the well-known methods in literature. It also demonstrates the improvement over its predecessor in several types of scenarios covered by a large number of datasets. Both qualitative and quantitative comparisons are carried out to demonstrate the applicability and accuracy of the proposed method.
Dibyendu Mukherjee, Ashirbani Saha, Q. M. Jonathan Wu, Wei Jiang 0009
SMC3
2014 Robust logo watermarking using biometrics inspired key generation
Gaurav Bhatnagar, Q. M. Jonathan Wu, Pradeep K. Atrey
Expert Syst. Appl.2
2014 Image segmentation by dirichlet process mixture model with generalised mean
abstract
The Dirichlet process mixture model (DPMM) with spatial constraints – e.g. hidden Markov random field (HMRF) model – has been considered as an effective algorithm for image processing application. However, the HMRF model is complex and time‐consuming for implementation. A new DPMM has been introduced, where a generalised mean (GDM) is selected as the spatial constraints function. The GDM is applied not only on prior probability (and posterior probability) to incorporate local spatial information and component information, but also on conditional probability to incorporate local spatial information and observation information. The purpose of the HMRF model and GDM are the same for incorporating some spatial constraints into the system. However, compared to HMRF, GDM is easier, faster and simpler to implement. Finally, a variational Bayesian approach has been adopted for parameters estimation and model selection. Experimental results on image segmentation application demonstrate the improved performance of the proposed approach.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
IET Image Process.2
2014 Effective fuzzy clustering algorithm with Bayesian model and mean template for image segmentation
abstract
Fuzzy c‐means (FCMs) with spatial constraints have been considered as an effective algorithm for image segmentation. The well‐known Gaussian mixture model (GMM) has also been regarded as a useful tool in several image segmentation applications. In this study, the authors propose a new algorithm to incorporate the merits of these two approaches and reveal some intrinsic relationships between them. In the authors model, the new objective function pays more attention on spatial constraints and adopts Gaussian distribution as the distance function. Thus, their model can degrade to the standard GMM as a special case. Our algorithm is fully free of the empirically pre‐defined parameters that are used in traditional FCM methods to balance between robustness to noise and effectiveness of preserving the image sharpness and details. Furthermore, in their algorithm, the prior probability of an image pixel is influenced by the fuzzy memberships of pixels in its immediate neighbourhood to incorporate the local spatial information and intensity information. Finally, they utilise the mean template instead of the traditional hidden Markov random field (HMRF) model for estimation of prior probability. The mean template is considered as a spatial constraint for collecting more image spatial information. Compared with HMRF, their method is simple, easy and fast to implement. The performance of their proposed algorithm, compared with state‐of‐the‐art technologies including extensions of possibilistic fuzzy c‐means (PFCM), GMM, FCM, HMRF and their hybrid models, demonstrates its improved robustness and effectiveness.
Hui Zhang 0015, Q. M. Jonathan Wu, Yuhui Zheng, Thanh Minh Nguyen 0001, Dingcheng Wang
IET Image Process.2
2014 Analysis and extension of multiresolution singular value decomposition
Gaurav Bhatnagar, Ashirbani Saha, Q. M. Jonathan Wu, Pradeep K. Atrey
Inf. Sci.3
2014 Expert content-based image retrieval system using robust local patterns
M. Subrahmanyam 0001, Q. M. Jonathan Wu
J. Vis. Commun. Image Represent.2
2014 Enhancing the transmission security of biometric images using chaotic encryption
Gaurav Bhatnagar, Q. M. Jonathan Wu
Multim. Syst.2
2014 Bounded generalized Gaussian mixture model
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu, Hui Zhang 0015
Pattern Recognit.2
2014 MRI and CT image indexing and retrieval using local mesh peak valley edge patterns
M. Subrahmanyam 0001, Q. M. Jonathan Wu
Signal Process. Image Commun.2
2014 Bounded Asymmetrical Student's-t Mixture Model
abstract
The finite mixture model based on the Student's-t distribution, which is heavily tailed and more robust than the Gaussian mixture model (GMM), is a flexible and powerful tool to address many computer vision and pattern recognition problems. However, the Student's-t distribution is unbounded and symmetrical around its mean. In many applications, the observed data are digitalized and have bounded support. The distribution of the observed data usually has an asymmetric form. A new finite bounded asymmetrical Student's-t mixture model (BASMM), which includes the GMM and the Student's-t mixture model (SMM) as special cases, is presented in this paper. We propose an extension of the Student's-t distribution in this paper. This new distribution is sufficiently flexible to fit different shapes of observed data, such as non-Gaussian, nonsymmetric, and bounded support data. Another advantage of the proposed model is that each of its components can model the observed data with different bounded support regions. In order to estimate the model parameters, previous models represent the Student's-t distributions as an infinite mixture of scaled Gaussians. We propose an alternate approach in order to minimize the higher bound on the data negative log-likelihood function, and directly deal with the Student's-t distribution. As an application, our method has been applied to image segmentation with promising results.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Cybern.2
2014 Synthetic Aperture Radar Image Segmentation by Modified Student's t-Mixture Model
abstract
Synthetic aperture radar (SAR) data are often affected by speckle noise, which originates in the SAR system's coherent nature. In this paper, we introduce a simple and effective algorithm to make the traditional Student's t-mixture model (SMM) more robust to noise. The proposed new modified SMM (MSMM) is applied for SAR image segmentation. SMM has come to be regarded as an alternative to the Gaussian mixture model (GMM) as it is heavy tailed and more robust to outliers. However, a major shortcoming of this method is that it does not take into account the spatial dependencies in the image. Although some existing methods incorporate the spatial relationship between neighboring pixels, they are still not robust enough to noise. The advantages of our method are as follows. First, we introduce MSMM to incorporate the local spatial information and pixel intensity value by considering the conditional probability of an image pixel influenced by the probabilities of pixels in its immediate neighborhood. Furthermore, we introduce the additional parameter α to control the extent of this influence. The larger α indicates the heavier extent of influence in the neighborhoods. Second, the prior probability of an image pixel is influenced by the probabilities of pixels in its immediate neighborhood, which incorporates local spatial and component information. Third, our model is based on the finite mixture model (FMM); it is simple and easy to implement, and the expectation maximization algorithm can be applied for estimation of optimal parameters. Finally, the traditional SMM can be considered as a special case of our model. Thus, our method is general enough for FMM-based techniques. Experimental results on both simulated and real SAR images demonstrate the improved robustness and effectiveness of our approach.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001, Xingming Sun
IEEE Trans. Geosci. Remote. Sens.2
2014 Gaussian Mixture Model With Advanced Distance Measure Based on Support Weights and Histogram of Gradients for Background Suppression
abstract
The proposed work is targeted toward improving the Gaussian mixture model (GMM) for the background suppression-based moving object detection. The GMM has been widely used for moving object detection due to its high applicability. However, the GMM cannot properly model noisy or nonstationary backgrounds and fails to discriminate between the foreground and background modes. The extensions to GMM provide increased accuracy in expense of complex implementation and reduced applicability. In response, this work proposes two simple improvements: 1) a novel distance measure based on local support weights and histogram of gradients to provide distinct cluster values; and 2) use of background layer concept to properly segment the foreground. The method also uses variable number of clusters for generalization. The main advantages of the method are implicit use of pixel relationships through distance measure with least modification to the conventional GMM and effective background noise removal through the use of background layer concept with no postprocessing involved. The extensive experimentations on various types of video sequences are performed to validate the improvement in accuracy compared to the GMM and a number of state-of-the-art methods.
Dibyendu Mukherjee, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
IEEE Trans. Ind. Informatics2
2014 An Unsupervised Feature Selection Dynamic Mixture Model for Motion Segmentation
abstract
The automatic clustering of time-varying characteristics and phenomena in natural scenes has recently received great attention. While there exist many algorithms for motion segmentation, an important issue arising from these studies concerns that for which attributes of the data should be used to cluster phenomena with a certain repetitiveness in both space and time. It is difficult because there is no knowledge about the labels of the phenomena to guide the search. In this paper, we present a feature selection dynamic mixture model for motion segmentation. The advantage of our method is that it is intuitively appealing, avoiding any combinatorial search, and allowing us to prune the feature set. Numerical experiments on various phenomena are conducted. The performance of the proposed model is compared with that of other motion segmentation algorithms, demonstrating the robustness and accuracy of our method.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Image Process.2
2014 A Bayesian Bounded Asymmetric Mixture Model With Segmentation Application
abstract
Segmentation of a medical image based on the modeling and estimation of the tissue intensity probability density functions via a Gaussian mixture model has recently received great attention. However, the Gaussian distribution is unbounded and symmetrical around its mean. This study presents a new bounded asymmetric mixture model for analyzing both univariate and multivariate data. The advantage of the proposed model is that it has the flexibility to fit different shapes of observed data such as non-Gaussian, nonsymmetric, and bounded support data. Another advantage is that each component of the proposed model has the ability to model the observed data with different bounded support regions, which is suitable for application on image segmentation. Our method is intuitively appealing, simple, and easy to implement. We also propose a new method to estimate the model parameters in order to minimize the higher bound on the data negative log-likelihood function. Numerical experiments are presented where the proposed model is tested in various images from simulated to real 3- D medical ones.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu, Dibyendu Mukherjee, Hui Zhang 0015
IEEE J. Biomed. Health Informatics2
2014 Local Mesh Patterns Versus Local Binary Patterns: Biomedical Image Indexing and Retrieval
abstract
In this paper, a new image indexing and retrieval algorithm using local mesh patterns are proposed for biomedical image retrieval application. The standard local binary pattern encodes the relationship between the referenced pixel and its surrounding neighbors, whereas the proposed method encodes the relationship among the surrounding neighbors for a given referenced pixel in an image. The possible relationships among the surrounding neighbors are depending on the number of neighbors, P. In addition, the effectiveness of our algorithm is confirmed by combining it with the Gabor transform. To prove the effectiveness of our algorithm, three experiments have been carried out on three different biomedical image databases. Out of which two are meant for computer tomography (CT) and one for magnetic resonance (MR) image retrieval. It is further mentioned that the database considered for three experiments are OASIS-MRI database, NEMA-CT database, and VIA/I-ELCAP database which includes region of interest CT images. The results after being investigated show a significant improvement in terms of their evaluation measures as compared to LBP, LBP with Gabor transform, and other spatial and transform domain methods.
M. Subrahmanyam 0001, Q. M. Jonathan Wu
IEEE J. Biomed. Health Informatics2
2014 Biometric Inspired Multimedia Encryption Based on Dual Parameter Fractional Fourier Transform
abstract
In this paper, a novel biometric inspired multimedia encryption technique is proposed. For this purpose, a new advent in the definition of fractional Fourier transform, namely, dual parameter fractional Fourier transform (DP-FrFT) is proposed and used in multimedia encryption. The core idea behind the proposed encryption technique is to obtain biometrically encoded bitstream followed by the generation of the keys used in the encryption process. Since the key generation process of encryption technique directly determines the security of the technique. Therefore, this paper proposes an efficient method for generating biometrically encoded bitstream from biometrics and its usage to generate the keys. Then, the encryption of multimedia data is done in the DP-FrFT domain with the help of Hessenberg decomposition and nonlinear chaotic map. Finally, a reliable decryption process is proposed to construct original multimedia data from the encrypted data. Theoretical analyses and computer simulations both confirm high security and efficiency of the proposed encryption technique.
Gaurav Bhatnagar, Q. M. Jonathan Wu
IEEE Trans. Syst. Man Cybern. Syst.2
2014 Low-resolution face recognition: a review
Zhenjiang Miao, Q. M. Jonathan Wu, Yanli Wan
Vis. Comput.3
2013 Image segmentation by a robust Modified Gaussian Mixture Model
abstract
The Gaussian Mixture Model (GMM) with a spatial constraint, e.g. a Hidden Markov Random Field (HMRF), has been proven effective for image segmentation. However, the determination of parameter β in the HMRF model is, in fact, noise dependent to some degree. In this paper, we propose a simple and effective algorithm to make the traditional Gaussian Mixture Model more robust to noise, with consideration of the relationship between the local spatial information and the pixel intensity value information. The conditional probability of an image pixel is influenced by the probabilities of pixels in its immediate neighborhood to incorporate the spatial and intensity information. In this case, the parameter β can be assigned to a small value to preserve image sharpness and detail in non-noise images. At the same time, the neighborhood window is used to tolerate the noise for heavy-noised images. Thus, the parameter β is independent of image noise degree in our model. Finally, our algorithm is not limited to GMM-it is general enough so that it can be applied to other distributions based on the construction of the Finite Mixture Model (FMM) technique.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
ICASSP2
2013 Multivariate Student's-t mixture model for bounded support data
abstract
The finite mixture model based on the Student's-t distribution, which is heavily tailed and more robust than the Gaussian mixture model (GMM), is a flexible and powerful tool to address many pattern recognition problems. However, the Student's-t distribution is unbounded. In many applications, the observed data are digitalized and have bounded support. A new finite multivariate Student's-t mixture model for bounded support data, which includes the GMM and the Student's-t mixture model (SMM) as special cases, is presented in this paper. We propose an extension of the Student's-t distribution in this paper. This new distribution is sufficiently flexible to fit different shapes of observed data, such as non-Gaussian, non-symmetric, and bounded support data. Another advantage of the proposed model is that each of its components can model the observed data with different bounded support regions. In order to estimate the model parameters, previous models represent the Student's-t distributions as an infinite mixture of scaled Gaussians. We propose an alternate approach in order to minimize the higher bound on the data negative log-likelihood function, and directly deal with the Student's-t distribution.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
ICASSP2
2013 Bounded asymmetric mixture model for medical image segmentation
abstract
Segmentation of medical image based on the modeling and estimation of the tissue intensity probability density functions via Gaussian mixture model (GMM) has recently received great attention. However, Gaussian distribution is unbounded and symmetrical around its mean. This study presents a new bounded asymmetric mixture model for analyzing both univariate and multivariate data. The advantage of the proposed model is that it has the flexibility to fit different shapes of observed data such as non-Gaussian, non-symmetric, and bounded support data. Another advantage is that each component of the proposed model has the ability to model the observed data with different bounded support regions, which is suitable for application on image segmentation. Our method is intuitively appealing, simple, and easy to implement. We also propose a new method to estimate the model parameters in order to minimize the higher bound on the data negative log-likelihood function. Numerical experiments are presented where the proposed model is tested in various images from simulated to real 3D medical ones.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu, Dibyendu Mukherjee, Hui Zhang 0015
ICASSP2
2013 A detail-preserving mixture model for image segmentation
abstract
This paper presents a new finite mixture model for image segmentation. First, in order to take into account the spatial dependencies in an image, existing mixture models use a constant temperature parameter (β) throughout the image for every label. The constant value of β reduces the impact of noise in homogeneous regions but negatively affects segmentation along the border of two regions. We propose a new way to use a different value of β throughout the image. Secondly, in order to incorporate the correlation between each centre pixel and its neighboring pixels, existing mixture model gives the same importance to all pixels in a neighborhood window. We assign different weights to different pixels appearing in the window, which is based on the fact that the clique strength should be reduced with distance. Thirdly, our model is based on the Student's-t distribution, which is heavily tailed and more robust than Gaussian. We exploit Dirichlet distribution and Dirichlet law to incorporate the spatial relationships between pixels in an image. Finally, expectation maximization (EM) algorithm is adopted to maximize the data log-likelihood and to optimize the parameters. The performance is compared to other existing models based on the model-based techniques, demonstrating superiority of the proposed model for image segmentation.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu, Dibyendu Mukherjee, Hui Zhang 0015
ICASSP2
2013 A study on using spectral saliency detection approaches for image quality assessment
abstract
Recent developments in the field of full reference image quality assessment (FR-IQA) have witnessed the use of spectral residual (SR) based index as a fast measure with high accuracy. Following SR, several variants of spectral measures for visual saliency have come up. These new measures differ in their computational times as well as in performances and have established themselves better than or competitive with SR as measures of visual saliency. The effectiveness of these measures in FR-IQA is still an open question. In this paper, a study to evaluate the performance of the recent spectral approaches for visual saliency (hence spectral saliency) for FR-IQA is presented. We have fixed a framework for FR-IQA to maintain uniformity in the evaluation process. Also, the parameters required by the framework are chosen to bring out the best potential of each measure. Our experiments on six benchmark databases reveal some insightful details about the usage of these measures to form an FR-IQA measure.
Ashirbani Saha, Q. M. Jonathan Wu
ICASSP2
2013 An effective fuzzy clustering algorithm for image segmentation
abstract
Fuzzy c-means (FCM) with spatial constraints has been considered as an effective algorithm for image segmentation. In this paper, we propose a new algorithm to incorporate the local spatial information with the consideration of mean template. Our algorithm is fully free of the empirically predefined parameters that are used in other FCM methods to balance between robustness to noise and effectiveness of preserving the image sharpness and details. Furthermore, in our algorithm, the prior probability of an image pixel is influenced by the fuzzy memberships of pixels in its immediate neighborhood to incorporate the local spatial information and intensity information. Finally, we utilize the mean template instead of the traditional hidden Markov random field (HMRF) model for estimation of prior probability. Compared to HMRF, our method is simple, easy and fast to implement.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
ICASSP2
2013 Image segmentation by a robust generalized fuzzy c-means algorithm
abstract
Fuzzy c-means (FCM) has been considered as an effective algorithm for image segmentation. However, it lacks of sufficient robustness to image noise. In this paper, we propose a simple and effective method to make the traditional FCM more robust to noise, with the help of generalized mean. Traditional FCM can be considered as a linear combination of membership and distance (function) from the expression of its mathematical formula. The proposed generalized FCM (GFCM) is generated by applying generalized mean on these two items. We impose generalized mean on membership to incorporate local spatial information and cluster information, and on distance function to incorporate local spatial information and observation information (image intensity value). Thus, our GFCM is more robust to image noise with the spatial constraints: the generalized mean. The performance of our proposed algorithm, compared with state-of-the-art technologies including modified FCM, HMRF and their hybrid models, demonstrates its improved robustness and effectiveness.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
ICIP2
2013 Robust rank-4 affine factorization for structure from motion
abstract
The paper focuses on 3D structure and motion factorization from uncalibrated image sequences. A rank-4 affine factorization algorithm and a robust structure and motion factorization scheme are proposed to handle outlying and missing data. The novelty and main contribution of the paper are as follows: (i) The rank-4 factorization algorithm is a new addition to previous affine factorization family using rank-3 constraint; (ii) the outliers and image uncertainty are estimated directly from the image reprojection residuals; and (iii) the robust factorization scheme is proved empirically to be more efficient and accurate than other robust algorithms. Extensive experiments on synthetic data and real images validate the proposed approach.
Guanghui Wang 0001, John S. Zelek, Q. M. Jonathan Wu, Ruzena Bajcsy
WACV3
2013 Human visual system inspired multi-modal medical image fusion framework
Gaurav Bhatnagar, Q. M. Jonathan Wu, Zheng Liu 0002
Expert Syst. Appl.2
2013 Biometrics inspired watermarking based on a fractional dual tree complex wavelet transform
Gaurav Bhatnagar, Q. M. Jonathan Wu
Future Gener. Comput. Syst.2
2013 Image segmentation by a new weighted student's t-mixture model
abstract
In this study, the authors introduce a new weighted Student's t ‐mixture model (WSMM) for image segmentation. Gaussian distribution and Student's t ‐distribution are the two commonly used probabilities in the finite mixture model (FMM). The Student's t ‐mixture model has come to be regarded as an alternative to Gaussian mixture models, as it is heavily tailed and more robust for outliers. Moreover, the pixels are considered independent of each other in the FMM. Although some existing methods incorporate the spatial relationship between neighbouring pixels, they do not consider the relationship between spatial information and clustering information, thus those reported methods remain sensitive to noise. The advantages of the authors method are as follows: first, the authors introduce WSMM to incorporate the local spatial information, pixel intensity value and clustering information in an image. Second, the authors model is simple, easy to implement and has a good balance between noise insensitiveness and image detail preservation. Third, they adopt the gradient method and expectation maximisation algorithm, which allow for simultaneous estimation of optimal parameters. Finally, the most useful statistical tool for image segmentation, the well‐known hidden Markov random field model, is a special case of their model. Thus, their method is general enough for model‐based techniques construction. Experimental results on synthetic and real images demonstrate the improved robustness and effectiveness of their approach.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
IET Image Process.2
2013 Modified student's t-hidden Markov model for pattern recognition and classification
abstract
The Gaussian hidden Markov model has been successfully used in pattern recognition and classification applications; however, recently the Student's t ‐mixture model is regarded as an alternative to Gaussian mixture models, as it is more robust for outliers. The model using Student's t ‐mixture distribution as its hidden state is the Student's t ‐hidden Markov model (SHMM). The authors propose a novel Student's t ‐hidden Markov model, which considers the relationship among Markov states, latent components and observations by introducing a regularising scalar exponent in the component densities of the model's emission densities. Moreover, the standard SHMM can be considered as a special case of the modified SHMM with the selection of proper parameter values. Finally, the authors adopt the gradient method to estimate optimal weight parameters. Simultaneously, the expectation–maximisation algorithm is used to fit the modified SHMM. Thus, our model is simple and easy to implement. The experimental results using synthetic and real data demonstrate the improved robustness of the proposed approach.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
IET Signal Process.2
2013 Local ternary co-occurrence patterns: A new feature descriptor for MRI and CT image retrieval
M. Subrahmanyam 0001, Q. M. Jonathan Wu
Neurocomputing2
2013 Variational Bayes and Localized Feature Selection for Student's T-Mixture Models
abstract
In this paper, we propose a novel algorithm for feature selection and model detection using Student's t-distribution based on the variational Bayesian (VB) approach. First, our method is based on the Student's t-mixture model (SMM) which has heavier tail than the Gaussian distribution and is therefore less sensitive to small numbers of data points and consequent precision-estimates of the components number. Second, the number of components, the local feature saliency and the parameters of the mixture model are simultaneously estimated by Bayesian variational learning. Experimental results using synthetic and real data demonstrate the improved robustness of our approach.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
Int. J. Pattern Recognit. Artif. Intell.2
2013 Discrete fractional wavelet transform and its application to multiple encryption
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Inf. Sci.2
2013 A new aspect in robust digital watermarking
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Multim. Tools Appl.2
2013 An efficient illumination invariant face recognition framework via illumination enhancement and DD-DTCWT filtering
Aryaz Baradarani, Q. M. Jonathan Wu, Majid Ahmadi
Pattern Recognit.2
2013 A finite mixture model for detail-preserving image segmentation
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu, Dibyendu Mukherjee, Hui Zhang 0015
Signal Process.2
2013 Perceptual image quality assessment using phase deviation sensitive energy features
Ashirbani Saha, Q. M. Jonathan Wu
Signal Process.2
2013 A Robust Fuzzy Algorithm Based on Student's t-Distribution and Mean Template for Image Segmentation Application
abstract
Fuzzy c-means (FCM) with spatial constraints has been considered as an effective algorithm for image segmentation. Student's t-distribution has come to be regarded as an alternative to Gaussian distribution, as it is heavily tailed and more robust for outliers. In this letter, we propose a new algorithm to incorporate the merits of these two approaches. The advantages of our method are as follows: First, we incorporate the local spatial information and pixel intensity value by considering the labeling of an image pixel influenced by the labels in its immediate neighborhood. Second, we introduce additional parameterato control the extent of this influence. The largeraindicates heavier extent of influence in the neighborhoods. Finally, we utilize a mean template instead of the traditional hidden Markov random field (HMRF) model for estimation of prior probability. Compared with HMRF, our method is simple, easy and fast to implement. Experimental results on synthetic and real images demonstrate the improved robustness and effectiveness of our approach.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
IEEE Signal Process. Lett.2
2013 Fast and Robust Spatially Constrained Gaussian Mixture Model for Image Segmentation
abstract
In this paper, a new mixture model for image segmentation is presented. We propose a new way to incorporate spatial information between neighboring pixels into the Gaussian mixture model based on Markov random field (MRF). In comparison to other mixture models that are complex and computationally expensive, the proposed method is fast and easy to implement. In mixture models based on MRF, the M-step of the expectation-maximization (EM) algorithm cannot be directly applied to the prior distribution πijfor maximization of the log-likelihood with respect to the corresponding parameters. Compared with these models, our proposed method directly applies the EM algorithm to optimize the parameters, which makes it much simpler. Experimental results obtained by employing the proposed method on many synthetic and real-world grayscale and colored images demonstrate its robustness, accuracy, and effectiveness, compared with other mixture models.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Circuits Syst. Video Technol.2
2013 A Nonsymmetric Mixture Model for Unsupervised Image Segmentation
abstract
Finite mixture models with symmetric distribution have been widely used for many computer vision and pattern recognition problems. However, in many applications, the distribution of the data has a non-Gaussian and nonsymmetric form. This paper presents a new nonsymmetric mixture model for image segmentation. The advantage of our method is that it is simple, easy to implement, and intuitively appealing. In this paper, each label is modeled with multiple D-dimensional Student's t-distribution, which is heavily tailed and more robust than Gaussian distribution. Expectation-maximization algorithm is adopted to estimate model parameters and to maximize the lower bound on the data log-likelihood from observations. Numerical experiments on various data types are conducted. The performance of the proposed model is compared with that of other mixture models, demonstrating the robustness, accuracy, and effectiveness of our method.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Cybern.2
2013 Multidimensional Latent Semantic Analysis Using Term Spatial Information
abstract
In this paper, we consider the problem of in-depth document analysis. In particular, we propose a novel document analysis method, named multidimensional latent semantic analysis (MDLSA), which enables us to mine local information efficiently from a document with respect to term associations and spatial distributions. MDLSA works by first partitioning each document into paragraphs and building a term affinity graph, which represents the frequency of term cooccurrence in a paragraph. We then conduct a 2-D principal component analysis to achieve an optimal semantic mapping. This analysis involves finding the leading eigenvectors of the sample covariance matrix of a training set to characterize the lower dimensional semantic space. A hybrid document similarity measure is designed to further improve the performance of this framework. Our algorithm is examined in two document applications: retrieval and classification. Experimental results demonstrate that the proposed technique outperforms current algorithms with respect to accuracy and computational efficiency.
Haijun Zhang 0002, John K. L. Ho, Q. M. Jonathan Wu, Yunming Ye
IEEE Trans. Cybern.3
2013 Dynamic Fuzzy Clustering and Its Application in Motion Segmentation
abstract
Dynamic textures are common in natural scenes and have recently received great attention in video content analysis. A dynamic fuzzy clustering to automatically segment time-varying characteristics and phenomena is presented in this paper. First, compared with the existing models that assume a common prior distribution, which independently generates the labels, the prior distribution in our model is different for each observation and depends on the labels. In addition, in order to properly account for the neighboring observations during the learning step, we introduce the explicit assumptions of the hidden Markov random field model into the dynamic fuzzy clustering. Second, in order to model the observed dynamic texture data, only grayscale information is taken into consideration of the existing models. We use different visual properties by proposing a new distribution in this paper. Finally, to estimate the model parameters, the gradient method is adopted to minimize the fuzzy objective function with the Kullback–Leibler divergence information. Numerical experiments are presented, where the proposed model is tested on various simulated and real dynamic textures.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Fuzzy Syst.2
2013 Multiresolution Based Gaussian Mixture Model for Background Suppression
abstract
This paper aims toward improving background suppression from video frames by incorporating multiresolution features in Gaussian mixture model (GMM). GMM has proven its place for background modeling due to its better applicability and robustness compared with other popular methods in literature. However, GMM fails in a number of situations such as noisy and non-stationary background, slow foregrounds, and illumination variation. Extensions to GMM have also been proposed to increase accuracy in expense of increased complexity, decrease in execution speed, and reduced applicability. In view of the above, this paper aims to provide a methodology to assimilate useful multiresolution features with GMM that considerably improves the performance. The contributions of this paper are: 1) a novel framework to incorporate wavelet subbands in GMM to improve its performance; 2) an approach to incorporate variable number of clusters in the aforesaid framework; and 3) a generic platform to use any multiresolution decomposition based GMM for background suppression. Extensive experimentations on several video sequences are performed to verify the improvement in accuracy compared with conventional GMM as well as a number of state-of-the-arts approaches. Along with qualitative and quantitative analysis, justification on the use of multiresolution is provided for clarification.
Dibyendu Mukherjee, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
IEEE Trans. Image Process.2
2013 Directive Contrast Based Multimodal Medical Image Fusion in NSCT Domain
abstract
Multimodal medical image fusion, as a powerful tool for the clinical applications, has developed with the advent of various imaging modalities in medical imaging. The main motivation is to capture most relevant information from sources into a single output, which plays an important role in medical diagnosis. In this paper, a novel fusion framework is proposed for multimodal medical images based on non-subsampled contourlet transform (NSCT). The source medical images are first transformed by NSCT followed by combining low- and high-frequency components. Two different fusion rules based on phase congruency and directive contrast are proposed and used to fuse low- and high-frequency coefficients. Finally, the fused image is constructed by the inverse NSCT with all composite coefficients. Experimental results and comparative study show that the proposed fusion framework provides an effective way to enable more accurate analysis of multimodality images. Further, the applicability of the proposed framework is carried out by the three clinical examples of persons affected with Alzheimer, subacute stroke and recurrent tumor.
Gaurav Bhatnagar, Q. M. Jonathan Wu, Zheng Liu 0002
IEEE Trans. Multim.2
2013 Editorial A Successful Change From TNN to TNNLS and a Very Successful Year
abstract
This issue marks the first anniversary issue of IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS after it changed its name from IEEE TRANSACTIONS ON NEURAL NETWORKS. I am happy to report that we had a great year! The number of new submissions in a year exceeded 1,000 for the first time in the history of TNN/TNNLS. IEEE TNN had a very successful development for 22 years from 1990 to 2011, and we have good reasons to believe that IEEE TNNLS will have many more years of successful growth.
Derong Liu 0001, Charles W. Anderson, Ahmad Taher Azar, Giorgio Battistelli, Eduardo Bayro-Corrochano, Cristiano Cervellera, David A. Elizondo, Maurizio Filippone, Giorgio Gnecco, Tingwen Huang, Weifeng Liu 0016, Wenlian Lu, Ana Madureira, Igor Skrjanc, Thomas Villmann, Q. M. Jonathan Wu, Shengli Xie 0001, Dong Xu 0001
IEEE Trans. Neural Networks Learn. Syst.17
2013 Incorporating Mean Template Into Finite Mixture Model for Image Segmentation
abstract
The well-known finite mixture model (FMM) has been regarded as a useful tool for image segmentation application. However, the pixels in FMM are considered independent of each other and the spatial relationship between neighboring pixels is not taken into account. These limitations make the FMM more sensitive to noise. In this brief, we propose a simple and effective method to make the traditional FMM more robust to noise with the help of a mean template. FMM can be considered a linear combination of prior and conditional probability from the expression of its mathematical formula. We calculate these probabilities with two mean templates: a weighted arithmetic mean template and a weighted geometric mean template. Thus, in our model, the prior probability (or conditional probability) of an image pixel is influenced by the probabilities of pixels in its immediate neighborhood to incorporate the local spatial and intensity information for eliminating the noise. Finally, our algorithm is general enough and can be extended to any other FMM-based models to achieve super performance. Experimental results demonstrate the improved robustness and effectiveness of our approach.
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
IEEE Trans. Neural Networks Learn. Syst.2
2013 Secure randomized image watermarking based on singular value decomposition
abstract
In this article, a novel logo watermarking scheme is proposed based on wavelet frame transform, singular value decomposition and automatic thresholding. The proposed scheme essentially rectifies the ambiguity problem in the SVD-based watermarking. The core idea is to randomly upscale the size of host image using reversible random extension transform followed by the embedding of logo watermark in the wavelet frame domain. After embedding, a verification phase is casted with the help of a binary watermark and toral automorphism. At the extraction end, the binary watermark is first extracted followed by the verification of watermarked image. The logo watermark is extracted if and only if the watermarked image is verified. The security, attack and comparative analysis confirm high security, efficiency and robustness of the proposed watermarking system.
Gaurav Bhatnagar, Q. M. Jonathan Wu, Pradeep K. Atrey
ACM Trans. Multim. Comput. Commun. Appl.2
2012 Bilateral filter based mixture model for image segmentation
abstract
This paper introduces a bilateral filtering based mixture model for image segmentation. The mixture model uses Markov Random Field (MRF) to incorporate spatial relationship among neighboring pixels into the Gaussian Mixture Model (GMM) in order to perform a segmentation that is robust against noise and other environmental factors. The bilateral filtering is used to smooth the posterior probability map as part of the MRF used. The advantage of the proposed model is its simplified structure so that the Expectation Maximization algorithm can be directly applied to the log-likelihood function to compute the optimum parameters of the mixture model. The method has been extensively tested on synthetic and natural images and compared with some of the state-of-the-arts algorithms currently available. The experimental results show that the proposed method is comparable to the other methods in terms of accuracy and quality and simpler in terms of implementation.
Dibyendu Mukherjee, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
ICIP2
2012 A robust non-symmetric mixture models for image segmentation
abstract
Finite mixture model with symmetric distribution has been widely used for many computer vision and pattern recognition problems. However, in many applications, the distribution of the data has a non-Gaussian and non-symmetric form. This study presents a new non-symmetric mixture model for image segmentation. The advantage of our method is that it is simple, easy to implement and intuitively appealing. In this paper, each label is modeled with multiple D-dimensional Student's-t distribution, which is heavily tailed and more robust than Gaussian distribution. Expectation maximization (EM) algorithm is adopted to estimate model parameters and to maximize the lower bound on the data log-likelihood from observations. Numerical experiments on various data types are conducted. The performance of the proposed model is compared to other mixture models, demonstrating the robustness, accuracy and effectiveness of our method.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
ICIP2
2012 Illumination suppression for illumination invariant face recognition
Aryaz Baradarani, Q. M. Jonathan Wu, Majid Ahmadi
ICPR2
2012 Bayesian feature selection and model detection for student's t-mixture distributions
Hui Zhang 0015, Q. M. Jonathan Wu, Thanh Minh Nguyen 0001
ICPR2
2012 Image and Video Encryption based on Dual Space-Filling Curves
abstract
In this paper, a new encryption scheme with three different modes of operations is proposed based on dual space-filling curves (SFSs) and a fractional wavelet transform (FrWT). This scheme is initially proposed for images and then extended to videos. The core idea behind the proposed schemes is to decompose an image/video first by the FrWT followed by the shuffling of each sub-band coefficients by means of a dual SFC. At last, an inverse FrWT is performed to get the encrypted image/video. A reliable decryption process is also proposed to construct the original image from the encrypted image. The experimental results, security and comparative analysis demonstrate the efficiency and robustness of the proposed scheme. Further, this paper also proposes an efficient implementation of an FrWT based on chaotic maps.
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Comput. J.2
2012 A new robust adjustable logo watermarking scheme
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Comput. Secur.2
2012 Fractional dual tree complex wavelet transform and its application to biometric security during communication and transmission
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
Future Gener. Comput. Syst.2
2012 Tunable halfband-pair wavelet filter banks and application to multifocus image fusion
Aryaz Baradarani, Q. M. Jonathan Wu, Majid Ahmadi, Pankajkumar Mendapara
Pattern Recognit.2
2012 A two-dimensional Neighborhood Preserving Projection for appearance-based face recognition
Haijun Zhang 0002, Q. M. Jonathan Wu, Tommy W. S. Chow, Ming-Bo Zhao
Pattern Recognit.2
2012 An Efficient Algorithm for Focus Measure Computation in Constant Time
abstract
This letter presents an efficient algorithm for focus measure computation, in constant time, to estimate depth map using image sequences acquired at varying focus. Two major factors that complicate focus measure computation include neighborhood support and gradient detection for oriented intensity variations. We present a distinct focus measure based on steerable filters that is invariant to neighborhood size and accomplishes fast depth map estimation at a considerably faster speed compared to other well-documented methods. Steerable filters represent architecture to synthesize filters of arbitrary orientation using a linear combination of basis filters. Such synthesis is helpful to analytically determine the filter output as a function of orientation. Steerable filters remove inherent limitations of traditional gradient detection techniques which perform inadequately for oriented intensity variations and low textured regions.
Rashid Minhas, Abdul Adeel Mohammed, Q. M. Jonathan Wu
IEEE Trans. Circuits Syst. Video Technol.3
2012 Incremental Learning in Human Action Recognition Based on Snippets
abstract
In this paper, we present a systematic framework for recognizing human actions without relying on impractical assumptions, such as processing of an entire video or requiring a large look-ahead of frames to label an incoming video. As a secondary goal, we examine incremental learning as an overlooked obstruction to the implementation of reliable real-time recognition. Assuming weak appearance constancy, the shape of an actor is approximated by adaptively changing intensity histograms to extract pyramid histograms of oriented gradient features. As action progresses, the shape update is carried out by adjustment of a few blocks within a tracking window to closely track evolving contours. The nonlinear dynamics of an action are learned using a recursive analytic approach, which transforms training into a simple linear representation. Such a learning strategy has two advantages: 1) minimized error rates, and significant savings in computational time; and 2) elimination of the widely accepted limitations of batch-mode training for action recognition. The effectiveness of our proposed framework is corroborated by experimental validation against the state of the art.
Rashid Minhas, Abdul Adeel Mohammed, Q. M. Jonathan Wu
IEEE Trans. Circuits Syst. Video Technol.3
2012 Structure and Motion Recovery Based on Spatial-and-Temporal-Weighted Factorization
abstract
This paper focuses on the problem of structure and motion recovery from uncalibrated image sequences. It has been empirically proven that image measurement uncertainties can be modeled spatially and temporally by virtue of reprojection residuals. Consequently, a spatial-and-temporal-weighted factorization (STWF) algorithm is proposed to handle significant noise contained in the tracking data. This paper presents three novelties and contributions. First, the image reprojection residual of a feature point is demonstrated to be generally proportional to the error magnitude associated with the image point. Second, the error distributions are estimated from a different perspective, that of the reprojection residuals. The image errors are modeled both spatially and temporally to cope with different kinds of uncertainties. Previous studies have considered only the spatial information. Third, based on the estimated error distributions, an STWF algorithm is proposed to improve the overall accuracy and robustness of traditional approaches. Unlike existing approaches, the proposed technique does not require prior information of image measurement and is easy to implement. Extensive experiments on synthetic data and real images validate the proposed method.
Guanghui Wang 0001, John S. Zelek, Q. M. Jonathan Wu
IEEE Trans. Circuits Syst. Video Technol.3
2012 Tracking and Pairing Vehicle Headlight in Night Scenes
abstract
Traffic surveillance is an important topic in computer vision and intelligent transportation systems and has intensively been studied in the past decades. However, most of the state-of-the-art methods concentrate on daytime traffic monitoring. In this paper, we propose a nighttime traffic surveillance system, which consists of headlight detection, headlight tracking and pairing, and camera calibration and vehicle speed estimation. First, a vehicle headlight is detected using a reflection intensity map and a reflection suppressed map based on the analysis of the light attenuation model. Second, the headlight is tracked and paired by utilizing a simple yet effective bidirectional reasoning algorithm. Finally, the trajectories of the vehicle's headlight are employed to calibrate the surveillance camera and estimate the vehicle's speed. Experimental results on typical sequences show that the proposed method can robustly detect, track, and pair the vehicle headlight in night scenes. Extensive quantitative evaluations and related comparisons demonstrate that the proposed method outperforms state-of-the-art methods.
Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001, Xinge You
IEEE Trans. Intell. Transp. Syst.2
2012 Robust Student's-t Mixture Model With Spatial Constraints and Its Application in Medical Image Segmentation
abstract
Finite mixture model based on the Student's-t distribution, which is heavily tailed and more robust than Gaussian, has recently received great attention for image segmentation. A new finite Student's-t mixture model (SMM) is proposed in this paper. Existing models do not explicitly incorporate the spatial relationships between pixels. First, our model exploits Dirichlet distribution and Dirichlet law to incorporate the local spatial constrains in an image. Secondly, we directly deal with the Student's-t distribution in order to estimate the model parameters, whereas, the Student's-t distributions in previous models are represented as an infinite mixture of scaled Gaussians that lead to an increase in complexity. Finally, instead of using expectation maximization (EM) algorithm, the proposed method adopts the gradient method to minimize the higher bound on the data negative log-likelihood and to optimize the parameters. The proposed model is successfully compared to the state-of-the-art finite mixture models. Numerical experiments are presented where the proposed model is tested on various simulated and real medical images.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Medical Imaging2
2012 A New Fractional Random Wavelet Transform for Fingerprint Security
abstract
In this correspondence paper, the wavelet transform, which is an important tool in signal and image processing, has been generalized by coalescing wavelet transform and fractional random transform. The new transform, i.e., fractional random wavelet transform (FrRnWT) inherits the excellent mathematical properties of wavelet transform and fractional random transform. Possible applications of the proposed transform are in biometrics, image compression, image transmission, transient signal processing, etc. In this correspondence paper, biometrics is chosen as the primary application; and hence, a new technique is proposed for securing fingerprints during communication and transmission over insecure channel.
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
IEEE Trans. Syst. Man Cybern. Part A2
2012 Gaussian-Mixture-Model-Based Spatial Neighborhood Relationships for Pixel Labeling Problem
abstract
In this paper, we present a new algorithm for pixel labeling and image segmentation based on the standard Gaussian mixture model (GMM). Unlike the standard GMM where pixels themselves are considered independent of each other and the spatial relationship between neighboring pixels is not taken into account, the proposed method incorporates this spatial relationship into the standard GMM. Moreover, the proposed model requires fewer parameters compared with the models based on Markov random fields. In order to estimate model parameters from observations, instead of utilizing an expectation-maximization algorithm, we employ gradient method to minimize a higher bound on the data negative log-likelihood. The performance of the proposed model is compared with methods based on both standard GMM and Markov random fields, demonstrating the robustness, accuracy, and effectiveness of our method.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
IEEE Trans. Syst. Man Cybern. Part B2
2011 Human face classification based on localized blur descriptors
abstract
In our proposed work localized patch based geometric blur point descriptors are accumulated to generate a global similarity matrix for every pair of focal and query image. A focal image is a randomly selected template image that represents each class and a query image symbolizes all other images that belong to the same class. The similarity matrix is dimensionally reduced using the proposed bidirectional 2-dimensional principal component analysis technique to generate distinctive feature sets. These feature sets are used for training and testing an extreme learning machine classifier. The proposed face recognition structure handles variations in head positions, lighting conditions, facial expressions and cluttered background by exclusively matching template and query images. Extensive experiments are performed using challenging face databases and significant improvements in recognition accuracy were achieved.
Abdul Adeel Mohammed, Q. M. Jonathan Wu, Maher A. Sid-Ahmed
ICIP2
2011 A novel rate-distortion optimization method of H.264/AVC intra coder
abstract
In H.264/AVC, the concept of rate-distortion optimization (RDO) mode decision has proven to be an important coding tool. But the complexity and computation load of RDO technique is extremely high. In this paper, we propose an enhanced low complexity cost function for H.264/AVC intra 4x4 mode selections. The enhanced cost function uses sum of absolute Hadamard-transformed differences (SATD) and variance of the residual block to estimate distortion part of the cost function. A threshold based large coefficients count is also used for estimating the bit-rate part. The proposed method improves the rate-distortion (RD) performance of the conventional fast cost functions while maintaining low complexity requirement.
Mohammed Golam Sarwer, Q. M. Jonathan Wu, Xiao-Ping Zhang 0002
ICIP2
2011 Pattern recognition by affine Legendre moment invariants
abstract
Affine moment invariants are important shape descriptors in pattern recognition and computer vision. Existing affine invariants methods are based on geometric and complex moments. In this paper, we propose a set of affine invariants extracted from Legendre moments. These invariants are derived by the relationship between the Legendre moment of the affine transformed image and that of the original image. The performance of the proposed descriptor is evaluated with a set of binary and gray images. Experimental results show that the proposed method behaves better than existing methods in terms of pattern recognition accuracy.
Hui Zhang 0015, Q. M. Jonathan Wu
ICIP2
2011 Dirichlet Gaussian mixture model: Application to image segmentation
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
Image Vis. Comput.2
2011 Shape from focus using fast discrete curvelet transform
Rashid Minhas, Abdul Adeel Mohammed, Q. M. Jonathan Wu
Pattern Recognit.3
2011 Human face recognition based on multidimensional PCA and extreme learning machine
Abdul Adeel Mohammed, Rashid Minhas, Q. M. Jonathan Wu, Maher A. Sid-Ahmed
Pattern Recognit.3
2011 Affine Legendre Moment Invariants for Image Watermarking Robust to Geometric Distortions
abstract
Geometric distortions are generally simple and effective attacks for many watermarking methods. They can make detection and extraction of the embedded watermark difficult or even impossible by destroying the synchronization between the watermark reader and the embedded watermark. In this paper, we propose a new watermarking approach which allows watermark detection and extraction under affine transformation attacks. The novelty of our approach stands on a set of affine invariants we derived from Legendre moments. Watermark embedding and detection are directly performed on this set of invariants. We also show how these moments can be exploited for estimating the geometric distortion parameters in order to permit watermark extraction. Experimental results show that the proposed watermarking scheme is robust to a wide range of attacks: geometric distortion, filtering, compression, and additive noise.
Hui Zhang 0015, Huazhong Shu, Gouenou Coatrieux, Q. M. Jonathan Wu, Hongqing Zhu, Limin Luo 0001
IEEE Trans. Image Process.5
2011 Low Power Chien Search for BCH Decoder Using RT-Level Power Management
abstract
As a major contributor to the Bose-Chaudhuri-Hocquenghem (BCH) decoder's power consumption, Chien search is a critical step in the binary BCH decoding process for many portable applications. This paper proposes a new low-power design strategy by applying register transfer level (RTL) power management with significant power savings. The proposed Chien search is implemented for a (255, 187, 9) code in CMOS 0.18-μm technology, and simulations show 34% power improvement over the conventional method.
Shu-Yi Wong, Chunhong Chen, Q. M. Jonathan Wu
IEEE Trans. Very Large Scale Integr. Syst.3
2010 Human action recognition based on Self Organizing Map
abstract
This paper proposes a novel neural network approach for human action recognition based on Self Organizing Map (SOM). The SOM acts as a tool to cluster feature data and to reduce data dimensionality. The key poses in action sequences are extracted by the trained SOM. After the mapping of SOM, a human action sequence is represented as a trajectory of map units. For action recognition, a longest common subsequence algorithm is utilized to match action trajectories on the map robustly. The experiments are carried out on a well known human action dataset, viz.: the Weizmann dataset. We obtain promising results which show the potential of this SOM based action recognition method.
Wei Huang 0029, Q. M. Jonathan Wu
ICASSP2
2010 Stereo matching algorithm based on curvelet decomposition and modified support weights
abstract
We present a novel multiresolution analysis based stereo matching method using curvelets and modified adaptive support weight. Multiresolution analysis has long been applied to stereo correspondence. However, previous methods suffer from false matches arising from textureless region or repetitive textures and fattening effect due to area based matching. In the proposed approach, we have reduced false matches by using curvelet coefficients in different scales and orientations. Curvelet coefficients can uniquely represent different image points and increase matching accuracy. The fattening effect is reduced using support weights modified for curvelets. The proposed method is verified and compared with state-of-the art methods by extensive tests, and good results are obtained.
Dibyendu Mukherjee, Guanghui Wang 0001, Q. M. Jonathan Wu
ICASSP3
2010 Facial expression recognition using curvelet based local binary patterns
abstract
This paper proposes the use of the combination of digital curvelet transform and local binary patterns for recognizing facial expressions from still images. The curvelet transform is applied to the image of a face at a specific scale and orientation. Local binary patterns are extracted from the selected curvelet sub-bands to form the descriptive feature set of the expressions. The average of the features of a particular class of expression is considered as the representative feature set of that class. The expression recognition is performed using a nearest neighbor classifier with Chi-square as the dissimilarity metric. Experiments show that our method yields recognition rates of 93% and 90% in JAFFE and Cohn-Kanade databases respectively.
Ashirbani Saha, Q. M. Jonathan Wu
ICASSP2
2010 Real Time Human Visual System Based Framework for Image Fusion
Gaurav Bhatnagar, Q. M. Jonathan Wu, Balasubramanian Raman
ICISP2
2010 Performance Evaluation of Multiresolution Methods in Disparity Estimation
Dibyendu Mukherjee, Gaurav Bhatnagar, Q. M. Jonathan Wu
ICISP3
2010 An efficient depth map estimation technique using complex wavelets
abstract
A new focus measure system is proposed based on complex wavelet transform and quadrature pair of steerable filters. In shape from focus (SFF), noise, illumination variation and oriented features degrade the performance of focus measure operator. This paper introduces the use of complex wavelets due to shift-invariance and directionality of the transformation suitable for detecting various types of features which plays a pivotal role in depth estimation of a scene. A quadrature pair of steerable filters is employed to measure focus by calculating the local oriented energy of the detected features. Experimental examples are provided to illustrate the effectiveness of the approach and the results compare favorably to well-documented methods in literature.
Pankajkumar Mendapara, Aryaz Baradarani, Q. M. Jonathan Wu
ICME3
2010 On the Design of a Class of Odd-Length Biorthogonal Wavelet Filter Banks for Signal and Image Processing
abstract
In this paper, we introduce an approach to the design of odd-length biorthogonal wavelet filter banks based on semidefinite programming employing Bernstein polynomials. The method is systematic and renders a simple optimization problem, yet it offers wavelet filters ranging from maximally flat to maximal passband/stopband width. The odd-length biorthogonal filter pairs are then used in multi-focus imaging to obtain a fully-focused image from a set of registered semi-focused input images at varying focus employing the distance transform and exponentially decaying function on the subbands in wavelet domain. Various images are tested and experimental results compare favorably to recent results in literature.
Aryaz Baradarani, Pankajkumar Mendapara, Q. M. Jonathan Wu
ICPR3
2010 Quasi-perspective Projection Model: Theory and Application to Structure and Motion Factorization from Uncalibrated Image Sequences
Guanghui Wang 0001, Q. M. Jonathan Wu
Int. J. Comput. Vis.2
2010 Human action recognition using extreme learning machine based on visual vocabularies
Rashid Minhas, Aryaz Baradarani, Sepideh Seifzadeh, Q. M. Jonathan Wu
Neurocomputing4
2010 A fast recognition framework based on extreme learning machine using hybrid object information
Rashid Minhas, Abdul Adeel Mohammed, Q. M. Jonathan Wu
Neurocomputing3
2010 Image matching using enclosed region detector
Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001, Xinge You
J. Vis. Commun. Image Represent.2
2010 The quasi-perspective model: Geometric properties and 3D reconstruction
Guanghui Wang 0001, Q. M. Jonathan Wu
Pattern Recognit.2
2010 An Adaptive Computational Model for Salient Object Detection
abstract
Salient object detection is a basic technique for many computer vision applications. In this paper, we propose an adaptive computational model to detect the salient object in color images. Firstly, three human observation behaviors and scalable subtractive clustering techniques are used to construct attention Gaussian mixture model (AGMM) and background Gaussian mixture model (BGMM). Secondly, the Bayesian framework is employed to classify each pixel into salient object or background object. Thirdly, expectation-maximization (EM) algorithm is utilized to update the parameters of AGMM, BGMM, and Bayesian framework based on the detection results. Finally, the classification and update procedures are repeated until the detection results evolve to a steady state. Experiments on a variety of images demonstrate the robustness of the proposed method. Extensive quantitative evaluations and comparisons demonstrate that the proposed method significantly outperforms state-of-the-art methods.
Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001, Hai Bing Yin
IEEE Trans. Multim.2
2010 An extension of the standard mixture model for image segmentation
abstract
Standard gaussian mixture modeling (GMM) is a well-known method for image segmentation. However, the pixels themselves are considered independent of each other, making the segmentation result sensitive to noise. To reduce the sensitivity of the segmented result with respect to noise, Markov random field (MRF) models provide a powerful way to account for spatial dependences between image pixels. However, their main drawback is that they are computationally expensive to implement, and require large numbers of parameters. Based on these considerations, we propose an extension of the standard GMM for image segmentation, which utilizes a novel approach to incorporate the spatial relationships between neighboring pixels into the standard GMM. The proposed model is easy to implement and compared with MRF models, requires lesser number of parameters. We also propose a new method to estimate the model parameters in order to minimize the higher bound on the data negative log-likelihood, based on the gradient method. Experimental results obtained on noisy synthetic and real world grayscale images demonstrate the robustness, accuracy and effectiveness of the proposed model in image segmentation, as compared to other methods based on standard GMM and MRF models.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu, Siddhant Ahuja
IEEE Trans. Neural Networks2
2009 Human Action Recognition Using Non-separable Oriented 3D Dual-Tree Complex Wavelets
Rashid Minhas, Aryaz Baradarani, Sepideh Seifzadeh, Q. M. Jonathan Wu
ACCV (3)4
2009 Two-View Geometry and Reconstruction under Quasi-perspective Projection
Guanghui Wang 0001, Q. M. Jonathan Wu
ACCV (2)2
2009 Vehicle Headlights Detection Using Markov Random Fields
Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001
ACCV (1)2
2009 A generic fingerprint image compression technique based on wave atoms decomposition
abstract
Modern fingerprint image compression and reconstruction standards used by the US Federal Bureau of Investigation (FBI) are based upon the popular 9/7 discrete wavelet transform. Multiresolution analysis tools have been successfully applied for fingerprint image compression for more than a decade; we propose a novel fingerprint image compression technique based on recently proposed wave atoms decomposition. Wave atoms decomposition has specifically been designed for enhanced representation of oscillatory patterns to convey temporal and spatial information. Our proposed compression scheme is based upon linear vector quantization of decomposed wave atoms representation of fingerprint images. Later quantized information is encoded with arithmetic entropy scheme. The proposed image compression standard outperforms the FBI fingerprint image compression standard, the wavelet scalar quantization (WSQ). Data mining, law enforcement, border security, and forensic applications can potentially benefit from our proposed compression scheme.
Abdul Adeel Mohammed, Rashid Minhas, Q. M. Jonathan Wu, Maher A. Sid-Ahmed
ICIP3
2009 Fast block motion estimation by edge based partial distortion search
abstract
In order to reduce the computations of motion estimation, this paper proposes a novel edge based partial distortion search (EPDS) algorithm in which the entire macroblock (MB) is divided into different sub-blocks and the calculation order of partial distortion is determined based on the edge strength of sub-blocks. In the proposed method, only selected numbers of search points are considered for candidate motion vectors. An efficient early termination method, which is based on the dynamic threshold, is also proposed to decide whether a search point has met the distortion criterion. Simulation results show that the proposed method is 108 times faster than Full Search (FS), 10 times faster than Normalized Partial Distortion search (NPDS) and 2 times faster than the dual halfway stop NPDS (DHS-NPDS) on an average. PSNR degradation of the proposed algorithm is negligible and in the region of 0.01 dB.
Mohammed Golam Sarwer, Q. M. Jonathan Wu
ICIP2
2009 Power-management-based Chien search for low power BCH decoder
abstract
Chien search is a computation-intensive VLSI design in BCH (Bose-Chaudhuri-Hocquenghem) decoders for a variety of applications such as digital video and data storage. Existing low power approaches to this subject are effective only for single-bit error case, and their power efficiency decreases dramatically for multiple-bit-correcting code. This paper presents a novel approach of using register-transfer-level (RTL) power management in the search process, leading to significant power savings for BCH codes with higher correction capability. A (255, 187, 9) BCH code is implemented in 0.18μm CMOS technology as an example.
Shu-Yi Wong, Chunhong Chen, Q. M. Jonathan Wu
ISLPED3
2009 Adaptive Search Area Selection of Variable Block-Size Motion Estimation of H.264/AVC Video Coding Standard
abstract
The variable block size motion estimation is the most time consuming part of H.264/AVC encoder. In order to reduce the computation of motion estimation module, this paper presents a novel adaptive search area decision method by utilizing the information of the previously computed motion vectors (MVs). The direction of picture movement of the previously computed blocks is also considered for search area selection. Experimental results show that the proposed algorithm provides significant improvement of coding speed with negligible objective quality degradation compared to the full search motion estimation method adopted by H.264/AVC reference software.
Mohammed Golam Sarwer, Q. M. Jonathan Wu
ISM2
2009 Depth Map Estimation Using Exponentially Decaying Focus Measure Based on Susan Operator
abstract
This paper presents a novel technique for depth map estimation using a sequence of images acquired at varying focus. In depth map estimation noise, illumination variations and types of extracted features significantly affect the performance of a focus measure. This paper proposes the use of SUSAN operator, to extract features, because of its structure preserving noise filtering which plays a pivotal role in depth estimation of a scene. We introduce a new focus measure based on exponentially decaying function to use neighborhood information of an extracted feature point that assigns more weight to the closer pixel points. Experiments validate superior performance of our proposed algorithm in comparison to other well-documented methods.
Pankajkumar Mendapara, Rashid Minhas, Q. M. Jonathan Wu
SMC3
2009 Application of bidirectional two-dimensional principal component analysis to curvelet feature based face recognition
abstract
A bidirectional two-dimensional principal component analysis (2DPCA) is proposed for human face recognition using curvelet feature subspace. Traditionally multiresolution analysis tools namely wavelets and curvelets have been used in the past for extracting and analyzing still images for recognition and classification tasks. Curvelet transform has gained significant popularity over wavelet based techniques due to its improved directional and edge representation capability. In the past features extracted from curvelet subbands were dimensionally reduced using linear principal component analysis (PCA) for obtaining a representative feature set. The novelty of the proposed method lies in the application of 2DPCA to curvelet feature subspace by computing image covariance matrices of square training sample matrices in their original form and transposed form respectively to generate a more meaningful and enhanced feature vectors. Extensive experiments were performed using the proposed bidirectional 2DPCA based face recognition algorithm and superior performance is obtained in comparison with state of the art techniques.
Abdul Adeel Mohammed, Q. M. Jonathan Wu, M. A. Ahmed
SMC2
2009 A Real-Time Ellipse Detection Based on Edge Grouping
abstract
In this paper, we present a efficient algorithm for real-time ellipse detection. Unlike Hough transform algorithm that is computationally intense and requires a higher dimensional parameter space, our proposed method reduces the computational complexity significantly, and accurately detects ellipses in realtime. We present a new method of detecting arc-segments from the image, based on the properties of ellipse. We then group the arc-segments into elliptical arcs in order to estimate the parameters of the ellipse, which are calculated using the least-square method. Our method has been tested and implemented on synthetic and real-world images containing both complete and incomplete ellipses. The performance is compared to existing ellipse detection algorithms, demonstrating the robustness, accuracy and effectiveness of our algorithm.
Thanh Minh Nguyen 0001, Siddhant Ahuja, Q. M. Jonathan Wu
SMC3
2009 Human action recognition using Recursive Self Organizing map and longest common subsequence matching
abstract
A little attention has been given to the use of recursive self organizing map (SOM) for human action recognition in the past years. This paper introduces an action recognition framework using the recursive SOM, a temporal extension of SOM that learns adapted representations of temporal context associated with a time series. We demonstrate the effectiveness of recursive SOM for data clustering, dimensionality reduction and context learning in human action recognition. The atomic poses in motion sequences and their contextual information are extracted and encoded by the trained recursive SOM. A human action sequence is represented as a trajectory of map units. To classify a new action, a longest common subsequence algorithm using dynamic programming is employed to robustly match action trajectories on the map. To the best of our knowledge, we are the first to try recursive SOM approach for human action recognition. We test the approach on a well known benchmark action dataset and achieve promising results.
Wei Huang 0029, Q. M. Jonathan Wu
WACV2
2009 What can we learn about the scene structure from three orthogonal vanishing points in images
Guanghui Wang 0001, Hung-Tat Tsui, Q. M. Jonathan Wu
Pattern Recognit. Lett.3
2009 Curvelet based face recognition via dimension reduction
Tanaya Mandal, Q. M. Jonathan Wu, Yuan Yuan 0001
Signal Process.2
2009 Adaptive Variable Block-Size Early Motion Estimation Termination Algorithm for H.264/AVC Video Coding Standard
abstract
The variable block-size motion estimation (ME) process is the H.264/AVC encoder's most time-consuming function. This letter proposes to reduce the complexity of the ME process with an early termination algorithm that features an adaptive threshold based on the statistical characteristics of rate-distortion (RD) cost regarding current block and previously processed blocks and modes. In this method, most motion searches can be stopped early, with a large number of search points saved. A region-based search is also suggested to further reduce the computation required for full search ME. A search point reduction scheme for the fast motion estimation of H.264/AVC is also introduced, and the experimental results illustrate how the proposed method reduces ME times for full search and fast motion estimation by about 77 and 31%, respectively, despite the insignificant degradation of RD performance.
Mohammed Golam Sarwer, Q. M. Jonathan Wu
IEEE Trans. Circuits Syst. Video Technol.2
2009 Perspective 3-D Euclidean Reconstruction With Varying Camera Parameters
abstract
The paper addresses the problem of 3-D Euclidean structure and motion recovery from video sequences based on perspective factorization. It is well known that projective depth recovery and camera calibration are two essential and difficult steps in metric reconstruction. We focus on the difficulties and propose two new algorithms to improve the performance of perspective factorization. First, we propose to initialize the projective depths via a projective structure reconstructed from two views with large camera movement, and optimize the depths iteratively by minimizing reprojection residues. The algorithm is more accurate than previous methods and converges quickly. Second, we propose a self-calibration method based on the Kruppa constraint to deal with more general camera model. The Euclidean structure can be recovered from factorization of the normalized tracking matrix. Extensive experiments on synthetic data and real sequences are performed to validate the proposed method and good improvements are observed.
Guanghui Wang 0001, Q. M. Jonathan Wu
IEEE Trans. Circuits Syst. Video Technol.2
2008 Quasi-perspective projection with applications to 3D factorization from uncalibrated image sequences
abstract
The paper addresses the problem of factorization-based 3D reconstruction from uncalibrated image sequences. We propose a quasi-perspective projection model and apply the model to structure and motion recovery of rigid and nonrigid objects based on factorization of tracking matrix. The novelty and contribution of the paper lies in three aspects. First, under the assumption that the camera is far away from the object with small rotations, we propose and prove that the imaging process can be modeled by quasi-perspective projection. The model is more accurate than affine since the projective depths are implicitly embedded. Second, we apply the model to the factorization algorithm and establish the framework of rigid and nonrigid factorization under quasi-perspective assumption. Third, we propose a new and robust method to recover the transformation matrix that upgrades the factorization to the Euclidean space. The proposed method is validated and evaluated on synthetic and real image sequences and good improvements over existing solutions are observed.
Guanghui Wang 0001, Q. M. Jonathan Wu
CVPR2
2008 Maximum likelihood neural network based on the correlation among neighboring pixels for noisy image segmentation
abstract
In this paper, we will present a new algorithm which is extended from the standard Gaussian mixture model to segment the noisy image based on the correlation among neighboring pixels. Firstly, we use the correlation between each centre pixel and its neighboring pixels in 3 times 3 window in building the prior probability, and this centre pixel is used to construct the conditional density function. Finally, to estimate the posterior probabilities of each pixel, instead of using expectation maximization algorithm as usual, we present a new maximum likelihood neural network (MaxNet) to optimize the parameters by using the error back propagation. Extensive experimental results illustrate the better performance compared to mixture model based on Markov random fields.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
ICIP2
2008 A Combination of Positive and Negative Fuzzy Rules for Image Classification Problem
abstract
In this paper, we propose a new fuzzy rule-based system for application in image classification problem. Each rule in our proposed system can represent more than one class. While traditional fuzzy systems consider positive fuzzy rules only, in this paper, we focus on combining negative fuzzy rules with traditional positive ones leading to fuzzy inference systems. This new approach has been tested on image classification problem consisting of multiple images with excellent results.
Thanh Minh Nguyen 0001, Q. M. Jonathan Wu
ICMLA2
2008 Face recognition using curvelet based PCA
abstract
This paper identifies a novel feature space to address the problem of human face recognition from still images. This is based on the PCA space of the features extracted by a new multiresolution analysis tool called Fast Discrete Curvelet Transform. Curvelet Transform has better directional and edge representation abilities than widely used wavelet transform. Inspired by these attractive attributes of curvelets, we introduce the idea of decomposing images into its curvelet subbands and applying PCA (Principal Component Analysis) on the selected subbands in order to create a representative feature set. Experiments have been designed for both single and multiple training images per subject. A comparative study with wavelet-based and traditional PCA techniques is also presented. High accuracy rate achieved by the proposed method for two well-known databases indicates the potential of this curvelet based feature extraction method.
Tanaya Mandal, Q. M. Jonathan Wu
ICPR2
2008 Structure and motion factorization under quasi-perspective projection with missing data in tracking matrix
abstract
The paper is focused on the problem of structure and motion factorization from uncalibrated image sequences. Based on our early study on quasi-perspective projection, we give an analysis on the imaging errors of different projection models and propose to adopt power factorization algorithm to deal with missing data problem. The main contribution lies in two aspects. First, we carry out an error analysis of the quasi-perspective projection and prove that it is more accurate than affine model under small camera movements. Second, we propose to utilize power factorization to factorize the tracking matrix. Compared with SVD-based method, the algorithm can work with incomplete tracking data, and it is computationally cheaper than other methods. The proposed method is evaluated on synthetic and real image sequences and better results are observed.
Guanghui Wang 0001, Q. M. Jonathan Wu, Wei Huang 0029
ICPR2
2008 Adaptive semantic Bayesian framework for image attention
abstract
Image attention is the basic technique for many computer vision applications. In this paper, we propose an adaptive Bayesian framework to detect the image attention in color image. Firstly, three simple semantics and subtractive clustering are used to construct attention Gaussians mixture model (AGMM) and background Gaussians mixture model (BGMM). Secondly, the Bayesian framework is utilized to classify each pixel into attention objects and background objects. Thirdly, EM algorithm is used to update the parameters of AGMM, BGMM, and Bayesian framework according to the detection results. Finally, the above classification and update procedures are repeated until the detection results become steady. Experimental results on typical images exhibit the robustness of the proposed method.
Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001
ICPR2
2008 Real-Time Approach for Adaptive Object Segmentation in Time-of-Flight Sensors
abstract
In this paper, a high-speed, adaptive depth segmentation method is proposed, which results in superior performance over current employed segmentation algorithms when applied in real-time tracking applications. Existing segmentation methods are difficult to implement in real-time due to their slow performance, whereby enhancing their run-time they are better applicable in real-time approaches, i.e., tracking. The proposed method leverages the depth distributions of range images for segmentation of objects of interest, without having a priori knowledge about the scene. This approach has been tested with real data in unconstrained environments, under varying conditions. The experimental results demonstrate the speed efficiency, as well as robustness of the proposed technique.
Ehsan Parvizi, Q. M. Jonathan Wu
ICTAI (1)2
2008 Protocol-level performance analysis for anti-collision protocols in RFID systems
abstract
This paper provides an analytical approach to evaluate the performance of anti-collision protocols in radio-frequency identification (RFID) systems. The analysis is given at protocol-level using such performance metrics as the number of state transitions and clock cycles involved in the protocols’ state diagrams, their power and energy dissipation. Discussion and comparison are based on worst-case scenario that represents the process of identifying the last tag when multiple tags are simultaneously available in the system. A new protocol is also proposed for performance improvement.
Mohammed Berhea, Chunhong Chen, Q. M. Jonathan Wu
ISCAS3
2008 A small scale fingerprint matching scheme using Digital Curvelet Transform
abstract
This paper proposes a new method for fingerprint matching based on the features extracted using a new multiresolution analysis tool called digital curvelet transform. The curvelet coefficients extracted from enhanced fingerprint images act as the feature-set for a k-nearest neighbor classifier. The performance of this scheme has been evaluated on a small database of 120 images. A comparative study between the wavelet-based and the curvelet-based techniques has also been included. The high recognition rate achieved by using this method suggests an efficient solution for a small scale fingerprint recognition system. Note that, this paper is the first documented attempt to explore the possibilities of a new multiresolution analysis tool called curvelet transform to address the problem of fingerprint identification.
Tanaya Mandal, Q. M. Jonathan Wu
SMC2
2008 Rotation constrained power factorization for structure from motion of nonrigid objects
Guanghui Wang 0001, Hung-Tat Tsui, Q. M. Jonathan Wu
Pattern Recognit. Lett.3
2008 Single view based pose estimation from circle or parallel lines
Guanghui Wang 0001, Q. M. Jonathan Wu, Zhengqiao Ji
Pattern Recognit. Lett.2
2008 Kruppa equation based camera calibration from homography induced by remote plane
Guanghui Wang 0001, Q. M. Jonathan Wu, Wei Zhang 0025
Pattern Recognit. Lett.2
2008 Fast sum of absolute transformed difference based 4×4 intra-mode decision of H.264/AVC video coding standard
Mohammed Golam Sarwer, Lai-Man Po, Q. M. Jonathan Wu
Signal Process. Image Commun.3
2008 Multilevel Framework to Detect and Handle Vehicle Occlusion
abstract
This paper presents a multilevel framework to detect and handle vehicle occlusion. The proposed framework consists of the intraframe, interframe, and tracking levels. On the intraframe level, occlusion is detected by evaluating thecompactness ratioandinterior distance ratioof vehicles, and the detected occlusion is handled by removing a “cutting region” of the occluded vehicles. On the interframe level, occlusion is detected by performing subtractive clustering on the motion vectors of vehicles, and the occluded vehicles are separated according to the binary classification of motion vectors. On the tracking level, occlusion layer images are adaptively constructed and maintained, and the detected vehicles are tracked in both the captured images and the occlusion layer images by performing a bidirectional occlusion reasoning algorithm. The proposed intraframe, interframe, and tracking levels are sequentially implemented in our framework. Experiments on various typical scenes exhibit the effectiveness of the proposed framework. Quantitative evaluation and comparison demonstrate that the proposed method outperforms state-of-the-art methods.
Wei Zhang 0025, Q. M. Jonathan Wu, Xiaokang Yang 0001, Xiangzhong Fang
IEEE Trans. Intell. Transp. Syst.2
2008 Stratification Approach for 3-D Euclidean Reconstruction of Nonrigid Objects From Uncalibrated Image Sequences
abstract
This paper addresses the problem of 3-D reconstruction of nonrigid objects from uncalibrated image sequences. Under the assumption of affine camera and that the nonrigid object is composed of a rigid part and a deformation part, we propose a stratification approach to recover the structure of nonrigid objects by first reconstructing the structure in affine space and then upgrading it to the Euclidean space. The novelty and main features of the method lies in several aspects. First, we propose a deformation weight constraint to the problem and prove the invariability between the recovered structure and shape bases under this constraint. The constraint was not observed by previous studies. Second, we propose a constrained power factorization algorithm to recover the deformation structure in affine space. The algorithm overcomes some limitations of a previous singular-value-decomposition-based method. It can even work with missing data in the tracking matrix. Third, we propose to separate the rigid features from the deformation ones in 3-D affine space, which makes the detection more accurate and robust. The stratification matrix is estimated from the rigid features, which may relax the influence of large tracking errors in the deformation part. Extensive experiments on synthetic data and real sequences validate the proposed method and show improvements over existing solutions.
Guanghui Wang 0001, Q. M. Jonathan Wu
IEEE Trans. Syst. Man Cybern. Part B2
2007 Pose Estimation from Circle or Parallel Lines in a Single Image
Guanghui Wang 0001, Q. M. Jonathan Wu, Zhengqiao Ji
ACCV (2)2
2007 Real-Time 3D Head Tracking Based on Time-of-Flight Depth Sensor
abstract
This paper presents a novel head detection algorithm based on contour analysis on depth images. A sequence of depth-valued images is used as the system input using a 3D time-of-flight depth sensor. The background and foreground layers of the image are segmented using a straightforward depth thresholding technique. Moving regions are further processed in each frame, and contour analysis is performed on the depth maps to extract the curves of moving regions. Finally, ellipse fitting is performed to determine the objective head targets in the image. This information will be passed to the tracker in order to accomplish tracking the targets in the scene. Experimental results demonstrate the efficiency of the proposed method.
Ehsan Parvizi, Q. M. Jonathan Wu
ICTAI (1)2
2007 An improved immune Q-learning algorithm
abstract
Reinforcement learning is a framework in which an agent can learn behavior without knowledge on a task or an environment by exploration and exploitation. Striking a balance between exploration and exploitation is one of the key problems of action selection in reinforcement learning. Exploitation causes the agent to reach a locally optimal policy quickly, whereas excessive exploration degrades the performance of the algorithm, though it may improve the learning performance and escape from a locally optimal policy. Recently the human immune systems have aroused researcher's interest due to its useful mechanisms which can be exploited for information processing in a complex cognition system. In this paper, we transplant some immune mechanisms into the basic Q-learning algorithm and convert Q-learning algorithm into a search for the optimum solution in combinatorial optimization. Experiments show that the improved Q-learning converges more quickly than Q-learning or Boltzmann exploration, and easily obtains the global solution set.
Zhengqiao Ji, Q. M. Jonathan Wu, Maher A. Sid-Ahmed
SMC2
2007 Invariant feature set in convex hull for fast image registration
abstract
Abstract-In this paper, a novel feature set in images for registration is identified. Unique, geometrically invariant and easily extractable features in images called convex diagonal, convex quadrilateral are used for accurate image registration. Convex diagonals, convex quadrilaterals have attractive properties like easy extraction, geometric invariance and frequent occurrence. Coordinates, length and orientation information of corresponding convex diagonals in different images is used for initial transformation estimate. Corresponding convex hulls of scene objects are matched using Hausdorff distance as similarity measure operator. Coarse level estimate facilitates efficient, real time computation for final registration process. Initial transformation estimate based on convex diagonals, extracted from convex hull of scene objects, is refined using fine level image details to minimize errors originating from quantization and same convex hull information for different object shapes. The behavior of reference quadrilateral is robust against noise, outliers and broken edges.
Rashid Minhas, Q. M. Jonathan Wu
SMC2
2007 High-speed skin color segmentation for real-time human tracking
abstract
This paper proposes an original low-complex image classification technique for fast skin color segmentation proper for human face and limbs tracking. Experimental results demonstrate promising performance achievements compared to other existing skin color segmentation methods. Furthermore, the simplicity of this method is an attractive feature for real-time applications. The proposed method is independent of skin distribution shape and unlike the majority of techniques does not require exhaustive training.
Leila Sabeti, Q. M. Jonathan Wu
SMC2
2007 Moving Cast Shadows Detection Using Ratio Edge
abstract
Moving objects segmentation plays a very important role in real-time image analysis. However, as one of the common parts in the natural scenes, shadows severely interfere with the accuracy of moving objects detection in video surveillance. In this paper, we present a novel method for moving cast shadows detection. Based on the analysis of the physical model of moving shadows, we prove that the ratio edge is illumination invariant. The distribution of the ratio edge is discussed and a significance test is performed to classify each moving pixel into foreground object or moving shadow. Intensity constraint and geometric heuristics are imposed to further improve the performance. Experiments on various typical scenes exhibit the robustness of the proposed method. Extensively quantitative evaluation and comparison demonstrate that the proposed method significantly outperforms state-of-the-art methods.
Wei Zhang 0025, Xiangzhong Fang, Xiaokang Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Multim.4
2006 Application of artificial immune algorithms in multiple sensor system
abstract
Recently the human immune system arouse researchers interest since it has several useful mechanisms which can be used to information processing as a complex cognition system. Here it is not our concern to reproduce any immune phenomenon accurately, but to show that immune concepts can be applied to develop powerful computational tools for data processing. From this viewpoint, an improved artificial immune algorithm is presented in applying in the problems of image registration and configuration of multiple sensor system. The simulation results show that the immune algorithm can successfully obtain the global optimum with much less computational cost compared to other traditional algorithms. Hereby, we can say this new method can also be applied to other optimization problems.
Zhengqiao Ji, Q. M. Jonathan Wu
SMC2
2005 An intelligent dual mode vision guided robotic system
abstract
Industrial robotics have looked to vision systems for flexibility. This promise has largely been unrealized because existing systems are either too slow or too inaccurate. Both visual servoing and traditional look and move are insufficient because visual servoing requires too much bandwidth, and look and move requires very accurate calibration. To mitigate these effects, we have designed a hybrid system. Our hybrid system is composed of a roughly calibrated look-and-move system using a linear approximation, and a gain scheduled PD controller which performs visual servoing. The system performs markedly better than visual servoing or look-and-move techniques in isolation. This system have many potential applications including bin-picking, sorting, and tele-operation.
Kevin G. Stanley, Q. M. Jonathan Wu, William A. Gruver
SMC2
2004 A Hybrid Stereo Feature Matching Algorithm For Stereo Vision-Based Bin Picking
abstract
Stereo vision-based bin picking systems require accurate 3D information to be recovered from 2D stereo images. To achieve this goal, we have developed a hybrid coarse-to-fine algorithm for stereo feature matching, which is based on the 2D six-parameter affine transformation and local similarity evaluation. With this algorithm, the coarse matching is performed by the 2D six-parameter affine transformation to get rough feature matches, imposing a strong constraint to further search instead of the traditional epipolar constraint. To obtain precise matches, the perspective effect is dealt with fine stereo feature matching by performing local similarity evaluation on the attribute vectors of features. Experimental results proving the performance of the stereo feature matching algorithm are also presented.
Aiqiu Zuo, Jason Z. Zhang, Kevin G. Stanley, Q. M. Jonathan Wu
Int. J. Pattern Recognit. Artif. Intell.4
2004 A survey of motion-parallax-based 3-D reconstruction algorithms
abstract
The task of recovering three-dimensional (3-D) geometry from two-dimensional views of a scene is called 3-D reconstruction. It is an extremely active research area in computer vision. There is a large body of 3-D reconstruction algorithms available in the literature. These algorithms are often designed to provide different tradeoffs between speed, accuracy, and practicality. In addition, even the output of various algorithms can be quite different. For example, some algorithms only produce a sparse 3-D reconstruction while others are able to output a dense reconstruction. The selection of the appropriate 3-D reconstruction algorithm relies heavily on the intended application as well as the available resources. The goal of this paper is to review some of the commonly used motion-parallax-based 3-D reconstruction techniques and make clear the assumptions under which they are designed. To do so efficiently, we classify the reviewed reconstruction algorithms into two large categories depending on whether a prior calibration of the camera is required. Under each category, related algorithms are further grouped according to the common properties they share.
Jason Z. Zhang, Q. M. Jonathan Wu, Ze-Nian Li
IEEE Trans. Syst. Man Cybern. Part C3
2003 Active Head Tracking Based on Chromatic Shape Fitting
abstract
This paper presents a method for tracking a human head based on the integration of camera saccade and chromatic shape fitting, which are implemented as functional modules in an active tracking system. Head motion is detected in the saccade module by extracting edges from two successive images. The position of the head in the current image is approximated as the centroid of the apparition formed by the moving edges of the target. A visual position cue is used to drive a pan/tilt camera to perform real-time saccade keeping the target in the foveal area in the image. The shape-fitting module is invoked to extract more information from the target. The shape of the target is modeled as an ellipse whose position, orientation and size are dynamically determined by shape fitting, and implemented with a color registration technique. In the proposed method, quasi real-time pursuit is achieved using a Pentium II computer in an uncontrolled environment with arbitrary relative motion between the target and camera.
Jason Z. Zhang, Q. M. Jonathan Wu, William A. Gruver
Int. J. Pattern Recognit. Artif. Intell.2
2003 Correction to "Binocular transfer method for point-feature tracking of image sequences"
Jason Z. Zhang, Q. M. Jonathan Wu, William A. Gruver
IEEE Trans. Syst. Man Cybern. Part C2
2002 Binocular transfer methods for point-feature tracking of image sequences
abstract
Image transfer is a method for projecting a 3D scene from two or more reference images. Typically, the correspondences of target points to be transferred and the reference points must be known over the reference images. We present two new transfer methods that eliminate the target point correspondence requirement. We show that five reference points matched across two reference images are sufficient to linearly resolve transfer under affine projection using two views instead of three views as needed by other techniques. Furthermore, given the correspondences of any four of the five reference points in any other view, we can transfer a target point to a third view from any one of the two original reference views. To improve the robustness of the affine projection method, we incorporate an orthographic camera model. A factorization method is applied to the reference points matched over two reference views. Experiments with real image sequences demonstrate the application of both methods for motion tracking.
Jason Z. Zhang, Q. M. Jonathan Wu, Hung-Tat Tsui, William A. Gruver
IEEE Trans. Syst. Man Cybern. Part C2
2000 An intelligent vision guided telerobotic system for file manipulation and office automation
abstract
Describes a vision-guided telerobotic system that enables people with disabilities to perform clerical or office tasks. By adding a light-duty robot to the office workspace, the operator can manipulate files and perform other work-related tasks. To increase the effectiveness of the robot, vision can be used to verify that the robot is correctly positioned. In addition, vision can be be coupled with the telerobotic system to allow the user more intuitive control over the robot. Visual servoing and traditional computed kinematics actions are inappropriate for this application because visual servoing requires an excessive number of iterations and computed kinematics requires accurate calibration. To counteract these difficulties and to provide user functionality, we have designed a hybrid computed-kinematics telerobotic system with an initial coarsely-calibrated computed-kinematics step followed by a more accurate visual-servoing step. We show that there are significant performance benefits from this approach. Finally, we describe how the hybrid system may be utilized in an office environment.
Kevin G. Stanley, Q. M. Jonathan Wu, William A. Gruver
SMC2
2000 Machine vision system for curved surface inspection
Min-Fan Ricky Lee, Clarence W. de Silva, Elizabeth A. Croft, Q. M. Jonathan Wu
Mach. Vis. Appl.4
2000 Implementation of vision-based planar grasp planning
abstract
This research describes the implementation of a vision-based algorithm that is capable of rapidly determining robotic grasp points for planar objects. A representation of the target and a quadtree expansion generate candidate grasps that are compared using a cost function. The approach returns the first acceptable grasp point at a given tree resolution. The system has an execution time on the order of seconds and it is suitable for a large number of planar or near planar objects.
Kevin G. Stanley, Q. M. Jonathan Wu, William A. Gruver
IEEE Trans. Syst. Man Cybern. Part C2
1999 Content-Based Retrieval of Video Sequences Under Partial Occlusion
abstract
In this paper, we present a spatio-temporal segmentation method for content-based video retrieval. Our method is effective in the presence of partial occlusion or when new areas become exposed. In addition, periodic spatial segmentation is not performed. Rather, it is invoked as needed depending on whether or not new regions have been detected. A discussion of how this method can be used for content-based retrieval is also provided. Finally, experimental results that demonstrate the performance of the proposed spatio-temporal segmentation method are presented.
Shahram Shirani, Ali Jerbi, Faouzi Kossentini, Rabab K. Ward, Q. M. Jonathan Wu
ICIP (3)5
1999 Neural Network-Based Vision Guided Robotics
abstract
An essential problem of image based visual servoing is evaluating the inverse Jacobian which, relates changes in image features to the change in robot position. Neural networks can learn to approximate the inverse feature Jacobian. In addition, neural networks have been used in dimensionality reduction of image input. We show that it is possible to use neural networks for both feature extraction using compression and for feature Jacobian approximation in the visual servoing problem. In our system, we consider the following feature extraction methods: geometric features, averaging compression, vector quantization, and principal component extraction.
Kevin G. Stanley, Q. M. Jonathan Wu, Ali Jerbi, William A. Gruver
ICRA2
1999 A fast two dimensional image based grasp planner
abstract
This research concerns a grasp-planning algorithm that is fast and capable of determining grasp points for planar nondegenerate objects. We use a novel representation of the target and a quadtree based sampling scheme to generate a set of candidate grasps which are evaluated using a cost function. This function returns the first acceptable grasp point it finds. The resulting system has an execution time of seconds and is suitable for a large number of planar grasp planning problems.
Kevin G. Stanley, Q. M. Jonathan Wu, Ali Jerbi, William A. Gruver
IROS2
1997 Modular neural-visual servoing using a neural-fuzzy decision network
abstract
Visual servoing is a growing research area. One of the key problems of feature based visual servoing is calculating the inverse Jacobian, relating change in features to change in robot position. Neural networks can learn to approximate the inverse feature Jacobian. However, the neural network approach can only approximate the feature Jacobian for a small workspace. In order to overcome this problem, we propose using a modular approach, where several networks are trained over a small area. Furthermore, we use a neural-fuzzy counterpropagation network to decide which subspace the robot is currently occupying. The neural fuzzy network provides smoother transitions between subspaces than hard switching. Preliminary results of the system's operation are also presented.
Q. M. Jonathan Wu, Kevin G. Stanley
ICRA1
1991 Fast boundary extraction for industrial inspection
Q. M. Jonathan Wu, Michael G. Rodd
Pattern Recognit. Lett.1