EDBT 2026 Demo / reviewers in the wild / expert
Yuan Yan Tang
dblp:t/YuanYanTang · also YuanYan Tang, Yuanyan Tang
· DBLP profile ↗
416ranked-venue papers
43as first author
82since 2021 · last 2026
0000-0002-6887-130XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 246 · 30 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 94 · 9 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 51 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 43 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 38 · 10 first-author · 2 since 2021Security and privacy · 14 · 8 since 2021Systems, architecture and hardware · 3Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSRM: Efficient spectral reconstruction Mamba with multiscale spectral-spatial correlation
Huafu Xu, Thomas Wu 0001, Yifeng Tan, Yuan Yan Tang |
Expert Syst. Appl. | 7 |
| 2026 | Prototype similarity-constraint enhancement network: A few-Shot class-Incremental learning for hyperspectral image classification
Yifeng Tan, Lianhui Liang, Huafu Xu, Thomas Wu 0001, Xichun Li, Yuan Yan Tang |
Expert Syst. Appl. | 8 |
| 2026 | Cross-modal dual-branch parallel hybrid matching for text-to-image person retrieval
Mian Hu, Jie Xu 0006, Chen Xu 0007, Yuan Yan Tang |
Neurocomputing | 5 |
| 2026 | Unlocking the Potential of Auxiliary Captions via Dual-Branch Multi-Scale Network for Composed Image RetrievalabstractComposed image retrieval (CIR) aims to retrieve target images by combining a reference image with a modification text. Traditional CIR methods often struggle with feature-level multimodal fusion, leading to deviations from the original embedding space. To address this, we propose a Dual-Branch Multi-Scale Network (DMN) that integrates a combining branch and a complete text branch. To enhance the use of captions generated by advanced image captioning models for CIR, the DMN leverages an attribute-driven disentanglement layer to separate features into distinct latent factors and employs a dual-path multimodal fusion module for effective feature integration. Additionally, a multi-scale matching module incorporating both global and local matching strategies is introduced to enhance fine-grained feature discrimination. Experimental results on the FashionIQ, Shoes, and CIRR datasets demonstrate that our DMN model consistently outperforms state-of-the-art methods, achieving improvements of up to 1.43% in mean recall metrics. Jinhong Xu, Xichun Li, Thomas Wu 0001, Yuan Yan Tang, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2026 | Similarity-Guided Denoising Reconstruction for Unsupervised Image CaptioningabstractImage captioning aims to generate natural and accurate textual descriptions of given images. Although significant progress has been made in image captioning models in recent years, most existing approaches heavily rely on high-quality image-text paired datasets that require expensive human annotation, thus limiting model scalability. Current unsupervised image captioning methods primarily focus on leveraging zero-shot learning capabilities of large pre-trained models (e.g. CLIP, GPT-2), yet still face persistent challenges including modality gaps, inefficient inference, and excessive noise incorporation, which constrain model accuracy and generalization capabilities. To address these limitations, we propose SGDR-Cap ( Similarity- Guided Denoising Reconstruction for Captioning), a novel unsupervised image captioning method that bridges the vision-language modality gap through a similarity-guided denoising reconstruction module. Our method leverages similarity information to guide the reconstruction of authentic text features during caption generation while simultaneously forcing the model to learn how to extract crucial image-relevant features and filter out unnecessary noise information. This enhances both coarse- and fine-grained cross-modal alignment. Furthermore, our approach jointly optimizes denoising reconstruction loss and language modeling loss, ensuring accuracy and fluency, and promoting greater diversity. Extensive evaluations on the MSCOCO and Flickr30K benchmarks demonstrate that our method achieves state-of-the-art results across all major metrics, with the most notable gain on the CIDEr score, improving from 101.1 to 104.4. Dongnan Yang, Thomas Wu 0001, Xichun Li, Yuan Yan Tang, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2026 | Tensor singular value-preserving norm for robust visual data recovery
Thomas Wu 0001, Yulong Wang 0002, Tianchuan Yang, Yuan Yan Tang |
Knowl. Based Syst. | 6 |
| 2026 | Robust graph neural networks via supervised block diagonal regularizer
Yulong Wang 0002, Huiwu Luo, Jinyu Tian 0001, Yuan Yan Tang |
Pattern Recognit. | 8 |
| 2026 | Emo-DiT: Emotional Speech Synthesis With a Diffusion Model Approach to Enhance Naturalness and Emotional ExpressivenessabstractCurrent emotional text-to-speech tasks have achieved high-quality emotional speech by incorporating emotion modules into text-to-speech models. However, there has been limited in-depth research on embedding emotion modules within TTS models, and the expression of emotion in synthesized speech is constrained by both the TTS model and the emotion module, often preventing optimal results. This paper presents a novel TTS model, Grad-DiT, based on the DiT architecture of diffusion models, aimed at enhancing the naturalness and expressiveness of TTS. Unlike traditional U-net architectures, Grad-DiT leverages the Transformer architecture to better capture contextual information in text, resulting in more natural speech generation. Building on this model, we propose Emo-DiT, which incorporates an Emotion Feature Reconstruction (EFR) module to enable the synthesis of speech with specific emotional expressions. Experimental results show that Grad-DiT not only surpasses existing TTS models in speech quality but also significantly improves real-time performance and inference speed. Compared to traditional emotion generation methods, Emo-DiT offers more precise emotional expression, enabling the synthesis of speech with distinctive emotional characteristics and providing substantial support for future applications in emotional speech synthesis. Bingzhen Wang, Jinhong Xu, Dongnan Yang, Miao Zhou, Yuan Yan Tang |
IEEE Trans. Affect. Comput. | 7 |
| 2026 | Efficient Diffusion-Based 3D Human Pose Estimation With Hierarchical Temporal PruningabstractDiffusion models have demonstrated strong capabilities in generating high-fidelity 3D human poses, yet their iterative nature and multi-hypothesis requirements incur substantial computational cost. In this paper, we propose an efficient diffusion-based 3D human pose estimation framework with a Hierarchical Temporal Pruning (HTP) strategy, which dynamically prunes redundant pose tokens across both frame and semantic levels while preserving critical motion dynamics. HTP operates in a staged, top-down manner: (1) Temporal Correlation-Enhanced Pruning (TCEP) identifies essential frames by analyzing inter-frame motion correlations through adaptive temporal graph construction; (2) Sparse-Focused Temporal MHSA (SFT MHSA) leverages the resulting frame-level sparsity to reduce attention computation, focusing on motion-relevant tokens; and (3) Mask-Guided Pose Token Pruner (MGPTP) performs fine-grained semantic pruning via clustering, retaining only the most informative pose tokens. Experiments on Human3.6M and MPI-INF-3DHP show that HTP reduces training MACs by 38.5%, inference MACs by 56.8%, and improves inference speed by an average of 81.1% compared to prior diffusion-based methods, while achieving state-of-the-art performance. Yuquan Bi, Hongsong Wang 0001, Xinli Shi, Zhipeng Gui, Jie Gui, Yuan Yan Tang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Gradient Perturbation Guidance for Boosting Sparse Adversarial Attack TransferabilityabstractSparse adversarial attacks perturb only a few pixels to achieve an attack, making them harder to detect and more dangerous. Recently, generative sparse attacks decouple the generation of sparse adversarial examples (AEs) into dense perturbations and sparse masks. By modeling the data distribution from clean examples to sparse AEs, generative sparse attacks mitigate the poor transferability that arises from over-reliance on gradients. These methods put effort into deriving optimal sparse masks on the generated perturbation. However, the quality of perturbation generation has always been overlooked, which limits the transferability of sparse AEs. To explore the influence of perturbation quality, we conduct empirical analyses of sparse gradient-based perturbations. The results show that directly applying sparsity to gradient-based perturbations disrupts their holistic adversarial information, leading to degraded attack performance. Therefore, it is critical to extract key adversarial knowledge from gradient-based perturbations while preserving their overall integrity to guide sparse adversarial attacks. Motivated by this observation, we propose to extract essential adversarial information from gradient-based AEs to guide the generator to produce higher-quality dense perturbations and stronger transferable sparse AEs. Specifically, we introduce the Gradient Perturbation Guidance (GPG) sparse adversarial attack, which integrates gradient adversarial feature guidance and gradient perturbation guidance regularization. The former guides the generator to capture gradient-based adversarial features during encoding, while the latter refines adversarial knowledge from gradient-based perturbations during decoding. Extensive experiments on ImageNet-1K show that our GPG significantly boosts transferability compared to state-of-the-art methods under consistent sparsity constraints. Our code is available at Github. Chengze Jiang, Minjing Dong, Jie Gui, Lu Dong 0002, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Improving Question Embeddings With Cognitive Representation Optimization for Knowledge TracingabstractThe knowledge tracing (KT) aims to track changes in students' knowledge status and predict their future answers based on their historical answer records. Current research on KT modeling focuses on predicting student' future performance based on existing, unupdated records of student learning interactions. However, these approaches ignore the distractors (such as slipping and guessing) in the answering process and overlook that static cognitive representations are temporary and limited. Most of them assume that there are no distractors in the answering process and that the record representations fully represent the students' level of understanding and proficiency in knowledge. In this case, it may lead to many lack of synergy and incoordination issue in the original records. Therefore we propose a cognitive representation optimization for KT (CRO-KT) model, which utilizes a dynamic programming algorithm to optimize structure of cognitive representations. This ensures that the structure matches the students' cognitive patterns in terms of the difficulty of the exercises. Furthermore, we use the co-optimization algorithm to optimize the cognitive representations of the subtarget exercises in terms of the overall situation of exercises responses by considering all the exercises with co-relationships as a single goal. Meanwhile, the CRO-KT model fuses the learned relational embeddings from the bipartite graph with the optimized record representations in a weighted manner, enhancing the expression of students' cognition. Finally, experiments are conducted on three publicly available datasets respectively to validate the effectiveness of the proposed cognitive representation optimization model. The source code of CRDP-KT is available at https://github.com/bigdata-graph/CRO-KT. Lixiang Xu, Xianwei Ding, Xin Yuan 0008, Zhanlong Wang, Lu Bai 0001, Enhong Chen, Philip S. Yu, Yuan Yan Tang |
IEEE Trans. Cybern. | 8 |
| 2026 | Improving Fast Adversarial Training Paradigm: An Example Taxonomy PerspectiveabstractWhile adversarial training is an effective defense method against adversarial attacks, it notably increases the training cost. To this end, fast adversarial training (FAT) is presented for efficient training and has become a hot research topic. However, FAT suffers from catastrophic overfitting, which leads to a performance drop compared with multi-step adversarial training. However, the cause of catastrophic overfitting remains unclear and lacks exploration. In this paper, we present an example taxonomy in FAT, which suggests that catastrophic overfitting is correlated with the imbalance between the inner and outer optimization in FAT. Furthermore, we investigated the impact of varying degrees of training loss, revealing a correlation between training loss and catastrophic overfitting. Based on these observations, we redesign the loss function in FAT with the proposed dynamic label relaxation to concentrate the loss range and reduce the impact of misclassified examples. Meanwhile, we introduce batch momentum initialization to enhance diversity and prevent catastrophic overfitting in an efficient manner. Furthermore, we also propose Catastrophic Overfitting aware Loss Adaptation (COLA), which employs a separate training strategy for examples based on their loss degree. Our proposed method, named example taxonomy aware FAT (ETA), establishes an improved paradigm for FAT. Experiment results demonstrate that our ETA achieves higher robust accuracy than all other evaluated methods. Comprehensive experiments on four standard datasets demonstrate the competitiveness of our method. The source code and model checkpoints will be publicly released. Jie Gui, Chengze Jiang, Minjing Dong, Kun Tong, Xinli Shi, Yuan Yan Tang, Dacheng Tao |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | Axial-View-Oriented Contrastive Adversarial Training for Robust Point Cloud RecognitionabstractContrastive adversarial training emerges as an effective approach to enhancing model robustness in safety-critical applications, particularly point cloud recognition for autonomous driving and medical imaging. However, existing point cloud adversarial training methods mainly emphasize global contrastive learning while overlooking local geometric variations induced by adversarial perturbations. Motivated by the spatial and intensity variations of perturbations across axial views, we propose AVOC, a novel local-global adversarial training framework that utilizes axial-view-oriented contrastive learning. This framework leverages the smallest axial view for local contrastive learning, as it exhibits the highest perturbation differences, and utilizes the largest axial view for global contrastive learning, as it preserves global structural consistency. We conduct comprehensive experiments across four representative architectures, demonstrating significant robustness improvements on widely-adopted recognition benchmarks, including ModelNet40, ShapeNetPart, ModelNet40-C, and ScanObjectNN-C, and further validate its effectiveness on the large-scale KITTI benchmark for 3D object detection. Our results across diverse perturbation scenarios, encompassing white-box attacks, black-box attacks, and natural perturbations, demonstrate the consistent and significant model robustness enhancement of our proposed method. Jie Gui, Yu-Xin Zhang 0004, Xiaofeng Cong, Baosheng Yu, Zhipeng Gui, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | Rethinking Frequency Modeling: Tail-Aware Dynamic Adversarial Training for Long-Tailed RobustnessabstractAdversarial training (AT) is among the most effective defenses against adversarial attacks on deep neural networks. However, in real-world scenarios where data often follow long-tailed distributions, conventional AT methods struggle to handle such imbalance, resulting in severe robustness disparities across classes and limited overall robustness. Although recent efforts attempt to improve robustness through class frequency-aware weighting or distribution adjustments, our empirical analysis reveals that class frequency alone is an insufficient indicator of adversarial vulnerability, as robust accuracy does not correlate with the number of examples per class. Furthermore, AT under long-tailed distributions exhibits optimization instability, particularly for tail classes with limited data. To address these challenges, we present Tail-Aware Dynamic Adversarial Training (TAD-AT), which integrates three complementary components targeting the training loss, attack strategy, and weight average. TAD-AT captures data imbalance and performance disparity, improving adversarial robustness under long-tailed distributions. First, our training loss incorporates frequency- and accuracy-aware regularization to emphasize learning for vulnerable classes. Second, our attack adjusts perturbations based on class-wise vulnerability, encouraging robust feature learning around vulnerable regions, thereby mitigating robustness overfitting and improving clean accuracy. Third, our weight average improves robust generalization and training stability by adaptively controlling the decay rate across classes. Experiments on long-tailed benchmarks demonstrate that our TAD-AT significantly improves adversarial robustness, offering a systematic and practical solution to robustness challenges under long-tail distributions. Our code is publicly available on https://github.com/bookman233/TADAT. Chengze Jiang, Minjing Dong, Jie Gui, Ju Jia, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | PANDA: Diffusion-Guided Purification and Adaptation for Robust Point Cloud Classification Against Adversarial Attack
Yu-Xin Zhang 0004, Xiaofeng Cong, Minjing Dong, Zhipeng Gui, Jie Gui, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | From Individual to Universal: Regularized Multi-view Joint Representation for Multi-view Subspace-Preserving RecoveryabstractRecent years have witnessed an explosion of Multi- view Subspace Classification (MSCla) and Multi-view Subspace Clustering (MSClu) methods for various applications. However, their theoretical foundation have not been well explored and understood. In this paper, we investigate the multi-view subspace-preserving recovery theory, which is the theoretical underpinnings for MSCla and MSClu methods. Specifically, we derive novel geometrically interpretable conditions for the success of multi-view subspace-preserving recovery. Compared with prior related works, we make the following innovations: First, our theory does not require the equality constraint, which is a common requirement in prior theoretical works and may be too restrictive in reality. Second, we provide both Individual Theoretical Guarantee (ITG) and Universal Theoretical Guarantee (UTG) for multi-view subspace-preserving recovery while prior works only give the UTG. Third, we also apply the proposed theory to establish theoretical guarantees for MSCla and MSClu, respectively. Numerical results validate the proposed theory for multi-view subspace-preserving recovery. Yulong Wang 0002, Xinwei He 0001, Qiwei Xie, Kit Ian Kou, Yuan Yan Tang |
IJCAI | 6 |
| 2025 | Self-supervised extracted contrast network for facial expression recognition
Lingyu Yan, Jinquan Yang, Jinyao Xia, Rong Gao 0001, Li Zhang 0013, Yuan Yan Tang |
Multim. Tools Appl. | 7 |
| 2025 | A novel perturbation-based degraded image super-resolution method for object recognition in intelligent transportation system
Shan Zeng, Zhiguang Yang, Hao Li 0034, Yuan Yan Tang |
Neural Comput. Appl. | 6 |
| 2025 | ALR-HT: A fast and efficient Lasso regression without hyperparameter tuning
Bin Zou 0002, Jie Xu 0006, Chen Xu 0007, Yuan Yan Tang |
Neural Networks | 5 |
| 2025 | A Privacy-Preserving Large-Scale Image Retrieval Framework With Vision GNN HashingabstractWith the growing popularity of cloud services, companies and individuals outsource images to cloud servers to reduce storage and computing burdens. The images are encrypted before outsourcing for privacy protection. It has become urgent to solve the privacy-preserving image retrieval problem on the cloud. There are three main challenges in this area. First, how can we achieve high retrieval accuracy on the encryption domain? Second, how can we improve efficiency in large-scale encrypted image retrieval? Third, how can we ensure the reliability of the retrieval results? The existing schemes only consider some of these characteristics and the retrieval accuracy is insufficient. In this paper, we propose a privacy-preserving large-scale image retrieval framework with vision graph convolutional neural network hashing (ViGH). To the best of our knowledge, this is the first framework that is able to address all the above challenges with more advanced accuracy performance. To be specific, cycle-consistent adversarial networks and vision graph convolutional networks (ViG) are utilized to increase retrieval accuracy. By embedding encrypted images into hash codes, we can obtain high retrieval efficiency by Hamming distances. Cloud servers store the hash codes on the blockchain (Ethereum). The retrieval algorithm on the smart contracts and the consensus mechanism of blockchain ensure reliability of the retrieval results. The experimental results on three common datasets verify the effectiveness and efficiency of the proposed privacy-preserving image retrieval framework. The reliability of the retrieval results is ensured by the consensus mechanism of blockchain with no need for verification. Yuan Cao 0005, Fanlei Meng, Xinzheng Shang, Jie Gui, Yuan Yan Tang |
IEEE Trans. Big Data | 5 |
| 2025 | Adjustable Jacobi-Fourier Moment for Image RepresentationabstractThe widely adopted Jacobi-Fourier moment (JFM) is limited by its inability to effectively capture spatial information. Although fractional-order JFM (FOJFM) introduces spatial information through a fractional-order parameter, the control of spatial information remains inadequate. This limitation stems from the insufficient control over zeros distribution associated with the used moment's radial kernel. To address this issue, we generalize both JFM and FOJFM into a transformed JFM. A transformed function with four parameters is designed, and adjustable JFM (AJFM) is proposed. Two parameters correlate to increasing velocities on the left and right parts of the transformed functions, enabling zeros quantities of radial kernel fall in the left and right parts of the interval. The other two parameters segment the transformed function, adjusting regions where different quantities of zeros fall in. This refined control over the radial kernel's zero distribution enhances the versatility of feature extraction by the AJFM, governed by the introduced parameters. Experimental results demonstrate that AJFM, with properly chosen parameters, can emphasize specific regions within an image more effectively. Xin Yuan 0008, Xiaoqi Lu, Yuan Yan Tang |
IEEE Trans. Cybern. | 4 |
| 2025 | Improving Fast Adversarial Training via Self-Knowledge GuidanceabstractAdversarial training has achieved remarkable advancements in defending against adversarial attacks. Among them, fast adversarial training (FAT) is gaining attention for its ability to achieve competitive robustness with fewer computing resources. Existing FAT methods typically employ a uniform strategy that optimizes all training data equally without considering the influence of different examples, which leads to an imbalanced optimization. However, this imbalance remains unexplored in the field of FAT. In this paper, we conduct a comprehensive study of the imbalance issue in FAT and observe an obvious class disparity regarding their performances. This disparity could be embodied from a perspective of alignment between clean and robust accuracy. Based on the analysis, we mainly attribute the observed misalignment and disparity to the imbalanced optimization in FAT, which motivates us to optimize different training data adaptively to enhance robustness. Specifically, we take disparity and misalignment into consideration. First, we introduce self-knowledge guided regularization, which assigns differentiated regularization weights to each class based on its training state, alleviating class disparity. Additionally, we propose self-knowledge guided label relaxation, which adjusts label relaxation according to the training accuracy, alleviating the misalignment and improving robustness. By combining these methods, we formulate the Self-Knowledge Guided FAT (SKG-FAT), leveraging naturally generated knowledge during training to enhance the adversarial robustness without compromising training efficiency. Extensive experiments on four standard datasets demonstrate that the SKG-FAT improves the robustness and preserves competitive clean accuracy, outperforming the state-of-the-art methods. Code and checkpoints are available at SFG-FAT Code Implementation. Chengze Jiang, Minjing Dong, Jie Gui, Xinli Shi, Yuan Cao 0005, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | ColorVein: Colorful Cancelable Vein BiometricsabstractVein recognition technologies have become one of the primary solutions for high-security identification systems. However, the issue of biometric information leakage can still pose a serious threat to user privacy and anonymity. Currently, there is no cancelable biometric template generation scheme specifically designed for vein biometrics. Therefore, this paper proposes an innovative cancelable vein biometric generation scheme: ColorVein. Unlike previous cancelable template generation schemes, ColorVein does not destroy the original biometric features and introduces additional color information to grayscale vein images. This method significantly enhances the information density of vein images by transforming static grayscale information into dynamically controllable color representations through interactive colorization. ColorVein allows users/administrators to define a controllable pseudo-random color space for grayscale vein images by editing the position, number, and color of hint points, thereby generating protected cancelable templates. Additionally, we propose a new secure center loss to optimize the training process of the protected feature extraction model, effectively increasing the feature distance between enrolled users and any potential impostors. Finally, we evaluate ColorVein’s performance on all types of vein biometrics, including recognition performance, unlinkability, irreversibility, and revocability, and conduct security and privacy analyses. ColorVein achieves competitive performance compared with state-of-the-art methods. Yifan Wang 0036, Jie Gui, Xinli Shi, Linqing Gui, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Divide and Conquer: Frequency-Aware Contrastive Adversarial Training for Robust Point Cloud ClassificationabstractContrastive adversarial training has shown great potential in enhancing model robustness and has been adopted in point cloud classification. There are varying spatial distributions and densities across different regions in point cloud data, which makes adversarial perturbations always exhibit non-uniform patterns of attack intensity and distribution in different regions. However, existing approaches always rely on uniform feature contrast without considering the granularity in the context of point cloud data, limiting their capacities to counter adversarial perturbations effectively. To address this issue, we propose a novel frequency-aware contrastive adversarial training framework, which considers feature contrast via a “divide-and-conquer” method. Specifically, we systematically “divide” point clouds into distinct frequency components and “conquer” feature contrast within each frequency band, which fosters fine-grained feature consistency learning and leads to more informative as well as robust representations. Besides, existing methods typically apply group-level contrastive learning, which emphasizes category-wise similarity but often overlooks the nuanced structural variations among instances. To remedy this, we incorporate instance-level contrastive learning to capture per-instance geometric variations. Moreover, a frequency-specific hard-masked sample generation module is designed to construct challenging sample pairs by masking keypoint features in each frequency band, thereby promoting the model to learn more robust feature representations. Extensive experiments on multiple benchmark datasets demonstrate that our proposed method significantly outperforms existing state-of-the-art approaches in adversarial robustness for point cloud classification. The code is available on DiCon-FAT. Yu-Xin Zhang 0004, Jie Gui, Minjing Dong, Xiaofeng Cong, Yuan Cao 0005, Xin Gong 0001, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Exploring the Coordination of Frequency and Attention in Masked Image ModelingabstractRecently, masked image modeling (MIM), which learns visual representations by reconstructing the masked patches of an image, has become a popular self-supervised paradigm. However, the pre-training of MIM always takes massive time due to the large-scale data and large-size backbones. We mainly attribute it to the random patch masking in previous MIM works, which fails to leverage the crucial semantic information for effective visual representation learning. To tackle this issue, we propose the Frequency & Attention-driven Masking and Throwing Strategy (FAMT), which can detect semantic patches and reduce the number of training patches to boost model performance and training efficiency simultaneously. Specifically, FAMT utilizes the self-attention mechanism to extract semantic information from the image for masking during training in an unsupervised manner. However, attention alone could sometimes focus on inappropriate areas regarding the semantic information. Thus, we are motivated to incorporate the information from the frequency domain into the self-attention mechanism to derive the sampling weights for masking, which captures semantic patches for visual representation learning. Furthermore, we introduce a patch throwing strategy based on the derived sampling weights to reduce the training cost. FAMT can be seamlessly integrated as a plug-and-play module and surpasses previous works, e.g. reducing the training phase time by nearly 50% and improving the linear probing accuracy of MAE by $1.8$ % ~ $ 6.3$ % across various datasets, including CIFAR-10/100, Tiny ImageNet, and ImageNet-1K. FAMT also demonstrates superior performance in downstream detection and segmentation tasks. Jie Gui, Tuo Chen, Minjing Dong, Zhengqi Liu, Hao Luo 0004, James T. Kwok, Yuan Yan Tang |
IEEE Trans. Image Process. | 7 |
| 2025 | Tensor Nuclear Norm-Based Multi-Channel Atomic Representation for Robust Face RecognitionabstractNumerous representation-based classification (RC) methods have been developed for face recognition due to their decent model interpretability and robustness against noise. Most existing RC methods primarily characterize the gray-scale reconstruction error image (single-channel data) in two ways: the one-dimensional (1D) pixel-based error model and the two-dimensional (2D) gray-scale image-matrix-based error model. The former measures the reconstruction error pixel by pixel, while the latter leverages 2D structural information of the gray-scale error image, such as the low-rank property. However, when applying these methods to different color channels of a test color face image (multi-channel data) separately and independently, they neglect the three-dimensional (3D) structural correlations among distinct color channels. In real-world scenarios, face images are often contaminated with complex noise, including contiguous occlusion and random pixel corruption, which pose significant challenges to these approaches and can lead to a decline in performance. In this paper, we propose a Tensor Nuclear Norm based Robust Multi-channel Atomic Representation (TNN-RMAR) framework with application to color face recognition. The proposed method has the following three critical ingredients: 1) We propose a 3D color image-tensor-based error model, which can take full advantage of the 3D structural information of the color error image. 2) To leverage the 3D structural information of the color error image, we model it as a 3-order tensor and exploit its low-rank property with the tensor nuclear norm. Given that multiple color channels in a color image are generally corrupted at the same positions, we design a tube-wise tailored loss function to further leverage its tube-wise structure. 3) We devise the multi-channel atomic norm (MAN) regularization for the representation coefficient matrix, which allows us to jointly harness the correlation information of coefficients in different color channels. In addition, we also devise an efficient algorithm to solve the TNN-RMAR framework based on the alternating direction method of multipliers (ADMM) framework. By leveraging TNN-RMAR as a general platform, we also develop several novel robust multi-channel RC methods. Experimental results on benchmark real-world databases validate the effectiveness and robustness of the proposed framework for robust color face recognition. Yulong Wang 0002, Hong Chen 0004, Yuan Yan Tang |
IEEE Trans. Image Process. | 6 |
| 2025 | Unrevealed Threats: Adversarial Robustness Analysis of Underwater Image Enhancement ModelsabstractLearning-based methods for underwater image enhancement (UWIE) have undergone extensive exploration. However, learning-based models are usually vulnerable to adversarial examples so as the UWIE models. To the best of our knowledge, there is no comprehensive study on the adversarial robustness of UWIE models, which indicates that UWIE models are potentially under the threat of adversarial attacks. In this paper, we propose a general adversarial attack protocol. We make a first attempt to conduct adversarial attacks on five well-designed UWIE models on three common underwater image benchmark datasets. Considering the scattering and absorption of light in the underwater environment, there exists a strong correlation between color correction and underwater image enhancement. On the basis of that, we also design two effective UWIE-oriented adversarial attack methods, Pixel Attack and Color Shift Attack targeting different color spaces. The results show that five models exhibit varying degrees of vulnerability to adversarial attacks and well-designed small perturbations on degraded images are capable of preventing UWIE models from generating enhanced results. In addition, we conduct adversarial training on these models and successfully mitigated the effectiveness of adversarial attacks. In summary, we reveal the adversarial vulnerability of UWIE models and propose a new evaluation dimension of UWIE models. Siyu Zhai, Zhibo He, Xiaofeng Cong, Junming Hou, Jie Gui, Jian Wei You, Xin Gong 0001, James T. Kwok, Yuan Yan Tang |
IEEE Trans. Multim. | 9 |
| 2025 | Layer-Wise Mutual Information Meta-Learning Network for Few-Shot SegmentationabstractThe goal of few-shot segmentation (FSS) is to segment unlabeled images belonging to previously unseen classes using only a limited number of labeled images. The main objective is to transfer label information effectively from support images to query images. In this study, we introduce a novel meta-learning framework called layer-wise mutual information (LayerMI), which enhances the propagation of label information by maximizing the mutual information (MI) between support and query features at each layer. Our approach involves the utilization of a LayerMI Block based on information-theoretic co-clustering. This block performs online co-clustering on the joint probability distribution obtained from each layer, generating a target-specific attention map. The LayerMI Block can be seamlessly integrated into the meta-learning framework and applied to all convolutional neural network (CNN) layers without altering the training objectives. Notably, the LayerMI Block not only maximizes MI between support and query features but also facilitates internal clustering within the image. Extensive experiments demonstrate that LayerMI significantly enhances the performance of baseline and achieves competitive performance compared to state-of-the-art methods on three challenging benchmarks: PASCAL- $5^{i}$ , COCO- $20^{i}$ , and FSS-1000. Xiaoliu Luo, Zhao Duan, Anyong Qin, Zhuotao Tian, Ting Xie 0004, Taiping Zhang, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Graph Augmentation Empowered Contrastive Learning for RecommendationabstractThe application of contrastive learning (CL) to collaborative filtering (CF) in recommender systems has achieved remarkable success. CL-based recommendation models mainly focus on creating multiple augmented views by employing different graph augmentation methods and utilizing these views for self-supervised learning. However, current CL methods for recommender systems usually struggle to fully address the problem of noisy data. To address this problem, we propose the G raph A ugmentation E mpowered C ontrastive L earning (GAECL) for recommendation framework, which uses graph augmentation based on topological and semantic dual adaptation and global co-modeling via structural optimization to co-create contrasting views for better augmentation of the CF paradigm. Specifically, we strictly filter out unimportant topologies by reconstructing the adjacency matrix and mask unimportant attributes in nodes according to the PageRank centrality principle to generate an augmented view that filters out noisy data. Additionally, GAECL achieves global collaborative modeling through structural optimization and generates another augmented view based on the PageRank centrality principle. This helps to filter the noisy data while preserving the original semantics of the data for more effective data augmentation. Extensive experiments are conducted on five datasets to demonstrate the superior performance of our model over various recommendation models. Lixiang Xu, Yusheng Liu 0003, Tong Xu 0001, Enhong Chen, Yuan Yan Tang |
ACM Trans. Inf. Syst. | 5 |
| 2024 | Discriminative latent subspace learning with adaptive metric learning
Yuan Yan Tang, Zhaowei Shang |
Neural Comput. Appl. | 2 |
| 2024 | PFENet++: Boosting Few-Shot Semantic Segmentation With the Noise-Filtered Context-Aware Prior MaskabstractIn this work, we revisit the prior mask guidance proposed in “Prior Guided Feature Enrichment Network for Few-Shot Segmentation”. The prior mask serves as an indicator that highlights the region of interests of unseen categories, and it is effective in achieving better performance on different frameworks of recent studies. However, the current method directly takes the maximum element-to-element correspondence between the query and support features to indicate the probability of belonging to the target class, thus the broader contextual information is seldom exploited during the prior mask generation. To address this issue, first, we propose the Context-aware Prior Mask (CAPM) that leverages additional nearby semantic cues for better locating the objects in query images. Second, since the maximum correlation value is vulnerable to noisy features, we take one step further by incorporating a lightweight Noise Suppression Module (NSM) to screen out the unnecessary responses, yielding high-quality masks for providing the prior knowledge. Both two contributions are experimentally shown to have substantial practical merit, and the new model named PFENet++ significantly outperforms the baseline PFENet as well as all other competitors on three challenging benchmarks PASCAL-5$^{i}$, COCO-20$^{i}$and FSS-1000. The new state-of-the-art performance is achieved without compromising the efficiency, manifesting the potential for being a new strong baseline in few-shot semantic segmentation. Xiaoliu Luo, Zhuotao Tian, Taiping Zhang, Bei Yu 0001, Yuan Yan Tang, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Response Generation in Social Network With Topic and Emotion ConstraintsabstractResponse generation is the task of automatically generating human-like content based on the provided context. One of its prominent applications is to simulate realistic response content for social network posts. In the digital age, social network platforms play a vital role in information exchange and social interaction. This study focuses on response generation techniques for the platform of public opinion evolution simulation that simulate realistic response content, enabling a deeper understanding of the emotional expressions of network users. Recent advancements in deep learning techniques, particularly the sequence-to-sequence (Seq2Seq) model, have shown promise in the response generation field. However, we still face two challenges: content variety, topic and emotion relevancy. To this end, we propose the EmoTG-ETRS model which comprises three parts. The first is a response generation module based on Transformer architecture. Then, an auxiliary emotion improvement module is incorporated to enhance the emotional expressiveness of the response candidates. Finally, a reverse selection module, which combines maximum mutual information (MMI) evaluation, emotional expression evaluation, and topic consistency evaluation, is devised to select the highest-scoring response. Extensive experiments have been conducted to evaluate the effectiveness of the proposed model and the results demonstrate that the EmoTG-ETRS model improves the quality of produced replies in terms of topic consistency and emotional accuracy rate when compared with the SOTA research works. Biwei Cao, Jiuxin Cao, Bo Liu 0004, Jie Gui, Jun Zhou 0027, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | Fooling the Image Dehazing Models by First Order GradientabstractThe research on the single image dehazing task has been widely explored. However, as far as we know, no comprehensive study has been conducted on the robustness of the well-trained dehazing models. Therefore, there is no evidence that the dehazing networks can resist malicious attacks. In this paper, we focus on designing a group of attack methods based on first order gradient to verify the robustness of the existing dehazing algorithms. By analyzing the general purpose of image dehazing task, four attack methods are proposed, which are predicted dehazed image attack, hazy layer mask attack, haze-free image attack and haze-preserved attack. The corresponding experiments are conducted on six datasets with different scales. Further, the defense strategy based on adversarial training is adopted for reducing the negative effects caused by malicious attacks. In summary, this paper defines a new challenging problem for the image dehazing area, which can be called as adversarial attack on dehazing networks (AADN). Code is available at https://github.com/Xiaofeng-life/AADN_Dehazing. Jie Gui, Xiaofeng Cong, Chengwei Peng, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | GraKerformer: A Transformer With Graph Kernel for Unsupervised Graph Representation LearningabstractWhile highly influential in deep learning, especially in natural language processing, the Transformer model has not exhibited competitive performance in unsupervised graph representation learning (UGRL). Conventional approaches, which focus on local substructures on the graph, offer simplicity but often fall short in encapsulating comprehensive structural information of the graph. This deficiency leads to suboptimal generalization performance. To address this, we proposed the GraKerformer model, a variant of the standard Transformer architecture, to mitigate the shortfall in structural information representation and enhance the performance in UGRL. By leveraging the shortest-path graph kernel (SPGK) to weight attention scores and combining graph neural networks, the GraKerformer effectively encodes the nuanced structural information of graphs. We conducted evaluations on the benchmark datasets for graph classification to validate the superior performance of our approach. Lixiang Xu, Haifeng Liu 0004, Xin Yuan 0008, Enhong Chen, Yuan Yan Tang |
IEEE Trans. Cybern. | 5 |
| 2024 | Soft Multiprototype Clustering Algorithm via Two-Layer Semi-NMFabstractThis article proposes a novel soft multiprototype clustering algorithm (SMP) for high-dimensional data clustering with noisy and complex structural patterns. SMP integrates dimensionality reduction, multiprototype clustering, and multiprototype merge clustering under a two-layer seminonnegative matrix factorization (semi-NMF) architecture. Specifically, the first semi-NMF layer performs multiprototype clustering, which solves the problem that a single prototype cannot represent complex data structures. Meanwhile, the multiprototype fuzzy clustering constraints ensure that the multiprototypes better characterize the original data structure. The second semi-NMF layer performs multiprototype merge clustering to mitigate the issues of heavy computation burden and poor antinoise performance of the spectral clustering algorithm. The introduction of the Laplace graph matrix regularization constraint in this layer assists SMP in completing the merging of multiprototypes with complex data structures. Comprehensive experiments demonstrate that the proposed method outperforms the state-of-the-art algorithms. Shan Zeng, Xiangjun Duan, Kun Hu 0008, Yuan Yan Tang |
IEEE Trans. Fuzzy Syst. | 6 |
| 2024 | CFVNet: An End-to-End Cancelable Finger Vein Network for RecognitionabstractFinger vein recognition technology has become one of the primary solutions for high-security identification systems. However, it still has information leakage problems, which seriously jeopardizes user’s privacy and anonymity and cause great security risks. In addition, there is no work to consider a fully integrated secure finger vein recognition system. So, different from the previous systems, we integrate preprocessing and template protection into an integrated deep learning model. We propose an end-to-end cancelable finger vein network (CFVNet), which can be used to design an secure finger vein recognition system. It includes a plug-and-play BWR-ROIAlign unit, which consists of three sub-modules: Localization, Compression and Transformation. The localization module achieves automated localization of stable and unique finger vein ROI. The compression module losslessly removes spatial and channel redundancies. The transformation module uses the proposed BWR method to introduce unlinkability, irreversibility and revocability to the system. BWR-ROIAlign can directly plug into the model to introduce the above features for DCNN-based finger vein recognition systems. We perform extensive experiments on four public datasets to study the performance and cancelable biometric attributes of the CFVNet-based recognition system. The average accuracy, EERs and$D_{\leftrightarrow } ^{sys}$on the four datasets are 99.82%, 0.01% and 0.025, respectively, and achieves competitive performance compared with the state-of-the-arts. Yifan Wang 0036, Jie Gui, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Illumination Controllable Dehazing Network based on Unsupervised Retinex EmbeddingabstractOn the one hand, the dehazing task is an ill-posedness problem, which means that no unique solution exists. On the other hand, the dehazing task should take into account the subjective factor, which is to give the user selectable dehazed images rather than a single result. Therefore, this paper proposes a multi-output dehazing network by introducing illumination controllable ability, called IC-Dehazing. The proposed IC-Dehazing can change the illumination intensity by adjusting the factor of the illumination controllable module, which is realized based on the interpretable Retinex model. Moreover, the backbone dehazing network of IC-Dehazing consists of a Transformer with double decoders for high-quality image restoration. Further, the prior-based loss function and unsupervised training strategy enable IC-Dehazing to complete the parameter learning process without the need for paired data. To demonstrate the effectiveness of the proposed IC-Dehazing, quantitative and qualitative experiments are conducted. Code is available athttps://github.com/Xiaofeng-life/ICDehazing. Jie Gui, Xiaofeng Cong, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Multim. | 4 |
| 2024 | Group Multi-View Transformer for 3D Shape Analysis With Spatial EncodingabstractIn recent years, the results of view-based 3D shape recognition methods have saturated, and models with excellent performance cannot be deployed on memory-limited devices due to their huge size of parameters. To address this problem, we introduce a compression method based on knowledge distillation for this field, which largely reduces the number of parameters while preserving model performance as much as possible. Specifically, to enhance the capabilities of smaller models, we design a high-performing large model called Group Multi-view Vision Transformer (GMViT). In GMViT, the view-level ViT first establishes relationships between view-level features. Additionally, to capture deeper features, we employ the grouping module to enhance view-level features into group-level features. Finally, the group-level ViT aggregates group-level features into complete, well-formed 3D shape descriptors. Notably, in both ViTs, we introduce spatial encoding of camera coordinates as innovative position embeddings. Furthermore, we propose two compressed versions based on GMViT, namely GMViT-simple and GMViT-mini. To enhance the training effectiveness of the small models, we introduce a knowledge distillation method throughout the GMViT process, where the key outputs of each GMViT component serve as distillation targets. Extensive experiments demonstrate the efficacy of the proposed method. The large model GMViT achieves excellent 3D classification and retrieval results on the benchmark datasets ModelNet, ShapeNetCore55, and MCB. The smaller models, GMViT-simple and GMViT-mini, reduce the parameter size by 8 and 17.6 times, respectively, and improve shape recognition speed by 1.5 times on average, while preserving at least 90% of the recognition performance. Lixiang Xu, Qingzhe Cui, Richang Hong, Enhong Chen, Xin Yuan 0008, Chenglong Li 0002, Yuan Yan Tang |
IEEE Trans. Multim. | 8 |
| 2023 | GENet: Guidance Enhancement Network for 3D Shape RecognitionabstractBoth point cloud-based and view-based deep learning methods for 3D shape recognition have achieved relatively remarkable results in recent years. However, there are few methods to jointly represent 3D shapes from both point cloud and multi-view modal data. Therefore, we propose a guidance enhancement network (GENet) for 3D shape recognition based on multimodal data. On the one hand, the point cloud is encoded with features from both explicit and implicit aspects, and on the other hand, all views are encoded and constructed as a graph. In the multilayer guidance enhancement module, graph convolutional neural network (GCN) enhances each view feature, and then temporary high-level features (initially point cloud global feature) guide multiple low-level view features to obtain correlation coefficients, through which the views with higher importance are filtered as inputs for the next layer of the structure and the view features in the current layer are weighted and aggregated. The aggregated view features are then connected to the high-level features with residuals to form the enhanced high-level features. The 3D shape descriptor is finally obtained after several guidance and enhancements. The proposed GENet achieves state-of-the-art results on the 3D benchmark dataset ModelNet. Xiaofeng Wang 0009, Qingzhe Cui, Lixiang Xu, Haifeng Liu 0004, Lixin He, Bin Luo 0001, Sibao Chen 0001, Yuan Yan Tang |
IJCNN | 8 |
| 2023 | GLCNet: Global-Local Complementary Network for 3D Shape RecognitionabstractBoth point cloud-based and multi-view-based methods have achieved remarkable results in 3D shape recognition, yet there are few methods that combine the two types of data. In this paper, a novel Global-Local Complementary Network (GLCNet) based on multimodal data is proposed. The network obtains more powerful shape descriptors by stacking multiple layers of Global-Local Complementary Module (GLC Module). More specifically, the Global-Local Relation Score Module is first used to obtain the relationship between view features and global feature. The relationship is then utilized to facilitate the aggregation of view features and to filter out the more important ones. Finally, the aggregated view features are fused with the global features to form a stronger global feature. GLCNet enables the characteristics of various data to be fully utilized and achieves a true sense of complementarity of strengths and weaknesses. Extensive experiments on the benchmark dataset ModelNet show that GLCNet achieves state-of-the-art results in 3D shape classification and retrieval. Xiaofeng Wang 0009, Qingzhe Cui, Lixiang Xu, Haifeng Liu 0004, Lixin He, Bin Luo 0001, Sibao Chen 0001, Yuan Yan Tang |
IJCNN | 8 |
| 2023 | UGTransformer: Unsupervised Graph Transformer Representation LearningabstractThis paper mainly studies graph representation learning in unsupervised scenarios combined with Transformer models. Transformer network models have been widely used in many fields of machine learning and deep learning, and the application of transformer architectures to graph data has been very popular recently. For graph data, the field of graph representation learning has recently attracted a lot of attention. Graph-level representation is widely used in the real world, such as drug molecule design and disease classification in biochemistry. Traditional graph kernel methods, which design different graph kernels for different substructures, are simple but have poor generalization performance. Recently methods based on language models, such as graph2vec, use a particular substructure as the graph representation, which is also similar to the hand-crafted approach and also leads to poor generalization ability. In this paper, we propose the UGTransformer model, which builds on the standard Transformer architecture. We introduce several simple and effective structural encoding methods in order to encode the structural information of the graph into the model efficiently. The unsupervised representation of graphs is learned through a multi-headed attention mechanism and by using powerful aggregation functions. We conducted experiments on a benchmark date set for graph classification, and the experimental results validate the effectiveness of our proposed model. Lixiang Xu, Haifeng Liu 0004, Qingzhe Cui, Bin Luo 0001, Yan Chen 0037, Yuan Yan Tang |
IJCNN | 7 |
| 2023 | Distribution preserving-based deep semi-NMF for data representation
Anyong Qin, Zhuolin Tan, Xingli Tan, Cheng Jing, Yuan Yan Tang |
Neurocomputing | 6 |
| 2023 | MASK-CNN-Transformer for real-time multi-label weather recognition
Shengchao Chen, Ting Shu 0001, Huan Zhao 0004, Yuan Yan Tang |
Knowl. Based Syst. | 4 |
| 2023 | Generalization capacity of multi-class SVM based on Markovian resampling
Zijie Dong, Chen Xu 0007, Jie Xu 0006, Bin Zou 0002, Jingjing Zeng, Yuan Yan Tang |
Pattern Recognit. | 6 |
| 2023 | Semantic-based conditional generative adversarial hashing with pairwise labels
Qi Li 0005, Weining Wang 0001, Yuan Yan Tang, Cheng-Zhong Xu 0001, Zhenan Sun |
Pattern Recognit. | 3 |
| 2023 | Simultaneous Robust Matching Pursuit for Multi-view Learning
Yulong Wang 0002, Kit Ian Kou, Hong Chen 0004, Yuan Yan Tang, Luoqing Li |
Pattern Recognit. | 4 |
| 2023 | Adaptive reweighted quaternion sparse learning for data recovery and classification
Cuiming Zou, Kit Ian Kou, Yuan Yan Tang |
Pattern Recognit. | 3 |
| 2023 | Probabilistic quaternion collaborative representation and its application to robust color face identification
Cuiming Zou, Kit Ian Kou, Yuan Yan Tang |
Signal Process. | 3 |
| 2023 | Learning Performance of Weighted Distributed Learning With Support Vector MachinesabstractThe divide-and-conquer strategy is a very effective method of dealing with big data. Noisy samples in big data usually have a great impact on algorithmic performance. In this article, we introduce Markov sampling and different weights for distributed learning with the classical support vector machine (cSVM). We first estimate the generalization error of weighted distributed cSVM algorithm with uniformly ergodic Markov chain (u.e.M.c.) samples and obtain its optimal convergence rate. As applications, we obtain the generalization bounds of weighted distributed cSVM with strong mixing observations and independent and identically distributed (i.i.d.) samples, respectively. We also propose a novel weighted distributed cSVM based on Markov sampling (DM-cSVM). The numerical studies of benchmark datasets show that the DM-cSVM algorithm not only has better performance but also has less total time of sampling and training compared to other distributed algorithms. Bin Zou 0002, Chen Xu 0007, Jie Xu 0006, Xinge You, Yuan Yan Tang |
IEEE Trans. Cybern. | 6 |
| 2023 | Tensorial Multiview Representation for Saliency Detection via Nonconvex ApproachabstractIn the study of salient object detection, multiview features play an important role in identifying various underlying salient objects. As to current common patch-based methods, all different features are handled directly by stacking them into a high-dimensional vector to represent related image patches. These approaches ignore the correlations inhering in the original spatial structure, which may lead to the loss of certain underlying characterization such as view interaction. In this article, different from currently available approaches, a tensorial feature representation framework is developed for the salient object detection in order to better explore the complementary information of multiview features. Under the tensor framework, a tensor low-rank constraint is applied to the background to capture its intrinsic structure, a tensor group sparsity regularization is posed on the salient part, and a tensorial sliced Laplacian regularization is then introduced to enlarge the gap between the subspaces of the background and salient object. Moreover, a nonconvex tensor Log-determinant function, instead of the tensor nuclear norm, is adopted to approximate the tensor rank for effectively suppressing the confusing information resulted from underlying complex backgrounds. Further, we have deduced the closed-form solution of this nonconvex minimization problem and established a feasible algorithm whose convergence is mathematically proven. Experiments on five well-known public datasets are provided and the simulations demonstrate that our method outperforms the latest unsupervised handcrafted features-based methods in the literature. Furthermore, our model is flexible with various deep features and is competitive with the state-of-the-art approaches. Chen Xu 0004, Mingqing Xiao 0001, Yuan Yan Tang |
IEEE Trans. Cybern. | 5 |
| 2023 | A Sparse Framework for Robust Possibilistic K-Subspace ClusteringabstractClustering noisy, high-dimensional, and structurally complex data have always been a challenging task. As most existing clustering methods are not able to deal with both the adverse impact of noisy samples and the complex structures of data, in this article, we propose a novel robust and sparse possibilistic K-subspace (RSPKS) clustering algorithm to integrate subspace recovery and possibilistic clustering algorithms under a unified sparse framework. First, the proposed method sparsifies the membership matrix and the subspace projection vector under a dual-sparse framework to handle high-dimensional noisy data. This unifies dimensionality reduction and clustering using one objective function for which the optimization can be realized through synchronous iteration. Second, the reconstruction error of each sample in the local subspace is used as the distance metric for classification. That is, each sample itself is treated as a clustering prototype so as not to be affected by the structure of the overall data distribution. Therefore, the clustering prototype construction problem of the data with complex structures can be better addressed. Finally, to deal with nonlinear regions, our RSPKS method is further extended into a kernelized version, namely the kernelized RSPKS clustering algorithm. The experimental results on both synthetic and real-world datasets demonstrate that our proposed method outperforms state-of-the-art algorithms in terms of clustering accuracy. Shan Zeng, Xiangjun Duan, Hao Li 0034, Yuan Yan Tang, Zhiyong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2023 | A Novel Spatial-Spectral Pyramid Network for Hyperspectral Image ClassificationabstractAs the research on deep learning methods gradually progresses, more and more classification models are applied in the classification of hyperspectral image. High-dimensional and low-resolution characteristics of hyperspectral image (HSI), however, make it difficult for conventional models to process its data effectively. In this paper, a novel HSI classification model, namely Spatial Spectral Pyramid Network (SSPN), is designed by combining 3D Convolutional Neural Network (3D CNN) with feature pyramid structure. SSPN taking advantage of 3D convolution coupled with multi-scale convolutional extraction is used to obtain a large set of diverse spatial-spectral features. Multi-scale interfusion is also applied in SSPN to enrich the features contained in a single feature map and to improve the sensitivity on HSI spatial-spectral information, allowing it to better learn spatial-spectral features. Moreover, the losses of each combination based on multi-scale interfusion are calculated via weighted average, which enables SSPN to avoid the excessive influence of single combination in the updating of model parameters. Four HSI public datasets and several comparison models are employed to validate the classification effect of SSPN. Experimental results show that SSPN achieves the highest overall accuracy (OA) in all datasets compared with other classification models, with 100%, 98.8%, 99.8% and 98.7% on the datasets of Chikusei, Pavia University, Botswana and Houston 2013, respectively. SSPN is demonstrated to possess higher classification accuracy and better generalization performance on HSI. Junbo Zhou, Shan Zeng, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Double Auto-Weighted Tensor Robust Principal Component AnalysisabstractTensor Robust Principal Component Analysis (TRPCA), which aims to recover the low-rank and sparse components from their sum, has drawn intensive interest in recent years. Most existing TRPCA methods adopt the tensor nuclear norm (TNN) and the tensor ℓ1 norm as the regularization terms for the low-rank and sparse components, respectively. However, TNN treats each singular value of the low-rank tensor L equally and the tensor ℓ1 norm shrinks each entry of the sparse tensor S with the same strength. It has been shown that larger singular values generally correspond to prominent information of the data and should be less penalized. The same goes for large entries in S in terms of absolute values. In this paper, we propose a Double Auto-weighted TRPCA (DATRPCA) method. Instead of using predefined and manually set weights merely for the low-rank tensor as previous works, DATRPCA automatically and adaptively assigns smaller weights and applies lighter penalization to significant singular values of the low-rank tensor and large entries of the sparse tensorsimultaneously. We have further developed an efficient algorithm to implement DATRPCA based on the Alternating Direction Method of Multipliers (ADMM) framework. In addition, we have also established the convergence analysis of the proposed algorithm. The results on both synthetic and real-world data demonstrate the effectiveness of DATRPCA for low-rank tensor recovery, color image recovery and background modelling. Yulong Wang 0002, Kit Ian Kou, Hong Chen 0004, Yuan Yan Tang, Luoqing Li |
IEEE Trans. Image Process. | 4 |
| 2023 | Local Orthogonal Moments for Local FeaturesabstractBy introducing parameters with local information, several types of orthogonal moments have recently been developed for the extraction of local features in an image. But with the existing orthogonal moments, local features cannot be well-controlled with these parameters. The reason lies in that zeros distribution of these moments' basis function cannot be well-adjusted by the introduced parameters. To overcome this obstacle, a new framework, transformed orthogonal moment (TOM), is set up. Most existing continuous orthogonal moments, such as Zernike moments, fractional-order orthogonal moments (FOOMs), etc. are all special cases of TOM. To control the basis function's zeros distribution, a novel local constructor is designed, and local orthogonal moment (LOM) is proposed. Zeros distribution of LOM's basis function can be adjusted with parameters introduced by the designed local constructor. Consequently, locations, where local features extracted from by LOM, are more accurate than those by FOOMs. In comparison with Krawtchouk moments and Hahn moments etc., the range, where local features are extracted from by LOM, is order insensitive. Experimental results demonstrate that LOM can be utilized to extract local features in an image. Zezhi Zeng, Timothy C. H. Kwong, Yuan Yan Tang, Yuepeng Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2023 | Sliced Sparse Gradient Induced Multi-View Subspace Clustering via Tensorial Arctangent Rank MinimizationabstractMulti-view clustering method tries to improve the performance of clustering by using the information existing in different views. The tensorial representation is more suitable to capture the high order correlations across different views while keep local geometrical structure in specific view. In this paper, we propose a sliced sparse gradient induced multi-view subspace clustering method via tensorial arctangent rank minimization, named SSG-TAR method. Firstly, a tensorial arctangent rank (TAR) is defined, which is a tighter surrogate of the tensor rank and more effective to explore the consistency among multiple views. Secondly, a sliced sparse gradient regularization (SSG) is firstly proposed to enhance the discrimination between clusters and better capture the complementary information in view-specific feature space. Finally, we unify these two terms together and establish an efficient algorithm to optimize the proposed model. Furthermore, the constructed sequence was proved to converge to the stationary KKT point. We have carried out extensive experiments on ten datasets across different types and sizes to verify the performance of our model. The experimental results show that our method have achieved the state-of-the-art performance. Rui Zhu 0021, Ming Yang 0024, Yuan Yan Tang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | AlignVE: Visual Entailment Recognition Based on Alignment RelationsabstractVisual entailment (VE) is to recognize whether the semantics of a hypothesis text can be inferred from the given premise image, which is one special task among recent emerged vision and language understanding tasks. Currently, most of the existing VE approaches are derived from the methods of visual question answering. They recognize visual entailment by quantifying the similarity between the hypothesis and premise in the content semantic features from multi modalities. Such approaches, however, ignore the VE's unique nature of relation inference between the premise and hypothesis. Therefore, in this paper, a new architecture called AlignVE is proposed to solve the visual entailment problem with a relation interaction method. It models the relation between the premise and hypothesis as an alignment matrix. Then it introduces a pooling operation to get feature vectors with a fixed size. Finally, it goes through the fully-connected layer and normalization layer to complete the classification. Experiments show that our alignment-based architecture reaches 72.45% accuracy on SNLI-VE dataset, outperforming previous content-based models under the same settings. Biwei Cao, Jiuxin Cao, Jie Gui, Jiayun Shen, Bo Liu 0004, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Multim. | 7 |
| 2022 | Garbage Classification Detection Model Based on YOLOv4 with Lightweight Neural Network Feature Fusion
Xiaofeng Wang 0009, Jian-Tao Wang, Li-Xiang Xu, Jing Yang 0041, Yuan Yan Tang |
ICIC (3) | 6 |
| 2022 | Efficient residual attention network for single image super-resolution
Fangwei Hao, Taiping Zhang, Linchang Zhao, Yuan Yan Tang |
Appl. Intell. | 4 |
| 2022 | Gaussian process image classification based on multi-layer convolution kernel function
Lixiang Xu, Xinlu Li, Zhize Wu, Yan Chen 0037, Xiaofeng Wang 0009, Yuan Yan Tang |
Neurocomputing | 7 |
| 2022 | Multi-view unsupervised feature selection with tensor low-rank minimization
Junyu Li 0001, Yuan Yan Tang |
Neurocomputing | 4 |
| 2022 | LMSVCR: novel effective method of semi-supervised multi-classification
Zijie Dong, Yimo Qin, Bin Zou 0002, Jie Xu 0006, Yuan Yan Tang |
Neural Comput. Appl. | 5 |
| 2022 | Siamese networks with an online reweighted example for imbalanced data learning
Linchang Zhao, Zhaowei Shang, Mingliang Zhou 0001, Mu Zhang 0010, Dagang Gu, Taiping Zhang, Yuan Yan Tang |
Pattern Recognit. | 8 |
| 2022 | Generalized and Discriminative Collaborative Representation for Multiclass ClassificationabstractThis article presents a generalized collaborative representation-based classification (GCRC) framework, which includes many existing representation-based classification (RC) methods, such as collaborative RC (CRC) and sparse RC (SRC) as special cases. This article also advances the GCRC theory by exploring theoretical conditions on the general regularization matrix. A key drawback of CRC and SRC is that they fail to use the label information of training data and are essentially unsupervised in computing the representation vector. This largely compromises the discriminative ability of the learned representation vector and impedes the classification performance. Guided by the GCRC theory, we propose a novel RC method referred to as discriminative RC (DRC). The proposed DRC method has the following three desirable properties: 1) discriminability: DRC can leverage the label information of training data and is supervised in both representation and classification, thus improving the discriminative ability of the representation vector; 2) efficiency: it has a closed-form solution and is efficient in computing the representation vector and performing classification; and 3) theory: it also has theoretical guarantees for classification. Experimental results on benchmark databases demonstrate both the efficacy and efficiency of DRC for multiclass classification. Yulong Wang 0002, Yap-Peng Tan, Yuan Yan Tang, Hong Chen 0004, Cuiming Zou, Luoqing Li |
IEEE Trans. Cybern. | 3 |
| 2022 | Efficient Unsupervised Dimension Reduction for Streaming Multiview DataabstractMultiview learning has received substantial attention over the past decade due to its powerful capacity in integrating various types of information. Conventional unsupervised multiview dimension reduction (UMDR) methods are usually conducted in an offline manner and may fail in many real-world applications, where data arrive sequentially and the data distribution changes periodically. Moreover, satisfying the requirements of high memory consumption and expensive retraining of the time cost in large-scale scenarios are difficult. To remedy these drawbacks, we propose an online UMDR (OUMDR) framework. OUMDR aims to seek a low-dimensional and informative consensus representation for streaming multiview data. View-specific weights are also learned in this article to reflect the contributions of different views to the final consensus presentation. A specific model called OUMDR-E is developed by introducing the exclusive group LASSO (EG-LASSO) to explore the intraview and interview correlations. Then, we develop an efficient iterative algorithm with limited memory and time cost requirements for optimization, where the convergence of each update is theoretically guaranteed. We evaluate the proposed approach in video-based expression recognition applications. The experimental results demonstrate the superiority of our approach in terms of both effectiveness and efficiency. Weili Guo, Haikun Wei, Yuan Yan Tang, Dacheng Tao |
IEEE Trans. Cybern. | 4 |
| 2022 | An Efficient Cross-Modality Self-Calibrated Network for Hyperspectral and Multispectral Image FusionabstractRecently, deep convolutional neural network based hyperspectral and multispectral image fusion methods have shown significant performance. Nevertheless, the rich spatial and spectral details of hyperspectral images (HSIs) have not been fully explored, leaving room for further improve the representation ability of the model. In this paper, we propose an efficient cross-modality self-calibrated network (CMSCN) for hyperspectral and multispectral image fusion. Specifically, we use a cross-modality non-local module to fuse a high-resolution multispectral image (HR-MSI) and a low-resolution hyperspectral image (LR-HSI) to get an enhanced LR-HSI. In addition, a novel cross-scale self-calibrated convolution structure is proposed to explore and exploit multi-scale and hierarchical spatial-spectral features, which can improve the learning ability of the model. The introduced efficient spatial-spectral attention mechanism can calibrate the feature representation at different dimensions, thereby providing more efficient and accurate information for hyperspectral image reconstruction. Extensive experimental results on various hyperspectral images demonstrate the superiority of our method in comparison with the state-of-the-art image fusion methods. Huapeng Wu, Jie Gui, Yang Xu 0006, Zebin Wu 0001, Yuan Yan Tang, Zhihui Wei |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Weighted Error Entropy-Based Information Theoretic Learning for Robust Subspace RepresentationabstractIn most of the existing representation learning frameworks, the noise contaminating the data points is often assumed to be independent and identically distributed (i.i.d.), where the Gaussian distribution is often imposed. This assumption, though greatly simplifies the resulting representation problems, may not hold in many practical scenarios. For example, the noise in face representation is usually attributable to local variation, random occlusion, and unconstrained illumination, which is essentially structural, and hence, does not satisfy the i.i.d. property or the Gaussianity. In this article, we devise a generic noise model, referred to as independent and piecewise identically distributed (i.p.i.d.) model for robust presentation learning, where the statistical behavior of the underlying noise is characterized using a union of distributions. We demonstrate that our proposed i.p.i.d. model can better describe the complex noise encountered in practical scenarios and accommodate the traditional i.i.d. one as a special case. Assisted by the proposed noise model, we then develop a new information-theoretic learning framework for robust subspace representation through a novel minimum weighted error entropy criterion. Thanks to the superior modeling capability of the i.p.i.d. model, our proposed learning method achieves superior robustness against various types of noise. When applying our scheme to the subspace clustering and image recognition problems, we observe significant performance gains over the existing approaches. Yuanman Li, Jiantao Zhou 0001, Jinyu Tian 0001, Xianwei Zheng, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Effective Multiplatform Advertising PolicyabstractMultiplatform advertising (MPA) is recognized as an effective means of enhancing marketing revenue. In the context, we refer to the scheme of dynamically allocating the advertising expenditure among the selected media platforms as an MPA policy, and we refer to the problem of developing an MPA policy with maximum benefit as the MPA problem. This article is devoted to the solution of the MPA problem. An evolutionary model for the expected market state, in which the influence of both advertising and word-of-mouth (WOM) propagation is accounted for, is established. On this basis, the expected benefit of an MPA policy is calculated. Thereby, the MPA problem is reduced to an optimal control problem we refer to as the MPA model, where the objective functional stands for the expected benefit of an MPA strategy. The optimality system for the MPA model is derived. We refer to the MPA policy obtained by solving the optimality system as the promising MPA policy. The structure of the promising MPA policy is inspected. Through extensive comparative experiments, it is concluded that the promising MPA policy is superior to the majority of MPA policies in terms of expected benefit. Finally, how the expected benefit of the promising MPA policy is influenced by some factors is investigated. Kaifan Huang, Lu-Xing Yang, Xiaofan Yang 0001, Yuan Yan Tang |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | Channel Hourglass Residual Network For Single Image Super-ResolutionabstractDeep convolutional neural networks (CNNs) for Super-Resolution (SR) from low-resolution (LR) images have achieved remarkable reconstruction performance with the utilization of residual networks and visual attention mechanism. However, the existing single image super-resolution (SISR) methods with deeper or wider network architectures encounter module representation bottleneck and neglect module efficiency in real-world applications. To solve these issues, in this paper, we design channel hourglass residual structure (CHRS) consisted of several nested residual modules for reducing parameters and extracting more representational features. Furthermore, we integrate channel attention (CA) mechanism into CHRS to generate channel hourglass residual block (CHRB) which can be easily extended to other methods for improving performance. We also propose channel hourglass residual network (CHRN) which not only pays attention to network learning efficiency but also learns more discriminative expressions. Extensive experiments demonstrate the effectiveness of our CHRN and the generalization ability of our CHRB. Fangwei Hao, XinDi Ma, Taiping Zhang, Yuan Yan Tang |
IJCNN | 4 |
| 2021 | Indian Buffet Process-Based on Nonnegative Matrix Factorization with Single Binary ComponentabstractNonnegative matrix factorization seeks to find a basic matrix and a weight matrix to approximate the nonnegative matrix. It has proven to be a powerful low-rank decomposition technique for nonnegative multivariate data. However, its performance largely depends on the assumption of a fixed number of features. In this work, we propose a new probabilistic nonnegative matrix factorization which factorizes a nonnegative matrix into a low-rank factor matrix with {0,1} constraints and a nonnegative weight matrix. In order to automatically learn the potential binary features and feature number. A deterministic Indian buffet process variational inference is introduced to obtain the binary factor matrix. And the weight matrix is set to satisfy the exponential prior. In order to obtain the real posterior distribution of the two factor matrices, a variational Bayesian exponential Gaussian inference model is established. The comparative experiments on both the synthetic and real-world data sets show the efficacy of the proposed method. XinDi Ma, Taiping Zhang, Yuan Yan Tang |
SMC | 5 |
| 2021 | Distribution Preserving Deep Semi-Nonnegative Matrix FactorizationabstractDeep semi-nonnegative matrix factorization can obtain the hidden hierarchical representations according to the unknown attributes of the given data. On the other hand, the inherent structure of the each data cluster can be described by the distribution of the intra-class data. Then one hopes to learn a new low dimensional representation which can preserve the intrinsic structure embedded in the original high dimensional data space perfectly. Here we propose a novel distribution preserving deep semi-nonnegative matrix factorization method (DPNMF) to achieve this goal. As a result, the manifold structures in the raw data are well preserved in the feature space being from the top layer. The experimental results on the real-world datasets show that the proposed algorithm has good performance in terms of cluster accuracy and normalized mutual information (NMI). Zhuolin Tan, Anyong Qin, Yongqing Sun, Yuan Yan Tang |
SMC | 4 |
| 2021 | Ultrarobust support vector registration
Yuyi Wang 0001, Bin Zou 0002, Yuan Yan Tang |
Appl. Intell. | 5 |
| 2021 | OAA-SVM-MS: A fast and efficient multi-class classification algorithm
Yuze Duan, Bin Zou 0002, Jie Xu 0006, Jiaolong Wei, Yuan Yan Tang |
Neurocomputing | 6 |
| 2021 | A diversified shared latent variable model for efficient image characteristics extraction and modelling
Hao Xiong 0001, Yuan Yan Tang, Fionn Murtagh, Leszek Rutkowski, Shlomo Berkovsky |
Neurocomputing | 2 |
| 2021 | Smart Home Privacy Protection Based on the Improved LSB Information HidingabstractSmart home is an emerging form of the Internet of Things (IoT), enabling people to enjoy a convenient and intelligent life. The data generated by smart home devices are transmitted through the public channel, which is not secure enough, so the secret data in smart home are easily intercepted by malicious adversaries. In order to solve this problem, this paper proposes a smart home privacy protection method combining DES encryption and the improved Least Significant Bit (LSB) information hiding algorithm, changing the practice of directly exposing smart home secret information to the Internet, first, using Data Encryption Standard (DES) encryption to encrypt the smart home information and second, the improved LSB information hiding algorithm is used to hide the ciphertext, so that the adversary cannot detect the smart home secret information. The goal of the scheme is to provide a double protection for the secure transmission of the smart home secret information. If an attacker wants to carry out an attack, it has to break through at least two defense lines, which seems impossible to do. Experiment results show that the improved LSB algorithm is more robust than the existing algorithms, and it is very safe. Therefore, the scheme proposed in this paper is very practical for protecting the smart home secret information. Haiyu Deng, Ren Ping Liu 0001, Patrick Shen-Pei Wang, Xiaocui Dang, Yuan Yan Tang, Xichun Li |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2021 | A robust image representation method against illumination and occlusion variations
Taiping Zhang, Linchang Zhao, Xiaoliu Luo, Yuan Yan Tang |
Image Vis. Comput. | 5 |
| 2021 | Semi-supervised multi-Layer convolution kernel learning in credit evaluation
Lixiang Xu, Lixin Cui, Thomas Weise 0001, Xinlu Li, Zhize Wu, Feiping Nie 0001, Enhong Chen, Yuan Yan Tang |
Pattern Recognit. | 8 |
| 2021 | Quaternion block sparse representation for signal recovery and classification
Cuiming Zou, Kit Ian Kou, Yulong Wang 0002, Yuan Yan Tang |
Signal Process. | 4 |
| 2021 | Multi-focus image fusion with Geometrical Sparse Representation
Taiping Zhang, Linchang Zhao, Xiaoliu Luo, Yuan Yan Tang |
Signal Process. Image Commun. | 5 |
| 2021 | Self-Adaptive Multiprototype-Based Competitive Learning Approach: A k-Means-Type Algorithm for Imbalanced Data ClusteringabstractClass imbalance problem has been extensively studied in the recent years, but imbalanced data clustering in unsupervised environment, that is, the number of samples among clusters is imbalanced, has yet to be well studied. This paper, therefore, studies the imbalanced data clustering problem within the framework of k -means-type competitive learning. We introduce a new method called self-adaptive multiprototype-based competitive learning (SMCL) for imbalanced clusters. It uses multiple subclusters to represent each cluster with an automatic adjustment of the number of subclusters. Then, the subclusters are merged into the final clusters based on a novel separation measure. We also propose a new internal clustering validation measure to determine the number of final clusters during the merging process for imbalanced clusters. The advantages of SMCL are threefold: 1) it inherits the advantages of competitive learning and meanwhile is applicable to the imbalanced data clustering; 2) the self-adaptive multiprototype mechanism uses a proper number of subclusters to represent each cluster with any arbitrary shape; and 3) it automatically determines the number of clusters for imbalanced clusters. SMCL is compared with the existing counterparts for imbalanced clustering on the synthetic and real datasets. The experimental results show the efficacy of SMCL for imbalanced clusters. Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang |
IEEE Trans. Cybern. | 3 |
| 2021 | Robust Sparse Representation in Quaternion SpaceabstractSparse representation has achieved great success across various fields including signal processing, machine learning and computer vision. However, most existing sparse representation methods are confined to the real valued data. This largely limit their applicability to the quaternion valued data, which has been widely used in numerous applications such as color image processing. Another critical issue is that their performance may be severely hampered due to the data noise or outliers in practice. To tackle the problems above, in this work we propose a robust quaternion valued sparse representation (RQVSR) method in a fully quaternion valued setting. To handle the quaternion noises, we first define a new robust estimator referred as quaternion Welsch estimator to measure the quaternion residual error. Compared to the conventional quaternion mean square error, it can largely suppress the impact of large data corruption and outliers. To implement RQVSR, we have overcome the difficulties raised by the noncommutativity of quaternion multiplication and developed an effective algorithm by leveraging the half-quadratic theory and the alternating direction method of multipliers framework. The experimental results show the effectiveness and robustness of the proposed method for quaternion sparse signal recovery and color image reconstruction. Yulong Wang 0002, Kit Ian Kou, Cuiming Zou, Yuan Yan Tang |
IEEE Trans. Image Process. | 4 |
| 2021 | A Contour Co-Tracking Method for Image PairsabstractWe proposed a contour co-tracking method for co-segmentation of image pairs based on active contour model. Our method comprehensively re-models objects and backgrounds signified by level set functions, and leverages Hellinger distance to measure the similarity between image regions encoded by probability distributions. The main contribution are as follows. 1) The new energy functional, combining a rewarding and a penalty term, relaxes the assumptions of co-segmentation methods. 2) Hellinger distance, fulfilling the triangle inequality, ensures a coherence measurement between probability distributions in metric space, and contributes to finding a unique solution to the energy functional. The proposed contour co-tracking method was carefully verified against five representative methods on four popular datasets, i.e., the images pair dataset (105 pairs), MSRC dataset (30 pairs), iCoseg dataset (66 pairs) and Coseg-rep dataset (25 pairs). The comparison experiments suggest that our method achieves the competitive and even better performance compared to the state-of-the-art co-segmentation methods. Bin Wang 0027, Dapeng Tao, Yuan Yan Tang, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Learning to Hash With Dimension Analysis Based Quantizer for Image RetrievalabstractThe last few years have witnessed the rise of the big data era in which approximate nearest neighbor search is a fundamental problem in many applications, such as large-scale image retrieval. Recently, many research results have demonstrated that hashing can achieve promising performance due to its appealing storage and search efficiency. Since complex optimization problems for loss functions are difficult to solve, most hashing methods decompose the hash code learning problem into two steps: projection and quantization. In the quantization step, binary codes are widely used because ranking them by the Hamming distance is very efficient. However, the massive information loss produced by the quantization step should be reduced in applications where high search accuracy is required, such as in image retrieval. Since many two-step hashing methods produce uneven projected dimensions in the projection step, in this paper, we propose a novel dimension analysis-based quantization (DAQ) on two-step hashing methods for image retrieval. We first perform an importance analysis of the projected dimensions and select a subset of them that are more informative than others, and then we divide the selected projected dimensions into several regions with our quantizer. Every region is quantized with its corresponding codebook. Finally, the similarity between two hash codes is estimated by the Manhattan distance between their corresponding codebooks, which is also efficient. We conduct experiments on three public benchmarks containing up to one million descriptors and show that the proposed DAQ method consistently leads to significant accuracy improvements over state-of-the-art quantization methods. Yuan Cao 0005, Heng Qi, Jie Gui, Keqiu Li, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Multim. | 5 |
| 2020 | Adaptive parameter estimation of GMM and its application in clustering
Linchang Zhao, Zhaowei Shang, Xiaoliu Luo, Taiping Zhang, Yuan Yan Tang |
Future Gener. Comput. Syst. | 7 |
| 2020 | Reliable asynchronous sampled-data filtering of T-S fuzzy uncertain delayed neural networks with stochastic switched topologies
Kaibo Shi, Jun Wang 0128, Yuan Yan Tang, Shouming Zhong |
Fuzzy Sets Syst. | 3 |
| 2020 | Non-fragile memory filtering of T-S fuzzy delayed neural networks based on switched fuzzy sampled-data control
Kaibo Shi, Jun Wang 0128, Shouming Zhong, Yuan Yan Tang, Jun Cheng 0004 |
Fuzzy Sets Syst. | 4 |
| 2020 | Hybrid-driven finite-time H∞ sampling synchronization control for coupling memory complex networks with stochastic cyber attacks
Kaibo Shi, Jun Wang 0128, Shouming Zhong, Yuan Yan Tang, Jun Cheng 0004 |
Neurocomputing | 4 |
| 2020 | Modal regression based greedy algorithm for robust sparse signal recovery, clustering and classification
Yulong Wang 0002, Yuan Yan Tang, Cuiming Zou, Luoqing Li, Hong Chen 0004 |
Neurocomputing | 2 |
| 2020 | Low-rank matrix regression for image feature extraction and feature selection
Junyu Li 0001, Loi Lei Lai, Yuan Yan Tang |
Inf. Sci. | 4 |
| 2020 | SVM-Boosting based on Markov resampling: Theory and algorithm
Bin Zou 0002, Chen Xu 0007, Jie Xu 0006, Yuan Yan Tang |
Neural Networks | 5 |
| 2020 | Modal Regression-Based Atomic Representation for Robust Face Recognition and ReconstructionabstractRepresentation-based classification (RC) methods, such as sparse RC, have shown great potential in face recognition (FR) in recent years. Most previous RC methods are based on the conventional regression models, such as lasso regression, ridge regression, or group lasso regression. These regression models essentially impose a predefined assumption on the distribution of the noise variable in the query sample, such as the Gaussian or Laplacian distribution. However, the complicated noises in practice may violate the assumptions and impede the performance of these RC methods. In this paper, we propose a modal regression (MR)-based atomic representation and classification (MRARC) framework to alleviate such limitations. MR is a robust regression framework which aims to reveal the relationship between the input and response variables by regressing toward the conditional mode function. Atomic representation is a general atomic norm regularized linear representation framework which includes many popular representation methods, such as sparse representation, collaborative representation, and low-rank representation as special cases. Unlike previous RC methods, the MRARC framework does not require the noise variable to follow any specific predefined distributions. This gives rise to the capability of MRARC in handling various complex noises in reality. Using MRARC as a general platform, we also develop four novel RC methods for unimodal and multimodal FR, respectively. In addition, we devise a general optimization algorithm for the unified MRARC framework based on the alternating direction method of multipliers and half-quadratic theory. The experiments on real-world data validate the efficacy of MRARC for robust FR and reconstruction. Yulong Wang 0002, Yuan Yan Tang, Luoqing Li, Hong Chen 0004 |
IEEE Trans. Cybern. | 2 |
| 2020 | A Risk Management Approach to Defending Against the Advanced Persistent ThreatabstractThe advanced persistent threat (APT) as a new kind of cyber attack has posed a severe threat to modern organizations. When the APT has been detected, the organization has to deal with the APT response problem, i.e., to allocate the available response resources to fix her insecure hosts so as to mitigate her potential loss. This paper addresses the APT response problem by using the risk management approach. First, we introduce a model characterizing the evolution of the organization's expected state. By analyzing this model, we find the organization's expected state approaches a common limit expected state. Then, we use the organization's expected loss per unit time to measure her potential loss, and we find this measure is determined by the organization's limit expected state. On this basis, we model the APT response problem as a game-theoretic problem (the APT response game) in which the organization seeks a Nash equilibrium. We present a greedy algorithm for solving the game. Comparative experiments show that the algorithm is effective. Therefore, we recommend the response strategy generated by performing the algorithm. These findings contribute to defending against the APT. To our knowledge, this is the first time the APT response problem is addressed. Lu-Xing Yang, Pengdeng Li, Xiaofan Yang 0001, Yuan Yan Tang |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2020 | Quasi Fourier-Mellin Transform for Affine Invariant FeaturesabstractFourier-Mellin transform (FMT) has been widely used for the extraction of rotation- and scale-invariant features. However, affine transform is a more reasonable approximation model for real viewpoint change. Due to shearing, the integral along the angular direction in the calculation of FMT cannot be used to extract the inherent features of an image undergoing affine transform. To eliminate the effect of shearing, whitening transform should be conducted on the integral along the radial direction. FMT can hardly be modified by conventional whitening-based methods with low computational cost due to additional processes. In this paper, two factors are constructed and embedded into FMT. Quasi Fourier-Mellin transform (QFMT) is proposed. The embedding of these factors is equivalent to whitening transform and can eliminate the effect of shearing in the affine transform. In particular, QFMT can also be calculated by integrating along the radial direction followed by integrating along the angular direction, as in FMT. Based on QFMT, the quasi Fourier-Mellin descriptor (QFMD) is constructed for the extraction of affine invariant features. Some experiments have also been conducted to test the performance of the proposed method. Zhengda Lu, Yuan Yan Tang, Zhou Yuan |
IEEE Trans. Image Process. | 3 |
| 2020 | Adaptive Chunk-Based Dynamic Weighted Majority for Imbalanced Data Streams With Concept DriftabstractOne of the most challenging problems in the field of online learning is concept drift, which deeply influences the classification stability of streaming data. If the data stream is imbalanced, it is even more difficult to detect concept drifts and make an online learner adapt to them. Ensemble algorithms have been found effective for the classification of streaming data with concept drift, whereby an individual classifier is built for each incoming data chunk and its associated weight is adjusted to manage the drift. However, it is difficult to adjust the weights to achieve a balance between the stability and adaptability of the ensemble classifiers. In addition, when the data stream is imbalanced, the use of a size-fixed chunk to build a single classifier can create further problems; the data chunk may contain too few or even no minority class samples (i.e., only majority class samples). A classifier built on such a chunk is unstable in the ensemble. In this article, we propose a chunk-based incremental learning method called adaptive chunk-based dynamic weighted majority (ACDWM) to deal with imbalanced streaming data containing concept drift. ACDWM utilizes an ensemble framework by dynamically weighting the individual classifiers according to their classification performance on the current data chunk. The chunk size is adaptively selected by statistical hypothesis tests to access whether the classifier built on the current data chunk is sufficiently stable. ACDWM has four advantages compared with the existing methods as follows: 1) it can maintain stability when processing nondrifted streams and rapidly adapt to the new concept; 2) it is entirely incremental, i.e., no previous data need to be stored; 3) it stores a limited number of classifiers to ensure high efficiency; and 4) it adaptively selects the chunk size in the concept drift environment. Experiments on both synthetic and real data sets containing concept drift show that ACDWM outperforms both state-of-the-art chunk-based and online methods. Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Bayes Imbalance Impact Index: A Measure of Class Imbalanced Data Set for Classification ProblemabstractRecent studies of imbalanced data classification have shown that the imbalance ratio (IR) is not the only cause of performance loss in a classifier, as other data factors, such as small disjuncts, noise, and overlapping, can also make the problem difficult. The relationship between the IR and other data factors has been demonstrated, but to the best of our knowledge, there is no measurement of the extent to which class imbalance influences the classification performance of imbalanced data. In addition, it is also unknown which data factor serves as the main barrier for classification in a data set. In this article, we focus on the Bayes optimal classifier and examine the influence of class imbalance from a theoretical perspective. We propose an instance measure called the Individual Bayes Imbalance Impact Index (IBI3) and a data measure called the Bayes Imbalance Impact Index (BI3). IBI3and BI3reflect the extent of influence using only the imbalance factor, in terms of each minority class sample and the whole data set, respectively. Therefore, IBI3can be used as an instance complexity measure of imbalance and BI3as a criterion to demonstrate the degree to which imbalance deteriorates the classification of a data set. We can, therefore, use BI3to access whether it is worth using imbalance recovery methods, such as sampling or cost-sensitive methods, to recover the performance loss of a classifier. The experiments show that IBI3is highly consistent with the increase of the prediction score obtained by the imbalance recovery methods and that BI3is highly consistent with the improvement in the F1 score obtained by the imbalance recovery methods on both synthetic and real benchmark data sets. Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Sparse Supervised Representation-Based Classifier for Uncontrolled and Imbalanced ClassificationabstractThe sparse representation-based classification (SRC) has been utilized in many applications and is an effective algorithm in machine learning. However, the performance of SRC highly depends on the data distribution. Some existing works proved that SRC could not obtain satisfactory results on uncontrolled data sets. Except the uncontrolled data sets, SRC cannot deal with imbalanced classification either. In this paper, we proposed a model named sparse supervised representation classifier (SSRC) to solve the above-mentioned issues. The SSRC involves the class label information during the test sample representation phase to deal with the uncontrolled data sets. In SSRC, each class has the opportunity to linearly represent the test sample in its subspace, which can decrease the influences of the uncontrolled data distribution. In order to classify imbalanced data sets, a class weight learning model is proposed and added to SSRC. Each class weight is learned from its corresponding training samples. The experimental results based on the AR face database (uncontrolled) and 15 KEEL data sets (imbalanced) with an imbalanced rate ranging from 1.48 to 61.18 prove SSRC can effectively classify uncontrolled and imbalanced data sets. Ting Shu 0001, Bob Zhang 0001, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Robust Subspace Clustering With Independent and Piecewise Identically Distributed Noise ModelingabstractMost of the existing subspace clustering (SC) frameworks assume that the noise contaminating the data is generated by an independent and identically distributed (i.i.d.) source, where the Gaussianity is often imposed. Though these assumptions greatly simplify the underlying problems, they do not hold in many real-world applications. For instance, in face clustering, the noise is usually caused by random occlusions, local variations and unconstrained illuminations, which is essentially structural and hence satisfies neither the i.i.d. property nor the Gaussianity. In this work, we propose an independent and piecewise identically distributed (i.p.i.d.) noise model, where the i.i.d. property only holds locally. We demonstrate that the i.p.i.d. model better characterizes the noise encountered in practical scenarios, and accommodates the traditional i.i.d. model as a special case. Assisted by this generalized noise model, we design an information theoretic learning (ITL) framework for robust SC through a novel minimum weighted error entropy (MWEE) criterion. Extensive experimental results show that our proposed SC scheme significantly outperforms the state-of-the-art competing algorithms. Yuanman Li, Jiantao Zhou 0001, Xianwei Zheng, Jinyu Tian 0001, Yuan Yan Tang |
CVPR | 5 |
| 2019 | Distribution Preserving Network EmbeddingabstractThe deep autoencoder network which is based on constraining non-negative weights, can learn a low dimensional part-based representation. On the other hand, the inherent structure of the each data cluster can be described by the distribution of the intraclass sample. Then one hopes to learn a new low dimensional feature which can preserve the intrinsic structure embedded in the high dimensional data space perfectly. In this paper, by preserving data distribution, a deep part-based representation can be learned, and the novel algorithm is called Distribution Preserving Network Embedding (DPNE). In DPNE, we first need to estimate the distribution of the original data, and then we seek a part-based representation which respects the distribution. The experimental results on real-world data sets show that the proposed algorithm has good performance in terms of cluster accuracy and adjusted mutual information (AMI). Anyong Qin, Zhaowei Shang, Taiping Zhang, Yuan Yan Tang |
ICASSP | 4 |
| 2019 | A cost-sensitive meta-learning classifier: SPFCNN-Miner
Linchang Zhao, Zhaowei Shang, Anyong Qin, Taiping Zhang, Yuan Yan Tang |
Future Gener. Comput. Syst. | 7 |
| 2019 | Software defect prediction via cost-sensitive Siamese parallel fully-connected neural networks
Linchang Zhao, Zhaowei Shang, Taiping Zhang, Yuan Yan Tang |
Neurocomputing | 5 |
| 2019 | Spectral-Spatial Sparse Subspace Clustering Based on Three-Dimensional Edge-Preserving Filtering for Hyperspectral ImageabstractIntegrating spatial information into the sparse subspace clustering (SSC) models for hyperspectral images (HSIs) is an effective way to improve clustering accuracy. Since HSI is a three-dimensional (3D) cube datum, 3D spectral-spatial filtering becomes a simple method for extracting the spectral-spatial information. In this paper, a novel spectral-spatial SSC framework based on 3D edge-preserving filtering (EPF) is proposed to improve the clustering accuracy of HSI. First, the initial sparse coefficient matrix is obtained in the sparse representation process of the classical SSC model. Then, a 3D EPF is conducted on the initial sparse coefficient matrix to obtain a more accurate coefficient matrix by solving an optimization problem based on ADMM, which is used to build the similarity graph. Finally, the clustering result of HSI data is achieved by applying the spectral clustering algorithm to the similarity graph. Specifically, the filtered matrix can not only capture the spectral-spatial information but the intensity differences. The experimental results on three real-world HSI datasets demonstrated that the potential of including the proposed 3D EPF into the SSC framework can improve the clustering accuracy. Ailin Li, Anyong Qin, Zhaowei Shang, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2019 | A Fractal Dimension and Empirical Mode Decomposition-Based Method for Protein Sequence AnalysisabstractIn bioinformatics, the biological functions of proteins and their interactions can often be analyzed by the similarity of their sequences. In this paper, the authors combine the fractal dimension, empirical mode decomposition (EMD), and sliding window for protein sequence comparison. First, the protein sequence is characterized and digitized into a signal, and then the signal characteristics are obtained by using EMD and fractal dimension. Each protein sequence can be decomposed into Intrinsic Mode Functions (IMFs). The fixed window’s fractal dimension is applied to each IMF and the original signal to extract the protein sequence characteristics. Experiments have shown that the feature extracted by this hybrid method is superior to the EMD method alone. Pu Wei, Zuqiang Meng, Patrick Shen-Pei Wang, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2019 | Multi-Level Downsampling of Graph Signals via Improved Maximum Spanning TreesabstractGraph signal processing (GSP) is an emerging field in the signal processing community. Novel GSP-based transforms, such as graph Fourier transform and graph wavelet filter banks, have been successfully utilized in image processing and pattern recognition. As a rapidly developing research area, graph signal processing aims to extend classical signal processing techniques to signals with irregular underlying structures. One of the hot topics in GSP is to develop multi-scale transforms such that novel GSP-based techniques can be applied in image processing or other related areas. For designing graph signal multi-scale frameworks, downsampling operations that ensuring multi-level downsampling should be specifically constructed. Among the existing downsampling methods in graph signal processing, the state-of-the-art method was constructed based on the maximum spanning tree (MST). However, when using this method for multi-level downsampling of graph signals defined on unweighted densely connected graphs, such as social network data, the sampling rates are not close to [Formula: see text]. This phenomenon is summarized as a new problem and called downsampling unbalance problem in this paper. Due to the unbalance, MST-based downsampling method cannot be applied to construct graph signal multi-scale transforms. In this paper, we propose a novel and efficient method to detect and reduce the downsampling unbalance generated by the MST-based method. For any given graph signal, we apply the graph density to construct a measurement of the downsampling unbalance generated by the MST-based method. If a graph signal has large unbalance possibility, the multi-level downsampling is conducted after the MST is improved. The experimental results on synthetic and real-world social network data show that downsampling unbalance can be efficiently detected and then reduced by our method. Xianwei Zheng, Yuan Yan Tang, Jiantao Zhou 0001, Jianjia Pan, Shouzhi Yang, Youfa Li, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2019 | Spectral-Spatial Graph Convolutional Networks for Semisupervised Hyperspectral Image ClassificationabstractCollecting labeled samples is quite costly and time-consuming for hyperspectral image (HSI) classification task. Semisupervised learning framework, which combines the intrinsic information of labeled and unlabeled samples, can alleviate the deficient labeled samples and increase the accuracy of HSI classification. In this letter, we propose a novel semisupervised learning framework that is based on spectral-spatial graph convolutional networks (S2GCNs). It explicitly utilizes the adjacency nodes in graph to approximate the convolution. In the process of approximate convolution on graph, the proposed method makes full use of the spatial information of the current pixel. The experimental results on three real-life HSI data sets, i.e., Botswana Hyperion, Kennedy Space Center, and Indian Pines, show that the proposed S2GCN can significantly improve the classification accuracy. For instance, the overall accuracy on Indian data is increased from 66.8% (GCN) to 91.6%. Anyong Qin, Zhaowei Shang, Jinyu Tian 0001, Yulong Wang 0002, Taiping Zhang, Yuan Yan Tang |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2019 | Atomic Representation-Based Classification: Theory, Algorithm, and ApplicationsabstractRepresentation-based classification (RC) methods such as sparse RC (SRC) have attracted great interest in pattern recognition recently. Despite their empirical success, few theoretical results are reported to justify their effectiveness. In this paper, we establish the theoretical guarantees for a general unified framework termed as atomic representation-based classification (ARC), which includes most RC methods as special cases. We introduce a new condition called atomic classification condition (ACC), which reveals important geometric insights for the theory of ARC. We show that under such condition ARC is provably effective in correctly recognizing any new test sample, even corrupted with noise. Our theoretical analysis significantly broadens the range of conditions under which RC methods succeed for classification in the following two aspects: (1) prior theoretical advances of RC are mainly concerned with the single SRC method while our theory can apply to the general unified ARC framework, including SRC and many other RC methods; and (2) previous works are confined to the analysis of noiseless test data while we provide theoretical guarantees for ARC using both noiseless and noisy test data. Numerical results are provided to validate and complement our theoretical analysis of ARC and its important special cases for both noiseless and noisy test data. Yulong Wang 0002, Yuan Yan Tang, Luoqing Li, Hong Chen 0004, Jianjia Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Block sparse representation for pattern classification: Theory, extensions and applications
Yulong Wang 0002, Yuan Yan Tang, Luoqing Li, Xianwei Zheng |
Pattern Recognit. | 2 |
| 2019 | Mellin polar coordinate moment and its affine invariance
Yuan Yan Tang |
Pattern Recognit. | 3 |
| 2019 | Joint sparse matrix regression and nonnegative spectral analysis for two-dimensional unsupervised feature selection
Junyu Li 0001, Loi Lei Lai, Yuan Yan Tang |
Pattern Recognit. | 4 |
| 2019 | Seeking Best-Balanced Patch-Injecting Strategies through Optimal Control ApproachabstractTo restrain escalating computer viruses, new virus patches must be constantly injected into networks. In this scenario, the patch-developing cost should be balanced against the negative impact of virus. This article focuses on seeking best-balanced patch-injecting strategies. First, based on a novel virus-patch interactive model, the original problem is reduced to an optimal control problem, in which (a) each admissible control stands for a feasible patch-injecting strategy and (b) the objective functional measures the balance of a feasible patch-injecting strategy. Second, the solvability of the optimal control problem is proved, and the optimality system for solving the problem is derived. Next, a few best-balanced patch-injecting strategies are presented by solving the corresponding optimality systems. Finally, the effects of some factors on the best balance of a patch-injecting strategy are examined. Our results will be helpful in defending against virus attacks in a cost-effective way. Kaifan Huang, Pengdeng Li, Lu-Xing Yang, Xiaofan Yang 0001, Yuan Yan Tang |
Secur. Commun. Networks | 5 |
| 2019 | Cauchy greedy algorithm for robust sparse recovery and multiclass classification
Yulong Wang 0002, Cuiming Zou, Yuan Yan Tang, Luoqing Li, Zhaowei Shang |
Signal Process. | 3 |
| 2019 | Hyperspectral Unmixing via Total Variation Regularized Nonnegative Tensor FactorizationabstractHyperspectral unmixing decomposes a hyperspectral imagery (HSI) into a number of constituent materials and associated proportions. Recently, nonnegative tensor factorization (NTF)-based methods have been proposed for hyperspectral unmixing thanks to their capability in representing an HSI without any information loss. However, tensor factorization-based HSI processing approaches often suffer from low-signal-to-noise ratio condition of HSI and nonuniqueness of the solution. This problem can be effectively alleviated by introducing various spatial constraints into tensor factorization to suppress the noise and decrease the number of extreme, stationary, and saddle points. On the other hand, total variation (TV) adaptively promotes piecewise smoothness while preserving edges. In this paper, we propose a TV regularized matrix-vector NTF method. It takes advantage of tensor factorization in preserving global spectral-spatial information and the merits of TV in exploiting local spatial information, thus generating smooth abundance maps with preserved edges. Experimental results on synthetic and real-world data show that the proposed method outperforms the state-of-the-art methods. Fengchao Xiong, Yuntao Qian, Jun Zhou 0001, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | A Novel Rank Approximation Method for Mixture Noise Removal of Hyperspectral ImagesabstractMixture noise removal is a fundamental problem in hyperspectral images' (HSIs) processing that holds significant practical importance for subsequent applications. This problem can be recast as an approximation issue of a low-rank matrix. In this paper, a novel smooth rank approximation (SRA) model is proposed to cope with these mixture noises for HSIs. The crux idea is to devise a general smooth function under some assumptions to directly approximate the rank function, which attempts to explore a closer approximation than conventional methods. This new optimization model can be easily solved by the convex analysis tool and can remove the mixture noises of HSIs quickly and effectively. Subsequently, we give a feasible iterative algorithm, and the corresponding convergence analysis is discussed mathematically. Experimental results from the simulated data set as well as real data sets illustrate that the proposed SRA method significantly outperforms the state-of-the-art methods on HSI denoising. Hailiang Ye, Hong Li 0009, Feilong Cao, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Content-Adaptive Noise Estimation for Color Images With Cross-Channel Noise ModelingabstractNoise estimation is crucial in many image processing tasks such as denoising. Most of the existing noise estimation methods are specially developed for grayscale images. For color images, these methods simply handle each color channel independently, without considering the correlation across channels. Moreover, these methods often assume a globally fixed noise model throughout the entire image, neglecting the adaptation to the local structures. In this work, we propose a contentadaptive multivariate Gaussian approach to model the noise in color images, in which we explicitly consider both the contentdependence and the inter-dependence among color channels. We design an effective method for estimating the noise covariance matrices within the proposed model. Specifically, a patch selection scheme is first introduced to select weakly textured patches via thresholding the texture strength indicators. Noticing that the patch selection actually depends on the unknown noise covariance, we present an iterative noise covariance estimation algorithm, where the patch selection and the covariance estimation are conducted alternately. For the remaining textured regions, we estimate a distinct covariance matrix associated with each pixel using a linear shrinkage estimator, which adaptively fuses the estimate coming from the weakly textured region and the sample covariance estimated from the local region. Experimental results show that our method can effectively estimate the noise covariance. The usefulness of our method is demonstrated with several image processing applications such as color image denoising and noise-robust superpixel. Li Dong 0006, Jiantao Zhou 0001, Yuan Yan Tang |
IEEE Trans. Image Process. | 3 |
| 2019 | Trajectory Data Classification: A ReviewabstractThis article comprehensively surveys the development of trajectory data classification. Considering the critical role of trajectory data classification in modern intelligent systems for surveillance security, abnormal behavior detection, crowd behavior analysis, and traffic control, trajectory data classification has attracted growing attention. According to the availability of manual labels, which is critical to the classification performances, the methods can be classified into three categories, i.e., unsupervised, semi-supervised, and supervised. Furthermore, classification methods are divided into some sub-categories according to what extracted features are used. We provide a holistic understanding and deep insight into three types of trajectory data classification methods and present some promising future directions. Jiang Bian 0006, Dayong Tian, Yuan Yan Tang, Dacheng Tao |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2019 | Geometry and Topology Preserving Hashing for SIFT FeatureabstractIn recent years, content-based image retrieval has been of concern because of practical needs on Internet services, especially methods that can improve retrieving speed and accuracy. The SIFT feature is a well-designed local feature. It has mature applications in feature matching and retrieval, whereas the raw SIFT feature is high dimensional, with high storage cost as well as computational cost in feature similarity measurements. Thus, we propose a hashing scheme for fast SIFT feature-based image matching and retrieval. First, a training process of the hashing function involves geometric and topological information being introduced; second, a geometry-enhanced similarity evaluation that considers both the global and details of images in evaluation is explained. Compared with state-of-the-art methods, our method achieves better performance. Chen Kang, Li Zhu 0003, Xueming Qian, Junwei Han 0001, Meng Wang 0001, Yuan Yan Tang |
IEEE Trans. Multim. | 6 |
| 2019 | Maximum Likelihood Estimation-Based Joint Sparse Representation for the Classification of Hyperspectral Remote Sensing ImagesabstractA joint sparse representation (JSR) method has shown superior performance for the classification of hyperspectral images (HSIs). However, it is prone to be affected by outliers in the HSI spatial neighborhood. In order to improve the robustness of JSR, we propose a maximum likelihood estimation (MLE)-based JSR (MLEJSR) model, which replaces the traditional quadratic loss function with an MLE-like estimator for measuring the joint approximation error. The MLE-like estimator is actually a function of coding residuals. Given some priors on the coding residuals, the MLEJSR model can be easily converted to an iteratively reweighted JSR problem. Choosing a reasonable weight function, the effect of inhomogeneous neighboring pixels or outliers can be dramatically reduced. We provide a theoretical analysis of MLEJSR from the viewpoint of recovery error and evaluate its empirical performance on three public hyperspectral data sets. Both the theoretical and experimental results demonstrate the effectiveness of our proposed MLEJSR method, especially in the case of large noise. Jiangtao Peng, Luoqing Li, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | New Incremental Learning Algorithm With Support Vector MachinesabstractIncremental learning is one of the most effective methods of learning accumulated data and large-scale data. The newly increased samples of the previously known works on incremental learning are usually independent and identically distributed. To study how dependent sampling methods influence the learning ability of incremental support vector machines (ISVM) algorithm, in this paper we introduce an ISVM based on Markov resampling (MR-ISVM), and give the experimental research on the learning ability of the MR-ISVM algorithm. The experimental results indicate that the MR-ISVM algorithm has not only smaller misclassification rates and sparser of the obtained classifiers, but also less total time of sampling and training compared to ISVM based on randomly independent sampling. We also compare it with other ISVM algorithms. Jie Xu 0006, Chen Xu 0007, Bin Zou 0002, Yuan Yan Tang, Jiangtao Peng, Xinge You |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2018 | Cauchy Matching Pursuit for Robust Sparse Representation and ClassificationabstractVarious greedy algorithms have been developed for sparse signal recovery in recent years. However, most of them utilize the l2 norm based loss function and sensitive to non-Gaussian noises and outliers. This paper proposes a Cauchy matching pursuit (CauchyMP) algorithm for robust sparse representation and classification. By leveraging a Cauchy estimator based loss function, the proposed approach can robustly learn the sparse representation of noisy data corrupted by various severe noises. As a greedy algorithm, CauchyMP is also computationally efficient. We also develop a CauchyMP based classifier for robust classification with application to face recognition. The experiments on the datasets with gross corruptions demonstrate the efficacy and robustness of CauchyMP for learning robust sparse representation. Yulong Wang 0002, Cuiming Zou, Yuan Yan Tang, Luoqing Li |
ICPR | 3 |
| 2018 | Distribution preserving learning for unsupervised feature selection
Ting Xie 0004, Taiping Zhang, Yuan Yan Tang |
Neurocomputing | 4 |
| 2018 | Graph-based multiple rank regression for image classification
Junyu Li 0001, Loi Lei Lai, Yuan Yan Tang |
Neurocomputing | 4 |
| 2018 | A collaborative-competitive representation based classifier model
Xuecong Li, Loi Lei Lai, Yuan Yan Tang |
Neurocomputing | 6 |
| 2018 | An improved noninvasive method to detect Diabetes Mellitus using the Probabilistic Collaborative Representation based Classifier
Ting Shu 0001, Bob Zhang 0001, Yuan Yan Tang |
Inf. Sci. | 3 |
| 2018 | A constrained least squares regression model
Loi Lei Lai, Yuan Yan Tang |
Inf. Sci. | 4 |
| 2018 | Sparse structural feature selection for multitarget regression
Loi Lei Lai, Yuan Yan Tang |
Knowl. Based Syst. | 4 |
| 2018 | A fast convex hull algorithm inspired by human visual perception
Runzong Liu, Yuan Yan Tang, Patrick P. K. Chan |
Multim. Tools Appl. | 2 |
| 2018 | Multi-source fusion based geo-tagging for web images
Yisi Zhao, Xueming Qian, Yuan Yan Tang |
Multim. Tools Appl. | 4 |
| 2018 | Joint medical image fusion, denoising and enhancement via discriminative low-rank sparse dictionaries learning
Huafeng Li 0001, Xiaoge He, Dapeng Tao, Yuan Yan Tang, Ruxin Wang 0002 |
Pattern Recognit. | 4 |
| 2018 | Defending against the Advanced Persistent Threat: An Optimal Control ApproachabstractThe new cyberattack pattern of advanced persistent threat (APT) has posed a serious threat to modern society. This paper addresses the APT defense problem, that is, the problem of how to effectively defend against an APT campaign. Based on a novel APT attack-defense model, the effectiveness of an APT defense strategy is quantified. Thereby, the APT defense problem is modeled as an optimal control problem, in which an optimal control stands for a most effective APT defense strategy. The existence of an optimal control is proved, and an optimality system is derived. Consequently, an optimal control can be figured out by solving the optimality system. Some examples of the optimal control are given. Finally, the influence of some factors on the effectiveness of an optimal control is examined through computer experiments. These findings help organizations to work out policies of defending against APTs. Pengdeng Li, Xiaofan Yang 0001, Qingyu Xiong, Junhao Wen 0001, Yuan Yan Tang |
Secur. Commun. Networks | 5 |
| 2018 | Semi-supervised graph-based retargeted least squares regression
Loi Lei Lai, Yuan Yan Tang |
Signal Process. | 4 |
| 2018 | Tackling class overlap and imbalance problems in software defect prediction
Lin Chen 0023, Bin Fang 0001, Zhaowei Shang, Yuan Yan Tang |
Softw. Qual. J. | 4 |
| 2018 | Mixed Noise Removal via Robust Constrained Sparse RepresentationabstractIn recent years, the sparse coding-based techniques have been widely used for image denoising. However, most of the sparse coding-based mixed noise reduction methods fail to take full advantage of the geometric structure of data samples. In other words, they neglect the common information shared by the similar patches in sparse coding. To address this concern, in this paper, we propose a robust constrained sparse representation (RCSR) method to remove mixed noise. By using the center coefficient of similar patches as the guider which is approximated by the coefficient of query patch in sparse coding, the geometric structure of data can be well preserved. Moreover, different from most existing two-stage mixed noise reduction methods that use explicit detectors to restrain impulse noise, the proposed RCSR adaptively adjusts the contribution of each pixel in the loss function to eliminate the influences of outliers. Experiments on the reconstruction of synthetic data and the removal of mixed noise in real images demonstrate the effectiveness of our proposed method. Licheng Liu, C. L. Philip Chen, Xinge You, Yuan Yan Tang, Yushu Zhang 0001, Shutao Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | A Regularization Approach for Instance-Based Superset Label LearningabstractDifferent from the traditional supervised learning in which each training example has only one explicit label, superset label learning (SLL) refers to the problem that a training example can be associated with a set of candidate labels, and only one of them is correct. Existing SLL methods are either regularization-based or instance-based, and the latter of which has achieved state-of-the-art performance. This is because the latest instance-based methods contain an explicit disambiguation operation that accurately picks up the groundtruth label of each training example from its ambiguous candidate labels. However, such disambiguation operation does not fully consider the mutually exclusive relationship among different candidate labels, so the disambiguated labels are usually generated in a nondiscriminative way, which is unfavorable for the instance-based methods to obtain satisfactory performance. To address this defect, we develop a novel regularization approach for instance-based superset label (RegISL) learning so that our instance-based method also inherits the good discriminative ability possessed by the regularization scheme. Specifically, we employ a graph to represent the training set, and require the examples that are adjacent on the graph to obtain similar labels. More importantly, a discrimination term is proposed to enlarge the gap of values between possible labels and unlikely labels for every training example. As a result, the intrinsic constraints among different candidate labels are deployed, and the disambiguated labels generated by RegISL are more discriminative and accurate than those output by existing instance-based algorithms. The experimental results on various tasks convincingly demonstrate the superiority of our RegISL to other typical SLL methods in terms of both training accuracy and test accuracy. Chen Gong 0002, Tongliang Liu, Yuan Yan Tang, Jian Yang 0003, Jie Yang 0002, Dacheng Tao |
IEEE Trans. Cybern. | 3 |
| 2018 | Video Saliency Detection Using Object ProposalsabstractIn this paper, we introduce a novel approach to identify salient object regions in videos via object proposals. The core idea is to solve the saliency detection problem by ranking and selecting the salient proposals based on object-level saliency cues. Object proposals offer a more complete and high-level representation, which naturally caters to the needs of salient object detection. As well as introducing this novel solution for video salient object detection, we reorganize various discriminative saliency cues and traditional saliency assumptions on object proposals. With object candidates, a proposal ranking and voting scheme, based on various object-level saliency cues, is designed to screen out nonsalient parts, select salient object regions, and to infer an initial saliency estimate. Then a saliency optimization process that considers temporal consistency and appearance differences between salient and nonsalient regions is used to refine the initial saliency estimates. Our experiments on public datasets (SegTrackV2, Freiburg-Berkeley Motion Segmentation Dataset, and Densely Annotated Video Segmentation) validate the effectiveness, and the proposed method produces significant improvements over state-of-the-art algorithms. Wenguan Wang, Jianbing Shen, Ling Shao 0001, Jian Yang 0009, Dacheng Tao, Yuan Yan Tang |
IEEE Trans. Cybern. | 7 |
| 2018 | Robust Face Hallucination via Locality-Constrained Bi-Layer RepresentationabstractRecently, locality-constrained linear coding (LLC) has been drawn great attentions and been widely used in image processing and computer vision tasks. However, the conventional LLC model is always fragile to outliers. In this paper, we present a robust locality-constrained bi-layer representation model to simultaneously hallucinate the face images and suppress noise and outliers with the assistant of a group of training samples. The proposed scheme is not only able to capture the nonlinear manifold structure but also robust to outliers by incorporating a weight vector into the objective function to subtly tune the contribution of each pixel offered in the objective. Furthermore, a high-resolution (HR) layer is employed to compensate the missed information in the low-resolution (LR) space for coding. The use of two layers (the LR layer and the HR layer) is expected to expose the complicated correlation between the LR and HR patch spaces, which helps to obtain the desirable coefficients to reconstruct the final HR face. The experimental results demonstrate that the proposed method outperforms the state-of-the-art image super-resolution methods in terms of both quantitative measurements and visual effects. Licheng Liu, C. L. Philip Chen, Shutao Li 0001, Yuan Yan Tang, Long Chen 0001 |
IEEE Trans. Cybern. | 4 |
| 2018 | Simultaneous Spectral-Spatial Feature Selection and Extraction for Hyperspectral ImagesabstractIn hyperspectral remote sensing data mining, it is important to take into account of both spectral and spatial information, such as the spectral signature, texture feature, and morphological property, to improve the performances, e.g., the image classification accuracy. In a feature representation point of view, a nature approach to handle this situation is to concatenate the spectral and spatial features into a single but high dimensional vector and then apply a certain dimension reduction technique directly on that concatenated vector before feed it into the subsequent classifier. However, multiple features from various domains definitely have different physical meanings and statistical properties, and thus such concatenation has not efficiently explore the complementary properties among different features, which should benefit for boost the feature discriminability. Furthermore, it is also difficult to interpret the transformed results of the concatenated vector. Consequently, finding a physically meaningful consensus low dimensional feature representation of original multiple features is still a challenging task. In order to address these issues, we propose a novel feature learning framework, i.e., the simultaneous spectral-spatial feature selection and extraction algorithm, for hyperspectral images spectral-spatial feature representation and classification. Specifically, the proposed method learns a latent low dimensional subspace by projecting the spectral-spatial feature into a common feature space, where the complementary information has been effectively exploited, and simultaneously, only the most significant original features have been transformed. Encouraging experimental results on three public available hyperspectral remote sensing datasets confirm that our proposed method is effective and efficient. Lefei Zhang, Qian Zhang 0009, Bo Du 0001, Xin Huang 0002, Yuan Yan Tang, Dacheng Tao |
IEEE Trans. Cybern. | 5 |
| 2018 | Corrections to "Dictionary Learning-Based Feature-Level Domain Adaptation for Cross-Scene Hyperspectral Image Classification"abstractIn the above paper[1], there is an error inFig. 14.Fig. 14should include$3\times3$matrices rather than$7\times7$, since the Shanghai-Hangzhou dataset has three land-cover classes. The corrected figure appears here. Minchao Ye, Yuntao Qian, Jun Zhou 0001, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Effective and Fast Estimation for Image Sensor Noise Via Constrained Weighted Least SquaresabstractNoise estimation is crucial in many image processing algorithms such as image denoising. Conventionally, the noise is assumed as a signal-independent additive white Gaussian process. However, for the real raw data of image sensor, the present noise should be practically modeled as signal dependent. In this paper, we propose an effective and fast image sensor noise estimation method for a single raw image. The noise model parameters are estimated via constrained weighted least squares (WLS) fitting on a number of data samples, each of which is generated from a group of weakly textured patches. Specifically, we first design a fast scheme for selecting weakly textured patches, with the guidance of image histogram. To robustly fit the data samples, we then explicitly account for the credibility of each sample by measuring the texture strength of the grouped patches. The image sensor noise estimation is finally formulated as a constrained WLS optimization problem, which can be solved efficiently. Experimental results demonstrate that our method could run much faster than the existing schemes, while retaining the state-of-the-art estimation performance. Li Dong 0006, Jiantao Zhou 0001, Yuan Yan Tang |
IEEE Trans. Image Process. | 3 |
| 2018 | Learning With Coefficient-Based Regularized Regression on Markov ResamplingabstractBig data research has become a globally hot topic in recent years. One of the core problems in big data learning is how to extract effective information from the huge data. In this paper, we propose a Markov resampling algorithm to draw useful samples for handling coefficient-based regularized regression (CBRR) problem. The proposed Markov resampling algorithm is a selective sampling method, which can automatically select uniformly ergodic Markov chain (u.e.M.c.) samples according to transition probabilities. Based on u.e.M.c. samples, we analyze the theoretical performance of CBRR algorithm and generalize the existing results on independent and identically distributed observations. To be specific, when the kernel is infinitely differentiable, the learning rate depending on the sample size $m$ can be arbitrarily close to $\mathcal {O}(m^{-1})$ under a mild regularity condition on the regression function. The good generalization ability of the proposed method is validated by experiments on simulated and real data sets. Luoqing Li, Weifu Li, Bin Zou 0002, Yulong Wang 0002, Yuan Yan Tang, Hua Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | k-Times Markov Sampling for SVMCabstractSupport vector machine (SVM) is one of the most widely used learning algorithms for classification problems. Although SVM has good performance in practical applications, it has high algorithmic complexity as the size of training samples is large. In this paper, we introduce SVM classification (SVMC) algorithm based on -times Markov sampling and present the numerical studies on the learning performance of SVMC with -times Markov sampling for benchmark data sets. The experimental results show that the SVMC algorithm with -times Markov sampling not only have smaller misclassification rates, less time of sampling and training, but also the obtained classifier is more sparse compared with the classical SVMC and the previously known SVMC algorithm based on Markov sampling. We also give some discussions on the performance of SVMC with -times Markov sampling for the case of unbalanced training samples and large-scale training samples. Bin Zou 0002, Chen Xu 0007, Yang Lu 0009, Yuan Yan Tang, Jie Xu 0006, Xinge You |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Robust Privacy-Preserving Image Sharing over Online Social Networks (OSNs)abstractSharing images online has become extremely easy and popular due to the ever-increasing adoption of mobile devices and online social networks (OSNs). The privacy issues arising from image sharing over OSNs have received significant attention in recent years. In this article, we consider the problem of designing a secure, robust, high-fidelity, storage-efficient image-sharing scheme over Facebook, a representative OSN that is widely accessed. To accomplish this goal, we first conduct an in-depth investigation on the manipulations that Facebook performs to the uploaded images. Assisted by such knowledge, we propose a DCT-domain image encryption/decryption framework that is robust against these lossy operations. As verified theoretically and experimentally, superior performance in terms of data privacy, quality of the reconstructed images, and storage cost can be achieved. Weiwei Sun 0009, Jiantao Zhou 0001, Shuyuan Zhu, Yuan Yan Tang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2018 | Fast and Accurate Vanishing Point Detection and Its Application in Inverse Perspective Mapping of Structured RoadabstractFast and accurate visual scene understanding in autonomous vehicles is necessary but still very challenging. An autonomous vehicle must be taught to read the road like a human driver for better controlling the vehicle, so it is important to efficiently detect the road area and road markings. In this paper, we mainly focus on the vanishing point detection and its application in inverse perspective mapping (IPM) for road marking understanding. We first propose a fast and accurate vanishing point detection method for various types of roads, by adopting and improving Weber local descriptor to obtain salient representative texture and orientation information of the road area, and then voting for the dominant vanishing point with a simple line-voting scheme. Experimental results demonstrate that the proposed vanishing point detection approach gains a better performance than some state-of-the-art methods in terms of accuracy and computation time. Furthermore, we introduce the detected vanishing point into the IPM algorithm in the structured road environment, since some important calibration parameters can be automatically calculated by the vanishing point, especially on the rough road. Experiments also show that our proposed vanishing point-based IPM method is adaptive and accurate, which is conducive to the subsequent road marking detection and recognition. Weibin Yang, Bin Fang 0001, Yuan Yan Tang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2017 | Dynamic Weighted Majority for Incremental Learning of Imbalanced Data Streams with Concept DriftabstractConcept drifts occurring in data streams will jeopardize the accuracy and stability of the online learning process. If the data stream is imbalanced, it will be even more challenging to detect and cure the concept drift. In the literature, these two problems have been intensively addressed separately, but have yet to be well studied when they occur together. In this paper, we propose a chunk-based incremental learning method called Dynamic Weighted Majority for Imbalance Learning (DWMIL) to deal with the data streams with concept drift and class imbalance problem. DWMIL utilizes an ensemble framework by dynamically weighting the base classifiers according to their performance on the current data chunk. Compared with the existing methods, its merits are four-fold: (1) it can keep stable for non-drifted streams and quickly adapt to the new concept; (2) it is totally incremental, i.e. no previous data needs to be stored; (3) it keeps a limited number of classifiers to ensure high efficiency; and (4) it is simple and needs only one thresholding parameter. Experiments on both synthetic and real data sets with concept drift show that DWMIL performs better than the state-of-the-art competitors, with less computational cost. Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang |
IJCAI | 3 |
| 2017 | Maximum correntropy criterion for convex anc semi-nonnegative matrix factorizationabstractMatrix factorization is a popular low dimensional representation approach that plays an important role in many pattern recognition and computer vision domains. Among them, convex and semi-nonnegative matrix factorizations have attracted considerable interest, owing to its clustering interpretation. On the other hand, the generalized correlation function (correntropy) as the error measure does not depend on the assumption of Gaussianity, which the mean square error (MSE) heavily depends on. In this paper, we propose two novel algorithms, called Maximum Correntropy Criterion based Convex and Semi-Nonnegative Matrix Factorization (MCC-ConvexNMF, MCC-SemiNMF). Compared with the mean square error based convex and semi-nonnegative matrix factorization, the proposed methods can extract more information from the data and produce more accurate solutions. Experimental results on both synthetic dataset and the popular face database illustrate the effectiveness of our methods. Anyong Qin, Zhaowei Shang, Jinyu Tian 0001, Ailin Li, Yulong Wang 0002, Yuan Yan Tang |
SMC | 6 |
| 2017 | Information-theoretic generalized orthogonal matching pursuit for robust pattern classificationabstractOwing to its simplicity and efficacy, orthogonal matching pursuit (OMP) has been a popular sparse representation method for compressed sensing and pattern classification. As a recent extension of OMP, generalized OMP (GOMP) improves the efficiency of OMP by identifying multiple atoms each iteration. Nonetheless, GOMP utilizes the mean square error (MSE) criterion as the loss function, which has been proven to rely on the Gaussianity assumption of the noise distribution and sensitive to non-Gaussian noise. In this paper, we propose a robust sparse representation method, called information-theoretic generalized OMP (ITGOMP), to reduce the limitation of GOMP. The key idea is to minimize the correntropy based information-theoretic loss function, which is independent of the noise distribution. We also devise a half-quadratic based algorithm to tackle the optimization problem. Finally, an ITGOMP based classifier is developed for robust pattern classification. The experiments on public real-world databases verify the effectiveness and robustness of the proposed method for classification. Yulong Wang 0002, Yuan Yan Tang, Cuiming Zou |
SMC | 2 |
| 2017 | Efficient single image dehazing and denoising: An efficient multi-scale correlated wavelet approach
Xin Liu 0011, Yiu-Ming Cheung, Xinge You, Yuan Yan Tang |
Comput. Vis. Image Underst. | 5 |
| 2017 | Empirical Mode Decomposition - Window Fractal (EMDWF) Algorithm in Classification of Fingerprint of Medicinal HerbsabstractThis paper presents a new approach called the empirical mode decomposition — window fractal (EMDWF) algorithm in classification of fingerprint of medicinal herbs. In this way, we consider a glycyrrhiza fingerprint of medicinal herb as a signal sequence, and apply empirical mode decomposition (EMD) and Hiaguchis fractal dimension to construct a feature vector. By using EMD, the glycyrrhiza fingerprint of medicinal herb can be decomposed into some intrinsic mode functions (IMFs). As window fractal dimension (WFD) is applied to each IMF and original signal, the features of the glycyrrhiza fingerprint of medicinal herb can be obtained. Thereafter, SVM is applied as a classifier. The results of the experiments state clearly that the feature extracted by EMDWF is better than that of the existing methods including the pure EMD. With the increase of the number of training samples and the increase of the number of layers in EMD, the classification result achieves more stability. Jianwei Du, Zhengguang Xu, Zhichun Mu, Patrick Shen-Pei Wang, Yuan Yan Tang, Huiwu Luo |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2017 | Quaternionic Weber Local Descriptor of Color ImagesabstractThis paper proposes a simple but effective framework named quaternionic Weber local descriptor (QWLD) for color image feature extraction. Integrating quaternionic representation (QR) of the color image and Weber's law (WL), QWLD possesses both their superiorities. It uses QR to handle all color channels of the image in a holistic way while preserving their relations, and applies WL to ensure that the derived descriptors are robust and discriminative. Using the QWLD framework, we further develop the quaternionic-increment-based Weber descriptor and quaternionic-distance-based Weber descriptor in terms of different perspectives. Extensive experiments on different color image recognition problems demonstrate that the proposed framework and descriptors outperform state-of-the-art local descriptors. Rushi Lan, Yicong Zhou, Yuan Yan Tang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Robust Object Tracking via Key Patch Sparse RepresentationabstractMany conventional computer vision object tracking methods are sensitive to partial occlusion and background clutter. This is because the partial occlusion or little background information may exist in the bounding box, which tends to cause the drift. To this end, in this paper, we propose a robust tracker based on key patch sparse representation (KPSR) to reduce the disturbance of partial occlusion or unavoidable background information. Specifically, KPSR first uses patch sparse representations to get the patch score of each patch. Second, KPSR proposes a selection criterion of key patch to judge the patches within the bounding box and select the key patch according to its location and occlusion case. Third, KPSR designs the corresponding contribution factor for the sampled patches to emphasize the contribution of the selected key patches. Comparing the KPSR with eight other contemporary tracking methods on 13 benchmark video data sets, the experimental results show that the KPSR tracker outperforms classical or state-of-the-art tracking methods in the presence of partial occlusion, background clutter, and illumination change. Zhenyu He 0001, Shuangyan Yi, Yiu-Ming Cheung, Xinge You, Yuan Yan Tang |
IEEE Trans. Cybern. | 5 |
| 2017 | Cross-Domain Recognition by Identifying Joint Subspaces of Source Domain and Target DomainabstractThis paper introduces a new method to solve the cross-domain recognition problem. Different from the traditional domain adaption methods which rely on a global domain shift for all classes between the source and target domains, the proposed method is more flexible to capture individual class variations across domains. By adopting a natural and widely used assumption that the data samples from the same class should lay on an intrinsic low-dimensional subspace, even if they come from different domains, the proposed method circumvents the limitation of the global domain shift, and solves the cross-domain recognition by finding the joint subspaces of the source and target domains. Specifically, given labeled samples in the source domain, we construct a subspace for each of the classes. Then we construct subspaces in the target domain, called anchor subspaces, by collecting unlabeled samples that are close to each other and are highly likely to belong to the same class. The corresponding class label is then assigned by minimizing a cost function which reflects the overlap and topological structure consistency between subspaces across the source and target domains, and within the anchor subspaces, respectively. We further combine the anchor subspaces to the corresponding source subspaces to construct the joint subspaces. Subsequently, one-versus-rest support vector machine classifiers are trained using the data samples belonging to the same joint subspaces and applied to unlabeled data in the target domain. We evaluate the proposed method on two widely used datasets: 1) object recognition dataset for computer vision tasks and 2) sentiment classification dataset for natural language processing tasks. Comparison results demonstrate that the proposed method outperforms the comparison methods on both datasets. Yuewei Lin, Jing Chen 0008, Yu Cao 0003, Youjie Zhou, Lingfeng Zhang 0001, Yuan Yan Tang, Song Wang 0002 |
IEEE Trans. Cybern. | 6 |
| 2017 | Weighted Joint Sparse Representation for Removing Mixed Noise in ImageabstractJoint sparse representation (JSR) has shown great potential in various image processing and computer vision tasks. Nevertheless, the conventional JSR is fragile to outliers. In this paper, we propose a weighted JSR (WJSR) model to simultaneously encode a set of data samples that are drawn from the same subspace but corrupted with noise and outliers. Our model is desirable to exploit the common information shared by these data samples while reducing the influence of outliers. To solve the WJSR model, we further introduce a greedy algorithm called weighted simultaneous orthogonal matching pursuit to efficiently approximate the global optimal solution. Then, we apply the WJSR for mixed noise removal by jointly coding the grouped nonlocal similar image patches. The denoising performance is further improved by incorporating it with the global prior and the sparse errors into a unified framework. Experimental results show that our denoising method is superior to several state-of-the-art mixed noise removal methods. Licheng Liu, Long Chen 0001, C. L. Philip Chen, Yuan Yan Tang, Chi-Man Pun |
IEEE Trans. Cybern. | 4 |
| 2017 | Correntropy Matching Pursuit With Application to Robust Digit and Face RecognitionabstractAs an efficient sparse representation algorithm, orthogonal matching pursuit (OMP) has attracted massive attention in recent years. However, OMP and most of its variants estimate the sparse vector using the mean square error criterion, which depends on the Gaussianity assumption of the error distribution. A violation of this assumption, e.g., non-Gaussian noise, may lead to performance degradation. In this paper, a correntropy matching pursuit (CMP) method is proposed to alleviate this problem of OMP. Unlike many other matching pursuit methods, our method is independent of the error distribution. We show that CMP can adaptively assign small weights on severely corrupted entries of data and large weights on clean ones, thus reducing the effect of large noise. Our another contribution is to develop a robust sparse representation-based recognition method based on CMP. Experiments on synthetic and real data show the effectiveness of our method for both sparse approximation and pattern recognition, especially for noisy, corrupted, and incomplete data. Yulong Wang 0002, Yuan Yan Tang, Luoqing Li |
IEEE Trans. Cybern. | 2 |
| 2017 | Spectral-Spatial Shared Linear Regression for Hyperspectral Image ClassificationabstractClassification of the pixels in hyperspectral image (HSI) is an important task and has been popularly applied in many practical applications. Its major challenge is the high-dimensional small-sized problem. To deal with this problem, lots of subspace learning (SL) methods are developed to reduce the dimension of the pixels while preserving the important discriminant information. Motivated by ridge linear regression (RLR) framework for SL, we propose a spectral-spatial shared linear regression method (SSSLR) for extracting the feature representation. Comparing with RLR, our proposed SSSLR has the following two advantages. First, we utilize a convex set to explore the spatial structure for computing the linear projection matrix. Second, we utilize a shared structure learning model, which is formed by original data space and a hidden feature space, to learn a more discriminant linear projection matrix for classification. To optimize our proposed method, an efficient iterative algorithm is proposed. Experimental results on two popular HSI data sets, i.e., Indian Pines and Salinas demonstrate that our proposed methods outperform many SL methods. Yuan Yan Tang |
IEEE Trans. Cybern. | 2 |
| 2017 | Matrix-Vector Nonnegative Tensor Factorization for Blind Unmixing of Hyperspectral ImageryabstractMany spectral unmixing approaches ranging from geometry, algebra to statistics have been proposed, in which nonnegative matrix factorization (NMF)-based ones form an important family. The original NMF-based unmixing algorithm loses the spectral and spatial information between mixed pixels when stacking the spectral responses of the pixels into an observed matrix. Therefore, various constrained NMF methods are developed to impose spectral structure, spatial structure, and spectral-spatial joint structure into NMF to enforce the estimated endmembers and abundances preserve these structures. Compared with matrix format, the third-order tensor is more natural to represent a hyperspectral data cube as a whole, by which the intrinsic structure of hyperspectral imagery can be losslessly retained. Extended from NMF-based methods, a matrix-vector nonnegative tensor factorization (NTF) model is proposed in this paper for spectral unmixing. Different from widely used tensor factorization models, such as canonical polyadic decomposition CPD) and Tucker decomposition, the proposed method is derived from block term decomposition, which is a combination of CPD and Tucker decomposition. This leads to a more flexible frame to model various application-dependent problems. The matrix-vector NTF decomposes a third-order tensor into the sum of several component tensors, with each component tensor being the outer product of a vector (endmember) and a matrix (corresponding abundances). From a formal perspective, this tensor decomposition is consistent with linear spectral mixture model. From an informative perspective, the structures within spatial domain, within spectral domain, and cross spectral-spatial domain are retreated interdependently. Experiments demonstrate that the proposed method has outperformed several state-of-the-art NMF-based unmixing methods. Yuntao Qian, Fengchao Xiong, Shan Zeng, Jun Zhou 0001, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | Hyperspectral Image Classification Using Principal Components-Based Smooth Ordering and Multiple 1-D InterpolationabstractThis paper proposes a spectral-spatial classification algorithm based on principal components (PCs)-based smooth ordering and multiple 1-D interpolation, which can alleviate the general classification problems effectively. Because of the characteristics of hyperspectral image, there always exist easily separable samples (ESSs) and difficultly separable samples (DSSs) in view of the different sets of labeled samples. In this paper, the PC analysis is first used for reducing features and extracting the few first PCs of a hyperspectral image. Then, PC-based smooth ordering is designed for the separation of ESSs and DSSs, and multiple 1-D interpolation is used for the accurate classification of the ESSs. Next, the highly confident samples are selected from the ESSs by the spatial neighborhood information, which are added into the training set for the classification of DSSs. In the case of sufficient training samples, a supervised spectral-spatial method is used for classifying the DSSs by combining the spatial information built with popular extended multiattribute profiles. The proposed algorithm is compared with some state-of-the-art methods on three hyperspectral data sets. The results demonstrate that the presented algorithm achieves much better classification performance in terms of the accuracy and the computation time. Zhijing Ye 0001, Hong Li 0009, Yalong Song, Jón Atli Benediktsson, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | Dictionary Learning-Based Feature-Level Domain Adaptation for Cross-Scene Hyperspectral Image ClassificationabstractA big challenge of hyperspectral image (HSI) classification is the small size of labeled pixels for training classifier. In real remote sensing applications, we always face the situation that an HSI scene is not labeled at all, or is with very limited number of labeled pixels, but we have sufficient labeled pixels in another HSI scene with the similar land cover classes. In this paper, we try to classify an HSI scene containing no labeled sample or only a few labeled samples with the help of a similar HSI scene having a relative large size of labeled samples. The former scene is defined as the target scene, while the latter one is the source scene. We name this classification problem as cross-scene classification. The main challenge of cross-scene classification is spectral shift, i.e., even for the same class in different scenes, their spectral distributions maybe have significant deviation. As all or most training samples are drawn from the source scene, while the prediction is performed in the target scene, the difference in spectral distribution would greatly deteriorate the classification performance. To solve this problem, we propose a dictionary learning-based feature-level domain adaptation technique, which aligns the spectral distributions between source and target scenes by projecting their spectral features into a shared low-dimensional embedding space by multitask dictionary learning. The basis atoms in the learned dictionary represent the common spectral components, which span a cross-scene feature space to minimize the effect of spectral shift. After the HSIs of two scenes are transformed into the shared space, any traditional HSI classification approach can be used. In this paper, sparse logistic regression (SRL) is selected as the classifier. Especially, if there are a few labeled pixels in the target domain, multitask SRL is used to further promote the classification performance. The experimental results on synthetic and real HSIs show the advantages of the proposed method for cross-scene classification. Minchao Ye, Yuntao Qian, Jun Zhou 0001, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2017 | Efficient Human Motion Retrieval via Temporal Adjacent Bag of Words and Discriminative Neighborhood Preserving Dictionary LearningabstractHuman motion retrieval from motion capture data forms the fundamental basis for computer animation. In this paper, the authors propose an efficient human motion retrieval approach via temporal adjacent bag of words (TA-BoW) and discriminative neighborhood preserving dictionary learning (DNP-DL). The retrieval process includes two phases: offline training and online retrieval. In the first phase, the original skeleton model is first simplified and then pairwise joint distances are computed to characterize each motion frame. Then, a novel motion descriptor, namely TABoW, is proposed to discriminatively code the motion appearances, through which the articulated complexity and spatiotemporal dimensionality can be greatly reduced. Subsequently, by considering the neighborhood relationships of intraclass structure and the advantage of Fisher criterion, a DNP-DL method is exploited through which each human action can be discriminatively and sparsely represented by a linear combination of such dictionary atoms. In the second phase, a hierarchical retrieval mechanism is used by incorporating the sparse classification and chi-square ranking, whereby the searching range is significantly reduced. The experimental results show that the proposed human motion retrieval approach performs better than the state-of-the-art competing approaches. Xin Liu 0011, Gao-Feng He, Shu-Juan Peng, Yiu-Ming Cheung, Yuan Yan Tang |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2017 | Noise Level Estimation for Natural Images Based on Scale-Invariant Kurtosis and Piecewise StationarityabstractNoise level estimation is crucial in many image processing applications, such as blind image denoising. In this paper, we propose a novel noise level estimation approach for natural images by jointly exploiting the piecewise stationarity and a regular property of the kurtosis in bandpass domains. We design a K-means-based algorithm to adaptively partition an image into a series of non-overlapping regions, each of whose clean versions is assumed to be associated with a constant, but unknown kurtosis throughout scales. The noise level estimation is then cast into a problem to optimally fit this new kurtosis model. In addition, we develop a rectification scheme to further reduce the estimation bias through noise injection mechanism. Extensive experimental results show that our method can reliably estimate the noise level for a variety of noise types, and outperforms some state-of-the-art techniques, especially for non-Gaussian noises. Li Dong 0006, Jiantao Zhou 0001, Yuan Yan Tang |
IEEE Trans. Image Process. | 3 |
| 2017 | Image Re-Ranking Based on Topic DiversityabstractSocial media sharing Websites allow users to annotate images with free tags, which significantly contribute to the development of the web image retrieval. Tag-based image search is an important method to find images shared by users in social networks. However, how to make the top ranked result relevant and with diversity is challenging. In this paper, we propose a topic diverse ranking approach for tag-based image retrieval with the consideration of promoting the topic coverage performance. First, we construct a tag graph based on the similarity between each tag. Then, the community detection method is conducted to mine the topic community of each tag. After that, inter-community and intra-community ranking are introduced to obtain the final retrieved results. In the inter-community ranking process, an adaptive random walk model is employed to rank the community based on the multi-information of each topic community. Besides, we build an inverted index structure for images to accelerate the searching process. Experimental results on Flickr data set and NUS-Wide data sets show the effectiveness of the proposed approach. Xueming Qian, Dan Lu 0003, Yaxiong Wang, Li Zhu 0003, Yuan Yan Tang, Meng Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2017 | Learning the Distribution Preserving Semantic Subspace for ClusteringabstractThis paper proposes a new clustering method for images called distribution preserving indexing (DPI). It aims to find a lower dimensional semantic space approximating the original image space in the sense of preserving the distribution of the data. In the theory, the intrinsic structure of the data clusters can be described by the distribution of the data effectively. Therefore, the cluster structure of the data in a lower dimensional semantic space derived by the DPI becomes clear. Unlike these distance-based clustering methods, which reveal the intrinsic Euclidean structure of data, our method attempts to discover the intrinsic cluster structure of the data space that actually is the union of some sub-manifolds. Moreover, we propose a revised kernel density estimator for the case of high-dimensional data, which is a crucial step in DPI. In addition, we provide a theoretical analysis of the bound of our method. Finally, the extensive experiments compared with other algorithms, on COIL20, CBCL, and MNIST demonstrate the effectiveness of our proposed approach. Jinyu Tian 0001, Taiping Zhang, Anyong Qin, Zhaowei Shang, Yuan Yan Tang |
IEEE Trans. Image Process. | 5 |
| 2017 | Image Location Inference by Multisaliency EnhancementabstractLocations of images have been widely used in many application scenarios for large geotagged image corpora. As to images that are not geographically tagged, we estimate their locations with the help of the large geotagged image set by content-based image retrieval. Bag-of-words image representation has been utilized widely. However, the individual visual word-based image retrieval approach is not effective in expressing the salient relationships of image region. In this paper, we present an image location estimation approach by multisaliency enhancement. We first extract region-of-interests (ROIs) by mean-shift clustering on the visual words and salient map of the image based on which we further determine the importance of the ROI. Then, we describe each ROI by the spatial descriptors of visual words. Finally, region-based visual phrases are generated to further enhance the saliency in image location estimation. Experiments show the effectiveness of our proposed approach. Xueming Qian, Huan Wang 0002, Yisi Zhao, Xingsong Hou, Richang Hong, Meng Wang 0001, Yuan Yan Tang |
IEEE Trans. Multim. | 7 |
| 2017 | A Hybrid of Local and Global Saliencies for Detecting Image Salient Region and AppearanceabstractThis paper presents a visual saliency detection approach, which is a hybrid of local feature-based saliency and global feature-based saliency (simply called local saliency and global saliency, respectively, for short). First, we propose an automatic selection of smoothing parameter scheme to make the foreground and background of an input image more homogeneous. Then, we partition the smoothed image into a set of regions and compute the local saliency by measuring the color and texture dissimilarity in the smoothed regions and the original regions, respectively. Furthermore, we utilize the global color distribution model embedded with color coherence, together with the multiple edge saliency, to yield the global saliency. Finally, we combine the local and global saliencies, and utilize the composition information to obtain the final saliency. Experimental results show the efficacy of the proposed method, featuring: 1) the enhanced accuracy of detecting visual salient region and appearance in comparison with the existing counterparts, 2) the robustness against the noise and the low-resolution problem of images, and 3) its applicability to multisaliency detection task. Qinmu Peng, Yiu-Ming Cheung, Xinge You, Yuan Yan Tang |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2016 | Information-theoretic atomic representation for robust pattern classificationabstractRepresentation-based classifiers (RCs) including sparse RC (SRC) have attracted intensive interest in pattern recognition in recent years. In our previous work, we have proposed a general framework called atomic representation-based classifier (ARC) including many popular RCs as special cases. Despite the empirical success, ARC and conventional RCs utilize the mean square error (MSE) criterion and assign the same weights to all entries of the test data, including both severely corrupted and clean ones. This makes ARC sensitive to the entries with large noise and outliers. In this work, we propose an information-theoretic ARC (ITARC) framework to alleviate such limitation of ARC. Using ITARC as a general platform, we develop three novel representation-based classifiers. The experiments on public real-world datasets demonstrate the efficacy of ITARC for robust pattern recognition. Yulong Wang 0002, Yuan Yan Tang, Luoqing Li, Patrick Shen-Pei Wang |
ICPR | 2 |
| 2016 | Single image super-resolution with non-local balanced low-rank matrix restorationabstractSingle image super-resolution (SR) has gained popularity to construct a high-resolution (HR) image from a single low-resolution (LR) version. More recently, non-local self similarity (NSS) has been attracted enormous interests in the field of SR, and the non-local means (NLM)-based methods are classical NSS-based SR methods. However, NLM-based methods neglect the structure information in the patches and structural similarity between patches, so it will be prone to introduce unexpected details into resultant HR images. In this paper, we propose a non-local balanced low rank matrix restoration model (NB-LRM) to improve the performance of SR which will overcome the drawbacks of NLM-based methods and take full advantage of the NSS prior. The proposed algorithm formulates the constrained optimization problem for HR image recovery. First, to take advantage of the local structure in the patch and the structural similarity between the non-local similar patches, we propose a measurement of the similarity based on both Euclidean distance and Pearson distance, then reconstruct the target patch by weighted average the similar patches. Second, to guarantee the structural similarity and linear correlation between the target patch and similar patches, we propose a new low rank regular term. Third, we introduce the iterative low rank regular algorithm to solve our model. Addition, this method doesn't need other image priors and can produce more robust reconstruction of image local structures. Compared with state-of-the-art SR methods, the proposed NB-LRM method achieves highly competitive PSNR and SSIM result, while demonstrating better edge and texture preservation performance. Xinge You, Weiyong Xue, Jiajia Lei, Peng Zhang 0040, Yiu-Ming Cheung, Yuan Yan Tang, Naiding Zhou |
ICPR | 6 |
| 2016 | Maximal level estimation and unbalance reduction for graph signal downsamplingabstractThe emerging field of graph signal processing requires a solid design of downsampling operation for graph signals to extend pattern recognition, machine learning and signal processing techniques into the graph setting. The state-of-the-art downsampling method is constructed upon the maximum spanning trees of the graphs. However, under the framework of this method, unbalanced downsampling often occurs for signals defined on densely connected unweighted graphs, such as social network data. The unbalance also significantly reduces the maximal downsampling level, making it smaller than the level we expect. In applications, the maximal level must be estimated to ensure that it is larger than the expected level; meanwhile, the unbalance has to be reduced, if it occurs. In this paper, we propose a novel method to jointly estimate the maximal level and reduce the downsampling unbalance. This method also offers an estimation of the possibility of unbalanced downsampling. If a graph signal is classified to be with high unbalance possibility, the maximum spanning tree will be updated to generate a balanced downsampling. The simulation results on synthesis and real world data support the theoretical analysis. Xianwei Zheng, Yuan Yan Tang, Jiantao Zhou 0001, Patrick Shen-Pei Wang |
ICPR | 2 |
| 2016 | Hybrid Sampling with Bagging for Class Imbalance Learning
Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang |
PAKDD (1) | 3 |
| 2016 | A hybrid swarm optimization for neural network training with application in stock price forecastingabstractA improved swarm optimization method based on particle swarm optimization (PSO) and simplified swarm optimization (SSO) is proposed to adjust the weight in artificial neural network. This method is a modification of traditional PSO and SSO, and combines them to a new optimization method (PSOSSO for short). The proposed method overcomes some of the drawbacks of SSO and improves its ability to train the weight of ANN. In the experiments, the PSOSSO is employed to train fuzzy wavelet neural network (FWNN) forecasting model to predict the prices of Hong Kong Hang Seng Index. The experimental results present that the PSOSSO is more efficient than traditional PSO and SSO methods. Jianjia Pan, Yuan Yan Tang, Yulong Wang 0002, Xianwei Zheng, Huiwu Luo, Patrick Shen-Pei Wang |
SMC | 2 |
| 2016 | Road curve fitting by multi-resolution analysisabstractIn this paper, we propose a new method for road curve fitting in urban environment based on multi-resolution analysis. The main technical contributions of the proposed method are the reconstructed approximation on the basis of the wavelet decomposition structure for curve fitting and the de-noising via the wavelet coefficients thresholding. The carried out experimental tests show promising results in a series of continuous driving images, validating our suggested method can fit the road curve effectively and computational efficiently. Yuan Yan Tang, Patrick Shen-Pei Wang |
SMC | 2 |
| 2016 | Improving unbalanced downsampling via maximum spanning trees for graph signalsabstractThe state-of-the-art downsampling method for graph signals has been constructed by using maximum spanning trees (MSTs) of the graphs. For the graph signals defined on unweighted densely connected graphs, such as social network data, the sampling rates via MST-based downsampling are not close to 1/2, leading to a unbalanced downsampling phenomenon on multi-level downsampling. The unbalance hinders the applications of MST-based downsampling on constructing graph signal multiscale transforms, such as graph wavelet decomposition and multiscale pyramid transform. In this paper, we propose a simple but efficient method to improve the performance of the MST-based method on downsampling balance. For every graph signal, we first propose an unbalance possibility to measure the unbalance of the MST-based downsampling. If the unbalance possibility is high, the downsampling will be conducted on an improved MST, which is constructed by rearranging the structure of the MST to reduce the downsampling unbalance. The experiment results on synthesis graph signal show that the proposed improved MST leads to balanced downsampling. That is, the sampling rates produced by the improved MST are closer to 1/2 in multi-level downsampling than the original MST-based method. Xianwei Zheng, Yuan Yan Tang, Jiantao Zhou 0001, Patrick Shen-Pei Wang |
SMC | 2 |
| 2016 | Lip event detection using oriented histograms of regional optical flow and low rank affinity pursuit
Xin Liu 0011, Yiu-Ming Cheung, Yuan Yan Tang |
Comput. Vis. Image Underst. | 3 |
| 2016 | Scene-adaptive single image dehazing via opening dark channel modelabstractMany traditional dark channel prior based haze removal schemes often suffer from the colour distortion and generate halo artefacts in the remote scenes. To tackle these issues, the authors present an efficient scene‐adaptive single image dehazing approach via opening dark channel model (ODCM). First, the authors detect the image depth information and separate it into close view and distant view. Then, an ODCM is proposed to optimise the whole atmospheric veil, in which the values of close view are regularised by a minimum channel image while the distant parts are estimated by an appropriate lower constant. Accordingly, the transmission map can be further optimised by guide filter and smoothed by domain transform filter. Finally, the haze degraded image can be well restored by the atmosphere scattering model. The extensive experiments have shown that the proposed image dehazing approach has significantly increased the perceptual visibility of the scene and achieved a better colour fidelity visually. Xin Liu 0011, Yuan Yan Tang, Jixiang Du |
IET Image Process. | 3 |
| 2016 | Recognition of leaf image set based on manifold-manifold distance
Jixiang Du, Mei-Wen Shao, Chuan-Min Zhai, Jing Wang 0049, Yuan Yan Tang, C. L. Philip Chen |
Neurocomputing | 5 |
| 2016 | A new prospective for Learning Automata: A machine learning approach
Wen Jiang 0001, Bin Li 0002, Shenghong Li 0001, Yuan Yan Tang, C. L. Philip Chen |
Neurocomputing | 4 |
| 2016 | An image classification method that considers privacy-preservation
Chong Wen Liu, Zhaowei Shang, Yuan Yan Tang |
Neurocomputing | 3 |
| 2016 | An efficient level set method based on multi-scale image segmentation and hermite differential operator
Xiaofeng Wang 0009, Hai Min, Le Zou, Yi-Gang Zhang, Yuan Yan Tang, C. L. Philip Chen |
Neurocomputing | 5 |
| 2016 | Some novel approaches on state estimation of delayed neural networks
Kaibo Shi, Xinzhi Liu, Yuan Yan Tang, Hong Zhu 0001, Shouming Zhong |
Inf. Sci. | 3 |
| 2016 | A thermodynamics-inspired feature for anomaly detection on crowd motions in surveillance videos
Xinfeng Zhang 0003, Su Yang 0001, Yuan Yan Tang, Weishan Zhang |
Multim. Tools Appl. | 3 |
| 2016 | Person Re-identification by Exploiting Spatio-Temporal Cues and Multi-view Metric LearningabstractIn this letter, we introduce a new spatio-temporal feature, namely optical flow energy image (OFEI), for video-based person re-identification. OFEI aims to exploit spatio-temporally stable regions across frames, which can capture discriminative cues such as human body parts and carry-on stuffs. Furthermore, we propose a novel matching method, denoted by multi-view relevance metric learning with list-wise constraints (mvRMLLC), to integrate the spatio-temporal (i.e., OFEI) and appearance features. Unlike previous works, mvRMLLC assumes that multiple features are generated from different views with distinct data distributions, while their similarities should be globally consistent. Multiple similarity metrics are then learned and fused by maximizing their global consistency and simultaneously allowing local discrepancies. Extensive experiments on two benchmarks demonstrate that OFEI outperforms the state-of-the-art spatio-temporal features, and mvRMLLC could further enhance the overall performance significantly. Jiaxin Chen 0002, Yunhong Wang 0001, Yuan Yan Tang |
IEEE Signal Process. Lett. | 3 |
| 2016 | Adaptive Multiscale Decomposition of Graph SignalsabstractThis paper proposes an adaptive multiscale decomposition algorithm for graph signals. We develop two types of graph signal cost functions: α-sparsity functional and graph signal entropies, to capture the energy compaction of the signal components. The adaptive decomposition can then be constructed by applying a minimum cost constraint during the full subband decomposition. The proposed adaptive decomposition is shown to outperform graph wavelet decomposition in compressing nonpiecewise constant graph signals. Xianwei Zheng, Yuan Yan Tang, Jianjia Pan, Jiantao Zhou 0001 |
IEEE Signal Process. Lett. | 2 |
| 2016 | Secure Reversible Image Data Hiding Over Encrypted Domain via Key ModulationabstractThis paper proposes a novel reversible image data hiding scheme over encrypted domain. Data embedding is achieved through a public key modulation mechanism, in which access to the secret encryption key is not needed. At the decoder side, a powerful two-class SVM classifier is designed to distinguish encrypted and nonencrypted image patches, allowing us to jointly decode the embedded message and the original image signal. Compared with the state-of-the-art methods, the proposed approach provides higher embedding capacity and is able to perfectly reconstruct the original image as well as the embedded message. Extensive experimental results are provided to validate the superior performance of our scheme. Jiantao Zhou 0001, Weiwei Sun 0009, Li Dong 0006, Xianming Liu 0005, Oscar C. Au, Yuan Yan Tang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | A Manifold Alignment Approach for Hyperspectral Image Visualization With Natural ColorabstractThe trichromatic visualization of hundreds of bands in a hyperspectral image (HSI) has been an active research topic. The visualized image shall convey as much information as possible from the original data and facilitate easy image interpretation. However, most existing methods display HSIs in false color, which contradicts with user experience and expectation. In this paper, we propose a new framework for visualizing an HSI with natural color by the fusion of an HSI and a high-resolution color image via manifold alignment. Manifold alignment projects several data sets to a shared embedding space where the matching points between them are pairwise aligned. The embedding space bridges the gap between the high-dimensional spectral space of the HSI and the RGB space of the color image, making it possible to transfer natural color and spatial information in the color image to the HSI. In this way, a visualized image with natural color distribution and fine spatial details can be generated. Another advantage of the proposed method is its flexible data setting for various scenarios. As our approach only needs to search a limited number of matching pixel pairs that present the same object, the HSI and the color image can be captured from the same or semantically similar sites. Moreover, the learned projection function from the hyperspectral data space to the RGB space can be directly applied to other HSIs acquired by the same sensor to achieve a quick overview. Our method is also able to visualize user-specified bands as natural color images, which is very helpful for users to scan bands of interest. Danping Liao, Yuntao Qian, Jun Zhou 0001, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2016 | Hyperspectral Image Classification Based on Spectral-Spatial One-Dimensional Manifold EmbeddingabstractA novel approach called Spectral-Spatial 1-D Manifold Embedding (SS1DME) is proposed in this paper for remotely sensed hyperspectral image (HSI) classification. This novel approach is based on a generalization of the recently developed smooth ordering model, which has gathered a great interest in the image processing area. In the proposed approach, first, we employ the spectral-spatial information-based affinity metric to learn the similarity of HSI pixels, where the contextual information is encoded into the affinity metric using spatial information. In our derived model, based on the obtained affinity metric, the created multiple 1-D manifold embeddings (1DMEs) consist of several different versions of 1DME of the same set of all HSI points. Since each 1DME of the data is a 1-D sequence, a label function on the data can be obtained by applying the simple 1-D signal processing tools (such as interpolation/regression). By collecting the predicted labels from these label functions, we build a subset of the current unlabeled points, on which the labels are correctly labeled with high confidence. Next, we add a proportion of the elements from this subset to the original labeled set to get the updated labeled set, which is used for the next running instance. Repeating this process for several loops, we get an extended labeled set, where the new members are correctly labeled by the label functions with much high confidence. Finally, we utilize the extended labeled set to build the target classifier for the whole HSI pixels. In the whole process, 1DME plays the role of learning data features from the given affinity metric. With the incrementation of learning features during iteration, the proposed scheme will gradually approximate the exact labels of all sample points. The proposed scheme is experimentally demonstrated using four real HSI data sets, exhibiting promising classification performance when compared with other recently introduced spatial analysis alternatives. Huiwu Luo, Yuan Yan Tang, Yulong Wang 0002, Jianzhong Wang 0004, Chunli Li, Tingbo Hu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | SIFT Keypoint Removal and Injection via Convex RelaxationabstractScale invariant feature transform (SIFT), as one of the most popular local feature extraction algorithms, has been widely employed in many computer vision and multimedia security applications. Although SIFT has been extensively investigated from various perspectives, its security against malicious attacks has rarely been discussed. In this paper, we show that the SIFT keypoints can be effectively removed with minimized distortion on the processed image. The SIFT keypoint removal is formulated as a constrained optimization problem, where the constraints are carefully designed to suppress the existence of local extrema and prevent generating new keypoints within a local cuboid in the scale space. To hide the traces of performing SIFT keypoint removal, we then propose to inject a large number of fake SIFT keypoints into the previously cleaned image with minimized distortion. As demonstrated experimentally, our proposed SIFT removal and injection algorithms significantly outperform the state-of-the-art techniques. Furthermore, it is shown that the combined SIFT keypoint removal and injection attack strategy is capable of defeating the most powerful forensic detector designed for SIFT keypoint removal. Our results suggest that an authorization mechanism is required for SIFT-based systems to verify the validity of the input data, so as to achieve high reliability. Yuanman Li, Jiantao Zhou 0001, An Cheng, Xianming Liu 0005, Yuan Yan Tang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2016 | Connected Component Model for Multi-Object TrackingabstractIn multi-object tracking, it is critical to explore the data associations by exploiting the temporal information from a sequence of frames rather than the information from the adjacent two frames. Since straightforwardly obtaining data associations from multi-frames is an NP-hard multi-dimensional assignment (MDA) problem, most existing methods solve this MDA problem by either developing complicated approximate algorithms, or simplifying MDA as a 2D assignment problem based upon the information extracted only from adjacent frames. In this paper, we show that the relation between associations of two observations is the equivalence relation in the data association problem, based on the spatial-temporal constraint that the trajectories of different objects must be disjoint. Therefore, the MDA problem can be equivalently divided into independent subproblems by equivalence partitioning. In contrast to existing works for solving the MDA problem, we develop a connected component model (CCM) by exploiting the constraints of the data association and the equivalence relation on the constraints. Based upon CCM, we can efficiently obtain the global solution of the MDA problem for multi-object tracking by optimizing a sequence of independent data association subproblems. Experiments on challenging public data sets demonstrate that our algorithm outperforms the state-of-the-art approaches. Zhenyu He 0001, Xin Li 0034, Xinge You, Dacheng Tao, Yuan Yan Tang |
IEEE Trans. Image Process. | 5 |
| 2016 | Quaternionic Local Ranking Binary Pattern: A Local Descriptor of Color ImagesabstractThis paper proposes a local descriptor called quaternionic local ranking binary pattern (QLRBP) for color images. Different from traditional descriptors that are extracted from each color channel separately or from vector representations, QLRBP works on the quaternionic representation (QR) of the color image that encodes a color pixel using a quaternion. QLRBP is able to handle all color channels directly in the quaternionic domain and include their relations simultaneously. Applying a Clifford translation to QR of the color image, QLRBP uses a reference quaternion to rank QRs of two color pixels, and performs a local binary coding on the phase of the transformed result to generate local descriptors of the color image. Experiments demonstrate that the QLRBP outperforms several state-of-the-art methods. Rushi Lan, Yicong Zhou, Yuan Yan Tang |
IEEE Trans. Image Process. | 3 |
| 2016 | Person Re-Identification by Dual-Regularized KISS Metric LearningabstractPerson re-identification aims to match the images of pedestrians across different camera views from different locations. This is a challenging intelligent video surveillance problem that remains an active area of research due to the need for performance improvement. Person re-identification involves two main steps: feature representation and metric learning. Although the keep it simple and straightforward (KISS) metric learning method for discriminative distance metric learning has been shown to be effective for the person re-identification, the estimation of the inverse of a covariance matrix is unstable and indeed may not exist when the training set is small, resulting in poor performance. Here, we present dual-regularized KISS (DR-KISS) metric learning. By regularizing the two covariance matrices, DR-KISS improves on KISS by reducing overestimation of large eigenvalues of the two estimated covariance matrices and, in doing so, guarantees that the covariance matrix is irreversible. Furthermore, we provide theoretical analyses for supporting the motivations. Specifically, we first prove why the regularization is necessary. Then, we prove that the proposed method is robust for generalization. We conduct extensive experiments on three challenging person re-identification datasets, VIPeR, GRID, and CUHK 01, and show that DR-KISS achieves new state-of-the-art performance. Dapeng Tao, Yanan Guo 0003, Mingli Song, Yaotang Li, Zhengtao Yu 0001, Yuan Yan Tang |
IEEE Trans. Image Process. | 6 |
| 2016 | Face Aging Effect Simulation Using Hidden Factor Analysis Joint Sparse RepresentationabstractFace aging simulation has received rising investigations nowadays, whereas it still remains a challenge to generate convincing and natural age-progressed face images. In this paper, we present a novel approach to such an issue using hidden factor analysis joint sparse representation. In contrast to the majority of tasks in the literature that integrally handle the facial texture, the proposed aging approach separately models the person-specific facial properties that tend to be stable in a relatively long period and the age-specific clues that gradually change over time. It then transforms the age component to a target age group via sparse reconstruction, yielding aging effects, which is finally combined with the identity component to achieve the aged face. Experiments are carried out on three face aging databases, and the results achieved clearly demonstrate the effectiveness and robustness of the proposed method in rendering a face with aging effects. In addition, a series of evaluations prove its validity with respect to identity preservation and aging effect generation. Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001, Yuan Yan Tang |
IEEE Trans. Image Process. | 5 |
| 2016 | Learning Proximity Relations for Feature SelectionabstractThis work presents a feature selection method based on proximity relations learning. Each single feature is treated as a binary classifier that predicts for any three objects X, A, and B whether X is close to A or B. The performance of the classifier is a direct measure of feature quality. Any linear combination of feature-based binary classifiers naturally corresponds to feature selection. Thus, the feature selection problem is transformed into an ensemble learning problem of combining many weak classifiers into an optimized strong classifier. We provide a theoretical analysis of the generalization error of our proposed method which validates the effectiveness of our proposed method. Various experiments are conducted on synthetic data, four UCI data sets and 12 microarray data sets, and demonstrate the success of our approach applying to feature selection. A weakness of our algorithm is high time complexity. Taiping Zhang, Yuan Yan Tang, C. L. Philip Chen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2016 | Sketch-Based Image Retrieval by Salient Contour ReinforcementabstractThe paper presents a sketch-based image retrieval algorithm. One of the main challenges in sketch-based image retrieval (SBIR) is to measure the similarity between a sketch and an image. To tackle this problem, we propose an SBIR-based approach by salient contour reinforcement. In our approach, we divide the image contour into two types. The first is the global contour map. The second, called the salient contour map, is helpful to find out the object in images similar to the query. In addition, based on the two contour maps, we propose a new descriptor, namely an angular radial orientation partitioning (AROP) feature. It fully utilizes the edge pixels' orientation information in contour maps to identify the spatial relationships. Our AROP feature based on the two candidate contour maps is both efficient and effective to discover false matches of local features between sketches and images, and can greatly improve the retrieval performance. The application of the retrieval system based on this algorithm is established. The experiments on the image dataset with 0.3 million images show the effectiveness of the proposed method and comparisons with other algorithms are also given. Compared to baseline performance, the proposed method achieves 10% higher precision in top 5. Yuting Zhang 0007, Xueming Qian, Xianglong Tan, Junwei Han 0001, Yuan Yan Tang |
IEEE Trans. Multim. | 5 |
| 2015 | Crowd Motion Monitoring with Thermodynamics-Inspired FeatureabstractCrowd motion in surveillance videos is comparable to heat motion of basic particles. Inspired by that, we introduce Boltzmann Entropy to measure crowd motion in optical flow field so as to detect abnormal collective behaviors. As a result, the collective crowd moving pattern can be represented as a time series. We found that when most people behave anomaly, the entropy value will increase drastically. Thus, a threshold can be applied to the time series to identify abnormal crowd commotion in a simple and efficient manner without machine learning. The experimental results show promising performance compared with the state of the art methods. The system works in real time with high precision. Xinfeng Zhang 0003, Su Yang 0001, Yuan Yan Tang, Weishan Zhang |
AAAI | 3 |
| 2015 | Robust Discriminative Nonnegative Patch Alignment for Occluded Face Recognition
Weihua Ou, Gai Li, Shujian Yu, Fujia Ren, Yuan Yan Tang |
ICONIP (4) | 6 |
| 2015 | Webcam-Based Visual Gaze Estimation Under Desktop Environment
Shujian Yu, Weihua Ou, Xinge You, Xiubao Jiang, Yi Mou, Weigang Guo, Yuan Yan Tang, C. L. Philip Chen |
ICONIP (2) | 8 |
| 2015 | Generalized Kernel Normalized Mixed-Norm Algorithm: Analysis and Simulations
Shujian Yu, Xinge You, Xiubao Jiang, Weihua Ou, Yixiao Zhao, C. L. Philip Chen, Yuan Yan Tang |
ICONIP (2) | 8 |
| 2015 | Kernel normalized mixed-norm algorithm for system identificationabstractKernel methods provide an efficient nonparametric model to produce adaptive nonlinear filtering (ANF) algorithms. However, in practical applications, standard squared error based kernel methods suffer from two main issues: (1) a constant step size is used, which degrades the algorithm performance in non-stationary environment, and (2) additive noises are assumed to follow Gaussian distribution, while in practice the noises are generally non-Gaussian and follow other statistical distributions. To address these two issues simultaneously, this paper proposes a novel kernel normalized mixed-norm (KNMN) algorithm. Compared to the standard squared error based kernel methods, the KNMN algorithm extends the linear mixed-norm adaptive filtering algorithms to Reproducing Kernel Hilbert Space (RKHS) and introduces a normalized step size as well as adaptive mixing parameter. We also conduct the mean square convergence analysis and demonstrate the desirable performance of the KNMN algorithm in solving the system identification problem. Shujian Yu, Xinge You, Weihua Ou, Yuan Yan Tang |
IJCNN | 5 |
| 2015 | Human Heart Rate Estimation Using Ordinary Cameras under Natural MovementabstractNon-contact face-video based human heart rate (HR) estimation has attracted a lot of attentions in recent years. Almost all the state-of-the-art webcam or smartphone based HR estimation methods comprise three main steps: firstly, a region of interest (ROI) on the human face is detected in each video frame, then, the target signal is obtained by fusing multiple raw traces, which are extracted from the RGB channels across all the video frames, finally, HR is estimated by applying frequency analysis approach to the target signal. However, three major drawbacks impede the applicability of the current methods: (1) the performance of ROI detection is susceptible to head motion and facial expression, (2) there is still a lack of well-accepted method for fusing raw traces to form the target signal, and (3) the adopted frequency analysis approaches always provide estimation results with low resolution and high side lobes. To address these issues, we propose a novel HR estimation method which is applicable to ordinary cameras subject to natural head movement or facial expression. The proposed method features ROI detection via facial feature detection and tracking, target signal extraction via Independent Component Analysis (ICA) in the RGB channels, and HR estimation via real-valued iterative adaptive approach (RIAA). Experimental results validate the superiority of our proposed method. Shujian Yu, Xinge You, Xiubao Jiang, Yi Mou, Weihua Ou, Yuan Yan Tang, C. L. Philip Chen |
SMC | 7 |
| 2015 | NNMap: A method to construct a good embedding for nearest neighbor classification
Jing Chen 0008, Yuan Yan Tang, C. L. Philip Chen, Bin Fang 0001, Zhaowei Shang, Yuewei Lin |
Neurocomputing | 2 |
| 2015 | ApLeaf: An efficient android-based plant leaf identification system
Zhong-Qiu Zhao, Lin-Hai Ma, Yiu-Ming Cheung, Xindong Wu 0001, Yuan Yan Tang, C. L. Philip Chen |
Neurocomputing | 5 |
| 2015 | Negative samples reduction in cross-company software defects prediction
Lin Chen 0023, Bin Fang 0001, Zhaowei Shang, Yuan Yan Tang |
Inf. Softw. Technol. | 4 |
| 2015 | Feature Guided Biased Gaussian Mixture Model for image matching
Kun Sun 0002, Wenbing Tao, Yuan Yan Tang |
Inf. Sci. | 4 |
| 2015 | A novel item anomaly detection approach against shilling attacks in collaborative recommendation systems using the dynamic time interval segmentation technique
Bin Fang 0001, Yuan Yan Tang |
Inf. Sci. | 5 |
| 2015 | Learning With Hypergraph for Hyperspectral Image Feature ExtractionabstractIt is known that hyperspectral image (HSI) classification is a high-dimension low-sample-size problem. To ease this problem, one natural idea is to take the feature extraction as a preprocessing. A graph embedding model is a classic family of feature extraction methods, which preserves certain statistical or geometric properties of the data set. However, the graph embedding model considers only the pairwise relationship between two vertices, which cannot represent the complex relationships of the data. Utilizing the spatial structure of HSI, in this letter, we propose a spatial hypergraph embedding model for feature extraction. Experimental results demonstrate that our method outperforms many existing feature extract methods for HSI classification. Yuan Yan Tang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Mining near duplicate image groups
Jing Li 0049, Xueming Qian, Qing Li 0001, Yisi Zhao, Yuan Yan Tang |
Multim. Tools Appl. | 6 |
| 2015 | A Fractal Dimension and Wavelet Transform Based Method for Protein Sequence Similarity AnalysisabstractOne of the key tasks related to proteins is the similarity comparison of protein sequences in the area of bioinformatics and molecular biology, which helps the prediction and classification of protein structure and function. It is a significant and open issue to find similar proteins from a large scale of protein database efficiently. This paper presents a new distance based protein similarity analysis using a new encoding method of protein sequence which is based on fractal dimension. The protein sequences are first represented into the 1-dimensional feature vectors by their biochemical quantities. A series of Hybrid method involving discrete Wavelet transform, Fractal dimension calculation (HWF) with sliding window are then applied to form the feature vector. At last, through the similarity calculation, we can obtain the distance matrix, by which, the phylogenic tree can be constructed. We apply this approach by analyzing the ND5 (NADH dehydrogenase subunit 5) protein cluster data set. The experimental results show that the proposed model is more accurate than the existing ones such as Su's model, Zhang's model, Yao's model and MEGA software, and it is consistent with some known biological facts. Yuan Yan Tang, Yang Lu 0009, Huiwu Luo |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | Landmark Summarization With Diverse ViewpointsabstractLandmark summarization with diverse viewpoints is very important in landmark retrieval, as it can create a comprehensive description of a landmark for users. In this paper we present an approach for summarizing a collection of landmark images from diverse viewpoints. First, we group landmark images with content overlap by viewpoint album (VA) generation. Second, we model the relative viewpoint of each image within the VA based on the spatial layout of distinctive descriptors of a landmark. Third, we express the relative viewpoint of an image with a 4-D viewpoint vector, including horizontal, vertical, scale, and rotation. Finally, we summarize the landmarks in terms of viewpoints. Experimental results show the effectiveness of the proposed landmark summarization approach. Xueming Qian, Xiyu Yang, Yuan Yan Tang, Xingsong Hou, Tao Mei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Structural Atomic Representation for ClassificationabstractRecently, a large family of representation-based classification methods have been proposed and attracted great interest in pattern recognition and computer vision. This paper presents a general framework, termed as atomic representation-based classifier (ARC), to systematically unify many of them. By defining different atomic sets, most popular representation-based classifiers (RCs) follow ARC as special cases. Despite good performance, most RCs treat test samples separately and fail to consider the correlation between the test samples. In this paper, we develop a structural ARC (SARC) based on Bayesian analysis and generalizing a Markov random field-based multilevel logistic prior. The proposed SARC can utilize the structural information among the test data to further improve the performance of every RC belonging to the ARC framework. The experimental results on both synthetic and real-database demonstrate the effectiveness of the proposed framework. Yuan Yan Tang, Yulong Wang 0002, Luoqing Li, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2015 | The Generalization Ability of SVM Classification Based on Markov SamplingabstractUNLABELLED: The previously known works studying the generalization ability of support vector machine classification (SVMC) algorithm are usually based on the assumption of independent and identically distributed samples. In this paper, we go far beyond this classical framework by studying the generalization ability of SVMC based on uniformly ergodic Markov chain (u.e.M.c.) samples. We analyze the excess misclassification error of SVMC based on u.e.M.c. samples, and obtain the optimal learning rate of SVMC for u.e.M.c. SAMPLES: We also introduce a new Markov sampling algorithm for SVMC to generate u.e.M.c. samples from given dataset, and present the numerical studies on the learning performance of SVMC based on Markov sampling for benchmark datasets. The numerical studies show that the SVMC based on Markov sampling not only has better generalization ability as the number of training samples are bigger, but also the classifiers based on Markov sampling are sparsity when the size of dataset is bigger with regard to the input dimension. Jie Xu 0006, Yuan Yan Tang, Bin Zou 0002, Zongben Xu, Luoqing Li, Yang Lu 0009, Baochang Zhang 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | Hyperspectral Image Classification Based on Three-Dimensional Scattering Wavelet TransformabstractRecent research has shown that utilizing the spectral-spatial information can improve the performance of hyperspectral image (HSI) classification. Since HSI is a 3-D cube datum, 3-D spatial filtering becomes a simple and effective method for extracting the spectral-spatial information. In this paper, we propose a 3-D scattering wavelet transform, which filters the HSI cube data with a cascade of wavelet decompositions, complex modulus, and local weighted averaging. The scattering feature can adequately capture the spectral-spatial information for classification. In the classification step, a support vector machine based on Gaussian kernel is used as a classifier due to its capability to deal with high-dimensional data. Our method is fully evaluated on four classic HSIs, i.e., Indian Pines, Pavia University, Botswana, and Kennedy Space Center. The classification results show that our method achieves as high as 94.46%, 99.30%, 97.57%, and 95.20% accuracies, respectively, when only 5% of the total samples per class is labeled. Yuan Yan Tang, Yang Lu 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Weighted Couple Sparse Representation With Classified Regularization for Impulse Noise RemovalabstractMany impulse noise (IN) reduction methods suffer from two obstacles, the improper noise detectors and imperfect filters they used. To address such issue, in this paper, a weighted couple sparse representation model is presented to remove IN. In the proposed model, the complicated relationships between the reconstructed and the noisy images are exploited to make the coding coefficients more appropriate to recover the noise-free image. Moreover, the image pixels are classified into clear, slightly corrupted, and heavily corrupted ones. Different data-fidelity regularizations are then accordingly applied to different pixels to further improve the denoising performance. In our proposed method, the dictionary is directly trained on the noisy raw data by addressing a weighted rank-one minimization problem, which can capture more features of the original data. Experimental results demonstrate that the proposed method is superior to several state-of-the-art denoising methods. C. L. Philip Chen, Licheng Liu, Long Chen 0001, Yuan Yan Tang, Yicong Zhou |
IEEE Trans. Image Process. | 4 |
| 2015 | Robust Face Recognition via Minimum Error Entropy-Based Atomic RepresentationabstractRepresentation-based classifiers (RCs) have attracted considerable attention in face recognition in recent years. However, most existing RCs use the mean square error (MSE) criterion as the cost function, which relies on the Gaussianity assumption of the error distribution and is sensitive to non-Gaussian noise. This may severely degrade the performance of MSE-based RCs in recognizing facial images with random occlusion and corruption. In this paper, we present a minimum error entropy-based atomic representation (MEEAR) framework for face recognition. Unlike existing MSE-based RCs, our framework is based on the minimum error entropy criterion, which is not dependent on the error distribution and shown to be more robust to noise. In particular, MEEAR can produce discriminative representation vector by minimizing the atomic norm regularized Renyi's entropy of the reconstruction error. The optimality conditions are provided for general atomic representation model. As a general framework, MEEAR can also be used as a platform to develop new classifiers. Two effective MEE-based RCs are proposed by defining appropriate atomic sets. The experimental results on popular face databases show that MEEAR can improve both the recognition accuracy and the reconstructed results compared with the state-of-the-art MSE-based RCs. Yulong Wang 0002, Yuan Yan Tang, Luoqing Li |
IEEE Trans. Image Process. | 2 |
| 2015 | The Generalization Ability of Online SVM Classification Based on Markov SamplingabstractIn this paper, we consider online support vector machine (SVM) classification learning algorithms with uniformly ergodic Markov chain (u.e.M.c.) samples. We establish the bound on the misclassification error of an online SVM classification algorithm with u.e.M.c. samples based on reproducing kernel Hilbert spaces and obtain a satisfactory convergence rate. We also introduce a novel online SVM classification algorithm based on Markov sampling, and present the numerical studies on the learning ability of online SVM classification based on Markov sampling for benchmark repository. The numerical studies show that the learning performance of the online SVM classification algorithm based on Markov sampling is better than that of classical online SVM classification based on random sampling as the size of training samples is larger. Jie Xu 0006, Yuan Yan Tang, Bin Zou 0002, Zongben Xu, Luoqing Li, Yang Lu 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Robust Nonnegative Patch Alignment for Dimensionality ReductionabstractDimensionality reduction is an important method to analyze high-dimensional data and has many applications in pattern recognition and computer vision. In this paper, we propose a robust nonnegative patch alignment for dimensionality reduction, which includes a reconstruction error term and a whole alignment term. We use correntropy-induced metric to measure the reconstruction error, in which the weight is learned adaptively for each entry. For the whole alignment, we propose locality-preserving robust nonnegative patch alignment (LP-RNA) and sparsity-preserviing robust nonnegative patch alignment (SP-RNA), which are unsupervised and supervised, respectively. In the LP-RNA, we propose a locally sparse graph to encode the local geometric structure of the manifold embedded in high-dimensional space. In particular, we select large p -nearest neighbors for each sample, then obtain the sparse representation with respect to these neighbors. The sparse representation is used to build a graph, which simultaneously enjoys locality, sparseness, and robustness. In the SP-RNA, we simultaneously use local geometric structure and discriminative information, in which the sparse reconstruction coefficient is used to characterize the local geometric structure and weighted distance is used to measure the separability of different classes. For the induced nonconvex objective function, we formulate it into a weighted nonnegative matrix factorization based on half-quadratic optimization. We propose a multiplicative update rule to solve this function and show that the objective function converges to a local optimum. Several experimental results on synthetic and real data sets demonstrate that the learned representation is more discriminative and robust than most existing dimensionality reduction methods. Xinge You, Weihua Ou, C. L. Philip Chen, Qiang Li 0024, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2014 | Estimation of capacity parameters for dynamic histogram shifting (DHS)-based reversible image watermarkingabstractDynamic histogram shifting (DHS) is a generation of the conventional histogram shifting (HS) technique for reversible image watermarking. Its superior embedding performance is achieved at the cost of significantly increased computational burden incurred by estimating the capacity parameters via multi-rounds of embedding iterations. In this work, we propose an analytical framework on estimating the optimal capacity parameters for DHS-based reversible image watermarking. We demonstrate that such parameter estimation can be cast as a convex optimization problem, which can be numerically solved in an efficient manner. The estimated values can then be utilized to facilitate a local search algorithm to obtain the truly optimal ones with much lowered complexity. Experimental results are provided to verify the validity of our findings. Li Dong 0006, Jiantao Zhou 0001, Yuan Yan Tang, Xianming Liu 0005 |
ICME | 3 |
| 2014 | Person reidentification using quaternionic local binary patternabstractPerson reidentification is to identify the persons observed in nonoverlapping camera networks. Most existing methods usually extract features from the red, green, and blue color channels of images individually. They, however, neglect the connections between each color component in the image. To overcome this problem, a novel quaternionic local binary pattern (QLBP) is proposed for person reidentification in this paper. In the proposed QLBP, each pixel in a color image is represented by a quaternion so that we can handle all color components in a holistic way. A novel pseudo-rotation of quaternion (PRQ) is proposed to rank two quaternions. Some properties of PRQ are also discussed. After a QLBP coding, the local histograms are extracted and used as features. Experiments on two public benchmarking datasets, ETHZ and i-LIDS MCTS, are carried out to evaluate the QLBP performance. Comparison results show that the QLBP outperforms several stat-of-art methods for person reidentification. Rushi Lan, Yicong Zhou, Yuan Yan Tang, C. L. Philip Chen |
ICME | 3 |
| 2014 | Dual Fuzzy Hypergraph Regularized Multi-label Learning for Protein Subcellular Location PredictionabstractWith the explosion of newly found proteins, it is necessary and urgent to develop automated computational methods for protein sub cellular location prediction. In particular, the problem of predictor construction for multi-location proteins is challenging. Considering the main limitations of the existing methods, we propose a hierarchical multi-label learning model FHML for both single-location proteins and multi-location proteins. In this model, feature space is firstly decomposed onto a set of nonnegative bases under the nonnegative data factorization framework. The nonnegative bases act as latent feature concepts and the corresponding coefficients on these bases are views as the new feature representation on the latent feature concepts. The similar decomposition is later performed in label space, and then the latent label concepts are extracted. Using these latent concepts as hyper edges, we construct dual fuzzy hyper graphs to exploit the intrinsic high-order relations embedded in both feature space and label space. Finally, the sub cellular location annotation information is propagated from the labeled proteins to the unlabeled proteins by performing dual fuzzy hyper graph Laplacian regularization. In this work, our proposed method is evaluated on eukaryotic protein benchmark dataset, and the experimental results have shown its effectiveness. Jing Gien, Yuan Yan Tang, C. L. Philip Chen, Yuewei Lin |
ICPR | 2 |
| 2014 | Multi-scale Tensor l1-Based Algorithm for Hyperspectral Image ClassificationabstractSparsity-based model has been successfully applied in hyper spectral image classification. However, previous ℓ1-based method fails to consider the spatial structure of each pixel. In this paper, we generalize the ℓ1-based method to its tensor form, which takes full advantage of the spatial structure of the pixel. To optimize the scale of the spatial structure, a multi-scale fusion framework based on the ensemble learning method is proposed to further improve classification performance. Experimental results demonstrate that our proposed method can achieve state-of-the-art classification performance. Yuan Yan Tang |
ICPR | 2 |
| 2014 | Impulse noise removal using sparse representation with fuzzy weightsabstractMany impulse noise removal algorithms do not reach good denoising performance mainly due to the imperfect filters they adopted. In this paper, the popular used sparse representation model is extended for impulse noise removal by using a fuzzy weight matrix. This fuzzy weight is used to describe the noise-like level of the current pixel, and to determine how much information of this pixel should be used in the sparse land model. Besides, a regularization term which counts the proximity between the reconstructed image and the noisy image is also added into the sparse model. This makes the proposed model more robust to the noise detector which generates the fuzzy weight matrix. Moreover, unlike other sparse model, the dictionary used in our model is trained from some reference images that keep the similar structure information of the original image. Therefore, it is more suitable for reconstructing the original image. Simulation results show that our method is superior to all the tested state-of-the-art impulse noise removal methods. Licheng Liu, C. L. Philip Chen, Yicong Zhou, Yuan Yan Tang |
SMC | 4 |
| 2014 | Cross-scene learning: Improving classification performance by using "gray sample"abstractMost object classification models considered in image only exist positive samples and negative samples. In this paper, another type of sample exists, named “gray sample”, which belongs to neither positive samples nor negative samples, contains knowledge in other domains or scenes. The degree of “gray” is defined by semantic similarity between annotations and scenes. While the local descriptors represent one object by visual feature, and the annotations are semantic description of this object. One class of objects has exclusive concept and different descriptions in different scenes, “gray samples” is belong to one concept, but exist in other scenes by different manifestation. Only the similar scenes have commonality, which is a bridge connect knowledge in different scenes. To achieve goal of cross-scene learning and using “gray sample”, the similar degree of scenes is needed to be conducted, by computing the co-occurrence probability of annotations. In our model, following the bags-of-features (BoF) approach, a plenty of local descriptors are extracted from the annotation areas of images, the visual words are got through clustering those descriptors, and the proposed model is built upon the visual words. Using EM algorithm, construct model under different thresholds of correlation degree, and classify objects. The thresholds are important in decision whether the “gray samples” are suitable for training data, and key factor in classification performance. The experiments of object classification based on LabelMe dataset, the proposed model exhibits superior performances compared to the other existing methods. Chong Wen Liu, Zhaowei Shang, Yuan Yan Tang |
SMC | 3 |
| 2014 | Automatic decision support by information energy decision tree algorithmabstractThe application of information entropy to decision tree algorithms has been shown to produce very accurate classifiers. Information entropy is utilized to ensure that the average distance of paths from the non-leaf node to each descendant leaf node of the decision tree is shortest. Therefore, it works well for data set which covers all the underlying rules. But it is lack of prediction ability when the training data set can not cover all the underlying rules. In this paper, we propose a novel indicator, information energy, to generate decision tree. Information energy describes the distance from the current state of a data set to its balance state. Proper selection of attribute can divide a data set into a state of higher information energy and produce classification rules of prediction ability. A generator of random sample sets and rules is designed to provide synthetic samples for experimental verification. Experimental results show that information energy outperforms information entropy in both speed and accuracy when the training data set can not cover all the underlying rules. Runzong Liu, Yuan Yan Tang, Bin Fang 0001 |
SMC | 2 |
| 2014 | A novel method for protein structure retrieval using tableau representation and sparse codingabstractProtein retrieval is a difficult task and has become a hot issue recently due to the complex structure and large data size of proteins. This work helps biologists investigate the link between structure and function of a protein in a deeper level and can be used in lots of biomedical applications. The retrieval system gives scores to all proteins in the database, e.g. SCOP or PDB, by given a query protein to compare with them. In this paper, we propose a novel algorithm based on sparse coding to retrieve proteins in the database using tableau representation. Both unsupervised and supervised methods are studied in the proposed algorithm where the sparse coefficient is regarded as similarity measurement. Experiments are conducted on ASTRAL 1.73 95% database and show that the proposed algorithms can improve the original feature extraction method which only uses cosine similarity. Yang Lu 0009, Yulong Wang 0002, Huiwu Luo, Yuan Yan Tang |
SMC | 6 |
| 2014 | Incorporating local and global geometric structure for hyperspectral image classificationabstractThe highly correlated data structure makes the computational cost of hyperspectral image (HSI) much complex. The need of effective processing and analyzing of HSI has met many difficulties and become an open topic in the community of high dimensional data analysis. Local structure has shown great efficiency in feature extraction. Yet recent progress has also demonstrated the importance of global geometric structure in discriminant analysis. Thus, both the locality and global geometric structure are critical for dimension reduction. In this paper, a novel linear supervised dimensionality reduction algorithm, called Locality and Global Geometric Structure Preserving (LGGSP) projection, is proposed for dimension reduction. LGGSP encodes not only the local discriminant information into the optimal objective functions, but also the global margin information. To be specific, two adjacent graph (viz., similarity matrix and variance matrix), are constructed to detect the local intrinsic structure, simultaneously, a graph matrix to capture the global margin of different classes. Experimental results on both benchmark data sets and the real hyperspectral image data set demonstrate the effectiveness and practicability of proposed scheme. Huiwu Luo, Yuan Yan Tang |
SMC | 2 |
| 2014 | Spectral-spatial hyperspectral image destriping using low-rank representation and Huber-Markov random fieldsabstractThis paper presents a novel spectral-spatial destriping method for hyperspectral images. The ubiquitous striping noise in hyperspectral images might degrade the quality of the imagery and bring difficulties in hyperspectral data processing. Although numerous methods have been proposed for striping noise reduction recently, most of them fail to consider the spectral correlation and spatial information of the hyperspectral images simultaneously. In order to remedy this drawback, the proposed method integrates the spectral and spatial information to remove the striping noise in the hyperspectral images. To this end, firstly, the low-rank representation (LRR) is used to take advantage of the spectral information. Then, the spatial information is included using a Huber-Markov random field (MRF) prior model, which is convex and can well preserve the edge and texture information while removing the noise. The experimental results on simulated and real hyperspectral data sets demonstrate the effectiveness of the proposed method. Yulong Wang 0002, Yuan Yan Tang, Huiwu Luo, Yang Lu 0009 |
SMC | 2 |
| 2014 | Protein sequence analysis based on fractal-wavelet schemeabstractIt is a significant issue to find similar proteins from a large scale of protein database efficiently. This paper presents a new algorithm of protein sequence which is based on fractal dimension and wavelet transform. A hybrid method consisting fractal dimension calculation, discrete wavelet transform and sliding window are applied to generate a new encoding feature. Through the computation between the feature vectors, we can obtain the distance matrix and the phylogenic tree can be constructed.We apply this approach by analyzing the ND5 (NADH dehydrogenase subunit 5) protein dataset. The experimental results show that the proposed model is more accurate than the existing ones such as Su's model, Zhang's model and Yao's model, and it is consistent with the result generated from MEGA software and some known facts. Yuan Yan Tang, Yang Lu 0009, Huiwu Luo, Yulong Wang 0002 |
SMC | 2 |
| 2014 | Feature extraction based on kernel sparse representation for hyperspectral image classificationabstractFeature extraction is a promising technique for hyperspectral image classification. Recent research has shown that the criterion of sparse representation classification (SRC) can help to design a feature extraction method. This method is called the SRC steered discriminative projection (SRCDP). Motivated by the fact that kernel trick can exploit the nonlinear case of features, this paper generalizes SRCDP to its kernel case named KSRCDP. Extensive experiments show that KSRCDP can obtain excellent classification performance on two classic hyperspectral images. Huiwu Luo, Yang Lu 0009, Yulong Wang 0002, Yuan Yan Tang |
SMC | 6 |
| 2014 | Multiview Hessian discriminative sparse coding for image annotation
Weifeng Liu 0001, Dacheng Tao, Jun Cheng 0002, Yuan Yan Tang |
Comput. Vis. Image Underst. | 4 |
| 2014 | On the circular-l(2, 1)-labelling for strong products of paths and cyclesabstractLet k be a positive integer. A k ‐circular‐ L (2, 1)‐labelling of a graph G is an assignment f from V ( G ) to {0, 1, …, k −1} such that, for any two vertices u and v , | f ( u ) − f ( v )| k ≥ 2 if u and v are adjacent, and | f ( u ) − f ( v )| k ≥ 1 if u and v are at distance 2, where | x | k = min{| x |, k −| x |}. The minimum k such that G admits a k ‐circular‐ L (2, 1)‐labelling is called the circular‐ L (2, 1)‐labelling number (or just the σ ‐number) of G , denoted by σ ( G ). The exact values of σ ( P m ⊠ C n ) and σ ( C m ⊠ C n ) for some m and n have been determined in this study. Finally, it has been concluded that σ ( C m ⊠ C n ) ≤ 13 for n ≥ m ≥ 220. Yuan Yan Tang, Zehui Shao, Fangnian Lang, Xiaodong Xu 0006, Roger K. Yeh |
IET Commun. | 1 |
| 2014 | Active contours with a joint and region-scalable distribution metric for interactive natural image segmentationabstractIn this study, we present an efficient active contour with a joint and region‐scalable distribution metric for interactive natural image segmentation. First, the authors project a red–green–blue image into the CIELab colour space and employ independent component analysis to select two subspace channels. Then, by initialising the evolving curve interactively in terms of a polygonal curve or multiple polygonal curves, they compute a joint probability distribution associated with a region‐scalable mask to model the regional statistics and propose a simple but effective distribution metric to regularise the active contours. Subsequently, they convert the resultant level set function into binary pattern and find the larger 8‐connected regions as the desired objects. Finally, the selected regions are smoothed with a circular averaging filter such that the final segmentation results can be obtained. The proposed approach not only can deal with the complex appearance and intensity in homogeneity, but also has the advantages of fast convergence and easy implementation. The experiments have shown the precise and reliable segmentation results in comparison with the state‐of‐the‐art competing approaches. Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Yuan Yan Tang, Jixiang Du |
IET Image Process. | 4 |
| 2014 | Recognizing complex events in real movies by combining audio and video features
Jixiang Du, Chuan-Min Zhai, Yi-Lan Guo, Yuan Yan Tang, C. L. Philip Chen |
Neurocomputing | 4 |
| 2014 | Sparse-based neural response for image classification
Hong Li 0009, Yantao Wei, Yuan Yan Tang |
Neurocomputing | 4 |
| 2014 | Conditional simultaneous localization and mapping: A robust visual SLAM system
Jigang Liu, Dongquan Liu, Jun Cheng 0002, Yuan Yan Tang |
Neurocomputing | 4 |
| 2014 | An enhanced version and an incremental learning version of visual-attention-imitation convex hull algorithm
Runzong Liu, Yuan Yan Tang, Bin Fang 0001, Jingrui Pi |
Neurocomputing | 2 |
| 2014 | An approximate closed-form solution to correlation similarity discriminant analysis
Taiping Zhang, Yuan Yan Tang, C. L. Philip Chen, Zhaowei Shang, Bin Fang 0001 |
Neurocomputing | 2 |
| 2014 | Generalization performance of Gaussian kernels SVMC based on Markov sampling
Jie Xu 0006, Yuan Yan Tang, Bin Zou 0002, Zongben Xu, Luoqing Li, Yang Lu 0009 |
Neural Networks | 2 |
| 2014 | Nonnegative class-specific entropy component analysis with adaptive step search criterion
Chi-Man Pun, Yuan Yan Tang |
Pattern Anal. Appl. | 3 |
| 2014 | Robust face recognition via occlusion dictionary learning
Weihua Ou, Xinge You, Dacheng Tao, Pengyue Zhang, Yuan Yan Tang |
Pattern Recognit. | 5 |
| 2014 | Hierarchical kernel-based rotation and scale invariant similarity
Yuan Yan Tang, Yantao Wei, Hong Li 0009, Luoqing Li |
Pattern Recognit. | 1 |
| 2014 | Hyperspectral Image Classification Using Functional Data AnalysisabstractThe large number of spectral bands acquired by hyperspectral imaging sensors allows us to better distinguish many subtle objects and materials. Unlike other classical hyperspectral image classification methods in the multivariate analysis framework, in this paper, a novel method using functional data analysis (FDA) for accurate classification of hyperspectral images has been proposed. The central idea of FDA is to treat multivariate data as continuous functions. From this perspective, the spectral curve of each pixel in the hyperspectral images is naturally viewed as a function. This can be beneficial for making full use of the abundant spectral information. The relevance between adjacent pixel elements in the hyperspectral images can also be utilized reasonably. Functional principal component analysis is applied to solve the classification problem of these functions. Experimental results on three hyperspectral images show that the proposed method can achieve higher classification accuracies in comparison to some state-of-the-art hyperspectral image classification methods. Hong Li 0009, Guangrun Xiao, Yuan Yan Tang, Luoqing Li |
IEEE Trans. Cybern. | 4 |
| 2014 | Topological Coding and Its Application in the Refinement of SIFTabstractPoint pattern matching plays a prominent role in the fields of computer vision and pattern recognition. A technique combining the circular onion peeling and the radial decomposition is proposed to analyze the topology structure of a point pattern. The analysis derives a feature which records the topological structure of a point pattern. This novel feature is free from isometric assumption. It can resist various deformations such as adding points, suppressing points, affine transformations, projective transformations and elastic transformations to some degree. A refinement solution of the well known scale invariant feature transform (SIFT) algorithm is also proposed based on the probabilistic analysis of this feature. Experimental results show that the proposed refinement solution for SIFT using this feature is effective and robust. Runzong Liu, Yuan Yan Tang, Bin Fang 0001 |
IEEE Trans. Cybern. | 2 |
| 2014 | Social Image Tagging With Diverse SemanticsabstractWe have witnessed the popularity of image-sharing websites for sharing personal experiences through photos on the Web. These websites allow users describing the content of their uploaded images with a set of tags. Those user-annotated tags are often noisy and biased. Social image tagging aims at removing noisy tags and suggests new relevant tags. However, most existing tag enrichment approaches predominantly focus on tag relevance and overlook tag diversity problem. How to make the top-ranked tags covering a wide range of semantic is still an opening, yet challenging, issue. In this paper, we propose an approach to retag social images with diverse semantics. Both the relevance of a tag to image as well as its semantic compensations to the already determined tags are fused to determine the final tag list for a given image. Different from existing image tagging approaches, the top-ranked tags are not only highly relevant to the image but also have significant semantic compensations with each other. Experiments show the effectiveness of the proposed approach. Xueming Qian, Xian-Sheng Hua 0001, Yuan Yan Tang, Tao Mei 0001 |
IEEE Trans. Cybern. | 3 |
| 2014 | High-Order Distance-Based Multiview Stochastic Learning in Image ClassificationabstractHow do we find all images in a larger set of images which have a specific content? Or estimate the position of a specific object relative to the camera? Image classification methods, like support vector machine (supervised) and transductive support vector machine (semi-supervised), are invaluable tools for the applications of content-based image retrieval, pose estimation, and optical character recognition. However, these methods only can handle the images represented by single feature. In many cases, different features (or multiview data) can be obtained, and how to efficiently utilize them is a challenge. It is inappropriate for the traditionally concatenating schema to link features of different views into a long vector. The reason is each view has its specific statistical property and physical interpretation. In this paper, we propose a high-order distance-based multiview stochastic learning (HD-MSL) method for image classification. HD-MSL effectively combines varied features into a unified representation and integrates the labeling information based on a probabilistic framework. In comparison with the existing strategies, our approach adopts the high-order distance obtained from the hypergraph to replace pairwise distance in estimating the probability matrix of data distribution. In addition, the proposed approach can automatically learn a combination coefficient for each view, which plays an important role in utilizing the complementary information of multiview data. An alternative optimization is designed to solve the objective functions of HD-MSL and obtain different views on coefficients and classification scores simultaneously. Experiments on two real world datasets demonstrate the effectiveness of HD-MSL in image classification. Jun Yu 0002, Yong Rui, Yuan Yan Tang, Dacheng Tao |
IEEE Trans. Cybern. | 3 |
| 2014 | The Generalization Performance of Regularized Regression Algorithms Based on Markov SamplingabstractThis paper considers the generalization ability of two regularized regression algorithms [least square regularized regression (LSRR) and support vector machine regression (SVMR)] based on non-independent and identically distributed (non-i.i.d.) samples. Different from the previously known works for non-i.i.d. samples, in this paper, we research the generalization bounds of two regularized regression algorithms based on uniformly ergodic Markov chain (u.e.M.c.) samples. Inspired by the idea from Markov chain Monto Carlo (MCMC) methods, we also introduce a new Markov sampling algorithm for regression to generate u.e.M.c. samples from a given dataset, and then, we present the numerical studies on the learning performance of LSRR and SVMR based on Markov sampling, respectively. The experimental results show that LSRR and SVMR based on Markov sampling can present obviously smaller mean square errors and smaller variances compared to random sampling. Bin Zou 0002, Yuan Yan Tang, Zongben Xu, Luoqing Li, Jie Xu 0006, Yang Lu 0009 |
IEEE Trans. Cybern. | 2 |
| 2014 | A Local Contrast Method for Small Infrared Target DetectionabstractRobust small target detection of low signal-to-noise ratio (SNR) is very important in infrared search and track applications for self-defense or attacks. Consequently, an effective small target detection algorithm inspired by the contrast mechanism of human vision system and derived kernel model is presented in this paper. At the first stage, the local contrast map of the input image is obtained using the proposed local contrast measure which measures the dissimilarity between the current location and its neighborhoods. In this way, target signal enhancement and background clutter suppression are achieved simultaneously. At the second stage, an adaptive threshold is adopted to segment the target. The experiments on two sequences have validated the detection capability of the proposed target detection method. Experimental evaluation results show that our method is simple and effective with respect to detection accuracy. In particular, the proposed method can improve the SNR of the image significantly. C. L. Philip Chen, Hong Li 0009, Yantao Wei, Yuan Yan Tang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2014 | Manifold-Based Sparse Representation for Hyperspectral Image ClassificationabstractA sparsity-based model has led to interesting results in hyperspectral image (HSI) classification. Sparse representation from a test sample is used to identify the class label. However, an ℓ1-based sparse algorithm sometimes yields unstable sparse representation. Inspired by recent progress in manifold learning, two manifold-based sparse representation algorithms are proposed to exploit the local structure of the test samples in corresponding sparse representations for enforcing smoothness across neighboring samples' sparse representations. Using techniques from regularization and local invariance, two manifold-based regularization terms are incorporated into the ℓ1-based objective function. Extensive experiments show that our proposed algorithms obtain excellent classification performance on three classic HSIs. Yuan Yan Tang, Luoqing Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Scalable Compression of Stream Cipher Encrypted Images Through Context-Adaptive SamplingabstractThis paper proposes a novel scalable compression method for stream cipher encrypted images, where stream cipher is used in the standard format. The bit stream in the base layer is produced by coding a series of nonoverlapping patches of the uniformly down-sampled version of the encrypted image. An off-line learning approach can be exploited to model the reconstruction error from pixel samples of the original image patch, based on the intrinsic relationship between the local complexity and the length of the compressed bit stream. This error model leads to a greedy strategy of adaptively selecting pixels to be coded in the enhancement layer. At the decoder side, an iterative, multiscale technique is developed to reconstruct the image from all the available pixel samples. Experimental results demonstrate that the proposed scheme outperforms the state-of-the-arts in terms of both rate-distortion performance and visual quality of the reconstructed images at low and medium rate regions. Jiantao Zhou 0001, Oscar C. Au, Guangtao Zhai, Yuan Yan Tang, Xianming Liu 0005 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2014 | Designing an Efficient Image Encryption-Then-Compression System via Prediction Error Clustering and Random PermutationabstractIn many practical scenarios, image encryption has to be conducted prior to image compression. This has led to the problem of how to design a pair of image encryption and compression algorithms such that compressing the encrypted images can still be efficiently performed. In this paper, we design a highly efficient image encryption-then-compression (ETC) system, where both lossless and lossy compression are considered. The proposed image encryption scheme operated in the prediction error domain is shown to be able to provide a reasonably high level of security. We also demonstrate that an arithmetic coding-based approach can be exploited to efficiently compress the encrypted images. More notably, the proposed compression approach applied to encrypted images is only slightly worse, in terms of compression efficiency, than the state-of-the-art lossless/lossy image coders, which take original, unencrypted images as inputs. In contrast, most of the existing ETC solutions induce significant penalty on the compression efficiency. Jiantao Zhou 0001, Xianming Liu 0005, Oscar C. Au, Yuan Yan Tang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2014 | Group Sparse Multiview Patch Alignment Framework With View Consistency for Image ClassificationabstractNo single feature can satisfactorily characterize the semantic concepts of an image. Multiview learning aims to unify different kinds of features to produce a consensual and efficient representation. This paper redefines part optimization in the patch alignment framework (PAF) and develops a group sparse multiview patch alignment framework (GSM-PAF). The new part optimization considers not only the complementary properties of different views, but also view consistency. In particular, view consistency models the correlations between all possible combinations of any two kinds of view. In contrast to conventional dimensionality reduction algorithms that perform feature extraction and feature selection independently, GSM-PAF enjoys joint feature extraction and feature selection by exploiting l(2,1)-norm on the projection matrix to achieve row sparsity, which leads to the simultaneous selection of relevant features and learning transformation, and thus makes the algorithm more discriminative. Experiments on two real-world image data sets demonstrate the effectiveness of GSM-PAF for image classification. Jie Gui, Dacheng Tao, Zhenan Sun, Yong Luo 0002, Xinge You, Yuan Yan Tang |
IEEE Trans. Image Process. | 6 |
| 2013 | A no-reference image sharpness estimation based on expectation of wavelet transform coefficientsabstractIn this work, the expectation of wavelet transform coefficients is used for estimating an image sharpness. It's based on the observation that the greater the probability of big detail coefficients, the more pixels appear sharply, and consequently, the sharper the image. Specifically, an input image is firstly decomposed into three directional sub-bands by a separable discrete wavelet transform. Then these directional sub-bands are viewed as three random variables, and their expectations are computed. Finally, The proposed sharpness index is the weighted sum of three expectations. The experiments show that, despite its simplicity, the proposed sharpness index is competitive with the current best-performance techniques for no-reference image sharpness estimation. Hengjun Zhao, Bin Fang 0001, Yuan Yan Tang |
ICIP | 3 |
| 2013 | GPS Estimation from Users' Photos
Jing Li 0049, Xueming Qian, Yuan Yan Tang, Linjun Yang, Chaoteng Liu |
MMM (1) | 3 |
| 2013 | Generalization performance of support vector classifiers for density level detection
Hong Chen 0004, Yicong Zhou, Yi Tang 0003, Yuan Yan Tang, Zhibin Pan |
Neurocomputing | 4 |
| 2013 | ISABoost: A weak classifier inner structure adjusting based AdaBoost algorithm - ISABoost based application in scene categorization
Xueming Qian, Yuan Yan Tang, Kaiyu Hang |
Neurocomputing | 2 |
| 2013 | Visual saliency detection with center shift
Weibin Yang, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang, Yuewei Lin |
Neurocomputing | 2 |
| 2013 | Learning orthogonal projections for Isomap
Bin Fang 0001, Yuan Yan Tang, Taiping Zhang, Ruizong Liu |
Neurocomputing | 3 |
| 2013 | Filtering Terms from the Web for Image AnnotationsabstractIn this paper, we propose a novel automatic image annotation model by mining the web. In our approach, the terms or words appearing in the associated text are extracted and filtered as labels or annotations for the corresponding web images. Sure, much noise exists in those selected labels. In order to reduce the influence caused by the noisy labels, for each label or potential word, we improve web image-word relationships using Mixture Gaussian Distribution Model. By doing so, the relationships between words and images are re-weighted both in terms of sematic relevance and in terms of visual feature similarity. In fact, all the words associated to an image are not semantically independent. We use co-occurrences between two words to describe their semantic relevance. Thus, we further use a method, called Word Promotion, to co-enhance the weights of all the words associated to a given image based on their co-occurrences. Our experiments are conducted in several ways and the results show that our annotation method can achieve a satisfactory performance in respects of system scalability and sematic evolution. Zhiguo Gong, Jingzhi Guo, Yuan Yan Tang, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2013 | A Fast and Complete Convex-Hull Algorithm Architecture Based on Ellipse and Elastic Ellipse MethodsabstractThe number of inner points excluded in an initial convex hull (ICH) is vital to the efficiency getting the convex hull (CH) in a planar point set. The maximum inscribed circle method proposed recently is effective to remove inner points in ICH. However, limited by density distribution of a planar point set, it does not always work well. Although the affine transformation method can be used, it is still hard to have a better performance. Furthermore, the algorithm mentioned above fails to deal with the exceptional distribution: the gravity centroid (GC) of a planar point set is outside or on the edge formed by the extreme points in ICH. This paper considers how to remove more inner points in ICH when GC is inside of ICH and completely process the case which mentioned above. Further, we presented a complete algorithm architecture: (1) using the ellipse and elasticity ellipse methods (EM and EEM) to remove more inner points in ICH and process the cases: GC is inside or outside of ICH. (2) Using the traditional methods to process the situation: the initial centroid is on the edge in ICH. It is adaptive to more data sets than other algorithms. The experiments under seven distributions show that the proposed method performs better than other traditional algorithms in saving time and space. Xuegang Wu, Bin Fang 0001, Yuan Yan Tang, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2013 | Error Analysis of Coefficient-Based Regularized Algorithm for Density-Level DetectionabstractIn this letter, we consider a density-level detection (DLD) problem by a coefficient-based classification framework with [Formula: see text]-regularizer and data-dependent hypothesis spaces. Although the data-dependent characteristic of the algorithm provides flexibility and adaptivity for DLD, it leads to difficulty in generalization error analysis. To overcome this difficulty, an error decomposition is introduced from an established classification framework. On the basis of this decomposition, the estimate of the learning rate is obtained by using Rademacher average and stepping-stone techniques. In particular, the estimate is independent of the capacity assumption used in the previous literature. Hong Chen 0004, Zhibin Pan, Luoqing Li, Yuan Yan Tang |
Neural Comput. | 4 |
| 2013 | Convergence rate of the semi-supervised greedy algorithm
Hong Chen 0004, Yicong Zhou, Yuan Yan Tang, Luoqing Li, Zhibin Pan |
Neural Networks | 3 |
| 2013 | A Visual-Attention Model Using Earth Mover's Distance-Based Saliency Measurement and Nonlinear Feature CombinationabstractThis paper introduces a new computational visual-attention model for static and dynamic saliency maps. First, we use the Earth Mover's Distance (EMD) to measure the center-surround difference in the receptive field, instead of using the Difference-of-Gaussian filter that is widely used in many previous visual-attention models. Second, we propose to take two steps of biologically inspired nonlinear operations for combining different features: combining subsets of basic features into a set of super features using the Lm-norm and then combining the super features using the Winner-Take-All mechanism. Third, we extend the proposed model to construct dynamic saliency maps from videos by using EMD for computing the center-surround difference in the spatiotemporal receptive field. We evaluate the performance of the proposed model on both static image data and video data. Comparison results show that the proposed model outperforms several existing models under a unified evaluation setting. Yuewei Lin, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang, Song Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Multi-focus image fusion based on the neighbor distance
Hengjun Zhao, Zhaowei Shang, Yuan Yan Tang, Bin Fang 0001 |
Pattern Recognit. | 3 |
| 2013 | Error Analysis of Stochastic Gradient Descent RankingabstractRanking is always an important task in machine learning and information retrieval, e.g., collaborative filtering, recommender systems, drug discovery, etc. A kernel-based stochastic gradient descent algorithm with the least squares loss is proposed for ranking in this paper. The implementation of this algorithm is simple, and an expression of the solution is derived via a sampling operator and an integral operator. An explicit convergence rate for leaning a ranking function is given in terms of the suitable choices of the step size and the regularization parameter. The analysis technique used here is capacity independent and is novel in error analysis of ranking learning. Experimental results on real-world data have shown the effectiveness of the proposed algorithm in ranking tasks, which verifies the theoretical analysis in ranking error. Hong Chen 0004, Yi Tang 0003, Luoqing Li, Yuan Yuan 0001, Xuelong Li 0001, Yuan Yan Tang |
IEEE Trans. Cybern. | 6 |
| 2013 | GPS Estimation for Places of Interest From Social Users' Uploaded PhotosabstractSocial media has become a very popular way for people to share their photos with friends. Because most of the social images are attached with GPS (geo-tags), a photo's GPS information can be estimated with the help of the large geo-tagged image set while using a visual searching based approach. This paper proposes an unsupervised image GPS location estimation approach with hierarchical global feature clustering and local feature refinement. It consists of two parts: an offline system and an online system. In the offline system, a hierarchical structure is constructed for a large-scale offline social image set with GPS information. Representative images are selected for each GPS location refined cluster, and an inverted file structure is proposed. In the online system, when given an input image, its GPS information can be estimated by hierarchical global clusters selection and local feature refinement in the online system. Both the computational cost and GPS estimation performance demonstrates the effectiveness of the proposed hierarchical structure and inverted file structure in our approach. Jing Li 0049, Xueming Qian, Yuan Yan Tang, Linjun Yang, Tao Mei 0001 |
IEEE Trans. Multim. | 3 |
| 2013 | Generalization Performance of Fisher Linear Discriminant Based on Markov SamplingabstractFisher linear discriminant (FLD) is a well-known method for dimensionality reduction and classification that projects high-dimensional data onto a low-dimensional space where the data achieves maximum class separability. The previous works describing the generalization ability of FLD have usually been based on the assumption of independent and identically distributed (i.i.d.) samples. In this paper, we go far beyond this classical framework by studying the generalization ability of FLD based on Markov sampling. We first establish the bounds on the generalization performance of FLD based on uniformly ergodic Markov chain (u.e.M.c.) samples, and prove that FLD based on u.e.M.c. samples is consistent. By following the enlightening idea from Markov chain Monto Carlo methods, we also introduce a Markov sampling algorithm for FLD to generate u.e.M.c. samples from a given data of finite size. Through simulation studies and numerical studies on benchmark repository using FLD, we find that FLD based on u.e.M.c. samples generated by Markov sampling can provide smaller misclassification rates compared to i.i.d. samples. Bin Zou 0002, Luoqing Li, Zongben Xu, Tao Luo 0006, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2012 | Texture Classification Based on BIMF Monogenic Signals
Jianjia Pan, Yuan Yan Tang |
ACCV (2) | 2 |
| 2012 | Visual saliency estimation using support value transformabstractThis paper proposes a novel method for estimating visual saliency based on a typical agreement that image saliency depends mainly on local and global contrast from various feature channels. We compute the contrast between image patches on different low-level feature maps which are generated by color space conversion and support value transform. To obtain the representative measurement effectively, we calculate the dissimilarity in a reduced dimensional principal component space. In addition, our method may be easily extended for more conspicuous feature channels in an efficient manner. Experimental results on two public available human eye fixation datasets demonstrate that our method outperforms other seven state-of-the-art saliency models. Weibin Yang, Bin Fang 0001, Yuan Yan Tang, Zhaowei Shang, Hengjun Zhao |
ICIP | 3 |
| 2012 | Object categorization based on hierarchical learning
Yuan Yan Tang, Yantao Wei, Hong Li 0009, Luoqing Li |
ICPR | 2 |
| 2012 | Orthogonal Isometric Projection
Yuan Yan Tang, Bin Fang 0001, Taiping Zhang |
ICPR | 2 |
| 2012 | An affine invariant discriminate analysis with canonical correlation analysis
Rushi Lan, Zhan Song, Yuan Yan Tang |
Neurocomputing | 5 |
| 2012 | A fast convex hull algorithm with maximum inscribed circle affine transformation
Runzong Liu, Bin Fang 0001, Yuan Yan Tang, Jiye Qian |
Neurocomputing | 3 |
| 2012 | Fragmented edge structure coding for Chinese writer identification
Bin Fang 0001, Junlin Chen, Yuan Yan Tang, Hengxin Chen |
Neurocomputing | 4 |
| 2012 | Autonomous Behaviors of Graphical Avatars Based on Machine LearningabstractGraphical avatars have gained popularity in many application domains such as three-dimensional (3D) animation movies and animated simulations for product design. However, the methods to edit avatars' behaviors in the 3D graphical environment remained to be a challenging research topic. Since the hand-crafted methods are time-consuming and inefficient, the automatic actions of the avatars are required. To achieve the autonomous behaviors of the avatars, artificial intelligence should be used in this research area. In this paper, we present a novel approach to construct a system of automatic avatars in the 3D graphical environments based on the machine learning techniques. Specific framework is created for controlling the behaviors of avatars, such as classifying the difference among the environments and using hierarchical structure to describe these actions. Because of the requirement of simulating the interactions between avatars and environments after the classification of the environment, Reinforcement Learning is used to compute the policy to control the avatar intelligently in the 3D environment for the solution of the problem of different situations. Thus, our approach has solved problems such as where the levels of the missions will be defined and how the learning algorithm will be used to control the avatars. In this paper, our method to achieve these goals will be presented. The main contributions of this paper are presenting a hierarchical structure to control avatars automatically, developing a method for avatars to recognize environment and presenting an approach for making the policy of avatars' actions intelligently. Yuesheng He, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2012 | Multi-Scale Gradient Invariant for Face Recognition under varying IlluminationabstractIn this paper, a novel approach derived from image gradient domain called multi-scale gradient faces (MGF) is proposed to abstract multi-scale illumination-insensitive measure for face recognition. MGF applies multi-scale analysis on image gradient information, which can discover underlying inherent structure in images and keep the details at most while removing varying lighting. The proposed approach provides state-of-the-art performance on Extended YaleB and PIE: Recognition rates of 99.11% achieved on PIE database and 99.38% achieved on YaleB which outperforms most existing approaches. Furthermore, the experimental results on noised Yale-B validate that MGF is more robust to image noise. Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2012 | Document Clustering in Correlation Similarity Measure SpaceabstractThis paper presents a new spectral clustering method called correlation preserving indexing (CPI), which is performed in the correlation similarity measure space. In this framework, the documents are projected into a low-dimensional semantic space in which the correlations between the documents in the local patches are maximized while the correlations between the documents outside these patches are minimized simultaneously. Since the intrinsic geometrical structure of the document space is often embedded in the similarities between the documents, correlation as a similarity measure is more suitable for detecting the intrinsic geometrical structure of the document space than euclidean distance. Consequently, the proposed CPI method can effectively discover the intrinsic structures embedded in high-dimensional document space. The effectiveness of the new method is demonstrated by extensive experiments conducted on various data sets and by comparison with existing document clustering methods. Taiping Zhang, Yuan Yan Tang, Bin Fang 0001, Yong Xiang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | Texture Analysis Based on Saddle Points-Based BEMD and LBP
Jianjia Pan, Yuan Yan Tang |
CAIP (2) | 2 |
| 2011 | Kernel-view based discriminant approach for embedded feature extraction in high-dimensional space
Bin Fang 0001, Chi-Man Pun, Yuan Yan Tang |
Neurocomputing | 4 |
| 2011 | A Multi-Layer Contrast Analysis Method for Texture Classification Based on LBPabstractTexture classification is one of the important fields in pattern recognition and machine vision research. LBP method,13–15 proposed by Ojala, can be used to classify texture images effectively. And the LBP method has rotation-invariant, illumination-invariant, multi-resolution characteristics. But, since the contrast is not considered between neighbor pixels, the correct classification rate produced by this method has been remarkably influenced by light source type and light source orientation. The LMLCP (Local Multiple Layer Contrast Pattern) method, proposed by this paper, maps the contrast value between two near pixels to a rank value, which represent a relative contrast value range, and computes the statistic histogram referring to the work in LBP method. The LMLCP method can bring out the rapid expansion of feature dimension, so a special feature encoding method used in 3DLBP6 is adopted by this paper. The experiment, which is built based on Outex_TC_00012,12 demonstrates that the LMLCP can evidently make a more accurate classification rate than LBP method. Hengxin Chen, Yuan Yan Tang, Bin Fang 0001, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2011 | Illumination Invariant Face Recognition Using Fabemd Decomposition with Detail Measure WeightabstractWith varying illumination conditions, facial features obtained from images are distorted nonlinearly by variant lighting intensity and direction, so face recognition becomes very difficult. According to the "common assumption", illumination varies slowly and the face intrinsic feature (including 3D surface and reflectance) varies rapidly in local area, we can then consider high frequency features that represent the face intrinsic structure. FABEMD8 (Fast and Adaptive Bidimensional Empirical Mode Decomposition) is a fast and adaptive method of BEMD22 (Bidimensional Empirical Mode Decomposition), and not using time-consuming plane interpolation computation, it can decompose the image into multilayer high frequency images representing detail features and low frequency images representing analogy features. But we cannot make a quantitative analysis of how many detail features can be used to eliminate illumination variation. So we propose two measures to quantify the detail features, and with these measure weights, we can activitate FABEMD based multilayer detail images matching for face recognition under varying illumination. With PCA, the experiments based on Yale face database B and MU PIE face database show that the method proposed in this paper can get remarkable performance. Hengxin Chen, Yuan Yan Tang, Bin Fang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2011 | A Computational and Theoretical Analysis of Local Null Space Discriminant Method for Pattern ClassificationabstractMany problems in pattern classification and feature extraction involve dimensionality reduction as a necessary processing. Traditional manifold learning algorithms, such as ISOMAP, LLE, and Laplacian Eigenmap, seek the low-dimensional manifold in an unsupervised way, while the local discriminant analysis methods identify the underlying supervised submanifold structures. In addition, it has been well-known that the intraclass null subspace contains the most discriminative information if the original data exist in a high-dimensional space. In this paper, we seek for the local null space in accordance with the null space LDA (NLDA) approach and reveal that its computational expense mainly depends on the quantity of connected edges in graphs, which may be still unacceptable if a great deal of samples are involved. To address this limitation, an improved local null space algorithm is proposed to employ the penalty subspace to approximate the local discriminant subspace. Compared with the traditional approach, the proposed method can achieve more efficiency so that the overload problem is avoided, while slight discriminant power is lost theoretically. A comparative study on classification shows that the performance of the approximative algorithm is quite close to the genuine one. Bin Fang 0001, Yuan Yan Tang, Hengxin Chen |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2011 | Bionic Face Recognition Using Gabor TransformationabstractIn this paper, we propose a bionic face recognition method based on Gabor feature. First, Gabor features are extracted from face images, followed by dimensionality reduction using 2DPCA algorithm, which serves as the feature vectors of the proposed method. Finally, the bionic classifier is trained for classification. The experiment on AR and PIE face database is reported to show the effectiveness of the proposed method and compare it with Gabor-2DPCA algorithm and Gabor-PCA algorithm. Yuan Yan Tang, Bin Fang 0001, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2011 | Wavelet Decomposition of Pseudo-Motion Image and Application to Frequency SegmentationabstractIn this paper, we propose a novel approach to frequency segmentation using wavelet decomposition of pseudo-motion image. In this way, a fixed image is translated such that a sequence of moving images is produced, which are called the 0 pseudo-motion images. In fact, we can consider the translation of a function to be a motion of eyeshot. When a function is translated, its wavelet coefficients will oscillate. From this property, we can detect the special areas of the image. Some experiments were conducted, which are used to find the position of license plates of vehicles in digital images. The experimental results demonstrate the performance of the method. Yuan Yan Tang, Limin Cui |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2011 | Biologically Inspired Features for Scene Classification in Video SurveillanceabstractInspired by human visual cognition mechanism, this paper first presents a scene classification method based on an improved standard model feature. Compared with state-of-the-art efforts in scene classification, the newly proposed method is more robust, more selective , and of lower complexity. These advantages are demonstrated by two sets of experiments on both our own database and standard public ones. Furthermore, occlusion and disorder problems in scene classification in video surveillance are also first studied in this paper. Kaiqi Huang, Dacheng Tao, Yuan Yan Tang, Xuelong Li 0001, Tieniu Tan |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2010 | A Computational Model for Saliency Maps by Using Local EntropyabstractThis paper presents a computational framework for saliency maps. It employs the Earth Mover's Distance based on weighted-Histogram (EMD-wH) to measure the center-surround difference, instead of the Difference-of-Gaussian (DoG) filter used by traditional models. In addition, the model employs not only the traditional features such as colors, intensity and orientation but also the local entropy which expresses the local complexity. The major advantage of combining the local entropy map is that it can detect the salient regions which are not complex regions. Also, it uses a general framework to integrate the feature dimensions instead of summing the features directly. This model considers both local and global salient information, in contrast to the existing models that consider only one or the other. Furthermore, the "large scale bias" and "central bias" hypotheses are used in this model to select the fixation locations in the saliency map of different scales. The performance of this model is assessed by comparing their saliency maps and human fixation density. The results from this model are finally compared to those from other bottom-up models for reference. Yuewei Lin, Bin Fang 0001, Yuan Yan Tang |
AAAI | 3 |
| 2010 | Wavelet Domain Local Binary Pattern Features For Writer IdentificationabstractThe representation of writing styles is a crucial step of writer identification schemes. However, the large intra-writer variance makes it a challenging task. Thus, a good feature of writing style plays a key role in writer identification. In this paper, we present a simple and effective feature for off-line, text-independent writer identification, namely wavelet domain local binary patterns (WD-LBP). Based on WD-LBP, a writer identification algorithm is developed. WD-LBP is able to capture the essence of characteristics of writer while ignoring the variations intrinsic to every single writer. Unlike other texture framework method, we do not assign any statistical distribution assumption to the proposed method. This prevent us from making any, possibly erroneous, assumptions about the handwritten image feature distributions. The experimental results show that the proposed writer identification method achieves high accuracy of identification and outperforms recent writer identification method such as wavelet-GGD model and Gabor filtering method. Xinge You, Zhifan Gao, Yuan Yan Tang |
ICPR | 5 |
| 2010 | A fractal-based BEMD method for image texture analysisabstractThis paper presents a new method for texture analysis through Bidimensional Empirical Mode Decomposition (BEMD). Although there have been many filter based methods for texture analysis, problems of non-adaptively and redundancy are still hard to solve. The BEMD is a locally adaptive method and suitable for the analysis of nonlinear or nonstationary signals. The texture image can be decomposed to several IMFs (intrinsic mode functions) by BEMD, which present new characters of the images. But for the BEMD, the boundary interference is a main limit for its application. In this paper, we proposed a new BEMD method based on the self-similar extend method and the neighbor local extremes to reduce the boundary interference. This new method can get a lower orthogonality index (OI) of the IMF, which present more clearly features of the texture images. The experiment result shown the new method also reduced the computation complex compared to other surface interpolation based methods. Jianjia Pan, Dan Zhang 0008, Yuan Yan Tang |
SMC | 3 |
| 2010 | Illumination invariant face recognition based on the new phase featuresabstractHilbert-Huang transform (HHT) is a novel signal processing method which can efficiently handle non-stationary and nonlinear signals. Two key parts are included: Empirical Mode Decomposition (EMD) and Hilbert transform. EMD decomposes signals into a complete series of Intrinsic Mode Functions (IMFs), which capture the intrinsic frequency components of the original signals. Hilbert transform is adopted on the IMFs to get the analytical local features. Due to its efficiency in signal processing, the bidimensional version has been studied for the advanced image processing. EMD has been extended to bidimensional EMD (BEMD), and the corresponding monogenic signals are studied. Phase information is an important local feature of signals in frequency domain because it is robust to contrast, brightness, noise, shading in the image. The quantity Phase congruency (PC) is invariant to changes in image illumination. In this paper, we firstly proposed an improved BEMD method based on the novel evaluation of local mean, then the Riesz transform is applied to get the corresponding monogenic signals. Finally, PC was calculated based on the new phase information and it then has been adopted as facial features to classify faces under variant illumination conditions. The experimental results demonstrated the efficiency of the proposed approach. Dan Zhang 0008, Jianjia Pan, Yuan Yan Tang |
SMC | 3 |
| 2010 | A Least-Squares Model to Orthogonal Linear Discriminant AnalysisabstractOrthogonal transformation can delete the correlations among candidate features such that the extracted features do not disturb each other. An orthogonal set of discriminant vectors is more powerful than the classical discriminant vectors. In this paper, we present a new orthogonal linear discriminant analysis (OLDA) model based on least-squares approximation called LS-OLDA for pattern classification, which aims to find an orthogonal transformation W and a diagonal matrix D such that the difference between [Formula: see text] and WDWT is minimized in the least-squares sense, and the trace of D is maximized simultaneously. Theoretical analysis shows that the proposed model coincides with classical OLDA criterion. The experimental results on different standard data sets compared with related methods show that LS-OLDA achieves or approximates closely to the best accuracy, and has lower computational cost. Taiping Zhang, Bin Fang 0001, Yuan Yan Tang, Zhaowei Shang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2010 | Embedding meshes/tori in faulty crossed cubes
Xiaofan Yang 0001, Qiang Dong, Yuan Yan Tang |
Inf. Process. Lett. | 3 |
| 2010 | Marginal discriminant projections: An adaptable margin discriminant approach to feature reduction and extraction
Bin Fang 0001, Yuan Yan Tang |
Pattern Recognit. Lett. | 4 |
| 2010 | Incremental Embedding and Learning in the Local Discriminant Subspace With Application to Face RecognitionabstractDimensionality reduction and incremental learning have recently received broad attention in many applications of data mining, pattern recognition, and information retrieval. Inspired by the concept of manifold learning, many discriminant embedding techniques have been introduced to seek low-dimensional discriminative manifold structure in the high-dimensional space for feature reduction and classification. However, such graph-embedding framework-based subspace methods usually confront two limitations: (1) since there is no available updating rule for local discriminant analysis with the additive data, it is difficult to design incremental learning algorithm and (2) the small sample size (SSS) problem usually occurs if the original data exist in very high-dimensional space. To overcome these problems, this paper devises a supervised learning method, called local discriminant subspace embedding (LDSE), to extract discriminative features. Then, the incremental-mode algorithm, incremental LDSE (ILDSE), is proposed to learn the local discriminant subspace with the newly inserted data, which applies incremental learning extension to the batch LDSE algorithm by employing the idea of singular value-decomposition (SVD) updating algorithm. Furthermore, the SSS problem is avoided in our method for the high-dimensional data and the benchmark incremental learning experiments on face recognition show that ILDSE bears much less computational cost compared with the batch algorithm. Bin Fang 0001, Yuan Yan Tang, Taiping Zhang |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2010 | Generalized Discriminant Analysis: A Matrix Exponential ApproachabstractLinear discriminant analysis (LDA) is well known as a powerful tool for discriminant analysis. In the case of a small training data set, however, it cannot directly be applied to high-dimensional data. This case is the so-called small-sample-size or undersampled problem. In this paper, we propose an exponential discriminant analysis (EDA) technique to overcome the undersampled problem. The advantages of EDA are that, compared with principal component analysis (PCA) + LDA, the EDA method can extract the most discriminant information that was contained in the null space of a within-class scatter matrix, and compared with another LDA extension, i.e., null-space LDA (NLDA), the discriminant information that was contained in the non-null space of the within-class scatter matrix is not discarded. Furthermore, EDA is equivalent to transforming original data into a new space by distance diffusion mapping, and then, LDA is applied in such a new space. As a result of diffusion mapping, the margin between different classes is enlarged, which is helpful in improving classification accuracy. Comparisons of experimental results on different data sets are given with respect to existing LDA extensions, including PCA + LDA, LDA via generalized singular value decomposition, regularized LDA, NLDA, and LDA via QR decomposition, which demonstrate the effectiveness of the proposed EDA method. Taiping Zhang, Bin Fang 0001, Yuan Yan Tang, Zhaowei Shang |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2009 | Classification of 3D Models for the 3D Animation EnvironmentsabstractClassification the 3D models in the graphical environment is a key problem with applications in computer graphics, virtual reality, especially the intelligent virtual human for humanoid animation. The challenging aspects of this problem are to find a suitable shapes' feature that can be used to compare them quickly and a proper method to classify the models on the features. We propose a method of classifying shapes' features for surface-based 3D shape models based on their shape similarity. The features of shape of 3D models are computed by first converting an input surface based model into an oriented point set model and then computing features histograms of distance and shape distributions. Then, support vector machines (SVM) are used to classify the features of the models. By the classification the models can be given semantic meanings in the 3D environment. Yuesheng He, Yuan Yan Tang |
SMC | 2 |
| 2009 | Weightiness image Partition in 3D Face RecognitionabstractIn this paper we present a novel algorithm suitable to improve the accuracy of 3D face recognition. In the proposed algorithm, we represent the 3D points by point signatures and partition the facial data into fifteen regions according to ¿three courtyards and five eyes¿ theory in pencil sketch on facial image in Chinese traditional art. Then in each partition we use ICA getting eigenvalues of feature and structure character and depth information to represent the 3D facial data. We assign different weightiness to each sub-image according to the result of sub-image variety. In order to match incomplete data under structural constraints, we proposed a reformative robust structural Hausdorff distance to handle these possible cases. Experiments on FRGC v2.0 data set show that the proposed algorithm is robust and effective to 3D face with expression, lighting and expression variance. Yuan Yan Tang, Bin Fang 0001, Taiping Zhang |
SMC | 2 |
| 2009 | Improving the discriminant ability of local margin based learning method by incorporating the global between-class separability criterion
Bin Fang 0001, Yuan Yan Tang |
Neurocomputing | 3 |
| 2009 | Combining Eodh and Directional Gradient Density for Offline Signature VerificationabstractThe main problem to identify skilled forgeries for offline signature verification lies in the fact that it is difficult to formalize distinguished feature representation of the signature patterns and design appropriate fusion scheme for various types of feature vectors. To tackle these problems, in this paper, we propose an approach to extract robust Edge Orientation Distance Histogram (EODH) descriptor which effectively reflects signature structure variations. In addition, directional gradient density features are employed for skilled forgery verification attempt. To exploit the full capacity of two sets of features, we designed the multilevel weighted fuzzy classifier and fuse match scores by way of selection priority. Experiments were conducted on a subcorpus of open MCYT signature database which is widely used for performance evaluation. It shows that the proposed method was able to improve verification accuracy. Bin Fang 0001, Yuan Yan Tang, Patrick Shen-Pei Wang, Taiping Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2009 | Facial Biometrics Using Nontensor Product Wavelet and 2D Discriminant TechniquesabstractA new facial biometric scheme is proposed in this paper. Three steps are included. First, a new nontensor product bivariate wavelet is utilized to get different facial frequency components. Then a modified 2D linear discriminant technique (M2DLD) is applied on these frequency components to enhance the discrimination of the facial features. Finally, support vector machine (SVM) is adopted for classification. Compared with the traditional tensor product wavelet, the new nontensor product wavelet can detect more singular facial features in the high-frequency components. Earlier studies show that the high-frequency components are sensitive to facial expression variations and minor occlusions, while the low-frequency component is sensitive to illumination changes. Therefore, there are two advantages of using the new nontensor product wavelet compared with the traditional tensor product one. First, the low-frequency component is more robust to the expression variations and minor occlusions, which indicates that it is more efficient in facial feature representation. Second, the corresponding high-frequency components are more robust to the illumination changes, subsequently it is more powerful for classification as well. The application of the M2DLD on these wavelet frequency components enhances the discrimination of the facial features while reducing the feature vectors dimension a lot. The experimental results on the AR database and the PIE database verified the efficiency of the proposed method. Dan Zhang 0008, Xinge You, Patrick Shen-Pei Wang, Svetlana N. Yanushkevich, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2009 | Learning semantics from multimedia content
Dacheng Tao, Xuelong Li 0001, Yuan Yan Tang |
Pattern Recognit. | 3 |
| 2009 | Model-based signature verification with rotation invariant features
Bin Fang 0001, Yuan Yan Tang, Taiping Zhang |
Pattern Recognit. | 3 |
| 2009 | Multiscale facial structure representation for face recognition under varying illumination
Taiping Zhang, Bin Fang 0001, Yuan Yuan 0001, Yuan Yan Tang, Zhaowei Shang, Fangnian Lang |
Pattern Recognit. | 4 |
| 2009 | A novel iris segmentation using radial-suppression edge detection
Jing Huang 0018, Xinge You, Yuan Yan Tang, Yuan Yuan 0001 |
Signal Process. | 3 |
| 2009 | Face Recognition Under Varying Illumination Using GradientfacesabstractIn this correspondence, we propose a novel method to extract illumination insensitive features for face recognition under varying lighting called the Gradientfaces. Theoretical analysis shows Gradientfaces is an illumination insensitive measure, and robust to different illumination, including uncontrolled, natural lighting. In addition, Gradientfaces is derived from the image gradient domain such that it can discover underlying inherent structure of face images since the gradient domain explicitly considers the relationships between neighboring pixel points. Therefore, Gradientfaces has more discriminating power than the illumination insensitive measure extracted from the pixel domain. Recognition rates of 99.83% achieved on PIE database of 68 subjects, 98.96% achieved on Yale B of ten subjects, and 95.61% achieved on Outdoor database of 132 subjects under uncontrolled natural lighting conditions show that Gradientfaces is an effective method for face recognition under varying illumination. Furthermore, the experimental results on Yale database validate that Gradientfaces is also insensitive to image noise and object artifacts (such as facial expressions). Taiping Zhang, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang |
IEEE Trans. Image Process. | 2 |
| 2008 | Palmprint identification based on non-separable wavelet filter banksabstractCreases, as a special salient feature of palmprint, are large in number and distributed at all directions. It changes slowly in a personpsilas whole life, which qualifies themselves as features in palmprint identification. In this paper, we devised a new algorithm of crease extraction by using non-separable bivariate wavelet filter banks with linear phase. Compared with the traditional wavelet, our research demonstrates that the three high frequency sub-images generated by Non-separable Discrete Wavelet Transform (NDWT) can extract more creases and no longer extensively focus on the three special directions. As a consequence, we proposed a new method by combining NDWT and Support Vector Machines (SVM) for palmprint identification. Tested by our experiment, this method achieves a satisfied identification result and computational efficiency as well. Xinge You, Yuan Yan Tang, Yiu-Ming Cheung |
ICPR | 3 |
| 2008 | Iris recognition based on non-separable waveletabstractThis paper focuses on the rotation noise of iris recognition. Current iris recognition systems are unable to deal with the rotation noise perfectly. We propose a novel method for iris matching that decompose iris picture into wavelet subband coefficients via 16 non-separable wavelet filters, and use generalized Gaussian density (GGD) modeling of each non-separable orthogonal wavelet coefficients as a means of feature extraction, then compute the Kullback-Leibler distance (KLD) between GGDs and compare the iris code using the Kullback-Leibler distance. Experiments show that the proposed method is rotation invariance, it does not decrease their recognition rate, when the iris image is rotated. Jing Huang 0018, Xinge You, Yuan Yan Tang |
SMC | 3 |
| 2008 | Status of pattern recognition with wavelet analysis
Yuan Yan Tang |
Frontiers Comput. Sci. China | 1 |
| 2008 | Writer identification using global wavelet-based features
Zhenyu He 0001, Xinge You, Yuan Yan Tang |
Neurocomputing | 3 |
| 2008 | Total variation norm-based nonnegative matrix factorization for identifying discriminant representation of image patterns
Taiping Zhang, Bin Fang 0001, Weining Liu, Yuan Yan Tang |
Neurocomputing | 4 |
| 2008 | Editorial
Xizhao Wang, Yuan Yan Tang, Daniel S. Yeung |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2008 | Embedding a family of disjoint 3D meshes into a crossed cube
Qiang Dong, Xiaofan Yang 0001, Juan Zhao 0011, Yuan Yan Tang |
Inf. Sci. | 4 |
| 2008 | Writer identification of Chinese handwriting documents using hidden Markov tree model
Zhenyu He 0001, Xinge You, Yuan Yan Tang |
Pattern Recognit. | 3 |
| 2008 | Topology Preserving Non-negative Matrix Factorization for Face RecognitionabstractIn this paper, a novel topology preserving non-negative matrix factorization (TPNMF) method is proposed for face recognition. We derive the TPNMF model from original NMF algorithm by preserving local topology structure. The TPNMF is based on minimizing the constraint gradient distance in the high-dimensional space. Compared with L(2) distance, the gradient distance is able to reveal latent manifold structure of face patterns. By using TPNMF decomposition, the high-dimensional face space is transformed into a local topology preserving subspace for face recognition. In comparison with PCA, LDA, and original NMF, which search only the Euclidean structure of face space, the proposed TPNMF finds an embedding that preserves local topology information, such as edges and texture. Theoretical analysis and derivation given also validate the property of TPNMF. Experimental results on three different databases, containing more than 12,000 face images under varying in lighting, facial expression, and pose, show that the proposed TPNMF approach provides a better representation of face patterns and achieves higher recognition rates than NMF. Taiping Zhang, Bin Fang 0001, Yuan Yan Tang |
IEEE Trans. Image Process. | 3 |
| 2007 | Face recognition based on discriminant waveletfacesabstractPrincipal Component Analysis (PCA) and Linear Discriminant Analysis (LDA) have been widely applied in the face detection and recognition, yet still they have some limitations such as poor discriminative power and large computational load. This paper presents a method for face recognition using discriminant waveletfaces. Firstly wavelet transform is used to decompose the face image into different frequency subbands for extracting feature - waveletfaces, and then the discriminant analysis of PCA plus LDA is performed on the chosen subband. Finally the nearest neighbor classifier is adopted to make decision. In comparison with the traditional discriminant analysis, the experiments show that the proposed approach has better recognition rate and can reduce the computational load. Limin Cui, Yuan Yan Tang, Fucheng Liao, Xiufeng Du |
SMC | 2 |
| 2007 | Image segmentation based on the method of the maximal variance and improved genetic algorithmabstractAiming for the problem of falling into local optimum when searching for the optimal threshold of the image using normal genetic algorithm, this paper presents a new method based on the maximal variance and improved genetic algorithm to segment the face image. This new method uses the maximal variance of the face gray image as the fitness and changes the problem of image segmentation into a problem of optimization. Adopting genetic algorithm which is characteristic of robustness and adaptability can increase efficiency. As a result, this new method can obtain the optimal segmentation result when applied to different face images. Experiments show that using this method to search for the global threshold can converge the optimal value and decrease the searching time Jianjia Pan, Lanyan Xue, Shengling Zheng, Yuan Yan Tang |
SMC | 4 |
| 2007 | Offline signature verification: A new rotation invariant approachabstractRotation problem is one of the major difficulties to distinguish signature patterns in off-line skilled signature verification. This paper presents a new approach utilizing Ring- Peripheral features to tackle this problem. In principle, Ring- Peripheral features are able to describe internal and external structure of signatures with different phase shift. In order to extract stable and consistent presentation of signature patterns for verification purpose, FFT is used to eliminate phase effects. In the training samples stage, we employ a selection function to pick up reasonable samples for better threshold estimation. Experiment results demonstrated that the proposed method was successful to improve verification accuracy. Bin Fang 0001, Yuan Yan Tang, Taiping Zhang |
SMC | 3 |
| 2007 | Face contour location based on multiple-step hiding genetic algorithmabstractResearch on the interference within the face aiming at face complexity when locating the face, this paper advances a new algorithm of multiple-step hiding genetic algorithm to locate the face. This algorithm searches for intermediate locating parameters through genetic algorithm and parameterized deformable ellipse template, locates a hiding region based on the parameter and hides this region. On the basis of the hided image, searching for the next hiding region until finding the optimal region, and locate the contour ultimately based on the processing region. This algorithm can solve the interference within the face effectively when searching the face edge adopting genetic algorithm. Experimental results show that this method can obtain a satisfied result in the interference killing feature and stability. Lanyan Xue, Jianjia Pan, Baochang Pan, Shengling Zheng, Yuan Yan Tang |
SMC | 5 |
| 2007 | Remarks on different reviews of Chinese character recognition
Yuan Yan Tang |
Frontiers Comput. Sci. China | 1 |
| 2007 | A (4n-9)/3 diagnosis algorithm on n-dimensional cube network
Xiaofan Yang 0001, Yuan Yan Tang |
Inf. Sci. | 2 |
| 2007 | Efficient Fault Identification of Diagnosable Systems under the Comparison ModelabstractDiagnosis-by-comparison is a realistic approach to the fault diagnosis of massive multicomputers. This paper addresses the fault identification of diagnosable multicomputer systems under the MM* comparison model. We find that the fault location task can be reduced to that under the classical PMC* model. On this basis, we present an Ο(n×Δ3×δ) time diagnosis algorithm for an n-node MM* diagnosable system, where Δ and δ denote the maximum and minimum degrees of a node, respectively. The proposed algorithm is much more effi-cient than the fastest known diagnosis algorithm (which consumes Ο(n5) time) because realistic massive multi-computers are sparsely interconnected and hence Δ, δ « n. Xiaofan Yang 0001, Yuan Yan Tang |
IEEE Trans. Computers | 2 |
| 2007 | Wavelet-Based Approach to Character SkeletonabstractCharacter skeleton plays a significant role in character recognition. The strokes of a character may consist of two regions, i.e., singular and regular regions. The intersections and junctions of the strokes belong to singular region, while the straight and smooth parts of the strokes are categorized to regular region. Therefore, a skeletonization method requires two different processes to treat the skeletons in theses two different regions. All traditional skeletonization algorithms are based on the symmetry analysis technique. The major problems of these methods are as follows. 1) The computation of the primary skeleton in the regular region is indirect, so that its implementation is sophisticated and costly. 2) The extracted skeleton cannot be exactly located on the central line of the stroke. 3) The captured skeleton in the singular region may be distorted by artifacts and branches. To overcome these problems, a novel scheme of extracting the skeleton of character based on wavelet transform is presented in this paper. This scheme consists of two main steps, namely: a) extraction of primary skeleton in the regular region and b) amendment processing of the primary skeletons and connection of them in the singular region. A direct technique is used in the first step, where a new wavelet-based symmetry analysis is developed for finding the central line of the stroke directly. A novel method called smooth interpolation is designed in the second step, where a smooth operation is applied to the primary skeleton, and, thereafter, the interpolation compensation technique is proposed to link the primary skeleton, so that the skeleton in the singular region can be produced. Experiments are conducted and positive results are achieved, which show that the proposed skeletonization scheme is applicable to not only binary image but also gray-level image, and the skeleton is robust against noise and affine transform. Xinge You, Yuan Yan Tang |
IEEE Trans. Image Process. | 2 |
| 2006 | Handwriting-based personal identificationabstractHandwriting-based personal identification, which is also called handwriting-based writer identification, is an active research topic in pattern recognition. Despite continuous effort, offline handwriting-based writer identification still remains as a challenging problem because writing features can only be extracted from the handwriting image. As a result, plenty of dynamic writing information, which is very valuable for writer identification, is unavailable for offline writer identification. In this paper, we present a novel wavelet-based Generalized Gaussian Density (GGD) method for offline writer identification. Compared with the 2-D Gabor model, which is currently widely acknowledged as a good method for offline handwriting identification, GGD method not only achieves a better identification accuracy but also greatly reduces the elapsed time on calculation in our experiments. Zhenyu He 0001, Xinge You, Yuan Yan Tang, Bin Fang 0001, Jianwei Du |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2006 | EditorialabstractInternational Journal of Pattern Recognition and Artificial IntelligenceVol. 20, No. 02, pp. 111-112 (2006) SPECIAL ISSUE: Intelligent Computing and Applications; Edited by Y. Y. Tang and D.-S. HuangNo AccessEDITORIALYUAN YAN TANG and DE-SHUANG HUANGYUAN YAN TANGDepartment of Computer Science, Hong Kong Baptist University, Kowloon Tong, Kowloon, Hong Kong, China and DE-SHUANG HUANGInstitute of Intelligent Machines, Chinese Academy of Sciences, P.O. Box 1130, Hefei Anhui 230031, Chinahttps://doi.org/10.1142/S0218001406004545Cited by:0 Next AboutSectionsPDF/EPUB ToolsAdd to favoritesDownload CitationsTrack CitationsRecommend to Library ShareShare onFacebookTwitterLinked InRedditEmail FiguresReferencesRelatedDetails Recommended Vol. 20, No. 02 Metrics History PDF download Yuan Yan Tang, De-Shuang Huang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | Clustering Algorithm Research Based on Self-organizing Feature Maps NetworksabstractSelf-organizing feature maps (SOFM) can learn both the distribution and topology of the input vectors they are trained on. According to this characteristic, we construct neural networks with a family of self-organizing feature maps to cluster the input data space. The proposed algorithm in this paper defines a novel similarity measure, topological similarity, and employs some new concepts, such as SOFM family, UsageFactor. The clustering algorithm handles the clusters with arbitrary shapes and avoid the limitations of the conventional clustering algorithms. We conclude our paper by several experiments with synthetic and standard data set of different characteristics, which show good performance of the proposed algorithm. Junhao Wen 0001, Zhongfu Wu, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2006 | Thinning Character Using Modulus Minima of Wavelet TransformabstractAn essential step in character recognition is to extract the skeleton characteristics of the character. In this paper, an efficient algorithm is proposed to extract visually satisfactory skeleton from printed and handwritten characters, which overcomes fundamental shortcomings of our previous skeletonization technique based on the maximum modulus symmetry of wavelet transform (WT). The proposed method is motivated from some desirable properties of the WT with constructed wavelet functions: namely, the local modulus minima of the WT are scale-independent at different level scales and are located at the medial axis of the symmetrical contours of character stroke. Thus the modulus minima of the WT are computed as the intrinsic skeletons of character strokes. To achieve faster implementation, a multiscale processing technique is employed. Thus major structures of the skeleton are extracted using the coarse scale, while fine structures are extracted using the fine scale. We have tested the algorithm on handwritten and printed character images. Experimental results show that the proposed algorithm is applicable to not only binary image but also gray-level image where it can be impractical to use other skeletonization techniques, such as thinning and distance transforms. Further, it can effectively remove unwanted artifacts and branches from the extracted skeletons at the intersections and junctions of character strokes and is robust against noises while most existing methods perform poorly. Xinge You, Qiuhui Chen, Bin Fang 0001, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2005 | A Novel Method for Off-line Handwriting-based Writer IdentificationabstractHandwriting-based writer identification is a hot research topic in the pattern recognition field. Nowadays, online handwriting-based writer identification is steadily growing toward its maturity. On the contrary, offline handwriting-based writer identification still remains as a challenging problem because writing features only can be extracted from the handwriting image in this situation. As a result, plenty of dynamic writing information, which is very valuable for writer identification, is lost. At present, 2D Gabor filter method is widely acknowledged as a good method for offline handwriting identification, however it still suffers from some inherent disadvantages, such as the high computational cost. In this paper, we present a novel wavelet-based GGD method to replace the traditional 2D Gabor filters. Shown in our experiments, this novel method not only achieves better experiment results but also greatly reduces the elapsed time on calculation. Zhenyu He 0001, Yuan Yan Tang, Bin Fang 0001, Jianwei Du, Xinge You |
ICDAR | 2 |
| 2005 | Similarity Measurement for Off-Line Signature Verification
Xinge You, Bin Fang 0001, Zhenyu He 0001, Yuan Yan Tang |
ICIC (1) | 4 |
| 2005 | Locating Vessel Centerlines in Retinal Images Using Wavelet Transform: A Multilevel Approach
Xinge You, Bin Fang 0001, Yuan Yan Tang, Zhenyu He 0001, Jian Huang 0009 |
ICIC (1) | 3 |
| 2005 | Existence and Stability of Periodic Solution in a Class of Impulsive Neural Networks
Xiaofan Yang 0001, David J. Evans 0001, Yuan Yan Tang |
ISNN (1) | 3 |
| 2005 | A contourlet-based method for writer identificationabstractHandwriting-based writer identification is a hot research topic in the field of pattern recognition. Typically, there are four modes of writer identification: on-line text-dependent, on-line text-independent, off-line text-dependent, off-line text-independent; and off-line text-independent is the most challenging problem among them because many valuable writing features are not available in this case, such as shape features, dynastic writing information and etc. In this paper, we focus on the text-independent writer identification based on off-line Chinese handwriting and present a new contourlet-based GGD (Generalized Gaussian Density) method. This novel method achieves a good experiment result in our experiments. Zhenyu He 0001, Yuan Yan Tang, Xinge You |
SMC | 2 |
| 2005 | An uncorrelated fisherface approach
Xiaoyuan Jing, Hau-San Wong, David Zhang 0001, Yuan Yan Tang |
Neurocomputing | 4 |
| 2005 | Morphological structure reconstruction of retinal vessels in fundus imagesabstractVessels in retinal fundus images are useful in revealing the severity of eye-related diseases. In addition, they can act as landmarks for localizing lesions or the central vision area, and guide laser treatment of neovascularization. In this paper, we propose a two-stage scheme to extract vessels and reconstruct the morphological structure of vessels in retinal images. First, we employ mathematical morphology techniques to highlight large and small vessels with respect to their spatial properties. Different curvature response between vessel and noise patterns allows the use of curvature evaluation to remove enhanced vessel-like noise. A set of linear filters finalize the vessel map. However, the resulting vascular structure is incomplete of some important features in bifurcation points and central reflex. In order to rectify the pitfall, a reconstruction process is performed using dynamic local region growth to recover the morphological structure of vessels. Average performance of our method to extract vessels is 83.7% of TPR(True positive rate) and 3.8% of FPR(False positive rate) for 35 retinal images which include 21 abnormal images. Bin Fang 0001, Xinge You, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2005 | Universal writing model for recovery of writing sequence of static handwriting imagesabstractOnline features have been proven to be more robust information for handwriting recognition than an offline static image due to dynamic aspects, such as the writing sequence of strokes. The estimation of temporal information from a static image becomes an important issue. This paper presents a new statistical method to reconstruct the writing order of a handwritten signature from a two-dimensional static image. The reconstruction process consists of two phases, namely the training phase and the testing phase. In the training phase, the writing order with other attributes, such as length and direction, are extracted and analyzed from a set of training online handwritten signatures. A Universal Writing Model (UWM), which consists of a set of distribution functions, is then constructed. In the testing phase, the UWM is applied to reconstruct the writing order of an offline signature. 300 offline signatures with ground truth are used for evaluation. Experimental results show that about one-eighth of the reconstructed writing sequences are the same as the actual writing sequences. Kai Kwong Lau, Pong C. Yuen, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2005 | Signal denoising using wavelets and block hidden Markov modelabstractThis paper presents a new framework for signal denoising based on wavelet-domain hidden Markov models (HMMs). The new framework enables us to concisely model the statistical dependencies and non-Gaussian statistics encountered in real-world signals, and enables us to get a more reliable and local model using blocks. Wavelet-domain HMMs are designed with the intrinsic properties of wavelet transform and provide powerful yet tractable probabilistic signal models. In this paper, we propose a novel wavelet domain HMM using blocks to strike a delicate balance between improving spatial adaptability of contextual HMM (CHMM) and modeling a more reliable HMM. Each wavelet coefficient is modeled as a Gaussian mixture model, and the dependencies among wavelet coefficients in each subband are described by a context structure, then the structure is modified by blocks which are connected areas in a scale conditioned on the same context. Before denoising a signal, efficient Expectation Maximization (EM) algorithms are developed for fitting the HMMs to observational signal data. Parameters of trained HMM are used to modify wavelet coefficients according to the rule of minimizing the mean squared error (MSE) of the signal. Then, reverse wavelet transformation is utilized to modified wavelet coefficients. Finally, experimental results are given. The results show that block hidden Markov model (BHMM) is a powerful yet simple tool in signal denoising. Zhiwu Melody Liao, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | A wavelet-based approach to ridge thinning in fingerprint imagesabstractAs a global feature of fingerprints, the thinning of ridges, extraction of minutiae and computation of orientation field are very important for automatic fingerprint recognition. Many algorithms have been proposed for their computation and estimation, but their results are unsatisfactory, especially for poor quality fingerprint images. In this paper, a robust wavelet-based method to create thinned ridge map of fingerprint for automatic recognition is proposed. Properties of modulus minima based on the spline wavelet function are substantially investigated. Desirable characteristics show that this method is suitable to describe the skeleton of the ridge of the fingerprint image. A multi-scale thinning algorithm based on the modulus minima of wavelet transform is presented. The proposed algorithm is able to improve the skeleton representation of the ridge of the fingerprint without side-effects and limitations of the existing methods. The thinned ridge map can facilitate the extraction of the minutiae for matching in fingerprint recognition. Experiments have been conducted to validate the effectiveness and efficiency of the proposed method. Xinge You, Bin Fang 0001, Yuan Yan Tang, Zhenyu He 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2005 | A Fourier-LDA approach for image recognition
Xiaoyuan Jing, Yuan Yan Tang, David Zhang 0001 |
Pattern Recognit. | 2 |
| 2005 | Directed connection measurement for evaluating reconstructed stroke sequence in handwriting images
Kai Kwong Lau, Pong C. Yuen, Yuan Yan Tang |
Pattern Recognit. | 3 |
| 2005 | Improved class statistics estimation for sparse data problems in offline signature verificationabstractSparse data problems are prominent in applications of offline signature verification. By using a small number of training samples, the class statistics estimation errors may be significant, resulting in worsened verification performance. In this paper, we propose two methods to improve the statistics estimation. The first approach employs an elastic distortion model to artificially generate additional training samples for pairs of genuine signatures. These additional samples, together with original genuine samples, are used to compute statistic parameters for a Mahalanobis distance threshold classifier. The other approach is to adopt regularization techniques to overcome the problem of inverting an ill-conditioned sample covariance matrix due to insufficient training samples. A ridge-like estimator is modeled to add some constant values for diagonal elements of the sample covariance matrix. Experimental results showed that both methods were able to improve verification accuracy when they were incorporated with a set of peripheral features. Effectiveness of the methods was validated by quantity analysis. Bin Fang 0001, Yuan Yan Tang |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2004 | Fingerprint Enhancement Using Wavelet Transform Combined With Gabor FilterabstractThe performance of automatic fingerprint identification system (AFIS) is heavily determined by the quality of the input image, thus an effective method to enhance the fingerprint image is essential in such a system. In this paper, we combine the filter-based method, which is mostly used nowadays with wavelet transform to achieve a more reliable and effective approach to fingerprint enhancement. This novel approach consists of five main steps, namely: (1) normalization, (2) decomposition, (3) wavelet coefficient adjustment, (4) Gabor filtering, and (5) reconstruction. Using this new method, a more clear fingerprint image can be obtained, which can distinctly improve the accuracy of the minutiae extraction module and finally achieve a better performance of the entire system. Experiments have been conducted in our study and positive experimental results have been received, which show that the proposed combined method is more effective and robust than other existing methods such as the filter-based and direct gray-level approaches. Yuan Yan Tang, Xinge You |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2004 | An improved LDA approachabstractLinear discrimination analysis (LDA) technique is an important and well-developed area of image recognition and to date many linear discrimination methods have been put forward. Despite these efforts, there persist in LDA at least three areas of weakness. The first weakness is that not all the discrimination vectors that are obtained are useful in pattern classification. Second, it remains computationally expensive to make the discrimination vectors completely satisfy statistical uncorrelation. The third weakness is that it is necessary to select the appropriate principal components. In this paper, we propose to improve discrimination technique in these three areas and to that end present an improved LDA (ILDA) approach which synthesizes these improvements. Experimental results on different image databases demonstrate that our improvements on LDA are efficient, and that ILDA outperforms other state-of-the-art linear discrimination methods. Xiaoyuan Jing, David Zhang 0001, Yuan Yan Tang |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2003 | Skeletonization of Character Based on Wavelet Transform
Xinge You, Yuan Yan Tang |
CAIP | 2 |
| 2003 | Recovery of Writing Sequence of Static Images of Handwriting using UWMabstractIt is generally agreed that an on-line recognition system is always reliable than an off-line one. It is due to the availability of the dynamic information, especially the writing sequence of the strokes. This paper presents a new statistical method to reconstruct the writing order of a handwritten script from a two-dimensional static image. The reconstruction process consists of two phases, named the training phase and the testing phase. In the training phase, the writing order with other attributes, such as length and direction, are extracted from a set of training on-line handwritten scripts statistically to form a universal writing model (UWM). In the testing phase, UWM is applied to reconstruct the drawing order of offline handwritten scripts by finding the highest total probability. 300 off-line signatures with ground truth are used for evaluation. Experimental results show that the reconstructed writing sequence using UWM is close to the actual writing sequence. 1. Kai Kwong Lau, Pong C. Yuen, Yuan Yan Tang |
ICDAR | 3 |
| 2003 | Feature Extraction of Radar Multiple-Target Echoes Using Wavelet Packet Transform with the Best BasesabstractExtraction of effective features plays a key role in pattern recognition. A large number of patterns, such as speech, radar signals, earthquake signals, handwriting, etc. are of non-stationary signals or exhibit time-varying behavior. The features of these patterns are often located in both the time and frequency domains. The traditional methods fail to extract such kind of features. Fortunately, wavelet packet transform (WPT) can provide an arbitrary time-frequency decomposition for the signals, because a wavelet packet (WP) library contains many WP bases, which can handle the different components of a signal. Therefore, by selecting a suitable basis, which is called "best basis", the effective features can be extracted. In this paper, three criteria are used to select the best WPT basis, namely: (1) distance criterion, (2) divergence criterion and (3) entropy criterion. Three algorithms to implement the above criteria are also provided. Experiments are conducted and the positive results are obtained. Shouyong Wang, Guangxi Zhu, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2003 | Skeletonization of Ribbon-Like Shapes Based on a New Wavelet FunctionabstractA wavelet-based scheme to extract skeleton of Ribbon-like shape is proposed in this paper, where a novel wavelet function plays a key role in this scheme, which possesses three significant characteristics, namely, 1) the position of the local maximum moduli of the wavelet transform with respect to the Ribbon-like shape is independent of the gray-levels of the image. 2) When the appropriate scale of the wavelet transform is selected, the local maximum moduli of the wavelet transform of the Ribbon-like shape produce two new parallel contours, which are located symmetrically at two sides of the original one and have the same topological and geometric properties as that of the original shape. 3) The distance between these two parallel contours equals to the scale of the wavelet transform, which is independent of the width of the shape. This new scheme consists of two phases: 1) Generation of wavelet skeleton-based on the desirable properties of the new wavelet function, symmetry analyses of the maximum moduli of the wavelet transform is described. Midpoints of all pairs of contour elements can be connected to generate a skeleton of the shape, which is defined as wavelet skeleton. 2) Modification of the wavelet skeleton. Thereafter, a set of techniques are utilized for modifying the artifacts of the primary wavelet skeleton. The corresponding algorithm is also developed in this paper. Experimental results show that the proposed scheme is capable of extracting exactly the skeleton of the Ribbon-like shape with different width as well as different gray-levels. The skeleton representation is robust against noise and affine transformation. Yuan Yan Tang, Xinge You |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Off-line signature verification by the tracking of feature and stroke positions
Bin Fang 0001, Cheung Hoi Leung, Yuan Yan Tang, K. W. Tse 0001, Paul C. K. Kwok |
Pattern Recognit. | 3 |
| 2003 | A width-invariant property of curves based on wavelet transform with a novel wavelet functionabstractThis paper is an improvement on the characterization of edges. Using a novel wavelet function, it is proven that the maximum moduli of the wavelet transform (MMWT) of a curve produces two new symmetrical curves on both sides of the original with the same direction. The distance between the two curves is shown to be independent of the width d of the original curve if the scale s of the wavelet transform satisfies s/spl ges/d. This property provides a novel method of obtaining the skeletons of the curves in an image. Lihua Yang 0001, Ching Y. Suen, Yuan Yan Tang |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2002 | Multi-agent oriented constraint satisfaction
Jiming Liu 0001, Han Jing, Yuan Yan Tang |
Artif. Intell. | 3 |
| 2002 | Order Statistic Filter (OSF): A Novel Approach to Document AnalysisabstractPage segmentation is one of the important and basic research subjects of document analysis. There are two major kinds of page segmentation methods, i.e. hierarchical and no-hierarchical ones. Most traditional techniques such as top–down and bottom–up approaches belong to the hierarchical method. Though these two approaches have been used till now, they are not effective for processing documents with high geometric complexity and the process of splitting document needs iterative operations which is time consuming. A non-hierarchical method called the modified fractal signature (MFS) was presented in recent years. It can overcome the above weaknesses, however the MFS needs to calculate modified fractal signature which makes the theory very complex. In this thesis, we present a new page segmentation approach: Median Order Statistic Filter (MedOSF) — Maximum Order Statistic Filter (MaxOSF) approach which is more direct and much simpler. We use the MedOSF to remove the salt–pepper noise of the document and use the MaxOSF to do the page segmentation. In practice, they not only can adaptively process the documents with high geometrical complexity, but also save a lot of computing time. Hong Ma 0001, Jie Zhou 0002, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2002 | Mathematical Representation of a Chinese Character and its ApplicationsabstractIn this paper, a novel method to express Chinese characters mathematically is presented based on the knowledge of the structure of Chinese characters. Each Chinese character can be denoted by a mathematical expression in which the operands are components of Chinese characters and the operators are the location relations between the components. Five hundred five components are selected and 6 operators are defined to express all the Chinese characters successfully. These mathematical expressions of Chinese characters are simple, natural, and can be operated like the common mathematical expression of numbers. It makes Chinese information processing much simpler than before. This theory has been applied successfully in fonts automation, Chinese information transmission among different platforms and different operating systems on Internet, and knowledge discovery of the structure of Chinese characters. It can also be applied extensively to many areas such as typesetting, advertising, packing design, virtual library, network transmission, pattern recognition and Chinese mobile communication. Xingming Sun, Huowang Chen, Lihua Yang 0001, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2002 | Sequential Combination Methods for Data Clustering Analysis
Yuntao Qian, Ching Y. Suen, Yuan Yan Tang |
J. Comput. Sci. Technol. | 3 |
| 2002 | New method for feature extraction based on fractal behavior
Yuan Yan Tang, Yu Tao 0001, Ernest C. M. Lam |
Pattern Recognit. | 1 |
| 2002 | Two-channel adaptive biorthogonal filterbanks via lifting
Penglang Shui, Zheng Bao 0001, Xian-Da Zhang, Yuan Yan Tang |
Signal Process. | 4 |
| 2001 | Text Area Localization under Complex-Background Using Wavelet DecompositionabstractIn this paper, we propose a novel approach to determine the positions of text areas in images with complex background using wavelet decomposition and pseudo-motion of images. In our method, a fixed image is translated and provides a sequence of moving images-the pseudomotion sequence of image. In fact, we can consider the translation of an image to be a motion of eyeshot. When an image is translated, its wavelet coefficients will oscillate. From this property, we can locate the text areas in complex-background images. Experiments were conducted to demonstrate the performance of the method In the experiments, we detect the text areas of several different types of characters in images with multi-gray level complex backgrounds. Yuan Yan Tang |
ICDAR | 2 |
| 2001 | Discrimination of Oriental and Euramerican Scripts Using Fractal FeatureabstractThis paper presents a new approach based on modified fractal signatures (MFS) and modified fractal features (MFF)for the discrimination of Oriental and Euramerican scripts. These methods will be useful in the measurement and classification of patterns. MFS do not need iterative breaking or merging, and can divide a document into blocks in a single step. MFF is also used in the identification and classification of a selected set of texture images with good results. It is anticipated that this approach could be widely used to process various types of documents, even including some with high geometrical complexity. Yu Tao 0001, Yuan Yan Tang |
ICDAR | 2 |
| 2001 | Extraction of rotation invariant signature based on fractal geometryabstractA new method of feature extraction with a rotation invariant property is presented. One of the main contributions of this study is that a rotation invariant signature of 2D contours is selected based on fractal theory. The rotation invariant signature is a measure of the fractal dimensions, which is rotation invariant based on a series of central projection transform (CPT) groups. As the CPT is applied to a 2D object, a unique contour is obtained. In the unfolding process, this contour is further spread into a central projection unfolded curve, which can be viewed as a periodic function due to the different orientations of the pattern. We consider the unfolded curves to be non-empty and bounded sets in IR/sup n/, and the central projection unfolded curves with respect to the box computing dimension are rotation invariant. Some experiments with positive results have been conducted. This approach is applicable to a wide range of areas such as image analysis, pattern recognition etc. Yu Tao 0001, Thomas R. Ioerger, Yuan Yan Tang |
ICIP (1) | 3 |
| 2001 | Offline Signature Verification by the Analysis of Cursive StrokesabstractIn this paper, a method is proposed for offline signature verification. It is based on a smoothness criterion. It is observed that the cursive segments of forgery signatures are generally less smooth and less natural than the genuine ones, especially for those signatures that consist of cursive graphic patterns. Two approaches are proposed to extract a smoothness feature: a crossing method and a fractal dimension method. When the proposed smoothness feature is combined with other global shape features for signature verification, satisfactory results are obtained. Bin Fang 0001, Y. Y. Wang, Cheung Hoi Leung, K. W. Tse 0001, Yuan Yan Tang, Paul C. K. Kwok |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2001 | Classification of Similar 2-D Objects by Wavelet-Sparse-Matrix (WSM) MethodabstractThis paper proposes a novel method called Wavelet-Sparse-Matrix (WSM) to extract the spatial features of 2-D objects for classifying objects that have subtle differences. The differences between these objects are present in the spatial orientations of the objects, or in the local positions of points on the contours of the objects. The separable wavelets are able to distinguish these differences and to separate them into three sparse subpatterns. Sparse matrix technique has the ability to rearrange nonzero elements in a sparse matrix by moving them as close together as possible. WSM method is a combination of these two techniques which can considerably improve the distinction of slightly dissimilar objects. Experiments are conducted, which include a series of discriminative simulations and comparisons with Fourier descriptor and Zernike moment invariant. These experiments verify the feasibility and effectiveness of the WSM method. L. Feng, Tien D. Bui, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2001 | Intelligent Agent Technology - Introduction
Jiming Liu 0001, Ning Zhong 0001, Yuan Yan Tang, Patrick Shen-Pei Wang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2001 | A New Stroke Extraction Method of Chinese CharactersabstractStroke extraction of Chinese characters plays an important role in Chinese character information processing such as character recognition, document analysis, document compression and storage, font automation and so on. By analyzing the structure of Chinese characters deeply, this paper developed a novel method to extract strokes of Chinese characters directly from the original character pattern image. Two theorems, eight rules and an algorithm for stroke extraction of Chinese characters are presented. This method can overcome the difficulties encountered in disposing the intersection or connection of different strokes, and can eliminate noises successfully. Our experiments have shown that this method can extract strokes both accurately and efficiently. Xingming Sun, Lihua Yang 0001, Yuan Yan Tang, Yunfa Hu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2001 | A Combination of Fractal and Wavelet for Feature ExtractionabstractIn this paper, a novel approach to feature extraction with wavelet and fractal theories is presented as a powerful technique in pattern recognition. The motivation behind using fractal transformation is to develop a high-speed feature extraction technique. A multiresolution family of the wavelets is also used to compute information conserving micro-features. In this study, a new fractal feature is reported. We employed a central projection method to reduce the dimensionality of the original input pattern, and a wavelet transform technique to convert the derived pattern into a set of subpatterns, from which the fractal dimensions can readily be computed. The new feature is a measurement of the fractal dimension, which is an important characteristic that contains information about the geometrical structure. This new scheme includes utilizing the central projection transformation to describe the shape, the wavelet transformation to aid the boundary identification, and the fractal features to enhance image discrimination. The proposed method reduces the dimensionality of a 2-D pattern by way of a central projection approach, and thereafter, performs Daubechies' wavelet transform on the derived 1-D pattern to generate a set of wavelet transform subpatterns, namely, curves that are non-self-intersecting. Further from the resulting non-self-intersecting curves, the divider dimensions are computed with a modified box-counting approach. These divider dimensions constitute a new feature vector for the original 2-D pattern, defined over the curve's fractal dimensions. We have conducted several experiments in which a set of printed Chinese characters, English letters of varying fonts and other images were classified. Based on the formulation of our new feature vector, the experiments have satisfying results. Yu Tao 0001, Ernest C. M. Lam, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2001 | Basic Processes of Chinese Character Based on Cubic B-Spline Wavelet TransformabstractA novel approach based on cubic B-spline wavelet transform is proposed to process Chinese character including character compression, type zooming-in, and typeface composition. The basic idea is that a Chinese character is described by its contours which are represented by cubic B-spline functions, and each contour is decomposed into the details or the control points (wavelet coefficients) at different resolution levels. For character compression, there are two methods, one directly treats the details of wavelet coefficients and the other considers the sub-curves piecing together at the different resolution levels. In the type zooming-in, the wavelet reconstruction is used to scale the Chinese character with arbitrary size and the wavelet filter is used to improve the quality of the enlarged type. For typeface composition, the new style typefaces of Chinese character are obtained by editing and modifying the details at different resolution levels. The concrete algorithms are also given as well as the experimental results. Yuan Yan Tang, Jiming Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Feature extraction using wavelet and fractal
Yu Tao 0001, Ernest C. M. Lam, Yuan Yan Tang |
Pattern Recognit. Lett. | 3 |
| 2000 | EDT Based Tracing Maximum Thinning Algorithm on Grey Scale ImagesabstractMost of the thinning algorithm nowadays are based on bilevel images. In recognition of hand-written words, such as signatures, the use of grey-scale image is better because much data is available from the image. In this article, we propose an efficient thinning algorithm based on Euclidean distance transformation (EDT) on grey level image. The output of our algorithm consists of the skeleton of the object as well as the intersection points. Moreover, the proposed algorithm is efficient and accurate in finding the skeleton. Kai Kwong Lau, Pong C. Yuen, Yuan Yan Tang |
ICPR | 3 |
| 2000 | Extraction of Fractal Feature for Pattern RecognitionabstractAn approach to feature extraction based on fractal theory is presented as a powerful technique in pattern recognition. It can be used to extract the features of 2D objects, and identify different scripts. A fractal feature and fractal signature are reported. A multiresolution family of the wavelets is also used to compute information conserving micro-features. We employed a central projection method to reduce the dimensionality of the original input pattern, and a wavelet transformation technique to transform the derived pattern into a set of sub-patterns, from which the fractal dimension can readily be computed. Moreover, we have proposed an approach to classify different language using the modified fractal signature. For all these cases, difference in fractal dimension can yield the significative values. We expect that the proposed fractal method can also be used for improving the extraction and classification of features in a pattern recognition system. Yu Tao 0001, Ernest C. M. Lam, Yuan Yan Tang |
ICPR | 3 |
| 2000 | Edge Extraction of Images by Reconstruction Using Wavelet Decomposition Details at Different Resolution LevelsabstractThis paper describes a novel method for edge feature detection of document images based on wavelet decomposition and reconstruction. By applying the wavelet decomposition technique, a document image becomes a wavelet representation, i.e. the image is decomposed into a set of wavelet approximation coefficients and wavelet detail coefficients. Discarding wavelet approximation, the edge extraction is implemented by means of the wavelet reconstruction technique. In consideration of the mutual frequency, overlapping will occur between wavelet approximation and wavelet details, a multiresolution-edge extraction with respect to an iterative reconstruction procedure is developed to ameliorate the quality of the reconstructed edges in this case. A novel combination of this multiresolution-edge results in clear final edges of the document images. This multi-resolution reconstruction procedure follows a coarser-to-finer searching strategy. The edge feature extraction is accompanied by an energy distribution estimation from which the levels of wavelet decomposition are adaptively controlled. Compared with the scheme of wavelet transform, our method does not incur any redundant operation. Therefore, the computational time and the memory requirement are less than those in wavelet transform. L. Feng, Ching Y. Suen, Yuan Yan Tang, Lihua Yang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2000 | An Improved Embedded Zerotree Wavelet Image Coding Method Based on Coefficient Partitioning Using Morphological OperationabstractIn recent years, wavelets have attracted great attention in both still image compression and video coding, and several novel wavelet-based image compression algorithms have been developed so far, one of which is Shapiro's embedded zerotree wavelet (EZW) image compression algorithm. However, there are still some deficiencies in this algorithm. In this paper, after the analysis of the deficiency in EZW, a new algorithm based on quantized coefficient partitioning using morphological operation is proposed. Instead of encoding the coefficients in each subband line-by-line, regions in which most of the quantized coefficients are significant are extracted by morphological dilation and encoded first. This is followed by using zerotrees to encode the remaining space which has mostly zeros. Experimental results show that the proposed algorithm is not only superior to the EZW, but also compares favorably with the most efficient wavelet-based image compression algorithms reported so far. Junmei Zhong, Cheung Hoi Leung, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2000 | Characterization of Dirac-structure edges with wavelet transformabstractThis paper aims at studying the characterization of Dirac-structure edges with wavelet transform, and selecting the suitable wavelet functions to detect them. Three significant characteristics of the local maximum modulus of the wavelet transform with respect to the Dirac-structure edges are presented: (1) slope invariant: the local maximum modulus of the wavelet transform of a Dirac-structure edge is independent on the slope of the edge; (2) grey-level invariant: the local maximum modulus of the wavelet transform with respect to a Dirac-structure edge takes place at the same points when the images with different grey-levels are processed; and (3) width light-dependent: for various widths of the Dirac-structure edge images, the location of maximum modulus of the wavelet transform varies lightly under the certain circumscription that the scale of the wavelet transform is larger than the width of the Dirac-structure edges. It is important, in practice, to select the suitable wavelet functions, according to the structures of edges. For example, Haar wavelet is better to represent brick-like images than other wavelets. A mapping technique is applied in this paper to construct such a wavelet function. In this way, a low-pass function is mapped onto a wavelet function by a derivation operation. In this paper, the quadratic spline wavelet is utilized to characterize the Dirac-structure edges and a novel algorithm to extract the Dirac-structure edges by wavelet transform is also developed. Yuan Yan Tang, Lihua Yang 0001, Jiming Liu 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1999 | A Smoothness Index based Approach for Off-line Signature VerificationabstractProposes a method to tackle the problem of detecting skilled forgeries in off-line signature verification. Inspired by the approach adopted by expert examiners, it is based on a smoothness criterion. From a collection of genuine and forged signatures, it is observed that, although skilled forgery signatures are very similar to genuine ones on a global scale, they are generally less smooth and natural on a detailed scale than the genuine ones, especially for those skilled forgery signatures which consist of cursive graphic patterns. A smoothness index is derived from such signatures. This is combined with other global shape features and used for verification. Satisfactory results are obtained. Bin Fang 0001, Y. Y. Wang, Cheung Hoi Leung, Yuan Yan Tang, Paul C. K. Kwok, K. W. Tse 0001 |
ICDAR | 4 |
| 1999 | Accelerating the 2-D Mallat Decomposition Algorithm with Cyclical Convolution and FNTTabstractThis paper presents a novel approach to accelerate the 2D Mallat decomposition algorithm. In particular the proposed approach performs the 2D Mallat decomposition algorithm by way of computing the 2D cyclical convolution, and further completes the 2D cyclical convolution with the fast number theoretical transform (FNTT). Therefore, the 2D Mallat decomposition algorithm can be speeded up. The results obtained from the theoretical analysis and the experiment show that the 2D Mallat decomposition algorithm is significantly improved. Yuan Yan Tang, Hong Ma 0001, De B. Ren, Zhi K. Chen |
ICDAR | 1 |
| 1999 | Feature Extraction by Fractal DimensionsabstractProposes a method that reduces the dimensionality of a 2D pattern by means of a central projection approach, and thereafter performs a Daubechies wavelet transformation on the derived 1D pattern to generate a set of wavelet transformation sub-patterns, namely curves that are non-self-intersecting. Further, from the resulting non-self-intersecting curves, the divider dimensions are compared with the modified box-counting approach. These divider dimensions constitute a new feature vector for the original 2D pattern, defined over the curve's fractal dimensions. Yuan Yan Tang, Yu Tao 0001 |
ICDAR | 1 |
| 1999 | The Feature Extraction of Chinese Character based on Contour InformationabstractA new method, called central projection transformation, is proposed in this paper for feature extraction. From our experiments, the new method is found to be efficient in extracting features based on the contours of Chinese characters. Chinese characters have complex structures, and some of them are composed of several separate components, so several contours are embedded in a character. This may obstruct the application of the contour approach in recognizing Chinese characters. Central projection transformation can convert such a multi-contour pattern into a solid, convex pattern whose contour is a unique polygon. Most of the information of this new pattern is still located around its periphery. This approach can greatly simplify the processing of Chinese characters and other multi-contour patterns. It is also a powerful tool for processing Arabic, Japanese and other characters. Yu Tao 0001, Yuan Yan Tang |
ICDAR | 2 |
| 1999 | A Novel Method for Discriminating between Oriental and European Languages by Fractal FeaturesabstractA new method that uses a modified fractal dimension theory to segment a document image and to discriminate between Oriental and European languages is presented in this paper. Two types of techniques have been usually adopted in language discrimination: token matching and statistical analysis. A modified fractal feature is used to discriminate the distinct textual structure complexities of Oriental and European languages. Experiments show that this method is effective and reliable for processing the document image even if it is skewed or contains noise that can not be removed clearly. Dihua Xi, Seong-Whan Lee, Yuan Yan Tang |
ICDAR | 3 |
| 1999 | The Application of Fractal Analysis to Feature ExtractionabstractAs the interest in fractal geometry rises, the applications are getting more and more numerous in many domains. The aim of the authors is that these concepts can also be applied to feature extraction of patterns and that they can help, to a certain extent, to ease the solution of many problems. In this paper, the proposed method reduces the dimensionality of a two-dimensional pattern by way of a central projection approach, and thereafter, performs Daubechies' wavelet transformation on the derived one-dimensional pattern to generate a set of wavelet transformation sub-patterns, namely, curves that are non-self-intersecting. Further from the resulting nonself-intersecting curves, the divider dimensions are computed with modified box-counting approach. These divider dimensions constitute a new feature vector for the original two-dimensional pattern, defined over the curve's fractal dimensions. Yuan Yan Tang, Yu Tao 0001, Ernest C. M. Lam |
ICIP (2) | 1 |
| 1999 | Wavelet Orthonormal Decompositions for Extracting Features in Pattern RecognitionabstractIn this paper, a novel approach based on the wavelet orthonormal decomposition is presented to extract features in pattern recognition. The proposed approach first reduces the dimensionality of a two-dimensional pattern, and thereafter performs wavelet transform on the derived one-dimensional pattern to generate a set of wavelet transform subpatterns, namely, several uncorrelated functions. Based on these functions, new features are readily computed to represent the original two-dimensional pattern. As an application, experiments were conducted using a set of printed characters with varying orientations and fonts. The results obtained from these experiments have consistently shown that the proposed feature vectors can yield an excellent classification rate in pattern recognition. Yuan Yan Tang, Jiming Liu 0001, Hong Ma 0001, Bing F. Li |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1999 | Adaptive Image Segmentation With Distributed Behavior-Based AgentsabstractPresents an autonomous agent-based image segmentation approach. In this approach, a digital image is viewed as a two-dimensional cellular environment which the agents inhabit and attempt to label homogeneous segments. In so doing, the agents rely on some reactive behaviors such as breeding and diffusion. The agents that are successful in finding the pixels of a specific homogeneous segment will breed offspring agents inside their neighboring regions. Hence, the offspring agents will become likely to find more homogeneous-segment pixels. In the mean time, the unsuccessful agents will be inactivated, without further search in the environment. Jiming Liu 0001, Yuan Yan Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Characterization and detection of edges by Lipschitz exponents and MASW wavelet transformabstractThis paper is an improvement of Mallat's work (1992). The characterization of edges by Lipschitz exponents and wavelet transform has been studied in this paper A significant property has been proved and applied to identify different structures of edges, and an algorithm with the modulus-angle-separated wavelets (MASW) has been developed in this paper to extract step-structure edges from a multistructure-edges image effectively. Experiments have been conducted and excellent results have been achieved. Yuan Yan Tang, Lihua Yang 0001, L. Feng |
ICPR | 1 |
| 1998 | An improved zerotree wavelet image coder based on significance checking in wavelet treesabstractThe Embedded Zerotree Wavelet (EZW) image compression algorithm has been widely used in real applications for its high compression performance. In this paper an improvement of EZW is presented. In the original EZW algorithm, when a new significant coefficient is generated, its children are all encoded, although its descendants maybe all insignificant, and thus its performance is affected. The improvement proposed in this paper is based on significance checking in wavelet trees (SCIWT). It is aimed to avoid encoding the children of each newly generated significant coefficient if it has no significant descendant. Experiments show that this proposed algorithm not only outperforms the original EZW over a wide range of compression ratios, but also completely retains all its key features without introducing any sophisticated and computationally complex method. J. M. Zhong, Cheung Hoi Leung, Yuan Yan Tang |
SMC | 3 |
| 1998 | A Reliability Design Methodology for Chinese Character RecognitionabstractThis paper proposes a novel method which enables a Chinese character recognition system to obtain reliable recognition. In this method, two thresholds, i.e. class region thresholdRk and disambiguity thresholdAk, are used by each Chinese character k when the classifier is designed based on the nearest neighbor rule, where Rk defines the pattern distribution region of character k, and Ak prevents the samples not belonging to character k from being ambiguously recognized as character k. A novel algorithm to derive the appropriate thresholds Ak and Rk is developed so that a better recognition reliability can be obtained through iterative learning. Experiments performed on the ITRI printed Chinese character database have achieved highly reliable recognition performance (such as 0.999 reliability with a 95.14% recognition rate), which shows the feasibility and effectiveness of the proposed method. Yea-Shuan Huang, Ching Y. Suen, Ke Liu 0009, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 1998 | Distributed Autonomous Agents for Chines Document Images SegmentationabstractIn Chinese document image processing, text and/or graphical block detection serves as an essential step in document layout analysis that in turn permits the effective reasoning about the logical relationships among various text paragraphs and graphical entities for the purpose of document understanding. This paper presents a novel computational paradigm for extracting text/graphic blocks from Chinese document images, which is based on a notion of distributed autonomous agents. The primary features of the agents lie in that they are (1) adaptive to the locality of given images and hence efficient in locating the homogeneous image blocks, (2) reliable in performing image processing as the computation proceeds simultaneously from different image locations, (3) less sensitive to the noise in the given images as the computation disperses gracefully when it is moving away from the homogeneous blocks, and (4) easy to represent in their behaviors and evolvable in their performance. The paper, first of all, describes the formalisms as well as the behavioral characteristics of the agents, which is followed by a demonstration of the agents in detecting document blocks from some real-life images. Jiming Liu 0001, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1998 | Printed Chinese Character Similarity Measurement Using Ring Projection and Distance TransformabstractThis paper presents a new Chinese character similarity measurement method based on the ring projection algorithm and distance transform. The ring projection algorithm is used to transform a character image with two independent variables into a function of one independent variable in the ring projection space. This representation of character in the ring projection space has been proved to be in orientation and scale invariant. However, this representation will be distorted nonlinearly in the presence of noise. Therefore, common linear metrics such as Euclidean distance, cannot be applied to measure distance. To solve the nonlinear distortion problem, distance transform is proposed as a nonlinear metric. The similarity measurement is performed using the distance transformed image in the ring projection space. A number of Chinese characters are selected to evaluate the capability of the proposed measurement scheme and the results are encouraging. Pong C. Yuen, Guo-Can Feng, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1998 | Offline Recognition of Chinese Handwriting by Multifeature and Multilevel ClassificationabstractIn this paper, an off-line recognition system based on multifeature and multilevel classification is presented for handwritten Chinese characters. Ten classes of multifeatures, such as peripheral shape features, stroke density features, and stroke direction features, are used in this system. The multilevel classification scheme consists of a group classifier and a five-level character classifier, where two new technologies, overlap clustering and Gaussian distribution selector are developed. Experiments have been conducted to recognize 5,401 daily-used Chinese characters. The recognition rate is about 90 percent for a unique candidate, and 98 percent for multichoice with 10 candidates. Yuan Yan Tang, Lo-Ting Tu, Jiming Liu 0001, Seong-Whan Lee, Win-Win Lin, Ing-Shyh Shyu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | Information Acquisition and Storage of Forms in Document ProcessingabstractAn automatic form information acquisition and storage system is presented. By semantic meaning analysis and registration of every king of forms in advance, the system as able to recognize incoming forms and extract information from them automatically. An efficient form image storage method is also proposed. Instead of keeping the entire form image, this system greatly reduces the memory needed to store the form images in a database by extracting and storing user filled data images only. Yuan Yan Tang, Jiming Liu 0001 |
ICDAR | 1 |
| 1997 | Quadratic Spline Wavelet Approach to Automatic Extraction of Baselines from Document ImagesabstractThe paper presents a wavelet based approach to edge detection in document processing. According to local analysis of the document images using wavelet theory, a novel method is developed to detect the edges in document processing, including extraction of the contours of characters and extraction of the reference lines in the form document images with gray levels. In this method, the quadratic spline wavelet is utilized. Experiments have been contacted. The positive results show the effectiveness of the application of the quadratic spline wavelet to edge detection, especially to extract the reference lines and image boundaries in document processing. Yuan Yan Tang, Jiming Liu 0001, Lihua Yang 0001 |
ICDAR | 1 |
| 1997 | Location and recognition of legal amounts on Chinese bank chequesabstractThis paper describes a Chinese cheque processing system currently under development at the Centre for Pattern Recognition and Machine Intelligence (CENPARMI). The information on Chinese bank cheques is not the same as that on alphanumeric bank cheques. The legal amount in a Chinese bank cheque is the Chinese character text associated with each currency unit. This paper discusses a technique using each currency unit as a key word to locate/extract the legal amount in bank cheques. In the analysis and recognition process, the system tries to locate the smallest currency units in the image and identifies it first. Then, the system tries to locate the image strings associated with each currency unit. Each image string is separated and recognized. Next, a set of rules and context are applied to recognize the characters. In order to choose the correct one, the recognized character string is accepted only if it satisfies all the conditions governed by rules. Chiu L. Yu, Ching Y. Suen, Yuan Yan Tang |
ICDAR | 3 |
| 1997 | On-line recognition handwritten mathematical symbolabstractThe paper aims at online recognition of handwritten mathematical symbols. It analyses the structures of 94 opt used mathematical symbols and concludes that all of them consist of 10 basic elements. It proposes a new method of basic element ordering and reduces the number of standard symbols by extracting three primary features of mathematical symbols, namely, basic element vector, relative positions between basic elements and basic element length vector. The traditional dynamic programming method is improved by means of classifying roughly 94 mathematical symbols, considering matching and unmatching value, adding geometric restraints and solving matching problem, through improved Kohn-Munkres algorithm. Correctness rate reaches 90.52%, incorrectness rate 5.03% and refusal rate 4.45%. Xuejun Zhao, Shengling Zheng, Baochang Pan, Yuan Yan Tang |
ICDAR | 5 |
| 1997 | Automatic Extraction of Baselines and Data from Check ImagesabstractA novel approach to extract data from check images is proposed based on the determination of baselines of checks, a priori information about the positions of data on checks, and a layout-driven item extraction method. Several techniques and algorithms have been developed in this approach including check image preprocessing, the extraction and identification of baselines, the extraction of the strokes of handwritten legal amounts, courtesy amounts and date, and the separation of strokes connected to baselines. A complete working system has been developed. The results of both testing experiments and on-line applications show that this approach is effective and the proposed techniques and algorithms perform well. Ke Liu 0009, Ching Y. Suen, Mohamed Cheriet, Joseph N. Said, Christine P. Nadal, Yuan Yan Tang |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 1997 | Multiresolution analysis in extraction of reference lines from documents with gray level backgroundabstractBased on wavelets, a theoretical method has been developed to process multi-gray level documents. In this method, two-dimensional multiresolution analysis, a wavelet decomposition algorithm, and compactly supported orthonormal wavelets are used to transform a document image into sub-images. According to these sub-images, the reference lines of a multi-gray level document can be extracted, and knowledge about the geometric structure of the document can be acquired. Particularly, this approach is more efficient to process form documents with gray level background. Experiments indicate that this new method can be applied to process documents with promising results. Yuan Yan Tang, Hong Ma 0001, Jiming Liu 0001, Bing F. Li, Dihua Xi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | Chinese document layout analysis based on adaptive split-and-merge and qualitative spatial reasoning
Jiming Liu 0001, Yuan Yan Tang, Ching Y. Suen |
Pattern Recognit. | 2 |
| 1997 | An evolutionary autonomous agents approach to image feature extractionabstractThis paper presents a new approach to image feature extraction which utilizes evolutionary autonomous agents. Image features are often mathematically defined in terms of the gray-level intensity at image pixels. The optimality of image feature extraction is to find all the feature pixels from the image. In the proposed approach, the autonomous agents, being distributed computational entities, operate directly in the 2-D lattice of a digital image and exhibit a number of reactive behaviors. To effectively locate the feature pixels, individual agents sense the local stimuli from their image environment by means of evaluating the gray-level intensity of locally connected pixels, and accordingly activate their behaviors. The behavioral repository of the agents consists of: 1) feature-marking at local pixels and self-reproduction of offspring agents in the neighboring regions if the local stimuli are found to satisfy feature conditions, 2) diffusion to adjacent image regions if the feature conditions are not held, or 3) death if the agents exceed their life span. As part of the behavior evolution, the directions in which the agents self-reproduce and/or diffuse are inherited from the directions of their selected high-fitness parents. Here the fitness of a parent agent is defined according to the steps that the agent takes to locate an image feature pixel. Jiming Liu 0001, Yuan Yan Tang, Y. C. Cao |
IEEE Trans. Evol. Comput. | 2 |
| 1997 | Modified Fractal Signature (MFS): A New Approach to Document Analysis for Automatic Knowledge AcquisitionabstractOne of the key technologies related to knowledge and data engineering is the acquisition of knowledge and data in the development and utilization of information system and the strategies to capture new knowledge and data. Actually, millions of documents, including technical reports, government files, newspapers, books, magazines, letters, bank checks, etc., have to be processed every day, and knowledge has to be acquired from them. This paper presents a new approach to document analysis for automatic knowledge acquisition. The traditional approaches have two major disadvantages: (1) They are not effective for processing documents with high geometrical complexity. Specially, the top-down approach can process only the simple documents which have specific format or contain some a priori information. (2) The top-down approach needs to split large components into small ones iteratively, while the bottom-up approach needs to merge small components into large ones iteratively. They are time consuming. This new approach is based on modified fractal signature. It can overcome the above weaknesses. Yuan Yan Tang, Hong Ma 0001, Dihua Xi, Xiaogang Mao, Ching Y. Suen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1996 | Adaptive document segmentation and geometric relation labeling: algorithms and experimental resultsabstractThis paper describes a generic document segmentation and geometric relation labeling method with applications to document analysis. Unlike the previous document segmentation methods where text spacing, border lines, and/or a priori layout models based template processing are performed, the present method begins with a hierarchy of partitioned image layers where inhomogeneous higher-level regions are recursively positioned into lower-level rectangular subregions and at the same time lower-level smaller homogeneous regions are merged into larger homogeneous regions. The present method differs from the traditional split-and-merge segmentation method in that it orthogonally splits regions using thresholds adaptively computed from projection profiles. Jiming Liu 0001, Yuan Yan Tang, Qichao He, Ching Y. Suen |
ICPR | 2 |
| 1996 | A novel approach to optical character recognition based on ring-projection-wavelet-fractal signaturesabstractIn this paper, we present a novel approach to optical character recognition that utilizes ring-projection-wavelet-fractal-signatures. In particular, the proposed approach reduces the dimensionality of a two-dimensional pattern by way of a ring-projection method, and thereafter, performs Daubechies' wavelet transform on the derived one-dimensional pattern to generate a set of wavelet sub-patterns, namely, curves that are non-self intersecting. Further from the resulting non-self intersecting curves, the divider dimensions are readily computed. These divider dimensions constitute a new characteristic vector for the original two-dimensional pattern, defined over the curves' fractal dimensions. Yuan Yan Tang, Bing F. Li, Hong Ma 0001, Jiming Liu 0001, Cheung Hoi Leung, Ching Y. Suen |
ICPR | 1 |
| 1996 | Multiresolution recognition of unconstrained handwritten numerals with wavelet transform and multilayer cluster neural network
Seong-Whan Lee, Hong Ma 0001, Yuan Yan Tang |
Pattern Recognit. | 4 |
| 1996 | Nonlinear shape restoration of distorted images with coons transformation
Seong-Whan Lee, Eun-Soon Kim, Yuan Yan Tang |
Pattern Recognit. | 3 |
| 1996 | Automatic document processing: A survey
Yuan Yan Tang, Seong-Whan Lee, Ching Y. Suen |
Pattern Recognit. | 1 |
| 1995 | Nonlinear shape restoration of distorted images with Coons transformationabstractImage shape restoration based on mathematical transformation is a successful approach to nonlinear distortions in computer vision, robot vision and pattern recognition. The key of this process is to find the distortion function and its inverse function. Usually, the distortion function is unknown or unclear. Even in the case where the function is known, it remains difficult to compute or estimate the parameters necessary for the restoration. To overcome this problem, Coons transformation utilizing boundary functions for the distorted images have been used to approximate the exact distortion function. The boundary functions are calculated using B-spline curve interpolation which is coincided with the necessary condition of major elements that constitute a Coons transformation. Seong-Whan Lee, Eun-Soon Kim, Yuan Yan Tang |
ICDAR | 3 |
| 1995 | A new approach to document analysis based on modified fractal signatureabstractThis paper presents a new approach to document analysis. The proposed approach is based on modified fractal signature. Instead of the time-consuming traditional approaches (top-down and bottom-up approaches) where iterative operations are necessary to break a document into blocks to extract its geometric (layout) structure, this new approach can divide a document into blocks in only one step. This approach can be used to process documents with high geometrical complexity. Experiments have been conducted to prove the proposed new approach for document processing. Yuan Yan Tang, Hong Ma 0001, Xiaogang Mao, Ching Y. Suen |
ICDAR | 1 |
| 1995 | Extraction of reference lines from documents with grey-level background using sub-images of waveletsabstractBased on wavelets, a new theoretical method has been developed to process form documents. In this method, two-dimensional multiresolution analysis (MSA), wavelet decomposition algorithm, and compactly supported orthonormal wavelets are used to transform a document image into sub-images. According to these sub-images, the reference lines of forms can be extracted, and knowledge about the geometric structure of the document can be acquired. Experiments prove that this new method can be applied to process documents with promising results. Yuan Yan Tang, Hong Ma 0001, Dihua Xi, Ching Y. Suen |
ICDAR | 1 |
| 1995 | Document skew detection based on the fractal and least squares methodabstractIn this paper, a simple and robust algorithm is presented to detect skew in a totally unconstrained document. It can discover the skew angle not only in the whole page of document but also in different document blocks which have their different skew angles. This method consists of four major phases, namely: (a) skew detection and correction for whole page; (b) segmentation of document into blocks; (c) identification of skewed text blocks, and (d) skew detection and correction for the skewed text blocks. To detect the skew in a document, the saw-tooth algorithm and least squares method are used. To segment a document into blocks, the fractal approach is applied. Promising experimental results are also provided to prove the effectiveness of the proposed method. Chiu L. Yu, Yuan Yan Tang, Ching Y. Suen |
ICDAR | 2 |
| 1995 | Four directional adjacency graphs (FDAG) and their application in locating fields in formsabstractA new non-hierarchical spatial data structure named four directional adjacency graphs (FDAG) is proposed. In the FDAG vertical and horizontal neighborhood relationship between rectangles is well represented so that structural information can be easily extracted. An application for structural analysis of forms is given, where experiments are conducted with positive results. Jianxing Yuan, Yuan Yan Tang, Ching Y. Suen |
ICDAR | 2 |
| 1995 | A structure-parameter-adaptive (SPA) neural tree for the recognition of large character set
Yuan Yan Tang, Luyuan Fang |
Pattern Recognit. | 2 |
| 1995 | Financial document processing based on staff line and description languageabstractMillions of financial transactions take place every day. Associated with them are documents such as bank cheques, payment slips and bills which have to be processed. A great deal of time, effort and money will be saved if they can be entered into the computer and processed automatically. According to the specific characteristics of financial documents, it can be concluded that it is possible to build a system for recognizing specific types of financial documents, instead of a complex and general one aiming at different kinds of documents. In this paper, a financial document recognition prototype system which can process bank cheques, payment slips and bills, is presented. It consists of four major parts: (a) document image acquisition including scanning and binarization, (b) fixed document processing subsystem based on the detection of staff lines, (c) flexible document processing subsystem operating in a form description language (FDL), and (d) character recognition. Numerous experimental results are presented and discussed.> Yuan Yan Tang, Ching Y. Suen, Chang De Yan, Mohamed Cheriet |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1994 | VLSI arrays for speech processing with linear predictive codingabstractThe covariance analysis of linear predictive coding has wide applications, especially in speech recognition and speech signal processing. Real-time applications demand very high processing speed for linear predictive coding analysis. VLSI technology which possesses properties of low-cost, high-speed and massive computing capabilities is a suitable candidate. In this paper, systolic array processors for the covariance analysis of linear predictive coding are developed. The covariance analysis of linear predictive coding contains a large set of irregular and nested recurrence equations. Systolizing the algorithm is a difficult task for such a complex problem, Existing methods of systematic design for systolic arrays are not much helpful to this problem. To overcome it, a break-combination method is presented in this paper. In this manner, the task is first decomposed and then mapped onto several interconnected systolic arrays. The resulting systolic arrays of the sub problems are then combined to form a complete solution. Yuan Yan Tang, Ching Y. Suen |
ICPR (3) | 1 |
| 1994 | Extraction of peripheral shape features in Chinese character recognitionabstractExtraction of a stable and representative set of features is the heart of the design of a pattern recognition system. Knowing the distribution of information on the pixels of a character will be of great assistance to the study of feature extraction. In this paper, an analysis of the distribution of information on the pixels of binarized Chinese characters is presented. From the analysis, it is obvious that the information of a Chinese character tends to concentrate around the peripheries of the character. Several methods to extract peripheral shape features are presented. Some experiments are conducted on Chinese character recognition and the results show the advantages of the peripheral shape features. Yuan Yan Tang, Ching Y. Suen |
ICPR (2) | 1 |
| 1994 | Document Structures: A SurveyabstractKnowing the structure of a document is the key to successful processing of a document. There exist a variety of definitions of document structures. This paper is a survey of methods describing document structures. Several novel concepts and theoretical analyses are also presented. A document not only has a concrete two-dimensional image but also a conceptual structure which corresponds to human’s thinking. The process of publishing or writing corresponds to encoding the conceptual structure into a concrete structure. Conversely, the concrete structure of the document is decoded into its conceptual one in document processing. In this paper, conceptual and concrete structures are introduced. A complete system for treating both of the conceptual and concrete structures is probably still decades away. As the first stage, this study puts some emphasis on concrete structures, for which, geometric, logical, textual, information, textural, and other structures are described. Yuan Yan Tang, Ching Y. Suen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1994 | New algorithms for fixed and elastic geometric transformation modelsabstractThis paper describes a new approach that leads to the discovery of substitutions or approximations for physical transformation by fixed and elastic geometric transformation models. These substitutions and approximations can simplify the solution of normalization and generation of shapes in signal processing, image processing, computer vision, computer graphics, and pattern recognition. In this paper, several new algorithms for fixed geometric transformation models such as bilinear, quadratic, bi-quadratic, cubic, and bi-cubic are presented based on the finite element theory. To tackle more general and more complicated problems, elastic geometric transformation models including Coons, harmonic, and general elastic models are discussed. Several useful algorithms are also presented in this paper. The performance of the proposed approach has been evaluated by a series of experiments with interesting results. Yuan Yan Tang, Ching Y. Suen |
IEEE Trans. Image Process. | 1 |
| 1994 | Document Processing for Automatic Knowledge AcquisitionabstractThe knowledge acquisition bottleneck has become the major impediment to the development and application of effective information systems. To remove this bottleneck, new document processing techniques must be introduced to automatically acquire knowledge from various types of documents. By presenting a survey on the techniques and problems involved, this paper aims at serving as a catalyst to stimulate research in automatic knowledge acquisition through document processing. In this study, a document is considered to have two structures: geometric structure and logical structure. These play a key role in the process of the knowledge acquisition, which can be viewed as a process of acquiring the above structures. Extracting the geometric structure from a document refers to document analysis; mapping the geometric structure into logical structure is regarded as document understanding. Both areas are described in this paper, and the basic concept of document structure and its measurement based on entropy analysis is introduced. Logical structure and geometric models are proposed. Both top-down and bottom-up approaches and their entropy analyses are presented. Different techniques are discussed with practical examples. Mapping methods, such as tree transformation, document formatting knowledge and document format description language, are described.> Yuan Yan Tang, Chang De Yan, Ching Y. Suen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1994 | RPCT Algorithm and its VLSI ImplementationabstractThis paper presents the regional projection contour transformation (RPCT) which transforms a compound pattern or multicontour pattern into a unique outer contour. Two RPCT's, (1) diagonal-diagonal regional projection contour transformation and (2) horizontal-vertical regional projection contour transformation, are presented. They are applicable to a wide range of areas such as image analysis, pattern recognition, etc. A very large scale integration (VLSI) architecture to implement the RPCT has also been designed based on a canonical methodology which maps homogeneous dependence graphs into processor arrays. In this paper, a linear array has been designed, where an N/2-element vector is used to process a pattern with a size of N/spl times/N. It can speed up the recognition process considerably with a time complexity of O(N) compared with O(N/sup 2/) when a uniprocessor is used.> Yuan Yan Tang, Ching Y. Suen |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 1993 | VLSI implementation for HVRI algorithm in pattern recognitionabstractA VLSI architecture to implement the horizontal- vertical region integration (HVRI) algorithm has been designed. The HVRI algorithm transforms a multi-contour pattern into a unique outer contour. It is applicable to a wide range of areas such as image analysis, pattern recognition, etc. A linear array has been designed based on a canonical methodology which maps homogeneous dependence graphs into processor arrays. An N/2-element vector is used to process a pattern with a size of N/spl times/N. It can speed up the recognition process considerably with a time complexity of O(N) compared with O(N/sup 2/) when a uniprocessor is used.> Yuan Yan Tang, Seong-Whan Lee |
ICDAR | 1 |