Wei Qi Yan 0001

dblp:150/1142 · also Wei-Qi Yan 0001, WeiQi Yan 0001, Weiqi Yan 0001 · DBLP profile ↗
← Back
111ranked-venue papers
13as first author
49since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 84 · 12 first-author · 36 since 2021Artificial intelligence and machine learning · 10 · 6 since 2021Security and privacy · 8 · 2 since 2021Computer networks · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 DSAN: a dual-scale spatial-temporal aggregation network for robust gait recognition in one-shot and cross-view scenarios
Liang Ren, Xiuhui Wang, Wei Qi Yan 0001
Appl. Intell.3
2026 Inductive multiple clustering based on weakly-supervised salient representation learning
Wei Qi Yan 0001
Expert Syst. Appl.2
2026 Multi-view co-clustering with dynamic feature-level clustering candidates
Wei Qi Yan 0001, Qinli Zhou
Inf. Process. Manag.2
2026 MrNEAD: Community detection attacks in social networks using modularity-regularized network embedding via adversarial decomposition
Wei Qi Yan 0001
Inf. Sci.3
2026 Balanced clustering-regularized bilateral spectral learning for unsupervised feature selection in internet of things environments
Yingjie Dong, Wei Qi Yan 0001
Inf. Sci.3
2026 Trademark detection and classification from digital images with self-attention mechanism
Chunyuan Miao, Xiuhui Wang, Wei Qi Yan 0001
Multim. Tools Appl.3
2026 Deep Inductive and Scalable Subspace Clustering via Nonlocal Contrastive Self-Distillation
abstract
Deep subspace clustering has demonstrated remarkable results by leveraging the nonlinear subspace assumption. However, it often encounters challenges in terms of computational cost and memory footprint in dealing with large-scale data due to its traditional single-batch training strategy. To address this issue, this paper proposes a deep subspace clustering framework that is regularized by nonlocal contrastive self-distillation, enabling a Deep Inductive and Scalable Subspace Clustering (DISSC) algorithm. In particular, our framework incorporates two subspace learning modules, namely subspace learning based on self-expression model and inductive subspace clustering. These modules generate affinities from different perspectives by extracting intermediate features from two augmentations of the input data using a weight-sharing neural network. By integrating the concept of self-distillation, our framework effectively exploits the clustering-friendly knowledge contained in these two affinities through a novel nonlocal contrastive prediction task, employing an empirical yet effective threshold. This allows the framework to facilitate complementary knowledge mining and scalability without compromising clustering performance. With an alternate branch that bypasses the self-expression computation, our framework can infer subspace membership of the out-of-sample data through the predicted soft labels, eliminating the need for ad-hoc postprocessing. In addition, the self-expression matrix computed using mini-batch data benefits from the distilled knowledge obtained from the inductive subspace clustering module, enabling our framework to scale to data of arbitrary size. Experiments conducted on large-scale MNIST, Fashion-MINST, STL-10, CIFAR-10 and Stanford Online Products datasets validate the superiority of the proposed DISSC algorithm over state-of-the-art subspace clustering methods.
Bo Peng 0028, Wei Qi Yan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Optimal Access Structure Partition Methods for Image Secret Sharing
abstract
Visual cryptography scheme (VCS) and polynomial-based secret image sharing (PSIS) are two primary types of secret sharing for protecting images. VCS and PSIS have their respective pros and cons. For VCS, the benefits of perfect security and easy decoding are provided. But it suffers from the limitations of lossy secret recovery and binary image-oriented. PSIS can deal with grayscale/color images and offers lossless secret reconstruction. Whereas, the secret decoding is computationally intensive (i.e.,O(klog2k) for (k,n) threshold) and the residual-image problem in PSIS compromises the security. In this paper, we are motivated to investigate a sharing technique that can preserve the advantages of both VCS and PSIS. Differing from existing VCS and PSIS, the proposed sharing method is accomplished based on the access structure partition (ASP) result. Essentially, an ASP guided image secret sharing approach is developed and three optimal ASP algorithms are designed. When compared with existing partition method, significant improvement is offered by our partition techniques especially for the (k,n) threshold with a largern. Take the (2; 15), (2; 18), and (4; 12) thresholds for example, the numbers of involved sub-access structures by our method are 4, 5, and 19, while the quantities by existing approach are 8, 10, and 45. The percentages of improvement are 100%, 100%, and 137%. Further, based on the partition result from ASP algorithms, we can employ (k,k) probabilistic VCS (PVCS) to constitute a (k,n) sharing method for encoding gray-level/color images. Experiments are demonstrated to confirm the effectiveness of the sharing method and ASP algorithms. Meanwhile, comparisons are included to show that the merits of perfect security, low decoding complexity (i.e.,O(d)), lossless secret recovery (i.e., PSNR= ∞, SSIM= 1), and grayscale/color image-oriented are provided by our sharing method.
Zhihua Xia, Ching-Nung Yang, Wei Qi Yan 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Multi-Level Structural Contrastive Subspace Clustering Network
abstract
Deep subspace clustering methods based on autoencoder (AE) have achieved impressive performance in various applications. However, these methods often place excessive reliance on the AE framework, which focuses primarily on pixel-level reconstruction while overlooking the structural information inherent in the data. To overcome this limitation, we propose a novel approach called the Multi-level Structural Contrastive Subspace Clustering Network (MSCSCN). Unlike traditional AE-based methods, MSCSCN departs from the AE paradigm and introduces multi-level contrastive prediction to improve feature learning. Specifically, MSCSCN integrates multi-level features from both original and augmented data within a self-expression learning process, enhancing the learned pairwise affinities. Additionally, we propose a structural contrastive loss, which strengthens cluster boundary discrimination by effectively utilizing pairwise affinities and structural information. Our experimental results on several benchmark datasets demonstrate that MSCSCN outperforms competitive deep subspace clustering methods, highlighting its superior capability in improving clustering performance and capturing the underlying structural information within the data.
Wei Qi Yan 0001
IEEE Signal Process. Lett.3
2025 CRP2-VCS: Contrast-Oriented Region-Based Progressive Probabilistic Visual Cryptography Schemes
abstract
Most visual cryptography schemes (VCSs) are condition-oriented which implies their designs focus on satisfying the contrast and security conditions in VCS. In this paper, we explore a new architecture of VCS: contrast-oriented region-based progressive probabilistic VCS (CRP2-VCS). The term contrast-oriented indicates the optimality of multi-contrast is taken into consideration when producing shadows. First of all, new requirements for CRP2-VCS, described by probabilities, are introduced. As a non-interference requirement is proposed, the secret interference problem in existing region-based progressive VCS can be avoided. Then, a construction of CRP2-VCS based on a multi-contrast-maximizing model is provided. The multi-contrast-maximizing problem is essentially a probabilistic VCS model that fuses region-based sharing, multi-contrast optimization, and general access structure (GAS) together. Finally, a Max-Min based technique is adopted to solve the multi-objective optimization problem. Moreover, to further boost the visual quality, the proposed method is extended to allow employing XOR operation for image recovery. Experimental results and comparisons are demonstrated to show the effectiveness and advantages, such as optimal visual quality, non-expansible shadow and GAS sharing policy, are provided by the proposed technique.
Bofan Song, Jia Fang, Wei Qi Yan 0001, Qing-Yu Peng
IEEE Trans. Circuits Syst. Video Technol.4
2025 On the Design of Distributed Multi-User Secret Image Sharing for General Access Structures
abstract
In this article, a secret image sharing (SIS) scheme for general access structures (GAS) is designed for distributed multi-user scenario. In the proposed distributed multi-user SIS (DM-SIS), multiple secret images are encoded into shadows which are then distributed to the corresponding storage nodes of a network. By collecting the shadows from nodes, each user is capable of decrypting the corresponding secret image. Fundamentally, we utilize an invertible target matrix, which is initially obtained from the GAS and a base matrix with Vandermonte coordinates, to construct shadows. To deal with the case of non-invertible target matrix, three matrix-adjusting procedures are further introduced. Theoretical analysis, numerical examples, and experiments are provided to verify the feasibility of the proposed technique. When compared to previous methods, the proposed approach can implement GAS sharing strategy in distributed multi-user environment. Meantime, significant improvements on storage overhead and sharing capacity are also achieved.
Yuyang Xiong, Bing Chen 0004, Ching-Nung Yang, Wei Qi Yan 0001, Qing-Yu Peng
IEEE Trans. Dependable Secur. Comput.5
2025 EVCS-DAS: Evolving Visual Cryptography Schemes for Dynamic Access Structures
abstract
A systematic investigation of evolving visual cryptography scheme (EVCS) is carried out in this article. The evolving scheme, denoted as \((k,\infty)\) , differs from the \((k,n)\) threshold in that it permits an arbitrary and perhaps unlimited number of participants. More importantly, the access structure can be updated dynamically by adding new users. First of all, a preliminary implementation strategy for the \((2,\infty)\) EVCS is introduced. Then, by employing the \((2,2)\) VCS recursively with the \((2,\infty)\) EVCS, a \((k,\infty)\) EVCS is created. In order to enhance the performance, an improved scheme is constructed based on the multi-secret VCS (MVCS) and a series of EVCS schemes with thresholds of \((1,\infty)\) , \(\cdots\) , \((k-1,\infty)\) . Moreover, Boolean XOR operation is adopted for secret recovery to further improve the visual quality. To facilitate the XOR decryption, a novel access structure partition algorithm is presented. Additionally, the proposed partition method can successfully solve the security issue in existing multi-secret XOR-based VCS (MXVCS). By integrating the more secure MXVCS into the improved scheme, XOR decryption is provided. The two proposed methods are shown to be effective and advantageous through extensive experiments and comparisons.
Xinjie Feng, Bing Chen 0004, Ching-Nung Yang, Qing-Yu Peng, Wei Qi Yan 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2025 TFFD-Net: an effective two-stage mixed feature fusion and detail recovery dehazing network
Wei Qi Yan 0001, Shihua Zhou, Yueping Wang
Vis. Comput.2
2024 Pose estimation for swimmers in video surveillance
abstract
Abstract Traditional models for pose estimation in video surveillance are based on graph structures, in this paper, we propose a method that breaks the limitation of template matching within a range of pose changes to obtain robust results. We implement our swimmer pose estimation method based on deep learning. We take use of High-Resolution Net (HRNet) to extract and fuse visual features of visual object and complete the object detection using the key points of human joint. The proposed model could be applied to all kinds of swimming styles throughout appropriate training. Compared with the methods that require multimodel combinations and training, the proposed method directly achieves the end-to-end prediction, which is easily to be implemented and deployed. In addition, a cross-fusion module is added between parallel networks, which assists the network to make use of the characteristics of multiple resolutions. The proposed network has achieved ideal results in the pose estimation of swimmers by comparing HRNet-W32 and HRNet-W48. In addition, we propose an annotated key point dataset of swimmers which was created from the view of underwater swimmers. Compared with side view, the torso of swimmers collected by the underwater view is much suitable for a broad spectrum of machine vision tasks.
Wei Qi Yan 0001
Multim. Tools Appl.2
2024 Moving vehicle tracking and scene understanding: A hybrid approach
Wei Qi Yan 0001, Nikola K. Kasabov
Multim. Tools Appl.2
2024 A privacy-preserving word embedding text classification model based on privacy boundary constructed by deep belief network
abstract
Abstract To effectively extract and classify the information from reports or documents and protect the privacy of the extracted results, we propose a privacy classification named Word Embedding Combination Privacy-preserving Support Vector Machine (WECPPSVM) model to classify the text. In addition, this paper also proposes the Privacy-preserving Distribution and Independent Frequent Subsequence Extraction Algorithm (PPDIFSEA), which calculates the degree of independence of the training data input to the classification model by training the Deep Belief Network(DBN) in PPDIFSEA, then obtains the Privacy Boundary(PB). PB is an indispensable condition for both data sampling and privacy noise generation. And this model can protect privacy by injecting the privacy noise into the classification result, this method can interfere with the background knowledge-based privacy attack. Our quantitative analysis shows that the WECPPSVM proposed in this paper can approach mainstream text classification algorithms in terms of text classification accuracy while preserving privacy without increasing computational complexity. In addition, the fusion study and privacy threat evaluation also verify that the proposed PPDIFSEA method combined with WECPPSVM achieves an acceptable level of classification accuracy and privacy protection.
Bo Ma 0008, Edmund M.-K. Lai, Wei Qi Yan 0001, Jinsong Wu 0001
Multim. Tools Appl.3
2024 CISO: Co-iteration semi-supervised learning for visual object detection
abstract
Abstract Semi-supervised learning offers a solution to the high cost and limited availability of manually labeled samples in supervised learning. In semi-supervised visual object detection, the use of unlabeled data can significantly enhance the performance of deep learning models. In this paper, we introduce an end-to-end framework, named CISO (Co-Iteration Semi-Supervised Learning for Object Detection), which integrates a knowledge distillation approach and a collaborative, iterative semi-supervised learning strategy. To maximize the utilization of pseudo-label data and address the scarcity of pseudo-label data due to high threshold settings, we propose a mean iteration approach where all unlabeled data is applied to each training iteration. Pseudo-label data with high confidence is extracted based on an ever-changing threshold (average intersection over union of all pseudo-labeled data). This strategy not only ensures the accuracy of the pseudo-label but also optimizes the use of unlabeled data. Subsequently, we apply a weak-strong data augmentation strategy to update the model. Lastly, we evaluate CISO using Swin Transformer model and conduct comprehensive experiments on MS-COCO. Our framework showcases impressive results, outperforms the state-of-the-art methods by 2.16 mAP and 1.54 mAP with 10% and 5% labeled data, respectively.
Jianchun Qi, Minh Nguyen 0001, Wei Qi Yan 0001
Multim. Tools Appl.3
2024 NUNI-Waste: novel semi-supervised semantic segmentation waste classification with non-uniform data augmentation
abstract
Abstract Waste categorization and recycling are critical approaches for converting waste into valuable and functional materials, thereby significantly aiding in land preservation, reducing pollution, and optimizing resource usages. However, real-world classification and identification of recyclable waste face substantial hurdles due to the intricate and unpredictable nature of wastes, as well as the limited availability of comprehensive waste datasets. These factors limit efficacy of the existing research work in the domain of waste management. In this paper, we utilize semantic segmentation at individual pixel level and introduce a semi-supervised metod for authentic waste classification scenarios, leveraging the Zerowaste dataset. We devise a non-standard data augmentation strategy that mimics the ever-changing conditions of real-world waste environments. Additionally, we introduce an adaptive weighted loss function and dynamically adjust the ratio of positive to negative samples through a masking method, ensuring the model learns from relevant samples. Lastly, to maintain consistency between predictions made on data-augmented images and the original counterparts, we remove input perturbations. Our method proves to be effective, as verified by an array of standard experiments and ablation studies, achieved an accuracy improvement of 3.74% over the baseline Zerowaste method.
Jianchun Qi, Minh Nguyen 0001, Wei Qi Yan 0001
Multim. Tools Appl.3
2024 Apple ripeness identification from digital images using transformers
abstract
Abstract We describe a non-destructive test of apple ripeness using digital images of multiple types of apples. In this paper, fruit images are treated as data samples, artificial intelligence models are employed to implement the classification of fruits and the identification of maturity levels. In order to obtain the ripeness classifications of fruits, we make use of deep learning models to conduct our experiments; we evaluate the test results of our proposed models. In order to ensure the accuracy of our experimental results, we created our own dataset, and obtained the best accuracy of fruit classification by comparing Transformer model and YOLO model in deep learning, thereby attaining the best accuracy of fruit maturity recognition. At the same time, we also combined YOLO model with attention module and gave the fast object detection by using the improved YOLO model.
Bingjie Xiao, Minh Nguyen 0001, Wei Qi Yan 0001
Multim. Tools Appl.3
2024 Fruit ripeness identification using YOLOv8 model
abstract
Abstract Deep learning-based visual object detection is a fundamental aspect of computer vision. These models not only locate and classify multiple objects within an image, but they also identify bounding boxes. The focus of this paper's research work is to classify fruits as ripe or overripe using digital images. Our proposed model extracts visual features from fruit images and analyzes fruit peel characteristics to predict the fruit's class. We utilize our own datasets to train two "anchor-free" models: YOLOv8 and CenterNet, aiming to produce accurate predictions. The CenterNet network primarily incorporates ResNet-50 and employs the deconvolution module DeConv for feature map upsampling. The final three branches of convolutional neural networks are applied to predict the heatmap. The YOLOv8 model leverages CSP and C2f modules for lightweight processing. After analyzing and comparing the two models, we found that the C2f module of the YOLOv8 model significantly enhances classification results, achieving an impressive accuracy rate of 99.5%.
Bingjie Xiao, Minh Nguyen 0001, Wei Qi Yan 0001
Multim. Tools Appl.3
2024 Dual Knowledge Distillation on Multiview Pseudo Labels for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) has made significant progress by leveraging valuable pseudo labels from completely unlabeled data. However, the predominant use of pseudo labels heavily relies on clustering results, which may lead to the accumulation of supervision deviation due to inevitable noise. In this paper, we propose a novel framework, namely Dual Knowledge Distillation on Multiview Pseudo Labels (DKD-MPL), to address this challenge. Specifically, the proposed DKD-MPL framework consists of two modules: Global Knowledge Distillation (GKD) and Self-Knowledge Distillation (SKD). In the GKD module, the pseudo labels obtained from the epoch-wise clustering procedure serve as the logits for the teacher model, while the mini-batch query images' pseudo labels act as the logits for the student model. Within the SKD module, we facilitate self-knowledge distillation by considering the pseudo labels generated by positive anchors and query images as two augmentations of the mini-batch data. As a result, DKD-MPL facilitates the exploitation of both global and local complementary knowledge across different views of pseudo labels, thereby mitigating supervision deviation. To demonstrate the effectiveness of DKD-MPL, we provide a theoretical analysis of the proposed loss and conduct extensive experiments on four popular datasets, e.g., Market-1501, DukeMTMC-reID, MSMT17, and VeRi-776. The results indicate that our method surpasses unsupervised approaches and achieves comparable performance to supervised person Re-ID methods.
Bo Peng 0028, Wei Qi Yan 0001
IEEE Trans. Multim.3
2023 A High-Accuracy Deformable Model for Human Face Mask Detection
Xinyi Gao 0002, Minh Nguyen 0001, Wei Qi Yan 0001
PSIVT3
2023 Enhancement of Human Face Mask Detection Performance by Using Ensemble Learning Models
Xinyi Gao 0002, Minh Nguyen 0001, Wei Qi Yan 0001
PSIVT3
2023 Multiscale Kiwifruit Detection from Digital Images
Minh Nguyen 0001, Raymond Lutui, Wei Qi Yan 0001
PSIVT4
2023 Computational Analysis of Table Tennis Matches from Real-Time Videos Using Deep Learning
Minh Nguyen 0001, Wei Qi Yan 0001
PSIVT3
2023 Human identification based on Gait Manifold
Xiuhui Wang, Wei Qi Yan 0001
Appl. Intell.2
2023 Fruit ripeness identification using transformers
abstract
Abstract Pattern classification has always been essential in computer vision. Transformer paradigm having attention mechanism with global receptive field in computer vision improves the efficiency and effectiveness of visual object detection and recognition. The primary purpose of this article is to achieve the accurate ripeness classification of various types of fruits. We create fruit datasets to train, test, and evaluate multiple Transformer models. Transformers are fundamentally composed of encoding and decoding procedures. The encoder is to stack the blocks, like convolutional neural networks (CNN or ConvNet). Vision Transformer (ViT), Swin Transformer, and multilayer perceptron (MLP) are considered in this paper. We examine the advantages of these three models for accurately analyzing fruit ripeness. We find that Swin Transformer achieves more significant outcomes than ViT Transformer for both pears and apples from our dataset.
Bingjie Xiao, Minh Nguyen 0001, Wei Qi Yan 0001
Appl. Intell.3
2023 A deep ensemble learning method for colorectal polyp classification with optimized network parameters
abstract
Abstract Colorectal Cancer (CRC), a leading cause of cancer-related deaths, can be abated by timely polypectomy. Computer-aided classification of polyps helps endoscopists to resect timely without submitting the sample for histology. Deep learning-based algorithms are promoted for computer-aided colorectal polyp classification. However, the existing methods do not accommodate any information on hyperparametric settings essential for model optimisation. Furthermore, unlike the polyp types, i.e., hyperplastic and adenomatous, the third type, serrated adenoma, is difficult to classify due to its hybrid nature. Moreover, automated assessment of polyps is a challenging task due to the similarities in their patterns; therefore, the strength of individual weak learners is combined to form a weighted ensemble model for an accurate classification model by establishing the optimised hyperparameters. In contrast to existing studies on binary classification, multiclass classification require evaluation through advanced measures. This study compared six existing Convolutional Neural Networks in addition to transfer learning and opted for optimum performing architecture only for ensemble models. The performance evaluation on UCI and PICCOLO dataset of the proposed method in terms of accuracy (96.3%, 81.2%), precision (95.5%, 82.4%), recall (97.2%, 81.1%), F1-score (96.3%, 81.3%) and model reliability using Cohen’s Kappa Coefficient (0.94, 0.62) shows the superiority over existing models. The outcomes of experiments by other studies on the same dataset yielded 82.5% accuracy with 72.7% recall by SVM and 85.9% accuracy with 87.6% recall by other deep learning methods. The proposed method demonstrates that a weighted ensemble of optimised networks along with data augmentation significantly boosts the performance of deep learning-based CAD.
Farah Younas, Muhammad Usman 0005, Wei Qi Yan 0001
Appl. Intell.3
2023 Sign language recognition from digital videos using feature pyramid network with detection transformer
abstract
Abstract Sign language recognition is one of the fundamental ways to assist deaf people to communicate with others. An accurate vision-based sign language recognition system using deep learning is a fundamental goal for many researchers. Deep convolutional neural networks have been extensively considered in the last few years, and a slew of architectures have been proposed. Recently, Vision Transformer and other Transformers have shown apparent advantages in object recognition compared to traditional computer vision models such as Faster R-CNN, YOLO, SSD, and other deep learning models. In this paper, we propose a Vision Transformer-based sign language recognition method called DETR (Detection Transformer), aiming to improve the current state-of-the-art sign language recognition accuracy. The DETR method proposed in this paper is able to recognize sign language from digital videos with a high accuracy using a new deep learning model ResNet152 + FPN (i.e., Feature Pyramid Network), which is based on Detection Transformer. Our experiments show that the method has excellent potential for improving sign language recognition accuracy. For instance, our newly proposed net ResNet152 + FPN is able to enhance the detection accuracy up to 1.70% on the test dataset of sign language compared to the standard Detection Transformer models. Besides, an overall accuracy 96.45% was attained by using the proposed method.
Parma Nand, Md. Akbar Hossain, Minh Nguyen 0001, Wei Qi Yan 0001
Multim. Tools Appl.5
2023 An ensemble framework of deep neural networks for colorectal polyp classification
Farah Younas, Muhammad Usman 0005, Wei Qi Yan 0001
Multim. Tools Appl.3
2023 Sharing Visual Secrets Among Multiple Groups With Enhanced Performance
abstract
An almost unexplored research topic in visual secret sharing (VSS), i.e., sharing different visual secrets among various groups, is investigated in this article. Generally, multiple groups are cooperated together to encrypt multiple visual secrets, where every group is correlated with a VSS scheme to share one secret and each participant receives only one shadow. Such a cooperative system is referred to as cooperative VSS (CVSS). In this paper, we develop a general construction for CVSS with improved computational overhead. The basis matrices for generating the shadows are easily obtained. To improve the recovered image quality, an efficient XOR-based CVSS (XCVSS) is given. Furthermore, to fulfill the requirement of implementing complex sharing policy, the XCVSS is extended to collaborate general access structure (GAS) VSS schemes together. Theoretical and experimental results and comparisons are illustrated to demonstrate the advantages of the proposed techniques.
Zishuo Xu, Wei Qi Yan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Contrast Optimization for Size Invariant Visual Cryptography Scheme
abstract
Visual cryptography scheme (VCS) serves as an effective tool in image security. Size-invariant VCS (SI-VCS) can solve the pixel expansion problem in traditional VCS. On the other hand, it is anticipated that the contrast of the recovered image in SI-VCS should be as high as possible. The investigation of contrast optimization for SI-VCS is carried out in this article. We develop an approach to optimize the contrast by stacking t ( k ≤ t ≤ n ) shadows in (k, n) -SI-VCS. Generally, a contrast-maximizing problem is linked with a (k, n) -SI-VCS, where the contrast by t shadows is considered as an objective function. An ideal contrast by t shadows can be produced by addressing this problem using linear programming. However, there exist (n-k+1) different contrasts in a (k, n) scheme. An optimization-based design is further introduced to provide multiple optimal contrasts. These (n-k+1) different contrasts are regarded as objective functions and it is transformed into a multi-contrast-maximizing problem. The ideal point method and lexicographic method are adopted to address this problem. Additionally, if the Boolean XOR operation is used for secret recovery, a technique is also provided to offer multiple maximum contrasts. The effectiveness of the proposed schemes is verified by extensive experiments. Comparisons illustrate significant advancement on contrast is provided.
Jia Fang, Wei Qi Yan 0001
IEEE Trans. Image Process.3
2022 A Method for Face Image Inpainting Based on Autoencoder and Generative Adversarial Network
Xinyi Gao 0002, Minh Nguyen 0001, Wei Qi Yan 0001
PSIVT3
2022 Depth Estimation of Traffic Scenes from Image Sequence Using Deep Learning
Wei Qi Yan 0001
PSIVT2
2022 Waste Classification from Digital Images Using ConvNeXt
Jianchun Qi, Minh Nguyen 0001, Wei Qi Yan 0001
PSIVT3
2022 Traffic Sign Recognition from Digital Images by Using Deep Learning
Jiawei Xing, Ziyuan Luo, Minh Nguyen 0001, Wei Qi Yan 0001
PSIVT4
2022 A hybrid CTC+Attention model based on end-to-end framework for multilingual speech recognition
abstract
Speech recognition is an important field in natural language processing. In this paper, the end-to-end framework for speech recognition with multilingual datasets is proposed. The end-to-end methods do not require complicated alignment and construction of the pronunciation dictionary, which show a promising prospect. In this paper, we implement a hybrid model of CTC and attention (CTC+Attention) model based on PyTorch. In order to compare speech recognition methods for multiple languages, we design and create three datasets: Chinese, English, and Code-Switch. We evaluate the proposed hybrid CTC+Attention model in multilingual environment. Throughout our experiments, we find that the proposed hybrid CTC+Attention model based on end-to-end framework achieves better performance compared with the HMM-DNN model in a single language and Code-Switch speaking environment. Moreover, the results of speech recognition with regard to different languages are compared in this paper. The CER(i.e., Character Error Rate) of the proposed hybrid CTC+Attention model based on the Chinese dataset defeated the traditional model and reached 10.22%.
Sendong Liang, Wei Qi Yan 0001
Multim. Tools Appl.2
2022 Flexible neural network for fast and accurate road scene perception
abstract
Abstract Accurate object detection on the road is the most important requirement of autonomous vehicles. Extensive work has been accomplished for car, pedestrian, and cyclist detection; however, comparatively, very few efforts have been put into 2D object detection. In this article, a dynamic approach is investigated to design a perfect unified neural network that could achieve the best results based on our available hardware. The proposed architecture is based on CSPNet for feature extraction in an end-to-end way. The net extracts visual features by using backbone subnet, visual object detection is based on a feature pyramid network (FPN). In order to increase the net flexibility, an auto-anchor generating method is applied to the detection layer that makes the net suitable for any datasets. For fine-tuning the net, activation, optimization, and loss functions are considered along with multiple check points. The proposed net is trained and tested based on the benchmark KITTI dataset. Our extensive experiments show that the proposed model for visual object detection is superior to others, where other nets output very low accuracy for pedestrian and cyclist detection, our proposed model achieves 99.3% recall rate based on our dataset.
Sabeeha Mehtab, Wei Qi Yan 0001
Multim. Tools Appl.2
2022 Colorizing Grayscale CT images of human lungs using deep learning methods
abstract
Image colorization refers to computer-aided rendering technology which transfers colors from a reference color image to grayscale images or video frames. Deep learning elevated notably in the field of image colorization in the past years. In this paper, we formulate image colorization methods relying on exemplar colorization and automatic colorization, respectively. For hybrid colorization, we select appropriate reference images to colorize the grayscale CT images. The colours of meat resemble those of human lungs, so the images of fresh pork, lamb, beef, and even rotten meat are collected as our dataset for model training. Three sets of training data consisting of meat images are analysed to extract the pixelar features for colorizing lung CT images by using an automatic approach. Pertaining to the results, we consider numerous methods (i.e., loss functions, visual analysis, PSNR, and SSIM) to evaluate the proposed deep learning models. Moreover, compared with other methods of colorizing lung CT images, the results of rendering the images by using deep learning methods are significantly genuine and promising. The metrics for measuring image similarity such as SSIM and PSNR have satisfactory performance, up to 0.55 and 28.0, respectively. Additionally, the methods may provide novel ideas for rendering grayscale X-ray images in airports, ferries, and railway stations.
Yuewei Wang, Wei Qi Yan 0001
Multim. Tools Appl.2
2022 Traffic sign recognition based on deep learning
abstract
Abstract Intelligent Transportation System (ITS), including unmanned vehicles, has been gradually matured despite on road. How to eliminate the interference due to various environmental factors, carry out accurate and efficient traffic sign detection and recognition, is a key technical problem. However, traditional visual object recognition mainly relies on visual feature extraction, e.g., color and edge, which has limitations. Convolutional neural network (CNN) was designed for visual object recognition based on deep learning, which has successfully overcome the shortcomings of conventional object recognition. In this paper, we implement an experiment to evaluate the performance of the latest version of YOLOv5 based on our dataset for Traffic Sign Recognition (TSR), which unfolds how the model for visual object recognition in deep learning is suitable for TSR through a comprehensive comparison with SSD (i.e., single shot multibox detector) as the objective of this paper. The experiments in this project utilize our own dataset. Pertaining to the experimental results, YOLOv5 achieves 97.70% in terms of [email protected] for all classes, SSD obtains 90.14% mAP in the same term. Meanwhile, regarding recognition speed, YOLOv5 also outperforms SSD.
Yanzhao Zhu, Wei Qi Yan 0001
Multim. Tools Appl.2
2021 An Adaptive Ant Colony Algorithm for Autonomous Vehicles Global Path Planning
abstract
In order to improve the robustness of the autonomous vehicle path planning algorithm and reduce the number of turns in the planned path, this paper proposes an adaptive ant colony algorithm path planning method. The algorithm optimizes the initial pheromone matrix based on the environment map, reduces the blindness of the initial ant colony in pathfinding, and improves the convergence speed. Then an adaptive heuristic function is used, which adaptively adjusts according to the different proportions of the heuristic function in the algorithm process, so as to avoid the algorithm being trapped in local optimum. The pheromone is updated according to the corners of the planned route, reducing the acute angle of the route and unnecessary turns to further optimize the route. The simulation results show that the proposed algorithm achieves good results. The simulation results show that the improved adaptive ant colony algorithm has faster convergence speed, higher path planning quality, and improved stability of planned paths than classical ant colony algorithms and other adaptive ant colony algorithms.
Yanqiang Li, Wei Qi Yan 0001
CSCWD4
2021 Traffic-light sign recognition using capsule network
Wei Qi Yan 0001
Multim. Tools Appl.2
2021 Banknote serial number recognition using deep learning
Wei Qi Yan 0001
Multim. Tools Appl.2
2021 Non-local gait feature extraction and human identification
Xiuhui Wang, Wei Qi Yan 0001
Multim. Tools Appl.2
2021 Fast-moving coin recognition using deep learning
Yufeng Xiang, Wei Qi Yan 0001
Multim. Tools Appl.2
2021 Mitigating severe over-parameterization in deep convolutional neural networks through forced feature abstraction and compression with an entropy-based heuristic
Nidhi Gowdra, Roopak Sinha, Stephen G. MacDonell, Wei Qi Yan 0001
Pattern Recognit.4
2021 Human Gait Recognition Based on Self-Adaptive Hidden Markov Model
abstract
Human gait recognition has numerous challenges due to view angle changing, human dressing, bag carrying, and pedestrian walking speed, etc. In order to increase gait recognition accuracy under these circumstances, in this paper we propose a method for gait recognition based on a self-adaptive hidden Markov model (SAHMM). First, we present a feature extraction algorithm based on local gait energy image (LGEI) and construct an observation vector set. By using this set, we optimize parameters of the SAHMM-based method for gait recognition. Finally, the proposed method is evaluated extensively based on the CASIA Dataset B for gait recognition under various conditions such as cross view, human dressing, or bag carrying, etc. Furthermore, the generalization ability of this method is verified based on the OU-ISIR Large Population Dataset. Both experimental results show that the proposed method exhibits superior performance in comparison with those existing methods.
Xiuhui Wang, Shiling Feng, Wei Qi Yan 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 Salient Object Detection Based on Visual Perceptual Saturation and Two-Stream Hybrid Networks
abstract
Inspired by the perceived saturation of human visual system, this paper proposes a two-stream hybrid networks to simulate binocular vision for salient object detection (SOD). Each stream in our system consists of unsupervised and supervised methods to form a two-branch module, so as to model the interaction between human intuition and memory. The two-branch module parallel processes visual information with bottom-up and top-down SODs, and output two initial saliency maps. Then a polyharmonic neural network with random-weight (PNNRW) is utilized to fuse two-branch's perception and refine the salient objects by learning online via multi-source cues. Depend on visual perceptual saturation, we can select optimal parameter of superpixel for unsupervised branch, locate sampling regions for PNNRW, and construct a positive feedback loop to facilitate perception saturated after the perception fusion. By comparing the binary outputs of the two-stream, the pixel annotation of predicted object with high saturation degree could be taken as new training samples. The presented method constitutes a semi-supervised learning framework actually. Supervised branches only need to be pre-trained initial, the system can collect the training samples with high confidence level and then train new models by itself. Extensive experiments show that the new framework can improve performance of the existing SOD methods, that exceeds the state-of-the-art methods in six popular benchmarks.
Wei Qi Yan 0001, Feilong Cao, Yongxia Zhou
IEEE Trans. Image Process.3
2021 Multitarget Tracking Using Siamese Neural Networks
abstract
In this article, we detect and track visual objects by using Siamese network or twin neural network. The Siamese network is constructed to classify moving objects based on the associations of object detection network and object tracking network, which are thought of as the two branches of the twin neural network. The proposed tracking method was designed for single-target tracking, which implements multitarget tracking by using deep neural networks and object detection. The contributions of this article are stated as follows. First, we implement the proposed method for visual object tracking based on multiclass classification using deep neural networks. Then, we attain multitarget tracking by combining the object detection network and the single-target tracking network. Next, we uplift the tracking performance by fusing the outcomes of the object detection network and object tracking network. Finally, we speculate on the object occlusion problem based on IoU and similarity score, which effectively diminish the influence of this issue in multitarget tracking.
Wei Qi Yan 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2020 Red-Green-Blue Augmented Reality Tags for Retail Stores
Minh Nguyen 0001, Wei Qi Yan 0001
ACIVS3
2020 Human Gait Recognition Based on Frame-by-Frame Gait Energy Images and Convolutional Long Short-Term Memory
abstract
Human gait recognition is one of the most promising biometric technologies, especially for unobtrusive video surveillance and human identification from a distance. Aiming at improving recognition rate, in this paper we study gait recognition using deep learning and propose a novel method based on convolutional Long Short-Term Memory (Conv-LSTM). First, we present a variation of Gait Energy Images, i.e. frame-by-frame GEI (ff-GEI), to expand the volume of available Gait Energy Images (GEI) data and relax the constraints of gait cycle segmentation required by existing gait recognition methods. Second, we demonstrate the effectiveness of ff-GEI by analyzing the cross-covariance of one person's gait data. Then, making use of the temporality of our human gait, we design a novel gait recognition model using Conv-LSTM. Finally, the proposed method is evaluated extensively based on the CASIA Dataset B for cross-view gait recognition, furthermore the OU-ISIR Large Population Dataset is employed to verify its generalization ability. Our experimental results show that the proposed method outperforms other algorithms based on these two datasets. The results indicate that the proposed ff-GEI model using Conv-LSTM, coupled with the new gait representation, can effectively solve the problems related to cross-view gait recognition.
Xiuhui Wang, Wei Qi Yan 0001
Int. J. Neural Syst.2
2020 Object detection based on saturation of visual perception
Wei Qi Yan 0001
Multim. Tools Appl.2
2020 Weighted visual cryptographic scheme with improved image quality
Xuehu Yan, Feng Liu 0032, Wei Qi Yan 0001, Guozheng Yang, Yuliang Lu
Multim. Tools Appl.3
2020 Penrose tiling for visual secret sharing
Xuehu Yan, Wei Qi Yan 0001, Lintao Liu, Yuliang Lu
Multim. Tools Appl.2
2020 Cross-view gait recognition through ensemble learning
Xiuhui Wang, Wei Qi Yan 0001
Neural Comput. Appl.2
2020 Gait recognition using multichannel convolution neural networks
Xiuhui Wang, Wei Qi Yan 0001
Neural Comput. Appl.3
2019 A Sequential CNN Approach for Foreign Object Detection in Hyperspectral Images
Mahmoud Al-Sarayreh, Marlon M. Reis, Wei Qi Yan 0001, Reinhard Klette
CAIP (1)3
2019 A Web-based Augmented Reality Plat-form using Pictorial QR Code for Educational Purposes and Beyond
abstract
Augmented Reality (AR) provides the capability to overlay virtual 3D information onto a 2D printed flat surface; for example, displaying a 3D model on a single flat card that accompanies with the diagram shown in a learning text-book. The student can zoom in and out, rotate, and perceive the animation of the figure in real-time. This will make the educational theory more attractive; hence, motivates students to learn. AR is a great tool; however, the setup and display are not straight-forward (there are many different AR markers with different encryption, decryption methods, and displaying flat-forms). In this paper, we proposed a portable browser-based platform which uses the advantages of AR along with scan-able QR Code on mobile phones to enhance instant 3D visualisation. The user only needs a smart-phone (Apple iPhone or Android) with Internet-enabled; no specific Apps are needed to install. The user scans the QR Code embedded in a colour image, the code will link to a public website, and the website will produce AR Experience right on top of the browser. As a result, it provides a stress-free, low-cost, portable, and promising solution for not only educational purposes but also many other fields such as gaming, property selling, e-commerce, reporting. The set up is convenient: the user uploads a picture (e.g. a racing car), and what actions to be related to it (a 3D model to display, or a movie to play). The system will add on the picture one small colour QR code (to redirect to an online URL) and a thin black border. The user also uploads the 3D model (GLTF files) that he wants to display on top of the card to finish the set-up. At the display, the user can print the AR card, point their smart-phone towards the card, and pre-setup AR models or actions will appear on it. To students, these 3D graphics or animations will allow them to learn and understand the lessons in a much more intuitive way.
Minh Nguyen 0001, Minh Phu Lai, Wei Qi Yan 0001
VRST4
2019 BIIIA: a bioinformatics-inspired image identification approach
Abhimanyu Singh Garhwal, Wei Qi Yan 0001
Multim. Tools Appl.2
2019 BIIGA: Bioinformatics inspired image grouping approach
Abhimanyu Singh Garhwal, Wei Qi Yan 0001
Multim. Tools Appl.2
2019 A new visual evaluation criterion of visual cryptography scheme for character secret image
YaWei Ren, Feng Liu 0001, Wei Qi Yan 0001, Wen Wang 0008
Multim. Tools Appl.3
2018 Human Behaviour Recognition Using Deep Learning
abstract
Traditional human behaviour recognition is mostly based on global features of digital images. Nowadays, with the increase of computing power and processing capacity, deep neural networks (DNNs) acquire a high possibility to detect any objects, which have effectively led to a new era of machine learning. In this paper, we investigated a human behaviour recognition using deep learning based on YOLOv3 model. After a number of experiments conducted, our YOLOv3 model had shown to achieve 80.20% of accuracy in human behaviour recognition with the speed of approximate 15 fps using GPU acceleration. Our direct contributions are: (1) data augment and collection, (2) adjusting deep neural network structures, and (3) superior performance in evaluations for our proposed deep learning model.
Wei Qi Yan 0001, Minh Nguyen 0001
AVSS2
2018 Enhancing Visualisation of Anatomical Presentation and Education Using Marker-based Augmented Reality Technology on Web-based Platform
abstract
The domain of teaching medicine involves the mastery of many complex skills that almost always need to be performed in real life situations following very high professional standards. However, this training is not always possible for various reasons such as ethics, safety and costs. Virtual reality (VR) and augmented reality (AR) have started to be widely used as alternative medical teaching practices. Our proposed AR system works online, so no installation is required on a users device; it only requires a generic colour web-cam to track a pictorial AR tag which contains a hidden QR code. These QR codes contain data, such as the ID of a three-dimensional (3D) model, and merge it with text so that it can be used as both an identity and tag pattern in the AR marker. The system can then show the corresponding computer-generated 3D anatomical models of organs; relevant text information about the subject is displayed above the AR Tag. The tag is numerically encrypted and decrypted and detectable by shape and orientation. Different from other similar techniques, our AR Tag is both a bar-code and a template marker, QR code is used to load a previously setup website, and then the detail of that QR code is used as a template to identify the border and orientation of the marker. The system is thus faster and more robust that allows users to control and navigate the 3D environment by zooming in and out and rotating left and right. It is hoped that this virtual environment will help reduce the need for real-life surgical practice, instead of increasing intuition, the direct 3D perception of the human body and other 3D medical imaging data (mimesis). This system could even be further developed to present the framework of a patient's anatomy.
Minh Nguyen 0001, Hui Le, Wei Qi Yan 0001, Steffan Hooper
AVSS4
2018 Comparative Evaluations of Privacy on Digital Images
abstract
Privacy preservation on social networks is nowadays a societal issue. In this paper, our contributions are to establish such a model for privacy preservation. We use differential privacy for personal privacy analysis and measurement. Our conclusion is that privacy could be measured and preserved if the corresponding approaches could be taken.
Wei Qi Yan 0001
AVSS2
2018 Currency Detection and Recognition Based on Deep Learning
abstract
In recent years, deep learning has become the most popular research direction. It mainly trains the dataset through neural networks. There are many different models that can be used in this research project. Throughout these models, accuracy of currency recognition can be improved. Obviously, such research methods are in line with our expectations. In this paper, we mainly use Single Shot MultiBox Detector (SSD) model based on deep learning as the framework, employ Convolutional Neural Network (CNN) model to extract the features of paper currency, so that we can more accurately recognize the denomination of the currency, both front and back. Our main contribution is through using CNN and SSD, the average accuracy of currency recognition is up to 96.6%.
Wei Qi Yan 0001
AVSS2
2018 An effective method for plate number recognition
Boris Bacic, Wei Qi Yan 0001
Multim. Tools Appl.3
2018 Adopting secret sharing for reversible data hiding in encrypted images
Jian Weng 0001, Wei Qi Yan 0001
Signal Process.3
2017 AndroCon: An Android-Based Context-Aware Middleware Framework for Data Provisioning
abstract
Mobile devices have become major sources of context-aware data due to their ubiquity and sensing capabilities. However, deploying mobile devices as dynamic, unabridged context data provider either locally or remotely is still challenging due to their limited computing capability. Furthermore, integrating physical sensor data with social context data from online social networks is necessary for rich context data provisioning. In this paper, we present AndroCon, an Android-based context-aware middleware framework that enables mobile devices to acquire, integrate, manage, and provision context data. We have applied AndroCon to manage social and physical context data from various sources and have evaluated its performance in terms of power consumption and CPU utilization.
Jian Yu 0002, Quan Z. Sheng, Wei Qi Yan 0001, Olayinka Adeleye
MobiQuitous3
2017 Detection of Adulteration in Red Meat Species Using Hyperspectral Imaging
Mahmoud Al-Sarayreh, Marlon M. Reis, Wei Qi Yan 0001, Reinhard Klette
PSIVT3
2017 Integrated Multi-scale Event Verification in an Augmented Foreground Motion Space
Qin Gu, Jianyu Yang 0001, Wei Qi Yan 0001, Reinhard Klette
PSIVT3
2017 A tile based colour picture with hidden QR code for augmented reality and beyond
abstract
Most existing Augmented Reality (AR) applications use either template (picture) markers or bar-code markers to overlay computer-generated graphics on the real world surfaces. The use of template markers is computationally expensive and unreliable. On the other hand, bar-code markers display only black and white blocks; thus, they look uninteresting and uninformative. In this short paper, we describe a new way to optically hide a QR code inside a tile based colour picture. Each AR marker is built from hundreds of small tiles (just like tiling a bathroom), and the unique gaps between the tiles are used to determine the elements of the hidden QR Code. This novel type of AR marker presents not only a realistic-looking colour picture but also contains self-Correcting information (stored in QR code). In this article, we demonstrate that this tile based colour picture with hidden QR code is relatively robust under various conditions and scaling. We believe many nowadays' AR challenges could be solved with this type of marker. AR-enabled medias could then be easily generated. For instance, it would be capable of storing and displaying virtual figures of an entire book or magazine. Thus, it provides a promising AR approach to be used in many different AR applications; and beyond, it may even replace the barcodes and QR Codes in some cases.
Minh Nguyen 0001, Wei Qi Yan 0001
VRST4
2017 Content based authentication of visual cryptography
Wei Qi Yan 0001, Mohan Kankanhalli
Multim. Tools Appl.2
2016 Adaptive and compressive target tracking based on feature point matching
abstract
In compressive tracking algorithms, a feature reduction projection matrix is constructed by using compressed sensing theory. Target and non-target objects are discriminated by using a naive Bayesian classifier. Such an algorithm may ensure accuracy of target tracking in real-time. But it is not adaptive for tracking with respect to scales and rotations. In this paper, we propose a novel adaptive algorithm based on feature point matching for tracking objects which appear with various changes. We combine weight-average and improved compressive tracking algorithms together for tracking objects, then calculate the corresponding feature points between two subsequent frames of the same object for obtaining the target changes related to various scales and rotations. Our experimental results show that the improved algorithm effectively improves the accuracy of target tracking and ensures adaptability of the tracking algorithm.
Fengjiao Li, Wei Qi Yan 0001, Reinhard Klette
ICPR3
2016 2D Barcodes for visual cryptography
Feng Liu 0001, Wei Qi Yan 0001
Multim. Tools Appl.3
2015 An Improved Aspect Ratio Invariant Visual Cryptography Scheme with Flexible Pixel Expansion
Wen Wang 0008, Feng Liu 0001, Wei Qi Yan 0001, Teng Guo 0005
IWDW3
2015 Face Search in Encrypted Domain
Wei Qi Yan 0001, Mohan Kankanhalli
PSIVT1
2015 Currency security and forensics: a survey
Jarrett Chambers, Wei Qi Yan 0001, Abhimanyu Singh Garhwal, Mohan Kankanhalli
Multim. Tools Appl.2
2015 An empirical approach for currency identification
Wei Qi Yan 0001, Jarrett Chambers, Abhimanyu Singh Garhwal
Multim. Tools Appl.1
2014 Braille for Visual Cryptography
abstract
Visual Cryptography (VC) has been studied as a significant way of information security. In VC, original secret is divided into two images called shares. VC shares show no clue for secret perceptually, whereas participants are able to obtain the secret by simply superimposing the shares. Despite the obvious advantages of VC in crucial secret protection, one of its issues appears to be the authentication method for VC shares. It is likely to seek assistance from other areas of digital image processing. As an international standard reading guidance for the visually impaired people, Braille has been widely used as an effective communication channel. In this paper, we will explain Braille encoding and explain how it is applied to handle the authentication problem in VC. Our contribution is to use Braille for VC. To the best of our knowledge, this is the first time the Braille has been employed to the authentication of VC.
Feng Liu 0001, Wei Qi Yan 0001
ISM3
2014 iNavigation: an image based indoor navigation system
Wei Qi Yan 0001
Multim. Tools Appl.2
2013 An empirical approach for digital currency forensics
abstract
The banknote manufacturing industry is shrouded in secrecy, fundamental mechanics of security components are closely guarded trade secrets. Currency forensics is the application of systematic methods to determine authenticity of questioned currency. However, forensic analysis is a difficult task requiring specially trained examiners, the most important challenge is automating the analysis process reducing human error and time. In this study, an empirical approach for automated currency forensics is formulated and a prototype is developed. A two parts feature vector is defined comprised of color features and texture features. Finally the note in question is classified by a Feedforward Neural Network (FNN) and a measurement of the similarity between template and suspect note is output.
Wei Qi Yan 0001, Jarrett Chambers
ISCAS1
2012 A Secret Enriched Visual Cryptography
Feng Liu 0001, Wei Qi Yan 0001, Chuan Kun Wu
IWDW2
2012 A collusion attack optimization strategy for digital fingerprinting
abstract
Collusion attack is a cost-efficient attack for digital fingerprinting. In this article, we propose a novel collusion attack strategy, Iterative Optimization Collusion Attack (IOCA) , which is based upon the gradient attack and the principle of informed watermark embedding. We evaluate the performance of the proposed collusion attack strategy in defeating four typical fingerprinting schemes under a well-constructed evaluation framework. The simulation results show that the proposed strategy performs more effectively than the gradient attack, and adopting no more than three fingerprinted copies can sufficiently collapse examined fingerprinting schemes. Meanwhile, the content resulted from the proposed attack still preserves high perceptual quality.
Hui Feng 0002, Fuhao Zou, Wei Qi Yan 0001, Zhengding Lu
ACM Trans. Multim. Comput. Commun. Appl.4
2012 Image hatching for visual cryptography
abstract
Image hatching (or nonphotorealistic line-art) is a technique widely used in the printing or engraving of currency. Diverse styles of brush strokes have previously been adopted for different areas of an image to create aesthetically pleasing textures and shading. Because there is no continuous tone within these types of images, a multilevel scheme is proposed, which uses different textures based on a threshold level. These textures are then applied to the different levels and are then combined to build up the final hatched image. The proposed technique allows a secret to be hidden using Visual Cryptography (VC) within the hatched images. Visual cryptography provides a very powerful means by which one secret can be distributed into two or more pieces known as shares. When the shares are superimposed exactly together, the original secret can be recovered without computation. Also provided is a comparison between the original grayscale images and the resulting hatched images that are generated by the proposed algorithm. This reinforces that the overall quality of the hatched scheme is sufficient. The Structural SIMilarity index (SSIM) is used to perform this comparison.
Jonathan Weir, Wei Qi Yan 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.2
2011 Authenticating Visual Cryptography Shares Using 2D Barcodes
Jonathan Weir, Wei Qi Yan 0001
IWDW2
2011 Fine-search for image copy detection based on local affine-invariant descriptor and spatial dependent matching
Liyun Wang, Fuhao Zou, Wei Qi Yan 0001
Multim. Tools Appl.4
2011 A comprehensive study of visual event computing
Wei Qi Yan 0001, Declan F. Kieran, Setareh Rafatirad, Ramesh Jain 0001
Multim. Tools Appl.1
2010 A Framework for an Event Driven Video Surveillance System
abstract
In this paper we present an event driven surveillance system. The purpose of this system is to enable thorough exploration of surveillance events. The system uses a client-server web architecture as this provides scalability for further development of the system infrastructure. The system is designed to be accessed by surveillance operators who can review and comment on events generated by our event detection processing modules. The presentation interface is based around a cross between Gmail and YouTube, as we believe these interfaces to be intuitive for ordinary computer operators. Our motivation is to fully utilize the events archived in our database and to further refine the relevant events. We do not just focus on event detection, but are working towards the optimization of event detection. To the best of our knowledge this system provides a novel approach to the technological surveillance paradigm.
Declan F. Kieran, Wei Qi Yan 0001
AVSS2
2010 Intelligent Sensor Information System For Public Transport - To Safely Go
abstract
The Intelligent Sensor Information System (ISIS) is described. ISIS is an active CCTV approach to reducing crime and anti-social behavior on public transport systems such as buses. Key to the system is the idea of event composition, in which directly detected atomic events are combined to infer higher-level events with semantic meaning. Video analytics are described that profile the gender of passengers and track them as they move about a 3-D space. The overall system architecture is described which integrates the on-board event recognition with the control room software over a wireless network to generate a real-time alert. Data from preliminary data-gathering trial is presented.
Paul Miller 0003, Weiru Liu, Chris Fowler, Huiyu Zhou 0001, Jiali Shen, Jianbing Ma, Jianguo Zhang 0001, Wei Qi Yan 0001, Kieran McLaughlin, Sakir Sezer
AVSS8
2010 Human Localization in a Cluttered Space Using Multiple Cameras
abstract
The use of single and dual-camera approaches to locating a subject in a 3-D cluttered space is investigated. Specifically, we investigate the case where the lower portion of the body may be occluded, e.g., by a chair on a bus. Experiments were conducted involving eleven subjects moving along a pre-designated route within a cluttered space. For each time instant the position of each subject was manually estimated and compared to that produced automatically. The dual camera approach was found to give significantly better performance than the single camera approach. It was found that inaccurate bounding of the lowest part of the subject, due to occlusion, led to localisation errors in range as large as 10m for the latter. Using the side bounds of the detected object, which were found to be robust, accurate azimuth estimates can be obtained for a single camera. The dual-camera approach exploits the greater degree of accuracy in azimuth to estimate the range through triangulation, giving average localisation errors of 40cm over the space of interest.
Jiali Shen, Wei Qi Yan 0001, Paul Miller 0003, Huiyu Zhou 0001
AVSS2
2010 Resolution variant visual cryptography for street view of Google Maps
abstract
Resolution variant visual cryptography takes the idea of using a single share of visual cryptography (VC) to recover a secret from an image at multiple resolutions. That means, viewing the image on a one-to-one basis and superimposing the share will recover the secret. However, if the image is zoomed, using that same share we can recover other secrets at different levels. The same share is used at these varying resolutions in order to recover a large amount of hidden secrets. This process is quite similar to watermarking an image, whereby nothing can be seen while fully zoomed out, but as the zoom level is increased the watermark becomes visible. This would also be associated with a recursive style of secret sharing. This type of secret sharing scheme would be appropriate for recovering specific types of censored information, such as vehicle registration numbers within certain types of images. This adds an additional dimension to our scheme: content based visual cryptography.
Jonathan Weir, Wei Qi Yan 0001
ISCAS2
2010 A Novel Collusion Attack Strategy for Digital Fingerprinting
Hui Feng 0002, Fuhao Zou, Wei Qi Yan 0001, Zhengding Lu
IWDW4
2010 Plane Transform Visual Cryptography
Jonathan Weir, Wei Qi Yan 0001
IWDW2
2010 Optimal collusion attack for digital fingerprinting
abstract
The collusion attack is a cost-efficient attack against digital finger-printing where classes of users combine their fingerprinted content for the purpose of attenuating or removing the fingerprints. A recently introduced gradient attack which appeared in ACM MM 2004, demonstrated its efficacy in defeating most spread-spectrum based fingerprints. In this paper, we propose a novel collusion attack strategy, Iterative Optimization Collusion Attack (IOCA), which is based upon the gradient attack and the geometric principal of a Voronoi diagram. The simulation results, under the assumption that orthogonal fingerprints are used, show that the proposed collusion attack performs more effectively than the gradient attack. Less than five fingerprinted pieces of content can sufficiently interrupt orthogonal fingerprints accommodating many thousands of users, meanwhile, high perceptual quality of the attacked content is obtained after the proposed collusion attack.
Hui Feng 0002, Fuhao Zou, Wei Qi Yan 0001, Zhengding Lu
ACM Multimedia4
2009 Event Composition with Imperfect Information for Bus Surveillance
abstract
Demand for bus surveillance is growing due to the increased threats of terrorist attack, vandalism and litigation. However, CCTV systems are traditionally used in forensic mode, precluding an in-time reaction to an event. In this paper, we introduce a real-time event composition framework which can support the instant recognition of emergent events based on uncertain or imperfect information gathered from multiple sources. This framework deploys a rule-based reasoning component that can infer malicious situations (composite events) from a set of correlated atomic events. These are recognized by applying analytic algorithms to the multimedia contents of bus surveillance data. We demonstrate the significance and usefulness of our framework with a case study of an on-going bus surveillance project.
Jianbing Ma, Weiru Liu, Paul Miller 0003, Wei Qi Yan 0001
AVSS4
2009 Sharing Multiple Secrets using Visual Cryptography
abstract
Visual cryptography provides a very powerful technique by which one secret can be distributed into two or more pieces known as shares. When the shares on transparencies are superimposed exactly together the original secret can be discovered without computer participation. In this paper, we take multiple secrets into consideration, and generate a master key for all the secrets; correspondingly, we share each secret using the master key and obtain multiple shares. We merge these shares into a combined share, we adjust the master key and generate a new key. The secrets are revealed when the key is superimposed on the combined share in different locations using the proposed scheme. We provide the corresponding results in this paper.
Jonathan Weir, Wei Qi Yan 0001
ISCAS2
2009 Dot-Size Variant Visual Cryptography
Jonathan Weir, Wei Qi Yan 0001
IWDW2
2008 A cross-modal approach for karaoke artifacts correction
Wei Qi Yan 0001, Mohan Kankanhalli
Multim. Tools Appl.1
2008 Progressive Audio Scrambling in Compressed Domain
abstract
Audio scrambling can be employed to ensure confidentiality in audio distribution. We first describe scrambling for raw audio using the discrete wavelet transform (DWT) first and then focus on MP3 audio scrambling. We perform scrambling based on a set of keys which allows for a set of audio outputs having different qualities. During descrambling, the number of keys provided and the number of rounds of descrambling performed will decide the audio output quality. We also perform scrambling by using multiple keys on the MP3 audio format. With a subset of keys, we can descramble to obtain a low quality audio. However, we can obtain the original quality audio by using all of the keys. Our experiments show that the proposed algorithms are effective, fast, simple to implement while providing flexible control over the progressive quality of the audio output. The security level provided by the scheme is sufficient for protecting MP3 music content.
Wei Qi Yan 0001, Wei-Gang Fu, Mohan Kankanhalli
IEEE Trans. Multim.1
2007 A scalable signature scheme for video authentication
Pradeep K. Atrey, Wei Qi Yan 0001, Mohan Kankanhalli
Multim. Tools Appl.2
2007 Multimedia simplification for optimized MMS synthesis
abstract
We propose a novel transcoding technique called multimedia simplification which is based on experiential sampling. Multimedia simplification helps optimize the synthesis of MMS (multimedia messaging service) messages for mobile phones. Transcoding is useful in overcoming the limitations of these compact devices. The proposed approach aims at reducing the redundancy in the multimedia data captured by multiple types of media sensors. The simplified data is first stored into a gallery for further usage. Once a request for MMS is received, the MMS server makes use of the simplified media from the gallery. The multimedia data is aligned with respect to the timeline for MMS message synthesis. We demonstrate the use of the proposed techniques for two applications, namely, soccer video and home care monitoring video. The MMS sent to the receiver can basically reflect the gist of important events of interest to the user. Our technique is targeted towards users who are interested in obtaining salient multimedia information via mobile devices.
Wei Qi Yan 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.1
2005 Analogies based video editing
Wei Qi Yan 0001, Mohan Kankanhalli, Jun Wang 0012
Multim. Syst.1
2005 Automatic video logo detection and removal
Wei Qi Yan 0001, Jun Wang 0012, Mohan Kankanhalli
Multim. Syst.1
2004 Mosaic based view enlargement for moving objects in moving pictures
abstract
Conventional mosaicing techniques convert a video from frame-based representation to scene-based representation, but they usually lack dynamic information so that their mosaic is not complete. In this paper, we present a novel method to detect moving objects in the video sequences, then add them into the static background mosaic to represent the scene completely. This novel algorithm separates static and dynamic information in a video sequence, builds the background mosaic from static part and reconstructs moving objects on the static mosaic. We have implemented our techniques and the experimental results demonstrate the effectiveness of our approach
Mohan Kankanhalli, S. H. Srinivasan, Wei Qi Yan 0001
ICME4
2004 A Hierarchical Signature Scheme for Robust Video Authentication using Secret Sharing
abstract
Ensuring the integrity of a digital video is an important and challenging research problem arising out of many video applications. In this paper, we present a hierarchical framework for video authentication based on cryptographic secret sharing that protects a video from spatial cropping and temporal jittering, yet is robust against frame dropping in the streaming video scenario. Our algorithm provides a tradeoff between security and robustness by having configurable inputs. The authentication signature is compact and very sensitive against spatial attacks such as region tampering, and interframe attacks like frame replacement, major frame dropping, and frame reordering. Given a video, we identify the key frames based on different energy between the frames. Considering video frames as shares, we compute the secret at three hierarchical levels. The master secret is used as digital signature to authenticate the video. We present extensive experimental results which show the utility of our technique.
Pradeep K. Atrey, Wei Qi Yan 0001, Ee-Chien Chang, Mohan Kankanhalli
MMM2
2003 Colorizing infrared home videos
abstract
A color video always conveys more vivid sentiments than a grayscale one. Obtaining a grayscale video from a color video is almost trivial but the converse is known to be hard. Nowadays, digital camcorders come equipped with an infrared device for night shot that enables one to shoot home videos in the dark. Unfortunately, the infrared lighting device used generates a "green-scale" video which is akin to a grayscale video albeit possessing all tints of green. In this paper, we present a novel technique for colorizing infrared home videos. We first convert the green scale video into grayscale, afterwards our technique involves generating key-frames for every shot and then building up a one to one correspondence map between the key frames and the designated color images. These pairs are used to generate the color palette table for the video segment, which is then utilized to colorize that segment of the home video. Our novel technique could also be applied for colorizing X-ray videos generated by diagnostic imaging devices as well as surveillance videos generated by baggage scanners at airports.
Wei Qi Yan 0001, Mohan Kankanhalli
ICME1
2003 Scrambling of engineering drawings
abstract
Engineering drawings are ubiquitously used for capturing, conveying and archiving innovative engineering designs. Many engineering companies' core intellectual property resides in their proprietary engineering drawings. Therefore, protection of such vital data is extremely important. This paper provides a swap-transformation matrix based approach to scramble engineering drawings in order to enable confidentiality. An engineering drawing involves the topological information and vertex information. The vertex information is more valuable than the topological information, since the vertices information primarily determines the content of engineering drawings. We argue that the vertex information is more valuable than the topological information, even if some topological information is lost, a drawing may be reconstructed from the vertex positions. We provide for three keys to ensure the security of the drawing. The technique can facilitate digital rights management of engineering drawings. The advantages of our technique are that scrambling is computationally less intensive than encryption and it allows for partial obfuscation.
Wei Qi Yan 0001, Mohan Kankanhalli
ICME1
2002 Erasing video logos based on image inpainting
abstract
A video logo is usually a declaration of the video copyright. However it sometimes causes visual discomfort due to the presence of multiple logos in videos that have been filed and exchanged by different channels. We present an approach to erase logos from video clips. Based on the histogram energy analysis of the relevant video frames, we obtain the best quality logo frame that can be easily processed in the selected region of video frames. After that, we mark the logo area in the entire sequence of frames and inpaint each frame of the video logo based on color interpolation. We describe our technique and also provide experimental results.
Wei Qi Yan 0001, Mohan Kankanhalli
ICME (2)1
2002 Detection and removal of lighting & shaking artifacts in home videos
abstract
Many amateur videographers, like home video enthusiasts, may capture videos that are not of a professional quality. Many minor but visually annoying distortions like lighting imbalance and shaking artifacts could be introduced by the unskilled operations of the video camcorder. Since home videos constitute footage of great sentimental value, such videos cannot be summarily discarded. Unlike movies and sitcoms, shot re-takes of important events, such as wedding ceremonies are just not possible. Therefore, such distortions need to be corrected. In this paper, we present a novel method to detect segments of videos that have lighting and shaking artifacts. These segments can then be subjected to a restoration process that can remove these artifacts. We present techniques to correct lighting artifacts by appropriately adjusting the luminance. In order to remove the shaking artifact, image mosaicing is first employed to build a mosaic frame for the segment with the aid of edge blending techniques. Subsequently a Bezier-curve based blending of motion trajectory is employed to perform motion-compensated filtering of the shaking artifact. The restored video is then created by appropriately cropping the mosaic frame based on the compensated motion trajectory. We have implemented the developed techniques and the experimental results on home videos demonstrate the effectiveness of our approach. Detection and removal of artifacts are significant in other videos as well as those obtained from autonomous vehicles, robots and remote sensing.
Wei Qi Yan 0001, Mohan Kankanhalli
ACM Multimedia1
2002 Digital Image Watermarking Based on Discrete Wavelet Transform
Wei Qi Yan 0001, Dongxu Qi
J. Comput. Sci. Technol.2
2000 A Novel Digital Image Hiding Technology Based on Tangram and Conway's Game
abstract
We present a novel technology which could hide digital image information based on tangram and Conways' game. For explaining the scheme, we firstly discuss the old tangram puzzle and present how to transform between two different images. Then based on Conway' game, we present a scrambling technique which helps to hide the transform information. The advantage of this method is its security.
Wei Qi Yan 0001, Dongxu Qi
ICIP2