VLDB 2026 Research / reviewers in the wild / expert
Wanqing Li 0001
dblp:57/3330
· DBLP profile ↗
144ranked-venue papers
8as first author
45since 2021 · last 2026
0000-0002-4427-2687ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 8 first-author · 21 since 2021Artificial intelligence and machine learning · 65 · 27 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MACRec: A Multi-View Subspace Alignment Framework for Contrastive Sampling Calibration in RecommendationabstractGraph Contrastive Learning (GCL) has proven effective in mitigating data sparsity and enhancing representation learning for recommendation. Yet, most GCL frameworks indiscriminately treat all non-anchor nodes as negatives during contrastive sampling, often leading to the false negative problem where semantically similar nodes are incorrectly repelled. Previous attempts to mitigate this issue rely on predetermined heuristics or local neighborhood mining, which struggle to reliably identify false negatives. More critically, they often overlook authentic user-item interactions for anchoring sample relationships. As a result, this paper presents MACRec, a Multi-View subspace-Alignment framework designed to Calibrate contrastive sampling in GCLbased Recommendation. MACRec comprises three core components: (1) a Multi-View Affinity (MVA) module that captures consistent semantic relations across multiple augmentations via self-expression modeling; (2) a Cross-Subspace Alignment (CSA) mechanism that leverages authentic useritem behavioral interactions to enforce semantic consistency across user and item subspaces; and (3) a Calibrationbased Contrastive Reweighting (CCR) strategy to dynamically down-weight potential false negatives during the contrastive learning process. Extensive experiments on three realworld benchmarks demonstrate that MACRec consistently improves performance across various augmentation backbones, achieving up to 14.55% relative gains. Junping Liu, Mingchao Yu, Xinrong Hu, Wanqing Li 0001, Jie Yang 0009, Yi Guo 0001 |
AAAI | 5 |
| 2026 | FLAIR: Steering LLM Mathematical Problem Solving based on A Fuzzy-Logic-AssIsted ReasonerabstractHao Wu, Hongru Sun, Wanqing Li, Xinguo Yu, Hao Ming, Xiao Luo, Wenbin Zhang, Jiahong Zhao, Yi Guo, Jie Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Wu 0094, Hongru Sun, Wanqing Li 0001, Xinguo Yu, Hao Ming 0001, Xiao Luo 0001, Wenbin Zhang 0002, Jiahong Zhao, Yi Guo 0001, Jie Yang 0009 |
ACL (1) | 3 |
| 2026 | Generative Data Augmentation for Skeleton Action Recognition
Wanqing Li 0001, Anthony Adeyemi-Ejeye, Andrew Gilbert |
FG | 2 |
| 2026 | UIL-AQA: Uncertainty-Aware Clip-Level Interpretable Action Quality AssessmentabstractAbstract This work proposes UIL-AQA for long-term Action Quality Assessment AQA designed to be clip-level interpretable and uncertainty-aware. AQA evaluates the execution quality of actions in videos. However, the complexity and diversity of actions, especially in long videos, increase the difficulty of AQA. Existing AQA methods solve this by limiting themselves generally to short-term videos. These approaches lack detailed semantic interpretation for individual clips and fail to account for the impact of human biases and subjectivity in the data during model training. Moreover, although query-based Transformer networks demonstrate strong capabilities in long-term modelling, their interpretability in AQA remains insufficient. This is primarily due to a phenomenon we identified, termed Temporal Skipping , where the model skips self-attention layers to prevent output degradation. We introduce an Attention Loss function and a Query Initialization Module to enhance the modelling capability of query-based Transformer networks. Additionally, we incorporate a Gaussian Noise Injection Module to simulate biases in human scoring, mitigating the influence of uncertainty and improving model reliability. Furthermore, we propose a Difficulty-Quality Regression Module, which decomposes each clip’s action score into independent difficulty and quality components, enabling a more fine-grained and interpretable evaluation. Our extensive quantitative and qualitative analysis demonstrates that our proposed method achieves state-of-the-art performance on three long-term real-world AQA datasets. Our code is available at: https://github.com/dx199771/Interpretability-AQA Wanqing Li 0001, Anthony Adeyemi-Ejeye, Andrew Gilbert |
Int. J. Comput. Vis. | 3 |
| 2026 | Unrolling operator splitting in learning PDEs for object detectionabstractObject detection presents significant challenges due to the variability in object scale, location, and orientation within images. Most state-of-the-art detectors are based on convolutional or Transformer architectures, which, while effective, often result in deep, opaque models that generalise poorly and lack interpretability. In contrast, iterative algorithms offer greater transparency and generalisation, albeit at the cost of efficiency and accuracy. In this work, we reformulate object detection as a partial differential equation (PDE)-constrained optimal control problem. This formulation exploits linear combinations of fundamental differential invariants—such as translation and rotation invariance—to embed structural priors into the learning process. We solve this problem using operator splitting via the Alternating Direction Method of Multipliers (ADMM), and unroll each optimisation step into a network layer, yielding a novel architecture: ADMM-ODNet. This approach provides a principled and interpretable alternative to conventional deep networks. Experimental results on the Corel, Pascal VOC and COCO datasets demonstrate that ADMM-ODNet outperforms leading models such as Cascade Mask R-CNN, Swin Transformer, Deformable DETR, DINO, DN-and RT-DETR, and achieves performance comparable to Plain DETR and YOLO, while requiring significantly fewer parameters. Banu Wirawan Yohanes, Philip Ogunbona, Wanqing Li 0001 |
Neurocomputing | 3 |
| 2026 | Learning Disentangled Representations for Generalized Multi-View ClusteringabstractMulti-View Clustering (MVC) has gained significant attention for its ability to leverage complementary information across diverse views. However, existing deep MVC methods often struggle with view-distribution entanglement during cross-view fusion, which hampers the quality of the shared latent space and leads to suboptimal clustering performance. To address this issue, we propose the Generalized Multi-view Auto-Encoder (GMAE), a framework designed to preserve cross-view complementarity through disentangled representation learning. Specifically, GMAE employs dual-path autoencoders to decouple source features into view-specific and view-common embeddings, facilitating the discovery of clearer clustering structures. We further construct cross-view adversarial discriminators to guide view-specific encoders in capturing more discriminative features. By strategically modulating mutual information, GMAE effectively aligns distributions and prevents representation collapse, ensuring the generation of robust, non-trivial embeddings. Comprehensive experiments on 13 benchmark datasets demonstrate that GMAE consistently outperforms state-of-the-art methods in both complete and incomplete MVC tasks. Xin Zou 0001, Ruimeng Liu, Chang Tang, Zhenglai Li, Xinwang Liu 0002, Kunlun He, Wanqing Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Efficient spatial pyramid for object detectionabstractOne-stage object detectors emerge as a trade-off between detection accuracy and speed. However, they do not exploit long-range context relationship and their performance can easily drop in complex scenes. In this work, we propose efficient spatial pyramid based on graph model of Laplacian with a novel graph correlation filter. This filter is designed to measure symmetric uncertainty among features to learn long-range context relationship, while preserving important features. Furthermore, we employ Winograd algorithm to reduce floating point operations significantly without decreasing detection performance, by trading-off costly multiplication operations with more additions. They enable an efficient object detection. Extensive experiments were conducted on two challenging object detection datasets, COCO and KITTI. The proposed network was compared to state-of-the-art efficient object detectors, MobileNet-SSD Lite, YOLO, MobileVIT, Tiny DSOD, and EfficientDet. Detailed convergence proof and training epoch analysis provide strong support and evidence for the achieved results improving overall detection accuracy in complex scenes with less computational resources. • Hierarchical Laplacian-driven design capturing both fine-grained and global context. • Symmetrical uncertainty filter for capturing long-range feature interactions. • Accelerated convolutional layers via Winograd minimal filtering. • Thorough empirical validation and ablation study on COCO and KITTI datasets. • ESP attains the best-in-class accuracy-to-efficiency trade-off. Banu Wirawan Yohanes, Philip Ogunbona, Wanqing Li 0001 |
Pattern Recognit. | 3 |
| 2026 | Bayesian Joint Nonlinear System Model Learning, Sensing and Signal Detection in ISAC With Hardware ImperfectionsabstractThis work addresses the challenges of communication signal detection and direction of arrival (DOA) estimation in integrated sensing and communications (ISAC) systems with hardware imperfections. Conventional signal processing techniques often fail to effectively manage the complex nonlinearities caused by hardware imperfections, such as those introduced by power amplifiers and local oscillators. Recently, deep neural networks (DNNs) have been employed to mitigate the hardware imperfections, which however require a substantial amount of pilot signals for training, leading to unacceptable overhead and impracticality in fast time-varying channels. In this work, we employ an NN to characterize the nonlinear system, and propose a novel iterative approach to joint NN-based nonlinear system model learning, signal detection and DOA estimation. Instead of relying on pilot signals for NN learning, the proposed approach utilizes communication data signals as virtual training samples, enabling more accurate nonlinear model learning, which subsequently enhances signal detection and DOA estimation. A Bayesian framework is applied to the joint problem, wherein the NN parameters, the communication signals and the DOAs are jointly obtained by developing a message passing based inference algorithm. In particular, we impose sparse priors on the weights of the NN, so that overfitting can be better handled, resulting in significant improvement in system modeling performance. Extensive simulation results show that, compared to the state-of-the-art approaches, the proposed one delivers significantly better performance. Qinghua Guo 0001, Ming Jin 0001, Zhengdao Yuan, Guisheng Liao, Wanqing Li 0001, Yuntao Wu |
IEEE Trans. Wirel. Commun. | 6 |
| 2025 | Negative-Free Graph Contrastive Learning for RecommendationabstractGraph Contrastive Learning (GCL) emerges as a powerful approach in recommendation systems, leveraging graph structures to learn effective representations. However, existing contrastive sampling strategies often introduce unintended biases, most notably, the misclassification of genuine positive samples as negatives, which undermines representation quality and overall recommendation performance. Accordingly, this paper revisits the conventional contrastive sampling and introduces Negative-Free Sampling for Graph Contrastive Learning (NFS). NFS adopts a two-stage sampling strategy that selectively identifies and utilizes only positive instances during training. By removing reliance on negative samples, it effectively mitigates misclassification bias and improves the semantic alignment between related representations. In addition, a comprehensive theoretical analysis is also provided to establish the robustness of NFS against representation collapse. Experimental results on three benchmarks demonstrate that NFS consistently outperforms or performs state-of-the-art methods, achieving up to a 14.2% relative improvement across evaluated datasets. In addition, a detailed ablation study is also provided to examine how exclusively leveraging positive samples contributes to the efficiency of GCL. The results further demonstrate the plug-and-play nature of the proposed method and its resilience to noisy data. Junping Liu, Mingchao Yu, Xinrong Hu, Jie Yang 0009, Yi Guo 0001, Wanqing Li 0001, Wenbin Zhang 0002 |
ICDM | 6 |
| 2025 | A new active learning method based on uncertainty sampling and redundancy alleviation strategyabstractActive learning (AL) is a semi-supervised learning algorithm based on human-computer interaction. It aims to achieve satisfactory performance with less manual annotation costs. How to select the most informative samples from unlabeled data for manual annotation is crucial for AL. In this paper, we propose a hybrid sampling AL method FSP-Coreset, which is based on the uncertainty and diversity of unlabeled samples. Firstly, we design a new method First Second Probability (FSP) for measuring the uncertainty of unlabeled samples based on the ratio of the first largest prediction probability to the sum of the first and the second largest probability. Compared with the entropy-based method, the FSP reduces the impact of lower probability on uncertainty measurement and can select samples with significant. Secondly, we adopt the diversity-based method Coreset to alleviate the redundant samples. Through these methods, we obtain unlabeled samples with high uncertainty and low redundant. Experimental results show that our FSP-Coreset not only alleviates redundant uncertainty samples but also accelerates the convergence of Coreset. Furthermore, the performance of FSP-Coreset is superior to comparison methods. The ablation experimental results show that our FSP performs better than the entropy-based uncertainty sampling method, and the hybrid sampling method FSP-Coreset is better than the single method. Ruihua Wen, Mingliang Xu 0001, Wanqing Li 0001 |
ISCAS | 5 |
| 2025 | Skeleton-based action recognition through attention guided heterogeneous graph neural network
Tianchen Li, Pei Geng, Xuequan Lu, Wanqing Li 0001, Lei Lyu 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Asynchronous Joint-Based Temporal Pooling for Skeleton-Based Action RecognitionabstractDeep neural networks for skeleton-based human action recognition (HAR) often utilize traditional averaging or maximum temporal pooling to aggregate features by treating all joints and frames equally. However, this approach can excessively aggregate less discriminative or even indiscriminative features into the final feature vectors for recognition. To address this issue, a novel method called asynchronous joint adaptive temporal pooling (AJTP) is introduced in this paper. The method aims to enhance action recognition by identifying a set of informative joints across the temporal dimension and applying a joint-based and asynchronous motion-preservative pooling rather than conventional frame-based pooling. The effectiveness of the proposed AJTP has been empirically validated by integrating it with popular Graph Convolutional Network (GCN) models on three benchmark datasets: NTU RGB+D 120, PKUMMD, and Kinetic400. The results have shown that a GCN model with AJTP substantially improves performance compared to its counterpart GCN model with conventional temporal pooling techniques. The source code is available athttps://github.com/ShanakaRG/AJTP. Shanaka Ramesh Gunasekara, Wanqing Li 0001, Jie Yang 0009, Philip Ogunbona |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Unsupervised Feature Enrichment and Fidelity Preservation Learning Framework for Skeleton-Based Action RecognitionabstractUnsupervised skeleton-based action recognition has achieved remarkable progress recently. Existing unsupervised learning methods suffer from severe overfitting problem, and thus small networks are used, significantly reducing the representation capability. To address this problem, the overfitting mechanism behind the unsupervised learning for skeleton-based action recognition is first investigated. It is observed that skeleton is already a relatively high-level and low-dimension feature, but not in the same manifold as the features for action recognition. Simply applying the existing unsupervised learning method tends to produce features that discriminate the different samples rather than action classes, resulting in the overfitting problem. To address this problem, this paper proposes an Unsupervised spatial-temporal Feature Enrichment and Fidelity Preservation (U-FEFP) learning framework to generate rich distributed features that contain all the information of a skeleton sample. A spatial-temporal feature transformation subnetwork is developed using channel-wise topology refinement graph convolutional block and graph convolutional gated recurrent unit block as the basic feature extraction network. The unsupervised Bootstrap Your Own Latent-based learning is utilized to generate rich distributed features, and the unsupervised pretext task-based learning is employed to preserve the information contained in the skeleton. The two unsupervised learning ways are collaborated as U-FEFP to produce robust and discriminative representations. Experimental results on four widely used benchmarks, namely NTU-RGB+D-60, PKU-MMD, NTU-RGB+D-120 and AAV-Human dataset, demonstrate that the proposed U-FEFP obtains the best result compared with the state-of-the-art unsupervised learning methods. Chuankun Li, Shuai Li 0005, Yanbo Gao, Xingyu Gao 0001, Ping Chen 0004, Wanqing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Guest Editorial: When Multimedia Meets Food: Multimedia Computing for Food Data Analysis and Applications
Weiqing Min, Shuqiang Jiang, Petia Radeva, Vladimir Pavlovic 0001, Chong-Wah Ngo, Kiyoharu Aizawa, Wanqing Li 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Boosting Generalization of Semantic Segmentation With Unseen Style Seeking-Based Meta-LearningabstractThis article considers a worst and most challenging scene in domain generalization (DG), where a model aims to generalize well on unseen domains while only one single domain is available for training. Existing randomization-based methods achieve this goal by enriching the style of the training data. However, they fail to guarantee the diversity of newly generated data required for generalization and thus lead to insufficient expansion of the training distribution. Thus, we propose a novel single DG (SDG) framework, unseen style seeking-based meta-learning (USSML). In USSML, multiple plausible domains with various styles are first constructed from a single source domain and the combination is performed across generated domains to emulate unseen images, extending the distribution boundaries of the source domain. The domain combination is performed at two levels, i.e., global and instance, to meet the generalization challenge in semantic segmentation. Then, the generated diverse domains are further exploited to force the model to optimize in an unbiased manner across all domains by relearning regions lacking domain-invariant representation capability, driving the model toward domain invariance. A point worth mentioning is that the proposed method is easily integrated into existing segmentation methods with little computational cost to improve their generalization. Extensive experiments are conducted on five popular segmentation datasets and the results have verified the effectiveness of USSML in improving the model's generalization and the superiority of USSML over existing works. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Wanqing Li 0001, Dou Quan, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Interpretable Long-term Action Quality Assessment
Wanqing Li 0001, Anthony Adeyemi-Ejeye, Andrew Gilbert |
BMVC | 3 |
| 2024 | Extractive Question Answering with Contrastive Puzzles and Reweighted Clues
Jie Yang 0009, Wanqing Li 0001 |
ICDAR (6) | 3 |
| 2024 | Frequency-Domain Transformation-Based Dynamic Gesture Recognition with Skeleton
Chuankun Li, Shuai Li 0005, Wanqing Li 0001, Danyan Xie |
PRCV (3) | 4 |
| 2024 | Static graph convolution with learned temporal and channel-wise graph topology generation for skeleton-based action recognition
Chuankun Li, Shuai Li 0005, Yanbo Gao, Lijuan Zhou 0002, Wanqing Li 0001 |
Comput. Vis. Image Underst. | 5 |
| 2024 | An attention-based CNN for automatic whole-body postural assessmentabstractFully automatic postural assessment is highly useful, but has been challenging. Conventional methods either require manual assessment by ergonomists or depend on special devices that are intrusive, thus being hardly feasible in daily activities and workplaces. In this work, an attention-based convolutional neural network (CNN) is developed for automatic whole-body postural assessment. The proposed network learns to identify highly relevant regions (or body parts) and extract features automatically. Risk of the posture is estimated from the extracted features accordingly. To evaluate the proposed method, a postural dataset, referred to as pH36M, is created by re-targeting Human3.6M, one of the largest publicly available datasets for pose estimation using the Rapid Entire Body Assessment (REBA) criteria. Experimental results on pH36M demonstrate that proposed method achieves promising performance in comparison to baselines and the average assessment scores are substantially aligned with human assessment with a Kappa value of 0.73. Zewei Ding, Wanqing Li 0001, Jie Yang 0009, Philip Ogunbona |
Expert Syst. Appl. | 2 |
| 2024 | DFN: A deep fusion network for flexible single and multi-modal action recognitionabstractMulti-modal action recognition methods can be generally classified into two categories: (1) fusing multi-modal features with simple concatenation or fusing the classification scores of individual modalities without considering the interaction among the multi-modalities; (2) using one of the modalities as privileged information in training to boost the recognition on the other modalities in inference. The former approach usually is not able to deal with the cases where one of the modalities is missing. In the latter, the trained classifier does not work on the privileged modality. To address these shortcomings, this paper presents a novel end-to-end trainable deep fusion network (DFN) that is able to improve the performance not only in the cases where all modalities are available and also in the cases where there is a missing modality. The DFN is simple yet effective with the capability of retrieving an estimation of one modality by using another modality through a Multilayer Perceptron (MLP). In order to better preserve structure information, the DFN first maps the individual modality features to a high dimensional Kronecker-product space and subsequently learns a low-dimensional discriminative space for classification. The effectiveness of the proposed DFN has been verified on three benchmark datasets: the large NTU RGB+D, UTD-MHAD, and SYSU-3D datasets and it has achieved state-of-the-art results. Chuankun Li, Yonghong Hou, Wanqing Li 0001, Zewei Ding, Pichao Wang |
Expert Syst. Appl. | 3 |
| 2024 | MGC: MLP-Guided CNN Pretraining Using a Small-Scale Dataset for Remote Sensing ImagesabstractTo overcome the inherent domain gap between natural images and remote sensing images (RSIs), it is highly desirable to develop pretraining methods specifically for RSIs. Considering the lack of widely recognized large-scale benchmarks like ImageNet in the RSI community and limited computational resources, this article proposes multilayer perceptron (MLP)-guided convolutional neural network (CNN) (MGC), a method that employs an MLP to guide the pretraining of a CNN from small-scale datasets for RSIs. MGC has two encoders, each consisting of a CNN branch and an MLP branch. We first contrast pairwise samples from the same type of branches or different types of branches across the encoders and employ a positive-pair guidance strategy to explore consistency. Due to the inherent locality issue of shallow layers in a CNN, the CNN branches often do not attend to correct foreground regions such as objects, regions of interest, and land coverage. Therefore, we further propose an attention guidance strategy to guide the CNN branches to focus on foreground regions and learn discriminative representations effectively. The proposed MGC method is validated by pretraining a CNN model using the MGC and applying it to different downstream tasks including scene classification, rotated object detection, semantic segmentation, and change detection on ten datasets. Results have confirmed the effectiveness of the proposed MGC. Our code will be released at:https://github.com/benesakitam/MGC. Zhihao Li 0005, Biao Hou, Wanqing Li 0001, Zitong Wu, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Hierarchical Aggregated Graph Neural Network for Skeleton-Based Action RecognitionabstractSupervised human action recognition methods based on skeleton data have achieved impressive performance recently. However, many current works emphasize the design of different contrastive strategies to gain stronger supervised signals, ignoring the crucial role of the model's encoder in encoding fine-grained action representations. Our key insight is that a superior skeleton encoder can effectively exploit the fine-grained dependencies between different skeleton information (e.g., joint, bone, angle) in mining more discriminative fine-grained features. In this paper, we devise an innovative hierarchical aggregated graph neural network (HA-GNN) that involves several core components. In particular, the proposed hierarchical graph convolution (HGC) module learns the complementary semantic information among joint, bone, and angle in a hierarchical manner. The designed pyramid attention fusion mechanism (PAFM) fuses the skeleton features successively to compensate for the action representations obtained by the HGC. We use the multi-scale temporal convolution (MSTC) module to enrich the expression capability of temporal features. In addition, to learn more comprehensive semantic representations of the skeleton, we construct a multi-task learning framework with simple contrastive learning and design the learnable data-enhanced strategy to acquire different data representations. Extensive experiments on NTU RGB+D 60/120, NW-UCLA, Kinetics-400, UAV-Human, and PKUMMD datasets prove that the proposed HA-GNN without contrastive learning achieves state-of-the-art performance in skeleton-based action recognition, and it achieves even better results with contrastive learning. Pei Geng, Xuequan Lu, Wanqing Li 0001, Lei Lyu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Self-Supervised Monocular Depth Estimation via Binocular Geometric Correlation LearningabstractMonocular depth estimation aims to infer a depth map from a single image. Although supervised learning-based methods have achieved remarkable performance, they generally rely on a large amount of labor-intensively annotated data. Self-supervised methods, on the other hand, do not require any annotation of ground-truth depth and have recently attracted increasing attention. In this work, we propose a self-supervised monocular depth estimation network via binocular geometric correlation learning. Specifically, considering the inter-view geometric correlation, a binocular cue prediction module is presented to generate the auxiliary vision cue for the self-supervised learning of monocular depth estimation. Then, to deal with the occlusion in depth estimation, an occlusion interference attenuated constraint is developed to guide the supervision of the network by inferring the occlusion region and producing paired occlusion masks. Experimental results on two popular benchmark datasets have demonstrated that the proposed network obtains competitive results compared to state-of-the-art self-supervised methods and achieves comparable results to some popular supervised methods. Bo Peng 0007, Jianjun Lei 0001, Bingzheng Liu, Haifeng Shen, Wanqing Li 0001, Qingming Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Neural network model based on global and local features for multi-view mammogram classification
Jianpeng An, Chao Ma 0015, Hongjun Hou, Yanpeng Hou, Linyang Cui, Xuheng Jiang, Wanqing Li 0001, Zhongke Gao |
Neurocomputing | 8 |
| 2023 | Motion saliency based hierarchical attention network for action recognition
Zihui Guo, Yonghong Hou, Renyi Xiao, Chuankun Li, Wanqing Li 0001 |
Multim. Tools Appl. | 5 |
| 2023 | Sign language recognition via dimensional global-local shift and cross-scale aggregation
Zihui Guo, Yonghong Hou, Wanqing Li 0001 |
Neural Comput. Appl. | 3 |
| 2023 | FT-HID: a large-scale RGB-D dataset for first- and third-person human interaction analysis
Zihui Guo, Yonghong Hou, Pichao Wang, Zhimin Gao, Mingliang Xu 0001, Wanqing Li 0001 |
Neural Comput. Appl. | 6 |
| 2023 | Improved Shift Graph Convolutional Network for Action Recognition With SkeletonabstractShift graph convolutional network (Shift-GCN) achieves remarkable performance for skeleton based action recognition with lower computational complexity than other GCN based methods. However, the current Shift-GCN, with one spatial shift, a static mask and a local temporal convolution, cannot fully explore the spatial-temporal features among skeleton joints of different frames. In order to address these problems, an improved shift graph convolutional network (Ishift-GCN) is proposed in this letter. The Ishift-GCN consists of two parts including a bidirectional spatial shift graph convolution with a dynamic mask, and a multi-scale temporal shift graph convolution. The bidirectional spatial shift graph convolution exploits more spatial information among joints, and the dynamic mask with stronger generalization ability can learn different correlations among features of different joints for different actions. The multi-scale temporal shift graph convolution captures more temporal information by complementing the shifted features with multi-scale convolution. Furthermore, knowledge distillation is used to reduce computational complexity. Compared with Shift-GCN, the proposed Ishift-GCN achieves better results with less computation complexity on two widely used benchmarks, namely the NTU-RGB+D and UAV-Human dataset. Chuankun Li, Shuai Li 0005, Yanbo Gao, Wanqing Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Region-Aware Hierarchical Latent Feature Representation Learning-Guided Clustering for Hyperspectral Band SelectionabstractHyperspectral band selection aims to identify an optimal subset of bands for hyperspectral images (HSIs). For most existing clustering-based band selection methods, they directly stretch each band into a single feature vector and employ the pixelwise features to address band redundancy. In this way, they do not take full consideration of the spatial information and deal with the importance of different regions in HSIs, which leads to a nonoptimal selection. To address these issues, a region-aware hierarchical latent feature representation learning-guided clustering (HLFC) method is proposed. Specifically, in order to fully preserve the spatial information of HSIs, the superpixel segmentation algorithm is adopted to segment HSIs into multiple regions first. For each segmented region, the similarity graph is constructed to reflect the bands-wise similarity, and its corresponding Laplacian matrix is generated for learning low-dimensional latent features in a hierarchical way. All latent features are then fused to form a unified feature representation of HSIs. Finally, k -means clustering is utilized on the unified feature representation matrix to generate multiple clusters from which the band with maximum information entropy is selected to form the final subset of bands. Extensive experimental results demonstrate that the proposed clustering method can achieve superior performance than the state-of-the-art representative methods on the band selection. The demo code of this work is publicly available at https://github.com/WangJun2023/HLFC. Jun Wang 0118, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, Wanqing Li 0001, Xinzhong Zhu, Lizhe Wang 0001, Albert Y. Zomaya |
IEEE Trans. Cybern. | 5 |
| 2023 | Reducing Background Induced Domain Shift for Adaptive Person Re-IdentificationabstractCross-domain person re-identification (Re-ID) is a challenging and important task in monitoring safety and procedure compliance of industrial work places. In this article, a novel method is proposed to reduce background induced domain shift for adaptive person Re-ID. Specifically, a foreground-background joint clustering module is proposed to extract discriminative foreground and background features and an attention-based feature disentanglement module is designed to reduce the interference of background with the extraction of discriminative foreground features. Experimental results on three widely used person Re-ID benchmarking datasets (Market-1501, DukeMTMC-reID, and MSMT17) have demonstrated that the proposed method achieves promising performance compared with the state-of-the-art methods. Jianjun Lei 0001, Tianyi Qin, Bo Peng 0007, Wanqing Li 0001, Zhaoqing Pan, Haifeng Shen, Sam Kwong |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Novel View Synthesis from a Single Unposed Image via Unsupervised LearningabstractNovel view synthesis aims to generate novel views from one or more given source views. Although existing methods have achieved promising performance, they usually require paired views with different poses to learn a pixel transformation. This article proposes an unsupervised network to learn such a pixel transformation from a single source image. In particular, the network consists of a token transformation module that facilities the transformation of the features extracted from a source image into an intrinsic representation with respect to a pre-defined reference pose and a view generation module that synthesizes an arbitrary view from the representation. The learned transformation allows us to synthesize a novel view from any single source image of an unknown pose. Experiments on the widely used view synthesis datasets have demonstrated that the proposed network is able to produce comparable results to the state-of-the-art methods despite the fact that learning is unsupervised and only a single source image is required for generating a novel view. The code will be available upon the acceptance of the article. Bingzheng Liu, Jianjun Lei 0001, Bo Peng 0007, Chuanbo Yu, Wanqing Li 0001, Nam Ling |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Modeling Long-range Dependencies and Epipolar Geometry for Multi-view StereoabstractThis article proposes a network, referred to as Multi-View Stereo TRansformer (MVSTR) for depth estimation from multi-view images. By modeling long-range dependencies and epipolar geometry, the proposed MVSTR is capable of extracting dense features with global context and 3D consistency, which are crucial for reliable matching in multi-view stereo (MVS). Specifically, to tackle the problem of the limited receptive field of existing CNN-based MVS methods, a global-context Transformer module is designed to establish intra-view long-range dependencies so that global contextual features of each view are obtained. In addition, to further enable features of each view to be 3D consistent, a 3D-consistency Transformer module with an epipolar feature sampler is built, where epipolar geometry is modeled to effectively facilitate cross-view interaction. Experimental results show that the proposed MVSTR achieves the best overall performance on the DTU dataset and demonstrates strong generalization on the Tanks & Temples benchmark dataset. Bo Peng 0007, Wanqing Li 0001, Haifeng Shen, Qingming Huang, Jianjun Lei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Learning Using Privileged Information for Zero-Shot Action Recognition
Zhiyi Gao, Yonghong Hou, Wanqing Li 0001, Zihui Guo |
ACCV (4) | 3 |
| 2022 | Focal and Global Spatial-Temporal Transformer for Skeleton-Based Action Recognition
Zhimin Gao, Peitao Wang, Pei Lv, Xiaoheng Jiang, Qidong Liu 0001, Pichao Wang, Mingliang Xu 0001, Wanqing Li 0001 |
ACCV (4) | 8 |
| 2022 | Deep Stereo Image Compression via Bi-directional CodingabstractExisting learning-based stereo compression methods usually adopt a unidirectional approach to encoding one image independently and the other image conditioned upon the first. This paper proposes a novel bidirectional coding-based end-to-end stereo image compression network (BCSIC-Net). BCSIC-Net consists of a novel bidirectional contextual transform module which performs nonlinear transform conditioned upon the inter-view context in a latent space to reduce inter-view redundancy, and a bidirectional conditional entropy model that employs interview correspondence as a conditional prior to improve coding efficiency. Experimental results on the InStereo2K and KITTI datasets demonstrate that the proposed BCSIC-Net can effectively reduce the inter-view redundancy and out-performs state-of-the-art methods. Jianjun Lei 0001, Xiangrui Liu, Bo Peng 0007, Dengchao Jin, Wanqing Li 0001, Jingxiao Gu |
CVPR | 5 |
| 2022 | Contrastive Positive Mining for Unsupervised 3D Action Representation Learning
Yonghong Hou, Wanqing Li 0001 |
ECCV (4) | 4 |
| 2022 | Sparse fuzzy classification for profiling online users and relevant user-generated content
Jie Yang 0009, Brian Yecies, Jun Ma 0002, Wanqing Li 0001 |
Expert Syst. Appl. | 4 |
| 2022 | DeFusionNET: Defocus Blur Detection via Recurrently Fusing and Refining Discriminative Multi-Scale Deep FeaturesabstractAlbeit great success has been achieved in image defocus blur detection, there are still several unsolved challenges, e.g., interference of background clutter, scale sensitivity and missing boundary details of blur regions. To deal with these issues, we propose a deep neural network which recurrently fuses and refines multi-scale deep features (DeFusionNet) for defocus blur detection. We first fuse the features from different layers of FCN as shallow features and semantic features, respectively. Then, the fused shallow features are propagated to deep layers for refining the details of detected defocus blur regions, and the fused semantic features are propagated to shallow layers to assist in better locating blur regions. The fusion and refinement are carried out recurrently. In order to narrow the gap between low-level and high-level features, we embed a feature adaptation module before feature propagating to exploit the complementary information as well as reduce the contradictory response of different feature layers. Since different feature channels are with different extents of discrimination for detecting blur regions, we design a channel attention module to select discriminative features for feature refinement. Finally, the output of each layer at last recurrent step are fused to obtain the final result. We collect a new dataset consists of various challenging images and their pixel-wise annotations for promoting further study. Extensive experiments on two commonly used datasets and our newly collected one are conducted to demonstrate both the efficacy and efficiency of DeFusionNet. Chang Tang, Xinwang Liu 0002, Wanqing Li 0001, Jian Xiong 0002, Lizhe Wang 0001, Albert Y. Zomaya, Antonella Longo |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Deep Affine Motion Compensation Network for Inter Prediction in VVCabstractIn video coding, it is a challenge to deal with scenes with complex motions, such as rotation and zooming. Although affine motion compensation (AMC) is employed in Versatile Video Coding (VVC), it is still difficult to handle non-translational motions due to the adopted hand-craft block-based motion compensation. In this paper, we propose a deep affine motion compensation network (DAMC-Net) for inter prediction in video coding to effectively improve the prediction accuracy. To the best of our knowledge, our work is the first attempt to deal with the deformable motion compensation based on CNN in VVC. Specifically, a deformable motion-compensated prediction (DMCP) module is proposed to compensate the current encoding block through a learnable way to estimate accurate motion fields. Meanwhile, the spatial neighboring information and the temporal reference block as well as the initial motion field are fully exploited. By effectively fusing the multi-channel feature maps from DMCP, an attention-based fusion and reconstruction (AFR) module is designed to reconstruct the output block. The proposed DAMC-Net is integrated into VVC and the experimental results demonstrate that the proposed method considerably enhances the coding performance. Dengchao Jin, Jianjun Lei 0001, Bo Peng 0007, Wanqing Li 0001, Nam Ling, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | A Central Difference Graph Convolutional Operator for Skeleton-Based Action RecognitionabstractThis paper proposes a new graph convolutional operator called central difference graph convolution (CDGC) for skeleton based action recognition. It is not only able to aggregate node information like a vanilla graph convolutional operation but also gradient information. Without introducing any additional parameters, CDGC can replace vanilla graph convolution in any existing Graph Convolutional Networks (GCNs). In addition, an accelerated version of the CDGC is developed which greatly improves the speed of training. Experiments on two popular large-scale datasets NTU RGB+D 60 & 120 have demonstrated the efficacy of the proposed CDGC. Code is available athttps://github.com/iesymiao/CD-GCN. Shuangyan Miao, Yonghong Hou, Zhimin Gao, Mingliang Xu 0001, Wanqing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | C2FNet: A Coarse-to-Fine Network for Multi-View 3D Point Cloud GenerationabstractGeneration of a 3D model of an object from multiple views has a wide range of applications. Different parts of an object would be accurately captured by a particular view or a subset of views in the case of multiple views. In this paper, a novel coarse-to-fine network (C2FNet) is proposed for 3D point cloud generation from multiple views. C2FNet generates subsets of 3D points that are best captured by individual views with the support of other views in a coarse-to-fine way, and then fuses these subsets of 3D points to a whole point cloud. It consists of a coarse generation module where coarse point clouds are constructed from multiple views by exploring the cross-view spatial relations, and a fine generation module where the coarse point cloud features are refined under the guidance of global consistency in appearance and context. Extensive experiments on the benchmark datasets have demonstrated that the proposed method outperforms the state-of-the-art methods. Jianjun Lei 0001, Bo Peng 0007, Wanqing Li 0001, Zhaoqing Pan, Qingming Huang |
IEEE Trans. Image Process. | 4 |
| 2022 | MIEGAN: Mobile Image Enhancement via a Multi-Module Cascade Neural NetworkabstractVisual quality of images captured by mobile devices is often inferior to that of images captured by a Digital Single Lens Reflex (DSLR) camera. This paper presents a novel generative adversarial network-based mobile image enhancement method, referred to as MIEGAN. It consists of a novel multi-module cascade generative network and a novel adaptive multi-scale discriminative network. The multi-module cascade generative network is built upon a two-stream encoder, a feature transformer, and a decoder. In the two-stream encoder, a luminance-regularizing stream is proposed to help the network focus on low-light areas. In the feature transformation module, two networks effectively capture both global and local information of an image. To further assist the generative network to generate the high visual quality images, a multi-scale discriminator is used instead of a regular single discriminator to distinguish whether an image is fake or real globally and locally. To balance the global and local discriminators, an adaptive weight allocation is proposed. In addition, a contrast loss is proposed, and a new mixed loss function is developed to improve the visual quality of the enhanced images. Extensive experiments on the popular DSLR photo enhancement dataset and MIT-FiveK dataset have verified the effectiveness of the proposed MIEGAN. Zhaoqing Pan, Jianjun Lei 0001, Wanqing Li 0001, Nam Ling, Sam Kwong |
IEEE Trans. Multim. | 4 |
| 2021 | Beyond Covariance: SICE and Kernel Based Visual Feature Representation
Jianjia Zhang, Lei Wang 0001, Luping Zhou, Wanqing Li 0001 |
Int. J. Comput. Vis. | 4 |
| 2021 | Transformer guided geometry model for flow-based unsupervised visual odometry
Xiangyu Li 0009, Yonghong Hou, Pichao Wang, Zhimin Gao, Mingliang Xu 0001, Wanqing Li 0001 |
Neural Comput. Appl. | 6 |
| 2020 | Semi-Dynamic Hypergraph Neural Network for 3D Pose EstimationabstractThis paper proposes a novel Semi-Dynamic Hypergraph Neural Network (SD-HNN) to estimate 3D human pose from a single image. SD-HNN adopts hypergraph to represent the human body to effectively exploit the kinematic constrains among adjacent and non-adjacent joints. Specifically, a pose hypergraph in SD-HNN has two components. One is a static hypergraph constructed according to the conventional tree body structure. The other is the semi-dynamic hypergraph representing the dynamic kinematic constrains among different joints. These two hypergraphs are combined together to be trained in an end-to-end fashion. Unlike traditional Graph Convolutional Networks (GCNs) that are based on a fixed tree structure, the SD-HNN can deal with ambiguity in human pose estimation. Experimental results demonstrate that the proposed method achieves state-of-the-art performance both on the Human3.6M and MPI-INF-3DHP datasets. Pei Lv, Junjin Cheng, Wanqing Li 0001, Bing Zhou 0003, Mingliang Xu 0001 |
IJCAI | 6 |
| 2020 | SAR-NAS: Skeleton-based action recognition via neural architecture searching
Yonghong Hou, Pichao Wang, Zihui Guo, Wanqing Li 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2020 | ConvNets-based action recognition from skeleton motion maps
Yanfang Chen, Chuankun Li, Yonghong Hou, Wanqing Li 0001 |
Multim. Tools Appl. | 5 |
| 2020 | A Review of Dynamic Maps for 3D Human Motion Recognition Using ConvNets and Its Improvement
Zhimin Gao, Pichao Wang, Huogen Wang, Mingliang Xu 0001, Wanqing Li 0001 |
Neural Process. Lett. | 5 |
| 2020 | Jointly Learning Visual Poses and Pose Lexicon for Semantic Action RecognitionabstractA novel method for semantic action recognition through learning a pose lexicon is presented in this paper. A pose lexicon comprises a set of semantic poses, a set of visual poses, and a probabilistic mapping between the visual and semantic poses. This paper assumes that both the visual poses and mapping are hidden and proposes a method to simultaneously learn a visual pose model that estimates the likelihood of an observed video frame being generated from hidden visual poses, and a pose lexicon model establishes the probabilistic mapping between the hidden visual poses and the semantic poses parsed from textual instructions. Specifically, the proposed method consists of two-level hidden Markov models. One level represents the alignment between the visual poses and semantic poses. The other level represents a visual pose sequence, and each visual pose is modeled as a Gaussian mixture. An expectation-maximization algorithm is developed to train a pose lexicon. With the learned lexicon, action classification is formulated as a problem of finding the maximum posterior probability of a given sequence of video frames that follows a given sequence of semantic poses, constrained by the most likely visual pose and the alignment sequences. The proposed method was evaluated on MSRC-12, WorkoutSU-10, WorkoutUOW-18, Combined-15, Combined-17, and Combined-50 action datasets using cross-subject, cross-dataset, zero-shot, and seen/unseen protocols. Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Self-Attention Guided Deep Features for Action RecognitionabstractSkeleton based human action recognition is an important task in computer vision. However, it is very challenging due to the complex spatio-temporal variations of skeleton joints. In this work, we propose an end-to-end trainable network consisting of a Deep Convolutional Model (DCM) and a Self-Attention Model (SAM) for human action recognition from skeleton data. Specifically, skeleton sequences are encoded into color images and fed into DCM to extract deep features. In the SAM, handcrafted features representing the motion degree of joints are extracted and the attention weights are learned by a simple yet effective linear mapping. The effectiveness of proposed method has been verified on NTU RGB+D, SYSU-3D and UTD-MHAD datasets and achieved state-of-the-art results. Renyi Xiao, Yonghong Hou, Zihui Guo, Chuankun Li, Pichao Wang, Wanqing Li 0001 |
ICME | 6 |
| 2019 | Light Weight Stereo Matching via Deep Extraction and Integration of Low and High Level InformationabstractDeep convolutional neural networks (CNN) have demonstrated remarkable progress in stereo matching recently. However, disparity estimation in the ill-posed regions is still difficult. In addition, CNN based stereo matching methods often have impractical computational complexity and memory consumption. To address these problems we propose an end-to-end light weight CNN architecture to effectively learn and integrate low and high level information. To achieve this, a novel enhancement block built upon group convolution and dilated-convolution is proposed. Compared with state-of-the-art methods, the proposed method achieved competitive performance with the least number of network parameters on the Flyingthings3d and KITTI datasets. Yonghong Hou, Pichao Wang, Zhongyu Jiang, Wanqing Li 0001 |
ICME | 5 |
| 2019 | DVONet: Unsupervised Monocular Depth Estimation and Visual OdometryabstractThis paper proposes an unsupervised learning framework for monocular depth estimation and visual odometry (VO), referred to as DVONet. The framework is trained using stereo image sequences and is able to estimate absolute-scale scene depth and camera poses from monocular images. To mitigate the effect of stereo occlusions in training and improve the depth estimation, left-right occlusion mask is introduced. In addition, a novel VO network is proposed where the feature extraction network is shared between pose estimation and optical flow estimation. The proposed DVONet achieves state-of-the-art results for both depth estimation and VO tasks on the KITTI driving dataset, outperforming the existing unsupervised methods and being comparable to the traditional ones. Xiangyu Li 0009, Yonghong Hou, Pichao Wang, Wanqing Li 0001 |
VCIP | 5 |
| 2019 | A fully trainable network with RNN-based pooling
Shuai Li 0005, Wanqing Li 0001, Chris Cook, Ce Zhu, Yanbo Gao |
Neurocomputing | 2 |
| 2019 | Learning attentive dynamic maps (ADMs) for Understanding Human Actions
Chuankun Li, Yonghong Hou, Wanqing Li 0001, Pichao Wang |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Unsupervised domain adaptation: A multi-task learning-based method
Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona |
Knowl. Based Syst. | 2 |
| 2019 | A real-time webcam-based method for assessing upper-body postures
Zewei Ding, Wanqing Li 0001, Philip Ogunbona |
Mach. Vis. Appl. | 2 |
| 2019 | Multiview-Based 3-D Action Recognition Using Deep NetworksabstractIn multiview learning, views may be obtained from multiple sources or extracted from a single source as different features. In this paper, effective multiple views from skeleton sequences are proposed to learn the discriminative features using multiple networks for three-dimensional human action recognition. Specifically, three views are constructed in the spatial domain and fed to a stack of long short-term memory networks to exploit temporal information and three views are constructed using the improved joint trajectory maps and fed to three convolutional neural networks to exploit spatial information. Multiply fusion is used to combine the recognition scores of all views. The proposed method has been verified and achieved the state-of-the-art results on the widely used UTD-MHAD, MSRC-12 Kinect Gesture, and NTU red, green, blue (RGB)+D datasets. Chuankun Li, Yonghong Hou, Pichao Wang, Wanqing Li 0001 |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2018 | Cooperative Training of Deep Aggregation Networks for RGB-D Action RecognitionabstractA novel deep neural network training paradigm that exploits the conjoint information in multiple heterogeneous sources is proposed. Specifically, in a RGB-D based action recognition task, it cooperatively trains a single convolutional neural network (named c-ConvNet) on both RGB visual features and depth features, and deeply aggregates the two kinds of features for action recognition. Differently from the conventional ConvNet that learns the deep separable features for homogeneous modality-based classification with only one softmax loss function, the c-ConvNet enhances the discriminative power of the deeply learned features and weakens the undesired modality discrepancy by jointly optimizing a ranking loss and a softmax loss for both homogeneous and heterogeneous modalities. The ranking loss consists of intra-modality and cross-modality triplet losses, and it reduces both the intra-modality and cross-modality feature variations. Furthermore, the correlations between RGB and depth data are embedded in the c-ConvNet, and can be retrieved by either of the modalities and contribute to the recognition in the case even only one of the modalities is available. The proposed method was extensively evaluated on two large RGB-D action recognition datasets, ChaLearn LAP IsoGD and NTU RGB+D datasets, and one small dataset, SYSU 3D HOI, and achieved state-of-the-art results. Pichao Wang, Wanqing Li 0001, Jun Wan 0001, Philip Ogunbona, Xinwang Liu 0002 |
AAAI | 2 |
| 2018 | Independently Recurrent Neural Network (IndRNN): Building a Longer and Deeper RNNabstractRecurrent neural networks (RNNs) have been widely used for processing sequential data. However, RNNs are commonly difficult to train due to the well-known gradient vanishing and exploding problems and hard to learn long-term patterns. Long short-term memory (LSTM) and gated recurrent unit (GRU) were developed to address these problems, but the use of hyperbolic tangent and the sigmoid action functions results in gradient decay over layers. Consequently, construction of an efficiently trainable deep network is challenging. In addition, all the neurons in an RNN layer are entangled together and their behaviour is hard to interpret. To address these problems, a new type of RNN, referred to as independently recurrent neural network (IndRNN), is proposed in this paper, where neurons in the same layer are independent of each other and they are connected across layers. We have shown that an IndRNN can be easily regulated to prevent the gradient exploding and vanishing problems while allowing the network to learn long-term dependencies. Moreover, an IndRNN can work with non-saturated activation functions such as relu (rectified linear unit) and be still trained robustly. Multiple IndRNNs can be stacked to construct a network that is deeper than the existing RNNs. Experimental results have shown that the proposed IndRNN is able to process very long sequences (over 5000 time steps), can be used to construct very deep networks (21 layers used in the experiment) and still be trained robustly. Better performances have been achieved on various tasks by using IndRNNs compared with the traditional RNN and LSTM. Shuai Li 0005, Wanqing Li 0001, Chris Cook, Ce Zhu, Yanbo Gao |
CVPR | 2 |
| 2018 | Importance Weighted Adversarial Nets for Partial Domain AdaptationabstractThis paper proposes an importance weighted adversarial nets-based method for unsupervised domain adaptation, specific for partial domain adaptation where the target domain has less number of classes compared to the source domain. Previous domain adaptation methods generally assume the identical label spaces, such that reducing the distribution divergence leads to feasible knowledge transfer. However, such an assumption is no longer valid in a more realistic scenario that requires adaptation from a larger and more diverse source domain to a smaller target domain with less number of classes. This paper extends the adversarial nets-based domain adaptation and proposes a novel adversarial nets-based partial domain adaptation method to identify the source samples that are potentially from the outlier classes and, at the same time, reduce the shift of shared classes between domains. Jing Zhang 0017, Zewei Ding, Wanqing Li 0001, Philip Ogunbona |
CVPR | 3 |
| 2018 | RGB-D-based human motion recognition with deep learning: A survey
Pichao Wang, Wanqing Li 0001, Philip Ogunbona, Jun Wan 0001, Sergio Escalera |
Comput. Vis. Image Underst. | 2 |
| 2018 | Online human action recognition based on incremental learning of weighted covariance descriptors
Chang Tang, Wanqing Li 0001, Pichao Wang, Lizhe Wang 0001 |
Inf. Sci. | 2 |
| 2018 | Robust unsupervised feature selection via dual self-representation and manifold regularization
Chang Tang, Xinwang Liu 0002, Miaomiao Li 0001, Pichao Wang, Jiajia Chen 0010, Lizhe Wang 0001, Wanqing Li 0001 |
Knowl. Based Syst. | 7 |
| 2018 | Action recognition based on joint trajectory maps with convolutional neural networks
Pichao Wang, Wanqing Li 0001, Chuankun Li, Yonghong Hou |
Knowl. Based Syst. | 2 |
| 2018 | Skeleton Optical Spectra-Based Action Recognition Using Convolutional Neural NetworksabstractThis letter presents an effective method to encode the spatiotemporal information of a skeleton sequence into color texture images, referred to as skeleton optical spectra, and employs convolutional neural networks (ConvNets) to learn the discriminative features for action recognition. Such spectrum representation makes it possible to use a standard ConvNet architecture to learn suitable “dynamic” features from skeleton sequences without training millions of parameters afresh and it is especially valuable when there is insufficient annotated training video data. Specifically, the encoding consists of four steps: mapping of joint distribution, spectrum coding of joint trajectories, spectrum coding of body parts, and joint velocity weighted saturation and brightness. Experimental results on three widely used datasets have demonstrated the efficacy of the proposed method. Yonghong Hou, Pichao Wang, Wanqing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | A Fusion Framework for Camouflaged Moving Foreground Detection in the Wavelet DomainabstractDetecting camouflaged moving foreground objects has been known to be difficult due to the similarity between the foreground objects and the background. Conventional methods cannot distinguish the foreground from background due to the small differences between them and thus suffer from underdetection of the camouflaged foreground objects. In this paper, we present a fusion framework to address this problem in the wavelet domain. We first show that the small differences in the image domain can be highlighted in certain wavelet bands. Then the likelihood of each wavelet coefficient being foreground is estimated by formulating foreground and background models for each wavelet band. The proposed framework effectively aggregates the likelihoods from different wavelet bands based on the characteristics of the wavelet transform. Experimental results demonstrated that the proposed method significantly outperformed existing methods in detecting camouflaged foreground objects. Specifically, the average F-measure for the proposed algorithm was 0.87, compared to 0.71 to 0.8 for the other stateof- the-art methods. Shuai Li 0005, Dinei A. F. Florêncio, Wanqing Li 0001, Yaqin Zhao, Chris Cook |
IEEE Trans. Image Process. | 3 |
| 2018 | Detection and Separation of Smoke From Single Image FramesabstractThis paper proposes novel methods for detecting and separating smoke from a single image frame. Specifically, an image formation model is derived based on the atmospheric scattering models. The separation of a frame into quasi-smoke and quasi-background components is formulated as convex optimization that solves a sparse representation problem using dual dictionaries for the smoke and background components, respectively. A novel feature is constructed as a concatenation of the respective sparse coefficients for detection. In addition, a method based on the concept of image matting is developed to separate the true smoke and background components from the smoke detection results. Extensive experiments on detection were conducted and the results showed that the proposed feature significantly outperforms existing features for smoke detection. In particular, the proposed method is able to differentiate smoke from other challenging objects (e.g. fog/haze, cloud, and so on) with similar visual appearance in a gray-scale frame. Experiments on smoke separation also demonstrated that the proposed separation method can effectively estimate/separate the true smoke and background components. Hongda Tian, Wanqing Li 0001, Philip Ogunbona, Lei Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Depth Pooling Based Large-Scale 3-D Action Recognition With Convolutional Neural NetworksabstractThis paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as dynamic depth images (DDI), dynamic depth normal images (DDNI), and dynamic depth motion normal images (DDMNI), for both isolated and continuous action recognition. These dynamic images are constructed from a segmented sequence of depth maps using hierarchical bidirectional rank pooling to effectively capture the spatial-temporal information. Specifically, DDI exploits the dynamics of postures over time, and DDNI and DDMNI exploit the 3-D structural information captured by depth maps. Upon the proposed representations, a convolutional neural network (ConvNet)-based method is developed for action recognition. The image-based representations enable us to fine-tune the existing ConvNet models trained on image data without training a large number of parameters from scratch. The proposed method achieved the state-of-art results on three large datasets, namely, the large-scale continuous gesture recognition dataset (means the Jaccard index 0.4109), the large-scale isolated gesture recognition dataset (59.21%), and the NTU RGB+D dataset (87.08% cross-subject and 84.22% cross-view) even though only the depth modality was used. Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Philip Ogunbona |
IEEE Trans. Multim. | 2 |
| 2017 | Scene Flow to Action Map: A New Representation for RGB-D Based Action Recognition with Convolutional Neural NetworksabstractScene flow describes the motion of 3D objects in real world and potentially could be the basis of a good feature for 3D action recognition. However, its use for action recognition, especially in the context of convolutional neural networks (ConvNets), has not been previously studied. In this paper, we propose the extraction and use of scene flow for action recognition from RGB-D data. Previous works have considered the depth and RGB modalities as separate channels and extract features for later fusion. We take a different approach and consider the modalities as one entity, thus allowing feature extraction for action recognition at the beginning. Two key questions about the use of scene flow for action recognition are addressed: how to organize the scene flow vectors and how to represent the long term dynamics of videos based on scene flow. In order to calculate the scene flow correctly on the available datasets, we propose an effective self-calibration method to align the RGB and depth data spatially without knowledge of the camera parameters. Based on the scene flow vectors, we propose a new representation, namely, Scene Flow to Action Map (SFAM), that describes several long term spatio-temporal dynamics for action recognition. We adopt a channel transform kernel to transform the scene flow vectors to an optimal color space analogous to RGB. This transformation takes better advantage of the trained ConvNets models over ImageNet. Experimental results indicate that this new representation can surpass the performance of state-of-the-art methods on two large public datasets. Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Philip Ogunbona |
CVPR | 2 |
| 2017 | Joint Geometrical and Statistical Alignment for Visual Domain AdaptationabstractThis paper presents a novel unsupervised domain adaptation method for cross-domain visual recognition. We propose a unified framework that reduces the shift between domains both statistically and geometrically, referred to as Joint Geometrical and Statistical Alignment (JGSA). Specifically, we learn two coupled projections that project the source domain and target domain data into low-dimensional subspaces where the geometrical shift and distribution shift are reduced simultaneously. The objective function can be solved efficiently in a closed form. Extensive experiments have verified that the proposed method significantly outperforms several state-of-the-art domain adaptation methods on a synthetic dataset and three different real world cross-domain visual recognition tasks. Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona |
CVPR | 2 |
| 2017 | Foreground detection in camouflaged scenesabstractForeground detection has been widely studied for decades due to its importance in many practical applications. Most of the existing methods assume foreground and background show visually distinct characteristics and thus the foreground can be detected once a good background model is obtained. However, there are many situations where this is not the case. Of particular interest in video surveillance is the camouflage case. For example, an active attacker camouflages by intentionally wearing clothes that are visually similar to the background. In such cases, even given a decent background model, it is not trivial to detect foreground objects. This paper proposes a texture guided weighted voting (TGWV) method which can efficiently detect foreground objects in camouflaged scenes. The proposed method employs the stationary wavelet transform to decompose the image into frequency bands. We show that the small and hardly noticeable differences between foreground and background in the image domain can be effectively captured in certain wavelet frequency bands. To make the final foreground decision, a weighted voting scheme is developed based on intensity and texture of all the wavelet bands with weights carefully designed. Experimental results demonstrate that the proposed method achieves superior performance compared to the current state-of-the-art results. Shuai Li 0005, Dinei A. F. Florêncio, Yaqin Zhao, Chris Cook, Wanqing Li 0001 |
ICIP | 5 |
| 2017 | Weakly structured information aggregation for upper-body posture assessment using ConvNetsabstractPosture assessment aims to determine the risk associated with poor posture and thus avoid injury in subjects. Upper-body posture assessment from images offers an attractive alternative to manual methods by directly extracting relevant features for classification. A deep convolutional neural network is proposed to extract structured features from different body parts and learn shared features that are used to determine the appropriate assessment. The structured features are learned with triplet-based rank constraints based on head and torso separately. The shared feature and assessment function are learned with soft-max constraints based on posture risk measurements. Experimental evaluation on a self-collected upper-body posture dataset has verified the efficacy of the proposed method and network architecture. Zewei Ding, Wanqing Li 0001, Pichao Wang, Philip Ogunbona |
ICME | 2 |
| 2017 | Semantic action recognition by learning a pose lexicon
Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona, Zhengyou Zhang |
Pattern Recognit. | 2 |
| 2017 | Joint Distance Maps Based Action Recognition With Convolutional Neural NetworksabstractMotivated by the promising performance achieved by deep learning, an effective yet simple method is proposed to encode the spatio-temporal information of skeleton sequences into color texture images, referred to as joint distance maps (JDMs), and convolutional neural networks are employed to exploit the discriminative features from the JDMs for human action and interaction recognition. The pair-wise distances between joints over a sequence of single or multiple person skeletons are encoded into color variations to capture temporal information. The efficacy of the proposed method has been verified by the state-of-the-art results on the large RGB+D Dataset and small UTD-MHAD Dataset in both single-view and cross-view settings. Chuankun Li, Yonghong Hou, Pichao Wang, Wanqing Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2017 | Salient Object Detection via Weighted Low Rank Matrix RecoveryabstractImage-based salient object detection is a useful and important technique, which can promote the efficiency of several applications such as object detection, image classification/retrieval, object co-segmentation, and content-based image editing. In this letter, we present a novel weighted low-rank matrix recovery (WLRR) model for salient object detection. In order to facilitate efficient salient objects-background separation, a high-level background prior map is estimated by employing the property of the color, location, and boundary connectivity, and then this prior map is ensembled into a weighting matrix which indicates the likelihood that each image region belongs to the background. The final salient object detection task is formulated as the WLRR model with the weighting matrix. Both quantitative and qualitative experimental results on three challenging datasets show competitive results as compared with 24 state-of-the-art methods. Chang Tang, Pichao Wang, Changqing Zhang 0002, Wanqing Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2017 | Optimization of Camera Arrangement Using Correspondence Field to Improve Depth EstimationabstractStereo matching algorithms attempt to estimate depth from the images obtained by two cameras. In most cases, the arrangement of cameras (their locations and orientations with respect to the scene) is determined based on human experience. In this paper, it is shown that the camera arrangement can be optimized using the concept of correspondence field (CF) for better acquisition of depth. Specifically, this paper demonstrates the relationship between the CF of a pair of cameras and depth estimation accuracy and presents a method to optimize their arrangement based on the gradient of the CF. The experimental results show that a pair of cameras optimized by the proposed method can improve the accuracy of depth estimation by as much as 30% compared with the conventional camera arrangements. Shichao Fu, Farzad Safaei, Wanqing Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Learning structured dictionary based on inter-class similarity and representative marginsabstractWe consider the problem of learning a structured and discriminative dictionary based on sparse representation for classification task. The structure comprises class-shared and class-specific partitions which allows the separation of common and class-specific information in the data for classification. The resulting optimization problem was a max margin formulation that exploits the hinge loss function property. Comparative evaluation of the proposed classifier against four recent alternatives in a gender classification task indicates a 3-percenatge point improvement. Philip Ogunbona, Wanqing Li 0001, Gordon G. Wallace |
ICASSP | 3 |
| 2016 | Learning a pose lexicon for semantic action recognitionabstractThis paper presents a novel method for learning a pose lexicon comprising semantic poses defined by textual instructions and their associated visual poses defined by visual features. The proposed method simultaneously takes two input streams, semantic poses and visual pose candidates, and statistically learns a mapping between them to construct the lexicon. With the learned lexicon, action recognition can be cast as the problem of finding the maximum translation probability of a sequence of semantic poses given a stream of visual pose candidates. Experiments evaluating pre-trained and zero-shot action recognition conducted on MSRC-12 gesture and WorkoutSu-10 exercise datasets were used to verify the efficacy of the proposed method. Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona |
ICME | 2 |
| 2016 | Large-scale Isolated Gesture Recognition using Convolutional Neural NetworksabstractThis paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as Dynamic Depth Images (DDI), Dynamic Depth Normal Images (DDNI) and Dynamic Depth Motion Normal Images (DDMNI). These dynamic images are constructed from a sequence of depth maps using bidirectional rank pooling to effectively capture the spatial-temporal information. Such image-based representations enable us to fine-tune the existing ConvNets models trained on image data for classification of depth sequences, without introducing large parameters to learn. Upon the proposed representations, a convolutional Neural networks (ConvNets) based method is developed for gesture recognition and evaluated on the Large-scale Isolated Gesture Recognition at the ChaLearn Looking at People (LAP) challenge 2016. The method achieved 55.57% classification accuracy and ranked 2ndplace in this challenge but was very close to the best performance even though we only used depth data. Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Philip Ogunbona |
ICPR | 2 |
| 2016 | Large-scale Continuous Gesture Recognition Using Convolutional Neural NetworksabstractThis paper addresses the problem of continuous gesture recognition from sequences of depth maps using Convolutional Neural networks (ConvNets). The proposed method first segments individual gestures from a depth sequence based on quantity of movement (QOM). For each segmented gesture, an Improved Depth Motion Map (IDMM), which converts the depth sequence into one image, is constructed and fed to a ConvNet for recognition. The IDMM effectively encodes both spatial and temporal information and allows the fine-tuning with existing ConvNet models for classification without introducing millions of parameters to learn. The proposed method is evaluated on the Large-scale Continuous Gesture Recognition of the ChaLearn Looking at People (LAP) challenge 2016. It achieved the performance of 0.2655 (Mean Jaccard Index) and ranked 3rdplace in this challenge. Pichao Wang, Wanqing Li 0001, Zhimin Gao, Philip Ogunbona |
ICPR | 2 |
| 2016 | Action Recognition Based on Joint Trajectory Maps Using Convolutional Neural NetworksabstractRecently, Convolutional Neural Networks (ConvNets) have shown promising performances in many computer vision tasks, especially image-based recognition. How to effectively use ConvNets for video-based recognition is still an open problem. In this paper, we propose a compact, effective yet simple method to encode spatio-temporal information carried in 3D skeleton sequences into multiple 2D images, referred to as Joint Trajectory Maps (JTM), and ConvNets are adopted to exploit the discriminative features for real-time human action recognition. The proposed method has been evaluated on three public benchmarks, i.e., MSRC-12 Kinect gesture dataset (MSRC-12), G3D dataset and UTD multimodal human action dataset (UTD-MHAD) and achieved the state-of-the-art results. Pichao Wang, Yonghong Hou, Wanqing Li 0001 |
ACM Multimedia | 4 |
| 2016 | Guest Editorial: Human Activity Understanding from 2D and 3D Data
Junsong Yuan 0001, Wanqing Li 0001, Zhengyou Zhang, David J. Fleet, Jamie Shotton |
Int. J. Comput. Vis. | 2 |
| 2016 | Efficient 2D viewpoint combination for human action recognition
Behrouz Saghafi, Deepu Rajan, Wanqing Li 0001 |
Pattern Anal. Appl. | 3 |
| 2016 | Human detection from images and videos: A survey
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona |
Pattern Recognit. | 2 |
| 2016 | RGB-D-based action recognition datasets: A survey
Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona, Pichao Wang, Chang Tang |
Pattern Recognit. | 2 |
| 2016 | A Spectral and Spatial Approach of Coarse-to-Fine Blurred Image Region DetectionabstractBlur exists in many digital images, it can be mainly categorized into two classes: defocus blur which is caused by optical imaging systems and motion blur which is caused by the relative motion between camera and scene objects. In this letter, we propose a simple yet effective automatic blurred image region detection method. Based on the observation that blur attenuates high-frequency components of an image, we present a blur metric based on the log averaged spectrum residual to get a coarse blur map. Then, a novel iterative updating mechanism is proposed to refine the blur map from coarse to fine by exploiting the intrinsic relevance of similar neighbor image regions. The proposed iterative updating mechanism can partially resolve the problem of differentiating an in-focus smooth region and a blurred smooth region. In addition, our iterative updating mechanism can be integrated into other image blurred region detection algorithms to refine the final results. Both quantitative and qualitative experimental results demonstrate that our proposed method is more reliable and efficient compared to various state-of-the-art methods. Chang Tang, Yonghong Hou, Pichao Wang, Wanqing Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2016 | Action Recognition From Depth Maps Using Deep Convolutional Neural NetworksabstractThis paper proposes a new method, i.e., weighted hierarchical depth motion maps (WHDMM) + three-channel deep convolutional neural networks (3ConvNets), for human action recognition from depth maps on small training datasets. Three strategies are developed to leverage the capability of ConvNets in mining discriminative features for recognition. First, different viewpoints are mimicked by rotating the 3-D points of the captured depth maps. This not only synthesizes more data, but also makes the trained ConvNets view-tolerant. Second, WHDMMs at several temporal scales are constructed to encode the spatiotemporal motion patterns of actions into 2-D spatial structures. The 2-D spatial structures are further enhanced for recognition by converting the WHDMMs into pseudocolor images. Finally, the three ConvNets are initialized with the models obtained from ImageNet and fine-tuned independently on the color-coded WHDMMs constructed in three orthogonal planes. The proposed algorithm was evaluated on the MSRAction3D, MSRAction3DExt, UTKinect-Action, and MSRDailyActivity3D datasets using cross-subject protocols. In addition, the method was evaluated on the large dataset constructed from the above datasets. The proposed method achieved 2-9% better results on most of the individual datasets. Furthermore, the proposed method maintained its performance on the large dataset, whereas the performance of existing methods decreased with the increased number of actions. Pichao Wang, Wanqing Li 0001, Zhimin Gao, Jing Zhang 0017, Chang Tang, Philip Ogunbona |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2016 | Learning Discriminative Stein Kernel for SPD Matrices and Its ApplicationsabstractStein kernel (SK) has recently shown promising performance on classifying images represented by symmetric positive definite (SPD) matrices. It evaluates the similarity between two SPD matrices through their eigenvalues. In this paper, we argue that directly using the original eigenvalues may be problematic because: 1) eigenvalue estimation becomes biased when the number of samples is inadequate, which may lead to unreliable kernel evaluation, and 2) more importantly, eigenvalues reflect only the property of an individual SPD matrix. They are not necessarily optimal for computing SK when the goal is to discriminate different classes of SPD matrices. To address the two issues, we propose a discriminative SK (DSK), in which an extra parameter vector is defined to adjust the eigenvalues of input SPD matrices. The optimal parameter values are sought by optimizing a proxy of classification performance. To show the generality of the proposed method, three kernel learning criteria that are commonly used in the literature are employed as a proxy. A comprehensive experimental study is conducted on a variety of image classification tasks to compare the proposed DSK with the original SK and other methods for evaluating the similarity between SPD matrices. The results demonstrate that the DSK can attain greater discrimination and better align with classification tasks by altering the eigenvalues. This makes it produce higher classification performance than the original SK and other commonly used methods. Jianjia Zhang, Lei Wang 0001, Luping Zhou, Wanqing Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Beyond Covariance: Feature Representation with Nonlinear Kernel MatricesabstractCovariance matrix has recently received increasing attention in computer vision by leveraging Riemannian geometry of symmetric positive-definite (SPD) matrices. Originally proposed as a region descriptor, it has now been used as a generic representation in various recognition tasks. However, covariance matrix has shortcomings such as being prone to be singular, limited capability in modeling complicated feature relationship, and having a fixed form of representation. This paper argues that more appropriate SPD-matrix-based representations shall be explored to achieve better recognition. It proposes an open framework to use the kernel matrix over feature dimensions as a generic representation and discusses its properties and advantages. The proposed framework significantly elevates covariance representation to the unlimited opportunities provided by this new representation. Experimental study shows that this representation consistently outperforms its covariance counterpart on various visual recognition tasks. In particular, it achieves significant improvement on skeleton-based human action recognition, demonstrating the state-of-the-art performance over both the covariance and the existing non-covariance representations. Lei Wang 0001, Jianjia Zhang, Luping Zhou, Chang Tang, Wanqing Li 0001 |
ICCV | 5 |
| 2015 | Optimization of the number of rays in interpolation for light field based free viewpoint systemsabstractLight field (LF) rendering is widely used in free viewpoint video systems (FVV). Different methods have been proposed to employ depth maps to improve the rendering quality. However, estimation of depth is often error-prone. In this paper, a new method based on the concept of effective sampling density (ESD) is proposed for evaluating the depth-based LF rendering algorithms at different levels of errors in the depth estimation. In addition, for a given rendering quality, we provide an estimation of number of rays required in the interpolation algorithm to compensate for the adverse effect caused by errors in depth maps. The proposed method is particularly useful in designing a rendering algorithm with inaccurate knowledge of depth to achieve the required rendering quality. Both the theoretical study and numerical simulations have verified the efficacy of the proposed method. Hooman Shidanshidi, Farzad Safaei, Wanqing Li 0001 |
ICME | 3 |
| 2015 | ConvNets-Based Action Recognition from Depth Maps through Virtual Cameras and PseudocoloringabstractIn this paper, we propose to adopt ConvNets to recognize human actions from depth maps on relatively small datasets based on Depth Motion Maps (DMMs). In particular, three strategies are developed to effectively leverage the capability of ConvNets in mining discriminative features for recognition. Firstly, different viewpoints are mimicked by rotating virtual cameras around subject represented by the 3D points of the captured depth maps. This not only synthesizes more data from the captured ones, but also makes the trained ConvNets view-tolerant. Secondly, DMMs are constructed and further enhanced for recognition by encoding them into Pseudo-RGB images, turning the spatial-temporal motion patterns into textures and edges. Lastly, through transferring learning the models originally trained over ImageNet for image classification, the three ConvNets are trained independently on the color-coded DMMs constructed in three orthogonal planes. The proposed algorithm was extensively evaluated on MSRAction3D, MSRAction3DExt and UTKinect-Action datasets and achieved the state-of-the-art results on these datasets. Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Jing Zhang 0017, Philip Ogunbona |
ACM Multimedia | 2 |
| 2015 | Estimation of Signal Distortion Using Effective Sampling Density for Light Field-Based Free Viewpoint VideoabstractIn a light field-based free viewpoint video (LF-based FVV) system, effective sampling density (ESD) is defined as the number of rays per unit area of the scene that has been acquired and is selected in the rendering process for reconstructing an unknown ray. This paper extends the concept of ESD and shows that ESD is a tractable metric that quantifies the joint impact of the imperfections of LF acquisition and rendering. By deriving and analyzing ESD for the commonly used LF acquisition and rendering methods, it is shown that ESD is an effective indicator determined by system parameters and can be used to directly estimate output video distortion without access to the ground truth. This claim is verified by extensive numerical simulations and comparison to PSNR. Furthermore, an empirical relationship between the output distortion (in PSNR) and the calculated ESD is established to allow direct assessment of the overall video distortion without an actual implementation of the system. A small scale subjective user study is also conducted which indicates a correlation of 0.91 between ESD and perceived quality. Hooman Shidanshidi, Farzad Safaei, Wanqing Li 0001 |
IEEE Trans. Multim. | 3 |
| 2014 | Single Image Smoke Detection
Hongda Tian, Wanqing Li 0001, Philip Ogunbona, Lei Wang 0001 |
ACCV (2) | 2 |
| 2014 | Smoke Detection in Video: An Image Separation Approach
Hongda Tian, Wanqing Li 0001, Lei Wang 0001, Philip Ogunbona |
Int. J. Comput. Vis. | 2 |
| 2014 | Food image classification using local appearance and global structural information
Duc Thanh Nguyen, Zhimin Zong, Philip Ogunbona, Yasmine C. Probst, Wanqing Li 0001 |
Neurocomputing | 5 |
| 2014 | Introduction to the special issue on visual understanding and applications with RGB-D cameras
Zicheng Liu 0001, Michael Beetz, Daniel Cremers, Juergen Gall, Wanqing Li 0001, Dejan Pangercic, Jürgen Sturm, Yu-Wing Tai |
J. Vis. Commun. Image Represent. | 5 |
| 2013 | A Fast Approximate AIB Algorithm for Distributional Word ClusteringabstractDistributional word clustering merges the words having similar probability distributions to attain reliable parameter estimation, compact classification models and even better classification performance. Agglomerative Information Bottleneck (AIB) is one of the typical word clustering algorithms and has been applied to both traditional text classification and recent image recognition. Although enjoying theoretical elegance, AIB has one main issue on its computational efficiency, especially when clustering a large number of words. Different from existing solutions to this issue, we analyze the characteristics of its objective function-the loss of mutual information, and show that by merely using the ratio of word-class joint probabilities of each word, good candidate word pairs for merging can be easily identified. Based on this finding, we propose a fast approximate AIB algorithm and show that it can significantly improve the computational efficiency of AIB while well maintaining or even slightly increasing its classification performance. Experimental study on both text and image classification benchmark data sets shows that our algorithm can achieve more than 100 times speedup on large real data sets over the state-of-the-art method. Lei Wang 0001, Jianjia Zhang, Luping Zhou, Wanqing Li 0001 |
CVPR | 4 |
| 2013 | Non-uniform sampling of plenoptic signal based on the scene complexity variations for a free viewpoint video systemabstractAcquisition of a free viewpoint video (FVV) system is theoretically considered as a problem of plenoptic signal sampling. It is typically performed by using a regular camera grid. While a regular acquisition itself results in non-uniform sampling density, this non-uniformity does not match the scene complexity and frequency variations. This paper proposes an irregular acquisition method for optimum non-uniform plenoptic sampling corresponding to the variations of the scene complexity. Specifically, scene complexity is measured through analyzing DCT coefficients of reference images of the scene, describing the frequency behavior of the plenoptic signal over the scene space. An optimization model is formulated to calculate the optimum configurations of the acquisition cameras including positions and orientations. The theoretical analysis and numerical simulations demonstrate that the rendered video quality can be significantly improved (around 40% in mean PSNR) by employing the proposed irregular acquisition compared with the regular camera grid. Hooman Shidanshidi, Farzad Safaei, Arash Zamani-Farahani, Wanqing Li 0001 |
ICIP | 4 |
| 2013 | Scene-adaptive configuration of two cameras using the correspondence field functionabstractIn a free viewpoint video system, the scene is captured by a number of cameras and it would be desirable to optimize the configuration of cameras, such as their location or orientation, to improve the rendering quality. This paper introduces a mathematical representation of the multi-camera geometry, called the correspondence field (CF), which can be used to quantify the suitability of a camera configuration for a given arrangement of objects in the scene. The correspondence field describes the spatial topology of the intersecting rays of cameras, arranged as a number of layers or surfaces in the field of view of cameras. The paper derives the topology of CF for certain camera arrangements and analyzes the impact of changes in camera location or orientation on this topology. It demonstrates that CF can be used to find the optimum camera configuration for a given objective. It also presents simulation results of this method using our light field simulator. Farzad Safaei, Payam Mokhtarian, Hooman Shidanshidi, Wanqing Li 0001, M. Namazi-Rad, Amir Moosavie Nia |
ICME | 4 |
| 2013 | A method for calculating the minimum number of cameras in a light field based free viewpoint video systemabstractCalculation of the number of cameras required to capture the scene is an essential problem in a practical light field based free viewpoint video (FVV) system. Existing methods calculate the Nyquist rate by assuming a band-limited signal and perfect reconstruction of an arbitrary view using linear interpolation, which often results in an impractically high number of cameras. This paper proposes a new method based on the concept of effective sampling density (ESD). It is demonstrated that there is a trade-off between the depth information accuracy, the required number of cameras, and the desired rendering quality, which could be exploited to minimize the number of cameras for a given objective. Theoretical and numerical results show that the resulting number of cameras is significantly lower than what was reported in other studies with only a few percent reduction in the output PSNR. Moreover, it is shown that the methods proposed in those studies are special cases of the one presented in this paper. Hooman Shidanshidi, Farzad Safaei, Wanqing Li 0001 |
ICME | 3 |
| 2013 | Inter-occlusion reasoning for human detection based on variational mean field
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona |
Neurocomputing | 2 |
| 2013 | A novel shape-based non-redundant local binary pattern descriptor for object detection
Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001 |
Pattern Recognit. | 3 |
| 2013 | Measuring the degree of face familiarity based on extended NMFabstractGetting familiar with a face is an important cognitive process in human perception of faces, but little study has been reported on how to objectively measure the degree of familiarity. In this article, a method is proposed to quantitatively measure the familiarity of a face with respect to a set of reference faces that have been seen previously. The proposed method models the context-free and context-dependent forms of familiarity suggested by psychological studies and accounts for the key factors, namely exposure frequency, exposure intensity and similar exposure, that affect human perception of face familiarity. Specifically, the method divides the reference set into nonexclusive groups and measures the familiarity of a given face by aggregating the similarities of the face to the individual groups. In addition, the nonnegative matrix factorization (NMF) is extended in this paper to learn a compact and localized subspace representation for measuring the similarities of the face with respect to the individual groups. The proposed method has been evaluated through experiments that follow the protocols commonly used in psychological studies and has been compared with subjective evaluation. Results have shown that the proposed measurement is highly consistent with the subjective judgment of face familiarity. Moreover, a face recognition method is devised using the concept of face familiarity and the results on the standard FERET evaluation protocols have further verified the efficacy of the proposed familiarity measurement. Ce Zhan, Wanqing Li 0001, Philip Ogunbona |
ACM Trans. Appl. Percept. | 2 |
| 2012 | A Novel Video-Based Smoke Detection Method Using Image SeparationabstractIn the state-of-the-art video-based smoke detection methods, the representation of smoke mainly depends on the visual information in the current image frame. In the case of light smoke, the original background can be still seen and may deteriorate the characterization of smoke. The core idea of this paper is to demonstrate the superiority of using smoke component for smoke detection. In order to obtain smoke component, a blended image model is constructed, which basically is a linear combination of background and smoke components. Smoke opacity which represents a weighting of the smoke component is also defined. Based on this model, an optimization problem is posed. An algorithm is devised to solve for smoke opacity and smoke component, given an input image and the background. The resulting smoke opacity and smoke component are then used to perform the smoke detection task. The experimental results on both synthesized and real image data verify the effectiveness of the proposed method. Hongda Tian, Wanqing Li 0001, Lei Wang 0001, Philip Ogunbona |
ICME | 2 |
| 2012 | Measuring face familiarity and its application to face recognitionabstractThe familiarity of faces is one of the key factors that come into play during human face analysis. However, there is very little research that studies face familiarity. In this paper, two methods are proposed to quantitatively measure the degree of familiarity of a face with respect to a known set. The methods are in accordance with the psychological study. In particular, non-negative matrix factorization (NMF) is extended to learn a localized non-overlapping subspace representation of commonly experienced facial patterns from known faces. The familiarity of a given face is then measured based on its reconstruction error after being projected into the learned extended NMF subspaces. A subjective study involving 50 subjects indicates the proposed familiarity measurement is in line with human judgments. Furthermore, the familiarity vector generated during the measuring process is employed for face recognition. Experiments based on the standard FERET evaluation protocol demonstrates the efficacy of the familiarity based representation for face recognition. Ce Zhan, Wanqing Li 0001, Philip Ogunbona |
WACV | 2 |
| 2012 | Societally connected multimedia across culturesabstractThe advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu |
J. Zhejiang Univ. Sci. C | 8 |
| 2011 | Detecting humans under occlusion using variational mean field methodabstractThis paper proposes a human detection method using variational mean field approximation for occlusion reasoning. In the method, parts of human objects are detected individually using template matching. Initial detection hypotheses with spatial layout information are represented in a graphical model and refined through a Bayesian estimation. In this paper, mean field method is employed for such an estimation. The proposed method was evaluated on the popular CAVIAR-INRIA dataset. Experimental results show that the proposed algorithm is able to detect humans in severe occlusion within reasonable processing time. Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001 |
ICIP | 3 |
| 2011 | Human detection with contour-based local motion binary patternsabstractThis paper presents a human detection method using contour- based local motion features. The local motion is encoded using a variant of the popular Local Binary Pattern (LBP) called Non-Redundant Local Binary Pattern (NRLBP) descriptor computed on the difference image of two consecutive frames. In addition, the local motion features are extracted along the human's boundary contour. Localising features on the contours has the advantage of utilizing a precise human shape description. A motivation of the proposed method is that most of informative movements are performed on boundary contours of the body parts, e.g. legs of pedestrians. Evaluation of the proposed method was conducted on the INRIA and ETH datasets. Apart from showing the importance of motion information, experimental results also showed that localising features along the object boundary contours improves the detection performance. Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001 |
ICIP | 3 |
| 2011 | A quantitative approach for comparison and evaluation of light field rendering techniquesabstractLight field rendering (LFR) is an active research area in computer vision and computer graphics. LFR plays a crucial role in free viewpoint video systems (FVV). Although several rendering algorithms have been suggested for LFR but the lack of appropriate datasets with known ground truth has prevented a comparison and evaluation study of LFR algorithms. In most of the LFR papers the method is applied to several test cases for validation and as a result, just a subjective visualized output is given. To overcome this problem, this paper presents a quantitative approach for comparison and evaluation of LFR algorithms. The core of the proposed methodology is a simulation model and a 3D engine. The platform produces the reference images and ground truth data for a given 3D model. Subsequently, data are injected to a comparison engine to compare synthesized images from light field engine with original images from simulation, generating objective results for evaluation. The methodology is highly flexible and efficient to automatically generate LFR datasets and objectively compare and analyze any subset of LFR methods for any given experiment design scheme. Five key rendering algorithms are evaluated with proposed methodology to validate it. Overall, it is shown that the proposed quantitative methodology could be used for LFR objective evaluation and comparison. Hooman Shidanshidi, Farzad Safaei, Wanqing Li 0001 |
ICME | 3 |
| 2011 | Objective evaluation of light field rendering methods using effective sampling densityabstractLight field rendering (LFR) is an active research area in computer vision and computer graphics. LFR plays a crucial role in free viewpoint video systems (FVV). Several rendering algorithms have been suggested for LFR. However, comparative evaluation of these methods is often limited to subjective assessment of the output. To overcome this problem, this paper presents a geometric measurement, Effective Sampling Density of the scene, referred to as effective sampling for brevity, for objective comparison and evaluation of LFR algorithms. We have derived the effective sampling for the well-known LFR methods. Both theoretical study and numerical simulation have shown that the proposed effective sampling is an effective indicator of the performance for LFR methods. Hooman Shidanshidi, Farzad Safaei, Wanqing Li 0001 |
MMSP | 3 |
| 2011 | Smoke detection in videos using Non-Redundant Local Binary Pattern-based featuresabstractThis paper presents a novel and low complexity method for real-time video-based smoke detection. As a local texture operator, Non-Redundant Local Binary Pattern (NRLBP) is more discriminative and robust to illumination changes in comparison with original Local Binary Pattern (LBP), thus is employed to encode the appearance information of smoke. Non-Redundant Local Motion Binary Pattern (NRLMBP), which is computed on the difference image of consecutive frames, is introduced to capture the motion information of smoke. Experimental results show that NRLBP outperforms the original LBP in the smoke detection task. Furthermore, the combination of NRLBP and NRLMBP, which can be considered as a spatial-temporal descriptor of smoke, can lead to remarkable improvement on detection performance. Hongda Tian, Wanqing Li 0001, Philip Ogunbona, Duc Thanh Nguyen, Ce Zhan |
MMSP | 2 |
| 2011 | Age estimation based on extended non-negative matrix factorizationabstractPrevious studies suggested that local appearance-based methods are more efficient than geometric-based and holistic methods for age estimation. This is mainly due to the fact that age information are usually encoded by the local features such as wrinkles and skin texture on the forehead or at the eye corners. However, the variations of theses features caused by other factors such as identity, expression, pose and lighting may be larger than that caused by aging. Thus, one of the key challenges of age estimation lies in constructing a feature space that could successfully recovers age information while ignoring other sources of variations. In this paper, non-negative matrix factorization (NMF) is extended to learn a localized non-overlapping subspace representation for age estimation. To emphasize the appearance variation in aging, one individual extended NMF subspace is learned for each age or age group. The age or age group of a given face image is then estimated based on its reconstruction error after being projected into the learned age subspaces. Furthermore, a coarse to fine scheme is employed for exact age estimation, so that the age is estimated within the pre-classified age groups. Cross-database tests are conducted using FG-NET and MORPH databases to evaluate the proposed method. Experimental results have demonstrated the efficacy of the method. Ce Zhan, Wanqing Li 0001, Philip Ogunbona |
MMSP | 2 |
| 2010 | Human detection using local shape and Non-Redundant binary patternsabstractMotivated by the advantages of using shape matching technique in detecting objects in various postures and viewpoints and the discriminative power of local patterns in object recognition, this paper proposes a human detection method combining both shape and appearance cues. In particular, local shapes of the body parts are detected using template matching. Based on body parts' shapes, local appearance features are extracted. We introduce a novel local binary pattern (LBP) descriptor, called Non-Redundant LBP (NRLBP), to encode local appearance of human. The proposed method was evaluated and compared with other state-of-the-art human detection methods on two commonly used datasets: MIT and INRIA pedestrian test sets. We also performed extensive experiments on selecting appropriate parameters as well as verifying the improvement of the proposed method through all stages of the framework. Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona |
ICARCV | 2 |
| 2010 | Finding distinctive facial areas for face recognitionabstractOne of the key issues for local appearance based face recognition methods is that how to find the most discriminative facial areas. Most of the existing methods take the assumption that anatomical facial components, such as the eyes, nose, and mouth, are the most useful areas for recognition. Other more elaborate methods locate the most salient parts within the face according to a pre-specified criterion. In this paper, a novel method is proposed to identify the discriminative facial areas for face recognition. Unlike the existing methods that only analyze the given face, the proposed method identifies the distinctive areas of each individual's face by its comparison to the general population. In particular, non-negative matrix factorization (NMF) is extended to learn a localized non-overlapping subspace representation of the facial patterns from a generic face image database. In the learned subspace, the degree of distinctiveness for any facial area is measured depends on the probability of this area is belong to a general face. For evaluation, the proposed method is tested on exaggerated face images and applied in exiting face recognition systems. Experimental results demonstrate the efficiency of the proposed method. Ce Zhan, Wanqing Li 0001, Philip Ogunbona |
ICARCV | 2 |
| 2010 | Object detection using Non-Redundant Local Binary PatternsabstractLocal Binary Pattern (LBP) as a descriptor, has been successfully used in various object recognition tasks because of its discriminative property and computational simplicity. In this paper a variant of the LBP referred to as Non-Redundant Local Binary Pattern (NRLBP) is introduced and its application for object detection is demonstrated. Compared with the original LBP descriptor, the NRLBP has advantage of providing a more compact description of object's appearance. Furthermore, the NRLBP is more discriminative since it reflects the relative contrast between the background and foreground. The proposed descriptor is employed to encode human's appearance in a human detection task. Experimental results show that the NRLBP is robust and adaptive with changes of the background and foreground and also outperforms the original LBP in detection task. Duc Thanh Nguyen, Zhimin Zong, Philip Ogunbona, Wanqing Li 0001 |
ICIP | 4 |
| 2010 | Event Recognition Based on Top-Down Motion AttentionabstractHow to fuse static and dynamic information is a key issue in event analysis. In this paper, a top-down motion guided fusing method is proposed for recognizing events in an unconstrained news video. In the method, the static information is represented as a Bag-of-SIFT-features and motion information is employed to generate event specific attention map to direct the sampling of the interest points. We build class-specific motion histograms for each event so as to give more weight on the interest points that are discriminative to the corresponding event. Experimental results on TRECVID 2005 video corpus demonstrate that the proposed method can improve the mean average accuracy of recognition. Li Li 0010, Weiming Hu 0004, Bing Li 0001, Chunfeng Yuan, Pengfei Zhu 0001, Wanqing Li 0001 |
ICPR | 6 |
| 2010 | On the Combination of Local Texture and Global Structure for Food ClassificationabstractThis paper proposes a food image classification method using local textural patterns and their global structure to describe the food image. In this paper, a visual codebook of local textural patterns is created by employing Scale Invariant Feature Transformation (SIFT) interest point detector with the Local Binary Pattern (LBP) feature. In addition to describing the food image using local texture, the global structure of the food object is represented as the spatial distribution of the local textural structures and encoded using shape context. We evaluated the proposed method on the Pittsburgh Fast-Food Image (PFI) dataset. Experimental results showed that the proposed method could obtain better performance than the baseline experiment on the PFI dataset. Zhimin Zong, Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001 |
ISM | 4 |
| 2010 | Video Scene Segmentation Using Time Constraint Dominant-Set Clustering
Xianglin Zeng, Xiaoqin Zhang 0002, Weiming Hu 0004, Wanqing Li 0001 |
MMM | 4 |
| 2010 | Reconstruction From Limited-Angle Projections Based on delta-u Spectrum AnalysisabstractThis paper proposes a sparse representation of an image using discrete delta-u functions. A delta-u function is defined as the product of a Kronecker delta function and a step function. Based on the sparse representation, we have developed a novel and effective method for reconstructing an image from limited-angle projections. The method first estimates the parameters of the sparse representation from the incomplete projection data, and then directly calculates the image to be reconstructed. Experiments have shown that the proposed method can effectively recover the missing data and reconstruct images more accurately than the total-variation (TV) regularized reconstruction method. Jianhua Luo, Wanqing Li 0001, Yue Min Zhu |
IEEE Trans. Image Process. | 2 |
| 2009 | An Improved Template Matching Method for Object Detection
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona |
ACCV (3) | 2 |
| 2009 | Face detection using generalised integral image featuresabstractThis paper proposes generalised integral image features (GIIFs) for face detection. GIIFs provide a richer and more flexible set of features than Haar-like features. Due to the large set of possible GIIFs, a genetic algorithm is developed to select the feature space for the optimal weak classifiers. Experimental results have shown that this method is able to improve face detection accuracy. Alister Cordiner, Philip Ogunbona, Wanqing Li 0001 |
ICIP | 3 |
| 2009 | A novel template matching method for human detectionabstractThis paper proposes a novel weighted template matching method. It employs a generalized distance transform (GDT) and an orientation map (OM). The GDT allows us to weight the distance transform more on the strong edge points and the OM provides supplementary local orientation information for matching. Based on the matching method, a two-stage human detection method consisting of template matching and Bayesian verification is developed. Experimental results have shown that the proposed method can effectively reduce the false positive and false negative detection rates and perform superiorly in comparison to the conventional Chamfer matching method. Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona |
ICIP | 2 |
| 2009 | Human detection based on weighted template matchingabstractThis paper proposes a new two-stage human detection method involving matching and verification. A Bayesian framework is developed to verify the matching score obtained from a weighted distance measure. Performance evaluation indicates that the proposed method is able to utilize the flexible matching scheme and produce superior true positive, true negative and low misclassification rates. Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001 |
ICME | 3 |
| 2008 | Key-frame extraction using dominant-set clusteringabstractKey frames play an important role in video abstraction. Clustering is a popular approach for key-frame extraction. In this paper, we propose a novel method for key-frame extraction based on dominant-set clustering. Compared with the existing clustering-based methods, the proposed method dynamically decides the number of key frames depending on the complexity of video shots, produces key frames in a progressive manner and requires less computation. Experimental results on different types of video shots have verified the effectiveness of the method. Xianglin Zeng, Weiming Hu 0004, Wanqing Li 0001, Xiaoqin Zhang 0002 |
ICME | 3 |
| 2008 | Graphical modeling and decoding of human actionsabstractThis paper presents a graphical model for learning and recognizing human actions. Specifically, we propose to encode actions in a weighted directed graph, referred to as action graph, where nodes of the graph represent salient postures that are used to characterize the actions and shared by all actions. The weight between two nodes measures the transitional probability between the two postures. An action is encoded as one or multiple paths in the action graph. The salient postures are modeled using Gaussian Mixture Models (GMM). Both the salient postures and action graph are automatically learned from training samples through unsupervised clustering and expectation and maximization (EM) algorithm. Experimental results have verified the performance of the proposed model, its tolerance to noise and viewpoints and its robustness across different subjects and datasets. Wanqing Li 0001, Zhengyou Zhang, Zicheng Liu 0001 |
MMSP | 1 |
| 2008 | The HCM for perceptual image segmentation
Jonathan Randall, Ling Guan, Wanqing Li 0001 |
Neurocomputing | 3 |
| 2008 | An efficient iterative algorithm for image thresholding
Liju Dong, Ge Yu 0001, Philip Ogunbona, Wanqing Li 0001 |
Pattern Recognit. Lett. | 4 |
| 2008 | Expandable Data-Driven Graphical Modeling of Human Actions Based on Salient PosturesabstractThis paper presents a graphical model for learning and recognizing human actions. Specifically, we propose to encode actions in a weighted directed graph, referred to as action graph, where nodes of the graph represent salient postures that are used to characterize the actions and are shared by all actions. The weight between two nodes measures the transitional probability between the two postures represented by the two nodes. An action is encoded asoneor multiple paths in the action graph. The salient postures are modeled using Gaussian mixture models (GMMs). Both the salient postures and action graph are automatically learned from training samples through unsupervised clustering and expectation and maximization (EM) algorithm. The proposed action graph not only performs effective and robust recognition of actions, but it can also be expanded efficiently with new actions. An algorithm is also proposed for adding a new action to a trained action graph without compromising the existing action graph. Extensive experiments on widely used and challenging data sets have verified the performance of the proposed methods, its tolerance to noise and viewpoints, its robustness across different subjects and data sets, as well as the effectiveness of the algorithm for learning new actions. Wanqing Li 0001, Zhengyou Zhang, Zicheng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | A Maximum Likelihood Watermark Decoding SchemeabstractBased on the observation that an attack applied on a watermarked image, from a decoding point of view, modifies the distribution of the detection values away from the ideal distribution (without attack) for corresponding watermarking scheme, we propose a generic maximum likelihood decoding scheme by approximating the distribution with a finite Gaussian mixture model. The parameters of the model are estimated using expectation-maximization algorithm. The scheme allows the decoding to be automatically adapted to attacks that the watermarked images have undergone and, in consequence, to improve the decoding accuracy. Experiments on a QIM based watermarking system have clearly verified the significant improvement of the decoding accuracy achieved by the proposed maximum likelihood decoding in comparison to conventional threshold decoding. Wenming Lu, Wanqing Li 0001, Reihaneh Safavi-Naini, Philip Ogunbona |
ICME | 2 |
| 2006 | Feature Selection for Complex Patterns
Peter Schenkel, Wanqing Li 0001, Wanquan Liu |
ADMA | 2 |
| 2006 | A Prototype of Autonomous Intelligent Surveillance CamerasabstractThis paper presents an architecture and an FPGAbased prototype of an autonomous intelligent video surveillance camera. The camera takes the advantage of high resolution of CMOS image sensors and enables instantly automatic pan, tilt and zoom adjustment based upon motion activity. It performs automated scene analysis and provides immediate response to suspicious events by optimizing camera capturing parameters. The video output of the camera can be optimized to any region of interest while the camera continues to monitor the entire scene. Field trials of the prototyped camera have verified the proposed architecture. Wanqing Li 0001, Igor Kharitonenko, Serge Lichman, Chaminda Weerasinghe |
AVSS | 1 |
| 2006 | Indexing and Matching of Video Shots Based on Motion and Color AnalysisabstractThis paper concerns two fundamental issues in video shots retrieval: key frame identification and similarity measurement between the key frames. We propose a simple key frame extraction algorithm based on optical flow. The algorithm emphasizes the motion extremum in the shot. Color histograms in HSV color space are adopted to describe the content of the extracted key frames and a new model is proposed to measure the similarity between the key frames from different shots. Preliminary experiments have shown that the proposed method outperforms existing ones in retrieving sport video shots Ying Chen 0018, Weiming Hu 0004, Xianglin Zeng, Wanqing Li 0001 |
ICARCV | 4 |
| 2006 | A Pixel-Based Robust Imagewatermarking SystemabstractRobust image watermarking systems are required to be resistant to geometric attacks in addition to common image processing tasks, such as JPEG compression. However, robustness against geometric attacks, such as rotation, scaling and translation, still remains one of the most challenging research topics in image watermarking. We propose a new pixel-based watermarking system in which a binary logo is embedded, a bit per pixel, in the pixel domain of an image. The encoder of the proposed system is based on a sliding window embedding scheme that applies the local average quantization index modulation (QEM), to achieve geometric attack robustness. The decoder employs a maximum a posteriori (MAP) estimation supported by Markov random field (MRF) model to achieve robust decoding. Additionally, we demonstrate that the proposed scheme is also robust against possible watermark removal due to JPEG compression Wenming Lu, Wanqing Li 0001, Reihaneh Safavi-Naini, Philip Ogunbona |
ICME | 2 |
| 2005 | Simultaneous MAP estimation of inhomogeneity and segmentation of brain tissues from MR imagesabstractIntrascan and interscan intensity inhomogeneities have been identified as a common source of making many advanced segmentation techniques fail to produce satisfactory results in separating brains tissues from multi-spectral magnetic resonance (MR) images. A common solution is to correct the inhomogeneity before applying the segmentation techniques. This paper presents a method that is able to achieve simultaneous semi-supervised MAP (maximum a-posterior probability) estimation of the inhomogeneity field and segmentation of brain tissues, where the inhomogeneity is parameterized. Our method can incorporate any available incomplete training data and their contribution can be controlled in a flexible manner and therefore the segmentation of the brain tissues can be optimised. Experiments on both simulated and real MR images have demonstrated that the proposed method estimated the inhomogeneity field accurately and improved the segmentation. Wanqing Li 0001, Chris de Silver, Yianni Attikiouzel |
ICIP (2) | 1 |
| 2005 | A New Divide and Conquer Algorithm for Graph-based Image and Video SegmentationabstractThe concept of the shortest (or minimum) spanning tree (SST)and recursive SST (RSST) of an undirected weighted graph has been successfully applied in image segmentation and edge detection. This paper presents a divide-and-conquer approach for (R)SST based image segmentation in order to overcome the problem of high computational complexity associated with conventional graph algorithms. In the simplest form, the proposed approach, block-based RSST (BRSST), first divides the image into rectangular blocks, finds the (R)SST of each block individually using conventional graph algorithms and, then, merges the (R)SSTs of all image blocks to form an (R)SST of the entire image. Efficient merging algorithms are presented in this paper. We proved a theorem showing that the (R)SST obtained by the merging algorithms is one of the (R)SST that would be found by applying to the entire image the same algorithm used for finding the (R)SST of each image block. Theoretical analysis and experimental results have shown that BRSST has significantly reduced the computational cost. In addition, an incremental BRSST is proposed for video segmentation and results are presented Wanqing Li 0001, Mingren Shi, Philip Ogunbona |
MMSP | 1 |
| 2003 | Stereoscopic panoramic video generation using centro-circular projection techniqueabstractThe paper presents a method of stereoscopic panoramic video generation including techniques for panorama projection, stitching and calibration for various depth planes. The methods described can be used on video sequences captured by an arrangement of multiple pairs of cameras or multiple stereoscopic cameras mounted on a regular polygonal shaped camera rig. Algorithms can also be used in combination or separately, for generating both stereoscopic and monoscopic video and still panoramas. Chaminda Weerasinghe, Wanqing Li 0001, Philip Ogunbona |
ICASSP (3) | 2 |
| 2002 | Modelling of color cross-talk in CMOS image sensorsabstractThis paper presents a way to model the cross-talk effect in CMOS image sensors. Two algorithms are derived from the model; both of them work on the Bayer raw data and have low computational complexity. Experiments on Macbeth color chart and real images have shown the effectiveness of the modeling to eliminate the cross-talk effect and produce better quality images with traditional color interpolation and correction algorithms designed for CCD image sensors. Wanqing Li 0001, Philip Ogunbona, Igor Kharitonenko |
ICASSP | 1 |
| 2002 | Hierarchical cluster model for perceptual image processingabstractWe propose a method for extracting object symmetries from a digital image. To achieve this we examine the way in which the human visual system processes and organises visual information. Psychological evidence is combined with physiological processing. The evidence is based on image structure and the processing is based on the Hierarchical Cluster Model (HCM) which is used to model the human brain. Jonathan Randall, Ling Guan, Wanqing Li 0001 |
ICASSP | 4 |
| 2002 | The hierarchical cluster model for image region segmentationabstractThe hierarchical cluster model (HCM), a neural network inspired by the human brain (see Sutton, J., Harvard Medical School, MIT, Neural Systems Group, Technical Report, 1995), is demonstrated for the purpose of region segmentation in digital images. Starting with an over segmented image, regions are merged based on evidence of a valid edge between the two regions. Unlike Sutton's work, in which the HCM is used to recall a set of pre-trained memory patterns, the HCM in our work demonstrates unsupervised decision making capabilities. Jonathan Randall, Ling Guan, Wanqing Li 0001 |
ICME (1) | 4 |
| 2001 | 2D to pseudo-3D conversion of "head and shoulder" images using feature based parametric disparity mapsabstractThis paper presents a method of converting a 2D still photo containing the head & shoulders of a human (e.g. a passport photo) to pseudo-3D, so that the depth can be perceived via stereopsis. This technology has the potential to be included in self-serve photo booths and, also as an added accessory (i.e. software package) for digital still cameras and scanners. The basis of the algorithm is to exploit the ability of the human visual system in combining monoscopic and stereoscopic cues for depth perception. Common facial features are extracted from the 2D photograph, in order to create a parametric depth map that conforms to the available monoscopic depth cues. The original 2D photograph and the created depth map are used to generate left and right views for stereoscopic viewing. The algorithm is implemented in software, and promising results are obtained. Chaminda Weerasinghe, Philip Ogunbona, Wanqing Li 0001 |
ICIP (3) | 3 |
| 1997 | A geometric approach to cluster validity for normal mixtures
James C. Bezdek, Wanqing Li 0001, Yianni Attikiouzel, Michael P. Windham |
Soft Comput. | 2 |
| 1995 | Unsupervised segmentation of dual-echo MR images by a sequentially learned Gaussian mixture modelabstractThis paper proposes a method for unsupervised segmentation of brain tissues from dual-echo MR images without any prior knowledge about the number of tissues and their density distributions on each MRI echo. The brain tissues are described by a finite Gaussian mixture model (FGMM). The FGMM parameters are learned by sequentially applying the expectation maximization (EM) algorithm to a stream of data sets which are specifically organized according to the global spatial relationship of the brain tissues. Preliminary results on actual MRI slices have shown the method to be promising. Wanqing Li 0001, Mark W. Morrison, Yianni Attikiouzel |
ICIP (3) | 1 |
| 1994 | An Efficient Method of Volume Rendering for Medical SlicesabstractThe volume rendering method uses a special light source, slit-light source, to illuminate the volume and simulates the interactions of light with objects, such as light absorption, scattering and reflection. In comparison with previous methods, our method has several advantages. This includes flexibility for implementation, efficiency for computation, high reality for display and easiness of parallelization. The algorithm can be parallelized easily either using the data space subdivision method or the image space subdivision method. We describe some experimental results of applying our method to several sets of MRI and CT slices of human heads, which show the effectiveness of our method.> Wanqing Li 0001, Yianni Attikiouzel |
ICIP (1) | 1 |