Wankou Yang

dblp:99/3602 · DBLP profile ↗
← Back
177ranked-venue papers
21as first author
87since 2021 · last 2026
0000-0002-6385-6776ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 127 · 19 first-author · 58 since 2021Graphics, computer vision, multimedia, augmented reality and games · 58 · 5 first-author · 30 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
abstract
Efficient Multimodal Large Language Models (MLLMs) compress vision tokens to reduce resource consumption, but the loss of visual information can degrade comprehension capabilities. Although some priors introduce Knowledge Distillation to enhance student models, they overlook the fundamental differences in fine-grained vision comprehension caused by unbalanced vision tokens between the efficient student and vanilla teacher. In this paper, we propose EM-KD, a novel paradigm that enhances the Efficient MLLMs with Knowledge Distillation. To overcome the challenge of unbalanced vision tokens, we first calculate the Manhattan distance between the vision logits of teacher and student, and then align them in the spatial dimension with the Hungarian matching algorithm. After alignment, EM-KD introduces two distillation strategies: 1) Vision-Language Affinity Distillation (VLAD) and 2) Vision Semantic Distillation (VSD). Specifically, VLAD calculates the affinity matrix between text tokens and aligned vision tokens, and minimizes the smooth L1 distance of the student and the teacher affinity matrices. Considering the semantic richness of vision logits in the final layer, VSD employs the reverse KL divergence to measure the discrete probability distributions of the aligned vision logits over the vocabulary space. Comprehensive evaluation on diverse benchmarks demonstrates that EM-KD trained model outperforms prior Efficient MLLMs on both accuracy and efficiency with a large margin, validating its effectiveness. Compared with previous distillation methods, which are equipped with our proposed vision token matching strategy for fair comparison, EM-KD also achieves better performance.
Ze Feng, Boqiang Duan, Wankou Yang, Jingdong Wang 0001
AAAI4
2026 Cross-modal semantic token alignment via contrastive learning for weakly-supervised referring image segmentation
Congwei Zhang, Zhibin Quan, Wankou Yang
Expert Syst. Appl.3
2026 HoloHand: enhancing real-time neural hand rendering with self-occlusion-aware appearance fields and depth-aware learning
Qinghan Xiao, Menglei Zhang, Wankou Yang
Multim. Syst.5
2026 Self-visual prompting for training-free recursive weakly supervised semantic segmentation
Congwei Zhang, Zhibin Quan, Yuncong Yao, Wankou Yang
Multim. Syst.4
2026 Shadow-DETR: Alleviating matching conflicts through shadow queries
Jie Li 0040, Lingfeng Yang, Yifei Su, Yingpeng Li, Wankou Yang
Neural Networks6
2026 LoongTrack: Exploring long-sequence modeling for visual tracking
Tianyang Xu 0001, Mu Nie, Wankou Yang
Neural Networks5
2026 TATrack: Target-oriented adaptive vision transformer for UAV tracking
Tianyang Xu 0001, Wankou Yang
Neural Networks5
2026 Improving Generalized Visual Grounding With Instance-Aware Joint Learning
abstract
Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target scenarios. Specifically, GREC focuses on accurately identifying all referential objects at the coarse bounding box level, while GRES aims for achieve fine-grained pixel-level perception. However, existing approaches typically treat these tasks independently, overlooking the benefits of jointly training GREC and GRES to ensure consistent multi-granularity predictions and streamline the overall process. Moreover, current methods often treat GRES as a semantic segmentation task, neglecting the crucial role of instance-aware capabilities and the necessity of ensuring consistent predictions between instance-level boxes and masks. To address these limitations, we propose InstanceVG, a multi-task generalized visual grounding framework equipped with instance-aware capabilities, which leverages instance queries to unify the joint and consistency predictions of instance-level boxes and masks. To the best of our knowledge, InstanceVG is the first framework to simultaneously tackle both GREC and GRES while incorporating instance-aware capabilities into generalized visual grounding. To instantiate the framework, we assign each instance query a prior reference point, which also serves as an additional basis for target matching. This design facilitates consistent predictions of points, boxes, and masks for the same instance. Extensive experiments obtained on ten datasets across four tasks demonstrate that InstanceVG achieves state-of-the-art performance, significantly surpassing the existing methods in various evaluation metrics. The code and model will be publicly available at https://github.com/Dmmm1997/InstanceVG.
Wenxuan Cheng, Jiang-Jiang Liu 0001, Lingfeng Yang, Zhenhua Feng 0001, Wankou Yang, Jingdong Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Probabilistic modeling of disparity uncertainty for robust and efficient stereo matching
Wenxiao Cai, Dongting Hu, Ruoyan Yin, Jiankang Deng, Huan Fu, Wankou Yang, Mingming Gong
Pattern Recognit.6
2026 PLRVG: Progressive layer-wise refinement for visual grounding via deep-to-shallow decoding
Wenxuan Cheng, Wankou Yang
Pattern Recognit.3
2026 DRL: An efficient heterogeneous spatial feature interaction framework for UAV self-localization
Enhui Zheng, Wenxuan Cheng, Zhenhua Feng 0001, Wankou Yang
Pattern Recognit.6
2026 PoseAdapter: Efficiently transferring 2D human pose estimator to 3D whole-body task via adapter
Ze Feng, Jiang-Jiang Liu 0001, Wankou Yang
Pattern Recognit.4
2026 IRDFusion: Iterative relation-map difference guided feature fusion for multispectral object detection
Jifeng Shen, Haibo Zhan, Heng Fan 0001, Xiaohui Yuan 0001, Jun Li 0033, Wankou Yang
Pattern Recognit.7
2026 DTSI: Towards faster convergence of query-based detectors for rotated dense aerial images
Xiang Li 0041, Yuncong Yao, Wankou Yang
Pattern Recognit.5
2026 GC3VG: Generalized Multi-Task Visual Grounding With Coarse-to-Fine Consistency Constraints
abstract
In this work, we propose an efficient and streamlined paradigm to address the challenge of consistency prediction in generalized multi-task visual grounding. While most existing approaches primarily focus on integrating multi-modal information and employing multi-task learning to enhance both visual and linguistic understanding, they often rely on joint supervision at the region and pixel levels to exploit task complementarities. In contrast, C3VG explores the relatively under-addressed problem ofconsistency across multi-task predictions. To this end, a multi-task visual grounding framework based on a coarse-to-fine architecture is introduced. Empirical studies demonstrate that the incorporation of both implicit and explicit consistency constraints substantially enhances the coherence between detection and segmentation outputs. However, C3VG is restricted to single-referent visual grounding scenarios and exhibits limited generalizability to real-world applications, which often involve multi-referents or even absent referent. To overcome these limitations, we proposeGC3VG, which incorporates three key advancements: (1) extension to generalized scenarios, including both multi-referent and non-referent cases; (2) aUnified Coherent Refinement Modulethat implicitly encodes region- and instance-level features while explicitly modeling their relational alignment through an IoUbased constraint; and (3) aGranularity-aware Hard-mining Alignmentstrategy that enforces prediction consistency in the feature space and simultaneously enhances the discriminative power of visual and linguistic representations. Extensive experiments on RefCOCO/+/g and gRefCOCO demonstrate the effectiveness and generalizability of the proposed framework.
Kai Chen 0037, Wenxuan Cheng, Jiedong Zhuang, Zhenhua Feng 0001, Pengfei Zhu 0001, Wankou Yang
IEEE Trans. Circuits Syst. Video Technol.7
2025 Object-level Geometric Structure Preserving for Natural Image Stitching
abstract
The topic of stitching images with globally natural structures holds paramount significance, with two main goals: pixel-level alignment and distortion prevention. The existing approaches exhibit the ability to align well, yet fall short in maintaining object structures. In this paper, we endeavour to safeguard the overall OBJect-level structures within images based on Global Similarity Prior (OBJ-GSP), on the basis of good alignment performance. Our approach leverages semantic segmentation models like the family of Segment Anything Model to extract the contours of any objects in a scene. Triangular meshes are employed in image transformation to protect the overall shapes of objects within images. The balance between alignment and distortion prevention is achieved by allowing the object meshes to strike a balance between similarity and projective transformation. We also demonstrate that object-level semantic information is necessary in low-altitude aerial image stitching. Additionally, we propose StitchBench, the largest image stitching benchmark with most diverse scenarios. Extensive experimental results demonstrate that OBJ-GSP outperforms existing methods in both pixel alignment and shape preservation.
Wenxiao Cai, Wankou Yang
AAAI2
2025 Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
abstract
Multi-task visual grounding involves the simultaneous execution of localization and segmentation in images based on textual expressions. The majority of advanced methods predominantly focus on transformer-based multimodal fusion, aiming to extract robust multimodal representations. However, ambiguity between referring expression comprehension (REC) and referring image segmentation (RIS) is error-prone, leading to inconsistencies between multi-task predictions. Besides, insufficient multimodal understanding directly contributes to biased target perception. To overcome these challenges, we propose a Coarse-to-fine Consistency Constraints Visual Grounding architecture (C3VG), which integrates implicit and explicit modeling approaches within a two-stage framework. Initially, query and pixel decoders are employed to generate preliminary detection and segmentation outputs, a process referred to as the Rough Semantic Perception (RSP) stage. These coarse predictions are subsequently refined through the proposed Mask-guided Interaction Module (MIM) and a novel explicit bidirectional consistency constraint loss to ensure consistent representations across tasks, which we term the Refined Consistency Interaction (RCI) stage. Furthermore, to address the challenge of insufficient multimodal understanding, we leverage pre-trained models based on visual-linguistic fusion representations. Empirical evaluations on the RefCOCO, RefCOCO+, and RefCOCOg datasets demonstrate the efficacy and soundness of C3VG, which significantly outperforms state-of-the-art REC and RIS methods by a substantial margin.
Jiedong Zhuang, Wankou Yang
AAAI5
2025 DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation Through Loopback Synergy
abstract
Referring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions. While prior studies have predominantly concentrated on improving vision-language interactions and achieving fine-grained localization, a systematic analysis of the fundamental bottlenecks in existing RIS frameworks remains underexplored. To bridge this gap, we propose DeRIS, a novel framework that decomposes RIS into two key components: perception and cognition. This modular decomposition facilitates a systematic analysis of the primary bottlenecks impeding RIS performance. Our findings reveal that the predominant limitation lies not in perceptual deficiencies, but in the insufficient multi-modal cognitive capacity of current models. To mitigate this, we propose a Loopback Synergy mechanism, which enhances the synergy between the perception and cognition modules, thereby enabling precise segmentation while simultaneously improving robust image-text comprehension. Additionally, we analyze and introduce a simple non-referent sample conversion data augmentation to address the long-tail distribution issue related to target existence judgement in general scenarios. Notably, DeRIS demonstrates inherent adaptability to both non- and multi-referents scenarios without requiring specialized architectural modifications, enhancing its general applicability. The codes and models are available at https://github.com/Dmmm1997/DeRIS.
Wenxuan Cheng, Jiang-Jiang Liu 0001, Wenxiao Cai, Yanpeng Sun, Wankou Yang
ICCV7
2025 PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
abstract
Recent advances in visual grounding have largely shifted away from traditional proposal-based two-stage frameworks due to their inefficiency and high computational complexity, favoring end-to-end direct reference paradigms. However, these methods rely exclusively on the referred target for supervision, overlooking the potential benefits of prominent prospective targets. Moreover, existing approaches often fail to incorporate multi-granularity discrimination, which is crucial for robust object identification in complex scenarios. To address these limitations, we propose PropVG, an end-to-end proposal-based framework that, to the best of our knowledge, is the first to seamlessly integrate foreground object proposal generation with referential object comprehension without requiring additional detectors. Furthermore, we introduce a Contrastive-based Refer Scoring (CRS) module, which employs contrastive learning at both sentence and word levels to enhance the capability in understanding and distinguishing referred objects. Additionally, we design a Multi-granularity Target Discrimination (MTD) module that fuses object- and semantic-level information to improve the recognition of absent targets. Extensive experiments on gRefCOCO (GREC/GRES), Ref-ZOM, R-RefCOCO, and RefCOCO (REC/RES) benchmarks demonstrate the effectiveness of PropVG. The codes and models are available at https://github.com/Dmmm1997/PropVG.
Wenxuan Cheng, Jiedong Zhuang, Jiang-jiang Liu, Hongshen Zhao, Zhenhua Feng 0001, Wankou Yang
ICCV7
2025 SpatialBot: Precise Spatial Understanding with Vision Language Models
abstract
Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding; however, they still struggle with spatial understanding, which is fundamental to embodied AI. In this paper, we propose SpatialBot, a model designed to enhance spatial understanding by utilizing both RGB and depth images. To train VLMs for depth perception, we introduce the SpatialQA and SpatialQA$\boldsymbol{E}$datasets, which include multi-level depth-related questions spanning various scenarios and embodiment tasks. SpatialBench is also developed to comprehensively evaluate VLMs' spatial understanding capabilities across different levels. Extensive experiments on our spatial-understanding benchmark, general VLM benchmarks, and embodied AI tasks demonstrate the remarkable improvements offered by SpatialBot. The model, code, and datasets are available at https://github.com/BAAI-DCAI/SpatialBot.
Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li 0020, Wankou Yang, Hao Dong 0003, Bo Zhao 0037
ICRA5
2025 Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving
abstract
The autoregressive world model exhibits robust generalization capabilities in vectorized scene understanding but encounters difficulties in deriving actions due to insufficient uncertainty modeling and self-delusion. In this paper, we explore the feasibility of deriving decisions from an autoregres-sive world model by addressing these challenges through the formulation of multiple probabilistic hypotheses. We propose LatentDriver, a framework models the environment's next states and the ego vehicle's possible actions as a mixture distribution, from which a deterministic control signal is then derived. By incorporating mixture modeling, the stochastic nature of decision-making is captured. Additionally, the self-delusion problem is mitigated by providing intermediate actions sampled from a distribution to the world model. Experimen-tal results on the recently released closed-loop benchmark Waymax demonstrate that LatentDriver surpasses state-of-the-art reinforcement learning and imitation learning methods, achieving expert-level performance. The code and models will be made available at https://github.com/Sephirex-X/LatentDriver.
Lingyu Xiao, Jiang-Jiang Liu 0001, Xiaoqing Ye, Wankou Yang, Jingdong Wang 0001
ICRA6
2025 Precise GPS-Denied UAV Self-positioning via Context-Enhanced Cross-View Geo-Localization
Yuanze Xu, Wankou Yang
PRCV (15)3
2025 Synthesizing Spreading-out features for generative zero-shot image classification
Jingren Liu, Zheng Zhang 0006, Yang Long 0001, Wankou Yang, Yunyang Yan, Haofeng Zhang 0001
Eng. Appl. Artif. Intell.5
2025 Efficient template-separable hierarchical transformer tracking for edge computing
Wankou Yang
Eng. Appl. Artif. Intell.2
2025 VDD: Varied Drone Dataset for semantic segmentation
Wenxiao Cai, Jinyan Hou, Letian Wu, Wankou Yang
J. Vis. Commun. Image Represent.6
2025 Exploring efficient appearance prompts for light-weight object tracking
Wankou Yang
J. Vis. Commun. Image Represent.3
2025 Improving underwater semantic segmentation with underwater image quality attention and muti-scale aggregation attention
Jiaran Jiang, Jifeng Shen, Wankou Yang
Pattern Anal. Appl.4
2025 Enhancing knowledge distillation for semantic segmentation through text-assisted modular plugins
Letian Wu, Chuankai Zhang, Jiajun Liang, Wankou Yang
Pattern Recognit.6
2025 Domain adaptive depth completion via spatial-error consistency
Lingyu Xiao, Junjie Hu 0003, Wankou Yang
Pattern Recognit.5
2025 Global-local information sensitivity adjustment factor
Zinan Cheng, Letian Wu, Lei Qi 0001, Wankou Yang
Pattern Recognit. Lett.5
2025 SFFR: Spatial-Frequency Feature Reconstruction for Multispectral Aerial Object Detection
abstract
Recent multispectral object detection methods have primarily focused on spatial-domain feature fusion based on CNNs or Transformers, while the potential of frequency-domain feature remains underexplored. In this work, we propose a novel Spatial and Frequency Feature Reconstruction method (SFFR) method, which leverages the spatial-frequency feature representation mechanisms of the Kolmogorov–Arnold Network (KAN) to reconstruct complementary representations in both spatial and frequency domains prior to feature fusion. The core components of SFFR are the proposed Frequency Component Exchange KAN (FCEKAN) module and Multi-Scale Gaussian KAN (MSGKAN) module. The FCEKAN introduces an innovative selective frequency component exchange strategy that effectively enhances the complementarity and consistency of cross-modal features based on the frequency feature of RGB and IR images. The MSGKAN module demonstrates excellent nonlinear feature modeling capability in the spatial domain. By leveraging multi-scale Gaussian basis functions, it effectively captures the feature variations caused by scale changes at different UAV flight altitudes, significantly enhancing the model’s adaptability and robustness to scale variations. It is experimentally validated that our proposed FCEKAN and MSGKAN modules are complementary and can effectively capture the frequency and spatial semantic features respectively for better feature fusion. Extensive experiments on the SeaDroneSee, DroneVehicle and DVTOD datasets demonstrate the superior performance and significant advantages of the proposed method in UAV multispectral object perception task. Code will be available at https://github.com/qchenyu1027/SFFR.
Chenyu Qu, Haibo Zhan, Jifeng Shen, Wankou Yang
IEEE Trans. Geosci. Remote. Sens.5
2024 SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
abstract
Visual grounding is a common vision task that involves grounding descriptive sentences to the corresponding regions of an image. Most existing methods use independent image-text encoding and apply complex hand-crafted modules or encoder-decoder architectures for modal interaction and query reasoning. However, their performance significantly drops when dealing with complex textual expressions. This is because the former paradigm only utilizes limited downstream data to fit the multi-modal feature fusion. Therefore, it is only effective when the textual expressions are relatively simple. In contrast, given the wide diversity of textual expressions and the uniqueness of downstream training data, the existing fusion module, which extracts multimodal content from a visual-linguistic context, has not been fully investigated. In this paper, we present a simple yet robust transformer-based framework, SimVG, for visual grounding. Specifically, we decouple visual-linguistic feature fusion from downstream tasks by leveraging existing multimodal pre-trained models and incorporating additional object tokens to facilitate deep integration of downstream and pre-training tasks. Furthermore, we design a dynamic weight-balance distillation method in the multi-branch synchronous learning process to enhance the representation capability of the simpler branch. This branch only consists of a lightweight MLP, which simplifies the structure and improves reasoning speed. Experiments on six widely used VG datasets, i.e., RefCOCO/+/g, ReferIt, Flickr30K, and GRefCOCO, demonstrate the superiority of SimVG. Finally, the proposed method not only achieves improvements in efficiency and convergence speed but also attains new state-of-the-art performance on these benchmarks. Codes and models are available at https://github.com/Dmmm1997/SimVG.
Lingfeng Yang, Zhenhua Feng 0001, Wankou Yang
NeurIPS5
2024 Fusing spatial and frequency features for compositional zero-shot image classification
Suyi Li 0005, Chenyi Jiang, Qiaolin Ye, Wankou Yang, Haofeng Zhang 0001
Expert Syst. Appl.5
2024 Unsupervised cross domain semantic segmentation with mutual refinement and information distillation
Dexin Ren, Zheng Zhang 0006, Wankou Yang, Mingwu Ren, Haofeng Zhang 0001
Neurocomputing4
2024 M-YOLOv8s: An improved small target detection algorithm for UAV aerial photography
Siyao Duan, Wankou Yang
J. Vis. Commun. Image Represent.4
2024 Efficient object tracking on edge devices with MobileTrack
Jiang Zhai, Zinan Cheng, Dejun Zhu, Wankou Yang
J. Vis. Commun. Image Represent.5
2024 Adaptive Angle Module and Radian Regression Method for Rotated Object Detection
abstract
Rotated object detectors commonly encounter instability during the training process, primarily due to background noise and angular periodicity. Targets with elongated or nonconvex shapes may introduce background noise during convolution, hindering accurately extracting features. Meanwhile, the periodicity of angles leads to predictions beyond the defined range, subsequently impeding the convergence. This letter introduces an angle adaptive module (AAM) designed for the backbone, enhancing the ability of the model to accurately extract object features and dynamically select the optimal angle. Moreover, to mitigate the effect of angle periodicity, a method called radian regression method (RRM) is proposed for predicting proper angles. It avoids directly regressing the value and instead produces the probability density distribution of the offset. We elaborately design numerous experiments to demonstrate the effectiveness of the proposed modules. As a result, the proposed method attains competitive results across various datasets, including DOTAv1.0, DOTAv1.5, DOTAv2.0, HRSC2016, and DIOR-R.
Dejun Zhu, Wankou Yang
IEEE Geosci. Remote. Sens. Lett.4
2024 Correlation-Embedded Transformer Tracking: A Single-Branch Framework
abstract
Developing robust and discriminative appearance models has been a long-standing research challenge in visual object tracking. In the prevalent Siamese-based paradigm, the features extracted by the Siamese-like networks are often insufficient to model the tracked targets and distractor objects, thereby hindering them from being robust and discriminative simultaneously. While most Siamese trackers focus on designing robust correlation operations, we propose a novel single-branch tracking framework inspired by the transformer. Unlike the Siamese-like feature extraction, our tracker deeply embeds cross-image feature correlation in multiple layers of the feature network. By extensively matching the features of the two images through multiple layers, it can suppress non-target features, resulting in target-aware feature extraction. The output features can be directly used to predict target locations without additional correlation steps. Thus, we reformulate the two-branch Siamese tracking as a conceptually simple, fully transformer-based Single-Branch Tracking pipeline, dubbed SBT. After conducting an in-depth analysis of the SBT baseline, we summarize many effective design principles and propose an improved tracker dubbed SuperSBT. SuperSBT adopts a hierarchical architecture with a local modeling layer to enhance shallow-level features. A unified relation modeling is proposed to remove complex handcrafted layer pattern designs. SuperSBT is further improved by masked image modeling pre-training, integrating temporal modeling, and equipping with dedicated prediction heads. Thus, SuperSBT outperforms the SBT baseline by 4.7%,3.0%, and 4.5% AUC scores in LaSOT, TrackingNet, and GOT-10K. Notably, SuperSBT greatly raises the speed of SBT from 37 FPS to 81 FPS. Extensive experiments show that our method achieves superior results on eight VOT benchmarks.
Wankou Yang, Chunyu Wang 0001, Yue Cao 0001, Chao Ma 0004, Wenjun Zeng 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 ICAFusion: Iterative cross-attention guided feature fusion for multispectral object detection
Jifeng Shen, Heng Fan 0001, Wankou Yang
Pattern Recognit.6
2024 SSPNet: Scale and spatial priors guided generalizable and interpretable pedestrian attribute recognition
Jifeng Shen, Teng Guo 0003, Heng Fan 0001, Wankou Yang
Pattern Recognit.5
2024 PARDet: Dynamic point set alignment for rotated object detection
Jifeng Shen, Wankou Yang
Pattern Recognit.4
2024 SED: Searching Enhanced Decoder with switchable skip connection for semantic segmentation
Zhibin Quan, Qiang Li 0024, Dejun Zhu, Wankou Yang
Pattern Recognit.5
2024 CRTrack: Learning Correlation-Refine network for visual object tracking
Tianyang Xu 0001, Jiang Zhai, Wankou Yang
Pattern Recognit.5
2024 The MorPhEMe Machine: An Addressable Neural Memory for Learning Knowledge-Regularized Deep Contextualized Chinese Embedding
abstract
Deep contextualized embeddings, as learned by large pre-training models, have proven highly effective in various downstream natural language processing tasks. However, the embedding space in these large models lacks explicit regularization, leading to underfitting and substantial costs during large-scale training on huge corpora. In this paper, we present a novel approach to learning deep contextualized embeddings, introducing linguistic knowledge regularization. Specifically, our proposed model, MorPhEMe (Morphology and Phonology Embedding Memory), features an external addressable memory with two additional addressable memories for storing morphology and phonology knowledge. MorPhEMe can be seamlessly stacked into a deep architecture. Notably different from existing pre-training models, MorPhEMe boasts two distinctive features: (1) compositional encoding and decompositional decoding facilitated by a dynamic addressing mechanism; and (2) explicit memory embedding regularization through cross-layer memory sharing. Theoretical analysis suggests that the inclusion of morphology and phonology enables MorPhEMe to reduce the modeling complexity of natural language sequences. We evaluate MorPhEMe across a diverse set of Chinese natural language processing tasks, including language modeling, word similarity computation, word analogy reasoning, relation extraction, and machine reading comprehension. Experimental results demonstrate that MorPhEMe, in contrast to state-of-the-art models, achieves remarkable improvements with fewer parameters and rapid convergence.
Zhibin Quan, Chi-Man Vong, Wankou Yang
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 A Novel Hybrid Transformer-Based Framework for Solar Irradiance Forecasting Under Incomplete Data Scenarios
abstract
Accurate prediction of solar irradiance is crucial for the effective utilization of solar energy. However, in real-world scenarios, complex irradiance patterns and prevalent incomplete data pose challenges to precise forecasting, resulting in additional uncertainties and instability. To address these issues, this study proposes a novel irradiance forecasting model that integrates a Mask-Transformer data imputation module and a prediction module centered around the typical patterns representation mechanism. The Mask-Transformer leverages a mask modeling mechanism to model the context of missing data, facilitating accurate estimation of missing values and reducing noise and uncertainty in the input data. The typical patterns representation mechanism comprises a series decomposition module and a feature fusion module, providing the module with the capability to mitigate nonlinearity and nonstationarity in solar irradiance data. This enhancement leads to improved short-term forecasting performance while maintaining long-term forecasting capabilities. Experimental results on two datasets demonstrate that the proposed model exhibits sufficient robustness and accuracy, making it effective in scenarios with incomplete data.
Hanjin Zhang, Bin Li 0085, Shun-Feng Su, Wankou Yang
IEEE Trans. Ind. Informatics4
2024 Vision-Based UAV Self-Positioning in Low-Altitude Urban Environments
abstract
Unmanned Aerial Vehicles (UAVs) rely on satellite systems for stable positioning. However, due to limited satellite coverage or communication disruptions, UAVs may lose signals for positioning. In such situations, vision-based techniques can serve as an alternative, ensuring the self-positioning capability of UAVs. However, most of the existing datasets are developed for the geo-localization task of the objects captured by UAVs, rather than UAV self-positioning. Furthermore, the existing UAV datasets apply discrete sampling to synthetic data, such as Google Maps, neglecting the crucial aspects of dense sampling and the uncertainties commonly experienced in practical scenarios. To address these issues, this paper presents a new dataset, DenseUAV, that is the first publicly available dataset tailored for the UAV self-positioning task. DenseUAV adopts dense sampling on UAV images obtained in low-altitude urban areas. In total, over 27K UAV- and satellite-view images of 14 university campuses are collected and annotated. In terms of methodology, we first verify the superiority of Transformers over CNNs for the proposed task. Then we incorporate metric learning into representation learning to enhance the model's discriminative capacity and to reduce the modality discrepancy. Besides, to facilitate joint learning from both the satellite and UAV views, we introduce a mutually supervised learning approach. Last, we enhance the Recall@K metric and introduce a new measurement, SDM@K, to evaluate both the retrieval and localization performance for the proposed task. As a result, the proposed baseline method achieves a remarkable Recall@1 score of 83.01% and an SDM@1 score of 86.50% on DenseUAV. The dataset and code have been made publicly available on https://github.com/Dmmm1997/DenseUAV.
Enhui Zheng, Zhenhua Feng 0001, Lei Qi 0001, Jiedong Zhuang, Wankou Yang
IEEE Trans. Image Process.6
2024 Attentional Composition Networks for Long-Tailed Human Action Recognition
abstract
The problem of long-tailed visual recognition has been receiving increasing research attention. However, the long-tailed distribution problem remains underexplored for video-based visual recognition. To address this issue, in this article we propose a compositional learning based solution for video-based human action recognition. Our method, named Attentional Composition Networks (ACN), first learns verb-like and preposition-like components, then shuffles these components to generate samples for the tail classes in the feature space to augment the data for the tail classes. Specifically, during training, we represent each action video by a graph that captures the spatial-temporal relations (edges) among detected human/object instances (nodes). Then, ACN utilizes the position information to decompose each action into a set of verb and preposition representations using the edge features in the graph. After that, the verb and preposition features from different videos are combined via an attention structure to synthesize feature representations for tail classes. This way, we can enrich the data for the tail classes and consequently improve the action recognition for these classes. To evaluate the compositional human action recognition, we further contribute a new human action recognition dataset, namely NEU-Interaction (NEU-I). Experimental results on both Something-Something V2 and the proposed NEU-I demonstrate the effectiveness of the proposed method for long-tailed, few-shot, and zero-shot problems in human action recognition. Source code and the NEU-I dataset are available at https://github.com/YajieW99/ACN .
Haoran Wang 0001, Baosheng Yu, Yibing Zhan, Chunfeng Yuan, Wankou Yang
ACM Trans. Multim. Comput. Commun. Appl.6
2023 Enhancing Defect Diagnosis and Localization in Wafer Map Testing Through Weakly Supervised Learning
abstract
Defect diagnosis and localization in wafer maps are crucial tasks in semiconductor manufacturing. Existing deep learning methods often require pixel-level annotations, making them impractical for large-scale deployment. In this paper, we propose a novel weakly supervised learning approach to achieving high-precision defect identification and effective localization with only image-level labels. By leveraging the information of defect types and locations, we introduce a weighted fusion of activation maps, called Class Activation Map (CAM), to highlight classspecific regions. We further enhance defect localization accuracy and completeness by employing optimized region growing operations to eliminate noise in defect regions. Moreover, we present an optimized inference method that provides meaningful visual explanations for defect recognition. Experimental results on real-world wafer map images demonstrate the effectiveness of our approach in accurately segmenting defect patterns with no pixel-level annotations. By training the model solely on wafer map image classification labels, our proposed model significantly improves defect recognition, facilitating efficient defect analysis in semiconductor manufacturing. The proposed weakly supervised learning approach offers a practical solution for defect diagnosis and localization, with the potential of widespread adoption in the semiconductor industry.
Mu Nie, Wankou Yang, Senling Wang, Xiaoqing Wen, Tianming Ni
ATS3
2023 GPA-3D: Geometry-aware Prototype Alignment for Unsupervised Domain Adaptive 3D Object Detection from Point Clouds
abstract
LiDAR-based 3D detection has made great progress in recent years. However, the performance of 3D detectors is considerably limited when deployed in unseen environments, owing to the severe domain gap problem. Existing domain adaptive 3D detection methods do not adequately consider the problem of the distributional discrepancy in feature space, thereby hindering generalization of detectors across domains. In this work, we propose a novel unsupervised domain adaptive 3D detection framework, namely Geometry-aware Prototype Alignment (GPA-3D), which explicitly leverages the intrinsic geometric relationship from point cloud objects to reduce the feature discrepancy, thus facilitating cross-domain transferring. Specifically, GPA-3D assigns a series of tailored and learnable prototypes to point cloud objects with distinct geometric structures. Each prototype aligns BEV (bird’s-eye-view) features derived from corresponding point cloud objects on source and target domains, reducing the distributional discrepancy and achieving better adaptation. The evaluation results obtained on various benchmarks, including Waymo, nuScenes and KITTI, demonstrate the superiority of our GPA-3D over the state-of-the-art approaches for different adaptation scenarios. The MindSpore version code will be publicly available at https://github.com/Liz66666/GPA3D.
Tongtong Cao, Wankou Yang
ICCV5
2023 ADNet: Lane Shape Prediction via Anchor Decomposition
abstract
In this paper, we revisit the limitations of anchor-based lane detection methods, which have predominantly focused on fixed anchors that stem from the edges of the image, disregarding their versatility and quality. To overcome the inflexibility of anchors, we decompose them into learning the heat map of starting points and their associated directions. This decomposition removes the limitations on the starting point of anchors, making our algorithm adaptable to different lane types in various datasets. To enhance the quality of anchors, we introduce the Large Kernel Attention (LKA) for Feature Pyramid Network (FPN). This significantly increases the receptive field, which is crucial in capturing the sufficient context as lane lines typically run throughout the entire image. We have named our proposed system the Anchor Decomposition Network (ADNet). Additionally, we propose the General Lane IoU (GLIoU) loss, which significantly improves the performance of ADNet in complex scenarios. Experimental results on three widely used lane detection benchmarks, VIL-100, CU-Lane, and TuSimple, demonstrate that our approach outperforms the state-of-the-art methods on VIL-100 and exhibits competitive accuracy on CULane and TuSimple. Code and models will be released on https://github.com/Sephirex-X/ADNet.
Lingyu Xiao, Xiang Li 0041, Wankou Yang
ICCV4
2023 LiftedCL: Lifting Contrastive Learning for Human-Centric Perception
Qiang Li 0024, Wankou Yang
ICLR4
2023 Capturing the Motion of Every Joint: 3D Human Pose and Shape Estimation with Independent Tokens
Wen Heng, Guozhong Luo, Wankou Yang, Gang Yu 0002
ICLR5
2023 CLIP for Lightweight Semantic Segmentation
Wankou Yang
PRCV (10)2
2023 Enhancing motion visual cues for self-supervised video representation learning
Mu Nie, Zhibin Quan, Weiping Ding 0001, Wankou Yang
Eng. Appl. Artif. Intell.4
2023 Compact global association based adaptive routing framework for personnel behavior understanding
Yimin Zhou 0002, Juan Wang 0017, Zuli Wang 0001, Wankou Yang, Edward Szczerbicki
Future Gener. Comput. Syst.7
2023 Improving multispectral pedestrian detection with scale-aware permutation attention and adjacent feature aggregation
abstract
Abstract High quality feature fusion module is one of the key components for multispectral pedestrian detection system in challenging situations, such as large‐scale variance and occlusion. Although attention mechanism is one of the most effective ways for feature refining, the correlation between attention and scales in feature pyramid still remains unknown. Therefore, a scale‐aware permutated attention module is proposed to enhance features of objects with different scales adaptively in the feature pyramid. Specifically, four different local and global attention sub‐modules are investigated to refine feature maps with different permutations in the Feature Pyramid Networks, improving the quality of the feature fusion. Besides, to address the high miss‐rate issue for small‐sized pedestrians, an adjacent‐branch feature aggregation module is proposed to aggregate features across different scales, taking both semantic context and spatial resolution into consideration. Both modules can benefit from each other with significant performance improvement in terms of efficiency and accuracy, when equipped with the dual‐branch CenterNet detection framework. Experiments on the KAIST and FLIR datasets demonstrate its superior performance compared with other state‐of‐the‐arts.
Jifeng Shen, Wankou Yang
IET Comput. Vis.4
2023 Prototype-guided Instance matching for multiple pedestrian tracking
Qiang Wang 0023, Wankou Yang, Chunyan Xu, Zhen Cui 0001
Neurocomputing3
2023 Enhancing time series forecasting: A hierarchical transformer with probabilistic decomposition representation
Junlong Tong, Wankou Yang, Kan-Jian Zhang, Junsheng Zhao
Inf. Sci.3
2023 UAV image stitching by estimating orthograph with RGB cameras
abstract
In the field of image stitching, cases with large camera optical center movement and large parallax have been the virgin territory of research. The goal of image stitching is to overcome the parallax and stitch a natural image. We look into this problem in the context of ultra-low altitude flight of a UAV . We model the 3D world in this scenario and quickly estimate orthographic projection by pairs of homography matrices. Our stitching method can achieve precise alignment since it takes parallax well into consideration. The stitching results are natural and the extra time consumed is short.
Wenxiao Cai, Songlin Du, Wankou Yang
J. Vis. Commun. Image Represent.3
2023 LiDAR-only 3D object detection based on spatial context
Qiang Wang 0023, Dejun Zhu, Wankou Yang
J. Vis. Commun. Image Represent.4
2023 BFANet: Effective segmentation network for low altitude high-resolution urban scene image
Letian Wu, Dejun Zhu, Wankou Yang
J. Vis. Commun. Image Represent.4
2023 EfficientFace: an efficient deep network with feature enhancement for accurate face detection
Guangtao Wang, Jun Li 0033, Zhijian Wu, Jifeng Shen, Wankou Yang
Multim. Syst.6
2023 Flow Learning Based Dual Networks for Low-Light Image Enhancement
Changhui Hu 0001, Weilin Yi, Ziyun Cai, Mingliang Zhai, Wankou Yang
Neural Process. Lett.6
2023 Keep an eye on faces: Robust face detection with heatmap-Assisted spatial attention and scale-Aware layer attention
abstract
Modern anchor-based face detectors learn discriminative features using large-capacity networks and extensive anchor settings. In spite of their promising results, they are not without problems. First, most anchors extract redundant features from the background. As a consequence, the performance improvements are achieved at the expense of a disproportionate computational complexity. Second, the predicted face boxes are only distinguished by a classifier supervised by pre-defined positive, negative and ignored anchors. This strategy may ignore potential contributions from cohorts of anchors labeled negative/ignored during inference simply because of their inferior initialisation, although they can regress well to a target. In other words, true positives and representative features may get filtered out by unreliable confidence scores. To deal with the first concern and achieve more efficient face detection, we propose a Heatmap-assisted Spatial Attention (HSA) module and a Scale-aware Layer Attention (SLA) module to extract informative features using lower computational costs. To be specific, SLA incorporates the information from all the feature pyramid layers, weighted adaptively to remove redundant layers. HSA predicts a reshaped Gaussian heatmap and employs it to facilitate a spatial feature selection by better highlighting facial areas. For more reliable decision-making, we merge the predicted heatmap scores and classification results by voting. Since our heatmap scores are based on the distance to the face centres, they are able to retain all the well-regressed anchors. The experiments obtained on several well-known benchmarks demonstrate the merits of the proposed method.
Lei Ju 0005, Josef Kittler, Muhammad Awais Rana, Wankou Yang, Zhenhua Feng 0001
Pattern Recognit.4
2023 Detecting and grouping keypoints for multi-person pose estimation using instance-aware attention
Ze Feng, Zhicheng Wang 0001, Shoukui Zhang, Zhibin Quan, Shutao Xia, Wankou Yang
Pattern Recognit.8
2023 Pyramid Geometric Consistency Learning For Semantic Segmentation
Qiang Li 0024, Zhibin Quan, Wankou Yang
Pattern Recognit.4
2023 Joint Image-to-Image Translation for Traffic Monitoring Driver Face Image Enhancement
abstract
The real traffic monitoring driver face (TMDF) images are with complex multiple degradations, which decline face recognition accuracy in real intelligent transportation systems (ITS). This paper is the first to propose joint image-to-image (I2I) translation to enhance TMDF images of ITS. First, as TMDF images are without corresponding clear ones, identity preserving is critical for TMDF images under unpaired I2I translation. This paper proposes a fast diagonal symmetry pattern (FDSP) to preserve identity structure under unpaired I2I translation. Second, FDSP is introduced into CycleGAN to form FDSP-CG, which aims to learn the degradation mapping (i.e., FDSP-CG-d) from the clarity domain to the degradation domain. FDSP-CG-d can generate massive degradation/clarity image pairs for paired I2I translation training. Third, this paper proposes the dual residual block (DRB) to strengthen Pix2pix for rich face detail features learning (i.e., DRB-P2P), which learns the enhancement mapping from the degradation image to its clear version under paired I2I translation. Finally, the experiments on TMDF (i.e., the brevity name of the face database collected from real ITS) and Chinese famous face (CFF) databases, as well as CelebA and MegaFace databases, indicate that the proposed method can efficiently enhance TMDF images whose degradation variations are learned by FDSP-CG.
Changhui Hu 0001, Lin-Tao Xu, Xiaoyuan Jing, Xiaobo Lu, Wankou Yang, Pan Liu 0013
IEEE Trans. Intell. Transp. Syst.6
2022 Correlation-Aware Deep Tracking
abstract
Robustness and discrimination power are two fundamental requirements in visual object tracking. In most tracking paradigms, we find that the features extracted by the popular Siamese-like networks cannot fully discriminatively model the tracked targets and distractor objects, hindering them from simultaneously meeting these two requirements. While most methods focus on designing robust correlation operations, we propose a novel target-dependent feature network inspired by the self-/cross-attention scheme. In contrast to the Siamese-like feature extraction, our network deeply embeds cross-image feature correlation in multiple layers of the feature network. By extensively matching the features of the two images through multiple layers, it is able to suppress non-target features, resulting in instance-varying feature extraction. The output features of the search image can be directly used for predicting target locations without extra correlation step. Moreover, our model can be flexibly pre-trained on abundant unpaired images, leading to notably faster convergence than the existing methods. Extensive experiments show our method achieves the state-of-the-art results while running at real-time. Our feature networks also can be applied to existing tracking pipelines seamlessly to raise the tracking performance.
Chunyu Wang 0001, Guangting Wang, Yue Cao 0001, Wankou Yang, Wenjun Zeng 0001
CVPR5
2022 SimCC: A Simple Coordinate Classification Perspective for Human Pose Estimation
Peidong Liu 0003, Shoukui Zhang, Zhicheng Wang 0001, Wankou Yang, Shutao Xia
ECCV (6)7
2022 Semi-supervised cross-modal hashing with multi-view graph representation
Haofeng Zhang 0001, Lunbo Li, Wankou Yang, Li Liu 0004
Inf. Sci.4
2022 Subspace-based self-weighted multiview fusion for instance retrieval
Zhijian Wu, Jun Li 0033, Wankou Yang
Inf. Sci.4
2022 Siamese Transformer Network: Building an autonomous real-time target tracking system for UAV
Xiaolou Sun, Zhibin Quan, Hao Wang 0144, Yuncong Yao, Wankou Yang
J. Syst. Archit.8
2022 Learning discriminative and representative feature with cascade GAN for generalized zero-shot learning
Jingren Liu, Liyong Fu, Haofeng Zhang 0001, Qiaolin Ye, Wankou Yang, Li Liu 0004
Knowl. Based Syst.5
2022 Spatial information enhancement network for 3D object detection from point cloud
Yuncong Yao, Zhibin Quan, Wankou Yang
Pattern Recognit.5
2022 Searching part-specific neural fabrics for human pose estimation
Wankou Yang, Zhen Cui 0001
Pattern Recognit.2
2022 Multiview Learning With Robust Double-Sided Twin SVM
abstract
Multiview learning (MVL), which enhances the learners' performance by coordinating complementarity and consistency among different views, has attracted much attention. The multiview generalized eigenvalue proximal support vector machine (MvGSVM) is a recently proposed effective binary classification method, which introduces the concept of MVL into the classical generalized eigenvalue proximal support vector machine (GEPSVM). However, this approach cannot guarantee good classification performance and robustness yet. In this article, we develop multiview robust double-sided twin SVM (MvRDTSVM) with SVM-type problems, which introduces a set of double-sided constraints into the proposed model to promote classification performance. To improve the robustness of MvRDTSVM against outliers, we take L1-norm as the distance metric. Also, a fast version of MvRDTSVM (called MvFRDTSVM) is further presented. The reformulated problems are complex, and solving them are very challenging. As one of the main contributions of this article, we design two effective iterative algorithms to optimize the proposed nonconvex problems and then conduct theoretical analysis on the algorithms. The experimental results verify the effectiveness of our proposed methods.
Qiaolin Ye, Zhao Zhang 0001, Yuhui Zheng, Liyong Fu, Wankou Yang
IEEE Trans. Cybern.6
2022 Learning Robust Discriminant Subspace Based on Joint L₂, ₚ- and L₂, ₛ-Norm Distance Metrics
abstract
-norm as the distance metric. However, both of their robustness and discriminant power are limited. In this article, we present a new robust discriminant subspace (RDS) learning method for feature extraction, with an objective function formulated in a different form. To guarantee the subspace to be robust and discriminative, we measure the within-class distances based on [Formula: see text]-norm and use [Formula: see text]-norm to measure the between-class distances. This also makes our method include rotational invariance. Since the proposed model involves both [Formula: see text]-norm maximization and [Formula: see text]-norm minimization, it is very challenging to solve. To address this problem, we present an efficient nongreedy iterative algorithm. Besides, motivated by trace ratio criterion, a mechanism of automatically balancing the contributions of different terms in our objective is found. RDS is very flexible, as it can be extended to other existing feature extraction techniques. An in-depth theoretical analysis of the algorithm's convergence is presented in this article. Experiments are conducted on several typical databases for image classification, and the promising results indicate the effectiveness of RDS.
Liyong Fu, Zechao Li, Qiaolin Ye, Qingwang Liu, Xiaobo Chen 0001, Xijian Fan, Wankou Yang, Guowei Yang 0002
IEEE Trans. Neural Networks Learn. Syst.8
2021 Separable Batch Normalization for Robust Facial Landmark Localization
Shuangping Jin, Zhenhua Feng 0001, Wankou Yang, Josef Kittler
BMVC3
2021 TokenPose: Learning Keypoint Tokens for Human Pose Estimation
abstract
Human pose estimation deeply relies on visual clues and anatomical constraints between parts to locate keypoints. Most existing CNN-based methods do well in visual representation, however, lacking in the ability to explicitly learn the constraint relationships between keypoints. In this paper, we propose a novel approach based on Token representation for human Pose estimation (TokenPose). In detail, each keypoint is explicitly embedded as a token to simultaneously learn constraint relationships and appearance cues from images. Extensive experiments show that the small and large TokenPose models are on par with state-of-the-art CNN-based counterparts while being more lightweight. Specifically, our TokenPose-S and TokenPose-L achieve 72.5 AP and 75.8 AP on COCO validation dataset respectively, with significant reduction in parameters (↓80.6% ; ↓ 56.8%) and GFLOPs (↓ 75.3%; ↓24.7%). Code is publicly available1.
Shoukui Zhang, Zhicheng Wang 0001, Wankou Yang, Shutao Xia, Erjin Zhou
ICCV5
2021 TransPose: Keypoint Localization via Transformer
abstract
While CNN-based models have made remarkable progress on human pose estimation, what spatial dependencies they capture to localize keypoints remains unclear. In this work, we propose a model called Trans-Pose, which introduces Transformer for human pose estimation. The attention layers built in Transformer enable our model to capture long-range relationships efficiently and also can reveal what dependencies the predicted key-points rely on. To predict keypoint heatmaps, the last attention layer acts as an aggregator, which collects contributions from image clues and forms maximum positions of keypoints. Such a heatmap-based localization approach via Transformer conforms to the principle of Activation Maximization [19]. And the revealed dependencies are image-specific and fine-grained, which also can provide evidence of how the model handles special cases, e.g., occlusion. The experiments show that TransPose achieves 75.8 AP and 75.0 AP on COCO validation and test-dev sets, while being more lightweight and faster than mainstream CNN architectures. The TransPose model also transfers very well on MPII benchmark, achieving superior performance on the test set when fine-tuned with small training costs. Code and pre-trained models are publicly available1.
Zhibin Quan, Mu Nie, Wankou Yang
ICCV4
2021 OPLS-SR: A novel face super-resolution learning method using orthonormalized coherent features
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jipeng Qiang, Bin Li 0006, Wankou Yang, Furong Peng
Inf. Sci.6
2021 An Interconnected Feature Pyramid Networks for object detection
Qiang Wang 0023, Lukuan Zhou, Yuncong Yao, Yong Wang 0032, Jun Li 0033, Wankou Yang
J. Vis. Commun. Image Represent.6
2021 Beyond ITQ: Efficient binary multi-view subspace learning for instance retrieval
Zhijian Wu, Jun Li 0033, Wankou Yang
J. Vis. Commun. Image Represent.4
2021 Knowledge memorization and generation for action recognition in still images
Wankou Yang, Yazhou Yao, Fatih Porikli
Pattern Recognit.2
2021 Robust gait recognition using hybrid descriptors based on Skeleton Gait Energy Image
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang, Wankou Yang
Pattern Recognit. Lett.6
2021 Unsupervised Eyeglasses Removal in the Wild
abstract
Eyeglasses removal is challenging in removing different kinds of eyeglasses, e.g., rimless glasses, full-rim glasses, and sunglasses, and recovering appropriate eyes. Due to the significant visual variants, the conventional methods lack scalability. Most existing works focus on the frontal face images in the controlled environment, such as the laboratory, and need to design specific systems for different eyeglass types. To address the limitation, we propose a unified eyeglass removal model called the eyeglasses removal generative adversarial network (ERGAN), which could handle different types of glasses in the wild. The proposed method does not depend on the dense annotation of eyeglasses location but benefits from the large-scale face images with weak annotations. Specifically, we study the two relevant tasks simultaneously, that is, removing eyeglasses and wearing eyeglasses. Given two face images with and without eyeglasses, the proposed model learns to swap the eye area in two faces. The generation mechanism focuses on the eye area and invades the difficulty of generating a new face. In the experiment, we show the proposed method achieves a competitive removal quality in terms of realism and diversity. Furthermore, we evaluate ERGAN on several subsequent tasks, such as face verification and facial expression recognition. The experiment shows that our method could serve as a preprocessing method for these tasks.
Bingwen Hu, Zhedong Zheng, Ping Liu 0004, Wankou Yang, Mingwu Ren
IEEE Trans. Cybern.4
2021 Subspace-based multi-view fusion for instance-level image retrieval
Jun Li 0033, Bo Yang 0019, Wankou Yang, Changyin Sun 0001
Vis. Comput.3
2020 Multi-model Network for Fine-Grained Cross-Media Retrieval
Jiemi Bai, Yazhou Yao, Qiong Wang 0003, Wankou Yang, Fumin Shen
PRCV (2)5
2020 A Novel CNN Architecture for Real-Time Point Cloud Recognition in Road Environment
Duyao Fan, Yazhou Yao, Yunfei Cai, Xiangbo Shu, Wankou Yang
PRCV (1)6
2020 Hierarchical Representations with Discriminative Meta-filters in Dual Path Network for Tracking
Ning Wang 0020, Yuncong Yao, Wankou Yang, Kaihua Zhang 0001, Bo Liu 0005
PRCV (2)4
2020 Multilayer deep features with multiple kernel learning for action recognition
Biyun Sheng, Fu Xiao 0001, Wankou Yang
Neurocomputing4
2020 A local multiple patterns feature descriptor for face recognition
Wankou Yang, Jun Li 0011
Neurocomputing1
2020 Face image super-resolution with pose via nuclear norm regularized structural orthogonal Procrustes regression
Guangwei Gao, Meng Yang 0001, Huimin Lu 0001, Wankou Yang, Hao Gao 0005
Neural Comput. Appl.5
2020 Inverse Visual Question Answering: A New Benchmark and VQA Diagnosis Tool
abstract
In recent years, visual question answering (VQA) has become topical. The premise of VQA's significance as a benchmark in AI, is that both the image and textual question need to be well understood and mutually grounded in order to infer the correct answer. However, current VQA models perhaps 'understand' less than initially hoped, and instead master the easier task of exploiting cues given away in the question and biases in the answer distribution [1]. In this paper we propose the inverse problem of VQA (iVQA). The iVQA task is to generate a question that corresponds to a given image and answer pair. We propose a variational iVQA model that can generate diverse, grammatically correct and content correlated questions that match the given answer. Based on this model, we show that iVQA is an interesting benchmark for visuo-linguistic understanding, and a more challenging alternative to VQA because an iVQA model needs to understand the image better to be successful. As a second contribution, we show how to use iVQA in a novel reinforcement learning framework to diagnose any existing VQA model by way of exposing its belief set: the set of question-answer pairs that the VQA model would predict true for a given image. This provides a completely new window into what VQA models 'believe' about images. We show that existing VQA models have more erroneous beliefs than previously thought, revealing their intrinsic weaknesses. Suggestions are then made on how to address these weaknesses going forward.
Feng Liu 0036, Tao Xiang 0002, Timothy M. Hospedales, Wankou Yang, Changyin Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Advanced deep learning for image super-resolution
Pourya Shamsolmoali, Abdul Hamid Sadka, Huiyu Zhou 0001, Wankou Yang
Signal Process. Image Commun.4
2020 Discriminative Multi-View Subspace Feature Learning for Action Recognition
abstract
Although deep features have achieved the state-of-the-art performance in action recognition recently, the hand-crafted shallow features still play a critical role in characterizing human actions for taking advantage of visual contents in an intuitive way such as edge features. Therefore, the shallow features can serve as auxiliary visual cues supplementary to deep representations. In this paper, we propose a discriminative subspace learning model (DSLM) to explore the complementary properties between the hand-crafted shallow feature representations and the deep features. As for the RGB action recognition, this is the first work attempting to mine multi-level feature complementaries by the multi-view subspace learning scheme. To sufficiently capture the complementary information among heterogeneous features, we construct the DSLM by integrating the multi-view reconstruction error and classification error into an unified objective function. To be specific, we first use Fisher Vector to encode improved dense trajectories (iDT+FV) for shallow representations and two-stream convolutional neural network models (T-CNN) for generating deep features. Moreover, the presented DSLM algorithm projects multi-level features onto a shared discriminative subspace with the complementary information and discriminating capacity simultaneously incorporated. Finally, the action types of test samples are identified by the margins from the learned compact representations to the decision boundary. The experimental results on three datasets demonstrate the effectiveness of the proposed method.
Biyun Sheng, Jun Li 0033, Fu Xiao 0001, Qun Li 0002, Wankou Yang, Junwei Han 0001
IEEE Trans. Circuits Syst. Video Technol.5
2020 Discriminative Multi-View Privileged Information Learning for Image Re-Ranking
abstract
Conventional multi-view re-ranking methods usually perform asymmetrical matching between the region of interest (ROI) in the query image and the whole target image for similarity computation. Due to the inconsistency in the visual appearance, this practice tends to degrade the retrieval accuracy particularly when the image ROI, which is usually interpreted as the image objectness, accounts for a smaller region in the image. Since Privileged Information (PI), which can be viewed as the image prior, is able to characterize well the image objectness, we are aiming at leveraging PI for further improving the performance of multi-view re-ranking in this paper. Towards this end, we propose a discriminative multi-view re-ranking approach in which both the original global image visual contents and the local auxiliary PI features are simultaneously integrated into a unified training framework for generating the latent subspaces with sufficient discriminating power. For the on-the-fly re-ranking, since the multi-view PI features are unavailable, we only project the original multi-view image representations onto the latent subspace, and thus the re-ranking can be achieved by computing and sorting the distances from the multi-view embeddings to the separating hyperplane. Extensive experimental evaluations on the two public benchmarks, Oxford5k and Paris6k, reveal that our approach provides further performance boost for accurate image re-ranking, whilst the comparative study demonstrates the advantage of our method against other multi-view re-ranking methods.
Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001, Hong Zhang 0013
IEEE Trans. Image Process.3
2020 The Structure Transfer Machine Theory and Applications
abstract
Representation learning is a fundamental but challenging problem, especially when the distribution of data is unknown. In this paper, we propose a new representation learning method, named Structure Transfer Machine (STM), which enables feature learning process to converge at the representation expectation in a probabilistic way. We theoretically show that such an expected value of the representation (mean) is achievable if the manifold structure can be transferred from the data space to the feature space. The resulting structure regularization term, named manifold loss, is incorporated into the loss function of the typical deep learning pipeline. The STM architecture is constructed to enforce the learned deep representation to satisfy the intrinsic manifold structure from the data, which results in robust features that suit various application scenarios, such as digit recognition, image classification and object tracking. Compared with state-of-the-art CNN architectures, we achieve better results on several commonly used public benchmarks.
Baochang Zhang 0001, Wankou Yang, Ze Wang 0008, Lian Zhuo, Jungong Han, Xiantong Zhen
IEEE Trans. Image Process.2
2020 Recurrent Neural Networks With External Addressable Long-Term and Working Memory for Learning Long-Term Dependences
abstract
Learning long-term dependences (LTDs) with recurrent neural networks (RNNs) is challenging due to their limited internal memories. In this paper, we propose a new external memory architecture for RNNs called an external addressable long-term and working memory (EALWM)-augmented RNN. This architecture has two distinct advantages over existing neural external memory architectures, namely the division of the external memory into two parts-long-term memory and working memory-with both addressable and the capability to learn LTDs without suffering from vanishing gradients with necessary assumptions. The experimental results on algorithm learning, language modeling, and question answering demonstrate that the proposed neural memory architecture is promising for practical applications.
Zhibin Quan, Yandong Liu 0002, Yunxiu Yu, Wankou Yang
IEEE Trans. Neural Networks Learn. Syst.6
2020 A Probabilistic Zero-Shot Learning Method via Latent Nonnegative Prototype Synthesis of Unseen Classes
abstract
Zero-shot learning (ZSL), a type of structured multioutput learning, has attracted much attention due to its requirement of no training data for target classes. Conventional ZSL methods usually project visual features into semantic space and assign labels by finding their nearest prototypes. However, this type of nearest neighbor search (NNS)-based method often suffers from great performance degradation because of the nonuniform variances between different categories. In this article, we propose a probabilistic framework by taking covariance into account to deal with the above-mentioned problem. In this framework, we define a new latent space, which has two characteristics. The first is that the features in this space should gather within the classes and scatter between the classes, which is implemented by triplet learning; the second is that the prototypes of unseen classes are synthesized with nonnegative coefficients, which are generated by nonnegative matrix factorization (NMF) of relations between the seen classes and the unseen classes in attribute space. During training, the learned parameters are the projection model for triplet network and the nonnegative coefficients between the unseen classes and the seen classes. In the testing phase, visual features are projected into latent space and assigned with the labels that have the maximum probability among unseen classes for classic ZSL or within all classes for generalized ZSL. Extensive experiments are conducted on four popular data sets, and the results show that the proposed method can outperform the state-of-the-art methods in most circumstances.
Haofeng Zhang 0001, Huaqi Mao, Yang Long 0001, Wankou Yang, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2019 Clustering-driven unsupervised deep hashing for image retrieval
Haofeng Zhang 0001, Yazhou Yao, Wankou Yang, Li Liu 0004
Neurocomputing5
2019 Dual-verification network for zero-shot learning
Haofeng Zhang 0001, Yang Long 0001, Wankou Yang, Ling Shao 0001
Inf. Sci.3
2019 Exploiting aggregate channel features for urine sediment detection
Changyin Sun 0001, Wankou Yang
Multim. Tools Appl.4
2019 A Face Detection Method Based on Cascade Convolutional Neural Network
Wankou Yang, Lukuan Zhou, Tianhuang Li, Haoran Wang 0001
Multim. Tools Appl.1
2019 Pattern Recognition Techniques for Non Verbal Human Behavior (NVHB)
Wankou Yang, Pritee Khanna, Xiong Li 0002
Pattern Recognit. Lett.1
2019 Exploiting textual and visual features for image categorization
Yazhou Yao, Wankou Yang, Qiong Wang 0003, Yunfei Cai, Zhenmin Tang
Pattern Recognit. Lett.2
2019 Differential Features for Pedestrian Detection: A Taylor Series Perspective
abstract
Differential features are popularly used in computer vision tasks, such as object detection. In this paper, we revisit these features from a functional approximation perspective. In particular, we view an image as a 2-D functional and investigate its Taylor series approximation. Differential features are derived from the approximation coefficients and, therefore, are naturally collected for appearance representation. Thus motivated, we propose to use the zeroth-, first-, and second-order differential features for pedestrian detection and call such features Taylor feature transform (TAFT). In practice, the TAFT features are computed by discrete sampling to address scale issues and meanwhile achieve computational efficiency. In addition, orientation insensitivity is handled by using directional versions of differentials. When applied to pedestrian detection, the TAFT is sampled on grid pixels and calculated from multiple channels following previous solutions. In our extensive experiments on the INRIA, Caltech, TUD-Brussel, and KITTI data sets, the TAFT achieves state-of-the-art results. It outperforms all handcrafted features and performs on par with many deep-learning solutions. Moreover, when a low false-positive rate is requested, the TAFT generates results that are better than or comparable to the state-of-the-art deep learning-based methods. Meanwhile, our implementation runs at 33 fps for 640×480 images without GPU, making TAFT favorable in many practical scenarios.
Jifeng Shen, Wankou Yang, Danil V. Prokhorov, Xue Mei, Haibin Ling
IEEE Trans. Intell. Transp. Syst.3
2019 Pedestrian Proposal and Refining Based on the Shared Pixel Differential Feature
abstract
We design a pedestrian proposal and refining system tailored for fast pedestrian detection. The pedestrian proposal is based on pixel differential feature (PDF), which is a light weighted feature with a high recall rate. For the pedestrian refining, we propose an aggregated region feature (ARF) to distill the co-existing dominant pixel differential patterns in a local region to reject hard false positives. Albeit discriminative, ARF largely relies on the size of anchored regions and the scale of the PDF, which hinders its performance in real-world applications. Although multi-scale PDF with spatial pyramid somewhat alleviates this problem, it is computationally expensive and thus infeasible in practice. To address this issue, we further propose a directional radius pooling method to extract discriminative information in each orientation of PDF while reducing the feature dimensionality with a more compact size. The pedestrian proposal and refining framework is built on the shared pixel differential feature map which is very computationally efficient. More specifically, a set of pedestrian proposals generated from the single-scale PDF are first obtained in images. Second, multi-scale ARF in spatial pyramid is used to fuse information from different scales and spatial resolutions for anchor regions. Third, the directional radius pooling method is proposed to extract dominant information of each orientation in the anchor regions. The pedestrian proposal and refining are finally integrated for accurate pedestrian detection. The extensive experimental evaluations on five public benchmarks show that our method achieves state-of-the-art results while running at 18 fps for $480\times640$ images.
Jifeng Shen, Lei Zhu 0010, Jun Li 0033, Wankou Yang, Haibin Ling
IEEE Trans. Intell. Transp. Syst.5
2019 ROMIR: Robust Multi-View Image Re-Ranking
abstract
In multi-view re-ranking, multiple heterogeneous visual features are usually projected onto a low-dimensional subspace, and thus the resulting latent representation can be used for the subsequent similarity-based ranking. Albeit effective, this standard mechanism underplays the intrinsic structure underlying the latent subspace and does not take into account the substantial noise in the original spaces. In this paper, we propose a robust multi-view image re-ranking strategy. Due to the dramatic variability in image visual appearance, it is necessary to uncover the shared components underlying those query-related instances that are visually unlike for improving the re-ranking accuracy. Consequently, it is reasonable to assume the latent subspace enjoys the low-rank property and thus the subspace recovery can be achieved via the low-rank modeling accordingly. In addition, since the real-world data are usually partially contaminated, we employ `2;1-norm based sparsity constraint to appropriately model the sample-specific mapping noise for enhancing the model robustness. In order to produce discriminative representations, we encode a similarity preserving term in our multi-view embedding framework. As a result, the sample separability is maximally maintained in the latent subspace with sufficient discriminative power. The extensive evaluations on public landmark benchmarks demonstrate the efficacy and superiority of the proposed method.
Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001, Kotagiri Ramamohanarao, Dacheng Tao
IEEE Trans. Knowl. Data Eng.3
2019 Nonpeaked Discriminant Analysis for Data Representation
abstract
Of late, there are many studies on the robust discriminant analysis, which adopt L1-norm as the distance metric, but their results are not robust enough to gain universal acceptance. To overcome this problem, the authors of this article present a nonpeaked discriminant analysis (NPDA) technique, in which cutting L1-norm is adopted as the distance metric. As this kind of norm can better eliminate heavy outliers in learning models, the proposed algorithm is expected to be stronger in performing feature extraction tasks for data representation than the existing robust discriminant analysis techniques, which are based on the L1-norm distance metric. The authors also present a comprehensive analysis to show that cutting L1-norm distance can be computed equally well, using the difference between two special convex functions. Against this background, an efficient iterative algorithm is designed for the optimization of the proposed objective. Theoretical proofs on the convergence of the algorithm are also presented. Theoretical insights and effectiveness of the proposed method are validated by experimental tests on several real data sets.
Qiaolin Ye, Zechao Li, Liyong Fu, Zhao Zhang 0001, Wankou Yang, Guowei Yang 0002
IEEE Trans. Neural Networks Learn. Syst.5
2019 ℓ1-Norm Heteroscedastic Discriminant Analysis Under Mixture of Gaussian Distributions
abstract
Fisher’s criterion is one of the most popular discriminant criteria for feature extraction. It is defined as the generalized Rayleigh quotient of the between-class scatter distance to the within-class scatter distance. Consequently, Fisher’s criterion does not take advantage of the discriminant information in the class covariance differences, and hence, its discriminant ability largely depends on the class mean differences. If the class mean distances are relatively large compared with the within-class scatter distance, Fisher’s criterion-based discriminant analysis methods may achieve a good discriminant performance. Otherwise, it may not deliver good results. Moreover, we observe that the between-class distance of Fisher’s criterion is based on the$\ell _{2}$-norm, which would be disadvantageous to separate the classes with smaller class mean distances. To overcome the drawback of Fisher’s criterion, in this paper, we first derive a new discriminant criterion, expressed as amixture of absolute generalized Rayleigh quotients, based on a Bayes error upper bound estimation, where mixture of Gaussians is adopted to approximate the real distribution of data samples. Then, the criterion is further modified by replacing$\ell _{2}$-norm with$\ell _{1}$one to better describe the between-class scatter distance, such that it would be more effective to separate the different classes. Moreover, we propose a novel$\ell _{1}$-norm heteroscedastic discriminant analysis method based on the new discriminant analysis (L1-HDA/GM) for heteroscedastic feature extraction, in which the optimization problem of L1-HDA/GM can be efficiently solved by using the eigenvalue decomposition approach. Finally, we conduct extensive experiments on four real data sets and demonstrate that the proposed method achieves much competitive results compared with the state-of-the-art methods.
Wenming Zheng, Cheng Lu 0005, Zhouchen Lin, Tong Zhang 0021, Zhen Cui 0001, Wankou Yang
IEEE Trans. Neural Networks Learn. Syst.6
2018 Discovering and Distinguishing Multiple Visual Senses for Polysemous Words
abstract
To reduce the dependence on labeled data, there have been increasing research efforts on learning visual classifiers by exploiting web images. One issue that limits their performance is the problem of polysemy. To solve this problem, in this work, we present a novel framework that solves the problem of polysemy by allowing sense-specific diversity in search results. Specifically, we first discover a list of possible semantic senses to retrieve sense-specific images. Then we merge visual similar semantic senses and prune noises by using the retrieved images. Finally, we train a visual classifier for each selected semantic sense and use the learned sense-specific classifiers to distinguish multiple visual senses. Extensive experiments on classifying images into sense-specific categories and re-ranking search results demonstrate the superiority of our proposed approach.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Wankou Yang, Zhenmin Tang
AAAI4
2018 IVQA: Inverse Visual Question Answering
abstract
We propose the inverse problem of Visual question answering (iVQA), and explore its suitability as a benchmark for visuo-linguistic understanding. The iVQA task is to generate a question that corresponds to a given image and answer pair. Since the answers are less informative than the questions, and the questions have less learnable bias, an iVQA model needs to better understand the image to be successful than a VQA model. We pose question generation as a multi-modal dynamic inference process and propose an iVQA model that can gradually adjust its focus of attention guided by both a partially generated question and the answer. For evaluation, apart from existing linguistic metrics, we propose a new ranking metric. This metric compares the ground truth question's rank among a list of distractors, which allows the drawbacks of different algorithms and sources of error to be studied. Experimental results show that our model can generate diverse, grammatically correct and content correlated questions that match the given answer.
Feng Liu 0036, Tao Xiang 0002, Timothy M. Hospedales, Wankou Yang, Changyin Sun 0001
CVPR4
2018 Extracting Privileged Information from Untagged Corpora for Classifier Learning
abstract
The performance of data-driven learning approaches is often unsatisfactory when the training data is inadequate either in quantity or quality. Manually labeled privileged information (PI), \eg attributes, tags or properties, is usually incorporated to improve classifier learning. However, the process of manually labeling is time-consuming and labor-intensive. To address this issue, we propose to enhance classifier learning by extracting PI from untagged corpora, which can effectively eliminate the dependency on manually labeled data. In detail, we treat each selected PI as a subcategory and learn one classifier for per subcategory independently. The classifiers for all subcategories are then integrated together to form a more powerful category classifier. Particularly, we propose a new instance-level multi-instance learning (MIL) model to simultaneously select a subset of training images from each subcategory and learn the optimal classifiers based on the selected images. Extensive experiments demonstrate the superiority of our approach.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Wankou Yang, Xian-Sheng Hua 0001, Zhenmin Tang
IJCAI4
2018 Gabor Convolutional Networks
abstract
Steerable properties dominate the design of traditional filters, e.g., Gabor filters, and endow features the capability of dealing with spatial transformations. However, such excellent properties have not been well explored in the popular deep convolutional neural networks (DCNNs). In this paper, we propose a new deep model, termed Gabor Convolutional Networks (GCNs or Gabor CNNs), which incorporates Gabor filters into DCNNs to enhance the resistance of deep learned features to the orientation and scale changes. By only manipulating the basic element of DCNNs based on Gabor filters, i.e., the convolution operator, GCNs can be easily implemented and are compatible with any popular deep learning architecture. Experimental results demonstrate the super capability of our algorithm in recognizing objects, where the scale and rotation changes occur frequently. The proposed GCNs have much fewer learnable network parameters, and thus is easier to train with an endtoend pipeline. The source code will be here1.
Shangzhen Luan, Baochang Zhang 0001, Siyue Zhou, Chen Chen 0001, Jungong Han, Wankou Yang, Jianzhuang Liu
WACV6
2018 Action unit detection and key frame selection for human activity prediction
Haoran Wang 0001, Chunfeng Yuan, Jifeng Shen, Wankou Yang, Haibin Ling
Neurocomputing4
2018 PCANet: An energy perspective
Jiasong Wu, Shijie Qiu, Youyong Kong, Longyu Jiang, Yang Chen 0008, Wankou Yang, Lotfi Senhadji, Huazhong Shu
Neurocomputing6
2018 Crowd Counting via Weighted VLAD on a Dense Attribute Feature Map
abstract
Crowd counting is an important task in computer vision, which has many applications in video surveillance. Although the regression-based framework has achieved great improvements for crowd counting, how to improve the discriminative power of image representation is still an open problem. Conventional holistic features used in crowd counting often fail to capture semantic attributes and spatial cues of the image. In this paper, we propose integrating semantic information into learning locality-aware feature (LAF) sets for accurate crowd counting. First, with the help of a convolutional neural network, the original pixel space is mapped onto a dense attribute feature map, where each dimension of the pixelwise feature indicates the probabilistic strength of a certain semantic class. Then, LAF built on the idea of spatial pyramids on neighboring patches is proposed to explore more spatial context and local information. Finally, the traditional vector of locally aggregated descriptor (VLAD) encoding method is extended to a more generalized form weighted-VLAD (W-VLAD) in which diverse coefficient weights are taken into consideration. Experimental results validate the effectiveness of our presented method.
Biyun Sheng, Chunhua Shen, Guosheng Lin, Jun Li 0033, Wankou Yang, Changyin Sun 0001
IEEE Trans. Circuits Syst. Video Technol.5
2017 Semantic Regularisation for Recurrent Image Annotation
Feng Liu 0036, Tao Xiang 0002, Timothy M. Hospedales, Wankou Yang, Changyin Sun 0001
CVPR4
2017 Robust optimal control for time-delay systems with dynamic uncertainties via ADP
abstract
This paper considers a robust optimal control design for a class of nonlinear discrete-time systems with unknown time-varying delays and dynamic uncertainties. An iterative control strategy based on adaptive dynamic programming (ADP) has been proposed. Neural networks are applied to realize the state prediction, the control input estimation and the performance index function approximation. The estimated control input and performance index function are updated iteratively. Furthermore, it has been proven that the approximated performance index function can converge to the optimal solution of the Hamilton-Jacobia-Bellman (HJB) equation. Finally, the proposed algorithm has been conducted to a numerical simulation. The simulation results demonstrate the effectiveness of the new design.
Lu Dong 0002, Jun Li 0033, Wankou Yang, Changyin Sun 0001
IJCNN3
2017 SPA: Spatially Pooled Attributes for image retrieval
Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001
Neurocomputing3
2017 Filtered shallow-deep feature channels for pedestrian detection
Biyun Sheng, Qichang Hu, Jun Li 0033, Wankou Yang, Baochang Zhang 0001, Changyin Sun 0001
Neurocomputing4
2017 Human activity prediction using temporally-weighted generalized time warping
Haoran Wang 0001, Wankou Yang, Chunfeng Yuan, Haibin Ling, Weiming Hu 0004
Neurocomputing2
2017 Gender classification using 3D statistical models
Wankou Yang, Changyin Sun 0001, Wenming Zheng, Karl Ricanek
Multim. Tools Appl.1
2017 Supervised Local High-Order Differential Channel Feature Learning for Pedestrian Detection
Jifeng Shen, Haoran Wang 0001, Wankou Yang, Chengshan Qian
Neural Process. Lett.5
2017 Learning robust and discriminative low-rank representations for face recognition with occlusion
Guangwei Gao, Jian Yang 0003, Xiaoyuan Jing, Fumin Shen, Wankou Yang, Dong Yue 0001
Pattern Recognit.5
2017 A novel pixel neighborhood differential statistic feature for pedestrian and face detection
Jifeng Shen, Jun Li 0033, Wankou Yang, Haibin Ling
Pattern Recognit.4
2017 Discriminative Multi-View Interactive Image Re-Ranking
abstract
Given an unreliable visual patterns and insufficient query information, content-based image retrieval is often suboptimal and requires image re-ranking using auxiliary information. In this paper, we propose a discriminative multi-view interactive image re-ranking (DMINTIR), which integrates user relevance feedback capturing users' intentions and multiple features that sufficiently describe the images. In DMINTIR, heterogeneous property features are incorporated in the multi-view learning scheme to exploit their complementarities. In addition, a discriminatively learned weight vector is obtained to reassign updated scores and target images for re-ranking. Compared with other multi-view learning techniques, our scheme not only generates a compact representation in the latent space from the redundant multi-view features but also maximally preserves the discriminative information in feature encoding by the large-margin principle. Furthermore, the generalization error bound of the proposed algorithm is theoretically analyzed and shown to be improved by the interactions between the latent space and discriminant function learning. Experimental results on two benchmark data sets demonstrate that our approach boosts baseline retrieval quality and is competitive with the other state-of-the-art re-ranking strategies.
Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001, Dacheng Tao
IEEE Trans. Image Process.3
2016 SERVE: Soft and Equalized Residual VEctors for image retrieval
Jun Li 0033, Chang Xu 0002, Mingming Gong, Junliang Xing, Wankou Yang, Changyin Sun 0001
Neurocomputing5
2016 Discriminative low-rank dictionary learning for face recognition
Hoangvu Nguyen, Wankou Yang, Biyun Sheng, Changyin Sun 0001
Neurocomputing2
2016 Learning discriminative shape statistics distribution features for pedestrian detection
Jifeng Shen, Wankou Yang, Hualong Yu, Guohai Liu
Neurocomputing3
2016 A regularized least square based discriminative projections for feature extraction
Wankou Yang, Changyin Sun 0001, Wenming Zheng
Neurocomputing1
2016 Face recognition using adaptive local ternary patterns method
Wankou Yang, Baochang Zhang 0001
Neurocomputing1
2016 Sparsity analysis versus sparse representation classifier
Baochang Zhang 0001, Suli Ji, Shengping Zhang, Wankou Yang
Neurocomputing5
2016 ODOC-ELM: Optimal decision outputs compensation-based extreme learning machine for classifying imbalanced data
Hualong Yu, Changyin Sun 0001, Xibei Yang, Wankou Yang, Jifeng Shen, Yunsong Qi
Knowl. Based Syst.4
2016 Robust regression based face recognition with fast outlier removal
Fumin Shen, Wankou Yang, Hanwang Zhang, Heng Tao Shen
Multim. Tools Appl.2
2016 Face Recognition Using A Low Rank Representation Based Projections Method
Wankou Yang, Fumin Shen
Neural Process. Lett.2
2016 A novel label learning algorithm for face recognition
Shigang Liu, Xianye Ben, Wankou Yang, Guoyong Qiu
Signal Process.4
2015 A supervised dictionary learning and discriminative weighting model for action recognition
Changyin Sun 0001, Wankou Yang
Neurocomputing3
2015 Kernel Low-Rank Representation for face recognition
Hoangvu Nguyen, Wankou Yang, Fumin Shen, Changyin Sun 0001
Neurocomputing2
2015 Action recognition using direction-dependent feature pairs and non-negative low rank sparse model
Biyun Sheng, Wankou Yang, Changyin Sun 0001
Neurocomputing2
2015 AL-ELM: One uncertainty-based active learning algorithm using extreme learning machine
Hualong Yu, Changyin Sun 0001, Wankou Yang, Xibei Yang
Neurocomputing3
2015 Support vector machine-based optimized decision threshold adjustment strategy for classifying imbalanced data
Hualong Yu, Chaoxu Mu, Changyin Sun 0001, Wankou Yang, Xibei Yang
Knowl. Based Syst.4
2015 Low-resolution degradation face recognition over long distance based on CCA
Wankou Yang, Xianye Ben
Neural Comput. Appl.2
2015 A collaborative representation based projections method for feature extraction
Wankou Yang, Changyin Sun 0001
Pattern Recognit.1
2014 Spatial modeling via feature co-pooling and SG grafting
Feng Liu 0036, Yongzhen Huang, Liang Wang 0001, Wankou Yang, Changyin Sun 0001
Neurocomputing4
2014 Image classification using local linear regression
Wankou Yang, Karl Ricanek, Fumin Shen
Neural Comput. Appl.1
2014 Action Recognition Using Nonnegative Action Component Representation and Sparse Basis Selection
abstract
In this paper, we propose using high-level action units to represent human actions in videos and, based on such units, a novel sparse model is developed for human action recognition. There are three interconnected components in our approach. First, we propose a new context-aware spatial-temporal descriptor, named locally weighted word context, to improve the discriminability of the traditionally used local spatial-temporal descriptors. Second, from the statistics of the context-aware descriptors, we learn action units using the graph regularized nonnegative matrix factorization, which leads to a part-based representation and encodes the geometrical information. These units effectively bridge the semantic gap in action recognition. Third, we propose a sparse model based on a joint l2,1-norm to preserve the representative items and suppress noise in the action units. Intuitively, when learning the dictionary for action representation, the sparse model captures the fact that actions from the same class share similar units. The proposed approach is evaluated on several publicly available data sets. The experimental results and analysis clearly demonstrate the effectiveness of the proposed approach.
Haoran Wang 0001, Chunfeng Yuan, Weiming Hu 0004, Haibin Ling, Wankou Yang, Changyin Sun 0001
IEEE Trans. Image Process.5
2013 Real-time human detection based on gentle MILBoost with variable granularity HOG-CSLBP
Jifeng Shen, Wankou Yang, Changyin Sun 0001
Neural Comput. Appl.2
2013 Face recognition using fuzzy maximum scatter discriminant analysis
Jianguo Wang 0002, Wankou Yang, Jing-Yu Yang 0001
Neural Comput. Appl.2
2012 Large margin null space discriminant analysis with applications to face recognition
Xiaobo Chen 0001, Jian Yang 0003, Wankou Yang
ICPR3
2012 Kernel sparse representation based classification
Jun Yin 0003, Zhong Jin, Wankou Yang
Neurocomputing4
2012 Sequential Row-Column 2DPCA for face recognition
Wankou Yang, Changyin Sun 0001, Karl Ricanek
Neural Comput. Appl.1
2012 Feature extraction using 2DIFDA with fuzzy membership
Zhongxi Sun, Changyin Sun 0001, Wankou Yang, Jifeng Shen
Soft Comput.3
2011 Facial feature fusion and model selection for age estimation
abstract
Automatic face age estimation is challenging due to its complexity owing to genetic difference, behavior and environmental factors, the dynamics of facial aging between different individuals, etc. In this work we propose to fuse the global facial feature extracted from Active Appearance Model (AAM) and the local facial features extracted from Local Binary Pattern (LBP), as the representation of faces. Furthermore, we introduce an advanced age estimation system combining feature fusion and model selection schemes such as Least Angle Regression (LAR) and sequential approaches. Due to the fact that different facial feature representations may come with various types of measurement scales, we compare multiple normalization schemes for both facial features. We demonstrate that the feature fusion with model selection can achieve significant improvement in age estimation over single feature representation alone. Our experiment on multi-ethnicity UIUC-PAL database suggests that age estimation with feature fusion and model selection outperforms the single feature, or the full feature model.
Cuixian Chen, Wankou Yang, Yishi Wang, Karl Ricanek, Khoa Luu
FG2
2011 Fast Human Detection Based on Enhanced Variable Size HOG Features
Jifeng Shen, Changyin Sun 0001, Wankou Yang, Zhongxi Sun
ISNN (2)3
2011 Finger-Knuckle-Print Recognition Using LGBP
Ming Xiong, Wankou Yang, Changyin Sun 0001
ISNN (2)2
2011 Ensemble of Global and Local Features for Face Age Estimation
Wankou Yang, Cuixian Chen, Karl Ricanek, Changyin Sun 0001
ISNN (2)1
2011 Gender Classification Using the Profile
Wankou Yang, Amrutha Sethuram, Eric Patterson, Karl Ricanek, Changyin Sun 0001
ISNN (2)1
2011 A novel distribution-based feature for rapid object detection
Jifeng Shen, Changyin Sun 0001, Wankou Yang, Zhongxi Sun
Neurocomputing3
2011 Feature Extraction Using Laplacian Maximum Margin Criterion
Wankou Yang, Changyin Sun 0001, Helen S. Du, Jing-Yu Yang 0001
Neural Process. Lett.1
2011 Face Recognition Using Kernel UDP
Wankou Yang, Changyin Sun 0001, Jing-Yu Yang 0001, Helen S. Du, Karl Ricanek
Neural Process. Lett.1
2011 A multi-manifold discriminant analysis method for image feature extraction
Wankou Yang, Changyin Sun 0001, Lei Zhang 0006
Pattern Recognit.1
2010 Learning Discriminative Features Based on Distribution
abstract
In this paper, a novel feature named adaptive projection LBP (APLBP) is proposed for face detection. To promote discriminative power, the distribution information of training samples is embedded into the proposed feature. APLBP is generated by LDA which maximizes the margin between positive and negative samples adaptively, utilizing characteristics of similarity to Gaussian distribution of the training samples. Asymmetric Gentle Adaboost is utilized to train strong classifier and nested cascade is applied to construct the final detector. Experimental results based on MIT+CMU database demonstrate that APLBP feature outperforms several well-existing features due to its excellent discriminative power with less feature number.
Jifeng Shen, Wankou Yang, Changyin Sun 0001
ICPR2
2010 Face Recognition Using a Multi-manifold Discriminant Analysis Method
abstract
In this paper, we propose a Multi-Manifold Discriminant Analysis (MMDA) method for face feature extraction and face recognition, which is based on graph embedded learning and under the Fisher discriminant analysis framework. In MMDA, the within-class graph and between-class graph are designed to characterize the within-class compactness and the between-class separability, respectively, seeking for the discriminant matrix that simultaneously maximizing the between-class scatter and minimizing the within-class scatter. In addition, the within-class graph can also represent the sub-manifold information and the between-class graph can also represent the multi-manifold information. The proposed MMDA is examined by using the FERET face database, and the experimental results demonstrate that MMDA works well in feature extraction and lead to good recognition performance.
Wankou Yang, Changyin Sun 0001, Lei Zhang 0006
ICPR1
2010 Laplacian bidirectional PCA for face recognition
Wankou Yang, Changyin Sun 0001, Lei Zhang 0006, Karl Ricanek
Neurocomputing1
2010 Feature extraction based on fuzzy 2DLDA
Wankou Yang, Xiaoyong Yan, Lei Zhang 0006, Changyin Sun 0001
Neurocomputing1
2009 Discriminant feature extraction based on center distance
abstract
In this paper, a novel discriminant feature extraction algorithm employing center-based distance is proposed for face recognition. This new method, which is a supervised linear dimensionality reduction and feature extraction approach, computes the center-based distance between each training sample-pairs in the same class and the distance between each training sample-pair belonging to different classes. Then the high-dimensional data are embedded into a low-dimensional space, preserving the within-class geometric structure on a submanifold via maximum variance projection. Many experiments on ORL and Yale face database indicate that this method is highly effective.
Wankou Yang, Jian Yang 0003, Jing-Yu Yang 0001
ICIP2
2009 Feature extraction using fuzzy inverse FDA
Wankou Yang, Jianguo Wang 0002, Mingwu Ren, Lei Zhang 0006, Jing-Yu Yang 0001
Neurocomputing1
2009 Feature extraction based on Laplacian bidirectional maximum margin criterion
Wankou Yang, Jianguo Wang 0002, Mingwu Ren, Jing-Yu Yang 0001, Lei Zhang 0006, Guanghai Liu 0001
Pattern Recognit.1
2008 Fuzzy maximum scatter discriminant analysis and its application to face recognition
abstract
In this paper, a reformative scatter difference discriminant criterion (SDDC) with fuzzy set theory is studied. The scatter difference between between-class and within-class as discriminant criterion is effective to overcome the singularity problem of the within-class scatter matrix due to small sample size problem occurred in classical Fisher discriminant analysis. However, the conventional SDDC assumes the same level of relevance of each sample to the corresponding class. So, a fuzzy maximum scatter difference analysis (FMSDA) algorithm is proposed, in which the fuzzy k-nearest neighbor (FKNN) is implemented to achieve the distribution information of original samples, and this information is utilized to redefine corresponding scatter matrices which are different to the conventional SDDC and effective to extract discriminative features from overlapping (outlier) samples. Experiments conducted on FERET face databases demonstrate the effectiveness of the proposed method.
Jianguo Wang 0002, Wankou Yang, Jing-Yu Yang 0001
ICPR2
2008 Feature Extraction base on Local Maximum Margin Criterion
abstract
Maximum margin criterion (MMC) based feature extraction method is more efficient than LDA for calculating the discriminant vectors since it does not need to calculate the inverse within-class scatter matrix. However, MMC ignores the discriminative information within the local structures of samples. In this paper, we develop a novel criterion to address the issue, namely local maximum margin criterion (Local MMC). We define the total Laplacian matrix, within-class Laplacian matrix and between-class Laplacian matrix using the samples similar weighting. Local MMC gets the discriminant vectors by maximizing the difference between between-class laplacian matrix and within-class laplacian matrix. Experiments on FERET face database show the effectiveness of the proposed local MMC based feature extraction method.
Wankou Yang, Jianguo Wang 0002, Mingwu Ren, Jing-Yu Yang 0001
ICPR1
2008 Face recognition using Complete Fuzzy LDA
abstract
In this paper, we propose a novel method for feature extraction and recognition, namely, complete fuzzy LDA (CFLDA). CFLDA combines the complete LDA and fuzzy set theory. CFLDA redefines the fuzzy between-class scatter matrix and fuzzy within-class scatter matrix that make fully of the distribution of sample and simultaneously extract the irregular discriminative information and regular discriminative information. Experiments on the Yale and FERET face databases show that CFLDA can work well and surpass fuzzy Fisherface.
Wankou Yang, Jianguo Wang 0002, Jing-Yu Yang 0001
ICPR1
2008 Two-directional maximum scatter difference discriminant analysis for face recognition
Jianguo Wang 0002, Wankou Yang, Yusheng Lin, Jing-Yu Yang 0001
Neurocomputing2
2008 Kernel maximum scatter difference based feature extraction and its application to face recognition
Jianguo Wang 0002, Yusheng Lin, Wankou Yang, Jing-Yu Yang 0001
Pattern Recognit. Lett.3
2006 Face Detection Using Binary Template Matching and SVM
Qiong Wang 0003, Wankou Yang, Huan Wang 0013, Jing-Yu Yang 0001, Yu-Jie Zheng
PRICAI2
2005 A new and fast contour-filling algorithm
Mingwu Ren, Wankou Yang, Jing-Yu Yang 0001
Pattern Recognit.2