Shuai Wang 0049

dblp:42/1503-49 · DBLP profile ↗
← Back
24ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0001-7253-785XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 since 2021Computer networks · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Invariant and Discriminative Representations for Cross-Domain Deepfake Detection
Wenzhong Tang, Shijun Gao, Zhenyuan Huang, Shuai Wang 0049
ICIC (15)5
2025 Reliable and Balanced Transfer Learning for Generalized Multimodal Face Anti-Spoofing
abstract
Face Anti-Spoofing (FAS) is essential for securing face recognition systems against presentation attacks. Recent advances in sensor technology and multimodal learning have enabled the development of multimodal FAS systems. However, existing methods often struggle to generalize to unseen attacks and diverse environments due to two key challenges: (1) Modality unreliability, where sensors such as depth and infrared suffer from severe domain shifts, impairing the reliability of cross-modal fusion; and (2) Modality imbalance, where over-reliance on a dominant modality weakens the model's robustness against attacks that affect other modalities. To overcome these issues, we propose MMDG++, a multimodal domain-generalized FAS framework built upon the vision-language model CLIP. In MMDG++, we design the Uncertainty-Guided Cross-Adapter++ (U-Adapter++) to filter out unreliable regions within each modality, enabling more reliable multimodal interactions. Additionally, we introduce Rebalanced Modality Gradient Modulation (ReGrad) for adaptive gradient modulation to balance modality convergence. To further enhance generalization, propose Asymmetric Domain Prompts (ADPs) that leverage CLIP's language priors to learn generalized decision boundaries across modalities. We also develop a novel multimodal FAS benchmark to evaluate generalizability under various deployment conditions. Extensive experiments across this benchmark show our method outperforms state-of-the-art FAS methods, demonstrating superior generalization capability.
Xun Lin, Ajian Liu 0001, Zitong Yu, Rizhao Cai, Shuai Wang 0049, Yi Yu 0011, Jun Wan 0001, Zhen Lei 0001, Xiaochun Cao, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 MorFormer: Morphology-Aware Transformer for Generalized Pavement Crack Segmentation
abstract
Cracks are common on pavements. Accurate crack detection plays a vital role in pavement maintenance. However, cracks have rich and varied morphological features and fine edges, making this task challenging. Additionally, noise factors such as stains, scratches, and complex textures in the pavement background can easily be confused with cracks, increasing the risk of false prediction in the segmentation process. Therefore, we propose Background Morphology Learning (BML) to reconstruct morphological features of the pavement background noise, extract background morphological dissimilarity maps to suppress interference and reduce false alarms. In addition, we propose Crack Morphology-aware Attention (CMA), which adaptively learns the morphological shape of cracks and dynamically adjusts the shape of the attention receptive field to the topological features of the cracks. This significantly improves the completeness of segmentation. Our method mitigates the problems of false alarms and incomplete segmentation results in the crack segmentation task. Therefore, we propose a Morphology-Aware Transformer (MorFormer) that achieves state-of-the-art results on five public datasets. Moreover, we propose a large-scale cross-domain benchmark for crack segmentation, where MorFormer exhibits excellent domain generalization.
Wenzhong Tang, Shuai Wang 0049, Xiaolei Qu, Xun Lin
IEEE Trans. Intell. Transp. Syst.5
2025 Propagation Based Recycling Contrastive Learning for Coupled Noisy Visible-Infrared Person Re-Identification
abstract
Visible-Infrared Person Re-Identification (VI-ReID) plays a crucial role in round-the-clock security surveillance systems, aiming to detect consistent identity recognition across transitions from day to night. A significant challenge in this field is the variation in the appearance of the same identity across visible and infrared modalities, which often leads to coupled noisy labels, referring to both Noisy Annotation (NA) and Noisy Correspondence (NC). Therefore, learning noisy-tolerant and discriminative representations is the primary objective in VI-ReID. However, existing research typically faces two principal limitations: (1) Learning strategies for noisy labeled scenarios usually rely on analyzing the distribution of loss response while ignoring the rich semantic information from neighboring samples. (2) When dealing with identified noisy samples, most previous approaches usually employ filtering strategies to mitigate the impact of noisy samples but fail to consider the valuable information in the noisy samples. To address these challenges, we propose a Propagation based Recycling Contrastive Learning (PRCL) approach. This method utilizes a label propagation strategy to distinguish clean annotations to learn identity-wise semantic information and recycles filtered noisy samples to capture the geometric-wise representation. Thus, even in the presence of noisy labels, the method can help learn robust representations across visible and infrared modalities. Specifically, we design a Noisy-aware Heterogeneous Graph Propagation module, which identifies noisy samples by aggregating the effects of neighboring labels using a graph propagation strategy. In addition, we develop a Cross Modality Recycling Debiased Contrastive Learning algorithm, which leverages the identity-wise information from clean samples and geometry-wise information from noisy samples. This approach utilizes identity-wise and geometric-wise information to mitigate the effect of noisy labels and retain as much valuable information as possible. Extensive experiments on two VI-ReID benchmark datasets demonstrate that our proposed method achieves highly competitive performance.
Yongxi Li, Wenzhong Tang, Shuai Wang 0049, Shengsheng Qian, Quan Fang, Changsheng Xu
IEEE Trans. Multim.3
2025 Super-Ellipse Formation Tracking of Uncertain Vehicles: A Simplified Reinforcement Learning Energy Optimization Method
abstract
This article deals with the optimal super-ellipse formation tracking control problem for multiple unmanned vehicles (MUVs), where each vehicle contains nonlinear uncertainties of unmodeled basic resistance, and the objective of energy optimization includes the super-ellipse orbit tracking energy and formation motion energy on the normal and tangent directions along the super-ellipse orbits, respectively. The communication topology is the directed leader-following structure. To avoid using the inputs of neighboring MUVs and the global communication information, a novel augmented formation input is designed and integrated into the formation motion subsystem. To deal with the uncertain nonlinearity, the uncertain virtual leader information, and the limited information of neighboring MUVs in the Hamilton-Jacobi–Bellman equations, a simplified reinforcement learning (RL) energy optimization method is designed based on identifier neural networks (NNs) and optimized backstepping technique. Theoretical stability analysis of system errors are given in detail. Simulation results show that the super-ellipse formation tracking energy consumption is significantly saved and the algorithm run time is decreased through comparison.
Yang-Yang Chen 0001, Guanghui Wen, Shuai Wang 0049, Tingwen Huang
IEEE Trans. Syst. Man Cybern. Syst.4
2024 Suppress and Rebalance: Towards Generalized Multi-Modal Face Anti-Spoofing
abstract
Face Anti-Spoofing (FAS) is crucial for securing face recognition systems against presentation attacks. With ad-vancements in sensor manufacture and multi-modal learning techniques, many multi-modal FAS approaches have emerged. However, they face challenges in generalizing to unseen attacks and deployment conditions. These chal-lenges arise from (1) modality unreliability, where some modality sensors like depth and infrared undergo signifi-cant domain shifts in varying environments, leading to the spread of unreliable information during cross-modal feature fusion, and (2) modality imbalance, where training overly relies on a dominant modality hinders the conver-gence of others, reducing effectiveness against attack types that are indistinguishable by sorely using the dominant modality. To address modality unreliability, we propose the Uncertainty-Guided Cross-Adapter (U-Adapter) to recognize unreliably detected regions within each modality and suppress the impact of unreliable regions on other modal-ities. For modality imbalance, we propose a Rebalanced Modality Gradient Modulation (ReGrad) strategy to rebal-ance the convergence speed of all modalities by adaptively adjusting their gradients. Besides, we provide the first large-scale benchmark for evaluating multi-modal FAS per-formance under domain generalization scenarios. Exten-sive experiments demonstrate that our method outperforms state-of-the-art methods. Source codes and protocols are released on https://github.com/OMGGGGG/mmdg.
Xun Lin, Shuai Wang 0049, Rizhao Cai, Yizhong Liu, Ying Fu 0001, Wenzhong Tang, Zitong Yu, Alex Chichung Kot
CVPR2
2024 HideMIA: Hidden Wavelet Mining for Privacy-Enhancing Medical Image Analysis
Xun Lin, Yi Yu 0011, Zitong Yu, Ruohan Meng, Jiale Zhou 0001, Ajian Liu 0001, Yizhong Liu, Shuai Wang 0049, Wenzhong Tang, Zhen Lei 0001, Alex Chichung Kot
ACM Multimedia8
2024 Temperature compensation for humidity sensors using ISSA-BP neural network
abstract
High-precision humidity detection is essential in various fields. However, the sensor's output signal is often affected by complex environmental temperature changes. To address these challenges, this paper propose an improved sparrow search algorithm based on the back propagation neural network (ISSA-BP). This method optimizes the initialization process of the traditional algorithm and introduces an edge evolution strategy to enhance its iterative update process, significantly improving both efficiency and accuracy. Experimental results demonstrate that the proposed algorithm reduces the mean absolute percentage error (MAPE) from 4.35% to 0.92% and decreases the convergence time from 1.25 s to 0.65 s, compared to traditional methods.
Hechu Zhang, Wei Li 0181, Shuai Wang 0049, Dezhi Zheng
MobiCom5
2024 Multi-UAV Collaborative Surveillance Network Recovery via Deep Reinforcement Learning
abstract
As a typical nonterrestrial network (NTN)-enabled Internet of Things (IoT), the multi-Unmanned aerial vehicle (UAV) collaborative surveillance network boasts efficient capabilities in information collection and transmission. However, manufacturing techniques and environmental conditions can lead to UAV failures, thereby impacting network performance. To recover the performance of the multi-UAV collaborative surveillance network, the effective movement of multiple UAVs is under investigation in order to improve target coverage and data backhaul efficiency. In this article, we present a novel multiagent deep reinforcement learning-based algorithm to accomplish network recovery. The proposed algorithm employs a multihead attention network to facilitate coupled multiobjective learning and overcome the limitations imposed by local information. Additionally, a stable learning method is introduced to address the difficult convergence problem caused by dynamic topology changes due to UAV motion. Experimental results show that the proposed algorithm can generate feasible multi-UAV motion strategies, effectively facilitating network recovery and improving the performance of the multi-UAV collaborative surveillance network in different scenarios.
Tao Wang 0151, Jingjing Wang 0001, Wenbo Du 0001, Dezhi Zheng, Shuai Wang 0049
IEEE Internet Things J.6
2024 Distribution-Guided Hierarchical Calibration Contrastive Network for Unsupervised Person Re-Identification
abstract
The person re-identification task aims to retrieve the same identity under different cameras. The main difficulties of the task lie in the collection of a large amount of annotated data and the diversity of pedestrians. Therefore, how to learn a robust and discriminative representation feature with unlabeled data is the key to this task. The pseudo label based methods have shown significant effectiveness in the field by generating pseudo labels from unlabeled data instead of ground-truth labels. However, existing researches typically suffer two limitations: (1) The extracted features are insufficient to reflect the subtle local semantics; (2) The pseudo labels generated by clustering methods cannot avoid introducing noise, which will seriously affect the performance of the discriminative feature. In this paper, to address the above problems, we propose a Distribution-Guided Hierarchical Calibration Contrastive Network (DHCCN) to better exploit local clues and hierarchical representation, which can consider cross-granularity consistency and reduce the noise of pseudo labels by the calibrated feature distribution. A Hierarchical Feature Extractor is employed to capture the multi-granularity response of each image, and fuse both global salience and local subtle texture information of a pedestrian to generate the hierarchical feature. In addition, to reduce the error of the pseudo labels, we introduce a Feature Distribution Corrector to calibrate noisy features of low-confidence samples evaluated by a Gaussian Mixture Model. At last, we integrate cross-granularity consistency constraint by the difference between the global and local feature, which can help generate more accurate feature embedding and improve robustness of the model. Therefore, we can receive a performance that is close to the supervised person re-identification task by narrowing the gap between the pseudo and ground-truth label. Experiments on four standard benchmarks demonstrate the effectiveness of our method against the state-of-the-art unsupervised re-identification methods. The code is available at https://github.com/Li-Yongxi/2023-DHCCN.
Yongxi Li, Wenzhong Tang, Shuai Wang 0049, Shengsheng Qian, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Category-Level Band Learning-Based Feature Extraction for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is a classical task in remote sensing image analysis. With the development of deep learning, schemes based on deep learning have gradually become the mainstream of HSI classification. However, existing HSI classification schemes either lack the exploration of category-specific information in the spectral bands and the intrinsic value of information contained in features at different scales, or are unable to extract multiscale spatial information and global spectral properties simultaneously. To solve these problems, in this article, we propose a novel HSI classification framework named CL-MGNet, which can fully exploit the category-specific properties in spectral bands and obtain features with multiscale spatial information and global spectral properties. Specifically, we first propose a spectral weight learning (SWL) module with a category consistency loss to achieve the enhancement of information in important bands and the mining of category-specific properties. Then, a multiscale backbone is proposed to extract the spatial information at different scales and the cross-channel attention via multiscale convolution and a grouping attention module. Finally, we employ an attention multilayer perceptron (attention-MLP) block to exploit the global spectral properties of HSI, which is helpful for the final fully connected layer to obtain the classification result. The experimental results on five representative hyperspectral remote sensing datasets demonstrate the superiority of our method.
Ying Fu 0001, Hongrong Liu, Yunhao Zou, Shuai Wang 0049, Zhongxiang Li, Dezhi Zheng
IEEE Trans. Geosci. Remote. Sens.4
2024 Distributed Multiagent Reinforcement Learning With Action Networks for Dynamic Economic Dispatch
abstract
A new class of distributed multiagent reinforcement learning (MARL) algorithm suitable for problems with coupling constraints is proposed in this article to address the dynamic economic dispatch problem (DEDP) in smart grids. Specifically, the assumption made commonly in most existing results on the DEDP that the cost functions are known and/or convex is removed in this article. A distributed projection optimization algorithm is designed for the generation units to find the feasible power outputs satisfying the coupling constraints. By using a quadratic function to approximate the state-action value function of each generation unit, the approximate optimal solution of the original DEDP can be obtained by solving a convex optimization problem. Then, each action network utilizes a neural network (NN) to learn the relationship between the total power demand and the optimal power output of each generation unit, such that the algorithm obtains the generalization ability to predict the optimal power output distribution on an unseen total power demand. Furthermore, an improved experience replay mechanism is introduced into the action networks to improve the stability of the training process. Finally, the effectiveness and robustness of the proposed MARL algorithm are verified by simulation.
Chengfang Hu, Guanghui Wen, Shuai Wang 0049, Junjie Fu, Wenwu Yu
IEEE Trans. Neural Networks Learn. Syst.3
2023 Image manipulation detection by multiple tampering traces and edge artifact enhancement
Xun Lin, Shuai Wang 0049, Jiahao Deng, Ying Fu 0001, Xiao Bai 0001, Xinlei Chen, Xiaolei Qu, Wenzhong Tang
Pattern Recognit.2
2023 CDS-Net: Cooperative dual-stream network for image manipulation detection
Jiahao Deng, Xun Lin, Wenzhong Tang, Shuai Wang 0049
Pattern Recognit. Lett.5
2023 Blind Super-Resolution of Single Remotely Sensed Hyperspectral Image
abstract
Hyperspectral image (HSI) super-resolution has recently advanced with significant progress by utilizing the powerful representation capabilities of deep neural networks. These approaches, however, inevitably rely on a sizable amount of training data which can be difficult to acquire for remotely sensed HSIs. In many cases, these methods are designed and tailored for only one or a few specific super-resolution scenarios, making them inflexible for handling images with different unknown degradations. In this paper, we introduce a two-step framework for blind remotely sensed HSI super-resolution, where the degradation is unknown. Specifically, in the first step, we propose to leverage the abundant remotely sensed color images to address the data insufficiency for remotely sensed HSI super-resolution. It is achieved by exploring the spatial knowledge from remotely sensed color images with a super-resolution network for a predefined degradation, which is then transferred to HSIs via band-by-band super-resolution. Direct use of the results from the transferred super-resolution network is suboptimal as it neglects the spectral correlations of different bands and the gap between predefined degradation and the real one. To make further refinements, we present an unsupervised scheme that simultaneously refines the super-resolved HSI and the unknown degradation by a non-negative matrix factorization network and a learnable degradation prior. To validate the effectiveness of our method, we conducted extensive experiments on a variety of remotely sensed HSI datasets. The results demonstrate that our method could generalize on various unknown degradations with superior performance against the state-of-the-art methods.
Zhiyuan Liang, Shuai Wang 0049, Tao Zhang 0042, Ying Fu 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Riemannian Geometric Instance Filtering for Transfer Learning in Brain-Computer Interfaces
abstract
Due to the inter-subject variability of Electroencephalogram(EEG) signals, a long calibration time is required to collect a large number of labeled trials to calibrate classifier parameters before using the Brain-computer Interface(BCI). This challenge greatly limits the practical roll-out of BCIs. To address this problem, we propose a novel instance-based transfer learning framework named Riemannian Geometric Instance Filtering (RGIF) to reduce calibration time without sacrificing accuracy. A new inter-subject similarity metric based on Riemannian geometry is proposed to measure the similarity between a few trials from the target subject and adequate trials from source subjects. The classification model for the target subject is then trained with the help of abundant trials from similar source subjects with high similarity to the target subject. We evaluate our method on two open-source EEG datasets. The results show that our approach improves significantly compared with other baselines. Furthermore, compared with using all source subjects data, our method reduces the training time by at least half and achieves slightly better accuracy.
Qianxin Hui, Yang Li 0104, Susu Xu, Shuailei Zhang, Ying Sun 0012, Shuai Wang 0049, Xinlei Chen, Dezhi Zheng
SenSys7
2022 Non-Acoustic Speech Sensing System Based on Flexible Piezoelectric
abstract
Speech is one of the most important biological signals to complement human-human and human-computer interaction. Traditional speech datasets were collected by air microphones, but using these datasets in noisy environments such as factories is practically challenging. Therefore, speech recognition in noisy environments poses higher requirements. The non-acoustic speech dataset plays a significant role in robust speech recognition under high background noise. Existing datasets suffered from dull sound, low intelligibility and poor recognition accuracy due to hardware and computer technology limitations. This paper presents a non-acoustic speech sensing system based on flexible piezoelectric. The system collected vibration signals from the jaws of six males and five females, and the corpus contained ten different control commands at 90 dB of background noise. The dataset is reliable with high intelligibility and capable of achieving 93.7% recognition accuracy by calculation. With the aforementioned benefits, this dataset is an essential tool for studying human-computer interaction in high-noise environments, analyzing human acoustic properties, and aiding medical rehabilitation.
Shiji Yuan, Ying Sun 0012, Shuai Wang 0049, Xinlei Chen, Dezhi Zheng, Shangchun Fan
SenSys3
2022 A Wearable Low-Power Collaborative Sensing System for High-Quality SSVEP-BCI Signal Acquisition
abstract
The brain–computer interface (BCI) technology improves the communication efficiency between people and Internet of Things (IoT) devices. BCI based on the steady-state visual evoked potential (SSVEP-BCI) is the preferred scheme for controlling devices because of its convenient operation, low training requirement, and high information transmission rate (ITR). Most signal acquisition devices for BCIs are used for medical diagnosis and scientific research and utilize multiple channels and wet electrodes to obtain high-quality signals. However, the practicability, wearability, and cost of the signal acquisition devices for real-life applications need to be considered, resulting in new requirements for the acquisition mode, the number of electrodes, power consumption, and signal processing methods. This article presents a wearable low-power collaborative sensing system based on a time mask window canonical correlation analysis method (TMW-CCA). An 8-array spring dry electrode signal acquisition device based on a flexible circuit board is designed to address the shortcomings of traditional wet electrode acquisition devices, such as high-power consumption, discomfort, and being unsuitable for long-time use. The proposed TMW-CCA method, which uses a dry electrode sensor to evaluate the time domain’s signal quality dynamically, exhibits 12.5% higher steady-state visual evoked potential recognition accuracy and 40% lower average power consumption (only 740 mW) than the benchmark.
Rui Na, Dezhi Zheng, Ying Sun 0012, Mingzhe Han, Shuai Wang 0049, Shuailei Zhang, Qianxin Hui, Xinlei Chen, Jun Zhang 0007, Chun Hu
IEEE Internet Things J.5
2022 TCACNet: Temporal and channel attention convolutional network for motor imagery classification of EEG-based BCI
abstract
Brain–computer interface (BCI) is a promising intelligent healthcare technology to improve human living quality across the lifespan, which enables assistance of movement and communication, rehabilitation of exercise and nerves, monitoring sleep quality, fatigue and emotion. Most BCI systems are based on motor imagery electroencephalogram (MI-EEG) due to its advantages of sensory organs affection, operation at free will and etc. However, MI-EEG classification, a core problem in BCI systems, suffers from two critical challenges: the EEG signal’s temporal non-stationarity and the nonuniform information distribution over different electrode channels. To address these two challenges, this paper proposes TCACNet, a temporal and channel attention convolutional network for MI-EEG classification. TCACNet leverages a novel attention mechanism module and a well-designed network architecture to process the EEG signals. The former enables the TCACNet to pay more attention to signals of task-related time slices and electrode channels, supporting the latter to make accurate classification decisions. We compare the proposed TCACNet with other state-of-the-art deep learning baselines on two open source EEG datasets. Experimental results show that TCACNet achieves 11.4% and 7.9% classification accuracy improvement on two datasets respectively. Additionally, TCACNet achieves the same accuracy as other baselines with about 50% less training data. In terms of classification accuracy and data efficiency, the superiority of the TCACNet over advanced baselines demonstrates its practical value for BCI systems.
Rongye Shi, Qianxin Hui, Susu Xu, Shuai Wang 0049, Rui Na, Ying Sun 0012, Wenbo Ding 0001, Dezhi Zheng, Xinlei Chen
Inf. Process. Manag.5
2020 Spectral bounding: Strictly satisfying the 1-Lipschitz property for generative adversarial networks
Zhihong Zhang 0001, Yangbin Zeng, Lu Bai 0001, Yiqun Hu, Meihong Wu, Shuai Wang 0049, Edwin R. Hancock
Pattern Recognit.6
2019 Cross-modal hashing with semantic deep embedding
Xiao Bai 0001, Shuai Wang 0049, Jun Zhou 0001, Edwin R. Hancock
Neurocomputing3
2019 Multiscale Visual Attention Networks for Object Detection in VHR Remote Sensing Images
abstract
Object detection plays an active role in remote sensing applications. Recently, deep convolutional neural network models have been applied to automatically extract features, generate region proposals, and predict corresponding object class. However, these models face new challenges in VHR remote sensing images due to the orientation and scale variations and the cluttered background. In this letter, we propose an end-to-end multiscale visual attention networks (MS-VANs) method. We use skip-connected encoder-decoder model to extract multiscale features from a full-size image. For feature maps in each scale, we learn a visual attention network, which is followed by a classification branch and a regression branch, so as to highlight the features from object region and suppress the cluttered background. We train the MS-VANs model by a hybrid loss function which is a weighted sum of attention loss, classification loss, and regression loss. Experiments on a combined data set consisting of Dataset for Object Detection in Aerial Images and NWPU VHR-10 show that the proposed method outperforms several state-of-the-art approaches.
Chen Wang 0026, Xiao Bai 0001, Shuai Wang 0049, Jun Zhou 0001, Peng Ren 0001
IEEE Geosci. Remote. Sens. Lett.3
2019 A novel pattern with high-level commands for encoding motor imagery-based brain computer interface
Shuailei Zhang, Shuai Wang 0049, Dezhi Zheng, Mengxi Dai
Pattern Recognit. Lett.2
2018 Discriminate Cross-modal Quantization for Efficient Retrieval
abstract
Efficient cross-modal retrieval involves searching similar items across different modalities, e.g., using an image(text) to search for texts(images). To speed up cross-modal retrieval, hashing-based methods threshold continuous embeddings into binary codes, inducing substantial loss of accuracy retrieval. To further improve retrieval performance, several quantization-based methods quantize embeddings into real-valued codewords to maximumlly preserve inter-modal and intra-modal similarity relation, while the discrimination between dissimilar data is ignored. To address these challenges, we propose, for the first time, a novel discriminate cross-modal quantization(DCMQ) which nonlinearly maps different modalities into a common space where ir-relevant data points are semantically separable: the points belonging to a class lie in a cluster that is not overlapped with other clusters corresponding to other classes. An effective optimization algorithm is developed for the proposed method to jointly learn the modality-specific mapping functions, the sharing codebooks, the unified binary codes and a linear classifier. Experimental comparison with state-of-the-art algorithms over three benchmark datasets demonstrates that DCMQ achieves significant improvement in search accuracy.
Shuai Wang 0049, Xiao Bai 0001
ICPR3