Xiushan Nie

dblp:03/8117 · DBLP profile ↗
← Back
141ranked-venue papers
15as first author
93since 2021 · last 2026
0000-0001-9644-9723ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 87 · 10 first-author · 55 since 2021Artificial intelligence and machine learning · 41 · 31 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 7 since 2021Computer networks · 8 · 1 first-author · 7 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 PEOCH: Online Cross-Modal Hashing with Semi-Supervised Streaming Data Driving Prototype Evolution
abstract
The exponential growth of streaming multi-modal data presents critical challenges for cross-modal retrieval: distribution shifts, modality gap, and scarce labels. Semi-supervised online cross-modal hashing has gained increasing interest due to its ability to encode complex streaming data and update hash functions simultaneously. Nevertheless, existing methods can hardly generate high-quality unsupervised hash codes, which fundamentally limits diversity and flexibility during the retrieval process. To this end, we propose a novel method named Prototype Evolution Online Cross-modal Hashing (PEOCH). By driving prototype evolution with semi-supervised streaming data, precise and stable hash codes are generated for both labeled and unlabeled data. Specifically, two prototype updates with stability guarantee are conducted: labeled samples push semantic knowledge into the supervised prototypes, while unlabeled samples perform clustering to generate unsupervised prototypes. Simultaneously, a co-optimization mechanism is designed to ensure the prototypes continuously evolve and preserve the consistency of the entire streaming data. Besides, an elasticity regularizer integrates discriminability and smoothness constraints, improving the reliability of prototypes. Extensive experiments on three benchmark datasets demonstrate that PEOCH outperforms state-of-the-art methods, achieving an average improvement of 6.7% in mAP@all across various retrieval tasks.
Xiao Kang, Xingbo Liu, Shuo Pan, Xuening Zhang, Xiushan Nie, Yilong Yin
AAAI5
2026 Exploiting dynamic spatio-temporal correlations for origin-destination demand prediction
abstract
Accurate Origin-Destination (OD) demand prediction is fundamental to intelligent transportation systems (ITS), enabling real-time traffic management, dynamic vehicle dispatch, and efficient resource allocation in urban environments. However, OD demand exhibits complex, dynamic, and highly coupled spatio-temporal patterns that remain challenging for existing models. We propose a novel Dynamic Spatio-Temporal Correlation Network (DSTCN) for OD demand forecasting. DSTCN features three key components: (1) a bidirectional demand trend modeling module (Glstm2D) that jointly learns demand evolution from both origin and destination perspectives; (2) a Transformer-based spatial similarity module (Simformer) to dynamically extract and integrate inter-regional correlations across the OD matrix; and (3) a temporal fusion and modeling module (FF-TM) that combines and processes multi-source spatio-temporal features for next-step prediction. Extensive experiments on three large-scale real-world datasets (NYC-TOD2018, NYC-TOD2019, and HZMetro) demonstrate that DSTCN consistently outperforms state-of-the-art baselines across diverse urban scenarios.
Yongshun Gong, Piao Yu, Xu Zhang 0039, Xinxin Zhang 0004, Xiushan Nie, Haoliang Sun
Expert Syst. Appl.5
2026 STADNN: Spatio-temporal adaptive decomposition neural network for traffic prediction
Yongshun Gong, Xiushan Nie, Yilong Yin
Neurocomputing5
2026 Class-mismatched semi-supervised learning from a new perspective
Rundong He, Zhongyi Han, Xiushan Nie, Qi Wei 0004, Yilong Yin
Pattern Recognit.4
2026 OSAD: Open-Set Supervised Anomaly Detection in Surveillance Videos Based on Margin Metric Learning
abstract
Open-set Supervised Anomaly Detection (OSAD) strategy seeks to detect novel anomalies that are unseen during training. However, existing OSAD works fail to learn a comprehensive margin that separates normal and anomalous samples in the feature space owing to the lack of restriction on abnormal distribution. To this end, we propose a new OSAD scheme based on Margin Metric Learning, which fully exploits the differences between normal and abnormal events in the latent feature dimension. The developed framework embodies three major components. First, video clips including normal events and restricted anomalies are input into an Attention-embedded Spatial Convolutional Network to extract the spatial feature sequences. Then, the spatial feature sequences are input into a Temporal Convolutional Siamese Network to further obtain the temporal features. Second, a novel Quadruplet Contrastive Loss is designed based on the spatial-temporal features of normal and abnormal video clips to conduct the margin learning, which enlarges the inter-class distances between anomalous and normal instances as well as reduces the intra-class distances inside anomalous instances and normal instances. This contributes to acquire compact normal and abnormal distributions. Finally, during the testing stage, a simplified metric distance-based anomaly detection strategy is proposed to calculate the anomaly score of testing clip. Extensive experiments and ablation studies on the Avenue and ShanghaiTech datasets demonstrate the effectiveness and efficiency of our proposed method for discovering unseen anomalous events via limited types of aware anomalies.
Nanjun Li, Xiushan Nie
IEEE Trans. Circuits Syst. Video Technol.3
2026 Internal-External Context Interaction Network for Person Re-Identification
abstract
Capturing discriminative cues with attention mechanisms is crucial for solving the high inter-class similarity problem of person re-identification (Re-ID). Self-attention (SA) learns its own contextual information within a single sample using self-affinity between elements, and some works have demonstrated its superiority in person Re-ID. However, SA weakens some subtle semantic cues and additional visual cues such as backpacks, which makes it difficult to distinguish similar-looking persons. In this paper, we propose an internal-external context interaction (IEI) attention mechanism, which aims to exploit the interaction of inter-sample latent context information and intra-sample local context information to enhance the feature representation of each element. The mechanism is able to capture subtle differences between persons and additional visual cues using inter-sample difference information and rich detail information within the element neighborhood, improving the ability to distinguish similar persons. Based on this mechanism, we propose an internal-external context interaction network (IEINet) for extracting discriminative features from multiple dimensions. In addition, to capture more discriminative information, we propose a region-diverse loss to constrain the network. Many experiments validate the effectiveness of our IEINet and demonstrate that our approach attains state-of-the-art performance on several large-scale person Re-ID datasets.
Tongxin Liu, Xiyu Pang, Gangwu Jiang, Xiushan Nie, Meifeng Zheng, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.4
2026 Prior Distribution Guided Gaussian Mixture Variational Autoencoder (PDGM-VAE) for Image Generation
abstract
Variational Autoencoder(VAE) combines the ideas of autoencoders and variational inference, introducing the concept of latent space and variational inference to endow autoencoders to generate new images. VAE typically assumes that data follows a Gaussian distribution, but real data may follow other distributions. This inconsistency between the assumption and the true distribution can affect the modeling and reconstruction capabilities of VAE, which makes it difficult for traditional models to accurately capture the true distribution. To address the aforementioned issues, we propose a Prior Distribution Guided Gaussian Mixture Variational Autoencoder(PDGM-VAE). Specifically, we construct a Gaussian Mixture Prior Learner (GMPL) to capture complex features of the data distribution, enabling the model to learn and obtain a Gaussian mixture distribution that is reasonable and close to the real data distributions, which is then used as the prior distribution in the network. Furthermore, we build a Semantic-Aware Module with Embedded Prior Distribution (SAMEPD), integrating data and label information to learn the distribution parameters, enabling the network to learn and utilize the semantic knowledge contained in the labels. During training, by approximating the posterior distribution to the prior distribution, we enhance the model’s modeling and reconstruction capabilities, improving the quality of generated images. We evaluated the image generation task on five public datasets, and based on the FID metric, our proposed method outperformed other VAE methods.
Jingqi Song, Yipeng Ning, Xiaoming Xi, Jie Guo 0012, Xiushan Nie, Lishan Qiao, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.7
2026 Federated Domain Generalization via Prompt Learning and Aggregation
abstract
Federated domain generalization (FedDG) aims to improve the global model’s generalization ability in unseen domains by addressing data heterogeneity under privacy-preserving constraints. A common strategy in existing FedDG studies involves sharing domain-specific knowledge among clients, such as spectrum information, class prototypes, and data styles. However, this knowledge is extracted directly from local client samples, and sharing such sensitive information poses a potential risk of data leakage, which might not fully meet the FedDG requirements. In this paper, we introduce prompt learning to adapt pretrained vision-language models (VLMs) in the FedDG scenario, and leverage locally learned prompts as a more secure bridge to facilitate knowledge transfer among clients. Specifically, we propose a novel FedDG framework through Prompt Learning and AggregatioN (PLAN), which comprises two training stages to collaboratively generate local prompts and global prompts at each federated round. First, each client performs both text and visual prompt learning using their own data, with local prompts indirectly synchronized by regarding the global prompts as a common reference. Second, all domain-specific local prompts are exchanged among clients and selectively aggregated into global prompts using lightweight attention-based aggregators. The global prompts are finally applied to adapt the VLMs to unseen target domains. As our PLAN framework requires training only a limited number of prompts and lightweight aggregators, it offers notable advantages in terms of computational and communication efficiency for FedDG. Extensive experiments demonstrate the superior generalization ability of PLAN across four benchmark datasets. We have released our code at https://github.com/GongShuai8210/PLAN.
Shuai Gong, Chaoran Cui, Chunyun Zhang, Wenna Wang, Xiushan Nie, Lei Zhu 0002
IEEE Trans. Inf. Forensics Secur.5
2026 Token-Level Prompt Mixture With Parameter-Free Routing for Federated Domain Generalization
abstract
Federated Domain Generalization (FedDG) aims to train a globally generalizable model on data from decentralized, heterogeneous clients. While recent work has adapted vision-language models for FedDG using prompt learning, the prevailing "one-prompt-fits-all" paradigm struggles with sample diversity, causing a marked performance decline on personalized samples. The Mixture of Experts (MoE) architecture offers a promising solution for specialization. However, existing MoE-based prompt learning methods suffer from two key limitations: coarse image-level expert assignment and high communication costs from parameterized routers. To address these limitations, we propose TRIP, a Token-level pRompt mIxture with Parameter-free routing framework for FedDG. TRIP treats prompts as multiple experts, and assigns individual tokens within an image to distinct experts, facilitating the capture of fine-grained visual patterns. To ensure communication efficiency, TRIP introduces a parameter-free routing mechanism based on capacity-aware clustering and Optimal Transport (OT). First, tokens are grouped into capacity-aware clusters to ensure balanced workloads. These clusters are then assigned to experts via OT, stabilized by mapping cluster centroids to static, non-learnable keys. The final instance-specific prompt is synthesized by aggregating experts, weighted by the number of tokens assigned to each. Extensive experiments across four benchmarks demonstrate that TRIP achieves optimal generalization results, with communicating as few as 1K parameters. Our code is available at https://github.com/GongShuai8210/TRIP.
Shuai Gong, Chaoran Cui, Xiaolin Dong, Xiushan Nie, Lei Zhu 0002, Xiaojun Chang
IEEE Trans. Image Process.4
2026 Local Refinement and Global Strengthening Network for Vehicle Re-Identification
abstract
Vehicle re-identification (Re-ID) aims to retrieve vehicle images with the same identity as the query from an image library. Currently, the vehicle Re-ID task mainly faces two challenges: large intra-class variance and small inter-class variance. Learning discriminative local features and global features of vehicles is crucial to address both challenges, and the attention mechanism adequately learns local features and global features in vehicle images without the aid of an auxiliary model. Self-attention mechanism as a special kind of attention mechanism, it mainly contains two forms of local self-attention for extracting local features and Full self-attention for extracting global features. However, these two approaches have their own limitations: the window mode of local self-attention hinders adequate learning of the local detailed information of vehicles; the remote connections in the global context modeled by full self-attention are usually weak, which limits the full learning of the overall information about vehicles. To address the above problem, we propose two complementary modules: local refinement module (LRM) and global strengthening module (GSM). The LRM aims to learn the refined local representation, which captures the rich correlation information between adjacent pixels through the interactions of the target pixel with its nearest pixels. The GSM aims to learn the strengthened global representation, which first disperses attention at the target pixel into various windows to emphasize important remote dependence within each region and then aggregates globally meaningful remote connections by cross-window interaction. In addition, we construct a multi-branch network, local refinement and global strengthening network (LRGS-Net), which uses LRM and GSM to learn discriminative local features and global features to address vehicle Re-ID challenges. We validate the effectiveness of our method on three datasets, VeRi-776, VehicleID, and VERI-Wild.
Meifeng Zheng, Xiyu Pang, Xiushan Nie, Houren Zhou, Yilong Yin
IEEE Trans. Intell. Transp. Syst.5
2026 SOR-BDNet: Semantic-Optical Representation for Boundary-Aware Video Anomaly Detection with GPT-4o
abstract
In recent years, Video Anomaly Detection (VAD) has shifted from conventional appearance-based modeling to semantically driven frameworks empowered by LLMs. Traditional reconstruction- and prediction-based methods, relying on motion or appearance patterns learned from normal data, often misclassify previously unseen yet semantically normal events as anomalies. To address this limitation, we propose SOR-BDNet (Semantic-Optical Representation with Boundary Detection Network), an annotation-free multimodal VAD framework that jointly leverages visual appearance and motion dynamics to generate interpretable semantic representations at the frame level. Specifically, we employ RAFT to estimate dense motion fields and concatenate the resulting flow maps with RGB images to form unified spatiotemporal inputs. These fused representations are fed into a GPT-4o-based module that generates semantic captions capturing object semantics and motion cues. Anomalies are detected by measuring semantic deviations from a memory bank constructed from normal captions. To further refine temporal boundaries, we design a boundary refinement module that integrates visual continuity constraints with contrastive feature learning based on a Swin Transformer backbone. Extensive experiments on four challenging benchmarks—UCSD-Ped2, Avenue, ShanghaiTech, and UCF-Crime—demonstrate that SOR-BDNet achieves frame-level accuracies of 97.96%, 82.86%, 87.36%, and 85.64%, respectively. These results highlight the robustness and scalability of the proposed framework, while significantly improving interpretability and generalization across diverse real-world surveillance scenarios. The source code and pretrained models are available at https://github.com/syi-coder/SOR-BDNet-Semantic-Optical-Representation-for-Boundary-Aware-Video-Anomaly-Detection-with-GPT-4o .
Bryan W. Scotney, Xiushan Nie, Xingbo Liu, Shuai Zhang 0001, Lanting Qiu
ACM Trans. Multim. Comput. Commun. Appl.3
2026 Learning Prediction-aware Prior in Transformer Network for Accurate Spatio-Temporal Video Grounding
abstract
Spatio-temporal video grounding (STVG) aims to precisely locate a spatio-temporal tube in an untrimmed video corresponding to a given language description. Many existing methods decouple spatial and temporal grounding as separate tasks, missing the strong interdependencies between the two, which are crucial for accurately aligning spatial regions (such as objects) with their motion over time. Thus, to enhance spatio-temporal associations, we introduce a new Prior-Driven Transformer Network (PDTNet) with predicted temporal boundaries as priors to guide object bounding boxes for improved spatial grounding over time. Firstly, PDTNet employs a temporal prior, termed reference query, to enhance discriminability between language-related and language-irrelevant visual content, improving temporal boundary localization. Further, the context within predicted temporal boundaries serves as another prior knowledge to modulate spatial features. We also introduce a prediction-aware Gaussian prior to precise object localization. This ensures consistent tube construction and accurate object localization. Extensive experiments on STVG benchmarks validate the effectiveness of PDTNet. Code is available at https://github.com/tongzhang111/PDTNet .
Yongshun Gong, Jialin Gao, Yanyu Xu 0001, Xiushan Nie, Li-Zhen Cui 0001, Chengqi Zhang
ACM Trans. Multim. Comput. Commun. Appl.7
2025 Semi-Supervised Online Cross-Modal Hashing
abstract
Online cross-modal hashing has gained increasing interest due to its ability to encode streaming data and update hash functions simultaneously. Existing online methods often assume either fully supervised or completely unsupervised settings. However, they overlook the prevalent and challenging scenario of semi-supervised cross-modal streaming data, where diverse data types, including labeled/unlabeled, paired/unpaired, and multi-modal, are intertwined. To address this issue, we propose Semi-Supervised Online Cross-modal Hashing (SSOCH). It presents an alignment-free pseudo-labeling strategy that extracts semantic information from unlabeled streaming data without relying on pairing relations. Furthermore, we design an online tri-consistent preserving scheme, integrating pseudo-labeled data regularization, discriminative label embedding, and fine-grained similarity preservation. This scheme fully explores consistency across data annotation, modalities, and streaming chunks, improving the model's adaptiveness in these challenging scenarios. Extensive experiments on benchmark datasets demonstrate the superiority of SSOCH under various scenarios, highlighting the importance of semi-supervised learning for online cross-modal hashing.
Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
AAAI5
2025 Generalized Debiased Semi-Supervised Hashing for Large-Scale Image Retrieval
abstract
Semi-supervised hashing has shown promising efficacy in large-scale image retrieval, which learns similarity-preserving codes from both labeled and unlabeled data. To enable the use of advanced supervised hashing techniques, pseudo labels are widely applied. However, existing methods typically suffer from a biased learning issue due to pseudo label noise, which can be further aggravated during optimization. Although such a bias can adversely affect hashing accuracy, it has not been investigated sufficiently. In view of this, we present a comprehensive discussion on potential causes of biases, involving processes of pseudo-labeling, hash learning and optimization. Accordingly, a novel Generalized Debiased Semi-supervised Hashing (GDSH) method is proposed as a unified solution to mitigate the biases. Specifically, reliable pseudo labels are first predicted via a robust label completion strategy. Secondly, a debiased hash learning module is designed by combining label denoising and similarity updating. This can not only refine the supervision, but also obtain hash codes that are semantically debiased in both category and sample levels. Finally, a discrete semi-supervised hashing algorithm is proposed to alleviate the bias arising from optimization. Experimental results on three single-label and three multi-label image benchmarks demonstrate that GDSH remarkably outperforms the state-of-the-arts in different semi-supervised settings.
Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
AAAI3
2025 Binary Continual Stream-View Clustering
abstract
Multi-view clustering is valued for uncovering latent common semantics lying in multi-view data, which has been a hot topic in unsupervised learning. However, when dealing with incremental streaming views, existing approaches typically require reconstructing the view data and aggregating streaming representations, leading to misalignment between representation and clusters. More importantly, conducting the clustering process frequently results in significant time consumption. To address these issues, we propose a novel method called Binary Continual Stream-View Clustering (BCSVC). Specifically, we design a continual clustering method that seamlessly unifies streaming representation learning and cluster assignment within a single framework. We also introduce a variance-weighted center updating mechanism to smooth the frequent clustering operation and absorb the semantics of previous views. In addition, to reduce the time and space expenditure on computation and storage, binary code for clustering representations is introduced, which can also significantly improve the computational efficiency of continuous updates in streaming scenarios. Last but not least, comprehensive theoretical analysis and extensive experimental results demonstrate its superior performance under various scenarios.
Xingbo Liu, Kang Xiao, Xuening Zhang, Xiushan Nie
ECAI5
2025 Delving Into Coarse-Fine Feature Interaction Alignment for UAV Object Detection
abstract
Due to limited features and dense object layouts, object detection in UAV images is challenging. Given that existing feature fusion methods have not fully explored the relationship between fine- and coarse-grained features, direct feature fusion can result in poor correlation between them, hindering the representative capability of fine-grained semantic information. To alleviate this issue, we introduce a method of Coarse-fine Feature Interaction Alignment (CFIA), which enhances the correlation between coarse-grained and fine-grained features across multi-scale feature maps through their interactive alignment. Firstly, we present the Wavelet-based High-frequency Preserving Down-sampling (WHPD), utilizing wavelet transform to extract high-frequency information to enhance object boundaries, minimizing crucial fine-grained information loss. Secondly, we propose the Feature Refinement and Interaction Alignment Strategy (FRIAS), which achieves feature interaction alignment by establishing the association of feature maps between coarse-grained and fine-grained features. This enhances the representative capability of feature maps at various scales for detecting small objects. Extensive experiments on the VisDrone, CARPK, and Drone-vs-Bird datasets have demonstrated the effectiveness of the CFIA method, which is highly competitive with state-of-the-art methods. The code is available at https://github.com/b-yanchao/CFIA.git.
Yanchao Bi, Yang Ning, Xiushan Nie
ICASSP3
2025 Spatial Frequency-Aware Self-Distillation for Weakly-Supervised Semantic Segmentation
abstract
Weakly-supervised semantic segmentation (WSSS) aims to achieve pixel-level classification under image-level supervision. Recent class activation map (CAM)-based methods seek to expand foreground activation while suppressing background. However, they often overlook the uncertainty of CAM, where non-salient activation in some regions complicates semantic classification. These regions are typically dismissed as noise, resulting in inappropriate activations due to inadequate regularization. To resolve this, we introduce a Spatial Frequency-Aware Self-Distillation strategy (SFS). Firstly, to enhance the perception of high-frequency spatial information in uncertain regions, we propose a boundary self-distillation and uncertain region reconstruction strategy, which captures high-frequency boundary information and fine-grained spatial context in these regions. Secondly, to enhance the discrimination of low-frequency semantic features, we propose a contrastive attention mechanism that guides the Vision Transformer (ViT) to focus more on the foreground, thereby improving the distinction between foreground and background. Finally, our SFS demonstrates outstanding performance on both the VOC 2012 and COCO 2014 datasets, attributed to its superior spatial frequency perception capabilities. The code is available at https://github.com/fjoybest/SFS.
Jingyuan Fang, Yang Ning, Xiushan Nie
ICASSP3
2025 Towards Region-Adaptive Feature Disentanglement and Enhancement for Small Object Detection
abstract
Current feature fusion strategies often fail to adequately account for the influence of activation intensity across different scales on small object features, which impedes the effective detection of small objects. To address this limitation, we propose the Region-Adaptive Feature Disentanglement and Enhancement (RAFDE) strategy, which improves both downsampling and feature fusion by leveraging activation intensity variations at multiple scales. First, we introduce the Boundary Transitional Region-enhanced Downsampling (BTRD) module, which enhances boundary transitional regions containing both strongly and weakly activated features, thereby mitigating the loss of crucial boundary information for small objects. Second, we present the Regional-Adaptive Feature Fusion (RAFF) module, which adaptively disentangles and fuses co-activated and uni-activated regions from adjacent levels into the current level, effectively reducing the risk of small objects being overwhelmed. Extensive experiments on several public datasets demonstrate that the RAFDE strategy is highly effective and outperforms state-of-the-art methods. The code is available at https://github.com/b-yanchao/RAFDE.git.
Yanchao Bi, Yang Ning, Xiushan Nie, Xiankai Lu, Yongshun Gong, Leida Li
IJCAI3
2025 VLHP: Learning Discriminative Vision-Language Hybrid Prototypes for Weakly Supervised Semantic Segmentation
abstract
Recent advances in Weakly Supervised Semantic Segmentation (WSSS) focus on generating high-quality Class Activation Maps (CAMs) using image-level labels. However, the co-occurrence of foreground-background concepts in a single image often induces semantic confusion, which degrades the quality of conventional CAM-based approaches. In this paper, we propose VLHP, a novel framework that leverages vision-language hybrid prototypes to overcome semantic confusion. Specifically, VLHP constructs hybrid prototypes through cross-modal association between textual embeddings and visual features, generating discriminative semantic representations while effectively bridging the modality gap. To further improve discriminability, we introduce two dedicated strategies: Discriminative Explicit Alignment (DEA) to explore cross-modal consistent discrimination and Confounding Background Decoupling (CBD) to model co-occurring backgrounds and decouple them. Finally, a Prototype-driven Class-aware Decoder (PCD) employs these refined prototypes as category-specific priors to generate precise segmentation masks in a single-stage framework. Extensive experiments on PASCAL VOC and MS COCO benchmarks demonstrate that VLHP outperforms state-of-the-art alternatives. The code is available at https://github.com/fjy0105/VLHP.
Jingyuan Fang, Yang Ning, Xiushan Nie, Xinfeng Liu, Zhiyong Cheng 0001
ACM Multimedia3
2025 MAPformer: Multi-periodic Transformer with Adaptive Padding for Time Series Forecasting
Longtao Chang, Xiushan Nie, Xinyu Qin, Xinfeng Liu
PRCV (3)2
2025 Hierarchical Feature Alignment and Disentanglement for Cross-Domain Keyhole Penetration Prediction
Xiushan Nie, Xinfeng Liu, Fangzheng Zhou, Yunan Liu 0001
PRCV (3)2
2025 Dynamic Balance Sorting and Co-evolutionary Algorithm for Expensive Many-Objective Optimization
Xiushan Nie, Jie Tian 0004
PRICAI (4)2
2025 Cross-graph meta matching correction for noisy graph matching
Fangkai Li, Feiyu Pan, Wenjia Meng, Haoliang Sun, Xiushan Nie, Yilong Yin, Xiankai Lu
Comput. Vis. Image Underst.5
2025 Survey on deep learning-based weakly supervised salient object detection
Lina Du, Chunjie Ma, Huimin Zheng, Xiushan Nie, Zan Gao 0001
Expert Syst. Appl.5
2025 Leveraging spatio-temporal multi-task learning for potential urban flow prediction in newly developed regions
Wenqian Mu, Yongshun Gong, Xiushan Nie, Yilong Yin
Expert Syst. Appl.4
2025 Supervised Discriminative Transformer Hashing for Large-Scale Remote Sensing Image Retrieval
abstract
With the advancement of remote sensing technology and the exponential growth of remote sensing visual data, efficiently retrieving remote sensing images from extensive databases has become increasingly important. Deep hashing, which combines the advantages of deep learning and hashing techniques, has emerged as a significant research direction in remote sensing image retrieval (RSIR). However, remote sensing images often contain substantial amounts of complex background information that are unrelated to the target. This noise can obscure or interfere with the target features, making it challenging for the model to effectively distinguish the target objects. To address these challenges, we propose a novel method called Supervised Discriminative Transformer Hashing (SDTH) for large-scale RSIR task, designed to enhance the retrieval of remote sensing images by extracting more distinguishable features. Our approach utilizes the Swin Transformer V2 architecture to improve feature extraction capabilities, thereby acquiring richer global contextual information and multi-scale features. To mitigate the impact of image noise and generate more discriminative hash codes, we propose to integrate batch-hard triplet loss with symmetric cross entropy loss. Experimental results on three benchmark datasets for remote sensing demonstrate the effectiveness and superiority of the proposed method.
Xingbo Liu, Xuening Zhang, Xiushan Nie
IEEE Geosci. Remote. Sens. Lett.4
2025 GeM: Gaussian embeddings with Multi-hop graph transfer for next POI recommendation
Wenqian Mu, Jiyuan Liu 0013, Yongshun Gong, Ji Zhong, Wei Liu 0007, Haoliang Sun, Xiushan Nie, Yilong Yin, Yu Zheng 0004
Neural Networks7
2025 Consistency and label constrained transfer low-rank representation for cross-light finger vein recognition
Lu Yang 0005, Kuikui Wang, Xiaoming Xi, Xiushan Nie, Gongping Yang 0001, Yilong Yin
Pattern Recognit.5
2025 Adaptive division and priori reinforcement part learning network for vehicle re-identification
Xiaoying Zhou, Houren Zhou, Xiyu Pang, Jiachen Tian 0003, Xiushan Nie, Yilong Yin
Pattern Recognit.6
2025 Dual Difficulty-Aware Adaptive Pseudo Labeling for Semi-Supervised CNV Segmentation
abstract
In clinical practice, obtaining a large amount of labeled CNV data is very difficult. Semi-supervised learning can effectively utilize a large amount of unlabeled CNV data. Since CNV has complex features such as blurred and unevenly distributed pixels on the edges, there are differences in the segmentation difficulty between pixels in the same image. Existing semi-supervised segmentation methods do not consider the segmentation difficulty of pixels, which will reduce the segmentation accuracy. To address this problem, we propose a dual difficulty-aware adaptive pseudo-label learning (D2APL) method for semi-supervised CNV segmentation. The proposed dual difficulty awareness includes segmentation difficulty perception of pixels in labeled and unlabeled data. For labeled data, we propose a classification confidence-guided difficulty perception method. For unlabeled data, we propose a model stability-guided difficulty perception method. Finally, we propose a difficulty-aware self-training method to dynamically adjust the threshold of pseudolabels according to the difficulty, thereby improving the utilization of difficult-to-segment pixels in unlabeled data. Experimental results show that our method outperforms the state-of-the-art method in CNV segmentation.
Jie Guo 0012, Liangyun Sun, Lishan Qiao, Xiushan Nie, Jixin Yang, Weicui Li, Ying Guo 0030, Xiaoming Xi, Xinjian Chen 0001, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.4
2025 STDA: Spatio-Temporal Deviation Alignment Learning for Cross-City Fine-Grained Urban Flow Inference
abstract
Fine-grained urban flow inference (FUFI) is crucial for traffic management, as it infers high-resolution urban flow maps from coarse-grained observations. Existing FUFI methods typically focus on a single city and rely on comprehensive training with large-scale datasets to achieve precise inferences. However, data availability in developing cities may be limited, posing challenges to the development of well-performing models. To address this issue, we propose cross-city fine-grained urban flow inference, which aims to transfer spatio-temporal knowledge from data-rich cities to data-scarce areas using meta-transfer learning. This paper devises a Spatio-Temporal Deviation Alignment (STDA) framework to mitigate spatio-temporal distribution deviations and urban structural deviations between multiple source cities and the target city. Furthermore, STDA presents a cross-city normalization method that adaptively combines batch and instance normalization to maintain consistency between city-variant and city-invariant features. Besides, we design an urban structure alignment module to align spatial topological differences across cities. STDA effectively reduces distribution and structural deviations among different datasets while avoiding negative transfer. Extensive experiments conducted on three real-world datasets demonstrate that STDA consistently outperforms state-of-the-art baselines.
Min Yang 0006, Xiushan Nie, Muming Zhao, Chengqi Zhang, Yu Zheng 0004, Yongshun Gong
IEEE Trans. Knowl. Data Eng.4
2025 Learning Efficient and Adaptive Cross-Channel Dependencies for Weakly-Supervised Object Detection
Xiushan Nie, Yang Ning
IEEE Trans. Multim.2
2025 Online Hashing with Discriminative Attribute Embedding
abstract
Online hashing has emerged as a powerful tool for efficiently processing large-scale and streaming data. However, existing approaches often struggle with scalability limitations in similarity relations and inadequate discrimination provided by one-hot labels. To address these challenges, we propose Online Hashing with Discriminative Attribute Embedding (OHDAE). This novel method leverages a triple-matrix decomposition framework to dynamically decompose features into a dictionary, attributes, and category representations, effectively capturing semantic consistency without relying on accumulated data. To enhance the consistency and discriminability of attributes, we introduce an attribute construction strategy that integrates dictionary constraints with an online optimization strategy. Additionally, fine-grained semantic labels are embedded to improve the discriminability of hash codes by incorporating both semantic and similarity relationships. Experiments conducted on three benchmark datasets validate the superior performance, scalability, and robustness of OHDAE compared to existing state-of-the-art methods.
Xingbo Liu, Zhijie Zhao, Xuening Zhang, Xiao Kang, Xiushan Nie
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Unsupervised Online Cross-modal Hashing With Multiple Association Exploitation
abstract
Unsupervised online cross-modal hashing has gained increasing attention for its effectiveness in streaming data retrieval. However, existing methods primarily focus on exploiting shared properties, overlooking semantic shifts among chunks and specific properties of each modality. To address these challenges, we propose a novel method called Unsupervised Online Cross-Modal Hashing with multiple association exploitation, UOCMH in short. Specifically, we design a hierarchical matrix factorization framework. It skillfully constructs robust orthogonal bases, multi-modality specific representations, and unified common representations, thereby capturing semantic associations among multi-modality streaming data more sufficiently. Additionally, we present a semantic auto-encoder scheme as hash functions. It builds the association between features and hash codes, facilitating the stability of the hashing process. Extensive experiments on the widely-used benchmark datasets demonstrate the superiority of the proposed UOCMH.
Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
ICME5
2024 Fast Multi-view Clustering With Binary Anchor Graph
abstract
Multi-view clustering has achieved remarkable efficacy in integrating multi-view information, and received much research interest. Although anchor-based clustering algorithms have been well-investigated in past years, the separation of graph construction and category partitioning, can lead to suboptimal clustering performance and learning efficiency. To address these challenges, we propose a novel fast clustering algorithm named FAST-BAG. The proposed method can integrate the anchor graph construction and clustering partitioning seamlessly, breaking the separation between data fusion and task processes. Specifically, the multi-view data is unified into a consistent binary anchor graph with linear time complexity. Additionally, we leverage the high efficiency of binary distance computation to expedite the category partitioning process. Experiments conducted on five benchmark datasets validate the effectiveness and efficiency of the proposed method.
Xingbo Liu, Xiao Kang, Xuening Zhang, Xiushan Nie, Yilong Yin
ICME5
2024 Completely Unpaired Cross-Modal Hashing Based on Coupled Subspace
abstract
Unpaired cross-modal hashing which requires no supervision is a promising candidate to support large-scale retrieval across heterogeneous data. However, existing works focus on recovering pairwise relationships, which are usually time-consuming and sensitive to outliers. To tackle this issue, we propose a novel method termed Completely Unpaired Crossmodal Hashing (CUCH), which is applicable to scenarios where neither pairwise correspondence nor label information is available. The proposed CUCH creatively combines the merits of subspace recovery and cross-modal hashing, producing an effective subspace with both robustness and high efficiency. It first discovers robust subspace from each modality by excluding outliers. Then latent space translation is elaborated to obtain coupled subspace, based on which intermodal similarities can be captured. Moreover, the similarity-preservation property for CUCH is guaranteed. By manipulating subspaces rather than pairwise relations, CUCH reduces computational cost significantly. Experimental results demonstrate its advantages in various settings.
Xuening Zhang, Xingbo Liu, Xiao Kang, Xiushan Nie, Yilong Yin
ICME5
2024 Unsupervised Deep Hashing via Sample Weighted Contrastive Learning
abstract
Researchers face a demanding task in the context of unsupervised image retrieval, particularly concerning the challenge of learning discriminative hash codes for the similarity of each sample. The low utilization of the similarity between samples often results in the suboptimal construction of pseudo-labels during contrastive learning, leading to difficulties in accurately differentiating between positive and negative sample pairs. To address this issue, this study proposes a novel contrastive learning method called sample-weighted contrastive hashing (SWCH). In SWCH, a similarity attention module is introduced to exploit the similarity information between samples. By leveraging this approach, a similarity matrix can be constructed, resulting in more-accurate representation of the similarity between samples. Additionally, inspired by the design concept of PolyLoss, a loss term focusing on sample similarity is designed to enhance the model’s ability to capture similarities among data points, thereby generating superior hash codes. In comprehensive experiments, the proposed method demonstrates significant performance improvements in unsupervised image retrieval tasks. Compared with existing methods, our SWCH approach excels in learning discriminative hash codes and enhances the capability of capturing similarities, ultimately achieving superior results in unsupervised image retrieval tasks.
Jingyuan Fang, Xiushan Nie
IJCNN3
2024 Multi-Scale Temporal Relations and Segmented Channel Attention for Video Anomaly Detection
abstract
In recent years, the rapid advancement in video surveillance technology has significantly enhanced public safety and security. In conventional video anomaly detection approaches, there is often an exclusive focus on local information, with key temporal dynamics being overlooked. This oversight could potentially lead to a failure in recognizing dynamic anomalies, such as the sudden running of a person or the rapid movement of objects. Therefore, this study proposes a model framework structure called MTR-SCA. By utilizing widerresnet38 and Multi-Scale Temporal Relations (MTR) to capture the multi-scale temporal relationships in video time series, the framework achieves an understanding of spatial and temporal information. It introduces the Segmented Channel Attention (SCA) to enhance key information in the input feature maps and suppress less important channels for refined feature selection. We conducted experiments with the MTR-SCA network on three datasets: Avenue, ped2, and ShanghaiTech, achieving results of 97.8%, 86.8%, and 74.1% respectively.
Xiushan Nie, Bryan W. Scotney, Shuai Zhang 0001, Xingbo Liu
IJCNN2
2024 Visual Out-of-Distribution Detection in Open-Set Noisy Environments
Rundong He, Zhongyi Han, Xiushan Nie, Yilong Yin, Xiaojun Chang
Int. J. Comput. Vis.3
2024 Multi-axis interactive multidimensional attention network for vehicle re-identification
Xiyu Pang, Yanli Zheng, Xiushan Nie, Yilong Yin
Image Vis. Comput.3
2024 Weighted cross-modal hashing with label enhancement
Yongxin Wang 0001, Kuikui Wang, Xiushan Nie, Zhen-Duo Chen 0001
Knowl. Based Syst.4
2024 A survey of micro-video analysis
Jie Guo 0012, Yuling Ma, Meng Liu 0006, Xiaoming Xi, Xiushan Nie, Yilong Yin
Multim. Tools Appl.6
2024 Heterogeneous context interaction network for vehicle re-identification
Xiyu Pang, Meifeng Zheng, Xiushan Nie, Houren Zhou, Yilong Yin
Neural Networks4
2024 Demsasa: micro-video scene classification based on denoising multi-shots association self-attention
Jie Guo 0012, Xiushan Nie
Pattern Anal. Appl.6
2024 Discrete online cross-modal hashing with consistency preservation
Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
Pattern Recognit.5
2024 Discriminative atoms embedding relation dual network for classification of choroidal neovascularization in OCT images
Xiaoming Xi, Longsheng Xu, Xiushan Nie, Jianhua Nie, Xianjing Meng, Xinjian Chen 0001, Yilong Yin
Pattern Recognit.5
2024 Scalable Unsupervised Hashing via Exploiting Robust Cross-Modal Consistency
abstract
Unsupervised cross-modal hashing has received increasing attention because of its efficiency and scalability for large-scale data retrieval and analysis. However, existing unsupervised cross-modal hashing methods primarily focus on learning shared feature embedding, ignoring robustness and consistency across different modalities. To this end, this study proposes a novel method called scalable unsupervised hashing (SUH) for large-scale cross-modal retrieval. In the proposed method, latent semantic information and common semantic embedding within heterogeneous data are simultaneously exploited using multimodal clustering and collective matrix factorization, respectively. Furthermore, the robust norm is seamlessly integrated into the two processes, making SUH insensitive to outliers. Based on the robust consistency exploited from the latent semantic information and feature embedding, hash codes can be learned discretely to avoid cumulative quantitation loss. The experimental results on five benchmark datasets demonstrate the effectiveness of the proposed method under various scenarios.
Xingbo Liu, Jiamin Li 0003, Xiushan Nie, Xuening Zhang, Yilong Yin
IEEE Trans. Big Data3
2024 Online Discriminative Cross-Modal Hashing
abstract
Online cross-modal hashing has received increasing research attention due to its capability of encoding streaming data and updating hash functions simultaneously. Despite significant progress, there is still room for further improving accuracy from two aspects,i.e., 1) enhancing discrimination of hash codes with an efficient training process; 2) elevating generalization performance by harmonizing the training and retrieval process. Inspired by this, we propose an Online Discriminative Cross-modal Hashing method, called ODCH. To enlarge the inter-class margin and magnify the intra-class similarity, ODCH skillfully constructs a discriminative semantic space and seamlessly integrates bit balance and uncorrelation constraints, discrete optimization, and asymmetric strategy for embedding the discriminative semantic information into hamming space. Furthermore, ODCH attempts to boost the generalization process by bridging the gap between learning and generalization. It develops adaptive bit-wise weights to reflect different learning conditions among bits and transmits them into the generalization process. Besides, the proposed discriminative embedding and adaptive weighting can be adopted by existing supervised cross-modal hashing methods, achieving more precise performance than the original versions. Extensive experiments on three benchmarked datasets show that ODCH achieves up to an average of 4.17% mAP score gains compared to state-of-the-art online cross-modal hashing methods, indicating its superiority.
Xiao Kang, Xingbo Liu, Xuening Zhang, Xiushan Nie, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.4
2024 Multi-Branch GAN-Based Abnormal Events Detection via Context Learning in Surveillance Videos
abstract
Video anomaly detection is an important task in the field of intelligent security. However, existing methods mainly detect and analyze videos from a single time direction, ignoring the semantic information of the video context, which adversely affects the detection accuracy. To address this issue, we design a multi-branch generative adversarial network with context learning (MGAN-CL) to detect abnormal events. In particular, we combine video context information to generate predicted frames, and determine whether an anomaly occurs by comparing the predicted frame with the actual frame. Different from the existing GAN-based methods, in the anomaly event detection stage, we use the discriminator to judge the video frames generated by the generator, which improves the accuracy of anomaly detection. In order to improve the ability of the discriminator, a pseudo-anomaly module is added to the discriminator for data augmentation to improve the robustness of the model. An extensive set of experiments performed on public datasets demonstrate the method’s superior performance.
Daoheng Li, Xiushan Nie, Ximing Lin, Hui Yu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Semi-Supervised Semi-Paired Cross-Modal Hashing
abstract
Large-scale cross-modal hashing has drawn extensive attention due to its attractive efficiency in both storage and retrieval. Existing methods exhibit poor performance when exploiting the semantic correlations implied in unsupervised and unpaired data during training process. To deal with this issue, we propose a novel hashing method, named Semi-supervised Semi-paired Cross-modal Hashing (SSCH). By leveraging a general and flexible two-step scheme, the proposed method can handle the complex training data effectively and efficiently, where both the common semantics and the modality-specific optimal pseudo semantics are well captured. Specifically, the proposed SSCH performs an alignment-free pseudo-labeling process to get strengthened semantic information. Furthermore, hash representations for various data are learned via a label-enhanced strategy, through which the cross-modal correlations are strengthened and preserved with considering efficiency. The semantic-preserving proof of SSCH is given based on statistical analysis. Also, we prove the stability of the proposed time-saving algorithm using properties of Bregman divergence. Experimental results on three benchmark datasets show that SSCH can obtain satisfactory precision and scalability in various scenarios.
Xuening Zhang, Xingbo Liu, Xiushan Nie, Xiao Kang, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.3
2024 Biomarkers-Aware Asymmetric Bibranch GAN With Adaptive Memory Batch Normalization for Prediction of Anti-VEGF Treatment Response in Neovascular Age-Related Macular Degeneration
abstract
The emergence of anti-vascular endothelial growth factor (anti-VEGF) therapy has revolutionized neovascular age-related macular degeneration (nAMD). Post-therapeutic optical coherence tomography (OCT) imaging facilitates the prediction of therapeutic response to anti-VEGF therapy for nAMD. Although the generative adversarial network (GAN) is a popular generative model for post-therapeutic OCT image generation, it is realistically challenging to gather sufficient pre- and post-therapeutic OCT image pairs, resulting in overfitting. Moreover, the available GAN-based methods ignore local details, such as the biomarkers that are essential for nAMD treatment. To address these issues, a Biomarkers-aware Asymmetric Bibranch GAN (BAABGAN) is proposed to efficiently generate post-therapeutic OCT images. Specifically, one branch is developed to learn prior knowledge with a high degree of transferability from large-scale data, termed the source branch. Then, the source branch transfer knowledge to another branch, which is trained on small-scale paired data, termed the target branch. To boost the transferability, a novel Adaptive Memory Batch Normalization (AMBN) is introduced in the source branch, which learns more effective global knowledge that is impervious to noise via memory mechanism. Also, a novel Adaptive Biomarkers-aware Attention (ABA) module is proposed to encode biomarkers information into latent features of target branches to learn finer local details of biomarkers. The proposed method outperforms traditional GAN models and can produce high-quality post-treatment OCT pictures with limited data sets, as shown by the results of experiments.
Peng Zhao 0016, Xian Song, Xiaoming Xi, Xiushan Nie, Xianjing Meng, Yilong Yin
IEEE J. Biomed. Health Informatics4
2024 Learning Feature Semantic Matching for Spatio-Temporal Video Grounding
abstract
Spatio-temporal video grounding (STVG) aims to localize a spatio-temporal tube, including temporal boundaries and object bounding boxes, that semantically corresponds to a given language description in an untrimmed video. The existing onestage solutions in this task face two significant challenges, namely, vision-text semantic misalignment and spatial mislocalization, which limit their performance in grounding. These two limitations are mainly caused by neglect of fine-grained alignment in crossmodality fusion and the reliance on a text-agnostic query in sequentially spatial localization. To address these issues, we propose an effective model with a newly designed Feature Semantic Matching (FSM) module based on a Transformer architecture to address the above issues. Our method introduces a crossmodal feature matching module to achieve multi-granularity alignment between video and text while preventing the weakening of important features during the feature fusion stage. Additionally, we design a query-modulated matching module to facilitate text-relevant tube construction by multiple query generation and tubulet sequence matching. To ensure the quality of tube construction, we employ a novel mismatching rectify contrastive loss to rectify the mismatching between the learnable query and the objects corresponding to the text descriptions by restricting the generated spatial query. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods on two challenging STVG benchmarks.
Hao Fang 0010, Hao Zhang 0048, Jialin Gao, Xiankai Lu, Xiushan Nie, Yilong Yin
IEEE Trans. Multim.6
2024 Relational Network via Cascade CRF for Video Language Grounding
abstract
Video Language Grounding is one of the most challenging cross-modal video understanding tasks. This task aims to localize a target moment semantically corresponding to a given language query in an untrimmed video. Many existing VLG methods rely on the proposal-based framework, despite the dominant performance achieved, they usually focus on interacting a few internal frames with the query to score segment proposals, trapping in the long-range dependencies when the proposal feature is limited. Meanwhile, adjacent proposals share similar visual semantics, making VLG models hard to align the accurate semantics of video-query contents and degenerating the ranking performance. To remedy the above limitations, we propose VLG-CRF by introducing the conditional random fields (CRFs) to handle the discrete yet indistinguishable proposals. Specifically, VLG-CRF consists of two cascade CRF-based modules. The AttentiveCRFs is developed for multi-modal feature fusion to better integrate temporal and semantic relation between modalities. We also devise a new variant of ConvCRFs to capture the relation of discrete segments and rectify the predicting scores to make relatively high prediction scores clustered in a range. Experiments on three benchmark datasets,i.e., Charades-STA, ActivityNet-Caption, and TACoS, show the superiority of our method and the state-of-the-art performance is achieved.
Xiankai Lu, Hao Zhang 0048, Xiushan Nie, Yilong Yin, Jianbing Shen
IEEE Trans. Multim.4
2024 Online Cross-modal Hashing With Dynamic Prototype
abstract
Online cross-modal hashing has received increasing attention due to its efficiency and effectiveness in handling cross-modal streaming data retrieval. Despite the promising performance, these methods mainly focus on the supervised learning paradigm, demanding expensive and laborious work to obtain clean annotated data. Existing unsupervised online hashing methods mostly struggle to construct instructive semantic correlations among data chunks, resulting in the forgetting of accumulated data distribution. To this end, we propose a Dynamic Prototype-based Online Cross-modal Hashing method, called DPOCH. Based on the pre-learned reliable common representations, DPOCH generates prototypes incrementally as sketches of accumulated data and updates them dynamically for adapting streaming data. Thereafter, the prototype-based semantic embedding and similarity graphs are designed to promote stability and generalization of the hashing process, thereby obtaining globally adaptive hash codes and hash functions. Experimental results on benchmarked datasets demonstrate that the proposed DPOCH outperforms state-of-the-art unsupervised online cross-modal hashing methods.
Xiao Kang, Xingbo Liu, Xiushan Nie, Yilong Yin
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Fast Unsupervised Cross-Modal Hashing with Robust Factorization and Dual Projection
abstract
Unsupervised hashing has attracted extensive attention in effectively and efficiently tackling large-scale cross-modal retrieval task. Existing methods typically try to mine the latent common subspace across multimodal data without any category annotation. Despite the exciting progress, there are still three challenges that need to be further addressed: (1) efficiently improving the robustness during latent common subspace learning; (2) harmoniously embedding the intra-modal inherence and inter-modal relevance of multimodal data into Hamming space; and (3) effectively reducing the training time complexity and making the model scalable for large-scale datasets. To well address the above challenges, this study proposes a method named Fast Unsupervised Cross-Modal Hashing (FUCH). Specifically, FUCH proposes a semantic-aware collective matrix factorization to learn robust representation via exploiting latent category-specific attributes, and introduces Cauchy loss to measure the factorization process. Accordingly, the above process can effectively embed potential discriminative information into common space, while making the model insensitive for outliers. Moreover, FUCH designs a dual projection learning scheme, which not only learns modality-unique hash functions to excavate individual properties but also learns modality-mutual hash functions to multimodal correlational properties. Experimental results on three benchmark datasets verify the effectiveness of FUCH under various scenarios.
Xingbo Liu, Jiamin Li 0003, Xiushan Nie, Xuening Zhang, Yilong Yin
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Complex Scenario Image Retrieval via Deep Similarity-aware Hashing
abstract
When performing hashing-based image retrieval, it is difficult to learn discriminative hash codes especially for the multi-label, zero-shot and fine-grained settings. This is due to the fact that the similarities vary, even within the same category, under the conditions of complex scenario settings. To address this problem, this study develops a deep similarity-aware hashing method for complex scenario image retrieval (DEPISH). DEPISH more focuses on the samples that are difficult to distinguish from other images (i.e., “difficult samples”), such as images that contain multiple semantics. It dynamically divides attention among samples according to their difficulty levels with a margin weighting strategy. Furthermore, by adding special terms in the model, DEPISH is capable of avoiding the inconsistency between the hash code representation and true similarity among negative samples. In addition, unlike the existing methods that use a pre-defined similarity matrix with fixed values, the DEPISH adopts an adaptive similarity matrix, which accurately captures the various similarities among all samples. The results of our experiment on multiple benchmark datasets containing complex scenarios (i.e., multi-label, zero-shot, and fine-grained datasets) verify the effectiveness of this method.
Xiushan Nie, Weili Guan, Yilong Yin
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Exposing the Self-Supervised Space-Time Correspondence Learning via Graph Kernels
abstract
Self-supervised space-time correspondence learning is emerging as a promising way of leveraging unlabeled video. Currently, most methods adapt contrastive learning with mining negative samples or reconstruction adapted from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence predictive models require mining more inherent properties in videos, such as structural information. In this work, we propose the VideoHiGraph, a space-time correspondence framework based on a learnable graph kernel. Concerning the video as the spatial-temporal graph, the learning objectives of VideoHiGraph are emanated in a self-supervised manner for predicting unobserved hidden graphs via graph kernel manner. We learn a representation of the temporal coherence across frames in which pairwise similarity defines the structured hidden graph, such that a biased random walk graph kernel along the sub-graph can predict long-range correspondence. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The self-supervision of the model training is formed by the structural and temporal consistency of the graph. VideoHiGraph achieves superior performance and demonstrates its robustness across the benchmark of label propagation tasks involving objects, semantic parts, keypoints, and instances. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/VideoHiGraph.
Zheyun Qin, Xiankai Lu, Xiushan Nie, Yilong Yin, Jianbing Shen
AAAI3
2023 Multi-View Representation Learning via View-Aware Modulation
abstract
Multi-view (representation) learning derives an entity's representation from its multiple observable views to facilitate various downstream tasks. The most challenging topic is how to model unobserved entities and their relationships to specific views. To this end, this work proposes a novel multi-view learning method using a View-Aware parameter Modulation mechanism, termed VAM. The key idea is to use trainable parameters as proxies for unobserved entities and views, such that modeling entity-view relationships is converted into modeling the relationship between proxy parameters. Specifically, we first build a set of trainable parameters to learn a mapping from multi-view data to the unified representation as the entity proxy. Then we learn a prototype for each view and design a Modulation Parameter Generator (MPG) that learns a set of view-aware scale and shift parameters from prototypes to modulate the entity proxy and obtain view proxies. By constraining the representativeness, uniqueness, and simplicity of the proxies and proposing an entity-view contrastive loss, parameters are alternatively updated. We end up with a set of discriminative prototypes, view proxies, and an entity proxy that are flexible enough to yield robust representations for out-of-sample entities. Extensive experiments on five datasets show that the results of our VAM outperform existing methods in both classification and clustering tasks.
Ren Wang 0011, Haoliang Sun, Xiushan Nie, Yuxiu Lin, Xiaoming Xi, Yilong Yin
ACM Multimedia3
2023 Unified 3D Segmenter As Prototypical Classifiers
abstract
The task of point cloud segmentation, comprising semantic, instance, and panoptic segmentation, has been mainly tackled by designing task-specific network architectures, which often lack the flexibility to generalize across tasks, thus resulting in a fragmented research landscape. In this paper, we introduce ProtoSEG, a prototype-based model that unifies semantic, instance, and panoptic segmentation tasks. Our approach treats these three homogeneous tasks as a classification problem with different levels of granularity. By leveraging a Transformer architecture, we extract point embeddings to optimize prototype-class distances and dynamically learn class prototypes to accommodate the end tasks. Our prototypical design enjoys simplicity and transparency, powerful representational learning, and ad-hoc explainability. Empirical results demonstrate that ProtoSEG outperforms concurrent well-known specialized architectures on 3D point cloud benchmarks, achieving 72.3%, 76.4% and 74.2% mIoU for semantic segmentation on S3DIS, ScanNet V2 and SemanticKITTI, 66.8% mCov and 51.2% mAP for instance segmentation on S3DIS and ScanNet V2, 62.4% PQ for panoptic segmentation on SemanticKITTI, validating the strength of our concept and the effectiveness of our algorithm. The code and models are available at https://github.com/zyqin19/PROTOSEG.
Zheyun Qin, Cheng Han 0001, Qifan Wang 0001, Xiushan Nie, Yilong Yin, Xiankai Lu
NeurIPS4
2023 Vehicle re-identification based on grouping aggregation attention and cross-part interaction
Xiyu Pang, Xiushan Nie, Yilong Yin, Gangwu Jiang
J. Vis. Commun. Image Represent.3
2023 Deep regional detail-aware hashing
Yuling Ma, Jie Guo 0012, Xiushan Nie, Yilong Yin
Multim. Syst.5
2023 Detail enhancement-based vehicle re-identification with orientation-guided re-ranking
Ziruo Sun, Xiushan Nie, Xiaopeng Bi, Yilong Yin
Pattern Recognit.2
2023 Triple-attention interaction network for breast tumor classification based on multi-modality images
Xiaoming Xi, Kesong Wang, Liangyun Sun, Lingzhao Meng, Xiushan Nie, Lishan Qiao, Yilong Yin
Pattern Recognit.6
2023 Supervised Discrete Multiple-Length Hashing for Image Retrieval
abstract
Hashing can facilitate efficient retrieval and storage for large-scale images due to the binary representation. In the real applications, the trade-off between retrieval accuracy and speed is essential for designing a hashing framework, which is reflected by variable hash code lengths. In light of this, the existing hashing methods need to train different models for different lengths of hash codes, leading to considerable training time cost and hashing flexibility reduction. Given that a sample can be represented by various hash codes with different lengths, there are some helpful relationships that can boost the performance of hashing methods. However, the existing hashing methods do not fully utilize these relationships. To address the aforementioned issues, we propose a new model, known as supervised discrete multiple-length hashing (SDMLH), to simultaneously learn hash codes with multiple lengths. In this proposed SDMLH method, three types of information are respectively derived, from the hash codes with different lengths. The original features of the samples, and the label, are applied for hash learning. Unlike the existing hashing methods, SDMLH can fully employ the assistance among hash codes with different lengths and learn them in one step. Furthermore, given a hash length meeting the demand of users, we propose a hash fusion strategy to obtain the hash code with this desirable length by fusing the multiple-length hash codes. This obtained hash code outperforms the one learned directly. In addition, SDMLH can generate the hash code of any length that is shorter than the sum length of given multiple hash codes with the fusion strategy. To the best of our knowledge, SDMLH is one of the first attempts for learning multiple-length hash codes simultaneously. We conduct extensive experiments based on three benchmark datasets, demonstrating the superiority of this proposed method.
Xiushan Nie, Xingbo Liu, Jie Guo 0012, Yilong Yin
IEEE Trans. Big Data1
2023 Reformulating Graph Kernels for Self-Supervised Space-Time Correspondence Learning
abstract
Self-supervised space-time correspondence learning utilizing unlabeled videos holds great potential in computer vision. Most existing methods rely on contrastive learning with mining negative samples or adapting reconstruction from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence prediction models need to uncover more inherent properties of the video, such as structural information. In this work, we propose HiGraph+, a sophisticated space-time correspondence framework based on learnable graph kernels. By treating videos as a spatial-temporal graph, the learning objective of HiGraph+ is issued in a self-supervised manner, predicting the unobserved hidden graph via graph kernel methods. First, we learn the structural consistency of sub-graphs in graph-level correspondence learning. Furthermore, we introduce a spatio-temporal hidden graph loss through contrastive learning that facilitates learning temporal coherence across frames of sub-graphs and spatial diversity within the same frame. Therefore, we can predict long-term correspondences and drive the hidden graph to acquire distinct local structural representations. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The structural and temporal consistency of the graph forms the self-supervision of model training. HiGraph+ achieves excellent performance and demonstrates robustness in benchmark tests involving object, semantic part, keypoint, and instance labeling propagation tasks. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/HiGraph.
Zheyun Qin, Xiankai Lu, Dongfang Liu, Xiushan Nie, Yilong Yin, Jianbing Shen, Alexander C. Loui
IEEE Trans. Image Process.4
2023 Zero-Shot Hashing via Asymmetric Ratio Similarity Matrix
abstract
Zero-shot hashing targets to learn the hash codes of images in unseen classes based on the limited training data provided by seen classes. In zero-shot hashing, transferring the supervised knowledge, such as attributes and semantic relations, from seen classes to unseen ones is a widely employed method, where the performance is always subject to the ability to capture these supervised knowledge (which is always difficult to obtain). Therefore, in this study, we propose a new methodology for zero-shot hashing via an asymmetric ratio similarity matrix (ASZH), which only needs to calculate the semantic similarity among seen classes for hash learning. Specifically, we use an asymmetric ratio matrix in the similarity calculation to further explore the influence of similarity, where the values of positive weights for similar samples are not equivalent to those of negative ones for dissimilar samples. Additionally, a theoretical analysis regarding the utilization of an asymmetric ratio matrix is provided in this study. The experiments on three large benchmark datasets indicate that the proposed method achieves excellent performance than several state-of-the-art hashing methods.
Xiushan Nie, Xingbo Liu, Lu Yang 0005, Yilong Yin
IEEE Trans. Knowl. Data Eng.2
2022 Local Detail Enhancement Network for CNV Typing in OCT Images
abstract
Choroidal neovascularization (CNV) is one of the severe eye disease. The severe results will cause of loss of acuity, scotomata, and distortion of vision. Automatic and accurate classification of CNV with optical coherence tomography (OCT) images can assist doctors in treatment. However, the existing methods ignore the fact that semantic feature maps, used for classification, lose much feature detail information. Therefore, we proposed a local detail enhancement network for CNV classification, which includes both progressive training mode and local detail enhancement (LDE) module. With the progressive training mode, the learned features fuse shallow and stable fine-grained information with high-level semantic information, which promote the diversity of the learned features. In LDE module, the detail feature learn (DFL) module is introduced to learn the underlying detail information and embed it into the semantic feature map. The semantic feature map with detail information is propitious to capture the subtle discrepancy between different CNV types and promote the classification performance. Sufficient experiments are performed on our self-build CNV dataset. Our method excelled existing methods and in evaluation indicators ACC, AUC, SEN, and SPE are 92.3%, 87.1%, 91.5%, and 90.9%.
Chuanzhen Xu, Xiaoming Xi, Liangyun Sun, Lingzhao Meng, Xiushan Nie
HSI6
2022 Attention-based Interactions Network for Breast Tumor Classification with Multi-modality Images
abstract
Benefiting from the development of medical imaging, the automatic breast image classification has been extensively studied in a variety of breast cancer diagnosis tasks recently. The multi-modality image fusion was helpful to further improve classification performance. However, existing multi-modality fusion methods focused on the fusion of modalities, ignoring the interactions between modalities, which caused the inefficient performance. To address the above issues, we proposed a novel attention-based interactions network for breast tumor classification by using diffusion-weighted imaging (DWI) and apparent dispersion coefficient (ADC) images. Specifically, we proposed a multi-modality interaction mechanism, including relational interaction, channel interaction, and discriminative interaction, to design an attention-based interaction module, which enhanced the abilities of inter-modal interactions. Extensive ablation studies have been carried out, which provably affirmed the advantages of each component. The area under the receiver operating characteristic curve (AUC), accuracy (ACC), specificity (SPC), and sensitivity (SEN) were 87.0%, 87.0%, 88.0%, and 86.0%, respectively, also verifying its effectiveness.
Xiaoming Xi, Chuanzhen Xu, Liangyun Sun, Lingzhao Meng, Xiushan Nie
HSI6
2022 Exploring Linear Feature Disentanglement for Neural Networks
abstract
Non-linear activation functions, e.g., Sigmoid, ReLU, and Tanh, have achieved great success in neural networks (NNs). Due to the complex non-linear characteristic of samples, the objective of those activation functions is to project samples from their original feature space to a linear separable feature space. This phenomenon ignites our interest in exploring whether all features need to be transformed by all nonlinear functions in current typical NNs, i.e., whether there exists a part of features arriving at the linear separable feature space in the intermediate layers, that does not require further non-linear variation but an affine transformation instead. To validate the above hypothesis, we explore the problem of linear feature disentanglement for neural networks in this paper. Specifically, we devise a learnable mask module to distinguish between linear and non-linear features. Through our designed experiments we found that some features reach the linearly separable space earlier than the others and can be detached partly from the NNs. The explored method also provides a readily feasible pruning strategy which barely affects the performance of the original model. We conduct our experiments on four datasets and present promising results.
Tiantian He 0004, Zhibin Li 0002, Yongshun Gong, Yazhou Yao, Xiushan Nie, Yilong Yin
ICME5
2022 Series Photo Selection via Multi-View Graph Learning
abstract
Series photo selection (SPS) is an important branch of the image aesthetics quality assessment, which focuses on finding the best one from a series of nearly identical photos. While a great progress has been observed, most of the existing SPS approaches concentrate solely on extracting features from the original image, neglecting that multiple views, e.g, saturation level, color histogram and depth of field of the image, will be of benefit to successfully reflecting the subtle aesthetic changes. Taken multi-view into consideration, we leverage a graph neural network to construct the relationships between multi-view features. Besides, multiple views are aggregated with an adaptive-weight self-attention module to verify the significance of each view. Finally, a siamese network is proposed to select the best one from a series of nearly identical photos. Experimental results demonstrate that our model accomplish the highest success rates compared with competitive methods.
Lu Zhang 0062, Yongshun Gong, Jian Zhang 0002, Xiushan Nie, Yilong Yin
ICME5
2022 Difficulty-aware bi-network with spatial attention constrained graph for axillary lymph node segmentation
Xiaoming Xi, Xianjing Meng, Zheyun Qin, Xiushan Nie, Yongjian Wu 0001, Chenglong Li 0004, Yilong Yin
Sci. China Inf. Sci.5
2022 SNIP-FSL: Finding task-specific lottery jackpots for few-shot learning
Ren Wang 0011, Haoliang Sun, Xiushan Nie, Yilong Yin
Knowl. Based Syst.3
2022 Context-related video anomaly detection via generative adversarial network
Daoheng Li, Xiushan Nie, Yilong Yin
Pattern Recognit. Lett.2
2022 Supervised discrete hashing for hamming space retrieval
Xiao Kang, Fasheng Liu, Xiushan Nie, Xingbo Liu
Pattern Recognit. Lett.4
2022 Supervised Adaptive Similarity Matrix Hashing
abstract
Compact hash codes can facilitate large-scale multimedia retrieval, significantly reducing storage and computation. Most hashing methods learn hash functions based on the data similarity matrix, which is predefined by supervised labels or a distance metric type. However, this predefined similarity matrix cannot accurately reflect the real similarity relationship among images, which results in poor retrieval performance of hashing methods, especially in multi-label datasets and zero-shot datasets that are highly dependent on similarity relationships. Toward this end, this study proposes a new supervised hashing method called supervised adaptive similarity matrix hashing (SASH) via feature-label space consistency. SASH not only learns the similarity matrix adaptively, but also extracts the label correlations by maintaining consistency between the feature and the label space. This correlation information is then used to optimize the similarity matrix. The experiments on three large normal benchmark datasets (including two multi-label datasets) and three large zero-shot benchmark datasets show that SASH has an excellent performance compared with several state-of-the-art techniques.
Xiushan Nie, Xingbo Liu, Yilong Yin
IEEE Trans. Image Process.2
2022 Learning Binary Semantic Embedding for Large-Scale Breast Histology Image Analysis
abstract
With the progress of clinical imaging innovation and machine learning, the computer-assisted diagnosis of breast histology images has attracted broad attention. Nonetheless, the use of computer-assisted diagnoses has been blocked due to the incomprehensibility of customary classification models. In view of this question, we propose a novel method for Learning Binary Semantic Embedding (LBSE). In this study, bit balance and uncorrela-tion constraints, double supervision, discrete optimization and asymmetric pairwise similarity are seamlessly integrated for learning binary semantic-preserving embedding. Moreover, a fusion-based strategy is carefully designed to handle the intractable problem of parameter setting, saving huge amounts of time for boundary tuning. Based on the above-mentioned proficient and effective embedding, classification and retrieval are simultaneously performed to give interpretable image-based deduction and model helped conclusions for breast histology images. Extensive experiments are conducted on three benchmark datasets to approve the predominance of LBSE in different situations.
Xingbo Liu, Xiao Kang, Xiushan Nie, Jie Guo 0012, Yilong Yin
IEEE J. Biomed. Health Informatics3
2022 Regularized Two Granularity Loss Function for Weakly Supervised Video Moment Retrieval
abstract
Weakly supervised video moment retrieval or weakly supervised language moment retrieval aims to search the most relevant moment given a language query. In order to guide the model to capture the most matching video segments with the text description, we design a two-granularity loss function that simultaneously considers both video-level and instance-level relationships. Specifically, we first generate coarse video segments and regard each video segment as an instance. For video-level regularized multiple instance loss (MIL), we leverage the latent alignment between all intra-video segments (ie., positive bag) and text descriptions. Then, we classify these segments by regarding this procedure as a supervised learning task under noisy labels. With the instance-level regularized loss function, our model can learn to correct noisy instance-level labels so as to locate the more accurate frame boundary from all the positive instances. Comprehensive experimental results onActivityNetandDiDeModemonstrate that the proposed loss function sets a new state-of-the-art.
Junya Teng, Xiankai Lu, Yongshun Gong, Xinfang Liu, Xiushan Nie, Yilong Yin
IEEE Trans. Multim.5
2021 Learning Binary Semantic Embedding for Breast Histology Image Classification and Retrieval
abstract
With the development of medical imaging technology and machine learning, the computer-assisted diagnosis has attracted extensive research attention, which can provide beneficial reference to pathologists. However, the exponential growth of medical images and uninterpretability of traditional classification models have hindered the applications of the computer-assisted diagnosis. To address this issues, we propose a novel method for Learning Binary Semantic Embedding (LBSE). Based on this efficient and effective embedding, classification and retrieval are performed to provide interpretable computer-assisted diagnosis for histology images. Furthermore, double supervision, bit uncorrelation and balance constraint, asymmetric strategy and discrete optimization are seamlessly integrated in the proposed method for learning binary embedding. Experiments conducted on three benchmark datasets validate the superiority of LBSE under various scenarios.
Xiao Kang, Xingbo Liu, Xiushan Nie, Yilong Yin
ICASSP3
2021 Joint Learning of Image Aesthetic Quality Assessment and Semantic Recognition Based on Feature Enhancement
abstract
Aesthetic quality assessment and semantic recognition are the two fundamental aspects of image perception and understanding tasks. Though these two tasks are related, most of the current research generally treats them as independent problems without any interaction. In this paper, we explore the relationships between aesthetic quality assessment and semantic recognition task, and employ a multi-task convolutional neural network with feature enhancement mechanism to effectively integrate these two tasks. A novel Enhanced Aggregation of Features Network (EAFNet) for joint learning of the two tasks is proposed to enhance the valid features and suppress the invalid features of each task in both channel and spatial dimensions. Experiments conducted on two benchmark datasets well verify the superior performance of EAFNet in handling aesthetic quality assessment and semantic recognition tasks.
Xiangfei Liu, Xiushan Nie, Zhen Shen 0001, Yilong Yin
ICASSP2
2021 ECCL: Explicit Correlation-Based Convolution Boundary Locator for Moment Localization
abstract
Moment localization in videos using natural language refers to finding the most relevant segment from the video with given a query in natural language form. In this paper, we present a new boundary-determining strategy called explicit correlation-based convolution boundary locator (ECCL), which can handle any lengths of videos and moments while leveraging fine-grained matching relationships. In this method, we first train a deep network to obtain the correlation scores between video clips and query statements. Subsequently, with the correlation scores, we utilize a convolution kernel to generate the boundary probability distribution. Finally, the start and end time indexes of the video moment are calculated with an optimization problem. Experiments on two publicly available datasets demonstrate the feasibility of ECCL.
Xinfang Liu, Xiushan Nie, Junya Teng, Fanchang Hao, Yilong Yin
ICASSP2
2021 Learning Hierarchical Embedding for Video Instance Segmentation
abstract
In this paper, we address video instance segmentation using a new generative model that learns effective representations of the target and background appearance. We propose to exploit hierarchical structural embedding over spatio-temporal space, which is compact, powerful, and flexible in contrast to current tracking-by-detection methods. Specifically, our model segments and tracks instances across space and time in a single forward pass, which is formulated as hierarchical embedding learning. The model is trained to locate the pixels belonging to specific instances over a video clip. We firstly take advantage of a novel mixing function to better fuse spatio-temporal embeddings. Moreover, we introduce normalizing flows to further improve the robustness of the learned appearance embedding, which theoretically extends conventional generative flows to a factorized conditional scheme. Comprehensive experiments on the video instance segmentation benchmark, i.e., YouTube-VIS, demonstrate the effectiveness of the proposed approach. Furthermore, we evaluate our method on an unsupervised video object segmentation dataset to demonstrate its generalizability.
Zheyun Qin, Xiankai Lu, Xiushan Nie, Xiantong Zhen, Yilong Yin
ACM Multimedia3
2021 An Efficient Bus Crowdedness Classification System
abstract
We propose an efficient bus crowdedness classification system that can be used in daily life. In particular, we analyze and study the data collected from real bus, aiming to deal with the difficulty of bus congestion classification. Besides, we combine deep learning and computer vision technology to extract images or videos from the internal surveillance cameras of the bus. The information of crowd will finally be integrated with algorithms into a complete classification system. As a consequence, when the user enters the system and submits the image or video to be detected, the system will display the classification results in turn. The classification results include passenger density distribution, number of passengers, date, and algorithm running time. In addition, the user can use the mouse to delineate an area in the passenger density distribution map and count any image area.
Lingcan Meng, Xiushan Nie, Zhifang Tan
MMAsia2
2021 Deep Adaptive Attention Triple Hashing
abstract
Recent studies have verified that learning compact hash codes can facilitate big data retrieval processing. In particular, learning the deep hash function can greatly improve the retrieval performance. However, the existing deep supervised hashing algorithm treats all the samples in the same way, which leads to insufficient learning of difficult samples. Therefore, we cannot obtain the accurate learning of the similarity relation, making it difficult to achieve satisfactory performance. In light of this, this work proposes a deep supervised hashing model, called deep adaptive attention triple hashing (DAATH), which weights the similarity prediction scores of positive and negative samples in the form of triples, thus giving different degrees of attention to different samples. Compared with the traditional triple loss, it places a greater emphasis on the difficult triple, dramatically reducing the redundant calculation. Extensive experiments have been conducted to show that DAAH consistently outperforms the state-of-the-arts, confirmed its the effectiveness.
Xiushan Nie, Yilong Yin
MMAsia2
2021 Deep Multiple Length Hashing via Multi-task Learning
abstract
Hashing can compress heterogeneous high-dimensional data into compact binary codes. For most existing hash methods, they first predetermine a fixed length for the hash code and then train the model based on this fixed length. However, when the task requirements change, these methods need to retrain the model for a new length of hash codes, which increases time cost. To address this issue, we propose a deep supervised hashing method, called deep multiple length hashing(DMLH), which can learn multiple length hash codes simultaneously based on a multi-task learning network. This proposed DMLH can well utilize the relationships with a hard parameter sharing-based multi-task network. Specifically, in DMLH, the multiple hash codes with different lengths are regarded as different views of the same sample. Furthermore, we introduce a type of mutual information loss to mine the association among hash codes of different lengths. Extensive experiments have indicated that DMLH outperforms most existing models, verifying its effectiveness.
Xiushan Nie, Xingbo Liu
MMAsia2
2021 Attention based consistent semantic learning for micro-video scene recognition
Jie Guo 0012, Xiushan Nie, Yuling Ma, Kashif Shaheed, Inam Ullah 0002, Yilong Yin
Inf. Sci.2
2021 Discrete hashing with triple supervision learning
Xiao Kang, Fasheng Liu, Xiushan Nie, Xingbo Liu
J. Vis. Commun. Image Represent.4
2021 Supervised discrete hashing through similarity learning
Xingbo Liu, Xiushan Nie
Multim. Tools Appl.3
2021 Reinforced Short-Length Hashing
abstract
Given that retrieval and storage have compelling efficiency, similarity-preserving hashing has been extensively employed to approximate nearest neighbor search in large-scale image retrieval. Hash codes that are extremely compact not only can further lower the storage cost, but also accelerate the retrieval speed. However, existing methods perform poorly in retrieval based on an extremely short-length hash code, which attributes to the weak ability of classification and poor distribution of hash bit. To tackle this issue, in this study, we propose a novel reinforced short-length hashing (RSLH). In particular, this proposed method applies the mutual reconstruction between the hash representation and semantic label to retain the semantic information. Furthermore, to enhance the accuracy of hash representation, a pairwise similarity matrix is designed to make a balance between accuracy and training expenditure on memory. Besides, we integrate a parameter boosting strategy to strengthen the precision with the consideration of bit balance and uncorrelation constraints. Extensive experiments on three large-scale image benchmarks demonstrate the superior performance of RSLH under various short-length hashing scenarios.
Xingbo Liu, Xiushan Nie, Qi Dai 0001, Yupan Huang, Li Lian, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.2
2021 Fast Unmediated Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing is for the purpose of compressing heterogeneous multi-modal data into compact binary codes for the cross-modal retrieval, where accuracy and efficiency are two primary issues. To achieve high accuracy and efficiency, we put forward a novel method named Fast Unmediated Hashing (FUH) for cross-modal retrieval. For this method, motivated by the fact that label vector is a natural binary representation of samples for retrieval, we directly learn the cross-modal hash codes from semantic labels without any intermediate representation. This will capture more relations among different modalities, and reduce the number of variables. However, directly learning hash codes from labels would weaken the discrimination of hash codes. To address this issue, double supervision involving label information and pairwise similarity is proposed to enhance the discrimination. In addition, to decrease the training time, we present a strategy to bypass the similarity matrix-related operation in each iteration of optimization, thus some other related terms can also be computed offline to lower training complexity. Compared to several state-of-the-art techniques on three public datasets, the experimental results have manifested the superiority of FUH concerning efficiency and accuracy.
Xiushan Nie, Xingbo Liu, Xiaoming Xi, Chenglong Li 0004, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.1
2021 Deep Multiscale Fusion Hashing for Cross-Modal Retrieval
abstract
Owing to the rapid development of deep learning and the high efficiency of hashing, hashing methods based on deep learning models have been extensively adopted in the area of cross-modal retrieval. In general, in existing deep model-based methods, modality-specific features play an important role during the hash learning. However, most existing methods only use the modality-specific features from the final fully connected layer, ignoring the semantic relevance among modality-specific features with different scales in multiple layers. To address this issue, in this study, we put forward an end-to-end deep hashing method called deep multiscale fusion hashing (DMFH) for cross-modal retrieval. For the proposed DMFH, we first design different network branches for two modalities and then adopt multiscale fusion models for each branch network to fuse the multiscale semantics, which can be used to explore the semantic relevance. Furthermore, the multi-fusion models also embed the multiscale semantics into the final hash codes, making the final hash codes more representative. In addition, the proposed DMFH can learn common hash codes directly without a relaxation, thereby avoiding a loss in accuracy during hash learning. Experimental results on three benchmark datasets prove the relative superiority of the proposed method.
Xiushan Nie, Bowei Wang, Fanchang Hao, Muwei Jian, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.1
2021 Normality Learning in Multispace for Video Anomaly Detection
abstract
Video anomaly detection is a challenging task owing to the rare and diverse nature of abnormal events. However, most of the existing methods only learn the normality in a single space, focusing on low-level detailed features, which is easily affected by unimportant pixels. To address this issue, in this study, we propose a semi-supervised method based on the generative adversarial network and frame prediction, wherein the normality is learned in both the original image space and latent space, and the events deviating from the normality are detected as anomalies. In particular, given a video clip, we first predict a future frame and minimize the prediction errors between the generated frame and its ground truth. Thereafter, we encode the predicted frames and their ground truths in the latent space and minimize their differences. In the testing phase, we calculate the normal scores of each frame in both the image and latent spaces to obtain a comprehensive evaluation. Utilizing the multispace can capture more normality distribution information of the data, which can benefit anomaly detection. The results of experiments on three benchmark datasets demonstrate the effectiveness of the proposed method.
Xiushan Nie, Rundong He, Meng Chen 0003, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.2
2021 Deep Hashing With Weighted Spatial Importance
abstract
Hashing method has been widely used in big data retrieval because of its low computational complexity. Most of existing hashing methods learn the final hash code from the semantic information of the whole image. However, different spatial regions of an image have different influences during the hash learning. To tackle this issue, we propose a new deep hashing with weighted spatial importance (DWSH) in this paper. Specifically, the proposed DWSH first utilizes a spatial attention model to learn the importance of different spatial regions in the original image, and then assigns different weights to these spatial regions according to their importance. The final hash codes are learned based on the weighted spatial information. In addition, two strategies are designed to utilize the spatial importance, including discrete weight strategy and continuous weight strategy, which weight the spatial information with discrete and continuous values, respectively. The results of extensive experiments conducted on three benchmark datasets show that the proposed DWSH method is superior to the state-of-the-art hashing method based on different evaluation protocols.
Xiushan Nie, Meng Chen 0003, Li Lian, Yilong Yin
IEEE Trans. Multim.2
2021 Single-shot Semantic Matching Network for Moment Localization in Videos
abstract
Moment localization in videos using natural language refers to finding the most relevant segment from videos given a natural language query. Most of the existing methods require video segment candidates for further matching with the query, which leads to extra computational costs, and they may also not locate the relevant moments under any length evaluated. To address these issues, we present a lightweight single-shot semantic matching network (SSMN) to avoid the complex computations required to match the query and the segment candidates, and the proposed SSMN can locate moments of any length theoretically. Using the proposed SSMN, video features are first uniformly sampled to a fixed number, while the query sentence features are generated and enhanced by GloVe, long-term short memory (LSTM), and soft-attention modules. Subsequently, the video features and sentence features are fed to an enhanced cross-modal attention model to mine the semantic relationships between vision and language. Finally, a score predictor and a location predictor are designed to locate the start and stop indexes of the query moment. We evaluate the proposed method on two benchmark datasets and the experimental results demonstrate that SSMN outperforms state-of-the-art methods in both precision and efficiency.
Xinfang Liu, Xiushan Nie, Junya Teng, Li Lian, Yilong Yin
ACM Trans. Multim. Comput. Commun. Appl.2
2020 Focusing on Detail: Deep Hashing Based on Multiple Region Details (Student Abstract)
abstract
Fast retrieval efficiency and high performance hashing, which aims to convert multimedia data into a set of short binary codes while preserving the similarity of the original data, has been widely studied in recent years. Majority of the existing deep supervised hashing methods only utilize the semantics of a whole image in learning hash codes, but ignore the local image details, which are important in hash learning. To fully utilize the detailed information, we propose a novel deep multi-region hashing (DMRH), which learns hash codes from local regions, and in which the final hash codes of the image are obtained by fusing the local hash codes corresponding to local regions. In addition, we propose a self-similarity loss term to address the imbalance problem (i.e., the number of dissimilar pairs is significantly more than that of the similar ones) of methods based on pairwise similarity.
Xiushan Nie, Xingbo Liu, Yilong Yin
AAAI2
2020 Discrete Spatial Importance-Based Deep Weighted Hashing
Xiushan Nie, Xiaoming Xi, Yilong Yin
ACCV (3)2
2020 Deep Multi-Region Hashing
abstract
Hashing has been widely used for large-scale approximate nearest neighbors retrieval own to its high efficiency. In the existing hashing methods, deep supervised hashing methods have achieved the best performance by utilizing the semantic labels on data with deep learning. However, most of these methods only consider the semantics of whole image but ignore the local information which contains much more semantic details. Evidently, the semantic details are beneficial for hash learning. To address this issue, in this paper, we proposed a novel Deep Multi-Region Hashing (DMRH) method to fully utilize the semantic details, which uses overlapping N × N regions of an image to learn N2hash codes for getting a final hash code. Extensive experimental results with three datasets show that DMRH can achieve state-of-the-art performance.
Xiushan Nie, Xingbo Liu, Yilong Yin
ICASSP2
2020 CFVMNet: A Multi-branch Network for Vehicle Re-identification Based on Common Field of View
abstract
Vehicle re-identification (re-ID) aims to retrieve the image of the same vehicles across multiple cameras. It has attracted wide attention in the field of computer vision owing to the deployment of surveillance system. However, some unfavorable factors restrict the retrieval accuracy of re-ID; minor inter-class difference and orientation variation are two main issues. In this study, we proposed a multi-branch network based on common field of view (CFVMNet) to address these issues. In the proposed method, we extracted and fused the global and local detail features using four branches and the Batch DropBlock (BDB) strategy to accentuate inter-class difference. We also considered some other attributes (i.e., color, type, and model) in the feature extraction process to make the final features more recognizable. For the issue of orientation variation that could lead to large intra-class difference, we learned two different metrics according to whether there is common field of view of two vehicle images, respectively, which can enable the proposed CFVMNet to focus on different regions. Extensive experiments on two public datasets, VeRi-776 and VehicleID, show that the proposed method outperformed the state-of-the-art approaches to vehicle re-ID.
Ziruo Sun, Xiushan Nie, Xiaoming Xi, Yilong Yin
ACM Multimedia2
2020 Multi-view feature selection via Nonnegative Structured Graph Learning
Xiangpin Bai, Lei Zhu 0002, Cheng Liang 0001, Jingjing Li 0001, Xiushan Nie, Xiaojun Chang
Neurocomputing5
2020 Personalized image quality assessment with Social-Sensed aesthetic preference
Chaoran Cui, Wenya Yang, Meng Wang 0001, Xiushan Nie, Yilong Yin
Inf. Sci.5
2020 Saliency detection using multiple low-level priors and a propagation mechanism
Muwei Jian, Junyu Dong, Chaoran Cui, Xiushan Nie, Yilong Yin
Multim. Tools Appl.5
2020 Modality correlation-based video summarization
Xingrun Wang, Xiushan Nie, Xingbo Liu, Binze Wang, Yilong Yin
Multim. Tools Appl.2
2020 Efficient weakly-supervised discrete hashing for large-scale social image retrieval
Hui Cui 0004, Lei Zhu 0002, Chaoran Cui, Xiushan Nie, Huaxiang Zhang 0001
Pattern Recognit. Lett.4
2020 Model Optimization Boosting Framework for Linear Model Hash Learning
abstract
Efficient hashing techniques have attracted extensive research interests in both storage and retrieval of highdimensional data, such as images and videos. In existing hashing methods, a linear model is commonly utilized owing to its efficiency. To obtain better accuracy, linear-based hashing methods focus on designing a generalized linear objective function with different constraints or penalty terms that consider the inherent characteristics and neighborhood information of samples. Differing from existing hashing methods, in this study, we propose a self-improvement framework called Model Boost (MoBoost) to improve model parameter optimization for linear-based hashing methods without adding new constraints or penalty terms. In the proposed MoBoost, for a linear-based hashing method, we first repeatedly execute the hashing method to obtain several hash codes to training samples. Then, utilizing two novel fusion strategies, these codes are fused into a single set. We also propose two new criteria to evaluate the goodness of hash bits during the fusion process. Based on the fused set of hash codes, we learn new parameters for the linear hash function that can significantly improve the accuracy. In general, the proposed MoBoost can be adopted by existing linear-based hashing methods, achieving more precise and stable performance compared to the original methods, and adopting the proposed MoBoost will incur negligible time and space costs. To evaluate the proposed MoBoost, we performed extensive experiments on four benchmark datasets, and the results demonstrate superior performance.
Xingbo Liu, Xiushan Nie, Liqiang Nie, Yilong Yin
IEEE Trans. Image Process.2
2020 Joint Multi-View Hashing for Large-Scale Near-Duplicate Video Retrieval
abstract
Multi-view hashing can well support large-scale near-duplicate video retrieval, due to its desirable advantages of mutual reinforcement of multiple features, low storage cost, and fast retrieval speed. However, there are still two limitations that impede its performance. First, existing methods only consider local structures in multiple features. They ignore the global structure that is important for near-duplicate video retrieval, and cannot fully exploit the dependence and complementarity of multiple features. Second, existing works always learn hashing functions bit by bit, which unfortunately increases the time complexity of hash function learning. In this paper, we propose a supervised hashing scheme, termed as joint multi-view hashing (JMVH), to address the aforementioned problems. It jointly preserves the global and local structures of multiple features while learning hashing functions efficiently. Specially, JMVH considers features of video as items, based on which an underlying Hamming space is learned by simultaneously preserving their local and global structures. In addition, a simple but efficient multi-bit hash function learning based on generalized eigenvalue decomposition is devised to learn multiple hash functions within a single step. It can significantly reduce the time complexity of conventional hash function learning processes that sequentially learn multiple hash functions bit by bit. The proposed JMVH is evaluated on two public databases: CC_WEB_VIDEO and UQ_VIDEO. Experimental results demonstrate that the proposed JMVH achieves more than a 5 percent improvement compared to several state-of-the-art methods which indicates the superior performance of JMVH.
Xiushan Nie, Weizhen Jing, Chaoran Cui, Chen Zhang 0013, Lei Zhu 0002, Yilong Yin
IEEE Trans. Knowl. Data Eng.1
2020 Robust Structured Graph Clustering
abstract
Graph-based clustering methods have achieved remarkable performance by partitioning the data samples into disjoint groups with the similarity graph that characterizes the sample relations. Nevertheless, their learning scheme still suffers from two important problems: 1) the similarity graph directly constructed from the raw features may be unreliable as real-world data always involves adverse noises, outliers, and irrelevant information and 2) most graph-based clustering methods adopt two-step learning strategy that separates the similarity graph construction and clustering into two independent processes. Under such circumstance, the generated graph is unstructured and fixed. It may suffer from a low-quality clustering structure and thus lead to suboptimal clustering performance. To alleviate these limitations, in this article we propose a robust structured graph clustering (RSGC) model. We formulate a novel learning framework to simultaneously learn a robust structured similarity graph and perform clustering. Specifically, the structured graph with proper probabilistic neighborhood assignment is adaptively learned on a robust latent representation that resists the noises and outliers. Furthermore, an explicit rank constraint is imposed on the Laplacian matrix to structurize the graph such that the number of the connected components is exactly equal to the ground-truth cluster number. To solve the challenging objective formulation, we propose to first transform it into an equivalent one that can be tackled more easily. An iterative solution based on the augmented Lagrangian multiplier is then derived to solve the model. In RSGC, the discrete cluster labels can be directly obtained by partitioning the learned similarity graph without reliance on label discretization strategy as most graph-based clustering methods. Experiments on both synthetic and real data sets demonstrate the superiority of the proposed method compared with the state-of-the-art clustering techniques.
Dan Shi 0003, Lei Zhu 0002, Jingjing Li 0001, Xiushan Nie
IEEE Trans. Neural Networks Learn. Syst.5
2020 Social-sensed Image Aesthetics Assessment
abstract
Image aesthetics assessment aims to endow computers with the ability to judge the aesthetic values of images, and its potential has been recognized in a variety of applications. Most previous studies perform aesthetics assessment purely based on image content. However, given the fact that aesthetic perceiving is a human cognitive activity, it is necessary to consider users’ perception of an image when judging its aesthetic quality. In this article, we regard users’ social behavior as the reflection of their perception of images and harness these additional clues to improve image aesthetics assessment. Specifically, we first merge the raw social interactions between users and images into clusters as the social labels of images, so the collective social behavioral information associated with an image can be well represented over a structured and compact space. Then, we develop a novel deep multi-task network to jointly learn social labels in different modalities from social images and apply it to common web images. In this manner, our approach is readily generalized to web images without social behavioral information. Finally, we introduce a high-level fusion sub-network to the aesthetics model, in which the social and visual representations of images are well balanced for aesthetics assessment. Experimental results on two benchmark datasets well verify the effectiveness of our approach and highlight the benefits of different types of social behavioral information for image aesthetics assessment.
Chaoran Cui, Peiguang Lin, Xiushan Nie, Muwei Jian, Yilong Yin
ACM Trans. Multim. Comput. Commun. Appl.3
2019 Jointly Multiple Hash Learning
abstract
Hashing can compress heterogeneous high-dimensional data into compact binary codes while preserving the similarity to facilitate efficient retrieval and storage, and thus hashing has recently received much attention from information retrieval researchers. Most of the existing hashing methods first predefine a fixed length (e.g., 32, 64, or 128 bit) for the hash codes before learning them with this fixed length. However, one sample can be represented by various hash codes with different lengths, and thus there must be some associations and relationships among these different hash codes because they represent the same sample. Therefore, harnessing these relationships will boost the performance of hashing methods. Inspired by this possibility, in this study, we propose a new model jointly multiple hash learning (JMH), which can learn hash codes with multiple lengths simultaneously. In the proposed JMH method, three types of information are used for hash learning, which come from hash codes with different lengths, the original features of the samples and label. In contrast to the existing hashing methods, JMH can learn hash codes with different lengths in one step. Users can select appropriate hash codes for their retrieval tasks according to the requirements in terms of accuracy and complexity. To the best of our knowledge, JMH is one of the first attempts to learn multi-length hash codes simultaneously. In addition, in the proposed model, discrete and closed-form solutions for variables can be obtained by cyclic coordinate descent, thereby making the proposed model much faster during training. Extensive experiments were performed based on three benchmark datasets and the results demonstrated the superior performance of the proposed method.
Xingbo Liu, Xiushan Nie, Yingxin Wang, Yilong Yin
AAAI2
2019 MoBoost: A Self-improvement Framework for Linear-based Hashing
abstract
The linear model is commonly utilized in hashing methods owing to its efficiency. To obtain better accuracy, linear-based hashing methods focus on designing a generalized linear objective function with different constraints or penalty terms that consider neighborhood information. In this study, we propose a novel generalized framework called Model Boost (MoBoost), which can achieve the self-improvement of the linear-based hashing. The proposed MoBoost is used to improve model parameter optimization for linear-based hashing methods without adding new constraints or penalty terms. In the proposed MoBoost, given a linear-based hashing method, we first execute the method several times to get several different hash codes for training samples, and then combine these different hash codes into one set utilizing one novel fusion strategy. Based on this set of hash codes, we learn some new parameters for the linear hash function that can significantly improve accuracy. The proposed MoBoost can be generally adopted in existing linear-based hashing methods, achieving more precise and stable performance compared to the original methods while imposing negligible added expenditure in terms of time and space. Extensive experiments are performed based on three benchmark datasets, and the results demonstrate the superior performance of the proposed framework.
Xingbo Liu, Xiushan Nie, Xiaoming Xi, Lei Zhu 0002, Yilong Yin
CIKM2
2019 Variable-Length Quantization Strategy for Hashing
abstract
Hashing is widely used to solve fast Approximate Nearest Neighbor (ANN) search problems, involves converting the original real-valued samples to binary-valued representations. The conventional quantization strategies, such as Single-Bit Quantization and Multi-Bit quantization, are considered ineffective, because of their serious information loss. To address this issue, we propose a novel variable-length quantization (VLQ) strategy for hashing. In the proposed VLQ technique, we divide all samples into different regions in each dimension firstly given the real-valued features of samples. Then we compute the dispersion degrees of these regions. Subsequently, we attempt to optimally assign different number of bits to each dimensions to obtain the minimum dispersion degree. Our experiments show that the VLQ strategy achieves not only superior performance over the state-of-the-art methods, but also has a faster retrieval speed on public datasets.
Xiushan Nie, Xiaoming Xi, Yilong Yin
ICIP2
2019 Supervised Short-Length Hashing
abstract
Hashing can compress high-dimensional data into compact binary codes, while preserving the similarity, to facilitate efficient retrieval and storage. However, when retrieving using an extremely short length hash code learned by the existing methods, the performance cannot be guaranteed because of severe information loss. To address this issue, in this study, we propose a novel supervised short-length hashing (SSLH). In this proposed SSLH, mutual reconstruction between the short-length hash codes and original features are performed to reduce semantic loss. Furthermore, to enhance the robustness and accuracy of the hash representation, a robust estimator term is added to fully utilize the label information. Extensive experiments conducted on four image benchmarks demonstrate the superior performance of the proposed SSLH with short-length hash codes. In addition, the proposed SSLH outperforms the existing methods, with long-length hash codes. To the best of our knowledge, this is the first linear-based hashing method that focuses on both short and long-length hash codes for maintaining high precision.
Xingbo Liu, Xiushan Nie, Xiaoming Xi, Lei Zhu 0002, Yilong Yin
IJCAI2
2019 Supervised Discrete Hashing With Mutual Linear Regression
abstract
Supervised linear hashing can compress high-dimensional data into compact binary codes owing to its efficiency. Generally, the relation between label and hash codes is widely used in the existing hashing methods because of its effectiveness of improving the accuracy. The existing hashing methods always use two different projections to represent the mutual regression between hash codes and class labels. In contrast to the existing methods, we propose a novel learning-based hashing method termed supervised discrete hashing with mutual linear regression (SDHMLR) in this study, where only one stable projection is used to describe the linear correlation between hash codes and corresponding labels. To the best of our knowledge, this strategy has not been used for hashing previously. In addition, we further use a boosting strategy to improve the final performance of the proposed method without adding extra constraints and with little extra expenditure in terms of time and space. Extensive experiments conducted on three image benchmarks demonstrate the superior performance of the proposed method.
Xingbo Liu, Xiushan Nie, Yilong Yin
ACM Multimedia2
2019 Flexible Online Multi-modal Hashing for Large-scale Multimedia Retrieval
abstract
Multi-modal hashing fuses multi-modal features at both offline training and online query stage for compact binary hash learning. It has aroused extensive attention in research filed of efficient large-scale multimedia retrieval. However, existing methods adopt batch-based learning scheme or unsupervised learning paradigm. They cannot efficiently handle the very common online streaming multi-modal data (for batch-learning methods), or learn the hash codes suffering from limited discriminative capability and less flexibility for varied streaming data (for existing online multi-modal hashing methods). In this paper, we develop a supervised Flexible Online Multi-modal Hashing (FOMH) method to adaptively fuse heterogeneous modalities and flexibly learn the discriminative hash code for the newly coming data, even if part of the modalities is missing. Specifically, instead of adopting the fixed weights, the modalities weights in FOMH are automatically learned with the proposed flexible multi-modal binary projection to timely capture the variations of streaming samples. Further, we design an efficient asymmetric online supervised hashing strategy to enhance the discriminative capability of the hash codes, while avoiding the challenging symmetric semantic matrix decomposition and storage cost. Moreover, to support fast hash updating and avoid the propagation of binary quantization errors in online learning process, we propose to directly update the hash codes with an efficient discrete online optimization. Experiments on several public multimedia retrieval datasets validate the superiority of the proposed method from various aspects.
Xu Lu 0004, Lei Zhu 0002, Zhiyong Cheng 0001, Jingjing Li 0001, Xiushan Nie, Huaxiang Zhang 0001
ACM Multimedia5
2019 Pre-course student performance prediction with multi-instance multi-label learning
Yuling Ma, Chaoran Cui, Xiushan Nie, Gongping Yang 0001, Kashif Shaheed, Yilong Yin
Sci. China Inf. Sci.3
2019 Multi-view face hallucination using SVD and a mapping model
Muwei Jian, Chaoran Cui, Xiushan Nie, Huaxiang Zhang 0001, Liqiang Nie, Yilong Yin
Inf. Sci.3
2019 Automated segmentation of choroidal neovascularization in optical coherence tomography images using multi-scale convolutional neural networks with structure prior
Xiaoming Xi, Xianjing Meng, Lu Yang 0005, Xiushan Nie, Gongping Yang 0001, Haoyu Chen 0002, Yilong Yin, Xinjian Chen 0001
Multim. Syst.4
2019 Binary feature representation learning for scene retrieval in micro-video
Jie Guo 0012, Xiushan Nie, Muwei Jian, Yilong Yin
Multim. Tools Appl.2
2019 Assessment of feature fusion strategies in visual attention mechanism for saliency detection
Muwei Jian, Chaoran Cui, Xiushan Nie, Hanjiang Luo, Yilong Yin
Pattern Recognit. Lett.4
2019 Global-view hashing: harnessing global relations in near-duplicate video retrieval
Weizhen Jing, Xiushan Nie, Chaoran Cui, Xiaoming Xi, Gongping Yang 0001, Yilong Yin
World Wide Web2
2018 Modality-Specific Structure Preserving Hashing for Cross-Modal Retrieval
abstract
Hashing-based methods have made great advancements in cross-modal retrieval in both computational efficiency and storage. Learning a common space from different modalities is the common strategy of hashing-based methods, however, relational and structural information between samples in each modality, namely, a modality-specific structure, is always discarded during learning. In addition, cross-modality samples sometimes suffer from inter-class ambiguity and intra-class variability because of the uncertainty of manual labeling. To address these issues, we propose a novel method named Modality-specific structure Preserving Hashing (MsPH), which learns hashes by preserving the local structure and relations between samples in each modality. Moreover, label enhancement is utilized in MsPH to address label ambiguity and variability. Extensive experiments conducted on three benchmark datasets demonstrate the superiority of MsPH under various cross-modal scenarios.
Xingbo Liu, Haoliang Sun, Xiushan Nie, Chaoran Cui, Yilong Yin
ICASSP3
2018 Structural Compact Core Tensor Dictionary Learning for Multispec-Tral Remote Sensing Image Deblurring
abstract
The multispectral remote sensing image (MS-RSI) is blurred existing multispectral camera due to various hardware limitations. In this paper, we propose a novel structural compact core tensor dictionary learning (SCCTDL) model for MS-RSI deblurring. First, the multispectral patch is modeled by three-order tensor and high-order singular value decomposition is applied to the tensor. Then the task of MS-RSI deblurring is formulated as a minimum sparse core tensor estimation problem. To improve the accuracy of core tensor coding, the core tensor estimation based on the structural compact principle is introduced into the SCCTDL model to exploit abundant structural similarity in image. Experimental results suggest that our method outperforms several existing MS-RSI deblurring methods in both subjective image quality and visual perception.
Leilei Geng, Xiushan Nie, Sijie Niu, Yilong Yin
ICIP2
2018 Fast Discrete Cross-modal Hashing With Regressing From Semantic Labels
abstract
Hashing has recently received great attention in cross-modal retrieval. Cross-modal retrieval aims at retrieving information across heterogeneous modalities (e.g., texts vs. images). Cross-modal hashing compresses heterogeneous high-dimensional data into compact binary codes with similarity preserving, which provides efficiency and facility in both retrieval and storage. In this study, we propose a novel fast discrete cross-modal hashing (FDCH) method with regressing from semantic labels to take advantage of supervised labels to improve retrieval performance. In contrast to existing methods that learn the projection from hash codes to semantic labels, the proposed FDCH regresses the semantic labels of training examples to the corresponding hash codes with a drift. It not only accelerates the hash learning process, but also helps generate stable hash codes. Furthermore, the drift can adjust the regression and enhance the discriminative capability of hash codes. Especially in the case of training efficiency, FDCH is much faster than existing methods. Comparisons with several state-of-the-art techniques on three benchmark datasets have demonstrated the superiority of FDCH under various cross-modal retrieval scenarios.
Xingbo Liu, Xiushan Nie, Wenjun Zeng 0001, Chaoran Cui, Lei Zhu 0002, Yilong Yin
ACM Multimedia2
2018 Image Aesthetic Distribution Prediction with Fully Convolutional Network
Huidi Fang, Chaoran Cui, Xiang Deng 0002, Xiushan Nie, Muwei Jian, Yilong Yin
MMM (1)4
2018 Learned local similarity prior embedding active contour model for choroidal neovascularization segmentation in optical coherence tomography images
Xiaoming Xi, Xianjing Meng, Lu Yang 0005, Xiushan Nie, Zhilou Yu, Chunyun Zhang, Haoyu Chen 0002, Yilong Yin, Xinjian Chen 0001
Sci. China Inf. Sci.4
2018 Cross-modal hashing based on category structure preserving
Xiushan Nie, Xingbo Liu, Leilei Geng
J. Vis. Commun. Image Represent.2
2018 Saliency detection based on directional patches extraction and principal local color contrast
Muwei Jian, Wenyin Zhang, Hui Yu 0001, Chaoran Cui, Xiushan Nie, Huaxiang Zhang 0001, Yilong Yin
J. Vis. Commun. Image Represent.5
2018 Robust Image Fingerprinting Based on Feature Point Relationship Mining
abstract
Local feature points have been widely employed in robust image fingerprinting. One of their intrinsic advantages is their invariance under geometric transforms. However, their robustness against certain attacks that modify the positions of points, such as additive noising and blurring, is limited. In addition, local-feature-point-based approaches ignore the distribution of the feature points. In this paper, we harness feature point relationships, including local structures and global relevance, to overcome these limitations. In the relationship mining strategy, Delaunay triangulation is first applied to the feature points to capture their geometric structures. Subsequently, local structures are represented by searching for an independent set in the mapping graph constructed via Delaunay triangulation, whereas the global relevance is represented by the Laplacian of the graph. Finally, the local structures and global relevance are used as input to the quantization process of the image fingerprinting system. In the process of quantization, we propose an unsupervised quantization strategy called between-cluster distance-based quantization to preserve the neighborhood structure between the binary fingerprint space and the original feature space. Experimental results show that the proposed method achieves effective performance under common modifications.
Xiushan Nie, Yane Chai, Chaoran Cui, Xiaoming Xi, Yilong Yin
IEEE Trans. Inf. Forensics Secur.1
2017 Personalized Image Aesthetics Assessment
abstract
Automatically assessing image quality from an aesthetic perspective is of great interest to the high-level vision research community. Existing methods are typically non-personalized and quantify image aesthetics with a universal label. However, given the fact that aesthetics is a subjective perception, how to understand user aesthetic perceptions poses a formidable challenge to image aesthetics assessment. In this paper, we propose to model user aesthetic perceptions using a set of exemplar images from social media platforms, and realize personalized aesthetics assessment by transferring this knowledge to adapt the results of the trained generic model. In this way, image aesthetics is measured from both aspects of visual quality and user tastes. Extensive experiments on two benchmark datasets well verified the potential of our approach for personalized image aesthetics assessment.
Xiang Deng 0002, Chaoran Cui, Huidi Fang, Xiushan Nie, Yilong Yin
CIKM4
2017 Distribution-oriented Aesthetics Assessment for Image Search
abstract
Aesthetics has become increasingly prominent for image search to enhance user satisfaction. Therefore, image aesthetics assessment is emerging as a promising research topic in recent years. In this paper, distinguished from existing studies relying on a single label, we propose to quantify the image aesthetics by a distribution over quality levels. The distribution representation can effectively characterize the disagreement among the aesthetic perceptions of users regarding the same image. Our framework is developed on the foundation of label distribution learning, in which the reliability of training examples and the correlations between quality levels are fully taken into account. Extensive experiments on two benchmark datasets well verified the potential of our approach for aesthetics assessment. The role of aesthetics in image search was also rigorously investigated.
Chaoran Cui, Huidi Fang, Xiang Deng 0002, Xiushan Nie, Hongshuai Dai, Yilong Yin
SIGIR4
2017 Hybrid textual-visual relevance learning for content-based image retrieval
Chaoran Cui, Peiguang Lin, Xiushan Nie, Yilong Yin, Qingfeng Zhu
J. Vis. Commun. Image Represent.3
2017 Comprehensive Feature-Based Robust Video Fingerprinting Using Tensor Model
abstract
Content-based near-duplicate video detection (NDVD) is essential for effective search and retrieval, and robust video fingerprinting is a good solution for NDVD. Most existing video fingerprinting methods use a single feature or concatenate different features to generate video fingerprints, and show good performance under single-mode modifications such as noise addition and blurring. However, when they suffer combined modifications, the performance is degraded to a certain extent because such features cannot characterize the video content completely. By contrast, the assistance and consensus among different features can improve the performance of video fingerprinting. Therefore, in the present study, we mine the assistance and consensus among different features based on a tensor model, and we present a new comprehensive feature to fully use them in the proposed video fingerprinting framework. We also analyze what the comprehensive feature really is for representing the original video. In this framework, the video is initially set as a high-order tensor that consists of different features, and the video tensor is decomposed via the Tucker model with a solution that determines the number of components. Subsequently, the comprehensive feature is generated by the low-order tensor obtained from tensor decomposition. Finally, the video fingerprint is computed using this feature. A matching strategy used for narrowing the search is also proposed based on the core tensor. The robust video fingerprinting framework is resistant not only to single-mode modifications but also to their combination.
Xiushan Nie, Yilong Yin, Jiande Sun 0001, Chaoran Cui
IEEE Trans. Multim.1
2016 Spherical torus-based video hashing for near-duplicate video detection
Xiushan Nie, Yane Chai, Jiande Sun 0001, Yilong Yin
Sci. China Inf. Sci.1
2015 Graph-based video fingerprinting using double optimal projection
Xiushan Nie, Wenjun Zeng 0001
J. Vis. Commun. Image Represent.1
2014 Structural similarity-based video fingerprinting for video copy detection
abstract
The authors propose a video fingerprinting method based on structural similarity and a graph model. Structural similarity‐based fingerprint generation and double‐layer matching are the two main contributions. The video is mapped to a graph with frames as its vertices, and structural similarity is proposed to compute the weights of the edges. Then, the fingerprint consisting of a match tag (a coarse fingerprint) and a fine fingerprint is generated by this graph. The match tag is generated by an independent set of this graph, and the fine fingerprint is generated by the weight matrix of the graph based on the two‐block‐dimensional discrete cosine transform. During the matching, the video can be matched at the first‐layer using the match tag to obtain a candidate set, whereas the second‐layer matching is performed in this candidate set using the fine fingerprint to find a final match. The proposed video fingerprinting method is shown to be resistant to geometric attacks on frames and impairment of transmission channels.
Xiushan Nie, Wenjun Zeng 0001, Jiande Sun 0001
IET Image Process.1
2014 Video fingerprinting based on graph model
Xiushan Nie, Jiande Sun 0001, Zhihui Xing, Xiaocui Liu
Multim. Tools Appl.1
2013 Logarithmic Spread-Transform Dither Modulation watermarking Based on Perceptual Model
abstract
Logarithmic Quantization Index Modulation (LQIM) is an important extension of the original quantization-based watermarking method. However, it is well known that it is sensitive to valumetric scaling attack and easy to result in sign error after quantization and attacks. For that, in this paper, we propose a new method, namely Logarithmic Spread-Transform Dither Modulation Based on Perceptual Model (LSTDM-WM). In this regard the host signal is first projected onto a random vector and transformed using a novel Logarithmic Quantization function. Then the transformed signal is quantized regarding the watermark data and the watermarked signal is obtained by applying inverse transform to the quantized signal. The perceptual model is further exploited to adjust the quantization step adaptively for watermark embedding. Experimental results indicate that our proposed scheme overcomes two challenges cited above and has superior performance in comparison with conventional LQIM and former proposed schemes of STDM.
Wenbo Wan, Jiande Sun 0001, Xiushan Nie
ICIP5
2013 Robust video hashing based on representative-dispersive frames
Xiushan Nie, Jiande Sun 0001, Lianqi Wang
Sci. China Inf. Sci.1
2012 A visual saliency based video hashing algorithm
abstract
A novel video hashing algorithm is proposed, which takes account of visual saliency during hash generation. In the proposed algorithm, the video hash is fused by two hashes, which are spatio-temporal hash (ST-Hash) and visual hash (V-Hash). The ST-Hash is generated based on the ordinal feature, which is formed according to the intensity difference between adjacent blocks of the temporally informative representation image (TIRI). At the same time, the representative saliency map (RSM) is constructed by the visual saliency maps in video segments. The V-Hash is formed according to the intensity difference between adjacent blocks of the RSM, and used to modulate the ST-Hash to form the final video hash. Experiments on different kinds of videos with different kinds of attacks verify that the proposed algorithm has better performance on robustness and discrimination.
Jiande Sun 0001, Xiushan Nie
ICIP4
2012 Video Hashing Algorithm With Weighted Matching Based on Visual Saliency
abstract
In this letter, a novel video hashing algorithm is proposed, in which the weighted hash matching is defined in video hashing for the first time. In the proposed algorithm, the video hash is generated based on the ordinal feature derived from the temporally informative representation image (TIRI). At the same time the representative saliency map (RSM) is constructed by the visual saliency maps in video segments, and it generates the hash weights for hash matching. During hash matching, the traditional bit error rate (BER) is weighted with hash weights to form the weighted error rate (WER). WER is used to measure the similarity between different hashes. Experiments on different kinds of videos with different kinds of attacks verify the robustness and discrimination of the proposed algorithm.
Jiande Sun 0001, Xiushan Nie
IEEE Signal Process. Lett.4
2011 Robust Video Hashing Based on Double-Layer Embedding
abstract
A robust video hashing scheme for video content identification and authentication is proposed, which is called Double-Layer Embedding scheme. Intra-cluster Locally Linear Embedding (LLE) and inter-cluster Multi-Dimensional Scaling (MDS) are used in the scheme. Some dispersive frames of the video are first selected through graph model, and the video is partitioned into clusters based on the dispersive frames and the K-Nearest Neighbor method during the hashing. Then, the intra-cluster LLE and inter-cluster MDS are used to generate local and global hash sequences which can inherently describe the corresponding video. Experimental results show that the video hashing is resistant to geometric attacks of frames and channel impairments of transmission.
Xiushan Nie, Jiande Sun 0001, Wei Liu 0031
IEEE Signal Process. Lett.1
2010 Robust video hashing for identification based on MDS
abstract
Video identification is extremely important in video browsing, database search and security. In this paper, we present a video hashing based on MDS (Multi-Dimensional Scaling) which is able to work under variable video transmission impairments and resistant to signal processing. In this method, each frame of the video is divided into blocks and compute its low and middle frequency DCT coefficients of luminance component as a disparities measurement for MDS. Then the video is mapped to two-dimensional space using MDS, and generate a robust hashing as a video signature utilizing the distances between two points mapping from frames. It found that this video hashing is resistant frame geometric attacks (rotation, shift), random noises, lossy compression and other video transmission impairments. It can be instrumental in building database search, video copy detection and watermarking applications for video.
Xiushan Nie, Jiande Sun 0001
ICASSP1
2009 Independent triangles set-based watermarking for 3D models
abstract
An independent triangles set-based watermarking algorithm using singular value decomposition (SVD) is presented in this paper. In this algorithm, a 3D model is taken as a model that consists of triangular meshes, and we select the triangles called independent triangles set (ITS) to generate a Hankel matrix and apply SVD on it, then watermarks are embedded by modulating singular values. The watermarks that embedded by this method are resistant to similarity transformations and random noises. We prove that the proposed algorithm using ITS as embedding primitive is more robust than those using vertex coordinates discussed in other papers, and the experimental results corroborate it as well. In addition, when detect the watermarks, we only use some keys instead of the original model, so it is a blind watermarking algorithm too.
Xiushan Nie, Jiande Sun 0001, Jianping Qiao
ICME1