Xiang Zhang 0008

dblp:91/4353-8 · DBLP profile ↗
← Back
94ranked-venue papers
7as first author
49since 2021 · last 2026
0000-0002-5201-3802ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 4 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 3 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author
YearPublicationVenuePosition
2026 HAGC: A Hardware-Aware Gradient Compression framework for distributed deep learning
Aiqiang Yang, Jie Liu 0002, Bo Yang 0021, Xiang Zhang 0008, Zeyao Mo, Keqin Li 0001
J. Syst. Archit.4
2026 GNNRL-smoothing: A prior-free reinforcement learning model for mesh optimization
Xinhai Chen 0001, Chunye Gong, Bo Yang 0023, Liang Deng, Yufei Pang, Xiang Zhang 0008, Jie Liu 0002
Neural Networks8
2026 Phrase Grounding-Based Style Transfer for Single-Domain Generalized Object Detection
abstract
Single-domain generalized object detection aims to enhance a model’s generalization to multiple unseen target domains using only data from a single source domain during training. This is a practical yet challenging scenario, as it requires the model to address domain shift without incorporating target domain data into the training process. In this paper, we propose a novel phrase-grounding-based style transfer (PGST) approach for the task. Specifically, we first define textual prompts to describe objects for potential unseen target domains. Then, we leverage the grounded language-image pre-training (GLIP) model to capture the styles of these target domains and perform style transfer from the source to the target domains. The style-transferred visual features from the source domain are semantically rich and closely approximate those of their hypothetical counterparts in the target domain. Finally, we employ these style-transferred visual features to fine-tune GLIP. By introducing these imaginary counterparts, the detector can be effectively generalized to unseen target domains using only a single source domain during training. Our method significantly improves mean average precision (mAP), with an average increase of 8.8% across five diverse weather-driving benchmarks. Notably, our approach outperforms or matches the performance of domain-adaptive object detection methods, which require target domain data for training, in several challenging scenarios.
Wei Wang 0335, Cong Wang 0018, Mengzhu Wang, Xiang Zhang 0008, Long Lan, Xinwang Liu 0002, Kenli Li 0001, Xiaochun Cao
IEEE Trans. Circuits Syst. Video Technol.5
2026 Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and Segmentation
abstract
Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a language expression in a video. This intricate task involves the reasoning of linguistic and visual modalities, along with the temporal association of target objects. However, the seminal work relies on loose feature fusion and neglects long-term information. In this study, we introduce a compact Transformer-based method, termed TenRMOT. We conduct feature fusion at both encoding and decoding stages to fully exploit the advantages of Transformer architecture. Specifically, we incrementally perform cross-modal fusion layer-by-layer during the encoding phase. In the decoding phase, we utilize language-guided queries to probe memory features for accurate prediction of the desired objects. Moreover, we introduce a query update module that explicitly leverages temporal prior information of the tracked objects to enhance the consistency of their trajectories. In addition, we introduce a novel task called Referring Multi-Object Tracking and Segmentation (RMOTS) and construct a new dataset named Ref-KITTI Segmentation. Our dataset consists of 18 videos with 818 expressions, and each expression averages 10.7 masks, which poses a greater challenge compared to the typical single mask in most existing referring video segmentation datasets. TenRMOT demonstrates superior performance on both the referring multi-object tracking and the segmentation tasks.
Changcheng Xiao, Qiong Cao, Xiang Zhang 0008, Tao Wang 0006, Canqun Yang, Long Lan
IEEE Trans. Circuits Syst. Video Technol.4
2026 Exploring Direction Alignment and Discrepancy Standardization for Knowledge Distillation
abstract
Knowledge Distillation (KD) is a widely popular model compression technique that can effectively transfer knowledge from a pre-trained, large-scale teacher model to a more compact and lightweight student model. Traditional KD methods aim to improve the student’s representation capability by mimicking the teacher’s features, e.g., minimizing the \(\mathcal{L}_{2}\) distance between their intermediate features. However, due to the capacity gap between the student and the teacher, student often struggles to precisely mimic the features of the teacher. To address this challenge, we propose to boost the knowledge distillation for the visual recognition tasks via Direction Alignment and Discrepancy Standardization ( DADS) , which exploits the feature scaling technique to distill from both the feature direction and feature discrepancy. To this end, we devise an efficient feature alignment module to align the dimensions of teacher and student features. Moreover, we align the direction of student features and teacher features, which are pre-processed by normalization. Furthermore, we leverage the Kullback–Leibler (KL) divergence to refine the features alignment, minimizing discrepancy in the distribution of features across samples, which is pre-processed by \(\mathcal{Z}\) -score standardization. In this way, our proposed approach can effectively transfer the knowledge from the teacher to the student, facilitating the downstream visual recognition applications, such as image classification and semantic segmentation. Extensive experimental analyses clearly validate the effectiveness of DADS . Compared with previous KD methods, our approach sets a new benchmark, achieving state-of-the-art results on visual recognition tasks.
Dingyao Chen, Xiao Teng, Xiang Zhang 0008, Xun Yang 0001, Long Lan
ACM Trans. Knowl. Discov. Data3
2025 Informative Discrimination Network for Efficient Single Image Super-Resolution
abstract
Deploying convolutional neural networks on low-resource mobile devices for single image super-resolution (SISR) faces the issue of how to balance the parameter amount and performance. The default solution is simultaneously condensing both hierarchical representation and attention features into their respective light proxies. The insight underlying this solution lies in the fact that features are redundant since the super-resolution needs plenty of similar pixels. This work takes it to the next step from the viewpoint of informativeness and discrimination. In detail, we propose an informative disrcimination network (IDNet) for SISR. For informativeness, a multi-scale residual block (MRB) is explored to capture informative spatial details via the scale-in-scale structure. It mines rich intra-layer spatial details based on inter-layer ones of the default hierarchical representation. However, it also incurs feature redundancy. Though attention serves to reduce this redundancy, feature discrimination and pixel-wise structural preservation cannot be guaranteed. Here spatial discrimination attention behaves like the biased discriminant classifier to induce spatial discrimination, while the nuclear-norm regularization recovers the image low-rank structure to reduce artifacts or noises. Importantly, no extra network weights are introduced for model efficiency. Experiments show that IDNet delivers sound performance with fewer parameters, as compared to its cousins.
Yuzheng Tu, Xinhai Chen 0001, Chunye Gong, Jie Liu 0002, Bo Yang 0023, Xiang Gao 0020, Xiang Zhang 0008
IJCNN7
2025 UGM2N: An Unsupervised and Generalizable Mesh Movement Network via M-Uniform Loss
abstract
Partial differential equations (PDEs) form the mathematical foundation for modeling physical systems in science and engineering, where numerical solutions demand rigorous accuracy-efficiency tradeoffs. Mesh movement techniques address this challenge by dynamically relocating mesh nodes to rapidly-varying regions, enhancing both simulation accuracy and computational efficiency. However, traditional approaches suffer from high computational complexity and geometric inflexibility, limiting their applicability, and existing supervised learning-based approaches face challenges in zero-shot generalization across diverse PDEs and mesh topologies. In this paper, we present an $\textbf{U}$nsupervised and $\textbf{G}$eneralizable $\textbf{M}$esh $\textbf{M}$ovement $\textbf{N}$etwork (UGM2N). We first introduce unsupervised mesh adaptation through localized geometric feature learning, eliminating the dependency on pre-adapted meshes. We then develop a physics-constrained loss function, M-Uniform loss, that enforces mesh equidistribution at the nodal level. Experimental results demonstrate that the proposed network exhibits equation-agnostic generalization and geometric independence in efficient mesh adaptation. It demonstrates consistent superiority over existing methods, including robust performance across diverse PDEs and mesh geometries, scalability to multi-scale resolutions and guaranteed error reduction without mesh tangling.
Xinhai Chen 0001, Xiang Gao 0020, Qingyang Zhang 0009, Menghan Jia, Xiang Zhang 0008, Jie Liu 0002
NeurIPS7
2025 Fine-grained vectorized merge sorting on RISC-V: from register to cache
abstract
Abstract Merge sort as a divide-sort-merge paradigm has been widely applied in computer science fields. As modern reduced instruction set computing architectures like the fifth generation (RISC-V) regard multiple registers as a vector register group for wide instruction parallelism, optimizing merge sort with this vectorized property is becoming increasingly common. In this paper, we overhaul the divide-sort-merge paradigm, from its register-level sort to the cache-aware merge, to develop a fine-grained RISC-V vectorized merge sort (RVMS). From the register-level view, the inline vectorized transpose instruction is missed in RISC-V, so implementing it efficiently is non-trivial. Besides, the vectorized comparisons do not always work well in the merging networks. Both issues primarily stem from the expensive data shuffle instruction. To bypass it, RVMS strides to take register data as the proxy of data shuffle to accelerate the transpose operation, and meanwhile replaces vectorized comparisons with scalar cousin for more light real value swap. On the other hand, as cache-aware merge makes larger data merge in the cache, most merge schemes have two drawbacks: the in-cache merge usually has low cache utilization, while the out-of-cache merging network remains an ineffectively symmetric structure. To this end, we propose the half-merge scheme to employ the auxiliary space of in-place merge to halve the footprint of naïve merge sort, and meanwhile copy one sequence to this space to avoid the former data exchange. Furthermore, an asymmetric merging network is developed to adapt to two different input sizes. Experiments on the RISC-V processor SG2042 show that four fine-grained optimization schemes including register strided transpose, hybrid merging network, half-merge strategy, and asymmetric merging network, improve performance by 4.05%, 19.88%, 12.23%, and 11.04% respectively. Importantly, the overall performance is 1.34x faster than the parallel sorting in the Boost C++ library, and 1.85x faster than std::sort.
Jin Zhang 0018, Jincheng Zhou, Xiang Zhang 0008, Chunye Gong
CCF Trans. High Perform. Comput.3
2025 HCL: A Hierarchical Contrastive Learning Framework for Zero-Shot Relation Extraction
abstract
Zero-shot relation extraction (ZSRE) is shown to become more significant in the current information extraction system, which aims at predicting relation classes that lack annotations or have just never appeared during training. Previous works focus on projecting sentences with their corresponding relation descriptions to an intermediate semantic space and searching the nearest semantic for predicting unseen classes. Though these methods can achieve sound performance, they only obtain inferior semantic information via a trivial distance metric and neglect the interaction in the instance representations. We are thus motivated to tackle these issues and propose a hierarchical contrastive learning (HCL) framework for ZSRE including projection-level and instance-level modules. Specifically, the projection-level component replaces the distance score function by contrastive loss to connect the input sentence with the relation semantic space. And the instance-level component integrates the external knowledge from sentence entities to establish new contrastive pairs for efficiently learning representations from mutual information. The experimental results on three well-known datasets demonstrate that our model surpasses the existing SOTA by at most 18.97% improvement on the F1 score when unseen classes are 15. Moreover, our model can achieve more competitive performance alone with the increasing number of unseen classes.
Tianwei Yan 0001, Shan Zhao 0002, Minghao Hu 0001, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Foreground Enhanced Network for Weakly Supervised Temporal Language Grounding
Hongzhou Wu, Xuechen Zhao, Xiang Zhang 0008
CogSci4
2024 A Hybrid Vectorized Merge Sort on ARM NEON
Jincheng Zhou, Jin Zhang 0018, Xiang Zhang 0008, Tiaojie Xiao, Chunye Gong
ICA3PP (6)3
2024 CAMLB-SpMV: An Efficient Cache-Aware Memory Load-Balancing SpMV on CPU
abstract
Sparse Matrix-Vector Multiplication (SpMV) plays a crucial role in scientific computing, but severe load imbalance among threads restricts its performance. Previous load-balancing methods have primarily ignored the CPU’s cache line-based memory access characteristics and the impact of data locality during workload evaluation and partitioning, leading to limited effect in load balancing.
Jihu Guo, Jie Liu 0002, Xiaoxiong Zhu, Xiang Zhang 0008
ICPP5
2024 Discriminative object tracking by domain contrast
Huayue Cai, Xiang Zhang 0008, Long Lan, Changcheng Xiao, Chuanfu Xu, Jie Liu 0002, Zhigang Luo
Comput. Vis. Image Underst.2
2024 IoUformer: Pseudo-IoU prediction with transformer for visual tracking
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Yibing Zhan, Zhigang Luo
Neural Networks4
2024 MotionTrack: Learning motion predictor for multiple object tracking
Changcheng Xiao, Qiong Cao, Long Lan, Xiang Zhang 0008, Zhigang Luo, Dacheng Tao
Neural Networks5
2024 SiamATTRPN: Enhance Visual Tracking With Channel and Spatial Attention
abstract
Visual tracking is an important research topic in the field of computer vision. The current Siamese tracker based on the region proposal network (SiamRPN) has achieved promising tracking results in terms of efficiency and performance. However, through our empirical study, we have observed that deep features learned by SiamRPN are of substandard quality, as the salient regions within the deep features fail to correspond accurately with meaningful objects. To address this limitation, we propose an approach to enhance the quality of the learned deep features through the incorporation of an attention mechanism. Attention mechanisms have been shown to be effective in distinguishing similar objects, as they suppress background objects while highlighting target information that is most relevant. As a result, a new tracking method with channel and spatial attention termed SiamATTRPN is explored. To verify the effectiveness of SiamATTRPN, experiments on benchmark datasets demonstrate that our proposed tracker outperforms the baseline tracker significantly.
Huayue Cai, Xiang Zhang 0008, Long Lan, Wenxin Shen, Junyang Chen 0001, Victor C. M. Leung
IEEE Trans. Comput. Soc. Syst.2
2024 Dual-View Learning Based on Images and Sequences for Molecular Property Prediction
abstract
The prediction of molecular properties remains a challenging task in the field of drug design and development. Recently, there has been a growing interest in the analysis of biological images. Molecular images, as a novel representation, have proven to be competitive, yet they lack explicit information and detailed semantic richness. Conversely, semantic information in SMILES sequences is explicit but lacks spatial structural details. Therefore, in this study, we focus on and explore the relationship between these two types of representations, proposing a novel multimodal architecture named ISMol. ISMol relies on a cross-attention mechanism to extract information representations of molecules from both images and SMILES strings, thereby predicting molecular properties. Evaluation results on 14 small molecule ADMET datasets indicate that ISMol outperforms machine learning (ML) and deep learning (DL) models based on single-modal representations. In addition, we analyze our method through a large number of experiments to test the superiority, interpretability and generalizability of the method. In summary, ISMol offers a powerful deep learning toolbox for drug discovery in a variety of molecular properties.
Xiang Zhang 0008, Hongxin Xiang, Xixi Yang, Jingxin Dong 0002, Xiangzheng Fu, Xiangxiang Zeng, Keqin Li 0001
IEEE J. Biomed. Health Informatics1
2023 Atomic-action-based Contrastive Network for Weakly Supervised Temporal Language Grounding
abstract
As one knows, an event often consists of several actions while each action is atomic. Inspired by this insight, we propose a novel framework named Atomic-action-based Contrastive Network model (ACN) for weakly supervised temporal language grounding task to localize the query-related event moment in an untrimmed video, without access to any temporal annotations. Specifically, ACN first determines the accurate moment boundary of each action in a query-agnostic way. This can adequately exploit homogeneous visual cues while impeding the heterogeneity of the query from hurting the atomicity of visual action, i.e., action boundary. To effectively localize the query-related event, we seek the discriminative words in the given query, and explore a composite-grained contrastive module to retrieve those corresponding atomic actions in the common latent space across modalities. This boosts feature discrimination of visual event segment to remove irrelevant action video segments. Experiments on two popular datasets show the efficacy of our model.
Hongzhou Wu, Xuechen Zhao, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo
ICME6
2023 Semantics-Enriched Cross-Modal Alignment for Complex-Query Video Moment Retrieval
abstract
Video moment retrieval (VMR) aims to search for a video segment that matches the search intent in a query sentence, which has received increasing attention in recent years, due to its practical values in various fields. Existing efforts devoted to this interesting yet challenging task typically encode the query sentence and video segments into unstructured global representations for cross-modal interaction and fusion, which may fail to accurately capture the search intent in complex queries with multi-granularity semantics.
Xiang Zhang 0008, Xun Yang 0001, Yibing Zhan, Long Lan, Jianfeng Dong, Hongzhou Wu
ACM Multimedia2
2023 Domain-specific feature recalibration and alignment for multi-source unsupervised domain adaptation
abstract
Abstract Traditional unsupervised domain adaptation (UDA) usually assumes that the source domain has labels and the target domain has no labels. In a real environment, labelled source domain data usually comes from multiple different distributions. To handle this problem, multi‐source unsupervised domain adaptation (MUDA) is proposed. Multi‐source unsupervised domain adaptation aims to adapt the model trained on multi‐labelled source domains to the unlabelled target domain. In this paper, a novel MUDA method by domain‐specific feature recalibration and alignment (FRA) is proposed. Specifically, to achieve feature recalibration, the authors leverage channel attention to pick out significant channels and spatial attention to focus on important features in different channels. Such integration of channel and spatial attention can lead to effective domain‐specific feature recalibration that may be of great importance to MUDA. In addition, to achieve better MUDA, the authors propose domain‐specific feature alignment which consists of Maximum Mean Discrepancy and JS‐divergence loss. Maximum Mean Discrepancy can reduce the difference between the source domain and target domain. Meanwhile, JS‐divergence loss may ensure the prediction consistency of different classifiers in the source domains. Four experiments have proved that FRA can achieve significantly better results in popular benchmarks for MUDA.
Mengzhu Wang, Dingyao Chen, Fangzhou Tan, Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Zhigang Luo
IET Comput. Vis.6
2023 Online intervention siamese tracking
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Changcheng Xiao, Zhigang Luo
Inf. Sci.4
2023 SiamDF: Tracking training data-free siamese tracker
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Zhigang Luo
Neural Networks4
2023 Class-specific and self-learning local manifold structure for domain adaptation
Wei Wang 0335, Mengzhu Wang, Long Lan, Quannan Zu, Xiang Zhang 0008, Cong Wang 0018
Pattern Recognit.6
2023 Reducing bi-level feature redundancy for unsupervised domain adaptation
Mengzhu Wang, Shanshan Wang 0008, Wei Wang 0335, Li Shen 0008, Xiang Zhang 0008, Long Lan, Zhigang Luo
Pattern Recognit.5
2023 Discriminative Geometric-Structure-Based Deep Hashing for Large-Scale Image Retrieval
abstract
Deep hashing reaps the benefits of deep learning and hashing technology, and has become the mainstream of large-scale image retrieval. It generally encodes image into hash code with feature similarity preserving, that is, geometric-structure preservation, and achieves promising retrieval results. In this article, we find that existing geometric-structure preservation manner inadequately ensures feature discrimination, while improving feature discrimination of hash code essentially determines hash learning retrieval performance. This fact principally spurs us to propose a discriminative geometric-structure-based deep hashing method (DGDH), which investigates three novel loss terms based on class centers to induce the so-called discriminative geometrical structure. In detail, the margin-aware center loss assembles samples in the same class to the corresponding class centers for intraclass compactness, then a linear classifier based on class center serves to boost interclass separability, and the radius loss further puts different class centers on a hypersphere to tentatively reduce quantization errors. An efficient alternate optimization algorithm with guaranteed desirable convergence is proposed to optimize DGDH. We theoretically analyze the robustness and generalization of the proposed method. The experiments on five popular benchmark datasets demonstrate superior image retrieval performance of the proposed DGDH over several state of the arts.
Guohua Dong, Xiang Zhang 0008, Xiaobo Shen 0001, Long Lan, Zhigang Luo, Xiaomin Ying
IEEE Trans. Cybern.2
2023 Learning to Purification for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification is a challenging and promising task in computer vision. Nowadays unsupervised person re-identification methods have achieved great progress by training with pseudo labels. However, how to purify feature and label noise is less explicitly studied in the unsupervised manner. To purify the feature, we take into account two types of additional features from different local views to enrich the feature representation. The proposed multi-view features are carefully integrated into our cluster contrast learning to leverage more discriminative cues that the global feature easily ignored and biased. To purify the label noise, we propose to take advantage of the knowledge of teacher model in an offline scheme. Specifically, we first train a teacher model from noisy pseudo labels, and then use the teacher model to guide the learning of our student model. In our setting, the student model could converge fast with the supervision of the teacher model thus reduce the interference of noisy labels as the teacher model greatly suffered. After carefully handling the noise and bias in the feature learning, our purification modules are proven to be very effective for unsupervised person re-identification. Extensive experiments on two popular person re-identification datasets demonstrate the superiority of our method. Especially, our approach achieves a state-of-the-art accuracy 85.8% @mAP and 94.5% @Rank-1 on the challenging Market-1501 benchmark with ResNet-50 under the fully unsupervised setting. Code has been available at: https://github.com/tengxiao14/Purification_ReID.
Long Lan, Xiao Teng, Jing Zhang 0037, Xiang Zhang 0008, Dacheng Tao
IEEE Trans. Image Process.4
2023 Local-to-Global Deep Clustering on Approximate Uniform Manifold
abstract
Deep clustering usually treats the clustering assignments as supervisory signals to learn a more compact representation with deep neural networks, under the guidance of clustering-oriented losses. Nevertheless, we observe that, without reliable supervision, such losses for global clustering would destroy the locally geometric structure underlying data. In this paper, we propose a local-to-global deep clustering method based on approximate uniform manifold (LGC-AUM) to address this issue in a two-stage fashion. In the local stage, an intra-manifold preservation loss is proposed to preserve intra-manifold structures locally on basis of approximate uniform manifold, and an inter-manifold discrimination loss is for global inter-manifold structure. Thus, this stage serves to learn more discriminative structure-preserving features by reducing the correlations between different manifolds, which paves the way for the final clustering. Build off the learned features, the second stage explores a clustering loss based on approximate uniform manifold to establish stable network training for effective clustering with two auxiliary distributions. Experiments on five benchmark datasets verify the efficacy of our LGC-AUM as compared to several well-behaved clustering counterparts.
Xiang Zhang 0008, Long Lan, Zhigang Luo
IEEE Trans. Knowl. Data Eng.2
2022 Dynamic Hypergraph Convolutional Network
abstract
Hypergraph Convolutional Network (HCN) has be-come a proper choice for capturing high-order relationships. Existing HCN methods are tailored for static hypergraphs, which are unsuitable for the dynamic evolution in real-world scenarios. In this paper, we explore a dynamic HCN based on the attention mechanism (DyHCN) for time series prediction. It not only effectively exploits the spatial and temporal relationships in the dynamic hypergraph, but also continuously aggregates the temporal evolution cues of time-varying hypergraphs with the global and local embeddings. Specifically, these merits can be attributed to 1) dynamic hypergraph construction (DHC), which captures the feature of historical context content and provides a guideline for dynamic hypergraph construction; 2) spatio-temporal hypergraph convolution module (STHC), responsible for extracting the spatial and temporal relationships among nodes and hyperedges, and 3) collaborative prediction module (CP), for the overall time-varying hypergraphs embedding aggregation. Such modules endeavor to well learn feature embedding from nodes, hyperedges, and hypergraphs, which produces informative representations for downstream tasks. Experiments on three datasets including Tiingo, Stocktwits, and NYC-Taxi demonstrate that the proposed DyHCN achieves sound performance over existing cousins, and both STHC and CP modules play a key role in modeling the dynamic evolution property of hypergraphs.
Fuli Feng, Zhigang Luo, Xiang Zhang 0008, Wenjie Wang 0007, Xiao Luo 0001, Chong Chen 0002, Xian-Sheng Hua 0001
ICDE4
2022 Counterfactual Causal Adversarial Networks for Domain Adaptation
Yan Jia 0001, Xiang Zhang 0008, Long Lan, Zhigang Luo
ICONIP (6)2
2022 Logit Distillation via Student Diversity
Dingyao Chen, Long Lan, Mengzhu Wang, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo
ICONIP (5)4
2022 Self-Reinforcing Feedback Domain Adaptation Channel
Yan Jia 0001, Xiang Zhang 0008, Long Lan, Zhigang Luo
ICONIP (1)2
2022 Frustratingly Easy Knowledge Distillation via Attentive Similarity Matching
abstract
Knowledge distillation is an effective approach to transferring knowledge from the large teacher network to its small proxy student one, thereby letting the proxy student work on those resource-limited mobile devices. Most previous arts manually select the paired intermediate layers of teacher and student networks to align their pertinent features by dimension reduction. This sort of approach may confront information loss and insufficient layer-wise alignment that limit knowledge transferability. In this paper, we propose a simple and effective knowledge distillation method named attentive similarity matching (ASM). ASM at first concatenates the teacher’s intermediate features and the student’s ones together to enhance similarity representation of all the student’s layers, without involving dimension reduction, then align all cross-layer advanced similarities in an attentively weighted manner for semantic calibration. Experiments of image classification on three popular datasets show the effectiveness of the proposed method as compared to its previous cousins.
Dingyao Chen, Huibin Tan, Long Lan, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo
ICPR4
2022 Implicit Feature Alignment For Knowledge Distillation
abstract
Knowledge distillation is a technique of transferring knowledge from a large teacher network to a light student one. Existing studies purely use immediate layers' features for distillation and may fail to gain insufficient semantic knowledge from the teacher. Inspired by recent advances in contrastive learning, we propose to introduce extra light embedding layers of the teacher to enforce its generalization ability and further align the mixup-type features for knowledge distillation in an implicit fashion (IFKD). IFKD allows the student to learn richer structural knowledge, thanks to the learned embedding layers of the teacher. Crucially, benefitting from a plethora of mixed samples, we can further adequately mine much semantic knowledge of the teacher. For efficiency, we propose a simple reversed mixup scheme to organize images and implicitly ensure complete positive information comparisons. Extensive experiments on image classification on two popular datasets including CIFAR-100 and ImageNet verify the effectiveness of our approach as compared to the previous methods.
Dingyao Chen, Mengzhu Wang, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo
ICTAI3
2022 Joint Modality Synergy and Spatio-temporal Cue Purification for Moment Localization
abstract
Currently, many approaches to the sentence query based moment location (SQML) task emphasize (inter-)modality interaction between video and language query via transformer-based cross-attention or contrastive learning. However, they could still face two issues: 1) modality interaction could be unexpectedly friendly to modality specific learning that merely learns modality specific patterns, and 2) modality interaction easily confuses spatio-temporal cues and ultimately makes time cues in the original video ambiguous. In this paper, we propose a modality synergy with spatio-temporal cue purification method (MS2P) for SQML to address the above two issues. Particularly, a conceptually simple modality synergy strategy is explored to keep features modality specific while absorbing the other modality complementary information with both carefully designed cross-attention unit and non-contrastive learning. As a result, modality specific semantics can be calibrated progressively in a safer way. To preserve time cues in original video, we further purify video representation into spatial and temporal parts to enhance localization resolution by the proposed two light-weight sentence-aware filtering operations. Experiments on Charades-STA, TACoS, and ActivityNet Caption datasets show our model outperforms the state-of-the-art approaches by a large margin.
Long Lan, Huibin Tan, Xiang Zhang 0008, Xurui Ma, Zhigang Luo
ICMR4
2022 Multi-scale local cues and hierarchical attention-based LSTM for stock price trend prediction
Xiao Teng, Xiang Zhang 0008, Zhigang Luo
Neurocomputing2
2022 Informative pairs mining based adaptive metric learning for adversarial domain adaptation
Mengzhu Wang, Paul Li, Li Shen 0008, Ye Wang 0023, Shanshan Wang 0008, Wei Wang 0335, Xiang Zhang 0008, Junyang Chen 0001, Zhigang Luo
Neural Networks7
2022 Label Propagated Nonnegative Matrix Factorization for Clustering
abstract
Semi-supervised learning (SSL) that utilizes plenty of unlabeled examples to boost the performance of learning from limited labeled examples is a powerful learning paradigm with widely real-world applications such as information retrieval and document clustering. Label propagation (LP) is a popular SSL method which propagates labels through the dataset along high density areas defined by unlabeled examples, but it is fragile to bridge examples. Semi-supervised K-Means uses labeled examples to initialize clustering centers to separate different examples, however, semi-supervised K-Means fails in the situation of imbalanced issues, that is, the example size of each class varies significantly. This paper proposes a novel label propagated nonnegative matrix factorization method (LPNMF) to handle clean labeled but biased data and its extension LPNMF-E to handle noisy labeled data based on the framework of NMF. LPNMF decomposes the whole dataset into the product of a basis matrix and a coefficient matrix. To propagate labels to unlabeled examples, LPNMF regards the class indicators of labeled examples as their coefficients and iteratively updates both basis matrix and coefficients of unlabeled examples. LPNMF absorbs the merits from both semi-supervised K-Means and label propagation to handle their respective shortages. Specifically, on the one hand, LPNMF learns representative clustering centers based on the distribution of the dataset, similar to semi-supervised K-means, and thus is robust to the bridge examples. On the other hand, LPNMF pushes labels according to the affinity between examples, similar to label propagation, and thus relieves the biased problem. Moreover, we introduce a LPNMF extension to handle the noisy label case. LPNMF-E relaxes the constraint of labeled examples. Since the label of each labeled example also obtains label information from the global distribution of the whole dataset and local manifold of its neighbors, LPNMF-E outputs reliable class indicators even if a portion of examples are incorrectly labeled. Theoretical analyses for the generalization ability of our proposed models are also provided. Experimental results on both clean and noisy labeled datasets confirm the effectiveness of LPNMF and LPNMF-E compared with both LP and the representative semi-supervised K-Means algorithms.
Long Lan, Tongliang Liu, Xiang Zhang 0008, Chuanfu Xu, Zhigang Luo
IEEE Trans. Knowl. Data Eng.3
2022 An Integrated Multi-Task Model for Fake News Detection
abstract
Fake news detection attracts many researchers’ attention due to the negative impacts on the society. Most existing fake news detection approaches mainly focus on semantic analysis of news’ contents. However, the detection performance will dramatically decrease when the content of news is short. In this paper, we propose a novelfake news detection multi-task learning (FDML)model based on the following observations: 1) some certain topics have higher percentages of fake news; and 2) some certain news authors have higher intentions to publish fake news. FDML model investigates the impact of topic labels for the fake news and introduce contextual information of news at the same time to boost the detection performance on the short fake news. Specifically, the FDML model consists of representation learning and multi-task learning parts to train the fake news detection task and the news topic classification task, simultaneously. As far as we know, this is the first fake news detection work that integrates the above two tasks. The experiment results show that the FDML model outperforms state-of-the-art methods on real-world fake news dataset.
Qing Liao 0001, Heyan Chai 0001, Xiang Zhang 0008, Xuan Wang 0002, Wen Xia, Ye Ding 0002
IEEE Trans. Knowl. Data Eng.4
2021 Entity-Aware Biaffine Attention for Constituent Parsing
Xinyi Bai, Xiang Zhang 0008, Zhigang Luo
ICANN (1)3
2021 Joint motion context and clip augmentation for spatio-temporal action detection
abstract
This paper endeavors to leverage spatio-temporal visual cues to improve video-based action detection. As a result, a NOn-Local Action detector based on anchor-free called NOLA is proposed, which is built off a recent moving center detector (MOC) and further extends it by efficiently aggregating long-range spatio-temporal information. In detail, a significantly efficient spatio-temporal motion-aware non-local block is explored to provide global motion contexts for the entire predictive branches of MOC. This byproduct can make the large batch samples run on a resource limited device. Besides, a light-weighted data augmentation method termed clip augmentation designed for video-based tasks is proposed, which serves to improve the generalization ability of the detector with economical scale-and-addition operation. NOLA works with two above schemes in real-time as well. Experiments on two benchmark datasets show that NOLA significantly exceeds MOC. Compared to other existing methods,, NOLA reaches the state-of-the-art, in terms of video-level mean of average precision (video mAP).
Xurui Ma, Xiang Zhang 0008, Chengkun Wu, Chuanfu Xu, Jie Liu 0002, Zhigang Luo
ICMV2
2021 Spatio-Temporal Action Detector with Self-Attention
abstract
In the field of spatio-temporal action detection, some current studies attempt to solve the problem of action detection by using the one-stage object detectors based on anchor-free. Albeit efficiency, more performance boosts are expected. Towards this goal, a Self-Attention MovingCenter Detector (SAMOC) is proposed, which is blessed with two attractive aspects: 1) to effectively capture motion cues, a spatio-temporal self-attention block is explored to reinforce feature representation by aggregating motion-dependent global contexts, and 2) a link branch serves to model the frame-level object dependency, which promotes the confidence scores of correct actions. Experiments on two benchmark datasets show that SAMOC with the proposed two aspects achieves the state-of-the-art and works in real-time as well.
Xurui Ma, Zhigang Luo, Xiang Zhang 0008, Qing Liao 0001, Mengzhu Wang
IJCNN3
2021 InterBN: Channel Fusion for Adversarial Unsupervised Domain Adaptation
abstract
A classifier trained on one dataset rarely works on other datasets obtained under different conditions because of domain shifting. Such a problem is usually solved by domain adaptation methods. In this paper, we propose a novel unsupervised domain adaptation (UDA) method based on Interchangeable Batch Normalization (InterBN) to fuse different channels in deep neural networks for adversarial domain adaptation.Specifically, we first observe that the channels with small batch normalization scaling factor have less influence on the whole domain adaption, followed by a theoretical proof that the scaling factors for some channels will definitely come close to zero when imposing a sparsity regularization. Then, we replace the channels that have smaller scaling factors in the source domain with the mean of the channels which have larger scaling factors in the target domain or vice versa. Such a simple but effective channel fusion scheme can drastically increase the domain adaption ability.Extensive experimental results show that our InterBN significantly outperforms the current adversarial domain adaptation methods by a large margin on four visual benchmarks. In particular, InterBN achieves a remarkable improvement of 7.7% over the conditional adversarial adaptation networks (CDAN) on VisDA-2017 benchmark.
Mengzhu Wang, Wei Wang 0335, Baopu Li, Xiang Zhang 0008, Long Lan, Huibin Tan, Tianyi Liang 0001, Wei Yu 0029, Zhigang Luo
ACM Multimedia4
2021 Enhancing the association in multi-object tracking via neighbor graph
abstract
Most modern multi-object tracking (MOT) systems for videos follow the tracking-by-detection paradigm, where objects of interest are first located in each frame then associated correspondingly to form their intact trajectories. In this setting, the appearance features of objects usually provide the most important cues for data association, but it is very susceptible to occlusions, illumination variations, and inaccurate detections, thus easily resulting in incorrect trajectories. To address this issue, in this study we propose to make full use of the neighboring information. Our motivations derive from the observations that people tend to move in a group. As such, when an individual target's appearance is remarkably changed, the observer can still identify it with its neighbor context. To model the contextual information from neighbors, we first utilize the spatiotemporal relations among trajectories to efficiently select suitable neighbors for targets. Subsequently, we construct neighbor graph for each target and corresponding neighbors then employ the graph convolutional networks (GCNs) to model their relations and learn the graph features. To the best of our knowledge, it is the first time to explicitly leverage neighbor cues via GCN in MOT. Finally, standardized evaluations on the MOT16 and MOT17 data sets demonstrate that our approach can remarkably reduce the identity switches whilst achieve state-of-the-art overall performance.
Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Xindong Peng, Zhigang Luo
Int. J. Intell. Syst.3
2021 Semantic-consistent cross-modal hashing for large-scale image retrieval
Xuesong Gu, Guohua Dong, Xiang Zhang 0008, Long Lan, Zhigang Luo
Neurocomputing3
2021 A generic MOT boosting framework by combining cues from SOT, tracklet and re-identification
Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Zhigang Luo
Knowl. Inf. Syst.3
2021 Improving knowledge distillation via an expressive teacher
Jie Liu 0002, Xiang Zhang 0008
Knowl. Based Syst.3
2021 Learning deep discriminative embeddings via joint rescaled features and log-probability centers
Huayue Cai, Xiang Zhang 0008, Long Lan, Guohua Dong, Chuanfu Xu, Xinwang Liu 0002, Zhigang Luo
Pattern Recognit.2
2021 Near-Online Multi-Pedestrian Tracking via Combining Multiple Consistent Appearance Cues
abstract
An important cue for multi-pedestrian tracking in video is the consistent appearance of an individual for quite a while. In this paper, we address multi-pedestrian tracking by learning a robust appearance model from the paradigm of tracking by detection. To separate detections of different pedestrians while assembling detections of the same pedestrian, we take advantage of the cue of consistent appearance and exploit three types of evidence from the recent, past and near-future. Existing online approaches only exploit the detection-to-detection and sequence-to-detection metrics, which focus on the recent and past appearance patterns respectively, while the future pedestrian appearance is simply ignored. This drawback is remedied in this paper by further considering the sequence-to-sequence metric, which resorts to near-future appearance presentation. Adaptive combination weights are learned to fuse these three different metrics. Moreover, we propose a novel Focal Triplet Loss to make the model focus more on hard examples than the easy ones. We demonstrate that this can significantly enhance the discriminating power of the model compared with treating every sample equally. Effectiveness and efficiency of the proposed method is verified by conducting comprehensive ablation studies and comparing with many competitive (offline/online/near-online) counterparts on the MOT16 and MOT17 Challenges.
Weijiang Feng, Long Lan, Yong Luo 0002, Yue Yu 0001, Xiang Zhang 0008, Zhigang Luo
IEEE Trans. Circuits Syst. Video Technol.5
2021 Nocal-Siam: Refining Visual Features and Response With Advanced Non-Local Blocks for Real-Time Siamese Tracking
abstract
Siamese trackers contain two core stages, i.e., learning the features of both target and search inputs at first and then calculating response maps via the cross-correlation operation, which can also be used for regression and classification to construct typical one-shot detection tracking framework. Although they have drawn continuous interest from the visual tracking community due to the proper trade-off between accuracy and speed, both stages are easily sensitive to the distracters in search branch, thereby inducing unreliable response positions. To fill this gap, we advance Siamese trackers with two novel non-local blocks named Nocal-Siam, which leverages the long-range dependency property of the non-local attention in a supervised fashion from two aspects. First, a target-aware non-local block (T-Nocal) is proposed for learning the target-guided feature weights, which serve to refine visual features of both target and search branches, and thus effectively suppress noisy distracters. This block reinforces the interplay between both target and search branches in the first stage. Second, we further develop a location-aware non-local block (L-Nocal) to associate multiple response maps, which prevents them inducing diverse candidate target positions in the future coming frame. Experiments on five popular benchmarks show that Nocal-Siam performs favorably against well-behaved counterparts both in quantity and quality.
Huibin Tan, Xiang Zhang 0008, Long Lan, Wenju Zhang, Zhigang Luo
IEEE Trans. Image Process.2
2020 Robust Normalized Squares Maximization for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) attempts to transfer specific knowledge from one domain with labeled data to another domain without labels. Recently, maximum squares loss has been proposed to tackle UDA problem but it does not consider the prediction diversity which has proven beneficial to UDA. In this paper, we propose a novel normalized squares maximization (NSM) loss in which the maximum squares is normalized by the sum of squares of class sizes. The normalization term enforces the class sizes of predictions to be balanced to explicitly increase the diversity. Theoretical analysis shows that the optimal solution to NSM is one-hot vectors with balanced class sizes, i.e., NSM encourages both discriminate and diverse predictions. We further propose a robust variant of NSM, RNSM, by replacing the square loss with L2,1-norm to reduce the influence of outliers and noises. Experiments of cross-domain image classification on two benchmark datasets illustrate the effectiveness of both NSM and RNSM. RNSM achieves promising performance compared to state-of-the-art methods. The code is available at https://github.com/wj-zhang/NSM.
Wenju Zhang, Xiang Zhang 0008, Qing Liao 0001, Wenjing Yang 0002, Long Lan, Zhigang Luo
CIKM2
2020 Towards Making Unsupervised Graph Hashing Robust
abstract
Unsupervised hashing without supervision easily deteriorates in the case of grossly corrupted data. Motivated by robust optimization, this paper proposes a dual-graph regularized robust hashing (DGRH) based on both manifold smoothness and robust estimators in a more intuitive manner. Orthogonal to existing robust hashing methods, DGRH directly removes the outliers of datasets with M-estimator to exert robustness. In specific, it intends to recover low-rank representation from corrupted data via l1loss while preserving neighborhood relationships among samples with dual-graph regularization. Although DGRH seems a simple extension of robust PCA on graphs with hashing trick, it is easy to implement yet effective. Theory analysis is provided to support our claim. Experiments of image retrieval on three popular benchmark datasets show the efficacy of DGRH as compared to several well-behaved representative counterparts.
Xuesong Gu, Guohua Dong, Xiang Zhang 0008, Long Lan, Zhigang Luo
ICME3
2020 Learning sequence-to-sequence affinity metric for near-online multi-object tracking
Weijiang Feng, Long Lan, Xiang Zhang 0008, Zhigang Luo
Knowl. Inf. Syst.3
2020 Enhancing unsupervised domain adaptation by discriminative relevance regularization
Wenju Zhang, Xiang Zhang 0008, Long Lan, Zhigang Luo
Knowl. Inf. Syst.2
2020 Object-aware semantics of attention for image captioning
Long Lan, Xiang Zhang 0008, Guohua Dong, Zhigang Luo
Multim. Tools Appl.3
2020 GateCap: Gated spatial and semantic attention model for image captioning
Long Lan, Xiang Zhang 0008, Zhigang Luo
Multim. Tools Appl.3
2020 Maximum Mean and Covariance Discrepancy for Unsupervised Domain Adaptation
Wenju Zhang, Xiang Zhang 0008, Long Lan, Zhigang Luo
Neural Process. Lett.2
2019 Attentional Residual Dense Factorized Network for Real-Time Semantic Segmentation
Long Lan, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICANN (3)3
2019 Person re-identification via adaptive verification loss
Hui Tian 0005, Xiang Zhang 0008, Long Lan, Zhigang Luo
Neurocomputing2
2019 Label guided correlation hashing for large-scale cross-modal retrieval
Guohua Dong, Xiang Zhang 0008, Long Lan, Zhigang Luo
Multim. Tools Appl.2
2019 Stacked Marginal Time Warping for Temporal Alignment
Xiang Zhang 0008, Liquan Nie, Long Lan, Xuhui Huang, Zhigang Luo
Neural Process. Lett.1
2019 Nonnegative Constrained Graph Based Canonical Correlation Analysis for Multi-view Feature Learning
Huibin Tan, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
Neural Process. Lett.2
2018 Margin-Embedding Canonical Correlation Analysis with Feature Selection for Person Re-Identification
abstract
Canonical correlation analysis (CCA) is a classical subspace learning method of capturing the common semantic information underlying multi-view data. It has been used in person re-identification (re-ID) task by treating the task of matching identical individuals across non-overlapping multi-cameras as a multi-view learning problem. However, CCA-based reID methods still achieve unsatisfactory results because few jointly consider discriminative margin information and selecting importantly relevant features. To address this issue, we propose a novel l2,1-norm regularized margin-embedding CCA ( l2,1-MCCA), which learns a generalized discriminative subspace by employing more discriminative margin information. Moreover, the new method enforces the l2,1-norm regularization term over the learned subspace to identify the relevant features. Both lightweight and effective schemes can benefit from each other and endeavor to enlarge the interclass variations whilst reducing the intra-class variations. Experiments on three popular datasets show the efficacy of l2,1-MCCA as compared with recently representative re-ID methods.
Linfei Ma, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
ICIP2
2018 Graph-Laplacian Correlated Low-Rank Representation for Subspace Clustering
abstract
Subspace clustering seeks to segment a given unlabeled data into clusters with the hope of each cluster corresponding to a union of low-dimensional subspaces. Among them, low-rank representation (LRR) is a promising potential method which intends to build a good affinity matrix by using the self-expression of inputs. However, it completely ignores the important data locality. Although several works in this regard have considered the local geometric structure through the Laplacian regularizer, they also neglect the correlation of the data. In this paper, we propose a graph-Laplacian correlated low-rank representation model (GCLRR) to address such an issue. Particularly, GCLRR factorizes the self-expression as the product of two low-dimensional matrices, of which one is the latent representation of the self-expression. On the basis of the latent representation, the Laplacian regularizer is integrated with the orthogonal constraint together and behaves like the spectral clustering. Moreover, we devise a Frobenius norm based trace loss and use it to constrain both the latent representation and the self-expression to capture the correlation of the data. Our improved trace loss is more efficient than the original one. More importantly, GCLRR provides an effective unified framework to seamlessly integrate both aspects above. Then, we optimize GCLRR in the frame of alternating direction method (ADM) and fortunately derive the analytical solution to each subproblem. Experiments of motion segmentation and image clustering confirm the efficacy of the proposed GCLRR.
Huayue Cai, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
ICIP3
2018 Discrete Graph Hashing via Affine Transformation
abstract
In unsupervised graph-based hashing for large-scale image retrieval, many efforts have been made to bridge the gap between the learned graph embedding and the corresponding binary codes. Relatively, few studies focus on the issue of the discrimination of graph embedding. In this paper, we firstly devise a discrete graph hashing model (DGH) that smooths graph embedding and simultaneously solving binary codes under the balanced discrete constraint, which equals a novel method of jointly learning graph embedding and spectral rotation, theoretically. To further induce discriminant graph embedding, we substitute affine transformation for spectral rotation in our DGH (abbreviated as ADGH). This is because affine transformation can accommodate both rotational angle and distance of graph embedding, while respecting the neighborhood structure among most samples. Besides, each subproblem of ADGH can yield the closed-form solution. Experiments of image retrieval on three benchmark datasets show that ADGH outperforms the representative hashing methods in quantity.
Guohua Dong, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
ICME2
2018 Cross-Layer Convolutional Siamese Network for Visual Tracking
Yanyin Chen, Huibin Tan, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
ICONIP (2)4
2018 Background Subtraction via 3D Convolutional Neural Networks
abstract
Background subtraction can be treated as the binary classification problem of highlighting the foreground region in a video whilst masking the background region, and has been broadly applied in various vision tasks such as video surveillance and traffic monitoring. However, it still remains a challenging task due to complex scenes and for lack of the prior knowledge about the temporal information. In this paper, we propose a novel background subtraction model based on 3D convolutional neural networks (3D CNNs) which combines temporal and spatial information to effectively separate the foreground from all the sequences in an end-to-end manner. Different from conventional models, we view background subtraction as three-class classification problem, i.e., the foreground, the background and the boundary. This design can obtain more reasonable results than existing baseline models. Experiments on the Change Detection 2012 dataset verify the potential of our model in both quantity and quality.
Yongqiang Gao, Huayue Cai, Xiang Zhang 0008, Long Lan, Zhigang Luo
ICPR3
2018 Box-constrained Discriminant Projective Non-negative Matrix Factorization through Augmented Lagrangian Multiplier Method
abstract
Projective non-negative matrix factorization (PNMF) learns a non-negative projection matrix to project high-dimensional examples onto a lower-dimensional space spanned by the transpose of the learned projection matrix. Since PNMF can learn parts-based representation, it has attracted ample attention from computer vision community. However, existing PNMF methods either completely ignore labels of the dataset or endure the slow convergent optimization algorithm. In this paper, we propose a box-constrained discriminant PNMF (BDPNMF) method to address these issues. Specifically, BDPNMF jointly exploits the Fisher's criterion and the augmented Lagrangian multiplier (ALM) method into PNMF to boost discriminative capacity of the learned subspace and its efficiency. Experimental results on four popular face image datasets confirm the efficacy of BDPNMF compared to previous PNMF methods in quantity.
Huayue Cai, Xiang Zhang 0008, Zhigang Luo, Xuhui Huang
IJCNN2
2018 Multi-granularity Hierarchical Attention Siamese Network for Visual Tracking
abstract
Speed and accuracy are the two most important focuses for many visual tracking methods. Recently, siamese networks based trackers have shown very promising potentials in both aspects, which develop a twin network to measure the responses between target and hypotheses with a fully convolutional operation. However, the learned response maps are vulnerable to background clutters and scale changes as they ignore priori knowledge such as the object salience and multi-granularity cues. To explore the benefits of priori, this paper devises a multi-granularity hierarchical attention siamese network tracker (MHA-Siam) to further enhance the tracking stability without sacrificing real-time speed. Particularly, the channel-wise attention mechanism is exploited here to filter out the background while remain the salient object region; then, the response maps of the coarse-to-finer multi-layer features are fused to capture multi-granularity location information helpful for improvement in tracking stability. To make full use of them, MHA-Siam imposes the element-wise max-and-sum operation on them to induce a reliable response map for accurate location. Experiments of visual tracking on OTB benchmark shows the superiority of MHA-Siam with the competitive efficiency to its counterpart trackers.
Xiang Zhang 0008, Huibin Tan, Long Lan, Zhigang Luo, Xuhui Huang
IJCNN2
2018 Ranking-Embedded Transfer Canonical Correlation Analysis for Person Re-Identification
abstract
Person re-identification (re-ID) seeks to match the identical individuals across different cameras and is still a challenging visual task due to substantial variances of person appearance in complex scenarios. Different from most of conventional person re-ID methods, which generally reduce person re-ID task to either a multi-view learning problem or a multi- domain learning problem alone, this paper treats such a task as a multi-view multi-domain (MVMD) learning problem to exploit the both benefits by refreshing canonical correlation analysis (CCA) with two improvements, termed as ranking-embedded transfer CCA (RTCCA). Specifically, to bridge the semantic gap between different views, we first embed a ranking weight matrix into CCA to strength the correlations among the multi-view images of the same identity and simultaneously to weaken that of different identities. Furthermore, we utilize the well-known distribution metric maximum mean discrepancy (MMD) as a regularization term to reduce the domain shift between training set and testing set. More importantly, the two improvements benefit from each other and the joint merit can further boost the re-ID performance. Experiments on three benchmarks verify the efficacy of the proposed RTCCA when compared with the recently representative baseline person re-ID methods.
Linfei Ma, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
IJCNN2
2018 Collaborative Subspace Graph Hashing for Cross-modal Retrieval
abstract
Current hashing methods for cross-modal retrieval generally attempt to learn the separate modality-specific transformation matrices to embed multi-modality data into a latent common subspace, and usually ignore the fact that respecting the diversity of multi-modality features in the latent subspace could be beneficial for retrieval improvements. To this, we propose a collaborative subspace graph hashing method (CSGH) to perform a two-stage collaborative learning framework for cross-modal retrieval. Particularly, CSGH first embeds multi-modality data into separate latent subspaces through individual modality-specific transformation matrices, and then connects these latent subspaces to a common Hamming space through a shared transformation matrix. In this framework, CSGH considers the modality-specific neighborhood structure and the cross-modal correlation within multi-modality data through the Laplacian regularization and the graph based correlation constraint, respectively. To solve CSGH, we develop an alternative procedure to optimize it, and fortunately, each sub-problem of CSGH has the elegant analytical solution. Experiments of cross-modal retrieval on Wiki, NUS-WIDE, Flickr25K and Flickr1M datasets show the effectiveness of CSGH compared with the state-of-the-art cross-modal hashing methods.
Xiang Zhang 0008, Guohua Dong, Yimo Du, Chengkun Wu, Zhigang Luo, Canqun Yang
ICMR1
2018 Low-Rank Matrix Recovery via Continuation-Based Approximate Low-Rank Minimization
Xiang Zhang 0008, Yongqiang Gao, Long Lan, Xuhui Huang, Zhigang Luo
PRICAI (1)1
2017 Attention Focused Spatial Pyramid Pooling for Boxless Action Recognition in Still Images
Weijiang Feng, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICANN (2)2
2017 GNMF Revisited: Joint Robust k-NN Graph and Reconstruction-Based Graph Regularization for Image Clustering
Wenju Zhang, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICANN (2)3
2017 Unsupervised domain adaptation with joint supervised sparse coding and discriminative regularization term
abstract
Domain adaptation (DA) attempts to enhance the generalization capability of classifier through narrowing the gap of the distributions across domains. This paper focuses on unsupervised domain adaptation where labels are not available in target domain. Most existing approaches explore the domain-invariant features shared by domains but ignore the discriminative information of source domain. To address this issue, we propose a discriminative domain adaptation method (DDA) to reduce domain shift by seeking a common latent subspace jointly using supervised sparse coding (SSC) and discriminative regularization term. Particularly, DDA adapts SSC to yield discriminative coefficients of target data and further unites with discriminative regularization term to induce a common latent subspace across domains. We show that both strategies can boost the ability of transferring knowledge from source to target domain. Experiments on two real world datasets demonstrate the effectiveness of our proposed method over several existing state-of-the-art domain adaptation methods.
Xiang Zhang 0008, Wenju Zhang, Xuhui Huang, Naiyang Guan, Zhigang Luo
ICIP2
2017 Boxless Action Recognition in Still Images via Recurrent Visual Attention
Weijiang Feng, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICONIP (2)2
2017 Audio visual speech recognition with multimodal recurrent neural networks
abstract
Studies on nowadays human-machine interface have demonstrated that visual information can enhance speech recognition accuracy especially in noisy environments. Deep learning has been widely used to tackle such audio visual speech recognition (AVSR) problem due to its astonishing achievements in both speech recognition and image recognition. Although existing deep learning models succeed to incorporate visual information into speech recognition, none of them simultaneously considers sequential characteristics of both audio and visual modalities. To overcome this deficiency, we proposed a multimodal recurrent neural network (multimodal RNN) model to take into account the sequential characteristics of both audio and visual modalities for AVSR. In particular, multimodal RNN includes three components, i.e., audio part, visual part, and fusion part, where the audio part and visual part capture the sequential characteristics of audio and visual modalities, respectively, and the fusion part combines the outputs of both modalities. Here we modelled the audio modality by using a LSTM RNN, and modelled the visual modality by using a convolutional neural network (CNN) plus a LSTM RNN, and combined both models by a multimodal layer in the fusion part. We validated the effectiveness of the proposed multimodal RNN model on a multi-speaker AVSR benchmark dataset termed AVletters. The experimental results show the performance improvements comparing to the known highest audio visual recognition accuracies on AVletters, and confirm the robustness of our multimodal RNN model.
Weijiang Feng, Naiyang Guan, Yuan Li 0007, Xiang Zhang 0008, Zhigang Luo
IJCNN4
2017 Dependency-based long short term memory network for drug-drug interaction extraction
abstract
BACKGROUND: Drug-drug interaction extraction (DDI) needs assistance from automated methods to address the explosively increasing biomedical texts. In recent years, deep neural network based models have been developed to address such needs and they have made significant progress in relation identification. METHODS: We propose a dependency-based deep neural network model for DDI extraction. By introducing the dependency-based technique to a bi-directional long short term memory network (Bi-LSTM), we build three channels, namely, Linear channel, DFS channel and BFS channel. All of these channels are constructed with three network layers, including embedding layer, LSTM layer and max pooling layer from bottom up. In the embedding layer, we extract two types of features, one is distance-based feature and another is dependency-based feature. In the LSTM layer, a Bi-LSTM is instituted in each channel to better capture relation information. Then max pooling is used to get optimal features from the entire encoding sequential data. At last, we concatenate the outputs of all channels and then link it to the softmax layer for relation identification. RESULTS: To the best of our knowledge, our model achieves new state-of-the-art performance with the F-score of 72.0% on the DDIExtraction 2013 corpus. Moreover, our approach obtains much higher Recall value compared to the existing methods. CONCLUSIONS: The dependency-based Bi-LSTM model can learn effective relation information with less feature engineering in the task of DDI extraction. Besides, the experimental results show that our model excels at balancing the Precision and Recall values.
Wei Wang 0130, Xi Yang 0020, Canqun Yang, Xiang Zhang 0008, Chengkun Wu
BMC Bioinform.5
2016 Enhancing temporal alignment with autoencoder regularization
abstract
Temporal alignment aligns two temporal sequences and is quite challenging due to drastic differences among temporal sequences and source data from different views. Canonical time warping (CTW) has shown great potential in temporal alignment tasks because it can reduce data redundancy by transforming high-dimensional data to a lower-dimensional subspace via canonical correlation analysis (CCA). However, CTW cannot uncover the underlying nonlinear structure embedded in the dataset. In this paper, we propose an autoencoder regularized canonical time warping method (AECTW) to overcome this drawback. Specifically, AECTW enhances lower-dimensional representation of each sequence by incorporating an autoencoder regularization, meanwhile reveals the nonlinear structure of features by explicit nonlinear transformation. By these strategies, AECTW significantly boosts CTW in temporal alignment tasks. Experiments on both synthetic data and two practical human action datasets demonstrate that AECTW outperforms the representative DTW-based methods.
Liquan Nie, Yuanyuan Wang 0004, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
IJCNN3
2016 Correntropy induced metric based graph regularized non-negative matrix factorization
Yuanyuan Wang 0004, Shuyi Wu, Bin Mao, Xiang Zhang 0008, Zhigang Luo
Neurocomputing4
2016 Nonparametric Statistical Active Contour Based on Inclusion Degree of Fuzzy Sets
abstract
In this paper, inclusion degree of fuzzy sets is introduced to image segmentation. The image segmentation problem is novelly modeled as the minimization of the overlapping rates between the inside and outside regions, subject to a constraint on the total length of the region boundaries. Considering the similar properties of fuzzy sets and statistical image domain, we use fuzzy membership functions to represent the inside and outside regions and utilize nonparametric density estimates to estimate them. Then, the inclusion degree of fuzzy sets is adopted to formulate the overlapping rates between the inside and outside regions. We solve the inclusion-degree-based optimization problem by deriving the associated gradient flow and applying curve evolution techniques. Experimental results on both synthetic and real images confirm the effectiveness of the proposed method. Compared with the previous active contour models formulated to solve the same nonparametric statistical segmentation problem, our method performs well in efficiency and evolution time.
Maoguo Gong, Hao Li 0009, Xiang Zhang 0008, Qiunan Zhao, Bin Wang 0027
IEEE Trans. Fuzzy Syst.3
2015 Labelwalking nonnegative matrix factorization
abstract
Semi-supervised learning (SSL) utilizes plenty of unlabeled examples to boost the performance of learning from limited labeled examples. Due to its great discriminant power, SSL has been widely applied to various real-world tasks such as information retrieval, pattern recognition, and speech separa- tion. Label propagation (LP) is a popular SSL method which propagates labels through the dataset along high density areas defined by unlabeled examples, LP assumes nearby examples should share the same label, thus, it unavoidably pushes the labels to the wrong examples, especially when different la- beled examples are not strictly separated. Seed K-means uses labeled examples to initialize class centers, and avoid getting stuck in poor local optima comparing to traditional K-means, however the hard constraint of each example's membership makes Seed K-means failed in many real world applications. This paper proposes a novel label walking nonnegative matrix factorization method (LWNMF) to handle labeled examples in SSL based on the framework of NMF. LWNMF decomposes the whole dataset into the product of a basis matrix and a coefficient matrix, and to travel labels to unlabeled examples, LWNMF regards the class indicators of labeled examples as their coefficients and iteratively updates both basis matrix and coefficients of unlabeled examples. Since LWNMF learns comprehensive class centroids, labels iteratively walk to unlabeled examples through these significant centroids.
Long Lan, Naiyang Guan, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICASSP3
2015 Local Coordinate Projective Non-negative Matrix Factorization
abstract
Non-negative matrix factorization (NMF) decomposes a group of non-negative examples into both lower-rank factors including the basis and coefficients. It still suffers from the following deficiencies: 1) it does not always ensure the decomposed factors to be sparse theoretically, and 2) the learned basis often stays away from original examples, and thus lacks enough representative capacity. This paper proposes a local coordinate projective NMF (LCPNMF) to overcome the above deficiencies. Particularly, LCPNMF induces sparse coefficients by relaxing the original PNMF model meanwhile encouraging the basis to be close to original examples with the local coordinate constraint. Benefitting from both strategies, LCPNMF can significantly boost the representation ability of the PNMF. Then, we developed the multiplicative update rule to optimize LCPNMF and theoretically proved its convergence. Experimental results on three popular frontal face image datasets verify the effectiveness of LCPNMF comparing to the representative methods.
Qing Liao 0001, Xiang Zhang 0008, Naiyang Guan, Qian Zhang 0001
ICMLA2
2015 Constrained Projective Non-negative Matrix Factorization for Semi-supervised Multi-label Learning
abstract
This paper formulates multi-label learning as a constrained projective non-negative matrix factorization (CPNMF) problem which concentrates on a variant of the original projective NMF (PNMF) and explicitly introduces an auxiliary basis to learn the semantic subspace and boosts its discriminating ability by exploiting labeled and unlabeled examples together. Particularly, it propagates labels of the labeled examples to the unlabeled ones by enforcing coefficients of examples sharing identical semantic contents to be identical based on a hard constraint, i.e., embedding the class indicator of labeled examples into their coefficients. CPNMF preserves the geometrical structure of dataset via manifold regularization meanwhile captures the inherent structure of labels by using label correlations. We developed a multiplicative update rule (MUR) based algorithm to optimize CPNMF and proved its convergence. Experiments of image annotation on Corel dataset, text categorization on Rcv1v2 dataset, and text clustering on two popular text corpuses suggest the effectiveness of CPNMF.
Xiang Zhang 0008, Naiyang Guan, Zhigang Luo, Xuejun Yang
ICMLA1
2015 Semi-supervised Non-negative Local Coordinate Factorization
Cherong Zhou, Xiang Zhang 0008, Naiyang Guan, Xuhui Huang, Zhigang Luo
ICONIP (2)2
2015 Two-Dimensional Euler PCA for Face Recognition
Huibin Tan, Xiang Zhang 0008, Naiyang Guan, Dacheng Tao, Xuhui Huang, Zhigang Luo
MMM (2)2
2015 Non-negative Low-Rank and Group-Sparse Matrix Factorization
Shuyi Wu, Xiang Zhang 0008, Naiyang Guan, Dacheng Tao, Xuhui Huang, Zhigang Luo
MMM (2)2
2015 Robust Local Coordinate Non-negative Matrix Factorization via Maximum Correntropy Criteria
abstract
Non-negative matric factorization (NMF) decomposes a given data matrix X into the product of two lower dimensional non-negative matrices U and V. It has been widely applied in pattern recognition and computer vision because of its simplicity and effectiveness. However, existing NMF methods often fail to learn the sparse representation on high-dimensional dataset, especially when some examples are heavily corrupted. In this paper, we propose a robust local coordinate NMF method (RLCNMF) by using the maximum correntropy criteria to overcome such deficiency. Particularly, RLCNMF induces sparse coefficients by imposing the local coordinate constraint over both factors. To solve RLCNMF, we developed a multiplicative update rules and theoretically proved its convergence. Experimental results on popular image datasets verify the effectiveness of RLCNMF comparing with the representative methods.
Qing Liao 0001, Xiang Zhang 0008, Naiyang Guan, Qian Zhang 0001
SMC2
2015 Symmetric Non-negative Matrix Factorization Based Link Partition Method for Overlapping Community Detection
abstract
Partitioning links rather than nodes is effective in overlapping community detection (OCD) on complex networks. However, it consumes high CPU and memory overheads because the volume of links is huge especially when the network is rather complex. In this paper, we proposes a symmetric non-negative matrix factorization (SNMF) based link partition method called SNMF-Link to overcome this deficiency. In particular, SNMF-Link represents data in a lower-dimensional space spanned by the node-link incidence matrix. By solving a lighter SNMF problem, SNMF-Link learns the clustering indicators of each links. Since traditional multiplicative update rule (MUR) based optimization algorithm for SNMF suffers from slow convergence, we applied the augmented Lagrangian method (ALM) to efficiently optimize SNMF. Experimental results show that SNMF-Link is much more efficient than the representative clustering algorithms without reducing the OCD performance.
Xiang Zhang 0008, Naiyang Guan, Wenju Zhang, Xuhui Huang, Shuyi Wu, Zhigang Luo
SMC1
2014 Soft-constrained nonnegative matrix factorization via normalization
abstract
Semi-supervised clustering aims at boosting the clustering performance on unlabeled samples by using labels from a few labeled samples. Constrained NMF (CNMF) is one of the most significant semi-supervised clustering methods, and it factorizes the whole dataset by NMF and constrains those labeled samples from the same class to have identical encodings. In this paper, we propose a novel soft-constrained NMF (SCNMF) method by softening the hard constraint in CNMF. Particularly, SCNMF factorizes the whole dataset into two lower-dimensional factor matrices by using multiplicative update rule (MUR). To utilize the labels of labeled samples, SCNMF iteratively normalizes both factor matrices after updating them with MURs to make encodings of labeled samples close to their label vectors. It is therefore reasonable to believe that encodings of unlabeled samples are also close to their corresponding label vectors. Such strategy significantly boosts the clustering performance even when the labeled samples are rather limited, e.g., each class owns only a single labeled sample. Since the normalization procedure never increases the computational complexity of MUR, SCNMF is quite efficient and effective in practices. Experimental results on face image datasets illustrate both efficiency and effectiveness of SCNMF compared with both NMF and CNMF.
Long Lan, Naiyang Guan, Xiang Zhang 0008, Dacheng Tao, Zhigang Luo
IJCNN3
2014 Box-constrained projective nonnegative matrix factorization via augmented Lagrangian method
abstract
Projective non-negative matrix factorization (P-NMF) projects a set of examples onto a subspace spanned by a non-negative basis whose transpose is regarded as the projection matrix. Since PNMF learns a natural parts-based representation, it has been successfully used in text mining and pattern recognition. However, it is non-trivial to analyze the convergence of the optimization algorithms for PNMF because its objective function is non-convex. In this paper, we propose a Box-constrained PNMF (BPNMF) method to overcome this deficiency of PNMF. In particular, BPNMF introduces an auxiliary variable, i.e., the coefficients of examples, and incorporates the following two types of constraints: 1) each entry of the basis is non-negative and upper-bounded, i.e., box-constrained, and 2) the coefficients equal to the projected points of the examples. The first box constraint makes the basis to be bound and the second equality constraint keeps its equivalence to PNMF. Similar to PNMF, BPNMF is difficult because the objective function is non-convex. To solve BPNMF, we developed an efficient algorithm in the frame of augmented Lagrangian multiplier (ALM) method and proved that the ALM-based algorithm converges to local minima. Experimental results on two face image datasets demonstrate the effectiveness of BPNMF compared with the representative methods.
Xiang Zhang 0008, Naiyang Guan, Long Lan, Dacheng Tao, Zhigang Luo
IJCNN1
2013 Orthogonal Nonnegative Locally Linear Embedding
abstract
Nonnegative matrix factorization (NMF) decomposes a nonnegative dataset X into two low-rank nonnegative factor matrices, i.e., W and H, by minimizing either Kullback-Leibler (KL) divergence or Euclidean distance between X and WH. NMF has been widely used in pattern recognition, data mining and computer vision because the non-negativity constraints on both W and H usually yield intuitive parts-based representation. However, NMF suffers from two problems: 1) it ignores geometric structure of dataset, and 2) it does not explicitly guarantee parts-based representation on any datasets. In this paper, we propose an orthogonal nonnegative locally linear embedding (ONLLE) method to overcome aforementioned problems. ONLLE assumes that each example embeds in its nearest neighbors and keeps such relationship in the learned subspace to preserve geometric structure of a dataset. For the purpose of learning parts-based representation, ONLLE explicitly incorporates an orthogonality constraint on the learned basis to keep its spatial locality. To optimize ONLLE, we applied an efficient fast gradient descent (FGD) method on Stiefel manifold which accelerates the popular multiplicative update rule (MUR). The experimental results on real-world datasets show that FGD converges much faster than MUR. To evaluate the effectiveness of ONLLE, we conduct both face recognition and image clustering on real-world datasets by comparing with the representative NMF methods.
Lei Weit, Naiyang Guan, Xiang Zhang 0008, Zhigang Luo, Dacheng Tao
SMC3
2012 Graph Based Semi-supervised Non-negative Matrix Factorization for Document Clustering
abstract
Non-negative matrix factorization (NMF) approximates a non-negative matrix by the product of two low-rank matrices and achieves good performance in clustering. Recently, semi-supervised NMF (SS-NMF) further improves the performance by incorporating part of the labels of few samples into NMF. In this paper, we proposed a novel graph based SS-NMF (GSS-NMF). For each sample, GSS-NMF minimizes its distances to the same labeled samples and maximizes the distances against different labeled samples to incorporate the discriminative information. Since both labeled and unlabeled samples are embedded in the same reduced dimensional space, the discriminative information from the labeled samples is successfully transferred to the unlabeled samples, and thus it greatly improves the clustering performance. Since the traditional multiplicative update rule converges slowly, we applied the well-known projected gradient method to optimizing GSS-NMF and the proposed algorithm can be applied to optimizing other manifold regularized NMF efficiently. Experimental results on two popular document datasets, i.e., Reuters21578 and TDT-2, show that GSS-NMF outperforms the representative SS-NMF algorithms.
Naiyang Guan, Xuhui Huang, Long Lan, Zhigang Luo, Xiang Zhang 0008
ICMLA (1)5
2012 Sparse Representation Based Discriminative Canonical Correlation Analysis for Face Recognition
abstract
Canonical correlation analysis (CCA) has been widely used in pattern recognition and machine learning. However, both CCA and its extensions sometimes cannot give satisfactory results. In this paper, we propose a new CCA-type method termed sparse representation based discriminative CCA (SPDCCA) by incorporating sparse representation and discriminative information simultaneously into traditional CCA. In particular, SPDCCA not only preserves the sparse reconstruction relationship within data based on sparse representation, but also preserves the maximum-margin based discriminative information, and thus it further enhances the classification performance. Experimental results on Yale, Extended Yale B, and ORL datasets show that SPDCCA outperforms both CCA and its extensions including KCCA, LPCCA and LDCCA in face recognition.
Naiyang Guan, Xiang Zhang 0008, Zhigang Luo, Long Lan
ICMLA (1)2
2012 Semi-supervised Non-negative Patch Alignment Framework
abstract
Non-negative matrix factorization (NMF) learns the latent semantic space more direct and reliable than the latent semantic indexing (LSI) and the spectral clustering methods, thus performs well in document clustering. Recently, semi-supervised NMF such as N2S2L, CNMF and unsupervised method such as GNMF significantly improve the face recognition performance, but they are designed for classification. In this paper, we combine both geometric structure and label information with NMF under the non-negative patch alignment framework (NPAF) to form SS-NPAF. Due to this combination, it greatly improves the clustering performance. To optimize SS-NPAF, we apply the well-known projected gradient method to overcome the slow convergence problem of the mostly used multiplicative update rule. Experimental results on two popular document datasets, i.e., Reuters21578 and TDT-2, show that SS-NPAF outperforms the representative SS-NMF algorithms.
Long Lan, Xuhui Huang, Naiyang Guan, Zhigang Luo, Xiang Zhang 0008
ICMLA (1)5