Zhigang Luo

dblp:02/2039 · DBLP profile ↗
← Back
119ranked-venue papers
0as first author
59since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 65 · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 15 since 2021Databases, data management, data science and information retrieval · 14 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 since 2021Computer networks · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 5Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Divide-Then-Rule: A Cluster-Driven Hierarchical Interpolator for Attribute-Missing Graphs
abstract
Deep graph clustering (DGC) for attribute-missing graphs is an unsupervised task aimed at partitioning nodes with incomplete attributes into distinct clusters. Existing imputation methods for attribute-missing graphs often fail to account for the varying amounts of information available across node neighborhoods, leading to unreliable results. To address this issue, we propose a novel method named Divide-Then-Rule Graph Completion (DTRGC). This method first addresses nodes with sufficient known neighborhood information and treats the imputed results as new knowledge to iteratively impute more challenging nodes, while leveraging clustering information to correct imputation errors. Specifically, Dynamic Cluster-Aware Feature Propagation initializes missing node attributes by adjusting propagation weights based on the clustering structure. Subsequently, Hierarchical Neighborhood-Aware Imputation categorizes attribute-missing nodes into three groups based on the completeness of their neighborhood attributes. The imputation is performed hierarchically, prioritizing the groups with nodes that have the most available neighborhood information. The cluster structure is then used to refine the imputation and correct potential errors. Finally, Hop-wise Representation Enhancement integrates information across multiple hops, thereby enriching the expressiveness of node representations. Experimental results on 6 widely used graph datasets show that DTRGC significantly improves the clustering performance of various DGC methods under attribute-missing graphs.
Yaowen Hu, Wenxuan Tu, Yue Liu 0008, Miaomiao Li 0001, Wenpeng Lu, Zhigang Luo, Xinwang Liu 0002, Ping Chen 0004
ACM Multimedia6
2025 Measurements and Analysis of Millimeter-Wave Propagation for 500 km/h Ultrahigh-Speed Maglev Train Communications Between Train and Trackside
abstract
Millimeter-Wave (mmWave) has emerged as a competitive solution to meet the high-data-rate requirements of future high speed train (HST) communications due to its abundant spectrum resources. However, existing research on actual measurements of mmWave channels for HST is relatively scarce, let alone at speed up to 500 km/h. This paper reports the world’s first mmWave HST channel measurement campaigns with a maglev train at a maximum speed of 500 km/h in the semi-closed station scenario. A large number of large-and small-scale fading characteristics are obtained and analyzed, such as path loss, shadow fading, root-mean-squared delay spread (RMS-DS), K factor, coherence bandwidth, and decorrelation distance. We propose an improved cluster-based Saleh-Valenzuela channel model which is more suitable for mmWave channel measurements. Numerical results and best fitting distribution results are obtained for seven key cluster parameters. The simulated and measured channels are compared to verify the accuracy of this improved model. Finally, the significance of the study is discussed, and the measurement and analysis results are compared with the existing measurement results in the train station scenario. The measurement results fill the gap of the 500 km/h mmWave HST measurements, and help to guide and design the future mmWave HST communication system.
Xichen Liu, Lin Yang 0004, Zhigang Luo, Guangrong Yue
IEEE Internet Things J.4
2025 STFormer: Spatial-Temporal-Aware Transformer for Video Instance Segmentation
abstract
Video instance segmentation (VIS) is a challenging task, requiring handling object classification, segmentation, and tracking in videos. Existing Transformer-based VIS approaches have shown remarkable success, combining encoded features and instance queries as decoder inputs. However, their decoder inputs are low-resolution due to computational cost, resulting in a loss of fine-grained information, sensitivity to background interference, and poor handling of small objects. Moreover, the queries are randomly initialized without location information, hindering convergence efficiency and accurate object instance localization. To address these issues, we propose a novel VIS approach, STFormer, with a spatial-temporal feature aggregation (STFA) module and spatial-temporal-aware Transformer (STT). Specifically, STFA obtains robust high-resolution masked features efficiently for the decoder, while STT's location-guided instance query (LGIQ) improves initial instance queries. STFormer preserves more fine-grained information, improves convergence efficiency, and localizes object instance features accurately. Extensive experiments on YouTube-VIS 2019, YouTube-VIS 2021, and OVIS datasets show that STFormer outperforms mainstream VIS methods.
Wei Wang 0335, Mengzhu Wang, Huibin Tan, Long Lan, Zhigang Luo, Xinwang Liu 0002, Kenli Li 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 HCL: A Hierarchical Contrastive Learning Framework for Zero-Shot Relation Extraction
abstract
Zero-shot relation extraction (ZSRE) is shown to become more significant in the current information extraction system, which aims at predicting relation classes that lack annotations or have just never appeared during training. Previous works focus on projecting sentences with their corresponding relation descriptions to an intermediate semantic space and searching the nearest semantic for predicting unseen classes. Though these methods can achieve sound performance, they only obtain inferior semantic information via a trivial distance metric and neglect the interaction in the instance representations. We are thus motivated to tackle these issues and propose a hierarchical contrastive learning (HCL) framework for ZSRE including projection-level and instance-level modules. Specifically, the projection-level component replaces the distance score function by contrastive loss to connect the input sentence with the relation semantic space. And the instance-level component integrates the external knowledge from sentence entities to establish new contrastive pairs for efficiently learning representations from mutual information. The experimental results on three well-known datasets demonstrate that our model surpasses the existing SOTA by at most 18.97% improvement on the F1 score when unseen classes are 15. Moreover, our model can achieve more competitive performance alone with the increasing number of unseen classes.
Tianwei Yan 0001, Shan Zhao 0002, Minghao Hu 0001, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 FRCL-MNER: A Finer Grained Rank-Based Contrastive Learning Framework for Multimodal NER
abstract
Multimodal named entity recognition (MNER) is an emerging field that aims to automatically detect named entities and classify their categories, utilizing input text and auxiliary resources such as images. While previous studies have leveraged object detectors to preprocess images and fuse textual semantics with corresponding image features, these methods often overlook the potential finer grained information within each modality and may exacerbate error propagation due to predetection. To address these issues, we propose a finer grained rank-based contrastive learning (FRCL) framework for MNER. This framework employs a global-level contrastive learning to align multimodal semantic features and a Top-K rank-based mask strategy to construct positive-negative pairs, thereby learning a finer grained multimodal interaction representation. Experimental results from three well-known social media datasets reveal that our approach surpasses existing strong baselines, and achieves up to a 1.54% improvement on the Twitter2015 dataset. Extensive discussions further confirm the effectiveness of our approach. We will release the source code on https://github.com/augusyan/FRCL.
Tianwei Yan 0001, Shan Zhao 0002, Wentao Ma 0003, Shezheng Song, Chengyu Wang 0008, Zhibo Rao, Shizhao Chen, Zhigang Luo, Xinwang Liu 0002
IEEE Trans. Neural Networks Learn. Syst.8
2025 Sparse Low-Rank Multi-View Subspace Clustering With Consensus Anchors and Unified Bipartite Graph
abstract
Anchor technology is popularly employed in multi-view subspace clustering (MVSC) to reduce the complexity cost. However, due to the sampling operation being performed on each individual view independently and not considering the distribution of samples in all views, the produced anchors are usually slightly distinguishable, failing to characterize the whole data. Moreover, it is necessary to fuse multiple separated graphs into one, which leads to the final clustering performance heavily subject to the fusion algorithm adopted. What is worse, existing MVSC methods generate dense bipartite graphs, where each sample is associated with all anchor candidates. We argue that this dense-connected mechanism will fail to capture the essential local structures and degrade the discrimination of samples belonging to the respective near anchor clusters. To alleviate these issues, we devise a clustering framework named SL-CAUBG. Specifically, we do not utilize sampling strategy but optimize to generate the consensus anchors within all views so as to explore the information between different views. Based on the consensus anchors, we skip the fusion stage and directly construct the unified bipartite graph across views. Most importantly, norm and Laplacian-rank constraints employed on the unified bipartite graph make it capture both local and global structures simultaneously. norm helps eliminate the scatters between anchors and samples by constructing sparse links and guarantees our graph to be with clear anchor-sample affinity relationship. Laplacian-rank helps extract the global characteristics by measuring the connectivity of unified bipartite graph. To deal with the nondifferentiable objective function caused by norm, we adopt an iterative re-weighted method and the Newton's method. To handle the nonconvex Laplacian-rank, we equivalently transform it as a convex trace constraint. We also devise a four-step alternate method with linear complexity to solve the resultant problem. Substantial experiments show the superiority of our SL-CAUBG.
Shengju Yu, Suyuan Liu, Siwei Wang 0001, Chang Tang, Zhigang Luo, Xinwang Liu 0002, En Zhu
IEEE Trans. Neural Networks Learn. Syst.5
2024 MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space Model
abstract
Tracking by detection has been the prevailing paradigm in the field of Multi-object Tracking (MOT). These methods typically rely on the Kalman Filter to estimate the future locations of objects, assuming linear object motion. However, they fall short when tracking objects exhibiting nonlinear and diverse motion in scenarios like dancing and sports. In addition, there has been limited focus on utilizing learning-based motion predictors in MOT. To address these challenges, we resort to exploring data-driven motion prediction methods. Inspired by the great expectation of state space models (SSMs), such as Mamba, in long-term sequence modeling with near-linear complexity, we introduce a Mamba-based motion model named Mamba moTion Predictor (MTP). MTP is designed to model the complex motion patterns of objects like dancers and athletes. Specifically, MTP takes the spatial-temporal location dynamics of objects as input, captures the motion pattern using a bi-Mamba encoding layer, and predicts the next motion. In real-world scenarios, objects may be missed due to occlusion or motion blur, leading to premature termination of their trajectories. To tackle this challenge, we further expand the application of MTP. We employ it in an autoregressive way to compensate for missing observations by utilizing its own predictions as inputs, thereby contributing to more consistent trajectories. Our proposed tracker, MambaTrack, demonstrates advanced performance on benchmarks such as Dancetrack and SportsMOT, which are characterized by complex motion and severe occlusion.
Changcheng Xiao, Qiong Cao, Zhigang Luo, Long Lan
ACM Multimedia3
2024 Discriminative object tracking by domain contrast
Huayue Cai, Xiang Zhang 0008, Long Lan, Changcheng Xiao, Chuanfu Xu, Jie Liu 0002, Zhigang Luo
Comput. Vis. Image Underst.7
2024 Consistency-constrained unsupervised video anomaly detection framework based on Co-teaching
Wenhao Shao, Praboda Rajapaksha, Noël Crespi, Xuechen Zhao, Mengzhu Wang, Xinwang Liu 0002, Zhigang Luo
Neurocomputing8
2024 IoUformer: Pseudo-IoU prediction with transformer for visual tracking
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Yibing Zhan, Zhigang Luo
Neural Networks6
2024 MotionTrack: Learning motion predictor for multiple object tracking
Changcheng Xiao, Qiong Cao, Long Lan, Xiang Zhang 0008, Zhigang Luo, Dacheng Tao
Neural Networks6
2024 Compressing the Multiobject Tracking Model via Knowledge Distillation
abstract
Recent multiobject tracking (MOT) methods usually use very deep neural networks to achieve competitive accuracy, which inevitably results in degraded inference speed. To strike a better balance between tracking accuracy and speed, in this work, we propose to compress the MOT model via knowledge distillation (KD), enabling the more lightweight student model to obtain similar performance as the teacher model. Nonetheless, despite KD has been well studied for simpler tasks such as image classification, the complexity of MOT poses new challenges because the MOT model is more sensitive to foreground information than the classification model. To deal with that, we first propose attention-guided feature distillation, which focuses the student model on the crucial region (foreground and the region with strong discrepancy against itself) of the teacher’s feature map. Moreover, we propose foreground mask, which leverages the knowledge from the teacher model to filter out the low-quality soft labels from the background, thereby reducing their negative effects for distillation. Evaluations on several benchmarks demonstrate that the proposed KD method can make the student network achieve leading performance, meanwhile running faster than the teacher network 20.0%–27.4% and reducing the parameters 28.5%–87.1%. To the best of our knowledge, this is the first work to compress the MOT model via KD.
Tianyi Liang 0001, Mengzhu Wang, Junyang Chen 0001, Dingyao Chen, Zhigang Luo, Victor C. M. Leung
IEEE Trans. Comput. Soc. Syst.5
2024 How to Construct Corresponding Anchors for Incomplete Multiview Clustering
abstract
Anchor based incomplete multiview clustering has grasped growing interest recently because of its great success in effectively partitioning multimodal data. However, due to the absence of label information, the constructed anchors could be mismatched. Such an Anchor Mismatching Problem (AMP) will cause the structure of generated bipartite graph to be chaotic, degrading the clustering performance. To tackle this issue, we design an algorithm termed Constructing Corresponding Anchors for Incomplete Multiview Clustering (CCA-IMC). Specifically, we first devise a permutation strategy to transform anchors on each view. Subsequently, we directly generate the consensus bipartite graph, which is shared for all incomplete views, by the transformed anchors rather than by fusing each view-specific bipartite graph. Afterwards, all anchors and permutation matrices as well as the consensus bipartite graph are jointly optimized in one common framework so as to promote each other. In such ways, anchors are rearranged towards correct matching relationship according to the consensus graph structure. In addition to these, our CCA-IMC has also been proven to be with linear time and memory overheads, which makes it able to scale up to work with large-scale tasks. Massive experiments implemented on ten popular datasets give evidence of our superiorities compared to current strong IMC competitors.
Shengju Yu, Siwei Wang 0001, Yi Wen 0001, Zhigang Luo, En Zhu, Xinwang Liu 0002
IEEE Trans. Circuits Syst. Video Technol.5
2024 Empirical Study on Millimeter-Wave Vehicle-to-Vehicle Channel Characteristics in Open and Underground Parking Lot Scenarios
abstract
Parking lots, as one of the most significant vehicle-to-vehicle (V2V) communication scenarios, have rarely been studied for their millimeter-wave (mmWave) V2V channel characteristics. This paper presents the 41 GHz mmWave channel measurements and analysis results in open and underground parking lot scenarios for the first time. A large number of channel measurement data are obtained, including received signal strength indications (RSSIs) and channel impulse responses (CIRs) at different distances and different arrival angles for three cases. By clustering the power delay profile (PDP), cluster parameters such as cluster number, inter-cluster duration, cluster arrival rate and decay factor are obtained. The channel non-stationarity and consistency are revealed and discussed by various methods and from different perspectives. Large- and small-scale fading characteristics parameters, such as path loss, shadow fading, root-mean-square delay spread (RMS-DS), angle spread (AS), fade depth, and Ricean K factor are analyzed and discussed. The mmWave V2V channel characteristics of three cases are compared and analyzed. The measurement and modeling results fill the gaps of mmWave V2V channel measurements in open and underground parking lot scenarios and provide valuable suggestions for the design and modeling of mmWave V2V communication systems.
Xichen Liu, Dingrui Ke, Lin Yang 0004, Zhigang Luo, Guangrong Yue
IEEE Trans. Wirel. Commun.4
2024 Measurements and Analysis of Millimeter-Wave Propagation From In-Station to Out-Station in High-Speed Railway Between Train and Trackside
abstract
Train stations, as one of the most significant buildings on a high-speed railway (HSR), have been rarely studied for their millimeter-wave (mmWave) propagation characteristics. This paper presents the first 41 GHz mmWave band channel measurements on a HSR closed station at ultra-high speeds (up to 288 km/h). Multi-dimensional channel measurement data at different azimuth angles and speeds from inside to outside the station are obtained. It is found that the received signal strength indication (RSSI) experiences a significant attenuation caused by the influence of the station gate. In combination with the Saleh-Valenzuela (SV) model, an improved time delay clustering algorithm is used to cluster the power delay profile (PDP). According to the clustering results, parameters such as inter-cluster duration, cluster arrival rate, and intra-cluster delay spread are obtained. Most of the large- and small-scale fading characteristics, including path loss, shadow fading, PDP, angle spread, small-scale fading characteristic quantity, small-scale amplitude fading, and Rician K factor are investigated. The measurement and modeling results fill the gaps of channel measurements in closed HSR station scenarios and provide valuable suggestions for the design and modelling of HSR station communication systems.
Xichen Liu, Lin Yang 0004, Zhigang Luo, Daizhong Yu, Guangrong Yue
IEEE Trans. Wirel. Commun.4
2023 Atomic-action-based Contrastive Network for Weakly Supervised Temporal Language Grounding
abstract
As one knows, an event often consists of several actions while each action is atomic. Inspired by this insight, we propose a novel framework named Atomic-action-based Contrastive Network model (ACN) for weakly supervised temporal language grounding task to localize the query-related event moment in an untrimmed video, without access to any temporal annotations. Specifically, ACN first determines the accurate moment boundary of each action in a query-agnostic way. This can adequately exploit homogeneous visual cues while impeding the heterogeneity of the query from hurting the atomicity of visual action, i.e., action boundary. To effectively localize the query-related event, we seek the discriminative words in the given query, and explore a composite-grained contrastive module to retrieve those corresponding atomic actions in the common latent space across modalities. This boosts feature discrimination of visual event segment to remove irrelevant action video segments. Experiments on two popular datasets show the efficacy of our model.
Hongzhou Wu, Xuechen Zhao, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo
ICME7
2023 When content-centric networking meets multi-criteria group decision-making: Optimal cache placement policy achieved by MARCOS with q-rung orthopair fuzzy set pair analysis
abstract
Assessing the cache placement policy (CPP) is a crucial step in the deployment process of content-centric networking (CCN), which is still an open question. In this paper, assessing the CPP is formulated to be a multi-criteria group decision-making (MCGDM) issue since it involves the consideration of multiple experts, and a new aggregated MCGDM algorithm is presented for dealing with this issue. On that account, an assessment criteria system is to formulate to portray these experts’ considerations for assessing CPPs, and then the concepts of q-rung orthopair fuzzy set (q-ROFS) and set pair analysis (SPA) are presented to directly and indirectly define the group preference information of CPPs with respect to the criteria, respectively. Later, some revised aggregation operators are employed for aggregating the group preference information. Then, a new integrated objective criteria weights (IOCW) method based on score value and combined criteria weights based on non-linear comprehensive method are given for determining the important degrees of criteria. Based on IOCW method, non-linear comprehensive method, score function, distance measure, and aggregation operators, a new integrated q-rung orthopair fuzzy MCGDM model based on MARCOS (Measurement of Alternatives and Ranking according to COmpromise Solution) method is developed to rank CPPs. Finally, an example is given to illustrate the evaluation process of CPP, and the advantages of developed method are verified by comparative analysis.
Xindong Peng, Harish Garg, Zhigang Luo
Eng. Appl. Artif. Intell.3
2023 Domain-specific feature recalibration and alignment for multi-source unsupervised domain adaptation
abstract
Abstract Traditional unsupervised domain adaptation (UDA) usually assumes that the source domain has labels and the target domain has no labels. In a real environment, labelled source domain data usually comes from multiple different distributions. To handle this problem, multi‐source unsupervised domain adaptation (MUDA) is proposed. Multi‐source unsupervised domain adaptation aims to adapt the model trained on multi‐labelled source domains to the unlabelled target domain. In this paper, a novel MUDA method by domain‐specific feature recalibration and alignment (FRA) is proposed. Specifically, to achieve feature recalibration, the authors leverage channel attention to pick out significant channels and spatial attention to focus on important features in different channels. Such integration of channel and spatial attention can lead to effective domain‐specific feature recalibration that may be of great importance to MUDA. In addition, to achieve better MUDA, the authors propose domain‐specific feature alignment which consists of Maximum Mean Discrepancy and JS‐divergence loss. Maximum Mean Discrepancy can reduce the difference between the source domain and target domain. Meanwhile, JS‐divergence loss may ensure the prediction consistency of different classifiers in the source domains. Four experiments have proved that FRA can achieve significantly better results in popular benchmarks for MUDA.
Mengzhu Wang, Dingyao Chen, Fangzhou Tan, Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Zhigang Luo
IET Comput. Vis.7
2023 Online intervention siamese tracking
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Changcheng Xiao, Zhigang Luo
Inf. Sci.6
2023 Boosting unsupervised domain adaptation: A Fourier approach
Mengzhu Wang, Shanshan Wang 0008, Ye Wang 0023, Wei Wang 0335, Tianyi Liang 0001, Junyang Chen 0001, Zhigang Luo
Knowl. Based Syst.7
2023 SiamDF: Tracking training data-free siamese tracker
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Zhigang Luo
Neural Networks5
2023 Video anomaly detection with NTCN-ML: A novel TCN for multi-instance learning
Wenhao Shao, Ruliang Xiao, Praboda Rajapaksha, Mengzhu Wang, Noël Crespi, Zhigang Luo, Roberto Minerva
Pattern Recognit.6
2023 Reducing bi-level feature redundancy for unsupervised domain adaptation
Mengzhu Wang, Shanshan Wang 0008, Wei Wang 0335, Li Shen 0008, Xiang Zhang 0008, Long Lan, Zhigang Luo
Pattern Recognit.7
2023 q-Rung orthopair fuzzy inequality derived from equality and operation
Xindong Peng, Zhigang Luo
Soft Comput.3
2023 Discriminative Geometric-Structure-Based Deep Hashing for Large-Scale Image Retrieval
abstract
Deep hashing reaps the benefits of deep learning and hashing technology, and has become the mainstream of large-scale image retrieval. It generally encodes image into hash code with feature similarity preserving, that is, geometric-structure preservation, and achieves promising retrieval results. In this article, we find that existing geometric-structure preservation manner inadequately ensures feature discrimination, while improving feature discrimination of hash code essentially determines hash learning retrieval performance. This fact principally spurs us to propose a discriminative geometric-structure-based deep hashing method (DGDH), which investigates three novel loss terms based on class centers to induce the so-called discriminative geometrical structure. In detail, the margin-aware center loss assembles samples in the same class to the corresponding class centers for intraclass compactness, then a linear classifier based on class center serves to boost interclass separability, and the radius loss further puts different class centers on a hypersphere to tentatively reduce quantization errors. An efficient alternate optimization algorithm with guaranteed desirable convergence is proposed to optimize DGDH. We theoretically analyze the robustness and generalization of the proposed method. The experiments on five popular benchmark datasets demonstrate superior image retrieval performance of the proposed DGDH over several state of the arts.
Guohua Dong, Xiang Zhang 0008, Xiaobo Shen 0001, Long Lan, Zhigang Luo, Xiaomin Ying
IEEE Trans. Cybern.5
2023 A Closer Look at the Joint Training of Object Detection and Re-Identification in Multi-Object Tracking
abstract
Unifying object detection and re-identification (ReID) into a single network enables faster multi-object tracking (MOT), while this multi-task setting poses challenges for training. In this work, we dissect the joint training of detection and ReID from two dimensions: label assignment and loss function. We find previous works generally overlook them and directly borrow the practices from object detection, inevitably causing inferior performance. Specifically, we identify a qualified label assignment for MOT should: 1) have the assignment cost aware of ReID cost, not just detection cost; 2) provide sufficient positive samples for robust feature learning while avoiding ambiguous positives (i.e., the positives shared by different ground-truth objects). To achieve the above goals, we first propose Identity-aware Label Assignment, which jointly considers the assignment cost of detection and ReID to select positive samples for each instance without ambiguities. Moreover, we advance a novel Discriminative Focal Loss that integrates ReID predictions with Focal Loss to focus the training on the discriminative samples. Finally, we upgrade the strong baseline FairMOT with our techniques and achieve up to 7.0 MOTA / 54.1% IDs improvements on MOT16/17/20 benchmarks under favorable inference speed, which verifies our tailored label assignment and loss function for MOT are superior to those inherited from object detection.
Tianyi Liang 0001, Baopu Li, Mengzhu Wang, Huibin Tan, Zhigang Luo
IEEE Trans. Image Process.5
2023 Local-to-Global Deep Clustering on Approximate Uniform Manifold
abstract
Deep clustering usually treats the clustering assignments as supervisory signals to learn a more compact representation with deep neural networks, under the guidance of clustering-oriented losses. Nevertheless, we observe that, without reliable supervision, such losses for global clustering would destroy the locally geometric structure underlying data. In this paper, we propose a local-to-global deep clustering method based on approximate uniform manifold (LGC-AUM) to address this issue in a two-stage fashion. In the local stage, an intra-manifold preservation loss is proposed to preserve intra-manifold structures locally on basis of approximate uniform manifold, and an inter-manifold discrimination loss is for global inter-manifold structure. Thus, this stage serves to learn more discriminative structure-preserving features by reducing the correlations between different manifolds, which paves the way for the final clustering. Build off the learned features, the second stage explores a clustering loss based on approximate uniform manifold to establish stable network training for effective clustering with two auxiliary distributions. Experiments on five benchmark datasets verify the efficacy of our LGC-AUM as compared to several well-behaved clustering counterparts.
Xiang Zhang 0008, Long Lan, Zhigang Luo
IEEE Trans. Knowl. Data Eng.4
2023 OMG: Towards Effective Graph Classification Against Label Noise
abstract
Graph classification is a fundamental problem with diverse applications in bioinformatics and chemistry. Due to the intricate procedures of manual annotations in graphical domains, there may be abundant noisy labels of graphs in practice, resulting in poor performance for existing supervised methods. Thus, it is necessary and urgent to study the problem of graph classification with label noise. However, this problem is challenging due to the overfitting of noisy data as well as complicated relational structures of graphs. To handle this problem, we present a simple but effective approach called cOupledMix forGraph Contrast (OMG), which combines coupled Mixup with graph contrastive learning in the feature space. On the one hand, to improve the model generalization, we take convex combination of sample pairs in the feature space for positive pair construction. On the other hand, to accomplish effective optimization, we offer challenging negatives by multiple sample Mixup with different emphasis. To further reduce the impact of noisy data, we develop a neighbour-aware noise removal strategy, which promotes the smoothness in the neighbourhood of samples following the principle of curriculum learning. Extensive experiments on a range of benchmark datasets demonstrate the superiority of our proposed OMG.
Li Shen 0008, Mengzhu Wang, Xiao Luo 0001, Zhigang Luo, Dacheng Tao
IEEE Trans. Knowl. Data Eng.5
2023 Low-Latency Dimensional Expansion and Anomaly Detection Empowered Secure IoT Network
abstract
The Internet of Things (IoT) consists of a myriad of smart devices and offers tremendous innovation opportunities in industry, homes, and businesses to enhance the productivity and the quality of life. However, ecosystem of infrastructures and the services associated with IoT devices have introduced a new set of vulnerabilities and threats, resulting in abnormal values of information collected by sensors, jeopardizing system security. To secure sensor networks, it must be possible to detect such anomalies or sequences of patterns in IoT devices that significantly deviate from normal behavior. To perform this task, this paper proposes a real-time streaming anomaly detection method based on a Bloom filter combined with hashing. This method expands the data dimensions through a hashing algorithm, and then adopts competitive learning (Winner-Take-All) to build a multi-layer Bloom Filter anomaly detection model. The feasibility of the proposed algorithm is verified theoretically using two datasets, KDD (to detect anomalies at the TCP/IP network level) and Credit (to detect anomalies during credit card transactions). The simulation results show that the proposed in this paper can effectively identify anomalies in the simulation data streams, with almost 95% accuracy for both datasets.
Wenhao Shao, Yanyan Wei, Praboda Rajapaksha, Dun Li, Zhigang Luo, Noël Crespi
IEEE Trans. Netw. Serv. Manag.5
2023 COVAD: Content-Oriented Video Anomaly Detection using a Self-Attention based Deep Learning Model
abstract
Video anomaly detection has always been a hot topic and attracting an increasing amount of attention. Much of the existing methods on video anomaly detection depend on processing the entire video rather than considering only the significant context. This paper proposes a novel video anomaly detection method named COVAD, which mainly focuses on the region of interest in the video instead of the entire video. Our proposed COVAD method is based on an auto-encoded convolutional neural network and coordinated attention mechanism, which can effectively capture meaningful objects in the video and dependencies between different objects. Relying on the existing memory-guided video frame prediction network, our algorithm can more effectively predict the future motion and appearance of objects in the video. Our proposed algorithm obtained better experimental results on multiple data sets and outperformed the baseline models considered in our analysis. At the same time we improve a visual test that can provide pixel-level anomaly explanations.
Wenhao Shao, Praboda Rajapaksha, Yanyan Wei, Dun Li, Noël Crespi, Zhigang Luo
Virtual Real. Intell. Hardw.6
2022 Attention-based Adversarial Partial Domain Adaptation
abstract
With the rapid development of vision-based deep learning (DL), it is an effective method to generate large-scale synthetic data to supplement real data to train the DL models for domain adaptation. However, previous vanilla domain adaptation methods generally assume the same label space, and such an assumption is no longer valid for a more realistic scenario where it requires adaptation from a larger and more diverse source domain to a smaller target domain with less number of classes. To handle this problem, we propose an attention-based adversarial partial domain adaptation (AADA). Specifically, we leverage adversarial domain adaptation to augment the target domain by using source domain, then we can readily turn this task into a vanilla domain adaptation. Meanwhile, to accurately focus on the transferable features, we apply attention-based method to train the adversarial networks to obtain better transferable semantic features. Experiments on four benchmarks demonstrate that the proposed method outperforms existing methods by a large margin, especially on the tough domain adaptation tasks, e.g. VisDA-2017.
Mengzhu Wang, Shan An, Xiao Luo 0001, Xiong Peng, Wei Yu 0029, Junyang Chen 0001, Zhigang Luo
ICASSP7
2022 Dynamic Hypergraph Convolutional Network
abstract
Hypergraph Convolutional Network (HCN) has be-come a proper choice for capturing high-order relationships. Existing HCN methods are tailored for static hypergraphs, which are unsuitable for the dynamic evolution in real-world scenarios. In this paper, we explore a dynamic HCN based on the attention mechanism (DyHCN) for time series prediction. It not only effectively exploits the spatial and temporal relationships in the dynamic hypergraph, but also continuously aggregates the temporal evolution cues of time-varying hypergraphs with the global and local embeddings. Specifically, these merits can be attributed to 1) dynamic hypergraph construction (DHC), which captures the feature of historical context content and provides a guideline for dynamic hypergraph construction; 2) spatio-temporal hypergraph convolution module (STHC), responsible for extracting the spatial and temporal relationships among nodes and hyperedges, and 3) collaborative prediction module (CP), for the overall time-varying hypergraphs embedding aggregation. Such modules endeavor to well learn feature embedding from nodes, hyperedges, and hypergraphs, which produces informative representations for downstream tasks. Experiments on three datasets including Tiingo, Stocktwits, and NYC-Taxi demonstrate that the proposed DyHCN achieves sound performance over existing cousins, and both STHC and CP modules play a key role in modeling the dynamic evolution property of hypergraphs.
Fuli Feng, Zhigang Luo, Xiang Zhang 0008, Wenjie Wang 0007, Xiao Luo 0001, Chong Chen 0002, Xian-Sheng Hua 0001
ICDE3
2022 Toward to Real Low-Resolution Person Re-identification: A New Dataset and Baseline
abstract
Person re-identification(re-id) aims at querying and identifying the same target pedestrian in multiple non-overlapping cameras. However, in real scenarios, many person images captured by surveillance cameras tend to have low resolution due to camera hardware, shooting distance, viewing angle, etc. Many existing re-id methods mainly address the crossresolution person re-id problem, i.e., the high-low resolution mismatch problem. In contrast, the problem of matching low-resolution (LR) gallery images and query images with each other has been less studied. In this paper, we address this problem by developing a new gun-ball camera-based person re-id dataset and designing a LR re-id baseline model for this dataset to tackle the LR person matching problem. Extensive experiments validate the effectiveness of the proposed LR baseline model.
Dongting Sun, Long Lan, Zhigang Luo
ICME4
2022 Counterfactual Causal Adversarial Networks for Domain Adaptation
Yan Jia 0001, Xiang Zhang 0008, Long Lan, Zhigang Luo
ICONIP (6)4
2022 Logit Distillation via Student Diversity
Dingyao Chen, Long Lan, Mengzhu Wang, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo
ICONIP (5)6
2022 Self-Reinforcing Feedback Domain Adaptation Channel
Yan Jia 0001, Xiang Zhang 0008, Long Lan, Zhigang Luo
ICONIP (1)4
2022 AAT: Non-local Networks for Sim-to-Real Adversarial Augmentation Transfer
Mengzhu Wang, Shanshan Wang 0008, Tianwei Yan 0001, Zhigang Luo
ICONIP (4)4
2022 Frustratingly Easy Knowledge Distillation via Attentive Similarity Matching
abstract
Knowledge distillation is an effective approach to transferring knowledge from the large teacher network to its small proxy student one, thereby letting the proxy student work on those resource-limited mobile devices. Most previous arts manually select the paired intermediate layers of teacher and student networks to align their pertinent features by dimension reduction. This sort of approach may confront information loss and insufficient layer-wise alignment that limit knowledge transferability. In this paper, we propose a simple and effective knowledge distillation method named attentive similarity matching (ASM). ASM at first concatenates the teacher’s intermediate features and the student’s ones together to enhance similarity representation of all the student’s layers, without involving dimension reduction, then align all cross-layer advanced similarities in an attentively weighted manner for semantic calibration. Experiments of image classification on three popular datasets show the effectiveness of the proposed method as compared to its previous cousins.
Dingyao Chen, Huibin Tan, Long Lan, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo
ICPR6
2022 Implicit Feature Alignment For Knowledge Distillation
abstract
Knowledge distillation is a technique of transferring knowledge from a large teacher network to a light student one. Existing studies purely use immediate layers' features for distillation and may fail to gain insufficient semantic knowledge from the teacher. Inspired by recent advances in contrastive learning, we propose to introduce extra light embedding layers of the teacher to enforce its generalization ability and further align the mixup-type features for knowledge distillation in an implicit fashion (IFKD). IFKD allows the student to learn richer structural knowledge, thanks to the learned embedding layers of the teacher. Crucially, benefitting from a plethora of mixed samples, we can further adequately mine much semantic knowledge of the teacher. For efficiency, we propose a simple reversed mixup scheme to organize images and implicitly ensure complete positive information comparisons. Extensive experiments on image classification on two popular datasets including CIFAR-100 and ImageNet verify the effectiveness of our approach as compared to the previous methods.
Dingyao Chen, Mengzhu Wang, Xiang Zhang 0008, Tianyi Liang 0001, Zhigang Luo
ICTAI5
2022 Joint Modality Synergy and Spatio-temporal Cue Purification for Moment Localization
abstract
Currently, many approaches to the sentence query based moment location (SQML) task emphasize (inter-)modality interaction between video and language query via transformer-based cross-attention or contrastive learning. However, they could still face two issues: 1) modality interaction could be unexpectedly friendly to modality specific learning that merely learns modality specific patterns, and 2) modality interaction easily confuses spatio-temporal cues and ultimately makes time cues in the original video ambiguous. In this paper, we propose a modality synergy with spatio-temporal cue purification method (MS2P) for SQML to address the above two issues. Particularly, a conceptually simple modality synergy strategy is explored to keep features modality specific while absorbing the other modality complementary information with both carefully designed cross-attention unit and non-contrastive learning. As a result, modality specific semantics can be calibrated progressively in a safer way. To preserve time cues in original video, we further purify video representation into spatial and temporal parts to enhance localization resolution by the proposed two light-weight sentence-aware filtering operations. Experiments on Charades-STA, TACoS, and ActivityNet Caption datasets show our model outperforms the state-of-the-art approaches by a large margin.
Long Lan, Huibin Tan, Xiang Zhang 0008, Xurui Ma, Zhigang Luo
ICMR6
2022 DEAL: An Unsupervised Domain Adaptive Framework for Graph-level Classification
abstract
Graph neural networks (GNNs) have achieved state-of-the-art results on graph classification tasks. They have been primarily studied in cases of supervised end-to-end training, which requires abundant task-specific labels. Unfortunately, annotating labels of graph data could be prohibitively expensive or even impossible in many applications. An effective solution is to incorporate labeled graphs from a different, but related source domain, to develop a graph classification model for the target domain. However, the problem of unsupervised domain adaptation for graph classification is challenging due to potential domain discrepancy in graph space as well as the label scarcity in the target domain. In this paper, we present a novel GNN framework named DEAL by incorporating both source graphs and target graphs, which is featured by two modules, i.e., adversarial perturbation and pseudo-label distilling. Specifically, to overcome domain discrepancy, we equip source graphs with target semantics by applying to them adaptive perturbations which are adversarially trained against a domain discriminator. Additionally, DEAL explores distinct feature spaces at different layers of the GNN encoder, which emphasize global and local semantics respectively. Then, we distill the consistent predictions from two spaces to generate reliable pseudo-labels for sufficiently utilizing unlabeled data, which further improves the performance of graph classification. Extensive experiments on a wide range of graph classification datasets reveal the effectiveness of our proposed DEAL.
Li Shen 0008, Baopu Li, Mengzhu Wang, Xiao Luo 0001, Chong Chen 0002, Zhigang Luo, Xian-Sheng Hua 0001
ACM Multimedia7
2022 Multi-scale local cues and hierarchical attention-based LSTM for stock price trend prediction
Xiao Teng, Xiang Zhang 0008, Zhigang Luo
Neurocomputing3
2022 Online Multiple-Pedestrian Tracking With Detection-Pair-Based Graph Convolutional Networks
abstract
The typical Internet of Things application, unattended driving systems, will need the ability to recognize relevant traffic participants and detect dangerous situations ahead of time. An important component of these systems is one that is able to distinguish pedestrians and track their motion to make intelligent driving decisions. This article develops a high-accuracy multiple pedestrian tracking algorithm which is vital for intelligent transportation. Here, we use the off-the-shelf detectors and explore the benefits of modeling pedestrian interactions, such as the interaction of two pedestrians simultaneously matched to two pedestrians in another frame, for robust detection association. Explicitly studying interactions is nontrivial. Previous works often manually selected interacting detections (or “tracklets”) to simplify the association process. In this article, we propose a novel association method based on deep graph convolutional affinity networks (DGCANs) and extend detection-level interactions to the association-level, which treats a potential association of a detection pair as a node in the graph, and explicitly modeling the interactions among potential associations. Specifically, with the novel node, two corresponding edges are readily designed to model the compatible and colliding interactions between related associations. Our proposed method, by redefining nodes and edges, enables us to blend sufficient interaction cues from appearance and motion and learns a robust affinity measure in an end-to-end fashion. Using the Hungarian algorithm as an online tracker, our method archives state-of-the-art performance on benchmark data sets 2-D MOT15, MOT16, and MOT17.
Weijiang Feng, Long Lan, Michael Buro, Zhigang Luo
IEEE Internet Things J.4
2022 Informative pairs mining based adaptive metric learning for adversarial domain adaptation
Mengzhu Wang, Paul Li, Li Shen 0008, Ye Wang 0023, Shanshan Wang 0008, Wei Wang 0335, Xiang Zhang 0008, Junyang Chen 0001, Zhigang Luo
Neural Networks9
2022 Pythagorean fuzzy inequality derived by operation, equality and aggregation operator
Xindong Peng, Zhigang Luo
Soft Comput.2
2022 Label Propagated Nonnegative Matrix Factorization for Clustering
abstract
Semi-supervised learning (SSL) that utilizes plenty of unlabeled examples to boost the performance of learning from limited labeled examples is a powerful learning paradigm with widely real-world applications such as information retrieval and document clustering. Label propagation (LP) is a popular SSL method which propagates labels through the dataset along high density areas defined by unlabeled examples, but it is fragile to bridge examples. Semi-supervised K-Means uses labeled examples to initialize clustering centers to separate different examples, however, semi-supervised K-Means fails in the situation of imbalanced issues, that is, the example size of each class varies significantly. This paper proposes a novel label propagated nonnegative matrix factorization method (LPNMF) to handle clean labeled but biased data and its extension LPNMF-E to handle noisy labeled data based on the framework of NMF. LPNMF decomposes the whole dataset into the product of a basis matrix and a coefficient matrix. To propagate labels to unlabeled examples, LPNMF regards the class indicators of labeled examples as their coefficients and iteratively updates both basis matrix and coefficients of unlabeled examples. LPNMF absorbs the merits from both semi-supervised K-Means and label propagation to handle their respective shortages. Specifically, on the one hand, LPNMF learns representative clustering centers based on the distribution of the dataset, similar to semi-supervised K-means, and thus is robust to the bridge examples. On the other hand, LPNMF pushes labels according to the affinity between examples, similar to label propagation, and thus relieves the biased problem. Moreover, we introduce a LPNMF extension to handle the noisy label case. LPNMF-E relaxes the constraint of labeled examples. Since the label of each labeled example also obtains label information from the global distribution of the whole dataset and local manifold of its neighbors, LPNMF-E outputs reliable class indicators even if a portion of examples are incorrectly labeled. Theoretical analyses for the generalization ability of our proposed models are also provided. Experimental results on both clean and noisy labeled datasets confirm the effectiveness of LPNMF and LPNMF-E compared with both LP and the representative semi-supervised K-Means algorithms.
Long Lan, Tongliang Liu, Xiang Zhang 0008, Chuanfu Xu, Zhigang Luo
IEEE Trans. Knowl. Data Eng.5
2021 Entity-Aware Biaffine Attention for Constituent Parsing
Xinyi Bai, Xiang Zhang 0008, Zhigang Luo
ICANN (1)5
2021 Joint motion context and clip augmentation for spatio-temporal action detection
abstract
This paper endeavors to leverage spatio-temporal visual cues to improve video-based action detection. As a result, a NOn-Local Action detector based on anchor-free called NOLA is proposed, which is built off a recent moving center detector (MOC) and further extends it by efficiently aggregating long-range spatio-temporal information. In detail, a significantly efficient spatio-temporal motion-aware non-local block is explored to provide global motion contexts for the entire predictive branches of MOC. This byproduct can make the large batch samples run on a resource limited device. Besides, a light-weighted data augmentation method termed clip augmentation designed for video-based tasks is proposed, which serves to improve the generalization ability of the detector with economical scale-and-addition operation. NOLA works with two above schemes in real-time as well. Experiments on two benchmark datasets show that NOLA significantly exceeds MOC. Compared to other existing methods,, NOLA reaches the state-of-the-art, in terms of video-level mean of average precision (video mAP).
Xurui Ma, Xiang Zhang 0008, Chengkun Wu, Chuanfu Xu, Jie Liu 0002, Zhigang Luo
ICMV6
2021 A Focally Discriminative Loss for Unsupervised Domain Adaptation
Dongting Sun, Mengzhu Wang, Xurui Ma, Tianming Zhang, Wei Yu 0029, Zhigang Luo
ICONIP (1)7
2021 Spatio-Temporal Action Detector with Self-Attention
abstract
In the field of spatio-temporal action detection, some current studies attempt to solve the problem of action detection by using the one-stage object detectors based on anchor-free. Albeit efficiency, more performance boosts are expected. Towards this goal, a Self-Attention MovingCenter Detector (SAMOC) is proposed, which is blessed with two attractive aspects: 1) to effectively capture motion cues, a spatio-temporal self-attention block is explored to reinforce feature representation by aggregating motion-dependent global contexts, and 2) a link branch serves to model the frame-level object dependency, which promotes the confidence scores of correct actions. Experiments on two benchmark datasets show that SAMOC with the proposed two aspects achieves the state-of-the-art and works in real-time as well.
Xurui Ma, Zhigang Luo, Xiang Zhang 0008, Qing Liao 0001, Mengzhu Wang
IJCNN2
2021 InterBN: Channel Fusion for Adversarial Unsupervised Domain Adaptation
abstract
A classifier trained on one dataset rarely works on other datasets obtained under different conditions because of domain shifting. Such a problem is usually solved by domain adaptation methods. In this paper, we propose a novel unsupervised domain adaptation (UDA) method based on Interchangeable Batch Normalization (InterBN) to fuse different channels in deep neural networks for adversarial domain adaptation.Specifically, we first observe that the channels with small batch normalization scaling factor have less influence on the whole domain adaption, followed by a theoretical proof that the scaling factors for some channels will definitely come close to zero when imposing a sparsity regularization. Then, we replace the channels that have smaller scaling factors in the source domain with the mean of the channels which have larger scaling factors in the target domain or vice versa. Such a simple but effective channel fusion scheme can drastically increase the domain adaption ability.Extensive experimental results show that our InterBN significantly outperforms the current adversarial domain adaptation methods by a large margin on four visual benchmarks. In particular, InterBN achieves a remarkable improvement of 7.7% over the conditional adversarial adaptation networks (CDAN) on VisDA-2017 benchmark.
Mengzhu Wang, Wei Wang 0335, Baopu Li, Xiang Zhang 0008, Long Lan, Huibin Tan, Tianyi Liang 0001, Wei Yu 0029, Zhigang Luo
ACM Multimedia9
2021 Enhancing the association in multi-object tracking via neighbor graph
abstract
Most modern multi-object tracking (MOT) systems for videos follow the tracking-by-detection paradigm, where objects of interest are first located in each frame then associated correspondingly to form their intact trajectories. In this setting, the appearance features of objects usually provide the most important cues for data association, but it is very susceptible to occlusions, illumination variations, and inaccurate detections, thus easily resulting in incorrect trajectories. To address this issue, in this study we propose to make full use of the neighboring information. Our motivations derive from the observations that people tend to move in a group. As such, when an individual target's appearance is remarkably changed, the observer can still identify it with its neighbor context. To model the contextual information from neighbors, we first utilize the spatiotemporal relations among trajectories to efficiently select suitable neighbors for targets. Subsequently, we construct neighbor graph for each target and corresponding neighbors then employ the graph convolutional networks (GCNs) to model their relations and learn the graph features. To the best of our knowledge, it is the first time to explicitly leverage neighbor cues via GCN in MOT. Finally, standardized evaluations on the MOT16 and MOT17 data sets demonstrate that our approach can remarkably reduce the identity switches whilst achieve state-of-the-art overall performance.
Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Xindong Peng, Zhigang Luo
Int. J. Intell. Syst.5
2021 q-Rung orthopair fuzzy decision-making framework for integrating mobile edge caching scheme preferences
abstract
Mobile edge caching scheme (MECS) can determine where, how, and what to cache on user equipment by employing its own storage. When considering the performance of MECS, it is often full of uncertainty. The q-rung orthopair fuzzy set (q-ROFS), characterized by membership and nonmembership degrees with adjustable parameter q, is quite a high-efficiency way to capture uncertainty. In this paper, first, information measure (entropy, distance measure, and similarity measure)-based area difference under the q-rung orthopair fuzzy (q-ROF) circumstance is studied along with their detailed proofs. Then, we present a comprehensive weight-determination method by combining objective weights (determining by entropy) and subjective weights (given by experts) as combined weights, which can effectually alleviate the unconscionable influence of extreme data on evaluation results and simultaneously reflect objective data and subjective emotion. Moreover, q-ROF score function-based distance measure is presented for dealing with a value comparison problem. Later, q-ROF multicriteria decision-making (MCDM) method called total area based on orthogonal vector (TAOV) is introduced. Moreover, its feasibility is illustrated by MECS selection problem. Finally, a comparison of some existing MCDM methods and the proposed method is constructed for displaying their effectiveness. This proposed method can effectively avoid counterintuitive phenomena, eliminate antilogarithm by negative and zero issue, and has no division by zero issue.
Xindong Peng, Hai-Hui Huang, Zhigang Luo
Int. J. Intell. Syst.3
2021 Semantic-consistent cross-modal hashing for large-scale image retrieval
Xuesong Gu, Guohua Dong, Xiang Zhang 0008, Long Lan, Zhigang Luo
Neurocomputing5
2021 A generic MOT boosting framework by combining cues from SOT, tracklet and re-identification
Tianyi Liang 0001, Long Lan, Xiang Zhang 0008, Zhigang Luo
Knowl. Inf. Syst.4
2021 Learning deep discriminative embeddings via joint rescaled features and log-probability centers
Huayue Cai, Xiang Zhang 0008, Long Lan, Guohua Dong, Chuanfu Xu, Xinwang Liu 0002, Zhigang Luo
Pattern Recognit.7
2021 ChimST: An Efficient Spectral Library Search Tool for Peptide Identification from Chimeric Spectra in Data-Dependent Acquisition
abstract
Accurate and sensitive identification of peptides from MS/MS spectra is a very challenging problem in computational shotgun proteomics. To tackle this problem, spectral library search has been one of the competitive solutions. However, most existing library search tools were developed on the basis of one peptide per spectrum, which prevents them from working properly on chimeric spectra where two or more peptides are co-fragmented. In this work, we present a new library search tool called ChimST, which is particularly capable of reliably identifying multiple peptides from a chimeric spectrum. It starts with associating each query MS/MS spectrum with MS precursor features. For each precursor feature, there is a list of peptide candidates extracted from an input spectral library. Then, it takes one peptide candidate from each associated feature and scores how well they could collectively interpret the query spectrum. The highest-scoring set of peptide candidates are finally reported as the identification of the query spectrum. Our experimental tests show that ChimST could significantly outperform the three state-of-the-art library search tools, SpectraST, reSpect, and MSPLIT, in terms of the numbers of both peptide-spectrum matches and unique peptides, especially when the acquisition isolation window is broad.
Wenju Zhang, Zhewei Liang, Lei Xin, Baozhen Shan, Zhigang Luo, Ming Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.6
2021 Near-Online Multi-Pedestrian Tracking via Combining Multiple Consistent Appearance Cues
abstract
An important cue for multi-pedestrian tracking in video is the consistent appearance of an individual for quite a while. In this paper, we address multi-pedestrian tracking by learning a robust appearance model from the paradigm of tracking by detection. To separate detections of different pedestrians while assembling detections of the same pedestrian, we take advantage of the cue of consistent appearance and exploit three types of evidence from the recent, past and near-future. Existing online approaches only exploit the detection-to-detection and sequence-to-detection metrics, which focus on the recent and past appearance patterns respectively, while the future pedestrian appearance is simply ignored. This drawback is remedied in this paper by further considering the sequence-to-sequence metric, which resorts to near-future appearance presentation. Adaptive combination weights are learned to fuse these three different metrics. Moreover, we propose a novel Focal Triplet Loss to make the model focus more on hard examples than the easy ones. We demonstrate that this can significantly enhance the discriminating power of the model compared with treating every sample equally. Effectiveness and efficiency of the proposed method is verified by conducting comprehensive ablation studies and comparing with many competitive (offline/online/near-online) counterparts on the MOT16 and MOT17 Challenges.
Weijiang Feng, Long Lan, Yong Luo 0002, Yue Yu 0001, Xiang Zhang 0008, Zhigang Luo
IEEE Trans. Circuits Syst. Video Technol.6
2021 Nocal-Siam: Refining Visual Features and Response With Advanced Non-Local Blocks for Real-Time Siamese Tracking
abstract
Siamese trackers contain two core stages, i.e., learning the features of both target and search inputs at first and then calculating response maps via the cross-correlation operation, which can also be used for regression and classification to construct typical one-shot detection tracking framework. Although they have drawn continuous interest from the visual tracking community due to the proper trade-off between accuracy and speed, both stages are easily sensitive to the distracters in search branch, thereby inducing unreliable response positions. To fill this gap, we advance Siamese trackers with two novel non-local blocks named Nocal-Siam, which leverages the long-range dependency property of the non-local attention in a supervised fashion from two aspects. First, a target-aware non-local block (T-Nocal) is proposed for learning the target-guided feature weights, which serve to refine visual features of both target and search branches, and thus effectively suppress noisy distracters. This block reinforces the interplay between both target and search branches in the first stage. Second, we further develop a location-aware non-local block (L-Nocal) to associate multiple response maps, which prevents them inducing diverse candidate target positions in the future coming frame. Experiments on five popular benchmarks show that Nocal-Siam performs favorably against well-behaved counterparts both in quantity and quality.
Huibin Tan, Xiang Zhang 0008, Long Lan, Wenju Zhang, Zhigang Luo
IEEE Trans. Image Process.6
2020 CGD: Multi-View Clustering via Cross-View Graph Diffusion
abstract
Graph based multi-view clustering has been paid great attention by exploring the neighborhood relationship among data points from multiple views. Though achieving great success in various applications, we observe that most of previous methods learn a consensus graph by building certain data representation models, which at least bears the following drawbacks. First, their clustering performance highly depends on the data representation capability of the model. Second, solving these resultant optimization models usually results in high computational complexity. Third, there are often some hyper-parameters in these models need to tune for obtaining the optimal results. In this work, we propose a general, effective and parameter-free method with convergence guarantee to learn a unified graph for multi-view data clustering via cross-view graph diffusion (CGD), which is the first attempt to employ diffusion process for multi-view clustering. The proposed CGD takes the traditional predefined graph matrices of different views as input, and learns an improved graph for each single view via an iterative cross diffusion process by 1) capturing the underlying manifold geometry structure of original data points, and 2) leveraging the complementary information among multiple graphs. The final unified graph used for clustering is obtained by averaging the improved view associated graphs. Extensive experiments on several benchmark datasets are conducted to demonstrate the effectiveness of the proposed method in terms of seven clustering evaluation metrics.
Chang Tang, Xinwang Liu 0002, Xinzhong Zhu, En Zhu, Zhigang Luo, Lizhe Wang 0001, Wen Gao 0001
AAAI5
2020 Robust Normalized Squares Maximization for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) attempts to transfer specific knowledge from one domain with labeled data to another domain without labels. Recently, maximum squares loss has been proposed to tackle UDA problem but it does not consider the prediction diversity which has proven beneficial to UDA. In this paper, we propose a novel normalized squares maximization (NSM) loss in which the maximum squares is normalized by the sum of squares of class sizes. The normalization term enforces the class sizes of predictions to be balanced to explicitly increase the diversity. Theoretical analysis shows that the optimal solution to NSM is one-hot vectors with balanced class sizes, i.e., NSM encourages both discriminate and diverse predictions. We further propose a robust variant of NSM, RNSM, by replacing the square loss with L2,1-norm to reduce the influence of outliers and noises. Experiments of cross-domain image classification on two benchmark datasets illustrate the effectiveness of both NSM and RNSM. RNSM achieves promising performance compared to state-of-the-art methods. The code is available at https://github.com/wj-zhang/NSM.
Wenju Zhang, Xiang Zhang 0008, Qing Liao 0001, Wenjing Yang 0002, Long Lan, Zhigang Luo
CIKM6
2020 Towards Making Unsupervised Graph Hashing Robust
abstract
Unsupervised hashing without supervision easily deteriorates in the case of grossly corrupted data. Motivated by robust optimization, this paper proposes a dual-graph regularized robust hashing (DGRH) based on both manifold smoothness and robust estimators in a more intuitive manner. Orthogonal to existing robust hashing methods, DGRH directly removes the outliers of datasets with M-estimator to exert robustness. In specific, it intends to recover low-rank representation from corrupted data via l1loss while preserving neighborhood relationships among samples with dual-graph regularization. Although DGRH seems a simple extension of robust PCA on graphs with hashing trick, it is easy to implement yet effective. Theory analysis is provided to support our claim. Experiments of image retrieval on three popular benchmark datasets show the efficacy of DGRH as compared to several well-behaved representative counterparts.
Xuesong Gu, Guohua Dong, Xiang Zhang 0008, Long Lan, Zhigang Luo
ICME5
2020 Learning sequence-to-sequence affinity metric for near-online multi-object tracking
Weijiang Feng, Long Lan, Xiang Zhang 0008, Zhigang Luo
Knowl. Inf. Syst.4
2020 Enhancing unsupervised domain adaptation by discriminative relevance regularization
Wenju Zhang, Xiang Zhang 0008, Long Lan, Zhigang Luo
Knowl. Inf. Syst.4
2020 Object-aware semantics of attention for image captioning
Long Lan, Xiang Zhang 0008, Guohua Dong, Zhigang Luo
Multim. Tools Appl.5
2020 GateCap: Gated spatial and semantic attention model for image captioning
Long Lan, Xiang Zhang 0008, Zhigang Luo
Multim. Tools Appl.4
2020 Maximum Mean and Covariance Discrepancy for Unsupervised Domain Adaptation
Wenju Zhang, Xiang Zhang 0008, Long Lan, Zhigang Luo
Neural Process. Lett.4
2019 Attentional Residual Dense Factorized Network for Real-Time Semantic Segmentation
Long Lan, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICANN (3)5
2019 Person re-identification via adaptive verification loss
Hui Tian 0005, Xiang Zhang 0008, Long Lan, Zhigang Luo
Neurocomputing4
2019 Label guided correlation hashing for large-scale cross-modal retrieval
Guohua Dong, Xiang Zhang 0008, Long Lan, Zhigang Luo
Multim. Tools Appl.5
2019 Stacked Marginal Time Warping for Temporal Alignment
Xiang Zhang 0008, Liquan Nie, Long Lan, Xuhui Huang, Zhigang Luo
Neural Process. Lett.5
2019 Nonnegative Constrained Graph Based Canonical Correlation Analysis for Multi-view Feature Learning
Huibin Tan, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
Neural Process. Lett.5
2018 Margin-Embedding Canonical Correlation Analysis with Feature Selection for Person Re-Identification
abstract
Canonical correlation analysis (CCA) is a classical subspace learning method of capturing the common semantic information underlying multi-view data. It has been used in person re-identification (re-ID) task by treating the task of matching identical individuals across non-overlapping multi-cameras as a multi-view learning problem. However, CCA-based reID methods still achieve unsatisfactory results because few jointly consider discriminative margin information and selecting importantly relevant features. To address this issue, we propose a novel l2,1-norm regularized margin-embedding CCA ( l2,1-MCCA), which learns a generalized discriminative subspace by employing more discriminative margin information. Moreover, the new method enforces the l2,1-norm regularization term over the learned subspace to identify the relevant features. Both lightweight and effective schemes can benefit from each other and endeavor to enlarge the interclass variations whilst reducing the intra-class variations. Experiments on three popular datasets show the efficacy of l2,1-MCCA as compared with recently representative re-ID methods.
Linfei Ma, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
ICIP5
2018 Graph-Laplacian Correlated Low-Rank Representation for Subspace Clustering
abstract
Subspace clustering seeks to segment a given unlabeled data into clusters with the hope of each cluster corresponding to a union of low-dimensional subspaces. Among them, low-rank representation (LRR) is a promising potential method which intends to build a good affinity matrix by using the self-expression of inputs. However, it completely ignores the important data locality. Although several works in this regard have considered the local geometric structure through the Laplacian regularizer, they also neglect the correlation of the data. In this paper, we propose a graph-Laplacian correlated low-rank representation model (GCLRR) to address such an issue. Particularly, GCLRR factorizes the self-expression as the product of two low-dimensional matrices, of which one is the latent representation of the self-expression. On the basis of the latent representation, the Laplacian regularizer is integrated with the orthogonal constraint together and behaves like the spectral clustering. Moreover, we devise a Frobenius norm based trace loss and use it to constrain both the latent representation and the self-expression to capture the correlation of the data. Our improved trace loss is more efficient than the original one. More importantly, GCLRR provides an effective unified framework to seamlessly integrate both aspects above. Then, we optimize GCLRR in the frame of alternating direction method (ADM) and fortunately derive the analytical solution to each subproblem. Experiments of motion segmentation and image clustering confirm the efficacy of the proposed GCLRR.
Huayue Cai, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
ICIP6
2018 Discrete Graph Hashing via Affine Transformation
abstract
In unsupervised graph-based hashing for large-scale image retrieval, many efforts have been made to bridge the gap between the learned graph embedding and the corresponding binary codes. Relatively, few studies focus on the issue of the discrimination of graph embedding. In this paper, we firstly devise a discrete graph hashing model (DGH) that smooths graph embedding and simultaneously solving binary codes under the balanced discrete constraint, which equals a novel method of jointly learning graph embedding and spectral rotation, theoretically. To further induce discriminant graph embedding, we substitute affine transformation for spectral rotation in our DGH (abbreviated as ADGH). This is because affine transformation can accommodate both rotational angle and distance of graph embedding, while respecting the neighborhood structure among most samples. Besides, each subproblem of ADGH can yield the closed-form solution. Experiments of image retrieval on three benchmark datasets show that ADGH outperforms the representative hashing methods in quantity.
Guohua Dong, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
ICME5
2018 Cross-Layer Convolutional Siamese Network for Visual Tracking
Yanyin Chen, Huibin Tan, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
ICONIP (2)7
2018 Background Subtraction via 3D Convolutional Neural Networks
abstract
Background subtraction can be treated as the binary classification problem of highlighting the foreground region in a video whilst masking the background region, and has been broadly applied in various vision tasks such as video surveillance and traffic monitoring. However, it still remains a challenging task due to complex scenes and for lack of the prior knowledge about the temporal information. In this paper, we propose a novel background subtraction model based on 3D convolutional neural networks (3D CNNs) which combines temporal and spatial information to effectively separate the foreground from all the sequences in an end-to-end manner. Different from conventional models, we view background subtraction as three-class classification problem, i.e., the foreground, the background and the boundary. This design can obtain more reasonable results than existing baseline models. Experiments on the Change Detection 2012 dataset verify the potential of our model in both quantity and quality.
Yongqiang Gao, Huayue Cai, Xiang Zhang 0008, Long Lan, Zhigang Luo
ICPR5
2018 Box-constrained Discriminant Projective Non-negative Matrix Factorization through Augmented Lagrangian Multiplier Method
abstract
Projective non-negative matrix factorization (PNMF) learns a non-negative projection matrix to project high-dimensional examples onto a lower-dimensional space spanned by the transpose of the learned projection matrix. Since PNMF can learn parts-based representation, it has attracted ample attention from computer vision community. However, existing PNMF methods either completely ignore labels of the dataset or endure the slow convergent optimization algorithm. In this paper, we propose a box-constrained discriminant PNMF (BDPNMF) method to address these issues. Specifically, BDPNMF jointly exploits the Fisher's criterion and the augmented Lagrangian multiplier (ALM) method into PNMF to boost discriminative capacity of the learned subspace and its efficiency. Experimental results on four popular face image datasets confirm the efficacy of BDPNMF compared to previous PNMF methods in quantity.
Huayue Cai, Xiang Zhang 0008, Zhigang Luo, Xuhui Huang
IJCNN3
2018 Multi-granularity Hierarchical Attention Siamese Network for Visual Tracking
abstract
Speed and accuracy are the two most important focuses for many visual tracking methods. Recently, siamese networks based trackers have shown very promising potentials in both aspects, which develop a twin network to measure the responses between target and hypotheses with a fully convolutional operation. However, the learned response maps are vulnerable to background clutters and scale changes as they ignore priori knowledge such as the object salience and multi-granularity cues. To explore the benefits of priori, this paper devises a multi-granularity hierarchical attention siamese network tracker (MHA-Siam) to further enhance the tracking stability without sacrificing real-time speed. Particularly, the channel-wise attention mechanism is exploited here to filter out the background while remain the salient object region; then, the response maps of the coarse-to-finer multi-layer features are fused to capture multi-granularity location information helpful for improvement in tracking stability. To make full use of them, MHA-Siam imposes the element-wise max-and-sum operation on them to induce a reliable response map for accurate location. Experiments of visual tracking on OTB benchmark shows the superiority of MHA-Siam with the competitive efficiency to its counterpart trackers.
Xiang Zhang 0008, Huibin Tan, Long Lan, Zhigang Luo, Xuhui Huang
IJCNN5
2018 Ranking-Embedded Transfer Canonical Correlation Analysis for Person Re-Identification
abstract
Person re-identification (re-ID) seeks to match the identical individuals across different cameras and is still a challenging visual task due to substantial variances of person appearance in complex scenarios. Different from most of conventional person re-ID methods, which generally reduce person re-ID task to either a multi-view learning problem or a multi- domain learning problem alone, this paper treats such a task as a multi-view multi-domain (MVMD) learning problem to exploit the both benefits by refreshing canonical correlation analysis (CCA) with two improvements, termed as ranking-embedded transfer CCA (RTCCA). Specifically, to bridge the semantic gap between different views, we first embed a ranking weight matrix into CCA to strength the correlations among the multi-view images of the same identity and simultaneously to weaken that of different identities. Furthermore, we utilize the well-known distribution metric maximum mean discrepancy (MMD) as a regularization term to reduce the domain shift between training set and testing set. More importantly, the two improvements benefit from each other and the joint merit can further boost the re-ID performance. Experiments on three benchmarks verify the efficacy of the proposed RTCCA when compared with the recently representative baseline person re-ID methods.
Linfei Ma, Xiang Zhang 0008, Long Lan, Xuhui Huang, Zhigang Luo
IJCNN5
2018 Collaborative Subspace Graph Hashing for Cross-modal Retrieval
abstract
Current hashing methods for cross-modal retrieval generally attempt to learn the separate modality-specific transformation matrices to embed multi-modality data into a latent common subspace, and usually ignore the fact that respecting the diversity of multi-modality features in the latent subspace could be beneficial for retrieval improvements. To this, we propose a collaborative subspace graph hashing method (CSGH) to perform a two-stage collaborative learning framework for cross-modal retrieval. Particularly, CSGH first embeds multi-modality data into separate latent subspaces through individual modality-specific transformation matrices, and then connects these latent subspaces to a common Hamming space through a shared transformation matrix. In this framework, CSGH considers the modality-specific neighborhood structure and the cross-modal correlation within multi-modality data through the Laplacian regularization and the graph based correlation constraint, respectively. To solve CSGH, we develop an alternative procedure to optimize it, and fortunately, each sub-problem of CSGH has the elegant analytical solution. Experiments of cross-modal retrieval on Wiki, NUS-WIDE, Flickr25K and Flickr1M datasets show the effectiveness of CSGH compared with the state-of-the-art cross-modal hashing methods.
Xiang Zhang 0008, Guohua Dong, Yimo Du, Chengkun Wu, Zhigang Luo, Canqun Yang
ICMR5
2018 Low-Rank Matrix Recovery via Continuation-Based Approximate Low-Rank Minimization
Xiang Zhang 0008, Yongqiang Gao, Long Lan, Xuhui Huang, Zhigang Luo
PRICAI (1)6
2017 Attention Focused Spatial Pyramid Pooling for Boxless Action Recognition in Still Images
Weijiang Feng, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICANN (2)4
2017 GNMF Revisited: Joint Robust k-NN Graph and Reconstruction-Based Graph Regularization for Image Clustering
Wenju Zhang, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICANN (2)6
2017 Unsupervised domain adaptation with joint supervised sparse coding and discriminative regularization term
abstract
Domain adaptation (DA) attempts to enhance the generalization capability of classifier through narrowing the gap of the distributions across domains. This paper focuses on unsupervised domain adaptation where labels are not available in target domain. Most existing approaches explore the domain-invariant features shared by domains but ignore the discriminative information of source domain. To address this issue, we propose a discriminative domain adaptation method (DDA) to reduce domain shift by seeking a common latent subspace jointly using supervised sparse coding (SSC) and discriminative regularization term. Particularly, DDA adapts SSC to yield discriminative coefficients of target data and further unites with discriminative regularization term to induce a common latent subspace across domains. We show that both strategies can boost the ability of transferring knowledge from source to target domain. Experiments on two real world datasets demonstrate the effectiveness of our proposed method over several existing state-of-the-art domain adaptation methods.
Xiang Zhang 0008, Wenju Zhang, Xuhui Huang, Naiyang Guan, Zhigang Luo
ICIP6
2017 Boxless Action Recognition in Still Images via Recurrent Visual Attention
Weijiang Feng, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICONIP (2)4
2017 Audio visual speech recognition with multimodal recurrent neural networks
abstract
Studies on nowadays human-machine interface have demonstrated that visual information can enhance speech recognition accuracy especially in noisy environments. Deep learning has been widely used to tackle such audio visual speech recognition (AVSR) problem due to its astonishing achievements in both speech recognition and image recognition. Although existing deep learning models succeed to incorporate visual information into speech recognition, none of them simultaneously considers sequential characteristics of both audio and visual modalities. To overcome this deficiency, we proposed a multimodal recurrent neural network (multimodal RNN) model to take into account the sequential characteristics of both audio and visual modalities for AVSR. In particular, multimodal RNN includes three components, i.e., audio part, visual part, and fusion part, where the audio part and visual part capture the sequential characteristics of audio and visual modalities, respectively, and the fusion part combines the outputs of both modalities. Here we modelled the audio modality by using a LSTM RNN, and modelled the visual modality by using a convolutional neural network (CNN) plus a LSTM RNN, and combined both models by a multimodal layer in the fusion part. We validated the effectiveness of the proposed multimodal RNN model on a multi-speaker AVSR benchmark dataset termed AVletters. The experimental results show the performance improvements comparing to the known highest audio visual recognition accuracies on AVletters, and confirm the robustness of our multimodal RNN model.
Weijiang Feng, Naiyang Guan, Yuan Li 0007, Xiang Zhang 0008, Zhigang Luo
IJCNN5
2016 Online Multi-Object Tracking by Quadratic Pseudo-Boolean Optimization
Long Lan, Dacheng Tao, Chen Gong 0002, Naiyang Guan, Zhigang Luo
IJCAI5
2016 Enhancing temporal alignment with autoencoder regularization
abstract
Temporal alignment aligns two temporal sequences and is quite challenging due to drastic differences among temporal sequences and source data from different views. Canonical time warping (CTW) has shown great potential in temporal alignment tasks because it can reduce data redundancy by transforming high-dimensional data to a lower-dimensional subspace via canonical correlation analysis (CCA). However, CTW cannot uncover the underlying nonlinear structure embedded in the dataset. In this paper, we propose an autoencoder regularized canonical time warping method (AECTW) to overcome this drawback. Specifically, AECTW enhances lower-dimensional representation of each sequence by incorporating an autoencoder regularization, meanwhile reveals the nonlinear structure of features by explicit nonlinear transformation. By these strategies, AECTW significantly boosts CTW in temporal alignment tasks. Experiments on both synthetic data and two practical human action datasets demonstrate that AECTW outperforms the representative DTW-based methods.
Liquan Nie, Yuanyuan Wang 0004, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
IJCNN5
2016 Distributed graph regularized non-negative matrix factorization with greedy coordinate descent
abstract
Graph regularized non-negative matrix factorization (GNMF) decomposes a high-dimensional non-negative data matrix into two low-dimensional matrices with the non-negativity property kept and the geometric structure preserved. Due to its effectiveness, GNMF has been widely used in many fields such as computer vision and data mining. However, GNMF cannot process large-scale datasets on distributed system because the gradient of the graph regularization term costs huge amount of communication overheads among computing nodes. In this paper, we proposed a distributed GNMF (DGNMF) algorithm to overcome this deficiency. Particularly, DGNMF reformulates the graph regularization term to avoid multiplying graph Laplacian by factor matrix through introducing an auxiliary variable and incorporating an equality constraint over it. We optimize DGNMF by using greedy coordinate descent method in the frame of augmented Lagrange method and implement this algorithm on a distributed system. Since DGNMF requires quite few communication overheads among computing nodes, it can be applied to large scale dataset. The preliminary results illustrate efficiency, scalability, and effectiveness of DGNMF.
Ziheng Gao, Naiyang Guan, Xuhui Huang, Xuefeng Peng, Zhigang Luo, Yuhua Tang
SMC5
2016 Correntropy induced metric based graph regularized non-negative matrix factorization
Yuanyuan Wang 0004, Shuyi Wu, Bin Mao, Xiang Zhang 0008, Zhigang Luo
Neurocomputing5
2015 Labelwalking nonnegative matrix factorization
abstract
Semi-supervised learning (SSL) utilizes plenty of unlabeled examples to boost the performance of learning from limited labeled examples. Due to its great discriminant power, SSL has been widely applied to various real-world tasks such as information retrieval, pattern recognition, and speech separa- tion. Label propagation (LP) is a popular SSL method which propagates labels through the dataset along high density areas defined by unlabeled examples, LP assumes nearby examples should share the same label, thus, it unavoidably pushes the labels to the wrong examples, especially when different la- beled examples are not strictly separated. Seed K-means uses labeled examples to initialize class centers, and avoid getting stuck in poor local optima comparing to traditional K-means, however the hard constraint of each example's membership makes Seed K-means failed in many real world applications. This paper proposes a novel label walking nonnegative matrix factorization method (LWNMF) to handle labeled examples in SSL based on the framework of NMF. LWNMF decomposes the whole dataset into the product of a basis matrix and a coefficient matrix, and to travel labels to unlabeled examples, LWNMF regards the class indicators of labeled examples as their coefficients and iteratively updates both basis matrix and coefficients of unlabeled examples. Since LWNMF learns comprehensive class centroids, labels iteratively walk to unlabeled examples through these significant centroids.
Long Lan, Naiyang Guan, Xiang Zhang 0008, Xuhui Huang, Zhigang Luo
ICASSP5
2015 Constrained Projective Non-negative Matrix Factorization for Semi-supervised Multi-label Learning
abstract
This paper formulates multi-label learning as a constrained projective non-negative matrix factorization (CPNMF) problem which concentrates on a variant of the original projective NMF (PNMF) and explicitly introduces an auxiliary basis to learn the semantic subspace and boosts its discriminating ability by exploiting labeled and unlabeled examples together. Particularly, it propagates labels of the labeled examples to the unlabeled ones by enforcing coefficients of examples sharing identical semantic contents to be identical based on a hard constraint, i.e., embedding the class indicator of labeled examples into their coefficients. CPNMF preserves the geometrical structure of dataset via manifold regularization meanwhile captures the inherent structure of labels by using label correlations. We developed a multiplicative update rule (MUR) based algorithm to optimize CPNMF and proved its convergence. Experiments of image annotation on Corel dataset, text categorization on Rcv1v2 dataset, and text clustering on two popular text corpuses suggest the effectiveness of CPNMF.
Xiang Zhang 0008, Naiyang Guan, Zhigang Luo, Xuejun Yang
ICMLA3
2015 Semi-supervised Non-negative Local Coordinate Factorization
Cherong Zhou, Xiang Zhang 0008, Naiyang Guan, Xuhui Huang, Zhigang Luo
ICONIP (2)5
2015 New insights into the landscape relationships of host response to bacterial pathogens
abstract
Modern understanding of microbiology largely lays foundation in the biological characterization of microorganisms. However, the landscape relationships of host transcriptional response (HTR) to different bacterial pathogens have not yet been systematically explored. Here, we established the first generation of HTR network (HTRN) according to the HTR similarities among 21 different human pathogenic bacterial species by integrating 258 pairs of host cellular gene expression profiles upon infections. Further, the network was dissected into five bacterial communities of more consensus internal HTR. Interestingly, analysis of signature genes across different communities revealed that distinct community signatures (CS) present differential gene expression patterns. Functional annotation suggested a common feature of host cell response to bacterial infections that specific functional gene clusters (BPs and/or signaling pathways) were preferentially elicited or subverted by community bacterial pathogens. Notably, community signatures (especially key associators participating dissimilar functional profiles) were highly enriched of GWAS disease-related genes, which associated bacterial infections with common and specific non-infectious human disease(s). About 40% of the associations were confirmed by literature investigation that further indicated possible/potential association directionality. Our characterization and analysis were the first to feature differential community HTRs upon bacterial pathogen infections and suggested new perspective of understanding infection-disease associations and underlying pathogenesis.
Xiaoyao Yin, Xiaochen Bo, Cong Niu, Naiyang Guan, Zhigang Luo
IJCNN8
2015 Correntropy supervised non-negative matrix factorization
abstract
Non-negative matrix factorization (NMF) is a powerful dimension reduction method and has been widely used in many pattern recognition and computer vision problems. However, conventional NMF methods are neither robust enough as their loss functions are sensitive to outliers, nor discriminative because they completely ignore labels in a dataset. In this paper, we proposed a correntropy supervised NMF (CSNMF) to simultaneously overcome aforementioned deficiencies. In particular, CSNMF maximizes the correntropy between the data matrix and its reconstruction in low-dimensional space to inhibit outliers during learning the subspace, and narrows the minimizes the distances between coefficients of any two samples with the same class labels to enhance the subsequent classification performance. To solve CSNMF, we developed a multiplicative update rules and theoretically proved its convergence. Experimental results on popular face image datasets verify the effectiveness of CSNMF comparing with NMF, its supervised variants, and its robustified variants.
Wenju Zhang, Naiyang Guan, Dacheng Tao, Bin Mao, Xuhui Huang, Zhigang Luo
IJCNN6
2015 Two-Dimensional Euler PCA for Face Recognition
Huibin Tan, Xiang Zhang 0008, Naiyang Guan, Dacheng Tao, Xuhui Huang, Zhigang Luo
MMM (2)6
2015 Non-negative Low-Rank and Group-Sparse Matrix Factorization
Shuyi Wu, Xiang Zhang 0008, Naiyang Guan, Dacheng Tao, Xuhui Huang, Zhigang Luo
MMM (2)6
2015 Constraint-Relaxation Approach for Nonnegative Matrix Factorization: A Case Study
abstract
Nonnegative matrix factorization (NMF) is a powerful technique for dimensionality reduction. Conventional NMF algorithms usually keep the matrices W and H nonnegative while iterating. However, to get the NMF of a matrix, it's unnecessary to force the temporary solutions in iterations nonnegative. In this paper, we propose a two-staged approach for NMF. At the relaxation stage, the nonnegative constraint of temporary solutions is relaxed and a real valued matrix factorization is generated. At the constraint stage, the real valued matrix factorization is transformed to a nonnegative matrix factorization by an invertible linear transformation. Based on this approach, we study on exact nonnegative matrix factorization when rank=2. We proved that, given two real valued matrices of rank=2, there exists an invertible linear transformation which can transform the real valued matrices to nonnegative matrices with their product stable. We propose an algorithm to find out the transformation. When rank is higher than 2, this kind of transformation may not exist. In the experiments, it's showed that this approach can reach a nonnegative matrix factorization with lower reconstruction error than conventional methods, and the technique for rank=2 exact NMF works well.
Naiyang Guan, Xuhui Huang, Zhigang Luo
SMC4
2015 Symmetric Non-negative Matrix Factorization Based Link Partition Method for Overlapping Community Detection
abstract
Partitioning links rather than nodes is effective in overlapping community detection (OCD) on complex networks. However, it consumes high CPU and memory overheads because the volume of links is huge especially when the network is rather complex. In this paper, we proposes a symmetric non-negative matrix factorization (SNMF) based link partition method called SNMF-Link to overcome this deficiency. In particular, SNMF-Link represents data in a lower-dimensional space spanned by the node-link incidence matrix. By solving a lighter SNMF problem, SNMF-Link learns the clustering indicators of each links. Since traditional multiplicative update rule (MUR) based optimization algorithm for SNMF suffers from slow convergence, we applied the augmented Lagrangian method (ALM) to efficiently optimize SNMF. Experimental results show that SNMF-Link is much more efficient than the representative clustering algorithms without reducing the OCD performance.
Xiang Zhang 0008, Naiyang Guan, Wenju Zhang, Xuhui Huang, Shuyi Wu, Zhigang Luo
SMC6
2014 Transductive nonnegative matrix factorization for semi-supervised high-performance speech separation
abstract
Regarding the non-negativity property of the magnitude spectrogram of speech signals, nonnegative matrix factorization (NMF) has obtained promising performance for speech separation by independently learning a dictionary on the speech signals of each known speaker. However, traditional NM-F fails to represent the mixture signals accurately because the dictionaries for speakers are learned in the absence of mixture signals. In this paper, we propose a new transductive NMF algorithm (TNMF) to jointly learn a dictionary on both speech signals of each speaker and the mixture signals to be separated. Since TNMF learns a more descriptive dictionary by encoding the mixture signals than that learned by NMF, it significantly boosts the separation performance. Experiments results on a popular TIMIT dataset show that the proposed TNMF-based methods outperform traditional NMF-based methods for separating the monophonic mixtures of speech signals of known speakers.
Naiyang Guan, Long Lan, Dacheng Tao, Zhigang Luo, Xuejun Yang
ICASSP4
2014 Soft-constrained nonnegative matrix factorization via normalization
abstract
Semi-supervised clustering aims at boosting the clustering performance on unlabeled samples by using labels from a few labeled samples. Constrained NMF (CNMF) is one of the most significant semi-supervised clustering methods, and it factorizes the whole dataset by NMF and constrains those labeled samples from the same class to have identical encodings. In this paper, we propose a novel soft-constrained NMF (SCNMF) method by softening the hard constraint in CNMF. Particularly, SCNMF factorizes the whole dataset into two lower-dimensional factor matrices by using multiplicative update rule (MUR). To utilize the labels of labeled samples, SCNMF iteratively normalizes both factor matrices after updating them with MURs to make encodings of labeled samples close to their label vectors. It is therefore reasonable to believe that encodings of unlabeled samples are also close to their corresponding label vectors. Such strategy significantly boosts the clustering performance even when the labeled samples are rather limited, e.g., each class owns only a single labeled sample. Since the normalization procedure never increases the computational complexity of MUR, SCNMF is quite efficient and effective in practices. Experimental results on face image datasets illustrate both efficiency and effectiveness of SCNMF compared with both NMF and CNMF.
Long Lan, Naiyang Guan, Xiang Zhang 0008, Dacheng Tao, Zhigang Luo
IJCNN5
2014 Box-constrained projective nonnegative matrix factorization via augmented Lagrangian method
abstract
Projective non-negative matrix factorization (P-NMF) projects a set of examples onto a subspace spanned by a non-negative basis whose transpose is regarded as the projection matrix. Since PNMF learns a natural parts-based representation, it has been successfully used in text mining and pattern recognition. However, it is non-trivial to analyze the convergence of the optimization algorithms for PNMF because its objective function is non-convex. In this paper, we propose a Box-constrained PNMF (BPNMF) method to overcome this deficiency of PNMF. In particular, BPNMF introduces an auxiliary variable, i.e., the coefficients of examples, and incorporates the following two types of constraints: 1) each entry of the basis is non-negative and upper-bounded, i.e., box-constrained, and 2) the coefficients equal to the projected points of the examples. The first box constraint makes the basis to be bound and the second equality constraint keeps its equivalence to PNMF. Similar to PNMF, BPNMF is difficult because the objective function is non-convex. To solve BPNMF, we developed an efficient algorithm in the frame of augmented Lagrangian multiplier (ALM) method and proved that the ALM-based algorithm converges to local minima. Experimental results on two face image datasets demonstrate the effectiveness of BPNMF compared with the representative methods.
Xiang Zhang 0008, Naiyang Guan, Long Lan, Dacheng Tao, Zhigang Luo
IJCNN5
2014 Translation non-negative matrix factorization with fast optimization
abstract
Non-negative matrix factorization (NMF) reconstructs the original samples in a lower dimensional space and has been widely used in pattern recognition and data mining because it usually yields sparse representation. Since NMF leads to unsatisfactory reconstruction for the datasets that contain translations of large magnitude, it is required to develop translation NMF (TNMF) to first remove the translation and then conduct a decomposition. However, existing multiplicative update rule based algorithm for TNMF is not efficient enough. In this paper, we reformulate TNMF and show that it can be efficiently solved by using the state-of-the-art solvers such as NeNMF. Experimental results on face image datasets confirm both efficiency and effectiveness of the reformulated TNMF.
Yuanyuan Wang 0004, Naiyang Guan, Bin Mao, Xuhui Huang, Zhigang Luo
SMC5
2014 Merge-Weighted Dynamic Time Warping for Speech Recognition
Xianglilan Zhang, Zhigang Luo, Ming Li 0001
J. Comput. Sci. Technol.2
2013 Confidence index dynamic time warping for language-independent embedded speech recognition
abstract
Language-independent embedded speech recognition is a necessary and important application. Considering personal privacy, collection difficulty of all the reference words, and limited storage space of mobile devices, language-independent (LI) embedded speech recognition should be classified into lightweight speaker-dependent (SD) cases. Dynamic time warping (DTW) is the state-of-the-art algorithm for small foot-print SD automatic speech recognition. To decrease the high computational complexity of DTW, and to avoid constraints-induced coarse approximation and inaccuracy problems, we introduce a novel confidence index dynamic time warping (CIDTW) approach. CIDTW defines a new cost function, called the confidence index cost function (CICF), to measure the similarity between merged speech training and testing data, while follows the same DTW process. With extensive experiments on three representative SD datasets, CIDTW achieves better accuracy and overall six times faster speeds compared with DTW.
Xianglilan Zhang, Jiping Sun, Zhigang Luo, Ming Li 0001
ICASSP3
2013 Orthogonal Nonnegative Locally Linear Embedding
abstract
Nonnegative matrix factorization (NMF) decomposes a nonnegative dataset X into two low-rank nonnegative factor matrices, i.e., W and H, by minimizing either Kullback-Leibler (KL) divergence or Euclidean distance between X and WH. NMF has been widely used in pattern recognition, data mining and computer vision because the non-negativity constraints on both W and H usually yield intuitive parts-based representation. However, NMF suffers from two problems: 1) it ignores geometric structure of dataset, and 2) it does not explicitly guarantee parts-based representation on any datasets. In this paper, we propose an orthogonal nonnegative locally linear embedding (ONLLE) method to overcome aforementioned problems. ONLLE assumes that each example embeds in its nearest neighbors and keeps such relationship in the learned subspace to preserve geometric structure of a dataset. For the purpose of learning parts-based representation, ONLLE explicitly incorporates an orthogonality constraint on the learned basis to keep its spatial locality. To optimize ONLLE, we applied an efficient fast gradient descent (FGD) method on Stiefel manifold which accelerates the popular multiplicative update rule (MUR). The experimental results on real-world datasets show that FGD converges much faster than MUR. To evaluate the effectiveness of ONLLE, we conduct both face recognition and image clustering on real-world datasets by comparing with the representative NMF methods.
Lei Weit, Naiyang Guan, Xiang Zhang 0008, Zhigang Luo, Dacheng Tao
SMC4
2012 Graph Based Semi-supervised Non-negative Matrix Factorization for Document Clustering
abstract
Non-negative matrix factorization (NMF) approximates a non-negative matrix by the product of two low-rank matrices and achieves good performance in clustering. Recently, semi-supervised NMF (SS-NMF) further improves the performance by incorporating part of the labels of few samples into NMF. In this paper, we proposed a novel graph based SS-NMF (GSS-NMF). For each sample, GSS-NMF minimizes its distances to the same labeled samples and maximizes the distances against different labeled samples to incorporate the discriminative information. Since both labeled and unlabeled samples are embedded in the same reduced dimensional space, the discriminative information from the labeled samples is successfully transferred to the unlabeled samples, and thus it greatly improves the clustering performance. Since the traditional multiplicative update rule converges slowly, we applied the well-known projected gradient method to optimizing GSS-NMF and the proposed algorithm can be applied to optimizing other manifold regularized NMF efficiently. Experimental results on two popular document datasets, i.e., Reuters21578 and TDT-2, show that GSS-NMF outperforms the representative SS-NMF algorithms.
Naiyang Guan, Xuhui Huang, Long Lan, Zhigang Luo, Xiang Zhang 0008
ICMLA (1)4
2012 Sparse Representation Based Discriminative Canonical Correlation Analysis for Face Recognition
abstract
Canonical correlation analysis (CCA) has been widely used in pattern recognition and machine learning. However, both CCA and its extensions sometimes cannot give satisfactory results. In this paper, we propose a new CCA-type method termed sparse representation based discriminative CCA (SPDCCA) by incorporating sparse representation and discriminative information simultaneously into traditional CCA. In particular, SPDCCA not only preserves the sparse reconstruction relationship within data based on sparse representation, but also preserves the maximum-margin based discriminative information, and thus it further enhances the classification performance. Experimental results on Yale, Extended Yale B, and ORL datasets show that SPDCCA outperforms both CCA and its extensions including KCCA, LPCCA and LDCCA in face recognition.
Naiyang Guan, Xiang Zhang 0008, Zhigang Luo, Long Lan
ICMLA (1)3
2012 Semi-supervised Non-negative Patch Alignment Framework
abstract
Non-negative matrix factorization (NMF) learns the latent semantic space more direct and reliable than the latent semantic indexing (LSI) and the spectral clustering methods, thus performs well in document clustering. Recently, semi-supervised NMF such as N2S2L, CNMF and unsupervised method such as GNMF significantly improve the face recognition performance, but they are designed for classification. In this paper, we combine both geometric structure and label information with NMF under the non-negative patch alignment framework (NPAF) to form SS-NPAF. Due to this combination, it greatly improves the clustering performance. To optimize SS-NPAF, we apply the well-known projected gradient method to overcome the slow convergence problem of the mostly used multiplicative update rule. Experimental results on two popular document datasets, i.e., Reuters21578 and TDT-2, show that SS-NPAF outperforms the representative SS-NMF algorithms.
Long Lan, Xuhui Huang, Naiyang Guan, Zhigang Luo, Xiang Zhang 0008
ICMLA (1)4
2012 Classifying Stem Cell Differentiation Images by Information Distance
Xianglilan Zhang, Hongnan Wang, Tony J. Collins, Zhigang Luo, Ming Li 0001
ECML/PKDD (1)4
2012 Online Nonnegative Matrix Factorization With Robust Stochastic Approximation
abstract
Nonnegative matrix factorization (NMF) has become a popular dimension-reduction method and has been widely applied to image processing and pattern recognition problems. However, conventional NMF learning methods require the entire dataset to reside in the memory and thus cannot be applied to large-scale or streaming datasets. In this paper, we propose an efficient online RSA-NMF algorithm (OR-NMF) that learns NMF in an incremental fashion and thus solves this problem. In particular, OR-NMF receives one sample or a chunk of samples per step and updates the bases via robust stochastic approximation. Benefitting from the smartly chosen learning rate and averaging technique, OR-NMF converges at the rate of in each update of the bases. Furthermore, we prove that OR-NMF almost surely converges to a local optimal solution by using the quasi-martingale. By using a buffering strategy, we keep both the time and space complexities of one step of the OR-NMF constant and make OR-NMF suitable for large-scale or streaming datasets. Preliminary experimental results on real-world datasets show that OR-NMF outperforms the existing online NMF (ONMF) algorithms in terms of efficiency. Experimental results of face recognition and image annotation on public datasets confirm the effectiveness of OR-NMF compared with the existing ONMF algorithms.
Naiyang Guan, Dacheng Tao, Zhigang Luo, Bo Yuan 0003
IEEE Trans. Neural Networks Learn. Syst.3
2011 Manifold Regularized Discriminative Nonnegative Matrix Factorization With Fast Gradient Descent
abstract
Nonnegative matrix factorization (NMF) has become a popular data-representation method and has been widely used in image processing and pattern-recognition problems. This is because the learned bases can be interpreted as a natural parts-based representation of data and this interpretation is consistent with the psychological intuition of combining parts to form a whole. For practical classification tasks, however, NMF ignores both the local geometry of data and the discriminative information of different classes. In addition, existing research results show that the learned basis is unnecessarily parts-based because there is neither explicit nor implicit constraint to ensure the representation parts-based. In this paper, we introduce the manifold regularization and the margin maximization to NMF and obtain the manifold regularized discriminative NMF (MD-NMF) to overcome the aforementioned problems. The multiplicative update rule (MUR) can be applied to optimizing MD-NMF, but it converges slowly. In this paper, we propose a fast gradient descent (FGD) to optimize MD-NMF. FGD contains a Newton method that searches the optimal step length, and thus, FGD converges much faster than MUR. In addition, FGD includes MUR as a special case and can be applied to optimizing NMF and its variants. For a problem with 165 samples in R(1600), FGD converges in 28 s, while MUR requires 282 s. We also apply FGD in a variant of MD-NMF and experimental results confirm its efficiency. Experimental results on several face image datasets suggest the effectiveness of MD-NMF.
Naiyang Guan, Dacheng Tao, Zhigang Luo, Bo Yuan 0003
IEEE Trans. Image Process.3
2011 Non-Negative Patch Alignment Framework
abstract
In this paper, we present a non-negative patch alignment framework (NPAF) to unify popular non-negative matrix factorization (NMF) related dimension reduction algorithms. It offers a new viewpoint to better understand the common property of different NMF algorithms. Although multiplicative update rule (MUR) can solve NPAF and is easy to implement, it converges slowly. Thus, we propose a fast gradient descent (FGD) to overcome the aforementioned problem. FGD uses the Newton method to search the optimal step size, and thus converges faster than MUR. Experiments on synthetic and real-world datasets confirm the efficiency of FGD compared with MUR for optimizing NPAF. Based on NPAF, we develop non-negative discriminative locality alignment (NDLA). Experiments on face image and handwritten datasets suggest the effectiveness of NDLA in classification tasks and its robustness to image occlusions, compared with representative NMF-related dimension reduction algorithms.
Naiyang Guan, Dacheng Tao, Zhigang Luo, Bo Yuan 0003
IEEE Trans. Neural Networks3
2009 Managing Large-Scale Scientific Computing in Ensemble Prediction Using BPEL
abstract
The development of large-scale parallel scientific computing applications has put forward more urgent demands for powerful computing capacities and complex process managing technologies. Meanwhile, the scientific experiment processes become more and more complicated which makes it becomes a hard work for e-scientists to control the experiment analysis processes by hand. In this paper, we provide a scientific workflow system called EPSWFlow for the escientists in climate domain for services composition and workflow orchestration. In order to integrate the large number of the existing legacy applications into the system, we provide a service wrapping method and a unified interface for the workflow users to access to the services. The workflow system can process the experiment process dynamically and manage the heterogeneous grid resources transparently.
Cancan Liu, Zhigang Luo, Hai Liu 0014
ISPA3
2008 Service-Oriented Reliable Problem Solving Environment for Scientific Computation
abstract
Due to the large-scale and long running of scientific computation under the dynamic and unsteady grid architecture, the capability of fault-tolerance of scientific workflow management system becomes more and more important. In order to handle inevitable failures of activities in workflow, we present a three-level recovery strategy in this paper: in the service level, we provide a distributed Service Agent (SA) for each activity to monitor the execution status of workflow activities and implement the retry-based recovery strategy by submitting the failed activity multiple times; then in the workflow level, workflow engine implements replication-based strategy by request the Service Factory (SF) to create another service instance on a different node and invoke the new service instance for replacement; while in the user level, we provide a user interface for the users to handle the failure on demand. At last, a reliable Problem Solving Environment (PSE) in climate domain called Ensemble Prediction Scientific Workflow (EPSWFlow) is presented. This approach can seamlessly embed the complex control-flow intensive recovery strategies within the dataflow process network. Moreover, it can enable the prediction process more robust and more reusable.
Cancan Liu, Zhigang Luo, Hai Liu 0014
APSCC3
2008 Predicting RNA Secondary Structure Using Profile Stochastic Context-Free Grammars and Phylogenic Analysis
Xiaoyong Fang, Zhigang Luo, Zhenghua Wang
J. Comput. Sci. Technol.2
2007 Detecting and Assessing Conserved Stems for Accurate Structural Alignment of RNA Sequences
abstract
Since comparative sequence analysis is effective for non-coding RNA detection and RNA secondary structure prediction, efficient computational methods are expected for structuraf alignment of RNA sequences. However, it is very difficuit to construct a weft structuraf afignment without knowing the secondary structures. In this paper, we present a new method for constructing structural alignment of RNA sequences by detecting and assessing conserved stems. The method can be summarized by: 1) we detect conserved stems across multiple RNA sequences using the so-called position matrix with which some common paired positions are uncovered; 2) we assess the conserved stems using the Signal-to-Noise; 3) we build the structural alignment by incorporating some compatible conserved stems with the initial sequence alignment constructed by Clustal W program. We tested our method on the data sets composed of known structural alignments taken from the Rfam database. Experimentai resufts show that our method can buifd structuraf afignment of RNA sequences with much greater sensitivity and specificity than Clustal W.
Xiaoyong Fang, Zhigang Luo, Zhenghua Wang
BIBE2
2007 A Seed-Based Method for Predicting Common Secondary Structures in Unaligned RNA Sequences
Xiaoyong Fang, Zhigang Luo, Zhenghua Wang
MDAI2