Mulin Chen

dblp:176/7584 · DBLP profile ↗
← Back
46ranked-venue papers
12as first author
32since 2021 · last 2026
0000-0003-4634-3802ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 7 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1
YearPublicationVenuePosition
2026 Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval
abstract
In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in aligning with complex or ambiguous moments. To overcome these limitations, we explore Deep Evidential Regression (DER) to construct a vanilla Evidential baseline. However, this approach encounters two major issues: the inability to effectively handle modality imbalance and the structural differences in DER's heuristic uncertainty regularizer, which adversely affect uncertainty estimation. This misalignment results in high uncertainty being incorrectly associated with accurate samples rather than challenging ones. Our observations indicate that existing methods lack the adaptability required for complex video scenarios. In response, we propose Debiased Evidential Learning for Moment Retrieval (DEMR), a novel framework that incorporates a Reflective Flipped Fusion (RFF) block for cross-modal alignment and a query reconstruction task to enhance text sensitivity, thereby reducing bias in uncertainty estimation. Additionally, we introduce a Geom-regularizer to refine uncertainty predictions, enabling adaptive alignment with difficult moments and improving retrieval accuracy. Extensive testing on standard datasets and debiased datasets ActivityNet-CD and Charades-CD demonstrates significant enhancements in effectiveness, robustness, and interpretability, positioning our approach as a promising solution for temporal-semantic robustness in moment retrieval.
Haojian Huang, Kaijing Ma, Xianghao Zang, Han Fang 0002, Chao Ban, Hao Sun 0038, Mulin Chen, Zhongjiang He
AAAI10
2026 Locality-driven flexible consensus graph learning for multi-view clustering
Chusheng Zeng, Mulin Chen
Pattern Recognit.4
2025 Towards Learnable Anchor for Deep Multi-View Clustering
abstract
Deep multi-view clustering incorporating graph learning has presented tremendous potential. Most methods encounter costly square time consumption w.r.t. data size. Theoretically, anchor-based graph learning can alleviate this limitation, but related deep models mainly rely on manual discretization approaches to select anchors, which indicates that 1) the anchors are fixed during model training and 2) they may deviate from the true cluster distribution. Consequently, the unreliable anchors may corrupt clustering results. In this paper, we propose the Deep Multi-view Anchor Clustering (DMAC) model that performs clustering in linear time. Concretely, the initial anchors are intervened by the positive-incentive noise sampled from Gaussian distribution, such that they can be optimized with a newly designed anchor learning loss, which promotes a clear relationship between samples and anchors. Afterwards, anchor graph convolution is devised to model the cluster structure formed by the anchors, and the mutual information maximization loss is built to provide cross-view clustering guidance. In this way, the learned anchors can better represent clusters. With the optimal anchors, the full sample graph is calculated to derive a discriminative embedding for clustering. Extensive experiments on several datasets demonstrate the superior performance and efficiency of DMAC compared to state-of-the-art competitors.
Chusheng Zeng, Mulin Chen, Xuelong Li 0001
AAAI3
2025 Multi-Task Curriculum Graph Contrastive Learning with Clustering Entropy Guidance
abstract
Recent advances in unsupervised deep graph clustering have been significantly promoted by contrastive learning. Despite the strides, most graph contrastive learning models face challenges: 1) graph augmentation is used to improve learning diversity, but commonly used random augmentation methods may destroy inherent semantics and cause noise; 2) the fixed positive and negative sample selection strategy ignores the difficulty distribution of samples when deal with complex real data, thereby impeding the model’s capability to capture fine-grained patterns and trapping the model in sub-optimal for clustering. To reduce these problems, we propose the Clustering-guided Curriculum Graph contrastive Learning (CurGL) framework. CurGL uses clustering entropy as the guidance of the following graph augmentation and contrastive learning. Specifically, according to the clustering entropy, the intra-class edges and important features are emphasized in augmentation. Then, a multi-task curriculum learning scheme is proposed, which employs the clustering guidance to shift the focus from the discrimination task to the clustering task. In this way, the sample selection strategy of contrastive learning can be adjusted adaptively from early to late stage, which enhances the model's flexibility for complex data structure. Experimental results demonstrate that CurGL has achieved excellent performance compared to state-of-the-art competitors.
Chusheng Zeng, Jinghui Yuan, Mulin Chen, Xuelong Li 0001
IJCAI4
2025 Clustering-Oriented Generative Attribute Graph Imputation
abstract
Attribute-missing graph clustering has emerged as a significant unsupervised task, where only attribute vectors of partial nodes are available and the graph structure is intact. The related models generally follow the two-step paradigm of imputation and refinement. However, most imputation approaches fail to capture class-relevant semantic information, leading to sub-optimal imputation for clustering. Moreover, existing refinement strategies optimize the learned embedding through graph reconstruction, while neglecting the fact that some attributes are uncorrelated with the graph. To remedy the problems, we establish the Clustering-oriented Generative Imputation with reliable Refinement (CGIR) model. Concretely, the subcluster distributions are estimated to reveal the class-specific characteristics precisely, and constrain the sampling space of the generative adversarial module, such that the imputation nodes are impelled to align with the correct clusters. Afterwards, multiple subclusters are merged to guide the proposed edge attention network, which identifies the edge-wise attributes for each class, so as to avoid the redundant attributes in graph reconstruction from disturbing the refinement of overall embedding. To sum up, CGIR splits attribute-missing graph clustering into the search and mergence of subclusters, which guides to implement node imputation and refinement within a unified framework. Extensive experiments prove the advantages of CGIR over state-of-the-art competitors.
Mulin Chen, Zongcheng Miao, Xuelong Li 0001
ACM Multimedia1
2025 PIMG: Progressive Image-to-Music Generation With Contrastive Diffusion Models
abstract
The goal of Image-to-Music Generation is to create pure music according to the given image. Unlike existing tasks such as text-to-image generation, there is no explicit connection between image content and musical melody. Some existing studies attempt to generate music by directly mapping image features (such as color, edges, etc.) into musical notes, which may result in the melodic incoherence. Inspired by neuroscience, it is desirable to employ emotion to bridge these two modalities. However, the continuity and complexity of emotions make it difficult to capture the cross-modal correlation. Drawing from human perception mechanisms of emotions, a Progressive Image-to-Music Generation (PIMG) framework is proposed. The framework designs a mean-teacher based association network to guide the music generation process progressively, starting from highly correlated image-music pairs. The generation network receives more challenging sample pairs gradually, eventually capturing complex cross-modal emotional correspondences. Additionally, a contrastive learning strategy is introduced into the diffusion models to better capture the consistency between pieces of music with the similar emotions. Extensive experimental results demonstrate that the proposed framework is able to generate high-quality and emotionally consistent music from images.
Mulin Chen, Xuelong Li 0001
IEEE Trans. Multim.1
2024 Deep Contrastive Graph Learning with Clustering-Oriented Guidance
abstract
Graph Convolutional Network (GCN) has exhibited remarkable potential in improving graph-based clustering. To handle the general clustering scenario without a prior graph, these models estimate an initial graph beforehand to apply GCN. Throughout the literature, we have witnessed that 1) most models focus on the initial graph while neglecting the original features. Therefore, the discriminability of the learned representation may be corrupted by a low-quality initial graph; 2) the training procedure lacks effective clustering guidance, which may lead to the incorporation of clustering-irrelevant information into the learned graph. To tackle these problems, the Deep Contrastive Graph Learning (DCGL) model is proposed for general data clustering. Specifically, we establish a pseudo-siamese network, which incorporates auto-encoder with GCN to emphasize both the graph structure and the original features. On this basis, feature-level contrastive learning is introduced to enhance the discriminative capacity, and the relationship between samples and centroids is employed as the clustering-oriented guidance. Afterward, a two-branch graph learning mechanism is designed to extract the local and global structural relationships, which are further embedded into a unified graph under the cluster-level contrastive guidance. Experimental results on several benchmark datasets demonstrate the superiority of DCGL against state-of-the-art algorithms.
Mulin Chen, Xuelong Li 0001
AAAI1
2024 CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning
Haojian Huang, Xiaozhen Qiao, Zhuo Chen 0007, Bingyu Li 0002, Zhe Sun 0007, Mulin Chen, Xuelong Li 0001
ACM Multimedia7
2024 Robust Subcluster Search and Mergence Clustering
abstract
In recent years, graph-based clustering presents outstanding performance and has been widely investigated. It segments the data similarity graph into multiple subgraphs as final clusters. Many methods integrate graph learning and segmentation into a unified optimization problem to explore the graph structure. However, existing research 1) attempts to derive the final clusters from the learned graph directly, which relies on a highly tight internal distribution within each cluster, and is too strict for the real-world data; 2) generally constructs a holistic full sample graph, which means the outliers are involved in graph learning explicitly, and may corrupt the graph quality. To overcome the above limitations, a new clustering model called robust subcluster search and mergence (RSSM) is established in this article. Inspired by the positive-incentive noise (Pi-Noise), RSSM assumes that the outliers are useful for learning the data structure. Considering a few samples with large errors as outliers, RSSM finds the subcentroids by searching an imbalanced residue distribution. In this way, the subcentroids pull the normal samples together and push the outliers far away. Compared with the traditional clusters, the subclusters indicated by the subcentroids are more explicit, where the normal samples are tightly connected. After that, a subcluster similarity graph is constructed to guide the mergence of subclusters. To sum up, RSSM performs the search and mergence of subclusters simultaneously with the help of outliers, and generates a graph that is more suitable for clustering. Experiments on several datasets demonstrate the rationality and superiority of RSSM.
Mulin Chen, Xuelong Li 0001
IEEE Trans. Cybern.2
2024 Traffic Sign Interpretation via Natural Language Description
abstract
Most existing traffic sign-related works are dedicated to detecting and recognizing part of traffic signs separately, which fails to analyze the global semantic logic among signs and may convey inaccurate traffic instruction information. Following the above issues, we propose a traffic sign interpretation (TSI) task, which aims to interpret global semantic interrelated traffic signs (e.g., driving instruction-related texts, symbols, and guide panels) into a natural language for providing complete traffic instruction support to autonomous or assistant driving. Meanwhile, considering the lack of an effective framework for the proposed TSI task in existing works, we design a multi-task learning architecture (TSI-arch) to detect and recognize various traffic signs with drastic changes in sizes and aspect ratios. Meanwhile interpreting these signs into a natural language like a human according to Chinese design criteria of road traffic signs. Furthermore, the absence of a public TSI available dataset prompts us to build a traffic sign interpretation dataset, namely TSI-CN. The dataset consists of real road scene images, which are captured from the highway and the urban way in China from a driver’s perspective. It contains rich location labels of texts, symbols, and guide panels, and the corresponding natural language description labels. Experiments on our TSI-CN dataset demonstrate that the TSI task is achievable and the TSI architecture can interpret traffic signs from scenes successfully even if there is a complicated semantic logic among signs.
Chuang Yang 0003, Kai Zhuang, Mulin Chen, Haozhao Ma, Xu Han 0019, Tao Han 0002, Changxing Guo, Bingxuan Zhao, Qi Wang 0009
IEEE Trans. Intell. Transp. Syst.3
2024 Continuous Emotion-Based Image-to-Music Generation
abstract
Image-to-music generation aims to generate realistic pure music according to a given image. Although many previous works are conducted on bridging image and music, they mainly focus on the content-based cross-modal matching. For example, matching the Christmas song to an image that contains a Christmas tree. By comparison, image-to-music generation is a more challenging task due to its ambiguity and subjectivity. Specifically, there is no explicit correlation between the image content and music melody, without any lyric and human sound. Meanwhile, the perception of generated music varies from person to person. Inspired by the synesthesia phenomenon, we think that if an image tends to elicit a certain emotion on human, the generated music should also leave a similar impression. Therefore, in this paper, we propose a continuous emotion-based image-to-music generation framework, which uses emotion as the key for cross-modal generation. Specifically, a new image-music dataset is established, which uses valence-arousal (VA) space to capture the complex and nuanced nature of emotions. After that, a plug and play model is proposed to translate an image into a piece of music with similar emotion, which projects the emotions into continuous-valued labels, and explores both the intra-modal and inter-modal emotional consistency with contrastive learning. To our best knowledge, this is the first end-to-end framework towards the task of pure music generation from natural images. Extensive experiments show that the generated music achieves satisfactory emotional consistency with the input images, as well as impressive quality.
Mulin Chen, Xuelong Li 0001
IEEE Trans. Multim.2
2024 Zoom Text Detector
abstract
To pursue comprehensive performance, recent text detectors improve detection speed at the expense of accuracy. They adopt shrink-mask-based text representation strategies, which leads to a high dependence of detection accuracy on shrink-masks. Unfortunately, three disadvantages cause unreliable shrink-masks. Specifically, these methods try to strengthen the discrimination of shrink-masks from the background by semantic information. However, the feature defocusing phenomenon that coarse layers are optimized by fine-grained objectives limits the extraction of semantic features. Meanwhile, since both shrink-masks and the margins belong to texts, the detail loss phenomenon that the margins are ignored hinders the distinguishment of shrink-masks from the margins, which causes ambiguous shrink-mask edges. Moreover, false-positive samples enjoy similar visual features with shrink-masks. They aggravate the decline of shrink-masks recognition. To avoid the above problems, we propose a zoom text detector (ZTD) inspired by the zoom process of the camera. Specifically, zoomed-out view module (ZOM) is introduced to provide coarse-grained optimization objectives for coarse layers to avoid feature defocusing. Meanwhile, zoomed-in view module (ZIM) is presented to enhance the margins recognition to prevent detail loss. Furthermore, sequential-visual discriminator (SVD) is designed to suppress false-positive samples by sequential and visual features. Experiments verify the superior comprehensive performance of ZTD.
Chuang Yang 0003, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Neural Networks Learn. Syst.2
2023 One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance Field
abstract
Talking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encountered. Recent works instead employ explicit 3D structural representations or implicit neural rendering to improve performance under large pose changes. Nevertheless, the fidelity of identity and expression is not so desirable, especially for novel-view synthesis. In this paper, we propose HiDe-NeRF, which achieves high-fidelity and free-view talking-head synthesis. Drawing on the recently proposed Deformable Neural Radiance Fields, HiDe-NeRF represents the 3D dynamic scene into a canonical appearance field and an implicit deformation field, where the former comprises the canonical source face and the latter models the driving pose and expression. In particular, we improve fidelity from two aspects: (i) to enhance identity expressiveness, we design a generalized appearance module that leverages multi-scale volume features to preserve face shape and details; (ii) to improve expression preciseness, we propose a lightweight deformation module that explicitly decouples the pose and expression to enable precise expression modeling. Extensive experiments demonstrate that our proposed approach can generate better results than previous works. Project page: https://www.waytron.net/hidenerf/
Weichuang Li, Longhao Zhang, Dong Wang 0028, Bin Zhao 0001, Zhigang Wang 0002, Mulin Chen, Bang Zhang, Zhongjian Wang, Liefeng Bo, Xuelong Li 0001
CVPR6
2023 Fully Self-Supervised Depth Estimation from Defocus Clue
abstract
Depth-from-defocus (DFD), modeling the relationship between depth and defocus pattern in images, has demonstrated promising performance in depth estimation. Recently, several self-supervised works try to overcome the difficulties in acquiring accurate depth ground-truth. However, they depend on the all-in-focus (AIF) images, which cannot be captured in real-world scenarios. Such limitation discourages the applications of DFD methods. To tackle this issue, we propose a completely self-supervised framework that estimates depth purely from a sparse focal stack. We show that our framework circumvents the needs for the depth and AIF image ground-truth, and receives superior predictions, thus closing the gap between the theoretical success of DFD works and their applications in the real world. In particular, we propose (i) a more realistic setting for DFD tasks, where no depth or AIF image ground-truth is available; (ii) a novel self- supervision framework that provides reliable predictions of depth and AIF image under the challenging setting. The proposed framework uses a neural model to predict the depth and AIF image, and utilizes an optical model to validate and refine the prediction. We verify our framework on three benchmark datasets with rendered focal stacks and real focal stacks. Qualitative and quantitative evaluations show that our method provides a strong baseline for self- supervised DFD tasks. The source code is publicly avail- able at https://github.com/Ehzoahis/DEReD.
Haozhe Si, Bin Zhao 0001, Dong Wang 0028, Mulin Chen, Zhigang Wang 0002, Xuelong Li 0001
CVPR5
2023 Propagate and Calibrate: Real-Time Passive Non-Line-of-Sight Tracking
abstract
Non-line-of-sight (NLOS) tracking has drawn increasing attention in recent years, due to its ability to detect object motion out of sight. Most previous works on NLOS tracking rely on active illumination, e.g., laser, and suffer from high cost and elaborate experimental conditions. Besides, these techniques are still far from practical application due to oversimplified settings. In contrast, we propose a purely passive method to track a person walking in an invisible room by only observing a relay wall, which is more in line with real application scenarios, e.g., security. To excavate imperceptible changes in videos of the relay wall, we introduce difference frames as an essential carrier of temporal-local motion messages. In addition, we propose PAC-Net, which consists of alternating propagation and calibration, making it capable of leveraging both dynamic and static messages on a frame-level granularity. To evaluate the proposed method, we build and publish the first dynamic passive NLOS tracking dataset, NLOS-Track, which fills the vacuum of realistic NLOS datasets. NLOS-Track contains thousands of NLOS video clips and corresponding trajectories. Both real-shot and synthetic data are included. Our codes and dataset are available at https://againstentropy.github.io/NLOS-Track/.
Zhigang Wang 0002, Bin Zhao 0001, Dong Wang 0028, Mulin Chen, Xuelong Li 0001
CVPR5
2023 A Multitask Framework for Graffiti-to-Image Translation
abstract
Recently, image-to-image translation models have achieved great success in terms of content consistency and visual fidelity. However, in most of these tasks, the inaccuracy of sketches and the high cost of fine semantic masks acquisition limit the large-scale use of image translation models. Therefore, we propose to use graffiti that combines the advantages of sketches and semantic masks as model input. Graffiti reflects the general content of an image using lines and color distinctions, with some unlabeled regions. However, due to the large number of unknown areas in the graffiti, the generated results may be blurred, resulting in poor visual effects. To address these challenges, this paper proposes a multi-task framework that can predict unknown regions by learning semantic mask from graffiti, thereby improving the quality of generated real scene images. Furthermore, by introducing an edge activation module, which utilizes semantic and edge information to optimize the object boundaries of the generated images, the details of the generated images can be improved. Experiments on the Cityscapes dataset demonstrate that our multi-task framework achieves competitive performance on graffiti-based image generation task.
Mulin Chen, Xuelong Li 0001
ACM Multimedia2
2023 Projection concept factorization with self-representation for data clustering
Chenyu Shao, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009
Neurocomputing2
2023 Feature Weighted Non-Negative Matrix Factorization
abstract
Non-negative matrix factorization (NMF) is one of the most popular techniques for data representation and clustering and has been widely used in machine learning and data analysis. NMF concentrates the features of each sample into a vector and approximates it by the linear combination of basis vectors, such that the low-dimensional representations are achieved. However, in real-world applications, the features usually have different importance. To exploit the discriminative features, some methods project the samples into the subspace with a transformation matrix, which disturbs the original feature attributes and neglects the diversity of samples. To alleviate the above problems, we propose the feature weighted NMF (FNMF) in this article. The salient properties of FNMF can be summarized as three-fold: 1) it learns the weights of features adaptively according to their importance; 2) it utilizes multiple feature weighting components to preserve the diversity; and 3) it can be solved efficiently with the suggested optimization algorithm. The performance on synthetic and real-world datasets demonstrates that the proposed method obtains the state-of-the-art performance.
Mulin Chen, Maoguo Gong, Xuelong Li 0001
IEEE Trans. Cybern.1
2023 Reinforcement Shrink-Mask for Text Detection
abstract
Existing real-time text detectors reconstruct text contours by shrink-masks only. Though they simplify the framework and can make the model run fast, the strong dependence on shrink-masks leads to unreliable detection results (e.g., miss detection and overdetection). Moreover, these methods ignore the information from surrounding pixels, which causes sensitive shrink-masks and accelerates the reliability decline of detection results. Considering the above problems, we construct an effective and efficient text detection network, termed as Reinforcement Shrink-Mask for Text Detection (RSMTD), which strengthens the model's ability to recognize texts while enjoying a high detection speed. Specifically, an effective text representation strategy (Reinforcement Shrink-Mask, RSM) is designed to decouple texts and shrink-masks. RSM builds texts through shrink-masks and reinforcement offsets to ensure stable detection results encountering shrink-masks that deviate from the ground-truth. It is worth noting that reinforcement offsets can force our method to focus on the foreground shapes to bring precise shrink-mask edges. For the robustness improvement of shrink-masks, Super-pixel Window (SPW) is proposed to encourage RSMTD to utilize the surroundings of each pixel to predict shrink-masks. Particularly, SPW treats the interval regions between texts and shrink-masks as background, which helps to suppress interval regions and to avoid text adhesion. Moreover, a lightweight feature merging branch is constructed to further accelerate the inference process. As demonstrated in the experiments, our method is superior to existing state-of-the-art (SOTA) methods in both detection accuracy and speed on multiple benchmarks.
Chuang Yang 0003, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Multim.2
2023 Text Growing on Leaf
abstract
Irregular-shaped texts bring challenges to Scene Text Detection (STD). Although existing regression-based approaches achieve comparable performances, they fail to cover some highly curved ribbon-like text lines. Inspired by morphology, we found that the leaf vein can easily cover various geometries. Specifically, lateral and thin veins are emitted to margin along main vein gradually with the leaf growth. This process can decompose a concave object into consecutive convex regions, which are easier to fit. Hence, the leaf vein is suitable for representing highly curved texts. Considering the aforementioned advantage, we design a leaf vein-based text representation method (LVT), where text contour is treated as leaf margin and represented through main, lateral, and thin veins. We further construct a detection framework based on LVT, namely LeafText. In the text reconstruction stage, LeafText simulates the leaf growth process to rebuild text contours. It grows main veins in Cartesian coordinates to locate texts roughly at first. Then, lateral and thin veins are generated along the main vein growth direction in polar coordinates. They are responsible for generating the coarse contour and refining it, respectively. Meanwhile, Multi-Oriented Smoother (MOS) is designed to smooth the main vein for ensuring reliable growth directions of lateral and thin veins. Additionally, a global incentive loss is proposed to enhance the predictions of lateral and thin veins. Ablation experiments demonstrate LVT can fit irregular-shaped texts precisely and verify the effectiveness of MOS and global incentive loss. Comparisons show that LeafText is superior to existing state-of-the-art (SOTA) methods on MSRA-TD500, CTW1500, Total-Text, and ICDAR2015 datasets.
Chuang Yang 0003, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Multim.2
2023 Entropy Minimizing Matrix Factorization
abstract
Nonnegative matrix factorization (NMF) is a widely used data analysis technique and has yielded impressive results in many real-world tasks. Generally, existing NMF methods represent each sample with several centroids and find the optimal centroids by minimizing the sum of the residual errors. However, outliers deviating from the normal data distribution may have large residues and then dominate the objective value. In this study, an entropy minimizing matrix factorization (EMMF) framework is developed to tackle the above problem. Considering that outliers are usually much less than the normal samples, a new entropy loss function is established for matrix factorization, which minimizes the entropy of the residue distribution and allows a few samples to have large errors. In this way, the outliers do not affect the approximation of normal samples. Multiplicative updating rules for EMMF are derived, and the convergence is proven theoretically. In addition, a Graph regularized version of EMMF (G-EMMF) is also presented, which uses a data graph to capture the data relationship. Clustering results on various synthetic and real-world datasets demonstrate the advantages of the proposed models, and the effectiveness is also verified through the comparison with state-of-the-art methods.
Mulin Chen, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 BiP-Net: Bidirectional Perspective Strategy Based Arbitrary-Shaped Text Detection Network
abstract
Detecting irregular-shaped text instances is the main challenge for text detection. Existing approaches can be roughly divided into top-down and bottom-up perspective methods. The former encodes text contours into unified units, which always fails to fit highly curved text contours. The latter represents text instances by a number of local units, where the complicated network and post-processing lead to slow detection speed. In this paper, to detect arbitrary-shaped text instances with high detection accuracy and speed simultaneously, we propose a Bidirectional Perspective strategy based Network (BiP-Net). Specifically, a new text representation strategy is proposed to represent text contours from a topdown perspective, which can fit highly curved text contours effectively. Moreover, a contour connecting (CC) algorithm is proposed to avoid the information loss of text contours by rebuilding interval contours from a bottom-up perspective. The experimental results on MSRA-TD500, CTW1500, and ICDAR2015 datasets demonstrate the superiority of BiP-Net against several state-of-the-art methods.
Chuang Yang 0003, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009
ICASSP2
2022 Robust doubly stochastic graph clustering
Mulin Chen, Maoguo Gong, Xuelong Li 0001
Neurocomputing1
2022 Locality Adaptive Discriminant Analysis Framework
abstract
Linear discriminant analysis (LDA) is a well-known technique for supervised dimensionality reduction and has been extensively applied in many real-world applications. LDA assumes that the samples are Gaussian distributed, and the local data distribution is consistent with the global distribution. However, real-world data seldom satisfy this assumption. To handle the data with complex distributions, some methods emphasize the local geometrical structure and perform discriminant analysis between neighbors. But the neighboring relationship tends to be affected by the noise in the input space. In this research, we propose a new supervised dimensionality reduction method, namely, locality adaptive discriminant analysis (LADA). In order to directly process the data with matrix representation, such as images, the 2-D LADA (2DLADA) is also developed. The proposed methods have the following salient properties: 1) they find the principle projection directions without imposing any assumption on the data distribution; 2) they explore the data relationship in the desired subspace, which contains less noise; and 3) they find the local data relationship automatically without the efforts for tuning parameters. The performance of dimensionality reduction shows the superiorities of the proposed methods over the state of the art.
Xuelong Li 0001, Qi Wang 0009, Feiping Nie 0001, Mulin Chen
IEEE Trans. Cybern.4
2022 Autoweighted Multiview Feature Selection With Graph Optimization
abstract
In this article, we focus on the unsupervised multiview feature selection, which tries to handle high-dimensional data in the field of multiview learning. Although some graph-based methods have achieved satisfactory performance, they ignore the underlying data structure across different views. Besides, their predefined Laplacian graphs are sensitive to the noises in the original data space and fail to obtain the optimal neighbor assignment. To address the above problems, we propose a novel unsupervised multiview feature selection model based on graph learning, and the contributions are three-fold: 1) during the feature selection procedure, the consensus similarity graph shared by different views is learned. Therefore, the proposed model can reveal the data relationship from the feature subset; 2) a reasonable rank constraint is added to optimize the similarity matrix to obtain more accurate information; and 3) an autoweighted framework is presented to assign view weights adaptively, and an effective alternative iterative algorithm is proposed to optimize the problem. Experiments on various datasets demonstrate the superiority of the proposed method compared to the state-of-the-art methods.
Qi Wang 0009, Mulin Chen, Xuelong Li 0001
IEEE Trans. Cybern.3
2022 Robust Rank-Constrained Sparse Learning: A Graph-Based Framework for Single View and Multiview Clustering
abstract
Graph-based clustering aims to partition the data according to a similarity graph, which has shown impressive performance on various kinds of tasks. The quality of similarity graph largely determines the clustering results, but it is difficult to produce a high-quality one, especially when data contain noises and outliers. To solve this problem, we propose a robust rank constrained sparse learning (RRCSL) method in this article. The$L_{2,1}$-norm is adopted into the objective function of sparse representation to learn the optimal graph with robustness. To preserve the data structure, we construct an initial graph and search the graph within its neighborhood. By incorporating a rank constraint, the learned graph can be directly used as the cluster indicator, and the final results are obtained without additional postprocessing. In addition, the proposed method cannot only be applied to single-view clustering but also extended to multiview clustering. Plenty of experiments on synthetic and real-world datasets have demonstrated the superiority and robustness of the proposed framework.
Qi Wang 0009, Mulin Chen, Xuelong Li 0001
IEEE Trans. Cybern.3
2022 Spatial-Spectral Clustering With Anchor Graph for Hyperspectral Image
abstract
Hyperspectral image (HSI) clustering, which aims at dividing hyperspectral pixels into clusters without labeled training data, has drawn significant attention in practical applications. Recently, many graph-based clustering methods, which construct an adjacent graph to model the data relationship, have shown dominant performance. However, the high dimensionality of HSI data makes it hard to construct the pairwise adjacent graph. Besides, abundant spatial structures are often overlooked during the clustering procedure. In order to better handle the high dimensionality problem and preserve the spatial structures, this paper proposes a novel unsupervised approach called spatial-spectral clustering with anchor graph (SSCAG) for HSI data clustering. The SSCAG has the following contributions: 1) the multiscale filtering module is utilized to smooth the homogeneous regions, so that it can increase the similarity and consistency of neighboring pixels and capture the multiple views of a local region with different scales; 2) a new similarity metric is proposed to embed the spatial-spectral features into the combined adjacent graph, which can mine the intrinsic property structure of HSI data; 3) the AG-based strategy is adopted to construct the adjacent graph by a neighbor assignment scheme without hyperparameters, and its optimization employs SVD to replace eigenvalue decomposition to reduce the computational complexity. Extensive experiments on three public HSI datasets show that the proposed SSCAG is competitive against the state-of-the-art approaches.
Qi Wang 0009, Yanling Miao, Mulin Chen, Yuan Yuan 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 CM-Net: Concentric Mask Based Arbitrary-Shaped Text Detection
abstract
Recently fast arbitrary-shaped text detection has become an attractive research topic. However, most existing methods are non-real-time, which may fall short in intelligent systems. Although a few real-time text methods are proposed, the detection accuracy is far behind non-real-time methods. To improve the detection accuracy and speed simultaneously, we propose a novel fast and accurate text detection framework, namely CM-Net, which is constructed based on a new text representation method and a multi-perspective feature (MPF) module. The former can fit arbitrary-shaped text contours by concentric mask (CM) in an efficient and robust way. The latter encourages the network to learn more CM-related discriminative features from multiple perspectives and brings no extra computational cost. Benefiting the advantages of CM and MPF, the proposed CM-Net only needs to predict one CM of the text instance to rebuild the text contour and achieves the best balance between detection accuracy and speed compared with previous works. Moreover, to ensure that multi-perspective features are effectively learned, the multi-factor constraints loss is proposed. Extensive experiments demonstrate the proposed CM is efficient and robust to fit arbitrary-shaped text instances, and also validate the effectiveness of MPF and constraints loss for discriminative text features recognition. Furthermore, experimental results show that the proposed CM-Net is superior to existing state-of-the-art (SOTA) real-time text detection methods in both detection speed and accuracy on MSRA-TD500, CTW1500, Total-Text, and ICDAR2015 datasets.
Chuang Yang 0003, Mulin Chen, Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Image Process.2
2021 Spatial-Spectral Hyperspectral Image Classification Via Multiple Random Anchor Graphs Ensemble Learning
abstract
Graph-based semi-supervised learning methods, which deal well with the situation of limited labeled data, have shown dominant performance in practical applications. However, the high dimensionality of hyperspectral images (HSI) makes it hard to construct the pairwise adjacent graph. Besides, the fine spatial features that help improve the discriminability of the model are often overlooked. To handle the problems, this paper proposes a novel spatial-spectral HSI classification method via multiple random anchor graphs ensemble learning (RAGE). Firstly, the local binary pattern is adopted to extract the more descriptive features on each selected band, which preserves local structures and subtle changes of a region. Secondly, the adaptive neighbors assignment is introduced in the construction of anchor graph, to reduce the computational complexity. Finally, an ensemble model is built by utilizing multiple anchor graphs, such that the diversity of HSI is learned. Extensive experiments show that RAGE is competitive against the state-of-the-art approaches.
Yanling Miao, Qi Wang 0009, Mulin Chen, Xuelong Li 0001
IGARSS3
2021 Two-stream network for infrared and visible images fusion
Luolin Liu, Mulin Chen, Mingliang Xu 0001, Xuelong Li 0001
Neurocomputing2
2021 Concept Factorization With Local Centroids
abstract
Data clustering is a fundamental problem in the field of machine learning. Among the numerous clustering techniques, matrix factorization-based methods have achieved impressive performances because they are able to provide a compact and interpretable representation of the input data. However, most of the existing works assume that each class has a global centroid, which does not hold for data with complicated structures. Besides, they cannot guarantee that the sample is associated with the nearest centroid. In this work, we present a concept factorization with the local centroids (CFLCs) approach for data clustering. The proposed model has the following advantages: 1) the samples from the same class are allowed to connect with multiple local centroids such that the manifold structure is captured; 2) the pairwise relationship between the samples and centroids is modeled to produce a reasonable label assignment; and 3) the clustering problem is formulated as a bipartite graph partitioning task, and an efficient algorithm is designed for optimization. Experiments on several data sets validate the effectiveness of the CFLC model and demonstrate its superior performance over the state of the arts.
Mulin Chen, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Robust Matrix Factorization With Spectral Embedding
abstract
Nonnegative matrix factorization (NMF) and spectral clustering are two of the most widely used clustering techniques. However, NMF cannot deal with the nonlinear data, and spectral clustering relies on the postprocessing. In this article, we propose a Robust Matrix factorization with Spectral embedding (RMS) approach for data clustering, which inherits the advantages of NMF and spectral clustering, while avoiding their shortcomings. In addition, to cluster the data represented by multiple views, we present the multiview version of RMS (M-RMS), and the weights of different views are self-tuned. The main contributions of this research are threefold: 1) by integrating spectral clustering and matrix factorization, the proposed methods are able to capture the nonlinear data structure and obtain the cluster indicator directly; 2) instead of using the squared Frobenius-norm, the objectives are developed with the$\ell _{2,1}$-norm, such that the effects of the outliers are alleviated; and 3) the proposed methods are totally parameter-free, which increases the applicability for various real-world problems. Extensive experiments on several single-view/multiview data sets demonstrate the effectiveness of our methods and verify their superior clustering performance over the state of the arts.
Mulin Chen, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Robust Rank Constrained Sparse Learning: A Graph-Based Method for Clustering
abstract
Graph-based clustering is an advanced clustering techniuqe, which partitions the data according to an affinity graph. However, the graph quality affects the clustering results to a large extent, and it is difficult to construct a graph with high quality, especially for data with noises and outliers. To solve this problem, a robust rank constrained sparse learning method is proposed in this paper. The L2,1-norm objective function of sparse representation is introduced to learn the optimal graph with robustness. To preserve the data structure, the graph is searched within the neighborhood of the initial graph. By incorporating a rank constraint, the learned graph can be directly used as the cluster indicator and the final results is obtained without additional post-processing. Plenty of experiments on real-world data sets have proved the superiority and the robustness of the proposed approach.
Mulin Chen, Qi Wang 0009, Xuelong Li 0001
ICASSP2
2020 Detecting Coherent Groups in Crowd Scenes by Multiview Clustering
abstract
Detecting coherent groups is fundamentally important for crowd behavior analysis. In the past few decades, plenty of works have been conducted on this topic, but most of them have limitations due to the insufficient utilization of crowd properties and the arbitrary processing of individuals. In this study, a Multiview-based Parameter Free framework (MPF) is proposed. Based on the L1-norm and L2-norm, we design two versions of the multiview clustering method, which is the main part of the proposed framework. This paper presents the contributions on three aspects: (1) a new structural context descriptor is designed to characterize the structural properties of individuals in crowd scenes; (2) a self-weighted multiview clustering method is proposed to cluster feature points by incorporating their orientation and context similarities; and (3) a novel framework is introduced for group detection, which is able to determine the group number automatically without any parameter or threshold to be tuned. The effectiveness of the proposed framework is evaluated on real-world crowd videos, and the experimental results show its promising performance on group detection. In addition, the proposed multiview clustering method is also evaluated on a synthetic dataset and several standard benchmarks, and its superiority over the state-of-the-art competitors is demonstrated.
Qi Wang 0009, Mulin Chen, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Quantifying and Detecting Collective Motion in Crowd Scenes
abstract
People in crowd scenes always exhibit consistent behaviors and form collective motions. The analysis of collective motion has motivated a surge of interest in computer vision. Nevertheless, the effort is hampered by the complex nature of collective motions. Considering the fact that collective motions are formed by individuals, this paper proposes a new framework for both quantifying and detecting collective motion by investigating the spatio-temporal behavior of individuals. The main contributions of this work are threefold: 1) an intention-aware model is built to fully capture the intrinsic dynamics of individuals; 2) a structure-based collectiveness measurement is developed to accurately quantify the collective properties of crowds; 3) a multistage clustering strategy is formulated to detect both the local and global behavior consistency in crowd scenes. Experiments on real world data sets show that our method is able to handle crowds with various structures and time-varying dynamics. Especially, the proposed method shows nearly 10% improvement over the competitors in terms of NMI, Purity and RI. Its applicability is illustrated in the context of anomaly detection and semantic scene segmentation.
Xuelong Li 0001, Mulin Chen, Qi Wang 0009
IEEE Trans. Image Process.2
2020 Adaptive Consistency Propagation Method for Graph Clustering
abstract
Graph clustering plays an important role in data mining. Based on an input data graph, data points are partitioned into clusters. However, most existing methods keep the data graph fixed during the clustering procedure, so they are limited to exploit the implied data manifold and highly dependent on the initial graph construction. Inspired by the recent development on manifold learning, this paper proposes an Adaptive Consistency Propagation (ACP) method for graph clustering. In order to utilize the features captured from different perspectives, we further put forward the Multi-view version of the ACP model (MACP). The main contributions are threefold: (1) the manifold structure of input data is sufficiently exploited by propagating the topological connectivities between data points from near to far; (2) the optimal graph for clustering is learned by taking graph learning as a part of the optimization procedure; and (3) the negotiation among the heterogeneous features is captured by the multi-view clustering model. Extensive experiments on real-world datasets validate the effectiveness of the proposed methods on both single-and multi-view clustering, and show their superior performance over the state-of-the-arts.
Xuelong Li 0001, Mulin Chen, Qi Wang 0009
IEEE Trans. Knowl. Data Eng.2
2020 Discrimination-Aware Projected Matrix Factorization
abstract
Non-negative Matrix Factorization (NMF) has been one of the most popular clustering techniques in machine leaning, and involves various real-world applications. Most existing works perform matrix factorization on high-dimensional data directly. However, the intrinsic data structure is always hidden within the low-dimensional subspace. And, the redundant features within the input space may affect the final result adversely. In this paper, a new unsupervised matrix factorization method, Discrimination-aware Projected Matrix Factorization (DPMF), is proposed for data clustering. The main contributions are threefold: (1) The linear discriminant analysis is jointly incorporated into the unsupervised matrix factorization framework, so the clustering can be accomplished in the discriminant subspace. (2) The manifold regularization is introduced to perceive the geometric information, and the ℓ2,1-norm is utilized to improve the robustness. (3) An efficient optimization algorithm is designed to solve the proposed problem with proved convergence. Experimental results on one toy dataset and eight real-world benchmarks show the effectiveness of the proposed method.
Xuelong Li 0001, Mulin Chen, Qi Wang 0009
IEEE Trans. Knowl. Data Eng.2
2019 Self-Tuned Discrimination-Aware Method for Unsupervised Feature Selection
abstract
Unsupervised feature selection is fundamentally important for processing unlabeled high-dimensional data, and several methods have been proposed on this topic. Most existing embedded unsupervised methods just emphasize the data structure in the input space, which may contain large noise. Therefore, they are limited to perceive the discriminative information implied within the low-dimensional manifold. In addition, these methods always involve several parameters to be tuned, which is time-consuming. In this paper, we present a self-tuned discrimination-aware (STDA) approach for unsupervised feature selection. The main contributions of this paper are threefold: 1) it adopts the advantage of discriminant analysis technique to select the valuable features; 2) it learns the local data structure adaptively in the discriminative subspace to alleviate the effect of data noise; and 3) it performs feature selection and clustering simultaneously with an efficient optimization strategy, and saves the additional efforts to tune parameters. Experimental results on a toy data set and various real-world benchmarks justify the effectiveness of STDA on both feature selection and data clustering, and demonstrate its promising performance against the state of the arts.
Xuelong Li 0001, Mulin Chen, Qi Wang 0009
IEEE Trans. Neural Networks Learn. Syst.2
2018 Robust Adaptive Sparse Learning Method for Graph Clustering
abstract
Graph clustering aims to group the data into clusters according to a similarity graph, and has received sufficient attention in computer vision. As the basis of clustering, the quality of graph affects the results directly. In this paper, a Robust Adaptive Sparse Learning (RASL) method is proposed to improve the graph quality. The contributions made in this paper are three fold: (1) the sparse representation technique is employed to enforce the graph sparsity, and the l2,1 norm is introduced to improve the robustness; (2) the intrinsic manifold structure is captured by investigating the local relationship of data points; (3) an efficient optimization algorithm is designed to solve the proposed problem. Experimental results on various real-world benchmark datasets demonstrate the promising results of the proposed graph-based clustering method.
Mulin Chen, Qi Wang 0009, Xuelong Li 0001
ICIP1
2018 Adaptive Projected Matrix Factorization method for data clustering
Mulin Chen, Qi Wang 0009, Xuelong Li 0001
Neurocomputing1
2017 A Multiview-Based Parameter Free Framework for Group Detection
abstract
Group detection is fundamentally important for analyzing crowd behaviors, and has attracted plenty of attention in artificial intelligence. However, existing works mostly have limitations due to the insufficient utilization of crowd properties and the arbitrary processing of individuals. In this paper,we propose the Multiview-based Parameter Free (MPF) approach to detect groups in crowd scenes. The main contributions made in this study are threefold: (1) a new structural context descriptor is designed to characterize the structural property of individuals in crowd motions; (2) an self-weighted multiview clustering method is proposed to cluster feature points by incorporating their motion and context similarities;(3) a novel framework is introduced for group detection, which is able to determine the group number automatically without any parameter or threshold to be tuned. Extensive experiments on various real world datasets demonstrate the effectiveness of the proposed approach, and show its superiority against state-of-the-art group detection techniques.
Xuelong Li 0001, Mulin Chen, Feiping Nie 0001, Qi Wang 0009
AAAI2
2017 Quantifying and Detecting Collective Motion by Manifold Learning
abstract
The analysis of collective motion has attracted many researchers in artificial intelligence. Though plenty of works have been done on this topic, the achieved performance isstill unsatisfying due to the complex nature of collective motions. By investigating the similarity of individuals, this paper proposes a novel framework for both quantifying and detecting collective motions. Our main contributions are threefold: (1) the time-varying dynamics of individuals are deeply investigated to better characterize the individual motion; (2) a structure-based collectiveness measurement is designed toprecisely quantify both individual-level and scene-level properties of collective motions; (3) a multi-stage clustering strategy is presented to discover a more comprehensive understanding of the crowd scenes, containing both local and global collective motions. Extensive experimental results on realworld data sets show that our method is capable of handling crowd scenes with complicated structures and various dynamics, and demonstrate its superior performance against state-of-the-art competitors.
Qi Wang 0009, Mulin Chen, Xuelong Li 0001
AAAI2
2017 Anchor-based group detection in crowd scenes
abstract
Group detection aims to classify pedestrians into categories according to their motion dynamics. It's fundamental for analyzing crowd behaviors and involves a wide range of applications. In this paper, we propose a Anchor-based Manifold Ranking (AMR) method to detect groups in crowd scenes. Our main contributions are threefold: (1) the topological relationship of individuals are effectively investigated with a manifold ranking method; (2) global consistency in crowds are accurately recognized by a coherent merging strategy; (3) the number of groups is decided automatically based on the similarity graph of individuals. Experimental results show that the proposed framework is competitive against the state-of-the-art methods.
Mulin Chen, Qi Wang 0009, Xuelong Li 0001
ICASSP1
2017 Locality Adaptive Discriminant Analysis
abstract
Linear Discriminant Analysis (LDA) is a popular technique for supervised dimensionality reduction, and its performance is satisfying when dealing with Gaussian distributed data. However, the neglect of local data structure makes LDA inapplicable to many real-world situations. So some works focus on the discriminant analysis between neighbor points, which can be easily affected by the noise in the original data space. In this paper, we propose a new supervised dimensionality reduction method, Locality Adaptive Discriminant Analysis (LADA), to lean a representative subspace of the data. Compared to LDA and its variants, the proposed method has three salient advantages: (1) it finds the principle projection directions without imposing any assumption on the data distribution; (2) it’s able to exploit the local manifold structure of data in the desired subspace; (3) it exploits the points’ neighbor relationship automatically without introducing any additional parameter to be tuned. Performance on synthetic datasets and real-world benchmark datasets demonstrate the superiority of the proposed method.
Xuelong Li 0001, Mulin Chen, Feiping Nie 0001, Qi Wang 0009
IJCAI2
2017 Patch-based topic model for group detection
Mulin Chen, Qi Wang 0009, Xuelong Li 0001
Sci. China Inf. Sci.1
2016 Measuring Collectiveness via Refined Topological Similarity
abstract
Crowd system has motivated a surge of interests in many areas of multimedia, as it contains plenty of information about crowd scenes. In crowd systems, individuals tend to exhibit collective behaviors, and the motion of all those individuals is called collective motion. As a comprehensive descriptor of collective motion, collectiveness has been proposed to reflect the degree of individuals moving as an entirety. Nevertheless, existing works mostly have limitations to correctly find the individuals of a crowd system and precisely capture the various relationships between individuals, both of which are essential to measure collectiveness. In this article, we propose a collectiveness-measuring method that is capable of quantifying collectiveness accurately. Our main contributions are threefold: (1) we compute relatively accurate collectiveness by making the tracked feature points represent the individuals more precisely with a point selection strategy; (2) we jointly investigate the spatial-temporal information of individuals and utilize it to characterize the topological relationship between individuals by manifold learning; (3) we propose a stability descriptor to deal with the irregular individuals, which influence the calculation of collectiveness. Intensive experiments on the simulated and real world datasets demonstrate that the proposed method is able to compute relatively accurate collectiveness and keep high consistency with human perception.
Xuelong Li 0001, Mulin Chen, Qi Wang 0009
ACM Trans. Multim. Comput. Commun. Appl.2