Qian Guo 0005

dblp:22/2222-5 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0001-9189-7834ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Uncertainty-Guided View-Strength-Aware Feature Utilization for Multi-View Classification
abstract
In multi-view classification tasks (MVC), each view provides an unique perspective on the data, offering complementary information that can improve classification performance when properly integrated. However, traditional methods typically adopt a uniform processing strategy for all views before fusion, overlooking the fact that different views may require different treatments due to variations in their quality and informativeness. To address this limitation, we propose a novel framework called Uncertainty-Guided View-Strength-Aware Feature Utilization (UVF) for multi-view classification. Our approach introduces a view uncertainty estimation module to quantify the discriminative strength of each view. Based on this estimation, a Differentiated Feature Selector (DFS) adaptively selects features, retaining informative dimensions in weak views while preserving original features in strong views. Furthermore, we employ an uncertainty-guided fusion strategy that assigns dynamic weights to each view's contribution based on its uncertainty score, enhancing the robustness and reliability of the final decision. Experimental results on benchmark datasets demonstrate that our method significantly outperforms conventional approaches, achieving better classification accuracy and interpretability through strength-aware feature processing and fusion.
Qian Guo 0005, Li Zhang 0104, Liang Du 0003, Bingbing Jiang 0001, Lu Chen 0003, Xinyan Liang
AAAI2
2026 EvoFMVC: Trusted Federated Multi-View Clustering with Evolutionary Fusion
abstract
With the growing demand for decentralized collaborative analysis of privacy-sensitive data, federated multi-view clustering (FMVC) has attracted widespread attention due to its ability to balance privacy protection and collaborative modeling. However, current methods still face the following challenges: (1) Clients need to frequently upload high-dimensional data such as model parameters or graph structures, resulting in high communication costs; (2) The structured data uploaded often contains semantic features and has a high risk of being inverted; (3) The server usually merges the data from all clients with the fixed fusion rule, which may result in a suboptimized clustering result when there exist low-quality clients. To address the issues, we propose a new trusted federated multi-view clustering framework (EvoFMVC) that introduces three key innovations: First, lightweight trusted evidence serves as a compact communication medium, significantly reducing overhead compared to conventional model parameters or graph structures. Second, trusted evidences express clustering results in the form of probability distribution, which avoids the risk of structured information being easily inverted. Lastly, we formalize the server-side aggregation process as a neural architecture search (NAS) task where the server flexibly uses different fusion operators to filter and fuse necessary views through evolutionary algorithms, which significantly improves the fusion effect and model performance. Experimental results on multiple datasets show that our method is superior to existing FMVC methods in terms of clustering accuracy and communication efficiency.
Li Zhang 0104, Pinhan Fu, Qian Guo 0005, Liang Du 0003, Xinyan Liang
AAAI4
2026 Multi-Branch Tree-Based Fusion Neural Architecture Search With Zero-Cost Screen for Multi-Modal Classification
abstract
Multi-modal classification (MMC) leverages effective fusion of information from diverse modalities to achieve superior classification performance. Existing fusion methods, however, rely on expert-designed architectures that demand substantial domain expertise and computational resources, with fixed topologies that offer limited structural flexibility across tasks and datasets. Although neural architecture search (NAS)-based fusion methods have been proposed to automatically discover high-performing architectures, these approaches are computationally expensive and predominantly rely on pairwise modality combinations, which fail to capture complex multi-variable correlations, limiting the architectures' expressiveness and flexibility. Consequently, there remains a lack of multi-modal fusion frameworks that can simultaneously achieve high efficiency and high accuracy. To break through these bottlenecks, we propose a multi-branch tree-based fusion neural architecture search framework (MBTF-NAS). For performance enhancement, MBTF-NAS employs a multi-branch tree-structured encoding strategy that enables dynamic and computationally efficient exploration of fusion topologies and substantially strengthens cross-modal interaction. A learnable model-level attention weighting mechanism further emphasizes informative modalities, improving the overall quality of multi-modal feature fusion. For efficiency improvement, MBTF-NAS leverages zero-cost proxy metrics for architecture evaluation, enabling rapid identification of high-potential candidates while dramatically reducing computational overhead. We conducted a comprehensive evaluation of MBTF-NAS on seven representative multi-modal benchmarks. The experimental results demonstrate that MBTF-NAS consistently outperforms state-of-the-art approaches, highlighting its effectiveness and generalizability.
Qian Guo 0005, Quanchen Su, Xinyan Liang, Nan Li 0033, Zhihua Cui
IEEE Trans. Image Process.1
2025 Trusted Multi-View Classification via Evolutionary Multi-View Fusion
abstract
Multi-view classification based on the Dempster-Shafer theory is widely recognized for its reliability in safety-critical domains with multi-view data. However, the adoption of a late fusion strategy constrains information interaction among views, thereby leading to suboptimal utilization of multi-view data. A recent advancement addressing this limitation involves generating a pseudo view by concatenating individual views. Yet, the efficacy of this pseudo view may diminish when incorporating underperforming views like noisy views. Additionally, the integration of a pseudo view exacerbates the issue of imbalanced multi-view learning, as it contains a disproportionate amount of information compared to individual views. To address these issues, we propose the enhancing Trusted multi-view classification via Evolutionary multi-view Fusion (TEF) approach. TEF employs an evolutionary multi-view architecture search method to create a high-quality fusion architecture serving as the pseudo view, facilitating adaptive view and fusion operator selection. Furthermore, TEF enhances each view within the fusion architecture by concatenating the fusion architecture's decision output with its respective view. Our experimental results demonstrate the effectiveness of this straightforward yet powerful strategy in mitigating imbalanced multi-view learning issues, particularly on complex many-view datasets exceeding three views. Extensive evaluations across 13 multi-view datasets validate the superior performance of our proposed method compared to other trusted multi-view learning approaches. The code is available at https://github.com/fupinhan123/TEF.
Xinyan Liang, Pinhan Fu, Qian Guo 0005, Guoqing Liu 0001
ICLR4
2025 Robust Automatic Modulation Classification with Fuzzy Regularization
abstract
Automatic Modulation Classification (AMC) serves as a foundational pillar for cognitive radio systems, enabling critical functionalities including dynamic spectrum allocation, non-cooperative signal surveillance, and adaptive waveform optimization. However, practical deployment of AMC faces a fundamental challenge: prediction ambiguity arising from intrinsic similarity among modulation schemes and exacerbated under low signal-to-noise ratio (SNR) conditions. This phenomenon manifests as near-identical probability distributions across confusable modulation types, significantly degrading classification reliability. To address this, we propose Fuzzy Regularization-enhanced AMC (FR-AMC), a novel framework that integrates uncertainty quantification into the classification pipeline. The proposed FR has three features: (1) Explicitly model prediction ambiguity during backpropagation, (2) dynamic sample reweighting through adaptive loss scaling, (3) encourage margin maximization between confusable modulation clusters. Experimental results on benchmark datasets demonstrate that the FR achieves superior classification accuracy and robustness compared to compared methods, making it a promising solution for real-world spectrum management and communication applications.
Xinyan Liang, Ruijie Sang, Qian Guo 0005, Feijiang Li, Liang Du 0003
ICML4
2025 Trusted Multi-View Classification with Expert Knowledge Constraints
abstract
Multi-view classification (MVC) based on the Dempster-Shafer theory has gained significant recognition for its reliability in safety-critical applications. However, existing methods predominantly focus on providing confidence levels for decision outcomes without explaining the reasoning behind these decisions. Moreover, the reliance on first-order statistical magnitudes of belief masses often inadequately capture the intrinsic uncertainty within the evidence. To address these limitations, we propose a novel framework termed Trusted Multi-view Classification Constrained with Expert Knowledge (TMCEK). TMCEK integrates expert knowledge to enhance feature-level interpretability and introduces a distribution-aware subjective opinion mechanism to derive more reliable and realistic confidence estimates. The theoretical superiority of the proposed uncertainty measure over conventional approaches is rigorously established. Extensive experiments conducted on three multi-view datasets for sleep stage classification demonstrate that TMCEK achieves state-of-the-art performance while offering interpretability at both the feature and decision levels. These results position TMCEK as a robust and interpretable solution for MVC in safety-critical domains. The code is available at https://github.com/jie019/TMCEK_ICML2025.
Xinyan Liang, Qian Guo 0005, Liang Du 0003, Bingbing Jiang 0001, Tingjin Luo, Feijiang Li
ICML4
2025 A Fast Neural Architecture Search Method for Multi-Modal Classification via Knowledge Sharing
abstract
Neural architecture search-based multi-modal classification (NAS-MMC) aims to automatically find optimal network structures for improving the multi-modal classification performance. However, most current NAS-MMC methods are quite time-consuming during the training process. In this paper, we propose a knowledge sharing-based neural architecture search (KS-NAS) method for multi-modal classification. The KS-NAS optimizes the search process by introducing a dynamically updated knowledge base to reduce the consumption of computational resource. Specifically, during the deep evolutionary search, individuals in the initial population acquire initial parameters from a knowledge base, and then undergo training and optimization until convergence is reached, avoiding the need for training from scratch. The knowledge base is dynamically updated by aggregating the parameters of high-quality individuals trained within the population, thus progressively improving the quality of the knowledge base. As the population evolves, the knowledge base continues to optimize, ensuring that subsequent individuals can obtain higher-quality initialization parameters, which significantly accelerates the training speed of the population. Experimental results show that the KS-NAS method achieves state-of-the-art results in terms of classification performance and training efficiency across multiple popular multi-modal tasks.
Zhihua Cui, Shiwu Sun, Qian Guo 0005, Xinyan Liang, Zhixia Zhang
IJCAI3
2025 View-Association-Guided Dynamic Multi-View Classification
abstract
In multi-view classification tasks, integrating information from multiple views effectively is crucial for improving model performance. However, most existing methods fail to fully leverage the complex relationships between views, often treating them independently or using static fusion strategies. In this paper, we propose a View-Association-Guided Dynamic Multi-View Classification method (AssoDMVC) to address these limitations. Our approach dynamically models and incorporates the relationships between different views during the classification process. Specifically, we introduce a view-relation-guided mechanism that captures the dependencies and interactions between views, allowing for more flexible and adaptive feature fusion. This dynamic fusion strategy ensures that each view contributes optimally based on its contextual relevance and the inter-view relationships. Extensive experiments on multiple benchmark datasets demonstrate that our method outperforms traditional multi-view classification techniques, offering a more robust and efficient solution for tasks involving complex multi-view data.
Xinyan Liang, Qian Guo 0005, Bingbing Jiang 0001, Feijiang Li, Liang Du 0003, Lu Chen 0003
IJCAI3
2025 An Association-based Fusion Method for Speech Enhancement
abstract
Deep learning-based speech enhancement (SE) methods predominantly draw upon two architectural frameworks: generative adversarial networks and diffusion models. In the realm of SE, capturing the local and global relations between signal frames is crucial for the success of these methods. These frameworks typically employ a UNet architecture as their foundational backbone, integrating Long Short-Term Memory (LSTM) networks or attention mechanisms within the UNet to effectively model both local and global signal relations. However, the coupled relation modeling way may not fully harness the potential of these relations. In this paper, we propose an innovative Association-based Fusion Speech Enhancement method (AFSE), a decoupled method. AFSE first constructs a graph that encapsulates the association between each time window of the speech signal, and then models the global relations between frames by fusing the features of these time windows in a manner akin to graph neural networks. Furthermore, AFSE leverages a UNet with dilated convolutions to model the local relations, enabling the network to maintain a high-resolution representation while benefiting from a wider receptive field. Experimental results demonstrate that the AFSE method significantly improves performance in speech enhancement tasks, validating the effectiveness and superiority of our approach. The code is available at https://github.com/jie019/AFSE_IJCAI2025.
Qian Guo 0005, Lu Chen 0003, Liang Du 0003, Zikun Jin, Zhian Yuan, Xinyan Liang
IJCAI2
2025 Improving Evolutionary Multi-View Classification via Eliminating Individual Fitness Bias
abstract
Evolutionary multi-view classification (EMVC) methods have gained wide recognition due to their adaptive mechanisms. Fitness evaluation (FE), which aims to calculate the classification performance of each individual in the population and provide reliable performance ranking for subsequent operations, is a core step in such methods. Its accuracy directly determines the correctness of the evolutionary direction. That is, when FE fails to correctly reflect the superiority-inferiority relationship among individuals, it will lead to confusion in individual performance ranking, which in turn misleads the evolutionary direction and results in trapping into local optima. This paper is the first to identify the aforementioned issue in the field of EMVC and call it as fitness evaluation bias (FEB). FEB may be caused by a variety of factors, and this paper approaches the issue from the perspective of view information content: existing methods generally adopt joint training strategies, which restrict the exploration of key information in views with low information content. This makes it difficult for multi-view model (MVM) to achieve optimal performance during convergence, which in turn leads to FE failing to accurately reflect individual performance rankings and ultimately triggering FEB. To address this issue, we propose an evolutionary multi-view classification via eliminating individual fitness bias (EFB-EMVC) method, which alleviates the FEB issue by introducing evolutionary navigators for each MVM, thereby providing more accurate individual ranking. Experimental results fully verify the effectiveness of the proposed method in alleviating the FEB problem, and the EMVC method equipped with this strategy exhibits more superior performance compared with the original EMVC method. (The code is available at https://github.com/LiShuailzn/Neurips-2025-EFB-EMVC)
Xinyan Liang, Qian Guo 0005, Bingbing Jiang 0001, Tingjin Luo, Liang Du 0003
NeurIPS3
2025 A data representation method using distance correlation
Xinyan Liang, Qian Guo 0005, Keyin Zheng
Frontiers Comput. Sci.3
2025 Multi-Scale Features Are Effective for Multi-Modal Classification: An Architecture Search Viewpoint
abstract
Multi-modal neural architecture search (MNAS) is an effective approach to obtain task-adaptive multi-modal classification models. Deep neural networks, as currently main-stream feature extractors, can provide hierarchical features for each modality. Existing MNAS methods face difficulty in exploiting such hierarchical features due to their different form coexistence such as tensorial multi-scale features and vectorized penultimate features. Moreover, existing methods always focus on the evolution of fusion operators or vectorized features of all modalities, constraining search space. In this paper, a novel two-stage method called multi-modal multi-scale evolutionary neural architecture search (MM-ENAS) is proposed. The first stage unifies the representation form of hierarchical features by the proposed evolutionary statistics strategy. The second stage identifies the optimal combination of basic fusion operations for all unified hierarchical features by the evolutionary algorithm. MM-ENAS increases search space by simultaneously searching for feature statistical extraction methods, basic fusion operators and feature representation set consisting of tensorial multi-scale features and vectorized penultimate features. Experimental results on three multi-modal tasks demonstrate that the proposed method achieves competitive performance in terms of accuracy, search time, and number of parameters compared to existing representative MNAS methods. Additionally, the method exhibits fast adaptation to various multi-modal tasks.
Pinhan Fu, Xinyan Liang, Qian Guo 0005, Yayu Zhang, Qin Huang 0005, Ke Tang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 DC-NAS: Divide-and-Conquer Neural Architecture Search for Multi-Modal Classification
abstract
Neural architecture search-based multi-modal classification (NAS-MMC) methods can individually obtain the optimal classifier for different multi-modal data sets in an automatic manner. However, most existing NAS-MMC methods are dramatically time consuming due to the requirement for training and evaluating enormous models. In this paper, we propose an efficient evolutionary-based NAS-MMC method called divide-and-conquer neural architecture search (DC-NAS). Specifically, the evolved population is first divided into k+1 sub-populations, and then k sub-populations of them evolve on k small-scale data sets respectively that are obtained by splitting the entire data set using the k-fold stratified sampling technique; the remaining one evolves on the entire data set. To solve the sub-optimal fusion model problem caused by the training strategy of partial data, two kinds of sub-populations that are trained using partial data and entire data exchange the learned knowledge via two special knowledge bases. With the two techniques mentioned above, DC-NAS achieves the training time reduction and classification performance improvement. Experimental results show that DC-NAS achieves the state-of-the-art results in term of classification performance, training efficiency and the number of model parameters than the compared NAS-MMC methods on three popular multi-modal tasks including multi-label movie genre classification, action recognition with RGB and body joints and dynamic hand gesture recognition.
Xinyan Liang, Pinhan Fu, Qian Guo 0005, Keyin Zheng
AAAI3
2024 Core-Structures-Guided Multi-Modal Classification Neural Architecture Search
Pinhan Fu, Xinyan Liang, Tingjin Luo, Qian Guo 0005, Yayu Zhang
IJCAI4
2024 CoMO-NAS: Core-Structures-Guided Multi-Objective Neural Architecture Search for Multi-Modal Classification
abstract
Most existing NAS-based multi-modal classification (MMC-NAS) methods are optimized using the classification accuracy.They can not simultaneously provide multiple models with diverse perferences such as model complex and classification performance for meeting different users' demands. Combining NAS-MMC with multi-objective optimization is a nature way for this issue. However, the challenge problem of this solution is the high computation cost. For multi-objective optimization, the computing bottleneck is pareto front search. Some higher-quality MMC models (namely core structures, CSs) consisting of high-quality features and fusion operators are easier to identify. We find that CSs have a close relation with the pareto front (PF), i.e., the individuals lying in PF contain the CSs. Based on the finding, we propose an efficient multi-objective neural architecture search for multi-modal classification by applying CSs to guide the PF search (CoMO-NAS). In conclusion, experimental results thoroughly demonstrate the effectiveness of our CoMO-NAS. Compared to state-of-the-art competitors on benchmark multi-modal tasks, we achieve comparable performance with lower model complexity in shorter search time.
Pinhan Fu, Xinyan Liang, Qian Guo 0005, Zhifang Wei, Wen Li 0013
ACM Multimedia4
2024 A Progressive Skip Reasoning Fusion Method for Multi-Modal Classification
abstract
In multi-modal classification tasks, a good fusion algorithm can effectively integrate and process multi-modal data, thereby significantly improving its performance. Researchers often focus on the design of complex fusion operators and have proposed numerous fusion operators, while paying less attention to the design of feature fusion usage, specifically how features should be fused to better facilitate multi-modal classification tasks. In this article, we propose a progressive skip reasoning fusion network (PSRFN) to make some attempts to address this issue. Firstly, unlike most existing multi-modal fusion methods that only use one fusion operator in a single stage to fuse all view features, PSRFN utilizes the progressive skip reasoning (PSR) block to fuse all views with a fusion operator at each layer. Specifically, each PSR block utilizes all view features and the fused features from the previous layer to jointly obtain the fused features for the current layer. Secondly, each PSR block utilizes a dual-weighted fusion strategy with learnable parameters to adaptively allocate weights during the fusion process. The first level of weighting assigns weights to each view feature, while the second level assigns weights to the fused features from the previous layer and the fused features obtained from the first level of weighting in the current layer. This strategy ensures that the PSR block can dynamically adjust the weights based on the actual contribution of features. Finally, to enable the model to fully utilize feature information from different levels for feature fusion, the skip connections are adopted between PSR blocks. Extensive experiment results on six real multi-modal datasets show that a better usage for fusion operator is indeed able to improve performance.
Qian Guo 0005, Xinyan Liang, Zhihua Cui, Jie Wen 0008
ACM Multimedia1
2022 GLRM: Logical pattern mining in the case of inconsistent data distribution based on multigranulation strategy
Qian Guo 0005, Xinyan Liang
Int. J. Approx. Reason.1
2022 Corrigendum to "GLRM: Logical pattern mining in the case of inconsistent data distribution based on multigranulation strategy" [Int. J. Approx. Reason. 143 (2022) 78-101]
Qian Guo 0005, Xinyan Liang
Int. J. Approx. Reason.1
2022 AF: An Association-Based Fusion Method for Multi-Modal Classification
abstract
Multi-modal classification (MMC) aims to integrate the complementary information from different modalities to improve classification performance. Existing MMC methods can be grouped into two categories: traditional methods and deep learning-based methods. The traditional methods often implement fusion in a low-level original space. Besides, they mostly focus on the inter-modal fusion and neglect the intra-modal fusion. Thus, the representation capacity of fused features induced by them is insufficient. The deep learning-based methods implement the fusion in a high-level feature space where the associations among features are considered, while the whole process is implicit and the fused space lacks interpretability. Based on these observations, we propose a novel interpretative association-based fusion method for MMC, named AF. In AF, both the association information and the high-order information extracted from feature space are simultaneously encoded into a new feature space to help to train an MMC model in an explicit manner. Moreover, AF is a general fusion framework, and most existing MMC methods can be embedded into it to improve their performance. Finally, the effectiveness and the generality of AF are validated on 22 datasets, four typically traditional MMC methods adopting best modality, early, late and model fusion strategies and a deep learning-based MMC method.
Xinyan Liang, Qian Guo 0005, Honghong Cheng, Jiye Liang
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Image deep clustering based on local-topology embedding
Feijiang Li, Qian Guo 0005
Pattern Recognit. Lett.4
2021 Evolutionary Deep Fusion Method and its Application in Chemical Structure Recognition
abstract
Feature extraction is a critical issue in many machine learning systems. A number of basic fusion operators have been proposed and studied. This article proposes an evolutionary algorithm, called evolutionary deep fusion method, for searching an optimal combination scheme of different basic fusion operators to fuse multiview features. We apply our proposed method to chemical structure recognition. Our proposed method can directly take images as inputs, and users do not need to transform images to other formats. The experimental results demonstrate that our proposed method can achieve a better performance than those designed by human experts on this real-life problem.
Xinyan Liang, Qian Guo 0005, Weiping Ding 0001, Qingfu Zhang 0001
IEEE Trans. Evol. Comput.2
2019 Diversity-induced fuzzy clustering
Honghong Cheng, Qian Guo 0005
Int. J. Approx. Reason.4
2018 Local neighborhood rough set
Xinyan Liang, Qian Guo 0005, Jiye Liang
Knowl. Based Syst.4
2017 Local multigranulation decision-theoretic rough sets
Xinyan Liang, Guoping Lin, Qian Guo 0005, Jiye Liang
Int. J. Approx. Reason.4