Zhiqiang Kou

dblp:341/1380 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-3335-3959ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021
YearPublicationVenuePosition
2026 LoGoSeg: Integrating Local and Global Features for Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation (OVSS) extends traditional closed-set segmentation by enabling pixel-wise annotation for both seen and unseen categories using arbitrary textual descriptions. While existing methods leverage vision-language models (VLMs) like CLIP, their reliance on image-level pretraining often results in imprecise spatial alignment, leading to mismatched segmentations in ambiguous or cluttered scenes. However, most existing approaches lack strong object priors and region-level constraints, which can lead to object hallucination or missed detections, further degrading performance. To address these challenges, we propose LoGoSeg, an efficient single-stage framework that integrates three key innovations: (i) an object existence prior that dynamically weights relevant categories through global image-text similarity, effectively reducing hallucinations; (ii) a region-aware alignment module that establishes precise region-level visual-textual correspondences; and (iii) a dual-stream fusion mechanism that optimally combines local structural information with global semantic context. Unlike prior works, LoGoSeg eliminates the need for external mask proposals, additional backbones, or extra datasets, ensuring efficiency. Extensive experiments on six benchmarks (A-847, PC-459, A-150, PC-59, PAS-20, and PAS-20b) demonstrate its competitive performance and strong generalization in open-vocabulary settings.
Xiangbo Lv, Zhiqiang Kou, Xingdong Sheng, Yiguo Qiao
AAAI3
2026 ALPSB: Adaptive learngene with plastic and stable branches
Shuxia Lin, Xu Yang 0021, Qiufeng Wang 0002, Shunxin Guo, Zhiqiang Kou, Xin Geng 0001
Pattern Recognit.5
2026 Corrigendum to "ALPSB: Adaptive Learngene with Plastic and Stable Branches" [Pattern Recognition 172 (2026) 112623]
Shuxia Lin, Xu Yang 0021, Qiufeng Wang 0002, Shunxin Guo, Zhiqiang Kou, Xin Geng 0001
Pattern Recognit.5
2026 Tail-Aware Reconstruction of Incomplete Label Distributions With Low-Rank and Sparse Modeling
abstract
Label Distribution Learning (LDL) is a novel machine learning paradigm that addresses the problem of label ambiguity and has found widespread applications. However, obtaining complete label distributions in real-world scenarios is challenging, which has led to the emergence of Incomplete Label Distribution Learning (InLDL). Existing InLDL methods attempt to utilize low-rank label correlations to recover the complete label distribution. However, we find that real-world LDL datasets have animbalancednature; that is, the sum of the description degrees for normal labels is significantly larger than that for tail labels, which disrupts the low-rank assumption underlying the recovery of the label distribution. To solve the above problem, we propose Incomplete and Imbalance Label Distribution Learning (I2LDL), which makes the use of low-rank label correlations more reasonable for InLDL. Our method decomposes the recovered label distribution matrix into a low-rank component for frequent labels and a sparse component for tail labels, effectively capturing the structure of both head and tail labels. We further require that the entries in the observed positions of the recovered label distribution matrix be close to the observed values, and that the recovered label distribution for every instance forms a probability simplex (i.e., nonnegative entries summing to unity). Finally, the proposed model is optimized via the Alternating Direction Method of Multipliers (ADMM). We provide a theoretical analysis of its exact recovery guarantee under standard assumptions of incoherence, sparsity, and sufficient sampling. Furthermore, we establish a generalization error bound based on Rademacher complexity, offering theoretical insights into the learning performance of our method. Extensive experiments on 16 real-world datasets demonstrate the effectiveness and robustness of our framework compared to existing InLDL methods. The code is available at https://anonymous.4open.science/r/IncomLDL-tailaware-C021.
Zhiqiang Kou, Haoyuan Xuan, Hailin Wang 0001, Ming-Kun Xie, Changwei Wang 0001, Jing Wang 0113, Yuheng Jia, Xin Geng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Object-Level Correlation for Few-Shot Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment objects of novel categories in the query images given only a few annotated support samples. Existing methods primarily build the image-level correlation between the support target object and the entire query image. However, this correlation contains the hard pixel noise, \textit{i.e.}, irrelevant background objects, that is intractable to trace and suppress, leading to the overfitting of the background. To address the limitation of this correlation, we imitate the biological vision process to identify novel objects in the object-level information. Target identification in the general objects is more valid than in the entire image, especially in the low-data regime. Inspired by this, we design an Object-level Correlation Network (OCNet) by establishing the object-level correlation between the support target object and query general objects, which is mainly composed of the General Object Mining Module (GOMM) and Correlation Construction Module (CCM). Specifically, GOMM constructs the query general object feature by learning saliency and high-level similarity cues, where the general objects include the irrelevant background objects and the target foreground object. Then, CCM establishes the object-level correlation by allocating the target prototypes to match the general object feature. The generated object-level correlation can mine the query target feature and suppress the hard pixel noise for the final prediction. Extensive experiments on PASCAL-${5}^{i}$ and COCO-${20}^{i}$ show that our model achieves the state-of-the-art performance.
Chunlin Wen, Yu Zhang 0004, Hongyuan Zhu 0002, Xiu-Shen Wei, Zhiqiang Kou, Shuzhou Sun
ICCV7
2025 Label Distribution Learning with Biased Annotations Assisted by Multi-Label Learning
abstract
Multi-label learning (MLL) has gained attention for its ability to represent real-world data. Label Distribution Learning (LDL), an extension of MLL to learning from label distributions, faces challenges in collecting accurate label distributions. To address the issue of biased annotations, based on the low-rank assumption, existing works recover true distributions from biased observations by exploring the label correlations. However, recent evidence shows that the label distribution tends to be full-rank, and naive apply of low-rank approximation on biased observation leads to inaccurate recovery and performance degradation. In this paper, we address the LDL with biased annotations problem from a novel perspective, where we first degenerate the soft label distribution into a hard multi-hot label and then recover the true label information for each instance. This idea stems from an insight that assigning hard multi-hot labels is often easier than assigning a soft label distribution, and it shows stronger immunity to noise disturbances, leading to smaller label bias. Moreover, assuming that the multi-label space for predicting label distributions is low-rank offers a more reasonable approach to capturing label correlations. Theoretical analysis and experiments confirm the effectiveness and robustness of our method on real-world datasets.
Zhiqiang Kou, Si Qin, Hailin Wang 0001, Jing Wang 0113, Ming-Kun Xie, Shuo Chen 0003, Yuheng Jia, Tongliang Liu, Masashi Sugiyama, Xin Geng 0001
IJCAI1
2025 Collaboration Wins More: Dual-Modal Collaborative Attention Reinforcement for Mitigating Large Vision Language Models Hallucination
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual-language understanding for downstream multimodal tasks. However, these models often generate descriptions containing objects or details not present in the input image, a phenomenon commonly referred to as ''hallucination''. Existing methods focus solely on single-side hallucination mitigation: Intra-modal-only reinforcement (e.g. visual attention enhancement) ignores prompt-based guidance; Inter-modal-only correlation correction may introduce low-information visual tokens to mislead reasoning. To tackle this challenge, we propose Dual-Modal Collaborative Attention Reinforcement (DuCAR). Specifically, DuCAR is equipped with intra-visual CLS-driven sampling and cross-modal dynamic sampling, extracting important visual tokens guided by intra- and inter-modal joint information. During the multimodal fusion stage, DuCAR adaptively enhances the attention weights of these visual tokens. Our sampling and enhancement strategies in DuCAR simultaneously reinforces informative visual tokens, and suppresses attention dispersion towards question-irrelevant visual information. We conduct extensive experiments on the POPE and CHAIR hallucination benchmarks, demonstrating that our method outperforms existing state-of-the-art mitigation baselines and effectively reduces hallucinations in text generated by LVLMs. The code is available in the https://github.com/xjy2020/DuCAR.
Jiye Xie, Liangliang You, Zhiqiang Kou, Kexue Fu 0001, Youyang Qu, Wenjie Yang 0005, Jianwei Guo 0003, Weiliang Meng, Longxiang Gao, Haoran Yang 0003, Changwei Wang 0001, Yu Zhang 0133
ACM Multimedia6
2025 RankMatch: A Novel Approach to Semi-Supervised Label Distribution Learning Leveraging Rank Correlation between Labels
abstract
Pseudo label based semi-supervised learning (SSL) for single-label and multi-label classification tasks has been extensively studied; however, semi-supervised label distribution learning (SSLDL) remains a largely unexplored area. Existing SSL methods fail in SSLDL because the pseudo-labels they generate only ensure overall similarity to the ground truth but do not preserve the ranking relationships between true labels, as they rely solely on KL divergence as the loss function during training. These skewed pseudo-labels lead the model to learn incorrect semantic relationships, resulting in reduced performance accuracy. To address these issues, we propose a novel SSLDL method called \textit{RankMatch}. \textit{RankMatch} fully considers the ranking relationships between different labels during the training phase with labeled data to generate higher-quality pseudo-labels. Furthermore, our key observation is that a flexible utilization of pseudo-labels can enhance SSLDL performance. Specifically, focusing solely on the ranking relationships between labels while disregarding their margins helps prevent model overfitting. Theoretically, we prove that incorporating ranking correlations enhances SSLDL performance and establish generalization error bounds for \textit{RankMatch}. Finally, extensive real-world experiments validate its effectiveness.
Zhiqiang Kou, Yucheng Xie, Hailin Wang 0001, Jing Wang 0113, Ming-Kun Xie, Shuo Chen 0003, Yuheng Jia, Tongliang Liu, Xin Geng 0001
NeurIPS1
2025 Progressive label enhancement
Zhiqiang Kou, Jing Wang 0113, Yuheng Jia, Xin Geng 0001
Pattern Recognit.1
2025 Label enhancement by manifold fusion of feature and label spaces
Jing Wang 0113, Zhiqiang Kou, Yuheng Jia, Jianhui Lv, Xin Geng 0001
Pattern Recognit.2
2025 Addressing Skewed Heterogeneity via Federated Prototype Rectification With Personalization
abstract
Federated learning (FL) is an efficient framework designed to facilitate collaborative model training across multiple distributed devices while preserving user data privacy. A significant challenge of FL is data-level heterogeneity, i.e., skewed or long-tailed distribution of private data. Although various methods have been proposed to address this challenge, most of them assume that the underlying global data are uniformly distributed across all clients. This article investigates data-level heterogeneity FL with a brief review and redefines a more practical and challenging setting called skewed heterogeneous FL (SHFL). Accordingly, we propose a novel federated prototype rectification with personalization (FedPRP) which consists of two parts: federated personalization and federated prototype rectification. The former aims to construct balanced decision boundaries between dominant and minority classes based on private data, while the latter exploits both interclass discrimination and intraclass consistency to rectify empirical prototypes. Experiments on three popular benchmarks show that the proposed approach outperforms current state-of-the-art methods and achieves balanced performance in both personalization and generalization.
Shunxin Guo, Hongsong Wang 0001, Shuxia Lin, Zhiqiang Kou, Xin Geng 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Instance-Dependent Inaccurate Label Distribution Learning
abstract
Label distribution learning (LDL) is a novel learning paradigm that assigns each instance with a label distribution. Although many specialized LDL algorithms have been proposed, few of them have noticed that the obtained label distributions are generally inaccurate with noise due to the difficulty of annotation. Besides, existing LDL algorithms overlooked that the noise in the inaccurate label distributions generally depends on instances. In this article, we identify the instance-dependent inaccurate LDL (IDI-LDL) problem and propose a novel algorithm called low-rank and sparse LDL (LRS-LDL). First, we assume that the inaccurate label distribution consists of the ground-truth label distribution and instance-dependent noise. Then, we learn a low-rank linear mapping from instances to the ground-truth label distributions and a sparse mapping from instances to the instance-dependent noise. In the theoretical analysis, we establish a generalization bound for LRS-LDL. Finally, in the experiments, we demonstrate that LRS-LDL can effectively address the IDI-LDL problem and outperform existing LDL methods.
Zhiqiang Kou, Jing Wang 0113, Yuheng Jia, Xin Geng 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 Label Distribution Learning by Exploiting Fuzzy Label Correlation
abstract
Researchers have proposed to exploit label correlation to alleviate the exponential-size output space of label distribution learning (LDL). In particular, some have designed LDL methods to consider local label correlation. These methods roughly partition the training set into clusters and then exploit local label correlation on each one. Each sample belongs to one cluster and therefore has only one local label correlation. However, in real-world scenarios, the training samples may have fuzziness and belong to multiple clusters with blended local label correlations, which challenge these works. To solve this problem, we propose in LDL fuzzy label correlation (FLC)-each sample blends, with fuzzy membership, multiple local label correlations. First, we propose two types of FLCs, i.e., fuzzy membership-induced label correlation (FC) and joint fuzzy clustering and label correlation (FCC). Then, we put forward LDL-FC and LDL-FCC to exploit these two FLCs, respectively. Finally, we conduct extensive experiments to justify that LDL-FC and LDL-FCC statistically outperform state-of-the-art LDL methods.
Jing Wang 0113, Zhiqiang Kou, Yuheng Jia, Jianhui Lv, Xin Geng 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Building Variable-Sized Models via Learngene Pool
abstract
Recently, Stitchable Neural Networks (SN-Net) is proposed to stitch some pre-trained networks for quickly building numerous networks with different complexity and performance trade-offs. In this way, the burdens of designing or training the variable-sized networks, which can be used in application scenarios with diverse resource constraints, are alleviated. However, SN-Net still faces a few challenges. 1) Stitching from multiple independently pre-trained anchors introduces high storage resource consumption. 2) SN-Net faces challenges to build smaller models for low resource constraints. 3). SN-Net uses an unlearned initialization method for stitch layers, limiting the final performance. To overcome these challenges, motivated by the recently proposed Learngene framework, we propose a novel method called Learngene Pool. Briefly, Learngene distills the critical knowledge from a large pre-trained model into a small part (termed as learngene) and then expands this small part into a few variable-sized models. In our proposed method, we distill one pre-trained large model into multiple small models whose network blocks are used as learngene instances to construct the learngene pool. Since only one large model is used, we do not need to store more large models as SN-Net and after distilling, smaller learngene instances can be created to build small models to satisfy low resource constraints. We also insert learnable transformation matrices between the instances to stitch them into variable-sized models to improve the performance of these models. Exhaustive experiments have been implemented and the results validate the effectiveness of the proposed Learngene Pool compared with SN-Net.
Boyu Shi, Shiyu Xia, Xu Yang 0021, Zhiqiang Kou, Xin Geng 0001
AAAI5
2024 Exploiting Multi-Label Correlation in Label Distribution Learning
Zhiqiang Kou, Jing Wang 0113, Yuheng Jia, Boyu Shi, Xin Geng 0001
IJCAI1
2024 Inaccurate Label Distribution Learning
abstract
Label distribution learning (LDL) trains a model to predict the relevance of a set of labels (called label distribution (LD)) to an instance. The previous LDL methods all assumed the LDs of the training instances are accurate. However, annotating highly accurate LDs for training instances is time-consuming and extremely expensive, and in reality the collected LDs are often inaccurate. This paper first investigates the inaccurate LDL (ILDL) problem—learn an LDL method from the inaccurate LDs. We assume that the inaccurate LD blends the ground-truth LD and sparse noise. Consequently, the ILDL problem becomes an inverse problem, whose objective is to recover the ground-truth LD and noise from the inaccurate LD. We hypothesize that the ground-truth LD exhibits low rank due to label correlations. Besides, we leverage the local geometric structure of instances (represented as graph) to further recover the ground-truth LD. Finally, the proposed method is formulated as a graph-regularized low-rank and sparse decomposition problem. Next, we induce an LDL predictive method by learning from recovered LD. Extensive experiments conducted on multiple datasets demonstrate the better performance of our method, especially for ILDL problem.
Zhiqiang Kou, Jing Wang 0113, Yuheng Jia, Xin Geng 0001
IEEE Trans. Circuits Syst. Video Technol.1