Kai Niu 0002

dblp:67/229-2 · DBLP profile ↗
← Back
16ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-7997-9930ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cache-aided cross-modal correlation correction for unsupervised cross-domain text-based person search
Kai Niu 0002, Qinzi Zhao
Pattern Recognit.1
2026 Pseudo Sentences Evaluation and Quality-Aware Robust Learning for Unsupervised Text-Based Person Search
abstract
Unsupervised Text-Based Person Search (TBPS) eliminates the need for costly manual sentence annotations by generating pseudo sentences via Multi-modal Large Language Models (MLLMs). However, these pseudo sentences often face the quality defect issues, resulting in semantic misalignment across modalities, which will hinder discriminative representation learning. To address this problem, we propose the PSE-QRL (Pseudo Sentences Evaluation and Quality-aware Robust Learning), a unified framework that enhances robustness to pseudo sentences for unsupervised TBPS. The PSE-QRL dynamically couples an evolving TBPS model with MLLMs to assess pseudo sentences' reliability, and adaptively leverages high-quality ones during training. It consists of three key components: 1) Multi-granularity Sentence Augmentation, for enriching pseudo sentences with multiple granularities to broaden the diversity of image-sentence pairs; 2) Hybrid Quality Evaluation, to combine MLLM's cross-modal reasoning knowledge with TBPS model's person-specific distinguishing capabilities for effective sentence quality assessment; and 3) Quality-aware Robust Learning, for selecting and re-weighting samples based on quality scores to emphasize reliable sentence annotations while suppressing low-quality ones. Extensive experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReid benchmarks demonstrate the effectiveness of PSE-QRL for improving learning robustness, achieving state-of-the-art (SOTA) retrieval performance for unsupervised TBPS.
Kai Niu 0002, Xinyue Song, Yanning Zhang 0001
IEEE Trans. Image Process.1
2025 Test-Time Adaptation for Text-Based Person Search
abstract
Text-based person search (TBPS), aiming to retrieve target pedestrian images with natural language descriptions, has seen significant progress in recent years. However, severe domain shift remains a key challenge in this field, causing source-domain-trained models to degrade significantly when applied to an unseen target domain. To address this, we propose the Identity-preserving Cross-modal Alignment and Adaptation (ICAA) model, a novel test-time adaptation framework for TBPS that enables seamless domain adaptation using only unlabeled target samples. Our method tackles two key challenges: 1) Cross-modal domain-shift misalignment: textual and visual modalities exhibit inconsistent distributional shifts across domains. To this end, our Cross-Modal Alignment adaptation (CMA) module identifies pseudo-positive image-text pairs and minimizes their matching discrepancies in the target domain, adapting to new cross-modal distribution relationships. 2) Identity semantic absence: crucial identity annotations are usually unavailable in both target text and image data. To mitigate this, we introduce the Identity-Preserving Dynamic adaptation (IPD) module, which dynamically associates image-text pairs with potential identity prototypes to enhance identity consistency in cross-modal alignment during adaptation. Our method is simple yet effective, establishing new state-of-the-art cross-domain results for TBPS on three public benchmarks, i.e., CUHK-PEDES, ICFG-PEDES, and RSTPReid.
Kai Niu 0002, Liucun Shi, Qinzi Zhao, Yanning Zhang 0001
ACM Multimedia1
2024 An Adaptive Correlation Filtering Method for Text-Based Person Search
Mengyang Sun, Wei Suo, Peng Wang 0015, Kai Niu 0002, Le Liu 0008, Guosheng Lin, Yanning Zhang 0001, Qi Wu 0001
Int. J. Comput. Vis.4
2024 Contrastive Pedestrian Attentive and Correlation Learning Network for Occluded Person Re-Identification
abstract
Occluded person Re-identification (ReID) aims to match occluded and holistic pedestrian images across different camera views. This task presents two primary challenges. First, it is crucial to accurately capture pedestrian foregrounds from seriously occluded person images. Second, a noticeable information asymmetry exists between the partial body in occluded images and the complete body in corresponding holistic images, which could cause the ReID model to underestimate their similarities. To address these challenges, we introduce a contrastive pedestrian attentive and correlation learning (CpaCol) model. Within CpaCol, we first design a Contrastive Pedestrian Attention (ContrastAttn) module to capture pedestrian foregrounds from occluded images. In this process, we notice that most existing attention-based methods only supervise the final predictions with identity loss yet neglect its causality with the generated attention maps, which could mislead the model to capture some salient yet pedestrian-irrelevant noises as discriminative clues. To rectify this, we integrate contrastive learning into our ContrastAttn module to guide it to learn the semantic divergence between pedestrian foregrounds and noises, thereby capturing pedestrian foregrounds more accurately. Besides, we propose a correlation learning module, where we tailor an effective dense feature correlation learning tool, 4D convolution, to enable it to adapt to pedestrian images and capture corresponding clues between comparing images. By focusing more on corresponding clues, our model could avoid overemphasizing the inherent information asymmetry between occluded and holistic images, thereby improving re-identification. Empowered by these modules, our CpaCol achieves state-of-the-art performance on three relevant ReID settings,i.e., occluded, partial, and holistic ReID. Our code is available in https://github.com/nwpugaoliying/CpaCol.
Liying Gao, Bingliang Jiao, Yuzhou Long, Kai Niu 0002, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 An Overview of Text-Based Person Search: Recent Advances and Future Directions
abstract
Due to the practical significance in smart video surveillance systems, Text-Based Person Search (TBPS) has been one of the research hotspots recently, which refers to searching for the interested pedestrian images given natural language sentences. To help researchers quickly grasp the developments of this important task, we comprehensively summarize the recent research advances of TBPS from two perspectives,i.e., Feature Extraction (FE) and Semantic Alignments (SA). Specifically, the FE mainly consists of pre-processing approaches and end-to-end frameworks, and the SA could be briefly divided into cross-modal attention mechanism, non-attention alignments, training objectives, and generative approaches. Afterwards, we elaborate four widely-used benchmarks and also the evaluation criterion for TBPS. And comparisons and analyses among the state-of-the-art (SOTA) solutions are provided based on these large-scale benchmarks. At last, we point out some future research directions that need to be further addressed, which will greatly facilitate the practical applications of TBPS.
Kai Niu 0002, Yanyi Liu, Yuzhou Long, Yan Huang 0008, Liang Wang 0001, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Comprehensive Attribute Prediction Learning for Person Search by Language
abstract
Person search by language refers to searching for the interested pedestrian images given natural language sentences, which requires capturing fine-grained differences to accurately distinguish different pedestrians, while still far from being well addressed by most of the current solutions. In this paper, we propose the Comprehensive Attribute Prediction Learning (CAPL) method, which explicitly carries out attribute prediction learning, for improving the modeling capabilities of fine-grained semantic attributes and obtaining more discriminative visual and textual representations. First, we construct the semantic ATTribute Vocabulary (ATT-Vocab) based on sentence analysis. Second, the complementary context-wise and attribute-wise attribute predictions are simultaneously conducted to better model the high-frequency in-vocab attributes in our In-vocab Attribute Prediction (IAP) module. Third, to additionally consider the out-of-vocab semantics, we present the Attribute Completeness Learning (ACL) module for better capturing the low-frequency attributes outside the ATT-Vocab, obtaining more comprehensive representations. Combining the IAP and ACL modules together, our CAPL method has obtained the currently state-of-the-art retrieval performance on two widely-used benchmarks, i.e., CUHK-PEDES and ICFG-PEDES datasets. Extensive experiments and analyses have been carried out to validate the effectiveness and generalization capacities of our CAPL method.
Kai Niu 0002, Linjiang Huang, Yuzhou Long, Yan Huang 0008, Liang Wang 0001, Yanning Zhang 0001
IEEE Trans. Image Process.1
2023 Addressing Information Inequality for Text-Based Person Search via Pedestrian-Centric Visual Denoising and Bias-Aware Alignments
abstract
Text-based person search is an important task in video surveillance, which aims to retrieve the corresponding pedestrian images with a given description. In this fine-grained retrieval task, accurate cross-modal information matching is an essential yet challenging problem. However, existing methods usually ignore the information inequality between modalities, which could introduce great difficulties to cross-modal matching. Specifically, in this task, the images inevitably contain some pedestrian-irrelevant noise like background and occlusion, and the descriptions could be biased to partial pedestrian content in images. With that in mind, in this paper, we propose a Text-Guided Denoising and Alignment (TGDA) model to alleviate the information inequality and realize effective cross-modal matching. In TGDA, we first design a prototype-based denoising module, which integrates pedestrian knowledge from textual features into a prototype vector and uses it as guidance to filter out pedestrian-irrelevant noise from visual features. Thereafter, a bias-aware alignment module is introduced, which guides our model to focus on the description-biased pedestrian content in cross-modal features consistently. Through extensive experiments, the effectiveness of both modules has been validated. Besides, our TGDA achieves state-of-the-art performance on various related benchmarks.
Liying Gao, Kai Niu 0002, Bingliang Jiao, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Improving Inconspicuous Attributes Modeling for Person Search by Language
abstract
Person search by language aims to retrieve the interested pedestrian images based on natural language sentences. Although great efforts have been made to address the cross-modal heterogeneity, most of the current solutions suffer from only capturing salient attributes while ignoring inconspicuous ones, being weak in distinguishing very similar pedestrians. In this work, we propose the Adaptive Salient Attribute Mask Network (ASAMN) to adaptively mask the salient attributes for cross-modal alignments, and therefore induce the model to simultaneously focus on inconspicuous attributes. Specifically, we consider the uni-modal and cross-modal relations for masking salient attributes in the Uni-modal Salient Attribute Mask (USAM) and Cross-modal Salient Attribute Mask (CSAM) modules, respectively. Then the Attribute Modeling Balance (AMB) module is presented to randomly select a proportion of masked features for cross-modal alignments, ensuring the balance of modeling capacity of both salient attributes and inconspicuous ones. Extensive experiments and analyses have been carried out to validate the effectiveness and generalization capacity of our proposed ASAMN method, and we have obtained the state-of-the-art retrieval performance on the widely-used CUHK-PEDES and ICFG-PEDES benchmarks.
Kai Niu 0002, Linjiang Huang, Liang Wang 0001, Yanning Zhang 0001
IEEE Trans. Image Process.1
2022 A Simple and Robust Correlation Filtering Method for Text-Based Person Search
Wei Suo, Mengyang Sun, Kai Niu 0002, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001
ECCV (35)3
2022 Cross-modal Co-occurrence Attributes Alignments for Person Search by Language
abstract
Person search by language refers to retrieving the interested pedestrian images based on a free-form natural language description, which has important applications in smart video surveillance. Although great efforts have been made to align images with sentences, the challenge of reporting bias, i.e., attributes are only partially matched across modalities, still incurs large noise and influences the accurate retrieval seriously. To address this challenge, we propose a novel cross-modal matching method named Cross-modal Co-occurrence Attributes Alignments (C2A2), which can better deal with noise and obtain significant improvements in retrieval performance for person search by language. First, we construct visual and textual attribute dictionaries relying on matrix decomposition, and carry out cross-modal alignments using denoising reconstruction features to address the noise from pedestrian-unrelated elements. Second, we re-gather pixels of image and words of sentence under the guidance of learned attribute dictionaries, to adaptively constitute more discriminative co-occurrence attributes in both modalities. And the re-gathered co-occurrence attributes are carefully captured by imposing explicit cross-modal one-to-one alignments which consider relations across modalities, better alleviating the noise from non-correspondence attributes. The whole C_2A_2 method can be trained end-to-end without any pre-processing, i.e., requiring negligible additional computation overheads. It significantly outperforms the existing solutions, and finally achieves the new state-of-the-art retrieval performance on two large-scale benchmarks, CUHK-PEDES and RSTPReid datasets.
Kai Niu 0002, Linjiang Huang, Yan Huang 0008, Peng Wang 0015, Liang Wang 0001, Yanning Zhang 0001
ACM Multimedia1
2022 Actor and Action Modular Network for Text-Based Video Segmentation
abstract
Text-based video segmentation aims to segment an actor in video sequences by specifying the actor and its performing action with a textual query. Previous methods fail to explicitly align the video content with the textual query in a fine-grained manner according to the actor and its action, due to the problem of semantic asymmetry. The semantic asymmetry implies that two modalities contain different amounts of semantic information during the multi-modal fusion process. To alleviate this problem, we propose a novel actor and action modular network that individually localizes the actor and its action in two separate modules. Specifically, we first learn the actor-/action-related content from the video and textual query, and then match them in a symmetrical manner to localize the target tube. The target tube contains the desired actor and action which is then fed into a fully convolutional network to predict segmentation masks of the actor. Our method also establishes the association of objects cross multiple frames with the proposed temporal proposal aggregation mechanism. This enables our method to segment the video effectively and keep the temporal consistency of predictions. The whole model is allowed for joint learning of the actor-action matching and segmentation, as well as achieves the state-of-the-art performance for both single-frame segmentation and full video segmentation on A2D Sentences and J-HMDB Sentences datasets.
Yan Huang 0008, Kai Niu 0002, Linjiang Huang, Zhanyu Ma, Liang Wang 0001
IEEE Trans. Image Process.3
2021 Text-Guided Visual Feature Refinement for Text-Based Person Search
abstract
Text-based person search is a task to retrieve the corresponding person in a large-scale image database given a textual description, which has important value in various fields like video surveillance. In the inferring phase, language descriptions, serving as queries, guide to search the corresponding person images. Most existing methods apply cross-modal signals to guide feature refinement. However, they employ visual features from the gallery to refine textual features, which may cause high similarity between unmatched pairs. Besides, the similarity-based cross-modal attention could disturb the choice of interested areas for descriptions. In this paper, we analyze the deficiency of previous methods and carefully design a Text-guided Visual Feature Refinement network (TVFR), which utilizes text as reference to refine visual representations. Firstly, we divide each visual feature into several horizontal stripes for fine-grained refinement. After that, we employ a text-based filter generation module to generate description-customized filters, which are used to indicate the corresponding stripes mentioned in the textual input. Thereafter, we employ a text-guided visual feature refinement module to fuse part-level visual features adaptively for each description. In experiments, we validate our TVFR through extensive experiments on CUHK-PEDES, which is the only available dataset for text-based person search. To the best of our knowledge, the TVFR outperforms other state-of-the-art methods.
Liying Gao, Kai Niu 0002, Zehong Ma, Bingliang Jiao, Tonghao Tan, Peng Wang 0015
ICMR2
2020 Textual Dependency Embedding for Person Search by Language
abstract
Person search by language aims to associate the pedestrian images with free-form natural language descriptions. Although great efforts have been made to align images with sentences, most researchers neglect the difficulty of long-distance dependency modeling in textual encoding, which is very important for solving this problem because the description sentences are always long and have complex structures for distinguishing different pedestrians. In this work, we focus on the long-distance dependencies in a sentence for better textual encoding, and accordingly propose the Textual Dependency Embedding (TDE) method. We first employ the sentence analysis tools to figure out the long-distance syntactic dependencies from a dependent to its governor in a sentence. Then we embed the dependent representations to their governor adaptively in our Governor-guided Dependent Attention Module (GDAM) to model these long-distance relations. After that, we further consider the dependency types, which also tell the importance of different dependents semantically, and embed them together with the dependents' features to clarify their inequivalent contributions to their governor. Extensive experiments and analysis on person search by language and image-text matching have validated the effectiveness of our method, and we have obtained the state-of-the-art performance on the CUHK-PEDES and Flickr30K datasets.
Kai Niu 0002, Yan Huang 0008, Liang Wang 0001
ACM Multimedia1
2020 Re-ranking image-text matching by adaptive metric fusion
Kai Niu 0002, Yan Huang 0008, Liang Wang 0001
Pattern Recognit.1
2020 Improving Description-Based Person Re-Identification by Multi-Granularity Image-Text Alignments
abstract
Description-based person re-identification (Re-id) is an important task in video surveillance that requires discriminative cross-modal representations to distinguish different people. It is difficult to directly measure the similarity between images and descriptions due to the modality heterogeneity (the crossmodal problem). And all samples belonging to a single category (the fine-grained problem) makes this task even harder than the conventional image-description matching task. In this paper, we propose a Multi-granularity Image-text Alignments (MIA) model to alleviate the cross-modal fine-grained problem for better similarity evaluation in description-based person Re-id. Specifically, three different granularities, i.e., global-global, global-local and local-local alignments are carried out hierarchically. Firstly, the global-global alignment in the Global Contrast (GC) module is for matching the global contexts of images and descriptions. Secondly, the global-local alignment employs the potential relations between local components and global contexts to highlight the distinguishable components while eliminating the uninvolved ones adaptively in the Relation-guided Global-local Alignment (RGA) module. Thirdly, as for the local-local alignment, we match visual human parts with noun phrases in the Bi-directional Fine-grained Matching (BFM) module. The whole network combining multiple granularities can be end-to-end trained without complex preprocessing. To address the difficulties in training the combination of multiple granularities, an effective step training strategy is proposed to train these granularities step-by-step. Extensive experiments and analysis have shown that our method obtains the state-of-the-art performance on the CUHK-PEDES dataset and outperforms the previous methods by a significant margin.
Kai Niu 0002, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
IEEE Trans. Image Process.1