Linhui Xiao

dblp:241/9207 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0003-2592-5264ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 62% Transfer learning and domain adaptation · 18% Segmentation and scene understanding · 8%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
visual grounding
3.342026
Toward Visual Grounding: A Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2026
CLIP-VG: Self-Paced Curriculum Adapting of CLIP for Visual Grounding · IEEE Trans. Multim. 2024
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling · NeurIPS 2024
Computer vision › Vision and language › visual grounding
referring expression comprehension
1.822026
Toward Visual Grounding: A Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2026
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling · NeurIPS 2024
Machine learning › Learning paradigms
curriculum learning
0.812024
CLIP-VG: Self-Paced Curriculum Adapting of CLIP for Visual Grounding · IEEE Trans. Multim. 2024
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot image classification
0.812024
SgVA-CLIP: Semantic-Guided Visual Adapting of Vision-Language Models for Few-Shot Image Classification · IEEE Trans. Multim. 2024
Computer vision › Segmentation and scene understanding
referring image segmentation
0.812024
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling · NeurIPS 2024
Computer vision › Vision and language › vision-language model
vision-language pre-trained model
0.812024
SgVA-CLIP: Semantic-Guided Visual Adapting of Vision-Language Models for Few-Shot Image Classification · IEEE Trans. Multim. 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
visual adaptation
0.812024
SgVA-CLIP: Semantic-Guided Visual Adapting of Vision-Language Models for Few-Shot Image Classification · IEEE Trans. Multim. 2024
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.212024
SgVA-CLIP: Semantic-Guided Visual Adapting of Vision-Language Models for Few-Shot Image Classification · IEEE Trans. Multim. 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked visual-language modeling
0.212024
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.212024
HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

survey · 1.0one-tower transformer · 0.8low-rank adaptation · 0.8knowledge distillation · 0.8dynamic image masking · 0.8cross-modal reconstruction · 0.8cross-modal bridge · 0.8contrastive learning · 0.8contrastive language-image pretraining · 0.8CLIP · 0.8
YearPublicationVenuePosition
2026 SHARP: Semantic Head-Aware Representation Pruning for Efficient MLLMs
abstract
Multimodal large language models (MLLMs) suffer from high inference costs, where visual tokens dominate the input sequence, often exceeding 90% of the total length. Current acceleration strategies typically employ inference-time token pruning, categorized into two main paradigms: internal LLM pruning and pre-LLM pruning. The former often undermines hardware optimizations like FlashAttention, while the latter, applied after the visual encoder, suffers from a lack of textual query guidance. In this study, we propose a Semantic Head-Aware Representation Pruning (SHARP) framework. The key idea is to identify pivotal attention heads that effectively capture cross-modal alignment by measuring text–image affinity derived from the visual encoder. Such a design not only leverages semantic alignment to preserve task-relevant information but also ensures significant end-to-end inference acceleration. Experiments on widely used vision–language benchmarks demonstrate that our approach achieves superior accuracy–efficiency trade-offs compared to previous token pruning strategies. Notably, on LLaVA-1.5-7B with FlashAttention, SHARP retains 95% of the original performance while requiring only 63% of the inference latency, underscoring its potential for deploying efficient MLLMs.
Mingyue Guo 0001, Linhui Xiao, Qingming Huang
ICMR3
2026 Toward Visual Grounding: A Survey
abstract
Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression text. This task simulates the common referential relationships between visual and linguistic modalities, enabling machines to develop human-like multimodal comprehension capabilities. Consequently, it has extensive applications in various domains. However, since 2021, visual grounding has witnessed significant advancements, with emerging new concepts such as grounded pre-training, grounding multimodal LLMs, generalized visual grounding, and giga-pixel grounding, which have brought numerous new challenges. In this survey, we first examine the developmental history of visual grounding and provide an overview of essential background knowledge, including fundamental concepts and evaluation metrics. We systematically track and summarize the advancements, and then meticulously define and organize the various settings to standardize future research and ensure a fair comparison. In the dataset section, we compile a comprehensive list of current relevant datasets, conduct a fair comparative analysis, and provide ultimate performance prediction to inspire the development of new standard benchmarks. Additionally, we delve into numerous applications and highlight several advanced topics. Finally, we outline the challenges confronting visual grounding and propose valuable directions for future research, which may serve as inspiration for subsequent researchers. By extracting common technical details, this survey encompasses the representative work in each subtopic over the past decade. To the best of our knowledge, this paper represents the most comprehensive overview currently available in the field of visual grounding. This survey is designed to be suitable for both beginners and experienced researchers, serving as an invaluable resource for understanding key concepts and tracking the latest research developments.
Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang 0001, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding
abstract
Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent grounding transfer works have advanced the field through one-tower architectures, they still suffer from two primary limitations: (1) over-entangled multimodal representations that exacerbate deceptive modality biases, and (2) insufficient semantic reasoning that hinders the comprehension of referential cues. In this paper, we propose BARE, a bias-aware and reasoning-enhanced framework for one-tower visual grounding. BARE introduces a mechanism that preserves modality-specific features and constructs referential semantics through three novel modules: (i) language salience modulator, (ii) visual bias correction and (iii) referential relationship enhancement, which jointly mitigate multimodal distractions and enhance referential comprehension. Extensive experimental results on five benchmarks demonstrate that BARE not only achieves state-of-the-art performance but also delivers superior computational efficiency compared to existing approaches. The code is publicly accessible at https://github.com/Marloweeee/BARE.
Hongbing Li, Linhui Xiao, Bo Xiao 0006, Zhanyu Ma
IEEE Trans. Circuits Syst. Video Technol.2
2025 Foregroundness-Aware Task Disentanglement and Self-Paced Curriculum Learning for Domain Adaptive Object Detection
abstract
Unsupervised domain adaptive object detection (UDA-OD) is a challenging problem since it needs to locate and recognize objects while maintaining the generalization ability across domains. Most existing UDA-OD methods directly integrate the adaptive modules into the detectors. This integration procedure can significantly sacrifice the detection performances, though it enhances the generalization ability. To solve this problem, we propose an effective framework, named foregroundness-aware task disentanglement and self-paced curriculum adaptation (FA-TDCA), to disentangle the UDA-OD task into four independent subtasks of source detector pretraining, classification adaptation, location adaptation, and target detector training. The disentanglement can transfer the knowledge effectively while maintaining the detection performance of our model. In addition, we propose a new metric, i.e., foregroundness, and use it to evaluate the confidence of the location result. We use both foregroundness and classification confidence to assess the label quality of the proposals. For effective knowledge transfer across domains, we utilize a self-paced curriculum learning paradigm to train adaptors and gradually improve the quality of the pseudolabels associated with the target samples. Experiment results indicate that our method achieves state-of-the-art results on four cross-domain object detection tasks.
Linhui Xiao, Chengliang Liu 0003, Zhihao Wu 0002, Yong Xu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
abstract
Visual grounding, which aims to ground a visual region via natural language, is a task that heavily relies on cross-modal alignment. Existing works utilized uni-modal pre-trained models to transfer visual or linguistic knowledge separately while ignoring the multimodal corresponding information. Motivated by recent advancements in contrastive language-image pre-training and low-rank adaptation (LoRA) methods, we aim to solve the grounding task based on multimodal pre-training. However, there exists significant task gaps between pre-training and grounding. Therefore, to address these gaps, we propose a concise and efficient hierarchical multimodal fine-grained modulation framework, namely HiVG. Specifically, HiVG consists of a multi-layer adaptive cross-modal bridge and a hierarchical multimodal low-rank adaptation (HiLoRA) paradigm. The cross-modal bridge can address the inconsistency between visual features and those required for grounding, and establish a connection between multi-level visual and text features. HiLoRA prevents the accumulation of perceptual errors by adapting the cross-modal features from shallow to deep layers in a hierarchical manner. Experimental results on five datasets demonstrate the effectiveness of our approach and showcase the significant grounding capabilities as well as promising energy efficiency advantages. The project page: https://github.com/linhuixiao/HiVG.
Linhui Xiao, Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu
ACM Multimedia1
2024 OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
abstract
Constrained by the separate encoding of vision and language, existing grounding and referring segmentation works heavily rely on bulky Transformer-based fusion en-/decoders and a variety of early-stage interaction technologies. Simultaneously, the current mask visual language modeling (MVLM) fails to capture the nuanced referential relationship between image-text in referring tasks. In this paper, we propose **OneRef**, a minimalist referring framework built on the modality-shared one-tower transformer that unifies the visual and linguistic feature spaces. To modeling the referential relationship, we introduce a novel MVLM paradigm called Mask Referring Modeling (**MRefM**), which encompasses both referring-aware mask image modeling and referring-aware mask language modeling. Both modules not only reconstruct modality-related content but also cross-modal referring content. Within MRefM, we propose a referring-aware dynamic image masking strategy that is aware of the referred region rather than relying on fixed ratios or generic random masking schemes. By leveraging the unified visual language feature space and incorporating MRefM's ability to model the referential relations, our approach enables direct regression of the referring results without resorting to various complex techniques. Our method consistently surpasses existing approaches and achieves SoTA performance on both grounding and segmentation tasks, providing valuable insights for future research. Our code and models are available at https://github.com/linhuixiao/OneRef.
Linhui Xiao, Xiaoshan Yang, Yaowei Wang 0001, Changsheng Xu
NeurIPS1
2024 SgVA-CLIP: Semantic-Guided Visual Adapting of Vision-Language Models for Few-Shot Image Classification
abstract
Although significant progress has been made in few-shot learning, most of existing few-shot image classification methods require supervised pre-training on a large amount of samples of base classes, which limits their generalization ability in real world application. Recently, large-scale Vision-Language Pre-trained models (VLPs) have been gaining increasing attention in few-shot learning because they can provide a new paradigm for transferable visual representation learning with easily available text on the Web. However, the VLPs may neglect detailed visual information that is difficult to describe by language sentences, but important for learning an effective classifier to distinguish different images. To address the above problem, we propose a new framework, named Semantic-guided Visual Adapting (SgVA), which can effectively extend vision-language pre-trained models to produce discriminative adapted visual features by comprehensively using an implicit knowledge distillation, a vision-specific contrastive loss, and a cross-modal contrastive loss. The implicit knowledge distillation is designed to transfer the fine-grained cross-modal knowledge to guide the updating of the vision adapter. State-of-the-art results on 13 datasets demonstrate that the adapted visual features can well complement the cross-modal features to improve few-shot image classification.
Xiaoshan Yang, Linhui Xiao, Yaowei Wang 0001, Changsheng Xu
IEEE Trans. Multim.3
2024 CLIP-VG: Self-Paced Curriculum Adapting of CLIP for Visual Grounding
abstract
Visual Grounding (VG) is a crucial topic in the field of vision and language, which involves locating a specific region described by expressions within an image. To reduce the reliance on manually labeled data, unsupervised methods have been developed to locate regions using pseudo-labels. However, the performance of existing unsupervised methods is highly dependent on the quality of pseudo-labels and these methods always encounter issues with limited diversity. In order to utilize vision and language pre-trained models to address the grounding problem, and reasonably take advantage of pseudo-labels, we propose CLIP-VG, a novel method that can conduct self-paced curriculum adapting of CLIP with pseudo-language labels. We propose a simple yet efficient end-to-end network architecture to realize the transfer of CLIP to the visual grounding. Based on the CLIP-based architecture, we further propose single-source and multi-source curriculum adapting algorithms, which can progressively find more reliable pseudo-labels to learn an optimal model, thereby achieving a balance between reliability and diversity for the pseudo-language labels. Our method outperforms the current state-of-the-art unsupervised method by a significant margin on RefCOCO/+/g datasets in both single-source and multi-source scenarios, with improvements ranging from 6.78% to 10.67% and 11.39% to 14.87%, respectively. Furthermore, our approach even outperforms existing weakly supervised methods.
Linhui Xiao, Xiaoshan Yang, Ming Yan 0008, Yaowei Wang 0001, Changsheng Xu
IEEE Trans. Multim.1