Hao Li 0093

dblp:17/5705-93 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0002-1758-5936ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence
abstract
Social Intelligence Queries (Social-IQ) serve as the primary multimodal benchmark for evaluating a model’s social intelligence level. While impressive multiple-choice question (MCQ) accuracy is achieved by current solutions, increasing evidence shows that they are largely, and in some cases entirely, dependent on language modality, overlooking visual context. Additionally, the closed-set nature further prevents the exploration of whether and to what extent the reasoning path behind selection is correct. To address these limitations, we propose the Visually Explainable and Grounded Artificial Social Intelligence (VEGAS) model. As a generative multimodal model, VEGAS leverages open-ended answering to provide explainable responses, which enhances the clarity and evaluation of reasoning paths. To enable visually grounded answering, we propose a novel sampling strategy to provide the model with more relevant visual frames. We then enhance the model’s interpretation of these frames through Generalist Instruction Fine-Tuning (GIFT), which aims to: i) learn multimodal language transformations for fundamental emotional social traits, and ii) establish multimodal joint reasoning capabilities. Extensive experiments, comprising modality ablation, open-ended assessments, and supervised MCQ evaluations, consistently show that VEGAS effectively utilizes visual information in reasoning to produce correct and also credible answers. We expect this work to offer a new perspective on Social-IQ and advance the development of human-like social AI.
Hao Li 0093, Hao Fei 0001, Zechao Hu 0003, Zhengwei Yang 0001, Zheng Wang 0007
AAAI1
2025 CCIN: Compositional Conflict Identification and Neutralization for Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) is a multi-modal task that seeks to retrieve target images by harmonizing a reference image with a modified instruction. A key challenge in CIR lies in compositional conflicts between the reference image (e.g., blue, long sleeve) and the modified instruction (e.g., grey, short sleeve). Previous works attempt to mitigate such conflicts through feature-level manipulation, commonly employing learnable masks to obscure conflicting features within the reference image. However, the inherent complexity of feature spaces poses significant challenges in precise conflict neutralization, thereby leading to uncontrollable results. To this end, this paper proposes the Compositional Conflict Identification and Neutralization (CCIN) framework, which sequentially identifies and neutralizes compositional conflicts for effective CIR. Specifically, CCIN comprises two core modules: 1) Compositional Conflict Identification module, which utilizes LLM-based analysis to identify specific conflicting attributes, and 2) Compositional Conflict Neutralization module, which first generates a kept instruction to preserve non-conflicting attributes, then neutralizes conflicts under collaborative guidance of both the kept and modified instructions. Extensive experiments demonstrate the superiority of CCIN over the state-of-the-arts. Code repository: https://github.com/LikaiTian/CCIN.
Likai Tian, Jian Zhao 0006, Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Lei Jin 0003, Zheng Wang 0007, Xuelong Li 0001
CVPR5
2025 Cross-Category Subjectivity Generalization for Style-Adaptive Sketch Re-ID
Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Zheng Wang 0007, Yixiong Zou
ICCV3
2025 The ACM Multimedia 2025 Grand Challenge of Truthful and Responsible Multimodal Learning
abstract
The Truthful and Responsible Multimodal Learning Challenge aims to foster advancements in the development of reliable and trustworthy multimodal AI systems by addressing two crucial tasks: multimodal hallucination detection and multimodal factuality detection. Task A focuses on detecting hallucinated elements in AI-generated image captions, such as fabricated objects or attributes. Task B targets verifying the factual accuracy of textual claims using visual and contextual cues. We establish benchmarks that support responsible multimodal AI in diverse real-world applications.
Kai Liu 0023, Yanlin Li 0014, Hao Li 0093, Zheng Wang 0007
ACM Multimedia4
2025 From Language to Instance: Generative Visual Prompting for Zero-shot Camouflaged Object Detection
abstract
Traditional Camouflaged Object Detection (COD) methods heavily depend on labor-intensive annotated datasets which require extensive manual effort, resulting in limited generalization. While recent studies have combined Multimodal Large Language Models (MLLMs) and Vision Foundation Models (VFMs) to achieve zero-shot COD, their performance is hindered by modality gap between linguistic semantics and fine-grained visual cues, especially in complex camouflage scenarios. In this paper, we propose Language-to-instance generative visual Prompting (LiP), a novel framework that addresses this limitation by transforming text prompts generated by MLLMs into instance-level visual prompts through a text-to-image generative process. Specifically, we introduce a Diffusion-driven Visual Prompt Generation (DVPG) module that leverages Stable Diffusion model to synthesize visual references, enabling robust homogeneous modality matching for COD. Additionally, we introduce Instruction Contrastive Reasoning (ICR) module to enhance the semantic reliability of prompts by suppressing hallucinated concepts during MLLM inference. To the best of our knowledge, LiP is the first framework that utilize text-to-image generative model to construct instance-level visual prompts in COD task. Extensive experiments on four benchmark datasets demonstrate the effectiveness and strong generalization ability of our approach.
Zihou Zhang, Hao Li 0093, Zhengwei Yang 0001, Zechao Hu 0003, Liang Li 0003, Zheng Wang 0007
ACM Multimedia2
2025 Unified Category and Style Generalization for Instance-Level Sketch Retrieval
abstract
Zero-shot instance-level sketch retrieval addresses a practical retrieval scenario in which sketches from unseen categories during training serve as queries to retrieve matching RGB images. The core challenges of this task lie in two aspects: unknown category generalization and subjective style adaptation. Existing methods either focus solely on category generalization or apply simplistic style elimination techniques within a specific category, leading to suboptimal performance when both challenges are present. To this end, we propose the Dual-Attentive Prompt (DAP) method, which unifies category generalization and style adaptation into a single, interpretable framework. Central to DAP is a dual-attentive prompt composer, consisting of two self-attention-based modules. This composer dynamically integrates pre-learned category-specific knowledge with instance-specific prompts that adapt to sketch-specific styles. By cooperating with additional style alignment loss, the proposed method ensures robust generalization of unseen categories while mitigating the impact of subjective style variations. Extensive experimental results demonstrate the state-of-the-art performance of the proposed method. Additionally, some insights are provided into the challenges of traditional training processes when handling multi-style sketches, along with quantitative and qualitative evidence showing how the proposed approach effectively mitigates the negative impact of subjective style variations.
Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Yixiong Zou, Fengbin Zhu, Zheng Wang 0007
SIGIR3
2025 Contrastive-Generative-Contrastive: Neutralize Subjectivity in Sketch Re-Identification
abstract
Sketch-based person re-identification (Sketch re-ID) aims to match pedestrian figures in hand-drawn sketches with their corresponding RGB photos. This technique allows for person retrieval or tracking in surveillance systems when the target person’s RGB photo is not available. While previous research predominantly focused on bridging the modality gap between sketches and RGB photos, the influence of the inherent subjectivity in hand-drawn sketches on re-ID performance remains under-explored. This subjectivity, originating from the artist’s unique style, perceptions, and interpretations, introduces inaccuracies in depicting pedestrian appearances, thereby posing additional challenges such as feature distortion and stylistic variation. This paper introduces a Contrastive-Generative-Contrastive (CGC) framework for subjective style-insensitive re-ID. The framework employs a generative model optimized through self-supervision by contrasting positive and negative pairs of pedestrian sketches and RGB photos. In this manner, it simulates an additional artist specializing in transforming original sketches from various subjective styles into uniform ones. Besides, a simple yet effective weighted contrastive learning loss is proposed to further enhance the model’s focus on pedestrian ID-relevant features. Experimental results demonstrate that the proposed method significantly reduces the influence of subjectivity in feature extraction, achieving new state-of-the-art results on benchmark datasets.
Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Zheng Wang 0007
IEEE Trans. Inf. Forensics Secur.3
2024 Expressiveness is Effectiveness: Self-supervised Fashion-aware CLIP for Video-to-Shop Retrieval
Likai Tian, Zhengwei Yang 0001, Zechao Hu 0003, Hao Li 0093, Yifang Yin, Zheng Wang 0007
IJCAI4
2023 Fuzzy regular least squares twin support vector machine and its application in fault diagnosis
Chengjiang Zhou, Hao Li 0093, Jintao Yang, Qihua Yang, Limiao Yang, Shanyou He, Xuyi Yuan
Expert Syst. Appl.2
2022 Real-Time Garbage Object Detection With Data Augmentation and Feature Fusion Using SUAV Low-Altitude Remote Sensing Images
abstract
Recently, a number of nature reserves have been shut down because of serious pollution from tourist garbage. Garbage monitoring in high-altitude natural reserves using small unmanned aerial vehicle (SUAV) remote sensing is an important and urgent need for environmental protection. In order to help cleaners to eliminate garbage more conveniently and quickly, a novel approach is proposed to detect scattered garbage regions in real time using low-altitude remote sensing videos captured by SUAVs. First, the high-resolution, low-altitude, multitemporal remote sensing images and videos containing scattered garbage were collected through SUAV and then proposed a data augmentation method to expand the training samples. Second, the Yolov4 detection network was used to classify the scattered garbage regions. Finally, the location of the object was roughly calculated according to the altitude, flight direction, global positioning system, and digital elevation model (DEM). Then, the garbage object was marked on the video, while the object location was marked on the map. Experimental results show that the proposed method achieves a mean accuracy of 91.34% and provides better performances on the real data set compared with state-of-the-art methods.
Hao Li 0093, Quanjing Li, Yang Yang 0032, Kun Yang 0007
IEEE Geosci. Remote. Sens. Lett.3
2022 Learning Relaxed Neighborhood Consistency for Feature Matching
abstract
Feature matching is a critical prerequisite in many applications of remote sensing, and its aim is to establish reliable correspondences between two sets of features. Existing attempts typically involve estimating the underlying image transformations to remove false matches in putative matches. However, the image transformation could vary with different application scenarios, which means that using a predefined geometrical model may lead to inferior matching accuracy, especially if the image transformation is nonrigid. This article casts the mismatch removal into a neighborhood consistency evaluation problem under a customized learning framework. With only seven training image pairs involving approximately 8000 putative matches, our method can handle different types of images or transformation models (affine, homography, piecewise-linear transformation, and others). Extensive experiments on feature matching and image registration are conducted to demonstrate the superiority of our method over the eight state-of-the-art competitors.
Shuang Chen 0008, Jiaxuan Chen 0002, Zenghui Xiong, Linjie Xing, Yang Yang 0032, Kai Yan 0002, Hao Li 0093
IEEE Trans. Geosci. Remote. Sens.8
2021 Progressive structure network-based multiscale feature fusion for object detection in real-time application
Lvjiyuan Jiang, Hao Li 0093, Kai Yan 0002, Yang Yang 0032, Yungang Zhang, Lianliu Qiao, Cuilian Fu
Eng. Appl. Artif. Intell.4