Xiaoshuai Hao

dblp:271/8403 · DBLP profile ↗
← Back
51ranked-venue papers
15as first author
51since 2021 · last 2026
0009-0007-4209-6695ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 9 first-author · 31 since 2021Artificial intelligence and machine learning · 28 · 8 first-author · 28 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 What You See Is What You Reach: Towards Spatial Navigation with High-Level Human Instructions
abstract
Embodied navigation is a fundamental capability that enables embodied agents to effectively interact with the physical world in various complex environments. However, a significant gap remains between current embodied navigation tasks and real-world requirements, as existing methods often struggle to integrate high-level human instructions with spatial understanding. To address this gap, we propose a new task of embodied navigation called spatial navigation, which encompasses two key components: spatial object navigation (SpON) for object-specific guidance and spatial area navigation (SpAN) for navigating to designated areas. Specifically, SpON guides agents to specific objects by leveraging spatial relationships and contextual understanding, while SpAN focuses on navigating to defined areas within complex environments. Together, these components significantly enhance agents’ navigation capabilities, enabling more effective interactions in real-world scenarios. To support this task, we have generated a spatial navigation dataset consisting of 10K trajectories within the simulator. This dataset includes high-level human instructions, detailed observations, and corresponding navigation actions, providing a comprehensive resource to enhance agent training and performance. Building on the spatial navigation dataset, we introduce SpNav, a hierarchical navigation framework. Specifically, SpNav employs vision-language model (VLM) to interpret high-level human instructions and accurately identify goal objects or areas within the observation range, achieving precise point-to-point navigation using a map and enhancing the agent’s ability to oper- ate effectively in complex environments by bridging the gap between perception and action. Extensive experiments show that SpNav achieves state-of-the-art (SOTA) performance in spatial navigation tasks across both simulated and real-world environments, validating the effectiveness of our method.
Haoxiang Fu, Xiaoshuai Hao, Qiang Zhang 0029, Long Chen 0015, Wenbo Ding 0001
AAAI3
2026 NavA³: Understanding Any Instruction, Navigating Anywhere, Finding Anything
abstract
Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang, Haoxiang Fu, Xinyu Zheng, Pengwei Wang, Zhongyuan Wang, Wenbo Ding, Shanghang Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaoshuai Hao, Yingbo Tang, Haoxiang Fu, Pengwei Wang 0005, Zhongyuan Wang 0006, Wenbo Ding 0001, Shanghang Zhang
ACL (1)2
2026 Question-Adaptive Graph Learning for Multi-hop Retrieval Augmented Generation
abstract
Retrieval-augmented generation (RAG) has demonstrated its ability to enhance Large Language Models (LLMs) by integrating external knowledge sources. However, multi-hop questions, which require the identification of multiple knowledge targets to form a synthesized answer, raise new challenges for RAG systems. Under the multi-hop settings, existing methods often struggle to fully understand the questions with complex semantic structures and are susceptible to irrelevant noise during the retrieval of multiple information targets. To address these limitations, we propose a novel graph representation learning framework for multi-hop question retrieval. We first introduce a Multi-information Level Knowledge Graph (Multi-L KG) to model various information levels for a more comprehensive understanding of multi-hop questions. Based on this, we design a Question-Adaptive Graph Neural Network (Quest-GNN) for representation learning on the Multi-L KG. Quest-GNN employs intra/inter-level message passing mechanisms, and in each message passing the information aggregation is guided by the question, which not only facilitates multi-granular information aggregation but also significantly reduces the impact of noise. To enhance its ability to learn robust representations, we further propose two synthesized data generation strategies for pre-training the Quest-GNN. Extensive experimental results demonstrate the effectiveness of our framework in multi-hop scenarios, especially in high-hop questions the improvement can reach 33.8%. The code is available at: https://github.com/Jerry2398/QSGNN.
Peiyan Zhang, Yatao Bian, Xiaoshuai Hao
SIGIR7
2026 DWCL: Dual-Weighted Contrastive Learning for robust multi-view clustering
Hanning Yuan, Lianhua Chi, Sijie Ruan, Wei Zhou 0021, Jinhui Pang, Xiaoshuai Hao
Eng. Appl. Artif. Intell.8
2026 H2R-BM: Can leveraging human videos enhance performance and generalizability in robotic bimanual manipulation?
Xiaoshuai Hao, Huaihai Lyu, Dayan Wu, Jing Zhang 0037, Long Chen 0015
Pattern Recognit.1
2026 Communication-efficient personalized federal graph learning via low-rank decomposition
Ruyue Liu, Rong Yin 0001, Xiangzhen Bo, Xiaoshuai Hao, Xingrui Zhou, Yong Liu 0018, Jinwen Zhong, Can Ma, Weiping Wang 0005
Pattern Recognit.4
2026 Pose-Guided Multi-Cue Explicit Query Construction for Disambiguating Human-Object Interactions
abstract
Human-Object Interaction (HOI) detection remains challenging due to the semantic ambiguity of interaction categories and the limited discriminability of their feature representations. Existing approaches often improve recognition by employing sophisticated models or auxiliary textual annotations. While effective in certain gains, these solutions incur additional computational or annotation costs and struggle to capture intrinsic interaction regularities. To address these issues, we propose Pose-Guided Multi-Cue Explicit Query Construction (PM-EQC), a unified Transformer-based framework that builds upon collaborative modeling of appearance, spatial, and pose cues for discriminative interaction reasoning. At its core, the Collaborative Multi-Cue Query Constructor (CM-CQC) jointly models dependencies among visual cues to generate explicit query embeddings. CM-CQC further incorporates a hierarchical pose contextualization mechanism: global body configurations adaptively guide attention to local critical joints, yielding fine-grained pose embeddings and more precise interaction disambiguation. Owing to its modular design, PM-EQC integrates seamlessly with diverse backbones and benefits from their advances. Extensive experiments on PhysLab, HICO-DET, and V-COCO datasets demonstrate that PM-EQC achieves state-of-the-art performance, and the code is publicly available at https://github.com/ZMHSDUST/ PM-EQC.
Minghao Zou, Qingtian Zeng, Xue Zhang 0008, Guiyuan Yuan, Xiaoshuai Hao, Jun Liu 0036, Wei Zhou 0021
IEEE Trans. Circuits Syst. Video Technol.6
2026 Embodied Spatial Affordance: Spatial-Aware Affordance Learning for Embodied Navigation and Manipulation
abstract
Embodied navigation and manipulation are fundamental capabilities for embodied agents operating in physical environments. A key challenge in this process is understanding the spatial context and the affordances of the environment, which involves recognizing how objects can be interacted with (object affordance) and identifying suitable locations for movement and object placement (free space affordance). While Vision-Language Models (VLMs) have shown promise in high-level task planning, their ability to translate reasoning into precise executable actions remains limited, particularly in image-based spatial understanding and precise affordance localization-a critical gap in image processing for robotics. To bridge this gap, we propose EspA, a novel image-to-keypoint model that leverages spatial-aware affordance learning to predict actionable affordances directly from 2D image inputs. Built on a hierarchical vision-language architecture, EspA jointly reasons about object affordances and free space affordances, enabling pixel-level localization of both types of interactions. Crucially, EspA translates language instructions into precise 2D affordance keypoints from observed images, which are then projected into 3D actionable coordinates using depth information. To support this unified affordance reasoning, we introduce the Embodied Spatial Affordance (ESA) dataset, which captures both object-centric interactions and free space contexts. By jointly modeling these affordances in a shared representation space, EspA overcomes the limitations of prior works that treat them independently. The dataset's fine-grained annotations enable our model to learn the intricate relationship between object functionality and spatial feasibility, significantly enhancing the spatial understanding in embodied tasks. Extensive experimental results demonstrate that EspA outperforms existing state-of-the-art Vision-Language Models (VLMs), both open-source and closed-source, in object and free space affordance prediction. Furthermore, it exhibits superior performance in real-world embodied navigation and manipulation experiments. Our work advances the field of image-based spatial reasoning by providing a scalable solution for translating high-level instructions into low-level actionable affordances. We believe this work paves the way for more robust and versatile embodied agents capable of effectively interacting with complex environments. The dataset, benchmark, and evaluation code will be publicly available to facilitate future research. Project website: https://embodied-spatial-affordance.github.io/.
Xiaoshuai Hao, Yingbo Tang, Long Chen 0015, Wei Zhou 0021, Jungong Han, Wenbo Ding 0001, Xiao-Ping Zhang 0002
IEEE Trans. Image Process.1
2026 Synergistic Prompting for Complementarity and Consistency in Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) aims to partition unlabeled multi-view data into semantically coherent groups, even when certain views are missing due to sensor failures, data collection constraints, or privacy concerns. Despite advancements in deep IMVC methods, two critical challenges remain unresolved: (i) the lack of explicit mechanisms to model cross-view complementarity and (ii) the absence of principled strategies to ensure global semantic consistency across views. To address these challenges, we propose SP-IMVC, a novel Synergistic Prompting framework that jointly models complementarity and consistency under view incompleteness. Specifically, we introduce two types of learnable prompts: the Cross-View Complementary Prompt (CVCP), which aggregates auxiliary representations from available views to enrich the semantics of the current view and mitigate information loss; and the Latent Anchor Prompt (LAP), which utilizes a global anchor prompt pool to provide adaptive semantic priors that promote globally consistent representations. These prompts are optimized jointly within a unified architecture to achieve synergistic prompting of cross-view complementarity and global semantic consistency. Extensive experiments on six public benchmarks demonstrate that SP-IMVC consistently outperforms 14 state-of-the-art IMVC approaches, particularly in scenarios with high missing-view ratios, validating the effectiveness and robustness of our synergistic prompt-guided clustering framework. The code will be released to facilitate future research.
Xiaoshuai Hao, Yingbo Tang, Peng Hao 0003, Yunfeng Diao, Guangyin Jin, Yu Liu 0023
IEEE Trans. Image Process.1
2026 DADA++: Dual Alignment Domain Adaptation for Unsupervised Video-Text Retrieval
abstract
Video-text retrieval aims at returning the most semantically relevant videos given a textual query, which is a thriving topic in both computer vision and natural language processing communities. This article focuses on a more challenging task, i.e., Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), wherein training and testing data come from different distributions. Previous approaches are mostly derived from classification-based domain adaptation methods, which are neither multi-modal nor suitable for retrieval tasks. They merely alleviate the domain shift while overlooking the pairwise misalignment issue in the target domain, i.e., there exist no semantic relationships between target videos and texts. While Foundation Models like CLIP perform well in in-domain video-text retrieval, their effectiveness significantly drops during domain shifts due to this lack of alignment. To tackle this, we propose a novel method named D ual A lignment D omain A daptation ( DADA ++). Specifically, we first introduce cross-modal semantic embedding to generate discriminative source features in a joint embedding space. Besides, we utilize cross-modal domain adaptations to balance the minimization of domain shift in a smooth manner. Furthermore, we empirically identify the pairwise misalignment in the target domain, and thus propose the i ntegrated D ual A lignment C onsistency (iDAC). The proposed iDAC adaptively aligns the video-text pairs, which are more likely to be relevant in the target domain, by verifying their cross-modal semantic proximity reciprocally in both hard and soft manners. This enables positive pairs to increase progressively while potentially aligning noisy pairs throughout the training procedure. We also provide insights into the functionality of DADA ++ through the lens of Foundation Models, explaining its superiority in a theoretical way. Compared with state-of-the-art methods, DADA ++ achieves 9.4% and 8.5% relative improvements on R@1 under the settings of TGIF \(\rightarrow\) MSR-VTT and TGIF \(\rightarrow\) MSVD, respectively, demonstrating its superior performance.
Xiaoshuai Hao, Yunfeng Diao, Rong Yin 0001, Guangyin Jin, Jing Zhang 0037, Wanqian Zhang, Wei Zhou 0021
ACM Trans. Multim. Comput. Commun. Appl.1
2025 KALAHash: Knowledge-Anchored Low-Resource Adaptation for Deep Hashing
abstract
Deep hashing has been widely used for large-scale approximate nearest neighbor search due to its storage and search efficiency. However, existing deep hashing methods predominantly rely on abundant training data, leaving the more challenging scenario of low-resource adaptation for deep hashing relatively underexplored. This setting involves adapting pre-trained models to downstream tasks with only an extremely small number of training samples available. Our preliminary benchmarks reveal that current methods suffer significant performance degradation due to the distribution shift caused by limited training samples. To address these challenges, we introduce Class-Calibration LoRA (CLoRA), a novel plug-and-play approach that dynamically constructs low-rank adaptation matrices by leveraging class-level textual knowledge embeddings. CLoRA effectively incorporates prior class knowledge as anchors, enabling parameter-efficient fine-tuning while maintaining the original data distribution. Furthermore, we propose Knowledge-Guided Discrete Optimization (KIDDO), a framework to utilize class knowledge to compensate for the scarcity of visual information and enhance the discriminability of hash codes. Extensive experiments demonstrate that our proposed method, Knowledge- Anchored Low-Resource Adaptation Hashing (KALAHash), significantly boosts retrieval performance and achieves a 4× data efficiency in low-resource scenarios.
Shu Zhao 0006, Xiaoshuai Hao, Narayanan Vijaykrishnan
AAAI3
2025 MapNav: A Novel Memory Representation via Annotated Semantic Maps for VLM-based Vision-and-Language Navigation
abstract
Lingfeng Zhang, Xiaoshuai Hao, Qinwen Xu, Qiang Zhang, Xinyao Zhang, Pengwei Wang, Jing Zhang, Zhongyuan Wang, Shanghang Zhang, Renjing Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiaoshuai Hao, Qinwen Xu, Qiang Zhang 0029, Pengwei Wang 0004, Jing Zhang 0037, Zhongyuan Wang 0006, Shanghang Zhang, Renjing Xu
ACL (1)2
2025 RPGCN: Relational Probabilistic Graphs for EEG-Based Emotion Mining
Xinliang Zhou, Jianheng Zhou, Jiaping Xiao, Xiaoshuai Hao, Jing Wang 0060, Badong Chen, Qingsong Wen
ADMA (1)5
2025 M3-Net: A Cost-Effective Graph-Free MLP-Based Model for Traffic Prediction
abstract
Achieving accurate traffic prediction is a fundamental but crucial task in the development of current intelligent transportation systems. These limitations pose significant challenges for the efficient deployment and operation of deep learning models on large-scale datasets. To address these challenges, we propose a cost-effective graph-free Multilayer Perceptron (MLP) based model M3-Net for traffic prediction. Extensive experiments conducted on multiple real datasets demonstrate the superiority of the proposed model in terms of prediction performance and lightweight deployment. Our code is available at https://github.com/jinguangyin/M3_NET
Guangyin Jin, Sicong Lai, Xiaoshuai Hao, Jinlei Zhang
CIKM3
2025 RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
abstract
Recent advancements in Multimodal Large Language Models (MLLMs) have shown remarkable capabilities across various multimodal contexts. However, their application in robotic scenarios, particularly for long-horizon manipulation tasks, reveals significant limitations. These limitations arise from the current MLLMs lacking three essential robotic brain capabilities: Planning Capability, which involves decomposing complex manipulation instructions into manageable sub-tasks; Affordance Perception, the ability to recognize and interpret the affordances of interactive objects; and Trajectory Prediction, the foresight to anticipate the complete manipulation trajectory necessary for successful execution. To enhance the robotic brain’s core capabilities from abstract to concrete, we introduce ShareRobot, a high-quality heterogeneous dataset that labels multi-dimensional information such as task planning, object affordance, and end-effector trajectory. ShareRobot’s diversity and accuracy have been meticulously refined by three human annotators. Building on this dataset, we developed RoboBrain, an MLLM-based model that combines robotic and general multi-modal data, utilizes a multi-stage training strategy, and incorporates long videos and high-resolution images to improve its robotic manipulation capabilities. Extensive experiments demonstrate that RoboBrain achieves state-of-the-art performance across various robotic tasks, highlighting its potential to advance robotic brain capabilities. Project website: RoboBrain.
Yuheng Ji, Huajie Tan, Xiaoshuai Hao, Yuan Zhang 0020, Pengwei Wang 0004, Mengdi Zhao, Yao Mu 0001, Pengju An, Xinda Xue, Qinghang Su, Huaihai Lyu, Xiaolong Zheng 0001, Jiaming Liu 0003, Zhongyuan Wang 0006, Shanghang Zhang
CVPR4
2025 Synergistic Prompting for Robust Visual Recognition with Missing Modalities
abstract
Large-scale multi-modal models have demonstrated remarkable performance across various visual recognition tasks by leveraging extensive paired multi-modal training data. However, in real-world applications, the presence of missing or incomplete modality inputs often leads to significant performance degradation. Recent research has focused on prompt-based strategies to tackle this issue; however, existing methods are hindered by two major limitations: (1) static prompts lack the flexibility to adapt to varying missing-data conditions, and (2) basic prompt-tuning methods struggle to ensure reliable performance when critical modalities are missing.To address these challenges, we propose a novel Synergistic Prompting (SyP) framework for robust visual recognition with missing modalities. The proposed SyP introduces two key innovations: (I) a Dynamic Adapter, which computes adaptive scaling factors to dynamically generate prompts, replacing static parameters for flexible multi-modal adaptation, and (II) a Synergistic Prompting Strategy, which combines static and dynamic prompts to balance information across modalities, ensuring robust reasoning even when key modalities are missing. The proposed SyP achieves significant performance improvements over existing approaches across three widely-used visual recognition datasets, demonstrating robustness under diverse missing rates and conditions. Extensive experiments and ablation studies validate its effectiveness in handling missing modalities, highlighting its superior adaptability and reliability.
Luanyuan Dai, Qika Lin, Yunfeng Diao, Guangyin Jin, Yufei Guo 0001, Jing Zhang 0037, Xiaoshuai Hao
ICCV8
2025 Training-Free Generation of Temporally Consistent Rewards from VLMs
abstract
Recent advances in vision-language models (VLMs) have significantly improved performance in embodied tasks such as goal decomposition and visual comprehension. However, providing accurate rewards for robotic manipulation without fine-tuning VLMs remains challenging due to the absence of domain-specific robotic knowledge in pre-trained datasets and high computational costs that hinder real-time applicability. To address this, we propose $\mathrm{T}^2$-VLM, a novel training-free, temporally consistent framework that generates accurate rewards through tracking the status changes in VLM-derived subgoals. Specifically, our method first queries the VLM to establish spatially aware subgoals and an initial completion estimate before each round of interaction. We then employ a Bayesian tracking algorithm to update the goal completion status dynamically, using subgoal hidden states to generate structured rewards for reinforcement learning (RL) agents. This approach enhances long-horizon decision-making and improves failure recovery capabilities with RL. Extensive experiments indicate that $\mathrm{T}^2$-VLM achieves state-of-the-art performance in two robot manipulation benchmarks, demonstrating superior reward accuracy with reduced computation consumption. We believe our approach not only advances reward generation techniques but also contributes to the broader field of embodied AI. Project website: https://t2-vlm.github.io/.
Yinuo Zhao, Jiale Yuan, Xiaoshuai Hao, Xinyi Zhang 0009, Kun Wu 0001, Zhengping Che, Chi Harold Liu, Jian Tang 0008
ICCV4
2025 TASAR: Transfer-based Attack on Skeletal Action Recognition
abstract
Skeletal sequence data, as a widely employed representation of human actions, are crucial in Human Activity Recognition (HAR). Recently, adversarial attacks have been proposed in this area, which exposes potential security concerns, and more importantly provides a good tool for model robustness test. Within this research, transfer-based attack is an important tool as it mimics the real-world scenario where an attacker has no knowledge of the target model, but is under-explored in Skeleton-based HAR (S-HAR). Consequently, existing S-HAR attacks exhibit weak adversarial transferability and the reason remains largely unknown. In this paper, we investigate this phenomenon via the characterization of the loss function. We find that one prominent indicator of poor transferability is the low smoothness of the loss function. Led by this observation, we improve the transferability by properly smoothening the loss when computing the adversarial examples. This leads to the first Transfer-based Attack on Skeletal Action Recognition, TASAR. TASAR explores the smoothened model posterior of pre-trained surrogates, which is achieved by a new post-train Dual Bayesian optimization strategy. Furthermore, unlike existing transfer-based methods which overlook the temporal coherence within sequences, TASAR incorporates motion dynamics into the Bayesian attack, effectively disrupting the spatial-temporal coherence of S-HARs. For exhaustive evaluation, we build the first large-scale robust S-HAR benchmark, comprising 7 S-HAR models, 10 attack methods, 3 S-HAR datasets and 2 defense models. Extensive results demonstrate the superiority of TASAR. Our benchmark enables easy comparisons for future studies, with the code available in the https://github.com/yunfengdiao/Skeleton-Robustness-Benchmark.
Yunfeng Diao, Baiqi Wu, Ajian Liu 0001, Xiaoshuai Hao, Meng Wang 0001, He Wang 0002
ICLR5
2025 Multi-Granularity Based Collaborative Learning for Semi-Supervised Hashing
abstract
Recently, deep semi-supervised hashing methods have achieved remarkable success by simultaneously leveraging sufficient unlabeled data and limited labeled data. However, these methods focus on the pseudo-label semantic similarity relation of unlabeled data, while ignoring multi-granularity similarity relations. The fine-grained instance-level and neighborhood similarity relations can help to learn a more detailed data distribution in the Hamming space, which is beneficial to improve the discrimination of hash codes. We thus propose a novel Multi-Granularity Based Collaborative Learning Hashing (MGCLH), which utilizes the complementary multi-granularity similarity relations of unlabeled data to collaboratively improve the discrimination of hash codes. Specifically, we introduce an Instance-Wise Contrastive Module (ICM), which embeds instance-wise similarity relation into hash codes to achieve coarse clustering of hash codes. Moreover, we design a Neighborhood Consistent Module (NCM) to capture neighborhood similarity relations for preserving the inherent neighborhood structure of hash codes. Furthermore, the Class-Wise Contrastive Module (CCM) embeds the class-wise semantic similarity relation between unlabeled data into the hash codes to improve its inter-class separability. Extensive experimental results on four image datasets demonstrate that the proposed method outperforms several state-of-the-art semi-supervised hashing methods.
Shuai Cheng 0002, Xiaoshuai Hao, Wanqian Zhang
ICME3
2025 MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving Perception
abstract
Multi-sensor fusion models play a crucial role in autonomous driving perception, particularly in tasks like 3D object detection and HD map construction. These models provide essential and comprehensive static environmental information for autonomous driving systems. While camera-LiDAR fusion methods have shown promising results by integrating data from both modalities, they often depend on complete sensor inputs. This reliance can lead to low robustness and potential failures when sensors are corrupted or missing, raising significant safety concerns. To tackle this challenge, we introduce the Multi-Sensor Corruption Benchmark (MSC-Bench), the first comprehensive benchmark aimed at evaluating the robustness of multi-sensor autonomous driving perception models against various sensor corruptions. Our benchmark includes 16 combinations of corruption types that disrupt both camera and LiDAR inputs, either individually or concurrently. Extensive evaluations of six 3D object detection models and four HD map construction models reveal substantial performance degradation under adverse weather conditions and sensor failures, underscoring critical safety issues. The benchmark toolkit and affiliated code and model checkpoints have been made publicly accessible. Project website: MSC-Bench.
Xiaoshuai Hao, Guanqun Liu 0008, Yuheng Ji, Mengchuan Wei, Haimei Zhao, Lingdong Kong, Rong Yin 0001, Yu Liu 0023
ICME1
2025 Uneven Event Modeling for Partially Relevant Video Retrieval
abstract
Given a text query, partially relevant video retrieval (PRVR) aims to retrieve untrimmed videos containing relevant moments, wherein event modeling is crucial for partitioning the video into smaller temporal events that partially correspond to the text. Previous methods typically segment videos into a fixed number of equal-length clips, resulting in ambiguous event boundaries. Additionally, they rely on mean pooling to compute event representations, inevitably introducing undesired misalignment. To address these, we propose an Uneven Event Modeling (UEM) framework for PRVR. We first introduce the Progressive-Grouped Video Segmentation (PGVS) module, to iteratively formulate events in light of both temporal dependencies and semantic similarity between consecutive frames, enabling clear event boundaries. Furthermore, we also propose the Context-Aware Event Refinement (CAER) module to refine the event representation conditioned the text’s cross-attention. This enables event representations to focus on the most relevant frames for a given text, facilitating more precise text-video alignment. Extensive experiments demonstrate that our method achieves state-of-the-art performance on two PRVR benchmarks. Code is available at https://github.com/Sasa77777779/UEM.git.
Sa Zhu, Huashan Chen, Wanqian Zhang, Jinchao Zhang 0002, Zexian Yang, Xiaoshuai Hao, Bo Li 0063
ICME6
2025 SafeMap: Robust HD Map Construction from Incomplete Observations
abstract
Robust high-definition (HD) map construction is vital for autonomous driving, yet existing methods often struggle with incomplete multi-view camera data. This paper presents SafeMap, a novel framework specifically designed to ensure accuracy even when certain camera views are missing. SafeMap integrates two key components: the Gaussian-based Perspective View Reconstruction (G-PVR) module and the Distillation-based Bird’s-Eye-View (BEV) Correction (D-BEVC) module. G-PVR leverages prior knowledge of view importance to dynamically prioritize the most informative regions based on the relationships among available camera views. Furthermore, D-BEVC utilizes panoramic BEV features to correct the BEV representations derived from incomplete observations. Together, these components facilitate comprehensive data reconstruction and robust HD map generation. SafeMap is easy to implement and integrates seamlessly into existing systems, offering a plug-and-play solution for enhanced robustness. Experimental results demonstrate that SafeMap significantly outperforms previous methods in both complete and incomplete scenarios, highlighting its superior performance and resilience.
Xiaoshuai Hao, Lingdong Kong, Rong Yin 0001, Pengwei Wang 0004, Jing Zhang 0037, Yunfeng Diao, Shu Zhao 0006
ICML1
2025 Open-Vocabulary Fine-Grained Hand Action Detection
abstract
In this work, we address the new challenge of open-vocabulary fine-grained hand action detection, which aims to recognize hand actions from both known and novel categories using textual descriptions. Traditional hand action detection methods are limited to closed-set detection, making it difficult for them to generalize to new, unseen hand action categories. While current open-vocabulary detection (OVD) methods are effective at detecting novel objects, they face challenges with fine-grained action recognition, particularly when data is limited and heterogeneous. This often leads to poor generalization and performance bias between base and novel categories. To address these issues, we propose a novel approach, Open-FGHA (Open-vocabulary Fine-Grained Hand Action), which learns to distinguish fine-grained features across multiple modalities from limited heterogeneous data. It then identifies optimal matching relationships among these features, enabling accurate open-vocabulary fine-grained hand action detection. Specifically, we introduce three key components: Hierarchical Heterogeneous Low-Rank Adaptation, Bidirectional Selection and Fusion Mechanism, and Cross-Modality Query Generator. These components work in unison to enhance the alignment and fusion of multimodal fine-grained features. Extensive experiments demonstrate that Open-FGHA outperforms existing OVD methods, showing its strong potential for open-vocabulary hand action detection. The source code is available at OV-FGHAD.
Ting Zhe, Mengya Han, Xiaoshuai Hao, Yong Luo 0002, Zheng He 0001, Xiantao Cai, Jing Zhang 0037
IJCAI3
2025 What Really Matters for Robust Multi-Sensor HD Map Construction?
abstract
High-definition (HD) map construction methods are crucial for providing precise and comprehensive static environmental information, which is essential for autonomous driving systems. While Camera-LiDAR fusion techniques have shown promising results by integrating data from both modalities, existing approaches primarily focus on improving model accuracy, often neglecting the robustness of perception models—a critical aspect for real-world applications. In this paper, we explore strategies to enhance the robustness of multi-modal fusion methods for HD map construction while maintaining high accuracy. We propose three key components: data augmentation, a novel multi-modal fusion module, and a modality dropout training strategy. These components are evaluated on a challenging dataset containing 13 types of multi-sensor corruption. Experimental results demonstrate that our proposed modules significantly enhance the robustness of baseline methods. Furthermore, our approach achieves state-of-the-art performance on the clean validation set of the NuScenes dataset. Our findings provide valuable insights for developing more robust and reliable HD map construction models, advancing their applicability in real-world autonomous driving scenarios. Project website: https://robomap-123.github.io/.
Xiaoshuai Hao, Yuheng Ji, Luanyuan Dai, Peng Hao 0003, Dingzhe Li, Shuai Cheng 0002, Rong Yin 0001
IROS1
2025 AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
abstract
Inferring the affordance of an object and grasping it in a task-oriented manner is crucial for robots to successfully complete manipulation tasks. Affordance indicates where and how to grasp an object by taking its functionality into account, serving as the foundation for effective task-oriented grasping. However, current task-oriented methods often depend on extensive training data that is confined to specific tasks and objects, making it difficult to generalize to novel objects and complex scenes. In this paper, we introduce AffordGrasp, a novel open-vocabulary grasping framework that leverages the reasoning capabilities of vision-language models (VLMs) for in-context affordance reasoning. Unlike existing methods that rely on explicit task and object specifications, our approach infers tasks directly from implicit user instructions, enabling more intuitive and seamless human-robot interaction in everyday scenarios. Building on the reasoning outcomes, our framework identifies task-relevant objects and grounds their part-level affordances using a visual grounding module. This allows us to generate task-oriented grasp poses precisely within the affordance regions of the object, ensuring both functional and context-aware robotic manipulation. Extensive experiments demonstrate that AffordGrasp achieves state-of-the-art performance in both simulation and real-world scenarios, highlighting the effectiveness of our method. We believe our approach advances robotic manipulation techniques and contributes to the broader field of embodied AI. Project website: https://eqcy.github.io/affordgrasp/.
Yingbo Tang, Shuaike Zhang, Xiaoshuai Hao, Pengwei Wang 0004, Jianlong Wu, Zhongyuan Wang 0006, Shanghang Zhang
IROS3
2025 Enhancing Adversarial Robustness of Vision-Language Models through Low-Rank Adaptation
abstract
Vision-Language Models (VLMs) play a crucial role in the advancement of Artificial General Intelligence (AGI). As AGI rapidly evolves, addressing security concerns has emerged as one of the most significant challenges for VLMs. In this paper, we present extensive experiments that expose the vulnerabilities of conventional adaptation methods for VLMs, highlighting significant security risks. Moreover, as VLMs grow in size, the application of traditional adversarial adaptation techniques incurs substantial computational costs. To address these issues, we propose a parameter-efficient adversarial adaptation method called AdvLoRA based on Low-Rank Adaptation. We investigate and reveal the inherent low-rank properties involved in adversarial adaptation for VLMs. Different from LoRA, we enhance the efficiency and robustness of adversarial adaptation by introducing a novel reparameterization method that leverages parameter clustering and alignment. Additionally, we propose an adaptive parameter update strategy to further bolster robustness. These innovations enable our AdvLoRA to mitigate issues related to model security and resource wastage. Extensive experiments confirm the effectiveness and efficiency of AdvLoRA.
Yuheng Ji, Yue Liu 0008, Zhao Zhang 0002, Xiaoshuai Hao, Gang Zhou 0001, Xingwei Zhang, Xiaolong Zheng 0001
ICMR6
2025 SVD: Spatial Video Dataset
abstract
Stereoscopic video has long been the subject of research due to its ability to deliver immersive three-dimensional content to a wide range of applications. The dual-view format inherently provides binocular disparity cues that enhance depth perception and realism, making it indispensable for fields such as telepresence, 3D mapping, and robotic vision. Until recently, however, end-to-end pipelines for capturing, encoding, and viewing high-quality stereoscopic video were neither widely accessible nor optimized for consumer-grade devices. Today's smartphones, such as the iPhone Pro, and modern Head-Mounted Displays (HMDs) like the Apple Vision Pro, offer built-in support for stereoscopic video capture, hardware-accelerated encoding, and seamless playback on devices like the Apple Vision Pro and Meta Quest 3, which require minimal user intervention. Apple refers to this streamlined workflow as spatial Video. Making the full stereoscopic video process available to everyone has made new applications possible. Despite these advances, there remains a notable absence of publicly available datasets that include the complete spatial video pipeline on consumer platforms, hindering reproducibility and comparative evaluation of emerging algorithms.
Mohammad Hossein Izadimehr, Milad Ghanbari, Guodong Chen 0004, Wei Zhou 0021, Xiaoshuai Hao, Mallesham Dasari, Christian Timmerer, Hadi Amirpour
ACM Multimedia5
2025 MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios
abstract
Rapid advances in Artificial Intelligence Generated Content (AIGC) have enabled increasingly sophisticated face forgeries, posing a significant threat to social security. However, current Deepfake detection methods are limited by constraints in existing datasets, which lack the diversity necessary in real-world scenarios. Specifically, these data sets fall short in four key areas: unknown of advanced forgery techniques, variability of facial scenes, richness of real data, and degradation of real-world propagation. To address these challenges, we propose the Multi-dimensional Face Forgery Image (MFFI ) dataset, tailored for real-world scenarios. MFFI enhances realism based on four strategic dimensions: 1) Wider Forgery Methods; 2) Varied Facial Scenes; 3) Diversified Authentic Data; 4) Multi-level Degradation Operations. MFFI integrates 50 different forgery methods and contains 1024K image samples. Benchmark evaluations show that MFFI outperforms existing public datasets in terms of scene complexity, cross-domain generalization capability, and detection difficulty gradients. These results validate the technical advance and practical utility of MFFI in simulating real-world conditions. The dataset and additional details are publicly available at https://github.com/inclusionConf/MFFI.
Changtao Miao, Weiwei Feng, Qi Chu 0001, Jianshu Li, Yunfeng Diao, Wei Zhou 0021, Joey Tianyi Zhou, Xiaoshuai Hao
ACM Multimedia12
2025 RoboAfford: A Dataset and Benchmark for Enhancing Object and Spatial Affordance Learning in Robot Manipulation
abstract
Robot manipulation is a fundamental capability of embodied intelligence, enabling effective robot interactions with the physical world. In robotic manipulation tasks, predicting precise grasping positions and object placement is essential. Achieving this requires object recognition to localize target object, predicting object affordances for interaction and spatial affordances for optimal arrangement. While Vision-Language Models (VLMs) provide insights for high-level task planning and scene understanding, they often struggle to predict precise action positions, such as functional grasp points and spatial placements. This limitation stems from the lack of annotations for object and spatial affordance data in their training datasets. To address this gap, we introduce RoboAfford , a novel large-scale dataset designed to enhance object and spatial affordance learning in robot manipulation. Our dataset comprises 819,987 images paired with 1.9 million question answering (QA) annotations, covering three critical tasks: object affordance recognition to identify objects based on attributes and spatial relationships, object affordance prediction to pinpoint functional grasping parts, and spatial affordance localization to identify free space for placement. Complementing this dataset, we propose RoboAfford-Eval , a comprehensive benchmark for assessing affordance-aware prediction in real-world scenarios, featuring 338 meticulously annotated samples across the same three tasks. Extensive experimental results reveal the deficiencies of existing VLMs in affordance learning, while fine-tuning on the RoboAfford dataset significantly enhances their affordance prediction in robot manipulation, validating the dataset's effectiveness. The dataset, benchmark and evaluation code will be made publicly available to facilitate future research. Project website: https://roboafford-dataset.github.io/.
Yingbo Tang, Yinuo Zhao, Xiaoshuai Hao
ACM Multimedia5
2025 Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
abstract
Video content comprehension is essential for various applications, ranging from video analysis to interactive systems. Despite advancements in large-scale vision-language models (VLMs), these models often struggle to capture the nuanced, spatiotemporal details essential for thorough video analysis. To address this gap, we introduce Video-CoT, a groundbreaking dataset designed to enhance spatiotemporal understanding using Chain-of-Thought(CoT) methodologies. Video-CoT contains 192,000 fine-grained spatiotemporal question-answer pairs and 23,000 high-quality CoT-annotated samples, providing a solid foundation for evaluating spatiotemporal understanding in video comprehension. Addition- ally, we provide a comprehensive benchmark for assessing these tasks, with each task featuring 750 images and tailored evaluation metrics. Our extensive experiments reveal that current VLMs face significant challenges in achieving satisfactory performance, high- lighting the difficulties of effective spatiotemporal understanding. Overall, the Video-CoT dataset and benchmark open new avenues for research in multimedia understanding and support future innovations in intelligent systems requiring advanced video analysis capabilities. By making these resources publicly available, we aim to encourage further exploration in this critical area. Project website: https://video-cot.github.io/ .
Xiaoshuai Hao, Yingbo Tang, Pengwei Wang 0005, Zhongyuan Wang 0006, Hongxuan Ma, Shanghang Zhang
ACM Multimedia2
2025 FastRSR: Efficient and Accurate Road Surface Reconstruction in Bird's Eye View
abstract
Road Surface Reconstruction (RSR) is crucial for autonomous driving, enabling the understanding of road surface conditions. Recently, RSR from the Bird's Eye View (BEV) has gained attention for its potential to enhance performance. However, existing methods for transforming perspective views to BEV face challenges such as information loss and representation sparsity. Moreover, stereo matching in BEV is limited by the need to balance accuracy with inference speed. To address these challenges, we propose two efficient and accurate BEV-based RSR models: FastRSR-mono and FastRSR-stereo. Specifically, we first introduce Depth-Aware Projection (DAP), an efficient view transformation strategy designed to mitigate information loss and sparsity by querying depth and image features to aggregate BEV data within specific road surface regions using a pre-computed look-up table. To optimize accuracy and speed in stereo matching, we design the Spatial Attention Enhancement (SAE) and Confidence Attention Generation (CAG) modules. SAE adaptively highlights important regions, while CAG focuses on high-confidence predictions and filters out irrelevant information. FastRSR achieves state-of-the-art performance, exceeding monocular competitors by over 6.0% in elevation absolute error and providing at least a 3.0× speedup by stereo methods on the RSRD dataset. The source code will be released.
Yuheng Ji, Xiaoshuai Hao, Shuxiao Li
ACM Multimedia3
2025 ESC-MISR: Enhancing Spatial Correlations for Multi-image Super-Resolution in Remote Sensing
Xiaoshuai Hao, Jianan Li 0001, Jinhui Pang
MMM (1)2
2025 SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed Graphs
abstract
Large-scale pre-trained models have revolutionized Natural Language Processing (NLP) and Computer Vision (CV), showcasing remarkable cross-domain generalization abilities. However, in graph learning, models are typically trained on individual graph datasets, limiting their capacity to transfer knowledge across different graphs and tasks. This approach also heavily relies on large volumes of annotated data, which presents a significant challenge in resource-constrained settings. Unlike NLP and CV, graph-structured data presents unique challenges due to its inherent heterogeneity, including domain-specific feature spaces and structural diversity across various applications. To address these challenges, we propose a novel structure-aware self-supervised learning method for Text-Attributed Graphs (SSTAG). By leveraging text as a unified representation medium for graph learning, SSTAG bridges the gap between the semantic reasoning of Large Language Models (LLMs) and the structural modeling capabilities of Graph Neural Networks (GNNs). Our approach introduces a dual knowledge distillation framework that co-distills both LLMs and GNNs into structure-aware multilayer perceptrons (MLPs), enhancing the scalability of large-scale TAGs. Additionally, we introduce an in-memory mechanism that stores typical graph representations, aligning them with memory anchors in an in-memory repository to integrate invariant knowledge, thereby improving the model’s generalization ability. Extensive experiments demonstrate that SSTAG outperforms state-of-the-art models on cross-domain transfer learning tasks, achieves exceptional scalability, and reduces inference costs while maintaining competitive performance.
Ruyue Liu, Rong Yin 0001, Xiangzhen Bo, Xiaoshuai Hao, Yong Liu 0018, Jinwen Zhong, Can Ma, Weiping Wang 0005
NeurIPS4
2025 Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
abstract
Visual reasoning abilities play a crucial role in understanding complex multimodal data, advancing both domain-specific applications and artificial general intelligence (AGI). Existing methods enhance Vision-Language Models (VLMs) through Chain-of-Thought (CoT) supervised fine-tuning using meticulously annotated data. However, this approach may lead to overfitting and cognitive rigidity, limiting the model’s generalization ability under domain shifts and reducing real-world applicability. To overcome these limitations, we propose Reason-RFT, a two-stage reinforcement fine-tuning framework for visual reasoning. First, Supervised Fine-Tuning (SFT) with curated CoT data activates the reasoning potential of VLMs. This is followed by reinforcement learning based on Group Relative Policy Optimization (GRPO), which generates multiple reasoning-response pairs to enhance adaptability to domain shifts. To evaluate Reason-RFT, we reconstructed a comprehensive dataset covering visual counting, structural perception, and spatial transformation, serving as a benchmark for systematic assessment across three key dimensions. Experimental results highlight three advantages: (1) performance enhancement, with Reason-RFT achieving state-of-the-art results and outperforming both open-source and proprietary models; (2) generalization superiority, maintaining robust performance under domain shifts across various tasks; and (3) data efficiency, excelling in few-shot learning scenarios and surpassing full-dataset SFT baselines. Reason-RFT introduces a novel training paradigm for visual reasoning and marks a significant step forward in multimodal research.
Huajie Tan, Yuheng Ji, Xiaoshuai Hao, Xiansheng Chen, Pengwei Wang 0004, Zhongyuan Wang 0006, Shanghang Zhang
NeurIPS3
2025 VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment: Methods and Results
abstract
This paper reviews the VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment. The competition aims to foster innovation in both subjective and objective VQA techniques tailored to live broadcasting videos, addressing the unique challenges posed by live streaming impairments while emphasizing the evaluation of QoE. The grand challenge used live broadcasting database LBVD which consists of 1013 videos focusing on distortion in live broadcasting videos. The competition had 14 participants and 5 teams submitted valid solutions for the final testing phase. The proposed solutions have shown significant progress in areas such as combining traditional feature engineering with deep learning models, achieved state-of-the-art performances for LBVD. Team ATHENA-Live-QoE and Team HZX Force tied for the first position. The dataset can be found at https://github.com/cpf0079/LBVD.
Wenqi Fei, Yuhua Zhang, MohammadAli Hamidi, Hadi Amirpour, Erjia Xiao, Zhenjie Su, Hao Cheng 0015, Yu Liu 0023, Wei Zhou 0021, Yanbiao Ma, Renjing Xu, Long Chen 0015, Xiaoshuai Hao, Yipo Huang, Tushar Shinde
VCIP16
2025 STViT+: improving self-supervised multi-camera depth estimation with spatial-temporal context and adversarial geometry regularization
abstract
Abstract Multi-camera depth estimation has gained significant attention in autonomous driving due to its importance in perceiving complex environments. However, extending monocular self-supervised methods to multi-camera setups introduces unique challenges that existing techniques often fail to address. In this paper, we propose STViT+ , a novel Transformer-based framework for self-supervised multi-camera depth estimation. Our key contributions include: 1) the Spatial-Temporal Transformer (STTrans) , which integrates local spatial connectivity and global context to capture enriched spatial-temporal cross-view correlations, resulting in more accurate 3D geometry reconstruction; 2) the Spatial-Temporal Photometric Consistency Correction (STPCC) strategy that mitigates the impact of varying illumination, ensuring brightness consistency across frames during photometric loss calculation; 3) the Adversarial Geometry Regularization (AGR) module, which employs Generative Adversarial Networks to impose spatial constraints by using unpaired depth maps, enhancing performance under adverse conditions such as rain and nighttime driving. Extensive evaluations on large-scale autonomous driving datasets, including Nuscenes and DDAD, confirm that STViT+ sets a new benchmark for multi-camera depth estimation.
Zhuo Chen 0040, Haimei Zhao, Xiaoshuai Hao, Bo Yuan 0003, Xiu Li 0001
Appl. Intell.3
2025 A universal sampling method based on feature and structural comprehensive proximity measure
Jinhui Pang, Cheng Shang, Ziyu Jia, Peng Hao 0003, Xiaoshuai Hao
Neurocomputing6
2025 A hierarchical reinforcement learning framework for multi-UAV combat using leader-follower strategy
Jinhui Pang, Jinglin He, Noureldin Mohamed Abdelaal Ahmed Mohamed, Xiaoshuai Hao
Knowl. Based Syst.6
2025 Multi-Modal Molecular Representation Learning via Structure Awareness
abstract
Accurate extraction of molecular representations is a critical step in the drug discovery process. In recent years, significant progress has been made in molecular representation learning methods, among which multi-modal molecular representation methods based on images, and 2D/3D topologies have become increasingly mainstream. However, existing these multi-modal approaches often directly fuse information from different modalities, overlooking the potential of intermodal interactions and failing to adequately capture the complex higher-order relationships and invariant features between molecules. To overcome these challenges, we propose a structure-awareness-based multi-modal self-supervised molecular representation pre-training framework (MMSA) designed to enhance molecular graph representations by leveraging invariant knowledge between molecules. The framework consists of two main modules: the multi-modal molecular representation learning module and the structure-awareness module. The multi-modal molecular representation learning module collaboratively processes information from different modalities of the same molecule to overcome intermodal differences and generate a unified molecular embedding. Subsequently, the structure-awareness module enhances the molecular representation by constructing a hypergraph structure to model higher-order correlations between molecules. This module also introduces a memory mechanism for storing typical molecular representations, aligning them with memory anchors in the memory bank to integrate invariant knowledge, thereby improving the model's generalization ability. Compared to existing multi-modal approaches, MMSA can be seamlessly integrated with any graph-based method and supports multiple molecular data modalities, ensuring both versatility and compatibility. Extensive experiments have demonstrated the effectiveness of MMSA, which achieves state-of-the-art performance on the MoleculeNet benchmark, with average ROC-AUC improvements ranging from 1.8% to 9.6% over baseline methods.
Rong Yin 0001, Ruyue Liu, Xiaoshuai Hao, Xingrui Zhou, Yong Liu 0018, Can Ma, Weiping Wang 0005
IEEE Trans. Image Process.3
2025 AS-GCL: Asymmetric Spectral Augmentation on Graph Contrastive Learning
abstract
Graph Contrastive Learning (GCL) has emerged as the foremost approach for self-supervised learning on graph-structured data. GCL reduces reliance on labeled data by learning robust representations from various augmented views. However, existing GCL methods typically depend on consistent stochastic augmentations, which overlook their impact on the intrinsic structure of the spectral domain, thereby limiting the model's ability to generalize effectively. To address these limitations, we propose a novel paradigm called AS-GCL that incorporates asymmetric spectral augmentation for graph contrastive learning. A typical GCL framework consists of three key components: graph data augmentation, view encoding, and contrastive loss. Our method introduces significant enhancements to each of these components. Specifically, for data augmentation, we apply spectral-based augmentation to minimize spectral variations, strengthen structural invariance, and reduce noise. With respect to encoding, we employ parameter-sharing encoders with distinct diffusion operators to generate diverse, noise-resistant graph views. For contrastive loss, we introduce an upper-bound loss function that promotes generalization by maintaining a balanced distribution of intra- and inter-class distance. To our knowledge, we are the first to encode augmentation views of the spectral domain using asymmetric encoders. Extensive experiments on eight benchmark datasets across various node-level tasks demonstrate the advantages of the proposed method.
Ruyue Liu, Rong Yin 0001, Yong Liu 0018, Xiaoshuai Hao, Haichao Shi, Can Ma, Weiping Wang 0005
IEEE Trans. Multim.4
2024 Enhancing 3D Hand Pose Estimation via Dense Ordinal Regression Network
Yamin Mao, SoonYong Cho, Xiaoshuai Hao
BMVC6
2024 MapDistill: Boosting Efficient Camera-Based HD Map Construction via Camera-LiDAR Fusion Model Distillation
Xiaoshuai Hao, Ruikai Li, Hui Zhang 0093, Dingzhe Li, Rong Yin 0001, Sangil Jung, Seung In Park, ByungIn Yoo, Haimei Zhao, Jing Zhang 0037
ECCV (3)1
2024 Customized Treatment Per Pixel for Blind Image Super-Resolution
abstract
Blind image super-resolution task aims at restoring high-resolution images from their low-resolution counterparts by reversing the unknown degradation. Existing methods have achieved promising results when handling degradation with isotropic or anisotropic gaussian blur, whereas suffer from performance drop when addressing degradation with motion blur. Compared with gaussian blur, motion blur is more diverse, i.e., each pixel moves a peculiar distance in distinct orientation, thereby each pixel requires individual treatment. To tackle this degradation with motion blur issue, we propose a novel blind image super-resolution method named deformAble receptive Super Resolution (ArcSR), which provides deformable receptive field and unique parameters for each pixel. Specifically, we propose Deformable Mutual convolution (DMconv) and Kernel Guided convolution (KGconv) for blur kernel estimation and super-resolution, respectively. The DMconv explores the correlation within channels of image features and achieves deformable receptive field by redesigning deformable convolution to generate kernels for each pixel. Meanwhile, the KGconv views the estimated kernel as an attention matrix for convolutional parameters and gives each pixel disparate convolutional parameters to rewrite the missing high-frequency information. Comprehensive experiments demonstrate the superiority of our method.
Guanqun Liu 0005, Xiaoshuai Hao
ICASSP2
2024 MBFusion: A New Multi-modal BEV Feature Fusion Method for HD Map Construction
abstract
HD map construction is a fundamental and challenging task in autonomous driving to understand the surrounding environment. Recently, Camera-LiDAR BEV feature fusion methods have attracted increasing attention in HD map construction task, which can significantly boost the benchmark. However, existing fusion methods ignore modal interaction and utilize very simple fusion strategy, which suffers from the problems of misalignment and information loss. To tackle this, we propose a novel Multi-modal BEV feature fusion method named MBFusion. Specifically, to solve the semantic misalignment problem between Camera and LiDAR features, we design Cross-modal Interaction Transform (CIT) module to make these two feature spaces interact knowledge with each other to enhance the feature representation by the cross-attention mechanism. Then, we propose a Dual Dynamic Fusion (DDF) module to automatically select valuable information from different modalities for better feature fusion. Moreover, MBFusion is simple, and can be plug-and-played into existing pipelines. We evaluate MBFusion on three architectures, including HDMapNet, VectorMapNet, and MapTR, to show its versatility and effectiveness. Compared with the state-of-the-art methods, MBFusion achieves 3.6% and 4.1% absolute improvements on mAP on the nuScenes and the Argoverse2 datasets, respectively, demonstrating the superiority of our method.
Xiaoshuai Hao, Hui Zhang 0093, Yifan Yang 0007, Yi Zhou 0020, Sangil Jung, Seung In Park, ByungIn Yoo
ICRA1
2024 FTF-ER: Feature-Topology Fusion-Based Experience Replay Method for Continual Graph Learning
abstract
Continual graph learning (CGL) is an important and challenging task that aims to extend static GNNs to dynamic task flow scenarios. As one of the mainstream CGL methods, the experience replay (ER) method receives widespread attention due to its superior performance. However, existing ER methods focus on identifying samples by feature significance or topological relevance, which limits their utilization of comprehensive graph data. In addition, the topology-based ER methods only consider local topological information and add neighboring nodes to the buffer, which ignores the global topological information and increases memory overhead. To bridge these gaps, we propose a novel method called Feature-Topology Fusion-based Experience Replay (FTF-ER) to effectively mitigate the catastrophic forgetting issue with enhanced efficiency. Specifically, from an overall perspective to maximize the utilization of the entire graph data, we propose a highly complementary approach including both feature and global topological information, which can significantly improve the effectiveness of the sampled nodes. Moreover, to further utilize global topological information, we propose Hodge Potential Score (HPS) as a novel module to calculate the topological importance of nodes. HPS derives a global node ranking via Hodge decomposition on graphs, providing more accurate global topological information compared to neighbor sampling. By excluding neighbor sampling, HPS significantly reduces buffer storage costs for acquiring topological information and simultaneously decreases training time. Compared with state-of-the-art methods, FTF-ER achieves a significant improvement of 3.6% in AA and 7.1% in AF on the OGB-Arxiv dataset, demonstrating its superior performance in the class-incremental learning setting.
Jinhui Pang, Xiaoshuai Hao, Rong Yin 0001, Zixuan Wang 0021, Jinglin He, Huang Tai Sheng
ACM Multimedia3
2024 Is Your HD Map Constructor Reliable under Sensor Corruptions?
abstract
Driving systems often rely on high-definition (HD) maps for precise environmental information, which is crucial for planning and navigation. While current HD map constructors perform well under ideal conditions, their resilience to real-world challenges, \eg, adverse weather and sensor failures, is not well understood, raising safety concerns. This work introduces MapBench, the first comprehensive benchmark designed to evaluate the robustness of HD map construction methods against various sensor corruptions. Our benchmark encompasses a total of 29 types of corruptions that occur from cameras and LiDAR sensors. Extensive evaluations across 31 HD map constructors reveal significant performance degradation of existing methods under adverse weather conditions and sensor failures, underscoring critical safety concerns. We identify effective strategies for enhancing robustness, including innovative approaches that leverage multi-modal fusion, advanced data augmentation, and architectural techniques. These insights provide a pathway for developing more reliable HD map construction methods, which are essential for the advancement of autonomous driving technology. The benchmark toolkit and affiliated code and model checkpoints have been made publicly accessible.
Xiaoshuai Hao, Mengchuan Wei, Yifan Yang 0007, Haimei Zhao, Hui Zhang 0093, Yi Zhou 0020, Lingdong Kong, Jing Zhang 0037
NeurIPS1
2023 Dual Alignment Unsupervised Domain Adaptation for Video-Text Retrieval
abstract
Video-text retrieval is an emerging stream in both computer vision and natural language processing communities, which aims to find relevant videos given text queries. In this paper, we study the notoriously challenging task, i.e., Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), wherein training and testing data come from different distributions. Previous works merely alleviate the domain shift, which however overlook the pairwise misalignment issue in target domain, i.e., there exist no semantic relationships between target videos and texts. To tackle this, we propose a novel method named Dual Alignment Domain Adaptation (DADA). Specifically, we first introduce the cross-modal semantic embedding to generate discriminative source features in a joint embedding space. Besides, we utilize the video and text domain adaptations to smoothly balance the minimization of the domain shifts. To tackle the pairwise misalignment in target domain, we propose the Dual Alignment Consistency (DAC) to fully exploit the semantic information of both modalities in target domain. The proposed DAC adaptively aligns the video-text pairs which are more likely to be relevant in target domain, enabling that positive pairs are increasing progressively and the noisy ones will potentially be aligned in the later stages. To that end, our method can generate more truly aligned target pairs and ensure the discriminability of target features. Compared with the state-of-the-art methods, DADA achieves 20.18% and 18.61% relative improvements on R@1 under the setting of TGIF→MSR-VTT and TGIF→MSVD respectively, demonstrating the superiority of our method.
Xiaoshuai Hao, Wanqian Zhang, Dayan Wu, Bo Li 0063
CVPR1
2023 Uncertainty-Aware Alignment Network for Cross-Domain Video-Text Retrieval
abstract
Video-text retrieval is an important but challenging research task in the multimedia community. In this paper, we address the challenge task of Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), assuming that training (source) data and testing (target) data are from different domains. Previous approaches are mostly derived from classification based domain adaptation methods, which are neither multi-modal nor suitable for retrieval task. In addition, as to the pairwise misalignment issue in target domain, i.e., no pairwise annotations between target videos and texts, the existing method assumes that a video corresponds to a text. Yet we empirically find that in the real scene, one text usually corresponds to multiple videos and vice versa. To tackle this one-to-many issue, we propose a novel method named Uncertainty-aware Alignment Network (UAN). Specifically, we first introduce the multimodal mutual information module to balance the minimization of domain shift in a smooth manner. To tackle the multimodal uncertainties pairwise misalignment in target domain, we propose the Uncertainty-aware Alignment Mechanism (UAM) to fully exploit the semantic information of both modalities in target domain. Extensive experiments in the context of domain-adaptive video-text retrieval demonstrate that our proposed method consistently outperforms multiple baselines, showing a superior generalization ability for target data.
Xiaoshuai Hao, Wanqian Zhang
NeurIPS1
2022 Listen and Look: Multi-Modal Aggregation and Co-Attention Network for Video-Audio Retrieval
abstract
Video is a natural source of multi-modal data with intrinsic correlations between different modalities, such as objects, motions and captions. Though intuitive, such inherent supervision has not been well explored in previous video-audio retrieval works. Besides, existing methods exploit the video stream and the audio stream seperately, whereas ignoring the mutual interactions between them. In this paper, we propose a two-stream model named Multi-modal Aggregation and Co-attention network (MAC), which processes video and audio inputs with co-attentional interactions. Specifically, our method takes raw videos as inputs and extracts aggregated features from multiple modalities to benefit the video representation learning. Then, we introduce the self-attention mechanism to make videos adaptively assign higher weights to the representative modalities. Moreover, we introduce a coattention transformer module to better capture the relations among videos and audios. By exchanging key-value pairs in the multi-headed attention, this module enables video-attended audio features to be incorporated into video representations and vice versa. Experiments show that our method significantly outperform other state-of-the-arts.
Xiaoshuai Hao, Wanqian Zhang, Dayan Wu, Bo Li 0063
ICME1
2021 What Matters: Attentive and Relational Feature Aggregation Network for Video-Text Retrieval
abstract
Cross-modal video-text retrieval has been an emerging task due to the rapid growth of user-generated videos on the Internet. Most existing approaches focus on extracting visual feature for the video, while audio and caption on the screen containing rich information are ignored. Recently, the aggregations of multi-modal features in videos boost the benchmark of video-text retrieval. However, since these multi-modal features are high-dimensional and heterogeneous, their intrinsically structural relations have not been attached with enough importance and are often overlooked in previous methods. To address this issue, we propose a novel Attentive and Relational Feature Aggregation Network (ARFAN). Specifically, we introduce the self-attention mechanism to make videos adaptively assign higher weights to the representative modalities. Then, the graph convolutional layers are inserted to capture the relations among the multi-modal features to combine them. Our method achieves 15% and 12.9% relative improvements on R@1 when compared with the state-of-the-art method on MSR-VTT and MSVD datasets, respectively.
Xiaoshuai Hao, Yucan Zhou, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002
ICME1
2021 Multi-Feature Graph Attention Network for Cross-Modal Video-Text Retrieval
abstract
Cross-modal retrieval between videos and texts has attracted growing attention due to the rapid growth of user-generated videos on the web. To solve this problem, most approaches try to learn a joint embedding space to measure the cross-modal similarities, while paying little attention to the representation of each modality. Video is more complicated than the commonly used visual feature, since the audio and caption on the screen also contain rich information. Recently, the aggregations of multiple features in videos boost the benchmark of the video-text retrieval system. However, they usually handle each feature independently, which ignores the interchange of high-level semantic relations among these multiple features. Moreover, despite the inter-modal ranking constraint where semantically-similar texts and videos should stay closer, the modality-specific requirement, i.e. two similar videos/texts should have similar representations, is also significant. In this paper, we propose a novel Multi-Feature Graph ATtention Network (MFGATN) for cross-modal video-text retrieval. Specifically, we introduce a multi-feature graph attention module, which enriches the representation of each feature in videos with the interchange of high-level semantic information among them. Moreover, we elaborately design a novel Dual Constraint Ranking Loss (DCRL), which simultaneously considers the inter-modal ranking constraint and the intra-modal structure constraint to preserve both the cross-modal semantic similarity and the modality-specific consistency in the embedding space. Experiments on two datasets, i.e. MSR-VTT and MSVD, demonstrate that our method achieves significant performance gain compared with the state-of-the-arts.
Xiaoshuai Hao, Yucan Zhou, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005
ICMR1