Hao Sun 0038

dblp:82/2248-38 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
13since 2021 · last 2026
0009-0007-7917-1628ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval
abstract
In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in aligning with complex or ambiguous moments. To overcome these limitations, we explore Deep Evidential Regression (DER) to construct a vanilla Evidential baseline. However, this approach encounters two major issues: the inability to effectively handle modality imbalance and the structural differences in DER's heuristic uncertainty regularizer, which adversely affect uncertainty estimation. This misalignment results in high uncertainty being incorrectly associated with accurate samples rather than challenging ones. Our observations indicate that existing methods lack the adaptability required for complex video scenarios. In response, we propose Debiased Evidential Learning for Moment Retrieval (DEMR), a novel framework that incorporates a Reflective Flipped Fusion (RFF) block for cross-modal alignment and a query reconstruction task to enhance text sensitivity, thereby reducing bias in uncertainty estimation. Additionally, we introduce a Geom-regularizer to refine uncertainty predictions, enabling adaptive alignment with difficult moments and improving retrieval accuracy. Extensive testing on standard datasets and debiased datasets ActivityNet-CD and Charades-CD demonstrates significant enhancements in effectiveness, robustness, and interpretability, positioning our approach as a promising solution for temporal-semantic robustness in moment retrieval.
Haojian Huang, Kaijing Ma, Xianghao Zang, Han Fang 0002, Chao Ban, Hao Sun 0038, Mulin Chen, Zhongjiang He
AAAI9
2026 Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing
abstract
Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and the domain gap between natural and RS images. To bridge these gaps, we first establish a standardized OVRSIS benchmark (OVRSISBench) based on widely-used RS segmentation datasets, enabling consistent evaluation across methods. Using this benchmark, we comprehensively evaluate several representative OVS/OVRSIS models and reveal their limitations when directly applied to remote sensing scenarios. Building on these insights, we propose RSKT-Seg, a novel open-vocabulary segmentation framework tailored for remote sensing. RSKT-Seg integrates three key components: (1) a Multi-Directional Cost Map Aggregation (RS-CMA) module that captures rotation-invariant visual cues by computing vision-language cosine similarities across multiple directions; (2) an Efficient Cost Map Fusion (RS-Fusion) transformer, which jointly models spatial and semantic dependencies with a lightweight dimensionality reduction strategy; and (3) a Remote Sensing Knowledge Transfer (RS-Transfer) module that injects pre-trained knowledge and facilitates domain adaptation via enhanced upsampling. Extensive experiments on the benchmark show that RSKT-Seg consistently outperforms strong OVS baselines by +3.8 mIoU and +5.9 mACC, while achieving 2× faster inference through efficient aggregation.
Bingyu Li 0002, Haocheng Dong, Da Zhang 0010, Zhiyuan Zhao 0005, Hao Sun 0038, Junyu Gao 0001
AAAI5
2026 VAV-R1: Difficulty-Aware Multimodal Reasoning for Video Anomaly Validation
abstract
Video anomaly validation (VAV) serves as the final alarm validation step for false positive filtering in video anomaly detection (VAD), requiring both higher accuracy in anomaly identification and stronger justifications. Despite rapid VAD advancements, existing methods still lack sufficient interpretability and struggle with challenging cases. To address this, we propose VAV-R1, a multimodal reasoning model tailored for VAV. We first construct ThinkVAV, a dedicated benchmark with fine-grained reasoning annotations across diverse anomaly types. Furthermore, we introduce DA-GRPO, a difficulty-aware reinforcement learning strategy that prioritizes learning from more challenging cases. Extensive experiments demonstrate that VAV-R1 achieves SOTA performance across multiple tasks.
Qianhao Ren, Yutong Wang 0001, Zhongjiang He, Jingmin Xin, Hao Sun 0038
ICMR8
2026 Unifying Granularity and Reliability: A Robust and Efficient Framework for Text-based Person Retrieval
abstract
Text-based person retrieval (TPR) has become a crucial task in cross-modal retrieval due to its broad application in fields such as public safety and criminal investigation. Existing TPR methods typically rely on fully fine-tuning large-scale pretrained vision-language models like CLIP, which incurs high computational costs and tends to exhibit poor generalization in unseen domains due to overfitting. Fortunately, Parameter-Efficient Transfer Learning (PETL) has emerged as a lightweight alternative. However, applying PETL to TPR remains challenging, as its limited adaptation capacity struggles to capture intricate identity cues and becomes highly susceptible to gradient interference from unreliable image-text pairs. To address these challenges, we present a PETL-based framework named UniGR that unifies granularity and reliability for robust and efficient TPR. Specifically, we design a multi-granularity relational adapter (MRA) to capture both coarse-grained global and fine-grained local relational features among tokens, equipping the generic backbone with the task-specific, precise understanding needed for TPR. To combat the noise sensitivity of PETL, a reliability-aware reweighting strategy (RRS) is introduced to adaptively down-weight unreliable samples during training. Furthermore, we propose a parameter-free cross-modal cyclic verification (CMCV) module to mitigate ambiguities in cross-modal matching computations and refine retrieval ranking further. Experiments on benchmarks corroborate the superiority of UniGR among parameter-efficient methods. Remarkably, with only 4.5% of trainable parameters, UniGR outperforms most fully fine-tuned methods while maintaining strong generalization.
Jingchen Hao, Zhen Peng 0005, Yuting Zhang 0007, Zhongjiang He, Weizhan Zhang, Hao Sun 0038
SIGIR7
2025 Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification
abstract
Multi-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly rely on the Dempster-Shafer combination rule, which is sensitive to conflicting evidence and often neglects the critical role of neighborhood structures within multi-view data. To address these limitations, we propose a Trusted Unified Feature-NEighborhood Dynamics (TUNED) model for robust MVC. This method effectively integrates local and global feature-neighborhood (F-N) structures for robust decision-making. Specifically, we begin by extracting local F-N structures within each view. To further mitigate potential uncertainties and conflicts in multi-view fusion, we employ a selective Markov random field that adaptively manages cross-view neighborhood dependencies. Additionally, we employ a shared parameterized evidence extractor that learns global consensus conditioned on local F-N structures, thereby enhancing the global integration of multi-view features. Experiments on benchmark datasets show that our method improves accuracy and robustness over existing approaches, particularly in scenarios with high uncertainty and conflicting views.
Haojian Huang, Chuanyu Qin, Zhe Liu 0041, Kaijing Ma, Han Fang 0002, Chao Ban, Hao Sun 0038, Zhongjiang He
AAAI8
2025 FASTER: Face Attribute Sliders with Semantic Rewards
abstract
Large-scale text-to-image generative models have demonstrated remarkable success in generating diverse and high-quality faces. However, current methods for face editing often unintentionally modify facial features that are intended to be preserved. Multi-step denoising methods necessitate storing multi-step gradients, leading to considerable time and memory consumption. In this study, we propose FASTER(Face Attribute Sliders wiTh sEmantic Rewards), an effective method that employs stable diffusion models for face attribute editing. The key idea is to identify a low-rank attribute editing direction by leveraging attribute reward and S-CLIP reward between the original face and the edited face. This process helps to establish the desired face attribute slider. To acquire the edited face, we introduce an efficient one-step reward technique by utilizing denoised results at random timesteps for learning. This technique reduces training time by 6x. FASTER achieves 98.67% editing accuracy while simultaneously improving attribute preservation by nearly 10% compared to other methods on the CelebA-HQ dataset, all without compromising identity information.
Jingyan Chen, Lanxiang Zhou, Han Fang 0002, Zerun Feng, Chao Ban, Hao Sun 0038, Jiani Hu
ICASSP7
2025 ViCo: A Multitask Video-enhanced and Cognition-preserving Modality Alignment Training Framework
abstract
The rapid development of multimodal large language models (MLLMs) has brought significant breakthroughs to this field. However, current MLLMs typically rely on vision instruction tuning based on large language models (LLMs) to endow them with multimodal capabilities, which may lead to low video utilization and catastrophic forgetting due to the task-specific nature of vision instruction tuning. Take case of multiple-choice Video Question Answering (VideoQA) as an example, requiring only selecting correct answers makes it easier for LLMs to shortcut rather than fully comprehend video content. Moreover, the fixed task format of ’’predicting the correct option" risks catastrophic forgetting, which can cause LLMs to lose their origin cognitive abilities. To address this, we propose ViCo, a training framework that enables LLMs to perform a specific auxiliary tasks during vision instruction tuning, which helps the model leverage video information more thoroughly while mitigating catastrophic forgetting, thereby improving modality alignment and enhancing performance. We evaluate our method on several strong VideoQA benchmarks, where it outperforms all state-of-the-art methods, demonstrating its effectiveness and generality. Additionally, we also provide extensive analysis on the video information utilization and catastrophic forgetting resulting. The code will be made available at https://github.com/dunknsabsw/ViCo.
Zhenda Yu, Jiayu Shen, Lanxiang Zhou, Han Fang 0002, Xianghao Zang, Chao Ban, Jingfeng Chen, Zhongjiang He, Hao Sun 0038, Yanmei Kang
ICASSP10
2025 Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security
abstract
The rapid advancement of multimodal large language models (MLLMs) has led to breakthroughs in various applications, yet their security remains a critical challenge. One pressing issue involves unsafe image-query pairs-jailbreak inputs specifically designed to bypass security constraints and elicit unintended responses from MLLMs. Compared to general multimodal data, such unsafe inputs are relatively sparse, which limits the diversity and richness of training samples available for developing robust defense models. Meanwhile, existing guardrail-type methods rely on external modules to enforce security constraints but fail to address intrinsic vulnerabilities within MLLMs. Traditional supervised fine-tuning (SFT), on the other hand, often over-refuses harmless inputs, compromising general performance. Given these challenges, we propose Secure Tug-of-War (SecTOW), an innovative iterative defense-attack training method to enhance the security of MLLMs. SecTOW consists of two modules: a defender and an auxiliary attacker, both trained iteratively using reinforcement learning (GRPO). During the iterative process, the attacker identifies security vulnerabilities in the defense model and expands jailbreak data. The expanded data are then used to train the defender, enabling it to address identified security vulnerabilities. We also design reward mechanisms used for GRPO to simplify the use of response labels, reducing dependence on complex generative labels and enabling the efficient use of synthetic data. Additionally, a quality monitoring mechanism is used to mitigate the defender's over-refusal of harmless inputs and ensure the diversity of the jailbreak data generated by the attacker. Experimental results on safety-specific and general benchmarks demonstrate that SecTOW significantly improves security while preserving general performance. Warning: This paper contains offensive and unsafe content.
Muzhi Dai, Zhiyuan Zhao 0005, Junyu Gao 0001, Hao Sun 0038, Xuelong Li 0001
ACM Multimedia5
2025 SVGen: Interpretable Vector Graphics Generation with Large Language Models
abstract
Scalable Vector Graphics (SVG) has become an indispensable technology in front-end development and UI/UX design, due to its inherent advantages in scalability, editability, and rendering efficiency. In the creation of vector graphics, while expressing creative concepts is straightforward, translating them into precise digital artworks is often challenging and time-consuming. To overcome this technical bottleneck and achieve intelligent conversion from concept to final product, we have constructed SVG-1M, a large-scale dataset of high-quality SVG samples with paired textual descriptions. Through innovative data augmentation and annotation processes, we built precisely aligned ''Text instruction-SVG code'' training pairs, with a subset enhanced by Chain-of-Thought (CoT) annotations. This provides rich semantic supervision signals for model learning. Based on this dataset, we propose SVGen, an end-to-end generative model capable of directly converting natural language descriptions into SVG code. This design addresses the challenges of generating semantically accurate vector graphics while preserving complete structural information. We explored various training strategies and introduced a progressive curriculum learning approach, optimized with reinforcement learning algorithms. Notably, this study innovatively applies the CoT paradigm to vector graphics generation, effectively enhancing both the accuracy and interpretability of SVG synthesis. Experimental validation demonstrates that SVGen exhibits significant advantages over general large models in terms of SVG generation quality, while also surpassing optimization-based rendering methods in generation efficiency. The proposed method enables intelligent conversion between natural language and vector graphics, enabling novel workflows like real-time AI-assisted design iteration. Code, model, and data is released at: https://github.com/gitcat-404/SVGen
Zhiyuan Zhao 0005, Yuandong Liu, Da Zhang 0010, Junyu Gao 0001, Hao Sun 0038, Xuelong Li 0001
ACM Multimedia6
2024 ProTA: Probabilistic Token Aggregation for Text-Video Retrieval
abstract
Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (ProTA) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, ProTA achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%).
Han Fang 0002, Xianghao Zang, Chao Ban, Zerun Feng, Lanxiang Zhou, Zhongjiang He, Hao Sun 0038
ICME8
2024 GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment Tuning
abstract
Recent text-to-image (T2I) synthesis models have demonstrated intriguing abilities to produce high-quality images based on text prompts. However, current models still face Text-Image Misalignment problem (e.g., attribute errors and relation mistakes) for compositional generation. Existing models attempted to condition T2I models on grounding inputs to improve controllability while ignoring the explicit supervision from the layout conditions. To tackle this issue, we propose Grounded jOint lAyout aLignment (GOAL), an effective framework for T2I synthesis. Two novel modules, discriminative semantic alignment (DSAlign) and masked attention alignment (MAAlign), are proposed and incorporated in this framework to improve the text-image alignment. DSAlign leverages discriminative tasks at the region-wise level to ensure low-level semantic alignment. MAAlign provides high-level attention alignment by guiding the model to focus on the target object. We also build a dataset GOAL2K for model fine-tuning, which composes 2000 semantically accurate image-text pairs and their layout annotations. Comprehensive evaluations on T2I-Compbench, NSR-1K, and Drawbench demonstrate the superior generation performance of our method. Especially, there are improvements of 19%, 13%, and 12% in color, shape, and texture metrics for T2I-Compbench. Additionally, Q-Align metrics demonstrate that our method can generate images of higher quality.
Han Fang 0002, Zerun Feng, Kaijing Ma, Chao Ban, Xianghao Zang, Lanxiang Zhou, Zhongjiang He, Jingyan Chen, Jiani Hu, Hao Sun 0038
ACM Multimedia11
2023 Mask to Reconstruct: Cooperative Semantics Completion for Video-text Retrieval
abstract
Recently, masked video modeling has been widely explored and improved the model's understanding ability of visual regions at a local level. However, existing methods usually adopt random masking and follow the same reconstruction paradigm to complete the masked regions, which do not leverage the correlations between cross-modal content. In this paper, we present MAsk for Semantics COmpleTion (MASCOT) based on semantic-based masked modeling. Specifically, after applying attention-based video masking to generate high-informed and low-informed masks, we propose Informed Semantics Completion to recover masked semantics information. The recovery mechanism is achieved by aligning the masked content with the unmasked visual regions and corresponding textual context, which makes the model capture more text-related details at a patch level. Additionally, we shift the emphasis of reconstruction from irrelevant backgrounds to discriminative parts to ignore regions with low-informed masks. Furthermore, we design co-learning to incorporate video cues under different masks and learn more aligned representation. Our MASCOT performs state-of-the-art performance on four text-video retrieval benchmarks, including MSR-VTT, LSMDC, ActivityNet, and DiDeMo.
Han Fang 0002, Zhifei Yang 0004, Xianghao Zang, Chao Ban, Zhongjiang He, Hao Sun 0038, Lanxiang Zhou
ACM Multimedia6
2023 A Baseline Investigation: Transformer-based Cross-view Baseline for Text-based Person Search
abstract
This paper investigates a baseline approach for text-based person search by using a transformer-based framework. Existing methods usually treat the visual and textual features as independent entities for speeding up the model inference process. However, the attention to the same images should be changed according to different texts. In this paper, we use a commonly employed framework with a fused feature as the baseline, which overcomes the misalignment problem introduced by fixed features. A thorough investigation is conducted in this paper. Moreover, we propose Cross-View Matching (CVM) to provide challenging, positive text-image pairs that enable the model to learn cross-view meta-information. Furthermore, we suggest a novel evaluation process to reduce the inference time and GPU memory demand. The experiments are conducted on CUHK-PEDES, ICFG-PEDES, and RSTPReid benchmarks. Through extensive parameter analysis, the potentials of a transformer-based framework are fully explored. Although the proposed scheme is a simple framework, it achieves significant performance improvements compared with other state-of-the-art methods.
Xianghao Zang, Wei Gao 0003, Ge Li 0002, Han Fang 0002, Chao Ban, Zhongjiang He, Hao Sun 0038
ACM Multimedia7