Haosen Sun

dblp:326/5306 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Flow-Based Knowledge Transfer for Efficient Large Model Distillation
abstract
Traditional knowledge distillation relies on simple MSE or KL divergence losses that fail to capture the complex distributional relationships between teacher and student model representations. We propose FlowDistill, a novel distillation framework that employs normalizing flows to model and transfer the intricate knowledge distributions from teacher to student models. Our approach introduces three key innovations: (1) Invertible Knowledge Mapping using continuous normalizing flows (CNFs) to learn bijective transformations between teacher and student representation spaces, enabling precise knowledge transfer without information loss, (2) Flow-Guided Progressive Distillation that gradually increases the complexity of knowledge transfer by learning hierarchical flow transformations from simple to complex distributions, and (3) Conditional Flow Networks that adapt knowledge transfer based on input context and task requirements. Unlike previous diffusion-based distillation methods such as DiffKD that suffer from computational overhead due to iterative denoising processes and information loss during noise addition, our flow-based approach provides exact invertible transformations with significantly reduced computational cost. Extensive experiments on ImageNet classification, COCO object detection, and Cityscapes semantic segmentation demonstrate that FlowDistill achieves superior performance with 2.1% accuracy improvement over DiffKD on ResNet-34 to ResNet-18 distillation while reducing inference time by 3.5×. Our method establishes new state-of-the-art results across multiple distillation benchmarks and provides theoretical guarantees for lossless knowledge transfer through invertible flow transformations.
Haosen Sun, Xuesheng Zhang, Zebang Liu, Gaochao Xu
AAAI4
2026 ProgressLM: Towards Progress Reasoning in Vision-Language Models
abstract
Jianshu Zhang, Chengxuan Qian, Haosen Sun, Haoran Lu, Dingcheng Wang, Letian Xue, Han Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chengxuan Qian, Haosen Sun, Dingcheng Wang, Letian Xue, Han Liu 0001
ACL (1)3
2025 Re-thinking Temporal Search for Long-Form Video Understanding
abstract
Efficiently understanding long-form videos remains a significant challenge in computer vision. In this work, we revisit temporal search paradigms for long-form video understanding and address a fundamental issue pertaining to all state-of-the-art (SOTA) long-context vision-language models (VLMs). Our contributions are twofold: First, we frame temporal search as a Long Video Haystack problem – finding a minimal set of relevant frames (e.g., one to five) from tens of thousands based on specific queries. Upon this formulation, we introduce LV-Haystack, the first dataset with 480 hours of videos, 15,092 human-annotated instances for both training and evaluation aiming to improve temporal search quality and efficiency. Results on LV-HAYSTACK highlight a significant research gap in temporal search capabilities, with current SOTA search methods only achieving 2.1% temporal F1score on the LongVideoBench subset.Next, inspired by visual search in images, we propose a lightweight temporal search framework, T* that reframes costly temporal search as spatial search. T* leverages powerful visual localization techniques commonly used in images and introduces an adaptive zooming-in mechanism that operates across both temporal and spatial dimensions. Extensive experiments show that integrating T* with existing methods significantly improves SOTA long-form video understanding. Under an inference budget of 32 frames, T* improves GPT-4o’s performance from 50.5% to 53.1% and LLaVAOneVision-OV-72B’s performance from 56.5% to 62.4% on the LongVideoBench XL subset. Our code, benchmark, and models are provided in the Supplementary material.
Jinhui Ye, Zihan Wang 0008, Haosen Sun, Keshigeyan Chandrasegaran, Zane Durante, Cristóbal Eyzaguirre, Yonatan Bisk, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001, Jiajun Wu 0001, Manling Li
CVPR3
2024 Auto-GAS: Automated Proxy Discovery for Training-Free Generative Architecture Search
Lujun Li 0001, Haosen Sun, Shiwen Li, Peijie Dong, Wenhan Luo, Wei Xue 0002, Yike Guo
ECCV (5)2
2024 Auto-DAS: Automated Proxy Discovery for Training-Free Distillation-Aware Architecture Search
Haosen Sun, Lujun Li 0001, Peijie Dong, Zimian Wei, Shitong Shao
ECCV (5)1
2024 Learning from Distinction: Mitigating Backdoors Using a Low-Capacity Model
abstract
Deep neural networks (DNNs) are susceptible to backdoor attacks due to their black-box nature and lack of interpretability. Backdoor attacks intend to manipulate the model's prediction when hidden backdoors are activated by predefined triggers. Although considerable progress has been made in backdoor detection and removal at the model deployment stage, an effective defense against backdoor attacks during the training time is still under-explored. In this paper, we propose a novel training-time backdoor defense method called Learning from Distinction (LfD), allowing training a backdoor-free model on the backdoor-poisoned data. LfD uses a low-capacity model as a teacher to guide the learning of a backdoor-free student model via a dynamic weighting strategy. Extensive experiments on CIFAR-10, GTSRB and ImageNet-subset datasets show that LfD significantly reduces attack success rates to 0.67%, 6.14% and 1.42%, respectively, with minimal impact on clean accuracy (less than 1%, 3% and 1%).
Haosen Sun, Xixiang Lyu
ACM Multimedia1
2024 A Faster Fire Detection Network with Global Information Awareness
Jinrong Cui, Haosen Sun, Ciwei Kuang, Yong Xu 0001
PRCV (12)2
2022 Measuring road safety achievement based on EWM-GRA-SVD: A decision-making support system for APEC countries
Faan Chen, Haosen Sun, Jiacheng Zu, Zhenwei Sun
Knowl. Based Syst.5