Qihan He

dblp:273/9410 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2025
0009-0002-2627-9049ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CCANet: A Cross-scale Context Aggregation Network for UAV object detection
Qihan He, Huan Lei
Comput. Vis. Image Underst.2
2025 VARFVV: View-Adaptive Real-Time Interactive Free-View Video Streaming With Edge Computing
abstract
Free-view video (FVV) allows users to explore immersive video content from multiple views. However, delivering FVV poses significant challenges due to the uncertainty in view switching, combined with the substantial bandwidth and computational resources required to transmit and decode multiple video streams, which may result in frequent playback interruptions. Existing approaches, either client-based or cloud-based, struggle to meet high Quality of Experience (QoE) requirements under limited bandwidth and computational resources. To address these issues, we propose VARFVV, a bandwidth- and computationally-efficient system that enables real-time interactive FVV streaming with high QoE and low switching delay. Specifically, VARFVV introduces a low-complexity FVV generation scheme that reassembles multiview video frames at the edge server based on user-selected view tracks, eliminating the need for transcoding and significantly reducing computational overhead. This design makes it well-suited for large-scale, mobile-based UHD FVV experiences. Furthermore, we present a popularity-adaptive bit allocation method, leveraging a graph neural network, that predicts view popularity and dynamically adjusts bit allocation to maximize QoE within bandwidth constraints. We also construct an FVV dataset comprising 330 videos from 10 scenes, including basketball, opera, etc. Extensive experiments show that VARFVV surpasses existing methods in video quality, switching latency, computational efficiency, and bandwidth usage, supporting over 500 users on a single edge server with a switching delay of 71.5ms. Our code and dataset are available at https://github.com/qianghu-huber/VARFVV.
Qiang Hu 0003, Qihan He, Houqiang Zhong, Guo Lu, Xiaoyun Zhang 0001, Guangtao Zhai, Yanfeng Wang 0001
IEEE J. Sel. Areas Commun.2
2025 Star-PMFI: Star-attention and pyramid multi-scale feature integration network for small object detection in drone imagery
Zhongxu Li, Qihan He
J. Vis. Commun. Image Represent.3
2025 PCAF: UAV scenarios detector via pyramid converge-and-assign fusion network
Zhongxu Li, Qihan He, Lingfei Ren, Wenyong Yao
Multim. Syst.2
2025 Slfmamba:a state space based vision foundation models fine-tuning for domain generalized semantic segmentations
Yongchao Qiao, Ya'nan Guan, Qihan He, Zhongxu Li, Jingmin Yang
Multim. Syst.3
2025 A lightweight multidimensional feature network for small object detection on UAVs
Qihan He, Zhongxu Li
Pattern Anal. Appl.2
2025 E-FPN: an enhanced feature pyramid network for UAV scenarios detection
Zhongxu Li, Qihan He
Vis. Comput.2
2024 LMFE-RDD: a road damage detector with a lightweight multi-feature extraction network
Qihan He, Zhongxu Li
Multim. Syst.1
2024 Lsf-rdd: a local sensing feature network for road damage detection
Qihan He, Zhongxu Li
Pattern Anal. Appl.1
2023 Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos
abstract
The success of the Neural Radiance Fields (NeRFs) for modeling and free-view rendering static objects has in-spired numerous attempts on dynamic scenes. Current techniques that utilize neural rendering for facilitating free-view videos (FVVs) are restricted to either offline rendering or are capable of processing only brief sequences with minimal motion. In this paper, we present a novel technique, Residual Radiance Field or ReRF, as a highly com-pact neural representation to achieve real-time FVV ren-dering on long-duration dynamic scenes. ReRF explicitly models the residual information between adjacent times-tamps in the spatial-temporal feature space, with a global coordinate-based tiny MLP as the feature decoder. Specif-ically, ReRF employs a compact motion grid along with a residual feature grid to exploit inter-frame feature similar-ities. We show such a strategy can handle large motions without sacrificing quality. We further present a sequential training scheme to maintain the smoothness and the spar-sity of the motion/residual grids. Based on ReRF, we design a special FVV codec that achieves three orders of magni-tudes compression rate and provides a companion ReRF player to support online streaming of long-duration FVVs of dynamic scenes. Extensive experiments demonstrate the effectiveness of ReRF for compactly representing dynamic radiance fields, enabling an unprecedented free-viewpoint viewing experience in speed and quality.
Qiang Hu 0003, Qihan He, Jingyi Yu 0001, Tinne Tuytelaars, Lan Xu 0003, Minye Wu
CVPR3