Qun Hao

dblp:03/6160 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hierarchical Fourier encoding for spatially-adaptive continuous super-resolution
Yang Cheng 0006, Haoyue Xing, Chaohui Li, Cancan Yao, Qun Hao
Pattern Recognit.5
2026 ObjSplat: Geometry-Aware Gaussian Surfels for Active Object Reconstruction
abstract
Autonomous high-fidelity object reconstruction is fundamental for creating digital assets and bridging the simulation-to-reality gap in robotics. We present ObjSplat, an active reconstruction framework that leverages Gaussian surfels as a unified representation to progressively reconstruct unknown objects with both photorealistic appearance and accurate geometry. Addressing the limitations of conventional opacity or depth-based cues, we introduce a geometry-aware viewpoint evaluation pipeline that explicitly models back-face visibility and occlusion-aware multi-view covisibility, reliably identifying under-reconstructed regions even on geometrically complex objects. Furthermore, to overcome the limitations of greedy planning strategies, ObjSplat employs a next-best-path (NBP) planner that performs multi-step lookahead on a dynamically constructed spatial graph. By jointly optimizing information gain and movement cost, this planner generates globally efficient trajectories. Extensive experiments in simulation and on real-world cultural artifacts demonstrate that ObjSplat produces physically consistent models within minutes, achieving superior reconstruction fidelity and surface completeness while significantly reducing scan time and path length compared to state-of-the-art approaches.
Yuetao Li, Zhizhou Jia, Qun Hao
IEEE Trans Autom. Sci. Eng.4
2025 Towards Accurate Tumor Budding Detection: A Benchmark Dataset and A Detection Approach Based on Implicit Annotation Standardization and Positive-Negative Feature Coupling
Ruiqing Sun, Zeng Fan, Boyang Dai, Yiyan Su, Qun Hao, Chuyang Ye
MICCAI (3)5
2025 Speed-Oriented Lightweight Salient Object Detection in Optical Remote Sensing Images
abstract
The lightweight model for salient object detection in optical remote sensing images (SOD-RSI) is a recent emerging topic. Due to the complexity of the task, recently published works have achieved effective model compression but have not yet achieved the desired detection speed. To truly release the detection speed of lightweight models while ensuring a favorable accuracy-efficiency tradeoff, we propose a new speed-oriented lightweight SOD-RSI network (SOLNet), which has significant advantages in detection speed. Specifically, we design a lightweight group attention (LGA) module to deconstruct–interact–recombine channel features and an enhanced dynamic encoding (EDE) module for dynamically capturing spatial information. On this basis, the dynamically enhanced aggregation module (DEAM) is further proposed, which mines the intrinsic correlation of feature information by decoding high-level feature maps, eliminating the need to pay additional attention to other scales. SOLNet completes lightweight and efficient decoding through simple cascade aggregation operations. Notably, we also propose an evaluation strategy that takes both speed and accuracy into account, extending a novel lightweight gain (Lg) metric for SOD-RSI. This not only effectively reveals the under-gain issue of lightweight models but also provides theoretical support for the evaluation of subsequent lightweight works. Experimental results on the challenging EORSSD and ORSSD datasets show that SOLNet achieves significant speed improvements and is the state-of-the-art (SOTA) lightweight SOD-RSI method. The code is available athttps://github.com/SpiritAshes/SOLNet.
Zhaoyang Li 0009, Yinxiao Miao, Xiongwei Li, Jie Cao 0004, Qun Hao, Dongxing Li, Yunlong Sheng
IEEE Trans. Geosci. Remote. Sens.6
2025 Progressive Gradient-Guided Self-Distillation Keypoint Detection
abstract
With the increasing popularity of autonomous driving and 3D reconstruction, keypoint detection, as a key link in visual localization, has become a hot topic in current research. However, existing keypoint detection methods rarely pay attention to the difficulty differences of samples and lack a progressive learning mechanism, which often leads to overfitting for simple samples and underfitting for complex samples, limiting the overall performance of the model. To address these issues, we propose a novel progressive gradient-guided self-distillation method (PG${^{2}}$SD) for keypoint detection, which possesses self-evolutionary learning capabilities. Specifically, we propose a progressive gradient constraint strategy (PGCS) that dynamically adjusts the gradient contributions of different samples, enabling the model to adapt to the evolving learning capability during training. On this basis, we propose a gradient-guided self-distillation strategy (G${^{2}}$SDS), which integrates seamlessly with PGCS to alleviate the insufficient feature representation of hard samples in the early training stage. We further design a novel loss function to achieve dynamic collaboration between PGCS and G${^{2}}$SDS, allowing G${^{2}}$SDS to adaptively adjust the self-distillation parameters through the PGCS. Experimental results on multiple benchmark datasets show that our method achieves state-of-the-art performance on image matching, visual localization, and 3D reconstruction tasks without designing a proprietary network, indicating broad application prospects.
Zhaoyang Li 0009, Jie Cao 0004, Qun Hao, Haifeng Yao
IEEE Trans. Multim.3
2024 Semi-supervised Medical Image Segmentation based on Coarse-Fine Dual Training Streams
abstract
Commonly used semi-supervised medical segmentation networks usually use consistent learning under different data perturbations to regularise training, ignoring the multiscale information of the data itself. Therefore, this paper proposes a new network based on coarse and fine dual training streams(CF-UNet), which consists of a backbone network and auxiliary learning branches(ALB). Our approach has the following two novel designs: 1) We design a simple and effective coarse- fine dual training streams. Specifically, in the coarse training stream, we improve the robustness and generalisation of the model by establishing regularisation between different strong and weak perturbation views. In the fine training stream, we introduce an auxiliary learning branch to improve the prediction performance of the backbone network.2) In the ALB module, we design the channel spatial fusion attention module (CSMA) and multiscale large kernel convolutional attention (MS-LKA) to perform feature extraction and fusion from a variety of scales. We evaluate our proposed method on ACDC and DRIVE datasets and numerous experiments have shown that our CF-UNet outperforms state-of-the-art networks. Code is available at https://github.com/slz-bit/CF-UNet.
Lizhi Sun, Zhengyu Qiao, Qun Hao
BIBM6
2024 Design of novel deformable mirror with large displacement
abstract
Summary A stabilized zoom system with deformable mirrors (DMs) was designed for continuous optical zoom. In order to adjust the focal length and expand the field angle, the characteristic surface shapes of the DMs composed of the low‐order and high‐order Zernike polynomials were provided. The maximum peak‐to‐valley (PV) value at the center of the surface shape is 80 μm. This work presents a piezoelectrically‐actuated deformable mirror (PADM) with large displacement for a stabilized zoom system. The COMSOL simulation was conducted to obtain the desired design by matching and optimizing structural parameters. The maximum displacement of PADM was more than 80 μm. The fitting results by the steepest descent algorithm (SD) showed that the RMS values of the residual surface shapes were 1 of 7 to 1 of 4 of their PV values, and the PV values of the residual surface shapes were less than 1 of 10 of the PV values of the target surface shapes.
Jinchao Li, Jinling Yang, Yinfang Zhu, Xuemin Cheng, Qun Hao
Concurr. Comput. Pract. Exp.7
2024 Multi-scale convolutional neural networks and saliency weight maps for infrared and visible image fusion
Chenxuan Yang, Yunan He, Bingkun Chen, Jie Cao 0004, Yongtian Wang, Qun Hao
J. Vis. Commun. Image Represent.7
2023 Lightweight feature point detection network with channel enhancement
Zhaoyang Li 0009, Chun Bao, Jie Cao 0004, Dongxing Li, Qun Hao
Comput. Vis. Image Underst.6
2023 Dynamic threshold integrate and fire neuron model for low latency spiking neural networks
Xiyan Wu, Yong Song 0002, Yurong Jiang, Yashuo Bai, Xin Yang 0021, Qun Hao
Neurocomputing9
2023 Improving the Generalization of Visual Classification Models Across IoT Cameras via Cross-Modal Inference and Fusion
abstract
The performance of visual classification models across Internet of Things devices is usually limited by the changes in local environments, resulted from the diverse appearances of the target objects and differences in light conditions and background scenes. To alleviate these problems, existing studies usually introduce the multimodal information to guide the learning process of the visual classification models, making the models extract the visual features from the discriminative image regions. Especially, cross-modal alignment between visual and textual features has been considered as an effective way for this task by learning a domain-consistent latent feature space for the visual and semantic features. However, this approach may suffer from the heterogeneity between multiple modalities, such as the multimodal features and the differences in the learned feature values. To alleviate this problem, this article first presents a comparative analysis of the functionality of various alignment strategies and their impacts on improving visual classification. Subsequently, a cross-modal inference and fusion framework (termed as CRIF) is proposed to align the heterogeneous features in both the feature distributions and values. More importantly, CRIF includes a cross-modal information enrichment module to improve the final classification and learn the mappings from the visual to the semantic space. We conduct experiments on four benchmarking data sets, i.e., the Vireo-Food172, NUS-WIDE, MSR-VTT, and ActivityNet Captions data sets. We report state-of-the-art results for basic classification tasks on the four data sets and conduct subsequent experiments on feature alignment and fusion. The experimental results verify that CRIF can effectively improve the learning ability of the visual classification models, and it is a model-agnostic framework that consistently improves the performance of state-of-the-art visual classification models.
Qing-Ling Guan, Yuze Zheng, Lei Meng 0001, Liquan Dong, Qun Hao
IEEE Internet Things J.5
2023 Rega-Net: Retina Gabor Attention for Deep Convolutional Neural Networks
abstract
Extensive research works demonstrate that the attention mechanism in convolutional neural networks (CNNs) effectively improves accuracy. Nevertheless, few works design attention mechanisms using large receptive fields. In this work, we propose a novel attention method named Rega-Net to increase CNN accuracy by enlarging the receptive field. To the best of our knowledge, increasing the receptive field of the convolutional neural network requires increasing the size of the convolution kernel, which also increases the number of parameters. For solving this problem, we design convolutional kernels to resemble the non-uniformly distributed structure inspired by the mechanism of the human retina. Then, we sample variable-resolution values in the Gabor function distribution and fill these values in retina-like kernels. This distribution allows essential features to be more visible in the center position of the receptive field. We further design an attention module including these retina-like kernels. Experiments demonstrate that our Rega-Net achieves 79.96% Top-1 accuracy on ImageNet-1K for classification and 43.1% mAP on COCO2017 for object detection. The mAP of the Rega-Net increased by up to 3.5% compared to baseline networks.
Chun Bao, Jie Cao 0004, Yaqian Ning, Yang Cheng 0006, Qun Hao
IEEE Geosci. Remote. Sens. Lett.5
2021 Single Haze Image Restoration Under Non-Uniform Dense Scattering Media
abstract
We demonstrate an efficient restoration method based on sky segmentation and block optimized transmission (SSBOT) to restore the visibility of single haze image under non-uniform scattering medium. The SSBOT method uses sky region segmentation to determine the haze-opaque region, and estimates the global atmospheric light using quadtree-decomposition. The proposed SSBOT estimates the transmission in every sub-images of the degraded image with dark channel method with coefficient correction. The experimental results show that SSBOT can efficiently restore images degraded by non-uniform dense scattering medium. The proposed method can be used for various security and surveillance applications.
Jie Cao 0004, Saad Rizvi, Qun Hao
IEEE Signal Process. Lett.4
2020 IR saliency detection via a GCF-SB visual attention framework
Yong Song 0002, Muhammad Sulaman, Zhengkun Guo, Xin Yang 0021, Fengning Wang, Qun Hao
J. Vis. Commun. Image Represent.8
2020 Reduced data set for multi-target recognition using compressed sensing frame
abstract
Oceanographic plankton classification for many images is time consuming, especially for low-contrast images obtained when the water is dirty and opaque due to impurities. A novel, overcomplete dictionary algorithm is studied by analyzing the sparse characteristics of the image matrix using pixel values. The features in hyperspace are mapped onto a specifically designed vector space. Thus, mathematically, the algorithm exhibits faster calculating convergence and has a strong expressivity for the selective signal in the vector space with a lower signal loss rate. The clustering method based on the dictionary can classify planktons for species counting, which can enable high-speed, multi-object recognition of planktons in turbid water. The experimental results demonstrate that if less data (up to 60%) is processed for each image, a recall rate and accuracy greater than 75% and a structural similarity for the reconstructed image greater than 0.9 can be achieved.
Xuemin Cheng, Changqing Dong, Kaichang Cheng, Yao Hu 0004, Qun Hao
Pattern Recognit. Lett.7
2019 Extended-depth-of-field object detection with wavefront coding imaging system
Liquan Dong, Haoyuan Du, Ming Liu 0029, Yuejin Zhao, Shijia Feng, Mei Hui, Lingqin Kong, Qun Hao
Pattern Recognit. Lett.10
2019 Study on the performance of three-dimensional ghost image affected by target
Fanghua Zhang, Jie Cao 0004, Yang Cheng 0006, Qun Hao, Zonglei Mou
Pattern Recognit. Lett.5
2013 Aptitude digging education in project-based course
abstract
Students from China are always intelligent but lack of creativity. These are somewhat stereotypes. This is partially because of the reserved or implicit culture. In an objective point of view, it is also because of the limited education resources. In the single assessment criterion education circumstance, students chase for the high marks even without knowing their interests or aptitudes. In a 12-week open experimental course, Optoelectronic Instrument Experiments (OIE), we try to encourage the students to dig their aptitudes and bring them into full play to earn more credits for the course. Self-assessment and mutual-evaluation for technical proficiency, communication skills, collaboration and leadership are carried out for the final evaluation. We also communicate with the students the speciality and skill a qualified engineer needs. We hope to help them prepare themselves for engineering-related jobs in the further.
Yao Hu 0004, Liquan Dong, Ming Liu 0029, Yuejin Zhao, Qun Hao
FIE6
2013 Let's do it OR deal with it: Teamwork in project-based learning
abstract
Project-based course is based on teamwork and most of work is done and presented as a team. Team grouping rule is one of the most important issues. In project-based experimental course Optoelectronic Instrument Experiments (OIE), several different rules were attempted, each of which produced complaints by some students. After several trials of different grouping rule, we realized that there is no perfect rule which can satisfy everyone. Instead of changing the rule, trying to find a way to persuade the students to accept and support their group willingly might be a better solution. In this semester, the project teams are entirely determined by lot and several teamwork inspirational approaches are introduced to inspire team spirit in the course. Our purpose is to find a way to make student learn the interpersonal skill of working in team. Let's do it, not just deal with it inactively.
Yao Hu 0004, Liquan Dong, Ming Liu 0029, Yuejin Zhao, Qun Hao
FIE6
2007 Broadcasting Protocols for Multi-Radio Multi-Channel and Multi-Rate Mesh Networks
abstract
A vast amount of broadcasting protocols has been developed for wireless ad hoc networks. To the best of our knowledge, however, these protocols assume a single-radio single-channel and single-rate network model and/or a generalized physical model, which does not take into account the impact of interference. In this paper, we present a set of broadcasting protocols to simultaneously achieve 100% reliability, minimum broadcasting latency, and minimum redundant transmissions. Our research distinguishes itself in a number of ways. First, a multi-radio multi-channel and multi-rate mesh network model is used. Second, the broadcasting tree is constructed by using local information without the global network topological information. Third, a comprehensive link quality metric is defined to fully take into account the interference. The link quality information is also made available to broadcasting protocols. Fourth, three performance metrics that include reliability, latency, and redundancy are simultaneously considered. Simulations are conducted to evaluate the proposed protocols and compare the performance improvement to other protocols.
Min Song 0002, Jun Wang 0016, Qun Hao
ICC3