Hongru Zhao

dblp:247/1788 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Less Is Better: Sparse Instance Learning for Cross-Domain Few-Shot Object Detection
abstract
Cross-Domain Few-Shot Object Detection (CD-FSOD) is an extremely challenging task due to the inherent data scarcity and substantial domain shift between the source and target domains. Existing methods often suffer from overfitting and noisy feature representations, which hinder the construction of discriminative class prototypes in the target domain. In this paper, we propose a novel framework with sparse instance learning (SI-ViTO) for CD-FSOD, which leverages instance sparsity to achieve a better detection with less representation. SI-ViTO adopts a dual-stage sparsity module, consisting of instance feature sparsity not only on the few-shot support images but also on the query images. This dual sparsity enables the model to effectively preserve salient foreground semantics and simultaneously to filter out redundant or noisy information. Furthermore, a new prototype calibration strategy is also used to dynamically refine the class prototypes with query instances to accelerate prototype adaptation. Extensive experimental results on CD-FSOD benchmarks show that SI-ViTO outperforms the state-of-the-art methods, demonstrating that less discriminative representations yield better cross-domain few-shot object detection performance than more abundant ones.
Yali Huang, Hongru Zhao, Mingyuan Jiu, Hichem Sahbi
AAAI5
2026 TrajKD: Distilling Knowledge via Adaptive Trajectory Curriculum and Dynamic Weighting
Mingyuan Jiu, Mi Guo, Hongru Zhao, Mingliang Xu 0001
ICPR (10)6
2026 Calibrate and Aggregate: Cross-Modal Retrieval with Distribution Alignment and Token Reduction
abstract
Cross-modal image-text retrieval remains a fundamental challenge at the intersection of vision and language. While existing methods leveraging large-scale pre-trained models such as CLIP and BERT have achieved notable progress, they often struggle with the inherent modality gap, background clutter in images, and computational inefficiency. In this paper, we propose an end-to-end framework for efficient cross-modal retrieval using Distribution Alignment and Token Reduction (DATR). The model employs a feature extraction network to extract multi-level representations, which are preliminarily aligned via a distribution calibration module. This module maps features into a Gaussian latent space and maximizes inter-modal mutual information through InfoNCE loss, effectively bridging the semantic gap. Furthermore, we incorporate a differentiable token reduction module inspired by GroupViT to dynamically cluster redundant visual tokens into semantic groups, significantly reducing computational overhead. Similarity computation integrates global and local alignments for robust matching. Extensive experiments on Flickr30K and MSCOCO datasets demonstrate state-of-the-art performance, with image-to-text retrieval R@1 reaching 89.6% and text-to-image retrieval R@1 achieving 77.2% on Flickr30K. The source codes are available at https://github.com/Wuziyi123/DATR.
Mingyuan Jiu, Hongru Zhao, Hichem Sahbi, Mingliang Xu 0001
ICMR4
2026 Deep Convolutional Primal-Dual Network for Image Deblurring
abstract
Image deblurring is a challenging image task, which is regarded as a classical inverse problem. Deep primal-dual proximal network (DeepPDNet) is recently proposed which unrolls the Condat-Vũ primal-dual splitting algorithm as a feed-forward network and it has demonstrated excellent restoration performance. However, the feature patterns in the DeepPDNet are well manually designed and thus the network is not implemented in an efficient convolutional fashion. In this work, we revisit the DeepPDNet and extend it in three respects: i) the convolution and pooling operators as well as their associating adjoint operations are studied in the primal-dual algorithm, and then a deep convolutional primal-dual network (DeepConvPDNet) and its full variant with skips are proposed to preserve the optimization consistence of primal-dual Condat-Vũ algorithm; ii) two (cascade vs parallel) variants of the networks are designed according to the structure of convolutional kernels; iii) rather than that the blur kernels are given as prior knowledge, they can be encoded by a set of convolutional layers and deconvolutional layers for their conjugate, resulting to a full learnable deep convolutional primal-dual neural network.We investigate the proposed networks on the MNIST dataset, the grayscale and color version of BSD dataset and GoPro dataset for image deblurring. Extensive experiments are conducted to validate the performance of the proposed networks, and promising results in term of PSNR and SSIM are obtained in comparison with twelve methods including state-of-the-art methods (e.g. Restormer, DRUNet, and DeblurGAN), which validated its effectiveness.
Mingyuan Jiu, Mingjing Peng, Fanfan Zhang, Shupan Li, Hongru Zhao, Rongrong Ji, Mingliang Xu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 BiDAFuse: Bimodal Differences-Aware Attentive Network for Infrared and Visible Image Fusion
Keyu Sun, Mingyuan Jiu, Shupan Li, Hongru Zhao
ICIG (3)4
2025 Towards Communication-Efficient Heterogeneous Collaborative Perception via Semantic Disentanglement
abstract
Heterogeneous collaborative perception enables interconnected agents to share and fuse intermediate features extracted from multimodal sensor data, thereby enhancing robustness and reducing perception blind spots. However, shared high-dimensional features often contain redundant and modality-specific information, leading to semantic inconsistencies and excessive communication overhead. To address these challenges, we propose an efficient communication framework for heterogeneous collaborative perception. This framework integrates a shared-private feature decoupling module that disentangles cross-modal features into shared and private components, sharing only shared feature to improve communication efficiency. Additionally, we design a spatial redundancy elimination module to remove background information from transmitted features, further reducing bandwidth consumption. Collectively, these modules enable compact and communication-efficient perception capabilities across heterogeneous modalities, models, and tasks. Experimental results demonstrate that the proposed framework significantly reduces communication overhead in complex heterogeneous scenarios while maintaining perception accuracy.
Shijie Feng, Tiange Fu, Hongru Zhao, Guiyang Luo, Quan Yuan 0004
ICPADS3
2025 SE-D3FNet: A LiDAR-Camera Fusion Network with SE Attention and Dynamic 3D Focal Loss for 3D Object Detection
abstract
3D object detection is a critical problem in the field of computer vision, widely applied in autonomous driving, robotic navigation, and other domains. Although modern detectors have achieved success in singlesensor object detection, they remain vulnerable to complex environments due to the limitations of single-sensor modalities. We propose SE-D3FNet, a multi-modal fusion framework for 3D object detection that integrates Squeeze-and-Excitation (SE) channel attention and a dynamic 3D focal loss to significantly improve detection accuracy. We present an enhanced feature extraction network termed SE-ResBlock, which demonstrates superior capability in capturing global contextual information. The loss function is also optimized to more accurately capture targets with poor recognition rates. Experimental results on the KITTI benchmark demonstrate that our proposed 3D object detection algorithm achieves superior performance for the car category compared to existing methods.
Mingyuan Jiu, Shupan Li, Hongru Zhao, Mingliang Xu 0001
ICPADS5
2025 Two grids are better than one: Hybrid indoor scene reconstruction framework with adaptive priors
abstract
Indoor scene reconstruction from multi-view images is a pivotal technology within the field of robotics and augmented reality . Previous researches have predominantly focused on neural radiance fields aided by geometric monocular priors. However, due to the inductive smoothness bias introduced by deep Multi-Layer Perceptron (MLP) networks, these methods struggle to recover the scene surface with complex and fine geometry details. Additionally, when used as additional supervision signals during optimization, priors in different regions make different contributions. Simply incorporating them in all regions may lead to a decrease in the accuracy. To tackle these issues, we present a generic end-to-end framework named AdaptSurf, which combines Signed Distance Field (SDF) voxel grids and feature voxel grids to enhance the capability of reconstructing accurate geometry details, respectively. Furthermore, we design a policy network to adaptively enable the estimated depth or normal priors to supervise the learning process, which improves the reconstruction accuracy and accelerates neural surface reconstruction. Qualitative and quantitative experiments show that AdaptSurf yields high-quality surfaces, especially for fine-grained details and smooth regions. Furthermore, the policy network exhibits an interpretable behavior that depends on the voxel features, which helps to improve the quality of surface reconstruction.
Boyuan Bai, Xiuquan Qiao, Hongru Zhao, Wenzhe Shi, Hengjia Zhang, Yakun Huang
Neurocomputing4
2025 Globally-Optimal Greedy Active Sequential Estimation
abstract
Motivated by modern applications such as computerized adaptive testing, sequential rank aggregation, and heterogeneous data source selection, we study the problem of active sequential estimation. The goal is to design an adaptive experiment selection rule and an estimator for more accurate parameter estimation. Greedy information-based experiment selection rules, which optimize information gain one step ahead, have been employed in practice thanks to their computational convenience, flexibility to context or task changes, and broad applicability. However, the optimality of greedy methods under a sequential decision theory framework is only established in the one-dimensional case, partly due to the problem’s combinatorial nature and the seemingly limited capacity of greedy algorithms. In this study, we close the gap for multidimensional problems. We cast the problem under a sequential decision theory framework with generalized risk measures for a large class of design-and-estimation methods. We propose adopting the maximum likelihood estimator with a class of greedy experiment selection rules. This class encompasses both existing methods and introduces new methods with improved numerical efficiency. We prove that these methods achieve asymptotic optimality when the risk measure aligns with the selection rule. Additionally, we establish that the proposed estimators are consistent and asymptotically normal, and further extend the results to allow early stopping rules. We also perform extensive numerical studies on both simulated and real data to illustrate the efficacy of the proposed methods.
Xiaoou Li 0002, Hongru Zhao
IEEE Trans. Inf. Theory2
2023 Graph attention network-optimized dynamic monocular visual odometry
Hongru Zhao, Xiuquan Qiao
Appl. Intell.1
2021 Towards Video Streaming Analysis and Sharing for Multi-Device Interaction with Lightweight DNNs
abstract
Multi-device interaction has attracted a growing interest in both mobile communication industry and mobile computing research community as mobile devices enabled social media and social networking continue to blossom. However, due to the stringent low latency requirements and the complexity and intensity of computation, implementing efficient multi-device interaction for real-time video streaming analysis and sharing is still in its infancy. Unlike previous approaches that rely on high network bandwidth and high availability of cloud center with GPUs to support intensive computations for multi-device interaction and for improving the service experience, we propose MIRSA, a novel edge centric multi-device interaction framework with a lightweight end-to-end DNN for on-device visual odometry (VO) streaming analysis by leveraging edge computing optimizations with three main contributions. First, we design MIRSA to migrate computations from the cloud to the device side, reducing the high overhead for large transmission of video streaming while alleviating the server load of the cloud. Second, we design a lightweight VO network by utilizing temporal shift module to support on-device pose estimation. Third, we provide on-device resource-aware scheduling algorithm to optimize the task allocation. Extensive experiments show MIRSA provides real-time high quality pose estimation as an interactive service and outperforms baseline methods.
Yakun Huang, Hongru Zhao, Xiuquan Qiao, Jian Tang 0008, Ling Liu 0001
INFOCOM2
2020 Cross-Domain Brain CT Image Smart Segmentation via Shared Hidden Space Transfer FCM Clustering
abstract
Clustering is an important issue in brain medical image segmentation. Original medical images used for clinical diagnosis are often insufficient for clustering in the current domain. As there are sufficient medical images in the related domains, transfer clustering can improve the clustering performance of the current domain by transferring knowledge across the related domains. In this article, we propose a novel shared hidden space transfer fuzzy c- means (FCM) clustering called SHST-FCM for cross-domain brain computed tomography (CT) image segmentation. SHST-FCM projects both the data samples of the source domain and target domain into the shared hidden space, such that the distributions of the two domains are as close as possible. In the learned shared subspace, the data samples of the source domain serve as the auxiliary knowledge to aid the clustering process in the target domain. Extensive experiments on brain CT medical image datasets indicate the effectiveness of the proposed method.
Kaijian Xia, Hongsheng Yin 0001, Yong Jin 0003, Hongru Zhao
ACM Trans. Multim. Comput. Commun. Appl.5