EDBT 2026 Demo / reviewers in the wild / expert
Runmin Zhang
dblp:342/8502
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-8545-9966ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to correct unevenly exposed face images using RGB-NIR pairs
Jiacheng Ying, Runmin Zhang, Zhu Yu 0001, Beinan Yu, Si-Yuan Cao, Heng Yu 0001, Bailin Yang |
Pattern Recognit. | 3 |
| 2025 | Learned Image Transmission with Hierarchical Variational AutoencoderabstractIn this paper, we introduce an innovative hierarchical joint source-channel coding (HJSCC) framework for image transmission, utilizing a hierarchical variational autoencoder (VAE). Our approach leverages a combination of bottom-up and top-down paths at the transmitter to autoregressively generate multiple hierarchical representations of the original image. These representations are then directly mapped to channel symbols for transmission by the JSCC encoder. We extend this framework to scenarios with a feedback link, modeling transmission over a noisy channel as a probabilistic sampling process and deriving a novel generative formulation for JSCC with feedback. Compared with existing approaches, our proposed HJSCC provides enhanced adaptability by dynamically adjusting transmission bandwidth, encoding these representations into varying amounts of channel symbols. Additionally, we introduce a rate attention module to guide the JSCC encoder in optimizing its encoding strategy based on prior information. Extensive experiments on images of varying resolutions demonstrate that our proposed model outperforms existing baselines in rate-distortion performance and maintains robustness against channel noise. Guangyi Zhang 0005, Hanlei Li, Yunlong Cai, Qiyu Hu, Guanding Yu, Runmin Zhang |
AAAI | 6 |
| 2025 | SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split OptimizationabstractWe propose a novel unsupervised cross-modal homography estimation learning framework, named Split Supervised Homography estimation Network (SSHNet). SSHNet reformulates the unsupervised cross-modal homography estimation into two supervised sub-problems, each addressed by its specialized network: a homography estimation network and a modality transfer network. To realize stable training, we introduce an effective split optimization strategy to train each network separately within its respective sub-problem. We also formulate an extra homography feature space supervision to enhance feature consistency, further boosting the estimation accuracy. Moreover, we employ a simple yet effective distillation training technique to reduce model parameters and improve cross-domain generalization ability while maintaining comparable performance. The training stability of SSHNet enables its cooperation with various homography estimation architectures. Experiments reveal that the SSHNet using IHN as homography estimation network, namely SSHNet-IHN, outperforms previous unsupervised approaches by a significant margin. Even compared to supervised approaches MHN and LocalTrans, SSHNet-IHN achieves 47.4% and 85.8% mean average corner errors (MACEs) reduction on the challenging OPT-SAR dataset. Source code is available at https://github.com/Junchen-Yu/SSHNet. Junchen Yu, Si-Yuan Cao, Runmin Zhang, Chenghao Zhang 0002, Zhu Yu 0001, Shujie Chen 0001, Bailin Yang |
CVPR | 3 |
| 2025 | Language Driven Occupancy PredictionabstractWe introduce LOcc, an effective and generalizable framework for open-vocabulary occupancy (OVO) prediction. Previous approaches typically supervise the networks through coarse voxel-to-text correspondences via image features as intermediates or noisy and sparse correspondences from voxel-based model-view projections. To alleviate the inaccurate supervision, we propose a semantic transitive labeling pipeline to generate dense and fine-grained 3D language occupancy ground truth. Our pipeline presents a feasible way to dig into the valuable semantic information of images, transferring text labels from images to LiDAR point clouds and ultimately to voxels, to establish precise voxel-to-text correspondences. By replacing the original prediction head of supervised occupancy models with a geometry head for binary occupancy states and a language head for language features, LOcc effectively uses the generated language ground truth to guide the learning of 3D language volume. Through extensive experiments, we demonstrate that our transitive semantic labeling pipeline can produce more accurate pseudo-labeled ground truth, diminishing labor-intensive human annotations. Additionally, we validate LOcc across various architectures, where all models consistently outperform state-of-the-art zero-shot occupancy prediction approaches on the Occ3D-nuScenes dataset. Zhu Yu 0001, Lizhe Liu, Runmin Zhang, Si-Yuan Cao, Maochun Luo, Mingxia Chen |
ICCV | 4 |
| 2025 | Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionabstractThis work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a geometry and context aware aggregation module to integrate geometric and contextual information within adaptive regions in each image and dynamically adjust the contributions from different views, enhancing the representation capability of voxel features. Furthermore, we propose a sparse volume construction strategy that adaptively identifies and selects voxels with high occupancy probabilities for feature refinement, minimizing redundant computation in free space. Benefiting from the above designs, our framework achieves effective and efficient volume construction in an adaptive way. Better still, our network can be supervised using only 3D bounding boxes, eliminating the dependence on ground-truth scene geometry. Experimental results demonstrate that SGCDet achieves state-of-the-art performance on the ScanNet, ScanNet200 and ARKitScenes datasets. The source code is available at https://github.com/RM-Zhang/SGCDet. Runmin Zhang, Zhu Yu 0001, Si-Yuan Cao, Lingyu Zhu 0006, Guangyi Zhang 0005, Xiaokai Bai |
ICCV | 1 |
| 2025 | EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image RegistrationabstractPrevious deep image registration methods that employ single homography, multi-grid homography, or thin-plate spline often struggle with real scenes containing depth disparities due to their inherent limitations. To address this, we propose an Exponential-Decay Free-Form Deformation Network (EDFFDNet), which employs free-form deformation with an exponential-decay basis function. This design achieves higher efficiency and performs well in scenes with depth disparities, benefiting from its inherent locality. We also introduce an Adaptive Sparse Motion Aggregator (ASMA), which replaces the MLP motion aggregator used in previous methods. By transforming dense interactions into sparse ones, ASMA reduces parameters and improves accuracy. Additionally, we propose a progressive correlation refinement strategy that leverages global-local correlation patterns for coarse-to-fine motion estimation, further enhancing efficiency and accuracy. Experiments demonstrate that EDFFDNet reduces parameters, memory, and total runtime by 70.5%, 32.6%, and 33.7%, respectively, while achieving a 0.5 dB PSNR gain over the state-of-the-art method. With an additional local refinement stage,EDFFDNet-2 further improves PSNR by 1.06 dB while maintaining lower computational costs. Our method also demonstrates strong generalization ability across datasets, outperforming previous deep learning methods. Haokai Zhu, Bo Qu, Si-Yuan Cao, Runmin Zhang, Shujie Chen 0001, Bailin Yang |
ICCV | 4 |
| 2025 | Structure-Aware Radar-Camera Depth EstimationabstractRadar has gained much attention in autonomous driving due to its accessibility and robustness. However, its standalone application for depth perception is constrained by issues of sparsity and noise. Radar-camera depth estimation offers a more promising complementary solution. Despite significant progress, current approaches fail to produce satisfactory dense depth maps, due to the unsatisfactory processing of the sparse and noisy radar data. They constrain the regions of interest for radar points in rigid rectangular regions, which may introduce unexpected errors and confusions. To address these issues, we develop a structure-aware strategy for radar depth enhancement, which provides more targeted regions of interest by leveraging the structural priors of RGB images. Furthermore, we design a Multi-Scale Structure Guided Network to enhance radar features and preserve detailed structures, achieving accurate and structure-detailed dense metric depth estimation. Building on these, we propose a structure-aware radar-camera depth estimation framework, named SA-RCD. Extensive experiments demonstrate that our SA-RCD achieves state-of-the-art performance on the nuScenes dataset. Our code will be available at https://github.com/FreyZhangYeh/SA-RCD. Fuyi Zhang, Zhu Yu 0001, Chunhao Li, Runmin Zhang, Xiaokai Bai, Si-Yuan Cao |
ICRA | 4 |
| 2025 | Beyond Registration: Self-supervised Unknown Border Completion
Xiaokai Bai, Leyuan Yu, Runmin Zhang, Beinan Yu, Si-Yuan Cao |
PRCV (8) | 4 |
| 2025 | STARNet: Low-light video enhancement using spatio-temporal consistency aggregation
Zehua Sheng, Si-Yuan Cao, Runmin Zhang, Beinan Yu, Chenghao Zhang 0002, Bailin Yang |
Pattern Recognit. | 5 |
| 2025 | TEFormer: Thermal Infrared Image Enhancement by Preserving Spatial Consistency and DetailsabstractThermal infrared (TIR) images suffer from low contrast due to the atmospheric thermal radiation effect, especially under extreme conditions like low temperature. TIR image enhancement aims to improve image contrast, but previous enhancement approaches usually produce enhanced results with two limitations: spatial inconsistency and detail blurring. To deal with the limitations, we propose a novel TIR image enhancement method, named TEFormer, to preserve spatial consistency and restore fine-grained details during image enhancement. To preserve spatial consistency, we devise the global enhancement module (GEM) to enhance the low-resolution representation. The GEM performs long-range interactions across spatial dimensions and channel dimensions to condition the enhancement curve fitting. To keep details clear, we design the local enhancement module (LEM) as the decoding unit. The LEM injects additional detail structures into the enhanced low-resolution representation for high-resolution reconstruction. Besides, we further apply histogram-based supervision to facilitate learning in intensity distribution of clear images. Extensive experimental results on three challenging benchmarks demonstrate that the proposed method outperforms other state-of-the-art approaches. Yunxin Li, Runmin Zhang, Si-Yuan Cao, Jiacheng Ying, Xiaokai Bai, Shujie Chen 0001, Bailin Yang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SCPNet: Unsupervised Cross-Modal Homography Estimation via Intra-modal Self-supervised Learning
Runmin Zhang, Si-Yuan Cao, Lun Luo, Beinan Yu, Shujie Chen 0001, Junwei Li 0009 |
ECCV (23) | 1 |
| 2024 | Context and Geometry Aware Voxel Transformer for Semantic Scene CompletionabstractVision-based Semantic Scene Completion (SSC) has gained much attention due to its widespread applications in various 3D perception tasks. Existing sparse-to-dense approaches typically employ shared context-independent queries across various input images, which fails to capture distinctions among them as the focal regions of different inputs vary and may result in undirected feature aggregation of cross-attention. Additionally, the absence of depth information may lead to points projected onto the image plane sharing the same 2D position or similar sampling points in the feature map, resulting in depth ambiguity. In this paper, we present a novel context and geometry aware voxel transformer. It utilizes a context aware query generator to initialize context-dependent queries tailored to individual input images, effectively capturing their unique characteristics and aggregating information within the region of interest. Furthermore, it extend deformable cross-attention from 2D to 3D pixel space, enabling the differentiation of points with similar image coordinates based on their depth coordinates. Building upon this module, we introduce a neural network named CGFormer to achieve semantic scene completion. Simultaneously, CGFormer leverages multiple 3D representations (i.e., voxel and TPV) to boost the semantic and geometric representation abilities of the transformed 3D volume from both local and global perspectives. Experimental results demonstrate that CGFormer achieves state-of-the-art performance on the SemanticKITTI and SSCBench-KITTI-360 benchmarks, attaining a mIoU of 16.87 and 20.05, as well as an IoU of 45.99 and 48.07, respectively. Remarkably, CGFormer even outperforms approaches employing temporal images as inputs or much larger image backbone networks. Zhu Yu 0001, Runmin Zhang, Jiacheng Ying, Junchen Yu, Xiaohai Hu, Lun Luo, Si-Yuan Cao |
NeurIPS | 2 |
| 2023 | Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus TransformerabstractWe propose the Recurrent homography estimation framework using Homography-guided image Warping and Focus transformer (FocusFormer), named RHWF. Both being appropriately absorbed into the recurrent framework, the homography-guided image warping progressively enhances the feature consistency and the attention-focusing mechanism in FocusFormer aggregates the intra-inter correspondence in a global→nonlocal→local manner. Thanks to the above strategies, RHWF ranks top in accuracy on a variety of datasets, including the challenging cross-resolution and cross-modal ones. Meanwhile, benefiting from the recurrent framework, RHWF achieves parameter efficiency despite the transformer architecture. Compared to previous state-of-the-art approaches LocalTrans and IHN, RHWF reduces the mean average corner error (MACE) by about 70% and 38.1% on the MSCOCO dataset, while saving the parameter costs by 86.5% and 24.6%. Similar to the previous works, RHWF can also be arranged in 1-scale for efficiency and 2-scale for accuracy, with the 1-scale RHWF already outperforming most of the previous methods. Source code is available at https://github.com/imdump178/RHWF. Si-Yuan Cao, Runmin Zhang, Lun Luo, Beinan Yu, Zehua Sheng, Junwei Li 0009 |
CVPR | 2 |