VLDB 2026 Research / reviewers in the wild / expert
Ruonan Zhang 0002
dblp:55/725-2
· DBLP profile ↗
18ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-4594-3654ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology-Guided Feature Integration for Structured Object DetectionabstractIn object detection for dynamic and complex scenarios, end-to-end set prediction paradigms primarily rely on visual appearance and texture features for object discrimination. However, in scenarios characterized by dense stacking, occlusion, and a high proportion of small objects, the general features extracted based on appearance suffer from poor robustness. This deficiency leads to systematic misdetections by the model for categories that exhibit similar appearances but distinct topological structures. To address this issue, we propose an improved framework: topo fine-grained distribution refinement detector (Topo-FINE). Specifically, a topological feature extraction module is introduced to perform multi-scale shape encoding on instance masks, thereby extracting high-order topological representations of objects through structured modeling. Experimental results on newly constructed topology-feature-enhanced datasets based on ImageNet and COCO demonstrate that Topo-FINE outperforms mainstream and state-of-the-art models across multiple detection metrics. Jiayi Ding, Ruonan Zhang 0002 |
ICMR | 3 |
| 2026 | Deconstructing Centrality: Scale-Hierarchy for Hubness in Text-Video RetrievalabstractText-Video Retrieval (TVR) faces a persistent challenge known as hubness. Specifically, the uneven distribution in high-dimensional embedding spaces degrades retrieval results. Traditional methods rely solely on retrieval frequency for sample classification. They often indiscriminately suppress high-frequency samples globally while blindly boosting low-frequency ones. To address this limitation, we propose an analytical perspective based on retrieval frequency and semantic relevance. Under this perspective, we reconstruct traditional classification paradigms to reveal distinct sample attributes. We find that high-frequency regions contain both valid semantic centers and noisy interference requiring suppression. Conversely, low-frequency regions contain irrelevant outliers alongside buried high-value samples that need excavation. Based on these findings, we propose the Scale-Hierarchy Centrality Balance (SHC) framework. First, the Scale-Hierarchy Coupling (SLC) mechanism explicitly aligns multi-scale video features with multi-level text semantics. It mitigates hubness originating from single-granularity feature inconsistencies at the source. Second, the Gated Zoom-Banzhaf Interaction (GZBI) module targets the high-frequency zone. It employs Banzhaf interaction values to strictly distinguish valid samples from geometric noise. Finally, the Neighborhood Local Balance (NLB) module targets the low-frequency zone. It optimizes centrality weights to recover buried relevant samples from irrelevant outliers. Extensive experiments on MSR-VTT, ActivityNet, and DiDeMo datasets demonstrate that SHC significantly mitigates hubness. It outperforms state-of-the-art methods in both retrieval accuracy and fairness metrics. Houlin Zhu, Ruonan Zhang 0002, Zhen Deng |
ICMR | 3 |
| 2025 | A Quantum-Inspired Framework in Leader-Servant Mode for Large-Scale Multi-Modal Place RecognitionabstractMulti-modal place recognition aims to grasp diversified information implied in different modalities to bring vitality to place recognition tasks. The key challenge is rooted in the representation gap in modalities, the feature fusion method, and their relationships. The majority of existing methods are based on uni-modal, leaving these challenges unsolved effectively. To address the problems, encouraged by double-split experiments in physics and cooperation modes, in this paper, we introduce a leader-servant multi-modal framework inspired by quantum theory for large-scale place recognition. Two key modules are designed, a quantum representation module and an interference-aware fusion module. The former is designed for multi-modal data to capture their diversity and bridge the gap, while the latter is proposed to effectively fuse the multi-modal feature with the guidance of the quantum theory. Besides, we propose a leader-servant training strategy for stable training, where three cases are considered with the multi-modal loss as the leader to preserve overall characteristics and other uni-modal losses as the servants to lighten the modality influence of the leader. Furthermore, The framework is compatible with uni-modal place recognition. At last, The experiments on three datasets witness the efficiency, generalization, and robustness of the proposed method in contrast to the other existing methods. Ruonan Zhang 0002, Ge Li 0002, Wei Gao 0003, Shan Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | SDE2D: Semantic-Guided Discriminability Enhancement Feature Detector and DescriptorabstractLocal feature detectors and descriptors serve various computer vision tasks, such as image matching, visual localization, and 3D reconstruction. To address the extreme variations of rotation and light in the real world, most detectors and descriptors capture as much invariance as possible. However, these methods ignore feature discriminability and perform poorly in indoor scenes. Indoor scenes have too many weak-textured and even repeatedly textured regions, so it is necessary for the extracted features to possess sufficient discriminability. Therefore, we propose a semantic-guided method (called SDE2D) enhancing feature discriminability to improve the performance of descriptors for indoor scenes. We develop a kind of semantic-guided discriminability enhancement (SDE) loss function that uses semantic information from indoor scenes. To the best of our knowledge, this is the first deep research that applies semantic segmentation to enhance discriminability. In addition, we design a novel framework that allows semantic segmentation network to be well embedded as a module in the overall framework and provides guidance information for training. Besides, we explore the impact of different semantic segmentation models on our method. The experimental results on indoor scenes datasets demonstrate that the proposed SDE2D performs well compared with the state-of-the-art models. Ruonan Zhang 0002, Ge Li 0002, Thomas H. Li |
IEEE Trans. Multim. | 2 |
| 2024 | Sketch-aided Interactive Fusion Point Cloud Place RecognitionabstractExisting point cloud place recognition methods ignore textureless descriptions of scenes by point clouds. This further leads to lower generalization and bottlenecks in performance improvement. To solve these problems, we propose a novel sketch-aided interactive fusion point cloud place recognition method, which involves two networks to separately deal with point clouds and sketches and an interaction feature fused module to fuse features mathematically. Specifically, this is the first time to introduce sketches to guide the point cloud place recognition task as far as we know. The sketch-aided part and the point cloud could enhance the texture structure of the scene which is omitted in only the point cloud scenario. Meanwhile, we devise an interactive feature fusion module for fusing two features, which is encouraged by square summation in math. This module reflects the communication between features as well as the non-linear influence on the fused feature without bringing dimension growth. The experiments on two datasets witness the effectiveness of the proposed method in performance improvement and generalization subjectively and objectively. Ruonan Zhang 0002, Ge Li 0002, Thomas H. Li |
ICMR | 1 |
| 2024 | ComPoint: Can Complex-Valued Representation Benefit Point Cloud Place Recognition?abstract“Where was this place?”-figuring out the location of a point cloud scene is a challenge that has attracted researchers in recent years, under the name of point cloud place recognition. Driven by the drastic acceleration of 3D data and corresponding technique forces, research in this field has witnessed remarkable progress. However, the existing methods are stuck in a dilemma, the limited capability of feature representations needs more complicated architectures to enhance the performance further. This inspires us to envision whether there is a better representation for this task. To explore its possibility, in this paper, we propose a new framework, dubbed ComPoint, in the form of complex-valued representations for large-scale point cloud place recognition. Theoretically, the framework is guided by two proven propositions where one implies that richer information provided by the complex-valued representations of point clouds can benefit the performance. Practically speaking, ComPoint is designed with three modules, with each module highlighting its different characteristics. First, Com-Transform guarantees informative data delivery by mining informative complex-valued representations of initial point clouds. Next, Com-Perception perceives and digs deeper into complex-valued features via a series of simple-design convolution blocks, i.e.,ComplexPointConvandComplexPointFT. Then, Com-Fusion dynamic aggregates and interacts with the above features to obtain compact global complex-valued ones based on devised effective soft-balancing block in the VLAD network without involving extra memory footprint. Finally, our method is trained with proper strategies that are analyzed in-depth. The proposed method is witnessed to outperform the prior methods on four large-scale benchmarks quantitatively and qualitatively. It is also flexible plug-and-play in other approaches to improve their performance. Ruonan Zhang 0002, Ge Li 0002, Wei Gao 0003, Thomas H. Li |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | A Point is A Wave: Point-Wave Network for Place RecognitionabstractPoint cloud place recognition is key to auto-driving, navigation, localization, and robotics. It targets finding a similar scene of the query in the database via extracted compact point features. The core challenge focuses on obtaining descriptive features to enhance retrieval performance. Existing methods concentrate on the multi-layer perception with intricate architectures, needing lots of parameters to learn with limited gains. Unlike these methods, we propose an innovative method by designing a point-wave module, modeling a point as a wave function to avoid losing the information of origin points. In this way, it dynamically promotes the intercommunication among point features to advance the ultimate performance. Meanwhile, our designed point-wave architecture benefits the existing point-based methods to improve performance and save half convergence time with fewer learned parameters. Experiments on four datasets also show that the proposed method brings performance gains and is an easy plug-and-play with a lightweight property. Ge Li 0002, Ruonan Zhang 0002 |
ICASSP | 2 |
| 2023 | Frequency-Aware Self-Supervised Monocular Depth EstimationabstractWe present two versatile methods to generally enhance self-supervised monocular depth estimation (MDE) models. The high generalizability of our methods is achieved by solving the fundamental and ubiquitous problems in photometric loss function. In particular, from the perspective of spatial frequency, we first propose Ambiguity-Masking to suppress the incorrect supervision under photometric loss at specific object boundaries, the cause of which could be traced to pixel-level ambiguity. Second, we present a novel frequency-adaptive Gaussian low-pass filter, designed to robustify the photometric loss in high-frequency regions. We are the first to propose blurring images to improve depth estimators with an interpretable analysis. Both modules are lightweight, adding no parameters and no need to manually change the network structures. Experiments show that our methods provide performance boosts to a large number of existing models, including those who claimed state-of-the-art, while introducing no extra inference computation at all. Thomas H. Li, Ruonan Zhang 0002, Ge Li 0002 |
WACV | 3 |
| 2023 | Self-Supervised Monocular Depth Estimation: Solving the Edge-Fattening ProblemabstractSelf-supervised monocular depth estimation (MDE) models universally suffer from the notorious edge-fattening issue. Triplet loss, as a widespread metric learning strategy, has largely succeeded in many computer vision applications. In this paper, we redesign the patch-based triplet loss in MDE to alleviate the ubiquitous edge-fattening issue. We show two drawbacks of the raw triplet loss in MDE and demonstrate our problem-driven redesigns. First, we present a min. operator based strategy applied to all negative samples, to prevent well-performing negatives sheltering the error of edge-fattening negatives. Second, we split the anchor-positive distance and anchor-negative distance from within the original triplet, which directly optimizes the positives without any mutual effect with the negatives. Extensive experiments show the combination of these two small redesigns can achieve unprecedented results: Our powerful and versatile triplet loss not only makes our model outperform all previous SoTA by a large margin, but also provides substantial performance boosts to a large number of existing models, while introducing no extra inference computation at all. Ruonan Zhang 0002, Ji Jiang, Ge Li 0002, Thomas H. Li |
WACV | 2 |
| 2022 | Pointivae: Invertible Variational Autoencoder Framework for 3D Point Cloud GenerationabstractPoint cloud generation is a challenging task and has drawn great attention in 3D vision community. However, existing methods rarely consider to exploit local features, leading to unsatisfactory generated results that lack of high frequency. In this paper, we put forward a novel point cloud generation framework called PointIVAE, which adopts VAE based framework to construct local relations and enhance generating capability. PointIVAE contains three components, including an encoder, a flow model and a decoder. Specially, the encoder aims to aggregate neighborhood relations and provides high-quality latent codes. We then propose the invertible residual coupling stack in the flow model, in order to learn from the latent codes via an invertible manner. Based on the shape latent codes generated by the flow, the decoder converts the input noises into point clouds in an inverse way. Experimental results demonstrate that PointIVAE obtains the SOTA results in both point cloud generation and autoencoding. Ge Li 0002, Ruonan Zhang 0002, Thomas H. Li, Wei Gao 0003 |
ICIP | 3 |
| 2022 | SparseARFM-SI: Rotary Point Cloud Place Recognition Based on Multi-Resolution and Attention MechanismabstractPlace recognition based on 3D point cloud can directly describe the scene of 3D world by acquiring 3D point clouds via LiDAR, which is robust with environmental changes. The core challenge locates at how to obtain compact and representative feature expression for positioning. In this paper, SparseARFM-SI framework is proposed to solve the problem of feature extraction caused by the large size difference of different objects and the difficulty of feature extraction based on the point cloud data collected by rotating LiDAR. SparseARFM-SI is mainly composed of four models: data representation model, sparse convolution model, ARFM model and NetVLAD model. Experiments show that SparseARFM-SI framework has good performance in USyd data set collected based on rotating Lidar. Ruonan Zhang 0002, Jie Wang 0003, Ge Li 0002 |
VCIP | 2 |
| 2022 | PointNetGeM: Simple and Efficient Point Cloud Based Network for Place RecognitionabstractPoint Cloud based place recognition is a popular area of current research. Most existing methods add additional structures to PointNet to further extract compact information from point clouds, e.g. PointNetVLAD, PCAN. These methods achieved good results but incur severe computational overhead, increase structural complexity and are not easily scalable. In this paper, we propose a simple and efficient method named PointNetGeM. We optimize the training process and achieve discriminative global feature descriptors. Experiments on several datasets demonstrate the effectiveness of our network structure. Keli Wen, Ruonan Zhang 0002, Ge Li 0002 |
VCIP | 2 |
| 2022 | ERINet: Effective Rotation Invariant Network for Point Cloud based Place RecognitionabstractPlace recognition task is a crucial part of 3D scene recognition in various applications. Nowadays, learning-based point cloud place recognition approaches have achieved remarkable success. However, these methods seldom consider the possible rotation of point cloud data in large-scale real-world place recognition tasks. To cope with this problem, in this work, we propose a novel effective rotation invariant network for large-scale place recognition named ERINet, which captures the recent successful deep network architecture and benefits from holding the rotation-invariant property of point clouds. In this network, we design a core effective rotation invariant module, which enhances the ability to extract rotation-invariant features of 3D point clouds. The benchmark experiments illustrate that our network boosts the performance of the recent works on all evaluation metrics with various rotations, even under challenging rotation cases. Shichen Weng, Ruonan Zhang 0002, Ge Li 0002 |
VCIP | 2 |
| 2022 | PointOT: Interpretable Geometry-Inspired Point Cloud Generative Model via Optimal TransportabstractPoint cloud generative models have aroused increasing concern for their realistic generation potentialities. However, most existing methods adopt deep-neural-network (DNN) models for continuous mapping. DNN usually induces mode collapse and mixture problems without clear interpretation. Consequently, in this paper, we design a geometry-inspired point cloud generative framework called PointOT. PointOT decouples the generative model into two separate sub-tasks: manifold learning of the point cloud and distribution transformation. Then, we propose corresponding instances according to the requirements in each sub-task, where they can be established by the point cloud auto-encoder (AE) and the semi-continuous optimal transportation (SCOT) mapping, respectively. In particular, the transportation map between the source and the target distributions is discrete rather than continuous in geometric view. The learned continuous shape model of the DNN point cloud does not conform with the discrete distribution transformation. Therefore, the proposed SCOT efficiently relieves these problems by connecting the continuous-to-discrete domain. Besides, we provide theoretical explanations from a geometric view and analyze the fundamental reason for mode collapse and mixture in point cloud generative models. The proposed SCOT algorithm without the DNN model is computationally efficient and makes the original black box semi-transparent. Final experiments validate the virtue of the proposed approach, including the designed decomposition framework and the rigorous theory. Ruonan Zhang 0002, Wei Gao 0003, Ge Li 0002, Thomas H. Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | QINet: Decision Surface Learning and Adversarial Enhancement for Quasi-Immune Completion of Diverse Corrupted Point CloudsabstractIn point cloud completion task, most previous works fail to deal with diverse corrupted point clouds with large missing areas. Meanwhile, they are restricted by discrete point clouds lacking smooth surfaces to represent an object, and the resolution of generated point clouds is fixed once their networks are determined. In addition, the evaluation metrics are not specific for this task. Thus, we propose an innovative quasi-immune completion architecture of point cloud calledQINetin this paper, which is inspired by the artificial immunization process in biology. Specifically, to increase robustness and adaptation of the model, we conceive a mask algorithm named onion-peeling to generate diverse corrupted inputs. Meanwhile, two proposed modules are combined together to produce flexible resolution of point clouds, namely the decision surface learning and adversarial enhancement for the latent representation recovery. The first module transforms point clouds to surfaces with a continuous decision boundary function, which the second module is applied to deduce complete surface from corrupted point cloud by the cooperation of reinforcement learning and latent generative adversarial network. Besides, we evaluate the shortcomings of the existing methods and present two novel metrics to support multi-faceted comparisons. Experimental results verify that our approach can generate continuous 3D shapes with optional resolutions compared to other approaches, and achieves competitive results both quantitatively and qualitatively. Ruonan Zhang 0002, Wei Gao 0003, Ge Li 0002, Thomas H. Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Vaccine-style-net: Point Cloud Completion in Implicit Continuous Function SpaceabstractThough recent advances in point cloud completion have shown exciting promise with learning-based methods, most of them still generate coarse point clouds with a fixed number of points (e.g. 2048). In this paper, we propose Vaccine-Style-Net, a new point cloud completion method that can produce high resolution 3D shapes with complete smooth surface. Vaccine-Style-Net performs point cloud completion in the function space of 3D surface, which represent the 3D surface as the continuous decision boundary function. Meanwhile, a reinforcement learning agent is embedded to deduce the complete 3D geometry from the incomplete point cloud. In contrast to the existing approaches, the completed 3D shapes produced by our method can be any resolution without excessive memory footprint. Moreover, to increase the diversity and adaptability of the method, we introduce two-type-free-form masks to simulate various corrupted inputs as well as a mask dataset called onion-peeling-mask (OPM). Finally, we discuss the limitations of existing evaluation metrics for shape completion tasks and explore a novel metric to supplement the existing ones. Experiments demonstrate that our method not only achieves competitive results qualitatively and quantitatively but also can produce a continuous 3D shape with any resolution. Ruonan Zhang 0002, Jing Wang 0115, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
ACM Multimedia | 2 |
| 2019 | Base-detail image inpainting
Ruonan Zhang 0002, Yurui Ren, Jingfei Qiu, Ge Li 0002 |
BMVC | 1 |
| 2019 | StructureFlow: Image Inpainting via Structure-Aware Appearance FlowabstractImage inpainting techniques have shown significant improvements by using deep neural networks recently. However, most of them may either fail to reconstruct reasonable structures or restore fine-grained textures. In order to solve this problem, in this paper, we propose a two-stage model which splits the inpainting task into two parts: structure reconstruction and texture generation. In the first stage, edge-preserved smooth images are employed to train a structure reconstructor which completes the missing structures of the inputs. In the second stage, based on the reconstructed structures, a texture generator using appearance flow is designed to yield image details. Experiments on multiple publicly available datasets show the superior performance of the proposed network. Yurui Ren, Xiaoming Yu, Ruonan Zhang 0002, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
ICCV | 3 |