Zhong Wang 0009

dblp:02/2025-9 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-6206-526XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images
abstract
Recent advances in generative AI have accelerated the production of ultra-high-resolution visual content. However, traditional image formats face significant limitations in efficient compression and real-time decoding, which restricts their applicability on end-user devices. Inspired by 3D Gaussian Splatting, 2D Gaussian image models have achieved notable progress in enhancing image representation efficiency and quality. Nevertheless, existing methods struggle to balance compression ratios and reconstruction fidelity in ultra-high-resolution scenarios. To address these challenges, we propose SmartSplat, a highly adaptive and feature-aware GS-based image compression framework that effectively supports arbitrary image resolutions and compression ratios. By leveraging image-aware features such as gradients and color variances, SmartSplat introduces a Gradient-Color Guided Variational Sampling strategy alongside an Exclusion-based Uniform Sampling scheme, significantly improving the non-overlapping coverage of Gaussian primitives in pixel space. Additionally, a Scale-Adaptive Gaussian Color Sampling method is proposed to enhance the initialization of Gaussian color attributes across scales. Through joint optimization of spatial layout, scale, and color initialization, SmartSplat can efficiently capture both local structures and global textures of images using a limited number of Gaussians, achieving superior reconstruction quality under high compression ratios. Extensive experiments on DIV8K and a newly created 16K dataset demonstrate that SmartSplat significantly outperforms state-of-the-art methods at comparable compression ratios and surpasses their compression limits, exhibiting strong scalability and practical applicability. This framework can effectively alleviate the storage and transmission burdens of ultra-high-resolution images, providing a robust foundation for future high-efficiency visual content processing.
Linfei Li, Lin Zhang 0014, Zhong Wang 0009, Ying Shen 0005
AAAI3
2026 SNI-SLAM++: Tightly-Coupled Semantic Neural Implicit SLAM
abstract
We propose SNI-SLAM++, a tightly-coupled semantic SLAM system utilizing neural implicit representation, that simultaneously performs accurate semantic mapping, high-quality surface reconstruction, and robust camera tracking. Our system tightly integrates visual appearance, geometry, and semantics through five key components: (i) We introduce hierarchical semantic representation to allow multi-level semantic comprehension for top-down structured semantic mapping of the scene. (ii) To fully utilize the correlation between multiple attributes of the environment, we integrate appearance, geometry and semantic features through cross-attention for feature collaboration. This strategy enables a more multifaceted understanding of the environment, thereby allowing SNI-SLAM++ to remain robust even when single attribute is defective. (iii) We design an internal fusion-based decoder to obtain semantic, RGB, and Truncated Signed Distance Field (TSDF) values from multi-level features for accurate decoding. (iv) We introduce a semantics-coupled tracking framework that tightly incorporates semantic constraints for camera pose estimation in neural implicit SLAM. This framework leverages the multi-view consistency of semantics to construct a pose graph and perform semantic loop closure optimization, enabling robust tracking. (v) We propose a feature loss to update the scene representation at the feature level. Compared with low-level losses such as RGB loss and depth loss, our feature loss is capable of guiding the network optimization on a higher level. Our SNI-SLAM++ demonstrates superior performance over all recent visual SLAM methods in terms of mapping and tracking accuracy on the datasets of Replica, ScanNet, TUM-RGBD, and ScanNet++, while also showing excellent capabilities in accurate semantic segmentation and 3D semantic mapping.
Siting Zhu 0001, Guangming Wang 0001, Hermann Blum, Zhong Wang 0009, Ganlin Zhang 0001, Daniel Cremers, Marc Pollefeys, Hesheng Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Representing Sounds as Neural Amplitude Fields: A Benchmark of Coordinate-MLPs and a Fourier Kolmogorov-Arnold Framework
abstract
Although Coordinate-MLP-based implicit neural representations have excelled in representing radiance fields, 3D shapes, and images, their application to audio signals remains underexplored. To fill this gap, we investigate existing implicit neural representations, from which we extract 3 types of positional encoding and 16 commonly used activation functions. Through combinatorial design, we establish the first benchmark for Coordinate-MLPs in audio signal representations. Our benchmark reveals that Coordinate-MLPs require complex hyperparameter tuning and frequency-dependent initialization, limiting their robustness. To address these issues, we propose Fourier-ASR, a novel framework based on the Fourier series theorem and the Kolmogorov-Arnold representation theorem. Fourier-ASR introduces Fourier Kolmogorov-Arnold Networks (Fourier-KAN), which leverage periodicity and strong nonlinearity to represent audio signals, eliminating the need for additional positional encoding. Furthermore, a Frequency-adaptive Learning Strategy (FaLS) is proposed to enhance the convergence of Fourier-KAN by capturing high-frequency components and preventing overfitting of low-frequency signals. Extensive experiments conducted on natural speech and music datasets reveal that: (1) well-designed positional encoding and activation functions in Coordinate-MLPs can effectively improve audio representation quality; and (2) Fourier-ASR can robustly represent complex audio signals without extensive hyperparameter tuning. Looking ahead, the continuity and infinite resolution of implicit audio representations make our research highly promising for tasks such as audio compression, synthesis, and generation.
Linfei Li, Lin Zhang 0014, Zhong Wang 0009, Ying Shen 0005
AAAI3
2025 Towards Autonomous Indoor Parking: A Globally Consistent Semantic SLAM System and A Semantic Localization Subsystem
abstract
We propose a globally consistent semantic SLAM system (GCSLAM) and a semantic-fusion localization subsystem (SF-Loc), which achieves accurate semantic mapping and robust localization in complex parking lots. Visual cameras (front-view and surround-view), IMU, and wheel encoder form the input sensor configuration of our system. The first part of our work is GCSLAM. GCSLAM introduces a semantic-constrained factor graph for the optimization of poses and semantic map, which incorporates innovative error terms based on multi-sensor data and BEV (bird’s-eye view) semantic information. Additionally, GCSLAM integrates a Global Slot Management module that stores and manages parking slot observations. SF-Loc is the second part of our work, which leverages the semantic map built by GCSLAM to conduct map-based localization. SF-Loc integrates registration results and odometry poses with a novel factor graph. Our system demonstrates superior performance over existing SLAM on two real-world datasets, showing excellent capabilities in robust global localization and precise semantic mapping.
Yichen Sha, Siting Zhu 0001, Hekui Guo, Zhong Wang 0009, Hesheng Wang 0001
IROS4
2025 S²KAN-SLAM: Elastic Neural LiDAR SLAM With SDF Submaps and Kolmogorov-Arnold Networks
abstract
Traditional LiDAR SLAM approaches prioritize localization over mapping, yet high-precision dense maps are essential for numerous applications involving intelligent agents. Recent advancements have introduced methods leveraging neural fields to enhance mapping capabilities; however, these approaches still face several limitations. Firstly, concerning scene representation, they typically employ neural fields with high-dimensional features and multi-layer perceptron decoders utilizing non-continuous activation functions. This results in low learning efficiency and challenges in capturing high-frequency signals. Secondly, in terms of scene organization, these methods often treat the entire scene as a singular neural field, leading to inefficiencies, inflexibility, and difficulties in rectifying accumulated errors when mapping large-scale environments over extended periods. To tackle the first issue, we propose a lightweight continuous SDF regression approach by encoding the scene in single-valued embeddings and decoding SDF values from a Kolmogorov-Arnold Network. By minimizing discrepancies in measuring range, sampling distance, and decoded SDF values, we facilitate iterative frame-to-model tracking and bundle adjustment neural mapping. To mitigate the second challenge, we propose structuring the whole scene into multiple neural SDF submaps. By establishing node-node, node-submap, and loop closure constraints into a global pose graph, the system can create dense neural maps with global consistency across large-scale scenes. Experimental evaluations in both real-world and simulated settings indicate that our system achieves superior mapping completeness and accuracy, enhanced learning efficiency, reduced memory consumption, and greater flexibility compared to its counterparts.
Zhong Wang 0009, Lin Zhang 0014, Hesheng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Skeleton-Aware Graph-Based Adversarial Networks for Human Pose Estimation from Sparse IMUs
abstract
Recently, sparse-inertial human pose estimation (SI-HPE) with only a few IMUs has shown great potential in various fields. The most advanced work in this area achieved fairish results using only six IMUs. However, there are still two major issues that remain to be addressed. First, existing methods typically treat SI-HPE as a temporal sequential learning problem and often ignore the important spatial prior of skeletal topology. Second, there are far more synthetic data in their training data than real data, and the data distribution of synthetic data and real data is quite different, which makes it difficult for the model to be applied to more diverse real data. To address these issues, we propose “Graph-based Adversarial Inertial Poser (GAIP),” which tracks body movements using sparse data from six IMUs. To make full use of the spatial prior, we design a multi-stage pose regressor with graph convolution to explicitly learn the skeletal topology. A joint position loss is also introduced to implicitly mine spatial information. To enhance the generalization ability, we propose supervising the pose regression with an adversarial loss from a discriminator, bringing the ability of adversarial networks to learn implicit constraints into full play. Additionally, we construct a real dataset that includes hip support movements and a synthetic dataset containing various motion categories to enrich the diversity of inertial data for SI-HPE. Extensive experiments demonstrate that GAIP produces results with more precise limb movement amplitudes and relative joint positions, accompanied by smaller joint angle and position errors compared to state-of-the-art counterparts. The datasets and codes are publicly available at https://cslinzhang.github.io/GAIP/ .
Kaixin Chen 0003, Lin Zhang 0014, Zhong Wang 0009, Shengjie Zhao 0001, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.3
2024 IMU-Assisted Target-Free Extrinsic Calibration of Heterogeneous Lidars Based on Continuous-Time Optimization
abstract
Data fusion of heterogeneous LiDAR systems has gained significant attention due to its potential for providing wide-range sensing and high-density measurements for robots. However, existing LiDAR calibration methods primarily focus on homogeneous LiDAR systems and yield suboptimal outcomes when applied to heterogeneous setups. To this end, this paper proposes an IMU-Assisted Heterogeneous LiDAR extrinsics Calibration method, namely IA-HeLiC, which is a target-free method based on continuous-time optimization. Specifically, IA-HeLiC utilizes two types of errors, namely geometric constraint error and motion constraint error, and minimizes them within a B-spline-based continuous-time framework to achieve accurate extrinsic calibration. Using a parameter loopback mechanism, this optimization process is performed iteratively to further improve calibration accuracy. IA-HeLiC’s performance is corroborated through experiments using a ground-truth-known handheld device, by which multiple data sequences were collected in diverse real-world scenes. To make our results reproducible, the source code and the collected dataset have been released at https://cslinzhang.github.io/IA-HeLiC.
Zehao Yan, Lin Zhang 0014, Zhong Wang 0009, Shenjie Zhao
ICIP3
2024 GS3LAM: Gaussian Semantic Splatting SLAM
abstract
Recently, the multi-modal fusion of RGB, depth, and semantics has shown great potential in the domain of dense Simultaneous Localization and Mapping (SLAM), as known as dense semantic SLAM. Yet a prerequisite for generating consistent and continuous semantic maps is the availability of dense, efficient, and scalable scene representations. To date, existing semantic SLAM systems based on explicit scene representations (points/meshes/surfels) are limited by their resolutions and inabilities to predict unknown areas, thus failing to generate dense maps. Contrarily, a few implicit scene representations (Neural Radiance Fields) to deal with these problems rely on time-consuming ray tracing-based volume rendering technique, which cannot meet the real-time rendering requirements of SLAM. Fortunately, the Gaussian Splatting scene representation has recently emerged, which inherits the efficiency and scalability of point/surfel representations while smoothly represents geometric structures in a continuous manner, showing promise in addressing the aforementioned challenges. To this end, we propose GS3LAM, a Gaussian Semantic Splatting SLAM framework, which takes multimodal data as input and can render consistent, continuous dense semantic maps in real-time. To fuse multimodal data, GS3LAM models the scene as a Semantic Gaussian Field (SG-Field), and jointly optimizes camera poses and the field by establishing error constraints between observed and predicted data. Furthermore, a Depth-adaptive Scale Regularization (DSR) scheme is proposed to tackle the problem of misalignment between scale-invariant Gaussians and geometric surfaces within the SG-Field. To mitigate the forgetting phenomenon, we propose an effective Random Sampling-based Keyframe Mapping (RSKM) strategy, which exhibits notable superiority over local covisibility optimization strategies commonly utilized in 3DGS-based SLAM systems. Extensive experiments conducted on the benchmark datasets reveal that compared with state-of-the-art competitors, GS3 LAM demonstrates increased tracking robustness, superior real-time rendering quality, and enhanced semantic reconstruction precision. To make the results reproducible, the source code is available at https://github.com/lif314/GS3LAM.
Linfei Li, Lin Zhang 0014, Zhong Wang 0009, Ying Shen 0005
ACM Multimedia3
2024 Style Reconstruction-Driven Networks for Occlusion-Aware License Plate Recognition
abstract
Deep learning-based automatic license plate recognition methods have made significant advancements and are now widely used in real-world applications. Currently, license plate character recognition primarily relies on classification and the Connectionist Temporal Classification approach. While these methods achieve high recognition accuracy, they face challenges in accurately estimating confidence in occlusion scenarios. To solve this problem, we propose a novel style reconstruction-based network that can transform input license plates into standardized images. It computes character prediction confidence using a lightweight matching module, which effectively reduces the confidence score for occluded sections. Our network ingeniously integrates deep learning with traditional character segmentation methods, offering a fresh perspective on license plate recognition. Besides, to address the class imbalance in existing license plate datasets, we propose a novel synthetic license plate generation method that exchanges styles between real and standardized license plates. In order to comprehensively evaluate license plate recognition models across different regions, we construct and release the Chinese Balanced License Plates (CBLP) dataset, which includes over 30,000 images from all provinces in mainland China. Experimental evaluations on multiple datasets demonstrate that our methods achieve state-of-the-art performance. The code and dataset are available athttps://github.com/tj-cvrsg/lpsrnet.
Shaoming Zhang, Zhong Wang 0009
IEEE Trans. Circuits Syst. Video Technol.4
2024 Global Localization in Large-Scale Point Clouds via Roll-Pitch-Yaw Invariant Place Recognition and Low-Overlap Global Registration
abstract
For autonomous ground vehicles, global localization with 3D LiDAR is an indispensable part of tasks such as navigation. Usually, global localization using LiDAR is subdivided into two sub-problems, place recognition and global registration. For place recognition, the recent emerging schemes based on deep learning either rely on 3D convolution with high complexity or need to learn features from various forward perspectives. To mitigate this, we propose a model with roll-pitch-yaw invariance that represents point clouds as probabilistic voxels and generates occupancy grids from a bird’s-eye view, fulfilling robust place recognition by learning aggregated embeddings from a fixed perspective. For low-overlap global registration, the traditional handcraft feature-based methods are mostly limited to dense object-level point clouds, while the state-of-the-art learning-based approaches often rely on complex 3D convolution and additional feature association learning. To fill this gap to some extent, we propose to estimate the relative roll-pitch angles and vertical translation by fitting and aligning the ground plane of the point clouds and to determine the horizontal translations and yaw angle by matching their projected occupancy grids. Extensive experiments corroborate the superior recall and generalization ability of our place recognition model, as well as the advanced success rate and accuracy of our 3D registration approach. Especially in the recognition and registration of hard samples, our results far exceed those of our counterparts by large margins. To ensure full reproducibility, the relevant codes and data are made available online.
Zhong Wang 0009, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.1
2024 Ct-LVI: A Framework Toward Continuous-Time Laser-Visual-Inertial Odometry and Mapping
abstract
Owing to the inherent complementarity among LiDAR, camera, and IMU, a growing effort has been paid to laser-visual-inertial SLAM recently. The existing approaches, however, are limited in two aspects. First, at the front-end, they usually employ a discrete-time representation that requires high-precision hardware/software synchronization and are based on geometric laser features, leading to low robustness and scalability. Second, at the backend, visual loop constraints suffer from scale ambiguity and the sparseness of the point cloud deteriorates the scan-to-scan loop detection. To solve these problems, for the front-end, we propose a continuous-time laser-visual-inertial odometry which formulates the carrier trajectory in continuous time, organizes point clouds in probabilistic submaps, and jointly optimizes the loss terms of laser anchors, visual reprojections, and IMU readings, achieving accurate pose estimation even with fast motion or in unstructured scenes where it is difficult to extract meaningful geometric features. At the backend, we propose building 5-DoF laser constraints by matching projected 2D submaps and 6-DoF visual constraints via laser-aided visual relocalization, ensuring mapping consistency in large-scale scenes. Results show that our framework achieves high-precision estimation and is more robust than its counterparts when the carrier works in large scenes or with fast motion. The relevant codes and data are open-sourced at https://cslinzhang.github.io/Ct-LVI/Ct-LVI.html.
Zhong Wang 0009, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.1
2024 Ultra-Lightweight Automatic License Plate Recognition System for Microcontrollers: A Cost-Effective and Energy-Efficient Solution
abstract
In the realm of intelligent transportation systems, Automatic License Plate Recognition (ALPR) is integral, yet is often hindered by heavy computational loads and massive memory footprints. Existing methods either enhance recognition accuracy through complex models or minimize resource requirements by utilizing low-performance models and low-resolution inputs, both of which pose challenges in real-world applications. To address these issues, this paper introduces a novel ultra-lightweight ALPR system specifically designed for deployment on microcontrollers, offering a cost-effective and energy-efficient solution for large-scale applications. We simplify the process by decoupling ALPR into three sub-tasks: detection, localization, and recognition. For each sub-task, specific improvement strategies such as Residual Log-likelihood Estimation (RLE)-based keypoint localization, adaptive data augmentation, and multi-voting recognition mechanism are designed. Our proposed system naturally mitigates the issues of large perspective transformations frequently encountered in real-world settings. As a supplementary contribution, we introduce and release the TJLP dataset. Practical application and superior performance of our method have been demonstrated on resource-limited devices like RV1106 (ARM Cortex-A7 with a 0.5 TOPS NPU), achieving an inference speed of 22 FPS with a power consumption of 629 mW. Notably, our fully optimized model exhibits a recognition accuracy of 99.23% in real-world environments.
Shaoming Zhang, Zhong Wang 0009
IEEE Trans. Intell. Transp. Syst.4
2024 I2P Registration by Learning the Underlying Alignment Feature Space from Pixel-to-Point Similarities
abstract
Estimating the relative pose between a camera and a LiDAR holds paramount importance in facilitating complex task execution within multi-agent systems. Nonetheless, current methodologies encounter two primary limitations. First, amid the cross-modal feature extraction, they typically employ separate modal branches to extract cross-modal features from images and point clouds. This approach results in the feature spaces of images and point clouds being misaligned, thereby reducing the robustness of establishing correspondences. Second, due to the scale differences between images and point clouds, one-to-many pixel-point correspondences are inevitably encountered, which will mislead the pose optimization. To address these challenges, we propose a framework named I mage-to- P oint cloud registration by learning the underlying alignment feature space from P ixel-to- P oint SIM imilarities (I2P \({}_{\mathbf{ppsim}}\) ) . Central to \(\text{I2P}_{\text{ppsim}}\) is a Shared Feature Alignment Module (SFAM). It is designed under on a coarse-to-fine architecture and uses a weight-sharing network to construct an alignment feature space. Benefiting from SFAM, \(\text{I2P}_{\text{ppsim}}\) can effectively identify the co-view regions between images and point clouds and establish high-reliability 2D-3D correspondences. Moreover, to mitigate the one-to-many correspondence issue, we introduce a similarity maximization strategy termed point-max. This strategy effectively filters out outliers, thereby establishing accurate 2D-3D correspondences. To evaluate the efficacy of our framework, we conduct extensive experiments on KITTI Odometry and Oxford Robotcar. The results corroborate the effectiveness of our framework in improving image-to-point cloud registration. To make our results reproducible, the source codes have been released at https://cslinzhang.github.io/I2P
Yunda Sun, Lin Zhang 0014, Zhong Wang 0009, Yang Chen 0037, Shengjie Zhao 0001, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.3
2023 D-LIOM: Tightly-Coupled Direct LiDAR-Inertial Odometry and Mapping
abstract
Simultaneous localization and mapping via LiDAR-Inertial fusion is a crucial technology in many automation-related applications. Recently, a number of approaches based on geometric features have evolved, yielding impressive results via tightly-coupled estimation. This sort of feature-based techniques, however, are inextricably linked to the scanning mechanism of the LiDAR, relying on stable feature detection, and thus are difficult to adapt to multi-LiDAR systems. A few “direct” solutions, on the other hand, register the raw point cloud with the built probability map, which is more computationally efficient and easy to be extended. But, the existing direct approaches are all loosely-coupled, lacking correction of the IMU biases, and thus only work well in 2D cases. To this end, we present D-LIOM, a tightly-coupled Direct LiDAR-Inertial Odometry and Mapping framework. In D-LIOM, a scan is directly registered to a probability submap, and the LiDAR odometry, the IMU pre-integration, and the gravity constraint are integrated to build a local factor graph in the submap's time window, allowing the system to perform real-time high-precision pose estimation. Furthermore, to eliminate accumulated errors in time, we detect loops and adjust the sparse pose graph based on mutual matching of projected 2D submaps, allowing D-LIOM to run stably in large-scale scenes. In addition, to improve its flexibility to varied sensor combinations, D-LIOM supports multi-LiDAR inputs and facilitates the initialization with a common 6-axis IMU. Extensive experiments demonstrate that D-LIOM largely outperforms the existing state-of-the-art counterparts in mapping effect and localization accuracy as well as with high time efficiency. Lastly, to ensure that our results are entirely reproducible, all necessary data and codes are made open-source available. One introduction video can also be found on the online website.
Zhong Wang 0009, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
IEEE Trans. Multim.1
2022 LVI-ExC: A Target-free LiDAR-Visual-Inertial Extrinsic Calibration Framework
abstract
Recently, the multi-modal fusion with 3D LiDAR, camera, and IMU has shown great potential in applications of automation-related fields. Yet a prerequisite for a successful fusion is that the geometric relationships among the sensors are accurately determined, which is called an extrinsic calibration problem. To date, the existing target-based approaches to deal with this problem rely on sophisticated calibration objects (sites) and well-trained operators, which is time-consuming and inflexible in practical applications. Contrarily, a few target-free methods can overcome these shortcomings, while they only focus on the calibrations of two types of the sensors. Although it is possible to obtain LiDAR-visual-inertial extrinsics by chained calibrations, problems such as cumbersome operations, large cumulative errors, and weak geometric consistency still exist. To this end, we propose LVI-ExC, an integrated LiDAR-Visual-Inertial Extrinsic Calibration framework, which takes natural multi-modal data as input and yields sensor-to-sensor extrinsics end-to-end without any auxiliary object (site) or manual assistance. To fuse multi-modal data, we formulate the LiDAR-visual-inertial extrinsic calibration as a continuous-time simultaneous localization and mapping problem, in which the extrinsics, trajectories, time differences, and map points are jointly estimated by establishing sensor-to-sensor and sensor-to-trajectory constraints. Extensive experiments show that LVI-ExC can produce precise results. With LVI-ExC's outputs, the LiDAR-visual reprojection results and the reconstructed environment map are all highly consistent with the actual natural scenes, demonstrating LVI-ExC's outstanding performance. To ensure that our results are fully reproducible, all the relevant data and codes have been released publicly at https://cslinzhang.github.io/LVI-ExC/.
Zhong Wang 0009, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
ACM Multimedia1