EDBT 2026 Demo / reviewers in the wild / expert
Lizhi Zhao
dblp:164/3840
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GSHOI Denoiser: Denoising Gaussian Hand-Object Interaction for Photorealistic RenderingabstractMany VR/AR applications require the photorealistic rendering of hand-object interactions. Virtual hands are driven by users' hand poses captured via motion tracking to interact with virtual objects. The driven pose can be very noisy due to the constraints of tracking hardware and computation accuracy. This noise may lead to distorted hand poses and penetration artifacts during rendering. In this paper, we introduce the Gaussian Hand-Object Interaction Denoiser, the Gaussian splatting-based hand-object interaction denoising method, which effectively denoises the input twisted and penetrated hand poses to produce photorealistic results. We first propose the innovative joint-to-Gaussian surface representation, which accurately models the spatial relationships between hand skeleton joints and object Gaussians while highlighting hand-object penetrations and generalizing well to new hand poses and objects. Then, we propose a geometry-aware de-penetration algorithm that eliminates penetrations by detecting intersections between skeleton bones and object Gaussians and reposing any penetrated fingers onto the estimated underlying surface of the object. Experiments demonstrate that our method not only effectively reduces hand-object penetration depth but also produces more realistic rendering quality compared to the state-of-the-art methods MANUS+GEARS, MANUS+GeneOH, and$2 \text{DGS}+\text{Gene} \text{OH}$. The user study results show that our method significantly improves the users' visual perceptual experience regarding penetration and stability metrics. Project page: https://github.com/ZhaoLizz/GSHOIDenoiser Lizhi Zhao, Xuequan Lu, Wei Ke 0001, Lili Wang 0006 |
ISMAR | 1 |
| 2025 | Detecting the Abnormal Attitude Variation of Space Target Through Residual Range Migration AnalysisabstractDetecting the abnormal attitude variation (AAV) of noncooperative space targets is one of the most challenging tasks in space situational awareness. This article analyses the residual range migration (RRM), which is the uncompensated residual of the migration through resolution cells (MTRCs), to detect target’s AAV. Unlike the target with known attitude variation, the RRM of the noncooperative target with AAV can be significantly larger than half of the range resolution unit, which stems from incomplete MTRC compensation because of target’s unexpected attitude variation. First, this article analyzes the difference in the RRM between complete and incomplete MTRC compensation, and deduces the mathematical expression of the RRM. Second, the RRM of the space target is estimated by the generalized radon transform (GRT) fitting method, and the upper bound of the fitting error is also deduced. Then, the target’s AAV detection method based on the RRM analysis is designed with a self-adaptive detection threshold. Finally, the simulations analyze the feasibility and effectiveness of the proposed method. Simulation results indicate that the accurately estimated RRM can be utilized for detecting AAV at low signal-to-noise ratio (SNR) and scatterers glinting scenarios, exhibiting a high detection rate even when the target rotates with a small angle. The comparison experiment with existing inverse synthetic aperture radar (ISAR) image-based methods demonstrates the robustness of the proposed method toward the target’s initial attitudes. Hao Yang 0059, Xiongkui Zhang, Junling Wang 0002, Lizhi Zhao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Masked Autoencoders in 3D Point Cloud Representation LearningabstractTransformer-based Self-supervised Representation Learning methods learn generic features from unlabeled datasets for providing useful network initialization parameters for downstream tasks. Recently, methods based upon masking Autoencoders have been explored in the fields. The input can be intuitively masked due to regular content, like sequence words and 2D pixels. However, the extension to 3D point cloud is challenging due to irregularity. In this paper, we propose masked Autoencoders in 3D point cloud representation learning (abbreviated as MAE3D), a novel autoencoding paradigm for self-supervised learning. We first split the input point cloud into patches and mask a portion of them, then use our Patch Embedding Module to extract the features of unmasked patches. Secondly, we employ patch-wise MAE3D Transformers to learn both local features of point cloud patches and high-level contextual relationships between patches, then complete the latent representations of masked patches. We use our Point Cloud Reconstruction Module with multi-task loss to complete the incomplete point cloud as a result. We conduct self-supervised pre-training on ShapeNet55 with the point cloud completion pre-text task and fine-tune the pre-trained model on ModelNet40 and ScanObjectNN (PB_T50_RS, the hardest variant). Comprehensive experiments demonstrate that the local features extracted by our MAE3D from point cloud patches are beneficial for downstream classification tasks, soundly outperforming state-of-the-art methods (93.4% and 86.2% classification accuracy, respectively).Our source codes are available at:https://github.com/Jinec98/MAE3D. Jincen Jiang, Xuequan Lu, Lizhi Zhao, Richard Dazeley, Meili Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Fov-GS: Foveated 3D Gaussian Splatting for Dynamic ScenesabstractRendering quality and performance greatly affect the user's immersion in VR experiences. 3D Gaussian Splatting-based methods can achieve photo-realistic rendering with speeds of over 100 fps in static scenes, but the speed drops below 10 fps in monocular dynamic scenes. Foveated rendering provides a possible solution to accelerate rendering without compromising visual perceptual quality. However, 3DGS and foveated rendering are not compatible. In this paper, we propose Fov-GS, a foveated 3D Gaussian splatting method for rendering dynamic scenes in real time. We introduce a 3D Gaussian forest representation that represents the scene as a forest. To construct the 3D Gaussian forest, we propose a 3D Gaussian forest initialization method based on dynamic-static separation. Subsequently, we propose a 3D Gaussian forest optimization method based on deformation field and Gaussian decomposition to optimize the forest and deformation field. To achieve real-time dynamic scene rendering, we present a 3D Gaussian forest rendering method based on HVS models. Experiments demonstrate that our method not only achieves higher rendering quality in the foveal and salient regions compared to the SOTA methods but also dramatically improves rendering performance, achieving up to 11.33X speedup. We also conducted a user study, and the results prove that the perceptual quality of our method has a high visual similarity with the ground truth. Runze Fan, Jian Wu 0033, Xuehuai Shi, Lizhi Zhao, Qixiang Ma, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | SGSG: Stroke-Guided Scene Graph Generationabstract3D scene graph generation is essential for spatial computing in Extended Reality (XR), providing structured semantics for task planning and intelligent perception. However, unlike instance-segmentation-driven setups, generating semantic scene graphs still suffer from limited accuracy due to coarse and noisy point cloud data typically acquired in practice, and from the lack of interactive strategies to incorporate users' spatialized and intuitive guidance. We identify three key challenges: designing controllable interaction forms, involving guidance in inference, and generalizing from local corrections. To address these, we propose SGSG, a Stroke-Guided Scene Graph generation method that enables users to interactively refine 3D semantic relationships and improve predictions in real time. We propose three types of strokes and a lightweight SGstrokes dataset tailored for this modality. Our model integrates stroke guidance representation and injection for spatio-temporal feature learning and reasoning correction, along with intervention losses that combine consistency-repulsive and geometry-sensitive constraints to enhance accuracy and generalization. Experiments and the user study show that SGSG outperforms state-of-the-art methods 3DSSG and SGFN in overall accuracy and precision, surpasses JointSSG in predicate-level metrics, and reduces task load across all control conditions, establishing SGSG as a new benchmark for interactive 3D scene graph generation and semantic understanding in XR. Implementation resources are available at: https://github.com/Sycamore-Ma/SGSG-runtime. Qixiang Ma, Runze Fan, Lizhi Zhao, Jian Wu 0033, Sio Kei Im, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | GaussianHand: Real-Time 3D Gaussian Rendering for Hand Avatar AnimationabstractRendering animatable and realistic hand avatars is pivotal for enhancing user experiences in human-centered AR/VR applications. While recent initiatives have utilized neural radiance fields to forge hand avatars with lifelike appearances, these methods are often hindered by high computational demands and the necessity for extensive training views. In this paper, we introduce GaussianHand, the first Gaussian-based real-time 3D rendering approach that enables efficient free-view and free-pose hand avatar animation from sparse view images. Our approach encompasses two key innovations. We first propose Hand Gaussian Blend Shapes that effectively models hand surface geometry while ensuring consistent appearance across various poses. Second, we introduce the Neural Residual Skeleton, equipped with Residual Skinning Weights, designed to rectify inaccuracies involved in Linear Blend Skinning deformations due to geometry offsets. Experiments demonstrate that our method not only achieves far more realistic rendering quality with as few as 5 or 20 training views, compared to the 139 views required by existing methods, but also excels in efficiency, achieving up to 125 frames per second for real-time rendering and remarkably surpassing recent methods. Lizhi Zhao, Xuequan Lu, Runze Fan, Sio Kei Im, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | DHGCN: Dynamic Hop Graph Convolution Network for Self-Supervised Point Cloud LearningabstractRecent works attempt to extend Graph Convolution Networks (GCNs) to point clouds for classification and segmentation tasks. These works tend to sample and group points to create smaller point sets locally and mainly focus on extracting local features through GCNs, while ignoring the relationship between point sets. In this paper, we propose the Dynamic Hop Graph Convolution Network (DHGCN) for explicitly learning the contextual relationships between the voxelized point parts, which are treated as graph nodes. Motivated by the intuition that the contextual information between point parts lies in the pairwise adjacent relationship, which can be depicted by the hop distance of the graph quantitatively, we devise a novel self-supervised part-level hop distance reconstruction task and design a novel loss function accordingly to facilitate training. In addition, we propose the Hop Graph Attention (HGA), which takes the learned hop distance as input for producing attention weights to allow edge features to contribute distinctively in aggregation. Eventually, the proposed DHGCN is a plug-and-play module that is compatible with point-based backbone networks. Comprehensive experiments on different backbones and tasks demonstrate that our self-supervised method achieves state-of-the-art performance. Our source codes are available at: https://github.com/Jinec98/DHGCN. Jincen Jiang, Lizhi Zhao, Xuequan Lu, Muhammad Imran Razzak, Meili Wang 0001 |
AAAI | 2 |
| 2024 | PainterAR: A Self-Painting AR Interface for Mobile DevicesabstractABSTRACT Painting is a complex and creative process that involves the use of various drawing skills to create artworks. The concept of training artificial intelligence models to imitate this process is referred to as neural painting. To enable ordinary people to engage in the process of painting, we propose PainterAR, a novel interface that renders any paintings stroke‐by‐stroke in an immersive and realistic augmented reality (AR) environment. PainterAR is composed of two components: the neural painting model and the AR interface. Regarding the neural painting model, unlike previous models, we introduce the Kullback–Leibler divergence to replace the original Wasserstein distance existed in the baseline paint transformer model, which solves an important problem of encountering different scales of strokes (big or small) during painting. We then design an interactive AR interface, which allows users to upload an image and display the creation process of the neural painting model on the virtual drawing board. Experiments demonstrate that the paintings generated by our improved neural painting model are more realistic and vivid than previous neural painting models. The user study demonstrates that users prefer to control the painting process interactively in our AR environment. Yinghan Shi, Lizhi Zhao, Xuequan Lu, Henry Been-Lirn Duh, Meili Wang 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2024 | Estimating the Hotspot Area Observed by Earth Imaging Satellite via ISAR Image SequenceabstractCompared with directly estimating the stable attitude of space target, estimating the hotspot area observed by an Earth imaging satellite using inverse synthetic aperture radar (ISAR) imaging is a more practical and challenging task. In this article, by parameterizing the time-varying attitude of an Earth imaging satellite in spotlight mode into a fixed hotspot area location in the geodetic coordinates, a novel method is proposed. First, the migrations of the scatterers on the rotating component with respect to reference scatterers on the satellite body in the range and Doppler dimensions (MRCRSRD) are analysed, which are used for recognizing the scatterers on the rotating component. Next, it is proven that for the same scatterer on the rotating component, the vector difference sequences of the projection vectors normalized using the range and Doppler coordinates are orthogonal to the position vector of the scatterer. Based on this characteristic, an objective function for estimating the hotspot area is constructed. Then, based on the error principle, expressions of the estimated hotspot area longitude and latitude are derived, along with expressions of the estimation accuracy. Finally, the effects of different factors on the MRCRSRD are analyzed by traversal, under the assumption that the scatterer coordinates obtained by radar are error-free, the feasibility of the proposed method is validated and the estimation error caused by the grid quantization is evaluated, and then the robustness of the proposed method is validated by comparative experiments using parameter search method (PSM) and indirect adjustment method (IAM) under different Gaussian noise conditions. Bo Li 0136, Junling Wang 0002, Tuo Fu, Lizhi Zhao, Shuo Zhang 0033, Peng Lv 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Batch Processing for Enhanced InISAR Imaging of Space TargetsabstractHigh-precision 3-D reconstruction of satellites is critical for automatic target recognition (ATR). This article introduces a batch processing-based method to enhance the performance of interferometric inverse synthetic aperture radar (InISAR) imaging. By utilizing the phase information within the inverse synthetic aperture radar (ISAR) image sequence, the main error components of height estimation are accurately estimated via the ordinary least squares (OLSs) method. In addition, an arc interval determination strategy for optimizing the batch processing performance is proposed. To ensure accurate phase unwrapping in challenging imaging scenarios, a phase unwrapping method based on the minimum standard deviation criterion is also proposed. Simulation experiments validate the effectiveness and superiority of the proposed algorithms. Compared with the traditional two-ISAR-image InISAR imaging method, the batch processing approach improves the precision of height estimation by approximately 50% with an arc length of approximately 15°. Wenshuo Qian, Junling Wang 0002, Lizhi Zhao, Haiguang Li, Fujie Tang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | In-Place Gestures Classification via Long-term Memory Augmented NetworkabstractIn-place gesture-based virtual locomotion techniques enable users to control their viewpoint and intuitively move in the 3D virtual environment. A key research problem is to accurately and quickly recognize in-place gestures, since they can trigger specific movements of virtual viewpoints and enhance user experience. However, to achieve real-time experience, only short-term sensor sequence data (up to about 300ms, 6 to 10 frames) can be taken as input, which actually affects the classification performance due to limited spatiotemporal information. In this paper, we propose a novel long-term memory augmented network for in-place gestures classification. It takes as input both short-term gesture sequence samples and their corresponding long-term sequence samples that provide extra relevant spatio-temporal information in the training phase. We store long-term sequence features with an external memory queue. In addition, we design a memory augmented loss to help cluster features of the same class and push apart features from different classes, thus enabling our memory queue to memorize more relevant long-term sequence features. In the inference phase, we input only short-term sequence samples to recall the stored features accordingly, and fuse them together to predict the gesture class. We create a large-scale in-place gestures dataset from 25 participants with 11 gestures. Our method achieves a promising accuracy of 95.1% with a latency of 192ms, and an accuracy of 97.3% with a latency of 312ms, and is demonstrated to be superior to recent in-place gesture classification techniques. User study also validates our approach. Our source code and dataset will be made available to the community. Lizhi Zhao, Xuequan Lu, Qianyue Bao, Meili Wang 0001 |
ISMAR | 1 |
| 2021 | Classifying In-Place Gestures with End-to-End Point Cloud LearningabstractWalking in place for moving through virtual environments has attracted noticeable attention recently. Recent attempts focused on training a classifier to recognize certain patterns of gestures (e.g., standing, walking, etc) with the use of neural networks like CNN or LSTM. Nevertheless, they often consider very few types of gestures and/or induce less desired latency in virtual environments. In this paper, we propose a novel framework for accurate and efficient classification of in-place gestures. Our key idea is to treat several consecutive frames as a “point cloud”. The HMD and two VIVE trackers provide three points in each frame, with each point consisting of 12-dimensional features (i.e., three-dimensional position coordinates, velocity, rotation, angular velocity). We create a dataset consisting of 9 gesture classes for virtual in-place locomotion. In addition to the supervised point-based network, we also take unsupervised domain adaptation into account due to inter-person variations. To this end, we develop an end-to-end joint framework involving both a supervised loss for supervised point learning and an unsupervised loss for unsupervised domain adaptation. Experiments demonstrate that our approach generates very promising outcomes, in terms of high overall classification accuracy (95.0%) and real-time performance (192ms latency). We will release our dataset and source code to the community. Lizhi Zhao, Xuequan Lu, Meili Wang 0001 |
ISMAR | 1 |
| 2021 | A Novel Improved Reversible Visible Image Watermarking Algorithm Based on Grad-CAM and JNDabstractWith the rapid access convenience of content brought by 5G technology, the integrity protection of content becomes more important. The reversible visible watermarking algorithm has attracted more attention due to its effective content protection. In this paper, a novel improved reversible visible image watermarking scheme based on gradient-weighted class activation mapping (Grad-CAM) and the just noticeable difference (JND) model has been presented. The proposed region of interest (ROI) selection strategy is used to locate the main protected body of images for watermark embedding. Divide the watermark and ROI into nonoverlapping blocks in the same way and then embed the classified two types of watermark blocks into corresponding ROI blocks with the JND model. The optimal bit positions for watermark embedding can be selected adaptively with JND threshold and achieve the tradeoff between the watermark visibility and watermarked image quality. For lossless image recovery and watermark extraction, the recovery information is reversibly hidden into watermarked image. In the experiments, the same process of grayscale images is used to each channel separately for color images watermarking. Besides, there are six aspects in this paper to estimate the proposed scheme; with the comparison to other reversible visible watermarking schemes, experimental results demonstrate the effectiveness of our proposed scheme. Jiasheng Qu, Xiangchun Liu, Lizhi Zhao |
Secur. Commun. Networks | 4 |