Yuzhe Luo

dblp:248/7472 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Lightmap Compression with Color-Coherent UV Clustering and Cascade Texture Optimization
abstract
Abstract To address the storage overhead of lightmaps and the limitations of existing compression techniques, we propose a novel UV‐space compression framework based on per‐triangle processing. By mapping triangles to a standardized domain, we cluster and repack color‐coherent regions into a compact atlas, generating a cascade texture refined via differentiable rendering. Experimental results show an average storage reduction of 83% with approximately 10 dB higher PSNR than existing methods. Our approach is the first dedicated lightmap compression framework compatible with standard block‐based formats, offering an effective solution for memory‐efficient 3D asset delivery.
Dehan Chen, Hongyu Huang 0001, Yuzhe Luo, Hao Xu 0049, Yuqing Zhang 0005, Sipeng Yang, Xifeng Gao, Heng Cai, Xiaogang Jin 0001
Comput. Graph. Forum3
2026 Feature-Preserving Offset Meshing
abstract
We introduce a new offset meshing method that handles clean 3D surface meshes of arbitrary geometry and topology—where “clean” refers to meshes that are watertight, manifold, and free of self-intersections. Our approach also extends to imperfect, or “dirty,” meshes that violate these conditions, although the problem becomes significantly more difficult in such scenarios, and faithful feature preservation near defective areas cannot always be assured. In contrast to prior techniques, which have largely focused on constant-radius offsets, our method is, to our knowledge, the first to support mitered offsets while effectively preserving sharp features. Our method is designed based on several core principles: (1) explicitly generating the offset vertices and triangles with feature-capturing energy and constraints; (2) prioritizing the generation of the offset geometry before establishing its connectivity, (3) employing exact algorithms in critical pipeline steps for robustness, balancing the use of floating-point computations for efficiency, (4) applying various conservative speed up strategies including early reject non-contributing computations to the final output. Our approach further uniquely supports variable offset distances on input surface elements, offering a wider range of practical applications compared to conventional methods. For benchmarking purposes, we performed an extensive comparison against state-of-the-art offset methods using a curated subset of the Thingi10K dataset. Our results demonstrate the superiority of our approach over current state-of-the-art methods in terms of element count, feature preservation, and non-uniform offset distances of the resulting offset mesh surfaces, marking a significant advancement in the field.
Hongyi Cao, Gang Xu 0001, Renshu Gu, Jinlan Xu, Timon Rabczuk, Yuzhe Luo, Xifeng Gao
ACM Trans. Graph.7
2025 RL-ACD: Reinforcement Learning-based Approximate Convex Decomposition
abstract
Approximate Convex Decomposition (ACD) aims to approximate complex 3D shapes with convex components, which is widely applied to create compact collision representations for real-time applications, including VR/AR, interactive games, and robotic simulations. Efficiency and optimality are critical for ACD algorithms in approximating large-scale, complex 3D shapes, enabling high-quality decompositions with minimal components. Unfortunately, existing methods either employ sub-optimal greedy strategies or rely on computationally intensive multi-step searches. In this work, we propose RL-ACD, a data-driven, reinforcement learning-based approach for efficient and near-optimal convex shape decomposition. We formulate ACD as a Markov Decision Process (MDP), where cutting planes are iteratively applied based on the current stage's mesh fragments rather than the entire fine-grained mesh, leading to a novel, efficient geometric encoding. To train near-optimal policies for ACD, we propose a novel dual-state Bellman loss and analyze its convergence using a Q-learning algorithm. Comprehensive evaluations across diverse datasets validate the efficiency and accuracy of RL-ACD for convex decomposition tasks. Our method outperforms the multi-step tree search by 15× in terms of computational speed, while reducing the number of resulting components by 16% compared to the current state-of-the-art greedy algorithms, significantly narrowing the sub-optimality gap and enhancing downstream task performance.
Yuzhe Luo, Zherong Pan, Kui Wu 0003, Xingyi Du, Xiangjun Tang, Xiaogang Jin 0001, Xifeng Gao
ACM Trans. Graph.1
2025 Chorus: Robust Multitasking Local Client-Server Collaborative Inference With Wi-Fi 6 for AIoT Against Stochastic Congestion Delay
abstract
The rapid growth of AIoT devices brings huge demands for DNNs deployed on resource-constrained devices. However, the intensive computation and high memory footprint of DNN inference make it difficult for the AIoT devices to execute the inference tasks efficiently. In many widely deployed AIoT use cases, multiple local AIoT devices launch DNN inference tasks randomly. Although local collaborative inference has been proposed to accelerate DNN inference on local devices with limited resources, multitasking local collaborative inference, which is common in AIoT scenarios, has not been fully studied in previous works. We consider multitasking local client-server collaborative inference (MLCCI), which achieves efficient DNN inference by offloading the inference tasks from multiple AIoT devices to a more powerful local server with parallel pipelined execution streams through Wi-Fi 6. Our optimization goal is to minimize the mean end-to-end latency of MLCCI. Based on the experiment results, we identify three key challenges: high communication costs, high model initialization latency, and congestion delay brought by task interference. We analyze congestion delay in MLCCI and its stochastic fluctuations with queuing theory and propose Chorus, a high-performance adaptive MLCCI framework for AIoT devices, to minimize the mean end-to-end latency of MLCCI against stochastic congestion delay. Chorus generates communication-efficient model partitions with heuristic search, uses a prefetch-enabled two-level LRU cache to accelerate model initialization on the server, reduces congestion delay and its short-term fluctuations with execution stream allocation based on the cross-entropy method, and finally achieves efficient computation offloading with reinforcement learning. We established a system prototype, which statistically simulated many virtual clients with limited physical client devices to conduct performance evaluations, for Chorus with real devices. The evaluation results for various workload levels show that Chorus achieved an average of$1.4\times$,$1.3\times$, and$2\times$speedup over client-only inference, and server-only inference with LRU and MLSH, respectively.
Yuzhe Luo, Ji Qi 0002, Ling Li 0001, Ruizhi Chen, Limin Cheng
IEEE Trans. Parallel Distributed Syst.1
2024 Privacy-preserving Compression for Efficient Collaborative Inference
abstract
Collaborative inference accelerates DNN inference tasks of resource-limited devices (e.g., clients) by offloading model slices to resource-rich devices (e.g., servers). During the inference procedure, outputs of model slices are transmitted among devices, causing significant intermediate data transmission overhead and posing a risk of privacy leakage of the client’s input data. Quantization has been widely used in collaborative inference to enhance communication efficiency. However, traditional quantization cannot prevent data privacy leaks. Besides, perturbation-based privacy protection methods, such as adding Laplace noise to the intermediate data of collaborative inference, do not consider communication efficiency. In this paper, we introduce Layered Laplace Random Quantization to simultaneously achieve communication efficiency and data privacy protection in collaborative inference by compressing the intermediate data with Laplace quantization noise. We also propose stability training to recover the accuracy loss caused by our method. Evaluation results show that our method achieved an average inference latency speedup of 1.2x-1.3x for different DNN models compared with the baseline methods while achieving comparable data privacy protection and recoverable accuracy loss.
Yuzhe Luo, Ji Qi 0002, Jiageng Yu, Ruizhi Chen, Ke Gao 0012, Ling Li 0001
ICPADS1
2024 MNSS: Neural Supersampling Framework for Real-Time Rendering on Mobile Devices
abstract
Although neural supersampling has achieved great success in various applications for improving image quality, it is still difficult to apply it to a wide range of real-time rendering applications due to the high computational power demand. Most existing methods are computationally expensive and require high-performance hardware, preventing their use on platforms with limited hardware, such as smartphones. To this end, we propose a new supersampling framework for real-time rendering applications to reconstruct a high-quality image out of a low-resolution one, which is sufficiently lightweight to run on smartphones within a real-time budget. Our model takes as input the renderer-generated low resolution content and produces high resolution and anti-aliased results. To maximize sampling efficiency, we propose using an alternate sub-pixel sample pattern during the rasterization process. This allows us to create a relatively small reconstruction model while maintaining high image quality. By accumulating new samples into a high-resolution history buffer, an efficient history check and re-usage scheme is introduced to improve temporal stability. To our knowledge, this is the first research in pushing real-time neural supersampling on mobile devices. Due to the absence of training data, we present a new dataset containing 57 training and test sequences from three game scenes. Furthermore, based on the rendered motion vectors and a visual perception study, we introduce a new metric called inter-frame structural similarity (IF-SSIM) to quantitatively measure the temporal stability of rendered videos. Extensive evaluations demonstrate that our supersampling model outperforms existing or alternative solutions in both performance and temporal stability.
Sipeng Yang, Yunlu Zhao, Yuzhe Luo, He Wang 0002, Hongyu Sun 0001, Chen Li 0062, Binghuang Cai, Xiaogang Jin 0001
IEEE Trans. Vis. Comput. Graph.3
2023 Texture Atlas Compression Based on Repeated Content Removal
abstract
Optimizing the memory footprint of 3D models can have a major impact on the user experiences during real-time rendering and streaming visualization, where the major memory overhead lies in the high-resolution texture data. In this work, we propose a robust and automatic pipeline to content-aware, lossy compression for texture atlas. The design of our solution lies in two observations: 1) mapping multiple surface patches to the same texture region is seamlessly compatible with the standard rendering pipeline, requiring no decompression before any usage; 2) a texture image has background regions and salient structural features, which can be handled separately to achieve a high compression rate. Accordingly, our method contains joint operations of image segmentation, re-meshing, UV unwrapping, and texture baking. To evaluate the efficacy of our approach, we batch-processed a dataset containing 100 models collected online. On average, our method achieves a texture atlas compression ratio of 81.41% with an averaged PSNR and MS-SSIM scores of 40.90 and 0.98, a marginal error in visual appearance.
Yuzhe Luo, Xiaogang Jin 0001, Zherong Pan, Kui Wu 0003, Qilong Kou, Xiajun Yang, Xifeng Gao
SIGGRAPH Asia1
2022 GlyphCreator: Towards Example-based Automatic Generation of Circular Glyphs
abstract
Circular glyphs are used across disparate fields to represent multidimensional data. However, although these glyphs are extremely effective, creating them is often laborious, even for those with professional design skills. This paper presents GlyphCreator, an interactive tool for the example-based generation of circular glyphs. Given an example circular glyph and multidimensional input data, GlyphCreator promptly generates a list of design candidates, any of which can be edited to satisfy the requirements of a particular representation. To develop GlyphCreator, we first derive a design space of circular glyphs by summarizing relationships between different visual elements. With this design space, we build a circular glyph dataset and develop a deep learning model for glyph parsing. The model can deconstruct a circular glyph bitmap into a series of visual elements. Next, we introduce an interface that helps users bind the input data attributes to visual elements and customize visual styles. We evaluate the parsing model through a quantitative experiment, demonstrate the use of GlyphCreator through two use scenarios, and validate its effectiveness through user interviews.
Lu Ying, Tan Tang, Yuzhe Luo, Lvkeshen Shen, Xiao Xie, Lingyun Yu 0001, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.3