EDBT 2026 Demo / reviewers in the wild / expert
Zhuoxiao Li
dblp:312/7168
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learnable mixture distribution prior for image denoising
Zhuoxiao Li, Jun Liu 0029 |
Knowl. Based Syst. | 1 |
| 2025 | Adversarial Attacks on Event-Based Pedestrian Detectors: A Physical ApproachabstractEvent cameras, known for their low latency and high dynamic range, show great potential in pedestrian detection applications. However, while recent research has primarily focused on improving detection accuracy, the robustness of event-based visual models against physical adversarial attacks has received limited attention. For example, adversarial physical objects, such as specific clothing patterns or accessories, can exploit inherent vulnerabilities in these systems, leading to misdetections or misclassifications. This study is the first to explore physical adversarial attacks on event-driven pedestrian detectors, specifically investigating whether certain clothing patterns worn by pedestrians can cause these detectors to fail, effectively rendering them unable to detect the person. To address this, we developed an end-to-end adversarial framework in the digital domain, framing the design of adversarial clothing textures as a 2D texture optimization problem. By crafting an effective adversarial loss function, the framework iteratively generates optimal textures through backpropagation. Our results demonstrate that the textures identified in the digital domain possess strong adversarial properties. Furthermore, we translated these digitally optimized textures into physical clothing and tested them in real-world scenarios, successfully demonstrating that the designed textures significantly degrade the performance of event-based pedestrian detection models. This work highlights the vulnerability of such models to physical adversarial attacks. Guixu Lin, Muyao Niu, Qingtian Zhu, Zhengwei Yin, Zhuoxiao Li, Shengfeng He, Yinqiang Zheng |
AAAI | 5 |
| 2025 | EndoWave: 4D Gaussian Splatting with Rational Wavelet for Endoscopic ReconstructionabstractIn robot-assisted minimally invasive surgery, accurate 3D reconstruction from endoscopic video is vital for downstream tasks and improved outcomes. However, endoscopic scenarios present unique challenges, including photometric inconsistencies, non-rigid tissue motion, and view-dependent highlights. Most 3DGS-based methods that rely solely on appearance constraints for optimizing 3DGS are often insufficient in this context, as these dynamic visual artifacts can mislead the optimization process and lead to inaccurate reconstructions. To address these limitations, we present EndoWave, a unified spatiotemporal Gaussian Splatting framework by incorporating an optical flow-based geometric constraint and a multi-resolution rational wavelet supervision. First, we adopt a unified spatiotemporal Gaussian representation that directly optimizes primitives in a 4D domain. Second, we propose a geometric constraint derived from optical flow to enhance temporal coherence and effectively constrain the 3D structure of the scene. Third, we propose a multi-resolution rational orthogonal wavelet as a constraint, which can effectively separate the details of the endoscope and enhance the rendering performance. Extensive evaluations on two real surgical datasets, EndoNeRF [1] and StereoMIS [2], demonstrate that our method EndoWave achieves state-of-theart reconstruction quality and visual accuracy compared to the baseline method. Taoyu Wu, Yiyi Miao, Sihang Zhao, Zhuoxiao Li, Baoru Huang, Limin Yu |
BIBM | 6 |
| 2025 | ArchMap: Arch-Flattening and Knowledge-Guided Vision Language Model for Tooth Counting and Structured Dental UnderstandingabstractA structured understanding of intraoral 3D scans is essential for digital orthodontics. However, existing deep-learning approaches rely heavily on modality-specific training, large annotated datasets, and controlled scanning conditions, which limit generalization across devices and hinder deployment in real clinical workflows. Moreover, raw intraoral meshes exhibit substantial variation in arch pose, incomplete geometry caused by occlusion or tooth contact, and a lack of texture cues, making unified semantic interpretation highly challenging. To address these limitations, we propose ArchMap, a training-free and knowledge-guided framework for robust structured dental understanding. ArchMap first introduces a geometry-aware arch-flattening module that standardizes raw 3D meshes into spatially aligned, continuity-preserving multi-view projections. We then construct a Dental Knowledge Base (DKB) encoding hierarchical tooth ontology, dentition-stage policies, and clinical semantics to constrain the symbolic reasoning space. We validate ArchMap on 1060 pre-/post-orthodontic cases, demonstrating robust performance in tooth counting, anatomical partitioning, dentition-stage classification, and the identification of clinical conditions such as crowding, missing teeth, prosthetics, and caries. Compared with supervised pipelines and prompted VLM baselines, ArchMap achieves higher accuracy, reduced semantic drift, and superior stability under sparse or artifact-prone conditions. As a fully training-free system, ArchMap demonstrates that combining geometric normalization with ontology-guided multimodal reasoning offers a practical and scalable solution for the structured analysis of 3D intraoral scans in modern digital orthodontics. Yiyi Miao, Taoyu Wu, Tong Chen 0005, Ji Jiang, Zhuoxiao Li, Limin Yu, Jionglong Su |
IEEE Big Data | 6 |
| 2025 | EndoFlow-SLAM: Real-Time Endoscopic SLAM with Flow-Constrained Gaussian Splatting
Taoyu Wu, Yiyi Miao, Zhuoxiao Li, Haocheng Zhao, Kang Dang, Jionglong Su, Limin Yu, Haoang Li |
MICCAI (9) | 3 |
| 2024 | Fooling Polarization-Based Vision Using Locally Controllable Polarizing ProjectionabstractPolarization is a fundamental property of light that encodes abundant information regarding surface shape, material, illumination and viewing geometry. The computer vision community has witnessed a blossom of polarization-based vision applications, such as reflection removal, shape-from-polarization (SfP), transparent object segmentation and color constancy, partially due to the emergence of single-chip mono/color polarization sensors that make polarization data acquisition easier than ever. However, is polarization-based vision vulnerable to adversarial attacks? If so, is that possible to realize these adversarial attacks in the physical world, without being perceived by human eyes? In this paper, we warn the community of the vulnerability of polarization-based vision, which can be more serious than RGB-based vision. By adapting a commercial LCD projector, we achieve locally controllable polarizing projection, which is successfully utilized to fool state-of-the-art polarization-based vision algorithms for glass segmentation and SfP. Compared with existing physical attacks on RGB-based vision, which always suffer from the trade-off between attack efficacy and eye conceivability, the adversarial attackers based on polarizing projection are contact-free and visually imperceptible, since naked human eyes can rarely perceive the difference of viciously manipulated polarizing light and ordinary illumination. This poses unprecedented risks on polarization-based vision, for which due attentions should be paid and counter measures be considered. Zhuoxiao Li, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
CVPR | 1 |
| 2024 | RS-NeRF: Neural Radiance Fields from Rolling Shutter Images
Muyao Niu, Yifan Zhan, Zhuoxiao Li, Xiang Ji 0005, Yinqiang Zheng |
ECCV (46) | 4 |
| 2024 | KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter
Yifan Zhan, Zhuoxiao Li, Muyao Niu, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
ECCV (45) | 2 |
| 2023 | Polarized Color Image DenoisingabstractSingle-chip polarized color photography provides both visual textures and object surface information in one snapshot. However, the use of an additional directional polarizing filter array tends to lower photon count and SNR, when compared to conventional color imaging. As a result, such a bilayer structure usually leads to unpleasant noisy images and undermines performance of polarization analysis, especially in low-light conditions. It is a challenge for traditional image processing pipelines owing to the fact that the physical constraints exerted implicitly in the channels are excessively complicated. In this paper, we propose to tackle this issue through a noise modeling method for realistic data synthesis and a powerful network structure inspired by vision Transformer. A real-world polarized color image dataset of paired raw short-exposed noisy images and long-exposed reference images is captured for experimental evaluation, which has demonstrated the effectiveness of our approaches for data synthesis and polarized color image denoising. The code and data can be found at https://github.com/bandasyou/pcdenoise. Zhuoxiao Li, Haiyang Jiang 0002, Mingdeng Cao, Yinqiang Zheng |
CVPR | 1 |
| 2023 | Visibility Constrained Wide-Band Illumination Spectrum Design for Seeing-in-the-DarkabstractSeeing-in-the-dark is one of the most important and challenging computer vision tasks due to its wide applications and extreme complexities of in-the- wild scenarios. Existing arts can be mainly divided into two threads: 1) RGB-dependent methods restore information using degraded RGB inputs only (e.g., low-light enhancement), 2) RGB-independent methods translate images captured under auxiliary near-infrared (NIR) illuminants into RGB domain (e.g., NIR2RGB translation). The latter is very attractive since it works in complete darkness and the illuminants are visually friendly to naked eyes, but tends to be unstable due to its intrinsic ambiguities. In this paper, we try to robustify NIR2RGB translation by designing the optimal spectrum of auxiliary illumination in the wide-band VIS-NIR range, while keeping visual friendliness. Our core idea is to quantify the visibility constraint implied by the human vision system and incorporate it into the design pipeline. By modeling the formation process of images in the VIS-NIR range, the optimal multiplexing of a wide range of LEDs is automatically designed in a fully differentiable manner, within the feasible region defined by the visibility constraint. We also collect a substantially expanded VIS-NIR hyperspectral image dataset for experiments by using a customized 50-band filter wheel. Experimental results show that the task can be significantly improved by using the optimized wide-band illumination than using NIR only. Codes Available: https://github.com/MyNiuuu/VCSD. Muyao Niu, Zhuoxiao Li, Zhihang Zhong, Yinqiang Zheng |
CVPR | 2 |
| 2023 | Physics-Based Adversarial Attack on Near-Infrared Human Detector for Nighttime Surveillance Camera SystemsabstractMany surveillance cameras switch between daytime and nighttime modes based on illuminance levels. During the day, the camera records ordinary RGB images through an enabled IR-cut filter. At night, the filter is disabled to capture near-infrared (NIR) light emitted from NIR LEDs typically mounted around the lens. While the vulnerabilities of RGB-based AI algorithms have been widely reported, those of NIR-based AI have rarely been investigated. In this paper, we identify fundamental vulnerabilities in NIR-based image understanding caused by color and texture loss due to the intrinsic characteristics of clothes' reflectance and cameras' spectral sensitivity in the NIR range. We further show that the nearly co-located configuration of illuminants and cameras in existing surveillance systems facilitates concealing and fully passive attacks in the physical world. Specifically, we demonstrate how retro-reflective and insulation plastic tapes can manipulate the intensity distribution of NIR images. We showcase an attack on the YOLO-based human detector using binary patterns designed in the digital space (via black-box query and searching) and then physically realized using tapes pasted onto clothes. Our attack highlights significant reliability concerns about nighttime surveillance systems, which are intended to enhance security. Codes Available: https://github.com/MyNiuuu/AdvNIR. Muyao Niu, Zhuoxiao Li, Yifan Zhan, Huy H. Nguyen, Isao Echizen, Yinqiang Zheng |
ACM Multimedia | 2 |
| 2022 | Target Oriented Perceptual Adversarial Fusion Network for Underwater Image EnhancementabstractDue to the refraction and absorption of light by water, underwater images usually suffer from severe degradation, such as color cast, hazy blur, and low visibility, which would degrade the effectiveness of marine applications equipped on autonomous underwater vehicles. To eliminate the degradation of underwater images, we propose a target oriented perceptual adversarial fusion network, dubbed TOPAL. Concretely, we consider the degradation factors of underwater images in terms of turbidity and chromatism. And according to the degradation issues, we first develop a multi-scale dense boosted module to strengthen the visual contrast and a deep aesthetic render module to perform the color correction, respectively. After that, we employ the dual channel-wise attention module and guide the adaptive fusion of latent features, in which both diverse details and credible appearance are integrated. To bridge the gap between synthetic and real-world images, a global-local adversarial mechanism is introduced in the reconstruction. Besides, perceptual information is also embedded into the process to assist the understanding of scenery content. To evaluate the performance of TOPAL, we conduct extensive experiments on several benchmarks and make comparisons among state-of-the-art methods. Quantitative and qualitative results demonstrate that our TOPAL improves the quality of underwater images greatly and achieves superior performance than others. Zhiying Jiang, Zhuoxiao Li, Shuzhou Yang, Xin Fan 0001, Risheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Dual Mixture Model Based CNN for Image DenoisingabstractNon-Gaussian residual error and noise are common in the real applications, and they can be efficiently addressed by some non-quadratic fidelity terms in the classic variational method. However, they have not been well integrated into the architectures design in the convolutional neural networks (CNN) based image denoising method. In this paper, we propose a deep learning approach to handle non-Gaussian residual error. Our method is developed on an universal approximation property for the probability density functions of the non-Gaussian error/noise. By considering the duality of the maximum likelihood estimation for the non-Gaussian error, an adaptive weighting strategy can be derived for image fidelity. To get a good image prior, a learnable regularizer is adopted. Solving such a problem iteratively can be unrolled as a weighted residual CNN architecture. The main advantage of our method is that the weighted residual block can well handle the non-Gaussian residual, especially for the noise with non-uniformly spatial distribution. Numerical results show that it has better performance on non-Gaussian noise (e.g. Gaussian mixture, random-valued impulse noise) removal than the related existing methods. Zhuoxiao Li, Jun Liu 0029 |
IEEE Trans. Image Process. | 1 |
| 2021 | Multiple Task-Oriented Encoders for Unified Image FusionabstractImage fusion methods have achieved incredible progress, but they are vulnerable to handling a certain type of fusion task rather than considering deeper relations between cross-realm task correlations. To achieve this, we integrate different image fusion tasks into a unified network. Our method is accomplished through multiple task-oriented encoders and a generic decoder, in addition to a self-adapting loss function. The taskoriented encoders are trained to learn task-specific features, while the generic decoder reconstructs the fused features to generate a comprehensive image. Subsequently, by introducing the self-adapting loss in our method, it can automatically adjust itself to source data characteristics on different tasks. Besides, we formulate a training strategy based on bilevel optimization to update the multi-encoder and generic decoder in an alternative manner. Extensive experimental results demonstrate the superior performance of our method over the stateof-the-art methods. Zhuoxiao Li, Jinyuan Liu 0001, Risheng Liu, Xin Fan 0001, Zhongxuan Luo, Wen Gao 0001 |
ICME | 1 |