EDBT 2026 Demo / reviewers in the wild / expert
Zhaoshuo Li
dblp:221/4280
· DBLP profile ↗
16ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cox: a reliable runtime switching protocol for BFT consensus algorithmsabstractDifferent BFT (Byzantine Fault Tolerance) consensus algorithms have distinct characteristics and their optimal use cases. Rarely does a protocol demonstrate outstanding performance across different environments. Sometimes, systems need to switch consensus algorithms to adapt to different environments. This paper aims to investigate how to achieve seamless consensus algorithm switching without system downtime and compromising the security and consistency of the original consensus module. To address this issue, we propose Cox, a reliable online switching protocol for BFT consensus algorithms, and design a corresponding recovery sub-protocol for the lagging nodes. The experimental results indicate that, under the assumption of a partially synchronous network, Cox can ensure both the liveness and safety of the consensus module, and achieve reliable switching in milliseconds. Zhaoshuo Li, Weiwei Qiu, Fanglei Huang |
Blockchain Res. Appl. | 1 |
| 2025 | ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image IntermediaryabstractDesigning 3D scenes is traditionally a challenging and laborious task that demands both artistic expertise and proficiency with complex software. Recent advances in text-to-3D generation have greatly simplified this process by letting users create scenes based on simple text descriptions. However, as these methods generally require extra training or in-context learning, their performance is often hindered by the limited availability of high-quality 3D data. In contrast, modern text-to-image models learned from web-scale images can generate scenes with diverse, reliable spatial layouts and consistent, visually appealing styles. Our key insight is that instead of learning directly from 3D scenes, we can leverage generated 2D images as an intermediary to guide 3D synthesis. In light of this, we introduce ArtiScene, a training-free automated pipeline for scene design that integrates the flexibility of free-form text-to-image generation with the diversity and reliability of 2D intermediary layouts. First, we generate 2D images from a scene description, then extract the shape and appearance of objects to create 3D models. These models are assembled into the final scene using geometry, position, and pose information derived from the same intermediary image. Being generalizable to a wide range of scenes and styles, ArtiScene outperforms state-of-the-art benchmarks by a large margin in layout and aesthetic quality by quantitative metrics. It also averages a 74.89 % winning rate in extensive user studies and 95.07 % in GPT-4o evaluation. Zeqi Gu, Yin Cui, Zhaoshuo Li, Fangyin Wei, Yunhao Ge, Jinwei Gu, Ming-Yu Liu 0001, Abe Davis, Yifan Ding 0002 |
CVPR | 3 |
| 2025 | CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsabstractVision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLAs primarily focus on direct input–output mappings, lacking the intermediate reasoning steps crucial for complex manipulation tasks. As a result, existing VLAs lack temporal planning or reasoning capabilities. In this paper, we introduce a method that incorporates explicit visual chain-of-thought (CoT) reasoning into vision-language-action models (VLAs) by predicting future image frames autoregressively as visual goals before generating a short action sequence to achieve these goals. We introduce CoT-VLA, a state-of-the-art 7B VLA that can understand and generate visual and action tokens. Our experimental results demonstrate that CoT-VLA achieves strong performance, outperforming the state-of-the-art VLA model by 17% in real-world manipulation tasks and 6% in simulation benchmarks. Videos are available at: https://cot-vla.github.io/. Yao Lu 0006, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Song Han 0003, Chelsea Finn, Ankur Handa, Tsung-Yi Lin, Gordon Wetzstein, Ming-Yu Liu 0001, Donglai Xiang |
CVPR | 7 |
| 2025 | EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh GenerationabstractCurrent auto-regressive mesh generation methods suffer from issues such as incompleteness, insufficient detail, and poor generalization.
In this paper, we propose an Auto-regressive Auto-encoder (ArAE) model capable of generating high-quality 3D meshes with up to 4,000 faces at a spatial resolution of $512^3$.
We introduce a novel mesh tokenization algorithm that efficiently compresses triangular meshes into 1D token sequences, significantly enhancing training efficiency.
Furthermore, our model compresses variable-length triangular meshes into a fixed-length latent space, enabling training latent diffusion models for better generalization.
Extensive experiments demonstrate the superior quality, diversity, and generalization capabilities of our model in both point cloud and image-conditioned mesh generation tasks. Jiaxiang Tang, Zhaoshuo Li, Zekun Hao, Ming-Yu Liu 0001, Qinsheng Zhang |
ICLR | 2 |
| 2024 | Ada-Tracker: Soft Tissue Tracking via Inter-Frame and Adaptive-template MatchingabstractSoft tissue tracking is crucial for computer-assisted interventions. Existing approaches mainly rely on extracting discriminative features from the template and videos to recover corresponding matches. However, it is difficult to adopt these techniques in surgical scenes, where tissues are changing in shape and appearance throughout the surgery. To address this problem, we exploit optical flow to naturally capture the pixel-wise tissue deformations and adaptively correct the tracked template. Specifically, we first implement an inter-frame matching mechanism to extract a coarse region of interest based on optical flow from consecutive frames. To accommodate appearance change and alleviate drift, we then propose an adaptive-template matching method, which updates the tracked template based on the reliability of the estimates. Our approach, Ada-Tracker, enjoys both short-term dynamics modeling by capturing local deformations and long-term dynamics modeling by introducing global temporal compensation. We evaluate our approach on the public SurgT benchmark, which is generated from Hamlyn, SCARED, and Kidney boundary datasets. The experimental results show that Ada-Tracker achieves superior accuracy and performs more robustly against prior works. Code is available at https://github.com/wrld/Ada-Tracker. Jiangliu Wang, Zhaoshuo Li, Tongyu Jia, Qi Dou 0001, Yun-Hui Liu 0001 |
ICRA | 3 |
| 2024 | SpaceMesh: A Continuous Representation for Learning Manifold Surface MeshesabstractMeshes are ubiquitous in visual computing and simulation, yet most existing machine learning techniques represent meshes only indirectly, e.g. as the level set of a scalar field or deformation of a template, or as a disordered triangle soup lacking local structure. This work presents a scheme to directly generate manifold, polygonal meshes of complex connectivity as the output of a neural network. Our key innovation is to define a continuous latent connectivity space at each mesh vertex, which implies the discrete mesh. In particular, our vertex embeddings generate cyclic neighbor relationships in a halfedge mesh representation, which gives a guarantee of edge-manifoldness and the ability to represent general polygonal meshes. This representation is well-suited to machine learning and stochastic optimization, without restriction on connectivity or topology. We first explore the basic properties of this representation, then use it to fit distributions of meshes from large datasets. The resulting models generate diverse meshes with tessellation structure learned from the dataset population, with concise details and high-quality mesh elements. In applications, this approach not only yields high-quality outputs from generative models, but also enables directly learning challenging geometry processing tasks such as mesh repair. Tianchang Shen, Zhaoshuo Li, Marc T. Law, Matan Atzmon, Sanja Fidler, James Lucas, Jun Gao 0004, Nicholas Sharp |
SIGGRAPH Asia | 2 |
| 2023 | Neuralangelo: High-Fidelity Neural Surface ReconstructionabstractNeural surface reconstruction has been shown to be powerful for recovering dense 3D surfaces via image-based neural rendering. However, current methods struggle to recover detailed structures of real-world scenes. To address the issue, we present Neuralangelo, which combines the representation power of multiresolution 3D hash grids with neural surface rendering. Two key ingredients enable our approach: (1) numerical gradients for computing higher-order derivatives as a smoothing operation and (2) coarse-to-fine optimization on the hash grids controlling different levels of details. Even without auxiliary inputs such as depth, Neuralangelo can effectively recover dense 3D surface structures from multiview images with fidelity significantly surpassing previous methods, enabling detailed large-scale scene reconstruction from RGB video captures. Zhaoshuo Li, Thomas Müller 0013, Alex Evans, Russell H. Taylor, Mathias Unberath, Ming-Yu Liu 0001, Chen-Hsuan Lin 0001 |
CVPR | 1 |
| 2023 | Improving Surgical Situational Awareness with Signed Distance Field: A Pilot Study in Virtual RealityabstractThe introduction of image-guided surgical navigation (IGSN) has greatly benefited technically demanding surgical procedures by providing real-time support and guidance to the surgeon during surgery. To develop effective IGSN, a careful selection of the surgical information and the medium to present this information to the surgeon is needed. However, this is not a trivial task due to the broad array of available options. To address this problem, we have developed an open-source library that facilitates the development of multimodal navigation systems in a wide range of surgical procedures relying on medical imaging data. To provide guidance, our system calculates the minimum distance between the surgical instrument and the anatomy and then presents this information to the user through different mechanisms. The real-time performance of our approach is achieved by calculating Signed Distance Fields at initialization from segmented anatomical volumes. Using this framework, we developed a multimodal surgical navigation system to help surgeons navigate anatomical variability in a skull base surgery simulation environment. Three different feedback modalities were explored: visual, auditory, and haptic. To evaluate the proposed system, a pilot user study was conducted in which four clinicians performed mastoidectomy procedures with and without guidance. Each condition was assessed using objective performance and subjective workload metrics. This pilot user study showed improvements in procedural safety without additional time or workload. These results demonstrate our pipeline's successful use case in the context of mastoidectomy. Hisashi Ishida, Juan Barragan Noguera, Adnan Munawar, Zhaoshuo Li, Andy S. Ding, Peter Kazanzides, Danielle Trakimas, Francis X. Creighton, Russell H. Taylor |
IROS | 4 |
| 2023 | Temporally Consistent Online Depth Estimation in Dynamic ScenesabstractTemporally consistent depth estimation is crucial for online applications such as augmented reality. While stereo depth estimation has received substantial attention as a promising way to generate 3D information, there is relatively little work focused on maintaining temporal stability. Indeed, based on our analysis, current techniques still suffer from poor temporal consistency. Stabilizing depth temporally in dynamic scenes is challenging due to concurrent object and camera motion. In an online setting, this process is further aggravated because only past frames are available. We present a framework named Consistent Online Dynamic Depth (CODD) to produce temporally consistent depth estimates in dynamic scenes in an online setting. CODD augments per-frame stereo networks with novel motion and fusion networks. The motion network accounts for dynamics by predicting a per-pixel SE3 transformation and aligning the observations. The fusion network improves temporal depth consistency by aggregating the current and past estimates. We conduct extensive experiments and demonstrate quantitatively and qualitatively that CODD outperforms competing methods in terms of temporal consistency and performs on par in terms of per-frame accuracy. Zhaoshuo Li, Dilin Wang, Francis X. Creighton, Russell H. Taylor, Ganesh Venkatesh, Mathias Unberath |
WACV | 1 |
| 2022 | Context-Enhanced Stereo Transformer
Weiyu Guo, Zhaoshuo Li, Yongkui Yang, Zheng Wang 0027, Russell H. Taylor, Mathias Unberath, Alan L. Yuille, Yingwei Li 0002 |
ECCV (32) | 2 |
| 2022 | SAGE: SLAM with Appearance and Geometry Prior for Endoscopyabstract., surgical navigation) would benefit from a real-time method that can simultaneously track the endoscope and reconstruct the dense 3D geometry of the observed anatomy from a monocular endoscopic video. To this end, we develop a Simultaneous Localization and Mapping system by combining the learning-based appearance and optimizable geometry priors and factor graph optimization. The appearance and geometry priors are explicitly learned in an end-to-end differentiable training pipeline to master the task of pair-wise image alignment, one of the core components of the SLAM system. In our experiments, the proposed SLAM system is shown to robustly handle the challenges of texture scarceness and illumination variation that are commonly seen in endoscopy. The system generalizes well to unseen endoscopes and subjects and performs favorably compared with a state-of-the-art feature-based SLAM system. The code repository is available at https://github.com/lppllppl920/SAGE-SLAM.git. Xingtong Liu, Zhaoshuo Li, Masaru Ishii, Gregory D. Hager, Russell H. Taylor, Mathias Unberath |
ICRA | 2 |
| 2021 | Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersabstractStereo depth estimation relies on optimal correspondence matching between pixels on epipolar lines in the left and right images to infer depth. In this work, we revisit the problem from a sequence-to-sequence correspondence perspective to replace cost volume construction with dense pixel matching using position information and attention. This approach, named STereo TRansformer (STTR), has several advantages: It 1) relaxes the limitation of a fixed disparity range, 2) identifies occluded regions and provides confidence estimates, and 3) imposes uniqueness constraints during the matching process. We report promising results on both synthetic and real-world datasets and demonstrate that STTR generalizes across different domains, even without fine-tuning. Zhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding, Francis X. Creighton, Russell H. Taylor, Mathias Unberath |
ICCV | 1 |
| 2021 | E-DSSR: Efficient Dynamic Surgical Scene Reconstruction with Transformer-Based Stereoscopic Depth Perception
Yonghao Long 0001, Zhaoshuo Li, Chi Hang Yee, Chi-Fai Ng, Russell H. Taylor, Mathias Unberath, Qi Dou 0001 |
MICCAI (4) | 2 |
| 2021 | An Interpretable Approach to Automated Severity Scoring in Pelvic Trauma
Anna Zapaishchykova, David Dreizin, Zhaoshuo Li, Jie Ying Wu, Shahrooz Faghih Roohi, Mathias Unberath |
MICCAI (3) | 3 |
| 2020 | Anatomical Mesh-Based Virtual Fixtures for Surgical Robots*abstractThis paper presents a dynamic constraint formulation to provide protective virtual fixtures of 3D anatomical structures from polygon mesh representations. The proposed approach can anisotropically limit the tool motion of surgical robots without any assumption of the local anatomical shape close to the tool. Using a bounded search strategy and Principle Directed tree, the proposed system can run efficiently at 180 Hz for a mesh object containing 989,376 triangles and 493,460 vertices. The proposed algorithm has been validated in both simulation and skull cutting experiments. The skull cutting experiment setup uses a novel piezoelectric bone cutting tool designed for the da Vinci research kit. The result shows that the virtual fixture assisted teleoperation has statistically significant improvements in the cutting path accuracy and penetration depth control. The code has been made publicly available at https://github.com/mli0603/PolygonMeshVirtualFixture. Zhaoshuo Li, Alex Gordon, Thomas Looi, James M. Drake, Christopher R. Forrest, Russell H. Taylor |
IROS | 1 |
| 2019 | A Novel Semi-Autonomous Control Framework for Retina Confocal Endomicroscopy Scanning*abstractIn this paper, a novel semi-autonomous control framework is presented for enabling probe-based confocal laser endomicroscopy (pCLE) scan of the retinal tissue. With pCLE, retinal layers such as nerve fiber layer (NFL) and retinal ganglion cell (RGC) can be scanned and characterized in real-time for an improved diagnosis and surgical outcome prediction. However, the limited field of view of the pCLE system and the micron-scale optimal focus distance of the probe, which are in the order of physiological hand tremor, act as barriers to successful manual scan of retinal tissue. Therefore, a novel sensorless framework is proposed for real-time semi-autonomous endomicroscopy scanning during retinal surgery. The framework consists of the Steady-Hand Eye Robot (SHER) integrated with a pCLE system, where the motion of the probe is controlled semi-autonomously. Through a hybrid motion control strategy, the system autonomously controls the confocal probe to optimize the sharpness and quality of the pCLE images, while providing the surgeon with the ability to scan the tissue in a tremor-free manner. Effectiveness of the proposed architecture is validated through experimental evaluations as well as a user study involving 9 participants. It is shown through statistical analyses that the proposed framework can reduce the work load experienced by the users in a statistically-significant manner, while also enhancing their performance in retaining pCLE images with optimized quality. Zhaoshuo Li, Guang-Zhong Yang, Russell H. Taylor, Mahya Shahbazi, Niravkumar A. Patel, Eimear O' Sullivan, Khushi Vyas, Preetham Chalasani, Peter Gehlbach, Iulian Iordachita |
IROS | 1 |