EDBT 2026 Demo / reviewers in the wild / expert
Taoyu Wu
dblp:194/8066
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
0009-0008-7991-6869ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dental3R: Geometry-Aware Pairing for Intraoral 3D Reconstruction from Sparse-View PhotographsabstractDigital orthodontics increasingly depends on accurate 3D dental models, yet conventional intraoral scanning remains inaccessible in remote tele-orthodontics, which typically relies on sparse smartphone imagery. While 3D Gaussian Splatting (3DGS) shows promise for novel view synthesis, its application to the standard clinical triad of unposed anterior and bilateral buccal photographs is challenging. The limitations of sparse-view photometric supervision, combined with large baselines, inconsistent illumination, and specular surfaces, can destabilize simultaneous pose and geometry estimation, often inducing frequency bias and over-smoothed reconstructions that lose critical diagnostic details. To address these issues, we propose Dental3R, a pose-free, graph-guided pipeline for robust, highfidelity reconstruction from sparse intraoral photographs. Our method first constructs a Geometry-Aware Pairing Strategy (GAPS) to select a compact subgraph of high-value image pairs, improving correspondence matching, stabilizing geometry initialization, and reducing memory usage. Leveraging on the recovered poses and point cloud, we train the 3DGS model with a wavelet-regularized objective. By enforcing band-limited fidelity via a discrete wavelet transform, our approach preserves fine enamel boundaries and interproximal edges while suppressing high-frequency artifacts. We validate our approach on a largescale dataset of 950 clinical cases and an additional video-based test set of 195 cases, demonstrating that Dental3R effectively handles sparse, unposed inputs and achieves superior novel-view synthesis quality over state-of-the-art methods. Yiyi Miao, Taoyu Wu, Tong Chen 0005, Ji Jiang, Zhengyong Jiang, Angelos Stefanidis, Limin Yu, Jionglong Su |
BIBM | 2 |
| 2025 | Silhouette-to-Contour Registration: Aligning Intraoral Scan Models with Cephalometric RadiographsabstractReliable 3D-2D alignment between intraoral scan (IOS) models and lateral cephalometric radiographs is critical for orthodontic diagnosis, yet conventional intensity-driven registration methods struggle under real clinical conditions, where cephalograms exhibit magnification, distortion, low-contrast dental crowns, and acquisition-dependent variation. These factors hinder the stability of appearance-based metrics and often lead to convergence failures or anatomically implausible alignments. To address these limitations, we propose DentalSCR, named for its silhouette-to-contour registration scheme, a pose-stable and contour-guided framework for accurate and interpretable alignment that achieves state-of-the-art performance. Our method constructs a U-Midline Dental Axis (UMDA) to establish a unified cross-arch anatomical coordinate system, stabilizing initialization and standardizing projection geometry across cases. Using this reference frame, we generate radiograph-like projections via a surface-based DRR (Digitally Reconstructed Radiograph) formulation with coronal-axis perspective and Gaussian splatting, which preserves clinically accurate magnification and emphasizes external silhouettes. Registration is formulated as a 2D similarity transform optimized with a symmetric bidirectional Chamfer distance under a hierarchical coarse-to-fine schedule, enabling both large capture range and subpixel-level contour agreement. We evaluate DentalSCR on 34 expert-annotated cases. Results demonstrate substantial reductions in landmark error, particularly at posterior teeth, tighter lower-jaw dispersion, and low Chamfer and controlled Hausdorff distances. These findings indicate that DentalSCR robustly handles real-world cephalograms and delivers high-fidelity, clinically inspectable 3D-2D alignment, consistently outperforming baselines and establishing a new state-of-the-art. Yiyi Miao, Taoyu Wu, Ji Jiang, Tong Chen 0005, Zhengyong Jiang, Angelos Stefanidis, Limin Yu, Jionglong Su |
BIBM | 2 |
| 2025 | EndoWave: 4D Gaussian Splatting with Rational Wavelet for Endoscopic ReconstructionabstractIn robot-assisted minimally invasive surgery, accurate 3D reconstruction from endoscopic video is vital for downstream tasks and improved outcomes. However, endoscopic scenarios present unique challenges, including photometric inconsistencies, non-rigid tissue motion, and view-dependent highlights. Most 3DGS-based methods that rely solely on appearance constraints for optimizing 3DGS are often insufficient in this context, as these dynamic visual artifacts can mislead the optimization process and lead to inaccurate reconstructions. To address these limitations, we present EndoWave, a unified spatiotemporal Gaussian Splatting framework by incorporating an optical flow-based geometric constraint and a multi-resolution rational wavelet supervision. First, we adopt a unified spatiotemporal Gaussian representation that directly optimizes primitives in a 4D domain. Second, we propose a geometric constraint derived from optical flow to enhance temporal coherence and effectively constrain the 3D structure of the scene. Third, we propose a multi-resolution rational orthogonal wavelet as a constraint, which can effectively separate the details of the endoscope and enhance the rendering performance. Extensive evaluations on two real surgical datasets, EndoNeRF [1] and StereoMIS [2], demonstrate that our method EndoWave achieves state-of-theart reconstruction quality and visual accuracy compared to the baseline method. Taoyu Wu, Yiyi Miao, Sihang Zhao, Zhuoxiao Li, Baoru Huang, Limin Yu |
BIBM | 1 |
| 2025 | ArchMap: Arch-Flattening and Knowledge-Guided Vision Language Model for Tooth Counting and Structured Dental UnderstandingabstractA structured understanding of intraoral 3D scans is essential for digital orthodontics. However, existing deep-learning approaches rely heavily on modality-specific training, large annotated datasets, and controlled scanning conditions, which limit generalization across devices and hinder deployment in real clinical workflows. Moreover, raw intraoral meshes exhibit substantial variation in arch pose, incomplete geometry caused by occlusion or tooth contact, and a lack of texture cues, making unified semantic interpretation highly challenging. To address these limitations, we propose ArchMap, a training-free and knowledge-guided framework for robust structured dental understanding. ArchMap first introduces a geometry-aware arch-flattening module that standardizes raw 3D meshes into spatially aligned, continuity-preserving multi-view projections. We then construct a Dental Knowledge Base (DKB) encoding hierarchical tooth ontology, dentition-stage policies, and clinical semantics to constrain the symbolic reasoning space. We validate ArchMap on 1060 pre-/post-orthodontic cases, demonstrating robust performance in tooth counting, anatomical partitioning, dentition-stage classification, and the identification of clinical conditions such as crowding, missing teeth, prosthetics, and caries. Compared with supervised pipelines and prompted VLM baselines, ArchMap achieves higher accuracy, reduced semantic drift, and superior stability under sparse or artifact-prone conditions. As a fully training-free system, ArchMap demonstrates that combining geometric normalization with ontology-guided multimodal reasoning offers a practical and scalable solution for the structured analysis of 3D intraoral scans in modern digital orthodontics. Yiyi Miao, Taoyu Wu, Tong Chen 0005, Ji Jiang, Zhuoxiao Li, Limin Yu, Jionglong Su |
IEEE Big Data | 3 |
| 2025 | A Novel Bi-environmental Intuitionistic Fuzzy C-Means Clustering Algorithm
Yihao Zhang 0005, Youpeng Yang, Hao Lan Zhang 0001, Dongming Lu, Taoyu Wu, Xi Yang 0008 |
ICONIP (1) | 5 |
| 2025 | UniBEVFusion: Unified Radar-Vision Bevfusion for 3D Object Detectionabstract4D millimeter-wave (MMW) radar, which provides both height information and dense point cloud data over 3D MMW radar, has become increasingly popular in 3D object detection. In recent years, radar-vision fusion models have demonstrated performance close to that of LiDAR-based models, offering advantages in terms of lower hardware costs and better resilience in extreme conditions. However, many radar-vision fusion models treat radar as a sparse LiDAR, underutilizing radar-specific information. Additionally, these multi-modal networks are often sensitive to the failure of a single modality, particularly vision. To address these challenges, we propose the Radar Depth Lift-Splat-Shoot (RDL) module, which integrates radar-specific data into the depth prediction process, enhancing the quality of visual Bird's-Eye View (BEV) features. We further introduce a Unified Feature Fusion (UFF) approach that extracts BEV features across different modalities using shared module. To assess the robustness of multimodal models, we develop a novel Failure Test (FT) ablation experiment, which simulates vision modality failure by injecting Gaussian noise. We conduct extensive experiments on the View-of-Delft (VoD) and TJ4D datasets. The results demonstrated that our proposed Unified BEVFusion (UniBEVFusion) network significantly outperforms state-of-the-art models on the TJ4D dataset, with improvements of 3.96% in 3D and 4.17% in BEV object detection accuracy. Haocheng Zhao, Runwei Guan, Taoyu Wu, Ka Lok Man, Limin Yu, Yutao Yue |
ICRA | 3 |
| 2025 | EndoFlow-SLAM: Real-Time Endoscopic SLAM with Flow-Constrained Gaussian Splatting
Taoyu Wu, Yiyi Miao, Zhuoxiao Li, Haocheng Zhao, Kang Dang, Jionglong Su, Limin Yu, Haoang Li |
MICCAI (9) | 1 |
| 2025 | Abstractive summarization-based academic paper title drafting
Taoyu Wu, Jiaqi Deng 0001, Kaize Shi |
CCF Trans. Pervasive Comput. Interact. | 1 |
| 2024 | LaGDif: Latent Graph Diffusion Model for Efficient Protein Inverse Folding with Self-EnsembleabstractProtein inverse folding aims to identify viable amino acid sequences that can fold into given protein structures, enabling the design of novel proteins with desired functions for applications in drug discovery, enzyme engineering, and biomaterial development. Diffusion probabilistic models have emerged as a promising approach in inverse folding, offering both feasible and diverse solutions compared to traditional energy-based methods and more recent protein language models. However, existing diffusion models for protein inverse folding operate in discrete data spaces, necessitating prior distributions for transition matrices and limiting smooth transitions and gradients inherent to continuous spaces, leading to suboptimal performance. Drawing inspiration from the success of diffusion models in continuous domains, we introduce the Latent Graph Diffusion Model for Protein Inverse Folding (LaGDif). LaGDif bridges discrete and continuous realms through an encoder-decoder architecture, transforming protein graph data distributions into random noise within a continuous latent space. Our model then reconstructs protein sequences by considering spatial configurations, biochemical attributes, and environmental factors of each node. Additionally, we propose a novel inverse folding self-ensemble method that stabilizes prediction results and further enhances performance by aggregating multiple denoised output protein sequence. Empirical results on the CATH dataset demonstrate that LaGDif outperforms existing state-of-the-art techniques, achieving up to 45.55% improvement in sequence recovery rate for single-chain proteins and maintaining an average RMSD of 1.96 Å between generated and native structures. These advancements of LaGDif in protein inverse folding have the potential to accelerate the development of novel proteins for therapeutic and industrial applications. The code is public available at https://github.com/TaoyuW/LaGDif. Taoyu Wu, Yu Guang Wang 0001, Yiqing Shen 0003 |
BIBM | 1 |
| 2022 | Reversible data hiding in JPEG images based on coefficient-first selection
Xie Yang, Taoyu Wu, Fangjun Huang |
Signal Process. | 2 |