VLDB 2026 Research / reviewers in the wild / expert
Tao Tu 0002
dblp:195/9507-2
· DBLP profile ↗
8ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-9191-7938ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OpenM3D: Open Vocabulary Multi-View Indoor 3D Object Detection without Human Annotations
Peng-Hao Hsu, Ke Zhang 0028, Fu-En Wang, Tao Tu 0002, Ming-Feng Li, Yu-Lun Liu 0001, Albert Chen 0001, Min Sun 0001, Cheng-Hao Kuo |
ICCV | 4 |
| 2025 | DreaMo: Articulated 3D Reconstruction from a Single Casual VideoabstractArticulated 3D reconstruction has valuable applications in various domains, yet it remains costly and demands intensive work from domain experts. Recent advancements in template-free learning methods show promising results with monocular videos. Nevertheless, these approaches necessitate a comprehensive coverage of all viewpoints of the subject in the input video, thus limiting their applicability to casually captured videos from online sources. In this work, we study articulated 3D shape reconstruction from a single and casually captured Internet video, where the subject's view coverage is incomplete. We propose DreaMo that jointly performs shape reconstruction while solving the challenging low-coverage regions with view-conditioned diffusion prior and several tailored regularizations. In addition, we introduce a skeleton generation strategy to create human-interpretable skeletons from the learned neural bones and skinning weights without any predefined skeleton structures. We conduct our study on a self-collected internet video collection characterized by incomplete view coverage. DreaMo shows promising quality in novel-view rendering, detailed articulated shape reconstruction, and skeleton generation. Extensive qualitative and quantitative studies validate the efficacy of each proposed component, and show existing methods are unable to solve correct geometry due to the incomplete view coverage. Tao Tu 0002, Ming-Feng Li, Chieh Hubert Lin, Yen-Chi Cheng, Min Sun 0001, Ming-Hsuan Yang 0001 |
WACV | 1 |
| 2025 | V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D AnnotationsabstractThe field of indoor monocular 3D object detection is gaining significant attention, fueled by the increasing demand in VR/AR and robotic applications. However, its advancement is impeded by the limited availability and diversity of 3D training data, owing to the labor-intensive nature of 3D data collection and annotation processes. In this paper, we present V-MIND (Versatile Monocular INdoor Detector), which enhances the performance of in-door 3D detectors across a diverse set of object classes by harnessing publicly available large-scale 2D datasets. By leveraging well-established monocular depth estimation techniques and camera intrinsic predictors, we can generate 3D training data by converting large-scale 2D images into 3D point clouds and subsequently deriving pseudo 3D bounding boxes. To mitigate distance errors inherent in the converted point clouds, we introduce a novel 3D self-calibration loss for refining the pseudo 3D bounding boxes during training. Additionally, we propose a novel ambiguity loss to address the ambiguity that arises when introducing new classes from 2D datasets. Finally, through joint training with existing 3D datasets and pseudo 3D bounding boxes derived from 2D datasets, V-MIND achieves state-of-the-art object detection performance across a wide range of classes on the Omni3D indoor dataset. Jin-Cheng Jhang, Tao Tu 0002, Fu-En Wang, Ke Zhang 0028, Min Sun 0001, Cheng-Hao Kuo |
WACV | 2 |
| 2023 | ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object DetectionabstractWe propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-view images to alleviate the confusion arising from voxels of free space, and during the inference phase, only images from multiple views are required. Besides, a powerful pre-trained 2D feature extractor can be leveraged by our representation, leading to a more robust performance. To evaluate the effectiveness of ImGeoNet, we conduct quantitative and qualitative experiments on three indoor datasets, namely ARKitScenes, ScanNetV2, and ScanNet200. The results demonstrate that ImGeoNet outperforms the current state-of-the-art multiview image-based method, ImVoxelNet, on all three datasets in terms of detection accuracy. In addition, ImGeoNet shows great data efficiency by achieving results comparable to ImVoxelNet with 100 views while utilizing only 40 views. Furthermore, our studies indicate that our proposed image-induced geometry-aware representation can enable image-based methods to attain superior detection accuracy than the seminal point cloud-based method, VoteNet, in two practical scenarios: (1) scenarios where point clouds are sparse and noisy, such as in ARKitScenes, and (2) scenarios involve diverse object classes, particularly classes of small objects, as in the case in ScanNet200. Project page: https://ttaoretw.github.io/imgeonet. Tao Tu 0002, Shun-Po Chuang, Yu-Lun Liu 0001, Cheng Sun 0004, Ke Zhang 0028, Donna Roy, Cheng-Hao Kuo, Min Sun 0001 |
ICCV | 1 |
| 2021 | Learning Better Visual Dialog Agents With Pretrained Visual-Linguistic RepresentationabstractGuessWhat?! is a visual dialog guessing game which incorporates a Questioner agent that generates a sequence of questions, while an Oracle agent answers the respective questions about a target object in an image. Based on this dialog history between the Questioner and the Oracle, a Guesser agent makes a final guess of the target object. While previous work has focused on dialogue policy optimization and visual-linguistic information fusion, most work learns the vision-linguistic encoding for the three agents solely on the GuessWhat?! dataset without shared and prior knowledge of vision-linguistic representation. To bridge these gaps, this paper proposes new Oracle, Guesser and Questioner models that take advantage of a pretrained vision-linguistic model, VilBERT. For Oracle model, we introduce a two-way background/target fusion mechanism to understand both intra and inter-object questions. For Guesser model, we introduce a state-estimator that best utilizes VilBERT’s strength in single-turn referring expression comprehension. For the Questioner, we share the state-estimator from pretrained Guesser with Questioner to guide the question generator. Experimental results show that our proposed models outperform state-of-the-art models significantly by 7%, 10%, 12% for Oracle, Guesser and End-to-End Questioner respectively. Tao Tu 0002, Qing Ping, Govindarajan Thattai, Gökhan Tür, Premkumar Natarajan |
CVPR | 1 |
| 2020 | Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation LearningabstractIn this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very close to phoneme sequences of speech utterances. This is achieved by proper temporal segmentation to make the representations phoneme-synchronized, and proper phonetic clustering to have total number of distinct representations close to the number of phonemes. Mapping between the distinct representations and phonemes is learned from a small amount of annotated paired data. Preliminary experiments on LJSpeech demonstrated the learned representations for vowels have relative locations in latent space in good parallel to that shown in the IPA vowel chart defined by linguistics experts. With less than 20 minutes of annotated speech, our method outperformed existing methods on phoneme recognition and is able to synthesize intelligible speech that beats our baseline model. Alexander H. Liu, Tao Tu 0002, Hung-yi Lee, Lin-Shan Lee |
ICASSP | 2 |
| 2020 | Semi-Supervised Learning for Multi-Speaker Text-to-Speech Synthesis Using Discrete Speech RepresentationabstractRecently, end-to-end multi-speaker text-to-speech (TTS) systems gain success in the situation where a lot of high-quality speech plus their corresponding transcriptions are available.However, laborious paired data collection processes prevent many institutes from building multi-speaker TTS systems of great performance.In this work, we propose a semi-supervised learning approach for multi-speaker TTS.A multi-speaker TTS model can learn from the untranscribed audio via the proposed encoder-decoder framework with discrete speech representation.The experiment results demonstrate that with only an hour of paired speech data, whether the paired data is from multiple speakers or a single speaker, the proposed model can generate intelligible speech in different voices.We found the model can benefit from the proposed semi-supervised learning approach even when part of the unpaired speech data is noisy.In addition, our analysis reveals that different speaker characteristics of the paired data have an impact on the effectiveness of semisupervised TTS. Tao Tu 0002, Yuan-Jui Chen, Alexander H. Liu, Hung-yi Lee |
INTERSPEECH | 1 |
| 2019 | End-to-End Text-to-Speech for Low-Resource Languages by Cross-Lingual Transfer LearningabstractEnd-to-end text-to-speech (TTS) has shown great success on large quantities of paired text plus speech data.However, laborious data collection remains difficult for at least 95% of the languages over the world, which hinders the development of TTS in different languages.In this paper, we aim to build TTS systems for such low-resource (target) languages where only very limited paired data are available.We show such TTS can be effectively constructed by transferring knowledge from a high-resource (source) language.Since the model trained on source language cannot be directly applied to target language due to input space mismatch, we propose a method to learn a mapping between source and target linguistic symbols.Benefiting from this learned mapping, pronunciation information can be preserved throughout the transferring procedure.Preliminary experiments show that we only need around 15 minutes of paired data to obtain a relatively good TTS system.Furthermore, analytic studies demonstrated that the automatically discovered mapping correlate well with the phonetic expertise. Yuan-Jui Chen, Tao Tu 0002, Cheng-chieh Yeh, Hung-yi Lee |
INTERSPEECH | 2 |