EDBT 2026 Demo / reviewers in the wild / expert
Ting Yu 0016
dblp:181/2866-16
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-6918-3157ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CloudCap3D: enhancing 3D in-scene descriptions via point cloud integration and efficient text filtering
Zhiyuan Niu, Zeyi Dong, Tieqiao Liu, Ting Yu 0016 |
Multim. Syst. | 5 |
| 2026 | SwiftCraft3D: semantic-enhanced multi-view prompting for efficient and high-fidelity text-to-3D generation
Zeyi Dong, Ting Yu 0016 |
Vis. Comput. | 2 |
| 2025 | Fine-grained Adaptive Visual Prompt for Generative Medical Visual Question AnsweringabstractMedical Visual Question Answering (MedVQA) serves as an automated medical assistant, capable of answering patient queries and aiding physician diagnoses based on medical images and questions. Recent advancements have shown that incorporating Large Language Models (LLMs) into MedVQA tasks significantly enhances the capability for answer generation. However, for tasks requiring fine-grained organ-level precise localization, relying solely on language prompts struggles to accurately locate relevant regions within medical images due to substantial background noise. To address this challenge, we explore the use of visual prompts in MedVQA tasks for the first time and propose fine-grained adaptive visual prompts to enhance generative MedVQA. Specifically, we introduce an Adaptive Visual Prompt Creator that adaptively generates region-level visual prompts based on image characteristics of various organs, providing fine-grained references for LLMs during answer retrieval and generation from the medical domain, thereby improving the model's precise cross-modal localization capabilities on original images. Furthermore, we incorporate a Hierarchical Answer Generator with Parameter-Efficient Fine-Tuning (PEFT) techniques, significantly enhancing the model's understanding of spatial and contextual information with minimal parameter increase, promoting the alignment of representation learning with the medical space. Extensive experiments on VQA-RAD, SLAKE, and DME datasets validate the effectiveness of our proposed method, demonstrating its potential in generative MedVQA. Ting Yu 0016, Zixuan Tong, Jun Yu 0002, Ke Zhang 0029 |
AAAI | 1 |
| 2025 | WP-CMA: Waypoint Prediction for Cross-modal Alignment of Vision-and-Language Navigation in Continuous EnvironmentsabstractVision-and-Language Navigation (VLN) tasks challenge an embodied agent to follow natural language instructions and reach a specified goal in visually rich environments. In continuous navigation settings, predicting accurate waypoints becomes particularly difficult due to the complex spatial dynamics and the semantic gap between language and vision modalities. In this work, we introduce WP-CMA, a novel cross-modal alignment framework designed for waypoint prediction in continuous VLN scenarios. Our framework consists of two primary modules: a Depth Space Reasoning Module (DSRM) and a Vision-Language Fusion Module (VLFM). The DSRM leverages 3D information from depth imagery to perform spatial reasoning and predict optimal waypoints. Concurrently, the VLFM integrates visual features with linguistic instructions to generate a unified representation, which provides essential context to the DSRM. Experiments on the R2R-CE dataset show that our proposed method significantly improves navigation efficiency. Siyang Fu, Ting Yu 0016 |
MMAsia | 3 |
| 2025 | Semi-Supervised RGB-D Hand Gesture Recognition via Mutual Learning of Self-Supervised ModelsabstractHuman hand gesture recognition is important to human–computer interaction. Gesture recognition based on RGB and Depth (RGB-D) data exploits both RGB and depth images to provide comprehensive results. However, the research under scenario with insufficient annotated data is not adequate. In view of the problem, our insight is to perform self-supervised learning with respect to each modality, transfer the learned information to modality-specific classifiers, and then fuse their results for final decision. To this end, we propose a semi-supervised hand gesture recognition method known as Mutual Learning of Rotation-Aware Gesture Predictors (MLRAGP), which exploits unlabeled training RGB and depth images via self-supervised learning and achieves multi-modal decision fusion through deep mutual learning. For each modality, we rotate both labeled and unlabeled images to fixed angles and train an angle predictor to predict the angles, then we use the feature extraction part of the angle predictor to construct the category predictor and train it through labeled data. We subsequently fuse the category predictors about both modalities by impelling each of them to simulate the probability estimation produced by the other, and making the prediction of labeled images to approach the ground truth annotation. During the training of category predictor and mutual learning, the parameters of feature extractors can be slighted fine-tuned to avoid under-fitting. Experimental results on NTU-Microsoft Kinect Hand Gesture dataset and Washington RGB-D dataset demonstrate the superiority of this framework to existing methods. Jian Zhang 0026, Kaihao He, Ting Yu 0016, Jun Yu 0002, Zhenming Yuan |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | 3D human pose estimation with multi-hypotheses gated transformer
Xiena Dong, Jian Zhang 0026, Jun Yu 0002, Ting Yu 0016 |
Multim. Syst. | 4 |
| 2024 | A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D ScenesabstractThree-Dimensional (3D) dense captioning is an emerging vision-language bridging task that aims to generate multiple detailed and accurate descriptions for 3D scenes. It presents significant potential and challenges due to its closer representation of the real world compared to 2D visual captioning, as well as complexities in data collection and processing of 3D point cloud sources. Despite the popularity and success of existing methods, there is a lack of comprehensive surveys summarizing the advancements in this field, which hinders its progress. In this paper, we provide a comprehensive review of 3D dense captioning, covering task definition, architecture classification, dataset analysis, evaluation metrics, and in-depth prosperity discussions. Based on a synthesis of previous literature, we refine a standard pipeline that serves as a common paradigm for existing methods. We also introduce a clear taxonomy of existing models, summarize technologies involved in different modules, and conduct detailed experiment analysis. Instead of a chronological order introduction, we categorize the methods into different classes to facilitate exploration and analysis of the differences and connections among existing techniques. We also provide a reading guideline to assist readers with different backgrounds and purposes in reading efficiently. Furthermore, we propose a series of promising future directions for 3D dense captioning by identifying challenges and aligning them with the development of related tasks, offering valuable insights and inspiring future research in this field. Our aim is to provide a comprehensive understanding of 3D dense captioning, foster further investigations, and contribute to the development of novel applications in multimedia and related domains. Ting Yu 0016, Shuhui Wang, Weiguo Sheng 0001, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Multi-Granularity Contrastive Cross-Modal Collaborative Generation for End-to-End Long-Term Video Question AnsweringabstractLong-term Video Question Answering (VideoQA) is a challenging vision-and-language bridging task focusing on semantic understanding of untrimmed long-term videos and diverse free-form questions, simultaneously emphasizing comprehensive cross-modal reasoning to yield precise answers. The canonical approaches often rely on off-the-shelf feature extractors to detour the expensive computation overhead, but often result in domain-independent modality-unrelated representations. Furthermore, the inherent gradient blocking between unimodal comprehension and cross-modal interaction hinders reliable answer generation. In contrast, recent emerging successful video-language pre-training models enable cost-effective end-to-end modeling but fall short in domain-specific ratiocination and exhibit disparities in task formulation. Toward this end, we present an entirely end-to-end solution for long-term VideoQA: Multi-granularity Contrastive cross-modal collaborative Generation (MCG) model. To derive discriminative representations possessing high visual concepts, we introduce Joint Unimodal Modeling (JUM) on a clip-bone architecture and leverage Multi-granularity Contrastive Learning (MCL) to harness the intrinsically or explicitly exhibited semantic correspondences. To alleviate the task formulation discrepancy problem, we propose a Cross-modal Collaborative Generation (CCG) module to reformulate VideoQA as a generative task instead of the conventional classification scheme, empowering the model with the capability for cross-modal high-semantic fusion and generation so as to rationalize and answer. Extensive experiments conducted on six publicly available VideoQA datasets underscore the superiority of our proposed method. Ting Yu 0016, Kunhao Fu, Jian Zhang 0026, Qingming Huang, Jun Yu 0002 |
IEEE Trans. Image Process. | 1 |
| 2024 | Token-Mixer: Bind Image and Text in One Embedding Space for Medical Image ReportingabstractMedical image reporting focused on automatically generating the diagnostic reports from medical images has garnered growing research attention. In this task, learning cross-modal alignment between images and reports is crucial. However, the exposure bias problem in autoregressive text generation poses a notable challenge, as the model is optimized by a word-level loss function using the teacher-forcing strategy. To this end, we propose a novel Token-Mixer framework that learns to bind image and text in one embedding space for medical image reporting. Concretely, Token-Mixer enhances the cross-modal alignment by matching image-to-text generation with text-to-text generation that suffers less from exposure bias. The framework contains an image encoder, a text encoder and a text decoder. In training, images and paired reports are first encoded into image tokens and text tokens, and these tokens are randomly mixed to form the mixed tokens. Then, the text decoder accepts image tokens, text tokens or mixed tokens as prompt tokens and conducts text generation for network optimization. Furthermore, we introduce a tailored text decoder and an alternative training strategy that well integrate with our Token-Mixer framework. Extensive experiments across three publicly available datasets demonstrate Token-Mixer successfully enhances the image-text alignment and thereby attains a state-of-the-art performance. Related codes are available at https://github.com/yangyan22/Token-Mixer. Jun Yu 0002, Zhenqi Fu, Ke Zhang 0029, Ting Yu 0016, Xianyun Wang, Hanliang Jiang, Junhui Lv, Qingming Huang, Weidong Han 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Long-Term Video Question Answering via Multimodal Hierarchical Memory Attentive NetworksabstractLong-term Video Question Answering plays an essential role in visual information retrieval, which aims at generating natural language answers to discretionary free-form questions about the referenced long-term video. Rather than remember the video as a sequence of visual content, humans have an innate cognitive ability to identify the critical moments related to the question at first glance, then tie together the specific evidence around these critical moments for further analysis and reasoning. Motivated by this intuition, we propose the multimodal hierarchical memory attentive networks with two heterogeneous memory subnetworks: the top guided memory network and the bottom enhanced multimodal memory attentive network. The top guided memory network serves as a shallow inference engine to pick relevant and informative moments of questions and obtain salient video content at a coarse-grained level. Subsequently, the bottom enhanced multimodal memory attentive network is designed as an in-depth reasoning engine to perform more accurate attention with cues from video bottom evidence in a fine-grained level to enhance question answering quality. We evaluate the proposed method on three publicly available video question answering benchmarks, namely ActivityNet-QA, MSRVTT-QA, and MSVD-QA. Experimental results demonstrate that the proposed approach significantly outperforms other state-of-the-art methods for long-term videos. Extensive ablation studies are carried out to explore the reasons behind the proposed model's effectiveness. Ting Yu 0016, Jun Yu 0002, Zhou Yu 0001, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Multi-task Compositional Network for Visual Relationship Detection
Yibing Zhan, Jun Yu 0002, Ting Yu 0016, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2020 | Compositional Attention Networks With Two-Stream Fusion for Video Question AnsweringabstractGiven a video, Video Question Answering (VideoQA) aims at answering arbitrary free-form questions about the video content in natural language. A successful VideoQA framework usually has the following two key components: 1) a discriminative video encoder that learns the effective video representation to maintain as much information as possible about the video and 2) a question-guided decoder that learns to select the most related features to perform spatiotemporal reasoning, as well as outputs the correct answer. We propose compositional attention networks (CAN) with two-stream fusion for VideoQA tasks. For the encoder, we sample video snippets using a two-stream mechanism (i.e., a uniform sampling stream and an action pooling stream) and extract a sequence of visual features for each stream to represent the video semantics with implementation. For the decoder, we propose a compositional attention module to integrate the two-stream features with the attention mechanism. The compositional attention module is the core of CAN and can be seen as a modular combination of a unified attention block. With different fusion strategies, we devise five compositional attention module variants. We evaluate our approach on one long-term VideoQA dataset, ActivityNet-QA, and two short-term VideoQA datasets, MSRVTT-QA and MSVD-QA. Our CAN model achieves new state-of-the-art results on all the datasets. Ting Yu 0016, Jun Yu 0002, Zhou Yu 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2019 | ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question AnsweringabstractRecent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA). Compared to the image domain where large scale and fully annotated benchmark datasets exists, VideoQA datasets are limited to small scale and are automatically generated, etc. These limitations restrict their applicability in practice. Here we introduce ActivityNet-QA, a fully annotated and large scale VideoQA dataset. The dataset consists of 58,000 QA pairs on 5,800 complex web videos derived from the popular ActivityNet dataset. We present a statistical analysis of our ActivityNet-QA dataset and conduct extensive experiments on it by comparing existing VideoQA baselines. Moreover, we explore various video representation strategies to improve VideoQA performance, especially for long videos. Zhou Yu 0001, Dejing Xu, Jun Yu 0002, Ting Yu 0016, Zhou Zhao 0001, Yueting Zhuang, Dacheng Tao |
AAAI | 4 |
| 2019 | On Exploring Undetermined Relationships for Visual Relationship DetectionabstractIn visual relationship detection, human-notated relationships can be regarded as determinate relationships. However, there are still large amount of unlabeled data, such as object pairs with less significant relationships or even with no relationships. We refer to these unlabeled but potentially useful data as undetermined relationships. Although a vast body of literature exists, few methods exploit these undetermined relationships for visual relationship detection. In this paper, we explore the beneficial effect of undetermined relationships on visual relationship detection. We propose a novel multi-modal feature based undetermined relationship learning network (MF-URLN) and achieve great improvements in relationship detection. In detail, our MF-URLN automatically generates undetermined relationships by comparing object pairs with human-notated data according to a designed criterion. Then, the MF-URLN extracts and fuses features of object pairs from three complementary modals: visual, spatial, and linguistic modals. Further, the MF-URLN proposes two correlated subnetworks: one subnetwork decides the determinate confidence, and the other predicts the relationships. We evaluate the MF-URLN on two datasets: the Visual Relationship Detection (VRD) and the Visual Genome (VG) datasets. The experimental results compared with state-of-the-art methods verify the significant improvements made by the undetermined relationships, e.g., the top-50 relation detection recall improves from 19.5% to 23.9% on the VRD dataset. Yibing Zhan, Jun Yu 0002, Ting Yu 0016, Dacheng Tao |
CVPR | 3 |