EDBT 2026 Demo / reviewers in the wild / expert
Yuhao Dong
dblp:232/7896
· DBLP profile ↗
21ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0003-0564-4369ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uni-MMMU: A Massive Multi-discipline Multimodal Unified BenchmarkabstractKai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian, Dian Zheng, Hongbo Liu, Jingwen He, Bin Liu, Yu Qiao, Ziwei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuhao Dong, Shulin Tian, Dian Zheng, Jingwen He, Bin Liu 0016, Yu Qiao 0001, Ziwei Liu 0002 |
ACL (1) | 3 |
| 2025 | A Coarse-to-Fine Progressive Ensemble Framework for Coronary Artery LabelingabstractAutomatic coronary artery labeling is essential for accurate vascular identification and the diagnosis of coronary disease. The task requires delineating the full vasculature and classifying each segment; however, preserving global topology and local demarcation line precision is difficult due to complex anatomy and blurry contours. We propose a coarse-to-fine ensemble framework with two modules: a Coarse-to-fine Topology Extraction (CTE) network using topology priors for global continuity, and a Progressive Vessel Labeling (PVL) module with multibranch fusion for segmentation and classification. Experiments on the ARCADE dataset achieve a mean F1-score of 0.6028, outperforming state-of-the-art methods and enhancing topological integrity and labeling accuracy. Code: https://github.com/IPMINWU/PGSMODEL. Guansheng Peng, Zhuo Jin, Shaoxuan Wu, Yuhao Dong, Xiao Zhang 0028, Jun Feng 0003 |
BIBM | 4 |
| 2025 | Structural Points Dependency-Aware Template-Free Learning for Cardiac Mesh ReconstructionabstractHigh-fidelity, patient-specific cardiac mesh reconstruction underpins diagnosis, surgical planning, and hemodynamic simulation. Accurate and topologically coherent reconstruction remains challenging due to large inter-individual anatomical variability and complex cardiac morphology. We propose a template-free framework, TFSG, that integrates an Adaptive Structural Point Generation (ASG) module and a Structural Consistency Constraint (SCC). ASG extracts patientspecific anatomical landmarks from a point-cloud representation to guide deformation, while SCC enforces multi-level consistency (point distance, normal alignment and structural-point relations) to suppress topological and structural errors. Experiments on the CARE2025 WHS dataset show TFSG improves segmentation and mesh reconstruction quality compared to prior methods. Code: https://github.com/IPMI-NWU/TFSG. Shaoxuan Wu, Peilin Zhang, Yuhao Dong, Xiao Zhang 0028, Jun Feng 0003 |
BIBM | 4 |
| 2025 | 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive DiffusionabstractThe increasing demand for high-quality 3D assets across various industries necessitates efficient and automated 3D content creation. Despite recent advancements in 3D generative models, existing methods still face challenges with optimization speed, geometric fidelity, and the lack of assets for physically based rendering (PBR). In this paper, we introduce 3DTopia-XL, a scalable native 3D generative model designed to overcome these limitations. 3DTopia-XL leverages a novel primitive-based 3D representation, PrimX, which encodes detailed shape, albedo, and material field into a compact tensorial format, facilitating the modeling of high-resolution geometry with PBR assets. On top of the novel representation, we propose a generative framework based on Diffusion Transformer (DiT), which comprises 1) Primitive Patch Compression, 2) and Latent Primitive Diffusion. 3DTopia-XL learns to generate high-quality 3D assets from textual or visual inputs. Extensive qualitative and quantitative evaluations are conducted to demonstrate that 3DTopia-XL significantly outperforms existing methods in generating high-quality 3D assets with fine-grained textures and materials, efficiently bridging the quality gap between generative models and real-world applications. Zhaoxi Chen 0009, Jiaxiang Tang, Yuhao Dong, Ziang Cao, Fangzhou Hong, Yushi Lan, Tengfei Wang 0002, Haozhe Xie, Shunsuke Saito, Liang Pan, Dahua Lin, Ziwei Liu 0002 |
CVPR | 3 |
| 2025 | Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language ModelsabstractLarge Language Models (LLMs) demonstrate enhanced capabilities and reliability by reasoning more, evolving from Chain-of-Thought prompting to product-level solutions like OpenAI o1. Despite various efforts to improve LLM reasoning, high-quality long-chain reasoning data and optimized training pipelines still remain inadequately explored in vision-language tasks. In this paper, we present InsightV, an early effort to 1) scalably produce long and robust reasoning data for complex multi-modal tasks, and 2) an effective training pipeline to enhance the reasoning capabilities of multi-modal large language models (MLLMs). Specifically, to create long and structured reasoning data without human labor, we design a two-step pipeline with a progressive strategy to generate sufficiently long and diverse reasoning paths and a multi-granularity assessment method to ensure data quality. We observe that directly supervising MLLMs with such long and complex reasoning data will not yield ideal reasoning ability. To tackle this problem, we design a multi-agent system consisting of a reasoning agent dedicated to performing long-chain reasoning and a summary agent trained to judge and summarize reasoning results. We further incorporate an iterative DPO algorithm to enhance the reasoning agent’s generation stability and quality. Based on the popular LLaVA-NeXT model and our stronger base MLLM, we demonstrate significant performance gains across challenging multi-modal benchmarks requiring visual reasoning. Benefiting from our multi-agent system, Insight-V can also easily maintain or improve performance on perception-focused multi-modal tasks. Yuhao Dong, Zuyan Liu, Hai-Long Sun, Winston Hu, Yongming Rao, Ziwei Liu 0002 |
CVPR | 1 |
| 2025 | Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language ModelabstractMultimodal language models (MLLMs) are increasingly being applied in real-world environments, necessitating their ability to interpret 3D spaces and comprehend temporal dynamics. Current methods often rely on specialized architectural designs or task-specific fine-tuning to achieve this. We introduce Coarse Correspondences, a simple lightweight method that enhances MLLMs’ spatial-temporal reasoning with 2D images as input, without modifying the architecture or requiring task-specific fine-tuning. Our method uses a lightweight tracking model to identify primary object correspondences between frames in a video or across different image viewpoints, and then conveys this information to MLLMs through visual prompting. We demonstrate that this simple training-free approach brings substantial gains to GPT4-V/O consistently on four benchmarks that require spatial-temporal reasoning, including +20.5% improvement on ScanQA, +9.7% on OpenEQA’s episodic memory subset, +6.0% on the long-form video benchmark EgoSchema, and +11% on the R2R navigation benchmark. Additionally, we show that Coarse Correspondences can also enhance open-source MLLMs’ spatial reasoning (by +6.9% on ScanQA) when applied in both training and inference and that the improvement can generalize to unseen datasets such as SQA3D (+3.1%). Taken together, we show that Coarse Correspondences effectively and efficiently boosts models’ performance on downstream tasks requiring spatial-temporal reasoning. Benlin Liu, Yuhao Dong, Zixian Ma, Yansong Tang, Luming Tang, Yongming Rao, Wei-Chiu Ma, Ranjay Krishna |
CVPR | 2 |
| 2025 | EgoLife: Towards Egocentric Life AssistantabstractWe introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data collection study where six participants lived together for one week, continuously recording their daily activities—including discussions, shopping, cooking, social-izing, and entertainment—using AI glasses for multimodal person-view video references. This effort resulted in EgoLife Dataset, a comprehensive 300-hour egocentric, terpersonal, multiview, and multimodal daily life with intensive annotation. Leveraging this dataset, we troduce EgoLifeQA, a suite of long-context, life-oriented question-answering tasks designed to provide meaningful sistance in daily life by addressing practical questions as recalling past relevant events, monitoring health and offering personalized recommendations.To address the key technical challenges of 1) developing robust visual-audio models for egocentric data, 2) enabling identity recognition, and 3) facilitating long-context question answering over extensive temporal information, we introduce EgoBulter, an integrated system comprising EgoGPT and EgoRAG. EgoGPT is an omni-modal model trained on egocentric datasets, achieving state-of-the-art performance on egocentric video understanding. EgoRAG is a retrieval-based component that supports answering ultra-long-context questions. Our experimental studies verify their working mechanisms and reveal critical factors and bottlenecks, guiding future improvements. By releasing our datasets, models, and benchmarks, we aim to stimulate further research in egocentric AI assistants. Shuai Liu 0002, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Bo Li 0080, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Jörg Widmer, Francesco Gringoli, Lei Yang 0059, Ziwei Liu 0002 |
CVPR | 4 |
| 2025 | Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric PerspectivesabstractRecent advancements in Vision-Language Models (VLMs) have sparked interest in their use for autonomous driving, particularly in generating interpretable driving decisions through natural language. However, the assumption that VLMs inherently provide visually grounded, reliable, and interpretable explanations for driving remains largely unexamined. To address this gap, we introduce DriveBench, a benchmark dataset designed to evaluate VLM reliability across 17 settings (clean, corrupted, and text-only inputs), encompassing 19,200 frames, 20,498 question-answer pairs, three question types, four mainstream driving tasks, and a total of 12 popular VLMs. Our findings reveal that VLMs often generate plausible responses derived from general knowledge or textual cues rather than true visual grounding, especially under degraded or missing visual inputs. This behavior, concealed by dataset imbalances and insufficient evaluation metrics, poses significant risks in safety-critical scenarios like autonomous driving. We further observe that VLMs struggle with multi-modal reasoning and display heightened sensitivity to input corruptions, leading to inconsistencies in performance. To address these challenges, we propose refined evaluation metrics that prioritize robust visual grounding and multi-modal understanding. Additionally, we highlight the potential of leveraging VLMs' awareness of corruptions to enhance their reliability, offering a roadmap for developing more trustworthy and interpretable decision-making systems in real-world autonomous driving contexts. The benchmark toolkit is publicly accessible. Shaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima, Qi Alfred Chen, Ziwei Liu 0002, Liang Pan |
ICCV | 3 |
| 2025 | Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and UnderstandingabstractHuman intelligence requires correctness and robustness, with the former being foundational for the latter. In video understanding, correctness ensures the accurate interpretation of visual content, and robustness maintains consistent performance in challenging conditions. Despite advances in video large language models (video LLMs), existing benchmarks inadequately reflect the gap between these models and human intelligence in maintaining correctness and robustness in video interpretation. We introduce the Video Thinking Test (Video-TT), to assess if video LLMs can interpret real-world videos as effectively as humans. Video-TT reflects genuine gaps in understanding complex visual narratives, and evaluates robustness against natural adversarial questions. Video-TT comprises 1,000 YouTube Shorts videos, each with one open-ended question and four adversarial questions that probe visual and narrative complexity. Our evaluation shows a significant gap between video LLMs and human performance. Yuanhan Zhang, Yunice Chew, Yuhao Dong, Aria Leo |
ICCV | 3 |
| 2025 | Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary ResolutionabstractVisual data comes in various forms, ranging from small icons of just a few pixels to long videos spanning hours. Existing multi-modal LLMs usually standardize these diverse visual inputs to fixed-resolution images or patches for visual encoders and yield similar numbers of tokens for LLMs. This approach is non-optimal for multimodal understanding and inefficient for processing inputs with long and short visual contents. To solve the problem, we propose Oryx, a unified multimodal architecture for the spatial-temporal understanding of images, videos, and multi-view 3D scenes. Oryx offers an on-demand solution to seamlessly and efficiently process visual inputs with arbitrary spatial sizes and temporal lengths through two core innovations: 1) a pre-trained OryxViT model that can encode images at any resolution into LLM-friendly visual representations; 2) a dynamic compressor module that supports 1x to 16x compression on visual tokens by request. These designs enable Oryx to accommodate extremely long visual contexts, such as videos, with lower resolution and high compression while maintaining high recognition precision for tasks like document understanding with native resolution and no compression. Beyond the architectural improvements, enhanced data curation and specialized training on long-context retrieval and spatial-aware data help Oryx achieve strong capabilities in image, video, and 3D multimodal understanding simultaneously. Zuyan Liu, Yuhao Dong, Ziwei Liu 0002, Winston Hu, Jiwen Lu, Yongming Rao |
ICLR | 2 |
| 2025 | Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasabstractEvent cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challenge. We introduce Talk2Event, the first large-scale benchmark for language-driven object grounding in event-based perception. Built from real-world driving data, Talk2Event provides over 30,000 validated referring expressions, each enriched with four grounding attributes -- appearance, status, relation to viewer, and relation to other objects -- bridging spatial, temporal, and relational reasoning. To fully exploit these cues, we propose EventRefer, an attribute-aware grounding framework that dynamically fuses multi-attribute representations through a Mixture of Event-Attribute Experts (MoEE). Our method adapts to different modalities and scene dynamics, achieving consistent gains over state-of-the-art baselines in event-only, frame-only, and event-frame fusion settings. We hope our dataset and approach will establish a foundation for advancing multimodal, temporally-aware, and language-driven perception in real-world robotics and autonomy. Lingdong Kong, Dongyue Lu, Alan Liang, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit Cottereau |
NeurIPS | 5 |
| 2025 | 3EED: Ground Everything Everywhere in 3DabstractVisual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce 3EED, a multi-platform, multi-modal 3D grounding benchmark featuring RGB and LiDAR data from vehicle, drone, and quadruped platforms. We provide over 128,000 objects and 22,000 validated referring expressions across diverse outdoor scenes -- 10x larger than existing datasets. We develop a scalable annotation pipeline combining vision-language model prompting with human verification to ensure high-quality spatial grounding. To support cross-platform learning, we propose platform-aware normalization and cross-modal alignment techniques, and establish benchmark protocols for in-domain and cross-platform evaluations. Our findings reveal significant performance gaps, highlighting the challenges and opportunities of generalizable 3D grounding. The 3EED dataset and benchmark toolkit are released to advance future research in language-driven 3D embodied perception. Yuhao Dong, Tianshuai Hu, Alan Liang, Youquan Liu, Dongyue Lu, Liang Pan, Lingdong Kong, Junwei Liang 0001, Ziwei Liu 0002 |
NeurIPS | 2 |
| 2025 | ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsabstractRecent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicated benchmark for assessing VLMs’ understanding of cinematic language. ShotBench includes 3,049 still images and 500 video clips drawn from more than 200 films, with each sample annotated by trained annotators or curated from professional cinematography resources, resulting in 3,608 high-quality question-answer pairs. We conduct a comprehensive evaluation of over 20 state-of-the-art VLMs across eight core cinematography dimensions. Our analysis reveals clear limitations in fine-grained perception and cinematic reasoning of current VLMs. To improve VLMs capability in cinematography understanding, we construct a large-scale multimodal dataset, named ShotQA, which contains about 70k Question-Answer pairs derived from movie shots.
Besides, we propose ShotVL and train this VLM model with a two-stage training strategy, integrating both supervised fine-tuning and Group Relative Policy Optimization (GRPO). Experimental results demonstrate that our model achieves substantial improvements, surpassing all existing strongest open-source and proprietary models evaluated on ShotBench, establishing a new state-of-the-art performance. Jingwen He, Dian Zheng, Yuhao Dong, Fan Zhang 0045, Yinan He, Weichao Chen 0001, Yu Qiao 0001, Wanli Ouyang, Shengjie Zhao 0001, Ziwei Liu 0002 |
NeurIPS | 5 |
| 2024 | Efficient Inference of Vision Instruction-Following Models with Elastic Cache
Zuyan Liu, Benlin Liu, Jiahui Wang 0001, Yuhao Dong, Guangyi Chen 0002, Yongming Rao, Ranjay Krishna, Jiwen Lu |
ECCV (17) | 4 |
| 2024 | Octopus: Embodied Vision-Language Programmer from Environmental Feedback
Yuhao Dong, Shuai Liu 0002, Bo Li 0080, Haoran Tan, Chencheng Jiang, Jiamu Kang, Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu 0002 |
ECCV (1) | 2 |
| 2023 | Complete-to-Partial 4D Distillation for Self-Supervised Point Cloud Sequence Representation LearningabstractRecent work on 4D point cloud sequences has attracted a lot of attention. However, obtaining exhaustively labeled 4D datasets is often very expensive and laborious, so it is especially important to investigate how to utilize raw unlabeled data. However, most existing self-supervised point cloud representation learning methods only consider geometry from a static snapshot omitting the fact that sequential observations of dynamic scenes could reveal more comprehensive geometric details. To overcome such issues, this paper proposes a new 4D self-supervised pretraining method called Complete-to-Partial 4D Distillation. Our key idea is to formulate 4D self-supervised representation learning as a teacher-student knowledge distillation framework and let the student learn useful 4D representations with the guidance of the teacher. Experiments show that this approach significantly outperforms previous pre-training approaches on a wide range of 4D point cloud sequence understanding tasks. Code is available at: https://github.com/dongyh20/C2P. Zhuoyang Zhang, Yuhao Dong, Li Yi 0001 |
CVPR | 2 |
| 2023 | A clinically applicable AI system for diagnosis of congenital heart diseases based on computed tomography images
Xiaowei Xu 0004, Qianjun Jia, Haiyun Yuan, Hailong Qiu, Yuhao Dong, Wen Xie 0008, Zeyang Yao, Zhiqaing Nie, Xiaomeng Li 0001, Yiyu Shi 0001, James Zou 0001, Meiping Huang, Jian Zhuang |
Medical Image Anal. | 5 |
| 2022 | Astrape: Anonymous Payment Channels with Boring Cryptography
Yuhao Dong, Ian Goldberg 0001, Sergey Gorbunov 0001, Raouf Boutaba |
ACNS | 1 |
| 2021 | Multi-Cycle-Consistent Adversarial Networks for Edge Denoising of Computed Tomography ImagesabstractAs one of the most commonly ordered imaging tests, the computed tomography (CT) scan comes with inevitable radiation exposure that increases cancer risk to patients. However, CT image quality is directly related to radiation dose, and thus it is desirable to obtain high-quality CT images with as little dose as possible. CT image denoising tries to obtain high-dose-like high-quality CT images (domain Y ) from low dose low-quality CT images (domain X ), which can be treated as an image-to-image translation task where the goal is to learn the transform between a source domain X (noisy images) and a target domain Y (clean images). Recently, the cycle-consistent adversarial denoising network (CCADN) has achieved state-of-the-art results by enforcing cycle-consistent loss without the need of paired training data, since the paired data is hard to collect due to patients’ interests and cardiac motion. However, out of concerns on patients’ privacy and data security, protocols typically require clinics to perform medical image processing tasks including CT image denoising locally (i.e., edge denoising). Therefore, the network models need to achieve high performance under various computation resource constraints including memory and performance. Our detailed analysis of CCADN raises a number of interesting questions that point to potential ways to further improve its performance using the same or even fewer computation resources. For example, if the noise is large leading to a significant difference between domain X and domain Y , can we bridge X and Y with a intermediate domain Z such that both the denoising process between X and Z and that between Z and Y are easier to learn? As such intermediate domains lead to multiple cycles, how do we best enforce cycle- consistency? Driven by these questions, we propose a multi-cycle-consistent adversarial network (MCCAN) that builds intermediate domains and enforces both local and global cycle-consistency for edge denoising of CT images. The global cycle-consistency couples all generators together to model the whole denoising process, whereas the local cycle-consistency imposes effective supervision on the process between adjacent domains. Experiments show that both local and global cycle-consistency are important for the success of MCCAN, which outperforms CCADN in terms of denoising quality with slightly less computation resource consumption. Xiaowei Xu 0004, Jinglan Liu, Yukun Ding, Hailong Qiu, Haiyun Yuan, Jian Zhuang, Wen Xie 0008, Yuhao Dong, Qianjun Jia, Meiping Huang, Yiyu Shi 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 10 |
| 2020 | ImageCHD: A 3D Computed Tomography Image Dataset for Classification of Congenital Heart Disease
Xiaowei Xu 0004, Jian Zhuang, Haiyun Yuan, Meiping Huang, Jianzheng Cen, Qianjun Jia, Yuhao Dong, Yiyu Shi 0001 |
MICCAI (4) | 8 |
| 2018 | Bitforest: a Portable and Efficient Blockchain-Based Naming System
Yuhao Dong, Woojung Kim, Raouf Boutaba |
CNSM | 1 |