VLDB 2026 Research / reviewers in the wild / expert
Yehao Lu
dblp:323/4954
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-7544-910XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniScene-MoTion: Unified Scene & Motion-aware Diffusion Transition FrameworkabstractVideo transitions are critical for ensuring temporal coherence in edited media, yet existing methods often rely on handcrafted effects or relative-scale trajectories that fail to capture the physical structure of real-world scenes. In this work, we introduce a scale-aware video transition framework that explicitly incorporates depth-aware 3D reasoning into a diffusion-based generation pipeline. Built upon a powerful I2V foundation, our method leverages single-image depth prediction to align camera motion with metric-scale geometry, enabling physically consistent transitions. To reduce reliance on precise camera inputs, we propose a bidirectional conditional control module and a progressive training strategy with conditional dropout, enhancing generalization to loosely specified or missing camera trajectories. Extensive experiments demonstrate that our approach achieves state-of-the-art performance, delivering realistic, geometrically coherent transitions across diverse scenes and applications with minimal input guidance. Chongmian Wang, Xinghe Fu, Yehao Lu |
AAAI | 4 |
| 2026 | Exploring Vision-Based Active 3D Object Detection by Informativeness CharacterizationabstractVision-based 3D object detection (3DOD) gains lots of attention due to its low cost for deployment compared to Lidar-based tasks, while it suffers from labor-expensive data annotations. At the same time, active learning (AL) has shown great potential in reducing annotation costs in related tasks, which can maximize model performance within very limited labeled data. In this paper, we explore active learning for vision-based 3DOD for the first time. Inspired by the entropy analysis, we involve three concerns to characterize the sample informativeness: sample diversity in input space, feature informativeness in BEV space, and result distribution in prediction space. Based on these concerns, we propose a novel AL framework named HMAD, which utilizes Height Modeling and Adaptive Diversity-based sampling for comprehensive informativeness characterization. In HMAD, we first propose a novel height-guided adversarial module in BEV space, which measures the informativeness of height modeling for 2D-to-3D mapping in an adversarial manner. Furthermore, Budget-aware SpatioTemporal diversity Sampling (BSTS) and Class Balance Sampling (CBS) are proposed to adaptively measure the sample informativeness in input and prediction space, respectively. Finally, the three components are integrated into a two-stage sampling strategy, with which the most informative samples can be selected and annotated for the next iteration. Experiments evidence that HMAD achieves comparable performances by only using 50% annotated training data, and can generalize well on different conditions. Yiming Wu 0006, Yehao Lu, Xuewei Li 0003, Xiubo Liang, Xi Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera ControlabstractRecent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges, as users often struggle to provide precise camera parameters when working with arbitrary real-world images without knowledge of their depth nor scene scale. To address these real-world application issues, we propose RealCam-I2V, a novel diffusion-based video generation framework that integrates monocular metric depth estimation to establish 3D scene reconstruction in a preprocessing step. During training, the reconstructed 3D scene enables scaling camera parameters from relative to metric scales, ensuring compatibility and scale consistency across diverse real-world images. In inference, RealCam-I2V offers an intuitive interface where users can precisely draw camera trajectories by dragging within the 3D scene. To further enhance precise camera control and scene consistency, we propose scene-constrained noise shaping, which shapes high-level noise and also allows the framework to maintain dynamic and coherent video generation in lower noise stages. RealCam-I2V achieves significant improvements in controllability and video quality on the RealEstate10K and out-of-domain images. We further enables applications like camera-controlled looping video generation and generative frame interpolation. Project page: https://zgctroy.github.io/RealCam-I2V. Guangcong Zheng, Shuigen Zhan, Yehao Lu, Yining Lin, Chuanyun Deng, Yepan Xiong |
ICCV | 6 |
| 2025 | Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-Time Open-Vocabulary Object DetectionabstractThe Mixture of Experts (MoE) architecture has excelled in Large Vision-Language Models (LVLMs), yet its potential in real-time open-vocabulary object detectors, which also leverage large-scale vision-language datasets but smaller models, remains unexplored. This work investigates this domain, revealing intriguing insights. In the shallow layers, experts tend to cooperate with diverse peers to expand the search space. While in the deeper layers, fixed collaborative structures emerge, where each expert maintains 2-3 fixed partners and distinct expert combinations are specialized in processing specific patterns. Concretely, we propose Dynamic-DINO, which extends Grounding DINO 1.5 Edge from a dense model to a dynamic inference framework via an efficient MoE-Tuning strategy. Additionally, we design a granularity decomposition mechanism to decompose the Feed-Forward Network (FFN) of base model into multiple smaller expert networks, expanding the subnet search space. To prevent performance degradation at the start of fine-tuning, we further propose a pre-trained weight allocation strategy for the experts, coupled with a specific router initialization. During inference, only the input-relevant experts are activated to form a compact subnet. Experiments show that, pretrained with merely 1.56M open-source data, Dynamic-DINO outperforms Grounding DINO 1.5 Edge, pretrained on the private Grounding20M dataset. Yehao Lu, Minghe Weng, Zekang Xiao, Guangcong Zheng, Ping Luo 0002 |
ICCV | 1 |
| 2025 | Context-based emotion recognition: A survey
Rizwan Abbas, Bingnan Ni, Ruhui Ma, Yehao Lu, Xi Li 0001 |
Neurocomputing | 5 |
| 2025 | Decoupling Discriminative Attributes for Few-Shot Fine-Grained RecognitionabstractFew-shot fine-tuning of pre-trained vision-language models (VLMs) for downstream tasks has gained widespread attention for reducing data annotation efforts while maintaining high performance. However, we observe that VLMs excel in excluding most incorrect classes in fine-grained recognition tasks, but struggles with a small set of confusing categories, which are typically highly similar subspecies. Existing few-shot fine-tuning methods attempt to directly recognize the correct category among all predefined classes, limiting their ability to capture discriminative features for those confusing categories. This raises an intriguing question: Can we specifically extract useful information from confusing classes to enhance fine-grained recognition performance? Based on this insight, we propose a hierarchical few-shot fine-tuning framework to address the severe confusion problem while ensuring the interpretability, namely Attribute-Decoupled Discriminator (AttrDD). Instead of thinking once among all classes, AttrDD employs a two-stage recognition, "think through" then "think smart". Specifically, in the first phase, a representative VLM, CLIP, is fine-tuned to select the Top-K confusing classes. In the second phase, we leverage the knowledge of large language models (LLMs) to generate fixed format descriptions of attribute differences between these confusing classes via in-context learning. Attribute-decoupled classifications are then conducted to capture fine-grained discriminative features. To achieve parameter-efficient fine-tuning, we introduce a lightweight attention adapter for each phase to align image features with task-specific textual features and LLM-generated textual features. Extensive experiments on 9 fine-grained recognition benchmarks demonstrate that AttrDD consistently outperforms existing baselines by wide margins. Yehao Lu, Chaoxiang Cai, Wei Su 0009, Guangcong Zheng, Xuewei Li 0003, Xi Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object DetectionabstractVision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it en-compasses inherent advantages in reducing blind spots and expanding perception range. While previous work mainly focuses on accurately estimating depth or height for 2D-to-3D mapping, ignoring the position approximation error in the voxel pooling process. Inspired by this insight, we propose a novel voxel pooling strategy to reduce such error, dubbed BEVSpread. Specifically, instead of bringing the image features contained in a frustum point to a single BEV grid, BEVSpread considers each frustum point as a source and spreads the image features to the surrounding BEV grids with adaptive weights. To achieve superior prop- agation performance, a specific weight function is designed to dynamically control the decay speed of the weights according to distance and depth. Aided by customized CUDA parallel acceleration, BEVSpread achieves comparable inference time as the original voxel pooling. Extensive experiments on two large-scale roadside benchmarks demonstrate that, as a plug-in, BEVSpread can significantly improve the performance of existing frustum-based BEV methods by a large margin of (1.12, 5.26, 3.01) AP in vehicle, pedestrian and cyclist. The source code will be made publicly available at BEVSpread. Yehao Lu, Guangcong Zheng, Shuigen Zhan, Xiaoqing Ye, Zichang Tan, Jingdong Wang 0001, Gaoang Wang, Xi Li 0001 |
CVPR | 2 |
| 2023 | PSO-Based Sparse Source Location in Large-Scale Environments With a UAV SwarmabstractLocating multiple sources in an unknown environment based on their signal strength is called a multi-source location problem. In recent years, there has been great interest in deploying autonomous devices to solve it. A particle swarm optimizer (PSO) is a widely employed source location method. Yet most work in this field focuses on a flat search space while ignoring height information. An unmanned aerial vehicle (UAV) has a coarser but wider view as it flies higher. Inspired by such facts, this paper focuses on improving the efficiency of locating sources by utilizing height information through UAVs. A novel source location model is designed where their sensing range gradually increases as their flying height rises, but their obtained signal strength fades away. It can be directly deployed to existing PSO-based multi-source location methods and improve their performance, especially in a large-scale environment with sparse sources. UAVs can spontaneously switch their search schemes between a rough search at a higher height and a fine one at a lower height. Experimental results of three PSO-based methods show their significant improvement after deploying our model. Given the same computation resources, its deployment leads to over 30% hike in both location accuracy and speed. This represents a great advance to the field of source location. Yehao Lu, Yunzhe Wu, Cheng Wang 0001, Di Zang, Abdullah Abusorrah, MengChu Zhou |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Moving-Distance-Minimized PSO for Mobile Robot SwarmabstractParticle swarm optimizer (PSO) and mobile robot swarm are two typical swarm techniques. Many applications emerge separately along both of them while the similarity between them is rarely considered. When a solution space is a certain region in reality, a robot swarm can replace a particle swarm to explore the optimal solution by performing PSO. In this way, a mobile robot swarm should be able to efficiently explore an area just like the particle swarm and uninterruptedly work even under the shortage of robots or in the case of unexpected failure of robots. Furthermore, the moving distances of robots are highly constrained because energy and time can be costly. Inspired by such requirements, this article proposes a moving-distance-minimized PSO (MPSO) for a mobile robot swarm to minimize the total moving distance of its robots while performing optimization. The distances between the current robot positions and the particle ones in the next generation are utilized to derive paths for robots such that the total distance that robots move is minimized, hence minimizing the energy and time for a robot swarm to locate the optima. Experiments on 28 CEC2013 benchmark functions show the advantage of the proposed method over the standard PSO. By adopting the given algorithm, the moving distance can be reduced by more than 66% and the makespan can be reduced by nearly 70% while offering the same optimization effects. Yehao Lu, Lei Che, MengChu Zhou |
IEEE Trans. Cybern. | 2 |