Weilong Yan

dblp:350/8550 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 52% Transfer learning and domain adaptation · 15% Video understanding and tracking · 15%
Computer graphics and multimedia
1 paper
Image and video processing · 50% Computational photography and imaging · 50%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
depth estimation
0.912025
Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors · CVPR 2025
Computer vision › 3D vision › depth estimation › self-supervised depth estimation
self-supervised monocular depth estimation
0.912025
Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors · CVPR 2025
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
synthetic-to-real domain adaptation
0.912025
Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors · CVPR 2025
Computer vision › Vision and language › video-language model
video large language model
0.912025
PVChat: Personalized Video Chat with One-Shot Learning · ICCV 2025
Computer vision › Video understanding and tracking
video question answering
0.912025
PVChat: Personalized Video Chat with One-Shot Learning · ICCV 2025
Computer vision › 3D vision › motion estimation
camera motion estimation
0.712023
Deep Homography Mixture for Single Image Rolling Shutter Correction · ICCV 2023
Computer vision › 3D vision › multi-view geometry
homography estimation
0.712023
Deep Homography Mixture for Single Image Rolling Shutter Correction · ICCV 2023
Image and video processing
image restoration
0.712023
Deep Homography Mixture for Single Image Rolling Shutter Correction · ICCV 2023
Computational photography and imaging › image signal processing
rolling shutter correction
0.712023
Deep Homography Mixture for Single Image Rolling Shutter Correction · ICCV 2023
Robotics › Autonomous driving › perception › perception robustness
perception in adverse weather
0.312025
Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors · CVPR 2025

Methods — techniques the papers use, named apart from their topics

motion basis learning · 1.3deep homography mixture · 1.3pseudo-labeling · 0.9progressive image-to-video learning · 0.9one-shot learning · 0.9motion and structure priors · 0.9mixture-of-heads attention · 0.9consistency reweighting · 0.9
YearPublicationVenuePosition
2025 Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors
abstract
Self-supervised depth estimation from monocular cameras in diverse outdoor conditions, such as daytime, rain, and nighttime, is challenging due to the difficulty of learning universal representations and the severe lack of labeled real-world adverse data. Previous methods either rely on synthetic inputs and pseudo-depth labels or directly apply daytime strategies to adverse conditions, resulting in suboptimal results. In this paper, we present the first synthetic-to-real robust depth estimation framework, incorporating motion and structure priors to capture real-world knowledge effectively. In the synthetic adaptation, we transfer motion-structure knowledge inside cost volumes for better robust representation, using a frozen daytime model to train a depth estimator in synthetic adverse conditions. In the innovative real adaptation, which targets to fix synthetic-real gaps, models trained earlier identify the weather-insensitive regions with a designed consistency-reweighting strategy to emphasize valid pseudo-labels. We introduce a new regularization by gathering explicit depth distribution to constrain the model facing real-world data. Experiments show that our method outperforms the state-of-the-art across diverse conditions in multi-frame and single-frame evaluations. We achieve improvements of 7.5% and 4.3% in Ab-sRel and RMSE on average for nuScenes and Robotcar datasets (daytime, nighttime, rain). In zero-shot evaluation of DrivingStereo (rain, fog), our method generalizes better than previous ones. The code is at Syn2Real-Depth.
Weilong Yan, Shuwei Shao, Robby T. Tan
CVPR1
2025 PVChat: Personalized Video Chat with One-Shot Learning
abstract
Video large language models (ViLLMs) excel in general video understanding, e.g., recognizing activities like talking and eating, but struggle with identity-aware comprehension, such as "Wilson is receiving chemotherapy" or "Tom is discussing with Sarah", limiting their applicability in smart healthcare and smart home environments. To address this limitation, we propose a one-shot learning framework PVChat, the first personalized ViLLM that enables subject-aware question answering (QA) from a single video for each subject. Our approach optimizes a Mixture-of-Heads (MoH) enhanced ViLLM on a synthetically augmented video-QA dataset, leveraging a progressive image-to-video learning strategy. Specifically, we introduce an automated augmentation pipeline that synthesizes identity-preserving positive samples and retrieves hard negatives from existing video corpora, generating a diverse training dataset with four QA types: existence, appearance, action, and location inquiries. To enhance subject-specific learning, we propose a ReLU Routing MoH attention mechanism, alongside two novel objectives: (1) Smooth Proximity Regularization for progressive learning through exponential distance scaling and (2) Head Activation Enhancement for balanced attention routing. Finally, we adopt a two-stage training strategy, transitioning from image pre-training to video fine-tuning, enabling a gradual learning process from static attributes to dynamic representations. We evaluate PVChat on diverse datasets covering medical scenarios, TV series, anime, and real-world footage, demonstrating its superiority in personalized feature understanding after learning from a single video, compared to state-of-the-art ViLLMs.
Weilong Yan, Zhenxi Li, Si Yong Yeo
ICCV2
2024 Remote Sensing Domain Adaptive Alignment via Student-Teacher Learning
abstract
Image Alignment between Synthetic Aperture Radar (SAR) and Electro-Optical (EO) imagery is a task that has comprehensive remote sensing capabilities. Traditional deep learning-based SAR-optical image matching models heavily rely on supervised learning with expensive annotated datasets, leading to reduced accuracy and overfitting when encountering insufficient data during training. To tackle this issue, this paper proposes a student-teacher framework for Domain Adaptation (DA) approach, transferring deep learning models from well-annotated source domains like normal outside optical images to non-annotated SAR-EO target domains. Additionally, in contrast to previous methods which usually use CNN or ordinary Transformer structure to extract features from image pairs, we use self and cross attention mechanisms in Transformer to obtain feature descriptors that are conditioned on both multimodal images. The larger global receptive field and better feature extraction capability provided by this Transformer shows its ability to accommodate large disparities of multimodal data and manage fewer textures such as rural areas with forest or desert in satellite images, where previous backbones usually struggle to produce repeatable and correct interest points.
Qiuhang Liu, Weilong Yan, Bo Wang 0019, Bharadwaj Veeravalli, Robby T. Tan
IGARSS2
2023 Deep Homography Mixture for Single Image Rolling Shutter Correction
abstract
We present a deep homography mixture motion model for single image rolling shutter correction. Rolling shutter (RS) effects are often caused by row-wise exposure delay in the widely adopted CMOS sensor. Previous methods often require more than one frame for the correction, leading to data quality requirements. Few approaches address the more challenging task of single image RS correction, which often adopt designs like trajectory estimation or long rectangular kernels, to learn the camera motion parameters of an RS image, to restore the global shutter (GS) image. In this work, we adopt a more straightforward method to learn deep homography mixture motion between an RS image and its corresponding GS image, without large solution space or strict restrictions on image features. We show that dividing an image into blocks with a Gaussian weight of block scanlines fits well for the RS setting. Moreover, instead of directly learning the motion mapping, we learn coefficients that assemble several motion bases to produce the correction motion, where these bases are learned from the consecutive frames of natural videos beforehand. Experiments show that our method outperforms existing single RS methods statistically and visually, in both synthesized and real RS images. Our code and dataset are available at https://github.com/DavidYan2001/Deep_HM.
Weilong Yan, Robby T. Tan, Bing Zeng 0001, Shuaicheng Liu
ICCV1