VLDB 2026 Research / reviewers in the wild / expert
Yun Wang 0053
dblp:36/3235-53
· DBLP profile ↗
17ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-8384-6981ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoSe: Robust Self-Supervised Stereo Matching Under Adverse Weather ConditionsabstractRecent self-supervised stereo matching methods have made significant progress, but their performance significantly degrades under adverse weather conditions such as night, rain, and fog. We identify two primary weaknesses contributing to this performance degradation. First, adverse weather introduces noise and reduces visibility, making CNN-based feature extractors struggle with degraded regions like reflective and textureless areas. Second, these degraded regions can disrupt accurate pixel correspondences, leading to ineffective supervision based on the photometric consistency assumption. To address these challenges, we propose injecting robust priors derived from the visual foundation model into the CNN-based feature extractor to improve feature representation under adverse weather conditions. We then introduce scene correspondence priors to construct robust supervisory signals rather than relying solely on the photometric consistency assumption. Specifically, we create synthetic stereo datasets with realistic weather degradations. These datasets feature clear and adverse image pairs that maintain the same semantic context and disparity, preserving the scene correspondence property. With this knowledge, we propose a robust self-supervised training paradigm, consisting of two key steps: robust self-supervised scene correspondence learning and adverse weather distillation. Both steps aim to align underlying scene results from clean and adverse image pairs, thus improving model disparity estimation under adverse weather effects. Extensive experiments demonstrate the effectiveness and versatility of our proposed solution, which outperforms existing state-of-the-art self-supervised methods. Codes are available at https://github.com/cocowy1/RoSe-Robust-Self-supervised-Stereo-Matching-under-Adverse-Weather-Conditions. Yun Wang 0053, Junjie Hu 0003, Junhui Hou, Chenghao Zhang 0003, Renwei Yang, Dapeng Oliver Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | General Diffusion Transformer for Arbitrary Medical Image TranslationabstractMedical image translation plays a crucial role in assisting clinical diagnosis by enabling cross-modal synthesis (e.g., Computed Tomography to Magnetic Resonance Imaging) and super-resolution, effectively addressing clinical challenges such as radiation exposure, prolonged scan times, and allergic reactions to contrast agents. However, existing approaches primarily focus on developing specialized models for specific tasks, limiting their adaptability across different applications. Developing a general model capable of handling arbitrary medical image translation tasks not only enhances cross-domain generalization but also aligns with the broader trend of artificial intelligence evolving from specialized to general-purpose solutions. Achieving such one-for-all model, however, presents three key challenges: (1) varying task complexity, (2) modality discrepancies, and (3) structural variations across anatomical regions within the same modality. To tackle the first challenge, we utilize an advanced diffusion-based training paradigm to endow the denoising model with extensive pattern coverage capabilities, thereby handling tasks of varying difficulty levels. Subsequently, a General Diffusion Transformer incorporating a Fuzzy Mixture-of-Experts (FMoE) module and an Entropy-guided Attention Soft Prompt (EASP) module is proposed. The FMoE module, equipped with nonlinear modeling capabilities, is designed to address modality discrepancies, while the EASP module is employed to enhance the model's perception of structural variations in images. Extensive qualitative and quantitative experiments demonstrate the effectiveness of the proposed model in the arbitrary medical image translation task. Jiahao Zheng 0001, Xiaoping Wang 0001, Yongcan Luo, Yun Wang 0053, Dapeng Oliver Wu |
IEEE Trans. Fuzzy Syst. | 4 |
| 2026 | SMFormer: Empowering Self-Supervised Stereo Matching via Foundation Models and Data AugmentationabstractRecent self-supervised stereo matching methods have made significant progress. They typically rely on the photometric consistency assumption, which presumes corresponding points across views share the same appearance. However, this assumption could be compromised by real-world disturbances, resulting in invalid supervisory signals and a significant accuracy gap compared to supervised methods. To address this issue, we propose SMFormer, a framework integrating more reliable self-supervision guided by the Vision Foundation Model (VFM) and data augmentation. We first incorporate the VFM with the Feature Pyramid Network (FPN), providing a discriminative and robust feature representation against disturbance in various scenarios. We then devise an effective data augmentation mechanism that ensures robustness to various transformations. The data augmentation mechanism explicitly enforces consistency between learned features and those influenced by illumination variations. Additionally, it regularizes the output consistency between disparity predictions of strong augmented samples and those generated from standard samples. Experiments on multiple mainstream benchmarks demonstrate that our SMFormer achieves state-of-the-art (SOTA) performance among self-supervised methods and even competes on par with supervised ones. Remarkably, in the challenging Booster benchmark, SMFormer even outperforms some SOTA supervised methods, such as CFNet. Yun Wang 0053, Zhengjie Yang, Jiahao Zheng 0001, Zhanjie Zhang, Dapeng Oliver Wu, Yulan Guo |
IEEE Trans. Image Process. | 1 |
| 2025 | DualNet: Robust Self-Supervised Stereo Matching with Pseudo-Label SupervisionabstractSelf-supervised stereo matching has drawn attention due to its ability to estimate disparity without needing ground-truth data. However, existing self-supervised stereo matching methods heavily rely on the photo-metric consistency assumption, which is vulnerable to natural disturbances, resulting in ambiguous supervision and inferior performance compared to the supervised ones. To relax the limitation of the photo-metric consistency assumption and even bypass this assumption, we propose a novel self-supervised framework named DualNet, which consists of two key steps: robust self-supervised teacher learning and pseudo-label supervised student training. Specifically, the teacher model is first trained in a self-supervised manner with a focus on feature-metric consistency and data augmentation consistency. Then, the output of the teacher model is geometrically constrained to obtain high-quality pseudo labels. Benefiting from these high-quality pseudo labels, the student model can outperform its teacher model by a large margin. With the two well-designed steps, the proposed framework DualNet ranks 1st among all self-supervised methods on multiple benchmarks, surprisingly even outperforming several supervised counterparts. Yun Wang 0053, Jiahao Zheng 0001, Chenghao Zhang 0003, Zhanjie Zhang, Kunhong Li 0001, Junjie Hu 0003 |
AAAI | 1 |
| 2025 | Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
Ao Ma 0005, Jiasong Feng, Ke Cao 0001, Jing Wang 0021, Yun Wang 0053, Quanwei Zhang, Zhanjie Zhang |
ICCV | 5 |
| 2025 | Learning Robust Stereo Matching in the Wild with Selective Mixture-of-ExpertsabstractRecently, learning-based stereo matching networks have advanced significantly. However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets. Leveraging Vision Foundation Models (VFMs) can intuitively enhance the model's robustness, but integrating such a model into stereo matching cost-effectively to fully realize their robustness remains a key challenge. To address this, we propose SMoEStereo, a novel framework that adapts VFMs for stereo matching through a tailored, scene-specific fusion of Low-Rank Adaptation (LoRA) and Mixture-of-Experts (MoE) modules. SMoEStereo introduces MoE-LoRA with adaptive ranks and MoE-Adapter with adaptive kernel sizes. The former dynamically selects optimal experts within MoE to adapt varying scenes across domains, while the latter injects inductive bias into frozen VFMs to improve geometric feature extraction. Importantly, to mitigate computational overhead, we further propose a lightweight decision network that selectively activates MoE modules based on input complexity, balancing efficiency with accuracy. Extensive experiments demonstrate that our method exhibits state-of-the-art cross-domain and joint generalization across multiple benchmarks without dataset-specific adaptation. The code is available at \textcolor{red}{https://github.com/cocowy1/SMoE-Stereo}. Yun Wang 0053, Longguang Wang, Chenghao Zhang 0003, Zhanjie Zhang, Ao Ma 0005, Chenyou Fan, Tin Lun Lam, Junjie Hu 0003 |
ICCV | 1 |
| 2025 | Instruction-aware Memory Network for Video RecognitionabstractThe rapid development of multimodal large language models (MLLMs) has highlighted their potential in video understanding. However, challenges remain in long video tasks, particularly in integrating visual features with prompt texts. Existing methods naively store processed video frames in a long-term memory bank, but neglect simple yet effective cross-modal integration. To address this, we introduce the instruction-aware memory construction (IaMC) model for long-term video understanding. By integrating visual and textual information, our model can obtain cross-modal features with robust understanding capabilities. These features are stored in a text-visual memory bank, enabling efficient long-term aggregation without surpassing LLM context or GPU memory limits. Experiments on the LVU dataset demonstrate state-of-the-art performance in video understanding and question answering, showcasing the IaMC model’s effectiveness and setting a new benchmark for long-term video analysis. The source code and trained models will be released publicly. Bimei Wang, Haijiang Li, Jisheng Dang, Yun Wang 0053, Zhixuan Chen, Jiyuan Lin, Teng Wang 0007 |
ICME | 4 |
| 2025 | Diff-LMM: Diffusion Teacher-Guided Spatio-Temporal Perception for Video Large Multimodal ModelsabstractDynamic spatio-temporal understanding is essential for video-based multimodal tasks, yet existing methods often struggle to capture fine-grained temporal and spatial relationships in long videos. Current approaches primarily rely on pre-trained CLIP encoders, which excel in semantic understanding but lack spatially-aware visual context. This leads to hallucinated results when interpreting fine-grained objects or scenes. To address these limitations, we propose a novel framework that integrates diffusion models into multimodal video models. By employing diffusion encoders at intermediate layers, we enhance visual representations through feature alignment and knowledge distillation losses, significantly improving the model's ability to capture spatial patterns over time. Additionally, we introduce a multi-level alignment strategy to learn robust feature correspondence from pre-trained diffusion models. Extensive experiments on benchmark datasets demonstrate our approach's state-of-the-art performance across multiple video understanding tasks. These results establish diffusion models as a powerful tool for enhancing multimodal video models in complex, dynamic scenarios. Jisheng Dang, Ligen Chen, Jingze Wu, Ronghao Lin, Bimei Wang, Yun Wang 0053, Nannan Zhu, Teng Wang 0007 |
IJCAI | 6 |
| 2025 | PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo MatchingabstractTemporally consistent depth estimation from stereo video is critical for real-world applications such as augmented reality, where inconsistent depth estimation disrupts the immersion of users.
Despite its importance, this task remains challenging due to the difficulty in modeling long-term temporal consistency in a computationally efficient manner.
Previous methods attempt to address this by aggregating spatio-temporal information but face a fundamental trade-off: limited temporal modeling provides only modest gains, whereas capturing long-range dependencies significantly increases computational cost.
To address this limitation, we introduce a memory buffer for modeling long-range spatio-temporal consistency while achieving efficient dynamic stereo matching.
Inspired by the two-stage decision-making process in humans, we propose a Pick-and-Play Memory (PPM) construction module for dynamic Stereo matching, dubbed as PPMStereo. PPM consists of a pick process that identifies the most relevant frames and a play process that weights the selected frames adaptively for spatio-temporal aggregation.
This two-stage collaborative process maintains a compact yet highly informative memory buffer while achieving temporally consistent information aggregation.
Extensive experiments validate the effectiveness of PPMStereo, demonstrating state-of-the-art performance in both accuracy and temporal consistency.Codes are available at \textcolor{blue}{https://github.com/cocowy1/PPMStereo}. Yun Wang 0053, Junjie Hu 0003, Qiaole Dong, Yanwei Fu 0001, Tin Lun Lam, Dapeng Oliver Wu |
NeurIPS | 1 |
| 2025 | VectorSketcher: Learning to create a vector-based free-hand sketch
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | LGAST: Towards high-quality arbitrary style transfer with local-global style learning
Zhanjie Zhang, Ruichen Xia 0002, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011, Wei Xing 0001 |
Neurocomputing | 5 |
| 2025 | DyArtbank: Diverse artistic style transfer via pre-trained stable diffusion and dynamic style prompt Artbank
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011 |
Knowl. Based Syst. | 6 |
| 2025 | SPAST: Arbitrary style transfer with style priors via pre-trained large-scale model
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011 |
Neural Networks | 5 |
| 2025 | ADStereo: Efficient Stereo Matching With Adaptive Downsampling and Disparity AlignmentabstractThe balance between accuracy and computational efficiency is crucial for the applications of deep learning-based stereo matching algorithms in real-world scenarios. Since matching cost aggregation is usually the most computationally expensive component, a common practice is to construct cost volumes at a low resolution for aggregation and then directly regress a high-resolution disparity map. However, current solutions often suffer from limitations such as the loss of discriminative features caused by downsampling operations that treat all pixels equally, and spatial misalignment resulting from repeated downsampling and upsampling. To overcome these challenges, this paper presents two sampling strategies: the Adaptive Downsampling Module (ADM) and the Disparity Alignment Module (DAM), to prioritize real-time inference while ensuring accuracy. The ADM leverages local features to learn adaptive weights, enabling more effective downsampling while preserving crucial structure information. On the other hand, the DAM employs a learnable interpolation strategy to predict transformation offsets of pixels, thereby mitigating the spatial misalignment issue. Building upon these modules, we introduce ADStereo, a real-time yet accurate network that achieves highly competitive performance on multiple public benchmarks. Specifically, our ADStereo runs over faster than the current state-of-the-art CREStereo (0.054s vs. ) under the same hardware while achieving comparable accuracy (1.82% vs. 1.69%) on the KITTI stereo 2015 benchmark. The codes are available at: https://github.com/cocowy1/ADStereo. Yun Wang 0053, Kunhong Li 0001, Longguang Wang, Junjie Hu 0003, Dapeng Oliver Wu, Yulan Guo |
IEEE Trans. Image Process. | 1 |
| 2024 | Learning Representations from Foundation Models for Domain Generalized Stereo Matching
Longguang Wang, Kunhong Li 0001, Yun Wang 0053, Yulan Guo |
ECCV (42) | 4 |
| 2024 | Cost Volume Aggregation in Stereo Matching Revisited: A Disparity Classification PerspectiveabstractCost aggregation plays a critical role in existing stereo matching methods. In this paper, we revisit cost aggregation in stereo matching from disparity classification and propose a generic yet efficient Disparity Context Aggregation (DCA) module to improve the performance of CNN-based methods. Our approach is based on an insight that a coarse disparity class prior is beneficial to disparity regression. To obtain such a prior, we first classify pixels in an image into several disparity classes and treat pixels within the same class as homogeneous regions. We then generate homogeneous region representations and incorporate these representations into the cost volume to suppress irrelevant information while enhancing the matching ability for cost aggregation. With the help of homogeneous region representations, efficient and informative cost aggregation can be achieved with only a shallow 3D CNN. Our DCA module is fully-differentiable and well-compatible with different network architectures, which can be seamlessly plugged into existing networks to improve performance with small additional overheads. It is demonstrated that our DCA module can effectively exploit disparity class priors to improve the performance of cost aggregation. Based on our DCA, we design a highly accurate network named DCANet, which achieves state-of-the-art performance on several benchmarks. Yun Wang 0053, Longguang Wang, Kunhong Li 0001, Dapeng Oliver Wu, Yulan Guo |
IEEE Trans. Image Process. | 1 |
| 2023 | CVCNet: Learning Cost Volume Compression for Efficient Stereo MatchingabstractState-of-the-art deep learning based stereo matching algorithms usually rely on full-size cost volumes for highly accurate disparity estimation. The full-size cost volume processes all possible disparity candidates equally without considering their different matching uncertainties. Consequently, considerable redundant computation is involved on those candidates with very low matching uncertainties, making these methods difficult to be deployed in real-time applications. To tackle this problem, we propose CVCNet featuring an adaptive disparity range prediction module (ADR) and a disparity refinement module (DRM). The ADR adaptively predicts pixel-wise disparity range to discard the “unimportant” disparity candidates. It enables our network to obtain a compressed cost volume. Besides, the DRM improves disparity range prediction and refines the predicted disparity map. With the proposed modules, our CVCNet learns to build a compressed cost volume to achieve efficient disparity estimation. Experimental results on the KITTI and SceneFlow datasets show that our method achieves state-of-the-art performance, and runs at a significant order of magnitude faster speed than existing 3D CNN based methods. Particularly, our method ranks$\mathbf {1}\mathrm{st}$on the KITTI 2012 and KITTI 2015 benchmarks among all published methods with running time shorter than 100 ms. Yulan Guo, Yun Wang 0053, Longguang Wang, Zi Wang 0008 |
IEEE Trans. Multim. | 2 |