Kun Wang 0042

dblp:05/1958-42 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-6390-2373ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021
YearPublicationVenuePosition
2026 Physics-Guided Posterior Sampling for Diffusion-Based Real-World Dehazing and Image Enhancement
abstract
Real-world image dehazing is a challenging task due to the collection of aligned hazy/clear image pairs under unpredictable and complex environments. To address this limitation, we propose a Physical-Guided Posterior Sampling (PGPS) method that designs a dehazing reconstruction posterior to sample an RGB and depth from pre-trained unconditional diffusion generation process. First, we introduce a Hybrid Degradation Atmospheric Scattering Model (HD-ASM) to adapt the diffusion model, enabling the generation of high-fidelity dehazed images from posterior samples without relying on the aligned hazy/clear image pairs. Second, we propose a two-stage sampling strategy with piecewise loss to improve sampling quality and stability, along with a post-processing technique to remove JPEG compression artifacts amplified by dehazing. Extensive experiments show that our method outperforms state-of-the-art techniques in image dehazing, and in the RTTS dataset’s complex human-vehicle environment. Additionally, our approach also surpasses other benchmarks in object detection, exhibiting superior generalization performance.
Junkai Fan, Kun Wang 0042, Zhiqiang Yan 0001, Jianjun Qian, Heyou Chang, Jun Li 0027, Jian Yang 0003
IEEE Trans. Circuits Syst. Video Technol.3
2025 Depth-Centric Dehazing and Depth-Estimation from Real-World Hazy Driving Video
abstract
In this paper, we study the challenging problem of simultaneously removing haze and estimating depth from real monocular hazy videos. These tasks are inherently complementary: enhanced depth estimation improves dehazing via the atmospheric scattering model (ASM), while superior dehazing contributes to more accurate depth estimation through the brightness consistency constraint (BCC). To tackle these intertwined tasks, we propose a novel depth-centric learning framework that integrates the ASM model with the BCC constraint. Our key idea is that both ASM and BCC rely on a shared depth estimation network. This network simultaneously exploits adjacent dehazed frames to enhance depth estimation via BCC and uses the refined depth cues to more effectively remove haze through ASM. Additionally, we leverage a non-aligned clear video and its estimated depth to independently regularize the dehazing and depth estimation networks. This is achieved by designing two discriminator networks: D_MFIR enhances high-frequency details in dehazed videos, and D_MDR reduces the occurrence of black holes in low-texture regions. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art techniques in both video dehazing and depth estimation tasks, especially in real-world hazy scenes.
Junkai Fan, Kun Wang 0042, Zhiqiang Yan 0001, Xiang Chen 0015, Shangbing Gao, Jun Li 0027, Jian Yang 0003
AAAI2
2025 Completion as Enhancement: A Degradation-Aware Selective Image Guided Network for Depth Completion
abstract
In this paper, we introduce the Selective Image Guided Network (SigNet), a novel degradation-aware framework that transforms depth completion into depth enhancement for the first time. Moving beyond direct completion using convolutional neural networks (CNNs), SigNet initially densifies sparse depth data through non-CNN densification tools to obtain coarse yet dense depth. This approach eliminates the mismatch and ambiguity caused by direct convolution over irregularly sampled sparse data. Subsequently, SigNet redefines completion as enhancement, establishing a self-supervised degradation bridge between the coarse depth and the targeted dense depth for effective RGB-D fusion. To achieve this, SigNet leverages the implicit degradation to adaptively select high-frequency components (e.g., edges) of RGB data to compensate for the coarse depth. This degradation is further integrated into a multi-modal conditional Mamba, dynamically generating the state parameters to enable efficient global high-frequency information interaction. We conduct extensive experiments on the NYUv2, DIML, SUN RGBD, and TOFDC datasets, demonstrating the state-of-the-art (SOTA) performance of SigNet.
Zhiqiang Yan 0001, Zhengxue Wang, Kun Wang 0042, Jun Li 0027, Jian Yang 0003
CVPR3
2025 Position: LLMs Can be Good Tutors in English Education
abstract
Jingheng Ye, Shen Wang, Deqing Zou, Yibo Yan, Kun Wang, Hai-Tao Zheng, Ruitong Liu, Zenglin Xu, Irwin King, Philip S. Yu, Qingsong Wen. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jingheng Ye, Shen Wang 0005, Deqing Zou, Kun Wang 0042, Hai-Tao Zheng 0002, Zenglin Xu, Irwin King, Philip S. Yu, Qingsong Wen
EMNLP5
2025 Can LLMs Outshine Conventional Recommenders? A Comparative Evaluation
abstract
Integrating large language models (LLMs) into recommender systems has created new opportunities for improving recommendation quality. However, a comprehensive benchmark is needed to thoroughly evaluate and compare the recommendation capabilities of LLMs with traditional recommender systems. In this paper, we introduce \recbench{}, which systematically investigates various item representation forms (including unique identifier, text, semantic embedding, and semantic identifier) and evaluates two primary recommendation tasks, i.e., click-through rate prediction (CTR) and sequential recommendation (SeqRec). Our extensive experiments cover up to 17 large models and are conducted across five diverse datasets from fashion, news, video, books, and music domains. Our findings indicate that LLM-based recommenders outperform conventional recommenders, achieving up to a 5% AUC improvement in CTR and up to a 170% NDCG@10 improvement in SeqRec. However, these substantial performance gains come at the expense of significantly reduced inference efficiency, rendering LLMs impractical as real-time recommenders. We have released our code and data to enable other researchers to reproduce and build upon our experimental results.
Qijiong Liu, Jieming Zhu, Kun Wang 0042, Hengchang Hu, Wei Guo 0006, Yong Liu 0018, Xiao-Ming Wu 0003
NeurIPS4
2025 Breaking the Discretization Barrier of Continuous Physics Simulation Learning
abstract
The modeling of complicated time-evolving physical dynamics from partial observations is a long-standing challenge. Particularly, observations can be sparsely distributed in a seemingly random or unstructured manner, making it difficult to capture highly nonlinear features in a variety of scientific and engineering problems. However, existing data-driven approaches are often constrained by fixed spatial and temporal discretization. While some researchers attempt to achieve spatio-temporal continuity by designing novel strategies, they either overly rely on traditional numerical methods or fail to truly overcome the limitations imposed by discretization. To address these, we propose CoPS, a purely data-driven methods, to effectively model continuous physics simulation from partial observations. Specifically, we employ multiplicative filter network to fuse and encode spatial information with the corresponding observations. Then we customize geometric grids and use message-passing mechanism to map features from original spatial domain to the customized grids. Subsequently, CoPS models continuous-time dynamics by designing multi-scale graph ODEs, while introducing a Markov-based neural auto-correction module to assist and constrain the continuous extrapolations. Comprehensive experiments demonstrate that CoPS advances the state-of-the-art methods in space-time continuous modeling across various scenarios. The source code is available at~\url{https://github.com/Sunxkissed/CoPS}.
Fan Xu 0009, Hao Wu 0094, Nan Wang 0015, Lilan Peng, Kun Wang 0042, Wei Gong 0001, Xibin Zhao
NeurIPS5
2025 Tri-Perspective View Decomposition for Geometry Aware Depth Completion and Super-Resolution
abstract
Depth completion and super-resolution are crucial tasks for comprehensive RGB-D scene understanding, as they involve reconstructing the precise 3D geometry of a scene from sparse or low-resolution depth measurements. However, most existing methods either rely solely on 2D depth representations or directly incorporate raw 3D point clouds for compensation, which are still insufficient to capture the fine-grained 3D geometry of the scene. In this paper, we introduce Tri-Perspective View Decomposition (TPVD) frameworks that can explicitly model 3D geometry. To this end, (1) TPVD ingeniously decomposes the original 3D point cloud into three 2D views, one of which corresponds to the sparse or low-resolution depth input. (2) For sufficient geometric interaction, TPV Fusion is designed to update the 2D TPV features through recurrent 2D-3D-2D aggregation. (3) By adaptively searching for TPV affinitive neighbors, two additional refinement heads are developed for these two tasks to further improve the geometric consistency. Meanwhile, we build novel datasets named TOFDC for depth completion and TOFDSR for depth super-resolution. Both datasets are acquired using time-of-flight (TOF) sensors and color cameras on smartphones. Extensive experiments on TOFDC, KITTI, NYUv2, SUN RGBD, VKITTI, TOFDSR, RGB-D-D, Lu, and Middlebury datasets indicate that our TPVD outperforms previous depth completion and super-resolution methods, reaching the state of the art.
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Guangwei Gao, Jun Li 0027, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 AltNeRF: Learning Robust Neural Radiance Field via Alternating Depth-Pose Optimization
abstract
Neural Radiance Fields (NeRF) have shown promise in generating realistic novel views from sparse scene images. However, existing NeRF approaches often encounter challenges due to the lack of explicit 3D supervision and imprecise camera poses, resulting in suboptimal outcomes. To tackle these issues, we propose AltNeRF---a novel framework designed to create resilient NeRF representations using self-supervised monocular depth estimation (SMDE) from monocular videos, without relying on known camera poses. SMDE in AltNeRF masterfully learns depth and pose priors to regulate NeRF training. The depth prior enriches NeRF's capacity for precise scene geometry depiction, while the pose prior provides a robust starting point for subsequent pose refinement. Moreover, we introduce an alternating algorithm that harmoniously melds NeRF outputs into SMDE through a consistence-driven mechanism, thus enhancing the integrity of depth priors. This alternation empowers AltNeRF to progressively refine NeRF representations, yielding the synthesis of realistic novel views. Extensive experiments showcase the compelling capabilities of AltNeRF in generating high-fidelity and robust novel views that closely resemble reality.
Kun Wang 0042, Zhiqiang Yan 0001, Huang Tian, Zhenyu Zhang 0005, Xiang Li 0041, Jun Li 0027, Jian Yang 0003
AAAI1
2024 Driving-Video Dehazing with Non-Aligned Regularization for Safety Assistance
abstract
Real driving-video dehazing poses a significant challenge due to the inherent difficulty in acquiring precisely aligned hazy/clear video pairs for effective model training, especially in dynamic driving scenarios with unpredictable weather conditions. In this paper, we propose a pioneering approach that addresses this challenge through a nonaligned regularization strategy. Our core concept involves identifying clear frames that closely match hazy frames, serving as references to supervise a video dehazing network. Our approach comprises two key components: reference matching and video dehazing. Firstly, we introduce a non-aligned reference frame matching module, leveraging an adaptive sliding window to match high-quality reference frames from clear videos. Video dehazing incorporates flow-guided cosine attention sampler and deformable cosine attention fusion modules to enhance spatial multi-frame alignment and fuse their improved information. To validate our approach, we collect a GoProHazy dataset captured effortlessly with GoPro cameras in diverse rural and urban road environments. Extensive experiments demonstrate the superiority of the proposed method over current state-of-the-art methods in the challenging task of real driving-video dehazing. Project page.
Junkai Fan, Jiangwei Weng, Kun Wang 0042, Jianjun Qian, Jun Li 0027, Jian Yang 0003
CVPR3
2024 Tri-Perspective view Decomposition for Geometry-Aware Depth Completion
abstract
Depth completion is a vital taskfor autonomous driving, as it involves reconstructing the precise 3D geometry of a scene from sparse and noisy depth measurements. How-ever, most existing methods either rely only on 2D depth representations or directly incorporate raw 3D point clouds for compensation, which are still insufficient to capture the fine-grained 3D geometry of the scene. To address this chal-lenge, we introduce Tri-Perspective View Decomposition (TPVD), a novel framework that can explicitly model 3D geometry. In particular, (1) TPVD ingeniously decomposes the original point cloud into three 2D views, one of which corresponds to the sparse depth input. (2) We design TPV Fusion to update the 2D TPV features through recurrent 2D-3D-2D aggregation, where a Distance-Aware Spherical Convolution (DASC) is applied. (3) By adaptively choosing TPVaffinitive neighbors, the newly proposed Geometric Spatial Propagation Network (GSPN) further improves the geometric consistency. As a result, our TPVD outperforms existing methods on KITTI, NYUv2, and SUN RGBD. Fur-thermore, we build a novel depth completion dataset named TOFDC, which is acquired by the time-of-flight (TOF) sen-sor and the color camera on smart phones. Project page.
Zhiqiang Yan 0001, Yuankai Lin, Kun Wang 0042, Yupeng Zheng, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
CVPR3
2024 Towards Robust Trajectory Representations: Isolating Environmental Confounders with Causal Learning
Yuanshao Zhu, Wei Chen 0070, Kun Wang 0042, Zhengyang Zhou, Sijie Ruan, Yuxuan Liang 0002
IJCAI4
2024 Predicting Carpark Availability in Singapore with Cross-Domain Data: A New Dataset and A Data-Driven Approach
Huaiwu Zhang, Yutong Xia, Siru Zhong, Kun Wang 0042, Zekun Tong, Qingsong Wen, Roger Zimmermann, Yuxuan Liang 0002
IJCAI4
2024 DCDepth: Progressive Monocular Depth Estimation in Discrete Cosine Domain
abstract
In this paper, we introduce DCDepth, a novel framework for the long-standing monocular depth estimation task. Moving beyond conventional pixel-wise depth estimation in the spatial domain, our approach estimates the frequency coefficients of depth patches after transforming them into the discrete cosine domain. This unique formulation allows for the modeling of local depth correlations within each patch. Crucially, the frequency transformation segregates the depth information into various frequency components, with low-frequency components encapsulating the core scene structure and high-frequency components detailing the finer aspects. This decomposition forms the basis of our progressive strategy, which begins with the prediction of low-frequency components to establish a global scene context, followed by successive refinement of local details through the prediction of higher-frequency components. We conduct comprehensive experiments on NYU-Depth-V2, TOFDC, and KITTI datasets, and demonstrate the state-of-the-art performance of DCDepth. Code is available at https://github.com/w2kun/DCDepth.
Kun Wang 0042, Zhiqiang Yan 0001, Junkai Fan, Wanlu Zhu, Xiang Li 0041, Jun Li 0027, Jian Yang 0003
NeurIPS1
2024 Learning Complementary Correlations for Depth Super-Resolution With Incomplete Data in Real World
abstract
Depth information is a significant ingredient to visually perceive the physical world. However, mainstream depth sensors, e.g., time-of-flight (ToF) cameras, often measure incomplete and low-resolution depth data, resulting in low-quality visual perception. In this article, we try to address a potentially valuable task, i.e., depth super-resolution (DSR) with incomplete data, which recovers dense and high-resolution depth map from incomplete and low-resolution one. To tackle this task, we introduce a novel incomplete DSR (IDSR) framework, including a primary branch for DSR to recover high-frequency details, and an auxiliary branch for depth completion (DC) to fill missing pixels. More importantly, we propose two modules, joint correlation learning (JCL) and iterative-cross (IC), to enhance the learning of complementary information flows between the two branches. The former module aims to learn the correlative relationships of the two branches, whilst the latter module adequately fuses higher level representations for more precise predictions. Extensive experiments show that our framework is effective and achieves the state-of-the-art performance on the real-world RGB-D-D and the synthetic NYUv2 datasets.
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
IEEE Trans. Neural Networks Learn. Syst.2
2023 DesNet: Decomposed Scale-Consistent Network for Unsupervised Depth Completion
abstract
Unsupervised depth completion aims to recover dense depth from the sparse one without using the ground-truth annotation. Although depth measurement obtained from LiDAR is usually sparse, it contains valid and real distance information, i.e., scale-consistent absolute depth values. Meanwhile, scale-agnostic counterparts seek to estimate relative depth and have achieved impressive performance. To leverage both the inherent characteristics, we thus suggest to model scale-consistent depth upon unsupervised scale-agnostic frameworks. Specifically, we propose the decomposed scale-consistent learning (DSCL) strategy, which disintegrates the absolute depth into relative depth prediction and global scale estimation, contributing to individual learning benefits. But unfortunately, most existing unsupervised scale-agnostic frameworks heavily suffer from depth holes due to the extremely sparse depth input and weak supervisory signal. To tackle this issue, we introduce the global depth guidance (GDG) module, which attentively propagates dense depth reference into the sparse target via novel dense-to-sparse attention. Extensive experiments show the superiority of our method on outdoor KITTI, ranking 1st and outperforming the best KBNet more than 12% in RMSE. Additionally, our approach achieves state-of-the-art performance on indoor NYUv2 benchmark as well.
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
AAAI2
2023 Distortion and Uncertainty Aware Loss for Panoramic Depth Completion
abstract
Standard MSE or MAE loss function is commonly used in limited field-of-vision depth completion, treating each pixel equally under a basic assumption that all pixels have same contribution during optimization. Recently, with the rapid rise of panoramic photography, panoramic depth completion (PDC) has raised increasing attention in 3D computer vision. However, the assumption is inapplicable to panoramic data due to its latitude-wise distortion and high uncertainty nearby textures and edges. To handle these challenges, we propose distortion and uncertainty aware loss (DUL) that consists of a distortion-aware loss and an uncertainty-aware loss. The distortion-aware loss is designed to tackle the panoramic distortion caused by equirectangular projection, whose coordinate transformation relation is used to adaptively calculate the weight of the latitude-wise distortion, distributing uneven importance instead of the equal treatment for each pixel. The uncertainty-aware loss is presented to handle the inaccuracy in non-smooth regions. Specifically, we characterize uncertainty into PDC solutions under Bayesian deep learning framework, where a novel consistent uncertainty estimation constraint is designed to learn the consistency between multiple uncertainty maps of a single panorama. This consistency constraint allows model to produce more precise uncertainty estimation that is robust to feature deformation. Extensive experiments show the superiority of our method over standard loss functions, reaching the state of the art.
Zhiqiang Yan 0001, Xiang Li 0041, Kun Wang 0042, Shuo Chen 0003, Jun Li 0027, Jian Yang 0003
ICML3
2023 Deciphering Spatio-Temporal Graph Forecasting: A Causal Lens and Treatment
abstract
Spatio-Temporal Graph (STG) forecasting is a fundamental task in many real-world applications. Spatio-Temporal Graph Neural Networks have emerged as the most popular method for STG forecasting, but they often struggle with temporal out-of-distribution (OoD) issues and dynamic spatial causation. In this paper, we propose a novel framework called CaST to tackle these two challenges via causal treatments. Concretely, leveraging a causal lens, we first build a structural causal model to decipher the data generation process of STGs. To handle the temporal OoD issue, we employ the back-door adjustment by a novel disentanglement block to separate the temporal environments from input data. Moreover, we utilize the front-door adjustment and adopt edge-level convolution to model the ripple effect of causation. Experiments results on three real-world datasets demonstrate the effectiveness of CaST, which consistently outperforms existing methods with good interpretability. Our source code is available at https://github.com/yutong-xia/CaST.
Yutong Xia, Yuxuan Liang 0002, Haomin Wen, Xu Liu 0014, Kun Wang 0042, Zhengyang Zhou, Roger Zimmermann
NeurIPS5
2022 Multi-modal Masked Pre-training for Monocular Panoramic Depth Completion
Zhiqiang Yan 0001, Xiang Li 0041, Kun Wang 0042, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
ECCV (1)3
2022 RigNet: Repetitive Image Guided Network for Depth Completion
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
ECCV (27)2
2021 Regularizing Nighttime Weirdness: Efficient Self-supervised Monocular Depth Estimation in the Dark
abstract
Monocular depth estimation aims at predicting depth from a single image or video. Recently, self-supervised methods draw much attention since they are free of depth annotations and achieve impressive performance on several daytime benchmarks. However, they produce weird outputs in more challenging nighttime scenarios because of low visibility and varying illuminations, which bring weak textures and break brightness-consistency assumption, respectively. To address these problems, in this paper we propose a novel framework with several improvements: (1) we introduce Priors-Based Regularization to learn distribution knowledge from unpaired depth maps and prevent model from being incorrectly trained; (2) we leverage Mapping-Consistent Image Enhancement module to enhance image visibility and contrast while maintaining brightness consistency; and (3) we present Statistics-Based Mask strategy to tune the number of removed pixels within textureless regions, using dynamic statistics. Experimental results demonstrate the effectiveness of each component. Mean-while, our framework achieves remarkable improvements and state-of-the-art results on two nighttime datasets. Code is available at https://github.com/w2kun/RNW.
Kun Wang 0042, Zhenyu Zhang 0005, Zhiqiang Yan 0001, Xiang Li 0041, Baobei Xu, Jun Li 0027, Jian Yang 0003
ICCV1