Kang Han

dblp:178/7281 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Real-time trajectory monitoring system and positioning accuracy test for manned helicopter crop spraying
Jiangkun Xue, Yu Ru, Zifan Rong, Shuping Fang, Chenming Hu, Kang Han
Eng. Appl. Artif. Intell.7
2026 A Dynamic Differential Privacy Mechanism Based on Feature Importance in Deep Learning
abstract
The extensive adoption of deep learning, coupled with the exponential growth of data, has raised concerns regarding potential privacy disclosure, particularly through membership inference attacks where adversaries attempt to determine whether specific data samples were used in model training. Differential privacy has emerged as a prominent technique to mitigate these concerns. However, its application often results in degraded model performance and significant utility loss. This paper proposes a dynamic differential privacy mechanism based on feature importance in deep learning (DPFI) to address this issue.Meanwhile, the introduction of superpixel segmentation not only mitigates the trade-off between accuracy and utility but also reduces the high complexity caused by high-dimensional features. The core concept of DPFI is that the noise level for each feature is determined based on its importance to the model. Specifically, we first initialize and train a model with differential privacy. Then, we perform superpixel segmentation on the dataset and apply Shapley Additive Explanations on the segmented images to calculate the feature importance. Next, we propose a noise addition strategy based on the importance of the features and their distribution. The privacy guarantees are rigorously analyzed through Rényi differential privacy. Experiments demonstrate that DPFI outperforms existing methods in terms of both model accuracy and resistance to membership inference attacks.
Mi Wen, Hailun Shen, Xiumin Li, Kang Han, Kejie Lu
IEEE Internet Things J.4
2026 DeDiff-4DGS: Fusing Temporal Correlations and Diffusion Priors for Dynamic 3D Scenes
abstract
Reconstructing dynamic 3D (4D) scenes is challenging due to complex temporal dynamics and viewpoint sparsity in monocular videos. Existing extensions of 3D Gaussian Splatting (3D-GS) with its temporal modeling often fail to capture temporal correlations across frames, leading to redundant 3D Gaussians and reduced efficiency. To address this limitation, we propose DeDiff-4DGS, a framework that integrates temporal correlations and diffusion priors through two novel modules. The Temporal 3D Gaussian Latent Fusion (T3DLF) module fuses temporal information from sparse reference frames to promote spatio-temporal coherence and reduce the number of required 3D Gaussians. The Latent Diffusion Converter for 3D Gaussians (LDC3D) module enriches reference frames with semantic priors, complementing T3DLF under sparse-view conditions. Experimental results on standard benchmarks demonstrate that DeDiff-4DGS delivers higher reconstruction quality and improved efficiency over current state-of-the-art approaches.
Hoang Nguyen Nguyen, Wei Xiang 0001, Kang Han, Phu Lai, Tianyu Chen 0004, Yi-Ping Phoebe Chen
IEEE Trans. Multim.3
2025 RobSense: A Robust Multi-modal Foundation Model for Remote Sensing with Static, Temporal, and Incomplete Data Adaptability
abstract
Foundation models for remote sensing have garnered increasing attention for their strong performance across various observation tasks. However, current models lack robustness in managing diverse input types and handling incomplete data in downstream tasks. In this paper, we propose RobSense, a robust multi-modal foundation model for Multi-spectral and Synthetic Aperture Radar data. RobSense is designed with modular components and pre-trained by a combination of temporal multi-modal alignment and masked autoencoder strategies on a huge-scale dataset. Therefore, it can effectively support diverse input types, from static to temporal, uni-modal to multi-modal. To further handle the incomplete data, we incorporate two uni-modal latent reconstructors that recover rich representations from incomplete inputs, addressing variability in spectral bands and temporal sequence irregularities. Extensive experiments demonstrate that RobSense consistently outperforms state-of-the-art baselines on complete datasets across four input types for segmentation, classification, and change detection. On incomplete datasets, RobSense outperforms the baselines by considerably larger margins when the missing rate increases. Project page: https://ikhado.github.io/robsense/
Minh Kha Do, Kang Han, Phu Lai, Khoa T. Phan, Wei Xiang 0001
CVPR2
2025 Enhancing CLIP for Pedestrian Image-Text Retrieval via Bi-level Alignment and Weighted Similarity Distribution Matching Loss
Fumiaoyue Jia, HaoRan Bi, Kang Han, BinEr Zuo
ICIC (6)5
2025 CUOM: A causal unbiased optimization method for federated domain generalization
Mi Wen, Kang Han, Hailun Shen
Knowl. Based Syst.2
2025 Comp-Diff: A Unified Pruning and Distillation Framework for Compressing Diffusion Models
abstract
Recently, generative models such as diffusion models (DMs) have gained prominence in various applications, and there is a growing demand for their deployment on resource-constrained devices. Model pruning provides an effective solution by reducing the model redundancy without significantly impacting performance. However, most existing model pruning methods are designed for classification models and often lead to substantial performance degradation when applied to generative models. To address this issue, we propose Comp-Diff, a novel two-stage framework of pruning and knowledge distillation tailored for diffusion models. In the pruning stage, we propose a new structured content-aware pruning (CaP) method within Comp-Diff to identify and preserve informative units (filters/channels) that actually contribute to the generative capability of the model. Specifically, we introduce input perturbations to the pre-trained model and measure each unit’s importance score using gradients induced by these perturbations. Units with higher importance scores are considered more informative and are retained to maintain the model’s generative power. In the fine-tuning stage of Comp-Diff, we propose the distribution-aware knowledge distillation (DaKD) method, which effectively transfers fine-grained knowledge from the original model to the pruned one on both attention and noise distribution levels. In addition, DaKD includes an adversarial loss to improve the quality and diversity of generated outputs. To verify and evaluate our method, we apply the proposed Comp-Diff on three representative tasks: unconditional image generation, conditional image generation, and text-to-image generation. Extensive experiments on both multi-step and one-step diffusion models demonstrate that the proposed framework consistently yields compact models and outperforms existing pruning techniques by a large margin.
Wei Xiang 0001, Kang Han, Gaowen Liu, Ramana Rao Kompella
IEEE Trans. Multim.3
2024 Degradation-Aware Self-Attention Based Transformer for Blind Image Super-Resolution
abstract
Compared to CNN-based methods, Transformer-based methods achieve impressive image restoration outcomes due to their ability to model remote dependencies. However, how to apply Transformer-based methods to the field of blind super-resolution (SR) and further make an SR network adaptive to degradation information is still an open problem. In this paper, we propose a new degradation-aware self-attention-based Transformer model, where we incorporate contrastive learning into the Transformer network for learning the degradation representations of input images with unknown noise. In particular, we integrate both CNN and Transformer components into the SR network, where we first use the CNN modulated by the degradation information to extract local features, and then employ the degradation-aware Transformer to extract global semantic features. We apply our proposed model to several popular large-scale benchmark datasets for testing, and achieve the state-of-the-art performance compared to existing methods. In particular, our method yields a PSNR of 32.43 dB on the Urban100 dataset at ×2 scale, 0.94 dB higher than DASR, and 26.62 dB on the Urban100 dataset at ×4 scale, 0.26 dB improvement over KDSR, setting a new benchmark in this area. The source code is available at:https://github.com/I2-Multimedia-Lab/DSAT/tree/main.
Qingguo Liu, Pan Gao 0001, Kang Han, Ningzhong Liu, Wei Xiang 0001
IEEE Trans. Multim.3
2023 Multiscale Tensor Decomposition and Rendering Equation Encoding for View Synthesis
abstract
Rendering novel views from captured multi-view images has made considerable progress since the emergence of the neural radiance field. This paper aims to further advance the quality of view synthesis by proposing a novel approach dubbed the neural radiance feature field (NRFF). We first propose a multiscale tensor decomposition scheme to organize learnable features so as to represent scenes from coarse to fine scales. We demonstrate many benefits of the proposed multiscale representation, including more accurate scene shape and appearance reconstruction, and faster convergence compared with the single-scale representation. Instead of encoding view directions to model view-dependent effects, we further propose to encode the rendering equation in the feature space by employing the anisotropic spherical Gaussian mixture predicted from the proposed multiscale representation. The proposed NRFF improves state-of-the-art rendering results by over 1 dB in PSNR on both the NeRF and NSVF synthetic datasets. A significant improvement has also been observed on the real-world Tanks & Temples dataset. Code can be found at https://github.com/imkanghan/nrff.
Kang Han, Wei Xiang 0001
CVPR1
2023 Volume Feature Rendering for Fast Neural Radiance Field Reconstruction
abstract
Neural radiance fields (NeRFs) are able to synthesize realistic novel views from multi-view images captured from distinct positions and perspectives. In NeRF's rendering pipeline, neural networks are used to represent a scene independently or transform queried learnable feature vector of a point to the expected color or density. With the aid of geometry guides either in the form of occupancy grids or proposal networks, the number of color neural network evaluations can be reduced from hundreds to dozens in the standard volume rendering framework. However, many evaluations of the color neural network are still a bottleneck for fast NeRF reconstruction. This paper revisits volume feature rendering (VFR) for the purpose of fast NeRF reconstruction. The VFR integrates the queried feature vectors of a ray into one feature vector, which is then transformed to the final pixel color by a color neural network. This fundamental change to the standard volume rendering framework requires only one single color neural network evaluation to render a pixel, which substantially lowers the high computational complexity of the rendering framework attributed to a large number of color neural network evaluations. Consequently, we can use a comparably larger color neural network to achieve a better rendering quality while maintaining the same training and rendering time costs. This approach achieves the state-of-the-art rendering quality on both synthetic and real-world datasets while requiring less training time compared with existing methods.
Kang Han, Wei Xiang 0001
NeurIPS1
2022 A Novel Occlusion-Aware Vote Cost for Light Field Depth Estimation
abstract
Capturing the directions of light by light field cameras powers next-generation immersive multimedia applications. A critical problem in taking advantage of the rich visual information in light field images is depth estimation. Conventional light field depth estimation methods build a cost volume that measures the photo-consistency of pixels refocused to a range of depths, and the highest consistency indicates the correct depth. This strategy works well in most regions but usually generates blurry edges in the estimated depth map due to occlusions. Recent work shows that integrating occlusion models to light field depth estimation can largely reduce blurry edges. However, existing occlusion handling methods rely on complex edge-aided processing and post-refinement, and this reliance limits the resultant depth accuracy and impacts on the computational performance. In this paper, we propose a novel occlusion-aware vote cost (OAVC) which is able to accurately preserve edges in the depth map. Instead of using photo-consistency as an indicator of the correct depth, we construct a novel cost from a new perspective that counts the number of refocused pixels whose deviations from the central-view pixel are less than a small threshold, and utilizes that number to select the correct depth. The pixels from occluders are thus excluded in determining the correct depth. Without the use of any explicit occlusion handling methods, the proposed method can inherently preserve edges and produces high-quality depth estimates. Experimental results show that the proposed OAVC outperforms state-of-the-art light field depth estimation methods in terms of depth estimation accuracy and computational complexity.
Kang Han, Wei Xiang 0001, Eric Wang 0001, Tao Huang 0008
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Inference-Reconstruction Variational Autoencoder for Light Field Image Reconstruction
abstract
Light field cameras can capture the radiance and direction of light rays by a single exposure, providing a new perspective to photography and 3D geometry perception. However, existing sub-aperture based light field cameras are limited by their sensor resolution to obtain high spatial and angular resolution images simultaneously. In this paper, we propose an inference-reconstruction variational autoencoder (IR-VAE) to reconstruct a dense light field image out of four corner reference views in a light field image. The proposed IR-VAE is comprised of one inference network and one reconstruction network, where the inference network infers novel views from existing reference views and viewpoint conditions, and the reconstruction network reconstructs novel views from a latent variable that contains the information of reference views, novel views, and viewpoints. The conditional latent variable in the inference network is regularized by the latent variable in the reconstruction network to facilitate information flow between the conditional latent variable and novel views. We also propose a statistic distance measurement dubbed the mean local maximum mean discrepancy (MLMMD) to enable the measurement of the statistic distance between two distributions with high-resolution latent variables, which can capture richer information than their low-resolution counterparts. Finally, we propose a viewpoint-dependent indirect view synthesis method to synthesize novel views more efficiently by leveraging adaptive convolution. Experimental results show that our proposed methods outperform state-of-the-art methods on different light field datasets.
Kang Han, Wei Xiang 0001
IEEE Trans. Image Process.1
2021 Deep Adaptive Blending Network for 3D Magnetic Resonance Image Denoising
abstract
The visual quality of magnetic resonance images (MRIs) is crucial for clinical diagnosis and scientific research. The main source of quality degradation is the noise generated during MRI acquisition. Although denoising MRI by deep learning methods shows great superiority compared with traditional methods, the deep learning methods reported to date in the literature cannot simultaneously leverage long-range and hierarchical information, and cannot adequately utilize the similarity in 3D MRI. In this paper, we address the two issues by proposing a deep adaptive blending network (DABN) characterized by a large receptive field residual dense block and an adaptive blending method. We first propose the large receptive field residual dense block that can capture long-range information and fuse hierarchical features simultaneously. Then we propose the adaptive blending method that produces denoised pixels by adaptively filtering 3D MRI, which explicitly utilizes the similarity in 3D MRI. Residual is also considered as a compensating item after adaptive filtering. The blending adaptive filter and residual are predicted by a network consisting of several large receptive field residual dense blocks. Experimental results show that the proposed DABN outperforms state-of-the-art denoising methods in both clinical and simulated MRI data.
Kang Han, Yongming Zhou, Wei Xiang 0001
IEEE J. Biomed. Health Informatics2