EDBT 2026 Demo / reviewers in the wild / expert
Junkang Zhang
dblp:190/8745
· DBLP profile ↗
22ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physics-Aware Accelerated Unrolling Model for Sparse-View CT ReconstructionabstractDeep unrolling models (DUMs) have shown great poten-tial in sparse-view CT reconstruction by combining itera-tive optimization and deep learning. However, most DUMsinsufficiently account for physical degradation from sparse-view imaging, leading to slow convergence and persistentartifacts. To address this, we propose PAUM, a Physics-Aware Accelerated Unrolling Model explicitly incorporatingCT imaging physics into the iterative reconstruction. PAUMfirst introduces a Dual-Domain Physics-Aware Extrapolation(DDPE) module. By modeling dual-domain degradations, itperforms row-wise extrapolation in the sinogram domain toimprove missing view recovery, and pixel-wise extrapolationin the image domain to address spatially variant degradationfrom incomplete backprojection. This physics-aware extrap-olation aligns optimization dynamics with underlying physi-cal imaging degradation, significantly enhances structural up-dates, thereby accelerating convergence. Subsequently, wedevelop a lightweight Block-Attention Deformable Regu-larization Network (BDRN), leveraging deformable convo-lutions and block-wise attention to model spatially variantand structured artifact physical characteristics. This enablesspatially adaptive regularization on extrapolated results, ef-fectively improving reconstruction quality. Extensive exper-iments demonstrate PAUM achieves over 1dB improvementcompared to SOTA methods, while reducing iteration countby 50%. Shaojie Guo, Yingying Fang, Junkang Zhang, Yan Wang 0033 |
AAAI | 3 |
| 2026 | Deep Algorithm Unrolling with Alignment Embedding for Guided Image Super-resolution
Faming Fang, Tingting Wang 0007, Junkang Zhang, Aimin Zhou, Riquan Zhang, Guixu Zhang |
Int. J. Comput. Vis. | 3 |
| 2026 | Task-aware all-in-one guided image super-resolution
Tingting Wang 0007, Jun Wang 0024, Qiuhai Yan, Junkang Zhang, Faming Fang, Guixu Zhang |
Pattern Recognit. | 4 |
| 2025 | Decoupling Scattering: Pseudo-Label Guided NeRF for Scenes with Scattering MediaabstractNeural Radiance Fields (NeRF) has been widely used in computer vision and graphics, achieving impressive results in novel view synthesis and multi-view 3D reconstruction. However, despite its excellent performance under ideal conditions, NeRF struggles in challenging environments such as hazy, foggy, and underwater scenes, primarily due to the difficulty in decoupling objects from the scattering medium. To mitigate this limitation, we proposed a novel approach for NeRF in scenes with scattering media. Specifically, we leverage pseudo-labels during the early stage of training to guide NeRF in decoupling the densities of objects and the scattering medium, guiding the model toward a more appropriate search space. Furthermore, we introduce a Cyclical Progressive Dimensional Optimization Strategy (CPDOS) that focuses on optimizing a single or a few variables during specific periods. Experimental results demonstrate that our method can effectively simulate hazy and underwater scenes, accurately decouple the scattering medium from objects, estimate atmospheric parameters, and outperform existing methods in novel view synthesis and image restoration tasks. Junkang Zhang, Faming Fang, Guixu Zhang |
AAAI | 2 |
| 2025 | Explicit Depth-Aware Blurry Video Frame Interpolation Guided by Differential CurvesabstractBlurry video frame interpolation (BVFI), which aims to generate high-frame-rate clear videos from low-frame-rate blurry inputs, is a challenging yet significant task in computer vision. Current state-of-the-art approaches typically rely on linear or quadratic models to estimate intermediate motion. However, these methods often overlook depth variations that occur during fast object motion, leading to changes in object size and hindering interpolation performance.This paper proposes the Differential Curves-guided Blurry Video Frame Interpolation (DC-BVFI) framework, which leverages the differential curves theory to analyze and mitigate the effects of depth variations caused by object motion. Specifically, DC-BVFI consists of UBNet and MPNet. Unlike prior approaches that rely on optical flow for frame interpolation, MPNet is designed to estimate the 3D scene flow, which facilitates a more precise awareness of depth and velocity variations. Since scene flow cannot be directly inferred in the 2D frame space, UBNet is introduced to transform them into 3D point maps. Extensive experiments demonstrate that the proposed DC-BVFI framework surpasses state-of-the-art performance in simulated and real-world datasets. Zaoming Yan, Pengcheng Lei, Tingting Wang 0007, Faming Fang, Junkang Zhang, Yaomin Huang |
CVPR | 5 |
| 2025 | FrDiff: Framelet-Based Conditional Diffusion Model for Multispectral and Panchromatic Image FusionabstractThe process of fusing low-resolution multispectral (LRMS) and high-resolution panchromatic (PAN) imagery, commonly referred to as pansharpening, is intended to generate high-resolution multispectral (HRMS) imagery. Typically, most pre-existing pansharpening frameworks mainly emphasize the straightforward learning of the mapping relationship among PAN and LRMS images to HRMS images. However, a key limitation of these frameworks is their potential overemphasis on spatial information, particularly the enhancement of low-frequency components. As a result, such an oversight potentially hinders the model's ability to simultaneously restore both spectral and spatial details. To address this issue, we propose a novel pansharpening model based on the denoising diffusion probabilistic model (DDPM), dubbed FrDiff. Specifically, we build a framelet-based conditional diffusion model that leverages the generative power of diffusion models to produce more refine results. Different from conventional methods directly inferring HRMS images, our strategy is designed to project their framelet coefficients, utilizing the available PAN and LRMS images as resources. This approach enables the separation of high-frequency and low-frequency components through framelet transformation, which are subsequently recombined to create a novel set of conditional embeddings that feed into the diffusion process. At the same time, the powerful predictive power of the diffusion model is exploited to simultaneously recover the high-frequency and low-frequency components of the HRMS. Moreover, we introduce a framelet-oriented cross-attention module dedicated to honing spectral fidelity. This module is crucial for improving the spectral precision of the HRMS images, ensuring a balanced emphasis on both spatial and spectral enhancements. Quantitative and qualitative experiments on multiple benchmark datasets demonstrate that the proposed method achieves more robustness and high-quality results than other state-of-the-art pansharpening methods. Junkang Zhang, Faming Fang, Tingting Wang 0007, Guixu Zhang |
IEEE Trans. Multim. | 1 |
| 2024 | UGNet: Uncertainty aware geometry enhanced networks for stereo matching
Zhengkai Qi, Junkang Zhang, Faming Fang, Tingting Wang 0007, Guixu Zhang |
Pattern Recognit. | 2 |
| 2023 | Indoor Depth Recovery Based on Deep Unfolding with Non-Local PriorabstractIn recent years, depth recovery based on deep networks has achieved great success. However, the existing state-of-the-art network designs perform like black boxes in depth recovery tasks, lacking a clear mechanism. Utilizing the property that there is a large amount of non-local common characteristics in depth images, we propose a novel model-guided depth recovery method, namely the DC-NLAR model. A non-local auto-regressive regular term is also embedded into our model to capture more non-local depth information. To fully use the excellent performance of neural networks, we develop a deep image prior to better describe the characteristic of depth images. We also introduce an implicit data consistency term to tackle the degenerate operator with high heterogeneity. We then unfold the proposed model into networks by using the half-quadratic splitting algorithm. This proposed method is experimented on the NYU-Depth V2 and SUN RGB-D datasets, and the experimental results achieve comparable performance to that of deep learning methods. Yuhui Dai, Junkang Zhang, Faming Fang, Guixu Zhang |
ICCV | 2 |
| 2023 | Accurate Registration between Ultra-Wide-Field and Narrow Angle Retina Images with 3D Eyeball Shape OptimizationabstractThe Ultra-Wide-Field (UWF) retina images have attracted wide attentions in recent years in the study of retina. However, accurate registration between the UWF images and the other types of retina images could be challenging due to the distortion in the peripheral areas of an UWF image, which a 2D warping can not handle. In this paper, we propose a novel 3D distortion correction method which sets up a 3D projection model and optimizes a dense 3D retina mesh to correct the distortion in the UWF image. The corrected UWF image can then be accurately aligned to the target image using 2D alignment methods. The experimental results show that our proposed method outperforms the state-of-the-art method by 30%. Junkang Zhang, Fritz Gerald P. Kalaw, Melina Cavichini, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen, Cheolhong An |
ICIP | 1 |
| 2023 | DMCSC: Deep Multisource Convolutional Sparse Coding Model for PansharpeningabstractPansharpening aims to produce a high-resolution multispectral (HRMS) image by combining a low-resolution multispectral (LRMS) image with a high-resolution panchromatic (PAN) image through a fusion process. Deep learning (DL)-based pansharpening methods have demonstrated impressive results in generating high-quality HRMS images. However, they suffer from a lack of interpretability due to their black-box network architectures. Recently, model-based deep unrolling networks have been proposed to improve the interpretability of networks. Among these approaches, the multi-source convolutional sparse coding (MCSC)-based models stand out by effectively learning common and unique features from both LRMS and PAN images, showing promising results. As the LRMS image provides limited information in MCSC-based models, it can result in weak feature response and even lead to incorrect fusion outcomes. To address this issue, we propose a novel deep MCSC-based method that enhances the robustness and performance. Specifically, we build an optimization model that integrates MCSC with a degradation model and a deep prior, which can sufficiently capture the common information shared by the latent HRMS images and PAN images, thereby enabling the recovery of more accurate spectral information. To optimize the proposed model, we adopt an iterative optimization strategy that unfolds the iterative solution into networks. Moreover, we propose an enhanced version of our method that utilizes multi-scale dictionaries to capture common and unique features at different scales, thereby facilitating the extraction of more abundant spectral and spatial details. We evaluate the effectiveness of our proposed method on multiple benchmark datasets. Experiment results demonstrate its effectiveness in improving the robustness and performance of MCSC-based models. Junkang Zhang, Yongxu Ye, Faming Fang, Tingting Wang 0007, Guixu Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Self-Supervised Rigid Registration for Multimodal Retinal ImagesabstractThe ability to accurately overlay one modality retinal image to another is critical in ophthalmology. Our previous framework achieved the state-of-the-art results for multimodal retinal image registration. However, it requires human-annotated labels due to the supervised approach of the previous work. In this paper, we propose a self-supervised multimodal retina registration method to alleviate the burdens of time and expense to prepare for training data, that is, aiming to automatically register multimodal retinal images without any human annotations. Specially, we focus on registering color fundus images with infrared reflectance and fluorescein angiography images, and compare registration results with several conventional and supervised and unsupervised deep learning methods. From the experimental results, the proposed self-supervised framework achieves a comparable accuracy comparing to the state-of-the-art supervised learning method in terms of registration accuracy and Dice coefficient. Cheolhong An, Junkang Zhang, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2022 | Two-Step Registration on Multi-Modal Retinal Images via Deep Neural NetworksabstractMulti-modal retinal image registration plays an important role in the ophthalmological diagnosis process. The conventional methods lack robustness in aligning multi-modal images of various imaging qualities. Deep-learning methods have not been widely developed for this task, especially for the coarse-to-fine registration pipeline. To handle this task, we propose a two-step method based on deep convolutional networks, including a coarse alignment step and a fine alignment step. In the coarse alignment step, a global registration matrix is estimated by three sequentially connected networks for vessel segmentation, feature detection and description, and outlier rejection, respectively. In the fine alignment step, a deformable registration network is set up to find pixel-wise correspondence between a target image and a coarsely aligned image from the previous step to further improve the alignment accuracy. Particularly, an unsupervised learning framework is proposed to handle the difficulties of inconsistent modalities and lack of labeled training data for the fine alignment step. The proposed framework first changes multi-modal images into a same modality through modality transformers, and then adopts photometric consistency loss and smoothness loss to train the deformable registration network. The experimental results show that the proposed method achieves state-of-the-art results in Dice metrics and is more robust in challenging cases. Junkang Zhang, Ji Dai, Melina Cavichini, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen, Cheolhong An |
IEEE Trans. Image Process. | 1 |
| 2021 | Foveal Avascular Zone Segmentation of Octa Images Using Deep Learning Approach with Unsupervised Vessel SegmentationabstractFoveal Avascular Zone (FAZ) is a crucial indicator for retinal disease detection and accurate automatic FAZ segmentation has a significant impact in clinical applications. Apart from the binary FAZ segmentation map, a vessel segmentation map can provide further information. To simultaneously implement vessel and accurate FAZ segmentation, an end-to-end trained network is proposed to achieve unsupervised vessel segmentation and supervised FAZ segmentation. Due to the lack of vessel labels, the style transfer with consistency loss is proposed to the vessel segmentation. Then FAZ segmentation is achieved with a U-Net structure based on vessel segmentation. Two superficial layer OCTA image datasets - OCTAGON3 [1] and sFAZDATA datasets [2] - are used to evaluate the proposed method. We achieve the Dice scores of 0.9263 and 0.9784, which are better than those from other approaches. Zhijin Liang, Junkang Zhang, Cheolhong An |
ICASSP | 2 |
| 2021 | A Variational Model for Spatially Weighting in Image FusionabstractIn order to retain as many valuable details from the input source images as possible during the process of fusion, this paper proposes an adaptive weight based total variation model for image fusion. The main idea is to employ a nonconvex energy functional to determine simultaneously the output fused image and weight functions by maximizing the local variance of the output image and preserving the brightness of the input images. In order to minimize the differences among the weight functions at the nearby pixel locations, the total variation regularization of the weight functions is incorporated in the functional for the fusion process. The existence of minimizers to the proposed variational model is established. Furthermore, we develop an efficient algorithm to solve the model numerically by using the primal-dual method, and prove the convergence of the algorithm. Experimental results are reported to illustrate the effectiveness of the proposed method, and its performance is competitive with the other testing methods. Zhengmeng Jin, Junkang Zhang, Lihua Min, Michael Kwok-Po Ng |
SIAM J. Imaging Sci. | 2 |
| 2021 | Robust Content-Adaptive Global Registration for Multimodal Retinal Images Using Weakly Supervised Deep-Learning FrameworkabstractMultimodal retinal imaging plays an important role in ophthalmology. We propose a content-adaptive multimodal retinal image registration method in this paper that focuses on the globally coarse alignment and includes three weakly supervised neural networks for vessel segmentation, feature detection and description, and outlier rejection. We apply the proposed framework to register color fundus images with infrared reflectance and fluorescein angiography images, and compare it with several conventional and deep learning methods. Our proposed framework demonstrates a significant improvement in robustness and accuracy reflected by a higher success rate and Dice coefficient compared with other methods. Junkang Zhang, Melina Cavichini, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen, Cheolhong An |
IEEE Trans. Image Process. | 2 |
| 2020 | A Segmentation Based Robust Deep Learning Framework for Multimodal Retinal Image RegistrationabstractMultimodal image registration plays an important role in diagnosing and treating ophthalmologic diseases. In this paper, a deep learning framework for multimodal retinal image registration is proposed. The framework consists of a segmentation network, feature detection and description network, and an outlier rejection network, which focuses only on the globally coarse alignment step using the perspective transformation. We apply the proposed framework to register color fundus images with infrared reflectance images and compare it with the state-of-the-art conventional and learning-based approaches. The proposed framework demonstrates a significant improvement in robustness and accuracy reflected by a higher success rate and Dice coefficient compared to other coarse alignment methods. Junkang Zhang, Cheolhong An, Melina Cavichini, Mahima Jhingan, Manuel J. Amador-Patarroyo, Christopher P. Long, Dirk-Uwe Bartsch, William R. Freeman, Truong Q. Nguyen |
ICASSP | 2 |
| 2020 | Boosting Feature Matching Accuracy With Pairwise Affine EstimationabstractLocal image feature matching lies in the heart of many computer vision applications. Achieving high matching accuracy is challenging when significant geometric difference exists between the source and target images. The traditional matching pipeline addresses the geometric difference by introducing the concept of support region. Around each feature point, the support region defines a neighboring area characterized by estimated attributes like scale, orientation, affine shape, etc. To correctly assign support region is not an easy job, especially when each feature is processed individually. In this paper, we propose to estimate the relative affine transformation for every pair of to-be-compared features. This "tailored" measurement of geometric difference is more precise and helps improve the matching accuracy. Our pipeline can be incorporated into most existing 2D local image feature detectors and descriptors. We comprehensively evaluate its performance with various experiments on a diversified selection of benchmark datasets. The results show that the majority of tested detectors/descriptors gain additional matching accuracy with proposed pipeline. Ji Dai, Shiwei Jin, Junkang Zhang, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2019 | Explicit Learning of Feature Orientation EstimationabstractWhile many learning-driven algorithms for local feature detection and description have submerged during recent years. One key component in the pipeline, namely orientation estimation, still remains underdeveloped. Among all sorts of difficulties, the impracticality and tedium of finding a "ground truth" feature orientation as a learning target is one big challenge. In this paper, we bypass this "thinking trap" and propose an unsupervised scheme that explicitly trains a simple convolutional neural network to predict orientations for feature points. Together with a carefully designed loss term, the network manages to provide accurate orientation estimations. We further evaluate the capability of this estimator in two experiments: orientation estimation and feature matching. Results showed the proposed method outperforms other compared methods on multiple benchmark datasets. The pretrained model is publicly available. Ji Dai, Junkang Zhang, Truong Q. Nguyen |
ICIP | 2 |
| 2019 | Joint Vessel Segmentation and Deformable Registration on Multi-Modal Retinal Images Based on Style TransferabstractIn multi-modal retinal image registration task, there are two major challenges, i.e., poor performance in finding correspondence due to inconsistent features, and lack of labeled data for training learning-based models. In this paper, we propose a joint vessel segmentation and deformable registration model based on CNN for this task, built under the framework of weakly supervised style transfer learning and perceptual loss. In vessel segmentation, a style loss guides the model to generate segmentation maps that look authentic, and helps transform images of different modalities into consistent representations. In deformable registration, a content loss helps find dense correspondence for multi-modal images based on their consistent representations, and improves the segmentation results simultaneously. Experiment results show that our model has better performance than other deformable registration methods in both quantitative and visual evaluations, and the segmentation results also help the rigid transformation1. Junkang Zhang, Cheolhong An, Ji Dai, Manuel Amador, Dirk-Uwe Bartsch, Shyamanga Borooah, William R. Freeman, Truong Q. Nguyen |
ICIP | 1 |
| 2017 | Family Photo Recognition via Multiple Instance LearningabstractFamily photo recognition is an important task in social media analytics. Previous methods use singleton global features and conventional binary classifiers to distinguish family group photos from non-family ones. Different from them, we propose a novel family recognition approach with three dedicated local representations under Multiple Instance Learning framework, where geometry, kinship and semantic features are integrated to overcome issues in the previous work. Experimental results show that our method achieves the state-of-the-art result among global-feature models. Junkang Zhang, Si-Yu Xia, Ming Shao, Yun Fu 0001 |
ICMR | 1 |
| 2016 | A genetics-motivated unsupervised model for tri-subject kinship verificationabstractGiven a child's and a couple's facial photos, tri-subject kinship verification aims to determine the existence of blood relation between the child and the couple. Different from existing methods which model the kinship inheritance process among three persons in separate stages and only use simple features, this work establishes a simple model inspired by genetics to measure tri-subject kinship similarity in one step. Meanwhile, high-dimensional features are incorporated into this simple model to seek for better performance. Experiment results demonstrate the effectiveness of our approach. Junkang Zhang, Si-Yu Xia, Hong Pan 0001, A. K. Qin 0001 |
ICIP | 1 |
| 2016 | Robust road detection from a single imageabstractRoad detection from images is a challenging task in computer vision. Previous methods are not robust, because their features and classifiers cannot adapt to different circumstances. To overcome this problem, we propose to apply unsupervised feature learning for road detection. Specifically, we develop an improved encoding function and add a feature selection process to obtain robust and discriminative road features. Besides, a road segmentation algorithm is proposed to extract road regions from the learned feature maps, in which a tree structure is established to represent the hierarchical relations of various regions segmented by multiple thresholds, and a two-loop optimization is then employed to select the most stable regions as road areas. Experimental results on several challenging datasets justify the effectiveness of our method. Junkang Zhang, Si-Yu Xia, Kaiyue Lu, Hong Pan 0001, A. K. Qin 0001 |
ICPR | 1 |