VLDB 2026 Research / reviewers in the wild / expert
Yakun Ju
dblp:221/9647
· DBLP profile ↗
41ranked-venue papers
14as first author
37since 2021 · last 2026
0000-0003-4065-4108ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 9 first-author · 20 since 2021Artificial intelligence and machine learning · 15 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PUFM: Efficient Point Cloud Upsampling via Flow MatchingabstractDiffusion models have recently been adopted for point cloud upsampling due to their effectiveness in solving ill-posed problems. However, existing upsampling methods often struggle with inefficiencies, as they generate dense point clouds by mapping Gaussian noise to data, overlooking the geometric information already present in sparse inputs. To address this, we propose PUFM, a novel Point Cloud Upsampling via Flow Matching, which learns to directly transform sparse point clouds into their high-fidelity dense counterparts. Our approach first applies midpoint interpolation to densify the sparse input. Then, we construct a continuous interpolant between sparse and dense point clouds and train a neural network to estimate the velocity field for flow matching. Given the unordered nature of point clouds, we introduce a pre-alignment step based on Earth Mover's Distance (EMD) optimization to ensure coherent and meaningful interpolation between sparse and dense representations. This results in a more stable and efficient learning trajectory during flow matching. Experiments on synthetic benchmarks demonstrate that our method delivers superior upsampling quality but with fewer sampling steps. Further experiments on ScanNet and KITTI also show that our approach generalizes well to real-world RGB-D and LiDAR point clouds, making it more practical for real-world applications. Chenhang He, Yakun Ju, Lei Li 0038 |
AAAI | 3 |
| 2026 | EHIN: Early-aware hierarchical interaction network for weakly-supervised referring image segmentation
Anqing Chen, Wanli Ma 0001, Weide Liu, Yakun Ju, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
Neurocomputing | 7 |
| 2026 | Enhancing decision boundaries in continual learning through a decoupled Gaussian frameworkabstractThe goal of continual learning (CL) is to acquire new knowledge while retaining previously learned information. CNN-based and prompt-based CL methods have achieved remarkable progress in recent years. However, most prior work has primarily focused on reducing forgetting from the perspective of the model itself. In this paper, we investigate CL from the perspective of decision boundaries, analyzing the impact of instance-level feature overlap. To address this issue, we propose a generic Decoupled Gaussian Softmax Classifier that enhances class discriminability during CL process. Specifically, we decouple the features extracted by the backbone into multiple Gaussian distributions, which are directly fused into the feature space through weighted integration. A regularization term is introduced to penalize the overlap of similar features, while an adaptive decision boundary is assigned to each class to encourage inter-class separation and intra-class compactness. Experiments on 4 widely used continual learning datasets and 12 CL scenarios show that our method has good plug-and-play capability. It improves the average accuracy by 1%–2.63% over the baseline models, while effectively reducing both the forgetting rate and the Expected Calibration Error. Our code is available at: https://anonymous.4open.science/r/DGSC-main-310D . Zhikun Feng, Liu Yu 0001, Ping Kuang, Mian Zhou, Kang Dang, Yakun Ju |
Inf. Process. Manag. | 9 |
| 2026 | Clinically-informed prompt learning for explainable diagnosis with biomedical vision-language models
Yu-Cheng Fan, Ngai-Fong Law, Sai-Ho Ling, Yakun Ju |
Pattern Recognit. | 7 |
| 2026 | Dual-gated transformer with local context aggregation for weakly-supervised medical image anomaly detection
Fan Zhang 0045, Demin Liu, Yakun Ju |
Pattern Recognit. | 3 |
| 2026 | Perception-Inspired Network for Stereo Image Quality AssessmentabstractExisting stereo image quality assessment (SIQA) methods generally have limitations in binocular fusion and fine-grained perception modeling. To address these issues, we propose a Perception-Inspired Network for SIQA that simulates binocular difference-guided fusion, high-frequency sensitivity, and hierarchical perception mechanisms of the human visual system (HVS). First, a difference-guided binocular fusion (DGBF) module is designed to mimic the binocular difference sensitivity mechanism, which exploits difference information at both the feature-level and image-level to optimize binocular fusion. Furthermore, the image distortion primarily affects the high-frequency components, which are critical for perceptual quality. To reflect this, we propose a high-frequency enhancement module (HFEM) to simulate the human eye's sensitivity to edge and texture distortions. Finally, to better achieve fine-grained perception modeling, we propose a hierarchical quality regression strategy that simulates the human perceptual process, from perceiving local details to forming a global quality judgment, thereby achieving a quality prediction more aligned with human subjective evaluation. Experimental results demonstrate that the proposed method outperforms mainstream approaches, achieving a PLCC of 0.9734 on the LIVE I database, and a PLCC of 0.9632 on the LIVE II database. Yongli Chang, Guanghui Yue 0001, Li Yu 0004, Yakun Ju, Hadi Amirpour, Moncef Gabbouj, Wei Zhou 0021 |
IEEE Trans. Image Process. | 5 |
| 2026 | DCD-UIE: Decoupled Chromatic Diffusion Model for Underwater Image EnhancementabstractColor distortion and structural degradation in underwater images are classic challenges in underwater image enhancement. The core goal is to restore degraded images to high-quality images with both color and structure that conform to visual perception. However, in the traditional RGB space, these two issues are highly coupled, resulting in existing enhancement methods often neglecting one over the other. To address this challenge, we propose a guided diffusion model based on the principle of decoupling. Our key insight is that in perceptual color spaces such as HSV, color (H, S) and structure (V) are naturally separated. To exploit this property, we first design an adaptive perceptual guidance module, which analyzes the degraded HSV image and generates two orthogonal guidance signals: a color guide and a structure guide, which guide the denoising process of the diffusion model. To ensure that this decoupled guidance is faithfully implemented, we propose a corresponding decoupled loss optimization module, which uses independent loss functions to supervise the final output color and structure. By combining the forward decoupled guidance with the backward decoupled supervision, we construct a closed-loop optimization framework. This framework enables the model to collaboratively optimize color and structure under various degradation scenarios. Extensive experiments demonstrate that our proposed method outperforms existing state-of-the-art approaches in a variety of underwater scenes, particularly those degraded by color casts and haze. Furthermore, it exhibits superior performance on no-reference image quality assessment metrics. The source code is available at https://github.com/zy-world/DCD-UIE. Jingchun Zhou, Yakun Ju, Guang-Yong Chen, Jinjiang Li 0001, Alex Chichung Kot |
IEEE Trans. Image Process. | 4 |
| 2025 | FNIN: A Fourier Neural Operator-based Numerical Integration Network for Surface-from-gradientsabstractSurface-from-gradients (SfG) aims to recover a three-dimensional (3D) surface from its gradients. Traditional methods encounter significant challenges in achieving high accuracy and handling high-resolution inputs, particularly facing the complex nature of discontinuities and the inefficiencies associated with large-scale linear solvers. Although recent advances in deep learning, such as photometric stereo, have enhanced normal estimation accuracy, they do not fully address the intricacies of gradient-based surface reconstruction. To overcome these limitations, we propose a Fourier neural operator-based Numerical Integration Network (FNIN) within a two-stage optimization framework. In the first stage, our approach employs an iterative architecture for numerical integration, harnessing an advanced Fourier neural operator to approximate the solution operator in Fourier space. Additionally, a self-learning attention mechanism is incorporated to effectively detect and handle discontinuities. In the second stage, we refine the surface reconstruction by formulating a weighted least squares problem, addressing the identified discontinuities rationally. Extensive experiments demonstrate that our method achieves significant improvements in both accuracy and efficiency compared to current state-of-the-art solvers. This is particularly evident in handling high-resolution images with complex data, achieving errors of fewer than 0.1 mm on tested objects. Jiaqi Leng 0002, Yakun Ju, Yuanxu Duan, Jiangnan Zhang, Qingxuan Lv, Zuxuan Wu, Hao Fan 0004 |
AAAI | 2 |
| 2025 | HRHuman: Tuning-Free Higher-Resolution Human Image Generation via Template KnowledgeabstractHigh-resolution human-centric image generation offers significant potential across various industries, such as entertainment, media, and fashion. Diffusion models for text-to-image generation have significantly improved the quality of human image synthesis. However, when scaling to higher resolutions (2K, 4K, and above), they often encounter issues such as object repetition and structural distortion, which appear especially unnatural in human images. To address these challenges, we propose HRHuman, a tuning-free framework for Higher-Resolution Human Image Generation. By leveraging an open-source large human vision model that incorporates rich template knowledge as prior, we first introduce the Prompt Discretization scheme to discretize user-input text prompts, mapping them to image elements and human body parts. Additionally, we implement a Template-guided Prompt Filtering mechanism to align these discretized prompts with regional image semantics, ensuring fine-grained prompt guidance. Extensive experiments demonstrate that HRHuman achieves state-of-the-art performance in human-centeric higher-resolution image generation, significantly addressing both issues of object repetition and structural distortion. Ling Li 0012, Lanqing Guo, Siyuan Yang 0001, Yakun Ju, Weisi Lin, Alex Chichung Kot |
ISCAS | 5 |
| 2025 | A Novel Stereo Matching Network for Underwater ScenesabstractWith the rapid development and widespread application of stereo matching, significant progress has been achieved for ground scenes. However, traditional stereo matching methods designed for ground-based environments often perform poorly in underwater settings due to the unique optical properties and physical characteristics of the underwater environment. In this paper, we propose a novel stereo matching network tailored specifically for underwater scenes. Experimental results validate the effectiveness of our model in addressing the challenges posed by underwater environments. Lvwei Zhu, Yakun Ju, Ying Gao 0005 |
ISCAS | 2 |
| 2025 | Face reconstruction with detailed skin features via three selfie images
Yakun Ju, Bandara Dissanayake, Rachel Ang, Ling Li 0012, Dennis Sng, Alex Chichung Kot |
J. Vis. Commun. Image Represent. | 1 |
| 2025 | Revisiting One-Stage Deep Uncalibrated Photometric Stereo via Fourier EmbeddingabstractThis paper introduces a one-stage deep uncalibrated photometric stereo (UPS) network, namely Fourier Uncalibrated Photometric Stereo Network (FUPS-Net), for non-Lambertian objects under unknown light directions. It departs from traditional two-stage methods that first explicitly learn lighting information and then estimate surface normals. Two-stage methods were deployed because the interplay of lighting with shading cues presents challenges for directly estimating surface normals without explicit lighting information. However, these two-stage networks are disjointed and separately trained so that the error in explicit light calibration will propagate to the second stage and cannot be eliminated. In contrast, the proposed FUPS-Net utilizes an embedded Fourier transform network to implicitly learn lighting features by decomposing inputs, rather than employing a disjointed light estimation network. Our approach is motivated from observations in the Fourier domain of photometric stereo images: lighting information is mainly encoded in amplitudes, while geometry information is mainly associated with phases. Leveraging this property, our method "decomposes" geometry and lighting in the Fourier domain as guidance, via the proposed Fourier Embedding Extraction (FEE) block and Fourier Embedding Aggregation (FEA) block, which generate lighting and geometry features for the FUPS-Net to implicitly resolve the geometry-lighting ambiguity. Furthermore, we propose a Frequency-Spatial Weighted (FSW) block that assigns weights to combine features extracted from the frequency domain and those from the spatial domain for enhancing surface reconstructions. FUPS-Net overcomes the limitations of two-stage UPS methods, offering better training stability, a concise end-to-end structure, and avoiding accumulated errors in disjointed networks. Experimental results on synthetic and real datasets demonstrate the superior performance of our approach, and its simpler training setup, potentially paving the way for a new strategy in deep learning-based UPS methods. Yakun Ju, Boxin Shi, Bihan Wen, Kin-Man Lam 0001, Xudong Jiang 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Forgery-Aware Adaptive Learning With Vision Transformer for Generalized Face Forgery DetectionabstractWith the rapid progress of generative models, the current challenge in face forgery detection is how to effectively detect realistic manipulated faces from different unseen domains. Though previous studies show that pre-trained Vision Transformer (ViT) based models can achieve some promising results after fully fine-tuning on the Deepfake dataset, their generalization performances are still unsatisfactory. To this end, we present a Forgery-aware Adaptive Vision Transformer (FA-ViT) under the adaptive learning paradigm for generalized face forgery detection, where the parameters in the pre-trained ViT are kept fixed while the designed adaptive modules are optimized to capture forgery features. Specifically, a global adaptive module is designed to model long-range interactions among input tokens, which takes advantage of self-attention mechanism to mine global forgery clues. To further explore essential local forgery clues, a local adaptive module is proposed to expose local inconsistencies by enhancing the local contextual association. In addition, we introduce a fine-grained adaptive learning module that emphasizes the common compact representation of genuine faces through relationship learning in fine-grained pairs, driving these proposed adaptive modules to be aware of fine-grained forgery-aware information. Extensive experiments demonstrate that our FA-ViT achieves state-of-the-arts results in the cross-dataset evaluation, and enhances the robustness against unseen perturbations. Particularly, FA-ViT achieves 93.83% and 78.32% AUC scores on Celeb-DF and DFDC datasets in the cross-dataset evaluation. The code and trained model have been released at:https://github.com/LoveSiameseCat/FAViT. Anwei Luo, Rizhao Cai, Chenqi Kong, Yakun Ju, Xiangui Kang, Jiwu Huang, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Aerial Multiview Stereo via Adaptive Depth Range Inference and Normal CuesabstractThree-dimensional digital urban reconstruction from multi-view aerial images is a critical application where deep multi-view stereo (MVS) methods outperform traditional techniques. However, existing methods commonly overlook the key differences between aerial and close-range settings, such as varying depth ranges along epipolar lines and insensitive feature-matching associated with low-detailed aerial images. To address these issues, we propose an Adaptive Depth Range MVS (ADR-MVS), which integrates monocular geometric cues to improve multi-view depth estimation accuracy. The key component of ADR-MVS is the depth range predictor, which generates adaptive range maps from depth and normal estimates using cross-attention discrepancy learning. In the first stage, the range map derived from monocular cues breaks through predefined depth boundaries, improving feature-matching discriminability and mitigating convergence to local optima. In later stages, the inferred range maps are progressively narrowed, ultimately aligning with the cascaded MVS framework for precise depth regression. Moreover, a normal-guided cost aggregation operation is specially devised for aerial stereo images to improve geometric awareness within the cost volume. Finally, we introduce a normal-guided depth refinement module that surpasses existing RGB-guided techniques. Experimental results demonstrate that ADR-MVS achieves state-of-the-art performance on the WHU, LuoJia-MVS, and München datasets, while exhibits superior computational complexity. Yimei Liu, Yakun Ju, Yuan Rao 0001, Hao Fan 0004, Junyu Dong, Feng Gao 0005, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
Jinjiang Li 0001, Yakun Ju, Alex Chichung Kot |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Cross-Frequency Attention and Color Contrast Constraint for Remote Sensing DehazingabstractCurrent deep learning-based methods for remote sensing image dehazing have developed rapidly, yet they still commonly struggle to simultaneously preserve fine texture details and restore accurate colors. The fundamental reason lies in the insufficient modeling of high-frequency information that captures structural details, as well as the lack of effective constraints for color restoration. To address the insufficient modeling of global high-frequency information, we first develop an omni-directional high-frequency feature in painting mechanism that leverages the wavelet transform to extract multi-directional high-frequency components. While maintaining the advantage of linear complexity, it models global long-range texture dependencies through cross-frequency perception. Then, to further strengthen local high-frequency representation, we design a high-frequency prompt attention module that dynamically injects wavelet-domain optimized high-frequency features as cross-level guidance signals, significantly enhancing the model's capability in edge sharpness restoration and texture detail reconstruction. Further, to alleviate the problem of inaccurate color restoration, we propose a color contrast loss function based on the HSV color space, which explicitly models the statistical distribution differences of brightness and saturation in hazy regions, guiding the model to generate dehazed images with consistent colors and natural visual appearance. Finally, extensive experiments on multiple benchmark datasets demonstrate that the proposed method outperforms existing approaches in both texture detail restoration and color consistency. Further results and code are available at: https://github.com/fyxnl/C4RSD. Jufeng Li, Yakun Ju, Chunxu Li, Weisheng Dong, Alex Chichung Kot |
IEEE Trans. Image Process. | 5 |
| 2025 | Fragrant: frequency-auxiliary guided relational attention network for low-light action recognition
Wenxuan Liu 0008, Xuemei Jia, Yihao Ju, Yakun Ju, Kui Jiang, Shifeng Wu, Luo Zhong, Xian Zhong |
Vis. Comput. | 4 |
| 2024 | Towards Progressive Multi-Frequency Representation for Image WarpingabstractImage warping, a classic task in computer vision, aims to use geometric transformations to change the appearance of images. Recent methods learn the resampling kernels for warping through neural networks to estimate missing values in irregular grids, which, however, fail to capture local variations in deformed content and produce images with distortion and less high-frequency details. To address this issue, this paper proposes an effective method, namely MFR, to learn Multi-Frequency Representations from in-put images for image warping. Specifically, we propose a progressive filtering network to learn image representations from different frequency subbands and generate deformable images in a coarse-to-fine manner. Furthermore, we employ learnable Gabor wavelet filters to improve the model's capability to learn local spatial-frequency representations. Comprehensive experiments, including homography trans-formation, equirectangular to perspective projection, and asymmetric image super-resolution, demonstrate that the proposed MFR significantly outperforms state-of-the-art image warping methods. Our method also showcases superior generalization to out-of-distribution domains, where the generated images are equipped with rich details and less distortion, thereby high visual quality. The source code is available at https://github.com/junxiao01/MFR. Jun Xiao 0010, Zihang Lyu, Yakun Ju, Changjian Shui, Kin-Man Lam 0001 |
CVPR | 4 |
| 2024 | Image Gradient-Aided Photometric Stereo Network
Lin Qi 0004, Shiyu Qin, Yakun Ju, Junyu Dong |
PRICAI (3) | 5 |
| 2024 | Deep Learning Methods for Calibrated Photometric Stereo and BeyondabstractPhotometric stereo recovers the surface normals of an object from multiple images with varying shading cues, i.e., modeling the relationship between surface orientation and intensity at each pixel. Photometric stereo prevails in superior per-pixel resolution and fine reconstruction details. However, it is a complicated problem because of the non-linear relationship caused by non-Lambertian surface reflectance. Recently, various deep learning methods have shown a powerful ability in the context of photometric stereo against non-Lambertian surfaces. This paper provides a comprehensive review of existing deep learning-based calibrated photometric stereo methods utilizing orthographic cameras and directional light sources. We first analyze these methods from different perspectives, including input processing, supervision, and network architecture. We summarize the performance of deep learning photometric stereo models on the most widely-used benchmark data set. This demonstrates the advanced performance of deep learning-based photometric stereo methods. Finally, we give suggestions and propose future research trends based on the limitations of existing models. Yakun Ju, Kin-Man Lam 0001, Wuyuan Xie, Huiyu Zhou 0001, Junyu Dong, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Flow-Edge-Net: Video Saliency Detection Based on Optical Flow and Edge-Weighted Balance LossabstractOptical flow networks have been widely utilized for video saliency detection (VSD) due to their effective performance in capturing the motion of objects. However, the use of optical flow blurs the edges of salient objects and leads to the problems of poorly defined object boundaries. To address this issue, we propose an optical flow-based edge-weighted loss function, to train a network called Flow-Edge-Net, which can balance the weights of the foreground and background information at the edges of video frames. It has achieved superior performance in detecting salient boundaries. Specifically, we propose two complementary encoding and decoding networks based on the concept of decoupling. That is, the optical flow network focuses on moving objects, while the edge network, based on the encoder-decoder structure, focuses on edge information. As the two networks output features of the same dimension and are from the same input, our proposed self-designed adaptive weighted feature fusion module can compare and integrate the edge information and location information from the two networks through adaptive weighting. The proposed method has been evaluated on five widely used databases. Experiment results demonstrate the superior performance of the proposed Flow-Edge-Net in locating salient objects, with accurate and refined edges. The proposed method achieves superior performance over the state-of-the-art methods in detecting salient objects in videos. Muwei Jian, Xiangwei Lu, Yakun Ju, Hui Yu 0001, Kin-Man Lam 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Estimating High-Resolution Surface Normals via Low-Resolution Photometric Stereo ImagesabstractAcquiring high-resolution 3D surface structures is a crucial task in computer vision as it provides more detailed surface textures and clearer structures. Photometric stereo can measure per-pixel surface normals of a 3D object using various shading cues. However, obtaining high-resolution images in a linear response photometric stereo imaging system can be challenging. Additionally, photometric stereo, as a per-pixel reconstruction method, requires higher-resolution surface normal maps to accurately depict complex surface structures, particularly in regions that demand more attention and precise reconstruction. Therefore, measuring high-resolution surface normals via low-resolution photometric stereo images is of great importance. Motivated by these, we propose a Super-resolution Photometric Stereo Network, namely SR-PSN. In order to address the issues of measuring the high-resolution surface normals from low-resolution photometric images, we mainly (1) apply a dual-position threshold normalization pre-processing scheme to effectively handle the spatially-varying reflectance of non-Lambertian surfaces, (2) adopt a local affinity feature module to learn the rich structural representation by explicitly revealing the neighbor relationships, (3) employ a parallel multi-scale feature extractor, which preserves high-resolution representations and deep feature extraction, and (4) propose a shared-weight regressor to handle the multi-scale features, to prevent the model collapsing into learning non-important features related to a certain fixed scale. Extensive ablation experiments validate the effectiveness of our proposed modules. Furthermore, quantitative experiments conducted on public benchmarks demonstrate that SR-PSN outperforms state-of-the-art calibrated photometric stereo methods. Notably, SR-PSN achieves superior results while utilizing photometric stereo images with only half the resolution of other methods. It effectively restores the structure of complex surfaces, producing a high-resolution normal map. Yakun Ju, Muwei Jian, Cong Wang 0018, Junyu Dong, Kin-Man Lam 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | GR-PSN: Learning to Estimate Surface Normal and Reconstruct Photometric Stereo ImagesabstractIn this paper, we propose a novel method, namely GR-PSN, which learns surface normals from photometric stereo images and generates the photometric images under distant illumination from different lighting directions and surface materials. The framework is composed of two subnetworks, named GeometryNet and ReconstructNet, which are cascaded to perform shape reconstruction and image rendering in an end-to-end manner. ReconstructNet introduces additional supervision for surface-normal recovery, forming a closed-loop structure with GeometryNet. We also encode lighting and surface reflectance in ReconstructNet, to achieve arbitrary rendering. In training, we set up a parallel framework to simultaneously learn two arbitrary materials for an object, providing an additional transform loss. Therefore, our method is trained based on the supervision by three different loss functions, namely the surface-normal loss, reconstruction loss, and transform loss. We alternately input the predicted surface-normal map and the ground-truth into ReconstructNet, to achieve stable training for ReconstructNet. Experiments show that our method can accurately recover the surface normals of an object with an arbitrary number of inputs, and can re-render images of the object with arbitrary surface materials. Extensive experimental results show that our proposed method outperforms those methods based on a single surface recovery network and shows realistic rendering results on 100 different materials. Yakun Ju, Boxin Shi, Yang Chen 0036, Huiyu Zhou 0001, Junyu Dong, Kin-Man Lam 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | A Structure-Affinity Dual Attention-based Network to Segment Spine for Scoliosis AssessmentabstractUltrasound volume projection imaging has shown great promise to visualize spine features and diagnose scoliosis thanks to its harmlessness, cheapness, and efficiency. The key to measuring spine deformity and assessing scoliosis is to accurately segment the spine bone features. In this paper, we propose a novel structure-affinity dual attention-based network (SADANet) for effective spine segmentation. Global channel attention module and spatial criss-cross attention module are combined in a parallel manner to generate rich global context of spine images. Meanwhile, we present a structure-affinity strategy to encode the structural knowledge of spine bones into the semantic representations. By this means, the network can capture both contextual and structural information. Experiments show that our proposed algorithm achieves promising performance on spine segmentation as compared with other state-of-the-art candidates, which makes it an appealing approach for intelligent scoliosis assessment. Zixun Huang, Frank H. F. Leung, Yakun Ju, Sai-Ho Ling |
BIBM | 4 |
| 2023 | Efficient Feature Fusion for Learning-Based Photometric StereoabstractHow to handle an arbitrary number for input images is a fundamental problem of learning-based photometric stereo methods. Existing approaches adopt max-pooling or observation map to fuse an arbitrary number of extracted features. However, these methods discard a large amount of the features from the input images, impacting the utilization and accuracy, or ignore the constraints from the intra-image spatial domain. In this paper, we explore how to efficiently fuse features from a variable number of input images. First, we propose a bilateral extraction module, which categorizes features into positive and negative, to maximally keep the useful feature in the fusion stage. Second, we adopt a top-k pooling to both the bilateral information, which selects the k maximum response value from all features. These two modules proposed are "plug-and-play" and can be used in different fusion tasks. We further propose a hierarchical photometric stereo network, namely HPS-Net, to handle bilateral extraction and top-k pooling for multiscale features. Experiments in the widely used benchmark illustrate the improvement of our proposed framework in the conventional max-pooling method and the proposed HPS-Net outperforms existing learning-based photometric stereo methods. Yakun Ju, Kin-Man Lam 0001, Jun Xiao 0010, Cuixin Yang, Junyu Dong |
ICASSP | 1 |
| 2023 | Improving Robustness of Single Image Super-Resolution Models with Monte Carlo MethodabstractDeep learning-based methods have achieved promising results in single image super-resolution (SISR). However, the performance of existing deep SISR methods is very sensitive to image degradation. In addition, these methods are deterministic and do not introduce any uncertainty to the generated images, so we have no way of knowing the reliability of these generated images. To address these two challenging issues, we propose a model-agnostic approach for existing deep SISR networks to improve their robustness under various degradations. Our proposed method follows a probabilistic framework and applies Monte Carlo dropout to existing deep SISR methods. Instead of performing point estimation, the proposed method predicts the posterior distribution of super-resolved images. Based on this, we can determine the uncertainty of the generated images. Experiment results show that the proposed method can effectively improve the robustness of existing deep SISR methods, leading to state-of-the-art performance when applied to images having different degradations. The code is available at https://github.com/YangTracy/MCD-SR. Cuixin Yang, Jun Xiao 0010, Yakun Ju, Guoping Qiu, Kin-Man Lam 0001 |
ICIP | 3 |
| 2023 | Pyramid Masked Image Modeling for Transformer-Based Aerial Object DetectionabstractTwo obstacles, the scarcity of annotated samples and the difficulty in preserving multi-scale hierarchical representations, hinder the advancement of vision Transformer-based aerial object detection. The emergence of self-supervised learning has inspired some solutions to the first issue. However, most solutions focus on single-scale features, conflicting with solving the second issue. To bridge this gap, this paper proposes a novel pyramid masked image modeling (MIM) framework, termed PyraMIM, for self-supervised pretraining in aerial scenarios. Without manual annotation, PyraMIM enables establishing pyramid representations during pretraining, which can be seamlessly adapted to downstream aerial object detection for performance improvement. Experimental results demonstrate the effectiveness and superiority of our method. Tianshan Liu, Yakun Ju, Kin-Man Lam 0001 |
ICIP | 3 |
| 2023 | Learning Deep Photometric Stereo Network with Reflectance PriorsabstractPhotometric stereo recovers the surface normals of an object from images with varying shading cues. Conventional photometric stereo methods attempt to use handcrafted reflectance models to approximate surface normals, while deep learning-based networks have shown a much more powerful ability to handle non-Lambertian objects. However, none of the existing deep learning methods explores how prior reflectance information can be used to optimize surface-normal prediction. In this paper, we first present the introduction of reflectance prior to deep photometric stereo models. Our explorations include how the reflectance prior can simplify the optimization of deep networks by reparametrizing the weights, and (2) eliminate the impacts of surfaces with spatially varying reflectance for all-pixel input photometric stereo methods. To achieve these goals, we propose a residual fusion module (RFM) in our method, which explicitly extracts features useful for surface-normal recovery and removes those features influenced by reflectance. Additionally, we design a shading extractor with multi-scale and global-local feature fusion operations, which can fuse features with different receptive fields and better utilize the non-maximum features missing in the max-pooling operation. Experiments and ablation studies verify the accuracy and effectiveness of the proposed reflectance prior network on a widely used benchmark. Yakun Ju, Songsong Huang, Yuan Rao 0001, Kin-Man Lam 0001 |
ICME | 1 |
| 2023 | pmBQA: Projection-based Blind Point Cloud Quality Assessment via Multimodal LearningabstractWith the increasing communication and storage of point cloud data, there is an urgent need for an effective objective method to measure the quality before and after processing. To address this difficulty, we propose a projection-based blind quality indicator via multimodal learning for point cloud data, which can perceive both geometric distortion and texture distortion by using four homogeneous modalities (i.e., texture, normal, depth and roughness). To fully exploit the multimodal information, we further develop a deformable convolutionbased alignment module and a graph-based feature fusion module, and investigate a graph node attention-based evaluation method to forecast the quality score. Extensive experimental results on three benchmark databases show that our method achieves more accurate evaluation performance in comparison with 12 competitive methods. Wuyuan Xie, Kaimin Wang, Yakun Ju, Miaohui Wang |
ACM Multimedia | 3 |
| 2023 | PromptRestorer: A Prompting Image Restoration Method with Degradation PerceptionabstractWe show that raw degradation features can effectively guide deep restoration models, providing accurate degradation priors to facilitate better restoration. While networks that do not consider them for restoration forget gradually degradation during the learning process, model capacity is severely hindered. To address this, we propose a Prompting image Restorer, termed as PromptRestorer. Specifically, PromptRestorer contains two branches: a restoration branch and a prompting branch. The former is used to restore images, while the latter perceives degradation priors to prompt the restoration branch with reliable perceived content to guide the restoration process for better recovery. To better perceive the degradation which is extracted by a pre-trained model from given degradation observations, we propose a prompting degradation perception modulator, which adequately considers the characters of the self-attention mechanism and pixel-wise modulation, to better perceive the degradation priors from global and local perspectives. To control the propagation of the perceived content for the restoration branch, we propose gated degradation perception propagation, enabling the restoration branch to adaptively learn more useful features for better recovery. Extensive experimental results show that our PromptRestorer achieves state-of-the-art results on 4 image restoration tasks, including image deraining, deblurring, dehazing, and desnowing. Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Jiangxin Dong, Mengzhu Wang, Yakun Ju, Junyang Chen 0001 |
NeurIPS | 6 |
| 2023 | Learning General Descriptors for Image Matching With Regression FeedbackabstractRecent advances on feature descriptors for image matching put more emphasis on encoding invariances (e.g. illumination invariance) to promote the descriptors’ discriminative power. However, according to the information entropy, more invariance implies greater certainty and less informativeness in a descriptor. Consequently, descriptors encoding too many invariances usually show poor generalization to unknown image changes, lacking enough informativeness to cover the large uncertainty in unseen scenes. This limits the application scenarios of learned descriptors. In this paper, we propose to alleviate this issue from the perspective of informativeness and we thus design hierarchical consistent constraint by introducing regression feedback in a self-supervised manner. Combined with the hardest-within-batch matching constraint, we form a novel dual supervision framework, to encourage the descriptor to learn an informative representation while maintaining a good discriminative power. Moreover, to fully mine the context information hidden in image and boost the informativeness in turn, we present AANet, a descriptor network that efficiently predicts dense description by the powerful Attentional Aggregation of multi-level features. Experiments across challenging feature matching on HPatches, RDNIM datasets, and visual localization tasks on Aachen Day-night dataset show that our method outperforms recent state-of-the-art descriptors while keeping encouraging efficiency. The application of visual 3D reconstruction on various scenarios also demonstrates the high generalization ability of our method. Yuan Rao 0001, Yakun Ju, Eric Rigall, Jian Yang 0036, Hao Fan 0004, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Efficient Inductive Vision Transformer for Oriented Object Detection in Remote Sensing ImageryabstractObject detection is a fundamental task in remote sensing image analysis and scene understanding. Previous remote sensing object detectors are typically based on convolutional neural networks (CNNs), whose performance is significantly limited by the intrinsic locality of convolution operations. The emergence of vision Transformers brings potential solutions to this problem, which have the capability to be a solid alternative to CNNs. However, three crucial obstacles hinder the application and performance of Transformers in the task of remote sensing object detection, i.e., 1) high computational complexity, especially for high-resolution remote sensing images, 2) training-and sample-inefficiency caused by lack of inductive bias, and 3) difficulty in learning arbitrary orientation knowledge of geospatial objects. To address these issues, in this paper, a novel efficient inductive vision Transformer framework is proposed for oriented object detection in remote sensing imagery. This framework follows the hierarchical feature pyramid structure and makes threefold contributions, as follows. 1) Spatial redundancy in remote sensing images is fully explored and an adaptive multi-grained routing mechanism is proposed to facilitate token sparsity in Transformers, which can dramatically reduce the computational cost without comprising the accuracy. 2) A compact dual-path encoding architecture, where both global long-range dependencies and local semantic relations are jointly and complementarily captured, is proposed to enhance inductive bias in Transformers. 3) An angle tokenization technique is proposed to promote the encoding, embedding, and learning of direction knowledge for oriented objects in remote sensing scenarios. In this work, the above three contributions are instantiated in an advanced Transformer-based object detector, namely EIA-PVT. Comprehensive experiments on two publicly available datasets have demonstrated its effectiveness and superiority for oriented object detection in remote sensing images. Jingran Su, Yakun Ju, Kin-Man Lam 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Learning conditional photometric stereo with high-resolution featuresabstractPhotometric stereo aims to reconstruct 3D geometry by recovering the dense surface orientation of a 3D object from multiple images under differing illumination. Traditional methods normally adopt simplified reflectance models to make the surface orientation computable. However, the real reflectances of surfaces greatly limit applicability of such methods to real-world objects. While deep neural networks have been employed to handle non-Lambertian surfaces, these methods are subject to blurring and errors, especially in high-frequency regions (such as crinkles and edges), caused by spectral bias: neural networks favor low-frequency representations so exhibit a bias towards smooth functions. In this paper, therefore, we propose a self-learning conditional network with multi-scale features for photometric stereo, avoiding blurred reconstruction in such regions. Our explorations include: (i) a multi-scale feature fusion architecture, which keeps high-resolution representations and deep feature extraction, simultaneously, and (ii) an improved gradient-motivated conditionally parameterized convolution (GM-CondConv) in our photometric stereo network, with different combinations of convolution kernels for varying surfaces. Extensive experiments on public benchmark datasets show that our calibrated photometric stereo method outperforms the state-of-the-art. Yakun Ju, Yuxin Peng 0001, Muwei Jian, Feng Gao 0005, Junyu Dong |
Comput. Vis. Media | 1 |
| 2022 | NormAttention-PSN: A High-frequency Region Enhanced Photometric Stereo Network with Normalized Attention
Yakun Ju, Boxin Shi, Muwei Jian, Lin Qi 0004, Junyu Dong, Kin-Man Lam 0001 |
Int. J. Comput. Vis. | 1 |
| 2022 | A deep-shallow and global-local multi-feature fusion network for photometric stereo
Yanru Liu, Yakun Ju, Muwei Jian, Feng Gao 0005, Yuan Rao 0001, Yeqi Hu, Junyu Dong |
Image Vis. Comput. | 2 |
| 2022 | 3D Hand Pose Estimation From Monocular RGB With Feature Interaction Moduleabstract3D hand pose estimation from a monocular RGB image is a highly challenging task due to self-occlusion, diverse appearances, and inherent depth ambiguities within monocular images. Most of the previous methods first employ deep neural networks to fit 2D joint location maps, then combines them with implicit or explicit pose-aware features to directly regress 3D hand joints positions using their designed network structure. However, the skeleton positions and corresponding skeleton-aware content information located in the latent space are invariably ignored. These skeleton-aware contents effectively bridge the gap between hand joint and hand skeleton information by associating the relationship between different hand joints features and the hand skeleton positions distribution in 2D space. To address this issue, we propose a simple yet efficient deep neural network to directly recover reliable 3D hand pose from monocular RGB images with faster estimation process. Our purpose is the reduction of the model computational complexity while maintaining high precision performance. Therefore, we design a novel Feature Chat Block (FCB) to complete feature boosting, which enables the intuitively enhanced interaction between joint and skeleton features. First, this FCB module updates joint features effectively based on semantic graph convolutional neural network and multi-head self-attention mechanism. The GCN-based structure focuses on the physical hand joints included in a binary adjacency matrix and the self-attention part pays attention to hand joints located in a complementary matrix. Then, the FCB module employs query and key mechanisms respectively representing joint and skeleton features to further implement feature interaction. After a set of FCB modules, our model updates the fused features in a coarse-to-fine manner and finally outputs a predicted 3D hand pose. We conducted a comprehensive set of ablation experiments on the InterHand2.6M dataset to validate the effectiveness and significance of the proposed method. Additionally, experimental results on Rendered Hand Dataset, Stereo Hand Datasets, First-Person Hand Action Dataset and FreiHAND Dataset show our model surpasses the state-of-the-art methods with faster inference speed. Shaoxiang Guo, Eric Rigall, Yakun Ju, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Recovering Surface Normal and Arbitrary Images: A Dual Regression Network for Photometric StereoabstractPhotometric stereo recovers three-dimensional (3D) object surface normal from multiple images under different illumination directions. Traditional photometric stereo methods suffer from the problem of non-Lambertian surfaces with general reflectance. By leveraging deep neural networks, learning-based methods are capable of improving the surface normal estimation under general non-Lambertian surfaces. These state-of-the-art learning-based methods however do not associate surface normal with reconstructed images and, therefore, they cannot explore the beneficial effect of such association on the estimation of the surface normal. In this paper, we specifically exploit the positive impact of this association and propose a novel dual regression network for both fine surface normals and arbitrary reconstructed images in calibrated photometric stereo. Our work unifies the 3D reconstruction and rendering tasks in a deep learning framework, with the explorations including: 1. generating specified reconstructed images under arbitrary illumination directions, which provides more intuitive perception of the reflectance and is extremely useful for visual applications, such as virtual reality, and 2. our dual regression scheme introduces an additional constraint on observed images and reconstructed images, which forms a closed-loop to provide additional supervision. Experiments show that our proposed method achieves accurate reconstructed images under arbitrarily specified illumination directions and it significantly outperforms the state-of-the-art learning-based single regression methods in calibrated photometric stereo. Yakun Ju, Junyu Dong, Sheng Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Pay Attention to Devils: A Photometric Stereo Network for Better DetailsabstractWe present an attention-weighted loss in a photometric stereo neural network to improve 3D surface recovery accuracy in complex-structured areas, such as edges and crinkles, where existing learning-based methods often failed. Instead of using a uniform penalty for all pixels, our method employs the attention-weighted loss learned in a self-supervise manner for each pixel, avoiding blurry reconstruction result in such difficult regions. The network first estimates a surface normal map and an adaptive attention map, and then the latter is used to calculate a pixel-wise attention-weighted loss that focuses on complex regions. In these regions, the attention-weighted loss applies higher weights of the detail-preserving gradient loss to produce clear surface reconstructions. Experiments on real datasets show that our approach significantly outperforms traditional photometric stereo algorithms and state-of-the-art learning-based methods. Yakun Ju, Kin-Man Lam 0001, Yang Chen 0036, Lin Qi 0004, Junyu Dong |
IJCAI | 1 |
| 2020 | Learning Photometric Stereo via Manifold-based MappingabstractThree-dimensional reconstruction technologies are fundamental problems in computer vision. Photometric stereo recovers the surface normals of a 3D object from varying shading cues, prevailing in its capability for generating fine surface normal. In recent years, deep learning-based photometric stereo methods are capable of improving the surface-normal estimation under general non-Lambertian surfaces, due to its powerful fitting ability on the non-Lambertian surface. These state-of-the-art methods however usually regress the surface normal directly from the high-dimensional features, without exploring the embedded structural information. This results in the underutilization of the information available in the features. Therefore, in this paper, we propose an efficient manifold-based framework for learning-based photometric stereo, which can better map combined high-dimensional feature spaces to low-dimensional manifolds. Extensive experiments show that our method, learning with the low-dimensional manifolds, achieves more accurate surface-normal estimation, outperforming other state-of-the-art methods on the challenging DiLiGenT benchmark dataset. Yakun Ju, Muwei Jian, Junyu Dong, Kin-Man Lam 0001 |
VCIP | 1 |
| 2020 | MPS-Net: Learning to recover surface normal for multispectral photometric stereo
Yakun Ju, Lin Qi 0004, Jichao He, Xinghui Dong, Feng Gao 0005, Junyu Dong |
Neurocomputing | 1 |
| 2020 | A dual-cue network for multispectral photometric stereo
Yakun Ju, Xinghui Dong, Yingyu Wang, Lin Qi 0004, Junyu Dong |
Pattern Recognit. | 1 |