Junyu Fan

dblp:208/7645 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DSPNet: Dual-Path Saliency Prediction With Reconstructed Feature Injection and Discrepancy-Aware KL Loss
abstract
Visual saliency prediction (VSP) is a fundamental task in computer vision, which enables efficient sensing and computation by prioritizing informative regions, such as object detection, scene understanding, and video compression. However, the VSP task still faces three challenges: (i) fine-grained texture is irreversibly lost during repeated down-sampling in single-stream convolutional networks; (ii) traditional attention modules apply fixed, simplistic weights that adapt poorly to complex scenes; and (iii) standard Kullback-Leibler (KL) divergence fails to sufficiently amplify gradients in weakly activated regions, thereby hindering the learning of hard samples. A dual-path saliency prediction network (DSPNet) is proposed to address the aforementioned challenges. A vector-quantized variational autoencoder (VQ-VAE) branch discretizes local textures and reinjects them into a U-Net backbone to recover lost details caused by down-sampling. The restored texture is integrated with high-level semantics via a bidirectional mutual transpose attention (BMTA) decoder, enabling nonlinear cross-layer interactions and enhancing spatial discrimination in complex scenes. A discrepancy-aware focal KL (F-KL) loss is designed to optimize the model training, which adaptively emphasizes discrepant regions to improve attention to areas with prediction gaps. Evaluations on four public benchmarks (MIT300, SALICON, PASCAL-S, and TORONTO) show state-of-the-art results across metrics such as Similarity (SIM), KL divergence, and Normalized Scanpath Saliency (NSS), confirming that DSPNet effectively balances local texture fidelity with global semantic understanding.
Chuanlin Liao, Xiaolin Gou, Junyu Fan, Yi Lin 0006
IEEE Trans. Multim.3
2026 MSG-Net: Structure-Guided Enhancement for Underwater Images Based on Multi-View Feature Interaction
abstract
Underwater images often suffer from blurring, low contrast, and color distortion. Although current multi-view methods show better potential than single-input methods in mitigating these degradations, they still struggle to fuse and extract features across views. Moreover, most existing methods struggle to capture fine texture details and preserve key structural information, resulting in images with insufficient detail and unclear structures. To address these issues, we propose a Multi-View Input and Structure-Guided Network, MSG-Net, to achieve underwater image enhancement. Specifically, White Balance (WB) and Contrast Limited Adaptive Histogram Equalization (CLAHE) are employed to construct multi-view input, enabling more effective handling of diverse underwater degradation factors. A Cross-Channel Attention Fusion (CCAF) module is designed to effectively fuse these multi-view features via a channel-wise attention mechanism. Additionally, a Channel-Group Axial Transformer (CGAT) is introduced to extract axial features within channel groups, enhancing cross-channel interactions while alleviating the patch boundary discontinuities in Transformer-based methods. Furthermore, a Structure-Guided Transformer (SGT) is proposed to incorporate structural information into the enhancement process, enriching texture details and highlighting key structures. Extensive experiments demonstrate that MSG-Net outperforms or matches state-of-the-art underwater image enhancement methods across various benchmarks. Moreover, MSG-Net exhibits outstanding performance on other enhancement tasks.
Junyu Fan, Chuanlin Liao, Yi Lin 0006
IEEE Trans. Multim.2
2025 Design and Analysis of MIMO Coil Based Dynamic Wireless Charging System for Electric Vehicle
abstract
Dynamic wireless charging (DWC) systems have emerged as a highly promising solution to enhance the convenience and efficiency of electric vehicle (EV) charging. In this paper, we present a novel Multiple-Input Multiple-Output coils-based DWC (MIMO-DWC) approach, conducting a discrete analysis model of the continuous dynamic wireless charging process while aiming to improve the performance and reliability of MIMO-DWC systems. By integrating MIMO coils with traditional inductive power transfer (IPT), the proposed system theoretically attains a notably higher power transfer efficiency and remarkable spatial flexibility, effectively reducing alignment sensitivity and significantly enhancing the robustness of the charging process. The simulation results indicate that under a reasonable coupling coefficient, the MIMO-DWC system can achieve a maximum power transfer efficiency 84.5%, surpassing that of the Single-Input Single-Output (SISO) DWC system 62.5%, with experimental validation carried out to firmly confirm the feasibility and effectiveness of the proposed method.
Zhishen Luo, Junyu Fan, Bowang Zhang
IECON2
2025 Multi-CV-Output Domino Wireless Power Transfer System With Reconfigurable Voltage Levels
abstract
This paper presents a domino wireless power transfer system designed to provide flexible constant-voltage (CV) output in each relay unit. Importantly, the voltage level of each CV output can be independently and flexibly preconfigured, establishing a reconfigurable output paradigm. The core innovation lies in the use of an inductor-capacitor-capacitor-series (LCC-S) compensation network in each relay unit, allowing for the intentional design of compensated inductance and mutual inductance to achieve various CV lift ratios. Additionally, four typical CV scenarios are thoroughly analyzed: constant, monotonically increasing, monotonically decreasing, and high-low jumps. Finally, an experimental prototype has been developed, and additional measured waveforms will be presented to demonstrate the effectiveness of the proposed system.
Yilin Zhang 0005, Haoyuan Xiao, Bowang Zhang, Junyu Fan, Wei Han 0007
IECON5
2025 AV-RISE: Hierarchical Cross-Modal Denoising for Learning Robust Audio-Visual Speech Representation
abstract
Audio-visual speech recognition (AVSR) leverages complementary visual cues to improve speech recognition. However, in real-world scenarios, both modalities may suffer from noise or occlusion. In such scenarios, most existing fusion strategies overlook the variation in modality-specific quality under different degradation conditions. This limitation may lead to dominance of corrupted modality in the fusion process, resulting in worse AVSR performance than unimodal systems, termed as Corrupted Modality Bias (CMB) in this work. To address this, a self-supervised speech representation learning framework, called AV-RISE, is proposed to employ teacher-student self-distillation to robustly reconstruct clean speech representations from corrupted audio-visual inputs. A hierarchical fusion mechanism is designed to progressively refine audio and visual representations by integrating the Suppression and Enhancement Interaction (SEI) module into each layer of the pre-trained encoder. In the SEI module, cross-modal suppression and modality-oriented enhancement are performed to mitigate noise-induced feature inconsistencies, which strengthens the modeling of complementary semantic representations. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that AV-RISE outperforms SOTA AVSR models, especially under extreme degradation. Most importantly, the hierarchical SEI-based fusion effectively enhances reliable semantic representations to mitigate CMB, by evaluating feature similarities between clean and noise samples.
Zhishuo Zhao, Yi Lin 0006, Dongyue Guo, Junyu Fan
ACM Multimedia4
2025 See Through Water: Heuristic Modeling Toward Color Correction for Underwater Image Enhancement
abstract
Color cast is one of the main degradations in underwater images. Existing data-driven methods, while capable of learning color correction rules from large datasets, often overlook the imaging characteristics and light behavior in underwater environments, making them unable to accurately restore colors in complex water bodies. To address this, we use color constancy and an underwater imaging model to heuristically model the underwater environment for accurate color restoration. On one hand, we propose a multi-scale joint prior network architecture to fully explore the rich feature-level information at different scales in underwater images. This is used to fit the complex parameters of the underwater imaging model, deriving high-quality potential undegraded images. On the other hand, to tackle the challenges of color distortion caused by complex imaging factors in different water environments, we estimate the background light of the water body through the color constancy of underwater objects and dynamically incorporate it into the underwater imaging model as a prior. This not only guides the learning process more effectively but also allows the model to consider key aspects of underwater optical propagation, making it adaptable to different water environments and improving the color accuracy of the enhanced images. We have also conducted extensive experiments to demonstrate the effectiveness of the proposed method, which not only achieves the best overall performance in qualitative analysis and quantitative comparison but also boasts the best color accuracy and the fastest inference speed. The code is available athttps://github.com/JunyuFan/MJPNet.
Junyu Fan, Jingchun Zhou, Danling Meng, Yi Lin 0006
IEEE Trans. Circuits Syst. Video Technol.1
2024 Frequency-aware robust multidimensional information fusion framework for remote sensing image segmentation
Junyu Fan, Jinjiang Li 0001, Yepeng Liu 0003, Fan Zhang 0045
Eng. Appl. Artif. Intell.1
2024 Elevation Information-Guided Multimodal Fusion Robust Framework for Remote Sensing Image Segmentation
abstract
Currently, the task of remote sensing image segmentation still faces some challenges, such as variations in illumination, shadows, and occlusions present in remote sensing images. Additionally, there may be similarities and confusions between different types of terrain features. In this paper, we aim to explore how to utilize information exchange between multiple modalities to reduce the impact of interfering factors. To fully exploit the complementary information between different modalities, we establish an information exchange mechanism between optical images (visible light + infrared) features and Digital Surface Model (DSM) features. This allows them to interact and express themselves in a shared feature space, facilitating the acquisition of complementary information from different modalities. Furthermore, through a multimodal fusion encoder and decoder based on Transformer design, the optical features and DSM features are integrated, enabling the learning of high-level semantic representations in different dimensions. Extensive subjective, objective comparative experiments, and ablation experiments are conducted on the ISPRS Vaihingen and Potsdam datasets to evaluate the proposed method. The mIoU on the Vaihingen and Potsdam datasets reached 85.06% and 87.6% respectively, while the OA reached 92.01% and 91.92% respectively. The source code will be available at https://github.com/JunyuFan/MIEFNet.
Junyu Fan, Jinjiang Li 0001, Zhen Hua, Fan Zhang 0045, Caiming Zhang 0001
IEEE Geosci. Remote. Sens. Lett.1
2023 Joint transformer progressive self-calibration network for low light enhancement
abstract
Abstract When the lighting conditions are poor and the environmental light is weak, the image captured by the imaging device often has lower brightness and is accompanied by a lot of noise. The paper designs a progressive self‐calibration network model (PSCNet) for recovering high‐quality low‐light‐enhanced images. First, shallow features in low‐light images can be better focused and extracted with the help of attention mechanism. Next, the feature mapping is passed to the encoder and decoder modules, where the transformer and encoder‐decoder jump connection structures can be better combined with the semantic information of the context to learn rich deep feature information. Finally, the self‐calibration module can adaptively cascade the features decoded by the decoder and input them into the residual attention module quickly and accurately. Meanwhile, the LBP features of the image are also fused into the feature information of the residual attention module to enhance the detailed texture information of the image. Qualitative analysis and quantitative comparison of a large number of experimental results show that this method outperforms existing methods.
Junyu Fan, Jinjiang Li 0001, Zhen Hua, Linwei Fan
IET Image Process.1
2023 Multi-scale dynamic fusion for correcting uneven illumination images
Junyu Fan, Jinjiang Li 0001, Zheng Chen 0018
J. Vis. Commun. Image Represent.1
2017 Detrended fluctuation analysis of daily rainfall records of the entire China
abstract
In our analysis, we consider daily rainfall records y(n) from 3825 areas covering entire China. The daily rainfall records is 50 years long from 1961 to 2011. We study the statistical properties of the daily rainfall data and return intervals Tqbetween two consecutive rainfall records above some threshold q. Using the detrended fluctuation analysis (DFA) method to analyze the long-term correlation properties of the rainfall record y(n) and return intervals Tq, we find that the long-term correlation is not exist in rainfall record, while a weak long-term correlation is found in return intervals, pointing to a random regulation of rainfall dynamics. Then, we further analyze the probability distribution function (PDF) of the rainfall record and return intervals of all the 3825 areas, and find that all the PDF obey a power-law distribution.
Choujun Zhan, Weiwen Cao, Junyu Fan, Chengda Wu, Ming-Bo Zhao
INDIN3