Jiafan Zhuang

dblp:227/4676 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0003-3708-4634ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Masked Genetic Operators with Causal Grouping for Constrained Multi-Objective Optimization
abstract
Uncovering the direct causal relationships between decision variables and optimization objectives can significantly simplify the complexity of optimization problems. However, most existing constrained multi-objective evolutionary algorithms (CMOEAs) fail to address constrained multi-objective optimization from this perspective. To bridge this gap, this study introduces a novel algorithm, CI-CMOEA (Causal Intervention-based CMOEA), which leverages causal intervention techniques to enhance optimization performance. CI-CMOEA begins by constructing a causal relationship network that captures the interactions between decision variables and optimization objectives. Using this network, a genetic operator with a causal relationship mask is designed to group decision variables based on their causal impact on the objectives. By focusing genetic operations on key variables with significant causal influence, the algorithm effectively guides the evolutionary optimization process towards better solutions. To further improve performance, CI-CMOEA employs a dual-population collaboration mechanism. One population operates under relaxed epsilon constraints to explore the solution space, while the other disregards constraints to enhance convergence. Preliminary experiments on the LIR-CMOP test suite demonstrate that CI-CMOEA not only accurately identifies the causal relationships between decision variables and objectives but also outperforms eight state-of-the-art CMOEAs in terms of IGD, IGD+ and HV metrics, showcasing its superior optimization performance and reliability.
Zhaojun Wang, Jiachun Huang, Wenji Li, Shunge Wang, Yifeng Qiu, Jiafan Zhuang, Zhun Fan
CEC7
2025 U-Shaped Network Based on Particle Swarm Optimization for Retinal Vessel Segmentation
abstract
Accurate retinal vessel segmentation plays a critical role in the early detection and monitoring of ophthalmic diseases. In this work, we propose a novel retinal vessel segmentation method that integrates Neural Architecture Search (NAS) with a U-shaped encoder-decoder network, optimized using particle swarm optimization (PSO). The framework automates the design of scalable architectures by exploring an extensible search space built with lightweight construction modules, including 3 × 3 convolutions, batch normalization, attention modules, and residual connections. Experimental results on the DRIVE and CHASE_DB1 datasets demonstrate that the searched model achieves superior segmentation accuracy with the fewest parameters (only 0.04M) compared to existing methods. Furthermore, the model exhibits competitive performance on the crack bench-mark dataset CrackLS315, highlighting the strong generalization capability of the searched architecture. In conclusion, the proposed method achieves an optimal balance between segmentation accuracy and model complexity, demonstrating its potential for clinical applications.
Guijie Zhu, Jiafan Zhuang, Wenji Li, Zhun Fan
CEC3
2025 Robust Policy Learning for Multi-UAV Collision Avoidance with Causal Feature Selection
Jiafan Zhuang, Gaofei Han, Zihao Xia, Che Lin, Boxi Wang, Wenji Li, Ruichu Cai, Zhun Fan
AAMAS1
2025 Causality-Inspired Graph Neural Network for Interpretable Strabismus Subtype Classification
Jiawen Zheng, Jiafan Zhuang, Peiwei Wei, Lihao Zhong, Xiaoling Xie, Jinming Guo, Meng Xie, Xiaoli Kang, Jie Cen, Lingyan Dong, Zhun Fan
MICCAI (8)3
2025 Automatic lightweight networks for real-time road crack detection with DPSO
Guijie Zhu, Shuilong Shen, Meihua Wang, Jiafan Zhuang, Zhun Fan
Adv. Eng. Informatics5
2025 Paying more attention on backgrounds: Background-centric attention for UAV detection
Xiuxiu Lin, Yusu Niu, Xinran Yu, Zhun Fan, Jiafan Zhuang, An-Min Zou
Neural Networks5
2024 Well Trajectory Design Based on Constrained Many-Objective Optimization Algorithms
abstract
In the field of drilling engineering, the design and optimization of well trajectories are crucial. This study focuses on optimizing key aspects such as the length of the well trajectory, drill string torque, the energy of the well-profile, and accuracy in reaching the target. This problem encompasses eleven complex nonlinear constraints and four conflicting objectives, presenting challenges for traditional mathematical programming methods. To tackle this problem, we introduce a novel constrained many-objective optimization algorithm, named PPS-NSGA-III. The proposed algorithm partitions the objective space into subspaces, using NSGA-III to find Pareto optimal solutions in each, enhancing diversity. The push-and-pull search framework is employed to overcome local optima in each subproblem, accelerating overall convergence. Through a comparative analysis with some evolutionary algorithms, PPS-NSGA-III has shown superior performance. It delivers more effective design solutions with lower risk, reduced cost, and a higher drilling encounter rate in the proposed well trajectory optimization model.
Zhaojun Wang, Chenwen Ding, Wenji Li, Yifeng Qiu, Jiafan Zhuang, Zhun Fan
CEC5
2024 Infer from What You Have Seen Before: Temporally-dependent Classifier for Semi-supervised Video Segmentation
abstract
Due to high expense of human labor, one major challenge for semantic segmentation in real-world scenarios is the lack of sufficient pixel-level labels, which is more serious when processing video data. To exploit unlabeled data for model training, semi-supervised learning methods attempt to construct pseudo labels or various auxiliary constraints as supervision signals. However, most of them just process video data as a set of independent images in a per-frame manner. The rich temporal relationships are ignored, which can serve as valuable clues for representation learning. Besides, this per-frame recognition paradigm is quite different from that of humans. Actually, benefited from the internal temporal relevance of video data, human would wisely use the distinguished semantic concepts in historical frames to aid the recognition of the current frame. Motivated by this observation, we propose a novel temporally-dependent classifier (TDC) to mimic the human-like recognition procedure. Comparing to the conventional classifier, TDC can guide the model to learn a group of temporally-consistent semantic concepts across frames, which essentially provides an implicit and effective constraint. We conduct extensive experiments on Cityscapes and Cam Vid, and the results demonstrate the superiority of our proposed method to previous state-of-the-art methods. The code is available at https://github.com/jfzhuang/TDC.
Jiafan Zhuang, Zilei Wang, Zhun Fan
CVPR1
2024 OUR-Net: A Multi-Frequency Network With Octave Max Unpooling and Octave Convolution Residual Block for Pavement Crack Segmentation
abstract
Cracks are among the most common, most likely, and earliest of all pavement distresses. Detecting and repairing cracks as early as possible can help extend the service life of pavements. However, Detecting cracks with precision can be challenging due to their varied structural characteristics and complex background interference. In this paper, a new convolutional neural network architecture, OUR-Net, is designed to more efficiently treat both high-and low-frequency visual image features. An Ocatve Convolution is incorporated into the proposed network as an enhancement to conventional convolution. In particular, an Octave Convolution Residual Block (OCRB) is embedded in the encoder to replace the convolutional layer of the classical encoder. Moerover, we propose Octave Max Unpooling (OMU) as the upsampling operation of the decoder, enabling the neural network to learn how to decode multi-spatial frequency features. Compared with models using traditional convolution, OUR-Net has better capability of processing multi-scale information, thus simultaneously improving model performance while saving computational costs by reducing spatial redundancy. We evaluate the superiority of the proposed method by comparing it to state-of-the-art crack segmentation methods on four public datasets (CrackLS315, CFD, Crack200, DeepCrack), which encompass cracks of various widths. Comprehensive experimental results reveal that the proposed method performs excellently, which achieves F1-score and mIoU of 0.9112, 0.9271, 0.8106, 0.9318, and 0.8369, 0.8644, 0.6815, 0.8723, respectively, on the four datasets. A lightweight version of the proposed network is constructed using depthwise separable convolution that achieves excellent performance with only 0.88M parameters.
Pengtao Li, Meihua Wang, Zhun Fan, Han Huang 0002, Guijie Zhu, Jiafan Zhuang
IEEE Trans. Intell. Transp. Syst.6
2023 Exploit Domain-Robust Optical Flow in Domain Adaptive Video Semantic Segmentation
abstract
Domain adaptive semantic segmentation aims to exploit the pixel-level annotated samples on source domain to assist the segmentation of unlabeled samples on target domain. For such a task, the key is to construct reliable supervision signals on target domain. However, existing methods can only provide unreliable supervision signals constructed by segmentation model (SegNet) that are generally domain-sensitive. In this work, we try to find a domain-robust clue to construct more reliable supervision signals. Particularly, we experimentally observe the domain-robustness of optical flow in video tasks as it mainly represents the motion characteristics of scenes. However, optical flow cannot be directly used as supervision signals of semantic segmentation since both of them essentially represent different information. To tackle this issue, we first propose a novel Segmentation-to-Flow Module (SFM) that converts semantic segmentation maps to optical flows, named the segmentation-based flow (SF), and then propose a Segmentation-based Flow Consistency (SFC) method to impose consistency between SF and optical flow, which can implicitly supervise the training of segmentation model. The extensive experiments on two challenging benchmarks demonstrate the effectiveness of our method, and it outperforms previous state-of-the-art methods with considerable performance improvement. Our code is available at https://github.com/EdenHazardan/SFC.
Zilei Wang, Jiafan Zhuang, Yixin Zhang 0007, Junjie Li 0002
AAAI3
2023 Towards Effective Instance Discrimination Contrastive Loss for Unsupervised Domain Adaptation
abstract
Domain adaptation (DA) aims to transfer knowledge from a label-rich source domain to a related but label-scarce target domain. Recently, increasing research has focused on exploring data structure of the target domain. In light of the recent success of Instance Discrimination Contrastive (IDCo) loss in self-supervised learning, we try directly applying it to domain adaptation tasks. However, the improvement is very limited, which motivates us to rethink its underlying limitations for domain adaptation tasks. An intuitive limitation is that a pair of samples belonging to the same class could be treated as negatives. Here we argue that using low-confidence samples to construct positive and negative pairs can alleviate this issue and is more suitable for IDCo loss. Another limitation is that IDCo loss cannot capture enough semantic information. We address this by introducing domain-invariant and accurate semantic information from classifier weights and input data. Specifically, we propose a class relationship enhanced features. It uses probability weighted class prototpyes as the input features of IDCo loss, which can implicitly transfer the domain-invariant class relationship. We further propose a target-dominated cross-domain mixup that can incorporate accurate semantic information from the source domain. We evaluate the proposed method in unsupervised DA and other DA settings, and extensive experimental results reveal that our method can make IDCo loss more effective and achieve state-of-the-art performance.1
Yixin Zhang 0007, Zilei Wang, Junjie Li 0002, Jiafan Zhuang
ICCV4
2023 Frequency and content dual stream network for image dehazing
abstract
Image dehazing can improve image clarity and visual effect, which plays a pivotal role in many computer vision tasks. Existing dehazing methods are mostly based on a single feature stream and tend to ignore the low-frequency characteristics of haze. In this paper, we propose a dual stream network for image dehazing. To enhance the edge information and texture detail of the image, we construct a frequency stream based on attention octave convolution. We decompose the features into high and low-frequency branches in the frequency stream to obtain different structural information. By adding a residual channel attention block, the attention octave convolution can extract frequency features more efficiently and effectively. Due to the lower resolution of low-frequency features in the frequency stream, the frequency stream features alone are insufficient for recovering the overall content of the image. Therefore, a content stream was added to compensate for the information lost in the frequency stream. By fusing the outputs of two feature streams, the network achieves an enhanced dehazing performance. The results show that our method is superior to other state-of-the-art algorithms in quantitative evaluation and visual impact.
Meihua Wang, De Huang, Zhun Fan, Jiafan Zhuang
Image Vis. Comput.5
2022 Semi-Supervised Video Semantic Segmentation with Inter-Frame Feature Reconstruction
abstract
One major challenge for semantic segmentation in realworld scenarios is only limited pixel-level labels available due to high expense of human labor though a vast volume of video data is provided. Existing semi-supervised methods attempt to exploit unlabeled data in model training, but they just regard video as a set of independent images. To better explore semi-supervised segmentation problem with video data, we formulate a semi-supervised video semantic segmentation task in this paper. For this task, we observe that the overfitting is surprisingly severe between labeled and unlabeled frames within a training video although they are very similar in style and contents. This is called inner-video overfitting, and it would actually lead to inferior performance. To tackle this issue, we propose a novel interframe feature reconstruction (IFR) technique to leverage the ground-truth labels to supervise the model training on unlabeled frames. IFR is essentially to utilize the internal relevance of different frames within a video. During training, IFR would enforce the feature distributions between labeled and unlabeled frames to be narrowed. Consequently, the inner-video overfitting issue can be effectively alleviated. We conduct extensive experiments on Cityscapes and CamVid, and the results demonstrate the superiority of our proposed method to previous state-of-the-art methods. The code is available at https://github.com/jfzhuang/IFR.
Jiafan Zhuang, Zilei Wang
CVPR1
2021 Efficient License Plate Recognition via Holistic Position Attention
abstract
License plate recognition (LPR) is a fundamental component of various intelligent transportation systems, and is always expected to be accurate and efficient enough in real-world applications. Nowadays, recognition of single character has been sophisticated benefiting from the power of deep learning, and extracting position information for forming a character sequence becomes the main bottleneck of LPR. To tackle this issue, we propose a novel holistic position attention (HPA) in this paper that consists of position network and shared classifier. Specifically, the position network explicitly encodes the character position into the maps of HPA, and then the shared classifier performs the character recognition in a unified and parallel way. Here the extracted features are modulated by the attention maps before feeding into the classifier to yield the final recognition results. Note that our proposed method is end-to-end trainable, character recognition can be concurrently performed, and no post-processing is needed. Thus our LPR system can achieve good effectiveness and efficiency simultaneously. The experimental results on four public datasets, including AOLP, Media Lab, CCPD, and CLPD, well demonstrate the superiority of our method to previous state-of-the-art methods in both accuracy and speed.
Yesheng Zhang, Zilei Wang, Jiafan Zhuang
AAAI3
2021 Video Semantic Segmentation With Distortion-Aware Feature Correction
abstract
Video semantic segmentation aims to generate an accurate semantic map for each frame in a video. For such a task, conducting per-frame image segmentation is generally unacceptable in practice due to high computation cost. To address this issue, many works perform the flow-based feature propagation to reuse the features of previous frames, which essentially exploits the content continuity of consecutive frames. However, the estimated optical flow would inevitably suffer inaccuracy and then make the propagated features distorted. In this article, we propose a distortion-aware feature correction method with the goal of improving video segmentation performance at a low price. Our core idea is to correct the features on distorted regions using the current frame while reserving the propagated features for other regions. In this way, a lightweight network is enough for achieving promising segmentation results. In particular, we propose to predict the distorted regions by utilizing the consistency of distortion patterns in images and features, such that the high-cost feature extraction from current frames can be avoided. We conduct extensive experiments on Cityscapes, CamVid, and UAVid, and the results show that our proposed method significantly outperforms previous methods and achieves the state-of-the-art performance on both segmentation accuracy and speed. Code and pretrained models are available at https://github.com/jfzhuang/DAVSS.
Jiafan Zhuang, Zilei Wang, Bingke Wang
IEEE Trans. Circuits Syst. Video Technol.1
2018 Towards Human-Level License Plate Recognition
Jiafan Zhuang, Saihui Hou, Zilei Wang, Zhengjun Zha
ECCV (3)1