Jun Zhang 0034

dblp:29/4190-34 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0003-1321-6022ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 LBMS-SAM: Segment anything model guided SEM image segmentation for lithium battery materials
Jun Zhang 0034, Jian Kuang 0005, Tingting Ren, Zhuanti Wu, Qiaqia Zhang
Neural Networks2
2025 Frequency Decoupled Domain-Irrelevant Feature Learning for Pan-Sharpening
abstract
Pan-sharpening aims to generate high-detail multi-spectral images (HRMS) through the fusion of panchromatic (PAN) and multi-spectral (MS) images. However, existing pan-sharpening methods often suffer from significant performance degradation when dealing with out-of-distribution data, as they assume the training and test datasets are independent and identically distributed. To overcome this challenge, we propose a novel frequency domain-irrelevant feature learning framework that exhibits exceptional generalization capabilities. Our approach involves parallel extraction and processing of domain-irrelevant information from the amplitude and phase components of the input images. Specifically, we design a frequency information separation module to extract the amplitude and phase components of the paired images. The learnable high-pass filter is then employed to eliminate domain-specific information from the amplitude spectrums. After that, we devised two specialized sub-networks (AFL-Net and PFL-Net) to perform targeted learning of the frequency domain-irrelevant information. This allows our method to effectively capture the complementary domain-irrelevant information contained in the amplitude and phase spectra of the images. Finally, the information fusion and restoration module dynamically adjusts the feature channel weights, enabling the network to output high-quality HRMS images. Through this frequency domain-irrelevant feature learning framework, our method balances generalization capability and network performance on the distribution of training dataset. Extensive experiments conducted on various satellite datasets demonstrate the effectiveness of our method for generalized pan-sharpening. Our proposed network outperforms state-of-the-art methods in terms of both quantitative metrics and visual quality, showcasing its superior ability to handle diverse, out-of-distribution data.
Jie Zhang 0033, Ke Cao 0001, Yunlong Lin, Xuanhua He, Yingying Wang 0005, Rui Li 0027, Chengjun Xie, Jun Zhang 0034, Man Zhou 0003
IEEE Trans. Circuits Syst. Video Technol.9
2024 Strong-Help-Weak: An Online Multi-Task Inference Learning Approach for Robust Advanced Driver Assistance Systems
abstract
Multi-task learning in advanced driver assistance systems aims to endow models with the capacity to jointly handle multiple related tasks, such as object detection, depth estimation, and more. However, existing multi-task learning models largely rely on the extensive number of labelled data. In practice, the process of annotating data for multi-task training proves to be exceedingly costly, yet not always accurate. This study introduces an innovative setting named online multi-task inference learning that updates the multi-task model during inference. And we propose a Strong-Help-Weak (SHW) framework which aims to enhance weaker (or more challenging) tasks by leveraging guidance from closely related stronger (or easier) tasks. Specifically, we first build two benchmarks based on KITTI and BDD with four tasks (object detection, object depth estimation, lane line segmentation, and driving area segmentation). Then, we propose two novel modules inspired by two priors: 1) Detection-guided Depth Inference Learning (DetDis) module that leverages the inverse relationship between object size and distance to refine the predicted object distance; and 2) Area-guided Lane Line Inference Learning (AreaLane) module that utilises inclusion relationship between driving area and lane line to infer more accurate lane line. Both modules are efficient and can provide more reliable supervision for the corresponding weaker tasks (object distance estimation and lane line segmentation), respectively. Extensive experiments on the two benchmarks show that our SHW can obtain consistent improvements on the weaker tasks during the inference stage with low computational costs.
Wenjing Li 0005, Jian Kuang 0005, Jun Zhang 0034, ZhongCheng Wu, Mahdi Rezaei 0001
IEEE Trans. Intell. Transp. Syst.4
2024 MIFI: MultI-Camera Feature Integration for Robust 3D Distracted Driver Activity Recognition
abstract
Distracted driver activity recognition plays a critical role in risk aversion-particularly beneficial in intelligent transportation systems. However, most existing methods make use of only the video from a single view and the difficulty-inconsistent issue is neglected. Different from them, in this work, we propose a novel MultI-camera Feature Integration (MIFI) approach for 3D distracted driver activity recognition by jointly modeling the data from different camera views and explicitly re-weighting examples based on their degree of difficulty. Our contributions are two-fold: (1) We propose a simple but effective multi-camera feature integration framework and provide three types of feature fusion techniques. (2) To address the difficulty-inconsistent problem in distracted driver activity recognition, a periodic learning method, named example re-weighting that can jointly learn the easy and hard samples, is presented. The experimental results on the 3MDAD dataset demonstrate that the proposed MIFI can consistently boost performance compared to single-view models.
Jian Kuang 0005, Wenjing Li 0005, Jun Zhang 0034, ZhongCheng Wu
IEEE Trans. Intell. Transp. Syst.4
2023 YOLOv4-dense: A smaller and faster YOLOv4 for real-time edge-device based object detection in traffic scene
abstract
Abstract Edge‐device‐based object detection is crucial in many real‐world applications, such as self‐driving cars, ADAS, driver behavior analysis. Although deep learning (DL) has become the de‐facto approach for object detection, the limited computing resources of embedded devices and the large model size of current DL‐based methods increase the difficulty of real‐time object detection on edge devices. To overcome these difficulties, in this work a novel YOLOv4‐dense model is proposed to detect objects in an accurate, fast manner, which is built on top of the YOLOv4 framework but with substantial improvements. More specifically, lots of CSP layers are pruned since it will decrease inference speed. And to address the losing small objects problem, a dense block is introduced. In addition, a lightweight two‐stream YOLO head is also designed to further reduce the computational complexity of the model. Experimental results on NVIDIA JETSON TX2 embedded platform demonstrate that YOLOv4‐dense can achieve a higher accuracy, faster speed with smaller model size. For instance, on the KITTI dataset, YOLOv4‐dense obtains 84.3% mAP and 22.6 FPS with only 20.3 M parameters, surpassing the state‐of‐the‐art models with comparable parameter budget such as YOLOv3‐tiny, YOLOv4‐tiny, PP‐YOLO‐tiny by a large margin.
Wenjing Li 0005, Jun Zhang 0034, ZhongCheng Wu
IET Image Process.3
2023 FPT: Fine-Grained Detection of Driver Distraction Based on the Feature Pyramid Vision Transformer
abstract
According to the surveys of the World Health Organization, distracted driving is one of main causes of road traffic accidents. To improve road traffic safety, real-time detection of drivers’ driving behavior is very important for the development of highly reliable Advanced Driver Assistance System (ADAS). At present, the deep learning architecture based on a Convolutional Neural Network (CNN) has disadvantages such as large number of parameters and weak global feature extraction ability. Therefore, this paper proposes an innovative driver distraction detection model based on the fusion of a transformer and a CNN, referred to as FPT, which is the first exploration in the field of driver distraction detection. First, we introduce the latest Twins transformer as a benchmark. Then, we design residual embedding to replace block embedding, which can further integrate the convolutional neural network with Transformer and improve the feature extraction ability. In addition, the Multilayer Perceptron (MLP) module with a large parameter occupancy rate in the original transformer structure is replaced with a lightweight group convolution module to reduce computational complexity. Finally, a cross-entropy loss function for label smoothing is designed to guide network learning with significantly differentiated features. Comparison results on two large-scale driver distraction detection datasets show that the proposed FPT offers a better compromise between computational cost and performance compared to the state-of-the-art CNN and Transformer architectures.
Jie Chen 0035, Zhixiang Huang, Bing Li 0033, Jianming Lv, Jingmin Xi, Bocai Wu, Jun Zhang 0034, ZhongCheng Wu
IEEE Trans. Intell. Transp. Syst.8
2023 100-Driver: A Large-Scale, Diverse Dataset for Distracted Driver Classification
abstract
Distracted driver classification (DDC) plays an important role in ensuring driving safety. Although many datasets are introduced to support the study of DDC, most of them are small in data size and are short of diversity in environmental variations. This largely limits the development of DDC since many practical problems such as the cross-modality setting cannot be fully studied. In this paper, we introduce 100-Driver, a large-scale, diverse posture-based distracted diver dataset, with more than 470K images taken by 4 cameras observing 100 drivers over 79 hours from 5 vehicles. 100-Driver involves different types of variations that closely meet real-world applications, including changes in the vehicle, person, camera view, lighting, and modality. We provide a detailed analysis of 100-Driver and present 4 settings for investigating practical problems of DDC, including the traditional setting without domain shift and 3 challenging settings (i.e., cross-modality, cross-view, and cross-vehicle) with domain shifts. We conduct comprehensive experiments on these 4 settings with state-the-of-art techniques and show several insights to the future study of DDC. Our 100-Driver will be publicly available offering new opportunities to advance the development of DDC. The 100-driver dataset, source code, and evaluation protocols are available athttps://100-driver.github.io.
Jing Wang 0092, Wenjing Li 0005, Jun Zhang 0034, ZhongCheng Wu, Zhun Zhong, Nicu Sebe
IEEE Trans. Intell. Transp. Syst.4
2022 Multilevel wavelet-based hierarchical networks for image compressed sensing
Zhu Yin, Wuzhen Shi, ZhongCheng Wu, Jun Zhang 0034
Pattern Recognit.4
2022 A New Unsupervised Deep Learning Algorithm for Fine-Grained Detection of Driver Distraction
abstract
Traffic accidents caused by distracted drivers account for a large proportion of traffic accidents each year, and monitoring the driving state of drivers to avoid traffic accidents caused by distracted driving has become a very important research direction. At present, the field of driver distraction detection mainly adopts supervised learning methods, which have problems such as poor generalization ability, large labeling cost, and weak artificial intelligence. This paper is oriented toward driver distraction fine-grained detection and innovatively proposes a new unsupervised deep learning algorithm, which is referred to as UDL, to achieve a more human-like level of intelligence. First, we build a new unsupervised deep learning algorithm; furthermore, we integrate the multilayer perceptron (MLP) architecture to build a new backbone and projection head to strengthen feature extraction capabilities; and finally, a new loss function based on contrast learning and a stop-gradient strategy is designed to guide the model to learn more robust features. The comparison results on large-scale driver distraction detection datasets show that our UDL method can accurately detect driver distraction without labels and exhibits excellent generalization performance with a linear evaluation accuracy of 97.38%; In addition, after fine-tuning with fewer labels, our UDL method can achieve superior performance close to state-of-the-art supervised learning methods, achieving 99.07% accuracy after fine-tuning using only 50% of the labeled data, which greatly reduces the cost and limitations of manual annotation.
Bing Li 0033, Jie Chen 0035, Zhixiang Huang, Jianming Lv, Jingmin Xi, Jun Zhang 0034, ZhongCheng Wu
IEEE Trans. Intell. Transp. Syst.7
2022 Learning Accurate, Speedy, Lightweight CNNs via Instance-Specific Multi-Teacher Knowledge Distillation for Distracted Driver Posture Identification
abstract
For deployment on an embedded processor for distracted driver classification, the model should satisfy the demand for both high accuracy, real-time inference, and limited storage resources. Conventional deep CNN models such as VGG, ResNet, DenseNet, often aim for high accuracy, making their model heavy for an embedded system with limited memory space and computing resources. In contrast, lightweight models are greatly compressed but at a significant sacrifice of accuracy. To bridge this gap, we propose an instance-specific multi-teacher knowledge distillation model (IsMt-KD) to learn more accurate, speedy, and lightweight CNNs for distracted driver posture classification. Specifically, in multi-teacher knowledge distillation, most of the current approaches either randomly select a teacher model and apply the prediction of such teacher model as the soft-label or allocate an equal weight to every teacher model and average all the predictions of the teachers as the soft label. In this paper, we observe that, when facing the same instance, the outputs of different teachers vary greatly, in which some teachers can predict it right whereas the others may give pretty high probabilities to the irrelevant classes. Thus, it is inappropriate to set fixed weights or the same weights for teachers. To this end, a simple yet effective instance-specific teacher grading module is designed to dynamically assign weights to teacher models based on individual instances. In this way, we can dynamically distill the knowledge from multiple teachers by considering both instance-specific high-level and instance-specific intermediate-level information. Our extensive experimental results on AUC and StateFarm datasets, and our implementation on edge hardware platforms including HUAWEI MediaPad c5 and Nvidia Jetson TX2, verify the effectiveness and feasibility of our approach.
Wenjing Li 0005, Jing Wang 0092, Tingting Ren, Jun Zhang 0034, ZhongCheng Wu
IEEE Trans. Intell. Transp. Syst.5
2021 Contextual similarity-based multi-level second-order attention network for semi-supervised few-shot learning
Wenjing Li 0005, Tingting Ren, Jun Zhang 0034, ZhongCheng Wu
Neurocomputing4
2020 LGSim: local task-invariant and global task-specific similarity for few-shot classification
Wenjing Li 0005, ZhongCheng Wu, Jun Zhang 0034, Tingting Ren
Neural Comput. Appl.3
2019 Mutual information-based dropout: Learning deep relevant feature representation architectures
Jie Chen 0035, ZhongCheng Wu, Jun Zhang 0034
Neurocomputing3
2019 Driving Safety Risk Prediction Using Cost-Sensitive With Nonnegativity-Constrained Autoencoders Based on Imbalanced Naturalistic Driving Data
abstract
A large number of studies have shown that most vehicle collisions are caused by drivers' abnormal operations. To ensure the safety of all people on the road network as much as possible, it is crucial to be able to predict the drivers' driving safety risks in real time. In this paper, we propose a novel cost-sensitive L1/L2-nonnegativity-constrained deep autoencoder network for driving safety risk prediction. Unfortunately, with existing research methods, the size of the sliding time window is too large, the feature extraction is relatively subjective, and class imbalances occur, which leads to low identification accuracy, long prediction times, and poor applicability. We first propose using a three-layer L1/L2-nonnegativity-constrained autoencoder to adaptively search the optimal size of the sliding window and then construct a deep L1/L2-nonnegativity-constrained autoencoder network to automatically extract the hidden features of the driving behaviors. Finally, we build a new L1/L2-nonnegativityconstrained focal loss classifier to predict the driving behaviors under different safety risk levels. The results from the public 100-Car naturalistic driving study dataset indicate that our method can effectively find the optimal window size, reduce the data volume and reconstruction error, and extract more distinctive features. Furthermore, this method effectively curbs the class imbalance, improves the driving safety risk prediction performance, reduces overfitting, shortens the prediction time, and improves the timeliness.
Jie Chen 0035, ZhongCheng Wu, Jun Zhang 0034
IEEE Trans. Intell. Transp. Syst.3
2018 Cross-covariance regularized autoencoders for nonredundant sparse feature representation
Jie Chen 0035, ZhongCheng Wu, Jun Zhang 0034, Wenjing Li 0005
Neurocomputing3