Kun Long

dblp:252/4871 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0006-7495-3201ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Computer networks · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Cross-scale hybrid attention network for enhancing performance prediction of modified asphalt binder preparation
Jiakang Zhang, Guoan Gan, Kun Long, Allen Zhang 0001, Chuanqi Yan, Changfa Ai
Eng. Appl. Artif. Intell.3
2026 LDEMF: A Lightweight Diffusion-Enhanced Multidomain Fusion Framework for Edge-Side Fault Diagnosis Within Distributed Device Clusters
abstract
Fault diagnosis in distributed device clusters (DDCs) is essential for ensuring safe and reliable operation of modern automated systems. However, practical deployments face two fundamental challenges: severe class imbalance from scarce device-level fault data and strict resource constraints on edge nodes. This paper proposes a Lightweight Diffusion-Enhanced Multi-Domain Fusion (LDEMF) framework that addresses both challenges through integrated data balancing, representation learning, and efficient deployment. A dual-attention conditional diffusion model synthesizes class-balanced vibration samples preserving temporal, spectral, and non-stationary fault characteristics. A multi-domain feature fusion network then integrates time-, frequency-, and time-frequency-domain representations via cross-attention for robust fault characterization. Finally, structured channel pruning, knowledge distillation, and FP16 post-training quantization enable resource-efficient edge deployment. Experimental results on the CWRU bearing dataset and a self-collected industrial robot dataset demonstrate that LDEMF achieves superior diagnostic accuracy under moderate-to-severe class imbalance. The framework reduces model size, computational cost, and inference latency by up to an order of magnitude with negligible performance loss, validating its effectiveness for edge-level fault diagnosis in DDCs.
Jiapeng Wu, Jianyu Long, Yaqiang Ji, Kun Long, Qiang Luo 0008, Chuan Li 0003
IEEE Internet Things J.4
2024 Supervised Hierarchical Online Hashing for Cross-modal Retrieval
abstract
Online cross-modal hashing has gained attention for its adaptability in processing streaming data. However, existing methods only define the hard similarity between data using labels. This results in poor retrieval performance, as they fail to exploit the semantic structure information of labels and miss the high-quality hash codes guided by the hierarchical relevance between labels. In addition, they ignore the bit-flipping problem, which leads to sub-optimal cross-modal retrieval performance. To address these issues, we propose Supervised Hierarchical Online Hashing (SHOH) for cross-modal retrieval. Our approach acquires hierarchical similarity via cross-layer affiliation of labels and explores its application to online hashing. We design a hierarchical similarity learning method in the online learning framework, which includes virtual center learning and hierarchical similarity embedding. Labels with soft similarity bridge the label hierarchy and cross-modal hash embedding. Furthermore, we propose a Weighted Retrieval Strategy (WRS) to mitigate the impact caused by bit-flipping errors. Extensive experiments and verification on hierarchical and non-hierarchical datasets demonstrate that SHOH preserves accurate inter-class distances and achieves performance improvements compared to state-of-the-art methods. The source code is available at https://github.com/HUST-IDSM-AI/SHOH .
Yu Liu 0040, Rukai Wei, Ke Zhou 0001, Kun Long
ACM Trans. Multim. Comput. Commun. Appl.6
2023 Multiple Domain-Adversarial Ensemble Learning for Domain Generalization
abstract
Domain generalization (DG) aims to train a model on multiple source domains which can be well generalized to the unseen target domains. Currently, most DG techniques in vision tasks mainly focus on one of the three research lines (i.e., data manipulation, representation learning, and learning strategy). The direction of combining multiple DG research lines still remains to be studied and explored. In this paper, we propose a unified framework for DG that combines multiple research lines to enhance the generalization ability. To be specific, we combine representative strategies of three research lines: ensemble learning, domain alignment, and data augmentation. A basic framework is built with an ensemble learning strategy to train an expert model in a domain unit. Besides, the ability to learn domain-independent features is enhanced with an adversarial-based domain alignment strategy. Furthermore, a domain transformation network (DoTNet) is introduced, which is expected to generate additional deviated images and further enhance the generalization capability. Extensive experiments on DG datasets demonstrate the effectiveness of our approach.
Ze-Yu Mi, Kun Long
ICASSP2
2022 AEBSR: Active-Sampling and Energy-Based Single Image Super-Resolution
abstract
Traditionally, single image super-resolution (SISR) methods randomly crop fixed-size patches in both low and high resolution (LR and HR) images as training samples, and obtain reconstruction model through the regression of LR-HR pixels pairs. However, these will lead to two problems. One is the negligence of the essential information of textures and edges leading to redundant and inefficient training. The other is the lack of the overall perception of data distribution causing poor generalization. To mitigate these issues, we propose Active Sampling and Energy-Based Single Image Super-Resolution (AEBSR), which introduces Active Sampling (AS) and Energy-Based Training (EBT) into SISR. Specifically, we first actively sample texture and edge patches through information entropy, and then align different data distributions between SR and HR images through free energy to perceive the overall distribution characteristics. Extensive experiments show that AS and EBT can further improve the SISR effect and our AEBSR also achieves competitive results compared to the current state-of-the-art SISR approaches.
Biao Jiang, Kun Long
ICIP2
2022 Unsupervised Unpaired Super-Resolution Using an Active Sampling Strategy Based on Edge Detection
abstract
Most existing super-resolution (SR) methods rely on pairs of low resolution (LR) and high resolution (HR) images and predetermined degradation operations (e.g., bicubic), usually trained by supervised learning. However, they often fail in real-world scenarios because of the occurrence of noise and blur. The key reason is that the degradation process is unknown, and no HR-LR pairs can be obtained directly. To address the above issues, this paper explores the optimization of an unsupervised unpaired SR method inspired by generative models such as Generative Adversarial Networks (GAN) and Cycle-Consistent Adversarial Networks (CycleGAN). We propose an active sampling strategy based on edge detection, and introduce a denoising network to construct a novel unsupervised unpaired SR framework. The active sampling strategy can perform image-level feature alignment by sampling image patches actively, thus optimizing the learning direction of the generator. At the same time, the denoising network can reduce the learning difficulty of the generator by preprocessing real-world LR images. Extensive experiments indicate that our method obtains better performance over other existing solutions to the unsupervised unpaired SR challenge,
Kun Long, Biao Jiang
IJCNN1
2022 Multi-feature Fusion VoteNet for 3D Object Detection
abstract
In this article, we propose a Multi-feature Fusion VoteNet (MFFVoteNet) framework for improving the 3D object detection performance in cluttered and heavily occluded scenes. Our method takes the point cloud and the synchronized RGB image as inputs to provide object detection results in 3D space. Our detection architecture is built on VoteNet with three key designs. First, we augment the VoteNet input with point color information to enhance the difference of various instances in a scene. Next, we integrate an image feature module into the VoteNet to provide a strong object class signal that can facilitate deterministic detections in occlusion. Moreover, we propose a Projection Non-Maximum Suppression (PNMS) method in 3D object detection to eliminate redundant proposals and hence provide more accurate positioning of 3D objects. We evaluate the proposed MFFVoteNet on two challenging 3D object detection datasets, i.e., ScanNetv2 and SUN RGB-D. Extensive experiments show that our framework can effectively improve the performance of 3D object detection.
Zhoutao Wang, Qian Xie 0001, Mingqiang Wei, Kun Long, Jun Wang 0039
ACM Trans. Multim. Comput. Commun. Appl.4
2021 MLVSNet: Multi-level Voting Siamese Network for 3D Visual Tracking
abstract
Benefiting from the excellent performance of Siamese-based trackers, huge progress on 2D visual tracking has been achieved. However, 3D visual tracking is still under-explored. Inspired by the idea of Hough voting in 3D object detection, in this paper, we propose a Multi-level Voting Siamese Network (MLVSNet) for 3D visual tracking from outdoor point cloud sequences. To deal with sparsity in outdoor 3D point clouds, we propose to perform Hough voting on multi-level features to get more vote centers and retain more useful information, instead of voting only on the fi-nal level feature as in previous methods. We also design an efficient and lightweight Target-Guided Attention (TGA) module to transfer the target information and highlight the target points in the search area. Moreover, we propose a Vote-cluster Feature Enhancement (VFE) module to exploit the relationships between different vote clusters. Extensive experiments on the 3D tracking benchmark of KITTI dataset demonstrate that our MLVSNet outperforms state-of-the-art methods with significant margins. Code will be available at https://github.com/CodeWZT/MLVSNet.
Zhoutao Wang, Qian Xie 0001, Yukun Lai, Jing Wu 0004, Kun Long, Jun Wang 0039
ICCV5
2021 Lesion-Inspired Denoising Network: Connecting Medical Image Denoising and Lesion Detection
abstract
Deep learning has achieved notable performance in the denoising task of low-quality medical images and the detection task of lesions, respectively. However, existing low-quality medical image denoising approaches are disconnected from the detection task of lesions. Intuitively, the quality of denoised images will influence the lesion detection accuracy that in turn can be used to affect the denoising performance. To this end, we propose a play-and-plug medical image denoising framework, namely Lesion-Inspired Denoising Network (LIDnet), to collaboratively improve both denoising performance and detection accuracy of denoised medical images. Specifically, we propose to insert the feedback of downstream detection task into existing denoising framework by jointly learning a multi-loss objective. Instead of using perceptual loss calculated on the entire feature map, a novel region-of-interest (ROI) perceptual loss induced by the lesion detection task is proposed to further connect these two tasks. To achieve better optimization for overall framework, we propose a customized collaborative training strategy for LIDnet. On consideration of clinical usability and imaging characteristics, three low-dose CT images datasets are used to evaluate the effectiveness of the proposed LIDnet. Experiments show that, by equipping with LIDnet, both of the denoising and lesion detection performance of baseline methods can be significantly improved.
Kecheng Chen, Kun Long, Yazhou Ren 0001, Xiaorong Pu
ACM Multimedia2
2021 Probability-based Mask R-CNN for pulmonary embolism detection
Kun Long, Xiaorong Pu, Yazhou Ren 0001, Mingxiu Zheng, Chunjiang Song, Su Han, Fengbin Deng
Neurocomputing1