Jun Kong 0004

dblp:95/203-4 · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0001-7095-2400ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 Temporal Action Localization with Cross Layer Task Decoupling and Refinement
abstract
Temporal action localization (TAL) involves dual tasks to classify and localize actions within untrimmed videos. However, the two tasks often have conflicting requirements for features. Existing methods typically employ separate heads for classification and localization tasks but share the same input feature, leading to suboptimal performance. To address this issue, we propose a novel TAL method with Cross Layer Task Decoupling and Refinement (CLTDR). Based on the feature pyramid of video, CLTDR strategy integrates semantically strong features from higher pyramid layers and detailed boundary-aware boundary features from lower pyramid layers to effectively disentangle the action classification and localization tasks. Moreover, the multiple features from cross layers are also employed to refine and align the disentangled classification and regression results. At last, a lightweight Gated Multi-Granularity (GMG) module is proposed to comprehensively extract and aggregate video features at instant, local, and global temporal granularities. Benefiting from the CLTDR and GMG modules, our method achieves state-of-the-art performance on five challenging benchmarks: THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, and HACS. Code:https://github.com/LiQiang0307/CLTDR-GMG
Di Liu 0004, Jun Kong 0004, Jianzhong Wang 0003
AAAI3
2025 BSDSGANet: Bidirectional Skip-stored Dual-Stream Gated Attention Network for multivariate time series classification
Yugen Yi, Panpan Zhao, Hui Sheng, Min Liu 0024, Jiangyan Dai, Jun Kong 0004, Shaojie Qiao
Knowl. Based Syst.7
2025 A Lightweight Hybrid Network for Object Detection in Remote Sensing Images Balancing Global and Local Information
abstract
In recent years, hybrid convolutional neural networks (CNNs) and Transformer-based object detection technologies have achieved remarkable success. In the field of remote sensing image detection, since remote sensing systems rely on large-scale deployment of edge devices, detection models need to be lightweight with low parameter complexity to adapt to resource-constrained environments. However, existing lightweight models often struggle with an imbalance in extracting low-frequency global and high-frequency local information. In particular, when processing high-frequency local information (such as edges, textures, and fine structures), these models often lack in-depth analysis, leading to insufficient extraction of local features and reduced detection accuracy. To address the imbalance between low-frequency global information and high-frequency local information in lightweight remote sensing models, we propose an efficient and lightweight hybrid network detection framework, which mainly consists of the Global-Local Balance (GLB) module and the Detail-Aware Feature Fusion (DAFF) module. The GLB module adopts dynamic weight adjustment and context-aware mechanisms to effectively aggregate high-frequency local information in the image. The DAFF module further enhances feature fusion and detail refinement, improving the model’s performance and generalization ability. Experimental results on remote sensing datasets, including RSOD, NWPU VHR-10, and LEVIR, demonstrate that our proposed method achieves a well-balanced trade-off between model size and detection accuracy, reaching state-of-the-art performance.
Shuting Huang, Huanzun Zhang, Guangzhen Yao, Sandong Zhu, Jun Kong 0004
IEEE Geosci. Remote. Sens. Lett.8
2025 SOD-YOLOv10: Small Object Detection in Remote Sensing Images Based on YOLOv10
abstract
YOLOv10, known for its efficiency in object detection methods, quickly and accurately detects objects in images. However, when detecting small objects in remote sensing imagery, traditional algorithms often encounter challenges like background noise, missing information, and complex multiobject interactions, which can affect detection performance. To address these issues, we propose an enhanced algorithm for detecting small objects, named SOD-YOLOv10. We design the Multidimensional Information Interaction for the Transformer Backbone (TransBone) Network, which enhances global perception capabilities and effectively integrates both local and global information, thereby improving the detection of small object features. We also propose a feature fusion technology using an attention mechanism, called aggregated attention in a gated feature pyramid network (AA-GFPN). This technology uses an efficient feature aggregation network and re-parameterization techniques to optimize information interaction between feature maps of different scales. Additionally, by incorporating the aggregated attention (AA) mechanism, it accurately identifies essential features of small objects. Moreover, we propose the adaptive focal powerful IoU (AFP-IoU) loss function, which not only prevents excessive expansion of the anchor box area but also significantly accelerates model convergence. To evaluate our method, we conduct thorough tests on the RSOD, NWPU VHR-10, VisDrone2019, and AI-TOD datasets. The findings indicate that our SOD-YOLOv10 model attains 95.90%, 92.46%, 55.61%, and 59.47% for [email protected] and 73.42%, 66.84%, 39.03%, and 42.67% for [email protected]:0.95.
Guangzhen Yao, Sandong Zhu, Jun Kong 0004
IEEE Geosci. Remote. Sens. Lett.6
2024 Self-ensembling with mask-boundary domain adaptation for optic disc and cup segmentation
Jun Kong 0004, Di Liu 0004, Caixia Zheng
Eng. Appl. Artif. Intell.2
2024 An Adaptive Dual Selective Transformer for Temporal Action Localization
abstract
Temporal action localization (TAL), which aims to identify and localize actions in long untrimmed videos, is a challenging task in video understanding. Recent studies have shown that the Transformer and its variants are effective at improving the performance of TAL. The success of the Transformer can be attributed to the use of multi-head self-attention (MHSA) as a token mixer to capture long-term temporal dependencies within the video sequence. However, in the existing Transformer architecture, the features obtained by multiple token mixing (i.e., self-attention) heads are treated equally, which neglects the distinct characteristics of different heads and hampers the exploitation of discriminative information. To this end, we present a new method called the adaptive dual selective Transformer (ADSFormer) for TAL in this paper. The key component in ADSFormer is the dual selective multihead token mixer (DSMHTM), which integrates multiple feature representations from different token mixing heads by adaptively selecting important features across both the head and channel dimensions. Moreover, we also incorporate our ADSFormer into a pyramid structure so that the multi-scale features obtained can be effectively combined to improve TAL performance. Benefiting from the dual selective multi-head token mixer (DSMHTM) and pyramid feature combination, ADSFormer outperforms several state-of-the-art methods on four challenging benchmark datasets: THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100 and ActivityNet-1.3. We publicly released our codes and pretrained models at: https://github.com/LiQiang0307/ADSFormer.
Guang Zu, Jun Kong 0004, Jianzhong Wang 0003
IEEE Trans. Multim.4
2023 Multi-Scale Frequency Separation Network for Image Deblurring
abstract
Image deblurring aims to restore the detailed texture information or structures from the blurry images, which has become an indispensable step in many computer vision tasks. Although various methods have been proposed to deal with the image deblurring problem, most of them treated the blurry image as a whole and neglected the characteristics of different image frequencies. In this paper, we present a new method called multi-scale frequency separation network (MSFS-Net) for image deblurring. MSFS-Net introduces the frequency separation module (FSM) into an encoder-decoder network architecture to capture the low- and high-frequency information of image at multiple scales. Then, a simple cycle-consistency strategy and a sophisticated contrastive learning module (CLM) are respectively designed to retain the low-frequency information and recover the high-frequency information during deblurring. At last, the features of different scales are fused by a cross-scale feature fusion module (CSFFM). Extensive experiments on benchmark datasets show that the proposed network achieves state-of-the-art performance.
Miao Qi, Di Liu 0004, Jun Kong 0004, Jianzhong Wang 0003
IEEE Trans. Circuits Syst. Video Technol.5
2022 Double Closed-Loop Network for Image Deblurring
abstract
In this paper, a deep learning network with double closed-loop structure is introduced to tackle the image deblurring problem. The first closed-loop in our model is composed of two networks which learn a pair of opposite mappings between the blurry and sharp images. By this way, the solution spaces of possible functions that map a blurry image to its sharp counterpart can be effectively reduced. Furthermore, the first closed-loop also helps our model to deal with the unpaired samples in the training set. The second closed-loop in the proposed approach employed a self-supervision mechanism to constrain the features of intermedia layers in the network, so that the detailed information of sharp images can be well exploited. Through combining the two closed-loops together, our model can address the limitations of existing methods and improve the deblurring performance. Extensive experiments on both benchmark and real-world datasets show that the proposed network achieves state-of-the-art performance. The code will be released in: https://github.com/LiQiang0307/DCLNet.
Jun Kong 0004, Miao Qi, Jianzhong Wang 0003
ICASSP4
2022 Bottom-up improved multistage temporal convolutional network for action segmentation
Miao Qi, Qi Pu, Jun Kong 0004, Caixia Zheng
Appl. Intell.6
2022 Gaze Estimation via the Joint Modeling of Multiple Cues
abstract
How to automatically predict people’s gaze has attracted attention in the field of computer vision and machine learning. Previous studies on this topic set many constraints, such as restricted scenarios and strict and complex inputs. To mitigate these constraints to predict the gaze of people in more general scenarios, we propose a three-pathway network (TPNet) to estimate gaze via the joint modeling of multiple cues. Specifically, we first design a human-centric relationship inference (HCRI) module to learn the object-level relationship between the target person and the surrounding persons/objects in a scene. To the best of our knowledge, this is the first time that the object-level relationship is introduced into the gaze estimation task. Then, we construct a novel deep network with three pathways to fuse multiple cues, including scene saliency, object-level relationships and head information, to predict the gaze target. In addition, to extract the multilevel features during network training, we build and embed a micropyramid module in TPNet. The performance of TPNet is evaluated on two gaze estimation datasets: GazeFollow and DLGaze. A large number of quantitative and qualitative experimental results verify that TPNet can obtain robust results and significantly outperform the existing state-of-the-art gaze estimation methods. The code of TPNet will be released later.
Chao Zhu 0002, Xiaoli Liu 0001, Yinghua Lu, Caixia Zheng, Jun Kong 0004
IEEE Trans. Circuits Syst. Video Technol.7
2021 Image Deblurring based on Lightweight Multi-Information Fusion Network
abstract
Recently, deep learning based image deblurring has been well developed. However, exploiting the detailed image features in a deep learning framework always requires a mass of parameters, which inevitably makes the network suffer from high computational burden. To solve this problem, we propose a lightweight multi-information fusion network (LMFN) for image deblurring. The proposed LMFN is designed as an encoder-decoder architecture. In the encoding stage, the image feature is reduced to various small-scale spaces for multi-scale information extraction and fusion without a large amount of information loss. Then, a distillation network is used in the decoding stage, which allows the network benefit the most from residual learning while remaining sufficiently lightweight. Meanwhile, an information fusion strategy between distillation modules and feature channels is also carried out by attention mechanism. Through fusing different information in the proposed approach, our network can achieve state-of-the-art image deblurring result with smaller number of parameters and outperforms existing methods in model complexity.
Miao Qi, Dahong Xu, Jun Kong 0004, Jianzhong Wang 0003
ICIP6
2020 Non-Negative Matrix Factorization With Locality Constrained Adaptive Graph
abstract
Non-negative matrix factorization (NMF) has recently attracted much attention due to its good interpretation in perception science and widely applications in various fields. In this paper, a novel graph regularized NMF algorithm called NMF with locality constrained adaptive graph (NMF-LCAG) is proposed. Compared with other NMF based algorithms, the proposed NMF-LCAG algorithm has the following advantages: 1) Unlike the traditional NMF method which neglects the geometric information of original data, the proposed algorithm introduces a locality constrained graph to discover the latent manifold structure of the data and 2) Different from most graph regularized NMF algorithms in which the graphs are predefined and kept unchanged during the NMF procedure, two locality constraint terms are employed in our NMF-LCAG to adaptively optimize the graph. Thus, the weight matrix of graph and low dimensional features of data can be simultaneously learned by our algorithm, which makes NMF-LCAG more flexible than other approaches. Moreover, an iterative updating strategy is developed to optimize the objective function of our algorithm and the convergence analysis is also given. Extensive experiments are conducted on four face image databases and three UCI datasets to demonstrate the effectiveness of the proposed NMF-LCAG algorithm. Compared with some other related algorithms, the proposed NMF-LCAG algorithm can achieve at least 1% ~ 3% accuracy improvement in most cases.
Yugen Yi, Jianzhong Wang 0003, Wei Zhou 0003, Caixia Zheng, Jun Kong 0004, Shaojie Qiao
IEEE Trans. Circuits Syst. Video Technol.5
2019 Joint graph optimization and projection learning for dimensionality reduction
Yugen Yi, Jianzhong Wang 0003, Wei Zhou 0003, Jun Kong 0004, Yinghua Lu
Pattern Recognit.5
2017 Locality constrained Graph Optimization for Dimensionality Reduction
Jianzhong Wang 0003, Caixia Zheng, Jun Kong 0004, Yugen Yi
Neurocomputing5
2015 Semi-supervised local ridge regression for local matching based face recognition
Yugen Yi, Chao Bi, Jianzhong Wang 0003, Jun Kong 0004
Neurocomputing5
2015 Label propagation based semi-supervised non-negative matrix factorization for feature extraction
Yugen Yi, Yanjiao Shi, Jianzhong Wang 0003, Jun Kong 0004
Neurocomputing5
2015 An improved locality sensitive discriminant analysis approach for feature extraction
Yugen Yi, Baoxue Zhang, Jun Kong 0004, Jianzhong Wang 0003
Multim. Tools Appl.3
2014 A novel image retrieval method based on hybrid information descriptors
Ke Zhang 0023, Qinghe Feng, Jianzhong Wang 0003, Jun Kong 0004, Yinghua Lu
J. Vis. Commun. Image Represent.5
2013 A Novel Image Retrieval Method Based on Mutual Information Descriptors
Gang Hou, Ke Zhang 0023, Jun Kong 0004
ICIC (2)4
2013 Maximum weight and minimum redundancy: A novel framework for feature subset selection
Jianzhong Wang 0003, Lishan Wu, Jun Kong 0004, Yuxin Li 0001, Baoxue Zhang
Pattern Recognit.3
2011 Content-Based Biometric Image Hiding Approach
abstract
Recently, the use of information hiding techniques to protect biometric data has been an active topic. This paper proposes a novel image hiding approach based on correlation analysis to protect network-based transmitted biometric image for identification. Firstly, the correlation between the biometric image and the cover image is analyzed using principal component analysis (PCA) and genetic algorithm (GA). The purpose of correlation analysis is to enable the cover image to represent the secret image in content as much as possible, not just as a carrier of hidden information. Then, the unrepresented part of the biometric image, as the secret image, is encrypted and hidden into the middle-significant-bit plane (MSB) of the cover image redundantly. Extensive experimental results demonstrate that the proposed hiding approach not only gains good imperceptibility, but also resists some common attacks validated by the biometric identification accuracy.
Miao Qi, Jun Kong 0004, Yinghua Lu, Ning Du, Zhiqiang Ma 0003
Int. J. Pattern Recognit. Artif. Intell.2
2011 A structure-preserved local matching approach for face recognition
Jianzhong Wang 0003, Zhiqiang Ma 0003, Baoxue Zhang, Miao Qi, Jun Kong 0004
Pattern Recognit. Lett.5
2010 Linear discriminant projection embedding based on patches alignment
Jianzhong Wang 0003, Baoxue Zhang, Miao Qi, Jun Kong 0004
Image Vis. Comput.4
2010 An adaptively weighted sub-pattern locality preserving projection for face recognition
Jianzhong Wang 0003, Baoxue Zhang, Shuyan Wang, Miao Qi, Jun Kong 0004
J. Netw. Comput. Appl.5
2008 Local-based fuzzy clustering for segmentation of MR brain images
abstract
Accurate segmentation of magnetic resonance images (MRI) corrupted by intensity inhomogeneity is a challenging problem and has received an enormous amount of attention lately. On the basis of the local image model, we propose a different segmentation method for MR brain images without estimation and correction for intensity heterogeneity. Firstly, we obtain clustering context based on the distributing disciplinarian in anatomy that gray matter (GM) is always between white matter (WM) and cerebrospinal fluid (CSF) in brain, which ensure the three tissues exist together in each one. Then the size of the context is optimized by a minimum entropy criterion. Finally, FCM algorithm is independently performed in each context to calculate the degree of membership of a pixel to each tissue class. The proposed methodology has been evaluated for simulated images and shown the better results.
Jianzhong Wang 0003, Lili Dou, Na Che, Di Liu 0004, Baoxue Zhang, Jun Kong 0004
BIBE6
2008 Graph theory based algorithm for magnetic resonance brain images segmentation
abstract
Image segmentation is often required as a preliminary and indispensable stage in the computer aided medical image process, particularly during the clinical analysis of magnetic resonance (MR) brain images. The segmentation of magnetic resonance image (MRI) is a challenging problem that has received an enormous amount of attention lately. In this paper, we propose a simple and effective segmentation method combining watershed algorithm and normalized cuts (CWNC) for MR brain images. An initial partitioning of the MRI into primitive regions is set by applying the watershed transform. The latter process uses a region similarity graph representation of the image regions. And then the graph is segmented by normalized cuts algorithm. The efficacy of the proposed algorithm is demonstrated by extensive segmentation experiments using both simulated and real MR images and by comparison with other published algorithms.
Jianzhong Wang 0003, Di Liu 0004, Lili Dou, Baoxue Zhang, Jun Kong 0004, Yinghua Lu
BIBE5
2007 A Modified Fuzzy Kohonen's Competitive Learning Algorithms Incorporating Local Information for MR Image Segmentation
abstract
A Modified FKCL (MFKCL) algorithm for automatic segmentation of MR brain images is proposed in this paper. This algorithm is an extension of traditional fuzzy Kohonen's competitive learning algorithm. In our method, a factor that can estimate the effect of the neighbor pixels to the central pixel is introduced into the objective function of the standard FKCL algorithm as the local information. The local information is applied to trail off the effect of noise to the result of MRI segmentation. Experiments with simulated MR data and real MR data show that our algorithm can resist not only the little, but also the heavy noise compared with standard FKCL segmentation and other reported methods.
Jun Kong 0004, Wenjing Lu, Jianzhong Wang 0003, Na Che, Yinghua Lu
BIBE1
2006 A Novel Color Image Watermarking Method Based on Genetic Algorithm and Neural Networks
Jialing Han, Jun Kong 0004, Yinghua Lu, Gang Hou
ICONIP (3)2
2006 A Novel Automated Hand-Based Personal Identification
Yinghua Lu, Yuru Wang, Jun Kong 0004, Longkui Jiang
IWCIA3