Jianzhong Wang 0003

dblp:94/178-3 · DBLP profile ↗
← Back
37ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-6867-3282ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 RCLEAF: Reliable contrastive learning-driven efficient adaptive fusion for multi-view clustering
Yugen Yi, Litao Huang, Jingkai Guo, Yali Peng 0001, Wei Zhou 0003, Jianzhong Wang 0003
Knowl. Based Syst.7
2026 MG-Mono: A lightweight multi-granularity method for self-supervised monocular depth estimation
Yugen Yi, Jianzhong Wang 0003
Pattern Recognit.5
2025 Temporal Action Localization with Cross Layer Task Decoupling and Refinement
abstract
Temporal action localization (TAL) involves dual tasks to classify and localize actions within untrimmed videos. However, the two tasks often have conflicting requirements for features. Existing methods typically employ separate heads for classification and localization tasks but share the same input feature, leading to suboptimal performance. To address this issue, we propose a novel TAL method with Cross Layer Task Decoupling and Refinement (CLTDR). Based on the feature pyramid of video, CLTDR strategy integrates semantically strong features from higher pyramid layers and detailed boundary-aware boundary features from lower pyramid layers to effectively disentangle the action classification and localization tasks. Moreover, the multiple features from cross layers are also employed to refine and align the disentangled classification and regression results. At last, a lightweight Gated Multi-Granularity (GMG) module is proposed to comprehensively extract and aggregate video features at instant, local, and global temporal granularities. Benefiting from the CLTDR and GMG modules, our method achieves state-of-the-art performance on five challenging benchmarks: THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, and HACS. Code:https://github.com/LiQiang0307/CLTDR-GMG
Di Liu 0004, Jun Kong 0004, Jianzhong Wang 0003
AAAI6
2025 EMLFCL: An Efficient Multilevel Fusion Contrastive Learning for Multiview Clustering
abstract
Multiview clustering (MVC) with contrastive learning (CL) has attracted considerable interest. Nevertheless, current methods have specific drawbacks since the coherence between views in them is limited either at the feature representation level or the cluster representation level. Besides, certain methods demonstrate subpar performance and limited robustness when handling noisy data. This article introduces an efficient multilevel fusion CL framework for MVC called EMLFCL. The EMLFCL model seamlessly incorporates a shared multi-layer perceptron (MLP) network (MNet) and a fusion network (FNet) to capture and merge common representation information, which effectively eliminates the impact of view-specific private information during the clustering process. Specifically, we establish an efficient multilevel CL strategy at both the feature representation level and the clustering representation level. Rather than rely on pairwise comparisons between views, our proposed CL strategy makes comparisons between different views and the anchor view. Since the anchor view contains abundant shared information, this strategy effectively mitigates the influence of view-specific and noisy view information on model performance. The proposed method outperforms numerous advanced approaches, as evidenced by extensive experiments conducted on eleven challenging multiview datasets. Particularly, it achieves 66.4%, 74.7%, 82.3%, and 86.4% clustering accuracies on the four Caltech datasets with different views, respectively.
Yugen Yi, Ningyi Zhang, Yijian Fu, Jianzhong Wang 0003
IEEE Trans. Neural Networks Learn. Syst.6
2025 Multigranularity Feature Aggregation and Cross-level Boundary Modeling for Temporal Action Detection
abstract
This article presents a Temporal Action Detection (TAD) method with Multigranularity (MG) feature aggregation and Cross-level Boundary Modeling (CBM). Compared with other methods, our proposed approach has the following advantages. First, different from most existing works which only consider the local temporal context, a simple and computationally efficient MG module is proposed to comprehensively extract video features in instant, local, and global temporal granularities. Second, unlike the methods that only employ the information from single feature pyramid level for action boundary regression, a CBM strategy that integrates the relative information from both the same and higher level features is designed to improve the accuracy of boundary prediction. At lastfere, benefiting from the MG module and CBM strategy, our method outperforms other state-of-the-art approaches on five challenging TAD datasets: THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, and HACS. We make our code and pre-trained model publicly available at: https://github.com/MGCBM/TAL-MGCBM
Di Liu 0004, Guang Zu, Jianzhong Wang 0003
ACM Trans. Multim. Comput. Commun. Appl.6
2024 GPONet: A two-stream gated progressive optimization network for salient object detection
Yugen Yi, Ningyi Zhang, Wei Zhou 0003, Yanjiao Shi, Gengsheng Xie, Jianzhong Wang 0003
Pattern Recognit.6
2024 An Adaptive Dual Selective Transformer for Temporal Action Localization
abstract
Temporal action localization (TAL), which aims to identify and localize actions in long untrimmed videos, is a challenging task in video understanding. Recent studies have shown that the Transformer and its variants are effective at improving the performance of TAL. The success of the Transformer can be attributed to the use of multi-head self-attention (MHSA) as a token mixer to capture long-term temporal dependencies within the video sequence. However, in the existing Transformer architecture, the features obtained by multiple token mixing (i.e., self-attention) heads are treated equally, which neglects the distinct characteristics of different heads and hampers the exploitation of discriminative information. To this end, we present a new method called the adaptive dual selective Transformer (ADSFormer) for TAL in this paper. The key component in ADSFormer is the dual selective multihead token mixer (DSMHTM), which integrates multiple feature representations from different token mixing heads by adaptively selecting important features across both the head and channel dimensions. Moreover, we also incorporate our ADSFormer into a pyramid structure so that the multi-scale features obtained can be effectively combined to improve TAL performance. Benefiting from the dual selective multi-head token mixer (DSMHTM) and pyramid feature combination, ADSFormer outperforms several state-of-the-art methods on four challenging benchmark datasets: THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100 and ActivityNet-1.3. We publicly released our codes and pretrained models at: https://github.com/LiQiang0307/ADSFormer.
Guang Zu, Jun Kong 0004, Jianzhong Wang 0003
IEEE Trans. Multim.6
2023 RRNMF-MAGL: Robust regularization non-negative matrix factorization with multi-constraint adaptive graph learning for dimensionality reduction
Yugen Yi, Shumin Lai, Jiangyan Dai, Wenle Wang, Jianzhong Wang 0003
Inf. Sci.6
2023 Multi-Scale Frequency Separation Network for Image Deblurring
abstract
Image deblurring aims to restore the detailed texture information or structures from the blurry images, which has become an indispensable step in many computer vision tasks. Although various methods have been proposed to deal with the image deblurring problem, most of them treated the blurry image as a whole and neglected the characteristics of different image frequencies. In this paper, we present a new method called multi-scale frequency separation network (MSFS-Net) for image deblurring. MSFS-Net introduces the frequency separation module (FSM) into an encoder-decoder network architecture to capture the low- and high-frequency information of image at multiple scales. Then, a simple cycle-consistency strategy and a sophisticated contrastive learning module (CLM) are respectively designed to retain the low-frequency information and recover the high-frequency information during deblurring. At last, the features of different scales are fused by a cross-scale feature fusion module (CSFFM). Extensive experiments on benchmark datasets show that the proposed network achieves state-of-the-art performance.
Miao Qi, Di Liu 0004, Jun Kong 0004, Jianzhong Wang 0003
IEEE Trans. Circuits Syst. Video Technol.6
2022 Double Closed-Loop Network for Image Deblurring
abstract
In this paper, a deep learning network with double closed-loop structure is introduced to tackle the image deblurring problem. The first closed-loop in our model is composed of two networks which learn a pair of opposite mappings between the blurry and sharp images. By this way, the solution spaces of possible functions that map a blurry image to its sharp counterpart can be effectively reduced. Furthermore, the first closed-loop also helps our model to deal with the unpaired samples in the training set. The second closed-loop in the proposed approach employed a self-supervision mechanism to constrain the features of intermedia layers in the network, so that the detailed information of sharp images can be well exploited. Through combining the two closed-loops together, our model can address the limitations of existing methods and improve the deblurring performance. Extensive experiments on both benchmark and real-world datasets show that the proposed network achieves state-of-the-art performance. The code will be released in: https://github.com/LiQiang0307/DCLNet.
Jun Kong 0004, Miao Qi, Jianzhong Wang 0003
ICASSP6
2022 Joint Semantic Segmentation and Object Detection Based on Relational Mask R-CNN
Jingxuan Fan, Miao Qi, Jianzhong Wang 0003
ICIC (1)6
2022 Deep sparse autoencoder integrated with three-stage framework for glaucoma diagnosis
abstract
Recently, end-to-end deep neural networks-based glaucoma diagnosis approaches have been gaining much attention. However, the feature extractor and classier in these approaches are trained together, which is known as coadaptation. Therefore, the feature distribution in them should adapt to particular decision boundaries. To learn generic data representations and improve the generalization ability of the model, this paper designs a three-stage framework for glaucoma diagnosis. In the first stage, preprocessing is utilized to extract the Region of Interesting around the Optic Disc to reduce the computational cost and nonobjective interference. In the second stage, Deep Sparse Autoencoder is designed to learn hybrid features between the deep features and the original features, which could improve the effectiveness of final high-level feature expression. Meanwhile, L1 regularization is introduced and applied on the hybrid features to obtain deep features with high complementarity under small sample problem. In the third stage, the obtained generic feature representations are fed into different classifiers, in which Support Vector Machine classifier achieves the best diagnosis performance. The proposed approach is evaluated on two publicly available databases. Extensive experimental results indicate that our approach outperforms the state-of-the-art approaches with the accuracy of 96.00%, 97.00% and Area Under Curve of 96.94%, 98.28% for REFUGE and Drishti-GS1 databases, respectively.
Wenle Wang, Wei Zhou 0003, Jianhang Ji, Jikun Yang, Wei Guo 0016, Zhaoxuan Gong, Yugen Yi, Jianzhong Wang 0003
Int. J. Intell. Syst.8
2022 SDNMF: Semisupervised discriminative nonnegative matrix factorization for feature learning
abstract
As one of the most effective feature learning methods, Nonnegative Matrix Factorization (NMF) has been widely used in many scientific fields, such as computer vision, data mining, and bioinformatics. However, NMF is an unsupervised method that cannot fully utilize the label information of data. Thus, its performance is limited in some recognition and classification problems. To remedy this shortcoming, this paper proposes a Semisupervised Discriminative NMF (SDNMF) method. First, we design a Soft-Labeled NMF (SLNMF) model by introducing a soft-label matrix-based regression term into the original NMF, so that the relationship between the soft-label matrix and low-dimensional features can be constructed to improve the discriminative ability of low-dimensional features. Second, to effectively estimate the soft-label matrix, a Label Propagation (LP) model is adopted to fully explore the spatial distribution relationship between the labeled and unlabeled samples. Third, an Adaptive Graph Learning (AGL) model is proposed to exploit the geometric relationship of samples well, which could enhance the performance of LP. Finally, the above three models (i.e., SLNMF, LP, and AGL) are integrated into a unified framework for effective feature learning, which can not only effectively explore the structural relationship matrix between data, but also predict the labels for unknown samples. Moreover, an iterative optimization algorithm is presented to solve our objective function. The convergence and computational complexity analysis of the proposed SDNMF method are also provided. Extensive experiments are conducted on several standard data sets. Compared with related methods, the experimental results verify that the proposed SDNMF method achieves better performance.
Yugen Yi, Shumin Lai, Wenle Wang, Renbo Zhang, Wei Zhou 0003, Jianzhong Wang 0003
Int. J. Intell. Syst.8
2021 Image Deblurring based on Lightweight Multi-Information Fusion Network
abstract
Recently, deep learning based image deblurring has been well developed. However, exploiting the detailed image features in a deep learning framework always requires a mass of parameters, which inevitably makes the network suffer from high computational burden. To solve this problem, we propose a lightweight multi-information fusion network (LMFN) for image deblurring. The proposed LMFN is designed as an encoder-decoder architecture. In the encoding stage, the image feature is reduced to various small-scale spaces for multi-scale information extraction and fusion without a large amount of information loss. Then, a distillation network is used in the decoding stage, which allows the network benefit the most from residual learning while remaining sufficiently lightweight. Meanwhile, an information fusion strategy between distillation modules and feature channels is also carried out by attention mechanism. Through fusing different information in the proposed approach, our network can achieve state-of-the-art image deblurring result with smaller number of parameters and outperforms existing methods in model complexity.
Miao Qi, Dahong Xu, Jun Kong 0004, Jianzhong Wang 0003
ICIP7
2020 Joint feature representation and classification via adaptive graph semi-supervised nonnegative matrix factorization
Yugen Yi, Yuqi Chen 0004, Jianzhong Wang 0003, Gang Lei 0002, Jiangyan Dai, Huihui Zhang 0003
Signal Process. Image Commun.3
2020 Non-Negative Matrix Factorization With Locality Constrained Adaptive Graph
abstract
Non-negative matrix factorization (NMF) has recently attracted much attention due to its good interpretation in perception science and widely applications in various fields. In this paper, a novel graph regularized NMF algorithm called NMF with locality constrained adaptive graph (NMF-LCAG) is proposed. Compared with other NMF based algorithms, the proposed NMF-LCAG algorithm has the following advantages: 1) Unlike the traditional NMF method which neglects the geometric information of original data, the proposed algorithm introduces a locality constrained graph to discover the latent manifold structure of the data and 2) Different from most graph regularized NMF algorithms in which the graphs are predefined and kept unchanged during the NMF procedure, two locality constraint terms are employed in our NMF-LCAG to adaptively optimize the graph. Thus, the weight matrix of graph and low dimensional features of data can be simultaneously learned by our algorithm, which makes NMF-LCAG more flexible than other approaches. Moreover, an iterative updating strategy is developed to optimize the objective function of our algorithm and the convergence analysis is also given. Extensive experiments are conducted on four face image databases and three UCI datasets to demonstrate the effectiveness of the proposed NMF-LCAG algorithm. Compared with some other related algorithms, the proposed NMF-LCAG algorithm can achieve at least 1% ~ 3% accuracy improvement in most cases.
Yugen Yi, Jianzhong Wang 0003, Wei Zhou 0003, Caixia Zheng, Jun Kong 0004, Shaojie Qiao
IEEE Trans. Circuits Syst. Video Technol.2
2019 Joint graph optimization and projection learning for dimensionality reduction
Yugen Yi, Jianzhong Wang 0003, Wei Zhou 0003, Jun Kong 0004, Yinghua Lu
Pattern Recognit.2
2018 Unsupervised feature selection by regularized matrix factorization
Miao Qi, Ting Wang 0015, Fucong Liu, Baoxue Zhang, Jianzhong Wang 0003, Yugen Yi
Neurocomputing5
2018 Adaptive multiple graph regularized semi-supervised extreme learning machine
Yugen Yi, Shaojie Qiao, Wei Zhou 0003, Caixia Zheng, Jianzhong Wang 0003
Soft Comput.6
2018 Ordinal preserving matrix factorization for unsupervised feature selection
Yugen Yi, Wei Zhou 0003, Guoliang Luo, Jianzhong Wang 0003, Caixia Zheng
Signal Process. Image Commun.5
2017 Excavation equipment classification based on improved MFCC features and ELM
Jiuwen Cao, Tuo Zhao, Jianzhong Wang 0003, Ruirong Wang, Yun Chen 0008
Neurocomputing3
2017 Locality constrained Graph Optimization for Dimensionality Reduction
Jianzhong Wang 0003, Caixia Zheng, Jun Kong 0004, Yugen Yi
Neurocomputing1
2017 Excavation Equipment Recognition Based on Novel Acoustic Statistical Features
abstract
Excavation equipment recognition attracts increasing attentions in recent years due to its significance in underground pipeline network protection and civil construction management. In this paper, a novel classification algorithm based on acoustics processing is proposed for four representative excavation equipments. New acoustic statistical features, namely, the short frame energy ratio, concentration of spectrum amplitude ratio, truncated energy range, and interval of pulse are first developed to characterize acoustic signals. Then, probability density distributions of these acoustic features are analyzed and a novel classifier is presented. Experiments on real recorded acoustics of the four excavation devices are conducted to demonstrate the effectiveness of the proposed algorithm. Comparisons with two popular machine learning methods, support vector machine and extreme learning machine, combined with the popular linear prediction cepstral coefficients are provided to show the generalization capability of our method. A real surveillance system using our algorithm is developed and installed in a metro construction site for real-time recognition performance validation.
Jiuwen Cao, Jianzhong Wang 0003, Ruirong Wang
IEEE Trans. Cybern.3
2015 Semi-supervised local ridge regression for local matching based face recognition
Yugen Yi, Chao Bi, Jianzhong Wang 0003, Jun Kong 0004
Neurocomputing4
2015 Label propagation based semi-supervised non-negative matrix factorization for feature extraction
Yugen Yi, Yanjiao Shi, Jianzhong Wang 0003, Jun Kong 0004
Neurocomputing4
2015 An improved locality sensitive discriminant analysis approach for feature extraction
Yugen Yi, Baoxue Zhang, Jun Kong 0004, Jianzhong Wang 0003
Multim. Tools Appl.4
2014 Structure Constrained Discriminative Non-negative Matrix Factorization for Feature Extraction
Lisi Wei, Yugen Yi, Jianzhong Wang 0003
ICIC (2)4
2014 A novel image retrieval method based on hybrid information descriptors
Ke Zhang 0023, Qinghe Feng, Jianzhong Wang 0003, Jun Kong 0004, Yinghua Lu
J. Vis. Commun. Image Represent.4
2013 Hardware design of a localization system for staff in high-risk manufacturing areas
abstract
In this paper, we propose a hardware design for an effective real time indoor localization system for staff working in high-risk manufacturing areas. Because of the special requirements of our system, the chirp spread spectrum (CSS) is the most suitable indoor localization technology. The details of the new localization system are described. The system involves several anchors, tags, and a gateway, all of which use the nanoLOC TRX transceiver (NA5TR1) RF chip (Nanotron Co., Germany), which is based on the CSS technology. To validate the effectiveness of our system, both ranging and localization tests were carried out. The difference between the ranging accuracy indoors and outdoors was small. The localization system can position indoor mobile staff precisely, enabling the establishment of an emergency rescue mechanism.
Ruirong Wang, Rong-Rong Ye, Cui-Fei Xu, Jianzhong Wang 0003, Anke Xue
J. Zhejiang Univ. Sci. C4
2013 Maximum weight and minimum redundancy: A novel framework for feature subset selection
Jianzhong Wang 0003, Lishan Wu, Jun Kong 0004, Yuxin Li 0001, Baoxue Zhang
Pattern Recognit.1
2011 Identification of Masses in Digital Mammogram Using an Optimal Set of Features
abstract
Recently, Digital mammogram has become one of the most effective techniques for early breast cancer detection. The aim of this study is to develop an automated system for digital mammogram analysis. In the proposed system, the regions of interest (ROIs) in the mammogram are firstly segmented by a topographic representation method called the is contour map. Subsequently, the textural, intensity and shape features are extracted from the ROIs. Then an optimal feature selection method (Correlation-based Feature Selection, CFS) is used to select some important features to classify the ROIs as either masses or non-masses. Finally, we use these selected features to train the cost-sensitive BP neural network. The experimental results show that the proposed method can produce better identification performance than some other algorithms.
Wenfeng Han, Jianzhong Wang 0003
TrustCom5
2011 A structure-preserved local matching approach for face recognition
Jianzhong Wang 0003, Zhiqiang Ma 0003, Baoxue Zhang, Miao Qi, Jun Kong 0004
Pattern Recognit. Lett.1
2010 Linear discriminant projection embedding based on patches alignment
Jianzhong Wang 0003, Baoxue Zhang, Miao Qi, Jun Kong 0004
Image Vis. Comput.1
2010 An adaptively weighted sub-pattern locality preserving projection for face recognition
Jianzhong Wang 0003, Baoxue Zhang, Shuyan Wang, Miao Qi, Jun Kong 0004
J. Netw. Comput. Appl.1
2008 Local-based fuzzy clustering for segmentation of MR brain images
abstract
Accurate segmentation of magnetic resonance images (MRI) corrupted by intensity inhomogeneity is a challenging problem and has received an enormous amount of attention lately. On the basis of the local image model, we propose a different segmentation method for MR brain images without estimation and correction for intensity heterogeneity. Firstly, we obtain clustering context based on the distributing disciplinarian in anatomy that gray matter (GM) is always between white matter (WM) and cerebrospinal fluid (CSF) in brain, which ensure the three tissues exist together in each one. Then the size of the context is optimized by a minimum entropy criterion. Finally, FCM algorithm is independently performed in each context to calculate the degree of membership of a pixel to each tissue class. The proposed methodology has been evaluated for simulated images and shown the better results.
Jianzhong Wang 0003, Lili Dou, Na Che, Di Liu 0004, Baoxue Zhang, Jun Kong 0004
BIBE1
2008 Graph theory based algorithm for magnetic resonance brain images segmentation
abstract
Image segmentation is often required as a preliminary and indispensable stage in the computer aided medical image process, particularly during the clinical analysis of magnetic resonance (MR) brain images. The segmentation of magnetic resonance image (MRI) is a challenging problem that has received an enormous amount of attention lately. In this paper, we propose a simple and effective segmentation method combining watershed algorithm and normalized cuts (CWNC) for MR brain images. An initial partitioning of the MRI into primitive regions is set by applying the watershed transform. The latter process uses a region similarity graph representation of the image regions. And then the graph is segmented by normalized cuts algorithm. The efficacy of the proposed algorithm is demonstrated by extensive segmentation experiments using both simulated and real MR images and by comparison with other published algorithms.
Jianzhong Wang 0003, Di Liu 0004, Lili Dou, Baoxue Zhang, Jun Kong 0004, Yinghua Lu
BIBE1
2007 A Modified Fuzzy Kohonen's Competitive Learning Algorithms Incorporating Local Information for MR Image Segmentation
abstract
A Modified FKCL (MFKCL) algorithm for automatic segmentation of MR brain images is proposed in this paper. This algorithm is an extension of traditional fuzzy Kohonen's competitive learning algorithm. In our method, a factor that can estimate the effect of the neighbor pixels to the central pixel is introduced into the objective function of the standard FKCL algorithm as the local information. The local information is applied to trail off the effect of noise to the result of MRI segmentation. Experiments with simulated MR data and real MR data show that our algorithm can resist not only the little, but also the heavy noise compared with standard FKCL segmentation and other reported methods.
Jun Kong 0004, Wenjing Lu, Jianzhong Wang 0003, Na Che, Yinghua Lu
BIBE3