Jianwu Li

dblp:09/5312 · DBLP profile ↗
← Back
50ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0002-8632-4334ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 9 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 MultiScale Knowledge Distillation
Zikang Yao, Jianwu Li
Mach. Learn.3
2026 Multimodal emotion recognition via unified granularity contrastive learning and similar negative discrimination
Jianwu Li
Pattern Recognit.3
2025 All-in-one weather removal via Multi-Depth Gated Transformer with gradient modulation
Jianwu Li
Pattern Recognit.2
2024 Label-Efficient Few-Shot Semantic Segmentation with Unsupervised Meta-Training
abstract
The goal of this paper is to alleviate the training cost for few-shot semantic segmentation (FSS) models. Despite that FSS in nature improves model generalization to new concepts using only a handful of test exemplars, it relies on strong supervision from a considerable amount of labeled training data for base classes. However, collecting pixel-level annotations is notoriously expensive and time-consuming, and small-scale training datasets convey low information density that limits test-time generalization. To resolve the issue, we take a pioneering step towards label-efficient training of FSS models from fully unlabeled training data, or additionally a few labeled samples to enhance the performance. This motivates an approach based on a novel unsupervised meta-training paradigm. In particular, the approach first distills pre-trained unsupervised pixel embedding into compact semantic clusters from which a massive number of pseudo meta-tasks is constructed. To mitigate the noise in the pseudo meta-tasks, we further advocate a robust Transformer-based FSS model with a novel prototype-based cross-attention design. Extensive experiments have been conducted on two standard benchmarks, i.e., PASCAL-5i and COCO-20i, and the results show that our method produces impressive performance without any annotations, and is comparable to fully supervised competitors even using only 20% of the annotations. Our code is available at: https://github.com/SSSKYue/UMTFSS.
Jianwu Li, Kaiyue Shi, Guosen Xie, Xiaofeng Liu 0006, Jian Zhang 0002, Tianfei Zhou
AAAI1
2023 Unified Mask Embedding and Correspondence Learning for Self-Supervised Video Segmentation
abstract
The objective of this paper is self-supervised learning of video object segmentation. We develop a unified framework which simultaneously models cross-frame dense correspondence for locally discriminative feature learning and embeds object-level context for target-mask decoding. As a result, it is able to directly learn to perform mask-guided sequential segmentation from unlabeled videos, in contrast to previous efforts usually relying on an oblique solution - cheaply “copying” labels according to pixel-wise correlations. Concretely, our algorithm alternates between i) clustering video pixels for creating pseudo segmentation labels ex nihilo; and ii) utilizing the pseudo labels to learn mask encoding and decoding for VOS. Unsupervised correspondence learning is further incorporated into this self-taught, mask embedding scheme, so as to ensure the generic nature of the learnt representation and avoid cluster degeneracy. Our algorithm sets state-of-the-arts on two standard benchmarks (i.e., DAVIS17and YouTube-VOS), narrowing the gap between self- and fully-supervised VOS, in terms of both performance and network architecture design.
Liulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li, Yi Yang 0001
CVPR4
2023 Volumetric memory network for interactive medical image segmentation
abstract
Despite recent progress of automatic medical image segmentation techniques, fully automatic results usually fail to meet clinically acceptable accuracy, thus typically require further refinement. To this end, we propose a novel Volumetric Memory Network, dubbed as VMN, to enable segmentation of 3D medical images in an interactive manner. Provided by user hints on an arbitrary slice, a 2D interaction network is firstly employed to produce an initial 2D segmentation for the chosen slice. Then, the VMN propagates the initial segmentation mask bidirectionally to all slices of the entire volume. Subsequent refinement based on additional user guidance on other slices can be incorporated in the same manner. To facilitate smooth human-in-the-loop segmentation, a quality assessment module is introduced to suggest the next slice for interaction based on the segmentation quality of each slice produced in the previous round. Our VMN demonstrates two distinctive features: First, the memory-augmented network design offers our model the ability to quickly encode past segmentation information, which will be retrieved later for the segmentation of other slices; Second, the quality assessment module enables the model to directly estimate the quality of each segmentation prediction, which allows for an active learning paradigm where users preferentially label the lowest-quality slice for multi-round refinement. The proposed network leads to a robust interactive segmentation engine, which can generalize well to various types of user annotations (e.g., scribble, bounding box, extreme clicking). Extensive experiments have been conducted on three public medical image segmentation datasets (i.e., MSD, KiTS19, CVC-ClinicDB), and the results clearly confirm the superiority of our approach in comparison with state-of-the-art segmentation models. The code is made publicly available at https://github.com/0liliulei/Mem3D.
Tianfei Zhou, Liulei Li, Gustav Bredell, Jianwu Li, Jan Unkelbach, Ender Konukoglu
Medical Image Anal.4
2022 Handwritten Mathematical Expression Recognition via Attention Aggregation Based Bi-directional Mutual Learning
abstract
Handwritten mathematical expression recognition aims to automatically generate LaTeX sequences from given images. Currently, attention-based encoder-decoder models are widely used in this task. They typically generate target sequences in a left-to-right (L2R) manner, leaving the right-to-left (R2L) contexts unexploited. In this paper, we propose an Attention aggregation based Bi-directional Mutual learning Network (ABM) which consists of one shared encoder and two parallel inverse decoders (L2R and R2L). The two decoders are enhanced via mutual distillation, which involves one-to-one knowledge transfer at each training step, making full use of the complementary information from two inverse directions. Moreover, in order to deal with mathematical symbols in diverse scales, an Attention Aggregation Module (AAM) is proposed to effectively integrate multi-scale coverage attentions. Notably, in the inference phase, given that the model already learns knowledge from two inverse directions, we only use the L2R branch for inference, keeping the original parameter size and inference speed. Extensive experiments demonstrate that our proposed approach achieves the recognition accuracy of 56.85 % on CROHME 2014, 52.92 % on CROHME 2016, and 53.96 % on CROHME 2019 without data augmentation and model ensembling, substantially outperforming the state-of-the-art methods. The source code is available in https://github.com/XH-B/ABM.
Xiaohang Bian, Xiaozhe Xin, Jianwu Li, Xuefeng Su
AAAI4
2022 Deep Hierarchical Semantic Segmentation
abstract
Humans are able to recognize structured relations in observation, allowing us to decompose complex scenes into simpler parts and abstract the visual world in multiple levels. However, such hierarchical reasoning ability of human perception remains largely unexplored in current literature of semantic segmentation. Existing work is often aware of flatten labels and predicts target classes exclusively for each pixel. In this paper, we instead address hierarchical semantic segmentation (HSS), which aims at structured, pixel-wise description of visual observation in terms of a class hierarchy. We devise HSSN, a general HSS framework that tackles two critical issues in this task: i) how to efficiently adapt existing hierarchy-agnostic segmentation networks to the HSS setting, and ii) how to leverage the hierarchy information to regularize HSS network learning. To address i), HSSN directly casts HSS as a pixel-wise multi-label classification task, only bringing minimal architecture change to current segmentation models. To solve ii), HSSN first explores inherent properties of the hierarchy as a training objective, which enforces segmentation predictions to obey the hierarchy structure. Further, with hierarchy-induced margin constraints, HSSNreshapes the pixel embedding space, so as to generate well-structured pixel representations and improve segmentation eventually. We conduct experiments on four semantic segmentation datasets (i.e., Mapillary Vistas 2.0, City-scapes, LIP, and PASCAL-Person-Part), with different class hierarchies, segmentation network architectures and backbones, showing the generalization and superiority of HSSN.
Liulei Li, Tianfei Zhou, Wenguan Wang, Jianwu Li, Yi Yang 0001
CVPR4
2022 Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence Learning
abstract
Our target is to learn visual correspondence from unlabeled videos. We develop Liir, a locality-aware inter-and intra-video reconstruction method that fills in three missing pieces, i.e., instance discrimination, location awareness, and spatial compactness, of self-supervised correspondence learning puzzle. First, instead of most existing efforts focusing on intra-video self-supervision only, we exploit cross-video affinities as extra negative samples within a unified, inter-and intra-video reconstruction scheme. This enables instance discriminative representation learning by contrasting desired intra-video pixel association against negative inter-video correspondence. Second, we merge position information into correspondence matching, and design a position shifting strategy to remove the side-effect of position encoding during inter-video affinity computation, making our Liir location-sensitive. Third, to make full use of the spatial continuity nature of video data, we impose a compactness-based constraint on correspondence matching, yielding more sparse and reliable solutions. The learned representation surpasses self-supervised state-of-the-arts on label propagation tasks including objects, semantic parts, and keypoints.
Liulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang 0006, Jianwu Li, Yi Yang 0001
CVPR5
2022 Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation
abstract
Learning semantic segmentation from weakly-labeled (e.g., image tags only) data is challenging since it is hard to infer dense object regions from sparse semantic tags. Despite being broadly studied, most current efforts directly learn from limited semantic annotations carried by individual image or image pairs, and struggle to obtain integral localization maps. Our work alleviates this from a novel perspective, by exploring rich semantic contexts synergistically among abundant weakly-labeled training data for network learning and inference. In particular, we propose regional semantic contrast and aggregation (RCA). RCA is equipped with a regional memory bank to store massive, diverse object patterns appearing in training data, which acts as strong support for exploration of dataset-level semantic structure. Particularly, we propose i) semantic contrast to drive network learning by contrasting massive categorical object regions, leading to a more holistic object pattern understanding, and ii) semantic aggregation to gather diverse relational contexts in the memory to enrich semantic repre-sentations. In this manner, RCA earns a strong capability of fine-grained semantic understanding, and eventually establishes new state-of-the-art results on two popular benchmarks, i.e., PASCAL VOC 2012 and COCO 2014.
Tianfei Zhou, Meijie Zhang, Fang Zhao 0006, Jianwu Li
CVPR4
2022 Multi-Granular Semantic Mining for Weakly Supervised Semantic Segmentation
abstract
This paper solves the problem of learning image semantic segmentation using image-level supervision. The task is promising in terms of reducing annotation efforts, yet extremely challenging due to the difficulty to directly associate high-level concepts with low-level appearance. While current efforts handle each concept independently, we take a broader perspective to harvest implicit, holistic structures of semantic concepts, which express valuable prior knowledge for accurate concept grounding. This raises multi-granular semantic mining, a new formalism allowing flexible specification of complex relations in the label space. In particular, we propose a heterogeneous graph neural network (Hgnn) to model the heterogeneity of multi-granular semantics within a set of input images. The Hgnn consists of two types of sub-graphs: 1) an external graph characterizes the relations across different images to mine inter-image contexts; and for each image, 2) an internal graph is constructed to mine inter-class semantic dependencies within each individual image. Through heterogeneous graph learning, our Hgnn is able to land a comprehensive understanding of object patterns, leading to more accurate semantic concept grounding. Extensive experimental results show that Hgnn outperforms the current state-of-the-art approaches on the popular PASCAL VOC 2012 and COCO 2014 benchmarks. Our code is available at: https://github.com/maeve07/HGNN.git.
Meijie Zhang, Jianwu Li, Tianfei Zhou
ACM Multimedia2
2022 Graph neural networks with global noise filtering for session-based recommendation
Lixia Feng, Yongqi Cai, Erling Wei, Jianwu Li
Neurocomputing4
2022 Group-Wise Learning for Weakly Supervised Semantic Segmentation
abstract
Acquiring sufficient ground-truth supervision to train deep visual models has been a bottleneck over the years due to the data-hungry nature of deep learning. This is exacerbated in some structured prediction tasks, such as semantic segmentation, which require pixel-level annotations. This work addresses weakly supervised semantic segmentation (WSSS), with the goal of bridging the gap between image-level annotations and pixel-level segmentation. To achieve this, we propose, for the first time, a novel group-wise learning framework for WSSS. The framework explicitly encodes semantic dependencies in a group of images to discover rich semantic context for estimating more reliable pseudo ground-truths, which are subsequently employed to train more effective segmentation models. In particular, we solve the group-wise learning within a graph neural network (GNN), wherein input images are represented as graph nodes, and the underlying relations between a pair of images are characterized by graph edges. We then formulate semantic mining as an iterative reasoning process which propagates the common semantics shared by a group of images to enrich node representations. Moreover, in order to prevent the model from paying excessive attention to common semantics, we further propose a graph dropout layer to encourage the graph model to capture more accurate and complete object responses. With the above efforts, our model lays the foundation for more sophisticated and flexible group-wise semantic mining. We conduct comprehensive experiments on the popular PASCAL VOC 2012 and COCO benchmarks, and our model yields state-of-the-art performance. In addition, our model shows promising performance in weakly supervised object localization (WSOL) on the CUB-200-2011 dataset, demonstrating strong generalizability. Our code is available at: https://github.com/Lixy1997/Group-WSSS.
Tianfei Zhou, Liulei Li, Xueyi Li 0006, Chun-Mei Feng 0001, Jianwu Li, Ling Shao 0001
IEEE Trans. Image Process.5
2021 Group-Wise Semantic Mining for Weakly Supervised Semantic Segmentation
abstract
Acquiring sufficient ground-truth supervision to train deep vi- sual models has been a bottleneck over the years due to the data-hungry nature of deep learning. This is exacerbated in some structured prediction tasks, such as semantic segmen- tation, which requires pixel-level annotations. This work ad- dresses weakly supervised semantic segmentation (WSSS), with the goal of bridging the gap between image-level anno- tations and pixel-level segmentation. We formulate WSSS as a novel group-wise learning task that explicitly models se- mantic dependencies in a group of images to estimate more reliable pseudo ground-truths, which can be used for training more accurate segmentation models. In particular, we devise a graph neural network (GNN) for group-wise semantic min- ing, wherein input images are represented as graph nodes, and the underlying relations between a pair of images are char- acterized by an efficient co-attention mechanism. Moreover, in order to prevent the model from paying excessive atten- tion to common semantics only, we further propose a graph dropout layer, encouraging the model to learn more accurate and complete object responses. The whole network is end-to- end trainable by iterative message passing, which propagates interaction cues over the images to progressively improve the performance. We conduct experiments on the popular PAS- CAL VOC 2012 and COCO benchmarks, and our model yields state-of-the-art performance. Our code is available at: https://github.com/Lixy1997/Group-WSSS.
Xueyi Li 0006, Tianfei Zhou, Jianwu Li, Yi Zhou 0007, Zhaoxiang Zhang 0001
AAAI3
2021 Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object Segmentation
abstract
This paper addresses the task of unsupervised video multi-object segmentation. Current approaches follow a two-stage paradigm: 1) detect object proposals using pre-trained Mask R-CNN, and 2) conduct generic feature matching for temporal association using re-identification techniques. However, the generic features, widely used in both stages, are not reliable for characterizing unseen objects, leading to poor generalization. To address this, we introduce a novel approach for more accurate and efficient spatio-temporal segmentation. In particular, to address instance discrimination, we propose to combine foreground region estimation and instance grouping together in one network, and additionally introduce temporal guidance for segmenting each frame, enabling more accurate object discovery. For temporal association, we complement current video object segmentation architectures with a discriminative appearance model, capable of capturing more fine-grained target-specific information. Given object proposals from the instance discrimination network, three essential strategies are adopted to achieve accurate segmentation: 1) target-specific tracking using a memory-augmented appearance model; 2) target-agnostic verification to trace possible tracklets for the proposal; 3) adaptive memory updating using the verified segments. We evaluate the proposed approach on DAVIS17and YouTube-VIS, and the results demonstrate that it outperforms state-of-the-art methods both in segmentation accuracy and inference speed.
Tianfei Zhou, Jianwu Li, Xueyi Li 0006, Ling Shao 0001
CVPR2
2021 Quality-Aware Memory Network for Interactive Volumetric Image Segmentation
Tianfei Zhou, Liulei Li, Gustav Bredell, Jianwu Li, Ender Konukoglu
MICCAI (2)4
2021 Conditional adversarial consistent identity autoencoder for cross-age face synthesis
Xiaohang Bian, Jianwu Li
Multim. Tools Appl.2
2020 Motion-Attentive Transition for Zero-Shot Video Object Segmentation
abstract
In this paper, we present a novel Motion-Attentive Transition Network (MATNet) for zero-shot video object segmentation, which provides a new way of leveraging motion information to reinforce spatio-temporal object representation. An asymmetric attention block, called Motion-Attentive Transition (MAT), is designed within a two-stream encoder, which transforms appearance features into motion-attentive representations at each convolutional stage. In this way, the encoder becomes deeply interleaved, allowing for closely hierarchical interactions between object motion and appearance. This is superior to the typical two-stream architecture, which treats motion and appearance separately in each stream and often suffers from overfitting to appearance information. Additionally, a bridge network is proposed to obtain a compact, discriminative and scale-sensitive representation for multi-level encoder features, which is further fed into a decoder to achieve segmentation results. Extensive experiments on three challenging public benchmarks (i.e., DAVIS-16, FBMS and Youtube-Objects) show that our model achieves compelling performance against the state-of-the-arts. Code is available at: https://github.com/tfzhou/MATNet.
Tianfei Zhou, Shunzhou Wang, Yi Zhou 0007, Yazhou Yao, Jianwu Li, Ling Shao 0001
AAAI5
2020 Occluded offline handwritten Chinese character inpainting via generative adversarial network and self-attention mechanism
Jianwu Li, Zheng Wang 0073
Neurocomputing2
2020 Occluded offline handwritten Chinese character recognition using deep convolutional generative adversarial network and improved GoogLeNet
Jianwu Li, Minhua Zhang
Neural Comput. Appl.1
2020 Disentangled representation learning and residual GAN for age-invariant face verification
Shuyang Zhao, Jianwu Li
Pattern Recognit.2
2020 MATNet: Motion-Attentive Transition Network for Zero-Shot Video Object Segmentation
abstract
In this paper, we present a novel end-to-end learning neural network, i.e., MATNet, for zero-shot video object segmentation (ZVOS). Motivated by the human visual attention behavior, MATNet leverages motion cues as a bottom-up signal to guide the perception of object appearance. To achieve this, an asymmetric attention block, named Motion-Attentive Transition (MAT), is proposed within a two-stream encoder network to firstly identify moving regions and then attend appearance learning to capture the full extent of objects. Putting MATs in different convolutional layers, our encoder becomes deeply interleaved, allowing for close hierarchical interactions between object apperance and motion. Such a biologically-inspired design is proven to be superb to conventional two-stream structures, which treat motion and appearance independently in separate streams and often suffer severe overfitting to object appearance. Moreover, we introduce a bridge network to modulate multi-scale spatiotemporal features into more compact, discriminative and scale-sensitive representations, which are subsequently fed into a boundary-aware decoder network to produce accurate segmentation with crisp boundaries. We perform extensive quantitative and qualitative experiments on four challenging public benchmarks, i.e., DAVIS16, DAVIS17, FBMS and YouTube-Objects. Results show that our method achieves compelling performance against current state-of-the-art ZVOS methods. To further demonstrate the generalization ability of our spatiotemporal learning framework, we extend MATNet to another relevant task: dynamic visual attention prediction (DVAP). The experiments on two popular datasets (i.e., Hollywood-2 and UCF-Sports) further verify the superiority of our model. Our implementations have been made publicly available at https://github.com/tfzhou/MATNet.
Tianfei Zhou, Jianwu Li, Shunzhou Wang, Ran Tao 0003, Jianbing Shen
IEEE Trans. Image Process.2
2019 DTDN: Dual-task De-raining Network
abstract
Removing rain streaks from rainy images is necessary for many tasks in computer vision, such as object detection and recognition. It needs to address two mutually exclusive objectives: removing rain streaks and reserving realistic details. Balancing them is critical for de-raining methods. We propose an end-to-end network, called dual-task de-raining network (DTDN), consisting of two sub-networks: generative adversarial network (GAN) and convolutional neural network (CNN), to remove rain streaks via coordinating the two mutually exclusive objectives self-adaptively. DTDN-GAN is mainly used to remove structural rain streaks, and DTDN-CNN is designed to recover details in original images. We also design a training algorithm to train these two sub-networks of DTDN alternatively, which share same weights but use different training sets. We further enrich two existing datasets to approximate the distribution of real rain streaks. Experimental results show that our method outperforms several recent state-of-the-art methods, based on both benchmark testing datasets and real rainy images.
Zheng Wang 0073, Jianwu Li
ACM Multimedia2
2019 Removing ring artifacts in CBCT images via generative adversarial networks with unidirectional relative total variation loss
Zheng Wang 0073, Jianwu Li, Mogendi Enoh
Neural Comput. Appl.2
2018 Removing Ring Artifacts in Cbct Images Via Generative Adversarial Network
abstract
Cone-beam computed tomography (CBCT) images often have some ring artifacts because of the inconsistent response of detector pixels. Removing ring artifacts in CBCT images without impairing the image quality is critical for the application of CBCT. In this paper, we explore this issue as an “adversarial problem” and propose a novel method to eliminate ring artifacts from CBCT images by using an image-to-image network based on Generative Adversarial Network (GAN). Through combining the generative adversarial loss and the proposed smooth loss, both of the generator and the discriminator can be trained to remove ring artifacts in CBCT images by means of image-to-image. Experimental results demonstrate that the proposed method is more effective on both simulated data and real-world CBCT images, compared with other algorithms.
Shuyang Zhao, Jianwu Li, Qirun Huo
ICASSP2
2018 Wavelet energy entropy and linear regression classifier for detecting abnormal breasts
Yi Chen 0023, Yin Zhang 0002, Huimin Lu 0001, Xian-Qing Chen, Jianwu Li, Shuihua Wang
Multim. Tools Appl.5
2018 Single image super-resolution via self-similarity and low-rank matrix recovery
Jianwu Li, Zhengchao Dong
Multim. Tools Appl.2
2018 Smart pathological brain detection by synthetic minority oversampling technique, extreme learning machine, and Jaya algorithm
Yudong Zhang 0001, Guihu Zhao, Junding Sun, Xiaosheng Wu, Zhiheng Wang 0001, Hongmin Liu 0001, Vishnuvarthanan Govindaraj, Tianming Zhan, Jianwu Li
Multim. Tools Appl.9
2017 Learning to generate video object segment proposals
abstract
This paper proposes a fully automatic pipeline to generate accurate object segment proposals in realistic videos. Our approach first detects generic object proposals for all video frames and then learns to rank them using a Convolutional Neural Networks (CNN) descriptor built on appearance and motion cues. The ambiguity of the proposal set can be reduced while the quality can be retained as highly as possible Next, high-scoring proposals are greedily tracked over the entire sequence into distinct tracklets. Observing that the proposal tracklet set at this stage is noisy and redundant, we perform a tracklet selection scheme to suppress the highly overlapped tracklets, and detect occlusions based on appearance and location information. Finally, we exploit holistic appearance cues for refinement of video segment proposals to obtain pixel-accurate segmentation. Our method is evaluated on two video segmentation datasets i.e. SegTrack v1 and FBMS-59 and achieves competitive results in comparison with other state-of-the-art methods.
Jianwu Li, Tianfei Zhou, Yao Lu 0001
ICME1
2017 Training Deep Autoencoder via VLC-Genetic Algorithm
Qazi Sami Ullah Khan, Jianwu Li, Shuyang Zhao
ICONIP (2)2
2017 Generating Low-Rank Textures via Generative Adversarial Network
Shuyang Zhao, Jianwu Li
ICONIP (3)2
2017 Texture Analysis Method Based on Fractional Fourier Entropy and Fitness-scaling Adaptive Genetic Algorithm for Detecting Left-sided and Right-sided Sensorineural Hearing Loss
abstract
To detect the sensorineural hearing loss (SNHL) from healthy people accurately, we used magnetic resonance imaging (MRI) to obtain the imaging data, and then proposed a new computer-aided diagnosis (CAD) system, on the basis of texture analysis method. In the first, we extracted 12-element feature from each brain image via fractional Fourier entropy (FRFE). Afterwards, multilayer perceptron (MLP) was employed as the classifier, which was trained by a novel fitness-scaling adaptive genetic algorithm (FSAGA). The statistical analysis over 49 subjects showed the overall accuracy of our method yielded 95.51%. Experimental results performed better than four state-of-the-art weight optimization methods, and this CAD system give significantly better performance than manual interpretation.
Shuihua Wang, Ming Yang 0011, Jianwu Li, Xueyan Wu, Hainan Wang, Bin Liu 0043, Zhengchao Dong, Yudong Zhang 0001
Fundam. Informaticae3
2016 Face hallucination scheme based on singular value content metric for K-NN selection and an iterative refining in a modified feature space
abstract
Numbers of neighbor embedding (NE) methods have been proposed, which use the image content metric based on the distance values such as Euclidean distance between the input image patch and the image patches in the training set to find the nearest neighbors. In contrast to these approaches we propose to use image content metric that uses the most effective singular values of the patch of interest. Singular value content metric give the effective and quantitative measure of the true image content and can search the most similar patches from the training set which possess the local similarity with the input patch. First we find the K most similar low resolution (LR) and corresponding high resolution (HR) patches by using the proposed image content metric. Secondly we project the K neighbor onto a modified feature space by employing easy partial least square estimation (EZ-PLS). In modified feature space we propose to explore the data structure of both LR and HR manifold and iteratively update Z nearest neighbors and reconstruction weights based on the results from previous iteration. The Rigorous experimentation with application to face hallucination demonstrate the effectiveness of the proposed method.
Javaria Ikram, Yao Lu 0001, Jianwu Li, Nie Hui
ICIP3
2016 Locality constraint neighbor embedding via KPCA and optimized reference patch for face hallucination
abstract
Given that the limitations of the manifold assumption that the low-resolution (LR) and high-resolution (HR) patch manifolds are locally isometric, the geometrical information of HR patch manifold, which is much more credible and discriminant than LR patch manifold, has been paid more attention to in the recent face super-resolution algorithms. In general, these algorithms first construct its initial HR patch using conventional face super-resolution methods and then update the K-nearest neighbors (K-NN) of the input patch as well as corresponding reconstruction weights based on the initial HR patch to generate the final HR patch. Whether or not we can effectively utilize the information of the HR manifold depends on the quality of the initial HR patch. In this paper, to capture the nonlinear similarity of face features, we apply kernel principal component analysis (KPCA) to the conventional face super-resolution method and achieve a better initial HR patch. Furthermore, we propose the concept “optimized reference patch” to deal with the variations in human facial features and find the best-matched neighbors of input patch. Experimental results show that the proposed method outperforms several state-of-the-art face super-resolution algorithms.
Qiang Tu, Jianwu Li, Javaria Ikram
ICIP2
2016 Removing Ring Artifacts in CBCT Images Using Smoothing Based on Relative Total Variation
Qirun Huo, Jianwu Li, Yao Lu 0001, Ziye Yan
ICONIP (1)2
2016 Face Hallucination Using Correlative Residue Compensation in a Modified Feature Space
Javaria Ikram, Yao Lu 0001, Jianwu Li, Nie Hui
ICONIP (2)3
2016 Fast Dual-Tree Wavelet Composite Splitting Algorithms for Compressed Sensing MRI
Jianwu Li, Jinpeng Zhou, Qiang Tu, Javaria Ikram, Zhengchao Dong
ICONIP (1)1
2016 Refining pre-image via error compensation for KPCA-based pattern de-noising
abstract
Finding pre-image is crucial for kernel principal component analysis (KPCA) based pattern de-noising. This paper proposes to learn the systematic error of some classical methods of pre-image finding, and to refine the obtained pre-image via error compensation. Experiments based on simulated data as well as real-world data demonstrate that the proposed approach can improve effectively the results from two classical pre-image methods: gradient decent and distance constraint.
Jianwu Li, Qiang Tu, Ziye Yan
ICPR1
2015 Locality constraint neighbour embedding via Reference Patch
abstract
Recently, face hallucination (FH) methods using position priors have gained popularity; however position priors might not always be the best due to the intrinsic rigidness of faces collected from uncontrollable environment. Therefore, we improve the search criteria for K-nearest neighbors (K-NN) to address the variations in human facial features. Meanwhile, the limitations of the manifold assumption are taken into consideration to refine the neighborhood of the low-resolution (LR) patch by using the information from the high-resolution (HR) patches. For each input patch, we search the local neighborhood of its corresponding position patch in each training image to find the best-matched neighbor “Reference Patches”. Reference patches and their HR counterparts are taken to construct LR and HR patch dictionaries. The proposed method is composed of two steps. For an input LR patch, first we construct its initial HR patch using conventional FH methods. Secondly, we search the initial HR patch's nearest neighbors in HR manifold to extract the discriminant locality constraints. Then the corresponding LR reference patches are taken as refined K-NN of the input patch. These refined reference patches better optimize the reconstruction weights, thus the performance is improved. Extensive experiments show that our method outperforms recent position patch schemes in reconstruction error and visual quality.
Javaria Ikram, Danfeng Wan, Jianwu Li
ICME4
2015 Learning kernel subspace for face recognition
Jianwu Li
Neurocomputing1
2013 Learning KPCA for Face Recognition
Jianwu Li
ICIC (3)2
2013 Credit Scoring Based on Kernel Matching Pursuit
Jianwu Li, Haizhou Wei, Chunyan Kong
ICIC (3)1
2013 Community detection in complex networks using extended compact genetic algorithm
Jianwu Li, YuLong Song
Soft Comput.1
2012 Relation regularized subspace recommending for related scientific articles
abstract
Recommending related scientific articles for a researcher is very important and useful in practice but also is full of challenges due to the latent complex semantic relations among scientific literatures. To deal with these challenges, this paper proposes a novel framework with link-missing data adaption, which casts the recommendation task to subspace embedding and similarity ranking problems. The relation regularized subspace in this framework is constructed via Relation Regularized Matrix Factorization (RRMF) for well modeling both content and link structure simultaneously. However, the link structure for an article is not always available in practical recommending. To solve this problem, we further propose two alternative approaches based on Latent Dirichlet Allocation (LDA) for link-missing articles recommendation as an extension of RRMF. Experiments on CiteSeer dataset demonstrate our method is more effective in comparison with some state-of-the-art approaches and is able to handle the link-missing case which the link-based methods never can fit.
Jianwu Li
CIKM2
2011 Super Resolution of Text Image by Pruning Outlier
Ziye Yan, Yao Lu 0001, Jianwu Li
ICONIP (3)3
2011 Efficient Semantic Kernel-Based Text Classification Using Matching Pursuit KFDA
Jianwu Li
ICONIP (2)2
2011 Constructing the Shortest ECOC for Fast Multi-classification
Jianwu Li, Haizhou Wei, Ziye Yan
KSEM1
2010 Constructing Sparse KFDA Using Pre-image Reconstruction
Jianwu Li
ICONIP (2)2
2010 Refining Kernel Matching Pursuit
Jianwu Li, Yao Lu 0001
ISNN (2)1
2008 Adapting radial basis function neural networks for one-class classification
abstract
One-class classification (OCC) is to describe one class of objects, called target objects, and discriminate them from all other possible patterns. In this paper, we propose to adapt radial basis function neural networks (RBFNNs) for OCC. First, target objects are mapped into a feature space by using neurons in the hidden layer of the RBFNNs. Then, we perform support vector domain description (SVDD) with linear kernel functions in the feature space to realize OCC. In addition, we also model, in the feature space, the closed sphere centered on the mean of target objects for OCC. Compared to the SVDD with nonlinear kernel functions, our methods can use flexible nonlinear mappings, which do not necessarily satisfy Mercerpsilas conditions. Moreover, we can also control the complexity of solutions easily by setting the number of neurons in the hidden layer of RBFNNs. Experimental results show that the classification accuracies of our methods can be close to, and even can reach those of the SVDD for most of results, but with typically much sparser models.
Jianwu Li, Zhanyong Mao, Yao Lu 0001
IJCNN1