VLDB 2026 Research / reviewers in the wild / expert
Feng Jiang 0001
dblp:75/1693-1
· DBLP profile ↗
123ranked-venue papers
14as first author
53since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 76 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 24 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 10 since 2021Systems, architecture and hardware · 8 · 4 first-author · 3 since 2021Computer networks · 8 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Perception balance with uncertainty-guided fusion and proposal-wise mixture-of-experts for robust multi-agent 3D object detection
Peizhou Ni, Wenkai Zhu, Feng Jiang 0001, Benwu Wang, Yue Hu 0009 |
Expert Syst. Appl. | 3 |
| 2026 | An unsupervised open-set recognition method for user-independent human activity recognition
Qi Zhang 0137, Baichun Wei, Haiqi Zhu, Feng Jiang 0001, Chunzhi Yi |
Neurocomputing | 4 |
| 2026 | Affective-aware fine-grained image quality assessment via multi-modal large language models
Chenyue Song, Xianzhu Liu, Haiqi Zhu, Yachun Mi, Kai Geng, Zhengyue Zhou, Feng Jiang 0001 |
Pattern Recognit. | 9 |
| 2025 | EGENN: An Efficient Graph-Enhanced Neural Network for Multivariate Time Series ForecastingabstractGraph Neural Network (GNN) has been widely applied in multivariate time series forecasting due to its excellent relationship modeling capabilities. However, current methods still face limitations in computational efficiency or time series expression capabilities. To address these issues, we propose an Efficient Graph-Enhanced Neural Network (EGENN), which consists of an adjacency matrix generator, GNN, and projection module. Firstly, EGENN designs a spectral similarity-based graph construction method and further enhances the expressive power of temporal features. Secondly, we introduce an inter-layer attention graph convolutional network, which adaptively aggregates information from different network depths to better capture complex patterns. Finally, a predictive projection strategy fusing wavelet convolutions and patch-wise transformation is proposed to produce compact parameterization and extended receptive fields. Experiments on five datasets from different domains show that our model achieves state-of-the-art prediction performance while maintaining low computational resource consumption. Haiqi Zhu, Chunzhi Yi, Baichun Wei, Feng Jiang 0001 |
ICASSP | 6 |
| 2025 | ALIVE: Asynchronous Lower Body Pose Estimation with Images, Visual-Inertial Odometry and ElectromyographyabstractHuman pose estimation (HPE) is a critical technology for multimedia applications such as virtual reality (VR) and other interactive systems, where efficient, accurate and cost-effective pose estimation is essential. However, high-frequency and precise pose estimation often requires advanced equipment like high-speed cameras or time-of-flight sensors and prohibitive amount of computational resources, which are expensive and hinder widespread adoption. Previous fusion based methods always rely on synchronized signal input which is throttle by low frequency sensors and vulnerable to signal drops. To address this challenge, we propose integrating low-cost electromyography (EMG) and visual-inertial odometry (VIO) data for lower-body HPE using a novel multi-modal neural network. Instead of relying on synchronized sensor inputs, we reformulate the fusion problem as an outdated-signal-guided HPE prediction task, achieving latency as low as 1.6 ms per prediction. We validate our approach on a dataset of 1,000 lower-body pose clips from 10 subjects, specifically curated for this task. Experimental results demonstrate that our method achieves accurate, high-frequency pose estimation. The implementation is publicly available at https://github.com/k9tming/ALIVE. Guoming Du, Zhen Ding, Xinrun Li, Wendi Peng, Feng Jiang 0001 |
ICME | 7 |
| 2025 | BPCLIP: A Bottom-up Image Quality Assessment from Distortion to Semantics Based on CLIPabstractImage Quality Assessment (IQA) aims to evaluate the perceptual quality of images based on human subjective perception. Existing methods generally combine multiscale features to achieve high performance, but most rely on straightforward linear fusion of these features, which may not adequately capture the impact of distortions on semantic content. To address this, we propose a bottom-up image quality assessment approach based on the Contrastive Language-Image Pre-training (CLIP, a recently proposed model that aligns images and text in a shared feature space), named BPCLIP, which progressively extracts the impact of low-level distortions on high-level semantics. Specifically, we utilize an encoder to extract multiscale features from the input image and introduce a bottom-up multiscale cross attention module designed to capture the relationships between shallow and deep features. In addition, by incorporating 40 image quality adjectives across six distinct dimensions, we enable the pre-trained CLIP text encoder to generate representations of the intrinsic quality of the image, thereby strengthening the connection between image quality perception and human language. Our method achieves superior results on most public Full-Reference (FR) and No-Reference (NR) IQA benchmarks, while demonstrating greater robustness. Chenyue Song, Wei Zhang 0192, Haiqi Zhu, Shaohui Liu, Feng Jiang 0001 |
ICME | 7 |
| 2025 | MS-IQA: A Multi-scale Feature Fusion Network for PET/CT Image Quality Assessment
Siqiao Li, Wei Zhang 0192, Chenyue Song, Feng Jiang 0001, Haiqi Zhu |
MICCAI (13) | 6 |
| 2025 | LVPNet: A Latent-Variable-Based Prediction-Driven End-to-End Framework for Lossless Compression of Medical Images
Chenyue Song, Wei Zhang 0192, Siqiao Li, Haiqi Zhu, Shengping Zhang, Shaohui Liu, Feng Jiang 0001 |
MICCAI (8) | 10 |
| 2025 | AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference UnderstandingabstractEmbodied reference understanding is crucial for intelligent agents to predict referents based on human intention through gesture signals and language descriptions. This paper introduces the Attention-Dynamic DINO, a novel framework designed to mitigate misinterpretations of pointing gestures across various interaction contexts. Our approach integrates visual and textual features to simultaneously predict the target object’s bounding box and the attention source in pointing gestures. Leveraging the distance-aware nature of nonverbal communication in visual perspective taking, we extend the virtual touch line mechanism and propose an attention-dynamic touch line to represent referring gesture based on interactive distances. The combination of this distance-aware approach and independent prediction of the attention source, enhances the alignment between objects and the gesture represented line. Extensive experiments on the YouRefit dataset demonstrate the efficacy of our gesture information understanding method in significantly improving task performance. Our model achieves 76.3% accuracy at the 0.25 IoU threshold and, notably, surpasses human performance at the 0.75 IoU threshold, marking a first in this domain. Comparative experiments with distance-unaware understanding methods from previous research further validate the superiority of the Attention-Dynamic Touch Line across diverse contexts. Hao Guo 0015, Baichun Wei, Jianfei Zhu, Chunzhi Yi, Feng Jiang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Think Locally and Act Globally: A Frequency-Spatial Fusion Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) remains challenging due to the extremely low signal-to-noise ratio (SNR). Existing methods struggle to balance accuracy and speed, especially under limited computational resources. To address these issues, we propose the frequency-spatial contextual fusion network (FSCFNet) based on You Only Look Once (YOLO) v10n architecture. Particularly, the novel frequency-spatial convolution (FSConv) is designed that decomposes input features via Haar Wavelet Transform. High-frequency cues focus on local details to highlight small targets, while low-frequency cues provide global information to complement spatial features. Subsequently, the asymmetric cross-domain attention (ACA) is developed to enhance the local central feature extraction, which reflects the typical spatial Gaussian pattern of small targets. Furthermore, we introduce the customized multi-scale receptive contextual Block (MRCB) to capture the long-range information by leveraging diverse dilated convolutions. In addition, the Wasserstein Distance Loss (WDL) is utilized to improve bounding box quality. Extensive experiments on three public datasets including IRSTD-1k, NUDT-SIRST, and NUAA-SIRST confirm the effectiveness of FSCFNet. Notably, FSCFNet surpasses the baseline by 4.7% in precision, 3.3% in recall, and 3.9% in AP@50 on IRSTD-1k, with only a 3.6% increase in parameters. FSCFNet provides a robust solution for real-time infrared surveillance systems under resource-constrained environments. More comparisons are shown in Fig. 1. Weijie Xu, Zhenglong Ding, Zhiqing Cui, Yifan Hu 0015, Feng Jiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Image Compressive Sensing With Scale-Variable Adaptive Sampling and Hybrid-Attention Transformer ReconstructionabstractRecently, a large number of image compressive sensing (CS) methods with deep unfolding networks (DUNs) have been proposed. However, existing methods either use fixed-scale blocks for sampling that leads to limited insights into the image content or employ a plain convolutional neural network (CNN) in each iteration that weakens the perception of broader contextual prior. In this paper, we propose a novel DUN (dubbed SVASNet) for image compressive sensing, which achieves scale-variable adaptive sampling and hybrid-attention Transformer reconstruction with a single model. Specifically, for scale-variable sampling, a sampling matrix-based calculator is first employed to evaluate the reconstruction distortion, which only requires measurements without access to the ground truth image. Then, a Block Scale Aggregation (BSA) strategy is presented to compute the reconstruction distortion under block divisions at different scales and select the optimal division scale for sampling. To realize hybrid-attention reconstruction, a dual Cross Attention (CA) submodule in the gradient descent step and a Spatial Attention (SA) submodule in the proximal mapping step are developed. The CA submodule introduces inter-phase inertial forces in the gradient descent, which improves the memory effect between adjacent iterations. The SA submodule integrates local and global prior representations of CNN and Transformer, and explores local and global affinities between dense feature representations. Extensive experimental results show that the proposed SVASNet achieves significant improvements over the state-of-the-art methods. Debin Zhao, Weisi Lin, Shaohui Liu, Feng Jiang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Progressively Learning to Reach Remote Goals by Continuously Updating Boundary GoalsabstractTraining an effective policy on complex goal-reaching tasks with sparse rewards is an open challenge. It is more difficult for the task of reaching remote goals (RRG), as the unavailability of the original rewards and large Wasserstein distance between the distributions of desired goals and initial states make existing methods for common goal-reaching tasks inefficient or even completely ineffective. In this article, we propose progressively learning to reach remote goals by continuously updating boundary goals (PLUB), which solves RRG tasks by reducing the Wasserstein distance between the distributions of boundary goals and desired goals. Specifically, the concept of boundary goal is introduced, which is the set of the closest achieved goals for each desired goal. In addition, to reduce the computational complexity caused by the Wasserstein distance, the closest moving distance is introduced, which is its upper bound, and also the expectation of the distance between the desired goal and the closest boundary goal. By selecting the appropriate intermediate goal from all boundary goals and continuously updating boundary goals, both the closest moving distance and the Wasserstein distance can be reduced. As a result, RRG tasks degenerate into common goal-reaching tasks that can be efficiently solved by a combination of hindsight relabeling and the learning from demonstrations (LfD) method. Extensive experiments on several robotic manipulation tasks demonstrate that PLUB can bring substantial improvements over the existing methods. Mengxuan Shao, Haiqi Zhu, Debin Zhao, Feng Jiang 0001, Shaohui Liu, Wei Zhang 0192 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | ADTAH: Neuron 3D Reconstruction Via Adaptive Distance Transformation and Adaptive Hessian MatrixabstractThree-dimensional (3D) reconstruction of neurons is a critical and evolving area in neuroscience, addressing the substantial challenges presented by weak signals, high noise levels, and heterogeneous signal distribution in neuronal optical images. Previous methodologies predominantly focused on reconstructing neuronal fibers but faced significant limitations in integrating both neuronal fibers and somas, making it difficult to handle large-scale neuronal image reconstruction. Furthermore, conventional techniques often employ fixed thresholding to eliminate background noise, inadvertently leading to the loss of valuable low-intensity neuronal signals, which are crucial for comprehensive neuronal analysis. In response to these limitations, we propose ADTAH, an innovative neuron image reconstruction method that leverages the adaptive distance transform combined with the Hessian matrix for robust and precise 3D reconstruction of neurons. ADTAH consists of two branches: nerve fiber reconstruction and soma reconstruction. The nerve fiber reconstruction branch begins with adaptive threshold distance transform, dynamically adjusting the Hessian matrix window size to effectively capture nerve fibers of varying thicknesses, thereby optimizing the extraction of tubular structures. The soma reconstruction branch employs high-threshold distance transform to accurately identify and fill somas, ensuring comprehensive coverage of both fibers and somas. Our experiments on publicly available 3D neuron image datasets demonstrate that ADTAH outperforms existing techniques, providing superior differentiation between neuronal and background signals, enhanced computational efficiency, and greater robustness. Our approach has demonstrated state-of-the-art performance on the publicly available 3D Neuron image dataset Big Neuron and an fMost dataset made available by the Allen Institute for Science. Chenyue Song, Wei Zhang 0192, Feng Jiang 0001, Deqin Zheng, Ruichen Gao, Haiqi Zhu |
BIBM | 3 |
| 2024 | SC-HVPPNet: Spatial and Channel Hybrid-Attention Video Post-Processing Network with CNN and TransformerabstractConvolutional Neural Network (CNN) and Transformer have attracted much attention recently for video post-processing (VPP). However, the interaction between CNN and Transformer in existing VPP methods is not fully explored, leading to inefficient communication between the local and global extracted features. In this paper, we explore the interaction between CNN and Transformer in the task of VPP, and propose a novel Spatial and Channel Hybrid-Attention Video Post-Processing Network (SC-HVPPNet), which can cooperatively exploit the image priors in both spatial and channel domains. Specifically, in the spatial domain, a novel spatial attention fusion module is designed, in which two attention weights are generated to fuse the local and global representations collaboratively. In the channel domain, a novel channel attention fusion module is developed, which can blend the deep representations at the channel dimension dynamically. Extensive experiments show that SC-HVPPNet notably boosts video restoration quality, with average bitrate savings of 5.29%, 12.42%, and 13.09% for Y, U, and V components in the VTM-11.0-NNVC RA configuration. Wenxue Cui, Shaohui Liu, Feng Jiang 0001 |
ICME | 4 |
| 2024 | Deep Ad-hoc Sub-Team Partition Learning for Multi-Agent Air Combat CooperationabstractIn the future, unmanned autonomous air combat will encounter large-scale confrontation scenarios, where agents must consider complex time-varying relationships among aircraft when making decisions. Previous works have already introduced Multi-Agent Reinforcement Learning (MARL) into air combat and succeeded in surpassing the human expert level. However, they mainly focus on small-scale air combat with low relationship complexity, e.g., 1-vs-1 or 2-vs-2. As more agents join the confrontation, existing algorithms tend to suffer significant performance degradation due to the increase in problem dimensions. In view of this, this paper proposes Deep Ad-hoc Sub-Team Partition Learning(DASPL) to address large-scale air combat problems. DASPL models multi-agent air combat as a graph to handle the complex relations and introduces an automatic partitioning mechanism to generate dynamic sub-teams, which converts the existing large-scale multi-agent air combat cooperation problem into multiple small-scale equivalence problems. Additionally, DASPL incorporates an efficient message passing method among the participating sub-teams. Songyuan Fan, Haiyin Piao, Feng Jiang 0001, Roushu Yang |
IROS | 4 |
| 2024 | S2-CSNet: Scale-Aware Scalable Sampling Network for Image Compressive SensingabstractDeep network-based image Compressive Sensing (CS) has attracted much attention in recent years. However, there still exist the following two issues: 1) Existing methods typically use fixed-scale sampling, which leads to limited insights into the image content. 2) Most pre-trained models can only handle fixed sampling rates and fixed block scales, which restricts the scalability of the model. In this paper, we propose a novel scale-aware scalable CS network (dubbed S2-CSNet), which achieves scale-aware adaptive sampling, fine granular scalability and high-quality reconstruction with one single model. Specifically, to enhance the scalability of the model, a structural sampling matrix with a predefined order is first designed, which is a universal sampling matrix that can sample multi-scale image blocks with arbitrary sampling rates. Then, based on the universal sampling matrix, a distortion-guided scale-aware scheme is presented to achieve scale-variable adaptive sampling, which predicts the reconstruction distortion at different sampling scales from the measurements and select the optimal division scale for sampling. Furthermore, a multi-scale hierarchical sub-network under a well-defined compact framework is put forward to reconstruct the image. In the multi-scale feature domain of the sub-network, a dual spatial attention is developed to explore the local and global affinities between dense feature representations for deep fusion. Extensive experiments manifest that the proposed S2-CSNet outperforms existing state-of-the-art CS methods. Haiqi Zhu, Shuya Yan, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
ACM Multimedia | 5 |
| 2024 | HLAIImaster: a deep learning method with adaptive domain knowledge predicts HLA II neoepitope immunogenic responsesabstractWhile significant strides have been made in predicting neoepitopes that trigger autologous CD4+ T cell responses, accurately identifying the antigen presentation by human leukocyte antigen (HLA) class II molecules remains a challenge. This identification is critical for developing vaccines and cancer immunotherapies. Current prediction methods are limited, primarily due to a lack of high-quality training epitope datasets and algorithmic constraints. To predict the exogenous HLA class II-restricted peptides across most of the human population, we utilized the mass spectrometry data to profile >223 000 eluted ligands over HLA-DR, -DQ, and -DP alleles. Here, by integrating these data with peptide processing and gene expression, we introduce HLAIImaster, an attention-based deep learning framework with adaptive domain knowledge for predicting neoepitope immunogenicity. Leveraging diverse biological characteristics and our enhanced deep learning framework, HLAIImaster is significantly improved against existing tools in terms of positive predictive value across various neoantigen studies. Robust domain knowledge learning accurately identifies neoepitope immunogenicity, bridging the gap between neoantigen biology and the clinical setting and paving the way for future neoantigen-based therapies to provide greater clinical benefit. In summary, we present a comprehensive exploitation of the immunogenic neoepitope repertoire of cancers, facilitating the effective development of "just-in-time" personalized vaccines. Qiang Yang 0015, Weihe Dong, Xiaokun Li, Kuanquan Wang, Suyu Dong, Xianyu Zhang 0004, Tiansong Yang, Feng Jiang 0001, Bin Zhang 0042, Gongning Luo, Xin Gao 0001, Guohua Wang 0001 |
Briefings Bioinform. | 9 |
| 2024 | ActiveSelfHAR: Incorporating Self-Training Into Active Learning to Improve Cross-Subject Human Activity RecognitionabstractDeep learning (DL)-based human activity recognition (HAR) methods have shown promise in the applications of health Internet of Things (IoT) and wireless body sensor networks (BSNs). However, adapting these methods to new users in real-world scenarios is challenging due to the cross-subject issue. To solve this issue, we propose ActiveSelfHAR, a framework that combines active learning’s benefit of sparsely acquiring informative samples with actual labels and self-training’s benefit of effectively utilizing unlabeled data to adapt the HAR model to the target domain, i.e., the new users. ActiveSelfHAR consists of several key steps. First, we utilize the model from the source domain to select and label the domain invariant samples, forming a self-training set. Second, we leverage the distribution information of the self-training set to identify and annotate samples located around the class boundaries, forming a core set. Third, we augment the core set by considering the spatiotemporal relationships among the samples in the nonself-training set. Finally, we combine the self-training set and augmented core set to construct a diverse training set in the target domain and fine-tune the HAR model. Through leave-one-subject-out validation on three IMU-based data sets and one EMG-based data set, our method achieves mean HAR accuracies of 95.20%, 82.06%, 89.52%, and 92.82%, respectively. Our method demonstrates similar HAR accuracies to the upper bound, i.e., fine-tuning framework with approximately 1% labeled data of the target data set, while significantly improving data efficiency and time cost. Our work highlights the potential of implementing user-independent HAR methods into health IoT and BSN. Baichun Wei, Chunzhi Yi, Qi Zhang 0137, Haiqi Zhu, Jianfei Zhu, Feng Jiang 0001 |
IEEE Internet Things J. | 6 |
| 2024 | An Interpretable Multivariate Time-Series Anomaly Detection Method in Cyber-Physical Systems Based on Adaptive MaskabstractThe high complexity and wide applications of Cyber-Physical Systems (CPSs) pose a large requirement on both accuracy and interpretability of the time-series anomaly detection algorithms. While a large number of deep learning algorithms have achieved excellent accuracy, the interpretability is often limited, especially when considering retaining correlations in multivariate time-series. In this paper, we propose a novel multivariate time-series anomaly detection method based on adaptive masking mechanism to improve both accuracy and interpretability, which contains a specially designed series saliency module. For more intuitive and interpretable results, a learnable adaptive mask is introduced in the series saliency module, which can disclose the influence on anomalies in both feature and temporal dimensions. The original time-series and their versions with adaptive perturbations added are then mixed via the mask forming an adaptive data augmentation method to improve the accuracy of anomaly detection. Furthermore, the anomaly detection module is model-agnostic, whether based on forecasting or reconstruction. The optimization of the training objectives will lead to more accurate and interpretable detection results. With four real-world datasets, we demonstrate that the adaptive mask can provide more accurate anomaly detection results with meaningful interpretations in the form of a mask matrix. Haiqi Zhu, Chunzhi Yi, Seungmin Rho, Shaohui Liu, Feng Jiang 0001 |
IEEE Internet Things J. | 5 |
| 2024 | Reducing Data Transmission Efficiency in Wireless Capsule Endoscopy through DL-CEndo Framework: Reconstructing Lossy Low-Resolution Luma Images and Improving Summarization
Abderrahmane Salmi, Wei Zhang 0192, Feng Jiang 0001 |
Mob. Networks Appl. | 3 |
| 2024 | Deep Learning and Dempster-Shafer Theory Based Insider Threat Detection
Zhihong Tian 0001, Wei Shi 0001, Zhiyuan Tan 0001, Jing Qiu 0002, Yanbin Sun, Feng Jiang 0001, Yan Liu 0014 |
Mob. Networks Appl. | 6 |
| 2024 | Stereo Image Restoration via Attention-Guided Correspondence LearningabstractAlthough stereo image restoration has been extensively studied, most existing work focuses on restoring stereo images with limited horizontal parallax due to the binocular symmetry constraint. Stereo images with unlimited parallax (e.g., large ranges and asymmetrical types) are more challenging in real-world applications and have rarely been explored so far. To restore high-quality stereo images with unlimited parallax, this paper proposes an attention-guided correspondence learning method, which learns both self- and cross-views feature correspondence guided by parallax and omnidirectional attention. To learn cross-view feature correspondence, a Selective Parallax Attention Module (SPAM) is proposed to interact with cross-view features under the guidance of parallax attention that adaptively selects receptive fields for different parallax ranges. Furthermore, to handle asymmetrical parallax, we propose a Non-local Omnidirectional Attention Module (NOAM) to learn the non-local correlation of both self- and cross-view contexts, which guides the aggregation of global contextual features. Finally, we propose an Attention-guided Correspondence Learning Restoration Network (ACLRNet) upon SPAMs and NOAMs to restore stereo images by associating the features of two views based on the learned correspondence. Extensive experiments on five benchmark datasets demonstrate the effectiveness and generalization of the proposed method on three stereo image restoration tasks including super-resolution, denoising, and compression artifact reduction. Shengping Zhang, Wei Yu 0004, Feng Jiang 0001, Liqiang Nie, Hongxun Yao, Qingming Huang, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Magnetoencephalography Decoding Transfer Approach: From Deep Learning Models to Intrinsically Interpretable ModelsabstractWhen decoding neuroelectrophysiological signals represented by Magnetoencephalography (MEG), deep learning models generally achieve high predictive performance but lack the ability to interpret their predicted results. This limitation prevents them from meeting the essential requirements of reliability and ethical-legal considerations in practical applications. In contrast, intrinsically interpretable models, such as decision trees, possess self-evident interpretability while typically sacrificing accuracy. To effectively combine the respective advantages of both deep learning and intrinsically interpretable models, an MEG transfer approach through feature attribution-based knowledge distillation is pioneered, which transforms deep models (teacher) into highly accurate intrinsically interpretable models (student). The resulting models provide not only intrinsic interpretability but also high predictive performance, besides serving as an excellent approximate proxy to understand the inner workings of deep models. In the proposed approach, post-hoc feature knowledge derived from post-hoc interpretable algorithms, specifically feature attribution maps, is introduced into knowledge distillation for the first time. By guiding intrinsically interpretable models to assimilate this knowledge, the transfer of MEG decoding information from deep models to intrinsically interpretable models is implemented. Experimental results demonstrate that the proposed approach outperforms the benchmark knowledge distillation algorithms. This approach successfully improves the prediction accuracy of Soft Decision Tree by a maximum of 8.28%, reaching almost equivalent or even superior performance to deep teacher models. Furthermore, the model-agnostic nature of this approach offers broad application potential. Yongdong Fan, Qiong Li 0001, Haokun Mao, Feng Jiang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | An Adaptively Weighted Averaging Method for Regional Time Series Extraction of fMRI-Based Brain DecodingabstractBrain decoding that classifies cognitive states using the functional fluctuations of the brain can provide insightful information for understanding the brain mechanisms of cognitive functions. Among the common procedures of decoding the brain cognitive states with functional magnetic resonance imaging (fMRI), extracting the time series of each brain region after brain parcellation traditionally averages across the voxels within a brain region. This neglects the spatial information among the voxels and the requirement of extracting information for the downstream tasks. In this study, we propose to use a fully connected neural network that is jointly trained with the brain decoder to perform an adaptively weighted average across the voxels within each brain region. We perform extensive evaluations by cognitive state decoding, manifold learning, and interpretability analysis on the Human Connectome Project (HCP) dataset. The performance comparison of the cognitive state decoding presents an accuracy increase of up to 5% and stable accuracy improvement under different time window sizes, resampling sizes, and training data sizes. The results of manifold learning show that our method presents a considerable separability among cognitive states and basically excludes subject-specific information. The interpretability analysis shows that our method can identify reasonable brain regions corresponding to each cognitive state. Our study would aid the improvement of the basic pipeline of fMRI processing. Jianfei Zhu, Baichun Wei, Jiaru Tian, Feng Jiang 0001, Chunzhi Yi |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Rate-Adaptive Neural Network for Image Compressive SensingabstractDeep learning-based image compressive sensing (CS) methods have achieved great success in the past few years. However, most of them are content-independent, with a spatially uniform sampling rate allocation for the entire image. Such practises may potentially degrade the performance of image CS with block-based sampling, since the content of different blocks in an image is different. In this article, we propose a novel rate-adaptive image CS neural network (dubbed RACSNet) to achieve adaptive sampling rate allocation based on the content characteristics of the image with a single model. Specifically, a measurement domain-based reconstruction distortion is first used to guide the sampling rate allocation for different blocks in an image without access to the ground truth image. Then, a step-wise training strategy is designed to train a reusable sampling matrix, which is capable of sampling image blocks to generate the compressed measurements under arbitrary sampling rates. Subsequently, a pyramid-shaped initial reconstruction sub-network and a hierarchical deep reconstruction sub-network that fuse the measurement information of different scales are put forward to reconstruct image blocks from the compressed measurements. Finally, a reconstruction distortion map and an improved loss function are developed to eliminate the blocking artifacts and further enhance the CS reconstruction. Experimental results on both objective metrics and subjective visual qualities show that the proposed RACSNet achieves significant improvements over the state-of-the-art methods. Shengping Zhang, Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Multim. | 5 |
| 2023 | Graph Convolutional Network with Neural Inductive Matrix Completion for Predicting Disease-Related LncRNA GenesabstractNumerous researches emphasized that long non-coding RNA (lncRNA) plays a vital factor in various biological processes, and its mismatched expression and dysfunction are tightly linked with the occurrence of human diseases. Thus, computational models were designed to identify lncRNA-disease interactions by merging heterogeneous biological data. However, most of them neglected the intrinsic structure of multi-source information, which limits the performance for potential lncRNA-disease association prediction. Here, GCN-NIMC is introduced to alleviate the dilemma for disease-associated lncRNA genes identification based on the graph convolutional network with neural inductive matrix. This method builds a feature matrix with multi-source heterogeneous data and then learn the various information contained in the feature matrix for the sake of acquiring better feature expressions of the lncRNA-disease interactions. Experimental results on 10-repeated 5-fold cross-validation demonstrated that our proposed GCN-NIMC is superior to existing cutting-edge methods for identifying disease-related lncRNA genes. Furthermore, case studies confirmed our computational method as a practical tool with clinical benefits to develop the therapeutic schedule at lncRNA-level. Qiang Yang 0015, Suyu Dong, Weihe Dong, Xiaokun Li, Pengzhong Sun, Feng Jiang 0001, Xianyu Zhang 0004, Gongning Luo |
BIBM | 8 |
| 2023 | EI2SR: Learning an Enhanced Intra-Instance Semantic Relationship for Arbitrary-Shaped Scene Text DetectionabstractText detection in natural scenarios, has made significant progress with the deep learning architecture. Towards arbitrary-shaped text detection, fracture detection is the major concern due to the lack of semantic relationship within an instance in existing methods. To circumvent this dilemma, we propose a novel network to learn an Enhanced Intra-Instance Semantic Relationship (EI2SR) which consists of Text-Specific Attention Mechanism (TAM) and Border Attraction Grouping (BAG). The former models the rich semantic information between different coarse-grained text regions to guide the fine-grained learning of corresponding text representations. The latter enhances the border-center semantic correlation by establishing high-dimension embedding space to attract and group the border at both ends to their corresponding center. Extensive experimental results show that the proposed EI2SR achieves state-of-the-art or competitive performance on existing benchmarks. Shaohui Liu, Yu Zhou 0015, Feng Jiang 0001 |
ICASSP | 5 |
| 2023 | Aprogressive Image Dehazing Framework with inter and Intra Contrastive LearningabstractImage dehazing, aims to estimate latent haze-free images from hazy images, suffering from a lot of lost information. Existing contrastive learning methods tend to utilize hazefree images as positive samples without consideration of negative samples. Even if negative samples are employed, the connection between patches within an image is always ignored. In addition, it is hard to train end-to-end dehazing networks due to the enormous gap between hazy images and corresponding clear images. In this paper, we propose a novel progressive image dehazing framework with inter and intra contrastive learning to solve the above problems. Specifically, the Inter and Intra Contrastive Learning (IICL) is proposed, in which the brightest and darkest patches within the same image are considered for contrastive learning. Furthermore, a progressive image dehazing framework consisting of an efficient Pre-restore Module (PRM) and an Alternative Restored Module (ARM) is proposed to facilitate the end-to-end model training. It is noted that our framework can be a complement to existing image dehazing methods. Extensive experiments on the dehazing benchmark demonstrate that our framework benefits various dehazing models which surpass previous state-of-the-art image dehazing methods. Shaohui Liu, Feng Jiang 0001 |
ICASSP | 4 |
| 2023 | Hierarchical Interactive Reconstruction Network for Video Compressive SensingabstractDeep network-based image and video Compressive Sensing (CS) has attracted increasing attentions in recent years. However, in the existing deep network-based CS methods, a simple stacked convolutional network is usually adopted, which not only weakens the perception of rich contextual prior knowledge, but also limits the exploration of the correlations between temporal video frames. In this paper, we propose a novel Hierarchical InTeractive Video CS Reconstruction Network(HIT-VCSNet), which can cooperatively exploit the deep priors in both spatial and temporal domains to improve the reconstruction quality. Specifically, in the spatial domain, a novel hierarchical structure is designed, which can hierarchically extract deep features from keyframes and non-keyframes. In the temporal domain, a novel hierarchical interaction mechanism is proposed, which can cooperatively learn the correlations among different frames in the multi-scale space. Extensive experiments manifest that the proposed HIT-VCSNet outperforms the existing state-of-the-art video and image CS methods in a large margin. Wenxue Cui, Feng Jiang 0001 |
ICASSP | 4 |
| 2023 | VVA: Video Values Analysis
Yachun Mi, Shaohui Liu, Feng Jiang 0001 |
PRCV (7) | 5 |
| 2023 | Scale-Aware Frequency Attention network for super-resolution
Wei Yu 0004, Zonglin Li 0004, Qinglin Liu, Feng Jiang 0001, Changyong Guo, Shengping Zhang |
Neurocomputing | 4 |
| 2023 | Mordo: Silent Command Recognition Through Lightweight Around-Ear BiosensorsabstractThe prevalence of smart devices encourages increasing requirements of wearable human–computer interactions. To improve user acceptance, such interactions require easy-to-manipulate and unobtrusive characteristics. In this article, we, for the first time, propose to recognize silent commands through a lightweight and around-ear biosensing system Mordo that can be easily integrated with earphones, manipulate smart devices, and minimize social awkwardness. In particular, we first determine the empirical principles of constructing commands and experimentally screen the commands based on the around-ear configuration. Second, we select the optimal around-ear sensor configuration according to the single-channel signal-to-noise ratios (SNRs) and classification accuracies. Third, we propose a multistream CNN-LSTM network to learn the spatiotemporal mapping between the around-ear signals and commands. Finally, extensive experiments have been conducted to evaluate the feasibility and stability. The results indicate an averaged accuracy of 89.66% that outperforms other algorithms of similar tasks. The stability tests show that our system presents sufficient stability under command deformations and head motions. We demonstrate the necessity of collecting such scale of data by gradually reducing training data size. We also validate the generalization ability of our method toward other sensing parameters by reducing the spatial and temporal resolutions. The proof-of-concept design will aim the further development of the commercial products for silent command recognition. Chunzhi Yi, Baichun Wei, Jianfei Zhu, Seungmin Rho, Zhiyuan Chen 0007, Feng Jiang 0001 |
IEEE Internet Things J. | 6 |
| 2023 | Image Compressed Sensing Using Non-Local Neural NetworkabstractDeep network-based image Compressed Sensing (CS) has attracted much attention in recent years. However, the existing deep network-based CS schemes either reconstruct the target image in a block-by-block manner that leads to serious block artifacts or train the deep network as a black box that brings about limited insights of image prior knowledge. In this paper, a novel image CS framework using non-local neural network (NL-CSNet) is proposed, which utilizes the non-local self-similarity priors with deep network to improve the reconstruction quality. In the proposed NL-CSNet, two non-local subnetworks are constructed for utilizing the non-local self-similarity priors in the measurement domain and the multi-scale feature domain respectively. Specifically, in the subnetwork of measurement domain, the long-distance dependencies between the measurements of different image blocks are established for better initial reconstruction. Analogically, in the subnetwork of multi-scale feature domain, the affinities between the dense feature representations are explored in the multi-scale space for deep reconstruction. Furthermore, a novel loss function is developed to enhance the coupling between the non-local representations, which also enables an end-to-end training of NL-CSNet. Extensive experiments manifest that NL-CSNet outperforms existing state-of-the-art CS methods, while maintaining fast computational speed. Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Multim. | 3 |
| 2022 | Source Camera Identification with Multi-Scale Feature Fusion NetworkabstractSource camera identification (SCI) technology has attracted increasing attentions over the past few years. However, the existing methods suppress image content with denoising filters that are largely agnostic to the specific sensor pattern noise (SPN) signal of interest. Such practices may potentially degrade the performance of SPN-based SCI due to un-reliable SPNs, especially when forensic images are transmitted through social networking platforms. In this paper, we address the problem of SPN-based device identification and propose a multi-scale feature fusion network (MSFFN) to boost the sensor-based source camera identification attribution. Specifically, several image patches of different scales are selected and input into the MSFFN to extract the SPN. The MSFFN is a multi-scale encoder-decoder structure, which is used to suppress image content and improve source attribution. Subsequently, the content-independent SPN features of different scales are fused. At last, the fused features are used for image source identification. Experimental results compared with the state-of-the-art demonstrate that the proposed scheme achieves significant improvements, especially in the accuracy of social networking image source identification. Feng Jiang 0001, Shaohui Liu, Debin Zhao |
ICME | 2 |
| 2022 | Multi-Channel Adaptive Partitioning Network for Block-Based Image Compressive SensingabstractImage compressive sensing (CS) technology has attracted increasing attentions in the past few years, and a great deal deep learning-based methods have been proposed. However, the existing methods use fixed-scale blocks for sampling and re-construction. Such practice will inevitably result in the in-ability to distinguish between significant regions and background regions, and even waste excessive sampling resources on the background ones to a large extent. In this paper, we propose a novel multi-channel adaptive partitioning network for block-based image CS, in which image blocks of different scales are utilized to distinguish regions of different saliency. Specifically, an adaptive block partitioning method based on image saliency is put forward, using which significant regions are divided into large blocks and background regions are divided into small blocks. Subsequently, blocks of different scales are fed to different-channel networks for sampling to yield the compressed measurements. To improve the re-construction quality of the image, a scalable multi-scale re-construction network is proposed to recover the compressed measurements into the reconstructed image. Experimental results compared with the state-of-the-art show that the proposed scheme achieves significant improvements in terms of objective metrics and subjective visual image quality. Shaohui Liu, Feng Jiang 0001 |
ICME | 3 |
| 2022 | Learning from Hindsight Demonstrations
Mengxuan Shao, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
ICONIP (5) | 2 |
| 2022 | Hindsight Balanced Reward Shaping
Mengxuan Shao, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
ICONIP (5) | 2 |
| 2022 | Adversarial training of LSTM-ED based anomaly detection for complex time-series in cyber-physical-social systems
Haiqi Zhu, Shaohui Liu, Feng Jiang 0001 |
Pattern Recognit. Lett. | 3 |
| 2022 | Hierarchical complementary learning for weakly supervised object localization
Sabrina Narimene Benassou, Wuzhen Shi, Feng Jiang 0001, Abdallah Benzine |
Signal Process. Image Commun. | 3 |
| 2022 | Continuous Prediction of Lower-Limb Kinematics From Multi-Modal Biomedical SignalsabstractThe fast-growing techniques of measuring and fusing multi-modal biomedical signals enable advanced motor intent decoding schemes of lower-limb exoskeletons, meeting the increasing demand for rehabilitative or assistive applications of take-home healthcare. Challenges of exoskeletons’ motor intent decoding schemes remain in making a continuous prediction to compensate for the hysteretic response caused by mechanical transmission. In this paper, we solve this problem by proposing an ahead-of-time continuous prediction of lower-limb kinematics, with the prediction of knee angles during level walking as a case study. Firstly, an end-to-end kinematics prediction network(KinPreNet),1consisting of a feature extractor and an angle predictor, is proposed and experimentally compared with features and methods traditionally used in ahead-of-time prediction of gait phases. Secondly, inspired by the electromechanical delay(EMD), we further explore our algorithm’s capability of compensating response delay of mechanical transmission by validating the performance of the different sections of prediction time. And we experimentally reveal the time boundary of compensating the hysteretic response. Thirdly, a comparison of employing EMG signals or not is performed to reveal the EMG and kinematic signals’ collaborated contributions to the continuous prediction. During the experiments, EMG signals of nine muscles and knee angles calculated from inertial measurement unit (IMU) signals are recorded from ten healthy subjects. Our algorithm can predict knee angles with the averaged RMSE of 3.98 deg which is better than the 15.95-deg averaged RMSE of utilizing the traditional methods of ahead-of-time prediction. The best prediction time is in the interval of 27ms and 108ms. To the best of our knowledge, this is the first study of continuously predicting lower-limb kinematics in an ahead-of-time manner based on the electromechanical delay (EMD). Chunzhi Yi, Feng Jiang 0001, Shengping Zhang, Hao Guo 0015, Chifu Yang, Zhen Ding, Baichun Wei, Xiangyuan Lan, Huiyu Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Muscular Human Cybertwin for Internet of Everything: A Pilot StudyabstractThe cybertwin-driven 6G that can obtain static and dynamic data stream of users provide an exciting potential for a novel muscular human cybertwin beyond traditonally used artificial neural networks (ANNs) and musculoskeletal models (MSMs). In this article, we propose the conceptual design of the muscular human cybertwin and construct a baseline model with an improved generalization ability over ANN and an easier adaptation to new data distributions over MSMs. In particular, we for the first time propose to combine ANN and MSM, which benefits from the combination of learning-based approaches and analytical approaches. We then experimentally compare different manners of the combination and demonstrate the better combining manner on our testing case. Finally, we evaluate our method on an open-sourced dataset and on data from wearable sensors from the aspects of joint moment prediction accuracy, data efficiency, generalization ability, and time efficiency of personalization. Our proposed method achieves accuracy similar with ANN and over 30$\%$better than MSM with sufficient training data. Compared with ANN, the improved data efficiency is presented by the better accuracies with a small amount of training data, and the generalization ability to unseen walking conditions and new subjects are demonstrated by the over 70$\%$accuracy improvements. Moreover, when fine-tuning the model, our algorithm is demonstrated by the time 75$\%$shorter than calibrated MSM and the accuracy improvements. Chunzhi Yi, Sang Oh Park, Chifu Yang, Feng Jiang 0001, Zhen Ding, Jianfei Zhu, Jie Liu 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Spatio-Temporal Context Based Adaptive Camcorder Recording WatermarkingabstractVideo watermarking technology has attracted increasing attention in the past few years, and a great deal of traditional and deep learning-based methods have been proposed. However, these existing methods usually suffer from the following two challenges: First, most algorithms cannot resist camcorder recording attack, which limits their practical application. Second, watermark embedding may cause substantial degradation of video quality. Through analyzing the unique distortions presented in the camcorder recording process, including geometric distortion, temporal sampling distortion, sensor distortion and processing distortion, this paper proposes a novel spatio-temporal context based adaptive camcorder recording watermarking scheme STACR. In STACR, considering the geometric distortion and video visual quality, we embed the watermark by constructing a spatio-temporal histogram and incorporate a content features based adaptive locating algorithm to select embedding blocks and embedding strengths. As for the temporal sampling attack, we put forward a watermark correlation-based synchronization algorithm and combine it with cross-validation. Moreover, to resist the sensor distortion, we design a local matching-based algorithm to improve the extraction accuracy. In addition, grouped and repeated embedding strategies are combined to cope with the processing distortion. Experimental results compared with the state-of-the-art show that the proposed scheme achieves high video quality and is robust to geometric attacks, compression, scaling, transcoding, recoding, frame rate changes and especially for camcorder recording. Shaohui Liu, Wuzhen Shi, Feng Jiang 0001, Debin Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | LinkingPark: An automatic semantic table interpretation system
Shuang Chen 0003, Alperen Karaoglu, Carina Negreanu, Jin-Ge Yao, Jack Williams 0001, Feng Jiang 0001, Andrew D. Gordon 0001, Chin-Yew Lin |
J. Web Semant. | 7 |
| 2021 | Adaptive Flexible 3D Histogram WatermarkingabstractWatermarking technology has attracted increasing attentions in the past few years, and a great deal traditional and deep learning-based methods have been proposed. However, these methods usually suffer from the following three challenges: First, the current algorithms are designed separately for images or videos, and there is no universal solution. Second, most algorithms cannot resist screen recording, which limits its application. Third, some algorithms can only embed fixed-length watermarks and cannot handle the embedding capacity flexibly. In this paper, a novel watermarking scheme is proposed based on spatial and temporal histograms, in which two types of histogram watermarking are designed: One is constructed in the spatial domain, using the low-frequency characteristics of the image to change the shape of the histogram and embed the watermark. The other is established in the time domain, which uses the similarity of adjacent frames and combines texture features to modify the shape of the temporal histogram to embed the watermark. Experimental results compared with the state-of-the-art demonstrate that the proposed scheme achieves superior performance. Shaohui Liu, Wenxue Cui, Jinghua Zeng, Feng Jiang 0001, Debin Zhao |
ICME | 5 |
| 2021 | Small Object Recognition Using a Spatio-Temporal Neural NetworkabstractObject recognition at different scales has been a fundamental problem in computer vision. In particular, small object recognition attracts increasing attention recently. However, because of working on a single frame only, many recognizers’ performances become unacceptable in many practical application scenarios: very low resolutions, invisible small targets, extremely similar appearances etc. Motivated by the way humans deal with these challenging scenarios of object recognition, this paper introduces frame sequence and attention mechanism to compensate for mutilated information. Specifically, this paper proposes a spatiotemporal neural network (dubbed STNet) for small object recognition. STNet fixes the regions of interest with a super-resolution module, and focuses on the discriminative region with a spatio-temporal attention module. In addition, STNet applies a double layer long short-term memory subnet to make full use of the inter-frame information. Furthermore, this paper presents a challenging air-target recognition dataset ATSETC4 for evaluating the performance of each method in identifying small targets. Our model outperforms many state-of-the-art models on ATSETC4, including MobileNetV2 and SENet. In particular, STNet surpasses VGG11 at an average of 3.67%, even reaches 87.50% and 82.50% on 28 scale and 14 scale on AT-SETC4 respectively. Zhibo Liang, Shaohui Liu, Wuzhen Shi, Feng Jiang 0001 |
ICME | 5 |
| 2021 | A Bipolar Myoelectric Sensor-Enabled Human-Machine Interface Based On Spinal Module ActivationsabstractThe surface electromyography (sEMG) signal-based human-machine interface (HMI) has been widely used for various scenarios of physical human-robot interaction. However, current HMIs based on bipolar myoelectric sensors are hindered by the limitations of global sEMG features, which are prone to variability and delay. In this letter, we define a HMI that takes advantage of the underlying neural information of spinal module activations from bipolar sEMG signals, inspired by recent findings of neural codes. Firstly, the spinal module activations are identified by the spiking trains of the muscle synergies extracted from bipolar sEMG signals. Secondly, we extract the information encoded in both firing rates and spike timings of the spinal module activation in a population coding manner, which follows the information encoding principle of neurons. Thirdly, we map the series of spinal module activations into gait phases, locomotion modes, joint moment and human identity in order to experimentally reveal the physiological information contained in the spinal module activations. The contained information and the benefit of our design are demonstrated and experimentally explained by the presented results and comparisons with the traditionally used global sEMG features. The proposed bipolar myoelectric sensor-enabled human-machine interface could contribute to various scenarios of physical human-robot interaction. Chunzhi Yi, Feng Jiang 0001, Guangming Lu 0001, Chifu Yang, Zhen Ding, Jianfei Zhu, Jie Liu 0001 |
ICRA | 2 |
| 2021 | Protecting the Ownership of Deep Learning Models with An End-to-End Watermarking FrameworkabstractDeep neural network (DNN), as a key component of deep learning technology, plays a vital role in its development. Most major technology companies use deep neural network as a key component to build their artificial intelligence products and service. Building a deep neural network model requires us to pay a huge price: large-scale labeled data sets, a large number of computing resources, and highly specialized domain knowledge. Therefore, we believe that the model owner owns the intellectual property rights of the model, and it is very important to design a technology that protects the intellectual property rights of the deep neural network model and allows the owner to externally verify its copyright. Through statistical analysis of a large number of pre-trained network parameters, we propose an end-to-end network model protection framework-Deep Water based on the distribution of network model parameters. First, we propose a new research problem: embedding watermarks into deep neural networks. We also define the requirements for watermarking in deep neural networks, the embedding situation, and the types of attacks. Secondly, we propose a general framework for embedding the watermark into the parameter distribution function of each layer of the convolutional network. Our method does not harm the performance of the network where the watermark is placed, because the watermark is embedded when the host network is trained. Finally, we conducted a comprehensive experiment to reveal the potential of watermarking deep neural networks as the basis for this new research work. We proved that our framework can embed watermarks in the process of training deep neural networks from scratch and in the process of fine-tuning and distillation without compromising its performance. Even after migration learning and watermark overlay operations, the embedded watermark will not disappear. Even if 65% of the parameters are trimmed, the watermark remains intact. Wei Zhang 0192, Wenxue Cui, Feng Jiang 0001, Chifu Yang |
TrustCom | 3 |
| 2021 | DFD-Net: lung cancer detection from denoised CT scan image using deep learning
Worku Jifara Sori, Feng Jiang 0001, Arero W. Godana, Shaohui Liu |
Frontiers Comput. Sci. | 2 |
| 2021 | Smart healthcare-oriented online prediction of lower-limb kinematics and kinetics based on data-driven neural signal decoding
Chunzhi Yi, Feng Jiang 0001, Md. Zakirul Alam Bhuiyan, Chifu Yang, Xianzhong Gao, Hao Guo 0015, Jiantao Ma, Shen Su |
Future Gener. Comput. Syst. | 2 |
| 2021 | Entropy guided adversarial model for weakly supervised object localization
Sabrina Narimene Benassou, Wuzhen Shi, Feng Jiang 0001 |
Neurocomputing | 3 |
| 2021 | Combining Fields of Experts (FoE) and K-SVD methods in pursuing natural image priors
Feng Jiang 0001, Zhiyuan Chen 0007, Amril Nazir, Wuzhen Shi, Wei Xiang Lim, Shaohui Liu, Seungmin Rho |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Low-light image enhancement via deep Retinex decomposition and bilateral learning
Xiaoqian Lv, Yujing Sun 0004, Jun Zhang 0017, Feng Jiang 0001, Shengping Zhang |
Signal Process. Image Commun. | 4 |
| 2021 | Video Compressed Sensing Using a Convolutional Neural NetworkabstractRecently, a few image compressed sensing (CS) methods based on deep learning have been developed, which achieve remarkable reconstruction quality with low computational complexity. However, these existing deep learning-based image CS methods focus on exploring intraframe correlation while ignoring interframe cues, resulting in inefficiency when directly applied to video CS. In this paper, we propose a novel video CS framework based on a convolutional neural network (dubbed VCSNet) to explore both intraframe and interframe correlations. Specifically, VCSNet divides the video sequence into multiple groups of pictures (GOPs), of which the first frame is a keyframe that is sampled at a higher sampling ratio than the other nonkeyframes. In a GOP, the block-based framewise sampling by a convolution layer is proposed, which leads to the sampling matrix being automatically optimized. In the reconstruction process, the framewise initial reconstruction by using a linear convolutional neural network is first presented, which effectively utilizes the intraframe correlation. Then, the deep reconstruction with multilevel feature compensation is proposed, which compensates the nonkeyframes with the keyframe in a multilevel feature compensation manner. Such multilevel feature compensation allows the network to better explore both intraframe and interframe correlations. Extensive experiments on six benchmark videos show that VCSNet provides better performance over state-of-the-art video CS methods and deep learning-based image CS methods in both objective and subjective reconstruction quality. Wuzhen Shi, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Improving Entity Linking by Modeling Latent Entity Type InformationabstractExisting state of the art neural entity linking models employ attention-based bag-of-words context model and pre-trained entity embeddings bootstrapped from word embeddings to assess topic level context compatibility. However, the latent entity type information in the immediate context of the mention is neglected, which causes the models often link mentions to incorrect entities with incorrect type. To tackle this problem, we propose to inject latent entity type information into the entity embeddings based on pre-trained BERT. In addition, we integrate a BERT-based entity similarity score into the local context model of a state-of-the-art model to better capture latent entity type information. Our model significantly outperforms the state-of-the-art entity linking models on standard benchmark (AIDA-CoNLL). Detailed experiment analysis demonstrates that our model corrects most of the type errors produced by the direct baseline. Shuang Chen 0003, Jinpeng Wang 0001, Feng Jiang 0001, Chin-Yew Lin |
AAAI | 3 |
| 2020 | Multi-Stage Residual Hiding for Image-Into-Audio SteganographyabstractThe widespread application of audio communication technologies has speeded up audio data flowing across the Internet, which made it a popular carrier for covert communication. In this paper, we present a cross-modal steganography method for hiding image content into audio carriers while preserving the perceptual fidelity of the cover audio. In our framework, two multi-stage networks are designed: the first network encodes the decreasing multilevel residual errors inside different audio subsequences with the corresponding stage sub-networks, while the second network decodes the residual errors from the modified carrier with the corresponding stage sub-networks to produce the final revealed results. The multi-stage design of proposed framework not only make the controlling of payload capacity more flexible, but also make hiding easier because of the gradual sparse characteristic of residual errors. Qualitative experiments suggest that modifications to the carrier are unnoticeable by human listeners and that the decoded images are highly intelligible. Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Yongliang Liu, Debin Zhao |
ICASSP | 3 |
| 2020 | Classify and Explain: An Interpretable Convolutional Neural Network For Lung Cancer DiagnosisabstractThe deep network-based computer-aided diagnosis systems have encountered many difficulties in practical applications because of its "black box" feature. The crux of the problem is that these models should be explainable - the model should provide doctors rationales that can explain the diagnosis. In this paper, we present a novel network structure for visually interpretable lung nodule diagnosis. Our proposed model works in an end-to-end manner, consisting of an importance estimation network and a classification network. The former produces a diagnostic visual interpretation for each case, and the latter diagnoses the case. Based on a computed tomography image dataset (LUNA16) on pulmonary nodule, extensive experiments have been conducted, demonstrating that the proposed model can produce state-of-the-art diagnostic visual interpretations compared with all baseline methods. Donghao Gu, Zhaojing Wen, Feng Jiang 0001, Shaohui Liu |
ICASSP | 4 |
| 2020 | Generalized Critic Policy Optimization: A Model For Combining Advantage Estimates In Actor Critic MethodsabstractWe present a general model for actor critic methods that represent the possibility of combining value function estimations as a means to further reduce the policy gradient’s variance and improve the learning result. We show the potential of this architecture by implementing an example case to learn some of the Pybullet continous control robotic tasks with OpenAI Gym. We show by experimenting with a special case the effect of the external parameters on the overall performance of the policy optimization algorithm. Roumeissa Kitouni, Abderrahim Kitouni, Feng Jiang 0001 |
ICIP | 3 |
| 2020 | The online estimation of the joint angle based on the gravity acceleration using the accelerometer and gyroscope in the wireless networks
Zhen Ding, Chifu Yang, Jiantao Ma, Jianguo Wei, Feng Jiang 0001 |
Multim. Tools Appl. | 5 |
| 2020 | Obstructive sleep apnea detection using ecg-sensor with convolutional neural networks
Maowei Cheng, Yefu Wang, Shaohui Liu, Zhihong Tian 0001, Feng Jiang 0001 |
Multim. Tools Appl. | 6 |
| 2020 | Siamese Local and Global Networks for Robust Face TrackingabstractConvolutional neural networks (CNNs) have achieved great success in several face-related tasks, such as face detection, alignment and recognition. As a fundamental problem in computer vision, face tracking plays a crucial role in various applications, such as video surveillance, human emotion detection and human-computer interaction. However, few CNN-based approaches are proposed for face (bounding box) tracking. In this paper, we propose a face tracking method based on Siamese CNNs, which takes advantages of powerful representations of hierarchical CNN features learned from massive face images. The proposed method captures discriminative face information at both local and global levels. At the local level, representations for attribute patches (i.e:, eyes, nose and mouth) are learned to distinguish a face from another one, which are robust to pose changes and occlusions. At the global level, representations for each whole face are learned, which take into account the spatial relationships among local patches and facial characters, such as skin color and nevus. In addition, we build a new largescale challenging face tracking dataset to evaluate face tracking methods and to facilitate the research forward in this field. Extensive experiments on the collected dataset demonstrate the effectiveness of our method in comparison to several state-of-theart visual tracking methods. Yuankai Qi, Shengping Zhang, Feng Jiang 0001, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Image Compressed Sensing Using Convolutional Neural NetworkabstractIn the study of compressed sensing (CS), the two main challenges are the design of sampling matrix and the development of reconstruction method. On the one hand, the usually used random sampling matrices (e.g. GRM) are signal independent, which ignore the characteristics of the signal. On the other hand, the state-of-the-art image CS methods (e.g. GSR and MH) achieve quite good performance, however with much higher computational complexity. To deal with the two challenges, we propose an image CS framework using convolutional neural network (dubbed CSNet) that includes a sampling network and a reconstruction network, which are optimized jointly. The sampling network adaptively learns the sampling matrix from the training images, which makes the CS measurements retain more image structural information for better reconstruction. Specifically, three types of sampling matrices are learned, i.e. floating-point matrix, {0,1}-binary matrix, and {-1,+1}-bipolar matrix. The last two matrices are specially designed for easy storage and hardware implementation. The reconstruction network, which contains a linear initial reconstruction network and a non-linear deep reconstruction network, learns an end-to-end mapping between the CS measurements and the reconstructed images. Experimental results demonstrate that CSNet offers state-of-the-art reconstruction quality, while achieving fast running speed. In addition, CSNet with {0,1}-binary matrix, and {-1,+1}-bipolar matrix gets comparable performance with the existing deep learning based CS methods, and outperforms the traditional CS methods. What's more, the experimental results further suggest that the learned sampling matrices can improve the traditional image CS reconstruction methods significantly. Wuzhen Shi, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
IEEE Trans. Image Process. | 2 |
| 2020 | VINet: A Visually Interpretable Image Diagnosis NetworkabstractRecently, due to the black box characteristics of deep learning techniques, the deep network-based computer-aided diagnosis (CADx) systems have encountered many difficulties in practical applications. The crux of the problem is that these models should be explainable the model should give doctors rationales that can explain the diagnosis. In this paper, we propose a visually interpretable network (VINet) which can generate diagnostic visual interpretations while making accurate diagnoses. VINet is an end-to-end model consisting of an importance estimation network and a classification network. The former produces a diagnostic visual interpretation for each case, and the classifier diagnoses the case. In the classifier, by exploring the information in the diagnostic visual interpretation, the irrelevant information in the feature maps is eliminated by our proposed feature destruction process. This allows the classification network to concentrate on the important features and use them as the primary references for classification. Through a joint optimization of higher classification accuracy and eliminating as many irrelevant features as possible, a precise, fine-grained diagnostic visual interpretation, along with an accurate diagnosis, can be produced by our proposed network simultaneously. Based on a computed tomography image dataset (LUNA16) on pulmonary nodule, extensive experiments have been conducted, demonstrating that the proposed VINet can produce state-of-the-art diagnostic visual interpretations compared with all baseline methods. Donghao Gu, Feng Jiang 0001, Zhaojing Wen, Shaohui Liu, Wuzhen Shi, Guangming Lu 0001, Changsheng Zhou |
IEEE Trans. Multim. | 3 |
| 2020 | Delving Deeper in Drone-Based Person Re-Id by Employing Deep Decision Forest and Attributes FusionabstractDeep learning has revolutionized the field of computer vision and image processing. Its ability to extract the compact image representation has taken the person re-identification (re-id) problem to a new level. However, in most cases, researchers are focused on developing new approaches to extract more fruitful image representation and use it in the re-id task. The extra information about images is rarely taken into account because the traditional person re-id datasets usually do not have it. Nevertheless, the research in multimodal machine learning has demonstrated that the utilization of the information from different sources leads to better performance. In this work, we demonstrate how a person re-id problem can benefit from the utilization of multimodal data. We have used the UAV drone to collect and label the new person re-id dataset, which is composed of pedestrian images and its attributes. We have manually annotated this dataset with attributes, and in contrast to the recent research, we do not use the deep network to classify them. Instead, we employ the continuous bag-of-words model to extract the word embeddings from text descriptions and fuse it with features extracted from images. Then the deep neural decision forest is used for pedestrians classification. The extensive experiments on the collected dataset demonstrate the effectiveness of the proposed model. Aleksei Grigorev, Shaohui Liu, Zhihong Tian 0001, Jianxin Xiong, Seungmin Rho, Feng Jiang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2020 | Deep Learning Based Multi-Channel Intelligent Attack Detection for Data SecurityabstractDeep learning methods, e.g., convolutional neural networks (CNNs) and Recurrent Neural Networks (RNNs), have achieved great success in image processing and natural language processing especially in high level vision applications such as recognition and understanding. However, it is rarely used to solve information security problems such as attack detection studied in this paper. Here, we move forward a step and propose a novel multi-channel intelligent attack detection method based on long short term memory recurrent neural networks (LSTM-RNNs). To achieve high detection rate, data preprocessing, feature abstraction, and multi-channel training and detection are seamlessly integrated into an end-to-end detection framework. Data preprocessing provides high-quality data for subsequent processing, then different types of features are extracted from the processed data. Multi-channel processing is used to generate classifiers by training neural networks with different types of features, which preserve attack features of input vectors and classify the attack from normal data. With the results of the classifier's attack detection, we introduce a voting algorithm to decide whether the input data is an attack or not. Experimental results validate that the proposed attack detection method greatly outperforms several attack detection methods that use feature detection and Bayesian or SVM classifiers. Feng Jiang 0001, Yunsheng Fu, Brij B. Gupta, Seungmin Rho, Fang Lou, Fanzhi Meng, Zhihong Tian 0001 |
IEEE Trans. Sustain. Comput. | 1 |
| 2019 | Scalable Convolutional Neural Network for Image Compressed SensingabstractRecently, deep learning based image Compressed Sensing (CS) methods have been proposed and demonstrated superior reconstruction quality with low computational complexity. However, the existing deep learning based image CS methods need to train different models for different sampling ratios, which increases the complexity of the encoder and decoder. In this paper, we propose a scalable convolutional neural network (dubbed SCSNet) to achieve scalable sampling and scalable reconstruction with only one model. Specifically, SCSNet provides both coarse and fine granular scalability. For coarse granular scalability, SCSNet is designed as a single sampling matrix plus a hierarchical reconstruction network that contains a base layer plus multiple enhancement layers. The base layer provides the basic reconstruction quality, while the enhancement layers reference the lower reconstruction layers and gradually improve the reconstruction quality. For fine granular scalability, SCSNet achieves sampling and reconstruction at any sampling ratio by using a greedy method to select the measurement bases. Compared with the existing deep learning based image CS methods, SCSNet achieves scalable sampling and quality scalable reconstruction at any sampling ratio with only one model. Experimental results demonstrate that SCSNet has the state-of-the-art performance while maintaining a comparable running speed with the existing deep learning based image CS methods. Wuzhen Shi, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
CVPR | 2 |
| 2019 | Enhancing Neural Data-To-Text Generation Models with External Background KnowledgeabstractShuang Chen, Jinpeng Wang, Xiaocheng Feng, Feng Jiang, Bing Qin, Chin-Yew Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shuang Chen 0003, Jinpeng Wang 0001, Feng Jiang 0001, Bing Qin 0001, Chin-Yew Lin |
EMNLP/IJCNLP (1) | 4 |
| 2019 | DSNET: Accelerate Indoor Scene Semantic SegmentationabstractIn this paper, we address the problem of real-time image semantic segmentation for indoor scene. As for scene parsing, both accuracy and speed are equally important. However, most of existing methods mainly focus on improving accuracy rather than speed. How to find a balance between accuracy and speed is crucial for real-time semantic segmentation tasks. To tackle this problem, we propose a lightweight framework with depthwise dilation residual module and multi-scale information integration module. A single DSNet yields the performance of mIoU accuracy 32.12% on SUN RGB-D dataset and accuracy 26.32% on ADE20K dataset, which is the most challenge scene parsing dataset. Besides, our system yields real-time inference on a single NVIDIA GPU. Feng Jiang 0001, Feng Guo 0005, Rongrong Ji |
ICASSP | 1 |
| 2019 | Continuous Bidirectional Optical Flow for Video Frame Sequence InterpolationabstractExisting optical flow-based frame interpolation frameworks usually suffer from two problems. First, it is difficult to accurately estimate both large motion and fine motion in the optical flow estimation stage. Second, the hole problem and occlusion problem cannot be efficiently solved in the pixel synthesis step. In this paper, we propose a novel optical flowbased frame interpolation framework, which consists of two submodules: optical flow network and pixel synthesis network. In the optical flow network, we estimate bidirectional optical flow sequences iteratively, which makes full use of the continuity of motion and therefore improves the accuracy of the optical flow estimation. Besides, a novel multi-scale architecture is developed to capture finer motions. In the pixel synthesis network, we fuse the statistical information generated during forward warping to solve the hole problem and the occlusion problem. Experimental results demonstrate that the proposed method achieves superior performance compared to state-of-the-art methods. Donghao Gu, Zhaojing Wen, Wenxue Cui, Rui Wang 0093, Feng Jiang 0001, Shaohui Liu |
ICME | 5 |
| 2019 | A Video Post-Filter Deblocking Method Based on Temporal Boosting Residual NetworksabstractBlock-based hybrid coding is widely used in video compression. As the bit rate decreases, the quantization becomes lossy resulting in unacceptable blocking artifacts. Although advanced video encoder with loop filter has achieved promising results, the problem still remains unsolved. Most of the previous solutions regard a video sequence as a group of independent frames, without considering their temporal relationships. To address the above issue, this paper designs a temporal boosting residual network aiming to integrate the temporal information and the structural information. And the information is represented as the residual values of previous frames which are extracted by the deep residual network. The entire framework is a cascading architecture that implies a coarse-to-fine processing. The proposed solution does not modify any module of the codec, it only takes the lossy frames as input, and outputs the enhanced frames, which is a typical end-to-end mapping. Experimental results validate that our framework has achieved 0.6-1.0 dB improvement on average based on HEVC software and outperformed the state-of-the-art methods in both objective and perceptual quality. Shaohui Liu, Feng Jiang 0001, Xiaoshuai Sun, Yongliang Liu |
ICME | 3 |
| 2019 | Hierarchical residual learning for image denoising
Wuzhen Shi, Feng Jiang 0001, Shengping Zhang, Rui Wang 0093, Debin Zhao, Huiyu Zhou 0001 |
Signal Process. Image Commun. | 2 |
| 2019 | Medical image denoising using convolutional neural network: a residual learning approach
Worku Jifara Sori, Feng Jiang 0001, Seungmin Rho, Maowei Cheng, Shaohui Liu |
J. Supercomput. | 2 |
| 2018 | An Efficient Deep Convolutional Laplacian Pyramid Architecture for Cs Reconstruction At Low Sampling RatiosabstractThe compressed sensing (CS) has been successfully applied to image compression in the past few years as most image signals are sparse in a certain domain. Several CS reconstruction models have been proposed and obtained superior performance. However, these methods suffer from blocking artifacts or ringing effects at low sampling ratios in most cases. To address this problem, we propose a deep convolutional Laplacian Pyramid Compressed Sensing Network (LapC-SNet) for CS, which consists of a sampling sub-network and a reconstruction sub-network. In the sampling sub-network, we utilize a convolutional layer to mimic the sampling operator. In contrast to the fixed sampling matrices used in traditional CS methods, the filters used in our convolutional layer are jointly optimized with the reconstruction sub-network. In the reconstruction sub-network, two branches are designed to reconstruct multi -scale residual images and muti -scale target images progressively using a Laplacian pyramid architecture. The proposed LapCSNet not only integrates multi-scale information to achieve better performance but also reduces computational cost dramatically. Experimental results on benchmark datasets demonstrate that the proposed method is capable of reconstructing more details and sharper edges against the state-of-the-arts methods. Wenxue Cui, Heyao Xu, Xinwei Gao, Shengping Zhang, Feng Jiang 0001, Debin Zhao |
ICASSP | 5 |
| 2018 | Deep Neural Network Based Sparse Measurement Matrix for Image Compressed SensingabstractGaussian random matrix (GRM) has been widely used to generate linear measurements in compressed sensing (CS) of natural images. However, there actually exist two disadvantages with GRM in practice. One is that GRM has large memory requirement and high computational complexity, which restrict the applications of CS. Another is that the CS measurements randomly obtained by GRM cannot provide sufficient reconstruction performances. In this paper, a Deep neural network based Sparse Measurement Matrix (DSMM) is learned by the proposed convolutional network to reduce the sampling computational complexity and improve the CS reconstruction performance. Two sub-networks are included in the proposed network, which are the sampling sub-network and the reconstruction sub-network. In the sampling sub-network, the sparsity and the normalization are both considered by the limitation of the storage and the computational complexity. In order to improve the CS reconstruction performance, a reconstruction sub-network are introduced to help enhance the sampling sub-network. So by the offline iterative training of the proposed end-to-end network, the DSMM is generated for accurate measurement and excellent reconstruction. Experimental results demonstrate that the proposed DSMM outperforms GRM greatly on representative CS reconstruction methods. Wenxue Cui, Feng Jiang 0001, Xinwei Gao, Wen Tao, Debin Zhao |
ICIP | 2 |
| 2018 | Multi-Scale Deep Networks for Image Compressed SensingabstractAs a successful deep model applied in image compressed sensing, the Compressed Sensing Network (CSNet) has demonstrated superior performance to the previous handcrafted models in both running speed and reconstruction quality. However, CSNet trains different models for different sampling rates that hinders it from practical usage since too many models need to store. In this paper, we propose multi-scale deep network for image compressed sensing. We still use a sampling network to learn the sampling operator and implement the compressed sampling process. Given the compressed measurements, the reconstruction network directly maps them to the desired reconstructed images. There are three main differences in comparison with CSNet. Firstly, this paper proposes to use an unified deep reconstruction network for all sampling rates that decreases large amount of storage requirements. Secondly, we redesign a better deep reconstruction network using the popular residual learning technology. Finally, we investigate an image local smooth prior based loss function to enhance image structural information. Extensive experimental results show that the proposed multi-scale deep network based image compressed sensing method outperforms many other state-of-the-art methods. Wuzhen Shi, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
ICIP | 2 |
| 2018 | Classification Guided Deep Convolutional Network for Compressed SensingabstractCompressed Sensing (CS) has been successfully applied to image compression in the past few years. However, there are still several challenges that restrict its applications in practice including large memory requirement and unsatisfactory reconstruction performance. To address these challenges, in this paper, we propose a classification guided deep convolutional network for image compressed sensing (CCSNet), which includes a sampling sub-network and a reconstruction sub-network. In the sampling sub-network, multiple convolutional layers are used to sample the original image, which significantly reduces the parameters of the sampling matrix while causes performance degradation moderately compared against existing convolution based sampling methods. In the reconstruction sub-network, a novel two-branch architecture is proposed to improve the adaptability of the model to various textures in natural images. The first branch, named the classification branch, is to classify the sampled measurements of the original image to one of the predefined textural classes. The second branch, named the reconstruction branch, consists of multiple sub-branches, which are responsible for reconstructing the original images belonging to the corresponding textural classes. By jointly utilizing two sub-networks, the entire network can be trained in the form of end-to-end metric with a joint loss function. Experimental results demonstrate that the proposed method provides a significant quality improvement in terms of PSNR compared against state-of-the-art methods. Wenxue Cui, Shaohui Liu, Shengping Zhang, Yashu Liu 0003, Heyao Xu, Xinwei Gao, Feng Jiang 0001, Debin Zhao |
ICPR | 7 |
| 2018 | An Efficient Deep Quantized Compressed Sensing Coding Framework of Natural ImagesabstractTraditional image compressed sensing (CS) coding frameworks solve an inverse problem that is based on the measurement coding tools (prediction, quantization, entropy coding, etc.) and the optimization based image reconstruction method. These CS coding frameworks face the challenges of improving the coding efficiency at the encoder, while also suffering from high computational complexity at the decoder. In this paper, we move forward a step and propose a novel deep network based CS coding framework of natural images, which consists of three sub-networks: sampling sub-network, offset sub-network and reconstruction sub-network that responsible for sampling, quantization and reconstruction, respectively. By cooperatively utilizing these sub-networks, it can be trained in the form of an end-to-end metric with a proposed rate-distortion optimization loss function. The proposed framework not only improves the coding performance, but also reduces the computational cost of the image reconstruction dramatically. Experimental results on benchmark datasets demonstrate that the proposed method is capable of achieving superior rate-distortion performance against state-of-the-art methods. Wenxue Cui, Feng Jiang 0001, Xinwei Gao, Shengping Zhang, Debin Zhao |
ACM Multimedia | 2 |
| 2018 | A hybrid framework of data hiding and encryption in H.264/SVC
Shaohui Liu, Seungmin Rho, Worku Jifara Sori, Feng Jiang 0001 |
Discret. Appl. Math. | 4 |
| 2018 | Hyperspectral classification based on spectral-spatial convolutional neural networks
Feng Jiang 0001, Chifu Yang, Seungmin Rho, Weizheng Shen, Shaohui Liu |
Eng. Appl. Artif. Intell. | 2 |
| 2018 | User-perceived quality aware adaptive streaming of 3D multi-view video plus depth over the internet
Nabin Kumar Karn, Hongli Zhang 0001, Feng Jiang 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Feature-preserving mesh denoising based on guided normal filtering
Shaohui Liu, Seungmin Rho, Feng Jiang 0001 |
Multim. Tools Appl. | 4 |
| 2018 | Plant identification based on very deep convolutional neural networks
Heyan Zhu, Qinglin Liu, Yuankai Qi, Feng Jiang 0001, Shengping Zhang |
Multim. Tools Appl. | 5 |
| 2018 | Very deep feature extraction and fusion for arrhythmias detection
Moussa Amrani, Mohamed Hammad, Feng Jiang 0001, Kuanquan Wang, Amel Amrani |
Neural Comput. Appl. | 3 |
| 2018 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractDeep learning, e.g., convolutional neural networks (CNNs), has achieved great success in image processing and computer vision especially in high-level vision applications, such as recognition and understanding. However, it is rarely used to solve low-level vision problems such as image compression studied in this paper. Here, we move forward a step and propose a novel compression framework based on CNNs. To achieve high-quality image compression at low bit rates, two CNNs are seamlessly integrated into an end-to-end compression framework. The first CNN, named compact convolutional neural network (ComCNN), learns an optimal compact representation from an input image, which preserves the structural information and is then encoded using an image codec (e.g., JPEG, JPEG2000, or BPG). The second CNN, named reconstruction convolutional neural network (RecCNN), is used to reconstruct the decoded image with high quality in the decoding end. To make two CNNs effectively collaborate, we develop a unified end-to-end learning algorithm to simultaneously learn ComCNN and RecCNN, which facilitates the accurate reconstruction of the decoded image using RecCNN. Such a design also makes the proposed compression framework compatible with existing image coding standards. Experimental results validate that the proposed compression framework greatly outperforms several compression frameworks that use existing image coding standards with the state-of-the-art deblocking or denoising post-processing methods. Feng Jiang 0001, Wen Tao, Shaohui Liu, Jie Ren 0016, Xun Guo 0002, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | BoMW: Bag of Manifold Words for One-Shot Learning Gesture Recognition From KinectabstractIn this paper, we study one-shot learning gesture recognition on RGB-D data recorded from Microsoft's Kinect. To this end, we propose a novel bag of manifold words (BoMW)-based feature representation on symmetric positive definite (SPD) manifolds. In particular, we use covariance matrices to extract local features from RGB-D data due to its compact representation ability as well as the convenience of fusing both RGB and depth information. Since covariance matrices are SPD matrices and the space spanned by them is the SPD manifold, traditional learning methods in the Euclidean space, such as sparse coding, cannot be directly applied to them. To overcome this problem, we propose a unified framework to transfer the sparse coding on SPD manifolds to the one on the Euclidean space, which enables any existing learning method to be used. After building BoMW representation on a video from each gesture class, a nearest neighbor classifier is adopted to perform the one-shot learning gesture recognition. Experimental results on the ChaLearn gesture data set demonstrate the outstanding performance of the proposed one-shot learning gesture recognition method compared against the state-of-the-art methods. The effectiveness of the proposed feature extraction method is also validated on a new RGB-D action recognition data set. Lei Zhang 0036, Shengping Zhang, Feng Jiang 0001, Yuankai Qi, Jun Zhang 0017, Yuliang Guo, Huiyu Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Point-to-Set Distance Metric Learning on Deep Representations for Visual TrackingabstractFor autonomous driving application, a car shall be able to track objects in the scene in order to estimate where and how they will move such that the tracker embedded in the car can efficiently alert the car for effective collision-avoidance. Traditional discriminative object tracking methods usually train a binary classifier via a support vector machine (SVM) scheme to distinguish the target from its background. Despite demonstrated success, the performance of the SVM-based trackers is limited because the classification is carried out only depending on support vectors (SVs) but the target's dynamic appearance may look similar to the training samples that have not been selected as SVs, especially when the training samples are not linearly classifiable. In such cases, the tracker may drift to the background and fail to track the target eventually. To address this problem, in this paper, we propose to integrate the point-to-set/image-to-imageSet distance metric learning (DML) into visual tracking tasks and take full advantage of all the training samples when determining the best target candidate. The point-to-set DML is conducted on convolutional neural network features of the training data extracted from the starting frames. When a new frame comes, target candidates are first projected to the common subspace using the learned mapping functions, and then the candidate having the minimal distance to the target template sets is selected as the tracking result. Extensive experimental results show that even without model update the proposed method is able to achieve favorable performance on challenging image sequences compared with several state-of-the-art trackers. Shengping Zhang, Yuankai Qi, Feng Jiang 0001, Xiangyuan Lan, Pong C. Yuen, Huiyu Zhou 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | A Data Leakage Prevention Method Based on the Reduction of Confidential and Context Terms for Smart Mobile DevicesabstractEarly data leakage protection methods for smart mobile devices usually focus on confidential terms and their context, which truly prevent some kinds of data leakage events. However, with the high dimensionality and redundancy of text data, it is difficult to detect the documents which contain confidential contents accurately. Our approach updates cluster graph structure based on CBDLP (Data Leakage Protection Based on Context) model by computing the importance of confidential terms and the terms within the range of their context. By applying CBDLP with pruning procedure which has been validated, we further remove the redundancy terms and noise terms. Actually, not only can confidential terms be accurately detected but also the sophisticated rephrased confidential contents are detected during the experiments. Zhihong Tian 0001, Jing Qiu 0002, Feng Jiang 0001 |
Wirel. Commun. Mob. Comput. | 4 |
| 2017 | Convolutional Neural Networks Based Intra Prediction for HEVCabstractSummary form only given. Traditional intra prediction methods for HEVC rely on using the nearest reference lines for predicting a block, which ignore much richer context between the current block and its neighboring blocks and therefore cause inaccurate prediction especially when weak spatial correlation exists between the current block and the reference lines. To overcome this problem, in this paper, an intra-prediction convolutional neural network (IPCNN) is proposed for intra prediction, which exploits the rich context of the current block and therefore is capable of improving the accuracy of predicting the current block. Meanwhile, the reconstruction of the three nearest blocks can also be refined. To the best of our knowledge, this is the first paper that directly applies CNNs to intra prediction for HEVC. Experimental results validate the effectiveness of applying CNNs to intra prediction and the proposed method can achieve 0.70% bitrate reduction compared to HEVC reference software HM-14.0. Wenxue Cui, Tao Zhang 0013, Shengping Zhang, Feng Jiang 0001, Wangmeng Zuo, Zhaolin Wan, Debin Zhao |
DCC | 4 |
| 2017 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractSummary form only given. Traditional image coding standards (such as JPEG and JPEG2000) make the decoded image suffer from many blocking artifacts or noises since the use of big quantization steps. To overcome this problem, we proposed an end-to-end compression framework based on two CNNs, as shown in Figure 1, which produce a compact representation for encoding using a third party coding standard and reconstruct the decoded image, respectively. To make two CNNs effectively collaborate, we develop a unified end-to-end learning framework to simultaneously learn CrCNN and ReCNN such that the compact representation obtained by CrCNN preserves the structural information of the image, which facilitates to accurately reconstruct the decoded image using ReCNN and also makes the proposed compression framework compatible with existing image coding standards. Wen Tao, Feng Jiang 0001, Shengping Zhang, Jie Ren 0016, Wuzhen Shi, Wangmeng Zuo, Xun Guo 0002, Debin Zhao |
DCC | 2 |
| 2017 | Single image super-resolution with dilated convolution based multi-scale information learning inception moduleabstractTraditional works have shown that patches in a natural image tend to redundantly recur many times inside the image, both within the same scale, as well as across different scales. Make full use of these multi-scale information can improve the image restoration performance. However, the current proposed deep learning based restoration methods do not take the multi-scale information into account. In this paper, we propose a dilated convolution based inception module to learn multi-scale information and design a deep network for single image super-resolution. Different dilated convolution learns different scale feature, then the inception module concatenates all these features to fuse multi-scale information. In order to increase the reception field of our network to catch more contextual information, we cascade multiple inception modules to constitute a deep network to conduct single image super-resolution. With the novel dilated convolution based inception module, the proposed end-to-end single image super-resolution network can take advantage of multi-scale information to improve image super-resolution performance. Experimental results show that our proposed method outperforms many state-of-the-art single image super-resolution methods. Wuzhen Shi, Feng Jiang 0001, Debin Zhao |
ICIP | 2 |
| 2017 | Deep networks for compressed image sensingabstractThe compressed sensing (CS) theory has been successfully applied to image compression in the past few years as most image signals are sparse in a certain domain. Several CS reconstruction models have been recently proposed and obtained superior performance. However, there still exist two important challenges within the CS theory. The first one is how to design a sampling mechanism to achieve an optimal sampling efficiency, and the second one is how to perform the reconstruction to get the highest quality to achieve an optimal signal recovery. In this paper, we try to deal with these two problems with a deep network. First of all, we train a sampling matrix via the network training instead of using a traditional manually designed one, which is much appropriate for our deep network based reconstruct process. Then, we propose a deep network to recover the image, which imitates traditional compressed sensing reconstruction processes. Experimental results demonstrate that our deep networks based CS reconstruction method offers a very significant quality improvement compared against state-of-the-art ones. Wuzhen Shi, Feng Jiang 0001, Shengping Zhang, Debin Zhao |
ICME | 2 |
| 2017 | Structured entropy of primitive: big data-based stereoscopic image quality assessmentabstractThe ultimate receiver of image and video is human visual system (HVS). It is an important problem in the domain of image and video processing that how to establish visual information representation model meeting the HVS perception property. In this study, authors give theory analysis and experiment results to prove that l_1 norm‐based entropy of primitive (EoP) is superior to the l_0 norm‐based EoP for the monocular cue in image quality assessment. By developing the concept of mutual information of primitive (MIP) as the binocular cue, an l_1 EoP‐based stereoscopic image quality assessment metric is proposed. With EoP as monocular cue and MIP as binocular cue, the relative entropy between the original stereoscopic image and the distorted one is explored to predict the quality score with support vector regression. To avoid destroying image's structured information, the structured EoP (SEoP) is further explored to measure the stereoscopic image information. Extensive experimental results demonstrate that the stereoscopic image quality assessment algorithm with SEoP as monocular cue and MIP as binocular cue outperforms many state‐of‐the‐art ones. Chifu Yang, Seungmin Rho, Shaohui Liu, Feng Jiang 0001 |
IET Image Process. | 5 |
| 2017 | Depth estimation from single monocular images using deep hybrid network
Aleksei Grigorev, Feng Jiang 0001, Seungmin Rho, Worku Jifara Sori, Shaohui Liu, Sergey V. Sai |
Multim. Tools Appl. | 2 |
| 2017 | Hyperspectral image compression based on online learning spectral features dictionary
Worku Jifara Sori, Feng Jiang 0001, Huapeng Wang, Aleksei Grigorev, Shaohui Liu |
Multim. Tools Appl. | 2 |
| 2016 | Group-based sparse representation for low lighting image enhancementabstractThe Group-based Sparse Representation (GSR) is able to sparsely represent natural images in the domain of group, which enforces the intrinsic local sparsity and nonlocal self-similarity of images simultaneously in a unified framework. And the GSR-driven L0 minimization method for image restoration has been proposed. This paper expands the application of GSR from image restoration to low lighting image enhancement. The GSR is not used to represent the natural images anymore, but representing the transmission map of the haze image and recovering it. Because the transmission map is very important for the low lighting image enhancement, the dark channel prior based enhancement method with the enhanced transmission map can get a better enhanced results. Different from other methods, we evaluate the quality of the enhanced images not only by qualitative analysis but also by quantitative results. Extensive experiments on low lighting image show that the GSR-based method gets a better enhancement result than many current state-of-the-art ones. Wuzhen Shi, Feng Jiang 0001, Debin Zhao, Weizheng Shen |
ICIP | 3 |
| 2016 | Image Entropy of Primitive and visual quality assessmentabstractRecently, the concept of Entropy of Primitive (EoP) has been proposed to measure the image visual information. Some successful EoP based application also be developed. In this paper, we further explore the concept of EoP and propose an improved version: the L1 norm based EoP. Our EoP takes full account of the properties of a dictionary's layered structure and the characteristic of a basis pursuit method. Experimental results show that the L1 norm based EoP is superior to the L0 norm based one in measuring the image visual information. The curve of L1 norm based EoP holds a more consistent monotonicity with SSIM, its values is not trapped in the local convergence and the convergence value is less than that of the L0 norm based one. With the convergence characteristics of EoP, we further explore its application in stereoscopic image quality assessment (SIQA). With EoP as monocular cue and mutual information of primitive (MIP) as binocular cue, the relative entropy between the original stereoscopic image and the distorted one is used to compute the quality score by a prediction function which is trained using support vector regression (SVR). Extensive experimental results show that our new EoP based SIQA outperforms many state-of-the-art on the LIVE phase II databases. Wuzhen Shi, Feng Jiang 0001, Debin Zhao |
ICIP | 2 |
| 2016 | Hierarchical frame based spatial-temporal recovery for video compressive sensing coding
Xinwei Gao, Feng Jiang 0001, Shaohui Liu, Wenbin Che, Xiaopeng Fan 0001, Debin Zhao |
Neurocomputing | 2 |
| 2016 | 3D object retrieval with multi-feature collaboration and bipartite graph matching
Yan Zhang 0109, Feng Jiang 0001, Seungmin Rho, Shaohui Liu, Debin Zhao, Rongrong Ji |
Neurocomputing | 2 |
| 2016 | Big data driven decision making and multi-prior models collaboration for media restoration
Feng Jiang 0001, Seungmin Rho, Bo-Wei Chen, Debin Zhao |
Multim. Tools Appl. | 1 |
| 2016 | Optimal filter based on scale-invariance generation of natural images
Feng Jiang 0001, Bo-Wei Chen, Seungmin Rho, Wen Ji 0003, Liqiang Pan, Debin Zhao |
J. Supercomput. | 1 |
| 2015 | Spatial-temporal recovery for hierarchical frame based video compressed sensingabstractIn this paper, the hierarchical frame based video compressed sensing (CS) framework is proposed, which outperforms the traditional framework through the better exploitation of frames correlation with reference frames, the unequal sample subrates setting among frames in different layers and the reduction of the error propagation. By considering the spatial and temporal correlations of the video sequence, a spatial-temporal sparse representation based recovery is proposed for this framework. The similar blocks in both the current frame and these recovered reference frames are composed as a spatial-temporal group, which is defined as the unit of the sparse representation. By exploiting the low dimensional subspace description of each group, the video CS recovery is converted as a low-rank matrix approximation problem, which can be solved by exploiting the hard thresholding and the gradient descent. Experimental results show that the proposed method achieves better performance against both the state-of-art still-image CS recovery algorithms and the existing residual domain based video CS reconstruction approaches. Wenbin Che, Xinwei Gao, Xiaopeng Fan 0001, Feng Jiang 0001, Debin Zhao |
ICIP | 4 |
| 2015 | Reference image based method of region of interest enhancement for haze imageabstractDifferent from general algorithms of haze removal and low lighting image enhancement, which only use the information of image to process, this paper adds a reference image to get more information for the algorithm and focuses on enhancing region of interest of an image based on the reference one. With the reference image, the haze one can be divided into Region of Interest (RoI) and Region of no Interest (non-RoI). Furthermore, the reference image can provide more useful information for computing the transmission map and atmospheric light. For the non-RoI region, a more robust transmission map and minimizing reconstruction error cost function based method to estimate atmospheric light has been proposed. Because the atmospheric light is a global variable, the optimized one is also suitable for the RoI region. With the global optimized atmospheric light, an optimized transmission map can be got for the RoI region. The RoI region can be enhanced via the optimal transmission map and atmosphere light. Theoretical analysis gives eloquent proof proving that the proposed method is definitely better than the traditional dark-channel-prior-based methods due to our better transmission map and atmosphere light. Extensive experiments also show the expected results. Wuzhen Shi, Xinwei Gao, Boqi Chen, Feng Jiang 0001, Debin Zhao |
ICIP | 4 |
| 2015 | Discriminating features learning in hand gesture classificationabstractThe advent and popularity of Kinect provides a new choice and opportunity for hand gesture recognition (HGR) research. In this study, the authors propose a discriminating features extraction for HGR, in which features from red, green and blue (RGB) images and depth images are both explored. More specifically, histogram of oriented gradient feature, local binary pattern feature, structure feature and three‐dimensional voxel feature are first extracted from RGB images and depth images, then these features are further reduced with a novel deflation orthogonal discriminant analysis, which enhances the discriminative ability of the features with supervised subspace projection. The extensive experimental results show that the proposed method improves the HGR performance significantly. Feng Jiang 0001, Cuihua Wang, Debin Zhao |
IET Comput. Vis. | 1 |
| 2015 | Multi-layered gesture recognition with Kinect
Feng Jiang 0001, Shengping Zhang, Debin Zhao |
J. Mach. Learn. Res. | 1 |
| 2015 | Robust Visual Tracking Using Structurally Random Projection and Weighted Least SquaresabstractSparse representation-based visual tracking approaches have attracted increasing interests in the community in recent years. The main idea is to linearly represent each target candidate using a set of target and trivial templates, while imposing a sparsity constraint onto the representation coefficients. After we obtain the coefficients using ℓ1-norm minimization methods, the candidate with the lowest error, when it is reconstructed using only the target templates and the associated coefficients, is considered as the tracking result. In spite of promising system performance widely reported, it is unclear if the performance of these trackers can be maximized. In addition, computational complexity caused by the dimensionality of the feature space limits these algorithms in real-time applications. In this paper, we propose a real-time visual tracking method based on structurally random projection (RP) and weighted least squares (WLS) techniques. In particular, to enhance the discriminative capability of the tracker, we introduce background templates to the linear representation framework. To handle appearance variations over time, we relax the sparsity constraint using a WLS method to obtain the representation coefficients. To further reduce the computational complexity, structurally RP is used to reduce the dimensionality of the feature space, while preserving the pairwise distances between the data points in the feature space. Experimental results show that the proposed approach outperforms several state-of-the-art tracking methods. Shengping Zhang, Huiyu Zhou 0001, Feng Jiang 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Game theory based no-reference perceptual quality assessment for stereoscopic images
Feng Jiang 0001, K. Bharanitharan, Shovan Barma, Debin Zhao |
J. Supercomput. | 1 |
| 2015 | Face hallucination and recognition in social network services
Feng Jiang 0001, Seungmin Rho, Bo-Wei Chen, Xiaodan Du 0003, Debin Zhao |
J. Supercomput. | 1 |
| 2014 | Image Restoration via Multi-prior Collaboration
Feng Jiang 0001, Shengping Zhang, Debin Zhao, Sun-Yuan Kung |
ACCV (3) | 1 |
| 2013 | Structural Group Sparse Representation for Image Compressive Sensing RecoveryabstractCompressive Sensing (CS) theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet, contour let and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for image compressive sensing recovery via structural group sparse representation (SGSR) modeling, which enforces image sparsity and self-similarity simultaneously under a unified framework in an adaptive group domain, thus greatly confining the CS solution space. In addition, an efficient iterative shrinkage/thresholding algorithm based technique is developed to solve the above optimization problem. Experimental results demonstrate that the novel CS recovery strategy achieves significant performance improvements over the current state-of-the-art schemes and exhibits nice convergence. Jian Zhang 0018, Debin Zhao, Feng Jiang 0001, Wen Gao 0001 |
DCC | 3 |
| 2013 | Estimation of end-to-end distortion of virtual view for error-resilient depth map codingabstractIn this paper, we propose to estimate the end-to-end distortion of virtual view for error-resilient depth map coding by analyzing the view-rendering process. The end-to-end distortion is estimated according to the expectation of depth distortion. The expectation of depth distortion is estimated based on the error-propagated distortion from the reference frame, which is calculated recursively for each pixel by considering the source characteristics, channel condition and error concealment method. The error-propagated distortion in terms of each frame is updated after the frame is encoded, which is used for the subsequent frames. Experimental results demonstrate that the proposed algorithm achieves substantial and consistent gains in comparison with the random intra updating method and the conventional rate-distortion method for error-resilient depth map coding. Min Gao 0002, Xiaopeng Fan 0001, Debin Zhao, Feng Jiang 0001, Wen Gao 0001 |
ICIP | 4 |
| 2013 | From relation between filter-based MRFs model and sparsity based method to the pursuit of natural images spaceabstractIn the pursuit of natural image prior, responses to the specific filter bank and the character of sparse representation are of the most important clues. Based on these clues, many effective and successful algorithms are proposed and widely used in low vision tasks. Up to now, the corresponding researches with these clues are developed in relatively independent ways. In this paper, taking K-SVD as an example of sparse representation and Fields of experts (FoE) as an example of responses to the specific filter bank, we demonstrate the inherent relationship between them. The filters of FoE stand for the components that fire rarely on natural images, while the redundant dictionary of K-SVD depicts the primary components to some extent. They are two complementary pursuits of natural images space. We further bridge the gap between these two methods by proposing a method to get adaptive filters for FoE from the redundant dictionary of K-SVD. Instead of pursuing the state-of-the-art performance, our research gives a suggestive and unique view point from the essence of natural image space pursuit. Feng Jiang 0001, Xulin Wang, Debin Zhao |
ICIP | 1 |
| 2013 | Spatially directional predictive coding for block-based compressive sensing of natural imagesabstractA novel coding strategy for block-based compressive sensing named spatially directional predictive coding (SDPC) is proposed, which efficiently utilizes the intrinsic spatial correlation of natural images. At the encoder, for each block of compressive sensing (CS) measurements, the optimal prediction is selected from a set of prediction candidates that are generated by four designed directional predictive modes. Then, the resulting residual is processed by scalar quantization (SQ). At the decoder, the same prediction is added onto the de-quantized residuals to produce the quantized CS measurements, which is exploited for CS reconstruction. Experimental results substantiate significant improvements achieved by SDPC-plus-SQ in rate distortion performance as compared with SQ alone and DPCM-plus-SQ. Jian Zhang 0018, Debin Zhao, Feng Jiang 0001 |
ICIP | 3 |
| 2013 | Multi-scale face hallucination based on frequency bands analysisabstractIn this paper, a multi-scale face hallucination method is proposed to produce high-resolution (HR) face images from low-resolution (LR) ones according to the specific face characteristics and priors based on frequency bands analysis. In the first scale, the middle-resolution (MR) images are generated based on a patch-based learning method in DCT domain. In this scale, the DC coefficients and AC coefficients are estimated separately. In the second scale, a DCT upsampling for low frequency band restoration and a high frequency band restoration are combined to generate the final high-resolution face images. Extensive experiments show that the proposed algorithm achieves significant improvement. Xiaodan Du 0003, Feng Jiang 0001, Debin Zhao |
VCIP | 2 |
| 2013 | Natural images scale invariance and high-fidelity image restorationabstractOne of the most striking properties of natural image statistics is their scale invariance. Intuitively, a natural image always contains the same contents of different scales and dually the same contents of same scale exist throughout scales of the image. Different from the previous scale invariance related work decomposing an image to its local band-pass filter components, this paper seeks a general model of the natural image paths distribution to describe the scale invariance in the visual world and then a novel strategy for high-fidelity image restoration is presented by characterizing nonlocal self-similarity of natural images throughout scales in a unified statistical manner, which offers a powerful mechanism of combining natural images scale invariance and nonlocal self-similarity simultaneously to ensure a more reliable and robust estimation. Extensive experiments on image restoration from partial random samples manifest that the proposed algorithm achieves significant performance improvements over the current state-of-the-art schemes. Feng Jiang 0001, Shaohui Liu, Debin Zhao |
VCIP | 2 |
| 2013 | An improved image compression scheme with an adaptive parameters set in encrypted domainabstractA growing societal awareness about privacy and security push the development of signal processing techniques in the encrypted domain. Data compression in encrypted domain attracts much attention recently years due to its avoiding the leakage of data source during compression. This paper proposes an improved block-by-block compression scheme of encrypted image with flexible compression ratio. The original image is encrypted by permuting the blocks of the image and then permuting the pixels in the blocks. In the compression, pixels chosen randomly used as reference information, and remaining pixels are compressed by coset code. At the decoder side, side information (SI) which is generated by combining correlation among blocks and image restoration from partial random samples (IRPRS) is utilized to assist the decompression. Moreover, an adaptive system parameters selection method is also given in this paper. The experimental results show that the proposed method can achieve a better reconstructed result compared with the earlier method. Guochao Zhang, Shaohui Liu, Feng Jiang 0001, Debin Zhao, Wen Gao 0001 |
VCIP | 3 |
| 2012 | High-quality image interpolation via local autoregressive and nonlocal 3-D sparse regularizationabstractIn this paper, we propose a novel image interpolation algorithm, which is formulated via combining both the local autoregressive (AR) model and the nonlocal adaptive 3-D sparse model as regularized constraints under the regularization framework. Estimating the high-resolution image by the local AR regularization is different from these conventional AR models, which weighted calculates the interpolation coefficients without considering the rough structural similarity between the low-resolution (LR) and high-resolution (HR) images. Then the nonlocal adaptive 3-D sparse model is formulated to regularize the interpolated HR image, which provides a way to modify these pixels with the problem of numerical stability caused by AR model. In addition, a new Split-Bregman based iterative algorithm is developed to solve the above optimization problem iteratively. Experiment results demonstrate that the proposed algorithm achieves significant performance improvements over the traditional algorithms in terms of both objective quality and visual perception. Xinwei Gao, Jian Zhang 0018, Feng Jiang 0001, Xiaopeng Fan 0001, Siwei Ma 0001, Debin Zhao |
VCIP | 3 |
| 2012 | Viewpoint-independent hand gesture recognition systemabstractIn this paper, we creatively present a viewpoint-free hand gesture recognition system based on Kinect sensor. Through depth image, we build Point Clouds of user. Then, we estimate the current optimal viewpoint, i.e., the front, and project Point Clouds to that direction. Through that process we in great extent overcome the viewpoint-dependency issue. To match hand types, we propose an improved shape context to describe each hand gesture and use the Hungarian algorithm to calculate match degree. Our method is quite straightforward, however the experimental results prove that by this means gestures can be recognized independent of viewpoints with great accuracy. Besides, it is fast and robust, thus can be applied under various realistic scenarios in realtime. Feng Jiang 0001, Debin Zhao, Shaohui Liu, Wen Gao 0001 |
VCIP | 2 |
| 2011 | Saliency Detection: A Self-Adaption Sparse Representation ApproachabstractSaliency detection is essential to visual attention modelling and various computer vision tasks. Representation and measurement are two important issues for saliency models. Good representation and reasonable measurement are both critical issues in modelling visual saliency mechanism. For every input image, we obtain a self-adaptive dictionary that describes the image content effectively and image prior that forces sparsity in every location in the image using the K-SVD algorithm. For saliency measurement, background firing rate (BFR) is defined for each sparse features and it is followed by feature activation rate (FAR) computation to measure the bottom-up visual saliency. Gaoxiang Zhang, Feng Jiang 0001, Debin Zhao, Xiaoshuai Sun, Shaohui Liu |
ICIG | 2 |
| 2010 | An Image Data Hiding Method Using Pixel-Based JND Model
Shaohui Liu, Feng Jiang 0001, Hongxun Yao, Debin Zhao |
ICIC (3) | 2 |
| 2010 | Compressed image restoration based on edge enhancement field of expertsabstractImages encoded at low bit rate may suffer from blocking artifacts, which dramatically degrade the visual quality. In this paper, we propose a postprocessing image restoration scheme for JPEG compressed images based on the maximum a posteriori criterion. A degradation model, represent by additive Gaussian noise model, is proposed to simulate JPEG compression process, while the original image is modeled as a high order Markov random field (MRF) based on the fields of experts framework. Meanwhile, an enhancement algorithm is proposed to restore the high frequency (HF) components. Experiment results have demonstrated that the proposed scheme can reproduce higher-quality images in terms of both objective and subjective quality. Feng Jiang 0001, Debin Zhao |
VCIP | 2 |
| 2009 | Synthetic data generation technique in Signer-independent sign language recognition
Feng Jiang 0001, Wen Gao 0001, Hongxun Yao, Debin Zhao, Xilin Chen 0001 |
Pattern Recognit. Lett. | 1 |
| 2005 | Static Gesture Quantization and DCT Based Sign Language Generation
Feng Jiang 0001, Hongxun Yao, Guilin Yao, Wen Gao 0001 |
ACII | 2 |
| 2004 | Based on HMM and SVM multilayer architecture classifier for Chinese sign language recognition with large vocabularyabstractThis paper has put forward a new architecture classifier method for Chinese sign language recognition (CSLR) to improve the performance of recognition. It is a signer-independent method, to recognize Chinese sign language with large vocabulary using multilayer architecture classifier and making use of the advantages both of HMM (hidden Markov model) and SVM (support vector machines). Because HMM is good at dealing with sequential inputs, while SVM shows superior performance in classifying with good generalization properties especially for limited samples. Therefore, they can be combined to yield a better and effective multilayer architecture classifier. We apply SVMs to resolve the uncertainties of the remaining which are in confusable sets after the first-stage HMM-based recognizer. And the confusable sets would be updated dynamically according to the results of a recognition performance to optimize the discernment performance next time. Experimental results proved that it is an effective method for CSLR with large vocabulary keywords sign language recognition, HMM, SVM, multilayer architecture classifier. Jianjun Ye, Hongxun Yao, Feng Jiang 0001 |
ICIG | 3 |
| 2004 | Multilayer architecture in sign language recognition systemabstractUp to now analytical or statistical methods have been used in sign language recognition with large vocabulary. Analytical methods such as Dynamic Time Wrapping (DTW) or Euclidian distance have been used for isolated word recognition, but the performance is not satisfactory enough because it is easily interfered by noise. Statistical methods, especially hidden Markov Models are commonly used, for both continuous sign language and isolated words and with the expansion of vocabulary the processing time becomes increasingly unacceptable. Therefore, a multilayer architecture of sign language recognition for large vocabulary is proposed in this paper for the purpose of speeding up the recognition process. In this method the gesture sequence to be recognized is first located at a set of words that are easy to be confused (confusion set) through a global cursory search and then the gesture is recognized through a latter local search and the generation of confusion set is realized by DTW/ISODATA algorithm. Experiment results indicate that it is an effective algorithm for Chinese sign language recognition. Feng Jiang 0001, Hongxun Yao, Guilin Yao |
ICMI | 1 |