Jing Li 0046

dblp:181/2820-46 · DBLP profile ↗
← Back
36ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0002-9132-6684ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A comprehensive survey of deep learning-based cognitive diagnosis models in education: Methods, applications, and outlook
Jing Li 0046, Yang Xu 0025, Enrique Herrera-Viedma, Hui Yu 0001, Jiande Sun 0001
Neurocomputing2
2025 Orientation-Aware Reversible Data Hiding With Brainstorming Optimization for UAV Aerial Images
abstract
In recent years, with the rapid development of unmanned aerial vehicle (UAV), aerial images have extended across various industries such as intelligent building, agriculture, transportation, and Industry 4.0. Notably, the security of UAV‐assisted data acquisition during transmission has become a critical concern. The reversible data hiding (RDH) method can hide data in aerial images for transmission and ensure secure communication. In general, an aerial image may exhibit substantially different orientation regularity from a natural scene image. This casts major challenges to the RDH method, for which existing approaches lack effective mechanisms to capture such content type variations, and thus are difficult to generalize from one type to another. In this paper, the orientation‐aware selectivity mechanism is introduced to achieve an accurate orientation‐aware prediction along different directions in local regions with different structure regularity. Furthermore, we propose a progressive brainstorming optimization algorithm (BSO)‐guided optimal PSNR value strategy, which can obtain a superior perceptual performance and the corresponding thresholds by further exploring the pixel correlations within the UAV aerial images. Experimental results on the USC‐SIPI Miscellaneous dataset and two challenging aerial datasets, including the USC‐SIPI High Altitude Aerial Imagery dataset and the Kaggle dataset, demonstrate that the proposed framework enhances the imperceptibility powerfully in marked UAV aerial images and ensures sufficient embedding capacity effectively. The average PSNR of the marked image obtained by the proposed method is 63.85 dB when embedded with 30,000 bits of data, which is an improvement of 0.59 dB compared to the current state‐of‐the‐art RDH methods.
Xiaodan Tai, Yannan Ren, Jing Li 0046, Jiande Sun 0001, Kai Zhang 0010, Wenbo Wan
Int. J. Intell. Syst.3
2025 Toward Practical Colorectal Cancer Diagnosis: A Bowel-Sound-Based System With Portable Sensor and On-Board Lightweight AI Model
abstract
Colorectal Cancer (CRC) is one of the leading causes of cancer-related deaths worldwide, and early screening plays a crucial role in improving patient outcomes. In this study, we present a novel AI-assisted CRC diagnostic system using Bowel Sound (BS) signals. We first develop two portable BS acquisition devices with distinct form factors for high-fidelity signal capture in both clinical and home-care scenarios. A total of 221 recordings were collected under expert-guided protocol, with 144 CRC recordings and 59 Non-CRC healthy controls using the developed device. To enable low-resource deployment, we design a lightweight deep learning model optimized for real-time, on-board inference. The model incorporates multiple training strategies, including transfer learning on a large-scale public BS dataset, self-supervised temporal feature learning, and a hybrid semi-and weakly-supervised approach that leverages both unlabeled and real-noise data. Furthermore, a Sound Event Detection (SED) attention mechanism and iterative consistency learning are introduced to enhance the model’s sensitivity to BS activity. The proposed model comprises only 264.7 K parameters and 253.2 M Floating-Point Operations (FLOPs), requiring 1.57 MB of RAM and 1.03 MB of FLASH when deployed on microcontroller. It performs inference in approximately 3.4 s with low power consumption, making it well-suited for low-resource environments. Despite its compact design, the model achieves 93.06% classification accuracy, 96.46% sensitivity, and 86.99% specificity for binary-classes in CRC diagnosis. These results demonstrate the system’s potential for accessible and cost-effective CRC screening in community, home, and rural healthcare scenarios.
Fuze Tian, Yang Tan 0003, Enze Li, Jiedong Ma, Jingyu Liu 0002, Kun Qian 0003, Jing Li 0046, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller
IEEE Internet Things J.9
2025 Dual Prototypes-Based Personalized Federated Adversarial Cross-Modal Hashing
abstract
With the rapid advances in wireless communication and IoT platforms, it is increasingly difficult to analyze relevant multi-modal data distributed across geographically diverse and heterogeneous platforms. One promising approach is to rely on federated learning to build compact cross-modal hash codes. However, existing federated learning methods easily exhibit degenerative performance in the global model due to the distributed data being derived from diverse domains. In addition, directly forcing each client to adopt the same global parameters as local parameters, without effective local training, significantly reduces the performance of each client. To overcome these challenges, we propose a novel federated adversarial cross-modal hashing, called Dual Prototypes-based personalized Federated Adversarial (DP-FeAd), which provides iterated training of shared dual prototypes. Specifically, aiming to expand local hashing models beyond their knowledge realms, DP-FeAd enables participating clients to engage in cooperative learning through two constructions: cluster prototypes and unbiased prototypes, instead of the traditional global prototypes, ensuring both generalization and stability. Specifically, the cluster prototypes are derived from local class-level prototypes and adversarially trained with local approximate hash codes to align their distributions. The unbiased prototypes are averaged from cluster prototypes and integrated into the training of local hashing models to maintain consistency across different local class-level prototypes further. The experiments conducted on two benchmark datasets demonstrate that our proposed method significantly enhances the performance of deep cross-modal hashing models in both IID (Independent and Identically Distributed) and non-IID scenarios.
Lingchen Gu, Xiaojuan Shen, Jiande Sun 0001, Jing Li 0046, Zhihui Li 0001, Sen-Ching S. Cheung, Wenbo Wan
IEEE Trans. Circuits Syst. Video Technol.5
2024 Reversible Adversarial Examples based on Self-Embedding Watermark for Image Privacy Protection
abstract
Personal images shared online are susceptible to malicious collection and misuse, posing significant privacy risks. Reversible Adversarial Example (RAE) provides an effective safeguard that prevents the analysis of unauthorized models without affecting authorized models. However, the vulnerability of image content extends beyond illegal analysis to deliberate tampering. Existing RAE methods struggle to integrate with tamper defense techniques, and exhibit limited attack ability due to the trade-off between perturbation strength (i.e. attack ability) and recoverability. To this end, we propose a novel approach for generating reversible adversarial examples with self-embedding watermarks (W-RAE). Specifically, we convert the generation of RAEs into a deep steganography task and decouple the constraints between perturbation strength and recoverability to enhance RAEs’ attack performance. Additionally, by embedding crafted self-embedding watermarks during the RAE construction process, our method supports both data access control and tamper defense, thereby protecting image privacy from multiple perspectives. Extensive experiments have demonstrated its effectiveness as a privacy-preserving mechanism.
Jinghui Yin, Xuejun Cheng, Jing Li 0046, Guanghui Luo
IJCNN5
2024 Deep image watermarking with loss-driven modification
Wenqing Yang, Jing Li 0046, Jiande Sun 0001, Wenbo Wan
Multim. Tools Appl.5
2024 Class-Specific Thresholding for Imbalanced Semi-Supervised Learning
abstract
Semi-supervised learning (SSL) has emerged as a powerful technique to mitigate the scarcity of labeled data. However, the effectiveness of most SSL methods relies on the assumption of a balanced class distribution, which often proves unrealistic, especially in medical imaging scenarios. To address this challenge, we propose Class-Specific Thresholding (CST) for imbalanced SSL. Specifically, CST dynamically integrates the model's learning status, class-specific learning effects, and data class distribution to estimate confidence thresholds for each class. These thresholds enhance the reliability of selected unlabeled data, particularly for minority classes. Additionally, we introduce a class-sensitive unsupervised loss function that further intensifies the model's focus on minority classes at minimal computational cost by leveraging predicted class distributions from previous inputs. Extensive experiments on highly imbalanced skin disease and endoscopy image datasets demonstrate that CST significantly outperforms state-of-the-art SSL methods, particularly in improving the accuracy for minority classes. These results validate the effectiveness of CST in addressing class imbalance.
Aixi Qu, Qiang Wu 0009, Luyue Yu, Jing Li 0046
IEEE Signal Process. Lett.4
2024 TNCB: Tri-Net With Cross-Balanced Pseudo Supervision for Class Imbalanced Medical Image Classification
abstract
In clinical settings, the implementation of deep neural networks is impeded by the prevalent problems of label scarcity and class imbalance in medical images. To mitigate the need for labeled data, semi-supervised learning (SSL) has gained traction. However, existing SSL schemes exhibit certain limitations. 1) They commonly fail to address the class imbalance problem. Training with imbalanced data makes the model's prediction biased towards majority classes, consequently introducing prediction bias. 2) They usually suffer from training bias arising from unreasonable training strategies, such as strong coupling between the generation and utilization of pseudo labels. To address these problems, we propose a novel SSL framework called Tri-Net with Cross-Balanced pseudo supervision (TNCB). Specifically, two student networks focusing on different learning tasks and a teacher network equipped with an adaptive balancer are designed. This design enables the teacher model to pay more focus on minority classes, thereby reducing prediction bias. Additionally, we propose a virtual optimization strategy to further enhance the teacher model's resistance to class imbalance. Finally, to fully exploit valuable knowledge from unlabeled images, we employ cross-balanced pseudo supervision, where an adaptive cross loss function is introduced to reduce training bias. Extensive evaluation on four datasets with different diseases, image modalities, and imbalance ratios consistently demonstrate the superior performance of TNCB over state-of-the-art SSL methods. These results indicate the effectiveness and robustness of TNCB in addressing imbalanced medical image classification challenges.
Aixi Qu, Qiang Wu 0009, Jing Wang 0215, Luyue Yu, Jing Li 0046
IEEE J. Biomed. Health Informatics5
2023 A cover selection-based reversible data hiding method by learning cross-modal hashing
Liming Zou, Jiande Sun 0001, Wenbo Wan, Jing Li 0046, Q. M. Jonathan Wu
Multim. Tools Appl.4
2023 TSINIT: A Two-Stage Inpainting Network for Incomplete Text
abstract
Although there are lots of studies on scene text recognition, few of them focus on the recognition of the incomplete text. The recognition performance of existing text recognition algorithms on the incomplete text is far from the expected, and the recognition of the incomplete text is still challenging. In this paper, an end-to-end Two-Stage Inpainting Network for Incomplete Text (TSINIT) is proposed to reconstruct the incomplete text into the complete one even when the text is in various styles and with various backgrounds, and the reconstructed text can be recognized by the existing text recognition algorithms correctly. The proposed TSINIT is divided into text extraction module (TEM) and text reconstruction module (TRM) to make the inpainting only focus on the text. TEM separates the incomplete text from the background and character-like regions at the pixel level, which can reduce the ambiguity of text reconstruction caused by the background. TRM reconstructs the incomplete text towards the most possible text with the consideration of the abstract and semantic structures of the text. Furthermore, we build a synthetic incomplete text dataset (SITD), which contains contaminated and abraded text images. SITD is divided into 6 incomplete levels according to the number of pixels in the incomplete regions and the ratio of the incomplete characters to all characters. The experimental results show that the proposed method has better inpainting ability for the incomplete text compared with traditional image inpainting algorithms on the proposed SITD and real images. When using the same text recognition method, the recognition accuracy of the incomplete text on SITD can be improved much more with the help of the proposed TSINIT than with the traditional image inpainting methods.
Jiande Sun 0001, Fanfu Xue, Jing Li 0046, Lei Zhu 0002, Huaxiang Zhang 0001, Jia Zhang 0028
IEEE Trans. Multim.3
2023 Robust Coverless Image Steganography Based on Neglected Coverless Image Dataset Construction
abstract
Most of the existing image selection-based coverless image steganography methods mainly focus on improving the capacity and robustness under the assumption that the corresponding dataset is available. But they ignore how to successfully construct the coverless image dataset, which is the foundation of such methods and has a critical impact on the capacity. In this paper, a coverless image steganography is proposed that considers how to efficiently construct the coverless image dataset. In the proposed method, the CNN-based deep hash is extracted from the image and a specific mapping rule is designed to map the high-dimensional deep hash to the low-dimensional secret message. In addition, an unsupervised clustering algorithm is adopted to construct the coverless image dataset, which makes the construction of the coverless image dataset efficient and improves the robustness of the proposed steganography method. To our best knowledge, this is the first attempt to improve the construction efficiency of the coverless image dataset in the field of coverless image steganography. Experimental results show that the construction of a large coverless image dataset is feasible and reliable, and the proposed method has better robustness and higher dataset utilization rate compared with the state-of-the-art methods.
Liming Zou, Jing Li 0046, Wenbo Wan, Q. M. Jonathan Wu, Jiande Sun 0001
IEEE Trans. Multim.2
2022 TAGAN: Texture and Attention Guided Generative Adversarial Network for Image Super Resolution
abstract
Super Resolution (SR) methods based on Generative Adversarial Networks (GANs) accomplish predominant execution in visual perception and image quality. These methods are mainly generated by traditional Peak-Signal-to-Noise-Ratio (PSNR)oriented or perceptual-driven. As the reconstruction process usually loses high frequency information, various methods aim to preserve more details. To make the details of the generated image richer, the Gradient Weight (GW) loss is introduced in the proposed method, because the gradient can reflect the texture of the image to a certain extent. The GW loss function is helpful to improve the edge and detailed texture of the generated image. Furthermore, we introduce attention mechanism to the image reconstruction block via Squeeze and Excitation Net (SENet). Attention mechanism can effectively aggregate the global features obtained by the nonlinear mapping network, and improve the channel sensitivity of the model. With the help of GW and attention mechanism, the proposed method can achieve better performance and visual quality in image texture detail restoration. The performance comparison between the state-of-the-art methods and our proposed method verifies the feasibility and reliability of the proposed method.
Haitao Wang 0023, Jiande Sun 0001, Wenxiu Diao, Jing Li 0046, Kai Zhang 0010
ISCAS4
2022 INIT: Inpainting Network for Incomplete Text
abstract
In recent years, scene text recognition algorithms have achieved great progress, but they still face some challenges in practical environment, such as the incomplete scene text, which includes structurally broken characters and occluded characters, as shown in Fig. 1. Existing scene text recognition algorithms can not accurately recognize such incomplete scene text. In this paper, we design an end-to-end Inpainting Network for Incomplete Text (INIT), which can reconstruct each incomplete character into complete character. INIT can separate the text from the background and just reconstruct the text regions, which reduces the influence caused by the background. And INIT is supervised by reconstruction loss and semantic loss to reach the most likely text. Furthermore, to compensate for the absence of incomplete scene text dataset, we propose an incomplete text synthesis method and build an incomplete text dataset (SITD), which is more suitable for practical scenarios. On SITD, the recognition accuracy can achieve 82.99% by the existing text recognition method with the help of INIT, while the accuracy is only 61.99% by the same recognition method with the conventional image inpainting method. Experimental results show that the proposed method has better reconstruction ability for incomplete text compared with the existing image inpainting algorithms.
Fanfu Xue, Jia Zhang 0028, Jiande Sun 0001, Jinghui Yin, Liming Zou, Jing Li 0046
ISCAS6
2022 A comprehensive survey on robust image watermarking
Wenbo Wan, Jun Wang 0061, Yunming Zhang, Jing Li 0046, Hui Yu 0001, Jiande Sun 0001
Neurocomputing4
2022 Discrete Fusion Adversarial Hashing for cross-modal retrieval
Jing Li 0046, En Yu, Xiaojun Chang, Huaxiang Zhang 0001, Jiande Sun 0001
Knowl. Based Syst.1
2022 Unsupervised Change Detection of Multispectral Images Based on PCA and Low-Rank Prior
abstract
In this letter, we propose a new unsupervised change detection method based on low-rank prior for multispectral images. It is assumed that the changed and unchanged pixels are from different subspaces due to different appearance and statistical properties. So, low-rank representation (LRR) is employed to find informative pixels from the superpixels of the difference image (DI). Besides, taking the sparsity of changed pixels in the observed scenes into consideration, the selection rule is designed to distinguish these pixels. Then, principal component analysis (PCA) is used for the training of changed and unchanged dictionaries from these pixels. Finally, the change map is estimated by comparing the reconstruction error of each pixel in DI on changed and unchanged dictionaries. By LRR, more representative pixels are found for subsequent dictionary learning, which can efficiently improve the performance of the proposed method. Experiments on multitemporal images from the Landsat satellite demonstrate the effectiveness of the proposed method.
Jing Li 0046, Feng Zhang 0028, Jiande Sun 0001, Kai Zhang 0010
IEEE Geosci. Remote. Sens. Lett.2
2021 Visual Security Assessment via Saliency-Weighted Structure and Orientation Similarity for Selective Encrypted Images
abstract
Selective encryption has been widely used in image privacy protection. Visual security assessment is necessary for the effectiveness and practicability of image encryption methods, and there have been a series of research studies on this aspect. However, these methods do not take into account perceptual factors. In this paper, we propose a new visual security assessment (VSA) by saliency-weighted structure and orientation similarity. Considering that the human visual perception is sensitive to the characteristics of selective encrypted images, we extract the structure and orientation feature maps, and then similarity measurements are conducted on these feature maps to generate the structure and orientation similarity maps. Next, we compute the saliency map of the original image. Then, a simple saliency-based pooling strategy is subsequently used to combine these measurements and generate the final visual security score. Extensive experiments are conducted on two public encryption databases, and the results demonstrate the superiority and robustness of our proposed VSA compared with the existing most advanced work.
Zhengguo Wu, Kai Zhang 0010, Yannan Ren, Jing Li 0046, Jiande Sun 0001, Wenbo Wan
Secur. Commun. Networks4
2021 Eye-based Recognition for User Identification on Mobile Devices
abstract
User identification is becoming more and more important for Apps on mobile devices. However, the identity recognition based on eyes, e.g., iris recognition, is rarely used on mobile devices comparing with those based on face and fingerprint due to its extra cost in hardware and complicated operations during recognition. In this article, an eye-based recognition method is designed for identity recognition on mobile devices, which can be implemented just like face recognition. In the proposed method, the eye feature is composed of the static and dynamic features, where the periocular feature extracted by deep neural network from the eye image is used as the static feature, and the motion feature of saccadic velocity is selected as the dynamic feature. The eye images can be captured by the normal camera on mobile devices just like faces, and dynamic features can provide living information to increase the difficulty of forgery. The GazeCapture dataset is used to test the proposed method, because the eye images in this dataset are captured by mobile devices during daily use. The recognition accuracy of the proposed method on the GazeCapture dataset can reach 96.87% only based on the periocular feature and can be enhanced to 97.99% when it is fused with the saccadic feature. The experiment results show that the performance of the proposed method can be comparative to that of iris recognition methods. It demonstrates that the proposed method is a practical reference for the eye-based identity recognition, and the proposed method provides one more biometric choice for mobile devices.
Huiru Shao, Jing Li 0046, Jia Zhang 0028, Hui Yu 0001, Jiande Sun 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2020 Image super-resolution based on two-level residual learning CNN
Min Gao 0001, Xianhua Han, Jing Li 0046, Huaxiang Zhang 0001, Jiande Sun 0001
Multim. Tools Appl.3
2020 Two-class 3D-CNN classifiers combination for video copy detection
Jing Li 0046, Huaxiang Zhang 0001, Wenbo Wan, Jiande Sun 0001
Multim. Tools Appl.1
2020 Hash length: a neglected element
Haifeng Qi, Jing Li 0046, Qiang Wu 0009, Wenbo Wan, Jiande Sun 0001
Multim. Tools Appl.2
2020 Saccadic trajectory-based identity authentication
Huiru Shao, Jing Li 0046, Wenbo Wan, Huaxiang Zhang 0001, Jiande Sun 0001
Multim. Tools Appl.2
2020 Hybrid JND model-guided watermarking method for screen content images
Wenbo Wan, Jun Wang 0061, Jing Li 0046, Jiande Sun 0001, Huaxiang Zhang 0001
Multim. Tools Appl.3
2020 Pattern complexity-based JND estimation for quantization watermarking
Wenbo Wan, Jun Wang 0061, Jing Li 0046, Lili Meng, Jiande Sun 0001, Huaxiang Zhang 0001
Pattern Recognit. Lett.3
2020 Tensor-based sparse representations of multi-phase medical images for classification of focal liver lesions
Jian Wang 0004, Jing Li 0046, Xianhua Han, Lanfen Lin, Hongjie Hu, Qingqing Chen 0001, Yutaro Iwamoto, Yen-Wei Chen 0001
Pattern Recognit. Lett.2
2020 Multi-class joint subspace learning for cross-modal retrieval
En Yu, Jing Li 0046, Li Wang 0148, Jia Zhang 0028, Wenbo Wan, Jiande Sun 0001
Pattern Recognit. Lett.2
2019 Hierarchical prediction based on two-level Gaussian mixture model clustering for bike-sharing system
Wenzhen Jia, Yanyan Tan, Li Liu 0031, Jing Li 0046, Huaxiang Zhang 0001, Kai Zhao 0011
Knowl. Based Syst.4
2019 Adaptive Semi-Supervised Feature Selection for Cross-Modal Retrieval
abstract
In order to exploit the abundant potential information of the unlabeled data and contribute to analyzing the correlation among heterogeneous data, we propose the semi-supervised model named adaptive semi-supervised feature selection for cross-modal retrieval. First, we utilize the semantic regression to strengthen the neighboring relationship between the data with the same semantic. And the correlation between heterogeneous data can be optimized via keeping the pairwise closeness when learning the common latent space. Second, we adopt the graph-based constraint to predict accurate labels for unlabeled data, and it can also keep the geometric structure consistency between the label space and the feature space of heterogeneous data in the common latent space. Finally, an efficient joint optimization algorithm is proposed to update the mapping matrices and the label matrix for unlabeled data simultaneously and iteratively. It makes samples from different classes to be far apart, while the samples from same class lie as close as possible. Meanwhile, the l2,1-norm constraint is used for feature selection and outlier reduction when the mapping matrices are learned. In addition, we propose learning different mapping matrices corresponding to different sub-tasks to emphasize the semantic and structural information of query data. Experiment results on three datasets demonstrate that our method performs better than the state-of-the-art methods.
En Yu, Jiande Sun 0001, Jing Li 0046, Xiaojun Chang, Xianhua Han, Alex Hauptmann 0001
IEEE Trans. Multim.3
2018 View-invariant gait recognition based on kinect skeleton feature
Jiande Sun 0001, Jing Li 0046, Wenbo Wan, De Cheng, Huaxiang Zhang 0001
Multim. Tools Appl.3
2017 Appearance-based gaze block estimation via CNN classification
abstract
Appearance-based gaze estimation methods have received increasing attention in the field of human-computer interaction (HCI). These methods tried to estimate the accurate gaze point via Convolutional Neural Network (CNN) model, but the estimated accuracy can't reach the requirement of gaze-based HCI when the regression model is used in the output layer of CNN. Given the popularity of button-touch-based interaction, we propose an appearance-based gaze block estimation method, which aims to estimate the gaze block, not the gaze point. In the proposed method, we relax the estimation from point to block, so that the gaze block can be estimated by CNN-based classification instead of the previous regression model. We divide the screen into square blocks to imitate the button-touch interface, and build an eye-image dataset, which contains the eye images labelled by their corresponding gaze blocks on the screen. We train the CNN model according to this dataset to estimate the gaze block by classifying the eye images. The experiments on 6- and 54-block classifications demonstrate that the proposed method has high accuracy in gaze block estimation without any calibration, and it is promising in button-touch-based interaction.
Xuemei Wu, Jing Li 0046, Qiang Wu 0009, Jiande Sun 0001
MMSP2
2017 Block-Wise Gaze Estimation Based on Binocular Images
Xuemei Wu, Jing Li 0046, Qiang Wu 0009, Jiande Sun 0001
PSIVT2
2016 Depth propagation in 2D-to-3D conversion based on frame clustering
abstract
Depth propagation plays an important role to improve the quality of converted video in 2D-to-3D conversion. During depth propagation, the propagation path is usually along the time. In this paper, a new depth propagation algorithm is proposed to improve the quality of depth propagation by changing the propagation path according to frame clustering. The frames are clustered via K-means clustering based on HSV color histogram. The frame at the center of each cluster is selected as the keyframe of the cluster and the other frames are taken as the non-keyframes attaching to the keyframe. The depth information of keyframe is propagated to the non-keyframes within the same cluster directly. The performance of the proposed depth propagation algorithm is evaluated by MSE of the propagated depth map comparing with several existing algorithms. The comparison results show that the depth propagation errors can be reduced a lot by the proposed clustering based propagation.
Zhenxiao Fu, Jiande Sun 0001, Qiang Wu 0009, Jing Li 0046
ICASSP4
2016 Gait recognition based on 3D skeleton joints captured by kinect
abstract
2D-video-based gait recognition techniques have been studied for decades, but there are still many challenges, one of which is the robustness against the variation of view angle. In this paper, the second generation Kinect (Kinect V2) is used as a tool to establish a 3D-skeleton-based gait database, which includes both 3D information of the skeleton joints and the corresponding 2D silhouette images captured by Kinect V2. Based on this dataset, a human walking model is built, and the static and dynamic features are extracted, which are verified to be view-invariant for gait recognition. Referring to the walking model, the gait recognition abilities for the static and dynamic features are investigated respectively and a gait recognition scheme based on the matching-level-fusion of the static and dynamic features is proposed, in which the recognition is achieved by the nearest neighbor classification method. Experiments show that the proposed scheme has robust recognition performance against the variation of view angle.
Jiande Sun 0001, Jing Li 0046, Dong Zhao 0017
ICIP3
2016 Video hashing based on appearance and attention features fusion via DBN
Jiande Sun 0001, Xiaocui Liu, Wenbo Wan, Jing Li 0046, Dong Zhao 0017, Huaxiang Zhang 0001
Neurocomputing4
2015 Investigation on the Influence of Visual Attention on Image Memorability
Wulin Wang, Jiande Sun 0001, Jing Li 0046, Qiang Wu 0009
ICIG (3)3
2015 A calibration simplified method for gaze interaction based on using experience
abstract
The step of calibration prevents the gaze interaction from interacting naturally, which usually needs five or more calibration points to obtain the user's calibration information. In this paper, a calibration simplified method is proposed, which is based on user experience and consists of the user experience accumulation stage and the calibration simplified stage. In the user experience accumulation stage, the calibration is still needed as usual before each gaze interaction and the calibration parameters and the position of the user are stored as the calibration experience, where the calibration parameters include the actual coordinates of the calibration points and their estimation errors. In the calibration simplified stage, the user needs to fixate on only one calibration point for calibration before the calibration information can be estimated according to the user experience stored in the first stage. Even the calibration information can be estimated according to only the position of the user and the stored calibration parameters, which means the user can interact with computer via gaze directly without calibration. The simulations show that the gaze estimation accuracy of the proposed calibration simplified method can be 1.4047° with one calibration point and 1.7489° without calibration point, which are much higher than that without the user experience.
Cong Niu, Jiande Sun 0001, Jing Li 0046
MMSP3