Ping Li 0006

dblp:62/5860-6 · also Patrick L. Lee · DBLP profile ↗
← Back
53ranked-venue papers
36as first author
27since 2021 · last 2026
0000-0002-8515-7773ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 20 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 8 first-author · 5 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Temporal prompt guided visual-text-object alignment for zero-shot video captioning
Ping Li 0006, Zeyu Pan
Comput. Vis. Image Underst.1
2026 Class semantics guided knowledge distillation for few-shot class incremental learning
Ping Li 0006, Shaoqi Tian
Inf. Sci.1
2026 Adaptive saliency based contextual metric learning for few-shot open-set recognition
Ping Li 0006, Lijie Shang, Chenhao Ping
Pattern Recognit.1
2026 Triple-View Knowledge Distillation for Semi-Supervised Semantic Segmentation
abstract
To alleviate the expensive human labeling problem, semi-supervised semantic segmentation utilizes a few labeled images along with an abundance of unlabeled images to predict the pixel-level label maps with the same size. Previous methods often rely on co-training with two convolutional networks with the same architecture but different initialization, which fails to capture sufficiently diverse features. This limitation motivates us to employ tri-training and design a triple-view encoder to utilize encoders with different architectures to derive diverse features, while leveraging knowledge distillation to capture complementary semantics among these encoders. Moreover, existing approaches simply concatenate features from both encoder and decoder, and the simple concatenation requires a large memory cost. This inspires us to present a dual-frequency decoder that selects those important features by projecting the spatial-domain features into the frequency domain, where a dual-frequency channel attention mechanism is applied to evaluate the feature importance. Therefore, we propose a Triple-view Knowledge Distillation framework, termed TriKD, for semi-supervised semantic segmentation. It comprises the triple-view encoder and the dual-frequency decoder. Extensive experiments conducted on two benchmarks,i.e., Pascal VOC 2012 and Cityscapes, validate the superiority of our method, achieving a satisfying tradeoff between precision and inference speed. Our code is available at GitHub.
Ping Li 0006, Li Yuan 0007, Xianghua Xu, Mingli Song
IEEE Trans. Multim.1
2025 Sample-level Adaptive Knowledge Distillation for Action Recognition
abstract
Knowledge Distillation (KD) compresses neural networks by learning a small network (student) via transferring knowledge from a pre-trained large network (teacher). Many endeavours have been devoted to the image domain, while few works focus on video analysis which desires training much larger model making it be hardly deployed in resource-limited devices. However, traditional methods neglect two important problems, i.e., 1) Since the capacity gap between the teacher and the student exists, some knowledge w.r.t. difficult-to-transfer samples cannot be correctly transferred, or even badly affects the final performance of student, and 2) As training progresses, difficult-to-transfer samples may become easier to learn, and vice versa. To alleviate the two problems, we propose a Sample-level Adaptive Knowledge Distillation (SAKD) framework for action recognition. In particular, it mainly consists of the sample distillation difficulty evaluation module and the sample adaptive distillation module. The former applies the temporal interruption to frames, i.e., randomly dropout or shuffle the frames during training, which increases the learning difficulty of samples during distillation, so as to better discriminate their distillation difficulty. The latter module adaptively adjusts distillation ratio at sample level, such that KD loss dominates the training with easy-to-transfer samples while vanilla loss dominates that with difficult-to-transfer samples. More importantly, we only select those samples with both low distillation difficulty and high diversity to train the student model for reducing computational cost. Experimental results on three video benchmarks and one image benchmark demonstrate the superiority of the proposed method by striking a good balance between performance and efficiency. Code is available at https://github.com/mlvccn/SAKD_ActionRec.
Ping Li 0006, Chenhao Ping, Wenxiao Wang 0001, Mingli Song
ACM Multimedia1
2025 Uncertainty Estimation with Self-Distillation for semi-supervised few-shot classification
Ping Li 0006, Renshu Gu
Knowl. Based Syst.2
2025 Pseudo-labeling with keyword refining for few-supervised video captioning
Ping Li 0006, Xinkui Zhao, Xianghua Xu, Mingli Song
Pattern Recognit.1
2025 Coarse-to-Fine Hypergraph Network for Spatiotemporal Action Detection
abstract
Spatiotemporal action detection localizes the action instances along both spatial and temporal dimensions, by identifying action start time and end time, action class, and object (e.g., actor) bounding boxes. It faces two primary challenges: 1) varying durations of actions and inconsistent tempo of action instances within the same class, and 2) modeling complex object interactions, which are not well handled by previous methods. For the former, we develop the coarse-to-fine attention module, which employs an efficient dynamic time warping to make a coarse estimation of action frames by eliminating context-agnostic features, and further adopts the attention mechanism to capture the first-order object relations within those action frames. This results in a finer-granularity of action estimation. For the latter, we design the ternary high-order hypergraph neural networks, which model the spatial relation, the motion dynamics, and the high-order relations of different objects across frames. This encourages the positive relation of the objects within the same actions, while suppressing the negative relation of those in different actions. Therefore, we present a Coarse-to-Fine Hypergraph Network, abbreviated as CFHN, for spatiotemporal action detection, by considering the object local context, the first-order object relations, and the high-order object relations together. It combines the spatiotemporal first-order and high-order features along the channel dimension to obtain satisfying detection results. Extensive experiments on several benchmarks including AVA, JHMDB-21, and UCF101-24 demonstrate the superiority of the proposed approach.
Ping Li 0006, Xingchao Ye
IEEE Trans. Circuits Syst. Video Technol.1
2024 Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-Dataset 3D Object Detection
abstract
Recent self-training techniques have shown notable improvements in unsupervised domain adaptation for 3D object detection (3D UDA). These techniques typically select pseudo labels, i.e., 3D boxes, to supervise models for the target domain. However, this selection process inevitably introduces unreliable 3D boxes, in which 3D points cannot be definitively assigned as foreground or background. Previous techniques mitigate this by reweighting these boxes as pseudo labels, but these boxes can still poison the training process. To resolve this problem, in this paper, we propose a novel pseudo label refinery framework. Specifically, in the selection process, to improve the reliability of pseudo boxes, we propose a complementary augmentation strategy. This strategy involves either removing all points within an unre-liable box or replacing it with a high-confidence box. More-over, the point numbers of instances in high-beam datasets are considerably higher than those in low-beam datasets, also degrading the quality of pseudo labels during the training process. We alleviate this issue by generating additional proposals and aligning RoI features across different domains. Experimental results demonstrate that our method effectively enhances the quality of pseudo labels and consistently surpasses the state-of-the-art methods on six autonomous driving benchmarks. Code will be available at https://github.com/Zhanwei-Z/PERE.
Zhanwei Zhang, Minghao Chen 0001, Hengjia Li, Binbin Lin 0001, Ping Li 0006, Wenxiao Wang 0001, Boxi Wu 0001, Deng Cai 0001
CVPR7
2024 Residual spatial fusion network for RGB-thermal semantic segmentation
Ping Li 0006, Binbin Lin 0001, Xianghua Xu
Neurocomputing1
2024 Fully Transformer-Equipped Architecture for end-to-end Referring Video Object Segmentation
Ping Li 0006, Li Yuan 0007, Xianghua Xu
Inf. Process. Manag.1
2024 Bridging knowledge distillation gap for few-sample unsupervised semantic segmentation
Ping Li 0006
Inf. Sci.1
2024 Efficient Long-Short Temporal Attention network for unsupervised Video Object Segmentation
Ping Li 0006, Li Yuan 0007, Huaxin Xiao, Binbin Lin 0001, Xianghua Xu
Pattern Recognit.1
2024 Adversarial Attacks on Video Object Segmentation With Hard Region Discovery
abstract
Video object segmentation has been applied to various computer vision tasks, such as video editing, autonomous driving, and human-robot interaction. However, the methods based on deep neural networks are vulnerable to adversarial examples, which are the inputs attacked by almost human-imperceptible perturbations, and the adversary (i.e., attacker) will fool the segmentation model to make incorrect pixel-level predictions. This will rise the security issues in highly-demanding tasks because small perturbations to the input video will result in potential attack risks. Though adversarial examples have been extensively used for classification, it is rarely studied in video object segmentation. Existing related methods in computer vision either require prior knowledge of categories or cannot be directly applied due to the special design for certain tasks, failing to consider the pixel-wise region attack. Hence, this work develops an object-agnostic adversary that has adversarial impacts on VOS by first-frame attacking via hard region discovery. Particularly, the gradients from the segmentation model are exploited to discover the easily confused region, in which it is difficult to identify the pixel-wise objects from the background in a frame. This provides a hardness map that helps to generate perturbations with a stronger adversarial power for attacking the first frame. Empirical studies on three benchmarks indicate that our attacker significantly degrades the performance of several state-of-the-art video object segmentation models.
Ping Li 0006, Li Yuan 0007, Jian Zhao 0006, Xianghua Xu, Xiaoqin Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2024 Fast Fourier Inception Networks for Occluded Video Prediction
abstract
Video prediction is a pixel-level task that generates future frames by employing the historical frames. There often exist continuous complex motions, such as object overlapping and scene occlusion in video, which poses great challenges to this task. Previous works either fail to well capture the long-term temporal dynamics or do not handle the occlusion masks. To address these issues, we develop the fully convolutional Fast Fourier Inception Networks for video prediction, termedFFINet, which includes two primary components, i.e., the occlusion inpainter and the spatiotemporal translator. The former adopts the fast Fourier convolutions to enlarge the receptive field, such that the missing areas (occlusion) with complex geometric structures are filled by the inpainter. The latter employs the stacked Fourier transform inception module to learn the temporal evolution by group convolutions and the spatial movement by channel-wise Fourier convolutions, which captures both the local and the global spatiotemporal features. This encourages generating more realistic and high-quality future frames. To optimize the model, the recovery loss is imposed to the objective, i.e., minimizing the mean square error between the ground-truth frame and the recovery frame. Both quantitative and qualitative experimental results on five benchmarks, including Moving MNIST, TaxiBJ, Human3.6 M, Caltech Pedestrian, and KTH, have demonstrated the superiority of the proposed approach.
Ping Li 0006, Chenhan Zhang, Xianghua Xu
IEEE Trans. Multim.1
2023 Prototype contrastive learning for point-supervised temporal action detection
Ping Li 0006, Jiachen Cao, Xingchao Ye
Expert Syst. Appl.1
2023 Time-frequency recurrent transformer with diversity constraint for dense video captioning
Ping Li 0006, Huaxin Xiao
Inf. Process. Manag.1
2023 Deep metric learning via group channel-wise ensemble
Ping Li 0006, Guopan Zhao, Xianghua Xu
Knowl. Based Syst.1
2023 Truncated attention-aware proposal networks with multi-scale dilation for temporal action detection
Ping Li 0006, Jiachen Cao, Li Yuan 0007, Qinghao Ye, Xianghua Xu
Pattern Recognit.1
2022 Graph convolutional network meta-learning with multi-granularity POS guidance for video captioning
Ping Li 0006, Xianghua Xu
Neurocomputing1
2022 Coarse-to-fine few-shot classification with deep metric learning
Ping Li 0006, Guopan Zhao, Xianghua Xu
Inf. Sci.1
2022 Community-Aware Photo Quality Evaluation by Deeply Encoding Human Perception
abstract
Computational photo quality evaluation is a useful technique in many tasks of computer vision and graphics, for example, photo retaregeting, 3-D rendering, and fashion recommendation. The conventional photo quality models are designed by characterizing the pictures from all communities (e.g., "architecture" and "colorful") indiscriminately, wherein community-specific features are not exploited explicitly. In this article, we develop a new community-aware photo quality evaluation framework. It uncovers the latent community-specific topics by a regularized latent topic model (LTM) and captures human visual quality perception by exploring multiple attributes. More specifically, given massive-scale online photographs from multiple communities, a novel ranking algorithm is proposed to measure the visual/semantic attractiveness of regions inside each photograph. Meanwhile, three attributes, namely: 1) photo quality scores; weak semantic tags; and inter-region correlations, are seamlessly and collaboratively incorporated during ranking. Subsequently, we construct the gaze shifting path (GSP) for each photograph by sequentially linking the top-ranking regions from each photograph, and an aggregation-based CNN calculates the deep representation for each GSP. Based on this, an LTM is proposed to model the GSP distribution from multiple communities in the latent space. To mitigate the overfitting problem caused by communities with very few photographs, a regularizer is incorporated into our LTM. Finally, given a test photograph, we obtain its deep GSP representation and its quality score is determined by the posterior probability of the regularized LTM. Comparative studies on four image sets have shown the competitiveness of our method. Besides, the eye-tracking experiments have demonstrated that our ranking-based GSPs are highly consistent with real human gaze movements.
Yongheng Shang, Ping Li 0006, Hao Luo 0001, Ling Shao 0001
IEEE Trans. Cybern.3
2021 Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual Fusion
abstract
Video highlight detection plays an increasingly important role in social media content filtering, however, it remains highly challenging to develop automated video highlight detection methods because of the lack of temporal annotations (i.e., where the highlight moments are in long videos) for supervised learning. In this paper, we propose a novel weakly supervised method that can learn to detect highlights by mining video characteristics with video level annotations (topic tags) only. Particularly, we exploit audio-visual features to enhance video representation and take temporal cues into account for improving detection performance. Our contributions are threefold: 1) we propose an audio-visual tensor fusion mechanism that efficiently models the complex association between two modalities while reducing the gap of the heterogeneity between the two modalities; 2) we introduce a novel hierarchical temporal context encoder to embed local temporal clues in between neighboring segments; 3) finally, we alleviate the gradient vanishing problem theoretically during model optimization with attention-gated instance aggregation. Extensive experiments on two benchmark datasets (YouTube Highlights and TVSum) have demonstrated our method outperforms other state-of-the-art methods with remarkable improvements.
Qinghao Ye, Xiyue Shen, Yuan Gao 0017, Qi Bi, Ping Li 0006, Guang Yang 0006
ICCV6
2021 Adaptive Deep Metric Ensemble Learning with Consensus
abstract
To measure semantic similarity of data pairs using multiple metrics has received much attention. It is natural to construct an ensemble consisting of base learners for modeling distinct properties of data distribution in the embedding space. However, previous works fail to correlate data subset generation with the performance of base learners and ignore possibly large deviations of different metrics on measuring the same data pairs. To address these issues, we propose the Consensus-aware Adaptive Ensemble (CAE) framework for deep metric learning. CAE adaptively provides data subsets for base learners which are empowered with both diversity and consensus. For each learner, the samples of the classes having larger intra-class distances in the embedding space will form the subset periodically. Moreover, to diversify learners, we employ determinant point process loss to make them capture various semantics of data distributions, promoting the generalization ability of ensemble. Meanwhile, different learners are expected to agree on measuring the same data pairs, so we develop the consensus loss to guide learners to reach a consensus as much as possible. Extensive experiments on several benchmarks demonstrate that CAE achieves state-of-the-art performance, validating its effectiveness.
Ping Li 0006, Guopan Zhao, Huaxin Xiao
ICME1
2021 Dynamic Cross Fusion Network for Building-Based Damage Assessment
abstract
Building-based damage assessment usually involves two different tasks, i.e., localizing buildings before disasters and scoring the damage levels after disasters. Recently, Convolutional Neural Network based machine learning methods achieve impressive performance by directly fusing the feature from pre- and post-disaster satellite imageries, which are collected at totally different dates and climatic conditions. The discrepancy of input sources may disturb the learning process, especially when simultaneously detecting buildings and assessing the damage levels by a unified model become a new trend. In this paper, we propose a Dynamic Cross Fusion Network (DCFNet) that makes each task bidirectionally access to multi-levels features from a shared backbone network. Moreover, a task-shared head is designed to adaptively exchange information cross tasks and avoid under-fitting for one specific task. Without bells and whistles, the proposed DCSNet achieves state-of-the-art single-model performance in the xView2 Challenge, which provides a new and large-scale damage assessment dataset xBD.
Huaxin Xiao, Hanlin Tan, Ping Li 0006
ICME4
2021 Video summarization with a graph convolutional attention network
abstract
Video summarization has established itself as a fundamental technique for generating compact and concise video, which alleviates managing and browsing large-scale video data. Existing methods fail to fully consider the local and global relations among frames of video, leading to a deteriorated summarization performance. To address the above problem, we propose a graph convolutional attention network (GCAN) for video summarization. GCAN consists of two parts, embedding learning and context fusion, where embedding learning includes the temporal branch and graph branch. In particular, GCAN uses dilated temporal convolution to model local cues and temporal self-attention to exploit global cues for video frames. It learns graph embedding via a multi-layer graph convolutional network to reveal the intrinsic structure of frame samples. The context fusion part combines the output streams from the temporal branch and graph branch to create the context-aware representation of frames, on which the importance scores are evaluated for selecting representative frames to generate video summary. Experiments are carried out on two benchmark databases, SumMe and TVSum, showing that the proposed GCAN approach enjoys superior performance compared to several state-of-the-art alternatives in three evaluation settings.
Ping Li 0006, Xianghua Xu
Frontiers Inf. Technol. Electron. Eng.1
2021 Exploring global diverse attention via pairwise temporal relation for video summarization
Ping Li 0006, Qinghao Ye, Li Yuan 0007, Xianghua Xu, Ling Shao 0001
Pattern Recognit.1
2020 Abnormal visual event detection based on multi-instance learning and autoregressive integrated moving average model in edge-based Smart City surveillance
abstract
Summary The abnormal visual event detection is an important subject in Smart City surveillance where a lot of data can be processed locally in edge computing environment. Real‐time and detection effectiveness are critical in such an edge environment. In this paper, we propose an abnormal event detection approach based on multi‐instance learning and autoregressive integrated moving average model for video surveillance of crowded scenes in urban public places, focusing on real‐time and detection effectiveness. We propose an unsupervised method for abnormal event detection by combining multi‐instance visual feature selection and the autoregressive integrated moving average model. In the proposed method, each video clip is modeled as a visual feature bag containing several subvideo clips, each of which is regarded as an instance. The time‐transform characteristics of the optical flow characteristics within each subvideo clip are considered as a visual feature instance, and time‐series modeling is carried out for multiple visual feature instances related to all subvideo clips in a surveillance video clip. The abnormal events in each surveillance video clip are detected using the multi‐instance fusion method. This approach is verified on publically available urban surveillance video datasets and compared with state‐of‐the‐art alternatives. Experimental results demonstrate that the proposed method has better abnormal event detection performance for crowded scene of urban public places with an edge environment.
Xianghua Xu, LiQiming Liu, Lingjun Zhang, Ping Li 0006, Jinjun Chen
Softw. Pract. Exp.4
2020 Unsupervised Video Summarization With Cycle-Consistent Adversarial LSTM Networks
abstract
Video summarization is an important technique to browse, manage and retrieve a large amount of videos efficiently. The main objective of video summarization is to minimize the information loss when selecting a subset of video frames from the original video, hence the summary video can faithfully represent the overall story of the original video. Recently developed unsupervised video summarization approaches are free of requiring tedious annotation on important frames to train a video summarization model and thus are practically attractive. However, their performance is still limited due to the difficulty of minimizing information loss between the summary and original videos. In this paper, we address unsupervised video summarization by developing a novel Cycle-consistent Adversarial LSTM architecture to effectively reduce the information loss in the summary video. The proposed model, named Cycle-SUM, consists of a frame selector and a cycle-consistent learning based evaluator. The selector is a bi-directional LSTM network to capture the long-range relationship between video frames. To overcome the difficulty of specifying a suitable information preserving metric between original video and summary video, the evaluator is introduced to “supervise” selector to improve the video summarization quality. Specifically, the evaluator is composed of two generative adversarial networks (GANs), in which the forward GAN component is learned to reconstruct the original video from summary video, while the backward GAN learns to invert the process. We establish the relation between mutual information maximization and such cycle learning procedure and further introduce cycle-consistent loss to regularize the summarization. Extensive experiments on three video summarization benchmark datasets demonstrate a state-of-the-art performance, and show the superiority of the Cycle-SUM model compared with other unsupervised approaches.
Li Yuan 0007, Francis E. H. Tay, Ping Li 0006, Jiashi Feng
IEEE Trans. Multim.3
2020 Flickr Image Community Analytics by Deep Noise-Refined Matrix Factorization
abstract
Accurately categorizing Flickr images into multiple pre-defined communities (e.g., “architecture” and “peaceful”) is an indispensable technique in multimedia analysis, graphic design, fashion recommendation, etc. In practice, these communities are constructed and updated manually, which is subjective and intolerably time consuming. To alleviate these shortcomings, a noise-refined deep matrix factorization (MF) framework is proposed to intelligently discover communities from million-scale Flickr users, wherein the semantic tag correlations and community correlations are simultaneously encoded. More specifically, it is believable that Flickr communities are high-level clues on the basis of human visual semantic perception. Thereby, a MF algorithm is employed to approximate the community label matrix by the product of pairwise factor matrices, which represent the latent representations of user-provided tags and the corresponding basis matrix respectively. Subsequently, an end-to-end deep model is formulated to hierarchically derive the latent deep representation from raw image pixels to semantic tags. To robustly handle contaminated image semantic tags and community labels, an l1norm constraint is encoded to enhance the MF. Meanwhile, to optimally exploit the rich context information of Flickr images, the intrinsic structure between image semantic tags and between communities are collaboratively captured. Finally, the upgraded MF and the deep model are seamlessly combined into a unified framework, which is solved by an iterative algorithm. Experiments on 2 M Flickr images have demonstrated the superiority of our approach. Besides, the discovered Flickr communities can improve photo retargeting and visual aesthetics assessment significantly.
Jianwei Yin, Ping Li 0006, Yongheng Shang, Roger Zimmermann, Ling Shao 0001
IEEE Trans. Multim.3
2019 Cycle-SUM: Cycle-Consistent Adversarial LSTM Networks for Unsupervised Video Summarization
abstract
In this paper, we present a novel unsupervised video summarization model that requires no manual annotation. The proposed model termed Cycle-SUM adopts a new cycleconsistent adversarial LSTM architecture that can effectively maximize the information preserving and compactness of the summary video. It consists of a frame selector and a cycle-consistent learning based evaluator. The selector is a bi-direction LSTM network that learns video representations that embed the long-range relationships among video frames. The evaluator defines a learnable information preserving metric between original video and summary video and “supervises” the selector to identify the most informative frames to form the summary video. In particular, the evaluator is composed of two generative adversarial networks (GANs), in which the forward GAN is learned to reconstruct original video from summary video while the backward GAN learns to invert the processing. The consistency between the output of such cycle learning is adopted as the information preserving metric for video summarization. We demonstrate the close relation between mutual information maximization and such cycle learning procedure. Experiments on two video summarization benchmark datasets validate the state-of-theart performance and superiority of the Cycle-SUM model over previous baselines.
Li Yuan 0007, Francis E. H. Tay, Ping Li 0006, Jiashi Feng
AAAI3
2019 Scene Categorization Using Deeply Learned Gaze Shifting Kernel
abstract
Accurately recognizing sophisticated sceneries from a rich variety of semantic categories is an indispensable component in many intelligent systems, e.g., scene parsing, video surveillance, and autonomous driving. Recently, there have emerged a large quantity of deep architectures for scene categorization, wherein promising performance has been achieved. However, these models cannot explicitly encode human visual perception toward different sceneries, i.e., the sequence of humans sequentially allocates their gazes. To solve this problem, we propose deep gaze shifting kernel to distinguish sceneries from different categories. Specifically, we first project regions from each scenery into the so-called perceptual space, which is established by combining color, texture, and semantic features. Then, a novel non-negative matrix factorization algorithm is developed which decomposes the regions' feature matrix into the product of the basis matrix and the sparse codes. The sparse codes indicate the saliency level of different regions. In this way, the gaze shifting path from each scenery is derived and an aggregation-based convolutional neural network is designed accordingly to learn its deep representation. Finally, the deep representations of gaze shifting paths from all the scene images are incorporated into an image kernel, which is further fed into a kernel SVM for scene categorization. Comprehensive experiments on six scenery data sets have demonstrated the superiority of our method over a series of shallow/deep recognition models. Besides, eye tracking experiments have shown that our predicted gaze shifting paths are 94.6% consistent with the real human gaze allocations.
Xiao Sun 0003, Zepeng Wang 0003, Jie Chang 0001, Yiyang Yao, Ping Li 0006, Roger Zimmermann
IEEE Trans. Cybern.6
2019 Learning Latent Stable Patterns for Image Understanding With Weak and Noisy Labels
abstract
This paper focuses on weakly supervised image understanding, in which the semantic labels are available only at image-level, without the specific object or scene location in an image. Existing algorithms implicitly assume that image-level labels are error-free, which might be too restrictive. In practice, image labels obtained from the pretrained predictors are easily contaminated. To solve this problem, we propose a novel algorithm for weakly supervised segmentation when only noisy image labels are available during training. More specifically, a semantic space is constructed first by encoding image labels through a graphlet (i.e., superpixel cluster) embedding process. Then, we observe that in the semantic space, the distribution of graphlets from images with a same label remains stable, regardless of the noises in image labels. Therefore, we propose a generative model, called latent stability analysis, to discover the stable patterns from images with noisy labels. Inferring graphlet semantics by making use of these mid-level stable patterns is much more secure and accurate than directly transferring noisy image-level labels into different regions. Finally, we calculate the semantics of each superpixel using maximum majority voting of its correlated graphlets. Comprehensive experimental results show that our algorithm performs impressively when the image labels are predicted by either the hand-crafted or deeply learned image descriptors.
Yiyang Yao, Luo Wang, Yi Yang 0001, Ping Li 0006, Roger Zimmermann, Ling Shao 0001
IEEE Trans. Cybern.5
2019 Online Robust Low-Rank Tensor Modeling for Streaming Data Analysis
abstract
Tensor data (i.e., the data having multiple dimensions) are quickly growing in scale in many practical applications, which poses new challenges for data modeling and analysis approaches, such as high-order relations of large complexity, gross noise, and varying data scale. Existing low-rank data analysis methods, which are effective at analyzing matrix data, may fail in the regime of tensor data due to these challenges. A robust and scalable low-rank tensor modeling method is heavily desired. In this paper, we develop an online robust low-rank tensor modeling (ORLTM) method to address these challenges. The ORLTM method leverages the high-order correlations among all tensor modes to model an intrinsic low-rank structure of streaming tensor data online and can effectively analyze data residing in a mixture of multiple subspaces by virtue of dictionary learning. ORLTM consumes a very limited memory space that remains constant regardless of the increase of tensor data size, which facilitates processing tensor data at a large scale. More concretely, it models each mode unfolding of streaming tensor data using the bilinear formulation of tensor nuclear norms. With this reformulation, ORLTM employs a stochastic optimization algorithm to learn the tensor low-rank structure alternatively for online updating. To capture the final tensors, ORLTM uses an average pooling operation on folded tensors in all modes. We also provide the analysis regarding computational complexity, memory cost, and convergence. Moreover, we extend ORLTM to the image alignment scenario by incorporating the geometrical transformations and linearizing the constraints. Extensive empirical studies on synthetic database and three practical vision tasks, including video background subtraction, image alignment, and visual tracking, have demonstrated the superiority of the proposed method.
Ping Li 0006, Jiashi Feng, Xiaojie Jin 0004, Xianghua Xu, Shuicheng Yan
IEEE Trans. Neural Networks Learn. Syst.1
2018 Engineering Deep Representations for Modeling Aesthetic Perception
abstract
Many aesthetic models in multimedia and computer vision suffer from two shortcomings: 1) the low descriptiveness and interpretability 1 of the hand-crafted aesthetic criteria (i.e., fail to indicate region-level aesthetics) and 2) the difficulty of engineering aesthetic features adaptively and automatically toward different image sets. To remedy these problems, we develop a deep architecture to learn aesthetically relevant visual attributes from Flickr, 2 which are localized by multiple textual attributes in a weakly supervised setting. More specifically, using a bag-of-words representation of the frequent Flickr image tags, a sparsity-constrained subspace algorithm discovers a compact set of textual attributes (i.e., each textual attribute is a sparse and linear representation of those frequent image tags) for each Flickr image. Then, a weakly supervised learning algorithm projects the textual attributes at image-level to the highly-responsive image patches. These patches indicate where humans look at appealing regions with respect to each textual attribute, which are employed to learn the visual attributes. Psychological and anatomical studies have demonstrated that humans perceive visual concepts in a hierarchical way. Therefore, we normalize these patches and further feed them into a five-layer convolutional neural network to mimic the hierarchy of human perceiving the visual attributes. We apply the learned deep features onto applications like image retargeting, aesthetics ranking, and retrieval. Both subjective and objective experimental results thoroughly demonstrate the superiority of our approach.1 In this paper, "describing" and "interpretability" means the ability of seeking region-level representation of each mined textual attribute, i.e., a sparse and linear representation of those frequent image tags. 2 https://www.flickr.com/.
Yuxing Hu, Ping Li 0006, Chao Zhang 0014
IEEE Trans. Cybern.4
2018 Camera-Assisted Video Saliency Prediction and Its Applications
abstract
Video saliency prediction is an indispensable yet challenging technique which can facilitate various applications, such as video surveillance, autonomous driving, and realistic rendering. Based on the popularity of embedded cameras, we in this paper predict region-level saliency from videos by leveraging human gaze locations recorded using a camera, (e.g., those equipped on an iMAC and laptop PC). Our proposed camera-assisted mechanism improves saliency prediction by discovering human attended regions inside a video clip. It is orthogonal to the current saliency models, i.e., any existing video/image saliency model can be boosted by our mechanism. First of all, the spatial-and temporal-level visual features are exploited collaboratively for calculating an initial saliency map. We notice that the current saliency models are not sufficiently adaptable to the variations in lighting, different view angles, and complicated backgrounds. Therefore, assisted by a camera tracking human gaze movements, a non-negative matrix factorization algorithm is designed to accurately localize the semantically/visually salient video regions perceived by humans. Finally, the learned human gaze locations as well as the initial saliency map are integrated to optimize video saliency calculation. Empirical results thoroughly demonstrated that: 1) our approach achieves the state-of-the-art video saliency prediction accuracy by outperforming 11 mainstream algorithms considerably and 2) our method can conveniently and successfully enhance video retargeting, quality estimation, and summarization.
Xiao Sun 0003, Yuxing Hu, Ping Li 0006, Zhao Xie, Zhenguang Liu
IEEE Trans. Cybern.5
2018 Perceptually Aware Image Retargeting for Mobile Devices
abstract
Retargeting aims at adapting an original high-resolution photograph/video to a low-resolution screen with an arbitrary aspect ratio. Conventional approaches are generally based on desktop PCs, since the computation might be intolerable for mobile platforms (especially when retargeting videos). Typically, only low-level visual features are exploited, and human visual perception is not well encoded. In this paper, we propose a novel retargeting framework that rapidly shrinks a photograph/video by leveraging human gaze behavior. Specifically, we first derive a geometry-preserving graph ranking algorithm, which efficiently selects a few salient object patches to mimic the human gaze shifting path (GSP) when viewing a scene. Afterward, an aggregation-based CNN is developed to hierarchically learn the deep representation for each GSP. Based on this, a probabilistic model is developed to learn the priors of the training photographs that are marked as aesthetically pleasing by professional photographers. We utilize the learned priors to efficiently shrink the corresponding GSP of a retargeted photograph/video to maximize its similarity to those from the training photographs. Extensive experiments have demonstrated that: 1) our method requires less than 35 ms to retarget a photograph (or a video frame) on popular iOS/Android devices, which is orders of magnitude faster than the conventional retargeting algorithms; 2) the retargeted photographs/videos produced by our method significantly outperform those of its competitors based on a paired-comparison-based user study; and 3) the learned GSPs are highly indicative of human visual attention according to the human eye tracking experiments.
Yinzuo Zhou, Chao Zhang 0014, Ping Li 0006, Xuelong Li 0001
IEEE Trans. Image Process.4
2017 Online Robust Low-Rank Tensor Learning
abstract
The rapid increase of multidimensional data (a.k.a. tensor) like videos brings new challenges for low-rank data modeling approaches such as dynamic data size, complex high-order relations, and multiplicity of low-rank structures. Resolving these challenges require a new tensor analysis method that can perform tensor data analysis online, which however is still absent. In this paper, we propose an Online Robust Low-rank Tensor Modeling (ORLTM) approach to address these challenges. ORLTM dynamically explores the high-order correlations across all tensor modes for low-rank structure modeling. To analyze mixture data from multiple subspaces, ORLTM introduces a new dictionary learning component. ORLTM processes data streamingly and thus requires quite low memory cost that is independent of data size. This makes ORLTM quite suitable for processing large-scale tensor data. Empirical studies have validated the effectiveness of the proposed method on both synthetic data and one practical task, i.e., video background subtraction. In addition, we provide theoretical analysis regarding computational complexity and memory cost, demonstrating the efficiency of ORLTM rigorously.
Ping Li 0006, Jiashi Feng, Xiaojie Jin 0004, Xianghua Xu, Shuicheng Yan
IJCAI1
2017 Constrained Low-Rank Learning Using Least Squares-Based Regularization
abstract
Low-rank learning has attracted much attention recently due to its efficacy in a rich variety of real-world tasks, e.g., subspace segmentation and image categorization. Most low-rank methods are incapable of capturing low-dimensional subspace for supervised learning tasks, e.g., classification and regression. This paper aims to learn both the discriminant low-rank representation (LRR) and the robust projecting subspace in a supervised manner. To achieve this goal, we cast the problem into a constrained rank minimization framework by adopting the least squares regularization. Naturally, the data label structure tends to resemble that of the corresponding low-dimensional representation, which is derived from the robust subspace projection of clean data by low-rank learning. Moreover, the low-dimensional representation of original data can be paired with some informative structure by imposing an appropriate constraint, e.g., Laplacian regularizer. Therefore, we propose a novel constrained LRR method. The objective function is formulated as a constrained nuclear norm minimization problem, which can be solved by the inexact augmented Lagrange multiplier algorithm. Extensive experiments on image classification, human pose estimation, and robust face recovery have confirmed the superiority of our method.
Ping Li 0006, Jun Yu 0002, Meng Wang 0001, Deng Cai 0001, Xuelong Li 0001
IEEE Trans. Cybern.1
2016 Towards robust subspace recovery via sparsity-constrained latent low-rank representation
Ping Li 0006, Jiajun Bu, Jun Yu 0002, Chun Chen 0001
J. Vis. Commun. Image Represent.1
2015 Multi-view based multi-label propagation for image annotation
Zhanying He, Chun Chen 0001, Jiajun Bu, Ping Li 0006, Deng Cai 0001
Neurocomputing4
2015 Graph-based local concept coordinate factorization
Ping Li 0006, Jiajun Bu, Lijun Zhang 0005, Chun Chen 0001
Knowl. Inf. Syst.1
2015 Sparse fixed-rank representation for robust visual analysis
Ping Li 0006, Jiajun Bu, Bin Xu 0005, Zhanying He, Chun Chen 0001, Deng Cai 0001
Signal Process.1
2014 Discriminative Orthogonal Nonnegative matrix factorization with flexibility for data representation
Ping Li 0006, Jiajun Bu, Yi Yang 0001, Rongrong Ji, Chun Chen 0001, Deng Cai 0001
Expert Syst. Appl.1
2014 Manifold optimal experimental design via dependence maximization for active learning
Ping Li 0006, Jiajun Bu, Chun Chen 0001, Deng Cai 0001
Neurocomputing1
2014 Hypergraph canonical correlation analysis for multi-label classification
Ping Li 0006
Signal Process.2
2013 Subspace learning via Locally Constrained A-optimal nonnegative projection
Ping Li 0006, Jiajun Bu, Chun Chen 0001, Can Wang 0001, Deng Cai 0001
Neurocomputing1
2013 Locally discriminative spectral clustering with composite manifold
Ping Li 0006, Jiajun Bu, Bin Xu 0005, Beidou Wang, Chun Chen 0001
Neurocomputing1
2013 Multi-label ensemble based on variable pairwise constraint projection
Ping Li 0006, Min Wu 0002
Inf. Sci.1
2013 Relational Multimanifold Coclustering
abstract
Coclustering targets on grouping the samples (e.g.,documents and users) and the features (e.g., words and ratings) simultaneously. It employs the dual relation and the bilateral information between the samples and features. In many real-world applications, data usually reside on a submanifold of the ambient Euclidean space, but it is nontrivial to estimate the intrinsic manifold of the data space in a principled way. In this paper, we focus on improving the coclustering performance via manifold ensemble learning, which is able to maximally approximate the intrinsic manifolds of both the sample and feature spaces. To achieve this, we develop a novel coclustering algorithm called relational multimanifold coclustering based on symmetric nonnegative matrix trifactorization, which decomposes the relational data matrix into three submatrices. This method considers the intertype relationship revealed by the relational data matrix and also the intratype information reflected by the affinity matrices encoded on the sample and feature data distributions. Specifically, we assume that the intrinsic manifold of the sample or feature space lies in a convex hull of some predefined candidate manifolds. We want to learn a convex combination of them to maximally approach the desired intrinsic manifold. To optimize the objective function, the multiplicative rules are utilized to update the submatrices alternatively. In addition, both the entropic mirror descent algorithm and the coordinate descent algorithm are exploited to learn the manifold coefficient vector. Extensive experiments on documents, images, and gene expression data sets have demonstrated the superiority of the proposed algorithm compared with other well-established methods.
Ping Li 0006, Jiajun Bu, Chun Chen 0001, Zhanying He, Deng Cai 0001
IEEE Trans. Cybern.1
2012 Relational co-clustering via manifold ensemble learning
abstract
Co-clustering targets on grouping the samples and features simultaneously. It takes advantage of the duality between the samples and features. In many real-world applications, the data points or features usually reside on a submanifold of the ambient Euclidean space, but it is nontrivial to estimate the intrinsic manifolds in a principled way. In this study, we focus on improving the co-clustering performance via manifold ensemble learning, which aims to maximally approximate the intrinsic manifolds of both the sample and feature spaces. To achieve this, we develop a novel co-clustering algorithm called Relational Multi-manifold Co-clustering (RMC) based on symmetric nonnegative matrix tri-factorization, which decomposes the relational data matrix into three matrices. This method considers the inter-type relationship revealed by the relational data matrix and the intra-type information reflected by the affinity matrices. Specifically, we assume the intrinsic manifold of the sample or feature space lies in a convex hull of a group of pre-defined candidate manifolds. We hope to learn an appropriate convex combination of them to approach the desired intrinsic manifold. To optimize the objective, the multiplicative rules are utilized to update the factorized matrices and the entropic mirror descent algorithm is exploited to automatically learn the manifold coefficients. Experimental results demonstrate the superiority of the proposed algorithm.
Ping Li 0006, Jiajun Bu, Chun Chen 0001, Zhanying He
CIKM1
2012 Clustering analysis using manifold kernel concept factorization
Ping Li 0006, Chun Chen 0001, Jiajun Bu
Neurocomputing1
2010 Combine multi-valued attribute decomposition with multi-label learning
Yuejian Guo, Min Wu 0002, Ping Li 0006, Yao Xiang
Expert Syst. Appl.4