Peiguang Jing

dblp:04/10628 · DBLP profile ↗
← Back
85ranked-venue papers
23as first author
54since 2021 · last 2027
0000-0003-2648-7358ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 14 first-author · 30 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 12 since 2021Computer networks · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2027 CProPNet: Chrono-progressive dietary synergy network with pathology decoupling for individualized HbA1c prediction
Jing Liu 0002, Peiguang Jing, Yu Liu 0004
Expert Syst. Appl.3
2026 Multi-stage Superpixel-guided Mamba-based Network for Change Detection
Jing Liu 0002, Yuting Su 0001, Peiguang Jing
ISCAS5
2026 IoT-Oriented EEG-Based Memory State Recognition With Music-Facilitated Hierarchical Temporal-Frequency Fusion
abstract
Music dynamically modulates neural activity in memory-related brain regions, enabling more precise retention of details, facilitating easier recall of information, and enhancing working memory capacity. However, most current electroencephalography (EEG)-based memory state recognition approaches neither directly embed music stimuli as inputs nor sufficiently quantify the influence of musical genres. To address this gap, we propose a hierarchical temporal–frequency fusion (HTFF) network that integrates EEG data with musical cues for memory state recognition. Specifically, we propose a music-facilitated EEG temporal–frequency fusion module that progressively fuses low-level perceptual cues into high-level memory representations via cascaded resonant EEG fusion reinforcers (REFRs). These REFR units allocate musical cues into the temporal and frequency domains, thereby enhancing EEG representations through a cross-domain coupling mechanism. We also propose a mixed-constraints mechanism that creates more compact intra-class sample distributions and sharpens decision boundaries for musical genres. The experimental results demonstrate that the HTFF network outperforms state-of-the-art methods in memory state recognition. Moreover, its compatibility with the Internet of Things (IoT) and wearable devices highlights its potential for real-time monitoring and personalized cognitive interventions, paving the way for practical applications in cognitive health.
Zhuang Miao, Yu Liu 0004, Peiguang Jing, Zhiju Huang
IEEE Internet Things J.4
2026 MSSFN: Multi-stimulus stereo spatiotemporal fusion network with pattern disentanglement for Alzheimer's disease diagnosis
Peiguang Jing, Yu Liu 0004, Sun-Yuan Kung
Inf. Process. Manag.2
2026 MIGF-Net: Multimodal interaction-guided fusion network for image aesthetics assessment
Yun Liu 0009, Zhipeng Wen, Leida Li, Peiguang Jing, Daoxin Fan
Pattern Recognit.4
2026 FacDNet: A low-rank factorized diffusion network with dual-U compensated attention for low-light enhancement
Peiguang Jing, Ningyuan Zhao, Lijun Lai, Weiming Wang 0002, Yuting Su 0001
Signal Process.1
2026 TFFN: Three-Branch Feature Fusion Network for Stereoscopic Omnidirectional Image Quality Assessment
abstract
Stereoscopic omnidirectional image (SOI) has both omnidirectional and stereoscopic perception features. Many previous models have proved the viewport characteristics and stereoscopic visual features are crucial for quality perception of SOI. However, effective monocular and binocular visual features extraction and fusion are difficult due to the size of SOI and inaccuracy of feature representation. In this paper, we proposed a three-branch feature fusion network (TFFN) by fusing two-stream binocular visual features and the important monocular features based on the viewport perspective. The hierarchical fusion module is first designed to fuse effective binocular visual features from different semantic scales, and the pseudo-difference information extraction module is built to obtain the accuracy monocular visual features to complement the binocular visual features. Finally, the above monocular and binocular visual features are fused together to measure the quality of SOI. The comparison experiments are conducted on three public datasets and the analysis of the results demonstrate the effectiveness of the proposed method.
Yun Liu 0009, Daoxin Fan, Huiyu Duan, Peiguang Jing, Guanghui Yue 0001, Guangtao Zhai
IEEE Trans. Multim.5
2025 MPBR: Multimodal Progressive Bidirectional Reasoning for Open-Set Fine-Grained Recognition
Junfu Tan, Peiguang Jing, Yu Liu 0004
ICCV2
2025 VIIS: Visible and Infrared Information Synthesis for Severe Low-Light Image Enhancement
abstract
Images captured in severe low-light circumstances often suffer from significant information absence. Existing singular modality image enhancement methods struggle to restore image regions lacking valid information. By leveraging light-impervious infrared images, visible and infrared image fusion methods have the potential to reveal information hidden in darkness. However, they primarily emphasize inter-modal complementation but neglect intra-modal enhancement, limiting the perceptual quality of output images. To address these limitations, we propose a novel task, dubbed visible and infrared information synthesis (VIIS), which aims to achieve both information enhancement and fusion of the two modalities. Given the difficulty in obtaining ground truth in the VIIS task, we design an information synthesis pretext task (ISPT) based on image augmentation. We employ a diffusion model as the framework and design a sparse attention-based dual-modalities residual (SADMR) conditioning mechanism to enhance information interaction between the two modalities. This mechanism enables features with prior knowledge from both modalities to adaptively and iteratively attend to each modality's information during the denoising process. Our extensive experiments demonstrate that our model qualitatively and quantitatively outperforms not only the state-of-the-art methods in relevant fields but also the newly designed baselines capable of both information enhancement and fusion. The code is available at https://github.com/Chenz418/VIIS.
Chen Zhao 0030, Mengyuan Yu, Peiguang Jing
WACV4
2025 Adversarial neighbor perception network with feature distillation for anomaly detection
Yuting Su 0001, Enqi Su, Weiming Wang 0002, Peiguang Jing, Dubuke Ma, Fu Lee Wang
Expert Syst. Appl.4
2025 RSDC-Net: Robust self-supervised dynamic collaboration network for infrared and visible image fusion
Yun Li 0006, Ningyuan Zhao, Peiguang Jing
Knowl. Based Syst.5
2025 LUFormer : A luminance-informed localized transformer with frequency augmentation for nighttime flare removal
Wei Lu 0026, Dubuke Ma, Peiguang Jing
Neural Networks4
2025 SubMap: A Partial Mapping Strategy for CGRA Based on sub-CGRA Exploration
abstract
Coarse-grained reconfigurable array (CGRA) is a quality hardware for compute-intensive loop kernels, with its excellent balance of performance, energy efficiency, and reconfigurability. However, the efficiency of CGRA depends heavily on how the compiler maps the data flow graph (DFG) extracted from application kernels onto the target architecture. Most existing CGRA compilers encounter the challenge of long compilation times due to excessive exploration space. To reduce the exploration space and compilation time, we propose SubMap, which adaptively explores a suitable sub-CGRA for different DFGs in a target CGRA and efficiently performs the mapping process. The experimental results show that SubMap greatly reduces the compilation time compared to the latest methods while maintaining the mapping quality. On HyCube$4\times 4$, SubMap has an average performance improvement of$9.47 \times $and$11.67 \times $, respectively, compared with Morpher (Pathfinder) and Morpher (SA). As the scale of the target CGRA increases, the performance improvement of SubMap becomes more pronounced.
Peiguang Jing, Sio-Hang Pun, Yu Liu 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 AMDANet: Augmented Multiscale Difference Aggregation Network for Image Change Detection
abstract
The field of remote sensing image change detection (CD) has made significant improvements with the rapid development of deep learning techniques. However, current methods often inadequately utilize difference features of bitemporal images, resulting in biased focus and insensitivity to change information. Furthermore, the classic challenges of pseudo-CD and edge recognition in complex scenes have also weakened CD performance. In this article, we propose an augmented multiscale difference aggregation network (AMDANet) for image CD, which incorporates a difference feature extractor (DFE) within a Siamese feature extractor to perceive changes by capturing differences between bitemporal features. To address the issue of biased focus, we propose a hierarchical feature aggregator (HFA) that captures intrascale interactions in parallel for multigranularity change perception while personalizing coarse-grained and fine-grained features to highlight the attention to change regions. To deepen the perception of complex dependency relationships, we further design an O-shape feature augmentor (OFA) that leverages an information feedback loop to achieve precise alignment of multigranularity features. The integration of information across different granularities improves the recognition of pseudo-changes and edges. Experimental results on three publicly available datasets demonstrate the superiority of AMDANet over current state-of-the-art (SOTA) methods. Our source code will be publicly available athttps://github.com/mp-st/AMDANet.
Yuting Su 0001, Peng Ma, Weiming Wang 0002, Shaochu Wang, Yun Li 0006, Peiguang Jing
IEEE Trans. Geosci. Remote. Sens.7
2025 Multimodal Dual-Graph Collaborative Network With Serial Attentive Aggregation Mechanism for Micro-Video Multi-Label Classification
abstract
The increasing commercial value of micro-videos has spurred a rising demand for grasping their contents. The abundant multimodal cues in micro-videos exhibit substantial potential in enhancing content comprehension. However, effectively harnessing the collaborative characteristics across different modalities remains a significant challenge, especially in multi-label scenarios due to inconsistent behaviors regarding label correlations. To better tackle this issue, in this paper, we first introduce a multimodal dual-graph collaborative network with serial attentive aggregation mechanism (MDGCN) for micro-video multi-label classification. In MDGCN, we exploit an asymmetric encoder-decoder framework, which incorporates multiple parallel encoders with complementary representations and a decoder to ensure the completeness of encoded results. Meanwhile, an adversarial constraint is used to ensure individual differences prominently featured within each modality. Furthermore, considering the inconsistency of label correlations across various modalities, we then construct a serial attentive graph convolutional network that employs an interactive dual-graph attention paradigm to sequentially integrate multimodal representations and dynamically explore label correlations. The experiments conducted on two datasets demonstrate that our proposed method outperforms state-of-the-art approaches.
Wei Lu 0026, Peiguang Jing, Weiming Wang 0002, Yuting Su 0001
IEEE Trans. Multim.3
2025 PADNet: Progressive-Difference-Aware Feature Reconstruction Mechanism for Anomaly Detection
Peiguang Jing, Weiming Wang 0002, Fu Lee Wang, Yuting Su 0001
IEEE Trans. Multim.2
2024 Graph Disentangled Contrastive Learning with Personalized Transfer for Cross-Domain Recommendation
abstract
Cross-Domain Recommendation (CDR) has been proven to effectively alleviate the data sparsity problem in Recommender System (RS). Recent CDR methods often disentangle user features into domain-invariant and domain-specific features for efficient cross-domain knowledge transfer. Despite showcasing robust performance, three crucial aspects remain unexplored for existing disentangled CDR approaches: i) The significance nuances of the interaction behaviors are ignored in generating disentangled features; ii) The user features are disentangled irrelevant to the individual items to be recommended; iii) The general knowledge transfer overlooks the user's personality when interacting with diverse items. To this end, we propose a Graph Disentangled Contrastive framework for CDR (GDCCDR) with personalized transfer by meta-networks. An adaptive parameter-free filter is proposed to gauge the significance of diverse interactions, thereby facilitating more refined disentangled representations. In sight of the success of Contrastive Learning (CL) in RS, we propose two CL-based constraints for item-aware disentanglement. Proximate CL ensures the coherence of domain-invariant features between domains, while eliminatory CL strives to disentangle features within each domains using mutual information between users and items. Finally, for domain-invariant features, we adopt meta-networks to achieve personalized transfer. Experimental results on four real-world datasets demonstrate the superiority of GDCCDR over state-of-the-art methods.
Jing Liu 0002, Lele Sun, Weizhi Nie, Peiguang Jing, Yuting Su 0001
AAAI4
2024 Collaborative representation based cross-domain semantic transfer for vehicle re-identification
Yun Li 0006, Yudou Tian, Peiguang Jing
Neurocomputing6
2024 Multimodal High-Order Relationship Inference Network for Fashion Compatibility Modeling in Internet of Multimedia Things
abstract
Recent progress in artificial intelligence (AI) have broadened various intelligent application scenarios on the Internet of Multimedia Things (IoMT). Due to the urgent demands for intelligence in the online fashion industry, using AI techniques to explore user’s clothing collocations from fashion data generated by the IoMT system is a challenging task. Under the background, fashion compatibility modeling (FCM), which aims to estimate the matching degree of a given outfit, has attracted great attention in the multimedia analysis field. However, most of the studies often fail to fully leverage multimodal content or ignore the sparse associations between fashion items. In this article, we propose a novel multimodal high-order relationship inference network (MHRIN) for FCM task. In MHRIN, we focus on enriching multimodal representations of fashion items by means of incorporating the category correlations and injecting high-order item–item connectivity. Concretely, considering that fashion collocations depend on the semantic relevance patterns between categories, we design a category correlations learning module to adaptively learn category representations. On this basis, multiple modality representations are aggregated by a hierarchical multimodal fusion module to generate visual-semantic embeddings. To address the item-item matching interactions issue, we further refine the final representations by a high-order message propagation module to absorb rich connection information. Experiments on the publicly available data set demonstrate the superiority of our MHRIN over state-of-the-art methods.
Peiguang Jing, Jing Zhang 0038, Yun Li 0006, Yuting Su 0001
IEEE Internet Things J.1
2024 Dual Preference Perception Network for Fashion Recommendation in Social Internet of Things
abstract
Nowadays, with the continuous development of information technology, the application scenarios of the Internet of Things (IoT) are progressively expanding to the social field, engendering widespread attention to the Social IoT (SIoT). Personalized fashion recommendation that possesses the potential to establish social relationships between clothing and humans has substantially broadened the scope of the SIoT, particularly with the flourishing fashion industry and the ascent of smart home. Compared to conventional recommendations, fashion recommendation generally suggests a collection of items rather than individual pieces for users. Additionally, considering the public acceptance alongside the user-specific preference is reasonable for fashion recommendation, however, current methods often overlook the former. To comprehensively capture the public acceptance and the user-specific preference, we propose a dual preference perception network (DP2Net) for fashion recommendation. First, a fashion corpus is constructed to facilitate the condensation of general taste, wherein adversarial learning and determinantal point process are leveraged to ensure representativeness and diversity of the corpus. Second, a user-general preference perception module is built based on a bottleneck transformer structure to generate aggregated representations for the corpus. Third, a user-specific preference perception module is constructed to acquire collaborative representations of users and outfits by employing an attentive heterogeneous graph embedding. The final loss functions of two preference perception modules are constructed by combining the representations of users, outfits, and the corpus. Experiments on large-scale real-world data sets demonstrate the effectiveness of the proposed method. To facilitate reproducible research, we have made our code publicly available athttps://github.com/KaiZhang1228/DP2Net.
Peiguang Jing, Xianyi Liu, Yun Li 0006, Yu Liu 0004, Yuting Su 0001
IEEE Internet Things J.1
2024 Multimodal deep hierarchical semantic-aligned matrix factorization method for micro-video multi-label classification
Fugui Fan, Yuting Su 0001, Yun Liu 0009, Peiguang Jing, Kaihua Qu, Yu Liu 0004
Inf. Process. Manag.4
2024 Multimodal semantic enhanced representation network for micro-video event detection
Yun Li 0006, Xianyi Liu, Peiguang Jing
Knowl. Based Syst.5
2024 A deep low-rank semantic factorization method for micro-video multi-label classification
Fugui Fan, Yuting Su 0001, Yu Liu 0004, Peiguang Jing, Kaihua Qu
Multim. Syst.4
2024 Research on type-aware fashion compatibility prediction based on a hybrid attention mechanism
Yun Li 0006, GuoXiang Li, Jing Zhang 0038, Peiguang Jing
Multim. Tools Appl.4
2024 Context-aware focal alignment network for micro-video multi-label classification
Weiheng Yao, Peiguang Jing, Jing Zhang 0038, Kim Fung Tsang, Shuqiang Wang
Pattern Anal. Appl.3
2024 Adaptive proposal network based on generative adversarial learning for weakly supervised temporal sentence grounding
Weikang Wang 0002, Yuting Su 0001, Jing Liu 0002, Peiguang Jing
Pattern Recognit. Lett.4
2024 Collaborative spatial-temporal video salient object detection with cross attention transformer
Yuting Su 0001, Weikang Wang 0002, Jing Liu 0002, Peiguang Jing
Signal Process.4
2024 Deep Matrix Factorization With Complementary Semantic Aggregation for Micro-Video Multi-Label Classification
abstract
Deep matrix factorization has been demonstrated in extracting hierarchical knowledge describing micro-video characteristics. However, the complementary information across distinct latent layers is often ignored. To address this issue, we propose a deep matrix factorization with complementary semantic aggregation (DMFCSA) method for micro-video multi-label classification, which consists of the multi-layer representation learning module and the semantic decoding module. We first employ deep hierarchical matrix factorization to learn the underlying semantic representations at each latent layer. Meanwhile, the semantic aggregation strategy is exploited to integrate complementary information from different layers into the output layer. To enhance discriminability, we construct a triplet term that effectively establishes relationships among features, labels, and attributes. Moreover, the semantic decoding module is designed to enhance both robustness and representation ability by reconstructing the original inputs. Experimental results on a real-world multi-label dataset show the effectiveness and robustness of our method compared to several state-of-the-art methods.
Peiguang Jing, Yuting Su 0001
IEEE Signal Process. Lett.1
2024 Deep Learning-Based Eye-Tracking Analysis for Diagnosis of Alzheimer's Disease Using 3D Comprehensive Visual Stimuli
abstract
Alzheimer's Disease (AD) is a neurodegenerative disorder that causes a continuous decline in cognitive functions and eventually results in death. An early AD diagnosis is important for taking active measures to slow its deterioration. Traditional diagnoses are usually based on clinical experience, which is limited by several realistic factors. In this paper, we focus on exploiting deep learning techniques to diagnose AD based on eye-tracking behaviors. Visual attention, as a typical eye-tracking behavior, is of great clinical value in detecting cognitive abnormalities in AD patients. To better analyze the differences in visual attention between AD patients and normals, we first conducted a 3D comprehensive visual task on a noninvasive eye-tracking system to collect visual attention heatmaps. Then a multilayered comparison convolutional neural network (MC-CNN) is proposed to distinguish the visual attention differences between AD patients and normals. In MC-CNN, the multilayered feature representations of heatmaps were obtained by hierarchical residual blocks to better encode eye movement behaviors, which were further integrated into a distance vector to benefit the comprehensive visual task. From evaluation, MC-CNN can distinguish AD patients from normals with 0.84 accuracy, 0.86 recall, 0.82 precision, 0.83 F1-score and 0.90 area under the curve (AUC). The above results demonstrate the effectiveness of the proposed MC-CNN in AD diagnosis based on the comprehensive 3D visual task.
Fangyu Zuo, Peiguang Jing, Jinglin Sun, Jizhong Duan, Yu Liu 0004
IEEE J. Biomed. Health Informatics2
2024 Deep Multi-Modal Hashing With Semantic Enhancement for Multi-Label Micro-Video Retrieval
abstract
The pressing need for low storage and high efficiency has significantly propelled the advancement of deep hashing techniques in the realm of large-scale search and retrieval tasks. As one of the most prevailing forms of user-generated contents, micro-videos usually represent more complicated multi-modal behaviors that are further challenged in multi-label retrieval. Existing multi-modal hashing methods tend to prioritize the complementarity and consistency in multi-modal fusion, while neglecting the completeness problem. In this paper, we propose a deep multi-modal hashing with semantic enhancement (DMHSE) method that effectively integrates complete multi-modal representation learning with discriminative binary coding by means of collaboration between two distinct encoders, FoldCoder and HashCoder. FoldCoder translates latent multi-modal representation learning to a degradation process through mimicking data transmitting. Further, it incorporates a prompt learning paradigm to maximize the utilization of multi-label semantics for guiding representation learning. HashCoder combines pairwise and central constraints to ensure more discriminative hashing results. Pairwise constraint preserves the original local relevance structure, while central constraint tackles the problem of semantic ambiguity in multi-label data by leveraging the global label distribution. Experimental results demonstrate that DMHSE achieves superior performance in multi-label micro-video retrieval tasks.
Peiguang Jing, Haoyi Sun, Liqiang Nie, Yun Li 0006, Yuting Su 0001
IEEE Trans. Knowl. Data Eng.1
2024 SADCMF: Self-Attentive Deep Consistent Matrix Factorization for Micro-Video Multi-Label Classification
abstract
Currently, there is a growing scholarly and industrial interest in micro-video-centric research. Within these domains, multi-label learning has emerged as a fundamental yet attractive subject. Existing methods primarily place emphasis on feature representations of individual micro-videos, while neglecting latent interdependencies between instance and label domains. To address this problem, in this paper, we propose a novel self-attentive deep consistent matrix factorization (SADCMF) method, which jointly explores dualdomain hierarchical representations and their inherent dependencies for micro-video multi-label classification. Specifically, SADCMF includes three primary characteristics: 1) A dualdomain deep collaborative factorization module is developed to explore the first-stage representations of instance features and the discriminative embeddings of label semantics in a mutually beneficial manner. 2) A correlation-driven selfattentive factorization module is devised to acquire the labelaware attentive outputs, which are further combined with original features through a residual structure to enrich the second-stage feature representations. 3) A dual-stream representation consistency module ensures the unidirectional and bidirectional representation consistency, meanwhile, narrows the discrepancies between the two-stage representations for improving the generalization ability of our method. Extensive experiments conducted on two publicly available micro-video multi-label datasets demonstrate its superior performance in comparison with state-of-the-art methods.
Fugui Fan, Peiguang Jing, Liqiang Nie, Haoyu Gu, Yuting Su 0001
IEEE Trans. Multim.2
2024 Dual-Domain Aligned Deep Hierarchical Matrix Factorization Method for Micro-Video Multi-Label Classification
abstract
Recently, with the growing popularity of micro-videos, multi-label learning has attracted increasing attention due to its potential commercial value in different scenarios. However, existing methods place more emphasis on the alignment between explicit semantics and visual features, while neglecting the exploration of interactions at fine-grained semantic levels. To address this problem, we propose a novel dual-domain aligned deep hierarchical matrix factorization (DADHMF) method for micro-video multi-label classification. Specifically, we construct a dual-stream deep matrix factorization framework to explore implicit hierarchical semantics and corresponding intrinsic feature representations in top-down and bottom-up ways, respectively. On this basis, we leverage the intralayer alignment strategy to narrow the semantic gap between label and instance domains by introducing adaptive semantic-aware embeddings. Moreover, we further utilize the inverse covariance estimation module to automatically capture latent semantic correlations, and project the structural information into the semantic-aware embeddings to ensure the stability of the intralayer alignment. Extensive experiments on two available micro-video multi-label datasets demonstrate that our proposed method outperforms the state-of-the-art methods.
Fugui Fan, Yuting Su 0001, Liqiang Nie, Peiguang Jing, Daozheng Hong, Yu Liu 0004
IEEE Trans. Multim.4
2024 Multimodal Progressive Modulation Network for Micro-Video Multi-Label Classification
abstract
Micro-videos, as an increasingly popular form of user-generated content (UGC), naturally include diverse multimodal cues. However, in pursuit of consistent representations, existing methods neglect the simultaneous consideration of exploring modality discrepancy and preserving modality diversity. In this paper, we propose a multimodal progressive modulation network (MPMNet) for micro-video multi-label classification, which enhances the indicative ability of each modality through gradually regulating various modality biases. In MPMNet, we first leverage a unimodal-centered parallel aggregation strategy to obtain preliminary comprehensive representations. We then integrate feature-domain disentangled modulation process and category-domain adaptive modulation process into a unified framework to jointly refine modality-oriented representations. In the former modulation process, we constrain inter-modal dependencies in a latent space to obtain modality-oriented sample representations, and introduce a disentangled paradigm to further maintain modality diversity. In the latter modulation process, we construct global-context-aware graph convolutional networks to acquire modality-oriented category representations, and develop two instance-level parameter generators to further regulate unimodal semantic biases. Extensive experiments on two micro-video multi-label datasets show that our proposed approach outperforms the state-of-the-art methods.
Peiguang Jing, Fugui Fan, Yun Li 0006, Yuting Su 0001
IEEE Trans. Multim.1
2024 VMemNet: A Deep Collaborative Spatial-Temporal Network With Attention Representation for Video Memorability Prediction
abstract
Video memorability measures the degree to which a video is remembered by different viewers and has shown great potential in various contexts, including advertising, education, and health care. While extensive research has been conducted on image memorability, the study of video memorability is still in its early stages. Existing methods in this field primarily focus on coarse-grained spatial feature representation and decision fusion strategies, overlooking the crucial interactions between spatial and temporal domains. Therefore, we propose an end-to-end collaborative spatial-temporal network called VMemNet, which incorporates targeted attention mechanisms and intermediation fusion strategies. This enables VMemNet to capture the intricate relationships between spatial and temporal information and uncover more elements of memorability within video visual features. VMemNet integrates spatially and semantically guided attention modules into a dual-stream network architecture, allowing it to simultaneously capture static local cues and dynamic global cues in videos. Specifically, the spatial attention module is used to aggregate more memorable elements from spatial locations, and the semantically guided attention module is used to achieve semantic alignment and intermediate fusion of the local and global cues. In addition, two types of loss functions with complementary decision rules are associated with the corresponding attention modules to guide the training process of the proposed network. Experimental results obtained on a publicly available dataset verify that the proposed VMemNet approach outperforms all current single- and multi-modal methods in terms of video memorability prediction.
Wei Lu 0026, Jiaze Han, Peiguang Jing, Yu Liu 0004, Yuting Su 0001
IEEE Trans. Multim.4
2024 Multimodal Attentive Representation Learning for Micro-video Multi-label Classification
abstract
As one of the representative types of user-generated contents (UGCs) in social platforms, micro-videos have been becoming popular in our daily life. Although micro-videos naturally exhibit multimodal features that are rich enough to support representation learning, the complex correlations across modalities render valuable information difficult to integrate. In this paper, we introduced a multimodal attentive representation network (MARNET) to learn complete and robust representations to benefit micro-video multi-label classification. To address the commonly missing modality issue, we presented a multimodal information aggregation mechanism module to integrate multimodal information, where latent common representations are obtained by modeling the complementarity and consistency in terms of visual-centered modality groupings instead of single modalities. For the label correlation issue, we designed an attentive graph neural network module to adaptively learn the correlation matrix and representations of labels for better compatibility with training data. In addition, a cross-modal multi-head attention module is developed to make the learned common representations label-aware for multi-label classification. Experiments conducted on two micro-video datasets demonstrate the superior performance of MARNET compared with state-of-the-art methods.
Peiguang Jing, Xianyi Liu, Yun Li 0006, Yu Liu 0004, Yuting Su 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 A Novel Channel Pruning Approach based on Local Attention and Global Ranking for CNN Model Compression
abstract
Channel pruning facilitates the acceleration and deployment of convolutional neural networks on resource-constrained devices. Nevertheless, existing related methods mainly focus on the importance of an individual channel, neglecting the intra-layer relationship and inter-layer influence. In this paper, we propose a novel local attention and global ranking (LAGR) method for channel pruning. Specifically, we first introduce the attention mechanism to explore the local correlation between channels of the intra-layer. On this basis, we evaluate the global ranking of all channels across the network by the normalization operation. Besides, we introduce a noisy training strategy in the pre-training stage to ensure a balanced weight distribution. Extensive experiments conducted on three representative networks, including VGGNet, GoogLeNet, and ResNet, have demonstrated the superior performance of the proposed method in comparison with several state-of-the-art methods.
Wei Lu 0026, Peiguang Jing, Jinghui Chu, Fugui Fan
ICME3
2023 StyleEDL: Style-Guided High-order Attention Network for Image Emotion Distribution Learning
abstract
Emotion distribution learning has gained increasing attention with the tendency to express emotions through images. As for emotion ambiguity arising from humans' subjectivity, substantial previous methods generally focused on learning appropriate representations from the holistic or significant part of images. However, they rarely consider establishing connections with the stylistic information although it can lead to a better understanding of images. In this paper, we propose a style-guided high-order attention network for image emotion distribution learning termed StyleEDL, which interactively learns stylistic-aware representations of images by exploring the hierarchical stylistic information of visual contents. Specifically, we consider exploring the intra- and inter-layer correlations among GRAM-based stylistic representations, and meanwhile exploit an adversary-constrained high-order attention mechanism to capture potential interactions between subtle visual parts. In addition, we introduce a stylistic graph convolutional network to dynamically generate the content-dependent emotion representations to benefit the final emotion distribution learning. Extensive experiments conducted on several benchmark datasets demonstrate the effectiveness of our proposed StyleEDL compared to state-of-the-art methods. The implementation is released at: https://github.com/liuxianyi/StyleEDL.
Peiguang Jing, Xianyi Liu, Yinwei Wei, Liqiang Nie, Yuting Su 0001
ACM Multimedia1
2023 Internet of Things for Diagnosis of Alzheimer's Disease: A Multimodal Machine Learning Approach Based on Eye Movement Features
abstract
Alzheimer’s disease (AD) is a degenerative neurological disease that occurs in the elderly with typical symptoms of decline in cognition, manifested by eye movement behaviors. The key to AD treatment requires early detection of cognitive impairment, which relies on frequent medical screening. This article proposes an Internet of Things (IoT) architecture constructed with eye-tracker (ET) nodes and cloud-based diagnosis enabled by machine learning (ML), which can provide convenient screening of oculomotor abnormalities and automatic identification of early-stage AD. The bespoke ET nodes collect 3-D oculomotor responses from diverse stereo video stimulation trials and transmit data into a dedicated multimodal ML (MMML) algorithm in the cloud. The algorithm incorporates multimodal features extracted from diverse oculomotor types to enhance the classification accuracy and optimized data dimension reduction in feature fusion to improve the classifier’s performance. From evaluation, the proposed method can distinguish AD patients from the control group with 86% accuracy (ACC), 78% true positive rate (TPR), and 90% positive predictive value (PPV). The results confirm the effectiveness of our MMML algorithm in AD diagnosis with the fusion of multimodal oculomotor features and prove the feasibility of our IoT-powered eye-tracking solution for AD screening.
Yunpeng Yin, Han Wang 0039, Jinglin Sun, Peiguang Jing, Yu Liu 0004
IEEE Internet Things J.5
2023 Self-supervised deep partial adversarial network for micro-video multimodal classification
Yun Li 0006, Shuyi Liu, Peiguang Jing
Inf. Sci.4
2023 Edge-aware object pixel-level representation tracking
Peiguang Jing, Zijian Huang 0010, Jing Liu 0002, Jiexiao Yu
J. Vis. Commun. Image Represent.1
2023 A Multimodal Aggregation Network With Serial Self-Attention Mechanism for Micro-Video Multi-Label Classification
abstract
Currently, micro-videos have attracted increasing attention due to their unique properties and great commercial value. Considering that micro-videos naturally incorporate multimodal information, a powerful representation method for distinct joint multimodal representations is essential for real applications. Inspired by the potential of attention neural network architectures over various tasks, we propose a multimodal aggregation network (MANET) with a serial self-attention mechanism to perform tasks of micro-video multi-label classification. Specifically, we first propose a parallel content-dependent graph neural networks (CDGNN) module, which explores category-related embeddings of micro-videos by disentangling category relations into modality-specific and modality-shared category dependency patterns. Then we introduce a serial self-attention (SSA) module to transmit the multimodal information in sequential order, in which an aggregation bottleneck is incorporated to better collect and condense the significant information. Experiments conducted on a large-scale multi-label micro-video dataset demonstrate that our proposed method has achieved competitive results compared with several state-of-the-art methods.
Wei Lu 0026, Peiguang Jing, Yuting Su 0001
IEEE Signal Process. Lett.3
2023 Category-Aware Multimodal Attention Network for Fashion Compatibility Modeling
abstract
Fashion compatibility modeling, which is used to estimate the matching degree of a given set of fashion items, has received increasing attention in recent years. However, existing studies often fail to fully leverage multimodal information or ignore the semantic guidance of clothing categories in elevating the reliability of multimodal information. In this paper, we propose a fashion compatibility modeling approach with a category-aware multimodal attention network, termed as FCM-CMAN. In FCM-CMAN, we focus on enriching and aggregating multimodal representations of fashion items by means of the dynamic representations of categories and a contextual attention mechanism simultaneously. Specifically, considering that category correlations are always dynamic and varied for different fashion items, we design a categorical dynamic graph convolutional network to adaptively learn the semantic correlations between categories. When combined with the multi-layered visual outputs of a convolutional neural network and the surrounding contextual information, multiple content-aware category representations and context-aware attention weights are obtained to better characterize fashion items from different aspects. On this basis, two pieces of aware information are integrated by a multimodal factorized bilinear pooling strategy to generate visual-semantic embeddings, which are further improved by a multi-head self-attention mechanism to capture significant elements related to fashion compatibility. Extensive experiments conducted on the FashionVC and ExpFashion datasets demonstrate the superiority of FCM-CMAN over state-of-the-art methods.
Peiguang Jing, Weili Guan, Liqiang Nie, Yuting Su 0001
IEEE Trans. Multim.1
2023 Learning Dual Low-Rank Representation for Multi-Label Micro-Video Classification
abstract
Currently, with the rapid development of mobile Internet, micro-video has become a prevailing format of user-generated contents (UGCs) on various social media platforms. Several studies have been conducted towards to understanding high-level micro-video semantics, such as venue categorization, memorability, and popularity. However, these approaches supported tasks with only a single output, which exhibited limitations when attempting to use them to resolve tasks with multiple outputs, especially the multi-label micro-video classification. To tackle this problem, in this paper, we propose a dual multi-modal low-rank decomposition (DMLRD) method for multi-label micro-video classification tasks. To learn more comprehensive micro-video representations, we first learn the low-rank-regularized modality-specific and modality-shared components by considering the consistency and the complementarity among modalities simultaneously. Meanwhile, the less descriptive power of each modality aroused by inherent properties can be solved to a certain extent. To obtain unseen label representations, we next construct a sparsity-regularized multi-matrix normal estimation term to jointly encode the latent relationship structures among labels and dimensions. Experiments on two datasets demonstrate the effectiveness of our proposed method over the state-of-art methods.
Wei Lu 0026, Liqiang Nie, Peiguang Jing, Yuting Su 0001
IEEE Trans. Multim.4
2023 Exploiting Low-Rank Latent Gaussian Graphical Model Estimation for Visual Sentiment Distributions
abstract
Currently, an increasing number of applications and services has encouraged users to openly express their emotions via images. Unlike visual sentiment classification, visual sentiment distribution learning exploits the overall distribution to represent the relative importance of sentiment labels. Considering that most relevant studies have failed to completely model correlation structures or explicitly apply them to unknown instances, in this paper, we proposed a low-rank latent Gaussian graphical model estimation (LGGME) method for visual sentiment distribution learning tasks. There are three main characteristics of LGGME: 1) an integrated inverse covariance matrix whose parameters characterize the latent correlation structures between and within features and sentiments is estimated based on the sparse Gaussian graphical model; 2) a multivariate normal assumption is assigned on the concatenated latent feature representations and the estimated sentiment distributions instead of the original observations for a reasonable surrogate; and 3) the latent feature representations are projected from a low-rank subspace, which is also available for unseen instances, and the estimated sentiment distributions are evaluated by KL divergence to ensure a suitable setting for distribution learning. We further developed an effective optimization algorithm based on the alternating direction method of multipliers (ADMM) for our objective function. The experimental results obtained on three publicly available datasets demonstrate the superiority of our proposed method.
Yuting Su 0001, Peiguang Jing, Liqiang Nie
IEEE Trans. Multim.3
2022 Video frame deletion detection based on time-frequency analysis
Yuting Su 0001, Peiguang Jing
J. Vis. Commun. Image Represent.3
2022 Residual-Guided Multiscale Fusion Network for Bit-Depth Enhancement
abstract
Bit-depth enhancement (BDE) is a challenging task due to stubborn false contour artifacts and disappeared detailed information. Given the mixture of structural distortions and real edges in low bit-depth (LBD) images, both large and small receptive fields (RFs) are critical for BDE tasks. However, even powerful state-of-the-art CNN-based methods can hardly capture sufficient LBD features under multiple RFs. This paper proposes a residual-guided multiscale fusion network (RMFNet) to explore multiscale features in a residual manner. We find that the shuffling operation provides desired multiscale inputs for effectively distinguishing false contours from real edges without any loss of information. Therefore, we shuffle LBD images to multiple scales and then fully extract residual features under different RFs with corresponding subnets. To facilitate interscale guidance from the global context to the local context, we progressively transfer the encoded residual features between adjacent subnets from top to bottom. We further propose a dual-branch depthwise group fusion (DDGF) module to fully capture inter- and inner correlations of multiscale features with fewer parameters. Finally, extensive experiments show that our algorithm achieves excellent performance improvement both quantitatively and qualitatively, verifying its effectiveness.
Jing Liu 0002, Xin Wen 0017, Weizhi Nie, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Tripartite Graph Regularized Latent Low-Rank Representation for Fashion Compatibility Prediction
abstract
In recent years, an increasing online shopping demand has greatly promoted the innovation and development of the fashion industry. Visual fashion analysis has become a prospective research topic in computer vision and multimedia fields. Among these studies, fashion compatibility analysis is required in many real applications, such as fashion recommendation, matching, and retrieval. However, learning fashion compatibility is nontrivial, not only due to the uncertain and sparse dependencies among fashion items but also the latent and mutual associations among multiple factors such as color, texture, style, and functionality. To better predict fashion compatibility, in this paper, we proposed a tripartite graph regularized latent low-rank representation method, named TGRLLR, for fashion compatibility prediction. In TGRLLR, to learn more low-dimensional and effective representations, we considered the latent low-rank representation by decomposing the original feature matrix in both the column and row directions to tackle the problem of insufficient observations. On this basis, we simultaneously exploited different regularization strategies to encode the structured correlations among features, the high-order relationships among items, and the geometrical structures of outfits for more informative representations. Extensive experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed method compared with state-of-the-art methods.
Peiguang Jing, Jing Zhang 0038, Liqiang Nie, Jing Liu 0002, Yuting Su 0001
IEEE Trans. Multim.1
2021 Joint Co-Attention And Co-Reconstruction Representation Learning For One-Shot Object Detection
abstract
One-shot object detection aims to detect all candidate instances in a target image whose label class is unavailable in training, and only one labeled query image is given in testing. Nevertheless, insufficient utilization of the only known sample is one significant reason causing the performance degradation of current one-shot object detection models. To tackle the problem, we develop joint co-attention and co-reconstruction (CoAR) representation learning for one-shot object detection. The main contributions are described as follows. First, we propose a high-order feature fusion operation to exploit the deep co-attention of each target-query pair, which aims to enhance the correlation of the same class. Second, we use a low-rank structure to reconstruct the target-query feature in channel level, which aims to remove the irrelevant noise and enhance the latent similarity between the region proposals in target image and the query image. Experiments on both PASCAL VOC and MS COCO datasets demonstrate that our method outperforms previous state-of-the-art algorithms.
Jinghui Chu, Peiguang Jing, Wei Lu 0026
ICIP3
2021 Joint nuclear- and ℓ2, 1-norm regularized heterogeneous tensor decomposition for robust classification
Peiguang Jing, Yuting Su 0001
Neurocomputing1
2021 Learning robust affinity graph representation for multi-view clustering
Peiguang Jing, Yuting Su 0001, Zhengnan Li, Liqiang Nie
Inf. Sci.1
2021 Deep low-rank matrix factorization with latent correlation estimation for micro-video multi-label classification
Yuting Su 0001, Junyu Xu, Daozheng Hong, Fugui Fan, Jing Zhang 0038, Peiguang Jing
Inf. Sci.6
2021 A Dual Rank-Constrained Filter Pruning Approach for Convolutional Neural Networks
abstract
Filter pruning has attracted increasing attentions to compress and accelerate the convolutional neural networks (CNNs) on computationally restricted devices. Existing related methods mainly focus on independently leveraging the spatial information of individual filters while ignoring the inner correlation among filters. In this letter, we propose a dual rank-constrained filter pruning approach for convolutional neural networks, in which the representation, clustering, and identification of representative filters are integrated into an adaptive graph regularization framework. Particularly, the proposed approach utilizes the low-rank constraint to capture the low-dimensional intrinsic representations of filters for the adaptive affinity graph construction and clustering. It is noteworthy that the original filters are projected as points on Grassmann manifold for geometrical structure preserving. Meanwhile, the high-rank constraint is employed to select the most informative filter for representing the cluster. Experimental results on CIFAR-10 dataset show that the proposed approach achieves competitive results compared with the state-of-the-art methods by 87.1% in VGGNet-16, 52.4% in GoogLeNet and 72.9% in ResNet-56 in terms of model parameters compression.
Fugui Fan, Yuting Su 0001, Peiguang Jing, Wei Lu 0026
IEEE Signal Process. Lett.3
2021 Learning Low-Rank Sparse Representations With Robust Relationship Inference for Image Memorability Prediction
abstract
Image memorability prediction aims to estimate the degree to which an image will be remembered by observers. Generally, the core problem in image memorability prediction is how to obtain effective representations to characterize the visual content of an image. In contrast to existing methods, which focus more on exploring the factors that make images memorable, in this paper, we first propose a general framework for learning joint low-rank and sparse principal feature representations, called the LSPFR framework, to obtain the lowest-rank intrinsic representation for image memorability prediction. By considering the joint optimization of the nuclear and$\ell _{1}$-norms, the global low-rank structure and the local patterns embedded in data can be exploited to make the learned features more robust and informative. To improve our framework based on the exploitation of sample relationship structure information, we present an extended version of LSPFR, named E-LSPFR, in which the underlying relationship structure matrix is inferred through a negative log-likelihood term with a sparsity constraint. The results of experiments conducted on four publicly available datasets confirm the superior performance of our proposed approaches.
Peiguang Jing, Yuechen Shang, Liqiang Nie, Yuting Su 0001, Jing Liu 0002, Meng Wang 0001
IEEE Trans. Multim.1
2021 Market2Dish: Health-aware Food Recommendation
abstract
With the rising incidence of some diseases, such as obesity and diabetes, the healthy diet is arousing increasing attention. However, most existing food-related research efforts focus on recipe retrieval, user-preference-based food recommendation, cooking assistance, or the nutrition and calorie estimation of dishes, ignoring the personalized health-aware food recommendation. Therefore, in this work, we present a personalized health-aware food recommendation scheme, namely, Market2Dish, mapping the ingredients displayed in the market to the healthy dishes eaten at home. The proposed scheme comprises three components, namely, recipe retrieval, user health profiling, and health-aware food recommendation. In particular, recipe retrieval aims to acquire the ingredients available to the users and then retrieve recipe candidates from a large-scale recipe dataset. User health profiling is to characterize the health conditions of users by capturing the textual health-related information crawled from social networks. Specifically, to solve the issue that the health-related information is extremely sparse, we incorporate a word-class interaction mechanism into the proposed deep model to learn the fine-grained correlations between the textual tweets and pre-defined health concepts. For the health-aware food recommendation, we present a novel category-aware hierarchical memory network–based recommender to learn the health-aware user-recipe interactions for better food recommendation. Moreover, extensive experiments demonstrate the effectiveness of the health-aware food recommendation scheme.
Wenjie Wang 0007, Ling-Yu Duan, Peiguang Jing, Xuemeng Song, Liqiang Nie
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Predicting the popularity of micro-videos via a feature-discrimination transductive model
Yuting Su 0001, Yang Li 0108, Peiguang Jing
Multim. Syst.4
2020 Single image super-resolution via low-rank tensor representation and hierarchical dictionary learning
Peiguang Jing, Weili Guan, Yuting Su 0001
Multim. Tools Appl.1
2020 Low-Rank Regularized Deep Collaborative Matrix Factorization for Micro-Video Multi-Label Classification
abstract
Deep matrix factorization can be regarded as an extension of traditional matrix factorization to help improve applications like social image tag refinement, image retrieval, and face clustering. Toward this tendency, in this letter, we proposed a low-rank regularized deep collaborative matrix factorization (LRDCMF) method to better tackle micro-video multi-label classification tasks. The proposed method aims to collaboratively learn two sets of factor matrices for characterization of latent attributes and two deep representations for instances and labels, respectively. During factorization process, the inverse covariance constraints are exploited to capture the latent correlation structures among latent attributes and labels and the low-dimensional intrinsic deep representations are ensured by further considering low-rank constraints. Moreover, a triplet term that connects instances representations, label representations, and labels is constructed to increase discrimination power of our method. Experimental results conducted on a large-scale micro-video dataset illustrate our model achieves superior performance in comparison with state-of-the-art methods.
Yuting Su 0001, Daozheng Hong, Yang Li 0108, Peiguang Jing
IEEE Signal Process. Lett.4
2020 Low-Rank Regularized Multi-Representation Learning for Fashion Compatibility Prediction
abstract
The currently flourishing fashion-oriented community websites and the continuous pursuit of fashion have attracted the increased research interest of the fashion analysis community. Many studies show that predicting the compatibility of fashion outfits is a nontrivial task due to the difficulty in capturing the implicit patterns affecting fashion compatibility prediction and the complex relationships presented by raw data. To address these problems, in this paper, we propose a transductive low-rank hypergraph regularizer multiple-representation learning framework (LHMRL), whereby we formulate the processes of feature representation and fashion compatibility prediction in a joint framework. Specifically, we first introduce a low-rank regularized multiple-representation learning framework, in which the lowest-rank multiple representations of samples can be learned to characterize samples from different perspectives. In this framework, we maximize the total difference among multiple representations based on Grassmann manifold theory and incorporate a common hypergraph regularizer to naturally encode the complex relationships between fashion items and an outfit. To enhance the representation ability of our model, we then develop a supervised learning term by exploiting two types of supervision information from labeled data. Experiments on a publicly available large-scale dataset demonstrate the effectiveness of our proposed model over the state-of-the-art methods.
Peiguang Jing, Liqiang Nie, Jing Liu 0002, Yuting Su 0001
IEEE Trans. Multim.1
2019 Personalized Capsule Wardrobe Creation with Garment and User Modeling
abstract
Recent years have witnessed a growing trend of building the capsule wardrobe by minimizing and diversifying the garments in their messy wardrobes. Thanks to the recent advances in multimedia techniques, many researches have promoted the automatic creation of capsule wardrobes by the garment modeling. Nevertheless, most capsule wardrobes generated by existing methods fail to consider the user profile, including the user preferences, body shapes and consumption habits, which indeed largely affects the wardrobe creation. To this end, we introduce a combinatorial optimization-based personalized capsule wardrobe creation framework, named PCW-DC, which jointly integrates both garment modeling (\textiti.e., wardrobe compatibility) and user modeling (\textiti.e., preferences, body shapes). To justify our model, we construct a dataset, named bodyFashion, which consists of $116,532$ user-item purchase records on Amazon involving 11,784 users and 75,695 fashion items. Extensive experiments on bodyFashion have demonstrated the effectiveness of our proposed model. As a byproduct, we have released the codes and the data to facilitate the research community.
Xuemeng Song, Fuli Feng, Peiguang Jing, Xin-Shun Xu, Liqiang Nie
ACM Multimedia4
2019 Photo-realistic image bit-depth enhancement via residual transposed convolutional neural network
Yuting Su 0001, Wanning Sun, Jing Liu 0002, Guangtao Zhai, Peiguang Jing
Neurocomputing5
2019 Wearable Computing for Internet of Things: A Discriminant Approach for Human Activity Recognition
abstract
With the rapid development of the wireless sensor network and the continuous improvement of its key technologies, the concept of Internet of Things has been encouraged and extended due to its wide applications in scenarios, such as smart homes and healthcare. Under the background, human activity recognition has drawn great attention in recent years. In this paper, we present a discriminant approach to recognize daily human activities recorded through accelerometer sensor. In the proposed approach, we first use S transform (ST) to extract features, and then introduce a supervised regularization-based robust subspace (SRRS) learning method to learn low-dimensional intrinsic feature representation from the original feature subspace. Particularly, ST has been described as a joint time-frequency representation, which is insensitive to noise. SRRS can learn more robust and discriminative features to reinforce the descriptions of samples while removing noise and redundancy. Experiments are conducted on three publicly available datasets, i.e., wireless sensor data mining, SCUT-NAA, and mHealth demonstrating the superior performance of our proposed scheme compared with state-of-the-art methods.
Wei Lu 0026, Fugui Fan, Jinghui Chu, Peiguang Jing, Yuting Su 0001
IEEE Internet Things J.4
2019 Tensor-driven low-rank discriminant analysis for image set classification
Jing Zhang 0038, Zhengnan Li, Peiguang Jing, Ye Liu 0002, Yuting Su 0001
Multim. Tools Appl.3
2019 A structure-transfer-driven temporal subspace clustering for video summarization
Jing Zhang 0038, Peiguang Jing, Jing Liu 0002, Yuting Su 0001
Multim. Tools Appl.3
2019 Low-rank regularized tensor discriminant representation for image set classification
Peiguang Jing, Yuting Su 0001, Zhengnan Li, Jing Liu 0002, Liqiang Nie
Signal Process.1
2019 A Framework of Joint Low-Rank and Sparse Regression for Image Memorability Prediction
abstract
Image memorability is to measure the degree to which an image is remembered. Generally image memorability prediction involves two steps: feature representation and prediction. Most previous work just focused on addressing the first step by investigating the factors of making an image memorable. They not only lack the use of a learning mechanism in feature representation, but also often neglect the second step. In this paper, we first propose a joint low-rank and sparse regression (JLRSR) framework to address this problem. JLRSR aims to jointly learn: 1) a low-rank projection matrix that enables us to decompose the original data into a component part and an error part and 2) a sparse regression coefficient vector for image memorability prediction. The projection matrix and the regression coefficients are bound by a sparse constraint to make our approach invariant to training samples. Moreover, a graph regularizer is constructed to improve the generalization performance and prevent overfitting. We then extend JLRSR to a multi-view version called Mv-JLRSR by imposing the block-wise constraint to ensure the group effect and the view correlation constraint to eliminate the heterogeneity among views. Experiment results validate the effectiveness of our proposed approaches.
Peiguang Jing, Yuting Su 0001, Liqiang Nie, Huimin Gu, Jing Liu 0002, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 High-Order Temporal Correlation Model Learning for Time-Series Prediction
abstract
Time-series prediction has become a prominent challenge, especially when the data are described as sequences of multiway arrays. Because noise and redundancy may exist in the tensor representation of a time series, we focus on solving the problem of high-order time-series prediction under a tensor decomposition framework and develop two novel multilinear models: 1) the multilinear orthogonal autoregressive (MOAR) model and 2) the multilinear constrained autoregressive (MCAR) model. The MOAR model is designed to preserve as much information as possible from the original tensorial data under orthogonal constraints. The MCAR model is an enhanced version that is developed by replacing orthogonal constraints with an inverse decomposition error term. For both models, we project the original tensor into subspaces spanned by basis matrices to facilitate the discovery of the intrinsic temporal structure embedded in the original tensor. To build connections among consecutive slices of the tensor, we generalize a traditional autoregressive model to tensor form to better preserve the temporal smoothness. Experiments conducted on four publicly available datasets demonstrate that our proposed methods converge within a small number of iterations during the training stage and achieve promising results compared with state-of-the-art methods.
Peiguang Jing, Yuting Su 0001, Chengqian Zhang
IEEE Trans. Cybern.1
2019 BE-CALF: Bit-Depth Enhancement by Concatenating All Level Features of DNN
abstract
There is a growing demand for monitors to provide high-quality visualization with more bits representing each rendered pixel. However, since most existing images and videos are of low bit-depth (LBD), transforming LBD images to visually pleasant high bit-depth (HBD) versions is of significant value. Most existing bit-depth enhancement methods generate unsatisfactory HBD images with annoying false contour artifacts or blurry details, and some algorithms are also time-consuming. To overcome these drawbacks, we propose a bit-depth enhancement framework via concatenating all level features of deep neural networks (DNNs). A novel deep learning network is proposed based on the deep convolutional variational auto-encoders (VAEs), and skip connections that concatenate every two layers are applied to pass low-level and high-level features to consequent layers, easing the gradient vanishing problem. Meanwhile, the proposed network is optimized to generate the residual between original images and its quantized ones, which performs better than recovering HBD images directly. The experimental results show that the proposed algorithm can eliminate false contour artifacts of the recovered HBD images with low time consumption, and can achieve dramatic restoration performance gains compared with state-of-the-art methods both subjectively and objectively.
Jing Liu 0002, Wanning Sun, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001
IEEE Trans. Image Process.4
2019 Spatiotemporal Symmetric Convolutional Neural Network for Video Bit-Depth Enhancement
abstract
In contrast to the high sensitivity of human eyes and rapid development of modern display devices in terms of dynamic range, mainstream multimedia sources are generally at relatively lower bit depths (BDs). Therefore, BD enhancement (BDE), which attempts to transform low-BD multimedia sources into high-BD sources, is considered of significant research value. Current BDE algorithms are based on images rather than videos. However, for massive numbers of videos, temporal continuity among frames should be considered. Thus, in this paper, we propose a spatiotemporal symmetric BDE network for videos based on an encoder-decoder network. Consecutive frames are input into five subnets in the encoder, where the convolutional filters in the temporal symmetric subnets share the same weights to achieve lower model complexity. In addition, symmetric skip connections are introduced between the symmetric convolutional/deconvolutional layers of the encoder/decoder to pass features and alleviate the gradient diffusion problem. The experimental results show that our model can efficiently eliminate false contours and chroma distortions. The model significantly outperforms state-of-the-art image BDE algorithms and single-frame baseline models in terms of PSNR and SSIM.
Jing Liu 0002, Pingping Liu, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001
IEEE Trans. Multim.4
2018 Towards a sparse low-rank regression model for memorability prediction of images
Jinghui Chu, Huimin Gu, Yuting Su 0001, Peiguang Jing
Neurocomputing4
2018 HyperSSR: A hypergraph based semi-supervised ranking method for visual search reranking
Peiguang Jing, Yuting Su 0001, Chuan-Zhong Xu
Neurocomputing1
2018 Low-rank regularized multi-view inverse-covariance estimation for visual sentiment distribution prediction
Anan Liu, Yingdi Shi, Peiguang Jing, Jing Liu 0002, Yuting Su 0001
J. Vis. Commun. Image Represent.3
2018 Graph regularized low-rank tensor representation for feature selection
Yuting Su 0001, Peiguang Jing, Jing Zhang 0038, Jing Liu 0002
J. Vis. Commun. Image Represent.4
2018 Video logo removal detection based on sparse representation
Yuting Su 0001, Liang Zou, Chengqian Zhang, Peiguang Jing, Xuemeng Song
Multim. Tools Appl.5
2018 Structured low-rank inverse-covariance estimation for visual sentiment distribution prediction
Anan Liu, Yingdi Shi, Peiguang Jing, Jing Liu 0002, Yuting Su 0001
Signal Process.3
2018 Low-Rank Regularized Heterogeneous Tensor Decomposition for Subspace Clustering
abstract
This letter proposes a low-rank regularized heterogeneous tensor decomposition (LRRHTD) algorithm for subspace clustering, in which various constrains in different modes are incorporated to enhance the robustness of the proposed model. Specifically, due to the presence of noise and redundancy in the original tensor, LRRHTD seeks a set of orthogonal factor matrices for all but the last mode to map the high-dimensional tensor into a low-dimensional latent subspace. Furthermore, by imposing a low-rank constraint on the last mode, which is relaxed by using a nuclear norm, the lowest rank representation that reveals the global structure of samples is obtained for the purpose of clustering. We develop an effective algorithm based on the augmented Lagrange multiplier to optimize our model. Experiments on two public datasets demonstrate that our method reaches convergence within a small number of iterations and achieves promising results in comparison with the state of the arts.
Jing Zhang 0038, Peiguang Jing, Jing Liu 0002, Yuting Su 0001
IEEE Signal Process. Lett.3
2018 Low-Rank Multi-View Embedding Learning for Micro-Video Popularity Prediction
abstract
Recently, a prevailing trend of user generated content (UGC) on social media sites is the emerging micro-videos. Microvideos afford many potential opportunities ranging from network content caching to online advertising, yet there are still little efforts dedicated to research on micro-video understanding. In this paper, we focus on popularity prediction of micro-videos by presenting a novel low-rank multi-view embedding learning framework. We name it as transductive low-rank multi-view regression (TLRMVR), and it is capable of boosting the performance of micro-video popularity prediction by jointly considering the intrinsic representations of the source and target samples. In particular, TLRMVR integrates low-rank multi-view embedding and regression analysis into a unified framework such that the lowest-rank representation shared by all views not only captures the global structure of all views, but also indicates the regression requirements. The framework is formulated as a regression model and it seeks a set of view-specific projection matrices with low-rank constraints to map multi-view features into a common subspace. In addition, a multi-graph regularization term is constructed to improve the generalization capability and further prevents the overfitting problem. Extensive experiments conducted on a publicly available dataset demonstrate that our proposed method achieve promising results as compared with state-of-the-art baselines.
Peiguang Jing, Yuting Su 0001, Liqiang Nie, Jing Liu 0002, Meng Wang 0001
IEEE Trans. Knowl. Data Eng.1
2017 A spatial-temporal iterative tensor decomposition technique for action and gesture recognition
Yuting Su 0001, Haiyi Wang, Peiguang Jing, Chuan-Zhong Xu
Multim. Tools Appl.3
2017 SnapVideo: Personalized Video Generation for a Sightseeing Trip
abstract
Leisure tourism is an indispensable activity in urban people's life. Due to the popularity of intelligent mobile devices, a large number of photos and videos are recorded during a trip. Therefore, the ability to vividly and interestingly display these media data is a useful technique. In this paper, we propose SnapVideo, a new method that intelligently converts a personal album describing of a trip into a comprehensive, aesthetically pleasing, and coherent video clip. The proposed framework contains three main components. The scenic spot identification model first personalizes the video clips based on multiple prespecified audience classes. We then search for some auxiliary related videos from YouTube according to the selected photos. To comprehensively describe a scenery, the view generation module clusters the crawled video frames into a number of views. Finally, a probabilistic model is developed to fit the frames from multiple views into an aesthetically pleasing and coherent video clip, which optimally captures the semantics of a sightseeing trip. Extensive user studies demonstrated the competitiveness of our method from an aesthetic point of view. Moreover, quantitative analysis reflects that semantically important spots are well preserved in the final video clip.
Peiguang Jing, Yuting Su 0001, Chao Zhang 0014, Ling Shao 0001
IEEE Trans. Cybern.2
2017 Predicting Image Memorability Through Adaptive Transfer Learning From External Sources
abstract
Remembering images is an innate human capability. Camera images are captured by different people under varying environmental conditions, which leads to highly diverse image memorability scores. However, the factors that make an image more or less memorable are unclear, and it remains unknown how we can more accurately predict image memorability by using such factors. In this paper, we propose a novel framework called multiview transfer learning from external sources (MTLES) to predict image memorability. In this framework, we simultaneously leverage different types of visual feature sets and multiple types of predefined image attributes derived from external sources. In particular, to enhance representation ability of visual features, we construct connections between visual feature sets and higher level image attributes by transferring attribute knowledge from external sources. MTLES integrates weak learning through external sources, transfer learning, and multiview consistency loss with different types of feature sets into a joint framework. To better solve this joint optimization problem, we further develop an alternating iterative algorithm to deal with it. Experiments performed on the publicly available LaMem dataset demonstrate the effectiveness of the proposed scheme.
Peiguang Jing, Yuting Su 0001, Liqiang Nie, Huimin Gu
IEEE Trans. Multim.1
2016 Visual search reranking with RElevant Local Discriminant Analysis
Peiguang Jing, Zhong Ji, Yunlong Yu 0001, Zhongfei Zhang
Neurocomputing1
2016 Quality biased multimedia data retrieval in microblogs
Shuhan Qi, Peiguang Jing, Xuan Wang 0002, Liqiang Nie
J. Vis. Commun. Image Represent.2
2016 A Tensor-Driven Temporal Correlation Model for Video Sequence Classification
abstract
The task of video sequence classification plays a critical role in the development of computer vision. Considering this fact, this letter proposes a novel tensor decomposition method called tensor-driven temporal correlation in which general tensors are used as input for video sequence classification. Because distortion and redundancy may exist in the tensor representations of video sequences, we project the original tensor into subspaces spanned by spatial basis matrices in the proposed formulation. Moreover, to better preserve the temporal smoothness between consecutive slices of the tensor, the basis matrices are jointly learned by introducing an autoregressive model. An experiment on the commonly used Cambridge hand-gesture database demonstrates that our proposed method reaches convergence within a small number of iterations during the training stage and achieves promising results compared with state-of-the-art methods.
Jing Zhang 0038, Chuan-Zhong Xu, Peiguang Jing, Chengqian Zhang, Yuting Su 0001
IEEE Signal Process. Lett.3
2013 Ranking Fisher discriminant analysis
Zhong Ji, Peiguang Jing, Tianshi Yu, Yuting Su 0001, Changshu Liu
Neurocomputing2
2013 Rank canonical correlation analysis and its application in visual search reranking
Zhong Ji, Peiguang Jing, Yuting Su 0001, Yanwei Pang
Signal Process.2
2013 Ranking Graph Embedding for Learning to Rerank
abstract
Dimensionality reduction is a key step to improving the generalization ability of reranking in image search. However, existing dimensionality reduction methods are typically designed for classification, clustering, and visualization, rather than for the task of learning to rank. Without using of ranking information such as relevance degree labels, direct utilization of conventional dimensionality reduction methods in ranking tasks generally cannot achieve the best performance. In this paper, we show that introducing ranking information into dimensionality reduction significantly increases the performance of image search reranking. The proposed method transforms graph embedding, a general framework of dimensionality reduction, into ranking graph embedding (RANGE) by modeling the global structure and the local relationships in and between different relevance degree sets, respectively. The proposed method also defines three types of edge weight assignment between two nodes: binary, reconstruction, and global. In addition, a novel principal components analysis based similarity calculation method is presented in the stage of global graph construction. Extensive experimental results on the MSRA-MM database demonstrate the effectiveness and superiority of the proposed RANGE method and the image search reranking framework.
Yanwei Pang, Zhong Ji, Peiguang Jing, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3