Yuting Su 0001

dblp:10/7033-1 · also Yu-Ting Su 0001 · DBLP profile ↗
← Back
200ranked-venue papers
30as first author
88since 2021 · last 2026
0000-0001-5165-204XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 143 · 20 first-author · 68 since 2021Artificial intelligence and machine learning · 43 · 8 first-author · 9 since 2021Computer networks · 9 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 SGoT-R1: Social Graph of Thought Reasoning-Enhanced Multimodal Large Language Model for Harmful Meme Detection
abstract
Internet memes serve as widely distributed multimodal social content that conveys complex ideas through metaphorical expressions, often containing harmful implications that make accurate harmful meme detection an important problem. Reasoning knowledge extracted from large language models plays a crucial role in recent advances in harmful meme detection. However, these methods only perform reasoning analysis on memes from a single opinion, ignoring that memes are essentially products of group consensus, where their true meaning interpretation highly depends on the collision and aggregation process of diverse user viewpoints. To address this problem, we propose a Social Graph of Thought Reasoning Enhancement (SGoTRE) framework for harmful meme detection. The SGoTRE contains three key steps: First, through multi-agent simulation technology, we obtain diverse chains of thought that represent the parsing logic of users from different backgrounds toward memes, authentically restoring the diversity characteristics of group cognition. Second, we construct a Social Graph of Thought (SGoT) that effectively integrates multi-chain reasoning processes and structurally expresses the consensus and diversity of viewpoints among users. Finally, we utilize the SGoT for cognitive distillation, internalizing multi-opinion reasoning logic into a single multimodal large model SGoT-R1 to achieve efficient and interpretable harmful meme detection. Experimental results show that SGoT-R1 significantly improves detection performance on mainstream datasets. Particularly on the most challenging FHM dataset, SGoT-R1 achieves an 8.9% improvement over state-of-the-art models.
Xiuxian Wang, Yuting Su 0001, Wenhui Li 0001, Zhuojun Li, Anan Liu
AAAI2
2026 Multi-stage Superpixel-guided Mamba-based Network for Change Detection
Jing Liu 0002, Yuting Su 0001, Peiguang Jing
ISCAS4
2026 FacDNet: A low-rank factorized diffusion network with dual-U compensated attention for low-light enhancement
Peiguang Jing, Ningyuan Zhao, Lijun Lai, Weiming Wang 0002, Yuting Su 0001
Signal Process.6
2026 Separating Domain-Private Classes for Universal Unsupervised Cross-Domain 3D Model Retrieval
Jiayu Li 0004, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, Zan Gao 0002, Anan Liu
IEEE Trans. Multim.2
2025 Adversarial neighbor perception network with feature distillation for anomaly detection
Yuting Su 0001, Enqi Su, Weiming Wang 0002, Peiguang Jing, Dubuke Ma, Fu Lee Wang
Expert Syst. Appl.1
2025 Dynamic Causal Disentanglement Model for Dialogue Emotion Detection
abstract
Emotion detection is a critical technology extensively employed in diverse fields. While the incorporation of commonsense knowledge has proven beneficial for existing emotion detection methods, dialogue-based emotion detection encounters numerous difficulties and challenges due to human agency and the variability of dialogue content. In dialogues, human emotions tend to accumulate in bursts. However, they are often implicitly expressed. This implies that many genuine emotions remain concealed within a plethora of unrelated words and dialogues. In this paper, we propose a Dynamic Causal Disentanglement Model founded on the separation of hidden variables, which effectively decomposes the content of dialogues and investigates the temporal accumulation of emotions, thereby enabling more precise emotion recognition. First, we introduce a novel Causal Directed Acyclic Graph (DAG) to establish the correlation between hidden emotional information and other observed elements. Subsequently, our approach utilizes pre-extracted personal states and utterance topics as guiding factors for the distribution of hidden variables, aiming to separate irrelevant ones. Specifically, we propose a Dynamic Causal Disentanglement Model to infer the propagation of utterances and hidden variables, enabling the accumulation of emotion-related information throughout the conversation. To guide this disentanglement process, we leverage the GPT4.0 and LSTM networks to extract utterance topics and personal states as observed information. Finally, we test our approach on popular datasets in dialogue emotion detection and relevant experimental results verified the model's superiority.
Yuting Su 0001, Weizhi Nie, Sicheng Zhao, Anan Liu
IEEE Trans. Affect. Comput.1
2025 Multi-Scale Spatial-Temporal Transformer for Meteorological Variable Forecasting
abstract
Frequent occurrences of marine extreme climate and weather events pose significant threats to human life and property, underscoring the practical significance of meteorological data forecasting methods. Notably, significant advancements in meteorological forecasting fields have been achieved by data-driven deep learning techniques, which leverage observed meteorological datasets and employ deep networks to capture complex patterns. However, challenges remain in accurately extracting local details and capturing spatial-temporal correlations when dealing with multiple meteorological forecasting tasks that exhibits diverse temporal and spatial scales. Hence, in this paper, we propose a Multi-Scale Spatial Temporal Transformer (MS-STT) framework to achieve efficient and accurate meteorological data forecasting. Specifically, to achieve more detailed and multi-scale representation of meteorological data, we design the regionally coherent encoding strategy and multi-scale feature aggregation for visual representation. To enhance the multi-scale ability in terms of learning spatial-temporal correlations, we propose a multi-scale spatial-temporal transformer network, which integrates a multi-scale spatial transformer to learn the spatial association between local patches and multi-scale regions and a temporal transformer to learn the temporal dynamic evolution properties. Extensive quantitative and qualitative experiments on three popular spatial temporal forecasting tasks validate the effectiveness of the proposed method. In particular, compared to the representative data-driven deep learning ENSO forecasting method Earthformer, our approach achieves a 3.7% performance improvement with only one-third of the parameters.
Tianbao Li 0001, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, Zhiqiang Wei 0002, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.2
2025 Progressive Contrastive Label Optimization for Source-Free Universal 3D Model Retrieval
abstract
Unsupervised Cross-Domain 3D Model Retrieval (UCD3DMR) has emerged as an effective tool for managing 3D model data recently. However, existing UCD3DMR algorithms typically demand accessibility to source data and cross-domain label consistency, limiting their deployment in real-world industrial scenarios. Therefore, we relax the two demanding constraints and explore to address a newly challenging task, source-free universal 3D model retrieval (SFU3DMR). However, the inaccessibility to source data results in significant label noise in target pseudo-labels, while cross-domain label inconsistency introduces interference from target-private models, presenting tremendous challenges to model transfer. To address these challenges, we propose a novel SFU3DMR algorithm, Progressive Contrastive Label Optimization (PCLO). Specifically, we introduce the Neighbor-based Soft Label Optimization (NSLO) strategy, which refines target pseudo-labels based on the pseudo-label confidence of their nearest neighbors. Additionally, we design the Adaptive Hybrid Label Optimization (AHLO) strategy, which conducts positive label optimization to maximize label semantics for target-common models and executes negative label optimization to minimize label noise for target-private models. Experimental results confirm that the combined NSLO and AHLO strategies effectively refine the target pseudo-labels, and our PCLO achieves state-of-the-art performance for SFU3DMR on two well-established cross-domain benchmarks (MI3DOR and NTU/PSB).
Jiayu Li 0004, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, You Yang 0002, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.2
2025 What and Where: Semantic Grasping and Contextual Scanning for Moment Retrieval and Highlight Detection
abstract
The current surge in video content highlights the tasks of moment retrieval (MR) and highlight detection (HD), which involve localizing video segments of events and predicting clip-wise saliency scores based on text queries. The recent methods, while effective, may overlook two aspects: 1) Multimodal features often show weak alignment from frozen encoders, hindering thorough semantic exploration of video clips through fine-grained cross-modal interaction. 2) Due to the absence of significant distinction between adjacent video clips, it is challenging for clip-level context modeling to accurately locate query-relevant content. To mitigate these gaps and inspired by the human routine in understanding visual events, we propose a progressive framework dubbed “what and where” to initially grasp the aligned semantics of each video clip, and then proceed to scan moment-level contextual features temporally to identify events matching the query. In the ‘what’ stage, to enable explicit alignment of modal features and achieve a thorough semantic understanding, we firstly devise the Initial Semantic Projection (ISP) loss to bring closer different modal features with similar semantics. Additionally, we develop a Clip Semantic Mining module to deeply mine the relevance of these identified semantics to the specific query (at both word- and sentence-level). In the ‘where’ stage, to enhance feature distinctiveness, we design a Multi-Context Perception module that models moment-level context. It includes an Event Context (EC) branch and a Chronological Context (CC) branch, focusing on possible query-relevant event moments and temporal moments of various lengths. Finally, extensive experiments validate the state-of-the-art performance of our W2W model on three benchmark datasets without additional pre-training. Codes are available athttps://github.com/TJUMMG/W2W.
Jing Liu 0002, Zhuo He, Weizhi Nie, Zongbing Zhang, Yuting Su 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 KA-MIN: Knowledge-Aware Multimodal Interaction Network for Emotion Recognition in Conversation
abstract
Emotion recognition in conversations (ERC) has garnered significant attention for its critical role in human-computer interaction systems. ERC benefits from multimodal data, which offers diverse perspectives on emotional states, and commonsense knowledge (CSK), which enriches the context by incorporating real-world understanding of human behavior. However, existing ERC studies have not fully exploited the potential of multimodal-CSK interactions for complementary information learning from these sources. To address this, we innovatively propose a Knowledge-Aware Multimodal Interaction Network (KA-MIN). KA-MIN is designed to capture complementary emotional information from CSK-multimodal interactions, thereby facilitating the ERC task. To achieve this, KA-MIN begins by combining six relation types of CSK, leveraging their differences between multimodal emotional information. The fused CSK features are then refined to incorporate context and emotional information using multimodal contextual guidance. Subsequently, we construct a novel knowledge-aware multimodal graph structure that allows the CSK information to interact with multimodal information, leading to more comprehensive multimodal and context modeling. During the graph learning process, the CSK-multimodal interactions capture the complementary emotional information between CSK and multimodal features. Finally, we dynamically fuse the multimodal emotional information with the informative CSK and textual guidance to obtain the final utterance representations, which encompass effective emotional information from both multimodal and CSK features. Extensive experiments on two popular multimodal ERC datasets demonstrate the superiority and effectiveness of the proposed KA-MIN framework.
Minjie Ren, Xiangdong Huang 0002, Jing Liu 0002, Zan Gao 0001, Yuting Su 0001, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.5
2025 Few-Shot In-Context Learning for Implicit Semantic Multimodal Content Detection and Interpretation
abstract
In recent years, the field of explicit semantic multimodal content research makes significant progress. However, research on content with implicit semantics, such as online memes, remains insufficient. Memes often convey implicit semantics through metaphors and may sometimes contain hateful information. To address this issue, researchers propose a task for detecting hateful memes, opening up new avenues for exploring implicit semantics. The hateful meme detection currently faces two main problems: 1) the rapid emergence of meme content makes continuous tracking and detection difficult; 2) current methods often lack interpretability, which limits the understanding and trust in the detection results. To make a better understanding of memes, we analyze the definition of metaphor from social science and identify the three key factors of metaphor: socio-cultural knowledge, metaphorical tenor, and metaphorical representation pattern. According to these key factors, we guide a multimodal large language model (MLLM) to infer the metaphors expressed in memes step by step. Particularly, we propose a hateful meme detection and interpretation framework, which has four modules. We first leverage a multimodal generative search method to obtain socio-cultural knowledge relevant to visual objects of memes. Then, we use socio-cultural knowledge to instruct the MLLM to assess the social-cultural relevance scores between visual objects and textual information, and identify the metaphorical tenor of memes. Meanwhile, we apply a representative interpretation method to provide representative cases of memes and analyze these cases to explore metaphorical representation pattern. Finally, a chain-of-thought prompt is constructed to integrate the output of the above modules, guiding the MLLM to accurately detect and interpret hateful memes. Our method achieves state-of-the-art performance on three hateful meme detection benchmarks and performs better than supervised training models on the hateful meme interpretation benchmark.
Xiuxian Wang, Lanjun Wang, Yuting Su 0001, Hongshuo Tian, Guoqing Jin, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.3
2025 AMDANet: Augmented Multiscale Difference Aggregation Network for Image Change Detection
abstract
The field of remote sensing image change detection (CD) has made significant improvements with the rapid development of deep learning techniques. However, current methods often inadequately utilize difference features of bitemporal images, resulting in biased focus and insensitivity to change information. Furthermore, the classic challenges of pseudo-CD and edge recognition in complex scenes have also weakened CD performance. In this article, we propose an augmented multiscale difference aggregation network (AMDANet) for image CD, which incorporates a difference feature extractor (DFE) within a Siamese feature extractor to perceive changes by capturing differences between bitemporal features. To address the issue of biased focus, we propose a hierarchical feature aggregator (HFA) that captures intrascale interactions in parallel for multigranularity change perception while personalizing coarse-grained and fine-grained features to highlight the attention to change regions. To deepen the perception of complex dependency relationships, we further design an O-shape feature augmentor (OFA) that leverages an information feedback loop to achieve precise alignment of multigranularity features. The integration of information across different granularities improves the recognition of pseudo-changes and edges. Experimental results on three publicly available datasets demonstrate the superiority of AMDANet over current state-of-the-art (SOTA) methods. Our source code will be publicly available athttps://github.com/mp-st/AMDANet.
Yuting Su 0001, Peng Ma, Weiming Wang 0002, Shaochu Wang, Yun Li 0006, Peiguang Jing
IEEE Trans. Geosci. Remote. Sens.1
2025 Learning to Generate Realistic Images for Bit-Depth Enhancement via Camera Imaging Processing
abstract
With the prevalence of advanced displays devices, many attempts have been successfully made in bit-depth enhancement (BDE) to restore the low bit-depth (LBD) images to visually pleasant high bit-depth (HBD) images. However, most methods are still far from satisfactory when addressing real-world LBD images owing to their heavy dependence on LBD-HBD data pairs through direct pixel quantization. Therefore, in this paper, we propose a novel network dubbed RealGAN to generate real-world LBD images by simulating the complex quantization procedure in camera imaging process. Particularly, we design a two-mode differentiable quantization block embedded in the synthesis network facilitating adaptively simulation of the complicated quantization distortions. Furthermore, a simple residual group network is proposed in order to learn the distribution of degradation and non-linear processing in the Image Signal Processing (ISP) pipeline. In the absence of paired HBD and LBD data, the synthesis model is trained end-to-end within the generative adversarial framework using non-paired LBD and HBD images. Finally, we demonstrate that a series of BDE models can benefit from the proposed synthetic dataset and exhibit improved visual quality with sharper edges and finer textures on real-world scenes compared with the original versions trained on directly quantized LBD-HBD pairs.
Jing Liu 0002, Huiyu Duan, Yuting Su 0001, Guangtao Zhai
IEEE Trans. Multim.5
2025 Aggregate and Discriminate: Pseudo Clips-Guided Boundary Perception for Video Moment Retrieval
abstract
Video moment retrieval (VMR) aims to localize a video segment in an untrimmed video that is semantically relevant to a language query. The challenge of this task lies in effectively aligning the intricate and information-dense video modality with the succinctly summarized textual modality, and further localizing the starting and ending timestamps of the target moments. Previous works have attempted to achieve multi-granularity alignment of video and query in a coarse-to-fine manner, yet these efforts still fall short in addressing the inherent disparities in representation and information density between videos and queries, leading to modal misalignments. In this paper, we propose a progressive video moment retrieval framework, initially retrieving the most relevant and irrelevant video clips to the query as semantic guidance, thereby bridging the semantic gap between video modality and language modality. Futhermore, we introduce a pseudo clips guided aggregation module to aggregate densely relevant moment clips closer together and propose a discriminative boundary-enhanced decoder with the guidance of pseudo clips to push the semantically confusing proposals away. Extensive experiments on the Charades-STA, ActivityNet Captions and TACoS datasets demonstrate that our method outperforms existing methods.
Jing Liu 0002, Zongbing Zhang, Yuting Su 0001, Bing Yang 0003, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Multim.3
2025 Multimodal Dual-Graph Collaborative Network With Serial Attentive Aggregation Mechanism for Micro-Video Multi-Label Classification
abstract
The increasing commercial value of micro-videos has spurred a rising demand for grasping their contents. The abundant multimodal cues in micro-videos exhibit substantial potential in enhancing content comprehension. However, effectively harnessing the collaborative characteristics across different modalities remains a significant challenge, especially in multi-label scenarios due to inconsistent behaviors regarding label correlations. To better tackle this issue, in this paper, we first introduce a multimodal dual-graph collaborative network with serial attentive aggregation mechanism (MDGCN) for micro-video multi-label classification. In MDGCN, we exploit an asymmetric encoder-decoder framework, which incorporates multiple parallel encoders with complementary representations and a decoder to ensure the completeness of encoded results. Meanwhile, an adversarial constraint is used to ensure individual differences prominently featured within each modality. Furthermore, considering the inconsistency of label correlations across various modalities, we then construct a serial attentive graph convolutional network that employs an interactive dual-graph attention paradigm to sequentially integrate multimodal representations and dynamically explore label correlations. The experiments conducted on two datasets demonstrate that our proposed method outperforms state-of-the-art approaches.
Wei Lu 0026, Peiguang Jing, Weiming Wang 0002, Yuting Su 0001
IEEE Trans. Multim.5
2025 Weakly Supervised Referring Video Object Segmentation With Object-Centric Pseudo-Guidance
abstract
Referring video object segmentation (RVOS) is an emerging task for multimodal video comprehension while the expensive annotating process of object masks restricts the scalability and diversity of RVOS datasets. To relax the dependency on expensive mask annotations and take advantage from large-scale partially annotated data, in this paper, we explore a novel extended RVOS task, namely weakly supervised referring video object segmentation (WRVOS), which employs multiple weak supervision sources, including object points and bounding boxes. Correspondingly, we propose a unified WRVOS framework. Specifically, an object-centric pseudo mask generation method is introduced to provide effective shape priors for the pseudo guidance of spatial object location. Then, a pseudo-guided optimization strategy is proposed to effectively optimize the object outlines in terms of spatial location and projection density with a multi-stage online learning strategy. Furthermore, a multimodal cross-frame level set evolution method is proposed to iteratively refine the object boundaries considering both temporal consistency and cross-modal interactions. Extensive experiments are conducted on four publicly available RVOS datasets, including A2D Sentences, J-HMDB Sentences, Ref-DAVIS, and Ref-YoutubeVOS. Performance comparison shows that the proposed method achieves state-of-the-art performance in both point-supervised and box-supervised settings.
Weikang Wang 0002, Yuting Su 0001, Jing Liu 0002, Wei Sun 0029, Guangtao Zhai
IEEE Trans. Multim.2
2025 Disentangled Denoising and Counterfactual Balance for Multimodal Recommendation
abstract
Recently, graph convolutional network-based dual-view multimodal recommendation methods have achieved great success. They extract multimodal and behavior features based on item-item and user-item graphs, respectively. However, they still have two- fold limitations. First, the relevance between multimodal semantics and user preferences is ignored, resulting in the propagation and coupling of preference-irrelevant noise. Second, the direct use of uneven factual user-item graphs is suboptimal, as both redundant noisy edges and missing positive interaction edges impair recommendations. To solve the above issues, we propose aDisentAngled deNoising andCounterfactual balancEmethod for multimodal recommendation, dubbed asDANCE. Specifically, for multimodal features, we explicitly disentangle them into preference-relevant and preference-irrelevant representations, to absorb and discard irrelevant noise via the latter. An orthogonal regularization and a contrastive learning task on preference relevance score prediction are proposed as the dual safeguard to prevent preference-relevant representations from encoding irrelevant noise. For behavior feature extraction, we construct a balanced user-item graph by integrating factual and counterfactual graphs. In this process, we pre-train a behavior simulator to build the counterfactual graph with full interactions. Top-$K$sampling is adopted to omit noisy edges and add missing edges in the graph. The final recommendation is performed upon the fused representation of preference-relevant multimodal and behavior representations. Extensive experiments on three public datasets verify the power of our DANCE.
Xin Wen 0017, Weizhi Nie, Jing Liu 0002, Yuting Su 0001, Anan Liu
IEEE Trans. Multim.5
2025 PADNet: Progressive-Difference-Aware Feature Reconstruction Mechanism for Anomaly Detection
Peiguang Jing, Weiming Wang 0002, Fu Lee Wang, Yuting Su 0001
IEEE Trans. Multim.5
2025 MarkPlugger: Generalizable Watermark Framework for Latent Diffusion Models Without Retraining
abstract
Today, the family of latent diffusion models (LDMs) has gained prominence for its high quality outputs and scalability. This has also raised security concerns on social media, as malicious users can create and disseminate harmful content. Existing approaches typically involve training specific components or entire generative models to embed a watermark in generated images for traceability and responsibility. However, in the fast-evolving era of AI-generated content (AIGC), the rapid iteration and modification of LDMs makes retraining with watermark models costly. To address the problem, we proposeMarkPlugger, a generalizable plug-and-play watermark framework without LDM retrain. In particular, to reduce the disturbance of the watermark on the semantic of the generated image, we try to identify a watermark representation that is approaching orthogonal to the semantic in latent space, and the theoretical study shows that we can achieve approximate orthogonality in a high-dimensional space. Moreover, the offset through an additive fusion strategy for the watermark and the semantic is also bounded. Without modifying any components of the LDMs, we embed diverse watermarks in latent space, adapting to the denoising process. Our experimental findings reveal that our method effectively harmonizes image quality and watermark recovery rate. We also have validated that our method is generalized to multiple official versions and modified variants of LDMs, even without retraining the watermark model. Furthermore, it performs robustly under various attacks of different intensities.
Lanjun Wang, Yuting Su 0001, Anan Liu
IEEE Trans. Multim.3
2024 Graph Disentangled Contrastive Learning with Personalized Transfer for Cross-Domain Recommendation
abstract
Cross-Domain Recommendation (CDR) has been proven to effectively alleviate the data sparsity problem in Recommender System (RS). Recent CDR methods often disentangle user features into domain-invariant and domain-specific features for efficient cross-domain knowledge transfer. Despite showcasing robust performance, three crucial aspects remain unexplored for existing disentangled CDR approaches: i) The significance nuances of the interaction behaviors are ignored in generating disentangled features; ii) The user features are disentangled irrelevant to the individual items to be recommended; iii) The general knowledge transfer overlooks the user's personality when interacting with diverse items. To this end, we propose a Graph Disentangled Contrastive framework for CDR (GDCCDR) with personalized transfer by meta-networks. An adaptive parameter-free filter is proposed to gauge the significance of diverse interactions, thereby facilitating more refined disentangled representations. In sight of the success of Contrastive Learning (CL) in RS, we propose two CL-based constraints for item-aware disentanglement. Proximate CL ensures the coherence of domain-invariant features between domains, while eliminatory CL strives to disentangle features within each domains using mutual information between users and items. Finally, for domain-invariant features, we adopt meta-networks to achieve personalized transfer. Experimental results on four real-world datasets demonstrate the superiority of GDCCDR over state-of-the-art methods.
Jing Liu 0002, Lele Sun, Weizhi Nie, Peiguang Jing, Yuting Su 0001
AAAI5
2024 Beyond Users: Denoising Behavior-based Contrastive Learning for Disentangled Cross-Domain Recommendation
Lele Sun, Jing Liu 0002, Shenyuan Zhang, Weizhi Nie, Anan Liu, Yuting Su 0001
DASFAA (2)6
2024 Multimodal High-Order Relationship Inference Network for Fashion Compatibility Modeling in Internet of Multimedia Things
abstract
Recent progress in artificial intelligence (AI) have broadened various intelligent application scenarios on the Internet of Multimedia Things (IoMT). Due to the urgent demands for intelligence in the online fashion industry, using AI techniques to explore user’s clothing collocations from fashion data generated by the IoMT system is a challenging task. Under the background, fashion compatibility modeling (FCM), which aims to estimate the matching degree of a given outfit, has attracted great attention in the multimedia analysis field. However, most of the studies often fail to fully leverage multimodal content or ignore the sparse associations between fashion items. In this article, we propose a novel multimodal high-order relationship inference network (MHRIN) for FCM task. In MHRIN, we focus on enriching multimodal representations of fashion items by means of incorporating the category correlations and injecting high-order item–item connectivity. Concretely, considering that fashion collocations depend on the semantic relevance patterns between categories, we design a category correlations learning module to adaptively learn category representations. On this basis, multiple modality representations are aggregated by a hierarchical multimodal fusion module to generate visual-semantic embeddings. To address the item-item matching interactions issue, we further refine the final representations by a high-order message propagation module to absorb rich connection information. Experiments on the publicly available data set demonstrate the superiority of our MHRIN over state-of-the-art methods.
Peiguang Jing, Jing Zhang 0038, Yun Li 0006, Yuting Su 0001
IEEE Internet Things J.5
2024 Dual Preference Perception Network for Fashion Recommendation in Social Internet of Things
abstract
Nowadays, with the continuous development of information technology, the application scenarios of the Internet of Things (IoT) are progressively expanding to the social field, engendering widespread attention to the Social IoT (SIoT). Personalized fashion recommendation that possesses the potential to establish social relationships between clothing and humans has substantially broadened the scope of the SIoT, particularly with the flourishing fashion industry and the ascent of smart home. Compared to conventional recommendations, fashion recommendation generally suggests a collection of items rather than individual pieces for users. Additionally, considering the public acceptance alongside the user-specific preference is reasonable for fashion recommendation, however, current methods often overlook the former. To comprehensively capture the public acceptance and the user-specific preference, we propose a dual preference perception network (DP2Net) for fashion recommendation. First, a fashion corpus is constructed to facilitate the condensation of general taste, wherein adversarial learning and determinantal point process are leveraged to ensure representativeness and diversity of the corpus. Second, a user-general preference perception module is built based on a bottleneck transformer structure to generate aggregated representations for the corpus. Third, a user-specific preference perception module is constructed to acquire collaborative representations of users and outfits by employing an attentive heterogeneous graph embedding. The final loss functions of two preference perception modules are constructed by combining the representations of users, outfits, and the corpus. Experiments on large-scale real-world data sets demonstrate the effectiveness of the proposed method. To facilitate reproducible research, we have made our code publicly available athttps://github.com/KaiZhang1228/DP2Net.
Peiguang Jing, Xianyi Liu, Yun Li 0006, Yu Liu 0004, Yuting Su 0001
IEEE Internet Things J.6
2024 Multimodal deep hierarchical semantic-aligned matrix factorization method for micro-video multi-label classification
Fugui Fan, Yuting Su 0001, Yun Liu 0009, Peiguang Jing, Kaihua Qu, Yu Liu 0004
Inf. Process. Manag.2
2024 A deep low-rank semantic factorization method for micro-video multi-label classification
Fugui Fan, Yuting Su 0001, Yu Liu 0004, Peiguang Jing, Kaihua Qu
Multim. Syst.2
2024 Adaptive proposal network based on generative adversarial learning for weakly supervised temporal sentence grounding
Weikang Wang 0002, Yuting Su 0001, Jing Liu 0002, Peiguang Jing
Pattern Recognit. Lett.2
2024 Collaborative spatial-temporal video salient object detection with cross attention transformer
Yuting Su 0001, Weikang Wang 0002, Jing Liu 0002, Peiguang Jing
Signal Process.1
2024 Duration-aware and mode-aware micro-expression spotting for long video sequences
Jing Liu 0001, Guangtao Zhai, Yuting Su 0001
Signal Process. Image Commun.5
2024 Deep Matrix Factorization With Complementary Semantic Aggregation for Micro-Video Multi-Label Classification
abstract
Deep matrix factorization has been demonstrated in extracting hierarchical knowledge describing micro-video characteristics. However, the complementary information across distinct latent layers is often ignored. To address this issue, we propose a deep matrix factorization with complementary semantic aggregation (DMFCSA) method for micro-video multi-label classification, which consists of the multi-layer representation learning module and the semantic decoding module. We first employ deep hierarchical matrix factorization to learn the underlying semantic representations at each latent layer. Meanwhile, the semantic aggregation strategy is exploited to integrate complementary information from different layers into the output layer. To enhance discriminability, we construct a triplet term that effectively establishes relationships among features, labels, and attributes. Moreover, the semantic decoding module is designed to enhance both robustness and representation ability by reconstructing the original inputs. Experimental results on a real-world multi-label dataset show the effectiveness and robustness of our method compared to several state-of-the-art methods.
Peiguang Jing, Yuting Su 0001
IEEE Signal Process. Lett.4
2024 Deep Multi-Modal Hashing With Semantic Enhancement for Multi-Label Micro-Video Retrieval
abstract
The pressing need for low storage and high efficiency has significantly propelled the advancement of deep hashing techniques in the realm of large-scale search and retrieval tasks. As one of the most prevailing forms of user-generated contents, micro-videos usually represent more complicated multi-modal behaviors that are further challenged in multi-label retrieval. Existing multi-modal hashing methods tend to prioritize the complementarity and consistency in multi-modal fusion, while neglecting the completeness problem. In this paper, we propose a deep multi-modal hashing with semantic enhancement (DMHSE) method that effectively integrates complete multi-modal representation learning with discriminative binary coding by means of collaboration between two distinct encoders, FoldCoder and HashCoder. FoldCoder translates latent multi-modal representation learning to a degradation process through mimicking data transmitting. Further, it incorporates a prompt learning paradigm to maximize the utilization of multi-label semantics for guiding representation learning. HashCoder combines pairwise and central constraints to ensure more discriminative hashing results. Pairwise constraint preserves the original local relevance structure, while central constraint tackles the problem of semantic ambiguity in multi-label data by leveraging the global label distribution. Experimental results demonstrate that DMHSE achieves superior performance in multi-label micro-video retrieval tasks.
Peiguang Jing, Haoyi Sun, Liqiang Nie, Yun Li 0006, Yuting Su 0001
IEEE Trans. Knowl. Data Eng.5
2024 Multi-Task Spatial-Temporal Transformer for Multi-Variable Meteorological Forecasting
abstract
This study delves into multi-variable meteorological spatial-temporal prediction, focusing on the simultaneous forecasting of key meteorological parameters such as temperature, wind speed, and atmospheric pressure. The core challenge of this task lies in identifying commonalities across different variables while capturing their unique features and the interactions among them. To address this, we propose a novel multi-task learning framework tailored for multi-variable meteorological forecasting. Our framework integrates a convolutional variable-specific visual representation module and a variable-interactive spatial-temporal inference module. The former extracts distinct variable information independently for each variable, while the latter employs a tri-level attention mechanism across space, time, and variables to uncover both commonalities and interactions among the variables. An adaptive multi-loss optimization strategy and a local information aggregation module are introduced to balance task optimization complexities and enhance representation stability. Comprehensive experiments across various meteorological prediction tasks confirm the effectiveness of our methods, showcasing superior performance over existing approaches.
Tianbao Li 0001, Anan Liu, Dan Song 0006, Wenhui Li 0001, Jing Zhang 0038, Zhiqiang Wei 0002, Yuting Su 0001
IEEE Trans. Knowl. Data Eng.7
2024 SADCMF: Self-Attentive Deep Consistent Matrix Factorization for Micro-Video Multi-Label Classification
abstract
Currently, there is a growing scholarly and industrial interest in micro-video-centric research. Within these domains, multi-label learning has emerged as a fundamental yet attractive subject. Existing methods primarily place emphasis on feature representations of individual micro-videos, while neglecting latent interdependencies between instance and label domains. To address this problem, in this paper, we propose a novel self-attentive deep consistent matrix factorization (SADCMF) method, which jointly explores dualdomain hierarchical representations and their inherent dependencies for micro-video multi-label classification. Specifically, SADCMF includes three primary characteristics: 1) A dualdomain deep collaborative factorization module is developed to explore the first-stage representations of instance features and the discriminative embeddings of label semantics in a mutually beneficial manner. 2) A correlation-driven selfattentive factorization module is devised to acquire the labelaware attentive outputs, which are further combined with original features through a residual structure to enrich the second-stage feature representations. 3) A dual-stream representation consistency module ensures the unidirectional and bidirectional representation consistency, meanwhile, narrows the discrepancies between the two-stage representations for improving the generalization ability of our method. Extensive experiments conducted on two publicly available micro-video multi-label datasets demonstrate its superior performance in comparison with state-of-the-art methods.
Fugui Fan, Peiguang Jing, Liqiang Nie, Haoyu Gu, Yuting Su 0001
IEEE Trans. Multim.5
2024 Dual-Domain Aligned Deep Hierarchical Matrix Factorization Method for Micro-Video Multi-Label Classification
abstract
Recently, with the growing popularity of micro-videos, multi-label learning has attracted increasing attention due to its potential commercial value in different scenarios. However, existing methods place more emphasis on the alignment between explicit semantics and visual features, while neglecting the exploration of interactions at fine-grained semantic levels. To address this problem, we propose a novel dual-domain aligned deep hierarchical matrix factorization (DADHMF) method for micro-video multi-label classification. Specifically, we construct a dual-stream deep matrix factorization framework to explore implicit hierarchical semantics and corresponding intrinsic feature representations in top-down and bottom-up ways, respectively. On this basis, we leverage the intralayer alignment strategy to narrow the semantic gap between label and instance domains by introducing adaptive semantic-aware embeddings. Moreover, we further utilize the inverse covariance estimation module to automatically capture latent semantic correlations, and project the structural information into the semantic-aware embeddings to ensure the stability of the intralayer alignment. Extensive experiments on two available micro-video multi-label datasets demonstrate that our proposed method outperforms the state-of-the-art methods.
Fugui Fan, Yuting Su 0001, Liqiang Nie, Peiguang Jing, Daozheng Hong, Yu Liu 0004
IEEE Trans. Multim.2
2024 Multimodal Progressive Modulation Network for Micro-Video Multi-Label Classification
abstract
Micro-videos, as an increasingly popular form of user-generated content (UGC), naturally include diverse multimodal cues. However, in pursuit of consistent representations, existing methods neglect the simultaneous consideration of exploring modality discrepancy and preserving modality diversity. In this paper, we propose a multimodal progressive modulation network (MPMNet) for micro-video multi-label classification, which enhances the indicative ability of each modality through gradually regulating various modality biases. In MPMNet, we first leverage a unimodal-centered parallel aggregation strategy to obtain preliminary comprehensive representations. We then integrate feature-domain disentangled modulation process and category-domain adaptive modulation process into a unified framework to jointly refine modality-oriented representations. In the former modulation process, we constrain inter-modal dependencies in a latent space to obtain modality-oriented sample representations, and introduce a disentangled paradigm to further maintain modality diversity. In the latter modulation process, we construct global-context-aware graph convolutional networks to acquire modality-oriented category representations, and develop two instance-level parameter generators to further regulate unimodal semantic biases. Extensive experiments on two micro-video multi-label datasets show that our proposed approach outperforms the state-of-the-art methods.
Peiguang Jing, Fugui Fan, Yun Li 0006, Yuting Su 0001
IEEE Trans. Multim.6
2024 Progressive Fourier Adversarial Domain Adaptation for Object Classification and Retrieval
abstract
Domain adaptation has been extensively explored as a means of transferring knowledge from the labeled source domain to the unlabeled target domain with disparate data distributions. However, the absence of target annotations and significant domain discrepancies pose a great challenge to transfer knowledge directly from source domain to target domain. To address this challenge, we propose a Progressive Fourier Adversarial Domain Adaptation (PFADA) framework, an effective and versatile framework which can generalize across multiple domain adaptation tasks. Firstly, we propose a Fourier-based style transfer strategy to generate a Fourier intermediate domain that incorporates source images with target domain-specific styles, while preserving the domain-invariant representations of the source data. Secondly, we introduce a progressive adversarial domain adaptation approach that utilizes the Fourier intermediate domain to facilitate the learning of domain-invariant representations. Finally, we present cross-domain semantic alignment and discriminative enhancement approach, which effectively guides the learning of discriminative cross-domain representations utilizing labeled source and intermediate domain data. Extensive experimental evaluations consistently validate the superior performance of the proposed method across diverse visual tasks, encompassing multiple domain adaptive image classification and retrieval scenarios.
Tianbao Li 0001, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, Zhiqiang Wei 0002, Anan Liu
IEEE Trans. Multim.2
2024 Multi-Stage Spatio-Temporal Fusion Network for Fast and Accurate Video Bit-Depth Enhancement
abstract
For video bit-depth enhancement (VBDE) tasks, inter-frame information is critical for removing false contours and recovering the details in low bit-depth (LBD) videos. However, due to different structural distortions and complex motions in the neighboring frames, it is difficult to effectively utilized inter-frame information. Most algorithms rely on alignment operations to provide information of neighboring frames, suffering from slow inference speed due to the complex alignment module design. Meanwhile, most existing methods sequentially perform the intra-frame feature extractions and inter-frame information fusions, but fail to efficiently fuse spatio-temporal information. Therefore, in this paper, we propose a two-stage progressive group (TSPG) network to find complementary information related to the target frame without adopting an alignment operation. To simultaneously achieve intra-frame feature extractions and inter-frame feature fusions, we propose a parallel spatio-temporal fusion (PSTF) module with a dual-branch spatial-temporal residual (DSTR) block to focus on more useful temporal information while ensuring a faster inference speeds. Extensive experiments on public datasets demonstrate that our proposed multi-stage spatio-temporal fusion network (named MSTFN) can quickly and effectively eliminate false contours and recover high quality target frames. Furthermore, our method outperforms the state-of-the-art methods in terms of both PSNR and SSIM, and can reach faster inference speeds.
Jing Liu 0002, Ziwen Yang, Yuting Su 0001, Xiaokang Yang 0001
IEEE Trans. Multim.4
2024 Pixel-Learnable 3DLUT With Saturation-Aware Compensation for Image Enhancement
abstract
The 3D Lookup Table (3DLUT)-based methods are gaining popularity due to their satisfactory and stable performance in achieving automatic and adaptive real time image enhancement. In this paper, we present a new solution to the intractability in handling continuous color transformations of 3DLUT due to the lookup via three independent color channel coordinates in RGB space. Inspired by the inherent merits of the HSV color space, we separately enhance image intensity and color composition. The Transformer-based Pixel-Learnable 3D Lookup Table is proposed to undermine contouring artifacts, which enhances images in a pixel-wise manner with non-local information to emphasize the diverse spatially variant context. In addition, noticing the underestimation of composition color component, we develop the Saturation-Aware Compensation (SAC) module to enhance the under-saturated region determined by an adaptive SA map with Saturation-Interaction block, achieving well balance between preserving details and color rendition. Our approach can be applied to image retouching and tone mapping tasks with fairly good generality, especially in restoring localized regions with weak visibility. The performance in both theoretical analysis and comparative experiments manifests that the proposed solution is effective and robust.
Jing Liu 0002, Xiongkuo Min, Yuting Su 0001, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Multim.4
2024 Inter- and Intra-Domain Potential User Preferences for Cross-Domain Recommendation
abstract
Data sparsity poses a persistent challenge in Recommender Systems (RS), driving the emergence of Cross-Domain Recommendation (CDR) as a potential remedy. However, most existing CDR methods often struggle to circumvent the transfer of domain-specific information, which are perceived as noise in the target domain. Additionally, they primarily concentrate on inter-domain information transfer, disregarding the comprehensive exploration of data within intra-domains. To address these limitations, we propose SUCCDR (SeparatingUser features withCompound samples), a novel approach that tackles data sparsity by leveraging both cross-domain knowledge transfer and comprehensive intra-domain analysis. Specifically, to ensure the exclusion of noisy domain-specific features during the transfer process, user preferences are separated into domain-invariant and domain-specific features through three efficient constraints. Furthermore, the unobserved items are leveraged to generate compound samples that intelligently merge observed and unobserved potential user-item interaction, utilizing a simple yet efficient attention mechanism to enable a comprehensive and unbiased representation of user preferences. We evaluate the performance of SUCCDR on two real-world datasets, Douban and Amazon, and compare it with state-of-the-art single-domain and cross-domain recommendation methods. The experimental results demonstrate that SUCCDR outperforms existing approaches, highlighting its ability to effectively alleviate data sparsity problem.
Jing Liu 0002, Lele Sun, Weizhi Nie, Yuting Su 0001, Yongdong Zhang 0001, Anan Liu
IEEE Trans. Multim.4
2024 VMemNet: A Deep Collaborative Spatial-Temporal Network With Attention Representation for Video Memorability Prediction
abstract
Video memorability measures the degree to which a video is remembered by different viewers and has shown great potential in various contexts, including advertising, education, and health care. While extensive research has been conducted on image memorability, the study of video memorability is still in its early stages. Existing methods in this field primarily focus on coarse-grained spatial feature representation and decision fusion strategies, overlooking the crucial interactions between spatial and temporal domains. Therefore, we propose an end-to-end collaborative spatial-temporal network called VMemNet, which incorporates targeted attention mechanisms and intermediation fusion strategies. This enables VMemNet to capture the intricate relationships between spatial and temporal information and uncover more elements of memorability within video visual features. VMemNet integrates spatially and semantically guided attention modules into a dual-stream network architecture, allowing it to simultaneously capture static local cues and dynamic global cues in videos. Specifically, the spatial attention module is used to aggregate more memorable elements from spatial locations, and the semantically guided attention module is used to achieve semantic alignment and intermediate fusion of the local and global cues. In addition, two types of loss functions with complementary decision rules are associated with the corresponding attention modules to guide the training process of the proposed network. Experimental results obtained on a publicly available dataset verify that the proposed VMemNet approach outperforms all current single- and multi-modal methods in terms of video memorability prediction.
Wei Lu 0026, Jiaze Han, Peiguang Jing, Yu Liu 0004, Yuting Su 0001
IEEE Trans. Multim.6
2024 MCDAN: A Multi-Scale Context-Enhanced Dynamic Attention Network for Diffusion Prediction
abstract
Information diffusion prediction aims at predicting the target users in the information diffusion path on social networks. Prior works mainly focus on the observed structure or sequence of cascades, trying to predict to whom this cascade will be infected passively. In this study, we argue that user intent understanding is also a key part of information diffusion prediction. We thereby propose a novel Multi-scale Context-enhanced Dynamic Attention Network (MCDAN) to predict which user will most likely join the observed current cascades. Specifically, to consider the global interactive relationship among users, we take full advantage of user friendships and global cascading relationships, which are extracted from the social network and historical cascades, respectively. To refine the model's ability to understand the user's preference for the current cascade, we propose a multi-scale sequential hypergraph attention module to capture the dynamic preference of users at different time scales. Moreover, we design a contextual attention enhancement module to strengthen the interaction of user representations within the current cascade. Finally, to engage the user's own susceptibility, we construct a susceptibility label for each user based on user susceptibility analysis and use the rank of this label for auxiliary prediction. We conduct experiments over four widely used datasets and show that MCDAN significantly overperforms the state-of-the-art models. The average improvements are up to 5.41% in terms of Hits@100 and 8.47% in terms of MAP@100, respectively.
Lanjun Wang, Yuting Su 0001, Yongdong Zhang 0001, Anan Liu
IEEE Trans. Multim.3
2024 CDCM: ChatGPT-Aided Diversity-Aware Causal Model for Interactive Recommendation
abstract
In recent years, interactive recommender systems (IRSs) have attracted extensive interest. Existing IRSs are typically implemented with offline reinforcement learning (RL). They are devoted to improving recommendation accuracy by optimizing the extraction of users' inherent preferences. However, there hasn't been much attention on recommendation diversity, which could result in the monotony effect,i.e., categories of recommended items are consistently fixed and unchanging. In this paper, we center on category diversification in IRSs while largely preserving or even boosting recommendation accuracy. To this end, we propose a ChatGPT-aided diversity-aware causal model (CDCM) to enhance the offline RL framework with causal inference and ChatGPT. Specifically, we first propose a diversity-aware causal user model (DCUM) to estimate user satisfaction. This model disentangles the causal effect of users' inherent preferences and the monotony effect to obtain user satisfaction with both accuracy and diversity. Then, DCUM is used to assist the RL agent in recommendation policy learning. A ChatGPT-aided state encoder (CSE) is proposed to provide user state representation for each time step of policy learning. With the help of ChatGPT, CSE incorporates multi-category information in line with users' potential preferences to promote diverse and relevant category recommendations. Extensive experiment results on two real-world datasets validate the superiority of our CDCM regarding both accuracy and diversity.
Xin Wen 0017, Weizhi Nie, Jing Liu 0002, Yuting Su 0001, Yongdong Zhang 0001, Anan Liu
IEEE Trans. Multim.4
2024 Multimodal Attentive Representation Learning for Micro-video Multi-label Classification
abstract
As one of the representative types of user-generated contents (UGCs) in social platforms, micro-videos have been becoming popular in our daily life. Although micro-videos naturally exhibit multimodal features that are rich enough to support representation learning, the complex correlations across modalities render valuable information difficult to integrate. In this paper, we introduced a multimodal attentive representation network (MARNET) to learn complete and robust representations to benefit micro-video multi-label classification. To address the commonly missing modality issue, we presented a multimodal information aggregation mechanism module to integrate multimodal information, where latent common representations are obtained by modeling the complementarity and consistency in terms of visual-centered modality groupings instead of single modalities. For the label correlation issue, we designed an attentive graph neural network module to adaptively learn the correlation matrix and representations of labels for better compatibility with training data. In addition, a cross-modal multi-head attention module is developed to make the learned common representations label-aware for multi-label classification. Experiments conducted on two micro-video datasets demonstrate the superior performance of MARNET compared with state-of-the-art methods.
Peiguang Jing, Xianyi Liu, Yun Li 0006, Yu Liu 0004, Yuting Su 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Privacy-preserving Multi-source Cross-domain Recommendation Based on Knowledge Graph
abstract
The cross-domain recommender systems aim to alleviate the data sparsity problem in the target domain by transferring knowledge from the auxiliary domain. However, existing works ignore the fact that the data sparsity problem may also exist in the single auxiliary domain, and sharing user behavior data is restricted by the privacy policy. In addition, their cross-domain models lack interpretability. To address these concerns, we propose a novel multi-source cross-domain model based on knowledge graph. Specifically, to avoid the insufficiency of single auxiliary domain, we construct a knowledge graph comprehensively leveraging items from multiple auxiliary domains. To avoid the leakage of user privacy when user information is transferred to multiple domains, we construct graph for information transfer between items to effectively avoid the propagation of users’ private information between different domains. We implicitly integrate the user–item interaction by transferring the learned item embeddings. To improve the interpretability of cross-domain knowledge transfer, we propose a knowledge graph-based retrieval and fusion method to transfer knowledge derived from multiple auxiliary domains. An attention-based fusion network is designed to enhance the representation of the targeted user and items with the transferred item embedding. We perform extensive experiments on three real-world datasets, demonstrating that our model outperforms the states of the art.
Jing Liu 0002, Litao Shang, Yuting Su 0001, Weizhi Nie, Xin Wen 0017, Anan Liu
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Semantic Embedding Uncertainty Learning for Image and Text Matching
abstract
Image and text matching measures the semantic similarity for cross-modal retrieval. The core of this task is semantic embedding, which mines the intrinsic characteristics of visual and textual for discriminative representation. However, cross-modal ambiguity of image and text (the existence of one-to-many associations) is prone to semantic diversity. The mainstream approaches utilized the fixed point embedding to represent semantics, which ignored the embedding uncertainty caused by semantic diversity leading to incorrect results. To address this issue, we propose a novel Semantic Embedding Uncertainty Learning (SEUL), which represents the embedding uncertainty of image and text as Gaussian distributions and simultaneously learns the salient embedding (mean) and uncertainty (variance) in the common space. We design semantic uncertainty embedding for facilitating the robustness of the representation in the semantic diversity context. A combined objective function is proposed, which optimizes the semantic uncertainty and maintains discriminability to enhance cross-modal associations. Extended experiments are performed on two datasets to demonstrate advanced performance.
Yan Wang 0114, Yuting Su 0001, Wenhui Li 0001, Chenggang Yan 0001, Bolun Zheng, Xuanya Li, Anan Liu
ICME2
2023 StyleEDL: Style-Guided High-order Attention Network for Image Emotion Distribution Learning
abstract
Emotion distribution learning has gained increasing attention with the tendency to express emotions through images. As for emotion ambiguity arising from humans' subjectivity, substantial previous methods generally focused on learning appropriate representations from the holistic or significant part of images. However, they rarely consider establishing connections with the stylistic information although it can lead to a better understanding of images. In this paper, we propose a style-guided high-order attention network for image emotion distribution learning termed StyleEDL, which interactively learns stylistic-aware representations of images by exploring the hierarchical stylistic information of visual contents. Specifically, we consider exploring the intra- and inter-layer correlations among GRAM-based stylistic representations, and meanwhile exploit an adversary-constrained high-order attention mechanism to capture potential interactions between subtle visual parts. In addition, we introduce a stylistic graph convolutional network to dynamically generate the content-dependent emotion representations to benefit the final emotion distribution learning. Extensive experiments conducted on several benchmark datasets demonstrate the effectiveness of our proposed StyleEDL compared to state-of-the-art methods. The implementation is released at: https://github.com/liuxianyi/StyleEDL.
Peiguang Jing, Xianyi Liu, Yinwei Wei, Liqiang Nie, Yuting Su 0001
ACM Multimedia6
2023 Progressive Positive Association Framework for Image and Text Retrieval
abstract
With the increasing amount of multimedia data, the demand for fast and accurate access to information is growing. Image and text retrieval learns visual and textual semantic relationships for multimedia data management and content recognition. The main challenge of this task is how to derive image and text similarity based on local associations under huge modal gap. However, the existing methods compute semantic relevance using associations of all fragments (visual regions and textual words), which underestimate the uncertainty of associations and discriminative positive associations leading to cross-modal correspondence ambiguity. To address these issues, we propose a novel Progressive Positive Association Framework (PPAF), which models association uncertainty as a normal distribution and progressively mines direct and potential positive associations according to the characteristics of the association distribution. We design positive association matching, which adaptively fuses multi-step associations for local matching depending on the relevance difference. In addition, we apply KL loss constraint on cross-modal association distribution in order to enhance local semantic alignment. Extended experiments demonstrate the leading performance of PPAF.
Wenhui Li 0001, Yan Wang 0114, Yuting Su 0001, Lanjun Wang, Weizhi Nie, Anan Liu
ACM Multimedia3
2023 Efficient Spatio-Temporal Video Grounding with Semantic-Guided Feature Decomposition
abstract
Spatio-temporal video grounding (STVG) aims to localize the spatio-temporal object tube in a video according to a given text query. Current approaches address the STVG task with end-to-end frameworks while suffering from heavy computational complexity and insufficient spatio-temporal interactions. To overcome these limitations, we propose a novel Semantic-Guided Feature Decomposition based Network (SGFDN). A semantic-guided mapping operation is proposed to decompose the 3D spatio-temporal feature into 2D motions and 1D object embedding without losing much object-related semantic information. Thus, the computational complexity in computationally expensive operations such as attention mechanisms can be effectively reduced by replacing the input spatio-temporal feature with the decomposed features. Furthermore, based on this decomposition strategy, a pyramid relevance filtering based attention is proposed to capture the cross-modal interactions at multiple spatio-temporal scales. In addition, a decomposition-based grounding head is proposed to locate the queried objects with less computational complexity. Extensive experiments on two widely-used STVG datasets (VidSTG and HC-STVG) demonstrate that our method enjoys state-of-the-art performance as well as less computational complexity. The code has been available at https://github.com/TJUMMG/SGFDN.
Weikang Wang 0002, Jing Liu 0002, Yuting Su 0001, Weizhi Nie
ACM Multimedia3
2023 Principal views selection based on growing graph convolution network for multi-view 3D model recognition
Qi Liang 0004, Qiang Li 0048, Weizhi Nie, Yuting Su 0001
Appl. Intell.4
2023 Rare-aware attention network for image-text matching
Yan Wang 0114, Yuting Su 0001, Wenhui Li 0001, Zhengya Sun, Zhiqiang Wei 0002, Jie Nie, Xuanya Li, Anan Liu
Inf. Process. Manag.2
2023 SMPC: boosting social media popularity prediction with caption
Anan Liu, Ning Xu 0003, Jing Liu 0002, Yuting Su 0001, Shenyuan Zhang, Yejun Tang, Junbo Guo, Guoqing Jin, Xuanya Li
Multim. Syst.5
2023 A Multimodal Aggregation Network With Serial Self-Attention Mechanism for Micro-Video Multi-Label Classification
abstract
Currently, micro-videos have attracted increasing attention due to their unique properties and great commercial value. Considering that micro-videos naturally incorporate multimodal information, a powerful representation method for distinct joint multimodal representations is essential for real applications. Inspired by the potential of attention neural network architectures over various tasks, we propose a multimodal aggregation network (MANET) with a serial self-attention mechanism to perform tasks of micro-video multi-label classification. Specifically, we first propose a parallel content-dependent graph neural networks (CDGNN) module, which explores category-related embeddings of micro-videos by disentangling category relations into modality-specific and modality-shared category dependency patterns. Then we introduce a serial self-attention (SSA) module to transmit the multimodal information in sequential order, in which an aggregation bottleneck is incorporated to better collect and condense the significant information. Experiments conducted on a large-scale multi-label micro-video dataset demonstrate that our proposed method has achieved competitive results compared with several state-of-the-art methods.
Wei Lu 0026, Peiguang Jing, Yuting Su 0001
IEEE Signal Process. Lett.4
2023 Focus on Hard Samples: Hierarchical Unbiased Constraints for Cross-Domain 3D Model Retrieval
abstract
Cross-domain 3D model retrieval facilitates the management of explosively emerging unlabeled 3D models with conveniently available 2D images or RGB-D objects, which has attracted more and more attention. The modality gap between query samples (2D images or RGB-D objects) and 3D models makes the task challenging, and adversarial domain adaptation techniques have achieved success in narrowing such gaps. However, existing methods always pay excessive attention to the samples with high discriminability and transferability, whereas the hard samples with rich information are neglected. Accordingly, we propose hierarchical unbiased constraints to make full use of data at semantic level, sample level and feature level to improve the retrieval performance. At semantic level, we utilize maximum F-norm loss to constrain the semantic prediction results of target domain, which takes advantage of more hard samples to reduce ambiguous predictions and enhance discriminability. At sample level, we propose an adaptive triplet center loss to assign less confident samples with a farther negative class, which reliably compacts samples within the same class and expands the distance across different classes. At feature level, we perform SVD (singular value decomposition) for both source features and target features and suppress the relative value of the largest singular value, so that the information of other eigenvectors can be fully utilized to improve transferability. Experiments on two public datasets validate the superiority of the proposed method, and the ablation study analyzes different roles played by these hierarchical unbiased constraints.
Tianbao Li 0001, Anan Liu, Dan Song 0006, Wenhui Li 0001, Xuanya Li, Yuting Su 0001
IEEE Trans. Circuits Syst. Video Technol.6
2023 Trajectory Guided Robust Visual Object Tracking With Selective Remedy
abstract
Siamese trackers have received a lot of attentions due to their promising performance and real-time high speed. However, the robustness is generally limited especially in challenging conditions, such as occlusion. Motivated by that tracking failure does not always occur and the target movement tends to follow certain patterns, in the paper, we propose a generic, fast and flexible approach to improve the robustness of Siamese trackers with two light-load novel modules: Trajectory Guidance Module (TGM) and Selective Refinement Module (SRM). Specifically, TGM encourages to pay a soft attention on possible target location based on short-term historical trajectory. SRM selectively remedies the tracking results at the risk of failure with little impact on the speed. The proposed algorithm can be easily establish upon state-of-the-art Siamese trackers and obtains better performance on seven benchmarks with high real-time tracking speed. The code is available athttps://github.com/TJUMMG/TGSR.
Han Wang 0034, Jing Liu 0002, Yuting Su 0001, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Dual-Path Rare Content Enhancement Network for Image and Text Matching
abstract
Image and text matching plays a crucial role in bridging the cross-modal gap between vision and language, and has achieved great progress due to the deep learning. However, the existing methods still suffer from the long-tail problem, where only a small proportion contains highly frequent semantics and a long tail proportion is constructed by rare semantics. In this paper, we propose a novel Dual-path Rare Content Enhancement Network (DRCE) to tackle the long-tail issue. Specifically, the Cross-modal Representation Enhancement (CRE) and Cross-modal Association Enhancement (CAE) are proposed to construct dual-path structure to enhance rare content representation and association with the benefit of cross-modal prior knowledge. This structure can effectively exploit the complementary cross-modal relation from different aspects and fuse these information in an adaptively manner by the proposed Adaptive Fusion Strategy (AFS). Moreover, we also propose an alternative re-ranking strategy (ARR) to explore the reciprocal contextual information to refine image-text matching results, which can further suppress the negative effect of long-tail effect. Extensive experiments on two large-scale datasets show the significant improvements and validate the superiority of our method.
Yan Wang 0114, Yuting Su 0001, Wenhui Li 0001, Jun Xiao 0001, Xuanya Li, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.2
2023 MRFT: Multiscale Recurrent Fusion Transformer Based Prior Knowledge for Bit-Depth Enhancement
abstract
Bit-depth enhancement (BDE) plays an important role in providing high bit-depth data support for high-dynamic range (HDR) display. Although convolutional neural network (CNN) based BDE methods have achieved top performance, multiscale feature extraction and fusion still suffer from some inherent architectural flaws. Moreover, the training-data-scarce scene has not been effectively explored. To this end, this paper proposes an innovative multiscale recurrent fusion transformer (MRFT) framework, which contains three key components, i.e. multiscale transformer feature encoder, recurrent feature fusion module, and prior knowledge injection. Specifically, the multiscale transformer feature encoder consists of a prior-injected context encoder (PICE) and a multiscale local feature encoder (MLFE). PICE leverages the vanilla self-attention mechanism to extract the global context correlating spatially-distant contents for distinguishing long-distance false contours. MLFE exploits the local self-attention mechanism with varied window sizes to capture different-scale detail features. Then, a hierarchical recurrent decoder (HRD) is proposed as the recurrent feature fusion module to fuse multiscale visual information with global guidance. Via the circular query-key mechanism, global-to-local information is progressively fused. Furthermore, we propose a two-stage alternating optimization strategy for prior knowledge injection. By pre-parameterizing the global auxiliary priors, the training dilemma on the data-scarce domain is significantly alleviated. Extensive analyses on multiple benchmark datasets demonstrate the superiority of our MRFT in terms of quantitative measures and aesthetic effects.
Xin Wen 0017, Weizhi Nie, Jing Liu 0002, Yuting Su 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Sequence as a Whole: A Unified Framework for Video Action Localization With Long-Range Text Query
abstract
Comprehensive understanding of video content requires both spatial and temporal localization. However, there lacks a unified video action localization framework, which hinders the coordinated development of this field. Existing 3D CNN methods take fixed and limited input length at the cost of ignoring temporally long-range cross-modal interaction. On the other hand, despite having large temporal context, existing sequential methods often avoid dense cross-modal interactions for complexity reasons. To address this issue, in this paper, we propose a unified framework which handles the whole video in sequential manner with long-range and dense visual-linguistic interaction in an end-to-end manner. Specifically, a lightweight relevance filtering based transformer (Ref-Transformer) is designed, which is composed of relevance filtering based attention and temporally expanded MLP. The text-relevant spatial regions and temporal clips in video can be efficiently highlighted through the relevance filtering and then propagated among the whole video sequence with the temporally expanded MLP. Extensive experiments on three sub-tasks of referring video action localization, i.e., referring video segmentation, temporal sentence grounding, and spatiotemporal video grounding, show that the proposed framework achieves the state-of-the-art performance in all referring video action localization tasks. The code has been available at https://github.com/TJUMMG/SAW.
Yuting Su 0001, Weikang Wang 0002, Jing Liu 0002, Xiaokang Yang 0001
IEEE Trans. Image Process.1
2023 Category-Aware Multimodal Attention Network for Fashion Compatibility Modeling
abstract
Fashion compatibility modeling, which is used to estimate the matching degree of a given set of fashion items, has received increasing attention in recent years. However, existing studies often fail to fully leverage multimodal information or ignore the semantic guidance of clothing categories in elevating the reliability of multimodal information. In this paper, we propose a fashion compatibility modeling approach with a category-aware multimodal attention network, termed as FCM-CMAN. In FCM-CMAN, we focus on enriching and aggregating multimodal representations of fashion items by means of the dynamic representations of categories and a contextual attention mechanism simultaneously. Specifically, considering that category correlations are always dynamic and varied for different fashion items, we design a categorical dynamic graph convolutional network to adaptively learn the semantic correlations between categories. When combined with the multi-layered visual outputs of a convolutional neural network and the surrounding contextual information, multiple content-aware category representations and context-aware attention weights are obtained to better characterize fashion items from different aspects. On this basis, two pieces of aware information are integrated by a multimodal factorized bilinear pooling strategy to generate visual-semantic embeddings, which are further improved by a multi-head self-attention mechanism to capture significant elements related to fashion compatibility. Extensive experiments conducted on the FashionVC and ExpFashion datasets demonstrate the superiority of FCM-CMAN over state-of-the-art methods.
Peiguang Jing, Weili Guan, Liqiang Nie, Yuting Su 0001
IEEE Trans. Multim.5
2023 Multi-Scale Fine-Grained Alignments for Image and Sentence Matching
abstract
Image and sentence matching is a critical task to bridge the visual and textual discrepancy due to the heterogeneous modalities. Great progress has been made by exploring the coarse-grained relationships between images and sentences or fine-grained relationships between regions and words. However, how to fully excavate and exploit corresponding relations between these two modalities is still challenging. In this work, we propose a novel Multi-scale Fine-grained Alignments Network (MFA), which can effectively explore multi-scale visual-textual correspondences to facilitate bridging cross-modal discrepancy. Specifically, word-scale matching module is firstly utilized to mine the basic but fundamental correspondences between a single word and independent region. Then, we propose a phrase-scale matching module to explore the relations between objects with the constraint of attribute and corresponding region, which can further reserve more associated information. To cope with the complex interactions among multiple phrases and images, we design the relation-scale matching module to capture high-order semantics between two modalities. Moreover, each matching module includes visual aggregation and textual aggregations, which can ensure the bi-directional coupling of multi-scale semantics. Extensive qualitative and quantitative experiments on two challenging datasets including Flickr30 K and MSCOCO, show that the proposed method achieves superior performance compared with the existing methods.
Wenhui Li 0001, Yan Wang 0114, Yuting Su 0001, Xuanya Li, Anan Liu, Yongdong Zhang 0001
IEEE Trans. Multim.3
2023 Learning Dual Low-Rank Representation for Multi-Label Micro-Video Classification
abstract
Currently, with the rapid development of mobile Internet, micro-video has become a prevailing format of user-generated contents (UGCs) on various social media platforms. Several studies have been conducted towards to understanding high-level micro-video semantics, such as venue categorization, memorability, and popularity. However, these approaches supported tasks with only a single output, which exhibited limitations when attempting to use them to resolve tasks with multiple outputs, especially the multi-label micro-video classification. To tackle this problem, in this paper, we propose a dual multi-modal low-rank decomposition (DMLRD) method for multi-label micro-video classification tasks. To learn more comprehensive micro-video representations, we first learn the low-rank-regularized modality-specific and modality-shared components by considering the consistency and the complementarity among modalities simultaneously. Meanwhile, the less descriptive power of each modality aroused by inherent properties can be solved to a certain extent. To obtain unseen label representations, we next construct a sparsity-regularized multi-matrix normal estimation term to jointly encode the latent relationship structures among labels and dimensions. Experiments on two datasets demonstrate the effectiveness of our proposed method over the state-of-art methods.
Wei Lu 0026, Liqiang Nie, Peiguang Jing, Yuting Su 0001
IEEE Trans. Multim.5
2023 Exploiting Low-Rank Latent Gaussian Graphical Model Estimation for Visual Sentiment Distributions
abstract
Currently, an increasing number of applications and services has encouraged users to openly express their emotions via images. Unlike visual sentiment classification, visual sentiment distribution learning exploits the overall distribution to represent the relative importance of sentiment labels. Considering that most relevant studies have failed to completely model correlation structures or explicitly apply them to unknown instances, in this paper, we proposed a low-rank latent Gaussian graphical model estimation (LGGME) method for visual sentiment distribution learning tasks. There are three main characteristics of LGGME: 1) an integrated inverse covariance matrix whose parameters characterize the latent correlation structures between and within features and sentiments is estimated based on the sparse Gaussian graphical model; 2) a multivariate normal assumption is assigned on the concatenated latent feature representations and the estimated sentiment distributions instead of the original observations for a reasonable surrogate; and 3) the latent feature representations are projected from a low-rank subspace, which is also available for unseen instances, and the estimated sentiment distributions are evaluated by KL divergence to ensure a suitable setting for distribution learning. We further developed an effective optimization algorithm based on the alternating direction method of multipliers (ADMM) for our objective function. The experimental results obtained on three publicly available datasets demonstrate the superiority of our proposed method.
Yuting Su 0001, Peiguang Jing, Liqiang Nie
IEEE Trans. Multim.1
2022 FCMNet: Frequency-aware cross-modality attention networks for RGB-D salient object detection
Chunle Guo, Jing Xu 0008, Yuting Su 0001
Neurocomputing6
2022 Video frame deletion detection based on time-frequency analysis
Yuting Su 0001, Peiguang Jing
J. Vis. Commun. Image Represent.2
2022 Tracking by dynamic template: Dual update mechanism
Jing Liu 0002, Xiangdong Huang 0002, Yuting Su 0001
J. Vis. Commun. Image Represent.4
2022 3DFP-FCGAN: Face completion generative adversarial network with 3D facial prior
Jing Liu 0002, Weikang Wang 0002, Jiexiao Yu, Chunping Zhang, Yuting Su 0001
J. Vis. Commun. Image Represent.5
2022 JFLN: Joint Feature Learning Network for 2D sketch based 3D shape retrieval
Yue Zhao 0042, Qi Liang 0004, Ruixin Ma, Weizhi Nie, Yuting Su 0001
J. Vis. Commun. Image Represent.5
2022 Semantically guided projection for zero-shot 3D model classification and retrieval
Yuting Su 0001, Jiayu Li 0004, Wenhui Li 0001, Zan Gao 0002, Haipeng Chen 0002, Xuanya Li, Anan Liu
Multim. Syst.1
2022 Video splicing detection and localization based on multi-level deep feature fusion and reinforcement learning
Jing Xu 0008, Yuting Su 0001
Multim. Tools Appl.5
2022 Iterative Residual Feature Refinement Network for Bit-Depth Enhancement
abstract
Bit-depth enhancement (BDE) restores high bit-depth (HBD) images from low bit-depth (LBD) ones, which has important applications. Recently, residual-optimized BDE algorithms based on convolutional neural networks (CNNs) have achieved top performance. However, they fail to use a single model to accurately recover all frequency information encoded by missing significant bits at one time on challenging large bit-depth recovery tasks. In this paper, we redefine BDE residual recovery from the perspective of image frequency characteristics. On this basis, we propose an iterative residual feature optimization strategy, which provides an implicit error correction mechanism and improves training and inference efficiency. Furthermore, we design a simple but effective iterative residual feature refinement network (IRFRN). By linking model complexity with the recovery of different frequency information, IRFRN enables a single model to simultaneously recover the missing low and high frequency information. Extensive experiments indicate that our method achieves the state-of-the-art quantitative and qualitative performance on large bit-depth recovery tasks.
Weizhi Nie, Xin Wen 0017, Jing Liu 0002, Yuting Su 0001
IEEE Signal Process. Lett.4
2022 Deep Reinforcement Learning-Based Progressive Sequence Saliency Discovery Network for Mitosis Detection In Time-Lapse Phase-Contrast Microscopy Images
abstract
Mitosis detection plays an important role in the analysis of cell status and behavior and is therefore widely utilized in many biological research and medical applications. In this article, we propose a deep reinforcement learning-based progressive sequence saliency discovery network (PSSD)for mitosis detection in time-lapse phase contrast microscopy images. By discovering the salient frames when cell state changes in the sequence, PSSD can more effectively model the mitosis process for mitosis detection. We formulate the discovery of salient frames as a Markov Decision Process (MDP)that progressively adjusts the selection positions of salient frames in the sequence, and further leverage deep reinforcement learning to learn the policy in the salient frame discovery process. The proposed method consists of two parts: 1)the saliency discovery module that selects the salient frames from the input cell image sequence by progressively adjusting the selection positions of salient frames; 2)the mitosis identification module that takes a sequence of salient frames and performs temporal information fusion for mitotic sequence classification. Since the policy network of the saliency discovery module is trained under the guidance of the mitosis identification module, PSSD can comprehensively explore the salient frames that are beneficial for mitosis detection. To our knowledge, this is the first work to implement deep reinforcement learning to the mitosis detection problem. In the experiment, we evaluate the proposed method on the largest mitosis detection dataset, C2C12-16. Experiment results show that compared with the state-of-the-arts, the proposed method can achieve significant improvement for both mitosis identification and temporal localization on C2C12-16.
Yuting Su 0001, Yao Lu 0005, Anan Liu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 Residual-Guided Multiscale Fusion Network for Bit-Depth Enhancement
abstract
Bit-depth enhancement (BDE) is a challenging task due to stubborn false contour artifacts and disappeared detailed information. Given the mixture of structural distortions and real edges in low bit-depth (LBD) images, both large and small receptive fields (RFs) are critical for BDE tasks. However, even powerful state-of-the-art CNN-based methods can hardly capture sufficient LBD features under multiple RFs. This paper proposes a residual-guided multiscale fusion network (RMFNet) to explore multiscale features in a residual manner. We find that the shuffling operation provides desired multiscale inputs for effectively distinguishing false contours from real edges without any loss of information. Therefore, we shuffle LBD images to multiple scales and then fully extract residual features under different RFs with corresponding subnets. To facilitate interscale guidance from the global context to the local context, we progressively transfer the encoded residual features between adjacent subnets from top to bottom. We further propose a dual-branch depthwise group fusion (DDGF) module to fully capture inter- and inner correlations of multiscale features with fewer parameters. Finally, extensive experiments show that our algorithm achieves excellent performance improvement both quantitatively and qualitatively, verifying its effectiveness.
Jing Liu 0002, Xin Wen 0017, Weizhi Nie, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Tripartite Graph Regularized Latent Low-Rank Representation for Fashion Compatibility Prediction
abstract
In recent years, an increasing online shopping demand has greatly promoted the innovation and development of the fashion industry. Visual fashion analysis has become a prospective research topic in computer vision and multimedia fields. Among these studies, fashion compatibility analysis is required in many real applications, such as fashion recommendation, matching, and retrieval. However, learning fashion compatibility is nontrivial, not only due to the uncertain and sparse dependencies among fashion items but also the latent and mutual associations among multiple factors such as color, texture, style, and functionality. To better predict fashion compatibility, in this paper, we proposed a tripartite graph regularized latent low-rank representation method, named TGRLLR, for fashion compatibility prediction. In TGRLLR, to learn more low-dimensional and effective representations, we considered the latent low-rank representation by decomposing the original feature matrix in both the column and row directions to tackle the problem of insufficient observations. On this basis, we simultaneously exploited different regularization strategies to encode the structured correlations among features, the high-order relationships among items, and the geometrical structures of outfits for more informative representations. Extensive experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed method compared with state-of-the-art methods.
Peiguang Jing, Jing Zhang 0038, Liqiang Nie, Jing Liu 0002, Yuting Su 0001
IEEE Trans. Multim.6
2022 TANet: Target Attention Network for Video Bit-Depth Enhancement
abstract
Video bit-depth enhancement (VBDE) reconstructs high-bit-depth (HBD) frames from a low-bit-depth (LBD) video sequence. As neighboring frames contain a considerable amount of complementary information related to the center frame, it is vital for VBDE to exploit neighboring frames as much as possible. Conventional VBDE algorithms with explicit alignment across frames attempt to warp each neighboring frame to the center frame with estimated optical flow, taking into account only pairwise correlation. Most spatiotemporal fusion approaches involve direct concatenation or 3D convolution and treat all features equally, failing to focus on information related to the center frame. Therefore, in this paper, we introduce an improved nonlocal block as a global attentive alignment (GAA) module, which takes the whole input video sequence into consideration to capture features that are globally correlated, to perform implicit alignment. Furthermore, given the bulk of features extracted from the center and neighboring frames, we propose target-guided attention (TGA). TGA can exploit more center-frame-related details and facilitate feature fusion. The proposed network (dubbed TANet) is capable of effectively eliminating false contours and recovering the center frame in high quality, as demonstrated by the experimental results. TANet outperforms state-of-the-art models in terms of both PSNR and SSIM with low time consumption.
Jing Liu 0002, Ziwen Yang, Yuting Su 0001, Xiaokang Yang 0001
IEEE Trans. Multim.3
2022 I-GCN: Incremental Graph Convolution Network for Conversation Emotion Detection
abstract
Sentiment analysis and emotion detection in conversation are becoming hot topics in regard to several applications. With the development of the social robot, social network, and intelligent voice assistant, emotion detection is attracting more attention as a key component in these research fields. Many approaches have been proposed to handle this problem in recent years. However, these previous approaches focus on either the temporal change information of the conversation or the semantic correlation information of the dialogue but ignore the combination of temporal information and semantic correlation information. In this paper, we propose an incremental graph convolution network (I-GCN) to handle emotion detection in conversation. We first utilize the graph structure to represent conversation at different times, which can represent the semantic correlation information of utterances. Then, we apply the incremental graph structure to imitate the process of dynamic conversation, which can preserve the temporal change information of conversation. Especially, for the first step of the process, we creatively propose utterance-level GCN (U-GCN) and speaker-level GCN (S-GCN) to learn the features of utterances for emotion detection. U-GCN focuses on the correlations among utterances and applies the multi-head attention model to find latent correlation information among utterances, which aims to further enhance the guidance of semantic relevance for feature learning. S-GCN focuses on the correlation between speaker and utterances, which can provide a different angle to guide the feature learning of utterances. In the learning of model parameters, we constantly utilize the new utterances to fine-tune the parameters of GNN for enhancement of the contribution of temporal change information. Detailed evaluations of the proposed method on three published conversation corpuses demonstrate the great effectiveness of our approach over several conventional competitive baselines.
Weizhi Nie, Rihao Chang, Minjie Ren, Yuting Su 0001, Anan Liu
IEEE Trans. Multim.4
2021 Object-Based Video Forgery Detection via Dual-Stream Networks
abstract
The object-based video forgery detection aims to expose tampered regions from video sequences without any codec information. However, existing methods mainly focus on manually selected features and models for a specific task, either splicing or copy-move, while the general representation ability of deep learning models and the correlation of different forensic features have not been fully explored. In this paper, we propose a dual-stream framework to jointly discover and integrate effective features for object-based video forgery detection. First, two different types of branches are employed to extract discriminative features. Then, after the dual-stream feature fusion, a Conditional Random Field (CRF) layer is utilized to further refine segmentation results. Finally, we consider temporal consistency by incorporating the video tracking strategy. Extensive experiments on four datasets show that the proposed method achieves competitive performance against the state-of-the-art methods.
Jing Xu 0008, Yuting Su 0001
ICME5
2021 Key Facial Components Guided Micro-Expression Recognition Based on First & Second-Order Motion
abstract
Although there have been many successful attempts in the field of micro-expression recognition, plenty of challenges remain due to the subtle spatio-temporal changes and high locality of micro-expressions. In this paper, to tackle such issues, we propose a novel key facial components guided micro-expression recognition approach (KFC-MER). Face semantic segmentation probability maps involving several key components provide a guidance for feature learning. With the Components-Aware Attention (CAA) module, expression-related areas are highlighted and the relationship between components will also be learned. To cope with the limited size of micro-expression datasets, we design a parallel shallow residual network as the MER network. Both the first- and second-order motion are exploited as the input data, for capturing motive information and non-rigid deformation, respectively. Extensive experiments on three benchmark datasets demonstrate that our method outperforms previous works and achieves state-of-the-art performance. The code is publicly available on GitHub: https://github.com/TJUMMG/KFC-MER.
Yuting Su 0001, Jing Liu 0002, Guangtao Zhai
ICME1
2021 SVHAN: Sequential View Based Hierarchical Attention Network for 3D Shape Recognition
abstract
As an important field of multimedia, 3D shape recognition has attracted much research attention in recent years. A lot of deep learning models have been proposed for effective 3D shape representation. The view-based methods show the superiority due to the comprehensive exploration of the visual characteristics with the help of established 2D CNN architectures. Generally, the current approaches contain the following disadvantages: First, the most majority of methods lack the consideration for sequential information among the multiple views, which can provide descriptive characteristics for shape representation. Second, the incomprehensive exploration for the multi-view correlations directly affects the discrimination of shape descriptors. Finally, roughly aggregating multi-view features leads to the loss of descriptive information, which limits the shape representation effectiveness. To handle these issues, we propose a novel sequential view based hierarchical attention network (SVHAN) for 3D shape recognition. Specifically, we first divide the view sequence into several view blocks. Then, we introduce a novel hierarchical feature aggregation module (HFAM), which hierarchically exploits the view-level, block-level, and shape-level features, the intra- and inter- view-block correlations are also captured to improve the discrimination of learned features. Subsequently, a novel selective fusion module (SFM) is designed for feature aggregation, considering the correlations between different levels and preserving effective information. Finally, discriminative and informative shape descriptors are generated for the recognition task. We validate the effectiveness of our proposed method on two public databases. The experimental results show the superiority of SVHAN against the current state-of-the-art approaches.
Yue Zhao 0042, Weizhi Nie, Anan Liu, Zan Gao 0002, Yuting Su 0001
ACM Multimedia5
2021 Joint nuclear- and ℓ2, 1-norm regularized heterogeneous tensor decomposition for robust classification
Peiguang Jing, Yuting Su 0001
Neurocomputing5
2021 Learning robust affinity graph representation for multi-view clustering
Peiguang Jing, Yuting Su 0001, Zhengnan Li, Liqiang Nie
Inf. Sci.2
2021 Deep low-rank matrix factorization with latent correlation estimation for micro-video multi-label classification
Yuting Su 0001, Junyu Xu, Daozheng Hong, Fugui Fan, Jing Zhang 0038, Peiguang Jing
Inf. Sci.1
2021 Exploring contextual information for view-wised 3D model retrieval
Wenhui Li 0002, Yuting Su 0001, Zhenlan Zhao
Multim. Tools Appl.2
2021 A Dual Rank-Constrained Filter Pruning Approach for Convolutional Neural Networks
abstract
Filter pruning has attracted increasing attentions to compress and accelerate the convolutional neural networks (CNNs) on computationally restricted devices. Existing related methods mainly focus on independently leveraging the spatial information of individual filters while ignoring the inner correlation among filters. In this letter, we propose a dual rank-constrained filter pruning approach for convolutional neural networks, in which the representation, clustering, and identification of representative filters are integrated into an adaptive graph regularization framework. Particularly, the proposed approach utilizes the low-rank constraint to capture the low-dimensional intrinsic representations of filters for the adaptive affinity graph construction and clustering. It is noteworthy that the original filters are projected as points on Grassmann manifold for geometrical structure preserving. Meanwhile, the high-rank constraint is employed to select the most informative filter for representing the cluster. Experimental results on CIFAR-10 dataset show that the proposed approach achieves competitive results compared with the state-of-the-art methods by 87.1% in VGGNet-16, 52.4% in GoogLeNet and 72.9% in ResNet-56 in terms of model parameters compression.
Fugui Fan, Yuting Su 0001, Peiguang Jing, Wei Lu 0026
IEEE Signal Process. Lett.2
2021 Scene Graph Inference via Multi-Scale Context Modeling
abstract
The scene graph generated for an image structurally represents its object interactions and it substantially aids image scene understanding. To the best of our knowledge, most current works on scene graph generation chiefly focus on pairwise object regions for object and relation inference while ignoring the global visual context outside of these regions. Guided by the intuition that object/relation inference can benefit from the visual context within an image, this paper proposes a multi-scale context modeling method, which can jointly discover and integrate the complementary object-centric and region-centric context for scene graph inference. While both the object-centric and region-centric contexts are separately modeled by their individual modules, a bi-directional message propagation strategy is designed to mutually reinforce the context modeling. A context-fused inference is then proposed to integrate the multi-scale context to guide scene graph inference. Extensive experiments establish that this method can achieve competitive performance compared to the state-of-the-art methods on three benchmarks. Additional ablation studies further validate its effectiveness. Code has been made available at: https://github.com/ningxu1990/MSCM.
Ning Xu 0003, Anan Liu, Yongkang Wong, Weizhi Nie, Yuting Su 0001, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.5
2021 Spatio-Temporal Mitosis Detection in Time-Lapse Phase-Contrast Microscopy Image Sequences: A Benchmark
abstract
In this paper, we report the results of the first international contest on mitosis detection in phase-contrast microscopy image sequences (https://www.iti-tju.org/mitosisdetection), which was held at the workshop of computer vision for microscopy image analysis (CVMI) in CVPR 2019. This contest aims to promote research on spatiotemporal mitosis detection under microscopy images. In this contest, we released a large-scale time-lapse phase-contrast microscopy image dataset (C2C12-16) for the mitosis detection task. Compared with the previous popular datasets (e.g., C2C12, C3H10), C2C12-16 contains more annotated mitotic events and more diverse cell culture environments. A total of ten different mitosis detection methods were submitted in the contest and evaluated on the test sets of four different cell culture environments in C2C12-16. In this benchmark, we describe all methods and conduct a thorough analysis based on their performances and discuss a feasible direction for mitosis detection. To the best of our knowledge, this is the first benchmark for the mitosis detection problem using a time-lapse phase-contrast microscopy spatiotemporal image sequence model.
Yuting Su 0001, Yao Lu 0005, Jing Liu 0002, Anan Liu
IEEE Trans. Medical Imaging1
2021 Learning Low-Rank Sparse Representations With Robust Relationship Inference for Image Memorability Prediction
abstract
Image memorability prediction aims to estimate the degree to which an image will be remembered by observers. Generally, the core problem in image memorability prediction is how to obtain effective representations to characterize the visual content of an image. In contrast to existing methods, which focus more on exploring the factors that make images memorable, in this paper, we first propose a general framework for learning joint low-rank and sparse principal feature representations, called the LSPFR framework, to obtain the lowest-rank intrinsic representation for image memorability prediction. By considering the joint optimization of the nuclear and$\ell _{1}$-norms, the global low-rank structure and the local patterns embedded in data can be exploited to make the learned features more robust and informative. To improve our framework based on the exploitation of sample relationship structure information, we present an extended version of LSPFR, named E-LSPFR, in which the underlying relationship structure matrix is inferred through a negative log-likelihood term with a sparsity constraint. The results of experiments conducted on four publicly available datasets confirm the superior performance of our proposed approaches.
Peiguang Jing, Yuechen Shang, Liqiang Nie, Yuting Su 0001, Jing Liu 0002, Meng Wang 0001
IEEE Trans. Multim.4
2021 Joint Intermediate Domain Generation and Distribution Alignment for 2D Image-Based 3D Objects Retrieval
abstract
2D image-based 3D object retrieval provides a convenient way to manage 3D big data with easily accessed 2D images. It is also a challenging task due to the significant differences between 2D images and 3D objects. In this paper, we propose a 2D image-based 3D object retrieval method, which can reduce the distribution discrepancy between 2D images and 3D objects and learn invariant features between them. Specifically, we first construct an intermediate domain module based on maximum mean discrepancy (MMD) in an unsupervised way, which can reduce the 2D and 3D distribution discrepancy by marginal distribution constraint. Second, to further reduce conditional distribution discrepancy and learn invariant features, we use source domain labels as semantic information to dynamically guide distribution alignment. Moreover, in order to support the research in 3D object retrieval, we contribute a new dataset, MDI3D. We conducted extensive experiments on MDI3D and some popular datasets, such as MI3DOR and SHREC2013. The experimental results demonstrate the superiority of the proposed method by comparing with the state-of-the-art methods.
Yuting Su 0001, Yuqian Li 0003, Dan Song 0006, Anan Liu, Jie Nie
IEEE Trans. Multim.1
2021 Generating Face Images With Attributes for Free
abstract
With superhuman-level performance of face recognition, we are more concerned about the recognition of fine-grained attributes, such as emotion, age, and gender. However, given that the label space is extremely large and follows a long-tail distribution, it is quite expensive to collect sufficient samples for fine-grained attributes. This results in imbalanced training samples and inferior attribute recognition models. To this end, we propose the use of arbitrary attribute combinations, without human effort, to synthesize face images. In particular, to bridge the semantic gap between high-level attribute label space and low-level face image, we propose a novel neural-network-based approach that maps the target attribute labels to an embedding vector, which can be fed into a pretrained image decoder to synthesize a new face image. Furthermore, to regularize the attribute for image synthesis, we propose to use a perceptual loss to make the new image explicitly faithful to target attributes. Experimental results show that our approach can generate photorealistic face images from attribute labels, and more importantly, by serving as augmented training samples, these images can significantly boost the performance of attribute recognition model. The code is open-sourced at this link.
Yaoyao Liu 0001, Qianru Sun, Xiangnan He 0001, Anan Liu, Yuting Su 0001, Tat-Seng Chua
IEEE Trans. Neural Networks Learn. Syst.5
2021 MMFN: Multimodal Information Fusion Networks for 3D Model Classification and Retrieval
abstract
In recent years, research into 3D shape recognition in the field of multimedia and computer vision has attracted wide attention. With the rapid development of deep learning, various deep models have achieved state-of-the-art performance based on different representations. There are many modalities for representing a 3D model, such as point cloud, multiview, and panorama view. Deep learning models based on these different modalities have different concerns, and all of them have achieved high performance for 3D shape recognition. However, all of these methods ignore the multimodality information in conditions where the same 3D model is represented by different modalities. Thus, we can obtain a better descriptor by guiding the training to consider these multiple representations. In this article, we propose MMFN, a novel multimodal fusion network for 3D shape recognition that employs correlations between the different modalities to generate a fused descriptor, which is more robust. In particular, we design two novel loss functions to help the model learn the correlation information during training. The first is correlation loss, which focuses on the correlations among different descriptors generated from different structures. This approach reduces the training time and improves the robustness of the fused descriptor of the 3D model. The second is instance loss, which preserves the independence of each modality and utilizes feature differentiation to guide model learning during the training process. More specifically, we use the weighted fusion method, which applies statistical methods to obtain robust descriptors that maximize the advantages of the information from the different modalities. We evaluated the proposed method on the ModelNet40 and ShapeNetCore55 datasets for 3D shape classification and retrieval tasks. The experimental results and comparisons with state-of-the-art methods demonstrate the superiority of our approach.
Weizhi Nie, Qi Liang 0004, Yuting Su 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2021 MV-LFN: Multi-view based local information fusion network for 3D shape recognition
abstract
3D shape recognition is a challenging task due to the difficulty of representing the complex structure of 3D shapes. Recently, the view-based approaches that utilize the multiple views rendered from the shape for visual information extraction and feature aggregation to generate a global shape descriptor , achieved promising performance. However, the view-based approaches commonly ignore the exploration and utilization of local information in the multiple views, which influences the effectiveness of generated features. In this paper, we design a novel Multi-view based Local Information Fusion Network (MV-LFN) for the 3D shape recognition task. The local correlation attention mechanism (LCAM) is introduced to exploit the local correlations in the feature maps for generating a more effective view descriptor. Then, we hierarchically aggregate the multi-view feature maps to generate a shape super matrix (SSM). The local information is effectively extracted and maintained during the multi-view aggregation process, and the discrimination of shape descriptors is significantly improved. We conduct comparative experiments on the ModelNet and ShapeNetCore55 databases. The experimental performances effectively validate the superiority of MV-LFN.
Jing Zhang 0038, Dangdang Zhou, Yue Zhao 0042, Weizhi Nie, Yuting Su 0001
Vis. Informatics5
2020 Mnemonics Training: Multi-Class Incremental Learning Without Forgetting
abstract
Multi-Class Incremental Learning (MCIL) aims to learn new concepts by incrementally updating a model trained on previous concepts. However, there is an inherent trade-off to effectively learning new concepts without catastrophic forgetting of previous ones. To alleviate this issue, it has been proposed to keep around a few examples of the previous concepts but the effectiveness of this approach heavily depends on the representativeness of these examples. This paper proposes a novel and automatic framework we call mnemonics, where we parameterize exemplars and make them optimizable in an end-to-end manner. We train the framework through bilevel optimizations, i.e., model-level and exemplar-level. We conduct extensive experiments on three MCIL benchmarks, CIFAR-100, ImageNet-Subset and ImageNet, and show that using mnemonics exemplars can surpass the state-of-the-art by a large margin. Interestingly and quite intriguingly, the mnemonics exemplars tend to be on the boundaries between different classes.
Yaoyao Liu 0001, Yuting Su 0001, Anan Liu, Bernt Schiele, Qianru Sun
CVPR2
2020 Consistent Domain Structure Learning and Domain Alignment for 2D Image-Based 3D Objects Retrieval
abstract
2D image-based 3D objects retrieval is a new topic for 3D objects retrieval which can be used to manage 3D data with 2D images. The goal is to search some related 3D objects when given a 2D image. The task is challenging due to the large domain gap between 2D images and 3D objects. Therefore, it is essential to consider domain adaptation problems to reduce domain discrepancy. However, most of the existing domain adaptation methods only utilize the semantic information from the source domain to predict labels in the target domain and neglect the intrinsic structure of the target domain. In this paper, we propose a domain alignment framework with consistent domain structure learning to reduce the large gap between 2D images and 3D objects. The domain structure learning module makes use of both the semantic information from the source domain and the intrinsic structure of the target domain, which provides more reliable predicted labels to the domain alignment module to better align the conditional distribution. We conducted experiments on two public datasets, MI3DOR and MI3DOR-2, and the experimental results demonstrate the proposed method outperforms the state-of-the-art methods.
Yuting Su 0001, Yuqian Li 0003, Dan Song 0006, Weizhi Nie, Wenhui Li 0001, Anan Liu
IJCAI1
2020 Multi-graph Convolutional Network for Unsupervised 3D Shape Retrieval
abstract
3D shape retrieval has attracted much research attention due to its wide applications in the fields of computer vision and multimedia. Various approaches have been proposed in recent years for learning 3D shape descriptor from different modalities. The existing works contain the following disadvantages: 1) the vast majority methods rely on the large scale of training data with clear category information; 2) many approaches focus on the fusion of multi-modal information but ignore the guidance of correlations among different modalities for shape representation learning; 3) many methods pay attention to the structural feature learning of 3D shape but ignore the guidance of structural similarity between every two shapes. To solve these problems, we propose a novel multi-graph network (MGN) for unsupervised 3D shape retrieval, which utilizes the correlations among modalities and structural similarity between two models to guide the shape representation learning process without category information. More specifically, we propose two novel loss functions: auto-correlation loss and cross-correlation loss. The auto-correlation loss utilizes information from different modalities to increase the discrimination of shape descriptor. The cross-correlation loss utilizes the structural similarity between two models to strengthen the intra-class similarity and increase the inter-class distinction. Finally, an effective similarity measurement is designed for the shape retrieval task. To validate the effectiveness of our proposed method, we conduct experiments on the ModelNet dataset. Experimental results demonstrate the effectiveness of our proposed method, and significant improvements have been achieved compared with state-of-the-art methods.
Weizhi Nie, Yue Zhao 0042, Anan Liu, Zan Gao 0002, Yuting Su 0001
ACM Multimedia5
2020 Domain-Specific Alignment Network for Multi-Domain Image-Based 3D Object Retrieval
abstract
2D image-based 3D object retrieval is a very important task in computer vision and big data management. Conventional image-based 3D object retrieval usually assumes that the images are from one single domain. However, for real applications, 2D images may be from multiple domains (e.g., real image, sketch, and quick draw). It raises significant challenges for this task since these 2D images have a great domain gap with each other as well as a great modality gap with 3D objects. To address these issues, we propose an unsupervised Domain-Specific Alignment Network (DSAN) for multi-domain image-based 3D object retrieval. The proposed method aims to reduce domain discrepancy by domain-specific alignment network with multi-level moment matching, including first-order moment and second-order moment. Based on the observation that for any given sample, different domain classifiers should output the same label, we design a domain-specific classifier alignment module. To our knowledge, the proposed method is the first unsupervised work to align multiple-domain 2D images with 3D objects in an end-to-end manner. The multi-domain dataset MDI3D is utilized to advocate the research on this task, and the extensive experimental results demonstrate the superiority of the proposed method.
Yuting Su 0001, Yuqian Li 0003, Dan Song 0006, Zhendong Mao 0001, Xuanya Li, Anan Liu
ACM Multimedia1
2020 Predicting the popularity of micro-videos via a feature-discrimination transductive model
Yuting Su 0001, Yang Li 0108, Peiguang Jing
Multim. Syst.1
2020 Single image super-resolution via low-rank tensor representation and hierarchical dictionary learning
Peiguang Jing, Weili Guan, Yuting Su 0001
Multim. Tools Appl.5
2020 An End-to-End Perceptual Quality Assessment Method via Score Distribution Prediction
Jing Liu 0002, Jingting Wang, Weizhi Nie, Yuting Su 0001, Anan Liu
Neural Process. Lett.4
2020 Hierarchical Deep Neural Network for Image Captioning
Yuting Su 0001, Yuqian Li 0003, Ning Xu 0003, Anan Liu
Neural Process. Lett.1
2020 Subgraph learning for graph matching
Weizhi Nie, Hai Ding, Anan Liu, Zonghui Deng, Yuting Su 0001
Pattern Recognit. Lett.5
2020 ABSNet: Aesthetics-Based Saliency Network Using Multi-Task Convolutional Network
abstract
As a smart visual attention mechanism to analyze visual scenes, visual saliency has been shown to closely correlate with semantic information such as faces. Although many semantic-information-guided saliency models have been proposed, to the best of our knowledge, no semantic information in affective domain has been employed for saliency detection. Aesthetic, the affective perceptual quality that integrates factors like scene composition and contrast, can certainly benefit visual attention that highly depends on these visual factors. In this letter, we propose an end-to-end multi-task framework called aesthetics-based saliency network (ABSNet). We use three commonly-used shared backbones and design two distinct branches for each task. Mean square error (MSE) loss and Earth Mover's Distance (EMD) loss are jointly adopted to alternately train the shared network and individual branch for different tasks, facilitating the proposed model to extract more effective features for visual perception. Moreover, our model is resolution-friendly to predict saliency for images of arbitrary size. It has been shown that the proposed multi-task method is superior over single-task version and outperforms state-of-the-art saliency methods.
Jing Liu 0002, Jincheng Lv, Jing Zhang 0038, Yuting Su 0001
IEEE Signal Process. Lett.5
2020 Low-Rank Regularized Deep Collaborative Matrix Factorization for Micro-Video Multi-Label Classification
abstract
Deep matrix factorization can be regarded as an extension of traditional matrix factorization to help improve applications like social image tag refinement, image retrieval, and face clustering. Toward this tendency, in this letter, we proposed a low-rank regularized deep collaborative matrix factorization (LRDCMF) method to better tackle micro-video multi-label classification tasks. The proposed method aims to collaboratively learn two sets of factor matrices for characterization of latent attributes and two deep representations for instances and labels, respectively. During factorization process, the inverse covariance constraints are exploited to capture the latent correlation structures among latent attributes and labels and the low-dimensional intrinsic deep representations are ensured by further considering low-rank constraints. Moreover, a triplet term that connects instances representations, label representations, and labels is constructed to increase discrimination power of our method. Experimental results conducted on a large-scale micro-video dataset illustrate our model achieves superior performance in comparison with state-of-the-art methods.
Yuting Su 0001, Daozheng Hong, Yang Li 0108, Peiguang Jing
IEEE Signal Process. Lett.1
2020 Joint Heterogeneous Feature Learning and Distribution Alignment for 2D Image-Based 3D Object Retrieval
abstract
2D image-based 3D object retrieval is a novel but challenging task for 3D object retrieval. In this paper, we propose a 2D image-based 3D object retrieval method via joint heterogeneous feature learning and distribution alignment. Specifically, we propose to learn a mapping function in the Grassmann manifold to reduce the divergence of heterogeneous features of 2D images and 3D objects. We further employ the data distribution alignment method to adaptively integrate both marginal and conditional distributions. We embed both terms into the objective function to learn a domain-invariant classifier based on structural risk minimization. The output domain-invariant features of 2D images and 3D objects can be utilized for 2D image-based 3D object retrieval. Since there is lack of large-scale dataset for the evaluation of this task, we build two new datasets, MI3DOR and MI3DOR-2. We compare the proposed method against the representative methods for domain adaption and explore the influence of different components of the objective functions and key parameter. Comparison experiments show the superiority of this method.
Yuting Su 0001, Yuqian Li 0003, Weizhi Nie, Dan Song 0006, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.1
2020 Sequential Saliency Guided Deep Neural Network for Joint Mitosis Identification and Localization in Time-Lapse Phase Contrast Microscopy Images
abstract
The analysis of cell mitotic behavior plays important role in many biomedical research and medical diagnostic applications. To improve the accuracy of mitosis detection in automated analysis systems, this paper proposes the sequential saliency guided deep neural network (SSG-DNN) to jointly identify and localize mitotic events in time-lapse phase contrast microscopy images. It consists of three key modules. First, the module of visual context learning extracts static visual feature and dynamic visual transition within individual volumetric cell regions. Secondly, with these information, the module of sequential saliency modeling aims to discover the saliency distribution over all successive frames in each volumetric region. Finally, the module of sequence structure modeling can leverage both visual context and saliency distribution for mitosis identification and localization. SSG-DNN can jointly realize visual feature learning and sequential structure modeling in the end-to-end framework. Moreover, the proposed method is independent of complicated preconditioning methods for mitotic candidate extraction and can be applied for mitosis detection in one-shot manner. To our knowledge, it is the first weakly supervised work to realize joint mitosis identification and localization only with sequence-wise labels. In our experiments, we evaluate its performances of both tasks on the popular C3H10 dataset and a novel and large-scale dataset, C2C12-16, which contains much more mitotic events and is more challenging owing to diverse cell culture conditions. Experimental results can demonstrate the superiority of the proposed method.
Yao Lu 0005, Anan Liu, Weizhi Nie, Yuting Su 0001
IEEE J. Biomed. Health Informatics5
2020 Low-Rank Regularized Multi-Representation Learning for Fashion Compatibility Prediction
abstract
The currently flourishing fashion-oriented community websites and the continuous pursuit of fashion have attracted the increased research interest of the fashion analysis community. Many studies show that predicting the compatibility of fashion outfits is a nontrivial task due to the difficulty in capturing the implicit patterns affecting fashion compatibility prediction and the complex relationships presented by raw data. To address these problems, in this paper, we propose a transductive low-rank hypergraph regularizer multiple-representation learning framework (LHMRL), whereby we formulate the processes of feature representation and fashion compatibility prediction in a joint framework. Specifically, we first introduce a low-rank regularized multiple-representation learning framework, in which the lowest-rank multiple representations of samples can be learned to characterize samples from different perspectives. In this framework, we maximize the total difference among multiple representations based on Grassmann manifold theory and incorporate a common hypergraph regularizer to naturally encode the complex relationships between fashion items and an outfit. To enhance the representation ability of our model, we then develop a supervised learning term by exploiting two types of supervision information from labeled data. Experiments on a publicly available large-scale dataset demonstrate the effectiveness of our proposed model over the state-of-the-art methods.
Peiguang Jing, Liqiang Nie, Jing Liu 0002, Yuting Su 0001
IEEE Trans. Multim.5
2020 Multi-Level Policy and Reward-Based Deep Reinforcement Learning Framework for Image Captioning
abstract
Image captioning is one of the most challenging tasks in AI because it requires an understanding of both complex visuals and natural language. Because image captioning is essentially a sequential prediction task, recent advances in image captioning have used reinforcement learning (RL) to better explore the dynamics of word-by-word generation. However, the existing RL-based image captioning methods rely primarily on a single policy network and reward function-an approach that is not well matched to the multi-level (word and sentence) and multi-modal (vision and language) nature of the task. To solve this problem, we propose a novel multi-level policy and reward RL framework for image captioning that can be easily integrated with RNN-based captioning models, language metrics, or visual-semantic functions for optimization. Specifically, the proposed framework includes two modules: 1) a multi-level policy network that jointly updates the word- and sentence-level policies for word generation; and 2) a multi-level reward function that collaboratively leverages both a vision-language reward and a language-language reward to guide the policy. Furthermore, we propose a guidance term to bridge the policy and the reward for RL optimization. The extensive experiments on the MSCOCO and Flickr30k datasets and the analyses show that the proposed framework achieves competitive performances on a variety of evaluation metrics. In addition, we conduct ablation studies on multiple variants of the proposed framework and explore several representative image captioning models and metrics for the word-level policy network and the language-language reward function to evaluate the generalization ability of the proposed framework.
Ning Xu 0003, Hanwang Zhang, Anan Liu, Weizhi Nie, Yuting Su 0001, Jie Nie, Yongdong Zhang 0001
IEEE Trans. Multim.5
2020 HGAN: Holistic Generative Adversarial Networks for Two-dimensional Image-based Three-dimensional Object Retrieval
abstract
In this article, we propose a novel method to address the two-dimensional (2D) image-based 3D object retrieval problem. First, we extract a set of virtual views to represent each 3D object. Then, a soft-attention model is utilized to find the weight of each view to select one characteristic view for each 3D object. Second, we propose a novel Holistic Generative Adversarial Network (HGAN) to solve the cross-domain feature representation problem and make the feature space of virtual characteristic view more inclined to the feature space of the real picture. This will effectively mitigate the distribution discrepancies across the 2D image domains and 3D object domains. Finally, we utilize the generative model of the HGAN to obtain the “virtual real image” of each 3D object and make the characteristic view of the 3D object and real picture possess the same feature space for retrieval. To demonstrate the performance of our approach, We established a new dataset that includes pairs of 2D images and 3D objects, where the 3D objects are based on the ModelNet40 dataset. The experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods.
Weizhi Nie, Weijie Wang 0002, Anan Liu, Jie Nie, Yuting Su 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2020 Multi-View Graph Matching for 3D Model Retrieval
abstract
3D model retrieval has been widely utilized in numerous domains, such as computer-aided design, digital entertainment, and virtual reality. Recently, many graph-based methods have been proposed to address this task by using multi-view information of 3D models. However, these methods are always constrained by many-to-many graph matching for the similarity measure between pairwise models. In this article, we propose a multi-view graph matching method (MVGM) for 3D model retrieval. The proposed method can decompose the complicated multi-view graph-based similarity measure into multiple single-view graph-based similarity measures and fusion. First, we present the method for single-view graph generation, and we further propose the novel method for the similarity measure in a single-view graph by leveraging both node-wise context and model-wise context. Then, we propose multi-view fusion with diffusion, which can collaboratively integrate multiple single-view similarities w.r.t. different viewpoints and adaptively learn their weights, to compute the multi-view similarity between pairwise models. In this way, the proposed method can avoid the difficulty in the definition and computation of the traditional high-order graph. Moreover, this method is unsupervised and does not require a large-scale 3D dataset for model learning. We conduct evaluations on four popular and challenging datasets. The extensive experiments demonstrate the superiority and effectiveness of the proposed method compared against the state of the art. In particular, this unsupervised method can achieve competitive performances against the most recent supervised and deep learning method.
Yuting Su 0001, Wenhui Li 0001, Weizhi Nie, Anan Liu
ACM Trans. Multim. Comput. Commun. Appl.1
2019 Photo-realistic image bit-depth enhancement via residual transposed convolutional neural network
Yuting Su 0001, Wanning Sun, Jing Liu 0002, Guangtao Zhai, Peiguang Jing
Neurocomputing1
2019 Wearable Computing for Internet of Things: A Discriminant Approach for Human Activity Recognition
abstract
With the rapid development of the wireless sensor network and the continuous improvement of its key technologies, the concept of Internet of Things has been encouraged and extended due to its wide applications in scenarios, such as smart homes and healthcare. Under the background, human activity recognition has drawn great attention in recent years. In this paper, we present a discriminant approach to recognize daily human activities recorded through accelerometer sensor. In the proposed approach, we first use S transform (ST) to extract features, and then introduce a supervised regularization-based robust subspace (SRRS) learning method to learn low-dimensional intrinsic feature representation from the original feature subspace. Particularly, ST has been described as a joint time-frequency representation, which is insensitive to noise. SRRS can learn more robust and discriminative features to reinforce the descriptions of samples while removing noise and redundancy. Experiments are conducted on three publicly available datasets, i.e., wireless sensor data mining, SCUT-NAA, and mHealth demonstrating the superior performance of our proposed scheme compared with state-of-the-art methods.
Wei Lu 0026, Fugui Fan, Jinghui Chu, Peiguang Jing, Yuting Su 0001
IEEE Internet Things J.5
2019 Scene graph captioner: Image captioning based on structural visual representation
Ning Xu 0003, Anan Liu, Jing Liu 0002, Weizhi Nie, Yuting Su 0001
J. Vis. Commun. Image Represent.5
2019 Pooled time series representation for mitosis event recognition
Yuting Su 0001, Weizhi Nie
Multim. Syst.1
2019 Multi-guiding long short-term memory for video captioning
Ning Xu 0003, Anan Liu, Weizhi Nie, Yuting Su 0001
Multim. Syst.4
2019 LSTM-based multi-label video event detection
Anan Liu, Yongkang Wong, Junnan Li 0001, Yuting Su 0001, Mohan Kankanhalli
Multim. Tools Appl.5
2019 Comprehensive image quality assessment via predicting the distribution of opinion score
Anan Liu, Jingting Wang, Jing Liu 0002, Yuting Su 0001
Multim. Tools Appl.4
2019 Smooth filtering identification based on convolutional neural networks
Anan Liu, Zhengyu Zhao 0001, Chengqian Zhang, Yuting Su 0001
Multim. Tools Appl.4
2019 The assessment of 3D model representation for retrieval with CNN-RNN networks
Weizhi Nie, Yuting Su 0001
Multim. Tools Appl.4
2019 Open-view human action recognition based on linear discriminant analysis
Yuting Su 0001, Yang Li 0108, Anan Liu
Multim. Tools Appl.1
2019 Tensor-driven low-rank discriminant analysis for image set classification
Jing Zhang 0038, Zhengnan Li, Peiguang Jing, Ye Liu 0002, Yuting Su 0001
Multim. Tools Appl.5
2019 Visual attribute detction for pedestrian detection
Jing Zhang 0038, Fuwu Li, Weizhi Nie, Wenhui Li 0001, Yuting Su 0001
Multim. Tools Appl.5
2019 A structure-transfer-driven temporal subspace clustering for video summarization
Jing Zhang 0038, Peiguang Jing, Jing Liu 0002, Yuting Su 0001
Multim. Tools Appl.5
2019 Low-rank regularized tensor discriminant representation for image set classification
Peiguang Jing, Yuting Su 0001, Zhengnan Li, Jing Liu 0002, Liqiang Nie
Signal Process.2
2019 A Framework of Joint Low-Rank and Sparse Regression for Image Memorability Prediction
abstract
Image memorability is to measure the degree to which an image is remembered. Generally image memorability prediction involves two steps: feature representation and prediction. Most previous work just focused on addressing the first step by investigating the factors of making an image memorable. They not only lack the use of a learning mechanism in feature representation, but also often neglect the second step. In this paper, we first propose a joint low-rank and sparse regression (JLRSR) framework to address this problem. JLRSR aims to jointly learn: 1) a low-rank projection matrix that enables us to decompose the original data into a component part and an error part and 2) a sparse regression coefficient vector for image memorability prediction. The projection matrix and the regression coefficients are bound by a sparse constraint to make our approach invariant to training samples. Moreover, a graph regularizer is constructed to improve the generalization performance and prevent overfitting. We then extend JLRSR to a multi-view version called Mv-JLRSR by imposing the block-wise constraint to ensure the group effect and the view correlation constraint to eliminate the heterogeneity among views. Experiment results validate the effectiveness of our proposed approaches.
Peiguang Jing, Yuting Su 0001, Liqiang Nie, Huimin Gu, Jing Liu 0002, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2019 3D Object Retrieval Based on Multi-View Latent Variable Model
abstract
View-based 3D object retrieval, in which multiple views are used for representation and retrieval, has attracted increasing attention due to its great flexibility. In this paper, we propose a discriminative multi-view latent variable model (MVLVM) for this task. Specifically, we design MVLVM to have an undirected graph structure in which the view set of a given 3D object is treated as the observations from which to discover the latent visual and spatial contexts. Then, we detail the learning and inference process of MVLVM for view-based 3D object retrieval. The proposed MVLVM has the following beneficial features: 1) it jointly learns visual and spatial contexts for 3D object modelling and 2) it avoids the difficulty of representative view extraction for model representation. Consequently, it can support flexible 3D model retrieval for real applications by avoiding camera array constraints, which severely constrain traditional methods. We report extensive experiments conducted on single-modal datasets (the NTU and ITI datasets) and a multi-modal dataset (MVRED-RGB and MVRED-Depth). These comparative experiments demonstrate the superiority of the proposed method.
Anan Liu, Weizhi Nie, Yuting Su 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 Hyper-Clique Graph Matching and Applications
abstract
This paper proposes a method for hyper-clique graph (HCG) generation, which can be considered an extension of classical graphs and hyper-graphs in which the node is replaced with the clique (a set of neighboring nodes in a specific feature space) and the hyper-edge linking multiple nodes is replaced with the hyper-edge linking multiple cliques. In addition, we propose the HCG matching method by preserving global and local structures. Specifically, we embed the clique relations of arbitrary orders in a high-order similarity tensor in a recursive manner. Then, we formulate the objective function of HCG matching with respect to two latent variables: the latent clique structure information in the original graph and the similarity measure of clique sets from pairwise HCGs. Since the objective function is not jointly convex with respect to both latent variables, we decompose it into two consecutive measurements for optimization: 1) a clique-to-clique similarity measurement by preserving local unary and pairwise correspondences and 2) a graph-to-graph similarity measurement by preserving global clique-to-clique correspondence. We suitably adopt the affinity-preserving reweighted random walks to optimize the objective function. We extensively evaluate the HCG matching performance on multiple applications: 1) we evaluate the robustness of HCG with respect to the deformation noise, the number of outliers, and the edge density on synthetic data and explore the effects of both the clique order and hyper-edge order on performance; 2) we explore HCG matching for feature point matching on multiple image data sets (CMU house sequence, Caltech+MSRC, and Car+Motor); and 3) we explore HCG matching for multi-view object retrieval, which is a much more challenging task since multi-view objects contain significant variations of illumination, viewpoint, and so on, using popular data sets (MV-RED and NTU). A comparison against the state-of-the-art methods demonstrates the superior performance of the proposed method.
Weizhi Nie, Anan Liu, Yue Gao 0002, Yuting Su 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 Dual-Stream Recurrent Neural Network for Video Captioning
abstract
Recent progress in using recurrent neural networks (RNNs) for video description has attracted an increasing interest, due to its capability to encode a sequence of frames for caption generation. While existing methods have studied various features (e.g., CNN, 3D CNN, and semantic attributes) for visual encoding, the representation and fusion of heterogeneous information from multi-modal spaces have not fully explored. Consider that different modalities are often asynchronous, frame-level multi-modal fusion (e.g., concatenation and linear fusion) will negatively influence each modality. In this paper, we propose a dual-stream RNN (DS-RNN) framework to jointly discover and integrate the hidden states of both visual and semantic streams for video caption generation. First, an encoding RNN is used for each stream to flexibly exploit the hidden states of respective modality. Specifically, we proposed an attentive multi-grained encoder module to enhance the local feature learning with global semantics feature. Then, a dual-stream decoder is deployed to integrate the asynchronous yet complementary sequential hidden states from both streams for caption generation. Extensive experiments on three benchmark datasets, namely, MSVD, MSR-VTT, and MPII-MD, show that DS-RNN achieves competitive performance against the state-of-the-art. Additional ablation studies were conducted on various variants of the proposed DS-RNN.
Ning Xu 0003, Anan Liu, Yongkang Wong, Yongdong Zhang 0001, Weizhi Nie, Yuting Su 0001, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.6
2019 High-Order Temporal Correlation Model Learning for Time-Series Prediction
abstract
Time-series prediction has become a prominent challenge, especially when the data are described as sequences of multiway arrays. Because noise and redundancy may exist in the tensor representation of a time series, we focus on solving the problem of high-order time-series prediction under a tensor decomposition framework and develop two novel multilinear models: 1) the multilinear orthogonal autoregressive (MOAR) model and 2) the multilinear constrained autoregressive (MCAR) model. The MOAR model is designed to preserve as much information as possible from the original tensorial data under orthogonal constraints. The MCAR model is an enhanced version that is developed by replacing orthogonal constraints with an inverse decomposition error term. For both models, we project the original tensor into subspaces spanned by basis matrices to facilitate the discovery of the intrinsic temporal structure embedded in the original tensor. To build connections among consecutive slices of the tensor, we generalize a traditional autoregressive model to tensor form to better preserve the temporal smoothness. Experiments conducted on four publicly available datasets demonstrate that our proposed methods converge within a small number of iterations during the training stage and achieve promising results compared with state-of-the-art methods.
Peiguang Jing, Yuting Su 0001, Chengqian Zhang
IEEE Trans. Cybern.2
2019 BE-CALF: Bit-Depth Enhancement by Concatenating All Level Features of DNN
abstract
There is a growing demand for monitors to provide high-quality visualization with more bits representing each rendered pixel. However, since most existing images and videos are of low bit-depth (LBD), transforming LBD images to visually pleasant high bit-depth (HBD) versions is of significant value. Most existing bit-depth enhancement methods generate unsatisfactory HBD images with annoying false contour artifacts or blurry details, and some algorithms are also time-consuming. To overcome these drawbacks, we propose a bit-depth enhancement framework via concatenating all level features of deep neural networks (DNNs). A novel deep learning network is proposed based on the deep convolutional variational auto-encoders (VAEs), and skip connections that concatenate every two layers are applied to pass low-level and high-level features to consequent layers, easing the gradient vanishing problem. Meanwhile, the proposed network is optimized to generate the residual between original images and its quantized ones, which performs better than recovering HBD images directly. The experimental results show that the proposed algorithm can eliminate false contour artifacts of the recovered HBD images with low time consumption, and can achieve dramatic restoration performance gains compared with state-of-the-art methods both subjectively and objectively.
Jing Liu 0002, Wanning Sun, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001
IEEE Trans. Image Process.3
2019 Multi-Domain and Multi-Task Learning for Human Action Recognition
abstract
Domain-invariant (view-invariant & modalityinvariant) feature representation is essential for human action recognition. Moreover, given a discriminative visual representation, it is critical to discover the latent correlations among multiple actions in order to facilitate action modeling. To address these problems, we propose a multi-domain & multi-task learning (MDMTL) method to (1) extract domain-invariant information for multi-view and multi-modal action representation and (2) explore the relatedness among multiple action categories. Specifically, we present a sparse transfer learning-based method to co-embed multi-domain (multi-view & multi-modality) data into a single common space for discriminative feature learning. Additionally, visual feature learning is incorporated into the multitask learning framework, with the Frobenius-norm regularization term and the sparse constraint term, for joint task modeling and task relatedness-induced feature learning. To the best of our knowledge, MDMTL is the first supervised framework to jointly realize domain-invariant feature learning and task modeling for multi-domain action recognition. Experiments conducted on the INRIA Xmas Motion Acquisition Sequences (IXMAS) dataset, the MSR Daily Activity 3D (DailyActivity3D) dataset, and the Multi-modal & Multi-view & Interactive (M2I) dataset, which is the most recent and largest multi-view and multi-model action recognition dataset, demonstrate the superiority of MDMTL over the state-of-the-art approaches.
Anan Liu, Ning Xu 0003, Weizhi Nie, Yuting Su 0001, Yongdong Zhang 0001
IEEE Trans. Image Process.4
2019 Spatiotemporal Symmetric Convolutional Neural Network for Video Bit-Depth Enhancement
abstract
In contrast to the high sensitivity of human eyes and rapid development of modern display devices in terms of dynamic range, mainstream multimedia sources are generally at relatively lower bit depths (BDs). Therefore, BD enhancement (BDE), which attempts to transform low-BD multimedia sources into high-BD sources, is considered of significant research value. Current BDE algorithms are based on images rather than videos. However, for massive numbers of videos, temporal continuity among frames should be considered. Thus, in this paper, we propose a spatiotemporal symmetric BDE network for videos based on an encoder-decoder network. Consecutive frames are input into five subnets in the encoder, where the convolutional filters in the temporal symmetric subnets share the same weights to achieve lower model complexity. In addition, symmetric skip connections are introduced between the symmetric convolutional/deconvolutional layers of the encoder/decoder to pass features and alleviate the gradient diffusion problem. The experimental results show that our model can efficiently eliminate false contours and chroma distortions. The model significantly outperforms state-of-the-art image BDE algorithms and single-frame baseline models in terms of PSNR and SSIM.
Jing Liu 0002, Pingping Liu, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001
IEEE Trans. Multim.3
2019 Multi-modal Correlated Network for emotion recognition in speech
abstract
With the growing demand of automatic emotion recognition system, emotion recognition is becoming more and more crucial for human–computer interaction (HCI) research. Recently, there is a continuous improvement in the performance of automatic emotion recognition due to the development of both hardware and deep learning methods. However, because of the abstract concept and multiple expressions of emotion, automatic emotion recognition is still a challenging task. In this paper, we propose a novel Multi-modal Correlated Network for emotion recognition aiming at exploiting the information from both audio and visual channels to achieve more robust and accurate detection. In the proposed method, the audio signals and visual signals are first preprocessed for the feature extraction. After preprocessing, we obtain the Mel-spectrograms, which can be treated as images, and the representative frames from visual segments. Then the Mel-spectrograms are fed to the convolutional neural network (CNN) to get the audio features and the representative frames are fed to the CNN and LSTM to get features. Specially, we employ the triplet loss to increase the differentiation of inter-class. Meanwhile, we propose a novel correlated loss to reduce the differentiation of intra-class. Finally, we apply the feature fusion method to fuse the audio and visual feature for emotion recognition classification. The experimental result on AEFW dataset demonstrates the correlation information of multiple modals is crucial for automatic emotion recognition and the proposed method can achieve the state-of-the-art performance on the classification task .
Minjie Ren, Weizhi Nie, Anan Liu, Yuting Su 0001
Vis. Informatics4
2018 Cross-Domain 3D Model Retrieval via Visual Domain Adaption
abstract
Recent advances in 3D capturing devices and 3D modeling software have led to extensive and diverse 3D datasets, which usually have different distributions. Cross-domain 3D model retrieval is becoming an important but challenging task. However, existing works mainly focus on 3D model retrieval in a closed dataset, which seriously constrain their implementation for real applications. To address this problem, we propose a novel crossdomain 3D model retrieval method by visual domain adaptation. This method can inherit the advantage of deep learning to learn multi-view visual features in the data-driven manner for 3D model representation. Moreover, it can reduce the domain divergence by exploiting both domainshared and domain-specific features of different domains. Consequently, it can augment the discrimination of visual descriptors for cross-domain similarity measure. Extensive experiments on two popular datasets, under three designed cross-domain scenarios, demonstrate the superiority and effectiveness of the proposed method by comparing against the state-of-the-art methods. Especially, the proposed method can significantly outperform the most recent method for cross-domain 3D model retrieval and the champion of Shrec’16 Large-Scale 3D Shape Retrieval from ShapeNet Core55.
Anan Liu, Shu Xiang, Wenhui Li 0001, Weizhi Nie, Yuting Su 0001
IJCAI5
2018 Multi-Level Policy and Reward Reinforcement Learning for Image Captioning
abstract
Image captioning is one of the most challenging hallmark of AI, due to its complexity in visual and natural language understanding. As it is essentially a sequential prediction task, recent advances in image captioning use Reinforcement Learning (RL) to better explore the dynamics of word-by-word generation. However, existing RL-based image captioning methods mainly rely on a single policy network and reward function that does not well fit the multi-level (word and sentence) and multi-modal (vision and language) nature of the task. To this end, we propose a novel multi-level policy and reward RL framework for image captioning. It contains two modules: 1) Multi-Level Policy Network that can adaptively fuse the word-level policy and the sentence-level policy for the word generation; and 2) Multi-Level Reward Function that collaboratively leverages both vision-language reward and language-language reward to guide the policy. Further, we propose a guidance term to bridge the policy and the reward for RL optimization. Extensive experiments and analysis on MSCOCO and Flickr30k show that the proposed framework can achieve competing performances with respect to different evaluation metrics.
Anan Liu, Ning Xu 0003, Hanwang Zhang, Weizhi Nie, Yuting Su 0001, Yongdong Zhang 0001
IJCAI5
2018 Hierarchical Graph Structure Learning for Multi-View 3D Model Retrieval
abstract
3D model retrieval has been widely utilized in numerous domains, such as computer-aided design, digital entertainment and virtual reality. Recently, many graph-based methods have been proposed to address this task by using multiple views of 3D models. However, these methods are always constrained by the many-to-many graph matching for similarity measure between pair-wise models. In this paper, we propose an hierarchical graph structure learning method (HGS) for 3D model retrieval. The proposed method can decompose the complicated multi-view graph-based similarity measure into multiple single-view graph-based similarity measures. In the bottom hierarchy, we present the method for single-view graph generation and further propose the novel method for similarity measure in single-view graph by leveraging both node-wise context and model-wise context. In the top hierarchy, we fuse the similarities in single-view graphs with respect to different viewpoints to get the multi-view similarity between pair-wise models. In this way, the proposed method can avoid the difficulty in definition and computation in the traditional high-order graph. Moreover, this method is unsupervised and is independent of large-scale 3D dataset for model learning. We conduct extensive evaluation on three popular and challenging datasets. The comparison demonstrates the superiority and effectiveness of the proposed method comparing with the state of the arts. Especially, this unsupervised method can achieve competing performance against the most recent supervised & deep learning method.
Yuting Su 0001, Wenhui Li 0001, Anan Liu, Weizhi Nie
IJCAI1
2018 Image Inpainting Detection Based on a Modified Formulation of Canonical Correlation Analysis
abstract
Image inpainting is a common image editing technique for filling the missing areas in images. It can be adopted to destroy the integrity of images by forgers with ulterior motives. Compared with other types of inpainting, sparsity-based inpainting assumes more general prior knowledge and is more widely used in practical applications. Although several methods for detecting exemplar-based and diffusion-based inpainting have been proposed, there is a shortage of effective scheme for detecting sparsity-based inpainting. In this paper, we proposed a novel algorithm for sparsity-based image inpainting detection. This type of inpainting has a strong effect on the coefficients of Canonical Correlation Analysis (CCA). Based on this observation, a modified objective function of CCA and a corresponding optimization algorithm are further developed to enhance the difference of inter-class in our feature set. The experiments implemented on two publicly available datasets demonstrated our method's superiority over other competitors. Particularly, unlike previous inpainting detection methods, the proposed framework has better performance in the case of JPEG compression.
Yuting Su 0001, Z. Jane Wang 0001
MMSP2
2018 A Review of Breast Cancer Detection in Medical Images
abstract
Breast cancer is a malignant tumor that occurs in the glandular epithelium of the breast. It is considered to be one of the most common cancers affecting women in the world. However, there is not an effective way to cure breast cancer yet, the key to reducing the risk of death is the early detection and diagnosis of breast cancer. Accurate diagnosis of breast cancer normally requires analysis of medical images of different modalities. There is a great need of automated system that could analyze these images accurately and rapidly. In this paper, we introduce some commonly used medical imaging methods for diagnosis of breast cancer, and based on them we investigate some recently proposed approaches for breast cancer detection with computer vision and machine learning techniques. Finally, we compare and analyze the detection performance of different methods on histological images and mammograph images respectively.
Yao Lu 0005, Jia-Yu Li, Yuting Su 0001, Anan Liu
VCIP3
2018 PANORAMA-Based Multi-Scale and Multi-Channel CNN for 3D Model Retrieval
abstract
In this paper, a novel method for the retrieval of 3D models is proposed. Firstly, the 2D panoramic view is extracted from each 3D model as representation. Secondly, a novel Multi-Scale and Multi-Channel CNN (MSMC-NN) is proposed to automatically generate the feature of each 3D model. Here, the Multi-Scale CNN can effectively save the local and global information of 3D model from panorama views. Finally, the features extracted from the full connect layer are used to handle the retrieval problem. In experiment section, the popular standard datasets ModelNet-10 and ModelNet-40 were used to demonstrate the performance of the proposed method. Meanwhile, some classic 3D model retrieval methods were leveraged as comparison methods in this paper. The corresponding experiments also demonstrate the superiority of our approach.
Weizhi Nie, Yao Lu 0005, Anan Liu, Yuting Su 0001
VCIP5
2018 Towards a sparse low-rank regression model for memorability prediction of images
Jinghui Chu, Huimin Gu, Yuting Su 0001, Peiguang Jing
Neurocomputing3
2018 HyperSSR: A hypergraph based semi-supervised ranking method for visual search reranking
Peiguang Jing, Yuting Su 0001, Chuan-Zhong Xu
Neurocomputing2
2018 Attention-in-Attention Networks for Surveillance Video Understanding in Internet of Things
abstract
In this paper, we propose an approach to generate the comprehensive video interpretation for the surveillance video understanding in Internet of Things. The key problem of many visual learning tasks is to adaptively select and fuse diverse and complimentary features for video representation. We design the attention-in-attention (AIA) network to hierarchically explore the attention fusion in an end-to-end manner, and demonstrate the value of this model on the multievent recognition and video captioning challenges. Particularly, it consists of multiple encoder attention modules (EAMs) and a fusion attention module (FAM). Each EAM aims to highlight the space-specific features by selecting the most salient visual features or semantic attributes and averages them into one attentive feature. The FAM can suppress or enhance the activation of multispace attentive features and adaptively co-embed them for comprehensive video representation. Then, one long short-term memory unit decodes the video representations to generate multiple event labels or video captions. This architecture is capable of: 1) adaptively learning the salient space-specific feature representation and 2) co-embedding multispace attentive features into one space for feature fusion. Experiments conducted on the surveillance video dataset (concurrent event dataset) and the popular video captioning datasets (Microsoft Research Video Description Corpus and MSR-Video to Text). It shows that the proposed AIA can achieve competitive performances against the state of the arts.
Ning Xu 0003, Anan Liu, Weizhi Nie, Yuting Su 0001
IEEE Internet Things J.4
2018 Low-rank regularized multi-view inverse-covariance estimation for visual sentiment distribution prediction
Anan Liu, Yingdi Shi, Peiguang Jing, Jing Liu 0002, Yuting Su 0001
J. Vis. Commun. Image Represent.5
2018 Graph regularized low-rank tensor representation for feature selection
Yuting Su 0001, Peiguang Jing, Jing Zhang 0038, Jing Liu 0002
J. Vis. Commun. Image Represent.1
2018 Video logo removal detection based on sparse representation
Yuting Su 0001, Liang Zou, Chengqian Zhang, Peiguang Jing, Xuemeng Song
Multim. Tools Appl.2
2018 3D model retrieval via single image based on feature mapping
Anan Liu, Nannan Liu, Weizhi Nie, Yuting Su 0001
Multim. Tools Appl.4
2018 View-based 3D model retrieval via supervised multi-view feature learning
Anan Liu, Weizhi Nie, Yuting Su 0001
Multim. Tools Appl.4
2018 Mitosis event recognition and detection based on evolution of feature in time domain
Weizhi Nie, Yuting Su 0001
Mach. Vis. Appl.5
2018 View-Based 3D Model Retrieval via Multi-graph Matching
Weizhi Nie, Anan Liu, Yahui Hao, Yuting Su 0001
Neural Process. Lett.4
2018 Structured low-rank inverse-covariance estimation for visual sentiment distribution prediction
Anan Liu, Yingdi Shi, Peiguang Jing, Jing Liu 0002, Yuting Su 0001
Signal Process.5
2018 Low-Rank Regularized Heterogeneous Tensor Decomposition for Subspace Clustering
abstract
This letter proposes a low-rank regularized heterogeneous tensor decomposition (LRRHTD) algorithm for subspace clustering, in which various constrains in different modes are incorporated to enhance the robustness of the proposed model. Specifically, due to the presence of noise and redundancy in the original tensor, LRRHTD seeks a set of orthogonal factor matrices for all but the last mode to map the high-dimensional tensor into a low-dimensional latent subspace. Furthermore, by imposing a low-rank constraint on the last mode, which is relaxed by using a nuclear norm, the lowest rank representation that reveals the global structure of samples is obtained for the purpose of clustering. We develop an effective algorithm based on the augmented Lagrange multiplier to optimize our model. Experiments on two public datasets demonstrate that our method reaches convergence within a small number of iterations and achieves promising results in comparison with the state of the arts.
Jing Zhang 0038, Peiguang Jing, Jing Liu 0002, Yuting Su 0001
IEEE Signal Process. Lett.5
2018 View-Based 3-D Model Retrieval: A Benchmark
abstract
View-based 3-D model retrieval is one of the most important techniques in numerous applications of computer vision. While many methods have been proposed in recent years, to the best of our knowledge, there is no benchmark to evaluate the state-of-the-art methods. To tackle this problem, we systematically investigate and evaluate the related methods by: 1) proposing a clique graph-based method and 2) reimplementing six representative methods. Moreover, we concurrently evaluate both hand-crafted visual features and deep features on four popular datasets (NTU60, NTU216, PSB, and ETH) and one challenging real-world multiview model dataset (MV-RED) prepared by our group with various evaluation criteria to understand how these algorithms perform. By quantitatively analyzing the performances, we discover the graph matching-based method with deep features, especially the clique graph matching algorithm with convolutional neural networks features, can usually outperform the others. We further discuss the future research directions in this field.
Anan Liu, Weizhi Nie, Yue Gao 0002, Yuting Su 0001
IEEE Trans. Cybern.4
2018 Low-Rank Multi-View Embedding Learning for Micro-Video Popularity Prediction
abstract
Recently, a prevailing trend of user generated content (UGC) on social media sites is the emerging micro-videos. Microvideos afford many potential opportunities ranging from network content caching to online advertising, yet there are still little efforts dedicated to research on micro-video understanding. In this paper, we focus on popularity prediction of micro-videos by presenting a novel low-rank multi-view embedding learning framework. We name it as transductive low-rank multi-view regression (TLRMVR), and it is capable of boosting the performance of micro-video popularity prediction by jointly considering the intrinsic representations of the source and target samples. In particular, TLRMVR integrates low-rank multi-view embedding and regression analysis into a unified framework such that the lowest-rank representation shared by all views not only captures the global structure of all views, but also indicates the regression requirements. The framework is formulated as a regression model and it seeks a set of view-specific projection matrices with low-rank constraints to map multi-view features into a common subspace. In addition, a multi-graph regularization term is constructed to improve the generalization capability and further prevents the overfitting problem. Extensive experiments conducted on a publicly available dataset demonstrate that our proposed method achieve promising results as compared with state-of-the-art baselines.
Peiguang Jing, Yuting Su 0001, Liqiang Nie, Jing Liu 0002, Meng Wang 0001
IEEE Trans. Knowl. Data Eng.2
2017 LingoSent - A Platform for Linguistic Aware Sentiment Analysis for Social Media Messages
Yuting Su 0001, Huijing Wang
MMM (1)1
2017 Multi-Camera Action Dataset for Cross-Camera Action Recognition Benchmarking
abstract
Action recognition has received increasing attention from the computer vision and machine learning communities in the last decade. To enable the study of this problem, there exist a vast number of action datasets, which are recorded under controlled laboratory settings, real-world surveillance environments, or crawled from the Internet. Apart from the "in-the-wild" datasets, the training and test split of conventional datasets often possess similar environments conditions, which leads to close to perfect performance on constrained datasets. In this paper, we introduce a new dataset, namely Multi-Camera Action Dataset (MCAD), which is designed to evaluate the open view classification problem under the surveillance environment. In total, MCAD contains 14,298 action samples from 18 action categories, which are performed by 20 subjects and independently recorded with 5 cameras. Inspired by the well received evaluation approach on the LFW dataset, we designed a standard evaluation protocol and benchmarked MCAD under several scenarios. The benchmark shows that while an average of 85% accuracy is achieved under the closed-view scenario, the performance suffers from a significant drop under the cross-view scenario. In the worst case scenario, the performance of 10-fold cross validation drops from 87.0% to 47.4%.
Wenhui Li 0001, Yongkang Wong, Anan Liu, Yang Li 0108, Yuting Su 0001, Mohan Kankanhalli
WACV5
2017 Hierarchical & multimodal video captioning: Discovering and transferring multimodal knowledge for vision to language
Anan Liu, Ning Xu 0003, Yongkang Wong, Junnan Li 0001, Yuting Su 0001, Mohan Kankanhalli
Comput. Vis. Image Underst.5
2017 3D models retrieval algorithm based on multimodal data
Anan Liu, Wenhui Li 0001, Weizhi Nie, Yuting Su 0001
Neurocomputing4
2017 Multi-view feature extraction based on slow feature analysis
Weizhi Nie, Anan Liu, Yuting Su 0001, Sha Wei
Neurocomputing3
2017 Multimedia venue semantic modeling based on multimodal data
Weizhi Nie, Wen-Juan Peng, Yi-Liang Zhao, Yuting Su 0001
J. Vis. Commun. Image Represent.5
2017 Hierarchical image resampling detection based on blind deconvolution
Yuting Su 0001, Chengqian Zhang, Yawei Chen
J. Vis. Commun. Image Represent.1
2017 Convolutional deep learning for 3D object retrieval
Weizhi Nie, Qun Cao, Anan Liu, Yuting Su 0001
Multim. Syst.4
2017 Median filtering forensics in digital images based on frequency-domain features
Anan Liu, Zhengyu Zhao 0001, Chengqian Zhang, Yuting Su 0001
Multim. Tools Appl.4
2017 3D object retrieval based on Spatial+LDA model
Weizhi Nie, Anan Liu, Yuting Su 0001
Multim. Tools Appl.4
2017 A spatial-temporal iterative tensor decomposition technique for action and gesture recognition
Yuting Su 0001, Haiyi Wang, Peiguang Jing, Chuan-Zhong Xu
Multim. Tools Appl.1
2017 Automatic report generation based on multi-modal information
Jing Zhang 0038, Weizhi Nie, Yuting Su 0001
Multim. Tools Appl.4
2017 Hierarchical Clustering Multi-Task Learning for Joint Human Action Grouping and Recognition
abstract
This paper proposes a hierarchical clustering multi-task learning (HC-MTL) method for joint human action grouping and recognition. Specifically, we formulate the objective function into the group-wise least square loss regularized by low rank and sparsity with respect to two latent variables, model parameters and grouping information, for joint optimization. To handle this non-convex optimization, we decompose it into two sub-tasks, multi-task learning and task relatedness discovery. First, we convert this non-convex objective function into the convex formulation by fixing the latent grouping information. This new objective function focuses on multi-task learning by strengthening the shared-action relationship and action-specific feature learning. Second, we leverage the learned model parameters for the task relatedness measure and clustering. In this way, HC-MTL can attain both optimal action models and group discovery by alternating iteratively. The proposed method is validated on three kinds of challenging datasets, including six realistic action datasets (Hollywood2, YouTube, UCF Sports, UCF50, HMDB51 & UCF101), two constrained datasets (KTH & TJU), and two multi-view datasets (MV-TJU & IXMAS). The extensive experimental results show that: 1) HC-MTL can produce competing performances to the state of the arts for action recognition and grouping; 2) HC-MTL can overcome the difficulty in heuristic action grouping simply based on human knowledge; 3) HC-MTL can avoid the possible inconsistency between the subjective action grouping depending on human knowledge and objective action grouping based on the feature subspace distributions of multiple actions. Comparison with the popular clustered multi-task learning further reveals that the discovered latent relatedness by HC-MTL aids inducing the group-wise multi-task learning and boosts the performance. To the best of our knowledge, ours is the first work that breaks the assumption that all actions are either independent for individual learning or correlated for joint modeling and proposes HC-MTL for automated, joint action grouping and modeling.
Anan Liu, Yuting Su 0001, Weizhi Nie, Mohan Kankanhalli
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Mitosis Detection in Phase Contrast Microscopy Image Sequences of Stem Cell Populations: A Critical Review
abstract
Detecting mitosis from cell population is a fundamental problem in many biological researches and biomedical applications. In modern researches, advanced imaging technologies have been applied to generate large amount of microscopy images of cells. However, detecting all mitotic cells from these images with human eye is tedious and time-consuming. In recent years, several approaches have been proposed to help humans finish this job automatically with high efficiency and accuracy. In this review paper, we first described some commonly used datasets for mitosis detection, and then discussed different kinds of methods for mitosis detection, like tracking based methods, tracking free methods, hybrid methods, and the most recently proposed works based on deep learning architecture. We compared these methods on same datasets, and found that deep learning based approaches have achieved a great improvement in performance. At last, we discussed the future possible approaches on mitosis detection, to combine the success of previous works and the advantage of big data in modern researches. Considering expertise is highly required in biomedical area, we will further discuss the possibility to learn information from biomedical big data with less expert annotation.
Anan Liu, Yao Lu 0005, Yuting Su 0001
IEEE Trans. Big Data4
2017 Modeling Temporal Information of Mitotic for Mitotic Event Detection
abstract
Due to the enormous potential and influence that stem cells may have in regenerative medicine, there has been a rapidly growing interest in developing tools to analyze and characterize the behaviors of these cells in vitro. Among these behaviors, mitosis, or cell division, is very important because stem cells proliferate and renew themselves through mitosis. However, current automated systems for mitosis detection often require traditional computer vision technology and machine learning methods; automated mitosis detection and recognition are difficult to achieve and mainly rely on manual annotation. In this paper, we proposed an effective method to capture video-wide temporal information for automated mitosis detection and recognition, which is a nondestructive imaging modality, thereby allowing continuous monitoring of cells in culture. In this approach, we postulate that a function capable of ordering the frames of a video temporally well captures the evolution of the appearance within the video. We learn such ranking functions per video via a ranking machine and use the parameters of these functions as a new video representation. Here, we utilized the CNN model (VGG-16) and some classic low-level feature extraction methods (HOG, SIFT, and GIST) to extract low-level features for each frame. The proposed method is easy to interpret and implement, fast to compute and effective in recognizing mitosis events. In a comparison experiment, our approach significantly outperformed previous approaches in terms of both detection accuracy and computational efficiency. The data that we validate the proposed method with includes C3H10 mesenchymal and C2C12 myoblastic stem cell populations. Our approach achieves an F-score of 95.8 percent on the C2C12 dataset and an F-score of 95.3 percent on the C3H10 dataset. The results on both datasets outperform traditional mitosis recognition methods based on probability models. These experiments all demonstrate the significance of our approach.
Weizhi Nie, Huiyun Cheng, Yuting Su 0001
IEEE Trans. Big Data3
2017 Benchmarking a Multimodal and Multiview and Interactive Dataset for Human Action Recognition
abstract
Human action recognition is an active research area in both computer vision and machine learning communities. In the past decades, the machine learning problem has evolved from conventional single-view learning problem, to cross-view learning, cross-domain learning and multitask learning, where a large number of algorithms have been proposed in the literature. Despite having large number of action recognition datasets, most of them are designed for a subset of the four learning problems, where the comparisons between algorithms can further limited by variances within datasets, experimental configurations, and other factors. To the best of our knowledge, there exists no dataset that allows concurrent analysis on the four learning problems. In this paper, we introduce a novel multimodal and multiview and interactive (M2I) dataset, which is designed for the evaluation of human action recognition methods under all four scenarios. This dataset consists of 1760 action samples from 22 action categories, including nine person-person interactive actions and 13 person-object interactive actions. We systematically benchmark state-of-the-art approaches on M2I dataset on all four learning problems. Overall, we evaluated 13 approaches with nine popular feature and descriptor combinations. Our comprehensive analysis demonstrates that M2I dataset is challenging due to significant intraclass and view variations, and multiple similar action categories, as well as provides solid foundation for the evaluation of existing state-of-the-art algorithms.
Anan Liu, Ning Xu 0003, Weizhi Nie, Yuting Su 0001, Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Cybern.4
2017 SnapVideo: Personalized Video Generation for a Sightseeing Trip
abstract
Leisure tourism is an indispensable activity in urban people's life. Due to the popularity of intelligent mobile devices, a large number of photos and videos are recorded during a trip. Therefore, the ability to vividly and interestingly display these media data is a useful technique. In this paper, we propose SnapVideo, a new method that intelligently converts a personal album describing of a trip into a comprehensive, aesthetically pleasing, and coherent video clip. The proposed framework contains three main components. The scenic spot identification model first personalizes the video clips based on multiple prespecified audience classes. We then search for some auxiliary related videos from YouTube according to the selected photos. To comprehensively describe a scenery, the view generation module clusters the crawled video frames into a number of views. Finally, a probabilistic model is developed to fit the frames from multiple views into an aesthetically pleasing and coherent video clip, which optimally captures the semantics of a sightseeing trip. Extensive user studies demonstrated the competitiveness of our method from an aesthetic point of view. Moreover, quantitative analysis reflects that semantically important spots are well preserved in the final video clip.
Peiguang Jing, Yuting Su 0001, Chao Zhang 0014, Ling Shao 0001
IEEE Trans. Cybern.3
2017 Multi-Grained Random Fields for Mitosis Identification in Time-Lapse Phase Contrast Microscopy Image Sequences
abstract
This paper proposes a multi-grained random fields (MGRFs) model for mitosis identification. To deal with the difficulty in hidden state discovery and sequential structure modeling in mitosis sequences only containing gradual visual pattern changes, we design the graphical structure to transform individual sequence into a set of coarse-to-fine grained sequencesconveying diverse temporal dynamics. Furthermore, we propose the corresponding probabilistic model for joint temporal learning and feature learning. To deal with the non-convex formulation of MGRF, we decomposemodel training into two sub-tasks, layer-wise sequential learning of both temporal dynamics and visual feature and new layer generation by graph-based sequential grouping, and optimize the model by alternating between them iteratively. The proposed method is validated on very challenging mitosis data set of C3H10T1/2 and C2C12 stem cells. Extensive comparison experiments demonstrate its superiority to the state of the arts.
Anan Liu, Jinhui Tang 0001, Weizhi Nie, Yuting Su 0001
IEEE Trans. Medical Imaging4
2017 Predicting Image Memorability Through Adaptive Transfer Learning From External Sources
abstract
Remembering images is an innate human capability. Camera images are captured by different people under varying environmental conditions, which leads to highly diverse image memorability scores. However, the factors that make an image more or less memorable are unclear, and it remains unknown how we can more accurately predict image memorability by using such factors. In this paper, we propose a novel framework called multiview transfer learning from external sources (MTLES) to predict image memorability. In this framework, we simultaneously leverage different types of visual feature sets and multiple types of predefined image attributes derived from external sources. In particular, to enhance representation ability of visual features, we construct connections between visual feature sets and higher level image attributes by transferring attribute knowledge from external sources. MTLES integrates weak learning through external sources, transfer learning, and multiview consistency loss with different types of feature sets into a joint framework. To better solve this joint optimization problem, we further develop an alternating iterative algorithm to deal with it. Experiments performed on the publicly available LaMem dataset demonstrate the effectiveness of the proposed scheme.
Peiguang Jing, Yuting Su 0001, Liqiang Nie, Huimin Gu
IEEE Trans. Multim.2
2016 HEp-2 cells Classification via clustered multi-task learning
Anan Liu, Yao Lu 0005, Weizhi Nie, Yuting Su 0001, Zhaoxuan Yang
Neurocomputing4
2016 Effective 3D object detection based on detector and tracker
Weizhi Nie, Anan Liu, Yuting Su 0001
Neurocomputing4
2016 Cross-view action recognition by cross-domain learning
Weizhi Nie, Anan Liu, Wenhui Li 0001, Yuting Su 0001
Image Vis. Comput.4
2016 3D object retrieval based on sparse coding in weak supervision
Weizhi Nie, Anan Liu, Yuting Su 0001
J. Vis. Commun. Image Represent.3
2016 Cross-domain semantic transfer from large-scale social media
Weizhi Nie, Anan Liu, Yuting Su 0001
Multim. Syst.3
2016 Geo-location driven image tagging via cross-domain learning
Weizhi Nie, Anan Liu, Yuting Su 0001
Multim. Syst.4
2016 Quality models for venue recommendation in location-based social network
Weizhi Nie, Anan Liu, Xiaorong Zhu, Yuting Su 0001
Multim. Tools Appl.4
2016 A Tensor-Driven Temporal Correlation Model for Video Sequence Classification
abstract
The task of video sequence classification plays a critical role in the development of computer vision. Considering this fact, this letter proposes a novel tensor decomposition method called tensor-driven temporal correlation in which general tensors are used as input for video sequence classification. Because distortion and redundancy may exist in the tensor representations of video sequences, we project the original tensor into subspaces spanned by spatial basis matrices in the proposed formulation. Moreover, to better preserve the temporal smoothness between consecutive slices of the tensor, the basis matrices are jointly learned by introducing an autoregressive model. An experiment on the commonly used Cambridge hand-gesture database demonstrates that our proposed method reaches convergence within a small number of iterations during the training stage and achieves promising results compared with state-of-the-art methods.
Jing Zhang 0038, Chuan-Zhong Xu, Peiguang Jing, Chengqian Zhang, Yuting Su 0001
IEEE Signal Process. Lett.5
2016 Multi-Modal Clique-Graph Matching for View-Based 3D Model Retrieval
abstract
Multi-view matching is an important but a challenging task in view-based 3D model retrieval. To address this challenge, we propose an original multi-modal clique graph (MCG) matching method in this paper. We systematically present a method for MCG generation that is composed of cliques, which consist of neighbor nodes in multi-modal feature space and hyper-edges that link pairwise cliques. Moreover, we propose an image set-based clique/edgewise similarity measure to address the issue of the set-to-set distance measure, which is the core problem in MCG matching. The proposed MCG provides the following benefits: 1) preserves the local and global attributes of a graph with the designed structure; 2) eliminates redundant and noisy information by strengthening inliers while suppressing outliers; and 3) avoids the difficulty of defining high-order attributes and solving hyper-graph matching. We validate the MCG-based 3D model retrieval using three popular single-modal data sets and one novel multi-modal data set. Extensive experiments show the superiority of the proposed method through comparisons. Moreover, we contribute a novel real-world 3D object data set, the multi-view RGB-D object data set. To the best of our knowledge, it is the largest real-world 3D object data set containing multi-modal and multi-view information.
Anan Liu, Weizhi Nie, Yue Gao 0002, Yuting Su 0001
IEEE Trans. Image Process.4
2015 Clique-graph matching by preserving global & local structure
abstract
This paper originally proposes the clique-graph and further presents a clique-graph matching method by preserving global and local structures. Especially, we formulate the objective function of clique-graph matching with respective to two latent variables, the clique information in the original graph and the pairwise clique correspondence constrained by the one-to-one matching. Since the objective function is not jointly convex to both latent variables, we decompose it into two consecutive steps for optimization: 1) clique-to-clique similarity measure by preserving local unary and pairwise correspondences; 2) graph-to-graph similarity measure by preserving global clique-to-clique correspondence. Extensive experiments on the synthetic data and real images show that the proposed method can outperform representative methods especially when both noise and outliers exist.
Weizhi Nie, Anan Liu, Zan Gao 0002, Yuting Su 0001
CVPR4
2015 Multi-modal & Multi-view & Interactive Benchmark Dataset for Human Action Recognition
abstract
Human action recognition is one of the most active research areas in both computer vision and machine learning communities. Several methods for human action recognition have been proposed in the literature and promising results have been achieved on the popular datasets. However, the comparison of existing methods is often limited given the different datasets, experimental settings, feature representations, and so on. In particularly, there are no human action dataset that allow concurrent analysis on three popular scenarios, namely single view, cross view, and cross domain. In this paper, we introduce a Multi-modal & Multi-view & Interactive (M2I) dataset, which is designed for the evaluation of the performances of human action recognition under multi-view scenario. This dataset consists of 1760 action samples, including 9 person-person interaction actions and 13 person-object interaction actions. Moreover, we respectively evaluate three representative methods for the single-view, cross-view, and cross domain human action recognition on this dataset with the proposed evaluation protocol. It is experimentally demonstrated that this dataset is extremely challenging due to large intraclass variation, multiple similar actions, significant view difference. This benchmark can provide solid basis for the evaluation of this task and will benefit advancing related computer vision and machine learning research topics.
Ning Xu 0003, Anan Liu, Weizhi Nie, Yongkang Wong, Fuwu Li, Yuting Su 0001
ACM Multimedia6
2015 Single/multi-view human action recognition via regularized multi-task learning
Anan Liu, Ning Xu 0003, Yuting Su 0001, Zhaoxuan Yang
Neurocomputing3
2015 Graph-based characteristic view set extraction and matching for 3D model retrieval
Anan Liu, Weizhi Nie, Yuting Su 0001
Inf. Sci.4
2015 Coupled hidden conditional random fields for RGB-D human action recognition
Anan Liu, Weizhi Nie, Yuting Su 0001, Zhaoxuan Yang
Signal Process.3
2015 Robust Image Hashing Based on Selective Quaternion Invariance
abstract
Robust image hashing, which maps the perceptual contents of image to a short digest, is a key technique for tackling the challenges of content-based indexing, searching and copyright protection. In this letter, we propose a quaternion invariance based hashing algorithm that can fuse complementary visual features to compact hash. The proposed algorithm leverages quaternion polar cosine transform to holistically capture the spatial and chromatic characteristics of digital image, and rotation-invariant features are derived from the phase information in the quaternion frequency domain. In addition, we also propose an information theoretic based metric for feature quality assessment and a greedy strategy to select the optimal subset of features for hash computation. Extensive experiments over a large database demonstrate that the proposed work shows higher content identification accuracy than most competing algorithms, and the resulting hash is compact and easy to compute.
Yuenan Li 0001, Yuting Su 0001
IEEE Signal Process. Lett.3
2015 Multipe/Single-View Human Action Recognition via Part-Induced Multitask Structural Learning
abstract
This paper proposes a unified framework for multiple/single-view human action recognition. First, we propose the hierarchical partwise bag-of-words representation which encodes both local and global visual saliency based on the body structure cue. Then, we formulate the multiple/single-view human action recognition as a part-regularized multitask structural learning (MTSL) problem which has two advantages on both model learning and feature selection: 1) preserving the consistence between the body-based action classification and the part-based action classification with the complementary information among different action categories and multiple views and 2) discovering both action-specific and action-shared feature subspaces to strengthen the generalization ability of model learning. Moreover, we contribute two novel human action recognition datasets, TJU (a single-view multimodal dataset) and MV-TJU (a multiview multimodal dataset). The proposed method is validated on three kinds of challenging datasets, including two single-view RGB datasets (KTH and TJU), two well-known depth dataset (MSR action 3-D and MSR daily activity 3-D), and one novel multiview multimodal dataset (MV-TJU). The extensive experimental results show that this method can outperform the popular 2-D/3-D part model-based methods and several other competing methods for multiple/single-view human action recognition in both RGB and depth modalities. To our knowledge, this paper is the first to demonstrate the applicability of MTSL with part-based regularization on multiple/single-view human action recognition in both RGB and depth modalities.
Anan Liu, Yuting Su 0001, Ping-Ping Jia, Zhaoxuan Yang
IEEE Trans. Cybern.2
2014 View-invariant feature discovering for multi-camera human action recognition
abstract
Intelligent video surveillance system is built to automatically detect events of interest, especially on object tracking and behavior understanding. In this paper, we focus on the task of human action recognition under surveillance environment, specifically in a multi-camera monitoring scene. Despite many approaches have achieved success in recognizing human action from video sequences, they are designed for single view and generally not robust against viewpoint invariant. Human action recognition across different views remains challenging due to the large variations from one view to another. We present a framework to solve the problem of transferring action models learned in one view (source view) to another view (target view). First, local space-time interest point feature and global shape-flow feature are extracted as low-level feature, followed by building the hybrid Bag-of-Words model for each action sequence. The data distribution of relevant actions from source view and target view are linked via a cross-view discriminative dictionary learning method. Through the view-adaptive dictionary pair learned by the method, the data from source and target view can be respectively mapped into a common space which is view-invariant. Furthermore, We extend our framework to transfer action models from multiple views to one view when there are multiple source views available. Experiments on the IXMAS human action dataset, which contains videos captured with five viewpoints, show the efficacy of our framework.
Lekha Chaisorn, Yongkang Wong, Anan Liu, Yuting Su 0001, Mohan Kankanhalli
MMSP5
2014 Multi-view action recognition by cross-domain learning
abstract
This paper proposes a novel multi-view human action recognition method by discovering and sharing common knowledge among different video sets captured in multiple viewpoints. To our knowledge, we are the first to treat a specific view as target domain and the others as source domains and consequently formulate the multi-view action recognition into the cross-domain learning framework. First, the classic bag-of-visual word framework is implemented for visual feature extraction in individual viewpoints. Then, we propose a cross-domain learning method with block-wise weighted kernel function matrix to highlight the saliency components and consequently augment the discriminative ability of the model. Extensive experiments are implemented on IXMAS, the popular multi-view action dataset. The experimental results demonstrate that the proposed method can consistently outperform the state of the arts.
Weizhi Nie, Anan Liu, Yuting Su 0001, Lekha Chaisorn, Yongkang Wong, Mohan Kankanhalli
MMSP4
2014 Single/cross-camera multiple-person tracking by graph matching
Weizhi Nie, Anan Liu, Yuting Su 0001, Huan-Bo Luan, Zhaoxuan Yang, Liujuan Cao, Rongrong Ji
Neurocomputing3
2013 Image Search Reranking with Semi-supervised LPP and Ranking SVM
Zhong Ji, Yanru Yu, Yuting Su 0001, Yanwei Pang
MMM (1)3
2013 An Effective Tracking System for Multiple Object Tracking in Occlusion Scenes
Weizhi Nie, Anan Liu, Yuting Su 0001, Zan Gao 0002
MMM (1)3
2013 Detection of JPEG double compression and identification of smartphone image source and post-capture manipulation
Qingzhong Liu, Peter A. Cooper, Lei Chen 0029, Hyuk Cho, Zhongxue Chen, Mengyu Qiao, Yuting Su 0001, Mingzhen Wei, Andrew H. Sung
Appl. Intell.7
2013 Ranking Fisher discriminant analysis
Zhong Ji, Peiguang Jing, Tianshi Yu, Yuting Su 0001, Changshu Liu
Neurocomputing4
2013 Balance between object and background: Object-enhanced features for scene image classification
Zhong Ji, Yuting Su 0001, Zhanjie Song, Shikai Xing
Neurocomputing3
2013 Detection of Double MPEG-2 Compression Based on Distributions of DCT coefficients
abstract
Detection of double compression in digital multimedia is considered to reveal the history of multimedia signal processing, which is very useful in digital forensics. In this paper, distributions of quantized DCT coefficients are analyzed in depth after double MPEG-2 compression with constant bit rate (CBR) mode, and a new algorithm is presented to detect double MPEG-2 compression based on convex patterns in the distribution of quantized DCT coefficients. In order to demonstrate the practical value of our proposed algorithm, two digital video cameras (DV) and two MPEG-2 software coders are utilized to simulate the general process of double MPEG-2 compression. Experiment results show that the proposed scheme can effectively detect doubly MPEG-2 compressed videos with CBR mode and the target output bit rate can vary in a large range.
Junyu Xu, Yuting Su 0001, Qingzhong Liu
Int. J. Pattern Recognit. Artif. Intell.2
2013 Rank canonical correlation analysis and its application in visual search reranking
Zhong Ji, Peiguang Jing, Yuting Su 0001, Yanwei Pang
Signal Process.3
2012 Multiple Person Tracking by Spatiotemporal Tracklet Association
abstract
In the field of video surveillance, multiple object tracking is a challenging problem in the real application. In this paper, we propose a multiple object tracking method by spatiotemporal tracklet association. Firstly, reliable tracklets, the fragments of the entire trajectory of individual object movement, are generated by frame-wise association between object localization results in the neighbor frames. To avoid the negative influence of occlusion on reliable tracklet generation, part-based similarity computation is performed. Secondly, the produced tracklets are associated considering both spatial and temporal constrains to output the entire trajectory for individual person. Especially, we formulate the task of spatiotemporal multiple tracklet matching into a Maximum A Posterior (MAP) problem in the form of Markov Chain with spatiotemporal context constraints. The experiment on PETS 2012 dataset demonstrates the superiority of the proposed method.
Weizhi Nie, Anan Liu, Yuting Su 0001
AVSS3
2011 Diversifying the Image Relevance Reranking with Absorbing Random Walks
abstract
Image visual reranking holds the simple search mechanism preferred by typical users, and exploits the visual information and image analysis methods in another way. Therefore, it integrates characteristics of real-time and accuracy, and has great importance to establish practical image search system. A novel reranking method named DIRRA is proposed in this paper, in which absorbing random walks is utilized to enhance the diversity as well as relevance of the initial search results. Four kinds of image visual features are extracted firstly, and then a graph is built, where nodes are images and edges are the similarities between images. Next, the first item is decided by teleporting random walks on the graph, and the other items are decided by absorbing random walks on the graph at last. Experiments are performed on a web image database including 10 queries, which prove the reranking results are both diverse and relevant, and practical to improve user's satisfaction in web search.
Zhong Ji, Yuting Su 0001, Yanwei Pang, Xiaojie Qu
ICIG2
2011 A video steganalytic algorithm against motion-vector-based steganography
Yuting Su 0001, Chengqian Zhang, Chuntian Zhang
Signal Process.1
2010 A Novel Source MPEG-2 Video Identification Algorithm
abstract
With the availability of powerful multimedia editing software, all types of personalized image and video resources are available in networks. Multimedia forensics technology has become a new topic in the field of information security. In this paper, a new source video system identification algorithm is proposed based on the features in the video stream; it takes full advantage of the different characteristics in the rate control module and the motion prediction module, which are two open parts in the MPEG-2 video compression standard, and combines a support vector machine classifier to build an intelligent computing system for video source identification. The experiments show this proposed algorithm can effectively identify video streams that come from a number of video coding systems.
Yuting Su 0001, Junyu Xu, Jing Zhang 0038, Qingzhong Liu
Int. J. Pattern Recognit. Artif. Intell.1
2010 Exposing Digital Video Logo-Removal Forgery by Inconsistency of Blur
abstract
A novel approach for detecting video logo-removal forgery is proposed by measuring inconsistency of blur. Our approach is based on the assumption that if a digital video undergoes logo-removal forgery; the blurriness of the forged region is expected to be different as compared to the nontampered parts of the video. Blurriness is first estimated by analyzing the spatial and temporal statistical property of logo areas, and suspicious areas are roughly located; then features are extracted and a fine classification is implemented by applying support vector machine (SVM) to extract features. If the suspicious areas and the reference areas are classified into different classes, the video is judged as a forged video. Experimental results show that our method is robust to video lossy compression for logo-removal forgery detection with the advantages of high classification accuracy and low computation cost.
Yuting Su 0001, Jing Zhang 0038, Qingzhong Liu
Int. J. Pattern Recognit. Artif. Intell.1
2007 Anchorperson Shot Detection in MPEG Domain
abstract
In this paper, a refined ASD algorithm in MPEG compressed domain is proposed. The new method is expected to outperform the existing strategies based on the following two improvements. One is that an effective face detection method is introduced. It aims to further remove the false alarms in the candidates generated from the previous modules, employing the chrominance DC coefficients. The other is the utilization of a new robust metric, which represents the dissimilarity of the video frames in the unsupervised clustering module. The proposed algorithm has been evaluated by six different TV channels. Compared with the state-of-the-art methods, the new algorithm is effective and computationally efficient.
Zhong Ji, Chuntian Zhang, Yuting Su 0001
ICME3
2007 Watermarking for Authentication of LZ-77 Compressed Documents
Yanfang Du, Jing Zhang 0038, Yuting Su 0001
IWDW3