Jin Liu 0009

dblp:01/2537-9 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0001-7249-698XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 4 first-author · 14 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Clustering with Self-Learned Graph Regression
abstract
Graph-based clustering algorithms aim to construct an affinity graph that accurately captures the intrinsic structure of a dataset. To achieve this goal, these algorithms often use the k-nearest-neighbor (k-nn) method to build a graph regularizer for the required affinity graph, enabling it to have a grouping effect. However, due to the complex nature of real-world data, the k-nn method often fails to capture the true neighborhood relationships of a dataset, which in turn limits the quality of the learned affinity graph. Motivated by the insight that a learned affinity graph itself can more effectively reflect the underlying data structure, we propose a new graph-based clustering method, termed Self-learned Graph Regression (SGR). Unlike traditional approaches, SGR constructs its graph regularizer directly from the affinity graph being learned, allowing the graph to adaptively capture more accurate structural information. To solve the proposed problem, we develop an optimization algorithm along with an acceleration strategy. We further analyze the convergence and computational complexity of the proposed algorithm. Extensive clustering experiments on various benchmark datasets demonstrate that our method outperforms the state-of-the-art graph-based clustering algorithms.
Lai Wei 0001, Jin Liu 0009
AAAI2
2026 Towards robust sentiment analysis with multimodal interaction graph and hybrid contrastive learning
Peizhu Gong, Jin Liu 0009, Xiliang Zhang, Xingye Li, Lai Wei 0001, Huihua He
Pattern Recognit.2
2026 Implicit Knowledge and Emotional Cues-Enhanced Multimodal Sarcasm Detection Model
abstract
Multimodal sarcasm detection aims to determine whether sarcasm exists in the interplay between text and accompanying images or audio. Effectively exploiting commonsense knowledge is crucial in multimodal sarcasm detection. However, previous methods only utilize the inherent attributes within modalities to construct superficial inconsistencies. They fail to fully leverage commonsense knowledge and implicit sarcasm cues and cannot uncover the deep contradictions between modalities. Accordingly, we propose a novel model named the Knowledge and Emotion Enhanced (KEE) model, which enhances sarcasm detection by integrating commonsense reasoning and cross-modal emotional contrast. Specifically, the KEE model leverages a generative language model to infer commonsense knowledge from the visual modality. It then selects highly relevant knowledge phrases through an attention mechanism and integrates them with text embeddings via a gated weighting mechanism to reveal sarcastic clues behind literal semantics. Additionally, the model captures emotional polarity inconsistencies across modalities and optimizes detection through a joint distribution framework. Experimental results demonstrate that the KEE model outperforms state-of-the-art methods on benchmark datasets.
Jin Liu 0009, Xingye Li, Huihua He
IEEE Trans. Affect. Comput.2
2026 Text-Centric Sparse Interaction Fusion Network With a Modality Calibrating Module for Multimodal Sentiment Analysis
abstract
Despite significant progress in Multimodal Sentiment Analysis (MSA), particularly with methods based on Transformers and Attentions, three major issues remain unresolved. First, existing unimodal representation learning methods overlook the noise and redundancy present within unimodal sequences. Second, current methods ignore the uneven distribution of emotional information across different modalities and fail to consider the highest contribution of text. Third, they neglect the issue of inter-modal noise redundancy in modal interactions. To address the aforementioned issues, this paper proposes a Text-Centric Sparse Interaction Fusion Network with a Modality Calibrating Module (TC-SIFN-MCM). A Modality Calibrating Module (MCM) is introduced during the unimodal representation learning phase to suppress noise within modalities. A Text-Centric Sparse Interaction Fusion Network (TC-SIFN) is designed to guide the learning of resource-poor non-text modalities with the resource-rich text modality, fully leveraging the highest contribution of the text modality. A Top-K Sparse Attention Transformer is utilized as a means of modal inter action, suppressing inter-modal noise generated during modal interactions. We conducted extensive experiments on the CMU MOSI, CMU-MOSEI, and SIMS datasets, and the results show that our TC-SIFN-MCM model outperforms current methods, with a notably impressive Acc-2 score of 86.07% on the MOSEI dataset.
Hongkun Zhou, Jin Liu 0009, Xingye Li, Huihua He
IEEE Trans. Affect. Comput.2
2025 Vessel re-identification by a hierarchical perceptual aggregation network with inclination-aware attention
abstract
Abstract Vessel re-identification (re-ID) is a crucial task in maritime supervision, enhancing maritime safety and improving the maritime situational awareness system. However, distinct from land-based scenarios involving vehicles or pedestrians, vessels, as enormous rigid bodies situated in the dynamic marine environment, face unique challenges such as significant variations in the scale of discriminative features and unpredictable sway. Furthermore, there is a limited number of publicly available datasets for vessel re-ID in complex backgrounds. In this paper, to overcome these challenges, a novel Hierarchical Perceptual Aggregation Network with Inclination-Aware Attention (HPAN-IAA) is proposed. HPAN-IAA comprises two main modules: the Hierarchical Perceptual Aggregation Block (HPAB) and the Inclination-Aware Attention Block (IAAB). Specifically, in HPAB, a hierarchical perceptual function is introduced to decompose visual information of vessels into discriminative features at multiple levels. These feature maps with different levels of detail from diverse network layers are then fused together by concatenation, resulting in a comprehensive feature representation that effectively integrates information across various scales. Conversely, to address the irregular variations and random omissions in discriminative feature distribution caused by unpredictable vessel sway, in IAAB, the Channel Collaborative Attention Module and the Pyramidal Spatial Attention Module are designed to adaptively extract potential discriminative features within each channel and spatial dimension, enhancing model’s ability in effectively extracting and utilizing irregularly changing discriminative features. Moreover, we propose a novel vessel re-ID dataset—VesselReID-2258. Extensive experiments conducted on VesselReID-2258 and the publicly available dataset VesselReID demonstrate that HPAN-IAA outperforms the current state-of-the-art methods,achieving superior performance with mean Average Precision scores of 0.861 and 0.823.
Yuetian Cao, Jin Liu 0009, Zijun Yu, Xingye Li, Lai Wei 0001, Zhongdai Wu
Comput. J.2
2025 Progressive feature refining with the deformable-guidance describer for dense video captioning
Jin Liu 0009, Huihua He
Expert Syst. Appl.2
2025 Goal-driven long-term marine vessel trajectory prediction with a memory-enhanced network
Xiliang Zhang, Jin Liu 0009, Chengcheng Chen, Lai Wei 0001, Zhongdai Wu, Wenjuan Dai
Expert Syst. Appl.2
2025 DNMCN: Dual-Stage Normalization Based Modality-Collaborative Fusion Network for Multimodal Sentiment Analysis
abstract
Due to the high-quality semantic information provided by the text modality, text-driven models have become the dominant approach for Multimodal Sentiment Analysis (MSA) in recent years. Despite notable progress in previous studies, two primary limitations remain: (i) aligning multimodal features often relies on simple matching of sequence length or feature dimension, which overlooks cross-modal heterogeneity. (ii) existing fusion techniques tend to over-rely on text, potentially diminishing the emotional data contributed by other modalities. To address these issues, in this paper, we propose a Dual-stage Normalization based Modality-Collaborative Fusion Network (DNMCN). Initially, to reduce modality discrepancies, we introduce a dual-stage normalization strategy, where features from different modalities were mapped into a common dimensional space in the first stage to facilitate effective cross-modal comparisons; sequence length inconsistencies caused by cropping and multiscale dimension reduction were addressed in the second stage. Additionally, to achieve high-quality cross-modal mapping without losing non-textual modality information, we propose an Adaptive modality-Collaborative Fusion Transformer (ACF-T) block. Specifically, in ACF-T block, textual semantics are first integrated into any non-text modality via multi-head attention. Next, a novel adaptive weighting strategy is introduced to balance the contribution of fused features and other non-textual modality features, thereby enhancing crossmodal interaction. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches on the public benchmark datasets CH-SIMS, CMU-MOSI and CMU-MOSEI.
Jin Liu 0009, Xingye Li, Jiajia Jiao, Huihua He
IEEE Trans. Affect. Comput.2
2025 AGT-Net: It Takes Two to Tango in Long-Term Person Reidentification
abstract
The key to address long-term person reidentification (Re-ID) in videos is to extract invariant spatio-temporal features (ISTF), which can be broadly categorized into two forms: 1) clothes-irrelevant appearance such as facial characteristics and body shape; 2) identity-distinctive motion such as posture and gait. However, existing studies mainly focus on mining either appearance- or motion-based features in the sequences without sufficient utilization of the ISTF. In this article, we propose an Appearance and gait-based tango network (AGT-Net) to comprehensively mine ISTF from appearance details and gait motions. Specifically, on the appearance detail branch, we introduce an invariance-aware video Swin transformer (IA-VST) to extract clothes-irrelevant appearance from the original RGB sequences. On the gait motion branch, we propose a motion-sensitive gait feature extractor (MS-GFE) to learn identity-distinctive motion from the pre-processed gait sequences. In addition, a score-level fusion strategy is introduced to integrate information from the two streams for prediction. Besides, since there is a lack of publicly available datasets, we propose a style-transferring synthetic long-term Video Re-ID (STYLE-VID) dataset, particularly for long-term Re-ID. Extensive experiments demonstrate that AGT-Net not only outperforms the state-of-the-art methods by up to 1.5% mAP on STYLE-VID, but also achieves comparable performance to other models on the traditional short-term Re-ID dataset MARS.
Zijun Yu, Jin Liu 0009, Peizhu Gong, Xingye Li, Lai Wei 0001, Huihua He, Zhongdai Wu
IEEE Trans. Comput. Soc. Syst.2
2025 Purely Contrastive Multiview Subspace Clustering
abstract
Multiview subspace clustering (MVSC) aims to integrate complementary information from different views to accurately reveal the subspace structure of a multiview dataset. Traditional MVSC methods often emphasize the aggregation of samples within the same subspace, while neglecting the separation of samples across different subspaces. In this article, we incorporate contrastive learning techniques into the MVSC framework, developing a contrastive data self-representation module, a contrastive regularizer for the reconstruction coefficient matrix in each view, and a contrastive alignment term to obtain a consensus coefficient matrix that fuses structural information from the reconstruction coefficient matrices. This leads to the framework of a purely contrastive MVSC (PCMVSC) approach. We elaborate on the superiority of the proposed modules in PCMVSC over similar ones in existing methods and show that the consensus reconstruction coefficient matrix obtained by PCMVSC can effectively uncover the underlying subspace structure of multiview datasets. Extensive subspace clustering experiments prove the effectiveness of PCMVSC and reveal that it outperforms various existing multiview clustering algorithms.
Lai Wei 0001, Rigui Zhou, Jin Liu 0009
IEEE Trans. Cybern.4
2025 SegCoT: Dependable Intrusion Detection System Based on Segment-Wise CoTransformer for Ship Communication Networks
abstract
Modern vessels integrate a massive digital infrastructure and navigation-dependent operating systems, allowing for ship-to-shore and ship-to-ship collaborative communication. However, the heightened interconnection of various maritime infrastructures inevitably amplifies the risk of vessel navigation and communication. Existing intrusion detection techniques were usually built on individual network events, failing to account for the multi-event long-term dependency problem caused by the high latency and low bandwidth of ship communication networks, therefore cannot tackle sophisticated cyber-ship attacks, resulting in lower accuracy in intrusion detection. In this paper, we propose a dependable Intrusion Detection System(IDS) based on Segment-wise CoTransformer(SegCoT) to detect cyber-ship intrusion events, which primarily contains a two-stage Network Pattern Extraction Component (NPEC) and an Intrusion Event Identification Component (IEIC). The NPEC automates the extraction of long-term dependency of massive intrusion events employing a SegEvent-wise Attention (SEA). Furthermore, the extracted dependencies are leveraged by the IEIC for specific intrusion type detection from a spatio-temporal feature fusion perspective. Based on a cyber-ship dataset collected from real ocean-going vessels, the proposed model achieves 99% intrusion detection accuracy, outperforming the existing state-of-the-art approaches.
Qiangqiang Shi, Jin Liu 0009, Lai Wei 0001, Jiajia Jiao, Bing Han 0009, Zhongdai Wu
IEEE Trans. Netw. Serv. Manag.2
2024 Ensemble based fully convolutional transformer network for time series classification
Yilin Dong 0001, Yuzhuo Xu, Rigui Zhou, Changming Zhu, Jin Liu 0009, Jiamin Song, Xinliang Wu
Appl. Intell.5
2024 MAGDRA: A Multi-modal Attention Graph Network with Dynamic Routing-By-Agreement for multi-label emotion recognition
Xingye Li, Jin Liu 0009, Yurong Xie, Peizhu Gong, Xiliang Zhang, Huihua He
Knowl. Based Syst.2
2024 MEDMCN: a novel multi-modal EfficientDet with multi-scale CapsNet for object detection
Xingye Li, Jin Liu 0009, Zhengyu Tang, Bing Han 0009, Zhongdai Wu
J. Supercomput.2
2024 Learning Idempotent Representation for Subspace Clustering
abstract
The critical point for the success of spectral-type subspace clustering algorithms is to seek reconstruction coefficient matrices that can faithfully reveal the subspace structures of data sets. An ideal reconstruction coefficient matrix should have two properties: 1) it is block-diagonal with each block indicating a subspace; 2) each block is fully connected. We find that a normalized membership matrix naturally satisfies the above two conditions. Therefore, in this paper, we devise an idempotent representation (IDR) algorithm to pursue reconstruction coefficient matrices approximating normalized membership matrices. IDR designs a new idempotent constraint. And by combining the doubly stochastic constraints, the coefficient matrices which are close to normalized membership matrices could be directly achieved. We present an optimization algorithm for solving IDR problem and analyze its computation burden as well as convergence. The comparisons between IDR and related algorithms show the superiority of IDR. Plentiful experiments conducted on both synthetic and real-world datasets prove that IDR is an effective subspace clustering algorithm.
Lai Wei 0001, Shiteng Liu, Rigui Zhou, Changming Zhu, Jin Liu 0009
IEEE Trans. Knowl. Data Eng.5
2023 Adaptive Graph Convolutional Subspace Clustering
abstract
Spectral-type subspace clustering algorithms have shown excellent performance in many subspace clustering applications. The existing spectral-type subspace clustering algorithms either focus on designing constraints for the reconstruction coefficient matrix or feature extraction methods for finding latent features of original data samples. In this paper, inspired by graph convolutional networks, we use the graph convolution technique to develop a feature extraction method and a coefficient matrix constraint simultaneously. And the graph-convolutional operator is updated iteratively and adaptively in our proposed algorithm. Hence, we call the proposed method adaptive graph convolutional subspace clustering (AGCSC). We claim that, by using AGCSC, the aggregated feature representation of original data samples is suitable for subspace clustering, and the coefficient matrix could reveal the subspace structure of the original data set more faithfully. Finally, plenty of subspace clustering experiments prove our conclusions and show that AGCSC11We present the codes of AGCSC and the evaluated algorithms on https://github.com/weilyshmtu/AGCSC. outperforms some related methods as well as some deep models.
Lai Wei 0001, Zhengwei Chen, Jun Yin 0003, Changming Zhu, Rigui Zhou, Jin Liu 0009
CVPR6
2023 Scale-Aware Regional Collective Feature Enhancement Network for Scene Object Detection
Yiyao Li, Jin Liu 0009
Neural Process. Lett.2
2022 Circulant-interactive Transformer with Dimension-aware Fusion for Multimodal Sentiment Analysis
Peizhu Gong, Jin Liu 0009, Xiliang Zhang, Xingye Li, Zijun Yu
ACML2
2021 Multiscale fully convolutional network-based approach for multilingual character segmentation
abstract
Abstract Character segmentation is a challenging task for optical character recognition systems. Traditional methods usually utilize rule‐based algorithms but most of them are not applicable in modern intelligent recognition applications that require high accuracy. It is especially the case for text containing Eastern Asian language characters with complex pictograph structures, such as Chinese. To alleviate this problem, this study proposes an encoder–decoder structure‐based multiscale fully convolutional network (MSFCN) model for optical character segmentation. Comparing with other methods, MSFCN can not only effectively extract semantic details from images but also exploit boundary information of intervals between characters, thereby distinguishing characters from a background in pixel level. Extensive experiments have been conducted on two benchmark data sets of ICDAR2013 and MLCS. Obtained results prove that MSFCN achieves state‐of‐the‐art segmentation performance and indicated its practical application value.
Jin Liu 0009, Yunhui Li
IET Comput. Vis.2
2020 Multi-level semantic representation enhancement network for relationship extraction
Jin Liu 0009, Yihe Yang, Huihua He
Neurocomputing1
2019 Multi-scale multi-class conditional generative adversarial network for handwritten character generation
Jin Liu 0009, Chenkai Gu, Jin Wang 0001, Geumran Youn, Jeong-Uk Kim
J. Supercomput.1
2018 Building neural network language model with POS-based negative sampling and stochastic conjugate gradient descent
Jin Liu 0009, Haoliang Ren, Minghao Gu, Jin Wang 0001, Geumran Youn, Jeong-Uk Kim
Soft Comput.1
2018 Multiple relations extraction among multiple entities in unstructured text
Jin Liu 0009, Haoliang Ren, Jin Wang 0001, Hye-Jin Kim 0003
Soft Comput.1
2018 Fine-grained entity type classification with adaptive context
Jin Liu 0009, Mingji Zhou, Jin Wang 0001, Sungyoung Lee 0001
Soft Comput.1