Jinlin Guo

dblp:57/5947 · DBLP profile ↗
← Back
34ranked-venue papers
4as first author
16since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Beyond content: A dual-channel approach for social bot detection via unmasking behavioral sequence camouflage
Hongshuo Tian, Jinlin Guo, Xianzhu Liu, Ning Xu 0003, Lanjun Wang
Expert Syst. Appl.4
2026 Fine-grained talking motion consistent network for realistic face generation
Jinlin Guo, Xueliang Liu
Eng. Appl. Artif. Intell.2
2026 Towards social-aware image captioning via chain-of-thought prompting
Shenyuan Zhang, Ning Xu 0003, Quanhan Wu, Jinlin Guo, Hongshuo Tian, Anan Liu
Expert Syst. Appl.5
2026 Versatile and harmless deepfake proactive forensics via conditional watermarking
Xiaoshuai Wu, Xin Liao 0001, Jie Zhang 0004, Jinlin Guo
Inf. Sci.6
2026 How to Understand Named Entities: Using Commonsense for News Captioning
abstract
News captioning aims to describe an image with its news article body as input. It greatly relies on a set of detected named entities, including real-world people, organizations, and places. This article exploits commonsense knowledge to understand named entities for news captioning. By “understand,” we mean correlating the news content with commonsense in the wild, which helps an agent to (1) distinguish semantically similar named entities and (2) describe named entities using words outside of training corpora. Our approach consists of three modules: (a) Filter Module aims to clarify the commonsense concerning a named entity from two aspects: what does it mean ? and what is it related to ?, which divide the commonsense into explanatory knowledge and relevant knowledge , respectively. (b) Distinguish Module aggregates explanatory knowledge from node-degree , dependency , and distinguish three aspects to distinguish semantically similar named entities. (c) Enrich Module attaches relevant knowledge to named entities to enrich the entity description by commonsense information (e.g., identity and social position). Finally, all of information is integrated into the large multimodal model to generate the news caption. Extensive experiments on two challenging datasets (i.e., GoodNews and NYTimes) demonstrate the superiority of our method. Ablation studies and visualization further validate its effectiveness in understanding named entities.
Shenyuan Zhang, Ning Xu 0003, Yanhui Wang 0001, Tongle Ma, Wu Liu 0005, Jinlin Guo, Anan Liu
ACM Trans. Multim. Comput. Commun. Appl.6
2025 EmoGaussian High-Fidelity Emotional Talking Head Generation with 3D Gaussian Splatting
Jinlin Guo, Xueliang Liu
ICIC (6)3
2025 DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model
abstract
Recent advances in talking face generation have significantly improved facial animation synthesis. However, existing approaches face fundamental limitations: 3DMM-based methods maintain temporal consistency but lack fine-grained regional control, while Stable Diffusion-based methods enable spatial manipulation but suffer from temporal inconsistencies. The integration of these approaches is hindered by incompatible control mechanisms and semantic entanglement of facial representations. This paper presents DisentTalk, introducing a data-driven semantic disentanglement framework that decomposes 3DMM expression parameters into meaningful subspaces for fine-grained facial control. Building upon this disentangled representation, we develop a hierarchical latent diffusion architecture that operates in 3DMM parameter space, integrating region-aware attention mechanisms to ensure both spatial precision and temporal coherence. To address the scarcity of high-quality Chinese training data, we introduce CHDTF, a Chinese high-definition talking face dataset. Extensive experiments show superior performance over existing methods across multiple metrics, including lip synchronization, expression quality, and temporal consistency. Project Page: https://kangweiiliu.github.io/DisentTalk.
Kangwei Liu 0003, Junwu Liu, Yun Cao 0001, Jinlin Guo, Xiaowei Yi
ICME4
2025 Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space
abstract
Audio-driven emotional 3D facial animation encounters two significant challenges: (1) reliance on single-modal control signals (videos, text, or emotion labels) without leveraging their complementary strengths for comprehensive emotion manipulation, and (2) deterministic regression-based mapping that constrains the stochastic nature of emotional expressions and non-verbal behaviors, limiting the expressiveness of synthesized animations. To address these challenges, we present a diffusion-based framework for controllable expressive 3D facial animation. Our approach introduces two key innovations: (1) a FLAME-centered multimodal emotion binding strategy that aligns diverse modalities (text, audio, and emotion labels) through contrastive learning, enabling flexible emotion control from multiple signal sources, and (2) an attention-based latent diffusion model with content-aware attention and emotion-guided layers, which enriches motion diversity while maintaining temporal coherence and natural facial dynamics. Extensive experiments demonstrate that our method outperforms existing approaches across most metrics, achieving a 21.6% improvement in emotion similarity while preserving physiologically plausible facial dynamics. Project Page: https://kangweiiliu.github.io/Control_3D_Animation.
Kangwei Liu 0003, Junwu Liu, Xiaowei Yi, Jinlin Guo, Yun Cao 0001
ICME4
2025 Text-Vision Embedding for Generalized Diffusion Generated Videos Detection
Jinchuan Li, Jinlin Guo, Yun Cao 0001, Kangwei Liu 0003
PRCV (13)2
2025 DA-CCQ: Visually Explainable Image Forgery Localization via Difference Amplification and Cross-Clue Querying
Jinlin Guo, Yun Cao 0001, Jinchuan Li, Chengcheng Ma
PRCV (12)2
2025 Flexible Partial Screen-Shooting Watermarking With Provable Robustness
abstract
Screen-shooting watermarking is an effective means of protecting screen content from unauthorized capture and illegal dissemination. However, existing methods are primarily designed for full-image capture, making them ineffective for partial screen-shooting prevalent in real-world scenarios. To address this limitation, we propose FPSMark, a flexible watermarking method tailored for partial screen-shooting that embeds consistent watermarks in multiple uniformly distributed cover blocks. Specifically, considering that robustness requirements vary according to the layout of each image, we model the mathematical relationship between the watermark block count and robustness, proving the flexibility of FPSMark in ensuring partial screen-shooting robustness. Moreover, partial screen-shooting disrupts watermark synchronization, posing challenges for precise watermark localization. To overcome this, we design an intrinsic signal localization network optimized with a hybrid loss. The localization network exploits the inherent distinctions between the watermark and non-watermark features, while the hybrid loss constrains the network at three dimensions: pixel-level, region-level, and sample-level. Experimental results demonstrate the superiority of FPSMark, showing robust performance across partial capture percentages. Its extraction accuracy exceeds 98% even with only half of the image captured, and it achieves 82% accuracy at a 40% capture ratio, whereas existing methods achieve only around 50% under the same conditions.
Xin Liao 0001, Han Fang 0004, Jinlin Guo, Xiaoshuai Wu
IEEE Trans. Circuits Syst. Video Technol.4
2025 WaveRecovery: Screen-Shooting Watermarking Based on Wavelet and Recovery
abstract
The demand for resilient watermarking technology in the context of the screen-shooting scenario is steadily on the rise. The principal objective of this technique is to embed messages into the cover image, with the ability to effectively recover the message from the screen-captured image at the extraction end. However, current watermarking methods result in low visual quality watermarked images and are insufficiently robust in screen-shooting scenarios. This is mainly because they only utilize spatial domain information during embedding, and they do not consider the impact of noise that introduced during screen capturing. This paper introduces an innovative network framework, including the wavelet domain concatenation and recovery mechanism, to overcome the dual challenges encountered in robust watermarking, namely visual fidelity and robustness. For fidelity, we present a cascade network operating in the wavelet domain. This network excel at detecting watermark information in the wavelet domain. This capability makes it more sensitive to high and low-frequency details. Discrete wavelet transform can make CNN focus on different frequency characteristics, and the use of discrete inverse wavelet transform in upsampling can make the information high fidelity. As a result, it can more accurately identify and preserve critical visual details in this frequency domain, leading to an overall enhancement in visual quality. For robustness, a recovery network is specifically designed to mitigate the influence of noise introduced during screen-shooting on watermark information extraction. Experimental validation of our proposed method substantiates its effectiveness in significantly enhancing the visual quality and the accuracy of the watermarked images.
Linbo Fu, Xin Liao 0001, Jinlin Guo, Li Dong 0006, Zheng Qin 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Talking-DiSSM: Enhancing Temporal Consistency in Talking Face Video Generation with Bidirectional SSMs
abstract
Generating temporally smooth and high-resolution videos is a crucial objective in talking face generation tasks. Diffusion-based generative models have emerged as a prime choice for these tasks due to their ability to produce high-quality outputs. To mitigate the impact of stochasticity in the diffusion process, recent research has predominantly utilized self-attention layers to extract temporal features, ensuring temporal consistency in the generated videos. However, self-attention mechanisms have computational complexity that scales quadratically with video length, leading to high computational costs. This limitation poses significant challenges when attempting to generate longer video sequences using diffusion models. To address this challenge, we propose Talking-DiSSM, an end-to-end method for generating audio-driven talking face videos using State-Space Models (SSMs). This novel framework for conditional video diffusion modeling integrates Bidirectional State-Space Models (Bi-SSM) as temporal modeling modules with linear complexity, effectively capturing complex sequential temporal information and intra-batch sequential interdependencies in videos. Additionally, we employ a simple yet effective batch-overlapped sampling strategy to process input video clips, constructing inter-batch correlations while incorporating reference face clips and landmarks as conditions to ensure stability in the generation process. Extensive experiments demonstrate that Talking-DiSSM generates temporally consistent, high-quality, and identity-preserving talking face videos synchronized with the driving audio, achieving state-of-the-art results compared to existing models.
Xueliang Liu, Jinlin Guo, Richang Hong, Meng Wang 0001
ACM Trans. Intell. Syst. Technol.3
2023 Local Self-attention-based Hybrid Multiple Instance Learning for Partial Spoof Speech Detection
abstract
The development of speech synthesis technology has increased the attention toward the threat of spoofed speech. Although various high-performance spoofing countermeasures have been proposed in recent years, a particular scenario is overlooked: partially spoofed audio, where spoofed utterances may contain both spoofed and bona fide segments. Currently, the research on partially spoofed speech detection is lacking. The existing methods either train with partially spoofed speech at utterance level, resulting in gradient conflicting at the segment level, or directly train with segment level data, which requires segment labels that are difficult to obtain in practice. In this study, to better detect partially spoofed speech when only utterance labels are available, we formulate partially spoofed speech detection into a multiple instance learning (MIL) problem. The typical MIL uses a pooling layer to fuse patch scores as a whole, and we propose a hybrid MIL (H-MIL) framework based on max and log-sum-exp pooling methods, which can learn better segment representations to improve partially spoofed speech detection performance. Theoretical and experimental verification shows that H-MIL can effectively relieve the gradient conflicting and gradient vanishing problems. In addition, we analyze the local correlations between segments and introduce a local self-attention mechanism to enhance segment features, which further promotes the detection performance. In our experiments, we provide not only detection results at the segment and utterance levels but also some detailed visualization analysis, including the effect of spoof ratio and cross-dataset detection. The experimental results demonstrate the effective detection performance of our method at both the utterance and segment levels, especially when dealing with low spoof ratio attacks. The results confirm that our approach can better deal with partially spoofed speech detection than previous methods.
Yupeng Zhu, Zuxing Zhao, Xueliang Liu, Jinlin Guo
ACM Trans. Intell. Syst. Technol.5
2022 Can you trust what you hear: Effects of audio-attacks on voice-to-face generation system
abstract
Owing to the widespread deployment of face and speaker recognition systems, research on attacks on neural-network-based biometric systems, which involves face or voice signal classification problems with a low-dimensional output vector, has drawn increasing attention. Recently, cross-modal voice-to-face (VTF) systems have learned to generate faces from voices by matching several biometric characteristics of the generated faces to those of speakers. However, attacks focusing on VTF systems with high-dimensional face image outputs have not yet been conducted. In this paper, we introduce various adversarial attack methods for the VTF system under different attack conditions. These methods can generate a fake face close to the target face or far from the original face, by adding subtle perturbations to the original voice. Under the white-box setting, we formulate a multiobjective optimization to generate target faces and improve the imperceptibility of the adversarial sample. Further a stepwise iterative optimization strategy is proposed to achieve faster and more effective attacks. Finally, the results of comparative experiments with various methods are demonstrated. Under the black-box setting, the adversarial samples generated from surrogate models are able to generate the fake face far from the original one. Qualitative and quantitative experimental results show the high target face-matching rate and irrelevance to the original face, as well as the imperceptibility of the adversarial audio. This study provides useful insights for privacy protection and improving generation robustness for information security.
Yupeng Zhu, Jinlin Guo
Int. J. Intell. Syst.4
2021 DCA-CLA: A scRNA-seq Classification Framework based on Deep Count Autoencoder
abstract
Identifying cell types is crucial for single-cell RNA sequencing (scRNA-seq) analysis and can be potentially utilized to understand high-level biological processes. Supervised models based on neural networks have recently been successfully applied in the scRNA-seq cell type classification problem and achieved promising results. While most existing works directly use the raw or transformed data, we argue that the original data are too sparse and high-dimensional, and extracting their effective low-dimensional features can better train downstream classifiers, thereby improving the cell type classification performance. In this paper, we propose a novel framework, named Deep Count Autoencoder-based Classifier (DCA-CLA), to leverage the discriminative low-dimensional features for classification. Specifically, DCA-CLA first denoises the original count matrix and extracts the data features from the hidden layer using a deep count autoencoder module, then it feeds these bottleneck features into the classifier network to train the learnable parameters and test the performance. Experimental results on eight separate datasets and four pairs of datasets demonstrate that the proposed DCA-CLA framework achieves competitive performance over the state-of-the-art frameworks.
Yanming Guo, Songyang Lao, Jinlin Guo
IJCNN5
2020 PFNet: a novel part fusion network for fine-grained visual categorization
Jingyun Liang, Jinlin Guo, Yanming Guo, Songyang Lao
Multim. Tools Appl.2
2019 Semantically-enhanced kernel canonical correlation analysis: a multi-label cross-modal retrieval
Yuhua Jia, Peng Wang 0012, Jinlin Guo, Yuxiang Xie
Multim. Tools Appl.5
2018 Deep Convolutional Neural Network for Correlating Images and Sentences
Yuhua Jia, Peng Wang 0012, Jinlin Guo, Yuxiang Xie
MMM (1)4
2018 Irrelevance reduction with locality-sensitive hash learning for efficient cross-media retrieval
Yuhua Jia, Peng Wang 0012, Jinlin Guo, Yuxiang Xie
Multim. Tools Appl.4
2017 Utilizing Locality-Sensitive Hash Learning for Cross-Media Retrieval
Yuhua Jia, Peng Wang 0012, Jinlin Guo, Yuxiang Xie
MMM (1)4
2017 Deep Convolutional Neural Network for Bidirectional Image-Sentence Mapping
Jinlin Guo, Yuxiang Xie
MMM (2)3
2015 A Semi-automatic Solution Archive for Cross-Cut Shredded Text Documents Reconstruction
Shuxuan Guo, Songyang Lao, Jinlin Guo, Hang Xiang
ICIG (1)3
2013 Who produced this video, amateur or professional?
abstract
As the increasing affordability for capturing and storing video and the proliferation of Web 2.0 applications, video content is no longer necessarily created and supplied by a limited number of professional producers; any amateur can produce and publish his/her video quickly. Therefore, the amount of both professional-produced as well as amateur-produced video on the web is ever increasing. In this work, we propose a question; whether we can automatically classify an Internet video clip as being either professional-produced or amateur-produced? Hence, we investigate features and classification methods to answer this question. Based on the differences in the production processes of these two video categories, four features including camera motion, structure, audio feature and combined feature are adopted and studied along with with four popular classifiers KNN, SVM GMM and C4.5. Extensive experiments over representative datasets, evaluate these features and classifiers under different settings and compare to existing techniques. Experimental results demonstrate that SVMs with multimodal features from multi-sources are more effective at classifying video type. Finally, for answering the proposed question, results also show that automatically classifying a clip as professional-produced video or amateur-produced video can be achieved with good accuracy.
Jinlin Guo, Cathal Gurrin, Songyang Lao
ICMR1
2013 Quality Assessment of User-Generated Video Using Camera Motion
Jinlin Guo, Cathal Gurrin, Frank Hopfgartner, Songyang Lao
MMM (1)1
2013 Helping the Helpers: How Video Retrieval Can Assist Special Interest Groups
Frank Hopfgartner, Jinlin Guo, David Scott, Yang Yang 0076, Lijuan Marissa Zhou, Cathal Gurrin
MMM (2)2
2013 DCU at MMM 2013 Video Browser Showdown
David Scott, Jinlin Guo, Cathal Gurrin, Frank Hopfgartner, Kevin McGuinness, Noel E. O'Connor, Alan F. Smeaton, Yang Yang 0076
MMM (2)2
2013 Evaluating Novice and Expert Users on Handheld Video Retrieval Systems
David Scott, Frank Hopfgartner, Jinlin Guo, Cathal Gurrin
MMM (2)3
2013 Browsing Linked Video Archives of WWW Video
Cathal Gurrin, Jinlin Guo
MMM (2)3
2012 Supporting browsing of user generated video on a tablet
abstract
In this demo paper, we describe our user-generated video search system, compromising of an iPad interface communicating with a remote server. The goal of this system is to provide an easy access to video content lacking textual annotations by clustering key frames. Moreover, the graphical user interface allows users to filter video content based on various semantic concepts.
Frank Hopfgartner, David Scott, Jinlin Guo, Yang Yang 0076, Cathal Gurrin, Alan F. Smeaton
ICMR3
2012 Clipboard: A Visual Search and Browsing Engine for Tablet and PC
David Scott, Jinlin Guo, Yang Yang 0076, Frank Hopfgartner, Cathal Gurrin
MMM2
2011 Semantic concept detection in imbalanced datasets based on different under-sampling strategies
abstract
Semantic concept detection is a very useful technique for developing powerful retrieval or filtering systems for multimedia data. To date, the methods for concept detection have been converging on generic classification schemes. However, there is often imbalanced dataset or rare class problems in classification algorithms, which deteriorate the performance of many classifiers. In this paper, we adopt three “under-sampling” strategies to handle this imbalanced dataset issue in a SVM classification framework and evaluate their performances on the TRECVid 2007 dataset and additional positive samples from TRECVid 2010 development set. Experimental results show that our well-designed “under-sampling” methods (method SAK) increase the performance of concept detection about 9:6% overall. In cases of extreme imbalance in the collection the proposed methods worsen the performance than a baseline sampling method (method SI), however in the majority of cases, our proposed methods increase the performance of concept detection substantially. We also conclude that method SAK is a promising solution to address the SVM classification with not extremely imbalanced datasets.
Jinlin Guo, Colum Foley, Cathal Gurrin, Songyang Lao
ICME1
2011 Localization and Recognition of the Scoreboard in Sports Video Based on SIFT Point Matching
Jinlin Guo, Cathal Gurrin, Songyang Lao, Colum Foley, Alan F. Smeaton
MMM (2)1
2010 Prioritized Broadcast Contention Control in VANET
abstract
Reliable and timely multi-hop propagation of messages among vehicles is essential for a safer and greener transportation system. Various broadcast-based forwarding strategies are envisioned for infrastructure-less vehicle-to-vehicle (v2v) communications. This paper proposes a prioritized broadcast contention control (PBCC) module/layer that provides reliable and low latency multi-hop connection. The PBCC forwarding algorithm optimizes the back-off distribution to improve the probability of successful broadcast and prioritizes forwarders based on location information. This module can be implemented in WAVE devices with minimum system modification. We integrate simple vehicular mobility models into ns-2 and implement a WAVE/802.11p communication protocol stack. Extensive simulations demonstrate PBCC's superiority in multi-hop delay.
Fei Ye 0001, Raymond Yim, Jinlin Guo, Jinyun Zhang, Sumit Roy 0001
ICC3