Keda Lu

dblp:294/2970 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2025
0009-0006-8974-3813ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Heterogeneous Encoder Fusion with KAN Decoder for Group Engagement Modeling via 8× Sliding Pipelines
abstract
Estimating engagement in group interactions is crucial for building socially intelligent systems, such as in human-agent and human-robot interaction. However, precisely modeling the continuous frame-level fluctuations of engagement remains challenging, particularly when considering the complex multi-party signal interactions within groups. Our method employs an encoder that integrates BiLSTM with Transformer to effectively capture both local and global temporal dependencies of multimodal features. Crucially, we explicitly fuse signals from both the target participant and their conversational partners in group to model the holistic group interaction dynamics. Furthermore, we introduce an 8x overlapped optimized sliding window strategy, constructing a ''sliding pipeline'', which significantly enhances the temporal smoothness, continuity, and stability of predictions. In the final regression stage, we replace the traditional multilayer perceptron(MLP) decoder with Kolmogorov-Arnold Network (KAN), leveraging their superior function approximation capability to achieve more accurate engagement predictions. Evaluated on the test sets of NoXi-base, NoXi-addition, MPIIGroupInteraction and NoXi-J datasets from the Multimediate'25 Engagement Challenge, our approach demonstrates significant performance improvements, achieving highly competitive Concordance Correlation Coefficients (CCC) of 0.678 for global, approximately 56.1% higher than the baseline, which shows a significant improvement.
Yuefeng Zou, Hui Zhang 0044, Jun Yu 0001, Keda Lu, Lingsi Zhu, Fengzhao Sun, Jianqing Sun, Jiaen Liang
ACM Multimedia4
2024 EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech
abstract
Recently, deep learning-based Text-to-Speech (TTS) systems have achieved high-quality speech synthesis results. Recurrent neural networks have become a standard modeling technique for sequential data in TTS systems and are widely used. However, training a TTS model which includes RNN components requires powerful GPU performance and takes a long time. In contrast, CNN-based sequence synthesis techniques can significantly reduce the parameters and training time of a TTS model while guaranteeing a certain performance due to their high parallelism, which alleviate these economic costs of training. In this paper, we propose a lightweight TTS system based on deep convolutional neural networks, which is a two-stage training end-to-end TTS model and does not employ any recurrent units. Our model consists of two stages: Text2Spectrum and SSRN. The former is used to encode phonemes into a coarse mel spectrogram and the latter is used to synthesize the complete spectrum from the coarse mel spectrogram. Meanwhile, we improve the robustness of our model by a series of data augmentations, such as noise suppression, time warping, frequency masking and time masking, for solving the low resource mongolian problem. Experiments show that our model can reduce the training time and parameters while ensuring the quality and naturalness of the synthesized speech compared to using mainstream TTS models. Our method uses NCMMSC2022-MTTSC Challenge dataset for validation, which significantly reduces training time while maintaining a certain accuracy.
Ziqi Liang, Haoxiang Shi, Keda Lu
CSCWD4
2024 Dialogue Cross-Enhanced Central Engagement Attention Model for Real-Time Engagement Estimation
Jun Yu 0001, Keda Lu, Ji Zhao 0020, Zhihong Wei, Iek-Heng Chu, Peng Chang 0002
IJCAI2
2024 Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global Interactions
abstract
The continual advancements in Generative Artificial Intelligence have created substantial hurdles for accurate deepfake detection, leading to limitations of currently popular detection methods across content-driven video-level deepfake detection scenarios. In this paper, we present the solutions to the Video-Level Deepfake Detection task. Our empirical findings demonstrate that modeling correlations of audio-visual modalities is important for video-level deepfake detection. Therefore, we introduce the model denoted Audio-Visual Local-Global Neural Network (i.e., AV-LGNN) in which the core design is the proposed AV-LGI Module (Audio-Visual Local-Global Interaction Module). The AV-LGI Module is composed of three stages: Local Intra-Region Interaction, Global Inter-Region Interaction, and Local-Global Interaction, which can better capture detailed information at local-level and efficiently learn the fine-grained correlations of inter-modalities in video deepfake detection under lower computational overheads. We further propose an adaptive modality selection strategy to facilitate model learning. Besides, a variety of data augmentation techniques are incorporated for audio-visual branches to enhance the robustness of the AV-LGNN. The experimental results verify the effectiveness of our model.
Jia Zhang 0016, Mohan Jing, Keda Lu, Jun Yu 0001, Wen Su 0004, Fang Gao 0001, Jianqing Sun, Jiaen Liang
ACM Multimedia5
2024 End-to-end Spatio-Temporal Information Aggregation For Micro-Action Detection
abstract
Micro-actions convey the emotions of characters in daily communication and offer richer semantic information compared to conventional actions. Accurate detection of these micro-actions is essential for video understanding. Due to their short duration, low intensity, and high overlap, micro-actions require more detailed video features, presenting a significant challenge for accurate detection. To address these challenges, we propose the 3D-SENet Adapter, which aggregates spatio-temporal information and enables end-to-end online video feature learning. We also find that incorporating background information significantly enhances the detection of small-scale micro-actions. Thus we develop the Cross-Attention Aggregation Detection Head, which integrates multi-scale features within the feature pyramid, thereby improving the detection accuracy of micro-actions occupying small regions in video frames. Our approach achieves first place in the Multi-label Micro-Action Detection (MMAD) and second place in the Micro-Action Recognition (MAR) of Micro-Action Analysis Grand Challenge.
Jun Yu 0001, Mohan Jing, Guopeng Zhao, Keda Lu, Feng Zhao 0005, Jiaqing Sun, Jiaen Liang
ACM Multimedia4
2023 Answer-Based Entity Extraction and Alignment for Visual Text Question Answering
abstract
As a variant of visual question answering (VQA), visual text question answering (VTQA) provides a text-image pair for each question. Text utilizes named entities to describe corresponding image. Consequently, the ability to perform multi-hop reasoning using named entities between text and image becomes critically important. However, existing models pay relatively less attention to this aspect. Therefore, we propose Answer-Based Entity Extraction and Alignment Model (AEEA) to enable a comprehensive understanding and support multi-hop reasoning. The core of AEEA lies in two main components: AKECMR and answer aware predictor. The former emphasizes the alignment of modalities and effectively distinguishes between intra-modal and inter-modal information, and the latter prioritizes the full utilization of intrinsic semantic information contained in answers during training. Our model outperforms the baseline by 2.24% on test-dev set and 1.06% on test set, securing the third place in VTQA2023(English).
Jun Yu 0001, Mohan Jing, Weihao Liu 0004, Tongxu Luo, Keda Lu, Fangyu Lei, Jianqing Sun, Jiaen Liang
ACM Multimedia6
2023 Sliding Window Seq2seq Modeling for Engagement Estimation
abstract
Engagement estimation in human conversations has been one of the most important research issues for natural human-robot interaction. However, previous datasets and studies mainly focus on the video-wise level of engagement estimation, therefore, can hardly reflect human's constantly changing engagement. Fortunately, the MultiMediate '23 challenge provides the frame-wise level of engagement estimation task. In this paper, we propose Sliding Window Seq2seq Modeling by BiLSTM and Transformer with powerful sequence modeling capabilities. Our method fully utilizes the global and local multi-modal feature information in the participants' videos and accurately expresses the engagement of the participants at each moment. Our method achieves the state-of-the-art CCC result of 0.71 for engagement estimation on the corresponding test sets.
Jun Yu 0001, Keda Lu, Mohan Jing, Ziqi Liang, Jianqing Sun, Jiaen Liang
ACM Multimedia2
2021 Aspect-Level Sentiment Analysis Approach via BERT and Aspect Feature Location Model
abstract
With the rapid development of Internet social platforms, buyer shows (such as comment text) have become an important basis for consumers to understand products and purchase decisions. The early sentiment analysis methods were mainly text‐level and sentence‐level, which believed that a text had only one sentiment. This phenomenon will cover up the details, and it is difficult to reflect people’s fine‐grained and comprehensive sentiments fully, leading to people’s wrong decisions. Obviously, aspect‐level sentiment analysis can obtain a more comprehensive sentiment classification by mining the sentiment tendencies of different aspects in the comment text. However, the existing aspect‐level sentiment analysis methods mainly focus on attention mechanism and recurrent neural network. They lack emotional sensitivity to the position of aspect words and tend to ignore long‐term dependencies. In order to solve this problem, on the basis of Bidirectional Encoder Representations from Transformers (BERT), this paper proposes an effective aspect‐level sentiment analysis approach (ALM‐BERT) by constructing an aspect feature location model. Specifically, we use the pretrained BERT model first to mine more aspect‐level auxiliary information from the comment context. Secondly, for the sake of learning the expression features of aspect words and the interactive information of aspect words’ context, we construct an aspect‐based sentiment feature extraction method. Finally, we construct evaluation experiments on three benchmark datasets. The experimental results show that the aspect‐level sentiment analysis performance of the ALM‐BERT approach proposed in this paper is significantly better than other comparison methods.
Guangyao Pang, Keda Lu, Zhiyi Mo, Zizhen Peng, Baoxing Pu
Wirel. Commun. Mob. Comput.2