EDBT 2026 Demo / reviewers in the wild / expert
Jianqing Sun
dblp:63/543
· DBLP profile ↗
10ranked-venue papers
0as first author
9since 2021 · last 2025
0009-0007-3598-8564ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimization of Multimodal Inputs Based on Diffusion Models: Zero-Shot Semantic Image GenerationabstractWith the continuous advancement of large models, the scale of data has become increasingly important in semantic segmentation tasks. However, the complexity and high cost of annotating semantic segmentation data pose significant challenges to the expansion of datasets. This study aims to leverage pre-trained diffusion generative models for conditional image generation, where labeled masks are used to generate corresponding synthetic images. This ensures a direct correspondence between input and output, effectively bypassing the annotation stage to reduce the cost of labor-intensive tasks. We employ multimodal conditions, to control the generation results. Additionally, we propose a multimodal alignment scheme to optimize the input control conditions, thereby improving the spatial structural accuracy of the generated results. Furthermore, we explore zero-shot generation tasks and successfully achieve zero-shot generation performance across multiple datasets, demonstrating the effectiveness of our approach. In the zero-shot generation experiments on the Cityscapes dataset, our method achieved a 0.6% improvement in the mIoU evaluation metric. On the ADE20K dataset, the performance improvement reached 2.52%, while on the COCO-Stuff dataset, the improvement was 2.43%. Leilei Wang, Renjie Lu 0001, Fengzhao Sun, Jun Yu 0001, Jianqing Sun, Jiaen Liang |
ICME | 7 |
| 2025 | Heterogeneous Encoder Fusion with KAN Decoder for Group Engagement Modeling via 8× Sliding PipelinesabstractEstimating engagement in group interactions is crucial for building socially intelligent systems, such as in human-agent and human-robot interaction. However, precisely modeling the continuous frame-level fluctuations of engagement remains challenging, particularly when considering the complex multi-party signal interactions within groups. Our method employs an encoder that integrates BiLSTM with Transformer to effectively capture both local and global temporal dependencies of multimodal features. Crucially, we explicitly fuse signals from both the target participant and their conversational partners in group to model the holistic group interaction dynamics. Furthermore, we introduce an 8x overlapped optimized sliding window strategy, constructing a ''sliding pipeline'', which significantly enhances the temporal smoothness, continuity, and stability of predictions. In the final regression stage, we replace the traditional multilayer perceptron(MLP) decoder with Kolmogorov-Arnold Network (KAN), leveraging their superior function approximation capability to achieve more accurate engagement predictions. Evaluated on the test sets of NoXi-base, NoXi-addition, MPIIGroupInteraction and NoXi-J datasets from the Multimediate'25 Engagement Challenge, our approach demonstrates significant performance improvements, achieving highly competitive Concordance Correlation Coefficients (CCC) of 0.678 for global, approximately 56.1% higher than the baseline, which shows a significant improvement. Yuefeng Zou, Hui Zhang 0044, Jun Yu 0001, Keda Lu, Lingsi Zhu, Fengzhao Sun, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 9 |
| 2024 | Micro-Expression Spotting Based on Optical Flow Feature with Boundary Calibration
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Zhongpeng Cai, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 8 |
| 2024 | Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global InteractionsabstractThe continual advancements in Generative Artificial Intelligence have created substantial hurdles for accurate deepfake detection, leading to limitations of currently popular detection methods across content-driven video-level deepfake detection scenarios. In this paper, we present the solutions to the Video-Level Deepfake Detection task. Our empirical findings demonstrate that modeling correlations of audio-visual modalities is important for video-level deepfake detection. Therefore, we introduce the model denoted Audio-Visual Local-Global Neural Network (i.e., AV-LGNN) in which the core design is the proposed AV-LGI Module (Audio-Visual Local-Global Interaction Module). The AV-LGI Module is composed of three stages: Local Intra-Region Interaction, Global Inter-Region Interaction, and Local-Global Interaction, which can better capture detailed information at local-level and efficiently learn the fine-grained correlations of inter-modalities in video deepfake detection under lower computational overheads. We further propose an adaptive modality selection strategy to facilitate model learning. Besides, a variety of data augmentation techniques are incorporated for audio-visual branches to enhance the robustness of the AV-LGNN. The experimental results verify the effectiveness of our model. Jia Zhang 0016, Mohan Jing, Keda Lu, Jun Yu 0001, Wen Su 0004, Fang Gao 0001, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 10 |
| 2024 | Temporal-Informative Adapters in VideoMAE V2 and Multi-Scale Feature Fusion for Micro-Expression Spotting-then-Recognize
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 8 |
| 2024 | RAG-Guided Large Language Models for Visual Spatial Description with Adaptive Hallucination CorrectorabstractVisual Spatial Description (VSD) is an emerging image-to-text task which aims at generating descriptions of the spatial relationships between given objects in an image. In this paper, we apply Retrieval-Augmented Generation (RAG) technology in guiding Multimodal Large Language Models (MLLMs) for the task of VSD, complemented by an Adaptive Hallucination Corrector, and further fine-tuning them to bolster semantic understanding and overall model efficacy. We found that our approach demonstrated higher accuracy and fewer hallucination errors in both spatial relationship classification and visual language description tasks within the VSD task, achieving state-of-the-art results. Jun Yu 0001, Gongpeng Zhao, Fengzhao Sun, Fanrui Zhang, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 9 |
| 2023 | M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech SynthesisabstractConversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational TTS systems only focus on extracting global information and omit local prosody features, which contain important fine-grained information like keywords and emphasis. Moreover, it is insufficient to only consider the textual features, and acoustic features also contain various prosody information. Hence, we propose M2-CTTS, an end-to-end multi-scale multi-modal conversational text-to-speech system, aiming to comprehensively utilize historical conversation and enhance prosodic expression. More specifically, we design a textual context module and an acoustic context module with both coarse-grained and fine-grained modeling. Experimental results demonstrate that our model mixed with fine-grained context information and additionally considering acoustic features achieves better prosody performance and naturalness in CMOS tests. Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li 0001, Yingming Gao, Jianhua Tao 0001, Jianqing Sun, Jiaen Liang |
ICASSP | 7 |
| 2023 | Answer-Based Entity Extraction and Alignment for Visual Text Question AnsweringabstractAs a variant of visual question answering (VQA), visual text question answering (VTQA) provides a text-image pair for each question. Text utilizes named entities to describe corresponding image. Consequently, the ability to perform multi-hop reasoning using named entities between text and image becomes critically important. However, existing models pay relatively less attention to this aspect. Therefore, we propose Answer-Based Entity Extraction and Alignment Model (AEEA) to enable a comprehensive understanding and support multi-hop reasoning. The core of AEEA lies in two main components: AKECMR and answer aware predictor. The former emphasizes the alignment of modalities and effectively distinguishes between intra-modal and inter-modal information, and the latter prioritizes the full utilization of intrinsic semantic information contained in answers during training. Our model outperforms the baseline by 2.24% on test-dev set and 1.06% on test set, securing the third place in VTQA2023(English). Jun Yu 0001, Mohan Jing, Weihao Liu 0004, Tongxu Luo, Keda Lu, Fangyu Lei, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 8 |
| 2023 | Sliding Window Seq2seq Modeling for Engagement EstimationabstractEngagement estimation in human conversations has been one of the most important research issues for natural human-robot interaction. However, previous datasets and studies mainly focus on the video-wise level of engagement estimation, therefore, can hardly reflect human's constantly changing engagement. Fortunately, the MultiMediate '23 challenge provides the frame-wise level of engagement estimation task. In this paper, we propose Sliding Window Seq2seq Modeling by BiLSTM and Transformer with powerful sequence modeling capabilities. Our method fully utilizes the global and local multi-modal feature information in the participants' videos and accurately expresses the engagement of the participants at each moment. Our method achieves the state-of-the-art CCC result of 0.71 for engagement estimation on the corresponding test sets. Jun Yu 0001, Keda Lu, Mohan Jing, Ziqi Liang, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 6 |
| 2020 | Speech Driven Talking Head Generation via Attentional Landmarks Based Representation
Jianqing Sun, Jiaen Liang |
INTERSPEECH | 3 |