EDBT 2026 Demo / reviewers in the wild / expert
Jiaen Liang
dblp:01/7824
· DBLP profile ↗
31ranked-venue papers
1as first author
17since 2021 · last 2026
0009-0001-8309-1301ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 12 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FocalOrder: Focal Preference Optimization for Reading Order DetectionabstractFuyuan Liu, Dianyu Yu, He Ren, Nayu Liu, Xiaomian Kang, Delai Qiu, Fa Zhang, Genpeng Zhen, Shengping Liu, Liang Jiaen, Weihuang, Yining Wang, Junnan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Fuyuan Liu, Dianyu Yu, Nayu Liu, Xiaomian Kang, Delai Qiu, Genpeng Zhen, Shengping Liu, Jiaen Liang, Junnan Zhu |
ACL (1) | 10 |
| 2026 | MetaDB: Metadata-Guided Diffusion Bridge Model for High-Fidelity Medical Image SynthesisabstractMedical image synthesis is pivotal in modern clinical workflows, addressing the issue of missing imaging modalities. While diffusion-based models have shown promise, existing approaches often neglect the rich clinical metadata, leading to synthesized images that lack semantic fidelity and fail to maintain strict consistency with the target modality. To address these challenges, we propose a metadata-guided diffusion bridge model, termed MetaDB, a novel framework that leverages textual clinical priors to steer the source-to-target translation process. Our method introduces two key innovations to ensure high-fidelity synthesis. First, we design a text-guided adaptive normalization layer, which dynamically modulates the feature statistics of the diffusion backbone using encoded clinical metadata. This mechanism explicitly aligns the synthesized features with the target modality’s attributes, ensuring semantic consistency throughout the generation process. Second, to prevent semantic degradation during the iterative denoising steps, we propose a semantics reconstruction network. This auxiliary module imposes a constraint that forces the network to preserve deep semantic representations, further reinforcing the semantic consistency between the generated output and the target description. Extensive experiments on multiple medical imaging datasets demonstrate that our approach achieves state-of-the-art performance in terms of quantitative metrics and visual quality, generating images that are both anatomically accurate and semantically faithful to clinical protocols. Yanjun Chi, Jiaen Liang, Jun Yu 0001 |
ICMR | 5 |
| 2026 | PRSE: A two-stage joint optimization approach for lightweight speech enhancement
Haixin Guan, Guanyong Wang, Yanhua Long, Jiaen Liang, Xiaobin Tan |
Speech Commun. | 4 |
| 2025 | Optimization of Multimodal Inputs Based on Diffusion Models: Zero-Shot Semantic Image GenerationabstractWith the continuous advancement of large models, the scale of data has become increasingly important in semantic segmentation tasks. However, the complexity and high cost of annotating semantic segmentation data pose significant challenges to the expansion of datasets. This study aims to leverage pre-trained diffusion generative models for conditional image generation, where labeled masks are used to generate corresponding synthetic images. This ensures a direct correspondence between input and output, effectively bypassing the annotation stage to reduce the cost of labor-intensive tasks. We employ multimodal conditions, to control the generation results. Additionally, we propose a multimodal alignment scheme to optimize the input control conditions, thereby improving the spatial structural accuracy of the generated results. Furthermore, we explore zero-shot generation tasks and successfully achieve zero-shot generation performance across multiple datasets, demonstrating the effectiveness of our approach. In the zero-shot generation experiments on the Cityscapes dataset, our method achieved a 0.6% improvement in the mIoU evaluation metric. On the ADE20K dataset, the performance improvement reached 2.52%, while on the COCO-Stuff dataset, the improvement was 2.43%. Leilei Wang, Renjie Lu 0001, Fengzhao Sun, Jun Yu 0001, Jianqing Sun, Jiaen Liang |
ICME | 8 |
| 2025 | Heterogeneous Encoder Fusion with KAN Decoder for Group Engagement Modeling via 8× Sliding PipelinesabstractEstimating engagement in group interactions is crucial for building socially intelligent systems, such as in human-agent and human-robot interaction. However, precisely modeling the continuous frame-level fluctuations of engagement remains challenging, particularly when considering the complex multi-party signal interactions within groups. Our method employs an encoder that integrates BiLSTM with Transformer to effectively capture both local and global temporal dependencies of multimodal features. Crucially, we explicitly fuse signals from both the target participant and their conversational partners in group to model the holistic group interaction dynamics. Furthermore, we introduce an 8x overlapped optimized sliding window strategy, constructing a ''sliding pipeline'', which significantly enhances the temporal smoothness, continuity, and stability of predictions. In the final regression stage, we replace the traditional multilayer perceptron(MLP) decoder with Kolmogorov-Arnold Network (KAN), leveraging their superior function approximation capability to achieve more accurate engagement predictions. Evaluated on the test sets of NoXi-base, NoXi-addition, MPIIGroupInteraction and NoXi-J datasets from the Multimediate'25 Engagement Challenge, our approach demonstrates significant performance improvements, achieving highly competitive Concordance Correlation Coefficients (CCC) of 0.678 for global, approximately 56.1% higher than the baseline, which shows a significant improvement. Yuefeng Zou, Hui Zhang 0044, Jun Yu 0001, Keda Lu, Lingsi Zhu, Fengzhao Sun, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 10 |
| 2024 | Reducing Speech Distortion and Artifacts for Speech Enhancement by Loss Function
Haixin Guan, Guanyong Wang, Xiaobin Tan, Jiaen Liang |
INTERSPEECH | 6 |
| 2024 | Micro-Expression Spotting Based on Optical Flow Feature with Boundary Calibration
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Zhongpeng Cai, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 9 |
| 2024 | Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global InteractionsabstractThe continual advancements in Generative Artificial Intelligence have created substantial hurdles for accurate deepfake detection, leading to limitations of currently popular detection methods across content-driven video-level deepfake detection scenarios. In this paper, we present the solutions to the Video-Level Deepfake Detection task. Our empirical findings demonstrate that modeling correlations of audio-visual modalities is important for video-level deepfake detection. Therefore, we introduce the model denoted Audio-Visual Local-Global Neural Network (i.e., AV-LGNN) in which the core design is the proposed AV-LGI Module (Audio-Visual Local-Global Interaction Module). The AV-LGI Module is composed of three stages: Local Intra-Region Interaction, Global Inter-Region Interaction, and Local-Global Interaction, which can better capture detailed information at local-level and efficiently learn the fine-grained correlations of inter-modalities in video deepfake detection under lower computational overheads. We further propose an adaptive modality selection strategy to facilitate model learning. Besides, a variety of data augmentation techniques are incorporated for audio-visual branches to enhance the robustness of the AV-LGNN. The experimental results verify the effectiveness of our model. Jia Zhang 0016, Mohan Jing, Keda Lu, Jun Yu 0001, Wen Su 0004, Fang Gao 0001, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 11 |
| 2024 | End-to-end Spatio-Temporal Information Aggregation For Micro-Action DetectionabstractMicro-actions convey the emotions of characters in daily communication and offer richer semantic information compared to conventional actions. Accurate detection of these micro-actions is essential for video understanding. Due to their short duration, low intensity, and high overlap, micro-actions require more detailed video features, presenting a significant challenge for accurate detection. To address these challenges, we propose the 3D-SENet Adapter, which aggregates spatio-temporal information and enables end-to-end online video feature learning. We also find that incorporating background information significantly enhances the detection of small-scale micro-actions. Thus we develop the Cross-Attention Aggregation Detection Head, which integrates multi-scale features within the feature pyramid, thereby improving the detection accuracy of micro-actions occupying small regions in video frames. Our approach achieves first place in the Multi-label Micro-Action Detection (MMAD) and second place in the Micro-Action Recognition (MAR) of Micro-Action Analysis Grand Challenge. Jun Yu 0001, Mohan Jing, Guopeng Zhao, Keda Lu, Feng Zhao 0005, Jiaqing Sun, Jiaen Liang |
ACM Multimedia | 9 |
| 2024 | Temporal-Informative Adapters in VideoMAE V2 and Multi-Scale Feature Fusion for Micro-Expression Spotting-then-Recognize
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 9 |
| 2024 | RAG-Guided Large Language Models for Visual Spatial Description with Adaptive Hallucination CorrectorabstractVisual Spatial Description (VSD) is an emerging image-to-text task which aims at generating descriptions of the spatial relationships between given objects in an image. In this paper, we apply Retrieval-Augmented Generation (RAG) technology in guiding Multimodal Large Language Models (MLLMs) for the task of VSD, complemented by an Adaptive Hallucination Corrector, and further fine-tuning them to bolster semantic understanding and overall model efficacy. We found that our approach demonstrated higher accuracy and fewer hallucination errors in both spatial relationship classification and visual language description tasks within the VSD task, achieving state-of-the-art results. Jun Yu 0001, Gongpeng Zhao, Fengzhao Sun, Fanrui Zhang, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 10 |
| 2023 | M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech SynthesisabstractConversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational TTS systems only focus on extracting global information and omit local prosody features, which contain important fine-grained information like keywords and emphasis. Moreover, it is insufficient to only consider the textual features, and acoustic features also contain various prosody information. Hence, we propose M2-CTTS, an end-to-end multi-scale multi-modal conversational text-to-speech system, aiming to comprehensively utilize historical conversation and enhance prosodic expression. More specifically, we design a textual context module and an acoustic context module with both coarse-grained and fine-grained modeling. Experimental results demonstrate that our model mixed with fine-grained context information and additionally considering acoustic features achieves better prosody performance and naturalness in CMOS tests. Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li 0001, Yingming Gao, Jianhua Tao 0001, Jianqing Sun, Jiaen Liang |
ICASSP | 8 |
| 2023 | Answer-Based Entity Extraction and Alignment for Visual Text Question AnsweringabstractAs a variant of visual question answering (VQA), visual text question answering (VTQA) provides a text-image pair for each question. Text utilizes named entities to describe corresponding image. Consequently, the ability to perform multi-hop reasoning using named entities between text and image becomes critically important. However, existing models pay relatively less attention to this aspect. Therefore, we propose Answer-Based Entity Extraction and Alignment Model (AEEA) to enable a comprehensive understanding and support multi-hop reasoning. The core of AEEA lies in two main components: AKECMR and answer aware predictor. The former emphasizes the alignment of modalities and effectively distinguishes between intra-modal and inter-modal information, and the latter prioritizes the full utilization of intrinsic semantic information contained in answers during training. Our model outperforms the baseline by 2.24% on test-dev set and 1.06% on test set, securing the third place in VTQA2023(English). Jun Yu 0001, Mohan Jing, Weihao Liu 0004, Tongxu Luo, Keda Lu, Fangyu Lei, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 9 |
| 2023 | Sliding Window Seq2seq Modeling for Engagement EstimationabstractEngagement estimation in human conversations has been one of the most important research issues for natural human-robot interaction. However, previous datasets and studies mainly focus on the video-wise level of engagement estimation, therefore, can hardly reflect human's constantly changing engagement. Fortunately, the MultiMediate '23 challenge provides the frame-wise level of engagement estimation task. In this paper, we propose Sliding Window Seq2seq Modeling by BiLSTM and Transformer with powerful sequence modeling capabilities. Our method fully utilizes the global and local multi-modal feature information in the participants' videos and accurately expresses the engagement of the participants at each moment. Our method achieves the state-of-the-art CCC result of 0.71 for engagement estimation on the corresponding test sets. Jun Yu 0001, Keda Lu, Mohan Jing, Ziqi Liang, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 7 |
| 2023 | Dual-model self-regularization and fusion for domain adaptation of robust speaker verification
Yibo Duan, Yanhua Long, Jiaen Liang |
Speech Commun. | 3 |
| 2022 | Selective Pseudo-labeling and Class-wise Discriminative Fusion for Sound Event DetectionabstractIn recent years, exploring effective sound separation (SSep) techniques to improve overlapping sound event detection (SED) attracts more and more attention.Creating accurate separation signals to avoid the catastrophic error accumulation during SED model training is very important and challenging.In this study, we first propose a novel selective pseudo-labeling approach, termed SPL, to produce high confidence separated target events from blind sound separation outputs.These target events are then used to fine-tune the original SED model that pre-trained on the sound mixtures in a multi-objective learning style.Then, to further leverage the SSep outputs, a class-wise discriminative fusion is proposed to improve the final SED performances, by combining multiple frame-level event predictions of both sound mixtures and their separated signals.All experiments are performed on the public DCASE 2021 Task 4 dataset, and results show that our approaches significantly outperforms the official baseline, the collar-based F 1, PSDS1 and PSDS2 performances are improved from 44.3%, 37.3% and 54.9% to 46.5%, 44.5% and 75.4%, respectively. Yunhao Liang, Yanhua Long, Yijie Li 0001, Jiaen Liang |
INTERSPEECH | 4 |
| 2021 | Attention-Based Scaling Adaptation for Target Speech ExtractionabstractThe target speech extraction has attracted widespread attention in recent years. In this work, we focus on investigating the dynamic interaction between different mixtures and the target speaker to exploit the discriminative target speaker clues. We propose a special attention mechanism without introducing any additional parameters in a scaling adaptation layer to better adapt the network towards extracting the target speech. Furthermore, by introducing a mixture embedding matrix pooling method, our proposed attention-based scaling adaptation (ASA) can exploit the target speaker clues in a more efficient way. Experimental results on the spatialized reverberant WSJ0 2-mix dataset demonstrate that the proposed method can improve the performance of the target speech extraction effectively. Furthermore, we find that under the same network configurations, the ASA in a single-channel condition can achieve competitive performance gains as that achieved from two-channel mixtures with inter-microphone phase difference (IPD) features. Jiangyu Han, Yanhua Long, Jiaen Liang |
ASRU | 4 |
| 2020 | Speech Driven Talking Head Generation via Attentional Landmarks Based Representation
Jianqing Sun, Jiaen Liang |
INTERSPEECH | 5 |
| 2020 | Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-Based LVCSRabstractTransformer has shown impressive performance in automatic speech recognition.It uses an encoder-decoder structure with self-attention to learn the relationship between high-level representation of source inputs and embedding of target outputs.In this paper, we propose a novel decoder structure that features a self-and-mixed attention decoder (SMAD) with a deep acoustic structure (DAS) to improve the acoustic representation of Transformer-based LVCSR.Specifically, we introduce a self-attention mechanism to learn a multi-layer deep acoustic structure for multiple levels of acoustic abstraction.We also design a mixed attention mechanism that learns the alignment between different levels of acoustic abstraction and its corresponding linguistic information simultaneously in a shared embedding space.The ASR experiments on Aishell-1 show that the proposed structure achieves CERs of 4.8% on the dev set and 5.1% on the test set, which are the best reported results on this task to the best of our knowledge. Xinyuan Zhou, Grandee Lee, Emre Yilmaz 0001, Yanhua Long, Jiaen Liang, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2018 | Active Learning for LF-MMI Trained Neural Networks in ASR
Yanhua Long, Yijie Li 0001, Jiaen Liang |
INTERSPEECH | 4 |
| 2017 | Speaker Direction-of-Arrival Estimation Based on Frequency-Independent Beampattern
Feng Guo 0002, Yuhang Cao, Zheng Liu 0011, Jiaen Liang, Baoqing Li, Xiaobing Yuan |
INTERSPEECH | 4 |
| 2011 | Exploring nuisance attribute projection and score normalization for GLDS-SVM based automatic mispronunciation detection methodabstractIn the task of mispronunciation detection, the cross-speaker degradation and some other confusing nuisances are the challenging problems demanding prompt solution. In this paper, we will attempt to remove the non-pronunciation variations in the GLDS-SVM expansion space by using nuisance attribute projection strategy, in order to increase the separating capacity between different phoneme instances. Moreover, different kinds of score normalization methods with softmax, posterior probability vector (PPV), Z-norm and T-norm are comparatively discussed. The experiments on three kinds of speech corpora demonstrate the effectiveness of the above methods, and the performance improvement is not very significant, but sustainable. Hongyan Li 0010, Shen Huang, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002 |
ICASSP | 4 |
| 2010 | Automatic reference independent evaluation of prosody quality using multiple knowledge fusions
Shen Huang, Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002 |
INTERSPEECH | 4 |
| 2010 | Exploring goodness of prosody by diverse matching templatesabstractIn automatic speech grading systems, rare research is followed through addressing the issue of GOR (Goodness Of pRosody). In this paper we propose a novel method by taking the advantage of our QBH (Query By Humming) techniques in 2008 MIREX evaluation task. A set of standard samples related to the top-cream students are initially picked up as templates, a cascade QBH structure is then taken from two metrics: the MOMEL stylization followed by DTW distance; the Fujisaki model followed by EMD distance. Sentence GOR is obtained by the fused confidence between target and each template, and forms a weighted sum as the goodness in the passage level. Experiment results indicate that performance increases with the count of template, and Fujisaki-EMD metric outperforms MOMEL-DTW one in terms of correlation. Their combination can be treated as template based GOR score, compensated with our previous feature based GOR score, the approach can achieve 0.432 in correlation and 17.90% in EER in our corpus. Index Terms: speech prosody, query by humming Shen Huang, Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002 |
INTERSPEECH | 4 |
| 2009 | Context Dependent Feature Based Bottom-up Rescoring SVM Classifier in Children's English Stress Mis-pronunciation DetectionabstractAutomatic assessment of word stress error is an integral part for oral language grading system. However, problems that the property of vowels depends on its context information and the data sparseness of different vowel class are yet to be solved. This paper shall briefly introduce a hybrid method consisting of both traditional prosodic features and proposed context dependent strategies. In classification word stress is determined by weighting a bottom-up fashioned group tree with modified distributed probability score. In experiment, the overall equal error rate of our proposed system achieves 9.41%, which exhibits relative reduction and its competence of use in stress error detection system. Shen Huang, Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002 |
ICALT | 4 |
| 2009 | An efficient mispronounciation detction method using GLDS-SVM and formant enhanced featuresabstractMispronunciation detection is an important component in computer assisted language learning (CALL) system. In this work, we introduce an efficient GLDS-SVM based detection method, which is successfully used in language and speaker identification systems, and combine it with traditional methods. The main ideas include: extended MFCC features with normalized formant trajectory information, and then propose a novel multi-model strategy for model training to make full use of samples and solve the problem of data unbalance, finally combine GLDS-SVM method with UBM-GMM system to further improve the performance. Experiments show that GLDS-SVM is highly efficient than traditional RBF-SVM, and the fused system can achieve a significant relative improvement of 17.5% in EER reduction, compared with the baseline UBM-GMM system. Hongyan Li 0010, Jiaen Liang, Shijin Wang 0001, Bo Xu 0002 |
ICASSP | 2 |
| 2009 | High performance automatic mispronunciation detection method based on neural network and TRAP features
Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Shen Huang, Bo Xu 0002 |
INTERSPEECH | 3 |
| 2008 | Improved phonotactic language identification using random forest language modelsabstractRecently a new language model, the random forest language model (RFLM), has been proposed and shown encouraging results in speech recognition tasks. In this paper we applied the RFLM to language identification tasks. We proposed a shared backoff smoothing to deal with data sparseness problem. Experiments were conducted on a subset of NIST 2003 language recognition evaluation data. The RFLM obtained 15.7% relative error rate reduction comparing with the standard trigram LM. The RFLM can be used as a counterpart to n-gram LM and BTLM for system fusion. We also empirically studied the relation between system performance and the tree numbers in a RFLM. Shijin Wang 0001, Jiaen Liang, Bo Xu 0002 |
ICASSP | 3 |
| 2008 | Improving searching speed and accuracy of query by humming system based on three methods: feature fusion, candidates set reduction and multiple similarity measurement rescoring
Lei Wang 0062, Shen Huang, Jiaen Liang, Bo Xu 0002 |
INTERSPEECH | 4 |
| 2007 | A Novel Phone-State Matrix Based Vocabulary-Indenendent Keyword Spotting Method for Spontaneous SpeechabstractKeyword spotting (KWS) is an essential technique for speech information retrieval. When doing offline keyword query on large volume spontaneous speech data, fast and accurate KWS methods are required. In this paper, a novel phone-state matrix based vocabulary-independent KWS method is proposed, which has merits of both hidden Markov model (HMM) based and lattice-based methods. Four KWS systems are compared in our experiments on conversational telephone speech test set. Result shows that compared to the high precision HMM-based KWS system the proposed phone-state matrix system has better equal-error-rate (EER) and false-alarm (FA) performance than the other two lattice-based systems. Jiaen Liang, Peng Ding 0003, Bo Xu 0002 |
ICASSP (4) | 2 |
| 2006 | An Improved Mandarin Keyword Spotting System Using MCE Training and Context-Enhanced VerificationabstractThe task of keyword spotting is to detect a set of keywords in the input continuous speech. The main goal of this work is to develop an improved Mandarin keyword spotting (KWS) system for conversational telephone speech (CTS). In this paper, we propose an efficient online-garbage model based KWS system, which integrated with a word-level minimum classification error (MCE) training method and a novel context-enhanced verification method. Experiment showed that the proposed methods can reduce the equal-error-rate (EER) of the system by 13.8% in relative Jiaen Liang, Peng Ding 0003, Bo Xu 0002 |
ICASSP (1) | 1 |