EDBT 2026 Demo / reviewers in the wild / expert
Cunli Mao
dblp:35/2229
· DBLP profile ↗
27ranked-venue papers
2as first author
23since 2021 · last 2026
0000-0002-8289-6036ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving cross-lingual dependency parsing via LLM progressive alignment
Jianghui He, Jianjian Liu, Ying Li 0127, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao, Cunli Mao |
Pattern Recognit. | 7 |
| 2025 | BiDeV: Bilateral Defusing Verification for Complex Claim Fact-CheckingabstractComplex claim fact-checking performs a crucial role in disinformation detection. However, existing fact-checking methods struggle with claim vagueness, specifically in effectively handling latent information and complex relations within claims. Moreover, evidence redundancy, where non-essential information complicates the verification process, remains a significant issue. To tackle these limitations, we propose Bilateral Defusing Verification (BiDeV), a novel fact-checking working-flow framework integrating multiple role-played LLMs to mimic the human-expert fact-checking process. BiDeV consists of two main modules: Vagueness Defusing identifies latent information and resolves complex relations to simplify the claim, and Redundancy Defusing eliminates redundant content to enhance the evidence quality. Extensive experimental results on two widely used challenging fact-checking benchmarks (Hover and Feverous-s) demonstrate that our BiDeV can achieve the best performance under both gold and open settings. This highlights the effectiveness of BiDeV in handling complex claims and ensuring precise fact-checking. Yuxuan Liu 0009, Hongda Sun 0001, Wenya Guo, Xinyan Xiao, Cunli Mao, Zhengtao Yu 0001, Rui Yan 0001 |
AAAI | 5 |
| 2025 | SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language ModelsabstractWith the rapid advancement of large language models (LLMs), discrete speech representations have become crucial for integrating speech into LLMs. Existing methods for speech representation discretization rely on a predefined codebook size and Euclidean distance-based quantization. However, 1) the size of codebook is a critical parameter that affects both codec performance and downstream task training efficiency. 2) The Euclidean distance-based quantization may lead to audio distortion when the size of the codebook is controlled within a reasonable range. In fact, in the field of information compression, structural information and entropy guidance are crucial, but previous methods have largely overlooked these factors. Therefore, we address the above issues from an information-theoretic perspective, we present SECodec, a novel speech representation codec based on structural entropy (SE) for building speech language models. Specifically, we first model speech as a graph, clustering the speech features nodes within the graph and extracting the corresponding codebook by hierarchically and disentangledly minimizing 2D SE. Then, to address the issue of audio distortion, we propose a new quantization method. This method still adheres to the 2D SE minimization principle, adaptively selecting the most suitable token corresponding to the cluster for each incoming original speech node. Furthermore, we develop a Structural Entropy-based Speech Language Model (SESLM) that leverages SECodec. Experimental results demonstrate that SECodec performs comparably to EnCodec in speech reconstruction, and SESLM surpasses VALL-E in zero-shot text-to-speech tasks. Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004, Ling Dong |
AAAI | 5 |
| 2025 | Graph Contrastive Learning with Decoupled AugmentationabstractGraph contrastive learning based on augmentation strategies has recently demonstrated remarkable performance. Existing methods typically jointly leverage attribute and structural augmentations to generate graph views, learning data invariance information through contrasting sample pairs. However, this joint approach may deviate from the expectation of semantically similar before and after augmentation. The propagation of attribute information in graphs usually occurs through their structure, meaning that structural and attribute augmentations can interfere with each other and potentially distort the graph’s semantics. To address this, we propose a decoupled augmentation framework for graph contrastive learning, which eliminates the mutual interference between the two levels of augmentation while fully exploring graph information. Specifically, our framework employs separate encoders to learn data invariance under different augmentation levels, and it considers the positive gains generated between these levels. Experimental results on five public datasets show that the proposed method is more competitive than state-of-the-art approaches. Shihao Gao, Caoshuo Li, Cunli Mao, Xulong Zhang 0001, Xiaoyang Qu, Taisong Jin, Jianzong Wang |
ICASSP | 3 |
| 2025 | Voice Conversion via Structural EntropyabstractVoice conversion (VC) aims to transform a person’s voice to resemble that of another person while maintaining the original linguistic content. Existing methods suffer from the blurring of speech representations and the leakage of prosody information. To address this issue, this study introduces SEVC, a novel neural structural entropy-based VC framework. First, we extract self-supervised representations from both the source and reference speech. The representations of the reference speech are structured as a graph. Using two-dimensional (2D) structural entropy (SE), semantically similar representations are clustered together. For the speaker conversion process, each frame of the source speech is treated as a new node, and the most appropriate semantic cluster for each node is identified using SE. Each frame of the source representation is then replaced by the center representation of its corresponding semantic cluster from the reference speech. Finally, a pretrained vocoder synthesizes audio from the transformed representations. Quantitative and qualitative evaluations demonstrate that SEVC improves speaker similarity while maintaining intelligibility scores comparable to existing methods. Additionally, experimental results indicate that SEVC effectively disentangles paralinguistic features from the reference speech, allowing the model to better focus on the semantic content. Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Ling Dong, Yuxin Huang 0004 |
ICASSP | 4 |
| 2025 | SAGEC: Syntax-Aware Grammatical Error Correction with Retrieval-Augmented Generation
Shichang Zhu, Ying Li 0127, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004 |
NLPCC (2) | 6 |
| 2025 | Fine-Grained Contrastive Learning for End-to-End Vietnamese Text Image Machine Translation
Cunli Mao, Ying Li 0127, Shengxiang Gao, Zhengtao Yu 0001 |
NLPCC (3) | 1 |
| 2025 | Fine-Grained Prosody-Controllable Lao Speech Synthesis Guided by Natural Language
Zhenbei Guo, Cunli Mao, Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao |
NLPCC (2) | 3 |
| 2025 | A Southeast Asian Language OCR Dataset and Evaluation for Large Multimodal Models
Cunli Mao, Ying Li 0127, Shengxiang Gao, Zhengtao Yu 0001 |
NLPCC (2) | 3 |
| 2025 | Time-frequency perception guided multi-level contrastive learning for rotating machinery fault diagnosis
Zhenqiu Shu, Dazheng Peng, Hongbin Wang 0002, Cunli Mao, Zhengtao Yu 0001 |
Expert Syst. Appl. | 4 |
| 2025 | MKE-PLLM: A benchmark for multilingual knowledge editing on pretrained large language model
Ran Song 0002, Shengxiang Gao, Xiaofei Gao, Cunli Mao, Zhengtao Yu 0001 |
Neurocomputing | 4 |
| 2024 | "In-Dialogues We Learn": Towards Personalized Dialogue Without Pre-defined Profiles through In-Dialogue LearningabstractPersonalized dialogue systems have gained significant attention in recent years for their ability to generate responses in alignment with different personas.However, most existing approaches rely on pre-defined personal profiles, which are not only time-consuming and labor-intensive to create but also lack flexibility.We propose In-Dialogue Learning (IDL), a fine-tuning framework that enhances the ability of pre-trained large language models to leverage dialogue history to characterize persona for personalized dialogue generation tasks without pre-defined profiles.Our experiments on three datasets demonstrate that IDL brings substantial improvements, with BLEU and ROUGE scores increasing by up to 200% and 247%, respectively.Additionally, the results of human evaluations further validate the efficacy of our proposed method. Chuanqi Cheng, Quan Tu, Wei Wu 0014, Shuo Shang, Cunli Mao, Zhengtao Yu 0001, Rui Yan 0001 |
EMNLP | 5 |
| 2024 | DETS: End-to-End Single-Stage Text-to-Speech Via Hierarchical Diffusion Gan ModelsabstractEnd-to-end single-stage text-to-speech models have garnered significant attention in recent research, surpassing the performance of conventional two-stage pipeline systems. While prior single-stage models have made substantial advancements, there remains room for improvement in addressing intermittent issues related to unnaturalness and prosody diversity. Unlike previous works, we propose a novel single-stage TTS framework to tackle these problems via hierarchical denoising diffusion generative adversarial networks (GAN) modeling, which parameterizes the denoising model by directly predicting latent variables to improve the naturalness and diversity of the generated speech. Specifically, a conditional GAN is adopted as a non-Gaussian multimodal function to model the denoising distribution, which construct the duration predictor and speech decoder respectively. As such, it allows the TTS model learn the more natural one-to-many relationship in which a text input can be spoken in multiple ways with different pitches and rhythms. In addition, We show that DETS can generate high-fidelity speech waveform with only 1 denoising step. Extensive experimental results on the LJSpeech benchmark dataset demonstrate the favourable performance of the proposed method. Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004 |
ICASSP | 4 |
| 2024 | StreamingDialogue: Prolonged Dialogue Learning via Long Context Compression with Minimal LossesabstractStandard Large Language Models (LLMs) struggle with handling dialogues with long contexts due to efficiency and consistency issues. According to our observation, dialogue contexts are highly structured, and the special token of End-of-Utterance (EoU) in dialogues has the potential to aggregate information. We refer to the EoU tokens as ``conversational attention sinks'' (conv-attn sinks). Accordingly, we introduce StreamingDialogue, which compresses long dialogue history into conv-attn sinks with minimal losses, and thus reduces computational complexity quadratically with the number of sinks (i.e., the number of utterances). Current LLMs already demonstrate the ability to handle long context window, e.g., a window size of 200K or more. To this end, by compressing utterances into EoUs, our method has the potential to handle more than 200K of utterances, resulting in a prolonged dialogue learning. In order to minimize information losses from reconstruction after compression, we design two learning strategies of short-memory reconstruction (SMR) and long-memory reactivation (LMR). Our method outperforms strong baselines in dialogue tasks and achieves a 4 $\times$ speedup while reducing memory usage by 18 $\times$ compared to dense attention recomputation. Quan Tu, Cunli Mao, Zhengtao Yu 0001, Ji-Rong Wen, Rui Yan 0001 |
NeurIPS | 3 |
| 2024 | Enhancing Cross-Lingual Topic-Essay Generation with Knowledge and Topic Consistency Constraints
Huailing Gu, Yuxin Huang 0004, Zhengtao Yu 0001, Cunli Mao |
NLPCC (4) | 4 |
| 2024 | Part-of-Speech and Confusion-Set Constrained Language Model for Vietnamese Spelling Correction Corpus Construction
Ying Li 0127, Ling Dong, Zhengtao Yu 0001, Cunli Mao |
NLPCC (4) | 6 |
| 2024 | View-interactive attention information alignment-guided fusion for incomplete multi-view clustering
Zhenqiu Shu, Yunwei Luo, Yuxin Huang 0004, Cunli Mao, Zhengtao Yu 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Structure-guided feature and cluster contrastive learning for multi-view clustering
Zhenqiu Shu, Bin Li 0006, Cunli Mao, Shengxiang Gao, Zhengtao Yu 0001 |
Neurocomputing | 3 |
| 2024 | Zero-Shot Text Normalization via Cross-Lingual Knowledge DistillationabstractText normalization (TN) is a crucial preprocessing step in text-to-speech synthesis, which pertains to the accurate pronunciation of numbers and symbols within the text. Existing neural network-based TN methods have shown significant success in rich-resource languages. However, these methods are data-driven and highly rely on a large number of labeled datasets, which are not practical in zero-resource settings. Rule-based weighted finite-state transducers (WFST) are a common measure for zero-shot TN, but WFST-based TN approaches encounter challenges with ambiguous input, particularly in cases where the normalized form is context-dependent. On the other hand, conventional neural TN methods suffer from unrecoverable errors. In this paper, we propose ZSTN, a novel zero-shot TN framework based on cross-lingual knowledge distillation, which utilizes annotated data to train the teacher model on rich-resource language and unlabelled data to train the student model on zero-resource language. Furthermore, it incorporates expert knowledge from WFST into a knowledge distillation neural network. Concretely, a TN model with WFST pseudo-labels augmentation is trained as a teacher model in the source language. Subsequently, the student model is supervised by soft-labels from the teacher model and WFST pseudo-labels from the target language. By leveraging cross-lingual knowledge distillation, we address contextual ambiguity in the text, while WFST mitigates unrecoverable errors of the neural model. Additionally, ZSTN is adaptable to different zero-resource languages by using the joint loss function for the teacher model and WFST constraints. We also release a zero-shot text normalization dataset in five languages. We compare ZSTN with seven zero-shot TN benchmarks on public datasets in four languages for the teacher model and zero-shot datasets in five languages for the student model. The results demonstrate that the proposed ZSTN excels in performance without the need for labeled data. Linqin Wang, Zhengtao Yu 0001, Hao Peng 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004, Ling Dong, Philip S. Yu |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Multi-view clustering via label-embedded regularized NMF with dual-graph constraints
Bin Li 0006, Zhenqiu Shu, Cunli Mao, Shengxiang Gao, Zhengtao Yu 0001 |
Neurocomputing | 4 |
| 2022 | Discrete asymmetric zero-shot hashing with application to cross-modal retrieval
Zhenqiu Shu, Kailing Yong, Jun Yu 0011, Shengxiang Gao, Cunli Mao, Zhengtao Yu 0001 |
Neurocomputing | 5 |
| 2022 | Scale-Insensitive Object Detection via Attention Feature Pyramid Transformer Network
Lingling Li 0004, Changwen Zheng, Cunli Mao, Haibo Deng, Taisong Jin |
Neural Process. Lett. | 3 |
| 2021 | A Neural Joint Model with BERT for Burmese Syllable Segmentation, Word Segmentation, and POS TaggingabstractThe smallest semantic unit of the Burmese language is called the syllable. In the present study, it is intended to propose the first neural joint learning model for Burmese syllable segmentation, word segmentation, and part-of-speech ( POS ) tagging with the BERT. The proposed model alleviates the error propagation problem of the syllable segmentation. More specifically, it extends the neural joint model for Vietnamese word segmentation, POS tagging, and dependency parsing [28] with the pre-training method of the Burmese character, syllable, and word embedding with BiLSTM-CRF-based neural layers. In order to evaluate the performance of the proposed model, experiments are carried out on Burmese benchmark datasets, and we fine-tune the model of multilingual BERT. Obtained results show that the proposed joint model can result in an excellent performance. Cunli Mao, Zhibo Man, Zhengtao Yu 0001, Shengxiang Gao, Zhenhan Wang, Hongbin Wang 0002 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2019 | Word Segmentation for Burmese Based on Dual-Layer CRFsabstractBurmese is an isolated language, in which the syllable is the smallest unit. Syllable segmentation methods based on matching lead to performance subject to the syllable segmentation effect. This article proposes a word segmentation method with fusion conditions of double syllable features. It combines word segmentation and segmentation of syllables into one process, thus reducing the impact of errors on the syllable segmentation of Burmese. In the first layer of the conditional random fields (CRF) model, Burmese characters as atomic features are integrated into the Burma section of the Barkis Speech Paradigm (Backus normal form) features to realize the Burma syllable sequence tags. In the second layer of the CRFs model, with the syllable marked as input, it realizes the sequence markers through building a feature template with syllables as atomic features. The experimental results show that the proposed method has a better effect compared with the method based on the matching of syllables. Shaoning Zhang, Cunli Mao, Zhengtao Yu 0001, Hongbin Wang 0002, Jiafu Zhang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | Fractional differential and variational method for image fusion and super-resolution
Huafeng Li 0001, Zhengtao Yu 0001, Cunli Mao |
Neurocomputing | 3 |
| 2016 | Multifocus image fusion by combining with mixed-order structure tensors and multiscale neighborhood
Huafeng Li 0001, Xiaosong Li 0004, Zhengtao Yu 0001, Cunli Mao |
Inf. Sci. | 4 |
| 2010 | Question classification based on co-training style semi-supervised learning
Zhengtao Yu 0001, Lei Su 0003, Cunli Mao, Jianyi Guo |
Pattern Recognit. Lett. | 5 |