EDBT 2026 Demo / reviewers in the wild / expert
Ming Liu 0028
dblp:20/2039-28
· DBLP profile ↗
22ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-2160-6111ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Token-Level Text Anomaly DetectionabstractDespite significant progress in text anomaly detection for web applications such as spam filtering and fake news detection, existing methods are fundamentally limited to document-level analysis, unable to identify which specific parts of a text are anomalous. We introduce token-level anomaly detection, a novel paradigm that enables fine-grained localization of anomalies within text. We formally define text anomalies at both document and token-levels, and propose a unified detection framework that operates across multiple levels. To facilitate research in this direction, we collect and annotate three benchmark datasets spanning spam, reviews and grammar errors with token-level labels. Experimental results demonstrate that our framework achieves better performance than other 6 baselines, opening new possibilities for precise anomaly localization in text. All the codes and data are publicly available on https://github.com/charles-cao/TokenCore. Yang Cao 0019, Bicheng Yu, Sikun Yang, Ming Liu 0028, Yujiu Yang 0001 |
WWW | 4 |
| 2026 | TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
Yang Cao 0019, Sikun Yang, Chen Li 0027, Haolong Xiang, Lianyong Qi, Bo Liu 0057, Rongsheng Li, Ming Liu 0028 |
Mach. Learn. | 8 |
| 2026 | FaaSAdapter: An Adaptive Resource Configuration Framework for Serverless Workflows at the EdgeabstractServerless computing has emerged as a promising deployment paradigm for edge scenarios, owing to its efficient resource utilization and flexible provisioning enabled by Function-as-a-Service (FaaS). In Serverless environment, developers are required to configure resources for functions to balance cost efficiency and performance. However, determining appropriate resource allocations for the functions running at the edge is a challenge due to the dynamic nature of the environment. This challenge is further compounded when managing serverless workflows composed of multiple interconnected functions with complex dependencies. To address such an challenge, we present FaaSAdapter, an efficient runtime resource configuration framework for workflow functions, aiming at conserving computational resources at the edge while ensuring timely response to user requests. Different from existing dynamic resource configuration methods that incrementally determine resource schemes for only the immediate subsequent workflow function, FaaSAdapter predicts the execution times of all the unexecuted functions across various resource configurations and determines an optimal configuration schema for the function instances based on the current execution progress. Then, it updates the configuration schema as needed during runtime. Comprehensive experiments demonstrate that FaaSAdapter ensures satisfactory response time of user requests with lowest resource consumption. Haoyu Luo, Ming Liu 0028, Shaojian Qiu, Xiao Liu 0004 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2025 | Beyond Technical Failures: Multimodal Time-Series Modelling for Detecting Social Breakdowns and User Repair Attempts in Human-Robot InteractionabstractReliable detection of conversational errors and user-initiated corrections is critical for effective human-robot interaction (HRI). In this study, we present a comprehensive multimodal approach leveraging temporal window processing, targeted feature engineering, and a MiniRocket + Ridge classification pipeline to address the challenges introduced by the ERR@HRI 2.0 dataset. Our methodology systematically integrates multimodal data streams, including facial expressions, acoustic features, and linguistic embeddings, to predict robot failures and user reactions. Experimental results demonstrate significant improvements over baseline models in event-level detection performance. Notably, linguistic features derived from transcript embeddings emerged as the most informative modality, substantially enhancing model performance. However, we observed challenges associated with managing false positives at the event level, suggesting avenues for future refinement in adaptive thresholding and sequential post-processing techniques. Our findings underscore the importance of careful feature selection and robust temporal modelling in developing effective real-time error detection systems for conversational robots. Our code is available online. https://github.com/Ruddy202/err-hri-2.0-armas.git. Rutherford Agbeshi Patamia, Ha Pham Thien Dinh, Ming Liu 0028, Akansel Cosgun |
ACM Multimedia | 3 |
| 2025 | Large Language Model and Variational Autoencoder Based Deep Neural Framework for Cyber Attack Detection
Jyotheesh Gaddam, Ishara Bandara, Ming Liu 0028, Sutharshan Rajasegarar, Muneeb Ul Hassan 0001, Lu-Xing Yang, Gang Li 0009, Maia Angelova |
PAKDD (4) | 5 |
| 2025 | Contrastive Learning and Feature Space Tactics: A Dual Approach to Strengthen Backdoor AttacksabstractContrastive Learning and Feature Space Tactics: A Dual Approach to Strengthen Backdoor Attacks Hao Fu 0022, Ming Liu 0028, Rongsheng Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Overview of the NLPCC2024 Shared Task 6: Scientific Literature Survey Generation
Yangjie Tian, Xungang Gu, Aijia Li, He Zhang 0034, Ruohua Xu, Ming Liu 0028 |
NLPCC (5) | 7 |
| 2024 | Prototype-Guided Memory Replay for Continual LearningabstractContinual learning (CL) is a machine learning paradigm that accumulates knowledge while learning sequentially. The main challenge in CL is catastrophic forgetting of previously seen tasks, which occurs due to shifts in the probability distribution. To retain knowledge, existing CL models often save some past examples and revisit them while learning new tasks. As a result, the size of saved samples dramatically increases as more samples are seen. To address this issue, we introduce an efficient CL method by storing only a few samples to achieve good performance. Specifically, we propose a dynamic prototype-guided memory replay (PMR) module, where synthetic prototypes serve as knowledge representations and guide the sample selection for memory replay. This module is integrated into an online meta-learning (OML) model for efficient knowledge transfer. We conduct extensive experiments on the CL benchmark text classification datasets and examine the effect of training set order on the performance of CL models. The experimental results demonstrate the superiority our approach in terms of accuracy and efficiency. Stella Ho, Ming Liu 0028, Lan Du 0002, Longxiang Gao, Yong Xiang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | A Survey on Out-of-Distribution Evaluation of Neural NLP ModelsabstractAdversarial robustness, domain generalization and dataset biases are three active lines of research contributing to out-of-distribution (OOD) evaluation on neural NLP models. However, a comprehensive, integrated discussion of the three research lines is still lacking in the literature. This survey will 1) compare the three lines of research under a unifying definition; 2) summarize their data-generating processes and evaluation protocols for each line of research; and 3) emphasize the challenges and opportunities for future work. Ming Liu 0028, Shang Gao 0003, Wray L. Buntine |
IJCAI | 2 |
| 2022 | Semi-supervised Continual Learning with Meta Self-trainingabstractContinual learning (CL) aims to enhance sequential learning by alleviating the forgetting of previously acquired knowledge. Recent advances in CL lack consideration of the real-world scenarios, where labeled data are scarce and unlabeled data are abundant. To narrow this gap, we focus on semi-supervised continual learning (SSCL). We exploit unlabeled data under limited supervision in the CL setting and demonstrate the feasibility of semi-supervised learning in CL. In this work, we propose a novel method, namely Meta-SSCL, which combines meta-learning with pseudo-labeling and data augmentations to learn a sequence of semi-supervised tasks without catastrophic forgetting. Extensive experiments on CL benchmark text classification datasets show that our method achieves promising results in SSCL. Stella Ho, Ming Liu 0028, Lan Du 0002, Longxiang Gao, Shang Gao 0003 |
CIKM | 2 |
| 2022 | How Far are We from Robust Long Abstractive Summarization?abstractAbstractive summarization has made tremendous progress in recent years.In this work, we perform fine-grained human annotations to evaluate long document abstractive summarization systems (i.e., models and metrics) with the aim of implementing them to generate reliable summaries.For long document abstractive models, we show that the constant strive for state-of-the-art ROUGE results can lead us to generate more relevant summaries but not factual ones.For long document evaluation metrics, human evaluation results show that ROUGE remains the best at evaluating the relevancy of a summary.It also reveals important limitations of factuality metrics in detecting different types of factual errors and the reasons behind the effectiveness of BARTScore.We then suggest promising directions in the endeavor of developing factual consistency metrics.Finally, we release our annotated long document dataset with the hope that it can contribute to the development of metrics across a broader range of summarization settings. Huan Yee Koh, Jiaxin Ju, Ming Liu 0028, Shirui Pan |
EMNLP | 4 |
| 2022 | Overview of NLPCC2022 Shared Task 5 Track 2: Named Entity Recognition
Borui Cai, He Zhang 0034, Fenghong Liu 0002, Ming Liu 0028, Tianrui Zong |
NLPCC (2) | 4 |
| 2022 | Overview of NLPCC2022 Shared Task 5 Track 1: Multi-label Classification for Scientific Literature
Ming Liu 0028, He Zhang 0034, Yangjie Tian, Tianrui Zong, Borui Cai, Ruohua Xu |
NLPCC (2) | 1 |
| 2022 | Graph Intelligence Enhanced Bi-Channel Insider Threat Detection
Jiao Yin 0003, Mingshan You, Hua Wang 0002, Jinli Cao, Jianxin Li 0001, Ming Liu 0028 |
NSS | 7 |
| 2022 | Mulan: A Multiple Residual Article-Wise Attention Network for Legal Judgment PredictionabstractLegal judgment prediction (LJP) is used to predict judgment results based on the description of individual legal cases. In order to be more suitable for actual application scenarios in which the case has cited multiple articles and has multiple charges, we formulate legal judgment prediction as a multiple label learning problem and present a deep learning model that can effectively encode the content of each legal case via a multi-residual convolution neural network and the semantics of law articles via an article encoder. An article-wise attention mechanism is proposed to fuse the two types of encoded information. Experimental results derived on the CAIL2018 datasets show that our model provides a significant performance improvement over the existing neural models in predicting relevant law articles and charges. Lan Du 0002, Ming Liu 0028, Xiabing Zhou |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2021 | Neural Attention-Aware Hierarchical Topic ModelabstractNeural topic models (NTMs) apply deep neural networks to topic modelling.Despite their success, NTMs generally ignore two important aspects: (1) only document-level word count information is utilized for the training, while more fine-grained sentence-level information is ignored, and (2) external semantic knowledge regarding documents, sentences and words are not exploited for the training.To address these issues, we propose a variational autoencoder (VAE) NTM model that jointly reconstructs the sentence and document word counts using combinations of bag-of-words (BoW) topical embeddings and pre-trained semantic embeddings.The pre-trained embeddings are first transformed into a common latent topical space to align their semantics with the BoW embeddings.Our model also features hierarchical KL divergence to leverage embeddings of each document to regularize those of their sentences, thereby paying more attention to semantically relevant sentences.Both quantitative and qualitative experiments have shown the efficacy of our model in 1) lowering the reconstruction errors at both the sentence and document levels, and 2) discovering more coherent topics from real-world datasets. He Zhao 0001, Ming Liu 0028, Lan Du 0002, Wray L. Buntine |
EMNLP (1) | 3 |
| 2021 | Variational auto-encoder based Bayesian Poisson tensor factorization for sparse and imbalanced count data
Ming Liu 0028, Ruohua Xu, Lan Du 0002, Longxiang Gao, Yong Xiang 0001 |
Data Min. Knowl. Discov. | 2 |
| 2020 | Multi-label Few/Zero-shot Learning with Knowledge Aggregated from Multiple Label GraphsabstractFew/Zero-shot learning is a big challenge of many classifications tasks, where a classifier is required to recognise instances of classes that have very few or even no training samples. It becomes more difficult in multilabel classification, where each instance is la-belled with more than one class. In this paper, we present a simple multi-graph aggregation model that fuses knowledge from multiple label graphs encoding different semantic label relationships in order to study how the aggregated knowledge can benefit multi-label zero/few-shot document classification. The model utilises three kinds of semantic information, i.e., the pre-trained word embeddings, label description, and pre-defined label relations. Experimental results derived on two large clinical datasets (i.e., MIMIC-II and MIMIC-III) and the EU legislation dataset show that methods equipped with the multi-graph knowledge aggregation achieve significant performance improvement across almost all the measures on few/zero-shot labels. Jueqing Lu, Lan Du 0002, Ming Liu 0028, Joanna Dipnall |
EMNLP (1) | 3 |
| 2020 | SummPip: Unsupervised Multi-Document Summarization with Sentence Graph CompressionabstractObtaining training data for multi-document Summarization (MDS) is time consuming and resource-intensive, so recent neural models can only be trained for limited domains. In this paper, we propose SummPip: an unsupervised method for multi-document summarization, in which we convert the original documents to a sentence graph, taking both linguistic and deep representation into account, then apply spectral clustering to obtain multiple clusters of sentences, and finally compress each cluster to generate the final summary. Experiments on Multi-News and DUC-2004 datasets show that our method is competitive to previous unsupervised methods and is even comparable to the neural supervised approaches. In addition, human evaluation shows our system produces consistent and complete summaries compared to human written ones. Jinming Zhao, Ming Liu 0028, Longxiang Gao, Lan Du 0002, He Zhao 0001, He Zhang 0034, Gholamreza Haffari |
SIGIR | 2 |
| 2019 | Learning How to Active Learn by DreamingabstractHeuristic-based active learning (AL) methods are limited when the data distribution of the underlying learning problems vary.Recent data-driven AL policy learning methods are also restricted to learn from closely related domains.We introduce a new sample-efficient method that learns the AL policy directly on the target domain of interest by using wake and dream cycles.Our approach interleaves between querying the annotation of the selected datapoints to update the underlying student learner and improving AL policy using simulation where the current student learner acts as an imperfect annotator.We evaluate our method on cross-domain and cross-lingual text classification and named entity recognition tasks.Experimental results show that our dream-based AL policy training strategy is more effective than applying the pretrained policy without further fine-tuning, and better than the existing strong baseline methods that use heuristics or reinforcement learning. Thuy-Trang Vu, Ming Liu 0028, Dinh Q. Phung, Gholamreza Haffari |
ACL (1) | 2 |
| 2018 | Learning How to Actively Learn: A Deep Imitation Learning ApproachabstractHeuristic-based active learning (AL) methods are limited when the data distribution of the underlying learning problems vary.We introduce a method that learns an AL policy using imitation learning (IL).Our IL-based approach makes use of an efficient and effective algorithmic expert, which provides the policy learner with good actions in the encountered AL situations.The AL strategy is then learned with a feedforward network, mapping situations to most informative query datapoints.We evaluate our method on two different tasks: text classification and named entity recognition.Experimental results show that our IL-based AL strategy is more effective than strong previous methods using heuristics and reinforcement learning. Ming Liu 0028, Wray L. Buntine, Gholamreza Haffari |
ACL (1) | 1 |
| 2018 | Learning to Actively Learn Neural Machine TranslationabstractTraditional active learning (AL) methods for machine translation (MT) rely on heuristics.However, these heuristics are limited when the characteristics of the MT problem change due to e.g. the language pair or the amount of the initial bitext.In this paper, we present a framework to learn sentence selection strategies for neural MT.We train the AL query strategy using a high-resource language-pair based on AL simulations, and then transfer it to the lowresource language-pair of interest.The learned query strategy capitalizes on the shared characteristics between the language pairs to make an effective use of the AL budget.Our experiments on three language-pairs confirms that our method is more effective than strong heuristic-based methods in various conditions, including cold-start and warm-start as well as small and extremely small data conditions. Ming Liu 0028, Wray L. Buntine, Gholamreza Haffari |
CoNLL | 1 |