EDBT 2026 Demo / reviewers in the wild / expert
Minghan Wang
dblp:228/4495
· DBLP profile ↗
17ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VIM-Net: A voxel-interaction multimodal network for 3D object detection
Minghan Wang, Xijiong Wang, Yonghuai Liu, Ardhendu Behera, Baowen Zhang |
Pattern Recognit. | 1 |
| 2025 | OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMsabstractThe increased use of large language models (LLMs) across a variety of real-world applications calls for mechanisms to verify the fac- tual accuracy of their outputs. Difficulties lie in assessing the factuality of free-form responses in open domains. Also, different pa- pers use disparate evaluation benchmarks and measurements, which renders them hard to compare and hampers future progress. To mitigate these issues, we propose OpenFactCheck, a unified framework for building customized automatic fact-checking systems, benchmarking their accuracy, evaluating factuality of LLMs, and verifying claims in a document. OpenFactCheck consists of three modules: (i) CUSTCHECKER allows users to easily customize an automatic fact-checker and verify the factual correctness of documents and claims, (ii) LLMEVAL, a unified evaluation framework assesses LLM’s factuality ability from various perspectives fairly, and (iii) CHECKEREVAL is an extensible solution for gauging the reliability of automatic fact-checkers’ verification results using human-annotated datasets. Data and code are publicly available at https://github.com/yuxiaw/openfactcheck. Yuxia Wang 0003, Minghan Wang, Georgi Georgiev 0001, Jiahui Geng, Iryna Gurevych, Preslav Nakov |
COLING | 2 |
| 2025 | SpeechDialogueFactory: A Framework for Natural Speech Dialogue Generation
Minghan Wang, Ye Bai 0002, Thuy-Trang Vu, Ehsan Shareghi, Gholamreza Haffari |
INTERSPEECH | 1 |
| 2025 | Learning Across the Gap: Hybrid Multi-armed Bandits with Heterogeneous Offline and Online DataabstractThe multi-armed bandit (MAB) is a fundamental online decision-making framework that has been extensively studied over the past two decades. To mitigate the high cost and slow convergence of purely online learning, modern MAB approaches have explored _hybrid_ paradigms that leverage offline data to warm-start online learning. However, existing approaches face a significant limitation by assuming that the offline and online data are homogeneous—they share the same feedback structure and are drawn from the same underlying distribution. This assumption is often violated in practice, where offline data often originate from diverse sources and evolving environments, resulting in feedback heterogeneity and distributional shifts. In this work, we tackle the challenge of learning across this offline-online gap by developing a general hybrid bandit framework that incorporates heterogeneous offline data to improve online performance. We study two hybrid settings: (1) using reward-based offline data to accelerate online learning in preference-based bandits (i.e., dueling bandits), and (2) using preference-based offline data to improve online standard MAB algorithms. For both settings, we design novel algorithms and derive tight regret bounds that match or improve upon existing benchmarks despite heterogeneity. Empirical evaluations on both synthetic and real-world datasets show that our proposed methods significantly outperform baseline algorithms. Qijia He, Minghan Wang, Xutong Liu 0002, Fang Kong 0002 |
NeurIPS | 2 |
| 2025 | Against The Achilles' Heel: A Survey on Red Teaming for Generative ModelsabstractGenerative models are rapidly gaining popularity and being integrated into everyday applications, raising concerns over their safe use as various vulnerabilities are exposed. In light of this, the field of red teaming is undergoing fast-paced growth, highlighting the need for a comprehensive survey covering the entire pipeline and addressing emerging topics. Our extensive survey, which examines over 120 papers, introduces a taxonomy of fine-grained attack strategies grounded in the inherent capabilities of language models. Additionally, we have developed the “searcher” framework to unify various automatic red teaming approaches. Moreover, our survey covers novel areas including multimodal attacks and defenses, risks around LLM-based agents, overkill of harmless queries, and the balance between harmlessness and helpfulness. Warning: This paper contains examples that may be offensive, harmful, or biased. Lizhi Lin, Honglin Mu, Zenan Zhai, Minghan Wang, Yuxia Wang 0003, Renxi Wang, Wanxiang Che, Timothy Baldwin, Haonan Li 0002 |
J. Artif. Intell. Res. | 4 |
| 2024 | Factuality of Large Language Models: A SurveyabstractYuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu, Georgi Nenkov Georgiev, Rocktim Jyoti Das, Preslav Nakov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yuxia Wang 0003, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu 0023, Georgi Georgiev 0001, Rocktim Jyoti Das, Preslav Nakov |
EMNLP | 2 |
| 2024 | RASU: Retrieval Augmented Speech Understanding through Generative Modeling
Hao Yang 0006, Min Zhang 0042, Minghan Wang |
INTERSPEECH | 3 |
| 2023 | Text Style Transfer Back-TranslationabstractDaimeng Wei, Zhanglin Wu, Hengchao Shang, Zongyao Li, Minghan Wang, Jiaxin Guo, Xiaoyu Chen, Zhengzhe Yu, Hao Yang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Daimeng Wei, Zhanglin Wu, Hengchao Shang, Minghan Wang, Xiaoyu Chen 0004, Zhengzhe Yu, Hao Yang 0006 |
ACL (1) | 5 |
| 2023 | UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error CorrectionabstractError correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER). Previous works usually adopt end-to-end models and has strong dependency on Pseudo Paired Data and Original Paired Data. But when only pre-training on Pseudo Paired Data, previous models have negative effect on correction. While fine-tuning on Original Paired Data, the source side data must be transcribed by a well-trained ASR model, which takes a lot of time and not universal. In this paper, we propose UCorrect, an unsupervised Detector-Generator-Selector framework for ASR Error Correction. UCorrect has no dependency on the training data mentioned before. The whole procedure is first to detect whether the character is erroneous, then to generate some candidate characters and finally to select the most confident one to replace the error character. Experiments on the public AISHELL-1 dataset and WenetSpeech dataset show the effectiveness of UCorrect for ASR error correction: 1) it achieves significant WER reduction, achieves 6.83% even without fine-tuning and 14.29% after fine-tuning; 2) it outperforms the popular NAR correction models by a large margin with a competitive low latency; and 3) it is an universal method, as it reduces all WERs of the ASR model with different decoding strategies and reduces all WERs of ASR models trained on different scale datasets. Minghan Wang, Xiaosong Qiao, Daimeng Wei, Hengchao Shang, Zhengzhe Yu, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
ICASSP | 2 |
| 2023 | Zephyr: Zero-Shot Punctuation RestorationabstractPunctuation restoration can be crucial for the cascade speech translation system. Traditional approaches typically treat it as a sequential tagging problem, predicting which punctuation mark should follow a given word. However, this often requires significant computational and storage resources for full-stage training or fine-tuning. Our argument is that pre-trained language models (PLMs) can directly leverage their learned knowledge for punctuation generation, making additional training unnecessary. In this paper, we propose the Zephyr algorithm, which utilizes PLMs to perform zero-shot and few-shot punctuation restoration for both offline and streaming scenarios. Our experimental results demonstrate that, in comparison to fine-tuning-based baselines, Zephyr achieves competitive performance while requiring little to no training cost and exhibiting better generalizability in zeroshot and few-shot settings. Minghan Wang, Yinglu Li, Xiaosong Qiao, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
ICASSP | 1 |
| 2023 | WhiSLU: End-to-End Spoken Language Understanding with Whisper
Minghan Wang, Yinglu Li, Xiaosong Qiao, Hengchao Shang, Daimeng Wei, Shimin Tao, Min Zhang 0042, Hao Yang 0006 |
INTERSPEECH | 1 |
| 2022 | Diformer: Directional Transformer for Neural Machine TranslationabstractAutoregressive (AR) and Non-autoregressive (NAR) models have their own superiority on the performance and latency, combining them into one model may take advantage of both. Current combination frameworks focus more on the integration of multiple decoding paradigms with a unified generative model, e.g. Masked Language Model. However, the generalization can be harmful on the performance due to the gap between training objective and inference. In this paper, we aim to close the gap by preserving the original objective of AR and NAR under a unified framework. Specifically, we propose the Directional Transformer (Diformer) by jointly modelling AR and NAR into three generation directions (left-to-right, right-to-left and straight) with a newly introduced direction variable, which works by controlling the prediction of each token to have specific dependencies under that direction. The unification achieved by direction successfully preserves the original dependency assumption used in AR and NAR, retaining both generalization and performance. Experiments on 4 WMT benchmarks demonstrate that Diformer outperforms current united-modelling works with more than 1.5 BLEU points for both AR and NAR decoding, and is also competitive to the state-of-the-art independent AR and NAR models. Minghan Wang, Yuxia Wang 0003, Daimeng Wei, Hengchao Shang, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
EAMT | 1 |
| 2022 | CCDC: A Chinese-Centric Cross Domain Contrastive Learning Framework
Hao Yang 0006, Shimin Tao, Minghan Wang, Min Zhang 0042, Daimeng Wei, Shuai Zhao 0001, Miaomiao Ma |
KSEM (2) | 3 |
| 2021 | HI-CMLM: Improve CMLM with Hybrid Decoder InputabstractMask-predict CMLM (Ghazvininejad et al., 2019) has achieved stunning performance among non-autoregressive NMT models, but we find that the mechanism of predicting all of the target words only depending on the hidden state of [MASK] is not effective and efficient in initial iterations of refinement, resulting in ungrammatical repetitions and slow convergence.In this work, we mitigate this problem by combining copied source with embeddings of [MASK] in decoder.Notably.it's not a straightforward copying that is shown to be useless, but a novel heuristic hybrid strategy -fence-mask.Experimental results show that it gains consistent boosts on both WMT14 En↔De and WMT16 En↔Ro corpus by 0.5 BLEU on average, and 1 BLEU for lessinformative short sentences.This reveals that incorporating additional information by proper strategies is beneficial to improve CMLM, particularly translation quality of short texts and speeding up early-stage convergence. Minghan Wang, Yuxia Wang 0003, Chang Su 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
INLG | 1 |
| 2021 | Make the Blind Translator See The World: A Novel Transfer Learning Solution for Multimodal Machine TranslationabstractBased on large-scale pretrained networks and the liability to be easily overfitting with limited labelled training data of multimodal translation (MMT) is a critical issue in MMT. To this end and we propose a transfer learning solution. Specifically and 1) A vanilla Transformer is pre-trained on massive bilingual text-only corpus to obtain prior knowledge; 2) A multimodal Transformer named VLTransformer is proposed with several components incorporated visual contexts; and 3) The parameters of VLTransformer are initialized with the pre-trained vanilla Transformer and then being fine-tuned on MMT tasks with a newly proposed method named cross-modal masking which forces the model to learn from both modalities. We evaluated on the Multi30k en-de and en-fr dataset and improving up to 8% BLEU score compared with the SOTA performance. The experimental result demonstrates that performing transfer learning with monomodal pre-trained NMT model on multimodal NMT tasks can obtain considerable boosts. Minghan Wang, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
MTSummit (1) | 1 |
| 2020 | Unified Humor Detection Based on Sentence-pair Augmentation and Transfer LearningabstractWe propose a unified multilingual model for humor detection which can be trained under a transfer learning framework. 1) The model is built based on pre-trained multilingual BERT, thereby is able to make predictions on Chinese, Russian and Spanish corpora. 2) We step out from single sentence classification and propose sequence-pair prediction which considers the inter-sentence relationship. 3) We propose the Sentence Discrepancy Prediction (SDP) loss, aiming to measure the semantic discrepancy of the sequence-pair, which often appears in the setup and punchline of a joke. Our method achieves two SoTA and a second-place on three humor detection corpora in three languages (Russian, Spanish and Chinese), and also improves F1-score by 4%-6%, which demonstrates the effectiveness of it in humor detection tasks. Minghan Wang, Hao Yang 0006, Shiliang Sun |
EAMT | 1 |
| 2020 | Efficient Transfer Learning for Quality Estimation with Bottleneck Adapter LayerabstractThe Predictor-Estimator framework for quality estimation (QE) is commonly used for its strong performance. Where the predictor and estimator works on feature extraction and quality evaluation, respectively. However, training the predictor from scratch is computationally expensive. In this paper, we propose an efficient transfer learning framework to transfer knowledge from NMT dataset into QE models. A Predictor-Estimator alike model named BAL-QE is also proposed, aiming to extract high quality features with pre-trained NMT model, and make classification with a fine-tuned Bottleneck Adapter Layer (BAL). The experiment shows that BAL-QE achieves 97% of the SOTA performance in WMT19 En-De and En-Ru QE tasks by only training 3% of parameters within 4 hours on 4 Titan XP GPUs. Compared with the commonly used NuQE baseline, BAL-QE achieves 47% (En-Ru) and 75% (En-De) of performance promotions. Hao Yang 0006, Minghan Wang |
EAMT | 2 |