EDBT 2026 Demo / reviewers in the wild / expert
Zhongqiang Huang
dblp:10/3565
· DBLP profile ↗
52ranked-venue papers
10as first author
27since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 8 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | McHirc: A Multimodal Benchmark for Chinese Idiom Reading ComprehensionabstractThe performance of various tasks of natural language processing has greatly improved with the emergence of large language models. However, there is still much room for improvement in understanding certain specific linguistic phenomena, such as Chinese idioms, which are usually composed of four characters. Chinese idioms are difficult to understand due to semantic gaps between their literal and actual meanings. Researchers have proposed the Chinese idiom reading comprehension task to examine the ability of large language models to represent and understand Chinese idioms. The task requires choosing the correct Chinese idiom from a list of candidates to complete the sentence. The current research mainly focuses on text-based idiom comprehension. Nevertheless, there are many idiom application scenarios that combine images and text, and we believe that the corresponding images are beneficial for the model's understanding of the idioms. Therefore, to address the above problems, we first construct a large-scale Multimodal Chinese Idiom Reading Comprehension dataset (MChIRC), which contains a total of 44,433 image-text pairs covering 2,926 idioms. Then, we propose a Dual-Contrastive Idiom Graph Network (DCIGN), which employs a dual-contrastive learning module to align the text and image features corresponding to the same Chinese idiom at both coarse and fine levels, while utilizing a graph structure to capture the semantic relationships between idiom candidates. Finally, we use a cross-attention module to fuse multimodal features with graph features of candidate idioms to predict correct answers. The authoritativeness of MChIRC and the effectiveness of DCIGN are demonstrated through a variety of experiments, which provides a new benchmark for the multimodal Chinese idiom reading comprehension task. Tongguan Wang, Mingmin Wu, Guixin Su, Dongyu Su, Yuxue Hu, Zhongqiang Huang, Ying Sha |
AAAI | 6 |
| 2025 | Multi-Hop Attention Diffusion Graph Neural Networks For Multimodal Fake News Detection
Zhongqiang Huang, Dongli Lu, Ying Sha |
ICASSP | 1 |
| 2024 | Divergence-Guided Simultaneous Speech TranslationabstractTo achieve high-quality translation with low latency, a Simultaneous Speech Translation (SimulST) system relies on a policy module to decide whether to translate immediately or wait for additional streaming input, along with a translation model capable of effectively handling partial speech input. Prior research has tackled these components separately, either using ``wait-k'' policies based on fixed-length segments or detected word boundaries, or dynamic policies based on different strategies (e.g., meaningful units), while employing offline models for prefix-to-prefix translation. In this paper, we propose Divergence-Guided Simultaneous Speech Translation (DiG-SST), a tightly integrated approach focusing on both translation quality and latency for streaming input. Specifically, we introduce a simple yet effective prefix-based strategy for training translation models with partial speech input, and develop an adaptive policy that makes read/write decisions for the translation model based on the expected divergence in translation distributions resulting from future input. Our experiments on multiple translation directions of the MuST-C benchmark demonstrate that our approach achieves a better trade-off between translation quality and latency compared to existing methods. Xinjie Chen, Kai Fan 0002, Xinggao Liu, Zhongqiang Huang |
AAAI | 7 |
| 2024 | Uncovering and Mitigating the Hidden Chasm: A Study on the Text-Text Domain Gap in Euphemism IdentificationabstractEuphemisms are commonly used on social media and darknet marketplaces to evade platform regulations by masking their true meanings with innocent ones. For instance, “weed” is used instead of “marijuana” for illicit transactions. Thus, euphemism identification, i.e., mapping a given euphemism (“weed”) to its specific target word (“marijuana”), is essential for improving content moderation and combating underground markets. Existing methods employ self-supervised schemes to automatically construct labeled training datasets for euphemism identification. However, they overlook the text-text domain gap caused by the discrepancy between the constructed training data and the test data, leading to performance deterioration. In this paper, we present the text-text domain gap and explain how it forms in terms of the data distribution and the cone effect. Moreover, to bridge this gap, we introduce a feature alignment network (FA-Net), which can both align the in-domain and cross-domain features, thus mitigating the domain gap from training data to test data and improving the performance of the base models for euphemism identification. We apply this FA-Net to the base models, obtaining markedly better results, and creating a state-of-the-art model which beats the large language models. Yuxue Hu, Junsong Li, Mingmin Wu, Zhongqiang Huang, Ying Sha |
AAAI | 4 |
| 2024 | Refining Idioms Semantics Comprehension via Contrastive Learning and Cross-AttentionabstractChinese idioms on social media demand a nuanced understanding for correct usage. The Chinese idiom cloze test poses a unique challenge for machine reading comprehension due to the figurative meanings of idioms deviating from their literal interpretations, resulting in a semantic bias in models’ comprehension of idioms. Furthermore, given that the figurative meanings of many idioms are similar, their use as suboptimal options can interfere with optimal selection. Despite achieving some success in the Chinese idiom cloze test, existing methods based on deep learning still struggle to comprehensively grasp idiom semantics due to the aforementioned issues. To tackle these challenges, we introduce a Refining Idioms Semantics Comprehension Framework (RISCF) to capture the comprehensive idioms semantics. Specifically, we propose a semantic sense contrastive learning module to enhance the representation of idiom semantics, diminishing the semantic bias between figurative and literal meanings of idioms. Meanwhile, we propose an interference-resistant cross-attention module to attenuate the interference of suboptimal options, which considers the interaction between the candidate idioms and the blank space in the context. Experimental results on the benchmark datasets demonstrate the effectiveness of our RISCF model, which outperforms state-of-the-art methods significantly. Mingmin Wu, Guixin Su, Yongcheng Zhang, Zhongqiang Huang, Ying Sha |
LREC/COLING | 4 |
| 2024 | BLSP-Emo: Towards Empathetic Large Speech-Language ModelsabstractThe recent release of GPT-4o showcased the potential of end-to-end multimodal models, not just in terms of low latency but also in their ability to understand and generate expressive speech with rich emotions.While the details are unknown to the open research community, it likely involves significant amounts of curated data and compute, neither of which is readily accessible.In this paper, we present BLSP-Emo (Bootstrapped Language-Speech Pretraining with Emotion support), a novel approach to developing an end-to-end speechlanguage model capable of understanding both semantics and emotions in speech and generate empathetic responses.BLSP-Emo utilizes existing speech recognition (ASR) and speech emotion recognition (SER) datasets through a two-stage process.The first stage focuses on semantic alignment, following recent work on pretraining speech-language models using ASR data.The second stage performs emotion alignment with the pretrained speech-language model on an emotion-aware continuation task constructed from SER data.Our experiments demonstrate that the BLSP-Emo model excels in comprehending speech and delivering empathetic responses, both in instruction-following tasks and conversations. 1 Minpeng Liao, Zhongqiang Huang, Junhong Wu, Chengqing Zong, Jiajun Zhang 0001 |
EMNLP | 3 |
| 2024 | Euphemism Identification via Feature Fusion and IndividualizationabstractEuphemisms are widely used on social media and darknet markets to evade supervision. For instance, "ice" serves as a euphemism for the target keyword "methamphetamine" in illicit transactions. Thus, euphemism identification which aims to map the euphemism to its secret meaning (target keyword) is a crucial task in ensuring social network security. However, this task poses significant challenges, including resource limitations due to the unavailable of annotated datasets and linguistic challenges arising from subtle differences in meaning between target keywords. Existing methods employed self-supervised schemes to automatically construct labeled training data, addressing the resource limitations. Yet, these methods rely on static embedding methods that fail to distinguish between target keywords with similar meanings. In addition, we observe that different euphemisms in similar contexts confuse the identification results. To overcome these obstacles, we propose a feature fusion and individualization (FFI) method for euphemism identification. First, we reformulate the task as a cloze task, making it more feasible. Next, we develop a feature fusion module to capture both dynamic global and static local features, enhancing discrimination between different euphemisms in similar contexts. Additionally, we employ a feature individualization module to ensure each target keyword has a unique feature representation by projecting features into their orthogonal space. As a result, FFI can effectively identify similar euphemisms that refer to target keywords with similar meanings. Experimental results demonstrate that our method outperforms state-of-the-art methods and large language models, providing robust support for its effectiveness. Yuxue Hu, Mingmin Wu, Zhongqiang Huang, Junsong Li, Xing Ge, Ying Sha |
WWW | 3 |
| 2023 | Better Simultaneous Translation with Monotonic Knowledge DistillationabstractSimultaneous machine translation (SiMT) presents a unique challenge as it requires generating target tokens before the source sentence is fully consumed.This can lead to the hallucination problem, where target tokens are generated without support from the source sentence.The prefix-to-prefix training data used to train SiMT models are not always parallel, due to divergent word order between the source and target languages, and can contribute to the problem.In this paper, we propose a novel approach that leverages traditional translation models as teachers and employs a two-stage beam search algorithm to generate monotonic yet accurate reference translations for sequence-level knowledge distillation.Experimental results demonstrate the significant improvements achieved by our approach over multiple strong SiMT baselines, leading to new state-of-the-art performance across various language pairs.Notably, when evaluated on a monotonic version of the WMT15 De→En test set, which includes references generated in a more monotonic style by professional translators, our approach achieves even more substantial improvement over the baselines.The source code and data are publicly available for further exploration 1 . Shushu Wang, Kai Fan 0002, Zhongqiang Huang |
ACL (1) | 6 |
| 2023 | A Prompt-Based Representation Individual Enhancement Method for Chinese Idiom Reading Comprehension
Ying Sha, Mingmin Wu, Zhi Zeng 0001, Xing Ge, Zhongqiang Huang, Huan Wang 0005 |
DASFAA (3) | 5 |
| 2023 | Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and InferenceabstractA popular approach to streaming speech translation is to employ a single offline model with a wait-k policy to support different latency requirements, which is simpler than training multiple online models with different latency constraints.However, there is a mismatch problem in using a model trained with complete utterances for streaming inference with partial input.We demonstrate that speech representations extracted at the end of a streaming input are significantly different from those extracted from a complete utterance.To address this issue, we propose a new approach called Future-Aware Streaming Translation (FAST) that adapts an offline ST model for streaming input.FAST includes a Future-Aware Inference (FAI) strategy that incorporates future context through a trainable masked embedding, and a Future-Aware Distillation (FAD) framework that transfers future context from an approximation of full speech to streaming input.Our experiments on the MuST-C EnDe, EnEs, and EnFr benchmarks show that FAST achieves better trade-offs between translation quality and latency than strong baselines.Extensive analyses suggest that our methods effectively alleviate the aforementioned mismatch problem between offline training and online inference.1 Biao Fu, Minpeng Liao, Kai Fan 0002, Zhongqiang Huang, Boxing Chen, Yidong Chen 0001, Xiaodong Shi |
EMNLP | 4 |
| 2023 | Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyabstractSimultaneous Speech Translation (SimulST) is a task focused on ensuring high-quality translation of speech in low-latency situations.Despite this, the modality gap (e.g., unknown word boundaries) between audio and text presents a challenge.This gap hinders the effective application of policies from simultaneous text translation (SimulMT) and compromises the performance of offline speech translation.To address this issue, we first leverage the Montreal Forced Aligner (MFA) and utilize audio transcription pairs in pre-training the acoustic encoder, and introduce a token-level crossmodal alignment that allows the wait-k policy from SimulMT to better adapt to SimulST.This token-level boundary alignment simplifies the decision-making process for predicting read/write actions, as if the decoder were directly processing text tokens.Subsequently, to optimize the SimulST task, we propose a robust and random wait-k-tokens strategy.This strategy allows a single model to meet various latency requirements and minimizes error accumulation of boundary alignment during inference.Our experiments on the MuST-C dataset show that our method achieves a better tradeoff between translation quality and latency. Kai Fan 0002, Jiajun Bu, Zhongqiang Huang |
EMNLP | 4 |
| 2023 | Adaptive Policy with Wait-k Model for Simultaneous TranslationabstractSimultaneous machine translation (SiMT) requires a robust read/write (R/W) policy in conjunction with a high-quality translation model.Traditional methods rely on either a fixed waitk policy coupled with a standalone wait-k translation model, or an adaptive policy jointly trained with the translation model.In this study, we propose a more flexible approach by decoupling the adaptive policy model from the translation model.Our motivation stems from the observation that a standalone multi-path waitk model performs competitively with adaptive policies utilized in state-of-the-art SiMT approaches.Specifically, we introduce DaP, a divergence-based adaptive policy, that makes read/write decisions for any translation model based on the potential divergence in translation distributions resulting from future information.DaP extends a frozen wait-k model with lightweight parameters, and is both memory and computation efficient.Experimental results across various benchmarks demonstrate that our approach offers an improved trade-off between translation accuracy and latency, outperforming strong baselines.1 Kai Fan 0002, Shushu Wang, Ziqian Zeng, Zhongqiang Huang |
EMNLP | 7 |
| 2023 | Multimodal Stacked Cross Attention Network for Fine-Grained Fake News DetectionabstractFake news is usually disseminated in a multimodal form, which incorporates natural language, visual language, and so on. Therefore, many deep learning approaches are proposed to detect multimodal fake news. However, a drawback of existing methods is that they simply fuse unimodal features and ignore the latent semantic alignment of image and text modalities. In this paper, we propose a novel Multimodal Stacked Cross Attention Network (MSCA) to better align and fuse multimodal token-level textual and visual features for fake news detection. Experiments conducted on two publicly available datasets show that our method can significantly improve performance compared with other models. Furthermore, experimental analysis shows that MSCA can effectively align and fuse token-level features of multiple modalities. Zhongqiang Huang, Yuxue Hu, Zhi Zeng 0001, Xiang Li 0111, Ying Sha |
ICME | 1 |
| 2023 | An Explainable Multi-view Semantic Fusion Model for Multimodal Fake News DetectionabstractThe existing models have been achieved great success in capturing and fusing miltimodal semantics of news. However, they paid more attention to the global information, ignoring the interactions of global and local semantics and the inconsistency between different modalities. Therefore, we propose an explainable multi-view semantic fusion model (EMSFM), where we aggregate the important inconsistent semantics from local and global views to compensate the global information. Inspired by various forms of artificial fake news and real news, we summarize four views of multimodal correlation: consistency and inconsistency in the local and global views. Integrating these four views, our EMSFM can interpretatively establish global and local fusion between consistent and inconsistent semantics in multimodal relations for fake news detection. The extensive experimental results show that the EMSFM can improve the performance of multimodal fake news detection and provide a novel paradigm for explainable multi-view semantic fusion. Zhi Zeng 0001, Mingmin Wu, Xiang Li 0046, Zhongqiang Huang, Ying Sha |
ICME | 5 |
| 2023 | Correcting the Bias: Mitigating Multimodal Inconsistency Contrastive Learning for Multimodal Fake News DetectionabstractMultimodal fake news detection has become a topical research of fake news detection. Existing models have made great efforts in capturing and fusing multimodal semantics of news for classification. However, they overlooked mitigating inconsistency between different modalities, which may result in learning biased statistical information. Therefore, we propose a mitigating multimodal inconsistency contrastive learning framework (MMICF), which mitigates inconsistency in multi-modal relations for fake news detection. Inspired by various forms of artificial fake news, we summarize two patterns of multimodal inconsistency: local and global inconsistency. To mitigate local inconsistency in multimodal relations, we use a causal-relation reasoning module by causally removing the direct effects of the textual and visual entities. Considering the influence of global inconsistency in multimodal semantics, our contrastive learning framework mitigates the semantic deviation of contrastive text-image objectives, which are constrained into a unified semantic space by a modal unified module. Thus, our MMICF can jointly mitigate local and global inconsistency for further maximally exploiting multimodal consistent semantics for fake news detection. The extensive experimental results show that the MMICF can improve the performance of multimodal fake news detection and provide a novel paradigm for mitigating multimodal inconsistency contrastive learning. Zhi Zeng 0001, Mingmin Wu, Xiang Li 0046, Zhongqiang Huang, Ying Sha |
ICME | 5 |
| 2022 | Discrete Cross-Modal Alignment Enables Zero-Shot Speech TranslationabstractEnd-to-end Speech Translation (ST) aims at translating the source language speech into target language text without generating the intermediate transcriptions.However, the training of end-to-end methods relies on parallel ST data, which are difficult and expensive to obtain.Fortunately, the supervised data for automatic speech recognition (ASR) and machine translation (MT) are usually more accessible, making zero-shot speech translation a potential direction.Existing zero-shot methods fail to align the two modalities of speech and text into a shared semantic space, resulting in much worse performance compared to the supervised ST methods.In order to enable zero-shot ST, we propose a novel Discrete Cross-Modal Alignment (DCMA) method that employs a shared discrete vocabulary space to accommodate and match both modalities of speech and text.Specifically, we introduce a vector quantization module to discretize the continuous representations of speech and text into a finite set of virtual tokens, and use ASR data to map corresponding speech and text to the same virtual token in a shared codebook.This way, source language speech can be embedded in the same semantic space as the source language text, which can be then transformed into target language text with an MT module.Experiments on multiple language pairs demonstrate that our zero-shot ST method significantly improves the SOTA, and even performs on par with the strong supervised ST baselines 1 . Yuchen Liu 0007, Boxing Chen, Jiajun Zhang 0001, Zhongqiang Huang, Chengqing Zong |
EMNLP | 6 |
| 2022 | ITA: Image-Text Alignments for Multi-Modal Named Entity RecognitionabstractXinyu Wang, Min Gui, Yong Jiang, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Kewei Tu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Xinyu Wang 0013, Min Gui, Yong Jiang 0005, Zixia Jia, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Kewei Tu |
NAACL-HLT | 7 |
| 2021 | Bridging the Domain Gap: Improve Informal Language Translation via Counterfactual Domain AdaptationabstractDespite the near-human performances already achieved on formal texts such as news articles, neural machine translation still has difficulty in dealing with "user-generated" texts that have diverse linguistic phenomena but lack large-scale high-quality parallel corpora. To address this problem, we propose a counterfactual domain adaptation method to better leverage both large-scale source-domain data (formal texts) and small-scale target-domain data (informal texts). Specifically, by considering effective counterfactual conditions (the concatenations of source-domain texts and the target-domain tag), we construct the counterfactual representations to fill the sparse latent space of the target domain caused by a small amount of data, that is, bridging the gap between the source-domain data and the target-domain data. Experiments on English-to-Chinese and Chinese-to-English translation tasks show that our method outperforms the base model that is trained only on the informal corpus by a large margin, and consistently surpasses different baseline methods by +1.12 ~ 4.34 BLEU points on different datasets. Furthermore, we also show that our method achieves competitive performances on cross-domain language translation on four language pairs. Guandan Chen, Zhongqiang Huang, Xiaojun Wan 0001, Fei Huang 0002 |
AAAI | 3 |
| 2021 | Multi-View Cross-Lingual Structured Prediction with Minimum SupervisionabstractZechuan Hu, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zechuan Hu, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 5 |
| 2021 | Risk Minimization for Zero-shot Sequence LabelingabstractZechuan Hu, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zechuan Hu, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 5 |
| 2021 | Improving Named Entity Recognition by External Context Retrieving and Cooperative LearningabstractXinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 5 |
| 2021 | Automated Concatenation of Embeddings for Structured PredictionabstractXinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 5 |
| 2021 | Structural Knowledge Distillation: Tractably Distilling Information for Structured PredictorabstractXinyu Wang, Yong Jiang, Zhaohui Yan, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Wang 0013, Yong Jiang 0005, Zhaohui Yan 0001, Zixia Jia, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 7 |
| 2021 | Word Reordering for Zero-shot Cross-lingual Structured PredictionabstractAdapting word order from one language to another is a key problem in cross-lingual structured prediction.Current sentence encoders (e.g., RNN, Transformer with position embeddings) are usually word order sensitive.Even with uniform word form representations (MUSE, mBERT), word order discrepancies may hurt the adaptation of models.This paper builds structured prediction models with bag-of-words inputs.It introduces a new reordering module to organize words following the source language order, which learns taskspecific reordering strategies from a generalpurpose order predictor model.Experiments on zero-shot cross-lingual dependency parsing, POS tagging, and morphological tagging show that our model can significantly improve target language performances, especially for languages that are distant from the source language.1 Yong Jiang 0005, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Yuanbin Wu |
EMNLP (1) | 4 |
| 2021 | A Unified Encoding of Structures in Transition SystemsabstractTransition systems usually contain various dynamic structures (e.g., stacks, buffers).An ideal transition-based model should encode these structures completely and efficiently.Previous works relying on templates or neural network structures either only encode partial structure information or suffer from computation efficiency.In this paper, we propose a novel attention-based encoder unifying representation of all structures in a transition system.Specifically, we separate two views of items on structures, namely structure-invariant view and structure-dependent view.With the help of parallel-friendly attention network, we are able to encoding transition states with O(1) additional complexity (with respect to basic feature extractors).Experiments on the PTB and UD show that our proposed method significantly improves the test speed and achieves the best transition-based model, and is comparable to state-of-the-art methods. 1 Yong Jiang 0005, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Yuanbin Wu |
EMNLP (1) | 4 |
| 2021 | MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity RepresentationsabstractEntity retrieval, which aims at disambiguating mentions to canonical entities from massive KBs, is essential for many tasks in natural language processing.Recent progress in entity retrieval shows that the dual-encoder structure is a powerful and efficient framework to nominate candidates if entities are only identified by descriptions.However, they ignore the property that meanings of entity mentions diverge in different contexts and are related to various portions of descriptions, which are treated equally in previous works.In this work, we propose Multi-View Entity Representations (MuVER), a novel approach for entity retrieval that constructs multi-view representations for entity descriptions and approximates the optimal view for mentions via a heuristic searching method.Our method achieves the state-ofthe-art performance on ZESHEL and improves the quality of candidates on three standard Entity Linking datasets 1 . Xinyin Ma, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Weiming Lu 0001 |
EMNLP (1) | 5 |
| 2021 | Automatically Paraphrasing via Sentence Reconstruction and Round-trip TranslationabstractParaphrase generation plays key roles in NLP tasks such as question answering, machine translation, and information retrieval. In this paper, we propose a novel framework for paraphrase generation. It simultaneously decodes the output sentence using a pretrained wordset-to-sequence model and a round-trip translation model. We evaluate this framework on Quora, WikiAnswers, MSCOCO and Twitter, and show its advantage over previous state-of-the-art unsupervised methods and distantly-supervised methods by significant margins on all datasets. For Quora and WikiAnswers, our framework even performs better than some strongly supervised methods with domain adaptation. Further, we show that the generated paraphrases can be used to augment the training data for machine translation to achieve substantial improvements. Zilu Guo, Zhongqiang Huang, Kenny Q. Zhu, Guandan Chen, Kaibo Zhang, Boxing Chen, Fei Huang 0002 |
IJCAI | 2 |
| 2020 | Alignment-Enhanced Transformer for Constraining NMT with Pre-Specified TranslationsabstractWe investigate the task of constraining NMT with pre-specified translations, which has practical significance for a number of research and industrial applications. Existing works impose pre-specified translations as lexical constraints during decoding, which are based on word alignments derived from target-to-source attention weights. However, multiple recent studies have found that word alignment derived from generic attention heads in the Transformer is unreliable. We address this problem by introducing a dedicated head in the multi-head Transformer architecture to capture external supervision signals. Results on five language pairs show that our method is highly effective in constraining NMT with pre-specified translations, consistently outperforming previous methods in translation quality. Heng Yu 0006, Yue Zhang 0004, Zhongqiang Huang, Weihua Luo, Xiangyu Duan, Min Zhang 0005 |
AAAI | 5 |
| 2020 | AIN: Fast and Accurate Sequence Labeling with Approximate Inference NetworkabstractThe linear-chain Conditional Random Field (CRF) model is one of the most widely-used neural sequence labeling approaches.Exact probabilistic inference algorithms such as the forward-backward and Viterbi algorithms are typically applied in training and prediction stages of the CRF model.However, these algorithms require sequential computation that makes parallelization impossible.In this paper, we propose to employ a parallelizable approximate variational inference algorithm for the CRF model.Based on this algorithm, we design an approximate inference network that can be connected with the encoder of the neural CRF model to form an end-to-end network, which is amenable to parallelization for faster training and prediction.The empirical results show that our proposed approaches achieve a 12.7-fold improvement in decoding speed with long sentences and a competitive accuracy compared with the traditional CRF approach. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
EMNLP (1) | 5 |
| 2020 | Towards Better Word Alignment in TransformerabstractWhile neural models based on the Transformer architecture achieve the State-of-the-Art translation performance, it is well known that the learned target-to-source attentions do not correlate well with word alignment. There is an increasing interest in inducing accurate word alignment in Transformer, due to its important role in practical applications such as dictionary-guided translation and interactive translation. In this article, we extend and improve the recent work on unsupervised learning of word alignment in Transformer on two dimensions: a) parameter initialization from a pre-trained cross-lingual language model to leverage large amounts of monolingual data for learning robust contextualized word representations, and b) regularization of the training objective to directly model characteristics of word alignments which results in favorable word alignments receiving more concentrated probabilities. Experiments on benchmark data sets of three language pairs show that the proposed methods can significantly reduce alignment error rate (AER) by at least 3.7 to 7.7 points on each language pair over two recent works on improving the Transformer's word alignment. Moreover, our methods can achieve better alignment results than GIZA++ on certain test sets. Xiaoqing Zhou, Heng Yu 0006, Zhongqiang Huang, Yue Zhang 0004, Weihua Luo, Xiangyu Duan, Min Zhang 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Neural-Network Lexical Translation for Cross-lingual IR from Text and SpeechabstractWe propose a neural network model to estimate word translation probabilities for Cross-Lingual Information Retrieval (CLIR). The model estimates better probabilities for word translations than automatic word alignments alone, and generalizes to unseen source-target word pairs. We further improve the lexical neural translation model (and subsequently CLIR), by incorporating source word context, and by encoding the character sequences of input source words to generate translations of out-of-vocabulary words. To be effective, neural network models typically need training on large amounts of data labeled directly on the final task, in this case relevance to queries. In contrast, our approach only requires parallel data to train the translation model, and uses an unsupervised model to compute CLIR relevance scores. Rabih Zbib, Lingjun Zhao, Damianos Karakos, William Hartmann, Jay DeYoung, Zhongqiang Huang, Zhuolin Jiang, Noah Rivkin, Le Zhang 0002, Richard M. Schwartz, John Makhoul |
SIGIR | 6 |
| 2018 | Automatic Speech Recognition and Topic Identification from Speech for Almost-Zero-Resource Languages
Matthew Wiesner, Chunxi Liu, Lucas Ondel Yang, Craig Harman, Vimal Manohar, Jan Trmal, Zhongqiang Huang, Najim Dehak, Sanjeev Khudanpur |
INTERSPEECH | 7 |
| 2018 | Dynamic Multi-objective Estimation of Distribution Algorithm based on Domain Adaptation and Nonparametric Estimation
Min Jiang 0005, Liming Qiu, Zhongqiang Huang, Gary G. Yen |
Inf. Sci. | 3 |
| 2018 | BBN's low-resource machine translation for the LoReHLT 2016 evaluation
Hendra Setiawan, Zhongqiang Huang, Rabih Zbib |
Mach. Transl. | 2 |
| 2018 | Transfer Learning-Based Dynamic Multiobjective Optimization AlgorithmsabstractOne of the major distinguishing features of the dynamic multiobjective optimization problems (DMOPs) is that optimization objectives will change over time, thus tracking the varying Pareto-optimal front becomes a challenge. One of the promising solutions is reusing “experiences” to construct a prediction model via statistical machine learning approaches. However, most existing methods neglect the nonindependent and identically distributed nature of data to construct the prediction model. In this paper, we propose an algorithmic framework, called transfer learning-based dynamic multiobjective evolutionary algorithm (EA), which integrates transfer learning and population-based EAs to solve the DMOPs. This approach exploits the transfer learning technique as a tool to generate an effective initial population pool via reusing past experience to speed up the evolutionary process, and at the same time any population-based multiobjective algorithms can benefit from this integration without any extensive modifications. To verify this idea, we incorporate the proposed approach into the development of three well-known EAs, nondominated sorting genetic algorithm II, multiobjective particle swarm optimization, and the regularity model-based multiobjective estimation of distribution algorithm. We employ 12 benchmark functions to test these algorithms as well as compare them with some chosen state-of-the-art designs. The experimental results confirm the effectiveness of the proposed design for DMOPs. Min Jiang 0005, Zhongqiang Huang, Liming Qiu, Wenzhen Huang, Gary G. Yen |
IEEE Trans. Evol. Comput. | 2 |
| 2017 | Integration of Global and Local Metrics for Domain Adaptation Learning Via Dimensionality ReductionabstractDomain adaptation learning (DAL) investigates how to perform a task across different domains. In this paper, we present a kernelized local-global approach to solve domain adaptation problems. The basic idea of the proposed method is to consider the global and local information regarding the domains (e.g., maximum mean discrepancy and intraclass distance) and to convert the domain adaptation problem into a bi-object optimization problem via the kernel method. A solution for the optimization problem will help us identify a latent space in which the distributions of the different domains will be close to each other in the global sense, and the local properties of the labeled source samples will be preserved. Therefore, classic classification algorithms can be used to recognize unlabeled target domain data, which has a significant difference on the source samples. Based on the analysis, we validate the proposed algorithm using four different sources of data: synthetic, textual, object, and facial image. The experimental results indicate that the proposed method provides a reasonable means to improve DAL algorithms. Min Jiang 0005, Wenzhen Huang, Zhongqiang Huang, Gary G. Yen |
IEEE Trans. Cybern. | 3 |
| 2016 | Sage: The New BBN Speech Processing Platform
Roger Hsiao, Ralf Meermeier, Tim Ng, Zhongqiang Huang, Maxwell Jordan, Enoch Kan, Tanel Alumäe, Jan Silovský, William Hartmann, Francis Keith, Omer Lang, Man-Hung Siu, Owen Kimball |
INTERSPEECH | 4 |
| 2015 | Statistical Machine Translation Features with Multitask Tensor NetworksabstractWe present a three-pronged approach to improving Statistical Machine Translation (SMT), building on recent success in the application of neural networks to SMT. First, we propose new features based on neural networks to model various non-local translation phenomena. Second, we augment the architecture of the neural network with tensor layers that capture important higher-order interaction among the network units. Third, we apply multitask learning to estimate the neural network parameters jointly. Each of our proposed methods results in significant improvements that are complementary. The overall improvement is +2.7 and +1.8 BLEU points for Arabic-English and Chinese-English translation over a state-of-the-art system that already includes neural network features. Hendra Setiawan, Zhongqiang Huang, Jacob Devlin, Thomas Lamar, Rabih Zbib, Richard M. Schwartz, John Makhoul |
ACL (1) | 2 |
| 2014 | Fast and Robust Neural Network Joint Models for Statistical Machine TranslationabstractJacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard Schwartz, John Makhoul. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2014. Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard M. Schwartz, John Makhoul |
ACL (1) | 3 |
| 2014 | Improving machine vision via incorporating expectation-maximization into Deep Spatio-Temporal learningabstractThe Deep Spatio-Temporal Inference Network (DeSTIN) is a deep learning architecture which combines un-supervised learning and Bayesian inference. The original version of DeSTIN incorporates k-means clustering inside each processing node. Here we propose to replace k-means with a more sophisticated algorithm, online EM (Expectation Maximization), and show that this improves DeSTIN's performance on image classification and restoration tasks. Min Jiang 0005, Ben Goertzel, Zhongqiang Huang, Changle Zhou, Fei Chao 0001 |
IJCNN | 4 |
| 2013 | Factored Soft Source Syntactic Constraints for Hierarchical Machine TranslationabstractThis paper describes a factored approach to incorporating soft source syntactic constraints into a hierarchical phrase-based translation system.In contrast to traditional approaches that directly introduce syntactic constraints to translation rules by explicitly decorating them with syntactic annotations, which often exacerbate the data sparsity problem and cause other problems, our approach keeps translation rules intact and factorizes the use of syntactic constraints through two separate models: 1) a syntax mismatch model that associates each nonterminal of a translation rule with a distribution of tags that is used to measure the degree of syntactic compatibility of the translation rule on source spans; 2) a syntax-based reordering model that predicts whether a pair of sibling constituents in the constituent parse tree of the source sentence should be reordered or not when translated to the target language.The features produced by both models are used as soft constraints to guide the translation process.Experiments on Chinese-English translation show that the proposed approach significantly improves a strong string-to-dependency translation system on multiple evaluation sets. Zhongqiang Huang, Jacob Devlin, Rabih Zbib |
EMNLP | 1 |
| 2011 | Feature-Rich Log-Linear Lexical Model for Latent Variable PCFG Grammars
Zhongqiang Huang, Mary P. Harper |
IJCNLP | 1 |
| 2010 | Lessons Learned in Part-of-Speech Tagging of Conversational Speech
Vladimir Eidelman, Zhongqiang Huang, Mary P. Harper |
EMNLP | 2 |
| 2010 | Soft Syntactic Constraints for Hierarchical Phrase-Based Translation Using Latent Syntactic Distributions
Zhongqiang Huang, Martin Cmejrek |
EMNLP | 1 |
| 2010 | Self-Training with Products of Latent Variable Grammars
Zhongqiang Huang, Mary P. Harper, Slav Petrov |
EMNLP | 1 |
| 2010 | Appropriately Handled Prosodic Breaks Help PCFG Parsing
Zhongqiang Huang, Mary P. Harper |
HLT-NAACL | 1 |
| 2009 | Self-Training PCFG Grammars with Latent Annotations Across Languages
Zhongqiang Huang, Mary P. Harper |
EMNLP | 1 |
| 2007 | Mandarin Part-of-Speech Tagging and Discriminative Reranking
Zhongqiang Huang, Mary P. Harper, Wen Wang 0001 |
EMNLP-CoNLL | 1 |
| 2007 | Semi-Supervised Learning for Part-of-Speech Tagging of Mandarin Transcribed SpeechabstractIn this paper, we investigate bootstrapping part-of-speech (POS) taggers for Mandarin broadcast news (BN) transcripts using co-training, by iteratively retraining two competitive POS taggers from a small set of labeled training data and a large set of unlabeled data. We compare co-training with self-training and our results show that the performance using co-training is significantly better than that from self-training and these semi-supervised learning methods significantly improve tagging accuracy over training only on the small labeled seed corpus. We also investigate a variety of example selection approaches for co-training and find that the computationally expensive, agreement-based selection approach and a more efficient selection approach based on maximizing training utility produce comparable tagging performance from resulting POS taggers. By applying co-training, we are able to build effective POS taggers for Mandarin transcribed speech with the tagging accuracy comparable to that obtained on newswire text. Wen Wang 0001, Zhongqiang Huang, Mary P. Harper |
ICASSP (4) | 2 |
| 2006 | Using maximum entropy (ME) model to incorporate gesture cues for SU detectionabstractAccurate identification of sentence units (SUs) in spontaneous speech has been found to improve the accuracy of speech recognition, as well as downstream applications such as parsing. In recent multimodal investigations, gestur]al features were utilized, in addition to lexical and prosodic cues from the speech channel, for detecting SUs in conversational interactions using a hidden Markov model (HMM) approach. Although this approach is computationally efficient and provides a convenient way to modularize the knowledge sources, it has two drawbacks for our SU task. First, standard HMM training methods maximize the joint probability of observations and hidden events, as opposed to the posterior probability of a hidden event given observations, a criterion more closely related to SU classification error. A second challenge for integrating gestural features is that their absence sanctions neither SU events nor non-events; it is only the co-timing of gestures with the speech channel that should impact our model. To address these problems, a Maximum Entropy (ME) model is used to combine multimodal cues for SU estimation. Experiments carried out on VACE multi-party meetings confirm that the ME modeling approach provides a solid framework for multimodal integration. Lei Chen 0004, Mary P. Harper, Zhongqiang Huang |
ICMI | 3 |
| 2006 | An Open Source Prosodic Feature Extraction Tool
Zhongqiang Huang, Lei Chen 0004, Mary P. Harper |
LREC | 1 |
| 2006 | Impact of Automatic Comma Prediction on POS/Name Tagging of speechabstractThis work looks at the impact of automatically predicted commas on part-of-speech (POS) and name tagging of speech recognition transcripts of Mandarin broadcast news. There is a significant gain in both POS and name tagging accuracy due to using automatically predicted commas over sentence boundary prediction alone. One difference between Mandarin and English is that there are two types of commas, and experiments here show that, while they can be reliably distinguished in automatic prediction, the distinction does not give a clear benefit for POS or name tagging. Dustin Hillard, Zhongqiang Huang, Heng Ji 0001, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Mari Ostendorf, Wen Wang 0001 |
SLT | 2 |