Bing Xiang

dblp:82/5456 · DBLP profile ↗
← Back
90ranked-venue papers
15as first author
30since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 72 · 7 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 11 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Adaptation of Embedding Models to Financial Filings Via LLM Distillation
Eliot Brenner, Dominic Seyler, Manjunath Hegde, Andrei Simion, Koustuv Dasgupta, Bing Xiang
IEEE Big Data6
2025 PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QA
abstract
While Large Language Models (LLMs) show great promise, their tendencies to hallucinate pose significant risks in high-stakes domains like finance, especially when used for regulatory reporting and decision-making. Existing hallucination detection benchmarks fail to capture the complexities of financial benchmarks, which require high numerical precision, nuanced understanding of the language of finance, and ability to handle long-context documents. To address this, we introduce PHANTOM, a novel benchmark dataset for evaluating hallucination detection in long-context financial QA. Our approach first generates a seed dataset of high-quality "query-answer-document (chunk)" triplets, with either hallucinated or correct answers - that are validated by human annotators and subsequently expanded to capture various context lengths and information placements. We demonstrate how PHANTOM allows fair comparison of hallucination detection models and provides insights into LLM performance, offering a valuable resource for improving hallucination detection in financial applications. Further, our benchmarking results highlight the severe challenges out-of-the-box models face in detecting real-world hallucinations on long context data, and establish some promising directions towards alleviating these challenges, by fine-tuning open-source LLMs using PHANTOM.
Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang
NeurIPS6
2025 FinIR: The 2nd Workshop on Financial Information Retrieval in the Era of Generative AI
abstract
Recent advancements in Generative AI, such as Large Language Models (LLMs), have demonstrated remarkable success across various general tasks. Extensive studies have explored leveraging generative models in finance, but significant challenges persist. This half-day workshop explores potential approaches and research directions to address these challenges by equipping generative models with advanced Information Retrieval (IR) models. Specifically, this workshop seeks to provide a platform for discussing innovative ideas that facilitate the advancement of IR technology to enrich generative models in finance from four key perspectives: (i) financial IR techniques (ii) financial IR benchmarking and evaluation (iii) financial systems and agents/assistants (iv) and trustworthiness, privacy and security when applying financial IR and generative models. This workshop aims to deepen understanding, accelerate progress, and support the advancement of IR technology to enhance generative models to address financial challenges.
Fengbin Zhu, Yunshan Ma 0002, Fuli Feng, Chao Wang 0049, Huan-Bo Luan, Guangnan Ye, Shuo Zhang 0006, Dhagash Mehta, Pingping Chen 0004, Bing Xiang, Tat-Seng Chua
SIGIR10
2024 Lightweight reranking for language model generations
abstract
Large Language Models (LLMs) can exhibit considerable variation in the quality of their sampled outputs.Reranking and selecting the best generation from the sampled set is a popular way of obtaining strong gains in generation quality.In this paper, we present a novel approach for reranking LLM generations.Unlike other techniques that might involve additional inferences or training a specialized reranker, our approach relies on easy to compute pairwise statistics between the generations that have minimal compute overhead.We show that our approach can be formalized as an extension of self-consistency and analyze its performance in that framework, theoretically as well as via simulations.We show strong improvements for selecting the best k generations for code generation tasks as well as robust improvements for the best generation for the tasks of autoformalization, summarization, and translation.While our approach only assumes black-box access to LLMs, we show that additional access to token probabilities can improve performance even further.
Siddhartha Jain 0001, Xiaofei Ma 0001, Anoop Deoras, Bing Xiang
ACL (1)4
2024 CoCoMIC: Code Completion by Jointly Modeling In-file and Cross-file Context
abstract
While pre-trained language models (LM) for code have achieved great success in code completion, they generate code conditioned only on the contents within the file, i.e., in-file context, but ignore the rich semantics in other files within the same project, i.e., project-level cross-file context, a critical source of information that is especially useful in modern modular software development. Such overlooking constrains code LMs’ capacity in code completion, leading to unexpected behaviors such as generating hallucinated class member functions or function calls with unexpected arguments. In this work, we propose CoCoMIC, a novel framework that jointly learns the in-file and cross-file context on top of code LMs. To empower CoCoMIC, we develop CCFinder, a static-analysis-based tool that locates and retrieves the most relevant project-level cross-file context for code completion. CoCoMIC successfully improves the existing code LM with a 33.94% relative increase in exact match and 28.69% in identifier matching for code completion when the cross-file context is provided. Finally, we perform a series of ablation studies and share valuable insights for future research on integrating cross-file context into code LMs.
Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang
LREC/COLING8
2024 Code Representation Learning at Scale
abstract
Recent studies have shown that code language model at scale demonstrate significant performance gains on downstream tasks, i.e., code generation. However, most of the existing works on code representation learning train models at a hundred million parameter scale using very limited pretraining corpora. In this work, we fuel code representation learning with a vast amount of code data via a two-stage pretraining scheme. We first train the encoders via a mix that leverages both randomness in masking language modeling and implicit structure and semantic aspects of programming language. We then enhance the representations via contrastive learning with hard negative and hard positive constructed in an unsupervised manner. We establish an off-the-shelf encoder model that persistently outperforms the existing models on a wide variety of downstream tasks by large margins. To comprehend the factors contributing to successful code representation learning, we conduct detailed ablations and share our findings on (i) a customized and effective token-level denoising scheme for source code; (ii) the importance of hard negatives and hard positives; (iii) how the proposed bimodal contrastive learning boost the cross-lingual semantic search performance; and (iv) how the pretraining schemes decide the downstream task performance scales with the model size.
Dejiao Zhang, Wasi Uddin Ahmad, Hantian Ding, Ramesh Nallapati, Dan Roth 0001, Xiaofei Ma 0001, Bing Xiang
ICLR8
2024 Bifurcated Attention for Single-Context Large-Batch Sampling
abstract
In our study, we present bifurcated attention, a method developed for language model inference in single-context batch sampling contexts. This approach aims to reduce redundant memory IO costs, a significant factor in latency for high batch sizes and long context lengths. Bifurcated attention achieves this by dividing the attention mechanism during incremental decoding into two distinct GEMM operations, focusing on the KV cache from prefill and the decoding process. This method ensures precise computation and maintains the usual computational load (FLOPs) of standard attention mechanisms, but with reduced memory IO. Bifurcated attention is also compatible with multi-query attention mechanism known for reduced memory IO for KV cache, further enabling higher batch size and context length. The resulting efficiency leads to lower latency, improving suitability for real-time applications, e.g., enabling massively-parallel answer generation without substantially increasing latency, enhancing performance when integrated with post-processing techniques such as reranking.
Ben Athiwaratkun, Sujan K. Gonugondla, Sanjay Krishna Gouda, Haifeng Qian, Hantian Ding, Qing Sun 0013, Jun Wang 0022, Jiacheng Guo, Liangfu Chen, Parminder Bhatia, Ramesh Nallapati, Sudipta Sengupta, Bing Xiang
ICML13
2024 Optimizing Response Time of Multi-factor Authentication Service in EV Charging System: A DRL-based Algorithm
abstract
Electric Vehicles (EVs) are being stepping into the spotlight as a sustainable alternative to gasoline vehicles. The security of the EV charging infrastructure is more and more critical, necessitating the deployment and application of Multifactor Authentication (MFA). Commonly, it is necessary to make a tradeoff between security and performance. This paper aims to improve the quality of user experience in a practical three-stage MFA process in a real EV charging system without sacrificing the MFA security capability. We first develop a novel analytical model to assess the performance of MFA service in terms of service response time. Then, we propose a Deep Reinforcement Learning (DRL)-based algorithm to optimize user convenience in terms of MFA service response time. Simulation results demonstrate the effectiveness of our algorithm.
Bing Xiang, Tianlin Yang, Qingwen Han, Xueming Zhao
ISPA1
2023 ReCode: Robustness Evaluation of Code Generation Models
abstract
Shiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang, Zijian Wang, Mingyue Shang, Varun Kumar, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shiqi Wang 0002, Haifeng Qian, Chenghao Yang 0001, Zijian Wang 0002, Mingyue Shang, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth 0001, Bing Xiang
ACL (1)14
2023 ContraCLM: Contrastive Learning For Causal Language Model
abstract
Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang, Feng Nan, Xiaopeng Li, Ming Tan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang 0002, Feng Nan, Xiaopeng Li 0002, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma 0001, Bing Xiang
ACL (1)12
2023 Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning
abstract
Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma, Patrick Ng, Zhiguo Wang, Bonan Min, William Yang Wang, Kathleen McKeown, Vittorio Castelli, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma 0005, Patrick Ng, Zhiguo Wang 0006, Bonan Min, William Yang Wang, Kathy McKeown, Vittorio Castelli, Dan Roth 0001, Bing Xiang
ACL (1)12
2023 Efficient Shapley Values Estimation by Amortization for Text Classification
abstract
Despite the popularity of Shapley Values in explaining neural text classification models, computing them is prohibitive for large pretrained models due to a large number of model evaluations.In practice, Shapley Values are often estimated with a small number of stochastic model evaluations.However, we show that the estimated Shapley Values are sensitive to random seed choices -the top-ranked features often have little overlap across different seeds, especially on examples with longer input texts.This can only be mitigated by aggregating thousands of model evaluations, which on the other hand, induces substantial computational overheads.To mitigate the trade-off between stability and efficiency, we develop an amortized model that directly predicts each input feature's Shapley Value without additional model evaluations.It is trained on a set of examples whose Shapley Values are estimated from a large number of model evaluations to ensure stability.Experimental results on two text classification datasets demonstrate that our amortized model estimates Shapley Values accurately with up to 60 times speedup compared to traditional methods.Furthermore, the estimated values are stable as the inference is deterministic.We release our code at https://github.com/yangalan123/ Amortized-Interpretability.
Chenghao Yang 0001, Fan Yin, He He 0001, Kai-Wei Chang 0001, Xiaofei Ma 0001, Bing Xiang
ACL (1)6
2023 Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness
Shuaichen Chang, Jun Wang 0122, Mingwen Dong, Lin Pan 0003, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang 0029, Jiarong Jiang, Joe Lilien, Steve Ash, William Yang Wang, Zhiguo Wang 0006, Vittorio Castelli, Patrick Ng, Bing Xiang
ICLR16
2023 STREET: A Multi-Task Structured Reasoning and Explanation Benchmark
Danilo Neves Ribeiro, Shen Wang 0005, Xiaofei Ma 0001, Henghui Zhu, Deguang Kong, Juliette Burger, Anjelica Ramos, Zhiheng Huang, William Yang Wang, George Karypis, Bing Xiang, Dan Roth 0001
ICLR12
2023 DecAF: Joint Decoding of Answers and Logical Forms for Question Answering over Knowledge Bases
Donghan Yu, Sheng Zhang 0029, Patrick Ng, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Yiqun Hu, William Yang Wang, Zhiguo Wang 0006, Bing Xiang
ICLR10
2023 CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion
abstract
Code completion models have made significant progress in recent years, yet current popular evaluation datasets, such as HumanEval and MBPP, predominantly focus on code completion tasks within a single file. This over-simplified setting falls short of representing the real-world software development scenario where repositories span multiple files with numerous cross-file dependencies, and accessing and understanding cross-file context is often required to complete the code correctly. To fill in this gap, we propose CrossCodeEval, a diverse and multilingual code completion benchmark that necessitates an in-depth cross-file contextual understanding to complete the code accurately. CrossCodeEval is built on a diverse set of real-world, open-sourced, permissively-licensed repositories in four popular programming languages: Python, Java, TypeScript, and C#. To create examples that strictly require cross-file context for accurate completion, we propose a straightforward yet efficient static-analysis-based approach to pinpoint the use of cross-file context within the current file. Extensive experiments on state-of-the-art code language models like CodeGen and StarCoder demonstrate that CrossCodeEval is extremely challenging when the relevant cross-file context is absent, and we see clear improvements when adding these context into the prompt. However, despite such improvements, the pinnacle of performance remains notably unattained even with the highest-performing model, indicating that CrossCodeEval is also capable of assessing model's capability in leveraging extensive context to make better code completion. Finally, we benchmarked various methods in retrieving cross-file context, and show that CrossCodeEval can also be used to measure the capability of code retrievers.
Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Hantian Ding, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang
NeurIPS11
2023 Towards Greener Yet Powerful Code Generation via Quantization: An Empirical Study
abstract
ML-powered code generation aims to assist developers to write code in a more productive manner by intelligently generating code blocks based on natural language prompts. Recently, large pretrained deep learning models have pushed the boundary of code generation and achieved impressive performance. However, the huge number of model parameters poses a significant challenge to their adoption in a typical software development environment, where a developer might use a standard laptop or mid-size server to develop code. Such large models cost significant resources in terms of memory, latency, dollars, as well as carbon footprint.
Xiaokai Wei, Sujan K. Gonugondla, Shiqi Wang 0002, Wasi Uddin Ahmad, Baishakhi Ray, Haifeng Qian, Xiaopeng Li 0002, Zijian Wang 0002, Qing Sun 0013, Ben Athiwaratkun, Mingyue Shang, Murali Krishna Ramanathan, Parminder Bhatia, Bing Xiang
ESEC/SIGSOFT FSE16
2022 Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question Answering
abstract
Question answering over semi-structured tables has attracted significant attention in the NLP community. However, most of the existing work focus on questions that can be answered with short-form answer, i.e. the answer is often a table cell or aggregation of multiple cells. This can mismatch with the intents of users who want to ask more complex questions that require free-form answers such as explanations. To bridge the gap, most recently, pre-trained sequence-to-sequence language models such as T5 are used for generating free-form answers based on the question and table inputs. However, these pre-trained language models have weaker encoding abilities over table cells and schema. To mitigate this issue, in this work, we present an intermediate pre-training framework, Generation-focused Table-based Intermediate Pre-training (GENTAP), that jointly learns representations of natural language questions and tables. GENTAP learns to generate via two training objectives to enhance the question understanding and table representation abilities for complex questions. Based on experimental results, models that leverage GENTAP framework outperform the existing baselines on FETAQA benchmark. The pre-trained models are not only useful for free-form question answering, but also for few-shot data-to-text generation task, thus showing good transfer ability by obtaining new state-of-the-art results.
Peng Shi 0010, Patrick Ng, Feng Nan, Henghui Zhu, Jun Wang 0122, Jiarong Jiang, Alexander Hanbo Li, Rishav Chakravarti, Donald Weidner, Bing Xiang, Zhiguo Wang 0006
AAAI10
2022 Learning Dialogue Representations from Consecutive Utterances
abstract
Zhihan Zhou, Dejiao Zhang, Wei Xiao, Nicholas Dingwall, Xiaofei Ma, Andrew Arnold, Bing Xiang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Zhihan Zhou 0001, Dejiao Zhang, Wei Xiao 0001, Nicholas Dingwall, Xiaofei Ma 0001, Andrew O. Arnold, Bing Xiang
NAACL-HLT7
2021 Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training
abstract
Most recently, there has been significant interest in learning contextual representations for various NLP tasks, by leveraging large scale text corpora to train powerful language models with self-supervised learning objectives, such as Masked Language Model (MLM). Based on a pilot study, we observe three issues of existing general-purpose language models when they are applied in the text-to-SQL semantic parsers: fail to detect the column mentions in the utterances, to infer the column mentions from the cell values, and to compose target SQL queries when they are complex. To mitigate these issues, we present a model pretraining framework, Generation-Augmented Pre-training (GAP), that jointly learns representations of natural language utterance and table schemas, by leveraging generation models to generate high-quality pre-train data. GAP Model is trained on 2 million utterance-schema pairs and 30K utterance-schema-SQL triples, whose utterances are generated by generation models. Based on experimental results, neural semantic parsers that leverage GAP Model as a representation encoder obtain new state-of-the-art results on both Spider and Criteria-to-SQL benchmarks.
Peng Shi 0010, Patrick Ng, Zhiguo Wang 0006, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Cícero Nogueira dos Santos, Bing Xiang
AAAI8
2021 Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip Prediction
abstract
Yifan Gao, Henghui Zhu, Patrick Ng, Cicero Nogueira dos Santos, Zhiguo Wang, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yifan Gao 0001, Henghui Zhu, Patrick Ng, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang
ACL/IJCNLP (1)10
2021 Dual Reader-Parser on Hybrid Textual and Tabular Evidence for Open Domain Question Answering
abstract
Alexander Hanbo Li, Patrick Ng, Peng Xu, Henghui Zhu, Zhiguo Wang, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Alexander Hanbo Li, Patrick Ng, Henghui Zhu, Zhiguo Wang 0006, Bing Xiang
ACL/IJCNLP (1)6
2021 Improving Factual Consistency of Abstractive Summarization via Question Answering
abstract
Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Feng Nan, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathy McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang 0006, Andrew O. Arnold, Bing Xiang
ACL/IJCNLP (1)10
2021 Entity-level Factual Consistency of Abstractive Text Summarization
abstract
Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, Bing Xiang. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Feng Nan, Ramesh Nallapati, Zhiguo Wang 0006, Cícero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathy McKeown, Bing Xiang
EACL8
2021 Retrieval, Re-ranking and Multi-task Learning for Knowledge-Base Question Answering
abstract
Question answering over knowledge bases (KBQA) usually involves three sub-tasks, namely topic entity detection, entity linking and relation detection.Due to the large number of entities and relations inside knowledge bases (KB), previous work usually utilized sophisticated rules to narrow down the search space and managed only a subset of KBs in memory.In this work, we leverage a retrieveand-rerank framework to access KBs via traditional information retrieval (IR) method, and re-rank retrieved candidates with more powerful neural networks such as the pre-trained BERT model.Considering the fact that directly assigning a different BERT model for each sub-task may incur prohibitive costs, we propose to share a BERT encoder across all three sub-tasks and define task-specific layers on top of the shared layer.The unified model is then trained under a multi-task learning framework.Experiments show that: (1) Our IRbased retrieval method is able to collect highquality candidates efficiently, thus enables our method adapt to large-scale KBs easily; (2) the BERT model improves the accuracy across all three sub-tasks; and (3) benefiting from multitask learning, the unified model obtains further improvements with only 1/3 of the original parameters.Our final model achieves competitive results on the SimpleQuestions dataset and superior performance on the FreebaseQA dataset.
Zhiguo Wang 0006, Patrick Ng, Ramesh Nallapati, Bing Xiang
EACL4
2021 Generative Context Pair Selection for Multi-hop Question Answering
abstract
Dheeru Dua, Cicero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner, Sameer Singh. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Dheeru Dua, Cícero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner 0001, Sameer Singh 0001
EMNLP (1)5
2021 Pairwise Supervised Contrastive Learning of Sentence Representations
abstract
Many recent successes in sentence representation learning have been achieved by simply fine-tuning on the Natural Language Inference (NLI) datasets with triplet loss or siamese loss.Nevertheless, they share a common weakness: sentences in a contradiction pair are not necessarily from different semantic categories.Therefore, optimizing the semantic entailment and contradiction reasoning objective alone is inadequate to capture the high-level semantic structure.The drawback is compounded by the fact that the vanilla siamese or triplet losses only learn from individual sentence pairs or triplets, which often suffer from bad local optima.In this paper, we propose PairSupCon, an instance discrimination based approach aiming to bridge semantic entailment and contradiction understanding with high-level categorical concept encoding.We evaluate PairSupCon on various downstream tasks that involve understanding sentence semantics at different granularities.We outperform the previous state-of-theart method with 10%-13% averaged improvement on eight clustering tasks, and 5%-6% averaged improvement on seven semantic textual similarity (STS) tasks.
Dejiao Zhang, Shang-Wen Li 0001, Wei Xiao 0001, Henghui Zhu, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang
EMNLP (1)7
2021 Structured Prediction as Translation between Augmented Natural Languages
Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Alessandro Achille, Rishita Anubhai, Cícero Nogueira dos Santos, Bing Xiang, Stefano Soatto
ICLR8
2021 Supporting Clustering with Contrastive Learning
abstract
Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li, Henghui Zhu, Kathleen McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li 0001, Henghui Zhu, Kathy McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang
NAACL-HLT9
2021 Mixed-Curvature Multi-Relational Graph Neural Network for Knowledge Graph Completion
abstract
Knowledge graphs (KGs) have gradually become valuable assets for many AI applications. In a KG, a node denotes an entity, and an edge (or link) denotes a relationship between the entities represented by the nodes. Knowledge graph completion infers and predicts missing edges in a KG automatically. Knowledge graph embeddings have shed light on addressing this task. Recent research embeds KGs in hyperbolic (negatively curved) space instead of conventional Euclidean (zero curved) space and is effective in capturing hierarchical structures. However, as multi-relational graphs, KGs are not structured uniformly and display intrinsic heterogeneous structures. They usually contain rich types of structures, such as hierarchical and cyclic typed structures. Embedding KGs in single-curvature space, such as Euclidean or hyperbolic space, overlooks the intrinsic heterogeneous structures of KGs, and therefore cannot accurately capture their structures. To address this issue, we propose Mixed-Curvature Multi-Relational Graph Neural Network (M2GNN), a generic approach that embeds multi-relational KGs in a mixed-curvature space for knowledge graph completion. Specifically, we define and construct a mixed-curvature space through a product manifold combining multiple single-curvature spaces (e.g., spherical, hyperbolic, or Euclidean) with the purpose of modeling a variety of structures. However, constructing a mixed-curvature space typically requires manually defining the fixed curvatures, which needs domain knowledge and additional data analysis. Improperly defined curvature space also cannot capture the structures of KGs accurately. To address this problem, we set mixed-curvatures as trainable parameters to better capture the underlying structures of the KGs. Furthermore, we propose a Graph Neural Updater by leveraging the heterogeneous relational context in mixed-curvature space to improve the quality of the embedding. Experiments on three KG datasets demonstrate that the proposed M2GNN can outperform its single geometry counterpart as well as state-of-the-art embedding methods on the KG completion task.
Shen Wang 0005, Xiaokai Wei, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang, Philip S. Yu, Isabel F. Cruz
WWW7
2020 Who Did They Respond to? Conversation Structure Modeling Using Masked Hierarchical Transformer
abstract
Conversation structure is useful for both understanding the nature of conversation dynamics and for providing features for many downstream applications such as summarization of conversations. In this work, we define the problem of conversation structure modeling as identifying the parent utterance(s) to which each utterance in the conversation responds to. Previous work usually took a pair of utterances to decide whether one utterance is the parent of the other. We believe the entire ancestral history is a very important information source to make accurate prediction. Therefore, we design a novel masking mechanism to guide the ancestor flow, and leverage the transformer model to aggregate all ancestors to predict parent utterances. Our experiments are performed on the Reddit dataset (Zhang, Culbertson, and Paritosh 2017) and the Ubuntu IRC dataset (Kummerfeld et al. 2019). In addition, we also report experiments on a new larger corpus from the Reddit platform and release this dataset. We show that the proposed model, that takes into account the ancestral history of the conversation, significantly outperforms several strong baselines including the BERT model on all datasets.
Henghui Zhu, Feng Nan, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang
AAAI5
2020 Template-Based Question Generation from Retrieved Sentences for Improved Unsupervised Question Answering
abstract
Question Answering (QA) is in increasing demand as the amount of information available online and the desire for quick access to this content grows.A common approach to QA has been to fine-tune a pretrained language model on a task-specific labeled dataset.This paradigm, however, relies on scarce, and costly to obtain, large-scale human-labeled data.We propose an unsupervised approach to training QA models with generated pseudotraining data.We show that generating questions for QA training by applying a simple template on a related, retrieved sentence rather than the original context sentence improves downstream QA performance by allowing the model to learn more complex context-question relationships.Training a QA model on this data gives a relative improvement over a previous unsupervised model in F1 score on the SQuAD dataset by about 14%, and 20% when the answer is a named entity, achieving stateof-the-art performance on SQuAD for unsupervised QA.
Alexander R. Fabbri, Patrick Ng, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang
ACL5
2020 Augmented Natural Language for Generative Sequence Labeling
abstract
We propose a generative framework for joint sequence labeling and sentence-level classification.Our model performs multiple sequence labeling tasks at once using a single, shared natural language output space.Unlike prior discriminative methods, our model naturally incorporates label semantics and shares knowledge across tasks.Our framework is general purpose, performing well on fewshot, low-resource, and high-resource tasks.We demonstrate these advantages on popular named entity recognition, slot labeling, and intent classification benchmarks.We set a new state-of-the-art for few-shot slot labeling, improving substantially upon the previous 5-shot (75.0%!90.9%) and 1-shot (70.4% !81.0%) state-of-the-art results.Furthermore, our model generates large improvements (46.27% !63.83%) in low-resource slot labeling over a BERT baseline by incorporating label semantics.We also maintain competitive results on high-resource tasks, performing within two points of the state-of-theart on all tasks and setting a new state-of-theart on the SNIPS dataset.
Ben Athiwaratkun, Cícero Nogueira dos Santos, Jason Krone, Bing Xiang
EMNLP (1)4
2020 Beyond [CLS] through Ranking by Generation
abstract
Generative models for Information Retrieval, where ranking of documents is viewed as the task of generating a query from a document's language model, were very successful in various IR tasks in the past.However, with the advent of modern deep neural networks, attention has shifted to discriminative ranking functions that model the semantic similarity of documents and queries instead.Recently, deep generative models such as GPT2 and BART have been shown to be excellent text generators, but their effectiveness as rankers have not been demonstrated yet.In this work, we revisit the generative framework for information retrieval and show that our generative approaches are as effective as state-of-the-art semantic similarity-based discriminative models for the answer selection task.Additionally, we demonstrate the effectiveness of unlikelihood losses for IR.
Cícero Nogueira dos Santos, Xiaofei Ma 0001, Ramesh Nallapati, Zhiheng Huang, Bing Xiang
EMNLP (1)5
2020 End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering Systems
abstract
Siamak Shakeri, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang, Ramesh Nallapati, Bing Xiang. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Siamak Shakeri, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang
EMNLP (1)8
2020 H2KGAT: Hierarchical Hyperbolic Knowledge Graph Attention Network
Shen Wang 0005, Xiaokai Wei, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang, Philip S. Yu
EMNLP (1)7
2020 Elastic Machine Learning Algorithms in Amazon SageMaker
abstract
There is a large body of research on scalable machine learning (ML). Nevertheless, training ML models on large, continuously evolving datasets is still a difficult and costly undertaking for many companies and institutions. We discuss such challenges and derive requirements for an industrial-scale ML platform. Next, we describe the computational model behind Amazon SageMaker, which is designed to meet such challenges. SageMaker is an ML platform provided as part of Amazon Web Services (AWS), and supports incremental training, resumable and elastic learning as well as automatic hyperparameter optimization. We detail how to adapt several popular ML algorithms to its computational model. Finally, we present an experimental evaluation on large datasets, comparing SageMaker to several scalable, JVM-based implementations of ML algorithms, which we significantly outperform with regard to computation time and cost.
Edo Liberty, Zohar S. Karnin, Bing Xiang, Laurence Rouesnel, Baris Coskun, Ramesh Nallapati, Julio Delgado, Amir Sadoughi, Yury Astashonok, Piali Das, Can Balioglu, Saswata Chakravarty, Madhav Jha, Philip Gautier, David Arpin, Tim Januschowski, Valentin Flunkert, Yuyang Wang 0001, Jan Gasthaus, Lorenzo Stella, Syama Sundar Rangapuram, David Salinas, Sebastian Schelter, Alexander J. Smola
SIGMOD Conference3
2019 Topic Modeling with Wasserstein Autoencoders
abstract
We propose a novel neural topic model in the Wasserstein autoencoders (WAE) framework.Unlike existing variational autoencoder based models, we directly enforce Dirichlet prior on the latent document-topic vectors.We exploit the structure of the latent space and apply a suitable kernel in minimizing the Maximum Mean Discrepancy (MMD) to perform distribution matching.We discover that MMD performs much better than the Generative Adversarial Network (GAN) in matching high dimensional Dirichlet distribution.We further discover that incorporating randomness in the encoder output during training leads to significantly more coherent topics.To measure the diversity of the produced topics, we propose a simple topic uniqueness metric.Together with the widely used coherence measure NPMI, we offer a more wholistic evaluation of topic quality.Experiments on several real datasets show that our model produces significantly better topics than existing topic models.
Feng Nan, Ramesh Nallapati, Bing Xiang
ACL (1)4
2019 OCGAN: One-Class Novelty Detection Using GANs With Constrained Latent Representations
abstract
We present a novel model called OCGAN for the classical problem of one-class novelty detection, where, given a set of examples from a particular class, the goal is to determine if a query example is from the same class. Our solution is based on learning latent representations of in-class examples using a de-noising auto-encoder network. The key contribution of our work is our proposal to explicitly constrain the latent space to exclusively represent the given class. In order to accomplish this goal, firstly, we force the latent space to have bounded support by introducing a tanh activation in the encoder's output layer. Secondly, using a discriminator in the latent space that is trained adversarially, we ensure that encoded representations of in-class examples resemble uniform random samples drawn from the same bounded space. Thirdly, using a second adversarial discriminator in the input space, we ensure all randomly drawn latent samples generate examples that look real. Finally, we introduce a gradient-descent based sampling technique that explores points in the latent space that generate potential out-of-class examples, which are fed back to the network to further train it to generate in-class examples from those points. The effectiveness of the proposed method is measured across four publicly available datasets using two one-class novelty detection protocols where we achieve state-of-the-art results.
Pramuditha Perera, Ramesh Nallapati, Bing Xiang
CVPR3
2019 Multi-passage BERT: A Globally Normalized BERT Model for Open-domain Question Answering
abstract
Zhiguo Wang, Patrick Ng, Xiaofei Ma, Ramesh Nallapati, Bing Xiang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zhiguo Wang 0006, Patrick Ng, Xiaofei Ma 0001, Ramesh Nallapati, Bing Xiang
EMNLP/IJCNLP (1)5
2018 Coherence-Aware Neural Topic Modeling
abstract
Topic models are evaluated based on their ability to describe documents well (i.e.low perplexity) and to produce topics that carry coherent semantic meaning.In topic modeling so far, perplexity is a direct optimization target.However, topic coherence, owing to its challenging computation, is not optimized for and is only evaluated after training.In this work, under a neural variational inference framework, we propose methods to incorporate a topic coherence objective into the training process.We demonstrate that such a coherenceaware topic model exhibits a similar level of perplexity as baseline models but achieves substantially higher topic coherence.
Ramesh Nallapati, Bing Xiang
EMNLP3
2017 Neural Models for Sequence Chunking
abstract
Many natural language understanding (NLU) tasks, such as shallow parsing (i.e., text chunking) and semantic slot filling, require the assignment of representative labels to the meaningful chunks in a sentence. Most of the current deep neural network (DNN) based methods consider these tasks as a sequence labeling problem, in which a word, rather than a chunk, is treated as the basic unit for labeling. These chunks are then inferred by the standard IOB (Inside-Outside- Beginning) labels. In this paper, we propose an alternative approach by investigating the use of DNN for sequence chunking, and propose three neural models so that each chunk can be treated as a complete unit for labeling. Experimental results show that the proposed neural sequence chunking models can achieve start-of-the-art performance on both the text chunking and slot filling tasks.
Feifei Zhai, Saloni Potdar, Bing Xiang, Bowen Zhou 0006
AAAI3
2017 Improved Neural Relation Detection for Knowledge Base Question Answering
abstract
Relation detection is a core component of many NLP applications including Knowledge Base Question Answering (KBQA).In this paper, we propose a hierarchical recurrent neural network enhanced by residual learning which detects KB relations given an input question.Our method uses deep residual bidirectional LSTMs to compare questions and relation names via different levels of abstraction.Additionally, we propose a simple KBQA system that integrates entity linking and our proposed relation detector to make the two components enhance each other.Our experimental results show that our approach not only achieves outstanding relation detection performance, but more importantly, it helps our KBQA system achieve state-of-the-art accuracy for both single-relation (SimpleQuestions) and multi-relation (WebQSP) QA benchmarks.
Mo Yu, Wenpeng Yin 0001, Kazi Saidul Hasan, Cícero Nogueira dos Santos, Bing Xiang, Bowen Zhou 0002
ACL (1)5
2017 GaDei: On Scale-Up Training as a Service for Deep Learning
abstract
Deep learning (DL) training-as-a-service (TaaS) is an important emerging industrial workload. TaaS must satisfy a wide range of customers who have no experience and/or resources to tune DL hyper-parameters (e.g., mini-batch size and learning rate), and meticulous tuning for each user's dataset is prohibitively expensive. Therefore, TaaS hyper-parameters must be fixed with values that are applicable to all users. Unfortunately, few research papers have studied how to design a system for TaaS workloads. By evaluating the IBM Watson Natural Language Classfier (NLC) workloads, the most popular IBM cognitive service used by thousands of enterprise-level clients globally, we provide empirical evidence that only the conservative hyper-parameter setup (e.g., small mini-batch size) can guarantee acceptable model accuracy for a wide range of customers. Unfortunately, smaller mini-batch size requires higher communication bandwidth in a parameter-server based DL training system. In this paper, we characterize the exceedingly high communication bandwidth requirement of TaaS using representative industrial deep learning workloads. We then present GaDei, a highly optimized shared-memory based scale-up parameter server design. We evaluate GaDei using both commercial benchmarks and public benchmarks and demonstrate that GaDei significantly outperforms the state-of-the-art parameter-server based implementation while maintaining the required accuracy. GaDei achieves near-best-possible runtime performance, constrained only by the hardware limitation. Furthermore, to the best of our knowledge, GaDei is the only scale-up DL system that provides fault-tolerance.
Wei Zhang 0057, Minwei Feng, Yunhui Zheng, Yufei Ren, Yandong Wang 0001, Peng Liu 0010, Bing Xiang, Li Zhang 0002, Bowen Zhou 0002, Fei Wang 0001
ICDM8
2017 A Structured Self-Attentive Sentence Embedding
Zhouhan Lin, Minwei Feng, Cícero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou 0002, Yoshua Bengio
ICLR (Poster)5
2017 Jointly Trained Sequential Labeling and Classification by Sparse Attention Neural Networks
abstract
Sentence-level classification and sequential labeling are two fundamental tasks in language understanding.While these two tasks are usually modeled separately, in reality, they are often correlated, for example in intent classification and slot filling, or in topic classification and named-entity recognition.In order to utilize the potential benefits from their correlations, we propose a jointly trained model for learning the two tasks simultaneously via Long Short-Term Memory (LSTM) networks.This model predicts the sentence-level category and the word-level label sequence from the stepwise output hidden representations of LSTM.We also introduce a novel mechanism of "sparse attention" to weigh words differently based on their semantic relevance to sentence-level classification.The proposed method outperforms baseline models on ATIS and TREC datasets.
Mingbo Ma, Kai Zhao 0003, Liang Huang 0001, Bing Xiang, Bowen Zhou 0006
INTERSPEECH4
2016 Improved Representation Learning for Question Answer Matching
abstract
Passage-level question answer matching is a challenging task since it requires effective representations that capture the complex semantic relations between questions and answers.In this work, we propose a series of deep learning models to address passage answer selection.To match passage answers to questions accommodating their complex semantic relations, unlike most previous work that utilizes a single deep learning structure, we develop hybrid models that process the text using both convolutional and recurrent neural networks, combining the merits on extracting linguistic information from both structures.Additionally, we also develop a simple but effective attention mechanism for the purpose of constructing better answer representations according to the input question, which is imperative for better modeling long answer sequences.The results on two public benchmark datasets, InsuranceQA and TREC-QA, show that our proposed models outperform a variety of strong baselines.
Cícero Nogueira dos Santos, Bing Xiang, Bowen Zhou 0006
ACL (1)3
2016 Distributed Deep Learning for Question Answering
abstract
This paper is an empirical study of the distributed deep learning for question answering subtasks: answer selection and question classification. Comparison studies of SGD, MSGD, ADADELTA, ADAGRAD, ADAM/ADAMAX, RMSPROP, DOWNPOUR and EASGD/EAMSGD algorithms have been presented. Experimental results show that the distributed framework based on the message passing interface can accelerate the convergence speed at a sublinear scale. This paper demonstrates the importance of distributed training. For example, with 48 workers, a 24x speedup is achievable for the answer selection task and running time is decreased from 138.2 hours to 5.81 hours, which will increase the productivity significantly.
Minwei Feng, Bing Xiang, Bowen Zhou 0006
CIKM2
2016 Simple Question Answering by Attentive Convolutional Neural Network
abstract
This work focuses on answering single-relation factoid questions over Freebase. Each question can acquire the answer from a single fact of form (subject, predicate, object) in Freebase. This task, simple question answering (SimpleQA), can be addressed via a two-step pipeline: entity linking and fact selection. In fact selection, we match the subject entity in a fact candidate with the entity mention in the question by a character-level convolutional neural network (char-CNN), and match the predicate in that fact with the question by a word-level CNN (word-CNN). This work makes two main contributions. (i) A simple and effective entity linker over Freebase is proposed. Our entity linker outperforms the state-of-the-art entity linker over SimpleQA task. (ii) A novel attentive maxpooling is stacked over word-CNN, so that the predicate representation can be matched with the predicate-focused question representation more effectively. Experiments show that our system sets new state-of-the-art in this task.
Wenpeng Yin 0001, Mo Yu, Bing Xiang, Bowen Zhou 0002, Hinrich Schütze
COLING3
2016 Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond
abstract
In this work, we model abstractive text summarization using Attentional Encoder-Decoder Recurrent Neural Networks, and show that they achieve state-of-the-art performance on two different corpora.We propose several novel models that address critical problems in summarization that are not adequately modeled by the basic architecture, such as modeling key-words, capturing the hierarchy of sentence-toword structure, and emitting words that are rare or unseen at training time.Our work shows that many of our proposed models contribute to further improvement in performance.We also propose a new dataset consisting of multi-sentence summaries, and establish performance benchmarks for further research.
Ramesh Nallapati, Bowen Zhou 0002, Cícero Nogueira dos Santos, Caglar Gulcehre, Bing Xiang
CoNLL5
2016 Leveraging Sentence-level Information with Encoder LSTM for Semantic Slot Filling
abstract
Recurrent Neural Network (RNN) and one of its specific architectures, Long Short-Term Memory (LSTM), have been widely used for sequence labeling.Explicitly modeling output label dependencies on top of RNN/LSTM is a widely-studied and effective extension.We propose another extension to incorporate the global information spanning over the whole input sequence.The proposed method, encoder-labeler LSTM, first encodes the whole input sequence into a fixed length vector with the encoder LSTM, and then uses this encoded vector as the initial state of another LSTM for sequence labeling.With this method, we can predict the label sequence while taking the whole input sequence information into consideration.In the experiments of a slot filling task, which is an essential component of natural language understanding, with using the standard ATIS corpus, we achieved the state-of-the-art F 1 -score of 95.66%.
Gakuto Kurata, Bing Xiang, Bowen Zhou 0006, Mo Yu
EMNLP2
2016 Labeled Data Generation with Encoder-Decoder LSTM for Semantic Slot Filling
Gakuto Kurata, Bing Xiang, Bowen Zhou 0006
INTERSPEECH2
2016 Improved Neural Network-based Multi-label Classification with Better Initialization Leveraging Label Co-occurrence
Gakuto Kurata, Bing Xiang, Bowen Zhou 0006
HLT-NAACL2
2016 ABCNN: Attention-Based Convolutional Neural Network for Modeling Sentence Pairs
abstract
How to model a pair of sentences is a critical issue in many NLP tasks such as answer selection (AS), paraphrase identification (PI) and textual entailment (TE). Most prior work (i) deals with one individual task by fine-tuning a specific system; (ii) models each sentence’s representation separately, rarely considering the impact of the other sentence; or (iii) relies fully on manually designed, task-specific linguistic features. This work presents a general Attention Based Convolutional Neural Network (ABCNN) for modeling a pair of sentences. We make three contributions. (i) The ABCNN can be applied to a wide variety of tasks that require modeling of sentence pairs. (ii) We propose three attention schemes that integrate mutual influence between sentences into CNNs; thus, the representation of each sentence takes into consideration its counterpart. These interdependent sentence pair representations are more powerful than isolated sentence representations. (iii) ABCNNs achieve state-of-the-art performance on AS, PI and TE tasks. We release code at: https://github.com/yinwenpeng/Answer_Selection .
Wenpeng Yin 0001, Hinrich Schütze, Bing Xiang, Bowen Zhou 0002
Trans. Assoc. Comput. Linguistics3
2015 Classifying Relations by Ranking with Convolutional Neural Networks
abstract
Cícero dos Santos, Bing Xiang, Bowen Zhou. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Cícero Nogueira dos Santos, Bing Xiang, Bowen Zhou 0006
ACL (1)2
2015 Applying deep learning to answer selection: A study and an open task
abstract
We apply a general deep learning framework to address the non-factoid question answering task. Our approach does not rely on any linguistic tools and can be applied to different languages or domains. Various architectures are presented and compared. We create and release a QA corpus and setup a new QA task in the insurance domain. Experimental results demonstrate superior performance compared to the baseline methods and various technologies give further improvements. For this highly challenging task, the top-1 accuracy can reach up to 65.3% on a test set, which indicates a great potential for practical use.
Minwei Feng, Bing Xiang, Michael R. Glass, Bowen Zhou 0006
ASRU2
2015 Efficient Hyper-parameter Optimization for NLP Applications
abstract
Hyper-parameter optimization is an important problem in natural language processing (NLP) and machine learning.Recently, a group of studies has focused on using sequential Bayesian Optimization to solve this problem, which aims to reduce the number of iterations and trials required during the optimization process.In this paper, we explore this problem from a different angle, and propose a multi-stage hyper-parameter optimization that breaks the problem into multiple stages with increasingly amounts of data.Early stage provides fast estimates of good candidates which are used to initialize later stages for better performance and speed.We demonstrate the utility of this new algorithm by evaluating its speed and accuracy against state-of-the-art Bayesian Optimization algorithms on classification and prediction tasks.
Minwei Feng, Bowen Zhou 0006, Bing Xiang, Sridhar Mahadevan
EMNLP4
2013 Two-Neighbor Orientation Model with Cross-Boundary Global Contexts
Hendra Setiawan, Bowen Zhou 0006, Bing Xiang, Libin Shen
ACL (1)3
2013 Enlisting the Ghost: Modeling Empty Categories for Machine Translation
Bing Xiang, Xiaoqiang Luo, Bowen Zhou 0006
ACL (1)1
2013 Anchor Graph: Global Reordering Contexts for Statistical Machine Translation
abstract
Reordering poses one of the greatest challenges in Statistical Machine Translation research as the key contextual information may well be beyond the confine of translation units.We present the "Anchor Graph" (AG) model where we use a graph structure to model global contextual information that is crucial for reordering.The key ingredient of our AG model is the edges that capture the relationship between the reordering around a set of selected translation units, which we refer to as anchors.As the edges link anchors that may span multiple translation units at decoding time, our AG model effectively encodes global contextual information that is previously absent.We integrate our proposed model into a state-of-the-art translation system and demonstrate the efficacy of our proposal in a largescale Chinese-to-English translation task.* This work was done when the authors were with IBM. 1 We define translation units as phrases in phrase-based SMT or as translation rules in syntax-based SMT.
Hendra Setiawan, Bowen Zhou 0006, Bing Xiang
EMNLP3
2013 The IBM speech-to-speech translation system for smartphone: Improvements for resource-constrained tasks
Bowen Zhou 0002, Songfang Huang, Martin Cmejrek, Wei Zhang 0057, Jia Cui, Bing Xiang, Gregg Daggett, Upendra V. Chaudhari, Sameer Maskey, Etienne Marcheret
Comput. Speech Lang.8
2011 A Correction Model for Word Alignments
J. Scott McCarley, Abraham Ittycheriah, Salim Roukos, Bing Xiang, Jian-Ming Xu
EMNLP4
2010 Feature-Rich Discriminative Phrase Rescoring for SMT
Fei Huang 0002, Bing Xiang
COLING2
2009 Towards integrated machine translation using structural alignment from syntax-augmented synchronous parsing
abstract
In current statistical machine translation, IBM model based word alignment is widely used as a starting point to build phrase-based machine translation systems. However, such alignment model is separated from the rest of machine translation pipeline and optimized independently. Furthermore, structural information is not taken into account in the alignment model, which sometimes leads to incorrect alignments. In this paper, we present a novel method to connect a re-alignment model with a translation model in an integrated framework. We conduct bilingual chart parsing based on syntax-augmented synchronous context-free grammar. A Viterbi derivation tree is generated for each sentence pair with multiple features employed in a log-linear model. A new word alignment is created under the structural constraint from the Viterbi tree. Extensive experiments are conducted in a Farsi-to-English translation task in conversational speech domain and also a German-to-English translation task in text domain. Systems trained on the new alignment provide significant higher BLEU scores compared to a state-of-the-art baseline.
Bing Xiang, Bowen Zhou 0006, Martin Cmejrek
ASRU1
2009 Advances in syntax-based Malay-English speech translation
abstract
In this paper, we present advanced techniques that improved the performance of IBM Malay-English speech translation system significantly. During this work, we generated linguistics-driven hierarchical rules to enhance the formal syntax-based translation model; designed an active learning approach with bi-directional translations that outperformed unsupervised training; utilized translation direction information in parallel training corpus to build direction-specific interpolated language models for machine translation. There is 20% relative improvement achieved in the translation performance through all these techniques. A state-of-the-art Malay speech recognition system was also established as one of the crucial modules in the rapidly developed Malay-English speech translation.
Bing Xiang, Bowen Zhou 0006, Martin Cmejrek
ICASSP1
2009 A study of bootstrapping with multiple acoustic features for improved automatic speech recognition
Bing Xiang, Bowen Zhou 0006
INTERSPEECH3
2008 Developing high performance asr in the IBM multilingual speech-to-speech translation system
abstract
This paper presents our recent development of the real-time speech recognition component in the IBM English/Iraqi Arabic speech-to-speech translation system for the DARPA Transtac project. We describe the details of the acoustic and language modeling that lead to high recognition accuracy and noise robustness and give the performance of the system on the evaluation sets of spontaneous conversational speech. We also introduce the streaming decoding structure and several speedup techniques that achieves best recognition accuracy at about 0.3×RT recognition speed.
Liang Gu, Bing Xiang, Wei Zhang 0022
ICASSP3
2008 Unsupervised training for farsi-english speech-to-speech translation
abstract
Speech-to-speech translation has evolved into an attractive area in recent years with significant progress made by various research groups. However, the translation engines usually suffer from the lack of bilingual training data, especially for low-resource languages. In this paper we present an unsupervised training technique to alleviate this problem by taking advantage of available source language data. Different approaches are proposed and compared through extensive experiments conducted on a speech-to-speech translation task between Farsi and English. The translation performance is significantly improved in both directions with the enhanced translation model. A state-of-the-art Farsi automatic speech recognition system is also established in this work.
Bing Xiang, Yonggang Deng
ICASSP1
2007 Integrating Speech Recognition and Machine Translation
abstract
This paper presents a set of experiments that we conducted in order to optimize the performance of an Arabic/English machine translation system on broadcast news and conversational speech data. Proper integration of speech-to-text (STT) and machine translation (MT) requires special attention to issues such as sentence boundary detection, punctuation, STT accuracy, tokenization, conversion of spoken numbers and dates to written form, optimization of MT decoding weights, and scoring. We discuss these issues, and show that a carefully tuned STT/MT integration can lead to significant translation accuracy improvements compared to simply feeding the regular STT output to a text MT system.
Spyridon Matsoukas, Ivan Bulyko, Bing Xiang, Kham Nguyen, Richard M. Schwartz, John Makhoul
ICASSP (4)3
2007 Combining Outputs from Multiple Machine Translation Systems
Antti-Veikko I. Rosti, Necip Fazil Ayan, Bing Xiang, Spyridon Matsoukas, Richard M. Schwartz, Bonnie J. Dorr
HLT-NAACL3
2006 Morphological Decomposition for Arabic Broadcast News Transcription
abstract
In this paper, we present a novel approach for morphological de-composition in large vocabulary Arabic speech recognition. It achieved low out-of-vocabulary (OOV) rate as well as high recognition accuracy in a state-of-the-art Arabic broadcast news transcription system. In this approach, the compound words are decomposed into stems and affixes in both language training and acoustic training data. The decomposed words in the recognition output are re-joined before scoring. Four algorithms are experimented and compared in this work. The best system achieved 1.9% absolute reduction (9.8% relative) in word error rate (WER) when compared to the 64K-word baseline. The recognition performance of this system is also comparable to a 300K-word recognition system trained on the normal words. In the meantime, the decomposed system is much faster in terms of speed and also needs less memory than the systems with larger than 64K vocabularies.
Bing Xiang, Kham Nguyen, Long Nguyen 0001, Richard M. Schwartz, John Makhoul
ICASSP (1)1
2006 Advances in transcription of broadcast news and conversational telephone speech within the combined EARS BBN/LIMSI system
abstract
This paper describes the progress made in the transcription of broadcast news (BN) and conversational telephone speech (CTS) within the combined BBN/LIMSI system from May 2002 to September 2004. During that period, BBN and LIMSI collaborated in an effort to produce significant reductions in the word error rate (WER), as directed by the aggressive goals of the Effective, Affordable, Reusable, Speech-to-text [Defense Advanced Research Projects Agency (DARPA) EARS] program. The paper focuses on general modeling techniques that led to recognition accuracy improvements, as well as engineering approaches that enabled efficient use of large amounts of training data and fast decoding architectures. Special attention is given on efforts to integrate components of the BBN and LIMSI systems, discussing the tradeoff between speed and accuracy for various system combination strategies. Results on the EARS progress test sets show that the combined BBN/LIMSI system achieved relative reductions of 47% and 51% on the BN and CTS domains, respectively.
Spyridon Matsoukas, Jean-Luc Gauvain, Gilles Adda, Thomas Colthurst, Chia-Lin Kao, Owen Kimball, Lori Lamel, Fabrice Lefèvre, Jeff Z. Ma, John Makhoul, Long Nguyen 0001, Rohit Prasad, Richard M. Schwartz, Holger Schwenk, Bing Xiang
IEEE Trans. Speech Audio Process.15
2005 Cluster-Dependent Acoustic Modeling
abstract
In this paper, we present cluster-dependent acoustic modeling for large-vocabulary speech recognition. With large amounts of acoustic training data, we build multiple cluster-dependent models (CDM), each focusing on a group of speakers in order to represent speaker-dependent characteristics. It is motivated by the fact that a sufficiently trained speaker-dependent (SD) model is better than the speaker-independent (SI) model. During decoding, we decode the data of each test speaker using CDMs selected under certain criteria to achieve high recognition accuracy. Various speaker clustering and model selection techniques are proposed and compared in the task of broadcast news (BN) transcription. The CDM provided more than 1% absolute gain in unadapted decoding and 0.5% gain in adapted decoding when compared to our baseline system on the EARS BN 2003 development test set.
Bing Xiang, Long Nguyen 0001, Spyridon Matsoukas, Richard M. Schwartz
ICASSP (1)1
2005 Recent progress in Arabic broadcast news transcription at BBN
Mohamed Afify, Long Nguyen 0001, Bing Xiang, Sherif M. Abdou, John Makhoul
INTERSPEECH3
2005 The effects of speech recognition and punctuation on information extraction performance
John Makhoul, Alex Baron, Ivan Bulyko, Long Nguyen 0001, Lance A. Ramshaw, David Stallard, Richard M. Schwartz, Bing Xiang
INTERSPEECH8
2005 The 2004 BBN 1xRT recognition systems for English broadcast news and conversational telephone speech
abstract
This paper describes the BBN real-time recognition systems used in the 2004 Rich Transcription (RT) benchmark test for the English Conversational Telephone Speech (CTS) and Broadcast News (BN) tasks. We describe the system architecture, along withthe algorithms weused inorder to reduce computation with minimal impact on recognition accuracy. Particular choices in the design of thefinal system are analyzed toshow the trade-offs between speed and accuracy. We also present recently developed new architecture for the real-time systems, which outperforms the systems we submitted for the RT04 benchmark tests for both domains.
Spyridon Matsoukas, Rohit Prasad, Srinivas Laxminarayan, Bing Xiang, Long Nguyen 0001, Richard M. Schwartz
INTERSPEECH4
2005 The BBN RT04 English broadcast news transcription system
abstract
This paper describes the BBN English Broadcast News transcription system developed for the EARS Rich Transcription 2004 (RT04) evaluation. In comparison to the BBN RT03 system, we achieved around 22% relative reduction in word error rate for all EARS BN development test sets. The use of additional acoustic training data acquired through Light Supervision based on thousands of hours of found data made the biggest contribution to the improvement. Better audio segmentation, through the use of an online speaker clustering algorithm and chopping speaker turns into moderately long utterances, also contributed substantially to the improvement. Other contributions, even of modest size but adding up nicely, include using discriminative training for all acoustic models, using word duration as an additional knowledge source during N-best rescoring, and using updated lexicon and language models.
Long Nguyen 0001, Bing Xiang, Mohamed Afify, Sherif M. Abdou, Spyridon Matsoukas, Richard M. Schwartz, John Makhoul
INTERSPEECH2
2005 The BBN Mandarin broadcast news transcription system
Bing Xiang, Long Nguyen 0001, Xuefeng Guo, Dongxin Xu
INTERSPEECH1
2004 Light supervision in acoustic model training
abstract
We present a new light supervision method to derive additional acoustic training data automatically for broadcast news transcription systems. A subset of the TDT corpus, which consists of broadcast audio with corresponding closed-caption (CC) transcripts, is identified by aligning the CC transcripts and the hypotheses generated by lightly-supervised decoding. Phrases of three or more contiguous words, on which both the CC transcripts and the decoder's hypotheses agree, are selected. The selection yields 702 hours, or 72% of the captioned data. When adding 700 hours of selected data to the baseline 141 hour broadcast news training data set, we achieved a 13% relative word error rate reduction. The key to the effectiveness of this light supervision method is the use of a biased language model (LM) in the lightly supervised decoding. The biased LM, in which the CC transcripts are added with heavy weighting, helps in selecting words the recognizer could have misrecognized if using a fair LM.
Long Nguyen 0001, Bing Xiang
ICASSP (1)2
2004 Speech recognition in multiple languages and domains: the 2003 BBN/LIMSI EARS system
abstract
We report on the results of the first evaluations for the BBN/LIMSI system under the new DARPA EARS program. The evaluations were carried out for conversational telephone speech (CTS) and broadcast news (BN) for three languages: English, Mandarin, and Arabic. In addition to providing system descriptions and evaluation results, the paper highlights methods that worked well across the two domains and those few that worked well on one domain but not the other. For the BN evaluations, which had to be run under 10 times real-time, we demonstrated that a joint BBN/LIMSI system with a time constraint achieved better results than either system alone.
Richard M. Schwartz, Thomas Colthurst, Nicolae Duta, Herbert Gish, Rukmini Iyer, Chia-Lin Kao, Daben Liu, Owen Kimball, Jeff Z. Ma, John Makhoul, Spyridon Matsoukas, Long Nguyen 0001, Mohammed Noamany, Rohit Prasad, Bing Xiang, Dongxin Xu, Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Langzhou Chen
ICASSP (3)15
2003 Using prosodic and conversational features for high-performance speaker recognition: report from JHU WS'02
abstract
While there has been a long tradition of research seeking to use prosodic features, especially pitch, in speaker recognition systems, results have generally been disappointing when such features are used in isolation and only modest improvements have been seen when used in conjunction with traditional cepstral GMM systems. In contrast, we report here on work from the JHU 2002 Summer Workshop exploring a range of prosodic features, using as testbed the 2001 NIST Extended Data task. We examined a variety of modeling techniques, such as n-gram models of turn-level prosodic features and simple vectors of summary statistics per conversation side scored by k/sup th/ nearest-neighbor classifiers. We found that purely prosodic models were able to achieve equal error rates of under 10%, and yielded significant gains when combined with more traditional systems. We also report on exploratory work on "conversational" features, capturing properties of the interaction across conversation sides, such as turn-taking patterns.
Barbara Peskin, Jirí Navrátil 0001, Joy S. Abramson, Douglas A. Jones, David Klusácek, Douglas A. Reynolds, Bing Xiang
ICASSP (4)7
2003 The SuperSID project: exploiting high-level information for high-accuracy speaker recognition
abstract
The area of automatic speaker recognition has been dominated by systems using only short-term, low-level acoustic information, such as cepstral features. While these systems have indeed produced very low error rates, they ignore other levels of information beyond low-level acoustics that convey speaker information. Recently published work has shown examples that such high-level information can be used successfully in automatic speaker recognition systems and has the potential to improve accuracy and add robustness. For the 2002 JHU CLSP summer workshop, the SuperSID project (http://www.clsp.jhu.edu/ws2002/groups/supersid/) was undertaken to exploit these high-level information sources and dramatically increase speaker recognition accuracy on a defined NIST evaluation corpus and task. The paper provides an overview of the structure, data, task, tools, and accomplishments of this project. Wide ranging approaches using pronunciation models, prosodic dynamics, pitch and duration features, phone streams, and conversational interactions were explored and developed. We show how these novel features and classifiers indeed provide complementary information and can be fused together to drive down the equal error rate on the 2001 NIST extended data task to 0.2% - a 71% relative reduction in error over the previous state of the art.
Douglas A. Reynolds, Walter D. Andrews, Joseph P. Campbell, Jirí Navrátil 0001, Barbara Peskin, André Adami, Qin Jin, David Klusácek, Joy S. Abramson, Radu Mihaescu, John J. Godfrey, Douglas A. Jones, Bing Xiang
ICASSP (4)13
2003 Text-independent speaker verification with dynamic trajectory model
abstract
A novel approach is presented for text-independent speaker verification. Based on the frequencies of occurrence of Gaussian component strings in the quantized acoustic trajectories, a universal background trajectory model and multiple target trajectory models are created for the background and target speakers separately. Analysis of the speaker entropy in the trajectory space demonstrates that the segmental dynamic trajectory catches speaker-specific information. Experiments are conducted on the telephony speech used in the NIST 1999 speaker verification evaluation and show that the bicomponent strings achieve good performance and may provide complementary information to a general speaker verification system.
Bing Xiang
IEEE Signal Process. Lett.1
2003 Efficient text-independent speaker verification with structural Gaussian mixture models and neural network
abstract
We present an integrated system with structural Gaussian mixture models (SGMMs) and a neural network for purposes of achieving both computational efficiency and high accuracy in text-independent speaker verification. A structural background model (SBM) is constructed first by hierarchically clustering all Gaussian mixture components in a universal background model (UBM). In this way the acoustic space is partitioned into multiple regions in different levels of resolution. For each target speaker, a SGMM can be generated through multilevel maximum a posteriori (MAP) adaptation from the SBM. During test, only a small subset of Gaussian mixture components are scored for each feature vector in order to reduce the computational cost significantly. Furthermore, the scores obtained in different layers of the tree-structured models are combined via a neural network for final decision. Different configurations are compared in the experiments conducted on the telephony speech data used in the NIST speaker verification evaluation. The experimental results show that computational reduction by a factor of 17 can be achieved with 5% relative reduction in equal error rate (EER) compared with the baseline. The SGMM-SBM also shows some advantages over the recently proposed hash GMM, including higher speed and better verification performance.
Bing Xiang, Toby Berger
IEEE Trans. Speech Audio Process.1
2002 Short-time Gaussianization for robust speaker verification
abstract
In this paper, a novel approach for robust speaker verification, namely short-time Gaussianization, is proposed. Short-time Gaussianization is initiated by a global linear transformation of the features, followed by a short-time windowed cumulative distribution function (CDF) matching. First, the linear transformation in the feature space leads to local independence or decorrelation. Then the CDF matching is applied to segments of speech localized in time and tries to warp a given feature so that its CDF matches normal distribution. It is shown that one of the recent techniques used for speaker recognition, feature warping [l] can be formulated within the framework of Gaussianization. Compared to the baseline system with cepstral mean subtraction (CMS), around 20% relative improvement in both equal error rate(EER) and minimum detection cost function (DCF) is obtained on NIST 2001 cellular phone data evaluation.
Bing Xiang, Upendra V. Chaudhari, Jirí Navrátil 0001, Ganesh N. Ramaswamy, Ramesh A. Gopinath
ICASSP1
2002 Speaker verification using Gaussian component strings in dynamic trajectory space
Bing Xiang
INTERSPEECH1
2002 Structural Gaussian mixture models for efficient text-independent speaker verification
Bing Xiang, Toby Berger
INTERSPEECH1
2001 Multiple mixture segmental HMM and its applications
abstract
A multiple mixture segmental hidden Markov model (MMSHMM) is presented. This model is extended from the linear probabilistic-trajectory segmental HMM. Each segment is characterized by a linear trajectory with slope and mid-point parameters, and also the residual error covariances around the trajectory, so that both extra-segmental and intra-segmental variation are represented. Instead of modeling single distribution for each model parameter as earlier work, we use multiple mixture components for model parameters to represent the variability due to the variation within each speaker and also the differences between speakers. This model is evaluated on two applications. One is a phonetic classification task with TIMIT corpus, which shows that MMSHMM has advantages over conventional HMM. Another one is a speaker-independent keyword spotting task with the Road Rally database. By rescoring putative events hypothesized by a primary HMM keyword spotter, the experiments show that the performance is improved through distinguishing true hits from false alarms.
Bing Xiang, Toby Berger
ICASSP1
2000 Multistage coarticulation model combining articulatory, formant and cepstral features
Raimo Bakis, Jing Huang 0019, Bing Xiang
INTERSPEECH4
1999 Auditory model based speech feature extraction and its application to speaker identification
abstract
According to the characteristics of the auditory periphery and cochlear nucleus, as well as attempting to simulate the mechanism of auditory system as a whole, two kinds of novel speech feature are presented in this paper, and a framework of neural network has been adopted. The two features considered are: the weighted average localized synchronized rate cepstrum, and the weighted firing rate cepstrum. Both of them are applied to speaker identification. The modular tree and modified linear opinion pools are used as classifiers to simulate the parallel processing mechanism of the upper level function of auditory system. Good recognition accuracy is obtained under both clean and noisy environments.
Bing Xiang, Xihong Wu, Huisheng Chi
IJCNN1