Bo-Hsiang Tseng

dblp:164/5802 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Question answering and dialogue systems · 44% Planning, search and constraint satisfaction · 20% Language models and text generation · 17%

Topics — the 15 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
1.642021
Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking · EMNLP (1) 2021
A Generative Model for Joint Natural Language Understanding and Generation · ACL 2020
Semi-Supervised Bootstrapping of Dialogue State Trackers for Task-Oriented Modelling · EMNLP/IJCNLP (1) 2019
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
1.032021
Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking · EMNLP (1) 2021
Semi-Supervised Bootstrapping of Dialogue State Trackers for Task-Oriented Modelling · EMNLP/IJCNLP (1) 2019
MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling · EMNLP 2018
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution · ACL (1) 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
planning evaluation
0.912025
ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
user simulation
0.512021
Transferable Dialogue Systems and User Simulators · ACL/IJCNLP (1) 2021
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.412020
A Generative Model for Joint Natural Language Understanding and Generation · ACL 2020
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.422019
Personalizing Recurrent-Neural-Network-Based Language Model by Social Network · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Machine Comprehension of Spoken Content: TOEFL Listening Test and Spoken SQuAD · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.412019
Machine Comprehension of Spoken Content: TOEFL Listening Test and Spoken SQuAD · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Machine learning › Learning paradigms
semi-supervised learning
0.412019
Semi-Supervised Bootstrapping of Dialogue State Trackers for Task-Oriented Modelling · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › natural language understanding › question answering
spoken question answering
0.412019
Machine Comprehension of Spoken Content: TOEFL Listening Test and Spoken SQuAD · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Natural language and speech › Question answering and dialogue systems
dialogue dataset
0.312018
MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling · EMNLP 2018
Natural language and speech › Language models and text generation › large language model
GPT-2
0.112021
Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking · EMNLP (1) 2021
Machine learning › Graph learning › graph neural network › attention-based graph neural network
graph attention network
0.112021
Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking · EMNLP (1) 2021
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
speech recognition errors
0.112019
Machine Comprehension of Spoken Content: TOEFL Listening Test and Spoken SQuAD · IEEE ACM Trans. Audio Speech Lang. Process. 2019
Natural language and speech › Language models and text generation › neural language model
recurrent neural network language model
0.112017
Personalizing Recurrent-Neural-Network-Based Language Model by Social Network · IEEE ACM Trans. Audio Speech Lang. Process. 2017

Methods — techniques the papers use, named apart from their topics

simulated environment · 0.9user simulation · 0.5transfer learning · 0.5graph attention network · 0.5causal sequential prediction · 0.5GPT-2 · 0.5semi-supervised learning · 0.4hierarchical structure · 0.4bootstrapping · 0.4attention mechanism · 0.4
YearPublicationVenuePosition
2025 ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution
abstract
Alexandru Coca, Mark Gaynor, Zhenxing Zhang, Jianpeng Cheng, Bo-Hsiang Tseng, Peter Boothroyd, Hector Martinez Alonso, Diarmuid O Seaghdha, Anders Johannsen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Alexandru Coca, Mark Gaynor, Jianpeng Cheng 0001, Bo-Hsiang Tseng, Peter Boothroyd, Héctor Martínez Alonso, Diarmuid Ó Séaghdha, Anders Johannsen
ACL (1)5
2025 PyTOD: Programmable Task-Oriented Dialogue with Execution Feedback
abstract
Programmable task-oriented dialogue (TOD) agents enable language models to follow structured dialogue policies, but their effectiveness hinges on accurate dialogue state tracking (DST). We present PyTOD, an agent that generates executable code to track dialogue state and uses policy and execution feedback for efficient error correction. To achieve this, PyTOD employs a simple constrained decoding approach, using a language model instead of grammar rules to follow API schemata. This leads to state-of-the-art DST performance on the challenging SGD benchmark. Our experiments show that PyTOD surpasses strong baselines in both accuracy and cross-turn consistency, demonstrating the effectiveness of execution-aware state tracking.
Alexandru Coca, Bo-Hsiang Tseng, Peter Boothroyd, Jianpeng Cheng 0001, Mark Gaynor, Joe Stacey, Tristan Guigue, Héctor Martínez Alonso, Diarmuid Ó Séaghdha, Anders Johannsen
SIGDIAL2
2024 SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
abstract
Atharva Kulkarni, Bo-Hsiang Tseng, Joel Ruben Antony Moniz, Dhivya Piraviperumal, Hong Yu, Shruti Bhargava. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Atharva Kulkarni, Bo-Hsiang Tseng, Joel Ruben Antony Moniz, Dhivya Piraviperumal, Shruti Bhargava
EACL (1)2
2023 5IDER: Unified Query Rewriting for Steering, Intent Carryover, Disfluencies, Entity Carryover and Repair
abstract
Providing voice assistants the ability to navigate multi-turn conversations is a challenging problem.Handling multi-turn interactions requires the system to understand various conversational use-cases, such as steering, intent carryover, disfluencies, entity carryover, and repair.The complexity of this problem is compounded by the fact that these use-cases mix with each other, often appearing simultaneously in natural language.This work proposes a non-autoregressive query rewriting architecture that can handle not only the five aforementioned tasks, but also complex compositions of these use-cases.We show that our proposed model has competitive single task performance compared to the baseline approach, and even outperforms a fine-tuned T5 model in use-case compositions, despite being 15 times smaller in parameters and 25 times faster in latency.
Jiarui Lu, Bo-Hsiang Tseng, Joel Ruben Antony Moniz, Site Li, Xueyun Zhu, Murat Akbacak
INTERSPEECH2
2023 Grounding Description-Driven Dialogue State Trackers with Knowledge-Seeking Turns
abstract
Alexandru Coca, Bo-Hsiang Tseng, Jinghong Chen, Weizhe Lin, Weixuan Zhang, Tisha Anders, Bill Byrne. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023.
Alexandru Coca, Bo-Hsiang Tseng, Jinghong Chen, Weizhe Lin, Weixuan Zhang, Tisha Anders, William J. Byrne
SIGDIAL2
2021 Transferable Dialogue Systems and User Simulators
abstract
Bo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, Bill Byrne. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Bo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, William J. Byrne
ACL/IJCNLP (1)1
2021 Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking
abstract
Dialogue State Tracking is central to multidomain task-oriented dialogue systems, responsible for extracting information from user utterances.We present a novel hybrid architecture that augments GPT-2 with representations derived from Graph Attention Networks in such a way to allow causal, sequential prediction of slot values.The model architecture captures inter-slot relationships and dependencies across domains that otherwise can be lost in sequential prediction.We report improvements in state tracking performance in Mul-tiWOZ 2.0 against a strong GPT-2 baseline and investigate a simplified sparse training scenario in which DST models are trained only on session-level annotations but evaluated at the turn level.We further report detailed analyses to demonstrate the effectiveness of graph models in DST by showing that the proposed graph modules capture inter-slot dependencies and improve the predictions of values that are common to multiple domains.
Weizhe Lin, Bo-Hsiang Tseng, William J. Byrne
EMNLP (1)2
2021 CREAD: Combined Resolution of Ellipses and Anaphora in Dialogues
abstract
Bo-Hsiang Tseng, Shruti Bhargava, Jiarui Lu, Joel Ruben Antony Moniz, Dhivya Piraviperumal, Lin Li, Hong Yu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Bo-Hsiang Tseng, Shruti Bhargava, Jiarui Lu, Joel Ruben Antony Moniz, Dhivya Piraviperumal
NAACL-HLT1
2020 A Generative Model for Joint Natural Language Understanding and Generation
abstract
Natural language understanding (NLU) and natural language generation (NLG) are two fundamental and related tasks in building task-oriented dialogue systems with opposite objectives: NLU tackles the transformation from natural language to formal representations, whereas NLG does the reverse.A key to success in either task is parallel training data which is expensive to obtain at a large scale.In this work, we propose a generative model which couples NLU and NLG through a shared latent variable.This approach allows us to explore both spaces of natural language and formal representations, and facilitates information sharing through the latent space to eventually benefit NLU and NLG.Our model achieves state-of-the-art performance on two dialogue datasets with both flat and tree-structured formal representations.We also show that the model can be trained in a semi-supervised fashion by utilising unlabelled data to boost its performance.
Bo-Hsiang Tseng, Jianpeng Cheng 0001, Yimai Fang, David Vandyke
ACL1
2020 Improving Sample-Efficiency in Reinforcement Learning for Dialogue Systems by Using Trainable-Action-Mask
abstract
By interacting with human and learning from reward signals, reinforcement learning is an ideal way to build conversational AI. Concerning the expenses of real-users' responses, improving sample-efficiency has been the key issue when applying reinforcement learning in real-world spoken dialogue systems (SDS). Handcrafted action masks are commonly used to rule out impossible actions and accelerate the training process. However, the handcrafted action mask can barely be generalized to unseen domains. In this paper, we propose trainable-action-mask (TAM) which learns from data automatically without handcrafting complicated rules. In our experiments in Cambridge Restaurant domain, TAM requires only 30% of training data, compared with the baseline, to reach the 80% success rate and it also shows robustness to noisy environments.
Yen-Chen Wu, Bo-Hsiang Tseng, Carl E. Rasmussen
ICASSP2
2019 Semi-Supervised Bootstrapping of Dialogue State Trackers for Task-Oriented Modelling
abstract
Bo-Hsiang Tseng, Marek Rei, Paweł Budzianowski, Richard Turner, Bill Byrne, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Bo-Hsiang Tseng, Marek Rei, Pawel Budzianowski, Richard E. Turner, William J. Byrne, Anna Korhonen
EMNLP/IJCNLP (1)1
2019 Tree-Structured Semantic Encoder with Knowledge Sharing for Domain Adaptation in Natural Language Generation
abstract
Domain adaptation in natural language generation (NLG) remains challenging because of the high complexity of input semantics across domains and limited data of a target domain.This is particularly the case for dialogue systems, where we want to be able to seamlessly include new domains into the conversation.Therefore, it is crucial for generation models to share knowledge across domains for the effective adaptation from one domain to another.In this study, we exploit a tree-structured semantic encoder to capture the internal structure of complex semantic representations required for multi-domain dialogues in order to facilitate knowledge sharing across domains.In addition, a layer-wise attention mechanism between the tree encoder and the decoder is adopted to further improve the model's capability.The automatic evaluation results show that our model outperforms previous methods in terms of the BLEU score and the slot error rate, in particular when the adaptation data is limited.In subjective evaluation, human judges tend to prefer the sentences generated by our model, rating them more highly on informativeness and naturalness than other systems.
Bo-Hsiang Tseng, Pawel Budzianowski, Yen-Chen Wu, Milica Gasic
SIGdial1
2019 Machine Comprehension of Spoken Content: TOEFL Listening Test and Spoken SQuAD
abstract
A user can scan through a text easily, but it is not the case for spoken content, because they cannot be directly displayed on-screen. As a result, accessing large collections of spoken content is much more difficult and time-consuming than doing so for the text content. It would therefore be helpful to develop machines that understand spoken content. In this paper, we propose two new tasks for machine comprehension of spoken content. The first is a listening comprehension test for TOEFL, a challenging academic English examination for English learners who are not the native English speakers. We show that the proposed model outperforms the naive approaches and other neural network based models by exploiting the hierarchical structures of natural languages and the selective power of attention mechanism. For the second listening comprehension task - spoken SQuAD - we find that speech recognition errors severely impair machine comprehension; we propose the use of subword units to mitigate the impact of these errors.
Chia-Hsuan Lee 0001, Hung-yi Lee, Szu-Lin Wu, Chi-Liang Liu, Juei-Yang Hsu, Bo-Hsiang Tseng
IEEE ACM Trans. Audio Speech Lang. Process.7
2018 MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling
abstract
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, Milica Gašić. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Pawel Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, Milica Gasic
EMNLP3
2018 Feudal Dialogue Management with Jointly Learned Feature Extractors
abstract
Reinforcement learning (RL) is a promising dialogue policy optimisation approach, but traditional RL algorithms fail to scale to large domains.Recently, Feudal Dialogue Management (FDM), has shown to increase the scalability to large domains by decomposing the dialogue management decision into two steps, making use of the domain ontology to abstract the dialogue state in each step.In order to abstract the state space, however, previous work on FDM relies on handcrafted feature functions.In this work, we show that these feature functions can be learned jointly with the policy model while obtaining similar performance, even outperforming the handcrafted features in several environments and domains.
Iñigo Casanueva, Pawel Budzianowski, Stefan Ultes, Florian Kreyssig, Bo-Hsiang Tseng, Yen-Chen Wu, Milica Gasic
SIGDIAL Conference5
2018 Variational Cross-domain Natural Language Generation for Spoken Dialogue Systems
abstract
Cross-domain natural language generation (NLG) is still a difficult task within spoken dialogue modelling.Given a semantic representation provided by the dialogue manager, the language generator should generate sentences that convey desired information.Traditional template-based generators can produce sentences with all necessary information, but these sentences are not sufficiently diverse.With RNN-based models, the diversity of the generated sentences can be high, however, in the process some information is lost.In this work, we improve an RNN-based generator by considering latent information at the sentence level during generation using the conditional variational autoencoder architecture.We demonstrate that our model outperforms the original RNN-based generator, while yielding highly diverse sentences.In addition, our model performs better when the training data is limited.
Bo-Hsiang Tseng, Florian Kreyssig, Pawel Budzianowski, Iñigo Casanueva, Yen-Chen Wu, Stefan Ultes, Milica Gasic
SIGDIAL Conference1
2018 Addressing Objects and Their Relations: The Conversational Entity Dialogue Model
abstract
Stefan Ultes, Paweł Budzianowski, Iñigo Casanueva, Lina M. Rojas-Barahona, Bo-Hsiang Tseng, Yen-Chen Wu, Steve Young, Milica Gašić. Proceedings of the 19th Annual SIGdial Meeting on Discourse and Dialogue. 2018.
Stefan Ultes, Pawel Budzianowski, Iñigo Casanueva, Lina Maria Rojas-Barahona, Bo-Hsiang Tseng, Yen-Chen Wu, Steve J. Young, Milica Gasic
SIGDIAL Conference5
2017 Recurrent Neural Network based language modeling with controllable external Memory
abstract
It is crucial for language models to model long-term dependency in word sequences, which can be achieved to some good extent by recurrent neural network (RNN) based language models with long short-term memory (LSTM) units. To accurately model the sophisticated long-term information in human languages, large memory in language models is necessary. However, the size of RNN-based language models cannot be arbitrarily increased because the computational resources required and the model complexity will also be increase accordingly, due to the limitation of the structure. To overcome this problem, inspired from Neural Turing Machine and Memory Network, we equip RNN-based language models with controllable external memory. With a learnable memory controller, the size of the external memory is independent to the number of model parameters, so the proposed language model can have larger memory without increasing the parameters. In the experiments, the proposed model yielded lower perplexities than RNN-based language models with LSTM units on both English and Chinese corpora.
Wei-Jen Ko, Bo-Hsiang Tseng, Hung-yi Lee
ICASSP2
2017 Personalizing Recurrent-Neural-Network-Based Language Model by Social Network
abstract
With the popularity of mobile devices, personalized speech recognizers have become more attainable and are highly attractive. Since each mobile device is used primarily by a single user, it is possible to have a personalized recognizer that well matches the characteristics of the individual user. Although acoustic model personalization has been investigated for decades, much less work has been reported on personalizing language models, presumably because of the difficulties in collecting sufficient personalized corpora. In this paper, we propose a general framework for personalizing recurrent-neural-network-based language models (RNNLMs) using data collected from social networks, including the posts of many individual users and friend relationships among the users. Two major directions for this are model-based and feature-based RNNLM personalization. In model-based RNNLM personalization, the RNNLM parameters are fine-tuned to an individual user's wording patterns by incorporating social texts posted by the target user and his or her friends. For the feature-based approach, the RNNLM model parameters are fixed across users, but the RNNLM input features are instead augmented with personalized information. Both approaches not only drastically reduce the model perplexity, but also moderately reduce word error rates in n-best rescoring tests.
Hung-yi Lee, Bo-Hsiang Tseng, Tsung-Hsien Wen, Yu Tsao 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Towards Machine Comprehension of Spoken Content: Initial TOEFL Listening Comprehension Test by Machine
abstract
Multimedia or spoken content presents more attractive information than plain text content, but it's more difficult to display on a screen and be selected by a user. As a result, accessing large collections of the former is much more difficult and time-consuming than the latter for humans. It's highly attractive to develop a machine which can automatically understand spoken content and summarize the key information for humans to browse over. In this endeavor, we propose a new task of machine comprehension of spoken content. We define the initial goal as the listening comprehension test of TOEFL, a challenging academic English examination for English learners whose native language is not English. We further propose an Attention-based Multi-hop Recurrent Neural Network (AMRNN) architecture for this task, achieving encouraging results in the initial tests. Initial results also have shown that word-level attention is probably more robust than sentence-level attention for this task with ASR errors.
Bo-Hsiang Tseng, Sheng-syun Shen, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH1
2015 Personalizing universal recurrent neural network language model with user characteristic features by social network crowdsourcing
abstract
With the popularity of mobile devices, personalized speech recognizer becomes more realizable today and highly attractive. Each mobile device is primarily used by a single user, so it's possible to have a personalized recognizer well matching to the characteristics of individual user. Although acoustic model personalization has been investigated for decades, much less work have been reported on personalizing language model, probably because of the difficulties in collecting enough personalized corpora. Previous work used the corpora collected from social networks to solve the problem, but constructing a personalized model for each user is troublesome. In this paper, we propose a universal recurrent neural network language model with user characteristic features, so all users share the same model, except each with different user characteristic features. These user characteristic features can be obtained by crowdsouring over social networks, which include huge quantity of texts posted by users with known friend relationships, who may share some subject topics and wording patterns. The preliminary experiments on Facebook corpus showed that this proposed approach not only drastically reduced the model perplexity, but offered very good improvement in recognition accuracy in n-best rescoring tests. This approach also mitigated the data sparseness problem for personalized language models.
Bo-Hsiang Tseng, Hung-yi Lee, Lin-Shan Lee
ASRU1