Ting-Wei Wu

dblp:205/5008 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
15since 2021 · last 2023
0000-0002-4927-653XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Towards Zero-Shot Multilingual Transfer for Code-Switched Responses
abstract
Ting-Wei Wu, Changsheng Zhao, Ernie Chang, Yangyang Shi, Pierce Chuang, Vikas Chandra, Biing Juang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ting-Wei Wu, Changsheng Zhao 0002, Ernie Chang, Yangyang Shi, Pierce Chuang, Vikas Chandra, Biing-Hwang Juang
ACL (1)1
2023 Choice Fusion As Knowledge For Zero-Shot Dialogue State Tracking
abstract
With the demanding need for deploying dialogue systems in new domains with less cost, zero-shot dialogue state tracking (DST), which tracks user’s requirements in task-oriented dialogues without training on desired domains, draws attention increasingly. Although prior works have leveraged question-answering (QA) data to reduce the need for in-domain training in DST, they fail to explicitly model knowledge transfer and fusion for tracking dialogue states. To address this issue, we propose CoFunDST, which is trained on domain-agnostic QA datasets and directly uses candidate choices of slot-values as knowledge for zero-shot dialogue-state generation, based on a T5 pre-trained language model. Specifically, CoFunDST selects highly-relevant choices to the reference context and fuses them to initialize the decoder to constrain the model outputs. Our experimental results show that our proposed model achieves outperformed joint goal accuracy compared to existing zero-shot DST approaches in most domains on the MultiWOZ 2.1. Extensive analyses demonstrate the effectiveness of our proposed approach for improving zero-shot DST learning from QA.
Ruolin Su, Jingfeng Yang 0001, Ting-Wei Wu, Biing-Hwang Juang
ICASSP3
2023 Expert-defined Keywords Improve Interpretability of Retinal Image Captioning
abstract
Automatic machine learning-based (ML-based) medical report generation systems for retinal images suffer from a relative lack of interpretability. Hence, such ML-based systems are still not widely accepted. The main reason is that trust is one of the important motivating aspects of interpretability and humans do not trust blindly. Precise technical definitions of interpretability still lack consensus. Hence, it is difficult to make a human-comprehensible ML-based medical report generation system. Heat maps/saliency maps, i.e., post-hoc explanation approaches, are widely used to improve the interpretability of ML-based medical systems. However, they are well known to be problematic. From an ML-based medical model’s perspective, the highlighted areas of an image are considered important for making a prediction. However, from a doctor’s perspective, even the hottest regions of a heat map contain both useful and non-useful information. Simply localizing the region, therefore, does not reveal exactly what it was in that area that the model considered useful. Hence, the post-hoc explanation-based method relies on humans who probably have a biased nature to decide what a given heat map might mean. Interpretability boosters, in particular expert-defined keywords, are effective carriers of expert domain knowledge and they are human-comprehensible. In this work, we propose to exploit such keywords and a specialized attention-based strategy to build a more human-comprehensible medical report generation system for retinal images. Both keywords and the proposed strategy effectively improve the interpretability. The proposed method achieves state-of-the-art performance under commonly used text evaluation metrics BLEU, ROUGE, CIDEr, and METEOR. Project website: https://github.com/Jhhuangkay/Expert-defined-Keywords-Improve-Interpretability-of-Retinal-Image-Captioning.
Ting-Wei Wu, Jia-Hong Huang, Marcel Worring
WACV1
2022 Knowledge Augmented Bert Mutual Network in Multi-Turn Spoken Dialogues
abstract
Modern spoken language understanding (SLU) systems rely on sophisticated semantic notions revealed in single utterances to detect intents and slots. However, they lack the capability of modeling multi-turn dynamics within a dialogue particularly in long-term slot contexts. Without external knowledge, depending on limited linguistic legitimacy within a word sequence may overlook deep semantic information across dialogue turns. In this paper, we propose to equip a BERT-based joint model with a knowledge attention module to mutually leverage dialogue contexts between two SLU tasks. A gating mechanism is further utilized to filter out irrelevant knowledge triples and to circumvent distracting comprehension. Experimental results in two complicated multi-turn dialogue datasets have demonstrate by mutually modeling two SLU tasks with filtered knowledge and dialogue contexts, our approach has considerable improvements compared with several competitive baselines.
Ting-Wei Wu, Biing-Hwang Juang
ICASSP1
2022 Efficient Privacy Preserving Nearest Neighboring Classification from Tree Structures and Secret Sharing
abstract
The k-nearest neighbor (kNN) algorithm is a very simple manner in the area of machine learning. It is a supervised method to classify according to the distance between different instances and is also widely used in solving some classification problems. It is expected to obtain better training with a larger dataset. However, how to perform kNN algorithm efficiently is an issue with privacy-preserving. In this paper, we proposed a privacy-preserving k -nearest neighboring scheme by secret sharing and improve the kNN classification by preprocessing with tree structures. Finally, the applicability of our method is shown by experiments with real datasets.
Jhe-Kai Yang, Kuan-Chun Huang, Cheng-Yang Chung, Yu-Chi Chen 0001, Ting-Wei Wu
ICC5
2022 Learning to rank with BERT-based confidence models in ASR rescoring
Ting-Wei Wu, I-Fan Chen, Ankur Gandhe
INTERSPEECH1
2022 Induce Spoken Dialog Intents via Deep Unsupervised Context Contrastive Clustering
Ting-Wei Wu, Biing-Hwang Juang
INTERSPEECH1
2022 Non-local Attention Improves Description Generation for Retinal Images
abstract
Automatically generating medical reports from retinal images is a difficult task in which an algorithm must generate semantically coherent descriptions for a given retinal image. Existing methods mainly rely on the input image to generate descriptions. However, many abstract medical concepts or descriptions cannot be generated based on image information only. In this work, we integrate additional information to help solve this task; we observe that early in the diagnosis process, ophthalmologists have usually written down a small set of keywords denoting important information. These keywords are then subsequently used to aid the later creation of medical reports for a patient. Since these keywords commonly exist and are useful for generating medical reports, we incorporate them into automatic report generation. Since we have two types of inputs expert-defined unordered keywords and images - effectively fusing features from these different modalities is challenging. To that end, we propose a new keyword-driven medical report generation method based on a non-local attention-based multi-modal feature fusion approach, TransFuser, which is capable of fusing features from different types of inputs based on such attention. Our experiments show the proposed method successfully captures the mutual information of keywords and image content. We further show our proposed keyword-driven generation model reinforced by the TransFuser is superior to baselines under the popular text evaluation metrics BLEU, CIDEr, and ROUGE. Trans-Fuser Github: https://github.com/Jhhuangkay/Non-local-Attention-ImprovesDescription-Generation-for-Retinal-Images.
Jia-Hong Huang, Ting-Wei Wu, Chao-Han Huck Yang, Zenglin Shi, I-Hung Lin, Jesper Tegnér, Marcel Worring
WACV2
2021 IM Receptivity and Presentation-type Preferences among Users of a Mobile App with Automated Receptivity-status Adjustment
abstract
Researchers have long attempted to estimate instant-messaging (IM) users’ attentiveness, responsiveness, and interruptibility. Yet, IM users’ self-presentation of their receptivity, and their perceptions of automated adjustment/revelation of their receptivity status (e.g., Facebook Messenger’s green dot that deems a user to be “active”), remain under-explored. We therefore told our 43 participants that our IM app, IMStatus, was capable of automatically estimating and adjusting their receptivity status to responsive, attentive, or interruptible based on their smartphone activity. These statuses were also presented to their IM contacts in three different styles. Over a two-week period, the participants rarely chose the status interruptible, and when they did, it was usually to indicate low availability. Textual presentation was usually chosen to express statuses precisely, especially at high and low extremes of receptivity; while graphical and numeric presentations were preferred when self-perceived receptivity levels were more ambiguous. Conflicts between recipients’ and senders’ perspectives are also discussed.
Ting-Wei Wu, Yu-Ling Chien, Hao-Ping Lee, Yung-Ju Chang
CHI1
2021 A Label-Aware BERT Attention Network for Zero-Shot Multi-Intent Detection in Spoken Language Understanding
abstract
With the early success of query-answer assistants such as Alexa and Siri, research attempts to expand system capabilities of handling service automation are now abundant.However, preliminary systems have quickly found the inadequacy in relying on simple classification techniques to effectively accomplish the automation task.The main challenge is that the dialogue often involves complexity in user's intents (or purposes) which are multiproned, subject to spontaneous change, and difficult to track.Furthermore, public datasets have not considered these complications and the general semantic annotations are lacking which may result in zero-shot problem.Motivated by the above, we propose a Label-Aware BERT Attention Network (LABAN) for zeroshot multi-intent detection.We first encode input utterances with BERT and construct a label embedded space by considering embedded semantics in intent labels.An input utterance is then classified based on its projection weights on each intent embedding in this embedded space.We show that it successfully extends to few/zero-shot setting where part of intent labels are unseen in training data, by also taking account of semantics in these unseen intent labels.Experimental results show that our approach is capable of detecting many unseen intent labels correctly.It also achieves the state-of-the-art performance on five multiintent datasets in normal cases.
Ting-Wei Wu, Ruolin Su, Biing-Hwang Juang
EMNLP (1)1
2021 Deep Context-Encoding Network For Retinal Image Captioning
abstract
Automatically generating medical reports for retinal images is one of the promising ways to help ophthalmologists reduce their workload and improve work efficiency. In this work, we propose a new context-driven encoding network to automatically generate medical reports for retinal images. The proposed model is mainly composed of a multi-modal input encoder and a fused-feature decoder. Our experimental results show that our proposed method is capable of effectively leveraging the interactive information between the input image and context, i.e., keywords in our case. The proposed method creates more accurate and meaningful reports for retinal images than baseline models and achieves state-of-the-art performance. This performance is shown in several commonly used metrics for the medical report generation task: BLEUavg (+16%), CIDEr (+10.2%), and ROUGE (+8.6%).
Jia-Hong Huang, Ting-Wei Wu, Chao-Han Huck Yang, Marcel Worring
ICIP2
2021 Act-Aware Slot-Value Predicting in Multi-Domain Dialogue State Tracking
abstract
As an essential component in task-oriented dialogue systems, dialogue state tracking (DST) aims to track human-machine interactions and generate state representations for managing the dialogue. Representations of dialogue states are dependent on the domain ontology and the user's goals. In several task-oriented dialogues with a limited scope of objectives, dialogue states can be represented as a set of slot-value pairs. As the capabilities of dialogue systems expand to support increasing naturalness in communication, incorporating dialogue act processing into dialogue model design becomes essential. The lack of such consideration limits the scalability of dialogue state tracking models for dialogues having specific objectives and ontology. To address this issue, we formulate and incorporate dialogue acts, and leverage recent advances in machine reading comprehension to predict both categorical and non-categorical types of slots for multi-domain dialogue state tracking. Experimental results show that our models can improve the overall accuracy of dialogue state tracking on the MultiWOZ 2.1 dataset, and demonstrate that incorporating dialogue acts can guide dialogue state design for future task-oriented dialogue systems.
Ruolin Su, Ting-Wei Wu, Biing-Hwang Juang
Interspeech2
2021 A Context-Aware Hierarchical BERT Fusion Network for Multi-Turn Dialog Act Detection
abstract
The success of interactive dialog systems is usually associated with the quality of the spoken language understanding (SLU) task, which mainly identifies the corresponding dialog acts and slot values in each turn.By treating utterances in isolation, most SLU systems often overlook the semantic context in which a dialog act is expected.The act dependency between turns is nontrivial and yet critical to the identification of the correct semantic representations.Previous works with limited context awareness have exposed the inadequacy of dealing with complexity in multiproned user intents, which are subject to spontaneous change during turn transitions.In this work, we propose to enhance SLU in multi-turn dialogs, employing a context-aware hierarchical BERT fusion Network (CaBERT-SLU) to not only discern context information within a dialog but also jointly identify multiple dialog acts and slots in each utterance.Experimental results show that our approach reaches new state-of-the-art (SOTA) performances in two complicated multi-turn dialogue datasets with considerable improvements compared with previous methods, which only consider single utterances for multiple intents and slot filling.
Ting-Wei Wu, Ruolin Su, Biing-Hwang Juang
Interspeech1
2021 Contextualized Keyword Representations for Multi-modal Retinal Image Captioning
abstract
Medical image captioning automatically generates a medical description to describe the content of a given medical image. Traditional medical image captioning models create a medical description based on a single medical image input only. Hence, an abstract medical description or concept is hard to be generated based on the traditional approach. Such a method limits the effectiveness of medical image captioning. Multi-modal medical image captioning is one of the approaches utilized to address this problem. In multi-modal medical image captioning, textual input, e.g., expert-defined keywords, is considered as one of the main drivers of medical description generation. Thus, encoding the textual input and the medical image effectively are both important for the task of multi-modal medical image captioning. In this work, a new end-to-end deep multi-modal medical image captioning model is proposed. Contextualized keyword representations, textual feature reinforcement, and masked self-attention are used to develop the proposed approach. Based on the evaluation of an existing multi-modal medical image captioning dataset, experimental results show that the proposed model is effective with an increase of +53.2% in BLEU-avg and +18.6% in CIDEr, compared with the state-of-the-art method. https://github.com/Jhhuangkay/Contextualized-Keyword-Representations-for-Multi-modal-Retinal-Image-Captioning
Jia-Hong Huang, Ting-Wei Wu, Marcel Worring
ICMR2
2021 DeepOpht: Medical Report Generation for Retinal Images via Deep Models and Visual Explanation
abstract
In this work, we propose an AI-based method that intends to improve the conventional retinal disease treatment procedure and help ophthalmologists increase diagnosis efficiency and accuracy. The proposed method is composed of a deep neural networks-based (DNN-based) module, including a retinal disease identifier and clinical description generator, and a DNN visual explanation module. To train and validate the effectiveness of our DNN-based module, we propose a large-scale retinal disease image dataset. Also, as ground truth, we provide a retinal image dataset manually labeled by ophthalmologists to qualitatively show the proposed AI-based method is effective. With our experimental results, we show that the proposed method is quantitatively and qualitatively effective. Our method is capable of creating meaningful retinal image descriptions and visual explanations that are clinically relevant.https://github.com/Jhhuangkay/DeepOpht-Medical-Report-Generation-for-Retinal-Images-via-Deep-Models-and-Visual-Explanation.
Jia-Hong Huang, Chao-Han Huck Yang, Fangyu Liu 0001, Meng Tian 0002, Yi-Chieh Liu, Ting-Wei Wu, I-Hung Lin, Hiromasa Morikawa, Hernghua Chang, Jesper Tegnér, Marcel Worring
WACV6
2019 Exploring the Design of Availability Status in Mobile IM Messaging with User Enactments
abstract
Current mobile instant messaging (IM) applications offer limited information on the availability status of IM users, particularly, their availability for reading and responding to IM messages. Research suggests a gap between what IM recipients want to disclose and what IM senders want to see to determine when to initiate a conversation. The advancement of IM users' receptivity prediction makes it possible to present IM users' predicted availability status. In this research, we conducted user enactment, a design approach for researchers to let participants experience and reflect on possible designs of future technologies, to explore designs of IM availability status in 72 IM conversation scenarios. We explore how IM users interpret different presentations of an uncertain IM status from both the senders' and recipients' perspectives, and what they need and will act upon these presentations.
Yu-Ling Chien, Ting-Wei Wu, Yung-Ju Chang
MobileHCI2