VLDB 2026 Research / reviewers in the wild / expert
Ting-Hao 'Kenneth' Huang
dblp:215/4581 · also Ting-Hao K. Huang
· DBLP profile ↗
39ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-7021-4627ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 2 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 15 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | "Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online ReviewsabstractCoded language is an important part of human communication.It refers to cases where users intentionally encode meaning so that the surface text differs from the intended meaning and must be decoded to be understood.Current language models handle coded language poorly.Progress has been limited by the lack of real-world datasets and clear taxonomies.This paper introduces CODEDLANG, a dataset of 7,744 Chinese Google Maps reviews, including 900 reviews with span-level annotations of coded language.We developed a seven-class taxonomy that captures common encoding strategies, including phonetic, orthographic, and cross-lingual substitutions.We benchmarked language models on coded language detection, classification, and review rating prediction.Results show that even strong models can fail to identify or understand coded language.Because many coded expressions rely on pronunciation-based strategies, we further conducted a phonetic analysis of coded and decoded forms.Our code and dataset are publicly available 1 .Together, our results highlight coded language as an important and underexplored challenge for real-world NLP systems. Ruyuan Wan, Changye Li 0001, Ting-Hao 'Kenneth' Huang |
ACL (1) | 3 |
| 2026 | Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023abstractAbstract Since the SciCap dataset’s launch in 2021, the research community has made significant progress in generating captions for scientific figures in scholarly articles. In 2023, the first SciCap Challenge took place, inviting global teams to use an expanded SciCap dataset to develop models for captioning diverse figure types across various academic fields. At the same time, text generation models advanced quickly, with many powerful pre-trained large multimodal models (LMMs) emerging that showed impressive capabilities in various vision-and-language tasks. This paper presents an overview of the first SciCap Challenge and details the performance of various models on its data, capturing a snapshot of the field’s state. We found that professional editors overwhelmingly preferred figure captions generated by GPT-4V over those from all other models and even the original captions written by authors. Following this key finding, we conducted detailed analyses to answer this question: Have advanced LMMs solved the task of generating captions for scientific figures? Ting-Yao Hsu, Yi-Li Hsu, Shaurya Rohatgi, Chieh-Yang Huang, Ho Yin Sam Ng, Ryan Rossi, Sungchul Kim, Tong Yu 0001, Lun-Wei Ku, C. Lee Giles, Ting-Hao 'Kenneth' Huang |
Trans. Assoc. Comput. Linguistics | 11 |
| 2025 | From Selection to Generation: A Survey of LLM-based Active LearningabstractYu Xia, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen, Franck Dernoncourt, Branislav Kveton, Tong Yu, Ruiyi Zhang, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang, Xiang Chen, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao, Nedim Lipka, Seunghyun Yoon, Ting-Hao Kenneth Huang, Zichao Wang, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee, Zhehao Zhang, Namyong Park, Thien Huu Nguyen, Jiebo Luo, Ryan A. Rossi, Julian McAuley. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Xia 0007, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li 0001, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen 0003, Franck Dernoncourt, Branislav Kveton, Tong Yu 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang 0160, Xiang Chen 0010, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao 0016, Nedim Lipka, Seunghyun Yoon 0002, Ting-Hao 'Kenneth' Huang, Zichao Wang 0001, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee 0001, Zhehao Zhang 0001, Namyong Park 0001, Thien Huu Nguyen, Jiebo Luo 0001, Ryan Rossi, Julian J. McAuley |
ACL (1) | 24 |
| 2025 | Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are AbsentabstractMillions of users prompt large language models (LLMs) for various tasks, but how good are people at prompt engineering? Do users actually get closer to their desired outcome over multiple iterations of their prompts? These questions are crucial when no gold-standard labels are available to measure progress. This paper investigates a scenario in LLM-powered data labeling, "prompting in the dark," where users iteratively prompt LLMs to label data without using manually-labeled benchmarks. We developed PromptingSheet, a Google Sheets add-on that enables users to compose, revise, and iteratively label data through spreadsheets. Through a study with 20 participants, we found that prompting in the dark was highly unreliable -- only 9 participants improved labeling accuracy after four or more iterations. Automated prompt optimization tools like DSPy also struggled when few gold labels were available. Our findings highlight the importance of gold labels and the needs, as well as the risks, of automated support in human prompt engineering, providing insights for future tool design. Saniya Naphade, Ting-Hao 'Kenneth' Huang |
CHI | 3 |
| 2025 | Hashtag Re-Appropriation for Audience Control on Recommendation-Driven Social Media Xiaohongshu (rednote)abstractAlgorithms have played a central role in personalized recommendations on social media. However, they also present significant obstacles for content creators trying to predict and manage their audience reach. This issue is particularly challenging for marginalized groups seeking to maintain safe spaces. Our study explores how women on Xiaohongshu (rednote), a recommendation-driven social platform, proactively re-appropriate hashtags (e.g., #Baby Supplemental Food) by using them in posts unrelated to their literal meaning. The hashtags were strategically chosen from topics that would be uninteresting to the male audience they wanted to block. Through a mixed-methods approach, we analyzed the practice of hashtag re-appropriation based on 5,800 collected posts and interviewed 24 active users from diverse backgrounds to uncover users' motivations and reactions towards the re-appropriation. This practice highlights how users can reclaim agency over content distribution on recommendation-driven platforms, offering insights into self-governance within algorithmic-centered power structures. Ruyuan Wan, Lingbo Tong, Tiffany Knearem, Toby Jia-Jun Li, Ting-Hao 'Kenneth' Huang, Qunfang Wu |
CHI | 5 |
| 2024 | If in a Crowdsourced Data Annotation Pipeline, a GPT-4abstractRecent studies indicated GPT-4 outperforms online crowd workers in data labeling accuracy, notably workers from Amazon Mechanical Turk (MTurk). However, these studies were criticized for deviating from standard crowdsourcing practices and emphasizing individual workers’ performances over the whole data-annotation process. This paper compared GPT-4 and an ethical and well-executed MTurk pipeline, with 415 workers labeling 3,177 sentence segments from 200 scholarly articles using the CODA-19 scheme. Two worker interfaces yielded 127,080 labels, which were then used to infer the final labels through eight label-aggregation algorithms. Our evaluation showed that despite best practices, MTurk pipeline’s highest accuracy was 81.5%, whereas GPT-4 achieved 83.6%. Interestingly, when combining GPT-4’s labels with crowd labels collected via an advanced worker interface for aggregation, 2 out of the 8 algorithms achieved an even higher accuracy (87.5%, 87.0%). Further analysis suggested that, when the crowd’s and GPT-4’s labeling strengths are complementary, aggregating them could increase labeling accuracy. Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Shaurya Rohatgi, Ting-Hao 'Kenneth' Huang |
CHI | 5 |
| 2024 | What Color Scheme is More Effective in Assisting Readers to Locate Information in a Color-Coded Article?abstractColor coding, a technique assigning specific colors to cluster information types, has proven advantages in aiding human cognitive activities, especially reading and comprehension. The rise of Large Language Models (LLMs) has streamlined document coding, enabling simple automatic text labeling with various schemes. This has the potential to make color-coding more accessible and benefit more users. However, the impact of color choice on information seeking is understudied. We conducted a user study assessing various color schemes’ effectiveness in LLM-coded text documents, standardizing contrast ratios to approximately 5.55:1 across schemes. Participants performed timed information-seeking tasks in color-coded scholarly abstracts. Results showed non-analogous and yellow-inclusive color schemes improved performance, with the latter also being more preferred by participants. These findings can inform better color scheme choices for text annotation. As LLMs advance document coding, we advocate for more research focusing on the "color" aspect of color-coding techniques. Ho Yin Ng, Ting-Hao 'Kenneth' Huang |
IEEE VIS | 3 |
| 2023 | Unmasking Nationality Bias: A Study of Human Perception of Nationalities in AI-Generated ArticlesabstractWe investigate the potential for nationality biases in natural language processing (NLP) models using human evaluation methods. Biased NLP models can perpetuate stereotypes and lead to algorithmic discrimination, posing a significant challenge to the fairness and justice of AI systems. Our study employs a two-step mixed-methods approach that includes both quantitative and qualitative analysis to identify and understand the impact of nationality bias in a text generation model. Through our human-centered quantitative analysis, we measure the extent of nationality bias in articles generated by AI sources. We then conduct open-ended interviews with participants, performing qualitative coding and thematic analysis to understand the implications of these biases on human readers. Our findings reveal that biased NLP models tend to replicate and amplify existing societal biases, which can translate to harm if used in a sociotechnical setting. The qualitative analysis from our interviews offers insights into the experience readers have when encountering such articles, highlighting the potential to shift a reader’s perception of a country. These findings emphasize the critical role of public perception in shaping AI’s impact on society and the need to correct biases in AI systems. Pranav Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao 'Kenneth' Huang, Shomir Wilson |
AIES | 4 |
| 2023 | Nationality Bias in Text GenerationabstractPranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, Shomir Wilson. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Pranav Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao 'Kenneth' Huang, Shomir Wilson |
EACL | 4 |
| 2023 | Location-Aware Visual Question Generation with Lightweight ModelsabstractThis work introduces a novel task, locationaware visual question generation (LocaVQG), which aims to generate engaging questions from data relevant to a particular geographical location.Specifically, we represent such location-aware information with surrounding images and a GPS coordinate.To tackle this task, we present a dataset generation pipeline that leverages GPT-4 to produce diverse and sophisticated questions.Then, we aim to learn a lightweight model that can address the Lo-caVQG task and fit on an edge device, such as a mobile phone.To this end, we propose a method which can reliably generate engaging questions from location-aware information.Our proposed method outperforms baselines regarding human evaluation (e.g., engagement, grounding, coherence) and automatic evaluation metrics (e.g., BERTScore, ROUGE-2).Moreover, we conduct extensive ablation studies to justify our proposed techniques for generating the dataset and solving the task. Nicholas Collin Suwono, Justin Chih-Yao Chen, Tun-Min Hung, Ting-Hao 'Kenneth' Huang, I-Bin Liao, Yung-Hui Li, Lun-Wei Ku, Shao-Hua Sun |
EMNLP | 4 |
| 2023 | Summaries as Captions: Generating Figure Captions for Scientific Documents with Automated Text SummarizationabstractChieh-Yang Huang, Ting-Yao Hsu, Ryan Rossi, Ani Nenkova, Sungchul Kim, Gromit Yeuk-Yin Chan, Eunyee Koh, C Lee Giles, Ting-Hao Huang. Proceedings of the 16th International Natural Language Generation Conference. 2023. Chieh-Yang Huang, Ting-Yao Hsu, Ryan Rossi, Ani Nenkova, Sungchul Kim, Gromit Yeuk-Yin Chan, Eunyee Koh, C. Lee Giles, Ting-Hao 'Kenneth' Huang |
INLG | 9 |
| 2022 | Learning to Rank Visual Stories From Human Ranking DataabstractChi-Yang Hsu, Yun-Wei Chu, Vincent Chen, Kuan-Chieh Lo, Chacha Chen, Ting-Hao Huang, Lun-Wei Ku. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Chi-Yang Hsu, Yun-Wei Chu, Kuan-Chieh Lo, Chacha Chen, Ting-Hao 'Kenneth' Huang, Lun-Wei Ku |
ACL (1) | 6 |
| 2022 | Multi-VQG: Generating Engaging Questions for Multiple ImagesabstractGenerating engaging content has drawn much recent attention in the NLP community.Asking questions is a natural way to respond to photos and promote awareness.However, most answers to questions in traditional questionanswering (QA) datasets are factoids, which reduce individuals' willingness to answer.Furthermore, traditional visual question generation (VQG) confines the source data for question generation to single images, resulting in a limited ability to comprehend time-series information of the underlying event.In this paper, we propose generating engaging questions from multiple images.We present MVQG 1 , a new dataset, and establish a series of baselines, including both end-to-end and dual-stage architectures.Results show that building stories behind the image sequence enables models to generate engaging questions, which confirms our assumption that people typically construct a picture of the event in their minds before asking questions.These results open up an exciting challenge for visual-and-language models to implicitly construct a story behind a series of photos to allow for creativity and experience sharing and hence draw attention to downstream applications.How would you act if you found yourself in a room filled with cans of free drinks?Have you ever gone to beer tastings and where would that be at?How long did the cat lounge around in the book room?What would this cat sit on next? Min-Hsuan Yeh, Ting-Hao 'Kenneth' Huang, Lun-Wei Ku |
EMNLP | 3 |
| 2021 | Project RISE: Recognizing Industrial Smoke EmissionsabstractIndustrial smoke emissions pose a significant concern to human health. Prior works have shown that using Computer Vision (CV) techniques to identify smoke as visual evidence can influence the attitude of regulators and empower citizens to pursue environmental justice. However, existing datasets are not of sufficient quality nor quantity to train the robust CV models needed to support air quality advocacy. We introduce RISE, the first large-scale video dataset for Recognizing Industrial Smoke Emissions. We adopted a citizen science approach to collaborate with local community members to annotate whether a video clip has smoke emissions. Our dataset contains 12,567 clips from 19 distinct views from cameras that monitored three industrial facilities. These daytime clips span 30 days over two years, including all four seasons. We ran experiments using deep neural networks to establish a strong performance baseline and reveal smoke recognition challenges. Our survey study discussed community feedback, and our data analysis displayed opportunities for integrating citizen scientists and crowd workers into the application of Artificial Intelligence for Social Impact. Yen-Chia Hsu, Ting-Hao 'Kenneth' Huang, Ting-Yao Hu, Paul Dille, Sean Prendi, Ryan Hoffman, Anastasia Tsuhlares, Jessica Pachuta, Randy Sargent, Illah R. Nourbakhsh |
AAAI | 2 |
| 2021 | ABCD: A Graph Framework to Convert Complex Sentences to a Covering Set of Simple SentencesabstractYanjun Gao, Ting-Hao Huang, Rebecca J. Passonneau. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yanjun Gao, Ting-Hao 'Kenneth' Huang, Rebecca J. Passonneau |
ACL/IJCNLP (1) | 2 |
| 2021 | FinQA: A Dataset of Numerical Reasoning over Financial DataabstractZhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, William Yang Wang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhiyu Chen 0002, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao 'Kenneth' Huang, Bryan R. Routledge, William Yang Wang |
EMNLP (1) | 9 |
| 2021 | Semantic Frame ForecastabstractThis paper introduces semantic frame forecast, a task that predicts the semantic frames that will occur in the next 10, 100, or even 1,000 sentences in a running story.Prior work focused on predicting the immediate future of a story, such as one to a few sentences ahead.However, when novelists write long stories, generating a few sentences is not enough to help them gain high-level insight to develop the follow-up story.In this paper, we formulate a long story as a sequence of "story blocks," where each block contains a fixed number of sentences (e.g., 10, 100, or 200).This formulation allows us to predict the follow-up story arc beyond the scope of a few sentences.We represent a story block using the term frequencies (TF) of semantic frames in it, normalized by each frame's inverse document frequency (IDF).We conduct semantic frame forecast experiments on 4,794 books from the Bookcorpus and 7,962 scientific abstracts from CODA-19, with block sizes ranging from 5 to 1,000 sentences.The results show that automated models can forecast the follow-up story blocks better than the random, prior, and replay baselines, indicating the task's feasibility.We also learn that the models using the frame representation as features outperform all the existing approaches when the block size is over 150 sentences.The human evaluation also shows that the proposed frame representation, when visualized as word clouds, is comprehensible, representative, and specific to humans.Our code is available at: https://github.com/ appleternity/FrameForecasting. Chieh-Yang Huang, Ting-Hao 'Kenneth' Huang |
NAACL-HLT | 2 |
| 2020 | Knowledge-Enriched Visual StorytellingabstractStories are diverse and highly personalized, resulting in a large possible output space for story generation. Existing end-to-end approaches produce monotonous stories because they are limited to the vocabulary and knowledge in a single training dataset. This paper introduces KG-Story, a three-stage framework that allows the story generation model to take advantage of external Knowledge Graphs to produce interesting stories. KG-Story distills a set of representative words from the input prompts, enriches the word set by using external knowledge graphs, and finally generates stories based on the enriched word set. This distill-enrich-generate framework allows the use of external resources not only for the enrichment phase, but also for the distillation and generation phases. In this paper, we show the superiority of KG-Story for visual storytelling, where the input prompt is a sequence of five photos and the output is a short story. Per the human ranking evaluation, stories generated by KG-Story are on average ranked better than that of the state-of-the-art systems. Our code and output stories are available at https://github.com/zychen423/KE-VIST. Chao-Chun Hsu, Zi-Yuan Chen, Chi-Yang Hsu, Chih-Chia Li, Tzu-Yuan Lin, Ting-Hao 'Kenneth' Huang, Lun-Wei Ku |
AAAI | 6 |
| 2020 | Heteroglossia: In-Situ Story Ideation with the CrowdabstractIdeation is essential for creative writing. Many authors struggle to come up with ideas throughout the writing process, yet modern writing tools fail to provide on-the-spot assistance for writers when they get stuck. This paper introduces Heteroglossia, an add-on for Google Docs that allows writers to elicit story ideas from the online crowd using their text editors. Writers can share snippets of their working drafts and ask the crowd to provide follow-up story ideas based on it. Heteroglossia employs a strategy called "role play", where each worker is assigned a fictional character in a story and asked to brainstorm plot ideas from that character's perspective. Our deployment with two experienced story writers shows that Heteroglossia is easy to use and can generate interesting ideas. Heteroglossia allows us to gain insight into how future technologies can be developed to support ideation in creative writing. Chieh-Yang Huang, Shih-Hong Huang, Ting-Hao 'Kenneth' Huang |
CHI | 3 |
| 2020 | Assessing the Helpfulness of Learning Materials with Inference-Based Learner-Like AgentabstractMany English-as-a-second language learners have trouble using near-synonym words (e.g., small vs.little; briefly vs.shortly) correctly, and often look for example sentences to learn how two nearly synonymous terms differ. Prior work uses hand-crafted scores to recommend sentences but has difficulty in adopting such scores to all the near-synonyms as near-synonyms differ in various ways. We notice that the helpfulness of the learning material would reflect on the learners’ performance. Thus, we propose the inference-based learner-like agent to mimic learner behavior and identify good learning materials by examining the agent’s performance. To enable the agent to behave like a learner, we leverage entailment modeling’s capability of inferring answers from the provided materials. Experimental results show that the proposed agent is equipped with good learner-like behavior to achieve the best performance in both fill-in-the-blank (FITB) and good example sentence selection tasks. We further conduct a classroom user study with college ESL learners. The results of the user study show that the proposed agent can find out example sentences that help students learn more easily and efficiently. Compared to other models, the proposed agent improves the score of more than 17% of students after learning. Yun-Hsuan Jen, Chieh-Yang Huang, Mei-Hua Chen, Ting-Hao 'Kenneth' Huang, Lun-Wei Ku |
EMNLP (1) | 4 |
| 2020 | How Useful Are the Machine-Generated Interpretations to General Users? A Human Evaluation on Guessing the Incorrectly Predicted LabelsabstractExplaining to users why automated systems make certain mistakes is important and challenging. Researchers have proposed ways to automatically produce interpretations for deep neural network models. However, it is unclear how useful these interpretations are in helping users figure out why they are getting an error. If an interpretation effectively explains to users how the underlying deep neural network model works, people who were presented with the interpretation should be better at predicting the model’s outputs than those who were not. This paper presents an investigation on whether or not showing machine-generated visual interpretations helps users understand the incorrectly predicted labels produced by image classifiers. We showed the images and the correct labels to 150 online crowd workers and asked them to select the incorrectly predicted labels with or without showing them the machine-generated visual interpretations. The results demonstrated that displaying the visual interpretations did not increase, but rather decreased, the average guessing accuracy by roughly 10%. Hua Shen 0005, Ting-Hao 'Kenneth' Huang |
HCOMP | 2 |
| 2020 | Smell Pittsburgh: Engaging Community Citizen Science for Air QualityabstractUrban air pollution has been linked to various human health concerns, including cardiopulmonary diseases. Communities who suffer from poor air quality often rely on experts to identify pollution sources due to the lack of accessible tools. Taking this into account, we developedSmell Pittsburgh, a system that enables community members to report odors and track where these odors are frequently concentrated. All smell report data are publicly accessible online. These reports are also sent to the local health department and visualized on a map along with air quality data from monitoring stations. This visualization provides a comprehensive overview of the local pollution landscape. Additionally, with these reports and air quality data, we developed a model to predict upcoming smell events and send push notifications to inform communities. We also applied regression analysis to identify statistically significant effects of push notifications on user engagement. Our evaluation of this system demonstrates that engaging residents in documenting their experiences with pollution odors can help identify local air pollution patterns and can empower communities to advocate for better air quality. All citizen-contributed smell data are publicly accessible and can be downloaded fromhttps://smellpgh.org. Yen-Chia Hsu, Jennifer L. Cross, Paul Dille, Michael Tasota, Beatrice Dias, Randy Sargent, Ting-Hao 'Kenneth' Huang, Illah R. Nourbakhsh |
ACM Trans. Interact. Intell. Syst. | 7 |
| 2019 | Visual Story Post-EditingabstractWe introduce the first dataset for human edits of machine-generated visual stories and explore how these collected edits may be used for the visual story post-editing task.The dataset, VIST-Edit 1 , includes 14,905 humanedited versions of 2,981 machine-generated visual stories.The stories were generated by two state-of-the-art visual storytelling models, each aligned to 5 human-edited versions.We establish baselines for the task, showing how a relatively small set of human edits can be leveraged to boost the performance of large visual storytelling models.We also discuss the weak correlation between automatic evaluation scores and human ratings, motivating the need for new automatic metrics. Ting-Yao Hsu, Chieh-Yang Huang, Yen-Chia Hsu, Ting-Hao 'Kenneth' Huang |
ACL (1) | 4 |
| 2019 | Smell Pittsburgh: community-empowered mobile smell reporting systemabstractUrban air pollution has been linked to various human health considerations, including cardiopulmonary diseases. Communities who suffer from poor air quality often rely on experts to identify pollution sources due to the lack of accessible tools. Taking this into account, we developed Smell Pittsburgh, a system that enables community members to report odors and track where these odors are frequently concentrated. All smell report data are publicly accessible online. These reports are also sent to the local health department and visualized on a map along with air quality data from monitoring stations. This visualization provides a comprehensive overview of the local pollution landscape. Additionally, with these reports and air quality data, we developed a model to predict upcoming smell events and send push notifications to inform communities. Our evaluation of this system demonstrates that engaging residents in documenting their experiences with pollution odors can help identify local air pollution patterns, and can empower communities to advocate for better air quality. Yen-Chia Hsu, Jennifer L. Cross, Paul Dille, Michael Tasota, Beatrice Dias, Randy Sargent, Ting-Hao 'Kenneth' Huang, Illah R. Nourbakhsh |
IUI | 7 |
| 2019 | Dixit: Interactive Visual Storytelling via Term ManipulationabstractIn this paper, we introduce Dixit, an interactive visual storytelling system that the user interacts with iteratively to compose a short story for a photo sequence. The user initiates the process by uploading a sequence of photos. Dixit first extracts text terms from each photo which describe the objects (e.g., boy, bike) or actions (e.g., sleep) in the photo, and then allows the user to add new terms or remove existing terms. Dixit then generates a short story based on these terms. Behind the scenes, Dixit uses an LSTM-based model trained on image caption data and FrameNet to distill terms from each image, and utilizes a transformer decoder to compose a context-coherent story. Users change images or terms iteratively with Dixit to create the most ideal story. Dixit also allows users to manually edit and rate stories. The proposed procedure opens up possibilities for interpretable and controllable visual storytelling, allowing users to understand the story formation rationale and to intervene in the generation process. Chao-Chun Hsu, Yu-Hua Chen, Zi-Yuan Chen, Hsin-Yu Lin, Ting-Hao 'Kenneth' Huang, Lun-Wei Ku |
WWW | 5 |
| 2018 | Evorus: A Crowd-powered Conversational Assistant Built to Automate Itself Over TimeabstractCrowd-powered conversational assistants have been shown to be more robust than automated systems, but do so at the cost of higher response latency and monetary costs. A promising direction is to combine the two approaches for high quality, low latency, and low cost solutions. In this paper, we introduce Evorus, a crowd-powered conversational assistant built to automate itself over time by (i) allowing new chatbots to be easily integrated to automate more scenarios, (ii) reusing prior crowd answers, and (iii) learning to automatically approve response candidates. Our 5-month-long deployment with 80 participants and 281 conversations shows that Evorus can automate itself without compromising conversation quality. Crowd-AI architectures have long been proposed as a way to reduce cost and latency for crowd-powered systems; Evorus demonstrates how automation can be introduced successfully in a deployed system. Its architecture allows future researchers to make further innovation on the underlying automated components in the context of a deployed open domain dialog system. Ting-Hao 'Kenneth' Huang, Joseph Chee Chang, Jeffrey P. Bigham |
CHI | 1 |
| 2018 | EmotionLines: An Emotion Corpus of Multi-Party Conversations
Chao-Chun Hsu, Sheng-Yeh Chen, Chuan-Chun Kuo, Ting-Hao 'Kenneth' Huang, Lun-Wei Ku |
LREC | 4 |
| 2017 | On How Deaf People Might Use Speech to Control DevicesabstractSmart devices connected to the Internet are proliferating.To reduce costs of devices that havetraditionally been inexpensive(toasters, microwaves, printers, etc), manyof these devices have chosen to use a speech interface rather than a visual one. This transition has been hastened by the increasing capabilities of speech interfaces,exemplifiedbyproducts likeAmazon Echo and Apple'sSiri.A consequence of these products moving to voice control is that people who are deaf and hard of hearing (DHH) may be unable to use them. In this paper, we briefly introduce two technical approaches we are pursuingfor enabling DHH people to provide input to these devices: (i) human computationworkflows for understanding "deaf speech," and (ii) mobile interfaces that can be instructed to speak on the user's behalf. Jeffrey P. Bigham, Raja S. Kushalnagar, Ting-Hao 'Kenneth' Huang, Juan Pablo Flores, Saiph Savage |
ASSETS | 3 |
| 2017 | A 10-Month-Long Deployment Study of On-Demand Recruiting for Low-Latency CrowdsourcingabstractA number of interactive crowd-powered systems have been developed to solve difficult problems out of reach for automated solutions. To work interactively, such systems need access to on-demand labor. To meet this demand, workers can be (i) recruited when needed directly from the crowd marketplace, or (ii) recruited in advance and asked to wait in a retainer pool until they are needed. Most of the evaluations of these systems have been over a short time period, even though we know that marketplaces change and adapt over time. In this paper, we present the results of a 10-month deployment of a crowd-powered system that uses a hybrid approach to fast recruitment of workers that we call Ignition. We describe the Ignition approach and the observed times required to recruit workers from the marketplace and retainer over this long period of time. Our results demonstrate that it is possible to recruit workers with low latency even over long periods, and suggest a number of opportunities for future work for recruitment strategies and modeling that may further improve on-demand recruitment for deployed systems. Ting-Hao 'Kenneth' Huang, Jeffrey P. Bigham |
HCOMP | 1 |
| 2017 | WearMail: On-the-Go Access to Information in Your Email with a Privacy-Preserving Human Computation WorkflowabstractEmail is more than just a communication medium. Email serves as an external memory for people---it contains our reservation numbers, meeting details, phone numbers, and more. Often, people need access to this information while on the go, which is cumbersome from mobile devices with limited I/O bandwidth. In this paper, we introduce WearMail, a conversational interface to retrieve specific information in email. WearMail is mostly automated but is made robust to information extraction tasks via a novel privacy-preserving human computation workflow. In WearMail, crowdworkers never have direct access to emails, but rather (i) generate an email filter to help the system find messages that may contain the desired information, and (ii) generate examples of the requested information that are then used to create custom, low-level information extractors that run automatically within the set of filtered emails. We explore the impact of varying levels of obfuscation on result quality, demonstrating that workers are able to deal with highly-obfuscated information nearly as well as with the original. WearMail introduces general mechanisms that let the crowd search and select private data without having direct access to the data itself. Sai Swaminathan, Raymond Fok, Ting-Hao 'Kenneth' Huang, Irene Lin, Rohan Jadvani, Walter S. Lasecki, Jeffrey P. Bigham |
UIST | 4 |
| 2016 | "Is There Anything Else I Can Help You With?" Challenges in Deploying an On-Demand Crowd-Powered Conversational AgentabstractIntelligent conversational assistants, such as Apple's Siri, Microsoft's Cortana, and Amazon's Echo, have quickly become a part of our digital life. However, these assistants have major limitations, which prevents users from conversing with them as they would with human dialog partners. This limits our ability to observe how users really want to interact with the underlying system. To address this problem, we developed a crowd-powered conversational assistant, Chorus, and deployed it to see how users and workers would interact together when mediated by the system. Chorus sophisticatedly converses with end users over time by recruiting workers on demand, which in turn decide what might be the best response for each user sentence. Up to the first month of our deployment, 59 users have held conversations with Chorus during 320 conversational sessions. In this paper, we present an account of Chorus' deployment, with a focus on four challenges: (i) identifying when conversations are over, (ii) malicious users and workers, (iii) on-demand recruiting, and (iv) settings in which consensus is not enough. Our observations could assist the deployment of crowd-powered conversation systems and crowd-powered systems in general. Ting-Hao 'Kenneth' Huang, Walter S. Lasecki, Amos Azaria, Jeffrey P. Bigham |
HCOMP | 1 |
| 2016 | Visual StorytellingabstractTing-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, Margaret Mitchell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Ting-Hao 'Kenneth' Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross B. Girshick, Xiaodong He 0001, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, Margaret Mitchell |
HLT-NAACL | 1 |
| 2015 | A Survey of Current Datasets for Vision and Language ResearchabstractFrancis Ferraro, Nasrin Mostafazadeh, Ting-Hao Huang, Lucy Vanderwende, Jacob Devlin, Michel Galley, Margaret Mitchell. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Francis Ferraro, Nasrin Mostafazadeh, Ting-Hao 'Kenneth' Huang, Lucy Vanderwende, Jacob Devlin, Michel Galley, Margaret Mitchell |
EMNLP | 3 |
| 2015 | Guardian: A Crowd-Powered Spoken Dialog System for Web APIsabstractNatural language dialog is an important and intuitive way for people to access information and services. However, current dialog systems are limited in scope, brittle to the richness of natural language, and expensive to produce. This paper introduces Guardian, a crowd-powered framework that wraps existing Web APIs into immediately usable spoken dialog systems. Guardian takes as input the Web API and desired task, and the crowd determines the parameters necessary to complete it, how to ask for them, and interprets the responses from the API. The system is structured so that, over time, it can learn to take over for the crowd. This hybrid systems approach will help make dialog systems both more general and more robust going forward. Ting-Hao 'Kenneth' Huang, Walter S. Lasecki, Jeffrey P. Bigham |
HCOMP | 1 |
| 2014 | Combining Non-Expert and Expert Crowd Work to Convert Web APIs to Dialog SystemsabstractThousands of web APIs expose data and services that would be useful to access with natural dialog, from weather and sports to Twitter and movies. The process of adapting each API to a robust dialog system is difficult and time-consuming, as it requires not only programming but also anticipating what is mostly likely to be asked and how it is likely to be asked. We present a crowd-powered system able to generate a natural languageinterface for arbitrary web APIs from scratch without domain-dependent training data or knowledge.Our approach combines two types of crowd workers: non-expert Mechanical Turk workers interpret the functions of the API and elicit information from the user, and expert oDesk workers provide a minimal sufficient scaffolding around the API to allow us to make general queries.We describe our multi-stage process and present results for each stage. Ting-Hao 'Kenneth' Huang, Walter S. Lasecki, Alan L. Ritter, Jeffrey P. Bigham |
HCOMP | 1 |
| 2011 | Predicting Opinion Dependency Relations for Opinion Analysis
Lun-Wei Ku, Ting-Hao 'Kenneth' Huang, Hsin-Hsi Chen |
IJCNLP | 2 |
| 2010 | Predicting Morphological Types of Chinese Bi-Character Words by Machine Learning Approaches
Ting-Hao 'Kenneth' Huang, Lun-Wei Ku, Hsin-Hsi Chen |
LREC | 1 |
| 2010 | Construction of a Chinese Opinion Treebank
Lun-Wei Ku, Ting-Hao 'Kenneth' Huang, Hsin-Hsi Chen |
LREC | 2 |
| 2009 | Using Morphological and Syntactic Structures for Chinese Opinion Analysis
Lun-Wei Ku, Ting-Hao 'Kenneth' Huang, Hsin-Hsi Chen |
EMNLP | 2 |