Asli Celikyilmaz

dblp:15/3724 · DBLP profile ↗
← Back
91ranked-venue papers
26as first author
31since 2021 · last 2025
0000-0002-2854-1445ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 77 · 19 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 10 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Efficient Tool Use with Chain-of-Abstraction Reasoning
abstract
To achieve faithful reasoning that aligns with human expectations, large language models (LLMs) need to ground their reasoning to real-world knowledge (e.g., web facts, math and physical rules). Tools help LLMs access this external knowledge, but there remains challenges for fine-tuning LLM agents (e.g., Toolformer) to invoke tools in multi-step reasoning problems, where inter-connected tool calls require holistic and efficient tool usage planning. In this work, we propose a new method for LLMs to better leverage tools in multi-step reasoning. Our method, Chain-of-Abstraction (CoA), trains LLMs to first decode reasoning chains with abstract placeholders, and then call domain tools to reify each reasoning chain by filling in specific knowledge. This planning with abstract chains enables LLMs to learn more general reasoning strategies, which are robust to shifts of domain knowledge (e.g., math results) relevant to different reasoning questions. It also allows LLMs to perform decoding and calling of external tools in parallel, which avoids the inference delay caused by waiting for tool responses. In mathematical reasoning and Wiki QA domains, we show that our method consistently outperforms previous chain-of-thought and tool-augmented baselines on both in-distribution and out-of-distribution test sets, with an average ~6% absolute QA accuracy improvement. LLM agents trained with our method also show more efficient tool use, with inference speed being on average ~1.4x faster than baseline tool-augmented LLMs.
Silin Gao, Jane Dwivedi-Yu, Xiaoqing Ellen Tan, Ramakanth Pasunuru, Olga Golovneva, Koustuv Sinha, Asli Celikyilmaz, Antoine Bosselut
COLING8
2025 reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs
abstract
Reward models have become a staple in modern NLP, serving as not only a scalable text evaluator, but also an indispensable component in many alignment recipes and inference-time algorithms.However, while recent reward models increase performance on standard benchmarks, this may partly be due to overfitting effects, which would confound an understanding of their true capability.In this work, we scrutinize the robustness of reward models and the extent of such overfitting.We build re-WordBench, which systematically transforms reward model inputs in meaning-or rankingpreserving ways.We show that state-of-theart reward models suffer from substantial performance degradation even with minor input transformations, sometimes dropping to significantly below-random accuracy, suggesting brittleness.To improve reward model robustness, we propose to explicitly train them to assign similar scores to paraphrases, and find that this approach also improves robustness to other distinct kinds of transformations.For example, our robust reward model reduces such degradation by roughly half for the Chat Hard subset in RewardBench.Furthermore, when used in alignment, our robust reward models demonstrate better utility and lead to higher-quality outputs, winning in up to 59% of instances against a standardly trained RM.
Zhaofeng Wu, Michihiro Yasunaga, Andrew Cohen, Asli Celikyilmaz, Marjan Ghazvininejad
EMNLP5
2025 Explore Theory of Mind: program-guided adversarial data generation for theory of mind reasoning
abstract
Do large language models (LLMs) have theory of mind? A plethora of papers and benchmarks have been introduced to evaluate if current models have been able to develop this key ability of social intelligence. However, all rely on limited datasets with simple patterns that can potentially lead to problematic blind spots in evaluation and an overestimation of model capabilities. We introduce ExploreToM, the first framework to allow large-scale generation of diverse and challenging theory of mind data for robust training and evaluation. Our approach leverages an A* search over a custom domain-specific language to produce complex story structures and novel, diverse, yet plausible scenarios to stress test the limits of LLMs. Our evaluation reveals that state-of-the-art LLMs, such as Llama-3.1-70B and GPT-4o, show accuracies as low as 0% and 9% on ExploreToM-generated data, highlighting the need for more robust theory of mind evaluation. As our generations are a conceptual superset of prior work, fine-tuning on our data yields a 27-point accuracy improvement on the classic ToMi benchmark (Le et al., 2019). ExploreToM also enables uncovering underlying skills and factors missing for models to show theory of mind, such as unreliable state tracking or data imbalances, which may contribute to models' poor performance on benchmarks.
Melanie Sclar, Jane Dwivedi-Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov, Yonatan Bisk, Yejin Choi 0001, Asli Celikyilmaz
ICLR7
2025 Collaborative Reasoner: Self-Improving Social Agents with Synthetic Conversations
abstract
With increasingly powerful large language models (LLMs) and LLM-based agents tackling an ever-growing list of tasks, we envision a future where numerous LLM agents work seamlessly with other AI agents and humans to solve complex problems and enhance daily life. To achieve these goals, LLM agents must develop collaborative skills such as effective persuasion, assertion and disagreement, which are often overlooked in the prevalent single-turn training and evaluation of LLMs. In this work, we present Collaborative Reasoner (Coral), a framework to evaluate and improve the collaborative reasoning abilities of language models. In particular, tasks and metrics in Coral necessitate agents to disagree with incorrect solutions, convince their partners of a correct solution, and ultimately agree as a team to commit to a final solution, all through a natural multi-turn conversation. Through comprehensive evaluation on six collaborative reasoning tasks covering domains of coding, math, scientific QA and social reasoning, we show that current models cannot effectively collaborate due to undesirable social behaviors, collapsing even on problems that they can solve singlehandedly. To improve the collaborative reasoning capabilities of LLMs, we propose a self-play method to generate synthetic multi-turn preference data and further train the language models to be better collaborators. Experiments with Llama-3.1, Ministral and Qwen-2.5 models show that our proposed self-improvement approach consistently outperforms finetuned chain-of-thought performance of the same base model, yielding gains up to 16.7% absolute. Human evaluations show that the models exhibit more effective disagreement and produce more natural conversations after training on our synthetic interaction data.
Ansong Ni, Ruta Desai, Xinjie Lei, Jiemin Zhang, Jane Dwivedi-Yu, Ramya Raghavendra, Gargi Ghosh, Shang-Wen Li 0001, Asli Celikyilmaz
NeurIPS11
2024 RLCD: Reinforcement Learning from Contrastive Distillation for LM Alignment
abstract
We propose Reinforcement Learning from Contrastive Distillation (RLCD), a method for aligning language models to follow principles expressed in natural language (e.g., to be more harmless) without using human feedback. RLCD creates preference pairs from two contrasting model outputs, one using a positive prompt designed to encourage following the given principles, and one using a negative prompt designed to encourage violating them. Using two different prompts causes model outputs to be more differentiated on average, resulting in cleaner preference labels in the absence of human annotations. We then use the preference pairs to train a preference model, which is in turn used to improve a base unaligned language model via reinforcement learning. Empirically, RLCD outperforms RLAIF (Bai et al., 2022b) and context distillation (Huang et al., 2022) baselines across three diverse alignment tasks—harmlessness, helpfulness, and story outline generation—and when using both 7B and 30B model scales for simulating preference data
Kevin Yang, Daniel Klein 0001, Asli Celikyilmaz, Nanyun Peng 0001, Yuandong Tian
ICLR3
2024 Open-Domain Text Evaluation via Contrastive Distribution Methods
abstract
Recent advancements in open-domain text generation, driven by the power of large pre-trained language models (LLMs), have demonstrated remarkable performance. However, assessing these models’ generation quality remains a challenge. In this paper, we introduce a novel method for evaluating open-domain text generation called Contrastive Distribution Methods (CDM). Leveraging the connection between increasing model parameters and enhanced LLM performance, CDM creates a mapping from the contrast of two probabilistic distributions – one known to be superior to the other – to quality measures. We investigate CDM for open-domain text generation evaluation under two paradigms: 1) Generative CDM, which harnesses the contrast of two language models’ distributions to generate synthetic examples for training discriminator-based metrics; 2) Discriminative CDM, which directly uses distribution disparities between two language models for evaluation. Our experiments on coherence evaluation for multi-turn dialogue and commonsense evaluation for controllable generation demonstrate CDM’s superior correlate with human judgment than existing automatic evaluation metrics, highlighting the strong performance and generalizability of our approach.
Sidi Lu, Asli Celikyilmaz, Nanyun Peng 0001
ICML3
2024 RESPROMPT: Residual Connection Prompting Advances Multi-Step Reasoning in Large Language Models
abstract
Song Jiang, Zahra Shakeri, Aaron Chan, Maziar Sanjabi, Hamed Firooz, Yinglong Xia, Bugra Akyildiz, Yizhou Sun, Jinchao Li, Qifan Wang, Asli Celikyilmaz. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Song Jiang 0002, Zahra Shakeri, Aaron Chan, Maziar Sanjabi, Hamed Firooz, Yinglong Xia, Bugra Akyildiz, Yizhou Sun, Jinchao Li, Qifan Wang 0001, Asli Celikyilmaz
NAACL-HLT11
2024 Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
abstract
Swarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, Xian Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Swarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, Xian Li 0003
NAACL-HLT3
2024 The ART of LLM Refinement: Ask, Refine, and Trust
abstract
Kumar Shridhar, Koustuv Sinha, Andrew Cohen, Tianlu Wang, Ping Yu, Ramakanth Pasunuru, Mrinmaya Sachan, Jason Weston, Asli Celikyilmaz. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Kumar Shridhar, Koustuv Sinha, Andrew Cohen, Ramakanth Pasunuru, Mrinmaya Sachan, Jason Weston, Asli Celikyilmaz
NAACL-HLT9
2023 Understanding In-Context Learning via Supportive Pretraining Data
abstract
Xiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz, Tianlu Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz
ACL (1)5
2023 ALERT: Adapt Language Models to Reasoning Tasks
abstract
Ping Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin, Gargi Ghosh, Mona Diab, Asli Celikyilmaz. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin 0001, Gargi Ghosh, Mona T. Diab, Asli Celikyilmaz
ACL (1)9
2023 Methods for Measuring, Updating, and Visualizing Factual Beliefs in Language Models
abstract
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Peter Hase, Mona T. Diab, Asli Celikyilmaz, Xian Li 0003, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer 0001
EACL3
2023 Crystal: Introspective Reasoners Reinforced with Self-Feedback
abstract
Extensive work has shown that the performance and interpretability of commonsense reasoning can be improved via knowledge-augmented reasoning methods, where the knowledge that underpins the reasoning process is explicitly verbalized and utilized.However, existing implementations, including "chain-of-thought" and its variants, fall short in capturing the introspective nature of knowledge required in commonsense reasoning, and in accounting for the mutual adaptation between the generation and utilization of knowledge.We propose a novel method to develop an introspective commonsense reasoner, CRYSTAL.To tackle commonsense problems, it first introspects for knowledge statements related to the given question, and subsequently makes an informed prediction that is grounded in the previously introspected knowledge.The knowledge introspection and knowledge-grounded reasoning modes of the model are tuned via reinforcement learning to mutually adapt, where the reward derives from the feedback given by the model itself.Experiments show that CRYSTAL significantly outperforms both the standard supervised finetuning and chain-of-thought distilled methods, and enhances the transparency of the commonsense reasoning process.Our work ultimately validates the feasibility and potential of reinforcing a neural model with self-feedback. 1
Jiacheng Liu 0010, Ramakanth Pasunuru, Hannaneh Hajishirzi, Yejin Choi 0001, Asli Celikyilmaz
EMNLP5
2023 Gender Biases in Automatic Evaluation Metrics for Image Captioning
abstract
Model-based evaluation metrics (e.g., CLIP-Score and GPTScore) have demonstrated decent correlations with human judgments in various language generation tasks.However, their impact on fairness remains largely unexplored.It is widely recognized that pretrained models can inadvertently encode societal biases, thus employing these models for evaluation purposes may inadvertently perpetuate and amplify biases.For example, an evaluation metric may favor the caption "a woman is calculating an account book" over "a man is calculating an account book," even if the image only shows male accountants.In this paper, we conduct a systematic study of gender biases in modelbased automatic evaluation metrics for image captioning tasks.We start by curating a dataset comprising profession, activity, and object concepts associated with stereotypical gender associations.Then, we demonstrate the negative consequences of using these biased metrics, including the inability to differentiate between biased and unbiased generations, as well as the propagation of biases to generation models through reinforcement learning.Finally, we present a simple and effective way to mitigate the metric bias without hurting the correlations with human judgments.Our dataset and framework lay the foundation for understanding the potential harm of model-based evaluation metrics, and facilitate future works to develop more inclusive evaluation metrics. 1
Haoyi Qiu, Zi-Yi Dou, Asli Celikyilmaz, Nanyun Peng 0001
EMNLP4
2023 Look-back Decoding for Open-Ended Text Generation
abstract
Given a prefix (context), open-ended generation aims to decode texts that are coherent, which do not abruptly drift from previous topics, and informative, which do not suffer from undesired repetitions.In this paper, we propose Look-back , an improved decoding algorithm that leverages the Kullback-Leibler divergence to track the distribution distance between current and historical decoding steps.Thus Lookback can automatically predict potential repetitive phrase and topic drift, and remove tokens that may cause the failure modes, restricting the next token probability distribution within a plausible distance to the history.We perform decoding experiments on document continuation and story generation, and demonstrate that Look-back is able to generate more fluent and coherent text, outperforming other strong decoding methods significantly in both automatic and human evaluations 1 .
Chunting Zhou, Asli Celikyilmaz, Xuezhe Ma
EMNLP3
2023 ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
Olga Golovneva, Moya Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, Asli Celikyilmaz
ICLR7
2023 RECKONING: Reasoning through Dynamic Knowledge Encoding
abstract
Recent studies on transformer-based language models show that they can answer questions by reasoning over knowledge provided as part of the context (i.e., in-context reasoning). However, since the available knowledge is often not filtered for a particular question, in-context reasoning can be sensitive to distractor facts, additional content that is irrelevant to a question but that may be relevant for a different question (i.e., not necessarily random noise). In these situations, the model fails to distinguish the necessary knowledge to answer the question, leading to spurious reasoning and degraded performance. This reasoning failure contrasts with the model’s apparent ability to distinguish its contextual knowledge from all the knowledge it has memorized during pre-training. Following this observation, we propose teaching the model to reason more robustly by folding the provided contextual knowledge into the model’s parameters before presenting it with a question. Our method, RECKONING, is a bi-level learning algorithm that teaches language models to reason by updating their parametric knowledge through back-propagation, allowing them to answer questions using the updated parameters. During training, the inner loop rapidly adapts a copy of the model weights to encode contextual knowledge into its parameters. In the outer loop, the model learns to use the updated weights to reproduce and answer reasoning questions about the memorized knowledge. Our experiments on three diverse multi-hop reasoning datasets show that RECKONING’s performance improves over the in-context reasoning baseline (by up to 4.5%). We also find that compared to in-context reasoning, RECKONING generalizes better to longer reasoning chains unseen during training, is more robust to distractors in the context, and is computationally more efficient when multiple questions are asked about the same knowledge.
Zeming Chen 0001, Gail Weiss, Eric Mitchell, Asli Celikyilmaz, Antoine Bosselut
NeurIPS4
2023 How Much Do Language Models Copy From Their Training Data? Evaluating Linguistic Novelty in Text Generation Using RAVEN
abstract
Abstract Current language models can generate high-quality text. Are they simply copying text they have seen before, or have they learned generalizable linguistic abstractions? To tease apart these possibilities, we introduce RAVEN, a suite of analyses for assessing the novelty of generated text, focusing on sequential structure (n-grams) and syntactic structure. We apply these analyses to four neural language models trained on English (an LSTM, a Transformer, Transformer-XL, and GPT-2). For local structure—e.g., individual dependencies—text generated with a standard sampling scheme is substantially less novel than our baseline of human-generated text from each model’s test set. For larger-scale structure—e.g., overall sentence structure—model-generated text is as novel or even more novel than the human-generated baseline, but models still sometimes copy substantially, in some cases duplicating passages over 1,000 words long from the training set. We also perform extensive manual analysis, finding evidence that GPT-2 uses both compositional and analogical generalization mechanisms and showing that GPT-2’s novel text is usually well-formed morphologically and syntactically but has reasonably frequent semantic issues (e.g., being self-contradictory).
Tom McCoy 0001, Paul Smolensky, Tal Linzen, Jianfeng Gao 0001, Asli Celikyilmaz
Trans. Assoc. Comput. Linguistics5
2022 ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech Detection
abstract
Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer, Veselin Stoyanov, Zornitsa Kozareva, Xian Li, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona Diab. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva, Xian Li 0003, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona T. Diab
EMNLP9
2022 Discourse-Aware Soft Prompting for Text Generation
abstract
Current efficient fine-tuning methods (e.g., adapters (Houlsby et al., 2019), prefix-tuning (Li and Liang, 2021), etc.) have optimized conditional text generation via training a small set of extra parameters of the neural language model, while freezing the rest for efficiency.While showing strong performance on some generation tasks, they don't generalize across all generation tasks.We show that soft-prompt based conditional text generation can be improved with simple and efficient methods that simulate modeling the discourse structure of human written text.We investigate two design choices: First, we apply hierarchical blocking on the prefix parameters to simulate a higherlevel discourse structure of human written text.Second, we apply attention sparsity on the prefix parameters at different layers of the network and learn sparse transformations on the softmax-function.We show that structured design of prefix parameters yields more coherent, faithful and relevant generations than the baseline prefix-tuning on all generation tasks.
Marjan Ghazvininejad, Vladimir Karpukhin, Vera Gor, Asli Celikyilmaz
EMNLP4
2022 STRUDEL: Structured Dialogue Summarization for Dialogue Comprehension
abstract
Borui Wang, Chengcheng Feng, Arjun Nair, Madelyn Mao, Jai Desai, Asli Celikyilmaz, Haoran Li, Yashar Mehdad, Dragomir Radev. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Borui Wang, Chengcheng Feng, Arjun Nair, Madelyn Mao, Jai Desai, Asli Celikyilmaz, Haoran Li 0007, Yashar Mehdad, Dragomir R. Radev
EMNLP6
2022 Investigating Crowdsourcing Protocols for Evaluating the Factual Consistency of Summaries
abstract
Xiangru Tang, Alexander Fabbri, Haoran Li, Ziming Mao, Griffin Adams, Borui Wang, Asli Celikyilmaz, Yashar Mehdad, Dragomir Radev. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Xiangru Tang, Alexander R. Fabbri, Haoran Li 0007, Ziming Mao, Griffin Adams, Borui Wang, Asli Celikyilmaz, Yashar Mehdad, Dragomir R. Radev
NAACL-HLT7
2022 CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning
abstract
Xiangru Tang, Arjun Nair, Borui Wang, Bingyao Wang, Jai Desai, Aaron Wade, Haoran Li, Asli Celikyilmaz, Yashar Mehdad, Dragomir Radev. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Xiangru Tang, Arjun Nair, Borui Wang, Bingyao Wang, Jai Desai, Aaron Wade, Haoran Li 0007, Asli Celikyilmaz, Yashar Mehdad, Dragomir R. Radev
NAACL-HLT8
2021 Data Augmentation for Abstractive Query-Focused Multi-Document Summarization
abstract
The progress in Query-focused Multi-Document Summarization (QMDS) has been limited by the lack of sufficient largescale high-quality training datasets. We present two QMDS training datasets, which we construct using two data augmentation methods: (1) transferring the commonly used single-document CNN/Daily Mail summarization dataset to create the QMDSCNN dataset, and (2) mining search-query logs to create the QMDSIR dataset. These two datasets have complementary properties, i.e., QMDSCNN has real summaries but queries are simulated, while QMDSIR has real queries but simulated summaries. To cover both these real summary and query aspects, we build abstractive end-to-end neural network models on the combined datasets that yield new state-of-the-art transfer results on DUC datasets. We also introduce new hierarchical encoders that enable a more efficient encoding of the query together with multiple documents. Empirical results demonstrate that our data augmentation and encoding methods outperform baseline models on automatic metrics, as well as on human evaluations along multiple attributes.
Ramakanth Pasunuru, Asli Celikyilmaz, Michel Galley, Chenyan Xiong, Yizhe Zhang 0002, Mohit Bansal, Jianfeng Gao 0001
AAAI2
2021 EmailSum: Abstractive Email Thread Summarization
abstract
Shiyue Zhang, Asli Celikyilmaz, Jianfeng Gao, Mohit Bansal. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Shiyue Zhang 0001, Asli Celikyilmaz, Jianfeng Gao 0001, Mohit Bansal
ACL/IJCNLP (1)2
2021 AREDSUM: Adaptive Redundancy-Aware Iterative Sentence Ranking for Extractive Document Summarization
abstract
Redundancy-aware extractive summarization systems score the redundancy of the sentences to be included in a summary either jointly with their salience information or separately as an additional sentence scoring step.Previous work shows the efficacy of jointly scoring and selecting sentences with neural sequence generation models.It is, however, not well-understood if the gain is due to better encoding techniques or better redundancy reduction approaches.Similarly, the contribution of salience versus diversity components on the created summary is not studied well.Building on the state-of-the-art encoding methods for summarization, we present two adaptive learning models: AREDSUM-SEQ that jointly considers salience and novelty during sentence selection; and a two-step AREDSUM-CTX that scores salience first, then learns to balance salience and redundancy, enabling the measurement of the impact of each aspect.Empirical results on CNN/DailyMail and NYT50 datasets show that by modeling diversity explicitly in a separate step, AREDSUM-CTX achieves significantly better performance than AREDSUM-SEQ as well as state-of-the-art extractive summarization baselines.
Keping Bi, Rahul Jha, W. Bruce Croft, Asli Celikyilmaz
EACL4
2021 Contrastive Multi-document Question Generation
abstract
Woon Sang Cho, Yizhe Zhang, Sudha Rao, Asli Celikyilmaz, Chenyan Xiong, Jianfeng Gao, Mengdi Wang, Bill Dolan. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Woon Sang Cho, Yizhe Zhang 0002, Sudha Rao, Asli Celikyilmaz, Chenyan Xiong, Jianfeng Gao 0001, Mengdi Wang 0001, William B. Dolan
EACL4
2021 Discourse Understanding and Factual Consistency in Abstractive Summarization
abstract
Saadia Gabriel, Antoine Bosselut, Jeff Da, Ari Holtzman, Jan Buys, Kyle Lo, Asli Celikyilmaz, Yejin Choi. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Saadia Gabriel, Antoine Bosselut, Jeff Da, Ari Holtzman, Jan Buys, Kyle Lo, Asli Celikyilmaz, Yejin Choi 0001
EACL7
2021 Enriching Transformers with Structured Tensor-Product Representations for Abstractive Summarization
abstract
Yichen Jiang, Asli Celikyilmaz, Paul Smolensky, Paul Soulos, Sudha Rao, Hamid Palangi, Roland Fernandez, Caitlin Smith, Mohit Bansal, Jianfeng Gao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Asli Celikyilmaz, Paul Smolensky, Paul Soulos, Sudha Rao, Hamid Palangi, Roland Fernandez, Caitlin Smith, Mohit Bansal, Jianfeng Gao 0001
NAACL-HLT2
2021 QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
abstract
Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, Dragomir Radev. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ming Zhong 0005, Da Yin, Tao Yu 0009, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Awadallah 0001, Asli Celikyilmaz, Yang Liu 0124, Xipeng Qiu, Dragomir R. Radev
NAACL-HLT8
2021 Vision-Language Navigation Policy Learning and Adaptation
abstract
Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms baseline methods by 10 percent on Success Rate weighted by Path Length (SPL) and achieves the state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore and adapt to unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7 to 11.7 percent).
Xin Wang 0061, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao 0001, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, Lei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks
abstract
Many high-level procedural tasks can be decomposed into sequences of instructions that vary in their order and choice of tools. In the cooking domain, the web offers many partially-overlapping text and video recipes (i.e. procedures) that describe how to make the same dish (i.e. high-level task). Aligning instructions for the same dish across different sources can yield descriptive visual explanations that are far richer semantically than conventional textual instructions, providing commonsense insight into how real-world procedures are structured. Learning to align these different instruction sets is challenging because: a) different recipes vary in their order of instructions and use of ingredients; and b) video instructions can be noisy and tend to contain far more information than text instructions. To address these challenges, we first use an unsupervised alignment algorithm that learns pairwise alignments between instructions of different recipes for the same dish. We then use a graph algorithm to derive a joint alignment between multiple text and multiple video recipes for the same dish. We release the Microsoft Research Multimodal Aligned Recipe Corpus containing 150K pairwise alignments between recipes across 4,262 dishes with rich commonsense information.
Angela S. Lin, Sudha Rao, Asli Celikyilmaz, Elnaz Nouri, Chris Brockett, Debadeepta Dey, William B. Dolan
ACL3
2020 Substance over Style: Document-Level Targeted Content Transfer
abstract
Existing language models excel at writing from scratch, but many real-world scenarios require rewriting an existing document to fit a set of constraints.Although sentence-level rewriting has been fairly well-studied, little work has addressed the challenge of rewriting an entire document coherently.In this work, we introduce the task of document-level targeted content transfer and address it in the recipe domain, with a recipe as the document and a dietary restriction (such as vegan or dairy-free) as the targeted constraint.We propose a novel model for this task based on the generative pretrained language model (GPT-2) and train on a large number of roughly-aligned recipe pairs. 1 Both automatic and human evaluations show that our model out-performs existing methods by generating coherent and diverse rewrites that obey the constraint while remaining close to the original document.Finally, we analyze our model's rewrites to assess progress toward the goal of making language generation more attuned to constraints that are substantive rather than stylistic.
Allison Hegel, Sudha Rao, Asli Celikyilmaz, William B. Dolan
EMNLP (1)3
2020 PlotMachines: Outline-Conditioned Generation with Dynamic Plot State Tracking
abstract
We propose the task of outline-conditioned story generation: given an outline as a set of phrases that describe key characters and events to appear in a story, the task is to generate a coherent narrative that is consistent with the provided outline.This task is challenging as the input only provides a rough sketch of the plot, and thus, models need to generate a story by interweaving the key points provided in the outline.This requires the model to keep track of the dynamic states of the latent plot, conditioning on the input outline while generating the full story.We present PLOTMACHINES, a neural narrative model that learns to transform an outline into a coherent story by tracking the dynamic plot states.In addition, we enrich PLOTMACHINES with high-level discourse structure so that the model can learn different writing styles corresponding to different parts of the narrative.Comprehensive experiments over three fiction and non-fiction datasets demonstrate that large-scale language models, such as GPT-2 and GROVER, despite their impressive generation performance, are not sufficient in generating coherent narratives for the given outline, and dynamic plot state tracking is important for composing narratives with tighter, more consistent plots.
Hannah Rashkin, Asli Celikyilmaz, Yejin Choi 0001, Jianfeng Gao 0001
EMNLP (1)2
2020 Working Memory Graphs
abstract
Transformers have increasingly outperformed gated RNNs in obtaining new state-of-the-art results on supervised tasks involving text sequences. Inspired by this trend, we study the question of how Transformer-based models can improve the performance of sequential decision-making agents. We present the Working Memory Graph (WMG), an agent that employs multi-head self-attention to reason over a dynamic set of vectors representing observed and recurrent state. We evaluate WMG in three environments featuring factored observation spaces: a Pathfinding environment that requires complex reasoning over past observations, BabyAI gridworld levels that involve variable goals, and Sokoban which emphasizes future planning. We find that the combination of WMG’s Transformer-based architecture with factored observation spaces leads to significant gains in learning efficiency compared to baseline architectures across all tasks. WMG demonstrates how Transformer-based models can dramatically boost sample efficiency in RL environments for which observations can be factored.
Ricky Loynd, Roland Fernandez, Asli Celikyilmaz, Adith Swaminathan, Matthew J. Hausknecht
ICML3
2019 Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
abstract
We propose a hierarchically structured reinforcement learning approach to address the challenges of planning for generating coherent multi-sentence stories for the visual storytelling task. Within our framework, the task of generating a story given a sequence of images is divided across a two-level hierarchical decoder. The high-level decoder constructs a plan by generating a semantic concept (i.e., topic) for each image in sequence. The low-level decoder generates a sentence for each image using a semantic compositional network, which effectively grounds the sentence generation conditioned on the topic. The two decoders are jointly trained end-to-end using reinforcement learning. We evaluate our model on the visual storytelling (VIST) dataset. Empirical results from both automatic and human evaluations demonstrate that the proposed hierarchically structured reinforced training achieves significantly better performance compared to a strong flat deep reinforcement learning baseline.
Qiuyuan Huang, Zhe Gan, Asli Celikyilmaz, Dapeng Oliver Wu, Xiaodong He 0001
AAAI3
2019 COMET: Commonsense Transformers for Automatic Knowledge Graph Construction
abstract
We present the first comprehensive study on automatic knowledge base construction for two prevalent commonsense knowledge graphs: ATOMIC (Sap et al., 2019) and Con-ceptNet (Speer et al., 2017).Contrary to many conventional KBs that store knowledge with canonical templates, commonsense KBs only store loosely structured open-text descriptions of knowledge.We posit that an important step toward automatic commonsense completion is the development of generative models of commonsense knowledge, and propose COMmonsEnse Transformers (COMET ) that learn to generate rich and diverse commonsense descriptions in natural language.Despite the challenges of commonsense modeling, our investigation reveals promising results when implicit knowledge from deep pre-trained language models is transferred to generate explicit knowledge in commonsense knowledge graphs.Empirical results demonstrate that COMET is able to generate novel knowledge that humans rate as high quality, with up to 77.5% (ATOMIC) and 91.7% (ConceptNet) precision at top 1, which approaches human performance for these resources.Our findings suggest that using generative commonsense models for automatic commonsense KB completion could soon be a plausible alternative to extractive methods.
Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, Yejin Choi 0001
ACL (1)5
2019 Sentence Mover's Similarity: Automatic Evaluation for Multi-Sentence Texts
abstract
For evaluating machine-generated texts, automatic methods hold the promise of avoiding collection of human judgments, which can be expensive and time-consuming.The most common automatic metrics, like BLEU and ROUGE, depend on exact word matching, an inflexible approach for measuring semantic similarity.We introduce methods based on sentence mover's similarity; our automatic metrics evaluate text in a continuous space using word and sentence embeddings.We find that sentence-based metrics correlate with human judgments significantly better than ROUGE, both on machine-generated summaries (average length of 3.4 sentences) and human-authored essays (average length of 7.5).We also show that sentence mover's similarity can be used as a reward when learning a generation model via reinforcement learning; we present both automatic and human evaluations of summaries learned in this way, finding that our approach outperforms ROUGE.
Elizabeth Clark, Asli Celikyilmaz, Noah A. Smith
ACL (1)2
2019 Learning Compressed Sentence Representations for On-Device Text Processing
abstract
Dinghan Shen, Pengyu Cheng, Dhanasekar Sundararaman, Xinyuan Zhang, Qian Yang, Meng Tang, Asli Celikyilmaz, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Dinghan Shen, Pengyu Cheng, Dhanasekar Sundararaman, Xinyuan Zhang 0001, Qian Yang 0003, Asli Celikyilmaz, Lawrence Carin
ACL (1)7
2019 Towards Generating Long and Coherent Text with Multi-Level Latent Variable Models
abstract
Variational autoencoders (VAEs) have received much attention recently as an end-toend architecture for text generation with latent variables.However, previous works typically focus on synthesizing relatively short sentences (up to 20 words), and the posterior collapse issue has been widely identified in text-VAEs.In this paper, we propose to leverage several multi-level structures to learn a VAE model for generating long, and coherent text.In particular, a hierarchy of stochastic layers between the encoder and decoder networks is employed to abstract more informative and semantic-rich latent codes.Besides, we utilize a multi-level decoder structure to capture the coherent long-term structure inherent in long-form texts, by generating intermediate sentence representations as highlevel plan vectors.Extensive experimental results demonstrate that the proposed multi-level VAE model produces more coherent and less repetitive long text compared to baselines as well as can mitigate the posterior-collapse issue.
Dinghan Shen, Asli Celikyilmaz, Yizhe Zhang 0002, Liqun Chen 0001, Xin Wang 0061, Jianfeng Gao 0001, Lawrence Carin
ACL (1)2
2019 Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
abstract
Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms previous methods by 10% on SPL and achieves the new state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7% to 11.7%).
Xin Wang 0061, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao 0001, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, Lei Zhang 0001
CVPR3
2019 Robust Navigation with Language Pretraining and Stochastic Sampling
abstract
Xiujun Li, Chunyuan Li, Qiaolin Xia, Yonatan Bisk, Asli Celikyilmaz, Jianfeng Gao, Noah A. Smith, Yejin Choi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiujun Li, Chunyuan Li, Qiaolin Xia, Yonatan Bisk, Asli Celikyilmaz, Jianfeng Gao 0001, Noah A. Smith, Yejin Choi 0001
EMNLP/IJCNLP (1)5
2018 Discourse-Aware Neural Rewards for Coherent Text Generation
abstract
Antoine Bosselut, Asli Celikyilmaz, Xiaodong He, Jianfeng Gao, Po-Sen Huang, Yejin Choi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Antoine Bosselut, Asli Celikyilmaz, Xiaodong He 0001, Jianfeng Gao 0001, Po-Sen Huang, Yejin Choi 0001
NAACL-HLT2
2018 Deep Communicating Agents for Abstractive Summarization
abstract
Asli Celikyilmaz, Antoine Bosselut, Xiaodong He, Yejin Choi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Asli Celikyilmaz, Antoine Bosselut, Xiaodong He 0001, Yejin Choi 0001
NAACL-HLT1
2017 Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning
abstract
Building a dialogue agent to fulfill complex tasks, such as travel planning, is challenging because the agent has to learn to collectively complete multiple subtasks.For example, the agent needs to reserve a hotel and book a flight so that there leaves enough time for commute between arrival and hotel check-in.This paper addresses this challenge by formulating the task in the mathematical framework of options over Markov Decision Processes (MDPs), and proposing a hierarchical deep reinforcement learning approach to learning a dialogue manager that operates at different temporal scales.The dialogue manager consists of: (1) a top-level dialogue policy that selects among subtasks or options, (2) a low-level dialogue policy that selects primitive actions to complete the subtask given by the top-level policy, and (3) a global state tracker that helps ensure all cross-subtask constraints be satisfied.Experiments on a travel planning task with simulated and real users show that our approach leads to significant improvements over three baselines, two based on handcrafted rules and the other based on flat deep reinforcement learning.
Baolin Peng, Xiujun Li, Lihong Li 0001, Jianfeng Gao 0001, Asli Celikyilmaz, Kam-Fai Wong
EMNLP5
2017 End-to-End Task-Completion Neural Dialogue Systems
abstract
One of the major drawbacks of modularized task-completion dialogue systems is that each module is trained individually, which presents several challenges. For example, downstream modules are affected by earlier modules, and the performance of the entire system is not robust to the accumulated errors. This paper presents a novel end-to-end learning framework for task-completion dialogue systems to tackle such issues. Our neural dialogue system can directly interact with a structured database to assist users in accessing information and accomplishing certain tasks. The reinforcement learning based dialogue manager offers robust capabilities to handle noises caused by other components of the dialogue system. Our experiments in a movie-ticket booking domain show that our end-to-end system not only outperforms modularized dialogue system baselines for both objective and subjective evaluation, but also is robust to noises as demonstrated by several systematic experiments with different error granularity and rates specific to the language understanding module.
Xiujun Li, Yun-Nung Chen, Lihong Li 0001, Jianfeng Gao 0001, Asli Celikyilmaz
IJCNLP(1)5
2017 Spoken language understanding and interaction: machine learning for human-like conversational systems
Milica Gasic, Dilek Hakkani-Tür, Asli Celikyilmaz
Comput. Speech Lang.3
2016 A New Pre-Training Method for Training Deep Learning Models with Application to Spoken Language Understanding
abstract
We propose a simple and efficient approach for pre-training deep learning models with application to slot filling tasks in spoken language understanding. The proposed approach leverages unlabeled data to train the models and is generic enough to work with any deep learning model. In this study, we consider the CNN2CRF architecture that contains Convolutional Neural Network (CNN) with Conditional Random Fields (CRF) as top layer, since it has shown great potential for learning useful representations for supervised sequence learning tasks. The proposed pre-training approach with this architecture learns the feature representations from both labeled and unlabeled data at the CNN layer, covering features that would not be observed in limited labeled data. At the CRF layer, the unlabeled data uses predicted classes of words as latent sequence labels together with labeled sequences. Latent labeled sequences, in principle, has the regularization effect on the labeled sequences, yielding a better generalized model. This allows the network to learn representations that are useful for not only slot tagging using labeled data but also learning dependencies both within and between latent clusters of unseen words. The proposed pre-training method with the CRF2CNN architecture achieves significant gains with respect to the strongest semi-supervised baseline.
Asli Celikyilmaz, Ruhi Sarikaya, Dilek Hakkani-Tür, Nikhil Ramesh, Gökhan Tür
INTERSPEECH1
2016 Multi-Domain Joint Semantic Frame Parsing Using Bi-Directional RNN-LSTM
abstract
Sequence-to-sequence deep learning has recently emerged as a new paradigm in supervised learning for spoken language understanding. However, most of the previous studies explored this framework for building single domain models for each task, such as slot filling or domain classification, comparing deep learning based approaches with conventional ones like conditional random fields. This paper proposes a holistic multi-domain, multi-task (i.e. slot filling, domain and intent detection) modeling approach to estimate complete semantic frames for all user utterances addressed to a conversational system, demonstrating the distinctive power of deep learning methods, namely bi-directional recurrent neural network (RNN) with long-short term memory (LSTM) cells (RNN-LSTM) to handle such complexity. The contributions of the presented work are three-fold: (i) we propose an RNN-LSTM architecture for joint modeling of slot filling, intent determination, and domain classification; (ii) we build a joint multi-domain model enabling multi-task deep learning where the data from each domain reinforces each other; (iii) we investigate alternative architectures for modeling lexical context in spoken language understanding. In addition to the simplicity of the single model framework, experimental results show the power of such an approach on Microsoft Cortana real user data over alternative methods based on single domain/task deep learning.
Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Yun-Nung Chen, Jianfeng Gao 0001, Li Deng 0001, Ye-Yi Wang
INTERSPEECH3
2016 Syntax or semantics? knowledge-guided joint semantic frame parsing
abstract
Spoken language understanding (SLU) is a core component of a spoken dialogue system, which involves intent prediction and slot filling and also called semantic frame parsing. Recently recurrent neural networks (RNN) obtained strong results on SLU due to their superior ability of preserving sequential information over time. Traditionally, the SLU component parses semantic frames for utterances considering their flat structures, as the underlying RNN structure is a linear chain. However, natural language exhibits linguistic properties that provide rich, structured information for better understanding. This paper proposes to apply knowledge-guided structural attention networks (K-SAN), which additionally incorporate non-flat network topologies guided by prior knowledge, to a language understanding task. The model can effectively figure out the salient substructures that are essential to parse the given utterance into its semantic frame with an attention mechanism, where two types of knowledge, syntax and semantics, are utilized. The experiments on the benchmark Air Travel Information System (ATIS) data and the conversational assistant Cortana data show that 1) the proposed K-SAN models with syntax or semantics outperform the state-of-the-art neural network based results, and 2) the improvement for joint semantic frame parsing is more significant, because the structured information provides rich cues for sentence-level understanding, where intent prediction and slot filling can be mutually improved.
Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Jianfeng Gao 0001, Li Deng 0001
SLT4
2016 Intent detection using semantically enriched word embeddings
abstract
State-of-the-art targeted language understanding systems rely on deep learning methods using 1-hot word vectors or off-the-shelf word embeddings. While word embeddings can be enriched with information from semantic lexicons (such as WordNet and PPDB) to improve their semantic representation, most previous research on word-embedding enriching has focused on improving intrinsic word-level tasks such as word analogy and antonym detection. In this work, we enrich word embeddings to force semantically similar or dissimilar words to be closer or farther away in the embedding space to improve the performance of an extrinsic task, namely, intent detection for spoken language understanding. We utilize several semantic lexicons, such as WordNet, PPDB, and Macmillan Dictionary to enrich the word embeddings and later use them as initial representation of words for intent detection. Thus, we enrich embeddings outside the neural network as opposed to learning the embeddings within the network, and, on top of the embeddings, build bidirectional LSTM for intent detection. Our experiments on ATIS and a real log dataset from Microsoft Cortana show that word embeddings enriched with semantic lexicons can improve intent detection.
Joo-Kyung Kim, Gökhan Tür, Asli Celikyilmaz, Ye-Yi Wang
SLT3
2016 An overview of end-to-end language understanding and dialog management for personal digital assistants
abstract
Spoken language understanding and dialog management have emerged as key technologies in interacting with personal digital assistants (PDAs). The coverage, complexity, and the scale of PDAs are much larger than previous conversational understanding systems. As such, new problems arise. In this paper, we provide an overview of the language understanding and dialog management capabilities of PDAs, focusing particularly on Cortana, Microsoft's PDA. We explain the system architecture for language understanding and dialog management for our PDA, indicate how it differs with prior state-of-the-art systems, and describe key components. We also report a set of experiments detailing system performance on a variety of scenarios and tasks. We describe how the quality of user experiences are measured end-to-end and also discuss open issues.
Ruhi Sarikaya, Paul A. Crook, Alex Marin, Minwoo Jeong, Jean-Philippe Robichaud, Asli Celikyilmaz, Young-Bum Kim, Alexandre Rochette, Omar Zia Khan, Daniel Boies, Tasos Anastasakos, Zhaleh Feizollahi, Nikhil Ramesh, Hisami Suzuki, Roman Holenstein, Elizabeth Krawczyk, Vasiliy Radostev
SLT6
2016 An Empirical Investigation of Word Class-Based Features for Natural Language Understanding
abstract
There are many studies that show using class-based features improves the performance of natural language processing (NLP) tasks such as syntactic part-of-speech tagging, dependency parsing, sentiment analysis, and slot filling in natural language understanding (NLU), but not much has been reported on the underlying reasons for the performance improvements. In this paper, we investigate the effects of the word class-based features for the exponential family of models specifically focusing on NLU tasks, and demonstrate that the performance improvements could be attributed to the regularization effect of the class-based features on the underlying model. Our hypothesis is based on empirical observation that shrinking the sum of parameter magnitudes in an exponential model tends to improve performance. We show on several semantic tagging tasks that there is a positive correlation between the model size reduction by the addition of the class-based features and the model performance on a held-out dataset. We also demonstrate that class-based features extracted from different data sources using alternate word clustering methods can individually contribute to the performance gain. Since the proposed features are generated in an unsupervised manner without significant computational overhead, the improvements in performance largely come for free and we show that such features provide gains for a wide range of tasks from semantic classification and slot tagging in NLU to named entity recognition (NER).
Asli Celikyilmaz, Ruhi Sarikaya, Minwoo Jeong, Anoop Deoras
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 A universal model for flexible item selection in conversational dialogs
abstract
Human-computer interaction and statistical natural language understanding has changed with the addition of a visual display screen in modern mobile devices, as visual rendering is used to communicate the dialog system's response. Onscreen item identification and resolution when interpreting the user utterances is one critical problem to achieve the natural and accurate human-machine communication. This problem, also called Flexible Item Selection (FIS), has been posed as a classification task to correctly identify intended on-screen item(s) from user utterances. This paper presents a universal FIS model that can be applied to dialog systems developed in different languages. We design a set of input features for the FIS model that makes it largely language-independent. We demonstrate that a single universal FIS model can be used in place of language specific FIS models with no loss in accuracy. We also show that such a model can generalize well to new unseen languages with minimal loss in accuracy on held out languages including English, French, Spanish, Italian, German, and Chinese. Eliminating the need for building and maintaining a separate FIS model for each new language, the universal FIS model helps scaling an existing dialogue system to new languages faster at a lower development cost.
Asli Celikyilmaz, Zhaleh Feizollahi, Dilek Hakkani-Tür, Ruhi Sarikaya
ASRU1
2015 Natural language understanding for partial queries
abstract
Typical natural language understanding systems are built based on the assumption that they have access to the fully formed complete queries. Today's natural user interfaces, however, enable users to interact with various services and agents (e.g. search engines, personal digital assistants) running on desktop computers and laptops. The system is expected to understand the user's intent while the user is typing the query with the goal of increasing system response rate and ultimately improving the user's productivity. Language understanding models built on fully formed queries perform poorly when tested on partial or incomplete queries. In this study, we consider the problem of domain detection for typed partial natural language queries. We design two sets of features in addition to lexical features to train a multi-valued domain classification model. The first feature set consists of character n-gram features, and the second is the class-based features extracted from clustering of word embeddings. Our experiments show that the two feature sets improve the model's performance by up to 52.8% in comparison to the lexical n-gram baselines.
Asli Celikyilmaz, Ruhi Sarikaya
ASRU2
2015 Investigation of ensemble models for sequence learning
abstract
While ensemble models have proven useful for sequence learning tasks there is relatively fewer work that provide insights into what makes them powerful. In this paper, we investigate the empirical behavior of the ensemble approaches on sequence modeling, specifically for the semantic tagging task. We explore this by comparing the performance of commonly used and easy to implement ensemble methods such as majority voting, linear combination and stacking to a learning based and rather complex ensemble method. Next, we ask the question: when models of different learning methods such as predictive and representation learning (e.g., deep learning) are aggregated, do we get performance gains over the individual baseline models. We explore these questions on a range of datasets on syntactic and semantic tagging tasks such as slot filling. Our findings show that a ranking based ensemble model outperforms all other well-known ensemble models.
Asli Celikyilmaz, Dilek Hakkani-Tür
ICASSP1
2014 Resolving Referring Expressions in Conversational Dialogs for Natural User Interfaces
abstract
Unlike traditional over-the-phone spoken dialog systems (SDSs), modern dialog systems tend to have visual rendering on the device screen as an additional modality to communicate the system's response to the user.Visual display of the system's response not only changes human behavior when interacting with devices, but also creates new research areas in SDSs.Onscreen item identification and resolution in utterances is one critical problem to achieve a natural and accurate humanmachine communication.We pose the problem as a classification task to correctly identify intended on-screen item(s) from user utterances.Using syntactic, semantic as well as context features from the display screen, our model can resolve different types of referring expressions with up to 90% accuracy.In the experiments we also show that the proposed model is robust to domain and screen layout changes.
Asli Celikyilmaz, Zhaleh Feizollahi, Dilek Hakkani-Tür, Ruhi Sarikaya
EMNLP1
2014 A variational Bayesian model for user intent detection
abstract
Intent detectors in state-of-the-art spoken language understanding systems are often trained with a small number of manually annotated examples collected from the application domain. Search query logs provide a large number of unlabeled queries that would be beneficial to improve such supervised classification. Furthermore, the contents of user queries as well as the clicked URLs provide information about user's intent. In this paper, we propose a variational Bayesian approach for modeling latent intents of user queries and clicked URLs when available. We use this model to enhance supervised intent classification of user queries from conversational interactions. Experiments were run with large volumes of search queries and show significant improvements over state-of-the-art systems.
Yangfeng Ji, Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür
ICASSP3
2014 Eye Gaze for Spoken Language Understanding in Multi-modal Conversational Interactions
abstract
When humans converse with each other, they naturally amalgamate information from multiple modalities (i.e., speech, gestures, speech prosody, facial expressions, and eye gaze). This paper focuses on eye gaze and its combination with speech. We develop a model that resolves references to visual (screen) elements in a conversational web browsing system. The system detects eye gaze, recognizes speech, and then interprets the user's browsing intent (e.g., click on a specific element) through a combination of spoken language understanding and eye gaze tracking. We experiment with multi-turn interactions collected in a wizard-of-Oz scenario where users are asked to perform several web-browsing tasks. We compare several gaze features and evaluate their effectiveness when combined with speech-based lexical features. The resulting multi-modal system not only increases user intent (turn) accuracy by 17%, but also resolves the referring expression ambiguity commonly observed in dialog systems with a 10% increase in F-measure.
Dilek Hakkani-Tür, Malcolm Slaney, Asli Celikyilmaz, Larry Heck
ICMI3
2014 Probabilistic enrichment of knowledge graph entities for relation detection in conversational understanding
abstract
Knowledge encoded in semantic graphs such as Freebase has been shown to benefit semantic parsing and interpretation of natural language user utterances. In this paper, we propose new methods to assign weights to semantic graphs that reflect common usage types of the entities and their relations. Such statistical information can improve the disambiguation of entities in natural language utterances. Weights for entity types can be derived from the populated knowledge in the semantic graph, based on the frequency of occurrence of each type. They can also be learned from the usage frequencies in real world natural language text, such as related Wikipedia documents or user queries posed to a search engine. We compare the proposed methods with the unweighted version of the semantic knowledge graph for the relation detection task and show that all weighting methods result in better performance in comparison to using the unweighted version.
Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür, Geoffrey Zweig
INTERSPEECH2
2014 Shrinkage based features for slot tagging with conditional random fields
abstract
In this paper we propose a set of class-based features that are generated in an unsupservised fashion to improve slot tagging with Conditional Random Fields (CRFs). The feature generation is based on the idea behind shrinkage based language models, where shrinking the sum of parameter magnitudes in an exponential model tends to improve performance. We use these features with CRFs and show that they consistently improve the slot tagging performance against baselines on several natural language understanding tasks. Since the proposed features are generated in an unsupervised manner without significant computational overhead, the improvements in performance comes for free and we expect that the same features may result in gains in other tagging tasks.
Ruhi Sarikaya, Asli Celikyilmaz, Anoop Deoras, Minwoo Jeong
INTERSPEECH2
2013 Semi-Supervised Semantic Tagging of Conversational Understanding using Markov Topic Regression
Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür, Ruhi Sarikaya
ACL (1)1
2013 Easy contextual intent prediction and slot detection
abstract
Spoken language understanding (SLU) is one of the main tasks of a dialog system, aiming to identify semantic components in user utterances. In this paper, we investigate the incorporation of context into the SLU tasks of intent prediction and slot detection. Using a corpus that contains session-level information, including the start and end of a session and the sequence of utterances within it, we experiment with the incorporation of information from previous intra-session utterances into the SLU tasks on a given utterance. For slot detection, we find that including features indicating the slots appearing in the previous utterances gives no significant increase in performance. In contrast, for intent prediction we find that a similar approach that incorporates the intent of the previous utterance as a feature yields relative error rate reductions of 6.7% on transcribed data and 8.7% on automatically-recognized data. We also find similar gains when treating intent prediction of utterance sequences as a sequential tagging problem via SVM-HMMs.
Aditya Bhargava, Asli Celikyilmaz, Dilek Hakkani-Tür, Ruhi Sarikaya
ICASSP2
2013 Latent semantic modeling for slot filling in conversational understanding
abstract
In this paper, we propose a new framework for semantic template filling in a conversational understanding (CU) system. Our method decomposes the task into two steps: latent n-gram clustering using a semi-supervised latent Dirichlet allocation (LDA) and sequence tagging for learning semantic structures in a CU system. Latent semantic modeling has been investigated to improve many natural language processing tasks such as syntactic parsing or topic tracking. However, due to several complexity problems caused by issues involving utterance length or dialog corpus size, it has not been analyzed directly for semantic parsing tasks. In this paper, we propose extending the LDA by introducing prior knowledge we obtain from semantic knowledge bases. Then, the topic posteriors obtained from the new LDA model are used as additional constraints to a sequence learning model for the semantic template filling task. The experimental results show significant performance gains on semantic slot filling models when features from latent semantic models are used in a conditional random field (CRF).
Gökhan Tür, Asli Celikyilmaz, Dilek Hakkani-Tür
ICASSP2
2013 IsNL? a discriminative approach to detect natural language like queries for conversational understanding
abstract
While data-driven methods for spoken language understanding (SLU) provide state of the art performances and reduce maintenance and model adaptation costs compared to handcrafted parsers, the collection and annotation of domain-specific natural language utterances for training remains a time-consuming task. A recent line of research has focused on enriching the training data with in-domain utterances by mining search engine query logs to improve the SLU tasks. However genre mismatch is a big obstacle as search queries are typically keywords. In this paper, we present an efficient discriminative binary classification method that filters large collection of online web search queries only to select the natural language like queries. The training data used to build this classifier is mined from search query click logs, represented as a bipartite graph. Starting from queries which contain natural language salient phrases, random graph walk algorithms are employed to mine corresponding keyword queries. Then an active learning method is employed for quickly improving on top of this automatically mined data. The results show that our method is robust to noise in search queries by improving over a baseline model previously used for SLU data collection. We also show the effectiveness of detected natural language like queries in extrinsic evaluations on domain detection and slot filling tasks.
Asli Celikyilmaz, Gökhan Tür, Dilek Hakkani-Tür
INTERSPEECH1
2013 A weakly-supervised approach for discovering new user intents from search query logs
abstract
State-of-the art spoken language understanding models that automatically capture user intents in human to machine dialogs are trained with manually annotated data, which is cumbersome and time-consuming to prepare. For bootstrapping the learning algorithm that detects relations in natural language queries to a conversational system, one can rely on publicly available knowledge graphs, such as Freebase, and mine corresponding data from the web. In this paper, we present an unsupervised approach to discover new user intents using a novel Bayesian hierarchical graphical model. Our model employs search query click logs to enrich the information extracted from bootstrapped models. We use the clicked URLs as implicit supervision and extend the knowledge graph based on the relational information discovered from this model. The posteriors from the graphical model relate the newly discovered intents with the search queries. These queries are then used as additional training examples to complement the bootstrapped relation detection models. The experimental results demonstrate the effectiveness of this approach, showing extended coverage to new intents without impacting the known intents. Index Terms: spoken language understanding, graphical models, search query click logs, intent discovery.
Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür
INTERSPEECH2
2013 Learning to Relate Literal and Sentimental Descriptions of Visual Properties
Mark Yatskar, Svitlana Volkova, Asli Celikyilmaz, William B. Dolan, Luke Zettlemoyer
HLT-NAACL3
2012 A Joint Model for Discovery of Aspects in Utterances
Asli Celikyilmaz, Dilek Hakkani-Tür
ACL (1)1
2012 A Discriminative Classification-Based Approach to Information State Updates for a Multi-Domain Dialog System
abstract
We propose a discriminative classification approach for updating the current information state of a multi-domain dialog system based on user responses. Our method uses a set of lexical and domain independent features to compare the spoken language understanding (SLU) output for the current user turn with the previous information state. We then update the information state accordingly, employing a discriminative machine learning approach. Using a data set collected from our conversational interaction system, we investigate the impact of features based on context dependent and context independent SLU tagging schemas. We show that the proposed approach outperforms two non-trivial baselines, one based on manually crafted rules and the other on classification with lexical features alone. Furthermore, such an approach allows the addition of new domains to the dialog manager in a seamless way.
Dilek Hakkani-Tür, Gökhan Tür, Larry Heck, Ashley Fidler, Asli Celikyilmaz
INTERSPEECH5
2012 Statistical semantic interpretation modeling for spoken language understanding with enriched semantic features
abstract
In natural language human-machine statistical dialog systems, semantic interpretation is a key task typically performed following semantic parsing, and aims to extract canonical meaning representations of semantic components. In the literature, usually manually built rules are used for this task, even for implicitly mentioned non-named semantic components (like genre of a movie or price range of a restaurant). In this study, we present statistical methods for modeling interpretation, which can also benefit from semantic features extracted from large in-domain knowledge sources. We extract features from user utterances using a semantic parser and additional semantic features from textual sources (online reviews, synopses, etc.) using a novel tree clustering approach, to represent unstructured information that correspond to implicit semantic components related to targeted slots in the user's utterances. We evaluate our models on a virtual personal assistance system and demonstrate that our interpreter is effective in that it does not only improve the utterance interpretation in spoken dialog systems (reducing the interpretation error rate by 36% relative compared to a language model baseline), but also unveils hidden semantic units that are otherwise nearly impossible to extract from purely manual lexical features that are typically used in utterance interpretation.
Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür
SLT1
2011 Discovery of Topically Coherent Sentences for Extractive Summarization
Asli Celikyilmaz, Dilek Hakkani-Tür
ACL1
2011 Exploiting distance based similarity in topic models for user intent detection
abstract
One of the main components of spoken language understanding is intent detection, which allows user goals to be identified. A challenging sub-task of intent detection is the identification of intent bearing phrases from a limited amount of training data, while maintaining the ability to generalize well. We present a new probabilistic topic model for jointly identifying semantic intents and common phrases in spoken language utterances. Our model jointly learns a set of intent dependent phrases and captures semantic intent clusters as distributions over these phrases based on a distance dependent sampling method. This sampling method uses proximity of words utterances when assigning words to latent topics. We evaluate our method on labeled utterances and present several examples of discovered semantic units. We demonstrate that our model outperforms standard topic models based on bag-of-words assumption.
Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür, Ashley Fidler, Dustin Hillard
ASRU1
2011 Employing web search query click logs for multi-domain spoken language understanding
abstract
Logs of user queries from a search engine (such as Bing or Google) together with the links clicked provide valuable implicit feedback to improve statistical spoken language understanding (SLU) models. In this work, we propose to enrich the existing classification feature set for domain detection with features computed using the click distribution over a set of clicked URLs from search query click logs (QCLs) of user utterances. Since the form of natural language utterances differs stylistically from that of keyword search queries, to be able to match natural language utterances with related search queries, we perform a syntax-based transformation of the original utterances, after filtering out domain-independent salient phrases. This approach results in significant improvements for domain detection, especially when detecting the domains of web-related user utterances.
Dilek Hakkani-Tür, Gökhan Tür, Larry Heck, Asli Celikyilmaz, Ashley Fidler, Dustin Hillard, Rukmini Iyer, Sarangarajan Parthasarathy
ASRU4
2011 Concept-based classification for multi-document summarization
abstract
Documents often contain inherently many concepts reflecting specific and generic aspects. To automatically generate a short summary text of documents on similar topics, it is imperative that we discover general aspects in documents be cause summaries usually contain general rather than specific concepts. This paper presents a semi-supervised extractive summarization model based upon latent concept classification that can differentiate between the two types of aspects as hidden concepts being mentioned in documents. A classifier is trained on hidden concepts discovered from documents and their corresponding human-generated summaries using a probabilistic Bayesian model: the summary-focused topic model. Experimental results based on ROUGE evaluations indicate that ranking sentences to be included in summary text based on the latent summary concept classification has improvements on the quality of the generated summaries.
Asli Celikyilmaz, Dilek Hakkani-Tür
ICASSP1
2011 Approximate Inference for Domain Detection in Spoken Language Understanding
abstract
This paper presents a semi-latent topic model for semantic domain detection in spoken language understanding systems. We use labeled utterance information to capture latent topics, which directly correspond to semantic domains. Additionally, we introduce an ’informative prior ’ for Bayesian inference that can simultaneously segment utterances of known domains into classes and divide them from out-of-domain utterances. We show that our model generalizes well on the task of classify-ing spoken language utterances and compare its results to those of an unsupervised topic model, which does not use labeled in-formation. Index Terms: spoken language understanding, generative mod-els, gibbs sampling.
Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür
INTERSPEECH1
2011 Learning Weighted Entity Lists from Web Click Logs for Spoken Language Understanding
abstract
Named entity lists provide important features for language understanding, but typical lists can contain many ambiguous or incorrect phrases. We present an approach for automatically learning weighted entity lists by mining user clicks from web search logs. The approach significantly outperforms multiple baseline approaches and the weighted lists improve spoken language understanding tasks such as domain detection and slot filling. Our methods are general and can be easily applied to large quantities of entities, across any number of lists. Index Terms: spoken language understanding, domain detection, slot filling, named entity lists, click logs
Dustin Hillard, Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür
INTERSPEECH2
2011 Towards Unsupervised Spoken Language Understanding: Exploiting Query Click Logs for Slot Filling
abstract
In this paper, we present a novel approach to exploit user queries mined from search engine query click logs to bootstrap or improve slot filling models for spoken language understanding. We propose extending the earlier gazetteer population techniques to mine unannotated training data for semantic parsing. The automatically annotated mined data can then be used to train slot specific parsing models. We show that this method can be used to bootstrap slot filling models and can be combined with any available annotated data to improve performance. Furthermore, this approach may eliminate the need for populating and maintaining in-domain gazetteers, in addition to providing complementary information if they are already available. Index Terms: spoken language understanding, slot filling, data mining, named entity extraction, unsupervised learning
Gökhan Tür, Dilek Hakkani-Tür, Dustin Hillard, Asli Celikyilmaz
INTERSPEECH4
2010 A Hybrid Hierarchical Model for Multi-Document Summarization
Asli Celikyilmaz, Dilek Hakkani-Tür
ACL1
2010 Extractive summarization using a latent variable model
Asli Celikyilmaz, Dilek Hakkani-Tür
INTERSPEECH1
2010 Probabilistic model-based sentiment analysis of twitter messages
abstract
We present a machine learning approach to sentiment classification on twitter messages (tweets). We classify each tweet into two categories: polar and non-polar. Tweets with positive or negative sentiment are considered polar. They are considered non-polar otherwise. Sentiment analysis of tweets can potentially benefit different parties, such as consumers and marketing researchers, for obtaining opinions on different products and services. We present methods for text normalization of the noisy tweets and their classification with respect to the polarity. We experiment with a mixture model approach for generation of sentimental words, which are later used as indicator features of the classification model. Based on a gold standard manually annotated ensemble of tweets, with the new approach, we obtain F-scores that are relatively 10% better than a classification baseline that uses raw word n-gram features.
Asli Celikyilmaz, Dilek Hakkani-Tür, Junlan Feng
SLT1
2010 Decision making with imprecise parameters
Asli Celikyilmaz, I. Burhan Türksen
Int. J. Approx. Reason.1
2009 A Graph-based Semi-Supervised Learning for Question-Answering
Asli Celikyilmaz, Marcus Thint, Zhiheng Huang
ACL/IJCNLP1
2009 Investigation of Question Classifier in Question Answering
Zhiheng Huang, Marcus Thint, Asli Celikyilmaz
EMNLP3
2009 Accurate Semantic Class Classifier for Coreference Resolution
Zhiheng Huang, Guangping Zeng, Weiqun Xu, Asli Celikyilmaz
EMNLP4
2009 Increasing accuracy of two-class pattern recognition with enhanced fuzzy functions
Asli Celikyilmaz, I. Burhan Türksen, Ramazan Aktas, M. Mete Doganay, N. Basak Ceylan
Expert Syst. Appl.1
2008 Uncertainty bounds of Fuzzy C-Regression Method
abstract
The fuzzy C-regression Method (FCRM) based on fuzzy C-means (FCM) clustering algorithm was proposed by Hathaway and Bezdek to solve the switching regression problems, and it was applied to fuzzy models by many to build more powerful fuzzy inference systems. The FCRM methods require initialization parameters which are in need for proper identification, since uncertain information can create imperfect expressions, which may hamper the predictive power of these models. This paper investigates the behavior of the FCRM models under uncertain parameters. The upper and lower bounds of the membership values can be identified based on the limits of level of fuzziness parameter around the certain information points such as local functions and ensemble point values. This is a further step to identify the footprint-of-uncertainty of membership values when FCRM is used. It is shown that the uncertainty of membership values induced by the level of fuzziness parameter can be identified based on first order approximations of the membership value calculation function.
Asli Celikyilmaz, I. Burhan Türksen
FUZZ-IEEE1
2008 Validation criteria for enhanced fuzzy clustering
Asli Celikyilmaz, I. Burhan Türksen
Pattern Recognit. Lett.1
2008 Enhanced Fuzzy System Models With Improved Fuzzy Clustering Algorithm
abstract
Although traditional fuzzy models have proven to have high capacity of approximating the real-world systems, they have some challenges, such as computational complexity, optimization problems, subjectivity, etc. In order to solve some of these problems, this paper proposes a new fuzzy system modeling approach based on improved fuzzy functions to model systems with continuous output variable. The new modeling approach introduces three features: i) an improved fuzzy clustering (IFC) algorithm, ii) a new structure identification algorithm, and iii) a nonparametric inference engine. The IFC algorithm yields simultaneous estimates of parameters of c-regression models, together with fuzzy c-partitioning of the data, to calculate improved membership values with a new membership function. The structure identification of the new approach utilizes IFC, instead of standard fuzzy c-means clustering algorithm, to fuzzy partition the data, and it uses improved membership values as additional input variables along with the original scalar input variables for two different choices of regression methods: least squares estimation or support vector regression, to determine ldquofuzzy functionsrdquo for each cluster. With novel IFC, one could learn the system behavior more accurately compared to other FSM models. The nonparametric inference engine is a new approach, which uses the alike -nearest neighbor method for reasoning. Empirical comparisons indicate that the proposed approach yields comparable or better accuracy than fuzzy or neuro-fuzzy models based on fuzzy rules bases, as well as other soft computing methods.
Asli Celikyilmaz, I. Burhan Türksen
IEEE Trans. Fuzzy Syst.1
2008 Uncertainty Modeling of Improved Fuzzy Functions With Evolutionary Systems
abstract
This paper introduce a type-2 fuzzy function system for uncertainty modeling using evolutionary algorithms (ET2FF). The type-1 fuzzy inference systems (FISs) with fuzzy functions, which do not entail if ... then rule bases, have demonstrated better performance compared to traditional FIS. Nonetheless, the performance of these approaches is usually affected by their uncertain parameters. The proposed method implements a three-phase learning strategy to capture the uncertainties in fuzzy function systems induced by learning parameters, as well as fuzzy function structures. The improved fuzzy clustering initially finds hidden structures, and the genetic learning algorithm optimizes interval type-2 fuzzy sets to capture their optimum uncertainty interval. The proposed ET2FF architecture is evaluated using an extensive suite of real-life applications such as manufacturing process and financial market modeling. The results show that the proposed ET2FF method is comparable--if not superior--to earlier FIS in terms of generalization performance and robustness.
Asli Celikyilmaz, I. Burhan Türksen
IEEE Trans. Syst. Man Cybern. Part B1
2007 Evolutionary fuzzy system models with improved fuzzy functions and its application to industrial process
abstract
This paper presents a new Evolutionary Fuzzy System Modeling strategy alternative to Fuzzy Rule Bases, and does not entail if…then rule base structure. The new approach, which is based on Improved Fuzzy Functions with Genetic algorithms, is proposed to reduce complexity of earlier fuzzy system models and improve modeling accuracy. Structure identification of the new approach is based on a supervised Improved Fuzzy Clustering (IFC) method with a dual optimization algorithm, which yields improved membership values. The merit of the proposed FSM is that uncertain information on natural grouping of data samples, i.e., membership values, is utilized as additional predictors while structuring fuzzy functions. Presented model is applied to desulphurization process of a steel company in Canada. It is shown that proposed approach is superior in comparison to earlier fuzzy, neuro-fuzzy, and non-fuzzy system models in terms of robustness and error reduction.
Asli Celikyilmaz, I. Burhan Türksen
SMC1
2007 Fuzzy functions with support vector machines
Asli Celikyilmaz, I. Burhan Türksen
Inf. Sci.1