EDBT 2026 Demo / reviewers in the wild / expert
Ferhan Ture
dblp:02/4537 · also Ferhan Türe
· DBLP profile ↗
29ranked-venue papers
13as first author
7since 2021 · last 2025
0000-0002-5585-157XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 15 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Content Moderation in TV Search: Balancing Policy Compliance, Relevance, and User ExperienceabstractMillions of people rely on search functionality to find and explore content on entertainment platforms. Modern search systems use a combination of candidate generation and ranking approaches, with advanced methods leveraging deep learning and LLM-based techniques to retrieve, generate, and categorize search results. Despite these advancements, search algorithms can still surface inappropriate or irrelevant content due to factors like model unpredictability, metadata errors, or overlooked design flaws. Such issues can misalign with product goals and user expectations, potentially harming user trust and business outcomes. In this work, we introduce an additional monitoring layer using Large Language Models (LLMs) to enhance content moderation. This additional layer flags content if the user did not intend to search for it. This approach serves as a baseline for product quality assurance, with collected feedback used to refine the initial retrieval mechanisms of the search model, ensuring a safer and more reliable user experience. Adeep Hande, Kishorekumar Sundararajan, Sardar Hamidian, Ferhan Ture |
SIGIR | 4 |
| 2024 | Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image GenerationabstractDiffusion models are the state of the art in textto-image generation, but their perceptual variability remains understudied.In this paper, we examine how prompts affect image variability in black-box diffusion-based models.We propose W1KP, a human-calibrated measure of variability in a set of images, bootstrapped from existing image-pair perceptual distances.Current datasets do not cover recent diffusion models, thus we curate three test sets for evaluation.Our best perceptual distance outperforms nine baselines by up to 18 points in accuracy, and our calibration matches graded human judgements 78% of the time.Using W1KP, we study prompt reusability and show that Imagen prompts can be reused for 10-50 random seeds before new images become too similar to already generated images, while Stable Diffusion XL and DALL-E 3 can be reused 50-200 times.Lastly, we analyze 56 linguistic features of real prompts, finding that the prompt's length, CLIP embedding norm, concreteness, and word senses influence variability most.As far as we are aware, we are the first to analyze diffusion variability from a visuolinguistic perspective.Our project page is at http://w1kp.com. Raphael Tang, Xinyu Zhang 0018, Lixinyu Xu, Wenyan Li 0001, Pontus Stenetorp, Jimmy Lin, Ferhan Ture |
EMNLP | 8 |
| 2024 | Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language ModelsabstractRaphael Tang, Crystina Zhang, Xueguang Ma, Jimmy Lin, Ferhan Ture. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Raphael Tang, Xinyu Zhang 0018, Xueguang Ma, Jimmy Lin, Ferhan Ture |
NAACL-HLT | 5 |
| 2024 | "Ask Me Anything": How Comcast Uses LLMs to Assist Agents in Real TimeabstractCustomer service is how companies interface with their customers. It can contribute heavily towards the overall customer satisfaction. However, high-quality service can become expensive, creating an incentive to make it as cost efficient as possible and prompting most companies to utilize AI-powered assistants, or "chat bots". On the other hand, human-to-human interaction is still desired by customers, especially when it comes to complex scenarios such as disputes and sensitive topics like bill payment. Scott Rome, Tianwen Chen, Raphael Tang, Luwei Zhou, Ferhan Ture |
SIGIR | 5 |
| 2023 | What the DAAM: Interpreting Stable Diffusion Using Cross AttentionabstractRaphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang, Gefei Yang, Karun Kumar, Pontus Stenetorp, Jimmy Lin, Ferhan Ture. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Raphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang, Gefei Yang, Karun Kumar, Pontus Stenetorp, Jimmy Lin, Ferhan Ture |
ACL (1) | 9 |
| 2023 | Simulating Humans at Scale to Evaluate Voice Interfaces for TVs: the Round-Trip System at ComcastabstractEvaluating large-scale customer-facing voice interfaces involves a variety of challenges, such as data privacy, fairness or unintended bias, and the cost of human labor. Comcast's Xfinity Voice Remote is one such voice interface aimed at users looking to discover content on their TVs. The artificial intelligence (AI) behind the voice remote currently powers multiple voice interfaces, serving tens of millions of requests every day, from users across the globe.In this talk, we introduce a novel Round-Trip system we have built to evaluate the AI serving these voice interfaces in a semi-automated manner, providing a robust and cheap alternative to traditional quality assurance methods. We discuss five specific challenges we have encountered in Round-Trip and describe our solutions in detail. Breck Baldwin, Lauren Reese, Jan Neumann, Taylor Cassidy, Michael Pereira, G. Craig Murray, Kishorekumar Sundararajan, Yidnekachew Endale, Pramod Kadagattor, Paul Wolfe, Brian Aiken, Tony Braskich, Donte Jiggetts, Adam Sloan, Esther Vaturi, Crystal Pender, Ferhan Ture |
WSDM | 18 |
| 2022 | Learning to Rank Instant Search Results with Multiple Indices: A Case Study in Search Aggregation for EntertainmentabstractAt Xfinity, an instant search system provides a variety of results for a given query from different sources. For each keystroke, new results are rendered on screen to the user, which could contain movies, television series, sporting events, music videos, news clips, person pages, and other result types. Users are also able to use the Xfinity Voice Remote to submit longer queries, some of which are more open-ended. Examples of queries include incomplete words which match multiple results through lexical matching (i.e., "ali"), topical searches ("vampire movies"), and more specific longer searches ("Movies with Adam Sandler"). Since results can be based on lexical matches, semantic matches, item-to-item similarity matches, or a variety of business logic driven sources, a key challenge is how to combine results into a single list. To accomplish this, we propose merging the lists via a Learning to Rank (LTR) neural model which takes into account the search query. This combined list can be personalized via a second LTR neural model with knowledge of the user's search history and metadata of the programs. Because instant search is under-represented in the literature, we present our learnings from research to aid other practitioners. Scott Rome, Sardar Hamidian, Richard Walsh, Kevin Foley, Ferhan Ture |
SIGIR | 5 |
| 2020 | Auto-annotation for Voice-enabled Entertainment SystemsabstractVoice-activated intelligent entertainment systems are prevalent in modern TVs. These systems require accurate automatic speech recognition (ASR) models to transcribe voice queries for further downstream language understanding tasks. Currently, labeling audio data for training is the main bottleneck in deploying accurate machine learning ASR models, especially when these models require up-to-date training data to adapt to the shifting customer needs. We present an auto-annotation system, which provides high quality training data without any hand-labeled audios by detecting speech recognition errors and providing possible fixes. Through our algorithm, the auto-annotated training data reaches an overall word error rate (WER) of 0.002; furthermore, we obtained a reduction of 0.907 in WER after applying the auto-suggested fixes. Wenyan Li 0001, Ferhan Ture |
SIGIR | 2 |
| 2019 | Multi-Perspective Relevance Matching with Hierarchical ConvNets for Social Media SearchabstractDespite substantial interest in applications of neural networks to information retrieval, neural ranking models have mostly been applied to “standard” ad hoc retrieval tasks over web pages and newswire articles. This paper proposes MP-HCNN (Multi-Perspective Hierarchical Convolutional Neural Network), a novel neural ranking model specifically designed for ranking short social media posts. We identify document length, informal language, and heterogeneous relevance signals as features that distinguish documents in our domain, and present a model specifically designed with these characteristics in mind. Our model uses hierarchical convolutional layers to learn latent semantic soft-match relevance signals at the character, word, and phrase levels. A poolingbased similarity measurement layer integrates evidence from multiple types of matches between the query, the social media post, as well as URLs contained in the post. Extensive experiments using Twitter data from the TREC Microblog Tracks 2011–2014 show that our model significantly outperforms prior feature-based as well as existing neural ranking models. To our best knowledge, this paper presents the first substantial work tackling search over social media posts using neural ranking models. Our code and data are publicly available.1 Jinfeng Rao, Wei Yang 0017, Yuhao Zhang 0004, Ferhan Ture, Jimmy Lin |
AAAI | 4 |
| 2019 | Yelling at Your TV: An Analysis of Speech Recognition Errors and Subsequent User Behavior on Entertainment SystemsabstractMillions of consumers issue voice queries through television-based entertainment systems such as the Comcast X1, the Amazon Fire TV, and Roku TV. Automatic speech recognition (ASR) systems are responsible for transcribing these voice queries into text to feed downstream natural language understanding modules. However, ASR is far from perfect, often producing incorrect transcriptions and forcing users to take corrective action. To better understand their impact on sessions, this paper characterizes speech recognition errors as well as subsequent user responses. We provide both quantitative and qualitative analyses, examining the acoustic as well as lexical attributes of the utterances. This work represents, to our knowledge, the first analysis of speech recognition errors from real users on a widely-deployed entertainment system. Raphael Tang, Ferhan Ture, Jimmy Lin |
SIGIR | 2 |
| 2019 | Challenges and Opportunities in Understanding Spoken Queries Directed at Modern Entertainment PlatformsabstractModern in-home entertainment platforms---representing the evolution of the humble television of yesteryear---are packed with features and content: they offer a dizzying array of programs spanning hundreds of channels as well as a catalog of on-demand programs offering tens of thousands of options. Furthermore, the entertainment platform may serve as an in-home hub, providing capabilities ranging from playing music to controlling the home security system. At a high level, our goal is to provide natural speech-based access to these myriad features as an alternative to physical button entry on a remote control. Ferhan Ture, Jinfeng Rao, Raphael Tang, Jimmy Lin |
SIGIR | 1 |
| 2018 | Multi-Task Learning with Neural Networks for Voice Query Understanding on an Entertainment PlatformabstractWe tackle the challenge of understanding voice queries posed against the Comcast Xfinity X1 entertainment platform, where consumers direct speech input at their "voice remotes". Such queries range from specific program navigation (i.e., watch a movie) to requests with vague intents and even queries that have nothing to do with watching TV. We present successively richer neural network architectures to tackle this challenge based on two key insights: The first is that session context can be exploited to disambiguate queries and recover from ASR errors, which we operationalize with hierarchical recurrent neural networks. The second insight is that query understanding requires evidence integration across multiple related tasks, which we identify as program prediction, intent classification, and query tagging. We present a novel multi-task neural architecture that jointly learns to accomplish all three tasks. Our initial model, already deployed in production, serves millions of queries daily with an improved customer experience. The novel multi-task learning model, first described here, is evaluated through carefully-controlled laboratory experiments, which demonstrates further gains in effectiveness and increased system capabilities. Jinfeng Rao, Ferhan Ture, Jimmy Lin |
KDD | 2 |
| 2018 | What Do Viewers Say to Their TVs?: An Analysis of Voice Queries to Entertainment SystemsabstractA recently-introduced product of Comcast, a large cable company in the United States, is a "voice remote" that accepts spoken queries from viewers. We present an analysis of a large query log from this service to answer the question: "What do viewers say to their TVs?" In addition to a descriptive characterization of queries and sessions, we describe two complementary types of analyses to support query understanding. First, we propose a domain-specific intent taxonomy to characterize viewer behavior: as expected, most intents revolve around watching programs---both direct navigation as well as browsing---but there is a non-trivial fraction of non-viewing intents as well. Second, we propose a domain-specific tagging scheme for labeling query tokens, that when combined with intent and program prediction, provides a multi-faceted approach to understand voice queries directed at entertainment systems. Jinfeng Rao, Ferhan Ture, Jimmy Lin |
SIGIR | 2 |
| 2017 | Talking to Your TV: Context-Aware Voice Search with Hierarchical Recurrent Neural NetworksabstractWe tackle the novel problem of navigational voice queries posed against an entertainment system, where viewers interact with a voice-enabled remote controller to specify the TV program to watch. This is a difficult problem for several reasons: such queries are short, even shorter than comparable voice queries in other domains, which offers fewer opportunities for deciphering user intent. Furthermore, ambiguity is exacerbated by underlying speech recognition errors. We address these challenges by integrating word- and character-level query representations and by modeling voice search sessions to capture the contextual dependencies in query sequences. Both are accomplished with a probabilistic framework in which recurrent and feedforward neural network modules are organized in a hierarchical manner. From a raw dataset of 32M voice queries from 2.5M viewers on the Comcast Xfinity X1 entertainment system, we extracted data to train and test our models. We demonstrate the benefits of our hybrid representation and context-aware model, which significantly outperforms competitive baselines that use learning to rank as well as neural networks. Jinfeng Rao, Ferhan Ture, Oliver Jojic, Jimmy Lin |
CIKM | 2 |
| 2017 | No Need to Pay Attention: Simple Recurrent Neural Networks Work!abstractFirst-order factoid question answering assumes that the question can be answered by a single fact in a knowledge base (KB).While this does not seem like a challenging task, many recent attempts that apply either complex linguistic reasoning or deep neural networks achieve 65%-76% accuracy on benchmark sets.Our approach formulates the task as two machine learning problems: detecting the entities in the question, and classifying the question as one of the relation types in the KB.We train a recurrent neural network to solve each problem.On the SimpleQuestions dataset, our approach yields substantial improvements over previously published results -even neural networks based on much more complex architectures.The simplicity of our approach also has practical advantages, such as efficiency and modularity, that are valuable especially in an industry setting.In fact, we present a preliminary analysis of the performance of our model on real queries from Comcast's X1 entertainment platform with millions of users every day. Ferhan Ture, Oliver Jojic |
EMNLP | 1 |
| 2016 | Learning to Translate for Multilingual Question AnsweringabstractIn multilingual question answering, either the question needs to be translated into the document language, or vice versa.In addition to direction, there are multiple methods to perform the translation, four of which we explore in this paper: word-based, 10-best, contextbased, and grammar-based.We build a feature for each combination of translation direction and method, and train a model that learns optimal feature weights.On a large forum dataset consisting of posts in English, Arabic, and Chinese, our novel learn-to-translate approach was more effective than a strong baseline (p < 0.05): translating all text into English, then training a classifier based only on English (original or translated) text. Ferhan Ture, Elizabeth Boschee |
EMNLP | 1 |
| 2016 | Ask Your TV: Real-Time Question Answering with Recurrent Neural NetworksabstractVoice-based interfaces are very popular in today's world, and Comcast customers are no exception. Usage stats show that our new X1 TV platform receives millions of voice queries per day. As a result, expanding the coverage of our voice interface provides a critical competitive advantage, allowing customers to speak freely instead of having to stick to a rigid set of commands. The ultimate objective is to provide a more natural user experience and increase access to our knowledge graph (KG) and entertainment platform. Ferhan Ture, Oliver Jojic |
SIGIR | 1 |
| 2014 | Learning to Translate: A Query-Specific Combination Approach for Cross-Lingual Information RetrievalabstractWhen documents and queries are pre-sented in different languages, the com-mon approach is to translate the query into the document language. While there are a variety of query translation approaches, recent research suggests that combining multiple methods into a single ”structured query ” is the most effective. In this pa-per, we introduce a novel approach for producing a unique combination recipe for each query, as it has also been shown that the optimal combination weights dif-fer substantially across queries and other task specifics. Our query-specific combi-nation method generates statistically sig-nificant improvements over other combi-nation strategies presented in the litera-ture, such as uniform and task-specific weighting. An in-depth empirical anal-ysis presents insights about the effect of data size, domain differences, labeling and tuning on the end performance of our ap-proach. 1 Ferhan Ture, Elizabeth Boschee |
EMNLP | 1 |
| 2014 | Exploiting Representations from Statistical Machine Translation for Cross-Language Information RetrievalabstractThis work explores how internal representations of modern statistical machine translation systems can be exploited for cross-language information retrieval. We tackle two core issues that are central to query translation: how to exploit context to generate more accurate translations and how to preserve ambiguity that may be present in the original query, thereby retaining a diverse set of translation alternatives. These two considerations are often in tension since ambiguity in natural language is typically resolved by exploiting context, but effective retrieval requires striking the right balance. We propose two novel query translation approaches: the grammar-based approach extracts translation probabilities from translation grammars, while the decoder-based approach takes advantage of n -best translation hypotheses. Both are context-sensitive , in contrast to a baseline context-insensitive approach that uses bilingual dictionaries for word-by-word translation. Experimental results show that by “opening up” modern statistical machine translation systems, we can access intermediate representations that yield high retrieval effectiveness. By combining evidence from multiple sources, we demonstrate significant improvements over competitive baselines on standard cross-language information retrieval test collections. In addition to effectiveness, the efficiency of our techniques are explored as well. Ferhan Ture, Jimmy Lin |
ACM Trans. Inf. Syst. | 1 |
| 2013 | Flat vs. hierarchical phrase-based translation models for cross-language information retrievalabstractAlthough context-independent word-based approaches remain popular for cross-language information retrieval, many recent studies have shown that integrating insights from modern statistical machine translation systems can lead to substantial improvements in effectiveness. In this paper, we compare flat and hierarchical phrase-based translation models for query translation. Both approaches yield significantly better results than either a token-based or a one-best translation baseline on standard test collections. The choice of model manifests interesting tradeoffs in terms of effectiveness, efficiency, and model compactness. Ferhan Ture, Jimmy Lin |
SIGIR | 1 |
| 2012 | Combining Statistical Translation Techniques for Cross-Language Information Retrieval
Ferhan Ture, Jimmy Lin, Douglas W. Oard |
COLING | 1 |
| 2012 | Why Not Grab a Free Lunch? Mining Large Corpora for Parallel Sentences to Improve Translation Modeling
Ferhan Ture, Jimmy Lin |
HLT-NAACL | 1 |
| 2012 | Encouraging Consistent Translation Choices
Ferhan Ture, Douglas W. Oard, Philip Resnik |
HLT-NAACL | 1 |
| 2012 | Looking inside the box: context-sensitive translation for cross-language information retrievalabstractCross-language information retrieval (CLIR) today is dominated by techniques that use token-to-token mappings from bilingual dictionaries. Yet, state-of-the-art statistical translation models (e.g., using Synchronous Context-Free Grammars) are far richer, capturing multi-term phrases, term dependencies, and contextual constraints on translation choice. We present a novel CLIR framework that is able to reach inside the translation "black box" and exploit these sources of evidence. Experiments on the TREC-5/6 English-Chinese test collection show this approach to be promising. Ferhan Ture, Jimmy Lin, Douglas W. Oard |
SIGIR | 1 |
| 2011 | No free lunch: brute force vs. locality-sensitive hashing for cross-lingual pairwise similarityabstractThis work explores the problem of cross-lingual pairwise similarity, where the task is to extract similar pairs of documents across two different languages. Solutions to this problem are of general interest for text mining in the multilingual context and have specific applications in statistical machine translation. Our approach takes advantage of cross-language information retrieval (CLIR) techniques to project feature vectors from one language into another, and then uses locality-sensitive hashing (LSH) to extract similar pairs. We show that effective cross-lingual pairwise similarity requires working with similarity thresholds that are much lower than in typical monolingual applications, making the problem quite challenging. We present a parallel, scalable MapReduce implementation of the sort-based sliding window algorithm, which is compared to a brute-force approach on German and English Wikipedia collections. Our central finding can be summarized as "no free lunch": there is no single optimal solution. Instead, we characterize effectiveness-efficiency tradeoffs in the solution space, which can guide the developer to locate a desirable operating point based on application- and resource-specific constraints. Ferhan Ture, Tamer Elsayed, Jimmy Lin |
SIGIR | 1 |
| 2009 | HAPLO-ASP: Haplotype Inference Using Answer Set Programming
Esra Erdem 0001, Ozan Erdem, Ferhan Ture |
LPNMR | 3 |
| 2008 | Efficient Haplotype Inference with Answer Set Programming
Esra Erdem 0001, Ferhan Ture |
AAAI | 2 |
| 2008 | Efficient Haplotype Inference with Answer Set Programming
Ferhan Ture, Esra Erdem 0001 |
AAAI | 1 |
| 2006 | Learning Morphological Disambiguation Rules for Turkish
Deniz Yuret, Ferhan Ture |
HLT-NAACL | 2 |